Thank you for your thoughtful and challenging response. Your question strikes at the heart of a persistent challenge in health informatics: not just whether we trust data, but which data we trust and why. I appreciate the opportunity to reflect more deeply on this.
The Question of Authority: Which Source Should We Trust?
Sir, you raise an excellent point. In my earlier reflection, I suggested prioritizing nursing triage notes for symptom-onset timelines. However, I agree that this is an oversimplification. The question is not simply "which profession is more reliable?" but rather "which source is most authoritative for this specific data element, and why?"
This distinction is critical because different sources bring different strengths and limitations to the table.
A Framework for Source Authority
To address this, I propose a simple framework for determining source authority based on proximity to the patient's lived experience and the purpose of the data element.
|
Data Element
|
Most Authoritative Source
|
Rationale
|
|
Symptom Onset Timeline
|
Patient Interview / Nursing Intake
|
The patient has direct access to their own experience; nursing intake often captures this in the patient's own words before clinical interpretation
|
|
Physical Examination Findings
|
Clinician's Progress Note
|
The clinician has specialized training to interpret physical signs; their documentation is the clinical record of what was observed
|
|
Diagnosis / Staging
|
Clinician's Summary / MDT Record
|
Diagnosis requires clinical synthesis; the clinician is the most qualified to interpret and document this
|
|
Treatment History
|
Pharmacy Records / Treatment Logs
|
These are objective; less subject to recall bias
|
|
Laboratory Results
|
Laboratory Information System
|
Objective; least subject to human interpretation or transcription error
|
|
Psychosocial Distress
|
Patient Interview / Screening Tool
|
The patient is the best source for their own emotional state; screening tools provide structure
|
What This Means for ARGO/OAUTHC Research
In our context at OAUTHC, where we abstract data for the ARGO cancer database, this framework suggests a hybrid approach:
1. Define Source Rules Before Abstraction
Each variable in the data dictionary should have a clearly defined source of truth. For example:
|
Variable
|
Primary Source
|
Secondary Source (if primary missing)
|
|
Date of symptom onset
|
Nursing triage note
|
Patient interview record
|
|
Date of diagnosis
|
Physician progress note
|
MDT summary
|
|
Staging (TNM)
|
MDT summary
|
Physician note
|
|
ER/PR/HER2 status
|
Pathology report
|
Laboratory information system
|
2. Document Provenance
For each abstracted data point, researchers should document:
- The source document used
- Any discrepancy between sources
- The rationale for the final entry
This allows for data auditability, the ability to trace a data point back to its source and understand the decision-making process behind it.
3. Train Abstractors on Source Hierarchy
Research assistants should be trained not just to enter data, but to critically evaluate which source is most authoritative for each variable. This requires:
- Understanding the clinical workflow
- Knowing where each data element originates
- Recognizing when to flag discrepancies for senior review
A Realistic Example from Our Work
Let me revisit the case of the patient with the left breast mass. In our current workflow, the physician's progress note is often treated as the default source for all data. But if we applied the framework above:
|
Data Element
|
Current Source
|
Proposed Source
|
Rationale
|
|
Symptom onset
|
Physician note: "~1 year"
|
Nursing intake: "18 months"
|
Nursing note was closer to the patient's actual words and less subject to rounding
|
|
Clinical stage
|
Physician note: cT4dN2M0
|
MDT summary (if available)
|
MDT summary reflects consensus; physician note may reflect initial impression
|
|
Treatment plan
|
Physician note
|
Treatment log / pharmacy record
|
Objective; less subject to documentation error
|
This approach would not have eliminated the discrepancy entirely, but it would have systematically reduced the risk of abstraction bias by ensuring that each data element was sourced from the most appropriate document.
The Balance Between Efficiency and Accuracy
I recognize that this framework introduces additional complexity. In a high-volume setting like OAUTHC, researchers face significant time pressure. There is a tension between data completeness (getting all the data) and data validity (getting the right data).
However, I would argue that getting the right data, even if slightly slower, is more valuable than getting incomplete or inaccurate data quickly. The cost of data error—in terms of misdirected research, flawed policy recommendations, and ultimately, patient outcomes—far outweighs the cost of careful abstraction.
Addressing the "Devil's Advocate" Task
To directly address your question: "If the nursing note, physician note, and patient interview give different symptom-onset dates, how should the researcher decide which one makes it into the dataset?"
My recommendation is a three-step process:
- Default to the source defined in the data dictionary. If the data dictionary specifies "nursing triage note" as the primary source for symptom onset, use that.
- Flag discrepancies for review. If the source documents conflict, the researcher should flag this and note the discrepancy. This is not a failure, it is an opportunity for quality improvement.
- Escalate when necessary. For critical variables (e.g., stage, treatment), if there is unresolved disagreement between sources, the researcher should escalate to the clinical lead for resolution.
This process ensures that decisions are made transparently and consistently, rather than left to individual researcher judgment.
Conclusion
Your question has pushed me to think more systematically about source authority and data provenance. I now see that the challenge is not simply "which profession do I trust more?" but rather "which source is most authoritative for this specific data element, and how do I document my decision?"
In our work at ARGO and OAUTHC, where we are building a research database that will inform cancer care and policy, this is not just an academic exercise. It is a practical necessity. The decisions we make about source rules today will determine the quality of the data we use for research tomorrow.
Thank you again for the opportunity to engage with this question.