Discussion - Do you trust the data?

Discussion - Do you trust the data?

Number of replies: 18

Think about the last time you collected, entered, reviewed, or used health data. How confident were you that the data truly represented what happened to the patient?

What could have gone wrong between the patient encounter and the dataset? Share one example from your experience or something you have observed.

In reply to First post

Re: Discussion - Do you trust the data?

by Nzenwa Magnus Nnaemeka -

Reflection on Health Data Integrity: The Gap Between Patient Encounter and Dataset

1. Confidence in the Data: A Recent Experience The last time I reviewed and abstracted health data was while updating the ARGO oncology database with patient treatment outcomes and staging information at OAUTHC. Overall, my confidence in the structured data (e.g., demographic details, histopathology results, and laboratory values like Absolute Neutrophil Count) was high, as these are objectively measured and clearly documented.

However, my confidence in the narrative data, specifically the patient’s history of present illness and symptom timeline, was moderately cautious. I recognized that while the data in the dataset was a faithful representation of what was written in the medical folder, it was not necessarily a perfect representation of what actually happened to the patient. In health informatics, we must constantly distinguish between "data quality" (is the data entered correctly?) and "data validity" (does the data reflect reality?). In this instance, the validity was slightly compromised by the layers of translation and abstraction between the patient’s lived experience and the final registry entry.

2. What Could Go Wrong: The "Information Loss" Pipeline Between the initial patient encounter and the final dataset, data passes through several human and systemic filters: the patient’s recall and language, the clinician’s time-constrained documentation, the Health Information Management (HIM) department’s coding, and finally, the research assistant’s data abstraction. At any of these nodes, information can be lost, generalized, or misinterpreted.

A Realistic Example from the OAUTHC/ARGO Context: During a recent data abstraction exercise for the ARGO breast cancer database, I reviewed the file of a patient presenting with locally advanced disease.

  • The Patient Encounter: During the initial nursing intake and clinician interview, the patient explained (in Yoruba) that she first noticed a "small, hard lump" approximately 18 months ago. She delayed seeking formal hospital care due to financial constraints and initially relied on traditional herbal remedies, a well-documented reality in our demographic that significantly impacts time-to-diagnosis metrics.
  • The Clinical Documentation: The attending clinician, managing a high-volume clinic, documented the history in the patient’s folder as: "Patient presents with left breast mass, duration ~1 year." The clinician likely rounded the number for brevity or based it on when the patient first mentioned seeking formal care, rather than the actual symptom onset.
  • The Dataset Abstraction: When I abstracted this data for the ARGO database, we relied on the physician’s progress note as the source of truth. The dataset was subsequently populated with "Duration of symptoms: 12 months."

The Impact: This seemingly minor discrepancy creates a ripple effect. In oncology research, "time-to-diagnosis" is a critical metric used to evaluate public health interventions and early detection campaigns.

By recording 12 months instead of 18, the dataset artificially shortens the patient’s diagnostic delay. When aggregated across hundreds of patients, this systemic abstraction bias skews the research data, potentially leading to underfunded or misdirected community outreach programs aimed at reducing late-stage presentations.

3. Informatics Takeaway and Mitigation This observation highlights a classic health informatics challenge: unstructured data degradation. To mitigate this in our workflow at ARGO/OAUTHC, we can implement a few targeted strategies:

  • Structured Intake Tools: Introduce standardized, patient-facing intake forms (available in both English and Yoruba) that specifically prompt for symptom onset timelines before the clinician encounter.
  • Source Data Verification (SDV): Train data abstractors to prioritize nursing triage notes or direct patient interviews for historical timelines, rather than relying solely on the physician’s summarized progress note.
  • Clinical Decision Support (CDS): Advocate for EHR dropdowns that require specific ranges (e.g., "0–6 months," "6–12 months," ">12 months") with a mandatory free-text field for context, reducing the temptation to use vague approximations like "~1 year."

 

In reply to Nzenwa Magnus Nnaemeka

Re: Discussion - Do you trust the data?

by Chamberlain Nwanne -

These are great examples of how easily patient health data can be misrepresented, even with the best intentions. Your mitigation strategies are also very relevant. I especially agree that structured intake forms could go a long way toward improving the quality and consistency of documentation.

Just to play devil’s advocate, though, I’m not sure I would automatically prioritize the nursing documentation over the physician’s. Rather than assuming one source is always more reliable, perhaps the bigger question is: Which source should we consider authoritative for a particular data element, and why?

For example, if the nursing note, physician note, and patient interview give different symptom-onset dates, how should the researcher decide which one makes it into the dataset? That’s where provenance and clearly defined source rules become really important.

In reply to Chamberlain Nwanne

Re: Discussion - Do you trust the data?

by Nzenwa Magnus Nnaemeka -

Thank you for your thoughtful and challenging response. Your question strikes at the heart of a persistent challenge in health informatics: not just whether we trust data, but which data we trust and why. I appreciate the opportunity to reflect more deeply on this.

 

The Question of Authority: Which Source Should We Trust?

Sir, you raise an excellent point. In my earlier reflection, I suggested prioritizing nursing triage notes for symptom-onset timelines. However, I agree that this is an oversimplification. The question is not simply "which profession is more reliable?" but rather "which source is most authoritative for this specific data element, and why?"

This distinction is critical because different sources bring different strengths and limitations to the table.

 

A Framework for Source Authority

To address this, I propose a simple framework for determining source authority based on proximity to the patient's lived experience and the purpose of the data element.

Data Element

Most Authoritative Source

Rationale

Symptom Onset Timeline

Patient Interview / Nursing Intake

The patient has direct access to their own experience; nursing intake often captures this in the patient's own words before clinical interpretation

Physical Examination Findings

Clinician's Progress Note

The clinician has specialized training to interpret physical signs; their documentation is the clinical record of what was observed

Diagnosis / Staging

Clinician's Summary / MDT Record

Diagnosis requires clinical synthesis; the clinician is the most qualified to interpret and document this

Treatment History

Pharmacy Records / Treatment Logs

These are objective; less subject to recall bias

Laboratory Results

Laboratory Information System

Objective; least subject to human interpretation or transcription error

Psychosocial Distress

Patient Interview / Screening Tool

The patient is the best source for their own emotional state; screening tools provide structure

 

What This Means for ARGO/OAUTHC Research

In our context at OAUTHC, where we abstract data for the ARGO cancer database, this framework suggests a hybrid approach:

1. Define Source Rules Before Abstraction

Each variable in the data dictionary should have a clearly defined source of truth. For example:

Variable

Primary Source

Secondary Source (if primary missing)

Date of symptom onset

Nursing triage note

Patient interview record

Date of diagnosis

Physician progress note

MDT summary

Staging (TNM)

MDT summary

Physician note

ER/PR/HER2 status

Pathology report

Laboratory information system

 

2. Document Provenance

For each abstracted data point, researchers should document:

  • The source document used
  • Any discrepancy between sources
  • The rationale for the final entry

This allows for data auditability, the ability to trace a data point back to its source and understand the decision-making process behind it.

3. Train Abstractors on Source Hierarchy

Research assistants should be trained not just to enter data, but to critically evaluate which source is most authoritative for each variable. This requires:

  • Understanding the clinical workflow
  • Knowing where each data element originates
  • Recognizing when to flag discrepancies for senior review

A Realistic Example from Our Work

Let me revisit the case of the patient with the left breast mass. In our current workflow, the physician's progress note is often treated as the default source for all data. But if we applied the framework above:

Data Element

Current Source

Proposed Source

Rationale

Symptom onset

Physician note: "~1 year"

Nursing intake: "18 months"

Nursing note was closer to the patient's actual words and less subject to rounding

Clinical stage

Physician note: cT4dN2M0

MDT summary (if available)

MDT summary reflects consensus; physician note may reflect initial impression

Treatment plan

Physician note

Treatment log / pharmacy record

Objective; less subject to documentation error

 

This approach would not have eliminated the discrepancy entirely, but it would have systematically reduced the risk of abstraction bias by ensuring that each data element was sourced from the most appropriate document.

The Balance Between Efficiency and Accuracy

I recognize that this framework introduces additional complexity. In a high-volume setting like OAUTHC, researchers face significant time pressure. There is a tension between data completeness (getting all the data) and data validity (getting the right data).

However, I would argue that getting the right data, even if slightly slower, is more valuable than getting incomplete or inaccurate data quickly. The cost of data error—in terms of misdirected research, flawed policy recommendations, and ultimately, patient outcomes—far outweighs the cost of careful abstraction.

Addressing the "Devil's Advocate" Task

To directly address your question: "If the nursing note, physician note, and patient interview give different symptom-onset dates, how should the researcher decide which one makes it into the dataset?"

My recommendation is a three-step process:

  1. Default to the source defined in the data dictionary. If the data dictionary specifies "nursing triage note" as the primary source for symptom onset, use that.
  2. Flag discrepancies for review. If the source documents conflict, the researcher should flag this and note the discrepancy. This is not a failure, it is an opportunity for quality improvement.
  3. Escalate when necessary. For critical variables (e.g., stage, treatment), if there is unresolved disagreement between sources, the researcher should escalate to the clinical lead for resolution.

This process ensures that decisions are made transparently and consistently, rather than left to individual researcher judgment.

Conclusion

Your question has pushed me to think more systematically about source authority and data provenance. I now see that the challenge is not simply "which profession do I trust more?" but rather "which source is most authoritative for this specific data element, and how do I document my decision?"

In our work at ARGO and OAUTHC, where we are building a research database that will inform cancer care and policy, this is not just an academic exercise. It is a practical necessity. The decisions we make about source rules today will determine the quality of the data we use for research tomorrow.

Thank you again for the opportunity to engage with this question.

 

In reply to Nzenwa Magnus Nnaemeka

Re: Discussion - Do you trust the data?

by Chamberlain Nwanne -
Thanks a lot for your well thought-out response Magnus. I quite agree with many of the positions you landed on although with various degrees of complexity.
In reply to First post

Re: Discussion - Do you trust the data?

by Olatunde Olaniyi Abiodun Oluwafemi -
During my most recent research study on endometrial cancer in a tertiary teaching hospital in Northern Nigeria, I was moderately confident that the final dataset reflected what happened to the patients. However, achieving this required careful review and cross-checking of information from different departments.
One major challenge was data fragmentation between the Histopathology Department and the Obstetrics and Gynaecology Department. Important information was stored in separate records: histopathology records contained details such as tumour type, grade, and specimen diagnosis, while clinical records contained patient presentation, treatment, and follow-up information. Obtaining and linking these data was difficult and time-consuming.
For example, a patient’s histopathology report could confirm endometrial carcinoma, but the corresponding clinical record might have incomplete treatment or follow-up information. Differences in patient identifiers, missing case-file numbers, or incomplete documentation could result in missing data, duplicate entries, or incorrect linkage of pathology and clinical information. These problems could affect the completeness and accuracy of the dataset and potentially influence the study findings.
To reduce these errors, I reviewed the records carefully and cross-checked patient identifiers and key clinical and pathology variables before entering the final data.
In reply to Olatunde Olaniyi Abiodun Oluwafemi

Re: Discussion - Do you trust the data?

by Chamberlain Nwanne -
This is a great example of a very common research informatics challenge where the information you need exists, but not necessarily in the same place or under the same patient identifier (Defragmented data).
Your example also shows why having good pathology data and good clinical data separately does not automatically give us a good research dataset. The linkage between them matters just as much.

I also like that you cross-checked multiple identifiers and variables before accepting the records into your final dataset. But here's something to think about: What happens when the identifiers don't agree and you can't confidently determine whether two records belong to the same patient? At what point should we accept missing information rather than risk making an incorrect match?

Sometimes, recognizing that we cannot reliably link two records may actually be better for data integrity than making our best guess. This is where clearly defined patient-matching rules and documenting how uncertain matches were handled become especially important.
In reply to Chamberlain Nwanne

Re: Discussion - Do you trust the data?

by Olatunde Olaniyi Abiodun Oluwafemi -
Thank you for the insightful feedback and the question.
In my view, as a research informatics professional, simply accepting missing information without investigating the underlying processes can compromise the quality and validity of the research. Instead, we should define and apply clear standards for patient identifiers within a cohort to ensure data quality, trustworthiness, accuracy, and fitness for purpose.

I strongly agree that, to protect data integrity, we must recognize when we cannot reliably link two records especially in my setting, where fragmented systems and inconsistent identifiers are common. In such cases, it is better to classify the linkage as uncertain or missing than to force a match that may be incorrect. This approach should be supported by clearly documented patient-matching rules and a transparent description of how uncertain matches were handled in the analysis.
In reply to First post

Re: Discussion - Do you trust the data?

by Aromolaran Precious Adebisola -
I estimated my confidence in UNIMEDTHC health data at just 60% to 70% during a recent review. The physical lag between clinical care and data abstraction creates systemic vulnerabilities where ground truth degrades into incomplete records.

Vulnerabilities Between Encounter and Dataset
• Recall Decay: Delayed charting causes clinicians to omit secondary diagnoses, vital sign trends, or subtle drug reactions.
• Transcription Errors: Illegible handwriting leads to misinterpreted dosages, wrong ICD/ICD-10 coding, or transposed patient identification numbers.
• Physical Loss: Loose lab slips, triage notes, and prescription carbons routinely detach from paper folders during inter-departmental transport.
• Aggregation Gaps: Health Information Management (HIM) staff manually tallies monthly registry figures from incomplete physical files, compounding individual errors into flawed health statistics.

Case Study: UNIMEDTHC Outpatient Clinic Flow
While reviewing health data at a crowded Tuesday morning shift at the UNIMED Teaching Hospital Complex (UNIMEDTHC) outpatient department:
1. The Encounter: The doctor evaluates a hypertensive diabetic patient presenting with severe fatigue. He adjusts their antihypertensive regimen, order a Fasting Blood Sugar (FBS) test, and write the details across a worn paper case note.
2. The Breakpoint: Overwhelmed by a waiting room of 40 more patients, the doctor abbreviates his notes. The lab result returns hours later on a small paper slip; a nurse slips it into the folder without gluing it down.
3. The Data Entry Gap: Days later, a records officer pulls the physical file to transcribe monthly indicators into the central registry. The loose lab slip has fallen out in transit. The doctor’s hurried handwriting of "Amlodipine 10mg" is misread as "5mg", and the secondary diagnosis of early-stage diabetic nephropathy is skipped entirely because it was squeezed into the margin.

By the time this encounter hits the hospital's aggregated monthly dataset, the patient is recorded merely as a routine, well-controlled hypertensive follow-up. The actual clinical risk, medication dosage, and lab marker vanished entirely between the physician’s desk and the registry page.
In reply to Aromolaran Precious Adebisola

Re: Discussion - Do you trust the data?

by Chamberlain Nwanne -
Hi Precious, this is a great example of how information can gradually lose meaning as it moves from the patient encounter to the final dataset. I especially like how you identified several different breakpoints in the information flow rather than treating poor data quality as simply a records or data-entry problem. The missing lab slip, abbreviated documentation, and medication transcription error all occurred at different points in the information pathway, but ultimately affected the same dataset. And the patient ends up bearing the brunt of the consequences!

Your example also makes me wonder how digitization, particularly an EHR, might fit into a more comprehensive solution. An EHR could certainly address problems like lost paper results, illegible handwriting, and some transcription errors, but would it solve the underlying problem of rushed or incomplete documentation? Probably not by itself.

Implementing provider education and using documentation standards are important for minimizing ambiguous abbreviations, ineligible handwriting and using a consistent, structured note format.
But just something to think about - if you were asked to redesign this workflow as an informaticist, which problems would you address with technology, which would you address through workflow or provider education, and which would require a combination of both?
In reply to First post

Re: Discussion - Do you trust the data?

by Olubola Titilope Adegbosin -
I recently used data from an oncology clinic for a study of survival among patients treated for cervical cancer. The dataset was extensive, containing about 50 columns covering sociodemographic characteristics, clinical presentation, investigation findings, treatment characteristics, and outcomes. However, I was not fully confident that the data accurately represented what had happened to every patient. The integrity of some important variables, particularly the total radiation dose received and the modality of radiation treatment (EBRT, brachytherapy, or both), was uncertain. Many patients had missing survival information, particularly those who received only EBRT; in some cases, the recorded disease stage did not correspond with the treatment modality documented; while in other cases, the mismatch was between total radiation dose received and the treatment modality.

Part of the problem was fragmentation of care. EBRT and brachytherapy were sometimes administered at different facilities: only a few hospitals have facilities for both forms of radiotherapy. In some cases, the primary oncologist and most of the patient's clinical information were in one facility, while the actual treatment was delivered elsewhere. Even within the same radiotherapy facility, different sub-units recorded data differently, resulting in inconsistent and missing data. Many patients were subsequently lost to follow-up. Generally, it was obvious that the data was not collected with the intention that it would be required for research and audits later.
To address these challenges, some patients had to be excluded from analysis, survival analysis could not be compared across some patient groups, and the whole dataset needed extensive cross-checking of records from different points in the care pathway, not once but several times during the study. We also had to call patients to obtain missing information, although some were deceased or could no longer be reached on the telephone numbers available.

This experience taught me that it is easier to put structures in place to ensure that data generated during routine care are complete, useful, and meet predefined standards for their intended use, than to try to perfect messy data after it has already been generated.
In reply to Olubola Titilope Adegbosin

Re: Discussion - Do you trust the data?

by Chamberlain Nwanne -
Olubola, this is a great example of the difference between having a lot of data and having research-ready data. Fifty variables may sound like a rich dataset, but if critical variables such as radiation modality, total dose, disease stage, and survival status cannot be reliably interpreted, the research questions you can confidently answer become much more limited.

I also like your observation that much of this data was originally collected to support patient care, not necessarily future research. Fragmentation across facilities and even across units within the same facility makes provenance and continuity particularly challenging. The amount of cross-checking and even calling patients that was required also shows how expensive it can be to try to reconstruct information after the fact.

Your final point is especially important: it is much easier to design for good data quality upfront than to repair a messy dataset later. But here's something to think about: If patients will continue to receive EBRT and brachytherapy at different facilities, what minimum information should follow the patient so that the complete treatment course can eventually be reconstructed? And who should be responsible for making sure that information gets captured?
This may be a case where improving the dataset starts not with the research database, but with redesigning the information pathway across the patient's entire care journey.
In reply to First post

Discussion - Do you trust the data?

by Dr Aminu Bello Liman -

I will make reference to a recent study on developing machine learning model for predicting survival among breast cancer patients. I participated in the data collection process which involved retrospective review of paper-based records over a five year period. The data elements collected include: sociodemography, clinical, pathology treatment received and outcomes.
I will say that I have moderate confidence level that the data truly represent what happend to the patients.
The sociodemographic data was straightforward as the second page of the case files carried the necessary information. Retrieving other aspects of the datasets was challenging as most of the papers in the case files were not serially arranged and some variations exist in the information. Much time was required to scrutinize the case files to obtain as close to perfect information as possible. Considering the fact that patients were referred to other centres for radiotherapy, we had limited information about details of the treatment received.
I will say that I trust the data collected as it reflected the management pathway for breast cancer cases in our centre. However, gaps were noted in the completeness of the cases as a whole and some fields in the datasets. Definitely, we were unable to retrieve some case files despite checking other relevant departments. This may affect the general outlook of the findings. Also, for some of the datasets, some information could not be retrieved despite multiple reviews of the available case files.
This course emphasized the fact that the process of data collection and entry must be scrutinized to ensure that accurate and reliable data gets to the stage of analysis.
I am hopeful that with the ongoing implementation of EMR in our centre and practical application of clinical documentation improvement, subsequent data collection process will be easier and more reliable.

In reply to Dr Aminu Bello Liman

Re: Discussion - Do you trust the data?

by Chamberlain Nwanne -
Hi - One thing that stands out from your experience is the amount of effort required to reconstruct the patient journey from paper records. When documents are not arranged chronologically (and sometimes, even if they are) and information varies across different parts of the record, determining what actually happened becomes an interpretation exercise rather than simple data abstraction. There is now a more complex and resource-intensive process required.

Your point about patients receiving radiotherapy elsewhere is also important. Even if your abstraction accurately reflects what is documented at your center, it may not fully represent the patient's treatment journey because the other facilities may not document as well as your facility does. That distinction becomes especially important when developing a machine-learning model, because the model will learn from whatever information is available in the dataset including its missingness and limitations.

Just out of curiosity, what's the EMR/EHR that you are implementing? The new system would certainly help with accessibility, organization, and legibility, but digitization alone may not guarantee better research data. Clinical documentation improvement, standardized data definitions, and consistent capture of treatment received outside the institution will still be important.

One thing I'd be very mindful of is how the survival predictions produced by your model are affected if patients whose records could not be retrieved, or whose external radiotherapy information was missing, were systematically different from those with complete records.
That takes the issue from simply “missing data” to whether the dataset and ultimately the model adequately represents the patient population.
In reply to First post

Re: Discussion - Do you trust the data?

by Janiel Johnson -
One example comes directly from my work with veterans. Our record data system, MCR, had a veteran marked as having Medicaid, but when I sat down with him for his healthcare consultation, he said he may have applied for it while he was incarcerated. He was unsure whether the application had gone through, and he did not have a card to confirm either way. The record said, "has Medicaid," but the actual, verified status was closer to "unknown, possibly never completed." This gap matters a great deal in practice. If that veteran has a chronic condition or has run out of medication, I do not have the luxury of treating "Medicaid" as a confirmed fact and moving on. I need to quickly find a clinic or program that will see him as a walk-in or get him a sooner appointment, because I cannot assume insurance coverage will hold up if he shows up somewhere expecting it to work.

Thinking about this through the lens of provenance and lineage, the "Medicaid" entry in MCR likely came from an intake form filled out at some earlier point, possibly self-reported by the veteran or copied from a previous case file, with no verification step and no timestamp explaining when that status was last confirmed. Nothing in the record captured that the application happened under unusual circumstances (while incarcerated), that it may never have been completed, or that the veteran himself was uncertain about it. A status field like this treats "applied" and "enrolled and active" as the same thing, but for someone trying to get a veteran seen quickly, that distinction is the entire difference between a smooth referral and a same-day scramble.

This is a good example of why missingness and uncertainty deserve their own space in a record. If MCR had a way to flag "unconfirmed" or "self-reported, not verified" alongside the Medicaid status, I would know at a glance that I need to verify before relying on it, instead of finding out the hard way when a veteran is sitting in front of me with no medication and nowhere to go.
In reply to Janiel Johnson

Re: Discussion - Do you trust the data?

by Chamberlain Nwanne -
Janiel, the distinction you make between “has Medicaid” and “Medicaid status unknown or unverified” is important. That value in the system looks definitive, but the information behind it is much less certain. Without knowing when it was entered, where it came from, or whether it was ever verified, someone using the record could easily make a decision based on information that is no longer, or perhaps never was, accurate.

You also raise an important point about representing uncertainty. Ideally, the system would distinguish between applied, self-reported, verified active, inactive, and unknown coverage. But in your current role, you may not have the ability to structurally change MCR intake forms and add those options yourself. So for me, the practical informatics question becomes: What can you change within the workflow when you cannot change the technology?

For example, could there be a consistent verification step before insurance information is used for referrals, or a standard way to document that verification outside the existing status field?
Sometimes improving data quality does not begin with redesigning the system itself; it begins with redesigning the workflow (how people create, verify, and use the information) within the constraints of that system.
Great observation.
In reply to Chamberlain Nwanne

Re: Discussion - Do you trust the data?

by Janiel Johnson -
Thank you, Dr. Chamberlain. This perspective is really helpful. You're right that I've been focused on system design (what fields MCR should have) rather than workflow (what I can control right now with the fields MCR already has), and thinking about it that way puts my existing practice in a clearer light: I do call to verify coverage when I can, but as the only healthcare navigator for my organization at this location. I am seeing everyone on my schedule plus walk-ins and that verification step does not happen consistently for every veteran before a referral, simply because the time is not always there. Given that constraint, the more realistic fix may not be verifying every case up front, but adopting a standardized note format ("Medicaid: unverified as of [date], veteran unsure of application status") so that whenever I have not had time to verify, anyone reading the record, including a colleague covering my caseload, sees that caveat immediately instead of assuming the status field tells the whole story. What I'm taking from this is that data quality is not only a property of the record but also a property of the habits around how that record gets created and used, and that even an unverified field becomes far less risky if a reliable workflow step, even an imperfect or partial one given my caseload, consistently flags the uncertainty before someone acts on it. If the structured field is unavailable and my time to verify is limited, the next best thing is building a workflow that reliably generates that same context somewhere accessible, even if it lives in a note rather than a dropdown.
In reply to First post

Discussion - Do you trust the data?

by Ishola Ridwan Femi -

The last time I reviewed health data, I was fairly confident in the information, but I also recognized that errors could occur before the data reached the dataset. For example, I have observed cases where a patient was given a new hospital number even though the patient already had an existing record. This can lead to duplicate records, with previous diagnoses, medications, and investigations remaining under the old number.

This experience showed me that data quality problems can begin at the point of registration and documentation, long before the data are analyzed. Therefore, proper patient identification and validation are important for ensuring that health data accurately represent what happened to the patient.

In reply to Ishola Ridwan Femi

Re: Discussion - Do you trust the data?

by Chamberlain Nwanne -
Hi Ishola, you’ve identified a problem that can have consequences well beyond the registration desk. It would have been great if you provided us with a bit more context: Where did you encounter this issue and how?
But regardless, although the issue may seem simple, it is hugely important because if the same patient receives two hospital numbers, both records may contain accurate information, making each individual record appear internally valid but neither record tells the complete story of that patient. That can negatively affect care provided and continuity of care as well as the quality of any dataset created from those records.

Your point that the problem begins before analysis is also a good one. However, I would take it one step further and say that proper patient identification requires more than simply reminding registration staff to check carefully. We should have a clear and consistent way to determine whether someone already has a record, especially when two patients may have exactly the same names, or names may be spelled differently or other demographic information does not match exactly.