Discussion - Welcome to Research Informatics & Data Management

Number of replies: 2

Welcome to Research Informatics & Data Management, a one-week intensive course built around an important question:

How do we know that the data we use to generate evidence can actually be trusted?

This is not a traditional research-methods or statistics course. Instead, we will focus on what happens before the analysis—how a research question becomes a set of data requirements, where those data come from, what happens to them along the way, and how data quality, meaning, governance, and bias can ultimately affect our findings.

One idea will guide us throughout the course:

A research finding is only as trustworthy as the information pathway that produced it.

We will work through five connected modules. We start by translating research questions into clearly defined, computable data requirements. From there, we explore healthcare data sources, provenance and lineage, data models and dictionaries, metadata, data quality, missingness, and fitness for purpose. We will also look at governance, privacy, ethical use of data, FAIR principles, and an issue that is sometimes overlooked: how bias can enter our data long before statistical analysis begins.

We will then bring these concepts into LMIC and resource-constrained settings, where research may depend on a mix of paper and digital records, fragmented systems, inconsistent patient identifiers, and limited interoperability. The goal is not to lower our standards, but to think practically about how we can create reliable and sustainable research-data processes within these realities.

You will put these ideas into practice through the “Can We Trust the Dataset?” exercise and conclude the course by developing an applied Research Informatics Plan.

As you work through the course, I encourage you to move beyond simply asking, “Do we have the data?”

Instead, ask:

Where did the data come from? What do they really mean? What happened to them along the way? Who might be missing? And are they actually fit for the question we are trying to answer?

That way of thinking is at the heart of Research Informatics.

In reply to First post

Re: Discussion - Welcome to Research Informatics & Data Management

by Aromolaran Precious Adebisola -
This course description hits on a critical reality that is too often overlooked in health research: data quality and reliability are forged long before we open a statistical software package.

In resource-constrained and low- to middle-income country (LMIC) settings, this focus on the "information pathway" is particularly vital. When you are navigating hybrid environments—where paper charts, fragmented digital tools, parallel reporting registries, and inconsistent patient identification systems overlap—the gap between what a variable is intended to measure and what is actually recorded can widen rapidly.

Tracing data lineage, evaluating fitness for purpose, and uncovering structural bias early on isn't just about protecting internal validity; it's an issue of health equity. If missingness disproportionately affects specific patient populations or if local documentation workflows distort clinical meaning, our downstream evidence will be compromised no matter how advanced the analytical modeling is.

I’m looking forward to diving into the "Can We Trust the Dataset?" exercise and building out the Research Informatics Plan. Thinking through practical governance and data-provenance strategies that maintain rigorous standards while remaining feasible within real-world clinical realities is exactly where the field needs to focus.
In reply to First post

Discussion - Welcome to Research Informatics & Data Management

by Ishola Ridwan Femi -

I find the idea that “a research finding is only as trustworthy as the information pathway that produced it” very relevant to Health Information Management. In malaria surveillance, having data is not enough; we need to know where the data came from, how it was recorded, validated, and reported. This is important at OAUTHC, where paper and electronic systems may coexist and issues such as missing information, duplicate records, inconsistent identifiers, and delayed reporting can affect data quality.

This course has helped me understand that reliable research depends on reliable data. These concepts will guide my applied informatics project on improving the quality of malaria surveillance data at OAUTHC.