Hospitals and trials are drowning in records. Too much of the information that matters is captured after the moment it could have been useful.
4 min read

A patient arrives at the emergency department in decompensated heart failure. By the time she's admitted, the episode may already be more than a week old.
In one study that interviewed patients shortly after admission and worked backward through their symptoms, worsening oedema had begun a mean of 12.4 days before admission, weight gain 11.3 days before, and breathlessness 8.4 days before. A separate cohort found a median of seven days between the onset of fatigue and oedema and arrival at the emergency department.
In other words, some of the most decision-relevant days of the episode happened before anyone in the health system knew it was underway.
Not because the information was impossible to interpret.
Because nobody had collected it.
The chart is a record of encounters, not of illness
Much of what enters a medical record is generated because a patient interacts with the healthcare system: a visit, an admission, a lab draw, an imaging study, a prescription, a billable event.
That's the sampling frame.
The result is a record that can be extraordinarily dense when the patient is being observed and remarkably thin everywhere else. Those gaps aren't necessarily clinically unimportant. They include the days and weeks when someone is at home, between appointments, potentially changing in ways their care team cannot see.
More encounter data does not solve that problem. It gives us a more detailed description of the moments we were already watching.
One recent analysis framed the imbalance starkly: clinical contact may account for roughly one hour of a person's year, compared with approximately 8,760 hours lived outside that setting.
We instrument the hour and infer the rest.
We spent fifteen years improving the plumbing
Since the HITECH era, healthcare has made enormous investments in digitizing and moving information: electronic health records, interoperability, health information exchanges, FHIR APIs, warehouses and data lakes.
That work was necessary.
But moving existing information is not the same thing as collecting missing information.
Pipes don't make water.
The current AI wave inherits the same constraint. A significant portion of clinically useful information lives in unstructured sources such as notes, imaging reports and discharge summaries. Increasingly sophisticated models can extract, organize and interpret that information.
That work is valuable. But no model, regardless of how capable, can reconstruct an observation that was never made.
You cannot model your way to a measurement nobody took.
Trials have the same shape
Clinical trials face a similar problem.
Source data can be dense around scheduled visits and much thinner between them. Symptoms, fatigue, sleep, tolerability, mobility and quality of life may change continuously, but they only become data when somebody asks the patient about them.
If the between-visit interval is four weeks and the thing you're measuring changes meaningfully in three days, the schedule isn't really sampling the phenomenon.
It's sampling the calendar.
That distinction becomes even more important in long-term studies, where maintaining frequent human follow-up for hundreds or thousands of participants can become operationally difficult.
The question isn't simply whether a trial can collect another patient-reported outcome.
It's whether it can collect the right information often enough, consistently enough, and with little enough burden to make that information useful.
What “the right information at the right time” actually requires
Cadence matched to the condition, not the calendar. If meaningful deterioration can develop over a week, a monthly touchpoint may miss the window in which that information is most useful. Collection frequency should reflect how quickly the thing being measured can change.
Structured at the point of capture. Information collected in a structured form when the patient provides it can be acted on immediately and analyzed longitudinally. That's fundamentally different from trying to reconstruct the same signal from free text months or years later.
As little access friction as possible. Every additional requirement — an app download, password, portal login, charged device or new interface — creates another opportunity for participation to fail. Collection systems should work for the broadest possible patient population, including people who may be comfortable doing little more than answering a phone.
A route to someone who can act. This is the difference between collection and accumulation. Information that enters a record but isn't surfaced until the next scheduled review may be useful historically, but it cannot change what happens today.
The point
Healthcare's instinct has increasingly been to treat its data problem as an analytics problem: better models, better extraction, better integration and more compute applied to the information we already have.
Those tools can make existing data enormously more useful.
But they cannot fill every gap.
The medical record remains exceptionally good at describing what happened when a patient interacted with the healthcare system. What happened on the ordinary Tuesday between two appointments can remain almost invisible.
In heart failure, we know from the literature that symptoms can worsen for days before hospitalization. Yet for an individual patient, those days may leave almost no trace in the record until the patient finally presents for care.
That's not primarily a modeling problem.
It's an asking problem.
And asking the same questions, consistently, across hundreds or thousands of patients is repetitive work that does not scale easily through human staffing alone.
That's the problem Cali was built to address: not simply finding more meaning in the data organizations already have, but helping collect information that otherwise may never enter the record.
Structured patient conversations between visits. On a cadence the organization defines. With meaningful responses available to the people who can act on them.
Because sometimes the most valuable data isn't hidden in the chart.
It hasn't been collected yet.
Sources
• Decompensated heart failure: symptoms, patterns of onset, and contributing factors, The American Journal of Medicine https://pubmed.ncbi.nlm.nih.gov/12798449/
• How long before hospital admission do the symptoms of heart failure decompensation arise? https://pmc.ncbi.nlm.nih.gov/articles/PMC6396952/
• Personal Care Utility: Health as Everyday Infrastructure, on time spent inside and outside clinical settings https://arxiv.org/pdf/2606.14145
• Healthcare analytics statistics on EHR adoption, unstructured healthcare data and interoperability https://www.knowi.com/blog/healthcare-analytics-statistics-2026/
