Back to Blog Therapeutic Areas

Rare Disease Enrollment Is a Different Problem

Sparse dot pattern representing rare patient population

When we talk about enrollment challenges in the context of most trials, the underlying problem is efficiency: there are enough patients who would qualify, but finding them and processing them through pre-screening takes longer than it should. The patient pool exists. The work is identification and verification.

Rare disease enrollment is a different problem. It is a scarcity problem first and an efficiency problem second. When a condition affects 1 in 10,000 people, a trial targeting 60 patients is drawing from a potential US population of around 30,000 to 40,000 diagnosed individuals, distributed across a geography that no single site network can efficiently cover. The tools and workflows that address efficiency in common disease trials are necessary but not sufficient here.

The math is different and it drives everything else

In a Phase 2 oncology trial with a solid tumor indication, a single academic medical center might see 200 to 400 potentially eligible patients in its existing population. The pre-screening challenge is manageable per site, and multi-site scale is achievable by adding more sites that each have meaningful local populations.

For a rare neurological condition with prevalence of 1 in 15,000, that same academic center might have 20 to 40 patients with the diagnosis in its system. Of those, a fraction will be in an appropriate disease stage, treatment line, or clinical status to qualify. The site might generate 3 to 6 consented patients over the course of a 24-month enrollment period if things go well.

A sponsor who needs 80 patients to power a Phase 2 rare disease trial cannot simply open 20 sites and expect proportional enrollment. The local populations are too small and too fragmented. The activation overhead per site is the same regardless of expected enrollment yield, and a network of 20 sites each generating 3 to 4 patients costs roughly the same to maintain as a network of 8 sites each generating 10. The economics force different site selection strategies and different sourcing approaches.

What distributed EHR querying actually means in this context

For common disease trials, EHR querying at a site is a tool for pre-screening a local population that is large enough to generate a manageable candidate list. For rare disease, a distributed querying approach, where the query runs across multiple sites simultaneously rather than sequentially, becomes more important because no single site has enough local patients to work from.

The technical challenge with distributed EHR querying for rare diseases is not query construction. It is data model variation. A rare neuromuscular diagnosis documented in Epic at one center may be coded using a different ICD hierarchy than the same diagnosis documented in Cerner at another. Specific biomarker or genetic confirmation data may live in different locations within each system. A query that returns correct results at Site A may miss a meaningful proportion of the same patient population at Site B.

Rare disease queries also tend to carry more complex criteria. The diagnosis itself may require confirmation from genetic testing, specific biomarker thresholds, or specialist assessment notes. Building a query that captures patients with confirmed diagnoses rather than suspected ones requires reading note content, not just checking structured fields. The same note-reading challenge that affects pre-screening for complex protocols is amplified in rare disease because there are fewer candidates to work from, so missing eligible patients is proportionally more costly.

Patient registries: the sourcing layer no site network replaces

In rare disease, disease-specific patient registries often hold more diagnosed patients than any single EHR network can surface. Registries maintained by patient advocacy organizations, condition-specific foundations, or academic disease-monitoring programs represent patients who have sought out the condition community, which tends to correlate with engagement and willingness to consider trials.

Working with registries introduces a different set of operational questions. Registry data is patient-reported and may not match the clinical documentation standards that eligibility review relies on. A patient who self-reports their diagnosis and treatment history in a registry needs to be cross-referenced against their medical records before eligibility can be confirmed. The registry is a sourcing layer, not a pre-screening tool.

Consent and data use arrangements with registries also add operational complexity. The registry holds patient data under a specific set of consents, and outreach for trial participation requires its own consent pathway. Sponsors who want to work with registries need to factor registry partnership negotiation time into their enrollment planning, which is often measured in months, not weeks.

The role of patient communities and social outreach

For some rare conditions, the diagnosed patient community is small enough that social network effects are relevant to enrollment. Patients and families affected by rare diseases tend to be highly networked, both through formal patient advocacy organizations and informally through disease-specific communities online. A trial that patients and advocates know about will reach corners of the diagnosed population that traditional site-based sourcing never will.

We are not saying that social outreach replaces clinical site operations. Site-based enrollment remains how eligibility is confirmed and patients are consented and followed. But the sourcing funnel for rare disease trials benefits from channels that are largely irrelevant in common disease trials. A patient who hears about a trial through a community forum and contacts the trial team directly still needs to be matched to a participating site, have their EHR records reviewed, and go through a formal screening visit. The community reach changes where patients enter the funnel, not how they move through it.

What this means for the tools we build

We work with rare disease protocols differently than we work with common indication protocols. The pre-screening volume at any single site is lower, which means the efficiency gains from faster chart review are proportionally smaller per site. The value we add in rare disease settings comes more from query comprehensiveness than from processing speed: making sure that patients in a site's population who might qualify are not missed due to documentation variation or note-reading limitations.

The distributed aspect matters. For sponsors running rare disease trials across a network of 15 to 25 sites, each with small local populations, the ability to run a consistent first-pass screen across all sites simultaneously, rather than site by site over weeks, has a meaningful effect on how quickly the trial builds a candidate list that is worth actioning.

We do not solve the scarcity problem. If a condition is rare enough that the global diagnosed population is smaller than the trial's enrollment target, no matching tool changes that. What we can affect is the proportion of the diagnosed population that actually gets surfaced within the site network, and the time it takes to work through that population once it is identified.