AI Digs Up 26-Year-Old Paper, Cracks 18 Rare Disease Cold Cases

Source: OpenAI | Published: 2026-08-04T20:29:24Z

After AI surfaced a forgotten 1998 paper on S1PR1, Catherine emailed a scientist in the next building — he showed up at 4 p.m. that afternoon.


Genomic data from 376 families, screened by AI, yielded answers for 18 long-undiagnosed rare disease cases. That's the result of a joint study between Boston Children's Hospital's Manton Center for Orphan Disease Research and OpenAI. The number sounds modest — but for every family that spent years lost in the medical system, it means everything.

Eight Years to a Diagnosis

Stav Rones, one of the study's participants, is a software engineer from Boston. As a teenager, he began noticing something was off with his body — certain pains and movement limitations that his peers simply didn't have, but that no one could explain.

It took him three to four years just to establish that the problem wasn't muscular or skeletal, but neurological. Then another five years to narrow it from a neurological condition to a specific rare genetic disease. Nearly a decade in total.

In the world of rare disease, that's not an outlier — it's the norm. Medicine even has a term for it: the "diagnostic odyssey," describing the years spent drifting between departments, tests, and dead-end diagnoses. The global average is six to seven years.

Only after receiving his diagnosis could Stav connect with drug development teams working specifically on his condition, and only then did he begin seriously thinking about genetic risk in family planning. Before a diagnosis, he said, you don't even know what you might be doing to harm yourself — because you don't know what you have.

A Genome Is a Dataset with Half a Million Rows

Alan Beggs, director of the Manton Center at Boston Children's Hospital, explained why this is technically so hard.

The human genome contains roughly three billion base pairs, encoding about 20,000 genes. Of those, between 8,000 and 9,000 are known to be disease-associated — and that number keeps growing. After sequencing, the raw dataset contains approximately 500,000 rows, each representing a genetic variant, each with 20 to 30 fields: population frequency, which gene is affected, predicted impact on protein function, and more.

After filtering out common variants, you're still left with several thousand to 10,000 candidate variants across 500 to 1,000 genes.

Manton Center scientific director Catherine Brownstein described her actual workflow: reviewing interesting variants one by one, cross-referencing databases, digging through the literature on each gene, and sometimes spending hours only to conclude it's a dead end and start over. "We know an enormous amount about a small fraction of genes, but no one can be an expert in all 8,000 disease genes," Alan said. That's precisely where AI can step in.

What the Model Does: Compress Half a Million Rows to Two to Six Candidates

The research team's approach was to use AI for literature retrieval and hypothesis generation — narrowing the candidate space before handing off to human experts for judgment.

Machine learning researcher Suyash Shringarpure walked through the process: they worked with Alan and Catherine to define the prompt structure — what information the model should receive, what format it should output, and which analytical pitfalls to avoid. The model's output isn't a black-box verdict but an evidence-backed candidate list: this variant, the literature supporting this hypothesis, and the reasoning chain. Human experts then decide whether to pursue follow-up testing, and confirmed findings are returned to patients.

Before analyzing the 376 undiagnosed cases, the team validated the model against solved cases with known answers, iteratively refining the prompts — identifying recurring error patterns, such as false positives from insufficient sequencing depth — until the model hit 80 to 90 percent accuracy on known cases. Only then did they feel confident enough to take the model's output seriously in real analysis.

A 26-Year-Old Paper, in the Building Next Door

Catherine described one case that stood out above all others.

It was an early case from the registry — the Manton Center now has over 3,000 cases, and this was roughly number 151. The patient presented with vitiligo, transposition of the great arteries, and pulmonary hypertension. No obvious genetic connection.

The model flagged a candidate gene: S1PR1. Catherine's first reaction was "what on earth is this?" But she had the model run a full literature search anyway. It surfaced a paper published roughly 26 years ago that explicitly proposed S1PR1 as a serious candidate for vitiligo — and then drew on multiple additional papers to piece together the full molecular pathway.

Catherine did a quick Google search and found that the scientist who cloned S1PR1 back in 1988 works in the building next door to Boston Children's Hospital. She sent an email; he said to come by at 4pm. The case is now moving into deeper investigation.

"Given unlimited time, I would have eventually found that paper," Catherine said. "Or a motivated postdoc would have. But the model just does it that fast. We pat ourselves on the back for getting to page four of PubMed — the model went to page thirty."

The Answer the Model Got Wrong Actually Found Two Diseases

There's one detail from the validation phase worth highlighting. The team gave the model 20 solved cases. It got 19 right.

When they reviewed the one it got "wrong," they found that the model's answer differed from the human expert's conclusion for a specific reason: this patient had two distinct genetic variants simultaneously, each causing different symptoms. The human experts at the time had focused on only one. The model flagged both.

This means that in some cases, the premise itself — "one diagnosis explains all symptoms" — is wrong. Without that assumption baked in, the model is more likely to catch compounding conditions.

2,000 Families Still Waiting

The Manton Center currently has approximately 5,000 registered families, around 2,000 of whom still have no diagnosis.

Alan said AI's most obvious value right now is acceleration and efficiency — narrowing possibilities faster within the bounds of existing knowledge. But the S1PR1 case Catherine described represents the next level: the model synthesized existing literature to surface a hypothesis that human researchers hadn't considered, and that hypothesis turned out to matter.

What Suyash is actually hoping for is a future where every time a new paper is published, the model automatically reruns all related undiagnosed cases — and pushes a notification to researchers when something new turns up. A genome doesn't change, but knowledge about genes updates every day. Periodic, low-cost reanalysis is itself a new diagnostic capability.

For now, the workflow still requires the Manton Center's infrastructure to operate. The team is working with OpenAI to turn it into a tool any clinician can use — not just top-tier medical centers, not just genomics specialists. Whole-genome sequencing now costs under $1,000, cheaper than a single MRI. Insurers are still watching from the sidelines, but the technical barriers are no longer the main obstacle.

More articles on TLDRio