A synthetic third-line colorectal cancer case study built on Claude Science — how the mechanical-but-hard work between a molecular report and a patient’s next step is done reliably, traceably, and closed to care.
The complete AiMOneHealth cycle on Claude Science runs ten stages from source-file ingest to a patient-ready action plan, each emitting a machine-readable artifact — with a multidisciplinary panel and the OncoNavigator roadmap closing the loop from analysis to care. Here is one hard case, worked end to end.
PX-SYN-CRC-LIVER-001 is a synthetic 59-year-old with left-sided sigmoid colorectal adenocarcinoma that has spread only to the liver. At diagnosis: cT4a N2 M1a, seven bilobar liver metastases, unresectable. The molecular picture is specific and consequential — microsatellite-stable (MSS/pMMR), low tumor mutational burden (4.1), RAS and BRAF wild-type on both tissue and a recent liquid biopsy, HER2 not amplified. The germline carries a UGT1A1 *1/*28 variant that changes how one chemotherapy drug is dosed; DPYD is normal.
The patient has already been through two lines of therapy. First-line chemotherapy plus an anti-EGFR antibody worked well enough to convert the liver to resectable and achieve a period with no evidence of disease. The cancer returned in the liver, second-line therapy held it briefly, and it is now progressing again — the tumor marker CEA has climbed from a nadir of 6 back to 95. The disease is still liver-limited, and the patient is fit (ECOG 1).
This is a hard call. The patient is running out of standard lines. The wild-type RAS/BRAF status opens a door that most colorectal patients don’t have, but the tumor is MSS — which closes the immunotherapy door that gets so much attention. And the liver-limited pattern raises a surgical question that a purely systemic analysis would miss entirely. Answering it well means holding molecular, medical-oncology, surgical, and pharmacologic reasoning in view at once — and being able to show the work.
Step 1 — Reconcile the record before trusting it. The five source files (clinical profile, molecular report, somatic mutation file, germline VCF, labs) were cross-checked against each other with 18 explicit consistency tests: do the somatic variant allele fractions match the molecular report? Does the germline VCF confirm the UGT1A1 status the profile claims? Does the CEA trajectory match the story of response and relapse? All 18 passed, zero conflicts. Only then did interpretation begin — every later conclusion rests on data shown to be internally coherent.
Step 2 — Map every biomarker to a therapy verdict, including the “no”s. Each option was classed with its driving biomarker named. The MSS/TMB-low status makes single-agent immunotherapy NOT INDICATED — a clear, cited “no” that spares the patient an ineffective drug. HER2-, BRAF-, NTRK- and KRAS-G12C-targeted agents are NOT APPLICABLE because those alterations are absent. The one biomarker that opens a door — durable prior anti-EGFR benefit on wild-type RAS/BRAF, still wild-type on liquid biopsy — is what makes anti-EGFR rechallenge SUPPORTED.
Step 3 — Match live trials, and surface the gate. A live ClinicalTrials.gov query screened 139 unique recruiting trials, ranked by biological fit and real-world accessibility, each with a per-criterion eligibility map and a single gating criterion. The top match’s gate is surgical candidacy; the second enrolls only at second line (this patient is third); another is gated purely by geography (Italy-only), and a fourth by a combination of geography and a first-line-progression pattern this patient doesn’t fit. The gate is the one line that tells the team whether a trial is worth pursuing.
Step 4 — Verify every citation against the live literature. Each landmark anchor carries a PubMed identifier and DOI checked at query time. The pipeline even corrected itself — one starting identifier resolved to an unrelated cystic-fibrosis paper and was replaced with the correct trial reference. Verified, not recalled.
Step 5 — Corroborate the drivers on GPU. On an NVIDIA RTX 5090, the ESMC-600M protein language model scored the tumor’s four missense driver mutations for predicted effect (zero-shot log-likelihood ratio, mutant vs. wild-type). All four came back deleterious — FBXW7 R465C −13.8, SMAD4 R361H −8.0, TP53 R175H −5.75, PIK3CA E545K −3.44 (more negative = more disruptive). An independent, biophysical second opinion, concordant with the molecular report — framed as research corroboration, not clinical classification.
The analysis produced a ranked, evidence-cited plan for tumor-board review:
And explicitly not the paths that look tempting but don’t fit: single-agent immunotherapy (MSS), BRAF-targeted therapy (wild-type), HER2- and other targeted agents (biomarkers absent).
Crucially, the answer names the tests that would most change it — a fresh pre-treatment ctDNA, a hepatobiliary MDT with liver volumetry, restaging imaging to confirm the liver-limited pattern, a HER2 re-test, and confirmation the MSI status is unchanged. A recommendation that knows what could overturn it is a more honest recommendation.
For a medical readership, the value of this cycle is not the software — it is that every decision is a testable hypothesis paired with the evidence that confirmed or refuted it. Below are the five clinical hypotheses the case turned on, each stated in medical terms, each validated against a datum or a simulation result, and each mapped to the decision it drove.
H1 — The EGFR axis is still druggable. Hypothesis: despite two prior lines, the tumor remains RAS/BRAF wild-type and has not acquired an EGFR-pathway resistance mechanism, so re-exposing it to anti-EGFR antibody should re-engage a validated target. Validation: circulating tumor DNA (2026-06-12) confirmed RAS/BRAF wild-type maintained, with no acquired EGFR-ectodomain (S492R), MET, or HER2 alterations — and tissue and liquid-biopsy calls agreed across all 18 cross-artifact consistency checks. Decision: anti-EGFR rechallenge became the SUPPORTED lead option, explicitly gated on a repeat pre-treatment ctDNA. This is the CRICKET/CHRONOS rechallenge paradigm applied to a patient whose biomarker profile fits it.
H2 — The somatic drivers are genuinely function-altering. Hypothesis: the four missense mutations called from the molecular report (TP53 R175H, PIK3CA E545K, SMAD4 R361H, FBXW7 R465C) are true drivers, not passenger noise. Validation: we scored each on GPU with the ESMC-600M protein language model as a zero-shot log-likelihood ratio (LLR) of the mutant versus wild-type residue. All four are deleterious (LLR ≤ −3.4) — and the underlying probabilities are stark: the model assigns the wild-type residue a probability of 0.62–1.00 and the mutant residue as little as 0.000001 (one in a million) for FBXW7 R465C.
H3 — Single-agent immunotherapy will not help. Hypothesis: because the tumor is microsatellite-stable and TMB-low, it lacks the neoantigen burden that makes checkpoint blockade effective. Validation: MSI status is MSS, TMB is 4.1 mut/Mb (low); the KEYNOTE-177 pembrolizumab benefit is defined in the MSI-high population — precisely the group this patient is not in. Decision: single-agent anti-PD-1 marked NOT INDICATED, a cited negative that spares an ineffective and costly therapy.
H4 — Irinotecan needs dose attenuation. Hypothesis: the germline UGT1A1 genotype predicts impaired SN-38 glucuronidation and elevated irinotecan toxicity. Validation: the germline VCF confirms UGT1A1 *1/*28 heterozygous (rs8175347, TA7), with DPYD normal; CPIC guidance calls for reduced-dose irinotecan. Decision: any irinotecan-containing rechallenge backbone carries a pre-specified dose reduction with hematologic monitoring — a pharmacogenomic safeguard applied before the first dose.
H5 — The disease is truly liver-limited. Hypothesis: the relapse is confined to the liver, keeping local therapy on the table. Validation: restaging shows two hepatic lesions (segments II/VI) with no extrahepatic disease, and the CEA trajectory — risen from a nadir of 6 to 95 — is consistent with a hepatic-only relapse rather than diffuse progression.
Decision: repeat liver-directed local therapy (redo resection / ablation / SBRT) marked CONSIDER, gated on a hepatobiliary-MDT future-liver-remnant volumetry after the prior right hepatectomy.
What the validation adds up to. Four of five hypotheses were confirmed and drove a positive or protective action; one (H3) was a confirmed negative that removed an option. No decision rests on an unverified assumption. The GPU simulation is deliberately positioned as research corroboration — a second, orthogonal line of evidence that the drivers are real — not as a clinical variant classifier; its role is to strengthen confidence in the molecular picture the therapy plan is built on, with full transparency about what it can and cannot claim.
A recommendation is only trustworthy if the plausible alternatives were actively considered and defeated — not merely un-mentioned. In this case the cycle behaved like an internal critic, taking the six assumptions that could most easily have derailed the plan and confronting each with evidence. Disproving the alternative is as decisive as proving the hypothesis: a confirmed “no” removes a wrong turn, and every wrong turn removed narrows the decision toward the one path the data actually support.
A1 — “The wild-type call is stale.” The single most dangerous assumption: that the RAS/BRAF wild-type status, established before first line, no longer holds — anti-EGFR rechallenge would then be futile or harmful. This was not assumed away. A fresh liquid biopsy re-tested RAS/BRAF and actively screened for the known acquired anti-EGFR resistance mechanisms (EGFR-ectodomain S492R, MET amplification, HER2). None were present. Only after this negative screen did rechallenge stay the lead option — and it remains gated on one more pre-treatment ctDNA. Ruled out.
A2 — “The drivers might be passengers.” If the missense mutations were incidental rather than functional, the molecular rationale would soften. The GPU protein-language model tested this independently: a sequence-only model that never saw the clinical annotation scored all four as deleterious. The alternative — that these are noise — is not consistent with the evidence. Ruled out.
A3 — “MSS tumors sometimes respond to immunotherapy anyway.” A tempting move, given how much attention checkpoint blockade receives. The critic checked the biomarker basis: TMB is low (4.1), the tumor is MSS/pMMR, and the pembrolizumab benefit (KEYNOTE-177) is defined in the MSI-high population. There is no biomarker rationale in this patient. Eliminated — and the elimination saves the patient an ineffective, costly, potentially toxic line.
A4 — “The UGT1A1 effect is marginal; full-dose irinotecan is fine.” Left unchallenged, this risks grade 3–4 neutropenia and diarrhea. The germline genotype (*1/*28) and CPIC evidence say otherwise — and the patient’s own second line had already required a 20% irinotecan reduction. The assumption is corrected to a pre-specified reduced dose.
A5 — “The disease may be more diffuse than it looks.” If true, local liver therapy would be pointless. Restaging showed two hepatic lesions and no extrahepatic disease, with a CEA pattern consistent with hepatic-only relapse — so the option is retained, but gated on a future-liver-remnant volumetry rather than assumed feasible.
A6 — “Any RAS/BRAF-wild-type trial is a fit.” The most common matching error. Per-trial eligibility maps surfaced the gating criterion for each candidate: two are geography-locked, one enrolls only at an earlier line, one requires surgical candidacy. Most apparent matches were filtered out before they reached the patient.
Why the negative results matter. Three of these six (A1, A2, A3) are cases where the most valuable output was a refutation — the plan is trusted precisely because the obvious objections were raised and answered, not skipped. This is the discipline a molecular tumor board applies by instinct; here it is made explicit and auditable, so a reviewer can see not only what was recommended but what was considered and rejected, and on what evidence. A recommendation that has survived its own strongest counter-arguments is a fundamentally different object from one that was simply asserted.
A ranked recommendation is where most analyses stop. This one continued through two layers that turn an answer into a plan a patient can follow.
Seven specialist reviewers — medical oncology, hepatobiliary surgery, radiation oncology, interventional radiology, molecular pathology, clinical pharmacology, and patient navigation — each reviewed the same case brief from their own discipline, independently. All seven concurred with the lead plan. The value was in the conditions each attached:
The panel consensus was handed to a patient-navigation layer that rewrites it as a care roadmap: a one-sentence plan headline, sequenced next steps (what happens, why it matters, who leads it, how to prepare), the checkpoints where the path might branch, questions to bring to the care team, and the logistics — travel, cost, second opinions. The molecular tumor board answers “what does the evidence support?” The navigator answers “what happens Monday, who do I call, and what should I be ready for?”
Each stage was a deliberate methodological choice. Here is what runs under the hood, the reasoning, and the specific thing each choice changes.
| Design choice | Why it matters | What it changes |
|---|---|---|
| Cross-artifact consistency, not trust. Five source files reconciled with 18 explicit checks; conflicts flagged, never silently resolved. | A single transcription mismatch between the VCF and the summary can flip a decision. | The recommendation rests on data proven internally coherent — the whole downstream chain inherits that guarantee instead of a hidden error. |
| A matrix that names the negatives. Every option classed SUPPORTED / STANDARD / CONSIDER / NOT INDICATED / NOT APPLICABLE with its driving biomarker, each anchored to a guideline and a verified landmark trial. | The expensive errors in oncology are drugs given against the biomarker. | An auditable “why this is wrong for this patient” prevents the mis-indicated therapy before it is proposed. |
| Live trial matching with a gating criterion. Recruiting trials pulled live (query date recorded), ranked by fit and accessibility, each with an eligibility map and one surfaced gate. | A trial the patient can’t enroll in is not a match. | It turns a list of NCT numbers into a triage decision in a single line, collapsing hours of eligibility reading. |
| Citations verified against the live literature. Every anchor carries a PMID/DOI confirmed at query time; a provenance manifest records method and date. | A recommendation is only as strong as its weakest citation. | Verifying rather than recalling is the line between auditable decision-support and plausible fiction — and self-catching a wrong citation is exactly the failure mode that erodes trust in AI. |
| GPU protein-language corroboration. ESMC-600M scores missense drivers as zero-shot LLRs on an RTX 5090 (PyTorch 2.11 / CUDA 12.8); ESMFold2-Fast structural folding optional. | An independent signal from a model that never saw the clinical annotation. | It adds a different kind of evidence — biophysical, not database lookup — as a second opinion on the drivers. |
| Provenance as a first-class output. Every artifact carries patient ID, sources, query date, uncertainty notes, and a not-medical-advice line. | A lighthouse project must be auditable by someone who wasn’t in the room. | Traceability built in, not bolted on, is what makes the whole cycle deployable rather than a demo. |
This case is a template, not a trophy. To reproduce or adapt it to your own disease and data, here is the concrete substance of each layer.
There is at present no live “Multidisciplinary” or “OncoNavigator” connector on the platform — the panel and navigator layers are produced as agent-simulated reviews plus the structured JSON those systems would consume. That is deliberate and transparent: the artifacts are built to the interface a real MDT-coordination or patient-navigation system would expect, so wiring them into a live instance is a matter of connecting an endpoint, not re-doing the science.
The work between a molecular report and a patient’s next step — reconciling data, screening trials, verifying evidence, convening perspectives, translating for the person living it — is exactly the work that today limits how many patients get a genuinely considered, evidence-anchored plan. Every expert hour a cycle like this saves goes back to patients; every mis-indicated therapy the matrix prevents, every inaccessible trial the gate filters out, every plan the navigator makes livable, is a step toward care that is faster, safer, and more humane. Onboarding into this cycle means contributing to a template whose entire purpose is to turn precision-oncology capability into quality of life for real people — auditably enough to trust.
Faced with a hard third-line question — what to give a fit patient with MSS, RAS/BRAF wild-type, liver-limited colorectal cancer who has exhausted two lines — Claude Science reconciled the record, mapped every biomarker to a verdict, screened 139 live trials with their gates, verified every citation, and corroborated the drivers on a GPU. The answer: anti-EGFR rechallenge, gated on a fresh liquid biopsy, with repeat liver-directed therapy on the table pending a volumetric assessment — reviewed and conditioned by a seven-discipline panel, then translated into a patient roadmap. Every step left a provenance-carrying artifact behind.
The through-line: do the mechanical-but-hard work reliably and traceably — reconcile, verify, match, corroborate, translate — so scarce human expertise is spent on judgment. That combination of auditable synthesis and a genuine research-to-care handoff is what makes this cycle a template for precision oncology, not a one-off demo.