A Physician's Journey: Bringing Clinical Expertise to Medical AI
By Dr. Andrew Marsh — 2026-02-19
When Dr. Andrew Marsh finished his residency at a major academic medical center on the East Coast, he didn't expect his most consequential work would happen outside a hospital. Fourteen years into a career spanning internal medicine, biotech research, and clinical trials, he's now one of IXO's most active medical evaluators — spending 10 to 15 hours a week doing something he describes as essential: teaching AI systems how to think like a physician.
The problem he kept seeing
"Medical AI was being built without enough clinical input," he says plainly. "You'd see diagnostic tools that any first-year resident would question. The models were statistically sophisticated but clinically naive. They didn't understand how a presenting symptom changes meaning depending on a patient's history, their age, a dozen contextual factors that experienced physicians read instinctively."
He started exploring AI training work after a colleague mentioned it. What he found surprised him. The projects at IXO weren't data labeling in the way he'd imagined — they were substantive clinical reasoning tasks. Evaluating whether an AI system's differential diagnosis was medically sound. Ranking responses to patient queries by safety and accuracy. Annotating edge cases where model outputs were plausible but wrong in ways that could cause harm.
What the work actually involves
Dr. Marsh works across four types of projects. Clinical reasoning evaluation — assessing whether AI recommendations meet the standard of care. Safety review — identifying outputs that are technically coherent but clinically dangerous. RLHF for medical assistants — ranking AI responses to patient queries from best to worst. And literature annotation — turning clinical research into structured training data.
One project stands out in his memory. He spent three weeks evaluating an AI system designed to assist with differential diagnosis in emergency settings. His annotations — flagging cases where the model prioritized the statistically common diagnosis over the contextually correct one — contributed to a 30% reduction in false positives in internal testing.
"That's the thing about medical AI. The errors aren't random — they're systematic. The model learns a pattern that works 90% of the time and then fails badly on the 10% where clinical judgment would have caught it. Those are exactly the cases where you need an experienced physician in the loop."
On compensation and flexibility
He's direct about the practical side. "The rates reflect the stakes. This isn't survey work — it's specialist consulting, and IXO pays accordingly. I earn more per hour here than from most traditional medical consulting engagements."
The flexibility matters too. His IXO work happens around clinical responsibilities — early mornings, evenings, time between patient blocks. "There's no minimum hours, no fixed schedule. I decide when the work fits. That's genuinely rare."
Why it matters
"We're making decisions now that will affect how AI handles healthcare for decades. The models being trained today will be deployed at scale. If the training data reflects real clinical judgment — nuanced, contextual, hard-won — the outputs will too. If it doesn't, the errors will be baked in."
Read A Physician's Journey: Bringing Clinical Expertise to Medical AI on the IXO blog