Cover image

The First Scribble Does Most of the Work

TL;DR for operators When an automated system can be corrected repeatedly, the product question is not simply whether correction helps. It is how many corrections are worth asking a user to make. Sherif and colleagues test that question in interactive multitracer PET/CT lesion segmentation.1 Mean Dice rises from 0.5539 with no correction to 0.7223 after one scribble and 0.7512 after five. Lesion-level F1 follows the same pattern: 0.5281, then 0.7036, then 0.7326. Roughly 85% of the five-round gain arrives after the first correction. ...

October 1, 2026 · 6 min · Zelina
Cover image

Make the Structure Survive: Forma Turns Synthetic Clinical Cases Into Auditable Outputs

TL;DR for operators If synthetic cases are going to support training or evaluation, surface plausibility is a weak control objective. The harder requirement is preserving the relationships that make each case meaningful. Forma tests that idea by specifying a person-specific psychological structure before generation and asking whether those directional relationships can be recovered afterward. In the full condition, an external probe reaches MCC +0.41 and AUC 0.70 for directed-edge recovery. When the structural formulation is removed, performance falls close to chance: MCC +0.03/AUC 0.52 with demographics and self-report still present, and +0.01/0.50 under zero-shot generation. ...

September 25, 2026 · 7 min · Zelina
Cover image

More Agents, More Rules: HiMA-MDD Treats Multi-Agent AI as a Governance Problem

TL;DR for operators Dividing a high-stakes decision among specialist agents does not specify who may see which evidence, who owns each subdecision, or who can revise it. HiMA-MDD1 turns those choices into explicit system rules for PHQ-8 assessment from completed multimodal clinical interviews. The strongest architectural signal is not “more agents perform better.” Before global verification, four specialists produce the lowest total-score error, while two specialists produce the highest screening kappa and Macro-F1. Removing cross-factor audit and targeted revision causes the largest screening-performance decline among the paper’s three component ablations. Giving four specialists bounded item-specific evidence also beats giving them a capped shared evidence pool on every reported pre-verification metric, although that experiment changes both evidence composition and context length. ...

September 25, 2026 · 6 min · Zelina
Cover image

The Leaderboard Is Not a Clinical Clearance

TL;DR for operators A healthcare team choosing a model, configuration, and safeguards for diagnostic support should treat leaderboard leadership as evidence of capability—not proof of clinical readiness. GPT-5-medium led ClinMM-Bench, yet produced a completely correct diagnosis in only 33.88% of cases. The benchmark tests a difficult but essential requirement: combining clinical details and images as evidence unfolds, revising earlier conclusions, and explaining a diagnosis without omitting decisive information or introducing unsupported claims. Medical specialization and explicit reasoning settings do not improve these abilities consistently across model scales, metrics, or specialties. ...

August 6, 2026 · 8 min · Zelina
Cover image

The Answer Looked Clinical. The Critical Steps Were Missing.

TL;DR for operators A clinical assistant can follow instructions, produce a complete-looking answer, and read like expert work without performing the reasoning needed for a safe clinical decision. Those surface qualities may justify further evaluation, but they do not justify procurement or decision authority. Across five deliberately difficult tasks, the models satisfied 80–90% of the least consequential criteria but only 32.4–41.7% of the most clinically consequential criteria. All three models missed 56 of 108 critical requirements. The systems were therefore strongest where omissions mattered least and weakest on reasoning and safety steps most closely tied to harm. ...

August 1, 2026 · 9 min · Zelina
Cover image

Count the Missing, Weight the Rare: A Better Bargain for Cardiac Phenotyping

TL;DR for operators CW-B is not interesting because it discovers a novel model architecture. It is interesting because it treats three routinely neglected design decisions as one system: Rare classes receive more influence during training. Class weights are computed separately inside each training fold, preventing common phenotypes from dominating the tree-building process. Missing values retain their provenance. The pipeline imputes a numerical value but also adds an indicator showing that the original measurement was absent. Clinical priorities are audited separately from aggregate performance. Stable coronary artery disease, acute coronary syndrome, and non-obstructive coronary disease are grouped into a predefined evaluation set because missing them can alter follow-up and treatment pathways. On a five-class dataset containing 4,354 patient records and 57 structured features, CW-B records the strongest accuracy, Macro-F1, balanced accuracy, and prioritized F1 among the tested tree, ensemble, and neural baselines. Its balanced accuracy reaches 0.73, compared with 0.66 for a larger XGBoost baseline. Its prioritized F1 is 0.69, compared with 0.67 for that baseline. ...

July 13, 2026 · 17 min · Zelina
Cover image

The Rule Is the Model: DEM’s Case for Bedside Anomaly Detection Without Explainer Theatre

Alerts are cheap; trusted alerts are not A hospital monitor that screams without explaining itself is not a decision-support system. It is a very expensive doorbell. That is the practical problem behind Singh, Roy, Bose, and Hota’s Distilled Explanation Model, or DEM, for physiological anomaly detection in wireless body area networks.1 The paper is nominally about clinical sensor data: heart rate, oxygen saturation, blood pressure, temperature, stress signals, sensor dropouts, and ICU monitoring. But the more interesting argument is architectural. DEM is not trying to make a black-box model more charming after it has already made a decision. It is trying to make the explanation part of the decision itself. ...

June 14, 2026 · 17 min · Zelina
Cover image

Chart Check: Why Clinical Summaries Need Detectors Before Alignment

Chart review is the boring part of medicine, which is exactly why AI systems should learn from it. A clinical discharge summary does not fail only when it sounds clumsy. It fails when it tells a patient something that did not happen, invents a medication change, adds a procedure, misstates a timing detail, or turns a vague note into a confident medical fact. The prose may still be smooth. The bedside manner may even be excellent. Unfortunately, a hallucination delivered in fluent patient-friendly language is not safer because it has better manners. ...

June 2, 2026 · 17 min · Zelina
Cover image

Twin Peaks: When Alzheimer’s AI Learns to Remember What Clinics Forget

Opening — Why this matters now Healthcare AI has spent years trying to look impressive in carefully lit laboratory conditions. Alzheimer’s disease, with its irregular follow-ups, missing scans, incomplete biomarkers, and deeply uneven patient trajectories, is less polite. It is not a clean benchmark. It is a bureaucracy of biology. That is why the arXiv paper “CognitiveTwin: Robust Multi-Modal Digital Twins for Predicting Cognitive Decline in Alzheimer’s Disease” deserves attention.1 It does not merely ask whether a model can classify Alzheimer’s disease from a snapshot. That problem is already crowded, noisy, and occasionally dressed up as clinical transformation. Instead, the paper asks a harder and more operationally relevant question: can an AI system model an individual patient’s cognitive trajectory over time, using fragmented clinical evidence, while remaining accurate, calibrated, and fair across demographic groups? ...

April 29, 2026 · 12 min · Zelina
Cover image

MARCH Orders: When AI Holds a CT Case Conference

The useful meeting, unfortunately, exists Meetings are usually where productivity goes to file a complaint. But there is one kind of meeting that high-stakes work still needs: the review session where a first draft is challenged, evidence is checked, and a senior decision-maker signs off. Radiology has long understood this. A resident may draft the report. A fellow may question the interpretation. An attending radiologist resolves the remaining uncertainty. The point is not ceremony. The point is controlled disagreement. ...

April 22, 2026 · 16 min · Zelina