Cover image

When Benchmarks Rot: Why Static ‘Gold Labels’ Are a Clinical Liability

Clinical AI has a paperwork problem. Not the usual paperwork problem, where doctors drown in documentation and everyone promises that software will save them. The more interesting problem sits one layer below: the paperwork used to judge the software may itself be wrong. That is the uncomfortable center of Scalable Stewardship of an LLM-Assisted Clinical Benchmark with Physician Oversight, a paper that audits MedCalc-Bench, a benchmark for testing whether language models can compute medical risk scores from patient narratives.1 The paper’s target is not a toy dataset. MedCalc-Bench covers 55 medical calculators and includes 10,053 training instances plus 1,047 test instances. Its labels were produced through an LLM-assisted pipeline: GPT-3.5 matched patient contexts to calculator questions, GPT-4 extracted clinical features, and Python scripts aggregated those features into final scores. ...

December 23, 2025 · 15 min · Zelina
Cover image

Painkillers with Foresight: Teaching Machines to Anticipate Cancer Pain

A patient says the pain is manageable. The medication chart looks stable. The latest score is not alarming. Then, sometime before the next formal reassessment, the pain breaks through. That is the operational problem behind Zhuang et al.’s study on predicting lung-cancer pain episodes with a hybrid machine-learning and large-language-model pipeline.1 The paper is not really about whether “AI can predict pain,” a sentence that sounds impressive until one remembers that dashboards have been predicting things since before consultants discovered the word “agentic.” The more interesting question is narrower and more useful: when should a hospital trust structured data, and when should it ask a language model to read the messy clinical story around the data? ...

December 19, 2025 · 15 min · Zelina
Cover image

Mutation Impossible? How Multimodal Agents Are Rewriting Glioma Diagnostics

Report First, Diagnosis Second A medical report usually arrives after the diagnostic work is done. It explains, records, justifies, and sometimes politely hides how messy the evidence really was. This paper asks a more interesting question: what if the report itself becomes a predictive object? In Multimodal Oncology Agent for IDH1 Mutation Prediction in Low-Grade Glioma, Hafsa Akebli and colleagues build a Multimodal Oncology Agent, or MOA, for predicting IDH1 mutation status in low-grade glioma using TCGA-LGG data, whole-slide histology, structured clinical variables, genomic context, and external biomedical knowledge sources.1 The immediate headline is easy enough: the full multimodal setup reaches the best reported performance, with an F1-score of 0.912. ...

December 8, 2025 · 15 min · Zelina
Cover image

Therapy, Transcribed: How LLMs Turn Conversation Into Clinical Insight

A therapist finishes a session. The call ends, the room becomes quiet, and the notes begin. There is the obvious record: what the client said, what the therapist asked, what homework was discussed. Then there is the harder record: what pattern kept returning? Was the client describing low motivation, fear of failure, family obligation, avoidance, self-criticism, or some collision among all of them? And if several patterns appeared, which one might be upstream of the others? ...

December 8, 2025 · 16 min · Zelina
Cover image

Timeline Triage: How LLMs Learn to Read Between Clinical Lines

Hospital notes are not databases that forgot to wear a spreadsheet costume. They are fragments of care: treatment names, planned cycles, delayed doses, discontinued regimens, relative dates, typos, abbreviations, and the occasional phrase that looks obvious until two clinicians disagree about what it actually means. For oncology, that mess matters. A chemotherapy timeline is not just a historical summary; it is the skeleton of a patient’s treatment journey. Get the timeline wrong, and downstream systems may misunderstand what was given, when it started, when it ended, and whether a patient fits a registry, audit, research cohort, or trial-matching rule. ...

December 7, 2025 · 16 min · Zelina
Cover image

Bridging the Clinical Gap: When Bayesian Networks Meet Messy Medical Text

Hospitals already have the data. That is the annoying part. They have diagnosis codes, medications, lab results, visit histories, and structured fields that look reassuringly database-friendly. They also have clinical notes: dense, abbreviated, unevenly written, and occasionally allergic to neat categories. A patient can have a symptom implied by the record, described vaguely in the note, omitted entirely, or mentioned in a way that conflicts with everything else. ...

November 24, 2025 · 17 min · Zelina
Cover image

Uncertainty, But Make It Clinical: How MedBayes‑Lite Teaches LLMs to Say 'I Might Be Wrong'

A hospital does not need a chatbot that sounds certain. It needs a system that knows when certainty would be irresponsible. That sounds obvious until one remembers how most AI demos behave: fluent answer first, caveat somewhere after the damage has already put on shoes. In clinical decision support, this is not a stylistic defect. It is an operating risk. A model can be wrong in many ways, but the most dangerous version is the confidently wrong one: the triage answer that should have been escalated, the medication suggestion that should have been checked, the risk score that looks clean only because the system has no vocabulary for doubt. ...

November 22, 2025 · 16 min · Zelina
Cover image

Filling the Gaps: How Bayesian Networks Learn to Guess Smarter in Intensive Care

ICU data has a habit of disappearing exactly when analysts would prefer it to behave. A blood gas is not measured. A pressure reading arrives late. A neurological score is absent because the patient is sedated, unstable, transferred, or simply surrounded by humans doing triage instead of satisfying a data scientist’s spreadsheet fantasies. Then, after the ward has produced this imperfect record, a model is asked to infer how the patient’s physiology evolved over time. ...

November 8, 2025 · 15 min · Zelina
Cover image

Knows the Facts, Misses the Plot: LLMs’ Knowledge–Reasoning Split in Clinical NLI

TL;DR for operators A model that can answer clinical fact-checking questions is not necessarily a model that can reason clinically. That is the inconvenient result of The Knowledge-Reasoning Dissociation: Fundamental Limitations of LLMs in Clinical Natural Language Inference, which introduces CTNLI, a controlled clinical NLI benchmark paired with Ground Knowledge and Meta-Level Reasoning Verification probes.1 ...

August 18, 2025 · 19 min · Zelina
Cover image

From Chaos to Care: Structuring LLMs with Clinical Guidelines

TL;DR for operators Patient records are not just long documents. They are timelines with consequences. CliCARE, the framework proposed in the paper, attacks that problem by turning longitudinal cancer EHRs into patient-specific temporal knowledge graphs, then aligning those patient trajectories with clinical guideline knowledge graphs before asking an LLM to generate a clinical summary and recommendation.1 That sounds architectural because it is. The useful lesson is not that “AI can help doctors,” a phrase now so overused it should probably be placed in quarantine. The lesson is that clinical AI improves when the model is given a structured representation of disease progression and a normative map of what should happen next. ...

July 31, 2025 · 16 min · Zelina