Cover image

Seeing Is Believing: Why Visual RAG Might Be the Missing Layer in Clinical AI

Guidelines are not novels. That sounds obvious until we remember how most retrieval-augmented generation systems treat them. A clinical guideline becomes text. The text becomes chunks. The chunks become embeddings. The embeddings become “context.” Somewhere in that mechanical conversion, a dosing table, a referral pathway, or a threshold hidden inside a flowchart quietly loses its shape. Then everyone acts surprised when the answer is fluent but clinically thin. Very mysterious. ...

March 24, 2026 · 13 min · Zelina
Cover image

The Memory That Thinks: When AI Stops Remembering and Starts Reasoning

A memory mistake is still a mistake Memory sounds comforting until it remembers the wrong thing. Imagine a clinical AI agent facing a patient whose disease appears to be regressing after prior treatment. A past case in memory says that conflicting cancer signals should not be trusted too quickly. That sounds relevant. It even sounds cautious, which is the preferred costume of many bad decisions. But in this case, the regression is not noise. It is the signal. Treating it as a conflict leads the agent toward unnecessary systemic therapy rather than watchful waiting. ...

March 24, 2026 · 17 min · Zelina
Cover image

When Models Know But Won’t Act: The Interpretability Illusion

Triage is a wonderfully cruel test for AI safety. A patient message arrives. Maybe it is routine. Maybe it contains a medication interaction, an allergic reaction, suicidal ideation, a pregnancy-related risk, or a pediatric emergency. The model is not being asked to compose poetry, summarize a quarterly report, or role-play as an overenthusiastic consultant. It has one job: notice the hazard and recommend action. ...

March 21, 2026 · 17 min · Zelina
Cover image

Diagnosis, But Make It Iterative: When AI Learns Like a Doctor

Diagnosis begins with a small nuisance: the patient does not arrive as a completed spreadsheet. They arrive with pain, fragments, missing context, contradictory clues, and a clock running somewhere in the background. A doctor does not usually receive the full record, press “classify,” and return a disease label. The doctor asks for a physical exam, orders labs, checks imaging, updates the differential, and decides whether the next test is useful or merely expensive decoration. ...

March 13, 2026 · 17 min · Zelina
Cover image

When the Brain Refuses to Tick: Continuous-Time AI for Seizure Forecasting

The brain is not a metronome A hospital monitor has a clock. A machine-learning pipeline has windows. A spreadsheet has rows. The brain, inconveniently, has none of these manners. Electroencephalography, or EEG, records electrical activity as a continuous stream across multiple scalp channels. Clinical AI systems then often chop that stream into fixed segments, transform each segment into features, and ask a classifier a familiar question: seizure or not seizure, abnormal or normal, risk or no risk. ...

February 27, 2026 · 16 min · Zelina
Cover image

Heartbeat in Stereo: Why ECG AI Needs Both Contrast and Context

ECG models have a deceptively simple job: read a heartbeat and infer what might be wrong. The real problem is that a heartbeat is not a single line of data. A standard 12-lead ECG is a coordinated view of cardiac electrical activity from multiple spatial angles. Meanwhile, the associated clinical report is not a clean label. It is a human-written summary: useful, compressed, inconsistent, and occasionally full of stylistic residue. Medicine, regrettably, still contains humans. ...

February 25, 2026 · 14 min · Zelina
Cover image

Mind the Gap: When Clinical LLMs Learn from Their Own Mistakes

Mistakes are usually treated as waste. In clinical AI, they are treated even more nervously: logged, redacted, escalated, converted into a slide deck, and then politely buried under the next benchmark table. Understandable. Nobody wants a medical agent whose product roadmap reads like “learning through patient-adjacent embarrassment.” But the paper Closing Reasoning Gaps in Clinical Agents with Differential Reasoning Learning makes a useful move: it treats mistakes not as isolated failures, but as a structured raw material for improving future reasoning.1 The core idea is not that a clinical LLM should “reflect” harder, nor that we should throw more guidelines into the prompt until the context window starts whimpering. The idea is more surgical: compare the model’s reasoning with a better reference reasoning trace, locate the precise gap, convert that gap into a reusable instruction, and retrieve that instruction when a similar case appears later. ...

February 11, 2026 · 17 min · Zelina
Cover image

When 100% Sensitivity Isn’t Safety: How LLMs Fail in Real Clinical Work

Clinic. That is where the comforting AI story starts to wobble. In a benchmark, a clinical model receives a clean question, enough context, and a scoring rule that usually rewards the right answer. In a clinic, the same model sees an elderly patient with multiple conditions, incomplete records, medication changes from years ago, possible specialist involvement, ambiguous prescribing history, and a problem that may not require action at all. The model is not merely being asked, “Can you spot a risk?” It is being asked, “Do you understand whether this risk is real, current, important, and safely actionable?” ...

December 25, 2025 · 20 min · Zelina
Cover image

When 1B Beats 200B: DeepSeek’s Quiet Coup in Clinical AI

Chest X-rays are not a glamorous AI benchmark. They are routine, repetitive, and brutally operational. A hospital does not need a model that can write poetry about radiology. It needs reports that are accurate enough, fast enough, structured enough, and cheap enough to run inside an actual clinical workflow without turning the IT department into a cloud-billing support group. ...

December 24, 2025 · 15 min · Zelina
Cover image

When Bigger Isn’t Smarter: Stress‑Testing LLMs in the ICU

A hospital does not buy “intelligence.” It buys a workflow. That distinction sounds obvious until an AI vendor arrives with a model that has billions of parameters, a clinical pretraining story, and the gentle implication that smaller models are now museum pieces. In the ICU, however, the useful question is not whether the model can talk like a doctor. It is whether it can detect tomorrow’s clinical deterioration from messy notes better than simpler systems that cost less, run faster, and attract fewer infrastructure headaches. ...

December 24, 2025 · 12 min · Zelina