Cover image

Grading the Doctor: How Health-SCORE Scales Judgment in Medical AI

Checklist is a boring word. That is why it is useful. In healthcare AI, the glamorous question is whether a model can “reason like a doctor.” The operational question is uglier: did it invent a lab value, miss an emergency referral, overstate certainty, ignore the requested format, recommend unsafe antibiotics, or fail to ask for missing context? ...

February 2, 2026 · 15 min · Zelina
Cover image

When LLMs Invent Languages: Efficiency, Secrecy, and the Limits of Natural Speech

Chatbots are trained to sound human. Enterprise AI agents are increasingly asked to behave like colleagues: pass information, coordinate actions, summarize context, and explain what they are doing in language people can read. That arrangement feels safe because natural language is familiar. It also feels efficient enough, at least until agents start talking to other agents. ...

January 31, 2026 · 15 min · Zelina
Cover image

When Alignment Is Not Enough: Reading Between the Lines of Modern LLM Safety

A chatbot refuses a dangerous request. Everyone relaxes. This is the small theatre of modern AI safety: the model says no, the dashboard records a refusal, the vendor presentation adds another green checkmark, and the compliance team moves on to the next risk register. Very tidy. Very comforting. Also, increasingly insufficient. The problem is not that refusal behavior is meaningless. It is not. The problem is that refusal behavior is only one visible symptom of safety alignment. Modern LLM safety now depends on a larger chain: training objectives, post-training choices, inference interfaces, prompt formats, tool access, evaluation design, and deployment context. When any part of that chain changes, the nice refusal seen in a benchmark may not survive contact with the product. ...

January 26, 2026 · 15 min · Zelina
Cover image

Triage by Token: When Context Clues Quietly Override Clinical Judgment

A patient walks into an emergency department. Or arrives by ambulance. Or lives far from the hospital. Or has private insurance. Or has missed prior appointments. Clinically, those details may be background noise. In triage, the core question is supposed to be sharper: how sick is this patient, how urgent is the risk, and what resources are likely needed? The Emergency Severity Index, or ESI, is not a lifestyle quiz with a stethoscope attached. ...

January 24, 2026 · 13 min · Zelina
Cover image

Prompt Wars: When Pedagogy Beats Cleverness

A prompt review meeting usually sounds more scientific than it is. One person likes the “coach” version. Another prefers the “Socratic” version because it sounds more educational. Someone says the prompt should mention metacognition. Someone else adds “be concise,” because apparently every prompt eventually becomes a corporate email with anxiety issues. Then the team ships the one that feels best. ...

January 23, 2026 · 15 min · Zelina
Cover image

Auditing the Illusion of Forgetting: When Unlearning Isn’t Enough

Deletion requests sound simple until the model answers politely. A user asks for data to be removed. A publisher demands that copyrighted passages stop being reproduced. A compliance team wants evidence that a fine-tuned model no longer carries traces of a forbidden dataset. The model is run through an unlearning method, the surface tests improve, the dashboard turns less red, and everyone enjoys the brief spiritual comfort of a green checkmark. ...

January 22, 2026 · 17 min · Zelina
Cover image

Pay to Think: Incentive Design Is the Hidden Variable in Human–AI Research

Payment sounds like the boring part of a user study. Recruit participants. Estimate task time. Set a base rate. Add a small bonus if the budget allows. Put the number in the methods section, preferably somewhere readers can skim past with dignity. Then move on to the interesting material: trust, reliance, explanations, fairness, error rates, cognitive load, and all the other variables that make human–AI decision-making sound like a serious field rather than a procurement spreadsheet. ...

January 22, 2026 · 18 min · Zelina
Cover image

Rebuttal Agents, Not Rebuttal Text: Why ‘Verify‑Then‑Write’ Is the Only Scalable Future

Rebuttal is where polite language goes to be cross-examined. A reviewer asks why the baseline is missing. Another says the theory is unclear. A third implies that the claimed novelty is, shall we say, generously interpreted. The authors have a few days to respond, and every sentence must do three jobs at once: answer the concern, avoid overclaiming, and preserve the paper’s strategic position. ...

January 21, 2026 · 16 min · Zelina
Cover image

When Benchmarks Break: Why Bigger Models Keep Winning (and What That Costs You)

Budget. That is where the benchmark story usually becomes less elegant. A vendor shows a model card with better reasoning scores, stronger multi-task accuracy, and a leaderboard position polished to a mirror finish. Then someone in operations asks the rude question: what does this improvement cost per customer case, per analyst hour, per compliance review, or per failed escalation? ...

January 21, 2026 · 12 min · Zelina
Cover image

Who’s Really in Charge? Epistemic Control After the Age of the Black Box

Control is a comforting word. It suggests a hand on the wheel, a dashboard of indicators, and a human being somewhere nearby who can still say no. Machine learning makes that picture look increasingly theatrical. In AI-assisted science, researchers often do not know exactly which internal representations a model has learned, why a high-dimensional classifier separates one tumor subtype from another, or whether a model’s “useful pattern” corresponds to anything a scientist would recognize as a meaningful mechanism. The black box does not merely sit inside the laboratory. It starts to participate in deciding what the laboratory can see. ...

January 20, 2026 · 15 min · Zelina