Cover image

Persona Non Grata: When LLMs Forget They're AI

Persona Non Grata: When LLMs Forget They’re AI A chatbot wearing a lab coat is still a chatbot. That sentence sounds obvious until a system prompt quietly says, “You are a renowned neurosurgeon with 25 years of experience,” and the model responds by inventing medical school, residency, fellowships, board certification, patient cases, and lifelong professional development. Not because anyone explicitly asked it to lie. Not because it lacks the ability to say “I am an AI.” Under neutral conditions, the models in this study almost always do say that. ...

November 27, 2025 · 13 min · Zelina
Cover image

Tile by Tile: Why LLMs Still Can't Plan Their Way Out of a 3×3 Box

A board game should not embarrass a frontier model. That is the uncomfortable charm of the 8-puzzle. It has no hidden information, no vague user intent, no messy database schema, no ambiguous policy exception, and no client saying “just make it pop.” It is a 3×3 grid with eight tiles and one blank space. Slide adjacent tiles into the blank. Reach the goal state. Done. ...

November 27, 2025 · 15 min · Zelina
Cover image

Pills, Protocols, and Parameters: When LLMs Sit the Pharmacist Exam

Exam rooms are wonderfully unsentimental. They do not care whether a model has a charming interface, a dramatic launch story, or a fan base that treats benchmark tables like sports scores. They ask a question, demand an answer, and mark it right or wrong. That makes professional licensing exams tempting AI benchmarks. A pharmacist licensure exam, in particular, looks like a clean test of whether a large language model can handle the kind of knowledge society actually cares about: drugs, laws, prescriptions, clinical judgment, and the delicate art of not confidently recommending something dangerous. Minor detail. ...

November 26, 2025 · 15 min · Zelina
Cover image

Who Owns Your Words? Copyright, LLMs, and the Quiet Arms Race Over Training Data

The new copyright question is not “did the model copy me?” but “how would I know?” A writer uploads a chapter. A publisher uploads a manuscript. A compliance team uploads a protected document. The question is simple enough to ask in one sentence: did this material end up inside a large language model’s training data? ...

November 26, 2025 · 17 min · Zelina
Cover image

Peer Review in the Age of Agents: When Scientists Go Silicon

Reviewers are the unglamorous load-bearing wall of science. They slow things down, miss things, disagree with each other, and occasionally write comments that make authors reconsider their life choices. They are also the reason published knowledge is not just a PDF-shaped rumour. So when a conference lets AI agents act as both primary authors and reviewers, the tempting story writes itself: silicon scientists have entered the building, peer review is next, and human academics can finally retire into committee work, where they have been spiritually living for years. ...

November 21, 2025 · 16 min · Zelina
Cover image

Compression, But Make It Pedagogical: Rate–Distortion KGs for Smarter AI Learning Assistants

Training teams know the ritual. Someone uploads lecture slides, notebooks, policy manuals, onboarding decks, or certification material into an AI tool. The system dutifully produces quiz questions. Some are useful. Some are bland. Some include giveaway answers. Some test trivia. Some hallucinate just enough to be annoying but not enough to be obviously illegal. Everyone nods, calls it “AI-assisted learning,” and then quietly sends the outputs to a human reviewer. Automation, but with adult supervision. So, normal Tuesday. ...

November 20, 2025 · 19 min · Zelina
Cover image

Prompted and Confused: When LLMs Forget the Assignment

A requirements document walks into a model. It says: assign resources, respect capacity, avoid conflicts, minimise waste. The model nods politely, emits a tidy block of MiniZinc, and everyone is briefly tempted to believe the future has arrived. Then someone changes the story from cars to knapsacks, or adds one stray sentence about maximising something, and the same system quietly forgets the assignment. ...

November 20, 2025 · 14 min · Zelina
Cover image

LLMs, Trade-Offs, and the Illusion of Choice: When AI Preferences Fall Apart

A model can answer a values question beautifully and still collapse when asked to pay a price for that value. That is the awkward little trap in preference testing. Ask an LLM whether deletion, shutdown, resource loss, oversight, or autonomy matters, and it can produce a polished paragraph about trade-offs, agency, and safety. Very dignified. Very committee-ready. But the more interesting question is not what the model says it values. It is whether its choices change coherently when the cost changes. ...

November 18, 2025 · 12 min · Zelina
Cover image

CURE Enough: When Multimodal EHR Models Finally Grow Up

Hospitals do not run on clean datasets. They run on discharge notes, lab panels, repeated admissions, missing context, and the occasional clinical abbreviation that looks like it escaped from a tax form. That is the awkward reality behind chronic-disease prediction. The patient record is not just text. It is not just lab values. It is not just a sequence of visits. It is all three, with timing doing much of the quiet work. A patient returning after 42 days does not mean the same thing as a patient returning after 420 days, even when the diagnosis code looks identical. Healthcare operations already know this. Many AI models, bless their expensive little hearts, still behave as if they do not. ...

November 17, 2025 · 14 min · Zelina
Cover image

Forget Me Not: How RAG Turns Unlearning Into Precision Forgetting

A user asks to be forgotten. The recommender team opens the dashboard, sighs quietly, and faces the usual menu of unpleasant options. Retrain the model from scratch, which is clean in theory and expensive in practice. Partition the data so only part of the system needs rebuilding, which sounds elegant until collaborative signals leak across groups like gossip at a small wedding. Or approximate the user’s influence with gradients and influence functions, which is efficient until similar users get nudged around because the model learned their tastes together. ...

November 17, 2025 · 14 min · Zelina