Cover image

When the Brain Becomes the Dataset: Teaching AI to Hear Music Like Humans

Music is an unusually good test for artificial intelligence because it punishes lazy definitions of “understanding.” A model can identify notes. It can classify genre. It can predict the next audio token with impressive fluency. None of that means it hears music the way a person does. Human listeners do not merely receive sound. They anticipate, mispredict, adjust, and continue listening. The brain is not a passive microphone with better branding. ...

March 4, 2026 · 13 min · Zelina
Cover image

Code-SHARP: When Agents Start Writing Their Own Ambitions

Automation has a boring failure mode: the moment the world becomes slightly more complicated than the workflow diagram, the system starts asking for a human. That is not because the model lacks vocabulary. It is because the automation system does not know how to grow its own capabilities. Most AI agents are still built around a fixed menu of actions, fixed task definitions, and fixed reward signals. They can optimize, but they rarely expand the set of things they know how to optimize for. Very impressive, in the way a microwave is impressive until you ask it to cook without buttons. ...

February 11, 2026 · 19 min · Zelina
Cover image

The Patient Is Not a Moving Document: Why Clinical AI Needs World Models

A patient chart looks like a document because hospitals make it look that way. There are notes, medication lists, lab panels, procedure codes, imaging references, adverse events, survival outcomes, and enough timestamps to make a database administrator feel briefly useful. So it is tempting to treat the electronic health record as a very long piece of text: serialize the events, train a model to predict the next token, extract an embedding, and hope that clinical meaning emerges somewhere inside the transformer fog. ...

January 30, 2026 · 14 min · Zelina
Cover image

Gen Z, But Make It Statistical: Teaching LLMs to Listen to Data

A pricing team gives an LLM several hundred property listings and asks a sensible question: Which characteristics help predict the selling price? The model returns an equally sensible list. Swimming pools. Granite countertops. Scenic views. Green lawns. Kitchen islands. Everything sounds plausible. That is the problem. The list describes what generally makes a house attractive. It does not necessarily describe what separated expensive from inexpensive houses in this particular collection, sold in particular locations, during a particular year. The LLM has supplied real-estate conventional wisdom when the business needed dataset-specific evidence. ...

January 1, 2026 · 17 min · Zelina
Cover image

It Takes a Village (of Models): Why Multi-Agent Intelligence Won't Emerge by Accident

Agents are easy to multiply. That is the attractive part. Give one model a browser. Give another a code editor. Add a planner, a critic, a memory layer, a few tools, a dashboard, and suddenly the product demo looks like a small digital office. Everyone has a job title. Everyone talks. Nobody asks whether the “team” actually knows how to be a team. ...

December 10, 2025 · 14 min · Zelina
Cover image

Mutation Impossible? How Multimodal Agents Are Rewriting Glioma Diagnostics

Report First, Diagnosis Second A medical report usually arrives after the diagnostic work is done. It explains, records, justifies, and sometimes politely hides how messy the evidence really was. This paper asks a more interesting question: what if the report itself becomes a predictive object? In Multimodal Oncology Agent for IDH1 Mutation Prediction in Low-Grade Glioma, Hafsa Akebli and colleagues build a Multimodal Oncology Agent, or MOA, for predicting IDH1 mutation status in low-grade glioma using TCGA-LGG data, whole-slide histology, structured clinical variables, genomic context, and external biomedical knowledge sources.1 The immediate headline is easy enough: the full multimodal setup reaches the best reported performance, with an F1-score of 0.912. ...

December 8, 2025 · 15 min · Zelina
Cover image

Maps, Models, and Mobility: GPT Goes for a Walk

The delivery route is not a sentence A delivery van does not move like a sentence. It stops. It waits. It turns left because a road exists, not because grammar allows it. Its next point depends on geography, time of day, congestion, driver behavior, business constraints, and occasionally the small civic miracle of a loading bay being available. A language model sees the world as tokens arranged in sequence. A trajectory model sees movement as a sequence too, but the symbols are less polite: latitude, longitude, timestamp, region, point of interest, dwell time, elapsed time, and missing segments. ...

November 26, 2025 · 17 min · Zelina
Cover image

Map Before You Train: Data Cartography to Defuse LLM Memorization

TL;DR for operators Training data does not become risky only after a model has memorised it. It often leaves signals while training is still happening. That is the useful idea behind Generative Data Cartography, or GenDataCarto: track how each pretraining sample behaves during early training, then use that behaviour to decide which data should be kept, up-sampled, down-weighted, or removed.1 The method uses two signals. The first is early loss, which approximates how difficult a sample is. The second is the frequency of “forget events”, where a sample appears learned and later becomes poorly fitted again. In the paper’s framing, frequent forget events are not just training noise. They are a warning that a sample may be unusually influential, repeatedly re-entering the model’s attention like a guest who refuses to leave the meeting. ...

September 4, 2025 · 16 min · Zelina
Cover image

MoE Money, MoE Problems? FinCast Bets Big on Foundation Models for Markets

TL;DR for operators FinCast is a finance-specific time-series foundation model that tries to do for market forecasting what large pretrained models did for language: absorb enough diverse data that new tasks require less bespoke engineering.1 The paper reports strong evidence on forecasting accuracy. In a zero-shot benchmark of 3,632 financial time series and more than 4.38 million scalar time points, FinCast beats general-purpose time-series foundation models on average, with roughly 20% lower MSE and 10% lower MAE. In supervised stock benchmarks, even the zero-shot version beats the listed supervised baselines; lightweight fine-tuning improves the gap further. ...

August 30, 2025 · 16 min · Zelina
Cover image

The Invisible Hand in the Machine: Rethinking AI Through a Collectivist Lens

TL;DR for operators Users do not experience an AI product as a theorem. They experience it as a bargain. They give data, attention, labour, trust, prompts, feedback, documents, creative work, behavioural traces, and sometimes money. In return, they expect useful output, lower friction, safer decisions, visibility, compensation, privacy, or at least not being quietly turned into unpaid infrastructure. The bargain may be explicit. More often, because apparently we enjoy building planetary-scale systems on implied consent and vibes, it is not. ...

July 10, 2025 · 17 min · Zelina