Cover image

Fast Without False Precision: Foundation Models for Partial Causal Identification

TL;DR for operators Bellot and Dhir’s Foundation Models for Partial Causal Identification1 targets a specific failure mode in automated causal analysis: observational data may narrow a causal answer without determining one unique value. The proposed model is trained once to return a distribution over still-compatible causal or counterfactual answers, rather than solving a new bound-optimization problem for every dataset and query. ...

September 24, 2026 · 6 min · Zelina
Cover image

When AUROC Agrees and the Decision Still Changes

TL;DR for operators CRS-Bench1 tests a decision that clean-test AUROC does not fully answer: which pretrained medical image encoder deserves the next round of engineering, labeling, adaptation, and validation investment. Across 15 encoder families, AUROC and the benchmark’s broader reliability score are positively associated, yet 21 of 105 pairwise model choices reverse. The disagreement is not evidence that AUROC is useless. It shows that discrimination can preserve the broad ordering while missing differences in calibration, label efficiency, and robustness that change actual selection decisions. ...

September 22, 2026 · 7 min · Zelina
Cover image

Aviation AI Needs a Shared Operational State Before It Needs a Bigger Model

TL;DR for operators AviationLMM1 is best read as a blueprint for how aviation AI systems might stop treating radio, surveillance tracks, telemetry, video, operational text, and sensor feeds as separate evidence streams. The paper’s central claim is architectural: before a system can reason across those inputs, it must encode each modality appropriately, align them across time, space, meaning, and reliability, fuse them into a coherent operational state, and only then generate task-specific outputs. ...

September 13, 2026 · 7 min · Zelina
Cover image

Grounding Is a Responsibility, Not a Benchmark Score

TL;DR for operators A robotics system can generate a readable plan, revise that plan after an error, and improve its overall task-completion rate without demonstrating that its language component is correctly grounded in the physical environment. That attribution problem is the focus of a review by Yifan Guo and colleagues.1 The authors audit 105 foundation-model-enabled embodied-agent papers by separating two questions: what responsibility does language carry inside the system, and what evidence actually tests that responsibility? ...

August 31, 2026 · 7 min · Zelina
Cover image

Perceive Once, Decide Per Query: CogVis Splits Change Detection by Decision Scope

TL;DR for operators When the same before-and-after imagery must answer several semantic questions, rerunning the full visual analysis for every query wastes computation that does not actually depend on the query. CogVis separates those decisions by scope. It computes category-independent temporal change evidence once, then repeats only the semantic calibration and candidate verification that must depend on the requested category. The paper reports 28.50% higher inference throughput than the next-fastest compared method; with 10 queries, CogVis takes 10.03 seconds versus 12.56–956.60 seconds for the category-wise baselines evaluated. ...

August 23, 2026 · 8 min · Zelina
Cover image

Context Is Not Free, So Stop Feeding the Whole Table

TL;DR for operators Many tabular foundation models behave like very competent consultants with a mildly expensive habit: they want the entire labelled training set placed in front of them at inference time. That works neatly on small datasets. It becomes rather less charming when the table grows to tens or hundreds of thousands of rows and the model’s attention cost starts behaving like it has discovered compound interest. ...

June 24, 2026 · 24 min · Zelina
Cover image

LoRA’s Rank Excuse Has a Gradient Problem

TL;DR for operators LoRA is usually sold as a rank-and-cost compromise: train a small low-rank adapter instead of updating the whole model, accept some performance gap, and enjoy the budget meeting. The paper behind SDS-LoRA argues that this explanation is incomplete. The gap is not only because the adapter is low-rank. It is also because standard LoRA can distort the training signal that flows into that adapter.1 ...

June 21, 2026 · 21 min · Zelina
Cover image

Less Label, More Light: What a 3D Microscopy Foundation Model Actually Buys

Microscopy has a labor problem. Not the photogenic kind where a scientist leans into a glowing instrument and discovers the secret architecture of life before lunch. The duller problem is that modern light sheet fluorescence microscopy can produce rich three-dimensional volumes faster than expert teams can label them. Segmentation requires voxel-level masks. Stain classification requires domain knowledge. Restoration needs paired degraded and high-quality images, which nature, unhelpfully, does not always provide in tidy folders. ...

June 5, 2026 · 16 min · Zelina
Cover image

One Pass to Forecast Them All: Toto 2.0 and the Scaling Recipe for Time-Series AI

Forecasting is where machine learning often learns humility. A language model can sound clever while being wrong. A forecasting model has fewer hiding places. Revenue arrives or it does not. CPU saturation happens or it does not. Demand spikes, latency drifts, inventories rot, turbines fail, and the spreadsheet smiles politely before punishing everyone involved. This is why time-series foundation models have been treated with a particular kind of suspicion: useful, interesting, sometimes impressive, but not yet comfortably scalable in the way large language models became scalable. ...

June 5, 2026 · 18 min · Zelina
Cover image

Heart of Scale: Why Bigger ECG Models Don’t Always Beat Better Biases

Heart of Scale: Why Bigger ECG Models Don’t Always Beat Better Biases A hospital does not buy an ECG model because it enjoys leaderboard furniture. It buys one because somebody wants a cheap, reliable signal from a noisy waveform: rhythm abnormality, structural heart disease, ICU risk, mortality risk, maybe a demographic or physiological clue that was not explicitly labeled during pre-training. ...

June 1, 2026 · 19 min · Zelina