Cover image

Query Filters Catch the Attack That Advertises Itself

TL;DR for operators A retrieval-security team can reasonably start by looking for poisoned documents that repeat a user’s question or arrive as an unusually tight cluster of near-duplicate vectors. CamoDocs shows why neither signal should be treated as a durable property of poisoning. In matched-budget experiments, a simple query detector keeps three earlier attacks between 5.4% and 7.6% attack success, while CamoDocs reaches 75.2%. Put the target query back into CamoDocs, and success falls to 7.8%. The paper’s ablations then isolate the harder problem: deliberately spreading poisoned-document embeddings raises attack success against the clustering-based TrustRAG defense from 11.5% to 28.7%. ...

September 25, 2026 · 7 min · Zelina
Cover image

The Multilingual Tax Is Logarithmic: What Actually Breaks as Language Coverage Grows

TL;DR for operators Adding languages to an embedding model carries a theoretical representation cost, but the paper argues that the cost grows only logarithmically with language count. Under its formal conditions, the minimum dimensionality required to preserve useful within-language semantic structure, align translations across languages, and keep language variants distinguishable is $$ D_{X^{L}}=\Theta(\log L). $$ That changes how multilingual quality loss should be diagnosed. A sharp decline after expanding language coverage is not, by itself, evidence that embedding dimensionality has hit an unavoidable capacity wall. ...

September 24, 2026 · 8 min · Zelina
Cover image

Black Box Is Too Blunt: What AI Interfaces Reveal to Attackers

TL;DR for operators What exactly should a deployed AI system reveal to users, vendors, insiders, or connected applications? Hiding model architecture and weights does not make every interface equally opaque. Mahbub and colleagues separate access into six operational categories—None, Metadata, Decision-Only, Score/Rank, Embedding, and White-Box—because each exposes a different signal to an adversary.1 A binary decision permits probing, while numerical confidence or similarity values provide directional feedback; internal feature representations expose still richer information. The paper’s synthesis suggests that richer signals generally reduce attacker uncertainty and query burden while enabling additional risks such as model extraction and biometric template inversion. ...

August 15, 2026 · 7 min · Zelina
Cover image

When AI Can Solve But Can't Search: The MathNet Equation

Search. That is the unglamorous part of AI work. The demo asks a model to solve a clean problem. The enterprise system asks a model to find the right prior case, retrieve the relevant precedent, avoid the misleading near-match, and then adapt the answer without making a confident mess of it. MathNet is interesting because it puts that distinction under pressure. The paper introduces a large multilingual, multimodal Olympiad mathematics benchmark, but the more useful business lesson is not merely that frontier models can solve hard math. We already have enough leaderboards wearing medals. The sharper finding is that models and embedding systems can still fail at recognizing when two problems are mathematically the same, or when one problem is structurally useful for another.1 ...

April 23, 2026 · 13 min · Zelina
Cover image

Mind the Units: Why LLMs Still Can't Count (And How CONE Fixes It)

Numbers look harmless until they enter a business database. A revenue field says 50. A dosage field says 50. An age field says 50. A follow-up period says 50. A unit may be present, missing, abbreviated, buried in the column header, or inconsistently written as ml, mL, or something the spreadsheet inherited from a PDF extraction pipeline during its villain era. ...

March 8, 2026 · 14 min · Zelina
Cover image

Unsupervised, Unaware, Unfair: When Your Embedding Knows Too Much

Segmentation is where many businesses go to feel mathematically innocent. No target label. No credit decision. No hiring decision. No explicit age column. Just customers grouped by behavior, employees mapped by survey responses, users visualized in an embedding dashboard, or applicants compressed into a neat latent space before the “real” model begins. ...

February 23, 2026 · 14 min · Zelina
Cover image

Ultra‑Sparse Embeddings Without Apology

Search gets expensive quietly. At small scale, an embedding is just a vector. At product scale, it becomes rent: storage rent, memory rent, GPU rent, latency rent, and the recurring emotional tax of explaining why a semantic search feature needs yet another infrastructure budget. Dense embeddings made this bargain feel natural. More dimensions, more semantic capacity. More semantic capacity, better retrieval. Better retrieval, more invoices. Elegant, if one enjoys expensive inevitability. ...

February 8, 2026 · 19 min · Zelina
Cover image

Beyond Cosine: When Order Beats Angle in Embedding Similarity

Search has a small ritual. Take two embeddings, compute cosine similarity, rank the results, and move on. The ritual is fast, familiar, and usually good enough. It is also so deeply embedded in AI infrastructure that many teams treat it less like a modeling choice and more like plumbing. That is convenient. It is not always innocent. ...

February 7, 2026 · 14 min · Zelina
Cover image

Lost Without a Map: Why Intelligence Is Really About Navigation

Lost Without a Map: Why Intelligence Is Really About Navigation Map. That is the word most AI product teams should probably put above their dashboards, agent logs, evaluation suites, and occasionally their office coffee machine. Not because maps are poetic. Because when an AI system fails in a live workflow, the failure often does not look like “the model forgot a fact.” It looks like the system was navigating the wrong space. ...

January 21, 2026 · 18 min · Zelina
Cover image

Words + Returns: Teaching Embeddings to Invest in Themes

TL;DR for operators The paper behind THEME is not really about asking an LLM to “find AI stocks” and hoping it returns a genius portfolio, because that would be the usual theatre with a Bloomberg terminal costume.1 It is about building a retrieval layer that understands investment themes as a special kind of search problem: cross-sector, text-heavy, time-sensitive, and annoyingly allergic to static classification. ...

August 26, 2025 · 16 min · Zelina