Cover image

ID Crisis, Resolved: When Semantic IDs Stop Fighting Hash IDs

Catalogs have a boring problem. Most items are nearly invisible. A platform may have millions of products, posts, videos, restaurants, songs, or ads, but user interaction is never evenly distributed. A small number of head items collect enough clicks, saves, purchases, and dwell time to become statistically legible. The rest live in the long tail, where the system is expected to recommend them intelligently despite barely having seen them. Very democratic. Very inconvenient. ...

December 14, 2025 · 16 min · Zelina
Cover image

Making Noise Make Sense: How FANoise Sharpens Multimodal Representations

Search systems fail in boring ways before they fail in spectacular ones. A customer uploads a product photo and receives visually similar items that miss the actual intent. A compliance analyst searches a scanned document and gets pages that look close but answer the wrong question. A visual QA system finds the right region but ranks the wrong evidence first. Nobody in the meeting says, “Ah yes, our embedding space has poor spectral noise allocation.” They say the search feels unreliable. Much more executive-friendly. Much less useful. ...

November 30, 2025 · 13 min · Zelina
Cover image

What We Don’t C: Why Latent Space Blind Spots Matter More Than Ever

A dataset rarely hides everything equally. In most organisations, the visible structure is already over-managed. Product images are labelled by category. Medical scans are labelled by diagnosis. Satellite imagery is indexed by region and timestamp. Customer records are sliced into the usual demographic trays. Scientific images come with whatever measurements the field has already agreed are worth writing down. ...

November 13, 2025 · 16 min · Zelina