Cover image

One Code, Three Budgets

TL;DR for operators A semantic-search service may want different retrieval budgets for different workloads: a cheap first-stage filter, a moderate-cost interactive search path, and a higher-quality path when more compute is available. Maintaining a separately encoded corpus for each operating point adds storage, indexing, and deployment complexity. Matryoshka Hash Representations (MHR) address that problem with one nested binary document code whose 8-, 16-, and 32-byte prefixes are independently searchable.1 The crucial detail is that these prefixes are not obtained by simply chopping bits off a conventionally trained 256-bit hash. The authors first train the 256-bit representation on MS MARCO, freeze it, and then train residual adaptors to reorganize the prefixes while preserving the full-width solution. ...

October 11, 2026 · 7 min · Zelina