Storage & ML systems
I measure how systems behave, then make the common path simpler, faster, and more efficient.
Latest post07.23.26
Your LRU cache is working too hard.
The archive · 2026
- Your LRU cache is working too hard Across 6,357 production traces, most LRU promotions did little useful work. Delaying the decision made caches faster without making them less effective.
- Authorship inflation at OSDI: counting papers two ways OSDI author lists have more than doubled since 1994. Whether the field's most prolific authors are actually publishing more depends on whether a paper counts as one, or as one divided by its author count.
- Same open weights, wildly different prompt caches The same open-weight model, served by 35 different backends, held a cached prompt prefix anywhere from under 30 seconds to more than a day. On open weights, your cache TTL is set by who serves the model, not by the model.
- Proprietary prompt caches expire in minutes, not hours OpenAI, Gemini, Anthropic, and xAI all cache your prompt prefix — but the probe that caught open-weight backends holding prefixes for hours catches the big proprietary APIs letting go in 5 to 30 minutes.
- Make model precision elastic MorphServe temporarily converts selected model layers into KV-cache headroom when traffic spikes, then restores full precision when the pressure passes.
- Why FIFO is (almost) all you need for cache eviction LRU tries to rank everything. Production traces showed that caches need something simpler: reject one-hit objects quickly and delay promotion until it matters.
- What I look for in systems research students I am looking for curious, persistent students who want to build real systems. Here is what the work requires and how to write an email I will read carefully.
Currently at Harvard
Start a conversation