Which cache answered, and what did it cost?
This replays 201 real requests from a measured run of a RAG gateway on Claude Sonnet 5.5 (first of 3 repetitions). Every request was sent to all arms in a random order, so the arms are comparable. Nothing here is simulated: the latency, cost and answers are the logged values. Code and logs.
Replay
Cache setup:
Model call
Exact cache
Semantic cache
Wrong answer served from cache
Last requests
Live: ask the running gateway
This page is being served by the gateway itself, so this box sends real requests (they use the API budget). The key is only kept in this tab.
Results for the whole run (median of 3 repetitions)
Same 201 requests per repetition. "none-twin" is an identical copy of "none": the gap between them is the noise floor, so differences smaller than that mean nothing.