Replay of logged runs

Which cache answered, and what did it cost?

This replays 201 real requests from a measured run of a RAG gateway on Claude Sonnet 5.5 (first of 3 repetitions). Every request was sent to all arms in a random order, so the arms are comparable. Nothing here is simulated: the latency, cost and answers are the logged values. Code and logs.

Replay

Cache setup:

Model call Exact cache Semantic cache Wrong answer served from cache
Last requests

    Live: ask the running gateway

    This page is being served by the gateway itself, so this box sends real requests (they use the API budget). The key is only kept in this tab.

    Results for the whole run (median of 3 repetitions)

    Same 201 requests per repetition. "none-twin" is an identical copy of "none": the gap between them is the noise floor, so differences smaller than that mean nothing.