ShadowReport
What the cache would have done for one prompt, at each threshold it was asked about.
ThresholdCalibrator answers "which threshold is right?" but needs a labelled set, so the honest answer to a team standing at the door is "measure it, on data you first have to build". That first step, not the cache, is what stops a semantic cache going in front of production traffic.
Shadow mode removes it. The cache runs the whole lookup against real traffic, serves nothing, and reports the decision at every threshold in one pass, so the output is the team's own precision and recall curve against their own questions rather than somebody else's corpus.