GPT-6 Astra cache diagnostics identify a cause, while usage records show reuse
OpenAI's cache comparison can reveal changed tools or settings, but its diagnostic label is separate from the token usage that determines input charges.

GPT-6 Astra developers investigating a sudden sharp drop in cached input can compare a request with an earlier completed response. OpenAI's diagnostic guide separates that comparison from the real token reuse reported for billing.
The Responses API accepts a baseline response ID in its cache options. This asks for diagnostics; it neither loads the earlier conversation nor changes cache behaviour. A changed tool name, tool order or processing tier can prevent the expected prefix from matching.
A reported cache hit does not mean every input token was reused. New material can still require processing. Developers must check the response's usage fields rather than treating the diagnostic label as a total cost.
The service reports only the first classified cause so fixing one difference may reveal another. Diagnostics add no separate fee, but extra test requests are billed normally. The guidance offers a troubleshooting method, not evidence that a particular application's bill has fallen.
프리즘코리아 편집국 > Daniel Reed



