Eleven Cache Reads for Every Fresh Token
Where the cache markers go in a four-step pipeline and what 128 edits cost.
Prompt caching is one of the things that made the rebuild worth it. I pick what gets cached, and the logs tell me if it actually hit.
Each rule takes four LLM calls, more when a rule needs a retry. Every one of them opens with the same system preamble of about 6KB that never changes. It covers the rule format, the X12 conventions, and the output schema. A payer guide can have a few thousand edits, so paying full price for the preamble on every call gets expensive fast.
Anthropic's prompt cache solves that as long as the cached text stays exactly the same. You mark a block and then you have to send that block byte-for-byte. One character difference and the cache misses and you pay to write it again.
With the raw SDK I build the message array so I control the marker:
def make_cached_text_block(text: str) -> dict:
if os.getenv("LLM_PROVIDER", "anthropic").lower() == "anthropic":
return {"type": "text", "text": text,
"cache_control": {"type": "ephemeral"}}
return {"type": "text", "text": text} # OpenAI: no marker
The pipeline builds these blocks once at startup and reuses the same objects so the preamble is written to the cache once and read after that. A cache entry dies after five minutes with no read. A steady batch keeps it warm, and a long pause only costs one extra write.
The preamble is not even the big block. Each step's instructions get a marker too. Step 3 is the one that writes the rule logic, and once the function catalog and the loop table are in there it lands around 40KB. On a 128-edit run, step 3 alone read 1.93 million of the 3.08 million cached tokens.
The LangChain version already had this builder and it logged cache tokens. But LangChain built the final request, so its code decided whether my cached text stayed stable. Now a small client I wrote passes the blocks to the SDK unchanged.
Every call logs cache reads and cache writes separately. If the cached text changes, I can usually see the misses in the logs before I notice the extra cost.
Here is a real run: 128 edits from the WPS Medicare 837P spreadsheet, on Claude Sonnet 4.6. The four steps made 526 calls, which read 3,084,872 tokens from the cache and wrote only 18,991, against 274,708 tokens of fresh input. That is about eleven cached tokens read for every fresh one. The run cost $4.09, a little over three cents an edit. The same tokens with no cache would have been $12.40.
The implementation itself is pretty small: one cache_control key and making sure not to let the text drift between requests. This job runs the same four prompts over every edit in a guide, so the cache only works if I control the request.