I added prompt caching to my Anthropic Batch API workflow. The hit rate was 0%.Each model has a minimum cacheable token count — 4,096 for Haiku 4.5. If your cache_control block is below that, the API silently ignores it. Successful response, zero cache reads, no warning.My IAB taxonomy prompt was 1,064 tokens. Well under the threshold.Full write-up:https://mikenoe.com/posts/prompt-caching-classivore/#AnthropicAPI #LLM #PromptCaching #AIEngineering
Related
Enterprise AI Agents in Production: Build and Secure Them by Thomas De Vos is the featured bundle of ebooks 📚 on Leanpub...
Enterprise AI Agents in Production: Build and Secure Them by Thomas De Vos is the featured bundle of ebooks 📚 on Leanpub!Two practical books for teams building enterprise AI agents...
Theo Conjecture solves 35-year-old math problem, finds a term no one predictedhttps://firstprinciples.com/blog-article/a...
Theo Conjecture solves 35-year-old math problem, finds a term no one predictedhttps://firstprinciples.com/blog-article/ai-system-theo-conjecture-solves-35-year-old-math-conjecture#...
EVA Destiny Is Not Here to Compete With Anyone https://www.youtube.com/watch?v=8QN1t5tpCQsSince releasing EVA Destiny Sa...
EVA Destiny Is Not Here to Compete With Anyone https://www.youtube.com/watch?v=8QN1t5tpCQsSince releasing EVA Destiny Sapphira as free software, we have noticed a small but recurri...