AI Dictionary › Prompting
Prompt caching is a feature offered by several APIs that stores the initial, stable portion of a prompt, system instructions, reference documents, examples, so subsequent calls do not reprocess it from scratch, drastically cutting cost and latency. Example: a support assistant has 5,000 tokens of instructions and knowledge base identical on every request; with caching, that part is processed once and reused, and you pay a reduced rate only for the variable part (the user's question). Savings can exceed 90% on cached tokens.
Use it when many calls share a long, unchanging prefix: chatbots with long instructions, repeated document processing, few-shot with many fixed examples. Structure the prompt with the stable part first and the variable part last to maximise the cacheable portion.
From our network
Kaimaki Web: Websites That Win Customers
Custom websites, web apps and digital marketing for growing businesses.
Visit kaimakiweb.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.