AI Dictionary › Prompting

Prompt Caching

Prompt caching is a feature offered by several APIs that stores the initial, stable portion of a prompt, system instructions, reference documents, examples, so subsequent calls do not reprocess it from scratch, drastically cutting cost and latency. Example: a support assistant has 5,000 tokens of instructions and knowledge base identical on every request; with caching, that part is processed once and reused, and you pay a reduced rate only for the variable part (the user's question). Savings can exceed 90% on cached tokens.

Definition

Use it when many calls share a long, unchanging prefix: chatbots with long instructions, repeated document processing, few-shot with many fixed examples. Structure the prompt with the stable part first and the variable part last to maximise the cacheable portion.

Related terms

More in Prompting

Put it into practice

From our network

Kaimaki Web: Websites That Win Customers

Custom websites, web apps and digital marketing for growing businesses.

Visit kaimakiweb.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.