AI Dictionary › AI Fundamentals
Context caching (cache del contesto)
Context caching is a technique that stores the processing of a frequently reused portion of text, so the model need not reprocess it from scratch on every request. Many applications always send the same long preamble: system instructions, a manual, a fixed context. Instead of making the model "read" this block anew each time, the system keeps its already-digested version and reuses it. It is like a cook who prepares a base stock in advance and keeps it ready: when an order arrives, they start from there rather than from zero, saving time. Only the variable part of the request is processed as new.
Context caching matters because it markedly cuts cost and latency in applications reusing large fixed contexts. The same instructions repeated across thousands of calls become much cheaper and faster, making practical assistants and agents that lean on extensive documents or system prompts.
From our network
INDACO TMS: Transport Management for European Logistics
Shipment tracking, multi-carrier EDI and automated invoicing in one cloud platform. Invoices generated in under 10 seconds.
Visit indacotms.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.