AI Dictionary › Prompting
The max tokens parameter caps how many tokens the model can generate in a single response. It is a technical ceiling that guards against overly long output, controls cost and ensures predictable latency. Example: you set max_tokens = 150 for a short summary task; the model stops when it hits the limit, even mid-sentence if it has not finished. Note: max tokens does NOT instruct the model to be concise, it only truncates. For genuinely short answers you must also say so in the prompt ("summarise in at most 3 sentences").
Use it to make an application safe: avoid oversized responses, cap per-call spend, respect interface limits. Remember the limit adds to input tokens within the context window: a very long input leaves less room for output.
From our network
INDACO TMS: Transport Management for European Logistics
Shipment tracking, multi-carrier EDI and automated invoicing in one cloud platform. Invoices generated in under 10 seconds.
Visit indacotms.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.