AI Dictionary › Prompting

Max Tokens

The max tokens parameter caps how many tokens the model can generate in a single response. It is a technical ceiling that guards against overly long output, controls cost and ensures predictable latency. Example: you set max_tokens = 150 for a short summary task; the model stops when it hits the limit, even mid-sentence if it has not finished. Note: max tokens does NOT instruct the model to be concise, it only truncates. For genuinely short answers you must also say so in the prompt ("summarise in at most 3 sentences").

Definition

Use it to make an application safe: avoid oversized responses, cap per-call spend, respect interface limits. Remember the limit adds to input tokens within the context window: a very long input leaves less room for output.

Related terms

More in Prompting

Put it into practice

From our network

INDACO TMS: Transport Management for European Logistics

Shipment tracking, multi-carrier EDI and automated invoicing in one cloud platform. Invoices generated in under 10 seconds.

Visit indacotms.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.