AI Dictionary › Modelli AI
Vocabolario
A language model's vocabulary is the finite, predefined set of tokens the model can recognize as input and produce as output. Every word, subword, character or special symbol the model can handle must have a corresponding entry in this list, associated with a unique numerical identifier.
The vocabulary is built before training through a tokenization algorithm applied to a large text corpus, which identifies the most frequent and useful units to include as separate entries. Vocabulary size is a design choice: a larger vocabulary reduces the number of tokens needed to represent a text but increases the size of the model's embedding and output layers.
The vocabulary directly determines how a text is broken down into tokens before being processed, affecting how efficiently the model handles different languages, technical terms, source code or non-Latin characters. Models designed to be multilingual tend to have larger vocabularies to efficiently cover multiple languages at once.
The concept borrows the everyday sense of "vocabulary" as a list of known terms, applied here not to human language as a whole but to the closed set of textual units a particular model has been trained to recognize.
From our network
Magellano GPS: Fleet Tracking Made Simple
Real-time GPS tracking, remote engine lock, fuel and CO₂ reporting for your fleet.
Visit magellanogps.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.