AI Dictionary › Modelli AI
Vocabolario
A language model's vocabulary is the finite, predefined set of tokens the model can recognize as input and produce as output. Every word, subword, character or special symbol the model can handle must have a corresponding entry in this list, associated with a unique numerical identifier.
A language model's vocabulary is the finite, predefined set of tokens the model can recognize as input and produce as output. Every word, subword, character or special symbol the model can handle must have a corresponding entry in this list, associated with a unique numerical identifier.
The vocabulary is built before training through a tokenization algorithm applied to a large text corpus, which identifies the most frequent and useful units to include as separate entries. Vocabulary size is a design choice: a larger vocabulary reduces the number of tokens needed to represent a text but increases the size of the model's embedding and output layers.
The vocabulary directly determines how a text is broken down into tokens before being processed, affecting how efficiently the model handles different languages, technical terms, source code or non-Latin characters. Models designed to be multilingual tend to have larger vocabularies to efficiently cover multiple languages at once.
The concept borrows the everyday sense of "vocabulary" as a list of known terms, applied here not to human language as a whole but to the closed set of textual units a particular model has been trained to recognize.
Il vocabolario di un modello linguistico è l'insieme finito e predefinito di token che il modello è in grado di riconoscere in ingresso e produrre in uscita. Ogni parola, sottoparola, carattere o simbolo speciale che il modello può gestire deve avere una voce corrispondente in questo elenco, a cui è associato un identificatore numerico univoco.
Il vocabolario viene costruito prima dell'addestramento tramite un algoritmo di tokenizzazione applicato a un ampio corpus di testo, che individua le unità più frequenti e utili da includere come voci separate. La dimensione del vocabolario è una scelta di progettazione: un vocabolario più ampio riduce il numero di token necessari per rappresentare un testo ma aumenta la dimensione degli strati di embedding e di output del modello.
Il vocabolario determina direttamente come un testo viene scomposto in token prima di essere elaborato, influenzando l'efficienza con cui il modello gestisce lingue diverse, termini tecnici, codice sorgente o caratteri non latini. Modelli pensati per essere multilingue tendono ad avere vocabolari più ampi per coprire in modo efficiente più lingue contemporaneamente.
Il concetto riprende il senso comune della parola "vocabolario" come elenco di termini noti, applicato qui non al linguaggio umano nel suo insieme ma all'insieme chiuso di unità testuali che un particolare modello è stato addestrato a riconoscere.
From our network
INDACO TMS — Transport Management for European Logistics
Shipment tracking, multi-carrier EDI and automated invoicing in one cloud platform. Invoices generated in under 10 seconds.
Visit indacotms.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.