AI Dictionary › Fondamenti AI

Tokenization

Tokenizzazione

Tokenization is the process of splitting text into minimal units called tokens that a model can process numerically. A token isn't always a whole word: it can be a word, a syllable, a single character, or a fragment like a suffix. Picture cutting a necklace into beads: each bead is a token, and the model counts and orders beads rather than reading the continuous thread. Methods like Byte-Pair Encoding build a vocabulary of recurring fragments, so even rare or invented words are composed from known pieces. Each token becomes an ID number, the bridge between human language and computation.

Definition

Understanding tokenization is practical: API costs and context limits are measured in tokens, not words. Languages with long compound words consume more tokens, and text heavy with symbols or code can fragment unexpectedly, affecting both price and performance.

Related terms

More in Fondamenti AI

Put it into practice

From our network

HSE Genius: AI for Safety Data Sheets

Extract SDS data, H phrases and ECHA compliance checks in seconds, powered by AI.

Visit hsegenius.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.