Tokenization is the process of breaking text into smaller units – so-called tokens – that an AI model can process. A token can be a word, a word part, or even a single character. The number of tokens affects processing speed and cost for AI models. For German, tokenization is often less efficient than for English because words are longer and more complex.
In short
What is tokenization in AI models?
Tokenization is the process of breaking text into smaller units (tokens) that an AI model can process. A token can be a word, a word part, or a character. The number of tokens affects processing speed and cost. For German, tokenization is often less efficient than for English.
Using AI in your organisation?
We show you which applications are worth it for your organisation.