Tokenization Bench

Random Word

Tokenized Results

Tokenizer Vocab Size Tokens Token Count
IP.appify - umtoken 96k* uzdot+ajiem 2
OpenAI - GPT-5 200k uzdotajiem 3
Google - Gemma 4 262k uzdotajiem 4
Alibaba - Qwen3.6 248k uzdotajiem 4

If the tokens contain unexpected characters or hexadecimal codes, this is not an error. It is the way in which the respective tokenizer encodes non-ASCII characters.

umtoken is trained for the 24 official languages of the EU (bg, cs, da, de, el, en, es, et, fi, fr, ga, hr, hu, it, lt, lv, mt, nl, pl, pt, ro, sk, sl, sv) on the allenai/nllb and HuggingFaceFW/clean-wikipedia datasets, together with morphological forms from Wiktionary.
For more information on umtoken and an explanation of its levels, please visit us on GitHub.

* Here, 'k' denotes a factor of 1024, not 1000.