Give Up Internet!
DICTIONARY/ Alive/ Token (AI)
● Alive n. 7 yrs & counting AI era

Token (AI)

The unit a language model actually reads and writes: not a word, not a letter, but a chunk of text chosen by a compression algorithm, so 'cat' is one token and 'unbelievable' might be three. English averages about three-quarters of a word per token. Models are priced, limited and benchmarked by the token, and their strangest failures, counting letters for instance, come from it.

Born
Peak
Died
Cause of death
First sighting
Byte-pair encoding for translation (Sennrich et al., 2016); GPT-2's tokeniser (2019); per-token pricing from 2020
Life & death · 2018–2027
2018
2019
2020
2021
2022
2023
2024
2025
2026
2027
PEAK 2024
BORN · 2019
NOW · 2026

In the blog

1 article mention this record

Testimony

4 entries · newest first
Memory @caps_lock_carl · Poster 10 Oct 2024

I remember the first time I pasted text into a tokeniser demo and watched it colour the fragments. My surname was three pieces, none of them pronounceable. It explained a great deal about how the models spelled it.

Sighting @stan_twitter_stu · Poster 24 Sep 2024

Sighted on every AI pricing page since 2020: dollars per million tokens, input and output priced differently, with a calculator for people who cannot picture a million of anything. The record notes that the token became a currency before most users could define one.

Definition @ranked_rage · Poster 14 Sep 2022

The tokeniser is a dictionary of common fragments built before training, by repeatedly merging the most frequent pair of characters until the vocabulary hits a target, around fifty thousand for GPT-2 and two hundred thousand for later models. The model never sees letters. It sees numbers that stand for fragments.

Correction @stitch_sasha · Digger 15 Jan 2022

The correction the record makes most often: tokens are not words. A number can be several tokens, a space can belong to the next word, and other languages pay more tokens for the same sentence than English does. When a company says a model read ten trillion tokens, divide by about 1.3 for words.

File an entry

Add to the record

Where did you see it, and when? A date and a source that still resolves.

Filed straight to the record — it stays after you reload.