Raw text in, ids out: the interface is the idea

Classical NLP pipelines assume a stage before the subword model: a pre-tokenizer that splits text into words, usually on whitespace and punctuation, sometimes with a language-specific segmenter (MeCab for Japanese, a dictionary for Thai). Byte-pair encoding then learns merges inside those word boundaries. The subword model is only half a tokenizer; the other half is a pile of assumptions that vary by language and by codebase.

SentencePiece deletes that stage. Its input is a raw Unicode string; its output is a list of ids. The segmentation model is trained directly on unsegmented text, so nothing upstream needs to know what a ‘word’ is. Two consequences follow immediately. First, the tokenizer is language-agnostic by construction. Second, it is reproducible: a single .model file carries the normalizer, the pieces, and their scores, so anyone who loads it gets byte-identical ids without reproducing your preprocessing script.

Advertisement

The meta-symbol and the lossless round trip

If you throw whitespace away, you cannot put it back. ‘New York’ and ‘NewYork’ produce the same token sequence under a whitespace-splitting tokenizer, and detokenization becomes guesswork about where spaces belong — guesswork that is different in English, French, and Japanese. SentencePiece instead escapes whitespace: every space becomes the meta-symbol ▁ (U+2581, LOWER ONE EIGHTH BLOCK) and is treated as an ordinary character the model may merge into pieces.

encode: "Hello world"
  → escape:  ▁Hello▁world
  → pieces:  [▁Hel, lo, ▁world]
  → ids:     [8221, 385, 1128]

decode: concat(pieces) → "▁Hello▁world"
  → replace ▁ with " " → "Hello world"

Decoding is ▁-substitution on a plain concatenation — the same three lines for every language. That is the lossless-detokenization property, and it is why the leading ▁ on ▁Hello matters: it encodes ‘a space preceded this’ in the token itself.

Advertisement

Why this matters where whitespace does not delimit words

The pre-tokenizer assumption is invisible until you leave English. Japanese, Chinese, Thai, Khmer, and Lao do not put spaces between words. A whitespace pre-tokenizer hands the subword model one enormous ‘word’ per sentence, and any merge-inside-words discipline becomes meaningless. The usual workaround — run a language-specific morphological analyzer first — means a different external dependency per language, each with its own version, dictionary, and licensing.

Because SentencePiece never assumed word boundaries, it needs none of that. It learns pieces from the raw character stream, so a Japanese corpus yields pieces corresponding to morphemes and common character runs with no analyzer in the loop. The property also matters within a single multilingual model: one vocabulary trained over mixed text handles all scripts through the same code path. The fairness of that shared budget across languages is a separate question — treated in the multilingual-tokenization article — but the mechanical obstacle is gone.

Normalization: NFKC, and where it bites

Before anything is segmented, SentencePiece normalizes. The default rule set, nmt_nfkc, is Unicode NFKC plus NMT-flavored tweaks: it collapses runs of whitespace, strips most control characters, and applies compatibility folding — full-width A → A, the ligature fi → fi, ① → 1, and various quote and dash unifications. The goal is to stop the vocabulary from wasting slots on visually identical variants.

The pitfall is that NFKC is not reversible. Once fi has become fi, decode cannot restore it, so the lossless guarantee holds up to normalization, not to the literal input bytes. For code, mathematics, or any text where full-width characters and exotic spaces are meaningful, that folding is data loss. SentencePiece therefore offers nmt_nfkc_cf (adds case folding), plain identity, and user-supplied rule files. Choose deliberately; the choice is baked into the model file forever.

Byte fallback: closing the unknown-character hole

Training sees a finite corpus, so the character alphabet it learns is finite. SentencePiece makes this explicit with character_coverage — typically 0.9995 for character-rich languages like Japanese and 1.0 for Latin scripts — meaning the alphabet keeps the most frequent characters covering that fraction of the corpus and discards the long tail. Without a safety net, any discarded character at inference time collapses to <unk>, and <unk> destroys the round trip.

Byte fallback closes the hole. With byte_fallback=true, the vocabulary reserves 256 pieces <0x00>…<0xFF>, and any character with no piece is emitted as its raw UTF-8 bytes. Nothing is ever unknown, and decode still reconstructs the string exactly. The cost is length: a rare CJK character or emoji that would have been one token becomes three or four byte tokens. That is a good trade — correctness always, with a length penalty only on genuinely rare input.

One library, several model types

SentencePiece is a container for segmentation algorithms, not one algorithm. model_type selects among four: unigram (the default), bpe, char, and word. Everything discussed so far — the meta-symbol, normalization, byte fallback, the .model file — is shared infrastructure that sits above whichever you pick.

The two that matter in practice are unigram and BPE. BPE builds the vocabulary bottom-up by repeatedly merging the most frequent adjacent pair, and encodes by replaying those merges in rank order; it is deterministic and greedy. The unigram model works top-down instead: it starts from a large candidate set, assigns each piece a log-probability, and prunes by expected loss, so encoding is a Viterbi search for the highest-scoring segmentation. The practical difference is that unigram gives you a distribution over segmentations rather than a single answer — which the sampling section below depends on. Both are covered in depth in their own articles.