pre-registration · interactive

Is the Language Hardto Model, or Its Tokenizer

A pre-registered test of whether cross-lingual efficiency gaps come from languages or from how we represent them. Sanskrit encodes 5.07 characters per token under an aligned tokenizer, 1.13 under GPT-4's. Five languages, three scripts, four tokenizers.

14 beats · 18 min to read in full

Your path

Where do you want to start?

1 of 14 beats walked · the glass is on your branch