pre-registration · interactive
Is the Language Hardto Model, or Its Tokenizer
A pre-registered test of whether cross-lingual efficiency gaps come from languages or from how we represent them. Sanskrit encodes 5.07 characters per token under an aligned tokenizer, 1.13 under GPT-4's. Five languages, three scripts, four tokenizers.
14 beats · 18 min to read in full
Your path
Where do you want to start?
1 of 14 beats walked · the glass is on your branch