Before the model guesses anything, your words get chopped into tokens: reusable chunks with id numbers. It never sees your spelling. Some famous AI faceplants start right here.
Splits here follow a real tokenizer’s style; the id numbers are stand-ins,
since every model has its own vocabulary of roughly 50,000 to 200,000 tokens.