For years, Large Language Models (LLMs) have relied on a middleman called the tokenizer. Before a model even sees a word, the tokenizer chops text into 'subwords.' It sounds efficient, but it's actually a hidden bottleneck that makes models struggle with spelling, complex code, and diverse scripts. But a groundbreaking new study published in Nature suggests we can finally cut out the middleman.
The Magic of Byteification
Researchers have introduced a method called "byteification" that retrofits existing subword-based models to operate directly on raw bytes. Instead of starting from scratch—which would cost millions in compute—this two-stage conversion process adapts a pre-trained model with minimal extra training. In fact, it costs less than 1% of a standard pretraining run.
Better Science, Better Code
This isn't just a technical curiosity; it delivers real-world performance gains. By operating at the byte level, models become more robust across diverse data types. For instance, Bolmo 7B (a byteified version of Olmo 3) showed a staggering 16.5% absolute improvement in STEM tasks. Because the model sees the raw encoding of data rather than arbitrary subword chunks, it handles scientific data and varied scripts with far greater precision.
A New Standard for Open AI
Ai2 has already released checkpoints like Bolmo 7B and Bwen 8B to prove that this approach generalizes across different model families, including Llama. By removing the tokenizer limits, we are moving toward models that truly understand the raw fabric of digital text.
As we move away from the constraints of subwords, the next generation of LLMs will likely be more flexible, more accurate in technical fields, and far more accessible to the global variety of human languages.
Sources
Media




