Subscribe
Sciences

Byte-level LLM conversion — under 1% extra training, but where are the savings?

Byte-level LLM conversion — under 1% extra training, but where are the savings?

Can a byte-level LLM reuse an already capable model while learning to read finer details? Benjamin Minixhofer of Ai2 and the University of Cambridge, together with Valentin Hofmann of LMU Munich and Ai2 and their collaborators, published a conversion method called byteification in Nature on October 7. The additional training used less than 1% of a typical pretraining budget. Source: Nature 논문

The economic question is whether this makes AI development cheaper. The study demonstrates a way to reuse existing models; it does not measure a company’s total bill. Following the change in how models read text reveals where development costs might fall and where testing costs remain.

Minixhofer, Hofmann and the team behind the work

Minixhofer, the first and corresponding author, carried out the experiments. Hofmann, Edoardo Ponti of Imperial College London and Luca Soldaini of Ai2 co-designed the core experiments with him. Tyler Murray contributed to implementation and testing. Other co-authors, including Cambridge’s Anna Korhonen, helped guide the research and experimental design. Source: Nature 논문

Bolmo first appeared in December 2025. The new event is its Nature publication and accompanying releases, rather than the first announcement of the model. Ai2 has extended conversion beyond Olmo to the Qwen3 and Llama3 families and released intermediate models trained through stage 1, giving other labs a reusable starting point. Source: Ai2 공식 발표

Why a byte-level LLM can see finer details

Hofmann’s motivation concerns the difference between knowing a word’s meaning and knowing its spelling. Counting the letter r in strawberry requires access to individual positions. Many language models instead process tokens representing words or word fragments, so a token’s internal spelling becomes something the model must learn separately. Source: LMU 연구진 해설

This does not mean tokenization deletes letters from the source file. Ordinary tokenization can be reversible. The difficulty is that individual character structure is not directly exposed to the model during computation. Byteification changes the representation available for learning; it does not recover missing source data.

A byte is a small unit of digital storage. In UTF-8, an ordinary English letter takes one byte, while a Korean syllable takes several. Treating every byte as a character would therefore give misleading estimates of Korean processing speed and cost.

Minixhofer’s architecture does not send every byte separately through the expensive model core. A small input module groups bytes into variable-length patches. The existing large model processes those patches, and an output module produces byte-level predictions. It also reuses subword embeddings from the original model. Grouping still exists, but its implementation moves inside the architecture. Source: Nature 논문

  1. 1
    UTF-8 bytes

    Fine-grained text encoding

  2. 2
    Encoder and patches

    Group bytes into variable-length patches

  3. 3
    Reused model core

    Process relationships between patches

  4. 4
    Byte output

    Decoder predicts the next byte

What “under 1%” measures: 49.1 billion training tokens

The conversion has two stages. First, the team freezes the original model’s core and trains the new input and output components to reproduce its behavior. It then trains the entire model to exploit byte-level information. The first stage establishes compatibility; the second adapts the combined system. Source: Nature 논문

The Nature version reports 9.8 billion tokens for stage 1 and 39.3 billion for stage 2, totaling 49.1 billion. The authors describe the total as less than 1% of a typical pretraining budget. The chart compares training volumes, not shares of GPU charges or electricity consumption. Source: Nature 논문

This is not evidence that the total cost of building AI falls by 99%. The original model’s pretraining has already been paid for. Data preparation, conversion experiments, security checks and service integration also remain. The result supports a lower incremental barrier to trying a byte-level model when a suitable source model already exists.

Where coding performance and inference speed improve

In Minixhofer’s evaluations, Bolmo outperformed earlier byte-level models of comparable size; on STEM evaluations, for example, it scored 16.5 percentage points above Meta’s BLT 7B. It did not beat its subword counterparts on every task. A control model continued on the same training data is particularly useful for separating the architecture change from the benefit of extra training. Source: Nature 논문

In coding evaluation, pass@1 measures success from a single candidate, while pass@16 asks whether a set of candidates contains a successful answer. Against the Olmo control trained on the same data, Bolmo declined substantially on the former and improved on the latter. A service needing one immediately useful answer may have different economics from a system that automatically tests many candidates. Source: Nature 논문

Character gains also cannot be credited to architecture alone. Training included synthetic English exercises such as reversing or editing letters. The team excluded overlap with test words and observed gains on multilingual evaluation, but the character-oriented training data still contributed to the result. Source: Nature 논문

Observed improvements

Improved character manipulation

Character tasks

Code candidates · Better success across multiple candidates

Limits to generalization

Single-candidate success declined

First code answer

Attribution and uses · Synthetic training contributes; DNA needs testing

The team also measured speed. Larger patches reduce how often the expensive core must run. Extended Data measurements used H100 GPUs with a batch size of one. At roughly 6.6 bytes per patch, Bolmo 7B overtook Olmo 3 7B in prefilling latency—the time spent processing the prompt and preparing the first output. Source: Nature 논문

That does not establish faster generation for every workload. Larger patches trade quality against efficiency, and concurrency and prompt length matter. A deployment should compare completed answers at equal quality on the same hardware. Raw tokens per second are not directly comparable when one system’s token is a byte and the other’s is a word fragment.

Where spending could fall—and where it remains

From here on, this is economic inference from the paper’s numbers, not measured business outcomes. The first possible economic effect is a lower cost of experimentation. A lab or company with a suitable model could test a byte-level alternative without launching a full pretraining project. Public code and intermediate checkpoints may reduce the work needed to reproduce a starting model.

There is an established competing research path. Meta FAIR introduced BLT in 2024, using dynamic patches of bytes. Byte-level modeling itself is not the new invention here. The distinctive contribution of Hofmann and Minixhofer’s team is making practical use of training already invested in existing subword models. Source: Meta FAIR BLT

Serving costs require a separate calculation. Savings on conversion can be offset by byte-level computation or engineering work on a new runtime. Ai2’s public implementation uses specialized neural components and optimization libraries, so deployment involves more than changing a setting in an existing service. Source: AllenAI bolmo-core

A useful business metric is total cost per validated task. Allocate conversion and integration costs across usage, add generation plus testing and correction costs, then divide by the number of validated results. If success improves by generating more code candidates, their generation and testing costs must be included too.

  1. 1
    Conversion and integration

    Allocate fixed costs across usage

  2. 2
    Generate, test and correct

    Include the cost of extra candidates

  3. 3
    Divide by validated results

    Compare tasks at the same quality

If these economics work, conversion services, domain-specific evaluation and specialized hosting could become revenue opportunities. Providers that train a fresh model for each customer could face competition from reusable foundations. The available evidence establishes research and deployment tools, however, not Bolmo customer revenue, contracts or the scale of industrial displacement.

What DNA and Korean applications still need

Hofmann points to code and biological sequences as possible applications where fine structure matters. A single change in a DNA string can be meaningful. But a general model that handles characters well does not automatically understand biology. The byteification results do not establish better drug discovery success or lower laboratory costs. Source: LMU 연구진 해설

As a separate industry example, NVIDIA offers the genomic model Evo 2 through NIM deployment tooling and documents commercial-use availability. This shows an existing route for deploying biological AI. It is not evidence for Bolmo: the two models have different training targets and validation requirements. Source: NVIDIA Evo 2 NIM

A Korean biotechnology company could begin by testing whether the method reduces errors when handling sequences alongside research documents. Exact sequence handling, biological function prediction and experimental reproducibility require separate evaluations. Korean workloads also need direct measurement of input length and latency because characters occupy multiple bytes.

Our article on MedGemma’s medical adaptation examined specialization for a task. This research changes the units through which a model receives text. As with STATE and the economics of virtual cells, improved computation must still be connected to validated savings in actual research.

Minixhofer and Hofmann have opened a way to try a different representation while reusing prior training investment. Business value will depend on what happens after conversion: retained accuracy on a company’s own data, faster responses at equal quality, and lower cost per validated result.


The value of conversion is measured in cost per validated result.

Sources and further reading

For information only — this is not a recommendation to buy or sell any asset.

💱 FX calculator Subscribe

Comments 0

  • No comments yet — be the first.