> ## Content Index
> Fetch the complete content index at: https://articles.akadata.co.uk/llms.txt
> Use this file to discover other available public pages before exploring further.

# Tech Scroll 134: Checking a 1,000x-faster tokenizer, then rebuilding it in C
- URL: https://articles.akadata.co.uk/tech-scroll-134-checking-a-1-000x-faster-tokenizer-then-rebuilding-it-in-c/
- Published: 2026-10-04T05:39:59.000Z
- Updated: 2026-10-05T23:34:52.000Z
- Description: "Do not praise the mill for the speed of its stones. Praise it when its weights are true, for fast flour mixed with chaff feeds no one."
- Author: AKADATA
- Tags: tokenizer, saphira-tokenizer, llm, localllm

A tokenizer library claims 24.53 gigabytes per second on GPT-2, or 989 times the speed of HuggingFace Tokenizers and 681 times the speed of tiktoken, with outputs that match exactly.

Those numbers come from the project’s own benchmark page, measured on a 144-core server processor.

> This Scroll describes what happened when we took the claim seriously: reproduced the comparison on our own hardware, found the implementation wrong on inputs its validation never covered, fixed what we could, rebuilt the whole thing independently in C, made it fast, and finally fed it 12 gigabytes of real text including 29 complete Bible translations.

Every number below was measured. Every bug below has a reproducer. Where we could not verify something, the article says so.

**Bottom line, stated up front:** the speed is real engineering on clean text, and the exactness claim is unproven off it.

Our C implementation now tokenizes at up to 1.70 gigabytes per second on our 32-thread desktop machine, agrees with the reference implementation on 1.23 million Bible verses to the last checksum digit, and trains new tokenizers with no Python involved.

It also disagrees with the reference on 54 inputs out of 4,821, and in every one of those cases the evidence says the reference is the one at fault.

Both halves of that sentence matter, and this article presents both.

## The problem in brief

Language models do not read text. They read token IDs: integers produced by a tokenizer that splits text into pieces and maps each piece to a number.

The GPT-2 tokenizer, still one of the most widely used byte-level designs, does this in three stages.

First it splits text with a regular expression into pretokens: small runs of letters, numbers, punctuation and whitespace.

Then it maps each byte to an alphabet symbol.

Then it merges symbols pairwise, in an order learned from training data, until each pretoken becomes one or more token IDs.

Tokenization sits on the critical path of both training and inference serving. Every byte of training data passes through it, and every prompt passes through it before generation starts.

A slow tokenizer taxes the whole pipeline, which is why a claim of a thousandfold speedup deserves attention and also deserves checking.

A wrong tokenizer is worse than a slow one, because a wrong token ID is silent.

No crash. No error. Just a slightly different model trained on slightly different numbers.

The only defence is exactness proven against independent implementations, with every disagreement explained at the operation where the two sides first diverge rather than voted away.

That standard shaped everything that follows.

We built a second implementation from the specification, in plain C, with no shared code, and required it to produce byte-identical token IDs.

Where the two disagreed, we treated neither as authoritative and traced each divergence to its first divergent operation.

That discipline found real defects on both sides, including one that crashes the reference implementation.

The short version is that exactness is a property you prove, not a property you assert, and the proof has to cover the inputs nobody thinks to test.

## The existing wheel: what the claimant says

Gigatoken presents itself as the fastest tokenizer for language modelling, supporting a wide range of CPUs and nearly all commonly used tokenizers.

Its benchmark page reports encoding throughput on a fixed 11.9 gigabyte web-text file, on a dual-socket 144-core server processor, with a table per tokenizer.

For GPT-2 the figures are:

**Gigatoken:** 24.53 GB/s  
**HuggingFace Tokenizers:** 24.8 MB/s  
**tiktoken:** 36.0 MB/s

Hence the headline claims of 989 times and 681 times.

Similar ratios are listed for around twenty other tokenizers.

The page also shows results on smaller machines, gives the exact commands to reproduce the benchmark, and reports a validation step in which 20,401 documents match between implementations.

The page’s own FAQ explains where the speed comes from: replacing the regular expression engine with hand-tuned SIMD pretokenization, caching the encoding of previously seen pretokens, and minimising the cost of moving data between Python and the native code.

That account is credible on its face.

Words repeat constantly in natural text, so remembering their token IDs can skip the merge work entirely. Avoiding an expensive general-purpose regular-expression engine can matter. Crossing the Python/native boundary less often can matter.

As this article shows, each of those three is indeed where a large share of the performance lives.

Nothing about the claim is prima facie absurd.

It is, however, exactly the kind of claim that has to be reproduced before it is repeated, because every element of it, including the hardware, corpus, API path and validation set, affects the number.

The headline figures all come from the vendor’s best configuration on the vendor’s best hardware.

There are also two things the benchmark page does not do.

First, it never names the input domain of its validation.

The 20,401 matching documents are web text, which is valid UTF-8 by construction. Nothing in the validation covers arbitrary bytes.

That matters because a byte-level tokenizer explicitly accepts arbitrary bytes, including values that are not valid UTF-8 at all, and behaviour there is part of the contract.

Second, the per-core arithmetic is left to the reader.

24.53 gigabytes per second over 144 cores is about **170 megabytes per second per core**.

> Keep that figure in mind. It becomes important later.

## Learn the existing wheel first: how we set up the comparison

Before writing any C, we established the reference behaviour on our own machine: a 32-thread desktop processor, shared with other workloads, with power management initially left alone.

We pinned the exact upstream commit of the reference implementation, verified its digest, and built it unmodified in an isolated environment.

We then ran three implementations on identical inputs:

**the reference Rust wheel, HuggingFace Tokenizers, and later our C implementation.**

The rule throughout was simple:

**A speed number without a matching checksum is void.**

A fast run that produces different token IDs from the reference has not tokenized faster. It has produced a different answer faster, which is worth nothing.

Every throughput figure in this article is accompanied by a checksum over all emitted token IDs, and every comparison asserts equality before it asserts speed.

The corpora increased in size and nastiness as the work progressed.

We began with a structural battery of 611 documents covering empty input, whitespace edge cases, contractions, punctuation, all 256 byte values, invalid UTF-8 and embedded NUL.

Then came a seeded fuzz corpus of several thousand documents mixing arbitrary bytes with valid text: 4,821 documents in the standard configuration and 8,921 expanded.

After that came a 20 megabyte web-text slice, 119 megabytes of Bible text, and finally the full 11.9 gigabyte web-text file used by the claimant.

The Bible corpus deserves more explanation because it became central to the work.

It contains 29 complete English translations, including King James, ESV, NIV, NASB, NLT, WEB, Young’s Literal and 22 more.

Each contains 66 books.

In total there are 901,939 verses and 1.23 million lines, extracted from per-book JSON files into verse text, with a manifest recording per-translation counts and book-level split files excluded so nothing is double-counted.

Verse text only. No JSON markup, no references and no duplication.

![](https://articles.akadata.co.uk/content/images/2026/10/volume-4.svg)

****FIGURE: What was tokenized: the three corpora by input megabytes — volume**

The first headline result was a validation, not a speed result.

On the full Bible set, our C implementation and the reference agree on all **1,231,328 lines and 28,913,671 tokens**, with an identical checksum to the digit.

HuggingFace Tokenizers, run over the same 1.23 million verses in parallel processes, produces the identical checksum a third time.

Three implementations.

Three codebases.

One answer on every verse.

That is what a validation looks like, and it is the foundation everything else stands on.

## What we tested: the bugs the validation corpus never covered

With agreement established on clean text, we turned to the input class the claimant’s validation never exercises:

**arbitrary bytes.**

The differential harness compares exact token IDs on seeded random byte strings, and it found 155 disagreements out of 4,821 documents.

Every one of them involved bytes outside valid UTF-8.

That alone would have been worth reporting.

What made it significant is where the investigation led: not to one bug, but to a small family of them in both implementations, each precisely localised by tracing the scanner directly rather than inferring behaviour from outputs.

It is worth pausing on that methodological point because it determined everything.

Token IDs and pretoken spans do not uniquely identify how many bytes the scanner consumed.

Byte-pair merging and same-class span merging can make different scanner histories observationally identical.

Three successive black-box probes gave confident, specific, wrong answers.

The fourth attempt instrumented the decoder itself to emit, per position, the code path taken, width proposed, width accepted, value assembled, class returned and next offset.

No token IDs.

That trace settled in one run what inference could not settle in three attempts.

Readers building audits should take note:

**measure the decision, never the shadow it casts.**

### Defect one: width inferred without validating continuation bytes

The reference decoder inferred a codepoint width from the lead byte’s range alone and never checked that the following bytes were genuine continuation bytes.

A byte merely resembling a UTF-8 lead swallowed its neighbours.

For example, 0xEF followed by an ASCII byte was decoded as a three-byte character. The scanner advanced three bytes over input containing no such character, the pretoken span moved, and letter runs that must stay together were split.

A minimal case is the six bytes:

```
\xe0@yp
```

which produce wrong IDs, while:

```
a@yp
```

is correct.

We fixed this in the reference.

Width is now a proposal validated against real continuation bytes.

The frozen valid-UTF-8 oracle hash did not move by a single bit, proving the fix touched only the previously untested domain.

That scoping matters. It is why the existing validation corpus stayed green throughout.

### Defect two: bytes 0x80 to 0xBF

More consequentially, both the reference and, initially, our C implementation mishandled bytes in the range 0x80 to 0xBF.

In a UTF-8 parser these look like continuation bytes.

In the GPT-2 byte alphabet they are ordinary characters.

All 64 map to real codepoints. Thirty-seven of them are letters, with a few numbers such as superscripts and vulgar fractions, plus others.

The reference proposes a three-byte width for them and assembles continuations that follow.

As a result, the three bytes:

```
89 88 81
```

become a single invented character instead of three separate letters.

Likewise:

```
BD 9B 9A
```

becomes one invented letter instead of the number ½ followed by two letters.

Minimal cases are:

```
\x89\x88\x81\xab
```

and:

```
\xbd\x9b\x9a\xe9
```

The reference is still wrong here at the time of writing.

Our C implementation applies the corrected rule:

**never a lead, width exactly one, byte-alphabet category.**

It matches the independently derived expectation in spans and IDs.

This is recorded as an upstream defect with reproducers, but deliberately not fixed in their tree.

The deliverable is our implementation, not a rehabilitation of theirs.

### Defect three: a crash

Feeding a single byte 0xFF to the reference’s SentencePiece path segfaults the release build because its UTF-8 guard was compiled out in release.

That defect class, found during hardening before any of the above, is why the audit record distinguishes the hardened branch from pristine upstream.

It belongs here because a benchmark page that never feeds its subject an invalid byte cannot claim to have validated a byte-level tokenizer.

And because “fast” is not interesting in code that falls over on the input domain it advertises.

### Defect four: ours

One of our own defects was equally instructive.

Our C implementation ignored special tokens entirely.

Real web text uses:

```
<|endoftext|>
```

as a document separator **2.35 million times** in the 11.9 gigabyte file.

That produced a 2.6 per cent divergence on real data.

We reduced the problem to the 13-byte input:

```
<|endoftext|>
```

and fixed it by splitting on leftmost-longest special matches before pretokenizing, while leaving the no-special fast path untouched.

The same fix exposed and repaired a use-after-free in our tokenizer-file loader.

Both are now regressed in selftest.

The scoreboard, stated without adornment:

The reference was wrong on 155 fuzz inputs, plus the crash, plus the 0x80 to 0xBF band.

Our C implementation was wrong on special tokens and briefly on two forensic cases during development.

The independent Python reference we built for adjudication turned out to be wrong on whitespace and was demoted to experimental.

Every implementation in this story has been caught being wrong, including ours.

Each occasion is recorded rather than smoothed over.

**An oracle that cannot say it was wrong cannot be trusted when it says it is right.**

## What we measured: from one megabyte per second to over a gigabyte

The starting point was deliberately unoptimised C:

```
-O0
```

No SIMD.

Plain loops.

One translation unit per pipeline stage.

On a 20 megabyte web-text slice it managed **0.98 megabytes per second** of input.

The reference wheel on the same machine did **7.23 megabytes per second**.

A naive reading says C lost.

The honest reading says the comparison had barely started, because essentially all of the 0.98 was accidental cost rather than tokenization work.

What followed was a sequence of measured steps, each kept or reverted on numbers, with the checksum enforced throughout so no speedup could smuggle in a wrong answer.

![](https://articles.akadata.co.uk/content/images/2026/10/speed-ladder-4.svg)

FIGURE: Each optimisation measured; checksums identical at every step — speed-ladder

The byte-to-codepoint map ran a 256-iteration count per input byte.

Replacing it with a generated 256-entry table changed nothing measurable: 2.80 versus 2.69.

That is itself a result worth reporting.

**Profile before optimising, because the obvious hotspot is not always the real one.**

The profiler said the vocabulary constructor owned 80 per cent of wall time: an O(n²) insertion sort over 50,257 entries, executed once per tokenizer load.

Replacing it with `qsort` more than doubled throughput, from 2.69 to 6.23.

That was a wrong algorithm, not a missed optimisation, and the distinction matters.

A 256 by 256 direct table for single-symbol merge pairs, the entire first round of every merge, added 13 per cent.

A hash index alongside the sorted vocabulary array added another 8 per cent, with the array retained as the auditable reference order.

AVX2 key comparison added 20 per cent, with a scalar fallback so the binary remains portable.

Link-time optimisation added 5 per cent for free.

Then came the single largest item after fixing the sort:

**deleting a debugging `getenv` call that ran once per pretoken, over a million times per run.**

Throughput went from **163 to 299 megabytes per second**, an 83 per cent gain from removing eleven characters of leftover instrumentation.

Profile your hot loop for system calls before reaching for intrinsics.

Threading came next, and it required a design decision to stay honest.

Twenty-four private tokenizers, one per worker and roughly five megabytes each, blew the shared cache and capped scaling.

The fix was `gt_tokenizer_encode_mt`: one read-only tokenizer shared across workers, with per-call statistics and scratch space.

It was proven race-free by checksums identical across one to thirty-two threads after two real races were found by the same mechanism: a torn count sizing a stack copy, and a multi-writer interleave of the lock-free sequence counter.

Both were fixed with clamping and striped spinlocks.

Threads alone took the 20 megabyte slice from 12 to over 300 megabytes per second, with the sweet spot at 24 threads.

Thirty-two threads slow down on this chip’s efficiency-core layout, and pinning experiments confirmed that one thread per physical core beats hyperthreaded sharing here.

Then came the cache.

It deserves its own section because it failed first.

A pretoken memo is the obvious optimisation for this workload.

Natural text repeats words constantly, so remembering a pretoken’s token IDs skips the merge and vocabulary work entirely.

Version one lost sevenfold.

It used two hashes per pretoken, unbounded linear probes on a table that filled up, and 256 kilobytes thrashing the fast cache.

It was reverted, not defended.

Version two is direct-mapped: exactly one slot per pretoken, overwrite on collision.

It computes one hash reused for insertion, keeps hot entries in fast cache, and verifies every hit by comparison so a hit is bit-identical to a miss by construction.

That version added a modest 12 per cent on web text.

The large version, sized by measurement up to the level-3 cache bound, uses 256 thousand slots.

That bound is enforced at load even over explicit requests, because a cache larger than L3 becomes a main-memory table that poisons everything else.

The result was dramatic.

Web text went from **362 to 880 megabytes per second**.

The Bible corpus reached **1.70 gigabytes per second**.

The sweep across cache sizes is monotonic until the cache spills out of L3 and then regresses, exactly as the hardware model predicts.

![](https://articles.akadata.co.uk/content/images/2026/10/rust-vs-c-3.svg)

****FIGURE: Without our cache, the reference wins on both corpora — rust-vs-c**

![](https://articles.akadata.co.uk/content/images/2026/10/cached-vs-cached-2.svg)

****FI**GURE: With our 256k-slot cache, C wins on both corpora — cached-vs-cached

![](https://articles.akadata.co.uk/content/images/2026/10/cache-sweep-4.svg)

****FIGURE: Cache size against speed: the level-3 bound is the sweet spot — cache-sweep**

The final pre-optimisation numbers on our hardware, using cold processes and a ramdisk so disk I/O is not being measured, with identical checksums everywhere:

| Implementation         | 20 MB web text | 119 MB Bible, 1.23M verses |
| ---------------------- | -------------- | -------------------------- |
| C, 256k cache          | 880 MB/s       | 1,701 MB/s                 |
| Reference Rust wheel   | 122 MB/s       | 714 MB/s                   |
| HuggingFace Tokenizers | 4.5 MB/s       | 46 MB/s, 24 processes      |

Two honest qualifications belong here, because a benchmark article that hides them is marketing.

First, corpus shape dominates.

The reference beats us on short repetitive verses because its large cache earns its keep, and loses three to seven times on diverse long lines where our per-byte speed wins.

Neither “C is faster” nor “Rust is faster” is true without naming the corpus and cache state.

Warm second runs in one process flatter the reference tenfold through its hot cache, and we report cold numbers throughout for that reason.

Second, the claimant’s 24.53 gigabytes per second is measured on 144 dedicated server cores against our 32 shared desktop threads.

Per core, that is roughly **170 megabytes per second versus our 70**.

The remaining gap is consistent with their vectorised splitter, working cache and full-power server silicon.

Plausible.

Consistent with everything we measured.

Not verifiable from here.

We claim what we measured and label the rest as extrapolation.

## Giving it real work: training a tokenizer from raw bytes, in C

An encoder is half a tokenizer.

The other half is the trainer: the program that reads a corpus and learns the merge table in the first place.

Ours is written in C, takes raw bytes, and writes the same tokenizer file format the encoder loads.

A trained model therefore round-trips with no other tool involved and no Python anywhere in the loop.

The method is standard byte-level byte-pair encoding.

The corpus is split into pretokens using the same splitter as the encoder.

Every pretoken begins as its byte-alphabet symbols.

Adjacent pairs are counted.

The most frequent pair is merged everywhere it occurs.

That repeats until the vocabulary is full.

Pair counts are maintained incrementally, so a merge touches only words containing that pair, with the best pair drawn from a heap.

Tie-breaking is deterministic and documented because ties are the one place where two correct trainers may legitimately differ.

The policy has to be pinned rather than accidental.

On 200 kilobytes of web text with a 2,000-token vocabulary, the C trainer finishes in under a tenth of a second, single-threaded, against the reference parallel trainer’s seven hundredths.

They are essentially tied in wall time at that size, with the C version using one core while the other implementation uses many.

The first five merges agree character for character, and each implementation loads the other’s output file and produces identical token IDs.

That interoperability check matters more than the timing.

It proves that the file format, byte alphabet and merge semantics are shared, not merely similar.

![](https://articles.akadata.co.uk/content/images/2026/10/train-2.svg)

****FIGURE: Training wall time, C against reference — train**

Scale found two more cliffs.

Counting distinct pretokens across 119 megabytes of Bible text, representing **26.1 million occurrences, 45,242 distinct pretokens and 8,862 appearing exactly once**, is effectively instantaneous with a hash table.

Training a 45,498-token vocabulary from it, one token per distinct word plus the 256 byte symbols, takes 65 seconds single-threaded.

At larger scale, the trainer produced a 16,384-token vocabulary from gigabytes of real web text in about a minute on one core.

The long tail is real and visible in the numbers.

Nearly a fifth of distinct words occur exactly once, and each consumes one merge to learn a token that appears a single time in a hundred megabytes.

Production vocabularies therefore stop before absorbing all of that, or prune afterwards.

The counts show exactly why.

The tail is long, flat and expensive.

None of this required leaving C.

The trainer, counter, encoder and benchmark share the same pretokenizer and byte alphabet.

That means a defect in either can poison both, which is precisely why they are tested together:

train a tiny controlled corpus, assert the exact merges, load the result, encode with it, and compare IDs.

That round-trip is in the selftest suite alongside the other 1,690 checks, and it runs in milliseconds.

## Commands, code and config, so you can check

Everything below runs as written, given the repository at the commit recorded in the evidence bundle and the fixture tokenizer it pins.

Build the oracle, the optimised binary and the selftest:

```
cd src

make            # -O0 oracle build, zero warnings tolerated
make opt        # -O2 -mavx2 -flto, same source
make asan       # sanitizers

./build/gt_selftest <tokenizer.json>     # 1690 checks
```

Encode a few strings and inspect the spans:

```
./build/encode_demo <tokenizer.json> "hello world" "a  b" "café"

# hello world -> 31373 995
# spans: [hello] [ world]
```

Run the differential comparison against the reference implementation, using the pinned wheel environment:

```
python3 examples/quickstart.py <tokenizer.json> ./build
```

Train a tokenizer directly from raw bytes, with no Python anywhere in the loop:

```
./build/gt_train corpus.txt 2000 model.json
./build/gt_dump_ids model.json
```

Count distinct pretokens as the trainer sees them:

```
./build/gt_words corpus.txt

# total=26102475 distinct=45242 singletons=8862 suggested_vocab=45498
```

The implementation keeps the pipeline deliberately visible, one stage per file.

![](https://articles.akadata.co.uk/content/images/2026/10/pipeline-1.svg)

****FIGURE: One stage per file: a mismatch names its stage — pipeline**

## Why it works: three decisions that did all the work

The first is **oracle discipline**.

Every optimisation above was gated on a checksum over millions of token IDs, computed independently by a second implementation and required to match exactly.

That gate caught a wrong optimisation: a whole-pretoken shortcut that changed output because byte-pair merging does not necessarily reproduce whole-pretoken vocabulary entries.

It was reverted before shipping.

That gate is the reason any speed number here can be believed.

**Speed without that gate is just faster output of unknown correctness.**

The second decision is **tables derived from rules, never hand-maintained**.

The byte alphabet, byte-to-category map and Unicode ranges are generated from specifications by scripts in the tree, and the selftest verifies every entry.

When the generator and consumer disagree, the test fails loudly instead of drifting silently.

Hand-written lookup tables are where transcription errors go to hide for years.

The third decision is **measure, then cut**.

Five of the seven attempted optimisations in this project either did nothing, hurt performance, or were simply wrong.

The byte-map table and ASCII fast path did nothing useful on this workload.

The first memo implementation and profile-guided layout trained on the wrong corpus hurt.

The whole-pretoken shortcut was wrong.

Only measurement distinguished them.

The discipline of reverting failures is what kept the codebase clean enough to continue.

The two-line summary of the entire speed programme is this:

**a quadratic sort, a debug system call and cache residency were each worth more than all the vector instructions combined.**

We only know that because each step was timed in isolation.

## Failure modes: what not to assume

Do not assume a validation corpus covers the input domain.

Ours did not, and neither does the claimant’s.

Both validated on clean text while the contract explicitly includes arbitrary bytes.

Any tokenizer benchmark whose corpus is valid UTF-8 throughout has proven nothing about a byte-level tokenizer’s edge behaviour, no matter how many documents it lists.

Do not assume one throughput number characterises an implementation.

Corpus shape, including line length, repetition and byte distribution, matters.

Cache state matters.

Thread count relative to core topology matters.

Power management matters.

Neighbouring load matters.

Each moved our numbers by factors of two to ten.

Report the corpus, cache state, thread layout and spread across repeats, or report nothing.

Do not assume vector instructions are where the speed is.

On this workload the ranking was:

algorithmic complexity, accidental system calls, cache residency and sharing, thread topology, and only then vectorised comparison.

There is no dot product in tokenization, so matrix extensions have no natural role here.

The integer vector instructions earned their modest share honestly.

The rest would be theatre.

Do not assume agreement is correctness.

Three implementations can agree and still share a wrong reading of the specification.

Ours did on whitespace handling in an early reference transcription until the contract was written down independently.

Agreement is evidence, and the more independent the implementations, the stronger that evidence becomes.

But the contract outranks any vote.

Do not assume training inherits encoding correctness.

The trainer and encoder share the pretokenizer and byte alphabet, so a defect in either poisons both.

Tie-breaking policy, documented here as lowest pair then first-seen, is the one place two correct trainers may legitimately differ.

Compare trained models by round-tripping them:

load, encode, compare IDs.

Do not eyeball merge lists and call them equivalent.

## Conclusion

A thousandfold claim, checked end to end, turned out to be real engineering on clean text with an unproven exactness story off it.

Rebuilding the thing independently in C produced an implementation that is faster on our hardware:

**880 megabytes per second on web text and 1.70 gigabytes per second on Bible verses, against 122 and 714 megabytes per second for the reference wheel.**

It is exact where independent authorities exist:

**1.23 million verses, three implementations, one checksum.**

It is documented where it diverges:

**54 inputs where the evidence says the reference, not us, is wrong.**

And it can learn its own tokenizers directly from raw bytes, with no Python in the loop.

It also trains a 16,000-token vocabulary from gigabytes of real text in about a minute on one thread, while the industrial parallel implementation wins wall-clock time by using many cores.

Nothing here required exotic hardware, exotic instructions or trust.

It required a checksum, a corpus with sharp edges, and the willingness to write down every occasion when the work proved itself wrong.

The repository, evidence and reproducers are linked below for anyone who wants to repeat any of it.

## References

Gigatoken benchmarks and FAQ, including the claimant figures of 24.53 GB/s for GPT-2, 989x versus HuggingFace Tokenizers and 681x versus tiktoken, along with its file API and validation procedure:

[https://github.com/marcelroed/gigatoken](https://github.com/marcelroed/gigatoken?ref=articles.akadata.co.uk)

GPT-2 byte-level BPE, including the original encoder, merge ranking and `bytes_to_unicode`:

[https://github.com/openai/gpt-2/blob/master/src/encoder.py](https://github.com/openai/gpt-2/blob/master/src/encoder.py?ref=articles.akadata.co.uk)

HuggingFace Tokenizers ByteLevel documentation, including pre-tokenizer semantics and `add_prefix_space`:

[https://huggingface.co/docs/tokenizers](https://huggingface.co/docs/tokenizers?ref=articles.akadata.co.uk)

Unicode 16.0 Character Database, used for the category tables:

[https://www.unicode.org/Public/16.0/ucd/](https://www.unicode.org/Public/16.0/ucd/?ref=articles.akadata.co.uk)

Saphira Tokenizer repository: C implementation, evidence, scripts, frozen fixtures and selftest.

```git
git clone https://github.com/SaphiraLinux/saphira-tokenizer.git
```

Bible corpus VIN: 29 translations, per-book JSON, `legacy-corpus-2025` branch of the Bible translations source repository. The extraction script and per-translation manifest are in the Saphira Tokenizer tree.

OpenWebText training sample, 11.9 GB, used as the claimant’s benchmark corpus:

[https://huggingface.co/datasets/stanford-cs336/owt-sample](https://huggingface.co/datasets/stanford-cs336/owt-sample?ref=articles.akadata.co.uk)