Actual (SN95) has released toks, its first publicly available software product. The tool converts ordinary text into numerical tokens while matching established tokenizers at significantly higher processing speeds.

Actual (SN95) builds infrastructure that helps AI workloads run efficiently across laptops, consumer GPUs, and smaller edge devices. toks extends that work into tokenization, one of the most common processes inside modern language model systems.
What toks Does
Language models cannot read ordinary text directly, so every sentence must first become smaller units called tokens. Each token receives a numerical identifier, allowing language models to process text during inference, training, and other workloads.

As AI systems handle more text, faster tokenization can reduce delays and improve overall computing efficiency. toks was designed for the coming “quadrillion-token era” of artificial intelligence workloads.
Why toks Stands Out
toks produces identical IDs to Hugging Face Tokenizers 0.23.2 whenever a tokenizer is supported. When perfect compatibility is impossible, toks refuses to load the tokenizer instead of silently producing different results.
On one CPU core, toks runs between 13 and 151 times faster than Hugging Face Tokenizers. It also performs between four and 23 times faster than tiktoken across the compatible benchmarks tested.

For Llama 3 English text, toks reach roughly 270 megabytes per second on a single processor core. That compares with around 30 megabytes for tiktoken and roughly five megabytes for Hugging Face Tokenizers.
With multiple processor cores, toks_par can reach around 1.6 to 1.8 gigabytes per second on larger workloads.

How Actual Made It Faster
Actual built the most performance-sensitive parts using low-level optimizations designed for modern ARM and x86 processors. The core library remains small, avoids external dependencies, and requires no additional memory allocation after loading.
This approach makes toks suitable for integration into inference engines and other performance-sensitive artificial intelligence applications.
Actual says the software supports tokenization methods used by Llama, Gemma, Qwen, DeepSeek, BERT, and GPT-style models. It can also load around 98% of leading text-generation and embedding models available through Hugging Face.
How to Use toks
toks is publicly available through GitHub, with packages supporting Python versions 3.10 through 3.14. C libraries are also available for Linux systems and Apple Silicon devices running modern macOS versions.
The software uses the Business Source License 1.1, allowing extensive free usage below its commercial threshold.
Organizations processing fewer than one quadrillion tokens yearly can freely use, modify, and distribute the software commercially. Larger users can purchase a commercial license, while each release becomes Apache 2.0 licensed after four years.
What This Means for Actual
toks gives Actual its first public demonstration of the software approach behind SN95. SN95 focuses on making demanding AI workloads practical across everyday hardware instead of requiring massive computing infrastructure.
Tokenization represents only one part of that stack, but nearly every language model depends upon it. With toks, Actual is beginning with a small infrastructure layer that could support much larger AI workloads.
➛ Access toks on GitHub Here.
Enjoyed this article? Join our newsletter
Get the latest TAO & Bittensor news straight to your inbox.
We respect your privacy. Unsubscribe anytime.
Enjoyed this article?
Join our newsletter
Get the latest TAO & Bittensor news straight to your inbox — every morning before markets open.





No comments yet — be the first.