LIVE · TAO
TAO$— SUBNETS— VALIDATORS256
Bittensor intelligence updates
Home / AI News/ toks, Actual (SN95)’s Tokenizer, Is…
AI NEWS

toks, Actual (SN95)’s Tokenizer, Is Up to 151x Faster Than Hugging Face

Actual (SN95) launches toks, a new tokenizer up to 151x faster than Hugging Face, built for efficient AI workloads across everyday computing hardware.

toks, Actual (SN95)’s Tokenizer, Is Up to 151x Faster Than Hugging Face

Actual (SN95) has released toks, its first publicly available software product. The tool converts ordinary text into numerical tokens while matching established tokenizers at significantly higher processing speeds.

Computers Console on Actual (SN95)

Actual (SN95) builds infrastructure that helps AI workloads run efficiently across laptops, consumer GPUs, and smaller edge devices. toks extends that work into tokenization, one of the most common processes inside modern language model systems.

What toks Does

Language models cannot read ordinary text directly, so every sentence must first become smaller units called tokens. Each token receives a numerical identifier, allowing language models to process text during inference, training, and other workloads.

How LLM Tokenization Works

As AI systems handle more text, faster tokenization can reduce delays and improve overall computing efficiency. toks was designed for the coming “quadrillion-token era” of artificial intelligence workloads.

Why toks Stands Out

toks produces identical IDs to Hugging Face Tokenizers 0.23.2 whenever a tokenizer is supported. When perfect compatibility is impossible, toks refuses to load the tokenizer instead of silently producing different results.

On one CPU core, toks runs between 13 and 151 times faster than Hugging Face Tokenizers. It also performs between four and 23 times faster than tiktoken across the compatible benchmarks tested.

Tokenizer Speed Rate on One GPU Core

For Llama 3 English text, toks reach roughly 270 megabytes per second on a single processor core. That compares with around 30 megabytes for tiktoken and roughly five megabytes for Hugging Face Tokenizers.

With multiple processor cores, toks_par can reach around 1.6 to 1.8 gigabytes per second on larger workloads.

toks on Multiple Processor Cores

How Actual Made It Faster

Actual built the most performance-sensitive parts using low-level optimizations designed for modern ARM and x86 processors. The core library remains small, avoids external dependencies, and requires no additional memory allocation after loading.

This approach makes toks suitable for integration into inference engines and other performance-sensitive artificial intelligence applications.

Actual says the software supports tokenization methods used by Llama, Gemma, Qwen, DeepSeek, BERT, and GPT-style models. It can also load around 98% of leading text-generation and embedding models available through Hugging Face.

How to Use toks

toks is publicly available through GitHub, with packages supporting Python versions 3.10 through 3.14. C libraries are also available for Linux systems and Apple Silicon devices running modern macOS versions.

The software uses the Business Source License 1.1, allowing extensive free usage below its commercial threshold.

Organizations processing fewer than one quadrillion tokens yearly can freely use, modify, and distribute the software commercially. Larger users can purchase a commercial license, while each release becomes Apache 2.0 licensed after four years.

What This Means for Actual

toks gives Actual its first public demonstration of the software approach behind SN95. SN95 focuses on making demanding AI workloads practical across everyday hardware instead of requiring massive computing infrastructure.

Tokenization represents only one part of that stack, but nearly every language model depends upon it. With toks, Actual is beginning with a small infrastructure layer that could support much larger AI workloads.

➛ Access toks on GitHub Here.

Enjoyed this article? Join our newsletter

Get the latest TAO & Bittensor news straight to your inbox.

We respect your privacy. Unsubscribe anytime.

The Daily Dispatch

Enjoyed this article?
Join our newsletter

Get the latest TAO & Bittensor news straight to your inbox — every morning before markets open.

IA
Ige A
Editor-in-Chief

No comments yet — be the first.

Leave a Reply