LIVE · TAO
TAO$— SUBNETS— VALIDATORS256
Bittensor intelligence updates
Home / SUBNETS/ Chutes (SN64)’ Kappa Matches Llama…
SUBNETS

Chutes (SN64)’ Kappa Matches Llama 3.2 1B on 94% Fewer Tokens

Chutes (SN64) trained Kappa on Bittensor’s decentralized GPU network, matching Llama 3.2 1B across several benchmarks with roughly 94% fewer training tokens.

Chutes (SN64)’ Kappa Matches Llama 3.2 1B on 94% Fewer Tokens

Chutes (SN64), a Bittensor subnet providing decentralized GPU infrastructure for AI inference and training, has completed its latest Kappa experiment.

At the 576 billion token checkpoint, Kappa matched and beat Meta’s Llama 3.2 1B across several standard benchmarks. Those included ARC-C, OpenBookQA and TruthfulQA, while results remained close across several other commonly used evaluation tasks.

The comparison is notable because Meta reports Llama 3.2 1B was pretrained using up to 9 trillion tokens. Kappa reached those results with roughly 94% fewer training tokens, although further evaluation and training are still required.

Kappa Was Trained Across Decentralized GPUs

The important part of Kappa is how Chutes trained the model across its decentralized GPU network. Large AI models are usually trained inside datacenters where expensive GPUs communicate through extremely fast internal connections.

Chutes instead distributes training workloads across independently operated GPUs located in different regions and running different hardware. That setup can include high-end consumer GPUs, reducing dependence on the tightly connected infrastructure used by major AI labs.

Kappa therefore tests whether competitive models can be trained across permissionless hardware without relying on one centralized computing cluster. Chutes also estimates that Kappa’s training cost per token was roughly 90% lower than conventional approaches.

Kappa Uses a More Efficient Model Design

Part of Kappa’s efficiency comes from combining traditional attention with a newer approach known as linear attention. Standard attention becomes increasingly expensive as context grows because models must process relationships across larger amounts of information.

Kappa uses Gated DeltaNet-2 (GDN2) alongside other attention layers, creating a hybrid architecture designed to handle memory more efficiently. The aim is to retain the strengths of transformer models while reducing the computing and memory required.

The Model Is Fast During Inference

Kappa performed well during inference, reaching around 23,000 tokens per second on a single Nvidia RTX 5090. Another setup reached roughly 28,000 tokens per second, while prefill performance approached 70,000 tokens per second.

Those early results showed Kappa could remain efficient when running on high-end consumer hardware after training. Chutes (SN64) also plans to release llama.cpp support, mobile inference tools and additional training code for the architecture.

The Training Run Also Exposed Problems

Kappa’s training run also exposed several problems that Chutes is already working to fix in future versions. One issue affected the model’s internal memory system, creating instability and occasional gradient spikes across the training network.

The team found problems with slower machines submitting data too late and with how training samples were packed. Chutes (SN64) is now adjusting those areas while reconsidering how some architectural layers should be arranged in future models.

Chutes Is Already Testing a New Architecture

Chutes is now testing Erase-then-Delta Attention (EDA) as a possible replacement for GDN2. The newer design separates how information is forgotten and written, which could improve stability during future training runs.

Early tests suggest EDA may produce better training loss while avoiding some problems seen during Kappa’s latest experiment. Chutes is considering a wider model and other architectural changes aimed at improving reasoning performance.

Parallax Makes Distributed Training Possible

Kappa relies on Parallax, Chutes’ framework for training AI models across GPUs located in different regions. Instead of constantly moving large amounts of data, Parallax splits model components across participating machines more efficiently.

That allows GPUs with different hardware and connection speeds to contribute without relying on hyperscale networking infrastructure. Earlier Parallax experiments trained multi-billion-parameter models across several countries at roughly $11 per billion training tokens.

Chutes Is Moving Beyond AI Inference

Chutes (SN64) first gained attention across Bittensor as a decentralized inference network powered by independently operated GPU infrastructure. That system gives users access to open AI models while miners provide the hardware needed to run them.

Parallax and Kappa now extend Chutes into model training, showing how the same network can support development work. The goal is to make Bittensor infrastructure useful not only for serving models, but also for building them.

Why Kappa Matters for Bittensor

Kappa does not yet prove that decentralized networks can replace the massive training clusters used by major AI companies. The results remain preliminary, and broader evaluations will be needed before stronger comparisons can be made.

Still, the experiment shows that serious model training can happen across independently operated hardware in different locations. For Bittensor, Kappa adds another example of subnets producing measurable technical work beyond emissions and token incentives.

Enjoyed this article? Join our newsletter

Get the latest TAO & Bittensor news straight to your inbox.

We respect your privacy. Unsubscribe anytime.

The Daily Dispatch

Enjoyed this article?
Join our newsletter

Get the latest TAO & Bittensor news straight to your inbox — every morning before markets open.

IA
Ige A
Editor-in-Chief

No comments yet — be the first.

Leave a Reply