LIVE · TAO
TAO$— SUBNETS VALIDATORS256
Bittensor intelligence updates
Home / SUBNETS/ SN81-Trained Reliquary-4B Beat Baseline Models…
SUBNETS

SN81-Trained Reliquary-4B Beat Baseline Models With 6.6 Million Rollouts

Reliquary (SN81) trains Reliquary-4B through decentralized reinforcement learning, using 6.6 million miner-generated rollouts to improve mathematics and coding performance.

SN81-Trained Reliquary-4B Beat Baseline Models With 6.6 Million Rollouts

Training an AI model is not about feeding it more data and giving it more computing power. As models become smarter, finding the examples that can still teach them something becomes increasingly difficult.

That is the problem Reliquary (SN81) is trying to solve on Bittensor. It opens the search for valuable training examples to independent miners using their own hardware.

The network then verifies those contributions and are rewarded when their work helps improve the model.

What Is Reliquary (SN81)?

Reliquary (SN81) is a decentralized market for reinforcement learning training, where contributors search for prompts that still challenge an AI model, generate possible answers, and submit the results for verification.

Reliquary (SN81)’s Dashboard

The network evaluates these contributions for legitimate and useful learning signals before incorporating selected examples into the training process and rewarding contributors for valuable work.

In simple terms, Reliquary creates a marketplace around the model’s learning frontier, where the most valuable prompts sit between being too easy and too difficult.

Why Does Reliquary (SN81) Exist?

Reinforcement learning becomes less efficient when a model becomes increasingly capable. Many prompts eventually become either too easy or too difficult to produce useful learning signals.

Reliquary focuses on this problem by identifying the narrow zone where training can still produce meaningful improvement.

1. A prompt is too easy when the model consistently produces correct answers.

2. A prompt is too difficult when the model consistently produces incorrect answers.

3. A prompt is valuable when different attempts produce mixed results.

4. That variation gives reinforcement learning a signal about what the model should improve.

This matters because generating thousands of responses can consume significant computing resources without necessarily improving the model.

Reliquary therefore moves the search for useful prompts outside the central training operation. Miners compete to find the examples that offer the strongest remaining learning opportunities.

How Does Reliquary Work?

The process turns AI training into a division of labour between miners, validators, and the trainer.

How SN81 Operates

1. Miners Search for Useful Prompts

Miners download the latest public model checkpoint and search for prompts involving areas such as mathematics and coding. They then generate groups of 16 responses for each selected prompt using their own computing resources.

2. Validators Check the Submissions

Validators examine submitted generations to determine whether miners genuinely produced the claimed work.

They verify the declared model, check reward calculations, measure response variation, and remove groups with weak learning signals.

3. Useful Groups Enter Training

Groups that pass validation become part of the training batch. The trainer then uses those examples to update the model and produces a new checkpoint for the network.

That new checkpoint gives miners a stronger model to challenge, effectively moving the learning frontier forward.

4. Miners Earn for Useful Contributions

Miners are rewarded based on the value of the training groups they provide. This creates a different incentive from producing the largest quantity of valuable responses. Miners are rewarded for finding solvable prompts, generating useful responses, improving inference efficiency, and identifying areas for model improvement.

Since participation is permissionless, anyone with the required hardware can compete.

4B: What Has Reliquary Trained?

Reliquary’s first major research result arrived with Reliquary-4B, a model trained through the subnet’s decentralized reinforcement-learning system.

Reliquary-4B on Hugging Face

The project reports the following figures from the training run:

1. Starting model: Qwen3-4B-Base

2. Training approach: Pure reinforcement learning without supervised fine-tuning

3. Training updates: 12,852

4. Training duration: Approximately 15 days

5. Rollouts supplied: Approximately 6.6 million

6. Trainer-generated rollouts: Zero

The final model recorded substantial improvements on held-out mathematics and coding tests. The project also reports improvements on external benchmarks like AMC23, AIME, HumanEval+, and MBPP+ that were not used during training.

Reliquary-4B on Benchmarks

The model weights are publicly available on Hugging Face alongside a technical report describing the training run.

The Bigger Picture

Reliquary’s idea is to let a network of independent miners compete to find what an AI model still needs to learn.

Instead of paying for training data simply because it exists, the system attempts to reward data that produces measurable learning value.

Reliquary-4B provides an early demonstration that this decentralized loop can train a real model and deliver significant improvements.

Enjoyed this article? Join our newsletter

Get the latest TAO & Bittensor news straight to your inbox.

We respect your privacy. Unsubscribe anytime.

The Daily Dispatch

Enjoyed this article?
Join our newsletter

Get the latest TAO & Bittensor news straight to your inbox — every morning before markets open.

IA
Ige A
Editor-in-Chief

No comments yet — be the first.

Leave a Reply