Training an AI model is not about feeding it more data and giving it more computing power. As models become smarter, finding the examples that can still teach them something becomes increasingly difficult.
That is the problem Reliquary (SN81) is trying to solve on Bittensor. It opens the search for valuable training examples to independent miners using their own hardware.
The network then verifies those contributions and are rewarded when their work helps improve the model.
What Is Reliquary (SN81)?
Reliquary (SN81) is a decentralized market for reinforcement learning training, where contributors search for prompts that still challenge an AI model, generate possible answers, and submit the results for verification.

The network evaluates these contributions for legitimate and useful learning signals before incorporating selected examples into the training process and rewarding contributors for valuable work.
In simple terms, Reliquary creates a marketplace around the model’s learning frontier, where the most valuable prompts sit between being too easy and too difficult.
Why Does Reliquary (SN81) Exist?
Reinforcement learning becomes less efficient when a model becomes increasingly capable. Many prompts eventually become either too easy or too difficult to produce useful learning signals.
Reliquary focuses on this problem by identifying the narrow zone where training can still produce meaningful improvement.
1. A prompt is too easy when the model consistently produces correct answers.
2. A prompt is too difficult when the model consistently produces incorrect answers.
3. A prompt is valuable when different attempts produce mixed results.
4. That variation gives reinforcement learning a signal about what the model should improve.
This matters because generating thousands of responses can consume significant computing resources without necessarily improving the model.
Reliquary therefore moves the search for useful prompts outside the central training operation. Miners compete to find the examples that offer the strongest remaining learning opportunities.
How Does Reliquary Work?
The process turns AI training into a division of labour between miners, validators, and the trainer.

1. Miners Search for Useful Prompts
Miners download the latest public model checkpoint and search for prompts involving areas such as mathematics and coding. They then generate groups of 16 responses for each selected prompt using their own computing resources.
2. Validators Check the Submissions
Validators examine submitted generations to determine whether miners genuinely produced the claimed work.
They verify the declared model, check reward calculations, measure response variation, and remove groups with weak learning signals.
3. Useful Groups Enter Training
Groups that pass validation become part of the training batch. The trainer then uses those examples to update the model and produces a new checkpoint for the network.
That new checkpoint gives miners a stronger model to challenge, effectively moving the learning frontier forward.
4. Miners Earn for Useful Contributions
Miners are rewarded based on the value of the training groups they provide. This creates a different incentive from producing the largest quantity of valuable responses. Miners are rewarded for finding solvable prompts, generating useful responses, improving inference efficiency, and identifying areas for model improvement.
Since participation is permissionless, anyone with the required hardware can compete.
4B: What Has Reliquary Trained?
Reliquary’s first major research result arrived with Reliquary-4B, a model trained through the subnet’s decentralized reinforcement-learning system.

The project reports the following figures from the training run:
1. Starting model: Qwen3-4B-Base
2. Training approach: Pure reinforcement learning without supervised fine-tuning
3. Training updates: 12,852
4. Training duration: Approximately 15 days
5. Rollouts supplied: Approximately 6.6 million
6. Trainer-generated rollouts: Zero
The final model recorded substantial improvements on held-out mathematics and coding tests. The project also reports improvements on external benchmarks like AMC23, AIME, HumanEval+, and MBPP+ that were not used during training.

The model weights are publicly available on Hugging Face alongside a technical report describing the training run.
The Bigger Picture
Reliquary’s idea is to let a network of independent miners compete to find what an AI model still needs to learn.
Instead of paying for training data simply because it exists, the system attempts to reward data that produces measurable learning value.
Reliquary-4B provides an early demonstration that this decentralized loop can train a real model and deliver significant improvements.
Enjoyed this article? Join our newsletter
Get the latest TAO & Bittensor news straight to your inbox.
We respect your privacy. Unsubscribe anytime.
Enjoyed this article?
Join our newsletter
Get the latest TAO & Bittensor news straight to your inbox — every morning before markets open.





No comments yet — be the first.