Almanac (SN41) has opened the beta for its second incentive mechanism, designed to make its Alma agent better at forecasting real-world events.
After six months of onboarding traders to its terminal, the team found that 58% cited the same issues with existing AI: hallucinations, stale information, and overconfidence.
Alma was built to solve those problems with live news and social sentiment tailored for trader reasoning, while the new mechanism turns miners into its continuous training pipeline. The beta launches with 5% of emissions allocated to forecasting and 95% remaining with the trading markets mechanism until load testing is complete.
Why Frontier Models Fall Short
The research behind Almanac’s decision to build Alma confirms what traders already know from using AI in real trading conditions.

1. Hallucinations persist: A 2026 replication study found frontier models invent nonexistent library names between 4.6% and 6.1% of the time. The floor has barely moved since 2025.
2. Recent claims break the models: RealFactBench found GPT-4o’s F1 score drops from 70 on pre-2024 claims to under 49 on post-2025 ones. Competent on settled facts, barely better than chance on new ones.
3. Confidence stays high when accuracy is chance: Frontier models are systematically overconfident on scientific forecasting, and giving them the missing background does not fix it.
4. The newest models are more uniform, not more reliable: They all fail in roughly the same way now.
Frontier models are not built to forecast world events, which is the specific gap Almanac closes.
How the Mechanism Works and Why Research Quality Is the Real Game
Miners build forecasting agents that compete on real-world predictions, but the real output is the research behind every forecast.

1. Agents ship as single Python scripts, with complete freedom over architecture.
2. Each agent forecasts real-world events, starting with geopolitical and macro questions from Almanac and Polymarket.
3. Every prediction includes a probability, reasoning, and supporting research.
4. Forecasts are scored on accuracy, calibration, edge, and research quality.
5. The edge comes from better research, not better headlines.
6. Every submission builds a verified research corpus tied to real-world outcomes.
Research quality is difficult to measure directly, so Almanac uses forecasting accuracy as its proxy.
Where This Leads and How to Mine Today
The research corpus is the real product. Almanac plans to fine-tune a reasoning model on miner-generated forecasts, then pair it with the best-performing miner harness to power Alma:
1. Alma keeps live news and social sentiment.
2. She gains a reasoning model trained on the miner corpus.
3. Users get calibrated probabilities backed by verified research.
4. The subnet is fully open source on GitHub.
5. LLMs make onboarding easier: feed the repo to an AI assistant and have it explain the code or help build an agent.
6. The beta allocates 5% of emissions to forecasting and 95% to trading markets while load testing continues.
With modern LLMs and agent frameworks, building a subnet miner has never been more accessible.
The Loop That Funds Itself
Almanac’s second mechanism reframes what a Bittensor subnet can build. Instead of scoring miners on a fixed task and paying for the output, SN41 scores them on a moving target and uses the aggregate output to train the next version of the product itself.
With this, Alma gets better, the corpus gets richer, traders get sharper research, and miners get paid for producing exactly what the team needs to keep improving her.
For anyone watching how mining incentives translate into product improvement over time, this is one of the sharper answers the ecosystem has shipped.
Enjoyed this article? Join our newsletter
Get the latest TAO & Bittensor news straight to your inbox.
We respect your privacy. Unsubscribe anytime.
Enjoyed this article?
Join our newsletter
Get the latest TAO & Bittensor news straight to your inbox — every morning before markets open.





Be the first to comment