ReadyAI (SN33) rolled out Skill Coverage Evaluation (v2.37.74), a decentralized system where miners generate AI agent skills alongside the tests needed to verify them.
Instead of relying on a single model to produce a skill and hoping it works, SN33 has six miners independently create competing implementations for each requested capability.
Each submission includes the skill, a test plan, and assertions covering the semantic area defined by the validator, with rewards based on how effectively those tests cover the skill.
The team calls this Proof of Task, where the work ships with its own evidence of correctness; the update is mandatory for all validators and miners.
How Proof of Task Works
The mechanism turns every skill generation request into a competitive verification challenge.
1. Six miners compete per request: Each independently produces a complete skill implementation.
2. Every submission includes three components: A working skill, a TDD-style test plan, and test assertions covering the capability the skill claims to deliver.
3. Rewards depend on test coverage: Validators evaluate whether the miner’s tests meaningfully cover the capability being tested, not just whether the skill looks plausible.
4. The full loop is Generate → Test → Measure Coverage → Compare → Reward: Skills that cannot prove they work do not earn.
Validators no longer compare generated text against reference text. They check whether the accompanying tests genuinely verify the capability the skill implements. Miners get paid for both quality of the work and quality of the proof.
Why This Matters: The Task Economy
The largest emerging market in AI is long-run task completion, where every task carries a prompt, an environment, and a verifier that determines pass or fail.
1. Frontier labs are reportedly spending over $1B a year on tasks.
2. Mercor built a ~$2B revenue business supplying them.
3. Epoch AI puts the going rate at $200 to $2,000 for a single well-verified task.
SN33’s lane is where correctness is testable, because tests ship with the skill, a library of generated skills carries its own evaluation built in.
The next generation of AI agents will need skills that can be verified, and SN33 is positioned to be the network that authors the tasks AI gets measured against.
From Skill Generation to Skill Verification
While producing a plausible-looking SKILL.md file is easy, determining whether it captures the intended capability is much harder. SN33 attacks the problem at generation time rather than after the fact.
1. Verification happens during creation, not after: Every skill comes with the tests that prove it works.
2. Six competing implementations per request: Verification-quality competition rather than a single model output.
3. Executable knowledge, not just structured knowledge: SN33’s architecture now produces skills that can be run and tested rather than just read.
4. The open RAG-ready skill dataset now grows with test-backed entries: Every new skill added to the library ships with its own proof.
Where This Points
Ready AI’s move from generation to verification puts SN33 at the specific intersection every serious agent framework is heading toward.
Anthropic’s SKILL.md pattern proved that skills are the future of agent capability. What nobody solved cleanly was how to trust that any given skill does what it claims.
SN33 answers that by pairing every generated skill with its own executable verification, then paying miners based on the quality of both halves.
For any team building on top of agent frameworks that consume skills, this is the specific subnet worth watching for verified skill libraries. Testnet is live now, and mainnet arrives August 18 at 2:00 PM UTC.
Enjoyed this article? Join our newsletter
Get the latest TAO & Bittensor news straight to your inbox.
We respect your privacy. Unsubscribe anytime.
Enjoyed this article?
Join our newsletter
Get the latest TAO & Bittensor news straight to your inbox — every morning before markets open.





Be the first to comment