Artificial intelligence is advancing rapidly, but one important question remains increasingly difficult to answer. What if an AI model is not truly demonstrating what it knows, but simply remembering what it has seen before?
TPN (SN65) is building a decentralized network that helps optimize AI models while preserving their capabilities under demanding real world constraints. As part of that effort, TPN has developed fluid_knowledge, a private benchmark designed to test whether models can demonstrate genuine knowledge and not rely on familiarity with publicly available questions.

Many widely used benchmarks have been available for years, meaning their questions and answers may already exist within the training data of modern AI models. That creates a growing challenge for anyone trying to determine whether impressive benchmark scores actually represent genuine understanding and reasoning.
fluid_knowledge Is a Stealth Benchmark by Design
fluid_knowledge is designed around the idea that makes conventional benchmark testing much harder to game. TPN (SN65) generates questions across hundreds of subjects and keeps them completely private, and do not publishing its questions.

Those questions are organized into private benchmark versions called epochs, which means models cannot prepare for the exact tests they will eventually encounter. When an evaluation takes place, the model is tested against questions it could not have studied specifically, while the outside world receives only the resulting score.
The aim is to make a strong score mean something more useful, because performance should reflect what a model can actually demonstrate rather than what it has previously memorized.
Fluid Bench: The Infrastructure Behind the Test
Behind fluid_knowledge is Fluid Bench, the infrastructure that allows TPN to create and protect these private evaluations. Fluid Bench manages the process of generating questions, checking their quality, freezing benchmark versions, testing models, and producing verified results.

This becomes particularly important inside TPN, where miners compete to make AI models smaller and more efficient while preserving their useful capabilities. If miners can optimize models against publicly known questions, they may improve benchmark scores without necessarily improving the underlying model.
Private evaluations make that shortcut considerably harder, creating a stronger connection between benchmark performance and genuine model capability.
Conclusion
The future of AI depends not only on building more powerful models, but also on finding better ways to determine what those models actually know. fluid_knowledge addresses one of the biggest weaknesses in traditional benchmarking by keeping its evaluation material private from the models being tested.
For TPN, this makes benchmark integrity part of the optimization process rather than an afterthought added after models have already been trained. As AI increasingly moves into real world applications, reliable evaluation will matter because organizations need to know whether impressive scores translate into dependable performance.
fluid_knowledge is built around one important principle that the best test is one the model never gets the chance to study.
Enjoyed this article? Join our newsletter
Get the latest TAO & Bittensor news straight to your inbox.
We respect your privacy. Unsubscribe anytime.
Enjoyed this article?
Join our newsletter
Get the latest TAO & Bittensor news straight to your inbox — every morning before markets open.





No comments yet — be the first.