LIVE · TAO
TAO$— SUBNETS VALIDATORS256
Bittensor intelligence updates
Home / AI News/ Bittensor’s Engy Subnet Brings Earth’s…
AI NEWS

Bittensor’s Engy Subnet Brings Earth’s Lowest-Cost Access to Kimi K3

Engy (SN53) just made a 2.8-trillion-parameter frontier model (Kimi K3) run on gaming graphics cards, and they are charging half what the company that built it charges.

Bittensor’s Engy Subnet Brings Earth’s Lowest-Cost Access to Kimi K3

Some cracked folks just made a 2.8-trillion-parameter model run on graphics cards you can buy for a gaming rig, and they are charging half what the company that built the model charges.

That is the short version of what Engy, subnet 53 on Bittensor, pulled off yesterday with Kimi K3. Their achievement questions the assumption the largest AI labs have leaned on for years, which is that owning the best model is what keeps you ahead.

If the weights are public and anyone can serve them, the advantage stops being the model and starts being whoever serves it best for the least money.

That’s why open-source would ultimately win the AI race.

A Public Model Changes the Argument

Moonshot AI released Kimi K3 in late July 2026 and put the weights out in the open under its own license.

Kimi K3 is built as a sparse mixture-of-experts model, which means it holds 2.8 trillion parameters total but only activates about 104 billion for any given token, routing each one through 16 of its 896 experts.

That structure is what lets a model this large stay affordable to run, because you are never firing the whole network at once. On top of that it reads up to a million tokens of context at a time, handles text, images, and video natively, and adds two design pieces called Kimi Delta Attention and Attention Residuals. Together, those give it real strength on long coding jobs, agent workflows, and reasoning tasks.

Where it stands against the field is roughly fourth on the Artificial Analysis Intelligence Index, trailing closed systems like Claude Fable 5 and GPT-5.6 Sol but clearing plenty of others, including earlier Claude Opus versions. It has taken the top spot on some coding and agentic leaderboards, and it is the biggest open-weight model anyone has shipped so far.

Kimi K3’s benchmark result on THE AI RANKINGS

The Hardware Trick That Makes It Cheap

The part worth slowing down on is how Engy achieved it. Instead of the HBM datacenter GPUs everyone fights over, the team served the full official MXFP4 weights across 80 RTX 5090 cards, wired together with GDDR7 memory and plain Ethernet, no exotic interconnect.

Throughput held up, and later tuning pushed decode speeds higher, with the network moving thousands of tokens a second in aggregate. Sidestepping the scarcest chips in the market is most of the reason the price can go where it goes, because self-hosting a model this size otherwise runs into tens of thousands of dollars a month.

Pooling ordinary hardware under incentives is exactly the gap decentralized networks are built to fill, and it’s evident in Engy’s cost savings compared to other models:

Engy vs the field. Courtesy of Mark Jeffrey

On money, the numbers are the pitch. Kimi K3 on Engy runs $1.50 per million input tokens, $7.50 output, and $0.15 cached, against Moonshot’s own $3, $15, and $0.30. GLM-5.2 and a cheaper Qwen option share the same service, which also promises zero data retention and features built for agents.

Proof Baked Into Every Response

Engy is a verified inference subnet, and the verification is not decorative. You reach it through an OpenAI-compatible API at engy.ai, and every response comes back with a cryptographic proof,

TOPLOC activation fingerprints for now, with recompute audits on the way, that confirms the exact model you asked for is the one that answered.

That closes a gap most people never think about, which is that a cheap provider has every reason to quietly swap in a smaller model and pocket the difference. Validators score all of this on-chain through Bittensor consensus, so the guarantee runs on the network rather than on trust.

Early User Feedback: Strong Work, Real Bill

The TAO Daily found one user who hooked Kimi K3 on Engy to a Hermes harness and finished a coding task that Codex had flat-out refused, calling the output strong and the speed acceptable, then flagged that the job cost about $10 for “a simple coding task”.

The user went back to a subsidized Codex setup that hands them near-unlimited tokens for a flat monthly fee ($400/month) across several accounts and resets, and said they would use Engy’s cheaper Qwen going forward.

So the open model completed a task that a closed one turned down, yet pay-per-token still stung next to a subsidy. That last point is the quiet warning that big labs still eat up chunks of AI costs today, which may not be the case long-term. He said: “If labs stop subsidizing tokens we are fucked guys. Enjoy this while it lasts!”

What It Means for TAO

None of this settles whether Engy becomes a durable business, but the direction is worth watching.

The subnet is already building toward enterprise use with operator whitelists for zero-retention guarantees, the kind of assurance a fully open pool cannot make on its own, while leaving that open pool running.

For TAO itself, an inference subnet with genuine demand feeds real emissions and real pull on its alpha token, which is the loop the whole ecosystem is betting on (NFA).

What Engy has shown is that public weights plus an open contest can put pressure on the cost structures that closed labs built their lead on.

Read more on Engy:

Enjoyed this article? Join our newsletter

Get the latest TAO & Bittensor news straight to your inbox.

We respect your privacy. Unsubscribe anytime.

The Daily Dispatch

Enjoyed this article?
Join our newsletter

Get the latest TAO & Bittensor news straight to your inbox — every morning before markets open.

IA
Ige A
Senior Editor

Be the first to comment

Leave a Reply

Your email address will not be published.


*