Pareton (SN10) and Engy (SN53) have teamed up to improve how open AI models are served on Bittensor, with the first deployment already running in production.
On October 1, Pareton announced that it was tuning Engy’s inference infrastructure to reduce latency and improve throughput without changing the API that developers already use. Engy later confirmed that a Pareton-tuned version of Qwen3.8-27B was already handling real customer traffic.
Two Subnets, Two Jobs

The two projects operate at different parts of the inference stack. Pareton, which runs Bittensor Subnet 10, focuses on finding better ways to run models by testing configurations involving kernels, batching, quantisation, and serving systems against specific hardware and workloads.
Engy, on the other hand, runs Bittensor Subnet 53 and provides verified inference for open models. Its system uses cryptographic proofs to let customers verify that the model they paid to run was the model used to process their request.
The partnership combines those capabilities by allowing Pareton to optimize the way models are served while Engy continues to provide the customer-facing inference service and verification layer.
The collaboration was also highlighted at Exploit Summit, where the focus was on bringing Pareton’s Subnet 10 optimisation work into Engy’s production inference environment.
To have a deeper background on Pareton’s approach, see below:
What Qwen3.8-27B Does
The model being optimized is Qwen3.8-27B, a dense 27-billion-parameter version from Alibaba’s Qwen team. It is a native vision-language model that can process text, images and video, with capabilities aimed at coding, professional work, research and long-running agentic tasks.
Its agentic capabilities allow it to handle workflows that require several steps, including planning, tool use and computer or browser interaction. That makes the model suitable for applications where an inference request can involve a longer sequence of actions instead of a single question and answer.
Qwen3.8-27B uses a hybrid Gated DeltaNet and gated attention architecture and provides a native context window of around 262,000 tokens, which can be extended toward one million tokens. Its reasoning effort can also be adjusted, while its weights are released under the Apache 2.0 licence.
Quantised versions can also run on consumer-grade GPUs, giving Engy another open model that can be served across a range of hardware configurations.
The Optimisation Gains
Pareton’s reported tests show what its optimisation layer can change at the serving level. On the same hardware and workload, its top Qwen3.8-27B configuration produced 23% to 54% higher output throughput than the baseline when measured with NVIDIA’s AIPerf tool across multiple loads.

Higher throughput means more tokens can be generated from the available compute over the same period, which can improve the economics of running inference because each generated token carries a compute cost.
The result is therefore relevant to both speed and cost. If the production gains seen in testing carry across Engy’s broader workload, the network can serve more inference without requiring a proportional increase in compute.
What Users Get
The optimization is designed to stay behind Engy’s existing API, so developers do not need to rebuild their integrations to use the tuned model. The change takes place in the serving layer while Engy continues to handle the inference request and its verification process.
For Engy customers, that translates into four practical changes:
- Higher throughput: based on Pareton’s reported 23% to 54% improvement over baseline for Qwen3.8-27B.
- Lower latency: as the optimisation is designed to make inference respond faster.
- Potentially lower serving costs: because more tokens can be produced from the same compute resources.
- No API changes: allowing existing applications to continue using Engy’s current interface.
Engy has already confirmed that the tuned Qwen3.8 configuration is processing production requests, so the partnership has moved beyond testing into live customer workloads.
From Optimization to Production
This partnership gives Pareton a direct workload on which to continue its optimisation process. Instead of tuning against an isolated benchmark, the system can work with the traffic patterns generated by an active inference network such as Engy.
That creates a feedback loop in which better inference can support faster service and greater usage, while increased production traffic gives Pareton more real workloads to optimise. Engy described this cycle as faster inference driving more usage, with that usage helping fund further tuning.
For Bittensor, the partnership shows how specialised subnet capabilities can connect at the infrastructure level. Pareton brings the optimization work, Engy brings verified production inference, and Qwen3.8-27B is the first model where the two are now working together in live traffic.
Enjoyed this article? Join our newsletter
Get the latest TAO & Bittensor news straight to your inbox.
We respect your privacy. Unsubscribe anytime.
Enjoyed this article?
Join our newsletter
Get the latest TAO & Bittensor news straight to your inbox — every morning before markets open.





No comments yet — be the first.