LIVE · TAO
TAO$— SUBNETS— VALIDATORS256
Bittensor intelligence updates
Home / SUBNETS/ IOTA’s Orion-100B Still Proves Frontier…
SUBNETS

IOTA’s Orion-100B Still Proves Frontier Training Doesn’t Need One Giant Data Centre

IOTA’s Orion-100B showed that a 100-billion-parameter model can train across five data centres using single A100s over ordinary internet connections, reaching 30.8% average MFU.

IOTA’s Orion-100B Still Proves Frontier Training Doesn’t Need One Giant Data Centre

What if training a 100-billion-parameter AI model didn’t require putting hundreds of GPUs in one highly specialised data centre?

That is the question behind IOTA’s Orion-100B experiment.

On 24 September, IOTA (SN9) brought attention back to the June 2026 run, in which Macrocosmos trained a 100-billion-parameter model across five US data centres, using individual A100 GPUs connected through ordinary internet infrastructure.

The experiment reached 30.8% average Model FLOPs Utilization (MFU), with a sustained peak of 38%, while achieving roughly 65% of the performance of an equivalent co-located setup.

It processed around 1.1 billion tokens over two days before being stopped because of cost.

This is significant because the training was done across geographically separated machines that were never connected by the kind of specialised networking normally associated with large-scale AI training.

What Orion-100B Did

IOTA scales to more GPUs, and converges more reliably

Macrocosmos, the team behind Bittensor Subnet 9 and IOTA, trained a modified Llama-3.2 architecture scaled to 100 billion parameters.

The model was divided across 16 pipeline-parallel stages, with three replicas, using a total of 48 single A100-80GB GPUs.

Those GPUs were spread across five US data centres and communicated over commodity internet connections rather than a specialised high-bandwidth fabric.

The reported results were:

  • 30.8% average MFU
  • 38% sustained peak MFU over six hours
  • Around 65% of the performance of the equivalent co-located setup
  • Approximately 9,000 tokens per second on average
  • Around 1.1 billion tokens processed before the run was stopped

Macrocosmos described the result as the highest MFU reported for a distributed pipeline-parallel run. The TAO Daily also covered Orion-100B at the time as a major milestone for decentralised AI training.

The source is an X article.

Why Training Over the Internet Is Difficult

Training a model this large is not as simple as connecting a collection of GPUs and pressing start.

In a conventional AI cluster, GPUs can communicate over extremely fast networking. Orion had to work with machines separated across different locations and connected through ordinary internet infrastructure.

IOTA addressed this by using pipeline parallelism.

Instead of requiring every GPU to handle the entire model, the model is divided into stages. Each group of GPUs works on a different part of the model, allowing the workload to be distributed across the network.

The problem is then keeping those stages supplied with data and synchronised despite the distance between machines.

Three parts of IOTA’s architecture were particularly important:

  • ResBM (Residual Bottleneck Model): a lossless activation-compression technique designed to reduce the amount of data transferred between machines.
  • A fault-tolerant peer-to-peer networking protocol: designed to keep training operating despite unreliable connections or individual failures.
  • Distributed synchronisation: used to keep the different parts of the training process coordinated.

The technology was developed through earlier experimentation, including a 1.5-billion-parameter testbed that was used for hundreds of controlled experiments before the team moved toward larger runs.

The Economics Matter Too

There was another reason Orion-100B was significant: the hardware did not need to be packaged into one expensive machine.

Macrocosmos reported that a single A100 could be provisioned for roughly $1.25 per hour. An Orion-style replica using 16 non-colocated A100s came in at around $20 per hour, which the team compared favourably with high-end multi-GPU systems used by some competing approaches.

The broader idea is important. If individual GPUs can contribute effectively without being physically located in the same expensive cluster, distributed networks can potentially turn otherwise underused hardware into useful AI training capacity.

That is one of the central ideas behind IOTA.

Orion Was Only the Beginning

Orion-100B was the opening stage of Project Orion, not the final destination.

Since then, IOTA has moved toward more difficult scenarios, including heterogeneous hardware, interruptible compute and permissionless participation.

One example is Train at Home, which allows Mac users to contribute computing resources to IOTA.

Current Runs on Train At Home 

The idea extends the same principle demonstrated by Orion-100B: AI training capacity does not necessarily have to come from a single centralised cluster.

Enjoyed this article? Join our newsletter

Get the latest TAO & Bittensor news straight to your inbox.

We respect your privacy. Unsubscribe anytime.

The Daily Dispatch

Enjoyed this article?
Join our newsletter

Get the latest TAO & Bittensor news straight to your inbox — every morning before markets open.

IA
Ige A
Editor-in-Chief

No comments yet — be the first.

Leave a Reply