Architecture
Sharded inference
How a model too big for one GPU is served by many.
A large model is a stack of layers. Sharded inference splits that stack into contiguous blocks and gives each block to a different machine. A request passes through the machines in order, and no single contributor needs to hold the whole model.
Swarms
Contributed GPUs are pooled into swarms. A swarm serves one open model by covering all of its layers across its members. There is no datacenter and no content-policy layer in between.
Why sharding alone is not private
The prompt never has to exist in one place, which is often mistaken for privacy. But the activations each node receives are a deterministic function of the prompt and the public weights. With the weights in hand, an operator can work backwards to the tokens that produced them.
In PRINET's measurements, a sharded multi-node setup leaks 94–100% of the prompt from a single node's view. The entry and exit nodes, closest to the tokens, see the most.