
11 years of making decentralized storage faster
Sia has been in the storage business since 2015. A few years after launch, a multi-terabyte upload could take a week and still stall before finishing. Today, our new client SDK can handle long uploads while keeping throughput steady. Getting there meant learning how to keep transfers fast across a large distributed network without routing every upload through a centralized gateway.
Centralized gateways simplify this problem by giving clients a single known endpoint, then caching and distributing the data across the decentralized network. With Sia, we took a different path. Our client SDKs run wherever they're needed—in a browser, on a phone, or on a laptop—and transfer data directly to storage providers. A small coordination service we call an indexer handles everything that isn't a user upload or download. Since uploads and downloads don’t pass through the indexer, it needs much less bandwidth than a typical gateway that relays data.
Sia's basic erasure-coding scheme hasn't changed since 2018. Uploads are broken down into four steps:
- Files are split into chunks
- Each chunk is encrypted
- Each chunk is erasure-coded into 30 unique shards
- Each shard is uploaded to a different storage provider.
Those shards consume 3x the original data size, but any 10 can reconstruct the chunk. Plenty of opportunity for parallelism with Sia's large network of distributed providers. The primary issue is that clients need to keep enough useful work in flight to saturate the connection without letting one slow storage provider stall a chunk that every other storage provider has already finished.
Two benchmarks from 2018
In March 2018 a community member named Mtlynch decided to back up a small DVD collection to Sia as a test. Almost 5TB of data on a 1 Gbps residential connection. It took 231.7 hours. That works out to 45.8 Mbps of file data, 137 Mbps counting redundancy. Somewhere past hour 100, siad (the Sia client at the time) stopped making progress on new files for 17 hours straight. The post ends that section with "I don't have an explanation for what happened."
Three months later another community member, Fornax, ran a similar synthetic load test: 512 files of 10 GB each. Fornax's run averaged 211.13 Mbps including redundancy, 15.43 TB sent to storage providers for 5.12 TB of files, zero crashes. It was terminated at 5.1 TB because siad quietly stopped working "for unknown reasons."
Test | Workload | Throughput | Result |
|---|---|---|---|
Mtlynch, | ISOs and video files | 45.8 Mbps overall | 4.3 TiB over 231.7 hours; a 17-hour stretch without new-file progress |
Fornax, | Synthetic 10 GB files | 122.85 Mbps through the first 1 TB; 119.88 Mbps through the first 2 TB; 70.08 Mbps overall | 5.12 TB across 512 files over 162.35 hours; stopped after uploads ceased progressing |
Let's break down Fornax's run into 24-hour windows.
Hours | Throughput (without redundancy) | Throughput (with redundancy) | Total uploaded |
|---|---|---|---|
0 to 24 | 122.3 Mbps | 367.8 Mbps | 1.32 TB |
24 to 48 | 112.0 Mbps | 337.0 Mbps | 2.53 TB |
48 to 72 | 108.8 Mbps | 327.1 Mbps | 3.71 TB |
72 to 96 | 86.8 Mbps | 261.0 Mbps | 4.64 TB |
96 to 120 | 27.8 Mbps | 83.8 Mbps | 4.94 TB |
120 to 144 | 12.0 Mbps | 36.9 Mbps | 5.07 TB |
144 to 162 | 5.8 Mbps | 19.4 Mbps | 5.12 TB |
The first 4.5 TB took 91 hours. The last 0.6 TB took 71. The client just stopped filling the pipe. Mtlynch was able to utilize about 16% of an 880 Mbps uplink and Fornax about 42% of a 500 Mbps line. In both runs, throughput declined even though bandwidth remained available, and both stalled mid-upload.
Eight years later...
With the launch of Sia Storage, we revisited Fornax's synthetic load test using our new client SDK: 500 files, 10GiB each, 5 concurrent uploads. We ran it on a machine with 8 cores, 32 GiB of RAM, and a 5 Gbps rated uplink. Here's the harness we wrote for this test.
Test | Workload | Throughput | Result |
|---|---|---|---|
| 500 synthetic 10 GiB files | 1,481 Mbps through the first 1 TB; 1,473 Mbps through the first 2 TB; 1,461 Mbps overall | 5.37 TB across 500 files over 8.15 hours; every file completed |
Hourly throughput:
Hours | Throughput (without redundancy) | Throughput (with redundancy) | Total uploaded |
|---|---|---|---|
0 to 1 | 1,490 Mbps | 4,471 Mbps | 0.67 TB |
1 to 2 | 1,468 Mbps | 4,403 Mbps | 1.33 TB |
2 to 3 | 1,459 Mbps | 4,378 Mbps | 1.99 TB |
3 to 4 | 1,451 Mbps | 4,353 Mbps | 2.64 TB |
4 to 5 | 1,467 Mbps | 4,401 Mbps | 3.30 TB |
5 to 6 | 1,465 Mbps | 4,396 Mbps | 3.96 TB |
6 to 7 | 1,450 Mbps | 4,349 Mbps | 4.61 TB |
7 to 8 | 1,452 Mbps | 4,356 Mbps | 5.27 TB |
8 to 8.15 | 1,379 Mbps | 4,136 Mbps | 5.37 TB |
Individual files ranged from 105 Mbps (13.6 minutes) to 736 Mbps (1.9 minutes), averaging 310 Mbps per file with five in flight.
File throughput averaged 1.46 Gbps. When including 3× redundancy, that corresponds to roughly 4.4 Gbps of shard payload, or 88% of the link. Throughput stayed consistent throughout the run, and every file completed.
What changed
We've learned that a transfer is really a long stream of small decisions: which storage provider gets this shard, how many shards are in flight, how long to wait on a slow one before asking someone else, how much memory to hold while waiting. Our SDKs rank and choose storage providers based on their recent failure rate, average throughput, and load. Every completed shard updates the storage provider's history and every shard still in flight counts against it. A fast provider with five shards pending can lose to a slightly slower idle provider instead of continuing to collect work. Concurrency adapts to network conditions, and each transfer has its own concurrency and memory budget instead of a global shared limit.
With dynamic upload concurrency, each transfer starts with a minimal number of shards in flight and doubles that number until more concurrency starts hurting useful throughput. Once concurrency settles, a sustained 25% throughput drop starts reducing concurrency until it stabilizes, then tries increasing it again. Adjustments use aggregate transfer throughput rather than individual storage provider latency so a slow shard isn't mistaken for network congestion. This makes Sia automatically adapt to network conditions rather than relying on static limits.
Every chunk waits on its slowest shard. With 30 providers per chunk, there are often a few that are slower than normal. Instead of blocking and waiting, the SDK sends a replacement shard to a second provider. Racing only starts once every shard has an attempt in flight and the timeout is 1.5x the average of recent shards, so only providers well outside the normal spread get raced. Races stay inside the concurrency budget and never take a slot from new work. This can increase bandwidth slightly, but over long transfers, slow providers are used less frequently.
The same core SDK runs in browsers, on iOS and Android, from Node.js, Flutter, Python, and Go. The Sia Storage app on your phone is talking to the same storage providers your desktop is, over the same protocol with no middleman.
Quickstarts for the SDKs are available on the Developer Portal. You can also try Sia yourself with 50 GB of free storage on Sia Storage.


