Accelerated storage for
AI training

Every second your GPUs wait for data is wasted money. Tigris Acceleration Gateway is a local caching proxy that delivers near-local throughput for AI training, with zero code changes.

85+
Gbps per node
99.4%
GPU utilization
5.7×
Faster warm epoch
0
Code changes

TAG sits inside the training instance, between your code and the storage.

Hot data lands on local NVMe.

A local cache that speaks S3, so the first epoch is the only slow one.

TAG runs as a sidecar on the training instance. Epoch one fetches from Tigris; every epoch after reads local NVMe at disk speed, through the same S3 API your script already uses.

Fig 02a read, epoch one and every epoch after

Near-local throughput
NVMe-speed reads after the first epoch. Training data is served from local disk, not the network.
Zero code changes
Drop-in S3 API compatibility. Point your training script at TAG and it handles the rest.
Intelligent prefetching
Anticipates data access patterns to keep your GPU pipeline full and idle time near zero.

Store once, and cache beside every GPU you rent.

Run a gateway in each zone and cloud. Each one caches locally, and Tigris keeps the data close to all of them.

Fig 03one gateway per zone, one store under all of them

85+Gbps

Measured

Per node, out of the cache

One gateway node serves training reads from local NVMe at more than 85 Gbps, faster than the GPUs behind it can take the data in.

Epochs after the first ran 5.7× faster, and GPU utilization held at 99.4%.

Single node on a 100 Gbps NIC, objects 1 MiB and larger. Warm epoch measured against the first epoch on the same job.

Keep GPUs saturated, with up to 200× the throughput they can take in.

Measured against the same job reading Tigris directly, shard size by shard size.

Fig 04throughput by shard size, direct against warm cache

0×50×100×150×200×17×195×4 MB16×174×8 MB24×200×16 MB37×149×32 MB46×145×64 MBTIGRIS DIRECTTAG, WARM CACHE1× = GPU DEMAND, 134 SAMPLES/S

Accelerate your training pipeline.

TAG is open source and available today, for training workloads on any cloud.