Skip to content

AI STORAGE CASE STUDIES

Eight AI Companies. One AI Storage Cloud.

AI TrainingData LakeSurveillance AI AI Archive

AI Training

Frontier-scale AI training

A multicloud platform serving multimodal training datasets to frontier AI labs

Who they are. A large-scale web data acquisition company that operates a distributed network to collect public web content at petabyte scale. The company scrapes, cleans, and curates multimodal datasets, including text, video, image, and audio content, and sells enriched datasets to LLM developers and multimodal AI labs.

The storage problem. Storing and serving scraped multimodal data during peak training delivery cycles requires sustained throughput that most storage platforms can't handle. The data lifecycle requires hot storage for active training alongside an archive tier for data not currently in rotation. Their infrastructure was split across multiple cloud environments, including Azure, which couldn't economically sustain the scale or the throughput variability of their largest workloads.

The Wasabi solution. Compute clusters at third-party GPU provider Nebius pull directly from Wasabi buckets at high throughput during active training delivery windows, moving petabytes of data per day. Because Wasabi doesn't charge egress fees, moving that much data to compute every day costs nothing extra beyond the storage itself.

S3 compatibility meant no pipeline changes were required to make the switch. The company now stores roughly 150 petabytes in Wasabi, with another 20 petabytes in discussion to migrate off Azure. The platform has grown without storage becoming a limiting factor on delivery capacity.

Media

150 PB

stored

40 GBps

sustained throughput

petabytes

moved per day

A frontier AI lab building production-grade text-to-image editing models

Who they are. An AI lab building and licensing image generation models. The company develops both open-weight and proprietary models distributed via API and direct licensing. Enterprise contracts span major creative software platforms and consumer applications. The company competes at the frontier of image generation and has recently expanded into video generation.

The storage problem. The lab had accumulated large amounts of hot training data in Microsoft Azure, and the high costs and constraints on data movement had become unsustainable. As the team evaluated GPU providers outside the hyperscaler’s ecosystem, the egress fees they faced made the move impractical.

The Wasabi solution. The lab needed pricing that would scale gracefully beyond its existing hot training data. A rigorous proof of concept, including ingress and egress throughput testing from a third-party GPU compute cluster, confirmed no bandwidth limitations per account or bucket, and no throttling as workloads scale. The lab also cited Wasabi's responsiveness during the POC and contracting process as a differentiating factor.

Training data is now migrating from Azure to Wasabi Hot Cloud Storage with compute on third-party GPU infrastructure reading directly from Wasabi. The lab moved 50 petabytes of hot training data within the first three months of deployment and are now at 175 PBs with room to grow. The move eliminated hyperscaler dependency for primary training data and meaningfully improved storage economics.

AI data management

A generative AI company training foundational 3D models for product design, manufacturing, and gaming

Who they are. A San Francisco-based generative AI startup that converts text and image inputs into high-resolution, 3D-printable models. The company is training a foundational 3D generation model, which requires large-scale 3D asset datasets with high read and write frequency during active training cycles. Designers and engineers across gaming, VFX, product design, manufacturing, and architecture use the platform.

The storage problem. 3D training data makes unusual demands on storage: high polygon meshes, texture maps, simulation outputs, and iterative geometry variations produce large objects that the team reads and writes constantly during active training.

The Wasabi solution. The company runs its training data on Wasabi at a cost that fits its scale, with S3 compatibility that required no changes to existing tooling. Wasabi charges no fees on the high-frequency reads and writes that training demands, so the team can iterate without API request costs on top of what they’re already paying for storage capacity.

AI

90 days

to 50 PB deployment

175 PB

migrated to date

Data Lake

Consolidating AI training datasets

Curating enterprise training datasets across multiple clouds

Who they are. A Silicon Valley-backed startup that builds multimodal data curation and model-training tooling for enterprise customers, including Fortune 500 companies. The platform transforms raw multimodal data (images, video, text, and sensor data) into optimized, AI-ready training datasets. The company runs its primary compute on Alibaba Cloud in Japan and distributes data across hyperscaler storage and Wasabi.

The storage problem. The company had spread its training data across four cloud environments (AWS, Azure, GCP, and Alibaba Cloud) without a cost-effective, cross-environment object storage layer. Hyperscaler pricing made it expensive to store inactive training data in the cloud between active use cycles, and the team needed to rationalize storage spend without disrupting existing pipelines.

The Wasabi solution. Cost drove the decision. Pay-as-you-go pricing matched the economics of a growth-stage startup scaling its first petabyte, and S3 compatibility meant the switch required no pipeline changes. The company now offloads inactive and semi-active training data from its Alibaba Cloud compute environment to Wasabi, keeping it accessible without retrieval delays or unpredictable fees. It signed on in late 2025 and has stored a full petabyte with Wasabi within its first year, with more expected to move over as the relationship grows.

A petabyte-scale shared data lake for robot foundation models

Who they are. A Japan-based non-profit building a Robot Data Ecosystem that acquires and integrates real-world operational data from retail, manufacturing, and logistics environments. The organization uses this data to enable cross-domain learning for foundation model development and real-world robot deployment. At the core of the ecosystem sits a petabyte-scale data hub that distributes raw data across GPU clusters and cloud environments.

The storage problem. The data hub continuously uploads dozens of terabytes and depends on extensive API activity — uploads, downloads, differential backup checks, and object listings — as its core usage pattern. Variable hyperscaler pricing due to burdensome API request fees didn't fit a non-profit budget. Beyond price predictability, the team required a platform that could guarantee integrity throughout the entire data lifecycle, as data loss anywhere would render it unusable for training.

The Wasabi solution. No transfer fees and no API fees made the difference, since high-frequency API operations are exactly how this organization uses storage day to day. A proof-of-concept validated Wasabi upload speed and confirmed the data remained unaltered through full upload, storage, and retrieval cycles. Wasabi has continued to perform reliably as usage scales dynamically. The organization projects savings of more than JPY 100 million compared to alternative providers and has stored more than 2.2 petabytes with Wasabi, with that figure still growing.

Read the Case Study
helper robot on laptop

JPY 100M+

projected savings

2.2+ PB

stored and growing


Wasabi is operating flawlessly despite dynamic usage that continuously scales up. We can trust the data integrity.

Engineering Project Lead

Surveillance AI

Video AI at scale

Who they are. An AI platform that converts raw surveillance footage into structured, queryable data. The platform supports loss prevention, customer behavior analysis, and operational compliance for retail and enterprise customers globally, many of whom operate under strict data sovereignty requirements.

The storage problem. As the platform scaled, its storage infrastructure couldn't keep pace. Data loss events disrupted the customer experience and put compliance records at risk, which was a serious liability for retail customers already operating under strict data compliance requirements. Limited regional storage options made those residency requirements even harder to satisfy, and unpredictable cloud fees made the whole solution difficult to budget for. On top of all that, slow access to video and training data kept stalling the AI retraining workflows the product depends on to keep improving.

The Wasabi solution. The company moved its tagged footage and AI training data to Wasabi Hot Cloud Storage, replacing scattered, unreliable infrastructure with one dependable layer. Data loss events stopped. Low-latency access ended the retraining bottleneck, so the model keeps improving on schedule. Wasabi's regional availability satisfied data residency requirements without complicated workarounds, and predictable pricing gave the company a lower and more stable cost baseline to build customer pricing on top of.

Read the Case Study

Having low-latency storage was critical for our AI model training. With Wasabi, accessing training data is no longer a bottleneck.

Co-Founder and CTO

Powering edge AI surveillance across a global fitness network

Who they are. An AI-powered SaaS platform that provides global fitness franchises with unified business and operations management tools across marketing, payments, support, digital displays, surveillance, and analytics. Edge devices running on mini-PCs with NVIDIA GPUs perform real-time facial recognition and metadata tagging at thousands of gym locations worldwide. 

The storage problem. The company originally built its platform on AWS S3, with a separate storage bucket for each customer across its franchise network. Surveillance footage and metadata volume drove cloud fees up fast as the business grew, and customers experienced timeouts and lag pulling up archived clips and freeze frames, undercutting the product user experience. On top of that, calculating storage usage for client billing was expensive and time-consuming. The team needed real savings and predictable pricing without requiring major rewrites or any perceived downtime to the platform. 

The Wasabi solution. The company migrated its entire storage environment (hundreds of terabytes) to Wasabi in a single overnight cutover with zero downtime and no disruption to customer service. Edge devices now upload surveillance footage and AI-enriched metadata directly to Wasabi Hot Cloud Storage, including vectorized footage data that powers real-time search. Wasabi returns matching clips in seconds, handling high volumes of metadata and video files with no performance issues. Cloud-native programmatic access lets the company configure new Wasabi buckets in local regions in under a minute, meeting data residency requirements as it expands into new markets. Flat-rate, predictable pricing did away with the billing-calculation problem entirely and let the company build simple, profitable storage plans for its customers. 

Read the Case Study
surveillance

4-6X

cloud storage cost reduction

20X

cheaper billing calculations vs. AWS

<1 min

bucket setup per region


Wasabi eliminated our direct-from-device upload timeouts and small file issues critical for managing AI metadata and media at scale.

CTO

AI Archive

Compliance-grade archive

A secondary archive for massive facial recognition dataset

Who they are. A facial recognition company that maintains a database of billions of publicly sourced facial images used to power an identity resolution platform that governments and police departments use for criminal investigations, victim identification, and public safety operations. 

The storage problem. The company needed a cost-effective secondary archive for a massive 5.6 PB biometric image dataset. A dataset this sensitive and this large can't depend on a single storage environment. If the primary environment were ever compromised, corrupted, or held hostage, the company would have no fallback. They needed a genuinely separate, resilient copy of its data, not just a cheaper place to put a backup.  

The Wasabi solution. Wasabi’s zero-trust cyber-resilience features and low per-terabyte pricing make it the ideal target for the company’s petabyte-scale secondary storage. However, the company is interested in moving more active data from hyperscaler environments to storage providers like Wasabi with compute partners that don’t charge to move data between them. See Wasabi-Megaport partnership.  

facial recognition

5,600 TB

(5.6 PB) archived


Why Wasabi for AI

AI data is always moving, from raw ingest to preprocessing to training clusters and beyond. The hyperscalers tax that movement with API request, egress, and other transaction fees throughout the AI data lifecycle. 

Wasabi is different. We’re the neocloud of S3 object storage, built for workloads that are always growing. And always moving:  

  • No egress fees. Data moves freely between GPU compute providers, cloud platforms, and edge locations—wherever your AI stack happens to need it next.  

  • No API fees. High-frequency reads or writes never become a hidden tax on scale or iteration. 

  • 80% less than the hyperscalers. Flat, predictable pricing ideal for petabyte-scale projects. 

  • S3 compatibility. Zero migration friction 

  • Built for resilience. Data residency, durability, and zero-trust security features without premium pricing. 

ai



Ready to take the next step?

Start a 30-day free trial or set up a call to discuss your specific workload requirements.

Try Free
Contact us