Skip to content

TECH PARTNERS

VDURA + Wasabi: Keep Active AI Data Fast and Retained Data Affordable

August 4, 2026
Jen NewmanDIrector of Global Alliances

AI infrastructure teams invest heavily in high-performance storage placed near their GPUs, so training, checkpointing, inference, and other demanding workloads don't get stuck waiting on data.

But over time, that performance storage can take on another job: long-term archive.

Completed checkpoints remain in place. Old dataset versions accumulate. Leftover files from finished experiments sit alongside active workloads, using up the same capacity. Before long, the organization is buying more of its most expensive storage, not because its AI workloads need more performance, but because no one has established where inactive data should go next.

That is an expensive way to retain data, and it becomes a bigger problem as AI data volumes grow. Teams need a lifecycle strategy that keeps active data fast while making retained data cheaper to store, easy to find, and usable when needed. That’s the case for pairing on-prem storage with cloud.

The AI storage problem isn’t only about performance

Performance remains critical. If storage cannot keep pace with compute, GPUs sit idle and training times increase. That said, not every dataset, checkpoint, and model artifact needs to live on the fastest available storage indefinitely.

AI environments generate data continuously:

  • Training datasets are updated and versioned.

  • Checkpoints are created throughout model training.

  • Experiments produce logs, outputs, and derived datasets.

  • Models and supporting artifacts are retained for validation, recovery, or reuse.

  • Completed projects may need to remain available for governance, research, or future development.

Some of this data is active and performance-sensitive. Much of it becomes less active as a project progresses.

When those two categories are treated the same, inactive data consumes capacity intended for active work. Infrastructure teams may respond by purchasing additional GPU-adjacent storage, even when much of the existing capacity is occupied by data that no longer requires high-performance access.

The answer is not to delete everything as soon as a training run ends. Checkpoint and dataset history can be highly valuable. It can support model recovery, comparison, reproducibility, compliance, and future reuse. The goal is to retain more of that value without paying performance-tier prices for every byte.

Define data by lifecycle stage, not by where it first lands

A more sustainable architecture begins by separating the AI data lifecycle into active and retained phases.

Active data may include:

  • Datasets being staged or prepared for training

  • Current training data

  • Recent or operationally significant checkpoints

  • Frequently accessed model artifacts

  • Data supporting active inference or fine-tuning

Retained data may include:

  • Older checkpoints that are no longer part of an active training run

  • Previous dataset versions

  • Completed project data

  • Historical model artifacts

  • Data preserved for governance, recovery or future reuse

There's no fixed rule for when a checkpoint crosses from one category to the other. The dividing line will differ by organization and workload. A checkpoint does not automatically become inactive after a fixed number of days, and not all retained data has the same value.

Organizations can classify data based on factors such as project status, checkpoint age, access frequency, recovery requirements, and retention policy. The important step is to make those decisions intentionally and then translate them into repeatable movement policies.

Without that lifecycle framework, the default is simple but expensive: data lands on the performance tier and stays there.

What makes AI data tiering difficult

Traditional tiering runs into the same problems every time: retrieval takes too long, cloud bills come in higher than expected, the tooling only works with one vendor, and nobody's confident the data that got moved is still complete.

These concerns often lead teams to postpone lifecycle planning and continue expanding performance storage instead.

A successful AI data strategy should provide:

  • High-performance access for active workloads

  • Policy-based movement of inactive data

  • Predictable retention economics

  • S3-compatible access that supports interoperability

  • Clear retrieval and reuse processes

  • Visibility into where datasets, checkpoints, and artifacts reside

  • Protection and governance aligned to the value of the data

The objective is to create a reliable path between performance, retention, and reuse.

A lifecycle architecture with VDURA and Wasabi

VDURA and Wasabi are working together to help AI factories, neoclouds, and enterprise high-performance computing (HPC) teams address this lifecycle challenge.

The VDURA Data Platform supports the performance-intensive phases of the AI data pipeline, keeping active datasets and checkpoints close to GPU infrastructure. Wasabi provides a cloud object storage destination for retained data that no longer requires that same level of performance.

Using S3-compatible interfaces, organizations can build a more flexible movement pattern between the active and retained phases of the lifecycle. This allows teams to free GPU-adjacent capacity for current work while keeping historical data available for future recovery, analysis, or reuse.

The architecture is designed to help organizations:

  • Reserve high-performance capacity for active AI workloads

  • Retain more checkpoint and dataset history

  • Reduce unnecessary expansion of performance storage

  • Establish repeatable policies for moving inactive data

  • Maintain access to historical data without keeping it permanently on the highest-cost tier

  • Budget cloud retention more predictably

Wasabi Hot Cloud Storage does not charge separately for egress or API requests under standard terms, helping organizations estimate retention costs without the request and retrieval fee structure associated with many traditional cloud storage services.

The result is a clearer division of responsibility: performance storage does the work that requires performance, while cloud object storage supports economical long-term retention and reuse.

Keep active data fast and retained data ready

GPU-adjacent storage shouldn’t become permanent storage by default.

As AI environments scale, organizations need to treat data movement as an architectural principle rather than a cleanup exercise triggered by the next capacity shortage. By placing data according to its lifecycle stage, teams can protect performance, manage storage growth, and preserve more of the information created by their AI initiatives.

With VDURA supporting performance-intensive AI workloads and Wasabi providing predictable cloud object storage for retention, organizations can create a more balanced foundation for AI data, one that keeps active data fast and retained data ready for what comes next.

Headed to Ai4 2026? Meet with Wasabi at booth #935, August 4–6 at The Venetian in Las Vegas.

VDURA + Wasabi

VDURA and Wasabi have unveiled a joint reference architecture pairing high-performance AI storage with predictable cloud economics. Get the full details in the press release.

Read More

It's a joint architecture that pairs VDURA's high-performance, GPU-adjacent storage with Wasabi's cloud object storage, giving AI teams a defined path for moving inactive checkpoints and datasets off performance storage without losing access to them.

There's no fixed rule. Organizations typically classify data based on project status, checkpoint age, access frequency, recovery requirements, and retention policy, then apply those thresholds consistently through repeatable movement policies.

No. Retained data is still fully accessible; it's just no longer sitting on the highest-cost, highest-performance tier. Wasabi is built for durability and retrieval, not the same level of GPU-adjacent performance as VDURA, which is why it fits the retained side of the lifecycle rather than the active side.

No. The architecture uses S3-compatible interfaces on both sides, so data movement doesn't require proprietary tooling and doesn't create new dependencies on a single provider.

Wasabi Hot Cloud Storage doesn't charge separately for egress or API requests under standard terms, which makes it easier to estimate retention costs up front instead of budgeting around per-request or per-retrieval fees.

Related article

cloud storage
TECH PARTNERSWhy every neocloud needs a cloud storage strategy

Most Recent

When a cyber attack stops the production line: Lessons from the Fairlife ransomware incident

A 2026 ransomware attack halted Fairlife's US production. See how manufacturers can build cyber resilience that survives the next attack.

GLM-5.2 just changed the ransomware conversation: When AI levels up the attacker

An open-weight AI model called GLM-5.2 is making ransomware attacks faster. Learn why defense in depth is critical to keeping backups recoverable.

Wasabi MCP Beta is live: Your AI agents now have direct access to cloud storage

Wasabi MCP is now in beta. Connect any AI agent to your Wasabi cloud storage with 140+ tools, no custom code, no egress fees, and no API charges. Start building today.

SUBSCRIBE

Storage Insights from the Storage Experts

Storage insights sent direct to your inbox.

Subscribe