AI
Building a Production RAG Pipeline on Wasabi: From Document Ingest to Continuous AI Improvement
In our previous article, we explored why Wasabi serves as the persistent data repository behind a production Retrieval-Augmented Generation (RAG) application. Establishing that foundation is only the first step.
As large language model (LLM) context windows have expanded, it’s fair to ask whether retrieval is still necessary. A context window is the amount of text a model can hold and work with in a single interaction. Larger context windows allow models to work with substantially more information, but capacity alone doesn’t identify the right enterprise knowledge, control what enters the model's context, or manage the cost of reprocessing large volumes of data. For enterprise workloads, retrieval and long context tend to be complementary, not competing.
RAG has changed too. Production implementations now combine retrieval methods, manage context more deliberately, evaluate retrieval quality, and continuously refine how enterprise knowledge is used. That makes the architecture around RAG increasingly important.
That leads to a core design principle: separate what changes frequently from what should remain persistent. Embedding models, vector databases, LLMs, and orchestration frameworks can each evolve independently. A production architecture should allow those components to change without requiring the enterprise corpus to be re-ingested or reconstructed.
This article walks through what that looks like in practice, from the moment enterprise content lands in Wasabi to the feedback loop that keeps improving the system. With Wasabi as the system of record, processing services, embedding models, retrieval indexes, and inference applications can be updated independently while continuing to work from the same authoritative enterprise corpus.
Stage 1: Establish the enterprise corpus in Wasabi
Every production RAG pipeline begins with enterprise data. That data may include documents, PDFs, images, application logs, knowledge bases, collaboration content, or structured exports from business systems.
Before embedding or retrieval begins, the first objective is to establish an authoritative corpus: the complete body of source content it will draw from, kept separate from anything derived from it later.
A production implementation can use one or more Wasabi buckets as the persistent landing zone for raw content. Ingestion services and AI application workers access Wasabi through regional S3 endpoints using standard S3-compatible tooling and AWS SDKs. Granular IAM policies control which services are authorized to read, write, or manage these underlying objects.
For application-driven ingestion, presigned URLs can provide temporary access to upload or retrieve objects without distributing permanent storage credentials. Large files and datasets can use multipart uploads, allowing individual parts to be transferred independently rather than restarting an entire upload after an interruption.
Once content is stored in Wasabi, downstream AI services can work from that authoritative corpus instead of repeatedly returning to production source systems.
Stage 2: Process, chunk, and enrich without changing the source
Enterprise content is rarely ready for retrieval immediately after ingestion.
Processing services extract text, normalize formatting, remove duplicates, perform optical character recognition (OCR) where necessary, and divide content into chunks suitable for embedding. Metadata and semantic context can then enrich those chunks with information useful for retrieval, governance, and downstream processing.
The original corpus should remain separate from these derived versions. Keeping the original corpus intact ensures it remains a trustworthy reference point, even as derived versions are reprocessed, regenerated, or discarded.
A logical Wasabi object layout might look like this:
Wasabi doesn’t require this schema, and the right prefix strategy depends on the application. A useful starting point is to organize objects around ownership, access, lifecycle, and reprocessing boundaries, such as business unit or data domain, rather than defaulting to a purely chronological hierarchy. Dates can be nested beneath those domains when time is a meaningful retrieval, retention, or processing boundary.
For example, source/finance/2026/ preserves both organizational ownership and time without making the year the primary structure of the corpus. Attributes that do not need to determine an object's location can instead be represented by object tags, metadata, or the application's retrieval metadata. For embeddings, the naming convention should capture enough information to reproduce the embedding set, such as the model and version and, where relevant, the associated chunking or processing configuration.
Processing services read objects from the source dataset, transform them, and write normalized content and chunks back to Wasabi. Object tags or metadata can provide additional classification where it’s useful, while richer semantic information can remain with the processed dataset or retrieval metadata.
This separation allows engineering teams to modify chunking strategies, enrich metadata, or introduce new processing techniques without re-ingesting the original enterprise content.
Wasabi's own Business Intelligence team follows this same architectural pattern, using Wasabi as the persistent store for source and processed content while purpose-built retrieval services are rebuilt from that durable foundation. That internal implementation reinforces the broader design principle: keep enterprise knowledge in Wasabi, and let retrieval and inference layers evolve around it.
Stage 3: Generate embeddings and build a replaceable retrieval layer
Once documents have been prepared, an embedding model converts chunks into numerical representations that capture semantic relationships.
Those embeddings can be persisted in Wasabi before being loaded into a vector database for indexing and similarity search.
The distinction is important: The vector database is the retrieval layer. Wasabi remains the durable system of record for the corpus and reusable AI artifacts.
Keeping the underlying embeddings outside the retrieval index means that engineering teams can rebuild an index, compare vector databases, or introduce a new retrieval platform without modifying the authoritative corpus. If a new embedding model becomes preferable, the pipeline can read existing chunks from Wasabi, generate a new embedding set, persist that version, and build a new index for evaluation.
Each chunk and embedding should also retain a stable identifier that maps back to its source object and processing version in Wasabi. This allows retrieved results to be traced to the authoritative content from which they were created and helps preserve reproducibility as processing and embedding strategies change.
Retrieval doesn’t have to rely exclusively on vector similarity, either. Modern RAG implementations increasingly combine semantic retrieval with keyword or other search techniques according to the characteristics of the corpus and the questions being asked. Retrieval strategies can change while Wasabi continues to preserve the data and artifacts from which those indexes are built.
Where historical comparison or rollback matters, versioning can also help preserve changes to selected objects as the pipeline evolves.
Design the retrieval path for latency and throughput
Persistent object storage and online retrieval solve different parts of the architecture. The vector database should serve the latency-sensitive similarity-search path, and frequently accessed chunk text may be stored with the vector payload or maintained in an application cache when the serving design calls for it. Wasabi remains the durable source of record behind those serving layers, preserving the source corpus, processed text, embeddings, and other artifacts needed to rebuild or refresh them.
This distinction also separates end-user latency from throughput-intensive work. Ingestion, document processing, re-embedding, index rebuilding, and evaluation can move larger volumes of data independently of the application's request path. Multipart uploads can improve the reliability of large transfers, while processing services can parallelize object-level work according to the requirements of the application and compute environment.
The goal is not to force object storage into every synchronous request. It is to ensure the low-latency serving layers can always be reconstructed from an authoritative data foundation.
Stage 4: Retrieve context and perform grounded inference
When a user submits a prompt, the retrieval layer searches its index for the information most relevant to the request.
The retrieval index returns the chunks or references most relevant to the request. Depending on the serving architecture, chunk text may be stored directly with the vector payload, maintained in a low-latency cache, or resolved from Wasabi when direct access to the durable object is appropriate. The retrieved content is then assembled into the context supplied to the LLM.
In a straightforward RAG workflow, this may involve a single retrieval operation. More advanced implementations can evaluate what was retrieved, refine the query, search another source, or perform additional retrieval before inference begins.
This becomes especially important as context windows grow. The objective isn’t simply to fill the available context with as much enterprise data as possible. It’s to provide the model with the information most relevant to the task.
Once that context has been assembled, the LLM performs grounded inference and returns the result through a RAG application, whether that is a chatbot, search interface, API, or embedded business workflow.
This serving architecture keeps the latency-sensitive retrieval path optimized for application performance while Wasabi preserves the data required to reproduce, rebuild, or change that path later.
Stage 5: Persist the feedback loop
A production RAG pipeline doesn’t end when the model returns an answer.
Retrieved context, cached responses, evaluation datasets, prompt versions, retrieval results, user feedback, custom AI skills, and operational telemetry can all become useful artifacts for measuring quality and improving subsequent versions of the application.
Those artifacts can be written back to Wasabi alongside the enterprise corpus.
Over time, this creates a reusable operational history of the RAG application. Engineering teams can evaluate a new retrieval strategy against existing test data, compare prompt changes, regenerate embeddings, rebuild indexes, or introduce new AI services without having to reconstruct previous work.
Consider a change in embedding models. The enterprise corpus doesn’t need to move. The pipeline reads existing processed data from Wasabi, generates a new set of embeddings, writes them back as a separate representation, builds a new retrieval index, and evaluates that index before replacing the previous version.
The same pattern applies to a new chunking strategy. Source objects remain in Wasabi while new chunks, embeddings, indexes, and evaluation results are created around them.
This is also where storage economics become part of the technical architecture. RAG workloads repeatedly move data between processing services, embedding models, retrieval infrastructure, and inference applications. Wasabi's no-egress-fee model lets those services repeatedly access data without adding storage egress charges every time the pipeline is reprocessed or evaluated.
Protecting the RAG system of record
As RAG moves into production, the persistent data layer also becomes part of the application's security and governance model.
Wasabi's S3-compatible IAM capabilities allow teams to restrict access based on each service's responsibilities. A processing service, for example, may require access to source and processed objects, while an application-facing service may only require access to specific derived content.
A simplified policy for the processing tier could follow this pattern:
This illustrative policy gives the processing service read access to the authoritative corpus and read/write access to its derived outputs. Production policies should be scoped to the application's actual buckets, prefixes, services, and administrative requirements rather than copied directly.
Versioning can preserve previous versions of selected objects where rollback or historical comparison is important. For data subject to stronger retention requirements, Object Lock can protect objects from modification or deletion for a defined retention period.
Presigned URLs provide another mechanism for limiting access by granting temporary object-level permissions without exposing permanent storage credentials.
These controls become increasingly important as the RAG data layer expands beyond source documents to include embeddings, evaluation results, prompts, responses, and other records of how the application operates.
Best practices for RAG on Wasabi
Several principles keep the architecture maintainable as the surrounding AI stack changes:
Keep the authoritative corpus persistent. Store original enterprise content separately from the representations derived from it.
Organize around operational boundaries. Structure prefixes around ownership, access, lifecycle, and reprocessing requirements, adding chronological partitions where they provide practical value.
Organize for reprocessing. Separate source data, processed content, chunks, embeddings, evaluation data, cache objects, and telemetry logically so individual stages can be regenerated without rebuilding everything.
Persist reusable AI artifacts. Embeddings, evaluation datasets, retrieval results, prompt versions, and operational data can become valuable inputs to future versions of the application.
Keep the serving path optimized for latency. Use the vector database or application cache for latency-sensitive retrieval while maintaining Wasabi as the durable source from which those layers can be rebuilt.
Maintain traceability. Use stable identifiers to connect chunks and embeddings back to their source objects and processing versions.
Decouple storage from retrieval. Treat the vector database as an index rather than the system of record, so the retrieval infrastructure can change independently.
Version selectively. Preserve the history of artifacts where rollback, comparison, or reproducibility provides operational value.
Apply least-privilege access. Use IAM policies and scoped access patterns appropriate to the services interacting with the RAG corpus.
Design for iteration. Assume embedding models, chunking strategies, retrieval methods, LLMs, and orchestration frameworks will change over the lifetime of the application.
Build the data layer to outlast the AI stack
Production RAG is becoming more dynamic, not less.
Longer context windows expand what models can consume. Hybrid and adaptive retrieval improve how applications identify relevant knowledge. Evaluation and feedback loops continuously reshape how that knowledge is processed and used.
A RAG architecture built on Wasabi gives those components room to evolve independently. Source content remains available through a standard S3-compatible interface. Derived datasets and embeddings can be regenerated. Retrieval indexes can be rebuilt. Models and compute environments can be replaced without relocating the enterprise corpus.
The architectural principle is straightforward: compute can be replaceable, retrieval methods evolve, but the data layer is persistent.
Wasabi provides that foundation, allowing production RAG applications to evolve without rebuilding the enterprise knowledge beneath them.
The storage layer for your AI workloads
Choose affordable, predictable, and secure hot cloud storage for the AI data lifecycle, from ingest and training data to checkpoints, model artifacts, and long-term retention.
It's the storage foundation for the enterprise corpus and everything derived from it: chunks, embeddings, and evaluation data. It stays independent of the compute, models, and retrieval infrastructure built on top, so those can change or get replaced without moving the underlying data.
Yes. A larger context window lets a model hold more text, but it doesn't identify which information is relevant or reduce the cost of reprocessing large volumes of data. Retrieval and long context work together in most enterprise RAG systems rather than replacing one another.
The vector database is a retrieval index, not the system of record. Keeping embeddings and source content in persistent storage like Wasabi means a team can rebuild an index, switch vector databases, or adopt a new embedding model without touching the original corpus.
RAG workloads move data repeatedly between processing, embedding, retrieval, and inference services. Wasabi's no-egress-fee model means that repeated access doesn't add storage egress charges each time the pipeline gets reprocessed or evaluated.
Related article
Most Recent
Learn how Wasabi Account Control Manager helps university IT centralize, isolate, govern, and optimize higher education cloud storage management across every department.
A fully sovereign stack is impossible today. Wasabi’s Kevin Dunn explains why operational control beats corporate domicile, and why independent storage is EMEA’s highest-leverage sovereignty decision.
Performance Hub, an AI-powered gym management platform, connected to Wasabi's MCP beta with zero integration work. Its agents now query storage in natural language.
SUBSCRIBE
Storage Insights from the Storage Experts
Storage insights sent direct to your inbox.
&w=1920&q=75)