Cloud 101
Dark Data 101: What it is and How to Unlock it
In the artificial intelligence (AI) era, being “kept in the dark” can be costly. Yet the 2026 Wasabi Cloud Storage Index found that 25% to 74% of corporate data is dark data—information organizations store but never use. Many organizations don’t even know the extent of their dark data. And while companies continue paying to store it, they are also missing valuable insights that could fuel AI initiatives and business growth.
Dark data can also be dangerous, with risks to security and compliance. Its opaqueness threatens to render AI adoption incomplete, incorrect, and insecure. This article explores these issues and discusses why it accumulates, and what you can do about it.
What is dark data?
Dark data is data that organizations collect or generate but do not analyze for decision-making, operational, or strategic purposes. Some of the most common types include device logs, recordings of calls, email archives, surveillance video files, document archives, and chat records. There’s a lot of it, too. For a sense of scale, consider that a firewall can generate up to 7 TB of log data in a year. If no one is parsing that log data for the purpose of security operations, that’s dark data, and the money spent storing it is wasted.
Types of dark data
Dark data falls into three categories:
Structured dark data — Data stored in databases and business applications, such as customer relationship management (CRM) records, enterprise resource planning (ERP) information, server logs, and Internet of Things (IoT) sensor data.
Unstructured dark data — Data without a predefined format, including emails, documents, call recordings, social media content, and video files.
Semi-structured dark data — Data organized by tags or metadata rather than a fixed schema, such as invoices, HTML files, XML documents, and JavaScript Object Notation (JSON) files.
Why data goes dark
Why does data go dark? The creation of dark data is usually a matter of not doing anything with the data. You set up a server, and it starts writing log data onto a storage solution. No one pays too much attention to it, and it lingers as dark data.
Some of the most common reasons why dark data exists include:
Data silos across departments — When departments such as sales and marketing operate separate CRM systems and software, they invariably create data silos. These silos contain dark data, to the extent that no one outside the department can see it. For example, the marketing department may have customer sentiment data that could benefit the sales team, but sales is locked out of the data source.
“Set-and-forget” systems — Many devices feature “set-and-forget” functionality, such as log generation or “heartbeat” signals sent to system management platforms. The problem here is the “forget” part, which can lead to the amassing of dark data.
Lack of administrative resources — Managing dark data needs to be someone’s job, and many IT departments don’t have the bandwidth for it. If you have 25 TB of firewall log data, it may be simpler to stick it on a storage array and tell yourself that you’ll slim it down or run analytics on it later rather than assign the tasks to an overworked admin.
Legacy systems that don't integrate with modern tools — Legacy systems tend to be silos in functional terms, but also as data repositories. For instance, if you’re running a zSeries mainframe, its data might be in the Extended Binary Coded Decimal Interchange Code (EBCDIC) format, with 8-bit character encoding that is incompatible with modern relational databases like MySQL. That data is going to be dark, and if you want to “shine a light on it” you’re in for a serious data integration project.
Poor or absent data governance — Data governance policies help your organization manage data for optimal business outcomes, as well as for security, compliance, and legal liability. Without clear, well-enforced policies, problematic dark data can quickly accumulate. For instance, if you don’t have strict data retention and deletion policies, you may be unaware of old records that could affect litigation or privacy law compliance.
Unresolved data quality issues — Data quality problems often translate into dark data. A lack of master data is one example. You might have duplicate data on customers that is dark because no one is aware of the copies. The fallout from this dark data set might include failed product deliveries, billing mistakes, and so forth.
Changing business priorities that leave datasets behind — Dark data often piles up when a business restructures or changes its priorities. For example, if your company launches a social media campaign and starts to log thousands of social media comments—but then abandons the effort—that social media dataset could keep growing well after anyone cares about it.
Mergers and acquisitions (M&A) — Organizational changes generate dark data. For example, when IT admins get laid off or moved to new business units, there might not be anyone left behind who is up to date on the data being created.
The costs of dark data
Dark data comes with its share of costs. Some are tangible, others are more indirect, but it can still negatively affect a company’s bottom line. One cost factor is inefficiency. Dark data causes employees to waste time searching for data that isn’t organized or accessible.
Dark data also means missed opportunities for data analysis that can drive strategic thinking and operational improvements. For example, if the data from a logistics solution is dark to users of the ERP system, the company would not see ways to optimize delivery routes or consolidate shipments. Important findings about consumer sentiment and brand reputation could also be hiding in dark social media data.
Security and compliance problems can be expensive consequences of dark data. Security might be an issue, for example, if a dark database contains employees’ personal information, which enables attackers to impersonate privileged users. Or hackers might breach and exfiltrate dark data that contains sensitive information, and it may not even be evident to IT admins that anything has happened until it’s too late.
Dark data stores can cause compliance violations, particularly those related to privacy laws. For example, your company might comply with a data deletion request filed in accordance with the California Consumer Protection Act (CCPA). However, if you have a dark dataset that retains that consumer’s data without you knowing it, you could be hit with a penalty.
Storage costs are another burden of dark data. A Petabyte (PB) of unused data is costly enough, but paying to store it adds to the waste. On-prem storage ties up valuable infrastructure and increases administrative overhead. Cloud storage can be even more expensive, especially in high-performance tiers. Hidden charges, such as application programming interface (API) fees for backups, can drive costs even higher. The result can be unexpectedly large storage bills for data your organization never uses.
Dark data and AI
Dark data undermines AI adoption because AI models can miss it or use it incorrectly. Because AI is only as effective as the data it relies on, dark data limits the quality, accuracy, and depth of model outputs. As a result, AI can produce incomplete results or draw incorrect conclusions. For example, if a model is trained on only part of a customer transaction history, its understanding of that customer will be incomplete. Dark data also raises the risk that proprietary or sensitive information will enter the model and appear in responses where it should not.
AI dark data problems stem in part from storage issues. The 2026 Wasabi Cloud Storage Index found that 47% of organizations cite data storage challenges as their top obstacle to AI implementation. As a result, 91% say activating dark data is a priority. Hybrid storage can help by connecting isolated dark data with the data used to train AI models. It is now standard in AI workflows, but it only delivers value when the data is accessible. Otherwise, data can sit in hybrid storage and remain dark.
Shine a light on your dark data
Bringing dark data into the light is possible, though it requires time, effort, and investment. The good news is that organizations can significantly reduce hidden data and uncover valuable insights by using existing tools, improving data management practices, and taking a more deliberate approach to data discovery.
Many of the ways to uncover dark data are not technical at all; rather they depend on policy and organizational decisions. You can create or strengthen data governance policies that define what data to keep, archive, or delete. You can also identify data silos and decide whether it is worth breaking them down. Doing so can make data accessible across teams and increase its value to the business.
Auditing and classifying data across systems and departments is another smart step. This work typically requires specialized data discovery tools. When done well, a data audit often reveals forgotten or overlooked data stores. AI and machine learning (ML) tools can support this effort by helping classify unstructured data and extract useful value from it.
Storage must be on the agenda if you’re intent on reducing dark data. The best practice is to align storage tiering with data value. Reviewing where you store data is a crucial step. Make sure you’re storing low-priority data in the lowest cost tier. However, be careful about hidden costs in low-cost tiers of cloud storage. These services may come with high fees for data egress.
Conclusion
Dark data is a solvable problem. Getting rid of it is not a push-button process, however. It requires intent, the right infrastructure, and careful review and improvement of policies, practices, and architecture.
While storage often emerges as a barrier to uncovering dark data, affordable, predictable storage can be a big part of the solution. By putting data into a predictably priced solution like Wasabi, you gain visibility into previously opaque datasets. And Wasabi offers further advantages, such as access to your data when you need it, without egress or API fees.
Frequently Asked Questions
Dark data is data that organizations collect or generate but do not analyze for decision-making, operational, or strategic purposes.
AI models are only as accurate as the data they can access. Dark data, information an organization stores but never actively manages or analyzes, sits outside the pipelines that feed AI training and inference.
According to the Wasabi 2026 Cloud Storage Index, 47% of organizations cite data storage challenges as their primary obstacle to AI implementation, and 91% say that activating dark data is a priority. Getting that data organized, accessible, and connected to hybrid storage environments is a prerequisite for effective AI adoption.
Unstructured data, such as call recordings, email archives, video files, and document repositories, makes up a significant share of what organizations classify as dark data. Wasabi Hot Cloud Storage is built for large-scale object storage workloads, including unstructured data of any format, without performance tiers or minimum storage requirements.
Storage costs are a direct driver of dark data accumulation. When retrieving data is expensive, organizations avoid accessing it, which deepens the cycle of keeping data without using it. Many cloud storage providers charge egress fees every time data leaves their platform, along with API fees for accessing it. Those costs create a disincentive to audit, analyze, or act on stored data. Wasabi Hot Cloud Storage eliminates those fees entirely, making it practical to retrieve and analyze data you've stored, rather than letting it continue to accumulate in the dark.