- 80% of enterprise data is unstructured, yet only <1% is actively used in AI/analytics.
- Komprise's Transparent File Tables enable querying petabytes of unstructured data without moving it.
- The solution integrates seamlessly with Snowflake and Databricks via Apache Iceberg tables.
Experts would likely conclude that Komprise’s approach addresses a critical gap in AI development by making previously inaccessible unstructured data usable without costly migration, potentially transforming enterprise analytics.
Komprise Aims to Fuel Enterprise AI by Illuminating 'Dark' Unstructured Data
CAMPBELL, CA – June 23, 2026 – Komprise, a leader in analytics-driven data management, today announced a technology that targets one of the biggest and most expensive roadblocks in enterprise AI: the inability to use the vast troves of unstructured data locked away in corporate storage. The new offering, called Transparent File Tables, creates a virtual, structured window into petabytes of files and objects, allowing them to be queried directly from major data lakehouse platforms like Snowflake and Databricks without moving the data first. The move could fundamentally alter the economics of AI development for businesses drowning in data they can't effectively use.
The 99% Problem: Confronting a Mountain of Untapped Data
For years, a stark paradox has defined the enterprise data landscape. Industry research consistently confirms that unstructured data—documents, images, research files, logs, and media archives—constitutes over 80% of an organization's total data footprint. Yet, its role in the AI revolution has been negligible. According to Komprise, less than 1% of this data is actively used in AI and analytics, a figure that, while startling, is directionally supported by numerous industry analyses highlighting a massive gap between data availability and practical application. This vast repository of potential insight remains "dark data."
The reasons for this are systemic. Unstructured data, by its nature, lacks a consistent schema. It's often of poor quality, siloed across multi-vendor NAS and cloud storage, and prohibitively expensive and complex to move. Traditional data ingestion pipelines, built for structured sources, buckle under the weight of petabyte-scale file transfers that can take weeks or months and incur massive cloud egress and infrastructure costs.
“The reason 99% of enterprise unstructured data has been dark to AI and analytics is because discovering and generating its schema and moving it is inherently complex and costly,” said Kumar K. Goswami, CEO and co-founder of Komprise. “Komprise brings to light the huge petabytes of enterprise unstructured data in a form that data teams can access easily and transparently for analytics. Komprise Transparent File Tables opens a whole new world to AI.”
Beyond Data Movement: A New Paradigm for AI Ingestion
The core of Komprise's strategy is to circumvent the data movement problem altogether. Instead of a costly "lift-and-shift" operation, Transparent File Tables (TFT) virtualizes access. The system operates on a principle of analyzing data in place. Komprise’s scale-out architecture indexes files across an enterprise's entire hybrid cloud environment, creating a single, searchable Global Metadatabase. This process doesn't move the files but instead builds a rich catalog of metadata, including file attributes, content-derived tags, and even the results of sensitive data scans.
This enriched metadata is then presented as an Apache Iceberg table to platforms like Snowflake or Databricks. A data scientist can run a query against this table to, for example, find all research files created by a specific instrument in the last six months that contain a certain protein name. The query runs against the metadata index, not the raw files, returning results in seconds.
The "no movement" claim is powered by the company's patented Transparent Move Technology (TMT). When a query identifies a specific subset of files needed for an AI model, TMT facilitates their retrieval on demand. The Iceberg table contains a pointer to each file's location, and Komprise's technology fetches only the required data, moving it at what the company claims is twice the speed of standard tools. This "just-in-time" approach avoids the colossal waste of ingesting entire datasets, most of which may be irrelevant, noisy, or low-quality.
Bridging Silos: Connecting Unstructured Data to the Business Analyst
Perhaps the most significant long-term impact of this technology is its potential to democratize access to unstructured data. By leveraging the open Apache Iceberg table format, Komprise ensures its solution plugs directly into the modern data stack. Both Snowflake and Databricks have invested heavily in first-class support for Iceberg, making the integration seamless for data teams.
This means a business analyst, not just a data engineer, can explore unstructured data within the same BI dashboard they use to analyze financial data. For instance, a pharmaceutical analyst could create a dashboard in Snowflake that joins a Komprise Transparent File Table (showing project files, instrument logs, and lab notes) with structured tables from an ERP system (tracking project costs) and a lab information system like Benchling. This unified view, combining disparate data types without complex ETL pipelines, has been a long-sought-after goal in data analytics.
Similarly, in media and entertainment, an AI agent tasked with checking narrative consistency could first query structured project data to identify relevant film archives, then join that with a Komprise table to pinpoint specific script versions for summarization and analysis. This targeted ingestion dramatically improves the efficiency and relevance of AI workflows. Crucially, the end-user requires no knowledge of Komprise; they simply query a table in their environment of choice.
A Strategic Play in a Crowded Field
Komprise is entering a competitive arena where data virtualization tools, lakehouse platforms, and other data management vendors are all vying to unify the enterprise data ecosystem. However, its approach is highly differentiated. While many tools focus on federating structured databases or providing a unified query layer, Komprise has built its architecture from the ground up to address the unique physics of managing petabyte-scale file and object data.
The company's advantage lies in the combination of its non-intrusive, scalable metadata indexing, its patented transparent access technology, and its workflow automation engine, KAPPA, which can tag and classify data to prepare it for AI. Furthermore, the platform maintains data governance by enforcing the original source access permissions, a critical capability as organizations become increasingly concerned with securing data exposed to generative AI models. By focusing on eliminating the initial data-gravity bottleneck, Komprise is making a strategic bet that the fastest way to unlock the value of unstructured data is to leave it where it is. For enterprises struggling under the weight of their own data, this approach could finally make the promise of holistic, data-driven intelligence an operational reality.
