Choosing a Lakehouse Architecture That Won't Box You In

Big data concept.
Written By
Kezia Jungco
Kezia Jungco
Aug 27, 2026
5 minute read
eWeek content and product recommendations are editorially independent. We may make money when you click on links to our partners. Learn More

Enterprise AI tools are changing much faster than most data architectures. A lakehouse that works well for today's analytics and machine learning projects can become harder to adapt to if its tables, catalogs, governance, or compute choices are too closely tied to a single platform.

No one can know which models, agents, analytics engines, or cloud services will matter several years from now. Organizations can still control the data foundation those tools depend on. Open table formats, interoperable metadata, consistent governance, and flexible deployment can make it easier to adopt new analytics and AI tools without repeatedly moving or rebuilding the underlying data.

Open formats keep the data layer portable

A lakehouse's table format affects more than how files are organized. It also shapes how different engines interpret data changes, manage transactions, and locate the information needed to run a query.

Apache Iceberg provides an open way to manage large analytic tables, with support for ACID transactions, schema and partition evolution, snapshots, and time travel. Cloudera's Iceberg migration guide also described support across Spark, Hive, Flink, Impala, Trino, and Presto. Different engines can therefore work with the same underlying tables, rather than requiring a dedicated copy for each workload.

Gartner's Market Guide for Data Lakehouse Platforms identified open table formats, object storage, separated compute and storage, and unified metadata and governance as core lakehouse characteristics. The firm recommended verifying that most lakehouse data is stored in an open, shareable format rather than primarily in proprietary storage.

Data lakehouses combine the flexibility and scale of a data lake with capabilities traditionally associated with data warehouses. An open table format can make that shared data less dependent on whichever processing engine is using it today.

Advertisement

Interoperability depends on more than Iceberg

Supporting Iceberg does not automatically make every part of a lakehouse portable. Catalogs, metadata services, security controls, APIs, and operational tools can create dependencies of their own.

The Iceberg REST Catalog can reduce another source of friction by standardizing how engines connect to a catalog. With a shared interface, different tools can discover and work with the same tables without requiring a custom catalog connection for every platform.

Shared files do not guarantee that different engines will consistently work with the data. Each engine also needs a consistent view of table snapshots and metadata, appropriate security controls, and a reliable way to coordinate changes when several workloads use the same tables.

Gartner similarly advised buyers to check whether lakehouse interfaces follow open standards so that data remains shareable across tools.

A platform can support Iceberg and still introduce dependencies elsewhere in the stack. If adding another engine requires a new copy of the data, a custom pipeline, or a separate governance model, changing tools later can still become difficult.

Interoperability also depends on governance that remains consistent across tools and environments. Metadata, lineage, policies, and access controls need to remain useful as new engines and AI workloads are added, rather than creating another set of disconnected rules for each platform.

AI flexibility starts with the data

As organizations add new AI workloads, different use cases may call for different models, retrieval tools, vector systems, agent frameworks, and processing engines. A lakehouse that can accommodate those choices without rebuilding the data layer gives teams more room to change tools as their requirements evolve.

Forrester's Q3 2026 Data Lakehouses Wave identified open table formats, interoperable metadata standards, zero-copy data sharing, and extensible APIs as key capabilities for lakehouses that support agentic AI. The report noted that models, tools, and workflows continue to evolve, increasing the value of an architecture that can connect with different technologies.

AI readiness also depends on the quality, context, and governance of the underlying data. In IDC research sponsored by Cloudera, 51.6% of 353 organizations prioritizing AI data readiness selected data intelligence, including data quality, cataloging, lineage, metadata, and master data, as one of their two top areas of focus. Another 37.9% selected data modernization across hybrid or cloud data lakes, lakehouses, warehouses, and databases.

Advertisement

Cloudera's agent-ready data series emphasized that AI applications need more than access to data. Effective use also depends on timely data, visibility into where it comes from and how it is used, and the right context and safeguards.

When governed data is already available through interoperable systems, adding an AI workload need not start with another isolated copy of the dataset. Teams have more room to bring different tools to the data as requirements change.

Hybrid environments make workload placement a design choice

Lakehouse flexibility also depends on where data and workloads need to run. Regulatory requirements, data sovereignty, existing infrastructure, performance demands, and cloud costs can all keep parts of the data estate in different environments.

IDC reported that 53% of analytic data repositories were stored in public cloud, hybrid cloud, or multicloud environments, while 47% remained exclusively on-premises or in private clouds. The report cited factors including regulation, sovereignty, cost optimization, and compute needs as reasons for that distribution.

A design that assumes every dataset will eventually move to a single environment can be difficult to apply across a large organization. Moving data simply to accommodate another processing tool can also introduce another pipeline to manage and another copy to govern.

Separating compute from storage and using open formats gives teams more freedom to choose where a workload runs while continuing to use governed data. Sensitive data might remain on-premises or in a private environment, for example, while other workloads use public cloud resources where that arrangement makes more sense.

Cloudera applies that approach across hybrid environments. Forrester’s 2026 lakehouse evaluation described the platform as supporting on-premises, private cloud, and public cloud deployments, with multiengine access across Spark, Hive, Impala, and Trino alongside governance, lineage, security, and metadata management.

The report also highlighted Cloudera’s open architecture and interoperability with non-Cloudera engines, allowing organizations to keep existing platforms where they fit rather than moving every workload onto a single stack.

Operational choices can still limit flexibility

Open standards do not remove the work involved in running a lakehouse at scale. Teams still have to account for concurrency, table maintenance, metadata growth, workload performance, security, and reliability.

Cloudera's production guidance for Iceberg noted that partitioning strategies, job scheduling, retention policies, and commit frequency affect how tables behave under concurrency and continuous change. Running ingestion, maintenance, and queries against the same tables makes that coordination especially important.

Advertisement

Gartner likewise recommended evaluating performance, reliability, capacity, maintainability, workload management, and the wider data lifecycle rather than selecting a lakehouse on features alone. It also noted that dedicated data warehouses may still perform better for some optimized warehouse workloads.

Avoiding lock-in therefore takes more than choosing an open-source table format. Architects need to consider how the table format, catalog, governance layer, compute engines, interfaces, and deployment model work together, as well as what would need to change if any one of those components were replaced.

Future AI projects will introduce new requirements, but the cost of adapting to them will depend heavily on the underlying data architecture. Open formats, interoperable interfaces, consistent governance, and flexible workload placement can make new tools easier to adopt without turning every technology change into another data migration.

Looking for the right AI governance platform? See eWeek’s guide to the best AI governance tools to compare leading options and their strengths.

Kezia Jungco

Kezia Jungco is a staff writer with five years of hands-on experience testing and analyzing generative AI platforms, chatbots, and NLP tools. She writes in-depth coverage for both enterprise and consumer audiences, focusing on artificial intelligence, data analytics, CRM solutions, cloud infrastructure, cybersecurity, and emerging tech trends. Her work appears in TechRepublic, eWEEK, Datamation, TechnologyAdvice, and Selling Signals.

eWeek Logo

eWeek has the latest technology news and analysis, buying guides, and product reviews for IT professionals and technology buyers. The site's focus is on innovative solutions and covering in-depth technical content. eWeek stays on the cutting edge of technology news and IT trends through interviews and expert analysis. Gain insight from top innovators and thought leaders in the fields of IT, business, enterprise software, startups, and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.