Data Lake Market Size and Share

Data Lake Market Analysis by Mordor Intelligence
The data lakes market size is expected to grow from USD 18.68 billion in 2025 to USD 22.8 billion in 2026 and is forecast to reach USD 61.84 billion by 2031 at 22.08% CAGR over 2026-2031. Growth stems from surging unstructured data volumes generated by generative-AI pipelines, expanding regulatory record-keeping mandates, and the shift toward lakehouse architectures that collapse lake and warehouse footprints into a single tier. Fortune 500 firms report 35-40% total-cost savings after embracing lakehouses, while real-time ESG and risk-stress workloads are extending use cases into industrial and financial domains. Serverless open-table formats now anchor multi-cloud portability strategies, and automated governance layers are emerging to prevent “swamp” pitfalls without throttling innovation.
Key Report Takeaways
- By offering, solutions led with 69.35% revenue share in 2025; services are projected to expand at a 24.77% CAGR through 2031.
- By deployment, cloud captured 64.20% of the data lakes market share in 2025, while hybrid/multi-cloud is forecast to grow at a 23.1% CAGR between 2026-2031.
- By organization size, large enterprises commanded 71.10% of the data lakes market size in 2025; SMEs are the fastest risers at a 26.1% CAGR through 2031.
- By business function, operations & supply chain held 29.40% share of the data lakes market in 2025, whereas finance & risk is advancing at a 25.2% CAGR to 2031.
- By end-user vertical, IT & telecom led with 21.60% revenue share in 2025; healthcare & life sciences is poised to expand at a 25.6% CAGR to 2031.
- By geography, North America dominated with 37.40% share in 2025, while Asia is set to accelerate at a 23.5% CAGR through 2031.
Note: Market size and forecast figures in this report are generated using Mordor Intelligence’s proprietary estimation framework, updated with the latest available data and insights as of 2026.
Global Data Lake Market Trends and Insights
Drivers Impact Analysis*
| Driver | (~) % Impact on CAGR Forecast | Geographic Relevance | Impact Timeline |
|---|---|---|---|
| Explosion of unstructured & multimodal data from GenAI workloads | +7.5% | Global with concentration in North America & Western Europe | Medium term (2-4 years) |
| Data-residency mandates in Europe accelerating cloud-based lake adoption | +5.2% | European Union, UK, Switzerland & APAC | Short term (≤ 2 years) |
| Lakehouse convergence driving 35–40% TCO savings for Fortune 500 firms | +6.3% | Global with early adoption in North America | Medium term (2-4 years) |
| Serverless table formats (Iceberg/Delta) unlocking multi-cloud portability | +4.8% | Global, strongest where multi-cloud strategies are active | Medium term (2-4 years) |
| Real-time ESG Scope-3 data-capture requirements in industrial sector | +3.2% | Europe, North America, advanced APAC economies | Long term (≥ 4 years) |
| Regulatory stress-testing in financial services demanding decade-scale tick-data retention | +2.9% | Global financial centers (New York, London, Singapore, Hong Kong) | Medium term (2-4 years) |
| Source: Mordor Intelligence | |||
Explosion of unstructured and multimodal data from GenAI workloads
Generative-AI applications create vast image, audio, and text payloads that demand schema-on-read storage. Enterprises expect 30% of the global 175 zettabyte data sphere to require real-time processing by 2025, a profile unsuited to rigid warehouses. Data lakes therefore become the default landing zone for multi-modal corpora used in prompt-engineering loops.[1]Acceldata, “Enterprise Data Lakes: Revolutionizing Business Data,” acceldata.ioGoogle Cloud’s lakehouse blueprint shows how native-format storage paired with vector indexing accelerates foundation-model fine-tuning while lowering storage bills. Firms delaying adoption risk slower innovation cycles and higher unit-costs on AI workloads.
Data-residency mandates in Europe accelerating cloud-based lake adoption
The EU Data Governance Act and Data Act compel organizations to localize sensitive workloads. Hyperscalers are responding: AWS is investing EUR 7.8 billion in a sovereign-cloud region that ships with embedded data-location controls.[2]Databricks, “Databricks Agrees to Acquire Tabular,” databricks.com Enterprises now deploy region-segmented data lakes that meet residency rules yet remain queryable through federated engines, sparking demand for lineage-rich metadata catalogs capable of surfacing cross-border data usage in audit reports.
Lakehouse convergence delivering 35-40% TCO savings
A single-tier lakehouse erases the duplication that once plagued separate lakes and warehouses. Surveyed enterprises moving analytical jobs onto lakehouse engines cite halved data-movement costs and compression-driven storage savings. Performance gains from vector-aware query planners further collapse compute runtimes, freeing budget for AI experimentation. Eighty-one percent of firms now train ML models directly on lakehouse tables, indicating convergence is no longer an edge practice but a mainstream pattern.
Serverless table formats unlocking multi-cloud portability
Apache Iceberg, Delta Lake, and Hudi introduce ACID transactions, schema evolution, and time-travel to object stores. The formats decouple compute from storage, letting analytics engines in rival clouds query the same datasets without replication. Databricks’ 2024 acquisition of Tabular underscores the strategic value of open table metadata, while Google BigLake’s Omni feature queries Iceberg partitions in rival clouds, validating the neutral-format thesis.[3]European Commission, “A European Strategy for Data,” digital-strategy.ec.europa.eu
Restraints Impact Analysis*
| Restraint | (~) % Impact on CAGR Forecast | Geographic Relevance | Impact Timeline |
|---|---|---|---|
| Metadata drift creating “data swamps” | -3.8% | Global, more acute in legacy deployments | Short term (≤ 2 years) |
| Skilled data-lake engineering talent shortfall | -2.9% | APAC, Latin America, Middle East & Africa | Medium term (2-4 years) |
| Latency-sensitive use cases still prefer warehouses | -2.1% | Finance, telecom hubs worldwide | Short term (≤ 2 years) |
| Opaque consumption-based cloud pricing | -1.7% | Mid-market firms globally | Medium term (2-4 years) |
| Source: Mordor Intelligence | |||
Metadata drift creating “data swamps”
When ingestion outpaces catalog updates, data lakes devolve into unsearchable repositories. By 2025, global data volume will reach 163 zettabytes, heightening the risk of siloed files with missing context. Enterprises are responding by adopting automated lineage trackers such as Unity Catalog, which logs every read-write and flags orphaned assets. Without similar controls, governance overhead can erase savings projected from lakehouse consolidation.
Skilled lake-engineering talent shortfall in emerging regions
APAC and Latin-American firms cite a scarcity of engineers who understand distributed filesystems, open-table formats, and cloud cost tuning. POPsights data shows AI-driven role creation outpacing local training supply. OECD research highlights a widening urban-rural gap in access to advanced data skills.[4]OECD, “Job Creation and Local Economic Development 2024,” oecd.org Managed services and low-code pipelines are mitigating shortages, yet talent scarcity still lengthens deployment cycles, slowing data lakes market penetration.
*Our forecasts treat driver/restraint impacts as directional, not additive. The impact forecasts reflect baseline growth, mix effects, and variable interactions.
Segment Analysis
By Offering: Solutions lead, services surge
Solutions generated 69.35% of data lakes market revenue in 2025, equating to a data lakes market size of USD 12.95 billion. The dominance comes from enterprises standardizing on storage engines, query accelerators, and governance suites that form the backbone of AI-ready environments. Vendors bundle cost-optimizer dashboards, automated tiering, and native open-table support, maintaining relevance as workloads evolve.
The services sub-segment is racing ahead at a 24.77% CAGR to 2031, reflecting demand for migration blueprints, performance tuning, and 24×7 managed operations. Many firms lack staff who can re-platform legacy Hadoop estates, so they contract specialists that promise predictable SLA outcomes. The tight talent market ensures professional-services bookings will keep growing faster than the overall data lakes market

By Deployment: Cloud rules, hybrid accelerates
Cloud deployments captured 64.20% of the data lakes market share in 2025 as organizations sought instant scalability and integrated security. Elastic object stores like Amazon S3 eliminate CapEx while delivering lifecycle automation that auto-tiers cold data to low-cost classes. Analytics engines then spin up on demand, keeping compute spend aligned with project tempo.
Hybrid and multi-cloud configurations are expanding at 23.1% CAGR to 2031. Open-table formats let one metadata definition span on-prem and public-cloud buckets, slashing replication needs. Regional compliance rules further fuel hybrid strategies, as firms pin regulated workloads in sovereign regions yet still query them through cross-cloud fabrics. As a result, the data lakes market size for hybrid environments is rising in lockstep with sovereign-cloud launches.
By Organization Size: Large enterprises dominate, SMEs gain pace
Large enterprises accounted for 71.10% of the data lakes market size in 2025, or approximately USD 13.28 billion. Their complex, petabyte-scale estates require advanced RBAC, automated lineage, and FinOps governance. Banks, manufacturers, and telecoms rely on lakehouses to consolidate silos and support real-time AI applications.
Small and medium enterprises log the fastest 26.1% CAGR because vendor-managed plans now offer “pay-as-processed” billing. Low-code orchestration and template-driven schemas shorten deployment cycles. Community editions of Iceberg and Delta expose enterprise-grade capability without license fees, letting resource-constrained firms join the data lakes market mainstream.
By Business Function: Operations steady, finance & risk surging
Operations and supply-chain workloads generated 29.40% of 2025 spend, with manufacturers blending IoT telemetry, supplier EDI, and logistics feeds for predictive maintenance. Schema-on-read flexibility makes lakes ideal for fusing semi-structured sensor files with ERP tables, supporting control-tower dashboards that slice downtime risk.
Finance and risk applications are growing at 25.2% CAGR. Regulators now expect decade-deep tick histories, and lakehouses store these volumes efficiently. The Federal Reserve’s April 2025 buffer-rule proposal underscores the need to model capital impacts under stressed conditions. Banks that centralize risk, treasury, and ESG records inside a governed lake eliminate reconciliation delays, gaining reporting agility.

By End-User Vertical: IT and telecom lead, healthcare advances
IT and telecom operators held 21.60% of 2025 revenue. Carriers ingest call-detail records, network KPIs, and support transcripts in lakes, then run fraud detection and churn analytics that improve lifetime value. Softteco notes Vodafone and AT&T use AI-driven lake architectures to optimize towers and personalize offers.
Healthcare and life sciences are projected to climb at a 25.6% CAGR. Hospitals marry electronic health records, imaging, and genomics in unified repositories that power precision-medicine studies. Microsoft Fabric deployments illustrate how unified ingestion pipelines cut data prep times, enabling real-time clinical alerts. Pharma firms exploit repeatable lake workflows to trim discovery cycles, driving sustained investment in the data lakes market.
Geography Analysis
North America generated 37.40% of 2025 revenue and continues to set benchmarks in architecture maturity. Financial institutions lengthen time-series retention to meet evolving stress-test templates, while hospital networks build multimodal patient graphs that underpin AI-driven diagnostics. Venture capital also fuels governance-start-up formation, ensuring a vibrant ecosystem.
Asia-Pacific is the fastest-expanding region, clocking a 23.5% CAGR through 2031. Governments in Japan, India, and Singapore sponsor sovereign-cloud projects, spurring demand for region-compliant lake zones. Telcos in China analyze massive 5G logs for capacity planning, whereas Indonesian fintechs share fraud-intelligence lakes to curb cybercrime. Vendors establishing APAC headquarters, such as Wasabi in Japan, aim to catch the projected 36% IaaS upturn.
Europe accelerates adoption under strict data-sovereignty mandates. The European Strategy for Data drives investment in local hosting, and AWS will open a Brandenburg region by late 2025 to satisfy residency rules. Manufacturers store real-time Scope-3 emissions for CSRD reporting, and banks refine Basel III calculations inside audit-ready lakes. The European Banking Authority’s 2025 stress-test templates reinforce technical requirements that lakehouses fulfill.

Regulatory Landscape
Regulatory requirements around data access, location, and governance are tightening in ways that shape how enterprises design and run data lakes. In the European Union, the Data Act (Regulation (EU) 2023/2854) enters into application in stages, with effectiveness from September 12, 2025 and key application dates starting September 12, 2026. This strengthens obligations around data access and portability, which increases demand for standardized metadata, auditable sharing controls, and multi-environment architectures. Separately, the EU Artificial Intelligence Act introduces explicit data governance expectations for high-risk AI. Article 10 sets requirements on training, validation, and testing data quality and governance, and full obligations for high-risk systems (including Articles 10 to 12) apply from August 2, 2026, raising the priority of lineage, provenance, and dataset documentation within lakehouse and lake governance layers.
Public-sector and cross-border compliance drivers also raise baseline expectations for security and information management across large-scale data platforms. In the United States, OMB memoranda M-25-04 (FY 2025 guidance on federal information security and privacy management requirements under FISMA) and M-25-05 (Phase 2 implementation guidance for the Foundations for Evidence-Based Policymaking Act of 2018 on open data access and management) reinforce expectations for disciplined data handling, access control, and governance artifacts. These map closely to cataloging, policy enforcement, and audit-ready storage patterns used in enterprise data lakes. In the United Kingdom, the government data asset management policy in government published in May 2026 requires departments to identify, record, and centrally report data assets to the Government Digital Service, increasing the emphasis on discoverability and standardized data inventories that favor enterprise catalog-and-lake operating models.
Competitive Landscape
The data lakes market is moderately fragmented. Hyperscalers-AWS, Microsoft Azure, Google Cloud-dominate infrastructure, leveraging global regions and integrated governance. Specialized platforms such as Databricks and Snowflake distinguish themselves on performance, notebook integration, and lakehouse completeness. Open-source communities steer Iceberg, Delta, and Hudi, giving buyers format options that loosen vendor grip.
Strategic acquisitions are reshaping value-chains. Databricks purchased Tabular in 2024 to tie Iceberg lineage into Delta workflows, signaling a bet on universal metadata. Fivetran bought Census in 2025, unifying ingestion and reverse ETL to close the activation loop-. Commvault’s 2024 Clumio deal adds ransomware-recovery snapshots for S3 lakes. These moves point to a future where integrated suites span ingestion, governance, protection, and activation.
Despite hyperscaler heft, the top five suppliers capture roughly 55% of total spend, leaving headroom for innovators that specialize in cost-optimization, cross-cloud query acceleration, and vertical-specific governance blueprints. AI-augmented data-quality observability and sovereign-cloud governance are two emerging white spaces likely to attract new entrants.
Data Lake Industry Leaders
Microsoft Corporation
Amazon.com Inc.
Capgemini SE
Oracle Corporation
Teradata Corporation
- *Disclaimer: Major Players sorted in no particular order

Market Opportunities and Future Outlook
Compliance-driven governance and cross-platform interoperability are expanding the commercial whitespace around data lakes beyond storage and query, especially in regulated AI and data-sovereignty settings. The EU AI Act introduces enforceable data governance obligations for high-risk AI from August 2, 2026 (including Article 10). That creates demand for solutions that operationalize dataset provenance, bias assessment workflows, lineage, and controls over data preparation and documentation. As a result, buyers are leaning toward lakehouse and open-table implementations where governance policies, catalogs, and fine-grained permissions can be audited, and vendors and service providers can package AI-governance-ready reference architectures, controlled data product publishing, and managed compliance operations.
Execution in telecom also points to active investment pathways tied to unified data foundations for AI-driven operations. In April 2026, IBM deployed an AI-ready data lakehouse using IBM watsonx for Tata Play Fiber to consolidate fragmented data spanning customer, marketing, finance, and network operations. In June 2026, Nokia and Databricks completed a joint proof of concept for a unified, substrate-agnostic data platform to support AI-driven autonomous networks. Together, these programs underline opportunities around automated data product generation, data fabric and mesh overlays, and compute-to-data patterns that reduce movement of sensitive data while improving reuse. As open table formats such as Apache Iceberg and Delta Lake become procurement defaults for portability, suppliers can differentiate through interoperability features, governance automation that prevents data-swamp outcomes, and services-led modernization for hybrid and multi-cloud estates where data residency and cost controls shape architecture choices.
Recent Industry Developments
- June 2026: AWS updated Lake Formation to enable reading and writing of underlying Amazon S3 data files for tables registered in the AWS Glue Data Catalog, aligning permissions across SQL and file-based access patterns. The change supports mixed-engine operating models where different tools interact with the same lake tables, raising the bar for unified governance and lowering friction in multi-tool data lake deployments.
- February 2026: Microsoft announced general availability of OneLake and Snowflake interoperability, enabling native storage of Snowflake-managed Apache Iceberg tables in OneLake. This strengthens Iceberg as a neutral table layer and helps enterprises reduce duplication by keeping data in a shared lake storage foundation while enabling multiple analytics platforms to operate on it.
- May 2025: Fivetran acquired Census, adding reverse-ETL capabilities that activate data in operational systems. The deal connects ingestion and activation into a more complete data lifecycle, increasing the strategic role of data lake and lakehouse environments as the governed source feeding downstream business applications.
Research Methodology Framework and Report Scope
Market Definition and Coverage
This market covers revenue earned from data lake software and related services that help store, organize, and analyze large volumes of structured and unstructured data, across cloud, on-premise, and hybrid environments.
Scope exclusions: We exclude general-purpose hardware-only spending and broad IT outsourcing that is not directly tied to implementing or running a data lake environment.
Segmentation Overview
- By Offering
- Solutions
- Data Discovery and Cataloging
- Data Integration and ETL/ELT
- Analytics and Visualization Tools
- Governance and Security Platforms
- Services
- Professional Services (Consulting, Integration)
- Managed Services
- Solutions
- By Deployment
- Cloud
- Public Cloud
- Private Cloud
- Hybrid/Multi-Cloud
- On-Premise
- Cloud
- By Organization Size
- Large Enterprises
- Small and Mid-Size Enterprises (SMEs)
- By Business Function
- Operations and Supply-Chain
- Finance and Risk
- Sales and Marketing
- Human Resources
- By End-User Vertical
- IT and Telecom
- BFSI
- Healthcare and Life Sciences
- Retail and E-commerce
- Manufacturing and Industrial
- Media and Entertainment
- Government and Public Sector
- Energy and Utilities
- Others (Education, Hospitality)
- By Geography
- North America
- United States
- Canada
- Mexico
- South America
- Brazil
- Argentina
- Chile
- Peru
- Rest of South America
- Europe
- Germany
- United Kingdom
- France
- Italy
- Spain
- Rest of Europe
- Asia-Pacific
- China
- Japan
- India
- Australia
- New Zealand
- Rest of Asia-Pacific
- Middle East
- United Arab Emirates
- Saudi Arabia
- Turkey
- Rest of Middle East
- Africa
- South Africa
- Rest of Africa
- North America
Data Sources, Market Sizing, and Validation
Desk Research
Desk work starts by building a plain fact base around enterprise data growth, cloud adoption, and data governance needs, since these factors track closely with data lake rollouts across IT departments. We refer to public sources such as the US Bureau of Labor Statistics for IT spending signals, the US Census Bureau for business counts by industry, and OECD and World Bank indicators for digital economy comparisons across regions.
We also scan technical and adoption markers from sources such as NIST publications, peer-reviewed journals on data management architectures, and public procurement portals that show tender language around analytics platforms and data modernization. Company annual reports, earnings call notes, and investor presentations help us understand product packaging (software versus services) and how cloud and subscription revenue is recognized. Where required, we use paid subscriptions for company financials and intelligence, news and financials, and patent databases to cross-check timelines and ownership of core platform capabilities. The sources listed here are illustrative, and many other public references were used to collect, validate, and clarify inputs.
Primary Interviews and Surveys
Primary inputs are used to test adoption assumptions and pricing ranges with people who buy, deploy, and operate data lake environments, including IT leaders, data engineering managers, cloud architects, and service partners. For a global view, we spread conversations across major demand regions and then re-check any large differences through follow-up questions, so the final model reflects typical enterprise rollouts.
Distribution of primary research fieldwork respondents
| Company type | Respondent position | Region |
|---|---|---|
| Top tier: 35% | CXOs: 16% | APAC: 43% |
| Mid tier: 44% | Functional/Unit leaders: 25% | EMEA: 37% |
| Smaller Players: 21% | Managers: 59% | Americas: 20% |
Market-Sizing & Forecasting
Sizing is built using a top-down approach where enterprise data-platform spend pools are reconstructed by region, then filtered using adoption and migration rates specific to data lake use cases. We corroborate the totals with selective bottom-up checks, such as sampling typical annual contract values, validating the solution-to-services mix, and pressure-testing volumes through channel conversations and implementation timelines.
Key model inputs include cloud versus on-premise deployment mix, average subscription and support price progression, services attach rates for implementation and managed operations, workload growth tied to analytics projects, and industry-level adoption differences across regulated sectors. When a data point is missing for a country or vertical, we fill gaps using proxy indicators like enterprise counts and cloud maturity, then adjust the output after it is reviewed with primary respondents. Forecasting uses scenario analysis supported by regression-style sensitivity checks on the strongest drivers (cloud migration pace, data growth, and compliance-driven modernization), and assumptions are updated when experts indicate a clear shift in buying behavior.
Data Validation & Update Cycle
Validation is done through repeated cross-checks between the model and independent signals, such as regional IT spend direction, cloud services growth cues, and changes in enterprise data management priorities. Outliers are flagged early, and the assumptions behind them are reviewed by another analyst before the numbers are finalized, which helps reduce avoidable variance.
The work is refreshed annually, and interim updates are triggered when there are material events such as large pricing changes, major regulatory shifts, or a clear step change in cloud adoption patterns. Before delivery, we run a final pass to ensure recent developments are reflected, and any large deltas versus prior editions are explained and re-tested through quick expert re-contacts.
Mordor Intelligence's Data Lakes Market Size Compared With Other Published Estimates
Published market sizes for data lakes can look far apart even when they describe similar buyer needs, because the scope choices are not always consistent and the pricing logic is not always explained. Differences also show up when one estimate anchors on a different base year or uses a longer forecast window, which can change what is counted as current revenue.
The main gap comes from whether adjacent data management categories are bundled into the total. In this work, Mordor Intelligence counts data lake solutions and related services, but avoids folding in broader data warehousing and general analytics platforms unless they are sold and deployed as part of a data lake environment.
Benchmark comparison
| Source | Market Size | Gaps in Research Methodology |
|---|---|---|
| Mordor Intelligence | USD 18.68 B (2025) | |
| Global Consultancy A | USD 11.07 B (2025) | Uses a narrower revenue capture in the base year, which can undercount services attach and hybrid deployments that still drive paid platform and integration work. |
| Industry Publisher B | USD 10.70 B (2025) | Applies a different component split and often leans on a more conservative base-year conversion to US dollars, which can reduce the reported 2025 total versus models that normalize pricing and deployment mix. |
Looking at the table, the spread is mainly explained by what gets counted inside the data lake spend pool and how services and hybrid use cases are treated in the base year. By keeping the inputs tied to clear adoption, deployment mix, and pricing assumptions, we can explain each step and re-check the totals when new market signals appear without rebuilding the model from scratch.
Key Questions Answered in the Report
Why are enterprises moving from warehouses to lakehouses?
Lakehouses lower analytics TCO by 35–40% and support AI model training on raw data while preserving ACID performance guarantees.
How big is the data lakes market in 2026?
The data lakes market is valued at USD 22.8 billion in 2026 and is forecast to reach USD 61.84 billion by 2031.
Which region is growing fastest for data lake adoption?
Asia-Pacific leads with a projected 23.5% CAGR between 2026 and 2031, driven by rapid digital transformation and sovereign-cloud investments.
What is the main challenge preventing data lakes from delivering value?
Metadata drift can turn lakes into “data swamps,” prompting investment in automated catalogs and lineage tracking to maintain trust.
How do open-table formats affect vendor lock-in?
Formats like Apache Iceberg and Delta Lake enable multi-cloud portability by decoupling storage from compute engines, letting teams query the same data across different clouds.
Which industry vertical is forecast to grow fastest?
Healthcare & life sciences is set to expand at a 25.6% CAGR through 2031, leveraging data lakes for precision medicine and real-time patient analytics.
Page last updated on:




