Why Databricks Won the Data Lakehouse Market
August 5, 2026 · 16 min read
In 2013, the data world was cleanly split into two camps. On one side, data lakes stored everything cheaply in object storage (S3, HDFS) but couldn't handle SQL queries or ACID transactions. On the other side, data warehouses (Snowflake, Redshift) ran fast SQL queries but forced you to load data through rigid schemas and charged a fortune for storage. Every data team faced the same impossible choice: cheap storage that couldn't query, or fast queries that couldn't scale. Then Databricks proposed a heretical idea: what if you didn't have to choose?
Today Databricks is valued at $43 billion after its Series J round in late 2024, serves over 10,000 customers including 60% of the Fortune 500, and has crossed $2.4 billion in annual recurring revenue. It dominates the data lakehouse market so thoroughly that the term "lakehouse" itself, which Databricks coined in 2020, is now an industry standard. But its dominance wasn't inevitable. Snowflake had the better data warehouse. Google BigQuery had the most innovative serverless architecture. AWS Redshift had the distribution advantage of being embedded in the world's largest cloud. We analyzed Databricks against its three primary competitors using Spyglass's competitive intelligence framework. The results reveal how Databricks won by redefining the category rather than competing within it.
The Competitive Landscape
All four platforms help organizations store, process, and analyze large volumes of data. But each competitor approaches the problem from a fundamentally different angle:
| Dimension | Databricks | Snowflake | Google BigQuery | AWS Redshift |
|---|---|---|---|---|
| Founded | 2013 | 2012 | 2011 | 2012 |
| Target User | Data engineers, ML teams, lakehouse architects | Data analysts, BI teams, data warehousing | Google Cloud users, serverless-first teams | AWS-native teams, traditional warehousing |
| Architecture | Lakehouse (Delta Lake + Spark) | Cloud data warehouse | Serverless data warehouse | Cloud data warehouse (provisioned) |
| Pricing Model | DBU-based (compute) + storage | Credit-based (compute + storage) | Per-query (TB scanned) or compute | Node-based (provisioned) or serverless |
| Free Tier | Community Edition (limited) | $400 in free credits | 10 GB storage + 1 TB queries/mo | 2-month free trial (DC2 nodes) |
| Paid Plans | From ~$0.07/DBU (Jobs) to ~$0.55/DBU (Premium) | From ~$2/credit (Standard to Business Critical) | $6.25/TB (on-demand) or $0.04/Slot-hr | From $0.25/hr (dc2.large) to Serverless |
| Key Strength | Unified analytics (SQL + ML + streaming), open formats (Delta Lake), best Spark platform | Easiest SQL warehouse, best BI tool integration, separation of compute/storage | Serverless scale, Google ecosystem, BigQuery ML | Deep AWS integration, Redshift Spectrum for lake queries |
| Key Weakness | Complex setup, steeper learning curve, Spark expertise required | Expensive at scale, limited ML capabilities, vendor lock-in | GCP lock-in, less enterprise adoption, query costs unpredictable | Legacy architecture, less innovative, concurrency limitations |
| Revenue / Valuation | ~$2.4B ARR, $43B valuation (2024) | ~$3.4B ARR, ~$50B market cap | Part of Google Cloud ($44B ARR) | Part of AWS ($100B+ ARR) |
Databricks isn't the easiest to use (Snowflake is simpler for SQL analysts). It isn't the cheapest (BigQuery's serverless model can be more cost-effective for sporadic workloads). It doesn't have the distribution advantage of being embedded in the largest cloud (Redshift is native to AWS). Yet Databricks dominates mindshare among data engineers and ML teams who want a unified platform for analytics and AI. How?
Databricks' Five Strategic Moats
1. The Lakehouse Category Moat
Databricks' single most important strategic move was inventing and defining the "lakehouse" category. In June 2020, Databricks CEO Ali Ghodsi and UC Berkeley researchers published a seminal paper titled "Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics." This wasn't just marketing, it was a genuine architectural innovation: a system that combines the data management features of a data warehouse (ACID transactions, schema enforcement, governance) with the low-cost, flexible storage of a data lake.
By coining the term and publishing the research, Databricks ensured that every conversation about the future of data architecture would reference their technology. When Gartner, Forrester, and other analysts created their lakehouse category rankings, Databricks was the default leader because they defined the category. When enterprises issued RFPs for "lakehouse platforms," Databricks was the only vendor that had been building toward this vision for 7+ years. Snowflake responded by calling itself a "data cloud" (avoiding the lakehouse label), and BigQuery and Redshift tried to add lakehouse features retroactively, but Databricks had already set the terms of the debate.
Competitive Insight: Category creation is the most powerful competitive strategy in enterprise software. When you define the category, you set the evaluation criteria, and you can ensure those criteria play to your strengths. Databricks didn't compete in the data warehouse category (where Snowflake was winning), it created a new category where its lakehouse architecture was the only option. For indie founders, the lesson is: if you can't win the existing category, create a new one.
2. The Open Format Moat (Delta Lake)
Databricks' second strategic moat is Delta Lake, an open-source storage layer that brings ACID transactions, schema enforcement, and time travel to data lakes. Delta Lake sits on top of existing cloud storage (S3, ADLS, GCS) and stores data in Parquet format, meaning customers own their data in open formats that can be read by any tool. This is a radical departure from Snowflake's approach, where data is stored in proprietary micro-partitions that lock you into Snowflake's ecosystem.
The open-format strategy creates a powerful dynamic: customers adopt Databricks knowing they can leave at any time because their data is in Parquet/Delta format on their own cloud storage. This reduces the perceived risk of adoption (a critical factor in enterprise sales) while creating actual lock-in through Databricks' superior Delta Lake optimizations, Photon query engine, and Unity Catalog governance layer. The open-source community around Delta Lake (with 7,000+ GitHub stars and contributions from hundreds of companies) creates additional momentum. When Microsoft, Amazon, and Google all support Delta Lake in their analytics services, Databricks has won the format war even if some customers use alternative compute engines.
3. The Unified Analytics Moat (SQL + ML + Streaming)
Databricks' third moat is its unified platform that combines SQL analytics, data engineering, machine learning, and real-time streaming in a single environment. No other competitor offers this breadth. Snowflake is excellent at SQL analytics but has limited ML capabilities (Snowpark is playing catch-up). BigQuery has BigQuery ML but it's GCP-only and less flexible than Databricks' MLflow integration. Redshift has almost no native ML capabilities.
This unification matters because modern data teams don't just run SQL queries. They build ETL pipelines (Databricks Workflows), train ML models (MLflow, Feature Store), deploy real-time inference (Model Serving), and operationalize data quality (Delta Live Tables). When all of these workloads run on the same platform with the same data, the same governance, and the same billing, the operational efficiency gains are enormous. A team that uses Databricks for data engineering can hand off a Delta table to the ML team without any data movement. The ML team trains a model using MLflow, deploys it via Model Serving, and monitors it in the same platform. This workflow is impossible with Snowflake (which requires separate ML infrastructure) or BigQuery (which requires stitching together Vertex AI, Dataflow, and BigQuery).
4. The Spark Ecosystem Moat
Databricks was founded by the creators of Apache Spark, and the company remains the largest contributor to the Spark project. This creates a unique distribution moat: every organization using Spark (which is the dominant distributed compute engine for big data) is a potential Databricks customer. The Spark ecosystem includes hundreds of thousands of data engineers who learned Spark in school or at previous jobs and prefer to work with the best Spark platform available.
While Spark itself is open-source, Databricks' proprietary optimizations (Photon query engine, Delta Cache, optimized autoscaling) make Databricks-managed Spark 2-10x faster than vanilla open-source Spark. Organizations that start with open-source Spark on EMR or Dataproc often migrate to Databricks when they hit performance or operational limits. This migration path is frictionless because the code is identical, just faster. No other competitor has this natural upgrade path from open-source to commercial. Snowflake requires rewriting queries in its proprietary SQL dialect. BigQuery requires moving data into GCP. Redshift requires a completely different architecture. Databricks just makes your existing Spark code faster.
5. The Data Intelligence Moat (AI/ML Integration)
Databricks' newest and fastest-growing moat is its deep integration with AI and ML workloads. The company has invested aggressively in becoming the default platform for AI teams: MLflow for experiment tracking (15,000+ companies use it), Feature Store for feature engineering, Model Serving for production inference, and most recently, Mosaic AI for training and fine-tuning foundation models. Databricks' $1.3B acquisition of MosaicML in 2023 was a direct bet that the future of the data platform is inseparable from AI.
This AI integration creates a powerful flywheel. Teams that use Databricks for data engineering naturally adopt it for ML because the data is already there. Teams that use Databricks for ML generate more data (training logs, inference metrics, feature usage) that needs to be stored and analyzed. Each use case reinforces the others. Snowflake has responded with Cortex AI and Snowpark ML, but these feel bolted on rather than native. BigQuery ML is limited to specific model types. Redshift ML is barely mentioned in AWS's AI strategy. Databricks' AI-native approach positions it to capture the growing wave of AI infrastructure spending.
Where Competitors Went Wrong
Snowflake bet on the warehouse when the market wanted a lakehouse. Snowflake's core innovation was separating compute from storage in the data warehouse, making it elastic and easy to use. This was revolutionary in 2014 and drove Snowflake's massive IPO in 2020 (the largest software IPO in history at the time). But Snowflake's architecture is fundamentally a data warehouse: data must be loaded through ETL pipelines, stored in proprietary formats, and queried with SQL. As the market shifted toward lakehouse architectures that support open formats, ML workloads, and streaming data, Snowflake found itself adding lakehouse features (Iceberg tables, Snowpark, Dynamic Tables) to a warehouse architecture. This is a classic innovator's dilemma: Snowflake can't fully embrace the lakehouse without cannibalizing its core warehouse business. Databricks, with no legacy warehouse revenue to protect, went all-in on the lakehouse from day one.
Google BigQuery was too serverless and too GCP-locked. BigQuery is arguably the most technically impressive data warehouse ever built. Its serverless architecture (no clusters to manage), columnar storage engine, and integration with Google's AI/ML services make it a compelling choice for GCP-native organizations. But BigQuery has two fatal flaws. First, it's locked to GCP, and GCP has only ~11% cloud market share vs. AWS's ~31% and Azure's ~25%. Most enterprises are multi-cloud or AWS-primary, making BigQuery a non-starter. Second, BigQuery's per-query pricing (charged per TB scanned) creates unpredictable costs that terrify finance teams. Databricks' DBU-based pricing is more transparent and predictable. BigQuery is a great product in search of a larger addressable market.
AWS Redshift was too legacy to innovate. Redshift was one of the first cloud data warehouses, launching in 2012, and it captured significant early market share through AWS's distribution advantage. But Redshift's architecture (provisioned clusters, local storage, manual vacuum/sort) feels ancient compared to Snowflake's elasticity or Databricks' lakehouse. AWS has tried to modernize with Redshift Serverless and Redshift Spectrum (for querying data lakes), but these feel like patches on a legacy foundation. More damaging, AWS has seemed confused about its data strategy: should customers use Redshift, Athena, EMR, or Lake Formation? Databricks offers one unified platform. AWS offers four overlapping services that require an architecture diagram to understand.
The AI Infrastructure Wave
The data platform market is entering a new phase driven by AI infrastructure. Training and deploying LLMs requires massive data processing (for training data pipelines), feature engineering (for model inputs), experiment tracking (for model versions), and production serving (for inference). Databricks is uniquely positioned to capture this wave because AI workloads are fundamentally data workloads. You can't train a model without processing data, you can't serve a model without a feature store, and you can't improve a model without analyzing inference logs.
Databricks' Mosaic AI acquisition and the subsequent launch of foundation model training capabilities (MPT models, fine-tuning APIs, GPU cluster management) represent a direct play for the AI infrastructure budget. When an enterprise asks "where should we train our AI models?", Databricks can answer "on the same platform where your data already lives." This eliminates the data movement costs and governance complexity of using a separate AI platform (like AWS SageMaker or Azure ML) on top of a separate data platform. Snowflake's Cortex AI is a credible alternative for SQL-native AI use cases, but it lacks the deep ML infrastructure that serious AI teams need. The AI wave is Databricks' to lose.
What Indie Founders Can Learn from Databricks
- Create the category, don't compete in it. Databricks didn't try to be a better data warehouse (Snowflake was winning that battle). It invented the "lakehouse" category and defined the evaluation criteria to favor its architecture. If you're competing against an entrenched incumbent, ask yourself: can I define a new category where my approach is the only option? Category creation is the ultimate competitive moat because it lets you set the rules of the game.
- Open formats create paradoxical lock-in. Delta Lake stores data in open Parquet format that customers own, reducing adoption risk. But this openness actually creates stronger lock-in through superior tooling, optimizations, and ecosystem integration. When your product is built on open standards, customers adopt it faster (because they can leave anytime) and stay longer (because the experience is so good they don't want to). Open formats are a Trojan horse: they get you in the door, and your product keeps you there.
- Unified platforms beat point solutions at scale. Databricks combines SQL analytics, data engineering, ML, and streaming in one platform. Each workload generates data and context that benefits the others. When building your own SaaS, think about which adjacent workloads your customers are doing in separate tools. A unified platform reduces context switching, data movement, and vendor management overhead. The platform player wins the enterprise budget conversation because one invoice is always easier than four.
- Founding team credibility is a distribution moat. Databricks was founded by the creators of Apache Spark, the most popular distributed compute engine. This credibility gave them instant access to every organization using Spark (hundreds of thousands of companies). If your founding team has created a widely-adopted open-source project, standard, or methodology, that credibility is a distribution moat that no amount of marketing spend can replicate. Build your personal brand and community before you build your company.
- AI is a data problem, not a model problem. Databricks bet that the hard part of AI is data preparation, feature engineering, and model operations, not model architecture. This bet is paying off as enterprises discover that 80% of AI project time is spent on data, not models. If you're building for the AI wave, focus on the data infrastructure layer. Models are increasingly commoditized; data pipelines are the bottleneck.
The data platform market isn't winner-take-all. Snowflake will continue to dominate SQL analytics and BI workloads where simplicity and ease of use matter most. BigQuery will serve GCP-native organizations that want serverless scale. Redshift will persist as the default for AWS shops that don't want to evaluate alternatives. But for the growing segment of the market that wants a unified platform for data engineering, analytics, ML, and AI, Databricks' combination of category creation, open formats, unified analytics, Spark ecosystem, and AI integration creates structural advantages that will take years for any competitor to erode. For indie founders, the lesson is clear: define the category, embrace open standards, and build the platform your customers grow into, not the point solution they grow out of.
Get a Competitive Analysis for Your Market
Want to find your own Databricks-like position in your market? Spyglass Snapshot analyzes up to 3 of your competitors across pricing, features, positioning, and SWOT, and delivers specific recommendations for where you can win. Get your report for $9.