Data Warehouse Guide

The 2026 update

Data warehouses in 2026Data Warehouses in 2026: The State of the MarketThe state of the market: who bought whom in 2024-2025, what won (Iceberg, ELT, lakehouses), what died, and where the AI stack fits.2026 warehouse pricing playbookThe 2026 Warehouse Pricing PlaybookHow Snowflake, BigQuery, Databricks and Redshift actually charge.

Cloud Data Warehousing

Data Warehouse ConceptsCloud vs Traditional Data Warehouse Concepts (2026 Update)What stayed the same and what the cloud rewrote: separation of storage and compute, elasticity, schema-on-read, ELT, and open table formats.Data Warehouse ArchitectureData Warehouse Architecture: Traditional vs. Cloud ModelsA data warehouse is an electronic system that gathers data from a wide range of sources within a company and uses the data to support management decis..Data Mart vs Data WarehouseData Mart vs. Data Warehouse: The Difference with Examples [2024 Updated]A data mart is a subset of a data warehouse oriented to a specific business line.Database vs Data WarehouseDifference Between Data Warehouse and DatabaseDoes your business deal with a lot of transactions each day?Data Lake vs Data WarehouseData Warehouse vs Data Lake vs Lakehouse: A 2026 ComparisonWarehouse vs lake vs lakehouse, explained plainly.Data Warehouse ToolsData Warehouse Tools in 2026: The Modern Stack, ShortlistedThe tools around the warehouse in 2026: ingestion, transformation, orchestration, BI, reverse ETL, observability and catalogs.Data Warehouse Testing 101Data Warehouse Testing 101: Strategy & Best PracticesAs organizations develop, migrate, or consolidate data warehouses, they must employ best practices for data warehouse testing.BI and Data WarehousingBI and Data Warehousing: Do You Need a Data Warehouse Anymore?For a long time, Business Intelligence and Data Warehousing were almost synonymous.Data Warehousing and Data Mining 101Data Warehousing and Data Mining 101In physical mining of minerals from the earth, miners use heavy machinery to break up rock formations, extract materials, and separate them from their..The New EDW: Meet the Big Data StackThe New EDW: Meet the Big Data Stack TechnologyEnterprise Data Warehouse Definition: Then and Now What is an EDW?

ETL

ETL Process3 Ways to Build ETL Process Pipelines with Examples [2024 Updated]Are you stuck in the past? Are you still using the slow and old-fashioned Extract, Transform, Load (ETL) paradigm to process data?Redshift ETLRedshift ETL: How to Load Data into AWS Redshift WarehouseIt’s easier than ever to load data into the Amazon Redshift data warehouse.Data Warehouse ETLData Warehouse Infrastructure: Full vs Incremental Loading in ETL [2024 Updated]One of the end-goals of having an effective ETL process and ETL Data Warehouse, is the ability to reliably query data, obtain insights, and generate v..ETL ToolsETL Tools in 2026: Fivetran, Airbyte, dbt and the ELT RealityThe modern shortlist for ETL and ELT in 2026: Fivetran, Airbyte, dbt, SQLMesh, Meltano, plus enterprise incumbents.ETL TutorialETL Tutorial: Get Started with ETL | Panoply GuideThe Extract Transform Load (ETL) process has a central role in data management at large enterprises.ETL TestingHow To Become a Great ETL Tester: 8 Tips for ETL TestingETL testing is an old, proud tradition, which has been practiced by large organizations for the past two decades.

Redshift

Redshift ArchitectureBuyer’s Guide to Redshift Architecture, Pricing, and PerformanceAmazon Redshift is one of the fastest growing and most popular cloud services from Amazon Web Services.Redshift Columnar Storage 101Redshift Columnar Storage 101: What is a Columnar Database?Redshift is a managed data warehouse provided by Amazon Web Services (AWS).Redshift Cluster 101Redshift Cluster 101Amazon Redshift is a popular cloud service from Amazon Web Services, which provides a fully-managed analytical data warehouse.

BigQuery

BigQuery ArchitectureA Deep Dive Into Google BigQuery Architecture: How It Works [2024 Updated]Google’s BigQuery is an enterprise-grade cloud-native data warehouse.

About this guide

2026 edition

Last full review: 25 September 2026. This guide was originally published between 2018 and 2020. It has been rewritten from the ground up in 2026 to reflect a market that now runs on Snowflake, BigQuery, Databricks and a shrinking cast of open-table formats instead of Redshift and Hadoop. The old URLs still work, and the old chapters are still there. What changed is everything about the answers.

Quick answer:

A data warehouse is a central database built for analysis rather than day-to-day operations. It pulls copies of data from every system a company runs (the product database, Stripe, Salesforce, ad platforms) into one place, cleans them up, and lets analysts query years of history without slowing anything down.

If you’ve ever asked “how much revenue did we do last quarter, by region?” and got 3 different answers from 3 different tools, the data warehouse is the fix. It’s the one place where those numbers get reconciled, and it’s the reason almost every tool in this directory exists.

What changed since the last edition

The old guide (2018 to 2020) treated Amazon Redshift as the default, ETL as the pattern, and the on-premises Teradata era as a live memory. Each of those has flipped.

  • The default warehouse is now Snowflake or BigQuery. Redshift is still fine, and best when the rest of your stack is on AWS. It stopped being the automatic starting point around 2020.
  • ELT beat ETL for analytics. Load the raw data first, transform inside the warehouse with SQL. Fivetran, Airbyte and dbt made this the standard pattern.
  • The lakehouse is real. Databricks and open table formats like Apache Iceberg let you keep files in your own cloud bucket while getting warehouse-style SQL on top. Snowflake, BigQuery and Databricks all read Iceberg now.
  • Consolidation is loud. Fivetran and dbt Labs announced a merger in October 2025 at roughly $600M ARR. Fivetran also bought Census in May 2025. Databricks bought Tabular in 2024 and Tecton in 2025. Metaplane went to Datadog. Weights and Biases went to CoreWeave. The middle of the modern data stack is shrinking to a handful of platforms.
  • The AI data stack is a real category now. Vector databases, embeddings, LLM evals and RAG pipelines sit on top of the warehouse. Every major vendor added vector search in the last 24 months.

What you’ll learn in this guide

Twenty-two chapters, grouped so you can jump in wherever you are:

  • The 2026 update (new): the state of the market and a plain-English pricing playbook for Snowflake, BigQuery, Databricks and Redshift.
  • Cloud data warehousing: the concepts, the architecture, how a warehouse differs from a database or a lake, and how to pick the right tools.
  • ETL and ELT: the movement patterns, popular tools, and how the “T” moved inside the warehouse.
  • Redshift: architecture, columnar storage, clusters. Still relevant, still where a lot of enterprise workloads run.
  • BigQuery: how Google’s serverless architecture actually works underneath.

The four jobs of a data warehouse

Every warehouse does the same 4 things, and the tools around it just do them at different quality levels.

  1. Move. Copy data from source systems on a schedule (see ETL, ELT, CDC).
  2. Store. Hold years of history in a shape that scans fast (columnar storage, lakehouse tables).
  3. Transform. Turn raw tables into clean, consistent models that everyone agrees on (see dbt and semantic layers).
  4. Serve. Answer questions fast from BI tools, apps, and now AI assistants (Metabase, Looker, embedded analytics, RAG).

The rest of this page is a pointer to the chapters. Read top to bottom if you’re new. Jump if you’re not.

Which warehouse do you actually pick in 2026?

Four names cover 90% of the market. Here they are with the honest short version, plus links to our full reviews.

WarehouseBest forPricing modelWatch out for
Snowflake Teams that want zero infrastructure work and a proper marketplace of tools Per-second compute credits, storage billed separately Warehouses left auto-suspending too long
BigQuery Google Cloud shops, event-heavy workloads, teams already in GA4 Per TB scanned (on-demand) or per slot-hour (reservations) Unfiltered SELECT * on huge tables
Databricks Machine learning shops, lakehouse-first architectures, petabyte scale DBU credits, tiered by workload type Cluster right-sizing and idle time
Amazon Redshift Teams already deep in AWS, workloads with predictable throughput RA3 node-hours or serverless RPU-hours Concurrency scaling if not tuned

Beyond the big four: ClickHouse, MotherDuck, Firebolt, Microsoft Fabric and DuckDB are all worth a look for specific jobs. We keep an up-to-date directory of warehouses with pricing side by side.

Do you actually need a warehouse?

Sooner than you’d think, and later than vendors suggest.

If your whole analytics life fits in one Postgres and Metabase pointed at it, wait. The trigger is usually the second or third data source: once you’re joining Stripe payments against product events against ad spend, doing it anywhere except a warehouse turns into duct tape.

The good news is entry costs collapsed. A small team’s warehouse bill can be under $100 a month. The horror stories about 6-figure Snowflake invoices are almost always about unmanaged compute, and that’s a governance problem you get plenty of warning about before it hits.

The honest cost picture, by team size

Rough monthly ranges from what we see on real invoices. Assumes ELT via Fivetran or Airbyte, transformation in dbt, BI in Metabase or Looker.

Team stageWarehouseIngestionBIAll-in
Solo / 2-5 people$50-200$0-500Free tier$100-800
5-20 people$300-2,000$500-2,500$0-500$1k-5k
20-100 people$2,000-10,000$2,000-8,000$1,000-6,000$5k-25k
100+ enterprise$10,000+$5,000+$3,000+Six figures and up

These are ballparks, not quotes. See the pricing playbook chapter for the actual mechanics: how Snowflake credits scale, how BigQuery bills bytes, how DBU tiers work, and where the invoice surprises hide.

The modern data stack, in one diagram

Every arrow below has 5 to 30 vendors competing to be your default. This guide covers the warehouse in the middle. The rest is in the directory.

  Sources           Move                 Store                Transform          Serve
  ─────────         ─────                ──────               ──────────         ─────
  Production DB  →  Fivetran         →   Snowflake        →   dbt            →   Metabase
  Stripe            Airbyte              BigQuery             SQLMesh            Looker
  Salesforce        Debezium (CDC)       Databricks           Coalesce           Tableau
  HubSpot           Kafka (stream)       Redshift             Semantic layer     Sigma
  GA4               Segment              ClickHouse           (Cube, dbt SL)     Hex
  Ad platforms      Snowplow             Iceberg tables                          AI assistants
                                                                                 (RAG on top)

Where the AI stack fits

The warehouse is the ground truth. Every serious “chat with your data” product either reads from the warehouse or pushes retrieval through a semantic layer sitting on top of it. Vector search is now available inside Snowflake, BigQuery and Databricks natively, and RAG pipelines are becoming a normal pipeline layer next to transformation.

See our chapters on vector databases, embeddings, and RAG in the glossary.

Who this guide is for

Three audiences, one guide:

How we write and update this guide

Every chapter carries a Last updated line. Every price, credit rate and product claim is verified against the vendor’s own docs on that date, not from memory. When a vendor changes a plan or ships a new engine, we update the chapter and note it in the changelog.

We don’t take money to move a tool up a list. We do take money for labelled sponsored placements in the directory. Any editorial position with a commercial interest attached says so on the page. When something we wrote turns out to be wrong, we fix it visibly rather than quietly.

Common questions

Is a data warehouse the same as a database?

No. Both are databases in the technical sense, but they’re built for opposite jobs. A production database handles thousands of small operations per second (fetch a user, update an order) and can’t be trusted to keep years of history. A warehouse handles a few enormous scans per second and holds everything. Full comparison: Database vs Data Warehouse.

Warehouse or data lake?

Both, most likely. A warehouse gives you fast SQL on structured tables; a lake gives you cheap storage for anything. Modern setups keep raw data in the lake and modeled tables in the warehouse. The lakehouse collapses that into one layer where possible. Full read: Data Lake vs Data Warehouse.

How much does a real warehouse cost?

See the table above and the 2026 pricing playbook. Short version: a startup can start under $100 a month. A 100-person company usually lands in the $5,000 to $25,000 a month range. Enterprise starts at 6 figures.

Do I need to hire a data engineer to run one?

Not to start. Managed ELT, dbt Cloud, and a hosted BI tool let a technical founder run the whole stack solo for a while. You’ll want a first data hire once you’re joining more than 5 sources or once dashboards start driving real decisions.

Is Snowflake really the default, or is that just marketing?

Broadly yes, at mid-market. BigQuery leads in GCP-first shops and event-heavy analytics. Databricks leads where ML is the point. Redshift leads inside AWS shops that already run half their stack there. Any of the 4 is a defensible choice; the wrong choice is usually about pricing model fit rather than technology.

Will AI replace the data warehouse?

No. AI needs a warehouse more than warehouses need AI. LLMs are terrible at inventing facts and great at summarizing what’s in front of them (see hallucination). Grounding an AI assistant in your warehouse and semantic layer turns a party trick into a product. That’s why every warehouse vendor is racing to add vector search, and every AI product is racing to add warehouse connectors.

Start reading

The chapter cards above are the whole guide, in order. If you want a curated path, follow the “who this guide is for” section. If you’re checking pricing, jump to the playbook. If you want the current lay of the land, read the state-of-the-market chapter.

In the directory
Warehouses, ETL and ELT tools in the directory

Snowflake, BigQuery, Redshift, Databricks, Fivetran, Airbyte and the rest, with pricing and what each one replaces.

Browse the tools