Cloud vs Traditional Data Warehouse Concepts (2026 Update)
Last updated: 25 September 2026. Original chapter published before Snowflake’s IPO. Rewritten to match the current reality.
Quick answer:
The concepts that mattered in the on-premises era (partitioning, indexing, denormalization, star schemas, MPP) still matter, but the cloud rewrote the economics underneath them. The single biggest change is that storage and compute are separately priced and independently scaled. That one shift explains most of what looks different about a modern warehouse.
What stayed the same
Some fundamentals are timeless. If you’re transitioning from a traditional warehouse to a cloud one, these still apply:
- Columnar storage is still the reason warehouses are fast. Row-by-row storage (like a production database) is wrong for analytics; column-by-column storage (Parquet, Snowflake internal, BigQuery Capacitor) is right.
- Denormalization is still often the answer. Wide, joined tables scan faster than deeply normalized ones. A star schema (facts + dimensions) is still a good default.
- Time-based partitioning still cuts query cost enormously. Partition by day, filter by day.
- ETL discipline still matters. Bad data in = bad reports out, regardless of how fast the warehouse is.
- The “one number, one definition” problem still isn’t a tech problem. It’s a governance problem.
What the cloud rewrote
Separation of storage and compute
The traditional warehouse (Teradata, Netezza, on-premises Oracle) tied storage to compute nodes. Bigger tables meant bigger clusters, whether you needed the query power or not. Upgrades meant buying capacity for peak load and paying for it at 3 a.m.
Cloud warehouses split those bills. Snowflake stores everything on S3-equivalent object storage, and spins up virtual warehouses (compute) only when queries run. BigQuery is even more radical: it’s serverless, and Google decides which physical machines run your query. Scale storage to petabytes, scale compute to zero when nobody’s querying.
The practical effect: you don’t over-provision anymore. And your bill mostly comes from compute, not storage.
Elasticity replaces sizing
Traditional warehouse projects started with capacity planning: how many nodes, how much RAM, how many years of headroom. Get it wrong and you either overpaid or ran out.
Cloud warehouses scale on demand. Snowflake’s warehouse sizes range XS to 6XL, resizable in seconds. BigQuery slots can be added and removed via reservations. Redshift Serverless auto-scales inside a spend limit.
The new discipline is cost monitoring, not sizing. See the pricing playbook.
Schema-on-read became viable
Traditional warehouses enforced schema on load. Structured, typed, planned. That works for known data; it breaks for exploratory or messy sources.
The cloud made schema-on-read practical. Cheap object storage means you can dump raw JSON or CSVs and figure out the structure later. That’s the data lake pattern, and increasingly the lakehouse pattern on top of it.
ELT flipped the order
Traditional ETL: extract, transform, load. Middleware server did the T. Only cleaned data reached the warehouse.
Cloud ELT: extract, load, transform. Load raw, transform in SQL inside the warehouse. Cheaper storage + powerful warehouse compute made this the new default. Fivetran and Airbyte handle EL; dbt handles T.
The consequence: replayability. Business logic changed? Re-run SQL. No need to re-extract from source. Full breakdown in the ELT glossary entry.
Storage started going open
Traditional warehouses used proprietary formats. Leaving meant exporting everything. Cloud warehouses inherited that pattern (Snowflake’s internal format, BigQuery’s Capacitor), but in 2024-2025 the model shifted.
Apache Iceberg emerged as the open table format that every warehouse can read. Snowflake, BigQuery, Databricks, Redshift, Trino all support it. Store data once in your own S3 bucket, query it from whichever engine you prefer. The vendor keeps the compute; you keep the storage.
This is the biggest architectural shift of the decade, and most companies haven’t caught up to it yet.
Traditional vs cloud, side by side
| Traditional (Teradata era) | Cloud (2026) | |
|---|---|---|
| Storage/compute | Tied to each node | Separately priced |
| Scaling | Buy hardware, wait weeks | API call, seconds |
| Pricing | CapEx, one-time + maintenance | OpEx, per second/query/TB |
| Schema | Enforced on write | Both, depending on layer |
| Pattern | ETL (transform first) | ELT (transform in warehouse) |
| Storage format | Proprietary | Increasingly Iceberg (open) |
| Setup time | Months to a year | Hours |
| Failure mode | Over-provisioning waste | Runaway query bills |
| Best for | Predictable, regulated, on-prem-required | Everything else |
The mental model that helps
Think of a traditional warehouse as owning a factory: capital investment upfront, fixed capacity, maintenance overhead, but predictable output. Think of a cloud warehouse as renting factory time: pay for exactly the runs you do, scale up or down instantly, but the meter is always running.
Both are legitimate. The choice is about which risk you’d rather manage: over-buying capacity vs runaway usage bills.
When traditional still wins
On-premises warehouses aren’t extinct. They’re the right call in specific cases:
- Regulatory requirements that prohibit cloud (some government, defense, healthcare).
- Data sovereignty constraints that make cross-border cloud awkward.
- Extreme, predictable workloads where 5 years of hardware amortization beats 5 years of cloud bills.
- Existing investments that haven’t fully depreciated. Migrating a working Teradata environment is a multi-year project.
For 90% of new projects in 2026, cloud is the answer. For most existing traditional workloads, a hybrid path (keep Teradata for stable stuff, use Snowflake for new stuff, unify via Iceberg later) beats a rip-and-replace.
Common questions
Is on-premises really dying?
The market is shrinking, not zero. New investment overwhelmingly goes to cloud. Existing on-premises workloads continue for years because migrations are hard. Expect the “traditional” category to slowly become a niche of regulated and heavy-compute use cases rather than the default.
Does cloud really pay off vs owned hardware over 5+ years?
Depends entirely on workload variability. Predictable, always-on workloads at scale can be cheaper on owned hardware after year 3. Spiky or growing workloads are almost always cheaper on cloud because you don’t pay for peak headroom. Committed cloud contracts narrow the gap further.
What about hybrid?
It’s real. Snowflake, Databricks, and BigQuery all offer hybrid or on-premises variants (Snowflake via BYOC arrangements, Databricks via anywhere-cloud, Google Distributed Cloud). Fivetran ships Hybrid Deployment for regulated sources. The pure-cloud vs pure-on-prem split isn’t the only choice.
Where to next
For architecture specifics, see Data Warehouse Architecture. For the tools, Data Warehouse Tools. For the current market, Data Warehouses in 2026.
Snowflake, BigQuery, Redshift, Databricks, Fivetran, Airbyte and the rest, with pricing and what each one replaces.
Browse the tools