ETL Tools in 2026: Fivetran, Airbyte, dbt and the ELT Reality

Last updated: 25 September 2026. Original chapter published in 2018 when ETL was still the default. Fully rewritten for the ELT era.

Quick answer:

The category is called “ETL tools” for historical reasons; most modern “ETL tools” are actually ELT tools. They load raw data into the warehouse first, then let you transform it there in SQL. If you’re picking one in 2026, the real question is: managed and expensive with usable pricing, or open-source and cheaper but with more operational overhead?

Fivetran, Airbyte, dbt (for the T part), and a handful of enterprise incumbents cover 90% of what teams pick.

ETL vs ELT: why the acronym flipped

Old world: Extract, Transform, Load. Data got extracted from source, cleaned on a middleware server, then loaded into the warehouse. Storage was expensive, compute was expensive, and the warehouse couldn’t do the T well anyway.

New world: Extract, Load, Transform. Cloud storage costs pennies, and cloud warehouses (Snowflake, BigQuery, Databricks) do the T faster than any dedicated middleware. So: load raw first, transform in SQL inside the warehouse. That’s the pattern.

Full backstory in the ETL glossary entry and the ELT glossary entry.

The modern shortlist

What you’re actually choosing between in 2026, ranked by how often we see them picked.

Fivetran

The managed default. Connectors to hundreds of SaaS APIs and databases, priced per monthly active row (each distinct primary key counted once per month, per connection, per destination). Now merging with dbt Labs (announced October 2025), so ingestion and transformation are being sold as one story.

Best for: teams that want zero connector maintenance and can budget for usage-based pricing.

Watch out for: a bulk upstream migration or a schema change that shifts millions of rows can spike the bill.

Full review: Fivetran.

Airbyte

Open-source alternative. Connector variety is enormous. Cloud tier priced in credits: 6 credits per million rows for API connectors, 4 credits per GB serialized for databases and files. Self-hosted is free plus your infrastructure.

Best for: teams that want to self-host, need long-tail connectors, or want to fork the code when a connector breaks.

Watch out for: connector quality varies. The 50 most-used connectors are solid; the long tail can need Python attention.

Full review: Airbyte.

dbt (for the T)

The transformation half. SQL models, tests, docs, lineage, all in git. dbt Core is open source; dbt Cloud adds a scheduler, IDE, semantic layer, and CI features. The dbt Fusion engine (Rust rewrite, beta May 2025) is where the technical direction is heading.

Not an ingestion tool. Sits on top of Fivetran or Airbyte to model the raw tables into clean ones. Full review: dbt.

SQLMesh, Coalesce, Dataform

Three real alternatives to dbt:

  • SQLMesh: virtual environments, better handling of complex changes, column-level lineage. Managed as Tobiko Cloud. Growing fast.
  • Coalesce: visual, column-aware, Snowflake-focused. Enterprise-oriented.
  • Dataform: Google’s dbt-like tool bundled with BigQuery. Reasonable if you’re already BigQuery-native.

Meltano, dlt, Sling

Open-source ingestion options for teams that want code-first control:

  • Meltano: Singer-based, CLI-first.
  • dlt (dltHub): Python library for writing pipelines directly. Popular with data engineers who don’t want a UI.
  • Sling: CLI for database-to-database replication, lightweight.

Enterprise ETL (still relevant)

For teams with mainframe sources, SAP, or heavy compliance requirements, the traditional ETL vendors still matter:

  • Informatica: the enterprise incumbent. Salesforce acquisition announced May 2025.
  • Talend: Qlik Talend Cloud since 2023.
  • Matillion: cloud-first visual ETL/ELT.
  • Cloud-native: AWS Glue, Azure Data Factory, Google Cloud Data Fusion. Cheap if you’re already deep in one cloud; awkward across clouds.

CDC tools (real-time)

For syncing production databases with sub-minute freshness, change data capture tools tap the database’s transaction log:

  • Debezium: the open-source standard, runs on Kafka Connect.
  • Fivetran HVR: Fivetran’s enterprise CDC engine (bought in 2021).
  • Estuary Flow: real-time CDC and batch ELT in one product.
  • Qlik Replicate: enterprise CDC, formerly Attunity.

How to pick one

Three honest questions:

Do you have the sources it needs? Every tool has a connector list. Check yours are on it, at production quality, before anything else. The #1 failure mode is “we picked Tool X, then discovered our SFDC custom objects aren’t supported.”

What’s your bill going to look like at 10x current volume? Usage pricing is friendly at low volume and can bite hard when your product grows. Model the bill at 3x and 10x your current data volume. If it doesn’t fit the budget, pick the flat-rate or open-source option now, not later.

Who owns it when it breaks? Managed tools fail sometimes and you file tickets. Self-hosted tools fail sometimes and you fix them. Neither is right; both are trade-offs against staff cost.

What to skip in 2026

Categories that used to matter and no longer justify a decision:

  • Old-school data integration hubs (MuleSoft, Boomi, Dell Boomi) for analytics use cases. They’re app integration tools; the modern ELT stack is faster and cheaper for warehouse loads.
  • Custom-built Python scripts as your primary pipeline. The maintenance debt catches up. Even a small team saves time by adopting a tool.
  • Traditional row-by-row ETL frameworks for warehouse ingestion. Warehouses want bulk loads.

The reverse question: getting data out

Reverse ETL is the pattern for pushing warehouse data back into the tools where people work: CRMs, ad platforms, email. Hightouch is the standalone leader; Census was acquired by Fivetran (May 2025) and now sells as Fivetran Activations. Full picture in the glossary entry.

Common questions

Do I need both an ELT tool and dbt?

Yes, for anything serious. ELT tools load. dbt transforms. They’re the E-L and the T of ELT, sold separately. A few tools try to do both (Matillion, Keboola, Rivery), but the dominant modern stack is one tool for each.

Can I self-host Fivetran?

Partly. Fivetran’s Hybrid Deployment (Enterprise and Business Critical editions) lets the connector agent run on your own Kubernetes/Docker infrastructure. Only metadata and monthly active row counts leave your network. The alternative if you want fully self-hosted is Airbyte.

Is Fivetran or Airbyte cheaper?

Airbyte at low volume. Fivetran can be cheaper at very high volume with committed contracts. The middle band (small to mid-market) is where the pricing gets real: Airbyte’s cloud gets close to Fivetran once you’re at 100M+ rows/month, and self-hosted Airbyte trades tool cost for engineer time. Model both against your actual sources.

Do I need CDC?

Only if freshness matters. Nightly batch is fine for weekly reports. Sub-minute latency to a dashboard needs CDC. Sub-second needs streaming (Kafka, Flink, Materialize). Don’t buy freshness you won’t consume.

Where to next

For the market context, see Data Warehouses in 2026. For prices, the pricing playbook. For the ETL process itself, the older chapters at 3 Ways to Build an ETL Process and ETL Tutorial still hold up conceptually.

In the directory
Warehouses, ETL and ELT tools in the directory

Snowflake, BigQuery, Redshift, Databricks, Fivetran, Airbyte and the rest, with pricing and what each one replaces.

Browse the tools