Debezium

Open-source change data capture software that streams database changes through Kafka Connect, Debezium Server, or an embedded Java engine.

Best for: Engineering teams building database event streams and application integrations, with the expertise to operate CDC infrastructure.

Editor’s note: This Debezium review assesses its capabilities, deployment choices, and operating requirements. It does not include performance benchmarks.

Quick verdict: I recommend Debezium for engineering teams that want control over how database changes reach their applications and data systems. Its open-source software gives you flexible deployment choices with no license fee, but you need engineers who can keep the pipeline running.

My main takeaway from this Debezium review is that its change data capture (CDC) suits applications that depend on database events. I would be more cautious if you simply want a few tables copied into a warehouse.

Key Takeaways

  • Debezium captures database changes so downstream systems can react to inserts, updates, and deletes.
  • The software has no license fee, but hosting and engineering work still need a budget.
  • You can use Kafka Connect, Debezium Server, or an embedded Java engine to suit your architecture.
  • Connector behavior varies by database, so source compatibility deserves more attention than the headline integration list.
  • Default delivery is at least once. Your downstream processing needs to handle possible duplicates.

In this review, I’ll take a closer look at the costs, setup choices, and production responsibilities that determine whether Debezium is a good fit for your team.

Debezium official homepage introducing change data capture.
Debezium presents itself as an open-source change data capture platform. Source: Debezium; screenshot captured by Panoply using PeekShot.

Debezium Pros and Cons

Pros

  • No software license fee under Apache License 2.0
  • Database change capture with inserts, updates, and deletes
  • Kafka Connect, Server, and embedded Engine deployment options
  • Supported incremental snapshots can run alongside change streaming
  • Outbox event routing supports application messaging

Cons

  • Self-managed deployments require infrastructure and database expertise
  • Default at-least-once delivery can produce duplicate events
  • Database connectors have different prerequisites and limitations
  • Capturing changes does not finish your destination integration

Debezium Pricing: Free Software, Paid Operations

Debezium costs $0 in software license fees. There is no paid Debezium tier you need to unlock before using the open-source project in production.

I like the absence of a subscription, but I would not approve a Debezium project with a budget that stops at zero.

  • Open-source Debezium ($0 license): for teams that want to deploy and operate the software themselves.
  • Infrastructure and operations (variable): compute, storage, messaging, monitoring, and the people responsible for the deployment.
  • Third-party managed services (separately priced): for buyers who want a provider to operate some of the infrastructure. These are separate products, not Debezium upgrade plans.
Cost areaPrice basisWhat you need to budget for
Debezium software$0 license feeConfiguration and integration work
Runtime and messagingDepends on deploymentHosting, storage, and the chosen messaging service
Engineering and maintenanceDepends on your teamMonitoring, upgrades, recovery, and destination behavior
Optional commercial serviceProvider-specificThe provider’s supported connectors and operating scope

Is Debezium Good Value for Money?

I see the strongest value when you already have engineers responsible for data infrastructure. Adding CDC to an established platform is a different commitment from creating a new operating team just to avoid a subscription.

For example, an application team that needs database events for several services has reasons to invest in reusable infrastructure. A small analytics team that only needs reporting data has a much weaker case for taking on that work.

Fivetran gives you a useful comparison: its managed connectors charge using monthly active rows, with distinct changed rows counted within each connection. Debezium’s own license does not meter those rows, but that does not prove it will be cheaper overall.

My recommendation is to choose the free software when control is part of the requirement and operational ownership is already assigned. If your priority is reducing maintenance, price a managed alternative before committing.

Getting Started With Debezium

Start with your source database and destination, then choose the runtime. I would resist building the entire pipeline around Kafka before checking whether you actually need it.

You will need a database owner involved in setup, particularly for replication settings and permissions.

My recommended setup sequence is:

  1. Check the source connector. Confirm that your database version and deployment support its prerequisites.
  2. Prepare database access. Configure the required replication or CDC facilities and permissions with the database owner.
  3. Choose the runtime. Use Kafka Connect for a Kafka pipeline, Server for a supported direct destination, or Engine for custom Java integration.
  4. Limit the initial scope. Select a small set of representative tables and decide how existing records should enter the pipeline.
  5. Check the receiving system. Verify how it interprets keys, updates, and deletions before expanding the capture scope.

PostgreSQL illustrates why the first two steps matter. Its connector uses logical replication, and the database’s replica identity settings affect which old values appear in update and delete events. You should not assume every update includes a complete copy of the previous row.

If your application compares old and new values, settle that requirement before building consumers around an event format.

I would also make the first evaluation include a delete and a connector restart. A pipeline that handles a new row correctly has only demonstrated one part of the job.

Reviewer’s Notes: Agree on the required behavior at the destination before selecting the connector configuration. Copying current table contents and delivering business events can require very different handling of the same source change.

Capturing Database Changes Without Repeated Full Exports

Debezium is most appealing when downstream systems need changes as they happen. Its source connectors include MySQL, PostgreSQL, MongoDB, SQL Server, and Oracle, although their requirements and capabilities differ.

Debezium MySQL documentation showing incremental snapshot chunk size and watermark settings.
MySQL incremental snapshots use configurable chunks and watermark settings. Source: Debezium; screenshot captured by Panoply using PeekShot.

I like the fit for updating a search index or feeding an application that reacts to changing records. Those are situations where a change stream has a purpose beyond keeping a reporting table reasonably fresh.

Snapshots handle the existing data that predates your connector. Ongoing streaming handles subsequent changes. That distinction matters: choosing a CDC tool does not eliminate the initial work of bringing a destination up to date.

The MySQL connector’s incremental snapshots are particularly useful. They can scan selected tables in chunks while streaming continues, and an interrupted incremental snapshot can resume. You can also request a snapshot when adding a table to the capture scope.

I would value that flexibility more than an unqualified speed claim. It gives you a practical way to expand a running pipeline without treating every added table as a fresh deployment. You still need to assess the load on your source database and check the equivalent behavior for your chosen connector.

Compared with Airbyte‘s managed offering, Debezium puts more of the pipeline design in your hands. That is an advantage when your consumers need a particular event structure, and extra work when you mainly want a configured source-to-destination connection.

Choosing Kafka Connect, Server, or Engine

Kafka is optional with Debezium. This is one of the most useful distinctions to understand before estimating the infrastructure involved.

Debezium architecture diagram showing MySQL and PostgreSQL source connectors feeding Kafka and sink connectors.
The documented Kafka Connect path moves database changes through Kafka to downstream systems. Source: Debezium; screenshot captured by Panoply using PeekShot.
RuntimeBest fitMain responsibility
Kafka ConnectTeams using Kafka topics and connectorsOperate Connect and arrange downstream delivery
Debezium ServerTeams sending changes to a supported messaging destinationConfigure and operate the Server deployment
Debezium EngineDevelopers embedding CDC in a Java applicationOwn event handling, offsets, and application lifecycle

Kafka Connect is the familiar choice if Apache Kafka is already part of your architecture. Debezium supplies the source-side change events, while sink connectors or consumers handle what happens afterward.

Server is worth considering if your destination is a supported service such as Amazon Kinesis, Google Cloud Pub/Sub, or Apache Pulsar. I like having that choice because it can avoid introducing Kafka solely to host a CDC pipeline. Server remains software you operate, however.

Engine is the option I would reserve for a clear application requirement. Embedding it in Java gives you direct event handling, but also makes persistent offsets, errors, and shutdown behavior part of your application’s responsibility.

My default recommendation is to match the runtime to infrastructure your team can already support. Fewer unfamiliar components can be more valuable than theoretical flexibility you never use.

Turning Row Changes Into Useful Application Events

I especially like the outbox event router for application teams. It turns captured outbox records into routed business messages.

Debezium outbox event router documentation showing the capture and routing steps.
The outbox event router captures an outbox table and applies a message transformation. Source: Debezium; screenshot captured by Panoply using PeekShot.

An order service, for example, can save an order and its event record in the same database transaction. Debezium captures the outbox record and routes the message.

This reduces the inconsistency that can arise when saving data and separately sending a message. Your team can define an “order created” event that consumers understand without interpreting raw table changes.

There is application work involved. Your developers need to create the outbox records and agree on their format. The transformation must be applied selectively to the appropriate records, and MongoDB uses its own outbox transformation.

I would shortlist Debezium for this use case even without a warehouse project. It gives application developers a reason to invest in CDC beyond the free license.

Running Debezium in Production

Recovery is where I would spend most of the evaluation effort. The default delivery guarantee is at least once, so a failure and restart can lead to repeated events.

Your destination must tolerate duplicates. For instance, repeatedly setting a record to its latest state is a different operation from repeatedly adding the same payment amount to a total.

Selected connectors support Kafka Connect’s exactly-once capabilities with configuration and correctness caveats. I would still assess the sink and consumers separately before trusting the entire pipeline with actions such as charging a customer.

Monitoring and Database Health

Debezium exposes monitoring through JMX and connector metrics, with documented Prometheus and Grafana options. I like having those integration points; you still need to turn them into alerts someone will act on.

My operational checklist would include:

  • Capture lag: whether downstream data is falling behind the source.
  • Snapshot progress: whether an initial load or backfill is advancing.
  • Errors and restarts: whether the connector is repeatedly failing.
  • Source log retention: whether the database is accumulating logs the connector still needs.

PostgreSQL replication slots can retain WAL, so a stalled or poorly progressing consumer can become a database storage concern. PostgreSQL’s connector also does not receive DDL change events through logical decoding. I would involve the database team in both recovery planning and schema-change procedures.

Support and Ownership

The community provides documentation, issue tracking, chat, and a mailing list. I would use these for troubleshooting and implementation guidance, but they are not a purchased response-time commitment.

For a self-managed deployment, I would name an owner for connector upgrades and incident response before expanding beyond the first source. Otherwise, a successful pilot can leave you with a service nobody has agreed to maintain.

If you need a provider accountable for connector operations, Confluent is worth considering. Its managed PostgreSQL CDC Source V2 connector uses Debezium internally, although database preparation and connectivity remain your responsibility.

How Does Debezium Compare to Alternatives?

I would compare the operating model first. Start with how much responsibility you want your team to retain.

  • Fivetran: My preference for a shortlist centered on managed data movement into analytics destinations. Compare its recurring usage bill with the engineering and infrastructure you would otherwise fund.
  • Airbyte: Worth considering when your requirement is a broader replication platform with self-managed and managed deployment options. Choose the deployment deliberately; self-hosting still needs an owner.
  • Confluent Cloud: A strong option when you want managed Kafka infrastructure and a supported managed CDC connector. Its Debezium-based PostgreSQL connector shows that choosing a service does not necessarily mean abandoning Debezium technology.
  • Apache Kafka: A companion technology, rather than a like-for-like CDC replacement. Choose it for the messaging architecture; use a source connector such as Debezium to capture database changes.

I would keep Debezium ahead of these options when owning the change stream is part of the product requirement. For routine warehouse ingestion with limited engineering capacity, I would start with the managed services.

How I Reviewed Debezium

I evaluated source compatibility, deployment choices, snapshot behavior, delivery guarantees, and operating costs, then compared the responsibilities with managed alternatives. My recommendations reflect those trade-offs rather than measured throughput or installation-time results. Pricing basis checked October 2026.

Should You Choose Debezium?

Choose Debezium if your team needs a configurable CDC foundation and can own it in production. I especially like its fit for application event streams and organizations with established data infrastructure.

I would skip a self-managed deployment if your main goal is getting reporting data into a warehouse with minimal maintenance. The free license is valuable, but it does not compensate for an operating model your team cannot support.

My deciding question is straightforward: who will recover the pipeline when it stops? If you have a clear answer, Debezium deserves a place on your shortlist.

Debezium FAQ

Is Debezium free?

Yes. Debezium’s open-source software has no license fee. You still pay for the infrastructure and engineering needed to run it, and any commercial service you choose has separate pricing.

Does Debezium require Kafka?

No. Kafka Connect is a common deployment, but Debezium Server supports other destinations and Debezium Engine can run inside a Java application.

Is Debezium an ETL tool?

Its primary job is change data capture. It can be part of an ETL or ELT pipeline, but you still need to arrange the transformations and destination behavior your use case requires.

Does Debezium guarantee exactly-once delivery?

Not by default. Debezium provides at-least-once delivery. The selected connector, runtime, and configuration determine further guarantees; your destination needs its own assessment.

Is Debezium better than Fivetran?

I prefer Debezium when control over database events and deployment is the priority. I would shortlist Fivetran when managed data movement and reduced operational work matter more.

Questions people ask

Is Debezium free?

Yes. Debezium's open-source software has no license fee. You still pay for the infrastructure and engineering needed to run it, and any commercial service you choose has separate pricing.

Does Debezium require Kafka?

No. Kafka Connect is a common deployment, but Debezium Server supports other destinations and Debezium Engine can run inside a Java application.

Is Debezium an ETL tool?

Its primary job is change data capture. It can be part of an ETL or ELT pipeline, but you still need to arrange the transformations and destination behavior your use case requires.

Does Debezium guarantee exactly-once delivery?

Not by default. Debezium provides at-least-once delivery. The selected connector, runtime, and configuration determine further guarantees; your destination needs its own assessment.

Is Debezium better than Fivetran?

I prefer Debezium when control over database events and deployment is the priority. I would shortlist Fivetran when managed data movement and reduced operational work matter more.

Spotted a wrong price or a missing integration? Send a correction. A human reads every one.