Union

A commercial platform built on Flyte for coordinating data, machine learning, and AI workloads in your own infrastructure. Best suited to engineering teams with demanding execution and recovery requirements.

Best for: Engineering teams running complex data and ML pipelines with demanding recovery and compute requirements

Editor’s note: This review weighs Union’s pricing, documented capabilities, and deployment responsibilities.

Quick verdict: I recommend Union for engineering teams whose machine learning pipelines have become expensive to recover and awkward to operate. I like the control it gives you over distributed work. The cost and infrastructure commitment make it harder to recommend for a handful of scheduled Python scripts.

In this Union review, I’ll look at the costs, setup, and recovery features that matter when choosing a managed orchestrator for data and ML workflows.

Key Takeaways

  • I like Union most for complex ML pipelines with different resource needs at each step.
  • Flyte provides the execution foundation, while Union adds commercial deployment and operational capabilities.
  • Recovery and reusable outputs can reduce repeated work, provided you configure tasks carefully.
  • This is an engineering platform. You still need to understand code, dependencies, and your infrastructure.
  • Small teams should be especially careful about the spending commitment and the work involved in running their own cluster.

Pros and Cons

Pros

  • Python tasks can have their own resource and dependency settings.
  • Task recovery and output caching help avoid repeating expensive work.
  • Production workloads can run inside your own cloud.
  • Reusable containers suit frequent tasks with expensive initialization.
  • Flyte provides an open-source foundation.

Cons

  • A substantial spending commitment for light workloads.
  • Self-managed deployment leaves infrastructure operations with your team.
  • Caching needs explicit configuration and careful invalidation.
  • The platform adds concepts that simple scheduled scripts may not need.

How Much Does Union Cost?

Union needs an infrastructure budget. The minimum makes occasional experimentation a harder sell.

  • Team ($950 monthly minimum): for a single-cluster team; credited against usage after a 30-day trial.
  • Enterprise (custom pricing): for larger deployments and more specialized controls.
PlanConcurrent actionsActions per runClusters
Team1,00025,0001
EnterpriseCustom50,000+3+

An action is a unit of work, including a cache hit. Retries do not add action counts, but compute is metered separately. Entry rates are $0.003 per action, $0.0288 per vCPU-hour, and $0.004 per GB-hour of memory; GPU charges vary by class.

Optional managed BYOC deployment costs another $2,500 per environment monthly, outside the minimum. Self-managed infrastructure has no deployment fee.

Is Union Good Value for Money?

I judge its value against two costs: the engineer time spent maintaining execution infrastructure and the compute wasted when pipelines fail. A team with a reliable, modest workload may have little to save in either category.

Before committing, check these three things:

  • Model a representative month: include quiet periods as well as the busiest pipeline.
  • Separate platform and infrastructure spending: ask for the complete bill for your proposed deployment.
  • Assign operational ownership: a lower vendor bill helps less if your engineers inherit substantial maintenance.

I recommend evaluating Team for a single production cluster with clear recovery problems. Discuss Enterprise before a rollout that needs multiple clusters or a self-hosted control plane. The deployment requirements should decide the plan.

Getting Started With Union

I like that you can begin with a locally executed, tracked Python task before committing to a production cluster. That makes it possible to learn the execution model separately from the infrastructure rollout. It does not mean the production deployment is equally lightweight.

The current Flyte 2 approach uses a TaskEnvironment to describe the environment a task needs. You write Python functions, declare their resources and dependencies, and connect work using ordinary Python control flow.

For someone who already owns Python pipelines, that is an attractive starting point. The bigger decision is where to draw the boundaries between tasks. Split too little, and a failure may force an unnecessarily large chunk of work to repeat. Split everything into tiny pieces, and the orchestration becomes harder to reason about.

Start your evaluation with one existing pipeline:

  1. Choose a pipeline with a known problem. An unreliable batch inference job is more useful than a toy example that never fails.
  2. Separate meaningful stages. Data preparation, inference, and result validation are sensible candidates when they have different resource needs.
  3. Declare dependencies and resources. Make the execution environment reproducible instead of relying on whatever happens to be installed locally.
  4. Run locally, then remotely. Compare the development workflow before expanding the infrastructure footprint.
  5. Exercise the failure path. Decide what should retry, what should reuse an earlier result, and what must stop for investigation.

That last step matters most to me. A successful run proves that the pipeline can execute. A useful evaluation also establishes whether another engineer can understand and recover it when something goes wrong.

Teams moving from older Flyte examples should also check which generation of the SDK they are following. Flyte 2 uses tasks calling tasks, replacing the older separate workflow decorator. Following the wrong generation creates avoidable migration work.

Recovery and Caching Are the Real Reasons to Buy

Union’s recovery model is more compelling to me than the promise of running Python in the cloud. Most teams already have a way to do that. The harder problem is keeping a long, distributed process understandable when only part of it fails.

Durable task execution gives you a way to retain progress instead of treating the entire pipeline as one disposable job. For an ML workflow, that can mean separating an expensive preparation stage from a later step that calls an unreliable external service.

Recoverability still needs safe application logic. If a task sends a notification or writes to another system, rerunning it safely requires thought about duplicate effects. An orchestrator does not automatically make your business logic safe to repeat.

Caching Is Useful, but It Is a Decision

Flyte’s task caching is disabled by default. You enable it for suitable tasks and choose how their results should be reused.

I prefer that to treating every output as reusable. Caching is helpful when the same inputs and code should produce the same result. It is much less reassuring when a task quietly reads changing external data or performs a side effect.

For example, consider a feature-generation step that reads a database table. Reusing its output may be sensible when the input identifies a fixed dataset version. If the input is simply the name of a table that changes throughout the day, the cached result could be stale even though that input has not changed.

Reviewer’s Notes: My first candidates for caching are expensive, repeatable transformations over versioned data. Live-data reads and external writes need a clear freshness or duplicate-handling rule before reusing their outputs.

Prefect is worth comparing if your priority is coordinating ordinary Python jobs. My stronger reason to choose Union is combining recovery with different execution environments for data preparation, model training, and inference.

GPU Workloads Benefit More Than Simple Schedules

A pipeline that prepares data on CPUs and then runs inference on GPUs is a better fit for Union than a collection of identical cron jobs. Task-level resource configuration lets the execution requirements follow the work.

I like that separation because the expensive hardware does not need to define every stage of the pipeline. It also makes resource decisions easier to review: the team can ask why one task needs a GPU or a larger memory allocation without debating the whole application at once.

Union’s reusable containers address another specific problem: repeatedly initializing short tasks. Keeping an environment available can be useful when loading dependencies or preparing the runtime takes a disproportionate share of each invocation.

I recommend reuse selectively. Infrequent, long-running jobs have less to gain, and tasks that require strong isolation deserve a different design. Reuse is an optimization with a suitable workload, not a reason to put everything into a shared environment.

Before increasing capacity, check:

  • Task size: is orchestration overhead large compared with useful work?
  • Initialization: does each invocation repeat costly setup?
  • Resource allocation: are tasks requesting substantially more than they use?
  • Isolation: can the workload safely reuse an execution environment?

If your requirement is primarily to run Python functions on managed compute, Modal deserves a place on your shortlist. Union becomes more interesting when execution is part of a larger pipeline with recovery, infrastructure placement, and operational controls to coordinate.

The same separation helps with agents. A lightweight agent can call an embedding tool that runs as its own GPU task, with its own resources and retry settings. Reports can show model turns, tool calls, errors, and token usage, which gives engineers something concrete to inspect when a run fails.

I see this as a useful extension for teams already deploying agents. It does not remove the work of designing the agent or checking its decisions. Union supplies execution controls; your team still owns the behavior.

Keeping Workloads in Your Cloud Comes With Choices

Union offers a route for teams that want their workloads to stay inside their own cloud environment. I see that as a practical advantage when infrastructure ownership and data location already shape purchasing decisions.

The deployment choice deserves as much attention as the workflow features. With a self-managed cluster, your team remains responsible for the cluster’s operation. With managed BYOC, Union takes on deployment operations inside your cloud account. Those are different divisions of responsibility, even when the application code is similar.

The platform team should own this decision. A data scientist may reasonably prefer the shortest path to running a model, while the infrastructure team needs to know who handles upgrades, capacity, and incidents.

The control plane coordinates execution while your code, data, and logs stay in your cloud. I like that separation for teams with strict data-location requirements. Managed BYOC still gives Union access for cluster maintenance, so review the permissions as well as where the data sits.

Staffing should settle the choice. A team with an established Kubernetes operating model may find self-managed deployment a sensible fit. A team buying Union specifically to reduce infrastructure work should be cautious about choosing an arrangement that leaves that same work on its own backlog.

How Does Union Compare With Other Orchestrators?

For smaller workloads, the entry price alone makes Prefect and Modal worth a look. For established infrastructure teams, the cost of changing platforms matters just as much.

  • Prefect: my first comparison for teams coordinating Python workflows that do not have a strong reason to adopt Union’s ML-oriented infrastructure model. Its free Hobby tier covers two users and five deployments; the $100/month Starter plan supports bringing your own compute.
  • Apache Airflow: worth keeping when scheduled batch workflows and an existing Airflow estate dominate your needs. Its scheduling, backfills, and operator ecosystem remain good reasons to stay put when the system works.
  • Modal: worth evaluating when managed execution for Python and GPU workloads is the main purchase. Starter has no monthly platform fee and includes $30 in monthly credits. Compare complete workload costs because the deployment models differ.
  • Flyte: the open-source foundation to consider if you want direct control and have the engineering capacity to operate it. Union earns its place when its commercial capabilities, such as reusable containers or managed cluster operations, address a specific need.

Switching orchestrators creates work of its own. If your current platform is dependable, the question is whether Union solves a recurring problem well enough to justify migrating code and retraining the team.

How I Reviewed Union

I evaluated Union’s documented execution model, deployment responsibilities, and commercial terms against common data and ML pipeline requirements. My recommendations weigh operational fit and cost; they do not rely on measured speedups or customer savings claims.

Prices current as of October 2026.

Should You Choose Union?

I recommend Union for teams facing expensive pipeline failures, mixed compute requirements, or infrastructure operations that are holding back ML delivery. Its strongest selling point is making those workloads easier to divide, recover, and place on suitable infrastructure.

I would skip it if you only need to schedule a few Python jobs. Start with the specific failure, maintenance task, or deployment constraint you want to remove, and make Union demonstrate its value against that requirement during the evaluation.

Frequently Asked Questions

Is Union the Same as Flyte?

No. Flyte is the open-source execution foundation. Union is the commercial platform built around it, with additional operational and deployment capabilities.

Can Union Run in My Own Cloud?

Yes. Deployment options include running workloads in your own cloud, with different choices for who operates the infrastructure.

Is Union Suitable for Beginners?

I do not recommend it as a first automation tool. It is a better fit for engineers who already understand Python pipelines and need more control over how they run.

Spotted a wrong price or a missing integration? Send a correction. A human reads every one.

Similar tools

Other tools in the same category, with the same card and the same honest pricing.

Usage-based paid service; free trial

Astronomer

Managed orchestrator

Managed Apache Airflow through Astro, with local development tools, dbt orchestration, observability options, and enterprise deployment controls.

Visit site