Purpose

This post gives you a reliable way to keep orders, payments, and inventory in sync when one service in a multi-step flow fails. The outcome we want is simple. No customer gets charged for something they won't receive, and nobody on the team has to fix production data by hand at 2 AM.

Background

Take a typical e-commerce Order-to-Delivery flow split across four services: Orders, Payments, Inventory, and Shipping. Each service owns its own database, which is standard practice in microservices for good reasons. Each team controls its own data and schema, picks the database that suits its workload, scales independently, and a failure in one service doesn't take the others down with it.

This kind of setup usually gets built quickly against a tight deadline. The services call each other over HTTP, and it all gets deployed on a Friday afternoon because the ticket is due.

That night, a customer orders a $2,000 espresso machine. The Order service creates the record. Payments charges the card. Then Inventory fails, because someone is running a backup during peak hours. The customer has paid, nothing is reserved, and the Order service reports success. Someone on the team spends the next four hours running SQL by hand against three production databases, trying to work out what to refund and what to delete.

Problem statement

Inside a single database, a transaction gives you four guarantees, known as ACID. Atomicity means everything happens or nothing does. Consistency means data only moves between valid states. Isolation means concurrent transactions can't see each other's half-finished work. Durability means committed changes survive a crash.

Split the flow across four databases and you lose those guarantees for the flow as a whole. Each service still commits its own work, but nothing ties the commits together. A failure halfway through leaves the system half-finished, with nothing in place to clean up.

The code behind this usually looks like this:

public class OrderService(
    IOrderRepository orders,
    IPaymentClient payments,
    IInventoryClient inventory,
    IShippingClient shipping)
{
    public async Task CreateOrderAsync(OrderRequest request)
    {
        var order = await orders.SaveAsync(new Order(request));
        await payments.ChargeAsync(order.Id, request.Amount);
        await inventory.ReserveAsync(order.Sku, order.Quantity);
        await shipping.ScheduleAsync(order.Id);
    }
}

If ReserveAsync() throws, the order is already saved and the money is already taken. Nothing undoes either one.

Losing isolation causes trouble even when nothing fails. While a flow is running, other requests can see its in-between state, which leads to three kinds of bugs:

  • Lost updates. Two flows change the same record, and one silently overwrites the other.
  • Dirty reads. A flow reads data that another flow changed but hasn't finished with, such as stock reserved for an order that's about to be rolled back.
  • Fuzzy reads. Two steps in the same flow read the same record and get different answers, because something changed it in between.

On top of that, the constraints are real. Most teams don't have the people or time that companies like Netflix put into distributed transaction tooling. They have a deadline next week, and chained HTTP calls are the fastest thing to ship. The data corruption shows up months later.

Possible solutions

Option 1: Distributed transactions (two-phase commit)

Two-phase commit is the textbook fix. In practice, most modern databases, cloud services, and message brokers don't support it, and where it is available it adds enough latency and locking to slow the whole flow down. A TransactionScope or EF Core transaction isn't a substitute either. It covers your own database, but it can't roll back an HTTP call you already made to another service.

Option 2: Fix it after the fact

A scheduled job, or a person, finds inconsistent records and reconciles them later. This treats the symptom, not the cause. Customers still get charged wrongly in the meantime, and when the cleanup job itself fails, you have two problems instead of one.

Option 3: Build our own saga framework

Some teams respond by writing a generic, reusable saga engine before solving the actual business problem:

public interface ISagaStep<TContext> where TContext : IBaseContext
{
    Task<SagaResult> ExecuteAsync(TContext context);
    Task CompensateAsync(TContext context);
}

public abstract class AbstractDistributedSagaManager<TContext>
    where TContext : SagaContext
{
    private readonly ITransactionCoordinatorStrategyFactory _factory;
    private readonly List<ISagaStep<TContext>> _steps = new();
    // ...
}

It's hard to test and harder to debug. The first network timeout leaves something stuck in PendingCompensationFinalRetryV2 forever.

Option 4: Choreographed saga

A saga breaks one big transaction into a sequence of local transactions, one per service. If a step fails, compensating transactions undo the steps that already succeeded.

With choreography, there's no central coordinator. Each service does its work, publishes an event, and the next service reacts to it. This works well for simple flows with two or three services. There's no extra component to build or run, and no single point of failure.

It gets harder as the flow grows. Nobody owns the full picture, so adding a step means tracing which service listens to which event. Services can end up depending on each other's events in a loop. Integration testing needs every service running at once, and debugging a failure means piecing together logs from all of them.

Option 5: Orchestrated saga

With orchestration, one component, the orchestrator, runs the steps in order, records progress, and runs the compensations if something fails. It suits complex flows and makes adding steps easy. The flow lives in one place, so cyclic dependencies can't sneak in, and each service only has to do its own job.

The trade-off is that you need coordination logic, and the orchestrator becomes a single point of failure if it isn't built to survive crashes. Here's what that looks like in practice.

Persist the saga state. If you don't record how far the saga got, you can't recover after a failure.

public enum SagaStatus
{
    Started,
    InventoryReserved,
    PaymentCompleted,
    OrderCompleted,
    Failed,
    Compensated
}

public class OrderSagaState
{
    [Key]
    public Guid SagaId { get; set; }
    public long OrderId { get; set; }
    public SagaStatus Status { get; set; }
    public string? LastError { get; set; }
}

Order the steps around the pivot. Every step falls into one of three groups:

  • Compensatable steps can be undone by another step with the opposite effect. Reserving inventory is one, because you can release it.
  • The pivot is the point of no return, usually charging the card. It can be the last step you can still undo or the first step you can only retry. Either way, once it succeeds, the saga only moves forward.
  • Retriable steps come after the pivot and must eventually succeed, like scheduling shipping. They have to be idempotent, so retrying them after a temporary failure is always safe.

That's why the order changes from the original code. Reserve inventory first, because it's cheap to undo, then charge.

public class OrderSagaOrchestrator(
    IInventoryService inventory,
    IPaymentService payments,
    IShippingService shipping,
    ISagaStateRepository sagaStates,
    IMessageQueue queue)
{
    public async Task ExecuteAsync(OrderRequest request)
    {
        var sagaId = Guid.NewGuid();
        await UpdateStatusAsync(sagaId, SagaStatus.Started);

        try
        {
            // Compensatable: easy to undo
            await inventory.ReserveAsync(request.Sku);
            await UpdateStatusAsync(sagaId, SagaStatus.InventoryReserved);

            // Pivot: once this succeeds, we only move forward.
            // The saga ID doubles as an idempotency key, so a retried charge never bills twice.
            await payments.ChargeAsync(request.Amount, idempotencyKey: sagaId.ToString());
            await UpdateStatusAsync(sagaId, SagaStatus.PaymentCompleted);

            // Retriable: must eventually succeed
            await shipping.ScheduleAsync(request.OrderId);
            await UpdateStatusAsync(sagaId, SagaStatus.OrderCompleted);
        }
        catch (Exception ex)
        {
            await HandleFailureAsync(sagaId, request, ex);
        }
    }
}

Handle failure based on where you stopped. If you failed before the pivot, undo what you did. If you failed after it, retry. Never refund.

private async Task HandleFailureAsync(Guid sagaId, OrderRequest request, Exception ex)
{
    var state = await sagaStates.GetAsync(sagaId);
    state.LastError = ex.Message;

    switch (state.Status)
    {
        case SagaStatus.Started:
            // Inventory reservation failed. Nothing to undo.
            state.Status = SagaStatus.Failed;
            break;

        case SagaStatus.InventoryReserved:
            // Payment failed. Release the stock.
            await inventory.ReleaseAsync(request.Sku);
            state.Status = SagaStatus.Compensated;
            break;

        case SagaStatus.PaymentCompleted:
            // Past the pivot. Don't refund, retry shipping.
            await queue.SendToRetryAsync(new ShippingTask(request.OrderId));
            break;
    }

    await sagaStates.SaveAsync(state);
}

Plan for compensations that fail. A compensation is just another call to another service, so it can fail too. Make compensations idempotent and retry them. If one keeps failing, move it to a dead-letter queue and alert someone, so a person handles it on purpose instead of finding it weeks later.

Close the isolation gap. Sagas give you atomicity, consistency, and durability, but not isolation. The most common fix is a semantic lock, which marks a resource as in progress so other code knows not to rely on it yet:

public class InventoryService(IInventoryRepository inventoryRepo) : IInventoryService
{
    public Task ReserveAsync(string sku) =>
        inventoryRepo.UpdateStatusAsync(sku, "LOCKED_BY_SAGA");

    public Task ReleaseAsync(string sku) =>
        inventoryRepo.UpdateStatusAsync(sku, "AVAILABLE");
}

A few other countermeasures help with the anomalies from the problem statement:

  • Commutative updates. Write changes so the order doesn't matter, like adding to a reserved count instead of overwriting it. Two sagas can then update the same record without one losing the other's work.
  • Pessimistic ordering. Move risky updates into the retriable steps after the pivot. Those never get rolled back, so nothing can read data that's about to disappear.
  • Reread before writing. Check the record hasn't changed since you read it, and restart the step if it has. In EF Core, a concurrency token does this for you by throwing DbUpdateConcurrencyException on a stale write.
  • Match the tool to the risk. Use sagas for everyday updates, and keep high-risk data, where even a brief inconsistency is unacceptable, inside a single database transaction.

Conclusion and Recommendation

Use an orchestrated saga (Option 5), and run it on an existing engine instead of building your own. In .NET, good options include MassTransit state machine sagas, NServiceBus sagas, Temporal's .NET SDK, and Azure Durable Functions. If you're on AWS, Step Functions works too. These engines persist saga state and pick up where they left off after a crash, which removes the orchestrator's biggest weakness.

Here's why it beats the alternatives. Two-phase commit usually isn't available, and when it is, it's too slow. Cleanup jobs leave customers exposed until the fix runs. A home-grown framework eats the time you don't have and fails in ways that are hard to debug. Choreography is a fine choice for a two- or three-service flow, but a four-service flow that moves money needs someone in charge. An orchestrator gives you one place to see exactly where an order stopped and which compensations ran. Persisting the saga state costs a few milliseconds per step, which is nothing next to a senior developer spending a Saturday reconciling bank statements against database records.

When to use a saga

  • A business process spans several services, each with its own database. Orders, bookings, and payments are the classic examples, where every step has to land or none of them should.
  • You need to keep data consistent without tightly coupling the services. Each service keeps its own data and only talks to the others through commands and events.
  • Each step can either be undone or safely retried. If you can release the stock, cancel the booking, or retry the shipping request until it works, a saga fits naturally.
  • The process is long-running or depends on outside systems. Payment gateways, shipping providers, and approvals by people can take seconds or days, and you can't hold a database lock that long.
  • The business can live with eventual consistency. An order sitting in "Pending" for a few seconds while the steps finish is acceptable to most customers.

When not to use a saga

  • All the data lives in one database. Use a normal local transaction. It's simpler and gives you full ACID guarantees for free.
  • Nobody can ever see intermediate state. Sagas don't give you isolation. If the business can't tolerate a half-finished transaction being visible even briefly, keep that data in one database or rethink the service boundaries.
  • More than one step is irreversible. A saga only has room for one pivot. If two steps both can't be undone, redesign the flow first. For example, authorize the card early and capture the payment only at the end.
  • The services are tightly coupled or depend on each other in a loop. If every request needs several services to coordinate, or service A waits on B while B waits on A, the boundaries are drawn in the wrong place. Merging those services is usually simpler than orchestrating them.
  • The flow is just two steps and the second can be retried. A transactional outbox with idempotent retries is enough. A full saga is overkill.

Putting it into practice

  1. Find your pivot. It's usually the step that moves money.
  2. Write compensations first. If a step can't be undone, it belongs after the pivot and must be retriable.
  3. Use an existing engine. These tools already handle the hard cases, like the orchestrator itself crashing mid-saga.
  4. Make every step idempotent. Retries mean a service will receive the same command more than once, and it has to apply it exactly once. If your RefundAsync() endpoint isn't idempotent, you'll find out the hard way. I walk through how to do this in C#, with dedup tables, database constraints, and idempotency keys for payment providers, in Idempotency in Distributed Systems: Building Safe Retries in C#.
  5. Monitor every saga. Track each saga's current state and alert on any that sit in one state too long. A stuck saga is much cheaper to fix on day one than on day thirty.

Sagas won't stop failures from happening. What they give you is a plan for when they do.