How SagaLLM proposes to preserve state, validate plans and recover from failures—and what its planning experiments establish about dependable AI agents.
An agent has already arranged part of a trip when something changes. A flight is delayed, a connection becomes impossible or a booking service stops responding. The agent can generate a new itinerary. The harder question is whether that itinerary respects what has already happened.
A revised plan cannot erase a confirmed reservation. It cannot move someone back to the airport after they have left. And a request that timed out may still have succeeded at the other end.
SagaLLM proposes an architecture for handling that gap between generating a plan and managing its consequences. The paper combines persistent records, validation outside the task-performing agents and explicit recovery actions. It also makes a more ambitious proposal: use language models to generate the coordination machinery that developers would otherwise build.
For technical leaders, these are two adoption questions. There is value in making an agent's work observable and recoverable. Trusting an agent to generate the rules governing that recovery needs evidence of its own.
The authors identify weaknesses in planning systems that rely on language-model conversations to carry the work forward. Important context can be lost, an agent's own check may miss its error, and several agents can make individually plausible decisions that conflict when combined.
Their design proposition is to make context, validation and recovery explicit responsibilities. A task agent performs an operation. A separate validator checks its output and constraints. A coordinator tracks dependencies and decides what affected work needs to be reconsidered when a step fails.
The proposal also goes beyond adding a saved conversation. Given a problem and its constraints, SagaLLM is intended to construct a workflow, define agent roles, specify records of their actions and generate compensating behavior. The distinction matters: useful coordination principles do not automatically establish that generated coordination code is dependable.
To understand what the authors are proposing, it helps to follow the travel architecture before looking at the experiments.
The paper separates travel planning into two phases. First, the user reviews and approves an itinerary with dates, preferences, budget and other constraints. The proposed automated phase then organizes the agents and transactions needed to carry it out.
In the authors' example, bookings connect a flight to Berlin, accommodation there, a train to Cologne, another hotel and the return flight. These are related commitments. A change to the arrival date can affect the hotel stay and train connection, so the system needs to represent those relationships.
The authors’ original diagram makes those commitments visible. The user first approves the itinerary; the handoff then leads into a chain of bookings. The red arrows show recovery paths when later work cannot proceed. A failed train arrangement can have consequences for the accommodation on either side of it.
The workflow-generation procedure has three stages. First, it extracts roles and maps the dependencies between them. Second, it specifies a normal agent, a compensation agent and a logging structure for each step and connection. Third, it checks and refines the resulting workflow for structural correctness, constraint compliance and compensation behavior.
This is the paper’s more ambitious contribution: generating the machinery that coordinates the agents, not simply generating the itinerary. The algorithm describes how that machinery should be constructed. It does not establish that generated code will handle every real booking failure correctly.
For a builder, the key idea is that recovery should be considered when defining each action. A booking step needs more than instructions for making a reservation. It needs a way to recognize success, preserve the resulting commitment and respond when later work makes that commitment unsuitable.
SagaLLM’s state model distinguishes three kinds of state. In ordinary language, these are the facts the application is working with, the record of its operations and the relationships between them.
In the travel design, application facts include dates, booking details and user constraints. The operation record includes transaction identifiers, inputs, outputs, times and recovery information. Dependencies describe which bookings or checks must precede other work.
Those records answer different questions. “The traveler prefers this hotel” is not confirmation that the room was booked. “The booking request was sent” is not proof that the provider accepted it. “The hotel stay depends on the flight dates” tells the coordinator what to reconsider when those dates change.
Keeping those meanings separate gives replanning something firmer than the agent's latest account of events. It also gives a reviewer a way to challenge a proposal: which fact changed, which completed operation still stands, and which dependency makes another action necessary?
The proposed execution sequence checks inputs and dependencies, performs the operation, validates the output and commits the accepted result to the system's state. The coordinator also records the corresponding recovery procedure. The validator's responsibilities include checking structure, constraints, consistency and relationships between agents.
The validation examples in Table 3 make this more concrete. A flight result can contain the required fields and still fail the trip. Hotel dates must cover the stay without gaps. A hotel-to-train transfer must leave the required 45 minutes. Combined bookings must fit the $5,000 budget. And a hotel should not be finalized as if its prerequisite flight were already confirmed when it is not.
These are different checks. Required fields can be verified directly. Date coverage, transfer time and budget require comparing the output with other records. A second model saying “looks good” is not a substitute for those comparisons.
When validation fails, the recovery design identifies affected transactions, applies compensation in reverse dependency order, checks the resulting state and replans the affected work using the preserved context.
Compensation means taking another action to deal with an earlier commitment. In this design it might involve canceling a reservation, requesting a refund and reconsidering dependent bookings. It does not mean that a real-world event never occurred. A cancellation may have conditions or consequences that a new plan must respect.
This is also why separate validation deserves scrutiny. Giving a second agent the title of validator does not establish that its judgment is correct. A production evaluation needs to show what the check compares against and whether it catches the relevant failure before more actions follow.
The paper’s evaluation examines four selected tasks from the REALM planning benchmark: two sequential problems and two involving disruptions. It reports examples involving Claude 3.7, DeepSeek R1, GPT-4o and GPT-o1 alongside the proposed SagaLLM approach. The experiments took place in March 2025; this article uses the July 2025 manuscript.
The concrete cases are useful because they show what can go wrong even when the revised plan reads sensibly.
In the dinner scenario, the family must coordinate arrivals, transportation and preparation for a 6 p.m. meal. Cooking has timing requirements, and someone must remain at home while the oven is in use.
A flight that was due at 1 p.m. is delayed until 4 p.m. In Claude 3.7’s revised schedule, the airport pickup is reassigned to Sarah. That resolves one transportation problem by creating a cooking problem: she is also responsible for the turkey, which is already in the oven.
Read the food-preparation branch on the right of the original diagram. Sarah leaves while the turkey is cooking. Side dishes start at 4:30 p.m. and finish at 6 p.m.—90 minutes for a task that requires two hours. Dinner moves to 6:30 p.m., missing the 6 p.m. requirement. The new plan still connects every box, but the commitments inside the boxes no longer fit together.
The authors' proposed remedy is to maintain explicit state and constraints, with separate validation for travel coordination and food preparation. A new pickup arrangement should be checked against the cooking responsibility before it is accepted. A plan can only be considered feasible if both parts work together.
The lesson transfers beyond scheduling dinner. In a business workflow, solving a delayed shipment could create a staffing conflict or violate a delivery commitment. The useful test is whether the revised plan preserves all relevant constraints, including ones the latest user message did not repeat.
The wedding scenario involves collecting people and completing errands before a fixed photo session. A traffic alert arrives after some travel has already happened.
In the DeepSeek R1 schedule, the first row places Pat at the wedding venue when the alert arrives at 1 p.m. The next sends him toward the airport at 1:05. But the established plan had already put him at the airport collecting passengers before the alert. The revision has changed the past, then calculated a new journey from that invented starting point.
A saved itinerary alone would not resolve this. At the moment of the alert, where is Pat, who is in the car and which errands are unfinished? Those facts must constrain the next plan. An agent must not turn “this was the original intention” into “this is still waiting to happen.”
The paper also reports a feasible GPT-o1 response. Its original schedule below shows the distinction clearly. Here W is the wedding venue, B the airport, T the tailor and G the gift shop.
Pat leaves the airport with Alex and Jamie at 12:50. By the 1 p.m. traffic alert, ten minutes of a normally 40-minute trip have already passed. The remaining 30 minutes are affected by the threefold delay: they now take 90 minutes, giving a 2:30 arrival. The completed ten minutes stay completed.
The rest of the schedule matters too. Chris becomes available at 1:30, collects the clothes before the tailor’s 2 p.m. cutoff, buys the gift and returns at 2:40. Both groups are back before the 3 p.m. photos. The solution handles the affected journey and the other deadline together.
SagaLLM’s proposed coordination aims to make that kind of state preservation and constraint checking an explicit system responsibility. But this example is GPT-o1’s response, not a measured SagaLLM improvement. It supports a narrower conclusion: some generated plans lose track of progress, and the team needs a way to detect that failure. The paper does not establish that every standalone model fails or every simpler workflow needs replacement.
The scenarios make the failure modes understandable. They are a narrower basis for confidence than repeated measurements of a deployed system interacting with booking, payment or delivery services.
The manuscript describes architectural remedies and selected schedules. That supports investigation of its approach, but it does not by itself tell a team how often generated recovery logic will work, how much review it will need or what happens across a broad range of service failures. The paper also explicitly relaxes strict database transaction guarantees. Its recovery proposal should not be read as a promise that arbitrary external actions can be undone completely.
Our reading is that the case for preserving state and checking constraints is more developed than the evidence for trusting automatically generated coordination in general. Those capabilities should be evaluated separately.
An agent's authority to propose a new plan should not include authority to rewrite the record of completed actions. The dinner example asks whether a change preserves existing obligations. The wedding example asks whether the starting point is still true. A useful coordination system needs to answer both.
For implementation, distinguish an intended action, an attempted action and a confirmed outcome. If a service times out, the outcome may be unknown. Repeating the attempt without resolving that uncertainty can create a second effect rather than repair the first.
The practical goal is to make the next action defensible. The system should be able to show which facts it used, which constraints it checked, what it believes is complete and what remains unresolved. A fluent explanation is useful only when it corresponds to that record.
For a delivery product, we would start with a workflow that has already collected a parcel when the customer changes the destination. Use the current system as the baseline and preserve the confirmed collection event.
The agent can propose a reroute. Before execution, check whether the new destination is allowed, whether the delivery promise still holds and whether additional charges require a decision. The revised plan must begin with the parcel's current status. It cannot quietly reset the job to “awaiting collection.”
Now introduce a more demanding case: the carrier accepts the rerouting request, but its response is lost. Require the workflow to establish the carrier's current status before repeating the action. If it cannot resolve the outcome, keep the job visibly unresolved and route it to an operator. Do not manufacture a confirmation merely to finish the conversation.
Finally, test failed recovery. If the destination cannot be changed after a cutoff, the system needs an allowed next step, such as continuing to the original destination or requesting a human decision. The ability to generate a new plan does not expand the carrier's capabilities or the customer's permissions.
These cases create two useful comparisons. First, test whether explicit state and validation improve the existing workflow. Then compare developer-written coordination with generated coordination under the same disruptions and acceptance rules. Include review and correction time in the second comparison so generated code does not appear cheaper simply because its checking work is excluded.
LangGraph provides checkpointers for state within a thread and stores for information shared across threads. Consider it when the agent needs continuity across steps, pauses and resumptions. Use persistent storage when the record must survive a process restart; the in-memory examples do not provide that durability.
For the delivery test, keep the accepted destination, confirmed collection and pending change distinct. Resume from the saved state and verify that the agent does not regenerate completed work as if it were new. The application still needs to reconcile that state with the carrier's actual record.
Saving an agent's state is one part of the design. It does not establish that a plan is feasible or reverse an external action.
Temporal’s Saga guidance describes compensating actions for multi-step workflows and requires compensation that is safe to repeat. It recommends registering recovery before an action that could partially succeed, so a failure before confirmation does not leave that action without a recovery path.
For delivery, define the normal request, how to determine its outcome and what recovery is allowed. Use request identifiers honored by the receiving service to prevent retries from creating duplicate effects. Your team must supply the actual recovery policy and actions; a workflow engine cannot make a carrier reverse an irreversible operation.
Choose this when durable execution and recovery across services are the missing capabilities. If your current workflow engine already provides them, use it as the baseline. LangGraph and Temporal can serve complementary roles, but the first evaluation need not introduce both.
Run the candidate through ordinary completion, a changed constraint, a lost response and a failed compensation. Check that completed actions remain intact, proposed actions remain within policy and uncertain outcomes remain visible until resolved.
Measure duplicate effects, violated constraints, unresolved jobs, recovery time and operator effort alongside completion. A system that reports more completed jobs by guessing about unknown outcomes has made the record less trustworthy.
For generated coordination, also review whether the logging and recovery rules cover each external action. Compare the effort of generating, checking and maintaining those rules with the developer-built alternative. Keep the simpler approach if the extra autonomy does not improve the outcomes that matter.
SagaLLM is most useful as a proposal to make responsibility explicit when an agent's plan changes. It connects memory to confirmed actions, validation to constraints and recovery to the dependencies between steps.
The research does not require a team to adopt all of its automation ambitions at once. Let an agent propose changes while established application logic controls execution and recovery. Test generated coordination separately before allowing it to take on that responsibility.
We would expand autonomy when the system reliably preserves what happened, detects invalid next steps and handles uncertainty with less total effort than the current workflow. Until then, producing a convincing revised plan is evidence of planning ability, not sufficient evidence of dependable execution.
SagaLLM: Context Management, Validation, and Transaction Guarantees for Multi-Agent LLM Planning — Edward Y. Chang and Longling Geng. Version 3, July 9, 2025. Read the full original research at the link.
Test whether an agent can carry out a plan, respect dependencies and reach the actual goal.
Explore coordination, ownership and recovery when several agents contribute to one workflow.
Explore how predicting action consequences differs from producing a plausible plan.