When AI Agents Begin Modifying Business State: The Need for a Runtime Framework Beyond the Model

Table of Contents
Imagine a scenario: in the early hours of the morning, a company’s checkout service experiences a large number of request timeouts. The next day, an engineer opens the internal AI assistant and types a single sentence:
Investigate last night’s checkout service failure. Find relevant incident records, design documents, and recent code changes. If a regression is confirmed, create a Bug and suggest a responsible person.
To a human, this is one sentence. To an AI system, however, it contains two different types of operations.
Searching incident records, reading design documents, and checking code commits primarily aim to gather evidence. Creating a defect ticket, on the other hand, changes the real state within enterprise systems. Even if the Agent merely suggests “who should be responsible,” the system still needs to decide where this suggestion should be written, who can see it, and who has the authority to convert it into a formal assignment.
The previous two articles discussed “Is Enterprise RAG Worth Doing: From Scenario Selection to Ongoing Governance” and “Does the Enterprise Really Need AI Agents: When Should Models Decide the Next Step” respectively. RAG organizes enterprise knowledge for the model, while Agents can decide the next step based on new information. This article will continue to explore a more specific engineering question:
When the next step chosen by the model changes the real business state, which system is responsible for executing this choice safely and reliably?
I call this the Agent Harness. The model is responsible for understanding the problem, planning tasks, and proposing actions; while identity authentication, permission validation, task state preservation, actual execution, fault recovery, and result verification are handled by systems outside the model.
This does not mean that every Agent needs its own complex platform. Read-only Q&A, short tasks, and fixed-path automation scenarios can often use simpler architectures. When tasks are longer-duration, require calling multiple systems, and cannot simply restart from scratch after failure, the system needs specialized management for execution and recovery. Even for short tasks, as long as they involve modifying business data, permission validation and mechanisms to prevent duplicate writes are still necessary.
How Long Horizon Organizes Context, and Why Task State Preservation is Also Needed #

Rovo is Atlassian’s AI product for enterprise teams, offering enterprise search, conversational assistants, and Agent capabilities. Atlassian also develops Jira and Confluence: the former is often used for tracking tasks and defects, while the latter is used for team documentation and knowledge collaboration. Rovo can combine information from these products, as well as other applications connected via connectors, to help users find information, understand work context, and execute tasks. The cross-incident record, design document, and code change investigation mentioned at the beginning are precisely the scenarios such assistants need to handle.
This article focuses on Rovo Chat, the conversational assistant, and the runtime architecture that supports its multi-step tasks. In June 2026, Atlassian publicly introduced the Long Horizon architecture used for Rovo Chat. [1]
Early Rovo Chat adopted a layered Multi-Agent architecture. A coordinating Agent would assign tasks to specialized Agents responsible for Jira, Confluence, Slack, etc., and then aggregate the results. In this call path, the main Agent received summaries returned by the specialized Agents and could not directly read the raw results, error messages, or intermediate processes returned by the tools. This design was suitable for tasks with fewer steps and shorter contexts, but the longer the task, the more pronounced the impact of detail loss in summaries.
Atlassian subsequently centralized the main reasoning process into a single loop: a single model would select tools, read results, and decide the next step within the same context. Tool calls no longer had to be relayed through specialized Agents for each product; the main model could directly call a unified tool interface and retrieve specific tool descriptions as needed. [1]
This does not mean putting all work into an ever-present context. Long Horizon still compresses older content and saves external tool results for the model to read on demand; for independent subtasks, it can also create child instances with independent contexts, and then have the parent instance aggregate the results. The main model can now directly read tool results while still being able to split tasks via child instances. [1]
When a task lasts only a few seconds and calls one or two read-only tools, how context is preserved is usually not a major engineering problem. But if a task lasts tens of minutes, involving multiple queries, code executions, and external writes, the system must answer:
- Where is the current task in progress?
- Which evidence comes from the source system, and which content is merely model inference?
- Which step has been executed, and which step is still awaiting confirmation?
- From where should the task resume if interrupted?
- If the user disconnects, can the task still continue, and where are the running state and results saved?
These questions go beyond how to organize the model’s context and also involve persistent storage of task state and execution recovery. The June 2026 Long Horizon article listed persistent task capabilities like continuing after disconnection and progress tracking as future directions, meaning that a unified reasoning loop should not be directly equated with these capabilities being already implemented. [1]
Why Separate the Control Plane from the Execution Environment #

In August 2026, Atlassian again publicly introduced the Agent Harness architecture supporting Rovo. It separates the Control Plane from the Data Plane: the former is responsible for conversation and task execution, while the latter provides an isolated Sandbox for code execution and temporary computation. This architecture also supports asynchronous operation and child Agents that maintain their own session histories. [2]
The following diagram is a conceptual abstraction based on Atlassian’s public information, not an official component diagram:
User → Control Plane (saves conversation history and execution state, organizes task execution and recovery)
Path 1: Direct tool call
Control Plane → Internal service verifies permissions, completes authentication → Business API / External MCP service
Path 2: Execute code via sandbox and call tools within the code
Control Plane → Isolated Sandbox (saves temporary computation state and files)
│
▼
Callback Bridge
sends tool requests back to Control Plane
│
▼
Internal service verifies permissions, completes authentication
│
▼
Business API / External MCP service
This separation first addresses fault isolation. If an Agent’s operation and state preservation rely entirely on the sandbox, a sandbox crash could disrupt the entire task. With the Control Plane independently retaining conversation records and execution state, even if the sandbox crashes, the system can recreate the execution environment without losing the entire conversation. [2]
After separation, the two parts can also scale independently. Simple information queries may not require launching a sandbox; large-scale data analysis, code execution, or file generation can use an isolated computing environment. Atlassian’s implementation chooses to either call tools directly or use a sandbox that can preserve computation state, depending on the task. This is Rovo’s specific engineering choice, not a component list that all enterprise Agents must replicate. [2]
The isolated environment should also not directly hold long-term credentials to access all enterprise systems. Rovo’s Callback Bridge mechanism sends tool requests from the sandbox back to the Control Plane, where an internal service verifies user permissions, completes authentication, and then calls external APIs. The sandbox therefore does not need direct access to external networks for these tool calls. [2]
There is no unified industry definition for Agent Framework and Agent Harness. Anthropic’s definition of Harness includes processing input, organizing tool calls, and returning results, not necessarily predicated on long tasks or isolated execution. [3]
This article focuses on Agent Harness in enterprise production environments, emphasizing runtime responsibilities such as task state, execution isolation, permission validation, recovery, and verification. This is the scope of this article’s discussion, not a unified industry distinction between Framework and Harness. Some Frameworks also provide persistence, permissions, or evaluation capabilities; what truly matters is whether these system responsibilities have clear ownership.
Why Not Just Retry After a Write Operation Times Out #

Let’s return to the initial failure investigation scenario. Assume the engineer has provided a test that reliably reproduces the timeout, and the Agent verifies in an isolated environment that the same test passed before the change, failed after the change, and passed again after reverting the change. These results must meet the regression confirmation conditions agreed upon by the team for this task before proceeding to the step of creating a ticket. If there is only a suspicious commit or a temporal coincidence, the Agent should report candidate causes and evidence gaps, and should not unilaterally downgrade “confirmed regression” to “suspected regression.”
At this point, the task has completed the first four steps:
1. Search incident records ✓
2. Read relevant design documents ✓
3. Check recent code commits ✓
4. Confirm code regression via tests ✓
5. Create defect ticket ✗ timeout
The fifth step returns a timeout error. The most direct approach is to re-call create_issue(), but a timeout only indicates that the Agent failed to receive a result in time, not that the server failed to complete the write.
Agent calls create_issue()
↓
Server successfully creates ticket
↓
Response lost in network
↓
Agent receives timeout error
↓
Retries and creates a second ticket
Therefore, recovery for long-running tasks requires several types of mechanisms to work together.
Checkpoints are used to save execution progress and state. Rovo saves snapshots of execution state at key nodes, providing a basis for recovery and retry after failure, reducing redundant execution of completed steps. [2] However, checkpoints alone cannot prove whether a particular external write was successful. Restoring the Agent’s execution state also does not equate to reverting writes that have already occurred in external systems.
Idempotency restricts duplicate writes, and state reconciliation helps confirm execution results. Idempotency means that executing the same request multiple times will not generate additional business changes. The prerequisite is that the target interface supports reliable idempotent processing: the execution layer should save a stable operation identifier for the same creation operation, reusing it during retries or task recovery, and not changing the identifier due to restarting the task. The server also needs to ensure consistency between deduplication records and business writes, and clearly define the effective period for deduplication. [4]
If the target interface does not support idempotent processing, a combination of business unique constraints and state reconciliation is required. Merely “if not found, then create” can still lead to duplicates due to concurrent requests or delays in query result updates. When results cannot be confirmed, the execution layer should pause automatic retries, retain a pending reconciliation state, and hand it over to designated personnel, rather than treating unknown results as failures.
Compensation flows and manual takeover handle impacts that cannot be directly rolled back. Some external operations have taken effect but cannot be directly rolled back via database transactions, requiring compensatory measures based on business rules. Compensation may not restore the original state; for irreversible impacts, manual intervention or mitigation is also needed. When involving notifications, payments, deletions, or cross-organizational operations, teams should pre-define which operations are compensable, when to stop automatic execution, and who takes over.
These mechanisms are not interchangeable. Checkpoints address “where to continue from,” idempotency addresses “whether repeated requests have repeated effects,” business reconciliation addresses “what is the real system state now,” and compensation flows handle impacts that have occurred but cannot be directly undone.
Anthropic’s engineering experiments on long-running Agents also show that relying solely on context compression makes it difficult for a model to consistently advance a task across multiple context windows. Progress files, Git history, and continuously saved code can help subsequent sessions understand what was previously completed. [5] These experiments primarily target software engineering Agents, only demonstrating one possible solution, and cannot directly prove that all enterprise Agents should adopt the same implementation.
These practices also remind us:
When models start deciding the next step for long-running tasks, engineers still have to deal with state, timeouts, retries, idempotency, and partial failures, and also consider that the model might make different choices each time.
Model Proposes Action, System Does Not Necessarily Authorize Execution #

A long-running task being recoverable does not mean every step should be executed.
Continuing with the regression-verified case, suppose the Agent proposes the following action. Here, create_issue() is a business wrapper assumed for this article, not a native Jira API; suggested owner and suggested severity are saved in separate structured fields for team review before deciding whether to adopt:
create_issue(
title="Checkout timeout regression",
source_incident=832,
suggested_owner="Bob",
suggested_severity="P1"
)
Here, suggested_owner and suggested_severity are merely suggested fields; they do not mean the system has assigned the ticket to Bob or that the formal severity level has become P1. If the Agent were to set the formal assignee or severity, the system would still need to check permissions and approval requirements for these operations.
An action proposed by the model can go through the following execution chain before taking effect:
Model proposes action
↓
Validate parameter structure, type, and values
↓
Verify identity and operation permissions
↓
Reconcile current business state and execution conditions
↓
Route based on check results:
Authorized and conditions met → Continue execution
Authorization or confirmation still needed → Pause; re-verify after approval, then continue execution
Explicitly forbidden → Reject, terminate execution
↓
Perform idempotent processing for allowed requests and execute
↓
Record audit information, reconcile actual results
Manual approval is not a fixed step for all tool calls. Reading ordinary documents, deleting production data, and granting administrator privileges should not have the same threshold. Authorization and policy checking mechanisms in the runtime system need to decide whether to allow execution based on user identity, task authorization, operation parameters, current business state, and impact scope, following explicit rules.
Read-only retrieval also requires permission validation. Atlassian’s Rovo Code Context reuses user permissions from the source code management system and re-validates them by calling the code repository’s permission validation interface during queries, rather than relying solely on old Access Control Lists (ACLs) saved during indexing. [6] Data that users do not have permission to access should, in principle, not enter the model’s context, and certainly not be seen by the model first, relying on system prompts to decide whether to reveal it to the user.
Therefore, a more accurate boundary is not “RAG manages knowledge, Harness manages actions,” but rather:
Data governance and access control determine which evidence can enter the context; action policies and authorization systems determine whether changes proposed by the model can take effect.
Prompts can guide the model to choose appropriate behavior, but they cannot replace identity systems, business state machines, and tool interfaces as the ultimate security boundary.
Accepting an Agent: Don’t Just Look at What It Said Last #

Suppose the Agent’s final response is:
The checkout timeout regression has been confirmed based on the agreed-upon tests, and a defect ticket has been created, with test results attached. The suggested fields recorded Bob and P1 for team evaluation; the formal assignee and severity have not yet been set.
This response claims to have fulfilled the user request, but the text itself still cannot prove that the task is truly completed. Acceptance involves both checking the accuracy of the response and verifying the actual state in the business system: first, check the test records to confirm that the conditions for creating the ticket were indeed met, and then inspect the write result via business API or database. Assuming this example allows the formal assignee and severity to be left blank, and there are no automatic assignment rules, then a successful completion check should include:
Number of tickets for this creation operation = 1
source_incident = 832
assignee = null
severity = unset
suggested_owner = Bob
suggested_severity = P1
“One ticket” here refers only to this specific creation operation, not a requirement that there can only be one ticket for the entire incident or business system. If the actual system automatically assigns an assignee or sets a default level, the acceptance should be based on authorized business rules and audit records, and not simply require the fields to be empty.
Anthropic distinguishes between execution trajectory (also called transcript or trace) and final outcome in Agent Evaluations (Agent Evals). The execution trajectory records model output, tool calls, and intermediate results; the final outcome is the actual state in the environment after the task is completed. [3]
Therefore, Agents in production environments require both conventional software testing and evaluation of the complete task:
| Check Object | Question to Answer | More Suitable Verification Method |
|---|---|---|
| Tool Implementation | Does create_issue() correctly validate parameters, check operation permissions, and handle idempotency keys? | Unit tests, integration tests |
| Retrieval and Permissions | Did the Agent obtain necessary evidence, and did it access unauthorized data? | Access logs, deterministic permission checks |
| Execution Trajectory | Did the Agent attempt unauthorized actions, get into infinite loops, or perform invalid retries? | Trajectory rule checks, manual or model review |
| Execution Conditions | Were the user-agreed regression confirmation conditions truly met? | Test record verification, engineer review if necessary |
| Final Outcome | Did the real system produce only one ticket with correct fields? | Database or business API state checks |
When evaluating the execution trajectory, a rigid requirement for a unique tool order should be avoided. The same investigation task may have multiple reasonable paths, but permission validation, necessary approvals, and business dependencies like “verify regression first, then create ticket” should still be hard constraints. Given these constraints, open-ended tasks are better suited to prioritizing outcome checks, then using the trajectory to explain reasons for failure.
In addition to business condition verification, this case also requires two types of fault and permission tests.
Permission Leakage Test aims to ensure that data inaccessible to the current engineer, and its derived content, does not enter the model’s context. This can use test data with known permissions and identifiable content to check retrieval results, tool responses, and what is actually fed into the model, covering scenarios like permission revocation and cache reuse. The check object should include document chunks and summaries, not just compare complete document IDs. These tests can uncover non-compliant paths, but cannot conclusively prove that the system will not leak in all cases.
Duplicate Write Test simulates scenarios where the server has already created a ticket and successfully responded, but the response was lost. It verifies that the execution layer reuses the original operation identifier within the idempotency agreement’s validity period, ultimately returning the original ticket without creating an additional one. It should also cover concurrent retries and task restarts; if the result still cannot be confirmed, the system should pause and report a pending reconciliation state, rather than claiming success.
Neither of these checks can rely solely on model scoring (LLM-as-a-Judge). Permission leakage tests require checking the content the model actually accessed and access records; duplicate write tests require verifying the write results in the business system. For the latter, a program can judge whether the results meet requirements based on preset conditions, which is a deterministic outcome grader.
When Is Such a Runtime Framework Needed #
Rovo demonstrates a vendor architecture designed for a large number of users, long-running tasks, multiple tools, and isolated computing. Enterprises should not mechanically replicate all components just because they see separation of control plane and execution environment, sandboxes, or multi-agent collaboration.
What is more valuable to learn is the approach to determining system responsibilities. An independent runtime framework becomes particularly important in the following situations:
- Tasks cannot be reliably completed within a single synchronous request and need to continue after disconnection or execution interruption;
- The Agent writes to external systems, and repeated execution would produce real side effects;
- The number of tools, credentials, and data permissions dynamically change with the user or task;
- Tasks require isolated code execution, temporary files, or substantial computation;
- The enterprise must reconstruct who authorized what, what the model proposed, what the system executed, and how the final state changed.
If an Agent only answers a short question, does not modify business state, and the result can be immediately judged by the user, simple model calls, RAG, or fixed workflows are often sufficient. Even if a task requires a Harness, enterprises do not necessarily have to develop their own; hosted Agent platforms, workflow systems, and existing business infrastructure can all bear some of these responsibilities.
When choosing a solution, one should first confirm what responsibilities existing systems already handle and what capabilities are still missing, then decide what needs to be supplemented, rather than first purchasing or building a platform called Harness.
Conclusion #

The initial failure investigation task did not become simpler with the introduction of an Agent. The user only needed to say one sentence, but the system still had to complete searching, judging, writing, and verifying.
Long Horizon allows the model to continuously read feedback and choose the next step around a task; Agent Harness provides state preservation, execution isolation, permission validation, and fault recovery for this process. Atlassian’s solution is just one specific implementation, but it reveals a universal problem: once a model begins to change real business states, enterprises cannot rely solely on prompt engineering to constrain its behavior.
Enterprises deploying Agents still need to follow fundamental software engineering requirements, with the addition of a model that chooses its next step based on the environment and whose output is not entirely deterministic.
The model can propose actions and explain why. Identity authentication, permission validation, execution, recovery, and acceptance should still be handled by verifiable and auditable system mechanisms.
References #
[1] Sean Culatana et al. Long Horizon: How Atlassian Built a Reasoning Engine for Complex AI Tasks. Atlassian, 2026-06-17. https://www.atlassian.com/blog/how-we-build/rovo-long-horizon-reasoning-engine
[2] Atlassian. Opening the Door to Agent Autonomy: The Architecture Behind Rovo’s Agent Harness. 2026-08-26. https://www.atlassian.com/blog/rovo/agent-autonomy
[3] Anthropic. Demystifying Evals for AI Agents. 2026-01-09. https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents
[4] Malcolm Featonby. Making Retries Safe with Idempotent APIs. Amazon Builders’ Library. Accessed 2026-09-30. https://aws.amazon.com/builders-library/making-retries-safe-with-idempotent-APIs/
[5] Anthropic. Effective Harnesses for Long-Running Agents. 2025-11-26. https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents
[6] Atlassian. Set Up Code Context for Your Site. Atlassian Support. Accessed 2026-09-30. https://support.atlassian.com/organization-administration/docs/set-up-code-context-for-your-site/