Skip to main content

AI & assistant-friendly summary

This section provides structured content for AI assistants and search engines. You can cite or summarize it when referencing this page.

Summary

OpenAI's 11 February 2026 Codex experiment merged about 1,500 pull requests. AgentCore Harness, GA 17 June 2026, still defaults to 75 iterations and a 3,600-second timeout — the harness sets those bounds, not the model.

Key Facts

  • •OpenAI's 11 February 2026 Codex experiment merged about 1,500 pull requests
  • •AgentCore Harness, GA 17 June 2026, still defaults to 75 iterations and a 3,600-second timeout — the harness sets those bounds, not the model
  • •Checked 11 October 2026
  • •The repository, the linters, the app boot, and the review loop did the rest
  • •On AWS, AgentCore Harness reached general availability on 17 June 2026

Entity Definitions

Amazon Bedrock
Amazon Bedrock is an AWS service discussed in this article.
Bedrock
Bedrock is an AWS service discussed in this article.
Lambda
Lambda is an AWS service discussed in this article.
S3
S3 is an AWS service discussed in this article.
CloudWatch
CloudWatch is an AWS service discussed in this article.
IAM
IAM is an AWS service discussed in this article.
EKS
EKS is an AWS service discussed in this article.
fine-tuning
fine-tuning is a cloud computing concept discussed in this article.

Harness Engineering for Production AI Agents (October 2026)

AI AgentsPalaniappan P22 min read

Quick summary: OpenAI's 11 February 2026 Codex experiment merged about 1,500 pull requests. AgentCore Harness, GA 17 June 2026, still defaults to 75 iterations and a 3,600-second timeout — the harness sets those bounds, not the model.

Key Takeaways

  • OpenAI's 11 February 2026 Codex experiment merged about 1,500 pull requests
  • AgentCore Harness, GA 17 June 2026, still defaults to 75 iterations and a 3,600-second timeout — the harness sets those bounds, not the model
  • Checked 11 October 2026
  • The repository, the linters, the app boot, and the review loop did the rest
  • On AWS, AgentCore Harness reached general availability on 17 June 2026
Reference architecture for a production AI agent harness: task contract, context builder, model loop, tool boundary, memory, policy, independent verification, telemetry, and recovery
Table of Contents

Checked 11 October 2026. An AI agent in production is a model inside a system that decides what it can see, what it can call, how it recovers, how the work is checked, and how you score the result. A stronger model does not install that system.

OpenAI’s harness engineering note (Ryan Lopopolo, 11 February 2026) describes one internal product built with Codex: on the order of a million lines, about 1,500 merged pull requests, three engineers later seven, about 3.5 pull requests per engineer per day, and an estimate of roughly one tenth the time of writing it by hand. Humans steered. The repository, the linters, the app boot, and the review loop did the rest. Those figures are OpenAI’s account of that experiment, not an industry benchmark and not a FactualMinds result.

On AWS, AgentCore Harness reached general availability on 17 June 2026. The get-started guide still says that if you omit a model, the harness uses Anthropic Claude Sonnet 4.6 on Amazon Bedrock, and that a single invocation stops at 75 iterations or 3,600 seconds unless you set a tighter ceiling (operations). The Strands harness quickstart, read the same day, names a different default: Claude Opus 5 on Bedrock, plus shell, file, and web tools. Same word. Different control plane.

Cited published figures (not a FactualMinds run) — OpenAI, 11 February 2026: about 1,500 merged PRs, about 3.5 PRs per engineer per day, AGENTS.md kept to roughly 100 lines, single runs of up to six hours. Source: Harness engineering.

First-party signals we reuse (not a new agent engagement) — Gateway server-side tools cut median tool round-trip ~180 ms to ~95 ms on a B2B CRM assistant (12 tools, ~8k turns/day) — Gateway post. Support-style AgentCore at 50K sessions/mo is ~$791/mo platform plus model — decision guide. There are still zero published FactualMinds case studies of a production AI agent.

Reproduce this — Clone the folder under examples/architecture-blog-2026/harness-engineering/. python3 -m py_compile harness_policy.py agentcore_invoke_sketch.py strands_harness_sketch.py then python3 harness_policy.py. Expected lines include approval:blocked, retry:reconcile_before_retry, and verify_ok:pass.

Opinionated take: write the task contract, the tool allowlist, and one deterministic check in application code before you pick a host. Use AgentCore Harness when you want AWS to run a single-agent loop and you will still author IAM, Cedar, and evals. Use Strands harness for a general-purpose coding or research agent you can run locally, and do not point its default shell at orders or payments. Trade-off: you give up the one-line local loop when you move a commerce workflow onto Gateway. You keep a boundary the model cannot talk its way through.

Four meanings of “harness”

Teams collide these. Separate them before you compare prices.

Harness engineering is the discipline. You design the task contract, the context the model sees, the tools it may call, the memory it may keep, the checks that accept the work, the permissions around side effects, and the record you use later. You can build that with application code, an agent framework, a managed service, or a mix.

Amazon Bedrock AgentCore Harness is a managed loop on AgentCore Runtime. You declare model, instructions, tools, memory, and limits. AWS runs the loop inside a Firecracker microVM per session. CloudTrail records harness operations under AWS::BedrockAgentCore::Runtime. There is no separate harness charge. Framework choice and graph-style workflows are not on Harness; export to code when you need them.

Strands harness is an open-source agent you install and run. Python: pip install strands-harness and create_harness() from strands_harness (Python 3.10+). TypeScript: npm install @strands-agents/harness and createHarness() from @strands-agents/harness (Node.js 20+). The CLI on-ramp is npm install -g @strands-agents/cli. It returns a plain Strands Agent. Providers documented on 11 October 2026: Amazon Bedrock, Anthropic, OpenAI, Google, Ollama, and LiteLLM. It is built on the Strands Harness SDK, which you can use directly when every default is wrong. It does not grant AWS microVMs, Gateway, or Cedar.

A custom Strands agent on AgentCore Runtime is your process, packaged for Runtime (BedrockAgentCoreApp, ARM64 container or CodeZip). Runtime isolates the session. Memory, Gateway, Browser, Code Interpreter, and outbound Identity are calls you make. Exporting a managed harness (agentcore export harness) generates Strands Python so you can take that path without rewriting the tool list by hand. Claude Agent SDK export is still listed as coming soon (export).

Side-by-side diagram: AgentCore Harness is configuration inside AWS, while a Strands agent on Runtime owns the loop and only gets Gateway if the code calls it

Managed execution and a local create_harness() can sit in one program over time. The overlap that hurts is two memories, two tool catalogs, or two trace streams for the same task. Pick one system of record for the contract, one for tool authorization, and one for the run log.

Where each approach fits

Custom application harnessStrands harness SDK and CLIAgentCore HarnessStrands on AgentCore Runtime
Who controls the loopYour codeStrands defaults, every one overridableAWS, from your configYour code on a microVM
Effort to first useful runHighestLow for a research or coding agentLow for a single-domain agentMedium after you have a container and a role
CustomizationTotalHigh, including throwing defaults awayConfig, plus Lambda hooks where the grid says mixedTotal, inside the Runtime contract
Who deploysYouYou (laptop, Lambda, Fargate, EKS, or a container host)AWS, from CreateHarness or agentcore deployYou deploy the artifact; AWS runs the microVM
MemoryWhatever you storeFiles under ./.agent unless you set session.dir and memory.dirAgentCore Memory: short-term, and long-term strategies (semantic, summarization, user preference, episodic) with an actor idYou call Memory, or you bring your own
ObservabilityYour tracesStrands telemetry, then wherever you export itAgentCore traces into CloudWatchRuntime plumbing plus what you emit
Security you still ownAll of itSandbox, tool permissions, untrusted inputIAM, Cedar, prompt checks, skill sources, evalsThe same, plus the framework
Cost metersModel, your compute, your logsModel, your compute, your logsRuntime CPU and memory, model, and any Memory, Gateway, Browser, Code Interpreter, or CloudWatch you useSame AWS meters if you call those services, plus your framework’s token use
Best fitA narrow workflow you already run in codeA coding or research agent with a human nearbyFirst production agent on AWS whose topology fits one loopMulti-step code you will not express as harness config

Employee-only knowledge work has a third door. Compare Quick Suite and AgentCore before you install a shell-capable agent for people who already have seats.

Why a stronger model still fails in production

The failure is usually outside the weights.

The task was vague, so the agent edited a workflow file. The tool list included deploy because the prototype needed it. The context window filled with old tool output and the instruction that mattered was summarized away. A retry repeated a refund whose first response had timed out. The model graded its own patch and called it done. A silent model update changed the tool-call shape and nobody pinned the id.

OpenAI’s write-up is explicit about the early weeks: progress was slow because the environment was underspecified, and the fix was almost never “try harder.” They made the app, the logs, and the metrics legible, and they kept rules in linters so the agent could not forget them. They also still ask Codex to review its own change locally and then request other reviewers. Self-review is one step in that loop. It is not the acceptance test.

What broke — OpenAI, 11 February 2026, on their own repository: one large AGENTS.md crowded out the task, went stale, and could not be checked mechanically. They cut it to a map of about 100 lines and moved the system of record into docs/, with linters for structure and freshness. Source: Harness engineering.

I looked for a study that says self-evaluating agents approve their own output 94% of the time and a separate verifier cuts that to 61%. I could not find a primary source that measures those two rates on agent acceptance. Nearby papers use 94 and 61 for other quantities (math accuracy, iterative self-correction, or “everything looks good” feedback). Those are not this claim. Leave the pair out. Use a test.

The same standard applies to “40% of the budget went to replaying old context” and to “optimize when the accepted-change rate falls below 50%.” I could not trace either as a measured result. Acceptance rate is a team choice. It depends on the value of a good change, the cost of a bad one, and the cost of review. A 50% line is not a standard.

Ten responsibilities

Reference architecture: application entry, task contract, context builder, model loop, tool boundary, memory, policy, independent verification, telemetry, and recovery

Each box in that figure is a responsibility. Prompt text can remind the model. The stop has to live in code, IAM, a sandbox, a policy engine, or a runtime limit.

1. Task contract

Write the objective, the success checks, the allowed surface, and the prohibited operations before the model plans. The plan is a proposal. The contract is the constraint. Record preconditions (clean git status, order id that exists, customer matched to the tenant) and the evidence required before anyone marks the task complete. If the contract cannot name a file, a tool, or a business rule, the task is not ready for an agent.

2. Tool orchestration

Give each agent a small set of tools with typed arguments. Keep planning and execution apart when a wrong call is expensive. Deterministic code should own predictable steps: tax, eligibility, a status lookup with one id. Timeouts, retry caps, and idempotency keys belong on the tool runner. Delegation to a subagent pays off when the extra coordination is cheaper than stuffing the whole job into one context. The model must not be able to emit a tool call that skips validation. On AgentCore Harness, AWS rejects a caller-supplied toolUse block in the final InvokeHarness message, so a client cannot name shell and have it dispatched directly (security). Your own Runtime entrypoint does not get that rejection unless you add it.

3. Context

Build the window from the instructions, files, and tool results this task needs. Do not paste the repository or the full history into every call. When you summarize, keep the contract, the safety rules, and the open decisions verbatim. Mark facts as authoritative, uncertain, stale, or untrusted. Issue bodies, web pages, and tool output are untrusted even when the user is trusted.

4. State and memory

Keep these apart:

KindLives forCorrection
Task stateThis attemptDiscard on cancel
Session historyThis conversationCompact, but keep the contract
Durable memoryAcross sessionsProvenance, expiry, delete
Project instructionsUntil a human changes themReview, do not let the agent edit its own guard file unnoticed
Workflow stateUntil the business process endsRequired to resume after a crash
Business dataIn the system of recordRead it again. Do not treat memory as the order

Tenant and user isolation, retention, and deletion are properties of the store. On AgentCore, per-user memory uses an actor id, and per-user credential brokering through Identity’s token vault requires inbound OAuth. SigV4 does not propagate per-user identity into downstream tools today; AWS says that support is planned (security). Storing a transcript is not the same as extracting a preference. See agent memory for commerce.

5. Verification

Run checks the model does not grade: unit and integration tests, typecheck, lint, policy, schema, migration dry-runs, commerce rules (order state, amount, owner), and a person for high-impact changes. Agent eval suites catch regressions in behavior; they do not replace those checks. A second model is another opinion. It shares failure modes with the first when the rubric is vague. Evaluations belong next to the deterministic suite, which is the subject of the ecommerce regression post.

6. Guardrails

Four layers, from the practical split we use in reviews:

LayerExamplesEnforced by
BehavioralOutput shape, refused topics, task limitsPrompt plus a schema check on the final message. Bedrock Guardrails if you want a managed filter
DataPII redaction, retention, tenant isolation, what may be shownApplication code and the trace sink. A prompt does not redact logs
Tool and actionAllowlist, argument constraints, draft-then-commit, approvalTool runner, IAM, Gateway Cedar
OperationalIteration cap, deadline, token budget, concurrency, cancellationRuntime settings or your loop counter

7. Observability

Record run id, task id, model id and configuration, tool name, redacted arguments, result or error, latency, retries, context truncation, memory hits, verifier results, human decisions, token counts, and the final outcome. Skip raw prompts, full customer payloads, and secrets. The observability post is the longer version of what to dashboard. Traces that contain everything are a second data store you now have to protect.

8. Routing

Send cheap, fast models at classification and extraction when you have eval evidence they hold. Send a stronger model at work where a wrong answer is expensive. Also account for latency, price, and whether that model id is enabled in the account. Do not freeze a nickname ranking from last quarter. A routing model is the wrong tool for “if the intent is WISMO, call getOrder.” That is a table.

AgentCore documents mid-session provider switching that rebuilds history into the next provider’s format. Use it for a measured shift, then re-run evals. The Strands effort argument accepts auto (default), low, medium, high, and off, mapped per provider (model configuration).

9. Reviewed feedback

When the verifier fails:

  1. Store the failure and the evidence.
  2. Name the cause. A bad test is a different fix from a bad tool schema.
  3. Propose one narrow change to instructions, tools, permissions, tests, context selection, routing, or recovery.
  4. A person reviews that change.
  5. Add or update a regression test.
  6. Run the harness on a representative set, not only the one case.
  7. Release in a controlled way and watch both misses and false rejects.

Do not append every rejection to CONSTRAINTS.md. A wrong verifier, or a rule that blocks the common case, makes the system less useful. This is harness maintenance. It is not fine-tuning unless you actually train weights.

10. Recovery

Handle a half-finished task, a process restart, a tool error, a contradictory check, stale memory, a broken task record, and an upstream outage. Cap retries. Open a circuit when the dependency is down. Escalate to a person with the evidence attached.

Before a retry, answer one question: did the first call change remote state? If you do not know, read the system of record. If it changed, do not send the write again. If it did not, retry with the same idempotency key.

Context rot, and what to persist

Context rot is the practical failure of a long window: the instruction you needed is still “in” the transcript and the model no longer behaves as if it were. OpenAI’s failed encyclopedia file is the clean public example. Compaction and summarization, which both AgentCore (sliding_window or summarization) and Strands (offload bulky tool results, compact as the window fills) will do for you, can drop the same line unless you pin it outside the summary.

Keep a reproducible record of the context and the configuration used for a decision you may have to defend: contract version, model id, tool catalog version, and which documents were retrieved. You do not need the full prompt body to do that.

Strands sessions default to ./.agent/sessions with a generated id. Pass session={"id": "..."} to resume, and on anything but a laptop set session={"dir": ...} and memory={"dir": ...} to durable storage (production). AgentCore Memory is a different store: short-term events and long-term records with the strategies listed in the harness versus Runtime grid. Runtime’s own microVM filesystem dies with the session. Memory is how you keep something after that.

Tools, permissions, and prompt injection

Sequence from a model tool proposal through schema validation, authorization, optional human approval, execution, and an independent check, with a deny path

Treat retrieved content as untrusted. OpenAI’s prompt-injection note describes the pattern: a third party hides instructions in content the agent was asked to read. I am not repeating an unverified story about a crafted GitHub issue executing a script. The control is the same either way. The reader of untrusted text should not hold the credentials that perform the side effect. AWS is explicit that Harness does not inspect prompt meaning, and that skills fetched from S3 or Git are treated as trusted input. If callers can set skills or additionalParams on InvokeHarness, they can point the agent at another repository or pass provider fields through, including endpoint overrides on some providers. Strip those fields at your application edge (security).

Strands ships shell, file (read, write, edit), and web tools, and a programmatic_tool_caller that runs model-authored code in Monty. Monty isolates that code. It does not isolate the tools the code calls. The production guide says to sandbox and to use interventions before untrusted input, or to drop the tool. The default system prompt says to confirm before irreversible actions. That sentence is not an authorization boundary.

On AWS, the Gateway path is the one we use for anything that writes. The Gateway post measured tool round-trip, not safety. Safety is Cedar plus the absence of the dangerous tool. InvokeHarness also requires bedrock-agentcore:InvokeAgentRuntime on the harness ARN, because the harness is a Runtime resource. The sample execution role is yours to narrow.

Verification and the feedback loop

Feedback loop from an independent check through a named cause, a reviewed harness change, a regression test, and a controlled release, separate from model training

A coding agent that says “tests passed” has not passed tests until your runner’s exit code says so. A support agent that says “refund issued” has not issued a refund until the order system says so on a fresh read.

Wire the seven steps above into the place you already review changes. The artifact’s verify_coding_task() returns failure codes. An empty list is the only pass. The model is not in that function.

What you pay for

Token price and run price are different bills.

There is no separate AgentCore Harness charge. The pricing page, checked 11 October 2026, bills the capabilities you use. For a harness session that is at least Runtime, plus the model, plus CloudWatch if you keep the traces. Add Memory, Gateway, Policy, Browser, Code Interpreter, or Web Search only when that architecture calls them. A managed harness is not automatically cheaper than a loop you run yourself. You trade engineering time and a platform meter against tokens and a container you operate. Model it on the AgentCore pricing calculator and read the twelve-component pricing note.

Runtime microVMs, consumption rates on that page:

MeterRate
v1 CPU$0.0895 per vCPU-hour
v1 memory$0.00945 per GB-hour
v2 CPU, consumption$0.1276 per vCPU-hour
v2 memory, consumption$0.0169 per GB-hour

CPU can drop during model and tool wait when nothing else is using it. Memory stays billable while the session is up. Runtime v2 reclaims idle memory after 120 seconds. There is a 128 MB minimum and a 1-second minimum. The worked example on the same page, a support-shaped session with 60 seconds of 1 vCPU active time and their stated memory profile, comes to about $0.006703 of Runtime compute per session before model tokens, Gateway, Memory, and logs. That is AWS’s example, not a measurement from this site. A harness session also counts against Runtime quotas, including 2 vCPU and 8 GB per session (limits).

Other meters, same page, include only what you enable:

  • Browser and Code Interpreter: $0.0895 per vCPU-hour and $0.00945 per GB-hour, 128 MB minimum. Leaving them on for ordinary chat is a common way to inflate the platform line. The August harness post already flagged that pattern.
  • Gateway: $0.005 per 1,000 API invocations (list, invoke, ping). Search is $0.025 per 1,000. Tool indexing is $0.02 per 100 tools per month.
  • Policy: $0.000025 per authorization request. Natural-language authoring is $0.13 per 1,000 input tokens, separate from the per-request charge.
  • Web Search: $7 per 1,000 queries.
  • Memory, as of 6 October 2026: short-term ingestion $1.00 per GB (each event billed at least 12 KB and at most 64 KB), retrieval $0.20 per GB, storage $0.10 per GB-month on actual size. Long-term built-in storage $0.75 per 1,000 records per month. Long-term retrieval $0.50 per 1,000 retrievals.
  • Identity: no extra charge when used through Runtime or Gateway.
  • Observability: CloudWatch ingestion, storage, and query. Do not copy a CloudWatch unit price out of an AWS example without checking the CloudWatch page.

What inflates tokens, on any host: a large system prompt, unused tool definitions, retrieved memory, full history, repeated calls, and retries. What inflates Runtime: a long maxLifetime (default 8 hours) and a long idle timeout (default 15 minutes) on sessions that sit warm.

Strands token comparisons, including the 21 September 2026 Harbor numbers, stay in the Strands post. They are not an AgentCore invoice.

Illustrative accounting only. Suppose 10 attempts cost C in total and 4 are accepted by the independent check. Cost per attempt is C/10. Cost per accepted change is C/4. Those denominators answer different questions. Nothing here sets C, and a 4-of-10 outcome is not a target.

AgentCore Harness, as documented

CLI (Node.js 20+), from the current get-started page. The older @aws/agentcore@preview line in the launch blog is not the command that page shows now.

npm install -g @aws/agentcore
agentcore create --name coding-task-demo --model-provider bedrock
agentcore deploy

Python sketch, boto3, DRY_RUN=1 by default. Poll get_harness until READY in real code. Full file: agentcore_invoke_sketch.py.

# Python 3.10+, boto3, region where Harness is GA. DRY_RUN prints; it does not call AWS.
control = boto3.client("bedrock-agentcore-control", region_name="us-west-2")
created = control.create_harness(
    harnessName="coding-task-demo",
    executionRoleArn="arn:aws:iam::123456789012:role/HarnessExecutionRole",
)
# Poll get_harness until status is READY. The get-started page calls this field arn.
harness_arn = created.get("arn") or created.get("harnessArn")
client = boto3.client("bedrock-agentcore", region_name="us-west-2")
response = client.invoke_harness(
    harnessArn=harness_arn,
    runtimeSessionId=str(uuid.uuid4()),  # 36 chars; minimum is 33
    messages=[{"role": "user", "content": [{"text": "Summarize git status. Do not edit files."}]}],
    maxIterations=8,
    timeoutSeconds=900,
)

Stream events follow messageStart, contentBlock*, messageStop. stopReason can be end_turn, tool_use, max_iterations_exceeded, timeout_exceeded, max_output_tokens_exceeded, or hook_stopped when a before_invocation or after_tool_call Lambda hook returns deny. Hooks are mixed: the configuration is yours, the Lambda is code.

Lifecycle defaults if you omit them: maxIterations 75, timeoutSeconds 3600, maxTokens unset, idleRuntimeSessionTimeout 900, maxLifetime 28800. Inline client-side tools need your code. Built-in shell and file_operations, skills, Gateway, Browser, and Code Interpreter are configuration. Skills are trusted content. Review them before you attach a Git URL.

Export when config is not enough:

agentcore export harness --name coding-task-demo --build CodeZip

Then read EXPORT_NOTES.md. The generated agent is a Runtime agent. Deploying it does not delete the harness you exported from. Turn one of them off for that task so you do not operate two loops.

Strands harness, as documented

This API is not the block above.

# Python 3.10+, pip install strands-harness. Confirm the model id in your account.
from strands_harness import create_harness

agent = create_harness(
    model="global.anthropic.claude-opus-5",
    session={"dir": "/var/lib/agent/sessions", "id": "api-design"},
    memory={"dir": "/var/lib/agent/memory"},
)

The model page shows global.anthropic.claude-opus-5 as a Bedrock id and anthropic/claude-sonnet-5 as a provider string. The quickstart’s unnamed default is Claude Opus 5 on Bedrock. Pin one id. effort="high" is optional. The sketch that matches this call is strands_harness_sketch.py. It does not import the AgentCore clients.

Out of the box, Strands also gives you a tuned system prompt (explore, then act, confirm before irreversible steps, verify before finishing), context offload, prompt caching on Bedrock and Anthropic, long-term memory, a generalist subagent, a todos tracker, and Agent Skills when they are present (harness overview). Production hosting targets in their deploy guide include Lambda, Fargate, EKS, and AgentCore. “Listed as a host” means the process can run there. It does not mean CreateHarness.

An earlier post on this site recorded Claude Opus 4.8 as the quickstart default in September 2026. The quickstart checked on 11 October 2026 names Claude Opus 5. Re-read that page when you pin.

Pattern A: a coding agent

The contract is a data type, not a prompt.

# Python 3.10+, stdlib. From harness_policy.py in the artifact folder.
contract = TaskContract(
    task_id="codemod-retry-copy",
    objective="Update the retry message in src/utils/retry.ts and keep tests green.",
    allowed_paths=("src/utils/retry.ts", "src/utils/retry.test.ts"),
    prohibited_operations=("git_push", "deploy", "add_dependency"),
    required_checks=("unit", "typecheck"),
    max_iterations=8,
    deadline_seconds=900,
)

Walk the run in this order.

  1. Read the project instructions and this contract. The instructions are a map. The contract is the scope.
  2. Record git status before any edit. A dirty tree is a precondition failure, not something the agent should tidy by force.
  3. Ask for a plan that names the files it expects to touch. Reject the plan when a path sits outside allowed_paths.
  4. Execute with a command allowlist: git status, git diff, the unit test, the typecheck. git push, dependency installs, and workflow edits are absent from the runner, not merely discouraged.
  5. Run the checks yourself. verify_coding_task() fails the run on an empty diff, a missing check, or a path outside scope. In the sample, a package.json edit returns path_outside_scope:package.json.
  6. Compare the diff to the objective. A green test on the wrong file is a failure.
  7. Write a RunRecord: run id, task id, model id, iterations, token counts, tool failures, outcome. Leave the prompt body out.
  8. Stop when iterations, the deadline, or an unknown remote side effect says stop. require_approval("git_push", None) returns blocked.

retry_decision(mutated_remote_state=None, ...) returns reconcile_before_retry. That is the rule for a git or API call whose HTTP status you never saw.

Pattern B: a commerce support agent

The customer asks for a refund. The agent may draft the reply. It may not be the thing that authorizes the money.

Identity first. The session carries a tenant id and a customer id from your authenticator, not from the message text. getOrder accepts those three ids and returns the minimum fields the reply needs: status, total, and a shipment state. Street address and full payment numbers stay in the order system. This is the same boundary as the support agent and the store security posts.

The refund tool is not on the tool list for that turn. The agent can submit a proposal: order id, amount, reason. require_approval("refund", approval_id) returns blocked until a person in the human-review path writes approval_id. After that, Gateway Cedar (or your API authorizer) checks the amount and the order state again. The model does not mint the approval id. When the write returns, read the order once more and attach that status to the trace. A timeout on the write is mutated_remote_state is None: reconcile, do not send a second refund.

Cancel, inventory adjustment, and any reply that would reveal another customer’s data use the same gate. A prompt that says the agent is a careful associate is the behavioral layer. It is not this gate.

A custom policy, and the line that enforces it

application-policy.yaml is a review document. The header says it is not an AgentCore schema and not a Strands config file. harness_policy.py does not load it. The enforcement map is the part to keep open while you implement.

Abbreviated:

# Custom application policy. Not applied by CreateHarness or create_harness().
schema: factualminds.application-harness-policy/v1
scope:
  allowed_paths:
    - src/utils/retry.ts
    - src/utils/retry.test.ts
iterations:
  max_iterations: 8
  deadline_seconds: 900
cost:
  stop_after_model_calls: 12
approvals:
  required_operations: [git_push, deploy, refund, cancel_order, disclose_pii]
verification:
  required_checks: [unit, typecheck]
  agent_self_approval_counts: false
YAML fieldWhat stops the action
allowed_pathsverify_coding_task(), or a sandbox that cannot read the rest of the disk
max_iterations / deadline_secondsYour loop, or AgentCore maxIterations and timeoutSeconds
stop_after_model_callsCode that refuses the next model call. Not AWS Budgets
required_operationsrequire_approval(), and IAM or Cedar on the tools you exposed
required_checksThe test runner’s exit code
agent_self_approval_counts: falseYou never add a “mark complete” tool the model can call alone

A CONSTRAINTS.md file is a reasonable place to keep rules the reviewer accepted. It becomes real when a linter or the tool runner fails the build for a violation. Until then it is a document the model can ignore, the same way OpenAI’s oversized AGENTS.md was a document the agent could not reliably follow.

How to tell whether it is working

Define the denominator before you celebrate a rate.

MetricDenominator
Task successTasks whose independent checks passed, over tasks you intended the agent to attempt
First-pass verificationAttempts that passed with no harness retry and no human edit
Regression ratePreviously passing cases that fail after a harness change
Acceptance rateChanges a reviewer or a check accepts, over changes proposed
Human correction timeMinutes a person spends per accepted task
Tool failure and retry rateTool calls that error or repeat, over tool calls
Escalation rateTasks handed to a person, over tasks started
Latencyp50, p95, p99 of the whole task, not of one model call
Cost per successful taskModel plus platform plus logs, over tasks that passed
Cost per accepted changeThat same total, over changes that were accepted
Policy violations caughtDenies that matched a real forbidden action
False-positive rejectsDenies or check failures on work that should have shipped
Cost of failed workSpend on attempts that were abandoned or repeated

A high first-pass rate on an easy set is a weak signal. Include cases the agent should refuse, cases with missing data, and cases where the tool fails halfway. Inject those on a schedule. Roll out a harness change to a slice of tasks and watch false rejects, not only the happy path. The evals post covers how a green dashboard hides that.

Mistakes that show up in the first month

  • The tool catalog from the prototype, including shell and a write API, is still attached.
  • Browser or Code Interpreter stays on for turns that only needed a Gateway read.
  • Session files sit on ephemeral disk, so a new task “forgets” and the model re-reads everything.
  • Compaction drops the contract. The agent expands scope.
  • Retries replay a payment or a refund.
  • The agent can edit the file that contains its own allowlist.
  • Two harnesses run the same task after an export, and the traces disagree.
  • Every verifier complaint becomes a new bullet in the prompt, and the prompt gets worse.
  • Model id is “latest”, so a provider-side change looks like a random quality drop.
  • AWS Budgets is the only spend control, and the bill arrives anyway.

The Monday checklist is the short version of the fixes.

What to Do This Week

  1. Pick one workflow. Write the contract: objective, allowed tools or paths, prohibited operations, evidence of done.
  2. Remove every tool the workflow does not need. Refunds, deploys, and pushes require an approval id your code checks.
  3. Add one deterministic check and fail the task when it fails. Do not ask the model whether it passed.
  4. Log model id, tool name, redacted arguments, tokens, and outcome. Set an iteration cap and a deadline in the runner or on InvokeHarness.
  5. If you are on AgentCore, confirm the execution role, turn Browser off, and put writes behind Gateway. If you are on Strands, point session and memory at durable storage and sandbox the shell before any untrusted text.
  6. Price the mix on the calculator. Keep the OpenAI and Strands benchmark numbers in their own columns.

If you want a second pair of eyes on the contract and the tool boundary, start with the AI agents hub or contact the team. We will not pretend a harness diagram is a finished commerce agent. Gateway policy, identity, and the eval file are still yours. The production AgentCore guide is the AWS map. This post is the discipline that map sits inside.

What This Post Doesn’t Cover

  • A worked fine-tune or a change to model weights. The feedback loop here edits the harness.
  • Current Bedrock token prices. They move. Read the model page for the id you pin.
  • A full Cedar syntax tutorial. The Gateway and store security posts carry that.
  • Multi-agent topologies (Graph, Swarm, Workflow). Use the August Strands 1.0 post when you outgrow one loop.
  • Any claim that FactualMinds has shipped a production agent for a named client. We have not published one.

Frequently asked questions

When should I not start on Amazon Bedrock AgentCore Harness?
Skip the managed harness when you already need a graph, a workflow with hard step order, bidirectional streaming, or a framework other than the one AgentCore exports. The harness-versus-Runtime page lists framework choice and non-agent-loop patterns as unsupported on Harness. Use Runtime and your own code, or keep the agent off AgentCore for a single Converse call.
What could go wrong if the system prompt is the only guardrail?
The model can still call any tool you attached. AWS documents that AgentCore Harness validates the shape of InvokeHarness, not the meaning of the prompt. A line that says do not refund does not remove the refund tool. Put refunds, cancels, inventory writes, and private-data reads behind IAM, Gateway Cedar, or application code, and require a human approval id before those tools exist for the call.
Is Strands harness the same product as AgentCore Harness?
No. AgentCore Harness is the managed AWS loop (CreateHarness and InvokeHarness, GA 17 June 2026). Strands harness is the open-source package strands-harness and @strands-agents/harness. create_harness() returns a Strands Agent you run yourself. Hosting that process on AgentCore Runtime does not attach Gateway, Identity, or Policy unless your code calls them.
What could go wrong if an agent approves its own output?
The same model that wrote the change is a poor judge of it. OpenAI's harness write-up still has Codex review its own diff, then request separate reviewers, and it relies on linters and CI for rules that must hold. Prefer a test, a schema check, or a second read of the system of record. A second model can help and can still be wrong.
Does AWS Budgets stop a runaway agent immediately?
No. Budgets and CloudWatch alarms notify you. They are not an instantaneous cap on the next model call. AgentCore documents per-invoke limits you can set: maxIterations (default 75), timeoutSeconds (default 3600), and maxTokens (unset unless you set it). Idle session timeout defaults to 900 seconds and max lifetime to 28800 seconds. An application counter that refuses another call is the other hard stop.
Should every verifier failure become a new permanent rule?
No. Record the failure, name the cause, review a narrow harness change, add a regression test, and watch false rejects after release. A wrong verifier or an over-tight rule reduces the tasks the agent can finish. This loop changes instructions, tools, tests, or permissions. It does not train model weights.

Reference bounds

How the Harness engineering stays bounded

Level 2 — reference architecture. Risk high. Oversight: approval required.

Put the model inside controls for observation, action, recovery, verification, and evaluation.

Explore, then act, then confirm, then verify. Explore gathers the context for the Harness engineering. Act stays reversible. Confirm stops before an irreversible step. Verify is a separate check of the outcome.

Starts when
A team is deciding how a production agent is allowed to act.
Tools
A write goes through a router, a permission check, a policy check, and a budget check. The model does not commit it. An approval token lives in tool context, not in the user message.
Stops for a person
Prompt text is not authorization. Refunds, pushes, deploys, and other writes stay behind a named gate. The model does not approve its own output.
Checked by
An independent check runs before the task is accepted. A second model is not the first check.
Untrusted data
Tool results, tickets, issues, and retrieved memory that try to change the task or the policy.
If it fails
If the Harness engineering stops, name the reason: completed, budget_exceeded, timed_out, cancelled, guardrail_blocked, approval_required, tool_failure, verification_failed, or partial_completion. Retry a timeout at most twice. Back off on rate limits. After repeated verification failure, escalate. Stop when the budget is exhausted or a permission is denied. No unbounded loop. Prompt text is not authorization. Refunds, pushes, deploys, and other writes stay behind a named gate. The model does not approve its own output.

This is the reference architecture for the page, not a published production deployment. The shared contract is the AWS store-agent architecture. Permissions and data boundaries are in securing store agents.

Palaniappan P
Palaniappan P

AWS Cloud Architect & AI Expert

AWS-certified cloud architect and AI expert with deep expertise in cloud migrations, cost optimization, and generative AI on AWS.

AWS ArchitectureCloud MigrationGenAI on AWSCost OptimizationDevOps

Recommended Reading

Explore All Articles »