Harness Engineering for Production AI Agents (October 2026)
Quick summary: OpenAI's 11 February 2026 Codex experiment merged about 1,500 pull requests. AgentCore Harness, GA 17 June 2026, still defaults to 75 iterations and a 3,600-second timeout — the harness sets those bounds, not the model.
Key Takeaways
- OpenAI's 11 February 2026 Codex experiment merged about 1,500 pull requests
- AgentCore Harness, GA 17 June 2026, still defaults to 75 iterations and a 3,600-second timeout — the harness sets those bounds, not the model
- Checked 11 October 2026
- The repository, the linters, the app boot, and the review loop did the rest
- On AWS, AgentCore Harness reached general availability on 17 June 2026

Table of Contents
Checked 11 October 2026. An AI agent in production is a model inside a system that decides what it can see, what it can call, how it recovers, how the work is checked, and how you score the result. A stronger model does not install that system.
OpenAI’s harness engineering note (Ryan Lopopolo, 11 February 2026) describes one internal product built with Codex: on the order of a million lines, about 1,500 merged pull requests, three engineers later seven, about 3.5 pull requests per engineer per day, and an estimate of roughly one tenth the time of writing it by hand. Humans steered. The repository, the linters, the app boot, and the review loop did the rest. Those figures are OpenAI’s account of that experiment, not an industry benchmark and not a FactualMinds result.
On AWS, AgentCore Harness reached general availability on 17 June 2026. The get-started guide still says that if you omit a model, the harness uses Anthropic Claude Sonnet 4.6 on Amazon Bedrock, and that a single invocation stops at 75 iterations or 3,600 seconds unless you set a tighter ceiling (operations). The Strands harness quickstart, read the same day, names a different default: Claude Opus 5 on Bedrock, plus shell, file, and web tools. Same word. Different control plane.
Cited published figures (not a FactualMinds run) — OpenAI, 11 February 2026: about 1,500 merged PRs, about 3.5 PRs per engineer per day,
AGENTS.mdkept to roughly 100 lines, single runs of up to six hours. Source: Harness engineering.
First-party signals we reuse (not a new agent engagement) — Gateway server-side tools cut median tool round-trip ~180 ms to ~95 ms on a B2B CRM assistant (12 tools, ~8k turns/day) — Gateway post. Support-style AgentCore at 50K sessions/mo is ~$791/mo platform plus model — decision guide. There are still zero published FactualMinds case studies of a production AI agent.
Reproduce this — Clone the folder under
examples/architecture-blog-2026/harness-engineering/.python3 -m py_compile harness_policy.py agentcore_invoke_sketch.py strands_harness_sketch.pythenpython3 harness_policy.py. Expected lines includeapproval:blocked,retry:reconcile_before_retry, andverify_ok:pass.
Opinionated take: write the task contract, the tool allowlist, and one deterministic check in application code before you pick a host. Use AgentCore Harness when you want AWS to run a single-agent loop and you will still author IAM, Cedar, and evals. Use Strands harness for a general-purpose coding or research agent you can run locally, and do not point its default shell at orders or payments. Trade-off: you give up the one-line local loop when you move a commerce workflow onto Gateway. You keep a boundary the model cannot talk its way through.
Four meanings of “harness”
Teams collide these. Separate them before you compare prices.
Harness engineering is the discipline. You design the task contract, the context the model sees, the tools it may call, the memory it may keep, the checks that accept the work, the permissions around side effects, and the record you use later. You can build that with application code, an agent framework, a managed service, or a mix.
Amazon Bedrock AgentCore Harness is a managed loop on AgentCore Runtime. You declare model, instructions, tools, memory, and limits. AWS runs the loop inside a Firecracker microVM per session. CloudTrail records harness operations under AWS::BedrockAgentCore::Runtime. There is no separate harness charge. Framework choice and graph-style workflows are not on Harness; export to code when you need them.
Strands harness is an open-source agent you install and run. Python: pip install strands-harness and create_harness() from strands_harness (Python 3.10+). TypeScript: npm install @strands-agents/harness and createHarness() from @strands-agents/harness (Node.js 20+). The CLI on-ramp is npm install -g @strands-agents/cli. It returns a plain Strands Agent. Providers documented on 11 October 2026: Amazon Bedrock, Anthropic, OpenAI, Google, Ollama, and LiteLLM. It is built on the Strands Harness SDK, which you can use directly when every default is wrong. It does not grant AWS microVMs, Gateway, or Cedar.
A custom Strands agent on AgentCore Runtime is your process, packaged for Runtime (BedrockAgentCoreApp, ARM64 container or CodeZip). Runtime isolates the session. Memory, Gateway, Browser, Code Interpreter, and outbound Identity are calls you make. Exporting a managed harness (agentcore export harness) generates Strands Python so you can take that path without rewriting the tool list by hand. Claude Agent SDK export is still listed as coming soon (export).

Managed execution and a local create_harness() can sit in one program over time. The overlap that hurts is two memories, two tool catalogs, or two trace streams for the same task. Pick one system of record for the contract, one for tool authorization, and one for the run log.
Where each approach fits
| Custom application harness | Strands harness SDK and CLI | AgentCore Harness | Strands on AgentCore Runtime | |
|---|---|---|---|---|
| Who controls the loop | Your code | Strands defaults, every one overridable | AWS, from your config | Your code on a microVM |
| Effort to first useful run | Highest | Low for a research or coding agent | Low for a single-domain agent | Medium after you have a container and a role |
| Customization | Total | High, including throwing defaults away | Config, plus Lambda hooks where the grid says mixed | Total, inside the Runtime contract |
| Who deploys | You | You (laptop, Lambda, Fargate, EKS, or a container host) | AWS, from CreateHarness or agentcore deploy | You deploy the artifact; AWS runs the microVM |
| Memory | Whatever you store | Files under ./.agent unless you set session.dir and memory.dir | AgentCore Memory: short-term, and long-term strategies (semantic, summarization, user preference, episodic) with an actor id | You call Memory, or you bring your own |
| Observability | Your traces | Strands telemetry, then wherever you export it | AgentCore traces into CloudWatch | Runtime plumbing plus what you emit |
| Security you still own | All of it | Sandbox, tool permissions, untrusted input | IAM, Cedar, prompt checks, skill sources, evals | The same, plus the framework |
| Cost meters | Model, your compute, your logs | Model, your compute, your logs | Runtime CPU and memory, model, and any Memory, Gateway, Browser, Code Interpreter, or CloudWatch you use | Same AWS meters if you call those services, plus your framework’s token use |
| Best fit | A narrow workflow you already run in code | A coding or research agent with a human nearby | First production agent on AWS whose topology fits one loop | Multi-step code you will not express as harness config |
Employee-only knowledge work has a third door. Compare Quick Suite and AgentCore before you install a shell-capable agent for people who already have seats.
Why a stronger model still fails in production
The failure is usually outside the weights.
The task was vague, so the agent edited a workflow file. The tool list included deploy because the prototype needed it. The context window filled with old tool output and the instruction that mattered was summarized away. A retry repeated a refund whose first response had timed out. The model graded its own patch and called it done. A silent model update changed the tool-call shape and nobody pinned the id.
OpenAI’s write-up is explicit about the early weeks: progress was slow because the environment was underspecified, and the fix was almost never “try harder.” They made the app, the logs, and the metrics legible, and they kept rules in linters so the agent could not forget them. They also still ask Codex to review its own change locally and then request other reviewers. Self-review is one step in that loop. It is not the acceptance test.
What broke — OpenAI, 11 February 2026, on their own repository: one large
AGENTS.mdcrowded out the task, went stale, and could not be checked mechanically. They cut it to a map of about 100 lines and moved the system of record intodocs/, with linters for structure and freshness. Source: Harness engineering.
I looked for a study that says self-evaluating agents approve their own output 94% of the time and a separate verifier cuts that to 61%. I could not find a primary source that measures those two rates on agent acceptance. Nearby papers use 94 and 61 for other quantities (math accuracy, iterative self-correction, or “everything looks good” feedback). Those are not this claim. Leave the pair out. Use a test.
The same standard applies to “40% of the budget went to replaying old context” and to “optimize when the accepted-change rate falls below 50%.” I could not trace either as a measured result. Acceptance rate is a team choice. It depends on the value of a good change, the cost of a bad one, and the cost of review. A 50% line is not a standard.
Ten responsibilities

Each box in that figure is a responsibility. Prompt text can remind the model. The stop has to live in code, IAM, a sandbox, a policy engine, or a runtime limit.
1. Task contract
Write the objective, the success checks, the allowed surface, and the prohibited operations before the model plans. The plan is a proposal. The contract is the constraint. Record preconditions (clean git status, order id that exists, customer matched to the tenant) and the evidence required before anyone marks the task complete. If the contract cannot name a file, a tool, or a business rule, the task is not ready for an agent.
2. Tool orchestration
Give each agent a small set of tools with typed arguments. Keep planning and execution apart when a wrong call is expensive. Deterministic code should own predictable steps: tax, eligibility, a status lookup with one id. Timeouts, retry caps, and idempotency keys belong on the tool runner. Delegation to a subagent pays off when the extra coordination is cheaper than stuffing the whole job into one context. The model must not be able to emit a tool call that skips validation. On AgentCore Harness, AWS rejects a caller-supplied toolUse block in the final InvokeHarness message, so a client cannot name shell and have it dispatched directly (security). Your own Runtime entrypoint does not get that rejection unless you add it.
3. Context
Build the window from the instructions, files, and tool results this task needs. Do not paste the repository or the full history into every call. When you summarize, keep the contract, the safety rules, and the open decisions verbatim. Mark facts as authoritative, uncertain, stale, or untrusted. Issue bodies, web pages, and tool output are untrusted even when the user is trusted.
4. State and memory
Keep these apart:
| Kind | Lives for | Correction |
|---|---|---|
| Task state | This attempt | Discard on cancel |
| Session history | This conversation | Compact, but keep the contract |
| Durable memory | Across sessions | Provenance, expiry, delete |
| Project instructions | Until a human changes them | Review, do not let the agent edit its own guard file unnoticed |
| Workflow state | Until the business process ends | Required to resume after a crash |
| Business data | In the system of record | Read it again. Do not treat memory as the order |
Tenant and user isolation, retention, and deletion are properties of the store. On AgentCore, per-user memory uses an actor id, and per-user credential brokering through Identity’s token vault requires inbound OAuth. SigV4 does not propagate per-user identity into downstream tools today; AWS says that support is planned (security). Storing a transcript is not the same as extracting a preference. See agent memory for commerce.
5. Verification
Run checks the model does not grade: unit and integration tests, typecheck, lint, policy, schema, migration dry-runs, commerce rules (order state, amount, owner), and a person for high-impact changes. Agent eval suites catch regressions in behavior; they do not replace those checks. A second model is another opinion. It shares failure modes with the first when the rubric is vague. Evaluations belong next to the deterministic suite, which is the subject of the ecommerce regression post.
6. Guardrails
Four layers, from the practical split we use in reviews:
| Layer | Examples | Enforced by |
|---|---|---|
| Behavioral | Output shape, refused topics, task limits | Prompt plus a schema check on the final message. Bedrock Guardrails if you want a managed filter |
| Data | PII redaction, retention, tenant isolation, what may be shown | Application code and the trace sink. A prompt does not redact logs |
| Tool and action | Allowlist, argument constraints, draft-then-commit, approval | Tool runner, IAM, Gateway Cedar |
| Operational | Iteration cap, deadline, token budget, concurrency, cancellation | Runtime settings or your loop counter |
7. Observability
Record run id, task id, model id and configuration, tool name, redacted arguments, result or error, latency, retries, context truncation, memory hits, verifier results, human decisions, token counts, and the final outcome. Skip raw prompts, full customer payloads, and secrets. The observability post is the longer version of what to dashboard. Traces that contain everything are a second data store you now have to protect.
8. Routing
Send cheap, fast models at classification and extraction when you have eval evidence they hold. Send a stronger model at work where a wrong answer is expensive. Also account for latency, price, and whether that model id is enabled in the account. Do not freeze a nickname ranking from last quarter. A routing model is the wrong tool for “if the intent is WISMO, call getOrder.” That is a table.
AgentCore documents mid-session provider switching that rebuilds history into the next provider’s format. Use it for a measured shift, then re-run evals. The Strands effort argument accepts auto (default), low, medium, high, and off, mapped per provider (model configuration).
9. Reviewed feedback
When the verifier fails:
- Store the failure and the evidence.
- Name the cause. A bad test is a different fix from a bad tool schema.
- Propose one narrow change to instructions, tools, permissions, tests, context selection, routing, or recovery.
- A person reviews that change.
- Add or update a regression test.
- Run the harness on a representative set, not only the one case.
- Release in a controlled way and watch both misses and false rejects.
Do not append every rejection to CONSTRAINTS.md. A wrong verifier, or a rule that blocks the common case, makes the system less useful. This is harness maintenance. It is not fine-tuning unless you actually train weights.
10. Recovery
Handle a half-finished task, a process restart, a tool error, a contradictory check, stale memory, a broken task record, and an upstream outage. Cap retries. Open a circuit when the dependency is down. Escalate to a person with the evidence attached.
Before a retry, answer one question: did the first call change remote state? If you do not know, read the system of record. If it changed, do not send the write again. If it did not, retry with the same idempotency key.
Context rot, and what to persist
Context rot is the practical failure of a long window: the instruction you needed is still “in” the transcript and the model no longer behaves as if it were. OpenAI’s failed encyclopedia file is the clean public example. Compaction and summarization, which both AgentCore (sliding_window or summarization) and Strands (offload bulky tool results, compact as the window fills) will do for you, can drop the same line unless you pin it outside the summary.
Keep a reproducible record of the context and the configuration used for a decision you may have to defend: contract version, model id, tool catalog version, and which documents were retrieved. You do not need the full prompt body to do that.
Strands sessions default to ./.agent/sessions with a generated id. Pass session={"id": "..."} to resume, and on anything but a laptop set session={"dir": ...} and memory={"dir": ...} to durable storage (production). AgentCore Memory is a different store: short-term events and long-term records with the strategies listed in the harness versus Runtime grid. Runtime’s own microVM filesystem dies with the session. Memory is how you keep something after that.
Tools, permissions, and prompt injection

Treat retrieved content as untrusted. OpenAI’s prompt-injection note describes the pattern: a third party hides instructions in content the agent was asked to read. I am not repeating an unverified story about a crafted GitHub issue executing a script. The control is the same either way. The reader of untrusted text should not hold the credentials that perform the side effect. AWS is explicit that Harness does not inspect prompt meaning, and that skills fetched from S3 or Git are treated as trusted input. If callers can set skills or additionalParams on InvokeHarness, they can point the agent at another repository or pass provider fields through, including endpoint overrides on some providers. Strip those fields at your application edge (security).
Strands ships shell, file (read, write, edit), and web tools, and a programmatic_tool_caller that runs model-authored code in Monty. Monty isolates that code. It does not isolate the tools the code calls. The production guide says to sandbox and to use interventions before untrusted input, or to drop the tool. The default system prompt says to confirm before irreversible actions. That sentence is not an authorization boundary.
On AWS, the Gateway path is the one we use for anything that writes. The Gateway post measured tool round-trip, not safety. Safety is Cedar plus the absence of the dangerous tool. InvokeHarness also requires bedrock-agentcore:InvokeAgentRuntime on the harness ARN, because the harness is a Runtime resource. The sample execution role is yours to narrow.
Verification and the feedback loop

A coding agent that says “tests passed” has not passed tests until your runner’s exit code says so. A support agent that says “refund issued” has not issued a refund until the order system says so on a fresh read.
Wire the seven steps above into the place you already review changes. The artifact’s verify_coding_task() returns failure codes. An empty list is the only pass. The model is not in that function.
What you pay for
Token price and run price are different bills.
There is no separate AgentCore Harness charge. The pricing page, checked 11 October 2026, bills the capabilities you use. For a harness session that is at least Runtime, plus the model, plus CloudWatch if you keep the traces. Add Memory, Gateway, Policy, Browser, Code Interpreter, or Web Search only when that architecture calls them. A managed harness is not automatically cheaper than a loop you run yourself. You trade engineering time and a platform meter against tokens and a container you operate. Model it on the AgentCore pricing calculator and read the twelve-component pricing note.
Runtime microVMs, consumption rates on that page:
| Meter | Rate |
|---|---|
| v1 CPU | $0.0895 per vCPU-hour |
| v1 memory | $0.00945 per GB-hour |
| v2 CPU, consumption | $0.1276 per vCPU-hour |
| v2 memory, consumption | $0.0169 per GB-hour |
CPU can drop during model and tool wait when nothing else is using it. Memory stays billable while the session is up. Runtime v2 reclaims idle memory after 120 seconds. There is a 128 MB minimum and a 1-second minimum. The worked example on the same page, a support-shaped session with 60 seconds of 1 vCPU active time and their stated memory profile, comes to about $0.006703 of Runtime compute per session before model tokens, Gateway, Memory, and logs. That is AWS’s example, not a measurement from this site. A harness session also counts against Runtime quotas, including 2 vCPU and 8 GB per session (limits).
Other meters, same page, include only what you enable:
- Browser and Code Interpreter: $0.0895 per vCPU-hour and $0.00945 per GB-hour, 128 MB minimum. Leaving them on for ordinary chat is a common way to inflate the platform line. The August harness post already flagged that pattern.
- Gateway: $0.005 per 1,000 API invocations (list, invoke, ping). Search is $0.025 per 1,000. Tool indexing is $0.02 per 100 tools per month.
- Policy: $0.000025 per authorization request. Natural-language authoring is $0.13 per 1,000 input tokens, separate from the per-request charge.
- Web Search: $7 per 1,000 queries.
- Memory, as of 6 October 2026: short-term ingestion $1.00 per GB (each event billed at least 12 KB and at most 64 KB), retrieval $0.20 per GB, storage $0.10 per GB-month on actual size. Long-term built-in storage $0.75 per 1,000 records per month. Long-term retrieval $0.50 per 1,000 retrievals.
- Identity: no extra charge when used through Runtime or Gateway.
- Observability: CloudWatch ingestion, storage, and query. Do not copy a CloudWatch unit price out of an AWS example without checking the CloudWatch page.
What inflates tokens, on any host: a large system prompt, unused tool definitions, retrieved memory, full history, repeated calls, and retries. What inflates Runtime: a long maxLifetime (default 8 hours) and a long idle timeout (default 15 minutes) on sessions that sit warm.
Strands token comparisons, including the 21 September 2026 Harbor numbers, stay in the Strands post. They are not an AgentCore invoice.
Illustrative accounting only. Suppose 10 attempts cost C in total and 4 are accepted by the independent check. Cost per attempt is C/10. Cost per accepted change is C/4. Those denominators answer different questions. Nothing here sets C, and a 4-of-10 outcome is not a target.
AgentCore Harness, as documented
CLI (Node.js 20+), from the current get-started page. The older @aws/agentcore@preview line in the launch blog is not the command that page shows now.
npm install -g @aws/agentcore
agentcore create --name coding-task-demo --model-provider bedrock
agentcore deployPython sketch, boto3, DRY_RUN=1 by default. Poll get_harness until READY in real code. Full file: agentcore_invoke_sketch.py.
# Python 3.10+, boto3, region where Harness is GA. DRY_RUN prints; it does not call AWS.
control = boto3.client("bedrock-agentcore-control", region_name="us-west-2")
created = control.create_harness(
harnessName="coding-task-demo",
executionRoleArn="arn:aws:iam::123456789012:role/HarnessExecutionRole",
)
# Poll get_harness until status is READY. The get-started page calls this field arn.
harness_arn = created.get("arn") or created.get("harnessArn")
client = boto3.client("bedrock-agentcore", region_name="us-west-2")
response = client.invoke_harness(
harnessArn=harness_arn,
runtimeSessionId=str(uuid.uuid4()), # 36 chars; minimum is 33
messages=[{"role": "user", "content": [{"text": "Summarize git status. Do not edit files."}]}],
maxIterations=8,
timeoutSeconds=900,
)Stream events follow messageStart, contentBlock*, messageStop. stopReason can be end_turn, tool_use, max_iterations_exceeded, timeout_exceeded, max_output_tokens_exceeded, or hook_stopped when a before_invocation or after_tool_call Lambda hook returns deny. Hooks are mixed: the configuration is yours, the Lambda is code.
Lifecycle defaults if you omit them: maxIterations 75, timeoutSeconds 3600, maxTokens unset, idleRuntimeSessionTimeout 900, maxLifetime 28800. Inline client-side tools need your code. Built-in shell and file_operations, skills, Gateway, Browser, and Code Interpreter are configuration. Skills are trusted content. Review them before you attach a Git URL.
Export when config is not enough:
agentcore export harness --name coding-task-demo --build CodeZipThen read EXPORT_NOTES.md. The generated agent is a Runtime agent. Deploying it does not delete the harness you exported from. Turn one of them off for that task so you do not operate two loops.
Strands harness, as documented
This API is not the block above.
# Python 3.10+, pip install strands-harness. Confirm the model id in your account.
from strands_harness import create_harness
agent = create_harness(
model="global.anthropic.claude-opus-5",
session={"dir": "/var/lib/agent/sessions", "id": "api-design"},
memory={"dir": "/var/lib/agent/memory"},
)The model page shows global.anthropic.claude-opus-5 as a Bedrock id and anthropic/claude-sonnet-5 as a provider string. The quickstart’s unnamed default is Claude Opus 5 on Bedrock. Pin one id. effort="high" is optional. The sketch that matches this call is strands_harness_sketch.py. It does not import the AgentCore clients.
Out of the box, Strands also gives you a tuned system prompt (explore, then act, confirm before irreversible steps, verify before finishing), context offload, prompt caching on Bedrock and Anthropic, long-term memory, a generalist subagent, a todos tracker, and Agent Skills when they are present (harness overview). Production hosting targets in their deploy guide include Lambda, Fargate, EKS, and AgentCore. “Listed as a host” means the process can run there. It does not mean CreateHarness.
An earlier post on this site recorded Claude Opus 4.8 as the quickstart default in September 2026. The quickstart checked on 11 October 2026 names Claude Opus 5. Re-read that page when you pin.
Pattern A: a coding agent
The contract is a data type, not a prompt.
# Python 3.10+, stdlib. From harness_policy.py in the artifact folder.
contract = TaskContract(
task_id="codemod-retry-copy",
objective="Update the retry message in src/utils/retry.ts and keep tests green.",
allowed_paths=("src/utils/retry.ts", "src/utils/retry.test.ts"),
prohibited_operations=("git_push", "deploy", "add_dependency"),
required_checks=("unit", "typecheck"),
max_iterations=8,
deadline_seconds=900,
)Walk the run in this order.
- Read the project instructions and this contract. The instructions are a map. The contract is the scope.
- Record
git statusbefore any edit. A dirty tree is a precondition failure, not something the agent should tidy by force. - Ask for a plan that names the files it expects to touch. Reject the plan when a path sits outside
allowed_paths. - Execute with a command allowlist:
git status,git diff, the unit test, the typecheck.git push, dependency installs, and workflow edits are absent from the runner, not merely discouraged. - Run the checks yourself.
verify_coding_task()fails the run on an empty diff, a missing check, or a path outside scope. In the sample, apackage.jsonedit returnspath_outside_scope:package.json. - Compare the diff to the objective. A green test on the wrong file is a failure.
- Write a
RunRecord: run id, task id, model id, iterations, token counts, tool failures, outcome. Leave the prompt body out. - Stop when iterations, the deadline, or an unknown remote side effect says stop.
require_approval("git_push", None)returnsblocked.
retry_decision(mutated_remote_state=None, ...) returns reconcile_before_retry. That is the rule for a git or API call whose HTTP status you never saw.
Pattern B: a commerce support agent
The customer asks for a refund. The agent may draft the reply. It may not be the thing that authorizes the money.
Identity first. The session carries a tenant id and a customer id from your authenticator, not from the message text. getOrder accepts those three ids and returns the minimum fields the reply needs: status, total, and a shipment state. Street address and full payment numbers stay in the order system. This is the same boundary as the support agent and the store security posts.
The refund tool is not on the tool list for that turn. The agent can submit a proposal: order id, amount, reason. require_approval("refund", approval_id) returns blocked until a person in the human-review path writes approval_id. After that, Gateway Cedar (or your API authorizer) checks the amount and the order state again. The model does not mint the approval id. When the write returns, read the order once more and attach that status to the trace. A timeout on the write is mutated_remote_state is None: reconcile, do not send a second refund.
Cancel, inventory adjustment, and any reply that would reveal another customer’s data use the same gate. A prompt that says the agent is a careful associate is the behavioral layer. It is not this gate.
A custom policy, and the line that enforces it
application-policy.yaml is a review document. The header says it is not an AgentCore schema and not a Strands config file. harness_policy.py does not load it. The enforcement map is the part to keep open while you implement.
Abbreviated:
# Custom application policy. Not applied by CreateHarness or create_harness().
schema: factualminds.application-harness-policy/v1
scope:
allowed_paths:
- src/utils/retry.ts
- src/utils/retry.test.ts
iterations:
max_iterations: 8
deadline_seconds: 900
cost:
stop_after_model_calls: 12
approvals:
required_operations: [git_push, deploy, refund, cancel_order, disclose_pii]
verification:
required_checks: [unit, typecheck]
agent_self_approval_counts: false| YAML field | What stops the action |
|---|---|
allowed_paths | verify_coding_task(), or a sandbox that cannot read the rest of the disk |
max_iterations / deadline_seconds | Your loop, or AgentCore maxIterations and timeoutSeconds |
stop_after_model_calls | Code that refuses the next model call. Not AWS Budgets |
required_operations | require_approval(), and IAM or Cedar on the tools you exposed |
required_checks | The test runner’s exit code |
agent_self_approval_counts: false | You never add a “mark complete” tool the model can call alone |
A CONSTRAINTS.md file is a reasonable place to keep rules the reviewer accepted. It becomes real when a linter or the tool runner fails the build for a violation. Until then it is a document the model can ignore, the same way OpenAI’s oversized AGENTS.md was a document the agent could not reliably follow.
How to tell whether it is working
Define the denominator before you celebrate a rate.
| Metric | Denominator |
|---|---|
| Task success | Tasks whose independent checks passed, over tasks you intended the agent to attempt |
| First-pass verification | Attempts that passed with no harness retry and no human edit |
| Regression rate | Previously passing cases that fail after a harness change |
| Acceptance rate | Changes a reviewer or a check accepts, over changes proposed |
| Human correction time | Minutes a person spends per accepted task |
| Tool failure and retry rate | Tool calls that error or repeat, over tool calls |
| Escalation rate | Tasks handed to a person, over tasks started |
| Latency | p50, p95, p99 of the whole task, not of one model call |
| Cost per successful task | Model plus platform plus logs, over tasks that passed |
| Cost per accepted change | That same total, over changes that were accepted |
| Policy violations caught | Denies that matched a real forbidden action |
| False-positive rejects | Denies or check failures on work that should have shipped |
| Cost of failed work | Spend on attempts that were abandoned or repeated |
A high first-pass rate on an easy set is a weak signal. Include cases the agent should refuse, cases with missing data, and cases where the tool fails halfway. Inject those on a schedule. Roll out a harness change to a slice of tasks and watch false rejects, not only the happy path. The evals post covers how a green dashboard hides that.
Mistakes that show up in the first month
- The tool catalog from the prototype, including shell and a write API, is still attached.
- Browser or Code Interpreter stays on for turns that only needed a Gateway read.
- Session files sit on ephemeral disk, so a new task “forgets” and the model re-reads everything.
- Compaction drops the contract. The agent expands scope.
- Retries replay a payment or a refund.
- The agent can edit the file that contains its own allowlist.
- Two harnesses run the same task after an export, and the traces disagree.
- Every verifier complaint becomes a new bullet in the prompt, and the prompt gets worse.
- Model id is “latest”, so a provider-side change looks like a random quality drop.
- AWS Budgets is the only spend control, and the bill arrives anyway.
The Monday checklist is the short version of the fixes.
What to Do This Week
- Pick one workflow. Write the contract: objective, allowed tools or paths, prohibited operations, evidence of done.
- Remove every tool the workflow does not need. Refunds, deploys, and pushes require an approval id your code checks.
- Add one deterministic check and fail the task when it fails. Do not ask the model whether it passed.
- Log model id, tool name, redacted arguments, tokens, and outcome. Set an iteration cap and a deadline in the runner or on
InvokeHarness. - If you are on AgentCore, confirm the execution role, turn Browser off, and put writes behind Gateway. If you are on Strands, point session and memory at durable storage and sandbox the shell before any untrusted text.
- Price the mix on the calculator. Keep the OpenAI and Strands benchmark numbers in their own columns.
If you want a second pair of eyes on the contract and the tool boundary, start with the AI agents hub or contact the team. We will not pretend a harness diagram is a finished commerce agent. Gateway policy, identity, and the eval file are still yours. The production AgentCore guide is the AWS map. This post is the discipline that map sits inside.
What This Post Doesn’t Cover
- A worked fine-tune or a change to model weights. The feedback loop here edits the harness.
- Current Bedrock token prices. They move. Read the model page for the id you pin.
- A full Cedar syntax tutorial. The Gateway and store security posts carry that.
- Multi-agent topologies (Graph, Swarm, Workflow). Use the August Strands 1.0 post when you outgrow one loop.
- Any claim that FactualMinds has shipped a production agent for a named client. We have not published one.
Frequently asked questions
When should I not start on Amazon Bedrock AgentCore Harness?
What could go wrong if the system prompt is the only guardrail?
Is Strands harness the same product as AgentCore Harness?
What could go wrong if an agent approves its own output?
Does AWS Budgets stop a runaway agent immediately?
Should every verifier failure become a new permanent rule?
Reference bounds
How the Harness engineering stays bounded
Level 2 — reference architecture. Risk high. Oversight: approval required.
Put the model inside controls for observation, action, recovery, verification, and evaluation.
Explore, then act, then confirm, then verify. Explore gathers the context for the Harness engineering. Act stays reversible. Confirm stops before an irreversible step. Verify is a separate check of the outcome.
- Starts when
- A team is deciding how a production agent is allowed to act.
- Tools
- A write goes through a router, a permission check, a policy check, and a budget check. The model does not commit it. An approval token lives in tool context, not in the user message.
- Stops for a person
- Prompt text is not authorization. Refunds, pushes, deploys, and other writes stay behind a named gate. The model does not approve its own output.
- Checked by
- An independent check runs before the task is accepted. A second model is not the first check.
- Untrusted data
- Tool results, tickets, issues, and retrieved memory that try to change the task or the policy.
- If it fails
- If the Harness engineering stops, name the reason: completed, budget_exceeded, timed_out, cancelled, guardrail_blocked, approval_required, tool_failure, verification_failed, or partial_completion. Retry a timeout at most twice. Back off on rate limits. After repeated verification failure, escalate. Stop when the budget is exhausted or a permission is denied. No unbounded loop. Prompt text is not authorization. Refunds, pushes, deploys, and other writes stay behind a named gate. The model does not approve its own output.
This is the reference architecture for the page, not a published production deployment. The shared contract is the AWS store-agent architecture. Permissions and data boundaries are in securing store agents.

AWS Cloud Architect & AI Expert
AWS-certified cloud architect and AI expert with deep expertise in cloud migrations, cost optimization, and generative AI on AWS.




