AI Harness
Notes while exploring the harness: everything around the model that turns it into a working agent.
One-line summary
Agent = Model + Harness. The model thinks. The harness lets it act safely and reliably.
Pages in this section:
- Overview (this page): what a harness is and how the loop works
- Building Blocks: the 8 parts of a harness, each with an example
- Failure Modes & Future: what goes wrong, enterprise use, where it's heading
Related: Agentic Workflows Catalog covers the workflow layer on top of the harness.
Model vs. harness vs. agent
flowchart TB
subgraph Agent["π€ Agent: the full working system"]
direction TB
M["π§ Model<br/>reasons, decides next step"]
subgraph H["π¦Ύ Harness: body + workspace"]
direction LR
T[Tools]
Mem[Memory]
W[Workspace / files]
G[Guardrails]
end
M <--> H
end
User([π€ You]) --> Agent
H <--> World[("π Real world<br/>code, APIs, DBs, email")]
| Part | What it does | Analogy |
|---|---|---|
| Model | Reads context, reasons, picks next action. Has no memory and can't act on its own | Brain |
| Harness | Runs tools, manages memory, enforces rules | Body + desk + rulebook |
| Agent | Model + harness together | A worker who can think and act |
Same question, with and without a harness
Task: "How many failed orders were there yesterday?"
- Model only: "I don't have access to your database, but you could run
SELECT COUNT(*) ..." - Model + harness: the model writes the SQL β the harness runs it on the DB β the result
42comes back β the model answers "42 failed orders, mostly timeouts from venue X."
Where the harness sits
flowchart TB
WF["π Workflow<br/>my process: plan β implement β review"]
HR["π¦Ύ Harness<br/>Claude Code, Codex CLI, Gemini CLI"]
MD["π§ Model<br/>Claude, GPT, Gemini, Llama"]
WF --> HR --> MD
- Model: bought or picked, not built.
- Harness: the runtime. You configure it (tools, permissions, hooks) or build your own.
- Workflow: your process on top. See the catalog.
The core loop: Reason β Act β Observe
This loop is at the center of every agent. It's called ReAct, from Yao et al., 2022.
flowchart LR
Task([Task]) --> R
R["π§ Reason<br/>model reads context,<br/>picks next action"] -->|tool call| A
A["π¦Ύ Act<br/>harness runs tool /<br/>code / API"] --> O
O["π Observe<br/>harness adds result<br/>to context"] --> R
R -->|no more actions| Done([β
Done])
| Step | Who | What |
|---|---|---|
| Reason | Model | Read task + memory + previous results β decide the next action |
| Act | Harness | Run the tool, sandboxed code, or API call |
| Observe | Harness | Capture the result and feed it back as new context |
| Repeat | Both | Until the model says it's done |
Coding agent fixing a bug
sequenceDiagram
participant U as π€ User
participant M as π§ Model
participant H as π¦Ύ Harness
participant S as π¦ Sandbox
U->>M: "Fix failing test in OrderService"
M->>H: read OrderService.java
H->>S: cat file
S-->>H: file content
H-->>M: file content
M->>H: edit line 42, run tests
H->>S: apply edit, mvn test
S-->>H: β 1 test failed (NPE)
H-->>M: test output
Note over M: Reason: null check missing
M->>H: edit line 40, run tests
H->>S: apply edit, mvn test
S-->>H: β
all tests pass
H-->>M: test output
M->>U: "Fixed: added null check for venue"
The model never touches the file system. It only asks, and the harness does.
Why the harness matters
Key idea
Models are getting similar in ability, so the harness now decides performance. The same model can rank very differently on a benchmark depending on its harness.
quadrantChart
title Model vs harness quality
x-axis Weak harness --> Strong harness
y-axis Weaker model --> Stronger model
quadrant-1 Best
quadrant-2 Wasted model
quadrant-3 Unreliable
quadrant-4 Often good enough
Strong model weak harness: [0.2, 0.8]
Mid model strong harness: [0.8, 0.35]
- A strong harness + mid-tier model can beat a weak harness + strong model on workflow-heavy tasks.
- Data point: Databricks OfficeQA Pro Agent Harness scored 36.10% with GPT-5.4 and 52.63% with GPT-5.5, almost halving errors. The harness is what turns a model upgrade into reliable results.
Prompt β context β harness engineering
The focus has moved outward from the model. Each stage sits inside the next.
flowchart TB
subgraph HE["Harness engineering: the whole system"]
subgraph CE["Context engineering: what the model sees"]
PE["Prompt engineering:<br/>the wording"]
end
T[Tools]
S[Sandbox]
L[Loops]
G[Guardrails]
end
| Discipline | Focus | Example |
|---|---|---|
| Prompt engineering | Wording of the input | "You are a senior Java dev. Answer in bullet points." |
| Context engineering | What info goes in the window, and when | RAG: fetch the 3 most relevant docs before answering |
| Harness engineering | Tools, sandbox, loops, guardrails | Run mvn test after every edit; block git push without approval |
Harness vs. workflow: where does a change go?
flowchart TD
Q{Must it happen<br/>**every** time?}
Q -->|Yes, deterministic| H["π¦Ύ Harness<br/>hook, permission, tool"]
Q -->|Usually, judgment call| W["π Workflow<br/>prompt, command, CLAUDE.md rule"]
| I want to⦠| Layer | How (Claude Code) |
|---|---|---|
Block rm -rf every time |
Harness | permission deny rule |
| Always run tests after an edit | Harness | PostToolUse hook |
| Plan before coding | Workflow | /plan command or prompt |
| Keep project rules | Both | CLAUDE.md (the harness loads it, you write the content) |
| Let the agent query KDB | Harness | MCP server / tool |
Exploration checklist
- Read the Databricks article and add its key points
- Map each building block to Claude Code features in detail
- Compare harnesses: Claude Code vs. Codex CLI vs. Gemini CLI
- Build a minimal harness (ReAct loop + 2 tools) to see the moving parts
- Context management deep dive: compaction, sub-agents, file hand-offs
- Sandboxing and permissions deep dive
- How to evaluate a harness change (same model, different harness)
- Read the ReAct paper