Skip to content

Harness Failure Modes & Future

Key idea

Most production agent failures come from the harness, not the model.


Common failure modes

Failure Symptom Fix
Context rot Gets worse the longer it runs Compact, split into phases
Tool overload Picks the wrong tool, slow to start Fewer, general tools; load on demand
Brittle tool wiring Silent wrong calls after a small change Clear descriptions, eval tool changes
Latency 10s+ per response Parallel calls, smaller model for sub-steps
Irrelevant retrieval Confident but wrong answer Better retrieval, cite sources
Weak verification "Done!" but it's not Tests/checks in the loop
Missing guardrails Irreversible action without approval Permission gates, human-in-the-loop

(The Fix column is my own notes, not from the article.)

Context rot

Reasoning quality drops as history grows.

xychart-beta
    title "Answer quality vs context used"
    x-axis "Context used %" [10, 30, 50, 70, 90]
    y-axis "Quality" 0 --> 100
    line [95, 92, 85, 65, 40]

Example

After 80 turns of debugging, the agent forgets the constraint "no allocations on the hot path" from turn 3 and adds new ArrayList<>() inside onFragment().

Fix: compact, or write findings to NOTES.md and start a fresh session.

Tool overload

Too many tools → confusion and slow decisions.

Example

60 MCP tools loaded: search_jira, search_confluence, search_slack, search_docs… For "find the design doc", the model tries 4 search tools in a row.

Fix: load only the tools the current task needs (skills / on-demand MCP).

Brittle tool wiring

A small change in a tool description or schema → wrong usage, silent failures.

Example

You rename the param date → trade_date. The description still says "date". The model keeps sending date, the tool ignores it and returns today's trades. No error is raised.

Latency

Many sequential tool calls → slow.

gantt
    title Sequential vs parallel tool calls
    dateFormat s
    axisFormat %Ss
    section Sequential
    read A :0, 2
    read B :2, 2
    read C :4, 2
    section Parallel
    read A :0, 2
    read B :0, 2
    read C :0, 2

Irrelevant retrieval

The harness fetches the wrong docs → the model answers confidently from them.

Example

"What's our FIX session timeout?" → retrieval returns the UAT config instead of prod → the agent answers "30s", but prod is 60s.

Weak verification

No tests or checks in the loop → the agent stops early or claims success.

Example

"Refactor done ✅", but it never compiled. A PostToolUse hook running mvn compile would have caught it.

Missing guardrails

Irreversible actions without oversight: sending messages, deleting data, buying things.

Example

The agent "cleans up test data" with DELETE FROM trades WHERE ... on the wrong DB connection. An ask rule on DELETE would have paused it for approval.


Enterprise: agent sprawl

Companies build dozens of agents across teams. Without shared infrastructure, nobody can govern, evaluate, or improve them all.

flowchart TB
    subgraph Sprawl["❌ Agent sprawl"]
        direction LR
        A1[Team A agent<br/>own tools, own logs] 
        A2[Team B agent<br/>own tools, no evals]
        A3[Team C agent<br/>no guardrails]
    end
    subgraph Shared["✅ Shared harness (control plane)"]
        direction TB
        CP["Governance · Evals · Observability · Model routing"]
        B1[Team A agent] --> CP
        B2[Team B agent] --> CP
        B3[Team C agent] --> CP
        CP --> Models[(OpenAI / Anthropic /<br/>Google / open-source)]
    end
    Sprawl ==>|consolidate| Shared

A shared layer provides:

  • Access control: which data and actions each agent can use
  • Evaluation: measure every agent the same way
  • Audit + observability: one place to look
  • Model swapping: change the provider without rebuilding the agent

Databricks Agent Bricks

Governance through Unity Catalog, observability/evals through MLflow, and it works with models from many providers.


Where it's heading

flowchart LR
    subgraph Now["Today: harness does a lot"]
        H1[Stay on task]
        H2[Verify work]
        H3[Recover from errors]
        H4[Sandbox / tools / guardrails / logs]
    end
    subgraph Future["Later"]
        M["🧠 Model absorbs:<br/>stay on task, self-verify,<br/>recover"]
        H["🦾 Harness keeps:<br/>sandbox, tools, guardrails,<br/>observability"]
    end
    H1 & H2 & H3 --> M
    H4 --> H

The harness won't disappear. Execution, tools, guardrails, and observability still decide how reliable an agent is.

Two emerging ideas

Idea What Example
Disposable harness Lightweight, task-specific, thrown away after one workflow Spin up a container with 2 tools to migrate one repo, then delete it
Natural-language agent harness (NLAH) Describe agent behavior in plain language; a shared runtime runs it A SKILL.md saying "when asked to release: bump version, run tests, tag, ask before push"

Quote

The model contains the intelligence. The harness turns that intelligence into reliable work.


Back: Overview · Building Blocks