AI Agent Harness: A Builder Checklist for Context Compaction and Memory Boundaries

Share





FuturPulse analysis · 7 October 2026

On 6 October 2026, a TRACE study of 590 compaction boundaries found that history predicts compaction harm only weakly, so builders should set explicit memory boundaries rather than trust a clever summariser.

At a glance

  • The TRACE paper’s best interpretable trigger avoided 21% of harmful boundaries while retaining 84% of compaction opportunities.
  • Microsoft Agent Framework shows a harness configuration with a 128,000-token context limit and a 16,384-token output limit.
  • agent-browser caps automatic WebMCP tool summaries at 16 tools and 4 KiB of JSON before the agent requests detail.
  • Leviathan reports a median 436 tokens per indexed-record query on its synthetic maintenance-log benchmark.

What does harness mean in AI agents?

An AI agent harness is the runtime around a language model. It chooses what context reaches the model, runs approved tools, records results, stores work, and stops or asks for help when a task crosses a boundary.

Databricks defines the harness as the infrastructure that connects a model to tools, memory, execution environments and guardrails. Microsoft’s implementation description is more operational: it combines the chat pipeline, context providers, approvals, observability and optional bounded loops into one working agent.

The distinction matters during debugging. A model can generate a plausible next step. The harness decides whether that step can read a file, call an API, run a shell command, write durable memory, or send a message. That is where reliability and security become engineering choices.

The core loop comes from ReAct, submitted on 6 October 2022: the model alternates between reasoning and actions, while tool results return as new context. A modern harness adds budgets, persistence and permission checks around that loop.

What should builders set before the first tool call?

Set boundaries before adding more tools. The useful question is not “what can the agent access?” It is “what can it access in this task, through this session, with this approval level?”

  • Define the task contract. State the user goal, allowed systems, prohibited actions, completion test and escalation path.
  • Separate instructions by authority. Keep platform rules, product policy, user request and untrusted web or file content distinct.
  • Give each tool a narrow purpose. Describe inputs, expected outputs, side effects and the person or system that approves consequences.
  • Persist evidence, not every utterance. Save source links, commands, test results, decisions and open questions outside the chat transcript.
  • Put a limit on the loop. Define a maximum number of tool attempts, repeated-error threshold and stop condition.
  • Record the state needed to resume. A new session should recover a task plan and verified findings without replaying every message.

Microsoft Agent Framework’s harness makes this separation visible. It persists history after model calls, enables compaction when limits or a strategy are supplied, and treats file access, shell tooling, background agents and looping as configurable capabilities. That is the right mental model: components are not “features” until their authority is defined.

Builder rule: treat every tool result as untrusted input until its source, scope and effect are checked. A web page can describe an action without being authorized to request it.

Where should memory stop and context begin?

Context is temporary working material that the model reads in the current turn. Memory is durable information that the harness stores for later retrieval. Mixing them makes compaction dangerous, because deleting chat text can silently delete a decision, permission or fact that the next action needs.

Information typeWhere it belongsWhat compaction may doBoundary to enforce
User’s immediate request and current planActive contextSummarise only after preserving the goal and unfinished work.Never replace an unresolved approval with a summary.
Tool output needed for the next actionActive context plus task logKeep the exact result or a stable reference.Do not summarise away error codes, file paths or identifiers.
Verified user preferences and durable project factsLong-term memoryRetrieve by relevance rather than retaining in every turn.Attach source, owner, date and revision rule.
Credentials, private keys and access tokensSecret manager, not agent memoryExclude from summaries and durable notes.Return only scoped, temporary access where possible.

This boundary map is a practical design comparison: it separates temporary reasoning from durable evidence, rather than treating “memory” as one undifferentiated store.

Agent Memory Repo’s published format offers a useful durable-memory pattern. Its short MEMORY.md file acts as an entry point, while linked notes hold detail and metadata can retain source and added date. The important design choice is not the file format. It is that a session loads a compact index first, then follows links only when needed.

For broad operational history, retrieval should return evidence cards rather than raw logs. Leviathan’s benchmark reports 436 median tokens per query across 1 million records, versus 107,122 for its listed grep strategy. That result is from a synthetic maintenance log, not a universal production claim, but it illustrates the harness principle: retrieve the few records that answer the question, then keep their citations.

How should context compaction work?

Compaction should be a controlled state transition, not a blind instruction to “summarise the conversation.” Before removal, extract the task state into named fields and write durable evidence to a store that survives the next turn.

  1. Freeze the active task. Capture the goal, current subtask, constraints, pending approval and completion test.
  2. Write exact references. Keep tool call IDs, paths, URLs, source snippets and error messages outside the free-form summary.
  3. Classify each item. Mark it as active context, durable memory, secret, disposable trace or untrusted content.
  4. Produce a short handoff. Include what changed, what was verified, what failed and the one next action.
  5. Validate after reload. Ask the resumed agent to locate its task state and cite the evidence before it acts.

The new TRACE result is a warning against overfitting that decision to recent chat behaviour. Its best held-out trigger reached an AUROC of 0.66, while the authors say the release cannot establish whether it beats a token-budget rule at matched retention. Use a token budget as a predictable safety valve, then evaluate compaction against your own repeated-call and tool-error logs.

agent-browser’s WebMCP design shows one practical form of progressive disclosure. It first announces compact tool names, descriptions, origins and frame IDs, then requires a separate request for a full schema. The automatic announcement stops at 16 tools and 4 KiB. After compaction, the agent must explicitly recover the catalog.

That pattern is valuable beyond browsers. Keep a compact index in context. Fetch the full contract only for the one tool, file or memory entry the agent has selected. Refresh the item when its catalog changes. Never let a previous summary stand in for a changed permission or schema.

AI Agent Harness: A Builder Checklist for Context Compaction and Memory Boundaries
Web Developer · photo libre de droits

What was promised versus what shipped?

The timeline shows how the idea moved from an action loop to runtime scaffolding and then to a measurable compaction problem. The gap is clear: production frameworks ship mechanisms, but no source here establishes a generally reliable predictor for harmful compaction.

Promise: ReAct proposed interleaving reasoning and actions, so language models could update plans from external observations. What shipped: a research method and code-linked paper, not a complete memory, approval or persistence layer.

Promise: Microsoft documented a batteries-included harness for long-running work. What shipped: a composition of history persistence, optional compaction, planning, memory, tool approvals, telemetry and bounded looping.

Promise tested: TRACE tested whether recent history could predict harmful compaction. What shipped: a public the public record of 590 replayed boundaries and a bounded result, not proof that adaptive compaction beats token-budget compaction.

What are the best AI harnesses?

The best AI harness is the smallest one that enforces your task boundaries and can show its work. There is no single winner because a coding agent, a browser agent and a customer-support agent need different tools, storage and approval rules.

For a production application that needs planning, persistence, approvals and telemetry in one framework, Microsoft Agent Framework is a credible starting point. For browser work, agent-browser provides compact text output, ref-based page interaction and persistent browser sessions. For large record collections, an indexed retriever such as Leviathan can keep raw history out of the prompt.

Choose by failure mode. If the agent forgets decisions, add durable evidence with ownership and source fields. If it repeats failed tools, retain exact errors and loop counters. If it takes unauthorized actions, reduce tool scope and require confirmation. Adding another model will not repair an unclear boundary.

Should I build my own AI harness?

Build your own harness when your authority model, audit needs or workflow state are the product. Start from a framework when you mainly need standard tool calling, session management and observability.

A custom layer is justified when the agent must preserve a domain-specific task ledger, operate inside a strict approval process, or retrieve evidence from systems that cannot expose raw records. It is also justified when you need reproducible evaluation: the same task, tool permissions and memory state should replay after a failure.

Do not begin by writing a grand autonomous runtime. Begin with one task, one read-only tool, one durable log and one stop condition. Then test a forced compaction in the middle of the task. If the restarted agent cannot name the goal, locate its evidence and explain its next permitted action, the memory boundary is not ready.

What we could not verify?

No public source here establishes a universal compaction threshold, a standard schema for durable agent memory, or a benchmark proving that one harness is best across coding, browser and enterprise workflows. The TRACE authors explicitly say their release cannot test whether their trigger beats a token-budget rule at matched retention.

Framework maintainers could settle implementation details through versioned documentation and release notes. Builders can settle the operational question by publishing replayable task traces, compaction points, tool-error rates, repeated-call rates and the exact memory state available after each restart.

Sources



Maya Chen
Maya Chen
Maya Chen covers AI agents, orchestration frameworks, tool-use, and evaluation. She focuses on what actually works in production—failure modes, safety boundaries, and measurable performance—without the hype.

Read more

Local News