---
title: "An Empirical Study of Harness Design for Coding Agents"
slug: an-empirical-study-of-harness-design-for-coding-agents
url: https://listedarticles.com/articles/an-empirical-study-of-harness-design-for-coding-agents
canonical_url: https://arxiv.org/abs/2609.20804
content_type: research
language: en
published_at: 2026-09-17T17:58:07.000Z
updated_at: 2026-09-18T15:42:16.359Z
author: "Run-Ze Fan, Zihao Zhang, Simin Ma, Yebowen Hu, Shouju Wang, Kaiqiang Song, Fei Liu, Hamed Zamani, Xiaoyang Wang"
author_url: https://arxiv.org/abs/2609.20804
authored_by: human
publisher: "arXiv"
publisher_url: https://arxiv.org/
topics: ["AI Agents", "AI", "Research", "Machine Learning", "Programming"]
license: all-rights-reserved
word_count: 291
reading_minutes: 1
citation: "Run-Ze Fan, Zihao Zhang, Simin Ma, Yebowen Hu, Shouju Wang, Kaiqiang Song, Fei Liu, Hamed Zamani, Xiaoyang Wang, arXiv. \"An Empirical Study of Harness Design for Coding Agents.\" 17 Sept 2026. https://arxiv.org/abs/2609.20804 (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# An Empirical Study of Harness Design for Coding Agents

> Fan et al. ablate planning, action space, and context management in a fixed coding-agent loop across 176 SWE-Bench/Terminal-Bench settings, finding when context management, planning, and predefined tools help—and when bash-only is enough.

# An Empirical Study of Harness Design for Coding Agents

**Authors:** Run-Ze Fan, Zihao Zhang, Simin Ma, Yebowen Hu, Shouju Wang, Kaiqiang Song, Fei Liu, Hamed Zamani, Xiaoyang Wang

**arXiv:** [2609.20804](https://arxiv.org/abs/2609.20804) · [PDF](https://arxiv.org/pdf/2609.20804)

## Abstract

Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic systems, leaving the effectiveness of individual components unclear. To enable component-level comparisons, we study this question with a lightweight coding harness whose execution loop is fixed while three components are varied: planning, action space, and context management. Across four models evaluated on SWE-Bench Verified and Terminal-Bench 2.1, we evaluate 176 matched settings spanning five context-management strategies, four context-window budgets, and targeted ablations of planning and action space. We find that: (1) Context management becomes increasingly valuable as the context-window budget tightens, with most of its benefit coming from preventing context-overflow failures. (2) Staging rule-based elision before LLM-based summarization provides the strongest overall efficiency among the context-management strategies, whereas making elided content recoverable adds machinery that models rarely use and yields no accuracy gain. (3) Planning shifts from an accuracy scaffold for weaker models to a cost saver for stronger models, with little change in accuracy. (4) Predefined tools improve performance for models with weaker bash proficiency, whereas bash-capable models can operate effectively with a bash-only interface and achieve substantially lower cost, especially on command-line-centric tasks. Trajectory-level analysis explains these effects: context management extends execution trajectories without substantially altering agent behavior, planning changes where trajectories stop, and the action space changes the granularity at which code is written. These findings inform model- and budget-aware harness design and provide a modular framework for evaluating future harness components.

## Categories

cs.AI, cs.CL, cs.LG, cs.SE

*Full paper (43 pages) available on arXiv.*
