LongHorizon-Harness Advancing Long-Horizon Agents for Real-World Tasks
LongHorizon-Harness is a system for long-horizon AI agents that addresses failure modes in extended task sequences by separating state management, execution, and audit into three independent roles. Instead of a single growing context that judges itself, it employs a verified Manage-Execute-Audit loop where an auditor independently verifies results against the environment and updates status only from independently verified facts.
When you assign a task to an AI, it often fails to complete it satisfactorily for various reasons—from AI alibism, where models trust their own claims instead of critically verifying them, to information dilution across large contexts where important details get lost and lose significance. Harness approaches these issues generally, and this one from Alibaba Dream X Team may help solve them.
What are the main failure modes of existing AI agents on long-horizon tasks?
Current systems suffer from error compounding (an early mistake distorts every later choice), context rot (as history grows, relevant information becomes harder to retrieve and performance degrades sharply), and task-state loss (no accurate record of what has been completed, produced, or exists in the environment).
How does LongHorizon-Harness prevent unverified claims from propagating?
The system separates the executor from the auditor and manager roles. Only the auditor, which observes the environment and never sees the executor's reasoning, can update the state record. An executor's claim becomes a recorded fact only after the auditor independently verifies it from the environment.
Which AI models and agent backends does LongHorizon-Harness support?
The system is backend-agnostic and allows combining different models (Claude Opus, GPT, Qwen, Gemini, Kimi, MiniMax) with different agents (Claude Code, Codex CLI, Gemini CLI, mini-SWE-agent) independently for each of the three roles.
What performance improvements were observed on standard benchmarks?
WeaveBench PassRate improved from 51.8% to 80.7% (+28.9%), Terminal-Bench 2.1 success rate increased from 69.7% to 77.2% (+7.5%), and OSWorld 2.0 binary completion tripled from 2.8% to 8.3%.
- Harness Engineering for Self-Improvement — lilianweng.github.io 80 % match
- Understanding is the new bottleneck — geoffreylitt.com 77 % match
- LFM2.5-2.6B: small and capable local AI model — liquid.ai 75 % match
- LongHorizon-Harness
- Alibaba Dream X Team
- Claude Opus
- GPT-5.5
- Qwen 3.7-Plus
- Gemini 3.1
- Kimi 2.6
- MiniMax M3
- Claude Code8
- Codex CLI