TECH FLOW Svět Androida
← Back to the stream
lh-harness.pages.dev · picked by Petr Mišák · 47d ago

LongHorizon-Harness Advancing Long-Horizon Agents for Real-World Tasks

Source preview: LongHorizon-Harness Advancing Long-Horizon Agents for Real-World Tasks
AI summary

LongHorizon-Harness is a system for long-horizon AI agents that addresses failure modes in extended task sequences by separating state management, execution, and audit into three independent roles. Instead of a single growing context that judges itself, it employs a verified Manage-Execute-Audit loop where an auditor independently verifies results against the environment and updates status only from independently verified facts.

The summary is written by AI from the source; it isn’t the newsroom’s opinion. For details, read the source.

5 people have already opened the source

Tip author’s note

When you assign a task to an AI, it often fails to complete it satisfactorily for various reasons—from AI alibism, where models trust their own claims instead of critically verifying them, to information dilution across large contexts where important details get lost and lose significance. Harness approaches these issues generally, and this one from Alibaba Dream X Team may help solve them.

AI questions & answers
What are the main failure modes of existing AI agents on long-horizon tasks?

Current systems suffer from error compounding (an early mistake distorts every later choice), context rot (as history grows, relevant information becomes harder to retrieve and performance degrades sharply), and task-state loss (no accurate record of what has been completed, produced, or exists in the environment).

How does LongHorizon-Harness prevent unverified claims from propagating?

The system separates the executor from the auditor and manager roles. Only the auditor, which observes the environment and never sees the executor's reasoning, can update the state record. An executor's claim becomes a recorded fact only after the auditor independently verifies it from the environment.

Which AI models and agent backends does LongHorizon-Harness support?

The system is backend-agnostic and allows combining different models (Claude Opus, GPT, Qwen, Gemini, Kimi, MiniMax) with different agents (Claude Code, Codex CLI, Gemini CLI, mini-SWE-agent) independently for each of the three roles.

What performance improvements were observed on standard benchmarks?

WeaveBench PassRate improved from 51.8% to 80.7% (+28.9%), Terminal-Bench 2.1 success rate increased from 69.7% to 77.2% (+7.5%), and OSWorld 2.0 binary completion tripled from 2.8% to 8.3%.

Questions and answers are written by AI about the topic, not taken from the source; they aren’t the newsroom’s opinion.
Related from the stream
Mentions
  • LongHorizon-Harness
  • Alibaba Dream X Team
  • Claude Opus
  • GPT-5.5
  • Qwen 3.7-Plus
  • Gemini 3.1
  • Kimi 2.6
  • MiniMax M3
  • Claude Code8
  • Codex CLI