Research update: Long-horizon reliability, can the system stay coherent over extended periods?
Ordinary AI Labs reports sustained performance gains in long-horizon agent evaluation
The question we want to answer is: Long-horizon reliability, can the system stay coherent over extended periods?
Ordinary AI Labs today published internal results from its long-horizon evaluation programme, covering agent performance on multi-step tasks carried out over extended periods without human intervention.
The evaluation suite measures how reliably a system maintains task coherence, recovers from errors, and requests clarification when instructions are incomplete. Across the current evaluation set, the company’s production agent family completed 94.2% of assigned multi-step tasks within defined operating parameters.
“Long-horizon reliability is the difference between a demonstration and a tool,” said Mara Voss, Senior Alignment and Systems Architect. “A system that performs well for ninety seconds tells you very little. We are interested in what happens on day four.”
Full methodology will be published in the company’s next technical report.