Research blog
Blogs
February 2026 · ADRS series
Automating Algorithm Discovery: A Case Study in Improving Multi-Agent Reasoning Systems using MAST (Part 2)
The evaluator is the bottleneck: with MAST-based feedback instead
of binary reward, evolutionary search finds higher-performing
multi-agent reasoning systems on Olympiad-level math.
December 2025 · with IBM Research
Why Do Enterprise Agents Fail? Insights from IT-Bench using MAST
310 SRE traces, three models, one lesson: success rates hide the
failure signature. MAST separates fatal failures from benign
friction and turns IT-Bench scores into an engineering roadmap.
December 2025 · ADRS series
Automating Algorithm Discovery: A Case Study in Improving Multi-Agent System Design using MAST
Replacing hand-tuning with OpenEvolve: MAST failure modes as the
fine-grained signal that lets evolution mutate a multi-agent
architecture toward a 7x better failure rate.