Adaptive failure taxonomies

AdaMAST

Agents already generate the evidence of their own failures on every run. AdaMAST converts those traces into a compact taxonomy of named failure codes, induced automatically, validated before anything trusts them, and reused as feedback by search, runtime monitoring, and trajectory selection.

01 Observe 02 Correlate 03 Map 04 Decide

Papers

ICML 2026 FAGEN Workshop · Best Paper

Fantastic Adaptive Taxonomies and How to Use Them

Mert Cemri*, Andrei Cojocaru*, Melissa Pan, Shu Liu, Shubham Agarwal, Alexander Krentsel, Jay Tang, Kannan Ramchandran, Joseph E. Gonzalez, Matei Zaharia, Alexandros G. Dimakis, Ion Stoica

NeurIPS 2025 · Spotlight

Why Do Multi-Agent LLM Systems Fail?

Mert Cemri*, Melissa Z. Pan*, Shuyi Yang*, Lakshya A. Agrawal, Bhavya Chopra, Rishabh Tiwari, Kurt Keutzer, Aditya Parameswaran, Dan Klein, Kannan Ramchandran, Joseph E. Gonzalez, Matei Zaharia, Ion Stoica

Blog

Jul 2026 AdaMAST: Adaptive Failure Taxonomies for Improving LLM Agents Feb 2026 Automating Algorithm Discovery: A Case Study in Improving Multi-Agent Reasoning Systems using MAST (Part 2) Dec 2025 Why Do Enterprise Agents Fail? Insights from IT-Bench using MAST Dec 2025 Automating Algorithm Discovery: A Case Study in Improving Multi-Agent System Design using MAST Nov 2025 How to Use MAST to Improve Your Agents

All posts →