Author test
jev-agent-failure-benchmark
Attribute failures in multi-agent traces
How this project uses Jev
- Input
- Failure traces and candidate agents, steps, and error types
- Jev decides
- Responsible agent, key step, and error category
- Code executes
- The official scorer evaluates predictions
Evidence and limitations
The author uses an injected-error dataset. Some comparisons come from papers; constrained Who/When choices are not equivalent to free generation.
This project has not been run independently here. Author-reported results are not independently verified results.