We hosted a live research talk presenting an early preview of MultiNet v2.0, our cross-domain, multimodal benchmark for evaluating long-horizon agents.
We examined how frontier reasoning and vision-language models performed in controlled interactive 2D maze environments designed to isolate failures in planning, action execution, error recovery, visual association, and causal reasoning. The results showed that even simple environments exposed significant weaknesses in long-horizon behavior, with different models failing in distinct ways.
The talk was presented by Sean Rivera, an Open Source Research Scientist at Manifold Research, and concluded with an open Q&A and discussion.
Interested in working with us? Check out our open opportunities here: