events

Frontiers: Why Frontier Reasoning Models fail on Interactive 2D Mazes

9-11-27
Frontiers: Why Frontier Reasoning Models fail on Interactive 2D Mazes
In: events

We hosted a live research talk presenting an early preview of MultiNet v2.0, our cross-domain, multimodal benchmark for evaluating long-horizon agents.

We examined how frontier reasoning and vision-language models performed in controlled interactive 2D maze environments designed to isolate failures in planning, action execution, error recovery, visual association, and causal reasoning. The results showed that even simple environments exposed significant weaknesses in long-horizon behavior, with different models failing in distinct ways.

The talk was presented by Sean Rivera, an Open Source Research Scientist at Manifold Research, and concluded with an open Q&A and discussion.

Interested in working with us? Check out our open opportunities here:

https://www.manifoldrg.com/opportunities/

More from Manifold Research
Great! You’ve successfully signed up.
Welcome back! You've successfully signed in.
You've successfully subscribed to Manifold Research.
Your link has expired.
Success! Check your email for magic link to sign-in.
Success! Your billing info has been updated.
Your billing was not updated.