publication

Goal Progress Decays with Task Horizon for Frontier VLMs in Interactive 2D Environments

Pranav Guruprasad, Sean Rivera, Helen Lu, Arushi Jain, Hangliang Ren, Harshvardhan Sikka (2026)
Goal Progress Decays with Task Horizon for Frontier VLMs in Interactive 2D Environments
In: publication

We evaluated 3 highly capable Vision-Language Models (VLMs) on 50 2D mazes. Across the 3 models, only 6 total mazes were solved, and 45 of the 50 mazes were not solved by any model.

The size of the mazes ranges from 8x8 to 14x14 grids. An agent has to navigate through walls, avoid decoys and distractors, and operate objects that are to be operated in a specific order (a key that opens a door, a switch that toggles open a gate), in order to reach a goal tile and solve the maze. The action space consists of six actions: turn left, turn right, move forward, pickup, toggle, and done. Apart from the action space and the instruction to complete the maze by reaching the goal tile, the models are not given any other information about the environment. They have to figure out what to do by exploring the environment, acting in it, and observing the changes in the environment.

When a benchmark like the one we present records a zero for a model failing an episode, there are multiple possible reasons for the failure:

  • The model never worked out what an object in the environment does.
  • It worked out what the objects are, but could not operate them in the right order.
  • It executed the wrong actions from the action space.
  • It made mistakes and could not recover from them to get back on the right path.

All of these are completely different failure modes, and each of them requires a very different fix. Going down the wrong path in a scenario such as this is expensive in both time and resources spent, and the cost of not identifying the right failure mode is high: a survey of agentic benchmark construction found flaws that mis-estimate performance by up to 100%.

Separating the modes of failure means changing the task in different ways: running a maze without the switch-gate mechanism, running it with the optimal path length controlled, and so on. Whichever change moves the score helps narrow down on the failure mode of the model or agent.

Read the full report here!

More from Manifold Research
Great! You’ve successfully signed up.
Welcome back! You've successfully signed in.
You've successfully subscribed to Manifold Research.
Your link has expired.
Success! Check your email for magic link to sign-in.
Success! Your billing info has been updated.
Your billing was not updated.