We evaluated 3 highly capable Vision-Language Models (VLMs) on 50 2D mazes. Across the 3 models, only 6 total mazes were solved, and 45 of the 50 mazes were not solved by any model.
The size of the mazes ranges from 8x8 to 14x14 grids. An agent has to navigate through walls, avoid decoys and distractors, and operate objects that are to be operated in a specific order (a key that opens a door, a switch that toggles open a gate), in order to reach a goal tile and solve the maze. The action space consists of six actions: turn left, turn right, move forward, pickup, toggle, and done. Apart from the action space and the instruction to complete the maze by reaching the goal tile, the models are not given any other information about the environment. They have to figure out what to do by exploring the environment, acting in it, and observing the changes in the environment.
When a benchmark like the one we present records a zero for a model failing an episode, there are multiple possible reasons for the failure:
- The model never worked out what an object in the environment does.
- It worked out what the objects are, but could not operate them in the right order.
- It executed the wrong actions from the action space.
- It made mistakes and could not recover from them to get back on the right path.
All of these are completely different failure modes, and each of them requires a very different fix. Going down the wrong path in a scenario such as this is expensive in both time and resources spent, and the cost of not identifying the right failure mode is high: a survey of agentic benchmark construction found flaws that mis-estimate performance by up to 100%.
Separating the modes of failure means changing the task in different ways: running a maze without the switch-gate mechanism, running it with the optimal path length controlled, and so on. Whichever change moves the score helps narrow down on the failure mode of the model or agent.