Frontier AI Models Take the Wheel of a Toyota Corolla — and Mostly Crash

A team of researchers decided to find out what happens when general-purpose chatbots get a steering wheel instead of a keyboard. The answer, according to their newly published benchmark, is not encouraging: most of the leading AI agents could not navigate even a simple coning-marked course in a parking lot without ending up in the cones.
The project, dubbed DrivingBench, was assembled by Aditya Ramabadran, Simon Mahns and Tobias Gessler. Their approach was deliberately stripped down. Rather than building custom driving software, they took commercially available vision hardware from the autonomy world — Comma equipment running OpenPilot — and bolted it onto a Toyota Corolla. The test route was a short point-to-point loop through a parking lot with a handful of gentle bends, bordered by small cones. Nothing about the course was meant to be difficult for a competent human driver.
Each participating agent — Claude, GPT and Grok among them — was given identical instructions and full authority over throttle, brake and steering. Commands traveled wirelessly to remote data centers and back, which explains why the car sometimes idled mid-run while waiting for its next instruction. Agents were granted three attempts apiece.
The results were sobering. Across 11 runs involving four agents, eight of them covered less than 11 percent of the route, failing at the first bend. Grok interpreted a gap in the boundary cones as an opening to drive through and exited the course almost immediately. One GPT instance hallucinated a rule about cone colors that the human-written prompt had explicitly forbidden. Many agents simply misjudged how much steering input a curve required, under-rotating and drifting off line.
Only one agent — GPT-6 Astra — completed the full course, and it took a second try. Its path was sloppy, drifting wide through a long right-hander and nearly running out of room before the finish box, but it crossed the line. The run also carried a price tag of roughly $7.74 in compute, about four times the cost of an earlier attempt that only made it halfway through the course.
The study underlines a distinction that autonomy veterans have been making for years: piloting a vehicle is not the same kind of problem as generating plausible text. Driving requires continuous spatial reasoning, calibration and real-time correction — precisely the areas where pattern-matching language models tend to fall apart. Whether token-billed AI agents will ever be a practical path to self-driving remains an open question, but this experiment suggests the gap is still wide.
For now, the Corolla survived, and the researchers have a benchmark that others can build on. The bar it sets is modest: get through a parking lot without hitting the cones.
What do you think?