OpenAI has posted a new leaderboard on drivingbench.com, and the numbers show that GPT-6 Astra has gained the ability to drive a real car. The test is simple: models are given control of a Toyota Corolla’s steering, accelerator, and brakes, then asked to complete a fixed cone course. GPT-6 Astra is the only model to reach 100% progress on its best run.
The leaderboard lets each model try up to three times in one continuous chat. Each run is scored on distance traveled, speed, finish time, and commands issued, with costs tracked at list prices. The results show a clear gap between GPT-6 Astra and the rest of the field.
The Numbers Behind the Course
GPT-6 Astra’s best run covered the full course at 100% progress, finishing in 5:22. Its other two attempts also reached 100% progress. The leaderboard displays finish times only for completed runs, so individual times for those two runs are not shown.
Claude Fable 5.1 came closest among the runners-up, reaching 45% progress on its best run. Its remaining two attempts managed only 9% and 10% progress. Claude Fable 5.1 never finished the course, ending each run before reaching the finish zone.
Grok 4.6 performed worse, reaching 11% progress on its best run and 8% on its second. Its third attempt reached only 10% progress. Like Claude Fable 5.1, Grok 4.6 never completed the course.
GPT-5.6 Sol fared worst of all, topping out at 6% progress across all three attempts.
What the Leaderboard Shows
The leaderboard ranks models by their best single attempt, not by average performance across all three runs. That means a model can score highly even if its other two attempts fall apart.
| Model | Best progress | Finish time |
|---|---|---|
| GPT-6 Astra | 100% | 5:22 |
| Claude Fable 5.1 | 45% | DNF |
| Grok 4.6 | 11% | DNF |
| GPT-5.6 Sol | 6% | DNF |
DNF stands for “did not finish,” meaning the model ended its run before reaching the finish zone. The table shows that GPT-6 Astra is the only model to reach 100% progress.
How the Test Works
The course itself is a fixed cone path, and the scoring system rewards staying close to the centerline. A model earns points for distance traveled while staying within 4 meters of the centerline, as a share of the centerline length to the finish zone. A collision stops the attempt immediately, with the score locked at the progress reached before impact.
Speed is integrated from the first accepted motion command to the end of the attempt’s last engagement. Finish time is measured from the first accepted motion command to the end of the attempt’s last engagement, shown only for completed runs.
Commands are tracked as accepted motion commands and stop commands during the attempt. Costs are totaled at list prices, including the reflection after each attempt.
The Performance Gap
The leaderboard shows a clear separation between the top performer and everyone else. Here is how the models rank by best progress:
- GPT-6 Astra — 100%
- Claude Fable 5.1 — 45%
- Grok 4.6 — 11%
- GPT-5.6 Sol — 6%
GPT-6 Astra’s 100% progress is a full course completion on its best run. Claude Fable 5.1 reached 45% on its best run. Grok 4.6 topped out at 11%, and GPT-5.6 Sol never got past 6%.
What This Means for AI
A car is a complex machine with many moving parts, and driving it requires constant attention to speed, direction, and obstacles. A model that can finish a cone course without crashing has demonstrated real skill, not just raw power.
The leaderboard also shows that the gap between top performers and the rest of the field is significant. GPT-6 Astra finished at 100% progress, while the next-best model reached only 45%. That is a large margin in a test that measures actual performance.
The results suggest that driving ability is not evenly distributed across AI models. Some models can handle the task, while others struggle badly. The leaderboard does not say which models failed for technical reasons and which failed for lack of skill, but the pattern is consistent: GPT-6 Astra succeeds, and the rest fail.
The Road Ahead
The leaderboard is a snapshot of current performance, not a final verdict. Models improve over time, and newer versions may close the gap or widen it.
For now, the takeaway is simple: GPT-6 Astra has gained the ability to drive a real car, and the test on drivingbench.com proves it.
The leaderboard is available for anyone to inspect, and the video traces show exactly how each model approached the course. GPT-6 Astra’s runs are the ones worth watching, since it is the only model to reach 100% progress.
The comparison is stark, and the conclusion is clear: GPT-6 Astra is the only model on this leaderboard that reached 100% progress.
Source material: “GPT-6 Astra has gained the ability to drive a car,” drivingbench.com.
Get the Notebook.
The day's best stories and every fresh verdict, in plain English, in your inbox by seven. One email a day, no more.

