This page is the test, not the opinion. Six models get the same tasks. A model wins a task only if the code runs. A fluent wrong answer scores zero.
The models in this run are the current coding-relevant names as of October 6, 2026:
| Lab | Model in the test | Why this one |
|---|---|---|
| OpenAI | GPT-6.1 Sol | Latest listed coding and computer-use upgrade, September 29, 2026. |
| Anthropic | Claude Opus 5.5 | Current high-end Opus, September 22, 2026. |
| Gemini 3.8 Flash | Latest main Flash model, September 2, 2026. | |
| DeepSeek | DeepSeek-V4.1-Flash | Current live DeepSeek model, September 10, 2026. Older Flash names route here. |
| Meta | Llama 4 Maverick | Latest larger Llama in the timeline, April 5, 2025. |
| xAI | Grok-4.7 | Latest Grok for long coding work, September 21, 2026. |
Dates and siblings live on the AI model release timeline. The short buying advice lives on Best AI for coding in October 2026. The older ChatGPT vs Claude page stays up. This test does not replace it until the scores exist.
The rules
Use a fresh chat for every model. Do not let a model see another model's answer. Paste the task and nothing else on the first try. If it fails, send one repair message: "This failed. Here is the error. Fix only the bug." Count that as one correction.
Score each task from 0 to 2.
- 0: it does not run, or it solves a different problem.
- 1: it runs after one correction.
- 2: it runs on the first answer and meets the limits in the task.
Also write down the wall-clock time and the price if the product shows it. Do not guess a token cost.
The six tasks
1. A React bug
Give this component and this symptom: a list renders, the filter input updates state, and the list does not change.
Ask for a fix that keeps the same component structure. The model must not rewrite the file into a new library.
Pass: the filtered list updates on each keypress, and the empty state still shows when nothing matches.
2. A Node API bug
Give a small Express or Next route that returns 500 when the JSON body is missing a field.
Pass: invalid input returns 400 with a clear message. Valid input still returns the old success shape.
3. A refactor with a fence
Give one 200-line file and this limit: "Move the date helpers out. Do not change the public function names."
Pass: the names and return values stay the same, and the helpers live in a second file.
4. A failing test
Give a Jest or Vitest file that fails for a known off-by-one bug.
Pass: the test passes, and the model does not delete the assertion to make it pass.
5. Read a strange file
Give a 400-line file and ask: "Where is the price calculated, and what breaks if quantity is zero?"
Pass: it names the function and the zero-quantity bug without inventing a second bug that is not in the file.
6. SQL
Give a table shape and this question: "Orders last 30 days, by customer, skip customers with no orders."
Pass: the query uses the right join and date filter, and it does not scan an unbounded history.
Score sheet
| Task | GPT-6.1 Sol | Claude Opus 5.5 | Gemini 3.8 Flash | DeepSeek-V4.1-Flash | Llama 4 Maverick | Grok-4.7 |
|---|---|---|---|---|---|---|
| React bug | Not run | Not run | Not run | Not run | Not run | Not run |
| Node API | Not run | Not run | Not run | Not run | Not run | Not run |
| Refactor fence | Not run | Not run | Not run | Not run | Not run | Not run |
| Failing test | Not run | Not run | Not run | Not run | Not run | Not run |
| Read the file | Not run | Not run | Not run | Not run | Not run | Not run |
| SQL | Not run | Not run | Not run | Not run | Not run | Not run |
| Total (max 12) | Not run | Not run | Not run | Not run | Not run | Not run |
Winner: not chosen. Fill this sentence only after the table is real: "On these six tasks, [model] scored [n]/12."
What to paste under each score
For every model, keep four lines:
- First answer: ran, or the exact error.
- Correction used: yes or no.
- What it broke that you did not ask it to touch.
- Time to a working answer.
Screenshots of the first answer and the terminal belong here. A paragraph that only says a model is "strong at coding" does not belong here.
What this test will not prove
Six tasks cannot rank a model for every job. A model can win this sheet and still be worse at a 50-file migration. Publish the limit in the opening paragraph when the scores go in.
Llama 4 Maverick is older than the other five. If it loses, say the date. Do not treat April 2025 and September 2026 as the same generation.