# Epoch AI reports rise from 28% to 80% in furniture assembly test

- Published: 2026-09-28T13:16:26.000Z
- Updated: 2026-09-28T14:25:47.100Z
- Section: Models
- Topics: Research
- Page: https://aidailyjournal.com/en/n/2026/09/28/epoch-ai-reports-rise-from-28-to-80-in-furniture-assembly-test/
- Sources: [Xataka](https://xataka.com/robotica-e-ia/tenemos-nuevo-benchmark-para-ia-preguntarle-has-montado-correctamente-tus-muebles-ikea-no)

## Summary

Epoch AI says performance on the Furniture Assembly Benchmark (FAB) rose from 28%, achieved by Claude 4,5 in November 2025, to 80%, achieved by GPT-6 Astra in September 2026. Models receive a photo of partially assembled IKEA furniture, the instruction manual, an image-zooming tool and a Python interpreter; they must match the diagrams to real parts, reconstruct the assembly steps and locate errors such as an inverted piece. According to Epoch AI, GPT-6 Astra took an average of three minutes per photo, two to ten times less than models that had previously ranked among the best. The benchmark measures visual and spatial reasoning, not whether robots can physically assemble furniture. Epoch also found biases: Gemini and Qwen tended to identify nonexistent errors, while earlier OpenAI and Anthropic models approved faulty assemblies.

## Why it matters (editorial interpretation)

The result suggests meaningful gains in visual and spatial reasoning, but it does not establish that models can physically assemble furniture. The diagnostic biases identified by Epoch AI also suggest that evaluations need to account for both incorrectly flagged errors and failures to detect faulty assemblies.

_Summary written by AI Daily Journal in its own words, based on the sources above. For details, read the original articles._
