September 28, 2026 Understand what changes. Find what to use.

ModelsNews1 source

Epoch AI reports rise from 28% to 80% in furniture assembly test

Epoch AI says performance on the Furniture Assembly Benchmark (FAB) rose from 28%, achieved by Claude 4,5 in November 2025, to 80%, achieved by GPT-6 Astra in September 2026. Models receive a photo of partially assembled IKEA furniture, the instruction manual, an image-zooming tool and a Python interpreter; they must match the diagrams to real parts, reconstruct the assembly steps and locate errors such as an inverted piece. According to Epoch AI, GPT-6 Astra took an average of three minutes per photo, two to ten times less than models that had previously ranked among the best. The benchmark measures visual and spatial reasoning, not whether robots can physically assemble furniture. Epoch also found biases: Gemini and Qwen tended to identify nonexistent errors, while earlier OpenAI and Anthropic models approved faulty assemblies.

Why it matters · editorial interpretation

The result suggests meaningful gains in visual and spatial reasoning, but it does not establish that models can physically assemble furniture. The diagnostic biases identified by Epoch AI also suggest that evaluations need to account for both incorrectly flagged errors and failures to detect faulty assemblies.

Sources

Research

One story, many sources

Coverage of the same story gathered in one item, with a link to each original.

Summary and interpretation apart

What the sources say stays in the summary; editorial context is labeled separately.

Made with AI, with sources

Summaries and translations generated with AI from the original stories, always linked.