Benchmark

Can AI fix a real design failure?

Eight tasks built from verified records, graded by re-running the physics. No language model judges anything.

Sec_001 · The tasks

Eight tasks, one headline

TaskThe AI mustGraded by
1 · Design repair ★Fix a failing design so it passesRe-run simulation vs. same limit
2 · Failure diagnosisSay what fails, where, by how muchSpec limit and change log
3 · Change predictionPredict the result after a changeReal after-fix results
4 · CAD editBuild the new geometry from a changeGeometry diff vs. the real fix
5 · Simulation auditing ★Find setup mistakes in a recordReal bugs our reviews found
6 · Trade-off reasoningPick the best option and say whyRejected options in change log
7 · Multi-step iterationSimulate, revise, retry until it passesConverged? Steps used?
8 · Production optimisationMake it cheaper to produce at volume and still passSimulation + cost and cycle time saved

★ Headline: design repair, run as a multi-step agent task. Simulation auditing is the hardest, because the mistakes are the real ones our reviews caught.

Sec_002 · The grader

The simulation is re-run. The limit is the same.

The harness packages each record into a task with a hidden answer. The AI returns a design. The physics grader meshes it, runs the same solver with the same loads and boundary conditions, and checks the same spec limit. Open solvers only: CalculiX for structural and thermal, OpenFOAM for flow.

A baseline agent that reads CAD, runs the simulation, edits geometry and retries sets the bar, so you can see how far a raw model is from a working loop.

Fig_005 [ Status · in build ]
HarnessTask per record · hidden answer · any model
GraderCalculiX · OpenFOAM · same limit
BaselineRead CAD · simulate · edit · retry
BoardPublic scores · before and after
PrivateHeld-out tasks, never published

Scores publish with the first public release. Nothing on this page is a result yet.

Sec_003 · The board

Public scores arrive with the first release

Fig_006 [ No scores yet ]
Design repair, live↓ Pass rate, highest first
  1. 01— awaiting first run
  2. 02— awaiting first run
  3. 03— awaiting first run
Score = designs that pass the re-run limit, 0–100 · higher is better