AERO / LOCAL-MODEL OPERATOR BENCHMARK

Can a local model pilot a guarded CFD workflow?

Two public OpenFOAM seeds with retained proposal failures. These trials test case handoff and evidence reading, not aerodynamic accuracy or design readiness.

COMPLETED WORKFLOWS

2 bounded passes

Airfoil v7 and internal-flow v2 produced completed solver runs and verified evidence manifests. Their observers kept design readiness at no.

7B LOCAL RUN

Solver completed

The 7B model authored the accepted proposal and read the terminal state. Its strict observer grade failed on evidence-key aliases. A Codex operator prepared and started the run.

Retained attempts

Airfoil v1-v7 are iterative revisions, not independent samples. The 14B airfoil rubric and the 89/100 internal-flow review use different scoring methods and cannot be averaged.

TrialModelOutcomeSolver?Finding
Airfoil v17BProposal rejectedNoConservation and wall-treatment criteria missing.
Airfoil v27BProposal rejectedNoReference, mesh-independence and model-fidelity criteria missing.
Airfoil v37BProposal rejectedNoLift, drag, mesh-independence and model-fidelity criteria missing.
Airfoil v47BSchema rejectedNoInvalid assumptions field.
Airfoil v514BProposal rejectedNoUnresolved setup questions.
Airfoil v614BFreeze rejectedNoWall-y-plus request unsupported for the template.
Airfoil v714BWorkflow passYesSolver and manifest verified; strict observer passed; design-ready no. Recorded workflow rubric: 100/100.
Internal flow 7B7BSolver completed; strict observer failedYesModel-authored proposal and numerical pass; two evidence-key aliases. Operator-assisted setup and start.
Internal flow v114BSchema rejectedNoCross-backend fea_case field in an OpenFOAM proposal.
Internal flow v214BWorkflow passYesCorrected proposal, numerical pass, verified manifest, strict observer pass; design-ready no. Independent workflow review: 89/100.

The engineering boundary

Aero's typed harness and operator approval decide what may run; OpenFOAM and retained evidence decide what happened. Numerical convergence is not validation. The successful internal-flow run still lacks an independent matched reference, demonstrated model fidelity, and mesh independence.

Was the 7B run fully solo?

No. The local model wrote the accepted proposal and assessed the solver evidence. Earlier proposal failures led to a clearer reusable example, and a Codex operator invoked the staged handoff and solver start. The receipt does not name that operator's model, so Luna's involvement in this specific run is unverified.

Read the complete report

The report records each retained attempt, the 7B model's role, the operator's role, strict observer limits, and audit hashes.