ALEX BLYTHE / DATED EVIDENCE ADDITION /

AeroFEA.
A specialist, not an authority.

The existing fine-tuned intent reader is followed by deterministic JSON, field-bound unit and solver-admission checks. A separate model reviewer can veto, but cannot overrule those checks. This is an orchestration pipeline, not a true mixture-of-experts architecture.

AeroFEA at a glance

Qwen fine-tuned for FEA vs. the base model: Qwen2.5-14B-Instruct and the AeroFEA fine-tune, trained on 3,240 synthetic FEA request-to-setup examples. Grounded extraction: base 227/240 (94.6%), tuned 240/240 (100%). Separate check comparison using the same tuned model outputs: legacy checks 7/20 (35%), source-bound checks 19/20 (95%). Training and charted results used RTX 3090. Cantilever sketch illustrates a setup, not a solver result.
Qwen2.5-14B-Instruct versus its AeroFEA FEA-setup fine-tune on 240 fresh synthetic requests. The second chart compares the same tuned model outputs with legacy versus source-bound checks on 20 additional cases—not base versus tuned models. The cantilever sketch illustrates a setup; actual solver-field renders appear below. Download the graphic.
BOUNDED ACCEPTANCE / REVIEWER FAILURES RETAINED

What the fresh gate established

19 / 20Strict end-to-end cases
12Actual CalculiX solves
0Unsafe specialist admissions
7.75 msMedian check overhead

The new 20-case packet contains 12 solvable requests and eight that must not be solved. The existing pipeline attempted 6 inappropriate dispatches; test assertions intercepted them before knowingly incorrect physics ran. The specialist’s outcome is measured separately from raw model correctness.

Scope: personal synthetic fixtures, four known linear-elastic static templates, same-author holdout. Analytical agreement is not experimental validation, broad aerospace knowledge, design certification, or autonomous engineering authority.

Accuracy, rejection, and regression are different questions

Fresh-case extraction and end-to-end accuracy beside inappropriate solver-dispatch attempts, comparing the old and specialist pipelines.
MeasurementResultMeaning
Raw adapted reader14 / 20All extracted fields correct before downstream checks.
Existing grounding10 / 20Same new model outputs, old interpretation checks.
Existing end-to-end7 / 20Old dispatch behavior, strict fields and numerical acceptance on those same outputs.
Specialist grounding19 / 20Explicit source-bound conversions and conflict handling.
Specialist end-to-end19 / 20Correct fields plus appropriate refusal or numerically adequate solve.
Legacy language replay615 → 615 / 6180 previously correct interpretations newly broken; cached predictions, not new inference.
Unit tests36 passed32 frozen core tests plus four additive opt-in MCP interface tests.

The previous 40 cases are explicitly development replay, not a reused fresh holdout. The implementation and scoring code were frozen before the new packet was constructed and before fresh inference. Model inference received original request text and IDs, never expected answers.

One strict failure remains: case pipeline_fresh_018 emitted the invalid load-type enum uniform. The new schema check blocked dispatch, but safe rejection does not erase the model's incorrect output or turn 19/20 into 20/20.

Actual numerical proofs

Relative error of each actual CalculiX result against its closed-form reference, with the frozen three-percent tolerance marked.

Rectangular Timoshenko beam displacement, signed restrained thermal stress -E α ΔT, and Lamé open-ended cylinder bore hoop stress are checked against independently calculated references. Relative numerical tolerance: 3%. Final mesh change: ≤1% where applicable. The uniform thermal bar is exact on one mesh; it has no mesh-convergence claim.

The publication audit verified 203 retained solver-artifact hashes and independently recovered 12 reported observables from the actual DAT outputs. The decks, results and refinement outputs are retained in the evidence download.

FEA renders drawn from solver output

Actual CalculiX cantilever field, with physical units and rendering assumptions annotated.
cantilever · 6.20682 mm
Original deck · Raw results
Actual CalculiX simply supported field, with physical units and rendering assumptions annotated.
simply supported · 0.390020 mm
Original deck · Raw results
Actual CalculiX restrained thermal bar field, with physical units and rendering assumptions annotated.
restrained thermal bar · 60.9336 MPa
Original deck · Raw results
Actual CalculiX thick cylinder field, with physical units and rendering assumptions annotated.
thick cylinder · 11.3285 MPa
Original deck · Raw results

Beam deformation is magnified by the annotated factor. Colour is a corner-node displacement average or an element integration-point mean stress, as labeled. Quadratic-element corner faces are a visualization approximation. The pressure render is the actual axisymmetric strip, not a decorative 3D cylinder. A field image does not itself establish validation.

The model checker experiment did not earn trust

False approvals and observed CPU latency for the stock 1.5B and 7B interpretation reviewers.
Stock reviewerCorrect decisions / 12Wrong interpretations approved / 8Median CPU time
qwen2.5:1.5b4 / 128 / 83.28 s
qwen2.5:7b6 / 126 / 816.49 s

Both reviewers used the same frozen prompt and diagnostic mutations: four correct interpretations and eight deliberately wrong ones. Execution was CPU-only, four threads, num_gpu=0, with no retries. Verdict correctness does not establish the accuracy of every explanation; the original feedback is retained.

No new checker fine-tune is claimed. These stock models are not enabled as a default check or treated as a safety boundary. An independently authored training set and a fresh acceptance gate are required before any future checker fine-tune is promoted.

Timing and implementation boundary

The adapted Qwen2.5-14B reader took 38.04 seconds median in native NF4/BF16/SDPA execution on an RTX 3090. The deterministic checks added 7.75 milliseconds median, excluding inference and solver time. This is not deployed GGUF performance, a GB10 comparison, or an isolated peak-throughput test: a separate CPU reviewer diagnostic ran alongside part of inference.

The implementation is available as an explicit opt-in local API, CLI and separate guarded MCP server. Existing private model routing and the public conversation gateway were not silently replaced. New solve folders preserve existing files, and the guarded MCP path writes a hash-identified trace.

Cloud assistance was used to author code and reporting. Model inference and numerical solves were local. No employer data, private corpus, credentials, model weights or host execution logs are published. Reserved answers were not ingested into training or RAG.