ALEX BLYTHE / DATED EVIDENCE ADDITION /
AeroFEA.
A specialist, not an authority.
The existing fine-tuned intent reader is followed by deterministic JSON, field-bound unit and solver-admission checks. A separate model reviewer can veto, but cannot overrule those checks. This is an orchestration pipeline, not a true mixture-of-experts architecture.
Read the PDF evidence reportDownload solver evidenceAudited measurementsSHA-256 manifest
AeroFEA at a glance

What the fresh gate established
The new 20-case packet contains 12 solvable requests and eight that must not be solved. The existing pipeline attempted 6 inappropriate dispatches; test assertions intercepted them before knowingly incorrect physics ran. The specialist’s outcome is measured separately from raw model correctness.
Scope: personal synthetic fixtures, four known linear-elastic static templates, same-author holdout. Analytical agreement is not experimental validation, broad aerospace knowledge, design certification, or autonomous engineering authority.
Accuracy, rejection, and regression are different questions

| Measurement | Result | Meaning |
|---|---|---|
| Raw adapted reader | 14 / 20 | All extracted fields correct before downstream checks. |
| Existing grounding | 10 / 20 | Same new model outputs, old interpretation checks. |
| Existing end-to-end | 7 / 20 | Old dispatch behavior, strict fields and numerical acceptance on those same outputs. |
| Specialist grounding | 19 / 20 | Explicit source-bound conversions and conflict handling. |
| Specialist end-to-end | 19 / 20 | Correct fields plus appropriate refusal or numerically adequate solve. |
| Legacy language replay | 615 → 615 / 618 | 0 previously correct interpretations newly broken; cached predictions, not new inference. |
| Unit tests | 36 passed | 32 frozen core tests plus four additive opt-in MCP interface tests. |
The previous 40 cases are explicitly development replay, not a reused fresh holdout. The implementation and scoring code were frozen before the new packet was constructed and before fresh inference. Model inference received original request text and IDs, never expected answers.
One strict failure remains: case pipeline_fresh_018 emitted the invalid load-type enum uniform. The new schema check blocked dispatch, but safe rejection does not erase the model's incorrect output or turn 19/20 into 20/20.
Actual numerical proofs

Rectangular Timoshenko beam displacement, signed restrained thermal stress -E α ΔT, and Lamé open-ended cylinder bore hoop stress are checked against independently calculated references. Relative numerical tolerance: 3%. Final mesh change: ≤1% where applicable. The uniform thermal bar is exact on one mesh; it has no mesh-convergence claim.
The publication audit verified 203 retained solver-artifact hashes and independently recovered 12 reported observables from the actual DAT outputs. The decks, results and refinement outputs are retained in the evidence download.
FEA renders drawn from solver output

Original deck · Raw results

Original deck · Raw results

Original deck · Raw results

Original deck · Raw results
Beam deformation is magnified by the annotated factor. Colour is a corner-node displacement average or an element integration-point mean stress, as labeled. Quadratic-element corner faces are a visualization approximation. The pressure render is the actual axisymmetric strip, not a decorative 3D cylinder. A field image does not itself establish validation.
Render source hashes and case IDsEvery request, original response, decision, unit trace and failure
The model checker experiment did not earn trust

| Stock reviewer | Correct decisions / 12 | Wrong interpretations approved / 8 | Median CPU time |
|---|---|---|---|
| qwen2.5:1.5b | 4 / 12 | 8 / 8 | 3.28 s |
| qwen2.5:7b | 6 / 12 | 6 / 8 | 16.49 s |
Both reviewers used the same frozen prompt and diagnostic mutations: four correct interpretations and eight deliberately wrong ones. Execution was CPU-only, four threads, num_gpu=0, with no retries. Verdict correctness does not establish the accuracy of every explanation; the original feedback is retained.
No new checker fine-tune is claimed. These stock models are not enabled as a default check or treated as a safety boundary. An independently authored training set and a fresh acceptance gate are required before any future checker fine-tune is promoted.
Timing and implementation boundary
The adapted Qwen2.5-14B reader took 38.04 seconds median in native NF4/BF16/SDPA execution on an RTX 3090. The deterministic checks added 7.75 milliseconds median, excluding inference and solver time. This is not deployed GGUF performance, a GB10 comparison, or an isolated peak-throughput test: a separate CPU reviewer diagnostic ran alongside part of inference.
The implementation is available as an explicit opt-in local API, CLI and separate guarded MCP server. Existing private model routing and the public conversation gateway were not silently replaced. New solve folders preserve existing files, and the guarded MCP path writes a hash-identified trace.
Cloud assistance was used to author code and reporting. Model inference and numerical solves were local. No employer data, private corpus, credentials, model weights or host execution logs are published. Reserved answers were not ingested into training or RAG.
Complete PDF reportImplementation and scoring hashesReturn to Aero