Skip to content

Compare named vehicles across the standard sims with opt-trade - #63

Merged
raorjun merged 2 commits into
mainfrom
feat/optsim-trade-study
Sep 24, 2026
Merged

raorjun merged 2 commits into
mainfrom
feat/optsim-trade-study

Conversation

@raorjun

@raorjun raorjun commented Sep 21, 2026 •

Copy link
Copy Markdown
Member

What

This PR adds make opt-trade. It compares vehicles that you name, on metrics that you choose, across more than one standard sim. Stacked on #61.

OptSim could sample a design space (opt-standard) and invert for a setup (opt-solve). It could not answer the question that a design review asks: what does this specific change give, and what does it cost? This PR replaces nothing. The docs now start with a table that shows how the three tools relate.

name: rear_roll_stiffness
candidates:
  stiff_rear_bar:    {rear.stabar.rate_n_m_per_rad: 961.495352}
  soft_front_spring: {front.actuation.spring_rate_n_per_m: 21015.2202}
  both:              {rear.stabar.rate_n_m_per_rad: 961.495352, front.actuation.spring_rate_n_per_m: 21015.2202}
metrics:
  SteadyStateEval: {understeer_gradient_deg_per_g: {resolution: 0.02}}
  TransientEval:   {yaw_rise_time_s: {resolution: 0.005}}

Support for the other standard sims needed only a small change. SteadyStateEval, RampSteerEval and TransientEval all run the same compiled model. _3_StandardSim already builds it once for all three. The three also share one interface: python -m <module> <config>, the same simulation and report keys, and <stem>_metrics.csv. So pipeline/standards.py is a registry of three entries and a generic runner. One compile for each vehicle serves every standard. FourPostEval is not included, on purpose. It runs FourPostSim, which is a different model.

The sweep's batch.py now sends each standard through that registry. Before this change, a second standard in compiler_config.yaml would have run SteadyStateEval's report against the wrong executable.

The trade study compiles each candidate. It never uses an override. opt-solve makes the opposite choice, and both choices are on purpose. A trade study must be able to change mass, CG, toe and camber. #61 showed that an override of those parameters silently fails. The solver's evaluator had a private cache of compiled vehicles. That cache is now pipeline/variants.py, and both tools share it. The cache key is the content. So two studies build and simulate a shared vehicle only once, for example the baseline.

The report has no score and no ranking, on purpose. How much understeer is worth how much settling time is an engineering decision. A weighted sum would hide that decision in a number. The report does make sure that each difference is worth reading:

  • Resolution. The report shows a delta smaller than the metric's resolution, but marks it ~. The resolution is the smallest change worth acting on.
  • Lost cases. A run whose simulation lost cases is n/c, and the report does not compare it. Its fits use fewer points, so its delta would mix the design change with missing data. A baseline that lost cases makes that whole standard invalid. Exit 2.
  • Stacking. Where one candidate is exactly two others combined, the report gives the interaction: the combined result minus the sum of the parts. So the study does not assume that "we'll do both" equals the sum.
  • Ambiguous names. TransientEval reports yaw_gain_dc once for the step and again for the frequency sweep. You can ask for it only with its group (step.yaw_gain_dc). The bare name is an error that lists the options. A metric that repeats with different values is an error. A metric that repeats with the same value is one metric. TransientEval really does write yaw_overshoot_pct twice.

How to check

python -m pytest tests/test_optsim_trade.py -q     # 16 passed, no OpenModelica needed
make opt-trade                                      # configs/trade_study.yaml
make opt-solve                                      # same behavior, now on the shared store

After the rebase on main: 566 passed, 18 skipped, and ruff is clean. Before the rebase, mypy was clean inside the container, which has CI's scipy-stubs.

I ran the example study for real: 4 vehicles × 2 standards, OpenModelica 1.26.3, 12-CPU container.

  • The cold run gave exit 0 in 1144 s: four compiles, eight runs and eight PDFs.
  • The fully cached run took 3 s.
  • The sweep's committed _doe_config.yaml was byte-identical after the runs.

The numbers are credible:

  • Baseline understeer is 0.2864, which matches the standard's own figure.
  • A stiffer rear bar reduces roll gradient from 0.886 to 0.794. It increases roll overshoot from 25.1 to 28.6 %.
  • A softer front spring does the opposite.
  • For every candidate, yaw rise time, yaw overshoot and settling time change by less than their resolution. The report marks them as such and does not show them as findings.

I verified opt-solve again end to end after its evaluator moved to the shared store. It converged on the same setup and metrics as before the move: 231.5 / 748.5 / 22578.7 / 52139.2 gives 0.3130 / 0.8502.

Review time is best spent on two items:

  • pipeline/trade.py, which decides when a delta is worth reading.
  • Whether the resolution semantics for the stacking verdict are what you want.

Notes

A decision that needs a second opinion. In the real run, both changed understeer by -0.032. Its parts predicted -0.018. So the interaction is -0.014, almost as large as the two parts together. That was enough to put both over the 0.02 resolution, although neither part crossed it. The interaction itself is below resolution. So the label reads additive within resolution, with the number next to it. My first label was "stacks", and I changed it. The correct claim is "cannot be told apart from zero", which is weaker than "adds". If you prefer to flag an interaction that is large relative to its parts, that is a one-line change.

Compare results from the same tool. Each metric here comes from its standard's own test matrix. So an opt-trade gradient matches make standard-eval-steady-state. It does not match opt-solve, which fits through its own denser isoline.

A candidate can change only a declared variable (sweep.variables in vehicle_architecture.yaml). Each variable needs its Modelica record mapping. To trade on a new variable, such as wheelbase, declare it there first. The study accepts a value outside a variable's sweep range and adds a note.

This PR also stops the solver from fitting through an evaluation that lost cases. The trade study refused to compare such a run, but the solver did not. So a star point that lost cases would have bent the surrogate silently. Because evaluations are cached, it would also have bent every later solve. Both tools now share standards.case_loss.

The compile-only knob path is now verified end to end. #61 listed it as not verified. The run used front toe, rear toe and the rear bar together:

  • One batch built five executables. The bar's star points shared the baseline executable by override.
  • The verification compiled exactly one more executable and skipped 5.
  • The solve converged on the first verification at 0.3578 / 0.8359, against targets of 0.35 / 0.85, in 617 s.

The run also shows what the bars and springs could not do. Front toe changed understeer from 0.325 to 0.358. So on this car, toe has the authority over balance that roll stiffness does not have.

The first try of this run crashed, and the crash was useful. Another process edited hashed tooling files between the star and the verification compile. The staleness check correctly refused to continue. But its message came from the sweep's compiler and said to run make clean-opt. That advice is wrong for a store that discards a stale cache by itself. This PR replaces the message.

Not verified: RampSteerEval through opt-trade end to end. It is registered, and a test checks its interface against its real config. I did not run it. There is no regression-baseline run. This PR touches no physics.

Follow-ups, not in this PR:

  1. FourPostEval needs a compile for each model. VariantStore builds one model for each vehicle today.
  2. The sweep still compiles once for each standard. So a second VehicleSim standard in the sweep would build the same model twice. The store shows the fix.
  3. All the follow-ups from Solve for a setup directly instead of searching a sweep #61 still apply. The main one: every limit_* metric is a fallback until the SteadyStateEval test matrix reaches a limit. A trade study reports those columns correctly. That does not make them mean what their names say.

@raorjun
raorjun added this pull request to stack #65 September 21, 2026 03:25
@raorjun
raorjun marked this pull request as ready for review September 21, 2026 03:31
@raorjun
raorjun force-pushed the feat/optsim-trade-study branch from 4e10dc5 to ac575af Compare September 23, 2026 06:26
@raorjun
raorjun force-pushed the feat/optsim-trade-study branch from ac575af to f3a445b Compare September 24, 2026 01:27
Base automatically changed from feat/optsim-setup-solver to main September 24, 2026 01:30
OptSim could sample a design space and invert for a setup. It could not
answer the question that a design review asks: what does this specific
change give, and what does it cost, across more than one study.
opt-trade reads a YAML of named candidates and the metrics to compare.
It compiles each candidate once and runs each requested standard against
that one executable. Then it writes a comparison table.

SteadyStateEval, RampSteerEval and TransientEval all run the same
compiled model, and _3_StandardSim already builds it once for all three.
They also share one interface. So pipeline/standards.py is a registry of
three entries and a generic runner. The sweep's batch step now uses the
registry. Before this change, a second standard in compiler_config.yaml
would have run SteadyStateEval's report against the wrong executable.

opt-trade compiles each candidate and does not use an override. opt-solve
makes the opposite choice. A trade study must be able to change mass, CG, toe
and camber, and an override of those silently does nothing. The solver's
evaluator had a private cache of compiled vehicles. That cache is now
pipeline/variants.py, and both tools share it. The cache key is the
content, so two studies build and simulate a shared vehicle only once.

The report has no score and no ranking. It makes sure that each
difference is worth reading:

- It shows a delta below the metric's stated resolution, but marks it.
- It does not compare a run that lost simulation cases.
- If one candidate is exactly two others combined, it reports the
  interaction. So "we'll do both" is not assumed to be the sum.

TransientEval reports some metrics once per group under one name. You
must ask for those by group. A metric that repeats with different values
is an error.

Real run on the example study, four vehicles by two standards: 1144 s
cold, 3 s cached.
The trade study refused to compare a run that lost simulation cases,
but the solver did not. A star point that lost a case still returns
finite gradients, fitted through fewer points. So it bent the surrogate,
and no number looked wrong. Evaluations are cached, so it would also
have bent every later solve. Both tools now share one definition,
standards.case_loss. The solver stops and reports the variant, the
counts and what to change.

This commit also replaces the error for tooling that changes during a
run. The error came from the sweep's compiler and said to run
make clean-opt. That advice is wrong here, because the store discards a
stale cache by itself on the next run. A real run hit this error when
another process edited hashed files between the star and the first
verification compile.

The first end-to-end solve over compile-only knobs found that problem.
The solve now passes. Front toe, rear toe and the rear bar together
built five executables in one batch, and the bar's star points shared
the baseline executable. The verification compiled exactly one more.
The solve converged on the first try at 0.3578 / 0.8359, against
0.35 / 0.85, in 617 s.
@raorjun
raorjun force-pushed the feat/optsim-trade-study branch from f3a445b to a36fbf7 Compare September 24, 2026 01:30
@raorjun
raorjun merged commit 212344d into main Sep 24, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant