Commit Graph
176 Commits
Author SHA1 Message Date
swaits 819014bf58 docs(mutants): note the 2026-05 profiling campaign's equivalent mutants
The profiling campaign's round-4 `pareto_front` change adds a
`dominated` bitset that is pure skip-bookkeeping — deleting the mark
write or the skip check leaves the returned front bit-identical. A
future mutation run will report those as MISSED; record here that they
are genuine equivalent mutants, not test gaps, so nobody chases them
with new tests.
2026-05-14 12:58:28 -06:00
swaits a4b11f286e perf(pareto): skip already-dominated points in pareto_front
`pareto_front` re-scanned every candidate from scratch. Track a
`dominated` bitset instead: whenever `i`'s scan finds `i` dominates `j`,
mark `j` so the outer loop skips `j` outright when it reaches it. The
inner check now reads both directions of `pareto_compare` — the
objective scan already computes both flags, so this is ~free.

Bit-identical, including under NaN-intransitive dominance: a mark is
only ever set from a direct pairwise `pareto_compare` result, never
inferred transitively. Also strict-or-neutral on work — the marks can
only ever let the outer loop *skip*, never add a scan.

Whole-program callgrind Ir for the compare_profile benchmark:
169,440,644,233 -> 165,311,562,939 (-2.44%); `pareto_front` self-Ir
44.1B -> 39.9B.
2026-05-14 12:49:47 -06:00
swaits 66b4d9fa6b perf(age_moea): score only the splitting front in environmental_selection
`prox` and `nearest` were computed for every member of `combined`, but
the scoring loop only ever reads the entries for the splitting front
(`remaining`). Fill just those, skipping the `lp_norm` / `lp_distance`
work — and the `powf` calls inside them — for the rest of `combined`.

Bit-identical: the skipped entries were never read. In the
compare_profile benchmark this cut `pow` + its libm kernel from ~16.2B
to ~14.3B Ir and `environmental_selection` self-Ir from 2.13B to 1.82B.

Round 3 whole-program: 173,803,642,945 -> 169,440,644,233 (-2.51%).
2026-05-14 12:37:42 -06:00
swaits 0ea09bdf82 perf(hype): reuse the per-sample dominators buffer
`estimate_contributions` heap-allocated a fresh `dominators: Vec<usize>`
on every Monte Carlo sample — thousands of alloc/free pairs per call.
Hoist it out of the sample loop and `clear()` it each iteration.

Bit-identical. In the compare_profile benchmark this removed ~4.2M
malloc/free pairs, dropping `malloc` + `free` self-Ir by ~0.24B.
2026-05-14 12:37:41 -06:00
swaits f3088e8353 perf(ibea): precompute the exp-transformed indicator matrix
`environmental_selection` recomputed `exp(-indicator[worst][i] / scale)`
in its removal loop — the exact value already computed when building the
initial fitness vector. Pre-exponentiate the indicator matrix once; both
the initial fitness sum and every per-removal update then read from it,
turning the O((pool-n) · pool) removal-loop `exp` sweep into additions.

Bit-identical: same input bits -> same `exp` -> same output bits, and
the fitness sum order is preserved. In the compare_profile benchmark
this cut `exp` + its libm kernel from ~6.95B to ~5.40B Ir and
`environmental_selection` self-Ir from 8.89B to 8.73B.
2026-05-14 12:37:41 -06:00
swaits 2bd8f8fc11 perf(pareto): precompute oriented buffers in pareto_front
`pareto_front` was still the naive O(n²) formulation: a raw double loop
calling `pareto_compare` for every ordered pair, re-deriving feasibility
and the minimization-oriented objective values on every comparison.

Apply the same precompute pattern `non_dominated_sort` already uses:
hoist per-individual `feasible` / `violation` / `oriented` (a flat n*m
buffer) out of the loop, then run a branchless inlined dominance check
over the contiguous buffer. The kept set and its order are unchanged.

Whole-program callgrind Ir for the `compare_profile` benchmark:
221,836,742,708 -> 173,803,642,945 (-21.65%); `pareto_front` self-Ir
92.1B -> 44.1B.
2026-05-14 12:15:04 -06:00
swaits 50501bb01e test(bench): disable callgrind cache simulation in compare_profile
The profiling campaign ranks functions on instruction count (Ir) only,
so callgrind's cache simulation is pure overhead — it roughly doubles
each run's wall time. Pass `--cache-sim=no` via the gungraun
`LibraryBenchmarkConfig` to halve the measure->optimize loop's latency.
2026-05-14 12:01:25 -06:00
swaits 740965958f perf(pareto): make pareto_compare allocation-free
The objective-comparison branch of `pareto_compare` materialized two
`Vec<f64>`s per call via `ObjectiveSpace::as_minimization`. Because
`pareto_compare` runs O(n²) times across the multi-objective algorithms,
that per-call allocation pair dominated the whole `compare` workload.

Replace it with an allocation-free per-objective scan that branches on
`Objective::direction` directly: for a Maximize axis "a beats b" is just
`av > bv`, bit-identical to `-av < -bv` after orientation. The result is
unchanged for every input.

Whole-program callgrind Ir for the `compare_profile` benchmark:
357,060,633,544 -> 221,836,742,708 (-37.87%).
2026-05-14 12:01:17 -06:00
swaitsandClaude Opus 4.7 1cecd44511 test(bench): profile the full compare workload with gungraun
Adds benches/compare_profile.rs: a gungraun library_benchmark that runs
the entire `compare` workload once at seed 0 under callgrind. It path-
includes the shared examples/_shared/compare_workload.rs module and calls
the new profile_workload() entry point, which invokes all 82 algorithm
runners and folds every result into a checksum so nothing is elided.

gungraun reports the whole-program instruction count and diffs it against
the previous run; the saved callgrind.out
(target/gungraun/heuropt/compare_profile/.../callgrind.full_compare_workload.out)
carries the per-function breakdown for callgrind_annotate. This is the
measurement harness for the function-level optimization campaign.

Round-0 baseline: 357,060,633,544 Ir. The shared module also carries the
example's presentation layer, so the bench gets a commented
`#![allow(dead_code)]` -- but the runner functions are deliberately not
allow-listed, so a runner missing from profile_workload still warns (that
check already caught one omission).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 11:24:19 -06:00
swaitsandClaude Opus 4.7 9df9c149d6 refactor(compare): split workload into a reusable module
Moves the problem definitions, the ~88 algorithm runners, and the
table-printing presentation out of examples/compare.rs into
examples/_shared/compare_workload.rs (path-included as `mod workload`).
examples/compare.rs is now a thin shim that calls `workload::run_all()`.

This makes the exact `compare` workload reusable: an upcoming gungraun
profiling benchmark (benches/compare_profile.rs) path-includes the same
module and drives the runner functions directly, so it exercises
identical code with no duplication. `examples/_shared/` has no `main.rs`,
so cargo does not auto-discover it as an example.

Pure reorganization: `cargo run --release --example compare` produces
identical output (every quality metric, front size, and row order
unchanged; only the non-deterministic wall-clock `ms` column jitters).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 10:23:35 -06:00
swaitsandClaude Opus 4.7 708f1ce9cd docs: promote MOEA/D to the multi/many-objective default
The compare harness shows MOEA/D is the single most consistent
performer: top-3 on every multi- and many-objective table (convex,
disconnected, spherical and linear fronts; 2 through 10 objectives) and
fastest or near-fastest every time. No other algorithm is close to that
consistency. This matches the literature view of MOEA/D as a strong,
robust, scalable baseline -- with the known caveat that weight-vector
spread can leave gaps on highly irregular fronts (the DTLZ/ZDT suite
doesn't stress that).

But both decision trees buried it: the README filed it under "Want
decomposition / weight-vector style" -- a stylistic branch -- and framed
it as a speed pick; the book left it out of the TL;DR table entirely.
Meanwhile NSGA-II was the listed 2-3-objective default despite losing to
MOEA/D on every table and collapsing past ~4 objectives.

Both trees now lead the multi- and many-objective branches with MOEA/D,
keep NSGA-II as the well-understood alternative and the combinatorial
go-to, add MOEA/D to the TL;DR / quick-reference tables, and soften the
NSGA-III "strong default" framing to match the data.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 09:06:18 -06:00
swaitsandClaude Opus 4.7 398c82326a feat(compare): add many-objective problems (DTLZ at 4, 10, 8 objectives)
Adds a many-objective section to the comparison harness, exercising the
regime where Pareto dominance stops discriminating: with enough
objectives almost every pair of solutions is mutually non-dominated.

- DTLZ2 4-objective: the entry point to many-objective.
- DTLZ2 10-objective: the curse of dimensionality in full.
- DTLZ1 8-objective: dominance collapse stacked on DTLZ1's deceptive
  multimodal g-term.

Implemented generically: the existing Dtlz1/Dtlz2 structs and distance
metrics are already objective-count agnostic, so a single `ManySpec` +
nine generic runners (RandomSearch, NSGA-II, NSGA-III, MOEA/D, RVEA,
GrEA, IBEA, HypE, AGE-MOEA) cover all three tables -- and any future M.

The results are a clean teaching story:
- NSGA-II collapses -- on DTLZ2-10 it finishes dead last, *worse than
  random search* (2.01 vs 0.63); its crowding distance actively
  misleads in 10-D.
- HypE / MOEA/D / GrEA / IBEA barely notice the 4 -> 10 jump.
- GrEA wins DTLZ1-8, consistent with the 3-objective DTLZ1 table.
- HypE reverses: #1 on both DTLZ2 tables, #6 on the deceptive DTLZ1-8.

Regenerated examples/compare-results.md with the three new sections.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 08:56:48 -06:00
swaitsandClaude Opus 4.7 a9d72b94f4 docs: correct disconnected-front and sequencing guidance from compare results
The `compare` harness contradicts two recommendations in the decision
trees:

- "Disconnected or non-convex front -> AGE-MOEA, KnEA, IBEA" had it
  backwards. Added KnEA to the ZDT3 table (the disconnected-front
  benchmark) so the claim is actually exercised: AGE-MOEA and KnEA
  finish *last and second-last*; IBEA wins, MOEA/D and NSGA-II follow.
  The trees now split "disconnected" from "non-convex contiguous",
  lead disconnected with IBEA, and note the geometry-aware methods
  trail when the front is in pieces.
- The book filed Simulated Annealing on permutations as a "one-decision
  baseline" and led the JSS row with GA. On the harness SA *wins* the
  FT06 job-shop table and ties for the TSP optimum; SA/Tabu edge out
  the GA. Reframed SA/Tabu as strong sequencing methods.

Also regenerated examples/compare-results.md for the new ZDT3 row.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 08:43:55 -06:00
swaitsandClaude Opus 4.7 7004a572c8 test(permutation): add ERX edge-preservation tests
ERX had five tests, all checking `is_strict_perm` validity -- none
verified the *point* of edge recombination: that children actually
inherit parent edges. A "valid permutation but edge-ignoring" ERX would
have passed every existing test.

Adds:
- erx_identical_parents_inherit_every_edge: with identical parents the
  child's edge set must equal the parent's exactly (zero foreign edges).
- erx_preserves_parent_edges_better_than_order_crossover: ERX must
  strand fewer non-parent edges than Order Crossover -- a direct test of
  ERX's reason to exist.
- erx_pinned_output: locks the adjacency-walk + min-degree tie-break.

Investigation result: ERX is correct and effective. It wins the
tsp_operators_compare showdown on KroAB-25 (hypervolume 638M vs OX 622M,
PMX 609M, CX 593M) and produces the most diverse front. The compare TSP
table's GA underperformance is an Order-Crossover-plus-generational-GA
artifact on a convex-position instance, not an ERX bug.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 08:36:44 -06:00
swaitsandClaude Opus 4.7 2b453464e7 refactor(compare): realign + sort tables, add combinatorial problems
The terminal output was misaligned: headers and data were right-aligned
with hardcoded column widths and separator lengths, and the mean ± std
cells contained the non-ASCII `±` (plus `ε`, `↑`, `↓`) -- on any terminal
that renders those at a non-1 column width the columns drift, and the
hardcoded `-`.repeat(n) separators didn't match the real table width
anyway.

Changes:
- New `print_table` helper: column widths derived from the actual cell
  contents (header + every row), separator length computed to match.
- All table cells are now ASCII: `+/-` instead of `±`, `eps-MOEA`
  instead of `ε-MOEA`, arrows dropped from headers.
- Every table is sorted best-first by its primary quality metric.
- Added three combinatorial / sequencing problems with their own
  (permutation- / bitstring-native) algorithm rosters: a convex-position
  ring TSP (known optimum), FT06 job-shop makespan (known optimum 55),
  and a bi-objective 0/1 knapsack scored by hypervolume.
- Expanded every problem's preamble: what it is, why it's hard, and the
  best-known / optimal result.
- Regenerated examples/compare-results.md to match.

Continuous-problem quality metrics are unchanged (bit-identical to prior
snapshots); only ms columns and row order move.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 07:51:12 -06:00
swaitsandClaude Opus 4.7 c50e390969 perf(bayesian_opt): reuse scratch buffers in the EI acquisition loop (2.40M -> 2.26M instr)
The acquisition loop ran acquisition_samples GP predictions per BO
iteration, each allocating three short-lived Vecs: the candidate point,
the k_star kernel vector, and solve_lower's output. Threading reused
buffers through new sample_uniform_in_bounds_into / predict_into /
solve_lower_into entry points removes ~3000 alloc/free pairs from
bayesian_opt_short.

bayesian_opt_short: 2_398_972 -> 2_255_904 (-6%). This is a structural
(allocation) win, not an algorithmic one -- the GP fit (Cholesky) and EI
prediction are inherently O(n^2)/O(n^3) with transcendental kernels, and
that work is unchanged. Output bit-identical -- all 606 tests pass,
including the run() snapshot; async builds clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 07:48:22 -06:00
swaitsandClaude Opus 4.7 9b1352e375 perf(tpe): compute KDE bandwidths once per iteration, not per call (383K -> 188K instr)
Each TPE iteration drew `candidate_samples` candidates; every candidate
triggered three scott_bandwidths calls (one in sample_from_kde, two in
log_kde_density) -- each an O(support) two-pass scan plus a powf(-0.2). But
the good / bad supports are fixed for the whole iteration, so only two
distinct bandwidth vectors exist. Deriving them once and threading them
through cuts ~34 of every 36 scott_bandwidths calls. Also hoists the
constant (2*pi).sqrt() out of the inner density loop.

tpe_short: 382_518 -> 187_898 (-51%, 2.04x). scott_bandwidths is
deterministic in its inputs, so the once-vs-many results are identical --
output bit-identical, all 606 tests pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 06:55:08 -06:00
swaitsandClaude Opus 4.7 cf26243765 perf(ant_colony): hoist powf out of the tour-building hot loop (1.04M -> 540K instr)
build_tour computed pheromone[i][j].powf(alpha) and eta[i][j].powf(beta) for
every candidate at every step of every ant -- two transcendental calls per
edge consideration. But eta is constant for the whole run and pheromone is
constant across a generation's ant loop. Pre-raising eta to beta once and
pheromone to alpha once per generation (into a reused buffer) turns the hot
per-candidate weight into a single multiply.

ant_colony_tsp_short: 1_036_680 -> 539_924 (-48%, 1.92x). build_tour now
takes the pre-raised matrices; the three direct-call tests pre-raise via a
`raise` helper. Output bit-identical -- all 606 tests pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 06:52:38 -06:00
swaitsandClaude Opus 4.7 3b4f765e03 perf(operators): make PMX and OX crossovers O(n) with value-indexed lookups
Both crossovers had an O(n) inner scan making them O(n^2): PMX located the
value to swap with `child.iter().position`, OX tested segment membership
with `segment.contains`. Both operate on strict permutations of 0..n, so a
value-indexed table (PMX: position kept in sync across swaps; OX: a static
membership bitmap) gives O(1) lookups. debug_asserts document the range
assumption, consistent with ERX and CX.

pmx_crossover_vary n=100: 21_507 -> 5_900 (-73%, 3.65x); n=30 -18%.
order_crossover_vary n=100: 18_265 -> 7_389 (-60%, 2.47x); n=30 -29%.
All four permutation crossovers are now O(n). Bit-identical for valid
permutations -- all 606 tests pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 06:42:35 -06:00
swaitsandClaude Opus 4.7 47939fc5cb perf(operators): make CycleCrossover O(n) with per-parent position tables (59K -> 12K instr)
cx_child located each cycle's next value with an O(n) `position` scan,
making the walk O(n^2). CX operates on strict permutations of 0..n, so a
direct value-indexed position table per parent gives O(1) lookups; a
debug_assert documents the range assumption (mirroring ERX).

cycle_crossover_vary n=100: 58_762 -> 12_084 (-79%, 4.86x); n=30 -36%.
Bit-identical for valid permutations -- all 606 tests pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 06:41:05 -06:00
swaitsandClaude Opus 4.7 4aab5c7029 perf(pareto): flatten non_dominated_sort objective buffer (1.29M -> 1.10M instr)
The O(n^2) pair loop reads oriented[j] for every j; with Vec<Vec<f64>>
that chased a separate heap allocation per individual. A flat n*m buffer
keeps those reads contiguous and sequential in j.

non_dominated_sort_2d n=200: 1_288_072 -> 1_096_738 (-15%, 1.17x); n=50
-19%. nsga2 one-generation -5.7%. Combined with the earlier antisymmetry
fix, n=200 is down 55% from the original 2.46M. Pure data-layout change --
output bit-identical, all 606 tests pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 06:38:03 -06:00
swaitsandClaude Opus 4.7 37821bdd3d perf(metrics): stop re-sorting hypervolume_nd prefixes per slice (361K -> 291K instr)
The HSO M=3 path called the generic 2-D base case for every last-axis
slice, which re-sorted the active prefix by axis 0 each time -- O(n^2 log n)
overall. Since `projected` is already in last-axis order, sorting the
projected indices by axis 0 once and sweeping them with a `pi > k` skip
gives O(n^2) with no per-slice allocation. The M>=4 path is unchanged
(lifted out of the inner branch verbatim).

hypervolume_nd_bench_3d n=100: 361_595 -> 291_247 (-19%, 1.24x); n=30 -16%.
The sweep visits points in the same (axis-0, then last-axis) order the
stable per-prefix sort produced -- output is bit-identical, all 606 tests
pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 06:35:52 -06:00
swaitsandClaude Opus 4.7 532d4a54db perf(pareto): sort crowding-distance keys without Vec<Vec<f64>> indirection (-4.5%)
crowding_distance sorted bare front indices with a comparator that chased
two Vec<Vec<f64>> indirections per comparison. Extracting (objective value,
front position) tuples into a buffer reused across objectives keeps the hot
comparator a single f64 compare.

crowding_distance_2d n=200: 181_493 -> 173_286 (-4.5%); n=50 -4.5%. Stable
sort over the (value, index) pairs preserves the original tie-order, so the
output is bit-identical -- all 606 tests pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 06:33:25 -06:00
swaitsandClaude Opus 4.7 3f104a4395 perf(operators): make ERX neighbor removal O(degree) per step (397K -> 140K instr)
Edge Recombination Crossover scrubbed `current` from every one of the n
adjacency lists on each step of the walk -- an O(n^2) pass. The parent-tour
adjacency relation is symmetric (b in adj[a] iff a in adj[b]), so `current`
only ever appears in the lists of its own neighbors. Taking adj[current]
out with mem::take and retaining only over those lists is O(degree).

edge_recombination_crossover_vary n=100: 397_214 -> 140_403 (-65%, 2.83x);
n=30: 62_708 -> 39_987 (-36%). Output bit-identical -- all 606 tests pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 06:29:23 -06:00
swaitsandClaude Opus 4.7 eb4c8a6a9f perf(pareto): halve non_dominated_sort dominance comparisons (2.46M -> 1.27M instr)
The dominance relation is antisymmetric, so the outcome of compare(i, j)
fully determines compare(j, i). Iterating only j > i and applying the
result in both directions does identical work in half the pair scans.

non_dominated_sort_2d n=200: 2_461_178 -> 1_268_372 (-48%); n=50 -1.65x.
Ripples into dependents: nsga2 one-generation -20%, nsga3 / sms_emoa ~-9%.
Output is bit-identical (dominates[] still ascending, first_front order
unchanged) -- all 606 tests including run() snapshots pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 06:26:08 -06:00
swaitsandClaude Opus 4.7 b1d339a869 test(bench): cover permutation operators, combinatorial problems, and remaining algorithms
Expand benches/hot_paths.rs so the instruction-count harness exercises the
whole library. Adds four groups — permutation_ops_group (all 10 permutation
operators at n=30/100), variation_ops_group (BitFlip, Levy, BoundedGaussian,
ClampToBounds, ProjectToSimplex), combinatorial_group (TSP/JSS/knapsack
end-to-end plus AntColonyTsp), and multi_fidelity_group (Hyperband) — and
folds tabu_search_short and umda_short into single_objective_group. All 33
algorithms and 20 operators are now on the benchmarking surface (69 benches).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 06:21:59 -06:00
swaits 54f9db5dc9 docs(mutants): document the mutation-testing campaign in mutants.toml
Records the outcome of the 2026-05 mutation-testing campaign and the
gotchas for future runs:
- always pass --all-features so the async runners and explorer module
  are compiled (otherwise their mutants are unviable/missed noise);
- --test-tool nextest needs --no-config because the libtest-style
  --test-threads=1 arg isn't accepted there;
- where the new invariant tests live, and what the residual MISSED /
  TIMEOUT categories actually represent.
2026-05-14 03:47:22 -06:00
swaits 6aee6d6318 style: rustfmt the Phase 1 test additions
The per-file Phase 1 test commits were written without running rustfmt
as I went; this pass formats the new test code (long assert_eq! lines
wrapped, etc.). Formatting-only — no behavioural change.
2026-05-14 03:41:38 -06:00
swaits 2ca8d71b94 test(snapshot): pin exact run() output for every algorithm
The full-codebase mutants run showed ~600 of the 902 surviving mutants
are arithmetic / comparison flips inside algorithm run() bodies — the
per-helper Phase 1 tests don't reach the optimization loop itself, and
the existing deterministic-with-same-seed tests can't catch them (both
the clean and mutated runs use the same seed, so they still match).

This extends the async-parity sweep: each of the 33 parity tests now
also asserts the sync run's result against an exact captured snapshot
(best objectives for single-objective algorithms; sorted pareto-front
objective tuples for multi-objective ones). Any arithmetic flip
anywhere in run() perturbs at least one f64 and breaks the snapshot.

Fixtures are deliberately multi-dimensional — SnapSphere (3-D sum of
squares), SnapMo (3-variable / 2-objective), a 6-city TinyTsp, a 3-D
SnapSpherePartial for Hyperband. A 1-D problem leaves the per-axis /
covariance-matrix / simplex machinery degenerate, so arithmetic
mutations there wouldn't change the result; 3-D exercises the full
loop body.

Snapshots captured from the un-mutated implementation; an intentional
algorithm change requires regenerating them, by design. The assertions
live in the async-gated module because they reuse its per-algorithm
constructions — active during the mutation campaign
(--features async,serde) and under cargo test --features async.
2026-05-14 03:41:38 -06:00
swaits 7aa9e627f0 test(sms_emoa,pesa2,paes,random_search): pin remaining algorithm helpers
Phase 1, final algorithm batch:
- sms_emoa: pick_drop_index returns the singleton worst front, and
  finds the least-HV-contributor at a non-zero index.
- pesa2: build_grid empty/corner-point boxing; region_tournament
  prefers the less-crowded grid box (statistical majority).
- paes: deterministic non-empty front + archive cap.
- random_search: evaluation count = iterations*batch; best is no
  worse than any sampled candidate.
2026-05-13 22:58:17 -06:00
swaits c5b003a9c9 test(tlbo,umda,simulated_annealing,tabu_search,tpe,snes,rvea,spea2): pin comparison and geometry helpers
Phase 1 tests for eight more algorithms — the feasibility-first
comparison helpers (better / better_than / compare_so / worse_than),
plus algorithm-specific pure functions:
- tlbo: best_index min/max/tie.
- tpe: oriented_target sign-flip + penalty; split_good_bad partition
  and clamp-to-at-least-one-each.
- snes: nes_utilities sum-to-zero + descending + positive-best.
- rvea: unit_normalize (3-4-5 → 0.6/0.8), zero-vector passthrough;
  closest_reference smallest-angle; smallest_neighbor_angle = π/2 for
  orthogonal refs.
- spea2: euclidean distance basics; binary_tournament prefers lower
  fitness.
2026-05-13 22:58:17 -06:00
swaits 6819ce4091 test(nelder_mead,nsga2,nsga3,one_plus_one_es,particle_swarm): pin selection/geometry helpers
Phase 1 tests:
- nelder_mead: compare / better feasibility-first + direction.
- nsga2: binary_tournament prefers lower rank, then higher crowding
  distance at equal rank (statistical majority over 200 seeds).
- nsga3: solve_intercepts on axis-aligned extremes / singular / empty;
  associate picks the closest reference direction with correct
  perpendicular distance.
- one_plus_one_es: worse_than across feasibility + direction + equal.
- particle_swarm: best_index min/max/tie/single-element.
2026-05-13 22:58:17 -06:00
swaits c2319116b8 test(hyperband,moead,knea,ibea,ipop_cma_es,mopso): pin helper functions
Phase 1 tests:
- hyperband: compare / better feasibility-first + direction branches.
- moead: tchebycheff (max weighted deviation from ideal) and
  weight_distance (Euclidean) pins.
- knea: perpendicular_distance to the simplex hyperplane, zero-on-plane,
  and the too-few-extremes degenerate fallback.
- ibea: compute_fitness empty/dominating/symmetric-tradeoff cases and
  binary_tournament fitness preference.
- ipop_cma_es: better feasibility-first + direction + equal-not-better.
- mopso: population/front sizing and determinism cross-check.
2026-05-13 22:58:17 -06:00
swaits 952d93ac85 test(grea,hill_climber,hype): pin selection sizing and tournament logic
Phase 1 tests:
- grea: environmental_selection truncates the 2N pool to exactly N
  across three population sizes.
- hill_climber: full-run never-worsens and decreases-sphere pins.
- hype: binary_tournament picks the higher-fitness index (statistical
  majority + valid-index invariant).
2026-05-13 22:58:17 -06:00
swaits 8ff43d0240 test(epsilon_moea): pin box_coords, corner_distance, box_dominates
Phase 1 tests for the ε-MOEA box-archive helpers: per-axis floor in
box_coords, Euclidean corner_distance (incl. zero at exact corner),
and box_dominates across the strict/boundary cross-product.
2026-05-13 22:58:17 -06:00
swaits 03b450f050 test(genetic_algorithm): pin compare_for_fitness and survival_selection
Phase 1 tests for GA — feasibility-first fitness comparison across all
branches, and survival_selection's exact elite + best-offspring
composition (including the zero-elitism case).
2026-05-13 22:58:17 -06:00
swaits 4569244a68 test(pareto,metrics,selection): pin shared-utility comparisons and arithmetic
Phase 1, tier 3 of the mutation-testing campaign — the shared Pareto /
metric / selection utilities used by every multi-objective algorithm.
A scoped cargo-mutants run found 75 survivors across these files; the
tests below target them.

- metrics/hypervolume.rs: dominates() boundary cases, non_dominated_
  projection retained-set pins, hso_recursive 1-D/2-D base cases,
  hypervolume_nd_from_evaluations empty/non-dominating skips.
- selection/tournament.rs: challenger_wins across the full feasibility
  cross-product + equal-objective tie; better_by_objective and
  better_by_feasibility branch pins; stochastic_ranking_select pf=0
  feasibility ordering and count-wraps-modulo-population.
- pareto/crowding.rs: exact interior crowding distance on symmetric
  and asymmetric fronts (pins the (next-prev)/span arithmetic).
- pareto/sort.rs: three-non-dominated-then-one-dominated and a strict
  3-chain producing three singleton fronts.
- pareto/dominance.rs: trade-off → NonDominated, better-on-one-equal-
  on-other → Dominates, identical → Equal.
- pareto/archive.rs: truncate boundary, trade-off kept alongside,
  equal candidate rejected, smaller-violation infeasible eviction.
- pareto/front.rs: best_candidate keeps the first of tied minima.
- metrics/spacing.rs: exact spacing for a varying-NN-distance front.

src/core/problem.rs's lone survivor (decision_schema default body
'replace with vec![]') is an equivalent mutant — Vec::new() and vec![]
are identical — and is left in the residue.
2026-05-13 22:58:17 -06:00
swaits 7b8b7170f4 test(differential_evolution): pin pick_three_distinct and convergence
Phase 1 tests for DE — distinct-index helper and a convergence sanity test.
2026-05-13 22:58:17 -06:00
swaits 44b70c34d5 test(cma_es): pin compare_so / better_than_so and exercise full convergence
Phase 1 tests for src/algorithms/cma_es.rs.

- compare_so: feasibility-first ordering, minimize/maximize inversion,
  infeasibility-violation comparison.
- better_than_so: matches compare_so == Less; equal evaluations are
  not strictly better.
- A 30-generation Sphere1D run pins that CMA-ES actually decreases
  the objective (catches mutants that collapse the update rules).
2026-05-13 22:58:17 -06:00
swaits 8003a97dfa test(bayesian_opt): pin GP / EI / erf helpers
Phase 1 tests for src/algorithms/bayesian_opt.rs. Adds 15 tests
pinning the GP regression and EI acquisition machinery:

- rbf_kernel: signal-variance return at zero distance, exp(-0.5) at
  unit distance, monotone in length scale, decays to 0 for far points.
- normal_pdf: symmetric about zero, value at zero equals 1/sqrt(2π).
- normal_cdf: 0.5 at z=0, symmetric tail sums to 1.
- erf: odd function and erf(0) ≈ 0 within the rational approximation's
  ~1e-7 accuracy.
- expected_improvement: zero at sigma=0, monotone in sigma, positive
  when mu < f_best.
- oriented_target: sign flips under direction, infeasible adds 1e6
  penalty.
- better: feasibility-first then objective ordering under both
  directions.
2026-05-13 22:58:16 -06:00
swaits 3456db6cd8 test(ant_colony_tsp): pin better_than_so branches and build_tour invariants
Phase 1 for src/algorithms/ant_colony_tsp.rs. Targets the ~30 remaining
mutants after Phase 0 sweeps — mostly arithmetic flips in build_tour
(pheromone × heuristic weighting) and the feasibility-comparison logic
in better_than_so.

Added:
- Four branch tests for better_than_so covering the full feasibility
  cross-product: feasible-vs-infeasible (both orders), two-infeasible
  (smaller violation wins), and two-feasible under both directions.
  Plus an equal-objectives test pinning the strict-less-than semantics.
- build_tour-is-a-permutation invariant across 20 seeds × 6 start cities.
- A strong-heuristic test: with eta favoring the next-city by 1000x and
  beta=5, build_tour walks the preferred path. Pins the .powf(beta)
  arithmetic.
- Zero-alpha/zero-beta degenerate-case test: uniform random fallback
  still returns a permutation.
2026-05-13 22:58:16 -06:00
swaits e5e979f02a test(age_moea): pin L_p helpers and add full-run snapshots
Phase 1 for src/algorithms/age_moea.rs. Targets the ~50 algorithm-internal
mutants surviving after the Phase 0 sweeps:

Pure helper-fn pins (lp_norm / lp_distance / nearest_neighbor_distance /
estimate_p):
- L_p norm at p ∈ {1, 2} on canonical vectors (unit, all-ones,
  Pythagorean 3-4-5, signed-via-abs).
- L_p distance: zero-to-itself = 0, symmetry, L_1/L_2 sanity values.
- nearest_neighbor_distance: empty selected → ∞, self-only → ∞, picks
  the closest of mixed-distance candidates.
- estimate_p: empty-front fallback to p=2; axis-aligned extremes (CV=0
  for all p, returns first candidate 0.25); corner-vs-diagonal extremes
  (CV minimized at large p).

Full-run pins:
- A 10-generation Schaffer-N1 run with seed 7 verifies the pareto front
  is non-empty and finite/nonneg (catches  body collapse mutants).
- Population size after run matches config across three pop sizes
  (catches  size mutants).
- Evaluation count falls in [pop, pop*(gens+1)] (catches comparison
  flips in the offspring-collection loop).
2026-05-13 22:58:16 -06:00
swaits f63b85900c test(real): pin exact outputs and algebraic invariants for every real-valued operator
Phase 1 of the mutation-testing campaign for src/operators/real.rs (the
file with the largest mutant surface — 128 missed mutants spread across
GaussianMutation, BoundedGaussianMutation, SBX, PolynomialMutation,
LevyMutation, and the Mantegna gamma/sigma helpers).

Added:
- Seed-pinned numerical snapshots for each operator's vary() output on
  a fixed parent and seed. Any arithmetic flip in the operator's math
  changes one of the snapshot values and fails the assertion. The
  snapshot tolerance is 1e-12 so even subtle FP drift is caught.
- An algebraic-identity test for SBX: c1 + c2 = p1 + p2 per dimension
  before clamping. This identity holds for any β and pins the
  (1+β)·p1 + (1-β)·p2 formula cleanly across 20 seeds.
- A scale-coupling test for PolynomialMutation: a 10× wider bound range
  produces a 10× larger perturbation step at the same seed. Catches any
  mutation that breaks the δ·(hi-lo) coupling.
- Three direct pin tests for mantegna_sigma_u (alpha = 1.5, 1.0, 2.0)
  exercising the gamma() Lanczos series and the formula's edge cases
  (alpha = 1.0 → Cauchy, alpha = 2.0 → Normal-limit where sin(π) ≈ 0).
- A monotonicity property test (sigma_u changes with alpha) to catch
  structural mutants that collapse the formula to a constant.
2026-05-13 22:58:16 -06:00
swaits aa8b6e46e6 test(permutation): pin operator outputs and prove they actually mutate
Phase 1 of the mutation-testing campaign for src/operators/permutation.rs.
Adds 13 tests targeting the 30 surviving mutants in the new permutation
toolkit:

Mutation operators (Inversion / Insertion / Scramble):
- Previously only checked that the output was a valid permutation, which
  passes trivially when the mutant 'replace >= 2 with < 2' skips the
  guard entirely (no mutation = identity output = still a permutation).
  New tests run 30 seeds on an 8-element parent and assert at least one
  seed produces a non-identity output. Kills the >= ↔ < flips.

ShuffledMultisetPermutation::initialize:
- Tightened to assert pop.len() == size up-front, killing the 'replace
  with vec![]' mutant.

Crossover operators (OX / PMX / CX / ERX):
- 'Recombines for n >= 3' tests: with 5-element distinct parents, some
  seed must produce a child differing from both parents. Kills the
  < ↔ > / == / <= guard flips that would early-return parents at n >= 3.
- CX-specific pinned tests: the single-cycle case (children = parents)
  and the two-cycle case (exactly known output). Pins the cycle-detection
  arithmetic and the parent-alternation logic — kills the ==↔!= and
  += ↔ *= mutants inside cx_child.
- ERX: 'distinct starts can yield distinct children' across 30 seeds —
  kills the prev/next-index arithmetic mutants in the adjacency table.

Some residual mutants in this file are equivalent (e.g., < ↔ <= when
n=2 still produces the same OX result because for length-2 inputs the
segment-and-fill recombination converges to the parents anyway).
Documented in test comments.
2026-05-13 20:25:43 -06:00
swaits 47cb1d5f79 test(repair): pin ProjectToSimplex argmax tie-breaking and position
The degenerate-magnitude shortcut in ProjectToSimplex::repair scans
`decision` for the argmax and concentrates all mass there. cargo
mutants found that the strict-greater scan was unpinned: replacing
`>` with `>=` (which would shift the argmax to the last tied
index) and `>` with `==` (which would silently skip larger
values further along) both survived.

Two tests:
- A 3-element vector with two tied maxima at the front pins that the
  scan keeps the first index on a tie.
- A 3-element vector whose argmax is at index 1 pins that the scan
  actually walks past the start when later values are larger.

Other mutants in this file (the `>` ↔ `>=` threshold check at line
106, the `*` ↔ `+` in the threshold constant, the `-` ↔ `+` /
`/` in the tau-fallback initializer, and the `>` ↔ `>=` in the
projection loop's rho update) are equivalent mutants for non-pathological
inputs: the normal-path and shortcut-path math converge to the same
projection result for any input the operator is documented to handle.
Leaving them in the residue.
2026-05-13 20:22:40 -06:00
swaits c94357abe3 test(explorer): pin ExplorerExport, builders, ToDecisionValues outputs
Phase 0.3 of the mutation-testing campaign. Extends the inline tests in
src/explorer/mod.rs with 24 new tests covering the gaps cargo-mutants
identified — about 30 surviving mutants in this one file.

Coverage added:
- Exact-output tests for ToDecisionValues impls on Vec<f64>, Vec<i64>,
  Vec<bool>, Vec<usize> (the previous tests only asserted lengths or
  spot-checked individual entries).
- from_result propagates evaluations / generations from the
  OptimizationResult into RunMeta.
- with_problem_name / with_wall_clock / with_timestamp each set their
  field and preserve the rest of the export.
- to_json emits a JSON containing schema_version, candidates, problem
  name, and algorithm name strings.
- to_writer and to_file round-trip the same bytes.
- Free top-level to_json / to_writer / to_file convenience functions
  exercised end-to-end (round-trip through tmp file).
- pad_decision_schema at all three boundaries (< target / == target /
  > target) to pin the < comparison.
- candidate_to_export's in_pareto_front toggles at front_rank == 0.
- candidate_to_export's feasible toggles at constraint_violation <= 0.
- candidate_to_export pads short objective vectors and truncates long
  ones (the defensive branch).
2026-05-13 20:22:40 -06:00
swaits 3085359d01 test(async): parity sweep — run_async must match run with same seed
Phase 0.2 of the mutation-testing campaign. Adds a module gated on
#[cfg(feature = "async")] that, for every algorithm with a run_async,
asserts that the async runner produces the same result as the sync
runner given the same Config + seed + problem.

Before: nothing exercised run_async, so cargo mutants survived
'replace run_async body with OptimizationResult::new()' and every
comparison/arithmetic mutant inside the async loop for every
async-capable algorithm — about 25-30 algorithms * 5-10 mutants each.
After: every such mutant is killed because the parity test detects
any divergence in best.evaluation.objectives or pareto-front
objective tuples.

Coverage:
- Single-objective real (Sphere1D fixture): RandomSearch,
  HillClimber, OnePlusOneEs, SimulatedAnnealing, GA, PSO, DE, CmaEs,
  IpopCmaEs, sNES, TLBO, NelderMead, BayesianOpt, TPE.
- Multi-objective real (SchafferN1 fixture): NSGA-II/III, SPEA2,
  MOEA/D, MOPSO, IBEA, SMS-EMOA, HypE, PESA-II, ε-MOEA, AGE-MOEA,
  GrEA, KnEA, RVEA, PAES.
- Binary (OneMax): UMDA.
- Permutation (TinyTsp fixture): AntColonyTsp.
- Integer (AbsInt fixture): TabuSearch.
- Multi-fidelity (Sphere1DPartial fixture): Hyperband.

Run with: cargo test --features async --test algorithm_properties async_parity
2026-05-13 19:47:33 -06:00
swaits a773a1eaf6 test(algorithm_info): pin name/full_name/seed for every algorithm
Phase 0.1 of the mutation-testing campaign: a sweep test per algorithm
(33 total) asserting the exact strings returned by AlgorithmInfo::name()
and AlgorithmInfo::full_name() plus the seed propagated through
AlgorithmInfo::seed().

Before: cargo mutants survived dozens of mutants per algorithm replacing
the name/full_name return values with "" or "xyzzy", and the seed
return with None/Some(0)/Some(1). After: every such mutant is caught
by an exact-equality assertion.

NelderMead is deterministic and has no seed override (intentionally);
its test asserts seed() == None to pin the default-trait-impl behavior.
2026-05-13 19:43:17 -06:00
swaits 6368db74bb docs(book): permutation toolkit and multi-objective combinatorial cookbook
- Rewrites cookbook/permutation.md to cover the new operator toolkit:
  initializers, crossovers (OX/PMX/CX/ERX), mutations, a 'what should I
  use' picker, and a worked GA-on-TSP example. JSS multiset section
  explains why the strict-permutation crossovers don't compose with
  operation-string encodings and shows the local POX pattern.

- Adds cookbook/multi-objective-combinatorial.md: bi-objective TSP via
  NSGA-II, bi-objective knapsack (binary encoding), 3-objective JSS
  via NSGA-III, and a hypervolume-based operator comparison.

- Updates choosing-an-algorithm.md to reference the new operators in
  the single- and multi-objective decision tables, plus a noting
  NSGA-II/III's genericity over Vec<usize> and Vec<bool> decisions.

- SUMMARY.md and cookbook.md updated to list the new recipe.
2026-05-13 19:43:17 -06:00
swaits ab545e268a docs(examples): add bi-objective TSP crossover-comparison demo
tsp_operators_compare.rs runs NSGA-II four times on the KroAB-25
bi-objective TSP, holding everything constant except the crossover
operator. Ranks OX, PMX, CX, ERX by hypervolume against a fixed
reference point, plus front size, unique-point count, and runtime.

Pedagogical demonstration that the right comparison metric for a
Pareto search is hypervolume, not single-objective fitness.
2026-05-13 19:33:51 -06:00