`pareto_front` re-scanned every candidate from scratch. Track a
`dominated` bitset instead: whenever `i`'s scan finds `i` dominates `j`,
mark `j` so the outer loop skips `j` outright when it reaches it. The
inner check now reads both directions of `pareto_compare` — the
objective scan already computes both flags, so this is ~free.
Bit-identical, including under NaN-intransitive dominance: a mark is
only ever set from a direct pairwise `pareto_compare` result, never
inferred transitively. Also strict-or-neutral on work — the marks can
only ever let the outer loop *skip*, never add a scan.
Whole-program callgrind Ir for the compare_profile benchmark:
169,440,644,233 -> 165,311,562,939 (-2.44%); `pareto_front` self-Ir
44.1B -> 39.9B.
`pareto_front` was still the naive O(n²) formulation: a raw double loop
calling `pareto_compare` for every ordered pair, re-deriving feasibility
and the minimization-oriented objective values on every comparison.
Apply the same precompute pattern `non_dominated_sort` already uses:
hoist per-individual `feasible` / `violation` / `oriented` (a flat n*m
buffer) out of the loop, then run a branchless inlined dominance check
over the contiguous buffer. The kept set and its order are unchanged.
Whole-program callgrind Ir for the `compare_profile` benchmark:
221,836,742,708 -> 173,803,642,945 (-21.65%); `pareto_front` self-Ir
92.1B -> 44.1B.
The objective-comparison branch of `pareto_compare` materialized two
`Vec<f64>`s per call via `ObjectiveSpace::as_minimization`. Because
`pareto_compare` runs O(n²) times across the multi-objective algorithms,
that per-call allocation pair dominated the whole `compare` workload.
Replace it with an allocation-free per-objective scan that branches on
`Objective::direction` directly: for a Maximize axis "a beats b" is just
`av > bv`, bit-identical to `-av < -bv` after orientation. The result is
unchanged for every input.
Whole-program callgrind Ir for the `compare_profile` benchmark:
357,060,633,544 -> 221,836,742,708 (-37.87%).
The O(n^2) pair loop reads oriented[j] for every j; with Vec<Vec<f64>>
that chased a separate heap allocation per individual. A flat n*m buffer
keeps those reads contiguous and sequential in j.
non_dominated_sort_2d n=200: 1_288_072 -> 1_096_738 (-15%, 1.17x); n=50
-19%. nsga2 one-generation -5.7%. Combined with the earlier antisymmetry
fix, n=200 is down 55% from the original 2.46M. Pure data-layout change --
output bit-identical, all 606 tests pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
crowding_distance sorted bare front indices with a comparator that chased
two Vec<Vec<f64>> indirections per comparison. Extracting (objective value,
front position) tuples into a buffer reused across objectives keeps the hot
comparator a single f64 compare.
crowding_distance_2d n=200: 181_493 -> 173_286 (-4.5%); n=50 -4.5%. Stable
sort over the (value, index) pairs preserves the original tie-order, so the
output is bit-identical -- all 606 tests pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The dominance relation is antisymmetric, so the outcome of compare(i, j)
fully determines compare(j, i). Iterating only j > i and applying the
result in both directions does identical work in half the pair scans.
non_dominated_sort_2d n=200: 2_461_178 -> 1_268_372 (-48%); n=50 -1.65x.
Ripples into dependents: nsga2 one-generation -20%, nsga3 / sms_emoa ~-9%.
Output is bit-identical (dominates[] still ascending, first_front order
unchanged) -- all 606 tests including run() snapshots pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The per-file Phase 1 test commits were written without running rustfmt
as I went; this pass formats the new test code (long assert_eq! lines
wrapped, etc.). Formatting-only — no behavioural change.
Phase 1, tier 3 of the mutation-testing campaign — the shared Pareto /
metric / selection utilities used by every multi-objective algorithm.
A scoped cargo-mutants run found 75 survivors across these files; the
tests below target them.
- metrics/hypervolume.rs: dominates() boundary cases, non_dominated_
projection retained-set pins, hso_recursive 1-D/2-D base cases,
hypervolume_nd_from_evaluations empty/non-dominating skips.
- selection/tournament.rs: challenger_wins across the full feasibility
cross-product + equal-objective tie; better_by_objective and
better_by_feasibility branch pins; stochastic_ranking_select pf=0
feasibility ordering and count-wraps-modulo-population.
- pareto/crowding.rs: exact interior crowding distance on symmetric
and asymmetric fronts (pins the (next-prev)/span arithmetic).
- pareto/sort.rs: three-non-dominated-then-one-dominated and a strict
3-chain producing three singleton fronts.
- pareto/dominance.rs: trade-off → NonDominated, better-on-one-equal-
on-other → Dominates, identical → Equal.
- pareto/archive.rs: truncate boundary, trade-off kept alongside,
equal candidate rejected, smaller-violation infeasible eviction.
- pareto/front.rs: best_candidate keeps the first of tied minima.
- metrics/spacing.rs: exact spacing for a varying-NN-distance front.
src/core/problem.rs's lone survivor (decision_schema default body
'replace with vec![]') is an equivalent mutant — Vec::new() and vec![]
are identical — and is left in the residue.
Completes the rustdoc audit — every public item now has at least one
```rust example block in its docstring, exercised by
`cargo test --doc` (55 doctests, all passing).
- Operators: BitFlipMutation, SwapMutation, RealBounds,
GaussianMutation, BoundedGaussianMutation,
SimulatedBinaryCrossover, PolynomialMutation, LevyMutation,
ClampToBounds, ProjectToSimplex.
- Metrics: hypervolume_2d, hypervolume_nd, spacing.
- Pareto utilities: pareto_compare, pareto_front, best_candidate,
non_dominated_sort, crowding_distance, das_dennis,
ParetoArchive.
Each example is short (5-15 lines) and self-contained — copy-paste
into a fresh project and it runs.
`ParetoArchive::insert` calls `pareto_compare` twice per existing
member (once per pass), and each call re-allocates two Vec<f64>s
via `as_minimization` — 4N allocations per insert. Cache the
candidate's oriented + feasibility/violation once, build each
member's oriented vector once for the call, then inline the
dominance test against those cached arrays.
Used by PESA-II (per offspring per generation), PAES (per child),
ε-MOEA, and any user code working through the archive directly.
Wall-clock (compare harness, 10-seed mean):
- PESA-II / DTLZ2: 498 → 426 ms (-14 %)
- PESA-II / ZDT1: 87 → 75 ms (-14 %)
Smaller wins on PAES / MOPSO / IBEA / HypE / ε-MOEA where the
archive isn't the dominant per-generation cost.
Bit-identical via the compare harness.
The Deb fast non-dominated sort calls `pareto_compare` twice for
every (i, j) pair, and each `pareto_compare` call invokes
`ObjectiveSpace::as_minimization` twice — so for an N-point
population that's 4·N·(N-1) fresh `Vec<f64>` allocations per sort.
At N=100 with thousands of generations across the compare harness,
this dominated the per-generation cost of every Pareto-based MOEA.
Cache `as_minimization`/feasibility/violation once per individual
up front, then inline the dominance test against those cached
arrays. The output (per-pair dominance outcome and the per-i
`dominates` lists) is bit-identical to `pareto_compare`.
gungraun (instructions):
- non_dominated_sort_2d n=50: 852 317 → 198 574 (-77 %, 4.3×)
- non_dominated_sort_2d n=200: 13 513 271 → 2 601 813 (-81 %, 5.2×)
Wall-clock (compare harness, 10-seed mean):
- NSGA-II / ZDT1: 268 → 65 ms (4.1×)
- NSGA-II / ZDT3: 267 → 65 ms (4.1×)
- NSGA-II / DTLZ2: 344 → 106 ms (3.2×)
- NSGA-II / Rastrigin: 260 → 71 ms (3.7×)
- NSGA-III / DTLZ2: 318 → 122 ms (2.6×)
- NSGA-III / DTLZ1: 303 → 122 ms (2.5×)
- SMS-EMOA / DTLZ2: 1413 → 1369 ms (small additional win on top of HV)
- AGE-MOEA / DTLZ1: 430 → 229 ms (1.9×, on top of the AGE-MOEA caching)
- HypE / DTLZ2: 80 → 44 ms (1.8×)
The standard structured weight/reference vector generator for
many-objective MOEAs (NSGA-III, MOEA/D). Generates (H+M-1 choose M-1)
points uniformly distributed on the unit simplex by enumerating all
integer compositions of `divisions` into `num_objectives` parts and
dividing each by `divisions`.
Lives in src/pareto/reference_points.rs. Re-exported from the prelude
as `das_dennis`.
Tests cover: M=2/H=4 → 5 points along the diagonal; M=3/H=12 → 91
points (the canonical NSGA-III 3-objective ref set); each generated
point has exactly M components summing to 1 within float tolerance.
- pareto/crowding.rs: rewrite the inner loop to iterate per-objective
via index_axis-style indexing on `oriented` rather than naming an
unused loop variable `k`.
- operators/{binary,permutation}.rs tests: pass parents via
`std::slice::from_ref` instead of `&[parent.clone()]` to avoid the
cloned_ref_to_slice_refs lint.
Pure cleanup — no behavior change, all 83 unit tests + 2 doctests still
pass.
A concrete archive (not a trait — spec §13). On insert it discards the
new candidate if any existing member dominates it, then removes existing
members the new candidate dominates. `truncate` does simple
tail-truncation in v1; the doc note flags that crowding-aware
truncation is a future improvement.
Computes per-point crowding distance over a single Pareto front (spec
§9.6):
- Returns Vec<f64> with the same length as the front index slice.
- Empty front → empty Vec.
- Front of length ≤ 2 → all f64::INFINITY.
- Boundary points along each objective receive INFINITY.
- Interior points get sum of normalized neighbor gaps; if max == min for
an objective the contribution is zero.
- Operates on minimization-oriented objective values.
Deb's fast non-dominated sort: returns Vec<Vec<usize>> of front indices
into the input population, with fronts[0] being the non-dominated set.
O(N²·M) is acceptable for v1 (spec §9.5).
Tests cover: small known population produces expected fronts; equal
candidates land on the same front; an empty population yields no
fronts.
`pareto_front` returns all candidates not dominated by any other in
input order (O(N²·M), acceptable for v1 per spec §9.3).
`best_candidate` is the single-objective "best" finder: returns None
unless there is exactly one objective; ignores infeasibles; returns None
if every candidate is infeasible (spec §9.4).
Both re-exported from the prelude.