Files
heuropt/docs/book/src/choosing-an-algorithm.md
T
swaits fa3f2e8fb0 feat: v0.5.0 — comprehensive documentation release
Theme: documentation and project polish. No public-API changes; this
is the v0.5 release that elevates heuropt's docs/onboarding/governance
to bar-setting status.

Adds:
- mdbook user guide at docs/book/ with intro, getting-started,
  defining-problems, choosing-an-algorithm, cookbook (7 recipes),
  comparison vs other libraries, stability/SemVer, migration guides.
  Deploys to https://swaits.github.io/heuropt/ via .github/workflows/
  docs.yml.
- Runnable rustdoc examples on every algorithm (35 of them), all
  exercised by cargo test --doc.
- Three real-world examples: portfolio.rs (multi-obj with budget
  constraint), hyperparam_tuning.rs (BO + TPE), scheduling.rs
  (permutation via SA + SwapMutation against Smith's-rule oracle).
- Governance: CONTRIBUTING.md, SECURITY.md, CODE_OF_CONDUCT.md
  (adopting builderscode.org's Builder's Code of Conduct), GitHub
  issue templates, PR template.

Polishes:
- README hero with badges + user-guide link.
- lib.rs crate-level docs.
- CHANGELOG entry for 0.5.0.

Bumps Cargo.toml to 0.5.0.
2026-05-05 14:33:12 -06:00

13 KiB
Raw Blame History

Choosing an algorithm

The README has a compact decision tree. This chapter expands it with the reasoning behind each branch.

Step 0: How expensive is one evaluation?

This is the first fork because it changes everything that comes after it.

Eval cost Budget you can afford Algorithm family
Microseconds (pure math) 10 000 1 000 000 evals Population-based
Milliseconds (sim, IO) 1 000 10 000 evals Population-based
Seconds (small training) 100 1 000 evals Sample-efficient (BO, TPE)
Minutes+ (full training) 50 500 evals Sample-efficient + multi-fidelity

For the cheap-eval branch, you have the run of the catalog. For the expensive branch, classical evolutionary methods waste your evaluation budget — go to BayesianOpt or Tpe. For the very expensive branch where each eval has a tunable budget (epochs, MC samples, sim steps), Hyperband over the PartialProblem trait is the move.

Step 1: How many objectives?

The biggest fork.

  • One — there's a single best answer. Pick from the single-objective branch.
  • Two or three — a Pareto front. Pick from the multi-objective branch.
  • Four or more — a many-objective Pareto front; classical multi-objective methods break down here because almost every pair of points is non-dominated. Pick from the many-objective branch.

Pareto front: the set of decisions where you cannot improve any objective without sacrificing another. In a 2-objective minimize problem, plot every solution; the Pareto front is the lower-left envelope.

If you found yourself staring at a single composite score that's a weighted sum of conflicting goals, you probably have a multi-objective problem in disguise. A weighted sum bakes in your preferences before you've seen the trade-off; running a multi-objective optimizer first and picking off the front later is almost always a better workflow (see Pick one answer off a Pareto front).

Step 2 — single-objective continuous

These all take Vec<f64> decisions.

Smooth, low-to-moderate dimension

CmaEs is the strong default. It adapts the search distribution's covariance to the local landscape. On the comparison harness it hits machine epsilon on Rosenbrock at 30 000 evaluations.

For very low-dimensional smooth problems (≤ 5 dim), NelderMead is deterministic and converges to f = 0 exactly on Rosenbrock.

High dimension, smooth

SeparableNes uses a diagonal covariance — cheaper per step than CmaEs at the cost of being unable to model rotated landscapes. Worth trying when CmaEs's O(d²) per-step cost hurts.

Multimodal landscapes

Multimodal = many local minima that aren't the global one. Rastrigin and Ackley are classic traps.

IpopCmaEs is CmaEs with an increasing-population restart strategy specifically designed for this. On the harness it drops vanilla CmaEs's Rastrigin score from f = 2.35 to f = 0.13.

DifferentialEvolution is rarely beaten on cheap multimodal continuous problems. On Rastrigin it ties with (1+1)-ES at f = 0.

SimulatedAnnealing is a cheap, generic baseline that escapes local optima via temperature decay.

Want parameter-free

Tlbo (Teaching-Learning-Based Optimization) has no F, CR, w, or σ to tune. Often a respectable middle-of-the-pack performer.

Smallest possible self-adapting baseline

OnePlusOneEs — Rechenberg's 1973 (1+1)-ES with the one-fifth success rule. On the harness it hits f = 0 on Rastrigin in 50 000 evaluations.

Just want a baseline

RandomSearch. Useful as a sanity check: if your fancy optimizer can't beat random search, something is wrong (with the fancy optimizer or with the problem).

Step 2 — single-objective other types

Decision type Algorithm Notes
Vec<bool> Umda Per-bit marginal EDA. Independent-bit assumption.
Vec<bool> GeneticAlgorithm + BitFlipMutation When bit interactions matter.
Vec<usize> (permutation) AntColonyTsp TSP-style with a distance matrix.
Vec<usize> (permutation) SimulatedAnnealing + SwapMutation Generic discrete baseline.
Vec<usize> or custom TabuSearch You supply the neighbor function.
Custom struct SimulatedAnnealing / HillClimber With your own Variation impl.

Step 2 — multi-objective (2 or 3)

Strong default

Nsga2 is the canonical Pareto-based EA. Fast, well-understood, maintains diversity via crowding distance. On the harness it lands on the Pareto front of every test problem.

Real-valued, smooth front, want best convergence

Mopso (multi-objective PSO with archive). On ZDT1 it wins hypervolume outright and converges 100× tighter than the dominance-based methods.

Better front quality than NSGA-II

Ibea (indicator-based) is consistently the best of the dominance-based methods on the harness — wins ZDT3 hypervolume and DTLZ2 mean distance by 24×. It uses an additive ε-indicator for selection rather than dominance + crowding.

Spea2 (strength + density) — solid alternative; explicit external archive separate from the population.

SmsEmoa uses exact hypervolume contribution for selection. Elegant in theory; in practice on the harness budgets here it underperforms NSGA-II. Worth the higher per-step cost only when exact HV contribution is the right discriminator.

Decomposition / weight-vector style

Moead decomposes the multi-objective problem into many scalar sub-problems (Tchebycheff or weighted sum) and solves them in parallel. Very fast per generation; scales naturally to many objectives.

Disconnected or non-convex front

AgeMoea estimates the front geometry adaptively (the L_p parameter p is fit from data each generation).

Knea favors knee points — the regions of the front where small gains in one objective cost large losses in another.

Ibea also handles disconnected fronts well.

Region-based diversity

PesaII uses grid hyperboxes to drive selection — divide the objective space into a grid, pick from the least-crowded boxes.

EpsilonMoea uses an ε-grid archive that auto-limits its size.

Just one starting decision (no population budget)

Paes(1+1)-ES with a Pareto archive. Cheap, simple, useful when your evaluations are expensive enough that you can't afford a population.

Step 2 — many-objective (4+)

Linear / simplex-shaped front (e.g., DTLZ1)

Grea — grid coords drive ranking. On DTLZ1 it beats NSGA-III by 3× and AGE-MOEA by 2.5×.

Moead — decomposition shines on linear fronts; second on DTLZ1 and among the fastest per generation.

Curved / unknown front geometry

Nsga3 — reference-point niching; canonical many-objective method; strong default when the front isn't simplex-shaped.

AgeMoea — estimates L_p geometry per generation.

Rvea — reference vectors with adaptive penalty.

Indicator-based selection

Ibea — additive ε-indicator; doesn't degrade at high obj count.

HypE — Monte Carlo hypervolume estimation; scales to arbitrary objective count where exact HV is too expensive.

Step 3: Are there hard constraints?

heuropt models constraints as a single scalar constraint_violation on each Evaluation. Three escalations when the feasibility region is hard to find:

  1. Penalty-only. Just set constraint_violation > 0 for infeasible decisions. The default tournament/Pareto comparisons prefer feasibles automatically.
  2. Repair. Implement Repair<D> (or use the provided ClampToBounds / ProjectToSimplex) to project infeasible decisions back into the feasible region. Pair with a Variation in a CompositeVariation for bounds-aware variants.
  3. Stochastic ranking. Use stochastic_ranking_select instead of tournament_select_single_objective. It probabilistically explores near-feasibility instead of strict feasibility-first ordering, which helps when feasible regions are narrow.

See Constrain your search with Repair for worked examples.

Step 4: Should you parallelize?

Enable the parallel feature flag if your evaluate takes more than ~50 µs. Population-based algorithms (RandomSearch, Nsga2, DifferentialEvolution, Spea2, Ibea, Mopso, …) batch- evaluate via rayon when the feature is on. Seeded runs stay bit-identical to serial mode.

heuropt = { version = "0.5", features = ["parallel"] }

TL;DR table

Situation Pick
Smooth single-objective continuous CmaEs
Multimodal single-objective continuous IpopCmaEs or DifferentialEvolution
Expensive single-objective BayesianOpt or Tpe
Multi-fidelity single-objective Hyperband
2- or 3-objective default Nsga2
2-objective real-valued smooth front Mopso
Disconnected / non-convex front Ibea
Many-objective default (curved front) Nsga3
Many-objective linear / simplex front Grea
Permutation problem AntColonyTsp
Binary problem Umda
Custom decision type SimulatedAnnealing + your Variation
Sanity baseline RandomSearch