better_optimize is a friendlier front-end to scipy's optimize.minimize and optimize.root functions. Features
include:
- Progress bar!
- Early stopping!
- Better propagation of common arguments (
maxiters,tol)! - A typed, documented configuration object for every method, so swapping optimizers is a one-argument change!
To install better_optimize, simply use conda:
conda install -c conda-forge better_optimizeOr, if you prefer pip:
pip install better_optimizeAll optimization routines in better_optimize can display a rich, informative progress bar using the rich library. This includes:
- Iteration counts, elapsed time, and objective values.
- Gradient and Hessian norms (when available).
- Separate progress bars for global (basinhopping) and local (minimizer) steps.
- Toggleable display for headless or script environments.
- No more nested
optionsdictionaries! You can passtol,maxiter, and other common options directly as top-level keyword arguments. better_optimizeautomatically sorts and promotes these arguments to the correct place for each optimizer.- Generalizes argument handling: always provides
tolandmaxiter(or their equivalents) to the optimizer, even if you forget.rootmethods disagree about whether that budget is calledmaxiter,maxfevormaxfun, and you can passmaxiterto any of them. - Seven of the ten
rootmethods take part of their configuration nested insidejac_options, which scipy splats into the jacobian it builds. Pass it as a mapping or as the typed object for that jacobian (BroydenJacOptions,KrylovJacOptions, and so on). An option that belongs there raises if you pass it at the top level, where scipy would warn and drop it.
Every method scipy supports has a dataclass whose fields are exactly the options that method accepts, typed, defaulted, and documented with scipy's own prose. help(BFGSConfig) tells you what c2 does without a trip to scipy.optimize.show_options.
- Pass one wherever a method name works:
minimize(f, x0, method=BFGSConfig(gtol=1e-8)). The string form keeps working unchanged. - A misspelled option raises
TypeErrornaming it, rather than reaching scipy and being dropped with a warning nobody reads. - Fields and defaults are checked against scipy's own function signatures by test, so a scipy release that moves a default fails a test instead of drifting quietly.
- One
minimizecall reaches local minimizers, basinhopping and differential evolution alike, because the configuration carries which solver runs it.
- Automatic checking of provided gradient (
jac), Hessian (hess), and Hessian-vector (hessp) functions. - Warns if you provide unnecessary or unused arguments for a given method.
- Detects and handles fused objective functions (e.g., functions returning
(loss, grad)or(loss, grad, hess)tuples). - Ensures that the correct function signatures and return types are used for each optimizer.
- Provides an
LRUCache1utility to cache the results of expensive objective/gradient/Hessian computations. - Especially useful for triple-fused functions that return value, gradient, and Hessian together, avoiding redundant computation.
- Totally invisible -- just pass a function with 3 return values. Seamlessly integrated into the optimization workflow.
- Enhanced
basinhoppingimplementation allows you to continue even if the local minimizer fails. - Optionally accepts and stores failed minimizer results if they improve the global minimum.
- Useful for noisy or non-smooth objective functions where local minimization may occasionally fail.
scipy passes callbacks a different argument for almost every method: callback(xk), callback(xk, state),
callback(intermediate_result), or callback(x, f) for root finding. better_optimize normalizes them to a single
callback(res), where res is an OptimizeResult with the current res.x, res.fun, and res.nit -- plus res.jac
when a gradient is available, and solver-specific fields like res.accept (basinhopping) or res.convergence
(differential evolution).
- The return value is ignored, so a callback can report a value (an ELBO, a logged loss) without stopping the run.
- Raise
StopOptimizationto stop early; the best result so far is returned withsuccess=False. - The same callback works across
minimize,root,basinhopping, anddifferential_evolution.hybrandlmare the exception -- scipy never calls a callback for them, so it is ignored with a warning.
import numpy as np
from better_optimize import minimize
from better_optimize.configuration import LBFGSBConfig
def rosenbrock(x):
return sum(100.0*(x[1:] - x[:-1]**2.0)**2.0 + (1 - x[:-1])**2.0)
result = minimize(
rosenbrock,
x0=np.array([-1.0, 2.0]),
method=LBFGSBConfig(tol=1e-6, maxiter=1000),
progressbar=True, # Show a rich progress bar!
) Minimizing Elapsed Iteration Objective ||grad||
──────────────────────────────────────────────────────────────────────────────────────────────────
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 0:00:00 721/721 0.34271757 0.92457651The result object is a standard OptimizeResult from scipy.optimize, so there are no surprises there!
A configuration object says which solver runs it, so trying a different optimizer family is a change to one argument rather than a rewrite. Nothing below changes but the loop variable.
import numpy as np
from better_optimize import minimize
from better_optimize.configuration import (
BasinHoppingConfig,
DifferentialEvolutionConfig,
LBFGSBConfig,
)
def rosenbrock(x):
return sum(100.0*(x[1:] - x[:-1]**2.0)**2.0 + (1 - x[:-1])**2.0)
x0 = np.array([-1.0, 2.0])
bounds = [(-5.0, 5.0), (-5.0, 5.0)]
for config in [
LBFGSBConfig(maxiter=500),
BasinHoppingConfig(niter=25, minimizer_config=LBFGSBConfig()),
DifferentialEvolutionConfig(popsize=20, rng=0),
]:
result = minimize(rosenbrock, x0, method=config, bounds=bounds, progressbar=False)
print(f"{config.method_name:25s} {result.fun:.3e}")L-BFGS-B 9.191e-12
basinhopping 3.696e-12
differential_evolution 4.980e-30L-BFGS-B is a local minimizer, basinhopping restarts one from perturbed points, and differential_evolution is a population method that never calls the local minimizer at all. They take different arguments, and scipy reaches them through different entry points. The configuration object is what absorbs that difference.
For users who have already memorized every combination of algorithm and option, none of the above is required. Method names still work, and so do flat keywords -- including the ones scipy keeps inside its options dictionary, which better_optimize gathers up for you. If you memorized that part too, pass options yourself and it is merged and checked the same way.
import numpy as np
from better_optimize import minimize
from better_optimize.configuration import LBFGSBConfig
def rosenbrock(x):
return sum(100.0*(x[1:] - x[:-1]**2.0)**2.0 + (1 - x[:-1])**2.0)
x0 = np.array([-1.0, 2.0])
# A configuration object
minimize(rosenbrock, x0, method=LBFGSBConfig(ftol=1e-6, gtol=1e-6, maxiter=1000, maxfun=1000))
# Flat keywords, including ftol and gtol, which scipy keeps in `options`
minimize(rosenbrock, x0, method="L-BFGS-B", ftol=1e-6, gtol=1e-6, maxiter=1000)
# Or the options dictionary itself
minimize(rosenbrock, x0, method="L-BFGS-B",
options={"ftol": 1e-6, "gtol": 1e-6, "maxiter": 1000, "maxfun": 1000})Those three are the same call and return the same answer. The flat form additionally spreads a top-level maxiter across every budget the method has, which for L-BFGS-B means maxfun -- that is why the first and third forms name it and the second does not.
All three are checked the same way, which is the part that is not backwards compatible. A misspelled option raises TypeError before the run starts, with a suggestion:
minimize(rosenbrock, x0, method="L-BFGS-B", options={"gtoll": 1e-6})
# TypeError: LBFGSBConfig.__init__() got an unexpected keyword argument 'gtoll'. Did you mean 'gtol'?scipy would have accepted that call, run to completion, and mentioned the dropped option in an OptimizeWarning afterwards.
Pass a callback and it runs after each iteration with an OptimizeResult. Its return value is ignored; raise
StopOptimization to stop and return the best result so far.
import numpy as np
from better_optimize import minimize, StopOptimization
from better_optimize.configuration import LBFGSBConfig
def rosenbrock(x):
return sum(100.0*(x[1:] - x[:-1]**2.0)**2.0 + (1 - x[:-1])**2.0)
history = []
def callback(res):
history.append(res.fun) # res.x, res.fun, res.nit; res.jac when available
if res.fun < 1e-8:
raise StopOptimization
result = minimize(
rosenbrock,
x0=np.array([-1.0, 2.0]),
method=LBFGSBConfig(),
callback=callback,
)The same callback works for root (res.fun is the residual vector), basinhopping (res.accept), and
differential_evolution (res.convergence).
from better_optimize import minimize
from better_optimize.configuration import NewtonCGConfig
import pytensor.tensor as pt
from pytensor import function
import numpy as np
x = pt.vector('x')
value = pt.sum(100.0*(x[1:] - x[:-1]**2.0)**2.0 + (1 - x[:-1])**2.0)
grad = pt.grad(value, x)
hess = pt.hessian(value, x)
fused_fn = function([x], [value, grad, hess])
x0 = np.array([1.3, 0.7, 0.8, 1.9, 1.2])
result = minimize(
fused_fn, # No need to set flags separately, `better_optimize` handles it!
x0=x0,
method=NewtonCGConfig(tol=1e-6, maxiter=1000),
progressbar=True, # Show a rich progress bar!
)Many sub-computations are repeated between the objective, gradient, and hessian functions. Scipy allows you to pass a
fused value_and_grad function, but better_optimize also lets you pass a triple-fused value_grad_and_hess function.
This avoids redundant computation and speeds up the optimization process.
differential_evolution searches a bounded region with a population instead of descending from a starting point, which suits objectives with many local minima and no useful gradient. It takes bounds rather than x0, and gets the same progress bar and callback treatment as everything else.
import numpy as np
from better_optimize import differential_evolution
def rastrigin(x):
return 10 * len(x) + sum(x**2 - 10 * np.cos(2 * np.pi * x))
result = differential_evolution(
rastrigin,
bounds=[(-5.12, 5.12)] * 5,
popsize=20,
rng=0,
)
print(result.x) # [-0. -0. -0. -0. -0.]
print(result.fun) # 0.0Rastrigin has a local minimum at roughly every integer lattice point, so a gradient method started anywhere but the middle converges to the wrong one.
It is often quicker to get close with a robust method and finish with a fast one. sequential_optimize runs a list of stages, forwarding each stage's best point into the next stage's starting point, and reports what every stage did.
import numpy as np
from better_optimize import sequential_optimize
from better_optimize.configuration import BFGSConfig, NelderMeadConfig
def rosenbrock(x):
return sum(100.0*(x[1:] - x[:-1]**2.0)**2.0 + (1 - x[:-1])**2.0)
result = sequential_optimize(
rosenbrock,
x0=np.array([-1.2, 1.0]),
stages=[NelderMeadConfig(maxiter=200), BFGSConfig(gtol=1e-10)],
)
print(result.x) # [0.99999552 0.99999104]
print(result.best_stage) # 1
result.summary() # Rich table of every stage, with the winner starredA stage can also be a dict carrying its own solver and keyword arguments, which is what you need when a stage wants its own bounds, constraints, or starting point. By default the chain stops at the first stage that fails. Pass on_failure="continue" to run the rest anyway and keep the best result from those that worked.
Real-world objectives often have multiple local minima. A common workaround is to throw many random starting points at
the optimizer and keep the best result. better_optimize makes this painless with multi_optimize:
import numpy as np
from better_optimize import minimize, multi_optimize
from better_optimize.configuration import LBFGSBConfig
def rosenbrock(x):
return sum(100.0*(x[1:] - x[:-1]**2.0)**2.0 + (1 - x[:-1])**2.0)
result = multi_optimize(
solver=minimize,
solver_kwargs=dict(f=rosenbrock, method=LBFGSBConfig(tol=1e-10)),
x0=np.zeros(5),
n_runs=16,
init_strategy="uniform",
bounds=(-5, 5),
backend="loky",
n_jobs=-1,
seed=42,
progressbar=True,
)
print(result.best) # Best OptimizeResult
print(result.x_best) # Best parameter vector
print(result.fun_best) # Best objective value
result.summary() # Rich table of all runs, rankedThe
loky(process) backend needs anif __name__ == "__main__":guard when run as a script on macOS or Windows. Notebooks and thethreading/sequentialbackends don't.
multi_optimize works with any solver that follows the (x0, **kwargs) -> OptimizeResult signature -- that includes
minimize, root, basinhopping, or your own custom wrapper. It just calls solver(x0=x0_i, **solver_kwargs) for
each starting point; it never inspects the solver internals.
A few highlights:
- Initialization strategies --
"uniform","normal","sobol","lhs", or pass your own callable. Bounded strategies (uniform,sobol,lhs) require aboundsargument;"normal"perturbs aroundx0with a configurableinit_scale. Or just pass an explicitlist[np.ndarray]asx0and skip the generation entirely. - Parallel backends --
"sequential"(for debugging),"loky"(CPU-bound work, default), or"threading"(GIL-releasing code). Under the hood this isjoblib, so the usualn_jobs=-1convention works. - BLAS thread control -- When many workers each spawn a full BLAS/OpenMP thread pool, you get noisy-neighbor
over-subscription. The
blas_coresargument (default"auto") caps the total thread budget so workers don't fight over the same cores.
The returned MultiStartResult gives you best, top_k(k), ranked(), success_rate, a summary() table, and
to_dataframe() for further analysis.
We welcome contributions! If you find a bug, have a feature request, or want to improve the documentation, please open an issue or submit a pull request on GitHub.