How to Use scipy.optimize.minimize and curve_fit in Python
Learn how to use scipy.optimize.minimize for general optimization and curve_fit for nonlinear least squares fitting, including bounds, constraints, and common pitfalls.
When you need to minimize a scalar function or fit a model to data, Python's scipy.optimize module provides two primary entry points: minimize for general optimization and curve_fit for nonlinear least squares. This article shows when to use each function, how to call them, and how to handle common optimization and fitting issues.
Using minimize for General Optimization
scipy.optimize.minimize solves problems of the form:
[ \min_x f(x) ]
where (f) is a scalar-valued function of one or more variables. It takes an initial guess x0 and returns an OptimizeResult containing the optimal values, the objective value, convergence status, and other diagnostics.
Here is a minimal example that minimizes the Rosenbrock function, a common test problem:
import numpy as np from scipy.optimize import minimize def rosen(x): return sum(100.0 * (x[1:] - x[:-1]**2)**2 + (1 - x[:-1])**2) x0 = np.array([-1.2, 1.0]) result = minimize(rosen, x0, method='BFGS') print(result.x) print(result.fun)
The method parameter selects the optimization algorithm. Common choices include 'BFGS' for unconstrained problems, 'L-BFGS-B' for bounded problems, 'SLSQP' for constrained problems, and 'trust-constr' for problems with both bounds and constraints. The default method depends on the problem structure: BFGS for unconstrained problems, L-BFGS-B when bounds are present, and SLSQP when constraints are present.
When the objective function is expensive to evaluate, you can supply a gradient via the jac parameter. For example:
def rosen_grad(x): grad = np.zeros_like(x) grad[0] = -400 * x[0] * (x[1] - x[0]**2) - 2 * (1 - x[0]) grad[1] = 200 * (x[1] - x[0]**2) return grad result = minimize(rosen, x0, jac=rosen_grad, method='BFGS')
Providing an analytic gradient often improves convergence speed and reliability.
Fitting a Model to Data with curve_fit
scipy.optimize.curve_fit is a specialized tool for nonlinear least squares fitting. Given a model function f(x, *params) and observed data (xdata, ydata), it finds the parameter values that minimize the sum of squared residuals:
[ \min_{\theta} \sum_i (y_i - f(x_i, \theta))^2 ]
The basic usage is straightforward:
from scipy.optimize import curve_fit def model(x, a, b, c): return a * np.exp(-b * x) + c xdata = np.linspace(0, 4, 50) ydata = model(xdata, 2.5, 1.3, 0.5) + 0.2 * np.random.normal(size=50) popt, pcov = curve_fit(model, xdata, ydata, p0=[1, 1, 1]) print(popt)
The function returns two values: popt, the optimal parameter values, and pcov, the estimated covariance of the parameters. The square root of the diagonal of pcov gives the standard errors of the fitted parameters.
You can constrain parameters using the bounds argument. It accepts a tuple of two lists: the lower and upper bounds for each parameter. Use -np.inf and np.inf for unbounded parameters:
popt, pcov = curve_fit(model, xdata, ydata, p0=[1, 1, 1], bounds=([0, 0, -np.inf], [np.inf, np.inf, np.inf]))
Key Differences Between minimize and curve_fit
| Aspect | minimize | curve_fit |
|---|---|---|
| Primary purpose | General scalar optimization | Nonlinear least squares fitting |
| Objective function | Any scalar function | Sum of squared residuals |
| Input data | No data required | Requires xdata and ydata |
| Output | OptimizeResult with x and fun | Parameter array and covariance matrix |
| Constraints | Supports bounds and general constraints | Supports parameter bounds only |
| Typical use | Resource allocation, design optimization | Model calibration, regression |
curve_fit is a convenience wrapper that handles residual computation and parameter unpacking; depending on the method, it calls scipy.optimize.least_squares or scipy.optimize.leastsq. If your problem is purely about fitting a model to data, curve_fit is usually more convenient. If you need to optimize an arbitrary objective function, minimize is the appropriate tool.
Handling Bounds and Constraints
For minimize, bounds are passed as a sequence of (min, max) pairs for each variable. For example:
result = minimize(rosen, x0, method='L-BFGS-B', bounds=[(-2, 2), (-2, 2)])
General constraints are defined using LinearConstraint or NonlinearConstraint objects and passed to the constraints parameter. Here is an example with a nonlinear inequality constraint:
from scipy.optimize import NonlinearConstraint def constraint(x): return x[0]**2 + x[1]**2 nlc = NonlinearConstraint(constraint, 0, 1) result = minimize(rosen, x0, method='trust-constr', constraints=[nlc])
curve_fit only supports box bounds, not general constraints. If you need to enforce a relationship between parameters, you must reparameterize the model or use minimize with a custom residual function.
Choosing the Right Method
The performance and reliability of minimize depend heavily on the chosen method. For smooth unconstrained problems, BFGS is a good default. For very large unconstrained problems or when bounds are present, L-BFGS-B uses limited memory and is often a good choice. SLSQP handles equality and inequality constraints, while trust-constr is designed for problems with both bounds and nonlinear constraints.
curve_fit uses the Levenberg-Marquardt algorithm by default for unconstrained problems, which is efficient for small and medium-sized fits. When bounds are supplied, the default becomes 'trf' (Trust Region Reflective). For difficult or bounded problems, method='trf' or method='dogbox' may be more reliable than the default.
Numerical Considerations and Performance
Scaling matters. If parameters differ by orders of magnitude, the optimizer may struggle. Consider rescaling variables or providing a custom gradient. For curve_fit, the sigma parameter can be used to weight residuals, which is important when data points have different uncertainties.
Tolerances control when the optimizer stops. minimize accepts tol, and method-specific options can be passed through options. curve_fit forwards extra keyword arguments to the underlying solver. For the default lm method, these include ftol, xtol, gtol, maxfev, and epsfcn; for trf and dogbox, use names such as ftol, xtol, gtol, and max_nfev. Excessively tight tolerances can cause extra iterations without meaningful improvement.
When the objective function is expensive, consider supplying a Jacobian to minimize or a model Jacobian to curve_fit via the jac parameter. This can reduce the number of function evaluations significantly.
Common Pitfalls and How to Avoid Them
A poor initial guess is the most common cause of failure. For curve_fit, start with parameter values that are plausible given the data. Plotting the model with the initial guess against the data can quickly reveal bad starting points.
Local minima are a risk in non-convex problems. Run the optimization from multiple starting points and compare the resulting objective values or residuals. For curve_fit, loop over different p0 arrays yourself; p0 is a single initial parameter vector, not a collection of alternative guesses.
Singular covariance matrices indicate that the model is overparameterized or that parameters are not identifiable from the data. Reduce the number of parameters or add regularization.
Finally, the covariance matrix returned by curve_fit is a meaningful uncertainty estimate only under the usual least-squares assumptions, such as independent residuals with known or constant variance. If those assumptions are violated, the standard errors may be misleading.