2. Relaxations and bounds

Draft

Section 1 placed the difficulty of a mixed-integer nonlinear program in two features of its feasible set: some coordinates must be integers, and some constraint functions curve the wrong way. Each is a nonconvexity. This section introduces the one device on which every exact method for either difficulty rests. The problem is replaced by a second problem that is easier to solve and whose optimal value is provably no worse than the original's. The second problem is a relaxation, and its value is a bound. A solver proves things with bounds and finds things with feasible points. The distance between its best bound and its best feasible point is the only quantity that summarizes an unfinished run, and the run has finished when that distance is zero.

Subsection 2.1 defines relaxations, incumbents and gaps, fixes the sign and denominator conventions that solver logs use, and introduces the vocabulary of formulation strength on the running two-variable example. Subsection 2.2 asks where a bound comes from once a relaxation has been written down, and answers with duality: weak duality makes every bound valid, and the duality gap of a nonconvex problem is the reason the search of Section 3 branches. Subsection 2.3 runs one complete solve of the running example, attaches each term of the vocabulary to an event in it, and defines the measure by which the primal side of a solver is judged. Subsections 2.4 to 2.6 then build relaxations for curved constraints and tighten them.

Minimization is the default in every display. The figures of this section drawn on the running examples R1 and R2, the relaxation, vocabulary, McCormick and FBBT figures, maximize, and the text says so where each appears. The duality figure minimizes.

The relaxation and the gap

This subsection defines the objects that every later section manipulates: the relaxation, the bound it gives, the incumbent, and the gap between them. It recalls the first running example, a program in two integer variables small enough to draw, and states on it the theorem that makes integer programming a matter of describing a polyhedron. It then settles two matters of convention that cause real confusion when solver logs are read, the sign of the inequalities and the denominator of the relative gap.

Definition 1.1.3 defined a relaxation in the original variables, as a second problem whose feasible set contains \(\mathcal F\) and whose objective lies under \(f\) on \(\mathcal F\). The general form adds extra coordinates \(w\).

Definition 2.1.1 (relaxation). Let (P) be the problem \(z^\star = \inf\{ f(x) : x \in \mathcal F \}\) with \(\mathcal F \subseteq \mathbb{R}^n\). A problem (R) with feasible set \(\mathcal F_R \subseteq \mathbb{R}^n \times \mathbb{R}^k\) and objective \(f_R(x, w)\) is a relaxation of (P) if every \(x \in \mathcal F\) extends to a point \((x, w) \in \mathcal F_R\) with \(f_R(x, w) \le f(x)\). Its value is \(z_R = \inf\{ f_R(x, w) : (x, w) \in \mathcal F_R \}\). The case \(k = 0\), with no extra coordinates, is Definition 1.1.3.

There are two ways to relax a problem, and most relaxations do both. The feasible set can be enlarged, by dropping a constraint or replacing it by a weaker one. The objective can be lowered, by replacing \(f\) with a function that lies under it. The extra coordinates \(w\) of the general form allow a third move, which Sections 2.4 and 4.7 use constantly. A new variable stands for a nonlinear expression, and linear inequalities in the new variable replace the expression's definition. The relaxation then lives in a larger space, and only its projection onto the original variables is compared with \(\mathcal F\). The bilinear example R2 of Section 1.6 is the instance to keep in mind: the new variable is a single \(w\) standing for the product \(xy\), the linear inequalities are the four McCormick planes that Section 2.4 derives, and the relaxation is a polyhedron in \((x, y, w)\) whose shadow on the \((x, y)\) plane contains the feasible set.

Proposition 2.1.2 (what a relaxation proves). Let (R) be a relaxation of (P). (a) \(z_R \le z^\star\). (b) If (R) is infeasible, so is (P). (c) If \((\bar x, \bar w)\) is optimal for (R), \(\bar x \in \mathcal F\) and \(f_R(\bar x, \bar w) = f(\bar x)\), then \(\bar x\) is optimal for (P) and \(z_R = z^\star\). (d) If (R\('\)) is a relaxation of (R) in the same variables, then \(z_{R'} \le z_R\).

Proof. (a) and (c) are Proposition 1.1.4, whose proof applies word for word to the extension \((x, w)\) of each \(x \in \mathcal F\). (b) and (d) are new. (b) If \(\mathcal F_R\) is empty, no \(x \in \mathcal F\) can have an extension, so \(\mathcal F\) is empty. (d) is (a) applied to (R) in the role of (P). ∎

Each clause is a rule a solver applies thousands of times. Clause (a) is the bound at every node of every search tree. Clause (b) discards a node whose relaxation has no solution. Clause (c) is the stopping rule: when the relaxed optimum is feasible for the original problem and its relaxed objective is its true objective, nothing remains to be proved. Clause (d) orders relaxations. A tighter relaxation gives a larger bound, and a relaxation of a relaxation is weaker than either. Four terms from convex geometry describe the relaxations of integer programs. A polyhedron is a finite intersection of closed half-spaces, and it is rational if its data are. A polytope is a bounded polyhedron. The faces of a polyhedron are its intersections with its supporting hyperplanes, and its facets are the faces of dimension one less than its own. For an integer program the natural relaxation is the one everyone meets first.

Definition 2.1.3 (LP relaxation, integer hull). Let \(\mathcal F = \{ (x, y) \in \mathbb{R}^n \times \mathbb{Z}^p : Ax + Gy \le b \}\) be the feasible set of a mixed-integer linear program with objective \(c^\top x + d^\top y\). Its LP relaxation is the linear program over the polyhedron \(P = \{ (x, y) \in \mathbb{R}^{n+p} : Ax + Gy \le b \}\), obtained by dropping the integrality of \(y\). Its value is \(z_{\mathrm{LP}} \le z^\star\). The integer hull of \(\mathcal F\) is its convex hull \(\operatorname{conv}(\mathcal F)\).

The running example

The first running example, R1 of Section 1.6, is the integer program of the plane figures. It maximizes, as the figure says, and its data, repeated from Section 1.6, are

\[\max\ \cos\theta \cdot x + \sin\theta \cdot y \quad \text{s.t.}\quad 2x + 5y \le 24.5,\ \ 5x + 2y \le 30.5,\ \ -3x + 4y \le 11,\ \ x - 2y \le 4.2,\ \ 0 \le x \le 7,\ \ 0 \le y \le 6,\ \ x, y \in \mathbb{Z},\]

with the angle \(\theta\) of the objective set by a slider. The polygon \(P\) cut out by the four rows and the nonnegativity constraints lies inside the box, whose upper bounds are nowhere active on it, and has six vertices, among them \((4.929, 2.929)\), \((5.783, 0.792)\) and \((1.870, 4.152)\). It contains 22 integer points, and the convex hull \(H\) of those points has seven facets. In the figure below the polygon is the LP relaxation, the blue dots are the integer points, the blue region is the hull, and a dashed grey polygon \(W\) is a second description of the same 22 points by looser inequalities. Three things are worth watching as the objective turns. The orange ring, where the relaxation stops, sits at a vertex of \(P\) and is fractional at most angles. The blue dot, where the integer program stops, is an integer point and a vertex of \(H\). The dashed polygon's own optimum, printed as the weak LP value, is worse than \(P\)'s at every angle.

A program in two variables, drawn: the polygon is the LP relaxation, the blue dots are the integer points inside it and the blue region is their convex hull. The objective is the arrow, which can be dragged or turned with the slider. The orange ring is where the relaxation stops, the blue dot is where the integer program stops, and the dashed grey polygon is a second formulation of the same integers with a weaker bound.

At the default angle \(\theta = 45^\circ\) the relaxation over \(P\) reaches \(5.56\) at the fractional vertex \((4.93, 2.93)\), while the best integer point, \((4, 3)\), is worth \(4.95\). The bound overstates the integer optimum by \(10.9\%\). The weaker polygon gives \(5.63\) at \((4.98, 2.98)\) and a gap of \(12.0\%\). The hull gives \(4.95\), the integer optimum. The case buttons show two more angles. At \(\theta = 14^\circ\) the bound is \(5.80\) against an integer optimum of \(5.34\) at \((5, 2)\), a gap of \(8.1\%\). At \(\theta = 76^\circ\) it is \(4.48\) against \(4.37\) at \((2, 4)\), a gap of \(2.6\%\). At every angle the LP over the hull lands on an integer point and equals the integer optimum, and the next theorem says that this is not an accident of the example. One convention needs stating at once as well. The figure's gap divides the overstatement by the bound, \((5.56 - 4.95)/5.56\). Gurobi and CPLEX divide by the incumbent, which gives \(12.2\%\) at the same angle. Both are correct statements of the same two numbers, and Definition 2.1.5 below records the conventions in use.

Proposition 1.5.6(ii) used the attainment half of the following theorem, and here is the full statement with its proof sketch.

Theorem 2.1.4 (Meyer, 1974). If \(A\), \(G\) and \(b\) are rational, \(\operatorname{conv}(\mathcal F)\) is a rational polyhedron. If \(\mathcal F \ne \emptyset\) and a linear objective is bounded below on \(\mathcal F\), its infimum over \(\mathcal F\) is attained, and \(\min\{ c^\top x + d^\top y : (x, y) \in \operatorname{conv}(\mathcal F) \} = z^\star\).R. R. Meyer, "On the existence of optimal solutions to integer and mixed-integer programming problems", Mathematical Programming 7 (1974), 223–235. The proof decomposes \(P\) into a polytope plus a rational cone and shows that the integer parts of the points of \(\mathcal F\) in a bounded region take finitely many values, which together with the cone's integral generators generate the hull.

Proof sketch. Write \(P = Q + C\) with \(Q\) a polytope and \(C\) a rational cone generated by integral vectors \(r^1, \dots, r^k\). The decomposition is needed because \(P\) may be unbounded. When \(P\) is a polytope, \(C = \{0\}\) and the rest of the argument needs no directions. Every point of \(P\) is a point of the bounded set \(Q + \{ \sum_i \mu_i r^i : 0 \le \mu_i \le 1 \}\) plus a nonnegative integer combination of the \(r^i\). The integer parts \(y\) of the points of \(\mathcal F\) in the bounded set take finitely many values, and for each value the continuous part \(x\) ranges over a polytope, so the convex hull of those points is a polytope. Together with the directions \(r^i\) it generates \(\operatorname{conv}(\mathcal F)\), which is therefore a rational polyhedron. A linear function bounded below on a polyhedron attains its minimum on a face. A face of \(\operatorname{conv}(\mathcal F)\) is the convex hull of the points of \(\mathcal F\) it contains, so the minimum is attained at a point of \(\mathcal F\) and equals \(z^\star\). ∎

The theorem says that an integer program is a linear program over its integer hull, which is what the figure shows at every angle. Among all convex relaxations written in the original variables the hull is the tightest possible, since any convex relaxation of \(\mathcal F\) contains \(\operatorname{conv}(\mathcal F)\). The whole difficulty of integer programming is that the inequality description of the hull is enormous and is not known in advance. Enormous is meant literally. The hull of a \(0\)–\(1\) program in \(n\) variables can have a number of facets that grows exponentially with \(n\), and even where the facets are known in principle, finding one that the current relaxed point violates is in general as hard as the integer program itself. A solver therefore never writes the hull down. It computes, at the point where its relaxation stopped, a few of the inequalities the hull implies, the cutting planes of Section 3.3, and leaves the rest to branching. The rationality hypothesis is needed. With irrational data the infimum can fail to be attained. In \(\inf\{ x - \sqrt 2\, y : x - \sqrt 2\, y \ge 0,\ x \ge 1,\ x, y \in \mathbb{Z} \}\) every feasible pair has a positive value, because \(\sqrt 2\) is irrational, yet the values \(0.17\) at \((3, 2)\), \(0.029\) at \((17, 12)\) and \(1.5 \times 10^{-4}\) at \((3363, 2378)\) approach zero. The infimum is \(0\), and no integer pair reaches it.

Incumbents, gaps and conventions

Definition 2.1.5 (incumbent, bounds, gap). For a minimization problem, an incumbent is a feasible point \(\bar x \in \mathcal F\), and its value \(z_{\mathrm{inc}} = f(\bar x)\) is an upper bound on \(z^\star\). Together with a relaxation value \(z_R\),

\[z_R \;\le\; z^\star \;\le\; z_{\mathrm{inc}} .\]

The absolute gap is \(z_{\mathrm{inc}} - z_R \ge 0\). A relative gap is the absolute gap divided by a scale of the problem, most often \(|z_{\mathrm{inc}}|\). The scale is a convention and differs between solvers. To prune a region of the search is to discard it because the relaxation's bound on that region is no better than the incumbent: nothing in it can improve on what is already in hand.

When the two bounds meet, the incumbent is proved optimal. The cost of a solve is the number of relaxations that had to be solved before the gap closed, and most of this series is about making that number small. Two conventions have to be fixed before any number can be read off a log. The first is the sign. Minimization is the default in every display of this series, and the figures of this section drawn on R1 and R2 (relaxation, vocabulary, McCormick and FBBT) maximize. The sign convention of Section 1.1 gives the dictionary, and the one item it leaves to this section is the chain of bounds, which for a maximization reads \(z_{\mathrm{inc}} \le z^\star \le z_R\). For the drawn figures of this section read every inequality of this section with its direction reversed.

The second convention is the denominator of the relative gap, and here the solvers disagree. Gurobi and CPLEX divide the absolute gap by the incumbent's magnitude. SCIP divides by the smaller of the two bounds' magnitudes, and it declares the gap infinite when the bounds have opposite signs or when either of them is zero. Xpress and BARON use the larger of the two. CPLEX adds \(10^{-10}\) to its denominator so that a zero incumbent does not divide by zero. GAMS, when it reports on behalf of any of them, uses its own formula and passes its own tolerances down to the solver.Gurobi Optimization, "What is the MIPGap?", Gurobi Help Center, support.gurobi.com, and Gurobi Optimizer Reference Manual, version 13.0, parameters MIPGap and MIPGapAbs, docs.gurobi.com; IBM ILOG CPLEX Optimization Studio 22.1.2, parameter pages EpGap and EpAGap; FICO Xpress Optimizer, controls MIPRELSTOP and MIPABSSTOP; SCIP, source file src/scip/set.c (parameters limits/gap and limits/absgap) and the gap formula of A. M. Gleixner, T. Berthold, B. Müller and S. Weltge, "Three enhancements for optimization-based bound tightening", Journal of Global Optimization 67 (2017), footnote 2; BARON User Manual, version 2026.9.10, options EpsA and EpsR, minlp.com, with the GAMS/BARON documentation for the GAMS defaults; GAMS User's Guide, options optCR and optCA; NVIDIA cuOpt User Guide, version 26.08, MIP settings. All read on 4 and 5 October 2026. The table collects the formulas and the default tolerances at which each solver declares a run finished. Every entry is a documented default, read from the vendor's documentation or source on the dates in the note.

solverrelative gaprel. tol.abs. tol.
Gurobi 13\(\vert z_{\mathrm{inc}} - z_R \vert / \vert z_{\mathrm{inc}} \vert\); 0 when both are 0; infinite when \(z_{\mathrm{inc}} = 0\) and \(z_R \ne 0\)1e-4 (MIPGap)1e-10 (MIPGapAbs)
CPLEX 22.1.2\(\vert z_{\mathrm{inc}} - z_R \vert / (10^{-10} + \vert z_{\mathrm{inc}} \vert)\)1e-4 (EpGap)1e-6 (EpAGap)
Xpress\(\vert z_{\mathrm{inc}} - z_R \vert \le \mathrm{tol} \cdot \max(\vert z_{\mathrm{inc}} \vert, \vert z_R \vert)\)1e-4 (MIPRELSTOP)0 (MIPABSSTOP)
SCIP 10\(\vert z_{\mathrm{inc}} - z_R \vert / \min(\vert z_{\mathrm{inc}} \vert, \vert z_R \vert)\); infinite when the signs differ or either bound is 0; the run ends when the open list is empty0 (limits/gap)0 (limits/absgap)
BARON 2026.9.10\(\vert z_{\mathrm{inc}} - z_R \vert \le \mathrm{EpsR} \cdot \max(\vert z_{\mathrm{inc}} \vert, \vert z_R \vert)\) (under GAMS the GAMS values below are passed in)1e-6 (EpsR)1e-6 (EpsA)
GAMS, reporting\(\vert \mathrm{PB} - \mathrm{DB} \vert / \max(\vert \mathrm{PB} \vert, \vert \mathrm{DB} \vert)\) (PB, DB: GAMS's names for the incumbent and the dual bound)1e-4 (optCR)0 (optCA)
cuOpt 26.08\(\vert z_{\mathrm{inc}} - z_R \vert / \vert z_{\mathrm{inc}} \vert\); infinite when \(z_{\mathrm{inc}} = 0\)1e-41e-10
How the solvers define the relative gap, and the default tolerances at which they stop.

Three remarks follow from the table. First, "a 1% gap" is not one number. The same pair of bounds is \(10.9\%\) apart when the bound is the scale and \(12.2\%\) when the incumbent is, and a benchmark that compares solvers must say which. Second, the relative gap is meaningless near zero. A problem whose optimal value is zero has an undefined or infinite relative gap. A feasibility problem with a zero objective is one example, and a tax objective in which realized gains and losses happen to net to nothing is another. This is why every solver carries an absolute tolerance as well, and why SCIP declares the gap infinite when a bound is zero or the bounds straddle zero. Third, the right scale is the problem's own unit. The tax problem of Section 9 is measured in dollars, its bounds differ by dollars, and the stopping rule that makes sense for it is an absolute gap of a few dollars. Section 8 returns to the three tolerances, feasibility, integrality and gap, that a solver works to, and to what each of them means in those units.

the absolute gap divided byused by\(P\): \(z_R = 5.556\), gap \(0.606\)\(W\): \(z_R = 5.627\), gap \(0.677\)
\(\max(\vert z_{\mathrm{inc}} \vert, \vert z_R \vert) = z_R\)the relaxation figure; Xpress, BARON, GAMS10.9%12.0%
\(\vert z_{\mathrm{inc}} \vert = 4.950\)Gurobi, CPLEX, cuOpt12.2%13.7%
\(\min(\vert z_{\mathrm{inc}} \vert, \vert z_R \vert) = 4.950\)SCIP12.2%13.7%
R1 at 45 degrees, a maximization, \(z_{\mathrm{inc}} \le z^\star \le z_R\): \(z_{\mathrm{inc}} = z^\star = 4.950\), and \(z_R = 5.556\) over \(P\) and \(5.627\) over \(W\). One pair of bounds, the relative gap three ways.

(What a global solver's answer certifies) Definition 1.5.20 stated what a global solver returns: a point \(\bar x\) that satisfies the constraints to \(\varepsilon_{\mathrm{feas}}\) and the integrality requirements to \(\varepsilon_{\mathrm{int}}\), together with a valid lower bound \(\underline z \le z^\star\) such that \(f(\bar x) - \underline z \le \max(\varepsilon_a, \varepsilon_r |f(\bar x)|)\). The underlined \(\underline z\) is the monograph's symbol for a proved lower bound on \(z^\star\), and Definition 2.3.1 uses it for the global bound of a branch-and-bound run. Call such an \(\bar x\) an \(\varepsilon\)-optimal solution. The tolerances \(\varepsilon_a\) and \(\varepsilon_r\) of that definition are the absolute and relative tolerances of the table above, each solver with its own denominator, and \(\underline z\) and \(f(\bar x)\) are the final values of \(z_R\) and \(z_{\mathrm{inc}}\).

(Feasibility is a matter of tolerance too) Conditions (i) and (ii) of Definition 1.5.20 say that "feasible" is itself a matter of tolerance, and Section 2.6 and Section 3.5 show that a spatial search would never terminate without them. Condition (iii), the validity of the bound in floating-point arithmetic at every node, is the one Section 7.3 is about, as Definition 1.5.20 noted, and the question there is what the guarantee costs when the relaxations are solved inexactly on a GPU. A local method certifies something much weaker than (iii), a point at which the first-order conditions of Section 1.3 hold approximately. The word "optimal" therefore makes different claims in a local solver's log and in a global solver's log.

Two formulations of one set of integers

The dashed polygon in the figure describes the same 22 integer points as \(P\), so as integer programs the two are identical. As relaxations they are not: \(W\) gives \(5.63\) where \(P\) gives \(5.56\) and \(H\) gives \(4.95\), at the same angle, for the same problem. The solver spends its time on the relaxation, and the relaxation is the formulation. Two formulations identical in what they permit can differ by orders of magnitude in what they cost to solve. The vocabulary for this is Vielma's, and it is used throughout Section 4.J. P. Vielma, "Mixed integer linear programming formulation techniques", SIAM Review 57 (2015), 3–57, Section 2, Definitions 2.2 and 2.3 and Propositions 2.4 and 2.5; Proposition 2.4, that the second property implies the first, is stated there and proved in Section 3. Vielma's term for the second property is "locally ideal", after Padberg; this series uses the shorter "ideal" throughout.

Definition 2.1.6 (formulation, sharp, ideal). Let \(S \subseteq \mathbb{R}^n \times \mathbb{Z}^p\) be a mixed-integer set. A formulation of \(S\) is a mixed-integer set \(Q = \{ (x, y, w) : A x + G y + D w \le b,\ y \in \mathbb{Z}^p,\ w \in \mathbb{R}^{k_1} \times \mathbb{Z}^{k_2} \}\) in the original variables and possibly auxiliary variables \(w\), whose projection onto \((x, y)\) is \(S\). Its LP relaxation \(Q_{\mathrm{LP}}\) drops every integrality requirement. The formulation is sharp if the projection of \(Q_{\mathrm{LP}}\) onto \((x, y)\) equals \(\operatorname{conv}(S)\), and ideal if every vertex of \(Q_{\mathrm{LP}}\) satisfies the integrality requirements, that is, lies in \(Q\).

Proposition 2.1.7. Let \(Q_{\mathrm{LP}}\) be bounded. (a) An ideal formulation is sharp. (b) A sharp formulation gives the LP value \(z_{\mathrm{LP}} = z^\star\) for every linear objective. (c) A formulation without auxiliary variables is sharp if and only if it is ideal, and then \(Q_{\mathrm{LP}} = \operatorname{conv}(S)\).

Proof. (a) A polytope is the convex hull of its vertices, so if every vertex lies in \(Q\) then \(Q_{\mathrm{LP}} \subseteq \operatorname{conv}(Q)\). Projection commutes with convex hulls, so the projection of \(Q_{\mathrm{LP}}\) lies in \(\operatorname{conv}(S)\). The reverse inclusion holds for every formulation, because the projection of \(Q_{\mathrm{LP}}\) is convex and contains \(S\). (b) A linear objective has the same minimum over \(S\) and over \(\operatorname{conv}(S)\). This is Theorem 2.1.4 in the bounded case, or directly: the minimum of a linear function over a polytope is attained at a vertex, and the vertices of \(\operatorname{conv}(S)\) lie in \(S\). Over a sharp formulation the LP computes the minimum over \(\operatorname{conv}(S)\). (c) Without auxiliary variables, sharp means \(Q_{\mathrm{LP}} = \operatorname{conv}(S)\), whose vertices are points of \(S\), hence integral. Ideal means every vertex of \(Q_{\mathrm{LP}}\) is in \(S\), so \(Q_{\mathrm{LP}} = \operatorname{conv}(\text{vertices}) \subseteq \operatorname{conv}(S) \subseteq Q_{\mathrm{LP}}\). ∎

With auxiliary variables the two notions separate. A formulation can project onto the hull while its own relaxation has fractional vertices, and Section 4 meets both kinds. In the figure the hull \(H\) is sharp and ideal, because its relaxation is itself and every vertex is an integer point. \(P\) and \(W\) are neither, because each has fractional vertices and each strictly contains \(H\). The figure's labels report what happens at the chosen angle, "equals the integer optimum, 4.95 at this angle", while sharpness is a property of the formulation at every angle. The weak polygon was built by pushing each facet of \(P\) outward until the next integer point would have entered. Recomputed, its right-hand sides are \(24.85\), \(30.85\), \(11.22\) and \(4.38\) in place of \(24.5\), \(30.5\), \(11\) and \(4.2\). A modeller who wrote \(W\) instead of \(P\) would have described the same feasible set and handed the solver a weaker bound at every angle. The integers are fixed by the problem. The polygon is chosen by whoever writes the model, and Section 4 is about choosing it well.

R1: one set of 22 integer points, three formulations

                                          LP value
  W   the rows of P pushed out to          5.627  neither sharp
  |   24.85, 30.85, 11.22, 4.38                   nor ideal
  |     contains
  P   2x + 5y <= 24.5, 5x + 2y <= 30.5,    5.556  neither sharp
  |   -3x + 4y <= 11, x - 2y <= 4.2               nor ideal
  |     contains
  H   the integer hull, 7 facets           4.950  sharp and ideal
  |     contains
  F   the 22 integer points                4.950  the integer
                                                  optimum

  the same integer points in every set; at 45 degrees, a
  maximization, the larger set gives the larger, weaker bound

The script below recomputes everything the figure displays, at the three angles of its case buttons. It prints the LP values over \(P\), \(W\) and \(H\), the integer optimum by enumeration of the 22 points, and the gap under the three conventions of the table.

# The running example R1: LP relaxation, weak formulation W, integer
# hull H, and the gap conventions.
#
# Maximize c.(x, y) over the integer points of P = {2x + 5y <= 24.5,
# 5x + 2y <= 30.5, -3x + 4y <= 11, x - 2y <= 4.2, x, y >= 0} inside the
# box [0, 7] x [0, 6]; c = (cos theta, sin theta).

import itertools
import math

import numpy as np

BOX = (0, 7, 0, 6)
# rows a x + b y <= r
CONS = [(2, 5, 24.5), (5, 2, 30.5), (-3, 4, 11), (1, -2, 4.2),
        (-1, 0, 0), (0, -1, 0)]

def inside(cons, p, tol=1e-9):
    return all(a * p[0] + b * p[1] <= r + tol for a, b, r in cons)

def lattice(cons):
    """The integer points of the polygon inside the box."""
    return [(x, y)
            for x in range(BOX[0], BOX[1] + 1)
            for y in range(BOX[2], BOX[3] + 1)
            if inside(cons, (x, y))]

def vertices(cons):
    """Pairwise intersections of the rows (and the box) that satisfy
    every row."""
    rows = list(cons) + [(1, 0, BOX[1]), (0, 1, BOX[3]),
                         (-1, 0, -BOX[0]), (0, -1, -BOX[2])]
    V = []
    for (a1, b1, r1), (a2, b2, r2) in itertools.combinations(rows, 2):
        d = a1 * b2 - a2 * b1
        if abs(d) < 1e-12:
            continue
        p = ((r1 * b2 - r2 * b1) / d, (a1 * r2 - a2 * r1) / d)
        if (inside(rows, p)
                and all(abs(p[0] - q[0]) + abs(p[1] - q[1]) > 1e-9
                        for q in V)):
            V.append(p)
    return V

def weaken(cons):
    """The figure's W: push each facet out while no new integer point
    enters."""
    base, out = len(lattice(cons)), [list(k) for k in cons]
    for i, k in enumerate(cons):
        # the nonnegativity rows stay
        if k[0] <= 0 and k[1] <= 0:
            continue
        s, t = 0.0, 0.05
        while t <= 4:
            trial = [(q[0], q[1], q[2] + t) if j == i else q
                     for j, q in enumerate(cons)]
            if len(lattice(trial)) != base:
                break
            s = t
            t += 0.05
        out[i][2] += max(0.0, s - 0.05)

    # if the pushes together let a point in, halve them
    for _ in range(30):
        if len(lattice([tuple(k) for k in out])) == base:
            break
        out = [[k[0], k[1], cons[i][2] + (k[2] - cons[i][2]) / 2]
               for i, k in enumerate(out)]
    return [tuple(k) for k in out]

def hull(pts):
    """Convex hull of a point set (Andrew's monotone chain),
    counter-clockwise."""
    pts = sorted(set(pts))

    def cross(o, a, b):
        return (a[0]-o[0])*(b[1]-o[1]) - (a[1]-o[1])*(b[0]-o[0])

    lo, up = [], []
    for p in pts:
        while len(lo) >= 2 and cross(lo[-2], lo[-1], p) <= 0:
            lo.pop()
        lo.append(p)
    for p in reversed(pts):
        while len(up) >= 2 and cross(up[-2], up[-1], p) <= 0:
            up.pop()
        up.append(p)
    return lo[:-1] + up[:-1]

def best(points, c):
    """A linear objective over a finite set: its maximum and where."""
    z = [c[0] * p[0] + c[1] * p[1] for p in points]
    i = int(np.argmax(z))
    return z[i], points[i]

W, L, H = weaken(CONS), lattice(CONS), hull(lattice(CONS))
print(f"P: {len(vertices(CONS))} vertices; {len(L)} integer points "
      f"inside; their hull H has {len(H)} facets")
rows = [(a, b, round(r, 2)) for a, b, r in W]
print("W rows (a, b, r):", ", ".join(map(str, rows[:3])) + ",")
print(" " * 17, ", ".join(map(str, rows[3:])))

for th in (45, 14, 76):
    c = (math.cos(math.radians(th)), math.sin(math.radians(th)))
    zP, pP = best(vertices(CONS), c)
    zW, pW = best(vertices(W), c)
    zI, pI = best(L, c)
    zH, pH = best(H, c)
    print(f"\ntheta = {th} deg")
    print(f"  LP over P  {zP:.3f} at ({pP[0]:.3f}, {pP[1]:.3f})     "
          f"LP over W  {zW:.3f} at ({pW[0]:.3f}, {pW[1]:.3f})")
    print(f"  LP over H  {zH:.3f} (at an integer vertex)   "
          f"integer    {zI:.3f} at ({pI[0]:.0f}, {pI[1]:.0f})")
    for name, zb in (("P", zP), ("W", zW)):
        print(f"  gap of {name}: absolute {zb - zI:.3f}; "
              f"(bound - inc)/bound {100*(zb - zI)/zb:.1f}% (the figure);")
        print(f"      (bound - inc)/|inc| "
              f"{100*(zb - zI)/abs(zI):.1f}% (Gurobi, CPLEX);")
        print(f"      /min(|inc|,|bound|) "
              f"{100*(zb - zI)/min(abs(zI), abs(zb)):.1f}% (SCIP)")
P: 6 vertices; 22 integer points inside; their hull H has 7 facets
W rows (a, b, r): (2, 5, 24.85), (5, 2, 30.85), (-3, 4, 11.22),
                  (1, -2, 4.38), (-1, 0, 0.0), (0, -1, 0.0)

theta = 45 deg
  LP over P  5.556 at (4.929, 2.929)     LP over W  5.627 at (4.979, 2.979)
  LP over H  4.950 (at an integer vertex)   integer    4.950 at (4, 3)
  gap of P: absolute 0.606; (bound - inc)/bound 10.9% (the figure);
      (bound - inc)/|inc| 12.2% (Gurobi, CPLEX);
      /min(|inc|,|bound|) 12.2% (SCIP)
  gap of W: absolute 0.677; (bound - inc)/bound 12.0% (the figure);
      (bound - inc)/|inc| 13.7% (Gurobi, CPLEX);
      /min(|inc|,|bound|) 13.7% (SCIP)

theta = 14 deg
  LP over P  5.803 at (5.783, 0.792)     LP over W  5.877 at (5.871, 0.748)
  LP over H  5.335 (at an integer vertex)   integer    5.335 at (5, 2)
  gap of P: absolute 0.468; (bound - inc)/bound 8.1% (the figure);
      (bound - inc)/|inc| 8.8% (Gurobi, CPLEX);
      /min(|inc|,|bound|) 8.8% (SCIP)
  gap of W: absolute 0.542; (bound - inc)/bound 9.2% (the figure);
      (bound - inc)/|inc| 10.2% (Gurobi, CPLEX);
      /min(|inc|,|bound|) 10.2% (SCIP)

theta = 76 deg
  LP over P  4.481 at (1.870, 4.152)     LP over W  4.547 at (1.882, 4.217)
  LP over H  4.365 (at an integer vertex)   integer    4.365 at (2, 4)
  gap of P: absolute 0.116; (bound - inc)/bound 2.6% (the figure);
      (bound - inc)/|inc| 2.7% (Gurobi, CPLEX);
      /min(|inc|,|bound|) 2.7% (SCIP)
  gap of W: absolute 0.182; (bound - inc)/bound 4.0% (the figure);
      (bound - inc)/|inc| 4.2% (Gurobi, CPLEX);
      /min(|inc|,|bound|) 4.2% (SCIP)

Each linear program here is solved by enumerating the vertices of a polygon, which costs a number of \(2 \times 2\) solves quadratic in the number of rows, and the integer program by enumerating 22 points. Both are the brute force that a solver cannot afford in more than a handful of dimensions, and the methods of Section 3 exist to avoid them. The three angles are independent, as are the three polygons at each angle, which is the trivial parallelism of evaluating many relaxations at once. Section 7 is about the case where the relaxations are many and each is large.

At \(\theta = 45^\circ\) the hull's maximum is attained at two integer points, \((4, 3)\) and \((5, 2)\), both worth \(4.950\). The figure breaks the tie toward \((4, 3)\), the first in lattice order.

Where this is used

Every solver in this series prints the two bounds of Definition 2.1.5 at intervals during a run and stops at its gap tolerance. The table above gives each solver's tolerance. SCIP, whose gap tolerance is zero, and Couenne, whose default is the same, stop only when the open list is empty, and the tolerances of Section 2.6 and Section 3.5 decide when a node may be discarded.

What parallelizes

For the GPU programme the gap tolerance is the budget a GPU bound may spend. A first-order method such as PDHG (primal–dual hybrid gradient, the first-order LP method of Section 7.2) returns an approximate solution of the relaxation, and an approximate bound is not a bound until it has been corrected. The correction of Section 7.3 costs tightness, and the tightness it costs is spent from the gap the user asked for. A solver that prunes on uncorrected first-order bounds can declare a wrong answer optimal. A solver that corrects them loses part of its tolerance at every node. How much is lost, and whether a cheap repair of the dual recovers it, is one of the open questions of Section 7.8.

Where bounds come from: duality

A relaxation gives a bound only once it has been solved, and "solved" means something specific for a convex problem. The solver returns not only a point but a certificate: a second object whose value is provably no greater than the objective at every feasible point. That object is a dual solution. This subsection defines the Lagrangian and its dual function and proves weak duality, which is the only reason any bound in this series is valid. It draws the picture in which the dual bound is a supporting line under the image of the problem. It states strong duality for convex problems and for linear programs, works the LP dual of a knapsack by hand, and proves Geoffrion's theorem on what the Lagrangian dual of an integer program computes. It ends with the fact that drives everything after it. For a nonconvex problem the dual bound stops short of the optimum by the height of a hole in a convex hull. No multiplier can close that gap, but branching can, and Section 3 does so.

Definition 2.2.1 (dual function, dual problem). Write the problem as

\[z^\star \;=\; \inf\ \{\, f(x) \;:\; g_i(x) \le 0,\ i = 1, \dots, m,\ x \in \mathcal X \,\},\]

where \(\mathcal X \subseteq \mathbb{R}^n\) collects the constraints that are kept inside the subproblem: a box, a polyhedron, a lattice, or a box intersected with a lattice. The Lagrangian \(L(x, \lambda) = f(x) + \lambda^\top g(x)\) with \(\lambda \in \mathbb{R}^m_{\ge 0}\) and the multipliers \(\lambda_i\) are those of Definition 1.3.1. The only change is that \(x\) is restricted to \(\mathcal X\). The dual function is \(q(\lambda) = \inf_{x \in \mathcal X} L(x, \lambda) \in [-\infty, \infty)\). The dual problem is \(d^\star = \sup_{\lambda \ge 0} q(\lambda)\), and the duality gap is \(z^\star - d^\star\).

Three remarks on the definition. First, \(q\) is concave whatever \(f\), \(g\) and \(\mathcal X\) are, because it is an infimum of functions affine in \(\lambda\). Maximizing it is a convex problem even when the original is not. Second, the choice of \(\mathcal X\) is a design decision. Everything in \(\mathcal X\) is handled exactly inside the subproblem, everything outside it is priced, and Geoffrion's theorem below says precisely what the choice buys. Third, Proposition 1.3.5 showed that the multipliers are prices, and the picture below is the same statement for the dual function. Read \(\lambda_i\) as the price per unit of violation of constraint \(i\). Then \(L(x, \lambda)\) is the cost of \(x\) when violations may be bought at those prices, and \(q(\lambda)\) is the cheapest cost attainable in a market that permits them. The next theorem says that no feasible point, which buys nothing, can cost less than the cheapest point of that market.

Theorem 2.2.2 (weak duality). For every \(\lambda \ge 0\) and every feasible \(x\), that is, every \(x \in \mathcal X\) with \(g(x) \le 0\),

\[q(\lambda) \;\le\; f(x) .\]

Hence \(d^\star \le z^\star\), and \(q(\lambda)\) is a valid lower bound on \(z^\star\) for every \(\lambda \ge 0\), optimal or not.

Proof. \(q(\lambda) = \inf_{x' \in \mathcal X} L(x', \lambda) \le L(x, \lambda) = f(x) + \lambda^\top g(x) \le f(x)\), since \(\lambda \ge 0\) and \(g(x) \le 0\) make the second term nonpositive. Taking the supremum over \(\lambda\) on the left and the infimum over feasible \(x\) on the right gives \(d^\star \le z^\star\). ∎

Paying a nonnegative price for a violation that does not occur cannot raise the cost of a feasible point, and the cheapest point of the enlarged market cannot cost more than any one feasible point. That is the whole proof. It uses neither convexity nor differentiability nor finiteness of anything: \(\mathcal X\) may be a lattice, \(f\) may be discontinuous, and \(g\) may be anything at all. This is why a Lagrangian bound of an integer program is valid without further argument. It is why an inexact dual vector from an iterative method gives a valid bound as long as \(q\) itself is evaluated correctly. And it is why the bounds that first-order GPU methods produce can be made rigorous at all (Section 7.3). Everything that is specific to convex problems concerns the other inequality, \(d^\star \ge z^\star\).

The picture

The dual function has a geometric meaning that makes every duality gap visible, and two objects carry it. The image set of the problem is \(\mathcal G = \{ (g(x), f(x)) : x \in \mathcal X \} \subset \mathbb{R}^m \times \mathbb{R}\), the set of pairs of constraint values and objective value that the points of \(\mathcal X\) produce. The perturbation function \(v(u) = \inf\{ f(x) : g(x) \le u,\ x \in \mathcal X \}\) of Section 1.3, now with \(x\) restricted to \(\mathcal X\), is the optimal value when the right-hand sides are moved from \(0\) to \(u\), so that \(z^\star = v(0)\). Fix \(\lambda\) and consider the line \(t + \lambda^\top u = q(\lambda)\) in the \((u, t)\) coordinates of the image set. It has slope \(-\lambda\), it lies on or below every point \((g(x), f(x))\) of \(\mathcal G\) by the definition of \(q\), and it touches \(\mathcal G\) at the minimizers of the Lagrangian. Its height at \(u = 0\) is \(q(\lambda)\). The primal optimum \(z^\star\) is the lowest point of \(\mathcal G\) on or left of the vertical axis \(u = 0\). Raising a supporting line until it can rise no further and reading its height at the axis is the dual problem.

A one-variable example shows the gap. Take \(f(x) = -x^2\), \(g(x) = x - 1\) and \(\mathcal X = [-1, 2]\). The feasible set is \([-1, 1]\) and \(z^\star = -1\), attained at both ends. The Lagrangian \(-x^2 + \lambda(x - 1)\) is concave in \(x\), so its minimum over the interval is at an endpoint: \(q(\lambda) = \min\{ -1 - 2\lambda,\ -4 + \lambda \}\), a concave piecewise-linear function whose two pieces cross at \(\lambda = 1\). There \(d^\star = -3\), and the duality gap is \(z^\star - d^\star = 2\). The image set is the arc of the parabola \(t = -(u + 1)^2\) from \((-2, -1)\) through the vertex \((-1, 0)\) to \((1, -4)\). The primal reads the arc itself over \(u \le 0\) and finds \(-1\). The dual reads the best supporting line, which is the chord from \((-2, -1)\) to \((1, -4)\) and crosses the axis at \((0, -3)\), so the dual value is \(-3\). The gap is the height of the hole between the arc and its chord at \(u = 0\), and no line of any slope can be pushed into that hole.

The one-variable example, minimize −x² subject to x − 1 ≤ 0 on X = [−1, 2]: its image set and best supporting line. The primal reads the arc over u ≤ 0 and finds −1. The dual reads the chord and finds −3. The gap, 2, is the height of the hole between them at u = 0. The slider moves the multiplier λ: the line of slope −λ rests on the image set from below and meets u = 0 at q(λ) = min{−1 − 2λ, −4 + λ}, which is largest, −3, at λ = 1.

Clause (iii) of the next proposition uses two notions from convex analysis. The conjugate of \(v\) is \(v^*(s) = \sup_u \{ s^\top u - v(u) \}\), and the biconjugate \(v^{**} = (v^*)^*\) is the largest closed convex function lying below \(v\), its closed convex envelope (Definition 2.4.2).

Proposition 2.2.3 (the geometry of the dual; Lemaréchal and Renaud, 2001). (i) For each \(\lambda \ge 0\), \(q(\lambda) = \inf\{ t + \lambda^\top u : (u, t) \in \mathcal G \}\). The hyperplane \(\{ (u, t) : t + \lambda^\top u = q(\lambda) \}\) supports \(\mathcal G\) from below and meets the axis \(u = 0\) at height \(q(\lambda)\). (ii) With \(\mathcal G^+ = \mathcal G + (\mathbb{R}^m_{\ge 0} \times \mathbb{R}_{\ge 0})\),

\[d^\star \;=\; \min\{ t : (0, t) \in \operatorname{cl}\operatorname{conv}(\mathcal G^+) \}, \qquad z^\star \;=\; \inf\{ t : (0, t) \in \mathcal G^+ \}.\]

(iii) In terms of the perturbation function, \(q(\lambda) = -v^*(-\lambda)\) and \(d^\star = v^{**}(0)\), so the duality gap \(v(0) - v^{**}(0)\) is the distance at \(0\) between \(v\) and its closed convex envelope.C. Lemaréchal and A. Renaud, "A geometric study of duality gaps, with applications", Mathematical Programming 90 (2001), 399–427, which develops the image-set picture and uses it to estimate gaps; the conjugate form is R. T. Rockafellar, Convex Analysis (Princeton University Press, 1970), Section 30. J. E. Falk, "Lagrange multipliers and nonconvex programs", SIAM Journal on Control 7 (1969), 534–545, proved that the dual value of a nonconvex program equals the value of the convexified primal, which is (ii) in another dress.

Proof sketch. (i) is the definition of \(q\) written in the coordinates \((u, t) = (g(x), f(x))\). For (ii), adding the orthant does not change the infimum defining \(q\), because \(\lambda \ge 0\) makes \(t + \lambda^\top u\) nondecreasing in both \(u\) and \(t\). A nonvertical hyperplane that lies below a set lies below the closure of its convex hull as well, so every \(q(\lambda)\) is at most the height \(t_0\) of the lowest point of \(\operatorname{cl}\operatorname{conv}(\mathcal G^+)\) on the axis, and \(d^\star \le t_0\). Conversely, let \(t < t_0\). The point \((0, t)\) lies outside the closed convex set \(\operatorname{cl}\operatorname{conv}(\mathcal G^+)\), so a hyperplane separates the two. Because the set recedes in every direction of the orthant, the hyperplane's normal \((\mu, \mu_0)\) has \(\mu \ge 0\) and \(\mu_0 \ge 0\), and \(\mu_0 > 0\) when the problem is feasible, since a vertical hyperplane \(\mu^\top u = \text{const}\) with \(\mu \ge 0\) cannot separate \((0, t)\) from a set that contains points with \(u \le 0\). Dividing by \(\mu_0\) gives \(\lambda = \mu/\mu_0 \ge 0\) with \(t' + \lambda^\top u > t\) on \(\mathcal G^+\), that is, \(q(\lambda) \ge t\). So \(d^\star \ge t\) for every \(t < t_0\), and \(d^\star = t_0\). The formula for \(z^\star\) is the definition. (iii) restates (i) with \(v(u) = \inf\{ t : (u', t) \in \mathcal G,\ u' \le u \}\): \(q(\lambda) = \inf_u \{ v(u) + \lambda^\top u \} = -\sup_u \{ (-\lambda)^\top u - v(u) \} = -v^*(-\lambda)\), and the supremum over \(\lambda\) is the biconjugate at \(0\). ∎

(What the dual sees, and what it cannot) Clause (iii) is the dual side of Proposition 1.3.5(ii). There, \(-\bar\lambda\) was a subgradient of \(v\) at \(0\), so the affine function \(v(0) - \bar\lambda^\top u\) lies under \(v\) and touches it at \(0\). Here the dual problem searches over all such affine minorants, and the best one touches \(v\) at \(0\) exactly when the duality gap is zero. The content of the proposition is one sentence: the dual sees only the convex hull of the image set, while the primal sees the set itself. Every duality gap is a picture of an image set with a hole under the axis, and the gap is the depth of the hole. The next figure draws this picture for three one-variable problems and lets the reader move the multiplier. In each, the left panel plots every \(x\) in the box as the point \((g(x), f(x))\), and the right half of the plane is infeasible. The blue dot is \(p^\star\), the figure's name for \(z^\star\). The orange line of slope \(-\lambda\) is raised until it rests on the image from below, and its height at \(g = 0\) is the bound \(q(\lambda)\). The right panel plots \(q\) against \(\lambda\) on \([0, 4]\), marks \(d^\star\) and \(p^\star\), and shades the gap in red. The three problems share the constraint \(g(x) = x - c \le 0\) and differ in \(f\) and in the box.

A one-variable problem, minimize f(x) subject to g(x) ≤ 0, seen through its Lagrangian dual. On the left every x in the box is drawn as the point (g(x), f(x)), the feasible points lie left of the line g = 0, the blue dot is the true minimum p*, and the orange line of slope −λ is raised until it touches the image from below, so that its height at g = 0 is the bound q(λ). On the right q(λ) is drawn against λ with the best bound d* marked, p* as a dashed blue line and the duality gap shaded red, and the slider moves λ while the buttons switch between a convex, a nonconvex and an integer problem.

The convex problem is \(f(x) = (x - 2)^2\), \(g(x) = x - 1\) on \(0 \le x \le 4\), which is Example 1.3.2 with a box added. The Lagrangian \((x-2)^2 + \lambda(x - 1)\) is minimized at \(x(\lambda) = 2 - \lambda/2\), which lies in the box for \(\lambda \le 4\), and substituting gives the closed form \(q(\lambda) = \lambda - \lambda^2/4\). At the slider's default \(\lambda = 1\) the line \(f = 0.75 - g\) rests on the image at \(x = 1.5\) and gives the bound \(0.75\). Dragging \(\lambda\) to \(2\) raises the line until it passes through the optimum itself: \(d^\star = 1.00\) at \(\lambda = 2\), \(z^\star = 1.00\) at \(x = 1\), and the gap is zero. The KKT multiplier \(\bar\lambda = 2\) of Example 1.3.2 is the maximizer of the dual function, as Proposition 1.3.5 predicts. The image set of a convex problem is a convex curve, and a supporting line can always be brought to touch it at the feasible optimum. This is strong duality, Theorem 2.2.5 below, in a picture.

The nonconvex problem is \(f(x) = 0.5(x - 2)^2 + 1.2 \sin 3x\) with the same constraint on \(-1 \le x \le 4\). Its image set is a curve with a dent. The feasible optimum is \(z^\star = 0.67\) at \(x = 1\), on the constraint boundary. The best bound is \(d^\star = -0.31\) at \(\lambda = 1.476\), and at that multiplier the supporting line rests on the image at two points, \(x \approx -0.44\) and \(x \approx 1.48\), one on each side of \(g = 0\). The line is the chord across the dent, and it passes \(0.98\) below \(z^\star\). That \(0.98\) is the duality gap, the red strip in the right panel, and the figure's right panel shows that no choice of \(\lambda\) reduces it. At the default \(\lambda = 1\) the bound is \(-0.55\).

The integer problem is \(f(x) = (x - 2.6)^2\), \(g(x) = x - 1.3\), with \(x\) an integer in \(\{0, \dots, 6\}\). Its image set is seven points. The feasible ones are \(x = 0\) and \(x = 1\), and \(z^\star = 2.56\) at \(x = 1\). The lower convex hull of the seven points crosses the axis on the chord from \(x = 1\) to \(x = 2\), at height \(1.90\), so \(d^\star = 1.90\) at \(\lambda = 2.2\) and the gap is \(0.66\). At the default \(\lambda = 1\) the line rests on \(x = 2\) and gives \(1.06\). The figure also prints the continuous relaxation, \(x\) anywhere in \([0, 6]\) with \(x \le 1.3\), whose value is \(1.69\). The Lagrangian bound \(1.90\) is the tighter of the two, and the chain \(1.69 \le 1.90 \le 2.56\) is an instance of Theorem 2.2.7 below. Only branching closes what is left, and the next subsection and Section 3 do exactly that.

Two consequences of weak duality are used so often that they deserve their own statement. The first is that shrinking \(\mathcal X\) can only raise the dual bound, which is what makes branching work. The second is that for a linear program over a box the dual function has a closed form in which any multiplier vector at all, exact or not, gives a bound. In the proposition below and in Theorem 2.2.6 the LP dual vector is written \(y\), as is usual in the linear programming literature and in the PDHG literature of Section 7. It is not the integer vector of Definition 2.1.3.

Proposition 2.2.4 (two consequences of weak duality). (a) If \(\mathcal X' \subseteq \mathcal X\), then \(q_{\mathcal X'}(\lambda) \ge q_{\mathcal X}(\lambda)\) for every \(\lambda \ge 0\), where the subscript names the set kept in the subproblem. Hence \(d^\star_{\mathcal X'} \ge d^\star_{\mathcal X}\). (b) For the linear program \(\min\{ c^\top x : Ax \ge b,\ l \le x \le u \}\) with finite \(l, u\), keep the box \(\mathcal X = [l, u]\) inside and price the rows. Then for every \(y \in \mathbb{R}^m_{\ge 0}\),

\[q(y) \;=\; b^\top y + \sum_{j=1}^n \min\{ r_j l_j,\ r_j u_j \}, \qquad r = c - A^\top y, \tag{2.2.1}\]

and \(q(y) \le z_{\mathrm{LP}}\) for every \(y \ge 0\), with equality when \(y\) is an optimal dual solution. The vector \(r = c - A^\top y\) is the one that Theorem 2.2.6 below names the reduced costs.

Proof. (a) The infimum of the same function over a smaller set is larger. (b) \(L(x, y) = c^\top x + y^\top (b - Ax) = b^\top y + r^\top x\), and \(r^\top x\) separates over the coordinates. On \([l_j, u_j]\) the linear function \(r_j x_j\) is minimized at \(l_j\) if \(r_j \ge 0\) and at \(u_j\) if \(r_j < 0\), which is the formula. Validity is Theorem 2.2.2. Tightness at an optimal dual is the strong duality of Theorem 2.2.6. ∎

Formula (2.2.1) is the first appearance of an object that Section 7 cannot do without. A first-order LP method stopped at a relative accuracy of \(10^{-4}\) returns a dual vector \(y\) that is not dual feasible, so \(b^\top y\) is not a bound. Formula (2.2.1) evaluated at the same \(y\), with negative components clipped to zero, is one. The theorem of Neumaier and Shcherbina in Section 7.3 makes the formula rigorous in floating-point arithmetic by directed rounding. The PDHG figure of Section 7.2 draws the maximization form, \(c^\top x \le b^\top y + \sum_j U_j \max(0, c_j - (A^\top y)_j)\), at every iteration.A. Neumaier and O. Shcherbina, "Safe bounds in linear and mixed-integer linear programming", Mathematical Programming 99 (2004), 283–296. The statement, proof and general form are in Section 7.3; here only the exact-arithmetic identity is used. Clause (a) says that a child node of a branch-and-bound tree inherits at least its parent's dual bound. This is why the global bound of Section 2.3 never falls as the search proceeds.

Formula (2.2.1): a bound from any dual vector y

  y from a first-order method: not dual feasible, b^T y is no bound
            |
            |  negative components clipped to zero
            v
          y >= 0
            |
            |  r = c - A^T y: one sparse matrix-vector product
            v
       [ r_1      r_2      ...      r_n ]
          |        |                 |
          v        v                 v
   min{r_j l_j, r_j u_j} for j = 1, ..., n     one term per column
          |        |                 |
          +--------+--------+--------+         a reduction: the sum
                            |
                            v
        q(y) = b^T y + sum_j min{r_j l_j, r_j u_j}  <=  z_LP

Strong duality, for convex problems and for linear programs

The theorem below needs one term from convex geometry. The relative interior of a convex set is its interior within the smallest affine subspace that contains it, so that a segment in the plane, or a face of a polytope, has a nonempty relative interior although its interior in \(\mathbb{R}^n\) is empty. Slater's condition asks for a point there at which every nonlinear constraint holds strictly, not merely for a feasible point.

Theorem 2.2.5 (strong duality under Slater's condition). Let \(f\) and \(g_1, \dots, g_m\) be convex, let \(\mathcal X\) be convex, and suppose some \(\tilde x\) in the relative interior of \(\mathcal X\) has \(g_i(\tilde x) < 0\) for every nonlinear \(g_i\) and \(g_i(\tilde x) \le 0\) for every affine \(g_i\). If \(z^\star\) is finite, then \(d^\star = z^\star\) and the dual optimum is attained.M. Slater, "Lagrange multipliers revisited", Cowles Commission Discussion Paper, Mathematics 403 (1950), reprinted in Traces and Emergence of Nonlinear Programming (Birkhäuser, 2014), 293–306. The proof sketched here follows S. Boyd and L. Vandenberghe, Convex Optimization (Cambridge University Press, 2004), Section 5.3.2.

Proof sketch. Let \(\mathcal A = \mathcal G^+\) be the image set with the nonnegative orthant added, as in Proposition 2.2.3(ii). Because \(f\) and \(g\) are convex, \(\mathcal A\) is convex, and \((0, z^\star)\) is not an interior point of \(\mathcal A\) by the definition of \(z^\star\). A supporting hyperplane at \((0, z^\star)\) gives \((\mu, \mu_0) \ne 0\) with \(\mu^\top u + \mu_0 t \ge \mu_0 z^\star\) on \(\mathcal A\). The set \(\mathcal A\) recedes in every direction of the orthant, so \(\mu \ge 0\) and \(\mu_0 \ge 0\). If some \(\mu_i\) were negative, moving far along the \(i\)-th coordinate inside \(\mathcal A\) would send \(\mu^\top u + \mu_0 t\) to \(-\infty\) and violate the inequality. The same argument in the \(t\) direction gives \(\mu_0 \ge 0\). If \(\mu_0 = 0\), then \(\mu^\top g(\tilde x) \ge 0\) with \(g(\tilde x) < 0\) forces \(\mu = 0\), a contradiction. So \(\mu_0 > 0\), and \(\lambda = \mu/\mu_0 \ge 0\) satisfies \(f(x) + \lambda^\top g(x) \ge z^\star\) for all \(x \in \mathcal X\), that is, \(q(\lambda) \ge z^\star\). With Theorem 2.2.2, \(q(\lambda) = d^\star = z^\star\). The sketch covers the case \(g(\tilde x) < 0\) for every constraint. The refinement that lets affine constraints hold with equality at \(\tilde x\) is Boyd and Vandenberghe, Section 5.2.3. ∎

The image set of a convex problem has no dent, so the point \((0, z^\star)\) on its lower boundary has a supporting hyperplane, and that hyperplane is a dual solution. Slater's condition rules out the one degenerate case, a vertical hyperplane that carries no information about \(f\). The convex panel of the figure is this proof with \(m = 1\). For linear programs the statement sharpens: no constraint qualification is needed, and both problems attain their optima whenever either is feasible and bounded.

Theorem 2.2.6 (linear programming duality). For the pair

\[\text{(LP)}\quad \min\{ c^\top x : Ax \ge b,\ x \ge 0 \}, \qquad\qquad \text{(DP)}\quad \max\{ b^\top y : A^\top y \le c,\ y \ge 0 \},\]

(i) \(b^\top y \le c^\top x\) for every pair of feasible points (weak duality); (ii) if either problem has an optimal solution, so does the other, and the optimal values are equal; (iii) at a pair of optimal solutions, \(y_i (A x - b)_i = 0\) for every row and \(x_j (c - A^\top y)_j = 0\) for every column (complementary slackness). (DP) is the Lagrangian dual of (LP) with \(\mathcal X = \mathbb{R}^n_{\ge 0}\): \(q(y) = b^\top y + \inf_{x \ge 0} (c - A^\top y)^\top x\) equals \(b^\top y\) when \(A^\top y \le c\) and \(-\infty\) otherwise.M. Conforti, G. Cornuéjols and G. Zambelli, Integer Programming, Graduate Texts in Mathematics 271 (Springer, 2014), Chapter 3, for the statement and a proof from Farkas' lemma; S. Boyd and L. Vandenberghe, Convex Optimization (2004), Section 5.2.4, for the derivation from the Lagrangian.

Statement (i) is Theorem 2.2.2. Statement (ii) is the strong duality of Theorem 2.2.5 in the case of affine constraints, with \(\mathcal X = \mathbb{R}^n\) and the sign constraints \(x \ge 0\) counted among the affine constraints, so that the refined Slater condition reduces to feasibility. The reference in the sidenote proves it directly from Farkas' lemma. Statement (iii) says that at the optimum a positive price is paid only for a row that is tight, and a column carries a positive value only if its reduced cost is zero. Two terms in that sentence are used throughout the series. The simplex method, the classical LP algorithm, to which Section 7.1 returns, walks from vertex to vertex of the polyhedron \(\{ Ax \ge b,\ x \ge 0 \}\), improving the objective at each step, and at the optimal vertex it also produces the multipliers \(y\) at no extra cost. The reduced cost of column \(j\) is \(r_j = (c - A^\top y)_j\), the change in the objective per unit increase of \(x_j\) at the prices \(y\). An LP relaxation's bound therefore comes with its certificate attached, and the bound-tightening devices of Section 2.6 work with that certificate.

A knapsack shows the whole mechanism by hand. The knapsack maximizes, so its dual minimizes. For a maximization \(\max\{ c^\top x : Ax \le b,\ x \ge 0 \}\) the dual is \(\min\{ b^\top y : A^\top y \ge c,\ y \ge 0 \}\), which is the pair (LP), (DP) with the roles of the two problems exchanged and the inequalities reversed. Three items have values \((10, 6, 4)\) and weights \((5, 4, 4)\), the capacity is \(8\), and each item is taken or not: \(\max\{ 10x_1 + 6x_2 + 4x_3 : 5x_1 + 4x_2 + 4x_3 \le 8,\ x \in \{0, 1\}^3 \}\). The LP relaxation lets \(x \in [0, 1]^3\). Its solution is Dantzig's: order the items by value per unit of weight, \(2\), \(1.5\) and \(1\), take them whole while they fit, and take the fitting fraction of the first that does not. That gives item 1 whole, three quarters of item 2, and the value \(10 + 4.5 = 14.5\).G. B. Dantzig, "Discrete-variable extremum problems", Operations Research 5 (1957), 266–288, where this bound and its greedy solution are given. The LP dual prices the capacity at \(u\) per unit and each upper bound \(x_j \le 1\) at \(s_j\):

\[\min\ 8u + s_1 + s_2 + s_3 \quad \text{s.t.}\quad 5u + s_1 \ge 10,\ \ 4u + s_2 \ge 6,\ \ 4u + s_3 \ge 4,\ \ u, s \ge 0 .\]

For a fixed price \(u\) the best \(s_j\) is \(\max(0, v_j - w_j u)\), so the dual objective is the one-variable convex piecewise-linear function \(D(u) = 8u + \max(0, 10 - 5u) + \max(0, 6 - 4u) + \max(0, 4 - 4u)\), with kinks at the ratios \(v_j / w_j\). At \(u = 1.5\), the kink of item 2, the value is \(D(1.5) = 12 + 2.5 + 0 + 0 = 14.5\), with \(s = (2.5, 0, 0)\). The values \(D(1) = 15\) and \(D(2) = 16\) on either side confirm the minimum. The two values agree, \(14.5 = 14.5\), and complementary slackness reads off the primal. Item 1 has \(s_1 = 2.5 > 0\), so \(x_1 = 1\). Item 2 has reduced cost zero, so it may be fractional. Item 3 has \(4u + s_3 = 6 > 4\), so \(x_3 = 0\). The integer optimum is \(10\), by item 1 alone or by items 2 and 3 together, so the root gap is \(4.5\) in absolute terms, \(45\%\) of the incumbent or \(31\%\) of the bound.

The three-item knapsack (values 10, 6, 4; weights 5, 4, 4; capacity 8): the LP bound 14.5 from both sides. Above, Dantzig's fill. Below, the dual function D(u), whose minimum is the same 14.5. At u = 1.5, s = (2.5, 0, 0), complementary slackness reads off x = (1, 0.75, 0). The slider moves the price u, and the sentence above the panels gives D(u) and s there.

The same function \(D(u)\) is the Lagrangian dual function of the knapsack with the box \(\{0,1\}^3\) kept inside: \(L(u) = 8u + \sum_j \max(0, v_j - w_j u)\), since for a fixed price each item is taken exactly when its value exceeds its priced weight. The two duals coincide, and the next theorem explains why: the box is integral. The script below recomputes Dantzig's bound, the dual function at three prices and its minimum, and the integer optimum.

# The LP dual and the Lagrangian dual of a three-item knapsack.
#
#   maximize  10 x1 + 6 x2 + 4 x3
#   s.t.      5 x1 + 4 x2 + 4 x3 <= 8,  x in {0,1}^3.

import itertools

import numpy as np

v = np.array([10.0, 6.0, 4.0])
w = np.array([5.0, 4.0, 4.0])
W = 8.0

# Dantzig's bound: items by value per unit weight, whole until one does
# not fit, then the fitting fraction.
order = np.argsort(-v / w)
cap, z_lp, x_lp = W, 0.0, np.zeros(3)
for j in order:
    take = min(1.0, cap / w[j])
    x_lp[j] = take
    z_lp += take * v[j]
    cap -= take * w[j]
    if take < 1.0:
        break
print(f"LP relaxation (Dantzig): {z_lp:.2f} at x = {x_lp}")

# LP dual:  min W u + s1 + s2 + s3  s.t.  w_j u + s_j >= v_j,  u, s >= 0.
# For fixed u the best s_j is max(0, v_j - w_j u), so the dual objective
# is the one-variable function D(u) = W u + sum_j max(0, v_j - w_j u),
# convex and piecewise linear with breakpoints at v_j / w_j; it is also
# the Lagrangian dual function L(u) with the box kept inside.
D = lambda u: W * u + np.maximum(0.0, v - w * u).sum()
for u in (1.0, 1.5, 2.0):
    print(f"L({u}) = D({u}) = {D(u):.2f}")
breaks = sorted(set([0.0] + list(v / w)))
u_star = min(breaks, key=D)
s_star = np.maximum(0.0, v - w * u_star)
print(f"dual optimum: u = {u_star}, s = {s_star}, value {D(u_star):.2f}")
print("  (equals the LP bound: strong duality)")

# the integer optimum, by enumeration of the eight subsets
best = max(((v @ x, x)
            for x in itertools.product((0, 1), repeat=3)
            if w @ np.array(x) <= W),
           key=lambda t: t[0])
z_int = best[0]
print(f"integer optimum: {z_int:.0f} (e.g. x = {best[1]})")
print(f"root gap: absolute {z_lp - z_int:.1f}; "
      f"relative to the incumbent {100 * (z_lp - z_int) / z_int:.0f}%;")
print(f"  relative to the bound {100 * (z_lp - z_int) / z_lp:.0f}%")
print("the box {0,1}^3 is integral, so the Lagrangian dual equals the LP "
      "bound")
print("  (Geoffrion, integrality property)")
LP relaxation (Dantzig): 14.50 at x = [1.   0.75 0.  ]
L(1.0) = D(1.0) = 15.00
L(1.5) = D(1.5) = 14.50
L(2.0) = D(2.0) = 16.00
dual optimum: u = 1.5, s = [2.5 0.  0. ], value 14.50
  (equals the LP bound: strong duality)
integer optimum: 10 (e.g. x = (0, 1, 1))
root gap: absolute 4.5; relative to the incumbent 45%;
  relative to the bound 31%
the box {0,1}^3 is integral, so the Lagrangian dual equals the LP bound
  (Geoffrion, integrality property)

The dual is minimized over its three kinks and \(u = 0\), and the primal enumerated over eight subsets, a few dozen operations in all. The \(n\) terms of \(D(u)\) are independent. This separability is what dual decomposition exploits on larger problems: pricing the few constraints that couple otherwise independent blocks makes the Lagrangian minimization separate into one problem per block, and the sum over the blocks is a reduction.

Dual decomposition on the knapsack: D(u) at u = 1.5

     the capacity row 5x_1 + 4x_2 + 4x_3 <= 8, priced at u = 1.5
                               |
       +-----------------------+-----------------------+
       |                       |                       |
       v                       v                       v
    item 1                  item 2                  item 3
  max(0, 10 - 5u)         max(0, 6 - 4u)          max(0, 4 - 4u)
    = 2.5                   = 0                     = 0
       |                       |                       |
       +-----------------------+-----------------------+
                               |  a reduction: the sum, plus 8u
                               v
                 D(1.5) = 12 + 2.5 + 0 + 0 = 14.5

  one term per item, each independent of the others

Lagrangian relaxation of an integer program

Theorem 2.2.7 (Geoffrion, 1974). Consider \(z^\star = \min\{ c^\top x : Ax \le b,\ x \in \mathcal X \}\) with \(\mathcal X = \{ x \in \mathbb{Z}^n_{\ge 0} : Dx \le d \}\) and rational data (\(\mathcal X\) may also be a mixed-integer set). Dualize the complicating constraints \(Ax \le b\):

\[q(\lambda) = \min_{x \in \mathcal X}\ c^\top x + \lambda^\top (Ax - b), \qquad z_{\mathrm{LD}} = \max_{\lambda \ge 0}\ q(\lambda) .\]

Then (i) \(z_{\mathrm{LD}} = \min\{ c^\top x : Ax \le b,\ x \in \operatorname{conv}(\mathcal X) \}\); (ii) \(z_{\mathrm{LP}} \le z_{\mathrm{LD}} \le z^\star\), where \(z_{\mathrm{LP}} = \min\{ c^\top x : Ax \le b,\ Dx \le d,\ x \ge 0 \}\) is the ordinary LP relaxation; (iii) \(z_{\mathrm{LD}} = z_{\mathrm{LP}}\) for every \(c\) and \(b\) if and only if \(\mathcal X\) has the integrality property, \(\operatorname{conv}(\mathcal X) = \{ x \ge 0 : Dx \le d \}\).A. M. Geoffrion, "Lagrangean relaxation for integer programming", Mathematical Programming Studies 2 (1974), 82–114. M. L. Fisher, "The Lagrangian relaxation method for solving integer programming problems", Management Science 27 (1981), 1–18, is the practitioner's account.

Proof. For fixed \(\lambda\) the inner problem minimizes a linear function over \(\mathcal X\), and a linear function has the same minimum over a set and over its convex hull, so \(q(\lambda) = \min_{x \in \operatorname{conv}(\mathcal X)} c^\top x + \lambda^\top (Ax - b)\). By Theorem 2.1.4, \(\operatorname{conv}(\mathcal X)\) is a rational polyhedron, so \(\max_{\lambda \ge 0} q(\lambda)\) is the Lagrangian dual of the linear program \(\min\{ c^\top x : Ax \le b,\ x \in \operatorname{conv}(\mathcal X) \}\) with the polyhedron \(\operatorname{conv}(\mathcal X)\) kept inside, and linear programming duality (Theorem 2.2.6, or Theorem 2.2.5 with affine constraints) gives equality. This is (i). For (ii), \(\operatorname{conv}(\mathcal X) \subseteq \{ x \ge 0 : Dx \le d \}\) gives \(z_{\mathrm{LD}} \ge z_{\mathrm{LP}}\) by Proposition 2.1.2(d), and \(\mathcal X \subseteq \operatorname{conv}(\mathcal X)\) gives \(z_{\mathrm{LD}} \le z^\star\). (iii) follows from (i): if the two polyhedra coincide the two linear programs coincide for every objective and right-hand side, and if they differ some linear objective separates a vertex of the larger from the smaller. ∎

The theorem says exactly what the Lagrangian dual convexifies: the part of the problem kept inside the subproblem, and no more. Dualizing constraints that leave behind a set whose LP relaxation is already its integer hull buys nothing over the LP bound. Dualizing constraints that leave behind a set whose hull is strictly inside its LP relaxation buys exactly the difference. The knapsack above is the first case: the box \(\{0, 1\}^3\) has the integrality property, and \(L(u) = D(u)\). The classical instance of the first case is the Held–Karp bound for the travelling salesman problem. Pricing the constraint that each city has exactly two incident edges leaves a spanning-tree-like subproblem whose LP relaxation is integral. By (i) the Lagrangian bound therefore equals the value of a linear program with exponentially many rows, one that could not be written down when the bound was invented.M. Held and R. M. Karp, "The traveling-salesman problem and minimum spanning trees", Operations Research 18 (1970), 1138–1162, and "Part II", Mathematical Programming 1 (1971), 6–25.

A two-variable example in which all three values differ makes (ii) concrete. Minimize \(-4x_1 - 3x_2\) subject to the complicating constraint \(3x_1 + 2x_2 \le 7.5\) and \(x \in \mathcal X = \{ x \in \mathbb{Z}^2 : 0 \le x \le 3,\ 2x_1 + 2x_2 \le 7 \}\). The set \(\mathcal X\) has ten points, and its hull is cut out by \(x_1 + x_2 \le 3\), which is strictly inside the LP relaxation's \(x_1 + x_2 \le 3.5\): \(\mathcal X\) lacks the integrality property. The integer optimum is \(-10\) at \((1, 2)\). The LP relaxation gives \(-11\) at \((0.5, 3)\). The Lagrangian dual gives \(-10.5\) at \(\lambda = 1\), and the linear program over \(\operatorname{conv}(\mathcal X)\) with the complicating row added gives the same \(-10.5\) at \((1.5, 1.5)\), as (i) requires. The dual function is concave and piecewise linear in \(\lambda\), with kinks where two points of \(\mathcal X\) tie for the Lagrangian minimum, so the script maximizes it exactly over the finitely many candidate kinks.

# Geoffrion's theorem on a two-variable integer program.
#
# The chain z_LP < z_LD < z* is strict on
#
#   minimize  -4 x1 - 3 x2
#   s.t.      3 x1 + 2 x2 <= 7.5                             (dualized)
#   and       x in X = {x in Z^2 : 0 <= x <= 3, 2 x1 + 2 x2 <= 7}.

import itertools
from fractions import Fraction as Fr

import numpy as np

c = np.array([-4.0, -3.0])
a = np.array([3.0, 2.0])
b = 7.5
X = np.array([(i, j) for i in range(4) for j in range(4)
              if 2 * i + 2 * j <= 7])

def lp_min(rows, cost):
    """min cost.x over {a x <= r}: enumerate vertices of the polygon."""
    V = []
    for (a1, r1), (a2, r2) in itertools.combinations(rows, 2):
        M = np.array([a1, a2])
        d = np.linalg.det(M)
        if abs(d) < 1e-12:
            continue
        x = np.linalg.solve(M, np.array([r1, r2]))
        if all(np.dot(q, x) <= r + 1e-9 for q, r in rows):
            V.append(x)
    V = np.array(V)
    z = V @ cost
    i = int(np.argmin(z))
    return z[i], V[i]

box = [((-1, 0), 0), ((0, -1), 0), ((1, 0), 3), ((0, 1), 3)]
feas = X[X @ a <= b]
i = int(np.argmin(feas @ c))
z_star, x_star = (feas @ c)[i], feas[i]

# drop integrality everywhere
z_lp, x_lp = lp_min(box + [((2, 2), 7), (tuple(a), b)], c)

# the dual function over X
q = lambda lam: (X @ c + lam * (X @ a - b)).min()

# q is concave piecewise linear; its kinks are where two points of X tie
cands = {Fr(0)}
for p, r in itertools.permutations(X, 2):
    da = float(r @ a - p @ a)
    if abs(da) > 1e-12:
        l = Fr(str(float(p @ c - r @ c))) / Fr(str(da))
        if l >= 0:
            cands.add(l)
lam_star = max(cands, key=lambda l: q(float(l)))
z_ld = q(float(lam_star))

# conv(X) = {0 <= x <= 3, x1 + x2 <= 3}
z_hull, x_hull = lp_min(box + [((1, 1), 3), (tuple(a), b)], c)

print(f"X has {len(X)} points; conv(X) is cut out by x1 + x2 <= 3,")
print("  strictly inside the LP relaxation x1 + x2 <= 3.5")
print(f"z*   = {z_star:.2f} at x = {x_star}")
print(f"z_LP = {z_lp:.2f} at x = {x_lp}")
print(f"z_LD = max_lam q(lam) = {z_ld:.2f} at lam = {lam_star}   "
      f"(q(0) = {q(0):.2f}, q(2) = {q(2):.2f})")
print(f"min over conv(X) with the dualized row kept = {z_hull:.2f} "
      f"at x = {x_hull}")
print("  (Geoffrion: equals z_LD)")
print(f"chain: z_LP = {z_lp:.2f} < z_LD = {z_ld:.2f} < z* = {z_star:.2f}")
X has 10 points; conv(X) is cut out by x1 + x2 <= 3,
  strictly inside the LP relaxation x1 + x2 <= 3.5
z*   = -10.00 at x = [1 2]
z_LP = -11.00 at x = [0.5 3. ]
z_LD = max_lam q(lam) = -10.50 at lam = 1   (q(0) = -12.00, q(2) = -15.00)
min over conv(X) with the dualized row kept = -10.50 at x = [1.5 1.5]
  (Geoffrion: equals z_LD)
chain: z_LP = -11.00 < z_LD = -10.50 < z* = -10.00

Each evaluation of \(q\) is a minimum over the ten points of \(\mathcal X\), and the candidates for the maximizer number at most \(\binom{10}{2}\). The candidates are independent of one another, and each evaluation of \(q\) is a reduction over the points of the easy set, which is the shape the knapsack had as well: one subproblem per block, then a sum. In a real Lagrangian relaxation the inner minimum is a shortest path, a tree or a knapsack, and the maximization over \(\lambda\) is Algorithm 2.2.8 below. The chain \(-11 < -10.5 < -10\) is the content of the theorem in one line. The LP sees the polygon \(x_1 + x_2 \le 3.5\). The Lagrangian dual sees the hull \(x_1 + x_2 \le 3\) of the easy set, but only the hull. The integer program sees the points.

Theorem 2.2.7 (Geoffrion) on a two-variable example: minimize −4x₁ − 3x₂ subject to 3x₁ + 2x₂ ≤ 7.5 over X = {x ∈ ℤ² : 0 ≤ x ≤ 3, 2x₁ + 2x₂ ≤ 7}. The LP sees the polygon, the Lagrangian dual sees the hull, and the integer program sees the points. The chain is z_LP = −11 < z_LD = −10.5 < z* = −10.

The duality gap of nonconvex problems is why we branch

(The first remedy: rewrite the problem) For a convex problem Theorem 2.2.5 closes the gap, and for a linear program Theorem 2.2.6 closes it with a certificate the simplex method produces for free. For a nonconvex problem, whether the nonconvexity is a dent in a curve or the holes of a lattice, Proposition 2.2.3 says the gap is the depth of a hole in the convex hull of the image set at \(u = 0\), and no multiplier can reach into the hole. Two things can be done about this, and the rest of the series is about both. The first is to change the image set by rewriting the problem. Lifting products into new variables, adding redundant constraints and their products, or replacing a function by a tighter convex underestimator (Definition 2.4.2) all move \(\operatorname{conv}(\mathcal G)\) closer to \(\mathcal G\) at the axis. Section 2.4 and Section 4 do this. Section 4.7 includes the one classical case, a quadratic objective with a single quadratic constraint, in which the image set is convex although neither function is and the gap is zero without any rewriting.I. Pólik and T. Terlaky, "A survey of the S-lemma", SIAM Review 49 (2007), 371–418. Section 4.7 states the lemma; the geometric reason is that the joint numerical range of two quadratic forms is convex.

(The second remedy: split the set, which is branching) The second is to split \(\mathcal X\). If \(\mathcal X = \mathcal X_1 \cup \mathcal X_2\), then \(z^\star = \min(z^\star_1, z^\star_2)\) for the two restricted problems, and the bound \(\min(d^\star_1, d^\star_2)\) is at least \(d^\star\) by Proposition 2.2.4(a). The image set of each piece is a subset of \(\mathcal G\), its convex hull has a shallower hole, and in the limit of pieces shrinking to points the hole vanishes for continuous \(f\) and \(g\). Section 3.5 states this as the consistency of bounding and proves that it makes spatial branch and bound converge. On the figure's nonconvex problem the numbers are these. On the whole box \([-1, 4]\) the dual bound is \(-0.31\) against \(z^\star = 0.67\). Split the box at \(x = 0.5\), between the two points on which the best supporting line rests. On \([-1, 0.5]\) every point is feasible, the Lagrangian with \(\lambda = 0\) is already exact, and the bound is \(1.70\), the minimum of \(f\) on that piece. The piece is pruned the moment the incumbent \(0.67\) is known. On \([0.5, 4]\) the bound is \(0.50\) at \(\lambda = 3.64\), so the global bound rises from \(-0.31\) to \(0.50\) and the gap falls from \(0.98\) to \(0.17\). One more split, at the constraint boundary \(x = 1\), gives \(0.67\) on both pieces, and the gap is zero. Integer problems branch in the same way. Fixing an integer variable to one of its values splits the lattice, each piece's hull is tighter, and Section 2.3 watches the pieces' bounds close on the running example. The script below recomputes the three problems of the figure, the one-variable example above, and the bounds on the pieces of the nonconvex box.

Splitting the box of the nonconvex problem closes the duality gap

                          [-1, 4]
              d* = -0.31 at lambda = 1.476,
              z* = 0.67: gap 0.98
                    split at x = 0.5
                     /              \
                    /                \
          [-1, 0.5]                    [0.5, 4]
   every point feasible,         d* = 0.50 at lambda = 3.64:
   lambda = 0 exact:             global bound -0.31 -> 0.50,
   bound 1.70 = min of f;        gap 0.98 -> 0.17
   pruned once the                 split at x = 1, the
   incumbent 0.67 is known         constraint boundary
                                    /              \
                                   /                \
                           [0.5, 1]                [1, 4]
                           d* = 0.67               d* = 0.67
                           at lambda = 0.00        at lambda = 4.57

  global bound 0.67 = z*: the gap is zero
# The three problems of the dual figure, and the one-variable gap example.
#
# Each is  minimize f(x)  s.t.  g(x) = x - c <= 0,  x in a box (integer
# in the third). For each the program prints p*, d*, the multiplier and
# the gap; then the bounds on pieces of the nonconvex problem's box.

import numpy as np

def dual(xs, f, g, lam_max=4.0, n=4001):
    """q(lam) = min over the box of f + lam g.

    d* = max over lam in [0, lam_max], by a grid and a golden refinement.
    """
    F, G = f(xs), g(xs)
    q = lambda lam: (F + lam * G).min()
    grid = np.linspace(0, lam_max, n)
    vals = np.array([q(l) for l in grid])
    k = int(np.argmax(vals))
    a, b = grid[max(k - 1, 0)], grid[min(k + 1, n - 1)]
    r = (np.sqrt(5) - 1) / 2
    for _ in range(80):
        c1, c2 = b - r * (b - a), a + r * (b - a)
        if q(c1) < q(c2):
            a = c1
        else:
            b = c2
    lam = 0.5 * (a + b)
    d = q(lam)
    L = F + lam * G
    # where the supporting line rests
    rest = xs[np.abs(L - d) <= 1e-6 * (F.max() - F.min()) + 1e-12]
    return lam, d, sorted(set(np.round(rest, 2)))

def report(name, xs, f, g):
    F, G = f(xs), g(xs)
    feas = G <= 1e-12
    p = F[feas].min()
    xp = xs[feas][np.argmin(F[feas])]
    lam, d, rest = dual(xs, f, g)
    q1 = (F + 1.0 * G).min()
    print(f"{name}:")
    print(f"   p* = {p:.2f} at x = {xp:.2f}; q(1) = {q1:.2f}; "
          f"d* = {d:.2f} at lam = {lam:.3f};")
    print(f"   gap p* - d* = {p - d:.2f}; the line rests on x = {rest}")
    return p, d

x_c = np.linspace(0, 4, 400001)
report("convex    f = (x-2)^2, g = x-1, x in [0,4]",
       x_c, lambda x: (x - 2) ** 2, lambda x: x - 1)
print("   closed form: x(lam) = 2 - lam/2, "
      "q(lam) = lam - lam^2/4 on [0, 4],")
print("                maximal at lam = 2")

x_n = np.linspace(-1, 4, 500001)
fn = lambda x: 0.5 * (x - 2) ** 2 + 1.2 * np.sin(3 * x)
gn = lambda x: x - 1
report("nonconvex f = 0.5(x-2)^2 + 1.2 sin 3x, x in [-1,4]", x_n, fn, gn)

x_i = np.arange(0, 7, dtype=float)
report("integer   f = (x-2.6)^2, g = x-1.3, x in {0..6}",
       x_i, lambda x: (x - 2.6) ** 2, lambda x: x - 1.3)
xr = np.linspace(0, 6, 600001)
print("   continuous relaxation of the integer problem:")
print(f"      min (x-2.6)^2 over x <= 1.3 is "
      f"{((xr[xr <= 1.3] - 2.6) ** 2).min():.2f}")

# the hand example: min -x^2 s.t. x - 1 <= 0 on X = [-1, 2]
x_h = np.linspace(-1, 2, 300001)
report("hand      f = -x^2, g = x-1, x in [-1,2]",
       x_h, lambda x: -x ** 2, lambda x: x - 1)

# why we branch: the nonconvex problem on pieces of its box: split at
# x = 0.5 (between the two resting points), then split the right piece
# at x = 1 (the constraint boundary)
for lo, hi in ((-1, 0.5), (0.5, 4), (0.5, 1), (1, 4)):
    xs = np.linspace(lo, hi, 250001)
    G = gn(xs)
    if (G <= 0).any():
        lam, d, _ = dual(xs, fn, gn, lam_max=40.0)
        print(f"   box [{lo}, {hi}]: d* = {d:.2f} at lam = {lam:.2f}")
    else:
        print(f"   box [{lo}, {hi}]: no feasible point; q(lam) -> +inf "
              f"(q(40) = {(fn(xs) + 40 * G).min():.1f}): "
              f"pruned by the dual")
convex    f = (x-2)^2, g = x-1, x in [0,4]:
   p* = 1.00 at x = 1.00; q(1) = 0.75; d* = 1.00 at lam = 2.000;
   gap p* - d* = 0.00; the line rests on x = [1.0]
   closed form: x(lam) = 2 - lam/2, q(lam) = lam - lam^2/4 on [0, 4],
                maximal at lam = 2
nonconvex f = 0.5(x-2)^2 + 1.2 sin 3x, x in [-1,4]:
   p* = 0.67 at x = 1.00; q(1) = -0.55; d* = -0.31 at lam = 1.476;
   gap p* - d* = 0.98; the line rests on x = [-0.43, 1.48]
integer   f = (x-2.6)^2, g = x-1.3, x in {0..6}:
   p* = 2.56 at x = 1.00; q(1) = 1.06; d* = 1.90 at lam = 2.200;
   gap p* - d* = 0.66; the line rests on x = [1.0, 2.0]
   continuous relaxation of the integer problem:
      min (x-2.6)^2 over x <= 1.3 is 1.69
hand      f = -x^2, g = x-1, x in [-1,2]:
   p* = -1.00 at x = -1.00; q(1) = -3.00; d* = -3.00 at lam = 1.000;
   gap p* - d* = 2.00; the line rests on x = [-1.0, 2.0]
   box [-1, 0.5]: d* = 1.70 at lam = 0.00
   box [0.5, 4]: d* = 0.50 at lam = 3.64
   box [0.5, 1]: d* = 0.67 at lam = 0.00
   box [1, 4]: d* = 0.67 at lam = 4.57

Each continuous problem's dual function is evaluated on a grid of several hundred thousand points, and the integer problem's on its seven points, so one evaluation of \(q\) is a reduction over the grid and the maximization over \(\lambda\) is a few thousand such reductions. The script's left resting point prints as \(-0.43\) where the figure, on its coarser grid, prints \(-0.44\). On the last piece, \([1, 4]\), the only feasible point is \(x = 1\), and the dual bound reaches \(0.67\) for every \(\lambda \ge 4.57\), since beyond that price the Lagrangian's minimum over the piece sits at the constraint boundary.

Maximizing the dual, and what a dual solution is good for

The dual function is concave but not smooth. A solver that uses Lagrangian bounds maximizes it by the subgradient method of Held, Wolfe and Crowder, with Polyak's step rule, in which every iterate's value is a valid bound and the best one is kept.M. Held, P. Wolfe and H. P. Crowder, "Validation of subgradient optimization", Mathematical Programming 6 (1974), 62–88. The step with a target value is Polyak's: B. T. Polyak, "Minimization of unsmooth functionals", USSR Computational Mathematics and Mathematical Physics 9 (1969), 14–29. Held, Wolfe and Crowder validated it computationally and added the halving of \(\theta\); the method itself goes back to Held and Karp (1971). At a multiplier \(\lambda\) with Lagrangian minimizer \(x(\lambda)\), the vector \(g(x(\lambda))\) is a supergradient of \(q\), because for every \(\lambda'\), \(q(\lambda') \le L(x(\lambda), \lambda') = q(\lambda) + (\lambda' - \lambda)^\top g(x(\lambda))\).

Algorithm 2.2.8 (subgradient ascent on the Lagrangian dual; Polyak 1969; Held, Wolfe and Crowder 1974).

Algorithm 2.2.8  Subgradient ascent on the Lagrangian dual
                 (Polyak 1969; Held, Wolfe and Crowder 1974)

Input   an oracle SUB(lambda) returning x(lambda) in
        argmin_{x in 𝒳} L(x, lambda) (the "easy" subproblem);
        an upper bound UB on z* (an incumbent value);
        theta_0 in (0, 2]; a halving schedule.
Output  q_best <= z*, a valid lower bound,
        and the multiplier that attains it.

1.  lambda_0 := 0;  q_best := -inf

2.  for t = 0, 1, 2, ...

3.     x_t := SUB(lambda_t)
       q_t := f(x_t) + lambda_t^T g(x_t)
       gamma_t := g(x_t)                        (a supergradient of q)

4.     q_best := max(q_best, q_t)
                        (valid by Theorem 2.2.2, whatever lambda_t is)

5.     if g(x_t) <= 0 and lambda_t^T g(x_t) = 0:
          x_t is feasible and complementary; q_t = z*; stop

6.     s_t := theta_t (UB - q_t) / ||gamma_t||^2
       lambda_{t+1} := max(0, lambda_t + s_t gamma_t)
       halve theta_t when q_best has not improved
       for a fixed number of steps

Invariant
    q_best <= z* at every step;
    the sequence lambda_t stays in R^m_{>=0}.
The loop of Algorithm 2.2.8: a bound and a step at every pass

  1. lambda_0 := 0;  q_best := -inf
     |
     v
  2. for t = 0, 1, 2, ... <-------------------------------------+
     |                                                          |
  3. x_t := SUB(lambda_t)                                       |
     q_t := f(x_t) + lambda_t^T g(x_t)                          |
     gamma_t := g(x_t), a supergradient of q                    |
     |                                                          |
  4. q_best := max(q_best, q_t)      valid by Theorem 2.2.2,    |
     |                               whatever lambda_t is       |
     |                                                          |
  5. g(x_t) <= 0 and lambda_t^T g(x_t) = 0 ? --yes--> stop:     |
     | no                                         q_t = z*      |
     |                                                          |
  6. s_t := theta_t (UB - q_t) / ||gamma_t||^2                  |
     lambda_{t+1} := max(0, lambda_t + s_t gamma_t) ------------+

  theta_t is halved when q_best has not improved for a fixed
  number of steps; q_best <= z* at every step

Each step costs one solve of the subproblem plus the inner products that form \(q_t\) and \(\gamma_t\), which is work proportional to the number of nonzeros in \(g\). The subproblem is a shortest path, a tree, a knapsack, or, when \(\mathcal X\) is a product of blocks, one small independent problem per block, and those run one per thread. The value \(q_t\) and the supergradient \(\gamma_t\) are sums over the blocks, that is, reductions. Convergence is sublinear and the multipliers oscillate. Bundle methods, which replace the step by a small quadratic program built from the supergradients seen so far, remove the oscillation at that price.

Step 5 is the stopping test that Everett's theorem explains: a Lagrangian minimizer that happens to satisfy the dualized constraints, with every positive price on a tight constraint, is optimal.

Proposition 2.2.9 (Everett, 1963). If \(\bar x \in \arg\min_{x \in \mathcal X} L(x, \bar\lambda)\) for some \(\bar\lambda \ge 0\), then \(\bar x\) is a global minimizer of \(f\) over \(\{ x \in \mathcal X : g_i(x) \le g_i(\bar x) \text{ for every } i \text{ with } \bar\lambda_i > 0 \}\). In particular, if \(g(\bar x) \le 0\) and \(\bar\lambda^\top g(\bar x) = 0\), then \(\bar x\) is optimal for the original problem and \(q(\bar\lambda) = z^\star\).H. Everett, "Generalized Lagrange multiplier method for solving problems of optimum allocation of resources", Operations Research 11 (1963), 399–417.

Proof. For such \(x\), \(f(x) = L(x, \bar\lambda) - \bar\lambda^\top g(x) \ge L(\bar x, \bar\lambda) - \bar\lambda^\top g(x) = f(\bar x) + \bar\lambda^\top (g(\bar x) - g(x)) \ge f(\bar x)\), because each term \(\bar\lambda_i (g_i(\bar x) - g_i(x))\) is nonnegative. The particular case is the first with the right-hand sides at \(0\). ∎

A Lagrangian subproblem's solution is therefore exactly optimal for the problem whose right-hand sides it happens to meet. This is how Lagrangian heuristics work: they read a subproblem solution as a feasible point of a nearby problem and repair it into a feasible point of the original. Section 3.2 catalogues such heuristics. Two further facts about the gap are used later and are stated where they are needed. For separable problems with few coupling constraints the gap is bounded by a quantity that does not grow with the number of blocks. The Shapley–Folkman lemma bounds \(z^\star - d^\star\) by the sum of the nonconvexities of at most as many blocks as there are coupling constraints that can be active at once. Section 5.4 states the theorem with its caveats and works an example on the tax lots.M. Udell and S. Boyd, "Bounding duality gap for separable problems with linear constraints", Computational Optimization and Applications 64 (2016), 355–378; the earlier estimate is J. P. Aubin and I. Ekeland, "Estimates of the duality gap in nonconvex optimization", Mathematics of Operations Research 1 (1976), 225–245. And for problems that decompose once a few constraints are priced, the Lagrangian bound can be used inside a branch-and-bound tree with cuts (Section 3.3) in place of a convex relaxation. That is the design of Karuppiah and Grossmann for nonconvex problems with decomposable structure, and Section 3 returns to it as one of the ways a node can be bounded.R. Karuppiah and I. E. Grossmann, "A Lagrangean based branch-and-cut algorithm for global optimization of nonconvex mixed-integer nonlinear programs with decomposable structures", Journal of Global Optimization 41 (2008), 163–186.

Where this is used

The solvers of this series use duality at every node without naming it. Every LP relaxation solved by the simplex method returns its dual, so every node bound in a MILP or spatial branch-and-bound tree comes with the certificate of Theorem 2.2.6. The reduced costs of that dual drive reduced-cost tightening and the Lagrangian variable bounds of Section 2.6, which are Theorem 2.2.2 applied to the rows of the LP with the incumbent as the upper bound.H. S. Ryoo and N. V. Sahinidis, "A branch-and-reduce approach to global optimization", Journal of Global Optimization 8 (1996), 107–138, for duality-based range reduction; A. M. Gleixner, T. Berthold, B. Müller and S. Weltge, "Three enhancements for optimization-based bound tightening", Journal of Global Optimization 67 (2017), 731–757, for Lagrangian variable bounds. Both are stated and proved in Section 2.6. Gradient cuts for convex constraints (Section 3.3) are dual objects too, supporting hyperplanes to a convex set.

What parallelizes

For the GPU programme the subsection leaves three facts. Weak duality needs only \(\lambda \ge 0\), so any dual iterate of a first-order LP method is a bound once formula (2.2.1) has been evaluated, which is one sparse matrix–vector product and a reduction. Section 7.3 makes it rigorous in floating point. Dual decomposition turns a problem with few coupling constraints into independent subproblems, one per block, with the multiplier update a reduction, which is the structure a GPU wants. The tax problem of Section 9 has exactly that shape: lots coupled through a few account-level constraints. The duality gap of a nonconvex problem is a quantity of the problem, not of the method. Branching is the only general way to close it. The cost of branching is the tree of Section 3.5, and every parallel design in Section 6 and Section 7 is an attempt to pay that cost faster.

The vocabulary, in one solve

Section 3 builds the search that turns bounds into proofs, one algorithm at a time. Before it begins, this subsection runs one complete solve of the running example and attaches every term of the vocabulary to an event in it. The events are the root relaxation and its bound, a heuristic that produces an incumbent, a cut, a branch, children pruned by bound and by infeasibility, and the moment the two bounds meet. It then defines the measure by which the primal side of a solver, the side that finds feasible points, is judged. The objects are defined first, in the minimal form the figure needs. Section 3.1 gives the algorithm in full with its invariant and its termination.

Definition 2.3.1 (the objects of a branch-and-bound run). The problem is a minimization. The figure below mirrors every inequality.

  1. A node \(N\) is the original problem with extra bound constraints added. The root is the node with none.
  2. The node's relaxation is a relaxation (Definition 2.1.1) of the node's problem. Its value is the node bound \(\bar z(N)\), a lower bound on the best feasible value within the node.
  3. The open list \(\mathcal L\) holds the nodes created and not yet processed. An open node carries its parent's bound until its own relaxation has been solved.
  4. The incumbent \(z_{\mathrm{inc}}\) is the best feasible value found so far, and \(+\infty\) while none is known.
  5. The global bound, written \(\underline z\) with the underline marking a proved lower bound on \(z^\star\), is \(\underline z = \min\{ z_{\mathrm{inc}},\ \min_{N \in \mathcal L} \bar z(N) \}\). When the open list is empty it equals \(z_{\mathrm{inc}}\).
  6. A node is pruned by bound when \(\bar z(N) \ge z_{\mathrm{inc}}\).
  7. A node is pruned by infeasibility when its relaxation has no solution.
  8. A node is closed by optimality when its relaxed solution is feasible for the original problem and its relaxed objective equals its true objective there. The second condition is automatic when the relaxation keeps the objective, as the LP relaxation of a MILP does. The node then updates the incumbent.
  9. To branch on a node is to replace it by two children whose feasible sets cover its own. For an integer variable with relaxed value \(\bar y_j \notin \mathbb{Z}\), the children add \(y_j \le \lfloor \bar y_j \rfloor\) and \(y_j \ge \lceil \bar y_j \rceil\).
  10. A cut is an inequality satisfied by every feasible point of the node and violated by its relaxed solution. Adding it tightens the relaxation without removing a feasible point.
  11. A heuristic is any procedure that tries to produce a feasible point without a guarantee of success.

The names are old. Branch and bound is Land and Doig's, the two-child split on a fractional variable is Dakin's, and cuts a solver can compute are Gomory's.A. H. Land and A. G. Doig, "An automatic method of solving discrete programming problems", Econometrica 28 (1960), 497–520, for branch and bound; R. J. Dakin, "A tree-search algorithm for mixed integer programming problems", The Computer Journal 8 (1965), 250–255, for the two-child dichotomy on a fractional variable; R. E. Gomory, "Outline of an algorithm for integer solutions to linear programs", Bulletin of the American Mathematical Society 64 (1958), 275–278, for cuts a solver can compute, the subject of Section 3.3.

Proposition 2.3.2 (the invariant). At every step of a run that prunes only by the three rules of Definition 2.3.1 and branches only by covering disjunctions, \(\underline z \le z^\star \le z_{\mathrm{inc}}\), and every feasible point of the original problem lies in some open node or has value at least \(z_{\mathrm{inc}}\). The global bound never decreases and the incumbent never increases. When the open list is empty, \(\underline z = z_{\mathrm{inc}} = z^\star\).

Proof. The chain of inequalities and the covering claim are Theorem 3.1.5, proved in Section 3.1 for the algorithm whose pruning rules are those of Definition 2.3.1. That proof uses Proposition 2.1.2(a), (b) and (c) for the three rules, and Proposition 2.1.2(d), or Proposition 2.2.4(a) for a Lagrangian bound, for the fact that a child inherits at least its parent's bound. Monotonicity follows from the covering claim. Solving a node replaces an inherited bound by one at least as large, and removing a node leaves the minimum over the open list unchanged or raises it. A new incumbent with value \(v\) lies in an open node whose bound is at most \(v\), so lowering \(z_{\mathrm{inc}}\) to \(v\) does not lower the minimum, and \(z_{\mathrm{inc}}\) itself only ever decreases. ∎

The figure below runs this on the running example, which maximizes, so the inequalities of the proposition are read in the mirror of the sign convention of Section 1.1. Node bounds are upper bounds. The global bound is the largest bound over the open nodes, or the incumbent when none is open, and it only ever falls. A node is pruned when its bound is at most the incumbent. The objective points at \(\theta = 20^\circ\). Nodes are processed in best-bound order, the open node with the largest bound first. The heuristic rounds the root vertex to each of its four neighbouring integer points and keeps the best feasible one. The cut is the facet of the integer hull most violated by the current vertex, the same device as in the cutting-plane figure of Section 3.3, which a solver would have to compute rather than look up. A node is branched on its most fractional coordinate. The left panel draws the plane with the current node's relaxation in orange, the incumbent in green, the cut in purple and the pruned regions hatched red. The right panel draws the two numbers every solver's log reports, the bound falling and the incumbent rising, with the gap between them as a red band. Below the optimum line it draws the green wash of the primal integral defined after the figure. Step through the run with the slider and watch each term happen.

One solve, step by step. In the plane, orange is the relaxation the current node is solving and the vertex it reaches, green is the incumbent, purple is a cut and the sliver it removed, and hatched red is what has been pruned and will never be looked at again. In the side panel the red band between the bound and the incumbent is the gap, and the green wash between the incumbent and the dotted blue optimum line is the primal integral. A hatched column marks a step with no incumbent yet, which counts 1. The "tolerance bands" switch draws an exaggerated ε = 0.01 band above the incumbent and marks the first step at which the bound falls inside it.

The default run has nine steps, and the figure's readout prints them as a table. It is reproduced here with the node's relaxation, the global bound and the incumbent after each step, the gap in the figure's convention (bound minus incumbent, divided by the incumbent), and the primal gap defined below.

stepeventnode and its relaxationglobal boundincumbentgapprimal gap
1root5.705 at \((5.78, 0.79)\)5.705none–1.000
2heuristicrounding \((5, 1)\) is feasible5.7055.04013.2%0.064
3cut\(x \le 5\) added: 5.639 at \((5, 2.75)\)5.6395.04011.9%0.064
4branchon \(y\): children \(y \le 2\) and \(y \ge 3\)5.6395.04011.9%0.064
5branch\(y \ge 3\): 5.490 at \((4.75, 3)\), branched on \(x\)5.6395.04011.9%0.064
6incumbent\(y \le 2\): 5.383 at \((5, 2)\), integral5.4905.3832.0%0.000
7infeasible\(y \ge 3\), \(x \ge 5\): empty5.4905.3832.0%0.000
8pruned\(y \ge 3\), \(x \le 4\): 4.887 at \((4, 3.3)\)5.3835.3830.0%0.000
9doneno open node5.3835.3830.0%0.000
The default run of the vocabulary figure: \(\theta = 20^\circ\), one cut at the root, rounding heuristic on, best-bound order.
The default run as a tree (theta = 20 degrees, best bound first)

                [1] root: 5.705 at (5.78, 0.79)
                [2] rounding (5, 1): incumbent 5.040
                [3] cut x <= 5: 5.639 at (5, 2.75)
                [4] branch on y = 2.75
                   y >= 3 /          \ y <= 2
                         /            \
    [5] 5.490 at (4.75, 3)            [6] 5.383 at (5, 2)
        branch on x = 4.75                integral: incumbent
         x >= 5 /     \ x <= 4            5.040 -> 5.383
               /       \
      [7] empty:       [8] 4.887 at (4, 3.3)
          5x + 2y >= 31    <= 5.383: pruned by bound
          > 30.5

  [9] no open node: 5.383 at (5, 2) is optimal
  [k] is step k of the table

Step 1 solves the root relaxation: the vertex \((5.78, 0.79)\) of the polygon is worth \(5.705\), and both coordinates are fractional. Step 2 is the heuristic: of the four roundings of the vertex only \((5, 1)\) lies in the polygon, so it becomes the incumbent, worth \(5.040\), and the gap is \(13.2\%\). Step 3 adds the cut \(x \le 5\), the hull facet the vertex violates most. The relaxation moves to \((5, 2.75)\) and the bound falls to \(5.639\). The cut removed a sliver of the polygon that contained no integer point, which is what a valid cut does. Step 4 branches on \(y\), the only fractional coordinate: the children are \(y \le 2\) and \(y \ge 3\), and each inherits the bound \(5.639\) until it is solved. Step 5 processes the child \(y \ge 3\): its relaxation is \(5.490\) at \((4.75, 3)\), fractional in \(x\), so it is branched in turn, into \(x \le 4\) and \(x \ge 5\). Step 6 processes the child \(y \le 2\), which now holds the best bound. Its relaxed vertex is the integer point \((5, 2)\), worth \(5.383\), so the node is closed by optimality and the incumbent improves from \(5.040\) to \(5.383\). The global bound, now the largest bound among the open grandchildren, falls to \(5.490\), and the gap is \(2.0\%\). Step 7 finds the grandchild \(y \ge 3,\ x \ge 5\) infeasible, since \(5x + 2y \ge 31 > 30.5\). Step 8 solves the grandchild \(y \ge 3,\ x \le 4\). Its bound \(4.887\), at \((4, 3.3)\), is below the incumbent, so it is pruned by bound without being branched. No open node remains, so the global bound equals the incumbent, and step 9 records that \(5.383\) at \((5, 2)\) is optimal. The record of bounds is the proof: every region not searched was shown to hold nothing better. Two nodes were pruned, one by infeasibility and one by bound.

The open list and the global bound after steps 4 to 8 (default run)

  after step     4          5          6          7          8
             +--------+ +--------+ +--------+ +--------+
  next -->   | y >= 3 | | y <= 2 | | x >= 5 | | x <= 4 |  (empty)
             | 5.639  | | 5.639  | | 5.490  | | 5.490  |
             +--------+ +--------+ +--------+ +--------+
             | y <= 2 | | x >= 5 | | x <= 4 |
             | 5.639  | | 5.490  | | 5.490  |
             +--------+ +--------+ +--------+
                        | x <= 4 |
                        | 5.490  |
                        +--------+

  global bound  5.639      5.639      5.490      5.490      5.383

  x >= 5 and x <= 4 are the children of y >= 3; an open node
  carries its parent's bound; the global bound is the largest
  bound over the open nodes, or the incumbent when none is open

The case buttons show what each device bought. With neither cuts nor heuristic, "branch and bound alone", the run also takes nine steps but prunes three nodes and has no incumbent until step 6, so five steps pass with nothing in hand. With cuts added at the root until the vertex is integral, "cuts until the root is integral", the run takes six steps and prunes nothing. The cutting-plane loop solves the integer program at the root, as the cuts figure of Section 3.3 shows in detail, and the tree is never needed. The answer is \(5.383\) in every case. The three runs differ in how early a feasible point appeared and in how many relaxations had to be solved, and the next definitions measure the first of those.

The primal gap and the primal integral

A solver's log reports the dual bound and the incumbent against time, and the gap between them says how far the proof has to go. It does not say how good the incumbent was before the proof finished. Many nonconvex MINLPs of practical size end at a time limit. For a user who will stop the solver there, the quality of the incumbent at every moment of the run matters more than the moment the gap closes, and Berthold's primal integral is the measure of it.T. Berthold, "Measuring the impact of primal heuristics", Operations Research Letters 41 (2013), 611–614, which defines the primal gap function and the primal integral and uses them to evaluate heuristics inside SCIP.

Definition 2.3.3 (primal gap; Berthold, 2013). Let \(z^\star\) be the optimal value, or the best value known when the optimum is not. The primal gap of a feasible value \(\tilde z\) is

\[\gamma(\tilde z) \;=\; \begin{cases} 0 & \text{if } \tilde z = z^\star, \\ 1 & \text{if no feasible point is known, or } \tilde z\, z^\star < 0, \\ \dfrac{|z^\star - \tilde z|}{\max\{ |z^\star|, |\tilde z| \}} & \text{otherwise}, \end{cases}\]

a number in \([0, 1]\). The primal gap function of a run is \(p(t) = \gamma(z_{\mathrm{inc}}(t))\), the primal gap of the incumbent held at time \(t\). It is a step function, equal to \(1\) before the first incumbent, that jumps down at each improvement.

Definition 2.3.4 (primal integral and its time average). For a run of length \(T\) with incumbent improvements at times \(t_1 < \dots < t_k\), the primal integral is

\[P(T) \;=\; \int_0^T p(t)\, dt \;=\; \sum_{i=0}^{k} p(t_i)\,(t_{i+1} - t_i), \qquad t_0 = 0,\ t_{k+1} = T,\]

in units of time, and the time-averaged primal integral is \(\bar P(T) = P(T)/T \in [0, 1]\). Before the first incumbent the integrand is \(p(t) = 1\). This is Berthold's convention, and the series uses it throughout. Some benchmark pages charge \(2\) there, and the text says so wherever it quotes one. The same definitions apply with steps, nodes or relaxations solved in place of time, in which case the integral is in step units.

Because \(p\) takes values in \([0, 1]\), \(0 \le P(T) \le T\) and \(0 \le \bar P(T) \le 1\), and \(P(T) = T\) when no feasible point was found. Two runs that reach the same final gap at the same time but find their first good solution at different times have different primal integrals, and the earlier one is smaller. Both normalizations are in use, and a quoted number must say which. Berthold's \(P(T)\) carries the unit of time and is bounded by \(T\). The time average \(\bar P(T)\) is dimensionless, and it is the form reported by the competitions and benchmark pages that score heuristics, which Section 7.7 lists with their dates, each with its own convention for the steps before the first incumbent.MIP Workshop 2026, "Computational Competition: GPU-Accelerated Primal Heuristics for MIP" (the Land–Doig competition; topic announced 13 October 2025, submissions due 20 March 2026, results presented May 2026), rules and results, mixedinteger.org; GAMS, "Expanding the focus: introducing the MIPFEAS benchmark", blog, March 2026, with a September 2026 update, gams.com. Both report the time average \(\bar P(T)\) with the integrand \(2\) while no incumbent exists, \(1\) when the incumbent and the optimum differ in sign, and otherwise \(|z(t) - z^\star| / \max(|z(t)|, |z^\star|, 1)\), the formula printed on Mittelmann's MIPFEAS page, plato.asu.edu/ftp/mipfeas.html. Berthold's \(P(T)\) uses \(1\) in both cases. All three pages read on 5 October 2026. Every primal-integral number in this series names its symbol.

A worked example makes the second observation concrete. The optimum is \(100\) and the budget \(10\) seconds. Run A finds \(90\) at \(t = 1\) and \(100\) at \(t = 5\): \(P = 1 \cdot 1 + 4 \cdot 0.10 + 5 \cdot 0 = 1.4\) gap-seconds, \(\bar P = 0.14\). Run B finds nothing until \(t = 3\) and then \(100\): \(P = 3.0\), \(\bar P = 0.30\). Run B proved optimality earlier if it also closed its bound at \(t = 3\). Run A was the better primal run, and during seconds 1 to 3 it could prune on a bar that B did not have. In the figure the integral is taken over steps. The default run has one step with no incumbent, which counts \(1\). It has four steps with the incumbent \(5.040\) against the final \(5.383\), each counting \((5.383 - 5.040)/5.383 = 0.064\), and four steps at \(0\). The sum is \(1.25\) step units, the number the figure prints and the area of its green wash. With the heuristic switched off the five empty steps give \(5.00\). With cuts until the root is integral the run is six steps and the integral \(1.19\). A smaller value means that good incumbents arrived earlier, and the figure's right panel shows it as the area between the incumbent curve and the optimum line. The script below computes \(P(T)\) and \(\bar P(T)\) for the two runs and recomputes the gap and primal-gap columns and the integral of the nine-step table.

The primal gap p(t) of runs A and B (optimum 100, T = 10 seconds), in Berthold's convention of Definition 2.3.3: p = 1 while no incumbent is held. The shaded area is the primal integral P(T): 1.4 gap-seconds for run A and 3.0 for run B.
# The primal gap and the primal integral (Berthold 2013).
#
# P(T) and P(T)/T for the two timed runs of the worked example, then
# the nine steps of the vocabulary figure's default run.

import numpy as np

def primal_gap(z_inc, z_star):
    """The primal gap of Definition 2.3.3, in Berthold's convention.

    1 with no incumbent or with opposite signs, else the ratio
    |z_star - z_inc| / max(|z_star|, |z_inc|).
    """
    if z_inc is None or z_inc * z_star < 0:
        return 1.0
    if z_inc == z_star:
        return 0.0
    return abs(z_star - z_inc) / max(abs(z_star), abs(z_inc))

def primal_integral(events, T, z_star):
    """P(T): the primal gap, a step function on [0, T], integrated.

    events: (time, incumbent value) pairs, in time order.
    """
    times = [t for t, _ in events] + [T]
    inc = None
    P = 0.0
    t_prev = 0.0
    for (t, z), t_next in zip([(0.0, None)] + events, times):
        if z is not None:
            inc = z
        P += primal_gap(inc, z_star) * (t_next - t_prev)
        t_prev = t_next
    return P

T, z_star = 10.0, 100.0
for name, ev in (("run A: 90 at t = 1, 100 at t = 5",
                  [(1.0, 90.0), (5.0, 100.0)]),
                 ("run B: nothing until t = 3, then 100",
                  [(3.0, 100.0)])):
    P = primal_integral(ev, T, z_star)
    print(f"{name}:\n"
          f"  P(T) = {P:.2f} gap-seconds, "
          f"time average P(T)/T = {P / T:.2f}")

# the vocabulary figure, default run (theta = 20 deg, one cut at the
# root, rounding heuristic on); the bound and the incumbent after each
# step are the figure's; the gap and primal-gap columns and the
# integral are recomputed here
steps = [("root", 5.705, None),
         ("heuristic", 5.705, 5.040),
         ("cut", 5.639, 5.040),
         ("branch", 5.639, 5.040),
         ("branch", 5.639, 5.040),
         ("incumbent", 5.490, 5.383),
         ("infeasible", 5.490, 5.383),
         ("pruned", 5.383, 5.383),
         ("done", 5.383, 5.383)]
zf = steps[-1][2]
total = 0.0
print("\nstep  event       bound   incumbent   "
      "gap (bound-inc)/|inc|   primal gap")
for k, (ev, bd, inc) in enumerate(steps, 1):
    gap = ("   -  " if inc is None
           else f"{100 * (bd - inc) / abs(inc):5.1f}%")
    pg = primal_gap(inc, zf)
    total += pg
    print(f"{k:3d}   {ev:<10}  {bd:.3f}   "
          f"{'  -  ' if inc is None else f'{inc:.3f}'}      "
          f"{gap}              {pg:.3f}")

print(f"primal integral over the 9 steps, in step units: {total:.3f}\n"
      f"  (1 for the step before any incumbent, "
      f"{primal_gap(5.040, zf):.3f} for each of 4 steps, 0 after)")
print("relative gap at step 2 under the figure's other convention,\n"
      f"  (bound-inc)/bound: {100 * (5.705 - 5.040) / 5.705:.1f}%")
run A: 90 at t = 1, 100 at t = 5:
  P(T) = 1.40 gap-seconds, time average P(T)/T = 0.14
run B: nothing until t = 3, then 100:
  P(T) = 3.00 gap-seconds, time average P(T)/T = 0.30

step  event       bound   incumbent   gap (bound-inc)/|inc|   primal gap
  1   root        5.705     -           -                1.000
  2   heuristic   5.705   5.040       13.2%              0.064
  3   cut         5.639   5.040       11.9%              0.064
  4   branch      5.639   5.040       11.9%              0.064
  5   branch      5.639   5.040       11.9%              0.064
  6   incumbent   5.490   5.383        2.0%              0.000
  7   infeasible  5.490   5.383        2.0%              0.000
  8   pruned      5.383   5.383        0.0%              0.000
  9   done        5.383   5.383        0.0%              0.000
primal integral over the 9 steps, in step units: 1.255
  (1 for the step before any incumbent, 0.064 for each of 4 steps, 0 after)
relative gap at step 2 under the figure's other convention,
  (bound-inc)/bound: 11.7%

The script's \(1.255\) differs from the figure's \(1.254\) in the third decimal because it starts from the three-decimal values the figure prints rather than the unrounded ones. Both display as \(1.25\). The computation is a single pass over the events of a run, a few operations per incumbent, and a solver keeps it as a running sum at no measurable cost. The sum runs over the events in the order of the run, a short list, and nothing in it is worth parallelizing. The last line records that the relaxation figure and the vocabulary figure use different denominators for the relative gap, the bound in the first and the incumbent here.

Stopping on a tolerance

A solver does not wait for the two bounds to coincide. It stops when they are within the gap tolerances \(\varepsilon_a\) and \(\varepsilon_r\) of Definition 1.5.20, and the figure's "tolerance bands" switch shows what that means. It draws a band of relative width \(\varepsilon = 0.01\) above the incumbent curve and marks the first step at which the bound falls inside it. In the default run that step is step 8, where the gap has already closed outright from \(2.0\%\) at step 7, so a tolerance of \(0.01\) would have changed nothing. Turn the objective to \(\theta = 17^\circ\) and the band is entered at step 6 with a gap of \(1.0\%\) still open. A solver with that tolerance would have stopped there, two steps before the gap closed at step 8, and returned its incumbent as \(\varepsilon\)-optimal with a certificate that nothing better by more than \(1\%\) exists. The band is drawn at \(0.01\) to be visible. The solvers of the table in Section 2.1 stop at \(10^{-4}\), BARON on its own at \(10^{-6}\) and SCIP at zero. Those two steps are what a tolerance buys. On a nonconvex problem whose bound converges slowly, Section 3.5 shows the last few digits of the gap costing more nodes than all the digits before them. The tolerance is where the user decides how many of those digits are worth paying for.

Stopping on a tolerance: the gap (bound − incumbent)/incumbent after each step of the default run (θ = 20°), plus the one step the post gives for θ = 17°, against a band of relative width ε. The text's ε = 0.01 is drawn large so that it is visible. The slider moves ε, and the sentence above the panel gives the first step at which the default run is inside the band.

The dual side has a matching measure. The area of the red band in the right panel, the gap integrated over the run, is the primal–dual integral. The developers of SCIP use its shifted geometric mean (the benchmark statistic of Section 5.2) over a test set to measure what a component of the solver is worth. Switching off domain propagation (bound propagation through the constraints, Section 2.6), for example, raised the shifted geometric mean of the primal–dual integral by 150% and the running time by 85% on the 456 affected instances of their test set.S. Vigerske and A. Gleixner, "SCIP: global optimization of mixed-integer nonlinear programs in a branch-and-cut framework", Optimization Methods and Software 33 (2018), 563–593, Table 2, row "domain propagation": 456 of 465 instances affected, time +85%, primal–dual integral +150%, nodes +86%. The paper credits the primal–dual integral to Berthold (2013), cited above. Section 2.6 quotes its OBBT row. For comparing a heuristic solver, which never closes the gap, against a global one, Berthold and Csizmadia defined a confined variant of the primal integral. This series does not use it.T. Berthold and Z. Csizmadia, "The confined primal integral: a measure to benchmark heuristic MINLP solvers against global MINLP solvers", Mathematical Programming 188 (2021), 523–537.

Where this is used

The vocabulary of this subsection is the vocabulary of every solver log in the series. Gurobi, CPLEX, Xpress, SCIP and BARON print, at intervals, the number of nodes processed and open, the global bound, the incumbent and the gap. A reader who can follow the nine-step table can follow any of them. One lesson carries forward to Section 3. The incumbent is half of pruning: step 8 discarded a node because \(5.383\) was in hand, and with no incumbent the same node would have been branched. Section 3.2 is about finding incumbents early, and the primal integral is its scoreboard.

What parallelizes

The heuristics that NVIDIA's documentation says run on its GPU, Feasibility Jump, feasibility pumps and local search, are primal devices, so the GPU solvers of Section 7.7 are judged by the primal integral before anything else.NVIDIA, cuOpt User Guide, version 26.08, "Introduction", docs.nvidia.com, read 5 October 2026: "Primal heuristics such as local search, feasibility pump, and feasibility jump run on the GPU." This is the vendor's description of its own software; Section 7.7 gives the independently measured numbers. In a search that bounds nodes in batches, as Section 6.4 describes, an incumbent found one batch earlier discards whole batches of nodes rather than one. The timing of incumbents therefore matters more on a GPU than on a CPU, not less.

Convex envelopes and factorable relaxations

The relaxation of an integer program is free: drop the integrality constraints and an LP remains. The relaxation of a nonconvex function has to be built. Spatial branch and bound (Section 3.5) needs, at every node \(N\) with box \(B_N\), a convex problem whose optimal value \(\bar z(N)\) is a valid lower bound on the node's true value. The objective \(f\) and the constraint functions \(g_i\) are nonconvex. Each must therefore be replaced by a convex function that lies under it on \(B_N\), and each constraint \(g_i \le 0\) by the relaxed constraint \(g_i^{\mathrm{cv}} \le 0\), which enlarges the feasible set. This subsection is about how such functions are built and how good they can be. The best possible replacement of a single function on a box is its convex envelope, the largest convex function below it. The envelope is the yardstick for everything else. What solvers compute is not the envelope but a factorable relaxation. The function is written as a graph of sums, products and library univariate functions, each node is relaxed on its own interval by a rule that looks only at that node, and the rules compose. This is mechanical, cheap and valid for any closed-form expression. It is in general not the envelope, it depends on how the expression is written, and it inherits the dependency problem of interval arithmetic. The subsection ends with the solvers that use these constructions and with what in them a GPU can do.

Minimization is the default throughout. The bilinear running example (R2) maximizes, and the text says so where it appears.

The convex envelope

Definition 1.2.1 defined the convex hull \(\operatorname{conv}(S)\) and the epigraph \(\operatorname{epi} f\). Two further facts about the hull are needed.

Proposition 2.4.1 (Carathéodory, and compactness of the hull). Let \(S \subseteq \mathbb{R}^n\). (a) Every point of \(\operatorname{conv}(S)\) is a convex combination of at most \(n + 1\) points of \(S\). (b) If \(S\) is compact then \(\operatorname{conv}(S)\) is compact.Proofs in R. T. Rockafellar, Convex Analysis (Princeton University Press, 1970), Section 17.

Definition 2.4.2. Let \(B \subset \mathbb{R}^n\) be compact and convex (a box \(B = [l, u]\) in everything that follows) and let \(f : B \to \mathbb{R}\) be continuous. Write \(\operatorname{epi}_B f = \{(x, t) : x \in B,\ t \ge f(x)\}\) for the epigraph of Definition 1.2.1 restricted to \(B\). The convex envelope of \(f\) on \(B\) is

\[\operatorname{vex}_B f\,(x) \;=\; \sup\{\, g(x) : g : B \to \mathbb{R} \text{ convex},\; g \le f \text{ on } B \,\}, \qquad x \in B,\]

and the concave envelope is \(\operatorname{cav}_B f = -\operatorname{vex}_B(-f)\). A convex underestimator of \(f\) on \(B\) is any convex function \(f^{\mathrm{cv}} \le f\) on \(B\). A concave overestimator is any concave \(f^{\mathrm{cc}} \ge f\). The supremum of convex functions is convex and each is below \(f\), so \(\operatorname{vex}_B f\) is itself a convex underestimator, and it is the largest one. For every underestimator and overestimator,

\[f^{\mathrm{cv}} \;\le\; \operatorname{vex}_B f \;\le\; f \;\le\; \operatorname{cav}_B f \;\le\; f^{\mathrm{cc}} \qquad \text{on } B .\]

Given the problem \(z^\star = \min\{f(x) : g_i(x) \le 0,\ i = 1, \dots, m,\ x \in B\}\), consider \(z_R = \min\{f^{\mathrm{cv}}(x) : g_i^{\mathrm{cv}}(x) \le 0,\ x \in B\}\). This problem is convex, its feasible set contains \(\mathcal F\), and \(f^{\mathrm{cv}} \le f\) on \(\mathcal F\). Hence \(z_R \le z^\star\): it is a relaxation in the sense of Section 2.1. An equality constraint \(h(x) = 0\) is relaxed to the pair \(h^{\mathrm{cv}}(x) \le 0 \le h^{\mathrm{cc}}(x)\). This is why both envelopes are needed and why a solver carries both sides of every node.

The envelope has three equivalent descriptions, and each is used below. It is the largest convex minorant (the definition), the lower boundary of the convex hull of the epigraph, and a minimum over weighted averages of function values.

Proposition 2.4.3 (three descriptions of the envelope). Let \(B\) be compact convex and \(f : B \to \mathbb{R}\) continuous. Then for every \(x \in B\)

\[\operatorname{vex}_B f\,(x) \;=\; \min\Big\{ \sum_{i=1}^{n+1} \theta_i f(x_i) \;:\; x_i \in B,\ \theta_i \ge 0,\ \sum_i \theta_i = 1,\ \sum_i \theta_i x_i = x \Big\}, \tag{2.4.1}\]

the minimum is attained, and \(\operatorname{epi}_B(\operatorname{vex}_B f) = \operatorname{conv}(\operatorname{epi}_B f)\).

Proof sketch. Call the right-hand side of (2.4.1) \(\phi(x)\). A point \((x, t)\) lies in \(\operatorname{conv}(\operatorname{epi}_B f)\) if and only if it is a convex combination of points \((x_i, t_i)\) with \(t_i \ge f(x_i)\). By Carathéodory's theorem in \(\mathbb{R}^{n+1}\) at most \(n + 2\) points are needed, and one more can be removed. For fixed points \(x_1, \dots, x_{n+2}\), minimizing \(\sum_i \theta_i f(x_i)\) over the weights subject to \(\sum_i \theta_i x_i = x\), \(\sum_i \theta_i = 1\) and \(\theta \ge 0\) is a linear program with \(n + 1\) equality constraints. A basic optimal solution of it has at most \(n + 1\) positive weights. Hence \((x, t) \in \operatorname{conv}(\operatorname{epi}_B f)\) if and only if \(t \ge \phi(x)\): the epigraph of \(\phi\) is \(\operatorname{conv}(\operatorname{epi}_B f)\), which is convex, and \(\phi\) is a convex function. Taking a single point in (2.4.1) gives \(\phi \le f\). If \(g\) is convex with \(g \le f\), then \(g(x) \le \sum_i \theta_i g(x_i) \le \sum_i \theta_i f(x_i)\) for every admissible representation, so \(g \le \phi\). Hence \(\phi\) is the largest convex minorant of \(f\), which is \(\operatorname{vex}_B f\). The minimum in (2.4.1) is attained because the representations range over a compact set and \(f\) is continuous. ∎

In a picture, the envelope is the floor of the convex hull of the region above the graph. Each value \(\operatorname{vex}_B f(x)\) is realized by a chord through at most \(n + 1\) points of the graph whose weighted average sits over \(x\). On an interval this is the familiar lower convex hull of a curve: pieces of the curve where it is convex, joined by tangent or secant segments where it is not.

The first theorem about envelopes says that relaxing a single function costs nothing at the minimum.

Theorem 2.4.4 (Falk, 1969). Let \(f\) be continuous on the compact convex set \(B\). Then

\[\min_{x \in B} \operatorname{vex}_B f\,(x) \;=\; \min_{x \in B} f(x), \qquad \operatorname{argmin}_B \operatorname{vex}_B f \;=\; \operatorname{conv}\big(\operatorname{argmin}_B f\big).\]

Proof. Let \(m = \min_B f\). The constant function \(m\) is a convex minorant of \(f\), so \(\operatorname{vex}_B f \ge m\) everywhere, and \(\operatorname{vex}_B f \le f\) gives \(\min_B \operatorname{vex}_B f \le m\). For the second claim, if \(\operatorname{vex}_B f(x^\star) = m\), then by (2.4.1) \(x^\star = \sum_i \theta_i x_i\) with \(\sum_i \theta_i f(x_i) = m\). Since every \(f(x_i) \ge m\), each \(x_i\) with \(\theta_i > 0\) is a minimizer of \(f\), so \(x^\star \in \operatorname{conv}(\operatorname{argmin}_B f)\). Conversely, \(\operatorname{vex}_B f = m\) on \(\operatorname{argmin}_B f\) and, by convexity, \(\operatorname{vex}_B f \le m\) on the convex hull of that set, hence \(= m\) there. ∎

Falk proved this on the way to showing that the Lagrangian dual of a nonconvex program equals the primal value of the convexified program. Section 2.2 stated that result as Proposition 2.2.3 and, for integer programs, as Theorem 2.2.7 (Geoffrion). Horst and Tuy state the theorem as a basic property of the envelope.J. E. Falk, "Lagrange multipliers and nonconvex programs", SIAM Journal on Control 7 (1969). R. Horst and H. Tuy, Global Optimization: Deterministic Approaches, 3rd ed. (Springer, 1996). The theorem holds in any dimension. Lowering a function to its convex floor never lowers the floor itself. What the envelope loses is the shape between the minimizers, the band between the curve and the envelope. Relaxing a single unconstrained function is therefore exact at the root of a branch-and-bound tree, whatever the number of variables. The gap in practice comes from two sources. The constraints couple variables and let a relaxed constraint exploit the band. And the relaxation a solver computes is not the envelope. Both appear below.

The second basic fact is that the envelope depends on the domain, and in a quantifiable way.

Proposition 2.4.5 (dependence on the domain). (a) If \(B' \subseteq B\) are compact convex sets then \(\operatorname{vex}_{B'} f \ge \operatorname{vex}_B f\) on \(B'\). (b) If \(B_k \ni \bar x\) are boxes whose widths \(w(B_k) = \max_i (u_i - l_i)\) tend to zero, then \(\operatorname{vex}_{B_k} f(\bar x) \to f(\bar x)\). (c) If \(f\) is twice continuously differentiable on a box \(B_0\) and \(\nabla^2 f(x) + 2\bar\alpha I \succeq 0\) for all \(x \in B_0\) and some \(\bar\alpha \ge 0\), then for every box \(B = [l, u] \subseteq B_0\) and every \(x \in B\)

\[0 \;\le\; f(x) - \operatorname{vex}_B f\,(x) \;\le\; \bar\alpha \sum_{i=1}^n \frac{(u_i - l_i)^2}{4} \;\le\; \frac{n\bar\alpha}{4}\, w(B)^2 .\]

Proof. (a) The restriction to \(B'\) of a convex minorant of \(f\) on \(B\) is a convex minorant on \(B'\), so the supremum over the larger class of minorants is larger. (b) By Theorem 2.4.4, \(\min_{B_k} f \le \operatorname{vex}_{B_k} f(\bar x) \le f(\bar x)\), and \(\min_{B_k} f \to f(\bar x)\) by continuity. (c) The function \(L_{\bar\alpha}(x) = f(x) + \bar\alpha \sum_i (x_i - l_i)(x_i - u_i)\) is convex on \(B\), because its Hessian is \(\nabla^2 f + 2\bar\alpha I \succeq 0\), and it lies below \(f\) because each added term is nonpositive on \(B\). So \(\operatorname{vex}_B f \ge L_{\bar\alpha}\), and \(f - L_{\bar\alpha} = \bar\alpha \sum_i (x_i - l_i)(u_i - x_i)\) is a sum of downward parabolas, each at most \(\bar\alpha (u_i - l_i)^2/4\). ∎

Shrinking the box removes candidate chords, so the floor rises. For a smooth function it rises quadratically in the width. This quadratic rate is the reason spatial branching on a few variables can close a gap in a modest number of levels. It is also the yardstick for the convergence-order discussion at the end of this subsection. The function \(L_{\bar\alpha}\) in the proof of (c) is the αBB underestimator, which has a part of its own below.

The third fact is the one that makes everything after it necessary: envelopes do not add.

Proposition 2.4.6 (sums and scalars). For \(f, g\) continuous on compact convex \(B\) and \(\mu \ge 0\): \(\operatorname{vex}_B(f + g) \ge \operatorname{vex}_B f + \operatorname{vex}_B g\), with equality whenever \(g\) is affine, and \(\operatorname{vex}_B(\mu f) = \mu \operatorname{vex}_B f\). The inequality is strict in general: on \([-1, 1]\), \(f = x^2\) and \(g = -x^2\) give \(\operatorname{vex} f + \operatorname{vex} g = x^2 - 1\) while \(\operatorname{vex}(f + g) = 0\).

Proof. The function \(\operatorname{vex}_B f + \operatorname{vex}_B g\) is convex and lies below \(f + g\), hence below the largest such function. If \(g\) is affine, \(\operatorname{vex}_B(f + g) - g\) is convex and lies below \(f\), so \(\operatorname{vex}_B(f + g) \le \operatorname{vex}_B f + g\). For the example, the convex envelope of \(-x^2\) on \([-1, 1]\) is the chord through its endpoint values, the constant \(-1\), and \(x^2 - 1 < 0\) on the open interval. ∎

Proposition 2.4.6 on [−1, 1] with f = x² and g = −x². Relax each term, then add: vex f + vex g = x² + (−1) = x² − 1, where vex g = −1 is the chord through the endpoint values of −x². Add, then relax: vex(f + g) = vex 0 = 0, since f + g = 0 is its own envelope. The shaded band is the gap between the two routes, positive on the open interval.

Relaxing terms separately and adding is always valid and usually loose, because the chords chosen for one term are not the chords chosen for the other. Every factorable relaxation relaxes term by term. The attempt to do better is called simultaneous convexification and is taken up at the end of the catalogue below.

The first figure draws the one-variable picture. The curve is the quartic \(f(x) = 0.25x^4 - 1.6x^2 + 0.35x + 2.2\) on \([-2.6, 2.6]\), which has two wells. Its minimum is \(-0.995\) at \(x = -1.84\). On the whole interval the envelope is a single chord lying under both wells, tangent to the curve near each well. The band between curve and chord is widest between the wells: \(2.560\) deep at \(x = 0\), with a total band area of \(4.885\). The minimum of the envelope is \(-0.995\), the minimum of \(f\), as Theorem 2.4.4 requires. Move the "pieces" slider and the interval is split. On each piece the envelope is recomputed and the band shrinks (Proposition 2.4.5). A piece whose own envelope minimum already exceeds the incumbent \(-0.995\) is hatched, because a branch and bound would never refine it. The "how to split" buttons choose the midpoint of the piece with the largest gap, the point where that gap is widest, or a uniform grid. The figure stops splitting once every live piece is convex. The "relaxation" control switches from the envelope to the αBB underestimator, the subject of the part "Relaxing a whole function at once: αBB" below. In envelope mode the "alphaBB α" slider is dimmed.

A quartic with two wells, relaxed on each piece either by its convex envelope or by the alphaBB underestimator L = f + α(x − a)(x − b). Orange is what the relaxation cannot see, hatching marks pieces the incumbent rules out, and a dashed red L has α below α_min and is not convex. The sliders set the number of pieces and α, and the buttons choose where each split goes and which relaxation is drawn.

What to look for is that nothing coupled to \(x\) sees the curve. Anything constrained through \(x\) is bounded by the chord, two and a half units below the curve between the wells, and that is the gap a search has to pay back. The next part adds a second variable and a constraint for exactly that reason.

The bilinear term

The product of two variables is the term that appears most in applications: a quantity times a price, a flow times a concentration, a weight times a return, the fraction of a lot sold times its gain. Its envelopes on a box are known in closed form, and almost every other relaxation in this series is built from them. In this part a two-dimensional box is written \(R = [x^L, x^U] \times [y^L, y^U]\), with its bounds as superscripts rather than as the \(l\) and \(u\) of Definition 2.4.2, so that the planes stay readable. On \(R\) the four inequalities

\[\begin{aligned} w &\ge x^L y + x\,y^L - x^L y^L, &\qquad w &\ge x^U y + x\,y^U - x^U y^U, \\ w &\le x^U y + x\,y^L - x^U y^L, &\qquad w &\le x^L y + x\,y^U - x^L y^U \end{aligned} \tag{2.4.2}\]

are McCormick's planes for \(w = xy\). The theorem that they are the envelopes is Al-Khayyal and Falk's.The planes: G. P. McCormick, "Computability of global solutions to factorable nonconvex programs: Part I—Convex underestimating problems", Mathematical Programming 10 (1976). The envelope theorem: F. A. Al-Khayyal and J. E. Falk, "Jointly constrained biconvex programming", Mathematics of Operations Research 8 (1983).

Theorem 2.4.7 (Al-Khayyal and Falk, 1983). On \(R = [x^L, x^U] \times [y^L, y^U]\),

\[\begin{aligned} \operatorname{vex}_R(xy) &= \max\{\, x^L y + x\,y^L - x^L y^L,\;\; x^U y + x\,y^U - x^U y^U \,\}, \\ \operatorname{cav}_R(xy) &= \min\{\, x^U y + x\,y^L - x^U y^L,\;\; x^L y + x\,y^U - x^L y^U \,\}. \end{aligned}\]

Moreover: (i) \(xy - \operatorname{vex}_R(xy)\) and \(\operatorname{cav}_R(xy) - xy\) both attain their maximum, \((x^U - x^L)(y^U - y^L)/4\), at the centre of \(R\); (ii) the band \(\operatorname{cav}_R(xy) - \operatorname{vex}_R(xy)\) equals \((x^U - x^L)(y^U - y^L)/2\) at the centre; (iii) both envelopes agree with \(xy\) exactly on the boundary of \(R\), that is on all four edges, and are strictly inexact at every interior point.

Proof sketch. Validity of the four planes comes from products of nonnegative bound factors. On \(R\),

\[(x - x^L)(y - y^L) \ge 0, \qquad (x^U - x)(y^U - y) \ge 0, \qquad (x^U - x)(y - y^L) \ge 0, \qquad (x - x^L)(y^U - y) \ge 0,\]

and expanding each with \(xy\) written as \(w\) gives the four lines of (2.4.2). Tightness is checked on the unit square first. There the lower planes read \(w \ge 0\) and \(w \ge x + y - 1\), so their maximum \(\psi(x, y) = \max\{0, x + y - 1\}\) is a convex underestimator of \(xy\). Let \(g\) be any convex function with \(g \le xy\) on the square. A point with \(x + y \le 1\) is the convex combination of the corners \((0,0), (1,0), (0,1)\) with weights \(1 - x - y\), \(x\), \(y\). The product vanishes at those three corners, so \(g(x, y) \le (1 - x - y)\,g(0,0) + x\,g(1,0) + y\,g(0,1) \le 0 = \psi(x, y)\). A point with \(x + y \ge 1\) is the combination of \((1,0), (0,1), (1,1)\) with weights \(1 - y\), \(1 - x\), \(x + y - 1\), so \(g(x, y) \le (x + y - 1) \cdot 1 = \psi(x, y)\). Hence \(g \le \psi\), and \(\psi\) is the envelope. The concave side is the same argument with the upper triangles \(\{(0,0),(1,0),(1,1)\}\) and \(\{(0,0),(0,1),(1,1)\}\), on which the interpolants of the corner values are \(y\) and \(x\), giving \(\min\{x, y\}\). In a picture, the four corner points of the graph span a tetrahedron. Its two lower faces are the convex envelope and its two upper faces the concave one. For a general rectangle substitute \(x = x^L + (x^U - x^L)\xi\) and \(y = y^L + (y^U - y^L)\eta\). Then \(xy\) is \((x^U - x^L)(y^U - y^L)\,\xi\eta\) plus an affine function of \((\xi, \eta)\). By Proposition 2.4.6 the affine part passes through the envelope and the positive factor scales it, and undoing the substitution gives the displayed formulas. For (i): on the unit square, \(xy - \max\{0, x + y - 1\}\) equals \(xy\) where \(x + y \le 1\) and \((1 - x)(1 - y)\) where \(x + y \ge 1\), both maximized at \((\tfrac12, \tfrac12)\) with value \(\tfrac14\). The concave side is symmetric, and scaling gives the general formula. For (ii): at the centre of the unit square the lower envelope is \(\max\{0, 0\} = 0\) and the upper is \(\min\{\tfrac12, \tfrac12\} = \tfrac12\). For (iii): \(xy - (x^L y + x y^L - x^L y^L) = (x - x^L)(y - y^L)\) vanishes exactly on the edges \(x = x^L\) and \(y = y^L\), and likewise for the other three planes. Each envelope is therefore exact exactly where one of its two planes is, which is the union of the four edges. ∎

Theorem 2.4.16 below generalizes the corner argument to every function that is concave along each coordinate direction.

A bilinear function is a saddle. The convex floor of its graph over a rectangle is a roof of two flat panels resting on the four corners. The function sags a quarter of the rectangle's area above the floor at the centre. Half the rectangle's area separates the floor from the ceiling at the centre. Two facts are often misstated. The envelopes are exact on the whole boundary of the box, not only at its corners. And the quarter of the area is the error of each envelope. The band between the two, which is what a relaxed constraint gets to exploit, is half the area. The separation is proportional to the area of the box. Tightening both factors from width \(10\) to width \(1\) therefore divides it by \(100\), and tightening one factor divides it by \(10\). Section 2.6 is about how solvers tighten.

McCormick's planes for w = xy on the unit square (Theorem 2.4.7). Above, the convex envelope max{0, x + y − 1}: w ≥ 0 on one side of the anti-diagonal and w ≥ x + y − 1 on the other. Below, the concave envelope min{x, y}: w ≤ x above the diagonal and w ≤ y below it. Both envelopes equal xy on all four edges. At the centre they are 0 and 0.5 around xy = 0.25, so each misses by 1/4 and the band is 1/2.

(The term alone is exact; the constraint reaches into the band) Two computations make this concrete. On the unit square the planes read \(w \ge 0\), \(w \ge x + y - 1\), \(w \le x\), \(w \le y\). At the centre \((\tfrac12, \tfrac12)\) the overestimator gives \(\min\{x, y\} = 0.5\), the underestimator \(\max\{0, x + y - 1\} = 0\), and the product is \(0.25\). Each envelope misses by \(\tfrac14\) and the band is \(\tfrac12\). Now add the coupling. The running bilinear example asks for the largest \(xy\) on the unit square under the single constraint \(2x + y \le 1.2\). This example maximizes. On the constraint line \(y = 1.2 - 2x\) the product is \(1.2x - 2x^2\), which is largest at \(x = 0.3\), so the true maximum is \(0.18\) at \((0.3, 0.6)\). The gradient of \(xy\) there is \((0.6, 0.3)\), parallel to the normal \((2, 1)\) of the line: the line is tangent to the level curve \(xy = 0.18\). The relaxation replaces \(xy\) by \(w\) with the four planes. In a maximization only the two upper planes bind, so the relaxation maximizes \(\min\{x, y\}\) subject to \(2x + y \le 1.2\). That is attained at \(x = y = 0.4\) with value \(0.40\), more than twice the truth. The overstatement is a gap in the sense of Section 2.1. The term alone would be relaxed exactly at its optimum. The constraint is what reaches into the band.

(Products of constraint factors: the ladder of Section 4.7) The proof of Theorem 2.4.7 contained a second idea, and Section 4.7 is built on it. The McCormick planes are the products of pairs of bound factors \(x - x^L \ge 0\), \(x^U - x \ge 0\), \(y - y^L \ge 0\), \(y^U - y \ge 0\), with the product \(xy\) renamed \(w\). A constraint can be multiplied by a bound factor in the same way. In the running example, \((1.2 - 2x - y)\,x \ge 0\) holds on the feasible set. Write \(w\) for \(xy\) and \(X_{11}\) for \(x^2\), in the lifted notation \(X = x x^\top\) of Section 4.7, where \(w\) is the entry \(X_{12}\). The product then reads \(1.2x - 2X_{11} - w \ge 0\), a linear inequality the McCormick planes do not imply. Adding all such products of the five factors \(x\), \(1 - x\), \(y\), \(1 - y\), \(1.2 - 2x - y\) lowers the relaxed maximum from \(0.40\) to \(0.30\). Adding the two convexity facts \(X_{11} \ge x^2\) and \(X_{22} \ge y^2\) lowers it to \(0.18\), the exact value, with no branching at all. The first of the two already suffices: with \(X_{11} \ge x^2\) the product row reads \(w \le 1.2x - 2X_{11} \le 1.2x - 2x^2 \le 0.18\). This is the reformulation-linearization technique of Sherali and Adams. The ladder \(0.40 \to 0.30 \to 0.18\) is worked through in Section 4.7 with its own figure, and the rest of this subsection refers to it as the ladder of Section 4.7.H. D. Sherali and W. P. Adams, "A hierarchy of relaxations between the continuous and convex hull representations for zero-one programming problems", SIAM Journal on Discrete Mathematics 3 (1990); H. D. Sherali and A. Alameddine, "A new reformulation-linearization technique for bilinear programming problems", Journal of Global Optimization 2 (1992). That the products of bound factors are exactly the McCormick planes, and that products with constraint factors are not implied by them, is the content of the running example's ladder, recomputed for this post: McCormick 0.40, full level-1 RLT 0.30, RLT with \(X_{11} \ge x^2\) and \(X_{22} \ge y^2\) 0.18.

The ladder on the running example R2: max xy, 2x + y <= 1.2

  relaxed
  maximum
   0.40 --+-- McCormick: the products of pairs of the bound factors
          |   x, 1 - x, y, 1 - y, with the product xy renamed w
          |   + the products with the constraint factor
          |     1.2 - 2x - y, e.g. (1.2 - 2x - y) x >= 0, which
          v     reads 1.2x - 2X_11 - w >= 0 with X_11 for x^2
   0.30 --+-- level-1 RLT: all products of the five factors
          |
          |   + X_11 >= x^2 (and X_22 >= y^2; the first suffices):
          |     w <= 1.2x - 2X_11 <= 1.2x - 2x^2 <= 0.18
          |
          v
   0.18 --+-- exact: the true maximum, at (0.3, 0.6), with no
              branching at all

The figure below is the two-variable picture. Its upper panel is a heat map of the band \(\operatorname{cav}(xy) - \operatorname{vex}(xy)\) over the unit square, darker where the band is wider. The band is zero along the edges of every sub-box and \(1/(2k^2)\) at the centre of each sub-box when each axis is cut into \(k\) pieces. The constraint \(2x + y \le 1.2\) is drawn with its far side hatched. The blue dot is the true maximum \(0.18\) at \((0.3, 0.6)\), and the orange ring is where the relaxation puts it, \(0.400\) at \((0.4, 0.4)\). The lower panel is a slice at fixed \(y\), by default \(y = 0.60\), showing \(w = xy\) against \(x\) with the under- and overestimators the relaxation uses. At this slice the band is at most \(0.4000\) wide, and the estimators meet the curve only where \(x\) or \(y\) sits on an edge of a sub-box. In the default view the stats read: \(1 \times 1\) sub-boxes, widest band \(0.5000\), relaxed maximum \(0.400\), true maximum \(0.180\), gap \(122\%\). The gap is relative to the true value \(0.180\), which is the Gurobi and CPLEX convention of Section 2.1 (relative to the bound \(0.400\) it is \(55\%\)), and the table's gap column uses the same convention. Move the "pieces per axis" slider and the figure recomputes the relaxed maximum exactly on every sub-box. Its table reads, for \(k = 1, \dots, 6\): widest band \(0.5000\), \(0.1250\), \(0.0556\), \(0.0313\), \(0.0200\), \(0.0139\); relaxed maximum \(0.400\), \(0.233\), \(0.193\), \(0.192\), \(0.187\), \(0.185\); gap \(122\%\), \(30\%\), \(7\%\), \(6\%\), \(4\%\), \(3\%\).

The width of the McCormick envelope of w = x·y over the unit square, darker where looser: zero along the edges of every sub-box and 1/(2k²) at its centre, with the one constraint drawn and the far side hatched. The blue dot is the true maximum of x·y under the line. The orange ring is where the relaxation puts it. Below, a slice at fixed y shows the band the relaxation permits. Partition each axis and watch the band narrow by the square of the count, and the relaxed maximum come down.

Two different quantities fall in that table, and they fall differently. The band is a property of the relaxation of one term on one cell, and it falls exactly like \(1/k^2\).

Proposition 2.4.8 (the band of a piecewise McCormick relaxation). Partition \(R = [x^L, x^U] \times [y^L, y^U]\), of area \(A = (x^U - x^L)(y^U - y^L)\), into \(k_x \times k_y\) congruent cells, and on each cell \(C\) relax \(xy\) by the envelopes of Theorem 2.4.7 taken on \(C\). Then on every cell

\[\max_C \big(\operatorname{cav}_C(xy) - \operatorname{vex}_C(xy)\big) \;=\; \frac{A}{2 k_x k_y}, \qquad \max_C \big(xy - \operatorname{vex}_C(xy)\big) \;=\; \max_C \big(\operatorname{cav}_C(xy) - xy\big) \;=\; \frac{A}{4 k_x k_y},\]

all attained at the centre of the cell, the band only there. On the unit square with \(k \times k\) cells the widest band is \(1/(2k^2)\). With \(k \times 1\) cells, one factor partitioned, it is \(1/(2k)\).

Proof. A cell has sides \((x^U - x^L)/k_x\) and \((y^U - y^L)/k_y\), hence area \(A/(k_x k_y)\), and Theorem 2.4.7(i) and (ii) applied to the cell give the three values at its centre. For the uniqueness of the band's maximizer, take the unit square, where the band is \(\min\{x, y\} - \max\{0, x + y - 1\}\). Where \(x \le y\) and \(x + y \le 1\) it equals \(x\), and there \(x \le \tfrac12\). Where \(x \le y\) and \(x + y \ge 1\) it equals \(1 - y\), and there \(y \ge \tfrac12\). The case \(x \ge y\) is symmetric. The band is therefore at most \(\tfrac12\), with equality only at \((\tfrac12, \tfrac12)\), and an affine map onto a cell preserves this. ∎

The bound of the relaxed problem does not fall like \(1/k^2\). In the table, \(4 \times 4\) cells give \(0.1917\), barely better than the \(0.1926\) of \(3 \times 3\), while the band has fallen from \(0.0556\) to \(0.0313\). The bound is the optimum of a relaxed problem with a constraint. Where that optimum sits inside its cell decides how much of the band the constraint can exploit. The following script computes the figure's table exactly, cell by cell, by maximizing the concave overestimator over the polygon that the constraint cuts from each cell. The overestimator is the minimum of two planes, so its maximum over a polygon is at a vertex of the polygon or where the crease between the planes crosses an edge. The script also partitions one axis only, which is the univariate partitioning that solvers usually implement. It prints the band to five decimals, so the \(0.0313\) of the figure's table at \(k = 4\) appears as \(0.03125\). Both are \(1/32\).

# Piecewise McCormick on the running bilinear example.
#
# Maximize x*y subject to 2x + y <= 1.2 on the unit square. On every
# cell of a kx x ky grid the program maximizes the cell's concave
# McCormick overestimator, exactly, over the polygon that the constraint
# cuts from the cell, and reports the best cell (numpy only).

import numpy as np

# the constraint  A x + B y <= C
A, B, C = 2.0, 1.0, 1.2

def clip(poly, a, b, c):
    """Sutherland-Hodgman: keep the part of poly with a x + b y <= c."""
    out = []
    for i in range(len(poly)):
        P, Q = poly[i], poly[(i + 1) % len(poly)]
        fP, fQ = a*P[0] + b*P[1] - c, a*Q[0] + b*Q[1] - c
        if fP <= 0:
            out.append(P)
        if (fP < 0 < fQ) or (fQ < 0 < fP):
            t = fP/(fP - fQ)
            out.append((P[0] + t*(Q[0] - P[0]),
                        P[1] + t*(Q[1] - P[1])))
    return out

def cell_max(xL, xU, yL, yU):
    """Exact maximum of the concave overestimator over the cell
    intersected with {2x + y <= 1.2}, as (value, x, y)."""
    # the two upper McCormick planes of the cell
    p1 = lambda x, y: xU*y + x*yL - xU*yL
    p2 = lambda x, y: xL*y + x*yU - xL*yU
    poly = clip([(xL, yL), (xU, yL), (xU, yU), (xL, yU)], A, B, C)
    if len(poly) < 3:
        return None

    # vertices, plus where the crease p1 = p2 crosses an edge
    cands = list(poly)
    ca, cb, cc = -(yU - yL), (xU - xL), xU*yL - xL*yU
    for i in range(len(poly)):
        P, Q = poly[i], poly[(i + 1) % len(poly)]
        fP, fQ = ca*P[0] + cb*P[1] - cc, ca*Q[0] + cb*Q[1] - cc
        if fP*fQ < 0:
            t = fP/(fP - fQ)
            cands.append((P[0] + t*(Q[0] - P[0]),
                          P[1] + t*(Q[1] - P[1])))
    return max((min(p1(x, y), p2(x, y)), x, y) for x, y in cands)

print("true maximum 0.18 at (0.3, 0.6); the relaxed maximum is the best")
print("cell's value, which is the optimum of the piecewise McCormick MILP")
print("and the LP bound of its formulation with the constraint row copied")
print("into every cell")
print()
print("cells    band at a    relaxed maximum             excess")
print("         cell centre")

grids = [(1, 1), (2, 2), (3, 3), (4, 4), (8, 8), (16, 16),
         (2, 1), (4, 1), (8, 1), (16, 1)]
for kx, ky in grids:
    best = max(r for i in range(kx) for j in range(ky)
               if (r := cell_max(i/kx, (i + 1)/kx,
                                 j/ky, (j + 1)/ky)) is not None)
    w, x, y = best
    print(f"{kx:2d} x {ky:<2d}  {1/(2*kx*ky):.5f}      "
          f"{w:.4f} at ({x:.3f}, {y:.3f})    over 0.18: {w - 0.18:.4f}")
true maximum 0.18 at (0.3, 0.6); the relaxed maximum is the best
cell's value, which is the optimum of the piecewise McCormick MILP
and the LP bound of its formulation with the constraint row copied
into every cell

cells    band at a    relaxed maximum             excess
         cell centre
 1 x 1   0.50000      0.4000 at (0.400, 0.400)    over 0.18: 0.2200
 2 x 2   0.12500      0.2333 at (0.233, 0.733)    over 0.18: 0.0533
 3 x 3   0.05556      0.1926 at (0.289, 0.622)    over 0.18: 0.0126
 4 x 4   0.03125      0.1917 at (0.317, 0.567)    over 0.18: 0.0117
 8 x 8   0.00781      0.1833 at (0.317, 0.567)    over 0.18: 0.0033
16 x 16  0.00195      0.1807 at (0.296, 0.608)    over 0.18: 0.0007
 2 x 1   0.25000      0.3000 at (0.300, 0.600)    over 0.18: 0.1200
 4 x 1   0.12500      0.2333 at (0.367, 0.467)    over 0.18: 0.0533
 8 x 1   0.06250      0.2100 at (0.320, 0.560)    over 0.18: 0.0300
16 x 1   0.03125      0.1944 at (0.289, 0.622)    over 0.18: 0.0144

The script's per-cell work is a handful of arithmetic operations, and the cells are independent. A partition into \(k^2\) cells is therefore \(k^2\) identical small problems that can be solved at once, an observation taken up at the end of the subsection. The excess over the truth falls by a factor of about four from \(k = 1\) to \(2\) and again from \(2\) to \(3\), then by about \(7\%\) from \(3\) to \(4\). The one-sided partition is roughly first order: \(0.30\), \(0.2333\), \(0.21\), \(0.1944\) for \(k = 2, 4, 8, 16\). The worst band per cell falls by \(k^2\). The bound of a particular problem falls irregularly, depending on where its optimum lies relative to the grid.

(A partition as one relaxation: piecewise McCormick) A partition can be built into a single relaxation rather than explored as a tree. Number the cells \(C_1, \dots, C_K\). For a partition of \(x\) alone into \(K\) pieces the cells are the strips \([a_{k-1}, a_k] \times [y^L, y^U]\), which give the \(k \times 1\) rows of the table. For a \(k \times k\) grid there are \(K = k^2\) cells, one selector per cell. Introduce one binary selector \(z_k\) per cell with \(\sum_k z_k = 1\), and copies \(x_k, y_k, w_k\) of the variables with \(x = \sum_k x_k\), \(y = \sum_k y_k\) and \(w = \sum_k w_k\). Each copy is confined to its cell scaled by its selector: \(x_k^L z_k \le x_k \le x_k^U z_k\) and \(y_k^L z_k \le y_k \le y_k^U z_k\). It satisfies the four planes (2.4.2) of its cell, written for \((x_k, y_k, w_k)\) with every bound multiplied by \(z_k\). Every linear constraint on \((x, y)\) is copied into the cells in the same way, with its right-hand side multiplied by the selector. In the running example this is \(2x_k + y_k \le 1.2\,z_k\) for each \(k\). For integral \(z\) the formulation is the McCormick relaxation of \(w = xy\) on the selected cell, intersected with the constraint. It is called piecewise McCormick.

Proposition 2.4.9 (piecewise McCormick as one relaxation). Let \(P_k = \{(x, y, w) : (x, y) \in C_k,\ (2.4.2) \text{ holds on } C_k\}\) be the McCormick set of cell \(C_k\), and let \(H\) be the set of \((x, y, w)\) satisfying the linear constraints. The continuous relaxation of the piecewise McCormick formulation, with \(z \in [0, 1]^K\), projects onto \(\operatorname{conv}\big(\bigcup_k (P_k \cap H)\big)\). A linear objective over it therefore attains the best of the \(K\) per-cell optima, so the values printed above are both the optimum of the mixed-integer program and the bound of its LP relaxation. If the linear constraints are kept only in the aggregated variables \((x, y)\), the relaxation projects onto \(\operatorname{conv}(\bigcup_k P_k) \cap H\) instead, which equals \(P_0 \cap H\) for the McCormick set \(P_0\) of the whole box. The bound is then \(0.40\) for every \(K\).

Proof sketch. The first claim is Balas' theorem on the convex hull of a union of polytopes, proved in Section 4.2. The formulation with a scaled copy of every constraint is the hull formulation of the disjunction \(\bigvee_k \big[(x, y, w) \in P_k \cap H\big]\). For the second claim, each \(P_k\) is contained in \(P_0\), because the envelopes on a cell are at least as tight as those on the whole box (Proposition 2.4.5(a)), so \(\operatorname{conv}(\bigcup_k P_k) \subseteq P_0\). Conversely, \(P_0\) is the convex hull of the graph \(\{(x, y, xy)\}\) over the whole box, by Theorem 2.4.7 and Proposition 2.4.3, and every point of that graph lies in some \(P_k\), so \(P_0 \subseteq \operatorname{conv}(\bigcup_k P_k)\). On the running example the point \((0.4, 0.4, 0.4) = 0.6\,(0, 0, 0) + 0.4\,(1, 1, 1)\) lies in \(P_0\) and on the line \(2x + y = 1.2\), so the aggregated relaxation still reports \(0.40\). ∎

Proposition 2.4.9: why the aggregated bound stays at 0.40. The graph points (0, 0, 0) and (1, 1, 1), each in the set P_k of its cell, mix to (0.4, 0.4, 0.4) = 0.6 (0, 0, 0) + 0.4 (1, 1, 1), which lies on the line 2x + y = 1.2 with w = 0.4. With the row 2x + y ≤ 1.2 on the sums (x, y) only, the mixture is feasible and the LP bound is 0.40 for every K. With the row copied into every cell, 2x_k + y_k ≤ 1.2 z_k, the copy of (1, 1, 1) is (0.4, 0.4, 0.4) with z_k = 0.4 and violates it, and the bound is the best cell's: 0.2333 on 2 × 2 cells.

(What a partition costs) The price is easy to count. The formulation has \(K\) binaries and \(3K\) continuous copies per bilinear term, plus a copy of every linear row per cell. With \(d\) variables partitioned into \(K\) pieces each there are \(K^d\) cells, and the branch and bound over the selectors is exponential in their number in the worst case. Logarithmic encodings reduce the binaries to \(\lceil \log_2 K \rceil\) per term. The normalized multiparametric disaggregation technique discretizes one variable digit by digit in base ten and relaxes only the residual with (2.4.2). Adaptive partitioning concentrates the pieces around the incumbent. None of them changes the count of cells a solver may have to visit.M. L. Bergamini, P. Aguirre and I. Grossmann, "Logic-based outer approximation for globally optimal synthesis of process networks", Computers & Chemical Engineering 29 (2005); D. S. Wicaksono and I. A. Karimi, "Piecewise MILP under- and overestimators for global optimization of bilinear programs", AIChE Journal 54 (2008); R. Misener, J. P. Thompson and C. A. Floudas, "APOGEE: Global optimization of standard, generalized, and extended pooling problems via linear and logarithmic partitioning schemes", Computers & Chemical Engineering 35 (2011); J. P. Vielma, S. Ahmed and G. Nemhauser, "Mixed-integer models for nonseparable piecewise-linear optimization: unifying framework and extensions", Operations Research 58 (2010); S. Kolodziej, P. M. Castro and I. E. Grossmann, "Global optimization of bilinear programs with a multiparametric disaggregation technique", Journal of Global Optimization 57 (2013); P. M. Castro, "Normalized multiparametric disaggregation: an efficient relaxation for mixed-integer bilinear problems", Journal of Global Optimization 64 (2016); H. Nagarajan, M. Lu, S. Wang, R. Bent and K. Sundar, "An adaptive, multivariate partitioning algorithm for global optimization of nonconvex programs", Journal of Global Optimization 74 (2019). Section 4.6 treats piecewise-linear models and these relaxations in detail. Spatial branching buys accuracy quadratically in the width of each cell and pays for it exponentially in the number of nonconvex variables. That ratio is the reason nonconvex problems with a few hundred bilinear terms are hard while linear problems with tens of thousands of integer variables are solved routinely (Section 1.1). The tree of Section 3.5 grows at the same rate.

(The pooling problem: two terms relaxed separately) Haverly's pooling problem, the running example R4, shows the same mechanism on the canonical bilinear instance. Section 3.5 gives the data, the three formulations and a figure that solves it in the page. Here is the one lesson. Two inputs with sulphur contents of \(3\) and \(1\) per cent are blended in one pool whose quality \(q \in [1, 3]\) is itself a variable. The pool sends \(y_X \in [0, 100]\) units to a product X and \(y_Y \in [0, 200]\) to a product Y. The sulphur the pool delivers is \(q\,y_X\) and \(q\,y_Y\), two bilinear terms. The problem maximizes profit, and its maximum is \(400\), at \(q = 1\). The McCormick relaxation replaces the two products by \(w_X\) and \(w_Y\) with the eight planes of (2.4.2) and reports \(500\), at \(q = 2\), \(y_X = 50\), \(y_Y = 100\), \(w_X = 150\) and \(w_Y = 100\). It has set \(w_X = 3\,y_X\) and \(w_Y = 1 \cdot y_Y\) at once: one pool at quality \(3\) feeding X and another at quality \(1\) feeding Y, from the same material, which no single pool quality can do. Each term is relaxed exactly on the edges of its own box, and the pair is not relaxed jointly. Splitting the interval of \(q\) at \(2\) and relaxing each half separately, which Section 3.5 calls a spatial branch, gives \(400\) on \([1, 2]\) and \(100\) on \([2, 3]\). The gap is closed.C. A. Haverly, "Studies of the behavior of recursion for the pooling problem", ACM SIGMAP Bulletin 25 (1978); M. Tawarmalani and N. V. Sahinidis, Convexification and Global Optimization in Continuous and Mixed-Integer Nonlinear Programming (Kluwer, 2002), Chapter 9, for the p-, q- and pq-formulations. The bounds were recomputed for this post: the McCormick relaxations of the p- and pq-formulations both give \(500\), \(1000\) and \(800\) on Haverly's instances 1, 2 and 3, against the optima \(400\), \(600\) and \(750\). The figure in Section 3.5 computes the same numbers, writes \(q\) for the pool quality as here, and splits \(q\) at \(2\) in the page.

Haverly's pool (R4), its McCormick relaxation, and one split of q

  the model: one pool whose quality q is a variable
    sulphur 3 % --\                         /--> X: y_X in [0, 100]
                   >--> pool, q in [1, 3] --
    sulphur 1 % --/                         \--> Y: y_Y in [0, 200]
    sulphur out of the pool: q y_X and q y_Y, two bilinear terms

  the relaxation: w_X, w_Y and the eight planes (2.4.2); its
  optimum 500 at q = 2, y_X = 50, y_Y = 100 delivers
    w_X = 150 = 3 y_X     to X, as if from a pool at quality 3
    w_Y = 100 = 1 * y_Y   to Y, as if from a pool at quality 1
  from the same material, which no single q can do

  one spatial branch on q, at 2:

                     q in [1, 3]: bound 500
                    /                      \
     q in [1, 2]: bound 400          q in [2, 3]: bound 100
     = the maximum 400, at q = 1:
     the gap is closed

Factorable functions and McCormick's rules

The bilinear term is one node of a larger construction. McCormick's 1976 paper relaxes any function assembled from sums, products and compositions with library functions. Each operation is relaxed on its own interval, and the relaxations compose.

Definition 2.4.10. Fix a library \(\mathcal U\) of univariate functions, for example \(\exp\), \(\log\), \(x^p\), \(\sqrt{\cdot}\), \(1/x\), \(\sin\), \(\cos\), \(|x|\). For each of them a convex underestimator and a concave overestimator on any interval of its domain are known. A function \(f : B \to \mathbb{R}\) is factorable if there is a finite sequence \(v_1, \dots, v_N\) with \(v_i = x_i\) for \(i \le n\) and, for \(k > n\), either \(v_k = v_i + v_j\), \(v_k = v_i v_j\), \(v_k = c\,v_i\) or \(v_k = T(v_i)\) with \(T \in \mathcal U\) and \(i, j < k\), such that \(f = v_N\). Implementations call the sequence a tape. Read as a graph, with an edge from each operand to the node that uses it, it is a directed acyclic graph, the expression DAG. Its leaves are variables and constants and its internal nodes are operations. The natural interval extension of the DAG evaluates the expression with every operation replaced by its range on intervals (Definition 2.6.3). It encloses the range of every node but overestimates it whenever a variable appears more than once, which Section 2.6 calls the dependency problem and treats in full.McCormick (1976), cited above. Every closed-form expression over the library is factorable; so are implicit functions and algorithms with a fixed number of iterations, which is the subject of A. Mitsos, B. Chachuat and P. I. Barton, "McCormick-based relaxations of algorithms", SIAM Journal on Optimization 20 (2009), discussed at the end of this subsection. Interval arithmetic with directed rounding, and the dependency problem in full, are in Section 2.6.

A solver that receives a model does two things to its DAG before relaxing anything. It rewrites it, merging repeated subexpressions, recognizing \(x \cdot x\) as \(x^2\), and reordering sums and products. Then it walks the graph to detect which nodes are convex or concave, so that a convex node can be kept as it is or linearized by tangents instead of relaxed. The same walk is where a solver recognizes that the constraint \(xy \ge 1\) with \(x, y \ge 0\) of Section 1.2 is a rotated second-order cone, and keeps it convex instead of relaxing the product (the SCIP 10 detection quoted in Section 1.2).R. Fourer, C. Maheshwari, A. Neumaier, D. Orban and H. Schichl, "Convexity and concavity detection in computational graphs: tree walks for convexity assessment", INFORMS Journal on Computing 22 (2010). Both steps belong to presolve and are treated in Section 2.5. Here the DAG is taken as given, and the question is how to relax each node.

Two architectures turn the relaxed nodes into a number, and the rest of this subsection keeps them apart.

Definition 2.4.11 (the two architectures). The auxiliary variable method introduces one new variable per DAG node and adds the node's envelope inequalities as constraints. For a product these are the four planes (2.4.2). For a univariate node they are its underestimator and overestimator, or tangent and secant cuts for them. The result is a linear or convex program in a lifted space, solved by an LP solver at every node of the tree. The pointwise method keeps the original variables. By one forward sweep over the DAG at a chosen point \(\bar x\) it evaluates the values of the convex and concave relaxations of every node together with a subgradient and a supergradient. Here a vector \(s\) is a subgradient of a convex function \(g\) at \(\bar x\) if \(g(x) \ge g(\bar x) + s^\top (x - \bar x)\) for all \(x\) in the domain. The set of subgradients is the subdifferential \(\partial g(\bar x)\). A supergradient of a concave function and its superdifferential are the mirror notions. The pointwise method obtains a bound from the supporting hyperplanes it has computed, in closed form or by a small LP in the original variables.

The auxiliary variable method is the architecture of BARON, Couenne and SCIP, and Section 3.5's spatial branch and bound is written around it.E. M. B. Smith and C. C. Pantelides, "A symbolic reformulation/spatial branch-and-bound algorithm for the global optimisation of nonconvex MINLPs", Computers & Chemical Engineering 23 (1999); M. Tawarmalani and N. V. Sahinidis, "A polyhedral branch-and-cut approach to global optimization", Mathematical Programming 103 (2005). The pointwise method is the architecture of MAiNGO and EAGO, cited below. Both are built from the same two rules, which follow.

Definition 2.4.11: the two architectures on f = x e^y, unit square

  the DAG:  x -----------------------+
                                     +--> v_4 = x v_3 = f
            y ---> v_3 = exp(y) -----+
                   v_3 in [1, e]

  auxiliary variable method        pointwise method
  -------------------------        ----------------
  variables x, y, v_3, v_4         variables x, y only
  v_3: tangent cuts below,         one forward sweep at a point
       the secant 1 + (e - 1) y    xbar: cv, cc, s and t at every
       above                       node; at (0.5, 0.5) the root
  v_4: the four planes (2.4.2)     has cv = 0.5000, s = (1, 0)
       on [0, 1] x [1, e]
           |                            |
           v                            v
  an LP in (x, y, v_3, v_4),       the affine bound (2.4.3), here
  solved at every node of the      0.0000, or a small LP in (x, y)
  tree: BARON, Couenne, SCIP       over several points: MAiNGO, EAGO

Theorem 2.4.12 (McCormick, 1976: the composition rule). Let \(B\) be compact convex, let \(v : B \to \mathbb{R}\) have a convex underestimator \(v^{\mathrm{cv}}\) and a concave overestimator \(v^{\mathrm{cc}}\) on \(B\), and let \([v^L, v^U] \supseteq v(B)\). Let \(T : [v^L, v^U] \to \mathbb{R}\) have a convex underestimator \(e\) and a concave overestimator \(E\) on \([v^L, v^U]\), with \(z_{\min} \in \operatorname{argmin} e\) and \(z_{\max} \in \operatorname{argmax} E\). Write \(\operatorname{mid}\{a, b, c\}\) for the median of three numbers. Then

\[u(x) = e\big(\operatorname{mid}\{v^{\mathrm{cv}}(x),\, v^{\mathrm{cc}}(x),\, z_{\min}\}\big), \qquad o(x) = E\big(\operatorname{mid}\{v^{\mathrm{cv}}(x),\, v^{\mathrm{cc}}(x),\, z_{\max}\}\big)\]

are a convex underestimator and a concave overestimator of \(T \circ v\) on \(B\).

Proof. Fix \(x\). The interval \(I(x) = [v^{\mathrm{cv}}(x), v^{\mathrm{cc}}(x)]\) contains \(v(x)\), and so does \(J(x) = I(x) \cap [v^L, v^U]\), on which \(e\) is defined. The convex function \(e\) is nonincreasing to the left of \(z_{\min}\) and nondecreasing to the right of it, so its minimum over \(J(x)\) is attained at the projection of \(z_{\min}\) onto \(J(x)\). That projection is the median \(\operatorname{mid}\{v^{\mathrm{cv}}(x), v^{\mathrm{cc}}(x), z_{\min}\}\). The median lies in \(I(x)\), and it lies in \([v^L, v^U]\) because it either equals \(z_{\min}\) or lies between \(z_{\min}\) and \(v(x)\), both of which do. Hence

\[u(x) \;=\; \min\{\, e(z) : (x, z) \in D \,\}, \qquad D = \{(x, z) : x \in B,\ \max\{v^{\mathrm{cv}}(x), v^L\} \le z \le \min\{v^{\mathrm{cc}}(x), v^U\}\}.\]

The set \(D\) is convex because \(v^{\mathrm{cv}}\) is convex and \(v^{\mathrm{cc}}\) is concave, and \((x, z) \mapsto e(z)\) is jointly convex. The partial minimum of a jointly convex function over a convex set is convex in the remaining variable, so \(u\) is convex. It underestimates: \(v(x) \in J(x)\) gives \(u(x) \le e(v(x)) \le T(v(x))\). The concave side is the same argument applied to \(-T\). ∎

Theorem 2.4.12: the median finds the minimum of e on an interval

    e(z), convex
        *                                                   *
          *                                               *
            * *                                       * *
                *                                   *
                  * *                           * *
                      * * *               * * *
                            * * * * * * *
                                  ^
                                z_min
    [v^cv(x), v^cc(x)] in three positions:
      z = v^cc(x),         z = z_min,           z = v^cv(x),
      the right end        inside               the left end
  z = mid{v^cv(x), v^cc(x), z_min}: the point of the interval
  nearest z_min; u(x) = e(z)

(What the composition rule says) Knowing only that the inner value lies in \([v^{\mathrm{cv}}(x), v^{\mathrm{cc}}(x)]\), the safest convex estimate of \(T\) of it is the smallest value \(e\) takes on that interval. In the monotone cases the rule simplifies. If \(e\) is nondecreasing, as it is when \(T\) is increasing and \(e\) is its envelope or its secant, then \(z_{\min} = v^L\) and \(u = e(\max\{v^{\mathrm{cv}}(x), v^L\})\). This is \(e(v^{\mathrm{cv}}(x))\) whenever the inner relaxation respects its interval. If \(e\) is nonincreasing, as for a decreasing \(T\) with the same choices, then \(u = e(v^{\mathrm{cc}}(x))\) in the same sense. McCormick's paper states the result through these monotonicity cases. The \(\operatorname{mid}\) form is the one implementations use.Mitsos, Chachuat and Barton (2009), cited above; J. K. Scott, M. D. Stuber and P. I. Barton, "Generalized McCormick relaxations", Journal of Global Optimization 51 (2011), which treats the interval and the relaxation as one object and gives the calculus in this form. The rule needs the interval \([v^L, v^U]\) to build \(e\) and \(E\) on. This is why McCormick relaxations always travel together with interval bounds, and why a loose interval, from the dependency problem, weakens the outer relaxation even when the inner relaxations are tight.

(The product rule, in words) The product rule keeps each of the four planes of (2.4.2) and, inside each plane, replaces every factor by whichever of the factor's two relaxations keeps the expression convex, for the two lower planes, or concave, for the two upper ones. The sign of the coefficient in front of the factor decides which relaxation that is.

Theorem 2.4.13 (McCormick, 1976: the product rule). Let \(v_1, v_2 : B \to \mathbb{R}\) have relaxations \((v_k^{\mathrm{cv}}, v_k^{\mathrm{cc}})\) and bounds \(v_1 \in [a_1, b_1]\), \(v_2 \in [a_2, b_2]\) on \(B\). Then \(v_1 v_2\) has the convex underestimator and concave overestimator

\[\begin{aligned} (v_1 v_2)^{\mathrm{cv}} &= \max\{\alpha_1 + \alpha_2 - a_1 a_2,\;\ \beta_1 + \beta_2 - b_1 b_2\}, \\ &\quad \alpha_1 = \min\{a_2 v_1^{\mathrm{cv}}, a_2 v_1^{\mathrm{cc}}\},\quad \alpha_2 = \min\{a_1 v_2^{\mathrm{cv}}, a_1 v_2^{\mathrm{cc}}\},\quad \beta_1 = \min\{b_2 v_1^{\mathrm{cv}}, b_2 v_1^{\mathrm{cc}}\},\quad \beta_2 = \min\{b_1 v_2^{\mathrm{cv}}, b_1 v_2^{\mathrm{cc}}\}, \\ (v_1 v_2)^{\mathrm{cc}} &= \min\{\gamma_1 + \gamma_2 - a_1 b_2,\;\ \delta_1 + \delta_2 - b_1 a_2\}, \\ &\quad \gamma_1 = \max\{b_2 v_1^{\mathrm{cv}}, b_2 v_1^{\mathrm{cc}}\},\quad \gamma_2 = \max\{a_1 v_2^{\mathrm{cv}}, a_1 v_2^{\mathrm{cc}}\},\quad \delta_1 = \max\{a_2 v_1^{\mathrm{cv}}, a_2 v_1^{\mathrm{cc}}\},\quad \delta_2 = \max\{b_1 v_2^{\mathrm{cv}}, b_1 v_2^{\mathrm{cc}}\}. \end{aligned}\]

Proof sketch. On the rectangle \([a_1, b_1] \times [a_2, b_2]\) the planes \(a_2 z_1 + a_1 z_2 - a_1 a_2\) and \(b_2 z_1 + b_1 z_2 - b_1 b_2\) lie below \(z_1 z_2\) by Theorem 2.4.7. In the first plane replace \(a_2 v_1(x)\) by \(\min\{a_2 v_1^{\mathrm{cv}}(x), a_2 v_1^{\mathrm{cc}}(x)\}\). When \(a_2 \ge 0\) this is \(a_2 v_1^{\mathrm{cv}}(x) \le a_2 v_1(x)\) and is convex. When \(a_2 < 0\) it is \(a_2 v_1^{\mathrm{cc}}(x) \le a_2 v_1(x)\) and is convex as a negative multiple of a concave function. Sums and maxima of convex functions are convex, so each of the two expressions is a convex underestimator of \(v_1 v_2\), and so is their maximum. The concave side is symmetric. ∎

A first check is the case where the factors are the variables themselves. With \(v_1 = x\) and \(v_2 = y\) on \([a_1, b_1] \times [a_2, b_2]\) the relaxations are the identity, \(v_k^{\mathrm{cv}} = v_k^{\mathrm{cc}} = v_k\). Then \(\alpha_1 = a_2 x\) and \(\alpha_2 = a_1 y\), so the first piece is \(a_2 x + a_1 y - a_1 a_2\), the first plane of (2.4.2). The other three pieces are the other three planes, and the rule collapses to Theorem 2.4.7. The second check applies the rule by hand to \(f(x, y) = x e^y\) on the unit square, factored as \(v_3 = \exp(y)\) and \(v_4 = x v_3\). On \(y \in [0, 1]\) the exponential is convex and increasing, so by Theorem 2.4.12 \(v_3^{\mathrm{cv}} = e^y\), \(v_3^{\mathrm{cc}} = 1 + (e - 1)y\), the secant, and \(v_3 \in [1, e]\). In the product rule with \(v_1 = x \in [0, 1]\) and \(v_2 = v_3 \in [1, e]\): \(\alpha_1 = \min\{x, x\} = x\), \(\alpha_2 = \min\{0, 0\} = 0\), \(\beta_1 = e\,x\), \(\beta_2 = \min\{e^y, 1 + (e - 1)y\} = e^y\), \(\gamma_1 = e\,x\), \(\gamma_2 = 0\), \(\delta_1 = x\) and \(\delta_2 = \max\{e^y, 1 + (e - 1)y\} = 1 + (e - 1)y\). Hence

\[f^{\mathrm{cv}}(x, y) = \max\{\, x,\; e\,x + e^y - e \,\}, \qquad f^{\mathrm{cc}}(x, y) = \min\{\, e\,x,\; x + (e - 1)\, y \,\}.\]

At \((0.5, 0.5)\) these give \(0.5\) and \(1.3591\) around \(f = 0.8244\).

Two remarks are in order. The rule is the bilinear envelope applied to the factors and then relaxed again through the factors' own relaxations. The result is in general not the envelope of \(v_1 v_2\): the true convex envelope of \(x e^y\), derived in the generating-set part below, exceeds \(f^{\mathrm{cv}}\) by up to \(0.137\) on the unit square. The rule also treats the two factors one at a time. Tsoukalas and Mitsos replace it by a multivariate rule, \(u(x) = \min\{\operatorname{vex} F(z) : z_k \in [v_k^{\mathrm{cv}}(x), v_k^{\mathrm{cc}}(x)]\}\) for any outer function \(F\) of several inner functions. It is convex by the same partial-minimization argument as in Theorem 2.4.12 and is at least as tight as the binary rule. For a product, the binary rule computes the maximum of two separately minimized planes and the multivariate rule the minimum of their maximum, and a min–max is never below a max–min.A. Tsoukalas and A. Mitsos, "Multivariate McCormick relaxations", Journal of Global Optimization 59 (2014); the convergence analysis of the multivariate rule is J. Najman and A. Mitsos, "Convergence analysis of multivariate McCormick relaxations", Journal of Global Optimization 66 (2016).

(From functions to a number: the pointwise bound) The rules give functions, but a bound needs a number. In the auxiliary variable method the number is the optimum of the lifted LP. In the pointwise method it comes from supporting hyperplanes, and the following theorem is what makes the sweep carry their slopes.

Theorem 2.4.14 (Mitsos, Chachuat and Barton, 2009: subgradient propagation). Fix a box \(B\) and a point \(\bar x \in B\). Suppose each node of a factorable DAG carries, besides its interval \([v^L, v^U]\) and its values \(v^{\mathrm{cv}}(\bar x) \le v(\bar x) \le v^{\mathrm{cc}}(\bar x)\), a subgradient \(s \in \partial v^{\mathrm{cv}}(\bar x)\) and a supergradient \(t\) of \(v^{\mathrm{cc}}\) at \(\bar x\). Then the following vectors are a subgradient of the node's convex relaxation and a supergradient of its concave relaxation at \(\bar x\). For a sum \(v_i + v_j\): \(s_i + s_j\) and \(t_i + t_j\). For a scalar multiple \(c\,v_i\): \((c\,s_i, c\,t_i)\) if \(c \ge 0\) and \((c\,t_i, c\,s_i)\) if \(c < 0\). For a product \(v_i v_j\): in the piece that attains the maximum in the formula for \((v_i v_j)^{\mathrm{cv}}\) of Theorem 2.4.13, replace each term \(c\,v_m^{\mathrm{cv}}\) by \(c\,s_m\) and each term \(c\,v_m^{\mathrm{cc}}\) by \(c\,t_m\), and add the results. The same replacement in the piece that attains the minimum in \((v_i v_j)^{\mathrm{cc}}\) gives the supergradient. For a composition \(T(v_i)\) with \(z = \operatorname{mid}\{v_i^{\mathrm{cv}}(\bar x), v_i^{\mathrm{cc}}(\bar x), z_{\min}\}\) as in Theorem 2.4.12: \(a\,s_i\) if \(z = v_i^{\mathrm{cv}}(\bar x)\), \(a\,t_i\) if \(z = v_i^{\mathrm{cc}}(\bar x)\) and \(0\) if \(z = z_{\min}\), where \(a\) is a one-sided derivative of \(e\) at \(z\) (the right one in the first case, the left one in the second), and likewise for \(E\) with \(z_{\max}\). Applied node by node along the tape, these rules give at the root a vector \(s_N\) that is a subgradient of \(f^{\mathrm{cv}}\) at \(\bar x\), hence

\[a(x) \;=\; f^{\mathrm{cv}}(\bar x) + s_N^\top (x - \bar x) \;\le\; f^{\mathrm{cv}}(x) \;\le\; f(x) \quad (x \in B), \qquad \min_{x \in B} a(x) \;=\; f^{\mathrm{cv}}(\bar x) + \sum_{i=1}^n \min\{ s_{N,i}(l_i - \bar x_i),\ s_{N,i}(u_i - \bar x_i) \}. \tag{2.4.3}\]

Proof sketch. Each rule is a finite composition of sums, nonnegative scalings, pointwise maxima and minima of convex functions, and compositions with univariate convex functions that are monotone on the relevant side. The subdifferential calculus for these operations gives an element of the subdifferential at each step: the sum rule, the selection rule for a pointwise maximum, and the chain rule for a monotone convex composition. Interior points of \(B\) need no further assumption and boundary points a mild one. The display is the subgradient inequality for the convex function \(f^{\mathrm{cv}}\), followed by the minimization of an affine function over a box, which is done coordinate by coordinate. ∎

A convex relaxation whose value and slope are known at one point already gives a valid supporting hyperplane, and a hyperplane is minimized over a box without any linear programming. Several points give several hyperplanes and a small LP in the \(n\) original variables, which is what MAiNGO and EAGO solve. The interval bounds can themselves be tightened inside the sweep by using the subgradients. Constraint information can be propagated backwards through the DAG to shrink the relaxations of intermediate nodes.Mitsos, Chachuat and Barton (2009), cited above. The linearization strategies and the hybrid with the auxiliary variable method: J. Najman, D. Bongartz and A. Mitsos, "Linearization of McCormick relaxations and hybridization with the auxiliary variable method", Journal of Global Optimization 80 (2021). Interval tightening from subgradients: J. Najman and A. Mitsos, "Tighter McCormick relaxations through subgradient propagation", Journal of Global Optimization 75 (2019). Reverse propagation: A. Wechsung, J. K. Scott, H. A. J. Watson and P. I. Barton, "Reverse propagation of McCormick relaxations", Journal of Global Optimization 63 (2015). EAGO: M. E. Wilhelm and M. D. Stuber, "EAGO.jl: easy advanced global optimization in Julia", Optimization Methods and Software 37 (2022). The reference implementation of the sweep is MC++: B. Chachuat, MC++, library for construction, manipulation and evaluation of factorable functions, version 5, github.com/omega-icl/mcpp (README read 5 October 2026).

The forward sweep, with subgradients, is the algorithm that the rest of this series refers to as "the McCormick relaxation of a factorable function". Its steps 3 to 6 are the rules of Theorems 2.4.12, 2.4.13 and 2.4.14.

Algorithm 2.4.15 (McCormick relaxation of a factorable function at a point, with subgradients).

Algorithm 2.4.15  McCormick relaxation of a factorable function at a
                  point, with subgradients

Input   tape v_1..v_N over variables x_1..x_n (Definition 2.4.10);
        box B = [l, u]; point xbar in B;
        for every library function T, routines giving on any interval
        [a, b] a convex underestimator e_T, a concave overestimator E_T,
        their minimizer z_min and maximizer z_max, one-sided
        derivatives, and the image interval T([a, b]).
Output  for every node k: lo_k <= v_k(x) <= hi_k on B;
        cv_k <= v_k(xbar) <= cc_k; a subgradient s_k of the convex
        relaxation and a supergradient t_k of the concave relaxation
        at xbar.

1  for i = 1..n:
      (lo_i, hi_i, cv_i,   cc_i,   s_i, t_i)
    = (l_i,  u_i,  xbar_i, xbar_i, e_i, e_i)           # e_i the unit vector

2  for k = n+1..N in topological order, with operands i, j:

3     if v_k = v_i + v_j:
         lo = lo_i + lo_j;  hi = hi_i + hi_j
         cv = cv_i + cv_j;  cc = cc_i + cc_j
         s = s_i + s_j;  t = t_i + t_j

4     if v_k = c * v_i:
         if c >= 0: scale all six fields by c
         else:      (lo, hi) = (c hi_i, c lo_i)
                    (cv, cc) = (c cc_i, c cv_i)
                    (s, t)   = (c t_i, c s_i)

5     if v_k = v_i * v_j:                                   # Theorem 2.4.13
         [lo, hi] = [lo_i, hi_i] * [lo_j, hi_j]           # interval product
         low(c, m) = (c cv_m, c s_m) if c >= 0 else (c cc_m, c t_m)
                                               # min{c cv_m, c cc_m}, convex
         upp(c, m) = (c cc_m, c t_m) if c >= 0 else (c cv_m, c s_m)
                                              # max{c cv_m, c cc_m}, concave
         (cv, s) = argmax over { low(lo_j, i) + low(lo_i, j) - lo_i lo_j ,
                                 low(hi_j, i) + low(hi_i, j) - hi_i hi_j }
         (cc, t) = argmin over { upp(hi_j, i) + upp(lo_i, j) - lo_i hi_j ,
                                 upp(lo_j, i) + upp(hi_i, j) - hi_i lo_j }

6     if v_k = T(v_i):                                      # Theorem 2.4.12
         [lo, hi] = T([lo_i, hi_i])
         build e_T, E_T, z_min, z_max on [lo_i, hi_i]
         z = mid{cv_i, cc_i, z_min};  cv = e_T(z)
         s = e_T'(z) * (s_i if z = cv_i; t_i if z = cc_i; 0 otherwise)
         w = mid{cv_i, cc_i, z_max};  cc = E_T(w)
         t = E_T'(w) * (s_i if w = cv_i; t_i if w = cc_i; 0 otherwise)

7     store (lo_k, hi_k, cv_k, cc_k, s_k, t_k)

8  return node N;
      a(x) = cv_N + s_N . (x - xbar)
   is a valid underestimator of v_N on B, and
      LB = cv_N + sum_i min{ s_N,i (l_i - xbar_i), s_N,i (u_i - xbar_i) }
   is a lower bound on min_B v_N                                     (2.4.3)

Invariant (after step k, for every x in B)
    lo_k <= v_k(x) <= hi_k; as functions of xbar, cv_k is convex and
    cc_k is concave with cv_k(x) <= v_k(x) <= cc_k(x); s_k is in the
    subdifferential of cv_k at xbar and t_k in the superdifferential of
    cc_k. Each step preserves it by Theorems 2.4.12, 2.4.13 and 2.4.14.

Cost per step
    O(1) arithmetic for the four scalar fields plus O(n) for the two
    gradient vectors; O(N n) per (box, point), or O(N) when the
    gradients of the single output are obtained by a reverse sweep.

What parallelizes
    everything across boxes and points, with an identical instruction
    stream per thread; within one tape only the dependency order of the
    DAG is sequential; the selections in steps 5 and 6 are branch-free
    selects. For a rigorous bound, round lo and cv down and hi and cc up
    (directed rounding, Section 2.6) in the arithmetic steps, and give
    each library function its own error bound or a correctly rounded
    implementation; or widen every result by an eps.
A run of Algorithm 2.4.15: f = x e^y, unit square, xbar = (0.5, 0.5)

  step 1  v_1 = x   [0, 1]   cv = cc = 0.5   s = t = (1, 0)
          v_2 = y   [0, 1]   cv = cc = 0.5   s = t = (0, 1)
             |         |
             |         v
  step 6     |     v_3 = exp(v_2)   [1, e]
             |     cv = e^y (exp itself)
             |     cc = 1 + (e - 1) y (the secant)
             |         |
             v         v
  step 5  v_4 = v_1 v_3, the product rule on [0, 1] x [1, e]:
          [0, e]   cv = max{x, e x + e^y - e}
                   cc = min{e x, x + (e - 1) y}
          at xbar: cv = 0.5000 <= f = 0.8244 <= cc = 1.3591,
                   s = (1, 0), the gradient of the piece x
             |
             v
  step 8  the bound (2.4.3) on the box:
          LB = 0.5 + min{1 (0 - 0.5), 1 (1 - 0.5)}
                   + min{0 (0 - 0.5), 0 (1 - 0.5)} = 0.0000

The following script is a complete numpy implementation of Algorithm 2.4.15 with operator overloading. It has a product rule that carries the (sub)gradient of the selected piece, the composition rule with the median selection, and the library functions \(\exp\) and \(x^2\). It relaxes \(f(x, y) = x e^y\) on the unit square at three points and computes the affine bound (2.4.3) at each. It then compares the McCormick underestimator with the true convex envelope of \(f\), derived in the generating-set part below, over a grid, and ends with two cases that show what the rules depend on.

# McCormick relaxations with subgradients through a factorable DAG.
#
# Algorithm 2.4.15 by operator overloading (numpy only): every Node
# carries an interval, the values of its convex and concave relaxations
# at one point, a subgradient and a supergradient. The program relaxes
# f(x, y) = x exp(y) on the unit square at three points, with the affine
# bound (2.4.3) of each, compares the McCormick underestimator with the
# true convex envelope of f on a grid, and ends with two cases that show
# what the rules depend on.

import numpy as np

def mid(a, b, c):
    return sorted((a, b, c))[1]

class Node:
    """One DAG node at one point xbar of one box.

    The interval [lo, hi], the values cv <= v(xbar) <= cc, a subgradient
    g of the convex relaxation and a supergradient h of the concave one
    at xbar.
    """

    def __init__(s, lo, hi, cv, cc, g, h):
        s.lo, s.hi, s.cv, s.cc = lo, hi, cv, cc
        s.g, s.h = np.asarray(g, float), np.asarray(h, float)

    @staticmethod
    def var(i, n, x, l, u):
        e = np.zeros(n)
        e[i] = 1.0
        return Node(l, u, x, x, e, e)

    def __neg__(s):
        return Node(-s.hi, -s.lo, -s.cc, -s.cv, -s.h, -s.g)

    def __add__(s, o):
        if not isinstance(o, Node):
            return Node(s.lo + o, s.hi + o, s.cv + o, s.cc + o, s.g, s.h)
        return Node(s.lo + o.lo, s.hi + o.hi, s.cv + o.cv, s.cc + o.cc,
                    s.g + o.g, s.h + o.h)

    __radd__ = __add__

    def __sub__(s, o):
        return s + (-o)

    def __mul__(s, o):
        """The product rule, each selected piece carrying its
        (sub)gradient."""
        if not isinstance(o, Node):
            if o >= 0:
                return Node(o*s.lo, o*s.hi, o*s.cv, o*s.cc, o*s.g, o*s.h)
            return Node(o*s.hi, o*s.lo, o*s.cc, o*s.cv, o*s.h, o*s.g)

        # min{c cv, c cc}: convex
        low = lambda c, w: (c*w.cv, c*w.g) if c >= 0 else (c*w.cc, c*w.h)
        # max{c cv, c cc}: concave
        upp = lambda c, w: (c*w.cc, c*w.h) if c >= 0 else (c*w.cv, c*w.g)

        # the two lower planes; the larger one, with its subgradient
        a1, ga1 = low(o.lo, s)
        a2, ga2 = low(s.lo, o)
        b1, gb1 = low(o.hi, s)
        b2, gb2 = low(s.hi, o)
        cv, g = max([(a1 + a2 - s.lo*o.lo, ga1 + ga2),
                     (b1 + b2 - s.hi*o.hi, gb1 + gb2)],
                    key=lambda t: t[0])

        # the two upper planes; the smaller one, with its supergradient
        c1, hc1 = upp(o.hi, s)
        c2, hc2 = upp(s.lo, o)
        d1, hd1 = upp(o.lo, s)
        d2, hd2 = upp(s.hi, o)
        cc, h = min([(c1 + c2 - s.lo*o.hi, hc1 + hc2),
                     (d1 + d2 - s.hi*o.lo, hd1 + hd2)],
                    key=lambda t: t[0])

        P = [s.lo*o.lo, s.lo*o.hi, s.hi*o.lo, s.hi*o.hi]
        return Node(min(P), max(P), cv, cc, g, h)

    __rmul__ = __mul__

def compose(u, e, de, zmin, E, dE, zmax, lo, hi):
    """The composition rule: T(u) with e <= T <= E on [u.lo, u.hi]."""
    z = mid(u.cv, u.cc, zmin)
    g = u.g if z == u.cv else (u.h if z == u.cc else 0*u.g)
    w = mid(u.cv, u.cc, zmax)
    h = u.g if w == u.cv else (u.h if w == u.cc else 0*u.g)
    return Node(lo, hi, e(z), E(w), de(z)*g, dE(w)*h)

def secant(f, a, b):
    m = (f(b) - f(a))/(b - a) if b > a else 0.0
    return (lambda z: f(a) + m*(z - a)), (lambda z: m)

def exp(u):
    """Convex and increasing: itself below, the secant above."""
    S, dS = secant(np.exp, u.lo, u.hi)
    return compose(u, np.exp, np.exp, u.lo, S, dS, u.hi,
                   np.exp(u.lo), np.exp(u.hi))

def sqr(u):
    """Convex: itself below, minimized at mid(lo, hi, 0); the secant
    above."""
    S, dS = secant(np.square, u.lo, u.hi)
    zmax = u.hi if u.lo + u.hi >= 0 else u.lo
    lo = 0.0 if u.lo <= 0 <= u.hi else min(u.lo**2, u.hi**2)
    return compose(u, np.square, lambda z: 2*z, mid(u.lo, u.hi, 0.0),
                   S, dS, zmax, lo, max(u.lo**2, u.hi**2))

# x exp(y) at three points of the unit square, and the affine bound
# (2.4.3) that each linearization gives on the box
box = np.array([[0.0, 1.0], [0.0, 1.0]])
for xb in [(0.5, 0.5), (0.75, 0.75), (0.3, 0.9)]:
    x, y = (Node.var(i, 2, xb[i], *box[i]) for i in range(2))
    r = x*exp(y)
    LB = r.cv + sum(min(r.g[i]*(box[i, 0] - xb[i]),
                        r.g[i]*(box[i, 1] - xb[i]))
                    for i in range(2))
    print(f"x*exp(y) at {xb}: f={xb[0]*np.exp(xb[1]):.4f}  "
          f"cv={r.cv:.4f}  cc={r.cc:.4f}  g={np.round(r.g, 4)}")
    print(f"    range [{r.lo:.3f}, {r.hi:.3f}]  "
          f"affine bound on the box {LB:.4f}")

# the McCormick underestimator in closed form against the true convex
# envelope of x exp(y), on a 201 x 201 grid
g = np.linspace(0, 1, 201)
X, Y = np.meshgrid(g, g, indexing='ij')
mc_cv = np.maximum(X, np.e*X + np.exp(Y) - np.e)
t = np.clip((X + Y - 1)/np.where(X > 0, X, 1.0), 0, 1)
vex = np.where(X + Y <= 1, X, X*np.exp(t))
D = vex - mc_cv
i = np.unravel_index(np.argmax(D), D.shape)
print("envelope minus McCormick underestimator:")
print(f"    largest {D[i]:.5f} at ({X[i]:.3f}, {Y[i]:.3f}); "
      f"smallest {D.min():.1e}")

# the same function written two ways: x*x by the product rule, x^2 as
# a library function
x = Node.var(0, 1, 0.0, -1.0, 1.0)
print(f"x*x at 0 on [-1,1]: cv={(x*x).cv:.3f};  "
      f"x^2 as a library function: cv={sqr(x).cv:.3f}")

# the dependency problem: the interval of the inner node y^2 - y
y = Node.var(0, 1, 0.3, 0.0, 2.0)
r = exp(sqr(y) - y)
print(f"exp(y^2 - y) at 0.3 on [0,2]: f={np.exp(0.09 - 0.3):.4f}  "
      f"cv={r.cv:.4f}  cc={r.cc:.4f}")
print(f"    (inner interval [{(sqr(y) - y).lo:.1f}, "
      f"{(sqr(y) - y).hi:.1f}] against the true range [-0.25, 2])")
x*exp(y) at (0.5, 0.5): f=0.8244  cv=0.5000  cc=1.3591  g=[1. 0.]
    range [0.000, 2.718]  affine bound on the box 0.0000
x*exp(y) at (0.75, 0.75): f=1.5878  cv=1.4374  cc=2.0387  g=[2.7183 2.117 ]
    range [0.000, 2.718]  affine bound on the box -2.1890
x*exp(y) at (0.3, 0.9): f=0.7379  cv=0.5568  cc=0.8155  g=[2.7183 2.4596]
    range [0.000, 2.718]  affine bound on the box -2.4723
envelope minus McCormick underestimator:
    largest 0.13669 at (0.565, 0.560); smallest -8.9e-16
x*x at 0 on [-1,1]: cv=-1.000;  x^2 as a library function: cv=0.000
exp(y^2 - y) at 0.3 on [0,2]: f=0.8106  cv=0.8106  cc=21.0127
    (inner interval [-2.0, 4.0] against the true range [-0.25, 2])

(Four things visible in the output) Each evaluation is one pass over the four-node tape, a few dozen floating-point operations, and the evaluations at different points or boxes are independent. Four things are visible in the output. First, at \((0.5, 0.5)\) the subgradient \((1, 0)\) gives the affine bound \(0\) over the box, which is the true minimum of \(f\). At \((0.75, 0.75)\) the steeper subgradient gives \(-2.19\), far below it. One linearization point is a weak certificate. Several points or a small LP are needed, and the choice of the point is a design variable, with the midpoint the usual default. Second, the McCormick underestimator is not the envelope: over the unit square the true convex envelope of \(x e^y\) exceeds it by up to \(0.137\) at \((0.565, 0.560)\), and is never below it. Third, the same function written two ways gives two relaxations. The product \(x \cdot x\) relaxed by the product rule on \([-1, 1]\) is \(\max\{-2x - 1, 2x - 1\} = 2|x| - 1\), whose minimum is \(-1\), while \(x^2\) as a library function is its own convex relaxation with minimum \(0\). In general the product rule applied to \(x \cdot x\) on \([l, u]\) gives below the two tangents of \(x^2\) at the ends of the interval, \(2lx - l^2\) and \(2ux - u^2\), and above the secant \((l + u)x - lu\), twice over. These are the three products of bound factors for a squared variable, which Section 4.7 writes as rows for the lifted entry \(X_{ii}\). Fourth, the dependency problem. The inner node \(y^2 - y\) on \([0, 2]\) gets the interval \([0, 4] - [0, 2] = [-2, 4]\) instead of its true range \([-0.25, 2]\). The secant of \(\exp\) built on that interval gives the concave bound \(21.0\) at \(y = 0.3\), where the function's largest value on the box is \(e^2 \approx 7.39\). McCormick relaxations are valid but not unique. They depend on the DAG, on the association of sums and products, and on the intervals carried along. This is why rewriting before relaxing is part of every solver and why the multilinear and generating-set results below matter. Section 2.6 draws both effects on the expression graph of \(x e^y - (x + y)^2\).

One function written two ways, and the dependency problem

  x*x on [-1, 1], the product rule     x^2 on [-1, 1], the library
                                       function sqr
     x --+
         +--> mul                         x ---> sqr
     x --+    cv = max{-2x - 1, 2x - 1}          cv = x^2
                 = 2|x| - 1                      at 0: 0.000
              at 0: -1.000

  exp(y^2 - y) on [0, 2], at y = 0.3

     y --> sqr --> y^2 in [0, 4] ---+
     |                              +--> y^2 - y in
     +---------->  y   in [0, 2] ---+    [0, 4] - [0, 2] = [-2, 4]
                                         (true range [-0.25, 2])
                                                  |
                                                  v
                     exp, its secant built on [-2, 4] above:
                     cc = 21.0127 at y = 0.3, where f = 0.8106;
                     the largest value of f on the box is e^2,
                     about 7.39

(The same sweep, laid out for a GPU) The same sweep in C++23, written as a straight-line tape over a batch of boxes and points, is the form that matters for a GPU. Every node of every batch element is a record of four numbers. The storage is node-major, so that the loads of one instruction across the batch are contiguous, and the inner loop over the batch index is the loop a kernel maps to threads. The program reproduces the Python values without gradients.

// Batched McCormick relaxations over a straight-line tape.
//
// The tape is a factorable DAG in evaluation order. Every node carries
// an interval [lo, hi] and the values cv <= v <= cc of its relaxations
// at one point of one box; the batch index b (one box and point per b)
// is the parallel dimension.

#include <algorithm>
#include <cmath>
#include <cstdio>
#include <span>
#include <vector>

struct Node {
    double lo, hi, cv, cc;
};

enum class Op { Var, Add, Mul, Exp, Sqr };

// a, b: operand indices (for Var, a is the variable index)
struct Instr {
    Op op;
    int a, b;
};

inline double mid(double a, double b, double c) {
    return std::max(std::min(a, b), std::min(std::max(a, b), c));
}

inline Node add(Node u, Node v) {
    return {u.lo + v.lo, u.hi + v.hi, u.cv + v.cv, u.cc + v.cc};
}

// McCormick's product rule
inline Node mul(Node u, Node v) {
    // min{c cv, c cc}
    auto lowc = [](double c, Node w) {
        return c >= 0 ? c * w.cv : c * w.cc;
    };
    // max{c cv, c cc}
    auto uppc = [](double c, Node w) {
        return c >= 0 ? c * w.cc : c * w.cv;
    };
    double cv = std::max(lowc(v.lo, u) + lowc(u.lo, v) - u.lo * v.lo,
                         lowc(v.hi, u) + lowc(u.hi, v) - u.hi * v.hi);
    double cc = std::min(uppc(v.hi, u) + uppc(u.lo, v) - u.lo * v.hi,
                         uppc(v.lo, u) + uppc(u.hi, v) - u.hi * v.lo);
    double p[4] = {u.lo * v.lo, u.lo * v.hi, u.hi * v.lo, u.hi * v.hi};
    return {*std::min_element(p, p + 4), *std::max_element(p, p + 4),
            cv, cc};
}

// exp is convex and increasing: itself below, the secant above
inline Node expo(Node u) {
    double el = std::exp(u.lo);
    double eh = std::exp(u.hi);
    double sec = u.hi > u.lo
                     ? el + (eh - el) * (u.cc - u.lo) / (u.hi - u.lo)
                     : el;
    return {el, eh, std::exp(u.cv), sec};
}

// x^2 is convex: evaluate at mid{cv, cc, argmin}, secant above
inline Node sqr(Node u) {
    double z = mid(u.cv, u.cc, mid(u.lo, u.hi, 0.0));
    double lo = (u.lo <= 0 && 0 <= u.hi)
                    ? 0.0
                    : std::min(u.lo * u.lo, u.hi * u.hi);
    double slope = u.lo + u.hi;
    double zc = slope >= 0 ? u.cc : u.cv;
    return {lo, std::max(u.lo * u.lo, u.hi * u.hi), z * z,
            u.lo * u.lo + slope * (zc - u.lo)};
}

// Node-major storage out[k * batch + b]: the inner loop over b is the
// one a GPU maps to threads.
void evaluate(std::span<const Instr> tape, std::span<const double> lo,
              std::span<const double> hi, std::span<const double> x,
              int nvar, int batch, std::vector<Node>& out) {
    out.resize(tape.size() * batch);
    for (std::size_t k = 0; k < tape.size(); ++k)
        for (int b = 0; b < batch; ++b) {
            const Instr& I = tape[k];
            Node r{};
            switch (I.op) {
                case Op::Var: {
                    double v = x[b * nvar + I.a];
                    r = {lo[b * nvar + I.a], hi[b * nvar + I.a], v, v};
                    break;
                }
                case Op::Add:
                    r = add(out[I.a * batch + b], out[I.b * batch + b]);
                    break;
                case Op::Mul:
                    r = mul(out[I.a * batch + b], out[I.b * batch + b]);
                    break;
                case Op::Exp:
                    r = expo(out[I.a * batch + b]);
                    break;
                case Op::Sqr:
                    r = sqr(out[I.a * batch + b]);
                    break;
            }
            out[k * batch + b] = r;
        }
}

int main() {
    // f(x, y) = x * exp(y) on the unit square at three points, then
    // x*x against x^2 on [-1, 1] at 0.

    // the tape of x * exp(y): node 0 = x, node 1 = y,
    // node 2 = exp(node 1), node 3 = node 0 * node 2
    std::vector<Instr> tape = {{Op::Var, 0, -1},
                               {Op::Var, 1, -1},
                               {Op::Exp, 1, -1},
                               {Op::Mul, 0, 2}};
    std::vector<double> lo(6, 0.0), hi(6, 1.0);
    std::vector<double> x = {0.5, 0.5, 0.75, 0.75, 0.3, 0.9};
    std::vector<Node> out;
    evaluate(tape, lo, hi, x, 2, 3, out);
    for (int b = 0; b < 3; ++b) {
        // node 3 of batch element b
        Node r = out[3 * 3 + b];
        std::printf("x*exp(y) at (%.2f, %.2f): "
                    "f=%.4f  cv=%.4f  cc=%.4f\n",
                    x[2 * b], x[2 * b + 1],
                    x[2 * b] * std::exp(x[2 * b + 1]), r.cv, r.cc);
        std::printf("    range [%.3f, %.3f]\n", r.lo, r.hi);
    }

    // node 0 = x, node 1 = x * x by the product rule, node 2 = x^2
    // by sqr; one batch element
    std::vector<Instr> t2 = {{Op::Var, 0, -1},
                             {Op::Mul, 0, 0},
                             {Op::Sqr, 0, -1}};
    std::vector<double> lo2 = {-1.0}, hi2 = {1.0}, x2 = {0.0};
    evaluate(t2, lo2, hi2, x2, 1, 1, out);
    std::printf("x*x at 0 on [-1,1]: cv=%.3f cc=%.3f;\n",
                out[1].cv, out[1].cc);
    std::printf("    x^2 as a library function: cv=%.3f cc=%.3f\n",
                out[2].cv, out[2].cc);
}
x*exp(y) at (0.50, 0.50): f=0.8244  cv=0.5000  cc=1.3591
    range [0.000, 2.718]
x*exp(y) at (0.75, 0.75): f=1.5878  cv=1.4374  cc=2.0387
    range [0.000, 2.718]
x*exp(y) at (0.30, 0.90): f=0.7379  cv=0.5568  cc=0.8155
    range [0.000, 2.718]
x*x at 0 on [-1,1]: cv=-1.000 cc=1.000;
    x^2 as a library function: cv=0.000 cc=1.000

(What the batched tape costs, and what rigour needs) The cost is \(O(N)\) per batch element with no allocation and no data-dependent branching beyond the selects inside min, max and mid. The loop nest, with the instruction index outside and the batch index inside, is the one a GPU kernel has when each thread takes one value of b. Section 7.1 explains the execution model. The node-major layout makes the loads of out[I.a * batch + b] contiguous across the group of threads that executes one instruction together, which is the access pattern the memory system serves fastest. Rigour needs one more ingredient. Rounding lo and cv down and hi and cc up, called directed rounding and defined in Section 2.6, makes the sums and products rigorous. On the host this is std::fesetround, and on the device the rounding-mode intrinsics such as __dadd_rd and __dadd_ru. Library functions such as exp are not correctly rounded and do not honour the rounding mode, so each needs its own error bound or a correctly rounded implementation before the bound is a proof. Without these steps a linearization computed in rounded arithmetic is not a proof.

The node-major layout of the C++ listing: out[k * batch + b]

  the tape of x * exp(y)          the batch, batch = 3
  k = 0  Var x                    b = 0  (0.50, 0.50)
  k = 1  Var y                    b = 1  (0.75, 0.75)
  k = 2  Exp of node 1            b = 2  (0.30, 0.90)
  k = 3  Mul of nodes 0 and 2

             k = 0       k = 1       k = 2       k = 3
         +---+---+---+---+---+---+---+---+---+---+---+---+
  out    |b=0|b=1|b=2|b=0|b=1|b=2|b=0|b=1|b=2|b=0|b=1|b=2|
         +---+---+---+---+---+---+---+---+---+---+---+---+
  index    0   1   2   3   4   5   6   7   8   9  10  11
          \_________/             \_________/ \_________/
          operand a               operand b   result

  for k = 3 the threads b = 0, 1, 2 read out[0..2] and out[6..8]
  and write out[9..11]: each access is contiguous across b

The catalogue of terms

Each library function needs, on an interval \([l, u]\), a convex underestimator and a concave overestimator. Each structured term needs the same on a box. The table collects the constructions this series uses, with the cost of the relaxation and whether it is the envelope of the term. A function that is concave along each coordinate direction is called edge-concave. The multilinear row refers to this class, and Theorem 2.4.16 is about it. The last rows are forward references to Section 4, where lifted and conic relaxations are built. They are listed here so that the whole ladder of relaxations is in one place. Three abbreviations appear in them. A second-order cone program (SOCP) is a convex program over the cone met in Section 1.2 and treated in Section 4.8. A semidefinite program (SDP) is a linear program over the cone of positive semidefinite matrices (Section 4.7). Special ordered sets of type 2 (SOS2) are the constraint that at most two consecutive weights of a piecewise-linear model are nonzero (Section 4.6).

termrelaxation on the interval or boxcost of the relaxationthe envelope?
affineitselffreeyes
convex (\(x^2\), \(\exp\), \(\vert x \vert^p\) with \(p \ge 1\))itself, or its tangents as cuts; concave side: the secant from \((l, f(l))\) to \((u, f(u))\)convex, or LP by tangentsyes
concave (\(\log\), \(\sqrt{\cdot}\), \(x^p\) with \(0 < p < 1\))convex side: the secant; concave side: itselfas aboveyes
\(1/x\), \(0 < l\)itself below, secant aboveas aboveyes
\(1/x\), \(u < 0\)secant below, itself aboveas aboveyes
\(1/x\), \(l < 0 < u\)unbounded: the interval must exclude \(0\) firstbranch or tighten–
\(x^p\), \(p\) odd, \(l < 0 < u\)line from \((l, l^p)\) tangent at \(c = r \vert l \vert\) (or the secant if \(c > u\)), then \(x^p\); mirror image above; \(r\) solves \((p-1) r^p + p r^{p-1} = 1\): \(r = 0.5\), \(0.6058\), \(0.6703\) for \(p = 3\), \(5\), \(7\)LP by tangentsyes
\(\sin\), \(\cos\)lower convex hull of the arc: convex pieces of the curve joined by tangent or secant segments, tangency points by Newton's methodLP by tangentsyes
bilinear \(w = xy\)the four planes (2.4.2)LP, 4 rowsyes: exact on the four edges
fractional \(x / y\) (\(x \ge 0\), \(y > 0\))one-dimensional convex minimization over the faces \(x = l_x\), \(x = u_x\); two planes aboveSOCP (convex side), LP (concave side)yes
trilinear \(xyz\) and multilinearlower hull of the 8 corner values: an LP with \(2^n\) columns; pairwise McCormick as a fallbackLP, \(2^n\) columnsyes (any multilinear or edge-concave term)
general \(f\), twice differentiable\(f(x) + \sum_i \alpha_i (x_i - l_i)(x_i - u_i)\), \(\alpha\) from an interval Hessian (αBB)convex NLPno: exact at the \(2^n\) vertices only
integer \(y \in \{0, 1\}\)\(y \in [0, 1]\): the chord of \(y - y^2 \le 0\)freeyes (Section 1.4)
piecewise McCormick, \(K\) cells\(K\) copies of the four planes with selectors; band \(A/(2 k_x k_y)\) per cell, \(A\) the box areaMILP, \(K\) binaries (\(\log_2 K\) with encoding)LP relaxation = hull of the union of cells
hull of a disjunctionBalas' lifted formulationLP in a lifted spaceyes (Section 4.2)
perspective \(z f(x/z)\)the perspective of a convex \(f\) with an indicatorSOCP for quadratic \(f\)yes (Section 4.3)
piecewise linear, λ modelconvex combination of breakpoints, SOS2MILP; LP relaxationLP relaxation = hull of the breakpoints (Section 4.6)
RLT, level 1all pairwise products of bound and constraint factors, linearizedLP, \(O(n^2)\) variablesno in general; equals McCormick for one term
Shor SDP\(\big(\begin{smallmatrix} 1 & x^\top \\ x & X \end{smallmatrix}\big)\) positive semidefiniteSDP, \((n+1) \times (n+1)\)exact for one quadratic constraint (Section 4.7)
Lasserre, order \(k\)moment matrix of order \(k\)SDP, \(\binom{n+k}{k}\) rowsconverges in \(k\) (Section 4.7)
copositiveBurer's reformulationexact, intractableyes, in principle (Section 4.7)
The catalogue: what replaces each kind of term, what it costs, and when it is the envelope.

The univariate rows are the lower convex hull of Proposition 2.4.3 computed in closed form. Two computed cases make the odd-power and trigonometric rows concrete. For \(x^3\) on \([-1, 2]\) the envelope leaves the left endpoint \((-1, -1)\) along the tangent line that touches the curve at \(c = -l/2 = 0.5\), with slope \(3c^2 = 0.75\), and follows \(x^3\) from there. So \(\operatorname{vex} x^3 = -1 + 0.75(x + 1)\) for \(x \le 0.5\), which is \(-0.25\) at \(x = 0\) against \(f(0) = 0\). The general constant \(r\) in the table, with \(c = r|l|\), is Liberti and Pantelides' result for monomials of odd degree.L. Liberti and C. C. Pantelides, "Convex envelopes of monomials of odd degree", Journal of Global Optimization 25 (2003). The constants \(r = 0.5, 0.605830, 0.670332\) for \(p = 3, 5, 7\) were recomputed for this post as the root in \((0, 1)\) of \((p-1) r^p + p r^{p-1} - 1 = 0\). For \(\sin x\) on \([0, 7]\) the envelope is the line from \((0, 0)\) tangent to the curve at \(c_1 = 4.49341\), where \(\tan c_1 = c_1\), then \(\sin x\) itself, then the line tangent at \(c_2 = 5.92710\) through \((7, \sin 7)\). At \(\pi/2\) the envelope is \(-0.3412\) against \(\sin(\pi/2) = 1\). SCIP's trigonometric handler builds exactly this hull, locating the tangency points by Newton's method and verifying them.SCIP Optimization Suite, source file src/scip/expr_trig.c (author F. Wegscheider), github.com/scipopt/scip (accessed 4 October 2026).

The convex envelope of x³ on [−1, 2] (the odd-power row of the catalogue, p = 3): the tangent line from (−1, −1) to the curve at c = r|l| = 0.5, then x³ itself. At x = 0 the envelope is −0.25, against f(0) = 0.
The convex envelope of sin x on [0, 7] (the trigonometric row of the catalogue): the line from (0, 0) tangent at c₁ = 4.49341, where tan c₁ = c₁, then sin x itself, then the tangent at c₂ = 5.92710 through (7, sin 7). At π/2 the envelope is −0.3412, against sin(π/2) = 1.

The fractional row is the first multivariate term in the table whose envelope is not polyhedral. For \(x/y\) on \([x^L, x^U] \times [y^L, y^U]\) with \(0 \le x^L\) and \(0 < y^L\), Tawarmalani and Sahinidis derived the convex envelope as a one-dimensional convex minimization over the two faces \(x = x^L\) and \(x = x^U\). In the case \(x^L = 0\) it closes to

\[\operatorname{vex}_R \frac{x}{y}\,(x, y) \;=\; \max\Big\{ \frac{x}{y^U},\ \frac{x^2}{x^U y - (x^U - x)\, y^L} \Big\},\]

a function whose second piece, \(t \ge s^2/y\) after a change of variables, is a rotated second-order cone. The concave envelope is the minimum of two planes through the four corner values. Relaxing \(x \cdot (1/y)\) by the product rule instead, with \(1/y\) relaxed by itself, is valid and weaker. On \([0.5, 4] \times [1, 3]\) the product rule falls short of the envelope by up to \(0.333\), at \((2.25, 1.5)\), where the function is \(1.5\).M. Tawarmalani and N. V. Sahinidis, "Semidefinite relaxations of fractional programs via novel convexification techniques", Journal of Global Optimization 20 (2001). The closed form above and the gap of \(0.333\) were recomputed for this post by a fine one-dimensional search over the faces and checked against a numerical envelope; the general-case formula was derived here from the generating-set argument of Theorem 2.4.18 and verified numerically, not compared line by line with the paper's printed form. This is the envelope derived for BARON's fractional terms; its use is described in the book cited below (Tawarmalani and Sahinidis, 2002) and in Tawarmalani and Sahinidis (2005), cited above.

Multilinear terms, generating sets and the hull of a sum

Theorem 2.4.7 said that the envelopes of a bilinear term are determined by its values at the four corners. The reason is that \(xy\) is affine in each variable separately, and the mechanism generalizes to every edge-concave function.

Theorem 2.4.16 (vertex polyhedral envelopes: Rikun, 1997; Tardella, 2008; Meyer and Floudas, 2005). Let \(f\) be continuous on the box \(B\) with vertex set \(V(B)\), \(|V(B)| = 2^n\), and suppose \(f\) is edge-concave. That is, for every \(i\) and every fixed value of the other coordinates, \(x_i \mapsto f(x)\) is concave on \([l_i, u_i]\). Then

\[\operatorname{vex}_B f\,(x) \;=\; \min\Big\{ \sum_{v \in V(B)} \theta_v f(v) \;:\; \theta \ge 0,\ \sum_v \theta_v = 1,\ \sum_v \theta_v v = x \Big\},\]

a polyhedral function whose graph is the lower convex hull of the \(2^n\) points \((v, f(v))\). If \(f\) is multilinear, affine in each coordinate separately, both \(\operatorname{vex}_B f\) and \(\operatorname{cav}_B f\) are of this form.

Proof. Let \(\psi(x)\) denote the right-hand side. Since the vertices lie in \(B\), (2.4.1) gives \(\operatorname{vex}_B f \le \psi\). For the reverse inequality take \(x \in B\). Write it as a convex combination \(x = \tau x' + (1 - \tau) x''\) of the two points obtained by moving along coordinate \(1\) to the faces \(x_1 = l_1\) and \(x_1 = u_1\), with the other coordinates unchanged. Concavity in \(x_1\) gives \(f(x) \ge \tau f(x') + (1 - \tau) f(x'')\). Repeat on each face with coordinate \(2\), and so on. After \(n\) steps, \(f(x) \ge \sum_v \theta_v f(v)\) for the product weights \(\theta_v\), which satisfy \(\sum_v \theta_v v = x\), so \(f(x) \ge \psi(x)\). The function \(\psi\) is convex, being the lower hull of finitely many points (the value function of a parametric linear program in \(x\)), so it is a convex minorant of \(f\) and \(\psi \le \operatorname{vex}_B f\). For multilinear \(f\) apply the argument to \(f\) and to \(-f\). ∎

(Why the corners determine the floor) If a function bends downward along every coordinate direction, every interior point of its graph lies above the tent spanned by the corners, so the corners determine the floor.A. D. Rikun, "A convex envelope formula for multilinear functions", Journal of Global Optimization 10 (1997); F. Tardella, "Existence and sum decomposition of vertex polyhedral convex envelopes", Optimization Letters 2 (2008), which characterizes the functions whose envelope over a polytope is vertex polyhedral; C. A. Meyer and C. A. Floudas, "Convex envelopes for edge-concave functions", Mathematical Programming 103 (2005), which constructs the facets from triangulations of the vertex set. On the unit hypercube the hull of a single monomial is the standard linearization \(w \ge 0\), \(w \ge \sum_i x_i - n + 1\), \(w \le x_i\): Y. Crama, "Concave extensions for nonlinear 0–1 maximization problems", Mathematical Programming 61 (1993). The value at a point is a linear program with \(2^n\) variables and \(n + 1\) equality constraints. For \(n = 2\) it is Theorem 2.4.7. For \(n = 3\) the facets are listed explicitly for every sign pattern of the bounds, and for \(n = 4\) partially. Beyond that, solvers fall back to applying the bilinear planes pairwise, which is what the product rule of Theorem 2.4.13 does to \((xy)z\).C. A. Meyer and C. A. Floudas, "Trilinear monomials with mixed sign domains: facets of the convex and concave envelopes", Journal of Global Optimization 29 (2004); S. Cafieri, J. Lee and L. Liberti, "On convex relaxations of quadrilinear terms", Journal of Global Optimization 47 (2010). The hypothesis matters. The envelope of a function that is not concave along the coordinates, such as \(x^2\) or \(x e^y\) in the direction of \(y\), is not vertex polyhedral. The vertex LP is then only an upper bound on the envelope, and for \(x^2\) it is the secant, which is the concave envelope.

Proposition 2.4.17 (recursive McCormick). Let \(B = [x^L, x^U] \times [y^L, y^U] \times [z^L, z^U]\). Relax the trilinear term \(xyz\) by introducing \(w = xy\) with the planes (2.4.2) on \([x^L, x^U] \times [y^L, y^U]\), bounding \(w\) by the interval product \([w^L, w^U]\), and introducing \(u = wz\) with the planes (2.4.2) on \([w^L, w^U] \times [z^L, z^U]\). The eight inequalities are valid: every \((x, y, z, xyz)\) with \((x, y, z) \in B\) satisfies them for \(w = xy\). Their projection onto \((x, y, z, u)\) is a polyhedron containing the set \(\{(x, y, z, u) : \operatorname{vex}_B(xyz) \le u \le \operatorname{cav}_B(xyz)\}\) of Theorem 2.4.16, and the containment can be strict.

Proof. Validity: for \((x, y, z) \in B\) put \(w = xy\). Then \((x, y, w)\) satisfies the first four planes by Theorem 2.4.7, \(w \in [w^L, w^U]\) by the definition of the interval product, and \((w, z, wz)\) satisfies the last four. Containment: the projection is convex, being the projection of a polyhedron, and by validity it contains the graph \(\{(x, y, z, xyz)\}\). It therefore contains the convex hull of the graph, which is the set between the two envelopes by Proposition 2.4.3 applied to \(xyz\) and to \(-xyz\). Strictness: on \([1, 2]^3\) take \((x, y, z) = (1, 2, 1.6)\), a point on an edge of the box, where the envelope of the multilinear function equals the function, \(3.2\). The planes for \(w\) on \([1, 2]^2\) fix \(w = xy = 2\) exactly, because \((1, 2)\) is a corner. But the planes for \(u = wz\) are built on \(w \in [w^L, w^U] = [1, 4]\), and at \((w, z) = (2, 1.6)\) they only give \(u \ge \max\{1 \cdot 1.6 + 2 \cdot 1 - 1,\ 4 \cdot 1.6 + 2 \cdot 2 - 8\} = 2.6\). So the projection contains \((1, 2, 1.6, 2.6)\), which lies below the envelope. ∎

Proposition 2.4.17: recursive McCormick for xyz on [1, 2]^3

     x in [1, 2]      y in [1, 2]
            \           /
         planes (2.4.2) on [1, 2] x [1, 2]
                  |
                  v
         w = xy, with its interval product        z in [1, 2]
         [w^L, w^U] = [1, 4]                          |
                  \                                   /
                   planes (2.4.2) on [1, 4] x [1, 2]
                                |
                                v
                             u = wz

  at (x, y, z) = (1, 2, 1.6), on an edge of the box, the envelope
  equals xyz = 3.2, but:
    stage 1: (1, 2) is a corner of [1, 2]^2, so w = xy = 2 exactly
    stage 2: (w, z) = (2, 1.6) is inside [1, 4] x [1, 2], and
             u >= max{1 * 1.6 + 2 * 1 - 1, 4 * 1.6 + 2 * 2 - 8}
               = 2.6
  so (1, 2, 1.6, 2.6) is in the projection, below the envelope

(How much the recursive relaxation loses) Luedtke, Namazifar and Linderoth study how much is lost. For a single multilinear monomial they prove that the recursive relaxation equals the envelopes when every variable's bounds are symmetric about zero. For bilinear functions, sums of bilinear terms, they show that the width of the termwise McCormick band is, in the words of their abstract, "within a constant of" the width of the band between the envelopes. The constant and the exact statement are in the paper. Speakman and Lee compute the volume of each of the three recursive relaxations of a trilinear monomial over a box with nonnegative bounds and identify which association order is best.J. Luedtke, M. Namazifar and J. Linderoth, "Some results on the strength of relaxations of multilinear functions", Mathematical Programming 136 (2012); E. Speakman and J. Lee, "Quantifying double McCormick", Mathematics of Operations Research 42 (2017). Computed for this post with a numpy script that compares the exact envelopes of \(xyz\) (the vertex LP over the eight corners) with the recursive relaxation on an \(11^3\) grid: on \([0, 1]^3\), on \([0, 1] \times [0, 2] \times [0, 3]\) and on \([-1, 1]^3\) the recursive relaxation coincides with the envelopes at every grid point for all three association orders; on \([1, 2]^3\) it is strictly weaker at \(775\) of the \(1331\) points, with a largest gap of \(0.6\), and the mean width of the band rises from \(0.521\) for the envelopes to \(0.721\), \(38\%\) wider; on \([1, 2] \times [0, 1]^2\) the order that multiplies the two zero-based variables first is exact and the other two are not. On \([1, 2]^3\) the recursive relaxation is strictly weaker than the exact envelope, by \(0.6\) at worst, and on a mixed box the association order decides whether it is exact. The lesson for the hull of a sum is this. The envelope of a single bilinear term is known, but the sum of the termwise envelopes is in general not the envelope of the sum. The gap between them is what Luedtke, Namazifar and Linderoth bound.

The engine behind every explicit envelope in this subsection is a statement about which points of the graph the envelope actually uses.

Theorem 2.4.18 (generating sets and convex extensions: Tawarmalani and Sahinidis, 2002). Let \(f\) be continuous on the box \(B\) and let the generating set \(G_B(f)\) be the set of \(x \in B\) for which \((x, f(x))\) is an extreme point of \(\operatorname{conv}(\operatorname{epi}_B f)\). (a) \(\operatorname{vex}_B f\) is determined by \(f|_{G_B(f)}\): it is the convex envelope of the function equal to \(f\) on \(G_B(f)\) and \(+\infty\) elsewhere, the convex extension of \(f\) restricted to its generating set. (b) If \(f\) is concave in the coordinate \(x_i\) for every fixed value of the other coordinates, then \(G_B(f) \subseteq \{x \in B : x_i \in \{l_i, u_i\}\}\). (c) Consequently, if \(f\) is concave in each of the coordinates in a set \(J\), the envelope is the convex extension of \(f\) restricted to the union of the faces on which every \(x_j\), \(j \in J\), is at a bound.

Proof sketch. (a) The extreme points of \(\operatorname{conv}(\operatorname{epi}_B f)\) lie on the graph, because a point strictly above the graph is the midpoint of two epigraph points on the same vertical line. The set \(\operatorname{conv}(\operatorname{epi}_B f)\) is closed, convex and contains no line, so it is the convex hull of its extreme points and extreme directions (Rockafellar, Theorem 18.5), and its only extreme direction is the vertical one. Hence the hull of the epigraph is generated by \(\{(x, f(x)) : x \in G_B(f)\}\) together with the vertical direction. (b) If \(l_i < x_i < u_i\), then \((x, f(x))\) lies on or above the segment between the two epigraph points at the ends of the \(i\)-th coordinate segment through \(x\), by concavity along that segment. So it is not an extreme point. (c) Apply (b) for each \(j \in J\). ∎

Wherever the function bends downward in some direction, the envelope ignores it and looks to the boundary in that direction. For \(xy\), affine in both variables, the generating set is the four corners, and Theorem 2.4.7 follows. For \(x/y\), affine in \(x\), it lies in the two faces \(x = x^L\) and \(x = x^U\), which is where the fractional envelope above came from. Tawarmalani and Sahinidis develop this calculus, the vocabulary of convex extensions, and conditions under which a function's envelope can be assembled from envelopes of simpler pieces. Tawarmalani, Richard and Xiong turn it into a technique that computes envelopes in closed form when the generating set is a finite union of polytopes. The tool is the polyhedral subdivision of the domain that the parametric LP (2.4.1) induces.M. Tawarmalani and N. V. Sahinidis, "Convex extensions and envelopes of lower semi-continuous functions", Mathematical Programming 93 (2002); M. Tawarmalani and N. V. Sahinidis, Convexification and Global Optimization in Continuous and Mixed-Integer Nonlinear Programming (Kluwer, 2002); M. Tawarmalani, J.-P. P. Richard and C. Xiong, "Explicit convex and concave envelopes through polyhedral subdivisions", Mathematical Programming 138 (2013); A. Khajavirad and N. V. Sahinidis, "Convex envelopes generated from finitely many compact convex sets", Mathematical Programming 137 (2013).

(The envelope of x e^y, by generating sets) The example the Python script computed is an instance. Its McCormick relaxations were derived by hand after Theorem 2.4.13: \(f^{\mathrm{cv}} = \max\{x, e x + e^y - e\}\) and \(f^{\mathrm{cc}} = \min\{e x, x + (e - 1)y\}\) for \(f(x, y) = x e^y\) on the unit square. The true envelopes come from Theorem 2.4.18. The function is affine in \(x\), so \(\operatorname{vex} f\) is generated on the faces \(x = 0\), where \(f = 0\), and \(x = 1\), where \(f = e^y\) is convex. A point \((x, y)\) is the combination of \((0, y_0)\) with weight \(1 - x\) and \((1, y_1)\) with weight \(x\), with \((1 - x) y_0 + x y_1 = y\). The value \(x e^{y_1}\) is smallest for the smallest admissible \(y_1 = \max\{0, (x + y - 1)/x\}\). So \(\operatorname{vex} f = x\) where \(x + y \le 1\) and \(\operatorname{vex} f = x\, e^{(x + y - 1)/x}\) beyond the anti-diagonal. The function is convex in \(y\), so \(\operatorname{cav} f\) is generated on the faces \(y = 0\) and \(y = 1\), where it is linear in \(x\). The maximization over the segment of admissible pairs is linear and ends at an endpoint, which gives \(\operatorname{cav} f = \min\{e x, x + (e - 1) y\}\), exactly the McCormick overestimator. On this example, then, the McCormick concave relaxation is the envelope and the McCormick convex relaxation is not. Where \(x + y \le 1\) both equal \(x\). Beyond the anti-diagonal the envelope is an exponential chord while McCormick keeps the plane \(e x + e^y - e\), up to \(0.137\) below it.

Theorem 2.4.18 on f = x eʸ over the unit square: vex f is generated on the faces x = 0, where f = 0, and x = 1, where f = eʸ. Below the anti-diagonal, vex f = x. Beyond it, a point (x, y) mixes (0, 1) and (1, y₁), with y₁ = (x + y − 1)/x and weights 1 − x and x, so vex f = x e^((x + y − 1)/x). The ring is (0.565, 0.560), where vex f exceeds the McCormick underestimator most, by 0.13669 (the script's output above).

In closed form, then, \(\operatorname{vex} f = x\) for \(x + y \le 1\) and \(\operatorname{vex} f = x\, e^{(x + y - 1)/x}\) otherwise, and the script's output above measures the gap to the McCormick underestimator at \(0.13669\), at \((0.565, 0.560)\). Theorem 2.4.4 is also visible: \(f\), the McCormick underestimator and \(\operatorname{vex} f\) all have minimum \(0\) on the square, attained on the edge \(x = 0\), and the three functions differ only away from the minimizers.

Two closing remarks concern sums, which is where the gap of a relaxed problem is born. First, Proposition 2.4.6 and the multilinear results say that even with the exact envelope of every function, the relaxed problem can have a gap. There are two reasons: \(\operatorname{vex}(f + g) \ne \operatorname{vex} f + \operatorname{vex} g\), and the feasible set of the relaxed constraints is larger than the convex hull of the feasible set. The convex hull of the feasible set is a different and stronger object than the collection of envelopes of the functions that define it. Simultaneous convexification means convexifying the graph of a vector of functions \((f_1, \dots, f_m)\) jointly, that is taking \(\operatorname{conv}\{(x, f_1(x), \dots, f_m(x))\}\) rather than the product of the individual hulls. The multiterm polyhedral cuts for quadratic constraints in BARON are an instance.X. Bao, N. V. Sahinidis and M. Tawarmalani, "Multiterm polyhedral relaxations for nonconvex, quadratically constrained quadratic programs", Optimization Methods and Software 24 (2009). The concept is treated in Tawarmalani and Sahinidis (2002), the book cited above, Chapter 2. Second, the lifted view of Section 4.7 says the same thing in matrix language. For a problem whose nonconvexity is quadratic, write \(X = x x^\top\). The McCormick planes for every pair are exactly the products of bound factors. The remaining gap is the gap of the lifted set \(\{(x, X) : X = x x^\top,\ x \in \mathcal F\}\) against its relaxation. Products with the constraints and the convexity cut \(X_{ii} \ge x_i^2\) are what close it on the running example, as the ladder of Section 4.7 shows.

Relaxing a whole function at once: αBB

Everything so far relaxed a function term by term. The αBB method relaxes a twice continuously differentiable function all at once. It adds a quadratic that is negative inside the box, zero at its vertices, and curved enough to make the sum convex. For \(f \in C^2(B)\), \(B = [l, u]\), and \(\alpha \in \mathbb{R}^n_{\ge 0}\), the αBB underestimator, written \(L_\alpha\) and not to be confused with the Lagrangian \(L(x, \lambda)\) of Section 2.2, is

\[L_\alpha(x) \;=\; f(x) + \sum_{i=1}^n \alpha_i\,(x_i - l_i)(x_i - u_i). \tag{2.4.4}\]

Each term \((x_i - l_i)(x_i - u_i)\) is nonpositive on \(B\), so \(L_\alpha \le f\) for \(\alpha \ge 0\). The same function is often written \(f(x) - \sum_i \alpha_i (x_i - l_i)(u_i - x_i)\), and the two forms are identical. Writing a plus sign with \((x_i - l_i)(u_i - x_i)\) produces an overestimator. An interval Hessian of \(f\) on \(B\) is a matrix of intervals \([H] = ([h_{ij}^L, h_{ij}^U])\) with \(\nabla^2 f(x) \in [H]\) for every \(x \in B\), obtained by evaluating the second-derivative expressions in interval arithmetic.

Theorem 2.4.19 (αBB: Maranas and Floudas, 1994; Androulakis, Maranas and Floudas, 1995; Adjiman, Dallwig, Floudas and Neumaier, 1998). Let \(f \in C^2(B)\) and \(L_\alpha\) as in (2.4.4) with \(\alpha \ge 0\).

(a) \(L_\alpha \le f\) on \(B\), with equality exactly where every coordinate \(i\) with \(\alpha_i > 0\) is at a bound, in particular at all \(2^n\) vertices.

(b) \(L_\alpha\) is convex on \(B\) if and only if \(\nabla^2 f(x) + 2\operatorname{diag}(\alpha) \succeq 0\) for all \(x \in B\).

(c) \(\max_{x \in B}\big(f(x) - L_\alpha(x)\big) = \sum_{i=1}^n \alpha_i (u_i - l_i)^2/4\), attained at the centre of \(B\).

(d) (Gershgorin, scaled.) Let \([H]\) be an interval Hessian of \(f\) on \(B\) and \(d > 0\) any vector. Then

\[\alpha_i \;=\; \max\Big\{ 0,\ -\tfrac12 \Big( h_{ii}^L - \sum_{j \ne i} \max\{|h_{ij}^L|, |h_{ij}^U|\}\,\frac{d_j}{d_i} \Big) \Big\}, \qquad i = 1, \dots, n, \tag{2.4.5}\]

makes \(L_\alpha\) convex on \(B\). The choice \(d = u - l\) is the scaled version, \(d = (1, \dots, 1)\) the plain one.

(e) Fix the scaling vector \(d\). If \(B\) is bisected along coordinate \(k\) and \(\alpha\) is recomputed from (2.4.5) with the child's interval Hessian and the same \(d\), then the child's \(\alpha\) is componentwise no larger than the parent's. Its maximum separation is at most the parent's with the \(k\)-th term divided by four. With the scaled choice \(d = u - l\) recomputed on the child, the ratios \(d_j/d_k\) double and \(\alpha_k\) can grow.

Proof. (a) Each term \(\alpha_i (x_i - l_i)(x_i - u_i)\) is \(\le 0\) on \(B\) and vanishes only at \(x_i \in \{l_i, u_i\}\). (b) \(\nabla^2 L_\alpha = \nabla^2 f + 2\operatorname{diag}(\alpha)\), and a \(C^2\) function on a convex set is convex if and only if its Hessian is positive semidefinite there. (c) \(f - L_\alpha = \sum_i \alpha_i (x_i - l_i)(u_i - x_i)\) is separable, each term a downward parabola with maximum \(\alpha_i (u_i - l_i)^2/4\) at the midpoint. (d) Let \(D = \operatorname{diag}(d)\). For \(x \in B\) the matrix \(M = \nabla^2 f(x) + 2\operatorname{diag}(\alpha)\) is similar to \(D^{-1} M D\), whose Gershgorin discs are centred at \(m_{ii} = h_{ii}(x) + 2\alpha_i\) with radii \(\sum_{j \ne i} |h_{ij}(x)|\, d_j/d_i\). With (2.4.5), \(m_{ii} \ge h_{ii}^L + 2\alpha_i \ge \sum_{j \ne i} \max\{|h_{ij}^L|, |h_{ij}^U|\}\, d_j/d_i\), which is at least the radius, so every disc lies in the closed right half-plane and \(M \succeq 0\). Apply (b). (e) With \(d\) fixed, the child's interval Hessian is contained in the parent's, so every term in (2.4.5) is no larger, and the \(k\)-th width is halved and enters (c) squared. With \(d = u - l\) recomputed, halving the box along \(k\) halves \(d_k\) and doubles every ratio \(d_j/d_k\) in the formula for \(\alpha_k\). For \(f = x y^2 - x^2\) on \([0, 2]^2\) the scaled \(\alpha\) is \((3, 2)\), and on the child \([0, 1] \times [0, 2]\) of a bisection in \(x\) it is \((5, 1)\), so \(\alpha_1\) grows from \(3\) to \(5\). The script below prints both. ∎

(What the perturbation does, and what it costs) The quadratic perturbation adds curvature \(2\alpha_i\) in each coordinate direction, enough to cancel the most negative curvature the Hessian can have anywhere in the box, and it costs exactly a parabola's height at the centre. Larger boxes have wider interval Hessians and need larger \(\alpha\), so under splitting with a fixed scaling the separation shrinks faster than quadratically. The simplest valid choice of a uniform \(\alpha\) is \(\alpha \ge \max\{0, -\tfrac12 \lambda_{\min}\}\) with \(\lambda_{\min}\) a lower bound on the smallest eigenvalue of the Hessian over the box. The formula (2.4.5) is the Gershgorin bound on that eigenvalue, coordinate by coordinate, and the paper of Adjiman, Dallwig, Floudas and Neumaier compares it with several other interval eigenvalue bounds.C. D. Maranas and C. A. Floudas, "Global minimum potential energy conformations of small molecules", Journal of Global Optimization 4 (1994); I. P. Androulakis, C. D. Maranas and C. A. Floudas, "αBB: a global optimization method for general constrained nonconvex problems", Journal of Global Optimization 7 (1995); C. S. Adjiman, S. Dallwig, C. A. Floudas and A. Neumaier, "A global optimization method, αBB, for general twice-differentiable constrained NLPs—I. Theoretical advances", Computers & Chemical Engineering 22 (1998). Variants: piecewise quadratic perturbations in C. A. Meyer and C. A. Floudas, "Convex underestimation of twice continuously differentiable functions by piecewise quadratic perturbation: spline αBB underestimators", Journal of Global Optimization 32 (2005), and non-diagonal perturbations in A. Skjäl, T. Westerlund, R. Misener and C. A. Floudas, "A generalization of the classical αBB convex underestimation via diagonal and nondiagonal quadratic terms", Journal of Optimization Theory and Applications 154 (2012).

(One variable, one split) A one-variable example shows the mechanism and what one split does to it. Take \(f(x) = x^3\) on \([-1, 1]\). The second derivative is \(6x \ge -6\) on the interval, so \(\alpha = 3\) and \(L(x) = x^3 + 3(x^2 - 1)\). Its derivative \(3x^2 + 6x = 3x(x + 2)\) vanishes at \(x = 0\), so \(\min_{[-1,1]} L = L(0) = -3\), against the true minimum \(f(-1) = -1\). The separation certificate of (c) is \(\alpha (u - l)^2/4 = 3\), and it is attained at the centre, where \(f(0) - L(0) = 3\). Now split at \(0\). On \([-1, 0]\) the interval Hessian is still \([-6, 0]\), so \(\alpha = 3\) and \(L(x) = x^3 + 3x(x + 1)\), whose derivative \(3x^2 + 6x + 3 = 3(x + 1)^2\) is nonnegative. The function \(L\) is therefore increasing and its minimum is \(L(-1) = -1\), exactly the minimum of \(f\). On \([0, 1]\) the function is convex, \(\alpha = 0\) and \(L = f\). After one split the bound equals the true minimum.

alphaBB on x^3 over [-1, 1], and one split at 0

                            [-1, 1]
             f'' = 6x >= -6, so alpha = 3 and
             L(x) = x^3 + 3(x^2 - 1)
             min L = L(0) = -3, against min f = f(-1) = -1;
             certificate alpha (u - l)^2/4 = 3, attained at 0
                    /                          \
                   /         split at 0         \
                  /                              \
             [-1, 0]                           [0, 1]
  interval Hessian [-6, 0]: alpha = 3     f convex: alpha = 0,
  L(x) = x^3 + 3x(x + 1)                  L = f
  L'(x) = 3(x + 1)^2 >= 0, so
  min L = L(-1) = -1 = min f

  after one split the bound equals the true minimum

(Two variables: the scaled and the plain α) In two variables the Gershgorin formula has something to do, and a computation for this post works one case through. Take \(f(x, y) = x y^2 - x^2\) on \(B = [0, 2] \times [0, 1]\). The Hessian is

\[\nabla^2 f(x, y) \;=\; \begin{pmatrix} -2 & 2y \\ 2y & 2x \end{pmatrix},\]

and evaluating its entries in interval arithmetic on \(B\) gives the interval Hessian \(h_{11} \in [-2, -2]\), \(h_{12} \in [0, 2]\), \(h_{22} \in [0, 4]\). With \(d = (1, 1)\), formula (2.4.5) gives \(\alpha_1 = \max\{0, -\tfrac12(-2 - 2)\} = 2\) and \(\alpha_2 = \max\{0, -\tfrac12(0 - 2)\} = 1\). The separation certificate of (c) is \(2 \cdot 2^2/4 + 1 \cdot 1^2/4 = 2.25\), and the bound is \(\min_B L_\alpha = -4.0835\) against the true minimum \(-4\) at \((2, 0)\). With the scaled choice \(d = u - l = (2, 1)\) the off-diagonal entry is weighted by \(d_2/d_1 = \tfrac12\) in the first row and by \(d_1/d_2 = 2\) in the second. So \(\alpha_1 = \max\{0, -\tfrac12(-2 - 1)\} = 1.5\) and \(\alpha_2 = \max\{0, -\tfrac12(0 - 4)\} = 2\). The certificate falls to \(1.5 \cdot 1 + 2 \cdot \tfrac14 = 2.0\), but the bound gets worse: \(L_\alpha\) is now minimized on the edge \(x = 2\), where it equals \(4y^2 - 2y - 4\), with minimum \(-4.25\) at \(y = \tfrac14\). Minimizing the worst-case separation is not the same as maximizing the bound, because the bound is the minimum of \(L_\alpha\) and depends on where the function is low. The following script performs the computation. It also runs it on \([0, 2]^2\) and on the child \([0, 1] \times [0, 2]\) of a bisection in \(x\), where the scaled \(\alpha\) grows from \((3, 2)\) to \((5, 1)\) as part (e) warned. Its last line compares αBB with the McCormick envelope on the bilinear term.

# alphaBB on f(x, y) = x y^2 - x^2.
#
# The interval Hessian, the Gershgorin alphas (2.4.5) in the plain
# (d = 1) and the scaled (d = u - l) variant, the separation certificate
# of Theorem 2.4.19(c) (column "sep.") and the bound min L, on
# [0, 2] x [0, 1], on [0, 2]^2 and on the child [0, 1] x [0, 2] of a
# bisection in x; then alphaBB against McCormick on the bilinear term
# x y (numpy only).

import numpy as np

f = lambda x, y: x*y**2 - x**2

def interval_hessian(xl, xu, yl, yu):
    """H = [[-2, 2y], [2y, 2x]] evaluated in interval arithmetic."""
    return {'11': (-2.0, -2.0), '12': (2*yl, 2*yu), '22': (2*xl, 2*xu)}

def gershgorin(H, d):
    """The alphas of (2.4.5):

    alpha_i = max{0, -1/2 (h_ii^L - sum_j max|h_ij| d_j/d_i)}.
    """
    m12 = max(abs(H['12'][0]), abs(H['12'][1]))
    return (max(0.0, -0.5*(H['11'][0] - m12*d[1]/d[0])),
            max(0.0, -0.5*(H['22'][0] - m12*d[0]/d[1])))

def alphabb(box, scaled, n=2001):
    """Steps 1 to 5 of Algorithm 2.4.20, the convex solve done on a
    fine grid."""
    xl, xu, yl, yu = box
    d = (xu - xl, yu - yl) if scaled else (1.0, 1.0)
    a = gershgorin(interval_hessian(*box), d)
    X, Y = np.meshgrid(np.linspace(xl, xu, n), np.linspace(yl, yu, n),
                       indexing='ij')
    # the underestimator (2.4.4)
    L = f(X, Y) + a[0]*(X - xl)*(X - xu) + a[1]*(Y - yl)*(Y - yu)
    i = np.unravel_index(np.argmin(L), L.shape)
    # the certificate of part (c)
    sep = a[0]*(xu - xl)**2/4 + a[1]*(yu - yl)**2/4
    return a, sep, L[i], (X[i], Y[i]), f(X, Y).min()

print("               d      alpha         sep.    min L    at"
      "              min f")
for box in [(0, 2, 0, 1), (0, 2, 0, 2), (0, 1, 0, 2)]:
    for scaled in (False, True):
        a, sep, lb, at, fmin = alphabb(box, scaled)
        dname = 'u - l' if scaled else '1    '
        print(f"[{box[0]},{box[1]}] x [{box[2]},{box[3]}]  {dname}  "
              f"({a[0]:.2f}, {a[1]:.2f})  {sep:.4f}  {lb:.4f}  "
              f"({at[0]:.3f}, {at[1]:.3f})  {fmin:.4f}")

# the bilinear term x y on the unit square
X, Y = np.meshgrid(np.linspace(0, 1, 1001), np.linspace(0, 1, 1001),
                   indexing='ij')
L = X*Y + 0.5*X*(X - 1) + 0.5*Y*(Y - 1)
print()
print(f"x*y on [0,1]^2: alpha = (0.5, 0.5), min L = {L.min():.4f};")
print("    min of the McCormick envelope max(0, x + y - 1) = "
      f"{np.maximum(0, X + Y - 1).min():.4f}")
               d      alpha         sep.    min L    at              min f
[0,2] x [0,1]  1      (2.00, 1.00)  2.2500  -4.0835  (1.986, 0.168)  -4.0000
[0,2] x [0,1]  u - l  (1.50, 2.00)  2.0000  -4.2500  (2.000, 0.250)  -4.0000
[0,2] x [0,2]  1      (3.00, 2.00)  5.0000  -5.6569  (1.414, 0.586)  -4.0000
[0,2] x [0,2]  u - l  (3.00, 2.00)  5.0000  -5.6569  (1.414, 0.586)  -4.0000
[0,1] x [0,2]  1      (3.00, 2.00)  2.7500  -2.6185  (0.602, 0.769)  -1.0000
[0,1] x [0,2]  u - l  (5.00, 1.00)  2.2500  -2.1874  (0.575, 0.635)  -1.0000

x*y on [0,1]^2: alpha = (0.5, 0.5), min L = -0.1250;
    min of the McCormick envelope max(0, x + y - 1) = 0.0000

(Three things visible in the output) The script's cost per box is the interval Hessian, the two Gershgorin sums and one convex minimization, which it does on a \(2001 \times 2001\) grid because the box is two-dimensional. A solver uses a few projected Newton steps instead. The boxes are independent of each other. Three things are visible in the output. On the square \([0, 2]^2\) the scaled and the plain choice coincide, because all ratios \(d_j/d_i\) are \(1\). On the child \([0, 1] \times [0, 2]\) the plain \(\alpha\) stays at \((3, 2)\), since the interval Hessian did not shrink in the entries that matter. The scaled \(\alpha\) moves to \((5, 1)\), and the scaled certificate falls from \(5\) to \(2.25\), by more than the plain one, which falls from \(5\) to \(2.75\). And on the last line, for the bilinear term \(xy\) on \([0, 1]^2\) the Gershgorin value \(\alpha = (\tfrac12, \tfrac12)\) is exact, since the Hessian has eigenvalue \(-1\), and \(L(x, y) = xy + \tfrac12 x(x - 1) + \tfrac12 y(y - 1) = \tfrac12 (x + y)(x + y - 1)\) has minimum \(-\tfrac18\). The McCormick envelope \(\max\{0, x + y - 1\}\) has minimum \(0\). The two have the same maximum separation \(\tfrac14\) and αBB has the weaker bound, so for a term with a known envelope αBB is dominated even when its \(\alpha\) is exact.

Theorem 2.4.19(e) on f = x y^2 - x^2: a bisection of [0, 2]^2 in x

  y
    0        1        2  x
                       d = 1 (plain)       d = u - l (scaled)
  parent  alpha        (3, 2)              (3, 2), d = (2, 2)
          separation   5                   5
          min L        -5.6569             -5.6569
  child   alpha        (3, 2)              (5, 1), d = (1, 2)
          separation   2.75                2.25
          min L        -2.6185             -2.1874
  with d fixed, alpha cannot grow, and the separation is at most the
  parent's with its x-term divided by four; with d = u - l
  recomputed, d_2/d_1 doubles and alpha_1 grows from 3 to 5

Algorithm 2.4.20 (αBB lower bound on a box).

Algorithm 2.4.20  αBB lower bound on a box

Input   f in C^2 on B = [l, u] as an expression DAG; interval arithmetic
        for the second-derivative DAG; a scaling d > 0 (d = u - l for
        the scaled variant, d = 1 for the plain one)
Output  a convex underestimator L_alpha and a bound beta <= min_B f

1  [H] = interval Hessian of f on B
   (interval arithmetic on each second-derivative expression)

2  for i = 1..n:                                                     (2.4.5)
      alpha_i = max{ 0, -1/2 ( h_ii^L - sum_{j != i}
                               max(|h_ij^L|, |h_ij^U|) d_j / d_i ) }

3  L_alpha(x) = f(x) + sum_i alpha_i (x_i - l_i)(x_i - u_i)

4  x_L = argmin_B L_alpha by any local method for a smooth convex
   box-constrained problem (projected gradient or projected Newton);
   every KKT point is a global minimizer because L_alpha is convex

5  return beta = L_alpha(x_L) and the certificate
      f(x) - L_alpha(x) <= sum_i alpha_i (u_i - l_i)^2 / 4  on B

Invariant
    nabla^2 L_alpha = nabla^2 f + 2 diag(alpha) >= 0 on B
    (Theorem 2.4.19(d)), so the output of step 4 is the global minimum
    of L_alpha on B and beta is a valid bound.
A run of Algorithm 2.4.20: f = x y^2 - x^2 over B = [0, 2] x [0, 1]

  Hessian of f:  [[-2, 2y], [2y, 2x]]
        |
        |  step 1: interval arithmetic on B
        v
  [H]:  h_11 in [-2, -2],  h_12 in [0, 2],  h_22 in [0, 4]
        |
        |  step 2: Gershgorin (2.4.5), two scalings d
        +------------------------------+
        |                              |
        v  d = (1, 1), plain           v  d = u - l = (2, 1), scaled
  alpha = (2, 1)                  alpha = (1.5, 2)
  certificate 2.25                certificate 2.0
        |                              |
        |  steps 3 to 5: L_alpha, its minimum over B, the bound
        v                              v
  beta = -4.0835                  beta = -4.25, at y = 1/4 on the
  at (1.986, 0.168)               edge x = 2
        \                              /
         +------------+---------------+
                      |
       both below min f = -4, at (2, 0); the smaller certificate
       came with the weaker bound

Per box, step 1 costs \(O(n^2 N)\) interval operations for a tape of length \(N\) and step 2 costs \(O(n^2)\). Step 4 is a small smooth convex problem, a few Newton steps on a dense \(n \times n\) system. The \(n^2\) interval entries and the \(n\) values of \(\alpha\) are independent of each other. Across boxes the whole algorithm is independent too, and the convex solves of many boxes are one batched Newton method.

(The envelope figure in αBB mode) The envelope figure above has an αBB mode. Switch its "relaxation" control to "alphaBB underestimator" and each piece \([a, b]\) of \([-2.6, 2.6]\) is relaxed by \(L(x) = f(x) + \alpha (x - a)(x - b)\) with the \(\alpha\) of the slider, drawn in orange with the band between \(f\) and \(L\) filled. For the quartic, \(f''(x) = 3x^2 - 3.2\) is smallest at \(x = 0\), where it is \(-3.2\), so on any piece containing \(0\) the smallest \(\alpha\) that makes \(L\) convex is \(\alpha_{\min} = 1.60\). A piece whose \(\alpha\) is below its \(\alpha_{\min}\) draws \(L\) dashed in red, because by part (b) of the theorem it is not convex. A local method can then no longer certify the minimum of \(L\). By part (a) \(L\) still lies below \(f\), and the minimum the figure finds on its sampled grid is still a valid bound. In the one-piece view at the default \(\alpha = 1.60\), the sag \(\alpha (b - a)^2/4\) is \(10.816\), against a band of \(2.560\) for the envelope, and the minimum of \(L\) is \(-8.801\), against \(\min f = -0.995\). With four equal pieces \(1.30\) wide the sag per piece is \(0.676\), a sixteenth of the root's, as part (c) predicts for a quarter of the width. The figure prunes and splits on the minimum of \(L\) in place of the envelope minimum, so the same tree logic runs with a weaker bound. The comparison with the envelope mode is the point of having both.

Relaxations of algorithms and reduced-space formulations

Definition 2.4.10 asked for a function given as a tape of operations. Mitsos, Chachuat and Barton observed that this includes far more than closed-form expressions. Any algorithm with a fixed number of iterations is a factorable function of its inputs once its arithmetic is recorded. Examples are a fixed-point iteration run for a set number of steps, the solution of a parametric linear system, and a numerical integrator with fixed steps. Algorithm 2.4.15 relaxes it. The relaxations so obtained are called McCormick relaxations of algorithms, and they are what allows a global solver to bound a model that contains a simulation.Mitsos, Chachuat and Barton (2009), cited above. Relaxations of implicit functions and of embedded neural networks follow the same route: A. M. Schweidtmann and A. Mitsos, "Deterministic global optimization with artificial neural networks embedded", Journal of Optimization Theory and Applications 180 (2019). The rules produce nonsmooth relaxations because of the \(\max\), \(\min\) and \(\operatorname{mid}\) selections; a variant whose relaxations are continuously differentiable, so that gradients come from ordinary automatic differentiation, is K. A. Khan, H. A. J. Watson and P. I. Barton, "Differentiable McCormick relaxations", Journal of Global Optimization 67 (2017).

(Full space and reduced space) The same observation separates two formulations of one model. In a full-space formulation every intermediate quantity of the model is an optimization variable with an equality constraint, and the auxiliary variable method relaxes each equality. The relaxation is an LP whose size grows with the model, and bound tightening acts on every intermediate. In a reduced-space formulation only the degrees of freedom are variables. The model equations are evaluated inside the objective and constraint functions, and McCormick relaxations are propagated through that evaluation. Branch and bound then operates in the small space. Bongartz and Mitsos compared the two on process flowsheets. They found substantial reductions in solution time for the reduced-space formulations in their case studies. They also noted that the pointwise relaxations in the reduced space can be weaker than the lifted LP relaxations, because auxiliary variables remove the dependency problem and a forward sweep does not.D. Bongartz and A. Mitsos, "Deterministic global optimization of process flowsheets in a reduced space using McCormick relaxations", Journal of Global Optimization 69 (2017). The solver is MAiNGO: D. Bongartz, J. Najman, S. Sass and A. Mitsos, "MAiNGO – McCormick-based Algorithm for mixed-integer Nonlinear Global Optimization", technical report, Process Systems Engineering (AVT.SVT), RWTH Aachen University (2018), permalink.avt.rwth-aachen.de/?id=729717; the software's authors also include C. Witte. Its documentation, avt-svt.pages.rwth-aachen.de/public/maingo (read 5 October 2026), describes a solver built on McCormick relaxations of factorable functions, computed through MC++, that can work in a reduced variable space. In Julia the same architecture is EAGO, Wilhelm and Stuber (2022), cited above. For parallel evaluation the reduced space is decisive. The per-node work is a fixed tape over few variables, with no LP of growing size, and it can be evaluated for thousands of boxes at once.

How good a relaxation has to be is governed by how fast its gap shrinks with the box, and there is a sharp statement about why a cheap relaxation can be too cheap.

Definition 2.4.21 (convergence order). Consider a scheme that assigns to every box \(B\) a convex underestimator \(f_B^{\mathrm{cv}}\) and a concave overestimator \(f_B^{\mathrm{cc}}\) of \(f\) on \(B\). The scheme has pointwise convergence of order \(\beta > 0\) at \(\bar x\) if there is a constant \(C\) with \(f(\bar x) - f_B^{\mathrm{cv}}(\bar x) \le C\, w(B)^\beta\) for all boxes \(B \ni \bar x\) of sufficiently small width, and likewise for \(f_B^{\mathrm{cc}}(\bar x) - f(\bar x)\). The interval \([\min_B f_B^{\mathrm{cv}}, \max_B f_B^{\mathrm{cc}}]\) encloses the range \(f(B)\). The scheme has Hausdorff convergence of order \(\beta\) if there is a constant \(C\) with \(\max\{\min_B f - \min_B f_B^{\mathrm{cv}},\ \max_B f_B^{\mathrm{cc}} - \max_B f\} \le C\, w(B)^\beta\) for all boxes of sufficiently small width. In words, the enclosure sticks out of the true range by at most \(C\,w(B)^\beta\) on either side.A. Bompadre and A. Mitsos, "Convergence rate of McCormick relaxations", Journal of Global Optimization 52 (2012), where both notions are defined and the propagation rules below are proved.

Proposition 2.4.22 (second order for envelopes and αBB). Let \(f\) be twice continuously differentiable on a box \(B_0\) with \(-2\bar\alpha I \preceq \nabla^2 f(x) \preceq 2\bar\alpha I\) for all \(x \in B_0\). The envelope scheme \(B \mapsto (\operatorname{vex}_B f, \operatorname{cav}_B f)\) has pointwise convergence of order \(2\) with constant \(n\bar\alpha/4\) on every box \(B \subseteq B_0\), and Hausdorff distance \(0\) to the true range. The αBB scheme, with \(L_\alpha\) below and the mirror overestimator \(f(x) - \sum_i \alpha'_i (x_i - l_i)(x_i - u_i)\) above, has pointwise and Hausdorff convergence of order \(2\) with constant \(n\bar\alpha'/4\). This holds for any rule that keeps the two estimators convex and concave and keeps \(\alpha, \alpha'\) componentwise below a fixed \(\bar\alpha'\), as the Gershgorin rule with a fixed scaling does.

Proof. By Proposition 2.4.5(c) applied to \(f\) and to \(-f\), \(f - \operatorname{vex}_B f\) and \(\operatorname{cav}_B f - f\) are at most \((n\bar\alpha/4)\,w(B)^2\) on \(B\), which is the pointwise order. By Theorem 2.4.4 the envelopes have exactly the minimum and the maximum of \(f\) on \(B\), so the Hausdorff distance is \(0\). For αBB, Theorem 2.4.19(c) bounds \(f - L_\alpha\) by \(\sum_i \alpha_i (u_i - l_i)^2/4 \le (n\bar\alpha'/4)\,w(B)^2\), which gives the pointwise order, and \(\min_B f - \min_B L_\alpha \le \max_B (f - L_\alpha)\) gives the Hausdorff order on the lower side. The upper side is the same argument with \(-f\). The Gershgorin value with fixed \(d\) is bounded on \(B_0\) because the interval Hessian entries are bounded by the largest second derivative of \(f\) on \(B_0\). ∎

(Convergence order and the cluster problem) Bompadre and Mitsos prove rules for how these orders propagate through the operations of Algorithm 2.4.15. With library relaxations of second order, the composite McCormick relaxation keeps pointwise convergence of order \(2\). Its Hausdorff order is limited by the order of the interval bounds carried for the factors, and the natural interval extension gives those only to first order in general. The exact hypotheses are in the paper.Bompadre and Mitsos (2012), cited above; the multivariate rule is analysed in Najman and Mitsos (2016), cited above. The reason the order matters is the cluster problem, which Du and Kearfott identified and Wechsung, Schaber and Barton analysed. In branch and bound a node is fathomed, or pruned, when its bound exceeds the incumbent, and the set of open boxes is the frontier (Section 6.4). Near a minimizer, a box of width \(\delta\) has a true range of about \(c\,\delta^2\) above the optimum. A first-order relaxation overestimates that range by \(C\delta\), which for small \(\delta\) dwarfs \(c\delta^2\). No box near the minimizer can then be fathomed until \(\delta\) is of the order of \(\varepsilon/C\), and there are of the order of \((1/\delta)^n\) such boxes. The number of boxes that cannot be fathomed therefore grows without limit as the tolerance \(\varepsilon\) shrinks. A second-order relaxation overestimates by \(C\delta^2\), comparable to the true range. Wechsung, Schaber and Barton show that second order avoids this when the prefactor \(C\) is below a threshold that depends on the curvature of \(f\) at the minimizer. The number of unfathomable boxes is then bounded independently of \(\varepsilon\), and the threshold is in the paper.K. Du and R. B. Kearfott, "The cluster problem in multivariate global optimization", Journal of Global Optimization 5 (1994); A. Wechsung, S. D. Schaber and P. I. Barton, "The cluster problem revisited", Journal of Global Optimization 58 (2014). For a GPU design the lesson is quantitative. A cheap relaxation that drops to first order, an interval bound for instance, or a McCormick relaxation whose inner intervals are loose, multiplies the frontier near the optimum rather than the work per node. The throughput of the device has to pay for both.

(The one GPU measurement) Pointwise McCormick evaluation is the one relaxation technology that has been run on a GPU with published numbers. Gottlieb, Xu and Stuber transform symbolic expressions by source-code generation into CUDA kernels that evaluate McCormick relaxations, interval extensions and subgradients for thousands of points or boxes at once. This is the batched evaluation of the C++ tape above in practice. The package reports, as the authors' numbers, about \(9\) ns per relaxation evaluation on the GPU against \(237\) ns for the CPU library it was compared with. It also reports speedups of \(11\) to \(22\) times of the accompanying branch and bound over EAGO on its test problems. The supported operation set was still small at the time of reading: the four arithmetic operations, squaring and the exponential.R. X. Gottlieb, P. Xu and M. D. Stuber, "Automatic source code generation for deterministic global optimization with parallel architectures", Optimization Methods and Software 41 (2026); software SourceCodeMcCormick.jl, github.com/PSORLab/SourceCodeMcCormick.jl. The timings and the speedups are the authors' own, from the package README and the paper, read 4 October 2026.

Where this is used

Every general-purpose global solver named in Section 5 (BARON, Couenne, SCIP, ANTIGONE) is, at bottom, a factorable relaxation inside a branch and bound, with a great deal of engineering about which pieces to relax how.

BARON's branch-and-reduce builds the auxiliary variable method over a DAG of elementary terms. The convex sides of univariate terms are outer-approximated by tangents, added by a sandwich algorithm that places each new tangent where the gap to the secant is largest, and the concave sides by secants. Bilinear terms get the McCormick planes, fractional terms the envelope above, and quadratic constraints the multiterm polyhedral cuts. The node bound is an LP, surrounded by heavy domain reduction.N. V. Sahinidis, "BARON: a general purpose global optimization software package", Journal of Global Optimization 8 (1996); M. Tawarmalani and N. V. Sahinidis, "Global optimization of mixed-integer nonlinear programs: a theoretical and computational study", Mathematical Programming 99 (2004); Tawarmalani and Sahinidis (2005), cited above; Bao, Sahinidis and Tawarmalani (2009), cited above; Y. Puranik and N. V. Sahinidis, "Domain reduction techniques for global NLP and MINLP optimization", Constraints 22 (2017).

SCIP's expression framework has one estimator per operator: products through the planes (2.4.2), powers, exponentials and logarithms by tangents and secants, trigonometric functions by the hull construction of the catalogue. Nonlinear handlers detect structure in the DAG and add stronger relaxations for convex and concave subexpressions, quadratics, bilinear terms, perspective and second-order-cone structure. Everything is a linear cut in the lifted variables inside branch and cut, that is branch and bound with cutting planes added at the nodes (Section 3.3).S. Vigerske and A. Gleixner, "SCIP: global optimization of mixed-integer nonlinear programs in a branch-and-cut framework", Optimization Methods and Software 33 (2018); K. Bestuzheva, A. Chmiela, B. Müller, F. Serrano, S. Vigerske and F. Wegscheider, "Global optimization of mixed-integer nonlinear programs with SCIP 8", Journal of Global Optimization 91 (2025).

Couenne, the open reference implementation of spatial branch and bound, introduces an auxiliary variable for every DAG node. It separates linearization cuts from the per-operator envelopes, including the odd-power envelopes of the catalogue.P. Belotti, J. Lee, L. Liberti, F. Margot and A. Wächter, "Branching and bounds tightening techniques for non-convex MINLP", Optimization Methods and Software 24 (2009).

ANTIGONE reformulates the model to a standard form and detects bilinear and multilinear terms, edge-concave aggregations and convex or concave univariate pieces. It then chooses a relaxation per term from a toolbox: McCormick and multilinear envelopes, edge-concave relaxations in the sense of Theorem 2.4.16, αBB-type underestimators for general terms, and piecewise-linear relaxations of bilinear terms.R. Misener and C. A. Floudas, "ANTIGONE: Algorithms for coNTinuous / Integer Global Optimization of Nonlinear Equations", Journal of Global Optimization 59 (2014).

Gurobi does not publish its relaxation rules. Its documentation describes nonlinear constraints \(y = f(x)\) built from arithmetic operations and a list of univariate functions, handled by spatial branch and bound, with bilinear terms as the special case its algorithms suit. For univariate functions the FuncNonlinear parameter chooses between a static piecewise-linear approximation and a dynamic outer approximation inside the tree, and the older function constraints are deprecated since version 13.0. These are vendor statements, and the documented ingredients are those of this subsection.Gurobi Optimization, Gurobi Optimizer Reference Manual, "Constraints" and "Parameters", docs.gurobi.com (accessed 4 October 2026).

The pointwise architecture is MAiNGO's and EAGO's, built on MC++ or on its Julia counterpart. The αBB codes relax whole functions at once rather than term by term.

What parallelizes

There are two architectures, and they parallelize differently. The auxiliary variable method produces one LP per node, in a lifted space whose size grows with the DAG. Its parallel form is a batch of LPs over a frontier of nodes, the subject of Sections 7.2 to 7.4. There the dual bounds of an inexact first-order solve are made safe before anything is pruned. The pointwise method produces, per node and per linearization point, one forward sweep of a straight-line tape with identical control flow for every box and every point. The C++ listing above showed the shape: \(O(N)\) arithmetic per element, selections that compile to branch-free selects, a node-major layout, and directed rounding for rigour. Its output is a closed-form bound (2.4.3), or a small dense LP in the \(n\) original variables when several points are linearized, and the per-node work does not grow with the model's intermediates. This is the work pattern a GPU executes efficiently, and the Gottlieb–Xu–Stuber numbers are the measurement that exists. Three things in this subsection are then design variables rather than theorems. The first is the number and placement of linearization points: one point gave the bound \(-2.19\) where the truth was \(0\), and how many points recover most of the LP bound is open. The second is the relaxation's convergence order. A scheme that drops to first order, through loose intervals or through a plain interval bound, multiplies the frontier near the optimum. A cheap relaxation in massive parallel therefore only wins if it keeps second order. The third is the choice between exact and recursive envelopes for multilinear terms. The exact envelope of a term in \(n \le 5\) variables is a dense LP with \(2^n\) columns and \(n + 1\) rows, small enough for a batched device solver. Current solvers use pairwise planes there and, on the one box \([1, 2]^3\) computed above, leave a band \(38\%\) wider on average. Partition cells and branch-and-bound nodes are the same work on different schedules, which is why a \(k^d\)-cell partition is a frontier evaluated at once. Section 7.8 ranks these as open problems with the first experiment for each.

Presolve

How tight a relaxation is depends on how the problem is written down. Section 2.1 showed two polygons around one set of integer points, and the solver that works on the looser polygon pays for it at every node. Section 2.4 built relaxations of nonconvex terms on the expression graph as the modeller wrote it. Between the two sits a step that every solver runs before it solves anything: it rewrites the formulation, without changing the optimum, so that the relaxation it then builds is smaller and tighter. That step is presolve. This subsection defines presolve and gives the reductions that matter, with proofs that they are valid. It then treats symmetry. It ends with the two parts of presolve that are specific to nonlinear problems: the rewriting of the expression graph into a standard form, and the detection of convexity. The second decides, for each node of the graph, whether it is relaxed by a tangent, by a chord or by an envelope.

The value of presolve has been measured. In the two CPLEX ablation studies, as they are usually summarized, presolve ranks second after cutting planes among the components whose removal hurts most. In the one open study, of 2026, it ranks first. The chapters of the two CPLEX studies could not be opened for this post. Their ranking is therefore quoted from secondary summaries, and their per-component speed-up factors, which are quoted everywhere, are not repeated here.R. E. Bixby, M. Fenelon, Z. Gu, E. Rothberg and R. Wunderling, "Mixed-integer programming: a progress report", in M. Grötschel (ed.), The Sharpest Cut (MPS-SIAM, 2004), 309–325; T. Achterberg and R. Wunderling, "Mixed integer programming: analyzing 12 years of progress", in Facets of Combinatorial Optimization (Springer, 2013). Both chapters are paywalled and their ablation tables were not re-read for this post. The ranking (cuts, then presolve, then branching) is taken from the secondary summaries in R. Bixby and E. Rothberg, "Progress in computational mixed integer programming: a look back from the other side of the tipping point", Annals of Operations Research 149 (2007), and A. Lodi, "Mixed integer programming computation", in M. Jünger et al. (eds), 50 Years of Integer Programming 1958–2008 (Springer, 2010), neither of which was re-read for this post either; it was not checked against the chapters' tables. Two ablations with open sources exist. On FICO's internal MINLP test set, disabling presolve altogether loses about 8% of the solved instances and slows the rest by about 40%. How that loss divides between the linear and the nonlinear halves of presolve is quoted below, where the two halves are compared.P. Belotti, T. Berthold, T. Gally, L. Gottwald and I. Pólik, "Solving MINLPs to global optimality with FICO Xpress Global", Optimization Online (July 2025), Section 4.5. The test set is the vendor's own, so the numbers are a vendor's measurement on a vendor's set. On SCIP 10, over 349 MILP instances and five seeds, switching presolving off slows the solver by a factor of 2.60, against 2.55 for random branching and 2.10 for switching cutting planes off.G. Mexi, The two faces of mixed-integer programming: primal and dual progress, doctoral thesis, Technische Universität Berlin (2026), doi 10.14279/depositonce-26588, read 5 October 2026; the thesis describes its factors as broadly consistent with Achterberg and Wunderling's for CPLEX 12.5. The canonical description of a modern presolve, with the measured value of each class of reduction in Gurobi, is Achterberg, Bixby, Gu, Rothberg and Weninger (2020), and the catalogue below follows theirs.T. Achterberg, R. E. Bixby, Z. Gu, E. Rothberg and D. Weninger, "Presolve reductions in mixed integer programming", INFORMS Journal on Computing 32 (2020). The LP-presolve classics are A. L. Brearley, G. Mitra and H. P. Williams, "Analysis of mathematical programming problems prior to applying the simplex algorithm", Mathematical Programming 8 (1975), and E. D. Andersen and K. D. Andersen, "Presolving in linear programming", Mathematical Programming 71 (1995); the MILP techniques are M. W. P. Savelsbergh, "Preprocessing and probing techniques for mixed integer programming problems", ORSA Journal on Computing 6 (1994), and G. Gamrath, T. Koch, A. Martin, M. Miltenberger and D. Weninger, "Progress in presolving for mixed integer programming", Mathematical Programming Computation 7 (2015).

What switching presolve off costs, in the two ablations with open sources. Top: FICO Xpress Global on FICO's internal MINLP test set, where disabling presolve altogether loses about 8% of the solved instances and slows the rest by about 40% (P. Belotti, T. Berthold, T. Gally, L. Gottwald and I. Pólik, "Solving MINLPs to global optimality with FICO Xpress Global", Optimization Online, July 2025, Section 4.5; the test set is the vendor's own, so the numbers are a vendor's measurement on a vendor's set). Bottom: SCIP 10 on 349 MILP instances and five seeds, where switching presolving off slows the solver by a factor of 2.60, against 2.55 for random branching and 2.10 for switching cutting planes off (G. Mexi, The two faces of mixed-integer programming: primal and dual progress, doctoral thesis, Technische Universität Berlin, 2026, doi 10.14279/depositonce-26588, read 5 October 2026). The two studies measure different solvers on different test sets in different units. The per-component factors of the two CPLEX studies are not drawn, as the text does not repeat them.

Reductions on the linear rows

Throughout this part the problem is written in the MILP form \(\min\{c^\top x : Ax \le b,\ l \le x \le u,\ x_j \in \mathbb{Z} \ (j \in I)\}\), with every nonlinear constraint of a MINLP either kept aside or, after the reformulation below, represented by linear rows plus defining constraints of library type. A row is \(\sum_j a_j x_j \le b\). An equality is two rows.

Definition 2.5.1. A presolve reduction is a transformation of the data \((A, b, c, l, u, I)\) together with a postsolve map that carries any solution of the transformed problem back to a solution of the original. A reduction is primal if it leaves the feasible set unchanged (it removes redundant rows, tightens bounds that every feasible point satisfies, fixes variables that every feasible point fixes). It is dual if it may remove feasible points but keeps at least one optimal solution (it uses the objective, as dual fixing does, or dominance between columns). Presolve applies reductions repeatedly until a round changes little, pushing each change on a postsolve stack, which is replayed in reverse once the reduced problem is solved.

Algorithm 2.5.2  Presolve to a fixed point

Input   MILP data (A, b, c, l, u, I); abort fraction kappa; round cap
Output  reduced data and a postsolve stack that maps any solution of
        the reduced problem to one of the original

 1. repeat

 2.    for each reduction R in
          [trivial rows and columns,
           activity-based bound tightening,
           coefficient tightening,
           parallel rows and columns,
           singleton columns and aggregation of doubleton equations,
           dual fixing,
           dominated columns,
           implied-integer detection,
           clique extraction and merging,
           probing,
           implied bounds,
           stuffing]:

 3.       apply R to every row or column it is defined for;
          push each change on the postsolve stack

 4. until the round changed less than a fraction kappa of the rows and
    columns, or the cap is reached

 5. return the reduced data and the stack

Invariant
    every primal reduction preserves the feasible set and every dual
    reduction preserves at least one optimal solution, so the reduced
    problem has the same optimal value; replaying the stack in reverse
    on any optimal solution of the reduced problem gives an optimal
    solution of the original.
Presolve, solve and postsolve: the loop of Algorithm 2.5.2

  data (A, b, c, l, u, I)
        |
        v
  +---> a round, steps 2 and 3:                   postsolve stack
  |     every reduction R on every row or column  +--------------+
  |     it is defined for -- push each change --> | last change  |
  |     |                                         |     ...      |
  |     |                                         | first change |
  |     v                                         +--------------+
  |     step 4: changed less than a fraction              |
  |     kappa of the rows and columns, or the             |
  |     cap reached?                                      |
  |     | no               | yes                          |
  +-----+                  v                              |
               step 5: the reduced data                   |
                           |                              |
                           v                              |
               solve: an optimal solution of the          |
               reduced problem                            |
                           |                              |
                           v                              |
               replay the stack in reverse <--------------+
               (last change first)
                           |
                           v
               an optimal solution of the original

SCIP's default abort fraction is \(\kappa = 8 \cdot 10^{-4}\), and after the root node it runs the whole loop again on the problem with the variables that the root fixed, a restart, when at least \(2.5\%\) of the integer variables were fixed.SCIP 10.0.0 source, src/scip/set.c: presolving/abortfac = 8e-4, presolving/restartfac = 0.025, presolving/maxrestarts = -1; T. Achterberg, Constraint integer programming, doctoral thesis, Technische Universität Berlin (2007), Section 10.9 for restarts. The reductions themselves are elementary. The three that matter most for a MINLP with indicator constraints, rows \(x \le M z\) with a binary \(z\) that force \(x = 0\) when \(z = 0\), are stated now with proofs.

Proposition 2.5.3 (activity bounds). For the row \(\sum_j a_j x_j \le b\) on the box \([l, u]\) let the minimal and maximal activities be

\[\underline{\alpha} \;=\; \sum_{a_j > 0} a_j l_j + \sum_{a_j < 0} a_j u_j, \qquad \overline{\alpha} \;=\; \sum_{a_j > 0} a_j u_j + \sum_{a_j < 0} a_j l_j ,\]

and let \(\underline{\alpha}_{-k}\) be the minimal activity of the row without the term \(a_k x_k\). Then (a) the row is redundant on the box if \(\overline{\alpha} \le b\) and infeasible on the box if \(\underline{\alpha} > b\). (b) Every point of the box that satisfies the row satisfies

\[x_k \;\le\; \frac{b - \underline{\alpha}_{-k}}{a_k} \quad (a_k > 0), \qquad x_k \;\ge\; \frac{b - \underline{\alpha}_{-k}}{a_k} \quad (a_k < 0),\]

and for \(k \in I\) the right-hand side may be rounded toward the inside of the interval.

Proof. The activity \(\sum_j a_j x_j\) is affine in each variable, so its extremes over the box are at the ends, which gives \(\underline{\alpha} \le \sum_j a_j x_j \le \overline{\alpha}\) on the box and (a). For (b), \(a_k x_k = \sum_j a_j x_j - \sum_{j \ne k} a_j x_j \le b - \underline{\alpha}_{-k}\), and dividing by \(a_k\) gives the two cases. An integer below a real number \(t\) is at most \(\lfloor t \rfloor\). ∎

The letters \(\underline{\alpha}\) and \(\overline{\alpha}\) denote row activities in this part only and are unrelated to the αBB parameter \(\alpha\) of Theorem 2.4.19. The row \(3x + 4y \le 10\) with \(x, y \in [0, 5]\) gives \(x \le 10/3\), hence \(x \le 3\) for integer \(x\). Iterated to a fixed point over all rows, (b) is feasibility-based bound tightening on linear rows, the subject of Section 2.6, and the solver runs it in presolve first because tighter bounds make every other reduction stronger.

Activity bounds on one row: 3x + 4y ≤ b on the box [0, 5]², the text's row at b = 10, with the bounds Proposition 2.5.3(b) reads off it, x ≤ (b − 4·l_y)/3 and y ≤ b/4, and their roundings for integer variables. A bound is drawn, with its value, only where it cuts the box. The lower bound l_y stands in for a bound another row has already tightened, the chart's own device for the text's remark that tighter bounds make every other reduction stronger: raising it tightens the bound on x. Values other than the text's b = 10, l_y = 0 are illustrative.

Proposition 2.5.4 (coefficient tightening; Savelsbergh, 1994). Let \(x_k\) be binary with \(a_k > 0\) in the row \(\sum_j a_j x_j \le b\), let \(\overline{\alpha}_{-k}\) be the maximal activity of the row without \(x_k\), and suppose \(d = b - \overline{\alpha}_{-k}\) satisfies \(0 < d < a_k\). Replace \(a_k\) by \(a_k - d\) and \(b\) by \(b - d\). Then the set of points of the box with \(x_k \in \{0, 1\}\) that satisfy the row is unchanged, and on the box with \(0 \le x_k \le 1\) the new row implies the old one and is strictly stronger wherever \(x_k < 1\).Savelsbergh (1994), cited above; Achterberg, Bixby, Gu, Rothberg and Weninger (2020), cited above, where the rule is one of the single-row reductions.

Proof. At \(x_k = 0\) the old row reads \(\sum_{j \ne k} a_j x_j \le b\) and the new one \(\sum_{j \ne k} a_j x_j \le b - d = \overline{\alpha}_{-k}\). Both hold at every point of the box, since the left side is at most \(\overline{\alpha}_{-k}\). At \(x_k = 1\) the old row reads \(\sum_{j \ne k} a_j x_j \le b - a_k\) and the new one \(\sum_{j \ne k} a_j x_j \le (b - d) - (a_k - d) = b - a_k\): the same inequality. So the integer sets coincide. On the box, the new row is \(\sum_{j \ne k} a_j x_j + a_k x_k \le b - d(1 - x_k)\), whose right-hand side is at most \(b\) for \(x_k \le 1\) and strictly less than \(b\) for \(x_k < 1\). ∎

A small example is the row \(10x + y \le 12\) with \(x\) binary and \(0 \le y \le 5\). Without \(x\) the activity is at most \(5\), so \(d = 12 - 5 = 7\), and the row becomes \(3x + y \le 5\). For \(x = 0\) both rows allow \(y \le 5\) and for \(x = 1\) both give \(y \le 2\). At \(y = 5\) the old relaxation allowed \(x \le 0.7\) and the new one forces \(x = 0\). The case that recurs in this series is the indicator constraint \(x \le M z\) with \(0 \le x \le u\), \(z\) binary and a loose \(M > u\): writing it on the complement \(z' = 1 - z\) as \(x + M z' \le M\), the maximal activity without \(z'\) is \(u\), so \(d = M - u\), and the row becomes \(x + u z' \le u\), that is \(x \le u z\). Presolve replaces a loose big-\(M\) by the smallest valid one. It cannot do more than that: the perspective reformulation of Section 4.3 is stronger than any big-\(M\), and no linear presolve reduction produces it.

Coefficient tightening on 10x + y ≤ 12 (x binary, 0 ≤ y ≤ 5): with d = 7 the row becomes 3x + y ≤ 5. The same integer points satisfy both rows, y ≤ 5 at x = 0 and y ≤ 2 at x = 1, and at y = 5 the LP allows x ≤ 0.7 before and x = 0 after.
The row before and after tightening, on the box 0 ≤ x ≤ 1, 0 ≤ y ≤ 5 with x binary: the LP region of 10x + y ≤ b, and of the row that coefficient tightening makes of it (Proposition 2.5.4; Savelsbergh, 1994; one of the single-row reductions of Achterberg, Bixby, Gu, Rothberg and Weninger, 2020). Without x the activity is at most 5, so d = b − 5, and replacing 10 by 10 − d and b by b − d gives (15 − b)x + y ≤ 5. For the text's b = 12, d = 7 and the row becomes 3x + y ≤ 5: for x = 0 both rows allow y ≤ 5 and for x = 1 both give y ≤ 2, and at y = 5 the old relaxation allowed x ≤ 0.7 and the new one forces x = 0. The other values of b, 6 to 14, are the same arithmetic on the text's row and are illustrative; the slider stops at 14 because at b = 15, d would reach a = 10, which the proposition excludes, and the row is redundant on the box (Proposition 2.5.3(a)).
Coefficient tightening on an indicator row x <= M z with M > u

  x <= M z                  0 <= x <= u, z binary, a loose M > u
     |
     |  written on the complement z' = 1 - z
     v
  x + M z' <= M             without z' the activity is at most u,
     |                      so d = M - u
     |  a_z' <- M - d = u,  b <- M - d = u
     v
  x + u z' <= u             that is x <= u z: the smallest valid M
A loose big-M and the smallest valid one: the indicator row x ≤ Mz, with 0 ≤ x ≤ u and z binary, drawn at u = 4, the bound of the second example below, and an illustrative M. Coefficient tightening (Proposition 2.5.4), applied on the complement z′ = 1 − z, turns it into x ≤ uz: the points with z = 0 or 1 are the same under both rows, and every point the relaxation loses has 0 < z < 1. Presolve can do no more than this: the perspective reformulation of Section 4.3 is stronger than any big-M, and no linear presolve reduction produces it. The slider moves M; the sentence above the panel gives the smallest z the relaxation of x ≤ Mz allows at x = 4.

Proposition 2.5.5 (dual fixing). In the minimization problem above, suppose \(c_j \ge 0\) and every row in which \(x_j\) appears has a coefficient \(a_{ij} \ge 0\) (rows written as \(\le\)). If the problem has an optimal solution and \(l_j\) is finite, some optimal solution has \(x_j = l_j\), so \(x_j\) may be fixed there. Symmetrically, if \(c_j \le 0\) and every coefficient \(a_{ij} \le 0\), some optimal solution has \(x_j = u_j\).

Proof. Take any feasible point and lower \(x_j\) to \(l_j\). Every row's activity decreases or stays, because its coefficient of \(x_j\) is nonnegative, so every row still holds, and the bounds hold by construction. The objective changes by \(c_j (l_j - x_j) \le 0\). Applied to an optimal solution this gives an optimal solution with \(x_j = l_j\). ∎

Dual fixing removes feasible points, so it is a dual reduction: it is correct for optimization and wrong for enumeration of all solutions. Dominated columns generalize it. If \(c_j \le c_k\) and \(a_{ij} \le a_{ik}\) in every row, and \(x_j\) is continuous or both variables are integer, then \(x_j\) is at least as good as \(x_k\) in the objective and no worse in every constraint. There is then an optimal solution with \(x_k\) at its lower bound unless \(x_j\) is at its upper bound. With finite bounds this fixes or tightens \(x_k\).Achterberg, Bixby, Gu, Rothberg and Weninger (2020), cited above, for the dual reductions; Gamrath, Koch, Martin, Miltenberger and Weninger (2015), cited above, for dominated columns and stuffing in SCIP. The remaining reductions of Algorithm 2.5.2 are listed in the table with the data they read and what they preserve.

reductionreadsdoeskindcost per round
activity boundsone row, the boxdrops redundant rows; detects infeasibilityprimal\(O(\mathrm{nnz})\)
bound tighteningone row, the box\(l_k\), \(u_k\) from the row (Prop. 2.5.3); rounds for integersprimal\(O(\mathrm{nnz})\), iterated
coefficient tighteningone row, a binary, the box\(a_k \leftarrow a_k - d\), \(b \leftarrow b - d\) (Prop. 2.5.4)primal\(O(\mathrm{nnz})\)
parallel rowshashes of coefficient patternsmerges \(a_k = \lambda a_i\) into one ranged rowprimal\(O(\mathrm{nnz})\) + hashing
parallel columnscolumn patterns, objectivemerges columns with proportional dataprimal\(O(\mathrm{nnz})\) + hashing
singleton / doubletona column in one row; an equality with two variablessubstitutes the variable out (aggregation)primal\(O(\text{row length})\)
dual fixingobjective sign, column signs\(x_j \leftarrow l_j\) or \(u_j\) (Prop. 2.5.5)dual\(O(\mathrm{nnz})\)
dominated columnspairs of columnsfixes or bounds the dominated onedualpairwise comparison
implied integersan equality of integers with integer data and one \(\pm 1\) termmarks a continuous variable integerprimal\(O(\mathrm{nnz})\)
probingone binary, propagationfixes, implications, implied boundsprimaltwo propagations per binary
clique mergingset-packing rows, implicationsmaximal cliques \(\sum_{j \in Q} x_j \le 1\) replace the rows they dominateprimalgraph work
stuffinga row and the objectivefixes several columns at once by their ratiosdual\(O(\text{row length})\)
The presolve catalogue (after Achterberg, Bixby, Gu, Rothberg and Weninger 2020)

Four terms in the table belong to the MILP trade. A singleton column is a variable that appears in one row, a doubleton equation is an equality with two variables, a ranged row is a two-sided row \(\ell \le a^\top x \le u\), and a set-packing row is a row \(\sum_{j \in Q} x_j \le 1\) over binary variables. Stuffing reads one row together with the objective and fixes several of the row's singleton columns at once, in the order of their objective-to-coefficient ratios. Three rows of the table need a sentence each. An implied integer is a continuous variable that is integer in every feasible solution, typically because it appears with coefficient \(\pm 1\) in an equality whose other variables are integer with integer coefficients and an integer right-hand side. Marking it integer lets cuts and branching use it, and SCIP 10 extends the detection to network submatrices of the constraint matrix.C. Hojny et al., "The SCIP Optimization Suite 10.0", arXiv 2511.18580 (2025), Section 3.2.1. Probing tentatively fixes a binary to \(0\) and to \(1\) and propagates each fixing with Proposition 2.5.3: if one side is infeasible the variable is fixed to the other. Bounds implied by both sides are globally valid. Implications of the form "\(x_j = 1\) implies \(x_k \le \beta\)" go into an implication graph that feeds the clique table and the implied-bound cuts of Section 3.6, inequalities derived from such implications. A second small example is a binary \(z\) with \(x \le 4z\), a row \(x + y \ge 5\) and \(y \le 3\): probing \(z = 0\) gives \(x \le 0\), hence \(y \ge 5 > 3\), which is infeasible, so \(z = 1\) is fixed at the root. A clique is a set of binaries of which at most one can be \(1\). Cliques read off set-packing rows and off the implication graph are merged into maximal ones, and one row \(\sum_{j \in Q} x_j \le 1\) replaces every row it dominates.

The block below runs the four reductions that have formulas above on the two examples, in exact arithmetic, and checks by brute force that the integer feasible set is unchanged where it should be.

# Four presolve reductions on two small systems.
#
# Activity-based bound tightening, coefficient tightening, dual fixing
# and probing, all but dual fixing checked by brute force on the integer
# points. The systems are the two examples of the text: (a) the row
# 10 x + y <= 12, (b) the binary z with x <= 4 z, x + y >= 5 and y <= 3.

from fractions import Fraction as F
from itertools import product

def act(a, lo, hi, skip=None, kind="max"):
    """Extreme activity of a row without variable `skip`."""
    if kind == "max":
        pick = lambda j: hi[j] if a[j] > 0 else lo[j]
    else:
        pick = lambda j: lo[j] if a[j] > 0 else hi[j]
    return sum(a[j] * pick(j) for j in a if j != skip)

def tighten_bounds(rows, lo, hi, integer):
    """Single-row bound tightening (FBBT on linear rows).

    Returns True if some bound moved.
    """
    changed = False
    for a, b in rows:
        for j in a:
            v = (b - act(a, lo, hi, skip=j, kind="min")) / a[j]
            if a[j] > 0 and v < hi[j]:
                # an upper bound, rounded down for an integer variable
                hi[j] = v if j not in integer else F(int(v // 1))
                changed = True
            if a[j] < 0 and v > lo[j]:
                # a lower bound, rounded up for an integer variable
                lo[j] = v if j not in integer else F(-int((-v) // 1))
                changed = True
    return changed

def coeff_tighten(a, b, lo, hi, binaries):
    """Savelsbergh 1994, for a binary j with a_j > 0.

    With d = b - maxact_{-j} in (0, a_j): a_j <- a_j - d and b <- b - d.
    """
    a = dict(a)
    for j in binaries:
        if a.get(j, 0) > 0:
            d = b - act(a, lo, hi, skip=j)
            if 0 < d < a[j]:
                a[j] -= d
                b -= d
    return a, b

def dual_fix(rows, c, lo, hi):
    """Min c.x: c_j >= 0 and every <= row has a_ij >= 0  ->  x_j = l_j."""
    return [j for j in c
            if c[j] >= 0 and all(a.get(j, 0) >= 0 for a, _ in rows)]

def feasible(rows, p):
    return all(sum(a[j] * p[j] for j in a) <= b for a, b in rows)

def probe(rows, lo, hi, integer, z):
    """Fix z = 0 and z = 1, propagate; an infeasible side fixes the other.

    verdict[v] is False when propagation empties the box of the side z = v.
    """
    verdict = {}
    for v in (0, 1):
        L, H = dict(lo), dict(hi)
        L[z] = H[z] = F(v)
        for _ in range(20):
            tighten_bounds(rows, L, H, integer)
        verdict[v] = all(L[j] <= H[j] for j in L)
    return verdict

# (a): the row 10 x + y <= 12 with x binary and 0 <= y <= 5
a, b = {"x": F(10), "y": F(1)}, F(12)
lo = {"x": F(0), "y": F(0)}
hi = {"x": F(1), "y": F(5)}
a2, b2 = coeff_tighten(a, b, lo, hi, ["x"])
print(f"(a)  d = b - maxact without x = {b - act(a, lo, hi, skip='x')};  "
      f"tightened row: {a2['x']} x + {a2['y']} y <= {b2}")

pts = [(x, y) for x in (0, 1) for y in range(6)]
same = ([p for p in pts if feasible([(a, b)], dict(zip("xy", p)))]
        == [p for p in pts if feasible([(a2, b2)], dict(zip("xy", p)))])
print(f"     integer points (x in {{0,1}}, y in 0..5): "
      f"same set under both rows: {same}")
print(f"     at y = 5 the LP allows x <= {(b - 5) / a['x']} before, "
      f"x <= {(b2 - 5) / a2['x']} after")

# (b): binary z, x <= 4 z, x + y >= 5, 0 <= x <= 4, 0 <= y <= 3;
# objective min 3 x + y + 2 z + 2 w with w in [0, 10] in a row x + w <= 6
rows = [({"x": F(1), "z": F(-4)}, F(0)),       # x - 4 z <= 0
        ({"x": F(-1), "y": F(-1)}, F(-5)),     # -x - y <= -5
        ({"x": F(1), "w": F(1)}, F(6))]        # x + w <= 6
lo = {"x": F(0), "y": F(0), "z": F(0), "w": F(0)}
hi = {"x": F(4), "y": F(3), "z": F(1), "w": F(10)}
c = {"x": F(3), "y": F(1), "z": F(2), "w": F(2)}

print("(b)  dual fixing: variables with c_j >= 0 "
      "and nonnegative coefficients")
print(f"       in every row: {dual_fix(rows, c, lo, hi)} "
      "-> fixed at their lower bound 0")

ver = probe(rows, lo, hi, {"z"}, "z")
print(f"     probing z: z = 0 feasible after propagation? {ver[0]}")
print(f"                z = 1 feasible? {ver[1]}"
      "  -> z fixed to 1 at the root")

# fix z = 1 and propagate to a fixed point
lo["z"] = hi["z"] = F(1)
while tighten_bounds(rows, lo, hi, {"z"}):
    pass
print("     bounds after fixing z = 1 and propagating: "
      f"x in [{lo['x']}, {hi['x']}], y in [{lo['y']}, {hi['y']}]")

brute = [(x, y, z)
         for x in range(5) for y in range(4) for z in (0, 1)
         if feasible(rows, {"x": x, "y": y, "z": z, "w": 0})]
print("     brute force over integer (x, y, z) with w = 0: "
      f"{len(brute)} feasible points,")
print(f"       all with z = {set(p[2] for p in brute)}, "
      f"x >= {min(p[0] for p in brute)}, y >= {min(p[1] for p in brute)}")
(a)  d = b - maxact without x = 7;  tightened row: 3 x + 1 y <= 5
     integer points (x in {0,1}, y in 0..5): same set under both rows: True
     at y = 5 the LP allows x <= 7/10 before, x <= 0 after
(b)  dual fixing: variables with c_j >= 0 and nonnegative coefficients
       in every row: ['w'] -> fixed at their lower bound 0
     probing z: z = 0 feasible after propagation? False
                z = 1 feasible? True  -> z fixed to 1 at the root
     bounds after fixing z = 1 and propagating: x in [2, 4], y in [1, 3]
     brute force over integer (x, y, z) with w = 0: 6 feasible points,
       all with z = {1}, x >= 2, y >= 1

(What a round costs, and PaPILO) Each reduction is a pass over the nonzeros of the rows it touches, and probing is two propagations per binary, which is why probing is the one reduction that solvers budget. Across rows and across binaries the work is independent. The parallel presolve library PaPILO is built on exactly that. Its presolvers run concurrently on an immutable copy of the problem, and each proposes a batch of reductions as a transaction. A sequential applier accepts a transaction if it touches no row or column already modified in the round, and discards it otherwise. The library is templated on the number type, so the same code runs in rational arithmetic for exact solving.A. Gleixner, L. Gottwald and A. Hoen, "PaPILO: a parallel presolving library for integer and linear optimization with multiprecision support", INFORMS Journal on Computing 35 (2023). SCIP 10's exact mode presolves with PaPILO in rational arithmetic (Hojny et al., cited above, Section 3.1.3). After fixing \(z = 1\) in the example, one more round of Proposition 2.5.3 gives \(x \in [2, 4]\) and \(y \in [1, 3]\), and the brute-force check confirms that the six integer points left all satisfy these bounds. A fixing found by probing feeds the next round of bound tightening, which is the reason the loop runs to a fixed point rather than once.

Probing z in the second example, x ≤ 4z, x + y ≥ 5, 0 ≤ x ≤ 4 and 0 ≤ y ≤ 3: with z = 0, x ≤ 4z gives x ≤ 0 and the row then needs y ≥ 5 > 3, which is infeasible, so z = 1 is fixed at the root; one more round of Proposition 2.5.3 then gives x ∈ [2, 4] and y ∈ [1, 3], and the brute-force check confirms that the six integer points left all satisfy these bounds. The slider moves the row's right-hand side r, and the sentence above the panels gives both sides' bounds and counts there; at r ≤ 3 neither side is empty.
Probing z in the second example, then one more round

  x <= 4z,  x + y >= 5,  0 <= x <= 4,  0 <= y <= 3,  z binary

                           probe z
                  z = 0 /          \ z = 1
                       /            \
        x <= 4z gives x <= 0,      feasible after
        so x + y >= 5 needs        propagation
        y >= 5 > 3: infeasible           |
                       |                 |
                       +--------+--------+
                                |
                                v
                     z = 1 fixed at the root
                                |
                                v
          one more round of Proposition 2.5.3 with z = 1:
            x + y >= 5 and y <= 3 give x >= 2:  x in [2, 4]
            x + y >= 5 and x <= 4 give y >= 1:  y in [1, 3]
                                |
                                v
          brute force: the six integer points left all
          satisfy these bounds

Symmetry

A formulation often has more symmetry than the problem it models. Four identical tax lots of which two are to be sold give \(\binom{4}{2} = 6\) selections that are indistinguishable in every constraint and in the objective. Ten lots of which five are sold give \(252\). A branch-and-bound tree that does not know this explores all of them, because pruning one does not prune the others. The remedy ranges from the trivial, an integer variable \(n \in \{0, \dots, 4\}\) counting the lots sold, which has one solution where the binaries had six, to the general machinery below.

Selections a tree cannot tell apart: the number of ways to sell two of m identical lots, C(m, 2), and five of m, C(m, 5), with the text's 6 at four lots and 252 at ten; each selection is indistinguishable in every constraint and in the objective, and a branch-and-bound tree that does not know this explores all of them. The text's integer variable n counting the lots sold has one solution per count, the line at 1. The letter m for the number of lots is the chart's, so that n keeps the text's meaning.

Definition 2.5.6. Let the binary variables of a formulation be indexed by a set \(J\). A permutation \(g\) of \(J\) acts on a point \(x\) by permuting its coordinates. The formulation group \(G\) is the group of permutations that map every feasible point to a feasible point with the same objective value. The orbit of a variable \(x_j\) under a subgroup \(H \le G\) is \(\{x_{h(j)} : h \in H\}\). At a node \(N\) of a branch-and-bound tree whose fixings are \(F_0\) (variables fixed to \(0\)) and \(F_1\) (fixed to \(1\)), the stabilizer \(\operatorname{stab}(N)\) is the subgroup of \(G\) that maps \(F_0\) onto itself and \(F_1\) onto itself.

Proposition 2.5.7 (orbital branching is valid; Ostrowski, Linderoth, Rossi and Smriglio, 2011). Let \(x_j\) be fractional in the relaxation at node \(N\) and let \(O\) be the orbit of \(x_j\) under \(\operatorname{stab}(N)\). Then the disjunction

\[x_j = 1 \qquad \text{or} \qquad \sum_{i \in O} x_i = 0\]

loses no optimal value below \(N\): the better of the two children's optima equals the optimum of \(N\).J. Ostrowski, J. Linderoth, F. Rossi and S. Smriglio, "Orbital branching", Mathematical Programming 126 (2011).

Proof. Take an optimal solution \(x\) of the subproblem at \(N\). If \(x_i = 0\) for every \(i \in O\), it lies in the second child. Otherwise \(x_i = 1\) for some \(i \in O\), and there is \(g \in \operatorname{stab}(N)\) with \(g(i) = j\). The point \(g x\) is feasible with the same objective value, it still satisfies the fixings of \(N\) because \(g\) maps \(F_0\) and \(F_1\) onto themselves, and it has \(x_j = 1\): it lies in the first child. ∎

The four-lot example shows the disjunction at work. Let \(x_1, \dots, x_4\) be the binary sale variables with the constraint \(\sum_i x_i = 2\), and suppose the root relaxation leaves \(x_1\) fractional. No variable is fixed at the root, so the stabilizer is the whole formulation group, the symmetric group on the four lots, and the orbit of \(x_1\) is \(\{x_1, x_2, x_3, x_4\}\). The two children are \(x_1 = 1\) and \(x_1 = x_2 = x_3 = x_4 = 0\). The second child violates \(\sum_i x_i = 2\) and is discarded at once. Ordinary branching on \(x_1\) would create the child \(x_1 = 0\) instead, which still contains the three selections that do not use lot 1, and the tree would go on to explore copies of selections it had already seen. In the surviving child the stabilizer is the symmetric group on lots 2, 3 and 4, the orbit of a fractional \(x_2\) is \(\{x_2, x_3, x_4\}\), and the same step leaves one child again, \(x_2 = 1\). The six equivalent selections are explored as one.

Branching at the root of the four-lot example, x_1 + ... + x_4 = 2

  {i,j}: the selection that sells lots i and j

  ordinary branching on x_1

                       root
             x_1 = 1 /      \ x_1 = 0
                    /        \
       {1,2} {1,3} {1,4}    {2,3} {2,4} {3,4}: the three selections
                            without lot 1, copies of selections
                            already seen

  orbital branching on x_1, fractional at the root; no variable is
  fixed, so its orbit under the whole group is {x_1, x_2, x_3, x_4}

                       root
             x_1 = 1 /      \ x_1 = x_2 = x_3 = x_4 = 0
                    /        \
                   |          violates x_1 + ... + x_4 = 2:
                   |          discarded at once
                   |
     stabilizer: all permutations of lots 2, 3, 4;
     orbit of a fractional x_2: {x_2, x_3, x_4}
             x_2 = 1 /      \ x_2 = x_3 = x_4 = 0
                    /        \
     {1,2}: x_1 = x_2 = 1     sum 1, not 2: discarded

  the six equivalent selections are explored as one

(Why the second child matters, and the alternatives) The point of the disjunction is the second child. Ordinary branching on \(x_j\) would create the child \(x_j = 0\), which still contains the symmetric copies with \(x_i = 1\) for \(i \in O \setminus \{j\}\). Orbital branching fixes the whole orbit to zero at once, because any solution with one of them at \(1\) is a copy of a solution in the other child. The group is found before the search by a graph automorphism computation on a coloured graph built from the formulation (one vertex per variable, per row and per distinct coefficient), and the stabilizers are recomputed as fixings accumulate. The same construction extends to nonlinear expressions by including the expression graph's nodes.F. Margot, "Symmetry in integer linear programming", in 50 Years of Integer Programming 1958–2008 (Springer, 2010), is the survey, and F. Margot, "Pruning by isomorphism in branch-and-cut", Mathematical Programming 94 (2002), the isomorphism-pruning method; L. Liberti, "Reformulations in mathematical programming: automatic symmetry detection and exploitation", Mathematical Programming 131 (2012), builds the detection graph for nonlinear expressions. The alternatives are polyhedral: symmetry-breaking inequalities that keep one representative point of each orbit of solutions, the lexicographically largest, and cut off the rest. Pfetsch and Rehn's computational comparison found that orbital fixing, the propagation form of the same idea, which fixes to zero the whole orbit of a variable that branching has fixed to zero, and the polyhedral methods each win on some classes and that detection is cheap enough to be on by default. SCIP's default enables the polyhedral handling and orbital reduction together with Schreier–Sims cuts, a family of inequalities built from a stabilizer chain of the group. SCIP 10 extends detection to reflection symmetries, maps that permute the variables and negate some of them, \(x_j \mapsto -x_j\).M. E. Pfetsch and T. Rehn, "A computational comparison of symmetry handling methods for mixed integer programs", Mathematical Programming Computation 11 (2019); SCIP 10.0.0 source, src/scip/set.c, misc/usesymmetry = 7; the polyhedral inequalities are those of SCIP's symresack and orbitope constraint handlers; Hojny et al. (2025), cited above, Section 3.3.1.

Reformulating the expression graph

For a MINLP, presolve has a second half that a MILP solver never sees. Section 2.4 defined a factorable function through its expression DAG (Definition 2.4.10) and built relaxations node by node. The intervals that this part and the next attach to the nodes of the DAG are those of Section 2.6, which defines them and proves their properties, and here they are taken as given. The DAG a solver relaxes is not the one the modeller wrote. It is the result of a symbolic reformulation that Smith and Pantelides introduced for spatial branch and bound and that every global solver performs in some form.E. M. B. Smith and C. C. Pantelides, "A symbolic reformulation/spatial branch-and-bound algorithm for the global optimisation of nonconvex MINLPs", Computers & Chemical Engineering 23 (1999). The DAG with shared subexpressions as the central data structure of a global solver is H. Schichl and A. Neumaier, "Interval analysis on directed acyclic graphs for global optimization", Journal of Global Optimization 33 (2005); the polyhedral version that solves an LP at every node is M. Tawarmalani and N. V. Sahinidis, "A polyhedral branch-and-cut approach to global optimization", Mathematical Programming 103 (2005).

Definition 2.5.8. The auxiliary-variable reformulation of a factorable problem introduces one new variable \(w_k\) for every nonlinear node \(v_k\) of the expression DAG and one defining constraint \(w_k = \mathrm{OP}_k(\cdot)\) of library type, so that the problem becomes

\[\min\Big\{ c^\top x + c_w^\top w \;:\; A x + A_w w \le b,\ \ w_k = \mathrm{OP}_k(w_{i(k)}, w_{j(k)}) \ \text{for each nonlinear node } k,\ \ l \le x \le u,\ \ x_j \in \mathbb{Z}\ (j \in I) \Big\},\]

in which every constraint is linear except the defining constraints, each of which is a product \(w_k = w_i w_j\), a power, an exponential, a logarithm, a quotient or another library function of one or two arguments. The box of each auxiliary variable is the natural interval extension of its node: the node's expression evaluated with each variable replaced by its interval and each operation by the operation on intervals that contains every possible result. On the unit square \(x + y\) gives \([0, 2]\) and \(e^{y}\) gives \([1, e]\). Section 2.6 defines this precisely and proves that it encloses the range (Theorem 2.6.4). The relaxation of the reformulated problem replaces each defining constraint by its envelope inequalities on the current box, which is the auxiliary variable method of Section 2.4.

Three things happen in the reformulation that affect every bound the solver will ever compute. First, common subexpressions are shared: the same product appearing in the objective and in three constraints becomes one node, one auxiliary variable and one set of envelope inequalities, and tightening its box tightens all four at once. Second, the expression is rewritten to a canonical form, with constants folded, sums flattened, products of a variable with itself recognized as a power and products reassociated. Section 2.4 showed why this matters. On \([-1, 1]\) the product \(x \cdot x\) relaxed by the product rule has minimum \(-1\), while \(x^2\) as a library function has minimum \(0\). On the box \([1, 2] \times [0, 1]^2\) the recursive McCormick relaxation of \(xyz\) is exact for the association order that multiplies the two zero-based variables first and strictly weaker for the other two (Section 2.4). Third, the reformulation decides what the solver branches on. SCIP and Couenne branch only on the original variables, never on the auxiliaries, so the choice of which intermediates become auxiliaries fixes the dimension of the spatial search. In SCIP 8 the expression is first simplified to a canonical form. Auxiliary variables are then introduced lazily, only for the subexpressions that a nonlinear handler cannot relax as a whole, which keeps the LP small; a nonlinear handler, named in Section 2.4 and described below, is a solver component that recognizes a structured subexpression and relaxes it as a unit.S. Vigerske and A. Gleixner, "SCIP: global optimization of mixed-integer nonlinear programs in a branch-and-cut framework", Optimization Methods and Software 33 (2018), Section 2.1; K. Bestuzheva, A. Chmiela, B. Müller, F. Serrano, S. Vigerske and F. Wegscheider, "Global optimization of mixed-integer nonlinear programs with SCIP 8", Journal of Global Optimization 91 (2025), Section 2.1. SCIP's branching/aux default never branches on auxiliary variables (src/scip/cons_nonlinear.c, SCIP 10.0.0). Xpress Global's ablation puts numbers on the two halves of presolve. Switching the nonlinear presolve off costs 138 solved instances on its test set, against 96 for the linear presolve, and within the nonlinear presolve formula simplification alone accounts for 82. Its authors read this as the nonlinear presolve contributing more than the linear one for MINLP.Belotti, Berthold, Gally, Gottwald and Pólik (2025), cited above, Section 4.5.

The two halves of presolve, measured: solved instances lost on Xpress Global's test set with the nonlinear presolve off (138, of which 82 from formula simplification alone) and with the linear presolve off (96). Numbers from Belotti, Berthold, Gally, Gottwald and Pólik, Optimization Online (July 2025), Section 4.5; the test set is the vendor's own, and the reading that the nonlinear half contributes more for MINLP is the authors'.
One product in the objective and in three constraints, shared

  as written                   after the reformulation

  objective -----> product     objective -----+
  constraint 1 --> product     constraint 1 --+
  constraint 2 --> product     constraint 2 --+--> one node,
  constraint 3 --> product     constraint 3 --+    one auxiliary
                                                   variable, one
                                                   set of envelope
                                                   inequalities

  tightening the box of the one auxiliary variable tightens all
  four at once
Two rewrites of the canonical form (the examples of Section 2.4)

  on [-1, 1]

        x * x                           x^2
         / \     recognized as a power   |
        x   x    -------------------->   x

    the product rule:               the library square:
    minimum -1                      minimum 0

  the recursive McCormick relaxation of xyz on [1, 2] x [0, 1]^2
  (x in [1, 2]; y and z in [0, 1]) in its three association orders

          *                    *                    *
         / \                  / \                  / \
        x   *                *   z                *   y
           / \              / \                  / \
          y   z            x   y                x   z

    y and z first,       x and y first:       x and z first:
    both zero-based:     strictly weaker      strictly weaker
    exact
The same square, two nodes: x · x relaxed by the product rule against x² as a library function, on a box [l, u] with l ≤ 0 < u. The product rule bounds x · x below by the tangents of x² at the two ends, 2lx − l² and 2ux − u², and above by the secant (l + u)x − lu, twice over; on [−1, 1] its minimum is −1, while x² as a library function is its own convex relaxation, with minimum 0 (Section 2.4). The canonical form recognizes the product of a variable with itself as a power. The sliders move l and u; the sentence above the panels gives where the two tangents meet and the product rule's minimum there.

Convexity detection: the tree walk

A spatial branch-and-bound solver treats a constraint differently according to its curvature. A convex constraint \(g(x) \le 0\) is relaxed exactly, in the limit, by its own tangent planes, which are the gradient cuts of Section 3.3. A concave function on the left of a \(\le\) is relaxed by its secant, which lies below it. The convex envelope of a concave function of one variable on an interval is the chord between its end values, and on a box in several variables it is the piecewise-affine function determined by the vertex values (Theorem 2.4.16). Anything else goes to the envelope machinery of Section 2.4. The labels have to be proved, not guessed. A tangent to a nonconvex function is not a valid inequality, and Section 3.4 shows outer approximation returning a wrong answer with a closed gap on the running example R3 for exactly this reason. The tree walk of Fourer, Maheshwari, Neumaier, Orban and Schichl proves the labels by applying the composition rules of convex analysis along the DAG. The interval of each node, from the forward pass of Section 2.6, supplies the monotonicity and sign information the rules need.R. Fourer, C. Maheshwari, A. Neumaier, D. Orban and H. Schichl, "Convexity and concavity detection in computational graphs: tree walks for convexity assessment", INFORMS Journal on Computing 22 (2010). The paper's abstract describes symbolic tools, implemented in the COCONUT environment and in the Dr. Ampl meta-solver, "which can be used to automatically detect the presence or absence of convexity and concavity in the objective and constraint functions, as well as convexity of the feasible set in some cases", supplemented "when it returns inconclusive results, with a numerical phase that may detect nonconvexity"; Dr. Ampl is R. Fourer and D. Orban, "DrAmpl: a meta solver for optimization problem analysis", Computational Management Science 7 (2010). The same rule set is the basis of disciplined convex programming: M. Grant, S. Boyd and Y. Ye, "Disciplined convex programming", in L. Liberti and N. Maculan (eds), Global Optimization: From Theory to Implementation (Springer, 2006).

Definition 2.5.9. Let \(B\) be a box and \(g : B \to \mathbb{R}\). The curvature label of \(g\) on \(B\) is one of aff (affine on \(B\)), cvx (convex on \(B\)), ccv (concave on \(B\)) and unk (nothing proved). The monotonicity label of \(g\) in the variable \(x_j\) on \(B\) is inc if \(g\) is nondecreasing in \(x_j\) for every fixed value of the other variables, dec if nonincreasing, cst if \(g\) does not depend on \(x_j\), and unk otherwise. A label other than unk is a certificate of the property it names. The label unk is not a claim of nonconvexity.

Proposition 2.5.10 (composition rules on a box). Let \(g, h : B \to \mathbb{R}\), \(c \in \mathbb{R}\), and let \(T\) be a univariate function defined on an interval \(I_g \supseteq g(B)\).

(a) An affine function is both convex and concave. The sum \(g + h\) is convex if both are convex and concave if both are concave. The multiple \(c\,g\) keeps the label of \(g\) if \(c > 0\), takes the opposite label if \(c < 0\), and is affine if \(c = 0\).

(b) A maximum of finitely many convex functions is convex. A minimum of finitely many concave functions is concave.

(c) \(T \circ g\) is convex on \(B\) if \(T\) is convex on \(I_g\) and either \(T\) is nondecreasing on \(I_g\) and \(g\) is convex, or \(T\) is nonincreasing on \(I_g\) and \(g\) is concave. It is concave on \(B\) if \(T\) is concave on \(I_g\) and either \(T\) is nondecreasing and \(g\) is concave, or \(T\) is nonincreasing and \(g\) is convex. If \(g\) is affine, \(T \circ g\) has the curvature of \(T\) on \(I_g\) with no monotonicity condition.

(d) \(1/g\) is convex if \(g\) is concave and positive on \(B\), and concave if \(g\) is convex and negative on \(B\).

Proof. (a) follows by adding the defining inequalities and multiplying them by \(c\). (b) The epigraph of a maximum is the intersection of the epigraphs, and an intersection of convex sets is convex. A minimum of concave functions is the negative of a maximum of convex ones. (c) Let \(x, y \in B\), \(\lambda \in [0, 1]\) and \(z = \lambda x + (1 - \lambda) y\). In the first case \(g(z) \le \lambda g(x) + (1 - \lambda) g(y)\), both sides lie in \(I_g\), and \(T\) nondecreasing on \(I_g\) gives \(T(g(z)) \le T(\lambda g(x) + (1 - \lambda) g(y)) \le \lambda T(g(x)) + (1 - \lambda) T(g(y))\), the second inequality by the convexity of \(T\) on \(I_g\). In the second case \(g(z) \ge \lambda g(x) + (1 - \lambda) g(y)\) and \(T\) nonincreasing gives the same first inequality. The concave cases follow by applying the convex cases to \(-T\). If \(g\) is affine then \(g(z) = \lambda g(x) + (1 - \lambda) g(y)\) and the monotonicity of \(T\) is not used. (d) is (c) with \(T(z) = 1/z\), which is convex and nonincreasing on \((0, \infty)\) and concave and nonincreasing on \((-\infty, 0)\). ∎

(Where the interval enters, and where the rule loses) The interval enters in (c) and (d). The convexity and monotonicity of \(T\) are needed only on \(I_g\), the node's forward interval, and a box that excludes zero turns \(x^2\) monotone and \(1/x\) convex or concave. The square is convex everywhere but monotone only on half-lines, and this is where the rule loses. The function \((e^{x} - 1)^2\) is convex on \([-0.5, 1]\), since its second derivative \(2e^{x}(2e^{x} - 1)\) is positive for \(x > -\ln 2\). But the inner node \(e^{x} - 1\) is convex and not affine, with interval \([e^{-0.5} - 1, e - 1] = [-0.393, 1.718]\), on which the square is not monotone, so step 5 of the algorithm below returns unk. The function \(x^4\) written as \((x^2)^2\), by contrast, is certified on every box: the inner interval of \(x^2\) always lies in \([0, \infty)\), where the square is convex and nondecreasing, and the inner node is convex. The monotonicity that matters is that of \(T\) on the inner interval, not that of the inner function in \(x\). Products need a rule of their own.

Step 5 on two squares: the table entry for z² (cvx; dec on I ≤ 0, inc on I ≥ 0, else unk) is read on the inner interval I, which straddles 0 for (eˣ − 1)² on [−0.5, 1] and lies in [0, ∞) for (x²)².

Proposition 2.5.11 (products). (a) If \(g, h \ge 0\) are convex on \(B\) and similarly ordered, that is \((g(x) - g(y))(h(x) - h(y)) \ge 0\) for all \(x, y \in B\), then \(gh\) is convex on \(B\). (b) In particular, if \(g\) and \(h\) are convex, nonnegative functions of one and the same variable \(x_j\) on \(B\), both nondecreasing or both nonincreasing in \(x_j\), then \(gh\) is convex on \(B\). (c) Nonnegativity, convexity and monotonicity in every variable do not suffice in two or more variables: on \([0, 1]^2\) the functions \(g = x\) and \(h = e^{y}\) have all three properties and \(gh = x e^{y}\) has Hessian determinant \(-e^{2y} < 0\), so it is neither convex nor concave on any open set.

Proof. (a) With \(z = \lambda x + (1 - \lambda) y\), convexity and nonnegativity give \(g(z) h(z) \le (\lambda g(x) + (1 - \lambda) g(y))(\lambda h(x) + (1 - \lambda) h(y))\). The right side equals \(\lambda g(x) h(x) + (1 - \lambda) g(y) h(y) - \lambda (1 - \lambda)(g(x) - g(y))(h(x) - h(y))\), and the last term is nonpositive by the ordering hypothesis. (b) For functions of one variable that are both nondecreasing, or both nonincreasing, the two differences in (a) have the same sign whenever \(x_j \le y_j\) and whenever \(x_j \ge y_j\). (c) At \(x = (1, 0)\) and \(y = (0, 1)\) the differences are \(1\) and \(1 - e < 0\), and the Hessian \(\begin{pmatrix} 0 & e^{y} \\ e^{y} & x e^{y} \end{pmatrix}\) has negative determinant. ∎

A product node therefore receives a label only when one factor is constant, or when both factors depend on one and the same variable and (b) applies, or when a dedicated rule recognizes the node as part of a quadratic form. The monotonicity labels make (b) checkable and the variable sets prevent the error of (c). The walk is the following algorithm.

Algorithm 2.5.12  Tree walk for curvature and monotonicity labels on a box B
                  (after Fourer, Maheshwari, Neumaier, Orban and
                   Schichl 2010; rules of Propositions 2.5.10 and 2.5.11)

Input   the expression DAG in topological order; the box B; forward
        intervals I_k of every node (the natural interval extension,
        Section 2.6); a table giving for each library function T and
        any interval I its curvature and monotonicity on I
           (exp:  cvx, inc;
            log:  ccv, inc on (0, inf);
            z^2:  cvx, dec on I <= 0, inc on I >= 0, else unk;
            sqrt: ccv, inc;
            1/z:  cvx, dec on I > 0; ccv, dec on I < 0;
            |z|:  cvx, dec on I <= 0, inc on I >= 0)
Output  a curvature label L_k in {aff, cvx, ccv, unk}, monotonicity
        labels M_kj in {inc, dec, cst, unk} for every variable j, and
        the variable set var(k), for every node k

 for each node k in topological order (nodes of equal depth in parallel):

 1. source x_j:  L = aff, M_jj = inc, M_jl = cst for l != j, var = {j};
    constant:    L = aff, all M = cst

 2. a + b:
       L   = aff if both aff;
             cvx if both in {aff, cvx};
             ccv if both in {aff, ccv};
             else unk.
       M_j = inc if both in {inc, cst};
             dec if both in {dec, cst};
             cst if both cst;
             else unk

 3. c * a (c constant):
       L = L_a if c > 0, the opposite label if c < 0, aff if c = 0;
       M reversed if c < 0

 4. a * b (both non-constant):
       L   = cvx if var(a) = var(b) = {j}, I_a >= 0, I_b >= 0,
                    L_a and L_b in {aff, cvx}
                    and M_aj = M_bj in {inc, dec};
             else L = unk.                              [Proposition 2.5.11]
       M_j = inc if I_a, I_b >= 0 and M_aj, M_bj in {inc, cst};
             dec symmetrically;
             cst if both cst;
             else unk

 5. T(a), with (curv, mono) = table(T, I_a):
       L = curv if L_a = aff;
       L = cvx if curv = cvx and ((mono = inc and L_a in {aff, cvx})
                               or (mono = dec and L_a in {aff, ccv}));
       L = ccv if curv = ccv and ((mono = inc and L_a in {aff, ccv})
                               or (mono = dec and L_a in {aff, cvx}));
       else L = unk.
       M_j = M_aj if mono = inc;
             reversed if mono = dec;
             cst if M_aj = cst;
             else unk
                                                     [Proposition 2.5.10(c)]

 6. a / b: treat as a * (1/b) with 1/b labelled by step 5
                                                     [Proposition 2.5.10(d)]

 7. max(a, b): cvx if both in {aff, cvx};
    min(a, b): ccv if both in {aff, ccv};
               else unk                              [Proposition 2.5.10(b)]

 8. a node recognized as a quadratic form x^T Q x + q^T x + r:
       cvx if Q is positive semidefinite,
       ccv if negative semidefinite,
       aff if Q = 0,
       else unk (an LDL^T factorization decides)

 9. var(k) = the union of the children's variable sets

Use
    a constraint g(x) <= 0 with L = cvx is relaxed by gradient cuts;
    a node labelled ccv under a <= is replaced by its secant, an
      underestimator;
    a node labelled cvx that enters with a negative sign, or stands
      under a >=, is overestimated by its secant;
    L = unk sends the node to the envelope rules of Section 2.4.

Cost
    O(|DAG| n) with the monotonicity vectors, O(|DAG|) without, after
    one forward interval sweep.

Theorem 2.5.13 (soundness of the tree walk). If Algorithm 2.5.12 assigns the label cvx (ccv, aff) to node \(k\), then \(v_k\) is convex (concave, affine) on \(B\). If it assigns inc (dec, cst) in \(x_j\), then \(v_k\) is nondecreasing (nonincreasing, constant) in \(x_j\) on \(B\).

Proof sketch. Induction along the topological order. Sources are affine with the stated monotonicity. Each curvature rule is one of Proposition 2.5.10(a) to (d) or Proposition 2.5.11(b), applied under the induction hypothesis to the children and with intervals that are valid by the fundamental theorem of interval arithmetic (Theorem 2.6.4). Each monotonicity rule is a true implication: a sum of functions nondecreasing in \(x_j\) is nondecreasing in \(x_j\). A positive multiple preserves and a negative multiple reverses monotonicity. If \(T\) is nondecreasing on \(I_g\) then \(T \circ g\) has the monotonicity of \(g\) and if nonincreasing the reversed one. A product of two nonnegative functions both nondecreasing in \(x_j\) is nondecreasing in \(x_j\). A function is constant in \(x_j\) when none of its children depends on \(x_j\). ∎

(The walk on the figure's function) The walk on the function of the next subsection's figure, \(f(x, y) = x e^{y} - (x + y)^2\) on the unit square, runs as follows. The node \(v_3 = x + y\) is aff and inc in both variables. The node \(v_4 = v_3^2\) has inner interval \(I_{v_3} = [0, 2]\), on which the square is convex and increasing, and an affine inner node, so it is cvx and inc in both variables. The node \(v_1 = e^{y}\) is cvx, cst in \(x\) and inc in \(y\). The node \(v_2 = x \cdot v_1\) has two nonconstant factors with variable sets \(\{x\}\) and \(\{y\}\), so step 4 returns unk, which Proposition 2.5.11(c) shows to be the truth. The monotonicity rule still gives inc in both variables, since both factors are nonnegative on the box. The root \(f = v_2 + (-1) v_4\) adds unk to ccv, which is unk, and inc to dec, which is unk. The verdict is correct and not a weakness of the rules: the Hessian of \(f\) has eigenvalues \(-3.000\) and \(-1.000\) at the origin and \(-2.178\) and \(0.896\) at \((1, 1)\), so \(f\) is concave in one corner of the box and indefinite in the opposite one. What the labels still buy is the choice of relaxation node by node. The ccv label of \(-v_4\) says that the overestimator of \(v_4\) inside \(f = v_2 - v_4\) is the secant of the square over \([0, 2]\), and the unk label of \(v_2\) sends it to the bilinear envelope of Theorem 2.4.7 with the interval \([1, e]\) of \(e^{y}\). Both are what the figure in Section 2.6 draws in its McCormick pass.

The tree walk on f = x e^y - (x + y)^2 over the unit square

  at each node: curvature label; monotonicity in x, in y

                         f = v_2 - v_4
                         unk; unk, unk
       term -v_4:   (-1) /           \
       ccv; dec, dec    /             \
       v_4 = v_3^2                     v_2 = x v_1
       cvx; inc, inc                   unk (step 4); inc, inc
            |                           /        \
            |                          /          \
       v_3 = x + y in [0, 2]          /        v_1 = e^y in [1, e]
       aff; inc, inc                 /         cvx; cst, inc
        |       \                   /               |
        |        +------> x <------+                |
        |             aff; inc, cst                 |
        +-----------------------------------------> y
                                              aff; cst, inc

  v_2 goes to the bilinear envelope with e^y in [1, e]; the
  overestimator of v_4 inside f is its secant over [0, 2]
  f: Hessian eigenvalues -3.000, -1.000 at (0, 0) and -2.178,
  0.896 at (1, 1), so the verdict unk at the root is correct
Tangent or secant: what a label buys. The node v₄ = v₃² of the next subsection's figure on its inner interval [0, 2u], which is [0, 2] on the unit square. Left, as a convex node under ≤ it is cut by a tangent, a gradient cut; right, inside f = v₂ − v₄ it enters with a negative sign and is overestimated by the secant, with the gap widest at z = u. At u = 1 and z = 1, the point (0.5, 0.5) of the Section 2.6 figure, the pair is v₄ : [1.000, 2.000], as that figure prints. The half-size box is illustrative.
Unk at the root is the truth: eigenvalues of the Hessian along the segment from the origin to (x₀, y₀). Left, f = x·eʸ − (x + y)²: −3.000 and −1.000 at the origin and −2.178 and 0.896 at (1, 1), the text's values, so f is concave in one corner of the box and indefinite in the opposite one; right, the node x·eʸ, whose Hessian determinant −e²ʸ is negative everywhere (Proposition 2.5.11(c)). The values at (1, 0), at (0, 1) and along the segments are computed for the chart from the same Hessians and are not stated in the text.

The walk is a sufficient test, and the function it misses that matters most for this series is the perspective.

Proposition 2.5.14 (the perspective defeats the walk). For a convex \(g : \mathbb{R}^n \to \mathbb{R}\) the perspective \(t\, g(x/t)\) is convex on \(\{t > 0\}\). In particular \(p(x, t) = x^2 / t\) is convex on \(\mathbb{R} \times (0, \infty)\): its Hessian

\[\nabla^2 p(x, t) \;=\; \begin{pmatrix} 2/t & -2x/t^2 \\ -2x/t^2 & 2x^2/t^3 \end{pmatrix}\]

has trace \(2/t + 2x^2/t^3 > 0\) and determinant \(0\), hence is positive semidefinite. On any expression of \(p\) as a product or quotient of its parts, such as \((x \cdot x) \cdot (1/t)\) or \(x^2 / t\) on a box \([-1, 1] \times [\underline{t}, \overline{t}]\) with \(\underline{t} > 0\), Algorithm 2.5.12 labels \(x^2\) cvx, \(1/t\) cvx, and the product unk, because no product rule applies to two nonconstant factors of different variables. The label of \(x^2\) assumes the library square. The literal product \(x \cdot x\) on \([-1, 1]\) already receives unk by step 4, since the factors' interval is not nonnegative.The convexity of the perspective is S. Boyd and L. Vandenberghe, Convex Optimization (Cambridge University Press, 2004), Section 3.2.6; the function \(x^2/t\) is the one behind the perspective reformulations of Section 4.3.

Proof. A symmetric \(2 \times 2\) matrix with nonnegative trace and zero determinant has eigenvalues \(0\) and the trace, hence is positive semidefinite, and a \(C^2\) function with positive semidefinite Hessian on an open convex set is convex. The labels are read off Algorithm 2.5.12. ∎

Proposition 2.5.14: the walk on p(x, t) = x^2 / t

  box [-1, 1] x [t_lo, t_hi] with t_lo > 0; a / b is a * (1/b)

                  x^2 * (1/t): unk, since no product rule
                    /       \   applies to two nonconstant
                   /         \  factors of different variables
                 x^2         1/t
                 cvx         cvx
                  |           |
                  x           t

  yet p is convex on t > 0: its Hessian has trace
  2/t + 2x^2/t^3 > 0 and determinant 0
  x^2 here is the library square; the literal product x * x on
  [-1, 1] is already unk by step 4, since the factors' interval
  is not nonnegative

(Two remedies: atoms and handlers) Two remedies exist, and both are used. Disciplined convex programming makes such functions atoms with known curvature, so that the walk never sees inside them.Grant, Boyd and Ye (2006), cited above. The MINLP solvers instead attach nonlinear handlers to the DAG, components that recognize structure a composition rule cannot. SCIP 8 has handlers for quotients, for perspective structure, for second-order cones and for quadratic forms. Its detection applies the composition rules "in reverse order": from the requirement that an expression be convex it derives the conditions its arguments must meet, and it introduces an auxiliary variable for any argument that cannot meet them.Bestuzheva, Chmiela, Müller, Serrano, Vigerske and Wegscheider (2025), cited above; the quadratic handler decides convexity by the eigenvalues of the coefficient matrix. The detection of the rotated cone behind \(xy \ge 1\) (Section 1.2) happens here, by the SCIP 10.0.0 changelog entry on SOC detection quoted in Section 1.2. BARON runs "convexity identification at each node of the branch-and-bound tree", because the labels are box-dependent and a node deep in the tree has a smaller box. Khajavirad and Sahinidis report that feeding the labels into a hybrid LP/NLP relaxation strategy "increases by 30% the number of problems that are solvable to global optimality within 500 s" on their test libraries.A. Khajavirad and N. V. Sahinidis, "A hybrid LP/NLP paradigm for global optimization relaxations", Mathematical Programming Computation 10 (2018), both quotations from the paper.

The limits of the walk are not an accident of its rules. Deciding convexity is hard in exactly the sense that matters here.

Theorem 2.5.15 (Ahmadi, Olshevsky, Parrilo and Tsitsiklis, 2013). Unless P = NP, there is no polynomial-time algorithm, and no pseudo-polynomial-time algorithm, that decides whether a multivariate polynomial of degree four, or of any higher even degree, with rational coefficients is convex on \(\mathbb{R}^n\). Deciding strict convexity, strong convexity, quasiconvexity and pseudoconvexity of polynomials of even degree at least four is strongly NP-hard. Quasiconvexity and pseudoconvexity of odd-degree polynomials can be decided in polynomial time.A. A. Ahmadi, A. Olshevsky, P. A. Parrilo and J. N. Tsitsiklis, "NP-hardness of deciding convexity of quartic polynomials and related problems", Mathematical Programming 137 (2013); the statement is the paper's abstract, which records that the question was posed by N. Z. Shor in 1992. Proof in the paper. Three terms of the statement are not in Definition 1.5.1. A problem is strongly NP-hard if it stays NP-hard when the numbers in its input are bounded by a polynomial in the input length, and a pseudo-polynomial-time algorithm is one whose running time is polynomial in the input length and in the magnitude of the largest input number. A differentiable function \(f\) is pseudoconvex if \(\nabla f(x)^\top (y - x) \ge 0\) implies \(f(y) \ge f(x)\), so that every stationary point is a global minimizer.M. R. Garey and D. S. Johnson, Computers and Intractability: A Guide to the Theory of NP-Completeness (W. H. Freeman, 1979), Chapter 4, for strong NP-hardness and pseudo-polynomial time; O. L. Mangasarian, "Pseudo-convex functions", Journal of the Society for Industrial and Applied Mathematics, Series A: Control 3 (1965), for pseudoconvexity.

(Quadratics are easy, transcendentals undecidable) By contrast a quadratic \(x^\top Q x + q^\top x + r\) is convex on \(\mathbb{R}^n\) exactly when \(Q\) is positive semidefinite, which an \(LDL^\top\) factorization in rational arithmetic decides in polynomial time. This is the test of step 8 and the one every solver runs on quadratic forms. With transcendental functions the question is not merely hard but undecidable.

Proposition 2.5.16 (undecidability over the line; an elementary consequence of Richardson's theorem). Let \(\mathcal E\) be Richardson's class of expressions in one real variable: the rational constants, \(\pi\) and \(\ln 2\), the variable \(x\), closed under \(+\), \(-\), \(\times\) and composition, and containing \(\sin\), \(\exp\) and \(|\cdot|\). Richardson proved that the problem "\(E(x) = 0\) for all real \(x\)" is undecidable for \(E \in \mathcal E\). Consequently the problem "the function represented by \(E\) is convex on \(\mathbb{R}\)" is undecidable for \(E \in \mathcal E\).D. Richardson, "Some undecidable problems involving elementary functions of a real variable", Journal of Symbolic Logic 33 (1968); P. S. Wang, "The undecidability of the existence of zeros of real elementary functions", Journal of the ACM 21 (1974), removes the absolute value for the related zero-existence problem. The reduction to convexity below is an elementary consequence, written out for this post.

Proof. Map \(E\) to \(h_E = -\big(x\,E(x)\big)^2\), which lies in \(\mathcal E\). If \(E \equiv 0\) then \(h_E \equiv 0\) is convex. Conversely, suppose \(h_E\) is convex on \(\mathbb{R}\). It is bounded above by \(0\), and a convex function on \(\mathbb{R}\) that takes two different values is unbounded above: if \(h(a) < h(b)\) with \(a < b\), convexity gives \(h(t) \ge h(b) + (h(b) - h(a))(t - b)/(b - a)\) for \(t > b\), which tends to infinity, and the case \(h(a) > h(b)\) is symmetric as \(t \to -\infty\). So \(h_E\) is constant, equal to \(h_E(0) = 0\), hence \(x E(x) = 0\) for all \(x\), \(E(x) = 0\) for \(x \ne 0\), and \(E(0) = 0\) by continuity. A decision procedure for convexity on \(\mathbb{R}\) would therefore decide Richardson's problem. ∎

A convex function bounded above on the whole line is constant: the step that does the work in the proof of Proposition 2.5.16, written out for this post as an elementary consequence of Richardson (1968). The line through two points of h, extended both ways; a convex h lies above it outside the two points, so h ≤ 0 on the whole line forces the slope to be 0. E ≡ 0 gives h ≡ 0; E ≡ 1, an illustrative choice not in the text, gives h = −x², which is not convex; h = x²/2 − 1 is an illustrative convex function, not in the text, and the bounded interval the remark after the proof mentions is not drawn.

(What the two results bracket) The argument uses the whole line, since a convex function bounded above on a bounded interval need not be constant, and the class is far from the polynomials of Theorem 2.5.15. The two results bracket the practical situation. With transcendental library functions no decision procedure exists for convexity over the line. With polynomials of degree four none is efficient unless P = NP. Convexity of a polynomial on a box is a first-order sentence about the reals, so it is decidable in principle by Tarski's quantifier elimination, at a cost that grows doubly exponentially with the number of variables.G. E. Collins, "Quantifier elimination for real closed fields by cylindrical algebraic decomposition", in Automata Theory and Formal Languages, Lecture Notes in Computer Science 33 (Springer, 1975), 134–183, is the algorithm; J. H. Davenport and J. Heintz, "Real quantifier elimination is doubly exponential", Journal of Symbolic Computation 5 (1988), 29–35, is the lower bound. Solvers therefore certify what the composition rules and the handlers reach in linear time, and relax the rest.

Where this is used

Every solver named in Section 5 presolves, and the differences are in the nonlinear half. SCIP runs the linear reductions of Algorithm 2.5.2 through its presolver plugins and PaPILO. It then simplifies every expression to a canonical form, runs the nonlinear handlers' detection and introduces auxiliary variables where a handler asked for them. Feasibility- and optimality-based bound tightening run on the result before the first LP. It repeats the linear presolve at a restart.Vigerske and Gleixner (2018), cited above, Section 2; Bestuzheva et al. (2025), cited above, Section 2.1. Gurobi's presolve is the one Achterberg, Bixby, Gu, Rothberg and Weninger describe, and its documentation says nothing further about the nonlinear part beyond the translation of nonconvex quadratics into bilinear form before spatial branching.Achterberg, Bixby, Gu, Rothberg and Weninger (2020), cited above; Gurobi Optimization, Gurobi Optimizer Reference Manual, version 13.0, parameter NonConvex, docs.gurobi.com (accessed 5 October 2026). Xpress Global is the Xpress MILP solver with a nonlinear presolve, convexification cuts and spatial branching added, and its ablations are the ones quoted above.Belotti, Berthold, Gally, Gottwald and Pólik (2025), cited above. BARON's preprocessing is range reduction on the factorable decomposition and a randomized multistart local search for an incumbent, with the convexity identification of Khajavirad and Sahinidis repeated at every node.BARON User Manual, The Optimization Firm, v. 2026.9.10 (10 September 2026), Section 5, minlp.com/baron-user-manual (accessed 5 October 2026); Khajavirad and Sahinidis (2018), cited above. On the GPU side, cuOpt's MIP presolve has been PaPILO-based and on by default since release 25.10 of October 2025, with a bound-propagation presolve in the releases before it.NVIDIA, cuOpt Release Notes, version 26.08, section "New Features (25.10)": "MIP presolve using Papilo (enabled by default)", docs.nvidia.com/cuopt (read 5 October 2026); the 25.05 and 25.08 notes already mention a conditional-bounds presolve.

SCIP's presolve of a MINLP before the first LP, and a restart

  linear reductions of Algorithm 2.5.2  <----------------------+
  (presolver plugins and PaPILO)                               |
        |                                                      |
        v                                                      |
  simplify every expression to a canonical form                |
        |                                                      |
        v                                                      |
  nonlinear handlers' detection; auxiliary variables           |
  where a handler asked for them                               |
        |                                                      |
        v                                                      |
  feasibility- and optimality-based bound tightening           |
        |                                                      |
        v                                                      |
  first LP, root node                                          |
        |                                                      |
        +-- at least 2.5% of the integer variables fixed: -----+
            a restart repeats the linear presolve

What parallelizes

Presolve is mostly independent work over rows, columns and binaries, and PaPILO's transactional design shows how to run it concurrently without changing the result: proposers in parallel, one applier that resolves conflicts, a round structure. The expensive reduction, probing, is a batch of propagations that share the row structure and differ in one fixing each, which on a GPU is the same shape as the batched bound tightening of Section 2.6. The convexity labels are box-dependent and are recomputed per node by BARON. Over a frontier of boxes the walk is one batched traversal of one DAG with different interval data and identical control flow for every box. How often labels flip deep in a tree, and what the switch of relaxation is worth, is a question left open here. Two things do not parallelize well. Symmetry detection is a graph automorphism computation, sequential and run once. The sequential applier at the heart of any deterministic presolve is the other, and Section 6 returns to why determinism is a product requirement and what it costs.

One PaPILO round: presolvers in parallel, one sequential applier

               an immutable copy of the problem
            /           |             |            \
           v            v             v             v
      presolver     presolver   ...  presolver    presolver
           |            |             |             |      in
           v            v             v             v      parallel
      transaction   transaction ...  transaction  transaction
      (a batch of reductions each)
            \           |             |            /
             v          v             v           v
        sequential applier, one transaction at a time:
        does it touch a row or column already modified
        in the round?
              | no                          | yes
              v                             v
           accept                        discard
              |
              v
        the round's result, from which the next round starts

Domain reduction

Every relaxation of Section 2.4 is built on a box, and its quality is a function of the box. Theorem 2.4.7 made this exact for the term that appears most: each McCormick envelope of \(xy\) on \([x^L, x^U] \times [y^L, y^U]\) misses the product by \((x^U - x^L)(y^U - y^L)/4\) at the centre of the box, so the separation is proportional to the box's area. Tightening both variables from \([0, 10]\) to \([2, 3]\) divides the area, and with it the separation, by \(100\). Tightening one of them divides it by \(10\). A global solver therefore spends much of its time shrinking boxes before it solves anything, and again at every node, and the techniques for doing so are the subject of this subsection. They are called domain reduction, or bound tightening, or range reduction, and they come in two kinds. A feasibility-based reduction uses only the constraints: it removes parts of the box that contain no feasible point. An optimality-based reduction uses the incumbent as well: it removes parts of the box that contain no feasible point better than the incumbent, which is all the search needs to keep. The second kind is why a global solver wants a good incumbent early, and why Section 3.2 comes before Section 3.5.

Tightening the box of xy: the McCormick separation at the centre, (x^U − x^L)(y^U − y^L)/4 (Theorem 2.4.7), is proportional to the box's area, so tightening one variable from [0, 10] to [2, 3] divides it by 10 and tightening both divides it by 100.

The idea that reduction should be a first-class step at every node, before the box is bounded and again after its relaxation is solved, is Ryoo and Sahinidis's, who named the result branch-and-reduce. Tawarmalani and Sahinidis gave the Lagrangian framework from which the reduction tests follow as corollaries. Belotti and co-authors fixed the vocabulary of FBBT and OBBT used below in their study of Couenne. Puranik and Sahinidis's survey is the account of how much of a global solver's performance comes from this step.H. S. Ryoo and N. V. Sahinidis, "Global optimization of nonconvex NLPs and MINLPs with applications in process design", Computers & Chemical Engineering 19 (1995), and "A branch-and-reduce approach to global optimization", Journal of Global Optimization 8 (1996); M. Tawarmalani and N. V. Sahinidis, "Global optimization of mixed-integer nonlinear programs: a theoretical and computational study", Mathematical Programming 99 (2004); P. Belotti, J. Lee, L. Liberti, F. Margot and A. Wächter, "Branching and bounds tightening techniques for non-convex MINLP", Optimization Methods and Software 24 (2009); Y. Puranik and N. V. Sahinidis, "Domain reduction techniques for global NLP and MINLP optimization", Constraints 22 (2017). The subsection begins with the arithmetic that makes a bound a proof. It then gives the three reduction techniques with their theorems, the two figures that run them, the certificates that interval methods add, the measured value in the solvers, and what in all of it a GPU can do.

Interval arithmetic

Interval arithmetic is arithmetic on sets. Every number in a computation is replaced by an interval that contains it, and every operation by an operation on intervals whose result contains every result the real operation could have produced. A computation then becomes a proof about a range, and with one further device, directed rounding, a floating-point computation becomes a proof that no rounding error can invalidate.R. E. Moore, Interval Analysis (Prentice-Hall, 1966), is the origin; the textbooks are A. Neumaier, Interval Methods for Systems of Equations (Cambridge University Press, 1990), and R. E. Moore, R. B. Kearfott and M. J. Cloud, Introduction to Interval Analysis (SIAM, 2009); the survey of its use in global optimization is A. Neumaier, "Complete search in continuous global optimization and constraint satisfaction", Acta Numerica 13 (2004).

Definition 2.6.1. An interval is a set \(X = [\underline{x}, \overline{x}] = \{x \in \mathbb{R} : \underline{x} \le x \le \overline{x}\}\) with \(\underline{x} \le \overline{x}\), and \(\mathbb{IR}\) is the set of intervals. Its width is \(w(X) = \overline{x} - \underline{x}\), its radius is \(\operatorname{rad} X = w(X)/2\), its midpoint is \(\operatorname{mid} X = (\underline{x} + \overline{x})/2\) and its magnitude is \(|X| = \max\{|\underline{x}|, |\overline{x}|\}\). A point \(x\) is identified with the point interval \([x, x]\). The interval hull of a bounded set \(S \subset \mathbb{R}\) is \(\square S = [\inf S, \sup S]\). A box is a product \(X = X_1 \times \dots \times X_n\) of intervals, with \(w(X) = \max_i w(X_i)\), and \(\operatorname{rad} X\) and \(\operatorname{mid} X\) are then the vectors of the sides' radii and midpoints. For \(X, Y \in \mathbb{IR}\) and \(\circ \in \{+, -, \cdot, /\}\), with \(0 \notin Y\) for division, the interval operation is the range

\[X \circ Y \;=\; \{\, x \circ y \;:\; x \in X,\ y \in Y \,\},\]

and for a continuous library function \(\varphi\) defined on \(X\), \(\varphi(X) = \{\varphi(x) : x \in X\}\).

Proposition 2.6.2 (the elementary operations compute exact ranges). For \(X, Y \in \mathbb{IR}\), with \(0 \notin Y\) for division, the sets \(X \circ Y\) are intervals, given by

\[\begin{aligned} X + Y &= [\underline{x} + \underline{y},\ \overline{x} + \overline{y}], &\qquad X - Y &= [\underline{x} - \overline{y},\ \overline{x} - \underline{y}], \\ X \cdot Y &= [\min P,\ \max P],\quad P = \{\underline{x}\underline{y}, \underline{x}\overline{y}, \overline{x}\underline{y}, \overline{x}\overline{y}\}, &\qquad X / Y &= X \cdot [1/\overline{y},\ 1/\underline{y}]. \end{aligned}\]

For a continuous \(\varphi\) monotone on \(X\), \(\varphi(X)\) is the interval between \(\varphi(\underline{x})\) and \(\varphi(\overline{x})\). For a function with interior extrema, such as the square, the range is read off the ends and the extrema, \(X^2 = [0, \max\{\underline{x}^2, \overline{x}^2\}]\) when \(0 \in X\).

Proof. The rectangle \(X \times Y\) is compact and connected and \((x, y) \mapsto x \circ y\) is continuous on it, so the image is a compact interval, and it suffices to locate its minimum and maximum. For fixed \(y\), each of \(x + y\), \(x - y\), \(xy\) and \(x/y\) is affine in \(x\), hence monotone, so its extremes over \(X\) are at \(\underline{x}\) or \(\overline{x}\). The same holds in \(y\) for fixed \(x\). The extremes over the rectangle are therefore attained at its four corners. For \(+\) and \(-\) the direction of monotonicity is fixed, which gives the two formulas. For \(\cdot\) the formula enumerates the corners. For \(/\) the identity \(x/y = x \cdot (1/y)\) and the monotonicity of \(1/y\) on an interval not containing \(0\) reduce it to the product. The statement for monotone \(\varphi\) is the intermediate value theorem. ∎

When \(0 \in Y\) the quotient set \(\{x/y : x \in X,\ y \in Y,\ y \ne 0\}\) is the whole line if \(0 \in X\), and otherwise a union of at most two unbounded intervals. Writing it out is Kahan's extended interval division, which Hansen's extended interval Newton method, given below, uses. For \(\overline{x} < 0\) and \(\underline{y} < 0 < \overline{y}\) it is \((-\infty, \overline{x}/\overline{y}] \cup [\overline{x}/\underline{y}, \infty)\), and the other cases are its mirror images.W. M. Kahan, "A more complete interval arithmetic", lecture notes for a summer course (1968), is the source usually cited for the extended division; the notes are not available online and are cited here from the secondary literature. The extended interval Newton method is E. R. Hansen, "A globally convergent interval method for computing and bounding real roots", BIT 18 (1978), 415–424. The textbook treatment of both is E. Hansen and G. W. Walster, Global Optimization Using Interval Analysis, 2nd edition (CRC Press/Marcel Dekker, 2004), Chapter 4.

Section 2.4 used the natural interval extension and the dependency problem informally (Definition 2.4.10). This part states and proves what they rest on.

Definition 2.6.3. Let \(f : D \subseteq \mathbb{R}^n \to \mathbb{R}\). An interval extension of \(f\) is a map \(F\) from boxes \(X \subseteq D\) to intervals with \(F([x, x]) = f(x)\) for every point \(x \in D\). It is inclusion isotone if \(X \subseteq Y\) implies \(F(X) \subseteq F(Y)\). The natural interval extension of an expression for \(f\) is obtained by evaluating the expression with every real operation and library function replaced by its interval version. An expression is single-use if every variable occurs in it at most once. An interval \(F(X) \supseteq f(X)\) is an enclosure of the range, and \(w(F(X)) - w(f(X))\) is its excess width.

Theorem 2.6.4 (the fundamental theorem of interval arithmetic; Moore, 1966). Let \(F\) be an inclusion-isotone interval extension of \(f\). Then \(f(X) \subseteq F(X)\) for every box \(X \subseteq D\). The natural interval extension of any expression for \(f\), with every library function evaluated by an inclusion-isotone enclosure of its range, is an inclusion-isotone interval extension, hence an enclosure of the range.Moore (1966), cited above; Neumaier (1990), cited above, Chapter 1.

Proof. Let \(x \in X\). Then \([x, x] \subseteq X\), so \(f(x) = F([x, x]) \subseteq F(X)\) by isotonicity, and \(f(X) \subseteq F(X)\) follows. For the natural extension, each elementary operation is inclusion isotone, because the range over a subset is a subset of the range: if \(X' \subseteq X\) and \(Y' \subseteq Y\) then \(\{x \circ y : x \in X', y \in Y'\} \subseteq \{x \circ y : x \in X, y \in Y\}\). A composition of isotone maps is isotone, so by induction along the expression every node's interval function is isotone. At a point box every node computes the real value of its operation, so the natural extension is an interval extension. ∎

(Where the excess width comes from) The picture is this. The interval of a node is the range of the node's operation over the product of its children's intervals. The true range is the range over the image of the box under the children's functions, which is a subset of that product and in general a proper one. The excess width of the natural extension is the gap between these two sets, accumulated node by node. The next proposition says when the gap is zero and how fast it closes.

Proposition 2.6.5 (single-use expressions are exact; the excess width is of first order; Moore, 1966). (a) If every variable occurs at most once in an expression tree for \(f\), and every library function in it is evaluated by its exact range on intervals, then \(F(X) = f(X)\) for every box \(X\). (b) If the operations of an expression for \(f\) are Lipschitz on a bounded box \(B\) (divisions with denominators bounded away from zero, library functions Lipschitz on the node intervals), there is a constant \(L\) with

\[w(F(X)) \;\le\; 2L\, w(X), \qquad\text{hence}\qquad w(F(X)) - w(f(X)) \;\le\; 2L\, w(X) \qquad \text{for every box } X \subseteq B .\]

Proof. (a) Induction on the tree. A leaf's interval is its range. At a binary node \(v = v_a \circ v_b\) the two subtrees have disjoint variable sets \(A\) and \(B\), because each variable occurs once, so the children's ranges are taken over independent coordinates \(X_A\) and \(X_B\) of the box. Hence \(\{v_a(x) \circ v_b(x) : x \in X\} = \{s \circ t : s \in v_a(X_A),\ t \in v_b(X_B)\} = v_a(X_A) \circ v_b(X_B)\) by Proposition 2.6.2, while the children's intervals are exactly \(v_a(X_A)\) and \(v_b(X_B)\) by the induction hypothesis. A unary library node preserves exactness by assumption. (b) Each elementary interval operation is Lipschitz for the Hausdorff distance \(q(X, X') = \max\{|\underline{x} - \underline{x}'|, |\overline{x} - \overline{x}'|\}\) on bounded sets, for instance \(q(X + Y, X' + Y') \le q(X, X') + q(Y, Y')\) and \(q(XY, X'Y') \le |Y|\, q(X, X') + |X'|\, q(Y, Y')\), so by induction \(q(F(X), F(X')) \le L\, q(X, X')\). Take \(X' = [x, x]\) for a point \(x \in X\): then \(q(X, [x, x]) \le w(X)\) and \(F([x, x]) = f(x)\), so \(F(X) \subseteq [f(x) - L\, w(X), f(x) + L\, w(X)]\). ∎

(The dependency problem and the wrapping effect) The failure of exactness when a variable occurs twice is the dependency problem of Definition 2.4.10, and it is intrinsic to interval evaluation rather than a bug. The simplest example is \(x - x\) on \([-1, 1]\), which evaluates to \([-2, 2]\) because the two occurrences are treated as independent. On \([0, 1]\) the expression \(x(1 - x)\) evaluates to \([0, 1] \cdot [0, 1] = [0, 1]\) against the range \([0, 1/4]\), while the algebraically equal single-use expression \(1/4 - (x - 1/2)^2\) evaluates to \([0, 1/4]\) exactly, as (a) promises. In the function of the second figure of this subsection, the expression graph of \(f(x, y) = x e^{y} - (x + y)^2\) on the unit square, the lower end \(-4\) of the natural interval pairs the minimum \(0\) of \(x e^{y}\), attained at \(x = 0\), with the maximum \(4\) of \((x + y)^2\), attained at \(x = y = 1\). No point of the box does both, and the natural interval \([-4, 2.718]\) is \(4.4\) times as wide as the range \([-1.282, 0.250]\). In vector-valued computations the same loss is called the wrapping effect. The interval image of the square \([-1, 1]^2\) under a rotation by \(45^\circ\) is the box \([-\sqrt 2, \sqrt 2]^2\) around the rotated square, of area \(8\) against \(4\), and after \(k\) rotations the box has side \(2(\sqrt 2)^k\) while the true image is still a square of side \(2\). Enclosing an intermediate vector by a box forgets the correlation between its components. The remedies are structural, not numerical: sharing subexpressions in the DAG, rewriting to single-use forms where possible, propagators that know the structure of a constraint, and the second-order enclosures (centred forms, Taylor models, affine arithmetic) that keep the first-order dependence between nodes exactly.Moore (1966), cited above, for the dependency problem and the wrapping effect; the second-order devices are K. Makino and M. Berz, "Higher order verified inclusions of multidimensional systems by Taylor models", Nonlinear Analysis 47 (2001), 3503–3514, and L. H. de Figueiredo and J. Stolfi, "Affine arithmetic: concepts and applications", Numerical Algorithms 37 (2004); their convergence orders are in A. Bompadre, A. Mitsos and B. Chachuat, "Convergence analysis of Taylor models and McCormick-Taylor models", Journal of Global Optimization 57 (2013). In the terminology of Definition 2.4.21, the natural extension converges to first order, and the cluster problem described in Section 2.4 (after Proposition 2.4.22, and stated as Theorem 3.5.7) says what that costs near a minimizer: the number of boxes that cannot be pruned grows as the tolerance shrinks. Interval bounds are cheap and rigorous and are not, on their own, the bounding rule of a competitive global solver.

The dependency problem: one variable, two occurrences

  x - x on [-1, 1]      x(1 - x) on [0, 1]      1/4 - (x - 1/2)^2
                                                on [0, 1]

      -  [-2, 2]            *  [0, 1]                 -  [0, 1/4]
     / \                   / \                       / \
     \ /                  |   -  [0, 1]           1/4   ^2
      x  [-1, 1]          |  / \                         |
                          | 1   |                        -
                           \   /                        / \
                             x  [0, 1]                 x   1/2

  the true ranges are {0}, [0, 1/4] and [0, 1/4]: the first two
  read the two occurrences of x as independent, the third is
  single-use and its interval is exact (Proposition 2.6.5 (a))
The wrapping effect: the interval image of the square [−1, 1]² under a rotation by 45° is the box [−√2, √2]² around the rotated square, of area 8 against 4, and after k rotations the box has side 2(√2)ᵏ while the true image is still a square of side 2. The slider sets k; at k = 2 the dashed square is the box of k = 1, rotated, and the new box is drawn around it.

Directed rounding, and why it makes a bound a proof

A floating-point number system \(\mathbb{F}\) (binary64 in what follows, together with \(\pm\infty\)) is a finite set, and the exact result of an operation on two of its elements is in general not one of them. IEEE 754 arithmetic returns the exact result rounded according to a rounding-direction attribute, and two of the attributes are rounding toward \(-\infty\) and toward \(+\infty\).IEEE, IEEE Standard for Floating-Point Arithmetic, IEEE Std 754-2019: each of the operations \(+\), \(-\), \(\times\), \(\div\), square root and fused multiply-add is computed as if the exact result were formed and then rounded under the current rounding-direction attribute, and roundTowardNegative and roundTowardPositive are two of the attributes. In C++ the attribute is set by std::fesetround(FE_DOWNWARD) and FE_UPWARD from <cfenv>, honoured by the compiler under #pragma STDC FENV_ACCESS ON.

Definition 2.6.6. For \(t \in \mathbb{R}\) let \(\mathrm{RD}(t) = \max\{a \in \mathbb{F} : a \le t\}\) and \(\mathrm{RU}(t) = \min\{a \in \mathbb{F} : a \ge t\}\) be the roundings toward \(-\infty\) and toward \(+\infty\). A machine interval has ends in \(\mathbb{F}\). The machine version of an elementary operation applies the formulas of Proposition 2.6.2 with every lower end rounded by \(\mathrm{RD}\) and every upper end by \(\mathrm{RU}\), for example \(X \oplus Y = [\mathrm{RD}(\underline{x} + \underline{y}),\ \mathrm{RU}(\overline{x} + \overline{y})]\) and \(X \odot Y = [\mathrm{RD}(\min P),\ \mathrm{RU}(\max P)]\). This is outward rounding. The machine version of a library function \(\varphi\) is any machine interval \(\tilde\varphi(X) \supseteq \varphi(X)\) that is monotone in \(X\), for instance the two end values widened by a certified bound on the error of the library's implementation.

Theorem 2.6.7 (machine interval arithmetic is an enclosure, whatever the rounding errors). Let \(f\) be given by an expression over \(+, -, \cdot, /\) and library functions, let \(F\) be its natural interval extension in exact arithmetic, and let \(\tilde F\) be its evaluation in machine interval arithmetic. Then for every box \(X\) with floating-point ends

\[f(X) \;\subseteq\; F(X) \;\subseteq\; \tilde F(X),\]

and \(\tilde F\) is inclusion isotone. The statement holds for any number of operations, any order of evaluation consistent with the expression and any floating-point format, and it is unaffected by overflow to \(\pm\infty\) at an end.

Proof. For one operation with exact result \(Z = X \circ Y = [\underline{z}, \overline{z}]\), the machine result \(\tilde Z\) has lower end \(\mathrm{RD}(\underline{z}) \le \underline{z}\) and upper end \(\mathrm{RU}(\overline{z}) \ge \overline{z}\), so \(Z \subseteq \tilde Z\). For the product the four candidates may be rounded individually, since \(\mathrm{RD}\) is monotone and so \(\mathrm{RD}(\min P) = \min \mathrm{RD}(P)\). For a library function \(\varphi(X) \subseteq \tilde\varphi(X)\) by assumption. Now induct along the expression. At the sources the exact and machine intervals coincide because the box has floating-point ends. At a node whose children carry machine intervals \(\tilde X_c \supseteq X_c\), isotonicity of the exact operation gives \(\mathrm{OP}(\tilde X_c) \supseteq \mathrm{OP}(X_c) = X_k\), and the machine rounding only enlarges: \(\tilde X_k \supseteq \mathrm{OP}(\tilde X_c)\). Hence \(\tilde X_k \supseteq X_k\) at every node, in particular at the root. Isotonicity of \(\tilde F\) follows from the monotonicity of \(\mathrm{RD}\) and \(\mathrm{RU}\) and of the exact formulas in the ends. The first inclusion is Theorem 2.6.4. Nothing in the argument counts operations or bounds the size of a rounding error, and an end that overflows to \(\pm\infty\) remains a bound. ∎

(Why a machine interval bound is a proof, and three conditions) This theorem is the reason interval bounds are called rigorous and tolerance-based bounds are not. A simplex solver's bound is accurate to within its tolerances on the data it saw, and nothing certifies the accumulated rounding of a long computation. A machine interval bound is correct by construction, in binary32 as in binary64, after a million operations as after one, on a GPU as on a CPU. Precision and operation count affect the width, not the validity. Three practical conditions attach to the hypotheses. Every rounding must be directed. A compiler that constant-folds or vectorizes under the default mode, or a library call that resets the mode, breaks the hypothesis silently. This is why the listing below uses volatile operands and the pragma, and why per-operation intrinsics are the cleaner route on a GPU. Subnormal results must not be flushed to zero, because an upper end flushed to \(+0\) lies below a positive exact value. The CUDA compiler's --use_fast_math turns that flushing on and rigorous kernels are compiled without it. And library functions need certified enclosures: the standard recommends but does not require correctly rounded \(\exp\), \(\log\) and trigonometric functions, so a rigorous implementation widens their values by a documented error bound.NVIDIA, NVIDIA CUDA Compiler Driver NVCC, options --ftz, --fmad, --prec-div, --prec-sqrt and --use_fast_math, docs.nvidia.com/cuda/cuda-compiler-driver-nvcc (read 5 October 2026): --use_fast_math implies --ftz=true, --prec-div=false, --prec-sqrt=false and --fmad=true. Under a directed mode a fused multiply-add is still a directed rounding of an exact quantity, so contraction does not break the enclosure; it breaks bitwise reproducibility and any analysis that assumed one rounding per written operation. S. M. Rump, "Verification methods: rigorous results using floating-point arithmetic", Acta Numerica 19 (2010), is the survey of how rounding modes are handled in verification methods.

(Directed rounding on a GPU, and on a CPU) On NVIDIA GPUs no global mode is set. Each operation carries its own direction through an intrinsic: __dadd_rd and __dadd_ru return a sum rounded down and rounded up, and the same pairs exist for multiplication, division and square root in double and in single precision, and for the fused multiply-add. A batched interval kernel over thousands of boxes is therefore a straight-line program with no mode switch at all.NVIDIA, CUDA Math API Reference Manual, CUDA 13, "Single precision intrinsics" and "Double precision intrinsics", docs.nvidia.com/cuda/cuda-math-api (read 5 October 2026); __fadd_rd is documented as "Add two floating-point values in round-down mode" and __fadd_ru as the same "in round-up mode". The full list is __fadd_rd, __fadd_ru, __fmul_rd, __fmul_ru, __fdiv_rd, __fdiv_ru, __fsqrt_rd and __fsqrt_ru in single precision and __dadd_rd, __dadd_ru, __dmul_rd, __dmul_ru, __ddiv_rd, __ddiv_ru, __dsqrt_rd, __dsqrt_ru, __fma_rd and __fma_ru in double precision. The implementation pattern is S. Collange, M. Daumas and D. Defour, "Interval arithmetic in CUDA", in GPU Computing Gems Jade Edition (Elsevier, 2012). On a CPU, switching the rounding mode is expensive and serializes the pipeline, and one switch suffices.

Proposition 2.6.8 (one rounding mode suffices). With the processor in roundTowardNegative, the upper end of each elementary operation can be computed without changing the mode: for floats \(a, b\),

\[\mathrm{RU}(a + b) = -\mathrm{RD}\big((-a) + (-b)\big), \quad \mathrm{RU}(a - b) = -\mathrm{RD}(b - a), \quad \mathrm{RU}(ab) = -\mathrm{RD}\big((-a)\, b\big), \quad \mathrm{RU}(a / b) = -\mathrm{RD}\big((-a) / b\big).\]

Proof. Negation is an order-reversing bijection of \(\mathbb{R}\) that maps \(\mathbb{F} \cup \{\pm\infty\}\) onto itself, so the largest float not exceeding \(-t\) is the negative of the smallest float not below \(t\): \(\mathrm{RD}(-t) = -\mathrm{RU}(t)\). Negation of a float is exact, so \((-a) + (-b)\), \(b - a\), \((-a) b\) and \((-a)/b\) are exact expressions for the negatives of \(a + b\), \(a - b\), \(ab\) and \(a/b\). ∎

The listing implements Definitions 2.6.1 and 2.6.6 in C++23 with this device: the mode is set once on entry and, for the arithmetic operations, every lower end is computed directly and every upper end is the negative of a rounded-down negative. It evaluates the dependency example, encloses \(1/3\), and computes the natural interval extension of that function, \(f(x, y) = x e^{y} - (x + y)^2\), on the unit square. The comment gives the shape of the same operations in CUDA.

// Machine interval arithmetic with directed rounding (C++23).
//
// One rounding mode, roundTowardNegative, is set once; every arithmetic
// upper end is the negative of a rounded-down negative, so no mode switch
// happens inside a computation. The program evaluates the dependency
// example, encloses 1/3, and computes the natural interval extension of
// f(x, y) = x exp(y) - (x + y)^2 on the unit square.

#include <cfenv>
#include <cmath>
#include <cstdio>

#pragma STDC FENV_ACCESS ON

struct I {
    double lo, hi;
};

// RAII guard: FE_DOWNWARD inside, the old mode restored on exit
struct RoundDown {
    int old;
    RoundDown() : old(std::fegetround()) { std::fesetround(FE_DOWNWARD); }
    ~RoundDown() { std::fesetround(old); }
};

// volatile keeps the compiler from folding or reordering the operations
// around the mode switch

static I add(I a, I b)
{
    volatile double lo = a.lo + b.lo;
    volatile double nh = (-a.hi) + (-b.hi);
    return {lo, -nh};
}

static I sub(I a, I b)
{
    volatile double lo = a.lo - b.hi;
    volatile double nh = (-a.hi) + b.lo;
    return {lo, -nh};
}

static I mul(I a, I b)
{
    // the four corner products, rounded down
    volatile double p1 = a.lo * b.lo, p2 = a.lo * b.hi;
    volatile double p3 = a.hi * b.lo, p4 = a.hi * b.hi;
    // their negatives, rounded down
    volatile double q1 = (-a.lo) * b.lo, q2 = (-a.lo) * b.hi;
    volatile double q3 = (-a.hi) * b.lo, q4 = (-a.hi) * b.hi;
    return {std::fmin(std::fmin(p1, p2), std::fmin(p3, p4)),
            -std::fmin(std::fmin(q1, q2), std::fmin(q3, q4))};
}

static I sqr(I a)
{
    if (a.lo >= 0) {
        volatile double lo = a.lo * a.lo;
        volatile double nh = (-a.hi) * a.hi;
        return {lo, -nh};
    }
    if (a.hi <= 0) {
        volatile double lo = a.hi * a.hi;
        volatile double nh = (-a.lo) * a.lo;
        return {lo, -nh};
    }
    volatile double n1 = (-a.lo) * a.lo, n2 = (-a.hi) * a.hi;
    return {0.0, -std::fmin(n1, n2)};
}

// libm's exp is not directed: widen each end by one ulp and rely on a
// documented sub-ulp error bound of the library
static I exp_(I a)
{
    double lo = std::exp(a.lo), hi = std::exp(a.hi);
    return {std::nextafter(lo, -INFINITY), std::nextafter(hi, INFINITY)};
}

/* On an NVIDIA GPU no mode is set; each operation carries its direction
   (CUDA Math API intrinsics):

   __device__ I add(I a, I b)
   {
       return {__dadd_rd(a.lo, b.lo), __dadd_ru(a.hi, b.hi)};
   }

   __device__ I mul(I a, I b)
   {
       return {fmin(fmin(__dmul_rd(a.lo, b.lo), __dmul_rd(a.lo, b.hi)),
                    fmin(__dmul_rd(a.hi, b.lo), __dmul_rd(a.hi, b.hi))),
               fmax(fmax(__dmul_ru(a.lo, b.lo), __dmul_ru(a.lo, b.hi)),
                    fmax(__dmul_ru(a.hi, b.lo), __dmul_ru(a.hi, b.hi)))};
   }
*/

int main()
{
    RoundDown guard;
    I x{0.0, 1.0}, y{0.0, 1.0};

    // the dependency problem: [-1, 1], not {0}
    I d = sub(x, x);

    // 1/3 enclosed by two adjacent floats
    I third{1.0 / 3.0, -((-1.0) / 3.0)};

    // the natural interval extension of f, node by node
    I v3 = add(x, y);
    I v1 = exp_(y);
    I v4 = sqr(v3);
    I v2 = mul(x, v1);
    I f = sub(v2, v4);

    std::printf("x - x on [0,1] = [%g, %g]\n", d.lo, d.hi);
    std::printf("1/3 in [%.17g, %.17g]\n", third.lo, third.hi);
    std::printf("f = x exp(y) - (x+y)^2 on [0,1]^2: [%.17g, %.17g]\n",
                f.lo, f.hi);
}
x - x on [0,1] = [-1, 1]
1/3 in [0.33333333333333331, 0.33333333333333337]
f = x exp(y) - (x+y)^2 on [0,1]^2: [-4, 2.7182818284590455]

The enclosure of \(1/3\) is two adjacent floats, the best any enclosure can do. The upper end of the natural interval of \(f\) is \(e\) rounded up to the next float, \(2.7182818284590455\), against the nearest float \(2.718281828459045\). The cost is a handful of instructions per operation with no branch except the sign test in the square, and the computation for one box is independent of every other box, so a kernel evaluates the same tape per box with identical control flow. The two things a GPU implementation changes are that the mode switch disappears, replaced by the intrinsics in the comment, and that the library functions need the error bounds the device documentation publishes.

Feasibility-based bound tightening

The forward pass of interval arithmetic computes, from the box, an interval for every node of the expression DAG, including the constraint functions at its sinks. A constraint \(g_i(x) \le 0\) says that the sink's interval may be cut at \(0\), and the cut can be pushed back down the graph: if \(v = a + b\) must lie in \(V\) and \(b\) lies in \(B\), then \(a\) must lie in \(V - B\). Pushing every constraint's interval back to the variables, through the inverse of every operation, and repeating until nothing moves, is feasibility-based bound tightening, FBBT, the nonlinear form of the single-row bound tightening of Proposition 2.5.3. In constraint programming the same procedure is called HC4. In the MINLP solvers it is the propagation step of Couenne and of SCIP.F. Benhamou, F. Goualard, L. Granvilliers and J.-F. Puget, "Revising hull and box consistency", in Logic Programming: Proceedings of the 1999 International Conference (MIT Press, 1999), is HC4; the DAG form is Schichl and Neumaier (2005), cited above, with the propagation and search algorithms of X.-H. Vu, H. Schichl and D. Sam-Haroud, "Interval propagation and search on directed acyclic graphs for numerical constraint solving", Journal of Global Optimization 45 (2009); Belotti, Lee, Liberti, Margot and Wächter (2009), cited above, Section 3.2, and Vigerske and Gleixner (2018), cited above, Section 2, describe it in Couenne and SCIP.

Definition 2.6.9. (a) A state assigns an interval \(X_k\) to every node \(k\) of the DAG, the variables carrying the sides of the current box. (b) A contractor for a set \(S\) of feasible points is a map \(C\) on states with three properties. It is contracting: \(C(X) \subseteq X\) componentwise. It is sound: every point of \(S\) whose node values lie in \(X\) has node values in \(C(X)\). It is monotone: \(X \subseteq X'\) implies \(C(X) \subseteq C(X')\). (c) The forward rule of an operation node \(k\) with children \(c_1, c_2\) replaces \(X_k\) by \(X_k \cap \mathrm{OP}_k(X_{c_1}, X_{c_2})\). (d) The backward rule of \(k\) with respect to its child \(c\) replaces \(X_c\) by \(X_c \cap R\), where \(R\) is an interval containing the exact inverse image \(\{s : \exists\, t \in X_{c'} \text{ with } \mathrm{OP}_k(s, t) \in X_k\}\), the hull of that image when it is not an interval. The node contractor \(C_k\) applies the forward rule and then the backward rules of node \(k\). (e) A sink contractor intersects a constraint's interval with its sides, and the objective's with \((-\infty, z_{\mathrm{inc}} - \varepsilon]\) when an incumbent exists. A state is a fixed point if no contractor changes it.

One node shows the rules at work. Let \(v = a + b\) with \(A = B = [0, 1]\), and let a sink contractor cut the node's interval to \(V = [1.5, 2]\). The forward rule gives \(V = [1.5, 2] \cap [0, 2] = [1.5, 2]\), no change. The backward rule for \(a\) gives \(A = [0, 1] \cap (V - B) = [0, 1] \cap [0.5, 2] = [0.5, 1]\), and the same for \(b\). The point \((a, b) = (0.3, 1)\), which has \(v = 1.3 < 1.5\), is removed, and no point with \(a + b \ge 1.5\) and \(a, b \le 1\) is removed, because such a point has \(a \ge 0.5\) and \(b \ge 0.5\). The inverse rules of the other operations are listed in the table after the algorithm.

One node of FBBT: v = a + b on A = B = [0, 1], with V cut to [1.5, 2]. The backward rule shrinks both A and B to [0.5, 1]: it removes the point (0.3, 1), where v = 1.3 < 1.5, and no point with a + b ≥ 1.5, all of which lie in the new box [0.5, 1] × [0.5, 1].

The node contractors are contractors in the sense of the definition. They contract because every update is an intersection with the old interval. They are sound by Theorem 2.6.4 for the forward rule, and for the backward rule because a feasible point with \(v_k(x) \in X_k\) and \(v_{c'}(x) \in X_{c'}\) has \(v_c(x)\) in the exact inverse image, which \(R\) contains. They are monotone because the exact range, the exact inverse image and their hulls are monotone in all their interval arguments, and intersection with a larger interval gives a larger result. The algorithm is the fair iteration of all of them.

Algorithm 2.6.10  FBBT: forward-backward interval propagation on the
                  expression DAG, to a fixed point

Input   DAG with waves D_0 (variables and constants), D_1, ..., D_H by
        depth (the depth of a node is one more than the largest depth
        of a child); variable intervals X_j = [l_j, u_j]; constraint
        sides [g_i^lo, g_i^up] for each sink, including the cutoff sink
        f <= z_inc - eps when an incumbent exists; round cap R;
        threshold tau
Output  tightened intervals for every node, or INFEASIBLE

 1. repeat up to R rounds:

 2.    forward sweep: for j = 1, ..., H, for all nodes k in D_j
       (in parallel):
          X_k <- X_k meet OP_k(X_children)
                                  [natural extension, outward rounding]

 3.    for each sink i: X_i <- X_i meet [g_i^lo, g_i^up]
       if any interval is empty: return INFEASIBLE

 4.    backward sweep: for j = H, ..., 1, for all k in D_j
       (in parallel), for each child c of k:
          X_c <- X_c meet OP_k^{-1,c}(X_k; X_siblings)
                              [inverse rule, hull of the inverse image]
          (a shared child receives every requirement: intersect
          them all)
       if any interval is empty: return INFEASIBLE

 5.    accept a tightening only if it shrinks an interval by more
       than tau (relative); stop if nothing changed

 6. return the intervals

Invariant
    every feasible x (with f(x) <= z_inc - eps when the cutoff is a
    sink) has v_k(x) in X_k for every node k, after every step
                                               [soundness of each step]

Fixed point
    the limit is the greatest common fixed point of the node
    contractors, the same for every fair schedule (Theorem 2.6.11);
    it may be reached only in the limit (Proposition 2.6.13)
node \(v = \mathrm{OP}(a, b)\)forward \(V = \mathrm{OP}(A, B)\)backward: \(a\) must lie in\(b\) must lie in
\(a + b\)\([\underline{a} + \underline{b},\ \overline{a} + \overline{b}]\)\(V - B\)\(V - A\)
\(a - b\)\([\underline{a} - \overline{b},\ \overline{a} - \underline{b}]\)\(V + B\)\(A - V\)
\(a \cdot b\)hull of the four corner products\(V / B\) (extended division when \(0 \in B\); the hull of the pieces unless \(0 \in V\))\(V / A\)
\(a / b\) (\(0 \notin B\))\(A \cdot [1 / \overline{b},\ 1 / \underline{b}]\)\(V \cdot B\)\(A / V\) (extended)
\(\exp(a)\)\([\exp(\underline{a}),\ \exp(\overline{a})]\)\(\log(V \cap (0, \infty))\) 
\(a^2\)[0, max of the two squares] if \(0 \in A\), else the two squares in order\([-\sqrt{\overline{v}},\ \sqrt{\overline{v}}]\); and \(a \ge \sqrt{\underline{v}}\) if \(A \ge 0\), \(a \le -\sqrt{\underline{v}}\) if \(A \le 0\); empty if \(\overline{v} < 0\) 
\(\varphi\) increasing\([\varphi(\underline{a}),\ \varphi(\overline{a})]\)\(\varphi^{-1}(V \cap \operatorname{range}(\varphi))\) 
The inverse rules of the backward sweep (\(V\) the node's interval, \(A\) and \(B\) the children's, all with outward rounding).

Two rules lose information deliberately. The inverse of the square is a union of two intervals when \(0 \in A\) and \(\underline{v} > 0\), and the rule keeps the hull unless the sign of \(a\) is known. The extended division returns two pieces whose hull is the whole line when \(0\) is interior to \(B\). A solver can keep the pieces as a disjunction and branch on them, which is the search of Vu, Schichl and Sam-Haroud. FBBT keeps the hull and stays a single box. For a linear row the forward and backward rules collapse to one formula, the one Proposition 2.5.3 gave: for \(\ell \le a^\top x \le u\) and \(a_i > 0\),

\[\frac{1}{a_i}\Big(\ell - \sum_{j \ne i} \max\{a_j l_j,\ a_j u_j\}\Big) \;\le\; x_i \;\le\; \frac{1}{a_i}\Big(u - \sum_{j \ne i} \min\{a_j l_j,\ a_j u_j\}\Big).\]

The cost of a round is two traversals of the DAG, linear in its size, with each node reading its children or its parent and siblings. The backward step at a node depends only on the intervals of that node and its neighbours, so a sweep is a sequence of waves each of which is a parallel map. What a round computes does not depend on the order in which the nodes are visited, which is the theorem that licenses that parallelism.

The round loop of Algorithm 2.6.10, each sweep a sequence of waves

  1. repeat up to R rounds  <------------------------------------+
     |                                                           |
  2. forward sweep, waves D_1, D_2, ..., D_H; in a wave, all     |
     nodes at once: X_k <- X_k meet OP_k(X_children)             |
     |                                                           |
  3. at the sinks: X_i <- X_i meet [g_i^lo, g_i^up]              |
     | empty? --yes--> return INFEASIBLE                         |
     |                                                           |
  4. backward sweep, waves D_H, ..., D_2, D_1; in a wave, all    |
     nodes at once, for each child c of k:                       |
     X_c <- X_c meet OP_k^{-1,c}(X_k; X_siblings)                |
     | empty? --yes--> return INFEASIBLE                         |
     |                                                           |
  5. an interval shrank by more than tau (relative)? --yes-------+
     | no, or R rounds done
  6. return the intervals

Theorem 2.6.11 (the FBBT fixed point exists and does not depend on the schedule; Apt, 1999; Belotti, Cafieri, Lee and Liberti, 2010). Let \(C_1, \dots, C_m\) be monotone, sound contractors on states contained in an initial state \(X^0\), each continuous on decreasing chains, meaning \(C_i(\bigcap_k X^k) = \bigcap_k C_i(X^k)\) for nested \(X^k\). Let \((i_k)\) be a fair sequence of indices, in which every index appears infinitely often, and \(X^{k+1} = C_{i_k}(X^k)\). Then \(X^k\) decreases to a state \(X^\infty\) that is the greatest common fixed point of the \(C_i\) contained in \(X^0\). In particular \(X^\infty\) contains the node values of every feasible point of the box, and it does not depend on the fair order. In floating-point arithmetic the lattice of states is finite, so the iteration terminates after finitely many steps at the same state.K. R. Apt, "The essence of constraint propagation", Theoretical Computer Science 221 (1999), gives the lattice-theoretic treatment; P. Belotti, S. Cafieri, J. Lee and L. Liberti, "Feasibility-based bounds tightening via fixed points", in Combinatorial Optimization and Applications (COCOA 2010), Lecture Notes in Computer Science 6508 (Springer, 2010), apply it to FBBT; the node contractors satisfy the chain-continuity hypothesis because the range and the inverse image of a continuous operation over nested compact sets commute with the intersection, and so does the hull of nested compact sets.

Proof sketch. The states are nested and compact, so \(X^\infty = \bigcap_k X^k\) is a state, possibly empty. Fix \(i\) and the infinitely many \(k\) with \(i_k = i\). Then \(X^{k+1} = C_i(X^k) \supseteq C_i(X^\infty)\) by monotonicity, so \(X^\infty \supseteq C_i(X^\infty)\), and chain continuity gives \(C_i(X^\infty) = \bigcap_k C_i(X^k) \supseteq \bigcap_k X^{k+1} = X^\infty\). Hence \(X^\infty\) is a common fixed point. If \(X'\) is any common fixed point in \(X^0\), induction gives \(X' = C_{i_k}(X') \subseteq C_{i_k}(X^k) = X^{k+1}\), so \(X' \subseteq X^\infty\), and \(X^\infty\) is the greatest. Soundness gives the inclusion of the feasible node values at every step. ∎

(What the fixed-point theorem says, and what it does not) The theorem says that FBBT computes a well-defined object, the greatest state closed under all the single-node contractors, whatever the schedule. The sequential sweeps of Algorithm 2.6.10, a schedule in which every node reads the latest intervals, and a Jacobi schedule in which every contractor reads the state at the start of a round and the results are intersected at the end, have the same limit. They do not have the same iterates.

Proposition 2.6.12 (Gauss–Seidel is at least as tight as Jacobi after the same number of rounds). Let \(C_1, \dots, C_m\) be monotone contracting maps on states, let \(J(X) = \bigcap_i C_i(X)\) be the Jacobi round and \(G(X) = C_m(C_{m-1}(\cdots C_1(X)))\) the Gauss–Seidel round. Then \(G(X) \subseteq J(X)\) for every state \(X\), and \(G^k(X) \subseteq J^k(X)\) for every \(k\). Under the hypotheses of Theorem 2.6.11 both sequences decrease to the same limit, so the Jacobi schedule needs at least as many rounds as the Gauss–Seidel schedule to reach any given state.

Proof. Put \(X_0 = X\) and \(X_i = C_i(X_{i-1})\). By contraction \(X_m \subseteq X_i\) for every \(i \le m\), and by monotonicity \(X_i = C_i(X_{i-1}) \subseteq C_i(X)\) because \(X_{i-1} \subseteq X\). Hence \(G(X) = X_m \subseteq \bigcap_i C_i(X) = J(X)\). For the induction, \(J\) is monotone as an intersection of monotone maps, so \(G^{k+1}(X) = G(G^k(X)) \subseteq J(G^k(X)) \subseteq J(J^k(X)) = J^{k+1}(X)\). ∎

Proposition 2.6.12: one round of the contractors C_1, ..., C_m

  Gauss-Seidel: each contractor reads the latest state

    X --> C_1 --> X_1 --> C_2 --> X_2 --> ... --> C_m --> X_m = G(X)

  Jacobi: every contractor reads X, and the results are intersected

           +--> C_1(X) --+
           |             |
    X -----+--> C_2(X) --+--> meet --> J(X)
           |     ...     |
           +--> C_m(X) --+

  G(X) lies in J(X), and G^k(X) in J^k(X); both decrease to the
  same fixed point (Theorem 2.6.11). FBBT rounds from the cutoff
  f <= -1.2 on the dag figure's f, to the tolerances 1e-3, 1e-6
  and 1e-9, in the listing below: Gauss-Seidel 12, 26 and 40,
  Jacobi 31, 90 and 149

The theorem does not say the computation is fast, and in general it is not finite.

Proposition 2.6.13 (the fixed point of linear rows is an LP, and the iteration may never reach it; Belotti, Cafieri, Lee and Liberti, 2010, 2012). For linear rows \(Ax \le b\) on a box, the greatest fixed point of the single-row contractors is the pair \((l, u)\), with \(l\) componentwise smallest and \(u\) componentwise largest, among those that satisfy the original bounds together with every propagation inequality of the display above, written as a constraint on \((l, u)\). It is the solution of one linear program of polynomial size, for instance maximizing \(\sum_j (u_j - l_j)\) subject to those inequalities, once the signs of the coefficients fix which bound enters each term. The iteration itself may approach that point only geometrically. On the two rows \(x \le y/2\) and \(y \le x\) with \(x, y \in [0, 1]\), whose only feasible point is the origin, each round halves both upper bounds. The fixed point \([0, 0]^2\) is therefore reached only in the limit, and in binary64 arithmetic the iteration takes \(1{,}075\) rounds to underflow to zero.Belotti, Cafieri, Lee and Liberti (2010), cited above, and P. Belotti, S. Cafieri, J. Lee and L. Liberti, "On feasibility based bounds tightening", Optimization Online report 3325 (January 2012).

Proposition 2.6.13: on the rows x ≤ y/2 and y ≤ x over [0, 1]², each FBBT round halves both upper bounds, and the fixed point (0, 0) is reached only in the limit: the change per round falls below 10⁻⁶ after 19 rounds, and binary64 underflows to 0 after 1,075 rounds, while one LP in (u_x, u_y) finds (0, 0) in one solve. The slider replays the first rounds.

(The slow example, and the LP that finds its fixed point) The slow example is checked by hand: from \(x \le y/2\) with \(y \le 1\) comes \(x \le 1/2\), then from \(y \le x\) comes \(y \le 1/2\), and the next round gives \(1/4\), \(1/8\), and so on, the change per round falling below \(10^{-6}\) only after 19 rounds. For these two rows only the upper bounds move, so the unknowns of the LP in the proposition are \(u_x\) and \(u_y\), and the LP reads

\[\max\{\, u_x + u_y \;:\; u_x \le u_y / 2,\ \ u_y \le u_x,\ \ 0 \le u_x \le 1,\ \ 0 \le u_y \le 1 \,\},\]

whose optimum is \((0, 0)\): the two row constraints give \(u_x \le u_y / 2 \le u_x / 2\), hence \(u_x \le 0\). The first constraint is the propagation inequality of the row \(x \le y/2\) for the upper bound of \(x\), read with \(y\) at its upper bound, and the second is that of the row \(y \le x\). In general each propagation inequality is linear in \((l, u)\) once the sign of every coefficient has fixed which of \(l_j\) and \(u_j\) enters each term, which is why the greatest fixed point is a linear program and is found in one solve. Solvers therefore cap the rounds and ignore small improvements. Couenne's default cap is 3 rounds, and SCIP's nonlinear constraint handler runs at most 10 propagation rounds per call and accepts a tightening only if it moves a bound by more than \(5\%\) of the smaller of the interval's width and the bound's magnitude. Both solvers apply the LP alternative, in the form of OBBT below, where propagation stalls.Couenne source, src/problem/problem.cpp, MAX_FBBT_ITER (default max_fbbt_iter = 3); SCIP 10.0.0 source, constraints/nonlinear/maxproprounds = 10 and numerics/boundstreps = 0.05 in src/scip/set.c and cons_nonlinear.c.

(The linear case on R1: three boxes) The first figure of this subsection runs the linear case on the running example R1. The polygon is the one of Section 2.1, with rows \(2x + 5y \le 24.5\), \(5x + 2y \le 30.5\), \(-3x + 4y \le 11\) and \(x - 2y \le 4.2\), the given box is \([0, 7] \times [0, 6]\) of area \(42.0\), and the objective at the default angle of \(45^\circ\) is maximized, as the drawn example always is. Three boxes are drawn around the polygon. The orange box is FBBT: each row is read on its own with the other variable at its worst, and after one round the box is \(x \le 6.10\), \(y \le 4.90\), of area \(29.9\). The second round moves nothing, so the figure reports two rounds. The fixed point stops there because the row \(5x + 2y \le 30.5\), read with \(y\) at its lower bound \(0\), allows \(x = 6.1\), while the row \(x - 2y \le 4.2\) would need \(y \ge 0.95\) at that \(x\). The two rows together exclude the point and neither does on its own. The dashed box is the exact projection of the polygon, four small LPs, \(x \le 5.78\) and \(y \le 4.15\), of area \(24.0\). The green box adds the incumbent, in the form of a cutoff row. The cutoff is treated in full under "Duality-based range reduction" below, and here it is only a fifth row. The best integer point at this angle is worth \(4.95\), so the constraint that the objective be at least \(4.95\), which here reads \(x + y \ge 7\), is added before the four LPs. The box collapses to \([3.5, 5.5] \times [1.5, 3.5]\), of area \(4.0\), with \(3\) of the polygon's \(22\) integer points left inside it. Every envelope built on that box would be tighter by the ratio of the areas, before a single branch.

Three boxes around the running polygon: the orange box is feasibility-based bound tightening, each row read on its own with the other variable at its worst, replayed round by round by the slider. The dashed box is the exact projection of the polygon, found by four small LPs, and the green box is the same projection once the objective is required to beat the incumbent, with the region that requirement rules out hatched. The angle slider turns the objective and the switch removes the incumbent.

The block below computes the three boxes. FBBT is the row-by-row propagation of the figure, one round being one Gauss–Seidel pass over the rows. The exact projection enumerates the vertices of the polygon, which is an LP in two variables. The cutoff is the integer optimum at \(45^\circ\), \(7/\sqrt 2\), added as a fifth row.

# Three boxes around the running polygon R1 (the fbbt figure).
#
# FBBT by row-by-row interval propagation, the exact projection by four
# small LPs, and the projection again with the incumbent cutoff added.

import itertools

import numpy as np

# the rows a x + b y <= r: the four of R1, then -x <= 0 and -y <= 0
rows = [(2, 5, 24.5), (5, 2, 30.5), (-3, 4, 11), (1, -2, 4.2),
        (-1, 0, 0), (0, -1, 0)]
# the given box [xl, xu, yl, yu]
box = [0.0, 7.0, 0.0, 6.0]
area = lambda B: (B[1] - B[0]) * (B[3] - B[2])

def fbbt(B, max_rounds=30):
    """Rounds of row-by-row propagation, until a round moves nothing.

    One round takes every row once (Gauss-Seidel), each row read with
    the other variable at its worst. Returns the box before the first
    round and after every round.
    """
    hist = [list(B)]
    for _ in range(max_rounds):
        N = list(hist[-1])
        for a, b, r in rows:
            min_by = b * (N[2] if b >= 0 else N[3])
            min_ax = a * (N[0] if a >= 0 else N[1])
            if a > 0:
                N[1] = min(N[1], (r - min_by) / a)
            if a < 0:
                N[0] = max(N[0], (r - min_by) / a)
            if b > 0:
                N[3] = min(N[3], (r - min_ax) / b)
            if b < 0:
                N[2] = max(N[2], (r - min_ax) / b)
        hist.append(N)
        if max(abs(N[i] - hist[-2][i]) for i in range(4)) < 1e-9:
            break
    return hist

def lp_max(c, cons):
    """Maximize c.p over the polygon by vertex enumeration.

    In two variables every vertex is the meeting point of two rows.
    """
    best = -np.inf
    for (a1, b1, r1), (a2, b2, r2) in itertools.combinations(cons, 2):
        M = np.array([[a1, b1], [a2, b2]], float)
        if abs(np.linalg.det(M)) < 1e-12:
            continue
        p = np.linalg.solve(M, [r1, r2])
        if all(a * p[0] + b * p[1] <= r + 1e-9 for a, b, r in cons):
            best = max(best, c @ p)
    return best

def projection(cons):
    """The smallest box around the polygon: four LPs."""
    return [0.0 - lp_max(np.array([-1.0, 0]), cons),
            lp_max(np.array([1.0, 0]), cons),
            0.0 - lp_max(np.array([0, -1.0]), cons),
            lp_max(np.array([0, 1.0]), cons)]

H = fbbt(box)
for k, B in enumerate(H):
    print(f"FBBT round {k}: x in [{B[0]:.3f}, {B[1]:.3f}], "
          f"y in [{B[2]:.3f}, {B[3]:.3f}], area {area(B):.2f}")
print(f"fixed point after {len(H) - 2} round(s); "
      f"the last round moved nothing")

E = projection(rows)
print("exact projection (4 LPs):")
print(f"    x in [{E[0]:.3f}, {E[1]:.3f}], y in [{E[2]:.3f}, {E[3]:.3f}], "
      f"area {area(E):.2f}")

c = np.array([np.cos(np.pi / 4), np.sin(np.pi / 4)])
lattice = [(x, y) for x in range(8) for y in range(7)
           if all(a * x + b * y <= r for a, b, r in rows)]
# the integer optimum at 45 degrees, 7/sqrt(2)
z_inc = max(c @ np.array(p) for p in lattice)
# add the cutoff row c.(x, y) >= z_inc
C = projection(rows + [(-c[0], -c[1], -z_inc)])
inside = [p for p in lattice
          if C[0] - 1e-9 <= p[0] <= C[1] + 1e-9
          and C[2] - 1e-9 <= p[1] <= C[3] + 1e-9]
print(f"with the incumbent {z_inc:.2f} (c.(x, y) >= {z_inc:.4f}, "
      f"i.e. x + y >= {z_inc * np.sqrt(2):.1f}):")
print(f"    x in [{C[0]:.2f}, {C[1]:.2f}], y in [{C[2]:.2f}, {C[3]:.2f}], "
      f"area {area(C):.2f}")
print(f"integer points of the polygon: {len(lattice)}; "
      f"left inside the last box: {len(inside)}")
print(f"    {inside}")
FBBT round 0: x in [0.000, 7.000], y in [0.000, 6.000], area 42.00
FBBT round 1: x in [0.000, 6.100], y in [0.000, 4.900], area 29.89
FBBT round 2: x in [0.000, 6.100], y in [0.000, 4.900], area 29.89
fixed point after 1 round(s); the last round moved nothing
exact projection (4 LPs):
    x in [0.000, 5.783], y in [0.000, 4.152], area 24.01
with the incumbent 4.95 (c.(x, y) >= 4.9497, i.e. x + y >= 7.0):
    x in [3.50, 5.50], y in [1.50, 3.50], area 4.00
integer points of the polygon: 22; left inside the last box: 3
    [(4, 2), (4, 3), (5, 2)]

Each FBBT round is one pass over the nonzeros of the rows, and the four LPs are each a vertex enumeration of a polygon with six sides. Under a Jacobi schedule the rows of a round are independent of one another, and the four LPs are independent in any case. The three integer points left in the green box are \((4, 2)\), \((4, 3)\) and \((5, 2)\), of which only the last two satisfy the cutoff. A box can only enclose the set of improving points, never coincide with it, and that limitation is shared by every bound-tightening method.

(The nonlinear case: the expression graph of Section 2.5) The second figure runs the nonlinear case on the function that the convexity walk of Section 2.5 labelled unk, \(f(x, y) = x e^{y} - (x + y)^2\), whose expression DAG has seven nodes in four waves: the leaves \(x\) and \(y\), then \(v_3 = x + y\) and \(v_1 = e^{y}\), then \(v_4 = v_3^2\) and \(v_2 = x v_1\), then \(f = v_2 - v_4\). The variable \(x\) is shared by \(v_3\) and \(v_2\), which is where the dependency problem enters. The figure's first pass is the forward sweep on the unit square: \(v_3 \in [0, 2]\), \(v_1 \in [1, 2.718]\), \(v_4 \in [0, 4]\), \(v_2 \in [0, 2.718]\) and \(f \in [-4.000, 2.718]\), of width \(6.718\). The true range, measured on a \(201 \times 201\) grid, is \([-1.282, 0.250]\), of width \(1.532\), so the natural interval is \(4.4\) times as wide. The second pass is FBBT from the cutoff \(f \le c\) with \(c = -1.20\) by default. The root's interval becomes \([-4.000, -1.200]\). The backward sweep asks \(v_4 \ge 1.2\), hence \(v_3 \ge 1.095\), hence \(x \ge 0.095\) and \(y \ge 0.095\) in the first round. The rounds then creep up geometrically and stop after \(12\) rounds, when no end moves by more than \(0.001\), at \(x, y \in [0.204, 1.000]\). The box area falls from \(1.000\) to \(0.633\) and the interval of \(f\) from \([-4.000, 2.718]\) to \([-3.750, -1.200]\). Each round moves the bounds by about \(0.61\) times the previous move, so the loop stops at its tolerance and not at the fixed point, which is why the figure labels its count "rounds to tolerance 0.001". The next proposition gives the exact mechanism.

Proposition 2.6.14 (the figure's fixed point is approached geometrically). On the unit square with the cutoff \(f \le -1.2\), the rounds of Algorithm 2.6.10 move only the lower bounds of \(x\) and \(y\), which remain equal. Their common value after round \(k\) is \(a_k\) with \(a_0 = 0\) and

\[a_{k+1} \;=\; \phi(a_k), \qquad \phi(a) \;=\; \sqrt{1.2 + a\, e^{a}} - 1 .\]

The sequence increases to the unique fixed point \(a^\star = 0.204589\) of \(\phi\) in \([0, 1]\), the root of \((a + 1)^2 = 1.2 + a e^{a}\), and the error contracts asymptotically by the factor \(\phi'(a^\star) = e^{a^\star}/2 = 0.6135\) per round.

Proof sketch. Consider a state in which \(x, y \in [a, 1]\) with \(a \in [0, a^\star]\) and the other intervals those of a forward sweep: \(v_3 \in [a + 1, 2]\) or wider, \(v_1 \in [e^{a}, e]\), \(v_2 \in [a e^{a}, e]\), \(v_4\) with upper end \(4\), and \(f\) with upper end \(c\). The backward sweep at \(f = v_2 - v_4\) requires \(v_4 \ge \underline{v_2} - c = a e^{a} + 1.2\), which at \(v_4 = v_3^2\) with \(v_3 \ge 0\) requires \(v_3 \ge \sqrt{a e^{a} + 1.2}\), which at \(v_3 = x + y\) requires \(x \ge \sqrt{a e^{a} + 1.2} - \overline{y} = \phi(a)\) and the same for \(y\). The requirements at \(v_2 = x v_1\) do not move the ends, because \(\underline{v_2} = \underline{x}\,\underline{v_1}\) already holds, and no upper bound moves in any round. The forward sweep then reproduces a state of the same form with \(a\) replaced by \(\phi(a)\). The function \(h(a) = (a + 1)^2 - 1.2 - a e^{a}\) has \(h(0) = -0.2\), is increasing on \([0, \ln 2)\) and positive at \(1\), so it has exactly one zero \(a^\star\) in \([0, 1]\), and \(\phi(a) > a\) on \([0, a^\star)\). The sequence is increasing and bounded by \(a^\star\), so it converges to the fixed point. At the fixed point \(\sqrt{1.2 + a^\star e^{a^\star}} = a^\star + 1\), whence \(\phi'(a^\star) = e^{a^\star}(1 + a^\star)/(2(a^\star + 1)) = e^{a^\star}/2\). ∎

Proposition 2.6.14: the FBBT rounds on the dag figure's f with the cutoff f ≤ −1.2 iterate a_{k+1} = φ(a_k), φ(a) = √(1.2 + a e^a) − 1, from a₀ = 0 toward the fixed point a* = 0.204589. The ratio of successive moves tends to e^a*/2 = 0.6135 (0.613 at round 12), and the loop stops at its tolerance, 0.001, after 12 rounds, at 0.204, not at a*. The control shows the listing's rounds 1, 2, 3 and 12, after which x, y ≥ 0.095, 0.142, 0.168 and 0.204.
One round from x, y in [a, 1]: the backward sweep runs down

  at f = v2 - v4         at v4 = v3^2                at v3 = x + y
  v4 >= a e^a + 1.2 -->  v3 >= sqrt(a e^a + 1.2) --> x, y >= phi(a)

  round 1, a = 0:
  v4 >= 1.2         -->  v3 >= 1.095             --> x, y >= 0.095

(Why propagation stops far from the hull) The fixed point is also far from the hull of the feasible region. The set \(\{f \le -1.2\}\) inside the unit square is a corner of area \(0.002\) near \((1, 1)\), fitting in \([0.946, 1]^2\), while propagation stops at \([0.2046, 1]^2\), of area \(0.633\). The reason is that every inverse rule reads the other variables at their worst, so the same dependency problem that widened the forward pass widens the backward pass. The figure's third pass evaluates the McCormick relaxation of Algorithm 2.4.15 at a point, by default \((0.5, 0.5)\), through the same graph: the pairs are \(v_3 : [1.000, 1.000]\), \(v_1 : [1.649, 1.859]\), \(v_4 : [1.000, 2.000]\), \(v_2 : [0.500, 1.359]\), and at the root \(f = -0.176\) sits in \([-1.500, 0.359]\), a band of width \(1.859\). On the box \([0.204, 1]^2\) that FBBT left, the band at the same point shrinks to \([-0.892, 0.168]\), of width \(1.060\). Tightening the box tightens every envelope built on it, and these numbers make that sentence quantitative. Three other settings are worth trying. With \(c = -1.00\) nothing moves at all, because \(f(0, 1) = -1\) exactly and the corner \((0, 1)\) of the box is feasible for the cutoff. With \(c = -2.50\) the box is proved empty by arithmetic alone after two rounds. On a \(0.2 \times 0.2\) box with \(c = -3\) the forward pass alone proves the box empty, with no round at all. And the toggle that replaces \(f\) by \(x e^{y} - x e^{y}\), the same node subtracted from itself, returns the natural interval \([-2.718, 2.718]\) for the zero function and, in the McCormick pass, a band \([-0.859, 0.859]\) of width \(1.718\). The convex relaxation of a difference pairs the convex side of one copy with the concave side of the other, so the relaxation inherits the dependency problem.

The expression graph of f(x, y) = x·exp(y) − (x + y)² on a box, with the interval each node carries, and beside it the box itself: the region f ≤ c in blue, the box after bound tightening in orange, and f on a number line. The pass control chooses forward interval propagation, backward propagation (FBBT) iterated to a fixed point, or the McCormick relaxation evaluated at a point, and the step slider replays the chosen pass wave by wave or round by round. The other sliders move c and the upper bounds of the box. In the McCormick pass the point can be dragged in the box or moved along the diagonal with its slider, which also selects that pass.

The block below is the figure's DAG with its forward and inverse rules. It prints the forward pass, the Gauss–Seidel rounds of FBBT from the cutoff, the fixed point of Proposition 2.6.14, and the number of rounds a Jacobi schedule needs for the same tolerances.

# Forward interval propagation and FBBT on an expression DAG.
#
# The DAG is that of f(x, y) = x exp(y) - (x + y)^2 (the dag figure).
# Gauss-Seidel rounds against Jacobi rounds, both converging to the same
# fixed point. numpy only.

import numpy as np

INF = float("inf")

# the DAG in evaluation order: (node, operation, children)
G = [("v3", "add", ("x", "y")),
     ("v1", "exp", ("y",)),
     ("v4", "sqr", ("v3",)),
     ("v2", "mul", ("x", "v1")),
     ("f", "sub", ("v2", "v4"))]

def fwd(k, A, B=None):
    """The natural interval extension of one node."""
    if k == "exp":
        return (np.exp(A[0]), np.exp(A[1]))
    if k == "add":
        return (A[0] + B[0], A[1] + B[1])
    if k == "sub":
        return (A[0] - B[1], A[1] - B[0])
    if k == "mul":
        c = (A[0] * B[0], A[0] * B[1], A[1] * B[0], A[1] * B[1])
        return (min(c), max(c))
    # sqr
    if A[0] >= 0:
        return (A[0] ** 2, A[1] ** 2)
    if A[1] <= 0:
        return (A[1] ** 2, A[0] ** 2)
    return (0.0, max(A[0] ** 2, A[1] ** 2))

def meet(P, Q):
    return (max(P[0], Q[0]), min(P[1], Q[1]))

def empty(P):
    return P[0] > P[1] + 1e-12

def idiv(V, B):
    """The hull of {v / b}.

    Four quotients when 0 is not in B, one-sided when B touches 0.
    """
    if B[0] > 0 or B[1] < 0:
        c = (V[0] / B[0], V[0] / B[1], V[1] / B[0], V[1] / B[1])
        return (min(c), max(c))
    if V[0] <= 0 <= V[1] or (B[0] < 0 < B[1]):
        return (-INF, INF)
    if B[0] == 0:
        return (V[0] / B[1], INF) if V[0] > 0 else (-INF, V[1] / B[1])
    return (-INF, V[0] / B[0]) if V[0] > 0 else (V[1] / B[0], INF)

def back(k, V, A, B=None):
    """The inverse rules: what each child must satisfy, given the
    node's interval V.
    """
    if k == "exp":
        return [(np.log(V[0]) if V[0] > 0 else -INF,
                 np.log(V[1]) if V[1] > 0 else -INF)]
    if k == "add":
        return [(V[0] - B[1], V[1] - B[0]), (V[0] - A[1], V[1] - A[0])]
    if k == "sub":
        return [(V[0] + B[0], V[1] + B[1]), (A[0] - V[1], A[1] - V[0])]
    if k == "mul":
        return [idiv(V, B), idiv(V, A)]
    # sqr with a negative upper end: empty
    if V[1] < 0:
        return [(1.0, -1.0)]
    r, s = np.sqrt(max(0.0, V[1])), np.sqrt(max(0.0, V[0]))
    lo, hi = -r, r
    # the inner gap is cut only when the sign of the child is known
    if A[0] >= 0:
        lo = max(lo, s)
    elif A[1] <= 0:
        hi = min(hi, -s)
    return [(lo, hi)]

def forward(S, c):
    """One forward sweep; the root's upper end is cut at c."""
    for n, k, a in G:
        J = fwd(k, S[a[0]], S[a[1]] if len(a) > 1 else None)
        if n == "f":
            J = (J[0], min(J[1], c))
        S[n] = meet(S[n], J) if n in S else J

def backward(S):
    """One backward sweep, from the root down to the variables."""
    for n, k, a in reversed(G):
        if empty(S[n]):
            return
        reqs = back(k, S[n], S[a[0]], S[a[1]] if len(a) > 1 else None)
        for child, R in zip(a, reqs):
            S[child] = meet(S[child], R)

def move(P, Q):
    """The largest change of any interval end between two states."""
    return max(max(abs(P[n][0] - Q[n][0]), abs(P[n][1] - Q[n][1]))
               for n in P)

def gauss_seidel(c, tol, cap=5000):
    """Gauss-Seidel rounds: every rule reads the latest intervals.

    A forward sweep with the cutoff, then rounds of one backward and one
    forward sweep. Returns the state after the first sweep and after
    every round.
    """
    S = {"x": (0.0, 1.0), "y": (0.0, 1.0)}
    forward(S, c)
    hist = [dict(S)]
    for _ in range(cap):
        P = dict(S)
        backward(S)
        forward(S, c)
        hist.append(dict(S))
        if move(P, S) <= tol:
            break
    return hist

def jacobi(c, tol, cap=5000):
    """Jacobi rounds.

    Every forward and backward rule reads the state at the start of the
    round. Returns the last state and the number of rounds.
    """
    S = {"x": (0.0, 1.0), "y": (0.0, 1.0)}
    forward(S, c)
    rounds = 0
    while rounds < cap:
        N = dict(S)
        for n, k, a in G:
            J = fwd(k, S[a[0]], S[a[1]] if len(a) > 1 else None)
            if n == "f":
                J = (J[0], min(J[1], c))
            N[n] = meet(N[n], J)
            reqs = back(k, S[n], S[a[0]], S[a[1]] if len(a) > 1 else None)
            for child, R in zip(a, reqs):
                N[child] = meet(N[child], R)
        rounds += 1
        m = move(S, N)
        S = N
        if m <= tol:
            break
    return S, rounds

S0 = {"x": (0.0, 1.0), "y": (0.0, 1.0)}
forward(S0, INF)
g = np.linspace(0, 1, 201)
X, Y = np.meshgrid(g, g, indexing="ij")
Fg = X * np.exp(Y) - (X + Y) ** 2

# the forward pass, wave by wave
print("forward pass on [0,1]^2:")
for wave in (("v3", "v1"), ("v4", "v2"), ("f",)):
    print("   ", "  ".join(f"{n} in [{S0[n][0]:.3f}, {S0[n][1]:.3f}]"
                           for n in wave))
print(f"natural width {S0['f'][1] - S0['f'][0]:.3f} against the grid "
      f"range [{Fg.min():.3f}, {Fg.max():.3f}]")
print(f"    of width {Fg.max() - Fg.min():.3f}: ratio "
      f"{(S0['f'][1] - S0['f'][0]) / (Fg.max() - Fg.min()):.1f}")

# Gauss-Seidel rounds of FBBT from the cutoff
c = -1.2
H = gauss_seidel(c, 1e-3)
print()
print("FBBT from            x, y >=  f in                move  ratio")
for r in (1, 2, 3, len(H) - 1):
    mv = move(H[r - 1], H[r])
    ratio = mv / move(H[r - 2], H[r - 1]) if r >= 2 else float("nan")
    print(f"f <= {c}, round {r:2d}{H[r]['x'][0]:9.3f}  "
          f"[{H[r]['f'][0]:.3f}, {H[r]['f'][1]:.3f}]{mv:8.4f}"
          f"{ratio:7.3f}")
E = H[-1]
a_star = gauss_seidel(c, 1e-13)[-1]["x"][0]
print(f"stopped after {len(H) - 1} rounds at tolerance 1e-3; "
      f"box area 1.000 -> "
      f"{(E['x'][1] - E['x'][0]) * (E['y'][1] - E['y'][0]):.3f}")
print(f"fixed point a* = {a_star:.6f}, "
      f"contraction e^a*/2 = {np.exp(a_star) / 2:.4f}")

# the same tolerances under the two schedules
print()
print("rounds to tolerance  Gauss-Seidel  Jacobi  Jacobi lower bound of x")
for tol in (1e-3, 1e-6, 1e-9):
    SJ, rJ = jacobi(c, tol)
    rG = len(gauss_seidel(c, tol)) - 1
    print(f"{tol:<19.0e}{rG:14d}{rJ:8d}{SJ['x'][0]:25.6f}")
forward pass on [0,1]^2:
    v3 in [0.000, 2.000]  v1 in [1.000, 2.718]
    v4 in [0.000, 4.000]  v2 in [0.000, 2.718]
    f in [-4.000, 2.718]
natural width 6.718 against the grid range [-1.282, 0.250]
    of width 1.532: ratio 4.4

FBBT from            x, y >=  f in                move  ratio
f <= -1.2, round  1    0.095  [-3.895, -1.200]  1.2000    nan
f <= -1.2, round  2    0.142  [-3.836, -1.200]  0.1050  0.088
f <= -1.2, round  3    0.168  [-3.801, -1.200]  0.0591  0.563
f <= -1.2, round 12    0.204  [-3.750, -1.200]  0.0006  0.613
stopped after 12 rounds at tolerance 1e-3; box area 1.000 -> 0.633
fixed point a* = 0.204589, contraction e^a*/2 = 0.6135

rounds to tolerance  Gauss-Seidel  Jacobi  Jacobi lower bound of x
1e-03                          12      31                 0.200314
1e-06                          26      90                 0.204585
1e-09                          40     149                 0.204589

A round costs two passes over seven nodes, and the "move" column is the largest change of any interval end in the round, which the first round reports as \(1.2\) because the lower end of \(v_4\) jumps from \(0\) to \(1.2\). The last three lines are Proposition 2.6.12 in numbers: the Jacobi schedule needs \(31\) rounds instead of \(12\) to reach tolerance \(10^{-3}\), and its box at that tolerance is looser (\(0.2003\) against \(0.204\)). At \(10^{-9}\) the ratio is \(149\) to \(40\). Within a round every node of a wave is independent, and across boxes every node of every box is, which is the GPU view taken up at the end of the subsection.

Optimality-based bound tightening

FBBT reads one constraint at a time. Optimality-based bound tightening, OBBT, reads them all at once, by solving, for each variable, two relaxed problems: the least and the greatest value the variable can take over the relaxation's feasible set, with the cutoff row added when an incumbent exists. The range-reduction literature writes the incumbent value as \(U\) and a node's lower bound as \(L\). This series keeps \(z_{\mathrm{inc}}\) for the incumbent and writes a Lagrangian lower bound as \(\underline z\), so that \(L\) remains the Lagrangian of Section 2.2.

Proposition 2.6.15 (what OBBT computes). Let \(R_N\) be the feasible set of the relaxation at node \(N\), \(\bar f\) its objective and \(z_{\mathrm{inc}}\) the incumbent value. The values \(\min\{x_j : x \in R_N,\ \bar f(x) \le z_{\mathrm{inc}}\}\) and \(\max\{x_j : x \in R_N,\ \bar f(x) \le z_{\mathrm{inc}}\}\) are exactly the sides of the smallest box containing \(R_N \cap \{\bar f \le z_{\mathrm{inc}}\}\). OBBT is therefore the tightest box-valued consequence of the relaxation and the cutoff: exact for the relaxation, not for the original feasible set, and in general incomparable with FBBT applied to the original nonlinear constraints, which can be strictly tighter.

Proof. The first sentence is the definition of a bounding box. For the last claim an example suffices. At the root of the bilinear example R2, maximize \(xy\) subject to \(2x + y \le 1.2\) on the unit square, the root relaxation's point \((0.4, 0.4)\) has true value \(0.16\), which becomes the incumbent, so the cutoff is \(xy \ge 0.16\). FBBT on the exact constraints alternates \(x \le (1.2 - \underline{y})/2\), \(y \le 1.2 - 2\underline{x}\) from the line and \(x \ge 0.16/\overline{y}\), \(y \ge 0.16/\overline{x}\) from the hyperbola. After one round \(x \in [0.16, 0.6]\) and \(y \in [0.2667, 1]\), after two \(x \in [0.1818, 0.4667]\) and \(y \in [0.3429, 0.88]\), after ten \(x \in [0.1999, 0.4002]\) and \(y \in [0.3998, 0.8003]\), converging to \(x \in [0.2, 0.4]\), \(y \in [0.4, 0.8]\). This is the bounding box of the exact set, since the hyperbola \(xy = 0.16\) meets the line at \((0.2, 0.8)\) and \((0.4, 0.4)\). OBBT on the McCormick relaxation of Theorem 2.4.7 with the cutoff \(w \ge 0.16\) sees only \(w \le x\), \(w \le y\) and the line, and returns \(x \in [0.16, 0.52]\), \(y \in [0.16, 0.88]\) from four LPs. ∎

Proposition 2.6.15 on R2: FBBT rounds against OBBT at the root. Maximize xy subject to 2x + y ≤ 1.2 on the unit square, with the cutoff xy ≥ 0.16 from the incumbent (0.4, 0.4). FBBT on the exact constraints gives [0.16, 0.6] × [0.2667, 1] after one round, [0.1818, 0.4667] × [0.3429, 0.88] after two and [0.1999, 0.4002] × [0.3998, 0.8003] after ten, converging to [0.2, 0.4] × [0.4, 0.8], the bounding box of the exact set, since the line meets xy = 0.16 at (0.2, 0.8) and (0.4, 0.4). OBBT on the McCormick relaxation returns [0.16, 0.52] × [0.16, 0.88].

(Neither FBBT nor OBBT dominates) After OBBT the McCormick band at the centre of the box is \(0.36 \times 0.72 / 2 \approx 0.13\). After FBBT's fixed point it is \(0.2 \times 0.4 / 2 = 0.04\): three times tighter without a branch. The example reverses the usual relation, in which OBBT is tighter because the relaxation sees all constraints at once while FBBT sees them one at a time. It reverses it because the single nonlinear constraint propagates exactly here. In general neither dominates, which is why solvers run both, and why the sentence "OBBT is exact" needs its qualifier: exact for the relaxation, with the cutoff, and no more.Belotti, Lee, Liberti, Margot and Wächter (2009), cited above, Section 4.2, write that FBBT "provides in general weaker bounds than the ones obtained from OBBT"; Puranik and Sahinidis (2017), cited above, problem (12), state OBBT as the optimization over the relaxation.

(What OBBT costs, and three ways to pay less) OBBT is expensive: up to \(2n\) LPs per node, each the size of the node's relaxation. Two terms from LP practice recur here. The reduced cost of a variable (Proposition 2.2.4) is the multiplier of its bound constraint in an LP, the vector \(r = c - A^\top y\) of (2.2.1), and it is the rate at which the LP value changes when that bound is moved. Re-solving an LP after a small change, starting from the previous optimal solution, is a warm-start (Section 1.3). The simplex method does it in a few steps, called pivots, as Section 3.1 describes (Section 7.1 gives the history). The dual simplex method is the variant that stays dual feasible while a bound changes or a row is added. The OBBT LPs differ from one another only in their objective, so it is the primal side that stays feasible: the previous optimal basis is primal feasible for the next LP and the primal simplex continues from it. The engineering that makes OBBT affordable is Gleixner, Berthold, Müller and Weltge's, and it has three parts: filtering candidates that cannot improve, ordering the LPs so that the simplex warm-start is cheap, and learning from each OBBT LP an inequality that stays valid in the whole subtree, a Lagrangian variable bound (LVB).A. M. Gleixner, T. Berthold, B. Müller and S. Weltge, "Three enhancements for optimization-based bound tightening", Journal of Global Optimization 67 (2017); the Lagrangian variable bounds were introduced in A. M. Gleixner and S. Weltge, "Learning and propagating Lagrangian variable bounds for mixed-integer nonlinear programming", CPAIOR 2013, Lecture Notes in Computer Science 7874 (Springer, 2013).

Algorithm 2.6.16  OBBT at a node, with filtering, ordering and
                  Lagrangian variable bounds
                  (Gleixner, Berthold, Müller and Weltge 2017;
                  SCIP's prop_obbt)

Input   LP relaxation rows D x <= d on the box [l, u]; cutoff row
        c^T x <= z_inc; candidate set K of variables that enter
        nonconvex terms; an LP iteration budget
Output  tightened bounds; a set of Lagrangian variable bounds
        (Theorem 2.6.17)

 1. candidates: for each k in K both "min x_k" and "max x_k", unless
    the bound is infinite or k is fixed

 2. trivial filtering: for any LP solution x~ at hand (the root LP,
    the OBBT LPs as they are solved), if x~_k = l_k then "min x_k"
    cannot improve: drop it; likewise for u_k

 3. aggressive filtering (optional): choose a direction v with
    v_k > 0 for pending "max x_k" and v_k < 0 for pending "min x_k";
    solve max v^T x over the relaxation; drop every candidate whose
    variable sits at its targeted bound; repeat while enough
    candidates were dropped

 4. ordering: process the remaining candidates in a greedy order that
    keeps consecutive LP objectives close, so that the simplex
    warm-start from one LP to the next is a few pivots

 5. for each candidate (k, direction) in that order:

 6.    solve the LP; if the bound improved by more than a threshold:
       set it

 7.    from the dual solution form the Lagrangian variable bound of
       Theorem 2.6.17 and keep it if it is nontrivial

 8.    apply trivial filtering with the new LP solution

 9. return the bounds and the LVBs; in the tree, propagate the LVBs
    like linear rows whenever a bound or the incumbent changes, in an
    order given by the strongly connected components of their
    dependency graph

Invariant
    every x in the relaxation with c^T x <= z_inc satisfies the new
    bounds (Proposition 2.6.15) and every LVB

Theorem 2.6.17 (Lagrangian variable bounds; Gleixner, Berthold, Müller and Weltge, 2017). Let the relaxation at the root be \(R = \{x : Dx \le d,\ x \in [l, u]\}\) and let \(z_{\mathrm{inc}}\) be the incumbent value, a valid upper bound on the optimal value. Suppose \((\tilde x, \tilde\lambda, \tilde\mu)\) is an optimal primal–dual solution of the OBBT problem \(\max\{x_k : Dx \le d,\ c^\top x \le z_{\mathrm{inc}},\ x \in [l, u]\}\), with \(\tilde\lambda \ge 0\) the multipliers of \(Dx \le d\), \(\tilde\mu \ge 0\) the multiplier of the cutoff row, and \(\tilde r = e_k - D^\top \tilde\lambda - \tilde\mu\, c\) the reduced costs. Then

\[x_k \;\le\; \tilde r^\top x + \tilde\lambda^\top d + \tilde\mu\, z_{\mathrm{inc}} \tag{2.6.1}\]

is valid for every \(x \in R\) with \(c^\top x \le z_{\mathrm{inc}}\), and it is tight at \(\tilde x\). Replacing each \(x_j\) on the right by \(u_j\) if \(\tilde r_j > 0\) and by \(l_j\) if \(\tilde r_j < 0\) gives a valid upper bound on \(x_k\) that equals the OBBT bound when the duals are optimal. The minimization case is symmetric.

Proof. Aggregating the rows \(Dx \le d\) with weights \(\tilde\lambda \ge 0\) and the cutoff row with weight \(\tilde\mu \ge 0\) gives \((D^\top \tilde\lambda + \tilde\mu c)^\top x \le \tilde\lambda^\top d + \tilde\mu z_{\mathrm{inc}}\) for every \(x\) in question. Substituting \(D^\top \tilde\lambda + \tilde\mu c = e_k - \tilde r\) yields (2.6.1). Tightness at \(\tilde x\) is complementary slackness. Bounding \(\tilde r^\top x\) over the box term by term gives the last claim. ∎

Theorem 2.6.17: one root OBBT LP, one bound for the whole tree

  root:  max x_k  s.t.  D x <= d,  c^T x <= z_inc,  x in [l, u]
         optimal duals lambda~ >= 0 and mu~ >= 0, reduced costs
         r~ = e_k - D^T lambda~ - mu~ c
                               |
                               v
         x_k <= r~^T x + lambda~^T d + mu~ z_inc            (2.6.1)
         valid for every x with D x <= d and c^T x <= z_inc,
         tight at x~
                               |
          +--------------------+--------------------+
          v                                         v
  a node below the root:                  the incumbent improves:
  bound r~^T x over the node's box,       the right-hand side
  x_j at u_j where r~_j > 0 and at        improves with z_inc
  l_j where r~_j < 0: the bound
  improves as the box tightens

  each update is a sparse inner product, not an LP

(Why a Lagrangian variable bound survives in the tree) The point of an LVB is that it stays valid at every descendant node, because \(Dx \le d\) remains valid, while its right-hand side improves automatically when bounds tighten or the incumbent improves, and propagating it costs a sparse inner product instead of an LP. The paper's measurements are the reason SCIP runs OBBT only at the root and relies on LVBs in the tree. Only about one OBBT LP in seven tightened a bound, but each run learned two to three times as many nontrivial LVBs as it tightened bounds, and those LVBs are what the tree then propagates. The table gives the figures.Gleixner, Berthold, Müller and Weltge (2017), cited above; on the single instance elec200 the root OBBT time fell from 714.3 s to 62.6 s with aggressive filtering and greedy ordering. SCIP 10.0.0 defaults: OBBT at the root only, an iteration budget of ten times the root LP's iterations and at least \(5{,}000\), trivial filtering on, aggressive filtering off, greedy ordering, propagating/obbt/creategenvbounds = TRUE, with the genvbounds propagator applying the LVBs in the tree (src/scip/prop_obbt.c, prop_genvbounds.c).

measurementset INTset GO
OBBT LPs that tightened a bound15.6%13.0%
nontrivial LVBs learned, per tightened bound2.92.3
root OBBT time, aggressive filtering with greedy ordering–−16%
average speed-up on hard instances, OBBT with LVB propagation17% to 19% 
instances: all but 3 of those plain SCIP solved; 16 more on which plain SCIP timed out  
Root OBBT with Lagrangian variable bounds (Gleixner, Berthold, Müller and Weltge 2017), two MINLPLib test sets.

Each OBBT LP shares the constraint matrix with every other and differs only in its objective, so a round of OBBT is a batch of LPs with one matrix. On a CPU the batch is made cheap by warm-starts in a good order. On a GPU it is a many-right-hand-side workload for a first-order method, with the LVBs as the cheap product that the tree then propagates.

Duality-based range reduction, the incumbent and probing

The LP that bounds a node has a dual, and the dual carries information that OBBT would otherwise buy with another LP: the sensitivity of the bound to each variable's bounds. Combined with the incumbent, that sensitivity tightens the box for free. The statement is Ryoo and Sahinidis's, generalized by Tawarmalani and Sahinidis to any Lagrangian bound.Ryoo and Sahinidis (1996), cited above, Theorem 2 and Corollary 4; Tawarmalani and Sahinidis (2004), cited above, Theorem 4.1, which derives the known reduction tests as corollaries of a general range-reduction problem. In MILP the same device is reduced-cost fixing: H. Crowder, E. L. Johnson and M. Padberg, "Solving large-scale zero-one linear programming problems", Operations Research 31 (1983).

Theorem 2.6.18 (duality-based range reduction; Ryoo and Sahinidis, 1996; Tawarmalani and Sahinidis, 2004). Let node \(N\) have the convex relaxation \(\min\{\bar f(x) : \bar g_i(x) \le 0,\ i = 1, \dots, m,\ l \le x \le u\}\) with \(\bar f \le f\) and \(\bar g_i \le g_i\) on \(B_N\). Let \(\lambda \in \mathbb{R}^m_{\ge 0}\), \(\mu^+, \mu^- \in \mathbb{R}^n_{\ge 0}\) and \(\underline z \in \mathbb{R}\) satisfy the Lagrangian lower bound

\[\bar f(x) + \sum_i \lambda_i \bar g_i(x) + \sum_j \mu^+_j (x_j - u_j) + \sum_j \mu^-_j (l_j - x_j) \;\ge\; \underline z \qquad \text{for all } x \in B_N ,\]

and let \(z_{\mathrm{inc}} \ge z^\star\) be the incumbent value. Then every feasible \(x \in B_N\) with \(f(x) \le z_{\mathrm{inc}}\) satisfies

\[\bar g_i(x) \;\ge\; -\frac{z_{\mathrm{inc}} - \underline z}{\lambda_i}\quad (\lambda_i > 0), \qquad x_j \;\ge\; u_j - \frac{z_{\mathrm{inc}} - \underline z}{\mu^+_j}\quad (\mu^+_j > 0), \qquad x_j \;\le\; l_j + \frac{z_{\mathrm{inc}} - \underline z}{\mu^-_j}\quad (\mu^-_j > 0).\]

Proof. Fix such an \(x\). Every term on the left of the Lagrangian inequality other than the one singled out is nonpositive: \(\lambda_k \bar g_k(x) \le \lambda_k g_k(x) \le 0\), \(\mu^+_k (x_k - u_k) \le 0\) and \(\mu^-_k (l_k - x_k) \le 0\). Dropping them can only decrease the left side, so \(\underline z \le \bar f(x) + \lambda_i \bar g_i(x) \le f(x) + \lambda_i \bar g_i(x) \le z_{\mathrm{inc}} + \lambda_i \bar g_i(x)\), which is the first claim. The other two are the same argument with the bound terms. ∎

(The picture: a tangent meets the incumbent) In a picture, the relaxation's optimal value is a convex function of each bound, its slope at the current bound is the multiplier, and the incumbent is a horizontal line. The tangent of slope \(\mu\) crosses the line at distance \((z_{\mathrm{inc}} - \underline z)/\mu\) from the bound, and no improving solution can lie beyond the crossing. When the relaxation is an LP solved to optimality, \(\underline z\) is its value \(\bar z(N)\) and \(\mu^\pm\) are the reduced costs \(r_j\) of the variables at their bounds, which is why MILP solvers call the step reduced-cost fixing. Two remarks matter for the GPU programme. The multipliers need not be optimal: any dual-feasible \((\lambda, \mu)\) together with its dual objective \(\underline z\) satisfies the hypothesis by weak duality, so inexact duals from a first-order LP solver are usable. The condition is that \(\underline z\) is computed from them and not from the primal objective, which is the safe bound of Section 7.3. And the step is empty without an incumbent: the width \((z_{\mathrm{inc}} - \underline z)/\mu\) is infinite when \(z_{\mathrm{inc}}\) is, which is one reason global solvers run local search before the tree begins.BARON User Manual, cited above, Section 5.3, gives this as the reason for front-loading local search; the manual calls the step marginals-based reduction (option MDo).

(Two numerical instances) Two numerical instances follow, in the maximization convention of the drawn examples. For a maximization node with relaxation value \(\bar z(N)\), incumbent \(z_{\mathrm{inc}}\) and a variable that the LP solution leaves at its lower bound \(l_j\) with reduced cost \(r_j < 0\), the theorem reads

\[x_j \;\le\; l_j + \frac{\bar z(N) - z_{\mathrm{inc}}}{|r_j|}, \qquad\text{and, for a variable at its upper bound } u_j \text{ with } r_j > 0, \qquad x_j \;\ge\; u_j - \frac{\bar z(N) - z_{\mathrm{inc}}}{r_j} .\]

Here the roles of the two bounds are exchanged, and \(z_{\mathrm{inc}} - \underline z\) becomes \(\bar z(N) - z_{\mathrm{inc}}\), the distance from the node's bound down to the incumbent. First, a schematic instance: the relaxation has value \(100\) and the incumbent is \(95\). An integer variable that the LP solution leaves at its lower bound \(0\), with reduced cost \(-6\), lowers the bound by \(6\) per unit it is raised, so every improving solution has \(x \le 5/6\), and the variable is fixed at \(0\). A variable with reduced cost \(-2\) and upper bound \(10\) gets \(x \le 5/2\), hence \(x \le 2\). Second, a schematic node of R2: suppose the relaxation at some node has value \(0.2333\), the incumbent is \(0.18\), and \(y\) sits at its upper bound \(1\) with reduced cost \(0.3\). Then every point that could still improve the incumbent has \(y \ge 1 - 0.0533/0.3 = 0.822\). In the actual trace of Section 3.5 the rule never fires, because the relaxed point is interior in the branched coordinate at every surviving node, so the multipliers of the bound rows are zero. The same reasoning with a temporary fixing is probing.

Reduced-cost tightening on the schematic instance (maximization): the node's bound 100 falls by the reduced cost per unit of x_j and meets the incumbent 95 at x_j = 5/6 for reduced cost −6 and at x_j = 5/2 for reduced cost −2; beyond a crossing no solution improves on the incumbent. The first integer variable is fixed at 0; the second, with upper bound 10, gets x_j ≤ 2. The slider moves the reduced cost, and the sentence above the panel gives the crossing and the integer bound there.

Corollary 2.6.19 (probing; Ryoo and Sahinidis, 1996). Fix \(x_j\) temporarily at \(u_j\), solve the relaxation of \(N\) with the extra constraint \(x_j \ge u_j\), and let \(\underline z^+_j\) be its value and \(\mu'_j > 0\) the multiplier of that constraint. Then every feasible \(x \in B_N\) with \(f(x) \le z_{\mathrm{inc}}\) satisfies

\[x_j \;\le\; u_j - \frac{\underline z^+_j - z_{\mathrm{inc}}}{\mu'_j} .\]

The bound is vacuous when \(\underline z^+_j \le z_{\mathrm{inc}}\). When \(\underline z^+_j > z_{\mathrm{inc}}\) it excludes a slab below \(u_j\), so in particular no improving point has \(x_j = u_j\), and for an integer variable the value \(u_j\) itself is excluded whenever the slab has positive width. The symmetric probe at \(l_j\), with value \(\underline z^-_j\) and multiplier \(\mu'_j\) of the constraint \(x_j \le l_j\), gives \(x_j \ge l_j + (\underline z^-_j - z_{\mathrm{inc}})/\mu'_j\). For a binary variable, a probe whose value exceeds \(z_{\mathrm{inc}}\) fixes the variable at the other value.

Proof. Let \(\varphi(t)\) be the optimal value of the relaxation with \(x_j \ge u_j - t\). Enlarging the feasible set lowers the value, so \(\varphi\) is nonincreasing, and as the value of a convex program under a perturbation of a right-hand side it is convex in \(t\), with \(\varphi(0) = \underline z^+_j\) and \(-\mu'_j\) a subgradient at \(0\). Hence \(\varphi(t) \ge \underline z^+_j - \mu'_j t\). An improving feasible \(x\) with \(t = u_j - x_j \ge 0\) is feasible for that relaxation, so \(\varphi(t) \le \bar f(x) \le f(x) \le z_{\mathrm{inc}}\). Together, \(\underline z^+_j - \mu'_j t \le z_{\mathrm{inc}}\), that is \(t \ge (\underline z^+_j - z_{\mathrm{inc}})/\mu'_j\), which is the claim. The same inequality follows from Theorem 2.6.18 applied to the probing relaxation, whose box has lower bound \(u_j\) in coordinate \(j\) with multiplier \(\mu^-_j = \mu'_j\). ∎

Algorithm 2.6.20  Duality-based range reduction at a node
                  (Theorem 2.6.18 and Corollary 2.6.19)

Input   the node LP's value zlow = zbar(N) (or the dual objective of
        any dual-feasible pair), the reduced costs r_j >= 0 (in
        magnitude) of the variables at their bounds, the incumbent
        z_inc, the box [l, u]
Output  a tightened box

 1. for each j with xbar_j = l_j and r_j > 0:
       u_j <- min(u_j, l_j + (z_inc - zlow) / r_j)

 2. for each j with xbar_j = u_j and r_j > 0:
       l_j <- max(l_j, u_j - (z_inc - zlow) / r_j)

 3. probing (optional, one LP per probe): for selected j, fix
    x_j = u_j and re-solve, obtaining the value zlow_j^+ and the
    multiplier mu'_j of the fixing;
    if zlow_j^+ > z_inc:
       u_j <- u_j - (zlow_j^+ - z_inc) / mu'_j     (Corollary 2.6.19);
    symmetrically at l_j with zlow_j^- :
       l_j <- l_j + (zlow_j^- - z_inc) / mu'_j

 4. for integer variables round the new bounds inward

Invariant
    no feasible x with f(x) <= z_inc is removed
                                    [Theorem 2.6.18, Corollary 2.6.19]

Steps 1 and 2 are free, since the duals are already there. Step 3 costs one LP per probe, and probing is the expensive step of BARON's node. Belotti and co-authors measured BARON 7.5 spending between \(41\%\) and \(99\%\) of its time in probing on the instances that needed more than ten seconds, with a geometric average of \(77\%\), in exchange for about five times fewer nodes than Couenne. Couenne's cheaper "aggressive FBBT" runs propagation rather than an LP on each half of a variable's range, and a later paper learns, from instance features, when a probe is likely to succeed.Belotti, Lee, Liberti, Margot and Wächter (2009), cited above, Sections 4.3 and 7.3; G. Nannicini, P. Belotti, J. Lee, J. Linderoth, F. Margot and A. Wächter, "A probing algorithm for MINLP with failure prediction by SVM", CPAIOR 2011, Lecture Notes in Computer Science 6697 (Springer, 2011).

(One principle behind every reduction with a cutoff) The common thread of OBBT with a cutoff row, FBBT on the cutoff constraint, reduced-cost tightening and probing against the incumbent is one principle. BARON's manual states it in one sentence: "Once a feasible solution with objective value \(U\) is known, any part of the search space that cannot contain a better solution than \(U\) may be discarded. BARON uses this principle not only to fathom entire nodes, but also to shrink the box at nodes that survive. This is the 'reduce' step in branch-and-reduce."BARON User Manual, cited above, Section 5.3. On the bilinear example the effect is drastic. The spatial branch and bound of Section 3.5, with midpoint branching and no tightening, processes \(13\) nodes, one LP each, to prove \(0.18\) optimal at tolerance \(0.01\), \(21\) at \(0.001\) and \(27\) at \(0.0001\). With FBBT on the constraints and the cutoff \(xy \ge z_{\mathrm{inc}}\) before every LP it solves only \(2\), \(4\) and \(4\) LPs, because after the root the child \(x \ge 0.5\) is proved empty by propagation before an LP is solved.Computed for this post with the spatial branch-and-bound script behind the figure of Section 3.5 (sbb_variants.py, numpy only, re-run 5 October 2026), which counts LPs solved: 13, 21 and 27 for the midpoint rule without tightening, matching the node counts of Section 3.5, and 2, 4 and 4 with FBBT on \(2x + y \le 1.2\) and \(xy \ge z_{\mathrm{inc}}\) before each LP. Section 3.5 does not report the second set of counts. Most of the tree is spent on the size of the box, not on the choice of branching rule.

Certificates: interval Newton and the Krawczyk test

Interval arithmetic adds one more kind of statement to the toolbox, and it is a statement no relaxation can make: that a box contains exactly one zero of a system of equations, or none. For a MINLP with the integer variables fixed, the system is the gradient or the KKT system of the remaining NLP, and the test turns a local solver's approximate stationary point into a verified one, or deletes a box that has none. The one-dimensional case shows the whole mechanism.

Theorem 2.6.21 (univariate interval Newton; Moore, 1966; extended form after Hansen, 1978). Let \(g\) be continuously differentiable on an open interval containing \(X = [a, b]\), let \(G'(X)\) be an interval containing \(g'(\xi)\) for every \(\xi \in X\), let \(m \in X\), and define

\[N(X) \;=\; m - \frac{g(m)}{G'(X)},\]

with the extended division when \(0 \in G'(X)\), so that \(N(X)\) is an interval or the union of two unbounded intervals. Then (i) every zero of \(g\) in \(X\) lies in \(N(X) \cap X\). (ii) If \(N(X) \cap X = \emptyset\), \(g\) has no zero in \(X\). (iii) If \(0 \notin G'(X)\) and \(N(X) \subseteq X\), \(g\) has exactly one zero in \(X\). (iv) If \(0 \notin G'(X)\) and \(g(m) \ne 0\), then \(N(X)\) lies strictly on one side of \(m\), so with \(m = \operatorname{mid} X\) the iteration \(X_{k+1} = N(X_k) \cap X_k\) either proves that \(X_0\) has no zero or produces nested intervals, of width at most halved at each step, shrinking to the unique zero in \(X_0\).Moore (1966), cited above; Hansen (1978), cited above, for the extended form with Kahan's division; the convergence is in fact quadratic once the interval is small, under a Lipschitz condition on \(G'\) (Moore, Kearfott and Cloud (2009), cited above; proof there).

Proof. (i) Let \(g(x^\star) = 0\) with \(x^\star \in X\). By the mean value theorem \(0 = g(m) + g'(\xi)(x^\star - m)\) for some \(\xi\) between \(m\) and \(x^\star\), so \(g'(\xi) \in G'(X)\). If \(g'(\xi) \ne 0\) then \(x^\star = m - g(m)/g'(\xi) \in N(X)\), because the extended division contains every quotient by a nonzero element of \(G'(X)\). If \(g'(\xi) = 0\) then \(g(m) = 0\), so \(0 \in G'(X)\) and the extended quotient \(0/G'(X)\) is the whole line. (ii) is the contrapositive of (i). (iii) Since \(0 \notin G'(X)\), \(g\) is strictly monotone on \(X\) and has at most one zero. Replacing \(g\) by \(-g\) if necessary, assume \(G'(X) = [d_1, d_2]\) with \(d_1 > 0\), and suppose \(g\) had no zero, so that it has constant sign on \(X\). If \(g > 0\), then \(N(X) = [m - g(m)/d_1,\ m - g(m)/d_2] \subseteq X\) gives \(g(m) \le d_1 (m - a)\), while the mean value theorem gives \(g(m) - g(a) \ge d_1 (m - a)\), so \(g(a) \le 0\), a contradiction. The case \(g < 0\) is symmetric at \(b\). (iv) With \(d_1 > 0\) and \(g(m) > 0\) every element of \(N(X)\) is \(m - g(m)/d\) with \(d > 0\), hence less than \(m\), and with \(g(m) < 0\) every element exceeds \(m\). So \(N(X) \cap X\) lies in \([a, m)\) or \((m, b]\), of width at most \(w(X)/2\) when \(m\) is the midpoint. The iterates are nested, each contains every zero of \(X_0\) by (i), and an empty intersection means there was none. ∎

(The Newton test on the figure's g) The picture is that \(N(X)\) is the set of Newton steps from \(m\) with every admissible slope, and the mean value theorem says the true zero is reached by one of those slopes. If all of them land outside \(X\) there is no zero, and if all of them land inside, the monotone function must cross zero in \(X\). On the function \(g(y) = e^{y} - 2y - 2\), which is the equation the stationary points of the figure's \(f\) satisfy once \(x = 1\) is known, the test on \([1.5, 2]\) gives \(N = [1.647408, 1.702756]\), inside the interval, so exactly one zero. On \([-1, -0.5]\) it gives \([-0.769831, -0.766931]\), again exactly one. On \([0, 0.5]\) it gives \([-3.21, -0.97]\), disjoint, so none. On \([-1, 2]\), where \(0 \in G'\), the extended division returns two unbounded pieces whose intersection with the interval removes the gap \((-0.328, 0.751)\) in one step and keeps one piece around each zero. The widths of the iteration from \([1.5, 2]\) fall as \(5.5 \cdot 10^{-2}\), \(2.9 \cdot 10^{-4}\), \(6.0 \cdot 10^{-9}\) and \(2.0 \cdot 10^{-15}\), each a constant times the square of the one before, as quadratic convergence predicts, until the floor of binary64 is reached. In several variables the division by an interval matrix has no direct analogue, and Krawczyk's operator replaces it by multiplication with an approximate inverse.

Interval Newton on g(y) = e^y − 2y − 2, with zeros −0.768039 and 1.678347. On X = [1.5, 2] the Newton image is N(X) = [1.647408, 1.702756], and on [−1, −0.5] it is [−0.769831, −0.766931], both inside X, so each holds exactly one zero; on [0, 0.5] it is [−3.21, −0.97], disjoint from X, so none; on [−1, 2], where 0 ∈ G′(X), the extended division returns two unbounded pieces whose intersection with X removes the gap (−0.328, 0.751) in one step and keeps one piece around each zero.

Theorem 2.6.22 (the Krawczyk test; Krawczyk, 1969; Moore, 1977; Neumaier, 1990). Let \(f : D \to \mathbb{R}^n\) be continuously differentiable on an open set containing the box \(X\), let \(\tilde x \in X\), let \(Y\) be any real \(n \times n\) matrix, and let \(F'(X)\) be an interval matrix with \(\partial_j f_i(\xi) \in F'(X)_{ij}\) for all \(\xi \in X\). Define

\[K(X) \;=\; \tilde x - Y f(\tilde x) + \big(I - Y F'(X)\big)\,(X - \tilde x).\]

Then (i) every zero of \(f\) in \(X\) lies in \(K(X) \cap X\). (ii) If \(K(X) \cap X = \emptyset\), \(f\) has no zero in \(X\). (iii) If \(K(X) \subseteq X\) and \(Y\) is nonsingular, \(f\) has at least one zero in \(X\). (iv) If \(K(X) \subseteq \operatorname{int} X\), then \(Y\) and every real matrix in \(F'(X)\) are nonsingular and \(f\) has exactly one zero in \(X\). (v) In the situation of (iv), with \(Y\) and \(F'(X_0)\) kept fixed and \(\tilde x_k = \operatorname{mid} X_k\), the iteration \(X_{k+1} = K(X_k) \cap X_k\) produces nested boxes containing the zero whose widths tend to zero.R. Krawczyk, "Newton-Algorithmen zur Bestimmung von Nullstellen mit Fehlerschranken", Computing 4 (1969); R. E. Moore, "A test for existence of solutions to nonlinear systems", SIAM Journal on Numerical Analysis 14 (1977); Neumaier (1990), cited above, where (iv) and (v) are proved and the re-evaluation of \(Y\) and \(F'\) on each box is shown to give quadratic convergence under a Lipschitz condition. E. Hansen and S. Sengupta, "Bounding solutions of systems of equations using interval analysis", BIT 21 (1981), replace the operator by an interval Gauss–Seidel step, usually tighter.

Proof of (i) to (iii), and a sketch of (iv) and (v). Let \(f(x^\star) = 0\) with \(x^\star \in X\). Apply the mean value theorem to each component along the segment from \(\tilde x\) to \(x^\star\), which lies in \(X\): \(f_i(x^\star) - f_i(\tilde x) = \nabla f_i(\xi_i)^\top (x^\star - \tilde x)\) with \(\xi_i \in X\). The real matrix \(J\) with rows \(\nabla f_i(\xi_i)^\top\) lies in \(F'(X)\) entrywise, and \(0 = f(\tilde x) + J(x^\star - \tilde x)\) gives \(x^\star = \tilde x - Y f(\tilde x) + (I - YJ)(x^\star - \tilde x) \in K(X)\) by Theorem 2.6.4, which is (i), and (ii) is its contrapositive. For (iii), the same computation without \(f(x) = 0\) shows that \(h(x) = x - Y f(x)\) maps \(X\) into \(K(X) \subseteq X\). The map \(h\) is continuous on a compact convex set mapped into itself, so Brouwer's theorem gives a fixed point \(x^\star\), and \(Y f(x^\star) = 0\) with \(Y\) nonsingular gives \(f(x^\star) = 0\). For (iv), take \(\tilde x = \operatorname{mid} X\), so that \(X - \tilde x = [-v, v]\) with \(v = \operatorname{rad} X > 0\), and write \(M = I - Y F'(X)\) and \(|M|\) for the matrix of entrywise magnitudes. Then \(K(X) = \tilde x - Y f(\tilde x) + [-|M| v, |M| v]\), and \(K(X) \subseteq \operatorname{int} X\) forces \(|M| v < v\) componentwise. Every real matrix \(J \in F'(X)\) satisfies \(|I - YJ| \le |M|\) entrywise, so the spectral radius of \(I - YJ\) is at most that of \(|M|\), which is below \(1\) because \(|M| v < v\) for the positive vector \(v\). Hence \(YJ = I - (I - YJ)\) is nonsingular, and so are \(Y\) and \(J\). Existence follows from (iii). For uniqueness, two zeros \(x^\star, y^\star \in X\) give, by the mean value theorem on the segment between them, \(0 = J (x^\star - y^\star)\) with \(J \in F'(X)\) nonsingular, so \(x^\star = y^\star\). For (v), with \(Y\) and \(F'(X_0)\) fixed, \(\operatorname{rad} K(X_k) = |M| \operatorname{rad} X_k\), so in the weighted norm \(\|z\|_v = \max_i |z_i| / v_i\) the radii contract by the factor \(\theta = \max_i (|M| v)_i / v_i < 1\) at every step, while every \(X_k\) contains the zero by (i). ∎

The usual choices are \(\tilde x = \operatorname{mid} X\) and \(Y\) the floating-point inverse of \(f'(\tilde x)\). The theorem allows any \(Y\), so an inaccurate inverse costs tightness and never validity, and the interval evaluation of \(K\) makes the whole test rigorous by Theorem 2.6.7. When \(Y\) and \(F'(X_k)\) are re-evaluated on each box, as Algorithm 2.6.23 does, every \(X_k\) still contains the zero by (i), and the convergence becomes quadratic under a Lipschitz condition on \(F'\) (proof in Neumaier 1990).

Algorithm 2.6.23  Krawczyk test on a box, and branch and prune for
                  all zeros of f(x) = 0 in a region

Input   f : R^n -> R^n, an interval enclosure F of f and an enclosure
        F' of its Jacobian (forward-mode interval differentiation on
        the DAG); a box X; a minimum width w_min
Output  Test(X) in {NONE, UNIQUE(enclosure), UNDECIDED(box)};
        BranchAndPrune(X_0): certified and undecided boxes

Test(X):

 1. x~ = mid X;  Y = floating-point inverse of f'(x~)
    if singular: return UNDECIDED(X)

 2. in machine interval arithmetic:
       K = x~ - Y f([x~, x~]) + (I - Y F'(X)) (X - x~)

 3. if K meet X is empty:  return NONE          [Theorem 2.6.22 (ii)]

 4. if K lies in int X:    iterate X <- K(X) meet X, re-evaluating
                           x~, Y and F'(X), until the width stops
                           shrinking; return UNIQUE(X)
                                           [Theorem 2.6.22 (iv), (v)]

 5. return UNDECIDED(K meet X)

BranchAndPrune(X_0):

 6. L = {X_0}; while L is not empty: pop X from L

 7.    if some component F_i(X) excludes 0, f has no zero in X:
       discard X  (for f = grad phi this is the monotonicity test)

 8.    r = Test(X):  NONE: discard;  UNIQUE: record;
       UNDECIDED(X'): if w(X') < w_min record X' as undecided,
       else bisect X' along its widest side and push both halves

Invariant
    every zero of f in X_0 lies in a box of L or in a recorded box
                                                 [Theorem 2.6.22 (i)]
The branch-and-prune loop of Algorithm 2.6.23

  6. L = {X_0}
     |
     +<-----------------------------------------------------------+
     |                                                            |
     L empty? --yes--> output: the recorded boxes, certified      |
     | no              and undecided                              |
     pop X from L                                                 |
     |                                                            |
  7. some F_i(X) excludes 0? --yes--> discard X ------------------+
     |  (the monotonicity test, when f = grad phi)                |
     | no                                                         |
  8. r = Test(X)                                                  |
     +-- NONE: K meet X is empty -------> discard X --------------+
     +-- UNIQUE: K lies in int X -------> record the enclosure ---+
     +-- UNDECIDED(X'): X' = K meet X                             |
           |  (X' = X if f'(x~) is singular)                      |
           +-- w(X') < w_min --> record X' as undecided ----------+
           +-- else: bisect X' along its widest side and push     |
                     both halves onto L --------------------------+

  the figure's f on the unit square, system grad f = 0: seven
  boxes examined, four leaves excluded by K(X) meet X empty, none
  by the monotonicity test; f has no stationary point there

(The Krawczyk iteration on the figure's f) The cost per box is one interval Jacobian, one floating-point \(n \times n\) inverse and one interval matrix–vector product, with identical control flow for every box up to the three-way verdict, so a frontier of boxes is one kernel launch with a batched inverse. On the figure's function the stationary points are \((1, 1.678347)\) and \((1, -0.768039)\), both outside the unit square. On the box \([0.8, 1.2] \times [1.5, 1.9]\) the Krawczyk image \([0.826, 1.174] \times [1.555, 1.802]\) lies inside at once, and the iteration's maximum widths fall as \(3.5 \cdot 10^{-1}\), \(1.5 \cdot 10^{-1}\), \(2.9 \cdot 10^{-2}\), \(9.8 \cdot 10^{-4}\), \(1.1 \cdot 10^{-6}\) and \(1.4 \cdot 10^{-12}\). On the wider box \([0.8, 1.2] \times [-0.9, -0.6]\) around the other point the first image protrudes below \(-0.9\) by \(0.009\) and the test is inconclusive. One intersection step shrinks the box, and the second image lies inside. The test is sufficient, and the iteration makes its own hypothesis true after a step or two when a zero is present. On the unit square itself the branch-and-prune loop examines seven boxes and excludes four leaves by \(K(X) \cap X = \emptyset\), with none excluded by the monotonicity test because both components of \(\nabla f\) change sign on every box the loop sees. The conclusion is a proof, in the sense of Theorem 2.6.7, that \(f\) has no stationary point in the unit square, so its minimum \(-1.282\) and maximum \(0.250\) over the square lie on the boundary, as the figure's number line shows.

The Krawczyk test around a stationary point of the figure's f: on X = [0.8, 1.2] × [1.5, 1.9] the image K(X) = [0.826, 1.174] × [1.555, 1.802] lies inside X at once, so X holds exactly one stationary point, (1, 1.678347).
The Krawczyk test around the stationary points of the figure's f

  X' = [0.8, 1.2] x [-0.9, -0.6]
     |  K(X') protrudes below -0.9 by 0.009: inconclusive
     v
  X' <- K(X') meet X'
     |  the second image lies inside
     v
  exactly one zero, (1, -0.768039)

(Three uses of the test) Three uses of the test matter for this series. A box \(X^\star\) with \(K(X^\star) \subseteq \operatorname{int} X^\star\) for the gradient system, or for the KKT system of a constrained problem with the multipliers as unknowns, certifies exactly one stationary point. Every other box inside \(X^\star\) that does not contain that point can then be deleted. Such a box is an exclusion region, which Schichl and Neumaier enlarge to fight the clustering of boxes around found minimizers, the cluster problem of Section 2.4 (after Proposition 2.4.22, and stated as Theorem 3.5.7).H. Schichl and A. Neumaier, "Exclusion regions for systems of equations", SIAM Journal on Numerical Analysis 42 (2004). In interval branch and bound for unconstrained problems the exclusion \(K(X) \cap X = \emptyset\) complements the monotonicity test, which deletes a box on which some partial derivative has constant sign.R. B. Kearfott, Rigorous Global Search: Continuous Problems (Kluwer, 1996); Hansen and Walster (2004), cited above. And the operator is a contractor for equality constraints in constraint propagation, used in Kearfott's GlobSol and in the Ibex library, whose contractor programming is the abstraction behind Definition 2.6.9.G. Chabert and L. Jaulin, "Contractor programming", Artificial Intelligence 173 (2009); L. Jaulin, M. Kieffer, O. Didrit and É. Walter, Applied Interval Analysis (Springer, 2001). What the certificate costs against the tolerance-based practice of the MINLP solvers is weighed in R. B. Kearfott, "Interval computations, rigour and non-rigour in deterministic continuous global optimization", Optimization Methods and Software 26 (2011). What is certified is a KKT point of an NLP with the integers fixed, not global optimality, which is still proved by the branch-and-bound tree.

The measured value

Domain reduction is the component whose removal hurts a global solver most, and the measurements agree across solvers. Puranik and Sahinidis turned off BARON's five range-reduction options on four test libraries totalling \(1{,}740\) problems: node counts rose by \(1{,}180\%\), \(261\%\), \(546\%\) and \(802\%\) on the four libraries and times by \(67\%\), \(70\%\), \(75\%\) and \(47\%\). On the MINLP library the paper reports that without reduction "the no-reduction-based algorithm is not even able to find good feasible solutions". The other two solvers show the same dependence: Couenne's node counts rose by \(129\%\), \(21\%\), \(186\%\) and \(171\%\) on the four libraries and SCIP's by \(174\%\), \(152\%\), \(417\%\) and \(56\%\).Puranik and Sahinidis (2017), cited above, Section 8 and Tables 3, 5 and 7; the quotation is from Section 8. Vigerske and Gleixner's measurement for SCIP, quoted in Section 2.3, is that disabling domain propagation on \(456\) MINLPLib instances raised the shifted geometric mean of the running time by \(85\%\), of the primal–dual integral by \(150\%\) and of the node count by \(86\%\). The same table gives, for disabling OBBT on \(466\) instances, increases of \(32\%\) in time, \(89\%\) in the primal–dual integral and \(46\%\) in nodes.Vigerske and Gleixner (2018), cited above, Table 2, rows "domain propagation" and "OBBT". The domain-propagation row is the one Section 2.3 checked against the Optimization Online preprint (2016/05/5433, p. 19), with the columns in the order time, primal–dual integral, nodes; the OBBT row is read in the same column order from the same table and was not re-read for this post. The shifted geometric mean is defined in Section 5.2. Gleixner, Berthold, Müller and Weltge's numbers for OBBT with Lagrangian variable bounds are quoted above. The probing cost of BARON 7.5 measured by Belotti and co-authors, \(77\%\) of the time for five times fewer nodes, is the clearest statement of the trade-off every reduction technique makes: fewer nodes against more work per node. On the small examples of this series the ratio is extreme, \(13\) nodes processed, one LP each, against \(2\) LPs, because a single nonlinear constraint propagates exactly. On real instances it is the factors just quoted.

Where this is used

BARON runs every reduction of this subsection at every node by default. The list is linear feasibility-based tightening on the rows, nonlinear feasibility-based tightening through the factorable decomposition, marginals-based reduction from the stored duals, optimality-based tightening, and probing with the number of probes decided by the solver. Its manual places the whole subject in one sentence: "Tight bounds based on physical or problem-specific knowledge are often more valuable than any algorithmic option in this manual."BARON User Manual, cited above, Sections 4.1 and 11.3; the five reductions are the options LBTTDo, TDo, MDo, OBTTDo and PDo, in the order listed, the last defaulting to an automatic decision; GAMS, BARON solver documentation, gams.com/latest/docs/S_BARON.html (accessed 5 October 2026). SCIP runs FBBT through the interval callbacks of its expression handlers and through structure-specific propagators. Among them is a propagator for bivariate quadratic constraints that is "best possible when considering the bounds", following Domes and Neumaier, which is one of the places where a structure-aware rule escapes the dependency problem. It runs OBBT at the root only, propagates the Lagrangian variable bounds in the tree, and applies reduced-cost tightening to integer variables only by default.Bestuzheva, Chmiela, Müller, Serrano, Vigerske and Wegscheider (2025), cited above, Section 2, with F. Domes and A. Neumaier, "Constraint propagation on quadratic constraints", Constraints 15 (2010); SCIP 10.0.0 source, propagating/redcost/continuous = FALSE, prop_redcost.c. Couenne runs FBBT at every node with its cap of three rounds, OBBT at every node up to a chosen depth and with a decreasing probability below it, aggressive FBBT to a shallow depth, and reduced-cost tightening.Couenne source, src/couenne.opt and src/problem/problem.cpp: feasibility_bt yes, max_fbbt_iter 3, optimality_bt yes with log_num_obbt_per_level 1, aggressive_fbbt yes with log_num_abt_per_level 2, redcost_bt yes; Belotti, Lee, Liberti, Margot and Wächter (2009), cited above, Section 4. Gurobi's domain-reduction rules are not documented and nothing is claimed about them here. The rigorous codes are a different class of software. GlobSol, the COCONUT environment, Ibex and INTLAB use directed rounding throughout and certify what they return. The MINLP solvers of this series do not pay that cost: all of them run with tolerances, and none of them, in its default mode, produces a bound that Theorem 2.6.7 covers.R. B. Kearfott, "GlobSol user guide", Optimization Methods and Software 24 (2009); Neumaier (2004), cited above, for COCONUT; Chabert and Jaulin (2009), cited above, for Ibex; S. M. Rump, "INTLAB — INTerval LABoratory", in Developments in Reliable Computing (Kluwer, 1999). Interval arithmetic itself is standardized in IEEE, IEEE Standard for Interval Arithmetic, IEEE Std 1788-2015. SCIP 10's exact mode, which does produce certified bounds, covers linear problems only (Section 5.6).

What parallelizes

Three of the four techniques of this subsection are naturally parallel, and the theorems above say in what sense.

FBBT parallelizes in two ways. Within a round, every node of a wave of the DAG is independent. Across constraints, the Jacobi schedule lets every contractor read the state at the start of the round and merges the candidate intervals at the end. Theorem 2.6.11 guarantees that this schedule reaches the same fixed point as the sequential sweep. Proposition 2.6.12 warns that it needs more rounds, \(31\) against \(12\) on the figure's instance at tolerance \(10^{-3}\). Sofranac, Gleixner and Pokutta run whole propagation rounds for linear rows on the GPU. One warp or block handles one row, the activities are computed by a reduction over the row's nonzeros, the candidate bounds are merged by atomic minimum and maximum, and the rounds are iterated to a cap of \(100\). They report fixed points identical to the sequential code's on \(893\) of \(987\) MIPLIB 2017 instances and geometric-mean speed-ups of \(10\) to \(20\) over single-threaded propagation, with the parallel schedule needing more rounds, as the proposition predicts.B. Sofranac, A. Gleixner and S. Pokutta, "Accelerating domain propagation: an efficient GPU-parallel algorithm over sparse matrices", Parallel Computing 109 (2022). Across boxes, FBBT is one tape run on many boxes with identical control flow, in a node-major layout like the McCormick tape of Section 2.4. Directed rounding comes per operation from the intrinsics named above, so the rigour of Theorem 2.6.7 costs no mode switch. The one GPU lower-bounding engine built inside a deterministic global solver is Zhang and co-authors' interval bounder in MAiNGO. It partitions each node's box into many sub-boxes and evaluates a second-order interval form on all of them at once. The authors report a speed-up of three orders of magnitude over interval arithmetic on the CPU without partitioning and, in some case studies, bounds competitive with or better than the solver's default McCormick bounder. The measurement is their own.H. Zhang, T. Kerkenhoff, N. Kichler, M. Dahmen, A. Mitsos, U. Naumann and D. Bongartz, "Accelerating deterministic global optimization via GPU-parallel interval arithmetic", arXiv 2507.20769 (2025); the numbers are the authors' from the abstract. The GPU McCormick evaluation of Gottlieb, Xu and Stuber, cited in Section 2.4, is the companion measurement.

Linear-row propagation on the GPU (Sofranac, Gleixner and Pokutta)

  +--> the bounds l, u
  |       |
  |       +--> row 1: one warp or block --+   in each row: the
  |       +--> row 2: one warp or block --+   activities by a
  |       |    ...                        |   reduction over its
  |       +--> row m: one warp or block --+   nonzeros, then the
  |                                       |   candidate bounds
  |                                       v
  |     candidates merged into l, u by atomic minimum and maximum
  |                                       |
  +---------------------------------------+
     the next round, up to a cap of 100 rounds

  fixed points identical to the sequential code's on 893 of 987
  MIPLIB 2017 instances; geometric-mean speed-ups of 10 to 20 over
  single-threaded propagation, with more rounds (Proposition 2.6.12)

OBBT is a batch of up to \(2|K|\) LPs with one constraint matrix and different objectives. A device has no simplex warm-start to offer such a batch, but the batch is exactly the many-objective workload that a first-order LP method handles as one problem. Theorem 2.6.18 says that the duals such a method returns, inexact as they are, give valid range reductions as soon as the bound \(\underline z\) is computed from them by the safe bound of Section 7.3. The Lagrangian variable bounds then carry the information through the tree at the cost of propagation. Probing is the same batch with one bound changed per LP.

The Krawczyk test is a dense \(n \times n\) computation per box with fixed control flow. Certifying the thousands of candidate minimizers that a GPU multistart produces, and building exclusion regions around them, is therefore one kernel with a batched factorization.

What does not parallelize is the dependence of everything on the incumbent. The cutoff that makes FBBT, OBBT and reduced-cost tightening bite is a shared scalar, and every worker must see it as soon as it improves, which is the synchronization problem of Section 6. Section 7.8 takes up one open question from this subsection, how many Jacobi rounds the DAG's structure forces. Two others are left open here: whether a binary32 interval kernel's wider bounds interact badly with the cluster problem, and how much of BARON's \(77\%\) a batched probing round recovers.

← Back to all posts