3. Search

Draft

The relaxations of Section 2 produce bounds. This section is about the search that turns bounds into a proof: a tree of subproblems, each with its own relaxation, grown until every leaf has been shown to hold nothing better than the best point in hand. Section 3.1 states the algorithm and proves that it is correct and that it terminates. It then examines the three decisions that shape the tree, and it shows on a one-line instance that the tree can be exponential however those decisions are made. Section 3.2 is about the other half of the proof, the incumbent, and the heuristics that find one. Sections 3.3 and 3.4 add cutting planes and the methods for convex MINLP. Section 3.5 extends the tree to continuous nonconvexities, and Section 3.6 lays out what a node of a production solver actually does. Section 3.7 adds the bounds that exploit block structure: Lagrangian decomposition and a nonconvex Benders decomposition. Minimization is the default in every display. The knapsack of the first figure maximizes, and the text says so where it does.

Branch and bound

Branch and bound is the one algorithm that every exact solver in this series runs. The MILP engine runs it on linear relaxations. The convex MINLP methods of Section 3.4 run it on nonlinear relaxations or on a growing linear outer approximation. The spatial method of Section 3.5 runs it on envelope relaxations with domain reduction between the nodes. Land and Doig introduced the method in 1960 for discrete programming. The two-child split on a fractional variable, which every implementation since has used, is Dakin's, from 1965.A. H. Land and A. G. Doig, "An automatic method of solving discrete programming problems", Econometrica 28 (1960); R. J. Dakin, "A tree-search algorithm for mixed integer programming problems", The Computer Journal 8 (1965). Land and Doig's branching fixes the variable to integer values, nearest first; Dakin's abstract says his algorithm "appears to offer some advantages over a similar algorithm proposed by Land and Doig, from which it was developed", and names its "modest storage requirements". The name "branch and bound" is usually traced to J. D. C. Little, K. G. Murty, D. W. Sweeney and C. Karel, "An algorithm for the traveling salesman problem", Operations Research 11 (1963); the first survey is E. L. Lawler and D. E. Wood, "Branch-and-bound methods: a survey", Operations Research 14 (1966). The idea takes four sentences to state. Solve a relaxation. If its solution is not feasible for the original problem, split the feasible set into pieces whose union is the whole, and repeat on each piece. Discard any piece whose relaxation proves that it cannot beat the best point found so far. The rest of this subsection makes this precise, proves that it works, and then asks what it costs.

The algorithm

The problem is the one of Section 1, written with a single variable vector and an index set of integer coordinates:

\[z^\star \;=\; \min\{\, f(x) : x \in \mathcal F \,\}, \qquad \mathcal F \;=\; \{\, x \in B_0 = [l^0, u^0] : g_i(x) \le 0,\ i = 1, \dots, m,\ x_j \in \mathbb Z \text{ for } j \in I \,\},\]

with \(z^\star = +\infty\) when \(\mathcal F\) is empty. For a MILP, \(f(x) = c^\top x\) and \(g(x) = Ax - b\) with rational data.

Definition 3.1.1 (node, relaxation, node bound). A node \(N\) is a subproblem \(\min\{f(x) : x \in \mathcal F_N\}\) with \(\mathcal F_N \subseteq \mathcal F\). In the version studied here \(\mathcal F_N = \mathcal F \cap B_N\) for a box \(B_N = [l^N, u^N] \subseteq B_0\), so a node is a box of bounds. A relaxation of \(N\) is a problem \(\min\{\bar f(x) : x \in R_N\}\) with \(R_N \supseteq \mathcal F_N\) and \(\bar f \le f\) on \(\mathcal F_N\). For a MILP it is the LP over \(R_N = \{x : Ax \le b,\ l^N \le x \le u^N\}\) with \(\bar f = f\). Its optimal value is the node bound \(\bar z(N)\), with \(\bar z(N) = +\infty\) when \(R_N\) is empty, and \(\bar x_N\) is an optimal point of the relaxation. A node whose relaxation has not yet been solved carries its parent's bound as its inherited bound. Until the node's own relaxation is solved, \(\bar z(N)\) denotes that inherited bound, as in Definition 2.3.1. Node bounds are taken monotone: a child's bound is the larger of its own relaxation value and its parent's bound. This costs nothing and is valid because the child's subproblem is contained in the parent's.

Definition 3.1.2 (branching, Dakin's dichotomy, split disjunction). Branching a node \(N\) replaces it by finitely many children \(N_1, \dots, N_r\) with \(\mathcal F_{N_1} \cup \dots \cup \mathcal F_{N_r} = \mathcal F_N\). Dakin's dichotomy takes an integer variable \(x_j\), \(j \in I\), whose relaxed value \(v = (\bar x_N)_j\) is fractional and creates the two children \(B_N \cap \{x_j \le \lfloor v \rfloor\}\) and \(B_N \cap \{x_j \ge \lceil v \rceil\}\). The pair of inequalities is the split disjunction \(x_j \le \lfloor v \rfloor \ \vee\ x_j \ge \lfloor v \rfloor + 1\). A general split disjunction is \(\pi^\top x \le \pi_0 \ \vee\ \pi^\top x \ge \pi_0 + 1\) with \(\pi \in \mathbb Z^n\) supported on \(I\) and \(\pi_0 \in \mathbb Z\). Every point of \(\mathcal F\) satisfies one side, so the two children again cover \(\mathcal F_N\).

Definition 3.1.3 (pruning rules and node statuses). After its relaxation is solved, a node receives one of four statuses. It is infeasible (pruned by infeasibility in Definition 2.3.1) if \(R_N = \emptyset\). It is pruned by bound if \(\bar z(N) \ge z_{\mathrm{inc}}\), where \(z_{\mathrm{inc}}\) is the value of the best feasible point found so far (\(+\infty\) before the first one). With a tolerance \(\varepsilon \ge 0\) the test is \(\bar z(N) \ge z_{\mathrm{inc}} - \varepsilon\), and when the objective takes integer values on \(\mathcal F\) the test \(\lceil \bar z(N) \rceil \ge z_{\mathrm{inc}}\) is valid as well. It is feasible (closed by optimality in Definition 2.3.1) if \(\bar x_N \in \mathcal F\), which for a MILP means that every integer variable is integral at \(\bar x_N\). Then \(\bar x_N\) replaces the incumbent if \(f(\bar x_N) < z_{\mathrm{inc}}\). If \(\bar f = f\) on \(\mathcal F_N\), which holds for the LP relaxation of a MILP and for every relaxation that relaxes constraints only, then \(\bar z(N) = f(\bar x_N)\). The node is then a leaf, since no point of \(\mathcal F_N\) can beat the node's own relaxed point. If the objective itself has been relaxed, as in the spatial method of Section 3.5, the node is a leaf only when \(f(\bar x_N) \le \bar z(N) + \varepsilon\), and is otherwise branched on a continuous variable. Every other node is branched. Infeasible, pruned and feasible nodes are the leaves of the tree. Under dichotomy the tree is a full binary tree, so a tree with \(\ell\) leaves has \(2\ell - 1\) nodes.

Definitions 3.1.1 to 3.1.3 restate the objects of Definition 2.3.1 for a general relaxation, with the two status names changed as noted. The algorithm follows. The open list \(\mathcal L\) holds the nodes created and not yet processed. A node carries its inherited bound so that it can be discarded without solving anything when the incumbent has improved since the node was created. The version below is lazy: a node's relaxation is solved when the node is selected, not when it is created.

Algorithm 3.1.4  BRANCH-AND-BOUND
                 (minimization; Dakin's dichotomy; lazy node evaluation)

Input   f, g, box B_0 = [l^0, u^0], integer index set I with finite
        bounds; a procedure RELAX(N) returning (status, zbar(N), xbar_N)
        for a node N; a node selection rule; a branching rule;
        a tolerance eps >= 0.
Output  z_inc = z* (within eps) and a feasible x_inc attaining it, or
        "infeasible" (z_inc = +inf). The processed tree is the
        certificate: every leaf is infeasible, pruned by bound or
        feasible.

 1. z_inc <- +inf;  x_inc <- none
    L <- {N_0} with B_{N_0} = B_0 and inherited bound -inf

 2. while L is nonempty:

 3.    select N in L by the selection rule and remove it

 4.    if the inherited bound of N is >= z_inc - eps:
          status PRUNED; continue              [no relaxation solved]

 5.    (status, z, xbar) <- RELAX(N)
       zbar(N) <- max(z, inherited bound)            [monotone bound]

 6.    if status = INFEASIBLE: continue

 7.    if zbar(N) >= z_inc - eps: status PRUNED; continue

 8.    if xbar is feasible for the original problem
       (for a MILP: xbar_j integral for all j in I):
          if f(xbar) < z_inc:
             z_inc <- f(xbar), x_inc <- xbar, and
             delete from L every node whose inherited bound
             is >= z_inc - eps
          if f(xbar) <= zbar(N) + eps:
             status FEASIBLE; continue
                                       [always, when fbar = f on F_N]

 9.    status BRANCHED. Split B_N by the branching rule.
       For a MILP: choose j in I with xbar_j fractional, v <- xbar_j,
       and create
          N_1 = N with u_j <- floor(v)   and
          N_2 = N with l_j <- ceil(v).
       When xbar is feasible but f(xbar) > zbar(N) + eps (a relaxed
       objective): split the range of a continuous variable
       (Section 3.5).
       Give both children the inherited bound zbar(N) and push them
       onto L.

10. return z_inc, x_inc

Invariant (Theorem 3.1.5)
    z_inc is +inf or the value of a feasible point; every feasible x
    with f(x) < z_inc - eps lies in the box of some node of L; hence
    the global bound of Definition 2.3.1,
        zlow := min(z_inc, min over L of zbar(N))
    (the underlined z of the text; the inner minimum is +inf for an
    empty L), satisfies zlow <= z* <= z_inc.

Termination (Theorem 3.1.6)
    every branching shrinks an integer range by at least one, so the
    depth is at most D = sum over I of (u^0_j - l^0_j) and the tree
    has at most 2^{D+1} - 1 nodes.
The node loop of Algorithm 3.1.4 (minimization)

  2. while L is nonempty <----------------------------------------+
     |                                                            |
  3. select N in L by the rule and remove it                      |
     |                                                            |
  4. inherited bound >= z_inc - eps ? --yes--> PRUNED ------------+
     | no                                     (nothing solved)    |
  5. (status, z, xbar) <- RELAX(N)                                |
     zbar(N) <- max(z, inherited bound)                           |
     |                                                            |
  6. status = INFEASIBLE ? -------------yes--> INFEASIBLE --------+
     | no                                                         |
  7. zbar(N) >= z_inc - eps ? ----------yes--> PRUNED ------------+
     | no                                                         |
  8. xbar feasible ? --yes--> if f(xbar) < z_inc:                 |
     | no               |       z_inc <- f(xbar); x_inc <- xbar   |
     |                  |       delete from L every node whose    |
     |                  |       inherited bound is >= z_inc - eps |
     |                  v                                         |
     |     f(xbar) <= zbar(N) + eps ? --yes--> FEASIBLE ----------+
     |                  | no (a relaxed objective)                |
     v                  v                                         |
  9. BRANCHED: N_1, N_2 inherit zbar(N), pushed onto L -----------+

  L empty: 10. return z_inc, x_inc

The cost of a node is one relaxation, the feasibility test, the two pruning tests and the branching decision. For a MILP the relaxation is a dual simplex re-solve from the parent's basis, which takes a few pivots in the common case (Proposition 3.1.12 below). The branching decision ranges from free, for a rule that reads the fractional values, to two child LPs per candidate for strong branching. Step 4 lets a node die without a relaxation when the incumbent improved after the node was created. The knapsack figure below counts those nodes among the "pruned by bound". What parallelizes is the body of the loop. Steps 5 to 9 for different open nodes depend on nothing but the node's own data and the incumbent value (Proposition 3.1.18). A frontier of open nodes can therefore be bounded in one batch. Steps 3, 8 and 9 touch the shared state, the open list and the incumbent, and are where parallel implementations synchronize.

Proposition 2.3.2 stated the invariant for the run of Section 2.3. The theorem below is its general form, for any valid relaxation.

Theorem 3.1.5 (the branch-and-bound invariant). Run Algorithm 3.1.4 with valid relaxations and \(\varepsilon = 0\). After every step the following hold.

(a) \(z_{\mathrm{inc}}\) is \(+\infty\) or the value \(f(x_{\mathrm{inc}})\) of a point \(x_{\mathrm{inc}} \in \mathcal F\), so \(z^\star \le z_{\mathrm{inc}}\).

(b) Every \(x \in \mathcal F\) with \(f(x) < z_{\mathrm{inc}}\) lies in \(\mathcal F_N\) for some open node \(N \in \mathcal L\).

(c) The global bound \(\underline z := \min\{ z_{\mathrm{inc}},\ \min_{N \in \mathcal L} \bar z(N) \}\) of Definition 2.3.1, in which the inner minimum is taken over the open nodes with their inherited bounds and is \(+\infty\) for an empty list, satisfies \(\underline z \le z^\star \le z_{\mathrm{inc}}\). Whenever \(\mathcal L\) is nonempty and \(z^\star < z_{\mathrm{inc}}\), the open-list minimum alone satisfies \(\min_{N \in \mathcal L} \bar z(N) \le z^\star\).

Solvers call the global bound \(\underline z\) the dual bound.

Proof. By induction on the steps. Initially \(\mathcal L = \{N_0\}\) and \(z_{\mathrm{inc}} = +\infty\). Conditions (a) to (c) hold because \(\mathcal F_{N_0} = \mathcal F\) and \(\bar z(N_0) \le z^\star\) by the validity of the relaxation. A step removes a node \(N\) and does one of four things.

Case 1: the node is infeasible. \(R_N = \emptyset\) implies \(\mathcal F_N = \emptyset\), so no point of \(\mathcal F\) leaves the cover in (b), and (c) holds because the removed node contained no point at all.

Case 2: the node is pruned by bound. For every \(x \in \mathcal F_N\), validity gives \(f(x) \ge \bar f(x) \ge \bar z(N) \ge z_{\mathrm{inc}}\), so no point with \(f(x) < z_{\mathrm{inc}}\) is lost from (b). For (c), if \(z^\star < z_{\mathrm{inc}}\) then an optimal \(x^\star\) has \(f(x^\star) < z_{\mathrm{inc}}\). By (b) it lies in some remaining open node \(N'\), and \(\underline z \le \min_{N \in \mathcal L} \bar z(N) \le \bar z(N') \le f(x^\star) = z^\star\).

Case 3: the relaxed point is feasible. The incumbent becomes \(\bar x_N\) if \(f(\bar x_N) < z_{\mathrm{inc}}\), which keeps (a). The node is a leaf only when \(f(\bar x_N) \le \bar z(N)\), and then every \(x \in \mathcal F_N\) has \(f(x) \ge \bar z(N) \ge f(\bar x_N) \ge z_{\mathrm{inc}}\), so removing \(N\) keeps (b). The nodes deleted from \(\mathcal L\) at the same time are removed by the argument of Case 2. (c) follows from (a) and (b) as before. If the node is not a leaf, Case 4 applies.

Case 4: the node is branched. The children cover \(\mathcal F_N\) by Definition 3.1.2, and each child's bound is at most the minimum of \(f\) over the child by validity, so (b) is kept. The children inherit \(\bar z(N) \le f(x)\) for every \(x \in \mathcal F_N\), so \(\min_{N \in \mathcal L} \bar z(N)\), and with it \(\underline z\), is at most the value of an optimal point in whichever child contains it, which is (c). ∎

In the picture, the open list covers the part of the feasible set that could still beat the incumbent. The node bounds are lower bounds on \(f\) over the pieces of that cover. Every pruning rule removes a piece that provably contains no improving point. Nothing in the proof uses linearity, integrality or finiteness. The invariant therefore holds word for word for spatial branch and bound in Section 3.5, whose relaxations relax the objective as well as the constraints. It holds for any relaxation that is valid in the sense of Definition 3.1.1. On the R1 run of the table above the invariant can be read off directly. After node 5 has delivered the incumbent, the open list holds nodes 6 and 7 with the inherited bound \(5.4896\) of their parent, and (c), read for a maximization, says \(5.3825 \le z^\star \le 5.4896\). Processing the two closes the bracket: node 6 is empty, and node 7's own bound \(4.8874\) cannot beat the incumbent, so the list empties with \(z_{\mathrm{inc}} = z^\star = 5.3825\).

Theorem 3.1.6 (Land and Doig 1960, Dakin 1965). Let the problem have finitely many integer variables with finite bounds \(l^0_j \le x_j \le u^0_j\), \(j \in I\). Let the relaxation satisfy \(\bar f = f\) on every \(\mathcal F_N\), and let every relaxed point that is integral on \(I\) belong to \(\mathcal F_N\). Both hold for the LP relaxation of a MILP, which enforces every constraint but integrality. Let branching be Dakin's dichotomy on an integer variable with fractional relaxed value. Then Algorithm 3.1.4 stops after finitely many steps, and at termination \(z_{\mathrm{inc}} = z^\star\), with \(z_{\mathrm{inc}} = +\infty\) exactly when \(\mathcal F = \emptyset\). The tree has depth at most \(D = \sum_{j \in I} (u^0_j - l^0_j)\) and at most \(2^{D+1} - 1\) nodes.

Proof. Termination. Every node with a nonempty relaxation has a relaxed point that is either integral on \(I\) or fractional in some \(x_j\), \(j \in I\). In the first case the point is feasible and the node is a leaf by Definition 3.1.3. In the second the node is branched on such a variable. Along any root-to-node path, each branching on \(x_j\) replaces the range \([l_j, u_j]\) by \([l_j, \lfloor v \rfloor]\) or \([\lceil v \rceil, u_j]\) with \(l_j < v < u_j\) fractional. The integer range \(u_j - l_j\) therefore decreases by at least one and never below zero. The ranges start at \(u^0_j - l^0_j\), so a path contains at most \(D\) branchings and the depth is at most \(D\). A node at which every integer variable has range zero has all of them fixed. Its relaxed point is then integral on \(I\) and the node is a leaf, or its relaxation is empty. A full binary tree of depth at most \(D\) has at most \(2^{D+1} - 1\) nodes, and each step processes one node, so the run is finite. Correctness. At termination \(\mathcal L = \emptyset\). By Theorem 3.1.5(b) there is no \(x \in \mathcal F\) with \(f(x) < z_{\mathrm{inc}}\), so \(z^\star \ge z_{\mathrm{inc}}\), and by (a) \(z^\star \le z_{\mathrm{inc}}\). If \(\mathcal F \ne \emptyset\) then \(z^\star < +\infty\) forces \(z_{\mathrm{inc}} < +\infty\). If \(\mathcal F = \emptyset\), no incumbent can be found. ∎

The bound \(2^{D+1} - 1\) is the size of the full enumeration tree. For \(n\) binary variables it is \(2^{n+1} - 1\), which is the "full tree" line that the figure below draws against the actual count: 8,191 nodes for twelve items. With the tolerance \(\varepsilon > 0\) in the pruning test the same proof gives termination with \(z_{\mathrm{inc}} \le z^\star + \varepsilon\). The finite bounds are needed. Proposition 3.1.7 shows what fails without them.

Proposition 3.1.7 (bounded integer variables are needed). Without finite bounds on the integer variables, branch and bound with Dakin's dichotomy need not terminate, even on an infeasible problem in two variables. The feasibility problem \(x_1 - x_2 = \tfrac12\), \(x_1, x_2 \in \mathbb Z_{\ge 0}\), with the objective \(\min x_1\) and the relaxed optimum taken at a vertex, generates an infinite path.

Proof. The problem is infeasible, because \(x_1 - x_2\) is an integer at integer points. The root LP is \(\min\{x_1 : x_1 - x_2 = 1/2,\ x \ge 0\}\) with optimum \((1/2, 0)\). Branching on \(x_1\), the child \(x_1 \le 0\) forces \(x_2 = -1/2\) and is infeasible, and the child \(x_1 \ge 1\) has optimum \((1, 1/2)\). Branching on \(x_2\), the child \(x_2 \le 0\) forces \(x_1 = 1/2 < 1\) and is infeasible, and the child \(x_2 \ge 1\) has optimum \((3/2, 1)\). Inductively, after \(k\) branchings the surviving node has a relaxed optimum with exactly one fractional coordinate, one child is infeasible and the other continues the path. There is no incumbent, so nothing is pruned by bound, and the run never ends. ∎

Proposition 3.1.7: an infinite path (x_1 - x_2 = 1/2, min x_1)

                    root (1/2, 0)
           x_1 <= 0 /          \ x_1 >= 1
                   /            \
  infeasible: x_2 = -1/2       (1, 1/2)
                      x_2 <= 0 /      \ x_2 >= 1
                              /        \
        infeasible: x_1 = 1/2 < 1       (3/2, 1)
                                        branch on x_1: one child
                                        infeasible, the other
                                        goes on, and so forever

  no incumbent is ever found, so no node is pruned by bound

With the bound \(x_1 \le k\) added, the run ends when \(x_1\) reaches its bound. The chain of surviving nodes then has \(2k\) nodes, every branching also produces one infeasible leaf, and the tree has \(4k + 1\) nodes and depth \(2k\). Theorem 3.1.6 applies once the implied bound \(x_2 \le k - 1\) is made explicit, with \(D = 2k - 1\), and the tree then has \(4k - 1\) nodes. Dadush and Tiwari use this bounded relative as the simplest example of a branching proof whose length grows with the bounds.D. Dadush and S. Tiwari, "On the complexity of branching proofs", arXiv 2006.04124 (2020), Section 1.1. The paper was presented at the Computational Complexity Conference 2020; its proceedings record was not verified for this series. Both versions are solved at the root by the single disjunction \(x_1 - x_2 \le 0 \ \vee\ x_1 - x_2 \ge 1\), whose two sides are both empty. The general lesson, that a disjunction aligned with the structure can replace a long tree, returns below with Jeroslow's instance. A rational MILP without explicit bounds is still finite in principle, because an optimal solution of polynomial size exists when any solution does (Section 1.5). Solvers rely on that only indirectly, through bound propagation and the bounds the user supplies.

The running example, and the knapsack

On the two-variable program R1 the algorithm is short enough to follow by hand. The polygon is \(2x + 5y \le 24.5\), \(5x + 2y \le 30.5\), \(-3x + 4y \le 11\), \(x - 2y \le 4.2\) inside the box \([0, 7] \times [0, 6]\), and the objective is to maximize \(\cos\theta\, x + \sin\theta\, y\) over integer \(x, y\). This example maximizes, so bounds are upper bounds and a node is pruned when its bound is at most the incumbent. At \(\theta = 20^\circ\), with best-bound selection, which takes next the open node whose bound is best (Definition 3.1.8 below), and branching on the more fractional variable, the one whose relaxed value is nearer to a half, the tree has seven nodes:

nodeboundsvaluestatus
1\(0 \le x \le 7\), \(0 \le y \le 6\)5.7053branch on \(x = 5.7833\)
2\(6 \le x \le 7\), \(0 \le y \le 6\)–LP infeasible
3\(0 \le x \le 5\), \(0 \le y \le 6\)5.6390branch on \(y = 2.7500\)
4\(0 \le x \le 5\), \(3 \le y \le 6\)5.4896branch on \(x = 4.7500\)
5\(0 \le x \le 5\), \(0 \le y \le 2\)5.3825integral: new incumbent \((5, 2)\)
6\(5 \le x \le 5\), \(3 \le y \le 6\)–LP infeasible
7\(0 \le x \le 4\), \(3 \le y \le 6\)4.8874pruned by bound (incumbent 5.3825)
 optimum5.3825at \((5, 2)\); 7 nodes
The same run as a tree (R1, theta = 20 degrees, best bound first)

                         (1) 5.7053
                     branch on x = 5.7833
                    x >= 6 /        \ x <= 5
                          /          \
          (2) LP infeasible          (3) 5.6390
                                 branch on y = 2.75
                                y >= 3 /       \ y <= 2
                                      /         \
                         (4) 5.4896              (5) 5.3825
                     branch on x = 4.75          integral (5, 2):
                    x >= 5 /      \ x <= 4       new incumbent
                          /        \
          (6) LP infeasible        (7) 4.8874 <= 5.3825:
                                       pruned by bound

In the symbols of Definitions 3.1.1 to 3.1.3, each row of the table is a node \(N\): the bounds column is its box \(B_N\), its relaxation is the LP over the polygon inside that box, with \(\bar f = f\), the value column is the node bound \(\bar z(N)\), the relaxed point \(\bar x_N\) is the vertex whose fractional coordinate the status column names, and the status is one of the four of Definition 3.1.3, read in the maximizing sense, so that \(z_{\mathrm{inc}}\) starts at \(-\infty\), becomes \(5.3825\) at node 5, and prunes node 7 because \(4.8874 \le 5.3825\). Node 1 is the root, with the fractional vertex \((5.7833, 0.7917)\). The child \(x \ge 6\) is empty and the child \(x \le 5\) has a new vertex \((5, 2.75)\). The branching on \(y\) produces the incumbent \((5, 2)\) at node 5 and the two nodes that close the tree. The vocabulary figure of Section 2.3 replays this tree, node for node, in its "branch and bound alone" mode at this angle. That figure counts steps rather than nodes. Its nine steps in that mode are the root relaxation, the root's branching, the six nodes below the root in the order above, and a final step that records the empty list. The three nodes it reports as pruned are nodes 2, 6 and 7 of this table. In its default mode the vocabulary figure adds a cut and a rounding heuristic to the run. At \(\theta = 45^\circ\) the tree has eleven nodes and the optimum found is \((4, 3)\), tied with \((5, 2)\), and at \(\theta = 70^\circ\) it has nine nodes and the optimum is \((2, 4)\).

The four leaves of the R1 run as boxes (θ = 20°): node 5 holds the incumbent (5, 2), node 7 is pruned by bound, and nodes 2 and 6 are LP infeasible. The gaps 5 < x < 6, 2 < y < 3 and 4 < x < 5 that the three branchings cut out hold no integer point.

A two-variable program never has a large tree. The knapsack problem, introduced with the tax lots of Section 1.6, does, and it is the instance of the first figure. There are \(n\) items with weights \(w_i\) and values \(v_i\) and a capacity \(W\), and the task is to choose the most valuable set that fits:

\[\max\Big\{ \sum_{i=1}^n v_i x_i \;:\; \sum_{i=1}^n w_i x_i \le W,\ x \in \{0, 1\}^n \Big\}.\]

Its LP relaxation has a closed form, Dantzig's bound of 1957.G. B. Dantzig, "Discrete-variable extremum problems", Operations Research 5 (1957). The first branch-and-bound code for the knapsack with this bound is P. J. Kolesar, "A branch and bound algorithm for the knapsack problem", Management Science 13 (1967). Sort the items by value per unit of weight. At a node, let \(F_1\) be the items fixed in and \(F_0\) the items fixed out. The node is infeasible when \(\sum_{i \in F_1} w_i > W\). Otherwise take the free items whole in ratio order until one does not fit, and take the fitting fraction of that one:

\[\bar z(N) \;=\; \sum_{i \in F_1} v_i \;+\; \sum_{\text{free } i < k} v_i \;+\; \frac{W - \sum_{i \in F_1} w_i - \sum_{\text{free } i < k} w_i}{w_k}\, v_k ,\]

where \(k\) is the first free item that does not fit. The root is the case \(F_0 = F_1 = \emptyset\). The relaxed solution has at most one fractional item, and that item is the branching variable, with the children "in" (\(x_k = 1\)) and "out" (\(x_k = 0\)). Because the values are integers the pruning test is \(\lfloor \bar z(N) \rfloor \le z_{\mathrm{inc}}\), the maximization form of the rounded test in Definition 3.1.3. The figure's default instance has twelve items drawn from a seeded generator so that it is the same on every visit. Sorted by value per unit of weight they are, as (weight, value) pairs, L (17, 36), I (18, 38), A (22, 41), F (24, 41), C (30, 49), B (32, 50), D (23, 35), J (34, 50), E (26, 38), K (36, 49), G (40, 54), H (40, 53). The total weight is 342 and the capacity is 171. The values sit a little above the weights, which is the family on which Dantzig's bound is least informative. The greedy solution, taking items in ratio order whenever they fit, is worth 290 against the optimum 293.

Dantzig's bound at the root of the knapsack (capacity 171)

  ratio order    L   I   A   F   C   B   D   J   E   K   G   H
  weight        17  18  22  24  30  32  23  34  26  36  40  40
  value         36  38  41  41  49  50  35  50  38  49  54  53
  root LP x      1   1   1   1   1   1   1  frac 0   0   0   0
               |<------- whole: 290 ----->|  ^ |<--- at 0 --->|
                                             |
     L to D fit whole, worth 290 as in the greedy solution; J is
     the first item that does not fit, so the fitting fraction of
     it is taken, and J is the branching variable of the root:
     "in" (x_J = 1) and "out" (x_J = 0)

The figure runs the algorithm on this instance and on a second instance introduced below, with a scrub that replays the run node by node. The Chvátal–Gomory cut named in the caption is an inequality obtained from the rows of the relaxation by taking a nonnegative combination of them and rounding its coefficients and its right-hand side down, which every integer point satisfies; it is defined below, before Proposition 3.1.16, and studied in Section 3.3. The main panel draws the tree in processing order: black for branched nodes, green for an incumbent, grey for pruned by bound, red for infeasible, with a green ring where the first incumbent arrived. The side panel is the chart a solver's log shows: the global bound coming down and the incumbent going up. The shaded area under the optimum is the primal integral of Definition 2.3.4, the gap between the incumbent and the optimum summed over the nodes processed, so that a run which finds good incumbents early has a small one; here it is counted in nodes rather than seconds. In the default view, depth-first search from no incumbent, the run takes 315 nodes against the 8,191 of the full tree, with 83 nodes pruned by bound and 70 infeasible. The optimum is 293, and the first incumbent, worth 269, arrives at node 19. Switch to best-bound search and the run takes 37 nodes, with the first incumbent, 290, at node 10 and the optimum at node 18. Choose "dive first" and the search is depth-first until that first incumbent at node 19, then best-bound, for 53 nodes in all. Start from the greedy incumbent instead and the depth-first run shrinks from 315 nodes to 57, with the primal integral falling from 36.9 to 0.3 node-units.

Branch and bound on a knapsack and on Jeroslow's parity instance 2(x₁ + … + xₙ) = n, x binary, n odd. Each dot is a node in processing order (black branched, green an incumbent, grey pruned by bound, red infeasible, a green ring where the first incumbent arrived), and the scrub replays the run. On the knapsack the node order and the starting incumbent change the tree, while on the parity instance no bound ever prunes, the tree has 2·C(n + 1, (n + 1)/2) − 1 nodes in any order, and one Chvátal–Gomory cut makes the root infeasible.

Here is the same computation as a program. It generates the figure's instance from the figure's seed, runs Algorithm 3.1.4 with the knapsack's conventions, and prints the six runs of the figure's comparison table: the three node orders, each from no incumbent and from the greedy one.

# Branch and bound for the 0-1 knapsack, as fig-bnb runs it.
#
# Dantzig's bound; branching on the fractional item with the "in" child
# first; pruning when the rounded bound is no better than the incumbent;
# the three node orders of the figure. The instance is the figure's
# default: 12 items, seed 9, capacity half the weight.

from math import floor, inf

def mulberry32(seed):
    """The figure's seeded generator, so the instance is the same."""
    a = seed & 0xFFFFFFFF

    def r():
        nonlocal a
        a = (a + 0x6D2B79F5) & 0xFFFFFFFF
        t = a
        t = ((t ^ (t >> 15)) * (t | 1)) & 0xFFFFFFFF
        t ^= ((t + (((t ^ (t >> 7)) * (t | 61)) & 0xFFFFFFFF))
              & 0xFFFFFFFF)
        return ((t ^ (t >> 14)) & 0xFFFFFFFF) / 4294967296

    return r

def instance(n=12, seed=9, cap_ratio=0.5):
    r = mulberry32(seed * 7919 + n * 104729 + 17)
    items = []
    for i in range(n):
        w = 10 + int(r() * 31)
        v = w + 12 + int(r() * 9)            # value a little over weight
        items.append((w, v, chr(65 + i)))
    items.sort(key=lambda it: -it[1] / it[0])   # by value per weight
    return items, round(cap_ratio * sum(w for w, _, _ in items))

def dantzig(items, cap, dec):
    """Forced items whole, free items in order, then a fraction.

    Returns (bound, j) with j the one fractional item, (bound, -1) when
    nothing is fractional (every free item fits), and (None, None) when
    the forced items do not fit (the node is infeasible).
    """
    w = sum(it[0] for it, d in zip(items, dec) if d == 1)
    v = sum(it[1] for it, d in zip(items, dec) if d == 1)
    if w > cap:
        return None, None
    rem = cap - w
    for j, (wj, vj, _) in enumerate(items):
        if dec[j] == -1 and wj <= rem:
            rem -= wj
            v += vj
        elif dec[j] == -1:
            return v + rem * vj / wj, j
    return v, -1

def greedy(items, cap):
    rem, v = cap, 0
    for wj, vj, _ in items:
        if wj <= rem:
            rem -= wj
            v += vj
    return v

def branch_and_bound(items, cap, order="dfs", start=-inf):
    n = len(items)
    inc, first, incs = start, None, []
    st = dict(branched=0, pruned=0, infeasible=0)
    open_ = [((-1,) * n, inf)]             # (decisions, inherited bound)
    while open_:
        best_first = order == "best" or (order == "dive" and inc > -inf)
        if best_first:                     # oldest among equals
            k = max(range(len(open_)), key=lambda i: open_[i][1])
        else:
            k = len(open_) - 1
        dec, pbound = open_.pop(k)
        if pbound < inf and floor(pbound + 1e-9) <= inc:
            st["pruned"] += 1              # inherited bound: nothing solved
        else:
            bound, frac = dantzig(items, cap, dec)
            if bound is None:
                st["infeasible"] += 1
            elif floor(bound + 1e-9) <= inc:
                st["pruned"] += 1
            elif frac < 0:                 # a new incumbent
                inc = bound
                first = first or (len(incs) + 1, bound)
            else:
                st["branched"] += 1
                d1 = dec[:frac] + (1,) + dec[frac + 1:]
                d0 = dec[:frac] + (0,) + dec[frac + 1:]
                # depth-first takes "in" first
                open_ += [(d0, bound), (d1, bound)]
        incs.append(inc)                   # the incumbent after this node
    # the primal integral, in node units
    P = sum(1.0 if z == -inf else (inc - z) / inc for z in incs)
    return dict(nodes=len(incs), opt=inc, first=first, P=P, **st)

items, cap = instance()
total = sum(w for w, _, _ in items)
print("items (weight, value):")
for row in range(0, len(items), 6):
    chunk = items[row:row + 6]
    print("   ", "  ".join(f"{c}({w},{v})" for w, v, c in chunk))
print(f"capacity {cap} of {total}")
print(f"greedy incumbent {greedy(items, cap)}; "
      f"full tree 2^13 - 1 = {2**13 - 1} nodes")
print()
print(f"{'order':<5}{'start':>6}{'nodes':>6}{'opt':>5}"
      f"{'first incumbent':>17}{'branch':>7}{'prune':>6}{'infeas':>7}"
      f"{'P(T)':>6}{'P(T)/T':>7}")
for start in (-inf, greedy(items, cap)):
    for order in ("dfs", "best", "dive"):
        r = branch_and_bound(items, cap, order, start)
        f = (f"{r['first'][1]} at node {r['first'][0]}"
             if r["first"] else "greedy")
        s = "none" if start == -inf else start
        print(f"{order:<5}{s:>6}{r['nodes']:>6}{r['opt']:>5}{f:>17}"
              f"{r['branched']:>7}{r['pruned']:>6}{r['infeasible']:>7}"
              f"{r['P']:>6.1f}{r['P'] / r['nodes']:>7.3f}")
items (weight, value):
    L(17,36)  I(18,38)  A(22,41)  F(24,41)  C(30,49)  B(32,50)
    D(23,35)  J(34,50)  E(26,38)  K(36,49)  G(40,54)  H(40,53)
capacity 171 of 342
greedy incumbent 290; full tree 2^13 - 1 = 8191 nodes

order start nodes  opt  first incumbent branch prune infeas  P(T) P(T)/T
dfs    none   315  293   269 at node 19    157    83     70  36.9  0.117
best   none    37  293   290 at node 10     18    17      0   9.1  0.245
dive   none    53  293   269 at node 19     26    18      6  18.7  0.354
dfs     290    57  293   291 at node 23     28    26      1   0.3  0.005
best    290    35  293   293 at node 18     17    17      0   0.2  0.005
dive    290    35  293   293 at node 18     17    17      0   0.2  0.005

The six rows are the figure's numbers. With the greedy start the dive has nothing to dive for, since an incumbent exists from the first node, so "dive first" coincides with best-bound search at 35 nodes. The first improvement the search itself finds is the optimum at node 18. The two last columns are the two normalizations of the primal integral of Definition 2.3.4: \(P(T)\) sums the primal gap over the \(T\) nodes processed and \(\bar P(T) = P(T)/T\) averages it. The two rank the orders differently. \(P(T)\) charges the depth-first run for its length, 36.9 against 9.1. The average rewards it for finding incumbents within a small fraction of its nodes, 0.117 against 0.245. The cost of the program is one Dantzig bound per node, a pass over at most \(n\) items, so the whole 315-node run is a few thousand arithmetic operations. The bounds of different open nodes are independent and could be computed together, which is the batching of Section 6.4.

Node selection

The size of the tree is the cost of the solve, and it is unknown until the solve is over. Three decisions shape it: which open node to process next, which variable to branch on, and how early a good incumbent arrives. The third is the subject of Section 3.2. The first two are treated here.

Definition 3.1.8 (selection rules). Depth-first search takes the most recently created node, so it dives from a node to a child. Among the two children it takes a fixed one first (the figure takes the "in" child). Best-first, or best-bound, search takes an open node of smallest bound and, among equals, the oldest. In the lazy form of Algorithm 3.1.4 the key is the inherited bound. In the eager form a child's relaxation is solved when the child is created, and the key is the node's own value \(\bar z(N)\). On a problem in which every bound is equal, best-first search becomes breadth-first. Best-estimate search takes the node that minimizes an estimate of the best feasible value in its subtree, \(\bar z(N) + \sum_j \min(\Psi^-_j f_j, \Psi^+_j (1 - f_j))\). Here \(f_j\) is the fractional part of \(\bar x_j\), and \(\Psi^\pm_j\) are the pseudocosts of Definition 3.1.11, the average bound gains per unit change of \(x_j\) observed in earlier branchings.J. J. H. Forrest, J. P. H. Hirst and J. A. Tomlin, "Practical solution of large mixed integer programming problems with UMPIRE", Management Science 20 (1974), for the estimate; J. T. Linderoth and M. W. P. Savelsbergh, "A computational study of search strategies for mixed integer programming", INFORMS Journal on Computing 11 (1999), for the comparison of depth-first, best-bound, best-estimate and the hybrids. Plunging is a depth-first descent from the selected node until a pruning rule or a depth limit stops it, after which the rule returns to best-bound or best-estimate selection. Dive-first is depth-first until an incumbent exists and best-first afterwards.

Proposition 3.1.9 (best-first search solves the fewest relaxations, up to ties). Fix the problem, a relaxation with \(\bar f = f\) on every \(\mathcal F_N\), the branching rule and \(\varepsilon = 0\), and suppose the children of a node depend only on the node. Let \(\mathcal T\) be the tree obtained by branching every node whose relaxation is feasible and whose relaxed point is not feasible, without pruning by bound. Call a node \(N\) of \(\mathcal T\) critical if \(\bar z(N) < z^\star\). Then (a) every run of Algorithm 3.1.4, with any selection rule, branches every critical node and solves the relaxation of the root and of every child of a critical node. (b) An eager best-first run branches only nodes with \(\bar z(N) \le z^\star\), so the relaxations it solves are those of the root and of the children of nodes with \(\bar z(N) \le z^\star\). The number of relaxations an eager best-first run solves is therefore minimal among all selection rules, up to the children of nodes whose bound equals \(z^\star\).

Proof. (a) Let \(N\) be critical. Its ancestors have bounds at most \(\bar z(N) < z^\star\) by monotonicity, so none of them is pruned by bound, since \(z_{\mathrm{inc}} \ge z^\star\) at all times by Theorem 3.1.5(a). None of them is a feasible leaf, because a feasible relaxed point at an ancestor \(A\) would have \(f(\bar x_A) = \bar z(A) < z^\star\), contradicting optimality. So every ancestor is branched, \(N\) is created, and the run does not stop while \(N\) is open. When \(N\) is selected its inherited bound is below \(z^\star \le z_{\mathrm{inc}}\), so its relaxation is solved and \(N\) is branched by the same argument. Its children inherit \(\bar z(N) < z^\star\), so they are never deleted or pruned without a solve, and their relaxations are solved in turn. (b) Suppose an eager best-first run selects a node \(N\) with \(\bar z(N) > z^\star\) and branches it. Then every open node has bound at least \(\bar z(N) > z^\star\). If \(z_{\mathrm{inc}} = z^\star\) already, \(N\) is pruned by the test \(\bar z(N) \ge z_{\mathrm{inc}}\), a contradiction. If \(z_{\mathrm{inc}} > z^\star\), Theorem 3.1.5(b) puts an optimal point in some open node \(N'\), whose bound is then at most \(z^\star < \bar z(N)\), contradicting the choice of \(N\). ∎

(Lazy against eager, counted on the knapsack) Algorithm 3.1.4 and the knapsack figure implement the lazy variant, in which the key is the inherited bound and the relaxation is solved after selection. The argument of (b), applied to inherited bounds, shows that the lazy variant solves a relaxation only when the inherited bound is at most \(z^\star\). Up to the children of nodes with \(\bar z(N) = z^\star\), it too solves only the root and the children of critical nodes. It can, however, branch a node whose own bound exceeds \(z^\star\) while no incumbent of value \(z^\star\) is in hand, which the eager variant never does. On the knapsack, which maximizes, the lazy best-bound run solves 33 relaxations: those of the root and of the children of its 16 critical nodes. It also branches two nodes whose bounds, 291.5 and 290.25, are worse than the optimum 293. Their four children are later discarded on their inherited bounds without a solve, which brings the figure's count to 37. An eager run on the same instance solves 39 relaxations, because three nodes with bound exactly 293 are selected before the incumbent 293 arrives and are branched. The figure's 37 is therefore a measured count, and the minimality of the proposition is a statement about relaxations up to ties, not about this number. The result itself is classical, and Ibaraki's 1976 analysis of search strategies is the reference.T. Ibaraki, "Theoretical comparisons of search strategies in branch-and-bound algorithms", International Journal of Computer and Information Sciences 5 (1976). Two caveats keep it from being a recommendation. The proposition says nothing about time. A best-first run holds many open nodes, and each selected node is far from the previous one in the tree. The warm start of Proposition 3.1.12 is cheapest when consecutive nodes are parent and child. And ties at \(z^\star\) can be numerous. On Jeroslow's instance below every node has the same bound, every rule processes the same nodes, and best-first degenerates into breadth-first.

Proposition 3.1.10 (depth-first search and memory). Under dichotomy with tree depth at most \(D\), depth-first search never holds more than \(D + 1\) open nodes, and it reaches a leaf within \(D + 1\) node evaluations from the root.

Proof. Processing a node at depth \(d\) removes it and adds at most two children at depth \(d + 1\), one of which is taken next. The open list therefore holds at most one deferred sibling per level of the current path, plus the current node, and the path has at most \(D + 1\) nodes. ∎

Proposition 3.1.10: what depth-first search keeps open

  depth 0        N_0
                /   \
  depth 1     N_1    S_1      deferred sibling, waits in L
             /   \
  depth 2  N_2    S_2         deferred sibling, waits in L
            :
          N_(d-1)
             /   \
  depth d  N_d    S_d         deferred sibling, waits in L
            ^ the node being processed

  L: at most one deferred sibling per level of the current path,
  plus the current node; the path has at most D + 1 nodes

The leaf reached first may be infeasible. On the knapsack the "in" child is tried first and may exceed the capacity. The "out" child forces nothing new, and its parent's forced items fit, since the parent's relaxation was feasible, so the "out" child always has a feasible relaxation. A dive therefore spends at most two evaluations per level, and the depth is at most \(n\). It reaches a feasible leaf, hence an incumbent, within \(2n + 1\) node evaluations. On the figure's instance the first incumbent arrives at node 19. The knapsack figure shows the two propositions in one instance. Best-bound processes 37 nodes, the fewest, and finds its first incumbent only at node 10. Depth-first finds an incumbent at node 19 and then spends 296 more nodes proving optimality, 83 of them pruned by bound and 70 infeasible. Dive-first takes the depth-first dive to its first incumbent and then switches, for 53 nodes.

Solvers choose a hybrid. SCIP's default node selector is best-estimate with plunging. After a node is processed, a child is selected as long as two conditions hold: the plunge is shallower than a dynamic depth limit, and the child's bound is within a quarter of the current gap above the global bound. Otherwise every tenth selection takes the node with the smallest bound and the others take the best estimate.SCIP 10.0.0 source, src/scip/nodesel_estimate.c (priority 200000, maxplungequot 0.25, bestnodefreq 10) and set.c, github.com/scipopt/scip. The design is T. Achterberg, "SCIP: solving constraint integer programs", Mathematical Programming Computation 1 (2009). BARON's NodeSel option offers its own strategy, best bound, LIFO (last in, first out, which is depth-first) and minimum infeasibilities.BARON User Manual, Section 11.4, "Tree management options" (NodeSel: 0 BARON's strategy, the default; 1 best bound; 2 LIFO; 3 minimum infeasibilities), The Optimization Firm, minlp.com/baron-user-manual, read 5 October 2026. Couenne dives from the best-bound node with Bonmin's probed-dive traversal. Section 3.6 lays SCIP's and BARON's node pipelines side by side. The cost of the two levers is easy to state. Best-first minimizes the number of relaxations up to ties and uses the most memory. Depth-first bounds memory, finds incumbents early and lowers the node count through pruning. Warm starts lower the cost per node and are cheapest along a dive. Berthold, Hendel and Koch make the hybrid explicit by phase: a feasibility phase until the first incumbent, an improvement phase, and a proof phase, each with its own node selection and heuristic schedule.T. Berthold, G. Hendel and T. Koch, "From feasibility to improvement to proof: three phases of solving mixed-integer programs", Optimization Methods and Software 33 (2018).

SCIP's default node selection: best estimate with plunging

  +----------> a node has been processed <-----------------+
  |                         |                              |
  |                         v                              |
  |   plunge shallower than the dynamic depth limit, and   |
  |   the child's bound within a quarter of the current    |
  |   gap above the global bound?                          |
  |           | yes                       | no             |
  |           v                           v                |
  +--- select a child:          every tenth selection: the |
       the plunge goes on       node with the smallest     |
                                bound; the others: the     |
                                best estimate -------------+

Branching

Definition 3.1.11 (pseudocosts, strong branching, reliability). For an integer variable \(x_j\) with relaxed value \(\bar x_j\) and fractionality \(f_j = \bar x_j - \lfloor \bar x_j \rfloor\), branching creates the children \(x_j \le \lfloor \bar x_j \rfloor\) and \(x_j \ge \lceil \bar x_j \rceil\) with relaxation values \(z^-_j\) and \(z^+_j\) and gains \(\Delta^-_j = z^-_j - \bar z(N)\), \(\Delta^+_j = z^+_j - \bar z(N)\). The unit gains \(\Delta^-_j / f_j\) and \(\Delta^+_j / (1 - f_j)\), averaged over all past branchings on \(x_j\), are the pseudocosts \(\Psi^-_j\) and \(\Psi^+_j\). The pseudocost estimates of the gains at the current node are \(\Psi^-_j f_j\) and \(\Psi^+_j (1 - f_j)\).M. Bénichou, J. M. Gauthier, P. Girodet, G. Hentgès, G. Ribière and O. Vincent, "Experiments in mixed-integer linear programming", Mathematical Programming 1 (1971). Strong branching computes \(\Delta^\pm_j\) by solving the two child relaxations, possibly with an iteration limit, before choosing.D. Applegate, R. Bixby, V. Chvátal and W. Cook, "Finding cuts in the TSP (a preliminary report)", DIMACS Technical Report 95-05 (1995); the account in their book, The Traveling Salesman Problem: A Computational Study (Princeton University Press, 2007), is where the name is explained. A score combines the two gains. SCIP's default is the product \(\max(\Delta^-_j, \delta) \cdot \max(\Delta^+_j, \delta)\) with \(\delta = 10^{-6}\), a floor on the gains in SCIP's code and not the optimality tolerance \(\varepsilon\), the alternative being the weighted sum \((1 - \mu)\min(\Delta^-_j, \Delta^+_j) + \mu \max(\Delta^-_j, \Delta^+_j)\) with \(\mu = 1/6\). Reliability branching uses strong branching on a variable until it has been branched on at least \(\eta_{\mathrm{rel}}\) times in each direction and its pseudocosts afterwards.

(Three facts about branching rules) Three facts organize the subject, all from Achterberg, Koch and Martin's 2005 comparison on the MIPLIB instances of the time.T. Achterberg, T. Koch and A. Martin, "Branching rules revisited", Operations Research Letters 33 (2005); the full tables are in ZIB-Report 04-13 (2004). The paper's term for branching on the variable closest to one half is "most infeasible branching". Branching on the most fractional variable, the one closest to one half, is "basically as good as random branching". Full strong branching produces the smallest trees and loses badly on time, since it solves two LPs for every fractional candidate at every node. Reliability branching consistently beat the hybrid strong/pseudocost rule that was then the state of the art, at every parameter setting the authors tried. It became the default in SCIP, where the rule relpscost has the highest priority of the branching rules. The best setting in time was a reliability threshold of \(\eta_{\mathrm{rel}} = 8\) with a lookahead of 4, the lookahead stopping the strong-branching loop after four consecutive candidates without a new best score. SCIP's current defaults differ from the paper's best setting and are given in the sidenote to the next paragraph. Le Bodic and Nemhauser later explained why fractionality carries no information about gains. Consider an abstract model in which every branching raises the bound by \(\ell\) in one child and by \(r \ge \ell\) in the other. The tree needed to close a gap \(G\) then grows like \(\varphi^G\), where \(\varphi\) is the root of \(\varphi^{-\ell} + \varphi^{-r} = 1\). The right score for a large gap is therefore that growth rate, and the product is its proxy when the gap is small.P. Le Bodic and G. Nemhauser, "An abstract model for branching and its application to mixed integer programming", Mathematical Programming 166 (2017). The product score is Achterberg's: T. Achterberg, Constraint Integer Programming, PhD thesis, Technische Universität Berlin (2007), Section 5.4; the hybrid of pseudocosts with inference, cutoff and conflict histories is T. Achterberg and T. Berthold, "Hybrid branching", CPAIOR 2009, LNCS 5547 (2009).

(Strong branching on R1) The two-variable example shows what strong branching buys. Because the example maximizes, the gains are read as \(\bar z(N) - z^\pm_j\), the amount by which a child's bound falls. For the objective \(x + y\) (the figure's \(45^\circ\)) the root vertex is \((69/14, 41/14)\) with value \(55/7 = 7.8571\) and both coordinates have fractional part \(13/14\). Branching on \(x\) gives the children \(x \le 4\), with bound \(7.3\), and \(x \ge 5\), with bound \(7.75\): gains \(0.5571\) and \(0.1071\), product score \(0.0597\), unit pseudocosts \(\Psi^- = 0.6\) and \(\Psi^+ = 1.5\). By the near-symmetry of the polygon branching on \(y\) gives the same numbers. For the objective \(11x + 4y\) (an angle of \(19.98^\circ\), the figure's \(20^\circ\) case to within rounding) the root is \((347/60, 19/24)\) with value \(66.78\). The child \(x \ge 6\) is infeasible, so the solver does not branch on \(x\) at all but tightens \(x \le 5\) at the node and re-solves. The children of \(y\) have gains \(20.58\) and \(0.08\). A rule that read fractionalities alone would have seen \(0.78\) against \(0.79\) and learned nothing. Deeper in the tree most variables are reliable and the rule reduces to pseudocost branching at a cost linear in the number of candidates. At the root, with hundreds of fractional candidates, strong branching dominates the cost of the node. SCIP caps it at half the LP iterations of the run so far, plus a fixed offset. The child LPs are independent of one another, which makes strong branching the most naturally parallel component of the engine. A batch of \(2|C|\) LPs that differ from the parent by one bound each is the shape a GPU bound kernel wants. cuOpt offers an option to run its root strong branching as one batched first-order solve for exactly this reason (Section 7.4).SCIP 10.0.0 source, src/scip/branch_relpscost.c (minreliable 1, maxreliable 5, maxlookahead 9, sbiterquot 0.5, initcand 100). Two refinements are in the SCIP 10 report: ancestral pseudocosts and a probabilistic lookahead that stops strong branching when a tree-size model says further candidates are unlikely to win; C. Hojny et al., "The SCIP Optimization Suite 10.0", arXiv 2511.18580 (2025), Section 3.7. G. Gamrath, T. Berthold and D. Salvagnin, "An exploratory computational analysis of dual degeneracy in mixed-integer programming", EURO Journal on Computational Optimization 8 (2020), explain why many candidate gains are zero at a dual-degenerate root, one at which many reduced costs vanish and the LP has many optimal vertices. The batched strong branching is the cuOpt release note of 26.02 (11 February 2026), "Added an option to use batch PDLP when running strong branching at the root", NVIDIA cuOpt User Guide, release notes, docs.nvidia.com/cuopt.

Strong branching at the root of R1, for two objectives

  x + y (45 degrees): root vertex (69/14, 41/14), value
  55/7 = 7.8571, both fractional parts 13/14; branching on x:

                         root 7.8571
                 x <= 4 /          \ x >= 5
                       /            \
                      /              \
             bound 7.3                bound 7.75
             gain 0.5571              gain 0.1071

    (the example maximizes: a gain is the fall of the bound)
    Psi- = 0.5571 / (13/14) = 0.6
    Psi+ = 0.1071 / (1 - 13/14) = 1.5
    product score 0.5571 x 0.1071 = 0.0597; y gives the same

  11x + 4y (19.98 degrees): root (347/60, 19/24), value 66.78

    fractional parts 0.78 (x) and 0.79 (y): no information
    x: the child x >= 6 is infeasible, so no branching on x;
       x <= 5 is tightened at the node, which is re-solved
    y: the two children have gains 20.58 and 0.08

The warm start

(Bases, basic solutions and reduced costs) The reason a child LP costs a few pivots is a fact about the dual simplex method. The same fact decides where a GPU can and cannot help in the tree today. Section 2.2 defined the reduced cost, in (2.2.1), and Section 2.6 the warm start, the pivot and the dual simplex method in one sentence each. The fuller account that those sentences rest on, and that Proposition 3.1.12 needs, is the following. Section 7.1 gives the method's history. Write the node LP in bounded-variable standard form, \(\min\{c^\top x : Ax = b,\ l \le x \le u\}\), after adding a slack variable to every row, so that \(A\) has \(m\) rows and \(n > m\) columns. A basis is a choice of \(m\) columns of \(A\) that form a nonsingular matrix \(B\). The other \(n - m\) columns form the nonbasic set \(\mathcal N\). A basic solution fixes every nonbasic variable at one of its two bounds, its bound status, and solves the rows for the basic variables:

\[x_B \;=\; B^{-1}\big(b - A_{\mathcal N}\, x_{\mathcal N}\big).\]

Geometrically a basic solution is a vertex of the polyhedron: \(n\) constraints hold with equality, the \(m\) rows and the \(n - m\) bounds at which the nonbasic variables sit. A pivot exchanges one basic and one nonbasic variable and moves along an edge to a neighbouring vertex. The reduced costs of the nonbasic variables are

\[r_{\mathcal N} \;=\; c_{\mathcal N} - A_{\mathcal N}^\top y, \qquad y \;=\; B^{-\top} c_B,\]

which is the vector \(r = c - A^\top y\) of (2.2.1) at the prices \(y\) that make the reduced costs of the basic variables zero. Each \(r_j\) is the change in the objective per unit move of the nonbasic variable \(x_j\) away from its bound.

(Primal and dual feasibility, and one pivot on R1) The basis is primal feasible if \(l_B \le x_B \le u_B\). It is dual feasible if every reduced cost has the sign that makes leaving the bound unprofitable: \(r_j \ge 0\) for a nonbasic \(j\) at its lower bound and \(r_j \le 0\) for a nonbasic \(j\) at its upper bound. A basis that is both is optimal. Dual feasibility has a meaning before optimality is reached. The objective value of a dual feasible basic solution is \(b^\top y + \sum_{j \in \mathcal N} r_j x_j\) with each \(x_j\) at its bound. This is the value of the dual function (2.2.1) at the prices \(y\), read for equality rows, whose prices carry no sign constraint. By weak duality, Theorem 2.2.2, it is a lower bound on the LP value. A dual feasible basis therefore carries a valid bound at every pivot, before it is optimal. The primal simplex keeps primal feasibility and improves the objective, walking downhill over feasible vertices. The dual simplex keeps dual feasibility and repairs primal infeasibility one bound at a time: a basic variable that violates a bound leaves the basis, and the dual ratio test chooses the entering variable so that every reduced cost keeps its required sign. The slack basis, in which the \(m\) slacks are basic and the original variables sit at bounds, is where a solve from scratch starts.V. Chvátal, Linear Programming (W. H. Freeman, 1983), for the revised simplex method, the bounded-variable form and the dual simplex method. The modern dual simplex, with steepest-edge pricing, bound flipping and sparse LU updates, is described in R. E. Bixby, "Solving real-world linear programs: a decade and more of progress", Operations Research 50 (2002). One pivot on the two-variable example shows the whole mechanism. At the root of the \(20^\circ\) run tabulated below, the basic variables are \(x\), \(y\) and the slacks of the rows \(2x + 5y \le 24.5\) and \(-3x + 4y \le 11\), and the slacks of the two binding rows \(5x + 2y \le 30.5\) and \(x - 2y \le 4.2\) are nonbasic at zero. The root vertex is \((5.7833, 0.7917)\). The child \(x \le 5\) changes the upper bound of the basic variable \(x\). The reduced costs do not change, so the basis is still dual feasible, and the only primal infeasibility is that \(x = 5.7833\) exceeds its new bound by \(0.7833\). The dual simplex removes \(x\) from the basis and fixes it at \(5\), and the ratio test brings in the slack of \(x - 2y \le 4.2\), the row that stops binding. The new basic solution is the vertex \((5, 2.75)\) of node 3 in the table, at which \(5x + 2y = 30.5\) still binds and the slack of the fourth row is \(4.7\). That is the single pivot the table records for node 3.

One dual simplex pivot: the R1 root to node 3 (theta = 20 degrees)

  slacks: s_1 of 2x + 5y <= 24.5     s_2 of 5x + 2y <= 30.5
          s_3 of -3x + 4y <= 11      s_4 of x - 2y <= 4.2

  root                               node 3, the child x <= 5
  basic     x, y, s_1, s_3           basic     y, s_1, s_3, s_4
  nonbasic  s_2 = 0, s_4 = 0         nonbasic  s_2 = 0, x = 5
  vertex    (5.7833, 0.7917)         vertex    (5, 2.75), s_4 = 4.7
     |                                  ^
     | new bound x <= 5: reduced costs  |
     | unchanged, still dual feasible;  |
     | x = 5.7833 exceeds it by 0.7833  |
     +----------------------------------+
       one pivot: x leaves the basis at its bound 5,
       the ratio test brings in s_4

Proposition 3.1.12 (the parent's optimal basis is dual feasible for the child). Let \((B, \mathcal N)\) with its bound statuses be an optimal basis of the parent LP, and let the child LP differ from the parent only in the bounds \(l, u\).

(a) The reduced costs \(r_{\mathcal N}\) depend on \((A, c, B)\) only, not on \(b\), \(l\) or \(u\), so the basis with the same bound statuses is dual feasible for the child.

(b) Let the changed bound belong to a basic variable \(x_j\), as the branching variable is, since its value lies strictly between integer bounds. Then \(x_B\) is unchanged and the basis is primal infeasible in that variable alone, by the amount \(v - \lfloor v \rfloor\) or \(\lceil v \rceil - v\).

(c) If the changed bound belongs to a nonbasic variable at that bound, the variable moves to the new bound value and \(x_B\) changes by \(-B^{-1} A_j \Delta\) for the shift \(\Delta\). The basis stays dual feasible and may lose primal feasibility in several basic variables.

In every case the dual simplex method can start from the parent's basis. Each of its pivots keeps dual feasibility and does not decrease the dual objective, which is at every pivot a valid lower bound on the child's LP value.

Proof. (a) is the formula for \(r_{\mathcal N}\), in which neither \(b\) nor the bounds appear. The sign conditions refer to the bound each nonbasic variable sits at, which the change of a basic variable's bound does not alter. (b) \(x_B = B^{-1}(b - A_{\mathcal N} x_{\mathcal N})\) does not involve the bounds of basic variables, so it is unchanged, and the only violated bound is the new one on \(x_j\). (c) \(x_{\mathcal N}\) changes in one component by \(\Delta\), so \(x_B\) changes by \(-B^{-1} A_j \Delta\). The sign conditions on \(r_{\mathcal N}\) are unchanged because the variable still sits at a bound of the same kind. The last sentence is the invariant of the dual simplex method. The dual ratio test chooses the entering variable so that every reduced cost keeps its required sign, and the dual objective is nondecreasing along dual pivots, equal to the LP value at termination. ∎

(Why the dual simplex is the tree algorithm) Adding a cut row to the parent LP also preserves dual feasibility, because the new row's slack enters the basis with reduced cost zero. The cut loop of Section 3.3 and the branching therefore both re-solve by dual simplex from the current basis. This is why every LP-based solver sets the dual simplex as its tree algorithm.SCIP: lp/initalgorithm = 's' and lp/resolvealgorithm = 's' in set.c. On the two-variable program the proposition can be measured. The program below runs branch and bound at three objective angles with a bounded-variable dual simplex, using the textbook ratio test and no bound flipping. It solves every node LP three ways: warm from the parent's optimal basis, from the root's optimal basis, and cold from the slack basis with \(x\) and \(y\) at their upper bounds. The root solve is counted in all three columns. It prints the pivot totals over each tree and the per-node counts at \(20^\circ\).

# The dual simplex warm start inside branch and bound on R1.
#
# Maximize cos(th) x + sin(th) y over the polygon with x, y integer in
# [0, 7] x [0, 6]. A bounded-variable dual simplex (textbook ratio test,
# no bound flipping) solves every node LP three ways: warm from the
# parent's optimal basis, from the root's optimal basis, and cold from
# the slack basis with x and y at their upper bounds (dual feasible,
# since both reduced costs are negative). Node selection and branching
# follow fig-vocab: best bound, oldest among equals, the more fractional
# variable, the ">=" child first.

import textwrap

import numpy as np

# The four rows of R1, each with its slack: columns x, y, s1, s2, s3, s4.
A = np.array([[2, 5, 1, 0, 0, 0],
              [5, 2, 0, 1, 0, 0],
              [-3, 4, 0, 0, 1, 0],
              [1, -2, 0, 0, 0, 1]], float)
b = np.array([24.5, 30.5, 11.0, 4.2])
INF = float("inf")

def dual_simplex(c, l, u, basis, at_upper):
    """Min c.x, A x = b, l <= x <= u, from a dual feasible basis.

    at_upper tells, for each nonbasic variable, whether it sits at its
    upper bound. Returns (status, x, basis, at_upper, pivots).
    """
    basis, at_upper, pivots = list(basis), dict(at_upper), 0
    while True:
        N = [j for j in range(6) if j not in basis]
        B = A[:, basis]
        xN = np.array([u[j] if at_upper[j] else l[j] for j in N])
        xB = np.linalg.solve(B, b - A[:, N] @ xN)

        # reduced costs: no b, l, u in them
        d = c[N] - A[:, N].T @ np.linalg.solve(B.T, c[basis])
        assert all(d[k] <= 1e-7 if at_upper[j] else d[k] >= -1e-7
                   for k, j in enumerate(N)), "dual feasibility lost"

        viol = np.array([max(l[j] - xB[i], xB[i] - u[j], 0.0)
                         for i, j in enumerate(basis)])
        r = int(np.argmax(viol))
        if viol[r] <= 1e-9:
            x = np.zeros(6)
            x[basis] = xB
            x[N] = xN
            return "optimal", x, basis, at_upper, pivots

        # row r of B^-1 A_N
        alpha = A[:, N].T @ np.linalg.solve(B.T, np.eye(4)[r])
        # x_r must rise to l_r or fall to u_r
        s = 1.0 if xB[r] < l[basis[r]] else -1.0
        elig = [k for k, j in enumerate(N)
                if s * (1 if at_upper[j] else -1) * alpha[k] > 1e-9]
        if not elig:
            # no entering variable: LP infeasible
            return "infeasible", None, basis, at_upper, pivots

        # the dual ratio test keeps every sign
        k = min(elig, key=lambda k: abs(d[k]) / abs(alpha[k]))
        leaving = basis[r]
        basis[r] = N[k]
        del at_upper[N[k]]
        at_upper[leaving] = s < 0
        pivots += 1

def print_row(node, bounds, val, what, pw, pr, pc):
    """One node of the table; a long status wraps onto a second line."""
    *first, last = textwrap.wrap(what, 23)
    head = f"{node:>4}  {bounds:<16}  {val}  "
    for part in first:
        print(head + part)
        head = " " * len(head)
    print(f"{head}{last:<23}  {pw:>4}  {pr:>4}  {pc:>4}")

def branch_and_bound(th_deg, table):
    """Best-bound search on R1 at th_deg degrees, counting pivots."""
    th = np.radians(th_deg)
    # maximize = minimize -c
    c = -np.array([np.cos(th), np.sin(th), 0, 0, 0, 0])
    # the slack basis
    cold = ([2, 3, 4, 5], {0: True, 1: True})
    st, x, bas, up, p0 = dual_simplex(c, np.zeros(6),
                                      np.array([7, 6, INF, INF, INF, INF]),
                                      *cold)
    rootstart = (bas, up)
    root = dict(l=np.zeros(6), u=np.array([7, 6, INF, INF, INF, INF]),
                start=cold, pbound=INF, order=0)
    inc, open_, nodes = -INF, [root], 0
    tot = dict(warm=0, root=0, cold=0)
    if table:
        print(f"theta = {th_deg} deg")
        print(" " * 63 + "pivots")
        print(f"node  {'bounds':<16}  {'value':>7}  {'status':<23}"
              f"  warm  root  cold")

    while open_:
        i = max(range(len(open_)),
                key=lambda i: (open_[i]["pbound"], -open_[i]["order"]))
        nd = open_.pop(i)
        nodes += 1

        if nd is root:
            st, pw = "optimal", p0
            pr = pc = p0
        else:
            st, x, bas, up, pw = dual_simplex(c, nd["l"], nd["u"],
                                              *nd["start"])
            s2, x2, _, _, pr = dual_simplex(c, nd["l"], nd["u"], *rootstart)
            s3, x3, _, _, pc = dual_simplex(c, nd["l"], nd["u"], *cold)
            # the three starts agree on the status and on the value
            assert s2 == s3 == st
            if st == "optimal":
                assert abs(c @ x - c @ x2) + abs(c @ x - c @ x3) < 1e-7
        tot["warm"] += pw
        tot["root"] += pr
        tot["cold"] += pc

        if st == "infeasible":
            what, val = "LP infeasible", "   --  "
        else:
            z = -c @ x
            val = f"{z:7.4f}"
            if z <= inc + 1e-9:
                what = f"pruned by bound (incumbent {inc:.4f})"
            elif all(abs(x[j] - round(x[j])) < 1e-7 for j in (0, 1)):
                what = "integral: new incumbent"
                inc = z
            else:
                # the more fractional of x and y: farther from an integer
                dist = [abs(x[k] - round(x[k])) for k in (0, 1)]
                j = 0 if dist[0] >= dist[1] else 1
                v = x[j]
                what = f"branch on {'xy'[j]} = {v:.4f}"
                # side 0 is the ">=" child, side 1 the "<=" child
                sides = [(np.ceil(v), nd["u"][j]),
                         (nd["l"][j], np.floor(v))]
                for side, (lo, hi) in enumerate(sides):
                    l2, u2 = nd["l"].copy(), nd["u"].copy()
                    l2[j], u2[j] = lo, hi
                    child = dict(l=l2, u=u2, start=(bas, up), pbound=z,
                                 order=len(open_) + 100 * nodes + side)
                    open_.append(child)

        bounds = (f"{nd['l'][0]:g}<=x<={nd['u'][0]:g}, "
                  f"{nd['l'][1]:g}<=y<={nd['u'][1]:g}")
        if table:
            print_row(nodes, bounds, val, what, pw, pr, pc)
    return nodes, tot

print("angle   nodes   pivots, warm from the parent   "
      "from the root basis   cold")
for th in (20, 45, 70):
    n, t = branch_and_bound(th, False)
    print(f"{th:>3} deg {n:>7} {t['warm']:>22} {t['root']:>21} "
          f"{t['cold']:>6}")
print()
branch_and_bound(20, True)
angle   nodes   pivots, warm from the parent   from the root basis   cold
 20 deg       7                      8                    13     17
 45 deg      11                     10                    14     17
 70 deg       9                      9                    12     18

theta = 20 deg
                                                               pivots
node  bounds              value  status                   warm  root  cold
   1  0<=x<=7, 0<=y<=6   5.7053  branch on x = 5.7833        3     3     3
   2  6<=x<=7, 0<=y<=6     --    LP infeasible               0     0     3
   3  0<=x<=5, 0<=y<=6   5.6390  branch on y = 2.7500        1     1     2
   4  0<=x<=5, 3<=y<=6   5.4896  branch on x = 4.7500        2     2     4
   5  0<=x<=5, 0<=y<=2   5.3825  integral: new incumbent     1     2     0
   6  5<=x<=5, 3<=y<=6     --    LP infeasible               0     2     4
   7  0<=x<=4, 3<=y<=6   4.8874  pruned by bound
                                 (incumbent 5.3825)          1     3     1

(Reading the pivot counts) At \(20^\circ\) the child \(x \ge 6\) is shown infeasible in zero pivots, because the ratio test finds no entering variable. The child \(x \le 5\) is solved in one pivot, the new bound on the basic variable \(x\) being its only primal infeasibility. The cold start sometimes wins a node, as at node 5, where the tightened bounds make the corner of the box feasible and the slack basis is already optimal. On a problem with four rows the cold start loses by a factor of about two over the whole tree. The factor grows with the number of rows, since a cold dual simplex must rebuild the active set from scratch while the warm start inherits it. The pivots are the sequential part of the engine. Each one is a backward solve with the basis factors, a row computation, a ratio test, a forward solve and a factor update, each depending on the one before. A first-order method has no basis to inherit. Its analogue of this proposition, a warm start from the parent's primal–dual iterate with a safe bound at every iterate, is the subject of Sections 7.3 and 7.4.

Branch and bound can be exponential

(Jeroslow's parity instance) No upper bound on the size of the tree follows from the theorems above, and Theorem 3.1.14 shows that none can. The simplest demonstration is Jeroslow's, from 1974. Let \(n\) be odd, write \(m = (n - 1)/2\), and consider the feasibility problem

\[2 \sum_{i=1}^{n} x_i \;=\; n, \qquad x \in \{0, 1\}^n . \tag{J}\]

It has no solution, because its left side is even and its right side odd. Its LP relaxation is feasible, at \(x_i = 1/2\) for every \(i\) and at every other point of the hyperplane \(\sum_i x_i = n/2\) inside the cube. Jeroslow's own program minimizes an extra variable \(x_{n+1}\) subject to \(2\sum_{i \le n} x_i + x_{n+1} = n\) over binaries. Its optimum is 1 and the LP value is 0. The child \(x_{n+1} \le 0\) of any tree is exactly (J), so what follows about (J) is a lower bound on proving optimality in his program.R. G. Jeroslow, "Trivial integer programs unsolvable by branch-and-bound", Mathematical Programming 6 (1974). The program is quoted here in the form given by B. Krishnamoorthy and G. Pataki, "Column basis reduction and decomposable knapsack problems", Discrete Optimization 6 (2009), equation (1.10), whose Corollary 4 proves the bound \(2^{(n-1)/2}\) on the number of nodes for (J). The bound \(2^{(n+1)/2}\) on the leaves stated in Theorem 3.1.14 is the one proved below; the exponent printed in Jeroslow's paper was not checked for this series, because the text was not accessible. The exact count is Proposition 3.1.15.

Lemma 3.1.13 (LP feasibility at a node of (J)). Let a node of a branch-and-bound tree for (J), branching on single variables, have \(a\) variables fixed to 1, \(b\) variables fixed to 0 and \(f = n - a - b\) free. Its LP relaxation is feasible if and only if \(a \le m\) and \(b \le m\).

Proof. The free variables must satisfy \(\sum_{\text{free}} x_i = n/2 - a\) with each \(x_i \in [0, 1]\), which is possible if and only if \(0 \le n/2 - a \le f\). The left inequality is \(a \le n/2\), that is \(a \le m\) since \(a\) is an integer and \(n\) is odd. The right inequality is \(n/2 - a \le n - a - b\), that is \(b \le n/2\), that is \(b \le m\). ∎

Theorem 3.1.14 (Jeroslow 1974). Every branch-and-bound tree that proves the infeasibility of (J) by branching on single variables, in any order, and pruning only by LP infeasibility has at least \(2^{(n+1)/2}\) leaves and at least \(2^{(n+1)/2 + 1} - 1\) nodes.

Proof. A branching on a free variable fixes it to 0 in one child and to 1 in the other, so a node at depth \(d\) has \(a + b = d\) fixed variables. By Lemma 3.1.13 a node is a leaf, that is LP infeasible, only if \(a \ge m + 1\) or \(b \ge m + 1\), hence only if \(d \ge m + 1 = (n+1)/2\). There is no objective, so no node is pruned by bound, and no relaxed point is integral, because \(\sum_i \bar x_i = n/2\) is not an integer, so every non-leaf is branched. In a full binary tree the leaves satisfy Kraft's equality \(\sum_{\text{leaves } v} 2^{-\operatorname{depth}(v)} = 1\), by induction on the tree: replacing a leaf at depth \(d\) by two children at depth \(d + 1\) preserves the sum. With every leaf at depth at least \((n+1)/2\) this gives \(\ell \cdot 2^{-(n+1)/2} \ge 1\) for the number \(\ell\) of leaves, so \(\ell \ge 2^{(n+1)/2}\), and the node count is \(2\ell - 1\). ∎

The theorem is a statement about every branching order and every selection rule at once, and for this instance the tree can be counted exactly.

Proposition 3.1.15 (the exact size of the tree). Every tree of Theorem 3.1.14 has exactly \(\ell_n = \binom{n+1}{(n+1)/2}\) leaves and \(T_n = 2\binom{n+1}{(n+1)/2} - 1\) nodes, whatever the order in which variables are branched on and whatever the selection rule. Writing \(T(a, b)\) for the size of the subtree below a node with counts \((a, b)\),

\[T(a, b) \;=\; 1 \ \text{ if } a > m \text{ or } b > m, \qquad T(a, b) \;=\; 1 + T(a + 1, b) + T(a, b + 1) \ \text{ otherwise}, \qquad T_n = T(0, 0).\]

Proof. By Lemma 3.1.13 the status of a node depends only on \((a, b)\), and branching on any free variable sends \((a, b)\) to \((a + 1, b)\) and \((a, b + 1)\). The shape of the tree is therefore the same for every choice of branching variables, and the subtree sizes satisfy the recursion. For the closed form, read a root-to-leaf path as a 0/1 string in the order the variables were fixed. A leaf with \(m + 1\) ones and \(j \le m\) zeros ends in a 1, and its first \(m + j\) symbols contain exactly \(m\) ones and \(j\) zeros in any order. There are therefore \(\binom{m+j}{j}\) such leaves. Summing over \(j\) gives \(\sum_{j=0}^{m} \binom{m+j}{j} = \binom{2m+1}{m}\) by the hockey-stick identity. Doubling for the leaves with \(m + 1\) zeros, \(\ell_n = 2\binom{2m+1}{m} = 2\binom{n}{m} = \binom{n}{m} + \binom{n}{m+1} = \binom{n+1}{m+1}\), using \(\binom{n}{m+1} = \binom{n}{n-m-1} = \binom{n}{m}\). A full binary tree with \(\ell_n\) leaves has \(2\ell_n - 1\) nodes. ∎

Every parity tree for n = 5 (m = 2), folded onto its states (a, b)

  a = variables fixed to 1, b = variables fixed to 0; branching
  on any free variable sends (a, b) to (a + 1, b) and (a, b + 1)

            b = 0       b = 1       b = 2       b = 3
  a = 0    (0,0) ----> (0,1) ----> (0,2) ----> [0,3]
             |           |           |
             v           v           v
  a = 1    (1,0) ----> (1,1) ----> (1,2) ----> [1,3]
             |           |           |
             v           v           v
  a = 2    (2,0) ----> (2,1) ----> (2,2) ----> [2,3]
             |           |           |
             v           v           v
  a = 3    [3,0]       [3,1]       [3,2]

  (a,b)  a <= m and b <= m: LP feasible, branched
  [a,b]  a = m + 1 or b = m + 1: LP infeasible, a leaf
  a tree node is a path from (0,0) in this grid: 39 nodes, 19 of
  them branched and 20 leaves, ten in the row a = 3 and ten in
  the column b = 3; the shallowest leaves are at depth 3

(The parity tree in the figure) For \(n = 5\) the tree has 39 nodes: 20 leaves, ten at which a third one arrives before a third zero and ten the other way round, and 19 branched nodes. The shallowest leaves are at depth 3, as Theorem 3.1.14 requires. The recursion is what the figure evaluates. The figure's parity instance is reached through its instance control or through the two case buttons, "Jeroslow parity, \(n = 7\)" and "parity with the Chvátal–Gomory cut". A Chvátal–Gomory cut is obtained from the rows \(Ax \le b\) of a relaxation in three steps: multiply them by a nonnegative vector \(u\), round every coefficient of \(u^\top A\) down to an integer, and round the right-hand side \(u^\top b\) down as well. The result is valid for every integer point with \(x \ge 0\), which is Theorem 3.3.3 of Section 3.3, where the family is treated. Proposition 3.1.16 applies it to (J). At \(n = 7\) the program \(2(x_1 + \dots + x_7) = 7\) has no binary solution, yet the LP is feasible at every node with at most three variables fixed to 1 and at most three fixed to 0. Branching on \(x_1, x_2, \dots\) in index order gives a tree of 139 nodes against 255 in the full tree, a share of 54.5 %, with 69 nodes branching and 70 LP-infeasible leaves. Under depth-first search the \(x = 1\) child of each node is visited before the \(x = 0\) child. Under best-bound every bound is equal, the oldest open node is taken, the run is breadth-first and the \(x = 0\) child comes first. All orders visit the same 139 nodes. The figure's side panel plots the exact size against \(2^{n+1} - 1\) for \(n = 3, 5, \dots, 15\) on a log scale. Its readout lists the seven pairs 11/15, 39/63, 139/255, 503/1,023, 1,847/4,095, 6,863/16,383 and 25,739/65,535. For \(n = 13\) and \(15\) the drawing stops at 3,000 nodes and the count comes from the recursion. By Stirling's formula \(\binom{2k}{k} \sim 4^k/\sqrt{\pi k}\), so the share of the full tree is asymptotically \(\sqrt{8/(\pi(n+1))}\), which is the 56.4 % the figure's take quotes next to the exact 54.5 % at \(n = 7\). The tree is a vanishing fraction of the full tree and still exponential in \(n\). A short program prints the seven counts and the shares.

# The size of every branch-and-bound tree for Jeroslow's parity instance.
#
# The instance is 2(x_1 + ... + x_n) = n, x binary, n odd (fig-bnb's
# second instance), and the trees branch on single variables and prune
# only by LP infeasibility. The program evaluates the recursion T(a, b),
# checks it against the closed form 2 C(n+1, (n+1)/2) - 1, and prints
# the share of the full tree next to its asymptote sqrt(8/(pi(n+1))).

from functools import lru_cache
from math import comb, pi, sqrt

def tree_size(n):
    """T(a, b) = 1 + T(a+1, b) + T(a, b+1) while a, b <= (n-1)/2, else 1.

    a and b count the variables fixed to 1 and to 0; the tree has
    T(0, 0) nodes.
    """
    m = (n - 1) // 2

    @lru_cache(None)
    def T(a, b):
        return 1 if a > m or b > m else 1 + T(a + 1, b) + T(a, b + 1)

    return T(0, 0)

print(" n  nodes  2C(n+1,(n+1)/2)-1  leaves  full tree  share"
      "  sqrt(8/(pi(n+1)))")
for n in range(3, 16, 2):
    T, L, full = tree_size(n), comb(n + 1, (n + 1) // 2), 2 ** (n + 1) - 1
    assert T == 2 * L - 1
    share = 100 * T / full
    asymptotic = 100 * sqrt(8 / (pi * (n + 1)))
    print(f"{n:2d}  {T:5d}  {2 * L - 1:17d}  {L:6d}  {full:9d}"
          f"  {share:4.1f}%  {asymptotic:16.1f}%")
 n  nodes  2C(n+1,(n+1)/2)-1  leaves  full tree  share  sqrt(8/(pi(n+1)))
 3     11                 11       6         15  73.3%              79.8%
 5     39                 39      20         63  61.9%              65.1%
 7    139                139      70        255  54.5%              56.4%
 9    503                503     252       1023  49.2%              50.5%
11   1847               1847     924       4095  45.1%              46.1%
13   6863               6863    3432      16383  41.9%              42.6%
15  25739              25739   12870      65535  39.3%              39.9%

The program costs nothing. One Chvátal–Gomory cut, or one disjunction, proves the infeasibility at the root.

Proposition 3.1.16 (one cut, or one disjunction, closes the root). Multiplying the equation of (J), written as the pair \(2\sum_i x_i \le n\) and \(-2\sum_i x_i \le -n\), by \(1/2\) and rounding the right-hand sides down gives the two Chvátal–Gomory cuts

\[\sum_{i=1}^n x_i \;\le\; \Big\lfloor \frac n2 \Big\rfloor = m \qquad \text{and} \qquad \sum_{i=1}^n x_i \;\ge\; \Big\lceil \frac n2 \Big\rceil = m + 1,\]

each valid for every binary point. Either one alone makes the root LP infeasible, since the LP forces \(\sum_i x_i = n/2\), which lies strictly between \(m\) and \(m + 1\). Equivalently, the single disjunction \(\sum_i x_i \le m \ \vee\ \sum_i x_i \ge m + 1\) at the root has two LP-infeasible children and proves infeasibility with three nodes.

Proof. Both inequalities are Chvátal–Gomory cuts, valid by Theorem 3.3.3, whose proof is three lines. Here the coefficients \(u^\top A\) are already integral, so only the right-hand sides round. With \(u = 1/2\) on \(2\sum_i x_i \le n\) the right side rounds to \(m\). With \(u = 1/2\) on \(-2\sum_i x_i \le -n\) it rounds to \(\lfloor -n/2 \rfloor = -(m+1)\). The rest is the observation that \(\sum_i x_i = n/2\) on the relaxation. ∎

Proposition 3.1.16: one cut, or one disjunction, closes (J)

  2 sum x_i <= n                    -2 sum x_i <= -n
        | times 1/2, round                 | times 1/2, round
        v the right side down              v the right side down
  sum x_i <= m                      -sum x_i <= -(m + 1),
                                    that is sum x_i >= m + 1

  the root LP forces sum x_i = n/2, strictly between m and m + 1:

                 root, sum x_i = n/2
     sum x_i <= m /                \ sum x_i >= m + 1
                 /                  \
         LP infeasible          LP infeasible

  either cut alone: 1 node; the disjunction: 3 nodes; branching
  on single variables at n = 7: 139 nodes

(One cut or one disjunction replaces the tree) In the figure, switching the Chvátal–Gomory cut on makes the root infeasible and the tree one node, 0.4 % of the full tree at \(n = 7\), against 139 without the cut. The side panel marks the single node with a purple line. This is the simplest example in the series of a cut replacing a tree. The reading that matters is the second one, through the disjunction: a tree that may branch on \(\sum_i x_i\) instead of on single variables has three nodes. Krishnamoorthy and Pataki make the reading systematic for knapsack problems. Yang, Boland and Savelsbergh report that on Chvátal's hard knapsack class an LP-based branch and bound with a well-chosen multivariable branching "explores either three or seven nodes".B. Krishnamoorthy and G. Pataki, Discrete Optimization 6 (2009), Theorem 3: if the infeasibility of a binary knapsack is proved by one split disjunction \(p^\top x \le k \,\vee\, p^\top x \ge k + 1\) with \(p\) positive and integral, ordinary branch and bound needs at least \(2^{\ell}\) nodes for an explicit \(\ell\) depending on \((p, k)\), equal to \(k\) for \(p = \mathbf 1\) and \(k < n/2\). Y. Yang, N. Boland and M. Savelsbergh, "Multivariable branching: a 0–1 knapsack problem case study", INFORMS Journal on Computing (2021), abstract.

(How general the phenomenon is) Two further theorems say how general the phenomenon is. Chvátal showed in 1980 that it survives a wider class of algorithms. For a class of 0–1 knapsack instances, every algorithm that combines branch and bound, dynamic programming and "rudimentary divisibility arguments" needs time that "grows exponentially with the square root" of the input length.V. Chvátal, "Hard knapsack problems", Operations Research 28 (1980); the quotations are from the abstract. The paper's text could not be consulted for this series. Secondary accounts describe the class as knapsacks with large random coefficients and the result as holding for almost all instances of it; neither detail, nor the constant in the exponent, is quoted here from the paper. Krishnamoorthy and Pataki (2009) quote two deterministic families from it, due to Todd and to Avis, and note that a single knapsack cover cut proves their infeasibility. Dey, Dubey and Molinaro showed in 2023 that it survives even when the tree may branch on arbitrary split disjunctions, which closes the escape route of Proposition 3.1.16. In such a general branch-and-bound tree for a polytope \(P\), each node is \(P\) intersected with the half-spaces chosen on the path from the root, one side of each split disjunction. These half-spaces are the node's branching constraints, and a leaf is a node whose intersection is empty. Their simplest instance, the cross-polytope, has a short proof.

Theorem 3.1.17 (Dey, Dubey and Molinaro 2023). Let

\[P_n \;=\; \Big\{ x \in [0, 1]^n : \sum_{i \in J} x_i + \sum_{i \notin J} (1 - x_i) \ \ge\ \tfrac12 \quad \text{for all } J \subseteq \{1, \dots, n\} \Big\}.\]

\(P_n\) contains no integer point, and every general branch-and-bound tree that certifies this using split disjunctions with arbitrary integer \(\pi\) has at least \(2^n\) leaves, hence at least \(2^{n+1} - 1\) nodes.

Proof. A 0/1 point \(\hat x\) violates the inequality indexed by \(J = \{i : \hat x_i = 0\}\), whose left side is 0 at \(\hat x\), so \(P_n\) has no integer point. Let \(v\) be a leaf whose branching constraints are satisfied by two distinct points \(y\) and \(y'\) of \(\{0, 1\}^n\). Their midpoint satisfies the branching constraints, which are linear, and lies in \(\{0, \tfrac12, 1\}^n\) with at least one coordinate equal to \(\tfrac12\). Such a point satisfies every defining inequality of \(P_n\), since that coordinate contributes \(\tfrac12\) to each and the others contribute nonnegative amounts. So the leaf's region meets \(P_n\) and the leaf cannot be infeasible. Hence the branching constraints of every leaf admit at most one point of \(\{0, 1\}^n\). Every such point satisfies one side of every disjunction on its way down the tree, so the \(2^n\) points of \(\{0, 1\}^n\) need at least \(2^n\) leaves. ∎

(What the lower bounds do and do not say) The same paper gives exponential lower bounds for general trees on packing and set-covering polytopes, on the subtour relaxation of the travelling salesman problem, and on a cross-polytope whose coefficients carry independent Gaussian noise. The last of these shows that smoothed analysis, the explanation of the simplex method's practical speed, cannot give a polynomial bound for branch and bound.S. S. Dey, Y. Dubey and M. Molinaro, "Lower bounds on the size of general branch-and-bound trees", Mathematical Programming 198 (2023); arXiv 2103.09807. The theorem above is their Proposition 3, which improves the \(2^n/n\) of Dadush and Tiwari (2020); the perturbed instance is their Theorem 2, the TSP bound \(2^{n/16 - 2}\) their Corollary 8. Their introduction notes that Jeroslow's and Chvátal's instances "can be solved with small (polynomial-size) general branch-and-bound trees". In the other direction, S. S. Dey, Y. Dubey and M. Molinaro, "Branch-and-bound solves random binary IPs in poly(n)-time", Mathematical Programming 200 (2023), show that simple branch and bound explores polynomially many nodes with good probability on random binary programs with a fixed number of constraints. Cuts do not always win either: outside the 0/1 setting, Basu, Conforti, Di Summa and Jiang give instances on which branch and bound finishes in constant time while cutting planes never finish.A. Basu, M. Conforti, M. Di Summa and H. Jiang, "Complexity of branch-and-bound and cutting planes in mixed-integer optimization", Mathematical Programming 198 (2023), abstract: for convex 0/1 problems with variable disjunctions cutting planes do at least as well as branch and bound, with stable-set instances where the gap is exponential, but "if one moves away from 0/1 sets, this advantage of CP over BB disappears; there are examples where BB finishes in \(O(1)\) time, but CP takes infinitely long to prove optimality".

(The consequence for a faster machine) Branch and bound is therefore not exponential on every instance, but it is on explicit families, for every processing order. The consequence for the parallel programme of this series is quantitative. On the parity instance the node count grows by a factor of about 3.7 for each step of two in \(n\), approaching 4. A machine that processes nodes \(10^3 \approx 2^{10}\) times faster, with the same tree, therefore solves instances larger by about ten variables and no more. The same arithmetic applies to every family in this subsection. A parallel machine divides the work by at most the number of workers. What changes the base of the exponential is the relaxation, the propagation, the cuts and the disjunctions branched on. That is why the rest of the series is about those. The computational literature on branching on general disjunctions reports smaller trees on structured instances and an open question about the cost per node. No production solver branches on them by default. The practical substitute is the cut loop of Section 3.3, whose Gomory cuts are the inequalities that a split disjunction implies for the LP, with Proposition 3.1.16 the case \(\pi = \mathbf 1\).M. Karamanov and G. Cornuéjols, "Branching on general disjunctions", Mathematical Programming 128 (2011); G. Cornuéjols, L. Liberti and G. Nannicini, "Improved strategies for branching on general disjunctions", Mathematical Programming 130 (2011). Dey, Dubey and Molinaro (2023, p. 3) note that "most commercial solvers" use simple branching. The modern survey of the search, branching and pruning literature is D. R. Morrison, S. H. Jacobson, J. J. Sauppe and E. C. Sewell, "Branch-and-bound algorithms: a survey of recent advances in searching, branching, and pruning", Discrete Optimization 19 (2016).

Where this is used

Every MILP and MINLP solver in this series runs Algorithm 3.1.4 with the components of Sections 2.5, 2.6, 3.2 and 3.3 inserted at the steps marked above. The relaxation at step 5 is a warm-started dual simplex solve in every LP-based code, which is Proposition 3.1.12 in production. The exceptions are the NLP-based convex solvers of Section 3.4, which warm-start interior-point or SQP methods less effectively, and the GPU first-order experiments of Section 7. Node selection is a hybrid everywhere. Branching is reliability branching on integer variables in the MILP engines and the violation-based scores of Section 3.5 on continuous ones. The pruning test uses the three tolerances of Section 3.5.

What parallelizes

Given the incumbent value, the processing of a node depends on nothing but the node itself. This is the fact on which every parallel branch and bound rests, and it is worth stating with its one subtlety.

Proposition 3.1.18 (node independence and asynchronous incumbents). (a) In Algorithm 3.1.4 the processing of a node (relaxation, pruning tests, branching) depends only on the node's data and on the incumbent value used in the pruning test. Two nodes neither of which is an ancestor of the other can be processed in either order or simultaneously with the same result for each. (b) Let several workers process open nodes concurrently, each using for its pruning test some value \(\tilde z\) that is the objective value of a feasible point: its own stale copy of the incumbent, or a global value received late. Then Theorem 3.1.5(a) and (b) hold at every moment for the union of the workers' open lists, with \(z_{\mathrm{inc}}\) the best value any worker has found. A stale \(\tilde z \ge z_{\mathrm{inc}}\) can only fail to prune a node that the sequential run would have pruned, never prune one it should not. (c) The run is complete when every worker's open list is empty and no node is in transit, and the final incumbent is then optimal by Theorem 3.1.6.

Proof. (a) is read off the steps of the algorithm, none of which consults another open node. (b) The proof of Theorem 3.1.5 used, in the pruned case, only that the pruning value is the value of a feasible point, so that \(f(x) \ge \bar z(N) \ge \tilde z \ge z^\star\) removes no improving point. The feasible and branched cases are local to the node. Lost pruning is the only effect of staleness. (c) is the correctness argument of Theorem 3.1.6 applied to the empty union of the open lists. ∎

(Stale incumbents, and what breaks the theorem) The subtlety is in (b). A stale incumbent value loses pruning but keeps correctness. A value that is not attained by a feasible point breaks Theorem 3.1.5(b). An implementation that prunes against a bound estimate, or against an incumbent whose feasibility was checked with a looser tolerance than the pruning test assumes, breaks the theorem silently. Everything else in parallel branch and bound follows from (a) to (c). A parallel run expands nodes that the sequential run would not, and speed-up anomalies result. Nodes are shared through a pool or by work stealing, incumbents are broadcast, a frontier of independent nodes is bounded in one batch on a GPU, and determinism has to be recovered. Sections 6.2 to 6.4 and 7.4 take these up.

Proposition 3.1.18 (b): two workers and a late incumbent

  worker 1  [N]--[N]--[N: new incumbent, value z_inc]
                       |
                       +--- broadcast ------------+
                                                  | received late
                                                  v
  worker 2  [N]--[N]--[N]--[N]--[N]--[N]--[N]--[N]--[N]--[N]--[N]
                       |<--- stale z~ >= z_inc -->|

  pruning with a stale z~, the value of a feasible point, worker 2
  can only fail to prune a node that the sequential run would have
  pruned, never prune one it should not; a z~ attained by no
  feasible point breaks Theorem 3.1.5(b). The run is complete
  when every open list is empty and no node is in transit.

For the GPU the lessons of this subsection are three. First, the tree's nodes are independent work items, so a frontier can be bounded in bulk. Second, the dual simplex warm start is the sequential kernel that a first-order method has to replace, bound by bound. Third, no amount of node throughput changes the exponent of a hard family. The bounding kernel, the propagation and the formulation are where the base of the exponential is decided.

Finding the incumbent

Every pruning rule in Section 3.1 compares a bound with \(z_{\mathrm{inc}}\), and until an incumbent exists no node is pruned by bound at all. The incumbent does three further jobs. The reduced-cost and duality-based range reductions of Section 2.6 shrink a box to a width proportional to \(z_{\mathrm{inc}} - \bar z(N)\). They are worthless before an incumbent exists and sharpen as it improves. The cutoff row \(f(x) \le z_{\mathrm{inc}}\) is what makes propagation with the incumbent effective. And the stopping rule of a solver is a gap between a bound and an incumbent. A run that has not found a feasible point cannot stop with a certificate at all. Berthold, Hendel and Koch's three phases (Section 3.1) are the solver's way of adapting its heuristics and node selection to this: feasibility until the first incumbent, improvement until the last, and proof. This subsection catalogues the heuristics solvers run and works through four of them on small examples. It ends with the heuristic that needs no LP at all, which is the one that has moved to the GPU first.

The three phases of a run (Berthold, Hendel and Koch, Section 3.1)

  start       first incumbent           last incumbent          end
    |                |                         |                  |
    +- feasibility --+------ improvement ------+------ proof -----+
                     |
                     from here on an incumbent exists, and with it
                     pruning by bound, the range reductions, the
                     cutoff row f(x) <= z_inc and a gap to stop on

Definition 3.2.1 (primal heuristic, large neighbourhood search). A primal heuristic is a procedure that tries to produce a feasible point of the problem without a guarantee of success and without a bound. A large neighbourhood search heuristic defines a neighbourhood of a reference point by adding constraints to the problem: fixing variables, bounding them, or bounding a distance. It solves the restricted problem with the solver itself under a work limit and keeps any improved point. The restricted problem is called the sub-MIP, or sub-MINLP.

The measure by which heuristics are judged is the primal integral. The primal gap and the primal integral \(P(T)\), with its time average \(\bar P(T)\), are Definitions 2.3.3 and 2.3.4. In the knapsack figure time is counted in nodes. The depth-first run's 36.9 node-units are \(P(T)\) with \(T = 315\), the best-bound run's 9.1 are \(P(37)\), and the time averages 0.117 and 0.245 rank the two runs the other way. A run that never finds a feasible point has \(P(T) = T\) and \(\bar P(T) = 1\). The integral is small when good solutions arrive early, and it sees what the time to optimality does not: a heuristic that changes the end of the run by a few percent can change the first half of it entirely.

heuristicwhat it doesneedssource
roundinground the relaxed point, coordinate by coordinate; simple rounding uses variable locks to stay feasibleLP pointBerthold 2006 (thesis); Wallace 2010 (ZI round)
divingfix one fractional variable, re-solve, repeat; fractional, coefficient, pseudocost, guided, vector-length; NLP diving does the same on the NLP relaxationLP re-solves (NLP re-solves)Berthold 2006; Bonami and Goncalves 2012 (NLP diving)
feasibility pumpalternate rounding with projection onto the relaxation until a rounding is feasible; MINLP version replaces rounding by a MILP over outer-approximation cutsLPs; NLPs and MILPs (MINLP)Fischetti, Glover and Lodi 2005; Bonami, Cornuejols, Lodi and Margot 2009
RINSfix the integers on which incumbent and LP agree; sub-MIP on the restincumbent, LPDanna, Rothberg and Le Pape 2005
local branchingsub-MIP inside the ball of \(k\) flips around the incumbentincumbentFischetti and Lodi 2003
RENSfix the integral LP coordinates, round the rest: a sub-MIP over all roundingsLP pointBerthold 2014
DINS, crossover, solution polishingdistance-induced, two-incumbent and evolutionary neighbourhoodsincumbentsGhosh 2007; Rothberg 2007
adaptive LNSa multi-armed bandit over eight neighbourhoodsincumbentHendel 2022
Undercoverfix a vertex cover of the nonlinear terms; the rest is a MIP, solved as a sub-MIPrelaxationBerthold and Gleixner 2014
sub-NLPfix the integers, solve the remaining NLP locallyinteger pointvirtually every global solver
multistartlocal NLP solves from sampled starts, clusterednothingUgray et al. 2007; SCIP 8
fix-and-propagatefix, propagate, repair, without an LPnothingGamrath et al. 2019; Salvagnin, Roberti and Fischetti 2025
Feasibility Jumpweighted local search on the integer lattice; no LPnothingLuteberget and Sartor 2023
penalty ADM, ADMMalternate over the blocks of a penalty problem; the pump is the special caseconvex blocksGeissler et al. 2017; Moehle, Gindi, Boyd and Kochenderfer 2023
The incumbent catalogue: what each heuristic fixes or searches, what it needs, and where it comes from.

Two rows of the table are not worked through below. Fix-and-propagate fixes the variables one at a time in a chosen order, propagates each fixing through the constraints, and repairs an infeasibility by undoing or altering fixings, all without an LP. It is also step 2 of Undercover.G. Gamrath, T. Berthold, S. Heinz and M. Winkler, "Structure-driven fix-and-propagate heuristics for mixed integer programming", Mathematical Programming Computation 11 (2019); D. Salvagnin, R. Roberti and M. Fischetti, "A fix-propagate-repair heuristic for mixed integer programming", Mathematical Programming Computation 17 (2025). The penalty methods are taken up in Section 5.4.

Rounding and diving

The cheapest heuristic rounds the relaxed point. For a row \(\sum_j a_{ij} x_j \le b_i\), increasing \(x_j\) can violate the row only if \(a_{ij} > 0\). The number of rows that can be violated by increasing \(x_j\) is its number of up-locks, and symmetrically for down-locks. A variable with no down-locks can be rounded down from any LP point without losing feasibility, and one with no up-locks rounded up. Simple rounding rounds every fractional variable that has a lock-free direction and gives up otherwise. ZI rounding and its relatives walk the fractional variables one at a time, rounding each in whichever direction the row slacks permit.T. Berthold, Primal Heuristics for Mixed Integer Programs, Diploma thesis, Technische Universität Berlin (2006), is the catalogue of the rounding and diving heuristics in SCIP; ZI rounding is C. Wallace, "ZI round, a MIP rounding heuristic", Journal of Heuristics 16 (2010). The book-length treatment is T. Berthold, A. Lodi and D. Salvagnin, Primal Heuristics in Integer Programming (Cambridge University Press, 2025). On the two-variable example the LP vertex \((4.93, 2.93)\) at \(45^\circ\) rounds to \((5, 3)\), which violates \(2x + 5y \le 24.5\). Both variables have up-locks in that row and down-locks elsewhere (\(x\) in \(-3x + 4y \le 11\), \(y\) in \(x - 2y \le 4.2\)), so simple rounding does nothing there. The roundings \((4, 3)\) and \((5, 2)\), both feasible and both worth \(7\), are found by the methods below. Diving is a depth-first probe without backtracking: fix one fractional variable to a rounded value, re-solve the LP, and repeat until the LP is integral or infeasible, with one or two levels of backtracking at most. The rules differ in the variable they fix: the least fractional (fractional diving), the one whose rounding violates the fewest rows (coefficient diving), the one with the smallest pseudocost estimate (pseudocost diving), the one that moves toward the incumbent (guided diving). NLP diving does the same on the NLP relaxation of a MINLP. SCIP's rule fixes the variable whose LP and NLP values are closest to a common integer, preferring binaries and variables in nonlinear terms. Bonami and Gonçalves found in Bonmin that diving rules and the pump "identify feasible solutions more rapidly than standard branch-and-bound" and reduce the total time.P. Bonami and J. P. M. Gonçalves, "Heuristics for convex mixed integer nonlinear programs", Computational Optimization and Applications 51 (2012); the SCIP rule is described in K. Bestuzheva, A. Chmiela, B. Müller, F. Serrano, S. Vigerske and F. Wegscheider, "Global optimization of mixed-integer nonlinear programs with SCIP 8", Journal of Global Optimization 91 (2025), Section 2.5.

Diving: a depth-first probe from the LP point

  LP point, fractional
     |
     v
  fix one fractional variable to a rounded value  <--------+
  (fractional, coefficient, pseudocost or guided diving)   |
     |                                                     |
     v                                                     |
  re-solve the LP                                          |
     |                                                     |
     +--> still fractional --------------------------------+
     |
     +--> integral: stop with a feasible point
     |
     +--> infeasible: stop, after one or two levels of
          backtracking at most

  NLP diving: the same loop on the NLP relaxation of a MINLP

The feasibility pump

The feasibility pump of Fischetti, Glover and Lodi alternates between two points, one in the relaxation and one on the integer lattice, until they coincide.M. Fischetti, F. Glover and A. Lodi, "The feasibility pump", Mathematical Programming 104 (2005). The general-integer version with auxiliary variables for the distance is L. Bertacco, M. Fischetti and A. Lodi, "A feasibility pump heuristic for general mixed-integer problems", Discrete Optimization 4 (2007); the objective feasibility pump, which blends the original objective into the distance with a decreasing weight, is T. Achterberg and T. Berthold, "Improving the feasibility pump", Discrete Optimization 4 (2007); Feasibility Pump 2.0 applies constraint propagation during the rounding, M. Fischetti and D. Salvagnin, "Feasibility pump 2.0", Mathematical Programming Computation 1 (2009); the survey is T. Berthold, A. Lodi and D. Salvagnin, "Ten years of feasibility pump, and counting", EURO Journal on Computational Optimization 7 (2019). The relaxation point is rounded. If the rounding is infeasible, the point of the relaxation nearest to the rounding in the 1-norm is computed, which is an LP, and rounded again.

Algorithm 3.2.2  FEASIBILITY-PUMP
                 (MILP; Fischetti, Glover and Lodi 2005)

Input   MILP with LP relaxation polyhedron P and integer index set I;
        a round limit; the flip range [T_lo, T_hi] (10 to 30).
Output  a feasible point of the MILP, or failure.

 1. x* <- an optimal point of the LP relaxation
    previous rounding <- none

 2. repeat up to the round limit:

 3.    x~ <- x* with every coordinate in I rounded to the nearest
       integer

 4.    if x~ in P: return x~       [integral on I and in P: feasible]

 5.    if x~ equals the previous rounding (a cycle of length one):
          flip the T coordinates of I with the largest |x*_j - x~_j|,
          T drawn from [T_lo, T_hi]
       after repeated or longer cycles: random restart

 6.    previous rounding <- x~

 7.    x* <- argmin { Delta(x, x~) : x in P },                [an LP]
       Delta(x, x~) = sum_{j in I} |x_j - x~_j|,
       which for binaries is
          sum_{j: x~_j = 0} x_j + sum_{j: x~_j = 1} (1 - x_j)

 8. return failure

Invariant
    x* always lies in P and x~ is always integral on I; the two
    coincide exactly when a feasible point has been found.
    Delta(x*, x~) is the 1-norm distance from the lattice point to the
    relaxation, and each projection is the point of P nearest to the
    current rounding.

Cost per round
    one rounding (O(n)) and one LP, warm-started from the previous one
    since only the objective changed.

Parallel
    several pumps from different starts or with different flips are
    independent.

(The pump as block coordinate descent) The invariant says what the pump is: a block coordinate descent on the distance between the polyhedron and the lattice, with rounding as the exact minimization over the lattice and the LP as the exact minimization over the polyhedron. Descent on a nonconvex function stalls at points that are minimal in each block separately, and that is the cycle the flips and restarts are for. Geißler, Morsi, Schewe and Schmidt made the observation precise. The feasibility pump is a penalty alternating direction method on the \(\ell_1\) penalty of the coupling \(v_I = u_I\) between a copy \(v\) in the relaxation and a copy \(u\) on the lattice. They prove that such methods converge to partial minima of the penalty problem and, with increasing penalty parameters, to feasible points under stated assumptions.B. Geißler, A. Morsi, L. Schewe and M. Schmidt, "Penalty alternating direction methods for mixed-integer optimization: a new view on feasibility pumps", SIAM Journal on Optimization 27 (2017); proof there. The alternating-direction heuristic for portfolio problems with separable nonconvex terms, N. Moehle, J. Gindi, S. Boyd and M. J. Kochenderfer, "Portfolio construction as linearly constrained separable optimization", Optimization and Engineering 24 (2023); arXiv 2103.05455, belongs to the same family and is taken up in Section 5.4.

The figure below runs the pump on the running example from a start the reader can drag. In a second mode it runs Feasibility Jump, the method of the last part of this subsection. From the default start \((6.40, 5.20)\) the pump rounds to \((6, 5)\), which is outside the polygon. It projects onto the polygon and lands on the vertex \((4.929, 2.929)\), where \(2x + 5y = 24.5\) and \(5x + 2y = 30.5\) meet. It rounds to \((5, 3)\), which violates \(2x + 5y \le 24.5\), projects back to the same vertex, and would round to \((5, 3)\) again. The cycle is broken by a flip. Both coordinates of the vertex have fractional part \(13/14\) and lie \(1/14\) from their roundings, so the flip rule of step 5, the coordinate farthest from its rounding, is a tie, which the figure resolves toward \(x\). The flip gives \((4, 3)\), which is feasible, and the figure reports three roundings, one cycle broken and the feasible point \((4, 3)\). Resolving the tie toward \(y\) gives \((5, 2)\) instead, also feasible and worth the same \(7\) under the objective \(x + y\). The figure projects in the Euclidean norm, which draws the same picture as the 1-norm LP of the original on this polygon: the two projections coincide at every step of this run. In the Feasibility Jump mode the figure never leaves the lattice and reaches \((5, 1)\) in two moves. The last part of this subsection explains them.

Two ways to find a feasible point from a start you can drag anywhere. The feasibility pump rounds the point you have (orange ring), projects back onto the polygon when the rounding is infeasible (red), and rounds again until a rounding is feasible (green). Feasibility Jump never leaves the integer lattice: from the rounded start it moves one variable at a time along the grey dashed path to the value that most reduces the weighted violation, raises the weights of the violated rows when no move helps (red ticks), and stops at a feasible point (green).
The feasibility pump on R1 from the default start (6.40, 5.20)

  continuous points                        lattice points

  start (6.40, 5.20) ---------- round ---> (6, 5): outside
                                              |
  (4.929, 2.929) <------------ project -------+
  the vertex where 2x + 5y = 24.5
  and 5x + 2y = 30.5 meet
     |
     +------------------------- round ---> (5, 3): violates
                                           2x + 5y <= 24.5
                                              |
  (4.929, 2.929) <------------ project -------+
  the same vertex
     |
     +------------------------- round ---> (5, 3) again: a cycle
                                              |
                  flip one coordinate: both fractional parts are
                  13/14, a tie
                                              |
                    toward x <----------------+--------> toward y
                       |                                    |
                       v                                    v
                    (4, 3) feasible                 (5, 2) feasible
                    (the figure's choice)

  both are worth 7 under the objective x + y; the figure reports
  three roundings, one cycle broken and the feasible point (4, 3)

The program below is the pump of the figure, with the Euclidean projection computed exactly by testing the foot of the perpendicular on every edge and every vertex of the polygon.

# The feasibility pump (Fischetti, Glover and Lodi 2005) on R1.
#
# The running example R1 is the polygon 2x + 5y <= 24.5, 5x + 2y <= 30.5,
# -3x + 4y <= 11, x - 2y <= 4.2, x, y >= 0, integer points wanted.
# Round; if the rounding is infeasible, project it back onto the polygon
# and round again; if the rounding repeats, flip the coordinate farther
# from its rounding (x on a tie), as fig-pump does.

import itertools
import math

ROWS = [(2, 5, 24.5), (5, 2, 30.5), (-3, 4, 11), (1, -2, 4.2),
        (-1, 0, 0), (0, -1, 0)]

def feasible(p, tol=1e-9):
    return all(a * p[0] + b * p[1] <= r + tol for a, b, r in ROWS)

def vertices():
    """The polygon's corners: pairwise intersections that are feasible."""
    V = []
    for (a1, b1, r1), (a2, b2, r2) in itertools.combinations(ROWS, 2):
        d = a1 * b2 - a2 * b1
        if abs(d) > 1e-12:
            p = ((r1 * b2 - r2 * b1) / d, (a1 * r2 - a2 * r1) / d)
            if feasible(p):
                V.append(p)
    return V

def project(q):
    """Nearest point of the polygon in the 2-norm (the figure's choice)."""
    if feasible(q):
        return q
    best = None
    # the foot of the perpendicular on each edge, if feasible
    for a, b, r in ROWS:
        t = (r - a * q[0] - b * q[1]) / (a * a + b * b)
        foot = (q[0] + a * t, q[1] + b * t)
        for p in [foot] + vertices():
            if not feasible(p):
                continue
            if best is None or math.dist(p, q) < math.dist(best, q) - 1e-12:
                best = p
    return best

def pump(start, max_rounds=12):
    """Run the pump from start; return its log of (step, point)."""
    x, last, log = start, None, [("start", start)]
    for _ in range(max_rounds):
        xr = [round(x[0]), round(x[1])]
        if xr == last:
            # the rounding repeats: flip the coordinate farther from its rounding
            fx, fy = abs(x[0] - xr[0]), abs(x[1] - xr[1])
            j = 0 if fx >= fy - 1e-9 else 1
            xr[j] += 1 if x[j] >= xr[j] else -1
            log.append(("flip", tuple(xr)))
        last = list(xr)
        if feasible(xr):
            log.append(("round: feasible", tuple(xr)))
            return log
        log.append(("round: infeasible", tuple(xr)))
        x = project(xr)
        log.append(("project", x))
    return log + [("gave up", None)]

for kind, p in pump((6.4, 5.2)):
    print(f"{kind:<18} ({p[0]:.3f}, {p[1]:.3f})" if p is not None else kind)
start              (6.400, 5.200)
round: infeasible  (6.000, 5.000)
project            (4.929, 2.929)
round: infeasible  (5.000, 3.000)
project            (4.929, 2.929)
flip               (4.000, 3.000)
round: feasible    (4.000, 3.000)

The cost here is trivial. In a solver each projection is an LP with the same rows as the relaxation and a new objective. The dual simplex warm start of Proposition 3.1.12 does not apply directly, since the objective changed and not the bounds, and a primal simplex from the previous optimal basis is used instead. Several pumps from different roundings are independent and run in parallel without interaction.

(The pump for a MINLP) For a MINLP the pump needs two changes, both due to Bonami, Cornuéjols, Lodi and Margot.P. Bonami, G. Cornuéjols, A. Lodi and F. Margot, "A feasibility pump for mixed integer nonlinear programs", Mathematical Programming 119 (2009). The nonconvex variants are C. D'Ambrosio, A. Frangioni, L. Liberti and A. Lodi, "A storm of feasibility pumps for nonconvex MINLP", Mathematical Programming 136 (2012). The projection onto the relaxation is an NLP, since the relaxed feasible set \(\mathcal F_{\mathrm{rel}}\) is now curved. The rounding step is replaced by a MILP over linearizations collected at the projection points, so that the rounding cannot return to a lattice point already cut off. The linearization of a constraint \(g_i(v) \le 0\) at a point \(\hat v\) is the affine inequality \(g_i(\hat v) + \nabla g_i(\hat v)^\top (v - \hat v) \le 0\), the inequality \(\ell_{\hat v}(v) \le 0\) of Proposition 1.2.4. For a convex \(g_i\) it is valid for the whole feasible set by that proposition. Section 3.4 builds the outer-approximation method on this fact.

Algorithm 3.2.3  FEASIBILITY-PUMP for convex MINLP
                 (Bonami, Cornuejols, Lodi and Margot 2009)

Input   convex MINLP with relaxed feasible set F_rel (Section 3.4),
        integer index set I, a lattice point v~^1 (for instance the
        rounding of the continuous relaxation's solution).
Output  a feasible point of the MINLP, or a certificate that none
        exists.

 1. K <- {};  k <- 1

 2. Projection (an NLP):
       vhat^k <- argmin { || v_I - v~^k_I ||_2 : v in F_rel }
    If vhat^k is integral on I: return vhat^k

 3. K <- K + {vhat^k}: add the linearizations
       g_i(vhat^k) + grad g_i(vhat^k)^T (v - vhat^k) <= 0
    of every constraint at vhat^k to the polyhedron P_K

 4. Rounding (a MILP):
       v~^{k+1} <- argmin { || v_I - vhat^k_I ||_1 :
                            v in P_K, v_I integral }
    If the MILP is infeasible: the MINLP is infeasible; stop.
    Else k <- k + 1 and go to 2.

Invariant
    every cut is valid for the MINLP, so an infeasible MILP certifies
    infeasibility; the cuts added at vhat^k exclude v~^k from P_K
    (Proposition 3.2.4), so no lattice point is visited twice.

Cost per round
    one smooth convex NLP and one MILP, which may be solved to a node
    limit or stopped at the first improving integer point.

Parallel
    independent pumps from different starts; the MILPs are the cost.

Proposition 3.2.4 (Bonami, Cornuéjols, Lodi and Margot 2009). For a convex MINLP, Algorithm 3.2.3 visits no lattice point twice. It therefore terminates after at most as many rounds as there are integer assignments, either with a feasible point or with an infeasible MILP that certifies the MINLP infeasible.

Proof sketch. The projection is a convex NLP: a convex objective over the convex constraints \(g_i \le 0\) and the box. Let \(\hat v\) be the projection of \(\tilde v\), with \(\hat v_I \ne \tilde v_I\). The vector \((\tilde v - \hat v)_I\) is the negative gradient of the objective \(\tfrac12 \|v_I - \tilde v_I\|^2\). Assume Slater's condition for the relaxed set, a point with every \(g_i < 0\). Theorem 1.3.4(ii) then makes the KKT conditions necessary at \(\hat v\), and they say that this vector is a nonnegative combination \(\sum_i \lambda_i \nabla_I g_i(\hat v)\) of the gradients of the active constraints, up to a vector in the normal cone of the box at \(\hat v\). That cone is the set of outward directions at the face of the box containing \(\hat v\): its vectors are nonpositive in the coordinates at their lower bound, nonnegative in those at their upper bound and zero elsewhere, so they have a nonpositive inner product with \(v - \hat v\) for every \(v\) in the box. Any \(v\) in the box satisfying the linearizations at \(\hat v\) has \(\nabla g_i(\hat v)^\top (v - \hat v) \le -g_i(\hat v) = 0\) for every active \(i\), hence \((\tilde v - \hat v)_I^\top (v_I - \hat v_I) \le 0\). At \(v = \tilde v\) this reads \(\|\tilde v_I - \hat v_I\|^2 \le 0\), which is impossible, so the cuts exclude \(\tilde v\) from every later MILP. The cuts are valid for the MINLP because the constraints are convex, so an infeasible MILP means that no feasible point exists. ∎

(Why the cuts never revisit a lattice point) In a picture, the vector from the projection to the lattice point is an outward normal of the relaxed set at the projection. The tangent plane there therefore separates the lattice point from the relaxation. For a nonconvex MINLP both steps weaken: the projection NLP can only be solved locally, and the linearizations are valid only for the constraints known to be convex, so the cuts must be restricted to those. D'Ambrosio, Frangioni, Liberti and Lodi study a large family of such variants, which they call a storm of pumps, on the MINLPLib instances. This is the form in which the pump runs inside the global solvers.

Proposition 3.2.4 in a picture (convex case): the vector from the projection v̂ᵏ to the lattice point ṽᵏ is an outward normal of F_rel at v̂ᵏ, so the tangent plane there, the linearization added to P_K, cuts ṽᵏ off from every later MILP over P_K, while F_rel stays on its near side and the cut is valid for the MINLP. Schematic, no scale.

Large neighbourhoods

Definition 3.2.5 (the standard neighbourhoods). Let \(x_{\mathrm{inc}}\) be the incumbent and \(\bar x\) the solution of the relaxation. RINS fixes \(x_j = x_{\mathrm{inc}}_j\) for every \(j \in I\) with \(x_{\mathrm{inc}}_j = \bar x_j\), the coordinates on which incumbent and relaxation agree, and solves the sub-MIP on the rest.E. Danna, E. Rothberg and C. Le Pape, "Exploring relaxation induced neighborhoods to improve MIP solutions", Mathematical Programming 102 (2005). RENS fixes \(x_j = \bar x_j\) where \(\bar x_j\) is integral and imposes \(\lfloor \bar x_j \rfloor \le x_j \le \lceil \bar x_j \rceil\) on the other integer coordinates, so its sub-MIP ranges over all roundings of \(\bar x\).T. Berthold, "RENS: the optimal rounding", Mathematical Programming Computation 6 (2014). Local branching adds, for binary variables, the row \(\sum_{j : x_{\mathrm{inc}}_j = 0} x_j + \sum_{j : x_{\mathrm{inc}}_j = 1} (1 - x_j) \le k\), the ball of radius \(k\) in the Hamming distance around the incumbent, with \(k\) about 10 to 20 in the original paper. For general integers it adds the 1-norm ball \(\sum_{j \in I} |x_j - x_{\mathrm{inc}}_j| \le k\).M. Fischetti and A. Lodi, "Local branching", Mathematical Programming 98 (2003).

Two small computations show the sizes. With four binaries and the incumbent \(x_{\mathrm{inc}} = (1, 0, 1, 0)\), local branching with \(k = 1\) adds \((1 - x_1) + x_2 + (1 - x_3) + x_4 \le 1\), which admits \(x_{\mathrm{inc}}\) and its four single flips, five points out of sixteen. If the LP solution is \(\bar x = (1, 0, 0.5, 1)\), RINS fixes \(x_1 = 1\) and \(x_2 = 0\), where the two agree, and leaves four candidates. With three general integers and the relaxed point \(\bar x = (2.4, 1.7, 3.0)\), RENS fixes \(x_3 = 3\) and leaves \(x_1 \in \{2, 3\}\), \(x_2 \in \{1, 2\}\), four candidates. The relaxed point \((2.4, 1.7, 2.6)\) leaves eight. A neighbourhood is small not because it has few variables but because the tree below it is shallow: every free variable in these examples has a range of one.

The small examples of Definition 3.2.5: what each sub-MIP keeps

  local branching, k = 1, around x_inc = (1, 0, 1, 0):
  (1 - x_1) + x_2 + (1 - x_3) + x_4 <= 1

          (0,0,1,0)                       (1,1,1,0)
           flip x_1 \                   / flip x_2
                     \                 /
                       (1,0,1,0) x_inc
                     /                 \
           flip x_3 /                   \ flip x_4
          (1,0,0,0)                       (1,0,1,1)

          x_inc and its four single flips: 5 of the 16 points

  RINS    x_inc     1       0       1       0
          xbar      1       0       0.5     1
                    agree   agree
          x_1 = 1, x_2 = 0 fixed; x_3, x_4 free: 4 candidates

  RENS    xbar      2.4     1.7     3.0
          values    {2,3}   {1,2}   {3}     4 candidates
          xbar      2.4     1.7     2.6
          values    {2,3}   {1,2}   {2,3}   8 candidates

  every free variable has a range of one: the tree below is shallow

Proposition 3.2.6 (Berthold 2014). The feasible set of the RENS sub-MIP is exactly the set of feasible points of the problem whose integer part is a rounding of \(\bar x\). Its optimal value is therefore the value of the best rounding of \(\bar x\), and it is infeasible if and only if no rounding of \(\bar x\) extends to a feasible point.

Proof. A vector with \(\lfloor \bar x_j \rfloor \le x_j \le \lceil \bar x_j \rceil\) and \(x_j \in \mathbb Z\) is by definition a rounding of \(\bar x\) in coordinate \(j\), and the sub-MIP keeps every other constraint of the problem. ∎

Berthold reports that across MIP, MIQCP and MINLP test sets "60 to 70% of the instances have roundable relaxation optima". He also reports that "the success rate of RENS does not depend on the percentage of fractional variables". As a root heuristic, RENS "complements nicely with existing primal heuristics in SCIP". The proposition says what one gets if the sub-MIP is solved to optimality. In practice it is solved under a node limit, like every sub-MIP in this subsection. The figure below draws the two neighbourhoods that need an incumbent on the running example. For local branching it draws the diamond \(|x - x_{\mathrm{inc}}_1| + |y - x_{\mathrm{inc}}_2| \le k\) in purple, and for RINS the line on which the coordinate shared by the incumbent and the rounded LP vertex is fixed. In the default view, at \(45^\circ\) with the incumbent \((4, 2)\) worth \(4.24\) and \(k = 2\), the diamond holds \(2k^2 + 2k + 1 = 13\) lattice points, 10 of them feasible. The best of them is \((4, 3)\), worth \(4.95\), an improvement of \(0.71\), tied with \((5, 2)\), and the two are the integer optima of the whole problem of 22 feasible points. The case "\(k = 1\) around \((1, 3)\)" holds 5 lattice points, 3 feasible, with best \((2, 3)\) worth \(3.54\), and the optimum lies outside the ball: a neighbourhood too small to reach it. The case "RINS from \((3, 3)\)" compares the incumbent with the rounded LP vertex \((5, 3)\): \(x\) differs, \(y\) agrees, so \(y = 3\) is fixed and the free line holds 4 feasible points, the best again \((4, 3)\). The case "RINS from \((4, 2)\)" agrees with \((5, 3)\) on neither coordinate, so RINS fixes nothing and the sub-MIP is the full problem. A solver would not call RINS there, because it requires a minimum fraction of the variables to be fixed before it starts a sub-MIP.

Two neighbourhoods on the running example. The green point is the incumbent, the orange ring the LP vertex, and the purple region the neighbourhood the sub-MIP searches: for local branching the diamond of 1-norm radius k about the incumbent, for RINS the line on which the coordinate shared by the incumbent and the rounded LP vertex is fixed. Feasible lattice points in the neighbourhood are ringed in purple, infeasible ones are marked in red, the best of them is blue, and a click on a lattice point makes it the incumbent while a drag elsewhere turns the objective.

The figure solves each sub-MIP by enumerating the lattice points of the neighbourhood, which is exact here. A solver would run Algorithm 3.1.4 on it, and the point of the neighbourhood is that this tree is shallow. The sub-MIP inherits the whole machinery of the solver, presolve, cuts, propagation and branching. This is why RINS and local branching are two of the ideas that Koch, Berthold, Pedersen and Vanaret single out in explaining the progress of MILP codes between 2001 and 2020.T. Koch, T. Berthold, J. Pedersen and C. Vanaret, "Progress in mathematical programming solvers from 2001 to 2020", EURO Journal on Computational Optimization 10 (2022), Section 1.1: "many new heuristic methods, such as RINS and local branching". The family has grown since. DINS defines the neighbourhood by distance from the relaxation, crossover fixes the coordinates on which two incumbents agree, and Rothberg's solution polishing runs an evolutionary algorithm over a pool of incumbents.S. Ghosh, "DINS, a MIP improvement heuristic", IPCO 2007, LNCS 4513 (2007); E. Rothberg, "An evolutionary algorithm for polishing mixed integer programming solutions", INFORMS Journal on Computing 19 (2007). Hendel's adaptive large neighbourhood search treats eight neighbourhoods (RINS, crossover, mutation, RENS, local branching, proximity search, zero objective and DINS) as the arms of a multi-armed bandit. The arm is selected by an \(\varepsilon\)-greedy, upper-confidence or Exp.3 rule. The reward combines whether a solution was found, the gap it closed and a failure penalty. The fixing rate of the sub-MIPs is adapted between 0.1 and 0.9 from their outcomes. The method was evaluated on 666 instances with about twenty calls per instance.G. Hendel, "Adaptive large neighborhood search for mixed integer programming", Mathematical Programming Computation 14 (2022).

Undercover: a sub-MIP for a MINLP

For a MINLP the natural neighbourhood is the one in which the nonlinearity disappears. Berthold and Gleixner's Undercover fixes just enough variables to make every constraint linear in the rest, and hands the remaining MIP to the MILP engine.T. Berthold and A. M. Gleixner, "Undercover: a primal MINLP heuristic exploring a largest sub-MIP", Mathematical Programming 144 (2014).

Definition 3.2.7 (co-occurrence graph, cover). Consider a problem whose nonlinear terms are products \(x_i x_j\) and squares \(x_i^2\). Its co-occurrence graph has the variables as vertices, an edge \(\{i, j\}\) for every product \(x_i x_j\) with \(i \ne j\), and a loop at \(i\) for every square \(x_i^2\). A set \(C\) of variables is a cover if fixing \(x_C\) makes every constraint linear in the remaining variables. For general factorable functions the same notion is read off the expression trees: a variable inside a nonlinear univariate term must be fixed, and of two variables in a product one must be.

Proposition 3.2.8 (Berthold and Gleixner 2014). \(C\) is a cover if and only if it is a vertex cover of the co-occurrence graph: every edge has an endpoint in \(C\) and every looped vertex is in \(C\). A minimum cover is therefore a minimum vertex cover, computed by the set-covering program \(\min\{\sum_j \alpha_j : \alpha_i + \alpha_j \ge 1 \text{ for every edge } \{i, j\},\ \alpha_i \ge 1 \text{ for every loop},\ \alpha \in \{0, 1\}^n\}\). After fixing \(x_C\) the remaining problem is a MIP, and every feasible point of that MIP is a feasible point of the original problem.

Proof. After fixing \(x_C\) the term \(x_i x_j\) is affine in the free variables exactly when \(i \in C\) or \(j \in C\), and \(x_i^2\) is constant exactly when \(i \in C\). Every other term was affine already. The sub-MIP's constraints are the original constraints with some variables fixed, so its feasible points satisfy the original constraints exactly. ∎

A small example: the constraints \(x_1 x_2 \le 1\), \(x_2 x_3 \le 2\) and \(x_3^2 + x_4 \le 3\) give the edges \(x_1\)–\(x_2\) and \(x_2\)–\(x_3\) and a loop at \(x_3\). The smallest cover is \(\{x_2, x_3\}\) or \(\{x_1, x_3\}\), two of four variables, and fixing either at its relaxed values leaves rows linear in the rest. In the tax problem of Section 9 the no-wash-sale row \(b_i \cdot S_i = 0\) is one edge per asset, and a cover is one variable per asset: fix the directions and the problem is a MIP, which is the structural fact that Section 9 builds on.

The small example of Definition 3.2.7: its co-occurrence graph

  x_1 x_2 <= 1         x_2 x_3 <= 2         x_3^2 + x_4 <= 3

                                       +-----+
                                       |     |  x_3^2: a loop
     x_1 ------------ x_2 ------------ x_3 --+          x_4
           x_1 x_2          x_2 x_3                (in no product)

  the smallest covers, two of the four variables:
     {x_2, x_3}  and  {x_1, x_3}

                 x_1 x_2 <= 1    x_2 x_3 <= 2    x_3^2 + x_4 <= 3
  fix x_2, x_3   x_1 [x_2]       [x_2] [x_3]     [x_3]^2 + x_4
  fix x_1, x_3   [x_1] x_2       x_2 [x_3]       [x_3]^2 + x_4

  [ ] fixed at its relaxed value: every row is linear in the rest

  the tax problem (Section 9): b_i S_i = 0, one edge per asset
     b_1 ---- S_1        b_2 ---- S_2        ...
  a cover is one variable per asset
Algorithm 3.2.9  UNDERCOVER
                 (Berthold and Gleixner 2014)

Input   a MINLP with factorable constraints; a reference point v_ref
        (the LP or NLP relaxation's solution, or an incumbent); work
        limits for the sub-MIP.
Output  a feasible point, or failure.

 1. Build the co-occurrence graph (Definition 3.2.7); solve the
    set-covering program of Proposition 3.2.8, or a greedy
    approximation, for a minimum cover C.

 2. Fix-and-propagate: for each j in C in order,
       fix x_j to the value of v_ref (rounded if j is integer) and run
       domain propagation;
       if propagation finds infeasibility, backtrack a bounded number
       of times, altering the order or the values (conflict analysis,
       Section 3.6, supplies the reason for the infeasibility).

 3. Every constraint is now linear in the free variables: solve the
    remaining MIP under a node or time limit; any feasible point of
    it is feasible for the MINLP.

 4. Polish: fix the integers of the best sub-MIP point and solve the
    remaining NLP locally (sub-NLP); keep the better of the two
    points.

Invariant
    after step 2 every constraint is affine in the free variables
    (Proposition 3.2.8), so step 3's points satisfy the nonlinear
    constraints exactly.

Cost
    a small covering program, one propagation pass per fixing, one
    truncated MIP, one NLP.

Parallel
    different covers, fixing values and orders give independent
    sub-MIPs.

The point of the method is that in many MINLPs the nonlinearity touches few variables, and a bilinear term needs only one of its two variables fixed. Hence "the majority of these instances allows for small covers". Berthold and Gleixner report that the heuristic "is most successful on MIQCPs" and "helps to significantly improve the overall performance of the MINLP solver SCIP". In SCIP 8 the cover is chosen by the set-covering program, and the fixing values come from the LP or NLP relaxation or from a known feasible point. The sub-MIP "does not need to be solved to proven optimality", and its best point is passed to the sub-NLP heuristic as a starting point.K. Bestuzheva, A. Chmiela, B. Müller, F. Serrano, S. Vigerske and F. Wegscheider, Journal of Global Optimization 91 (2025), Section 2.5. The rounding-based heuristics for Couenne, which round an NLP solution by solving a MILP at each step and improve it in neighbourhoods "defined by local branching cuts or box constraints", are G. Nannicini and P. Belotti, "Rounding-based heuristics for nonconvex MINLPs", Mathematical Programming Computation 4 (2012).

The NLP side: sub-NLP and multistart

Two heuristics of a global solver need no integer machinery at all. Sub-NLP takes any point whose integer coordinates are integral, from the LP, from a MIP heuristic or from Undercover. It fixes them, presolves, and solves the remaining NLP locally from that start. It is implemented, in the SCIP 8 paper's words, in "virtually any global MINLP solver", and SCIP sets its iteration limit adaptively from the average of previous runs. Multistart samples start points in the box and pushes each toward the feasible set by inexpensive gradient steps on the violated constraints. It clusters the pushed points and runs a local NLP solve from each cluster representative. BARON's preprocessing runs a randomized multistart whose size is set by its NumLoc option, and its DoLocal option governs the local searches launched at nodes of the tree. Knitro's multistart runs its local solves in parallel over threads and is documented as deterministic even then and as unable to "guarantee that multi-start will find the global optimum".Bestuzheva et al. (2025), Section 2.5, for SCIP's sub-NLP and multistart; BARON User Manual, Sections 5.5 and 11.5, "Local search options" (NumLoc, default −2, which lets BARON choose the number of local searches; DoLocal, default 1, which lets BARON decide when to run local search at nodes), The Optimization Firm, minlp.com/baron-user-manual, read 5 October 2026; Artelys, Knitro 16.0 User Guide, "Multi-start". The scatter-search multistart with a filter against repeated local solves is Z. Ugray, L. Lasdon, J. Plummer, F. Glover, J. Kelly and R. Martí, "Scatter search and local NLP solvers: a multistart framework for global optimization", INFORMS Journal on Computing 19 (2007). The arithmetic of multistart is that of Proposition 1.3.12, which gives the probability \((1 - \theta_1)^N\) that \(N\) independent starts all miss a global basin of share \(\theta_1\), together with the Bayesian stopping rule of Boender and Rinnooy Kan. With a basin of 20 % ten starts miss it with probability 0.107 and thirty with 0.0012. The cost grows linearly and the failure probability falls geometrically. This is why multistart is the first heuristic a solver runs and the first one to move to a parallel machine: the starts are independent, and a GPU can run thousands of them at once (Section 7.7). The stopping rule of Proposition 1.3.12(ii), with the thresholds cited in Section 1.3, is how a solver's multistart decides that the sample is large enough.

Multistart: sample, push, cluster, one local solve per cluster

  start points sampled       o     o    o       o      o      o
  in the box
       |
       |  push each toward the feasible set: inexpensive
       |  gradient steps on the violated constraints
       v
  pushed points                o  o  o            o   o      o
       |
       |  cluster the pushed points
       v
  clusters                   ( o  o  o )        ( o   o )  ( o )
                                  |                 |        |
  one representative each        [o]               [o]      [o]
       |
       |  a local NLP solve from each representative
       v
  local solutions                 *                 *        *

  pure multistart (Proposition 1.3.12(i)): N independent uniform
  starts all miss a global basin of share theta_1 with probability
  (1 - theta_1)^N; for a basin of 20 %:

       N                      10        30
       (1 - theta_1)^N     0.107    0.0012

Feasibility Jump: no LP at all

The most recent heuristic in the catalogue drops the LP entirely. Feasibility Jump is Luteberget and Sartor's entry to the 2022 MIP workshop's computational competition. It is a local search over the integer lattice guided by a weighted sum of constraint violations, with the weights playing the role of Lagrange multipliers: a row's weight rises each time the search is stuck with that row violated, as a multiplier rises in a dual ascent while its constraint is violated, which is the sense of the word Lagrangian in the paper's title.B. Luteberget and G. Sartor, "Feasibility Jump: an LP-free Lagrangian MIP heuristic", Mathematical Programming Computation 15 (2023); the reference implementation is at github.com/sintef/feasibilityjump. The competition is the Mixed Integer Programming Workshop 2022 computational competition, mixedinteger.org/2022/competition. Write the constraints as \(\sum_j a_{ij} x_j \le b_i\), \(i \in M\), with an equation written as two inequalities, give each a weight \(w_i \ge 0\), and define the weighted violation

\[F^w(x) \;=\; \sum_{i \in M} w_i \max\Big\{0,\ \sum_j a_{ij} x_j - b_i\Big\} .\]

A point is feasible exactly when \(F^w(x) = 0\). The method moves one variable at a time to the value that lowers \(F^w\) most, and raises the weights of the violated rows when no single move helps.

Proposition 3.2.10 (the one-variable violation function; Luteberget and Sartor 2023). Fix all variables but the integer variable \(x_j\) at the current point \(\bar x\) and put \(d_i = b_i - \sum_{k \ne j} a_{ik} \bar x_k\), the room row \(i\) leaves for \(a_{ij} x_j\). The function

\[G_j(t) \;=\; \sum_{i :\, a_{ij} \ne 0} w_i \max\{0,\ a_{ij} t - d_i\}\]

is \(F^w\) with \(x_j\) replaced by \(t\), up to the constant contributed by the rows that do not contain \(x_j\). It is convex and piecewise linear in \(t\), and its slope changes only at the breakpoints \(\beta_i = d_i / a_{ij}\), by \(\sigma_i = w_i |a_{ij}|\) at each. Over the real interval \([l_j, u_j]\) its minimizers form an interval whose left end \(t_{\mathrm{lo}}\) is a breakpoint or a bound. Over the integers of \([l_j, u_j]\) its minimizers are consecutive integers, and the smallest of them is \(\lfloor t_{\mathrm{lo}} \rfloor\) or \(\lceil t_{\mathrm{lo}} \rceil\). The jump value \(\mathrm{Jump}_j(\bar x)\), the smallest minimizer of \(G_j\) over the integers of \([l_j, u_j]\) other than \(\bar x_j\), lies in \(\{\lfloor t_{\mathrm{lo}} \rfloor, \lceil t_{\mathrm{lo}} \rceil, \bar x_j - 1, \bar x_j + 1\}\), and Algorithm 3.2.11 computes it exactly.

The costs, with \(\eta\) the largest number of rows a variable appears in and \(\mu\) the largest number of variables in a row, are these:

Proof sketch. Each summand of \(G_j\) is a convex piecewise-linear function of \(t\) with one breakpoint, \(\beta_i\), at which its slope rises by \(w_i |a_{ij}|\): from \(0\) to \(w_i a_{ij}\) if \(a_{ij} > 0\), and from \(w_i a_{ij} < 0\) to \(0\) if \(a_{ij} < 0\). A finite sum of such functions is convex and piecewise linear with these breakpoints, so its minimizers over an interval form a subinterval whose ends are breakpoints or bounds. The restriction of a convex function to the integers has nondecreasing differences, so its integer minimizers are consecutive. The smallest of them is the first integer at or after \(t_{\mathrm{lo}}\) if some integer minimizes \(G_j\) over the reals, and otherwise it is the better of the two integers around \(t_{\mathrm{lo}}\). If that smallest minimizer is \(\bar x_j\) itself, the differences are negative before \(\bar x_j\) and nonnegative after it, so the best integer other than \(\bar x_j\) is \(\bar x_j - 1\) or \(\bar x_j + 1\). The cost statements count the loops of Algorithms 3.2.11 and 3.2.12. ∎

Algorithm 3.2.11  JUMP-VALUE of the integer variable j at the point xbar
                  (exact)

Input   rows a_i x <= b_i with weights w_i; integer bounds
        l_j <= x_j <= u_j; the current xbar; d_i as above.
Output  Jump_j(xbar), the smallest minimizer of G_j over the integers
        of [l_j, u_j] other than xbar_j.

 1. for each row i with a_ij != 0:
       beta_i <- d_i / a_ij;  sigma_i <- w_i |a_ij|

 2. s <- sum of sigma_i over rows with a_ij > 0 and beta_i <= l_j,
         minus the sum of sigma_i over rows with a_ij < 0 and
         beta_i > l_j     [the slope of G_j just to the right of l_j]

 3. t <- l_j
    if s < 0:
       for each breakpoint beta_i in (l_j, u_j) in increasing order:
          s <- s + sigma_i
          if s >= 0: t <- beta_i; stop the loop
       If the loop ends with s < 0: t <- u_j               [t = t_lo]

 4. k <- the better of floor(t) and ceil(t), clamped to [l_j, u_j],
    under G_j; the smaller value on ties

 5. if k = xbar_j:
       k <- the better of xbar_j - 1 and xbar_j + 1, among those
       within the bounds; smaller on ties

 6. return k

Cost
    one sort of at most eta breakpoints and at most four evaluations
    of G_j at O(eta) each.

For a row with a positive coefficient \(a_{ij}\) the summand is zero up to its breakpoint \(\beta_i\) and rises after it; for a negative coefficient it falls until the breakpoint and is zero beyond. In both cases the slope of \(G_j\) rises by \(\sigma_i = w_i |a_{ij}|\) at \(\beta_i\), and the sweep of Algorithm 3.2.11 does nothing but accumulate these slope changes in breakpoint order until the slope turns nonnegative.

The two kinds of summand of G_j(t) in Algorithm 3.2.11. Each is zero on one side of its breakpoint β_i = d_i / a_ij and linear on the other, so at β_i the slope of G_j rises by σ_i = w_i |a_ij|.
The summands of G_j(t) and the sweep to t_lo in Algorithm 3.2.11

  steps 2 and 3 sweep the slope s of G_j from l_j to the right:

   l_j            beta              beta              beta       u_j
    |---------------|-----------------|-----------------|----------|
    s < 0           s <- s + sigma_i  s <- s + sigma_i  not reached
                    still < 0         now >= 0:
                                      t_lo = this beta

  step 4 then compares floor(t_lo) and ceil(t_lo) under G_j

(The paper's approximation to the jump value) The authors' own Algorithm 1, which their reference implementation follows, accumulates the weights \(w_i\) alone at breakpoints rounded to integers and takes the leftmost value on ties. It therefore minimizes \(\sum_i w_i \cdot \operatorname{dist}(t, T_i)\), where \(T_i\) is the rounded range of values of \(x_j\) that satisfy row \(i\): the violation is measured in units of \(x_j\) rather than in the units of each row. The two agree for binary variables, where only one candidate value exists, and for rows with unit coefficients and integer \(d_i\). Otherwise the paper's value is an approximation to the minimizer of \(G_j\). The scores below are exact in both.

Algorithm 3.2.12  FEASIBILITY-JUMP
                  (Luteberget and Sartor, Algorithm 2)

Input   rows a_i x <= b_i (an equation as two rows), bounds, a start x~
        (for instance a bound of each variable), a budget.
Output  a feasible point x*, or failure.

 1. x* <- none;  xbar <- x~;  w_i <- 1 for every row

 2. for every variable j:
       v_j <- Jump_j(xbar)
       s_j <- G_j(xbar_j) - G_j(v_j)   [the score: violation removed]
    P <- { j : s_j > 0 }

 3. while the budget is not exhausted:

 4.    if F^w(xbar) = 0: x* <- xbar; stop

 5.    if P is empty (a local minimum):
          w_i <- w_i + 1 for every violated row i
          recompute the scores of the variables in those rows and P
          pick a violated row at random and j* <- the variable in it
          with the largest score

 6.    else:
          sample up to 25 indices from P and let j* be the sampled
          index with the largest score

 7.    xbar_{j*} <- v_{j*};  recompute v_{j*}
       update s_j for every j sharing a row with j*;  update P

 8. return x*

Invariant
    a move with s_j > 0 strictly lowers F^w for the current weights;
    weights only grow, and only on rows violated at a local minimum,
    so the search is pushed toward the rows that are hard to satisfy
    together.

Cost per iteration
    the four items of Proposition 3.2.10,
    O(eta log eta + eta mu + m mu) in the worst case, and under a
    microsecond on many MIPLIB 2017 instances in the authors'
    implementation.

Parallel
    independent runs with different seeds and starts are a portfolio;
    a GPU variant recomputes every violation (one thread per row) and
    every jump value (one thread per variable) after each move or
    each batch of non-conflicting moves, trading O(nnz) work per step
    for full parallelism.

(The two mechanisms on a two-row example, and on R1) The smallest example shows the two mechanisms. Take the single row \(3x_1 + 2x_2 \ge 5\) over nonnegative integers at \(x = (0, 0)\), violated by 5 with weight 1. The jump \(x_1 \to 2\) and the jump \(x_2 \to 3\) both bring the violation to zero, so the two scores tie at 5. Add the row \(x_1 + x_2 \le 2\): the move \(x_2 \to 3\) now violates it by 1, and the jump of \(x_2\) becomes \(x_2 \to 2\), which leaves the first row violated by 1 instead, so its score falls to 4, while \(x_1 \to 2\) keeps the score 5 and wins. Each score touches only the rows that contain the variable, which is the whole reason the method is cheap. On the running example the figure's Feasibility Jump mode starts from the rounding \((6, 5)\) of the default start. There the first two rows are violated by \(12.5\) and \(9.5\) and the weighted violation is \(22.0\). The figure minimizes \(F^w\) exactly over each variable's range and takes the value nearer to the current one on ties. Moving \(y\) to \(1\) lowers the violation to \(1.5\), and \(1\) is the unique minimizer of \(G_y\) at \(x = 6\): the values \(0\) and \(2\) give \(1.8\) and \(3.5\). From \((6, 1)\) every value of \(x\) in \(\{0, \dots, 5\}\) brings the violation to zero, so \(G_x\) has six minimizers. The figure takes the nearest, \(x \to 5\), and stops at the feasible point \((5, 1)\). The smallest-on-ties rule of Algorithm 3.2.11 would take \(x \to 0\) and stop at \((0, 1)\), which is also feasible. On this polygon no local minimum is ever met, so the weight update never fires. The four-variable program below is a small one on which it does. The program is Algorithm 3.2.12 traced move by move, with the jump values computed by direct evaluation of \(G_j\), which is exact since every variable is binary.

Feasibility Jump on R1 from the rounding (6, 5) of the default start

  (6, 5)   2x + 5y <= 24.5 violated by 12.5, 5x + 2y <= 30.5 by 9.5
     |     F^w = 22.0
     |
     |     move y: G_y(t) at x = 6
     |         t       0      1      2
     |         G_y   1.8    1.5    3.5     1 is the unique minimizer
     v
  (6, 1)   F^w = 1.5
     |
     |     move x: G_x(t) at y = 1 is zero for t = 0, 1, 2, 3, 4, 5
     |
     +---- the figure: the nearest to 6 ---------> (5, 1) feasible
     |
     +---- Algorithm 3.2.11: the smallest -------> (0, 1) feasible

  no local minimum on this polygon: the weights are never raised
# Feasibility Jump (Luteberget and Sartor 2023), traced move by move.
#
# The instance is a four-variable binary program,
#   x1 + x2 + x3 = 1,  x1 + x4 >= 1,  x2 + x4 >= 1,  x3 + x4 <= 1
# (the equation is two inequalities). Rows are written a_i . x <= b_i;
# the feasible points are (1,0,0,1) and (0,1,0,1).

import random

random.seed(1)

ROWS = [({0: 1, 1: 1, 2: 1}, 1),
        ({0: -1, 1: -1, 2: -1}, -1),
        ({0: -1, 3: -1}, -1),
        ({1: -1, 3: -1}, -1),
        ({2: 1, 3: 1}, 1)]
NAMES = ["x1+x2+x3<=1", "x1+x2+x3>=1", "x1+x4>=1", "x2+x4>=1", "x3+x4<=1"]
n, m = 4, len(ROWS)
lo, up = [0] * n, [1] * n
w = [1.0] * m
cols = {j: [i for i, (row, _) in enumerate(ROWS) if j in row]
        for j in range(n)}

def viol(i, x):
    """The violation of row i at x."""
    row, b = ROWS[i]
    return max(0.0, sum(a * x[j] for j, a in row.items()) - b)

def Fw(x):
    """The weighted violation."""
    return sum(w[i] * viol(i, x) for i in range(m))

def G(j, t, x):
    """The weighted violation with x_j = t, rows containing x_j only."""
    y = list(x)
    y[j] = t
    return sum(w[i] * viol(i, y) for i in cols[j])

def jump(j, x):
    """The best value of x_j other than x[j], the smallest on ties."""
    return min((t for t in range(lo[j], up[j] + 1) if t != x[j]),
               key=lambda t: (G(j, t, x), t))

def score(j, x):
    """Violation removed by the best move of x_j."""
    return G(j, x[j], x) - G(j, jump(j, x), x)

x = [1, 1, 1, 0]
print("start x =", x, " weighted violation F_w =", Fw(x))
for step in range(1, 20):
    if Fw(x) == 0:
        print("feasible:", x)
        break
    P = [j for j in range(n) if score(j, x) > 1e-12]
    if not P:
        # a local minimum: raise the weights of the violated rows
        U = [i for i in range(m) if viol(i, x) > 0]
        for i in U:
            w[i] += 1
        istar = random.choice(U)
        jstar = max(ROWS[istar][0], key=lambda j: score(j, x))
        print(f"step {step}: local minimum, violated "
              f"{[NAMES[i] for i in U]}")
        print(f"        weights now {w}; best move inside it: x{jstar + 1}")
    else:
        jstar = max(random.sample(P, min(25, len(P))),
                    key=lambda j: score(j, x))
    v = jump(jstar, x)
    print(f"step {step}: scores "
          f"{[round(score(j, x), 1) for j in range(n)]}")
    print(f"        move x{jstar + 1}: {x[jstar]} -> {v}", end="")
    x[jstar] = v
    print(f"; x = {x}, F_w = {Fw(x)}")
start x = [1, 1, 1, 0]  weighted violation F_w = 2.0
step 1: scores [0.0, 0.0, 1.0, -1.0]
        move x3: 1 -> 0; x = [1, 1, 0, 0], F_w = 1.0
step 2: local minimum, violated ['x1+x2+x3<=1']
        weights now [2.0, 1.0, 1.0, 1.0, 1.0]; best move inside it: x1
step 2: scores [1.0, 1.0, -2.0, 0.0]
        move x1: 1 -> 0; x = [0, 1, 0, 0], F_w = 1.0
step 3: scores [-1.0, -2.0, -2.0, 1.0]
        move x4: 0 -> 1; x = [0, 1, 0, 1], F_w = 0.0
feasible: [0, 1, 0, 1]

(Reading the trace) At the start only \(x_3 \to 0\) has a positive score: setting \(x_1\) or \(x_2\) to 0 would repair one unit of the first row and break \(x_1 + x_4 \ge 1\) or \(x_2 + x_4 \ge 1\), for a net gain of zero. At \((1, 1, 0, 0)\) no single move improves the weighted violation, a local minimum. The violated row's weight rises to 2, after which \(x_1 \to 0\) has score \(2 - 1 = 1\) and is taken, and the newly violated row \(x_1 + x_4 \ge 1\) is repaired by \(x_4 \to 1\). Three moves, one weight update and a feasible point, with no LP anywhere. The program evaluates \(G_j\) on every row of the variable at every step. The incremental scheme of Algorithm 3.2.12 touches only the rows of the moved variable, and the GPU variant recomputes all rows and all jump values in bulk instead.

The loop of Algorithm 3.2.12 on the four-variable run

  +--> F^w(xbar) = 0 ? --- yes ---> x* <- xbar = (0, 1, 0, 1), stop
  |        | no
  |        v
  |    P empty ? --- yes ---> a local minimum, at step 2:
  |        | no               x = (1, 1, 0, 0); the weight of
  |        |                  x1+x2+x3 <= 1 goes from 1 to 2;
  |        |                  j* <- the best variable of that
  |        |                  row: x1
  |        v                               |
  |    j* <- the best of a sample of P     |
  |    (step 1: x3; step 3: x4)            |
  |        |                               |
  |        v                               |
  |    xbar_{j*} <- v_{j*}  <--------------+
  |    rescore the variables sharing a row with j*
  |        |
  +--------+

(What the method is worth) What Feasibility Jump achieves is measured, in the paper, by the primal integral and the time to a first feasible point rather than by the time to optimality. The method finds feasible solutions to 123 of the 240 instances of the MIPLIB 2017 benchmark set with no LP solve, in about 0.6 seconds on average. Inside FICO Xpress, where it has run by default since version 9.0, it changes the time to optimality by "about 3%" over the whole benchmark. It reduces the time to a first feasible solution by 25 % on average and by more than a factor of ten on about a tenth of the instances. The authors write that "Our C++ reference implementation is around 800 lines of code, with Xpress integration adding an additional 500 lines". The repository's header is 842 lines and its Xpress driver 665.Luteberget and Sartor (2023), Sections 3 and 4; the quotation is from Section 3. The line counts are from the repository at commit 93f1c2ae of 25 September 2026 (feasibilityjump.hh 842 lines, xpress_fj.cc 665 lines). The Xpress control is FEASIBILITYJUMP in the FICO Xpress Optimizer Reference Manual. HiGHS added the heuristic, on by default, in version 1.11.0 of 6 June 2025 (GitHub release notes); SCIP's development branch adds feasjump and local-search heuristics for SCIP 11 (SCIP CHANGELOG, master, read 4 October 2026); NVIDIA's cuOpt runs Feasibility Jump, the feasibility pump and local search on the GPU (cuOpt User Guide, version 26.08, "Introduction"). The method has since appeared in HiGHS, in SCIP's development branch, and on the GPU in cuOpt. Section 5.4 returns to the family of LP-free local searches it started.

Scheduling, and what heuristics are worth

A solver runs dozens of heuristics on a schedule. In SCIP each heuristic is a plugin with a frequency, an offset and a timing mask. These say at which depths of the tree and at which point of node processing it may run, before or after the LP. A priority orders the heuristics that are due at the same time. The schedule is shifted by phase. Rounding and diving heuristics are cheap enough to run at many nodes, and the sub-MIP heuristics are reserved for the root and for sparse calls in the tree.T. Achterberg, "SCIP: solving constraint integer programs", Mathematical Programming Computation 1 (2009), on the heuristic plugins and their timing; Berthold, Hendel and Koch (2018), cited in Section 3.1, on the phase-dependent schedule. SCIP 8's MINLP heuristics and their timing are in Bestuzheva et al. (2025), Section 2.5: sub-NLP, multistart, NLP diving, Undercover, and the MPEC heuristic, which writes the integrality of a binary \(y\) as the complementarity condition \(y(1 - y) = 0\) and solves the resulting mathematical program with equilibrium constraints locally. Gurobi exposes the schedule as a budget. The Heuristics parameter caps the share of the run spent in heuristics at 5 % by default, and RINS sets how often RINS runs. NoRelHeurTime and NoRelHeurWork allot time to a heuristic that searches for incumbents before the root relaxation has even been solved, which is the right regime when that LP takes minutes.Gurobi Optimizer Reference Manual, version 13.0, parameters Heuristics, RINS, NoRelHeurTime and NoRelHeurWork, docs.gurobi.com. In the global MINLP solvers the balance tilts toward the NLP side, because every pruning rule and every range reduction of Section 2.6 depends on the incumbent. BARON runs its multistart local search before the tree and local searches at nodes under DoLocal. SCIP runs sub-NLP whenever a heuristic or the LP hands it an integral point.

What the heuristics are worth depends on what is measured. In the component studies of the MILP engine (Section 5.1) the measure is time to optimality, and there switching the heuristics off costs a smaller factor than switching off cuts or presolve. This is the ranking of the two CPLEX studies as they are usually summarized, with the caveat Section 2.5 noted, and of the one open study, Mexi's ablation of SCIP 10, whose factors Section 5.1 quotes.R. E. Bixby, M. Fenelon, Z. Gu, E. Rothberg and R. Wunderling, "Mixed-integer programming: a progress report", in M. Grötschel (ed.), The Sharpest Cut (MPS-SIAM, 2004), 309–325, and T. Achterberg and R. Wunderling, "Mixed integer programming: analyzing 12 years of progress", in M. Jünger and G. Reinelt (eds), Facets of Combinatorial Optimization (Springer, 2013), 449–481, neither re-read for this series, as Section 2.5 noted; G. Mexi, The Two Faces of Mixed-Integer Programming: Primal and Dual Progress, doctoral thesis, Technische Universität Berlin (2026), doi 10.14279/depositonce-26588, Section 2.7.1 and Table 2.1, read 5 October 2026, cited in Section 2.5. Measured by the primal integral the picture changes. The Xpress numbers above, a 3 % change in total time against a 25 % change in the time to the first solution, are typical. Berthold introduced the primal integral (Definition 2.3.4) for exactly this reason, because a heuristic's contribution is invisible in a measure that only sees the end of the run.T. Berthold, "Measuring the impact of primal heuristics", Operations Research Letters 41 (2013). Koch, Berthold, Pedersen and Vanaret name RINS and local branching among the sources of the fifty-fold algorithmic speed-up of MILP codes between 2001 and 2020, in the sentence quoted above. They attribute no factor to them, because the features overlap.

What parallelizes

Every solver in this series runs the catalogue above, and the GPU reading is the one that opens Section 5.4 and Section 7.7. Almost everything on the primal side is independent work: fixing the integers and solving an NLP, projecting a rounding onto the relaxation, solving a sub-MIP on a neighbourhood, running a local search from a start point. The dual side of a node, the LP and its warm start, is sequential in the sense of Proposition 3.1.12. The primal side consists of independent tasks and can be run as a batch. This is why the heuristics moved to the GPU first, in cuOpt's Feasibility Jump, feasibility pump and local search. It is also why the natural first target for a parallel nonconvex MINLP solver is the incumbent: a GPU multistart over thousands of starts, a batched sub-NLP over every integer assignment in a solution pool, a bandit over neighbourhoods whose sub-MIPs run side by side. What such a solver cannot buy this way is the proof, which is the subject of the rest of this section.

What parallelizes: a chain on the dual side, a batch on the primal

  dual side, sequential in the sense of Proposition 3.1.12:
     parent's basis --> pivot --> pivot --> ... --> node bound

  primal side, independent tasks, run as a batch:
     fix the integers and solve an NLP ----------+
     project a rounding onto the relaxation -----+
     solve a sub-MIP on a neighbourhood ---------+---> incumbent
     run a local search from a start point ------+

Cutting planes

Branch and bound improves a bound by splitting the feasible set and relaxing each part separately. It leaves the rows of every node's relaxation as the formulation gave them, tightened only by the node's bounds and by the propagation of Section 2.6. A cutting plane improves the bound without splitting anything. It adds to the relaxation an inequality that every feasible point satisfies and that the current relaxed point violates, so that the relaxation shrinks toward the integer hull and its optimum moves.

Section 2.1 showed that the integer hull, the convex hull of the feasible integer points (Definition 2.1.3), is the ideal formulation, and that no solver can list its facets, the inequalities of its minimal description, in advance. This section is about the inequalities a solver can compute: how they are derived and how much of the gap they close. It also explains why the first thing a MILP solver does at the root is to run a loop of them before it branches at all.

The loop matters for the nonconvex case for two reasons. Every LP-based global solver named in this series (BARON, Couenne, SCIP, ANTIGONE, Xpress Global and Gurobi) bounds most of its nodes with a linear program, and that program is tightened by the same machinery. MAiNGO, which propagates McCormick relaxations through the expression graph instead (Section 3.5), is the exception. A reader who wants to understand what BARON or SCIP does at a node therefore has to understand the cut loop first. And nonlinearity supplies cuts of its own: the tangent planes of convex pieces, which are valid, and inequalities in the lifted variables of a quadratic problem, which are valid where the tangents are not. The section ends with those.

The running example is the two-variable program of the figures, which maximizes \(c^\top (x, y)\) over the integer points of the polygon \(2x + 5y \le 24.5\), \(5x + 2y \le 30.5\), \(-3x + 4y \le 11\), \(x - 2y \le 4.2\), \(x, y \ge 0\). Everywhere else in this section the problem is written as a minimization, as in the rest of the series.

Valid inequalities and split cuts

Throughout, the mixed-integer linear program is \(\min\{c^\top x : x \in \mathcal F\}\) with \(\mathcal F = P \cap (\mathbb{Z}^{I} \times \mathbb{R}^{n - |I|})\), where \(P = \{x : Ax \le b,\ l \le x \le u\}\) is the polyhedron of the LP relaxation and \(I\) is the set of integer coordinates. The integer hull is \(\operatorname{conv}(\mathcal F)\), a polyhedron when the data are rational (Theorem 2.1.4).

Definition 3.3.1 (valid inequality, cut, separation, efficacy, parallelism). An inequality \(\alpha^\top x \le \beta\) is valid for \(\mathcal F\) if every point of \(\mathcal F\) satisfies it. It is a cut for a point \(\bar x\) if \(\alpha^\top \bar x > \beta\). Its efficacy at \(\bar x\) is the Euclidean distance it cuts off, \((\alpha^\top \bar x - \beta)/\lVert \alpha \rVert_2\). Its objective parallelism is \(|\alpha^\top c| / (\lVert \alpha \rVert\, \lVert c \rVert)\), and the parallelism of two cuts is the cosine of the angle between their normals. To separate a point \(\bar x\) from a family of valid inequalities is to find one that \(\bar x\) violates, or to report that none exists, and the routine that does so for one family is a separator. Two further scores appear in Algorithm 3.3.10. The integer support of a cut is the share of its nonzero coefficients that fall on integer variables, and its directed cutoff distance at \(\bar x\) is the distance from \(\bar x\) to the cut's hyperplane measured along the segment from \(\bar x\) to the incumbent.

These scores are what a solver uses when it decides which of many candidate cuts to keep (Algorithm 3.3.10). A valid inequality with no violation is not useless, since it may become a cut at a later vertex, and solvers keep a pool of them for that reason.

Definition 3.3.2 (split cut). Let \((\pi, \pi_0) \in \mathbb{Z}^n \times \mathbb{Z}\) with \(\pi_j = 0\) for \(j \notin I\) give a split disjunction \(\pi^\top x \le \pi_0\) or \(\pi^\top x \ge \pi_0 + 1\) (Definition 3.1.2), which every point of \(\mathcal F\) satisfies. A split cut is an inequality valid for

\[\operatorname{conv}\Big( \big(P \cap \{\pi^\top x \le \pi_0\}\big) \cup \big(P \cap \{\pi^\top x \ge \pi_0 + 1\}\big) \Big).\]

A split disjunction is the branching of Definition 3.1.2 written as an inequality. Branching on a variable \(x_j\) at the fractional value \(\bar x_j\) is the disjunction with \(\pi = e_j\) and \(\pi_0 = \lfloor \bar x_j \rfloor\). The difference is what is done with it. Branching creates two children and solves both. A split cut keeps one problem and adds to it an inequality that both children satisfy. The convex hull of the two children's relaxations is in general strictly smaller than the parent's relaxation and strictly larger than the integer hull. Every cut family in this section is a way of describing part of that hull cheaply.

Chvátal–Gomory rounding

The Chvátal–Gomory cut has the shortest proof. Take a nonnegative combination of the rows, round its coefficients down, and round the right-hand side down.

Theorem 3.3.3 (Chvátal–Gomory cut; Gomory 1958, Chvátal 1973). Let \(P = \{x \in \mathbb{R}^n_+ : Ax \le b\}\) and \(X = P \cap \mathbb{Z}^n\). For every \(\lambda \in \mathbb{R}^m_+\) the inequality \(\lfloor \lambda^\top A \rfloor\, x \le \lfloor \lambda^\top b \rfloor\), with the floor taken in each coordinate, is valid for \(X\).R. E. Gomory, "Outline of an algorithm for integer solutions to linear programs", Bulletin of the American Mathematical Society 64 (1958); V. Chvátal, "Edmonds polytopes and a hierarchy of combinatorial problems", Discrete Mathematics 4 (1973).

Proof. For \(x \in P\), \(\lambda^\top A x \le \lambda^\top b\). Since \(x \ge 0\) and \(\lfloor \lambda^\top A \rfloor \le \lambda^\top A\) coordinatewise, \(\lfloor \lambda^\top A \rfloor x \le \lambda^\top A x \le \lambda^\top b\). If \(x\) is integral the left side is an integer, so it is at most \(\lfloor \lambda^\top b \rfloor\). ∎

The picture is that of rounding a line down until it passes through a lattice point. The combination \(\lambda^\top A x \le \lambda^\top b\) is a valid half-plane. Lowering its coefficients keeps it valid on the nonnegative orthant, and lowering its right-hand side to the next integer costs nothing on integer points, where the left side is an integer anyway.

Jeroslow's parity instance (J) of Section 3.1 is the cleanest example. The program \(2(x_1 + \dots + x_n) = n\) with \(n\) odd and \(x\) binary has no solution. Theorem 3.1.14, following Jeroslow, gives at least \(2^{(n+1)/2}\) leaves for every tree that branches on single variables, and Proposition 3.1.15 the exact count \(2\binom{n+1}{(n+1)/2} - 1\) nodes, however the nodes are ordered. Proposition 3.1.16 derived the two Chvátal–Gomory cuts that make the root infeasible. They are Theorem 3.3.3 with \(\lambda = \tfrac12\) on each of the two rows \(2\sum_j x_j \le n\) and \(-2\sum_j x_j \le -n\), and either alone settles the instance at the root. That is what the figure of Section 3.1 shows when its cut is switched on.

Two closures belong to this theorem. For a pure integer program with \(P = \{x \ge 0 : Ax \le b\}\), the Chvátal–Gomory closure of \(P\) is the intersection over all \(\lambda \ge 0\) of the half-spaces of Theorem 3.3.3. The split closure of \(P\) is the intersection over all \((\pi, \pi_0)\) of the convex hulls of Definition 3.3.2. Three facts frame them. Chvátal showed that for a bounded \(P\) finitely many rounds of the rounding operation, each applied to the result of the last, reach the integer hull: the Chvátal rank of a bounded rational polyhedron is finite.Chvátal (1973), cited above. Schrijver showed that the closure of a rational polyhedron under one round is again a polyhedron.A. Schrijver, "On cutting planes", Annals of Discrete Mathematics 9 (1980). And deciding whether a given point lies in the Chvátal–Gomory closure, which is what separating these cuts amounts to, is NP-hard, so a solver does not separate over the closure.F. Eisenbrand, "On the membership problem for the elementary closure of a polyhedron", Combinatorica 19 (1999).

A solver uses two cheap special cases instead. Gomory's 1958 fractional cut takes \(\lambda\) to be a row of the inverse of the optimal LP basis. From a tableau row \(x_i + \sum_j \bar a_{ij} x_j = \bar b_i\) of a pure integer program with fractional \(\bar b_i\) it gives the inequality \(\sum_j (\bar a_{ij} - \lfloor \bar a_{ij} \rfloor) x_j \ge \bar b_i - \lfloor \bar b_i \rfloor\). The \(\{0, \tfrac12\}\)-cuts of Caprara and Fischetti take \(\lambda \in \{0, \tfrac12\}^m\), so that separating them depends only on the parity of the coefficients and on the row slacks. Their separation is NP-complete in general, polynomial when each column or each row of the matrix has at most two odd entries, and heuristic in practice, under the name "zero-half cuts" in the logs of every major solver.A. Caprara and M. Fischetti, "{0, 1/2}-Chvátal–Gomory cuts", Mathematical Programming 74 (1996); the polynomial cases, via a reduction to a minimum-weight binary clutter problem, are as restated by L. Brandl and A. S. Schulz, arXiv 2311.03909 (2023); the practical separators are A. M. C. A. Koster, A. Zymolka and M. Kutschka, "Algorithms to separate {0, 1/2}-Chvátal–Gomory cuts", Algorithmica 55 (2009).

The Gomory mixed-integer cut

The cut every MILP solver generates first is Gomory's mixed-integer cut of 1960, not the fractional cut of 1958. It is read off a row of the simplex tableau, it handles continuous variables, and in the space of the nonbasic variables it dominates the fractional cut coefficient by coefficient. Write the relaxation in standard form \(Ax = b\), \(x \ge 0\) after adding a slack variable to each row. The basis \(B\), the nonbasic set, written \(N\) here, and the basic solution are those of Section 3.1 (the paragraph before Proposition 3.1.12), with every nonbasic variable at its lower bound zero, so that \(x_B = B^{-1} b\). One more object is needed, the tableau row of a basic variable \(x_i\), which expresses it in terms of the nonbasic variables,

\[x_i + \sum_{j \in N} \bar a_{ij}\, x_j = \bar b_i, \qquad \bar A = B^{-1} A_N,\quad \bar b = B^{-1} b,\]

and the row is an identity on the whole feasible set, not only at the vertex, because it is a linear combination of the equations \(Ax = b\). That is what makes a cut derived from it valid everywhere.R. E. Gomory, "An algorithm for the mixed integer problem", RAND Research Memorandum RM-2597 (1960); the title is as cited by G. Cornuéjols, "Valid inequalities for mixed integer linear programs", Mathematical Programming 112 (2008), whose survey gives the modern derivations; the RAND page itself was not reachable when this was checked.

Theorem 3.3.4 (Gomory mixed-integer cut; Gomory 1960). Consider a row

\[x_i + \sum_{j \in N_I} a_j x_j + \sum_{j \in N_C} a_j x_j = \beta, \qquad x_i \in \mathbb{Z},\quad x_j \ge 0\ (j \in N_I \cup N_C),\quad x_j \in \mathbb{Z}\ (j \in N_I),\]

in which \(N_I\) and \(N_C\) index the integer and the continuous nonbasic variables, and suppose \(f_0 = \beta - \lfloor \beta \rfloor \in (0, 1)\). Write \(f_j = a_j - \lfloor a_j \rfloor\). Then every feasible point satisfies

\[\sum_{\substack{j \in N_I \\ f_j \le f_0}} \frac{f_j}{f_0}\, x_j \;+\; \sum_{\substack{j \in N_I \\ f_j > f_0}} \frac{1 - f_j}{1 - f_0}\, x_j \;+\; \sum_{\substack{j \in N_C \\ a_j > 0}} \frac{a_j}{f_0}\, x_j \;+\; \sum_{\substack{j \in N_C \\ a_j < 0}} \frac{-a_j}{1 - f_0}\, x_j \;\ge\; 1. \tag{3.3.1}\]

Proof. Put

\[s^+ = \sum_{j \in N_I,\ f_j \le f_0} f_j x_j + \sum_{j \in N_C,\ a_j > 0} a_j x_j \ \ge 0, \qquad s^- = \sum_{j \in N_I,\ f_j > f_0} (1 - f_j) x_j + \sum_{j \in N_C,\ a_j < 0} (-a_j) x_j \ \ge 0 .\]

For an integer nonbasic variable in the second group write \(f_j x_j = x_j - (1 - f_j) x_j\). Then the row becomes \(s^+ - s^- = f_0 - k\) with

\[k = x_i - \lfloor \beta \rfloor + \sum_{j \in N_I} \lfloor a_j \rfloor x_j + \sum_{j \in N_I,\ f_j > f_0} x_j \ \in \mathbb{Z} .\]

If \(k \le 0\) then \(s^+ \ge s^+ - s^- = f_0 - k \ge f_0\) and so \(s^+ / f_0 \ge 1\). If \(k \ge 1\) then \(s^- \ge s^- - s^+ = k - f_0 \ge 1 - f_0\) and so \(s^- / (1 - f_0) \ge 1\). In both cases \(s^+/f_0 + s^-/(1 - f_0) \ge 1\), which is (3.3.1). ∎

Theorem 3.3.4: the GMI cut as a split on the integer k

     the row:  s+ - s- = f0 - k,  k an integer,  s+ >= 0, s- >= 0
                    /                                 \
           k <= 0  /                                   \  k >= 1
                  /                                     \
  s+ >= s+ - s- = f0 - k >= f0    s- >= s- - s+ = k - f0 >= 1 - f0
  so s+ / f0 >= 1                 so s- / (1 - f0) >= 1
                  \                                     /
                   +------------------+----------------+
                                      | in both cases
                                      v
                  s+ / f0 + s- / (1 - f0) >= 1,  that is (3.3.1)

  s+  the terms f_j x_j (j in N_I, f_j <= f0) and a_j x_j (j in N_C,
      a_j > 0)
  s-  the terms (1 - f_j) x_j (j in N_I, f_j > f0) and -a_j x_j
      (j in N_C, a_j < 0)
  at the vertex every nonbasic x_j is 0, so s+ = s- = 0: the vertex
  violates the cut by exactly one

(What the GMI cut is, and when it is weak) The proof says what the cut is. It is the split cut from the disjunction \(k \le 0\) or \(k \ge 1\) on the integer combination \(k\), with each nonbasic variable charged to the side of the disjunction on which it helps least. At the current vertex every nonbasic variable is zero, so the left side of (3.3.1) is zero and the vertex violates the cut by exactly one in this normalization: a GMI cut always separates the vertex it was derived from. Two features of the coefficients are worth noticing. Every coefficient of an integer variable is at most one, since \(f_j / f_0 \le 1\) when \(f_j \le f_0\) and \((1 - f_j)/(1 - f_0) < 1\) when \(f_j > f_0\). The fractional cut of 1958 divided by \(f_0\) has coefficients \(f_j / f_0\) throughout, so (3.3.1) is at least as strong coefficient by coefficient in the space of the nonbasic variables. And the continuous coefficients are \(a_j / f_0\) or \(-a_j / (1 - f_0)\). A row with large continuous coefficients, or a \(\beta\) whose fractional part is near \(0\) or \(1\), therefore gives a weak and numerically delicate cut, and SCIP refuses a row with \(\min(f_0, 1 - f_0) < 0.01\).SCIP 10.0.0, src/scip/sepa_gomory.c, parameter separating/gomory/away with default 0.01; the source is at github.com/scipopt/scip.

The GMI cut on the running example

At the \(45^\circ\) objective the LP maximizes \(x + y\), and its optimum is the vertex where the first two rows are tight. Scale those rows to integer data first, so that their slacks are integers at every integer point:

\[4x + 10y + s_1 = 49, \qquad 10x + 4y + s_2 = 61, \qquad s_1, s_2 \in \mathbb{Z}_+ .\]

The basis is \(B = \begin{pmatrix} 4 & 10 \\ 10 & 4 \end{pmatrix}\) with \(\det B = -84\) and \(B^{-1} = \begin{pmatrix} -1/21 & 5/42 \\ 5/42 & -1/21 \end{pmatrix}\), and the two tableau rows are

\[x - \tfrac{1}{21}\, s_1 + \tfrac{5}{42}\, s_2 = \tfrac{69}{14}, \qquad y + \tfrac{5}{42}\, s_1 - \tfrac{1}{21}\, s_2 = \tfrac{41}{14}.\]

(From the tableau row to the cut) The vertex is \((69/14, 41/14) = (4.929, 2.929)\) with bound \(55/7 = 7.857\), or \(5.556\) in the figure's units, where \(c = (\cos 45^\circ, \sin 45^\circ)\) divides every value by \(\sqrt 2\). Both coordinates have fractional part \(13/14\). From the row of \(x\), \(f_0 = 13/14\). The slack \(s_1\) has \(a = -1/21\), so \(f = 20/21 > f_0\) and its coefficient is \((1 - 20/21)/(1 - 13/14) = (1/21)/(1/14) = 2/3\). The slack \(s_2\) has \(a = 5/42 \le f_0\), so its coefficient is \((5/42)/(13/14) = 5/39\). The GMI cut is

\[\tfrac23\, s_1 + \tfrac{5}{39}\, s_2 \ \ge\ 1,\]

and substituting \(s_1 = 49 - 4x - 10y\) and \(s_2 = 61 - 10x - 4y\) and clearing denominators gives

\[11x + 20y \;\le\; 110. \tag{3.3.2}\]

(What one cut closes) At the vertex, \(11 \cdot 69/14 + 20 \cdot 41/14 = 1579/14 = 112.79\), so the cut is violated by \(39/14 = 2.786\), with efficacy \(2.786 / \sqrt{11^2 + 20^2} = 0.122\). Checking (3.3.2) against the 22 integer points of the polygon, none is cut off, and none lies on the cut: a GMI cut is valid, but it is not in general a facet of the integer hull. Re-solving the LP with (3.3.2) added moves the vertex to \((5, 11/4)\) with bound \(31/4 = 7.75\), which is \(5.480\) in the figure's units. The integer optimum is \(7\), or \(4.950\), so this one cut closes \((7.857 - 7.75)/(7.857 - 7) = 12.5\%\) of the root gap. The figure below prints the same cut scaled so that the vertex violates it by one, \(3.949x + 7.179y \le 39.487\), and rounds the share to \(13\%\).

(GMI against Chvátal–Gomory on the same row) The same row also gives a Chvátal–Gomory cut, and the comparison is instructive. The fractional cut \(\tfrac{20}{21} s_1 + \tfrac{5}{42} s_2 \ge \tfrac{13}{14}\) substitutes to \(x + 2y \le 10.6\), whose efficacy at the vertex is \(0.083\), less than the GMI cut's \(0.122\), as the coefficient comparison above predicts. But \(x + 2y\) is an integer at every integer point, so a second rounding gives \(x + 2y \le 10\), which is tight at \((2, 4)\) and \((4, 3)\) and has efficacy \(0.786/\sqrt 5 = 0.351\), nearly three times the GMI cut's. Dominance of one cut family over another is a statement about one derivation step in one space. Solvers scale their GMI cuts to integer coefficients and round again for exactly this reason.

One tableau row, two cuts (R1, maximize x + y)

     row of x:  x - (1/21) s_1 + (5/42) s_2 = 69/14,  f0 = 13/14
                  /                                \
   GMI, Theorem 3.3.4:                     fractional cut of 1958,
   s_1: f = 20/21 > f0, coefficient 2/3    with the fractional parts
   s_2: f = 5/42 <= f0, coefficient 5/39   20/21, 5/42 and f0
                |                                   |
 (2/3) s_1 + (5/39) s_2 >= 1     (20/21) s_1 + (5/42) s_2 >= 13/14
                |                                   |
                +-- s_1 = 49 - 4x - 10y,            |
                |   s_2 = 61 - 10x - 4y  -----------+
                v                                   v
 11x + 20y <= 110   (3.3.2)          x + 2y <= 10.6
 efficacy 0.122                      efficacy 0.083
                                                    | x + 2y is an
                                                    v integer: round
                                     x + 2y <= 10, efficacy 0.351,
                                     tight at (2, 4) and (4, 3)

  re-solved with (3.3.2): the vertex (69/14, 41/14), bound
  55/7 = 7.857, moves to (5, 11/4), bound 31/4 = 7.75, which closes
  12.5% of the root gap to the integer optimum 7
Rounding x + 2y ≤ 10.6 down to x + 2y ≤ 10 (R1, maximize x + y). No integer point lies on the fractional cut from the row of x; since x + 2y is an integer at every integer point, the second rounding is valid, and it is tight at (2, 4) and (4, 3). The tinted sliver is the part of R1 each cut removes, and neither holds an integer point.

SCIP's Gomory separator tries a strengthened Chvátal–Gomory cut from each aggregated row as well as the GMI cut, and by default adds both. With the flag genbothgomscg switched off it keeps the more efficacious of the two.SCIP 10.0.0, src/scip/sepa_gomory.c, parameters trystrongcg and genbothgomscg, both TRUE by default; the numbers of this paragraph are recomputed by the script below. The general equivalences between GMI, MIR and split cuts are in Cornuéjols (2008), cited above.

The derivation above, with the bookkeeping a solver adds, is the following algorithm.

Algorithm 3.3.5  GMI-CUT
                 (the Gomory mixed-integer cut from one tableau row)

Input   the optimal basis B of the node LP in standard form A x = b,
        l <= x <= u (slacks added); a basic integer variable x_i with
        fractional value; its tableau row
           x_i + sum_{j in N} abar_ij x_j = bbar_i ;
        for each nonbasic j, whether it sits at l_j or at u_j.
Output  a cut alpha^T x <= beta, valid for the mixed-integer set and
        violated by the LP point, or "reject".

1. Complement the nonbasics to zero:
      for j at its upper bound substitute x_j = u_j - x'_j
         (the coefficient changes sign);
      for j at its lower bound x_j = l_j + x'_j;
      move the constants into bbar_i.
   Now every x'_j >= 0 and every x'_j = 0 at the current point.

2. f0 := frac(bbar_i).
   If min(f0, 1 - f0) < away (SCIP: 0.01) return "reject".

3. For j in N_I: f_j := frac(abar_ij);
      g_j := f_j / f0 if f_j <= f0, else (1 - f_j)/(1 - f0).
   For j in N_C:
      g_j := abar_ij / f0 if abar_ij > 0, else -abar_ij / (1 - f0).

4. The cut is sum_j g_j x'_j >= 1. Undo the complementation of
   step 1; substitute the slacks s = b - A x to express the cut in
   the structural variables.

5. Clean up:
      move coefficients below 1e-9 to the right-hand side through
         the variable's bound (the safe side);
      scale so that the largest coefficient is of order one, or to
         integers;
      reject if the coefficient range exceeds 1e4
         (SCIP: separating/maxcoefratio) or if the efficacy is below
         1e-4 (SCIP: separating/minefficacy).

Invariant
    steps 1 to 4 are Theorem 3.3.4 applied to a row that is an
    identity on the feasible set, so the cut is valid; the current
    point has every x'_j = 0 and violates the normalized cut by 1.

The cost of one cut is the tableau row itself, \(e_i^\top B^{-1} A_N\). The simplex method keeps its basis as a triangular factorization, the LU factors of \(B\), and a solve with the transposed factors is called a BTRAN. The row costs one BTRAN followed by one sparse row-times-matrix product. That is \(O(\mathrm{nnz}(A))\) in the worst case and far less when the factors are hyper-sparse, that is, when they hold far fewer nonzeros than their dimension suggests. Generating cuts from all \(k\) fractional rows costs \(k\) independent solves with the same factorization, which is a natural batch for threads or for a device.

Mixed-integer rounding

The GMI cut is one member of a family with a two-line proof of its own. Mixed-integer rounding says that an integer variable which wants to exceed the floor of a bound must pay, for every unit it exceeds by, the continuous slack that the first unit costs: one minus the fractional part of the bound.

Theorem 3.3.6 (Mixed-integer rounding; Nemhauser and Wolsey 1990, in the form of Marchand and Wolsey 2001). (a) For \(X = \{(y, s) \in \mathbb{Z} \times \mathbb{R}_+ : y - s \le b\}\) with \(f_0 = b - \lfloor b \rfloor > 0\), the inequality \(y \le \lfloor b \rfloor + s/(1 - f_0)\) is valid.G. L. Nemhauser and L. A. Wolsey, "A recursive procedure to generate all cuts for 0–1 mixed integer programs", Mathematical Programming 46 (1990); H. Marchand and L. A. Wolsey, "Aggregation and mixed integer rounding to solve MIPs", Operations Research 49 (2001). (b) For \(X = \{(x, s) \in \mathbb{Z}^n_+ \times \mathbb{R}_+ : \sum_j a_j x_j - s \le b\}\) with \(f_j = a_j - \lfloor a_j \rfloor\),

\[\sum_j \Big( \lfloor a_j \rfloor + \frac{\max(f_j - f_0, 0)}{1 - f_0} \Big) x_j \;-\; \frac{s}{1 - f_0} \;\le\; \lfloor b \rfloor .\]

Proof of (a). If \(y \le \lfloor b \rfloor\) the inequality holds because \(s \ge 0\). Otherwise \(q = y - \lfloor b \rfloor \ge 1\) is an integer and \(s \ge y - b = q - f_0\), so \(s/(1 - f_0) \ge (q - f_0)/(1 - f_0) \ge q\), the last step because \(q - f_0 \ge q - q f_0\) when \(q \ge 1\). ∎

Proof sketch of (b). First relax the row. For every \(j\) with \(f_j \le f_0\) replace \(a_j\) by \(\lfloor a_j \rfloor\). Since \(x_j \ge 0\) this lowers the left side, so the relaxed row is still valid. For every \(j\) with \(f_j > f_0\) write \(a_j x_j = \lfloor a_j \rfloor x_j + x_j - (1 - f_j) x_j\). The relaxed row is now \(y - s'' \le b\) with the integer \(y = \sum_j \lfloor a_j \rfloor x_j + \sum_{f_j > f_0} x_j\) and the nonnegative \(s'' = s + \sum_{f_j > f_0} (1 - f_j) x_j\). Apply (a) to it: \(y \le \lfloor b \rfloor + s''/(1 - f_0)\). Expanding, the coefficient of \(x_j\) for \(f_j > f_0\) is \(\lfloor a_j \rfloor + 1 - (1 - f_j)/(1 - f_0) = \lfloor a_j \rfloor + (f_j - f_0)/(1 - f_0)\), and for \(f_j \le f_0\) it is \(\lfloor a_j \rfloor\). Nothing has to be dropped afterwards. ∎

An instance of (a): for \(y - s \le 2.5\) the inequality is \(y \le 2 + 2s\). At \(s = 0\) it forces \(y \le 2\), and each further unit of \(y\) costs half a unit of \(s\): \(y = 3\) needs \(s \ge 0.5\), which is what the row itself demands.

(The MIR inequality of a tableau row is the GMI cut) Two facts make this family central. Applied to a tableau row, with the nonbasic variables complemented to zero as in Algorithm 3.3.5, the MIR inequality is the GMI cut.Marchand and Wolsey (2001), cited above. On the running row this can be checked by hand. Read \(x - \tfrac{1}{21} s_1 + \tfrac{5}{42} s_2 = \tfrac{69}{14}\) as the inequality \(x - \tfrac{1}{21} s_1 + \tfrac{5}{42} s_2 \le \tfrac{69}{14}\) with every variable integer and no continuous slack, so that \(f_0 = 13/14\). The coefficient of \(x\) stays \(1\). For \(s_1\), \(\lfloor a \rfloor = -1\) and \(f = 20/21 > f_0\), so the coefficient is \(-1 + (20/21 - 13/14)/(1/14) = -1 + 1/3 = -2/3\). For \(s_2\), \(f = 5/42 \le f_0\) and the coefficient is \(\lfloor 5/42 \rfloor = 0\). The MIR inequality is \(x - \tfrac23 s_1 \le 4\). Substituting \(x = \tfrac{69}{14} + \tfrac{1}{21} s_1 - \tfrac{5}{42} s_2\) from the row gives \(\tfrac{13}{21} s_1 + \tfrac{5}{42} s_2 \ge \tfrac{13}{14}\), and dividing by \(13/14\) gives \(\tfrac23 s_1 + \tfrac{5}{39} s_2 \ge 1\), the GMI cut of the previous section.

The row of x, two derivations, one cut (R1, maximize x + y)

                x - (1/21) s_1 + (5/42) s_2 = 69/14
                   /                           \
  GMI, Theorem 3.3.4,             MIR, Theorem 3.3.6 (b), on the
  f0 = 13/14:                     row read as <=: every variable
  s_1: (1 - 20/21)/(1 - 13/14)    integer, no continuous slack,
       = 2/3                      f0 = 13/14:
  s_2: (5/42)/(13/14) = 5/39      x:   stays 1
                |                 s_1: -1 + 1/3 = -2/3
                |                 s_2: floor(5/42) = 0
                |                 x - (2/3) s_1 <= 4
                |                          | x from the row
                |                          v
                |                 (13/21) s_1 + (5/42) s_2 >= 13/14
                |                          | divided by 13/14
                 \                        /
                  v                      v
                 (2/3) s_1 + (5/39) s_2 >= 1

(Closures, and the c-MIR recipe) The second fact is about closures. The closure of all MIR inequalities over all nonnegative aggregations of rows equals the split closure, which is itself a polyhedron.Nemhauser and Wolsey (1990) for the equality; W. Cook, R. Kannan and A. Schrijver, "Chvátal closures for mixed integer programming problems", Mathematical Programming 47 (1990), for the polyhedrality of the split closure. Separating a split cut in general is NP-hard. But optimizing over the split closure offline, which can be done by a sequence of mixed-integer programs, closes most of the integrality gap on many benchmark instances. That is the theoretical reason the GMI and MIR families carry so much of the weight in a solver.A. Caprara and A. N. Letchford, "On the separation of split cuts and related inequalities", Mathematical Programming 94 (2003); E. Balas and A. Saxena, "Optimizing over the split closure", Mathematical Programming 113 (2008); S. Dash, O. Günlük and A. Lodi, "MIR closures of polyhedral sets", Mathematical Programming 121 (2010). Marchand and Wolsey's practical recipe, called c-MIR, is what solvers run. Aggregate a few rows, up to five or six, along a path through shared continuous variables. Complement the variables that sit at a bound. Scale the aggregated row by each of a small set of divisors, apply (b), and keep the cut with the largest violation.

Lift-and-project: cuts from a disjunction

The split disjunction of Definition 3.3.2 has, for a binary variable, only two pieces, and the convex hull of two polyhedra has an explicit description. Balas' theorem on that hull, which Section 4.2 states in general and proves, gives a cut family and an algorithm for finding the deepest cut.

Theorem 3.3.7 (Lift-and-project; Balas 1979, Balas, Ceria and Cornuéjols 1993). Let \(P = \{x \in \mathbb{R}^n : Ax \ge b\}\) include \(0 \le x_j \le 1\) for \(j \in I\) among its rows and let \(j \in I\). Then

\[P_j := \operatorname{conv}\big( (P \cap \{x_j = 0\}) \cup (P \cap \{x_j = 1\}) \big)\]

is a polyhedron. When both pieces are nonempty, an inequality \(\alpha^\top x \ge \beta\) is valid for \(P_j\) if and only if there are \(\lambda, \mu \in \mathbb{R}^m_+\) and \(\lambda_0, \mu_0 \ge 0\) with \(\alpha = \lambda^\top A - \lambda_0 e_j = \mu^\top A + \mu_0 e_j\) and \(\beta \le \min(\lambda^\top b,\ \mu^\top b + \mu_0)\). The most violated such inequality at a point \(\bar x\), under a normalization of \((\lambda, \mu)\), is the solution of a linear program, the cut-generating LP. Applying the operation to the binary variables in sequence gives \((\cdots (P_{j_1})_{j_2} \cdots)_{j_{|I|}} = \operatorname{conv}(P \cap \{x : x_j \in \{0, 1\} \text{ for } j \in I\})\).E. Balas, "Disjunctive programming", Annals of Discrete Mathematics 5 (1979); E. Balas, S. Ceria and G. Cornuéjols, "A lift-and-project cutting plane algorithm for mixed 0–1 programs", Mathematical Programming 58 (1993).

Proof sketch. If \(\alpha^\top x \ge \beta\) has multipliers of the stated form, it is implied by \(Ax \ge b\) and \(-x_j \ge 0\) on the first piece and by \(Ax \ge b\) and \(x_j \ge 1\) on the second. So it is valid on both pieces, and an inequality valid on two sets is valid on their convex hull. Conversely, an inequality valid for a nonempty polyhedron is a nonnegative combination of its defining inequalities by Farkas' lemma, applied to each piece. Sequential convexification holds because \(P_j \cap \{x_k = 0\}\) and \(P_j \cap \{x_k = 1\}\) contain the corresponding faces of the integer hull and the operation is monotone in its argument. The full argument is in Balas (1979). ∎

Theorem 3.3.7: one inequality, valid on both pieces

     piece x_j = 0 of P                 piece x_j = 1 of P
     A x >= b    times lambda >= 0      A x >= b   times mu >= 0
     -x_j >= 0   times lambda_0 >= 0    x_j >= 1   times mu_0 >= 0
               |                                  |
               v                                  v
  alpha = lambda^T A - lambda_0 e_j    alpha = mu^T A + mu_0 e_j
  beta <= lambda^T b                   beta <= mu^T b + mu_0
               \                                  /
                +-------> alpha^T x >= beta <----+
                     valid on both pieces, so on P_j

  one binary variable after another:
     P --> P_{j_1} --> (P_{j_1})_{j_2} --> ... --> after all of I,
     conv of the points of P with x_j in {0, 1} for every j in I

(Lift-and-project without the lifted program) The picture is that of a lifted space with one copy of the variables for each piece of the disjunction and a convex combination tying them together. The cut-generating LP searches that description for the inequality the current vertex violates most. Balas and Perregaard then showed that the optimal solutions of the cut-generating LP correspond to GMI cuts read from rows of the original tableau after a sequence of pivots. Lift-and-project can therefore be run without ever forming the lifted program, by pivoting in the simplex tableau, and that is how the solvers that have it implement it.E. Balas and M. Perregaard, "A precise correspondence between lift-and-project cuts, simple disjunctive cuts, and mixed integer Gomory cuts for 0–1 programming", Mathematical Programming 94 (2003). A computational study of optimizing over the lift-and-project closure is P. Bonami, "On optimizing over lift-and-project closures", Mathematical Programming Computation 4 (2012). For a convex MINLP the same construction works with convex pieces. Stubbs and Mehrotra wrote the hull of \(\{g(x) \le 0, x_j = 0\} \cup \{g(x) \le 0, x_j = 1\}\) through the perspective of \(g\), so that the cut-generating problem becomes a convex program. Kılınç, Linderoth and Luedtke replaced each piece by a polyhedral outer approximation, refined iteratively, so that the cut-generating problem is again a linear program, and its disjunctive cuts strengthen the outer-approximation master of Section 3.4.R. A. Stubbs and S. Mehrotra, "A branch-and-cut method for 0–1 mixed convex programming", Mathematical Programming 86 (1999); M. R. Kılınç, J. Linderoth and J. Luedtke, "Lift-and-project cuts for convex mixed integer nonlinear programs", Mathematical Programming Computation 9 (2017).

Covers and flow covers: cuts from the structure of a row

The families above read the tableau. Two further families read the original rows for structure that every solver recognizes: a knapsack row over binaries, and a row that caps a continuous flow by an on–off switch.

Proposition 3.3.8 (Cover inequalities; Balas 1975, Wolsey 1975, Hammer, Johnson and Peled 1975). Let \(K = \{x \in \{0,1\}^n : \sum_j a_j x_j \le b\}\) with \(a_j > 0\). A set \(C\) with \(\sum_{j \in C} a_j > b\) is a cover, and \(\sum_{j \in C} x_j \le |C| - 1\) is valid for \(K\). If \(C\) is minimal, so that every proper subset fits, the inequality defines a facet of \(\operatorname{conv}(K \cap \{x_j = 0 : j \notin C\})\) and can be lifted, one variable at a time, to a facet of \(\operatorname{conv}(K)\).E. Balas, "Facets of the knapsack polytope", Mathematical Programming 8 (1975); L. A. Wolsey, "Faces for a linear inequality in 0–1 variables", Mathematical Programming 8 (1975); P. L. Hammer, E. L. Johnson and U. N. Peled, "Facet of regular 0–1 polytopes", Mathematical Programming 8 (1975).

Proof of validity. If every \(x_j\) with \(j \in C\) were \(1\), the row would be violated. ∎

An example with numbers: take \(3x_1 + 4x_2 + 5x_3 + 6x_4 \le 9\). The set \(C = \{1, 2, 3\}\) has weight \(12 > 9\), so \(x_1 + x_2 + x_3 \le 2\) is valid. It is minimal because any two of the three fit: \(3 + 4\), \(3 + 5\) and \(4 + 5\) are all at most \(9\). The fractional point \((1, 1, 0.4, 0)\) satisfies the row with equality and violates the cover inequality by \(0.4\), so it is a cut there. Lifting \(x_4\): with \(x_4 = 1\) only three units of capacity remain and at most one of the three items fits, so \(x_1 + x_2 + x_3 + x_4 \le 2\) is valid for all of \(K\), and it is tight at \((1, 0, 0, 1)\).

items or pointweight, against the capacity 9what it shows
\(\{1, 2\}\)\(3 + 4 = 7\)fits
\(\{1, 3\}\)\(3 + 5 = 8\)fits
\(\{2, 3\}\)\(4 + 5 = 9\)fits, so \(C\) is minimal
\(\{1, 2, 3\}\)\(3 + 4 + 5 = 12 > 9\)a cover, so \(x_1 + x_2 + x_3 \le 2\)
\((1, 1, 0.4, 0)\)\(3 + 4 + 0.4 \cdot 5 = 9\)the row holds with equality; the cover inequality is violated by \(0.4\)
\((1, 0, 0, 1)\)\(3 + 6 = 9\)with \(x_4 = 1\) three units are left, and at most one of items 1, 2, 3 fits; lifting gives \(x_1 + x_2 + x_3 + x_4 \le 2\), tight here
Cover and lifting on 3x_1 + 4x_2 + 5x_3 + 6x_4 <= 9

The lifting statement is proved in the three papers. The computational recipe, a greedy cover built from the LP point and sequential lifting coefficients obtained by solving small knapsack problems, is due to Crowder, Johnson and Padberg. Their 1983 paper showed that preprocessing each row with cover inequalities inside an LP-based branch and bound solves real industrial 0–1 programs. It is the beginning of branch and cut as it is practised.H. Crowder, E. L. Johnson and M. Padberg, "Solving large-scale zero-one linear programming problems", Operations Research 31 (1983); Z. Gu, G. L. Nemhauser and M. W. P. Savelsbergh, "Lifted cover inequalities for 0–1 integer programs: computation", INFORMS Journal on Computing 10 (1998).

Theorem 3.3.9 (Flow cover; Padberg, Van Roy and Wolsey 1985). Let \(X = \{(x, y) \in \mathbb{R}^n_+ \times \{0,1\}^n : \sum_j x_j \le b,\ x_j \le u_j y_j\}\), a single node whose inflows are capped by switches. Let \(C\) with \(\sum_{j \in C} u_j = b + \delta\) and \(\delta > 0\) be a flow cover. Then

\[\sum_{j \in C} x_j \;+\; \sum_{j \in C} (u_j - \delta)^+ (1 - y_j) \;\le\; b\]

is valid for \(X\).M. W. Padberg, T. J. Van Roy and L. A. Wolsey, "Valid linear inequalities for fixed charge problems", Operations Research 33 (1985); lifted and complemented versions in Z. Gu, G. L. Nemhauser and M. W. P. Savelsbergh, "Lifted flow cover inequalities for mixed 0–1 integer programs", Mathematical Programming 85 (1999).

Proof sketch. Fix a feasible point and let \(S = \{j \in C : y_j = 0\}\). Then \(\sum_{j \in C} x_j \le \min\big(b,\ \sum_{j \in C \setminus S} u_j\big) = \min\big(b,\ b + \delta - \sum_{j \in S} u_j\big)\). If no \(j \in S\) has \(u_j > \delta\) the second sum of the inequality is zero and there is nothing to prove. Otherwise let \(T = \{j \in S : u_j > \delta\} \ne \emptyset\). Then \(b + \delta - \sum_{j \in S} u_j \le b + \delta - \sum_{j \in T} (u_j - \delta) - \delta |T| \le b - \sum_{j \in S} (u_j - \delta)^+\). ∎

An instance with numbers: \(b = 10\) and two arcs with capacities \(u = (6, 5)\), so that \(C = \{1, 2\}\) has \(6 + 5 = 11 = b + \delta\) with \(\delta = 1\), and the inequality reads \(x_1 + x_2 + 5(1 - y_1) + 4(1 - y_2) \le 10\). The LP point \(x = (6, 4)\), \(y = (1, 0.8)\), which sits on \(x_2 = 5 y_2\), violates it by \(0.8\), since \(6 + 4 + 0 + 0.8 = 10.8\). Closing switch 2 forces \(x_2 = 0\), and the inequality then asks \(x_1 \le 6\), which the capacity gives anyway; closing switch 1 forces \(x_1 = 0\) and it asks \(x_2 \le 1 + 4 y_2\), which is \(x_2 \le 5\) when switch 2 is open and \(x_2 = 0\) when it is closed; with both open it is the row itself. Closing a switch whose capacity exceeds \(\delta\) removes more capacity than the \(\delta\) the cover had in excess of \(b\), leaving the cover \(u_j - \delta\) short of \(b\), and the inequality charges for that shortfall. The rows \(x_j \le u_j y_j\) are the big-\(M\) form of an on–off switch: a continuous quantity bounded by a constant times a binary, which Section 4.2 treats in general. Flow covers are the family that matters most on models with that structure: fixed charges, lot sizing, network design, and the indicator constraints a portfolio problem produces when a position may be held or not. They are the linear ancestor of the perspective cuts of Section 4.3.

The cut loop and the selection of cuts

A solver does not add one cut. At the root it runs rounds of separation. Each round produces hundreds or thousands of candidates from a dozen separators, chooses a few of them, re-solves the LP, and repeats until the bound stops moving. The choice is as important as the generation. The following is the loop as SCIP runs it, with its default parameters, which are representative of the other solvers.Parameter values from SCIP 10.0.0, src/scip/set.c, cutsel_hybrid.c and sepa_gomory.c (separating/maxcutsroot 2000, maxcuts 100, maxstallroundsroot 10, maxstallrounds 1, cutagelimit 80, minortho 0.9, minactivityquot 0.8, the hybrid selector weights 1.0, 0.1, 0.1, 0.0, separating/gomory/maxroundsroot 10). The efficacy-plus-orthogonality selection is T. Achterberg, Constraint Integer Programming, PhD thesis, TU Berlin (2007), Section 8.9.

Algorithm 3.3.10  CUT-LOOP
                  (at one node; the root is the case that matters)

Input   the node LP; the incumbent value z_inc; separators
        S_1, ..., S_k (Gomory, c-MIR, flow cover, knapsack cover,
        zero-half, lift-and-project, clique, implied bound, ...);
        a cut pool.
Output  the tightened LP and its bound.

1. Solve the LP (dual simplex, warm). If it is infeasible or its
   bound is >= z_inc: prune the node.

2. round := 0; stall := 0.

3. repeat

      a. Collect candidates: violated cuts from the pool; then from
         each separator in priority order, up to its per-round
         limit (Gomory: 200 at the root, 50 in the tree).

      b. Score each candidate c:
            score(c) = w_e efficacy(c) + w_o objparallelism(c)
                       + w_i intsupport(c) + w_d dircutoffdist(c)
         (SCIP's hybrid selector: w_e = 1.0, w_o = 0.1, w_i = 0.1,
         w_d = 0.0).

      c. Select greedily: take the best-scored cut, discard every
         remaining candidate whose parallelism with it exceeds
         1 - minortho (minortho = 0.9), repeat until maxcuts have
         been chosen (2000 per round at the root, 100 in the tree)
         or none remain.

      d. Add the selected cuts to the LP as rows and put the rest in
         the pool. Re-solve by the dual simplex: the old basis stays
         dual feasible, because a new row only adds a slack column
         at zero (Proposition 3.1.12 and the remark after it).

      e. If the bound improved by less than a tolerance:
         stall += 1, else stall := 0.

      f. Age: remove rows whose slack has been basic (non-binding)
         for some rounds; delete pool cuts older than
         cutagelimit = 80.

      g. round += 1

   until stall > maxstallrounds (root: 10, tree: 1), or no separator
   produced a cut, or the vertex is integral.

4. At a restart, turn root cuts that were active in at least 80% of
   the LPs into constraints (SCIP: separating/minactivityquot 0.8).

Invariant
    every row of the LP is valid for the mixed-integer feasible set,
    so every bound in the loop is a valid bound and the integer hull
    is never cut.
The cut loop of Algorithm 3.3.10 at one node

  1. solve the node LP --- infeasible, or bound >= z_inc ---> prune
     |
  2. round := 0, stall := 0
     |
     v
  3. a. candidates: the pool, then the separators <-----------+
     |                                                        |
     b. score = 1.0 efficacy + 0.1 objparallelism             |
     |          + 0.1 intsupport + 0.0 dircutoffdist          |
     |                                                        |
     c. select greedily, best score first; drop a candidate   |
     |  whose parallelism to a chosen cut exceeds 1 - 0.9;    |
     |  stop at maxcuts (root 2000, tree 100)                 |
     |                                                        |
     d. add the chosen cuts as rows, the rest to the pool;    |
     |  re-solve by the dual simplex from the old basis       |
     |                                                        |
     e. f. g. stall count, aging, round += 1                  |
     |                                                        |
     stall > maxstallrounds (root 10, tree 1), no cut, or     |
     an integral vertex ? ------------------------- no -------+
     | yes
     v
  the tightened LP and its bound

The cost of one round is dominated by separation and by the re-solve. Separation means the Gomory solves and the c-MIR aggregation search, each a few rows and tens of divisors. The re-solve is tens to hundreds of dual pivots at the root. Selection is \(O(k^2)\) pairwise parallelism checks for \(k\) candidates, or \(O(k \cdot \text{selected})\) with the greedy filter. Achterberg's thesis introduced the efficacy-plus-orthogonality rule and documented that selection matters as much as generation. Wesselmann and Suhl compared a dozen measures and management policies. Dey and Molinaro collected what is known in theory about selecting cuts and listed what is not. Turner, Koch, Serrano and Winkler showed that the best selector weights are instance dependent and learned them from the instance, which is the one place in the cut loop where a learned component has an unambiguous role.F. Wesselmann and U. H. Suhl, "Implementing cutting plane management and selection techniques", technical report, Universität Paderborn, Optimization Online 2012/12/3714 (2012); S. S. Dey and M. Molinaro, "Theoretical challenges towards cutting-plane selection", Mathematical Programming 170 (2018); M. Turner, T. Koch, F. Serrano and M. Winkler, "Adaptive cut selection in mixed-integer linear programming", Open Journal of Mathematical Optimization 4 (2023); arXiv 2202.10962 (2022). Gurobi exposes a global aggressiveness level (Cuts, from −1 to 3) and one such level per cut family, but no selection weights: Gurobi Optimizer Reference Manual 13.0, "Parameters".

What does not parallelize in the loop is its outer structure: the cuts of round \(r + 1\) are computed from the vertex that round \(r\) produced, and the re-solve is a sequence of pivots. What does parallelize is everything inside a round. The separators are independent of one another and of each other's rows. Knapsack-cover, MIR and zero-half separation are procedures over single rows or small aggregations at a fixed LP point, so they run across rows at once. The \(k\) Gomory rows are \(k\) triangular solves with one factorization, and the scoring and the pairwise parallelism matrix are dense linear algebra over \(k\) sparse cut vectors. A GPU cut loop with a principled and deterministic selection is an open design problem, taken up in Section 7.

A pure cut loop on the running example

The script below runs the GMI loop of Theorem 3.3.4 on the running example in exact rational arithmetic, with the LP solved by enumerating vertices and the cut taken from the row of the most fractional basic variable. It prints the first cut, the bound after each of eight rounds, the share of the root gap closed, and the largest coefficient in each cut.

# Gomory mixed-integer cuts on the running example.
#
# Maximize x + y (the 45-degree objective) over 2x + 5y <= 24.5,
# 5x + 2y <= 30.5, -3x + 4y <= 11, x - 2y <= 4.2, x, y >= 0, with
# (x, y) integer. Exact rational arithmetic; each LP is solved by
# enumerating the vertices of the current polygon.

from fractions import Fraction as F
from itertools import combinations
from math import gcd, sqrt

# the rows a . (x, y) <= b of the polygon, then x >= 0 and y >= 0
rows = [([F(2), F(5)], F(49, 2)),
        ([F(5), F(2)], F(61, 2)),
        ([F(-3), F(4)], F(11)),
        ([F(1), F(-2)], F(21, 5)),
        ([F(-1), F(0)], F(0)),
        ([F(0), F(-1)], F(0))]
# 22 integer points, optimum 7
pts = [(p, q) for p in range(8) for q in range(7)
       if all(a[0]*p + a[1]*q <= b for a, b in rows)]
zint = max(p + q for p, q in pts)

def frac(q):
    """The fractional part of a Fraction q."""
    return q - (q.numerator // q.denominator)

def integral(a, b):
    """Scale a row to coprime integers.

    Its slack is then an integer at integer points.
    """
    L = 1
    for d in (a[0].denominator, a[1].denominator, b.denominator):
        L = L * d // gcd(L, d)
    v = [int(a[0]*L), int(a[1]*L), int(b*L)]
    g = gcd(gcd(abs(v[0]), abs(v[1])), abs(v[2]))
    return [F(v[0] // g), F(v[1] // g)], F(v[2] // g)

def lp_max(R):
    """Maximize x + y: the best feasible intersection of two rows."""
    best = None
    for (i, (ai, bi)), (j, (aj, bj)) in combinations(enumerate(R), 2):
        det = ai[0]*aj[1] - ai[1]*aj[0]
        if det == 0:
            continue
        x = (bi*aj[1] - ai[1]*bj) / det
        y = (ai[0]*bj - bi*aj[0]) / det
        if (all(a[0]*x + a[1]*y <= b for a, b in R)
                and (best is None or x + y > best[0])):
            best = (x + y, (x, y), (i, j))
    return best

def gmi(R, i, j, which):
    """The GMI cut from the tableau row of basic variable `which`.

    The vertex is the intersection of rows i and j.
    """
    (ai, bi), (aj, bj) = integral(*R[i]), integral(*R[j])
    det = ai[0]*aj[1] - ai[1]*aj[0]
    # x_B = Binv b - Binv s
    Binv = [[aj[1]/det, -ai[1]/det],
            [-aj[0]/det, ai[0]/det]]
    beta = [Binv[0][0]*bi + Binv[0][1]*bj,
            Binv[1][0]*bi + Binv[1][1]*bj]
    f0, g = frac(beta[which]), []
    # a = coefficient of the integer slack in the row; Theorem 3.3.4
    for a in Binv[which]:
        fj = frac(a)
        g.append(fj/f0 if fj <= f0 else (1 - fj)/(1 - f0))
    # g . s >= 1 with s = b - A x
    a_cut = [g[0]*ai[0] + g[1]*aj[0], g[0]*ai[1] + g[1]*aj[1]]
    return integral(a_cut, g[0]*bi + g[1]*bj - 1), f0, g

# the figure uses c = (cos 45, sin 45): divide x + y by sqrt 2
R, s2 = list(rows), sqrt(2)
z0, (x0, y0), _ = lp_max(R)
print(f"root: vertex ({x0}, {y0}), bound {float(z0):.4f} = "
      f"{float(z0)/s2:.3f} in the figure's units")
print(f"integer optimum {zint} = {zint/s2:.3f}")

for rnd in range(1, 9):
    z, (x, y), (i, j) = lp_max(R)
    # the most fractional row, ties to x
    dx, dy = abs(frac(x) - F(1, 2)), abs(frac(y) - F(1, 2))
    which = 0 if dx <= dy else 1
    (a, b), f0, g = gmi(R, i, j, which)
    valid = all(a[0]*p + a[1]*q <= b for p, q in pts)
    assert valid, "an integer point was cut off"

    R.append((a, b))
    znew, (xn, yn), _ = lp_max(R)

    if rnd == 1:
        print(f"row of {'xy'[which]}: f0 = {f0}, coefficients {g[0]} "
              f"and {g[1]}; cut {a[0]} x + {a[1]} y <= {b}")
        print()
        print(f"{'after':>5}   {'bound':>6}   {'in the':>6}   {'gap':>6}"
              f"  {'vertex':<21}  largest")
        print(f"{'cut':>5}   {'':>6}   {'figure':>6}   {'closed':>6}"
              f"  {'':<21}  coefficient")

    gap = 100 * float((z0 - znew) / (z0 - zint))
    vertex = f"({xn}, {yn})"
    coef = float(max(abs(a[0]), abs(a[1])))
    print(f"{rnd:>5}   {float(znew):>6.4f}   {float(znew)/s2:>6.3f}"
          f"   {gap:>5.1f}%  {vertex:<21}  {coef:.1e}")
root: vertex (69/14, 41/14), bound 7.8571 = 5.556 in the figure's units
integer optimum 7 = 4.950
row of x: f0 = 13/14, coefficients 2/3 and 5/39; cut 11 x + 20 y <= 110

after    bound   in the      gap  vertex                 largest
  cut            figure   closed                         coefficient
    1   7.7500    5.480    12.5%  (5, 11/4)              2.0e+01
    2   7.5455    5.335    36.4%  (50/11, 3)             1.1e+02
    3   7.3976    5.231    53.6%  (5, 199/83)            8.3e+02
    4   7.3324    5.185    61.2%  (3245/749, 3)          7.5e+03
    5   7.3101    5.169    63.8%  (5, 16741/7247)        7.2e+04
    6   7.3030    5.164    64.6%  (308705/71741, 3)      7.2e+05
    7   7.3009    5.163    64.9%  (5, 1645669/715223)    7.2e+06
    8   7.3003    5.162    65.0%  (30728345/7145669, 3)  7.1e+07
The pure GMI loop on R1 (maximize x + y): the bound after each cut falls quickly for three rounds and then crawls toward 7.3, without reaching the integer optimum 7. The slider replays the table cut by cut.

Each round of the script costs one vertex enumeration, \(O(m^2)\) pairwise intersections of the \(m\) current rows, and one exact tableau row from a \(2 \times 2\) inverse. The rounds are sequential, since each cut is read at the vertex the previous one produced, but the tableau rows of one round, two here and \(k\) in a solver, are independent solves with one factorization, as noted after Algorithm 3.3.5.

(Why the loop stalls at 7.3) Three things in this table are general. The bound falls quickly for three rounds and then crawls toward \(7.3\) without reaching the integer optimum \(7\). The value \(7.3\) is not an accident. The two pieces \(\{x \le 4\}\) and \(\{y \le 2\}\) of the relaxation both have LP value \(7.3\), and the piece \(\{x \ge 5,\ y \ge 3\}\) is empty because \(5 \cdot 5 + 2 \cdot 3 = 31 > 30.5\). So \(7.3\) is the bound that the two single-variable disjunctions, on \(x\) at \(4\) and on \(y\) at \(2\), give together. The loop's cuts, each read from a tableau row at the current vertex, approach that bound and never pass it. What they lack is the right disjunction, not rank. The Chvátal–Gomory cut with \(\lambda = (1/7, 1/7)\) on the two tight rows has \(\lambda^\top A = (1, 1)\) exactly and \(\lambda^\top b = 55/7\), so it is \(x + y \le 7\), which is the integer optimum in one step. It is the split cut of the disjunction \(x + y \le 7\) or \(x + y \ge 8\), whose second piece is empty because the LP maximum is \(7.857\). Branching does the same in a few steps. The children \(x \le 4\) and \(x \ge 5\) have bounds \(7.3\) and \(7.75\), as computed in Section 3.1, and the branch \(y \le 3\) under the first reaches the integer point \((4, 3)\), worth \(7\). Every remaining node has a bound below \(8\), so no integer point can beat \(7\).

Why the loop stalls at 7.3 (R1, maximize x + y). Its vertices stay in the strips 4 < x < 5 and 2 < y < 3 that the disjunctions on x at 4 and on y at 2 remove, and approach x + y = 7.3, the value of the pieces x ≤ 4 and y ≤ 2; the piece x ≥ 5, y ≥ 3 is empty, since 5 · 5 + 2 · 3 = 31 > 30.5, and the optimum x + y = 7 is at (4, 3) and (5, 2).

(The growth of the coefficients) The second general thing is the arithmetic. The coefficients grow by about a factor of ten per round, to eight digits by the eighth cut and to \(7 \times 10^{10}\) by the eleventh, which in double precision is already unsafe. This is the numerical reason for solvers' limits on the rank of a cut, the number of rounds of derivation behind it, and for safe GMI cuts. In floating point the tableau row is only approximately an identity. Cook, Dash, Fukasawa and Goycoolea showed how to compute a GMI cut that is provably valid, by evaluating it in directed rounding and moving the error terms to the safe side. SCIP's exact mode uses a variant of their construction.W. Cook, S. Dash, R. Fukasawa and M. Goycoolea, "Numerically safe Gomory mixed-integer cuts", INFORMS Journal on Computing 21 (2009); the exact mode is described in C. Hojny, M. Besançon, K. Bestuzheva, S. Borst, A. Chmiela, J. Dionísio, J. Ehls, L. Eifler, M. Ghannam, A. Gleixner, A. Göß, A. Hoen, J. von Holly-Ponientzietz, R. van der Hulst, D. Kamp, T. Koch, K. Kofler, J. Lentz, M. Lübbecke, S. J. Maher, P. M. Meinhold, G. Mexi, T. Mohr, E. Mühmer, K. K. Patel, M. E. Pfetsch, S. Pokutta, C. Reinartz Groba, F. Serrano, Y. Shinano, M. Turner, S. Vigerske, M. Walter, D. Weninger and L. Xu, "The SCIP Optimization Suite 10.0", arXiv 2511.18580 (2025), Section 3.1, with the safe GMI cuts in Section 3.1.5.

The largest coefficient of each cut of the GMI loop on R1, on a log scale: it grows by about a factor of ten per round, to eight digits by the eighth cut. The value for cut 11 is from the same computation continued, as the sidenote says; cuts 9 and 10 are not reported.

(When a solver stops) The third general thing is when a solver stops. The efficacies fall below SCIP's threshold of \(10^{-4}\) at round 10, and its Gomory separator stops after ten root rounds in any case, so a solver would end this loop at about that point and branch. Gomory's own proof that a pure cutting-plane algorithm terminates needs the lexicographic dual simplex rule and the cut from the first fractional row, which this loop does not use. No solver runs a pure cutting-plane algorithm.Gomory (1958), cited above; the finiteness proof in the modern form is in M. Conforti, G. Cornuéjols and G. Zambelli, Integer Programming, Graduate Texts in Mathematics 271 (Springer, 2014). The loop here runs only eight rounds; the eleventh-round coefficient and the round-10 efficacy are from the same computation continued, with the same rule.

The figure below replays this at the root of the running example, at any objective angle, and lets the reader compare two families. With "hull facets" each cut is the facet of the integer hull that the current vertex violates most, which a solver would have to discover rather than look up. With "Gomory mixed-integer cuts" each cut is Algorithm 3.3.5 applied to a fractional row of the current tableau, which is what a solver computes. At the default angle of \(45^\circ\) the root bound is \(5.556\) at \((4.93, 2.93)\), where the rows \(2x + 5y \le 24.5\) and \(5x + 2y \le 30.5\) are tight. The first Gomory cut is \(3.949x + 7.179y \le 39.487\), from the row of \(x\) with \(f_0 = 0.929\), which is the cut (3.3.2) scaled so that the vertex violates it by one. It moves the vertex to \(5.480\) at \((5.00, 2.75)\) and closes \(13\%\) of the gap, where one facet closes \(100\%\). After eight Gomory cuts the bound is \(5.162\) and \(65\%\) of the gap is closed with the vertex still fractional, where two hull facets reach the integer optimum \(4.950\) at \((5, 2)\). The angle matters. Below about \(27^\circ\) or above about \(68^\circ\) the Gomory cuts reach an integral vertex, in as few as two cuts close to the axes. Between those angles the loop stalls with between \(31\%\) and \(83\%\) of the gap closed. The readout prints the derivation of the current cut: the two tight rows, the tableau rows for \(x\) and \(y\), \(f_0\), the two coefficients, and the cut in the slacks and in \((x, y)\). It also checks that the cut is valid for all 22 integer points.

The root node, cut by cut, with two cut families to choose from. Under "hull facets" each cut is the facet of the integer hull that the current vertex violates most, which a solver would have to discover rather than look up, and under "Gomory mixed-integer cuts" each cut is read off the simplex tableau at the current vertex, which is what a solver computes. Purple is the cut and the hatched sliver it removes, orange is the relaxation after it, and the stats compare how much of the gap each family has closed after the same number of cuts.

(What a computable cut gives up) The hull facets show what every cut does: remove the current vertex and a sliver behind it, move the relaxation, and leave the integer points alone. They also show how few of the hull's facets a solver ever needs, since only the ones the objective is looking at are ever added. At \(45^\circ\) two of the seven facets finish the problem. The Gomory cuts show what a computable cut gives up. A cut the solver can derive from the tableau is valid and separates the vertex, but it is not a facet. A sequence of them approaches the hull only asymptotically in the directions the objective looks, and sometimes not at all.

What a cut loop is worth

The value of cutting planes to a MILP solver has been measured by switching them off. The ablation studies summarized in Section 3.2, and taken up in Section 5.1, rank cutting planes first among the components of the CPLEX engine. In the open SCIP 10 ablation, switching cuts off costs a factor of \(2.10\), behind presolve at \(2.60\) and the branching rule at \(2.55\).G. Mexi, The Two Faces of Mixed-Integer Programming: Primal and Dual Progress, doctoral thesis, Technische Universität Berlin (2026), doi 10.14279/depositonce-26588, read 5 October 2026; 349 instances, five seeds. The two CPLEX studies, Bixby et al. (2004) and Achterberg and Wunderling (2013), and the caveat that neither chapter was re-read for this series, are in the sidenote of Section 3.2. This is why a MILP solver closes a large part of its root gap before it branches, and why the phrase "root gap closed" is the first number a practitioner reads in a log.

For the MINLP solvers the same engine is underneath. BARON bounds its nodes with a linear program, a design decision of 2005 that the next subsection describes, and solves it with a MILP code's dual simplex. SCIP's nonlinear constraint handlers add their cuts to the same loop. FICO's Xpress Global is described by its authors as the Xpress MILP solver with a nonlinear presolve, convexification cuts and spatial branching added, with the MILP machinery carried over intact. The ablation they report on their own test set says the RLT cuts are essential.M. Tawarmalani and N. V. Sahinidis, "A polyhedral branch-and-cut approach to global optimization", Mathematical Programming 103 (2005); K. Bestuzheva, A. Chmiela, B. Müller, F. Serrano, S. Vigerske and F. Wegscheider, "Global optimization of mixed-integer nonlinear programs with SCIP 8", Journal of Global Optimization 91 (2025); P. Belotti, T. Berthold, T. Gally, L. Gottwald and I. Pólik, "Solving MINLPs to global optimality with FICO Xpress Global", Optimization Online (July 2025). The Xpress Global ablation is the vendor's own, on the developers' internal test set: it reports that its RLT cuts are "essential", with 34 fewer instances solved without them, while their effect on the time of the instances solved either way is about 1%.

Cuts for nonlinear constraints

Everything so far used integrality. Nonlinearity supplies cuts of its own, and the first family is the one that organizes the rest of this chapter: the tangent planes of a convex function.

Proposition 3.3.11 (gradient cuts; Proposition 1.2.4 restated as a cut). Let \(h\) be differentiable on a convex set \(C \subseteq \mathbb{R}^n\) and let \(\bar v \in C\). Write \(\ell_{\bar v}(v) = h(\bar v) + \nabla h(\bar v)^\top (v - \bar v)\) for the linearization of \(h\) at \(\bar v\), as in Proposition 1.2.4. (a) If \(h\) is convex then \(\ell_{\bar v} \le h\) on \(C\), so \(\{h \le 0\} \subseteq \{\ell_{\bar v} \le 0\}\): the linear inequality \(\ell_{\bar v}(v) \le 0\) is valid for the feasible set of \(h \le 0\), and it is a cut for every point with \(\ell_{\bar v} > 0\). (b) If \(h\) is concave then \(\ell_{\bar v} \ge h\) on \(C\), so \(\{\ell_{\bar v} \le 0\} \subseteq \{h \le 0\}\): the inequality \(\ell_{\bar v} \le 0\) is a restriction of the feasible set, not a relaxation, and adding it can remove feasible points.

Proof. Part (a) is Proposition 1.2.4. For concave \(h\) apply (a) to \(-h\). The set inclusions follow at once. ∎

(Tangents of convex pieces are cuts, of concave pieces are not) Part (a) is the gradient inequality of Proposition 1.2.4, read as a cut. A convex region \(\{g(x) \le 0\}\) is the intersection of all its tangent half-spaces. It can therefore be approximated from the outside by a polyhedron of tangents as closely as one likes, and a problem whose only nonlinear constraints are convex becomes a sequence of linear programs. Part (b) is the warning that governs how global solvers use the idea. The tangent of a concave constraint function lies above the function, so the half-space it defines is inside the feasible region rather than around it, and a method that adds it as a cut is solving a different problem. Section 3.4 shows the consequence on a library instance, where outer approximation terminates with a closed gap and the wrong answer.

Proposition 3.3.11 on a line. (a) A convex h lies above its tangent ℓ at v̄, so {h ≤ 0} lies inside {ℓ ≤ 0}: ℓ ≤ 0 is valid, and a cut for the points with ℓ > 0. (b) A concave h lies below its tangent, so {ℓ ≤ 0} lies inside {h ≤ 0}: ℓ ≤ 0 is a restriction, not a relaxation, and it removes the feasible points with ℓ > 0. The bars in the grey strip at the foot of each panel are the sets, with closed ends capped and unbounded ends arrowed; in (b) the red part of the h ≤ 0 bar is the feasible points that ℓ ≤ 0 removes. Schematic, no scale.

(How the global solvers use the division) Every global solver therefore adds gradient cuts only to expressions it has proved convex, and relaxes the others by the envelopes of Section 2.4. BARON made this division the organizing principle of a global solver in 2005: relax each nonconvex piece by its envelope, linearize each convex piece by tangents, and bound every node with a linear program. The linear program is cheap, warm-starts from node to node and gives a valid bound at every iteration, where an NLP solved locally does not.Tawarmalani and Sahinidis (2005), cited above; most of the global solvers of Section 5 followed. SCIP 8 does the same through its nonlinear handlers. A violated constraint that the handler has proved convex is separated by a tangent on its graph, with the gradient obtained by automatic differentiation, exact derivatives computed by the chain rule over the expression graph. A concave function is underestimated by the hyperplane that is tightest at the reference point among those valid at every vertex of the box.Bestuzheva et al. (2025), cited above. The convexity test on an expression graph is a sufficient condition, not a decision procedure: R. Fourer, C. Maheshwari, A. Neumaier, D. Orban and H. Schichl, "Convexity and concavity detection in computational graphs: tree walks for convexity assessment", INFORMS Journal on Computing 22 (2010). Gurobi's release notes, quoted in Section 4.6, describe spatial branching and a dynamic outer approximation for its general function constraints in place of the earlier static piecewise-linear approximation, which on a node's domain can only mean tangents where the piece is convex and secants where it is concave, the division above, although the notes publish no algorithm.

(Perspective, RLT and semidefinite cuts: a map) Three further families are specific to the structures of Section 4, and this is only a map to them. Perspective cuts are the gradient cuts of the perspective function \(z\, f(x/z)\) of a convex cost \(f\) switched on by an indicator \(z\). They are linear in \((x, z)\), and together they describe the convex hull of the on–off set, which is the subject of Section 4.3.A. Frangioni and C. Gentile, "Perspective cuts for a class of convex 0–1 mixed integer programs", Mathematical Programming 106 (2006). RLT cuts multiply pairs of linear constraints and replace each product of variables by a lifted variable \(X_{jk}\). The products of the bound constraints are McCormick's planes, and the products of the other rows are new. On the bilinear example \(\max xy\) under \(2x + y \le 1.2\) on the unit square, the McCormick relaxation gives \(0.40\) against the true \(0.18\). The full set of level-1 products lowers the bound to \(0.30\), and adding the single convexity fact \(X_{11} \ge x^2\) brings it to \(0.18\), the exact answer, with no branching at all. Section 4.7 derives the ladder and the figure there solves the three linear programs in the page.H. D. Sherali and W. P. Adams, "A hierarchy of relaxations between the continuous and convex hull representations for zero-one programming problems", SIAM Journal on Discrete Mathematics 3 (1990); H. D. Sherali and A. Alameddine, "A new reformulation-linearization technique for bilinear programming problems", Journal of Global Optimization 2 (1992). Semidefinite cuts read the lifted matrix \(M(x, X) = \begin{pmatrix} 1 & x^\top \\ x & X \end{pmatrix}\), which is positive semidefinite at every point with \(X = xx^\top\). If the LP point has an eigenvector \(v\) with \(v^\top M(\bar x, \bar X) v < 0\), the inequality \(v^\top M(x, X)\, v \ge 0\) is linear in \((x, X)\), valid, and violated. A solver can therefore harvest the strength of the semidefinite relaxation of Section 4.7 one linear cut at a time.H. D. Sherali and B. M. P. Fraticelli, "Enhancing RLT relaxations via a new class of semidefinite cuts", Journal of Global Optimization 22 (2002).

The newest family is the one that makes the nonconvex quadratic case look most like the integer case. An intersection cut needs a convex set whose interior contains the current vertex and no feasible point. The GMI cut is the case where that set is the strip \(\pi_0 < \pi^\top x < \pi_0 + 1\) of a split disjunction.

Definition 3.3.12 (S-free set, intersection cut). Let \(S\) be a closed set. A closed convex set \(C\) is \(S\)-free if \(\operatorname{int}(C) \cap S = \emptyset\), and maximal \(S\)-free if no \(S\)-free convex set properly contains it. Let \(\bar x \in \operatorname{int} C\) be a vertex of the LP relaxation with basis cone \(\bar x + \operatorname{cone}\{r^1, \dots, r^k\}\). Here \(r^j\) is the direction in which the basic variables move when the nonbasic \(s_j\) is raised from zero, the \(j\)-th column of \(-B^{-1} A_N\) in the notation of the tableau row before Theorem 3.3.4. The cone is therefore the set of points the tableau describes with \(s \ge 0\): every point of it is \(x = \bar x + \sum_j s_j r^j\) with \(s_j \ge 0\). Let \(\alpha_j = \sup\{\alpha : \bar x + \alpha r^j \in C\}\) be the step along each ray to the boundary of \(C\). The intersection cut is \(\sum_j s_j / \alpha_j \ge 1\), with the term dropped when \(\alpha_j = +\infty\). It is valid for \(S \cap (\text{cone})\) and it cuts off \(\bar x\).

Validity is one sentence. A point of the cone with \(\sum_j s_j/\alpha_j < 1\) is a convex combination of \(\bar x\) and points \(\bar x + \alpha_j r^j\) on the boundary of \(C\), with positive weight on \(\bar x\), so it lies in \(\operatorname{int} C\) and is not in \(S\). The depth of the cut is governed by the size of \(C\), which is why maximality matters. For a quadratic inequality the natural \(S\)-free set has a closed form.

Theorem 3.3.13 (Muñoz and Serrano 2022, Theorem 7). Let \(S_h = \{(p, q) \in \mathbb{R}^{n + m} : \lVert p \rVert \le \lVert q \rVert\}\) and, for \(\lVert \lambda \rVert = 1\), \(C_\lambda = \{(p, q) \in \mathbb{R}^{n + m} : \lambda^\top p \ge \lVert q \rVert\}\). Then \(C_\lambda\) is a maximal \(S_h\)-free set. Furthermore, if \(\lambda = \bar p / \lVert \bar p \rVert\), \(C_\lambda\) contains \((\bar p, \bar q)\) in its interior.G. Muñoz and F. Serrano, "Maximal quadratic-free sets", Mathematical Programming 192 (2022), where the two groups of coordinates are written \(x\) and \(y\). The non-homogeneous case, in which the quadratic set is intersected with a hyperplane after homogenization, is their Theorems 8 and 10. The earlier general theory of outer-product-free sets, whose maximal members are cones, is D. Bienstock, C. Chen and G. Muñoz, "Outer-product-free sets for polynomial optimization and oracle-based cuts", Mathematical Programming 183 (2020).

Proof of \(S_h\)-freeness. If \((p, q) \in \operatorname{int} C_\lambda\) then \(\lVert q \rVert < \lambda^\top p \le \lVert p \rVert\), so \((p, q) \notin S_h\). Maximality is proved in the paper through a criterion of theirs: an \(S\)-free convex set is maximal when each of its defining inequalities holds with equality at some point of \(S\) on the boundary of the set. For \(C_\lambda\) the defining inequalities are \(-\lambda^\top p + \beta^\top q \le 0\) with \(\lVert \beta \rVert = 1\), and each holds with equality at \((\lambda, \beta)\), a point of \(S_h\) on the boundary of \(C_\lambda\). ∎

(From a quadratic inequality to the theorem's form) The theorem, with its non-homogeneous companions (Theorems 8 and 10 of the paper), covers every quadratic inequality. Diagonalize the quadratic form and homogenize. Homogenization adds a coordinate \(x_0\) and replaces every affine term by its product with \(x_0\), so that the inequality becomes a homogeneous quadratic form whose section \(x_0 = 1\) is the original set. The eigen-decomposition of that form then gives the two groups of the theorem. The coordinates with positive eigenvalues form \(p\) and those with negative eigenvalues form \(q\), and the inequality reads "the norm of the first group is at most the norm of the second". The nonconvex region is the complement of the interior of the cone \(\lVert p \rVert \ge \lVert q \rVert\), the two nappes of a second-order cone when the first group is a single coordinate, and the convex set that avoids it is a rotated second-order cone meeting it only along the boundary cone \(\{(t\lambda, q) : \lVert q \rVert = t \ge 0\}\), the whole boundary of \(C_\lambda\) when \(n = 1\).

(The intersection cut on a lifted square) One worked instance shows the arithmetic. Take the LP relaxation with the two rows \(x \ge 0\) and \(w \le 5\) of a lifted square, whose vertex \((x, w) = (0, 5)\) violates the nonconvex side \(w \le x^2\). The rays of its basis cone are \(r^1 = (1, 0)\) and \(r^2 = (0, -1)\). With \(x_0 = 1\) the identity \(w x_0 = \big(\tfrac{w + x_0}{2}\big)^2 - \big(\tfrac{w - x_0}{2}\big)^2\) turns \(x^2 \ge w\) into \(|p| \le \lVert q \rVert\) with \(p = (w + 1)/2\) and \(q = (x, (w - 1)/2)\), the form of the theorem with \(n = 1\) and \(m = 2\). The vertex maps to \(p = 3\), \(q = (0, 2)\), inside \(C = \{\lVert q \rVert \le p\}\), which is \(C_\lambda\) with \(\lambda = 1\). Along \(r^1\) the boundary is reached when \(\sqrt{\alpha^2 + 4} = 3\), so \(\alpha_1 = \sqrt 5\). Along \(r^2\) it is reached when \(|2 - \alpha/2| = 3 - \alpha/2\), so \(\alpha_2 = 5\), at the point \((0, 0)\) on the parabola itself. The cut is

\[\frac{x}{\sqrt 5} + \frac{5 - w}{5} \;\ge\; 1, \qquad \text{that is} \qquad w \;\le\; \sqrt 5\, x \approx 2.236\, x .\]

It removes \((0, 5)\) and is valid for every point with \(w \le x^2\), \(w \le 5\) and \(x \ge 0\). It used no upper bound on \(x\), which is its point: the secant of \(x^2\) on \([0, 2]\), the envelope of Section 2.4, gives \(w \le 2x\) and needs the bound to exist. One caveat on the set used. The inequality \(w \le x^2\) is the section \(x_0 = 1\) of the homogenized set, for which the maximal free sets are those of the paper's Theorems 8 and 10. The homogeneous \(C\) used here is free for that section, so the cut is valid, but it need not be maximal there. The script below reproduces the two steps and checks the cut on a random sample of feasible points.

The intersection cut at the vertex (0, 5) of x ≥ 0, w ≤ 5. C = {‖q‖ ≤ p} is the side w ≥ x² of the parabola in this plane. The rays r¹ = (1, 0) and r² = (0, −1) of the vertex's basis cone leave C at (√5, 5) and (0, 0), so α₁ = √5 and α₂ = 5, and the cut x/√5 + (5 − w)/5 ≥ 1, that is w ≤ √5 x, removes (0, 5) and is valid for every point with w ≤ x², w ≤ 5 and x ≥ 0.
# An intersection cut for the nonconvex side w <= x^2.
#
# The cut comes from the maximal quadratic-free set
# C = {(p, q) : ||q|| <= p}, at the LP vertex (x, w) = (0, 5) with the
# rays r1 = (1, 0) and r2 = (0, -1).

import numpy as np

def lift(x, w):
    """Map (x, w) to (p, q), so that w <= x^2  <=>  |p| <= ||q||.

    p = (w + 1)/2 and q = (x, (w - 1)/2).
    """
    return (w + 1) / 2, np.array([x, (w - 1) / 2])

pb, qb = lift(0.0, 5.0)

def step(dx, dw):
    """The smallest alpha > 0 with ||q + alpha dq|| = p + alpha dp.

    That is where the ray (dx, dw) from the vertex leaves C.
    """
    dp, dq = dw / 2, np.array([dx, dw / 2])
    # ||qb + alpha dq||^2 = (pb + alpha dp)^2 as a alpha^2 + b alpha + c = 0
    a = dq @ dq - dp**2
    b = 2 * (qb @ dq - pb * dp)
    c = qb @ qb - pb**2
    if abs(a) > 1e-12:
        roots = np.roots([a, b, c])
    else:
        roots = np.array([-c / b])
    return min(r.real for r in roots
               if abs(r.imag) < 1e-9 and r.real > 1e-12
               and pb + r.real * dp >= 0)

a1, a2 = step(1.0, 0.0), step(0.0, -1.0)
print(f"vertex (0, 5) is interior to C: {np.linalg.norm(qb) < pb}")
print(f"steps alpha1 = {a1:.6f} (sqrt 5), alpha2 = {a2:.6f}")
print(f"cut: x/{a1:.4f} + (5 - w)/{a2:.4f} >= 1, i.e. w <= {a2/a1:.4f} x")
print(f"the vertex violates it by {5 - a2/a1*0:.0f}")

# validity on a random sample of feasible points
rng = np.random.default_rng(0)
X, W = rng.uniform(0, 4, 200000), rng.uniform(-5, 5, 200000)
feas = (W <= X**2) & (W <= 5)
bad = int((W[feas] > a2/a1 * X[feas] + 1e-9).sum())
print(f"validity: {bad} of {feas.sum()} sampled feasible points "
      f"violate the cut")
vertex (0, 5) is interior to C: True
steps alpha1 = 2.236068 (sqrt 5), alpha2 = 5.000000
cut: x/2.2361 + (5 - w)/5.0000 >= 1, i.e. w <= 2.2361 x
the vertex violates it by 5
validity: 0 of 162733 sampled feasible points violate the cut

The cost of one such cut is an eigendecomposition of the constraint's quadratic form, \(O(n^3)\) once per constraint and reusable across nodes, and then one scalar quadratic equation per ray. Chmiela, Muñoz and Serrano implemented these cuts in SCIP, strengthened them with the bound information and the hyperplane-adjusted sets of the general case, and measured their effect on quadratically constrained problems. The cuts shipped in SCIP 8's quadratic nonlinear handler.A. Chmiela, G. Muñoz and F. Serrano, "On the implementation and strengthening of intersection cuts for QCQPs", Mathematical Programming 197 (2023); Bestuzheva et al. (2025), cited above.

Where this is used

Every MILP solver runs Algorithm 3.3.10 at the root with the GMI, MIR, cover, flow-cover and zero-half families, and some add lift-and-project. In the tree it runs a reduced version, with far fewer rounds and cuts per node. Every global MINLP solver of Section 5 runs the same loop on the linear program that bounds its nodes. To the separators it adds the gradient cuts of its convex pieces, the RLT products of its linear rows and, in SCIP, the intersection cuts of Theorem 3.3.13.

What parallelizes

The loop is sequential in its rounds and in its re-solves, which are dual simplex pivots. Inside a round the work is wide. Separation runs row by row or aggregation by aggregation at one fixed LP point, and the Gomory rows are independent triangular solves with one factorization. The intersection cuts are independent scalar equations per ray after one cached eigendecomposition, and selection is a dense pairwise computation over the candidates. A frontier of nodes multiplies this by the number of nodes, since every node's LP point is a separate separation problem over the same rows. What a device cannot do is replace the re-solve, and what has not been designed is a selection rule for a batch of thousands of candidates that is both principled and deterministic. Both questions return in Section 7.

One round of the cut loop: wide inside, sequential across rounds

  the LP point of round r, fixed for the whole round
    |
    |  separation, all at once:
    +--> Gomory: k triangular solves with one factorization
    +--> knapsack cover, c-MIR, zero-half: single rows or small
    |    aggregations, across rows at once
    +--> intersection cuts: one scalar equation per ray, after
    |    one cached eigendecomposition
    v
  candidates --> selection: a dense pairwise computation
    |
    v
  re-solve: dual simplex pivots, one after another
    |
    v
  the LP point of round r + 1

  a frontier of nodes multiplies the wide part by the number of
  nodes: every node's LP point is a separate separation problem
  over the same rows

The convex case

The integers are one nonconvexity and the curves are another. This section removes the second and keeps the first. It treats the convex MINLP, in which every nonlinear function is convex and the integrality of \(y\) is the only thing that stands between the problem and a convex program. The reason to do this before the general case is that convexity buys exactly two facts, and the exact methods of this section are built from them. The first is that a local solve of the continuous problem with the integers fixed returns a global optimum of that problem, so it supplies a feasible point and an upper bound that can be trusted. The second is Proposition 3.3.11(a): every tangent of a convex constraint is a valid inequality, so a polyhedron of tangents is a relaxation and a master problem built from tangents gives a lower bound that can be trusted.

Branch and bound on the convex relaxation uses the first fact, outer approximation and its relatives use both, and the single-tree methods combine them inside one branch and bound. The section ends where the next one begins: on a nonconvex constraint the second fact fails, and the running example \(\mathsf{st\_e13}\) shows outer approximation terminating with a closed gap and a wrong answer. The figure of this section maximizes, as the figures do, and the displays minimize.

Definition 3.4.1 (convex MINLP, continuous relaxation). Write the problem as

\[\text{(P)} \qquad z^\star = \min\{\, f(x, y) : g_i(x, y) \le 0,\ i = 1, \dots, m,\ x \in X,\ y \in Y \,\},\]

with \(X \subset \mathbb{R}^n\) a compact polyhedron and \(Y \subset \mathbb{Z}^p\) finite, a box intersected with the integer lattice. (P) is a convex MINLP if \(f\) and every \(g_i\) are convex and continuously differentiable on an open set containing \(X \times \operatorname{conv}(Y)\). Equality constraints are allowed only if affine. The continuous relaxation \(\mathrm{(P_{rel})}\) replaces \(y \in Y\) by \(y \in \operatorname{conv}(Y)\). Its value is \(z_R \le z^\star\) and its feasible set \(\mathcal F_{\mathrm{rel}} = \{(x, y) : g(x, y) \le 0,\ x \in X,\ y \in \operatorname{conv}(Y)\}\) is convex.

Definition 3.4.2 (NLP subproblem, feasibility subproblem). For a fixed assignment \(\bar y \in Y\) the NLP subproblem is

\[\mathrm{NLP}(\bar y): \qquad \min_{x \in X}\ f(x, \bar y) \quad \text{subject to} \quad g(x, \bar y) \le 0,\]

a convex program whose minimizer \(\bar x\) is global and whose value \(f(\bar x, \bar y)\) is an upper bound on \(z^\star\). If \(\mathrm{NLP}(\bar y)\) is infeasible, the feasibility subproblem is

\[\mathrm F(\bar y): \qquad \min_{x \in X}\ \sum_{i=1}^m w_i \max\{0,\ g_i(x, \bar y)\}, \qquad w_i > 0,\]

a convex nonsmooth program that returns the point of least weighted violation. A KKT point of \(\mathrm{NLP}(\bar y)\) is as in Theorem 1.3.4, with the linear constraints of the polyhedron \(X\) carried by its normal cone. The normal cone of \(X\) at \(\bar x \in X\) is \(N_X(\bar x) = \{d : d^\top (x - \bar x) \le 0 \text{ for all } x \in X\}\), the set of directions that point out of \(X\) at \(\bar x\). When \(X\) is a box it consists of the vectors that vanish off the active bounds, are nonnegative at an active upper bound and nonpositive at an active lower bound, so that the condition below is the KKT system of Theorem 1.3.4 with the bounds written as the constraints \(x_j - u_j \le 0\) and \(l_j - x_j \le 0\). A KKT point is a feasible \(\bar x\) for which there is \(\lambda \ge 0\) with \(\nabla_x f(\bar x, \bar y) + \sum_i \lambda_i \nabla_x g_i(\bar x, \bar y) \in -N_X(\bar x)\), \(\lambda_i g_i(\bar x, \bar y) = 0\) and \(g(\bar x, \bar y) \le 0\). For a convex program every KKT point is a global minimizer, and the converse needs a constraint qualification such as Slater's, as Theorem 1.3.4 states.

Nonlinear branch and bound

The first method is the branch and bound of Section 3.1 with the LP replaced by the convex NLP.

Theorem 3.4.3 (NLP-based branch and bound; Dakin 1965, Gupta and Ravindran 1985). Let (P) be a convex MINLP with \(Y\) finite. Run branch and bound with the relaxation at node \(N\) the convex NLP obtained from \(\mathrm{(P_{rel})}\) by restricting \(y\) to the node's box \(B_N\), and with a node of fractional \(\bar y_j\) split into \(y_j \le \lfloor \bar y_j \rfloor\) and \(y_j \ge \lceil \bar y_j \rceil\). It terminates after finitely many nodes with an optimal solution of (P) or a proof that (P) is infeasible.R. J. Dakin, "A tree-search algorithm for mixed integer programming problems", The Computer Journal 8 (1965); O. K. Gupta and A. Ravindran, "Branch and bound experiments in convex nonlinear integer programming", Management Science 31 (1985).

Proof sketch. Each node relaxation is a convex program, so the local minimizer an NLP solver returns is global and its value \(\bar z(N)\) is a valid lower bound for the node. A node whose relaxed solution has integral \(y\) gives a feasible point and an upper bound. Pruning by \(\bar z(N) \ge z_{\mathrm{inc}}\) and by infeasibility discards nothing better than the incumbent, which is Theorem 3.1.5. Each branching shrinks an integer range by at least one, so the depth is at most \(\sum_j (u_j - l_j)\) and the tree is finite, as in Theorem 3.1.6. ∎

(What the tree needs from its NLP solver) The content of the theorem is that the only thing the tree needs from its NLP solver is a global solution of each node relaxation, and convexity supplies it. Dakin's branching rule is the split into the two integer neighbours of the fractional value. Gupta and Ravindran ran the method on convex nonlinear integer programs and compared node selection and branching rules, which made it a practical method rather than a remark. Two refinements avoid solving every node NLP to convergence. Borchers and Mitchell stop a node's NLP solve early when a valid lower bound available during the solve already exceeds the incumbent, so the node is pruned without the remaining iterations.B. Borchers and J. E. Mitchell, "An improved branch and bound algorithm for mixed integer nonlinear programs", Computers & Operations Research 21 (1994). Leyffer integrated the SQP method with the tree so that branching is allowed after each iteration of the NLP solver, using the quadratic subproblem's information to choose the branch. He reported a factor of up to three over the plain method.S. Leyffer, "Integrating SQP and branch-and-bound for mixed integer nonlinear programming", Computational Optimization and Applications 18 (2001). The modern descendants are Bonmin's B-BB, SBB, the branch-and-bound modes of Knitro and Minotaur, and Juniper.P. Bonami, L. T. Biegler, A. R. Conn, G. Cornuéjols, I. E. Grossmann, C. D. Laird, J. Lee, A. Lodi, F. Margot, N. Sawaya and A. Wächter, "An algorithmic framework for convex mixed integer nonlinear programs", Discrete Optimization 5 (2008); the solvers are compared in J. Kronqvist, D. E. Bernal, A. Lundell and I. E. Grossmann, "A review and comparison of solvers for convex MINLP", Optimization and Engineering 20 (2019). The cost is one NLP per node, which is tens to hundreds of times an LP re-solve, and the method has no warm start comparable to the dual simplex basis. It pays when the continuous relaxation is tight and the nonlinearity is heavy, where a polyhedron of tangents would need many rounds to describe the curve. The 2019 comparison found the NLP-based solvers outperforming the linearization-based ones on problems with a high degree of nonlinearity and losing on problems with a large relaxation gap.Kronqvist, Bernal, Lundell and Grossmann (2019), cited above.

Outer approximation

The second method keeps the integers in a MILP and the curves in an NLP, and alternates between them.

Definition 3.4.4 (polyhedral outer approximation, master problem). With the linearization \(\ell_{\bar v}\) of Proposition 1.2.4 (Proposition 3.3.11), the affine inequality \(\ell_{\bar v}(v) = h(\bar v) + \nabla h(\bar v)^\top (v - \bar v) \le 0\) taken at a point \(\bar v = (\bar x, \bar y)\) is valid for \(\{h \le 0\}\) whenever \(h\) is convex and differentiable. Given a finite set of points \(\{v^k\}_{k \in K}\), the polyhedral outer approximation of (P) is

\[\mathcal P_K = \Big\{ (x, y, \eta) : \eta \ge f(v^k) + \nabla f(v^k)^\top (v - v^k),\ \ g_i(v^k) + \nabla g_i(v^k)^\top (v - v^k) \le 0,\ \ i = 1, \dots, m,\ k \in K \Big\},\]

which contains \(\operatorname{epi} f \cap (\mathcal F_{\mathrm{rel}} \times \mathbb{R})\). The master problem over \(K\) is the MILP \(\mathrm{(M_K)}: z_{M_K} = \min\{\eta : (x, y, \eta) \in \mathcal P_K,\ x \in X,\ y \in Y\}\). Since \(\mathcal P_K\) contains the epigraph of (P), \(z_{M_K} \le z^\star\), and \(z_{M_K}\) is nondecreasing in \(K\).

Written out, with the linearization points \((\bar x^k, \bar y^k)\) collected so far, the incumbent value \(z_{\mathrm{inc}}\) and a tolerance \(\varepsilon \ge 0\), the master is

\[\begin{aligned} \min_{x \in X,\ y \in Y,\ \eta}\ & \eta \\ \text{subject to}\ & \eta \;\ge\; f(\bar x^k, \bar y^k) + \nabla f(\bar x^k, \bar y^k)^{\top}\!\begin{pmatrix} x - \bar x^k \\ y - \bar y^k \end{pmatrix}, & k \in K_{\mathrm{f}}, \\ & g(\bar x^k, \bar y^k) + \nabla g(\bar x^k, \bar y^k)^{\top}\!\begin{pmatrix} x - \bar x^k \\ y - \bar y^k \end{pmatrix} \;\le\; 0, & k \in K, \\ & \eta \;\le\; z_{\mathrm{inc}} - \varepsilon . \end{aligned}\]

(The feasibility cuts and the ε-cut) Two details of this display are easy to drop and both matter. When \(\mathrm{NLP}(\bar y^k)\) is infeasible there is no objective cut from that round (\(k \notin K_{\mathrm f}\)). The linearization point is the solution of \(\mathrm F(\bar y^k)\), and the constraint linearizations taken there make the column \(\bar y^k\) infeasible for every later master, as the proof below shows. And the last row, the \(\varepsilon\)-cut \(\eta \le z_{\mathrm{inc}} - \varepsilon\), asks the master for a solution strictly better than the incumbent, so that an infeasible master is the termination certificate: no integer assignment can improve on the incumbent by \(\varepsilon\). The equivalent test without the row is to stop when \(z_{M_K} \ge z_{\mathrm{inc}} - \varepsilon\).

Theorem 3.4.5 (finite termination of outer approximation; Duran and Grossmann 1986, Fletcher and Leyffer 1994). Assume (A1) (P) is a convex MINLP; (A2) \(X\) is a compact polyhedron and \(Y\) is finite; (A3) at every \(\bar y \in Y\) for which \(\mathrm{NLP}(\bar y)\) is feasible a constraint qualification holds, so that its minimizer \(\bar x\) is a KKT point, and at every \(\bar y\) for which it is infeasible the minimizer \(\bar x\) of \(\mathrm F(\bar y)\) satisfies the first-order condition of that problem: there are \(\mu_i \ge 0\), with \(\mu_i = 0\) unless \(g_i(\bar x, \bar y) \ge 0\), such that \(\sum_i \mu_i \nabla_x g_i(\bar x, \bar y) \in -N_X(\bar x)\) and \(\sum_i \mu_i g_i(\bar x, \bar y) > 0\). Let the outer approximation algorithm (Algorithm 3.4.6) generate master solutions \((x^k, y^k, \eta^k)\) and linearization points \((\bar x^k, y^k)\), with \(\bar x^k\) the solution of \(\mathrm{NLP}(y^k)\) or of \(\mathrm F(y^k)\). Then no integer assignment is the master's solution at two different iterations before the stopping test fires. The algorithm therefore terminates after at most \(|Y|\) master solves, and at termination \(z_{\mathrm{inc}} = z^\star\), or (P) is infeasible.M. A. Duran and I. E. Grossmann, "An outer-approximation algorithm for a class of mixed-integer nonlinear programs", Mathematical Programming 36 (1986), for problems in which the integer variables enter linearly; R. Fletcher and S. Leyffer, "Solving mixed integer nonlinear programs by outer approximation", Mathematical Programming 66 (1994), for the general convex case, the infeasible subproblem, the \(\varepsilon\)-cut and a shorter proof. The condition on \(\mathrm F(\bar y)\) is its KKT condition: the subgradient of \(w_i \max\{0, g_i\}\) at a point with \(g_i > 0\) is \(w_i \nabla g_i\), at a point with \(g_i = 0\) it is any \(t\, w_i \nabla g_i\) with \(t \in [0, 1]\), and at a point with \(g_i < 0\) it is zero, so \(\mu_i = w_i\) on the violated constraints, \(\mu_i \in [0, w_i]\) on the active ones and \(\mu_i = 0\) elsewhere; the sum \(\sum_i \mu_i g_i\) is then positive because some constraint is violated.

Proof sketch. Lower bound: by Proposition 3.3.11(a) every linearization is valid for \(\operatorname{epi} f \cap (\mathcal F_{\mathrm{rel}} \times \mathbb{R})\), so the master is a relaxation of (P) and \(z_{M_K} \le z^\star \le z_{\mathrm{inc}}\).

No repetition, feasible case. Suppose the master returns \(\bar y\) with \(\bar x\) the KKT solution of \(\mathrm{NLP}(\bar y)\) and multipliers \(\lambda\). Consider the master restricted to \(y = \bar y\) after the cuts at \(\bar v = (\bar x, \bar y)\) have been added: minimize \(\eta\) subject to \(\eta \ge f(\bar v) + \nabla_x f(\bar v)^\top (x - \bar x)\), \(g_i(\bar v) + \nabla_x g_i(\bar v)^\top (x - \bar x) \le 0\) and \(x \in X\). This is a linear program. At \((\bar x, \eta = f(\bar v))\) its objective and constraints have the same values and the same gradients as those of \(\mathrm{NLP}(\bar y)\) at \(\bar x\). Its stationarity, complementarity and feasibility equations are therefore the same equations as the NLP's, and the same \(\lambda\) satisfies them. For a linear program the KKT conditions are sufficient, so the restricted master's value is exactly \(f(\bar x, \bar y) \ge z_{\mathrm{inc}}\). If the master ever returns \(\bar y\) again, \(z_{M_K} \ge z_{\mathrm{inc}}\) and the stopping test fires.

No repetition, infeasible case. Let \(\bar x\) solve \(\mathrm F(\bar y)\) with the multipliers \(\mu\) of (A3). If some \(x \in X\) satisfied all the linearizations at \((\bar x, \bar y)\) with \(y = \bar y\), then \(0 \ge \sum_i \mu_i [g_i(\bar x, \bar y) + \nabla_x g_i(\bar x, \bar y)^\top (x - \bar x)] \ge \sum_i \mu_i g_i(\bar x, \bar y) > 0\), a contradiction. The middle inequality holds because \(\sum_i \mu_i \nabla_x g_i(\bar x, \bar y)^\top (x - \bar x) \ge 0\) for every \(x \in X\), by the normal-cone condition. So \(\bar y\) is infeasible for every later master. Finiteness follows from \(|Y| < \infty\), and optimality from \(z_{M_K} \le z^\star \le z_{\mathrm{inc}}\) at termination. ∎

(Why a visited column is never chosen again) In a picture, the master is the integer program over the polyhedron cut out by the tangent planes collected so far. Convexity puts the curved feasible set inside that polyhedron, so the master's value is a bound. The KKT conditions say that at the NLP solution the tangent planes already describe the problem exactly in the direction of the objective. So once an integer column has been visited, the tangents taken there pin the master's value on that column to the true value, and the column cannot be chosen again at a lower value. Fletcher and Leyffer also exhibited worst cases in which the method visits every integer assignment. They showed how to replace the feasibility problem by an exact penalty and how to extend the method to nonsmooth convex problems through subdifferentials. The constraint qualification in (A3) is needed. Without it the restricted master at \(\bar y\) can have a value strictly below \(f(\bar x, \bar y)\) and the same assignment can recur indefinitely, which is why implementations add an integer cut \(\sum_{j : \bar y_j = 1} (1 - y_j) + \sum_{j : \bar y_j = 0} y_j \ge 1\) to exclude a visited binary assignment explicitly.Fletcher and Leyffer (1994), cited above; the integer cuts are in Bonami et al. (2008) and in A. Lundell, J. Kronqvist and T. Westerlund, "The supporting hyperplane optimization toolkit for convex MINLP", Journal of Global Optimization 84 (2022).

Algorithm 3.4.6  OUTER-APPROXIMATION
                 (multi-tree; Duran and Grossmann 1986,
                  Fletcher and Leyffer 1994)

Input   a convex MINLP (P); a tolerance eps >= 0; a starting integer
        assignment y^1 (for instance the rounding of the continuous
        relaxation's solution).
Output  an eps-optimal solution (x*, y*) of (P), or a proof that (P)
        is infeasible.

1.  K := {};  z_inc := +inf;  k := 1.

2.  Solve NLP(y^k).

       2a. If it is feasible with solution xbar^k and
           f(xbar^k, y^k) < z_inc:
              z_inc := f(xbar^k, y^k) and (x*, y*) := (xbar^k, y^k).

       2b. If it is infeasible: solve F(y^k) for xbar^k, the point of
           least weighted violation.

3.  K := K + {(xbar^k, y^k)}: add the constraint linearizations at
    (xbar^k, y^k) to the master, and the objective linearization in
    case 2a. Optionally add the integer cut that excludes y^k.

4.  Solve the master MILP (M_K) with the row eta <= z_inc - eps.
    If it is infeasible:
       stop; (x*, y*) is eps-optimal, or (P) is infeasible when
       z_inc = +inf.
    Otherwise let (x^{k+1}, y^{k+1}, eta^{k+1}) be its solution.

5.  k := k + 1;  go to 2.

Invariant
    z_M <= z* <= z_inc after every step 4 (the master is a relaxation,
    the incumbent is feasible), and every assignment visited in step 2
    has its column pinned at its true value in the master
    (Theorem 3.4.5), so it cannot be returned again below z_inc - eps.
The loop of Algorithm 3.4.6: the NLP and the master MILP take turns

  2. NLP(y^k): the curves, with y fixed at y^k <-----------------+
     |                                                           |
     +-- feasible:   2a. xbar^k, a feasible point; a new         |
     |                   incumbent when f(xbar^k, y^k) < z_inc   |
     +-- infeasible: 2b. xbar^k from F(y^k)                      |
     |                                                           |
     v                                                           |
  3. linearizations at (xbar^k, y^k) join K                      |
     |                                                           |
     v                                                           |
  4. master MILP (M_K) with eta <= z_inc - eps: the integers     |
     over the tangent planes, a relaxation                       |
     |                                                           |
     +-- solution (x^{k+1}, y^{k+1}, eta^{k+1}):                 |
     |   5. k := k + 1 ------------------------------------------+
     |
     +-- infeasible: stop; (x*, y*) is eps-optimal,
                     or (P) is infeasible when z_inc = +inf

  z_M <= z* <= z_inc after every step 4; a visited column is
  pinned at its true value, so it is not returned again below
  z_inc - eps: at most |Y| master solves (Theorem 3.4.5)

The cost of one round is one convex NLP in \(n\) variables, or one feasibility NLP, and one MILP with \(p\) integer variables and \(|K|(m + 1)\) added rows, solved from scratch or warm-started from the previous master's tree. The NLP is small and independent of the master. The MILP is the serial bottleneck, and the reason the single-tree variant below exists. Several integer assignments from the master's solution pool, the set of integer-feasible points a MILP solver keeps during a run besides its incumbent, can have their NLPs solved at once, each yielding valid cuts and a candidate incumbent, which is the natural parallel version of a round.

(Cutting planes and Benders as variants of the round) The extended cutting plane method is Algorithm 3.4.6 with step 2 deleted. The linearization of step 3 is taken at the master's own point for the most violated constraints, and no NLP is solved. No incumbent exists until a master point is feasible to tolerance, and the stopping test is the constraint violation at the master point. Its convergence theorem is Theorem 3.4.8 below. Generalized Benders decomposition keeps the NLP but aggregates the cuts through the multipliers, and it is weaker for a reason that can be written in one line.

Proposition 3.4.7 (generalized Benders cuts are aggregated outer-approximation cuts; Duran and Grossmann 1986, Quesada and Grossmann 1992). Under (A1) to (A3), let \(\bar x\) solve \(\mathrm{NLP}(\bar y)\) with multipliers \(\bar \lambda\), and write \(L(x, y, \lambda) = f(x, y) + \lambda^\top g(x, y)\). The generalized Benders cut of Geoffrion's method,

\[\eta \;\ge\; f(\bar x, \bar y) + \nabla_y L(\bar x, \bar y, \bar\lambda)^\top (y - \bar y), \tag{3.4.1}\]

is implied by the outer-approximation cuts at \((\bar x, \bar y)\): it is the objective linearization plus \(\bar\lambda_i\) times the \(i\)-th constraint linearization. Consequently the Benders master's value never exceeds the outer-approximation master's value built from the same points, \(z^{\mathrm{GBD}}_M \le z^{\mathrm{OA}}_M \le z^\star\), and Benders can need strictly more iterations.A. M. Geoffrion, "Generalized Benders decomposition", Journal of Optimization Theory and Applications 10 (1972); the comparison of bounds is in Duran and Grossmann (1986), cited above, and in I. Quesada and I. E. Grossmann, "An LP/NLP based branch and bound algorithm for convex MINLP optimization problems", Computers & Chemical Engineering 16 (1992).

Proof. Add \(\eta \ge f(\bar v) + \nabla f(\bar v)^\top (v - \bar v)\) to \(\sum_i \bar\lambda_i [g_i(\bar v) + \nabla g_i(\bar v)^\top (v - \bar v)] \le 0\), which is valid because \(\bar\lambda \ge 0\):

\[\eta \;\ge\; f(\bar v) + \sum_i \bar\lambda_i g_i(\bar v) + \nabla_x L(\bar v, \bar\lambda)^\top (x - \bar x) + \nabla_y L(\bar v, \bar\lambda)^\top (y - \bar y).\]

Complementarity removes \(\sum_i \bar\lambda_i g_i(\bar v)\), and stationarity removes the \(x\)-term when \(\bar x\) is interior to \(X\). At a boundary point the normal-cone term only strengthens the inequality on \(X\). What remains is (3.4.1). A single implied inequality is weaker than the family that implies it. ∎

(Benders against outer approximation on two rows) A hand computation shows the gap. Minimize \(x\) subject to \(x \ge y\), \(x \ge 1 - y\), \(x \in [0, 2]\) and \(y \in \{0, 1\}\). The optimum is \(1\). At \(\bar y = 0\) the NLP gives \(\bar x = 1\) with multipliers \(\bar\lambda = (0, 1)\). The outer-approximation cuts are the two constraints themselves, since they are affine, so the master already has value \(1 = z_{\mathrm{inc}}\) and outer approximation stops after one round. The Benders cut (3.4.1) is \(\eta \ge 1 + (0 + 1 \cdot (-1))(y - 0) = 1 - y\), whose master returns \(y = 1\) with \(\eta = 0 < 1\). A second subproblem at \(\bar y = 1\) adds \(\eta \ge y\), and only then does the Benders master reach \(1\). Benders projects everything onto the integer space through one set of multipliers, so each cut carries one linear functional of the constraints. Outer approximation keeps the continuous variables and every constraint separately. Benders is preferred when the subproblem decomposes, not for the strength of its bound.

Generalized Benders on the hand example (minimize x subject to x ≥ y, x ≥ 1 − y, x ∈ [0, 2], y ∈ {0, 1}): the master's η over y. After round 1 the cut η ≥ 1 − y lets the master take y = 1 with η = 0 < 1 = z_inc, so a second subproblem runs at ȳ = 1; after round 2 the cut η ≥ y puts both columns at 1 = z_inc, and the method stops. The rings are the master's η on a column, the filled dot the master's point.

Cutting planes without NLP subproblems: ECP and ESH

Outer approximation pays one NLP per round to obtain a feasible point and a supporting tangent. Two methods drop the NLP and take their linearizations elsewhere. The extended cutting plane method takes them at the master's own point, which is infeasible. The extended supporting hyperplane method slides that point back to the boundary along a segment first. The first is Kelley's method of 1960 with integer variables in the master, and its convergence theorem is Kelley's. Both work on the problem in epigraph form: a new variable \(\eta\) replaces the objective, the constraint \(\eta \ge f(v)\) joins the others, and the objective becomes \(\min \eta\). Then \(f\) enters only through one convex constraint and the master's objective is linear.

Theorem 3.4.8 (convergence of the extended cutting plane method; Kelley 1960, Westerlund and Pettersson 1995). Let (P) be a convex MINLP with \(X \times \operatorname{conv}(Y)\) compact and the \(g_i\) convex with gradients bounded by \(L\) on it, written in epigraph form so that \(f\) is linear. Let \(v^k = (x^k, y^k)\) solve the master \(\mathrm{(M_{K_k})}\) built from the linearizations of the most violated constraints at the previous master points \(v^1, \dots, v^{k-1}\), with no NLP solved, and let \(G(v) = \max_i g_i(v)\). Then (i) \(z_{M_{K_k}}\) is nondecreasing and at most \(z^\star\); (ii) for every \(\varepsilon > 0\) there is a finite \(k\) with \(G(v^k) \le \varepsilon\), and that \(v^k\) is an \(\varepsilon\)-feasible point of (P) with objective value at most \(z^\star\); (iii) every accumulation point of \(\{v^k\}\) is an optimal solution of (P).J. E. Kelley, Jr., "The cutting-plane method for solving convex programs", Journal of the Society for Industrial and Applied Mathematics 8 (1960); T. Westerlund and F. Pettersson, "An extended cutting plane method for solving convex MINLP problems", Computers & Chemical Engineering 19 (1995).

Proof sketch. (i) is Definition 3.4.4. For (ii), suppose \(G(v^k) > \varepsilon\) for infinitely many \(k\). For \(j > k\) the point \(v^j\) satisfies the cut generated at \(v^k\) for the most violated constraint \(i_k\), so \(g_{i_k}(v^k) + \nabla g_{i_k}(v^k)^\top (v^j - v^k) \le 0\). Hence \(\varepsilon < G(v^k) \le \lVert \nabla g_{i_k}(v^k) \rVert\, \lVert v^j - v^k \rVert \le L \lVert v^j - v^k \rVert\). Infinitely many points pairwise at distance more than \(\varepsilon / L\) cannot lie in a compact set. For (iii), an accumulation point \(\bar v\) has \(G(\bar v) \le 0\) by (ii) and continuity, is integral in \(y\) because \(Y\) is finite, and has objective at most \(z^\star\) by (i), so it is optimal. ∎

(Convergence in the limit, and pseudoconvexity) This is convergence in the limit. The method is finite only up to a tolerance on the constraint violation, and until a master point is \(\varepsilon\)-feasible it has produced no feasible point at all, since every linearization is taken at an infeasible point. Its advantage is cost per round, not primal information. Westerlund and Pörn extended the scheme to pseudoconvex constraints, which is the basis of the GAMS solver AlphaECP. A differentiable function \(g\) is pseudoconvex when \(\nabla g(a)^\top (b - a) \ge 0\) implies \(g(b) \ge g(a)\), which is enough for a linearization at \(a\) to separate a point \(b\) with \(g(b) > g(a)\).T. Westerlund and R. Pörn, "Solving pseudo-convex mixed integer optimization problems by cutting plane techniques", Optimization and Engineering 3 (2002).

The supporting hyperplane method keeps the cost and recovers the quality of the cuts.

Proposition 3.4.9 (the extended supporting hyperplane cut; Veinott 1967, Kronqvist, Lundell and Westerlund 2016). Assume the \(g_i\) are convex and continuously differentiable, and that a point \(v_{\mathrm{int}}\) with \(G(v_{\mathrm{int}}) < 0\) is known. Let \(\hat v\) be a master solution with \(G(\hat v) > 0\). Then (a) there is a unique \(\tilde t \in (0, 1)\) with \(G(v_{\mathrm{int}} + \tilde t (\hat v - v_{\mathrm{int}})) = 0\), and we write \(\tilde v\) for that point. (b) For every \(i\) active at \(\tilde v\), the linearization \(g_i(\tilde v) + \nabla g_i(\tilde v)^\top (v - \tilde v) \le 0\) is valid for \(\mathcal F_{\mathrm{rel}}\), supports it at \(\tilde v\), and is violated by \(\hat v\). (c) The method that adds these supporting hyperplanes in place of the cuts at \(\hat v\) converges in the sense of Theorem 3.4.8.A. F. Veinott, Jr., "The supporting hyperplane method for unimodal programming", Operations Research 15 (1967), for the continuous ancestor; J. Kronqvist, A. Lundell and T. Westerlund, "The extended supporting hyperplane algorithm for convex mixed-integer nonlinear programming", Journal of Global Optimization 64 (2016).

Proof sketch. (a) The function \(\varphi(t) = G(v_{\mathrm{int}} + t(\hat v - v_{\mathrm{int}}))\) is convex and continuous on \([0, 1]\) with \(\varphi(0) < 0 < \varphi(1)\). A convex function that is negative at \(0\) and positive at \(1\) crosses zero exactly once, since it is strictly increasing from the crossing on. (b) Validity is Proposition 3.3.11(a) and support is the equality at \(\tilde v \in \mathcal F_{\mathrm{rel}}\). For the violation, let \(i\) be active at \(\tilde v\) and put \(\varphi_i(t) = g_i(v_{\mathrm{int}} + t(\hat v - v_{\mathrm{int}}))\), convex with \(\varphi_i(0) \le G(v_{\mathrm{int}}) < 0 = \varphi_i(\tilde t)\). Its derivative at \(\tilde t\) is at least the secant slope \((\varphi_i(\tilde t) - \varphi_i(0))/\tilde t > 0\), and that derivative equals \(\nabla g_i(\tilde v)^\top (\hat v - v_{\mathrm{int}})\). So \(\nabla g_i(\tilde v)^\top (\hat v - \tilde v) = (1 - \tilde t)\, \nabla g_i(\tilde v)^\top (\hat v - v_{\mathrm{int}}) > 0\), which is the violation. (c) The argument of Theorem 3.4.8 applies because the new cut separates \(\hat v\) from the current polyhedron by an amount bounded below in terms of \(G(\hat v)\). The details are in the 2016 paper. ∎

(Sliding the cut back to the boundary) The picture is that the cut at the infeasible master point is a valid plane floating outside the convex set. Sliding it inward along the segment to the interior point until it touches the set gives a supporting hyperplane that removes at least as much of the polyhedron in the direction of \(\hat v\). Outer approximation's cuts are also supporting hyperplanes, at NLP solutions, so the method obtains cuts of that quality at the cost of a line search, which needs only function evaluations, instead of an NLP solve. Serrano, Schwarz and Gleixner later showed that the iteration is Kelley's cutting-plane method applied to a reformulated problem. In it the constraint functions are replaced by the gauge of the feasible set with respect to the interior point, that is, the smallest \(t > 0\) such that \(v_{\mathrm{int}} + (v - v_{\mathrm{int}})/t\) lies in the set. Kelley's convergence proof therefore carries over, and the functions that describe the convex feasible set need not themselves be convex, only differentiable and satisfying a mild condition.F. Serrano, R. Schwarz and A. Gleixner, "On the relation between the extended supporting hyperplane algorithm and Kelley's cutting plane algorithm", Journal of Global Optimization 78 (2020). The interior point comes from one auxiliary convex program, \(\min\{\nu : g_i(v) - \nu \le 0\ \forall i\}\), which need not be solved to optimality, and SHOT stops it as soon as a point with \(\nu < 0\) is in hand.Lundell, Kronqvist and Westerlund (2022), cited above.

Algorithm 3.4.10  EXTENDED-SUPPORTING-HYPERPLANE
                  (Kronqvist, Lundell and Westerlund 2016)

Input   a convex MINLP (P) in epigraph form (linear objective);
        a tolerance eps > 0 on the constraint violation; a MILP solver.
Output  an eps-feasible point of (P) with objective <= z*, hence
        eps-optimal, or infeasibility.

1.  Interior point: solve, approximately,
       min nu  s.t.  g_i(v) - nu <= 0 for all i,  v in X x conv(Y);
    any solution with nu < 0 is v_int. If none exists, fall back to
    the extended cutting plane method.

2.  K := {}. Repeat:

       2a. Solve the master MILP (M_K); let vhat be its solution.
           If infeasible: (P) is infeasible; stop.

       2b. If G(vhat) <= eps: stop; vhat is eps-feasible and
           z_M <= z*.

       2c. Line search: find t~ in (0, 1) with
              G(v_int + t~ (vhat - v_int)) = 0
           by bisection (G is convex along the segment, so the root
           is unique); vtilde := the root.

       2d. For each i with g_i(vtilde) >= -tol (active at vtilde)
           add the supporting hyperplane
              g_i(vtilde) + grad g_i(vtilde)^T (v - vtilde) <= 0
           to K.

Invariant
    every cut is valid for F_rel and supports it (Proposition 3.4.9),
    so z_M <= z*; each cut is violated by the master point that
    produced it, which drives Kelley's argument.

One round is one MILP and a bisection of \(O(\log(1/\mathrm{tol}))\) evaluations of \(G\), with no NLP. In SHOT's single-tree mode the MILP is not re-solved from scratch. The cuts enter as lazy constraints, rows that the MILP solver asks for only when an integer-feasible node appears, through a callback, and one round is then one callback inside one tree. The bisections for a pool of master points are independent of one another, which the listing at the end of the section exploits.

One tree instead of many

Every method so far restarts a MILP master in each round and, in its basic form, throws away the tree the previous round built. Quesada and Grossmann's observation was that the master's branch and bound can be interrupted instead of restarted. Whenever a node's LP solution is integral in \(y\), solve the NLP on that column, add its linearizations to every open node, and continue the same tree.

Many trees or one: the master MILP in Algorithms 3.4.6 and 3.4.12

  multi-tree: every round restarts a MILP master

    NLP --> [ MILP tree 1 ] --> NLP --> [ MILP tree 2 ] --> NLP ...
            thrown away                 thrown away

  single tree: one branch and bound, interrupted, never restarted

    [ one MILP tree ------------------------------------------- ]
                 |                |                  |
                NLP              NLP                NLP
         at a node whose LP solution is integral in y; its
         linearizations go to every open node, and the same
         tree goes on

Proposition 3.4.11 (LP/NLP-based branch and bound; Quesada and Grossmann 1992, Bonami et al. 2008). Under (A1) to (A3), Algorithm 3.4.12 terminates finitely with an optimal solution of (P). Each cut it adds is globally valid. Each distinct integer assignment \(\hat y\) triggers at most one NLP solve before it is either certified optimal on its column or cut off, and the number of NLP solves is at most \(|Y|\).Quesada and Grossmann (1992), cited above; Bonami et al. (2008), cited above, whose B-QG and B-Hyb algorithms are this method and a hybrid with NLP solves at nodes of selected depths.

Proof sketch. The procedure is Theorem 3.4.5 with the master MILP solved by a branch and bound that is interrupted at integer-feasible nodes instead of being restarted. The outer-approximation cuts at \((\bar x, \hat y)\) raise the restricted master's value on the column \(\hat y\) to \(f(\bar x, \hat y)\) in the feasible case, or make the column infeasible in the infeasible case. So the node LP after the cuts either produces a new \(\hat y\), or is pruned, or has value at least \(z_{\mathrm{inc}}\). The cuts are globally valid by convexity. Finiteness follows from that of the tree and from \(|Y| < \infty\). ∎

Algorithm 3.4.12  LP/NLP-BASED-BRANCH-AND-BOUND
                  (single tree; Quesada and Grossmann 1992)

Input   a convex MINLP (P); a tolerance eps; an initial cut set K
        (for instance the linearizations at the continuous
        relaxation's solution, so that the root LP is bounded).
Output  an eps-optimal solution of (P).

1.  z_inc := +inf. Open list L := {root}, the LP relaxation of (M_K)
    over y in conv(Y).

2.  While L is nonempty:

       2a. Pop a node N (best bound first, or depth first); solve its
           LP; let (xhat, yhat, etahat) solve it.
           If infeasible or etahat >= z_inc - eps: discard N and
           continue.

       2b. If yhat is fractional: branch on a fractional y_j, push the
           two children, continue.

       2c. yhat is integral:
              solve NLP(yhat), or F(yhat) if it is infeasible, for xbar;
              update z_inc if the point is feasible and better;
              add the linearizations at (xbar, yhat) to EVERY open node
              (they are globally valid);
              re-solve N's LP with the new rows and return to 2a with
              the same node.

3.  Return the incumbent.

Invariant
    every node LP is a relaxation of (P) restricted to the node's box,
    since the cuts are globally valid by convexity, so pruning is
    safe; a column yhat on which NLP(yhat) has been solved has LP
    value >= its true value at every later visit (Proposition 3.4.11),
    so each yhat triggers at most one NLP solve.
The node loop of Algorithm 3.4.12, interrupted at integral nodes

  2. while L is nonempty <-------------------------------------+
     |                                                         |
  2a. pop a node N (best bound first, or depth first)          |
     |                                                         |
  +->solve N's LP: (xhat, yhat, etahat)                        |
  |  |                                                         |
  |  +-- infeasible, or etahat >= z_inc - eps: discard N ------+
  |  |                                                         |
  |  +-- 2b. yhat fractional: branch on a fractional y_j,      |
  |  |       push the two children ----------------------------+
  |  |
  |  +-- 2c. yhat integral: NLP(yhat), or F(yhat) if it is
  |          infeasible, for xbar; update z_inc if the point
  |          is feasible and better; add the linearizations at
  |          (xbar, yhat) to EVERY open node
  |                     |
  +---------------------+  the same node N, with the new rows

  L empty: 3. return the incumbent

One node costs one dual-simplex re-solve, warm-started from the parent's basis, plus an NLP at the integer-feasible nodes only. The tree parallelizes as a tree does, through the node pool of Section 6, with the extra rule that a cut from an integer-feasible node must reach every open node. The NLP solves at different integer-feasible nodes are independent of one another.

Quesada and Grossmann proposed the method to avoid solving a sequence of MILP masters. Modern implementations add the linearizations through the MILP solver's lazy-constraint callback, so that the MILP solver's presolve, cuts, branching and heuristics all work on the master unchanged. They include SHOT, AOA in AIMMS, Bonmin's B-QG, Minotaur's QG engine, Muriqui and Knitro's QG mode.Kronqvist, Bernal, Lundell and Grossmann (2019) and Lundell, Kronqvist and Westerlund (2022), both cited above. The 2019 comparison of sixteen convex MINLP solvers concluded that "SHOT and AOA were the overall fastest solvers. Both of the solvers are based on a single-tree approach closely integrated with the MILP solver by utilizing callbacks to add the linearizations as lazy constraints", and that the differences among solvers sharing an algorithm "are mainly due to different degrees of preprocessing, primal heuristics, cut generation procedures, and different strategies".Kronqvist, Bernal, Lundell and Grossmann (2019), cited above. The table collects the family.

methodper roundbound fromfeasible point fromterminates
NLP branch and boundone NLP per nodethe NLPintegral NLP solutionsfinitely, for bounded integers
outer approximation (OA)one NLP + one MILP masterthe masterthe NLP, every roundfinitely, at most \(\vert Y \vert\) masters (needs a KKT point at each NLP)
generalized Benders (GBD)one NLP + one MILP masterthe masterthe NLP; cuts are aggregations of OA'sfinitely; its bound is never better than OA's from the same points
extended cutting plane (ECP)one MILP master, no NLPthe masteronly at convergencein the limit; finite to a tolerance on the violation
extended supporting hyperplane (ESH)one MILP master + a line search, no NLPthe masteronly at convergencein the limit; finite to a tolerance, with OA-quality cuts
LP/NLP branch and boundone tree: an LP per node, an NLP at integer-feasible nodesthe treethe NLPsfinitely, at most \(\vert Y \vert\) NLPs
The convex-MINLP family: what each round solves, what it gives you, and how it terminates.

The three methods on one ellipse

The figure below runs outer approximation, the extended cutting plane method and the extended supporting hyperplane method on the smallest problem that shows the difference between them: one integer variable, one continuous variable and one convex constraint. It maximizes

\[\max\ \cos\theta\, x + \sin\theta\, y \quad \text{subject to} \quad g(x, y) = \Big(\frac{x - 4.2}{3.4}\Big)^2 + \Big(\frac{y - 3.6}{2.6}\Big)^2 - 1 \le 0, \qquad x \in \{0, \dots, 8\},\ y \in [0, 8],\]

so the feasible set is the part of an ellipse with centre \((4.2, 3.6)\) and semi-axes \(3.4\) and \(2.6\) that lies on the nine integer columns. The columns \(x = 0\) and \(x = 8\) miss the ellipse entirely. Here \(x\) plays the part of the integer variable \(y\) of the general notation, and because the maximization reverses the inequalities, the master gives an upper bound and the NLP a lower one. In the symbols of Definitions 3.4.1 and 3.4.2, \(X = [0, 8]\) is the continuous box, \(Y = \{0, \dots, 8\}\) is the set of nine columns, the objective is \(-(\cos\theta\, x + \sin\theta\, y)\) after the sign change, the one constraint is \(g\), the subproblem \(\mathrm{NLP}(\bar x)\) at a column \(\bar x\) is the one-variable problem of moving along that column as far as the objective wants while staying inside the ellipse, and the feasibility subproblem \(\mathrm F(\bar x)\) on a column that misses the ellipse returns the point of the column nearest to it, which on \(x = 8\) is \((8, 3.6)\). The master's optimum over nine columns is found exactly by scanning them. At the default angle \(\theta = 40^\circ\) outer approximation needs 3 iterations, the supporting hyperplane method 6 and the cutting plane method 7, all ending at \(8.624\) at \((7, 5.075)\).

(The three runs at 40 degrees) The outer-approximation run reads as follows. The empty master picks the box corner \((8, 8)\) with bound \(11.271\). The column \(x = 8\) lies outside the ellipse, so the NLP on it is infeasible. The feasibility problem returns \((8, 3.6)\), the point of least violation, and the linearization there is the feasibility cut \(x \le 7.62\), which removes the column for good (Theorem 3.4.5, infeasible case). The master moves to \((7, 8)\) with bound \(10.505\). The NLP on column \(7\) returns the top of the ellipse, \((7, 5.075)\), with value \(8.624\), which is the incumbent and, once the tangent there is added, the master's value as well. Bound and incumbent meet after three master solves, and every iteration from the second onward hands back a feasible point. The cutting plane method takes its first three cuts at the infeasible master points \((8, 8)\), \((8, 5.61)\) and \((8, 4.19)\), then three more on column \(7\), and has no feasible point until its last iteration. The supporting hyperplane method's first cut touches the ellipse at \((6.074, 5.770)\), on the segment from the centre to \((8, 8)\). It brings the master from \(11.27\) to \(9.21\), where the cutting plane method's first cut only reaches \(9.73\). The tables below carry the rest of both runs.

The first round at theta = 40 degrees, three ways

       master point (8, 8), bound 11.271, outside the ellipse
                             |
       +---------------------+---------------------+
       | OA                  | ECP                 | ESH
       v                     v                     v
  the NLP on column 8   cut at the master     cut at (6.074, 5.770),
  is infeasible; the    point (8, 8) itself   where the segment
  feasibility problem                         from the centre
  gives (8, 3.6): the                         (4.2, 3.6) to (8, 8)
  cut x <= 7.62                               meets the ellipse
       |                     |                     |
       v                     v                     v
  master (7, 8)         master (8, 5.609)     master (8, 4.797)
  bound 10.505          bound 9.734           bound 9.212

  ESH's line search on that segment, by bisection:

  (4.2, 3.6)           (6.074, 5.770)                (8, 8)
  o---------------------------*---------------------------o
  inside: g < 0        boundary: g = 0       outside: g > 0

Turn the objective to \(100^\circ\), almost straight up. In the script below, which runs the same iteration as the figure, outer approximation then walks across seven columns before its bound \(5.420\) at \((3, 6.033)\) meets the incumbent found at the third iteration, eight iterations in all, while the supporting hyperplane method again needs six. The figure has a case for this angle, but its readout there was not checked for this series, so these counts are the script's. The constraint control switches the ellipse to a crescent, which is not convex, and is discussed below.

Outer approximation on one integer variable, one continuous variable and one nonlinear constraint, with two other cut-generating methods to compare. The master picks the best point the cuts allow (orange ring), and the chosen method takes one tangent there (purple): outer approximation at the NLP solution on that column (green), the extended cutting plane method at the master point itself, and the extended supporting hyperplane method where the segment from an interior point to the master point crosses the boundary (purple dot). With the constraint switched to the crescent, which is not convex, a tangent taken where the bite is the active piece curves the wrong way, is drawn in red, and can remove feasible points, the true optimum (blue) among them.

The script below reproduces the three runs. The master is solved by the column scan, the NLP on a column has a closed form, and the line search is sixty steps of bisection. It prints the iteration tables at \(40^\circ\) and the iteration counts at \(100^\circ\).

# Outer approximation, ECP and ESH on the ellipse of fig-oa.
#
# Outer approximation, the extended cutting plane method and the
# extended supporting hyperplane method on the figure's problem (a
# maximization, as the figure is):
#
#   maximize cos(theta) x + sin(theta) y,  x in {0, ..., 8},  y in [0, 8],
#   g(x, y) = ((x - 4.2)/3.4)^2 + ((y - 3.6)/2.6)^2 - 1 <= 0.
#
# The master is a MILP in (x, y) over the box and the accumulated cuts;
# with one integer variable it is solved exactly by scanning the nine
# columns. The script prints the iteration tables at 40 degrees and the
# iteration counts at 100 degrees.

import math

# the ellipse: centre (CX, CY), semi-axes RX and RY
CX, CY, RX, RY = 4.2, 3.6, 3.4, 2.6

def g(x, y):
    return ((x - CX) / RX) ** 2 + ((y - CY) / RY) ** 2 - 1.0

def grad(x, y):
    return 2 * (x - CX) / RX ** 2, 2 * (y - CY) / RY ** 2

def cut_at(x, y):
    """The linearization of g <= 0 at (x, y), as a x + b y <= r."""
    a, b = grad(x, y)
    return a, b, a * x + b * y - g(x, y)

def master(c, cuts):
    """Max c.(x, y) over the box and the cuts, by column scan."""
    best = None
    for x in range(9):
        lo, hi = 0.0, 8.0
        for a, b, r in cuts:
            if b > 1e-12:
                hi = min(hi, (r - a * x) / b)
            elif b < -1e-12:
                lo = max(lo, (r - a * x) / b)
            elif r - a * x < -1e-9:
                # a cut with no y: the column is infeasible
                lo = 9.0
        if lo > hi + 1e-9:
            continue
        y = hi if c[1] >= 0 else lo
        z = c[0] * x + c[1] * y
        if best is None or z > best[0] + 1e-12:
            best = (z, x, y)
    return best

def nlp(c, x):
    """NLP(x): the best y on the column, and whether it is feasible.

    On an infeasible column, y is the point of least violation.
    """
    u = (x - CX) / RX
    if abs(u) > 1:
        return CY, False
    h = RY * math.sqrt(1 - u * u)
    return (min(8.0, CY + h) if c[1] >= 0 else max(0.0, CY - h)), True

def boundary(x, y):
    """ESH line search from the centre to (x, y), by bisection."""
    lo, hi = 0.0, 1.0
    for _ in range(60):
        t = 0.5 * (lo + hi)
        if g(CX + t * (x - CX), CY + t * (y - CY)) <= 0:
            lo = t
        else:
            hi = t
    return CX + hi * (x - CX), CY + hi * (y - CY)

def run(theta, method, tol=1e-6):
    """One method at angle theta: a row per master solve."""
    c = (math.cos(math.radians(theta)), math.sin(math.radians(theta)))
    cuts, rows, lower = [], [], -math.inf
    for it in range(1, 61):
        z, x, y = master(c, cuts)
        if g(x, y) <= tol:
            rows.append((it, x, y, z, lower, None,
                         "master point feasible: stop"))
            return rows
        if method == "OA":
            yn, feas = nlp(c, x)
            p = (x, yn)
            if feas:
                lower = max(lower, c[0] * x + c[1] * yn)
            note = ("tangent at the NLP point" if feas
                    else "column infeasible: feasibility cut")
        elif method == "ECP":
            p, note = (x, y), "cut at the master point"
        else:
            p = boundary(x, y)
            note = "supporting hyperplane at the boundary point"
        cuts.append(cut_at(*p))
        rows.append((it, x, y, z, lower, p, note))
    return rows

for i, (theta, method) in enumerate([(40, "OA"), (40, "ESH"), (40, "ECP"),
                                     (100, "OA"), (100, "ESH")]):
    rows = run(theta, method)
    if i:
        print()
    print(f"{method} at theta = {theta} deg: {len(rows)} iterations, "
          f"final value {rows[-1][3]:.3f} "
          f"at ({rows[-1][1]}, {rows[-1][2]:.3f})")
    if theta == 40:
        # each master solve on two lines: the master point, then the cut
        for it, x, y, z, lo, p, note in rows:
            ps = f"({p[0]:.3f}, {p[1]:.3f})" if p else "-"
            los = f"{lo:.3f}" if lo > -math.inf else "-"
            print(f"  it {it}: master ({x}, {y:.3f})  upper {z:6.3f}"
                  f"  incumbent {los:>6}")
            print(f"        cut at {ps:<16} {note}")
OA at theta = 40 deg: 3 iterations, final value 8.624 at (7, 5.075)
  it 1: master (8, 8.000)  upper 11.271  incumbent      -
        cut at (8.000, 3.600)   column infeasible: feasibility cut
  it 2: master (7, 8.000)  upper 10.505  incumbent  8.624
        cut at (7.000, 5.075)   tangent at the NLP point
  it 3: master (7, 5.075)  upper  8.624  incumbent  8.624
        cut at -                master point feasible: stop

ESH at theta = 40 deg: 6 iterations, final value 8.624 at (7, 5.075)
  it 1: master (8, 8.000)  upper 11.271  incumbent      -
        cut at (6.074, 5.770)   supporting hyperplane at the boundary point
  it 2: master (8, 4.797)  upper  9.212  incumbent      -
        cut at (7.344, 4.590)   supporting hyperplane at the boundary point
  it 3: master (7, 5.229)  upper  8.723  incumbent      -
        cut at (6.906, 5.174)   supporting hyperplane at the boundary point
  it 4: master (7, 5.080)  upper  8.627  incumbent      -
        cut at (6.997, 5.078)   supporting hyperplane at the boundary point
  it 5: master (7, 5.075)  upper  8.624  incumbent      -
        cut at (7.000, 5.075)   supporting hyperplane at the boundary point
  it 6: master (7, 5.075)  upper  8.624  incumbent      -
        cut at -                master point feasible: stop

ECP at theta = 40 deg: 7 iterations, final value 8.624 at (7, 5.075)
  it 1: master (8, 8.000)  upper 11.271  incumbent      -
        cut at (8.000, 8.000)   cut at the master point
  it 2: master (8, 5.609)  upper  9.734  incumbent      -
        cut at (8.000, 5.609)   cut at the master point
  it 3: master (8, 4.185)  upper  8.818  incumbent      -
        cut at (8.000, 4.185)   cut at the master point
  it 4: master (7, 5.291)  upper  8.764  incumbent      -
        cut at (7.000, 5.291)   cut at the master point
  it 5: master (7, 5.089)  upper  8.633  incumbent      -
        cut at (7.000, 5.089)   cut at the master point
  it 6: master (7, 5.075)  upper  8.624  incumbent      -
        cut at (7.000, 5.075)   cut at the master point
  it 7: master (7, 5.075)  upper  8.624  incumbent      -
        cut at -                master point feasible: stop

OA at theta = 100 deg: 8 iterations, final value 5.420 at (3, 6.033)

ESH at theta = 100 deg: 6 iterations, final value 5.420 at (3, 6.033)

(Reading the iteration tables) The incumbent column is the difference between the methods that the counts do not show. Outer approximation carries a feasible point from its second iteration, and a solver that is stopped early at any time returns one, while the two cutting plane methods carry nothing until the last line. The cut-point column is the other difference. The supporting hyperplane method's cut points lie on the ellipse, as outer approximation's NLP points do, and its master value is at or below the cutting plane method's at every iteration. On a problem with many integer variables the master is a MILP, and each row of these tables is one MILP solve, so the counts are the cost that matters. The script's five runs are independent of one another, and within one round of outer approximation the NLP solves on different columns are independent too, which is the parallel structure the listing at the end of the section uses.

Separable extended formulations make the cuts stronger

The strength of a tangent depends on the space it is taken in. When a convex constraint is a sum of convex functions of separate variables, \(g(x) = \sum_j g_j(x_j) \le 0\), one can introduce one variable per term,

\[t_j \;\ge\; g_j(x_j), \qquad \sum_j t_j \;\le\; 0,\]

and take the tangents of the one-dimensional functions \(g_j\) instead of the tangents of \(g\). A tangent of \(g\) at a point is the sum of the tangents of the \(g_j\) at that point's coordinates, so one round in the extended variables reproduces the cut in the original ones. The gain comes from combining rounds. Take the disc \(x^2 + y^2 \le 1\) and the tangent at \((0.6, 0.8)\): in the original variables it is the single cut \(0.6x + 0.8y \le 1\). In the extended variables the same round gives \(t_1 \ge 1.2x - 0.36\) and \(t_2 \ge 1.6y - 0.64\), which through \(t_1 + t_2 \le 1\) imply the same cut. Now add a second round with tangents at \(x = 1.0\) and \(y = 0.3\), that is \(t_1 \ge 2x - 1\) and \(t_2 \ge 0.6y - 0.09\). The projection of the extended polyhedron onto \((x, y)\) is the intersection of every pairwise sum, so two rounds give four cuts,

\[0.6x + 0.8y \le 1, \qquad 1.2x + 0.6y \le 1.45, \qquad 2x + 1.6y \le 2.64, \qquad 2x + 0.6y \le 2.09,\]

where two rounds in the original variables give two. After \(r\) rounds the extended formulation carries \(r^2\) tangent combinations for the price of \(2r\) rows, and in \(d\) dimensions \(r^d\) for the price of \(dr\). Hijazi, Bonami and Ouorou analysed the effect for the Euclidean ball, which after a change of variables is the risk constraint of the portfolio problems of Section 4.4, and built an outer-inner approximation algorithm on the extended formulation. Kronqvist, Lundell and Westerlund made the reformulation automatic in SHOT, including the detection of separable terms after a change of variables, and measured its effect across the convex instances of MINLPLib. Vielma, Dunning, Huchette and Lubin carried the same idea to mixed-integer conic quadratic programming, where a lifted formulation of the second-order cone gives the polyhedral outer approximation that an LP-based branch and bound uses.H. Hijazi, P. Bonami and A. Ouorou, "An outer-inner approximation for separable mixed-integer nonlinear programs", INFORMS Journal on Computing 26 (2014); J. Kronqvist, A. Lundell and T. Westerlund, "Reformulations for utilizing separability when solving convex MINLP problems", Journal of Global Optimization 71 (2018); Vielma, Dunning, Huchette and Lubin (2017), cited in full in Section 4.8. The same tangents describe a smaller set when they are taken in a better-chosen space. Section 4 develops this.

Two rounds of tangents on the disc x^2 + y^2 <= 1, extended

  x^2 + y^2 <= 1  becomes  t_1 >= x^2, t_2 >= y^2, t_1 + t_2 <= 1

                        round 1, y = 0.8      round 2, y = 0.3
                        t_2 >= 1.6y - 0.64    t_2 >= 0.6y - 0.09
                      +---------------------+---------------------+
  round 1, x = 0.6    |                     |                     |
  t_1 >= 1.2x - 0.36  | 0.6x + 0.8y <= 1    | 1.2x + 0.6y <= 1.45 |
                      +---------------------+---------------------+
  round 2, x = 1.0    |                     |                     |
  t_1 >= 2x - 1       | 2x + 1.6y <= 2.64   | 2x + 0.6y <= 2.09   |
                      +---------------------+---------------------+

  each cell is its row plus its column, through t_1 + t_2 <= 1:
  four rows in t, four cuts on (x, y), where two rounds in the
  original variables give two; r rounds give r^2 for 2r rows,
  and r^d for dr in d dimensions

Conic outer approximation

When the convex constraints are conic, that is, membership of an affine image of the variables in a second-order or semidefinite cone, the dual solution of the continuous conic subproblem at a fixed integer assignment supplies vectors of the dual cone whose half-spaces pin the master's value on that column, so they play the role of the KKT multipliers in Theorem 3.4.5, and a conic infeasibility certificate plays the role of the feasibility subproblem. Section 4.8 defines the dual cone, gives the algorithm (Algorithm 4.8.4) and names the solvers that implement it.

What breaks on a nonconvex set

Every theorem of this section used convexity in the same place: to make the linearization of a constraint a valid inequality. Proposition 3.3.11(b) says what happens to the linearization of a concave constraint, and the running example \(\mathsf{st\_e13}\) makes it concrete. The instance is R3 of Section 1.6, which the next section writes out as (3.5.1): minimize \(2x + y\) subject to the one nonlinear constraint \(g(x, y) = 1.25 - x^2 - y \le 0\), the row \(x + y \le 1.6\), the bounds \(0 \le x \le 1.6\) and \(y \in \{0, 1\}\). Its optimum is \(2.0\) at \((0.5, 1)\), with the other integer solution \(2\sqrt{1.25} = 2.236\) at \((1.118, 0)\). The function \(g\) is concave, so the feasible side \(y \ge 1.25 - x^2\) lies above a downward parabola, and the feasible set is two horizontal segments, one on each integer row.MINLPLib, instance st_e13, minlplib.org/st_e13.html; the library's record and the instance's sources are given with (3.5.1) in Section 3.5.

Proposition 3.4.13 (outer approximation on \(\mathsf{st\_e13}\)). Let \(\ell_{\bar x}(x, y) = 1.25 + \bar x^2 - 2\bar x\, x - y\) be the linearization of \(g\) at a point with first coordinate \(\bar x\). (i) \(\ell_{\bar x} - g = (x - \bar x)^2 \ge 0\) everywhere, so \(\{\ell_{\bar x} \le 0\} \subseteq \{g \le 0\}\): every outer-approximation cut is a restriction, and the master built from any set of them has value at least \(z^\star\) when it is finite. The optimum \((0.5, 1)\) lies on the parabola, so every cut with \(\bar x \ne 0.5\) removes it, with slack \((0.5 - \bar x)^2\). (ii) Started from \(y = 0\), Algorithm 3.4.6 solves the NLP, which returns \((1.118, 0)\) with value \(2.236\), the incumbent, and adds the cut \(2.236\, x + y \ge 2.5\). In the master the column \(y = 1\) requires \(x \ge 1.5/2.236 = 0.671\) from the cut and \(x \le 0.6\) from \(x + y \le 1.6\), so it is infeasible, and the column \(y = 0\) has value \(2.236\), equal to the incumbent. The stopping test fires after one NLP and one master, and the method reports \(2.236\) as optimal with a gap of zero. The true optimum is \(2.0\). (iii) Started from \(y = 1\), the NLP returns \((0.5, 1)\) with value \(2.0\), the cut is \(x + y \ge 1.5\), the master's column \(y = 1\) has value \(2.0\) and its column \(y = 0\) is overestimated at \(3.0\), and the method stops with the correct value. The master is still not a relaxation, and the correct answer is an accident of the starting point.

Proof. (i) \(\ell_{\bar x} - g = 1.25 + \bar x^2 - 2\bar x x - y - 1.25 + x^2 + y = (x - \bar x)^2\), which is Proposition 3.3.11(b) with the inequality made exact. At the optimum \(g(0.5, 1) = 0\), so \(\ell_{\bar x}(0.5, 1) = (0.5 - \bar x)^2\). (ii) The NLP with \(y = 0\) requires \(x^2 \ge 1.25\), so \(x = \sqrt{1.25} = 1.118\) and the value is \(2.236\). The cut at \((1.118, 0)\) is \(\ell_{1.118} \le 0\), that is \(2.236\, x + y \ge 1.25 + 1.25 = 2.5\), and at the optimum it reads \(2.236 \cdot 0.5 + 1 = 2.118 < 2.5\). The two columns of the master are one-dimensional and read as stated. The stopping test compares the master's value with the incumbent and finds them equal. (iii) The same computation with \(\bar x = 0.5\): the cut \(x + y \ge 1.5\) forces \(x \ge 1.5\) on the row \(y = 0\), where the true value is \(2.236\). ∎

Proposition 3.4.13: outer approximation on st_e13, started from y = 0 (above) and from y = 1 (below). The feasible set is one piece on each row, 0.5 ≤ x ≤ 0.6 on y = 1 and 1.118 ≤ x ≤ 1.6 on y = 0; z* = 2.0 at (0.5, 1), and the other integer solution is 2.236 at (1.118, 0). Each tangent of the concave g is a restriction; the tint is what it removes from the feasible side, red what it removes of each row's piece, blue what the master keeps and the two integer solutions, a hollow blue ring where the cut removes one, and an orange ring marks the master's best point on a row. Started from y = 0, the cut removes the optimum and the run reports 2.236 as optimal with a gap of zero, while the true optimum is 2.0; started from y = 1, it keeps the optimum and stops with the correct value, but the master is still not a relaxation: its row y = 0 is overestimated at 3.0.

(Where the failure comes from) Two remarks on what the proposition does and does not say. The stopping argument of Theorem 3.4.5 used only that \(\bar x\) is a KKT point of the NLP, and that still holds here, since the active constraint's gradient is nonzero. What convexity would have supplied is \(z_M \le z^\star\), and on this instance the opposite inequality holds. So the failure is not a failure of the NLP solver or of the KKT argument but of hypothesis (A1) alone, and nothing inside the algorithm detects it. The "gap of zero" it reports is the difference between an incumbent and a master value that was never a lower bound. The same happens to the cutting plane and supporting hyperplane methods, whose linearizations at boundary points with \(\bar x \ne 0.5\) remove the optimum by (i). The figure above shows the phenomenon with a draggable objective when its constraint is switched to the crescent, the ellipse with a disc of radius \(2.2\) centred at \((6.8, 5.4)\) removed. The figure computes the true optimum by enumeration and marks it in blue, draws every cut that removes a feasible point in red, and marks with a green ring a final master point that is feasible but not optimal. A cut taken where the bite is the active piece of the boundary curves the wrong way, exactly as \(\ell_{\bar x}\) does.

(The repair: the chord instead of the tangent) The repair is to linearize a convex underestimator of \(g\) instead of \(g\). On a box \([a, b]\) the concave envelope of \(x^2\) is its chord \((a + b)x - ab\), so the convex envelope of \(g\) is affine, and the tangent is replaced by the chord constraint \(y \ge 1.25 + ab - (a + b)x\), which is valid. Propositions 3.5.11 and 3.5.12 compute the resulting bounds and show that fixing the integer variable leaves them where they are and that only a spatial branch, which shrinks the interval of \(x\), moves them. That is spatial branch and bound, the subject of the next section.

(What the solvers do on a nonconvex constraint) Solvers divide the work as follows. The global solvers add gradient cuts only to constraints they have proved convex and relax the rest by envelopes, as Section 3.3 said. SHOT, which is built for the convex case, runs on nonconvex instances as a heuristic with a valid dual bound. It detects convexity per constraint and avoids cuts from nonconvex constraints for as long as it can. When an invalid cut makes the master infeasible it repairs the master by adding slack to the cuts that came from nonconvex constraints, and it states that it has no guarantee of closing the gap.A. Lundell and J. Kronqvist, "Polyhedral approximation strategies for nonconvex mixed-integer nonlinear programming in SHOT", Journal of Global Optimization 82 (2022). On Mittelmann's MINLP benchmark set, the same 200 instances in the runs of October 2025 and 26 February 2026, SHOT handles only the quadratic instances, as slide 21 of his INFORMS talk of 28 October 2025 notes: H. D. Mittelmann, "Latest progress in optimization software", INFORMS Annual Meeting, 28 October 2025, plato.asu.edu/talks/informs2025.pdf. DICOPT's augmented-penalty variant of outer approximation (Viswanathan and Grossmann 1990) is described in secondary sources as relaxing the linearizations with penalized slacks, so that violated cuts do not make the master infeasible, and AlphaECP's pseudoconvex extension does the analogous thing. Both are local methods on nonconvex problems and report no valid bound there. Kesavan, Allgor, Gatzke and Barton give outer-approximation algorithms for separable nonconvex MINLP, which is what their title says. The paper was not read for this series.J. Viswanathan and I. E. Grossmann, "A combined penalty function and outer-approximation method for MINLP optimization", Computers & Chemical Engineering 14 (1990), described here from secondary sources; Westerlund and Pörn (2002), cited above; P. Kesavan, R. J. Allgor, E. P. Gatzke and P. I. Barton, "Outer approximation algorithms for separable nonconvex mixed-integer nonlinear programs", Mathematical Programming 100 (2004).

Where this is used

The convex case is the one corner of MINLP where the solvers finish. Kronqvist, Bernal, Lundell and Grossmann compared sixteen solvers on the 335 instances of MINLPLib classified as convex, with a limit of 900 seconds and a relative gap of 0.1%. Among the branch-and-bound solvers BARON solved 306, SCIP 295 and Minotaur's branch and bound 258. Among the decomposition solvers SHOT solved 312, AOA 310 and Muriqui 301. The virtual best, the count of instances solved by at least one of the solvers (Section 5.2), was 326 and the virtual worst 39, and the two single-tree solvers were the fastest overall.Kronqvist, Bernal, Lundell and Grossmann (2019), cited above; the counts are the paper's, on its 2018 solver versions. SHOT's own later benchmark on 406 convex MINLPLib instances with discrete variables, same limit and gap, counts SHOT with CPLEX and CONOPT at 379, BARON at 375, AOA at 364, SCIP at 356 and DICOPT at 349.Lundell, Kronqvist and Westerlund (2022), cited above; the authors note that 67 of the 406 instances are pure MIQPs that a MILP solver accepts directly. In the global solvers the same machinery lives on the convex pieces. BARON identifies convexity at every node and switches between nonlinear and polyhedral relaxations, and SCIP's convex nonlinear handler separates by tangents.A. Khajavirad and N. V. Sahinidis, "A hybrid LP/NLP paradigm for global optimization relaxations", Mathematical Programming Computation 10 (2018); Bestuzheva et al. (2025), cited above.

What parallelizes

What parallelizes is the primal side and the cut generation, not the master. The NLP subproblems at different integer assignments are independent. A round of outer approximation can fix every assignment in the master's solution pool at once, solve the NLPs together, and add every cut and every incumbent. Theorem 3.4.5 says each fixed column is pinned, so a round with \(k\) columns is at least as strong as \(k\) sequential rounds. The supporting hyperplane method's line searches need only function and gradient evaluations and are independent across a pool of master points, which is the simplest batched kernel in this chapter. The listing below computes the supporting hyperplane cuts for a pool of master points on the figure's ellipse, one thread per strided slice of the pool.

// Batched supporting-hyperplane cuts on the figure's ellipse.
//
// For a pool of master points, each thread finds the boundary point on
// the segment from a common interior point by bisection and returns the
// tangent there. The points are independent, so the batch splits across
// threads; on a GPU each point is one thread.

#include <algorithm>
#include <cstdio>
#include <span>
#include <thread>
#include <vector>

struct Cut {
    double a, b, r;  // the cut  a x + b y <= r
};

// the figure's ellipse
constexpr double CX = 4.2, CY = 3.6, RX = 3.4, RY = 2.6;

double g(double x, double y) {
    const double u = (x - CX) / RX, v = (y - CY) / RY;
    return u * u + v * v - 1.0;
}

// One master point (px, py) with g(px, py) > 0. On the segment from the
// centre c to p, g(c + t (p - c)) < 0 at t = 0 and > 0 at t = 1.
Cut esh_cut(double px, double py) {
    double lo = 0.0, hi = 1.0;
    for (int k = 0; k < 60; ++k) {
        const double t = 0.5 * (lo + hi);
        if (g(CX + t * (px - CX), CY + t * (py - CY)) <= 0.0)
            lo = t;
        else
            hi = t;
    }

    // the boundary point, and the gradient of g there
    const double x = CX + hi * (px - CX), y = CY + hi * (py - CY);
    const double a = 2.0 * (x - CX) / (RX * RX);
    const double b = 2.0 * (y - CY) / (RY * RY);
    return {a, b, a * x + b * y - g(x, y)};
}

std::vector<Cut> esh_batch(std::span<const double> xs,
                           std::span<const double> ys) {
    std::vector<Cut> out(xs.size());
    const unsigned T = std::max(1u, std::thread::hardware_concurrency());
    {
        // worker w takes the points w, w + T, w + 2T, ...
        std::vector<std::jthread> workers;
        for (unsigned w = 0; w < T; ++w)
            workers.emplace_back([&, w] {
                for (std::size_t i = w; i < xs.size(); i += T)
                    out[i] = esh_cut(xs[i], ys[i]);
            });
    }   // the jthreads join here, before out is returned
    return out;
}

int main() {
    // the pool: the top of every integer column
    std::vector<double> xs, ys;
    for (int x = 0; x <= 8; ++x) {
        xs.push_back(x);
        ys.push_back(8.0);
    }

    const auto cuts = esh_batch(xs, ys);
    for (std::size_t i = 0; i < cuts.size(); ++i)
        std::printf("column x = %d: %.4f x + %.4f y <= %.4f\n",
                    int(xs[i]), cuts[i].a, cuts[i].b, cuts[i].r);
}

(The batch, and two theory gaps) Each cut costs sixty-one evaluations of \(g\), sixty in the bisection and one at the boundary point, and one gradient, with no shared state between points, so the batch scales with the number of points until the pool is exhausted. On a device the same loop body runs as one thread per master point, and the only sequential step left in a round is the master MILP that consumes the cuts. The single-tree method parallelizes as a tree does, through the node pool of Section 6, with the one extra rule that a cut generated at an integer-feasible node must reach every open node. Two theory gaps stand behind all of this. There are no iteration bounds for outer approximation, the cutting plane method or the supporting hyperplane method beyond finiteness and convergence to a tolerance. And there is no analysis of outer approximation with inexact NLP solutions, which is what a batched first-order NLP solver on a GPU would return. The early-termination rules of Borchers and Mitchell are empirical, and the regularized masters of Kronqvist, Bernal and Grossmann, which reduce the iteration count in practice, have no rate attached either.J. Kronqvist, D. E. Bernal and I. E. Grossmann, "Using regularization and second order information in outer approximation for convex MINLP", Mathematical Programming 180 (2020).

One round with the batched cuts: only the master MILP is serial

                      master MILP (serial)
                                |
        the pool of master points; in the listing (x, 8)
        for x = 0, ..., 8, the top of every integer column
                                |
        +-------+-------+-------+-------+-------+-------+
        |       |       |       |       |       |       |
        v       v       v       v       v       v       v
     (0, 8)  (1, 8)  (2, 8)    ...   (6, 8)  (7, 8)  (8, 8)
        |       |       |               |       |       |
        each: 60 bisection steps on g from the centre, then
        the tangent at the boundary point; no shared state
        |       |       |               |       |       |
        +-------+-------+-------+-------+-------+-------+
                                |
                      9 cuts  a x + b y <= r
                                |
                                v
                      master MILP (serial)

  worker w of T takes the points w, w + T, w + 2T, ...;
  on a GPU each point is one thread

Spatial branch and bound

Branch and bound on integer variables (Section 3.1) terminates because every branch fixes something: a bounded integer variable can be split only finitely often. When the nonconvexity sits in a continuous variable, the branch splits an interval, the two children get tighter envelopes (Section 2.4), and nothing is ever fixed. The tree is infinite in principle and finite only because the search stops at a tolerance. This subsection gives the algorithm that every general-purpose global solver runs, spatial branch and bound, and the theorem that says why it works. Because nothing is ever fixed, that theorem is a convergence theorem rather than a finiteness theorem, with three hypotheses on the bounding, the selection and the subdivision. The subsection then gives the full node loop with the domain reduction of Section 2.6 built in, the two branching decisions the integer case does not have, and the tolerances the solvers stop at. It works three problems: the bilinear running example R2 as a tree, the two-variable MINLPLib instance st_e13 (R3), on which integer branching alone cannot close the gap and one spatial branch can, and Haverly's pooling problem (R4), the canonical bilinear instance with two local optima. Minimization is the default in every display. The two figures that maximize say so.

The framework and the convergence theorem

The problem is the one of Section 1,

\[z^\star \;=\; \min\{\, f(x) \;:\; g_i(x) \le 0,\ i = 1,\dots,m,\ x \in B_0 = [l^0, u^0],\ x_j \in \mathbb{Z} \text{ for } j \in I \,\},\]

with \(f\) and \(g_i\) continuous on the compact box \(B_0\) and feasible set \(\mathcal F\).

Definition 3.5.1 (the bounds of a run, indexed by step). Definitions 2.1.5, 2.3.1 and 3.1.1 carry over. A node \(N\) carries a box \(B_N = [l^N, u^N] \subseteq B_0\) and the subproblem obtained by restricting \(x\) to \(B_N\). Its relaxation and its node bound \(\bar z(N)\), with \(\bar z(N) = +\infty\) when the relaxation is infeasible, are those of Definition 3.1.1, and node bounds are taken monotone as there. The incumbent \(z_{\mathrm{inc}}\) is that of Definition 2.1.5 and the open list \(\mathcal L\) that of Definition 2.3.1. New here is the step index, because the convergence theorem below is about sequences. Write \(\mathcal L_k\) for the open list after step \(k\), the set of nodes neither subdivided nor deleted by then, \(z^{k}_{\mathrm{inc}}\) for the incumbent after step \(k\), and

\[\underline z_k \;=\; \min_{N \in \mathcal L_k} \bar z(N), \qquad \underline z_k \;\le\; z^\star \;\le\; z^{k}_{\mathrm{inc}}\]

for the global lower bound after step \(k\). This is the open-list minimum \(\min_{N \in \mathcal L} \bar z(N)\) of Theorem 3.1.5(c) with the step made explicit, and it equals the \(\underline z\) of Definition 2.3.1 after every deletion step, since the surviving nodes then have bounds below the incumbent. The absolute gap after step \(k\) is \(z^{k}_{\mathrm{inc}} - \underline z_k\). The relative gap divides it by a solver-dependent denominator, as Definition 2.1.5 says, and the table of Section 2.1 lists the denominators.

Definition 3.5.2 (the prototype procedure: Horst and Tuy, 1996). At step \(k\) there is a finite collection \(\mathcal L_k\) of boxes covering the part of \(B_0\) not yet excluded, a bound \(\bar z(N)\) for each, and an incumbent \(z^{k}_{\mathrm{inc}}\). One step consists of four operations. Deletion: delete every \(N\) with \(\bar z(N) \ge z^{k}_{\mathrm{inc}} - \varepsilon\) (deletion by bound, with \(\varepsilon \ge 0\)) and every \(N\) whose relaxation is infeasible. Selection: choose a node \(N_k \in \mathcal L_k\). Subdivision: split \(B_{N_k}\) into finitely many boxes whose union is \(B_{N_k}\). Bounding: compute \(\bar z\) for each child and update the incumbent from any feasible point found. Three properties of these operations are what the theory needs.R. Horst and H. Tuy, Global Optimization: Deterministic Approaches, 3rd ed. (Springer, 1996), Chapter IV, "Branch and Bound", pp. 115–178. The definitions and the convergence theorem below are that chapter's; the theorem's number inside the book was not checked for this series, so none is quoted. The paraphrase in M. Tawarmalani and N. V. Sahinidis, "Global optimization of mixed-integer nonlinear programs: a theoretical and computational study", Mathematical Programming 99 (2004), Section 2, is the one most global-solver papers cite.

  1. The bounding operation is consistent if every undeleted node can be subdivided further and, along every infinite nested sequence \(N^1 \supset N^2 \supset \cdots\) of successively subdivided nodes with \(N^q\) subdivided at step \(k_q\), \(\lim_{q \to \infty} \big( z^{k_q}_{\mathrm{inc}} - \bar z(N^q) \big) = 0\).
  2. The selection operation is bound improving if there is a constant \(K\) such that among every \(K\) consecutive steps at least one selected node attains the current global lower bound, \(\bar z(N_k) = \underline z_k\). Best-bound selection has \(K = 1\). Depth-first search is not bound improving. A rule that dives and returns to the best node every \(K\) steps is bound improving with that \(K\).
  3. The subdivision is exhaustive if \(\operatorname{diam}(B_{N^q}) \to 0\) along every infinite nested sequence of successively subdivided nodes.
Definition 3.5.2: one step of the prototype procedure

  +--> deletion      zbar(N) >= z_inc - eps, or infeasible
  |       |
  |       v
  |    selection     N_k in L_k ........ bound improving: in every K
  |       |                              steps some zbar(N_k) = z_lb
  |       |                              (best bound: K = 1)
  |       v
  |    subdivision   B_{N_k} into ...... exhaustive:
  |       |          finitely many       diam(B_{N^q}) -> 0
  |       |          boxes
  |       v
  |    bounding      zbar of each ...... consistent:
  |       |          child; incumbent    z_inc - zbar(N^q) -> 0
  +-------+          update

  N^1, N^2, ...: any infinite nested sequence of successively
  subdivided nodes; z_lb is the global lower bound of the step

Theorem 3.5.3 (convergence of branch and bound: Horst and Tuy, 1996). Run the prototype procedure with \(\varepsilon = 0\), with valid monotone node bounds, a consistent bounding operation and a bound-improving selection. If it does not terminate, then

\[\lim_{k \to \infty} \underline z_k \;=\; \lim_{k \to \infty} z^{k}_{\mathrm{inc}} \;=\; z^\star ,\]

and every accumulation point of the sequence of incumbents is a global minimizer.

Proof sketch. The sequence \(\underline z_k\) is nondecreasing and \(z^{k}_{\mathrm{inc}}\) nonincreasing, both bounded by \(z^\star\), so both converge, to \(\underline z_\infty \le z^\star \le z^{\infty}_{\mathrm{inc}}\) say. Suppose \(z^{\infty}_{\mathrm{inc}} - \underline z_\infty = \gamma > 0\). By bound-improving selection infinitely many selected nodes attain the current lower bound. Each of them has \(\bar z(N) \le \underline z_\infty\), because the lower bounds increase to \(\underline z_\infty\). Each is subdivided rather than deleted, because \(\bar z(N) \le \underline z_\infty < z^{k}_{\mathrm{inc}} - \gamma/2\) for all large \(k\). Call a node deep if infinitely many of these selected nodes are among its descendants. The root is deep, and a deep node has a deep child, because it has finitely many children and infinitely many of the nodes in question lie below them. Choosing a deep child at every level from the root gives an infinite path \(N^1 \supset N^2 \supset \cdots\) of nodes that were all subdivided.The two sentences prove the case needed of König's lemma, that an infinite, finitely branching rooted tree has an infinite path: D. Kőnig, "Über eine Schlussweise aus dem Endlichen ins Unendliche", Acta Scientiarum Mathematicarum (Szeged) 3 (1927), 121–130. Every node on the path is an ancestor of a node of bound at most \(\underline z_\infty\), so by monotonicity \(\bar z(N^q) \le \underline z_\infty\) for every \(q\). Consistency gives \(z^{k_q}_{\mathrm{inc}} - \bar z(N^q) \to 0\), while \(z^{k_q}_{\mathrm{inc}} - \bar z(N^q) \ge z^{\infty}_{\mathrm{inc}} - \underline z_\infty = \gamma > 0\) for every \(q\). The contradiction shows \(\underline z_\infty = z^{\infty}_{\mathrm{inc}}\), and both equal \(z^\star\). An accumulation point \(\bar x\) of the incumbents is feasible, since \(\mathcal F\) is closed, and \(f(\bar x) = \lim z^{k}_{\mathrm{inc}} = z^\star\) by continuity. ∎

(The picture behind the proof) The picture is this. At every step some node carries the global lower bound. A bound-improving rule keeps returning to such nodes and splitting them. Consistency says that splitting a node forever drives its bound up to the incumbent. So the bound-carrying nodes cannot stay a fixed distance below the incumbent, and the two bounds meet. The theorem says nothing about how fast. It also says nothing about termination: with \(\varepsilon = 0\) the procedure can run forever on a continuous nonconvexity, and Proposition 3.6.5 in the next subsection gives a two-variable instance on which it does.

Corollary 3.5.4 (finite \(\varepsilon\)-termination). Under the hypotheses of Theorem 3.5.3, the procedure with the stopping rule \(z^{k}_{\mathrm{inc}} - \underline z_k \le \varepsilon\), \(\varepsilon > 0\), stops after finitely many steps, and the incumbent satisfies \(z_{\mathrm{inc}} - z^\star \le \varepsilon\).

Proof. If it ran forever, Theorem 3.5.3 would give \(z^{k}_{\mathrm{inc}} - \underline z_k \to 0 < \varepsilon\), so the rule would have fired. At termination \(z_{\mathrm{inc}} - z^\star = z^{k}_{\mathrm{inc}} - z^\star \le z^{k}_{\mathrm{inc}} - \underline z_k \le \varepsilon\). ∎

(Where the hypotheses come from) The hypotheses are not abstract. Consistency is supplied by two concrete properties of the relaxations of Section 2.4, and exhaustiveness by a rule for the branching point. The next proposition makes both precise and gives the worst-case size of the tree, the exponential count \(k^d\) of boxes made exact.

Proposition 3.5.5 (exhaustive subdivision and convergent relaxations give consistency, and the size of the tree). Suppose the subdivision is exhaustive and the relaxations are pointwise convergent of order \(\alpha \ge 1\) in the sense of Definition 2.4.21: on every node box \(B\) of width \(w\) the relaxed objective and constraints satisfy \(0 \le f(x) - \bar f_B(x) \le C w^\alpha\) and \(0 \le g_i(x) - \bar g_{i,B}(x) \le C w^\alpha\) for all \(x \in B\). Suppose incumbents are accepted when they are \(\varepsilon_f\)-feasible, that is, when every constraint holds within \(\varepsilon_f\).

(a) On every node of width \(w \le (\varepsilon_f / C)^{1/\alpha}\) the relaxed optimal point \(\bar x_N\) is \(\varepsilon_f\)-feasible and \(f(\bar x_N) - \bar z(N) \le C w^\alpha\). Hence the bounding operation is consistent, and Theorem 3.5.3 and Corollary 3.5.4 apply.

(b) If every split bisects the longest edge of a box in \(\mathbb{R}^d\) of initial width \(\Delta\), then every node at depth \(D\) has width at most \(\Delta\, 2^{-\lfloor D/d \rfloor}\). With absolute tolerance \(\varepsilon\) the tree has depth at most \(d\,\lceil \log_2 (\Delta (C/\varepsilon)^{1/\alpha}) \rceil\) and at most

\[2^{\,d\,\lceil \log_2 (\Delta (C/\varepsilon)^{1/\alpha}) \rceil + 1} \;\le\; 2\,(2\Delta)^d \Big(\frac{C}{\varepsilon}\Big)^{d/\alpha}\]

nodes, a number of order \((\Delta^\alpha C/\varepsilon)^{d/\alpha}\). For McCormick relaxations of bilinear terms \(\alpha = 2\) (Theorem 2.4.7 and Proposition 2.4.22) and the bound is of order \((\Delta^2 C/\varepsilon)^{d/2}\).

Proof sketch. (a) The relaxed point satisfies \(\bar g_{i,B}(\bar x_N) \le 0\), so \(g_i(\bar x_N) \le C w^\alpha \le \varepsilon_f\), and \(\bar z(N) = \bar f_B(\bar x_N) \ge f(\bar x_N) - C w^\alpha\). Along an infinite nested sequence the width tends to zero, so from some step on the incumbent satisfies \(z^{k_q}_{\mathrm{inc}} \le f(\bar x_{N^q}) \le \bar z(N^q) + C w_q^\alpha\), which is consistency. (b) After \(d\) consecutive bisections of the longest edge every edge has at most half the length it had before those splits. If after \(d\) consecutive splits some edge still exceeded half that length, it was never split. Every split halved an edge that was then at least as long as that one, so each of the other \(d - 1\) edges could be split at most once. That allows at most \(d - 1\) splits, a contradiction. A node whose relaxed point closes its own gap to within \(\varepsilon\) is never subdivided again, since it is deleted by bound once the incumbent is updated. The depth bound follows, a binary tree of depth \(D\) has at most \(2^{D+1}\) nodes, and \(\lceil \log_2 X \rceil \le \log_2 X + 1\) gives the closed form. ∎

Proposition 3.5.5(b) in the plane (d = 2): four cuts on one path

  +-----------+-----------+
  |           |           |   cut  box after the cut   width
  |           1           |    1   Delta/2 x Delta     Delta
  |           |           |    2   Delta/2 x Delta/2   Delta/2
  +-----+--2--+           |    3   Delta/4 x Delta/2   Delta/2
  |     3     |           |    4   Delta/4 x Delta/4   Delta/4
  +--4--+     |           |
  |  *  |     |           |
  +-----+-----+-----------+

  the root box is Delta x Delta; each cut halves the longest edge
  of the box the cut before it made, and * is the box at depth 4:
  at depth D the width is at most Delta 2^-floor(D/d)

(Reading the worst-case count) Part (b) can be read in two ways. On the bilinear example below, \(d = 2\), \(\Delta = 1\), \(\alpha = 2\) and \(C = 1/4\), since the McCormick envelope of \(xy\) on a square of width \(w\) is at most \(w^2/4\) from the product (Theorem 2.4.7), so at \(\varepsilon = 0.01\) the bound is \(2 \cdot 2^2 \cdot 25 = 200\) nodes, against the 13 the search processes. With the width halved \(k\) times per coordinate the McCormick band falls by \(4^k\), and the number of boxes needed to certify tolerance \(\varepsilon\) in the worst case is of order \((1/\sqrt{\varepsilon})^d\). Branch and bound does far better than the worst case because most boxes are deleted by bound long before they are small: the bilinear example below needs 13 nodes where the uniform grid of the figure below needs 25. What the worst case does capture is the behaviour near the optimum, where boxes cannot be deleted by bound until they are small, and that is the subject of Theorem 3.5.7.

Proposition 3.5.6 (exhaustiveness and the branching point). Let each split of a box in coordinate \(j\) be at a point \(b\) with \(l_j + \beta (u_j - l_j) \le b \le u_j - \beta (u_j - l_j)\) for a fixed \(\beta \in (0, 1/2]\). (a) Each child has width at most \((1 - \beta)(u_j - l_j)\) in coordinate \(j\), and if along every infinite path every coordinate is split infinitely often, the subdivision is exhaustive. (b) With \(\beta = 0\) the guarantee fails in two ways. A branching point at a bound returns a child identical to the parent, and the procedure can repeat that node forever. Branching points strictly inside the box with relative offsets \(\beta_q\) summable, \(\sum_q \beta_q < \infty\), leave \(\prod_q (1 - \beta_q) > 0\), so the widths do not tend to zero along that path.

Proof. (a) is immediate from the clamp: a coordinate split infinitely often with a factor at most \(1 - \beta\) each time has width tending to zero. (b) \(\prod_q (1 - \beta_q) > 0\) if and only if \(\sum_q \beta_q < \infty\). ∎

(Why the branching point is pulled inward) This is why no production solver branches at the relaxed point \(\bar x_j\) itself. Belotti and coauthors write that "selecting a branching point too close to the bounds is likely to create a very easy subproblem and a very hard one", and the rules of every solver pull the point toward the midpoint. The rules are given with Algorithm 3.5.9 below.P. Belotti, J. Lee, L. Liberti, F. Margot and A. Wächter, "Branching and bounds tightening techniques for non-convex MINLP", Optimization Methods and Software 24 (2009), Section 5.5. The paper is the reference treatment of branching variable, branching point and bound tightening together, with Couenne as the test bed.

(The cluster problem) The convergence order of the relaxation is not only a constant in Proposition 3.5.5. It decides whether the number of boxes the search cannot delete near the optimum stays bounded as the tolerance shrinks. This is the cluster problem. Definition 2.4.21 and Proposition 2.4.22 gave the picture, and the theorem below makes it quantitative. It is the precise reason second-order relaxations such as McCormick's and \(\alpha\)BB are preferred to first-order interval bounds inside a tree. Two terms of the statement are the interval community's: the midpoint test updates the incumbent from the midpoint of every box, and no acceleration means that no monotonicity or interval-Newton test is used to delete boxes.

Theorem 3.5.7 (the cluster problem: Du and Kearfott, 1994). Consider interval branch and bound for the unconstrained minimization of a twice continuously differentiable \(f\) on a box in \(\mathbb{R}^m\), with the midpoint test and no acceleration. Boxes are bisected until each has width at most \(\epsilon\), and a box is deleted when its lower bound exceeds the best function value found. Let \(x^\star\) be a global minimizer at a vertex of one of the boxes, let the Hessian be positive definite near \(x^\star\) with smallest eigenvalue at least \(\lambda_{1,0} > 0\), and let the bounding scheme have order \(\alpha\) with constant \(K\). Then the number of boxes of width \(\epsilon\) around \(x^\star\) that remain in the list is at most

\[N \;=\; \Big\{\, 2 \Big\lfloor \sqrt{\tfrac{2K}{\lambda_{1,0}}}\ \sqrt{\epsilon^{\alpha - 2}} \Big\rfloor + 1 \Big\}^{m} .\]

Proof sketch. Expand \(f\) to second order at \(x^\star\). A box at distance \(n\epsilon\) from \(x^\star\) in some direction has \(f - f(x^\star) \ge \tfrac12 \lambda_{1,0} (n\epsilon)^2\) at its nearest point, while its lower bound is below the true minimum over the box by at most \(K \epsilon^\alpha\). The box survives only if \(\tfrac12 \lambda_{1,0} (n\epsilon)^2 \le K\epsilon^\alpha\), that is, for at most \(\sqrt{2K/\lambda_{1,0}}\, \epsilon^{(\alpha-2)/2}\) values of \(n\) in each of the \(m\) directions. ∎

(What the cluster theorem means for a device) For first-order bounds (\(\alpha = 1\)) the cluster contains of order \(\epsilon^{-m/2}\) boxes, growing without limit as the tolerance shrinks. For second-order bounds (\(\alpha = 2\)) the count is bounded independently of \(\epsilon\), though still exponential in \(m\).K. Du and R. B. Kearfott, "The cluster problem in multivariate global optimization", Journal of Global Optimization 5 (1994), Theorem 1. A. Wechsung, S. D. Schaber and P. I. Barton, "The cluster problem revisited", Journal of Global Optimization 58 (2014), show that second-order convergence is required to remove the dependence on the tolerance and that a small enough second-order prefactor removes the cluster entirely. R. Kannan and P. I. Barton, "The cluster problem in constrained global optimization", Journal of Global Optimization 69 (2017), extend the analysis to constrained problems, where clusters form on nearly optimal and on nearly feasible regions. Nothing removes the dependence on \(m\) except a formulation with fewer branching variables, which is the subject of reduced-space formulations below. For the GPU programme the theorem is a design criterion. A cheap first-order interval bound evaluated on a thousand boxes at once is not a substitute for a second-order bound on ten, because near the optimum the first-order scheme multiplies the frontier it has to bound.

The algorithm

Everything of Sections 2.4 to 2.6 is now assembled into one loop. The algorithm is written once, as the solvers run it, with the domain reduction, the local search and the branching rules at the steps where they occur.

Algorithm 3.5.8  Spatial branch and bound with domain reduction
                 (minimization)

Input   f, g_i factorable; box B_0; integer index set I; tolerances
        eps_a (absolute gap), eps_r (relative gap), eps_f (feasibility)
Output  an eps-globally optimal point x_inc with value z_inc, and the
        certificate z_lb >= z_inc - eps_a

 0. Preprocess: FBBT on B_0 (Section 2.6); OBBT at the root with the
    cutoff row, learning Lagrangian variable bounds; a multistart local
    NLP search for a first incumbent z_inc.

 1. open <- {root}; compute the root bound by step 5.

 2. while open is nonempty:

 3.    select N in open: best bound, or dive and return to the best
       node every K steps.                            [bound improving]

 4.    FBBT on B_N with the cutoff f(x) <= z_inc - eps_a; propagate
       the Lagrangian variable bounds;
       if some interval is empty: delete N and continue.

 5.    build the relaxation on B_N: envelope rows and
       outer-approximation rows on the factorable DAG, the linear rows,
       the cutoff row; solve the LP (warm start from the parent's
       basis); if infeasible: delete N.
       Optionally add separating cuts, re-propagate, and re-solve while
       the bound still moves enough.

 6.    zbar(N) <- max(LP value, zbar(parent)).
       With a first-order LP solver replace the LP value by the safe
       bound D(y) of Theorem 7.3.1, computed from the inexact dual y.
       If zbar(N) >= z_inc - eps_a, or the relative test holds:
       delete N.

 7.    reduced-cost tightening of B_N from (y, reduced costs, zbar(N),
       z_inc) (the duality-based range reduction of Theorem 2.6.18);
       if the box shrank a lot: go to 5 (a bounded number of times).

 8.    upper bounding: if the relaxed point xbar is eps_f-feasible,
       evaluate f(xbar); run a local NLP solve from xbar with the
       integer variables fixed at rounded values; update z_inc and
       x_inc; on improvement re-run deletion by bound over open.

 9.    if xbar is eps_f-feasible for N and f(xbar) - zbar(N) <= eps_a:
       delete N (its gap is closed).

10.    branch: choose a variable j by Algorithm 3.5.9, among fractional
       integer variables first, otherwise among the variables of
       violated nonconvex terms;
       choose a point b in [l_j + beta w_j, u_j - beta w_j];
       children
          N_1 = {x_j <= b}  (floor(b) for an integer variable) and
          N_2 = {x_j >= b}  (ceil(b));
       both inherit zbar(N); push them onto open.

11. return x_inc, z_inc, and the global bound
       z_lb = min over open of zbar(N)
                                    (z_lb = z_inc when open is empty).

Invariant
    every feasible x with f(x) <= z_inc - eps_a lies in the box of some
    node of open, and for every N in open,
        zbar(N) <= min{ f(x) : x in F, x in B_N }.

Termination
    Corollary 3.5.4 under exhaustive subdivision (Proposition 3.5.6)
    and convergent relaxations (Proposition 3.5.5).
The node loop of Algorithm 3.5.8 (minimization)

    0. preprocess B_0: FBBT, OBBT, multistart local NLP
    1. open <- {root}
       |
    2. while open is nonempty <----------------------------------+
       |                                                         |
    3. select N: best bound, or dive and return                  |
       |                                  [bound improving]      |
    4. FBBT with the cutoff: interval empty? ---yes-> delete N --+
       | no                                                      |
+-> 5. solve the LP on B_N: infeasible? --------yes-> delete N --+
|      | no                                                      |
|   6. zbar(N) <- max(LP value, zbar(parent)); for a first-order |
|      LP solver, the safe bound D(y) of Theorem 7.3.1 instead   |
|      zbar(N) >= z_inc - eps_a? ---------------yes-> delete N --+
|      | no                                                      |
|   7. reduced-cost tightening of B_N                            |
+<-yes- the box shrank a lot?  (a bounded number of times)       |
       | no                                                      |
    8. upper bounding: f(xbar) if xbar is eps_f-feasible; local  |
       NLP from xbar; update z_inc, x_inc; on improvement,       |
       deletion by bound                                         |
       |                                                         |
    9. xbar eps_f-feasible and                                   |
       f(xbar) - zbar(N) <= eps_a? -------------yes-> delete N --+
       | no                                   (gap closed)       |
   10. branch on x_j at b (Algorithm 3.5.9): N_1 = {x_j <= b},   |
       N_2 = {x_j >= b} inherit zbar(N), pushed onto open -------+

    open empty: 11. return x_inc, z_inc and z_lb

(The invariant, the cost of a node, and what parallelizes) The invariant is the same as for integer branch and bound: nothing better than the incumbent is ever lost, and every node bound is valid for its box. Domain reduction is admissible because every tightening step removes only points that are infeasible or no better than the incumbent, which is proved in general in Theorem 3.6.2. The cost per node has four parts. The first is one LP, whose row count is proportional to the number of nonlinear operations in the DAG, the expression graph of Section 2.4, plus the cuts. The second is a few FBBT sweeps, the feasibility-based bound tightening of Section 2.6, each costing two traversals of the DAG, one forward to compute the interval of every operation and one backward to tighten its operands. The third is one local NLP solve when the heuristic fires, and the fourth is the branching decision. What parallelizes: across a frontier of open nodes everything in steps 4 to 9 is independent between nodes, which is the batched-bounding opportunity of Sections 6.4 and 7.4. Inside a node, FBBT is order-independent across constraints (the fixed point of propagation does not depend on the schedule, Section 2.6) and the LP's matrix–vector products are the parallel pieces. Step 3 and the incumbent are the shared state.

Which variable, and where

Integer branching has one decision, the variable. Spatial branching has two, and both are heuristics with measured consequences.

Algorithm 3.5.9  Branching variable and branching point for a
                 continuous variable

Candidates   the variables x_j that appear in a nonconvex term whose
             relaxation is violated at the relaxed point xbar (an
             auxiliary w with wbar != theta(xbar)); fractional integer
             variables come first.

Variable scores

 largest violation   score_j = sum over violated terms containing x_j
                     of |wbar - theta(xbar)|, the violation split among
                     the term's variables (SCIP: by how central xbar_j
                     is in [l_j, u_j]).

 longest edge        score_j = u_j - l_j
                     (exhaustive by construction; BARON's second rule).

 violation transfer  gamma_j = width of the smallest interval around
                     xbar_j on which every term containing x_j can be
                     made consistent with its relaxed value;
                     score_j = gamma_j * sum_i |pi_i a_ij|
                     with pi the LP duals (Tawarmalani and Sahinidis
                     2004; BARON's default).

 pseudocosts         psi_j^-, psi_j^+ = average bound improvement per
                     unit of distance from xbar_j to the branching
                     point over past branchings on x_j; estimated gains
                     phi^-, phi^+;
                     score_j = alpha max(phi^-, phi^+)
                               + (1 - alpha) min(phi^-, phi^+)
                     (Couenne alpha = 0.15).

 reliability         strong branching (solve both children's LPs) until
                     x_j has eta reliable observations, then its
                     pseudocosts  (SCIP: pscostreliable = 2).

 SCIP's default      score_j = 1.0 * violation_j + 1.0 * pseudocost_j
                               + 0.5 * vartype_j
                     (continuous 0, binary 1, integer 0.1), summed over
                     the expressions; candidates within 90 % of the
                     best score tie, and ties are broken at random with
                     a fixed seed.

Branching point, local bounds [l, u], LP value xbar, width w = u - l

 Couenne, paper      b = alpha xbar + (1 - alpha) (l + u)/2
                     clamped to [l + beta w, u - beta w]
                     (Belotti et al. 2009, eq. (2)); alpha = 0.25,
                     beta = 0.2, and with these values the clamp is
                     never active

 Couenne, source     the mid-point rule midInterval: alpha = 0.25,
                     raised toward 1 once the relative gap is below
                     1e-3; clamp closeToBounds = 0.05;
                     default_clamp = 0.2 applies to the lp-clamped
                     strategies

 SCIP                m = 0.75 (branching/midpull); if the local width
                     is below half the global width,
                     m <- m * (relative width);
                     b = m (l + u)/2 + (1 - m) xbar,
                     clamped to [l + 0.2 w, u - 0.2 w]
                     (branching/clamp = 0.2)

 BARON               BrPtStra 0 dynamic (default); 1 omega = xbar;
                     2 bisection; 3 a convex combination of the two

 Integer             floor(xbar) + 0.5, or u - 0.5 when xbar sits at the
                     upper bound

Violation and pseudocost scores cost one pass over the candidates. Strong branching costs two LPs per candidate, and those LPs are independent of one another, which is the batch of Section 7.4.

(The branching point on one interval) A small example fixes the arithmetic of the branching point. Let \(x \in [0, 1]\) have relaxed value \(0.9\). Splitting at \(0.9\) gives children of widths \(0.9\) and \(0.1\). The McCormick band of the wide child is barely smaller than the parent's, and the relaxed point lies on the boundary of both children, so neither child's relaxation is forced to move. Splitting at the midpoint gives widths \(0.5\) and \(0.5\) and ignores the relaxed point. Couenne's blend gives \(0.25 \cdot 0.9 + 0.75 \cdot 0.5 = 0.6\), children of widths \(0.6\) and \(0.4\), with the relaxed point strictly inside the right child, which is what makes its relaxation change. A common misreading of Couenne's rule exchanges the weights and obtains \(0.8\). Couenne's \(\alpha = 0.25\) is the weight on the relaxed point, which gives \(0.6\). At the root of the bilinear example the LP value of \(x\) is \(0.4\) on \([0, 1]\), and both Couenne's rule and SCIP's give \(0.25 \cdot 0.4 + 0.75 \cdot 0.5 = 0.475\). They differ deeper in the tree. At a node where \(x \in [0.25, 0.5]\), a quarter of its original width, SCIP scales its midpull to \(0.75 \cdot 0.25 = 0.1875\) and branches almost at the LP value. Its stated rationale is that once a variable has been branched on, the LP solution deserves more weight. Couenne keeps its weights fixed until the relative gap falls below \(10^{-3}\), and then moves its weight toward the LP value.Belotti et al. (2009), Section 5.5, equation (2), is the paper's rule, with \(\alpha = 0.25\) and \(\beta = 0.2\) and the remark that the clamp is then never active. In Couenne's COIN-OR source (CouenneObject.hpp and CouenneObject.cpp, master, read 5 October 2026) the default mid-point rule midInterval blends with default_alpha = 0.25, clamps at closeToBounds = 0.05, and raises the weight on the LP point toward 1 once the relative gap is below \(10^{-3}\); default_clamp = 0.2 (branch_lp_clamp) is the constant of the lp-clamped and lp-central strategies, which are not the default. SCIP's branching/midpull = 0.75, branching/midpullreldomtrig = 0.5 and branching/clamp = 0.2 are the defaults in src/scip/set.c, and the rule is SCIPbranchGetBranchingPoint in branch.c (SCIP 10.0.0 sources, github.com/scipopt/scip). BARON's BrPtStra and BrVarStra options are in the BARON user manual, The Optimization Firm, minlp.com/baron-user-manual, Sections 5.4 and 11.4 (vendor documentation).

Splitting x ∈ [0, 1] when its relaxed value is 0.9: at the relaxed value, children of widths 0.9 and 0.1 with the relaxed point on the boundary of both; at the midpoint, widths 0.5 and 0.5, the relaxed point ignored; and at Couenne's blend, with weight 0.25 on the relaxed point (Belotti et al. 2009, eq. (2)), 0.25 · 0.9 + 0.75 · 0.5 = 0.6, widths 0.6 and 0.4, the relaxed point strictly inside the right child.

The violation-transfer and pseudocost rules are not the obvious choice, and the experiments say why. The simplest rule, branch on the variable whose term is most violated, is the obvious first choice. Belotti and coauthors compared it with violation transfer, strong branching and two reliability rules on 33 instances with a two-hour limit. Strong branching had the fewest nodes on 17 of the 33 instances but was fastest only once. It spent 90 to 95 percent of its time in the child LPs. On easy instances the cheap rules were fastest and kept the tree shallower. On hard MINLPs the reliability rules carried over from the integer case dominated. Their conclusion is that "the comparison between branching rules on MINLP instances does not show a clear winner".Belotti et al. (2009), Section 7.2 and Table 6: instances solved within two hours, 19 for plain infeasibility branching, 22 for violation transfer, 23 for strong branching, 20 and 24 for the two reliability variants. The reliability rule itself is T. Achterberg, T. Koch and A. Martin, "Branching rules revisited", Operations Research Letters 33 (2005), whose finding for integer programs was that reliability branching consistently beat the hybrid strong/pseudocost rule at every parameter setting and that most-infeasible branching was no better than random. BARON's violation transfer is Tawarmalani and Sahinidis (2004), Section 5.1. S. Vigerske and A. Gleixner, "SCIP: global optimization of mixed-integer nonlinear programs in a branch-and-cut framework", Optimization Methods and Software 33 (2018), Section 3.6.2, report that changing SCIP's spatial rule matters only on the quarter of their instances that branch on continuous variables at all.

Tolerances

No major solver stops at a relative gap of \(10^{-6}\) by default, and a global solver works to three tolerances, not one. The feasibility tolerance \(\varepsilon_f\) decides when a point counts as feasible and so when it may become an incumbent. The integrality tolerance decides when a relaxed integer variable counts as integral. The gap tolerance, absolute or relative, decides when the search stops. Proposition 3.5.5(a) shows that finite termination needs \(\varepsilon_f > 0\) even with exhaustive subdivision. A relaxed point need not satisfy a nonlinear equality exactly, and without a feasibility tolerance or a local NLP solve the incumbent would never update. The solvers that set the gap tolerance to zero terminate because their pruning compares bounds within an epsilon and accepts \(\varepsilon_f\)-feasible incumbents, not because the gap is literally zero.

solverfeasibilityintegralitygap: absolute / relativesource
BARONAbsConFeasTol 1e-6AbsIntFeasTol 1e-5EpsA 1e-6 / EpsR 1e-6; under GAMS: optCA 0 / optCR 1e-4standalone manual; GAMS/BARON docs
SCIPfeastol 1e-6feastol 1e-6limits/absgap 0 / limits/gap 0; a node is cut off within epsilon 1e-9set.c, def.h, tree.c
Couennefeas_tolerance 1e-5integer_tolerance 1e-6 (Bonmin)allowable_gap 0 / allowable_fraction_gap 0; cutoff_decr 1e-5couenne.opt, Bonmin
GurobiFeasibilityTol 1e-6IntFeasTol 1e-5MIPGapAbs 1e-10 / MIPGap 1e-4parameter reference
Three tolerances, four solvers (defaults; the relative gap's denominators are in the table of Section 2.1)

Two consequences follow for reading a log. Relative gaps are not comparable across solvers, because the denominators differ (Section 2.1) and so do the defaults. The benchmarks of Section 5 run every solver at the same relative gap, \(10^{-4}\), for that reason. And a bound that is accurate to the LP solver's dual tolerance, \(10^{-7}\) say, is compared with an incumbent that is itself only \(10^{-6}\)-feasible. Pruning with floating-point LP bounds is therefore unsound in principle and is made sound in practice by the epsilons above. The rigorous alternative, the safe bound of Section 7.3, costs one matrix–vector product per node and is what exact codes use.BARON: BARON User Manual (The Optimization Firm, HTML edition v. 2026.9.10, 10 September 2026), options EpsA, EpsR, AbsConFeasTol, AbsIntFeasTol, and the GAMS/BARON documentation, gams.com/latest/docs/S_BARON.html, where EpsA and EpsR inherit GAMS's optCA and optCR. SCIP: src/scip/set.c, def.h and tree.c of SCIP 10.0.0. Couenne: src/couenne.opt and Bonmin's BonBabSetupBase.cpp (COIN-OR, master). Gurobi: Gurobi Optimizer Reference Manual, version 13.0, parameters MIPGap, MIPGapAbs, FeasibilityTol, IntFeasTol. The benchmark setting is K. Bestuzheva, A. Chmiela, B. Müller, F. Serrano, S. Vigerske and F. Wegscheider, "Global optimization of mixed-integer nonlinear programs with SCIP 8", Journal of Global Optimization 91 (2025), Section 3.2 (optcr \(10^{-4}\), optca \(10^{-6}\)).

The bilinear example as a tree

The figure below runs Algorithm 3.5.8 on the running example R2 of Section 2.4, maximize \(xy\) over the unit square under \(2x + y \le 1.2\), whose answer is \(0.18\) at \((0.3, 0.6)\). This problem maximizes, so node bounds are upper bounds and a node is pruned when its bound is at most the incumbent plus the tolerance. The node relaxation is the McCormick LP on the node box. The relaxed point is always feasible, since the only constraint is linear, so its own product \(xy\) is a candidate incumbent at every node. Watch three things. The root bound is \(0.400\) at \((0.4, 0.4)\), where the true product is only \(0.16\). The second node, the left half \([0, 0.5] \times [0, 1]\), has bound \(0.3\) at the relaxed point \((0.3, 0.6)\), which happens to be the true maximizer. The incumbent is therefore \(0.18\) after two nodes, and the remaining eleven nodes only bring the bound down. And no split rule is cheapest at every tolerance: watch the three counts in the readout change order between tolerance \(0.01\) and \(0.0001\). In the terms of Definition 3.5.2 the run has best-bound selection, which is bound improving with \(K = 1\), a midpoint subdivision, exhaustive by Proposition 3.5.6 with \(\beta = 1/2\), and McCormick bounding, consistent by Proposition 3.5.5(a) with \(\alpha = 2\), so Corollary 3.5.4 promises that the search stops at every positive tolerance; what it does not predict is the count.

Spatial branch and bound on the bilinear example: maximize w = x·y over the unit square under 2x + y ≤ 1.2, whose answer is 0.18 at (0.3, 0.6). The left panel shows the square cut into boxes as the search proceeds, with the box being processed outlined in orange, open boxes dashed orange, boxes pruned within the tolerance hatched grey with their McCormick bound written inside, boxes with no feasible point hatched red, the incumbent as a green dot and the true maximum as a blue ring. The right panel plots the global bound (orange) and the incumbent (green) against nodes processed. The scrub replays the search node by node, the tolerance sets when a box is pruned, and the split control chooses where the wider side is cut: at the midpoint, at the relaxed point (pulled into the middle 80 % of the side), or at the blend 0.25·(relaxed point) + 0.75·(midpoint), the weighting of Couenne's default mid-point rule, pulled into the middle 60 % of the side as SCIP's branching/clamp does.

(The counts, and the grid for comparison) In the default view (tolerance \(0.01\), midpoint rule) the search processes 13 nodes: 6 are branched, 5 are pruned within the tolerance and 2 are infeasible because their boxes lie entirely beyond the line. The incumbent is \(0.180\) at \((0.3, 0.6)\), the final global bound is \(0.188\), the gap is \(0.008\) and no box is open. At tolerance \(0.001\) the same rule needs 21 nodes and ends with bound \(0.1807\), at \(0.0001\) it needs 27. The readout's table gives the other two rules: at tolerance \(0.01\) the relaxed-point rule needs 9 nodes and the blend 15, at \(0.001\) they need 25 and 19, and at \(0.0001\) they need 31 and 29. The side panel shows the orange bound curve falling from \(0.400\) and entering the green band one tolerance above the incumbent, which is the stopping event. For comparison, the figure also reports a uniform \(k \times k\) partition of the unit square with the McCormick LP solved on every cell, the piecewise relaxation of Proposition 2.4.8, whose widest band is \(1/(2k^2)\). The grid needs \(k = 5\), that is 25 boxes, to bring its bound to within \(0.01\) of \(0.18\), and \(k = 10\), 100 boxes, for \(0.001\). At \(k = 10\) the grid's bound is exact because \(0.3\) and \(0.6\) are grid lines and the McCormick planes are exact on the edges of a box (Theorem 2.4.7). The search subdivides only boxes whose bound is above the incumbent, which is why it needs fewer boxes than the grid. The figure's blend rule splits the root on \(x\) at \(0.475\), as computed above. Its clamp never binds in any of the 32 blend splits across the four tolerances, so the figure's counts do not depend on which solver's clamp is used. The figure's rules are Couenne's weighting and SCIP's clamp, not either solver's exact rule. Couenne's mid-point rule clamps at 5 percent of the width (closeToBounds in its source), not at 20 percent, and raises its weight on the relaxed point once the relative gap falls below \(10^{-3}\). SCIP shrinks its midpull once a side falls below half its original width. Neither solver would therefore produce the figure's blend tree exactly.

The R2 search at tolerance 0.01, midpoint rule: its first split

                root: [0, 1] x [0, 1]
                McCormick bound 0.400 at (0.4, 0.4),
                where x*y = 0.16
           x <= 0.5 /               \ x >= 0.5
                   /                 \
   node 2: [0, 0.5] x [0, 1]         [0.5, 1] x [0, 1]
   bound 0.3 at (0.3, 0.6),
   where x*y = 0.18: the true
   maximizer, so the incumbent
   is 0.18 after two nodes

  the other 11 nodes only bring the bound down: 13 nodes in all,
  6 branched, 5 pruned within the tolerance, 2 infeasible (boxes
  beyond the line 2x + y = 1.2); final bound 0.188, gap 0.008

The following program is the search the figure runs, with the three split rules. The McCormick LP has three variables, so it is solved exactly by enumerating basic solutions. It prints the node counts quoted above.

# Spatial branch and bound on R2, the search that fig-spatial runs.
#
# Maximize x*y s.t. 2x + y <= 1.2, (x, y) in [0, 1]^2. Node relaxation:
# the McCormick LP on the node box, solved exactly by vertex enumeration.
# Best-bound selection, children inherit the parent's bound, prune when
# bound <= incumbent + tol. Split rules: midpoint; relaxed point (pulled
# into the middle 80% of the side); blend 0.25*relaxed + 0.75*midpoint
# (Couenne's weighting), clamped to the middle 60% (SCIP's clamp).

import heapq
import itertools

import numpy as np

def lp_max(c, A, b):
    """max c.v s.t. A v <= b for 3 variables.

    Solved exactly by enumerating basic solutions.
    """
    best, arg = -np.inf, None
    for rows in itertools.combinations(range(len(b)), 3):
        M = A[list(rows)]
        if abs(np.linalg.det(M)) < 1e-12:
            continue
        v = np.linalg.solve(M, b[list(rows)])
        if np.all(A @ v <= b + 1e-9) and c @ v > best + 1e-12:
            best, arg = c @ v, v
    return best, arg

def relax(box):
    """McCormick LP for max w on box = (xl, xu, yl, yu).

    The variables are v = (x, y, w); the rows are A v <= b.
    """
    xl, xu, yl, yu = box
    A = np.array([[ yl,  xl, -1.0],     # w >= xl*y + x*yl - xl*yl
                  [ yu,  xu, -1.0],     # w >= xu*y + x*yu - xu*yu
                  [-yl, -xu,  1.0],     # w <= xu*y + x*yl - xu*yl
                  [-yu, -xl,  1.0],     # w <= xl*y + x*yu - xl*yu
                  [2.0, 1.0, 0.0],      # 2x + y <= 1.2
                  [1, 0, 0], [-1, 0, 0], [0, 1, 0], [0, -1, 0]], float)
    b = np.array([xl*yl, xu*yu, -xu*yl, -xl*yu, 1.2, xu, -xl, yu, -yl])
    return lp_max(np.array([0.0, 0.0, 1.0]), A, b)

def split_point(lo, hi, rel, rule):
    """Where to cut the side [lo, hi], given the relaxed value rel."""
    w, mid = hi - lo, 0.5 * (lo + hi)
    if rule == "midpoint":
        return mid
    if rule == "relaxed":
        return min(max(rel, lo + 0.1 * w), hi - 0.1 * w)
    # blend
    return min(max(0.25 * rel + 0.75 * mid, lo + 0.2 * w), hi - 0.2 * w)

def sbb(tol, rule, cap=400):
    """Best-bound search to tolerance tol; returns (nodes, incumbent)."""
    inc, nodes, counter = -np.inf, 0, itertools.count()
    # heap entries: (-inherited bound, age, box)
    heap = [(-np.inf, next(counter), (0.0, 1.0, 0.0, 1.0))]
    while heap and nodes < cap:
        negpb, _, box = heapq.heappop(heap)
        # the inherited bound cannot beat the incumbent
        if -negpb <= inc + tol:
            continue

        ub, v = relax(box)
        nodes += 1
        # the box misses the line 2x + y <= 1.2
        if v is None:
            continue
        # the relaxed point is feasible: an incumbent
        inc = max(inc, v[0] * v[1])
        if ub <= inc + tol:
            continue                    # pruned within the tolerance

        # split the wider side (ties go to x)
        xl, xu, yl, yu = box
        if xu - xl >= yu - yl:
            m = split_point(xl, xu, v[0], rule)
            kids = [(xl, m, yl, yu), (m, xu, yl, yu)]
        else:
            m = split_point(yl, yu, v[1], rule)
            kids = [(xl, xu, yl, m), (xl, xu, m, yu)]
        for kid in kids:
            heapq.heappush(heap, (-ub, next(counter), kid))
    return nodes, inc

print("nodes processed (LPs solved) to prove  max x*y = 0.18  within tol:")
print(" rule       tol=0.1  tol=0.01  tol=0.001  tol=0.0001"
      "   incumbent at 0.0001")
for rule in ("midpoint", "relaxed", "blend"):
    counts = [sbb(t, rule) for t in (1e-1, 1e-2, 1e-3, 1e-4)]
    print(f" {rule:10s} {counts[0][0]:5d}    {counts[1][0]:5d}"
          f"     {counts[2][0]:5d}      {counts[3][0]:5d}"
          f"       {counts[3][1]:.6f}")
nodes processed (LPs solved) to prove  max x*y = 0.18  within tol:
 rule       tol=0.1  tol=0.01  tol=0.001  tol=0.0001   incumbent at 0.0001
 midpoint       5       13        21         27       0.180000
 relaxed        3        9        25         31       0.179998
 blend          5       15        19         29       0.179994

Each node costs one LP of nine rows, which the program solves by trying the \(\binom{9}{3} = 84\) bases. The counts are the figure's. What parallelizes here is small, because the depth-first and best-bound orders of this tiny tree never hold more than a handful of open boxes at once, and the next part says what a wide frontier looks like.

(What each device is worth) Three further experiments on the same program are instructive, and their results are in the table below. Reduced-cost tightening (step 7 of Algorithm 3.5.8) never fires on this instance, because at every surviving node the relaxed point is interior in the branched coordinate, so the multipliers of the bound rows are zero. A local search from the relaxed point does not help, because node 2's relaxed point already is the optimum. FBBT with the incumbent cutoff changes everything. After the root, the constraint \(xy \ge 0.16\) propagates to \(x \ge 0.16\) and \(y \ge 0.16\), then through \(2x + y \le 1.2\) to \(x \le 0.52\) and \(y \le 0.88\), and so on toward the fixed point \(x \in [0.2, 0.4]\), \(y \in [0.4, 0.8]\) computed in Section 2.6. The child \(x \ge 0.5\) is proved empty without an LP. The tree collapses to two to four LPs at every tolerance.The variants were run with a numpy script written for this series (sbb_variants.py, 5 October 2026), which adds FBBT on the two constraints, reduced-cost tightening from the LP duals and a one-dimensional local search to the program above.

varianttol 1e-2tol 1e-3tol 1e-4 
plain132127 
+ reduced-cost tightening132127(never fires: no bound is active)
+ local search from the relaxed point132127(node 2 already is the optimum)
+ FBBT with the incumbent cutoff244(\(x \ge 0.5\) proved empty without an LP)
+ all three224 
What each device is worth on the bilinear example (LPs solved; midpoint branching, best bound first)
FBBT with the incumbent cutoff x*y >= 0.16 on R2, after the root

        x                     y
     [0, 1]                [0, 1]          the root box
        |                     |
        v                     v
     x >= 0.16             y >= 0.16       from x*y >= 0.16
        |                     |
        v                     v
     x <= 0.52             y <= 0.88       through 2x + y <= 1.2
        :                     :
        v                     v
     [0.2, 0.4]            [0.4, 0.8]      the fixed point
                                           (Section 2.6)

  the child x >= 0.5 is then proved empty without an LP, and the
  tree collapses to two to four LPs at every tolerance

This is the mechanism behind two measured numbers that Section 2.6 reports. Turning off BARON's range reduction multiplied its node counts by a factor of nearly thirteen on one test library, and turning off SCIP's domain propagation increased its node counts by 86 percent. Domain reduction, not the choice of branching rule, determines the node count on this example.

The exponent, and reduced-space formulations

Proposition 3.5.5(b) says the tree is exponential in the number of coordinates that must be split. The scaling is visible on copies of the running example. Take \(d\) independent copies with objective \(\sum_i x_i y_i\), whose optimum is \(0.18\,d\). The McCormick LP separates over the copies, so each node's bound is a sum of \(d\) three-variable LPs, but the tree does not separate, since it branches on one of the \(2d\) variables at a time.

\(d\)tol 1e-2tol 1e-3optimum \(0.18\,d\)
113210.18
2631230.36
32496410.54
4129328490.72
Tree size against the number of nonconvex copies (midpoint branching, best bound first, absolute tolerance)

Each additional copy multiplies the node count by roughly four to five at \(\varepsilon = 10^{-2}\) and by four and a half to six at \(10^{-3}\), the per-copy tolerance tightening with \(d\) because the absolute tolerance is shared.The counts are from a numpy script run for this series (sbb_scaling.py, 5 October 2026, about 17 s), the program above with \(d\) copies of the McCormick rows. A separable structure that a decomposition would exploit, which is the subject of Section 3.7, is invisible to a monolithic tree.

(Reduced-space formulations) The exponent can be lowered by changing what is branched on. In the auxiliary-variable formulation that BARON, SCIP and Couenne use, every intermediate node of the factorable DAG becomes a variable with its own bounds. The relaxation is an LP in all of them, and the LP's bounds on intermediate quantities feed the domain reduction. In a reduced-space formulation only the original variables that enter nonconvex terms are branched on. The relaxation is evaluated by propagating McCormick relaxations through the DAG at a point (Theorem 2.4.12 and Algorithm 2.4.15) rather than by solving an LP in lifted variables. The tree then lives in a space of dimension equal to the number of nonconvex variables rather than the number of DAG nodes, at the price of weaker relaxations and no bounds on the intermediates. Epperly and Pistikopoulos proposed reduced-space branch and bound in 1997. MAiNGO is the solver built on it, with large time reductions reported for flowsheet problems in which most variables are determined by the few that enter the nonconvex terms.T. G. W. Epperly and E. N. Pistikopoulos, "A reduced space branch and bound algorithm for global optimization", Journal of Global Optimization 11 (1997). D. Bongartz and A. Mitsos, "Deterministic global optimization of process flowsheets in a reduced space using McCormick relaxations", Journal of Global Optimization 69 (2017). The solver: D. Bongartz, J. Najman, S. Sass and A. Mitsos, "MAiNGO – McCormick-based Algorithm for mixed-integer Nonlinear Global Optimization", technical report, Process Systems Engineering (AVT.SVT), RWTH Aachen University (2018), permalink.avt.rwth-aachen.de/?id=729717. The GPU evaluation of pointwise McCormick relaxations on thousands of boxes at once is R. X. Gottlieb, P. Xu and M. D. Stuber, "Automatic source code generation for deterministic global optimization with parallel architectures", Optimization Methods and Software 41 (2026), with the timings quoted in Section 2.4 (the authors' numbers, from the README of their package SourceCodeMcCormick.jl). It is the one relaxation technology other than interval arithmetic that has been run on a GPU inside a global solver; MAiNGO's GPU interval bounder is the other: H. Zhang, T. Kerkenhoff, N. Kichler, M. Dahmen, A. Mitsos, U. Naumann and D. Bongartz, "Accelerating deterministic global optimization via GPU-parallel interval arithmetic", arXiv 2507.20769 (2025). For a GPU the trade is attractive in one direction and dangerous in the other. A reduced-space relaxation is a straight-line program per box with identical control flow, which is the shape a kernel, one function executed by thousands of device threads at once (Section 7.1), wants. But the relaxation it evaluates is weaker, and Theorem 3.5.7 says that a weaker bound near the optimum costs a larger frontier. Nobody has measured where the balance lies.

st_e13: a gap that integer branching cannot close

The running example R3 is the MINLPLib instance st_e13,

\[z^\star \;=\; \min\{\, 2x + y \;:\; g(x, y) = 1.25 - x^2 - y \le 0,\ \ x + y \le 1.6,\ \ 0 \le x \le 1.6,\ \ y \in \{0, 1\} \,\}. \tag{3.5.1}\]

The library lists it as a two-variable MBQCP, a mixed-binary quadratically constrained program in MINLPLib's classification, with one binary, one linear and one quadratic constraint of concave curvature. Its page reports the primal bound \(2\) and the dual bound \(2\) from ANTIGONE, BARON, Couenne, Gurobi, LINDO and SCIP, and gives the source as "BARON book instance misc/e13".MINLPLib, instance page st_e13, minlplib.org/st_e13.html (read 5 October 2026; the page reports primal bounds at infeasibility at most \(10^{-8}\) and the dual bound per solver; SHOT, a convex-MINLP code, reports the trivial dual bound \(0\)). The page attributes the instance to M. Tawarmalani and N. V. Sahinidis, Convexification and Global Optimization in Continuous and Mixed-Integer Nonlinear Programming (Kluwer, 2002), and lists G. R. Kocis and I. E. Grossmann, "Global optimization of nonconvex mixed-integer nonlinear programming (MINLP) problems in process synthesis", Industrial & Engineering Chemistry Research 27 (1988), as a second reference. Whether the instance appears in that paper was not checked. The library itself: M. R. Bussieck, A. S. Drud and A. Meeraus, "MINLPLib—a collection of test models for mixed-integer nonlinear programming", INFORMS Journal on Computing 15 (2003). The nonlinear constraint is \(y \ge 1.25 - x^2\): the feasible side lies above a downward parabola, so \(g\) is concave and \(\{g \le 0\}\) is the complement of an open convex set. This is the simplest nonconvexity a factorable solver meets, a single square, and it is enough to defeat every method of Section 3.4. Section 3.4 showed outer approximation stopping on this instance at \(2.236\) with a closed gap and a wrong answer. The tangent of a concave function lies above it, so the "cut" is a restriction. Here the point is the other half: a method that branches only on \(y\) cannot close the gap either, and one spatial branch on \(x\) can.

Proposition 3.5.10 (the instance). (i) The feasible set of (3.5.1) is the union of the two segments \(S_0 = \{(x, 0) : \sqrt{1.25} \le x \le 1.6\}\) and \(S_1 = \{(x, 1) : 0.5 \le x \le 0.6\}\). (ii) \(z^\star = 2\), attained only at \((0.5, 1)\). The best point with \(y = 0\) is \((\sqrt{1.25}, 0) = (1.1180, 0)\) with value \(\sqrt 5 = 2.2361\). (iii) The continuous relaxation, \(y \in [0, 1]\) with the exact constraint kept, has optimal value \(2 = z^\star\), so the integrality gap is zero, and it has exactly two local minimizers, \((0.5, 1)\) with value \(2\) and \((1.1180, 0)\) with value \(2.2361\).

Proof. (i) With \(y = 0\) the nonlinear constraint reads \(x^2 \ge 1.25\), hence \(x \ge \sqrt{1.25}\), and the linear constraint gives \(x \le 1.6\). With \(y = 1\) it reads \(x^2 \ge 0.25\), hence \(x \ge 0.5\), and the linear constraint gives \(x \le 0.6\). (ii) On each segment \(2x + y\) increases with \(x\), so the minimum is at the left end. (iii) For fixed \(x\) the objective is smallest at the smallest feasible \(y\), namely \(y = \max(0, 1.25 - x^2)\), which is admissible exactly when \(x \ge 0.5\) (the second condition \(1.25 - x^2 \le 1.6 - x\) holds for every \(x\), since \(x^2 - x + 0.35\) has negative discriminant). The relaxation is therefore \(\min \varphi\) over \([0.5, 1.6]\) with \(\varphi(x) = 2x + 1.25 - x^2\) for \(x \le \sqrt{1.25}\) and \(\varphi(x) = 2x\) beyond. The first piece has \(\varphi' = 2 - 2x\), positive on \((0.5, 1)\) and negative on \((1, \sqrt{1.25})\), and the second has \(\varphi' = 2\). So \(\varphi\) rises from \(\varphi(0.5) = 2\) to \(\varphi(1) = 2.25\), falls to \(\varphi(\sqrt{1.25}) = \sqrt 5\) and rises again. ∎

Part (iii) is Section 1.3 on this instance: a local NLP solver started at \((1.6, 0)\) converges to the KKT point \((1.1180, 0)\) and returns \(2.2361\), an upper bound on \(z^\star\) and nothing more. Used as a root bound, as NLP-based branch and bound with a local solver would use it, it exceeds \(z^\star\) by \(0.2361\). The convex case of Section 3.4 needed convexity exactly so that the local solver's value is the global one.

Proposition 3.5.11 (the chord relaxation and the root bound). On \([a, b]\) with \(a < b\) the chord \(c(x) = (a + b)x - ab\) is the concave envelope of \(x^2\), with \(c(x) - x^2 = (x - a)(b - x) \le \tfrac14 (b - a)^2\), equality on the left only at the ends and on the right only at the midpoint. Consequently, on the node box \(B = [a, b] \times [y_l, y_u]\) the convex envelope of \(g\) is \(\operatorname{vex}_B g(x, y) = 1.25 + ab - (a + b)x - y\), and the chord relaxation of the node is the linear program

\[\bar z(B) \;=\; \min\{\, 2x + y \;:\; (a + b)\,x + y \ge 1.25 + ab,\ \ x + y \le 1.6,\ \ a \le x \le b,\ \ y_l \le y \le y_u \,\}, \tag{3.5.2}\]

a valid lower bound on the node. At the root (\(a = 0\), \(b = 1.6\), \(y \in [0, 1]\)) the relaxed constraint is \(y \ge 1.25 - 1.6x\), the feasible set of (3.5.2) is the quadrilateral with vertices \((0.78125, 0)\), \((1.6, 0)\), \((0.6, 1)\) and \((0.15625, 1)\), and \(\bar z(B_0) = 1.3125\) at \((0.15625, 1)\), a point with \(g = 0.2256 > 0\).

Proof. \((a + b)x - ab - x^2 = (x - a)(b - x)\) is nonnegative on \([a, b]\), zero exactly at the ends, and the product of two numbers with fixed sum \(b - a\) is largest when they are equal. The chord is affine, hence concave, and lies above \(x^2\). Any concave \(h \ge x^2\) on \([a, b]\) satisfies \(h((1-t)a + tb) \ge (1-t)a^2 + tb^2 = c(x)\) by concavity, so \(c\) is the least concave overestimator. Replacing \(-x^2\) by \(-c(x)\) in \(g\) gives the convex envelope because the other terms are affine, and \(g \ge \operatorname{vex}_B g\) makes \(\{g \le 0\} \subseteq \{\operatorname{vex}_B g \le 0\}\). At the root every point of the quadrilateral satisfies \(2x + y \ge 2x + (1.25 - 1.6x) = 1.25 + 0.4x\), and \(y \le 1\) with the relaxed constraint forces \(1.6x \ge 0.25\), so \(2x + y \ge 1.25 + 0.0625 = 1.3125\), with equality at \((0.15625, 1)\). ∎

(Why the root bound is loose) The root bound \(1.3125\) is \(0.6875\) below the optimum. It is not the best bound the geometry allows. The convex hull of the feasible set has lower edge \(y \ge 1.8090 - 1.6180x\), the line through \((0.5, 1)\) and \((1.1180, 0)\). The chord relaxation uses \(y \ge 1.25 - 1.6x\) instead, because the chord spans the whole range \([0, 1.6]\) although no feasible point has \(x < 0.5\). Domain reduction is what removes that difference, and the exact-propagation variant at the end of this st_e13 discussion closes the instance in three nodes.

Proposition 3.5.12 (integer branching cannot close the gap, and the first spatial branch can). (i) Fixing \(y = 1\) on the box \([0, 1.6]\) leaves the chord bound at \(1.3125\). After the linear row shrinks the box to \([0, 0.6]\) the chord is \(x^2 \le 0.6x\) and the bound is \(1.8333\) at \(x = 0.41667\). Fixing \(y = 0\) gives \(1.5625\) at \(x = 0.78125\). In each case the relaxed point violates \(g\), by \(0.2256\), \(0.0764\) and \(0.6396\), and no integer variable is left to branch on. (ii) At the node \(y = 1\) on \([0, 1.6]\), splitting \(x\) at the root's relaxed value \(0.15625\) gives a left child \([0, 0.15625]\) whose relaxed constraint \(y \ge 1.25 - 0.15625x \ge 1.2256 > 1\) is infeasible, and a right child \([0.15625, 1.6]\) with bound \(1.5694\). Splitting instead at the incumbent's \(x = 0.5\) gives two children with bound exactly \(2\): on \([0, 0.5]\) the chord \(x^2 \le 0.5x\) forces \(x \ge 0.5\), and on \([0.5, 1.6]\) the chord \(x^2 \le 2.1x - 0.8\) forces \(x \ge 0.5\).

Proof. (i) is (3.5.2) with \(y_l = y_u\), which makes the program one-dimensional with minimizer \(\tilde x = (1.25 + ab - y)/(a + b)\): \(0.25/1.6 = 0.15625\), \(0.25/0.6 = 0.41667\) and \(1.25/1.6 = 0.78125\). (ii) On \([0, 0.15625]\) with \(y = 1\) the chord row reads \(0.15625\,x \ge 0.25\), impossible. The other three bounds are the same formula with \(a = 0.15625\), with \(b = 0.5\), and with \(a = 0.5\). The last two are exact because the chord agrees with \(x^2\) at the endpoint \(0.5\), which is the optimizer. ∎

Proposition 3.5.12 on st_e13 (z* = 2): branching on y, then on x

               root: x in [0, 1.6], y in [0, 1]
               bound 1.3125 at (0.15625, 1)
              y = 1 /                      \ y = 0
                   /                        \
   x in [0, 1.6]: bound 1.3125          bound 1.5625 at x = 0.78125
   (g violated by 0.2256)               (g violated by 0.6396)
   after the linear row, x in [0, 0.6]:
   bound 1.8333 at x = 0.41667
   (g violated by 0.0764)
   no integer variable is left to branch on
        |
        |  splitting x at the node y = 1, x in [0, 1.6]:
        |
        +-- at 0.15625         [0, 0.15625]     infeasible
        |   (the root's x)     [0.15625, 1.6]   bound 1.5694
        |
        +-- at 0.5             [0, 0.5]         bound 2
            (the incumbent's)  [0.5, 1.6]       bound 2 = z*

(A nonconvexity gap, and the four methods in the figure) So the whole gap on this instance is a nonconvexity gap. The integrality gap is zero, fixing the binary variable does not move the bound, and the bound moves only when the interval of \(x\) shrinks, because the chord moves toward the parabola as its endpoints move. The figure below shows the four methods on the instance. The enumerate mode solves the two NLPs and reports \(2.0\) and \(2.236\). The outer-approximation mode takes the tangent at \((1.118, 0)\), the red cut \(y \ge 2.5 - 2.236x\). At the optimum the cut reads \(2.236 \cdot 0.5 + 1 = 2.118 < 2.5\), so the optimum is cut off, the column \(y = 1\) becomes infeasible, and the master returns \(2.236\) equal to the incumbent. The method stops with a false proof, and the side panel shows its "bound" sitting above the true optimum, which a valid bound cannot do. The envelope mode draws the chord relaxation and its root bound \(1.3125\) at \((0.15625, 1)\). The branch mode runs a spatial branch and bound with chords, closed-form NLPs with \(y\) fixed, and a choice of how many bound-tightening rounds precede a spatial branch.

MINLPLib st_e13: minimize 2x + y with y ∈ {0, 1}, x + y ≤ 1.6 and y ≥ 1.25 − x², a constraint whose feasible side lies above a concave curve, so the feasible set is the two blue segments, the blue dot is the optimum 2.0 at (0.5, 1) and the ring is the other integer solution 2.236. Choose a method and scrub the step: enumeration solves two NLPs, outer approximation takes one tangent cut (red, because the tangent of a concave function lies above it) and stops at 2.236 with a false proof, the chord relaxation gives the root bound 1.3125, and spatial branch and bound with chords and one round of bound tightening proves 2.0 with five nodes and seven relaxations. The lower panel plots the bound and the incumbent against problems solved.

(The figure's branch-and-bound run) The branch mode is Algorithm 3.5.8 specialized to the instance. It searches depth first, since the tree is finite and bound-improving selection is not needed. It propagates the linear row only, uses the chord LP as the relaxation and the closed-form NLP with \(y\) fixed as the local search, and prunes at \(\bar z(N) \ge z_{\mathrm{inc}} - 10^{-5}\). Its tightening step is the interval propagation of the chord row itself, \(x \ge (1.25 + ab - y_u)/(a + b)\), followed by a rebuilt chord on the smaller box. The spatial branch is at the incumbent's \(x\) when that lies strictly inside the box, otherwise at the relaxed point, otherwise at the midpoint. With the default rule, one tightening round before a spatial branch, the run is this.

nodebox of \(x\)chord \(x^2 \le s\,x - p\)\(x \ge\)boundrelaxed pointNLPstatus
1\([0, 1.6]\)\(x^2 \le 1.6x\)0.156251.3125\((0.15625, 1)\) branch on \(y\)
2\([0, 0.6]\), \(y = 1\)\(x^2 \le 0.6x\)0.4171.833\((0.417, 1)\) tightened: \(x \ge 0.41667\)
 \([0.417, 0.6]\)\(x^2 \le 1.017x - 0.25\)0.4921.984\((0.492, 1)\)2.0branch on \(x\) at 0.5
3\([0.417, 0.5]\), \(y = 1\)\(x^2 \le 0.917x - 0.208\)0.52.0\((0.5, 1)\) pruned
4\([0.5, 0.6]\), \(y = 1\)\(x^2 \le 1.1x - 0.3\)0.52.0\((0.5, 1)\) pruned
5\([0, 1.6]\), \(y = 0\)\(x^2 \le 1.6x\)0.781251.5625\((0.78125, 0)\) tightened: \(x \ge 0.78125\)
 \([0.78125, 1.6]\)\(x^2 \le 2.38125x - 1.25\)1.0502.100\((1.050, 0)\)2.236pruned
 totals     5 nodes, 7 relaxations, 2 NLPs, 2 tightening rounds; incumbent 2.0 at \((0.5, 1)\); least leaf bound 2.0
The default st_e13 run as a tree (depth first, one tightening round)

                    (1) x in [0, 1.6]: bound 1.3125
                        branch on y
                   y = 1 /           \ y = 0
                        /             \
   (2) [0, 0.6]: 1.833                 (5) [0, 1.6]: 1.5625
       NLP: incumbent 2.0 at (0.5, 1)      NLP 2.236
       tightened: x >= 0.41667             tightened: x >= 0.78125
       [0.417, 0.6]: 1.984                 [0.78125, 1.6]: 2.100
       branch on x at 0.5                  2.100 >= 2.0: pruned
      x <= 0.5 /      \ x >= 0.5
              /        \
   (3) [0.417, 0.5]    (4) [0.5, 0.6]
       bound 2.0:          bound 2.0:
       pruned              pruned

  5 nodes, 7 relaxations, 2 NLPs, 2 tightening rounds

(Two moments of the run) Two moments of the run matter. At node 2 the NLP with \(y = 1\) fixed finds the incumbent \(2.0\) at \(x = 0.5\), and after one tightening round the node splits at the incumbent's \(x\), where both children have bound exactly \(2.0\) by Proposition 3.5.12(ii) and are pruned. At node 5 the chord bound \(1.5625\) is below the incumbent, and the node is pruned only because one tightening round moves the chord: \(x \ge 0.78125\) gives the re-chord \(x^2 \le 2.38125x - 1.25\), which forces \(x \ge 2.5/2.38125 = 1.050\) and the bound \(2.100 \ge 2.0\). The orange curve of the lower panel stays at \(1.3125\) while the \(y = 1\) subtree is worked, because the open node \(y = 0\) still carries the inherited root bound. It rises to \(1.5625\) when node 5 is relaxed and meets the green incumbent line at \(2.0\) when node 5 is pruned. The figure's totals, which the prose quotes as the page displays them, are five nodes, seven relaxations, two NLPs and two tightening rounds. A count by a different convention, which skips node 2's own relaxation and its re-chord, since the incumbent decides them, and which counts the two implied bounds at the \(y = 0\) node as two rounds, gives five relaxations and two tightening rounds. Both are correct descriptions of the same run and both prove the same optimum.

The tightening control compares three rules, and the comparison is the figure's own lesson.

rulenodesrelaxationsNLPstightening roundsspatial branches
none7720two (at 0.5 and at 0.78125)
one round (default)5722one (at 0.5)
to a fixed point3926none
Tightening before a spatial branch, on st_e13

(Tightening trades relaxations for nodes) Without tightening the \(y = 0\) node splits at its relaxed point \(0.78125\) into an infeasible left child and a right child with the same bound \(2.100\), so the arithmetic of the tightening round reappears as one more node. With tightening to a fixed point the \(y = 1\) node never branches: its lower end moves \(0 \to 0.41667 \to 0.49180 \to 0.49925 \to 0.49993 \to 0.49999\), its bound moves \(1.8333 \to 1.9836 \to 1.9985 \to 1.99986 \to 1.99999 \to 1.9999989\), the last value is within \(10^{-5}\) of the incumbent, and the node is pruned after five rounds and six relaxations. Tightening trades relaxations for nodes, \(7/7 \to 7/5 \to 9/3\). The convergence of the iteration is geometric, and its rate can be computed.

Proposition 3.5.13 (the tightening iteration converges linearly to the exact bound). Fix \(y\), put \(c = 1.25 - y > 0\) and \(x^\star = \sqrt c\), and let \(b > \sqrt c\) and \(a_0 \in [0, \sqrt c)\). The tightening step is \(a_{k+1} = (c + a_k b)/(a_k + b)\). Then \(a_k\) increases strictly to \(\sqrt c\) without reaching it, and

\[\sqrt c - a_{k+1} \;=\; (\sqrt c - a_k)\, \frac{b - \sqrt c}{a_k + b}, \qquad \frac{\sqrt c - a_{k+1}}{\sqrt c - a_k} \;\to\; \frac{b - \sqrt c}{b + \sqrt c} .\]

For the node \(y = 1\) (\(c = 0.25\), \(b = 0.6\)) the asymptotic rate is \(0.1/1.1 = 1/11\). For the node \(y = 0\) (\(c = 1.25\), \(b = 1.6\)) it is \(0.1773\).

Proof. \(\sqrt c - a_{k+1} = [\sqrt c\, a_k + \sqrt c\, b - c - a_k b]/(a_k + b) = (\sqrt c - a_k)(b - \sqrt c)/(a_k + b)\). The factor lies in \((0, 1)\) because \(0 < b - \sqrt c < b \le a_k + b\), so the error decreases geometrically and stays positive, and \(a_{k+1} - a_k = (c - a_k^2)/(a_k + b) > 0\) while \(a_k < \sqrt c\). As \(a_k \to \sqrt c\) the factor tends to \((b - \sqrt c)/(b + \sqrt c)\). ∎

(The rate 1/11, and what a fixed point means) The figure's readout prints the moves of the lower end at the \(y = 1\) node, \(0.41667, 0.07514, 0.00745, 0.00068, 0.00006\), whose successive ratios approach \(1/11\). The fixed point of the chord iteration is the exact bound \(x \ge \sqrt{1.25 - y_u}\), the greatest fixed point of propagation in the sense of Section 2.6, approached only in the limit. Section 2.6 showed two linear rows on which the width halves every round. This is a second concrete rate, \(1/11\), from a nonlinear row relaxed by its chord. "Tightening to a fixed point" in a solver therefore means "until the box moves by less than a threshold, or a round cap is hit, or the bound reaches the cutoff", and here the last test fires first.

The chord iteration at the st_e13 node y = 1 (Proposition 3.5.13)

  +--> box [a_k, 0.6]: chord x^2 <= (a_k + 0.6) x - 0.6 a_k
  |       |
  |       v   y >= 1.25 - x^2 relaxed by the chord, at y = 1
  |    x >= a_(k+1) = (0.25 + 0.6 a_k) / (a_k + 0.6)
  |    bound 2 a_(k+1) + 1
  |       |
  |       +-- bound >= 2.0 - 1e-5 ? --yes--> pruned
  |       | no
  +-------+   one tightening round: the box becomes [a_(k+1), 0.6]

  relaxation 1        2        3        4        5        6
  box from   0        0.41667  0.49180  0.49925  0.49993  0.49999
  bound      1.8333   1.9836   1.9985   1.99986  1.99999  1.9999989
                                                          pruned

  the lower end climbs toward x* = sqrt(0.25) = 0.5 and never
  reaches it; its moves 0.41667, 0.07514, 0.00745, 0.00068,
  0.00006 shrink by ratios that tend to 0.1/1.1 = 1/11

(Exact propagation closes the instance in three nodes) The figure propagates only the linear row and the chord row, on purpose, so that the chord iteration can be seen. A solver propagates the quadratic constraint itself. The backward step through the square node of the expression graph, applied to \(y \ge 1.25 - x^2\) with \(y \le y_u\) and \(x \ge 0\), gives \(x \ge \sqrt{1.25 - y_u}\) in one step. Inserting that step before every relaxation changes the search completely. At the root \(y_u = 1\) gives \(x \ge 0.5\), the chord on \([0.5, 1.6]\) is \(x^2 \le 2.1x - 0.8\), and the root bound is \(1.9524\) at \((0.97619, 0)\) instead of \(1.3125\). At the child \(y = 0\) the propagation gives \(x \ge 1.1180\) and the relaxed point \((1.1180, 0)\) is feasible: the node is exact, incumbent \(2.2361\). At the child \(y = 1\) the box is \([0.5, 0.6]\), the relaxed point is \((0.5, 1)\) and feasible: exact, incumbent \(2.0\). Three nodes, three relaxations, no NLP and no tightening round. The reason is Proposition 3.5.12(ii) read backwards. After exact propagation the lower end of the box is \(x^\star\) itself, the chord is exact at the one point that matters, and every node with \(y\) fixed closes by its own relaxation. This is what Couenne's FBBT on the expression graph, SCIP's nonlinear propagators and BARON's nonlinear feasibility-based tightening do as a matter of course. A reader who knows solvers should take the figure's seven relaxations as a statement about the figure's deliberately weak propagation, not about the instance.All numbers of the st_e13 example are printed by a numpy script run for this series (ex_e13.py, Python 3.11.4, 5 October 2026), which reproduces the arithmetic of the figure's module e13.js: the two NLPs, the outer-approximation master, the root quadrilateral, the three branch-and-bound runs with the counts 7/7/2/0, 5/7/2/2 and 3/9/2/6, the exact-propagation variant 3/3/0/0 with root bound \(1.952381\), the tightening ratios tending to \(1/11\) and \(0.177322\), and the frontier table below. Belotti et al. (2009), Section 3.2, describe the expression-graph propagation that Couenne runs, and Bestuzheva et al. (2025), Section 2.3.1, SCIP's quadratic propagator.

st_e13 with exact propagation of the square: 3 nodes, 3 relaxations

                root: y_u = 1 gives x >= 0.5;
                chord on [0.5, 1.6]: x^2 <= 2.1x - 0.8;
                bound 1.9524 at (0.97619, 0), not 1.3125
               y = 0 /                    \ y = 1
                    /                      \
   x >= 1.1180; the relaxed           box [0.5, 0.6]; the relaxed
   point (1.1180, 0) is feasible:     point (0.5, 1) is feasible:
   exact, incumbent 2.2361            exact, incumbent 2.0

  no NLP and no tightening round: every node with y fixed closes
  by its own relaxation

A frontier of nodes is one array operation

After the branch on \(y\) every node of the instance is one-dimensional: a box \([a, b]\) on a row \(y\). Its chord bound is a single closed-form expression,

\[\bar z(a, b, y) \;=\; \begin{cases} 2 \max\!\Big(a,\ \dfrac{1.25 + ab - y}{a + b}\Big) + y, & \text{if } \max\!\Big(a,\ \dfrac{1.25 + ab - y}{a + b}\Big) \le \min(b,\ 1.6 - y), \\[8pt] +\infty & \text{otherwise}, \end{cases} \tag{3.5.3}\]

and the NLP with \(y\) fixed is \(2\max(a, \sqrt{1.25 - y}) + y\) when feasible. A frontier of \(N\) such nodes is therefore one array operation over three vectors of length \(N\): no pivoting, no factorization, no data-dependent control flow beyond a mask. The program below bounds a uniform frontier, \(k\) boxes of width \(1.6/k\) on each of the two rows, in one numpy call against the incumbent \(2.0\). It is the simplest instance of the batched-frontier pattern that Sections 6.4 and 7.4 develop in general.

The frontier of the program below at k = 4, as three arrays

  lane      0     1     2     3     4     5     6     7
  a       0.0   0.4   0.8   1.2   0.0   0.4   0.8   1.2
  b       0.4   0.8   1.2   1.6   0.4   0.8   1.2   1.6
  y         0     0     0     0     1     1     1     1
            |     |     |     |     |     |     |     |
            v     v     v     v     v     v     v     v
          chord_bounds(a, b, y): one call, the same arithmetic
          in every lane, no pivoting, no factorization
            |     |     |     |     |     |     |     |
            v     v     v     v     v     v     v     v
  masks   5 infeasible, 2 pruned (bound >= 2.0 - 1e-5),
          1 surviving: lane 5, [0.4000, 0.8000] on y = 1,
          bound 1.950000
# The chord bound of st_e13 on a whole frontier of nodes at once.
#
# After the branch on y every node is a box [a, b] on one row y, and its
# chord bound is one closed-form expression, so a frontier of nodes is
# one array operation. Here the frontier is a uniform grid of k boxes on
# each row, bounded in one numpy call against the incumbent 2.0.

import numpy as np

TOL, INC = 1e-5, 2.0

def chord_bounds(a, b, y):
    """Lower bound of 2x + y over the node (box [a, b], row y).

    From the chord x^2 <= (a+b)x - ab: x >= (1.25 + ab - y)/(a + b);
    the node is feasible iff that point is <= min(b, 1.6 - y).
    """
    xmin = np.maximum(a, (1.25 + a * b - y) / (a + b))
    feasible = xmin <= np.minimum(b, 1.6 - y) + 1e-12
    return np.where(feasible, 2 * xmin + y, np.inf), feasible

print(f"{'k':>4}{'nodes':>7}{'infeas':>8}{'pruned':>8}{'surv':>6}"
      f"  best surviving node and bound")
for k in (1, 2, 4, 8, 10, 16, 40, 1000):
    # k boxes on each of the two rows
    edges = np.linspace(0.0, 1.6, k + 1)
    a = np.tile(edges[:-1], 2)
    b = np.tile(edges[1:], 2)
    y = np.repeat([0.0, 1.0], k)

    # one vectorized evaluation
    bound, feasible = chord_bounds(a, b, y)
    pruned = feasible & (bound >= INC - TOL)
    alive = feasible & ~pruned

    if alive.any():
        i = np.argmin(np.where(alive, bound, np.inf))
        note = (f"[{a[i]:.4f}, {b[i]:.4f}] on y = {int(y[i])}: "
                f"{bound[i]:.6f}")
    else:
        note = "none: the single pass is a complete proof"
    print(f"{k:4d}{2 * k:7d}{int((~feasible).sum()):8d}"
          f"{int(pruned.sum()):8d}{int(alive.sum()):6d}  {note}")
   k  nodes  infeas  pruned  surv  best surviving node and bound
   1      2       0       0     2  [0.0000, 1.6000] on y = 1: 1.312500
   2      4       2       1     1  [0.0000, 0.8000] on y = 1: 1.625000
   4      8       5       2     1  [0.4000, 0.8000] on y = 1: 1.950000
   8     16      11       4     1  [0.4000, 0.6000] on y = 1: 1.980000
  10     20      15       4     1  [0.4800, 0.6400] on y = 1: 1.995000
  16     32      24       8     0  none: the single pass is a complete proof
  40     80      63      16     1  [0.4800, 0.5200] on y = 1: 1.999200
1000   2000    1634     366     0  none: the single pass is a complete proof

Every box left of the parabola's foot on its row is infeasible, every box right of the optimum on its row is pruned or, on the row \(y = 1\) beyond the cap \(x \le 0.6\), infeasible, and for \(k \ge 2\) at most one box survives. The survivor is the box holding \(x^\star = 0.5\) strictly inside on the row \(y = 1\). Its gap \(2(x^\star - a)(b - x^\star)/(a + b) \le w^2/(2(a + b))\) is of second order in the width \(w\), as Proposition 3.5.5 requires of a relaxation with \(\alpha = 2\). When \(0.5\) is a breakpoint of the grid, as for \(k = 16\), or when the gap falls below the tolerance, which the bound guarantees for \(k \ge 358\) and which in fact happens from \(k = 348\) on, nothing survives and the single pass is a complete proof. On this instance a GPU could prove optimality by brute force with one kernel over a few thousand boxes and no tree at all. The reason this does not generalize is the exponent of Proposition 3.5.5: with \(d\) variables entering nonconvex terms a uniform frontier has \(k^d\) boxes, and the tree exists to avoid most of them.

The same computation in C++23 is below. One node is a box and a row, and the bound is a pure function of the node. The frontier is a std::transform with identical control flow per element, which is the shape of a CUDA kernel with one thread per node. The comment inside gives that kernel. The listing was compiled with clang++ -std=c++23 -fsyntax-only and then built and run with Apple clang 17. It prints the same counts and bounds as the table above for \(k = 1, 2, 4, 8, 40, 1000\), one line per \(k\).

// Batched chord bound for a frontier of st_e13 nodes (C++23).
//
// One node is a box [a, b] on the row y; its bound is one closed-form
// expression, so the frontier is a map with identical control flow per
// lane: the shape of a CUDA kernel with one thread per node (the comment
// above bound_frontier gives it).

#include <algorithm>
#include <cmath>
#include <cstdio>
#include <limits>
#include <vector>

// box [a, b] for x, fixed integer row y
struct Node {
    double a, b, y;
};

struct Verdict {
    double bound;
    bool infeasible, pruned;
};

// The chord of x^2 on [a, b] is (a + b) x - a b; the relaxed constraint
// y >= 1.25 - x^2 becomes x >= (1.25 + a b - y) / (a + b); the node is
// infeasible if that exceeds min(b, 1.6 - y).
constexpr Verdict bound_node(Node n, double inc, double tol) {
    const double xmin =
        std::max(n.a, (1.25 + n.a * n.b - n.y) / (n.a + n.b));
    const bool infeasible = xmin > std::min(n.b, 1.6 - n.y) + 1e-12;
    const double z = infeasible ? std::numeric_limits<double>::infinity()
                                : 2.0 * xmin + n.y;
    return {z, infeasible, !infeasible && z >= inc - tol};
}

// The frontier kernel: one evaluation per node, no dependence between
// lanes. In CUDA:
//
//   __global__ void k(const Node* f, Verdict* v, int n,
//                     double inc, double tol) {
//       int i = blockIdx.x * blockDim.x + threadIdx.x;
//       if (i < n) v[i] = bound_node(f[i], inc, tol);
//   }
void bound_frontier(const std::vector<Node>& frontier,
                    std::vector<Verdict>& out, double inc, double tol) {
    out.resize(frontier.size());
    std::transform(frontier.begin(), frontier.end(), out.begin(),
                   [=](Node n) { return bound_node(n, inc, tol); });
}

int main() {
    for (int k : {1, 2, 4, 8, 40, 1000}) {
        // k boxes on each of the two rows y = 0 and y = 1
        std::vector<Node> frontier;
        for (double y : {0.0, 1.0}) {
            for (int i = 0; i < k; ++i) {
                frontier.push_back({1.6 * i / k, 1.6 * (i + 1) / k, y});
            }
        }

        std::vector<Verdict> v;
        bound_frontier(frontier, v, 2.0, 1e-5);

        int inf = 0;
        int pr = 0;
        int alive = 0;
        double best = std::numeric_limits<double>::infinity();
        for (const auto& r : v) {
            inf += r.infeasible;
            pr += r.pruned;
            if (!r.infeasible && !r.pruned) {
                ++alive;
                best = std::min(best, r.bound);
            }
        }

        std::printf("k = %4d: %5zu nodes, %4d infeasible, %3d pruned, "
                    "%d surviving",
                    k, frontier.size(), inf, pr, alive);
        if (alive) {
            std::printf(", best bound %.6f", best);
        }
        std::printf("\n");
    }
}

Cost: a dozen floating-point operations per node and no memory traffic beyond the three input numbers and the verdict. A real bounding kernel adds two things. The first is the rounding discipline of Section 7.3, so that the computed bound is a proof. The second is a relaxation that is an LP rather than a formula, which is what Section 7.4's batched first-order method supplies. The depth-first search of the figure never has more than three open nodes, so batching inside it buys little. A batched search needs a wide frontier, which is the breadth-first or best-first organization of Section 6.4. Its cost is the pruning loss: nodes are bounded before the incumbent that would have pruned them is known. On this instance that loss is zero, because both NLPs sit at depth one and a frontier holding both children of the root finds both incumbents in the same batch.

Haverly's pooling problem

The running example R4 is the smallest instance of the problem that made spatial branch and bound necessary in industry. A refinery blends crude streams of different sulphur content in pools and sells the pooled product to customers with sulphur limits. The pool's quality is a variable, the flow out of the pool is a variable, and the sulphur carried out is their product. Haverly's 1978 instance has three inputs, one pool and two products.C. A. Haverly, "Studies of the behavior of recursion for the pooling problem", ACM SIGMAP Bulletin 25 (1978), and "Behavior of recursion model – more studies", ACM SIGMAP Bulletin 26 (1979). The GAMS Model Library model haverly and MINLPLib's instance haverly (12 variables, 9 constraints, 3 of them quadratic; optimum \(-400\) in minimization form) carry the same numbers. L. S. Lasdon, A. D. Waren, S. Sarkar and F. Palacios, "Solving the pooling problem using generalized reduced gradient and successive linear programming algorithms", ACM SIGMAP Bulletin 27 (1979), studied the same instances with local methods. The modifications called Haverly 2 and 3 are the ones used throughout the literature; they are verified here only in that they reproduce the known optima 600 and 750.

inputcostsulphur %arcsproductpricesulphur limit %demand (max)
A (to pool)63A \(\to\) poolX92.5100
B (to pool)161B \(\to\) poolY151.5200
C (direct)102pool \(\to\) X, Y    
   C \(\to\) X, Y    
Haverly 1 (Haverly 1978). Maximize revenue minus cost. Sulphur in percent. Haverly 2: demand of X raised from 100 to 600. Haverly 3: cost of B lowered from 16 to 13.

The p-formulation writes the flows \(f_A, f_B\) into the pool, \(y_X, y_Y\) from the pool to the products, \(z_X, z_Y\) from C directly to the products, and the pool's sulphur \(q \in [1, 3]\) as variables. This problem maximizes.

\[\begin{aligned} \max\;& 9 (y_X + z_X) + 15 (y_Y + z_Y) - 6 f_A - 16 f_B - 10 (z_X + z_Y) \\ \text{s.t.}\;& f_A + f_B = y_X + y_Y, \qquad 3 f_A + 1\, f_B = q\,y_X + q\,y_Y, \\ & q\,y_X + 2 z_X \le 2.5\,(y_X + z_X), \qquad q\,y_Y + 2 z_Y \le 1.5\,(y_Y + z_Y), \\ & y_X + z_X \le 100, \qquad y_Y + z_Y \le 200, \qquad \text{all flows} \ge 0 . \end{aligned}\]

The first row is the material balance of the pool, the second its sulphur balance, the next two the quality limits of the products. The only nonlinear terms are the two products \(w_X = q\,y_X\) and \(w_Y = q\,y_Y\), with \(q \in [1, 3]\), \(y_X \in [0, 100]\) and \(y_Y \in [0, 200]\), the flow bounds coming from the demands. The McCormick relaxation of Theorem 2.4.7 replaces each product by a variable and adds the eight planes

\[\begin{aligned} w_X &\ge y_X, & w_X &\ge 3 y_X + 100\, q - 300, & w_X &\le 3 y_X, & w_X &\le y_X + 100\, q - 100, \\ w_Y &\ge y_Y, & w_Y &\ge 3 y_Y + 200\, q - 600, & w_Y &\le 3 y_Y, & w_Y &\le y_Y + 200\, q - 200 . \end{aligned}\]

The relaxation's optimum is \(500\), at \(q = 2\) with \(f_A = 50\), \(f_B = 100\), \(y_X = 50\), \(z_X = 50\), \(y_Y = 100\), \(z_Y = 100\), \(w_X = 150\) and \(w_Y = 100\): revenue \(9 \cdot 100 + 15 \cdot 200 = 3900\), cost \(6 \cdot 50 + 16 \cdot 100 + 10 \cdot 150 = 3400\). The mechanism of the overestimate is visible in the lifted variables. \(w_X = 150 = 3\,y_X\) says that the pool sends material of sulphur 3 to X, \(w_Y = 100 = 1 \cdot y_Y\) says that it sends material of sulphur 1 to Y, and the sulphur balance \(3 \cdot 50 + 100 = 250 = w_X + w_Y\) is satisfied. The relaxation has split the one pool into two pools of different quality, which the real plant cannot do. The true optimum is \(400\): \(f_B = 100\), \(y_Y = 100\), \(z_Y = 100\), \(q = 1\), product Y delivered at exactly \(1.5\) percent sulphur and nothing sold as X, with revenue \(3000\) and cost \(1600 + 1000\).

Haverly 1: the McCormick relaxation splits the one pool in two

  relaxed optimum, bound 500 (q = 2):

    A  50 --> [pool part, sulphur 3] --> y_X  50 -+
                                                  +--> X 100
    C  50 -----------------------------> z_X  50 -+

    B 100 --> [pool part, sulphur 1] --> y_Y 100 -+
                                                  +--> Y 200
    C 100 -----------------------------> z_Y 100 -+

    w_X = 150 = 3 y_X and w_Y = 100 = 1 y_Y; the sulphur balance
    3 * 50 + 100 = 250 = w_X + w_Y holds; revenue 3900, cost 3400

  true optimum, 400 (q = 1):

    B 100 --> [pool, sulphur 1] -------> y_Y 100 -+
                                                  +--> Y 200
    C 100 -----------------------------> z_Y 100 -+

    Y at exactly 1.5 % sulphur, nothing sold as X;
    revenue 3000, cost 1600 + 1000

(The p, q and pq formulations) Two reformulations matter, because the strength of the relaxation depends on which formulation is relaxed. The q-formulation of Ben-Tal, Eiger and Gershovitz replaces the pool quality by proportions. The variable \(q_i\) is the fraction of the pool's inflow that comes from input \(i\), so \(q_A + q_B = 1\), and \(y_J\) is the pool's outflow to product \(J\). The bilinear terms become \(q_i y_J\), the amount of input \(i\) that reaches product \(J\) through the pool, and the relaxation replaces each by a variable \(v_{iJ}\), so the pool buys \(\sum_J v_{iJ}\) of input \(i\). Relaxed on its own by McCormick planes, with the outflow \(y_J\) bounded by the demand \(D_J\) (\(100\) for X and \(200\) for Y), this formulation is useless. At \(q_i = 1/2\) the two lower planes read \(v_{iJ} \ge 0\) and \(v_{iJ} \ge y_J - D_J/2\), so for any outflow \(y_J \le D_J/2\) both lifted products may be zero. The relaxed model then sells half of each demand as pool material it never bought, and the bound on Haverly 1 is \(2450\). The pq-formulation of Tawarmalani and Sahinidis adds the rows \(v_{AX} + v_{BX} = y_X\) and \(v_{AY} + v_{BY} = y_Y\), obtained by multiplying \(q_A + q_B = 1\) by each outflow \(y_J\) and replacing the products by the lifted variables. Rows obtained by multiplying pairs of constraints and linearizing the products are RLT rows, after the reformulation–linearization technique that Section 4.7 treats in general. Here they restore the purchase: whatever leaves the pool was bought from A or from B. Tawarmalani and Sahinidis prove that the pq relaxation is at least as tight as the p relaxation and equals the Lagrangian bound that Ben-Tal, Eiger and Gershovitz derive from the q-formulation. On Haverly's three instances the dominance is an equality: p and pq both give \(500\), \(1000\) and \(800\) against the optima \(400\), \(600\) and \(750\). With one pool and no pool capacity, the only RLT products available are the two material-balance rows, and together with the McCormick planes of \(v_{BJ}\) they reproduce the McCormick planes of \(v_{AJ}\), so nothing is gained. The strict improvement of pq over p appears on larger instances with several pools and pool capacities.A. Ben-Tal, G. Eiger and V. Gershovitz, "Global minimization by reducing the duality gap", Mathematical Programming 63 (1994). M. Tawarmalani and N. V. Sahinidis, Convexification and Global Optimization in Continuous and Mixed-Integer Nonlinear Programming (Kluwer, 2002), Chapter 9, for the pq-formulation and its dominance. Strong formulations and their comparison on larger instances: M. Alfaki and D. Haugland, "Strong formulations for the pooling problem", Journal of Global Optimization 56 (2013); the tables of that paper were not consulted for this series. The formulations and solution methods are surveyed in C. Audet, J. Brimberg, P. Hansen, S. Le Digabel and N. Mladenović, "Pooling problem: alternate formulations and solution methods", Management Science 50 (2004). The bounds \(500 / 2450 / 500\), \(1000 / 4700 / 1000\) and \(800 / 2450 / 800\) for p, bare q and pq on Haverly 1 to 3 are printed by a numpy script run for this series (quadratic_lp_bounds.py, 5 October 2026).

Haverly 1 relaxed in three formulations: the McCormick LP bounds

  p:  pool quality q in [1, 3]; w_J = q y_J ......... bound  500
        |
        |  replace the quality by proportions, q_A + q_B = 1
        v
  q:  v_iJ = q_i y_J; at q_i = 1/2 both v_iJ may be 0
      for y_J <= D_J/2: pool material never bought .. bound 2450
        |
        |  multiply q_A + q_B = 1 by each y_J, linearize:
        |  v_AX + v_BX = y_X,  v_AY + v_BY = y_Y  (RLT rows)
        v
  pq: whatever leaves the pool was bought ........... bound  500

  optimum 400; p / q / pq on Haverly 2: 1000 / 4700 / 1000
  against 600, on Haverly 3: 800 / 2450 / 800 against 750

(The landscape over the pool quality, and Haverly's recursion) The instance has a second feature that the figure shows directly. For a fixed pool quality \(q\) the problem is a linear program in the six flows: the sulphur balance fixes the split of the inflow, and the products \(q\,y_J\) are linear. Sweeping \(q\) from \(1\) to \(3\) and solving the LP at every value gives the best profit as a function of \(q\), and the global optimum is its maximum. For Haverly 1 the curve falls from \(400\) at \(q = 1\) to \(300\) at \(q = 1.5\), drops to \(0\) on the plateau from \(q = 1.51\) to \(q = 2.4\), where no flow is profitable, and rises again to \(100\) at \(q = 3\). The landscape has two hills. The local maximum \(100\) at \(q = 3\) is the pool fed by A alone, selling \(50\) units of pool material and \(50\) of C as product X at exactly \(2.5\) percent: \(9 \cdot 100 - 6 \cdot 50 - 10 \cdot 50 = 100\). This is the trap of Haverly's recursion, the method the 1978 paper studied: fix the pool quality, solve the LP, re-estimate the quality from the resulting flows, repeat. Started at \(q = 3\) it stays there, because the LP at \(q = 3\) sends only A into the pool, and the local conditions at \(q = 3\) do not reveal the better hill at \(q = 1\). Haverly reported exactly this dependence on the starting point.Haverly (1978, 1979), cited above. The sweep is the one-pool case of the polynomial algorithm of N. Boland, T. Kalinowski and F. Rigterink, "A polynomially solvable case of the pooling problem", Journal of Global Optimization 67 (2017): with one pool and a bounded number of inputs the problem is solved by polynomially many LPs. The same paper's summary of the complexity border, drawing on Alfaki and Haugland and on Haugland's and Dey and Gupte's work, is that the pooling problem is strongly NP-hard even with a single pool, and NP-hard already with one quality, two inputs and two outputs. The individual attributions are taken from that summary and were not checked against the earlier papers.

Haverly's recursion on Haverly 1, started at q = 3

    +--> fix the pool quality: q = 3
    |       |
    |       v
    |    solve the LP in the six flows: only A enters the pool,
    |    f_A = 50, f_B = 0; X gets 50 of pool material and 50 of C;
    |    profit 9 * 100 - 6 * 50 - 10 * 50 = 100
    |       |
    |       v
    |    re-estimate the quality from the flows:
    |    q = (3 f_A + 1 f_B) / (f_A + f_B) = 3
    +-------+

  the recursion stays at the local maximum 100; the hill at q = 1,
  worth 400, is never revealed by the local conditions at q = 3

The figure below draws the network and the landscape, solves the LP for the chosen \(q\) in the page, and reports the relaxation bounds. This problem maximizes, so the orange bound lies above the blue optimum. The slider "pieces of \(q\)" cuts the quality range into equal pieces and relaxes each piece separately, which is what spatial branching on the pool quality does.

Haverly's pooling problem: above, the network, where inputs A and B blend in one pool of quality q, input C bypasses it, and the line widths are the best flows once q is fixed (green, or blue when that q is a global optimum's). Below, that best profit against q, with the McCormick bound of the chosen relaxation in orange, the global maxima in blue and any lower local maximum as a red ring. Drag in the plot or move the slider to change q, compare the bounds in the side panel, and raise "pieces of q" to relax each piece of the quality range separately, which is what spatial branching on the pool quality does.

(One spatial branch on the quality proves the optimum) In the default view (Haverly 1) the global maximum is \(400\) at \(q = 1\) percent, the other local maximum is \(100\) at \(q = 3\) percent, and both the p and the pq relaxation give the root bound \(500\). The figure's side panel shows the relaxed point with \(q = 2\), flows \(50, 100, 50, 100, 50, 100\) and the two sulphur loads \(150\) and \(100\) discussed above. Set "pieces of \(q\)" to \(2\) and the bound after the split falls to \(400\). On \([1, 2]\) the McCormick relaxation gives exactly \(400\) and on \([2, 3]\) it gives \(100\), so the largest piece bound equals the incumbent, and one spatial branch on the pool quality, at \(q = 2\) percent, proves the optimum. That is the figure's second lesson. The root relaxation is loose because the box of \(q\) is wide, and halving the box tightens the McCormick planes enough to close the gap. With eight pieces the piece bounds are \(400, 366.7, 300, 0, 0, 50, 83.3, 100\), an outer approximation of the landscape from above that touches it at the two hills. On Haverly 2 (demand of X \(600\)) the optimum is \(600\) at \(q = 3\), the other local maximum is \(400\) at \(q = 1\), both relaxations give \(1000\), and two pieces give \(400\) and \(600\). On Haverly 3 (B costs \(13\)) the optimum \(750\) is interior, at \(q = 1.5\) with \(50\) of A and \(150\) of B through the pool and all \(200\) units sold as Y. Its other local maximum is \(125\) at \(q = 2.5\), both relaxations give \(800\), and two pieces give \(750\) and \(125\). A recursion started at either end of the range misses the interior optimum of Haverly 3.

Haverly 1: one spatial branch on the pool quality q, at q = 2

                   q in [1, 3]: McCormick bound 500
            q <= 2 /                         \ q >= 2
                  /                           \
   q in [1, 2]: bound 400              q in [2, 3]: bound 100

  incumbent 400 at q = 1: the largest piece bound equals it, so
  the optimum is proved and the plateau and the right hill are
  pruned by the one relaxation of [2, 3]

  Haverly 2: root 1000; pieces 400 and 600; optimum 600 at q = 3
  Haverly 3: root 800; pieces 750 and 125; optimum 750 at q = 1.5

(What the instance says about incumbents and root bounds) Two remarks connect the instance to later sections. First, stronger root relaxations exist, and Section 4.7 computes them for this instance: the level-1 RLT bound is still \(500\), and only the order-2 moment relaxation closes the gap at the root. Second, the landscape shows why an incumbent is worth so much to a spatial search on this problem. With the incumbent \(400\) in hand, the plateau and the right hill are pruned by a single relaxation of \([2, 3]\). A local search started on the wrong hill would have produced an incumbent of \(100\) and pruned nothing.

Where this is used

Every general-purpose global solver named in Section 5 runs Algorithm 3.5.8 or a close variant. BARON runs it as branch-and-reduce with polyhedral relaxations, and SCIP and Couenne as spatial branch-and-cut over LP relaxations of an extended formulation. ANTIGONE runs it with term-by-term relaxations and dynamically generated cuts, Gurobi's nonconvex mode by translating quadratic constraints into bilinear form and applying spatial branching, and MAiNGO in the reduced space with McCormick relaxations propagated through the DAG.BARON, SCIP and Couenne are documented in Section 3.6 and the sidenotes above. ANTIGONE: R. Misener and C. A. Floudas, "ANTIGONE: Algorithms for coNTinuous / Integer Global Optimization of Nonlinear Equations", Journal of Global Optimization 59 (2014). Gurobi: Gurobi Optimizer Reference Manual, version 13.0, parameter NonConvex and the section on nonconvex quadratic optimization (vendor documentation). MAiNGO: the report of Bongartz, Najman, Sass and Mitsos cited above. Section 5.3 carries the full sources for each solver. Section 3.6 puts two of these node loops side by side. What differs between them is less the algorithm than the defaults: when to run optimality-based tightening, how many propagation rounds, which branching score, where the branching point goes. The two measured facts about those defaults that a reader should carry are that domain reduction is worth an order of magnitude in nodes (Section 2.6 quotes the numbers) and that no branching rule dominates.

What parallelizes

For the GPU programme this subsection supplies the shape of the work. The bound of a box does not depend on any other box, so a frontier of open nodes is a batch, and on the smallest example the batch is one array operation (3.5.3). The propagation inside a node does not depend on the order in which constraints are applied, so it can run as a Jacobi sweep over constraints at the price of more rounds (Section 2.6 and Section 7.4). The optimality-based tightening of a node is a set of LPs sharing one matrix, a natural many-right-hand-side batch. What does not parallelize is the chain. A node's tightening rounds depend on one another, and a child's relaxation waits for its parent's. Counting the longest chain of dependent relaxations in the st_e13 runs above, with a node's first relaxation waiting for its parent's last, gives the following depths.

tightening rulenodesrelaxationsdepth of the chain
none773
one round (default)574
to a fixed point397
exact propagation332
Critical path of the st_e13 runs (longest chain of dependent relaxations, counted by a script run for this series)
The longest chain of dependent relaxations in each st_e13 run

  step               1    2    3    4    5    6    7    depth
  none               o----o----o                          3
  one round          o----o====o----o                     4
  to a fixed point   o----o====o====o====o====o====o      7
  exact propagation  o----o                               2

  o     a relaxation; the first one is the root's
  ----  a child's first relaxation waits for its parent's last
  ====  the same node relaxed again after a tightening round

Under unlimited parallel bounding the shortest chain wins, so the rule with the fewest nodes has the longest critical path and the rule with the widest tree the shortest. Exact propagation is best under both readings because it shortens the chain without adding nodes. And the exponent remains. Proposition 3.5.5 says the worst-case tree is exponential in the number of coordinates that must be split, and Theorem 3.5.7 says a first-order bound multiplies the frontier near the optimum as the tolerance shrinks. No amount of per-node throughput changes either exponent. Parallel work must attack the base, through tighter boxes and tighter relaxations. That is why the domain reduction of Section 2.6, the reformulations of Section 4 and the decomposition of Section 3.7 matter more to the GPU agenda than the speed of one bound.

What a node actually does

Algorithm 3.5.8 lists the steps of a node in the order a textbook would. A solver's node is a pipeline with early exits, and most nodes never reach its end. They are deleted by an inherited bound before any work, or by propagation before the LP, or by the LP before the branching decision. This subsection puts the two best-documented pipelines side by side, SCIP's spatial branch-and-cut and BARON's branch-and-reduce. It then takes the three devices that distinguish a solver's node from the textbook's. The first is the reduce step, the family of range reductions that BARON runs at every node and that gives its algorithm its name. With it come the theorem that says reduction never costs correctness and the theorem that says when a nonconvex problem is nevertheless finite under integer branching alone. The second is conflict analysis, the MILP device that turns an infeasible or pruned node into a globally valid inequality. The third is the restart, which throws the tree away when the root has learned enough to deserve a second presolve. The subsection ends with the table of where a node's time goes, with its MILP and MINLP columns and a column for what parallelizes, step by step.

Two pipelines

stepBARON (branch-and-reduce)SCIP (spatial branch-and-cut)
node selectionNodeSel 0: a composite rule alternating between the node of minimum violation and the node of least lower bound, LIFO under memory pressure; the branching decision is stored and the node partitioned only when selected again (postponed partitioning)best estimate with plunging: dive while the child's bound is within a quarter of the gap, take the best-bound node every tenth selection (nodesel estimate, bestnodefreq 10, maxplungequot 0.25)
presolve and reformulationfactorable decomposition, one auxiliary per intermediate; recognition of multilinear, polynomial, convex-transformable and quadratic structure; user hints for convexityexpression simplification to a canonical form with shared subexpressions; fixing variables that appear in one convex constraint; products of binaries linearized; the MILP presolve
tighteningevery node: linear FBBT (LBTTDo), nonlinear FBBT through the decomposition (TDo), marginals-based (MDo), optimality-based and probing (OBTTDo, PDo) with the number of probes decided automatically; stored duals reused for reductionFBBT through the expression handlers at every node, with structure-specific propagators for quadratics; OBBT at the root only, with an LP iteration budget, filtering and Lagrangian variable bounds propagated in the tree; reduced-cost propagation for integer variables
relaxationpolyhedral outer approximation of the convex factorable relaxation (the default since 2005; 4 supporting lines per convex univariate term, 4 cut rounds); LP or NLP per node (hybrid rule); a MIP when integer variables are present and the activation rule says soLP relaxation of the extended formulation (one auxiliary per annotated subexpression), equalities relaxed to inequalities where monotonicity allows; an NLP relaxation kept for heuristics and separators
cutsouter-approximation rounds at the node; structure-specific relaxations; integer-programming cuts adapted to nonlinear problems (families not documented)gradient cuts for convex pieces, secants for concave ones, RLT, SDP-derived, intersection and perspective cuts, the integer chord, and the MILP separators with efficacy filters
heuristicsuser point; relaxation points that are feasible; multistart local search in preprocessing (NumLoc); local searches at nodes (DoLocal); MIP-relaxation solutions as starting pointsabout 40 heuristics; nonlinear ones: sub-NLP (fix integers, presolve, local NLP), NLP diving, multistart, MPEC rounding, Undercover
branchingvariable by violation transfer with priorities (BrVarStra 0; 1 largest violation; 2 longest edge); point by BrPtStra (0 dynamic; 1 omega; 2 bisection; 3 convex combination)integer variables first by reliability branching; spatial candidates scored by violation and pseudocosts; point = convex combination of LP value and midpoint
pruning and early exits\(\vert U - L \vert \le\) EpsA or EpsR \(\max(\vert L \vert, \vert U \vert)\); boxes below BoxTol 1e-8 discarded; CutOff and Target as user bounds; DeltaTerm on insufficient progress; partial optimization of a node's LPcutoff bound compared with the node bound within epsilon 1e-9; conflict analysis from dual proofs stored in a pool; root restarts when enough integers are fixed
parallelism and determinismthreads (default 1) passed to the MIP and NLP subsolvers; work units (defined below) as a deterministic clock (MaxWork, DeltaTerm); reproducible with threads = 1, a fixed seed and MaxWorkpresolving MIPs, convexity checks and Ipopt's linear algebra multithreaded; tree parallelism through FiberSCIP/UG, opportunistic by default (FiberSCIP also has a deterministic mode); a concurrent mode that runs several SCIP instances with different settings; version 10.1.0 (18 September 2026) adds a -t option that runs several SCIP configurations concurrently
One node, two solvers, as documented. BARON column: BARON User Manual, HTML edition v. 2026.9.10 (10 September 2026), vendor documentation; the node-management design, postponed partitioning and violation transfer as Tawarmalani and Sahinidis (2004) describe them, the polyhedral default as in their 2005 paper and the LP/NLP hybrid as in Khajavirad and Sahinidis (2018); the MIP-relaxation entries rest on the manual and two abstracts only. SCIP column: Bestuzheva et al., SCIP 8 (2025), with Vigerske and Gleixner (2018) and Achterberg (2009); the node-selection defaults from src/scip/set.c of SCIP 10.0.0; the concurrent mode and FiberSCIP's modes as Section 6.3 describes them; version 10.1.0 (18 September 2026) from the SCIP releases page.

Two design differences stand out, and everything else in the two pipelines is the same algorithm with different defaults. BARON reduces at every node by default, including optimality-based tightening and probing with an automatic budget, where SCIP runs OBBT, the optimality-based bound tightening of Section 2.6 that minimizes and maximizes each variable over the relaxation, at the root only and relies on the cheap Lagrangian variable bounds in the tree. And BARON's portfolio of relaxations includes a mixed-integer linear program, the polyhedral relaxation with the integrality of the original integer variables kept, where SCIP's does not.BARON column: The Optimization Firm, BARON User Manual, HTML edition v. 2026.9.10 (10 September 2026), Sections 5.1–5.7, 10.9 and 11, minlp.com/baron-user-manual, which is vendor documentation; the node-management design, postponed partitioning and violation transfer are M. Tawarmalani and N. V. Sahinidis, "Global optimization of mixed-integer nonlinear programs: a theoretical and computational study", Mathematical Programming 99 (2004), Sections 4 to 6; the polyhedral default is M. Tawarmalani and N. V. Sahinidis, "A polyhedral branch-and-cut approach to global optimization", Mathematical Programming 103 (2005); the LP/NLP hybrid is A. Khajavirad and N. V. Sahinidis, "A hybrid LP/NLP paradigm for global optimization relaxations", Mathematical Programming Computation 10 (2018); the MIP relaxations and integrality devices are M. R. Kılınç and N. V. Sahinidis, "Exploiting integrality in the global optimization of mixed-integer nonlinear programming problems with BARON", Optimization Methods and Software 33 (2018), whose body could not be read for this series, so only its abstract and the manual are relied on, and K. Zhou, M. R. Kılınç, X. Chen and N. V. Sahinidis, "An efficient strategy for the activation of MIP relaxations in a multicore global MINLP solver", Journal of Global Optimization 70 (2018), likewise from its abstract. SCIP column: K. Bestuzheva, A. Chmiela, B. Müller, F. Serrano, S. Vigerske and F. Wegscheider, "Global optimization of mixed-integer nonlinear programs with SCIP 8", Journal of Global Optimization 91 (2025), Sections 2.1 to 2.8; S. Vigerske and A. Gleixner, Optimization Methods and Software 33 (2018), cited above; T. Achterberg, "SCIP: solving constraint integer programs", Mathematical Programming Computation 1 (2009); the node-selection defaults are from src/scip/set.c of SCIP 10.0.0.

(The early exits) The early exits are where the time is saved. A node whose inherited bound is already at or above the cutoff is deleted when it is popped, without an LP, a node whose propagation empties an interval is deleted before its LP, and a node whose LP is infeasible or whose LP bound reaches the cutoff is deleted before any heuristic or branching work. BARON's postponed partitioning stores the branching decision with the node and partitions it only when the node is selected again. That halves the memory per node, and it allows a node whose bound after a few dual simplex iterations is already poor to be set aside before its LP is solved to optimality. Algorithm 3.5.8 reaches step 10 at a minority of nodes.

One node as a pipeline with early exits

  pop N from the open list
    |
    v
  inherited bound already at -----yes--> deleted when popped,
  or above the cutoff?                   without an LP
    | no
    v
  propagation empties an ---------yes--> deleted before its LP
  interval?
    | no
    v
  LP infeasible, or LP bound -----yes--> deleted before any
  reaches the cutoff?                    heuristic or branching work
    | no
    v
  heuristics, then branching; branching is step 10
  of Algorithm 3.5.8, reached at a minority of nodes

  BARON (postponed partitioning) stores the branching decision
  with N and partitions N only when it is selected again

The reduce step

Branch-and-reduce is the name Ryoo and Sahinidis gave in 1995 to a branch and bound in which range reduction is a first-class step at every node. Before the node is bounded, and again after its relaxation has been solved, the box is shrunk by tests that provably remove no point better than the incumbent. Only when reduction stalls is the box split.H. S. Ryoo and N. V. Sahinidis, "Global optimization of nonconvex NLPs and MINLPs with applications in process design", Computers & Chemical Engineering 19 (1995), and "A branch-and-reduce approach to global optimization", Journal of Global Optimization 8 (1996); N. V. Sahinidis, "BARON: a general purpose global optimization software package", Journal of Global Optimization 8 (1996). The range-reduction theorems themselves, duality-based reduction, probing, the FBBT fixed point, OBBT and Lagrangian variable bounds, are stated and proved in Section 2.6. The manual states the principle: "Once a feasible solution with objective value \(U\) is known, any part of the search space that cannot contain a better solution than \(U\) may be discarded. BARON uses this principle not only to fathom entire nodes, but also to shrink the box at nodes that survive. This is the 'reduce' step in branch-and-reduce." The general statement of what such a step is allowed to do, and the proof that it costs nothing, are short.

Branch-and-reduce at one node: reduce, bound, reduce, then split

  node N: box B_N; incumbent value U
    |
    v
  reduce: shrink B_N by tests that <-----------------------+
    |     provably remove no point better than U           |
    v                                                      |
  bound: solve the relaxation on B_N ---> fathomed if no   |
    |                                     better solution  |
    |                                     than U can lie   |
    v                                     in B_N           |
  reduce again, now that the relaxation                    |
    |     has been solved                                  |
    v                                                      |
  reduction stalls? -------no: B_N still shrinking---------+
    | yes
    v
  split B_N

Definition 3.6.1 (range-reduction operator). Let \(U\) be a valid upper bound on \(z^\star\) and \(\varepsilon \ge 0\) the pruning tolerance. A range-reduction operator at node \(N\) is a map \(C_N\) from boxes to boxes with \(C_N(B) \subseteq B\) that is sound for the improving set,

\[C_N(B) \;\supseteq\; \{\, x \in \mathcal F \cap B \;:\; f(x) \le U - \varepsilon \,\} .\]

It is feasibility based if it is sound without using \(U\), so that it loses no feasible point, and optimality based otherwise. FBBT and OBBT without the cutoff row are feasibility based. Marginals-based reduction, BARON's name for the duality-based range reduction of Theorem 2.6.18, probing against the incumbent, OBBT with the cutoff row and FBBT on the cutoff constraint are optimality based. Take R2 written as a minimization, \(\min\{-xy : 2x + y \le 1.2\}\) on the unit square, whose optimum is \(-0.18\) at \((0.3, 0.6)\). For an incumbent \(U\) and a tolerance \(\varepsilon\) the improving set is the lens between the hyperbola \(xy = \varepsilon - U\) and the line, and the smallest box around that lens is the tightest box a sound operator may return; everything outside it may be discarded, and once the lens is empty the node itself may be deleted.

What a sound reduction may remove: R2 of Section 1.6 written as a minimization, min −xy subject to 2x + y ≤ 1.2 on [0, 1]², with optimum −0.18 at (0.3, 0.6). The improving set {f ≤ U − ε} is the lens between the hyperbola xy = ε − U and the line, and the smallest box around it is the tightest box an operator sound in the sense of Definition 3.6.1 may return; the box's algebra is the chart's own, and 'sound' is the definition's term, not a claim about any solver's reduction. At U = 0 and ε = 0 nothing feasible is removed; at U = −0.18 with ε > 0 the set is empty and the node may be deleted. The text's incumbent is U = −0.18 and its tolerance zero; the other values of U and ε are illustrative.

Theorem 3.6.2 (reduction never costs correctness). Run the prototype procedure of Definition 3.5.2 with valid monotone node bounds, a consistent bounding operation and a bound-improving selection. Insert at every node, before and after bounding, any finite number of range-reduction operators sound for the improving set with tolerance \(\varepsilon\). Then (a) at every step every feasible \(x\) with \(f(x) < z^{k}_{\mathrm{inc}} - \varepsilon\) lies in the box of some open node, and \(\underline z_k \le z^\star\) whenever \(z^\star < z^{k}_{\mathrm{inc}} - \varepsilon\). (b) With \(\varepsilon = 0\) the conclusion of Theorem 3.5.3 holds. (c) With \(\varepsilon > 0\) the procedure stops after finitely many steps with \(z_{\mathrm{inc}} - z^\star \le \varepsilon\).

Proof sketch. (a) By induction over the steps. A point \(x\) with \(f(x) < z^{k}_{\mathrm{inc}} - \varepsilon\) satisfies \(f(x) \le z^{k'}_{\mathrm{inc}} - \varepsilon\) at every earlier step \(k'\), since the incumbent is nonincreasing. So no sound reduction ever removed it from the box of the node containing it, and no deletion by bound removed that node, whose valid bound is at most \(f(x) < z^{k}_{\mathrm{inc}} - \varepsilon\). Subdivision passes \(x\) to one child. If \(z^\star < z^{k}_{\mathrm{inc}} - \varepsilon\) a global minimizer lies in an open node whose bound is at most \(z^\star\), so \(\underline z_k \le z^\star\). (b) Reduction only shrinks boxes. Along any infinite nested sequence of subdivided nodes the widths therefore still tend to zero when the subdivision is exhaustive, the bounding operation stays consistent (Proposition 3.5.5 applies to the reduced boxes), and the selection rule is unchanged. The proof of Theorem 3.5.3 goes through with (a) supplying \(\underline z_k \le z^\star \le z^{k}_{\mathrm{inc}}\). (c) is Corollary 3.5.4. ∎

Proposition 3.6.3 (reduction never weakens a node bound). Let the relaxation scheme be inclusion isotone: for boxes \(B' \subseteq B\) the relaxed feasible set satisfies \(R(B') \subseteq R(B)\) and the relaxed objective satisfies \(\bar f_{B'} \ge \bar f_B\) on \(B'\). Then \(\bar z(C_N(B_N)) \ge \bar z(B_N)\) for every reduction operator \(C_N\). Factorable relaxations built from convex envelopes on boxes are inclusion isotone.

Proof. The relaxation on the smaller box minimizes a larger function over a smaller set. For envelopes, \(\operatorname{vex}_B f\) restricted to \(B' \subseteq B\) is a convex underestimator of \(f\) on \(B'\), hence at most \(\operatorname{vex}_{B'} f\) (Proposition 2.4.5), and the relaxed constraint sets are nested for the same reason. ∎

(What the theorem does and does not say) The theorem licenses BARON's design of applying several independent reduction mechanisms at every node. It does not say that reduction is free, and the proposition is about one node only. It does not say that the tree with reduction has fewer nodes than the tree without, because reduction changes the branching decisions and node counts are not monotone in the bounds. The measurements settle that question, and Section 2.6 reports them in full. Two headline figures suffice here. Turning off BARON's range reduction multiplied its node counts by a factor of nearly thirteen on one test library, and turning off SCIP's domain propagation increased its node counts by 86 percent.Y. Puranik and N. V. Sahinidis, "Domain reduction techniques for global NLP and MINLP optimization", Constraints 22 (2017), Table 3 (the Global library of 369 problems, node counts up 1180 percent), and S. Vigerske and A. Gleixner, Optimization Methods and Software 33 (2018), Table 2 (456 MINLPLib instances), both cited in full above. The complete ablation tables for BARON, Couenne and SCIP are quoted in Section 2.6.

(The price of probing) The cost of the reduce step is real, and the one independent measurement of it is old but telling. Belotti and coauthors compared Couenne with BARON 7.5. They found that "on the instances that required more than ten seconds of CPU time, a percentage between 41% and 99% of the CPU time was spent in probing, with a geometric average of 77%". In exchange BARON explored about five times fewer nodes than Couenne. On one box-constrained quadratic instance BARON spent 96 percent of its time probing and explored 22,567 nodes against Couenne's 100,726.Belotti et al. (2009), Section 7.3, with BARON 7.5 on a 3.2 GHz machine. The geometric mean of the node ratio was 1 : 5 on the 64 instances that needed at least two nodes for both solvers. Their verdict: "the performances of baron and couenne are comparable, although it is clear that baron is on average better both in CPU time and in lower bound." This is why BARON's default leaves the number of probes to an automatic decision. It is also why the question of Section 7.4, whether a batch of probing LPs sharing one matrix can run on a device, is a question about most of a global solver's time.

The yield and the price of reduction. Left: node counts with reduction switched off, relative to the default: BARON with range reduction off, the Global library of 369 problems, node counts up 1180 percent (Y. Puranik and N. V. Sahinidis, Constraints 22 (2017), Table 3); SCIP with domain propagation off, 456 MINLPLib instances, up 150 percent (S. Vigerske and A. Gleixner, Optimization Methods and Software 33 (2018), Table 2). Right: the share of CPU time BARON 7.5 spent in probing on the instances that required more than ten seconds, and one box-constrained quadratic instance (P. Belotti, J. Lee, L. Liberti, F. Margot and A. Wächter, Optimization Methods and Software 24 (2009), Section 7.3, BARON 7.5 on a 3.2 GHz machine). Published measurements on old versions, quoted, not rerun.

Finite or convergent

Theorem 3.5.3 is a convergence theorem. For integer branching, Section 3.1 had a finiteness theorem. The dividing line is whether a branch fixes something, and one class of nonconvex problems stays on the finite side.

Theorem 3.6.4 (finite termination when integer fixings make the relaxation exact). Let the integer variables \(y = x_I\) have finite bounds and put \(D = \sum_{j \in I} (u_j - l_j)\). Assume (H1) at every node in which all integer variables are fixed, the node relaxation is exact, that is, its optimal point is feasible for the node and \(\bar z(N)\) equals its true value. (H2) At every node neither deleted nor exact, the algorithm branches on an integer variable whose node range contains at least two integers, by the dichotomy \(x_j \le k\), \(x_j \ge k + 1\). (H3) Nodes are deleted by bound with tolerance \(\varepsilon \ge 0\) and incumbents are updated from exact nodes. Then the tree has depth at most \(D\) and at most \(2^{D+1} - 1\) nodes, and the final incumbent satisfies \(z_{\mathrm{inc}} - z^\star \le \varepsilon\), with \(\varepsilon = 0\) allowed. In particular, suppose every nonconvex term of the problem is a function of integer variables alone or a product \(y_j x_k\) of a bounded integer variable and a bounded continuous variable. Suppose the rest of the problem is convex and the products are relaxed by McCormick envelopes. Then (H1) holds and branch and bound on the integer variables terminates finitely at tolerance zero.

Proof. Each split of (H2) replaces a range containing \(u_j - l_j + 1\) integers by two ranges, \(\{l_j, \dots, k\}\) and \(\{k + 1, \dots, u_j\}\), of widths \(k - l_j\) and \(u_j - k - 1\), so the quantity \(D_N = \sum_j (u^N_j - l^N_j)\) decreases by at least one from a node to each child. A node with \(D_N = 0\) has all integer variables fixed and is exact by (H1), so it is closed without branching: the incumbent is updated to its value and the node is deleted by bound. Hence every branched node has \(D_N \ge 1\), the depth is at most \(D\), and a binary tree of depth \(D\) has at most \(2^{D+1} - 1\) nodes. Correctness is Theorem 3.6.2(a). For the last claim: if \(y_j\) is fixed at \(\hat y\), the four McCormick planes of \(w = y_j x_k\) on \([\hat y, \hat y] \times [l_k, u_k]\) all read \(w = \hat y x_k\) (set \(y^L = y^U = \hat y\) in Theorem 2.4.7). A function of integer variables alone is then a constant, and the remaining problem is convex, so the relaxation coincides with the node's subproblem. ∎

(What the theorem covers) This is the precise content of a sentence in BARON's manual, that finite termination "is also guaranteed for problems in which all nonconvexities involve discrete variables or products of discrete and bounded continuous variables", citing Kılınç and Sahinidis. The bound \(2^{D+1} - 1\) is crude: the example below has a tree of 5 nodes where the bound is 15. The connection to the reduce step is this: probing, described below, fixes integer variables at a node without branching, so it reaches the hypothesis (H1) sooner, but it does not change which problems the theorem covers. The theorem says nothing about continuous nonconvexities, and for them finiteness genuinely fails.BARON User Manual, Section 5.7, citing M. R. Kılınç and N. V. Sahinidis, Optimization Methods and Software 33 (2018), cited in full above, whose own formulation of the result could not be read for this series. Finite branch-and-bound schemes for continuous problems with extreme-point solutions exist for special classes: J. P. Shectman and N. V. Sahinidis, "A finite algorithm for global minimization of separable concave programs", Journal of Global Optimization 12 (1998), and F. A. Al-Khayyal and H. D. Sherali, "On finitely terminating branch-and-bound algorithms for some global optimization problems", SIAM Journal on Optimization 10 (2000). Both, as Tawarmalani and Sahinidis (2004, Remark 2.1) describe them, rely on branching at the incumbent or the relaxation point with an exhaustive rule and on the optimum being a vertex; the papers' bodies were not read for this series. Tawarmalani and Sahinidis (2004), Remark 2.1, draw the line between finite and convergent in the words quoted here in substance.

Proposition 3.6.5 (a continuous bilinear instance on which bisection never terminates at tolerance zero). Consider \(\min\{-xy : 2x + y \le 2.4,\ x \in [0, 1],\ y \in [0, 3]\}\) with the McCormick relaxation of \(w = xy\) on each node box, exact arithmetic, deletion by bound with \(\varepsilon = 0\), incumbents taken from relaxation points, and bisection of the relatively widest coordinate at its midpoint. The unique optimum is \((0.6, 1.2)\) with value \(-0.72\), and the procedure never terminates.

Proof. Since \(-xy\) decreases in both variables on the nonnegative orthant, every optimum lies on the line \(2x + y = 2.4\), where \(-x(2.4 - 2x)\) has the unique minimizer \(x = 0.6\) in \([0, 1]\), with \(y = 1.2\). Under bisection every box has \(x\)-bounds of the form \(k/2^m\) and \(y\)-bounds of the form \(3k/2^m\). The values \(0.6 = 3/5\) and \(1.2 = 6/5\) are not of these forms, so the optimum is interior to the box of every node that contains it and never lies on a splitting line. For such a box \(B\), Theorem 2.4.7 gives \(\operatorname{cav}_B(xy)(0.6, 1.2) - 0.72 = \min\{(x^U - 0.6)(1.2 - y^L), (0.6 - x^L)(y^U - 1.2)\} > 0\), so the node bound is strictly below \(-0.72 = z^\star \le U\) and the node is never deleted by bound. It is subdivided, and exactly one child again contains the optimum in its interior. By induction there is an infinite nested sequence of processed nodes. ∎

The box that never closes: Proposition 3.6.5’s instance, min −xy subject to 2x + y ≤ 2.4, x ∈ [0, 1], y ∈ [0, 3], with the McCormick relaxation of w = xy on each node’s box. Under bisection every box has x-bounds k/2^m and y-bounds 3k/2^m, so the optimum (0.6, 1.2) = (3/5, 6/5) is interior to the box of every node that contains it and never lies on a splitting line. The boxes, bounds and excesses are this chart’s evaluation of Theorem 2.4.7’s planes on the box that contains the optimum after m halvings of each axis, with the incumbent fixed at the optimum −0.72: the pruning depth shown is that idealized one. The proposition takes its incumbents from relaxation points, and the run’s node counts (19, 27, 37, 59 and 173 at ε = 10⁻², 10⁻³, 10⁻⁴, 10⁻⁶ and 0; a numpy script written for this series, branch_and_reduce_finiteness.py, 5 October 2026) are the text’s and are not computed here.
Proposition 3.6.5: the optimum (0.6, 1.2) never lies on a splitting line. Under bisection every box has x-bounds of the form k/2ᵐ and y-bounds of the form 3k/2ᵐ; the lines for m = 2 are drawn, and the control steps m from 1 to 4. The optimum (0.6, 1.2) = (3/5, 6/5), with value −0.72, is of neither form, so it is interior to the box of every node that contains it, and that node's bound stays strictly below −0.72 = z*; with deletion by bound at ε = 0 the node is never deleted.

(The instance and its integer twin, run) Running the instance and its integer twin, the same problem with \(y \in \{0, 1, 2, 3\}\), makes the two theorems concrete. The root McCormick bound is \(-1.44\) for both. The integer twin, branching on \(y\) only, closes in 5 nodes at every tolerance including zero. The root branches \(y \le 1\) against \(y \ge 2\). The child \(y \in [2, 3]\) has bound \(-0.4\) and is pruned against the incumbent \(-0.7\), and the child \(y \in [0, 1]\) branches into \(y = 0\) (bound \(0\)) and \(y = 1\), which is exact with value \(-0.7\). The continuous instance needs 19, 27, 37 and 59 nodes at \(\varepsilon = 10^{-2}, 10^{-3}, 10^{-4}, 10^{-6}\). At \(\varepsilon = 0\) it stops after 173 nodes, and only because the computed bound of a box of relative width \(4.8 \times 10^{-7}\) around the optimum equals the incumbent in double precision. A floating-point run that stops at tolerance zero is therefore not evidence of finiteness, which is one more reason to state the tolerance with every termination claim.The runs are from a numpy script written for this series (branch_and_reduce_finiteness.py, 5 October 2026): root bound \(-1.4400\) at \((0.48, 1.44)\); the integer twin 5 nodes and incumbent \(-0.700000\) at every tolerance; the continuous instance 19, 27, 37, 59 and 173 nodes with incumbents \(-0.719297\), \(-0.719956\), \(-0.719997\), \(-0.720000\), \(-0.720000\) and smallest pruned relative box widths \(1.25 \times 10^{-1}\), \(3.1 \times 10^{-2}\), \(1.6 \times 10^{-2}\), \(9.8 \times 10^{-4}\) and \(4.8 \times 10^{-7}\).

The integer twin, node by node: min −xy subject to 2x + y ≤ 2.4, x ∈ [0, 1], y ∈ {0, 1, 2, 3}, branching on y only, with the McCormick relaxation of w = xy on each node’s box. The root bound −1.44 at (0.48, 1.44), the bound −0.4 of y ≥ 2, the bound 0 of y = 0, the exact value −0.7 at y = 1 and the five nodes are the run’s (branch_and_reduce_finiteness.py, 5 October 2026); the bound −0.80 of the node y ≤ 1 and the relaxation points other than the root’s are this chart’s evaluation of Theorem 2.4.7’s planes, not numbers quoted from the run.
The integer twin, y in {0, 1, 2, 3}: 5 nodes at every tolerance

                    root, y in [0, 3]
                    McCormick bound -1.44
                y <= 1 /            \ y >= 2
                      /              \
             y in [0, 1]              y in [2, 3]: bound -0.4,
          y = 0 /     \ y = 1         pruned against the
               /       \              incumbent -0.7
        bound 0         exact, value -0.7:
                        the incumbent

  nodes                    eps = 1e-2   1e-3   1e-4   1e-6      0
  integer twin, y only              5      5      5      5      5
  continuous, bisection            19     27     37     59    173

  Theorem 3.6.4 bounds the twin's tree by 2^{D+1} - 1 = 15 nodes.
  At eps = 0 the continuous run stops only because, in double
  precision, the bound of a box of relative width 4.8e-7 around
  the optimum equals the incumbent.
Nodes against the tolerance, from a numpy script written for this series (branch_and_reduce_finiteness.py, 5 October 2026). The integer twin, branching on y only, closes in 5 nodes at every tolerance, under the bound 2^(D+1) − 1 = 15 of Theorem 3.6.4; the continuous instance, under bisection, needs 19, 27, 37 and 59 nodes at ε = 10⁻², 10⁻³, 10⁻⁴ and 10⁻⁶, and the smallest box it prunes shrinks with ε. At ε = 0 the run stops after 173 nodes only because, in double precision, the computed bound of a box of relative width 4.8 × 10⁻⁷ around the optimum equals the incumbent: a floating-point artefact, not evidence of finiteness.

Probing, implied-bound cuts and the MIP relaxation

The devices BARON added for integer variables are the ones a GPU design can least easily copy, and they are worth seeing on a small instance. Probing on an integer variable fixes it temporarily at one of its bounds (at both values for a binary) and applies a sound procedure to the restricted node. The procedure is either feasibility-based propagation or the node relaxation. The first is the MILP probing of Section 2.5, where the infeasibility of one fixing fixes the variable to the other value. The second yields a value \(L^v_j\) and multipliers that feed the duality-based reduction of Section 2.6. If the restricted node is infeasible or \(L^v_j \ge U - \varepsilon\), the value \(v\) is excluded and the bound moves inward by one. On a continuous variable probing reduces the range, on a binary it fixes the variable.Puranik and Sahinidis (2017), Section 6: "The probing tests are a generalization of probing for the integer programming case. However, they differ from the integer programming case in the sense that probing leads to variable fixing in the case of binary variables, whereas in the continuous case it leads to reduction in the variable domains." The MILP original is M. W. P. Savelsbergh, "Preprocessing and probing techniques for mixed integer programming problems", ORSA Journal on Computing 6 (1994).

Proposition 3.6.6 (McCormick exactness under a fixed factor, and the implied-bound cut). (a) If \(y^L = y^U = \hat y\), the McCormick envelopes of \(w = yx\) on \([\hat y, \hat y] \times [x^L, x^U]\) coincide with \(w = \hat y x\). (b) Let \(P\) be a convex relaxation of a node containing a binary variable \(s\), let \(\ell\) be a linear function of the relaxation's variables, and let \(\ell^0 \ge \max\{\ell(v) : v \in P, s = 0\}\) and \(\ell^1 \ge \max\{\ell(v) : v \in P, s = 1\}\). Then the implied-bound cut \(\ell(v) \le \ell^0 + (\ell^1 - \ell^0)\, s\) holds for every feasible point of the node and for every point of \(\operatorname{conv}\big((P \cap \{s = 0\}) \cup (P \cap \{s = 1\})\big)\). (c) The cut is violated by the relaxation point \(\bar v\) only if \(\bar s \in (0, 1)\), when \(\ell^0, \ell^1\) are the exact maxima.

Proof. (a) Substituting \(y^L = y^U = \hat y\) into the four planes of Theorem 2.4.7 gives \(\hat y x\) for each of them. (b) On \(P \cap \{s = 0\}\) the inequality reads \(\ell \le \ell^0\) and on \(P \cap \{s = 1\}\) it reads \(\ell \le \ell^1\), both valid by definition. A linear inequality valid on two sets is valid on the convex hull of their union, and every feasible point of the node lies in one of the two sets. (c) At \(\bar s \in \{0, 1\}\) the cut holds by (b). ∎

The cut is Savelsbergh's implied-bound inequality written for an arbitrary linear form of the relaxation, for instance for the auxiliary variable \(w\) of a bilinear term whose McCormick bound depends on the box that the binary controls. Its price is two LPs per binary and per linear form, which is the price of probing. The following instance, R2's bilinear objective with a binary that buys capacity,

\[\min\ -xy + 0.4\,s \quad \text{s.t.}\quad 2x + y \le 1.2 + 0.4\,s,\quad x \le 0.3 + 0.5\,s,\quad 0 \le x, y \le 1,\quad s \in \{0, 1\},\]

has optimum \(-0.18\) at \((0.3, 0.6, 0)\): with \(s = 0\) the feasible set is R2's triangle cut at \(x \le 0.3\), and with \(s = 1\) the best point is \((0.4, 0.8)\) with \(-0.32 + 0.4 = 0.08\). The program below runs one node's reduce step on it with the McCormick relaxation of \(w = xy\) on the unit square. Every LP has four variables and is solved exactly by enumerating basic solutions.

# The reduce step at one node.
#
# Instance: R2's bilinear objective with a binary s that buys capacity,
#
#     min  -w + 0.4 s
#     s.t. 2x + y <= 1.2 + 0.4 s,  x <= 0.3 + 0.5 s,
#          0 <= x, y <= 1,  s in {0, 1},
#
# with w standing for x*y and relaxed by the four McCormick planes on
# [0, 1]^2. Optimum -0.18 at (0.3, 0.6, 0). Every LP has four variables
# (x, y, s, w) and is solved exactly by vertex enumeration.

import itertools

import numpy as np

def lp_min(c, A, b):
    """Minimize c.v subject to A v <= b, in four variables, exactly.

    Every four rows with a nonsingular matrix meet in one basic
    solution; the best feasible one is returned as (value, point).
    """
    best, arg = np.inf, None
    for rows in itertools.combinations(range(len(b)), 4):
        M = A[list(rows)]
        if abs(np.linalg.det(M)) < 1e-12:
            continue
        v = np.linalg.solve(M, b[list(rows)])
        if np.all(A @ v <= b + 1e-9) and c @ v < best - 1e-12:
            best, arg = c @ v, v
    return best, arg

def rows(s_lo=0.0, s_hi=1.0):
    """The node LP as rows A v <= b in v = (x, y, s, w).

    s_lo = s_hi fixes s; the defaults relax s to [0, 1].
    """
    A = [[0, 0, 0, -1],                 # w >= 0          McCormick under
         [-1, -1, 0, 1],                # w >= x + y - 1  McCormick under
         [-1, 0, 0, 1],                 # w <= x          McCormick over
         [0, -1, 0, 1],                 # w <= y          McCormick over
         [2, 1, -0.4, 0],               # 2x + y <= 1.2 + 0.4 s
         [1, 0, -0.5, 0],               # x <= 0.3 + 0.5 s
         [1, 0, 0, 0], [-1, 0, 0, 0],   # 0 <= x <= 1
         [0, 1, 0, 0], [0, -1, 0, 0],   # 0 <= y <= 1
         [0, 0, 1, 0], [0, 0, -1, 0]]   # s_lo <= s <= s_hi
    b = [0, 1, 0, 0, 1.2, 0.3, 1, 0, 1, 0, s_hi, -s_lo]
    return np.array(A, float), np.array(b, float)

obj = np.array([0, 0, 0.4, -1.0])

# The root LP, with 0 <= s <= 1.
z0, v0 = lp_min(obj, *rows())
print(f"root LP: {z0:+.6f} at (x, y, s, w) = "
      f"({v0[0]:.4f}, {v0[1]:.4f}, {v0[2]:.4f}, {v0[3]:.4f})")

# Two LPs with s fixed; the better one is the MIP relaxation bound.
zs = {s: lp_min(obj, *rows(s, s)) for s in (0.0, 1.0)}
print(f"LP with s fixed: s = 0: {zs[0.0][0]:+.6f}   "
      f"s = 1: {zs[1.0][0]:+.6f}")
print(f"    (MIP relaxation bound {min(zs[0.0][0], zs[1.0][0]):+.4f})")

# Two probing LPs: the largest w with s fixed at 0 and at 1.
wmax = {s: -lp_min(np.array([0, 0, 0, -1.0]), *rows(s, s))[0]
        for s in (0.0, 1.0)}
print("implied-bound cut from two probing LPs: "
      f"w <= {wmax[0.0]:.4f} + {wmax[1.0] - wmax[0.0]:.4f} s;")
print("    root point violates it by "
      f"{v0[3] - (wmax[0.0] + (wmax[1.0] - wmax[0.0]) * v0[2]):.4f}")

# The root LP again, with the cut as one more row:
#     w - (wmax[1] - wmax[0]) s <= wmax[0].
A, b = rows()
A = np.vstack([A, [0, 0, -(wmax[1.0] - wmax[0.0]), 1.0]])
b = np.append(b, wmax[0.0])
zc, vc = lp_min(obj, A, b)
print(f"LP with the cut: {zc:+.6f}")
print("    at (x, y, s, w) = "
      f"({vc[0]:.4f}, {vc[1]:.4f}, {vc[2]:.4f}, {vc[3]:.4f})")

# The incumbent: round s to 0 and solve the rest exactly.
U = -0.18
print(f"incumbent U = {U:+.2f} at (0.3, 0.6, 0);")
print(f"    probing s = 1 gives {zs[1.0][0]:+.4f} >= U, "
      "so s = 1 is excluded without a node")

# Marginals-based reduction after the fixing: the binding row x <= 0.3
# has multiplier 1.0, and L is the LP value with s = 0.
print("after fixing s = 0 the box shrinks by marginals:")
print(f"    multiplier 1.0 on x <= 0.3 and U - L = {U - zs[0.0][0]:.2f} "
      f"give x >= {0.3 - (U - zs[0.0][0]) / 1.0:.2f}")
root LP: -0.327273 at (x, y, s, w) = (0.4364, 0.4364, 0.2727, 0.4364)
LP with s fixed: s = 0: -0.300000   s = 1: -0.133333
    (MIP relaxation bound -0.3000)
implied-bound cut from two probing LPs: w <= 0.3000 + 0.2333 s;
    root point violates it by 0.0727
LP with the cut: -0.300000
    at (x, y, s, w) = (0.3000, 0.3000, 0.0000, 0.3000)
incumbent U = -0.18 at (0.3, 0.6, 0);
    probing s = 1 gives -0.1333 >= U, so s = 1 is excluded without a node
after fixing s = 0 the box shrinks by marginals:
    multiplier 1.0 on x <= 0.3 and U - L = 0.12 give x >= 0.18

(Reading the output) The run solves six LPs of four variables each: the root, two with \(s\) fixed, two probing LPs that maximize \(w\), and one re-solve with the cut. The two probing LPs are independent of each other, as are the two LPs with \(s\) fixed, which is the batch shape that Section 7.4 exploits. The output reads as follows. The root LP buys 27 percent of the capacity and sits at no bound, so marginals-based reduction has nothing to work with at the root. The MIP relaxation, the same LP with \(s\) integer, is here the better of the two LPs with \(s\) fixed, \(-0.3\) against the root LP's \(-0.3273\). In general it is a tree. The two probing LPs that maximize \(w\) give \(w \le 0.3\) at \(s = 0\) and \(w \le 0.5333\) at \(s = 1\), so Proposition 3.6.6 gives the cut \(w \le 0.3 + 0.2333\,s\), violated by the root point by \(0.0727\). The LP with this one cut has value \(-0.3\). A single integrality-based cut from two probing LPs therefore recovers the whole MIP-relaxation bound without integer branching. Rounding \(s\) to \(0\) and solving the rest gives the incumbent \(-0.18\), which is the global optimum. Probing \(s = 1\) then gives \(-0.1333 \ge U\), so the value \(s = 1\) contains no improving point and \(s\) is fixed to \(0\) without creating a node. Integer branching would have created the same two children, one pruned at once, so the information is the same and probing obtained it without nodes. After the fixing, the binding row \(x \le 0.3 + 0.5s\) has multiplier \(1.0\) and the duality-based reduction of Theorem 2.6.18 with \(U - L = 0.12\) gives \(x \ge 0.18\). FBBT with the cutoff \(xy \ge 0.18\) then shrinks the box round by round toward the single point \((0.3, 0.6)\), since the incumbent is optimal and the two constraints \(xy \ge 0.18\) and \(2x + y \le 1.2\) touch there. The rate is sublinear, because the propagation map \(x \mapsto 0.18/(1.2 - 2x)\) has derivative \(1\) at the fixed point. With FBBT capped at ten rounds the root box's McCormick bound equals \(U\) to \(10^{-9}\), and one node closes the proof at every tolerance. The plain spatial tree on the remaining problem needs 5, 13 and 19 nodes at \(10^{-2}\), \(10^{-3}\) and \(10^{-4}\).The numbers are from a numpy script written for this series (branch_and_reduce_node_loop.py, 5 October 2026), which also runs the spatial proof after the fixing: nodes 5, 13, 19 plain; 5, 7, 9 with marginals-based reduction; 1, 1, 1 with FBBT capped at ten rounds, at tolerances \(10^{-2}\), \(10^{-3}\), \(10^{-4}\).

The implied-bound cut of Proposition 3.6.6 on the instance, in the plane of s and w. The two probing LPs give max w = 0.3000 at s = 0 and 0.5333 at s = 1, and the cut w ≤ 0.3 + 0.2333 s is the chord through them; the root LP point (s, w) = (0.2727, 0.4364) lies 0.0727 above it, and the LP with the cut ends at (0, 0.3). The orange curve, the largest w the root relaxation allows at each s, min(0.4 + 0.4s/3, 0.3 + 0.5s), is this chart’s evaluation of the instance’s rows between those quoted points, not a number from the run; the tinted sliver between curve and chord is what the cut removes.
The duality-based reduction of Theorem 2.6.18 after fixing s = 0. The black curve is the lowest objective a feasible point with a given x can have, −x · min(1, 1.2 − 2x), this chart’s own evaluation of the instance; the orange line of slope −1 through (0.3, −0.3) is the theorem’s underestimate L + μ(u − x), with the LP value L = −0.3 at the bound u = 0.3 and the multiplier μ = 1.0 of the run. No feasible point left of the line’s crossing with the incumbent U can improve on U: at U = −0.18, the optimum, the reduction gives x ≥ 0.18, as in the run; without an incumbent the step is empty.
FBBT round by round after the fixing: each round applies the propagation map x ↦ 0.18/(1.2 − 2x) of the text, from the cutoff xy ≥ 0.18 and y ≤ 1.2 − 2x, to the lower bound on x, starting from x ≥ 0.18 left by the marginals. The map’s derivative is 1 at its fixed point 0.3, so the steps shrink sublinearly. The closed form of the iterates, x_k = 0.3 + 1/(1/(x_0 − 0.3) − 10k/3), and the gap curve on the right are this chart’s own derivation from the text’s map (a Möbius map conjugate to a translation), checked against direct iteration from 0.18 (0.2143, 0.2333, 0.2455, …, 0.2760 after ten rounds); they are not output of the text’s script. The starting-bound slider stays at or above 0.1, where y ≤ 1 does not bind and the text’s map is the propagation step.
The reduce step at one node of the instance, step by step

                                                node bound
  root LP: s = 0.2727, at no bound                 -0.3273
     |  two LPs with s fixed: -0.3 (s = 0), -0.1333 (s = 1)
     v
  MIP relaxation bound                                -0.3
     |  two probing LPs: max w = 0.3 (s = 0), 0.5333 (s = 1)
     v
  root LP with the implied-bound cut                  -0.3
     |  round s to 0, solve the rest
     v
  incumbent U = -0.18 at (0.3, 0.6, 0)
     |  probing s = 1 gives -0.1333 >= U
     v
  s fixed to 0, without creating a node           L = -0.3
     |  marginals: multiplier 1.0 on x <= 0.3, U - L = 0.12
     v
  x >= 0.18
     |  FBBT with the cutoff xy >= 0.18, capped at ten rounds,
     |  shrinks the box toward the single point (0.3, 0.6)
     v
  McCormick bound = U to 1e-9:                       -0.18
  one node closes the proof

  nodes of the spatial proof after the fixing   1e-2  1e-3  1e-4
    plain                                          5    13    19
    with marginals-based reduction                 5     7     9
    with FBBT capped at ten rounds                 1     1     1
The bounds at one node of the capacity instance, from the program’s output above and from a numpy script written for this series (branch_and_reduce_node_loop.py, 5 October 2026). Left: the root LP −0.3273; the MIP relaxation bound −0.3, the better of the two LPs with s fixed; the root LP with the implied-bound cut, −0.3; the probe at s = 1, −0.1333, which is at least the incumbent U = −0.18 and so excludes s = 1 without a node; and the box after marginals and FBBT capped at ten rounds, whose McCormick bound equals U to 10⁻⁹. Right: nodes of the spatial proof after the fixing at ε = 10⁻², 10⁻³ and 10⁻⁴.

(BARON's MIP relaxation) Two facts about BARON's MIP relaxation are documented and one is not. It exists as the polyhedral relaxation with the integrality of the original integer variables kept, not as a piecewise-linear MIP model of the nonlinear functions of the kind Section 4.6 describes. It was added "aiming to improve dual bounds and offer good starting points for primal heuristics". "Since such relaxations necessitate the solution of NP-hard problems, their introduction to a branch-and-bound algorithm raises many practical issues", which the activation strategy of Zhou and coauthors addresses with a dynamic, parameter-free rule. What the rule is, and what the integrality devices measured on the public test set, is in the bodies of two papers that could not be read for this series, so no number from them is quoted. The manual adds that for difficult integer problems BARON "often spends most of its time solving MIP relaxations", which is why its threads option is passed to the MIP subsolver. The SCIP 8 benchmark observed that this was the only use BARON then made of threads, with "only moderate improvement of running time by up to 11%".Kılınç and Sahinidis (2018), abstract; Zhou, Kılınç, Chen and Sahinidis (2018), abstract; BARON User Manual, Sections 5.2 and 10.9; Bestuzheva et al. (2025), Section 3.4.2, for BARON 22.9.30. A GPU has today no MIP solver that proves optimality natively (Section 7.4 on cuOpt), so a MIP relaxation inside a device-resident spatial search would be a host-side task, and the implied-bound cuts of Proposition 3.6.6, which recover the MIP bound on the example with two LPs, are the device-friendly substitute to study.

Conflict analysis and dual proofs

When a node is infeasible or pruned, a MILP solver does not merely delete it. It extracts a certificate of the infeasibility that is valid everywhere in the tree, stores it, and uses it to prune other nodes. The certificate comes from LP duality, and it is as valid for the LP relaxation of a nonconvex MINLP as for a MILP, because the relaxation's rows are valid for the original problem.

Theorem 3.6.7 (dual proofs are globally valid: Achterberg, 2007; Witzig, Berthold and Heinz, 2021). Let the node LP be \(\min\{c^\top x : Ax \le b,\ l' \le x \le u'\}\) with local bounds \(l' \ge l\), \(u' \le u\), where the rows \(Ax \le b\) are valid for every feasible point of the original problem. If the LP is infeasible there is \(y \ge 0\) with \(\min\{(y^\top A)x : l' \le x \le u'\} > y^\top b\). The inequality \((y^\top A)x \le y^\top b\) is valid for the whole problem, and any local bound vector \((\tilde l, \tilde u)\) with \(\min\{(y^\top A)x : \tilde l \le x \le \tilde u\} > y^\top b\) is infeasible. If instead the LP is feasible with optimal value \(\bar z(N) > z_{\mathrm{inc}}\) and row duals \(y \ge 0\), the inequality \(c^\top x - y^\top(b - Ax) \le z_{\mathrm{inc}}\) is valid for every feasible solution better than the incumbent. Its minimum over the local box is \(\bar z(N) > z_{\mathrm{inc}}\), so every local box with the same property contains no improving solution.

Proof. Farkas' lemma gives \(y\) in the infeasible case, and \((y^\top A)x \le y^\top b\) is a nonnegative combination of valid rows. The bound-exceeding case is weak duality: for every \(x\) with \(Ax \le b\), \(c^\top x \ge c^\top x - y^\top(b - Ax)\), and the minimum of the right side over the local box is the LP dual bound, which exceeds \(z_{\mathrm{inc}}\). ∎

The content of conflict analysis is what one does with these certificates.

Algorithm 3.6.8  Conflict analysis from a dual proof
                 (Achterberg 2007; Witzig, Berthold and Heinz 2021)

Input   a node with local bounds (l', u'), its infeasible LP with the
        Farkas ray y >= 0 of Theorem 3.6.7 (or a bound-exceeding LP
        with dual solution y and incumbent z_inc); the chronological
        list of bound changes on the path from the root, each with its
        reason (a branching, or propagation by some row).
        minact(a; l, u) = sum_j min(a_j l_j, a_j u_j) is the smallest
        value of a^T x over the box [l, u].
Output  a globally valid row in the conflict pool; optionally a short
        bound disjunction

 1. a <- A^T y,  beta <- y^T b
    (bound-exceeding case: a <- c + A^T y,  beta <- z_inc + y^T b).
    Verify minact(a; l', u') > beta in floating point with a
    tolerance.                                        [Theorem 3.6.7]

 2. Relax the proof: walk the bound changes from the latest to the
    earliest and undo each one (replace the local bound by the
    global one) if minact stays > beta.
    The surviving bound changes are the conflict set.

 3. Store a^T x <= beta in the conflict pool; propagate it like any
    row at other nodes; delete it when its age (nodes since it last
    pruned or tightened anything) exceeds a limit.

 4. If the conflict set is small, also derive the bound disjunction
       OR over (j, bound) in the set of "x_j violates that bound"
    as a constraint, a clause when the variables are binary;
    optionally run the SAT-style first-unique-implication-point
    analysis on the implication graph of the propagations, so that
    propagated bound changes are replaced by the branching decisions
    that caused them.

Invariant
    a^T x <= beta is a nonnegative combination of global rows (and of
    the incumbent bound), hence valid for every improving feasible
    solution; every bound vector whose minimal activity exceeds beta
    is correctly declared infeasible.
Conflict analysis at a node N: steps 1 to 4 of Algorithm 3.6.8

  N: local bounds (l', u'), its LP infeasible or bound-exceeding
  step 1: the proof a^T x <= beta, with minact(a; l', u') > beta

  the bound changes on the path from the root to N, in order:

    root ---[1]-----[2]-----[3]---- ... ----[k]---> N
            earliest                     latest
            <----------------------------------
            step 2 walks back, from [k] to [1]

  at each [i]:  minact stays > beta with [i] undone?
                   |                            |
                   | yes                        | no
                   v                            v
                undo [i]: replace its        [i] survives: it is
                local bound by the           in the conflict set
                global one

  a^T x <= beta ---------------> the conflict pool (step 3)
  the conflict set, if small --> a bound disjunction (step 4)

Take step 2 of the algorithm on an invented row in two variables, the proof \(x_1 + x_2 \le 0.5\) on the global box \([0, 1]^2\). A local box \(x_1 \ge l_1\), \(x_2 \ge l_2\), reached by an earlier bound change and a later one, is infeasible for the row when its minimal activity \(l_1 + l_2\) exceeds \(0.5\). The step walks back from the later change to the earlier one and undoes each change that is not needed to keep the activity above \(0.5\); the changes that survive are the conflict set, and the box they describe is the larger region the proof excludes.

What a dual proof relaxes: step 2 of Algorithm 3.6.8 on an illustrative row in two variables. The proof x₁ + x₂ ≤ 0.5 (β = 0.5) is a globally valid row (Theorem 3.6.7); the global box is [0, 1]², over which its minimal activity is 0 ≤ β. A local box x₁ ≥ l₁, x₂ ≥ l₂, reached by an earlier bound change [1] and a later one [2], is declared infeasible exactly when its minimal activity, l₁ + l₂ at its lower-left corner, exceeds β. Step 2 walks back from [2] to [1] and undoes a change, replacing its local bound by the global bound 0, if the minimal activity stays above β without it; the surviving changes are the conflict set, and the dashed outline, drawn when a change is undone, is the larger box that the proof excludes with the surviving change alone. At (0.6, 0.6) the later change [2] is undone, since x₁ ≥ 0.6 by itself keeps the activity at 0.6 > β, and [1], tried next, survives; at (0.3, 0.3) neither can be undone; at (0.3, 0.6) only [2] survives. The row, β and the bounds are invented for the picture; they are not an example from Achterberg (2007) or from Witzig, Berthold and Heinz (2021).

(The cost, and a caution for nonconvex rows) Cost: one sparse matrix–vector product and a few passes over the bound changes, negligible next to the LP. Conflicts are generated independently at every node. The pool is shared state. Witzig, Berthold and Heinz compared the variants and found dual-proof analysis with a conflict pool and careful aging the most effective combination in SCIP, with measured speed-ups on the affected instances. A companion paper carried the same machinery to MINLP in SCIP, where the certificates come from the LP relaxation of the convexified problem and remain valid because the convexification is valid. For a nonconvex MINLP there is one caution the theorem makes visible. The rows \(Ax \le b\) must be valid for the original problem. A McCormick plane built on the node's local box is therefore not a global row and may not enter a dual proof as if it were, while a plane built on the global box may. SCIP's nonlinear constraint handler records for every cut whether it is valid globally or only at the node where it was built. A dual proof that uses a local row is itself valid only in the subtree where that row holds.T. Achterberg, "Conflict analysis in mixed integer programming", Discrete Optimization 4 (2007). J. Witzig, T. Berthold and S. Heinz, "Computational aspects of infeasibility analysis in mixed integer programming", Mathematical Programming Computation 13 (2021), with the measured speed-ups in their tables, and J. Witzig, T. Berthold and S. Heinz, "A status report on conflict analysis in mixed integer nonlinear programming", CPAIOR 2019, Lecture Notes in Computer Science 11494 (2019), for the MINLP case. The local/global flag on cuts is in SCIP's source (cons_nonlinear.c and the row flag SCIProwIsLocal, SCIP 10.0.0); the SCIP 8 paper does not discuss it. SCIP 10 adds a cut-based conflict analysis that derives the learned row by linear combinations, integer roundings and MIR cuts of the propagating rows rather than from the conflict graph; it is off by default (C. Hojny et al., "The SCIP Optimization Suite 10.0", arXiv 2511.18580 (2025), Section 3.4; G. Mexi, F. Serrano, T. Berthold, A. Gleixner and J. Nordström, "Cut-based conflict analysis in mixed integer programming", arXiv 2410.15110 (2024), record from the arXiv abstract page, journal venue not checked).

Where a dual proof is valid: global rows and local rows

                 root: the global box
                /                    \
              ...                    N_1  <-- a McCormick plane
                                    /   \     built on the local
                                  ...   N_2   box of N_1
                                       /   \
                                     ...   N_3: LP infeasible,
                                                a dual proof

  a proof from global rows only (planes on the global box):
      valid everywhere; a^T x <= beta enters the conflict pool
  a proof that uses the plane built at N_1:
      valid only in the subtree of N_1, where that row holds

Node selection, phases and restarts

Node selection in a spatial search follows MILP practice. The rules, depth-first, best-bound, best-estimate and plunging, are Definition 3.1.8, and the defaults of SCIP, BARON and Couenne are given with it in Section 3.1. Best-bound solves the fewest relaxations up to ties (Proposition 3.1.9), and it is bound improving in the sense of Definition 3.5.2. One thing is specific to the spatial search. Because the optimality-based reductions depend on the incumbent, early incumbents are worth more than in a MILP, which favours an initial dive before the best-bound phase. And the three phases of Berthold, Hendel and Koch, feasibility, improvement and proof (Section 3.1), apply unchanged: before the first incumbent the search dives, and once the incumbent is believed optimal only the bound remains to be raised, for which best-bound selection and no primal heuristics are right.

A restart is the other structural move. After the root is processed, the solver counts the integer variables fixed globally by reduced-cost fixing, propagation, strong branching and conflict analysis. If the fraction exceeds a threshold (SCIP: 2.5 percent) it transforms the problem with the fixings, presolves it again, keeps the cut pool and the pseudocosts, and starts over from a new root. The fixings enable reductions the first presolve could not make. Later restarts are triggered by tree-size estimation, when the current tree is predicted to be far larger than a restarted one would be.T. Achterberg, Constraint Integer Programming, PhD thesis, Technische Universität Berlin (2007), Section 10.5, depositonce.tu-berlin.de/handle/11303/1931, for root restarts, with SCIP's restartfac = 0.025 in set.c; D. Anderson, G. Hendel, P. Le Bodic and M. Viernickel, "Clairvoyant restarts in branch-and-bound search using online tree-size estimation", AAAI 2019, for restarts triggered by tree-size estimation. The invariant of a restart is that it solves the same problem with a reduced formulation, and the pseudocosts carried over remain unbiased estimates of per-unit gains. For a spatial search a restart has one more use that the MILP literature does not need. The global bounds tightened at the root by OBBT and by the Lagrangian variable bounds become the new box of the whole problem, and every envelope in the second tree is built on it.

A restart on a timeline: what the first root passes on

  [presolve] [root 1] restart [presolve again] [root 2] [tree ...
                |                    ^             ^        ^
                +-- fixings ---------+             |        |
                |   (the problem is transformed)   |        |
                |                                  |        |
                +-- global bounds tightened by ----+--------+
                |   OBBT and the Lagrangian        |        |
                |   variable bounds: the new box,  |        |
                |   on which every envelope of     |        |
                |   the second tree is built       |        |
                |                                  |        |
                +-- the cut pool, the pseudocosts -+--------+

  restart: when the fraction of integer variables fixed globally
  exceeds a threshold (SCIP: 2.5 percent), and later when tree-
  size estimation predicts the current tree far larger than a
  restarted one; the tree is thrown away

Where the time goes

The table below gives, step by step, where a node's time goes in a MILP solver and in a nonconvex MINLP solver, and what parallelizes in each step. One detail of the tightening rows deserves emphasis: propagation runs through the expression graph of every constraint, not through the envelopes, and envelopes enter only through the LP-based OBBT.

stepMILPnonconvex MINLPparallel?
inherited boundcompare with the cutoff: freethe sametrivially, across nodes
propagationactivity bounds on the linear rows: cheapforward-backward sweeps through the expression graph of every constraint; capped rounds; structure-specific propagators for quadraticsacross constraints (Jacobi rounds, same fixed point); across nodes
relaxationone LP, warm-started from the parent's basis, tens of pivotsone LP over envelopes and outer-approximation rows, warm-started; or an NLP solved locallythe simplex: no; first-order methods: matrix-vector products, and batched over sibling nodes with safe bounds
optimality-based tighteningreduced-cost fixing from the duals: freemarginals-based reduction, free; OBBT, two LPs per variable, dear; probing, two LPs per integer: dear\(2\vert K \vert\) LPs sharing one matrix: a batch; Lagrangian variable bounds: a sparse dot product
cutsseparators and the cut loopgradient, secant, RLT, intersection cuts plus the MILP families; efficacy filtersseparation across rows and nodes; selection sequential
heuristicsrounding, diving, local search on the LP pointthe same, plus local NLP solves from the relaxed point and sub-NLPsmultistarts independent; across nodes
branchingon an integer variable: two children, finite treeon an integer or a continuous variable: two children, tree finite only to a tolerance; the point is a second decisionstrong branching candidates as a batch of LPs
pruningbound against incumbentthe same; the bound is weakerfree
conflict analysisdual proof: one matrix-vector productthe same, from globally valid rows onlyacross nodes; the pool is shared state
node selectiona heap operationthe same; the incumbent matters more because reductions depend on itthe sequential core: a shared heap serializes
Where the time goes, per node

(Three measurements) Three measured facts place the rows. Vigerske and Gleixner's and Puranik and Sahinidis's ablations, quoted in Section 2.6 and summarized above, say that the tightening rows are worth up to an order of magnitude in nodes. Belotti and coauthors' measurement says that on BARON 7.5 the probing row was most of the time. And the SCIP 8 benchmark says that on instances solvable in serial within two hours, "enabling parallelization seldom has a considerable advantage". In parallel mode on the 200 unpermuted instances BARON solved 161, 160, 160 and 158 instances at 1, 4, 8 and 16 threads, and FiberSCIP 161, 145, 147 and 152.Bestuzheva et al. (2025), Section 3.4 and Table 3 of the arXiv version (arXiv 2301.00587, January 2023): GAMS 41.2.0 with SCIP 8.0.2 and BARON 22.9.30, two hours, Xeon E5-2670 v2; shifted geometric means of times 64.3, 58.2, 57.1 and 58.6 s for BARON and 76.9, 94.3, 77.8 and 74.8 s for FiberSCIP at 1, 4, 8 and 16 threads. The measurement that would settle how far the pipeline parallelizes on hard instances, those that no serial run finishes, has not been made (Section 8).

What sixteen threads bought, as measured then: the SCIP 8 benchmark in parallel mode, on the 200 unpermuted instances with a two-hour limit. Top, instances solved at 1, 4, 8 and 16 threads: BARON 161, 160, 160 and 158, FiberSCIP 161, 145, 147 and 152. Bottom, shifted geometric means of times: BARON 64.3, 58.2, 57.1 and 58.6 s, FiberSCIP 76.9, 94.3, 77.8 and 74.8 s. On instances solvable in serial within two hours the benchmark says that “enabling parallelization seldom has a considerable advantage”; the measurement that would settle how far the pipeline parallelizes on hard instances, those that no serial run finishes, has not been made (Section 8, What remains). Source: K. Bestuzheva, A. Chmiela, B. Müller, F. Serrano, S. Vigerske and F. Wegscheider, “Global optimization of mixed-integer nonlinear programs with SCIP 8”, Journal of Global Optimization 91 (2025), Section 3.4 and Table 3 of the arXiv version (arXiv 2301.00587, January 2023): GAMS 41.2.0 with SCIP 8.0.2 and BARON 22.9.30, two hours, Xeon E5-2670 v2. These are the versions and the machine of that run, not current releases.

What parallelizes

What parallelizes, step by step, is the right-hand column. Across nodes everything from propagation to branching is independent, and the shared state is the open list, the incumbent, the conflict pool and the stored duals. Inside a node, propagation across constraints is order-free (the fixed-point theorem of Section 2.6), so a Jacobi round over all constraints, in which every constraint is propagated from the same box and the results are intersected afterwards, converges to the same box as a Gauss–Seidel sweep, in which each constraint sees the tightenings of the ones before it, at the price of more rounds. Sofranac, Gleixner and Pokutta measured that price on a GPU for linear rows (Section 7.4). Probing and OBBT are batches of LPs sharing one matrix, differing in one bound or one objective each, and the strong-branching candidates of reliability branching are the same shape. The dual simplex that solves the node LP is the sequential core: its pricing, row computation, ratio test and update form a dependency chain, and its value in the tree is the warm start from the parent's basis. A first-order LP method replaces the chain by matrix–vector products and loses the warm start, and the safe bound of Section 7.3 is what makes its inexact dual usable for pruning and, through the duality-based reduction, for tightening (Section 7.4). Node selection is the one step that does not parallelize at all. A shared best-estimate heap serializes, work stealing with local plunging is the compromise of Section 6.3, and a frontier batched for a device is a different search order with a measurable pruning loss (Section 6.4).

What parallelizes: independent nodes, batches, and a chain

  across nodes: propagation to branching is independent per node

     node 1        node 2        node 3       ...
        \             |             /
         v            v            v
    +------------------------------------------+
    | shared state: the open list, the         |
    | incumbent, the conflict pool, the        |
    | stored duals                             |
    +------------------------------------------+
    node selection: a shared heap serializes

  inside a node

    probing, OBBT, strong branching:   the node LP by dual simplex:
    LPs sharing one matrix             a dependency chain

      +----- one matrix A -----+       pricing --> row computation
      |       |       |        |          ^               |
      LP      LP      LP      ...         |               v
                                       update <------ ratio test
    each LP differs in one bound
    or one objective: a batch          its value: the warm start
                                       from the parent's basis

                                       a first-order method replaces
                                       the chain by matrix-vector
                                       products and loses the warm
                                       start

Determinism is a product requirement for commercial solvers and an open question on a device. BARON measures effort in work units, "an internal, deterministic measure of computational effort that accumulates as BARON constructs relaxations, solves subproblems, performs range reduction, and processes nodes". Its manual gives the recipe for reproducibility: one thread, a fixed seed, and a work limit instead of a time limit, so that a stopped run "stops at exactly the same point on every machine". The commercial MILP engines synchronize their parallel tree searches on such a clock (Section 6.2). A GPU solver that batches a frontier has the batch composition and the atomics in its reductions as two sources of nondeterminism that no logical clock removes, and Section 6.4 says what is known about them.BARON User Manual, Section 5.6 (MaxWork, PrWorkFreq, DeltaTerm) and Section 10.9 on reproducibility with threads > 1, which "additionally also depends on the determinism of the parallel MIP subsolver". Deterministic parallel tree search by logical clock: T. Berthold, J. Farmer, S. Heinz and M. Perregaard, "Parallelization of the FICO Xpress-Optimizer", Optimization Methods and Software 33 (2018), described in Section 6.2, and the Xpress Global paper's measured cost of determinism of 2 to 9 percent (P. Belotti, T. Berthold, T. Gally, L. Gottwald and I. Pólik, "Solving MINLPs to global optimality with FICO Xpress Global", Optimization Online, July 2025, Table 8).

Lagrangian bounds and decomposition

A nonconvex MINLP is often a collection of small nonconvex problems held together by a few constraints or a few variables. A two-stage problem under uncertainty, defined in the next paragraph, has one set of decisions taken in advance and one copy of the operating problem per possible outcome, the copies coupled only through the advance decisions. Multi-site planning, unit commitment, a portfolio with a factor model and the tax-lot problem of Section 9 have the same shape. The spatial branch and bound of Section 3.5 does not see it. It relaxes every nonconvex term over its box, solves one large relaxation per node and branches on single variables. The resulting term-wise relaxation is weaker than the relaxation that convexifies each copy as a whole, by the amount Corollary 3.7.4 quantifies. This subsection gives the bound that exploits the shape, the Lagrangian bound of Section 2.2 applied to the constraints that couple the copies. It gives the theorem that says exactly what the bound convexifies and therefore how strong it is, the splitting device that makes it stronger, and the methods that maximize it. It gives the two algorithms that make decomposition exact for nonconvex MINLP: the Lagrangian branch and cut of Karuppiah and Grossmann and the nonconvex generalized Benders decomposition of Li, Tomasgard and Barton. It ends with the reason this pattern suits a GPU. The subproblems are independent, identical in structure and small. The multiplier update is a reduction. The master of a Benders scheme and the branching of a tree are the sequential part.

Three terms recur below and are fixed here. A two-stage problem has a decision \(y\) that is taken before an uncertain quantity is revealed and, for each of \(S\) scenarios \(s\) (the possible outcomes, with probabilities \(p_s\)), a recourse decision \(x_s\) taken afterwards. Its objective is \(f_0(y) + \sum_s p_s f_s(x_s, y)\), and each scenario's constraints involve only \(y\) and \(x_s\). Writing one copy \(y_s\) of \(y\) per scenario makes the scenarios independent except for the constraints \(y_1 = y_2 = \dots = y_S\), the nonanticipativity constraints, which say that the advance decision cannot depend on the outcome. Throughout this subsection, rows that involve the variables of more than one block are coupling constraints, and variables that appear in more than one block, such as \(y\) above, are linking variables. The literature also calls them complicating constraints and complicating, first-stage or design variables. Those names are used below only when a source's own model is described. Unit commitment, the problem of choosing which generating units to switch on in each hour of a planning horizon, is the oldest large application and is mentioned again at the end.

A two-stage problem before and after copying y (S scenarios)

  before: y is a linking variable, shared by every scenario

                              y
             +----------------+----------------+
             |                |                |
        scenario 1       scenario 2    ...    scenario S
        x_1, p_1         x_2, p_2             x_S, p_S

                              |  one copy y_s of y per scenario
                              v

  after: the scenarios are independent except for y_1 = ... = y_S

        scenario 1       scenario 2    ...    scenario S
        y_1, x_1         y_2, x_2             y_S, x_S
             |                |                |
             +------ = -------+------ = -------+
        y_1 = y_2 = ... = y_S, the nonanticipativity constraints:
        the advance decision cannot depend on the outcome

The Lagrangian bound for coupling constraints

Definition 3.7.1 (block-separable problem and its Lagrangian relaxation). The problem is

\[z^\star \;=\; \min\Big\{ \sum_{i=1}^{n} f_i(x_i) \;:\; \sum_{i=1}^{n} A_i x_i \le b,\quad x_i \in S_i\ (i = 1,\dots,n) \Big\}, \tag{3.7.1}\]

with blocks \(x_i \in \mathbb{R}^{n_i}\), compact block sets \(S_i\) (boxes intersected with the block's own constraints and with the integrality of its integer coordinates), lower semicontinuous block objectives \(f_i\), and \(m\) coupling constraints with matrices \(A_i\) and right-hand side \(b\). The block sets and objectives may be nonconvex. The coupling is linear. For \(\lambda \in \mathbb{R}^m_+\) the Lagrangian relaxation is

\[q(\lambda) \;=\; \min\Big\{ \sum_i f_i(x_i) + \lambda^\top\Big(\sum_i A_i x_i - b\Big) : x_i \in S_i \Big\} \;=\; \sum_{i=1}^n q_i(\lambda) - \lambda^\top b, \qquad q_i(\lambda) = \min_{x_i \in S_i}\ \big[ f_i(x_i) + \lambda^\top A_i x_i \big],\]

whose \(n\) block subproblems are independent. The function \(q\) is the dual function, \(d^\star = \sup_{\lambda \ge 0} q(\lambda)\) is the Lagrangian dual bound, and \(z^\star - d^\star \ge 0\) is the duality gap. Where the distinction from the split dual \(z_{\mathrm{LD}}\) of Definition 3.7.5 matters, the plain bound is written \(z_{\mathrm{LR}} = d^\star\). The convexified problem replaces each \(f_i\) by its closed convex envelope \(\hat f_i = \operatorname{vex}_{S_i} f_i\), the largest closed convex function at most \(f_i\) on \(S_i\) and \(+\infty\) outside \(\operatorname{conv}(S_i)\):

\[\hat z \;=\; \min\Big\{ \sum_i \hat f_i(x_i) : \sum_i A_i x_i \le b,\ x_i \in \operatorname{conv}(S_i) \Big\}. \tag{3.7.2}\]

(What the block subproblems require) This is the Lagrangian of Section 2.2 with the set of easy constraints equal to \(S_1 \times \dots \times S_n\) and the dualized constraints linear. The block subproblems are nonconvex MINLPs of the block's size, to be solved to global optimality, and the one requirement that matters in practice is that each \(q_i(\lambda)\) be a true global minimum, or a certified lower bound on it.

The shape of (3.7.1): n blocks held together by m coupling rows

               x_1      x_2     ...     x_n
            +--------+--------+-----+--------+
  coupling  |  A_1   |  A_2   | ... |  A_n   |  <= b      m rows
            +--------+--------+-----+--------+
  block 1   |  S_1   |
            +--------+--------+
  block 2            |  S_2   |
                     +--------+-----+
   ...                        | ... |
                              +-----+--------+
  block n                           |  S_n   |
                                    +--------+

  price the m coupling rows with lambda >= 0: n independent
  block subproblems

    q_1(lambda)       q_2(lambda)      ...      q_n(lambda)
         \                 |                        /
          q(lambda) = q_1 + ... + q_n - lambda^T b

    q_i(lambda) = min over S_i of [f_i(x_i) + lambda^T A_i x_i],
    a nonconvex MINLP of the block's size, solved globally

Proposition 3.7.2 (weak duality for blocks, and the shape of the dual function). (a) For every \(\lambda \ge 0\) and every feasible point \(x\) of (3.7.1), \(q(\lambda) \le \sum_i f_i(x_i)\), hence \(d^\star \le z^\star\), and every evaluation of \(q\) at a nonnegative multiplier, converged or not, is a valid lower bound. (b) \(q\) is finite and concave on \(\mathbb{R}^m\). If \(x(\lambda)\) solves the block subproblems at \(\lambda\), then \(g(\lambda) = \sum_i A_i x_i(\lambda) - b\) is a supergradient: \(q(\mu) \le q(\lambda) + g(\lambda)^\top(\mu - \lambda)\) for all \(\mu\). (c) Suppose every block has finitely many candidate solutions over all multipliers, for instance because every \(S_i\) is finite or because the block minima are attained at finitely many points. Then \(q\) is the minimum of finitely many affine functions of \(\lambda\), hence piecewise linear.

Proof. (a) For feasible \(x\), \(q(\lambda) \le \sum_i [f_i(x_i) + \lambda^\top A_i x_i] - \lambda^\top b = \sum_i f_i(x_i) + \lambda^\top(\sum_i A_i x_i - b) \le \sum_i f_i(x_i)\), the last step because \(\lambda \ge 0\) and the bracket is nonpositive. (b) \(q\) is a pointwise minimum of affine functions of \(\lambda\), one per \(x \in S_1 \times \dots \times S_n\), hence concave, and finite because the \(S_i\) are compact. For the supergradient, \(q(\mu) \le \sum_i f_i(x_i(\lambda)) + \mu^\top(\sum_i A_i x_i(\lambda) - b) = q(\lambda) + g(\lambda)^\top(\mu - \lambda)\), since \(x(\lambda)\) is admissible in the minimum defining \(q(\mu)\). (c) The minimum is attained at one of the finitely many candidates for every \(\lambda\). ∎

Theorem 3.7.3 (the Lagrangian dual is the convexified problem: Falk, 1969; Geoffrion, 1974; Lemaréchal, 2001). Let (3.7.1) be feasible with compact \(S_i\) and lower semicontinuous \(f_i\). Then for every \(\lambda\)

\[q_i(\lambda) \;=\; \min_{x_i \in \operatorname{conv}(S_i)} \big[ \hat f_i(x_i) + \lambda^\top A_i x_i \big],\]

so that \(q\) is also the dual function of the convexified problem (3.7.2), and

\[d^\star \;=\; \sup_{\lambda \ge 0} q(\lambda) \;=\; \hat z \;\le\; z^\star .\]

The same holds with equality couplings and free multipliers. The duality gap is therefore the price of replacing each block by its convex envelope along the coupling, and nothing else.J. E. Falk, "Lagrange multipliers and nonconvex programs", SIAM Journal on Control 7 (1969), for the envelope principle; A. M. Geoffrion, "Lagrangean relaxation for integer programming", Mathematical Programming Studies 2 (1974), for the integer-programming case, Theorem 2.2.7 (linear \(f_i\), \(S_i\) the integer points of a polyhedron, so that the identity reads \(z_{\mathrm{LD}} = \min\{c^\top x : Ax \le b,\ x \in \operatorname{conv}(X)\}\)); C. Lemaréchal, "Lagrangian relaxation", in Computational Combinatorial Optimization, Lecture Notes in Computer Science 2241 (Springer, 2001), for the convexification reading; C. Lemaréchal and A. Renaud, "A geometric study of duality gaps, with applications", Mathematical Programming 90 (2001), for the image-space version. The proof below is self-contained, since the texts of Geoffrion and Lemaréchal were not consulted for this series.

Proof. Write \(\varphi_i = f_i + \delta_{S_i}\), equal to \(f_i\) on \(S_i\) and \(+\infty\) elsewhere, and \(\ell(x_i) = \lambda^\top A_i x_i\). Adding an affine function commutes with the closed convex envelope: a closed convex \(h\) satisfies \(h \le \varphi_i + \ell\) exactly when \(h - \ell\) is closed convex with \(h - \ell \le \varphi_i\), so the largest such \(h\) is \(\hat f_i + \ell\). A function with compact domain and its closed convex envelope have the same minimum, since the constant \(c = \min(\varphi_i + \ell)\) is a closed convex minorant and so lies below the envelope (this is Falk's envelope principle, Theorem 2.4.4). Together \(q_i(\lambda) = \min(\varphi_i + \ell) = \min(\hat f_i + \ell)\), and summing over the blocks shows that \(q\) is the dual function of (3.7.2). For the second display, (3.7.2) is a convex program with closed convex objective, compact convex domain and linear constraints. Its perturbation function \(v(u) = \min\{\sum_i \hat f_i(x_i) : \sum_i A_i x_i \le b + u,\ x \in \prod_i \operatorname{conv}(S_i)\}\) is convex, finite at \(0\) since (3.7.1) is feasible, and lower semicontinuous at \(0\). For the last property, minimizers \(x^k\) of perturbed problems with \(u_k \to 0\) have a convergent subsequence in the compact domain. Its limit is feasible for \(u = 0\) by continuity of the constraints and has objective at most \(\liminf v(u_k)\) by lower semicontinuity. For a convex perturbation function the dual optimum equals the value of the closure of \(v\) at the origin (Section 2.2), and a convex function finite and lower semicontinuous at a point agrees with its closure there. Hence \(\sup_{\lambda \ge 0} q(\lambda) = v(0) = \hat z\), and \(\hat z \le z^\star\) because every feasible point of (3.7.1) is feasible for (3.7.2) with \(\hat f_i \le f_i\). ∎

(The dual sees the hull of the image) The theorem says that the dual sees each block only through the convex hull of its image \(\{(A_i x_i, f_i(x_i)) : x_i \in S_i\}\), exactly as the picture of Section 2.2 says, and that hull is the epigraph of the block's envelope along the coupling. The theorem has an immediate corollary that is the structural reason to decompose at all. On the two-sites instance of the worked example at the end of this section, two blocks coupled by one crew row, each of the nine crew assignments is a point \((g, f)\), with \(g\) the technicians needed beyond the crew and \(f\) the cost; the line of slope \(-\lambda\) is raised until it touches the points from below, and its height at \(g = 0\) is \(q(\lambda)\). It can only touch the convex hull of the points, which is the sum of the two blocks' hulls shifted by the crew's \(-7\), so the best height, \(d^\star = -0.254\), falls short of the optimum \(-0.16\) by a gap that no multiplier closes.

The dual sees the hull of the image (Theorem 3.7.3), on the two-sites instance of the worked example at the end of the section. Each of the nine crew assignments y ∈ {0, 1, 2}² is a point (g, f) = (3y₁ + 5y₂ − 7, v₁(y₁) + v₂(y₂)), with v₁ = (0, −0.15, −0.16) and v₂ = (0, −0.13, −0.12); the four left of g = 0 are the ones the crew of 7 admits. The line of slope −λ is raised until it touches the image from below, and its height at g = 0 is the bound q(λ); it can only touch the convex hull of the points, the sum of the two blocks' hulls (below) shifted by the crew's −7, so the dual sees each block through the hull of its image. The best height is d* = −0.254, at λ* = 0.026, where the line is the hull's edge from y = (1, 0) to y = (1, 1); the optimum is z* = −0.16, at y = (2, 0), 0.094 above. Both block value functions are convex on the integers, so each block's lower hull is its own polyline. The case λ = 0.0725, where both sites shut down and q = −0.5075, below the frame, is where the subgradient run in the table of further methods overshoots.

Corollary 3.7.4 (the Lagrangian bound dominates every term-wise convex relaxation). Let \(T_i \supseteq S_i\) be closed convex sets and \(\tilde f_i : T_i \to \mathbb{R}\) closed convex functions with \(\tilde f_i \le f_i\) on \(S_i\), for instance the McCormick relaxation of block \(i\) on its box in the lifted variables. Then

\[z_R \;=\; \min\Big\{ \sum_i \tilde f_i(x_i) : \sum_i A_i x_i \le b,\ x_i \in T_i \Big\} \;\le\; \hat z \;=\; d^\star .\]

Proof. \(\tilde f_i + \delta_{T_i}\) is closed convex and at most \(f_i + \delta_{S_i}\), hence at most the largest such function \(\hat f_i\), and \(\operatorname{conv}(S_i) \subseteq T_i\). The relaxation's feasible set contains that of (3.7.2) and its objective is pointwise smaller. ∎

(Why the Lagrangian bound beats a term-wise one) The McCormick relaxation of a block convexifies the product \(x_i w_i\) over the box and the integrality of \(y_i\) separately. The Lagrangian dual convexifies the whole block along \(y_i\), which is tighter whenever the block's envelope is not the sum of the envelopes of its terms, and Section 2.4 showed that it usually is not.

Variable splitting

When the coupling constraints are nonlinear, or involve integer variables whose structure the plain relaxation would lose, the remedy is to copy the coupled variables and price the agreement between the copies.

Definition 3.7.5 (variable splitting and Lagrangian decomposition). Let the problem be \(\min\{f(x) : x \in X \cap Y\}\) with two structured sets, \(X = S_1 \times \dots \times S_n\) the blocks and \(Y = \{x : Ax \le b,\ l \le x \le u,\ x_j \in \mathbb{Z}\ (j \in I)\}\) the coupling constraints with the bounds and integrality of the coupled coordinates. Introduce a copy \(x'\) of \(x\) and write the problem as \(\min\{f(x) : x \in X,\ x' \in Y,\ x_j = x'_j\ (j \in J)\}\), where \(J\) is the set of coordinates that appear in \(Y\). Dualizing the \(|J|\) copy constraints with free multipliers \(\lambda_j\), \(j \in J\), and writing \(\lambda \in \mathbb{R}^n\) for the vector with those entries and zeros elsewhere, gives the Lagrangian decomposition dual function

\[q_{\mathrm{LD}}(\lambda) \;=\; \min_{x \in X}\ \big[ f(x) - \lambda^\top x \big] \;+\; \min_{x' \in Y}\ \lambda^\top x', \qquad z_{\mathrm{LD}} = \sup_{\lambda} q_{\mathrm{LD}}(\lambda).\]

Only the coordinates in \(J\) need copies. The \(X\)-subproblem separates over the blocks with the price \(-\lambda\) in place of \(A_i^\top\lambda\). The \(Y\)-subproblem minimizes a linear function over the coupling set, an integer program when the coupled coordinates are integer. In the two-stage form of Karuppiah and Grossmann the linking variables get one copy per scenario, and the part of the objective that depends on them is split among the copies with weights summing to one. The nonanticipativity constraints, written as \(x^s = x^{s+1}\) between consecutive scenarios, are dualized. That leaves one independent subproblem per scenario.M. Guignard and S. Kim, "Lagrangean decomposition: a model yielding stronger Lagrangean bounds", Mathematical Programming 39 (1987), who named the device and proved the comparison below; its earlier form is the "variable splitting" of K. O. Jörnsten, M. Näsberg and P. A. Smeds (Linköping report LiTH-MAT-R-85-04, 1985, as cited by Karuppiah and Grossmann; not seen). P. Michelon and N. Maculan, "Lagrangean decomposition for integer nonlinear programming with linear constraints", Mathematical Programming 52 (1991), is the nonlinear extension. The scenario form is R. Karuppiah and I. E. Grossmann, "A Lagrangean based branch-and-cut algorithm for global optimization of nonconvex mixed-integer nonlinear programs with decomposable structures", Journal of Global Optimization 41 (2008), Section 2 and model (RP); the numbers quoted from that paper below are from the authors' preprint of July 2006, since the journal version could not be opened for this series.

Variable splitting (Definition 3.7.5): copy, then price the copies

  min f(x)  s.t.  x in X = S_1 x ... x S_n,  x in Y
                          |
                          |  a copy x' of x; only the coordinates
                          |  J that appear in Y need copies
                          v
  min f(x)  s.t.  x in X,  x' in Y,  x_j = x'_j  (j in J)
                          |
                          |  dualize x_j = x'_j with free lambda_j
                          v
      min over x in X                  min over x' in Y
      of f(x) - lambda^T x             of lambda^T x'
      separates over the blocks,       an integer program when
      price -lambda in place of        the coupled coordinates
      A_i^T lambda                     are integer
                  \                      /
                   v                    v
             q_LD(lambda) = the sum of the two minima
             z_LD = sup over lambda of q_LD(lambda)

Theorem 3.7.6 (Lagrangian decomposition is at least as strong: Guignard and Kim, 1987). Let \(X\) be compact, \(f\) lower semicontinuous on \(X\), \(Y\) as in Definition 3.7.5 and nonempty, and \(\hat f_X = \operatorname{vex}_X f\). Then (i) \(z_{\mathrm{LD}} = \min\{\hat f_X(x) : x \in \operatorname{conv}(Y)\}\). (ii) The plain Lagrangian relaxation of \(Ax \le b\) has \(z_{\mathrm{LR}} = \min\{\hat f_X(x) : Ax \le b,\ x \in \operatorname{conv}(X)\}\). (iii) \(z_{\mathrm{LR}} \le z_{\mathrm{LD}} \le z^\star\). (iv) If \(\operatorname{conv}(Y) = \{x : Ax \le b,\ l \le x \le u\}\), in particular if no coupled coordinate is integer or if \(Y\) has the integrality property, meaning that the vertices of its linear relaxation are integral, then \(z_{\mathrm{LD}} = z_{\mathrm{LR}}\).

Proof. (i) Apply Theorem 3.7.3 to the split problem, whose two blocks are \((X, f)\) and \((Y, 0)\) and whose coupling is the equality \(x - x' = 0\) with free multipliers. \(Y\) is a finite union of polytopes, hence compact, and the convex envelope of the zero function on \(Y\) is the indicator of \(\operatorname{conv}(Y)\), so the convexified problem is \(\min\{\hat f_X(x) + \delta_{\operatorname{conv}(Y)}(x') : x = x'\}\). (ii) is Theorem 3.7.3 for the plain relaxation. (iii) \(\operatorname{conv}(Y) \cap \operatorname{conv}(X)\) lies inside \(\{Ax \le b\} \cap \operatorname{conv}(X)\), the effective feasible set of (ii), since \(\hat f_X\) is \(+\infty\) outside \(\operatorname{conv}(X)\). So the minimum in (i) is over a smaller set and \(z_{\mathrm{LD}} \ge z_{\mathrm{LR}}\), and weak duality gives \(z_{\mathrm{LD}} \le z^\star\). (iv) is immediate from (i) and (ii). ∎

(What splitting keeps and what it gives up) The theorem says that the plain relaxation keeps the blocks' structure and replaces the coupling set by its linear relaxation. The split relaxation keeps both structures and gives up only the agreement between the copies. Whatever integrality or nonconvexity the coupling set has is preserved, and that is the only source of the improvement. With a single linear coupling over continuous coordinates the two bounds coincide, and the copies cost multipliers for nothing. The price is one free multiplier per split coordinate and one more subproblem, the integer program over \(Y\), per dual evaluation.

How many blocks are fractional

(Why the gap stays small: two answers) Why Lagrangian bounds on large separable problems are tight has a structural answer and a quantitative one. The quantitative one is the duality-gap bound of Udell and Boyd, stated and proved in Section 5.4 as Theorem 5.4.8, by way of the Shapley–Folkman theorem (Theorem 5.4.7). Write \(\rho(f_i) = \sup_{x_i \in S_i} \big(f_i(x_i) - \hat f_i(x_i)\big)\) for the nonconvexity of block \(i\) (Definition 5.4.6), the largest distance between the block objective and its convex envelope. Then the duality gap \(z^\star - d^\star\) is at most the sum of the \(\tilde m\) largest block nonconvexities, where \(\tilde m \le m\) is the largest number of coupling rows that can be active at one point. The gap therefore stays bounded as the number of blocks grows. The structural half needs no hypothesis beyond compactness, and for the convexified problems that are linear programs it has a direct proof. It uses the block value function along the coupled coordinate. When block \(i\) enters the coupling only through one coordinate \(y_i\), write \(v_i(y) = \min\{f_i(x_i) : x_i \in S_i,\ y_i = y\}\). On an integer range \(\{l_i, \dots, u_i\}\) the function \(v_i\) is a finite list of values, and \(\operatorname{vex} v_i\) is the piecewise-linear lower hull of that list, with breakpoints at integers.

Proposition 3.7.7 (few fractional blocks at a vertex). Suppose each block enters the coupling through one coordinate \(y_i\) with integer range \(\{l_i, \dots, u_i\}\), so that (3.7.2) is the linear program \(\min\{\sum_i s_i : s_i \ge (\text{pieces of } \operatorname{vex} v_i)(y_i),\ l_i \le y_i \le u_i,\ \text{coupling rows}\}\) in the \(2n\) variables \((y, s)\), with \(v_i\) the block value function along \(y_i\). At every vertex of its feasible region the number of blocks whose \(y_i\) lies strictly between two consecutive breakpoints of \(\operatorname{vex} v_i\) is at most the number of coupling rows active at that vertex, hence at most \(m\).

Proof. A vertex has \(2n\) linearly independent active constraints. Block \(i\) contributes at most two of its own when \(y_i\) is at a breakpoint (two adjacent pieces, or a bound and a piece). It contributes at most one when \(y_i\) is strictly inside a piece's interval, since then no bound and no second piece is active. If \(k\) blocks are of the second kind, the blocks supply at most \(2n - k\) active constraints, so at least \(k\) coupling rows must be active. ∎

(At most m fractional blocks) Breakpoints of \(\operatorname{vex} v_i\) are integers, so a block at a breakpoint sits at an integer point of its range. In the worked example below, with one coupling row, each of the two convexified optima has exactly one fractional block, and the whole gap is that one block. For blocks with integer coordinates the Shapley–Folkman bound is vacuous, because the nonconvexity \(\rho(f_i)\) of a block is infinite as soon as the convexified optimum places that block at a fractional point, where \(f_i\) is \(+\infty\). What survives is the proposition: at most \(m\) blocks are fractional, and branching should be aimed at them.J.-P. Aubin and I. Ekeland, "Estimates of the duality gap in nonconvex optimization", Mathematics of Operations Research 1 (1976), whose abstract states the consequence: "when the number of variables is very great with respect to the number of constraints, this duality gap is small in relative value". M. Udell and S. Boyd, "Bounding duality gap for separable problems with linear constraints", Computational Optimization and Applications 64 (2016), Theorem 1, is the modern form, stated and proved as Theorem 5.4.8 in Section 5.4. For blocks with integer coordinates the value-function argument of D. P. Bertsekas, G. S. Lauer, N. R. Sandell Jr. and T. A. Posbergh, "Optimal short-term scheduling of large-scale power systems", IEEE Transactions on Automatic Control 28 (1983), is what controls the gap in unit commitment; its exact hypothesis was not checked for this series.

Maximizing the dual

The dual is a concave, nonsmooth maximization in \(m\) variables, or in as many variables as there are split coordinates, whose value and one supergradient cost one round of block solves. Three families of methods are in use. Subgradient ascent is cheap per step and slow. Cutting planes terminate finitely on a polyhedral dual but are unstable. Bundle methods stabilize the cutting-plane model with a proximal term, a penalty on the distance from the best multiplier found so far, and are the method of choice when block solves are expensive.

Theorem 3.7.8 (subgradient ascent with the Polyak step: Polyak, 1969; Held, Wolfe and Crowder, 1974). Let \(q\) be concave and finite on \(\mathbb{R}^m\) with a nonempty optimal set \(\Lambda^\star\) on \(\Lambda = \mathbb{R}^m_+\) and optimal value \(q^\star\), and let \(P\) project onto \(\Lambda\). The iteration

\[\lambda_{t+1} = P\big( \lambda_t + s_t\, g_t \big), \qquad s_t = \theta_t\, \frac{q^\star - q(\lambda_t)}{\lVert g_t \rVert^2}, \qquad 0 < \varepsilon \le \theta_t \le 2 - \varepsilon,\]

with \(g_t\) a supergradient at \(\lambda_t\), satisfies \(\operatorname{dist}(\lambda_{t+1}, \Lambda^\star) \le \operatorname{dist}(\lambda_t, \Lambda^\star)\) and \(q(\lambda_t) \to q^\star\). With \(q^\star\) replaced by an upper estimate \(\bar q \ge q^\star\), such as an incumbent value, and \(\theta_t\) halved whenever the best value found has not improved for a fixed number of steps, the method is the one of Held, Wolfe and Crowder. Every \(q(\lambda_t)\) remains a valid bound, and the best one is kept.B. T. Polyak, "Minimization of unsmooth functionals", USSR Computational Mathematics and Mathematical Physics 9 (1969); M. Held, P. Wolfe and H. P. Crowder, "Validation of subgradient optimization", Mathematical Programming 6 (1974); M. L. Fisher, "The Lagrangian relaxation method for solving integer programming problems", Management Science 27 (1981), the practitioner's account. A. Nedić and A. Ozdaglar, "Approximate primal solutions and rate analysis for dual subgradient methods", SIAM Journal on Optimization 19 (2009), give rates and recover approximate primal points from averaged subproblem solutions.

Proof. Let \(\lambda^\star \in \Lambda^\star\). Projection onto a convex set containing \(\lambda^\star\) does not increase the distance to it, so \(\lVert \lambda_{t+1} - \lambda^\star \rVert^2 \le \lVert \lambda_t - \lambda^\star \rVert^2 + 2 s_t g_t^\top(\lambda_t - \lambda^\star) + s_t^2 \lVert g_t \rVert^2\). The supergradient inequality at \(\lambda_t\) gives \(g_t^\top(\lambda_t - \lambda^\star) \le -(q^\star - q(\lambda_t))\). Substituting the step, \(\lVert \lambda_{t+1} - \lambda^\star \rVert^2 \le \lVert \lambda_t - \lambda^\star \rVert^2 - \theta_t(2 - \theta_t)\,(q^\star - q(\lambda_t))^2 / \lVert g_t \rVert^2\). The distance is nonincreasing, summing over \(t\) gives \(\sum_t (q^\star - q(\lambda_t))^2/\lVert g_t \rVert^2 < \infty\), and \(\lVert g_t \rVert\) is bounded by the Lipschitz constant of \(q\), so \(q(\lambda_t) \to q^\star\). ∎

(Why the Polyak step works, and why the best value is reported) The step is chosen so that, if \(q\) were the affine function given by the current supergradient, one step would reach the optimal value exactly. Concavity makes the true function lie below that affine function, so the step overshoots but cannot increase the distance to any optimal multiplier. With an overestimate in place of \(q^\star\) the step is too long and the multipliers oscillate, which the halving repairs. Convergence is sublinear and the dual values are not monotone, so the quantity reported as the bound must be the best value seen, never the last.

The Polyak step of Theorem 3.7.8 on the dual q(λ) of the two-sites instance of the worked example below. From λₜ the step θ(q̄ − q(λₜ))/g follows the supergradient line toward the target q̄, and λₜ₊₁ = max(0, λₜ + step). With the exact q* = −0.254 and 0 < θ < 2 the distance to λ* = 0.026 cannot grow: with θ = 1 two steps from 0 land at 0.009 and then exactly at λ*. With an incumbent value as q̄ the step is too long: q̄ = 0 sends one step from 0 to 0.0725, where both sites shut down and q = −0.5075 (the point the methods table's run overshoots to at its second step), and q̄ = −0.16 makes the iterates from 0 bounce between 0.0325 and 0.0025 until θ is halved. The landing points are computed from the section's formula for q.

Proposition 3.7.9 (Kelley's cutting-plane method terminates finitely on a polyhedral dual). Let \(q(\lambda) = \min_{k \in K} (\alpha_k + \beta_k^\top \lambda)\) be the minimum of finitely many affine functions and \(\Lambda\) a compact polyhedron containing a maximizer. From any \(\lambda_1 \in \Lambda\), let the method evaluate \(q(\lambda_t)\) with an active piece \(k(t)\), form the model \(m_t(\lambda) = \min_{s \le t} (\alpha_{k(s)} + \beta_{k(s)}^\top \lambda) \ge q(\lambda)\), and take \(\lambda_{t+1} \in \arg\max_\Lambda m_t\), a linear program. If \(m_t(\lambda_{t+1}) = q(\lambda_{t+1})\) then \(\lambda_{t+1}\) maximizes \(q\) on \(\Lambda\). Otherwise the active piece at \(\lambda_{t+1}\) is new. Hence the method stops with an exact maximizer after at most \(\lvert K \rvert + 1\) evaluations, and at every step \(\max_{s \le t} q(\lambda_s) \le q^\star \le m_t(\lambda_{t+1})\) brackets the optimum.

Proof. \(m_t \ge q\) because every piece in the model is one of the pieces whose minimum is \(q\). If \(m_t(\lambda_{t+1}) = q(\lambda_{t+1})\) then \(q(\lambda_{t+1}) = \max_\Lambda m_t \ge \max_\Lambda q\). Otherwise \(q(\lambda_{t+1}) < m_t(\lambda_{t+1}) \le \alpha_{k(s)} + \beta_{k(s)}^\top \lambda_{t+1}\) for every collected piece, while the active piece attains \(q(\lambda_{t+1})\), so it is not among the collected ones. The bracket is weak duality for the model. ∎

Kelley's model closes the bracket: cutting planes on the dual q of the two-sites instance from λ₁ = 0 over Λ = [0, 0.08], a box chosen for this figure, computed from the formula for q. A box is needed because the model built from one piece is unbounded in its supergradient's direction: the first step stops at the box's edge, as the first two steps of the table's run do on the box [−1, 1]²; that run is Kelley's method on q_LD in two variables, seven evaluations.
Algorithm 3.7.10  Cutting-plane (Kelley) and proximal bundle maximization
                  of a Lagrangian dual q

Input   an oracle DUAL(lambda) returning q(lambda) and a supergradient
        g(lambda) = sum_i A_i x_i(lambda) - b (one global solve per
        block); a compact polyhedron Lambda containing a maximizer;
        tolerance eps; for the bundle variant a proximal weight tau > 0
        and a descent fraction kappa in (0, 1)
Output  q_best <= d* <= z*, the multiplier attaining it, and the model
        value max_Lambda m_t >= d* (a bracket)

 1. lambda_1 <- a point of Lambda (0 for the plain relaxation)
    centre lambda_hat <- lambda_1;  bundle B <- {};  q_best <- -inf

 2. for t = 1, 2, ...

 3.    (q_t, g_t) <- DUAL(lambda_t)
       add the piece l_t(lambda) = q_t + g_t^T (lambda - lambda_t) to B
       q_best <- max(q_best, q_t)               [a valid bound, Prop. 3.7.2]

 4.    Kelley:  lambda_{t+1} <- argmax over Lambda of m_t(lambda)
                   = min over l in B of l(lambda)                    (an LP)

       Bundle:  lambda_{t+1} <- argmax over Lambda of
                   m_t(lambda) - ||lambda - lambda_hat||^2 / (2 tau)
                                                                      (a QP)
                serious step if q(lambda_{t+1}) - q(lambda_hat)
                   >= kappa (m_t(lambda_{t+1}) - q(lambda_hat)):
                      lambda_hat <- lambda_{t+1}
                else a null step: the centre stays and the new piece
                   sharpens the model near it

 5.    stop when m_t(lambda_{t+1}) - q_best <= eps (Kelley)
       or the predicted increase is <= eps (bundle)

Invariant
    m_t >= q on Lambda, so q_best <= d* <= max_Lambda m_t at every step.

Cost/step
    n independent global block solves plus one LP or QP of dimension m
    with |B| rows.

Termination
    finite for polyhedral q with exact oracles (Prop. 3.7.9; Kiwiel 1990
    for the bundle); convergent for every concave q.

Parallel
    the block solves are independent, one block per thread or per batch
    lane; g_t is a reduction over blocks; the master LP or QP is small
    and sequential.
The loop of Algorithm 3.7.10: block solves, a reduction, a master

  +--> lambda_t
  |       |
  |       +--------------+--------------+
  |       v              v              v
  |    block 1        block 2   ...  block n     3. DUAL(lambda_t):
  |       |              |              |        one global solve
  |       +--------------+--------------+        per block, all
  |       v                                      independent
  |    a reduction over the blocks:
  |    q_t = q(lambda_t),  g_t = sum_i A_i x_i(lambda_t) - b
  |       |
  |       v
  |    the piece l_t joins B;  q_best <- max(q_best, q_t)
  |       |
  |       v
  |    4. master over Lambda, small and sequential:
  |       Kelley   argmax m_t                         (an LP)
  |       bundle   argmax m_t - ||lambda - lambda_hat||^2 / (2 tau)
  |                                                    (a QP)
  |       |
  +--- lambda_{t+1}    bundle: a serious step makes it the centre
                       lambda_hat; a null step keeps the centre

  5. stop when m_t(lambda_{t+1}) - q_best <= eps (Kelley), or when
     the predicted increase is <= eps (bundle); at every step
     q_best <= d* <= max_Lambda m_t

A bundle method keeps the cutting-plane model and a stability centre, the best multiplier found, and its serious/null-step logic, in which a step is taken only when the true dual value improves by a fraction of what the model promised and otherwise only the model is enriched, makes the reported bound monotone, which is what a branch-and-bound node wants from its bound. Kiwiel's proximal form with an adaptive weight, the level variant of Lemaréchal, Nemirovskii and Nesterov, and Frangioni's generalized bundle methods with aggregation of pieces are the references. None of them is needed to read the rest of this subsection.C. Lemaréchal, "An extension of Davidon methods to non differentiable problems", Mathematical Programming Studies 3 (1975); K. C. Kiwiel, "Proximity control in bundle methods for convex nondifferentiable minimization", Mathematical Programming 46 (1990); C. Lemaréchal, A. Nemirovskii and Yu. Nesterov, "New variants of bundle methods", Mathematical Programming 69 (1995); A. Frangioni, "Generalized bundle methods", SIAM Journal on Optimization 13 (2002); J.-B. Hiriart-Urruty and C. Lemaréchal, Convex Analysis and Minimization Algorithms II (Springer, 1993), for the textbook treatment. The surveys are A. Frangioni, "About Lagrangian methods in integer optimization", Annals of Operations Research 139 (2005), and M. Guignard, "Lagrangean relaxation", TOP 11 (2003).

Lagrangian cuts and Lagrangian branch and cut

A Lagrangian bound can be used in two ways inside a tree: as the node bound itself, or as a source of cuts for the node's convex relaxation. Karuppiah and Grossmann do both.

Theorem 3.7.11 (validity of the Lagrangian cuts: Karuppiah and Grossmann, 2008). For every \(\lambda \ge 0\) and every block \(i\), every feasible point of (3.7.1) satisfies the block cut

\[f_i(x_i) + \lambda^\top A_i x_i \;\ge\; q_i(\lambda), \tag{3.7.3}\]

and their sum, the aggregated cut \(\sum_i f_i(x_i) + \lambda^\top(\sum_i A_i x_i - b) \ge q(\lambda)\). Adding the cuts to any relaxation of (3.7.1), written with the relaxation's own expressions for the nonconvex terms (the McCormick variable \(t\) for a product \(xw\)), leaves a relaxation. If a block is solved only to a tolerance, the lower bound its solver certifies replaces \(q_i(\lambda)\) and the cut stays valid.

Proof. If \(x\) is feasible then \(x_i \in S_i\) and \(f_i(x_i) + \lambda^\top A_i x_i \ge \min_{x_i' \in S_i}[f_i(x_i') + \lambda^\top A_i x_i'] = q_i(\lambda)\). Summing and subtracting \(\lambda^\top b\) gives the aggregate. On the feasible set of (3.7.1), lifted to the relaxation's variables, the relaxation's expressions for the nonconvex terms take their true values, so every feasible point satisfies the added rows, and the intersection remains a relaxation. ∎

Proposition 3.7.12 (the cut-strengthened relaxation dominates both parents, and cuts inside a tree). (a) Let \(R\) be a convex relaxation of (3.7.1) with objective \(\sum_i \tilde f_i\) as in Corollary 3.7.4 and value \(z_R\), and let \(R(\lambda)\) be \(R\) with the block cuts at \(\lambda\) added as \(\tilde f_i(x_i) + \lambda^\top A_i x_i \ge q_i(\lambda)\). Then \(z_{R(\lambda)} \ge \max\{z_R, q(\lambda)\}\), and with \(\lambda\) optimal, \(z_{R(\lambda^\star)} \ge d^\star\). (b) Let node \(N\) restrict the blocks to \(S_i^N = S_i \cap B_N\) and let the cuts be generated from the node's subproblems. The cuts are valid for the node and every descendant, and cuts generated at an ancestor remain valid below it. If the parent's subproblem solution for block \(i\) lies in the child's box, it solves the child's subproblem for the same \(\lambda\), and that block need not be re-solved. (c) The node bound \(\max\{q^N(\lambda), z_{R^N(\lambda)}\}\) is valid and, made monotone, nondecreasing along every root-to-leaf path.

Proof. (a) \(R(\lambda)\) has the constraints of \(R\) and more, so \(z_{R(\lambda)} \ge z_R\). Summing the block cuts at any point feasible for \(R(\lambda)\) gives \(\sum_i \tilde f_i(x_i) \ge \sum_i q_i(\lambda) - \lambda^\top\sum_i A_i x_i \ge \sum_i q_i(\lambda) - \lambda^\top b = q(\lambda)\), using \(\lambda \ge 0\) and the coupling row, which \(R\) contains. (b) A descendant's feasible set is a subset of the node's, and Theorem 3.7.11 applies at the node with \(S_i^N\) in place of \(S_i\). Restricting a minimization to a subset that still contains a minimizer changes nothing. (c) is Proposition 3.7.2(a) and Theorem 3.7.11 at the node. ∎

(Why the cuts strengthen the relaxation) The inequality in (a) is often strict, because the cuts interact with the relaxation's other constraints. The cut carries the block's exact value into the relaxation as a linear row without changing the relaxation's constraint set. The relaxation has replaced the block's nonconvex structure by convex pieces, while the right-hand side \(q_i(\lambda)\) came from solving the block exactly, so the row is not implied by the relaxation's own constraints. That is the whole idea of the following algorithm, whose steps are those of Karuppiah and Grossmann with the invariant made explicit.

Algorithm 3.7.13  Lagrangian branch and cut for a decomposable nonconvex
                  MINLP (Karuppiah and Grossmann 2008)

Input   problem (3.7.1) in block or scenario form; a convex relaxation
        R(B) of it on any box B (McCormick or piecewise, in lifted
        variables); global solvers for the blocks; tolerance eps;
        a multiplier update (Algorithm 3.7.10 or the subgradient method
        of Theorem 3.7.8)
Output  an eps-global optimum: incumbent z_inc with z_inc - z* <= eps

 1. Root: B_0 <- initial box
    z_inc <- value of a local solve of the full problem (or +inf)
    open <- {B_0};  lambda <- an initial guess

 2. while open is nonempty:

 3.    select a node B from open (depth first in the paper; best bound
       makes the global bound monotone)

 4.    (optional) contract B by bound tightening (Section 2.6)

 5.    Lagrangian decomposition at B: for each block s (one scenario,
       in the two-stage form) solve its priced subproblem on B to
       global optimality, obtaining z_s* (or a certified lower bound)
       and the block's copy of the linking variables.
          If all copies agree, the block solutions form a feasible
          point of the full problem: update z_inc, close B.
          If some block is infeasible: close B (at the root, the
          problem is infeasible).
       Lagrangian bound q(lambda) <- sum_s z_s* - lambda^T b (the last
       term is zero in the scenario form).

 6.    Cuts: for each block add the Lagrangian cut (3.7.3) with
       right-hand side z_s* to the node's relaxation; optionally update
       lambda and repeat 5-6 to add more cuts

 7.    Bound: zbar(B) <- max(q(lambda), value of R(B) with all cuts),
       made monotone against the parent

 8.    Upper bound: fix the discrete linking variables at the
       relaxation's values, solve the remaining nonconvex NLP locally
       from the relaxation's point; update z_inc; if the NLP is solved
       globally, add the integer cut excluding that assignment below B

 9.    Fathom B if zbar(B) >= z_inc - eps, or if it was closed in 5

10.    Branch: choose the linking variable whose copies disperse most
       across the blocks (normalized spread); split B at the copies'
       average for a continuous variable, at y_j = 0 / y_j = 1 for a
       binary; the children inherit the cuts and reuse block solutions
       that fall inside their box (Prop. 3.7.12(b))

Invariant
    min over open nodes of zbar(B) <= z* <= z_inc, and every feasible
    point of the full problem lies in an open node's box or has value
    >= z_inc - eps  (Theorem 3.7.11, Prop. 3.7.12).

Cost/node
    n global block solves per multiplier value (small nonconvex
    MINLPs), one relaxation solve of the full size with the cuts, one
    local NLP solve.

Termination
    finite for binary linking variables and eps > 0 with exhaustive
    branching on the continuous ones (Theorem 3.5.3 applied to the node
    bounds, which are consistent because the relaxation tightens with
    the box).

Parallel
    step 5 is n independent global solves (the paper's Remark 1); the
    cut generation is a gather; step 7 is one sequential relaxation
    solve; nodes are independent as in Section 6.
One node B of Algorithm 3.7.13: what the block solves feed

                           lambda
            +----------------+----------------+
            v                v                v
         block 1          block 2    ...   block N    5. N global
            |                |                |       solves on B,
            +----------------+----------------+       independent
            |                                 |
     the values z_s*                 each block's copy of
            |                        the linking variables
     +------+--------+                        |
     v               v                        +--> all agree: a
  q(lambda)      6. cuts (3.7.3)              |    feasible point;
  = sum_s z_s*   with right-hand              |    update z_inc,
     |           sides z_s*, added            |    close B
     |           to R(B)                      |
     |               |                        +--> 10. they choose
     |               v                             the branching
     |           value of R(B)                     variable: the one
     |           with all cuts                     whose copies
     |               |                             disperse most
     +------+--------+
            v
  7. zbar(B) = max of the two, made monotone against the parent
            |
            v
  9. fathom B if zbar(B) >= z_inc - eps; otherwise 10. branch

(The water network: the regime the algorithm is built for) The paper's Example 2 shows the regime the algorithm is built for. It is the synthesis of an integrated water network with two process units and two treatment units in ten scenarios of contaminant loads and recoveries. The model is a multiscenario MINLP with 24 binary and 764 continuous variables, 928 constraints and 406 bilinear terms, in which the linking variables, the pipe capacities chosen before the scenario is known, couple the scenario flows. BARON 7.2.5 alone "could not verify global optimality of the upper bound of $651,653.06 that it generated, in more than 10 hours". With all multipliers set to one, the Lagrangian bound at the root was $644,856.82 against $610,092.61 for the plain MILP relaxation. With twenty Lagrangian cuts from two multiplier sets the root bound was $645,951.64, within one percent of the incumbent, so the one-percent run ended at the root. Two more nodes were needed for half a percent, with cut-strengthened bounds of $648,566.72 and $648,828.60 against $610,115.37 and $610,109.06 for the MILP relaxation alone. The total was 85.56 CPU seconds including the local solve that supplied the initial incumbent, on a 3.2 GHz machine with BARON for the subproblems and CPLEX 9.0 for the relaxations.Karuppiah and Grossmann (2008), Section 4.2 and Table 2 of the authors' preprint (July 2006, 33 pages, egon.cheme.cmu.edu); the journal version's numbering and figures may differ, and the preprint's text and its Table 2 differ by $3 in the root bound ($645,948.7 against $645,951.64) and by $0.59 in the upper bound ($651,653.06 against $651,653.65). The application papers are R. Karuppiah and I. E. Grossmann, "Global optimization for the synthesis of integrated water systems in chemical processes", Computers & Chemical Engineering 30 (2006), and "Global optimization of multiscenario mixed integer nonlinear programming models arising in the synthesis of integrated water networks under uncertainty", Computers & Chemical Engineering 32 (2008).

The root bound on the water network: Karuppiah and Grossmann (2008), Example 2, numbers from Table 2 of the authors' July 2006 preprint (the journal version could not be opened for the series; the preprint's text and its Table 2 differ by $3 in the root bound, $645,948.7 against $645,951.64, and by $0.59 in the upper bound, $651,653.06 against $651,653.65). Per node, a hollow grey dot is the MILP relaxation alone and a purple dot the cut-strengthened bound (twenty Lagrangian cuts from two multiplier sets at the root); the orange ring is the root's Lagrangian bound with all multipliers one (λ = 1); the green line is $651,653.06, the upper bound that BARON 7.2.5 alone generated and could not verify as globally optimal in more than 10 hours, used here as the incumbent U; the dashed line is U(1 − tol), which a bound ringed green reaches. The run used BARON for the subproblems and CPLEX 9.0 for the relaxations, on a 3.2 GHz machine. The percentages are computed here with U as the denominator. The chart reports the authors' numbers; it is not a comparison of solvers.

Nonconvex generalized Benders decomposition

Benders decomposition goes the other way. It fixes the linking variables \(y\), solves the blocks for the rest, and learns the value function \(v(y) = \min\{f(x, y) : g(x, y) \le 0,\ x \in X\}\) through cuts. Geoffrion's generalized Benders decomposition handles convex inner problems with nonlinear dependence on \(y\), and Proposition 3.4.7 showed its cuts to be aggregations of outer-approximation cuts.J. F. Benders, "Partitioning procedures for solving mixed-variables programming problems", Numerische Mathematik 4 (1962); A. M. Geoffrion, "Generalized Benders decomposition", Journal of Optimization Theory and Applications 10 (1972); R. Van Slyke and R. Wets, "L-shaped linear programs with applications to optimal control and stochastic programming", SIAM Journal on Applied Mathematics 17 (1969), the stochastic-programming form. Geoffrion's text was not consulted for this series, and the statement below is proved here under hypotheses stated in full.

Theorem 3.7.14 (generalized Benders decomposition: Geoffrion, 1972). Consider \(\min\{f(x, y) : g(x, y) \le 0,\ x \in X,\ y \in Y\}\) with \(X\) compact convex, \(Y\) finite, and suppose that for every fixed \(y \in Y\) the inner problem is feasible and convex in \(x\) with zero duality gap and attained dual. Write \(L(x, y, \lambda) = f(x, y) + \lambda^\top g(x, y)\) and \(\ell_\lambda(y) = \min_{x \in X} L(x, y, \lambda)\). Then (i) \(z^\star = \min_{y \in Y} v(y)\). (ii) For every \(\lambda \ge 0\) and every \(y\), \(\ell_\lambda(y) \le v(y)\), and if \(\bar\lambda\) is an optimal multiplier of the inner problem at \(\bar y\) then \(\ell_{\bar\lambda}(\bar y) = v(\bar y)\). (iii) The master \(\min\{\eta : \eta \ge \ell_{\lambda_k}(y)\ (k \in K),\ y \in Y\}\) over the cuts collected so far has value at most \(z^\star\). Consider the loop "solve the master, solve the inner problem at its \(\bar y\), add the cut from its optimal multiplier, update the incumbent with \(v(\bar y)\)". It returns no \(\bar y\) twice before the stopping test (master value \(\ge\) incumbent) fires, so it terminates with the optimum after at most \(\lvert Y \rvert\) inner solves.

Proof sketch. (i) is the definition of \(v\). (ii) is weak and strong duality for the inner problem: \(\ell_\lambda(y) \le L(x, y, \lambda) \le f(x, y)\) for every \(x\) feasible at \(y\), with equality at \((\bar y, \bar\lambda)\) by the zero-gap hypothesis. (iii) Every cut is valid for every \(y\) by (ii), so the master is a relaxation of the projected problem. After the cut at \(\bar y\) is added, the master restricted to \(y = \bar y\) has value at least \(\ell_{\bar\lambda}(\bar y) = v(\bar y) \ge z_{\mathrm{inc}}\), so if \(\bar y\) were returned again the stopping test would fire. \(Y\) is finite. ∎

If the inner problem can be infeasible at some \(\bar y\), a feasibility cut takes the place of the optimality cut. A Farkas-type certificate \(\mu \ge 0\) with \(\min_{x \in X} \mu^\top g(x, \bar y) > 0\) gives the row \(\min_{x \in X} \mu^\top g(x, y) \le 0\), which every \(y\) with finite \(v(y)\) satisfies and \(\bar y\) violates. The finiteness argument then goes through unchanged. Geoffrion proves the existence of \(\mu\) under a constraint qualification that is not restated here.

(What breaks on a nonconvex inner problem) When the inner problem is nonconvex two things break. The zero-gap hypothesis in (ii) fails, so the cut from an optimal multiplier is no longer tight at \(\bar y\), the point can recur, and finiteness is lost. Worse, a multiplier returned by a local solver at a local minimizer gives a "cut" \(\ell_{\bar\lambda}\) that is a valid lower bound on \(v\) only if the minimization in \(\ell_{\bar\lambda}\) is itself global. A local solver does not provide that, so the master can cut off the optimum. The nonconvex version of Li, Tomasgard and Barton repairs both failures by applying Benders decomposition to a convex relaxation of the problem and by enumerating the integers through integer cuts, using the nonconvex inner problem only for upper bounds.X. Li, A. Tomasgard and P. I. Barton, "Nonconvex generalized Benders decomposition for stochastic separable mixed-integer nonlinear programs", Journal of Optimization Theory and Applications 151 (2011). N. V. Sahinidis and I. E. Grossmann, "Convergence properties of generalized Benders decomposition", Computers & Chemical Engineering 15 (1991), analyse the convex method's convergence; its text was not read for this series. The paper's own names for its three subproblems and its tolerance bookkeeping were not accessible, and the theorem below is the structure its abstract describes, proved here.

Algorithm 3.7.15  Nonconvex generalized Benders decomposition
                  (Li, Tomasgard and Barton 2011)

Input   the problem, with Y finite (binary linking variables: the
        first-stage decisions of a two-stage problem) and the inner
        problem separable over scenarios for fixed y; a convex
        relaxation whose inner problem at fixed y is convex with zero
        gap and whose Benders cut is affine in y (separability in x
        and y); a global solver for the nonconvex inner problem; eps
Output  incumbent (x_inc, y_inc) with value U and U - z* <= eps

 1. U <- +inf;  visited V <- {};  cuts C <- {}
    (optionally seeded with a trivial lower bound on eta)

 2. repeat

 3.    Master: (z_lb, ybar) <- min { eta : eta >= l_k(y) for l_k in C,
                                     integer cuts excluding V, y in Y }
                                                                    (a MILP)
       If the master is infeasible (V = Y) or z_lb >= U - eps: stop

 4.    Relaxed inner problem at ybar: for each scenario solve the
       convex relaxed problem (independent), collect the multipliers,
       form the Benders cut as the sum of the scenario cuts (affine in
       y), add it to C

 5.    If the relaxed value at ybar is < U - eps:
          solve every scenario's nonconvex problem at ybar to global
          optimality (independent spatial branch-and-bound runs)
          U <- min(U, v(ybar));  record the incumbent

 6.    V <- V + {ybar};  add the integer cut for ybar to the master

Invariant
    z_lb <= min over Y \ V of v~(y) <= min over Y \ V of v(y), and
    min(U, min over Y \ V of v) = z*  (Theorem 3.7.16).

Cost/iter.
    one master MILP in the linking variables, S convex scenario
    relaxations, and S global scenario solves when the relaxed value
    does not already rule ybar out.

Termination
    at most |Y| iterations; eps-optimal at exit.

Parallel
    steps 4 and 5 are S independent problems each (the relaxed ones
    share one structure and batch naturally; the global ones are
    independent spatial branch-and-bound runs); the master is the
    sequential part.
The loop of Algorithm 3.7.15: the master proposes, scenarios answer

  3. master MILP over the cuts in C and the <-------------------+
     integer cuts: (z_lb, ybar)                                  |
     |                                                           |
     +-- infeasible (V = Y), or z_lb >= U - eps: stop            |
     |                                                           |
     v  ybar                                                     |
  4. relaxed inner problem at ybar: S convex scenario            |
     problems, independent                                       |
        scenario 1     scenario 2     ...     scenario S         |
              \             |                    /               |
               the sum of the scenario cuts: one                 |
               Benders cut, affine in y, added to C              |
     |                                                           |
     v                                                           |
  5. relaxed value at ybar < U - eps ?                           |
     | yes: S global scenario solves, independent spatial        |
     |      branch-and-bound runs; U <- min(U, v(ybar))          |
     | no:  ybar cannot improve U by more than eps               |
     v                                                           |
  6. V <- V + {ybar}; the integer cut for ybar joins the master -+

  the master's lower bound z_lb rises with the cuts, U never
  rises; at most |Y| iterations (Theorem 3.7.16)

Theorem 3.7.16 (nonconvex generalized Benders decomposition, its invariant and finite convergence: after Li, Tomasgard and Barton, 2011). Consider \(\min\{f(x, y) : g(x, y) \le 0,\ x \in X,\ y \in Y\}\) with \(X\) compact, \(Y\) finite, and \(f\) and \(g\) continuous. Let \((\tilde f, \tilde g, \tilde X)\) be a convex relaxation with \(\tilde f \le f\), \(\tilde g \le g\) and \(\tilde X \supseteq X\) compact convex, such that for every fixed \(y \in Y\) the relaxed inner problem is feasible and convex in \(x\) with zero duality gap. Let \(\tilde v(y) \le v(y)\) be the relaxed value function. Then Algorithm 3.7.15 with tolerance \(\varepsilon \ge 0\) maintains at every iteration, with \(V\) the visited set and \(\underline z\), \(U\) its bounds,

\[\underline z \;\le\; \min_{y \in Y \setminus V} \tilde v(y) \;\le\; \min_{y \in Y \setminus V} v(y), \qquad \min\big( U,\ \min_{y \in Y \setminus V} v(y) \big) = z^\star ,\]

so it stops after at most \(\lvert Y \rvert\) iterations and returns an incumbent with \(U - z^\star \le \varepsilon\).

Proof. The Benders cuts of \(\tilde v\) are valid for all \(y\) by Theorem 3.7.14(ii) applied to the relaxed problem, and the integer cuts remove only visited points. So the master of step 3 is a relaxation of \(\min_{Y \setminus V} \tilde v\), which is the first inequality, and \(\tilde v \le v\) is the second. Every visited point has either been solved globally in step 5, so that \(U \le v(\bar y)\), or been found to satisfy \(v(\bar y) \ge \tilde v(\bar y) \ge U - \varepsilon\), so it cannot improve the incumbent by more than \(\varepsilon\). Hence \(z^\star\) is attained within \(\varepsilon\) by the incumbent or by an unvisited point. Each iteration adds one point to \(V\). At termination either \(\underline z \ge U - \varepsilon\), which with the display gives \(\min_{Y \setminus V} v \ge U - \varepsilon\), or \(V = Y\). ∎

(What the relaxed inner problem does) The relaxed inner problem does two jobs. Its value rules out linking-variable points that cannot beat the incumbent, without paying for a global solve, and its multipliers give the only Benders cuts that are valid for a nonconvex problem. The rate at which the master's lower bound rises is governed by the relaxation's tightness. That is why Li, Chen and Barton replace the relaxation by a piecewise convex one and report solution times reduced "by up to an order of magnitude". It is also why the worked example below, whose relaxation is a loose McCormick one, ends up visiting every point of \(Y\). The abstract of the 2011 paper reports "a problem with almost 150,000 variables" solved "within 80 minutes of solver time" and a "dramatic computational advantage" over the general-purpose global solvers of the time.Li, Tomasgard and Barton (2011), abstract. X. Li, Y. Chen and P. I. Barton, "Nonconvex generalized Benders decomposition with piecewise convex relaxations for global optimization of integrated process design and operation problems", Industrial & Engineering Chemistry Research 51 (2012), abstract; X. Li, A. Tomasgard and P. I. Barton, "Decomposition strategy for the stochastic pooling problem", Journal of Global Optimization 54 (2012), applies the method to pooling under uncertainty.

A worked example: two sites, one crew

Two production sites \(i = 1, 2\). Site \(i\) installs \(y_i \in \{0, 1, 2\}\) production lines of capacity \(0.2\) each, so its output satisfies \(0 \le x_i \le 0.2\,y_i\). It buys a second input \(w_i \in [0, 1]\) under the market constraint \(2x_i + w_i \le 1.2\) and earns the product \(x_i w_i\). For \(y_i = 2\) the site's operating problem is the bilinear running example R2 with the added bound \(x \le 0.4\), \(\max\{xw : 2x + w \le 1.2,\ 0 \le x \le 0.4,\ 0 \le w \le 1\}\), with the same optimum \(0.18\) at \((0.3, 0.6)\); R2's McCormick bound \(0.40\) on the unit square becomes \(4/15\) on this box. A line costs \(c_1 = 0.01\) at site 1 and \(c_2 = 0.03\) at site 2. The sites share one crew of \(7\) technicians. A line at site 1 needs \(3\) of them and a line at site 2 needs \(5\):

\[z^\star = \min\ \sum_{i=1}^2 \big( c_i y_i - x_i w_i \big) \quad \text{s.t.}\quad 3 y_1 + 5 y_2 \le 7,\qquad (y_i, x_i, w_i) \in S_i = \{ y_i \in \{0, 1, 2\},\ 0 \le x_i \le 0.2 y_i,\ 0 \le w_i \le 1,\ 2x_i + w_i \le 1.2 \}.\]

Each block is a nonconvex MINLP with one integer variable and one bilinear term. In the symbols of Definition 3.7.1, \(n = 2\), block \(i\) is the triple \((y_i, x_i, w_i)\) with the block set \(S_i\) displayed above and the block objective \(f_i = c_i y_i - x_i w_i\), the one coupling row has \(m = 1\), \(A_1 = (3, 0, 0)\), \(A_2 = (5, 0, 0)\) and \(b = 7\), and the block subproblem \(q_i(\lambda)\) asks site \(i\) how many lines to install when a line costs \(c_i\) plus \(\lambda\) for each technician it needs. The crew constraint is the only coupling, the coupled coordinates are the integers \(y_i\), and the coupling set \(Y = \{y \in \{0, 1, 2\}^2 : 3y_1 + 5y_2 \le 7\}\) has integrality that the plain relaxation loses and the decomposition keeps. The block value functions are computed by hand. For capacity \(X\) the site's best profit is \(\pi(X) = \max\{xw : 0 \le x \le X,\ 0 \le w \le 1,\ 2x + w \le 1.2\}\). For fixed \(x\) the best \(w\) is \(\min(1, 1.2 - 2x)\), and \(x \mapsto x\min(1, 1.2 - 2x)\) increases up to \(x = 0.3\), so \(\pi(0) = 0\), \(\pi(0.2) = 0.16\) and \(\pi(0.4) = 0.18\). Hence \(v_i(y) = c_i y - \pi(0.2y)\) takes the values

\[v_1 = (0,\ -0.15,\ -0.16), \qquad v_2 = (0,\ -0.13,\ -0.12) \qquad \text{on } y = 0, 1, 2 .\]

The second line at site 1 adds only \(0.01\) because the market saturates, and the second line at site 2 loses money. Both value functions are convex on the integers, so the Lagrangian relaxations convexify them without loss along \(y_i\). The gaps below come from the integrality of the coupling and from the one fractional block that a single coupling constraint allows.

The optimum is found by enumerating the four feasible crew assignments. The crew admits \(y \in \{(0,0), (1,0), (2,0), (0,1)\}\), since \((1,1)\) needs eight technicians, with values \(0, -0.15, -0.16, -0.13\): \(z^\star = -0.16\) at \(y = (2, 0)\).

(The McCormick bound) The McCormick relaxation relaxes \(y_i\) to \([0, 2]\) and replaces \(t_i = x_i w_i\) by its McCormick envelope on \([0, 0.4] \times [0, 1]\), of which only the overestimators \(t_i \le x_i\) and \(t_i \le 0.4\,w_i\) matter for a minimization of \(-t_i\). With \(2x_i + w_i \le 1.2\) the largest \(t_i\) is reached at \(x_i = 0.4\,w_i\), \(1.8\,w_i = 1.2\), so \(t_i \le \min(0.2\,y_i, 4/15)\). The LP pays \(0.01 - 0.2 = -0.19\) per unit of \(y_1\) up to \(y_1 = 4/3\) and \(0.03 - 0.2 = -0.17\) per unit of \(y_2\) up to \(4/3\). Per technician site 1 is better, so it takes \(y_1 = 4/3\) (four technicians) and \(y_2 = 0.6\) (the remaining three): \(z_{\mathrm{LP}} = 0.01 \cdot 4/3 - 4/15 + 0.03 \cdot 0.6 - 0.12 = -0.3553\).

(The Lagrangian bound of the crew row) The Lagrangian relaxation of the crew row has the dual function \(q(\lambda) = \sum_i \min_{y \in \{0,1,2\}} [(c_i + 3\lambda\,[i = 1] + 5\lambda\,[i = 2])\,y - \pi(0.2y)] - 7\lambda\). At \(\lambda = 0.026\) block 1 gives \(\min(0,\ 0.088 - 0.16,\ 0.176 - 0.18) = -0.072\) at \(y_1 = 1\). Block 2 gives \(\min(0,\ 0.16 - 0.16,\ 0.32 - 0.18) = 0\), a tie between \(y_2 = 0\) and \(y_2 = 1\), which is why \(0.026\) is a kink. So \(q(0.026) = -0.072 + 0 - 0.182 = -0.254\). By Theorem 3.7.3 this is the LP \(\min \operatorname{vex} v_1(y_1) + \operatorname{vex} v_2(y_2)\) over the knapsack polytope \(3y_1 + 5y_2 \le 7\), \(0 \le y \le 2\), attained at \((1, 0.8)\): \(-0.15 + 0.8 \cdot (-0.13) = -0.254\).

q(0.026) on the two-sites instance: the crew row priced

      the crew row 3y_1 + 5y_2 <= 7, priced at lambda = 0.026
                                 |
            +--------------------+--------------------+
            v                                         v
  site 1: priced line cost              site 2: priced line cost
  0.01 + 0.078 = 0.088                  0.03 + 0.13 = 0.16
  y_1 = 0, 1, 2:                        y_2 = 0, 1, 2:
  0, 0.088 - 0.16, 0.176 - 0.18         0, 0.16 - 0.16, 0.32 - 0.18
  min -0.072 at y_1 = 1                 min 0, a tie: y_2 = 0 or 1
            |                                         |
            +--------------------+--------------------+
                                 |  the sum, minus 7 lambda = 0.182
                                 v
               q(0.026) = -0.072 + 0 - 0.182 = -0.254

  the tie at site 2 is why 0.026 is a kink of q; by Theorem 3.7.3
  the LP of vex v_1 + vex v_2 over the knapsack polytope gives the
  same value at (1, 0.8): -0.15 + 0.8 (-0.13) = -0.254
The dual function of the crew row on the two-sites instance. Top: q(λ) = q₁(λ) + q₂(λ) − 7λ for λ from 0 to 0.08, the minimum of finitely many affine functions and so concave and piecewise linear (Proposition 3.7.2), with pieces of slopes 4, 1, −4 and −7. Each kink is a price at which one site changes its number of lines: site 1 from two to one at λ = 1/300, site 2 from one to none at λ* = 0.026, site 1 from one to none at 0.05. The maximum is d* = q(0.026) = −0.254; the dashed line through the reader's point has the slope g = 3y₁ + 5y₂ − 7 of the blocks' choices and lies above the curve, the supergradient inequality. Every value is a lower bound on z* = −0.16, and the ladder z_LP = −0.3553 < d* < z_LD = −0.215 < z* is drawn across. Bottom: each site's priced values at 0, 1 and 2 lines, the line's cost plus λ for each technician it needs, times y, less π(0.2y), the least lit; at λ* = 0.026 site 2's values at no line and one line are both 0, the tie that makes 0.026 a kink. The case λ = 0.0725, where both sites shut down and q = −0.5075, is where the subgradient run in the table of further methods overshoots.

(The split bound) The Lagrangian decomposition copies \(y\) into the crew row and keeps the integrality of the copy. The hull of the four feasible crew assignments is the triangle \(\{y \ge 0 : y_1 + 2y_2 \le 2\}\). Minimizing \(\operatorname{vex} v_1 + \operatorname{vex} v_2\) over it: along the edge \(y_1 + 2y_2 = 2\) the value is \(-0.13 - 0.085\,y_1\) for \(y_1 \le 1\) and \(-0.27 + 0.055\,y_1\) for \(y_1 \ge 1\), both minimized at \(y_1 = 1\), \(y_2 = 0.5\): \(z_{\mathrm{LD}} = -0.15 - 0.065 = -0.215\).

What each bound sees on the two-sites instance, one bound at a time. The McCormick LP and the Lagrangian relaxation of the crew row minimize over the knapsack polygon 0 ≤ y ≤ 2, 3y₁ + 5y₂ ≤ 7, with corners (0, 0), (2, 0), (2, 0.2) and (0, 1.4); the Lagrangian decomposition minimizes over the triangle (0, 0), (2, 0), (0, 1), the hull of the four feasible crew assignments. z_LP = −0.3553 at (4/3, 0.6), z_LR = −0.2540 at (1, 0.8), z_LD = −0.2150 at (1, 0.5) and z* = −0.1600 at (2, 0); (1, 1) needs eight technicians. The ladder's gaps are 0.101 (Corollary 3.7.4), 0.039 (Theorem 3.7.6(iv) fails) and 0.055 (Proposition 3.7.7).

(The ladder, gap by gap) The ladder is \(-0.3553 < -0.2540 < -0.2150 < -0.1600\). The McCormick relaxation loses \(0.101\) to convexifying the product and the integrality term by term (Corollary 3.7.4). The plain Lagrangian relaxation loses a further \(0.039\) to the linear relaxation of the crew knapsack: Theorem 3.7.6(iv) fails because \(\operatorname{conv}(Y)\) is strictly inside the knapsack polytope. The decomposition's remaining gap of \(0.055\) is the one fractional block that a single coupling constraint allows (Proposition 3.7.7), here half a line at site 2. The program below computes every number. Each bound is the minimum of a separable convex piecewise-linear function over a polygon. Such a function is affine on each cell between its breakpoints, so the minimum is attained at a vertex of a cell clipped to the polygon, which is what the routine enumerates. The McCormick LP is written through its closed form \(t_i \le \min(0.2\,y_i, 4/15)\), derived above, and the Lagrangian cuts of Theorem 3.7.11 are added to it in the last line.

# Lagrangian bounds on the two-sites example.
#
# Two production sites share one crew. Site i installs y_i in {0, 1, 2}
# lines of capacity 0.2 each (x_i <= 0.2 y_i), buys w_i in [0, 1] under
# 2 x_i + w_i <= 1.2, earns x_i w_i and pays c_i per line; a line needs
# a_i technicians and 7 are available:
#
#     minimize sum_i (c_i y_i - x_i w_i)  s.t.  a.y <= 7.
#
# Every bound below is the minimum of a separable convex piecewise-linear
# function F(y) = sum_i F_i(y_i) over a polygon; F is affine on each cell
# between breakpoints, so the minimum is at a vertex of a clipped cell,
# which is what minimize() enumerates. numpy only.

import itertools

import numpy as np

c = np.array([0.01, 0.03])      # cost of a line at sites 1 and 2
a = np.array([3.0, 5.0])        # technicians per line
b = 7.0                         # technicians in the crew
kappa = 0.2                     # capacity of a line

def pi(X):
    """Best profit x*w with x <= X.

    For fixed x the best w is min(1, 1.2 - 2x), and x min(1, 1.2 - 2x)
    increases up to x = 0.3.
    """
    x = min(X, 0.3)
    return x * min(1.0, 1.2 - 2 * x)

def v(i, y):
    """The block value function of site i along y_i."""
    return c[i] * y - pi(kappa * y)

def clip(poly, nrm, rhs):
    """Sutherland-Hodgman: keep the part of poly with nrm.p <= rhs."""
    out = []
    for p, q in zip(poly, poly[1:] + poly[:1]):
        fp, fq = nrm @ p - rhs, nrm @ q - rhs
        if fp <= 1e-12:
            out.append(p)
        if fp * fq < -1e-24:
            # the edge pq crosses the line: keep the crossing point
            out.append(p + (q - p) * fp / (fp - fq))
    return out

def minimize(pieces, poly):
    """Min over poly of sum_i max_k (s_ik y_i + t_ik).

    pieces[i] is the list of (slope, intercept) pairs of F_i. Returns the
    minimum and a vertex attaining it.
    """
    # breakpoints of each F_i: the pairwise crossings of its pieces
    brk = []
    for P in pieces:
        pts = {0.0, 2.0}
        for (s1, t1), (s2, t2) in itertools.combinations(P, 2):
            if abs(s1 - s2) > 1e-12:
                yb = (t2 - t1) / (s1 - s2)
                if 0 < yb < 2:
                    pts.add(yb)
        brk.append(sorted(pts))

    F = lambda y: sum(max(s * y[i] + t for s, t in pieces[i])
                      for i in range(2))

    # every cell [l0, u0] x [l1, u1] between breakpoints, clipped to poly
    best, arg = np.inf, None
    for (l0, u0), (l1, u1) in itertools.product(zip(brk[0], brk[0][1:]),
                                                zip(brk[1], brk[1][1:])):
        cell = poly
        for nrm, rhs in ((np.array([-1., 0]), -l0),
                         (np.array([1., 0]), u0),
                         (np.array([0, -1.]), -l1),
                         (np.array([0, 1.]), u1)):
            cell = clip(cell, nrm, rhs)
            if not cell:
                break
        for p in cell:
            if F(p) < best - 1e-12:
                best, arg = F(p), p
    return best, arg

# the knapsack polytope {0 <= y <= 2, 3y1 + 5y2 <= 7}
knap = [np.array(p) for p in [(0., 0.), (2., 0.), (2., 0.2), (0., 1.4)]]
# conv of the 4 feasible integer y
hull = [np.array(p) for p in [(0., 0.), (2., 0.), (0., 1.)]]

feas = [y for y in itertools.product(range(3), range(3)) if a @ y <= b]
zstar = min(v(0, y[0]) + v(1, y[1]) for y in feas)

# (1) McCormick LP: the relaxed block earns at most min(0.2 y_i, 4/15),
#     so F_i = max(c_i y - 0.2 y, c_i y - 4/15)
zLP, yLP = minimize([[(c[i] - 0.2, 0.0), (c[i], -4 / 15)]
                     for i in range(2)], knap)

# (2) Lagrangian relaxation of the crew row = convexified problem
#     (Theorem 3.7.3): F_i = vex v_i, the lower hull of v_i
vex = [[(v(i, 1) - v(i, 0), 0.0),
        (v(i, 2) - v(i, 1), v(i, 1) - (v(i, 2) - v(i, 1)))]
       for i in range(2)]
zLR, yLR = minimize(vex, knap)

# the dual function itself, on a grid
lam = np.arange(0, 0.0601, 1e-4)
q = [sum(min(v(i, y) + l * a[i] * y for y in range(3)) for i in range(2))
     - b * l
     for l in lam]
lstar = lam[int(np.argmax(q))]

# (3) Lagrangian decomposition: the same envelopes over the hull of the
#     feasible integer points
zLD, yLD = minimize(vex, hull)

# (4) the McCormick LP with the two Lagrangian block cuts
#     (c_i + lambda a_i) y_i - t_i >= q_i(lambda*)
qi = [min(v(i, y) + lstar * a[i] * y for y in range(3)) for i in range(2)]
zLPc, yLPc = minimize([[(c[i] - 0.2, 0.0), (c[i], -4 / 15),
                        (-lstar * a[i], qi[i])]
                       for i in range(2)], knap)

print(f"feasible integer y: {feas};  z* = {zstar:+.4f}")
print(f"McCormick LP relaxation           z_LP = {zLP:+.4f} "
      f"at y = ({yLP[0]:.4f}, {yLP[1]:.4f})")
print(f"Lagrangian relaxation (crew row)  z_LR = {zLR:+.4f} "
      f"at y = ({yLR[0]:.4f}, {yLR[1]:.4f})")
print(f"    max_lambda q(lambda) = {max(q):+.4f} at lambda* = {lstar:.4f}")
print(f"Lagrangian decomposition          z_LD = {zLD:+.4f} "
      f"at y = ({yLD[0]:.4f}, {yLD[1]:.4f})")
print("LP + Lagrangian block cuts")
print(f"    {zLPc:+.4f} at y = ({yLPc[0]:.4f}, {yLPc[1]:.4f});  "
      f"q_1, q_2 = {qi[0]:+.4f}, {qi[1]:+.4f}")
feasible integer y: [(0, 0), (0, 1), (1, 0), (2, 0)];  z* = -0.1600
McCormick LP relaxation           z_LP = -0.3553 at y = (1.3333, 0.6000)
Lagrangian relaxation (crew row)  z_LR = -0.2540 at y = (1.0000, 0.8000)
    max_lambda q(lambda) = -0.2540 at lambda* = 0.0260
Lagrangian decomposition          z_LD = -0.2150 at y = (1.0000, 0.5000)
LP + Lagrangian block cuts
    -0.2540 at y = (0.6429, 1.0143);  q_1, q_2 = -0.0720, +0.0000

The program performs four minimizations of a piecewise-linear function over a polygon and 601 evaluations of the dual function on a grid, each evaluation two block solves by enumeration. The block solves within one evaluation are independent of each other, and in a larger instance they are the batch.

(The cuts in numbers) The last line is Proposition 3.7.12(a) in numbers. The two block cuts at \(\lambda^\star = 0.026\), \((0.01 + 0.078)\,y_1 - t_1 \ge -0.072\) and \((0.03 + 0.13)\,y_2 - t_2 \ge 0\) in the McCormick variables, lift the LP from \(-0.3553\) to exactly \(-0.2540\), the Lagrangian bound. The LP's optimal face with the cuts is flat, and the program lands on a different point of it than a simplex method would. A dense two-phase simplex run for this series reports \((2, 0.2)\), and both points have value \(-0.2540\). The cuts carried the blocks' exact values into the relaxation without changing its constraint set.

The block cut as a supporting line, one panel per site. Blue dots: the block's values v₁ = (0, −0.15, −0.16) and v₂ = (0, −0.13, −0.12) at 0, 1 and 2 lines; orange: their lower hulls vex vᵢ and, dashed, the McCormick block objective cᵢyᵢ − min(0.2yᵢ, 4/15); purple: the Lagrangian block cut (cᵢ + aᵢλ) yᵢ − tᵢ ≥ qᵢ(λ) of Theorem 3.7.11, drawn as the line qᵢ(λ) − aᵢλyᵢ, which touches vex vᵢ at the block's minimizers (rings). The purple wash is what the cut adds above the McCormick objective. At λ* = 0.026 the cuts are (0.01 + 0.078) y₁ − t₁ ≥ −0.072 and (0.03 + 0.13) y₂ − t₂ ≥ 0, and together they lift the McCormick LP from −0.3553 to exactly −0.2540 (Proposition 3.7.12(a)).

A Lagrangian branch and bound on \(y\), with the convexified LP of Theorem 3.7.3 on the node's integer ranges as its node bound, closes the instance in five nodes.

The Lagrangian branch and bound on y, node by node. The bound at a node is the convexified LP of Theorem 3.7.3 on the node's integer ranges, drawn as the node's box and the part of the knapsack polygon 3y₁ + 5y₂ ≤ 7 inside it. Node 0 gives −0.2540 at (1, 0.8) and branches on y₂; node 1 (y₂ = 0) gives −0.1600 at (2, 0), integral and tight, the incumbent; node 2 (y₂ ∈ {1, 2}) gives −0.2300 at (2/3, 1) and branches on y₁; node 3 (y₁ = 0) gives −0.1300 ≥ −0.16 and is pruned by bound; node 4 (y₁, y₂ ≥ 1) is infeasible, 3 + 5 > 7.
node\(y_1\) range\(y_2\) rangeboundat \(y\)action
0\(\{0, 1, 2\}\)\(\{0, 1, 2\}\)−0.2540\((1, 0.8)\)branch on \(y_2\): \(y_2 \le 0 \mid y_2 \ge 1\)
1\(\{0, 1, 2\}\)\(\{0\}\)−0.1600\((2, 0)\)integral and tight: incumbent −0.16
2\(\{0, 1, 2\}\)\(\{1, 2\}\)−0.2300\((2/3, 1)\)branch on \(y_1\): \(y_1 \le 0 \mid y_1 \ge 1\)
3\(\{0\}\)\(\{1, 2\}\)−0.1300\((0, 1)\)pruned by bound: \(-0.13 \ge -0.16\)
4\(\{1, 2\}\)\(\{1, 2\}\)infeasible \(3 + 5 > 7\): pruned
The Lagrangian branch and bound on y as a tree (five nodes)

                    (0) -0.2540 at (1, 0.8)
                        branch on y_2
           y_2 <= 0  /                  \  y_2 >= 1
                    /                    \
  (1) -0.1600 at (2, 0)            (2) -0.2300 at (2/3, 1)
      integral and tight:              branch on y_1
      incumbent -0.16         y_1 <= 0  /          \  y_1 >= 1
                                       /            \
                   (3) -0.1300 at (0, 1)          (4) infeasible:
                       pruned by bound:               3 + 5 > 7,
                       -0.13 >= -0.16                 pruned

  the bound at a node is the convexified LP of Theorem 3.7.3 on
  the node's integer ranges; no branching on x_i or w_i

(Why no spatial branching is needed) No spatial branching on \(x_i\) or \(w_i\) is needed, because the blocks are solved exactly inside the bound. A tree built on the McCormick relaxation would have had to branch on the continuous variables as well. A second numpy script run for this series also ran the two dual-ascent methods of this subsection and the nonconvex Benders loop on the instance, and the table summarizes the runs.

methodevaluations of \(q\)result
subgradient ascent, Polyak step (Theorem 3.7.8)200best value −0.254003; the second step overshoots to \(\lambda = 0.0725\), where both sites shut down and \(q = -0.5075\); the iterates at steps 10 and 50 are worse than the one at step 3
Kelley's method on \(q_{\mathrm{LD}}\) (Proposition 3.7.9)7exact maximum −0.2150; the first two steps land on the boundary of the box \([-1, 1]^2\), because the model built from one piece is unbounded in the supergradient's direction
nonconvex Benders (Algorithm 3.7.15)4 iterationsvisits all four points of \(Y\), each needing a global block solve; \(U = -0.16\) after the second iteration
Three further methods on the two-sites instance (a numpy script run for this series)

(What the three runs show) Each run illustrates a result above. The subgradient run shows why Theorem 3.7.8 reports the best value and not the last. Kelley's run stops after seven evaluations, as Proposition 3.7.9 says it must on a function with at most \(9 \times 4 = 36\) pieces. The Benders run visits every point of \(Y\) because the McCormick relaxation is too weak to rule any out without a global solve (the exception, \((0, 0)\), whose relaxed value \(0\) is exact, is visited before an incumbent exists), which is the behaviour Li, Chen and Barton describe and their piecewise relaxations are meant to fix.The numpy script run for this series (ex_lagrangian.py, with a dense two-phase simplex, 5 October 2026) prints the subgradient trace, the seven Kelley evaluations, the five-node tree and the four NGBD iterations in full. The numbers of the four bounds agree with the program above to the printed digits.

Which points the relaxation rules out: nonconvex Benders on the two-sites instance, the McCormick relaxed value against the global value at the four points of Y, with step 5 of Algorithm 3.7.15 tested against the final incumbent U = −0.16, which the global solve at (2, 0) supplied. The rings show that test, not what the run did: in the run all four points needed a global block solve. The relaxed values are computed from the section's closed form tᵢ ≤ min(0.2 yᵢ, 4/15) with y fixed.

Where this is used

The general-purpose global solvers of Section 5 run one spatial tree on the full problem and do not decompose it. Lagrangian duality enters them only through reduced-cost tightening and the Lagrangian variable bounds of Section 2.6. Decomposition is done outside the solver, with the solver as the engine for the blocks: Karuppiah and Grossmann used BARON for the subproblems and CPLEX for the relaxations. In stochastic integer programming the standard method is the dual decomposition of Carøe and Schultz: Lagrangian relaxation of the nonanticipativity constraints, one subproblem per scenario, and a tree on the linking variables. Kim and Zavala's DSP is its open parallel implementation. Boland and coauthors compute the same dual bound by a different route. They combine progressive hedging, an augmented-Lagrangian method that alternates scenario solves with an averaging of the scenario copies of the linking variables, with a Frank–Wolfe method, a convex minimization method that works with linearizations. In power systems the unit-commitment problem has been solved by Lagrangian relaxation of the demand constraints, one generating unit per block, since the early 1980s. There the duality gap shrinks relative to the objective as the number of units grows, which is the Shapley–Folkman phenomenon of Section 5.4 in its original application. The separable–affine form of the tax-lot problem of Section 9 is (3.7.1) with the factor model as the coupling.C. C. Carøe and R. Schultz, "Dual decomposition in stochastic integer programming", Operations Research Letters 24 (1999); K. Kim and V. M. Zavala, "Algorithmic innovations and software for the dual decomposition method applied to stochastic mixed-integer programs", Mathematical Programming Computation 10 (2018); N. Boland, J. Christiansen, B. Dandurand, A. Eberhard, J. Linderoth, J. Luedtke and F. Oliveira, "Combining progressive hedging with a Frank–Wolfe method to compute Lagrangian dual bounds in stochastic mixed-integer programming", SIAM Journal on Optimization 28 (2018); R. T. Rockafellar and R. J.-B. Wets, "Scenarios and policy aggregation in optimization under uncertainty", Mathematics of Operations Research 16 (1991), for progressive hedging; Bertsekas, Lauer, Sandell and Posbergh (1983), cited above, for unit commitment. Y. Cao and V. M. Zavala, "A scalable global optimization algorithm for stochastic nonlinear programs", Journal of Global Optimization 75 (2019), is a recent decomposition in the same spirit; the texts of these papers beyond their abstracts were not consulted for this series.

What parallelizes

The GPU reading is the one promised at the start. A dual iteration has three parts: the block solves, the reduction that forms \(q(\lambda)\) and its supergradient, and the multiplier update, a vector operation for the subgradient method or a small QP for a bundle method. The block solves are independent, identical in structure when the blocks are scenarios of one model, and small. That is the shape a GPU batches well. Each subproblem gets one thread block, a group of device threads that execute a kernel together, or one lane, a single thread's slot in a batched kernel. All of them run the same instruction stream. The data is laid out block-major, so that neighbouring threads read neighbouring addresses, which is the coalesced access of Section 7.1. The reduction is one parallel sum. The multiplier update is a short sequential step between batches. In a tree the block subproblems of all open nodes of a frontier can be batched together, frontier times blocks lanes, with Proposition 3.7.12(b) removing the lanes whose parent solution is still valid in the child. The Lagrangian cuts are then a clean device-to-host interface: the device produces block values, the host's relaxation consumes linear rows. The sequential part is the multiplier update, the bundle or Benders master, the node relaxation with cuts (unless it is itself solved by the batched first-order method of Section 7.4) and the branching.

A dual iteration on a device: frontier x blocks lanes

  device        block 1     block 2     ...     block n
              +-----------+-----------+-----+-----------+
  open node 1 |   lane    |   lane    | ... |   lane    |
              +-----------+-----------+-----+-----------+
  open node 2 |   lane    |   lane    | ... |   lane    |
              +-----------+-----------+-----+-----------+
      ...     |    ...    |    ...    | ... |    ...    |
              +-----------+-----------+-----+-----------+
              one subproblem per lane (or per thread block),
              the same instruction stream in every lane
                                |  block values
                                v
              one parallel sum: q(lambda) and its supergradient
                                |
  - - - - - - - - - - - - - - - | - - - - - - - - - - - - - - - -
  host                          v  the Lagrangian cuts: linear rows
              multiplier update (a vector operation, or a small QP),
              the master, the node relaxation with the cuts
              (unless solved by the batched first-order method of
              Section 7.4), the branching: the sequential part
                                |
                                +--> new multipliers: the next batch

  a lane whose parent solution is still valid in the child is
  removed (Proposition 3.7.12(b))

(ExaTron, and two cautions) The one published instance of this pattern at scale is ExaTron, which decomposes an AC optimal power flow problem by grid components into many small nonconvex nonlinear programs and solves them in batch on the device. Its authors report linear scaling in the batch size and the number of GPUs and "more than 35 times speedup on 6 GPUs than on 40 CPU cores available on a single node" on the Summit supercomputer. A companion paper demonstrates the scheme on a 70,000-bus system.Y. Kim, F. Pacaud, M. Schanen, K. Kim and M. Anitescu, "Leveraging GPU batching for scalable nonlinear programming through massive Lagrangian decomposition", SIAM Journal on Scientific Computing 47 (2025), B1133–B1157, abstract (authors' numbers); the arXiv preprint 2106.14995 (2021) has four authors. Y. Kim and K. Kim, "Accelerated computation and tracking of AC optimal power flow solutions using GPUs", ICPP Workshops 2022, abstract. Two cautions fix what this does and does not show. First, the subproblems there are solved locally, by a trust-region Newton method on bound-constrained nonconvex problems, so the scheme is a heuristic for the nonconvex problem rather than a global method. A Lagrangian bound is valid only if every block value is a global minimum or a certified lower bound (Proposition 3.7.2). A batched local solver gives neither, and no batched global block solver exists. What carries over to nonconvex MINLP is the batching pattern, not the guarantee, and Section 7.8 lists the batched block solver as an open item. Second, the same argument that makes any dual value a valid bound makes the scheme tolerant of inexact dual ascent. A multiplier update stopped early, computed in single precision or taken from a stale iterate still yields a bound, provided the block solves themselves are exact and the sum of block values is rounded toward \(-\infty\). This is the decomposition analogue of the safe dual bound from an inexact LP dual of Section 7.3. It is the reason decomposition is the most forgiving place to put a device. The multipliers may be inexact, the block values must be exact, and the block solves are the part that is batched.

← Back to all posts