4. Reformulation: the relaxation is a design

Draft

The relaxation figure of Section 2.1 drew two polygons around one set of integer points and showed that the same integers can carry two different bounds. Section 3 mostly took the formulation as given: the search solved the relaxation it was handed, tightened the box it was handed, and branched when the bound was not enough. The comparison of the p, q and pq formulations of the pooling problem in Section 3.5 and the variable splitting of Section 3.7 were previews of this section. This section is about the other lever. The relaxation is a property of the formulation, not of the problem, and a formulation is something the modeller chooses. Two formulations with identical feasible sets can differ by orders of magnitude in the size of the tree a solver needs. The reason is that the solver's time goes into solving relaxations, and the formulation determines the relaxation.

One set of integer points, two formulations: what the choice changes

                    one set of integer points
                  /                           \
                v      the modeller chooses     v
          formulation 1                   formulation 2
                |    identical feasible sets    |
                v                               v
           relaxation 1                    relaxation 2
          and its bound                   and its bound
                |                               |
                v                               v
     the tree a solver needs         the tree a solver needs

  the formulation determines the relaxation, and the solver's time
  goes into solving relaxations: the two trees can differ in size
  by orders of magnitude

The section proceeds from the general to the particular. Section 4.1 recalls what it means for a formulation to be strong, and sets out the ladder of relaxations that the rest of the section climbs. The rungs of that ladder are linear programs, second-order cone programs and semidefinite programs. A second-order cone program is a convex program whose constraints bound the Euclidean norm of an affine expression in the variables by another affine expression. A semidefinite program is a convex program whose variable is a matrix required to be positive semidefinite. Section 4.2 treats disjunctions, the "either this or that" structure that every binary variable, every branch and every piecewise relaxation is an instance of, and proves Balas' theorem on the convex hull of a union of polyhedra. Section 4.3 derives the perspective reformulation of an indicator constraint (Section 2.5) from that theorem. Sections 4.4 and 4.5 apply it to the quadratic problems of portfolio construction, first when the quadratic is separable and then when it is not. The later subsections treat piecewise-linear models, the lifting hierarchy of RLT and semidefinite relaxations, conic representability, reformulation as an algorithm in its own right, and what can be written as a mixed-integer convex program at all. The section closes with the ladder in one table, filled with the numbers computed in this series. Minimization is the default in every display. The disjunction figure and its worked example maximize, and the text says so where they appear.

Formulation strength

This subsection recalls the vocabulary the section uses to compare formulations, extends it from polyhedra to convex sets, and explains why the comparison matters beyond the root bound. The reader has seen the phenomenon in the relaxation figure of Section 2.1: the polygon \(P\), the weaker polygon \(W\) and the integer hull \(H\), the convex hull of the integer points, describe the same twenty-two integer points, and their linear relaxations give three different bounds. Definition 2.1.6 (formulation, sharp, ideal) made "better" precise, following Vielma's survey, and Proposition 2.1.7 proved that an ideal formulation is sharp when the relaxation is bounded and that the two notions coincide when there are no auxiliary variables.J. P. Vielma, "Mixed integer linear programming formulation techniques", SIAM Review 57 (2015), Definitions 2.1 to 2.3 and Propositions 2.4 and 2.5. The notions go back to R. G. Jeroslow and J. K. Lowe, "Modelling with integer variables", Mathematical Programming Studies 22 (1984), and, for "locally ideal", to M. Padberg, "Approximating separable nonlinear functions via mixed zero-one programs", Operations Research Letters 27 (2000).

(Three extensions of Definition 2.1.6) Three additions are needed here. First, the formulation \(Q\) of Definition 2.1.6 was a polyhedron. In this section it may be any closed convex set in the variables \((x, y, w)\), since the perspective and conic formulations of Sections 4.3 to 4.5 are not polyhedral, and its continuous relaxation drops the integrality of \(y\) and of the integer part of \(w\). The two definitions read unchanged: sharp means that the projection of the relaxation onto \((x, y)\) is \(\operatorname{conv}(S)\), and ideal means that every extreme point of the relaxation satisfies the integrality requirements. Second, the projection is written \(\operatorname{proj}_{(x, y)} Q = \{(x, y) : \exists\, w \text{ with } (x, y, w) \in Q\}\). A linear objective in \((x, y)\) minimized over \(Q\) gives the same value as the same objective minimized over the projection, so the projection is what the bound sees. Third, Padberg's term for ideal is locally ideal, the right word when \(S\) is one substructure of a larger model, and Vielma's definition also asks that \(Q\) have at least one extreme point. Proposition 2.1.7(a) assumed a bounded relaxation. The hull formulations of Section 4.2 may be unbounded, and the following proposition removes the assumption.

Proposition 4.1.1 (ideal implies sharp for unbounded rational polyhedra). Let \(Q\) be a formulation of \(S\) in the sense of Definition 2.1.6 whose relaxation \(Q_{\mathrm{LP}}\) is a rational polyhedron with at least one extreme point. If every extreme point of \(Q_{\mathrm{LP}}\) satisfies the integrality requirements, then \(\operatorname{proj}_{(x, y)} Q_{\mathrm{LP}} = \operatorname{conv}(S)\).

Proof sketch. By the Minkowski–Weyl theorem, which represents a polyhedron as the sum of a polytope and a cone, \(Q_{\mathrm{LP}} = \operatorname{conv}(\operatorname{ext} Q_{\mathrm{LP}}) + \operatorname{rec}(Q_{\mathrm{LP}})\). Here \(\operatorname{rec}(Q_{\mathrm{LP}}) = \{d : q + t d \in Q_{\mathrm{LP}} \text{ for all } q \in Q_{\mathrm{LP}} \text{ and } t \ge 0\}\) is the recession cone, the set of directions along which one can move inside the polyhedron for ever, so every point of \(Q_{\mathrm{LP}}\) is a convex combination of extreme points plus a recession direction. The extreme points satisfy the integrality requirements, so they lie in \(Q\) and project into \(S\). The recession cone of a rational polyhedron is generated by finitely many rational directions, and each of these can be scaled to be integral in the integer coordinates. Let \(q \in Q\), let \(d\) be such a direction and \(s > 0\), and pick an integer \(N \ge s\). Then \(q + s d = (1 - s/N)\, q + (s/N)(q + N d)\), and both \(q\) and \(q + N d\) lie in \(Q\), because \(N d\) is integral in the integer coordinates. Applying this to one generator at a time shows that every point of \(Q_{\mathrm{LP}}\) is a convex combination of points of \(Q\). Projection commutes with convex combinations, which gives \(\operatorname{proj}_{(x, y)} Q_{\mathrm{LP}} \subseteq \operatorname{conv}(S)\). The reverse inclusion holds for every formulation, as in the proof of Proposition 2.1.7. ∎

Proposition 4.1.1, a step along a recession direction (illustrative polyhedron): q + s d = (1 − s/N) q + (s/N)(q + N d) is a convex combination of two points of Q, since N d is integral in the integer coordinates.
Proposition 4.1.1: a step along a recession direction

  Q_LP = conv(ext Q_LP) + rec(Q_LP), and the extreme points lie
  in Q; d is a generator of rec(Q_LP), scaled to be integral in
  the integer coordinates; s > 0, and N is an integer with N >= s

     q                 q + s d                         q + N d
     o--------------------*-------------------------------o---> d
     |<------ s d ------->|                               |
     |<---------------------- N d ----------------------->|
     in Q                                                 in Q

     q + s d  =  (1 - s/N) q  +  (s/N) (q + N d),
     a convex combination of two points of Q; one generator at a
     time, every point of Q_LP is a convex combination of points
     of Q

(Sharp without ideal, and why the difference matters) Proposition 2.1.7(c) says that without auxiliary variables sharp and ideal coincide. The converse of Proposition 4.1.1 fails once auxiliary variables are present. Section 4.6 exhibits a formulation of a piecewise-linear function, the convex-combination model with one binary per piece, that is sharp but not ideal. Section 4.2 computes a big-M formulation of three squares in the plane, a formulation that switches each alternative's constraints off with a large constant \(M\) times a binary (Definition 4.2.2), whose projection is exactly the hull while its relaxation has vertices with fractional binaries. The distinction matters in practice because a solver does more with the relaxation than read off its value. It branches on the fractional coordinates of the relaxed point, so fractional vertices mean nodes. It rounds the relaxed point in its heuristics, so fractional vertices mean rounding failures. And it propagates the relaxed point's bounds, so a vertex that is integral in the binaries fixes the integer variables for the whole subtree.

The relaxation figure of Section 2.1 is the smallest illustration, with the numbers \(5.56\), \(5.63\) and \(4.95\) computed there: the gap of \(P\) is \(10.9\%\) on the figure's bound-relative convention and \(12.2\%\) on the Gurobi/CPLEX convention of Definition 2.1.5.

(The smallest illustration, in the definitions' terms) Both \(P\) and \(W\) are formulations of the same twenty-two points with no auxiliary variables, and both are neither sharp nor ideal: their relaxations are larger than the hull and their vertices are fractional. In the notation of Definition 2.1.6, \(S\) is the set of twenty-two integer points, the formulation \(P\) has no auxiliary variable \(w\), and its relaxation \(Q_{\mathrm{LP}}\) is the polygon \(P\) itself, so the projection is \(P\). It is strictly larger than \(\operatorname{conv}(S) = H\), which is the failure of sharpness, and the vertex \((4.929, 2.929)\) at which the relaxation attains \(5.556\) at \(45°\) is fractional, which is the failure of idealness. The hull \(H\), with seven facets, is both: its relaxation is \(\operatorname{conv}(S)\) by construction, and its seven vertices, \((0, 0)\), \((4, 0)\), \((5, 1)\), \((5, 2)\), \((4, 3)\), \((2, 4)\) and \((0, 2)\), are integer points, so its relaxation's value equals the integer optimum at every angle. The figure's point, repeated throughout this section, is that nobody can list the facets of \(H\) in advance for a problem of any size. The cutting planes of Section 3.3 discover some of them during the solve. This section finds them by writing the problem differently, and the central device is to add variables rather than inequalities.

Three polygons, one set of integers: the polygon P of Section 2.1, the weaker polygon W and the integer hull H describe the same twenty-two integer points, and their linear relaxations give three bounds, each drawn as the objective's level line at that polygon's LP value, in the polygon's colour (dashed for P and W); the strip puts the three bounds on one axis with the gap of P shaded. Turn the objective to see that the hull's value equals the integer optimum at every angle.

(Three warnings: substructure, size, recognition) Three warnings go with the definitions. First, strength is a property of a formulation of a selected substructure. A sharp formulation of the graph of a function \(f\) stops being sharp for the set \(\{(x, f(x)) : x \in X\}\) as soon as \(X\) couples \(x\) to other variables through constraints. The reason is that the hull of an intersection is in general smaller than the intersection of hulls. One dimension shows it: for \(S_1 = \{0, 2\}\) and \(S_2 = \{1, 2\}\) the intersection is the single point \(2\), and so is its hull, while the hulls \([0, 2]\) and \([1, 2]\) intersect in the whole segment \([1, 2]\). A relaxation assembled from the hulls of the parts admits every point of that segment, and only the hull of the whole excludes them. Vielma, Ahmed and Nemhauser state this precisely for piecewise-linear models, and Section 4.2 shows the device, called a basic step, that buys back some of the difference at an exponential price.J. P. Vielma, S. Ahmed and G. Nemhauser, "Mixed-integer models for nonseparable piecewise-linear optimization: unifying framework and extensions", Operations Research 58 (2010), Theorem 4. Second, sharp and ideal describe the relaxation at the root. The number of variables and constraints a formulation uses determines the cost of every relaxation in the tree, and a stronger formulation that is also much larger can lose on time what it gains on nodes. The hull formulations of Section 4.2 multiply the continuous variables by the number of alternatives, and the folklore that they perform worse than their strength predicts is documented.R. Anderson, J. Huchette, W. Ma, C. Tjandraatmadja and J. P. Vielma, "Strong mixed-integer programming formulations for trained neural networks", Mathematical Programming 183 (2020); J. Kronqvist, R. Misener and C. Tsay, "P-split formulations: a class of intermediate formulations between big-M and convex hull for disjunctive constraints", Mathematical Programming 218 (2026), online 2025. Both papers name the observation and build intermediate formulations to escape it. Third, deciding whether a given formulation is sharp is itself hard: for the simplest big-M-type formulation of a union of polyhedra, recognizing sharpness is NP-hard.C. Blair, "Representation for multiple right-hand sides", Mathematical Programming 49 (1990), cited as Theorem 6.5 of Vielma (2015). That is one more reason the formulations of this section are proved sharp by construction, as Theorem 4.2.4 proves the hull formulation, rather than tested after the fact.

(The two axes of the ladder) The relaxations of this section form a ladder, and the ladder has two axes. One axis is strength: how close the projection of the relaxation comes to the convex hull of the set it describes. The other is cost: what kind of convex program the relaxation is, and how many variables it carries, since that is what a node of the tree pays. The table below names the rungs that this section and Section 2.4 climb. One row names two relaxations that Section 4.7 defines: Shor's semidefinite relaxation of a quadratic problem, and the Lasserre hierarchy of semidefinite relaxations of a polynomial problem. Section 4.11 fills the table in with the numbers computed in this series.

relaxationdescribescost of one relaxationequals the hull when
weak LP formulation (\(W\))an integer setLP, no new variablesin general never
sharp LP formulation (\(H\))an integer setLP, often exponentially many rowsby definition, at the root
big-M, smallest valid \(M\)\(x \in A\) or \(x \in B\) (a union)LP, one binary, no copiestwo polytopes in the plane; boxes with parallel facets
hull (Balas) formulation\(x \in A\) or \(x \in B\)LP, one copy of \(x\) per alternativealways, up to closure (Theorem 4.2.4)
perspectivean indicator \(z\) with a separable convex costSOCP, or LP plus perspective cutsthe separable substructure, under any constraint on \(z\)
diagonal split + perspectivean indicator \(z\) with a coupled convex quadraticSOCPnot in general; the split decides how close
rank-one and \(2 \times 2\) hullsthe sameSOCP or SDP in an extended space, \(O(n^2)\) variablesnot in general; tighter than the perspective
full hull (Theorem 4.5.6)the sameSDP, exponentially many vertices of \(P\)exactly
McCormick, RLT, Shor, Lasserreproducts and polynomialsLP, LP, SDP, growing SDPssee Sections 2.4 and 4.7
piecewise modelsa curve, to a toleranceMILP (a tree of LPs)the breakpoints (Section 4.6)
The ladder of relaxations: what each describes, what it costs, and when it is the hull.

Where this is used

Every reformulation below is implemented somewhere. Modelling layers such as Pyomo.GDP, the disjunctive-programming extension of the Pyomo algebraic modelling language for Python, turn a logical model into the big-M or the hull formulation on request. The presolve of a MILP solver strengthens the formulation automatically by coefficient tightening (Proposition 2.5.4). The nonlinear handlers of SCIP detect indicator structure and strengthen their own cuts by the perspective. The folk calibration of strength against size, the rule that a smaller formulation with a weaker bound often wins, was measured on simplex-based solvers, for which every extra column is a cost at every pivot.

What parallelizes

The calibration has another side. A first-order LP method of the kind in Section 7.2 costs one sparse matrix–vector product per iteration. It is indifferent to the number of variables, while it is sensitive to conditioning. The hull formulations of the next subsection are large and block-structured, so they are a natural candidate for a GPU LP solver, and nobody has yet measured the strength-versus-size trade-off on one.

Disjunctions and their hull

This subsection is about the one structure that almost every discrete decision in a MINLP reduces to: the requirement that a point lie in at least one of several convex sets. It states the theorem that describes the convex hull of such a union, and it compares the two ways of writing the union as a mixed-integer program. A binary variable is the disjunction \(x_j \le 0 \ \vee\ x_j \ge 1\), where \(\vee\) reads "or". A general integer variable branched at \(\beta\) is \(x_j \le \lfloor \beta \rfloor \ \vee\ x_j \ge \lceil \beta \rceil\). A spatial branch on a continuous variable is \(x_j \le \beta \ \vee\ x_j \ge \beta\), and the two children of a spatial branch-and-bound node are two convex relaxations whose union the search must bound. Throughout this section the binary variables that select an alternative are written \(z\), as the disjunctive-programming literature writes them, rather than the \(y\) of the notation paragraph, and the continuous variables remain \(x\).

The list of disjunctions in disguise is long. A fixed charge is one: a trade is zero or it pays the charge. A minimum trade size, a limit on the number of positions and a piecewise-linear function with its pieces are others. So are a set of operating regimes in a process model, a rectified linear unit in a trained network, the function \(\max\{0, t\}\) that is either zero or its argument, and a tax rule with a one-year boundary. Nonconvex MINLP adds one more source. Whenever a solver relaxes a nonconvex set piecewise, for instance a bilinear term by McCormick envelopes on each of several sub-boxes, the relaxation is a union of convex pieces. Its quality is the quality of the convexification of that union.

Disjunctive programming is Balas' theory of these unions. Its central fact is that the convex hull of a union of polyhedra has a short description in a lifted space: one copy of the variables per alternative, one weight per alternative, and each alternative's constraints scaled by its weight.E. Balas, "Disjunctive programming: properties of the convex hull of feasible points", Discrete Applied Mathematics 89 (1998), a reprint of Management Science Research Report 348, Carnegie Mellon (1974); E. Balas, "Disjunctive programming", Annals of Discrete Mathematics 5 (1979); E. Balas, "Disjunctive programming and a hierarchy of relaxations for discrete optimization problems", SIAM Journal on Algebraic and Discrete Methods 6 (1985); E. Balas, Disjunctive Programming (Springer, 2018). Chapter 4 of M. Conforti, G. Cornuéjols and G. Zambelli, Integer Programming, GTM 271 (Springer, 2014), is the textbook account.

Definition 4.2.1 (disjunctive set). Let \(P_i = \{x \in \mathbb{R}^n : A^i x \le b^i\}\), \(i \in D\), \(|D| = k\), be polyhedra. The disjunctive set is the union \(S = \bigcup_{i \in D} P_i\). A set written as an intersection of unions of polyhedra, \(F = \bigcap_{j \in T} \bigcup_{i \in D_j} P_{ji}\), is in regular form, Balas' conjunctive normal form. A set written as a single union is in disjunctive normal form. The same definitions apply when each \(P_i\) is a closed convex set \(\{x : g_i(x) \le 0\}\) with \(g_i\) convex, and Section 4.3 needs that case.

Definition 4.2.2 (big-M formulation). Attach a binary \(z_i\) to each alternative, require \(\sum_{i \in D} z_i = 1\), and relax each row of each alternative by a constant that switches off when \(z_i = 0\):

\[A^i x \le b^i + M^i (1 - z_i) \quad (i \in D), \qquad \sum_{i \in D} z_i = 1, \qquad z \in \{0, 1\}^k .\]

Let \(a^i_l\) be row \(l\) of \(A^i\), and let \(h_{P}(a) = \max\{a^\top x : x \in P\}\) be the support function of a set \(P\) in the direction \(a\), the most that \(a^\top x\) can reach on \(P\). For the formulation to be valid, the constant on row \(l\) of alternative \(i\) must satisfy \(M^i_l \ge \max_{j \ne i} h_{P_j}(a^i_l) - b^i_l\): it must cover the most that any other alternative violates the row. The smallest valid constants take equality. Since \(b^i_l + M^i_l\) is then \(\max_{j \ne i} h_{P_j}(a^i_l)\), which can be smaller than \(b^i_l\), the smallest valid "big" \(M\) can be negative. A negative constant tightens the row for the other alternatives and is valid. The textbook convention clamps every constant at zero, which is valid but weaker.Vielma (2015), Proposition 6.1.

Definition 4.2.3 (hull formulation). With one copy \(\nu^i \in \mathbb{R}^n\) of the variables and one weight \(\lambda_i\) per alternative,

\[x = \sum_{i \in D} \nu^i, \qquad A^i \nu^i \le b^i \lambda_i \ \ (i \in D), \qquad \sum_{i \in D} \lambda_i = 1, \qquad \lambda \ge 0 . \tag{4.2.1}\]

With \(\lambda \in \{0, 1\}^k\) this is a formulation of \(\bigcup_i P_i\) in the sense of Definition 2.1.6, with \(\nu\) the continuous and \(\lambda\) the integer auxiliaries. With \(\lambda \in [0, 1]^k\) it is the hull relaxation. The names "disaggregated" and "extended" formulation are also used. The letter \(\lambda\) denotes a convex weight here and in Sections 4.3 and 4.5, as it does in Balas' papers, and not a Lagrange multiplier as it does elsewhere in this series.

(The worked example: two polygons, one objective) The smallest example that shows the difference between the two formulations can be checked by hand. Let \(P_1 = [0, 1]^2\) and \(P_2 = \{x : x_1 + x_2 \ge 3,\ 0 \le x \le 2\}\), a triangle with vertices \((1, 2)\), \((2, 1)\) and \((2, 2)\), and maximize \(x_1 - x_2\) over \(P_1 \cup P_2\). This example maximizes, as the figure below does. The maximum is \(1\), attained at \((1, 0)\) in \(P_1\) and at \((2, 1)\) in \(P_2\).

(The hull relaxation at weight one half) Take the hull formulation first. In the notation of (4.2.1), with \(D = \{1, 2\}\), \(\lambda_1 = \lambda\) and \(\lambda_2 = 1 - \lambda\), the block of \(P_1\) is \(0 \le \nu^1 \le \lambda\,(1, 1)\), the block of \(P_2\) is \(\nu^2_1 + \nu^2_2 \ge 3 (1 - \lambda)\) with \(0 \le \nu^2 \le 2 (1 - \lambda)\,(1, 1)\), and the linking row is \(x = \nu^1 + \nu^2\). Fix the weight \(\lambda = \tfrac12\). Then \(\nu^1 \in \tfrac12 P_1\) and \(\nu^2 \in \tfrac12 P_2\), so \(x = \nu^1 + \nu^2\) ranges over \(\tfrac12 P_1 + \tfrac12 P_2\), the set of midpoints of a point of \(P_1\) and a point of \(P_2\). That set is the square \([\tfrac12, \tfrac32]^2\) cut by \(x_1 + x_2 \ge \tfrac32\), and on it \(x_1 - x_2 \le 1\), with equality at \((\tfrac32, \tfrac12)\), the midpoint of \((1, 0)\) and \((2, 1)\). The same holds at every weight. The hull relaxation at weight \(\lambda\) is the Minkowski combination \(\lambda P_1 + (1 - \lambda) P_2\), the set of all sums of a point of \(\lambda P_1\) and a point of \((1 - \lambda) P_2\); the union of these sets over \(\lambda \in [0, 1]\) is \(\operatorname{conv}(P_1 \cup P_2)\) (Theorem 4.2.4 below), and \(x_1 - x_2 \le 1\) is a facet of that hull. So the hull formulation gives \(1\) at the root.

(The big-M relaxation, with and without negative constants) The textbook big-M formulation keeps the shared bounds \(0 \le x\) as they are and relaxes the other rows by the smallest nonnegative constants: \(x_i \le 2 - z\) and \(x_1 + x_2 \ge 3(1 - z)\), with \(z = 1\) selecting \(P_1\). In the notation of Definition 4.2.2 the row \(x_1 \le 1\) of \(P_1\) has \(a = (1, 0)\), \(b = 1\) and \(h_{P_2}(a) = 2\), so its smallest valid constant is \(2 - 1 = 1\) and the relaxed row is \(x_1 \le 1 + (1 - z) = 2 - z\); the row \(x_1 + x_2 \ge 3\) of \(P_2\), written \(-x_1 - x_2 \le -3\), has \(h_{P_1}(-1, -1) = 0\), so its constant is \(0 + 3 = 3\) and the relaxed row is \(x_1 + x_2 \ge 3 (1 - z)\). These are the constants the script below prints. At \(z = \tfrac12\) it admits \((1.5, 0)\), so its bound is \(1.5\), a gap of \(50\%\) relative to the true optimum. Allowing the negative constants of Definition 4.2.2 changes the picture. The row \(x_1 \ge 0\) of \(P_1\) has smallest valid constant \(-1\), because \(P_2\) never has \(x_1 < 1\), and becomes \(x_1 \ge 1 - z\). With all such rows the relaxation at \(z = \tfrac12\) is the square \([\tfrac12, \tfrac32]^2\) cut by \(x_1 + x_2 \ge \tfrac32\), the same set the hull relaxation gave, over which \(x_1 - x_2 \le 1\). The formulation with the smallest valid constants is then sharp, as Proposition 4.2.7 below says it must be for two polytopes in the plane. Under the convention that keeps every constant nonnegative, which the figure below also follows, the bound is \(1.5\) and the gap \(50\%\). The script prints both. It computes each bound by enumerating the vertices of a polygon, slicing the big-M relaxation at 2,001 values of \(z\).

Big-M against the hull: P1 = [0, 1]² and the triangle P2. Maximizing x1 − x2 over P1 ∪ P2, the optimum is 1, at (1, 0) in P1 and at (2, 1) in P2, and x1 − x2 ≤ 1 is a facet of the hull. The hull relaxation at λ = 1/2 reaches 1 at (3/2, 1/2), the midpoint of (1, 0) and (2, 1). The textbook big-M relaxation (every M ≥ 0) at z = 1/2 reaches 1.5 at (1.5, 0), a gap of 50%. With the smallest valid constants, negative allowed, the big-M bound is 1, as for the hull.
Big-M against the hull at z = λ = 1/2: the two slices. The textbook big-M relaxation (every M ≥ 0) is x_i ≤ 2 − z = 3/2, x1 + x2 ≥ 3(1 − z) = 3/2, x ≥ 0. It admits (1.5, 0), so its bound is 1.5. The hull relaxation is (1/2) P1 + (1/2) P2, the square [1/2, 3/2]² cut by x1 + x2 ≥ 3/2, on which x1 − x2 ≤ 1, with equality at (3/2, 1/2). The big-M relaxation with the smallest valid constants is this same set.
# Big-M against the hull, on the two polygons of the worked example.
#
# P1 = [0,1]^2 and P2 = {x1 + x2 >= 3, 0 <= x <= 2}; maximize x1 - x2
# over P1 u P2. Every set is a polygon, so each bound is an exact vertex
# enumeration; the big-M projection is sliced in z. Two conventions for
# the constants: the textbook one keeps every M >= 0; Vielma's smallest
# valid M may be negative.

import numpy as np

c = np.array([1.0, -1.0])

# rows a.x <= r
P1 = [((1, 0), 1), ((-1, 0), 0), ((0, 1), 1), ((0, -1), 0)]
P2 = [((-1, -1), -3), ((1, 0), 2), ((-1, 0), 0), ((0, 1), 2),
      ((0, -1), 0)]

def vertices(rows):
    """The vertices of {x : a.x <= r for every row}, by pairs of rows."""
    pts = []
    for i in range(len(rows)):
        for j in range(i + 1, len(rows)):
            A = np.array([rows[i][0], rows[j][0]], float)
            if abs(np.linalg.det(A)) < 1e-12:
                continue
            rhs = np.array([rows[i][1], rows[j][1]], float)
            x = np.linalg.solve(A, rhs)
            if all(np.dot(a, x) <= r + 1e-9 for a, r in rows):
                pts.append(x)
    return np.array(pts)

def support(rows, a):
    """max a.x over the polygon."""
    return (vertices(rows) @ np.array(a, float)).max()

V1, V2 = vertices(P1), vertices(P2)
true = max((V1 @ c).max(), (V2 @ c).max())
print(f"true optimum over P1 u P2 = hull formulation's bound: {true:.3f}")
print("  (attained at (1, 0) in P1 and (2, 1) in P2)")

# how much the other polygon can violate each row
M1 = [support(P2, a) - r for a, r in P1]
M2 = [support(P1, a) - r for a, r in P2]
print("smallest valid M, rows of P1:", np.round(M1, 2))
print("                  rows of P2:", np.round(M2, 2))

def big_m_bound(M1, M2):
    """The best vertex of the big-M relaxation over 2,001 slices in z.

    Returns (bound, z, x).
    """
    best = (-np.inf, None, None)
    # z = 1 selects P1, z = 0 selects P2
    for z in np.linspace(0, 1, 2001):
        rows = ([(a, r + M * (1 - z)) for (a, r), M in zip(P1, M1)]
                + [(a, r + M * z) for (a, r), M in zip(P2, M2)])
        V = vertices(rows)
        if len(V) == 0:
            continue
        k = (V @ c).argmax()
        if V[k] @ c > best[0]:
            best = (V[k] @ c, z, V[k])
    return best

for name, m1, m2 in (("textbook, M >= 0",
                      np.maximum(M1, 0), np.maximum(M2, 0)),
                     ("smallest valid, negative allowed", M1, M2)):
    b, z, x = big_m_bound(m1, m2)
    print(f"big-M relaxation ({name}):")
    print(f"  bound {b:.3f} at z = {z:.2f}, "
          f"x = ({x[0]:.2f}, {x[1]:.2f}); gap {(b - true) / true:.0%}")
true optimum over P1 u P2 = hull formulation's bound: 1.000
  (attained at (1, 0) in P1 and (2, 1) in P2)
smallest valid M, rows of P1: [ 1. -1.  1. -1.]
                  rows of P2: [ 3. -1.  0. -1.  0.]
big-M relaxation (textbook, M >= 0):
  bound 1.500 at z = 0.50, x = (1.50, -0.00); gap 50%
big-M relaxation (smallest valid, negative allowed):
  bound 1.000 at z = 0.00, x = (2.00, 1.00); gap 0%

The script costs one vertex enumeration per slice, each a few dozen \(2 \times 2\) solves, and the slices are independent of one another. The row that closes the gap, \(x_1 \ge 1 - z\), is an implied bound of \(P_2\) propagated into the formulation. That is the observation of Belotti and coauthors: bound tightening is the lever that makes indicator constraints work in a MILP solver.P. Belotti, P. Bonami, M. Fischetti, A. Lodi, M. Monaci, A. Nogales-Gómez and D. Salvagnin, "On handling indicator constraints in mixed integer programming", Computational Optimization and Applications 65 (2016).

The figure below draws the same comparison for two boxes in the plane and lets the reader change the objective, the constants and the geometry. The thing to look for is the orange region, the projection of the big-M relaxation, lying outside the blue outline of the hull even at the smallest valid constants, and growing as the constants grow. The set \(A = [0.5, 2.5] \times [0.5, 1.5]\) is axis-aligned and selected by \(z = 1\). The set \(B\) is a \(2 \times 2\) box centred at \((4.5, 3.5)\) and turned by \(30^\circ\) by default, selected by \(z = 0\). Both formulations also carry the bounds \(0 \le x_1 \le 6\) and \(0 \le x_2 \le 5\) of the drawing, which the big-M region reaches when the constants are large. The tilt is there because of Proposition 4.2.7 and its boundary case. Two boxes with parallel facets have a big-M formulation with the smallest valid constants that is exactly the hull: every relaxed facet is a bound on one coordinate, and the bounds at a given \(z\) are exactly those of \(zA + (1 - z)B\). With aligned boxes there would be nothing to see. The figure clamps a negative constant at zero, as the textbook convention does, and the slider multiplies every constant by a factor between \(1\) and \(4\).

In the default view the objective points at \(130^\circ\). The true maximum over \(A \cup B\) is \(0.947\), at the corner \((3.13, 3.87)\) of \(B\). The hull relaxation reaches exactly \(0.947\) at the same point, as it does at every angle, because the hull's vertices are corners of the two boxes. The big-M relaxation with the smallest valid constants reaches \(1.444\) at \((0.65, 2.43)\), a point in neither box. The gap is \(52.5\%\), measured relative to the true optimum as the figure's statistic computes it, and the orange region has \(10\%\) more area than the hull. At the default tilt the smallest valid constants are \(3.37\) for the two facets \(x_1 \le 2.5\) and \(x_2 \le 1.5\) of \(A\), \(3.96\) and \(0.60\) for the two lower facets of \(B\), and zero for the other four, which the other box already satisfies.

The other cases scale the constants, and the numbers that follow are the figure's own readouts. With the constants doubled the bound is \(1.475\) at \((0.50, 2.35)\), a gap of \(55.8\%\), and the area is \(28\%\) beyond the hull. At four times the smallest constants the bound is still \(1.475\), because the growth in that direction is stopped by the facet \(x_1 \ge 0.5\) of \(A\), which needed no constant at all, and by the bounds of the drawing. The area, however, grows to \(49\%\) beyond the hull. With the tilt set to zero the orange and blue outlines coincide and all three values are \(1.197\). The reader who wants to see the constants matter most should point the objective down and to the right, between \(280^\circ\) and \(350^\circ\). At \(310^\circ\) the smallest constants are exact at that angle, while twice and four times the smallest overshoot by \(39\%\) and \(116\%\).

One logical constraint, x in A or x in B, written two ways and relaxed. The two blue boxes are the feasible set, the blue outline is their convex hull, which is all the hull formulation's relaxation allows, and the orange region is what the big-M formulation's relaxation allows, with every M scaled by the slider from its smallest valid value and the faint dashed outlines marking the big-M set at z = ¼, ½ and ¾. The blue dot is the true optimum, which the hull relaxation also reaches, and the orange ring is the big-M relaxation's optimum. Drag in the plot to turn the objective, and set the tilt of B under the assumptions.

The example suggests the general statement. The hull relaxation at weight \(\lambda\) is the Minkowski combination \(\lambda P_1 + (1 - \lambda) P_2\), and the union of these sets over \(\lambda\) is the convex hull of \(P_1 \cup P_2\). The theorem makes this precise for any number of polyhedra, bounded or not.

Theorem 4.2.4 (Balas, 1974; the hull of a union of polyhedra). Let \(P_i = \{x : A^i x \le b^i\}\), \(i \in D\), be nonempty polyhedra, and let \(E\) be the polyhedron (4.2.1) in the variables \((x, \nu, \lambda)\). Then

\[\operatorname{proj}_x E \;=\; \operatorname{cl}\operatorname{conv}\Big(\bigcup_{i \in D} P_i\Big) \;=\; \operatorname{conv}\Big(\bigcup_{i \in D} P_i\Big) + \sum_{i \in D} \operatorname{rec}(P_i),\]

where \(\operatorname{rec}(P_i) = \{d : A^i d \le 0\}\) is the recession cone of \(P_i\). If all the \(P_i\) have the same recession cone, in particular if all are bounded, then \(\operatorname{conv}(\bigcup_i P_i)\) is closed and equals \(\operatorname{proj}_x E\).

Proof sketch. For the inclusion from right to left, take a point \(x = \sum_i \lambda_i p^i\) of the convex hull, with \(p^i \in P_i\) and \(\lambda\) in the simplex. Put \(\nu^i = \lambda_i p^i\). Then \(A^i \nu^i = \lambda_i A^i p^i \le \lambda_i b^i\), so \((x, \nu, \lambda) \in E\). Adding a recession direction \(d\) of \(P_i\) to the block \(\nu^i\) keeps \(A^i \nu^i \le \lambda_i b^i\), so \(x + d\) is also in the projection, and so is \(x\) plus any sum of recession directions. For the inclusion from left to right, take \((x, \nu, \lambda) \in E\). Where \(\lambda_i > 0\), the point \(p^i = \nu^i / \lambda_i\) lies in \(P_i\). Where \(\lambda_i = 0\), the block satisfies \(A^i \nu^i \le 0\), so \(\nu^i \in \operatorname{rec}(P_i)\). Hence \(x = \sum_{\lambda_i > 0} \lambda_i p^i + \sum_{\lambda_i = 0} \nu^i\) lies in the convex hull plus the sum of the recession cones. For the closure, \(\operatorname{proj}_x E\) is the projection of a polyhedron, hence a polyhedron, hence closed, and it contains the convex hull, so it contains the closure of the hull. Conversely, take \(p \in P_j\), \(q \in P_i\), \(d \in \operatorname{rec}(P_i)\) and \(t \ge 0\). The point \((1 - \varepsilon) p + \varepsilon (q + d\, t / \varepsilon)\) lies in the convex hull for every \(\varepsilon \in (0, 1]\) and tends to \(p + t d\) as \(\varepsilon \downarrow 0\), so the hull plus the recession cones lies in the closure of the hull. Finally, if every \(\operatorname{rec}(P_i)\) equals one cone \(C\) and \(d \in C\), then \(\sum_i \lambda_i p^i + d = \sum_i \lambda_i (p^i + d)\) with \(p^i + d \in P_i\), so the Minkowski sum adds nothing and the hull is already closed. ∎

(What the lifting does, and why the closure matters) In a picture, a point of the hull is a weighted average of one point from each alternative. The constraint \(A^i p^i \le b^i\) on the point from alternative \(i\) is nonlinear in the pair (weight, point) once the weight is a variable. Multiplying the point by its weight makes it the linear constraint \(A^i \nu^i \le b^i \lambda_i\) in the new variables. The nonconvex instruction "choose one alternative" has become the convex instruction "average over the alternatives", and the homogenization, the multiplication of each alternative's constraints through by its own weight, is exactly what makes the averaged constraint linear. The closure is not a technicality. Take \(P_1 = \{(0, 1)\}\), a point, and \(P_2 = \{(t, 0) : t \ge 0\}\), a ray. Their convex hull is \(\{(a, b) : a \ge 0,\ 0 \le b < 1\} \cup \{(0, 1)\}\), which is not closed. The formulation (4.2.1) with \(\lambda_1 = 1\) and \(\lambda_2 = 0\) allows \(\nu^2 = (t, 0)\) for any \(t \ge 0\), and so produces the missing points \((t, 1)\). An empty alternative must be removed before the theorem is applied, since an empty \(P_i\) can still have a nontrivial cone \(\{A^i d \le 0\}\) that (4.2.1) would wrongly add. Both facts are in Balas (1998) and in Chapter 4 of Conforti, Cornuéjols and Zambelli.

The closure is not a technicality: a point and a ray. The convex hull of P1 = {(0, 1)} and P2 = {(t, 0) : t ≥ 0} is {a ≥ 0, 0 ≤ b < 1} ∪ {(0, 1)}, which is not closed. The formulation (4.2.1), with λ1 = 1, λ2 = 0 and ν² = (t, 0), produces the missing points (t, 1).

Proposition 4.2.5 (the hull formulation is ideal). For polyhedra with a common recession cone, (4.2.1) with \(\lambda\) binary is a formulation of \(\bigcup_i P_i\), and every extreme point of its relaxation \(E\) has \(\lambda \in \{0, 1\}^k\).Jeroslow and Lowe (1984); Balas (1985); Vielma (2015), Proposition 4.2.

Proof. Suppose \((x, \nu, \lambda) \in E\) has \(0 < \lambda_i, \lambda_j < 1\) for two indices \(i \ne j\). For a small \(\varepsilon > 0\) define two points by \(\lambda_i^\pm = \lambda_i \pm \varepsilon\), \(\nu^{i, \pm} = \nu^i (1 \pm \varepsilon / \lambda_i)\), \(\lambda_j^\pm = \lambda_j \mp \varepsilon\), \(\nu^{j, \pm} = \nu^j (1 \mp \varepsilon / \lambda_j)\), every other block unchanged, and \(x^\pm = \sum_l \nu^{l, \pm}\). Both points satisfy (4.2.1), because \(A^i \nu^{i, \pm} = (1 \pm \varepsilon / \lambda_i) A^i \nu^i \le (1 \pm \varepsilon / \lambda_i) \lambda_i b^i = \lambda_i^\pm b^i\) and likewise for \(j\), and the original point is their midpoint. So a point with two fractional weights is not extreme. ∎

(The hull formulation is sharp, and what it costs) By Proposition 2.1.7(a) when the \(P_i\) are bounded, and by Proposition 4.1.1 for rational data in general, the hull formulation is therefore also sharp, which is Balas' theorem read in the language of Section 4.1. The price is the copies: \(k n\) continuous variables in place of \(n\), and the equations \(x = \sum_i \nu^i\) that tie them together. The big-M formulation keeps \(n\) continuous variables and pays in strength, and the next two results say how much.

Proposition 4.2.6 (validity and finiteness of the constants). If the \(P_i\) have a common recession cone, then \(h_{P_j}(a^i_l)\) is finite for every \(i, j, l\), and the big-M formulation of Definition 4.2.2 with any constants satisfying the inequality there is a formulation of \(\bigcup_i P_i\).Vielma (2015), Proposition 6.1.

Proof. If the linear program \(\max\{a^{i\top}_l x : x \in P_j\}\) were unbounded, \(P_j\) would have a recession direction \(r\) with \(a^{i\top}_l r > 0\). The direction is also a recession direction of \(P_i\), which contradicts \(a^{i\top}_l x \le b^i_l\) on \(P_i\). Validity is immediate: at \(z_i = 1\) the rows of alternative \(i\) are enforced, and at \(z_i = 0\) each row of alternative \(i\) is relaxed by at least as much as the chosen alternative can violate it. ∎

Proposition 4.2.7 (two polytopes in the plane; folklore, proof given here). Let \(A, B \subset \mathbb{R}^2\) be polytopes with facet descriptions \(A = \{a_l^\top x \le \alpha_l\}\) and \(B = \{c_m^\top x \le \beta_m\}\), and consider the big-M formulation with the smallest valid constants, negative values included:

\[a_l^\top x \le z\, \alpha_l + (1 - z)\, h_B(a_l) \ \ \forall l, \qquad c_m^\top x \le (1 - z)\, \beta_m + z\, h_A(c_m) \ \ \forall m, \qquad z \in \{0, 1\}.\]

Then the projection onto \(x\) of its relaxation (\(z \in [0, 1]\)) equals \(\operatorname{conv}(A \cup B)\): the formulation is sharp.

Proof. Fix \(z \in (0, 1)\) and let \(R(z)\) be the set of \(x\) feasible at that \(z\). Since \(h_A(a_l) = \alpha_l\) for a facet normal of \(A\), the right-hand side of the first family is \(z\, h_A(a_l) + (1 - z)\, h_B(a_l) = h_{zA + (1 - z)B}(a_l)\), because the support function of a Minkowski combination is the combination of the support functions. The second family is the same with the roles exchanged. So \(R(z)\) is the intersection of the supporting half-spaces of the polygon \(M(z) = zA + (1 - z)B\) in the directions \(\{a_l\} \cup \{c_m\}\), and \(R(z) \supseteq M(z)\). In the plane every edge of a Minkowski sum of two convex polygons is a translate of an edge of one of the summands, because walking round the sum turns through the edge directions of the two polygons merged in angular order and through no other direction, so every facet normal of \(M(z)\) belongs to \(\{a_l\} \cup \{c_m\}\). A polytope is the intersection of its facet-defining half-spaces, so \(R(z) = M(z)\). Finally \(\bigcup_{z \in [0, 1]} (zA + (1 - z)B) = \operatorname{conv}(A \cup B)\), the standard description of the hull of a union of two convex sets. ∎

(Sharp is not ideal, and the plane is special) Two remarks bound the reach of this proposition. First, sharp is not ideal. For three squares in the plane with binaries \(z_1 + z_2 + z_3 = 1\), the big-M formulation with the smallest single constant per row has a projection equal to the hull, while its relaxation in \((x, z)\) has vertices with fractional \(z\). The multiple-parameter form of Proposition 4.2.14 below is sharp and ideal on the same squares with no extra variables. Second, the plane is special. Vielma's Example 12 takes two "roofs" in \(\mathbb{R}^3\), polytopes with the same five facet normals and different right-hand sides. Their hull has one facet, the cap \(x_3 \le 1\), whose normal belongs to neither polytope. The big-M projection with the smallest valid constants carries the other nine facets of the hull and misses the cap, so its volume exceeds the hull's by about \(4.5\%\), a small pyramid above the cap. Vielma notes that adding \(x_3 \le 1\) restores the hull. The hull formulation adds one copy of \(x\), three continuous variables, and one weight, and removes the excess.Vielma (2015), Example 12 and Proposition 6.4, which shows that for two polyhedra without redundant rows the big-M and the hybrid formulation have the same relaxation. The following numbers were computed in exact rational arithmetic on 5 October 2026. Three squares \(P_1 = [0, 1]^2\), \(P_2 = [3, 4] \times [0, 1]\), \(P_3 = [1.5, 2.5] \times [3, 4]\): hull area \(23/2\); single-constant big-M projection of area \(23/2\) with 16 relaxation vertices, two of them fractional, \(z = (0, \tfrac12, \tfrac12)\) and \((\tfrac12, 0, \tfrac12)\); multiple-parameter big-M projection of area \(23/2\) with 12 vertices, all integral. Example 12: \(A\) has the rows \((1, 0, 1)\), \((-1, 0, 1)\), \((0, 1, 1)\), \((0, -1, 1)\) and \((0, 0, -1)\), with \(b^1 = (1, 1, 2, 2, 0)\) and \(b^2 = (2, 2, 1, 1, 0)\); each roof has volume \(10/3\); the hull has 12 vertices, 10 facets and volume \(22/3\); the big-M projection has 13 vertices and volume \(23/3\); the excess is the pyramid \(\int_1^{3/2} 2(3 - 2t)^2\, dt = 1/3\) above the cap, with its apex at \((0, 0, \tfrac32)\), the point the relaxation reaches at \(z = \tfrac12\). And as Section 4.1 said, recognizing sharpness is NP-hard in general, so the plane is also the last place where one should expect a clean rule.

(The cost of loose constants) Loose constants have a measurable cost. For two axis-aligned boxes \([0, 2] \times [0, 1]\) and \([3, 4] \times [2, 4]\), the hull has area \(19/2\) and the big-M projection with the smallest valid constants has area \(19/2\) as well. With every constant set to \(100\), a value a modeller might type without thinking, the projection is a polygon of area \(10{,}295.5\) with vertices near \(\pm 50\), more than a thousand times the hull. The relaxation's vertices then sit far outside any meaningful region. This geometry is the source of the numerical trouble with large constants that Section 8 returns to: a \(z\) of \(10^{-6}\) multiplied by an \(M\) of \(10^6\) is a quantity of one unit that the solver treats as zero. Computing the smallest valid constants is cheap, and the two algorithms below give the big-M and the hull reformulations as a solver or a modelling layer would execute them.

Algorithm 4.2.8 (tight big-M constants).

Algorithm TIGHT-BIG-M

Input   alternatives P_i = {A^i x <= b^i} (or convex g_i(x) <= 0) inside
        a bounded region X.
Output  constants for the single-parameter (BM) or multiple-parameter
        (MBM, Proposition 4.2.14) big-M formulation.

1. For each pair i != j and each row l of alternative i:
      M^{ij}_l := max { a^i_l . x - b^i_l : x in P_j, x in X }
      (one LP; an NLP when P_j is nonlinear; an interval bound computed
      from the box X is a cheap valid overestimate)

2. BM :  M^i_l := max_{j != i} M^{ij}_l;
         row  a^i_l . x <= b^i_l + M^i_l (1 - z_i).
   MBM:  row  a^i_l . x <= b^i_l + sum_{j != i} M^{ij}_l z_j.

Invariant
    both are valid formulations (Proposition 4.2.6); MBM's relaxation
    is contained in BM's (Proposition 4.2.14); a negative constant is
    valid and tightens, and clamping it at 0 is the textbook
    convention, valid but weaker.

Cost
    sum_i m_i (k - 1) small LPs, all independent of one another.

Parallel
    embarrassingly parallel across (i, j, l); on a device, batch the
    LPs, or replace them by batched interval or FBBT bounds computed
    from the node's current box.

Algorithm 4.2.9 (hull reformulation of a disjunction).

Algorithm HULL-REFORMULATE

Input   a disjunction OR_{i in D} [ g_i(x) <= 0 ] over x in [l, u]
        (bounded), each g_i affine (A^i x <= b^i) or convex; a
        tolerance eps in (0, 1) for nonlinear alternatives; the set V_i
        of variables that appear in alternative i.
Output  mixed-integer (conic) constraints whose relaxation projects
        onto conv of the disjunction.

1. For each i create a binary z_i and copies nu^i_j for j in V_i
   (variables not in V_i are not copied; they are linked only
   through x).

2. Add sum_{i in D} z_i = 1.

3. For each j: add x_j = sum_{i : j in V_i} nu^i_j.

4. For each i and each j in V_i: add z_i l_j <= nu^i_j <= z_i u_j.

5. For each i:
      if g_i is affine:
         add A^i nu^i <= b^i z_i.
      elif g_i is conic-representable (e.g. a convex quadratic):
         add its exact perspective in conic form, e.g.
         t_i z_i >= ||nu^i||^2 as a rotated second-order cone
         (Section 4.3).
      else:
         add ((1 - eps) z_i + eps) g_i( nu^i / ((1 - eps) z_i + eps) )
                - eps g_i(0) (1 - z_i) <= 0
         (Furman, Sawaya and Grossmann; Theorem 4.2.12).

Invariant
    for z in {0,1}^D the constraints are equivalent to the
    disjunction, exactly and for every eps; for z in [0,1]^D the
    projection onto x is conv of the union when step 5 is exact, and
    a convex relaxation of it that tightens as eps -> 0 otherwise.

Cost
    |D| binaries, sum_i |V_i| copies, sum_i (m_i + 2 |V_i|) + n + 1
    constraints.

Parallel
    the blocks i are independent except for steps 2 and 3; in a
    first-order LP solver the matrix is block-angular, and the
    per-block work (perspective evaluation, projection onto the scaled
    bounds) is elementwise over i.

(What the two reformulations cost a solver) Both algorithms are preprocessing, and their costs are easy to count. The constants of Algorithm 4.2.8 take one small linear program per row of each alternative and per other alternative, \(\sum_i m_i (k - 1)\) in all, and the programs are independent of one another. The hull model of Algorithm 4.2.9 carries \(|D|\) binaries and \(\sum_i |V_i|\) copied continuous variables, so its size grows with the number of alternatives times the number of variables they touch. Its constraint matrix is block-angular: the rows of alternative \(i\) involve only the copies \(\nu^i\) and the weight \(z_i\), so the matrix consists of \(|D|\) independent blocks, coupled only by the linking rows \(x_j = \sum_i \nu^i_j\) and \(\sum_i z_i = 1\). A simplex method pays for every added column at every pivot. The hull model is also degenerate in the LP sense: at a vertex of its relaxation many basic variables, here the copies of the inactive alternatives, sit at zero, and a simplex method stalls on such vertices. A first-order LP method of the kind in Section 7.2 pays one sparse matrix–vector product per iteration, which splits into one product per block plus the linking rows, and it is indifferent to the number of columns. That is why the hull formulation, costly for the one method and cheap for the other, is a natural candidate for the GPU solvers of Section 7.

The hull model's constraint matrix (Algorithm 4.2.9): block-angular

  columns         x     nu^1, z_1   nu^2, z_2   ...   nu^|D|, z_|D|
                +----+------------+-----------+-----+-------------+
  x = sum nu^i  | #  |     #      |     #     | ... |      #      |
  sum z_i = 1   |    |     #      |     #     | ... |      #      |
                +----+------------+-----------+-----+-------------+
  term 1        |    |  block 1   |           |     |             |
  term 2        |    |            |  block 2  |     |             |
  ...           |    |            |           | ... |             |
  term |D|      |    |            |           |     |  block |D|  |
                +----+------------+-----------+-----+-------------+

  #        the linking rows (steps 2 and 3), the only coupling
  block i  the rows of term i (steps 4 and 5): its copies nu^i and
           its weight z_i only
  a matrix-vector product splits into one product per block plus
  the linking rows

The hull formulation of a single two-term disjunction is the smallest instance of a much larger fact, which explains why branching on binary variables one at a time is enough to reach the integer hull. Theorem 3.3.7 stated the fact in its last clause and left the proof to this section.

Theorem 4.2.10 (Balas, 1979, 1985; sequential convexification; the last clause of Theorem 3.3.7). Let \(P \subseteq \mathbb{R}^n\) be a polyhedron for which \(0 \le x_j \le 1\) is valid for \(j = 1, \dots, p\), and for a convex set \(Q\) let \(\mathcal{P}_j(Q) = \operatorname{conv}\big((Q \cap \{x_j = 0\}) \cup (Q \cap \{x_j = 1\})\big)\). Then

\[\mathcal{P}_p\big(\mathcal{P}_{p-1}(\cdots \mathcal{P}_1(P) \cdots)\big) \;=\; \operatorname{conv}\big(P \cap \{x : x_j \in \{0, 1\},\ j = 1, \dots, p\}\big).\]

Proof sketch. The lemma is about faces. If \(F_1, \dots, F_m\) and \(F\) are faces of a polyhedron \(Q\), then \(\operatorname{conv}(\bigcup_i F_i) \cap F = \operatorname{conv}(\bigcup_i (F_i \cap F))\). The reason is that a point of \(F\) written as a convex combination of points of \(Q\) has all its support inside \(F\), by the definition of a face: a face is the set where some inequality valid for \(Q\) holds with equality, and a combination with positive weight on a point where the inequality is strict cannot land on it. Now \(F^0 = P \cap \{x_1 = 0\}\) and \(F^1 = P \cap \{x_1 = 1\}\) are faces of \(P_1 := \mathcal{P}_1(P)\). The inequality \(x_1 \ge 0\) is valid for \(P_1\), and a point of \(P_1\) with \(x_1 = 0\) is a convex combination of points with \(x_1 = 0\) and \(x_1 = 1\) that puts no weight on the latter. Since \(x_2 \ge 0\) is valid for \(P_1\), the set \(P_1 \cap \{x_2 = 0\}\) is a face of \(P_1\), and the lemma gives \(P_1 \cap \{x_2 = 0\} = \operatorname{conv}\big((P \cap \{x_1 = 0, x_2 = 0\}) \cup (P \cap \{x_1 = 1, x_2 = 0\})\big)\). The same holds for \(x_2 = 1\). Hence \(\mathcal{P}_2(\mathcal{P}_1(P))\) is the hull of the four faces \(P \cap \{x_1 = a, x_2 = b\}\), \(a, b \in \{0, 1\}\), and induction on \(p\) finishes the proof. ∎

(Why 0–1 disjunctions convexify one at a time) The 0–1 disjunctions are facial: both sides are faces of the current polyhedron. Convexifying one of them creates no new mixtures across another, so the hull can be built one variable at a time, in \(p\) rounds. For a general integer variable the split \(x_j \le \beta \ \vee\ x_j \ge \beta + 1\) is not facial and the finite statement fails: Cook, Kannan and Schrijver exhibit mixed-integer sets for which no finite number of rounds of split cuts gives the hull.W. Cook, R. Kannan and A. Schrijver, "Chvátal closures for mixed integer programming problems", Mathematical Programming 47 (1990). The sequential procedure is also the origin of the lift-and-project rank and of its comparison with the Lovász–Schrijver and Sherali–Adams hierarchies of Section 4.7: L. Lovász and A. Schrijver, "Cones of matrices and set-functions and 0–1 optimization", SIAM Journal on Optimization 1 (1991); H. D. Sherali and W. P. Adams, "A hierarchy of relaxations between the continuous and convex hull representations for zero-one programming problems", SIAM Journal on Discrete Mathematics 3 (1990).

Sequential convexification (Theorem 4.2.10), round by round

  P
  |  round 1: P_1(Q) = conv( (Q cap {x_1 = 0}) u (Q cap {x_1 = 1}) )
  v
  P_1(P)
  |  round 2: the same with x_2
  v
  P_2(P_1(P))
  :
  |  round p: the same with x_p
  v
  P_p(P_{p-1}( ... P_1(P) ... ))
      = conv( P cap {x : x_j in {0, 1}, j = 1, ..., p} )

  both sides of each 0-1 disjunction are faces of the current set,
  so p rounds reach the hull; a general integer split is not facial

(The dual reading: lift-and-project cuts) Balas' theorem also has a dual reading, which is where the lift-and-project cuts of Section 3.3 come from. Theorem 3.3.7 states the cut-generating program: with \(P = \{x : Ax \ge b\}\) as in that theorem, an inequality \(\alpha^\top x \ge \beta\) is valid for \(\mathcal{P}_j(P)\) exactly when \(\alpha = A^\top u - u_0 e_j = A^\top v + v_0 e_j\) and \(\beta \le \min(u^\top b,\ v^\top b + v_0)\) for multipliers \(u, v \ge 0\) and \(u_0, v_0 \ge 0\). Read in the lifted variables of (4.2.1), this is Balas' theorem in the dual. The hull formulation of the two-term disjunction has the blocks \(A \nu^1 \ge b \lambda_1,\ \nu^1_j \le 0\) and \(A \nu^2 \ge b \lambda_2,\ \nu^2_j \ge \lambda_2\), the multipliers \((u, u_0)\) and \((v, v_0)\) are Farkas certificates for the two blocks, nonnegative combinations of a block's rows that produce the inequality, which by Farkas' lemma is what every inequality valid for a polyhedron is, and the requirement that both certificates produce the same \(\alpha\) comes from the linking row \(x = \nu^1 + \nu^2\). What Theorem 3.3.7 left open is the normalization. Written as a linear program, the most violated cut at a point \(\bar x\) with \(0 < \bar x_j < 1\) is

\[\begin{aligned} \min_{\alpha, \beta, u, v, u_0, v_0}\ & \alpha^\top \bar x - \beta \\ \text{s.t.}\ & \alpha = A^\top u - u_0 e_j, \quad \beta \le b^\top u, \\ & \alpha = A^\top v + v_0 e_j, \quad \beta \le b^\top v + v_0, \\ & u, v \ge 0,\ \ u_0, v_0 \ge 0, \qquad \mathbf 1^\top u + u_0 + \mathbf 1^\top v + v_0 = 1 , \end{aligned}\]

(The normalization) The last row is the normalization. Without it the feasible set is a cone, every violated cut can be scaled to an arbitrarily large violation, and the minimum is \(-\infty\). With it the optimal value is negative if and only if \(\bar x \notin \mathcal{P}_j(P)\), and the facets of \(\mathcal{P}_j(P)\) correspond to extreme points of the program, up to the choice of the normalization. The row that sums all multipliers is the normalization of Balas, Ceria and Cornuéjols. Cornuéjols and Lemaréchal analyse the alternatives and show that the choice decides which facet the program returns.Balas, Ceria and Cornuéjols (1993), cited in Section 3.3; G. Cornuéjols and C. Lemaréchal, "A convex-analysis perspective on disjunctive cuts", Mathematical Programming 106 (2006). After \(\alpha\) is eliminated, the program has \(2m + 2\) multiplier variables and \(n\) linking rows \(A^\top u - u_0 e_j = A^\top v + v_0 e_j\), besides the two rows for \(\beta\) and the normalization. Balas and Perregaard's correspondence between its optimal solutions and Gomory mixed-integer cuts (Theorem 3.3.4) read from pivoted tableau rows, and the convex-MINLP variants of Stubbs and Mehrotra and of Kılınç, Linderoth and Luedtke, are described after Theorem 3.3.7. For nonconvex MINLP the disjunction is the spatial branch \(x_j \le \beta \ \vee\ x_j \ge \beta\) applied to the current LP relaxation, McCormick rows included, and the same program separates a cut from it. Belotti implemented this in Couenne.P. Belotti, "Disjunctive cuts for nonconvex MINLP", in J. Lee and S. Leyffer (eds), Mixed Integer Nonlinear Programming, IMA Volumes 154 (Springer, 2012).

Process engineers made the disjunction the primitive of a modelling language rather than a device inside a solver. A generalized disjunctive program writes a model with Boolean variables, disjunctions whose terms carry constraints and costs, and logic propositions among the Booleans, and the two reformulations above are its two compilation targets.

Definition 4.2.11 (generalized disjunctive program).

\[\begin{aligned} \min\ & f(x) + \sum_{k \in K} c_k \\ \text{s.t.}\ & g(x) \le 0, \\ & \bigvee_{i \in D_k} \begin{bmatrix} Y_{ki} \\ r_{ki}(x) \le 0 \\ c_k = \gamma_{ki} \end{bmatrix}, \quad k \in K, \\ & \Omega(Y) = \text{True}, \qquad x \in [l, u], \qquad Y_{ki} \in \{\text{True}, \text{False}\}, \end{aligned}\]

where exactly one \(Y_{ki}\) is true in each disjunction \(k\). The set \(\Omega\) of logic propositions on the Booleans is converted to linear inequalities \(Hz \ge h\) on binaries, and the \(r_{ki}\) are convex (a convex GDP) or general (a nonconvex GDP).R. Raman and I. E. Grossmann, "Modelling and computational techniques for logic based integer programming", Computers & Chemical Engineering 18 (1994); I. E. Grossmann and F. Trespalacios, "Systematic modeling of discrete-continuous optimization models through generalized disjunctive programming", AIChE Journal 59 (2013).

Theorem 4.2.12 (Lee and Grossmann, 2000; the hull relaxation of a convex GDP). For a convex GDP with bounded \(x\), the mixed-integer program

\[\begin{aligned} \min\ & f(x) + \sum_k \sum_{i \in D_k} \gamma_{ki} z_{ki} \quad \text{s.t. } g(x) \le 0, \\ & x = \sum_{i \in D_k} \nu^{ki}, \quad z_{ki}\, l \le \nu^{ki} \le z_{ki}\, u, \quad z_{ki}\, r_{ki}(\nu^{ki} / z_{ki}) \le 0, \quad \sum_{i \in D_k} z_{ki} = 1 \quad (k \in K), \\ & Hz \ge h, \quad z \in \{0, 1\}, \end{aligned}\]

is equivalent to the GDP, and its continuous relaxation is the intersection over \(k\) of the convex hulls of the individual disjunctions, which is at least as tight as the big-M relaxation.S. Lee and I. E. Grossmann, "New algorithms for nonlinear generalized disjunctive programming", Computers & Chemical Engineering 24 (2000); I. E. Grossmann and S. Lee, "Generalized convex disjunctive programming: nonlinear convex hull relaxation", Computational Optimization and Applications 26 (2003).

(The perspective, and the division by zero at \(z = 0\)) The function \(z\, r(\nu / z)\) is the perspective of \(r\), which Section 4.3 studies. It is convex, and at \(z \in \{0, 1\}\) it reproduces the disjunct exactly, but it is not differentiable at \(z = 0\) and a naive NLP code divides by zero there. Grossmann and Lee used \((z + \varepsilon)\, r(\nu / (z + \varepsilon)) \le 0\), which is correct only in the limit \(\varepsilon \to 0\) and produces violations of order \(\varepsilon\) at the integer points. Furman, Sawaya and Grossmann gave the form

\[\big((1 - \varepsilon) z + \varepsilon\big)\, r\!\Big(\frac{\nu}{(1 - \varepsilon) z + \varepsilon}\Big) - \varepsilon\, r(0)\,(1 - z) \;\le\; 0 ,\]

which is convex, exact at \(z \in \{0, 1\}\) for every \(\varepsilon \in (0, 1)\), and whose relaxation tightens as \(\varepsilon \downarrow 0\). It is the default of Pyomo.GDP's hull transformation, with \(\varepsilon = 10^{-4}\), and for quadratic disjuncts the exact conic perspective avoids \(\varepsilon\) altogether.K. C. Furman, N. W. Sawaya and I. E. Grossmann, "A computationally useful algebraic representation of nonlinear disjunctive convex sets using the perspective function", Computational Optimization and Applications 76 (2020). The implementation and its default are in Q. Chen, E. S. Johnson, D. E. Bernal, R. Valentin, S. Kale, J. Bates, J. D. Siirola and I. E. Grossmann, "Pyomo.GDP: an ecosystem for logic based modeling and optimization development", Optimization and Engineering 23 (2022), and in the source file pyomo/gdp/plugins/hull.py.

The hull relaxation of Theorem 4.2.12 convexifies each disjunction on its own, and the first warning of Section 4.1 applies: the intersection of the hulls is larger than the hull of the intersection. Balas' remedy is to intersect first and convexify afterwards.

Theorem 4.2.13 (Balas, 1985; basic steps). Let \(F = \bigcap_{j \in T} S_j\) be in regular form with \(S_j = \bigcup_{i \in Q_j} P_{ji}\). A basic step replaces two conjuncts \(S_k\) and \(S_l\) by the single disjunction \(S_{kl} = \bigcup_{i \in Q_k,\ i' \in Q_l} (P_{ki} \cap P_{li'})\), which leaves the set unchanged. Write \(\operatorname{h\text{-}rel}(F) = \bigcap_j \operatorname{cl}\operatorname{conv}(S_j)\) for the hull relaxation. If \(F_0, F_1, \dots, F_t\) is a sequence of regular forms, \(F_0\) the conjunctive normal form and \(F_t\) the disjunctive normal form, each obtained from the previous one by a basic step, then

\[\operatorname{h\text{-}rel}(F_0) \supseteq \operatorname{h\text{-}rel}(F_1) \supseteq \dots \supseteq \operatorname{h\text{-}rel}(F_t) = \operatorname{cl}\operatorname{conv}(F).\]

The inclusion holds because \(\operatorname{conv}(S_k \cap S_l) \subseteq \operatorname{conv}(S_k) \cap \operatorname{conv}(S_l)\) always and \(S_k \cap S_l = S_{kl}\). The price of a step is size, \(|Q_k| \cdot |Q_l|\) terms where there were \(|Q_k| + |Q_l|\). Grossmann and Ruiz report numbers for a strip-packing instance with four rectangles. Its hull formulation has 102 variables (24 binary) and 143 constraints with root bound 4. After basic steps it has 170 variables and 347 constraints with root bound 8, which is the optimum. For 25 rectangles the bound rises from 9 to 27 and for 31 from 10.64 to 33, with the number of binaries unchanged.Balas (1985), Theorem 4.3 as numbered by I. E. Grossmann and J. P. Ruiz, "Generalized disjunctive programming: a framework for formulation and alternative algorithms for MINLP optimization", in Mixed Integer Nonlinear Programming, IMA Volumes 154 (Springer, 2012), Table 2 (numbers read from the authors' PDF on 4 October 2026 and not re-verified). The linear case is N. W. Sawaya and I. E. Grossmann, "A hierarchy of relaxations for linear generalized disjunctive programming", European Journal of Operational Research 216 (2012); the convex case, which shows that the tightest relaxation allows a convex GDP to be solved as a single NLP in theory, is J. P. Ruiz and I. E. Grossmann, "A hierarchy of relaxations for nonlinear convex generalized disjunctive programming", European Journal of Operational Research 218 (2012). F. Trespalacios and I. E. Grossmann, "Cutting plane algorithm for convex generalized disjunctive programs", INFORMS Journal on Computing 28 (2016), use the strengthened hull to generate cuts for the big-M formulation instead of solving the large hull model.

Basic steps (Theorem 4.2.13): intersect first, convexify afterwards

  one basic step on two conjuncts S_k and S_l of a regular form:

    S_k = P_k1 or P_k2 or ...                     |Q_k| terms
    S_l = P_l1 or P_l2 or ...                     |Q_l| terms
                       |
                       v
    S_kl = (P_k1 cap P_l1) or (P_k1 cap P_l2) or ...
        or (P_k2 cap P_l1) or (P_k2 cap P_l2) or ...
        or ...                                    |Q_k| |Q_l| terms

  regular form                    hull relaxation
  F_0  conjunctive normal form    h-rel(F_0)
   |   a basic step                   contains
   v
  F_1                             h-rel(F_1)
   :   a basic step at a time         contains
   v
  F_t  disjunctive normal form    h-rel(F_t) = cl conv(F)

  h-rel(F) = the intersection over j of cl conv(S_j)
  strip packing, four rectangles (Grossmann and Ruiz): the root
  bound rises from 4 with the hull formulation to 8, the optimum,
  after basic steps

Proposition 4.2.14 (Trespalacios and Grossmann, 2015; multiple-parameter big-M). Keep one constant per row and per other alternative, \(M^{ij}_l = h_{P_j}(a^i_l) - b^i_l\), and write row \(l\) of alternative \(i\) as \(a^{i\top}_l x \le b^i_l + \sum_{j \ne i} M^{ij}_l z_j\). With the smallest valid constants, the continuous relaxation of this multiple-parameter big-M formulation is contained in that of the single-parameter one of Definition 4.2.2, with no new variables or constraints. By \(\sum_j z_j = 1\) the row reads \(a^{i\top}_l x \le \sum_j h_{P_j}(a^i_l)\, z_j\). Its right-hand side is the support function of the Minkowski combination \(\sum_j z_j P_j\) in the direction \(a^i_l\), which is why Vielma calls this the Balas–Blair–Jeroslow "hybrid" formulation. It is sharp wherever the single-parameter form is, since its relaxation is contained in that one and still contains the hull. In their numerical example the optimum is \(-9.472\), the single-parameter root bound \(-10.493\) and the multiple-parameter bound \(-9.735\).F. Trespalacios and I. E. Grossmann, "Improved Big-M reformulation for generalized disjunctive programs", Computers & Chemical Engineering 76 (2015), Theorem 3.1 and the example of Section 4; the three values were read from the authors' PDF on 4 October 2026 and could not be re-checked against the publisher's copy. Vielma (2015), Section 6.2, for the hybrid formulation. The constants require one optimization per row and per other alternative, an embarrassingly parallel preprocessing step.

The GDP framework also has its own convex-MINLP algorithm. Logic-based outer approximation fixes the Booleans and solves an NLP that contains only the constraints of the active terms, so the subproblem is smaller and better conditioned than the MINLP's own NLP relaxation. The MILP master contains linearizations of each term's constraints taken only at points where that term was active, placed inside the disjunction in big-M or hull form. A linearization of \(r_{ki}\) at a point where the term is inactive would mean nothing. An initial set-covering step makes sure every term is linearized at least once. The method inherits the finite convergence of outer approximation for convex problems (Theorem 3.4.5). Nonconvex variants solve the subproblems globally and use convex relaxations in the master, or branch on the disjunctions directly.M. Türkay and I. E. Grossmann, "Logic-based MINLP algorithms for the optimal synthesis of process networks", Computers & Chemical Engineering 20 (1996); S. Lee and I. E. Grossmann, "A global optimization algorithm for nonconvex generalized disjunctive programming and applications to process systems", Computers & Chemical Engineering 25 (2001). Pyomo.GDP (Chen et al. 2022) implements the big-M and hull transformations, cutting planes and basic steps, and its solver GDPopt implements logic-based outer approximation, its global variant, logic-based branch and bound and a relaxation-with-integer-cuts variant; H. D. Perez, S. Joshi and I. E. Grossmann, "DisjunctiveProgramming.jl: generalized disjunctive programming models and algorithms for JuMP", arXiv 2304.10492 (2023), brings the same reformulations to JuMP.

(How the solvers treat an indicator) Inside the MILP solvers the disjunction most often arrives as an indicator constraint (Section 2.5), \(z = 1 \Rightarrow a^\top x \le b\), and the solvers do not all turn it into a big-M row. SCIP's constraint handler adds a slack \(s \ge 0\) with \(a^\top x - s \le b\) and enforces \(z = 1 \Rightarrow s = 0\) by branching. This is a complementarity between \(z\) and \(s\) of the kind a special ordered set of type 1 expresses, a set of variables of which at most one may be nonzero. The handler separates cuts \(\sum_{i \in I} z_i \le |I| - 1\) from minimal infeasible subsystems, which it finds as vertices of an "alternative polyhedron". FICO Xpress removes indicator rows from the matrix after presolve, keeps them in a pool, and adds a row back only when its condition holds in the tree. Gurobi exposes indicators as a general constraint type whose internal treatment its documentation does not describe. The comparison of big-M, perspective and branching treatments by CPLEX and FICO developers concludes that aggressive bound tightening, often overlooked, is the key building block when indicator constraints and disjunctions are present. That is what the worked example above showed in miniature.SCIP source, src/scip/cons_indicator.c (master branch, accessed 4 October 2026; its header notes that the name "indicator constraint" apparently comes from CPLEX), github.com/scipopt/scip; FICO Xpress Optimizer Reference Manual, XPRSsetindicators; Gurobi Optimizer Reference Manual, "Constraints" (accessed 4 October 2026); Belotti et al. (2016); P. Bonami, A. Lodi, A. Tramontani and S. Wiese, "On mathematical programming with indicator constraints", Mathematical Programming 151 (2015), which surveys the subject and extends hull descriptions in the original space for two special two-term cases. GAMS states that indicator constraints are supported by COPT, CPLEX, Gurobi, SCIP and Xpress; BARON receives whatever big-M or hull formulation the modeller writes.

(Formulations between big-M and the hull) Between the big-M and the hull there is room for formulations of intermediate size, and two recent lines of work occupy it. Kronqvist, Misener and Tsay partition the variables of each term into \(P\) groups, give each group one bounded auxiliary that carries its share of the constraint, and apply (4.2.1) to the coarsened disjunction. The 1-split has the same relaxation as big-M, splitting a group never weakens the relaxation, and for affine terms over a box the \(n\)-split is the hull. On 344 instances from clustering, \(P\)-ball and trained-network problems the \(P\)-splits explored about as many nodes as the hull formulation in an order of magnitude less time, and beat big-M on both counts.J. Kronqvist, R. Misener and C. Tsay, "Between steps: intermediate relaxations between big-M and convex hull formulations", CPAIOR 2021, LNCS 12735 (Springer, 2021), and "P-split formulations: a class of intermediate formulations between big-M and convex hull for disjunctive constraints", Mathematical Programming 218 (2026), online 9 June 2025. Anderson and coauthors treat the rectified linear unit \(y = \max\{0, w^\top x + b\}\) over a box in \(\mathbb{R}^\eta\), a two-term disjunction whose big-M formulation is ideal only for \(\eta = 1\). They give an ideal formulation in the original variables as a family of \(2^\eta\) inequalities, the projection of (4.2.1), with a most violated member found in \(O(\eta)\) time.Anderson, Huchette, Ma, Tjandraatmadja and Vielma (2020). The partition-based formulations of C. Tsay, J. Kronqvist, A. Thebelt and R. Misener, "Partition-based formulations for mixed-integer optimization of trained ReLU neural networks", NeurIPS 2021, arXiv 2102.04373, interpolate between the big-M and this hull exactly as \(P\)-splits do.

(The simplest disjunction: the tax problem's buy-or-sell rule) The tax problem of Section 9 contains a disjunction of the simplest kind, and it is worth seeing now how little machinery it needs. For one asset, let \(b \in [0, B]\) be the shares bought and \(s \in [0, S]\) the shares sold, with the rule that a name is not bought and sold in the same trade list: \(b \cdot s = 0\). This is the disjunction \(\{b = 0\} \vee \{s = 0\}\). The McCormick envelope of Theorem 2.4.7 relaxes the product with a variable \(w\) and the plane \(w \ge B s + S b - B S\). Setting \(w = 0\) gives \(b / B + s / S \le 1\), the triangle with vertices \((0, 0)\), \((B, 0)\) and \((0, S)\), which is exactly \(\operatorname{conv}(\{b = 0\} \cup \{s = 0\})\) on the box. The indicator model \(b \le B z\), \(s \le S (1 - z)\), \(z \in \{0, 1\}\), relaxed to \(z \in [0, 1]\), gives the same triangle, since \(b / B + s / S \le z + (1 - z) = 1\). At the centre \((B/2, S/2)\) all three descriptions say "feasible, undecided", and one branch on \(z\), or one spatial branch on \(b\) at zero, resolves it. No product variable is needed, and the hull is one inequality.

One asset, the model's rule: not bought and sold in one trade list. On the box [0, B] × [0, S] the disjunction b s = 0 is the two edges b = 0 and s = 0. Its hull is the triangle b/B + s/S ≤ 1, one inequality. The McCormick plane with w = 0 gives it, and so does the indicator model b ≤ B z, s ≤ S (1 − z) with z relaxed to [0, 1]. The centre (B/2, S/2) is feasible in all three descriptions, undecided; one branch on z, or one spatial branch on b at zero, resolves it.

(Four pitfalls) Four pitfalls recur. The big-M formulation is not always weaker than the hull: with the smallest valid constants it is sharp for two polytopes in the plane, and for the three squares it is sharp without being ideal. The weakness appears in three or more dimensions, with several terms, and overwhelmingly with loose constants. The hull formulation is not always better in practice: its relaxation is degenerate, with many copies \(\nu^i\) at zero, and the folklore that it underperforms its strength is what the intermediate formulations respond to. The hull of a union of unbounded polyhedra with different recession cones is not closed, and (4.2.1) returns the closure, so a point the relaxation returns may be unattainable. And integer variables are not facial: the sequential-convexification argument is for 0–1 disjunctions, and for general integers the split closure of Section 3.3 may need infinitely many rounds.

Where this is used

Every MILP solver branches on disjunctions and separates lift-and-project or Gomory cuts, which are facets of the hull of a two-term disjunction. SCIP, Xpress, CPLEX, Gurobi and COPT accept indicator constraints directly. SCIP treats them by branching on a slack and Xpress by pooling the rows, and the other vendors do not document their internal treatment. Couenne separates disjunctive cuts from spatial branches on its McCormick relaxation. Pyomo.GDP and DisjunctiveProgramming.jl compile logical models to either formulation on request, and the piecewise relaxations of Section 4.6 are hull formulations of a structured union.

What parallelizes

The smallest valid constants are one small optimization per row and per other term, independent of each other and of the node. Recomputing them from a node's local bounds gives locally hull-like relaxations at big-M size, a kernel a GPU runs over a frontier of nodes at once. The hull formulation's matrix is block-angular, one block per term coupled only through \(x = \sum_i \nu^i\) and \(\sum_i \lambda_i = 1\) (the \(\sum_i z_i = 1\) of Algorithm 4.2.9), which maps to one block of threads per term in a first-order LP method. Cut separation for a two-term disjunction is one small LP per candidate disjunction and batches across candidates and across nodes. The GPU verifiers of trained neural networks of the CROWN family are complete branch-and-bound codes over exactly these ReLU disjunctions. Their node bounds come from a Lagrangian relaxation of the split constraints and, in the newest variant, from general cuts generated by a CPU MIP solver running alongside the GPU search. Whether that pattern transfers to disjunctions whose terms are McCormick boxes or perspective sets is one of the open questions of Section 7.8.S. Wang, H. Zhang, K. Xu, X. Lin, S. Jana, C.-J. Hsieh and J. Z. Kolter, "Beta-CROWN: efficient bound propagation with per-neuron split constraints for complete and incomplete neural network robustness verification", NeurIPS 2021, arXiv 2103.06624, which reports speed-ups of up to three orders of magnitude over LP-based branch and bound; H. Zhang, S. Wang, K. Xu, L. Li, B. Li, S. Jana, C.-J. Hsieh and J. Z. Kolter, "General cutting planes for bound-propagation-based neural network verification", NeurIPS 2022, arXiv 2208.05740.

The perspective

This subsection applies Balas' theorem to the disjunction that a portfolio model contains most often: a continuous quantity \(x\) that is either zero or carries a convex cost. Most of the integers in a portfolio are indicators, a binary \(z\) that is \(1\) if a position is held or a lot is touched and \(0\) if not, tied to a continuous \(x\) by the rule "\(x = 0\) unless \(z = 1\)". The textbook writes the rule as \(x \le M z\) with \(M\) a bound on \(x\). That is the big-M formulation of the previous subsection with one term the point \(\{0\}\), and the relaxation then sets \(z = x / M\) and pays almost none of whatever \(z\) costs. When \(x\) also carries a convex cost \(f(x)\), the hull of the two terms has a closed form, the perspective of \(f\). This subsection derives it, states when it is the whole convex hull, writes it as a second-order cone, and gives the cutting planes that an LP-based solver uses in its place. It also describes the two shapes the relaxation of a fixed charge takes, depending on how large the charge is relative to the bound.

Definition 4.3.1 (perspective). For a convex function \(g : \mathbb{R}^n \to \mathbb{R} \cup \{+\infty\}\) the perspective is

\[\tilde g(\nu, \lambda) \;=\; \lambda\, g(\nu / \lambda) \quad (\lambda > 0),\]

and its closure at \(\lambda = 0\) is \(\tilde g(\nu, 0) = \lim_{t \downarrow 0} t\, g(\nu / t)\), the recession function of \(g\) in the direction \(\nu\). When \(g\) has a bounded domain this limit is \(0\) at \(\nu = 0\) and \(+\infty\) at every \(\nu \ne 0\), which is the right value for "switched off".R. T. Rockafellar, Convex Analysis (Princeton University Press, 1970), Section 8; J.-B. Hiriart-Urruty and C. Lemaréchal, Fundamentals of Convex Analysis (Springer, 2001), Section B.2.2.

Lemma 4.3.2 (convexity of the perspective). If \(g\) is convex then \((\nu, \lambda) \mapsto \lambda\, g(\nu / \lambda)\) is convex on \(\{\lambda > 0\}\) and its closure is convex on \(\{\lambda \ge 0\}\). For \(g(x) = x^2\) the perspective is \(\nu^2 / \lambda\), and its epigraph \(\{(\nu, \lambda, t) : t \lambda \ge \nu^2,\ t, \lambda \ge 0\}\) is a rotated second-order cone,

\[t \lambda \ge \nu^2,\ t, \lambda \ge 0 \quad\Longleftrightarrow\quad \big\lVert (2\nu,\ t - \lambda) \big\rVert_2 \le t + \lambda .\]

Proof. The epigraph \(\{(\nu, \lambda, t) : t \ge \lambda\, g(\nu / \lambda),\ \lambda > 0\}\) equals \(\{\lambda\,(y, 1, s) : s \ge g(y),\ \lambda > 0\}\), the cone generated by \(\operatorname{epi} g \times \{1\}\), and a cone generated by a convex set is convex. For the cone identity, square both sides: \(4 \nu^2 + t^2 - 2 t \lambda + \lambda^2 \le t^2 + 2 t \lambda + \lambda^2\) reduces to \(\nu^2 \le t \lambda\), and the right-hand side \(t + \lambda \ge 0\) with \(|t - \lambda| \le t + \lambda\) gives \(t, \lambda \ge 0\). ∎

(The cone, in two numbers) The epigraph of a function is the set of points on or above its graph (Definition 1.2.1), so the lemma's cone is the set of triples \((\nu, \lambda, t)\) in which \(t\) is at least the perspective's value. Two values fix the picture. At \(\nu = 1\), \(\lambda = 1\) the cone requires \(t \ge 1\), the cost of the quantity itself. At \(\nu = 1\), \(\lambda = \tfrac12\) it requires \(t \ge 2\), since a half-switched-on unit must be charged as if two units were running at half the weight. The conic form is what the conic solvers consume directly, and it is the form in which Aktürk, Atamtürk and Gürel first wrote the quadratic perspective.M. S. Aktürk, A. Atamtürk and S. Gürel, "A strong conic quadratic reformulation for machine-job assignment with controllable processing times", Operations Research Letters 37 (2009). The cone identity is Proposition 3.3.2 of A. Ben-Tal and A. Nemirovski, Lectures on Modern Convex Optimization (SIAM, 2001), as cited by Günlük and Linderoth (2010).

With the lemma in hand, Balas' theorem extends from polyhedra to convex sets by replacing "scale the linear constraint" with "take the perspective of the convex constraint".

Theorem 4.3.3 (Ceria and Soares, 1999; the hull of a union of convex sets). Let \(K_t = \{x \in \mathbb{R}^n : G_t(x) \le 0\}\), \(t \in T\), be convex and bounded, with \(G_t\) convex vector functions. Then \(x \in \operatorname{conv}(\bigcup_t K_t)\) if and only if there exist \(x^t\) and \(\lambda_t \ge 0\) with

\[x = \sum_{t \in T} x^t, \qquad \sum_{t \in T} \lambda_t = 1, \qquad \tilde G_t(x^t, \lambda_t) \le 0 \quad (t \in T),\]

where \(\tilde G_t\) is the closed perspective of \(G_t\).S. Ceria and J. Soares, "Convex programming for disjunctive convex optimization", Mathematical Programming 86 (1999); the statement in this form is Theorem 2.1 of O. Günlük and J. Linderoth, "Perspective reformulation and applications", in Mixed Integer Nonlinear Programming, IMA Volumes 154 (Springer, 2012). Without boundedness the recession directions enter exactly as in Theorem 4.2.4: Stubbs and Mehrotra (1999); H. Hijazi, P. Bonami, G. Cornuéjols and A. Ouorou, "Mixed-integer nonlinear programs featuring 'on/off' constraints", Computational Optimization and Applications 52 (2012).

The proof is the proof of Theorem 4.2.4 with Lemma 4.3.2 in place of linearity: the scaled constraint on the point from term \(t\) is \(\lambda_t G_t(x^t / \lambda_t) \le 0\), which is convex in \((x^t, \lambda_t)\) by the lemma. Boundedness makes the hull closed and makes the closed perspective at \(\lambda_t = 0\) force \(x^t = 0\).

(The indicator pair: the hull in three lines) The indicator pair is the case of two terms, one of them a point, and the derivation is three lines. Take the cost constraint \(t \ge x^2\) with the switch, so the feasible set is \(\{(x, z, t) : z = 0, x = 0, t \ge 0\} \cup \{(x, z, t) : z = 1, t \ge x^2\}\). Write \(x = x^0 + x^1\) with \(x^0\) in \((1 - z) \cdot \{0\}\), which forces \(x^0 = 0\), and \((x^1, t^1)\) in \(z \cdot \{t \ge x^2\}\), which by the lemma is \(t^1 z \ge (x^1)^2\). Since \(x = x^1\), the hull is \(t z \ge x^2\). At \((x, z) = (\tfrac12, \tfrac12)\) the big-M relaxation, which keeps the cost constraint as \(t \ge x^2\) and only relaxes the switch, requires \(t \ge \tfrac14\). The hull requires \(t \ge x^2 / z = \tfrac12\), twice as much.

The indicator pair as a two-term disjunction (Theorem 4.3.3)

  term "off", weight 1 - z          term "on", weight z
  z = 0, x = 0, t >= 0              z = 1, t >= x^2
            |                                 |
            v                                 v
  x^0 in (1 - z) {0},               (x^1, t^1) in z {t >= x^2},
  which forces x^0 = 0              which by Lemma 4.3.2 is
            |                       t^1 z >= (x^1)^2
            |                                 |
            +----------------+----------------+
                             v
          x = x^0 + x^1 = x^1:  the hull is  t z >= x^2

  at (x, z) = (1/2, 1/2)
      big-M   t >= x^2      = 1/4
      hull    t >= x^2 / z  = 1/2, twice as much

Theorem 4.3.4 (Günlük and Linderoth, 2010; the hull of an on/off set). (a) Let \(W^0 = \{(x, z) : x = 0, z = 0\}\) and \(W^1 = \{(x, z) : f_i(x) \le 0\ (i \in I),\ l \le x \le u,\ z = 1\}\) with \(f_i\) convex and bounded on \([l, u]\). Then

\[\operatorname{conv}(W^0 \cup W^1) \;=\; \operatorname{cl}\big\{(x, z) : z f_i(x / z) \le 0\ (i \in I),\ l z \le x \le u z,\ 0 < z \le 1\big\}.\]

(b) For \(S = \{(x, y, z) \in \mathbb{R}^2 \times \{0, 1\} : y \ge x^2,\ u z \ge x \ge l z,\ x \ge 0\}\),

\[\operatorname{conv}(S) \;=\; \{(x, y, z) : y z \ge x^2,\ u z \ge x \ge l z,\ 0 \le z \le 1,\ x, y \ge 0\}.\]

(c) If \(P_1 \subseteq \mathbb{R}^{n_1}\) and \(P_2 \subseteq \mathbb{R}^{n_2}\) are closed, pointed and integral in their indicator coordinates, then so is \(P_1 \times P_2\), and \(\operatorname{conv}(P_1 \times P_2) = \operatorname{conv}(P_1) \times \operatorname{conv}(P_2)\). Adding a constraint \(w \ge q^\top y\) with \(q \ge 0\) preserves both properties. Consequently, for the separable set \(Q = \{(w, x, z) : w \ge \sum_i q_i x_i^2,\ u_i z_i \ge x_i \ge l_i z_i,\ z \in \{0, 1\}^n\}\), the extended set \(\{w \ge \sum_i q_i y_i,\ y_i z_i \ge x_i^2,\ u_i z_i \ge x_i \ge l_i z_i,\ 0 \le z \le 1\}\) is the convex hull of \(Q\) lifted, and its projection onto \((w, x, z)\) is described by the exponential family

\[w \prod_{i \in S} z_i \;\ge\; \sum_{i \in S} q_i x_i^2 \prod_{l \in S \setminus \{i\}} z_l \qquad (S \subseteq \{1, \dots, n\})\]

of inequalities, one for each subset \(S\).O. Günlük and J. Linderoth, "Perspective reformulations of mixed integer nonlinear programs with indicator variables", Mathematical Programming 124 (2010), Lemmas 1 to 7 and Corollary 1 of the Optimization Online preprint 2008/06/2014, whose numbering is used here; the survey form is Günlük and Linderoth (2012).

Proof sketch. We prove (b). Part (a) is the same argument with general convex \(f_i\), and (c) is in the reference. The set is \(S^0 \cup S^1\) with \(S^0 = \{(0, y, 0) : y \ge 0\}\), a ray, and \(S^1 = \{(x, y, 1) : y \ge x^2,\ l \le x \le u\}\). A convex combination \(\lambda (x_1, y_1, 1) + (1 - \lambda)(0, y_0, 0)\) has \(z = \lambda\), \(x = \lambda x_1\) and \(y = \lambda y_1 + (1 - \lambda) y_0 \ge \lambda x_1^2 = x^2 / z\), so \(y z \ge x^2\), and the bounds scale to \(l z \le x \le u z\). Conversely any point with \(y z \ge x^2\) and \(z \in (0, 1]\) decomposes as such a combination with \(x_1 = x / z\), \(y_1 = x^2 / z^2\) and the slack \(y - x^2 / z\) placed in \(y_0\). The set is closed because \(\{y z \ge x^2,\ x, y, z \ge 0\}\) is a closed cone. ∎

For \(n = 2\) and \(S = \{1, 2\}\) the member of the family is

\[w z_1 z_2 \;\ge\; q_1 x_1^2 z_2 + q_2 x_2^2 z_1 .\]

(The projected family for two pairs, and what part (c) buys) It is the lifted description multiplied out. Multiplying \(w \ge q_1 y_1 + q_2 y_2\) by \(z_1 z_2 \ge 0\) and using \(y_1 z_1 \ge x_1^2\) and \(y_2 z_2 \ge x_2^2\) gives \(w z_1 z_2 \ge q_1 (y_1 z_1) z_2 + q_2 (y_2 z_2) z_1 \ge q_1 x_1^2 z_2 + q_2 x_2^2 z_1\), and the auxiliary variables \(y\) have disappeared. The members for \(S = \{1\}\) and \(S = \{2\}\) are the single perspectives \(w z_1 \ge q_1 x_1^2\) and \(w z_2 \ge q_2 x_2^2\), and \(S = \emptyset\) gives \(w \ge 0\). Part (c) is the statement that matters for the quadratic problems of the next two subsections. It says that the hull of a Cartesian product is the product of the hulls, because the extreme points of a product are the pairs of extreme points, and that a sum of separable terms inherits this. The perspective reformulation of a separable convex objective with indicators is therefore the convex hull of the whole set, not merely of each term, as long as nothing else couples the pairs \((x_i, z_i)\). Wei, Gómez and Küçükyavuz later proved that it stays ideal under arbitrary constraints on the indicator vector \(z\) alone, a result Section 4.5 states. What breaks it is coupling through \(x\), which is the subject of that subsection.

Theorem 4.3.4(c) for n = 2: the lifted set and its projection

  lifted, with the auxiliary variables y_1 and y_2:
      w >= q_1 y_1 + q_2 y_2
      y_1 z_1 >= x_1^2          y_2 z_2 >= x_2^2
                     |
                     |  the lifted description multiplied out:
                     |  the auxiliary variables y disappear
                     v
  projected onto (w, x, z), one inequality per subset S of {1, 2}:
      S = {1, 2}   w z_1 z_2 >= q_1 x_1^2 z_2 + q_2 x_2^2 z_1
      S = {1}      w z_1 >= q_1 x_1^2
      S = {2}      w z_2 >= q_2 x_2^2
      S = {}       w >= 0

An LP-based solver cannot carry a cone, but it can carry the cone's tangent planes, and for the perspective these are the perspective cuts named in Section 3.3. A tangent plane of a convex function at a point is given by a subgradient there, a slope vector \(s\) with \(f(x) \ge f(\bar x) + s^\top (x - \bar x)\) for every \(x\), which for a differentiable function is the gradient; the proposition is stated for subgradients so that it covers piecewise costs too.

Proposition 4.3.5 (Frangioni and Gentile, 2006; perspective cuts). Consider \(\min\{f(x) + c z : A x \le b z,\ z \in \{0, 1\}\}\) with \(X = \{x : A x \le b\}\) bounded, \(f\) convex and finite on \(X\), \(f(0) = 0\), written in epigraph form as \(\min\{v : v \ge f(x) + c z,\ A x \le b z\}\). For every \(\bar x \in X\) and every subgradient \(s \in \partial f(\bar x)\) the inequality

\[v \;\ge\; s^\top x + \big(c + f(\bar x) - s^\top \bar x\big)\, z \tag{4.3.1}\]

is valid, and the family over all \(\bar x \in X\) describes the epigraph of the perspective \(z f(x / z) + c z\), which is the convex hull of the feasible set in \((x, z, v)\).A. Frangioni and C. Gentile, "Perspective cuts for a class of convex 0–1 mixed integer programs", Mathematical Programming 106 (2006); the cut in this form is equation (9) of Günlük and Linderoth (2010).

Proof sketch. The function \(\varphi(x, z) = z f(x / z) + c z\) is convex and positively homogeneous of degree one, by Lemma 4.3.2. A supporting hyperplane of a positively homogeneous convex function passes through the origin, so the tangent plane at \((\bar x, 1)\), where \(\partial \varphi\) contains \((s,\ c + f(\bar x) - s^\top \bar x)\), is \(v \ge s^\top x + (c + f(\bar x) - s^\top \bar x) z\). For \(f(x) = a x^2 + b x\) this reads \(v \ge (2 a \bar x + b)\, x + (c - a \bar x^2)\, z\). Separation is exact and closed-form. Given a relaxed point \((\bar x, \bar z, \bar v)\) with \(\bar z > 0\), set \(\hat x = \bar x / \bar z\). The cut at \(\hat x\) is tight for the perspective at \((\bar x, \bar z)\), so the point violates it exactly when \(\bar v < \bar z f(\bar x / \bar z) + c \bar z\), that is, exactly when it violates the perspective relaxation. ∎

The perspective relaxation of a fixed charge plus a quadratic is usually drawn as a tangent line from the origin followed by the curve itself. That is one of two regimes, and the figure's sliders reach the other. The proposition computes the convex envelope of the cost on \([0, M]\), the largest convex function lying below it there (Definition 2.4.2), and shows that the perspective relaxation is that envelope in both regimes.

Proposition 4.3.6 (the convex envelope of a fixed charge plus a quadratic). Let the cost be \(0\) at \(x = 0\) and \(c + x^2\) for \(0 < x \le M\), with \(c > 0\). Its convex envelope on \([0, M]\) is

\[\operatorname{vex}_{[0, M]}(x) \;=\; \begin{cases} 2\sqrt{c}\; x & 0 \le x \le \sqrt{c}, \\ c + x^2 & \sqrt{c} \le x \le M, \end{cases} \quad \text{if } \sqrt{c} \le M; \qquad \operatorname{vex}_{[0, M]}(x) \;=\; \Big(\frac{c}{M} + M\Big)\, x \quad \text{if } \sqrt{c} > M .\]

The perspective relaxation \(\min\{c z + x^2 / z : x / M \le z \le 1\}\) equals this envelope for every \(x \in [0, M]\), and the big-M relaxation \(\min\{c z + x^2 : x / M \le z \le 1\} = c x / M + x^2\) lies below it, with equality only at \(x = 0\) and \(x = M\).

Proof. Minimizing \(c z + x^2 / z\) over \(z > 0\) gives \(z^\ast = x / \sqrt{c}\) and the value \(2 \sqrt{c}\, x\). The constraint \(z \ge x / M\) is inactive exactly when \(x / \sqrt{c} \ge x / M\), that is, when \(M \ge \sqrt{c}\). The constraint \(z \le 1\) is inactive exactly when \(x \le \sqrt{c}\). In the first regime, \(\sqrt{c} \le M\), the value is \(2 \sqrt{c}\, x\) for \(x \le \sqrt{c}\) and, with \(z = 1\), \(c + x^2\) beyond. The two pieces meet at \(x = \sqrt{c}\) with the common slope \(2 \sqrt{c}\), so the function is continuously differentiable, convex, below the cost, and equal to it at \(0\) and on \([\sqrt{c}, M]\). A convex minorant that coincides with the cost wherever the envelope must is the envelope. In the second regime, \(\sqrt{c} > M\), the unconstrained minimizer violates \(z \ge x / M\) for every \(x > 0\), so \(z = x / M\) and the value is \(c x / M + M x = (c / M + M) x\), the chord from \((0, 0)\) to \((M, c + M^2)\). The tangent point \(\sqrt{c}\) lies outside the interval, so the chord is the envelope. The big-M relaxation gives \(z = x / M\) in both regimes. ∎

(The two regimes, checked by a script) The second regime occurs whenever the fixed charge is large relative to the square of the bound, which for a small position is the common case rather than the exception. The script below checks both closed forms against the perspective minimization on a grid of 20,001 points and prints the three values at \(x = 0.60\) that the figure's default view shows. It then runs the exact separation of Proposition 4.3.5 as a cutting-plane method at that one \(x\): at each round it takes the point \(\bar z\) that minimizes the current cuts, separates the cut at \(\hat x = \bar x / \bar z\), and repeats. Since the problem has one continuous variable once \(x\) is fixed, the "LP" is a minimization over \(z\) on a fine grid.

# The fixed-charge quadratic of the perspective figure.
#
# Cost 0 at x = 0, c + x^2 for 0 < x <= M. Closed forms of the two
# relaxations in both regimes, then perspective cuts separated exactly
# (Algorithm 4.3.7) until the cutting-plane value at x = 0.60 reaches the
# perspective relaxation. numpy only.

import numpy as np

def big_m(x, c, M):
    return c * x / M + x * x

def persp(x, c, M):
    """min over z in [x/M, 1] of c z + x^2 / z, in closed form."""
    z = np.clip(x / np.sqrt(c), x / M, 1.0)
    return np.where(x > 0, c * z + x * x / np.where(z > 0, z, 1), 0.0)

def envelope(x, c, M):
    """Proposition 4.3.6: tangent-then-curve, or the chord."""
    r = np.sqrt(c)
    if r <= M:
        return np.where(x <= r, 2 * r * x, c + x * x)
    return (c / M + M) * x

for c, M in ((1.5, 4.0), (4.0, 1.0)):
    xs = np.linspace(0, M, 20001)
    regime = 'tangent then curve' if np.sqrt(c) <= M else 'chord'
    print(f"c = {c}, M = {M}, sqrt(c) = {np.sqrt(c):.3f}: "
          f"regime '{regime}'")
    err = np.abs(persp(xs, c, M) - envelope(xs, c, M)).max()
    print(f"  max |perspective relaxation - closed-form envelope| "
          f"over [0, M] = {err:.1e}")
    x0 = 0.60
    print(f"  at x = {x0}: true {c + x0**2:.3f}, "
          f"big-M {big_m(x0, c, M):.3f}, "
          f"perspective {float(persp(np.array(x0), c, M)):.3f}")

# Perspective cuts, separated exactly:
#   v >= s x + (c + f(xh) - s xh) z  with f = x^2, s = 2 xh,
#   and xh = xbar / zbar clipped to [0, M].
c, M, x0 = 1.5, 4.0, 0.60
zs = np.linspace(x0 / M, 1.0, 40001)

# cut points xh; xh = 0 gives v >= c z (the cost of switching on)
cuts = [0.0]

def value(z):
    """max over the cuts at (x0, z)."""
    return max(2 * xh * x0 + (c + xh * xh - 2 * xh * xh) * z
               for xh in cuts)

bound = 2 * np.sqrt(c) * x0
print()
print("perspective cuts at x = 0.60 (c = 1.5, M = 4): "
      "the LP value climbs to")
print(f"  the perspective bound 2 sqrt(c) x = {bound:.6f}")
print("  k   zbar     xhat    LP value")
for k in range(1, 13):
    vals = np.array([value(z) for z in zs])
    i = vals.argmin()
    zbar = zs[i]
    print(f"  {k}  {zbar:.4f}  {min(x0 / zbar, M):.4f}   {vals[i]:.4f}")
    # Algorithm 4.3.7, step 2: the cut at xhat is tight for the
    # perspective at (xbar, zbar)
    cuts.append(min(x0 / zbar, M))
print(f"  after {k} cuts the LP value is {vals[i]:.6f} "
      f"against the bound {bound:.6f}:")
print(f"  within {bound - vals[i]:.1e}")
c = 1.5, M = 4.0, sqrt(c) = 1.225: regime 'tangent then curve'
  max |perspective relaxation - closed-form envelope| over [0, M] = 8.9e-16
  at x = 0.6: true 1.860, big-M 0.585, perspective 1.470
c = 4.0, M = 1.0, sqrt(c) = 2.000: regime 'chord'
  max |perspective relaxation - closed-form envelope| over [0, M] = 8.9e-16
  at x = 0.6: true 4.360, big-M 2.760, perspective 3.000

perspective cuts at x = 0.60 (c = 1.5, M = 4): the LP value climbs to
  the perspective bound 2 sqrt(c) x = 1.469694
  k   zbar     xhat    LP value
  1  0.1500  4.0000   0.2250
  2  0.3000  2.0000   0.4500
  3  0.6000  1.0000   0.9000
  4  0.4000  1.5000   1.4000
  5  0.4800  1.2500   1.4400
  6  0.5333  1.1250   1.4667
  7  0.5053  1.1875   1.4684
  8  0.4923  1.2187   1.4692
  9  0.4861  1.2343   1.4696
  10  0.4892  1.2265   1.4697
  11  0.4907  1.2226   1.4697
  12  0.4900  1.2246   1.4697
  after 12 cuts the LP value is 1.469692 against the bound 1.469694:
  within 1.8e-06

The two closed forms agree with the minimization to machine precision in both regimes. At \(x = 0.60\) with \(c = 1.5\) and \(M = 4\) the true cost is \(1.860\), the big-M relaxation charges \(0.585\) and the perspective charges \(1.470\). In the chord regime, \(c = 4\) and \(M = 1\), the perspective relaxation is the line of slope \((c + M^2) / M = 5\), so it charges \(3.000\) where the big-M relaxation charges \(2.760\) and the truth is \(4.360\). The cutting-plane run shows what an LP-based solver sees. After four cuts the value is within \(5\%\) of the perspective bound, and after six within \(0.3\%\). After twelve the value \(1.469692\) agrees with the bound \(2\sqrt{c}\, x = 1.469694\) to five decimals, and rounds eight to twelve gain only from the fourth decimal on. That slow tail is what a cutting-plane method always has on a smooth curve. Each round is one evaluation of \(f\) and one of \(f'\) at a single point, with no linear program solved for the separation itself. The grid minimization over \(z\) in each round is one vectorized pass, and across many indicator pairs the rounds are independent, one pair per thread, which is the form Algorithm 4.3.7 below takes.

The figure draws the same relaxations as functions of \(x\). Its second panel draws the set \(S = \{(0, 0)\} \cup \{(x, c + x^2) : 0 < x \le M\}\) with its convex hull, whose lower boundary is the perspective relaxation and whose upper boundary is the vertical edge from \((0, 0)\) to \((0, c)\) followed by the chord to \((M, c + M^2)\). The thing to look for is that the orange curve is the lower boundary of that hull while the grey big-M line runs below it. In the default view, \(c = 1.5\), \(M = 4\) and the reading at \(x = 0.60\), the stats are the three numbers of the script, \(0.58\), \(1.47\) and \(1.86\), so the big-M relaxation understates the cost by \(69\%\) and the perspective by \(21\%\). The condition \(\sqrt{c} = 1.225 \le M\) holds, the tangent point is \((\sqrt{c}, 2c) = (1.22, 3.00)\), the chord's slope would be \((c + M^2)/M = 4.375\), and the hull's section at \(x = 0.60\) is the interval \([1.470, 3.900]\).

The cases move \(M\) and \(c\). The case button "M = 1.25, nearly tight" moves the bound close to \(\sqrt{c}\), and "M = 8, loose" doubles it. The big-M relaxation moves each time and the perspective does not, because the envelope of Proposition 4.3.6 does not depend on \(M\) once \(M \ge \sqrt{c}\). Set \(c = 4\) and \(M = 1\) to see the other regime. The condition line turns red, and the hull panel labels the chord of slope \(5.00\) in place of the tangent. The readings at \(x = 0.60\) are then \(2.76\), \(3.00\) and \(4.36\), understatements of \(37\%\) and \(31\%\).

A fixed charge c to switch on, then a cost x². The true cost (blue) jumps at zero. The big-M relaxation (grey) pays the charge in proportion to x/M and understates the cost wherever x is small. The perspective relaxation (orange) is the convex envelope itself, exact at zero and exact beyond √c, or, when √c exceeds M, the chord from the origin to (M, c + M²). The second panel draws the set of (x, cost) pairs the switch allows and its convex hull, whose lower boundary is the perspective relaxation. Loosen M and only one of the two relaxations moves.

The separation step of Proposition 4.3.5, written out for a model with many indicator pairs, is the following algorithm. It is the form in which perspective strengthening lives inside an LP-based branch and bound, where the second-order cone of Lemma 4.3.2 cannot be carried directly.

Algorithm 4.3.7 (perspective-cut separation).

Algorithm PERSPECTIVE-CUT-SEPARATION

Input   indicator pairs (x_i, z_i), i = 1..n, each with a convex f_i,
        f_i(0) = 0, a charge c_i and bounds l_i z_i <= x_i <= u_i z_i;
        a relaxed point (xbar, zbar, vbar), vbar_i the epigraph
        variable.
Output  violated perspective cuts.

for each i, independently:

  1. if zbar_i <= tol: skip
     (the point is on the "off" side; a cut at any xhat would be valid
     but separates nothing).

  2. xhat := xbar_i / zbar_i, clipped to [l_i, u_i].

  3. s := f_i'(xhat)  (any subgradient).

  4. viol := zbar_i f_i(xhat) + c_i zbar_i - vbar_i.

  5. if viol > tol:
        add  v_i >= s x_i + (c_i + f_i(xhat) - s xhat) z_i.

Invariant
    every cut supports the epigraph of the perspective, so it is valid
    for the hull; at the returned point the cut is tight for the
    perspective, so the point violates the cut if and only if it
    violates the perspective relaxation (exact separation).

Cost
    one evaluation of f_i and one of f_i' per indicator; no linear
    program.

Parallel
    elementwise over i, one thread per indicator; the cuts are sparse
    rows with two nonzeros and the epigraph variable.

(The batched listing) The algorithm is one of the few steps in a MINLP solver whose parallel form is literally one thread per element. The listing below is the batched version for a frontier of nodes. Each node has its own local bounds, and the separation runs over every (node, indicator) pair. A second routine recomputes the smallest valid big-M constant of a facet from a node's box, which is Algorithm 4.2.8 with an interval bound in place of the LP.

// Batched perspective-cut separation and tight big-M recomputation
// for a frontier of nodes.
//
// Each indicator pair (x_i, z_i) has cost c_i z_i + a_i x_i^2 + b_i x_i
// and local bounds [l_i, u_i] that differ per node. One thread per
// (node, indicator): no LP, no communication, two evaluations of f and
// f'.

#include <algorithm>
#include <cmath>
#include <cstddef>
#include <span>
#include <vector>

// fixed charge c, quadratic a x^2 + b x
struct Indicator { double c, a, b; };

// v_i >= sx * x_i + sz * z_i
struct Cut { std::size_t node, i; double sx, sz; };

// Node-local bounds: a facet alpha.x <= r of one disjunct is relaxed by
// M = max over the other disjunct of alpha.x - r; over a box the maximum
// is a sum of the better corner of each coordinate. Negative M is valid
// and tightens.
double tight_M(std::span<const double> alpha, double r,
               std::span<const double> lo, std::span<const double> hi) {
    double s = -r;
    for (std::size_t j = 0; j < alpha.size(); ++j)
        s += alpha[j] > 0 ? alpha[j] * hi[j] : alpha[j] * lo[j];
    return s;
}

// Algorithm 4.3.7 for a batch: xbar, zbar, vbar are laid out
// node-major, n indicators per node.
std::vector<Cut> separate(std::span<const Indicator> ind,
                          std::span<const double> xbar,
                          std::span<const double> zbar,
                          std::span<const double> vbar,
                          std::span<const double> lo,
                          std::span<const double> hi,
                          std::size_t nodes, double tol) {
    const std::size_t n = ind.size();
    std::vector<Cut> out;
    // the (k, i) loop is the kernel: independent work items
    for (std::size_t k = 0; k < nodes; ++k)
        for (std::size_t i = 0; i < n; ++i) {
            const std::size_t p = k * n + i;
            const double z = zbar[p];
            // the point is on the "off" side; nothing to separate
            if (z <= tol)
                continue;
            const double xh = std::clamp(xbar[p] / z, lo[p], hi[p]);
            const double f = ind[i].a * xh * xh + ind[i].b * xh;
            const double s = 2 * ind[i].a * xh + ind[i].b;
            const double persp =
                z * (ind[i].a * (xbar[p] / z) * (xbar[p] / z)
                     + ind[i].b * (xbar[p] / z))
                + ind[i].c * z;
            // violated iff the point violates the perspective relaxation
            if (persp - vbar[p] > tol)
                out.push_back({k, i, s, ind[i].c + f - s * xh});
        }
    return out;
}

// What a device version changes: the two loops become a grid of
// threads, `out` becomes a per-thread flag array compacted by a prefix
// sum, and the per-node bounds lo, hi are the frontier's box data
// already resident on the device.

The listing costs a handful of floating-point operations per pair and touches each input once. Its violation test evaluates the perspective at the unclipped \(\bar x / \bar z\), which agrees with step 4 of the algorithm whenever the relaxed point satisfies \(l z \le x \le u z\), as a point of the relaxation does. On a device the only non-elementwise step is the compaction of the flags into a list of cuts, a prefix sum. What it does not decide is cut management, which of the many cuts a frontier produces to keep, and that remains open in the batched setting.

The listing's batch: node-major arrays, one thread per pair (k, i)

  xbar, zbar, vbar, lo, hi: one entry per (node k, indicator i), at
  p = k n + i, with n indicators per node

  p     0     1    ...   n-1    n    n+1   ...   2n-1   ...
     +-----+-----+-----+-----+-----+-----+-----+------+-----+
     |         node 0        |         node 1         | ... |
     +-----+-----+-----+-----+-----+-----+-----+------+-----+

  ind: one (c, a, b) per indicator i, shared by every node

  thread (k, i)   z = zbar[p]; nothing to separate if z <= tol
                  xh = clamp(xbar[p] / z, lo[p], hi[p])
                  a cut (k, i, sx, sz) if persp - vbar[p] > tol
                         |
                         v   on a device: a flag per thread
                  flags --prefix sum--> a compact list of cuts

(Three refinements) Three refinements of the perspective are in the literature and in the solvers. The projected perspective reformulation of Frangioni, Gentile, Grande and Pacifici applies when \(z\) appears only in the perspective term and in the bound \(x \le u z\). Then \(z\) can be minimized out in closed form, exactly as Proposition 4.3.6 did for one pair. The relaxation becomes a piecewise-quadratic convex program of the same size as the original continuous relaxation, which an NLP solver handles without cones or cuts.A. Frangioni, C. Gentile, E. Grande and A. Pacifici, "Projected perspective reformulations with applications in design problems", Operations Research 59 (2011). When the projection cannot be done exactly, the approximated perspective relaxation of Frangioni, Furini and Gentile projects and lifts in two steps.A. Frangioni, F. Furini and C. Gentile, "Approximated perspective relaxations: a project and lift approach", Computational Optimization and Applications 63 (2016). And the choice between the second-order cone and the cuts is an empirical one. Frangioni and Gentile compared the two inside a branch and bound and found the LP with perspective cuts competitive with the conic relaxation, which is why the MILP-engine solvers prefer (4.3.1).A. Frangioni and C. Gentile, "A computational comparison of reformulations of the perspective relaxation: SOCP vs. cutting planes", Operations Research Letters 37 (2009).

(Two pitfalls) Two pitfalls are specific to the perspective. The first is the point \(z = 0\), where \(z f(x / z)\) is undefined and a naive NLP code divides by zero or loses convexity. The remedies are the conic form when the cost is quadratic and otherwise the \(\varepsilon\)-form of Furman, Sawaya and Grossmann from Section 4.2, which is exact at the integers for every \(\varepsilon\). The \((z + \varepsilon)\) form of Grossmann and Lee is correct only in the limit and should be avoided. The second is scope: the perspective is the hull of the separable substructure. With a non-separable quadratic the per-term perspective is strictly weaker than the rank-one, \(2 \times 2\) and full semidefinite hulls of Section 4.5, and with coupling constraints through \(x\) it is a relaxation and not the hull.

Where this is used

SCIP's nonlinear handler perspective detects semicontinuous variables, those for which the model's bounds and linear rows imply \(z = 0 \Rightarrow x = x^0\) and \(z = 1 \Rightarrow l \le x \le u\). For a nonlinear expression all of whose variables are semicontinuous with a common indicator, it strengthens every linear estimator it generates. An underestimator \(\ell\) of a convex \(h\) is extended by \((h(x^0) - \ell(x^0))(1 - z)\) so that it is tight at the off point. For \(h = x^2\) and \(x^0 = 0\) this gives exactly \(w \ge 2 \hat x\, x - \hat x^2 z\), the cut (4.3.1) with \(a = 1\) and \(c = 0\). The computational study behind it found consistent improvements for convex constraints and, for nonconvex ones, smaller trees and stronger root bounds but no significant change in mean time.K. Bestuzheva, A. Chmiela, B. Müller, F. Serrano, S. Vigerske and F. Wegscheider, "Global optimization of mixed-integer nonlinear programs with SCIP 8", Journal of Global Optimization 91 (2025), online December 2023, Section 2.6; K. Bestuzheva, A. Gleixner and S. Vigerske, "A computational study of perspective cuts", Mathematical Programming Computation 15 (2023); SCIP source, src/scip/nlhdlr_perspective.c (parameters convexonly, tightenbounds, adjrefpoint; master branch, accessed 4 October 2026). Whether Gurobi derives perspective cuts internally is not described in its documentation and is unverified. In the conic solvers, and in the cardinality-constrained portfolio codes of the next subsection, the perspective is modelled directly as a rotated second-order cone. In the sparse-regression solver L0BnB it is the relaxation solved at every node by coordinate descent.H. Hazimeh, R. Mazumder and A. Saab, "Sparse regression at scale: branch-and-bound rooted in first-order optimization", Mathematical Programming 196 (2022).

What parallelizes

The separation parallelizes elementwise, over indicators and over nodes, as the listing shows. So does the conic relaxation itself when it is solved by a first-order method, since the projection onto a rotated second-order cone is a closed-form operation on three numbers per indicator (Section 7.6).

The quadratic corner: cardinality, thresholds, factor models

This subsection brings the perspective to the application this series builds toward. Portfolio construction with integer decisions is a convex mixed-integer quadratic program: the continuous relaxation is a convex QP, and the difficulty is purely combinatorial. Since Bienstock's 1996 study, every advance in solving it to proven optimality has come from a tighter continuous relaxation, and the table later in this subsection lists them. The subsection defines the problem family, shows that the big-M relaxation of a cardinality limit is no relaxation at all, applies the perspective to the part of the covariance it can see, and introduces the factor model. The factor model makes every QP in the subject cheap and supplies the diagonal that the perspective needs. The three-asset instance of the figure is the running example R5 for the rest of the post.

Definition 4.4.1 (portfolio problems with integer decisions). Given expected returns \(\mu \in \mathbb{R}^n\) and a positive semidefinite covariance \(\Sigma\), the mean–variance problem in its return-floor form is

\[\min_{w \in \mathbb{R}^n}\ w^\top \Sigma w \quad \text{s.t.} \quad e^\top w = 1, \quad \mu^\top w \ge R, \quad w \ge 0,\]

with \(e\) the all-ones vector. The scalarized form minimizes \(\tfrac{\kappa}{2} w^\top \Sigma w - \mu^\top w\) instead, with a risk-aversion weight \(\kappa \ge 0\).H. Markowitz, "Portfolio selection", The Journal of Finance 7 (1952). The figure uses the return-floor form; D. Bertsimas and R. Cory-Wright, "A scalable algorithm for sparse portfolio selection", INFORMS Journal on Computing 34 (2022), equation (1), use the scalarized form. The cardinality-constrained problem adds a limit \(\lVert w \rVert_0 \le k\) on the number of nonzero positions, which with indicators \(z \in \{0, 1\}^n\) and bounds \(0 \le w_i \le u_i\) is the convex MIQP

\[\min_{w, z}\ w^\top \Sigma w \quad \text{s.t.} \quad e^\top w = 1, \quad \mu^\top w \ge R, \quad 0 \le w_i \le u_i z_i, \quad e^\top z \le k, \quad z \in \{0, 1\}^n ,\]

possibly with linear side constraints \(A w \le b\) for sector exposures and position limits.D. Bienstock, "Computational study of a family of mixed-integer quadratic programming problems", Mathematical Programming 74 (1996); Bertsimas and Cory-Wright (2022), Problem (2), whose Table 2 records which side constraints each paper in the literature imposed. L. Mencarelli and C. D'Ambrosio, "Complex portfolio selection via convex mixed-integer quadratic programming: a survey", International Transactions in Operational Research 26 (2019), survey the full list of real-life constraints. A minimum buy-in, or semicontinuous position, is \(w_i \in \{0\} \cup [l_i, u_i]\) with a threshold \(l_i > 0\), written \(l_i z_i \le w_i \le u_i z_i\). A round-lot constraint is \(w_i = m_i \rho_i\) with \(\rho_i \in \mathbb{Z}_+\) and \(m_i\) the lot size in weight units. A fixed plus linear transaction cost on a trade \(x_i\) is \(\phi_i(x_i) = 0\) if \(x_i = 0\) and \(\beta_i + \alpha_i |x_i|\) otherwise, with \(\beta_i > 0\) the fixed charge and \(\alpha_i \ge 0\) the proportional cost. It is discontinuous at zero and enters the budget constraint \(e^\top x + \sum_i \phi_i(x_i) \le 0\) of a trade list.M. S. Lobo, M. Fazel and S. Boyd, "Portfolio optimization with linear and fixed transaction costs", Annals of Operations Research 152 (2007), equation (16).

Proposition 4.4.2 (the big-M relaxation forgets the cardinality limit). If \(k \ge 1\) and \(u_i \ge 1\) for every \(i\), the continuous relaxation of the cardinality-constrained problem above, with \(z \in [0, 1]^n\), has the same optimal value as the mean–variance problem without the limit.

Proof. Any feasible \(w\) of the unconstrained problem satisfies \(0 \le w_i \le 1\) and \(e^\top w = 1\). Set \(z = w\). Then \(w_i \le u_i z_i\) holds because \(u_i \ge 1\), and \(e^\top z = 1 \le k\). So every unconstrained portfolio is feasible for the relaxation, and the relaxation can be no better than the unconstrained minimum. It is also no worse, since dropping the limit only enlarges the feasible set. ∎

(Why the big-M bound is blind to the limit) The relaxation therefore returns the unconstrained optimum unchanged, and the bound it gives says nothing about the limit at all. The perspective of Section 4.3 sees the limit, but only through the part of the objective that is separable, and a covariance matrix is not separable unless the assets are uncorrelated. The device is to split off a diagonal part \(D\), leaving a remainder \(\Sigma - D\) that is still positive semidefinite, written \(\Sigma - D \succeq 0\), so that the diagonal terms \(D_{ii} w_i^2\) can each take the perspective while the remainder is kept as it is.

Proposition 4.4.3 (perspective reformulation with a diagonal split). Let \(D\) be diagonal with \(D \succeq 0\) and \(\Sigma - D \succeq 0\). Then \(w^\top \Sigma w = w^\top (\Sigma - D) w + \sum_i D_{ii} w_i^2\), and the cardinality-constrained problem is equivalent to the mixed-integer second-order cone program

\[\min_{w, z, \theta}\ w^\top (\Sigma - D) w + \sum_i D_{ii} \theta_i \quad \text{s.t.} \quad w_i^2 \le \theta_i z_i, \quad 0 \le w_i \le u_i z_i, \quad e^\top w = 1, \quad \mu^\top w \ge R, \quad e^\top z \le k, \quad z \in \{0, 1\}^n ,\]

whose continuous relaxation is at least as tight as the big-M relaxation.A. Frangioni and C. Gentile, "SDP diagonalizations and perspective cuts for a class of nonseparable MIQP", Operations Research Letters 35 (2007); Aktürk, Atamtürk and Gürel (2009); X. Zheng, X. Sun and D. Li, "Improving the performance of MIQP solvers for quadratic programs with cardinality and minimum threshold constraints: a semidefinite program approach", INFORMS Journal on Computing 26 (2014); Bertsimas and Cory-Wright (2022), equation (5).

Proof. The split is an identity. At \(z_i = 1\) the cone forces \(\theta_i \ge w_i^2\), and at \(z_i = 0\) it forces \(w_i = 0\) with \(\theta_i \ge 0\), so at integral \(z\) the program is the original one with \(\theta_i = w_i^2\) at the optimum. In the relaxation, \(\theta_i \ge w_i^2 / z_i \ge w_i^2\) for \(z_i \in (0, 1]\), so the relaxed objective is at least \(w^\top \Sigma w\) at every feasible point, which is what the big-M relaxation charges. ∎

Proposition 4.4.3: the diagonal split and the perspective cones

  D diagonal, D >= 0 and Sigma - D >= 0 (both semidefinite)

  w' Sigma w  =  w' (Sigma - D) w  +  sum_i D_ii w_i^2
                 |                    |
                 | kept as it is      | each D_ii w_i^2 becomes
                 |                    | D_ii theta_i, with the cone
                 v                    v w_i^2 <= theta_i z_i
  minimize       w' (Sigma - D) w  +  sum_i D_ii theta_i

  z_i = 1          theta_i >= w_i^2
  z_i = 0          w_i = 0 and theta_i >= 0
  z integral       the original program, with theta_i = w_i^2
                   at the optimum

  0 < z_i <= 1     theta_i >= w_i^2 / z_i >= w_i^2: the relaxed
  (relaxation)     objective is at least w' Sigma w, which is what
                   the big-M relaxation charges

How much of the covariance to put into \(D\), and how to choose it, is the subject of Section 4.5. In one case the choice is forced and the answer is exact.

Proposition 4.4.4 (uncorrelated assets). Let \(\Sigma = \operatorname{diag}(\sigma_1^2, \dots, \sigma_n^2)\) with every \(\sigma_i > 0\), take \(D = \Sigma\), \(u_i = 1\), and drop the return floor. Then the continuous relaxation of the program of Proposition 4.4.3 has an optimal solution with \(z \in \{0, 1\}^n\): it holds the \(k\) assets of smallest variance with weights proportional to \(1 / \sigma_i^2\), and its value is the true optimum.

Proof. With \(D = \Sigma\) the relaxed objective is \(\sum_i \sigma_i^2 \theta_i\) with \(\theta_i \ge w_i^2 / z_i\), so for fixed \(z\) the relaxation is \(\min\{\sum_i \sigma_i^2 w_i^2 / z_i : e^\top w = 1,\ 0 \le w_i \le z_i\}\). Dropping the upper bounds \(w_i \le z_i\) can only lower this minimum. Without them, the Cauchy–Schwarz inequality \((\sum_i w_i)^2 \le (\sum_i \sigma_i^2 w_i^2 / z_i)(\sum_i z_i / \sigma_i^2)\) with \(\sum_i w_i = 1\) gives the value \(1 / \sum_i (z_i / \sigma_i^2)\), attained at \(w_i \propto z_i / \sigma_i^2\). So for every \(z\) the relaxation's value is at least \(1 / \sum_i (z_i / \sigma_i^2)\), and the relaxation's optimum is at least the smallest value of this bound over \(e^\top z \le k\), \(0 \le z \le 1\). Making the bound small means making \(\sum_i z_i / \sigma_i^2\) large, a fractional knapsack with unit weights, that is, the LP relaxation of a knapsack problem of the kind Section 1.6 introduced, whose optimum puts \(z_i = 1\) on the \(k\) largest coefficients \(1 / \sigma_i^2\), that is, on the \(k\) smallest variances, and zero elsewhere. At that integral \(z\) the solution \(w_i \propto z_i / \sigma_i^2\) has \(w_i \le 1 = z_i\) on the held assets and \(w_i = 0 = z_i\) elsewhere. The dropped bounds hold, so the lower bound is attained and is the relaxation's value. The same \((w, z)\) is feasible for the original problem. Its value is the smallest variance any \(k\)-asset portfolio can have, because the minimum variance over a support \(T\) is \(1 / \sum_{i \in T} \sigma_i^{-2}\) and that sum is largest for the \(k\) smallest variances. ∎

(The three-asset instance, R5) The figure's instance has correlated assets, so Proposition 4.4.4 does not apply to it directly. The script below includes a check at zero correlation and a slack return floor, where it does apply: the relaxation closes the whole gap and lands on the two-asset optimum. The figure draws three assets with expected returns \(10\%\), \(7\%\) and \(4\%\) and volatilities \(20\%\), \(15\%\) and \(10\%\). One slider sets a common pairwise correlation \(\rho\), the other sets a return floor \(R\), and the limit is "at most two assets". Writing \(\sigma\) for the vector of volatilities, the covariance is \(S = \rho\, \sigma \sigma^\top + (1 - \rho)\operatorname{diag}(\sigma^2)\), a one-factor model in the sense of Definition 4.4.6 below, which Section 4.5 uses. In the notation of Definition 4.4.1, \(n = 3\), \(\mu = (0.10, 0.07, 0.04)\), \(\Sigma = S\), every \(u_i = 1\), \(k = 2\), and the return floor is \(\mu^\top w \ge R\). Every portfolio is a point of the triangle, and with at most two assets held the feasible set is the three edges above the return line, a union of three segments that is not convex. The figure's perspective relaxation uses the uniform split \(D = d I\) with \(d = 0.99\, \lambda_{\min}(S)\), the \(0.99\) keeping the remainder \(S - dI\) positive definite; in Proposition 4.4.3 the three cones then read \(w_i^2 \le \theta_i z_i\) and the objective is \(w^\top (S - dI)\, w + d\,(\theta_1 + \theta_2 + \theta_3)\). At \(\rho = 0.3\), \(\lambda_{\min}(S) = 0.00820\) and \(d = 0.00812\).

(Reading the figure: three marked points) The thing to look for is where the three marked points sit: the big-M relaxation's answer lies strictly inside the triangle, ignoring the limit, while the perspective bound lies between it and the edge. In the default view, \(R = 7\%\) and \(\rho = 0.3\), the unconstrained optimum holds all three assets, \(w = (0.32, 0.36, 0.32)\) with volatility \(11.12\%\), and by Proposition 4.4.2 that is also the big-M bound. The perspective relaxation gives \(w = (0.31, 0.38, 0.31)\) with volatility \(11.71\%\). Its indicators are \(z = 2w\), because the budget \(\sum_i z_i \le 2\) is tight and the inner minimization over \(z\) scales \(z\) in proportion to \(w\). The script below prints \(z = (0.617, 0.767, 0.617)\), which the figure's coarser search rounds to \((0.62, 0.76, 0.62)\). The best two-asset portfolio drops asset 2 and holds \(w = (0.50, 0, 0.50)\) with volatility \(12.45\%\). Measured on the variance, \(0.01237\), \(0.01371\) and \(0.01550\), the perspective closes \(43\%\) of the gap.

(The other cases, and why the bound decides everything) The other cases move the sliders, and the script's last two blocks reproduce them. At \(R = 9\%\) the unconstrained optimum already holds two assets, \((0.67, 0.33, 0)\) with volatility \(15.58\%\), so all three values coincide and the limit costs nothing. At \(\rho = 0.8\) the big-M bound is \(14.00\%\), the perspective bound \(14.15\%\) and the best two-asset portfolio \(14.32\%\). The perspective then closes \(46\%\) of the variance gap, and the whole gap is only \(0.32\) points of volatility, because assets that move together diversify little. A branch and bound on this instance has three leaves, one per asset dropped. For 500 assets and a limit of 20 the number of supports of exactly 20 assets is \(\binom{500}{20} \approx 2.7 \times 10^{35}\), and enumerating them at a billion a second would take about \(8.5 \times 10^{18}\) years. A search can only succeed if its bound prunes almost all of them, which is why the quality of the relaxation decides whether the problem can be solved at all.

Every portfolio of three assets with expected returns 10%, 7% and 4% and volatilities 20%, 15% and 10% is a point of the triangle, and the grey ellipses are lines of equal risk. With at most two assets held the feasible set shrinks to the blue edges above the return line: the orange ring is the unconstrained optimum, which the big-M relaxation returns unchanged, the orange dot is the perspective relaxation's bound, and the blue dot is the best two-asset portfolio. The right panel draws the frontiers without the limit (dashed orange, the big-M bound), with the perspective bound (orange) and with the limit (blue). The sliders set the target return R and the correlation between every pair.

The script recomputes the figure's default view, reproduces the figure's two other cases, and adds the second split the next subsection is about, \(D = (1 - \rho)\operatorname{diag}(\sigma^2)\), the specific-risk diagonal of the one-factor model. Each QP is solved exactly by enumerating active sets: with three variables, every choice of which bounds and which constraints hold with equality gives a small linear system, and the best feasible solution among them is the optimum. The perspective relaxation's inner minimization over \(z\) is solved in closed form. Its optimality conditions give \(z_i\) proportional to \(w_i\), clipped to \([w_i, 1]\) and scaled so that the budget \(\sum_i z_i = 2\) holds, a rule the code calls water-filling. The outer minimization over the triangle uses a grid followed by a compass search, a derivative-free descent that tries steps along fixed directions and halves the step when none improves. That is enough because the function of \(w\) is convex.

# Three assets, at most two: the cardinality figure.
#
# Minimum variance on the simplex with a return floor. The exact optimum
# by enumerating supports; the big-M relaxation (the unconstrained QP);
# the perspective relaxation with the figure's uniform split
# D = 0.99 lambda_min(S) I and with the factor split
# D = (1 - rho) diag(sigma^2); then the figure's two other cases,
# R = 9% and rho = 0.8. numpy only; deterministic.

import itertools

import numpy as np

MU = np.array([0.10, 0.07, 0.04])
SIG = np.array([0.20, 0.15, 0.10])
K = 2

def covariance(rho):
    return (1 - rho) * np.diag(SIG ** 2) + rho * np.outer(SIG, SIG)

def qp(S, R, support):
    """Min w'Sw, 1'w = 1, mu'w >= R, w >= 0, w = 0 off the support.

    Active-set enumeration: each choice of the constraints that hold
    with equality is one KKT system; the best feasible solution wins.
    """
    idx = list(support)
    k = len(idx)
    best = None
    opts = ([('ret', MU[idx], R)]
            + [('zero', np.eye(k)[j], 0.0) for j in range(k)])
    for mask in range(1 << len(opts)):
        rows, rhs = [np.ones(k)], [1.0]
        for t, (kind, row, b) in enumerate(opts):
            if mask >> t & 1:
                rows.append(row)
                rhs.append(b)
        if len(rows) > k:
            continue
        A = np.array(rows)
        Ss = S[np.ix_(idx, idx)]
        KKT = np.block([[2 * Ss, A.T],
                        [A, np.zeros((len(rows), len(rows)))]])
        try:
            w = np.linalg.solve(KKT, np.r_[np.zeros(k), rhs])[:k]
        except np.linalg.LinAlgError:
            continue
        if (w < -1e-9).any() or MU[idx] @ w < R - 1e-9:
            continue
        v = w @ Ss @ w
        if best is None or v < best[0] - 1e-15:
            best = (v, np.maximum(w, 0))
    return best

def exact(S, R):
    """The best two-asset portfolio."""
    best = None
    for sup in itertools.combinations(range(3), K):
        cand = qp(S, R, sup)
        if cand and (best is None or cand[0] < best[0] - 1e-15):
            best = cand + (sup,)
    return best

def water_fill(W, D, held):
    """z_i = clip(s w_i, w_i, 1), with s fixed by sum z = K.

    Bisection on the multiplier, for every row of W at once.
    """
    lo, hi = np.full(len(W), 1e-14), np.full(len(W), 1e4)
    for _ in range(60):
        mid = np.sqrt(lo * hi)
        z = np.where(held,
                     np.clip(np.sqrt(D / mid[:, None]) * W, W, 1.0),
                     0.0)
        over = z.sum(axis=1) > K
        lo = np.where(over, mid, lo)
        hi = np.where(over, hi, mid)
    return np.where(held,
                    np.clip(np.sqrt(D / hi[:, None]) * W, W, 1.0),
                    0.0)

def persp_z(W, D):
    """The z that attains the inner minimum.

    z_i = 1 on the held assets if at most K are held, else water-filling.
    """
    held = W > 1e-12
    z = np.where(held, 1.0, 0.0)
    need = held.sum(axis=1) > K
    if need.any():
        z[need] = water_fill(W[need], D, held[need])
    return z

def persp_penalty(W, D):
    """Min over z of sum D_i w_i^2 / z_i, w_i <= z_i <= 1, sum z <= K."""
    z = persp_z(W, D)
    terms = np.where(z > 0, D * W ** 2 / np.where(z > 0, z, 1), 0.0)
    return terms.sum(axis=1)

def persp_relax(S, D, R, N=240):
    """Grid over the triangle, then a compass search.

    The function of w is convex.
    """
    Sd = S - np.diag(D)
    val = lambda W: (np.einsum('ij,jk,ik->i', W, Sd, W)
                     + persp_penalty(W, D))

    # the grid: every point (i, j, N - i - j) / N above the return floor
    I, J = np.meshgrid(np.arange(N + 1), np.arange(N + 1), indexing='ij')
    keep = I + J <= N
    W = np.stack([I[keep] / N,
                  J[keep] / N,
                  1 - (I[keep] + J[keep]) / N], axis=1)
    W = W[W @ MU >= R - 1e-12]
    V = val(W)
    k = V.argmin()
    v, w = V[k], W[k]
    h = 1.0 / N

    # the compass search: 16 directions whose entries sum to zero, so
    # every step stays on the plane 1'w = 1; halve h when none improves
    angles = np.linspace(0, 2 * np.pi, 16, endpoint=False)
    dirs = np.array([[np.cos(t), np.sin(t), -np.cos(t) - np.sin(t)]
                     for t in angles])
    for _ in range(300):
        cand = w + h * dirs
        ok = (cand >= -1e-12).all(axis=1) & (cand @ MU >= R - 1e-12)
        if ok.any():
            vals = val(cand[ok])
            j = vals.argmin()
            if vals[j] < v - 1e-16:
                v, w = vals[j], cand[ok][j]
                continue
        h /= 2
        if h < 1e-9:
            break
    return v, w

def report(rho, R):
    S = covariance(rho)
    lam = np.linalg.eigvalsh(S)[0]
    Du = 0.99 * lam * np.ones(3)
    Df = (1 - rho) * SIG ** 2
    ex = exact(S, R)
    un = qp(S, R, (0, 1, 2))
    pu, wu = persp_relax(S, Du, R)
    pf, wf = persp_relax(S, Df, R)
    gap = ex[0] - un[0]

    def closed(p):
        if gap > 1e-12:
            return f"closes {(p - un[0]) / gap:.0%} of the variance gap"
        return "no gap to close"

    print(f"rho = {rho}, R = {R:.0%}: lambda_min(S) = {lam:.5f}")
    print(f"  big-M = unconstrained: sigma {np.sqrt(un[0]):.2%}, "
          f"w = {np.round(un[1], 2)}")
    print("  perspective, uniform split d = 0.99 lambda_min:")
    print(f"     sigma {np.sqrt(pu):.2%}, {closed(pu)}")
    print(f"     at w = {np.round(wu, 3)} with z = "
          f"{np.round(persp_z(wu[None, :], Du)[0], 3)}")
    print("  perspective, factor split D = (1 - rho) sigma^2:")
    print(f"     sigma {np.sqrt(pf):.2%}, {closed(pf)}")
    print(f"  best two-asset portfolio: sigma {np.sqrt(ex[0]):.2%}, "
          f"w = {np.round(ex[1], 2)} on assets "
          f"{tuple(i + 1 for i in ex[2])}")

# the figure's default view
report(0.3, 0.07)

# a diagonal covariance and a slack return floor: the factor split is
# the whole diagonal
report(0.0, 0.04)

# the figure's case R = 9%: the unconstrained optimum already holds two
# assets
report(0.3, 0.09)

# the figure's case rho = 0.8
report(0.8, 0.07)
rho = 0.3, R = 7%: lambda_min(S) = 0.00820
  big-M = unconstrained: sigma 11.12%, w = [0.32 0.36 0.32]
  perspective, uniform split d = 0.99 lambda_min:
     sigma 11.71%, closes 43% of the variance gap
     at w = [0.308 0.383 0.308] with z = [0.617 0.767 0.617]
  perspective, factor split D = (1 - rho) sigma^2:
     sigma 12.09%, closes 72% of the variance gap
  best two-asset portfolio: sigma 12.45%, w = [0.5 0.5] on assets (1, 3)
rho = 0.0, R = 4%: lambda_min(S) = 0.01000
  big-M = unconstrained: sigma 7.68%, w = [0.15 0.26 0.59]
  perspective, uniform split d = 0.99 lambda_min:
     sigma 8.08%, closes 61% of the variance gap
     at w = [0.103 0.245 0.653] with z = [0.295 0.705 1.   ]
  perspective, factor split D = (1 - rho) sigma^2:
     sigma 8.32%, closes 100% of the variance gap
  best two-asset portfolio: sigma 8.32%, w = [0.31 0.69] on assets (2, 3)
rho = 0.3, R = 9%: lambda_min(S) = 0.00820
  big-M = unconstrained: sigma 15.58%, w = [0.67 0.33 0.  ]
  perspective, uniform split d = 0.99 lambda_min:
     sigma 15.58%, no gap to close
     at w = [0.667 0.333 0.   ] with z = [1. 1. 0.]
  perspective, factor split D = (1 - rho) sigma^2:
     sigma 15.58%, no gap to close
  best two-asset portfolio: sigma 15.58%, w = [0.67 0.33] on assets (1, 2)
rho = 0.8, R = 7%: lambda_min(S) = 0.00250
  big-M = unconstrained: sigma 14.00%, w = [0.32 0.36 0.32]
  perspective, uniform split d = 0.99 lambda_min:
     sigma 14.15%, closes 46% of the variance gap
     at w = [0.308 0.383 0.308] with z = [0.617 0.767 0.617]
  perspective, factor split D = (1 - rho) sigma^2:
     sigma 14.23%, closes 72% of the variance gap
  best two-asset portfolio: sigma 14.32%, w = [0.5 0.5] on assets (1, 3)

The first block reproduces the figure: \(11.12\%\), \(11.71\%\) and \(12.45\%\), with \(43\%\) of the variance gap closed by the uniform split, and it prints the perspective minimizer \(w = (0.308, 0.383, 0.308)\) with \(z = 2w\). The factor split takes \(0.028\), \(0.0158\) and \(0.007\) into the perspective, against \(0.0081\) for every asset under the uniform split: \(3.4\) times, \(1.9\) times and \(0.86\) times as much, about twice as much in total. It closes \(72\%\) of the same gap, which is the point Section 4.5 makes in general. The second block is Proposition 4.4.4 in numbers. At zero correlation with a slack return floor the factor split is the whole diagonal, and the perspective relaxation lands exactly on the best two-asset portfolio, \(8.32\%\) on assets 2 and 3. The uniform split, with only \(d = 0.0099\), closes \(61\%\). The last two blocks are the figure's other cases: at \(R = 9\%\) all three values are \(15.58\%\), and at \(\rho = 0.8\) they are \(14.00\%\), \(14.15\%\) and \(14.32\%\), with \(46\%\) of the variance gap closed. The script's cost is dominated by the grid of about 29,000 points on the triangle, each a closed-form water-filling. The grid points are independent and the whole evaluation is one vectorized pass.

casebig-M bound B, the unconstrained optimumperspective, uniform split d = 0.99 λ_min(S): Uperspective, factor split D = (1 − ρ) diag(σ²): Fexact optimum, at most two assets held: E
ρ = 0.3, R = 7% (the figure's default view)11.12%11.71% (43%)12.09% (72%)12.45%
ρ = 0.0, R = 4% (Proposition 4.4.4)7.68%8.08% (61%)8.32% (100%)8.32%
ρ = 0.3, R = 9%15.58%15.58% (no gap to close)15.58% (no gap to close)15.58%
ρ = 0.8, R = 7%14.00%14.15% (46%)14.23% (72%)14.32%
The script's four runs: the perspective bounds in the variance gap

(Thirty years of relaxations, in one table) The methods that solved the cardinality problem to proven optimality differ mainly in their relaxation. Bienstock's 1996 method branched on subsets of the universe with surrogate constraints \(\sum_i w_i / u_i \le k\) and no binaries at all; the surrogate is the one row that the limit and the bounds imply together, since \(w_i / u_i \le z_i\) and \(\sum_i z_i \le k\). Bertsimas and Shioda built a branch and bound around Lemke's pivoting method, a simplex-like method for the linear complementarity problem that the optimality conditions of a QP form, with warm starts. It fathomed more nodes than CPLEX's MIQP solver on random instances but reached proven optimality only on small ones. The table below lists the largest instance each later method solved to proven optimality, as Bertsimas and Cory-Wright tabulate them, with a one-line description of each relaxation. Their own dualized outer approximation, which Section 4.9 derives, reached about 3,200 securities with certificates, on the Wilshire 5000 universe with cardinalities 10, 50, 100 and 200.Bienstock (1996), whose method is described here as Bertsimas and Cory-Wright (2022), Section 1.3, present it, because the original could not be read for this series and its problem sizes are unverified; D. Bertsimas and R. Shioda, "Algorithm for cardinality-constrained quadratic optimization", Computational Optimization and Applications 43 (2009); J. P. Vielma, S. Ahmed and G. L. Nemhauser, "A lifted linear programming branch-and-bound algorithm for mixed-integer conic quadratic programs", INFORMS Journal on Computing 20 (2008); P. Bonami and M. A. Lejeune, "An exact solution approach for portfolio optimization problems under stochastic and integer constraints", Operations Research 57 (2009); Frangioni and Gentile (2006, 2009); J. Gao and D. Li, "Optimal cardinality constrained portfolio selection", Operations Research 61 (2013); X. Cui, X. Zheng, S. Zhu and X. Sun, "Convex relaxations and MIQCQP reformulations for a class of cardinality-constrained portfolio selection problems", Journal of Global Optimization 56 (2013); Zheng, Sun and Li (2014); A. Frangioni, F. Furini and C. Gentile, "Approximated perspective relaxations: a project and lift approach", Computational Optimization and Applications 63 (2016). The sizes are those of Table 1 of Bertsimas and Cory-Wright (2022); the cardinalities are those of their Tables 7 to 9 in the arXiv version (v5, 2021) of the paper, which run the S&P 500, Russell 1000 and Wilshire 5000 universes with a 600-second limit on one thread. A lower bound explains why the sizes stalled for so long. Bienstock showed that on independent identically distributed securities a branch and bound must expand about \(2^{n/10}\) nodes to improve the sparsity-free bound, the bound that ignores the limit, by ten percent, so a method whose relaxation cannot see the limit cannot scale.D. Bienstock, "Eigenvalue techniques for convex objective, nonconvex optimization problems", IPCO 2010, LNCS 6080 (Springer, 2010), as reported by Bertsimas and Cory-Wright (2022).

methodthe relaxation at each nodesecurities
Vielma, Ahmed and Nemhauser (2008)lifted LP outer approximation of the cone100
Bonami and Lejeune (2009)exact approach under stochastic and integer constraints200
Frangioni and Gentile (2009)perspective cuts in an LP-based branch and bound400
Gao and Li (2013)Lagrangian and geometric (box, ball, ellipsoid) relaxations300
Cui, Zheng, Zhu and Sun (2013)convex relaxations and an MIQCQP reformulation300
Zheng, Sun and Li (2014)the perspective with a diagonal chosen by an SDP400
Frangioni, Furini and Gentile (2016)approximated perspective relaxations400
Bertsimas and Cory-Wright (2022)ridge dual and outer approximation (Section 4.9)about 3,200
The largest cardinality-constrained portfolio problem each method solved to proven optimality, as Bertsimas and Cory-Wright tabulate them (2022, Table 1); the last row's method adds a ridge term to the objective, which changes the problem (Section 4.9)

(Thresholds) Minimum buy-in thresholds are the on/off sets of Theorem 4.3.4(a) with the lower bound \(l_i\): the hull of \(\{w_i = 0, z_i = 0\} \cup \{l_i \le w_i \le u_i,\ z_i = 1\}\) together with a convex cost is described by the perspective and the scaled bounds \(l_i z_i \le w_i \le u_i z_i\). The threshold therefore costs nothing extra in the relaxation once the perspective is in place. A solver can find this structure itself: SCIP's perspective handler of Section 4.3 detects semicontinuous variables from the model's bounds and rows, and strengthens the cuts it generates for any convex term in them.

(Round lots) Round lots are a different kind of integer. With a lot of 100 shares at a price of 100 dollars, a 50,000-dollar account and a 3% target weight, the target is 1,500 dollars, or 15 shares, which is zero lots or one. One lot is 10,000 dollars, 20% of the account. The decision is a general integer \(\rho_i \in \{0, 1, 2, \dots\}\) rather than an indicator, its relaxation is the continuous weight, and the gap between the two is large exactly when the account is small relative to the lot. The early literature formulates minimum transaction lots, with and without fixed costs, as mixed-integer programs and solves realistic instances heuristically.R. Mansini and M. G. Speranza, "Heuristic algorithms for the portfolio selection problem with minimum transaction lots", European Journal of Operational Research 114 (1999); H. Kellerer, R. Mansini and M. G. Speranza, "Selecting portfolios with fixed costs and minimum transaction lots", Annals of Operations Research 99 (2000). The lots figure of Section 9 shows the same integrality for one stock's tax lots: the LP relaxation sells a third of a lot and promises 114 dollars more than any whole-lot sale can deliver.

(Fixed transaction costs) Fixed transaction costs have the shape of the perspective figure's cost, a jump at zero followed by a linear rather than a quadratic term, and their convex envelope is a line.

Proposition 4.4.5 (Lobo, Fazel and Boyd, 2007; the envelope of a fixed plus linear cost). On \([-l_i, u_i]\) with \(l_i, u_i > 0\), the convex envelope of \(\phi_i\) of Definition 4.4.1 is

\[\operatorname{vex}_{[-l_i, u_i]} \phi_i\,(x_i) \;=\; \begin{cases} (\beta_i / u_i + \alpha_i)\, x_i, & x_i \ge 0, \\ -(\beta_i / l_i + \alpha_i)\, x_i, & x_i \le 0, \end{cases}\]

a linear transaction cost. Replacing \(\phi_i\) by its envelope in the budget constraint enlarges the feasible set, so the relaxed problem's optimal expected wealth is an upper bound on the true optimum. The two agree only when every trade is \(-l_i\), \(0\) or \(u_i\).Lobo, Fazel and Boyd (2007), equation (17) and Sections 2.1 to 2.4, who write the envelope as \(\phi_i^{\mathrm{c.e.}}\).

Proof. On \([0, u_i]\) the cost is \(0\) at the origin and \(\beta_i + \alpha_i x_i\) elsewhere. The largest convex function below it is the chord from \((0, 0)\) to \((u_i, \beta_i + \alpha_i u_i)\), of slope \(\beta_i / u_i + \alpha_i\). The same holds on the left. ∎

Proposition 4.4.5 on [0, u_i]: the fixed plus linear cost and its convex envelope. The cost φ_i is 0 at x_i = 0 and β_i + α_i x_i elsewhere. Its convex envelope is the chord from (0, 0) to (u_i, β_i + α_i u_i), of slope β_i/u_i + α_i, a linear cost. On [−l_i, 0] the same holds with l_i in place of u_i.

(The chord regime, and the \(\ell_1\) relaxation) This is the chord regime of Proposition 4.3.6 again: when the fixed charge is large relative to the bound, the envelope is a chord, and it is exactly the \(\ell_1\) relaxation of statistics and signal processing. Lobo, Fazel and Boyd report that exhaustive search over the \(2^n\) sign patterns is impractical past about 10 to 20 assets, while the relaxation and their reweighting heuristic handle hundreds. They also derive the bounds \(u_i\) that make the envelope tight from the variance constraint itself, as \(x_i \le \bar\sigma \sqrt{(\Sigma^{-1})_{ii}} - w_i\) with \(\bar\sigma\) the portfolio's volatility limit, which is the bound-tightening logic of Section 2.6 applied by hand.

(The factor model) The last structure of this subsection is the one that makes everything else cheap. A covariance matrix estimated from returns is dense, and a dense \(n \times n\) matrix with \(n = 1{,}000\) has \(500{,}500\) distinct entries. Practitioners do not estimate it that way.

Definition 4.4.6 (factor risk model). A covariance has factor form if \(V = X \Omega X^\top + D\) with \(X \in \mathbb{R}^{n \times k}\) the exposures of the \(n\) assets to \(k \ll n\) factors, \(\Omega \in \mathbb{S}^k_{++}\) the factor covariance and \(D\) diagonal with \(D_{ii} > 0\) the specific, or idiosyncratic, variances. The risk of holdings \(h\) against a benchmark \(h_b\) splits into a systematic part \((h - h_b)^\top X \Omega X^\top (h - h_b)\) and a specific part \(\sum_i D_{ii} (h_i - h_{b, i})^2\), which is separable across assets.B. Rosenberg, "Extra-market components of covariance in security returns", Journal of Financial and Quantitative Analysis 9 (1974); R. C. Grinold and R. N. Kahn, Active Portfolio Management, 2nd ed. (McGraw-Hill, 1999). N. Moehle, M. J. Kochenderfer, S. Boyd and A. Ang, "Tax-aware portfolio construction via convex optimization", Journal of Optimization Theory and Applications 189 (2021), equations (2) and (3), use a commercial model with \(k = 72\) factors for about 1,000 securities.

With \(n = 1{,}000\) and \(k = 50\) the factor form stores \(1{,}000\) specific variances, \(50 \cdot 51 / 2 = 1{,}275\) entries of \(\Omega\) and \(50{,}000\) exposures, about \(52{,}000\) numbers in place of \(500{,}500\). The saving in arithmetic is larger than the saving in storage.

The factor form V = X Omega X' + D, n = 1,000 and k = 50

      V              X       Omega         X'             D
  +--------+       +---+     +---+     +--------+     +--------+
  |        |       |   |     |   |     | k x n  |     |\       |
  | n x n  |   =   |   |  x  +---+  x  +--------+  +  | \      |
  | dense  |       |   |     k x k                    |  \     |
  |        |       |   |                              |   \    |
  +--------+       +---+                              +--------+
                   n x k                              diagonal

  500,500          50,000    1,275                    1,000
  distinct         exposures entries                  specific
  entries                                             variances

  stored: about 52,000 numbers in place of 500,500

Proposition 4.4.7 (cost of a factor-model QP). Let \(V = F F^\top + D\) with \(F = X C \in \mathbb{R}^{n \times k}\), where \(C C^\top = \Omega\) is a Cholesky factor, and \(D \succ 0\) diagonal. Introducing \(y = F^\top h \in \mathbb{R}^k\) as \(k\) extra variables with \(k\) extra equality rows turns the Hessian of a QP in \(h\) with \(m_c\) linear equalities into the diagonal \(\operatorname{diag}(D, I_k)\). The KKT system of size \(n + 2k + m_c\) is then solved in \(O\big(n (k + m_c)^2 + (k + m_c)^3\big)\) operations by eliminating \(h\) first. Equivalently, by the Woodbury identity, \(V^{-1} = D^{-1} - D^{-1} F (I_k + F^\top D^{-1} F)^{-1} F^\top D^{-1}\), so every product \(V^{-1} r\) costs \(O(nk)\) after an \(O(nk^2 + k^3)\) factorization.

Proof sketch. With \(h\) and \(y\) eliminated through the diagonal block \(\operatorname{diag}(D, I_k)\), what remains is a dense Schur complement of order \(k + m_c\), the matrix that the remaining unknowns see once the eliminated block has been substituted out, formed in \(O(n (k + m_c)^2)\) operations and factorized in \(O((k + m_c)^3)\). ∎

Proposition 4.4.7: the factor-model QP, from dense to diagonal

  Hessian in h           V = F F' + D            n x n, dense
          |
          | y = F'h: k extra variables and k extra equality rows
          v
  Hessian in (h, y)      diag(D, I_k)            diagonal

  the KKT system, of size n + 2k + m_c:

                        n + k            k + m_c
                  +----------------+----------------+
           n + k  |  diag(D, I_k)  |  the rows      |
                  |                |  below,        |
                  |                |  transposed    |
                  +----------------+----------------+
         k + m_c  |  the k rows    |                |
                  |  y = F'h and   |       0        |
                  |  the m_c rows  |                |
                  +----------------+----------------+
                           |
                           | eliminate h and y first, through
                           | the diagonal block diag(D, I_k)
                           v
                  dense Schur complement of order k + m_c:
                  formed in O(n (k + m_c)^2) operations,
                  factorized in O((k + m_c)^3)

  equivalently (Woodbury): every V^-1 r in O(nk) after an
  O(nk^2 + k^3) factorization

(What the factor form buys: cheap QPs and the perspective's diagonal) A thousand assets and seventy factors therefore give a QP whose every step scales with \(n k^2\) rather than \(n^3\). The same fact describes a batch of such QPs that share one risk model and differ only in their bounds, which is what a branch-and-bound frontier on a portfolio problem consists of. Such a batch is a set of small dense Schur complements plus large elementwise work on the diagonal. The factor model also hands the perspective its diagonal. \(D\) itself is admissible in Proposition 4.4.3 because \(V - D = X \Omega X^\top \succeq 0\), and it is the diagonal one believes in, the variance that is specific to each name. In the figure's one-factor instance this is the split that closed \(72\%\) of the gap against \(43\%\) for the largest uniform diagonal. Section 4.5 explains why the uniform split is so much smaller, and what to do when no factor model is given.

assetuniform split: d = 0.99 λ_min(S), the same for every assetfactor split: D_ii = (1 − ρ) σ_i², the specific riskfactor / uniform
10.00810.0283.4 times
20.00810.01581.9 times
30.00810.0070.86 times
The diagonal each split hands the perspective, at rho = 0.3

Where this is used

With the perspective in conic form the cardinality problem is a mixed-integer second-order cone program, the class the conic branch-and-bound solvers of CPLEX, Gurobi, Xpress and MOSEK accept directly. Section 4.8 gives the outer approximation that the LP-based codes use for the cone. CPLEX incorporated the lifted polyhedral approximations of Vielma and coauthors from version 12.6.2, so a comparison against its mixed-integer conic solver is a comparison against that line of work.Bertsimas and Cory-Wright (2022), footnote 3, which is the source for the CPLEX version; Moehle et al. (2021), Section 6, who solved their 744 monthly tax-lot instances with \(n = 998\) and \(k = 72\) as MIQPs with CPLEX 12.9 under a 300-second limit, and list MOSEK's mixed-integer conic solver among the alternatives. The production record is older than the theory. Bertsimas, Darnell and Soucy describe a 1996–1998 system at GMO, ten thousand lines of Fortran 90 around CPLEX 4.0, for portfolios of 1,000 to 1,500 securities with eight to ten times as many variables. One instance took up to 15 hours before reformulation and strengthening brought a 240-problem monthly simulation to 0.4 to 4.0 minutes per problem.D. Bertsimas, C. Darnell and R. Soucy, "Portfolio construction through mixed-integer programming at Grantham, Mayo, Van Otterloo and Company", Interfaces 29 (1999), Table 1.

What parallelizes

The factor structure is what parallelizes. The \(n \times k\) products with a shared exposure matrix are a batched dense multiplication across the nodes of a frontier or across accounts that share the risk model. The specific-risk terms and the perspective terms are elementwise over assets, and the Schur complements of Proposition 4.4.7 are a batched small solve. Section 9 builds on exactly this shape.

When the quadratic is not separable

The perspective only sees the diagonal. A covariance matrix with correlated assets couples every pair of weights through its off-diagonal entries, and Proposition 4.4.3 relaxes it by splitting off a diagonal part \(D\) and applying the perspective to that part alone. This subsection asks how much to split off. It shows on a two-asset example that the right amount is exactly the smallest eigenvalue, and it explains why a factor model's specific-risk diagonal is usually far larger than any uniform diagonal. It then states what is known about relaxations that go beyond the diagonal: the rank-one, \(2 \times 2\) and full convex hulls of a quadratic with indicators, which are conic or semidefinite programs in an extended space. Throughout, \(A \preceq B\) between symmetric matrices means that \(B - A\) is positive semidefinite, so \(0 \preceq D \preceq \Sigma\) says that \(D\) and \(\Sigma - D\) are both valid covariances.

Proposition 4.5.1 (monotonicity of the diagonal split). Let \(\Sigma \succeq 0\) and let \(D\) and \(D'\) be diagonal with \(0 \preceq D \preceq D' \preceq \Sigma\). Then the continuous relaxation of the program of Proposition 4.4.3 with \(D'\) has a value at least as large as the one with \(D\), and both are at least the big-M relaxation's value. The condition \(\Sigma - D \succeq 0\) is what makes the relaxation a convex program. Past it, the remainder is indefinite and the relaxation certifies nothing.Frangioni and Gentile (2007); Zheng, Sun and Li (2014); A. Frangioni, C. Gentile and J. Hungerford, "Decompositions of semidefinite matrices and the perspective reformulation of nonseparable quadratic programs", Mathematics of Operations Research 45 (2020), which treats general semidefinite decompositions of which the diagonal split is the simplest case.

Proof. At a feasible point of either relaxation, with \(z_i \in (0, 1]\) and \(\theta_i = w_i^2 / z_i\) at the optimum, the relaxed objective is \(w^\top (\Sigma - D) w + \sum_i D_{ii} w_i^2 / z_i = w^\top \Sigma w + \sum_i D_{ii} w_i^2 (1 / z_i - 1)\). The second term is nonnegative and increases with every \(D_{ii}\), so moving mass from the remainder into the diagonal can only raise the value at every feasible \((w, z)\), hence also the minimum. With \(D = 0\) the value is \(w^\top \Sigma w\), the big-M relaxation's. If \(\Sigma - D\) has a negative eigenvalue the first term is a nonconvex quadratic, and the minimum of the relaxation over the relaxed set is no longer the value a convex solver computes. ∎

(Which diagonal: no single largest one) So the modeller wants the largest diagonal that keeps \(\Sigma - D\) positive semidefinite, and "largest" is not a single thing: the set of admissible diagonals is convex but has no greatest element, because two admissible diagonals need not have an admissible entrywise maximum. For the two-asset matrix \(S\) below, \(D = \operatorname{diag}(d_1, d_2)\) is admissible exactly when \((1 - d_1)(1 - d_2) \ge \rho^2\) with \(d_i \le 1\), so \((1 - \rho^2, 0)\) and \((0, 1 - \rho^2)\) are both admissible while their entrywise maximum \((1 - \rho^2, 1 - \rho^2)\) is not. The two-asset example shows the simplest case, where a uniform diagonal \(\delta I\) is the natural family and the right \(\delta\) is determined.

Proposition 4.5.2 (two correlated assets, at most one held). Let \(S = \begin{pmatrix} 1 & \rho \\ \rho & 1 \end{pmatrix}\) with \(|\rho| < 1\), weights \(w_1 + w_2 = 1\), \(w \ge 0\), and at most one asset held. The true optimum is \(1\), one asset held. The plain relaxation, which forgets the limit, holds both and gives \((1 + \rho) / 2\). The eigenvalues of \(S\) are \(1 + \rho\) along \((1, 1)\) and \(1 - \rho\) along \((1, -1)\). For \(0 \le \delta \le 1 - \rho\) the perspective relaxation after extracting \(\delta I\) has value

\[\operatorname{bound}(\delta) \;=\; \min_{t \in [0, 1]}\ \big[\, 1 - 2 t (1 - t)(1 - \rho - \delta) \,\big] \;=\; \frac{1 + \rho + \delta}{2},\]

a straight line in \(\delta\) that reaches the true optimum exactly at \(\delta = 1 - \rho\). For \(\rho \ge 0\) this is \(\lambda_{\min}(S)\) and the gap closes. For \(\rho < 0\) the smallest eigenvalue is \(1 + \rho\), along the direction \((1, 1)\) that the budget constraint fixes, and the extraction must stop at \(\delta = 1 + \rho\) with the gap still open.

Proof. The relaxation is \(\min\ w^\top (S - \delta I) w + \delta (w_1^2 / z_1 + w_2^2 / z_2)\) subject to \(w_i \le z_i \le 1\), \(z_1 + z_2 \le 1\) and \(w_1 + w_2 = 1\). Since \(z_1 + z_2 \ge w_1 + w_2 = 1\) and \(z_1 + z_2 \le 1\), the constraints force \(z = w\), and the perspective terms equal \(\delta (w_1 + w_2) = \delta\). The Cauchy–Schwarz inequality gives the same lower bound, \(\delta (w_1 + w_2)^2 / (z_1 + z_2) \ge \delta\). On the simplex \(w = (t, 1 - t)\) the remainder is \(w^\top (S - \delta I) w = 1 - \delta - 2 t (1 - t)(1 - \rho - \delta)\). Adding the perspective terms, which equal \(\delta\), gives \(1 - 2 t (1 - t)(1 - \rho - \delta)\), the displayed function of \(t\). If \(1 - \rho - \delta > 0\) it is convex in \(t\) with minimum at \(t = \tfrac12\), value \(1 - (1 - \rho - \delta) / 2 = (1 + \rho + \delta) / 2\). If \(1 - \rho - \delta = 0\) it is flat and equal to \(1\). If it is negative, the function is concave with minimum \(1\) at a vertex, but then \(S - \delta I\) has the negative eigenvalue \(1 - \rho - \delta\) and the relaxation is not a convex program. The eigenvalue statements are read off \(S\). ∎

The figure's default instance is \(\rho = 0.8\): the true optimum is \(1\), the plain relaxation is \(0.9\), extracting \(\delta = 0.1\) gives \(0.85 + 0.1 = 0.95\), and extracting \(\delta = \lambda_{\min} = 0.2\) gives \(0.8 + 0.2 = 1.0\), so the gap is closed at the root with no branching. The script checks the closed form against a search over the simplex, prints the remainder's eigenvalues so that the convexity condition is visible, and runs the negative-correlation case with and without the budget shift the figure offers.

# Two assets, at most one held: the split figure.
#
# S = [[1, rho], [rho, 1]], weights on the simplex. The perspective
# bound after extracting delta I from S, in closed form and by a 1-D
# search, against the plain relaxation (1 + rho)/2 and the true
# optimum 1; the remainder's eigenvalues say when the relaxation is
# convex.

import numpy as np

def bound(rho, delta, mu=0.0):
    """The bound by a search over the simplex and in closed form.

    Returns the two, the eigenvalues of the remainder S - delta I and
    the smallest eigenvalue of S.
    """
    # mu * 11^T is the constant mu on the simplex
    S = np.array([[1 + mu, rho + mu],
                  [rho + mu, 1 + mu]])
    R = S - delta * np.eye(2)

    # w = (t, 1 - t); the perspective terms add back exactly delta
    ts = np.linspace(0, 1, 100001)
    W = np.stack([ts, 1 - ts], axis=1)
    numeric = (np.einsum('ij,jk,ik->i', W, R, W) + delta - mu).min()
    closed = (1 + rho + delta) / 2 if 1 - rho - delta > 0 else 1.0
    return (numeric, closed,
            np.linalg.eigvalsh(R), np.linalg.eigvalsh(S)[0])

for rho in (0.8, -0.5):
    lam = np.linalg.eigvalsh(np.array([[1, rho], [rho, 1]]))
    head = f"rho = {rho}: "
    print(f"{head}eigenvalues of S {lam[0]:.3f} "
          f"(along (1,-1) if rho > 0, else (1,1))")
    print(f"{' ' * len(head)}and {lam[1]:.3f}; "
          f"plain relaxation (1 + rho)/2 = {(1 + rho) / 2:.3f}")
    print("  delta   bound  closed form  remainder eigenvalues")
    for delta in (0.0, 0.1, 0.2, 0.3, 0.5):
        num, clo, eig, lmin = bound(rho, delta)
        psd = delta <= lmin + 1e-12
        status = ("convex program" if psd
                  else "NOT convex: no certified bound")
        print(f"  {delta:5.2f}  {num:6.4f}  {clo:11.4f}  "
              f"{eig[0]:+.3f}, {eig[1]:+.3f}  {status}")
    print()

# the budget shift: S + mu 11^T with mu = 0.5
print("                                   remainder")
print("rho   with          delta   bound  eigenvalues     lambda_min")
rho = -0.5
for delta in (0.5, 1.0, 1.5):
    num, clo, eig, lmin = bound(rho, delta, mu=0.5)
    print(f"-0.5  S + 0.5*11^T  {delta:5.1f}  {num:6.4f}  "
          f"{eig[0]:+.3f}, {eig[1]:+.3f}  (S + mu 11^T) = {lmin:.3f}")
rho = 0.8: eigenvalues of S 0.200 (along (1,-1) if rho > 0, else (1,1))
           and 1.800; plain relaxation (1 + rho)/2 = 0.900
  delta   bound  closed form  remainder eigenvalues
   0.00  0.9000       0.9000  +0.200, +1.800  convex program
   0.10  0.9500       0.9500  +0.100, +1.700  convex program
   0.20  1.0000       1.0000  +0.000, +1.600  convex program
   0.30  1.0000       1.0000  -0.100, +1.500  NOT convex: no certified bound
   0.50  1.0000       1.0000  -0.300, +1.300  NOT convex: no certified bound

rho = -0.5: eigenvalues of S 0.500 (along (1,-1) if rho > 0, else (1,1))
            and 1.500; plain relaxation (1 + rho)/2 = 0.250
  delta   bound  closed form  remainder eigenvalues
   0.00  0.2500       0.2500  +0.500, +1.500  convex program
   0.10  0.3000       0.3000  +0.400, +1.400  convex program
   0.20  0.3500       0.3500  +0.300, +1.300  convex program
   0.30  0.4000       0.4000  +0.200, +1.200  convex program
   0.50  0.5000       0.5000  +0.000, +1.000  convex program

                                   remainder
rho   with          delta   bound  eigenvalues     lambda_min
-0.5  S + 0.5*11^T    0.5  0.5000  +1.000, +1.000  (S + mu 11^T) = 1.500
-0.5  S + 0.5*11^T    1.0  0.7500  +0.500, +0.500  (S + mu 11^T) = 1.500
-0.5  S + 0.5*11^T    1.5  1.0000  +0.000, +0.000  (S + mu 11^T) = 1.500

(What the script shows: the budget direction and the shift) At \(\rho = 0.8\) the closed form and the search agree at every \(\delta\), and the bound climbs from \(0.9\) to \(1.0\) as \(\delta\) goes from \(0\) to \(0.2\). At \(\delta = 0.3\) the remainder has the eigenvalue \(-0.1\). The search still prints \(1.0\), because the minimum of a concave quadratic on the simplex is at a vertex, but that number is no longer the value of a convex program and no convex solver would certify it. At \(\rho = -0.5\) the smallest eigenvalue, \(0.5\), belongs to the direction \((1, 1)\), which the budget constraint fixes. The extraction therefore stops at \(\delta = 0.5\) with the bound at \(0.5\), a third of the way from \(0.25\) to \(1\), although the curvature along the simplex, \(1 - \rho = 1.5\), has not been used up. The last three lines show the remedy. Adding \(\mu \mathbf{1}\mathbf{1}^\top\) to \(S\) with \(\mu = 0.5\) changes the objective by the constant \(\mu\) on the simplex, where \(w^\top \mathbf{1}\mathbf{1}^\top w = 1\), and raises the eigenvalue along \((1, 1)\) to \(1.5\). So \(\delta\) can reach \(1.5\) and the gap closes. Using the linear constraints to reshape the decomposition in this way is one of the devices Frangioni, Gentile and Hungerford study for general quadratics. The script's cost is a vector of 100,001 quadratic forms per \(\delta\), all independent.

caseeigenvalue along (1, 1), fixed by the budgeteigenvalue along (1, −1), along the simplexδ stops atbound there
ρ = 0.81.80.20.21.0: the gap closes
ρ = −0.50.51.50.50.5: a third of the gap; the 1.5 along the simplex not used up
ρ = −0.5, with S + 0.5·11ᵀ1.51.51.51.0: the gap closes
What extracting delta I takes from each direction of S

The figure plots the bound against \(\delta\) for the same instance, with sliders for \(\delta\) and \(\rho\) and a toggle for the budget shift. The thing to look for is that the orange curve is a straight line, as Proposition 4.5.2 says, and that it reaches the true optimum exactly where the hatching begins. In the default view, \(\rho = 0.80\) and \(\delta = 0.10\), the figure reports the eigenvalues \(1.800\) and \(0.200\) and the plain relaxation \(0.900\). The bound at the cursor is \(0.950\), which is \(50\%\) of the gap, and the end mark at \(\delta = \lambda_{\min} = 0.200\) reads \(1.000\) with the gap closed. Everything to the right of \(\lambda_{\min}\) is hatched, and dragging the cursor there turns it red with the reading "no convex bound".

The other cases move \(\rho\). At \(\rho = 0.95\) the plain relaxation is \(0.975\) and \(\lambda_{\min} = 0.050\), so the whole gap closes within a sliver of \(\delta\). At \(\rho = 0\) the assets are uncorrelated, \(\lambda_{\min} = 1\), the bound is \(0.750\) at \(\delta = 0.5\) and \(1.000\) at \(\delta = 1\), which is Proposition 4.4.4 again. At \(\rho = -0.50\) with \(S\) as given the end mark reads \(0.500\) and \(33\%\) of the gap. With the toggle on, \(S + 0.5\,\mathbf{1}\mathbf{1}^\top = 1.5\, I\), \(\lambda_{\min} = 1.500\), the axis extends and the bound reaches \(1.000\) at \(\delta = 1.5\). At \(\rho = -0.9\) as given the extraction stops at \(0.100\) with only \(5\%\) of the gap closed. With the toggle, \(\lambda_{\min} = 1.900\), the bound is \(0.825\) at \(\delta = 1.55\) and the gap closes at \(1.9\).

(Two phrases in the figure) The readout prints the eigen-decomposition, the argument that the constraints force \(z = w\), the one-dimensional reduction, and a live check of the two numbers of the example above, \(0.95\) and \(1.0\). Two phrases in the figure deserve a gloss. The caption's qualifier "when \(\rho \ge 0\)" is there because \(\lambda_{\min}(S) = 1 - |\rho|\): for \(\rho < 0\) the smallest eigenvalue is \(1 + \rho\), along \((1, 1)\), and the extraction stops short of \(1 - \rho\) unless the budget shift is on. The toggle's label "free on the simplex" means that the added term \(\mu \mathbf{1}\mathbf{1}^\top\) is constant on the simplex, so it changes no minimizer and no gap.

Two assets with covariance S = [[1, ρ], [ρ, 1]], weights on the simplex, at most one asset held: the orange curve is the perspective bound after a diagonal part δI is extracted from S, plotted against δ, rising from the plain relaxation's (1 + ρ)/2 (grey) to the true optimum 1 (blue), which it reaches at δ = λmin, the smallest eigenvalue of S, when ρ ≥ 0. The hatched part is where S − δI is no longer positive semidefinite and the relaxation stops being a convex program. The sliders move δ and ρ, and the toggle adds μ·11ᵀ to S, a constant on the simplex, which lets δ reach 1 − ρ when ρ is negative.

(Three ways to choose the diagonal: uniform, factor, semidefinite) For many assets there are three ways to choose the diagonal, and the two-asset picture explains the ranking between them. The uniform split \(D = \lambda_{\min}(\Sigma)\, I\) is always admissible and is what the cardinality figure of Section 4.4 uses, scaled by \(0.99\). A covariance with a strong common factor, however, has a very small smallest eigenvalue, since a portfolio that is long some assets and short others can nearly cancel the factor, so the uniform split takes very little. The factor split takes \(D\) to be the specific-risk diagonal of a factor model. It is admissible because the remainder \(X \Omega X^\top\) is positive semidefinite, and it takes out of each asset the variance that is genuinely its own. In the three-asset instance that was \(0.028\), \(0.0158\) and \(0.007\) against \(0.0081\) per asset for the uniform split, about twice as much in total, and it closed \(72\%\) of the gap against \(43\%\). When no factor model is given, the semidefinite split chooses \(D\) by an optimization. One natural choice is the diagonal of largest trace among the admissible ones,

\[\max\ \sum_i D_{ii} \quad \text{s.t.} \quad \Sigma - D \succeq 0, \quad D \text{ diagonal}, \quad D \ge 0,\]

a semidefinite program with \(n\) variables. Frangioni and Gentile choose \(D\) by an auxiliary semidefinite program of this kind. Zheng, Sun and Li show that the \(D\) giving the tightest continuous relaxation of the cardinality problem is itself the solution of a semidefinite program that incorporates the problem's constraints. That is what makes their relaxation stronger than any fixed split.Frangioni and Gentile (2007); Zheng, Sun and Li (2014); Frangioni, Gentile and Hungerford (2020), who treat general semidefinite decompositions of \(\Sigma\), of which the diagonal split is the simplest, and the use of the problem's constraints to enlarge the admissible set, as the budget shift of the figure does in the plane. Bertsimas and Cory-Wright (2022), Section 3.2, "boost" their ridge regularizer with Frangioni and Gentile's diagonal. The ridge regularizer of Bertsimas and Cory-Wright, which Section 4.9 takes up, is a diagonal split in the other direction. A term \(\tfrac{1}{2\gamma} \lVert w \rVert^2\) is added to the problem rather than extracted from it, which changes the problem by a controlled amount and supplies separable curvature where the covariance has none.

(Beyond the diagonal: three hulls) The diagonal is where the perspective stops. What the convex hull of a quadratic with indicators actually is has a sequence of answers from the last decade. Each of the three results below is an instance of Theorem 4.2.4: one weight and one copy of the variables per indicator pattern, a \(0\)–\(1\) vector saying which indicators are on, with the quadratic restricted to that pattern.

Theorem 4.5.3 (Atamtürk and Gómez, 2025; the rank-one hull). Let \(T\) be an index set, \(a_i \ne 0\) for \(i \in T\), write \(z(T) = \sum_{i \in T} z_i\), and let \(Q^{r1}_T = \{(z, \beta, t) \in \{0, 1\}^{|T|} \times \mathbb{R}^{|T|} \times \mathbb{R}_+ : (a_T^\top \beta)^2 \le t,\ \beta_i (1 - z_i) = 0\ \forall i\}\). Then

\[\operatorname{conv}(Q^{r1}_T) \;=\; \Big\{(z, \beta, t) \in [0, 1]^{|T|} \times \mathbb{R}^{|T|} \times \mathbb{R}_+ : (a_T^\top \beta)^2 \le t,\ \ \frac{(a_T^\top \beta)^2}{z(T)} \le t \Big\},\]

that is, \(t \ge (a_T^\top \beta)^2 / \min\{1, z(T)\}\).A. Atamtürk and A. Gómez, "Rank-one convexification for sparse regression", Journal of Machine Learning Research 26 (2025), paper 35, 1–50; the preprint is arXiv 1901.10334 (2019), whose Section 3 states the result. For \(|T| = 1\) the inequality is the perspective \(a^2 \beta^2 \le t z\). The authors decompose \(Q = R + \sum_h A_h\) with rank-one \(A_h\) and apply the inequality to each piece; the resulting relaxation is a semidefinite program in an extended space, stronger than the perspective relaxation. A. Atamtürk and A. Gómez, "Supermodularity and valid inequalities for quadratic optimization with indicators", Mathematical Programming 201 (2023), show the underlying set function is supermodular and derive the hull in the original space by lifting supermodular inequalities.

The proof optimizes an arbitrary linear function over both sets and checks that the optima agree. The content is that a rank-one term \((a^\top \beta)^2\) with several indicators is charged as if a single indicator \(\min\{1, z(T)\}\) switched it on: a half-on pair of assets is charged half as much as a fully-on one, not a quarter.

(Two hypotheses of the next theorem) The next theorem needs two hypotheses that deserve words before the statement. The finite set \(\mathcal F\) lists inequalities \(\pi^\top z \ge 1\) that cut the origin off from the hull of the admissible indicator patterns \(Q\) without cutting anything else off. For the cardinality set \(\{z : \mathbf 1^\top z \le q\}\) the single inequality \(\mathbf 1^\top z \ge 1\) does it, so \(\mathcal F = \{\mathbf 1\}\). The connectivity condition says that any two indicators can be on together in some admissible pattern. It fails, for instance, when \(Q\) forbids two particular indicators from being on at once.

Theorem 4.5.4 (Wei, Gómez and Küçükyavuz, 2022; constraints on the indicators). Let \(Q \subseteq \{0, 1\}^p\), let \(f\) be a one-dimensional convex function with \(f(0) = 0\), and let \(Z_Q = \{(z, \beta, t) \in Q \times \mathbb{R}^p \times \mathbb{R} : t \ge f(\mathbf 1^\top \beta),\ \beta_i (1 - z_i) = 0\}\). Let \(\mathcal F \subseteq \mathbb{R}^p\) be a finite set with \(\operatorname{conv}(Q \setminus \{0\}) = \operatorname{conv}(Q) \cap \{z : \pi^\top z \ge 1\ \forall \pi \in \mathcal F\}\). If the graph on \([p]\) with an edge \(ij\) whenever some \(z \in Q\) has \(z_i = z_j = 1\) is connected, then

\[\operatorname{cl}\operatorname{conv}(Z_Q) \;=\; \Big\{(z, \beta, t) : z \in \operatorname{conv}(Q),\ t \ge f(\mathbf 1^\top \beta),\ t \ge (\pi^\top z)\, f\Big(\frac{\mathbf 1^\top \beta}{\pi^\top z}\Big)\ \forall \pi \in \mathcal F\Big\}.\]

For \(Q = \{0, 1\}^p\) and for the cardinality set \(\{z : \mathbf 1^\top z \le q\}\) with \(q \ge 2\), \(\mathcal F = \{\mathbf 1\}\) and the description is the single inequality \(t \ge (\mathbf 1^\top z)\, f(\mathbf 1^\top \beta / \mathbf 1^\top z)\). For \(q = 1\) the hull is the separable perspective \(t \ge \sum_i z_i f(\beta_i / z_i)\). Moreover, for a separable objective \(t \ge \sum_i f_i(\beta_i)\) the perspective reformulation \(t \ge \sum_i z_i f_i(\beta_i / z_i)\) with \(z \in \operatorname{conv}(Q)\) is the closure of the convex hull for every \(Q\): the perspective is ideal independently of the constraints on the indicators.L. Wei, A. Gómez and S. Küçükyavuz, "Ideal formulations for constrained convex optimization problems with indicator variables", Mathematical Programming 192 (2022).

(Why the perspective survives a cardinality limit) The last sentence is what justifies the use of the perspective in Section 4.4 under a cardinality limit. A limit on the number of positions, a budget on the number of trades, or any other constraint on \(z\) alone leaves the perspective of a separable objective ideal. In the two-asset example the constraints force \(z = w\), and at \(\delta = \lambda_{\min}\) the remainder \(\rho \mathbf{1}\mathbf{1}^\top\) is constant on the simplex, so the objective is effectively the separable one and the separable perspective is exact there. The three-asset figure is the case \(q = 2\) with coupling through the return floor and through the off-diagonal entries of \(S - D\), which is why its gap stays open.

Theorem 4.5.5 (Han, Gómez and Atamtürk, 2023; the \(2 \times 2\) hull). Let \(d_1 d_2 \ge 1\) and \(Z_+ = \{(z, x, t) \in \{0, 1\}^2 \times \mathbb{R}^2_+ \times \mathbb{R} : t \ge d_1 x_1^2 + 2 x_1 x_2 + d_2 x_2^2,\ x_i (1 - z_i) = 0\}\), with \(z\) the indicators and \(x\) the continuous variables. Then

\[\operatorname{cl}\operatorname{conv}(Z_+) \;=\; \Big\{(z, x, t) \in [0, 1]^2 \times \mathbb{R}^2_+ \times \mathbb{R}_+ : \exists\, \lambda \ge 0,\ w \in \mathbb{R}^2_+ \text{ with } z_1 + z_2 - 1 \le \lambda \le \min\{z_1, z_2\},\ t \ge \frac{d_1 (x_1 - w_1)^2}{z_1 - \lambda} + \frac{d_2 (x_2 - w_2)^2}{z_2 - \lambda} + \frac{d_1 w_1^2 + 2 w_1 w_2 + d_2 w_2^2}{\lambda}\Big\},\]

a conic-quadratic extended formulation, with an explicit piecewise rational description of its projection onto the original space. Using the \(2 \times 2\) hull as a building block for every pair \(i < j\) gives an extended semidefinite relaxation for general \(n\) that is at least as strong as Shor's semidefinite relaxation, the optimal perspective relaxation and the optimal rank-one relaxation.S. Han, A. Gómez and A. Atamtürk, "2×2-convexifications for convex quadratic optimization with indicator variables", Mathematical Programming 202 (2023), Section 5 for the general-\(n\) relaxation. The paper writes the indicators as \(x\) and the continuous variables as \(y\); the statement above uses this section's convention.

Theorem 4.5.5: one point per indicator pattern, averaged. The patterns 00, 10, 01 and 11 are the corners of the unit square, with the weights of Theorem 4.2.4, where λ ≥ 0 and z1 + z2 − 1 ≤ λ ≤ min(z1, z2).
Theorem 4.5.5: one point per indicator pattern, averaged

  the term of each pattern in t, the perspective of the quadratic
  restricted to it (w is the copy of x for the 11 block):
    00   none
    10   d_1 (x_1 - w_1)^2 / (z_1 - lambda)
    01   d_2 (x_2 - w_2)^2 / (z_2 - lambda)
    11   (d_1 w_1^2 + 2 w_1 w_2 + d_2 w_2^2) / lambda

(The \(2 \times 2\) hull as a four-term disjunction) The formula is an instance of Theorem 4.2.4. The four terms of the disjunction are the indicator patterns \(00\), \(10\), \(01\) and \(11\). The weight of the pattern \(11\) is \(\lambda\), the weight of the pattern in which only \(i\) is on is \(z_i - \lambda\), and the weight of \(00\) is \(1 - z_1 - z_2 + \lambda\). Each fraction is the perspective of the quadratic restricted to that pattern, and \(w\) is the copy of \(x\) for the \(11\) block. The general statement is the following.

Theorem 4.5.6 (Wei, Atamtürk, Gómez and Küçükyavuz, 2024; the hull of a convex quadratic with indicators). Let \(Q \succ 0\), \(Z \subseteq \{0, 1\}^n\), and \(X = \{(x, z, t) \in \mathbb{R}^n \times Z \times \mathbb{R} : t \ge x^\top Q x,\ x \circ (\mathbf 1 - z) = 0\}\). Let \(P = \operatorname{conv}\{(\hat{\mathbf 1}_S, \hat Q_S^{-1}) : \hat{\mathbf 1}_S \in Z\}\), where \(\hat Q_S^{-1}\) is the inverse of the principal submatrix \(Q_S\) padded with zeros to size \(n \times n\). Then

\[\operatorname{cl}\operatorname{conv}(X) \;=\; \Big\{(z, x, t) \in [0, 1]^n \times \mathbb{R}^{n + 1} : \exists\, W \in \mathbb{R}^{n \times n},\ \begin{pmatrix} W & x \\ x^\top & t \end{pmatrix} \succeq 0,\ (z, W) \in P \Big\},\]

and the relaxation \(\min\ a^\top x + b^\top z + \tfrac12 t\) over this set has an optimal solution with integral \(z\).L. Wei, A. Atamtürk, A. Gómez and S. Küçükyavuz, "On the convex hull of convex quadratic optimization problems with indicators", Mathematical Programming 204 (2024). The polytope \(P\) has dimension at most \(n(n + 1)/2\) but exponentially many vertices in general; for low-rank or sparse \(Q\) the paper gives smaller representations. The case of an M-matrix \(Q\), with nonpositive off-diagonal entries, is treated in A. Atamtürk and A. Gómez, "Strong formulations for quadratic optimization with M-matrices and indicator variables", Mathematical Programming 170 (2018).

(The full hull: one matrix per pattern) The smallest case shows what the theorem says. For \(n = 1\) with \(Q = q > 0\) and \(Z = \{0, 1\}\), the two patterns give the pairs \((0, 0)\) and \((1, 1/q)\), so \(P\) is the segment \(\{(z, z/q) : 0 \le z \le 1\}\), and the matrix inequality with \(W = z/q\) says \(t \ge x^2 / W = q x^2 / z\), which is the perspective of Section 4.3. In general the inverses in \(P\) come from a Schur complement. Fix a support \(S\) and a vector \(x\) supported on \(S\), and put \(W = \hat Q_S^{-1}\). The block matrix \(\begin{pmatrix} W & x \\ x^\top & t \end{pmatrix}\) is positive semidefinite exactly when \(t \ge x^\top W^{+} x = x_S^\top Q_S x_S\), with \(W^{+}\) the pseudo-inverse, so the matrix inequality says precisely that \(t\) is at least the quadratic restricted to the pattern \(S\). Each admissible pattern \(\mathbf 1_S\) therefore comes with its own matrix \(\hat Q_S^{-1}\), and \(P\) is the set of averages, with the weights of Theorem 4.2.4, of the pairs (pattern, inverse). The theorem is Theorem 4.2.4 with one matrix copy per pattern, and its content is that averaging the matrices is all that is needed: no further inequality on \((x, t)\) appears. It explains why the perspective, the rank-one hull and the \(2 \times 2\) hull are all projections of one object, and it also explains why none of them is the end of the story at scale. The exact hull is a semidefinite program whose description grows with the number of supports in \(Z\).

(The ladder for this problem class) The ladder of Section 4.1 is now concrete for this problem class. The big-M relaxation is a QP that does not see the limit. The perspective with a diagonal split is a second-order cone program of the original size plus \(n\) cones, exact for the separable part and for every constraint on \(z\) alone. The rank-one and \(2 \times 2\) hulls are conic or semidefinite programs with \(O(n^2)\) extra variables, strictly tighter, and the full hull is exact and exponential. Each rung costs a different convex solver at every node of the tree: a QP that warm-starts, an interior-point conic solve that does not, or a semidefinite solve for which the only large-scale methods are first-order (Section 7.6).

The ladder of Section 4.1 for the cardinality problem

  tighter                              cost
     ^
     |  full hull, exact               semidefinite, exponential
     |  (Theorem 4.5.6)
     |
     |  rank-one and 2 x 2 hulls       conic or semidefinite,
     |  (Theorems 4.5.3 and 4.5.5)     O(n^2) extra variables
     |
     |  perspective, diagonal split    second-order cone,
     |  (Proposition 4.4.3)            original size + n cones
     |
     |  big-M relaxation               QP
     +

Where this is used

The diagonal split is what every certified cardinality-constrained portfolio code since Frangioni and Gentile (2007) does, with the diagonal chosen by a factor model when one is given and by a semidefinite program when it is not. Zheng, Sun and Li's SDP-chosen split and Bertsimas and Cory-Wright's boosted ridge are the two forms in which it reached 400 and 3,200 securities. The rank-one and \(2 \times 2\) relaxations are, as of this writing, research codes rather than solver features. For the tax problem of Section 9 the choice is made by the data. The risk model is a factor model with \(k = 72\) factors, and its specific-risk diagonal is the separable curvature that the two-stage relaxation of Moehle and coauthors borrows to convexify the tax term. The same diagonal is what a perspective reformulation of the fixed-charge pieces would use.Moehle, Kochenderfer, Boyd and Ang (2021), Section 4.2, whose relaxation adds the specific-risk parabola \(\gamma_{\mathrm{risk}} D_{ii} (h^{\mathrm{init}}_i - h_{b,i} + u_i)^2\), with \(h^{\mathrm{init}}\) the pre-trade holdings and \(u\) the trade, to each asset's nonconvex tax term before taking the convex envelope. On the three-asset instance of Section 4.4 the factor split closes the whole big-M gap at return floors of \(5\%\) and \(6\%\) and \(72\%\) of it at \(7\%\) and \(8\%\), against \(43\%\) to \(57\%\) for the uniform split.

What parallelizes

Choosing the split is a one-time semidefinite program, or no computation at all with a factor model. Once chosen, the perspective terms are elementwise over assets, and the remainder \(w^\top (\Sigma - D) w\) with \(\Sigma - D = X \Omega X^\top\) is a product with the shared exposure matrix. The relaxation at every node of a frontier is a conic program with the same data and different bounds, the batched shape of Section 7.6. The semidefinite rungs are the open end. Whether the rank-one or \(2 \times 2\) relaxations can be solved in batches on a device at tree scale has not been tried, and the first-order semidefinite solvers that now reach very large matrix variables are the tool it would take.

Piecewise-linear models and MIP relaxations

The previous subsections replaced a nonconvex structure by the convex hull of the set it describes. This subsection goes the other way. It replaces a nonlinear function by a function made of straight pieces, so that a mixed-integer linear solver can take over. It then asks what that costs, in bound quality and in binary variables. Two cases must be kept apart. In the first the function is piecewise linear to begin with: a tax schedule with brackets, a tiered transaction cost, or the tax on a lot, whose rate changes once the lot has been held for a year. (The holding-period and wash-sale rules are stated in Section 9, with their sources and with the unverified flag they carry.) Here the only question is how to write the pieces so that the LP relaxation is strong. In the second case the function is smooth and the pieces approximate it. The piecewise-linear problem is then not a relaxation of the original. The interpolant lies above the function in some places and below it in others, so a solver that works with the interpolant alone proves nothing about the original problem. Geißler, Martin, Morsi and Schewe showed how to turn it into a relaxation by widening each piece by its interpolation error. Burlacu, Geißler and Schewe showed how to refine the pieces until that relaxation is tight.B. Geißler, A. Martin, A. Morsi and L. Schewe, "Using piecewise linear functions for solving MINLPs", in J. Lee and S. Leyffer (eds), Mixed Integer Nonlinear Programming, IMA Volumes 154 (Springer, 2012), 287–314; R. Burlacu, B. Geißler and L. Schewe, "Solving mixed-integer nonlinear programmes using adaptively refined mixed-integer linear programmes", Optimization Methods and Software 35 (2020), 37–64. The same construction applied to a bilinear term is the piecewise McCormick relaxation. It closes the subsection and connects it back to spatial branching.

The λ model and its LP relaxation

Definition 4.6.1 (breakpoints, interpolant, the λ model, SOS2). Let \(f\) be a function on \([a, b]\) and let \(a = x_0 < x_1 < \dots < x_K = b\) be breakpoints. The piecewise-linear interpolant \(\phi\) is the function that equals \(f\) at every breakpoint and is affine on each piece \([x_{v-1}, x_v]\). The λ model of the graph of \(\phi\) introduces one weight per breakpoint and writes

\[x = \sum_{v=0}^{K} \lambda_v x_v, \qquad y = \sum_{v=0}^{K} \lambda_v f(x_v), \qquad \sum_{v=0}^{K} \lambda_v = 1, \qquad \lambda \ge 0,\]

together with the condition that at most two of the weights are nonzero and that those two are adjacent. A vector with that property is a special ordered set of type 2, or SOS2. Beale and Tomlin gave it the name when they built a branching rule for it into a general mathematical programming system.E. M. L. Beale and J. A. Tomlin, "Special facilities in a general mathematical programming system for non-convex problems using ordered sets of variables", in J. Lawrence (ed.), Proceedings of the Fifth International Conference on Operational Research (1970), 447–454. The convex-combination model itself goes back to G. B. Dantzig, "On the significance of solving linear programming problems with some integer variables", Econometrica 28 (1960), 30–44, and to H. M. Markowitz and A. S. Manne, "On the solution of discrete programming problems", Econometrica 25 (1957), 84–110. With the SOS2 condition the pair \((x, y)\) is a point of the graph of \(\phi\), and every point of the graph arises in this way.

The SOS2 condition is the nonconvex part of the model. Dropping it gives the LP relaxation. The following theorem says exactly what that relaxation is.

Theorem 4.6.2 (the LP relaxation of the λ model is the convex hull of the breakpoints). Let \(V = \{(x_v, f(x_v)) : v = 0, \dots, K\}\). Dropping the SOS2 condition from the λ model gives a set whose projection onto \((x, y)\) is exactly \(\operatorname{conv}(V)\), and \(\operatorname{conv}(V) = \operatorname{conv}(\operatorname{gr}\phi)\). In particular the relaxation allows, at each \(x\), every \(y\) between the lower and the upper convex hull of the breakpoints. The minimum of \(y\) over the relaxation equals \(\min_x \phi(x)\).

Proof. Without the SOS2 condition the constraints say that \((x, y) = \sum_v \lambda_v (x_v, f(x_v))\) with \(\lambda\) in the simplex, which is the definition of \(\operatorname{conv}(V)\). Every point of \(\operatorname{gr}\phi\) lies on a segment between two consecutive breakpoints, so \(\operatorname{gr}\phi \subseteq \operatorname{conv}(V)\). Also \(V \subseteq \operatorname{gr}\phi\), so the two convex hulls coincide. The lower boundary of \(\operatorname{conv}(V)\) is the convex envelope of \(\phi\) on \([a, b]\). By Falk's theorem (Theorem 2.4.4) a function and its convex envelope have the same minimum, which is the last claim. ∎

(What the hull of the breakpoints looks like) The theorem says that the relaxation constrains \(y\) only through the breakpoints. For a convex curve the hull is the sliver between the interpolant and the chord across the whole interval, and the relaxation is nearly exact. For a curve that bends the other way the hull is a fat region, and the relaxation may place \(y\) anywhere inside it. The last sentence of the theorem is Theorem 2.4.4 again. A piecewise-linear objective on its own is relaxed exactly at the root. The gap appears only when \(x\) is coupled to other variables through constraints.

(The four-breakpoint example, and the SOS2 branch) A small example makes the band visible. Take the four breakpoints \((0, 0)\), \((1, 2)\), \((2, 1)\) and \((3, 4)\). In the λ model of Definition 4.6.1 they give \(x = \lambda_1 + 2 \lambda_2 + 3 \lambda_3\) and \(y = 2 \lambda_1 + \lambda_2 + 4 \lambda_3\) with the four weights summing to one; at \(x = 1\) the SOS2 choice \(\lambda_1 = 1\) gives \(y = 2\), while the non-adjacent choice \(\lambda_0 = \lambda_2 = \tfrac12\) also has \(x = 1\) and gives \(y = 0.5\). So at \(x = 1\) the interpolant has the value \(2\), but the relaxation allows every \(y\) between \(0.5\) and \(2\). The lower hull of the four points passes through \((0, 0)\) and \((2, 1)\), which gives \(0.5\) at \(x = 1\). The upper hull passes through \((0, 0)\) and \((1, 2)\), which gives \(2\). The point \((1, 0.5)\) is the average of two non-adjacent breakpoints, which is exactly what SOS2 forbids. Beale and Tomlin's branch repairs it without any binary variable. Branching at breakpoint \(1\) creates one child with \(\lambda_2 = \lambda_3 = 0\), in which \(x = 1\) forces \(\lambda_1 = 1\) and \(y = 2\). It creates a second child with \(\lambda_0 = 0\), in which \(x = 1\) again forces \(\lambda_1 = 1\) and \(y = 2\). SCIP's SOS2 constraint handler enforces the condition by exactly this branching, with no binaries in the model.SCIP Optimization Suite, source file src/scip/cons_sos2.c (master branch, read 4 October 2026), which cites Beale and Tomlin for the branching rule.

The four-breakpoint example: the LP relaxation at x = 1. The interpolant φ runs through (0, 0), (1, 2), (2, 1) and (3, 4). The convex hull of the breakpoints is bounded by the pieces on [0, 1] and [2, 3] and the chords (0, 0)–(2, 1) and (1, 2)–(3, 4). At x = 1 the relaxation allows every y between 0.5 and 2, where φ gives 2. The point (1, 0.5), the average of the non-adjacent breakpoints (0, 0) and (2, 1), is what SOS2 forbids.
Beale and Tomlin's branch at breakpoint 1, seen at x = 1

               root: the LP allows (1, 0.5), the average of
               breakpoints 0 and 2: lambda_0 = lambda_2 = 1/2
                  /                                  \
   lambda_2 = lambda_3 = 0                        lambda_0 = 0
                 |                                     |
   x = 1 forces lambda_1 = 1               x = 1 forces lambda_1 = 1
   y = 2                                   y = 2

  both children give y = 2 = phi(1); no binary variable enters

Writing the SOS2 condition with binaries

Most models instead encode the SOS2 condition with binary variables, because a MILP solver's cuts, presolve and heuristics all work on binaries. For a function with \(K\) pieces the textbook choices are listed below. The counts and the strength classification are those of Vielma, Ahmed and Nemhauser.J. P. Vielma, S. Ahmed and G. Nemhauser, "Mixed-integer models for nonseparable piecewise-linear optimization: unifying framework and extensions", Operations Research 58 (2010), 303–315, Table 1 and Theorems 1 to 4. Sharp and ideal are the two grades of strength of Definition 2.1.6: sharp when the projection of the LP relaxation is the convex hull of the set formulated, ideal when every vertex of the LP relaxation has integral binaries.

formulationcontinuous variablesbinariesLP relaxationstrength
SOS2 (no binaries; branching)\(K + 1\) weights0hull of breakpointssharp
convex combination (CC)\(K + 1\) weights\(K\)hull of breakpointssharp, not ideal
disaggregated CC (DCC)\(2K\) weights\(K\)hull of breakpointsideal
multiple choice (MC)\(K\) copies of \(x\)\(K\)hull of breakpointsideal
incremental (Inc)\(K\) fill fractions\(K - 1\)hull of breakpointsideal
logarithmic (Log)\(K + 1\) weights\(\lceil \log_2 K \rceil\)hull of breakpointsideal
disaggregated logarithmic (DLog)\(2K\) weights\(\lceil \log_2 K \rceil\)hull of breakpointsideal
Formulations of a piecewise-linear function with \(K\) pieces (\(K + 1\) breakpoints), after Vielma, Ahmed and Nemhauser (2010)

(Two of the formulations, written out) The convex-combination formulation adds a binary \(z_v\) per piece with \(\sum_v z_v = 1\) and \(\lambda_v \le z_v + z_{v+1}\). A positive weight then forces one of its two incident pieces to be selected. The incremental formulation fills the pieces in order. With fractions \(\delta_1, \dots, \delta_K \in [0, 1]\) and binaries \(z_1, \dots, z_{K-1}\) it writes \(x = x_0 + \sum_v \delta_v (x_v - x_{v-1})\) and \(y = f(x_0) + \sum_v \delta_v \big(f(x_v) - f(x_{v-1})\big)\), together with \(\delta_{v+1} \le z_v \le \delta_v\). A piece can therefore start to fill only when the previous one is full. One binary per interior breakpoint suffices, which is \(K - 1\).

Theorem 4.6.3 (Vielma, Ahmed and Nemhauser, 2010). Among the six binary formulations in the table, all except the convex-combination formulation are ideal. All six are sharp. Sharpness is a property of the function's own graph. It is not preserved when \(x\) is linked to other variables by further constraints.

(Sharp against ideal, inside a tree) The failure of idealness for the convex-combination formulation was shown by Lee and Wilson and by Padberg. The fact that every textbook formulation's LP relaxation replaces the cost by its lower convex envelope is due to Croxton, Gendron and Magnanti.J. Lee and D. Wilson, "Polyhedral methods for piecewise-linear functions I: the lambda method", Discrete Applied Mathematics 108 (2001), 269–285; M. Padberg, "Approximating separable nonlinear functions via mixed zero-one programs", Operations Research Letters 27 (2000), 1–5; K. L. Croxton, B. Gendron and T. L. Magnanti, "A comparison of mixed-integer programming models for nonconvex piecewise linear cost minimization problems", Management Science 49 (2003), 1268–1273. The models with a general polytope family, and the proof that all six are sharp, are in Vielma, Ahmed and Nemhauser (2010); A. B. Keha, I. R. de Farias and G. L. Nemhauser, "Models for representing piecewise linear cost functions", Operations Research Letters 32 (2004), 44–48, compare them computationally. The difference between sharp and ideal matters inside a tree. A sharp formulation gives the hull bound at the root. An ideal one also keeps the LP vertices integral, which is what rounding heuristics and branching rules see.

The logarithmic formulation reduces the number of binaries from \(K\) to \(\lceil \log_2 K \rceil\).

Theorem 4.6.4 (Vielma and Nemhauser, 2011; the logarithmic formulation). Let \(r = \lceil \log_2 K \rceil\) and let \(B : \{1, \dots, K\} \to \{0, 1\}^r\) be an injective code in which consecutive pieces receive codes that differ in exactly one bit (a Gray code). For \(l = 1, \dots, r\) let \(J^+(l)\) be the set of breakpoints all of whose incident pieces have bit \(l\) equal to \(1\). Let \(J^0(l)\) be the set of breakpoints all of whose incident pieces have bit \(l\) equal to \(0\). Then the λ model together with the binaries \(z \in \{0, 1\}^r\) and the \(2r\) constraints

\[\sum_{v \in J^+(l)} \lambda_v \le z_l, \qquad \sum_{v \in J^0(l)} \lambda_v \le 1 - z_l, \qquad l = 1, \dots, r,\]

is an ideal formulation of the graph of \(\phi\) with \(\lceil \log_2 K \rceil\) binary variables.J. P. Vielma and G. L. Nemhauser, "Modeling disjunctive constraints with a logarithmic number of binary variables and constraints", Mathematical Programming 128 (2011), 49–72. The combinatorial generalization, with a characterization of which disjunctions admit such codes, is J. Huchette and J. P. Vielma, "A combinatorial approach for small and strong formulations of disjunctive constraints", Mathematics of Operations Research 44 (2019), 793–820; the practical formulations and the JuMP package PiecewiseLinearOpt are J. Huchette and J. P. Vielma, "Nonconvex piecewise linear functions: advanced formulations and simple modeling tools", Operations Research 71 (2023), 1835–1856.

(The logarithmic formulation on three pieces) The idea is that a binary code with \(r\) bits names \(2^r \ge K\) pieces. The Gray-code property makes the constraints consistent at a breakpoint shared by two pieces: the two pieces differ in one bit, and that bit is left free. The sets \(J^+(l)\) and \(J^0(l)\) are easiest to see on three pieces. Take the breakpoints \(0, 1, 2, 3\) and give the pieces \([0,1]\), \([1,2]\), \([2,3]\) the codes \(00\), \(01\), \(11\). Bit \(1\) is \(1\) only on the third piece, so \(J^+(1) = \{3\}\), and it is \(0\) on both pieces incident to breakpoints \(0\) and \(1\), so \(J^0(1) = \{0, 1\}\). Breakpoint \(2\) belongs to neither set. Bit \(2\) is \(1\) on the second and third pieces, so \(J^+(2) = \{2, 3\}\) and \(J^0(2) = \{0\}\). The four rows are

\[\lambda_3 \le z_1, \qquad \lambda_0 + \lambda_1 \le 1 - z_1, \qquad \lambda_2 + \lambda_3 \le z_2, \qquad \lambda_0 \le 1 - z_2 .\]

At \(z = (0, 1)\), the code of the second piece, the first and fourth rows force \(\lambda_3 = \lambda_0 = 0\) and leave \(\lambda_1, \lambda_2\) free: the point lies on the second piece. At \(z = (1, 0)\), the unused code, the second and third rows force every weight to zero, which contradicts \(\sum_v \lambda_v = 1\), so the unused code is infeasible. Two binaries do the work for which the convex-combination model needs three. The count is the best possible. Any mixed-integer convex formulation of a set containing \(w\) points, no two of which have their midpoint in the set, needs at least \(\lceil \log_2 w \rceil\) integer variables. This is the midpoint lemma of Lubin, Vielma and Zadik, Theorem 4.10.3(c), which is proved in the subsection on representability below. The \(K\) midpoints of the pieces of a strictly convex or strictly concave curve are such a set, so no formulation with fewer than \(\lceil \log_2 K \rceil\) binaries exists.M. Lubin, J. P. Vielma and I. Zadik, "Mixed-integer convex representability", Mathematics of Operations Research 47 (2022), 720–749, the midpoint lemma and its corollary. Embedding formulations, the smallest ideal formulations that use only binary auxiliaries, are studied by Vielma.J. P. Vielma, "Embedding formulations and complexity for unions of polyhedra", Management Science 64 (2018), 4721–4734; J. P. Vielma, "Small and strong formulations for unions of convex sets from the Cayley embedding", Mathematical Programming 177 (2019), 21–53.

Theorem 4.6.4 on three pieces: the code and the sets J

  K = 3 pieces, r = ceil(log2 K) = 2 binaries z_1, z_2

  breakpoint v    0            1            2            3
                  *------------*------------*------------*
  piece               [0,1]        [1,2]        [2,3]
  code                 0 0          0 1          1 1
  (bit 1, bit 2)

  bit 1 at v      J^0          J^0          -            J^+
  bit 2 at v      J^0          -            J^+          J^+

  J^0: every piece at v has the bit 0; J^+: every piece has it 1;
  -: the two pieces at v differ in that bit, which is left free

  rows   lambda_3 <= z_1        lambda_0 + lambda_1 <= 1 - z_1
         lambda_2 + lambda_3 <= z_2      lambda_0 <= 1 - z_2

  z = (0, 1), the code of [1,2]:  lambda_0 = lambda_3 = 0, the
                                  point lies on the second piece
  z = (1, 0), the unused code:    every weight 0, which contradicts
                                  sum_v lambda_v = 1: infeasible

How many pieces, and what they cost

When the pieces approximate a smooth function, the number of pieces is set by the accuracy wanted, and the binary count follows.

Proposition 4.6.5 (interpolation error and the number of binaries). Let \(f \in C^2[a, b]\) with \(|f''| \le L\), and let \(\phi\) interpolate \(f\) on a uniform grid of step \(h\). Then \(\|f - \phi\|_\infty \le L h^2 / 8\). For \(f(x) = x^2\) the error is exactly \(h^2/4\), attained at the midpoint of every piece. To reach a tolerance \(\varepsilon\) on \([a, b]\) one needs \(K \ge (b - a)\sqrt{L / (8\varepsilon)}\) pieces. A logarithmic formulation then needs about \(\tfrac12 \log_2 (1/\varepsilon)\) binaries, up to a constant, and the incremental or multiple-choice formulations need \(K - 1\) or \(K\).

Proof. Fix a piece \([x_{v-1}, x_v]\) and a point \(x\) in its interior. The error \(e = f - \phi\) vanishes at both ends of the piece. Consider the auxiliary function \(\psi(t) = e(t) - e(x)\,\dfrac{(t - x_{v-1})(t - x_v)}{(x - x_{v-1})(x - x_v)}\) on the piece. It vanishes at \(x_{v-1}\), at \(x\) and at \(x_v\). By Rolle's theorem \(\psi'\) vanishes at two interior points, and by Rolle's theorem again \(\psi''\) vanishes at some \(\eta\) between them. Since \(\phi\) is affine on the piece, \(e'' = f''\), and \(\psi''(\eta) = 0\) reads \(e(x) = \tfrac12 f''(\eta)\,(x - x_{v-1})(x - x_v)\). The product \(|(x - x_{v-1})(x - x_v)|\) is at most \(h^2/4\), attained at the midpoint, so \(|e(x)| \le L h^2/8\). For \(x^2\) we have \(f'' = 2\) and the formula is exact: the error is \((x - x_{v-1})(x_v - x)\), a parabola with maximum \(h^2/4\) at the midpoint. The remaining claims are arithmetic. ∎

On \([0, 1]\) with the breakpoints \(0, \tfrac12, 1\) the interpolant of \(x^2\) is off by at most \(h^2/4 = 1/16\) per piece. On \([0, 2]\) with the breakpoints \(0, 1, 2\) the error is \(1/4\) at the midpoints, and four pieces instead of two quarter it to \(1/16\). The table below carries this to the tolerances a solver would use, for \(x^2\) on \([0, 2]\), where \(K \ge 1/\sqrt{\varepsilon}\).

\(\varepsilon\)pieces \(K\)logarithmic binariesincremental binaries
1e-1423
1e-21049
1e-332531
1e-4100799
1e-6100010999
Pieces and binaries needed for an interpolation error of at most \(\varepsilon\), \(f(x) = x^2\) on \([0, 2]\)

Four decimal places cost seven binaries in the logarithmic formulation and ninety-nine in the incremental one. The LP relaxation is the same polygon in both cases. The choice of formulation therefore changes the tree and not the root bound.

The figure below draws all of this for a chosen curve. The black curve is \(f\) and the purple line through the breakpoints is the interpolant that the MILP sees. The orange region is the convex hull of the breakpoints, the LP relaxation of Theorem 4.6.2. The red band is the construction of the next paragraphs. Each piece of the interpolant is moved up by the most it falls under the curve and down by the most it rises over it, so that the curve lies inside the band. The default view is \(x^2\) on \([0, 2]\) with \(k = 5\) breakpoints, that is four pieces. The largest interpolation error is \(0.0625\), at \(x = 0.25\) and at the midpoints of the other three pieces, as the proposition says. The formulation needs \(4\) binaries as a convex-combination model, \(3\) as an incremental model and \(2\) as a logarithmic model. The hull has area \(1.250\). The areas quoted in this paragraph are computed by the script after the figure, not read from the figure. The breakpoints lie under the chord \(y = 2x\), whose area over \([0, 2]\) is \(4\), and the trapezoids under the interpolant sum to \(2.75\). With twice as many pieces the largest error is \(0.0156\), a factor of \(4.00\) smaller. The factor stays exactly four for \(x^2\) because the error is \(h^2/4\). Switch to \(\sin x\) on \([0, 2\pi]\) and the hull at \(k = 5\) is a quadrilateral of area \(6.283\) that covers far more than the curve. The factors between successive doublings are \(3.0\), then \(3.7\), then \(3.9\). They approach four from below because the pieces are still long compared with the bends of the curve. For the cubic \(x^3 - 2x\) on \([-1.6, 1.6]\) the area is \(4.915\) and the factors are \(3.4\), \(3.7\) and \(3.9\). The sliver for a convex curve and the fat hull for a nonconvex one show the general fact: piecewise-linear models of nonconvex functions have weak LP relaxations, whatever binaries are used.

A curve (black) and the piecewise-linear interpolant through k equally spaced breakpoints (purple, the linearization the MILP sees). The orange region is the convex hull of the breakpoints, which is the LP relaxation of the SOS2 model. The red band is each piece of the interpolant moved up by the most it falls under the curve and down by the most it rises over it, so the curve lies inside it. Choose the curve and move k, and watch the largest error (the red mark) fall by a factor of four each time the pieces double while the binaries double with them.

The following script recomputes the figure's numbers and the four-breakpoint example. It finds the hull of the breakpoints by Andrew's monotone chain, which sorts the points and sweeps them once for the lower chain and once for the upper. It measures the interpolation error on a grid of 201 points per piece. It computes the hull area by the shoelace formula, the halved sum of cross products of consecutive vertices, and it reports the three binary counts.

# pwl_models.py -- the lambda (SOS2) model.
#
# The hull of the breakpoints, the interpolation error and the binary
# counts: the four-breakpoint example, the three curves of fig-pwl,
# and the error h^2/4 of x^2.

import numpy as np

def hull(P):
    """Andrew's monotone chain; returns the lower and the upper chain."""
    P = sorted(map(tuple, P))
    cross = lambda o, a, b: ((a[0] - o[0]) * (b[1] - o[1])
                             - (a[1] - o[1]) * (b[0] - o[0]))
    lo, up = [], []
    for p in P:
        while len(lo) >= 2 and cross(lo[-2], lo[-1], p) <= 0:
            lo.pop()
        lo.append(p)
    for p in reversed(P):
        while len(up) >= 2 and cross(up[-2], up[-1], p) <= 0:
            up.pop()
        up.append(p)
    return lo, up[::-1]

def chain_at(chain, x):
    """The value of a piecewise-linear chain at x."""
    for (x0, y0), (x1, y1) in zip(chain[:-1], chain[1:]):
        if x0 <= x <= x1:
            return y0 + (y1 - y0) * (x - x0) / (x1 - x0)

def area(lo, up):
    """Area between the two chains (shoelace on the closed polygon)."""
    poly = lo + up[-2:0:-1]
    n = len(poly)
    return abs(sum(poly[i][0] * poly[(i + 1) % n][1]
                   - poly[(i + 1) % n][0] * poly[i][1]
                   for i in range(n))) / 2

# the four-breakpoint example: the LP relaxation at x = 1
B = [(0, 0), (1, 2), (2, 1), (3, 4)]
lo, up = hull(B)
print(f"breakpoints {B}: at x = 1 the LP relaxation allows f in "
      f"[{chain_at(lo, 1):.2f}, {chain_at(up, 1):.2f}], "
      f"the interpolant gives 2.00")

def model(f, a, b, k, m=200):
    """The figure's view of f on [a, b] with k breakpoints.

    The interpolant through k equally spaced breakpoints, its error on
    a grid of m + 1 points per piece, the hull area, and the binaries.
    """
    xs = np.linspace(a, b, k)
    ys = f(xs)
    h = (b - a) / (k - 1)
    err = 0.0
    at = None
    for i in range(k - 1):
        t = np.linspace(0, 1, m + 1)
        x = xs[i] + t * h
        d = np.abs(f(x) - (ys[i] + t * (ys[i + 1] - ys[i])))
        if d.max() > err + 1e-12:
            err, at = d.max(), x[d.argmax()]
    lo, up = hull(np.column_stack([xs, ys]))
    logb = 0
    while 2**logb < k - 1:
        logb += 1
    return err, at, area(lo, up), k - 1, k - 2, logb

# the figure's curves
curves = [('x^2 on [0, 2]', np.square, 0, 2),
          ('sin x on [0, 2 pi]', np.sin, 0, 2 * np.pi),
          ('x^3 - 2x on [-1.6, 1.6]', lambda x: x**3 - 2 * x, -1.6, 1.6)]
for name, f, a, b in curves:
    e5, at, A, cc, inc, lg = model(f, a, b, 5)
    print(f"{name:26s} k = 5: largest error {e5:.4f} at x = {at:.2f}; "
          f"hull area {A:.3f}; binaries {cc} (convex combination), "
          f"{inc} (incremental), {lg} (logarithmic)")
    errs = [model(f, a, b, k)[0] for k in (5, 9, 17, 33)]
    print(f"{'':26s} doubling the pieces 4 -> 8 -> 16 -> 32: "
          f"errors {errs[0]:.4f}, {errs[1]:.4f}, {errs[2]:.4f}, "
          f"{errs[3]:.4f}; ratios {errs[0] / errs[1]:.1f}, "
          f"{errs[1] / errs[2]:.1f}, {errs[2] / errs[3]:.1f}")

# x^2 with step h has error h^2/4 at the midpoints; the bound L h^2 / 8
# with L = 2 is attained
for a, b, k in [(0, 1, 3), (0, 2, 3), (0, 2, 5)]:
    e, at, *_ = model(np.square, a, b, k)
    h = (b - a) / (k - 1)
    print(f"x^2 on [{a}, {b}] with {k - 1} pieces of width {h}: "
          f"error {e:.4f} = h^2/4 = {h * h / 4:.4f}")

The script prints the interval \([0.50, 2.00]\) for the four-breakpoint example. For the default view it prints the error \(0.0625\) at \(x = 0.25\), the hull area \(1.250\) and the binary counts \(4\), \(3\) and \(2\). As the pieces double it prints the ratios \(4.0, 4.0, 4.0\) for \(x^2\), \(3.0, 3.7, 3.9\) for the sine and \(3.4, 3.7, 3.9\) for the cubic. It also prints the areas \(6.283\) and \(4.915\), and the errors \(0.0625\), \(0.2500\) and \(0.0625\) of the two-piece and four-piece interpolants of \(x^2\). Its cost is \(O(K m)\) function evaluations per model. The error computation is independent across pieces, which is the part a batched implementation would spread over threads.

From approximation to relaxation: the error band

An interpolant is not a relaxation. Replacing \(y = f(x)\) by \(y = \phi(x)\) removes every point of the graph of \(f\) that is not on the graph of \(\phi\), which is all of them except the breakpoints. The resulting MILP can therefore declare infeasible a problem that is feasible, and its optimal value is neither a lower nor an upper bound. Geißler, Martin, Morsi and Schewe repair this with one inequality per piece. The proposition speaks of a factorable MINLP, one whose functions are built from a finite list of elementary univariate functions, sums and products, which is the class Section 2.4 relaxes term by term.

Proposition 4.6.6 (MIP relaxations with an error band; Geißler, Martin, Morsi and Schewe, 2012). Let \(y = f(x)\) be a constraint of a MINLP, with \(x\) in a bounded interval, and let \(\phi\) be a piecewise-linear interpolant of \(f\). For each piece \(v\) let \(\underline\delta_v \ge \max_{x \in \text{piece } v} \big(f(x) - \phi(x)\big)\) and \(\overline\delta_v \ge \max_{x \in \text{piece } v}\big(\phi(x) - f(x)\big)\) be certified bounds on how far the curve rises over and falls under the interpolant. Then the set \(\{(x, y) : \phi(x) - \overline\delta_{v(x)} \le y \le \phi(x) + \underline\delta_{v(x)}\}\), where \(v(x)\) is the piece containing \(x\), contains the graph of \(f\). Replacing every nonlinearity of a factorable MINLP by such a band gives a MILP whose feasible set contains that of the MINLP. The MILP is therefore a relaxation, its optimal value is a valid bound, and a MILP solver alone can solve it.

Proof. For \(x\) in piece \(v\), \(f(x) - \phi(x)\) lies in \([-\overline\delta_v, \underline\delta_v]\) by the choice of the two constants, so \((x, f(x))\) is in the band. A problem whose feasible set has grown has an optimal value that is no larger, for a minimization. ∎

(The band, and what the MILP then is) The band in the figure is this construction with the two constants computed separately for every piece. For \(x^2\) the curve never rises over its chords, so \(\underline\delta_v = 0\) and the band hangs one-sidedly below the interpolant. The certified constants come from the curvature bound of Proposition 4.6.5, from interval arithmetic on the sub-box, or from an oracle for the maximum of the nonlinearity on a piece. The original paper favours the incremental formulation and gives error estimates for univariate and for multivariate pieces. The MILP it produces is a genuine relaxation. A MILP solver, with all its machinery for binaries, therefore becomes a global solver for the MINLP at a fixed accuracy. Refining the pieces where the MILP's solution violates the nonlinearity makes the method convergent.

Algorithm 4.6.7 (adaptive piecewise-linear MIP relaxation; Geißler et al. 2012; Burlacu, Geißler and Schewe 2020).

Algorithm ADAPTIVE-PL-RELAXATION

Input   a factorable MINLP whose nonlinearities y_j = f_j(x_j) have
        bounded arguments; a tolerance eps; for each f_j an oracle
        giving a certified bound on max |f_j - phi| over any
        sub-interval.
Output  an eps-feasible point and a valid lower bound on the MINLP's
        optimal value (minimization).

1. For every nonlinearity choose initial breakpoints and build the
   interpolant phi_j with a logarithmic or an incremental model; add
   the band
      phi_j(x_j) - delta_bar_j <= y_j <= phi_j(x_j) + delta_under_j
   per piece, with the two constants from the oracle
   (Proposition 4.6.6).

2. Solve the MILP relaxation. Its value LB is a valid lower bound;
   let (x*, y*) be its solution.

3. If max_j |y*_j - f_j(x*_j)| <= eps: stop. (x*, y*) is eps-feasible
   and LB is the global bound.

4. Otherwise insert new breakpoints at or near x*_j in every violated
   nonlinearity, recompute the constants of the pieces that changed
   (they shrink), and return to step 2.

Invariant
    every MILP solved is a relaxation of the MINLP
    (Proposition 4.6.6), so LB never exceeds the optimum; the bands
    shrink only where the current solution needs them to; for
    continuous f_j on bounded domains the sequence of violations
    tends to zero (Burlacu, Geißler and Schewe 2020).

Cost
    one MILP per round, growing by O(log K) binaries per refinement
    in the logarithmic model; the constants cost one oracle call per
    changed piece.

Parallel
    the MILP is the sequential part; the breakpoint insertion and the
    recomputation of the constants are independent across
    nonlinearities.
The loop of Algorithm 4.6.7 (ADAPTIVE-PL-RELAXATION)

  1. breakpoints, interpolants phi_j and a band per piece,
     with the constants from the oracle
        |
        v
  2. solve the MILP relaxation  <---------------------------+
     LB is a valid lower bound; (x*, y*) its solution        |
        |                                                    |
        v                                                    |
  3. max_j |y*_j - f_j(x*_j)| <= eps ? --yes--> stop:        |
        | no                         (x*, y*) eps-feasible,  |
        v                            LB the global bound     |
  4. new breakpoints at or near x*_j in every violated       |
     nonlinearity; the constants of the changed pieces       |
     recomputed (they shrink) -------------------------------+

  every MILP is a relaxation of the MINLP, so LB never exceeds
  the optimum; each round's MILP grows by O(log K) binaries per
  refinement in the logarithmic model

The method is demonstrated on gas network instances in the original papers, where the nonlinearities are the pressure-loss equations of pipes.Geißler, Martin, Morsi and Schewe (2012), cited above; Burlacu, Geißler and Schewe (2020), cited above, who prove convergence for continuous nonlinearities on bounded domains given the oracle for the maximum of each nonlinearity on a sub-box. Its cost per round is a MILP, which grows at every refinement. That is the trade it makes against spatial branch and bound. The branching over the pieces is done by the MILP solver's own tree, with the MILP solver's cuts and heuristics, instead of by a tree over boxes with LP relaxations at the nodes.

Piecewise McCormick

The same construction applied to a product \(w = xy\) is the piecewise McCormick relaxation. Partition the range of \(x\) into \(K\) pieces \([a_{k-1}, a_k]\) and introduce a binary \(z_k\) per piece with \(\sum_k z_k = 1\). Introduce copies \(x_k, y_k, w_k\) of the three variables with \(x = \sum_k x_k\), \(y = \sum_k y_k\) and \(w = \sum_k w_k\), and the bounds \(a_{k-1} z_k \le x_k \le a_k z_k\) and \(y^L z_k \le y_k \le y^U z_k\). On each piece add the four McCormick planes of Theorem 2.4.7 for \((x_k, y_k, w_k)\) with every bound multiplied by \(z_k\). This is the hull formulation (H) of Definition 4.2.3, whose projection Theorem 4.2.4 describes, applied to the disjunction over pieces. Section 2.4 built the same formulation for the bilinear example. What the guarantee says, and what it does not say, needs care.

Piecewise McCormick: the range of x cut into K pieces

  y^U +---------+---------+----   ----+---------+
      |         |         |           |         |
      | piece 1 | piece 2 |   ...     | piece K |
      |   z_1   |   z_2   |           |   z_K   |
      |   x_1,  |   x_2,  |           |   x_K,  |
      |   y_1,  |   y_2,  |           |   y_K,  |
      |   w_1   |   w_2   |           |   w_K   |
  y^L +---------+---------+----   ----+---------+---> x
     a_0       a_1       a_2       a_{K-1}     a_K

  sum_k z_k = 1;  x = sum_k x_k,  y = sum_k y_k,  w = sum_k w_k
  a_{k-1} z_k <= x_k <= a_k z_k   and   y^L z_k <= y_k <= y^U z_k
  on each piece: the four McCormick planes for (x_k, y_k, w_k),
  every bound multiplied by z_k

  z_k = 1: the copies of the other pieces are forced to zero, and
  what remains is the McCormick relaxation of w = xy on
  [a_{k-1}, a_k] x [y^L, y^U]

Proposition 4.6.8 (piecewise McCormick: the hull formulation and its price). (i) For integral \(z\) the formulation above is exactly the McCormick relaxation of \(w = xy\) on the selected piece. (ii) Propositions 2.4.8 and 2.4.9 give the band and the LP bound. The band on a cell falls like \(1/(2 k_x k_y)\) of the box's area, and with the coupling rows kept outside the disjunction, as solvers write them, the LP relaxation of the partition formulation is the McCormick bound of the whole box for every partition. The best of the per-piece optima is the MILP optimum of the partition formulation, reached by branching on the selectors or at once by copying the coupling rows into every disjunct. (iii) The number of cells is \(k^d\) when \(d\) variables are partitioned, and the MILP's tree over the selectors is exponential in their number in the worst case.The formulation is the disjunctive one of M. L. Bergamini, P. Aguirre and I. Grossmann, "Logic-based outer approximation for globally optimal synthesis of process networks", Computers & Chemical Engineering 29 (2005), 1914–1933, and D. S. Wicaksono and I. A. Karimi, "Piecewise MILP under- and overestimators for global optimization of bilinear programs", AIChE Journal 54 (2008), 991–1008; the formulations are compared on pooling problems in C. E. Gounaris, R. Misener and C. A. Floudas, "Computational comparison of piecewise-linear relaxations for pooling problems", Industrial & Engineering Chemistry Research 48 (2009), 5742–5766, and tightened in P. M. Castro, "Tightening piecewise McCormick relaxations for bilinear problems", Computers & Chemical Engineering 72 (2015), 300–311.

Proof. (i) With \(z_k = 1\) and the other selectors zero, every copy with index \(j \ne k\) is forced to zero by its scaled bounds, and the four planes for piece \(k\) are the McCormick planes on \([a_{k-1}, a_k] \times [y^L, y^U]\). (ii) The band is Proposition 2.4.8, and the two statements about the relaxation are Proposition 2.4.9, whose proof applies Theorem 4.2.4 to the union of the per-piece sets. (iii) is a count. ∎

(The bilinear example R2 under a partition) Return to the bilinear example of Section 2.4: maximize \(xy\) on the unit square under \(2x + y \le 1.2\), with true value \(0.18\). The ladder of per-cell bests is the one computed in Section 2.4, whose table gives the \(k \times k\) values from \(0.4000\) at \(k = 1\) to \(0.1807\) at \(k = 16\) and the one-sided values from \(0.30\) at \(k = 2\) to \(0.1944\) at \(k = 16\). Those are the MILP optima of the partition formulation, the values the tree over the selectors, the binaries \(z_k\) that pick a piece, reaches. The LP relaxation of the same formulation, with \(2x + y \le 1.2\) outside the disjunction, is \(0.40\) for every partition, by Proposition 2.4.9. The reason is worth a sentence: by Theorem 4.2.4 the relaxation's projection onto \((x, y, w)\) is the convex hull of the union of the per-piece McCormick sets, each of which is the hull of the graph of \(xy\) over its piece, so the union's hull is the hull of the whole graph, which is the box's own McCormick set, and a coupling row applied outside the disjunction sees nothing the whole-box relaxation did not. The band falls like \(1/k^2\). The bound does not, because the bound depends on where the optimum lies relative to the grid, and partitioning \(x\) alone, which is what most solvers do, is roughly first order.

(Fewer binaries) Three devices reduce the binary count. Logarithmic encodings of the selectors, as in Theorem 4.6.4, bring \(K\) binaries per term down to \(\lceil \log_2 K \rceil\). APOGEE and GloMIQO use them for pooling and for general mixed-integer quadratically constrained quadratic programs (MIQCQP).R. Misener, J. P. Thompson and C. A. Floudas, "APOGEE: global optimization of standard, generalized, and extended pooling problems via linear and logarithmic partitioning schemes", Computers & Chemical Engineering 35 (2011), 876–892; R. Misener and C. A. Floudas, "Global optimization of mixed-integer quadratically-constrained quadratic programs (MIQCQP) through piecewise-linear and edge-concave relaxations", Mathematical Programming 136 (2012), 155–182. The normalized multiparametric disaggregation technique writes one factor in base ten, \(x = \sum_{l} \sum_{q=0}^{9} 10^l q\, z_{l,q} + r\), where exactly one of the ten digit indicators \(z_{l,0}, \dots, z_{l,9}\) equals one at each position \(l\) and \(r\) is a residual. It linearizes each product \(y z_{l,q}\) exactly and relaxes only \(y r\) by McCormick, so that \(p\) decimal digits cost \(10p\) binaries for \(10^p\) pieces.S. Kolodziej, P. M. Castro and I. E. Grossmann, "Global optimization of bilinear programs with a multiparametric disaggregation technique", Journal of Global Optimization 57 (2013), 1039–1063; P. M. Castro, "Normalized multiparametric disaggregation: an efficient relaxation for mixed-integer bilinear problems", Journal of Global Optimization 64 (2016), 765–784. Adaptive partitioning, as in Alpine, places the pieces around the incumbent and tightens bounds between rounds instead of refining uniformly.H. Nagarajan, M. Lu, S. Wang, R. Bent and K. Sundar, "An adaptive, multivariate partitioning algorithm for global optimization of nonconvex programs", Journal of Global Optimization 74 (2019), 639–675.

(Partition against spatial tree) A partition and a spatial branch-and-bound tree do the same work in a different order. A \(k^d\)-cell partition evaluated at once is the frontier that a spatial tree reaches after \(d \log_2 k\) levels of bisection, and the per-cell LP of the partition is the node relaxation of the tree. The partition pays for every cell up front and hands the selection to a MILP solver's tree. The spatial tree pays only for the cells it visits and prunes the rest with the incumbent. Which is cheaper depends on how much of the frontier the incumbent would have pruned. That is also the question a GPU asks when it bounds a frontier in batches, in Section 6.4.

A partition and a spatial tree: the same cells, another order

  spatial tree: bisect level by level, prune with the incumbent

                      [          box          ]
                     /                         \
             [   half   ]                 [   half   ]
              /        \                   /        \
          [    ]     [    ]            [    ]     [    ]
           /  \       /  \              /  \       pruned by
          :    :     :    :            :    :      the incumbent
  after d log2 k levels, the frontier of k^d cells:
          [] [] ... [] [] ...          [] [] ...  -- --  never
                                                  paid for

  partition: all k^d cells at once, each with its LP, the node
  relaxation of the tree; a MILP solver's tree selects the cell

Where this is used

SOS2 constraints are enforced by branching in SCIP, as noted above, and CPLEX, Xpress and Gurobi accept them as a constraint type. Gurobi by default reformulates SOS constraints into linear constraints with binaries, and two parameters control the largest constant it will introduce and which encoding it uses. Xpress converts a piecewise-linear constraint \(y = f(x)\) given by breakpoints into linear constraints and an SOS2 set.Gurobi Optimizer Reference Manual, "Constraints" and "Parameters" (PreSOS2BigM for the largest constant and PreSOS2Encoding for the encoding; FuncPieces and FuncPieceError, default \(10^{-3}\), for the static approximation of function constraints), docs.gurobi.com, read 4 October 2026; FICO Xpress Optimizer Reference Manual, XPRSaddpwlcons. For smooth nonlinearities the vendors have moved from static pieces to dynamic ones. Gurobi handled its function constraints \(y = f(x)\) by a static piecewise-linear approximation through version 10. Version 11.0, released in November 2023, made a dynamic outer approximation inside the branch-and-bound tree available through the new parameter FuncNonlinear, with the static approximation still the default. Version 12.0, released in November 2024, made the dynamic approach the default. Version 13.0, released in November 2025, deprecates the function constraints in favour of general nonlinear constraints handled by spatial branch and bound.Gurobi Optimizer Reference Manual, "Additions, changes and removals in Gurobi 11.0", docs.gurobi.com/projects/optimizer/en/11.0/reference/releasenotes/changes.html: "The Gurobi Optimizer can now use spatial branching and outer approximation to solve models with non-linear functions, instead of using static piecewise-linear approximations", with FuncNonlinear listed as a new parameter "to control whether general function constraints shall be treated as nonlinear functions or via piecewise-linear approximation"; "Additions, changes and removals in Gurobi 12.0", docs.gurobi.com/projects/optimizer/en/12.0/reference/releasenotes/changes.html: "The new default (value 1) will use a dynamic outer-approximation approach for all univariate general functions constraints"; "Additions, changes and removals in Gurobi 13.0", docs.gurobi.com/projects/optimizer/en/current/reference/releasenotes/changes.html, which deprecates function constraints; all three read 5 October 2026. The release months are from Gurobi Optimization, "Gurobi release and support history", support.gurobi.com article 360048138771 (read 5 October 2026). Static pieces survive where the function is piecewise linear by nature, in Geißler-style MIP relaxations and in the piecewise relaxations of bilinear terms inside ANTIGONE and Alpine.

What parallelizes

What parallelizes here is the partition itself. The per-piece error constants, the per-cell McCormick bounds and the per-cell LPs are independent, which makes a \(k^d\) partition a batch. The MILP that selects among the cells is the sequential part, exactly as the tree is the sequential part of spatial branch and bound. Section 7.4 returns to the batched frontier.

Lifting: RLT, SDP and the hierarchy

Every relaxation of a product so far has treated the product as a function of its two factors and bounded it on a box. This subsection changes the point of view: it treats the product as a new variable. A quadratic program in \(x \in \mathbb{R}^n\) becomes a linear program in the pair \((x, X)\) once every monomial \(x_i x_j\) is replaced by a variable \(X_{ij}\). The whole nonconvexity is then concentrated in the single equation \(X = x x^\top\). Convex outer approximations of that one set are the subject of this subsection. They come in three families. The linear ones multiply constraints together before linearizing (the reformulation–linearization technique, RLT). The semidefinite ones ask that \(X - x x^\top\) be positive semidefinite (Shor's relaxation). The third family is a hierarchy of semidefinite relaxations over moment matrices that converges to the optimum (Lasserre's hierarchy). At the end of the hierarchy stands an exact reformulation over the completely positive cone (Burer), which moves the difficulty into a cone whose membership problem is NP-hard.

(Why the quadratic case matters for factorable MINLP) The quadratic case deserves this much space in a monograph on general nonconvex MINLP for one reason. A factorable relaxation introduces one auxiliary variable per intermediate operation, so every product of intermediates is a bilinear equation \(w = v_i v_j\). The McCormick planes of Theorem 2.4.7 are the weakest member of the RLT family. Everything in this subsection says how much a factorable relaxation leaves on the table and how to recover it.

Products as variables

Definition 4.7.1 (lifting; the Schur complement). For \(x \in \mathbb{R}^n\) write \(X = x x^\top \in \mathbb{S}^n\), the symmetric matrix with entries \(X_{ij} = x_i x_j\), and \(\langle Q, X \rangle = \sum_{ij} Q_{ij} X_{ij}\), so that \(x^\top Q x = \langle Q, X \rangle\). The lifted feasible set of a quadratically constrained quadratic program (QCQP) replaces every quadratic form by its linear image in \(X\) and adds the equation \(X = x x^\top\). The set

\[\mathcal S = \{ (x, X) \in \mathbb{R}^n \times \mathbb{S}^n : X = x x^\top \}\]

carries the whole nonconvexity, and everything else is linear. The Schur complement lemma gives its basic convex relaxation:

\[M(x, X) = \begin{pmatrix} 1 & x^\top \\ x & X \end{pmatrix} \succeq 0 \quad\iff\quad X - x x^\top \succeq 0,\]

and \(X = x x^\top\) holds if and only if, in addition, \(M(x, X)\) has rank one.

The lifted problem is a linear program over the convex hull of \(\mathcal S\) intersected with the constraints. That hull is not computable in general. Each relaxation below is a convex set containing it.

Lifting a QCQP (Definition 4.7.1) and the convex sets that replace S

  a QCQP in x in R^n          x^T Q_i x + c_i^T x + d_i <= 0
          |
          |  every monomial x_i x_j becomes a variable X_ij,
          |  so that x^T Q x = <Q, X>
          v
  linear in (x, X)            <Q_i, X> + c_i^T x + d_i <= 0
  plus one equation           X = x x^T, the set S, which carries
                              the whole nonconvexity
          |
          |  replace conv(S intersected with the constraints)
          |  by a convex set containing it
          v
  RLT        products of factors, linearized: linear rows
  Shor       M(x, X) PSD, i.e. X - x x^T PSD: one semidefinite cone
  Lasserre   moment matrices M_d(y) PSD: semidefinite programs that
             converge to the optimum as d grows
  Burer      M(x, X) completely positive: exact for the linearly
             constrained QPs of Theorem 4.7.13 (x >= 0, binaries),
             in a cone whose membership problem is NP-hard

The reformulation–linearization technique

Definition 4.7.2 (RLT factors and levels; Sherali and Adams). Given bounds \(l \le x \le u\) and linear constraints \(a_r^\top x \le b_r\), the bound factors, already met in the proof of Theorem 2.4.7, are the \(2n\) affine functions \(x_j - l_j \ge 0\) and \(u_j - x_j \ge 0\), and the constraint factors are \(b_r - a_r^\top x \ge 0\). The reformulation step multiplies factors pairwise (for level \(1\)) or in products of up to \(t\) factors (for level \(t\)). Every such product is a polynomial inequality valid on the feasible set. The linearization step replaces each monomial \(x_i x_j\) by \(X_{ij}\), and higher monomials by new variables. The level-\(t\) RLT relaxation is the resulting linear program. A linear equality \(a_r^\top x = b_r\) is multiplied by each variable, giving the equations \(\sum_j a_{rj} X_{jk} = b_r x_k\).H. D. Sherali and W. P. Adams, "A hierarchy of relaxations between the continuous and convex hull representations for zero-one programming problems", SIAM Journal on Discrete Mathematics 3 (1990), 411–430; H. D. Sherali and A. Alameddine, "A new reformulation-linearization technique for bilinear programming problems", Journal of Global Optimization 2 (1992), 379–410; the book is H. D. Sherali and W. P. Adams, A Reformulation-Linearization Technique for Solving Discrete and Continuous Nonconvex Problems (Kluwer, 1999). Products of equalities with variables as "reduction constraints" are L. Liberti, "Reduction constraints for the global optimization of NLPs", International Transactions in Operational Research 11 (2004), 33–41, and L. Liberti and C. C. Pantelides, "An exact reformulation algorithm for large nonconvex NLPs involving bilinear terms", Journal of Global Optimization 36 (2006), 161–189.

The first thing to see is that the McCormick planes are already RLT rows.

Proposition 4.7.3 (bound-factor products are the McCormick planes; constraint-factor products are new). On the box \([l_x, u_x] \times [l_y, u_y]\) the four linearized products of an \(x\)-factor with a \(y\)-factor,

\[(x - l_x)(y - l_y) \ge 0, \quad (u_x - x)(u_y - y) \ge 0, \quad (x - l_x)(u_y - y) \ge 0, \quad (u_x - x)(y - l_y) \ge 0,\]

with \(xy\) replaced by \(X_{12}\), are exactly the four McCormick inequalities. By the Al-Khayyal–Falk theorem (Theorem 2.4.7) they describe \(\operatorname{conv}\{(x, y, xy)\}\) over the box. The products of a factor of \(x_j\) with itself, \((x_j - l_j)^2 \ge 0\), \((u_j - x_j)^2 \ge 0\) and \((x_j - l_j)(u_j - x_j) \ge 0\), bound the diagonal entry: \(X_{jj} \ge 2 l_j x_j - l_j^2\), \(X_{jj} \ge 2 u_j x_j - u_j^2\) and \(X_{jj} \le (l_j + u_j) x_j - l_j u_j\). The products of a constraint factor with a bound factor are in general not implied by any of these.

Proof. Expanding \((x - l_x)(y - l_y) \ge 0\) gives \(X_{12} \ge l_y x + l_x y - l_x l_y\), which is the first McCormick plane, and the other three are the same computation. The squared factors expand to the three diagonal rows. For the last claim an example suffices, and the running bilinear example below is one: the McCormick rows alone give the bound \(0.40\), and the constraint-factor products lower it to \(0.30\). ∎

(The knapsack: three RLT levels by hand) A 0–1 knapsack shows the mechanism in two lines. Maximize \(x_1 + x_2\) subject to \(2x_1 + 2x_2 \le 3\) and \(x \in \{0, 1\}^2\). The LP relaxation gives \(1.5\). At level \(1\) introduce \(X_{12} = x_1 x_2\) and use \(x_j^2 = x_j\), which holds for binaries. Multiplying the constraint factor \(3 - 2x_1 - 2x_2 \ge 0\) by \(x_1\) gives \(3x_1 - 2x_1 - 2X_{12} \ge 0\), that is \(X_{12} \le x_1 / 2\), and likewise \(X_{12} \le x_2 / 2\). The bound-factor product \((1 - x_1)(1 - x_2) \ge 0\) gives \(X_{12} \ge x_1 + x_2 - 1\). Together, \(x_1 + x_2 - 1 \le x_1/2\) and \(x_1 + x_2 - 1 \le x_2/2\) add to \(\tfrac32 (x_1 + x_2) \le 2\), so the level-\(1\) bound is \(4/3\), at \(x_1 = x_2 = 2/3\) and \(X_{12} = 1/3\). At level \(2\) multiply the constraint by \(x_1 x_2\): \(3X_{12} - 2x_1^2 x_2 - 2 x_1 x_2^2 = 3X_{12} - 2X_{12} - 2X_{12} = -X_{12} \ge 0\), so \(X_{12} = 0\) and \(x_1 + x_2 \le 1\), which is the integer hull. The ladder \(1.5 \to 1.333 \to 1\) is the Sherali–Adams hierarchy on two variables.

The RLT levels on the 0–1 knapsack max x1 + x2, 2x1 + 2x2 ≤ 3: the LP relaxation (bound 1.5), level 1 (bound 4/3 at (2/3, 2/3)) and level 2 (bound 1, the integer hull). The blue points are the feasible 0–1 points; the red cross is (1, 1), cut off by 2x1 + 2x2 ≤ 3.

(The bilinear example R2: three linear programs) The running bilinear example gives the ladder that the next figure draws. Maximize \(xy\) on the unit square subject to \(2x + y \le 1.2\). The true value is \(0.18\) at \((0.3, 0.6)\), where the line is tangent to a level curve of the product. The lifted variables are \((x, y, X_{12}, X_{11}, X_{22})\) for \((x, y, xy, x^2, y^2)\), the entries of Definition 4.7.1, with \(X_{12}\) the product that Section 2.4 wrote as \(w\). The figure and the script below abbreviate them to \(W\), \(X\) and \(Y\). The factors are \(x\), \(1 - x\), \(y\), \(1 - y\) and \(1.2 - 2x - y\). Three linear programs, each maximizing \(X_{12}\):

The inequality \(X_{11} \ge x^2\) is the \(2 \times 2\) principal minor of the condition \(M(x, X) \succeq 0\), the nonnegativity of the determinant of the submatrix on the first two rows and columns, which is the diagonal of Shor's relaxation. One such fact, added to the products, closes the whole gap on this example.

The figure draws the three relaxations in the \((x, W)\) plane at a fixed \(y\), with \(W\) the figure's label for \(X_{12}\), which is the slice on which the three bounds can be compared. The ink curve is \(W = xy\) on the slice, solid where \(2x + y \le 1.2\) allows it. The orange band is the range of \(W\) the chosen relaxation permits at each \(x\), computed by two small linear programs per point, and the grey outlines are the other two. The side panel is the ladder of bounds, \(0.400\), \(0.300\) and \(0.180\), each a linear program solved in the page. The two drops are annotated: \(-0.100\) from the products, \(-0.120\) from \(X_{11} \ge x^2\). In the default slice at \(y = 0.40\) the McCormick band is the triangle between \(W = 0\) and \(W = x\), and the curve \(W = 0.4x\) reaches \(0.160\) at \(x = 0.40\). The other two bands reach \(0.240\) and \(0.160\) in that slice. The readout lists every row of the chosen relaxation with its product form, its linearized form and whether it binds. It reports the row counts \(7\), \(18\) and \(36\) and the pivot counts \(3\), \(4\) and \(12\) of the three solves. The gaps are \(122\) per cent, \(67\) per cent and zero.

The lifting ladder on the bilinear example, maximize x·y on the unit square under 2x + y ≤ 1.2. The main panel is the (x, W) plane in one slice at fixed y, with W the lifted product X12 = x·y. The ink curve is W = x·y, the orange band is the W the chosen relaxation permits, the grey outlines are the other two, and the blue dashed line is the true maximum 0.18. The slider moves the slice and the readout can list every row.

The script below builds the three linear programs from the factors and solves them with a dense tableau simplex. Every row has a nonnegative right-hand side, because each is a product of factors that are nonnegative at the origin, so the slack basis is feasible and no first phase is needed.

# rlt_ladder.py -- the lifting ladder on the bilinear example.
#
# Maximize x*y s.t. 2x + y <= 1.2 on [0,1]^2. Lifted variables
# v = (x, y, W, X, Y) for (x, y, xy, x^2, y^2). Three LPs, solved by a
# dense tableau simplex.

import numpy as np

def simplex_max(c, A, b, tol=1e-9):
    """Maximize c.v  s.t.  A v <= b,  v >= 0; Bland's rule.

    b >= 0, so the slack basis is feasible.
    """
    m, n = A.shape
    T = np.zeros((m + 1, n + m + 1))
    T[:m, :n] = A
    T[:m, n:n+m] = np.eye(m)
    T[:m, -1] = b
    T[m, :n] = -c
    basis = list(range(n, n + m))
    for pivots in range(1000):
        cols = np.where(T[m, :-1] < -tol)[0]
        if len(cols) == 0:
            v = np.zeros(n + m)
            for i, j in enumerate(basis):
                v[j] = T[i, -1]
            return T[m, -1], v[:n], pivots

        # lowest index with a negative reduced cost
        j = cols[0]
        rows = np.where(T[:m, j] > tol)[0]
        ratios = T[rows, -1] / T[rows, j]
        cand = rows[ratios <= ratios.min() + tol]
        # ties: lowest basic index
        i = min(cand, key=lambda r: basis[r])
        T[i] /= T[i, j]
        for r in range(m + 1):
            if r != i:
                T[r] -= T[r, j] * T[i]
        basis[i] = j
    raise RuntimeError("no convergence")

# affine factors p*x + q*y + r >= 0 on the feasible set
F = {'x': (1, 0, 0), '1-x': (-1, 0, 1), 'y': (0, 1, 0), '1-y': (0, -1, 1),
     '1.2-2x-y': (-2, -1, 1.2)}

def product(f, g):
    """Linearize (p1 x + q1 y + r1)(p2 x + q2 y + r2) >= 0
    as a.v + a0 >= 0."""
    p1, q1, r1 = f
    p2, q2, r2 = g
    a = np.array([p1*r2 + r1*p2, q1*r2 + r1*q2, p1*q2 + q1*p2,
                  p1*p2, q1*q2])
    return a, r1*r2

def tangent(var, t):
    """X >= x^2 outer-approximated by its tangent at t:
    X - 2 t x + t^2 >= 0."""
    a = np.zeros(5)
    a[3 if var == 'x' else 4] = 1
    a[0 if var == 'x' else 1] = -2*t
    return a, t*t

base = [(np.array([-1, 0, 0, 0, 0.]), 1.0),
        (np.array([0, -1, 0, 0, 0.]), 1.0),
        (np.array([-2, -1, 0, 0, 0.]), 1.2)]
mcc = [product(F[a], F[b]) for a in ('x', '1-x') for b in ('y', '1-y')]
names = list(F)
rlt = [product(F[names[i]], F[names[j]])
       for i in range(5) for j in range(i, 5)]
tang = [tangent(v, t/10) for v in ('x', 'y') for t in range(1, 10)]

c = np.array([0, 0, 1, 0, 0.])                  # maximize W
true = np.array([0.3, 0.6, 0.18, 0.09, 0.36])
ladder = [('McCormick', base + mcc),
          ('level-1 RLT', base + rlt),
          ('RLT + tangents of X >= x^2, Y >= y^2', base + rlt + tang)]
for name, rows in ladder:
    # a.v + a0 >= 0  <=>  (-a).v <= a0
    A = np.array([-a for a, a0 in rows])
    b = np.array([a0 for a, a0 in rows])
    z, v, piv = simplex_max(c, A, b)
    ok = np.all(A @ true <= b + 1e-12)
    print(f"{name:38s} rows {len(rows):2d}  bound {z:.4f}")
    print(f"    at (x, y, W, X, Y) = ({v[0]:.3f}, {v[1]:.3f}, "
          f"{v[2]:.3f}, {v[3]:.3f}, {v[4]:.3f})")
    print(f"    pivots {piv:2d}  true point feasible: {ok}")
print("true maximum 0.18 at (0.3, 0.6): x(1.2 - 2x) is largest at x = 0.3")

It prints the bounds \(0.4000\), \(0.3000\) and \(0.1800\) at the points \((0.400, 0.400, 0.400, 0, 0)\), \((0.300, 0.500, 0.300, 0, 0)\) and \((0.250, 0.550, 0.180, 0.060, 0.300)\), with \(7\), \(18\) and \(36\) rows and \(3\), \(4\) and \(12\) pivots. These are the numbers the figure displays. It also confirms that the true point satisfies every row of every relaxation. The cost is dominated by the pivots, each \(O(mn)\) on the dense tableau. The rows themselves are independent outer products, which is the part of RLT that parallelizes.

The hierarchy theorem says why level \(n\) reaches the hull for binary variables.

Theorem 4.7.4 (Sherali and Adams, 1990, 1994; the RLT hierarchy). For a mixed 0–1 linear program with \(n\) binary variables, let \(X_t\) be the projection onto the original variables of the level-\(t\) RLT relaxation. That relaxation multiplies every constraint by every product \(\prod_{j \in J_1} x_j \prod_{j \in J_2} (1 - x_j)\) with \(|J_1| + |J_2| = t\) and linearizes with \(x_j^2 = x_j\). Then

\[\operatorname{conv}(\mathcal F) = X_n \subseteq X_{n-1} \subseteq \dots \subseteq X_1 \subseteq X_0 = \text{the LP relaxation},\]

and the inclusions can be strict.Sherali and Adams (1990), cited above, for pure 0–1 programs; H. D. Sherali and W. P. Adams, "A hierarchy of relaxations and convex hull characterizations for mixed-integer zero-one programming problems", Discrete Applied Mathematics 52 (1994), 83–106, for the mixed case. The comparison with the Lovász–Schrijver and Lasserre hierarchies is M. Laurent, "A comparison of the Sherali–Adams, Lovász–Schrijver, and Lasserre relaxations for 0–1 programming", Mathematics of Operations Research 28 (2003), 470–496, which shows that the Lasserre relaxation is the strongest of the three at every level; L. Lovász and A. Schrijver, "Cones of matrices and set-functions and 0–1 optimization", SIAM Journal on Optimization 1 (1991), 166–190.

Proof sketch of \(X_n = \operatorname{conv}(\mathcal F)\). At level \(n\) there is one linearization variable \(w_J\) for each subset \(J\) of the binaries. For each partition \((J_1, J_2)\) of \(\{1, \dots, n\}\) let \(\lambda_{(J_1, J_2)} = L\big(\prod_{J_1} x_j \prod_{J_2} (1 - x_j)\big)\) be the linearization of the corresponding product. These quantities are linearizations of nonnegative products, so the RLT rows make them nonnegative. They sum to one because the products sum to \(\prod_j (x_j + 1 - x_j) = 1\). Expanding each product and collecting terms expresses every \(x_i\) as \(\sum_{(J_1, J_2) : i \in J_1} \lambda_{(J_1, J_2)}\), so \(x = \sum \lambda_{(J_1, J_2)} \chi_{J_1}\) is a convex combination of the 0–1 vectors \(\chi_{J_1}\). The multiplied constraints force \(\lambda_{(J_1, J_2)} = 0\) whenever \(\chi_{J_1}\) violates a constraint. So every point of \(X_n\) is a convex combination of feasible 0–1 points. ∎

(The hierarchy on two variables) For \(n = 2\) the four weights are \(\lambda_{12} = L(x_1 x_2)\), \(\lambda_1 = L(x_1 (1 - x_2))\), \(\lambda_2 = L((1 - x_1) x_2)\) and \(\lambda_\emptyset = L((1 - x_1)(1 - x_2))\). They are nonnegative and sum to one, and linearity of \(L\) gives \(x_1 = \lambda_{12} + \lambda_1\) and \(x_2 = \lambda_{12} + \lambda_2\). So \((x_1, x_2)\) is the convex combination of \((1,1)\), \((1,0)\), \((0,1)\) and \((0,0)\) with those four weights. In the knapsack example the level-\(2\) row \(-X_{12} \ge 0\) sets \(\lambda_{12} = X_{12} = 0\), which removes the point \((1, 1)\) from the combination.

(Case analysis done linearly, and its price) The intuition is that multiplying by \(x_j\) and by \(1 - x_j\) is a case analysis on the value of \(x_j\) carried out linearly. Level \(n\) performs the full case analysis. Its price is \(2^n\) variables, which is why solvers stop at level \(1\) or \(2\) and leave the rest to branching. For continuous variables on a box the RLT bound converges under branching at the same rate as the McCormick bound. On a cell of width \(h\) every \(X_{jk}\) is confined to an interval of width \(O(h^2)\) around \(x_j x_k\) by the bound-factor products. RLT improves the constant, not the order.H. D. Sherali and C. H. Tuncbilek, "A global optimization algorithm for polynomial programming problems using a reformulation-linearization technique", Journal of Global Optimization 2 (1992), 101–112, prove convergence of the RLT-based branch and bound for polynomial programs on a box.

Algorithm 4.7.5 (RLT, level 1).

Algorithm RLT-1
          (level-1 reformulation-linearization of a QCQP,
           minimization form)

Input   QCQP data (Q_i, c_i, d_i), i = 0..m; bounds l <= x <= u; the
        sets L of linear inequalities, written b_r - a_r^T x >= 0, and
        E of linear equalities a_r^T x = b_r.
Output  an LP in (x, X), X symmetric n x n, with value z_RLT <= z*.

 1. Lift: replace every x^T Q_i x by <Q_i, X>; keep the linear parts.

 2. Bound-factor products: for every pair j <= k and every choice of
       F_j in {x_j - l_j, u_j - x_j}  and
       F_k in {x_k - l_k, u_k - x_k},
    add the linearized inequality F_j F_k >= 0.
                        [For j != k these are the McCormick planes for X_jk;
                                        for j = k the three bounds on X_jj.]

 3. Constraint-bound products: for every r in L and every bound factor
    F_j, add (b_r - a_r^T x) F_j >= 0, linearized.

 4. Constraint-constraint products: for every r <= s in L, add
    (b_r - a_r^T x)(b_s - a_s^T x) >= 0, linearized.

 5. Equality products: for every r in E and every k, add
       sum_j a_rj X_jk = b_r x_k.

 6. Solve the LP. If the solution satisfies X = x x^T to tolerance, the
    bound is exact at this node.

Invariant
    every row is the linearization of a product of nonnegative affine
    factors, or of an equality with a variable, so (x, x x^T) satisfies
    it for every feasible x; hence z_RLT <= z*.

Cost
    n(2n+1) + 2n|L| + |L|(|L|+1)/2 + n|E| rows (four per pair j < k,
    three per j) with up to O(n^2) nonzeros each; building them is
    O(n^2 (n + |L|)^2); the LP has O(n^2) columns.

Parallel
    each row is an outer product a_r a_s^T, independent of every other
    row; the LP is the sequential part, unless a first-order LP method
    is used, whose matrix-vector products can be formed from the
    factors without storing the RLT matrix.
Algorithm RLT-1 on the running example: 5 factors, 15 products

              x        1-x      y        1-y      1.2-2x-y
  x           X11      X11      M        M        c
  1-x                  X11      M        M        c
  y                             X22      X22      c
  1-y                                    X22      c
  1.2-2x-y                                        cc

  M         an x-factor times a y-factor: the four
            McCormick planes for X12                       step 2
  X11, X22  the three bounds on X11 and the three on X22   step 2
  c         the constraint factor times a bound factor     step 3
  cc        the constraint factor squared                  step 4

  McCormick:    3 base rows + the 4 M                 =  7 rows
  level-1 RLT:  3 base rows + all 15 products         = 18 rows
  with the 9 + 9 tangents of X11 >= x^2, X22 >= y^2   = 36 rows

  y (1.2-2x-y) >= 0   X12 <= 0.6y - X22/2, which at y = 0.5 and
                      X22 = 0 is the level-1 bound 0.300
  x (1.2-2x-y) >= 0   X12 <= 1.2x - 2 X11, which with X11 >= x^2
                      reads X12 <= 1.2x - 2x^2 <= 0.18

(Equality products, and size) Two remarks on what the algorithm does for problems with structure. Equality rows multiplied by variables, step 5, are equalities in the lifted space. They replace McCormick inequalities rather than add to them. They are the strongest cheap tightening the pooling problem of Section 3.5 admits. The products of the pool's proportion row \(\sum_i q_{il} = 1\) with the outflows are exactly the RLT rows that turn the q-formulation into the pq-formulation, and the bounds they give on Haverly's instances are in Section 3.5. The second remark concerns size. A level-\(1\) RLT for a problem with hundreds of variables has tens of thousands of rows, most of them never binding. Solvers therefore generate RLT rows as cuts at the LP vertex rather than writing them all down.

Shor's semidefinite relaxation

The linear rows of RLT know the box and the constraints, but no finite set of linear inequalities in \((x, X)\) can say that \(X\) is an outer product: finitely many inequalities describe a polyhedron, and the convex hull of the lifted set has the curved boundary \(X_{jj} \ge x_j^2\), which no polyhedron matches. The Schur complement of Definition 4.7.1 says something weaker that is convex: \(X \succeq x x^\top\).

Definition 4.7.6 (the Shor relaxation). For the QCQP \(\min\{ x^\top Q_0 x + c_0^\top x : x^\top Q_i x + c_i^\top x + d_i \le 0,\ i = 1, \dots, m,\ l \le x \le u \}\), with symmetric \(Q_i\) of any sign, the Shor relaxation is

\[z_{\mathrm{SDP}} = \min\{ \langle Q_0, X \rangle + c_0^\top x : \langle Q_i, X \rangle + c_i^\top x + d_i \le 0,\ i = 1, \dots, m,\ l \le x \le u,\ M(x, X) \succeq 0 \},\]

a semidefinite program over the \((n+1) \times (n+1)\) matrix \(M(x, X)\). Adding the level-\(1\) RLT rows gives the SDP+RLT relaxation.N. Z. Shor, "Quadratic optimization problems", Soviet Journal of Computer and Systems Sciences 25 (1987), 1–11, as cited by L. Vandenberghe and S. Boyd, "Semidefinite programming", SIAM Review 38 (1996), 49–95 (details as given there). T. Fujie and M. Kojima, "Semidefinite programming relaxation for nonconvex quadratic programs", Journal of Global Optimization 10 (1997), 367–380, give the relaxation its modern form.

The relaxation is valid because \((x, x x^\top)\) satisfies \(M \succeq 0\) with equality in the Schur complement. It is also the Lagrangian dual of the QCQP seen from the other side.

Proposition 4.7.7 (the Shor relaxation is the dual of the Lagrangian dual). Write \(Q(\lambda) = Q_0 + \sum_i \lambda_i Q_i\), \(c(\lambda) = c_0 + \sum_i \lambda_i c_i\) and \(d(\lambda) = \sum_i \lambda_i d_i\) for \(\lambda \ge 0\), and drop the box for the statement. Then

\[d^\star = \max_{\lambda \ge 0}\ \inf_x \big( x^\top Q(\lambda) x + c(\lambda)^\top x + d(\lambda) \big) = \max_{\lambda \ge 0,\ \gamma}\Big\{ \gamma : N(\lambda, \gamma) = \begin{pmatrix} d(\lambda) - \gamma & \tfrac12 c(\lambda)^\top \\ \tfrac12 c(\lambda) & Q(\lambda) \end{pmatrix} \succeq 0 \Big\} \le z_{\mathrm{SDP}} \le z^\star,\]

with \(d^\star = z_{\mathrm{SDP}}\) when either program is strictly feasible.

Proof. The infimum of a quadratic \(x^\top A x + b^\top x + e\) over \(\mathbb{R}^n\) is at least \(\gamma\) if and only if the form \((1, x)^\top \big(\begin{smallmatrix} e - \gamma & b^\top / 2 \\ b/2 & A \end{smallmatrix}\big) (1, x)\) is nonnegative for all \(x\). A quadratic form that is nonnegative on every vector with first coordinate one is nonnegative everywhere, by homogeneity and continuity. So the condition is the positive semidefiniteness of that matrix, which gives the second expression for \(d^\star\). For the inequality \(d^\star \le z_{\mathrm{SDP}}\) attach to the constraint \(N(\lambda, \gamma) \succeq 0\) a multiplier matrix \(M = \big(\begin{smallmatrix} m_0 & x^\top \\ x & X \end{smallmatrix}\big) \succeq 0\) and form the Lagrangian of the maximization,

\[\gamma + \langle M, N(\lambda, \gamma) \rangle = \gamma\,(1 - m_0) + \sum_{i=1}^m \lambda_i \big( m_0 d_i + c_i^\top x + \langle Q_i, X \rangle \big) + c_0^\top x + \langle Q_0, X \rangle .\]

The semidefinite cone is self-dual (Definition 4.8.1): \(\langle M, N \rangle \ge 0\) for every \(N \succeq 0\) exactly when \(M \succeq 0\). So for every feasible \((\lambda, \gamma)\) and every \(M \succeq 0\) the Lagrangian is at least \(\gamma\), and \(d^\star\) is at most the infimum over \(M \succeq 0\) of the supremum of the Lagrangian over \(\gamma \in \mathbb{R}\) and \(\lambda \ge 0\). That supremum is finite only when \(m_0 = 1\) and \(\langle Q_i, X \rangle + c_i^\top x + d_i \le 0\) for every \(i\), and it then equals \(\langle Q_0, X \rangle + c_0^\top x\). Minimizing this over \(M \succeq 0\) with \(m_0 = 1\) is the Shor relaxation, so \(d^\star \le z_{\mathrm{SDP}}\). The inequality \(z_{\mathrm{SDP}} \le z^\star\) holds because \((x, x x^\top)\) is feasible for every feasible \(x\). Strong conic duality under strict feasibility, Theorem 2.2.5 in conic form (Definition 4.8.1), gives the equality. ∎

(What the Shor relaxation is, in one sentence) The Shor relaxation is therefore the Lagrangian dual of the QCQP written in the lifted variables, and its gap is the duality gap of the QCQP. Everything RLT adds corresponds to using quadratic rather than constant multipliers in the Lagrangian. There is one case in which the gap is always zero. It is the case that makes a single quadratic constraint special.

Theorem 4.7.8 (the S-lemma; Yakubovich 1971). Let \(f, g : \mathbb{R}^n \to \mathbb{R}\) be quadratic functions, not necessarily homogeneous or convex, and suppose \(g(\bar x) < 0\) for some \(\bar x\). Then the two statements below are equivalent.V. A. Yakubovich, "S-procedure in nonlinear control theory", Vestnik Leningrad University 1 (1971), 62–77 (in Russian), as cited in I. Pólik and T. Terlaky, "A survey of the S-lemma", SIAM Review 49 (2007), 371–418, which gives the history, several proofs and the counterexamples for two constraints. The convexity of the joint range of two quadratic forms is L. L. Dines, "On the mapping of quadratic forms", Bulletin of the American Mathematical Society 47 (1941), 494–498.

\[\big( g(x) \le 0 \;\Rightarrow\; f(x) \ge 0 \big) \quad\iff\quad \exists\, \lambda \ge 0 : \; f(x) + \lambda g(x) \ge 0 \ \text{ for all } x \in \mathbb{R}^n .\]

Proof sketch. The direction from right to left is immediate: if \(g(x) \le 0\) then \(f(x) \ge -\lambda g(x) \ge 0\). For the other direction take first the homogeneous case \(f(x) = x^\top F x\), \(g(x) = x^\top G x\). Dines' theorem says that the joint range \(\{ (x^\top F x, x^\top G x) : x \in \mathbb{R}^n \}\) of two quadratic forms is a convex cone in the plane. The hypothesis says that this cone does not meet the convex set \(\{ (u, v) : u < 0,\ v \le 0 \}\). Separate the two by a line through the origin. This gives \((\mu, \lambda)\), not both zero, with \(\mu u + \lambda v \ge 0\) on the cone and \(\mu u + \lambda v \le 0\) on the set. The points \((-1, 0)\) and \((0, -1)\) of the set give \(\mu \ge 0\) and \(\lambda \ge 0\), and the cone gives \(\mu f(x) + \lambda g(x) \ge 0\) for all \(x\). The Slater point forces \(\mu > 0\): otherwise \(\lambda g(\bar x) \ge 0\) with \(g(\bar x) < 0\) would give \(\lambda = 0\). Dividing by \(\mu\) gives the multiplier. The non-homogeneous case follows by homogenizing with one extra coordinate and a short limiting argument. ∎

The lemma says that for two quadratics "nonnegative on a set" and "nonnegative combination" coincide. They coincide for linear functions by Farkas' lemma, and they fail for three or more quadratics. The image set of Section 2.2 is convex for one quadratic constraint, which is why the gap closes.

Corollary 4.7.9 (exactness for one constraint). If \(m = 1\) and the constraint is strictly feasible, then \(z_{\mathrm{SDP}} = d^\star = z^\star\). In particular the trust-region subproblem, the minimization of a quadratic over a ball, \(\min\{ x^\top A x + 2 b^\top x : \|x\|^2 \le 1 \}\), has no duality gap for any symmetric \(A\).

Proof. Apply the S-lemma with \(f\) equal to the objective minus \(\gamma\) and \(g\) the constraint: \(z^\star \ge \gamma\) if and only if some \(\lambda \ge 0\) certifies it, and the best \(\gamma\) with a certificate is \(d^\star\). ∎

(The disc example: Shor exact, McCormick blind) A two-variable instance of the corollary is the one the next figure draws: maximize \(xy\) on the unit disc \(x^2 + y^2 \le 1\). The maximum is \(\tfrac12\) at \(x = y = 1/\sqrt 2\), since \(xy \le \tfrac12 (x^2 + y^2) \le \tfrac12\) with equality only there. Lift to \(X_{11} = x^2\), \(X_{22} = y^2\), \(X_{12} = xy\). The Shor relaxation maximizes \(X_{12}\) subject to \(X_{11} + X_{22} \le 1\) and \(M(x, X) \succeq 0\). From \(X \succeq x x^\top \succeq 0\) we get \(X_{12}^2 \le X_{11} X_{22} \le \big(\tfrac{X_{11} + X_{22}}{2}\big)^2 \le \tfrac14\), so \(X_{12} \le \tfrac12\), and the bound is exact. The multiplier of the S-lemma is \(\lambda = \tfrac12\), with \(f = \tfrac12 - xy\) and \(g = x^2 + y^2 - 1\): the certificate is \(f + \tfrac12 g = \tfrac12 (x^2 + y^2 - 1) - (xy - \tfrac12) = \tfrac12 (x - y)^2 \ge 0\). The McCormick relaxation of the same product on a box \([-l, l]^2\) containing the disc behaves very differently. Its two upper planes are \(W \le ly - lx + l^2\) and \(W \le -ly + lx + l^2\), so \(W \le l^2 - l\,|x - y|\), which allows \(W = l^2\) along the whole diagonal \(x = y\). The McCormick bound is \(l^2\), that is \(1\) for \(l = 1\) and \(4\) for \(l = 2\). The RLT products of the box constraints give the same line, because the only linear constraints are the bounds. The single positive-semidefiniteness fact can be isolated at one point. Fix \(x_1 = x_2 = \tfrac12\) and suppose \(X_{11} = X_{22} = \tfrac14\) are known. On the unit square \([0,1]^2\) the McCormick planes allow \(X_{12} \in [0, \tfrac12]\). Positive semidefiniteness of

\[\begin{pmatrix} 1 & \tfrac12 & \tfrac12 \\ \tfrac12 & \tfrac14 & X_{12} \\ \tfrac12 & X_{12} & \tfrac14 \end{pmatrix}\]

has the Schur complement \(X - x x^\top = \big(\begin{smallmatrix} 0 & X_{12} - \frac14 \\ X_{12} - \frac14 & 0 \end{smallmatrix}\big)\), which is positive semidefinite if and only if \(X_{12} = \tfrac14\). One condition pins the product to its true value.

The figure plots the largest \(X_{12}\) each relaxation allows against \(X_{11}\) on \([0, 1]\), with \(X_{22} = 1 - X_{11}\) so that the disc constraint is tight. The grey line at \(l^2\) is the McCormick bound from the box \([-l, l]^2\), the same at every \(X_{11}\). The RLT bound is drawn as white dashes on top of it because the two coincide. The orange semicircle \(\sqrt{X_{11}(1 - X_{11})}\) is Shor's bound, which peaks at the true optimum \(0.500\), marked in blue, at \(X_{11} = \tfrac12\). At the default \(l = 1\) the stats read McCormick bound \(1.000\) and gap \(100\) per cent, SDP bound \(0.500\) and gap zero. At the cursor's default \(X_{11} = 0.25\) the semidefinite relaxation allows \(X_{12} \le \sqrt{0.1875} = 0.433\) while the box allows \(1.000\). Moving the slider to \(l = 0.71\) brings the McCormick bound to \(0.504\), a gap of about one per cent. Moving it to \(l = 1.50\) sends the bound to \(2.250\), a gap of \(350\) per cent. The semidefinite bound never moves, because it does not use the box at all. The readout shows the lifted matrix at the semidefinite optimum with eigenvalues \(2, 0, 0\), a rank-one matrix and so an outer product. It also shows the matrix at a point the box relaxation admits with \(X_{12} = l^2\) (\(x = y = 0\), \(X_{11} = 0.25\), \(X_{22} = 0.75\)), whose eigenvalues \(1.531\), \(1.000\) and \(-0.531\) show that it is not positive semidefinite.

Maximize x·y on the unit disc, lifted to X11 = x², X22 = y², X12 = x·y. The grey line is the largest X12 the McCormick planes on the box [−l, l]² allow, the same for every X11 (RLT on the box gives the same line). The orange semicircle is Shor's semidefinite bound √(X11·X22), and the blue point is the true optimum 0.500. The slider sets the half-width l of the box, and dragging on the plot reads both bounds at one X11.

(Two cautions: rank, and SDP against RLT) Two cautions belong next to the figure. First, an interior-point solver returns the optimal solution of maximum rank. On the disc example that is \(x = 0\) with \(X = \tfrac12 \mathbf 1 \mathbf 1^\top\), not the rank-one solution. Pataki's theorem guarantees that a semidefinite program with \(m\) equality constraints that has an optimal solution has one of rank \(r\) with \(r(r+1)/2 \le m\). A rank-one optimum therefore exists for one constraint, but it has to be extracted, for instance from the top eigenvector of \(X\).G. Pataki, "On the rank of extreme matrices in semidefinite programs and the multiplicity of optimal eigenvalues", Mathematics of Operations Research 23 (1998), 339–358. Further polynomial-time checkable conditions for exactness of the Shor relaxation are in A. L. Wang and F. Kılınç-Karzan, "On the tightness of SDP relaxations of QCQPs", Mathematical Programming 193 (2022), 33–73, and S. Burer and Y. Ye, "Exact semidefinite formulations for a class of (random and non-random) nonconvex quadratic programs", Mathematical Programming 181 (2020), 1–17. Second, the semidefinite relaxation is not stronger than the linear one in general. On the running bilinear example the Shor relaxation alone is unbounded. Nothing bounds \(X_{11}\) and \(X_{22}\) from above, and the cone then lets \(X_{12}\) grow like \(\sqrt{X_{11} X_{22}}\). Adding the four McCormick rows and no other products gives \(0.40\), the McCormick bound itself. At \(x = y = X_{11} = X_{22} = X_{12} = 0.4\) the Schur complement is \(0.24 \cdot \mathbf 1 \mathbf 1^\top\), which is positive semidefinite, so the point the McCormick planes pick survives the cone. Adding instead only the diagonal rows \(X_{11} \le x\) and \(X_{22} \le y\) (the other diagonal rows of Proposition 4.7.3 are implied by \(X \succeq x x^\top\)) gives \(0.4136\), which is worse than McCormick's \(0.40\). Adding the full level-\(1\) RLT gives \(0.18\), exact. The two relaxations see different things, and that is a theorem.

Theorem 4.7.10 (Anstreicher 2009; Anstreicher and Burer 2010). Consider a QCQP with box constraints. The SDP+RLT relaxation is at least as tight as the Shor relaxation and as the RLT relaxation. Neither of the two dominates the other, and there are instances on which their intersection is strictly tighter than both. For \(n = 2\) and the box \([0, 1]^2\), the set of pairs \((x, X)\) with \(M(x, X) \succeq 0\) and the McCormick rows for every pair \(i \le j\), the diagonal rows \(X_{ii} \ge 0\), \(X_{ii} \ge 2x_i - 1\) and \(X_{ii} \le x_i\) included, is exactly \(\operatorname{conv}\{ (x, x x^\top) : x \in [0, 1]^2 \}\). For \(n = 3\) and a box the hull is obtained from a triangulation of the cube together with a doubly nonnegative representation over simplices. A matrix is doubly nonnegative when it is positive semidefinite and entrywise nonnegative. The hull is not obtained from the semidefinite and RLT rows alone.K. M. Anstreicher, "Semidefinite programming versus the reformulation-linearization technique for nonconvex quadratically constrained quadratic programming", Journal of Global Optimization 43 (2009), 471–484; K. M. Anstreicher and S. Burer, "Computable representations for convex hulls of low-dimensional quadratic forms", Mathematical Programming 124 (2010), 33–43. The \(n = 2\) statement is Theorem 6 of the authors' preprint (Optimization Online 2007/02/1586, read 5 October 2026), whose RLT set carries the McCormick rows for every pair \(i \le j\); the triangulation statement for \(n = 3\) is from the same preprint's abstract.

(Why neither dominates) The intuition is that RLT knows the box exactly, since it contains the corners and nothing beyond them in each pair of coordinates, but nothing about positive semidefiniteness. The Shor relaxation knows that \(X \succeq x x^\top\) but, without products, nothing about the box except through the linear terms. The \(n = 2\) statement is why the running example is exact under SDP+RLT. The \(n = 3\) statement is why the next table's three-variable instance is not.

instanceMcCormickRLT level 1Shor aloneShor + diagonalShor + RLTLasserre order 2optimum
max \(xy\), \(2x + y \le 1.2\), \([0,1]^2\)0.40000.3000unbounded0.41360.18000.18000.18
max \(xy\) on the unit disc, box \([-1,1]^2\)1.00001.00000.50000.50000.5000–0.5
Haverly 1 (profit)500.00500.00600.00532.30500.00400.0003400
boxQP, \(n = 3\) (sidenote below)–−3.000–−2.590−2.039–−2.000
Lifted relaxations on four small instances (maximize for the first three; minimize for the last).
Theorem 4.7.10 on the first instance (max xy, 2x + y <= 1.2)

   McCormick         0.4000          Shor alone       unbounded
       |                                  |
       v                                  v
   RLT level 1       0.3000          Shor + diagonal     0.4136
       |                                  |
       +--------------+    +--------------+
                      v    v
                    Shor + RLT   0.1800, the optimum 0.18

  an arrow adds rows or the cone: a bound at least as tight
  neither column contains the other: on the disc Shor alone gives
  0.5000 and RLT level 1 1.0000; on Haverly 1, 600.00 and 500.00

The semidefinite and moment values in the table come from interior-point solves of the relaxations, and the linear programs are the ones the figures solve.The McCormick and RLT values of the first row are those of rlt_ladder.py above. The McCormick and RLT values of the Haverly row were computed for this series with a numpy simplex, and the RLT value of the boxQP row with the HiGHS linear programming solver through cvxpy, both on 5 October 2026. Every semidefinite value and both moment values were computed with the Clarabel interior-point solver (version 0.11.1) through cvxpy on 4 October 2026 and reproduced on 5 October 2026. The three-variable instance is \(\min \tfrac12 x^\top Q x + c^\top x\) over \([0, 1]^3\) with \(Q = \big(\begin{smallmatrix} -2 & -2 & -6 \\ -2 & 2 & 4 \\ -6 & 4 & 0 \end{smallmatrix}\big)\) and \(c = (3, -3, 3)\), whose Hessian has eigenvalues \(-7.229\), \(-0.948\) and \(8.176\) and whose minimum \(-2\) is at \((0, 1, 0)\). On Haverly's pooling problem the Shor relaxation is useless on its own and adds nothing to RLT. On the disc the box relaxation is useless and the semidefinite one is exact. A solver that uses only the polyhedral family, which is every general-purpose global solver today, lives with the first row's \(0.30\) and closes the rest by branching.

(Semidefinite cuts) Semidefinite information can also enter a linear relaxation as cuts. If \(M(x, X)\) has a negative eigenvalue with eigenvector \(v\) at the current LP solution, then \(v^\top M(x, X) v \ge 0\) is a linear inequality in \((x, X)\). It is valid for the lifted set and violated by that solution. Sherali and Fraticelli introduced these semidefinite cuts into RLT. Saxena, Bonami and Lee studied them together with disjunctive cuts for nonconvex MIQCQP, and Qualizza, Belotti and Margot studied their effect inside an LP-based branch and bound.H. D. Sherali and B. M. P. Fraticelli, "Enhancing RLT relaxations via a new class of semidefinite cuts", Journal of Global Optimization 22 (2002), 233–261; A. Saxena, P. Bonami and J. Lee, "Convex relaxations of non-convex mixed integer quadratically constrained programs: extended formulations", Mathematical Programming 124 (2010), 383–411; A. Qualizza, P. Belotti and F. Margot, "Linear programming relaxations of quadratically constrained quadratic programs", in Mixed Integer Nonlinear Programming, IMA Volumes 154 (Springer, 2012), 407–426; arXiv 1206.1633. Gurobi's parameter reference lists three cut families of exactly this kind for its nonconvex quadratic solver, RLTCuts, BQPCuts and PSDCuts. That is the form in which RLT and semidefinite information reached a commercial MILP engine.Gurobi Optimizer Reference Manual, parameters RLTCuts, BQPCuts and PSDCuts, docs.gurobi.com, read 5 October 2026.

The moment hierarchy in a page

The Shor relaxation is the first member of a sequence of semidefinite relaxations that converges to the optimum of any polynomial program on a compact set. The sequence is indexed by a degree, and the price of each step is a matrix whose size is a binomial coefficient.

Definition 4.7.11 (moments, moment matrix, localizing matrix). For a sequence \(y = (y_\alpha)_{\alpha \in \mathbb{N}^n}\) indexed by exponent vectors, the Riesz functional is the linear map \(L_y\big(\sum_\alpha p_\alpha x^\alpha\big) = \sum_\alpha p_\alpha y_\alpha\) on polynomials. The moment matrix of order \(d\) is \(M_d(y)_{\beta\gamma} = y_{\beta + \gamma}\) for \(|\beta|, |\gamma| \le d\), and it has size \(s(d) = \binom{n + d}{d}\). For a polynomial \(g\) the localizing matrix is \(M_k(g\, y)_{\beta\gamma} = \sum_\alpha g_\alpha\, y_{\alpha + \beta + \gamma}\). If \(y\) is the sequence of moments of a probability measure \(\mu\) supported on \(\{ g \ge 0 \}\), then \(M_d(y) \succeq 0\) and \(M_k(g\, y) \succeq 0\) for every \(d, k\), because \(p^\top M_d(y) p = L_y(p(x)^2) = \int p^2\, d\mu \ge 0\) and likewise with the weight \(g\). A sum of squares is a polynomial \(\sigma = \sum_i q_i^2\). With \(m(x)\) the vector of monomials of degree at most \(d\), every sum of squares of degree at most \(2d\) can be written \(\sigma = m(x)^\top \Sigma\, m(x)\) for some \(\Sigma \succeq 0\), its Gram matrix. Conversely every such expression is a sum of squares. The quadratic module generated by \(g_1, \dots, g_m\) is \(\mathcal Q(g) = \{ \sigma_0 + \sum_j \sigma_j g_j : \sigma_j \text{ sums of squares} \}\). It is Archimedean if \(R^2 - \|x\|^2 \in \mathcal Q(g)\) for some \(R\).

The definitions are concrete on the running example, \(n = 2\), with the monomials ordered \(1, x, y\). The moment matrix of order \(1\) is

\[M_1(y) = \begin{pmatrix} 1 & y_{10} & y_{01} \\ y_{10} & y_{20} & y_{11} \\ y_{01} & y_{11} & y_{02} \end{pmatrix},\]

which is the matrix \(M(x, X)\) of Definition 4.7.1 with \(x = (y_{10}, y_{01})\) and \(X = \big(\begin{smallmatrix} y_{20} & y_{11} \\ y_{11} & y_{02} \end{smallmatrix}\big)\). The moments of degree one are the variables, the moments of degree two are the entries of \(X\), and \(M_1(y) \succeq 0\) is the Shor condition. For the linear constraint \(g = 1.2 - 2x - y\) the localizing matrix of order \(0\) is the scalar \(L_y(g) = 1.2 - 2y_{10} - y_{01} \ge 0\), the constraint itself. The order-\(1\) relaxation of a QCQP is therefore the Shor relaxation. At order \(2\) the localizing matrix of \(g\) is

\[M_1(g\,y) = \begin{pmatrix} L(g) & L(gx) & L(gy) \\ L(gx) & L(gx^2) & L(gxy) \\ L(gy) & L(gxy) & L(gy^2) \end{pmatrix} \succeq 0 .\]

Its diagonal entries are the products of \(g\) with the squares, \(L(g x^2) \ge 0\) and \(L(g y^2) \ge 0\), which are valid cubic inequalities, and its minors couple them to the entries \(L(gx)\) and \(L(gy)\). The RLT row \(x(1.2 - 2x - y) \ge 0\) of the figure, which is \(L(gx) \ge 0\), is an off-diagonal entry of this matrix and is not implied by it. That row is the product of two different constraint factors, \(g\) and \(x\). Such products belong to Schmüdgen's preordering, the set of sums \(\sum_J \sigma_J \prod_{j \in J} g_j\) over all subsets \(J\) of the constraint indices with sums of squares \(\sigma_J\), and not to the quadratic module, which admits each \(g_j\) only singly. The hierarchy below is built on the quadratic module.

Theorem 4.7.12 (Lasserre 2001; Putinar 1993; Parrilo 2003). Let \(f, g_1, \dots, g_m\) be polynomials, \(K = \{ x : g_j(x) \ge 0 \}\), and suppose \(\mathcal Q(g)\) is Archimedean, which can be arranged by adding a redundant ball constraint when \(K\) is bounded. For \(d \ge d_0 = \max\{ \lceil \deg f / 2 \rceil, \max_j \lceil \deg g_j / 2 \rceil \}\) define

\[\rho_d = \min_y \Big\{ L_y(f) : y_0 = 1,\ M_d(y) \succeq 0,\ M_{d - \lceil \deg g_j / 2 \rceil}(g_j\, y) \succeq 0,\ j = 1, \dots, m \Big\}.\]

Then \(\rho_d \le \rho_{d+1} \le \dots \le z^\star\) and \(\rho_d \to z^\star\). Here \(\rho_d\) is Lasserre's notation, unrelated to the nonconvexity \(\rho(f)\) of Sections 3.7 and 5.4. The dual of the semidefinite program defining \(\rho_d\) is the sum-of-squares program \(\max\{ \gamma : f - \gamma = \sigma_0 + \sum_j \sigma_j g_j,\ \deg \sigma_0,\ \deg (\sigma_j g_j) \le 2d \}\). Putinar's theorem, that every polynomial strictly positive on \(K\) has a representation in \(\mathcal Q(g)\), gives the convergence.J. B. Lasserre, "Global optimization with polynomials and the problem of moments", SIAM Journal on Optimization 11 (2001), 796–817; M. Putinar, "Positive polynomials on compact semi-algebraic sets", Indiana University Mathematics Journal 42 (1993), 969–984; P. A. Parrilo, "Semidefinite programming relaxations for semialgebraic problems", Mathematical Programming 96 (2003), 293–320. The book-length accounts are M. Laurent, "Sums of squares, moment matrices and optimization over polynomials", in Emerging Applications of Algebraic Geometry, IMA Volumes 149 (Springer, 2009), 157–270, and J. B. Lasserre, An Introduction to Polynomial and Semi-Algebraic Optimization (Cambridge University Press, 2015).

Proof sketch. Validity: the moment sequence of the Dirac measure at any feasible \(x\) is feasible for the program with objective \(f(x)\), so \(\rho_d \le z^\star\). Monotonicity holds because truncating a feasible \(y\) of order \(d + 1\) gives a feasible \(y\) of order \(d\). Weak duality: suppose \(f - \gamma = \sigma_0 + \sum_j \sigma_j g_j\) with Gram matrices \(\Sigma_j \succeq 0\) for the sums of squares. Then for any feasible \(y\), \(L_y(f) - \gamma = \operatorname{tr}(M_d(y)\, \Sigma_0) + \sum_j \operatorname{tr}(M(g_j y)\, \Sigma_j) \ge 0\), since each term is the trace of a product of two positive semidefinite matrices. Convergence: for \(\varepsilon > 0\) the polynomial \(f - z^\star + \varepsilon\) is strictly positive on \(K\), so by Putinar's theorem it has a representation of some degree \(2d(\varepsilon)\), and then \(\rho_d \ge z^\star - \varepsilon\) for all \(d \ge d(\varepsilon)\). ∎

(Five facts about the hierarchy) Five facts go with the theorem. First, the relation to the relaxations above. For a QCQP the order-\(1\) relaxation \(\rho_1\) is the Shor relaxation, as the example before the theorem shows. For 0–1 programs, where \(x_i^2 = x_i\) can be imposed, Laurent proved that the Lasserre relaxation of each order is at least as tight as the Sherali–Adams relaxation of the same order. For continuous variables no such statement holds. The order-\(2\) localizing matrices contain the products of each constraint with squares but not the products of two different constraints, which are the RLT rows, so \(\rho_2\) and SDP+RLT are incomparable in general. On Haverly 1 below \(\rho_2\) closes the gap that SDP+RLT leaves, which is a property of that example and not a theorem. Second, size. \(M_d(y)\) is \(\binom{n+d}{d} \times \binom{n+d}{d}\) with \(\binom{n + 2d}{2d}\) moments. For the running example, \(n = 2\) and \(d = 2\), that is a \(6 \times 6\) matrix with \(15\) moments. For Haverly's problem in its \(7\)-variable p-formulation it is a \(36 \times 36\) matrix with \(330\) moments. For \(n = 50\) it is \(1326 \times 1326\) with \(316{,}251\) moments, and an interior-point iteration costs of the order of the sixth power of the matrix size. Third, finite convergence. Nie proved that \(\rho_d = z^\star\) for some finite \(d\) whenever the standard nonlinear-programming optimality conditions (a constraint qualification, strict complementarity and the second-order sufficient condition) hold at every global minimizer. These conditions hold generically, outside a set of problem data of measure zero. Finite convergence can fail for specific instances.J. Nie, "Optimality conditions and finite convergence of Lasserre's hierarchy", Mathematical Programming 146 (2014), 97–121. Fourth, a certificate. The flat-truncation condition \(\operatorname{rank} M_{d - d_0}(y) = \operatorname{rank} M_d(y)\) at the solution certifies \(\rho_d = z^\star\) and allows the minimizers to be extracted from \(y\).J. Nie, "Certifying convergence of Lasserre's hierarchy via flat truncation", Mathematical Programming 142 (2013), 485–510. Fifth, sparsity. When the problem's variables interact through a chordal sparsity pattern, one whose maximal cliques \(I_1, \dots, I_p\) can be ordered so that each meets the union of its predecessors inside a single earlier clique (the running intersection property), the one large moment matrix can be replaced by one block per clique. Each block has size \(\binom{|I_k| + d}{d}\), and convergence is preserved. A further decomposition by the terms that actually appear gives smaller blocks still.H. Waki, S. Kim, M. Kojima and M. Muramatsu, "Sums of squares and semidefinite program relaxations for polynomial optimization problems with structured sparsity", SIAM Journal on Optimization 17 (2006), 218–242; J. B. Lasserre, "Convergent SDP-relaxations in polynomial optimization with sparsity", SIAM Journal on Optimization 17 (2006), 822–843; J. Wang, V. Magron and J. B. Lasserre, "TSSOS: a moment-SOS hierarchy that exploits term sparsity", SIAM Journal on Optimization 31 (2021), 30–58.

The order-2 moment matrix M_2(y) of the running example (n = 2)

              1     x     y     x^2   xy    y^2
      1    [ y00   y10   y01 | y20   y11   y02 ]
      x    [ y10   y20   y11 | y30   y21   y12 ]
      y    [ y01   y11   y02 | y21   y12   y03 ]
           [ ----------------+---------------- ]
      x^2  [ y20   y30   y21 | y40   y31   y22 ]
      xy   [ y11   y21   y12 | y31   y22   y13 ]
      y^2  [ y02   y12   y03 | y22   y13   y04 ]

  entry (beta, gamma) = y_{beta + gamma}, with y00 = 1
  top left 3 x 3: M_1(y) = M(x, X), the Shor matrix
  6 x 6 = C(4, 2) x C(4, 2), and C(6, 4) = 15 distinct moments

On the two examples the hierarchy does what the theorem promises. For the running bilinear problem \(\rho_1\) is unbounded and \(\rho_2 = 0.18\), exact. For Haverly's problem, with every variable scaled to \([0, 1]\) first so that the moments stay of order one, \(\rho_1 = 600\) and \(\rho_2 = 400.0003\), exact to solver tolerance. The McCormick, RLT and SDP+RLT bounds all stop at \(500\). The cost difference is a \(36 \times 36\) semidefinite program against a linear program in nine variables, and the pooling problem with fifty variables would need the \(1326 \times 1326\) matrix. One block per pool is what would make it usable. The hierarchy is a root-node tool and a tool for structured problems, not a per-node bound, unless first-order semidefinite methods on GPUs change the arithmetic.

The copositive endpoint

The last member of the family is not a relaxation but an exact reformulation.

Theorem 4.7.13 (Burer 2009). Consider the quadratic program

\[\min\{ x^\top Q x + 2 c^\top x : a_i^\top x = b_i\ (i = 1, \dots, m),\ x \ge 0,\ x_j \in \{0, 1\}\ (j \in \mathcal B) \}\]

and assume that \(x \ge 0\) together with \(a_i^\top x = b_i\) for all \(i\) implies \(0 \le x_j \le 1\) for \(j \in \mathcal B\). Let \(\mathcal{CP}_{n+1} = \operatorname{conv}\{ v v^\top : v \in \mathbb{R}^{n+1}_{\ge 0} \}\) be the completely positive cone. Then the optimal value equals that of the linear conic program below.S. Burer, "On the copositive representation of binary and continuous nonconvex quadratic programs", Mathematical Programming 120 (2009), 479–495. Surveys: M. Dür, "Copositive programming – a survey", in Recent Advances in Optimization and its Applications in Engineering (Springer, 2010), 3–20; S. Burer, "Copositive programming", in Handbook on Semidefinite, Conic and Polynomial Optimization (Springer, 2012), 201–218.

\[\min\Big\{ \langle Q, X \rangle + 2 c^\top x : a_i^\top x = b_i,\ a_i^\top X a_i = b_i^2,\ X_{jj} = x_j\ (j \in \mathcal B),\ \begin{pmatrix} 1 & x^\top \\ x & X \end{pmatrix} \in \mathcal{CP}_{n+1} \Big\}.\]

Proof sketch. A feasible matrix is a convex combination \(\sum_k \lambda_k (1, x^k)(1, x^k)^\top\) of rank-one completely positive matrices, up to a recession part that the constraints remove. The equations \(a_i^\top X a_i = b_i^2\) together with \(a_i^\top x = b_i\) say that \(\sum_k \lambda_k (a_i^\top x^k)^2 = \big(\sum_k \lambda_k a_i^\top x^k\big)^2\). That is a variance of zero, which forces every \(a_i^\top x^k\) to equal \(b_i\). The equations \(X_{jj} = x_j\) say that \(\sum_k \lambda_k (x^k_j)^2 = \sum_k \lambda_k x^k_j\), which with \(0 \le x^k_j \le 1\) forces \(x^k_j \in \{0, 1\}\). The objective is linear in \((x, X)\), so its value over the convex hull equals its value over the extreme points, which are feasible points of the original problem. ∎

(Where the difficulty went) Every nonconvexity, the quadratic objective and the binary constraints alike, has moved into the cone. The difficulty has moved with it. Membership in the dual cone, the cone of copositive matrices, is co-NP-complete (Murty and Kabadi), and membership in the completely positive cone itself is NP-hard (Dickinson and Gijben).K. G. Murty and S. N. Kabadi, "Some NP-complete problems in quadratic and nonlinear programming", Mathematical Programming 39 (1987), 117–129; P. J. C. Dickinson and L. Gijben, "On the computational complexity of membership problems for the completely positive cone and its dual", Computational Optimization and Applications 57 (2014), 403–415. Replacing \(\mathcal{CP}_{n+1}\) by the doubly nonnegative cone of matrices that are both positive semidefinite and entrywise nonnegative gives a tractable relaxation. It is the Shor relaxation plus nonnegativity of \(X\) plus the squared equalities. It is exact for \(n + 1 \le 4\), because the two cones coincide up to order four, and strictly weaker from order five on.P. H. Diananda, "On non-negative forms in real variables some or all of which are non-negative", Mathematical Proceedings of the Cambridge Philosophical Society 58 (1962), 17–25; the attribution of the order-four statement to this paper follows A. Berman, M. Dür and N. Shaked-Monderer, "Open problems in the theory of completely positive and copositive matrices", Electronic Journal of Linear Algebra 29 (2015), 46–58, p. 47. The doubly nonnegative relaxation is solved by first-order methods in S. Burer, "Optimizing a polyhedral-semidefinite relaxation of completely positive programs", Mathematical Programming Computation 2 (2010), 1–19, and S. Kim, M. Kojima and K.-C. Toh, "A Lagrangian–DNN relaxation: a fast method for computing tight lower bounds for a class of quadratic optimization problems", Mathematical Programming 156 (2016), 161–187. The doubly nonnegative relaxation is the strongest convex relaxation of binary quadratic programs in general use. It is the bound inside the exact max-cut and binary quadratic solvers of the BiqMac family.F. Rendl, G. Rinaldi and A. Wiegele, "Solving Max-Cut to optimality by intersecting semidefinite and polyhedral relaxations", Mathematical Programming 121 (2010), 307–335; N. Krislock, J. Malick and F. Roupin, "BiqCrunch: a semidefinite branch-and-bound method for solving binary quadratic problems", ACM Transactions on Mathematical Software 43 (2017), article 32; N. Gusmeroli, T. Hrga, B. Lužar, J. Povh, M. Siebenhofer and A. Wiegele, "BiqBin: a parallel branch-and-bound solver for binary quadratic problems with linear constraints", ACM Transactions on Mathematical Software 48 (2022), article 15.

Burer's reformulation (Theorem 4.7.13) and the DNN relaxation

  min x^T Q x + 2 c^T x  s.t.  a_i^T x = b_i,  x >= 0,
                               x_j in {0, 1} (j in B),
      where x >= 0 and a_i^T x = b_i imply 0 <= x_j <= 1 (j in B)
          |
          |  lift to (x, X): the same optimal value
          v
  min <Q, X> + 2 c^T x  s.t.  a_i^T x = b_i,  a_i^T X a_i = b_i^2,
      X_jj = x_j (j in B),  M(x, X) in CP_{n+1}
      a linear conic program; membership in CP_{n+1} is NP-hard
          |
          |  replace CP_{n+1} by the doubly nonnegative cone (DNN):
          |  PSD and entrywise >= 0
          v
  a tractable relaxation: exact for n + 1 <= 4, where the two cones
  coincide; strictly weaker from order five on

Finite branch and bound for nonconvex quadratic programs

Spatial branching on a continuous variable converges only to a tolerance. For a quadratic program with linear constraints there is a different disjunction to branch on, and it makes the tree finite.

Theorem 4.7.14 (Burer and Vandenbussche 2008; Chen and Burer 2012). Consider \(\min \tfrac12 x^\top Q x + c^\top x\) over the bounded polyhedron \(\{ Ax = b,\ x \ge 0 \}\). Every local minimizer satisfies the KKT conditions, with multipliers \(\lambda\) for the equations and \(\mu \ge 0\) for the nonnegativity constraints,

\[Q x + c - A^\top \lambda - \mu = 0, \qquad \mu \ge 0, \qquad x \ge 0, \qquad \mu_j x_j = 0 \quad (j = 1, \dots, n),\]

and at every KKT point the objective equals \(\tfrac12 (c^\top x + b^\top \lambda)\). Branching on the complementarities, \(x_j = 0\) or \(\mu_j = 0\), produces a tree of depth at most \(n\) whose leaves are linear programs, so the branch and bound is finite. The node relaxations are SDP+RLT relaxations in \((x, \lambda, \mu)\), or doubly nonnegative relaxations of the completely positive reformulation of Theorem 4.7.13.S. Burer and D. Vandenbussche, "A finite branch-and-bound algorithm for nonconvex quadratic programming via semidefinite relaxations", Mathematical Programming 113 (2008), 259–282; J. Chen and S. Burer, "Globally solving nonconvex quadratic programming problems via completely positive programming", Mathematical Programming Computation 4 (2012), 33–52, whose code is QuadProgBB. The box-constrained case by branch and cut is D. Vandenbussche and G. L. Nemhauser, "A branch-and-cut algorithm for nonconvex quadratic programs with box constraints", Mathematical Programming 102 (2005), 559–575.

Proof of the objective identity. Multiply the stationarity equation by \(x^\top\): \(x^\top Q x + c^\top x - b^\top \lambda - \mu^\top x = 0\), and \(\mu^\top x = 0\) by complementarity, so \(x^\top Q x = b^\top \lambda - c^\top x\) and \(\tfrac12 x^\top Q x + c^\top x = \tfrac12 (c^\top x + b^\top \lambda)\). ∎

(Complementarity branching on one variable) The nonconvexity of a quadratic program lives entirely in the complementarity conditions, which are a finite disjunction. Once they are fixed, what remains is linear. The one-variable instance shows the tree. Minimize \(-x^2\) on \([0, 1]\). The KKT conditions are \(-2x - \mu_0 + \mu_1 = 0\), \(\mu_0 x = 0\) and \(\mu_1 (1 - x) = 0\) with \(\mu_0, \mu_1 \ge 0\). Branching on the two complementarities gives four leaves. The leaf \(x = 0\) with \(x = 1\) is infeasible. The leaf \(x = 0\) with \(\mu_1 = 0\) gives the KKT point \(x = 0\) of value \(0\). The leaf \(\mu_0 = 0\) with \(x = 1\) gives \(\mu_1 = 2\) and the value \(-1\). The leaf \(\mu_0 = \mu_1 = 0\) gives \(x = 0\) and value \(0\) again. The tree is finite and the optimum \(x = 1\) is proved. This example is too small to show the method's advantage. A single concave term on an interval is relaxed exactly at the root by its chord, here \(-x\): by Falk's theorem (Theorem 2.4.4) the chord's minimum \(-1\) at \(x = 1\) is the minimum of \(-x^2\). The point of complementarity branching is finiteness when polyhedral constraints couple several variables, where spatial branching only converges to a tolerance.

Complementarity branching on min -x^2 over [0, 1] (Theorem 4.7.14)

  KKT: -2x - mu_0 + mu_1 = 0,  mu_0 x = 0,  mu_1 (1 - x) = 0

                               root
                  x = 0  /              \  mu_0 = 0
                        /                \
                     node                 node
          x = 1  /       \ mu_1 = 0   x = 1 /     \ mu_1 = 0
                /         \                /       \
       infeasible       x = 0,          x = 1,       x = 0,
                        value 0         mu_1 = 2,    value 0
                                        value -1:
                                        the optimum

  four leaves, all linear programs: the tree is finite

(The Boolean quadric polytope) The box-constrained case has its own polyhedral theory. The projection of \(\operatorname{conv}\{ (x, x x^\top) : x \in [0, 1]^n \}\) onto the variables \(x\) and the off-diagonal entries of \(X\) is Padberg's Boolean quadric polytope. Its facets, the triangle inequalities among them, are therefore valid cuts for box-constrained quadratic programs.M. Padberg, "The boolean quadric polytope: some characteristics, facets and relatives", Mathematical Programming 45 (1989), 139–172; S. Burer and A. N. Letchford, "On nonconvex quadratic programming with box constraints", SIAM Journal on Optimization 20 (2009), 1073–1089. Bonami, Günlük and Linderoth built a solver on this. They report that cuts from the Boolean quadric polytope reduce the optimality gap of box-constrained quadratic programs considerably, and they describe the Chvátal–Gomory closure of that polytope. Their solver combines the cuts with spatial branching, a branching rule based on integrality and a strengthened convex quadratic relaxation. Most of these techniques were implemented in CPLEX.P. Bonami, O. Günlük and J. Linderoth, "Globally solving nonconvex quadratic programming problems with box constraints via integer programming methods", Mathematical Programming Computation 10 (2018), 333–382; the summary follows the abstract of the Optimization Online preprint (2016/06/5488), which states that the Chvátal–Gomory closure of the Boolean quadric polytope is given by the odd-cycle inequalities even when the underlying graph is not complete, and that most of the techniques were implemented in CPLEX.

Where this is used

What the commercial solvers adopted follows this lineage. CPLEX 12.6 solved nonconvex mixed-integer quadratic programs with a nonconvex objective to global optimality. For a mixed-integer quadratic program with products of binaries, CPLEX also decides between linearizing the products with the McCormick rows, which are exact for binaries, and keeping the quadratic. Since version 12.10 that decision is made by a learned classifier.C. Bliek, P. Bonami and A. Lodi, "Solving mixed-integer quadratic programming problems with IBM-CPLEX: a progress report", Proceedings of the 26th RAMP Symposium (2014); P. Bonami, A. Lodi and G. Zarpellon, "A classifier to decide on the linearization of mixed-integer quadratic problems in CPLEX", Operations Research 70 (2022), 3303–3320. Gurobi 9.0, released in November 2019, introduced nonconvex quadratic optimization. Its documentation describes the method as translating the quadratics into bilinear form and solving by spatial branching, and its parameter reference lists the RLT, Boolean-quadric-polytope and semidefinite cut families named above. Version 11.0 (November 2023) extended spatial branching and outer approximation to general nonlinear constraints.Gurobi Optimization, "Gurobi release and support history", support.gurobi.com article 360048138771 (read 5 October 2026), which dates 9.0.0 to November 2019; the archived Gurobi home page of 11 December 2019 lists "Non-Convex Quadratic Optimization" as the release's first feature; GAMS 30.1.0 release notes (10 January 2020): "Gurobi 9 comes with a new bilinear solver, which allows to solve non-convex quadratic programming problems"; Gurobi Optimizer Reference Manual, "Constraints", for the bilinear translation and the NonConvex parameter; "Additions, changes and removals in Gurobi 11.0", cited in the previous subsection, for the extension to nonlinear functions. BARON's group developed multiterm polyhedral relaxations that treat several bilinear terms at once and reviewed semidefinite relaxations of QCQP against their polyhedral surrogates. BARON itself uses the polyhedral family and lists no semidefinite bounds.X. Bao, N. V. Sahinidis and M. Tawarmalani, "Multiterm polyhedral relaxations for nonconvex, quadratically constrained quadratic programs", Optimization Methods and Software 24 (2009), 485–504; X. Bao, N. V. Sahinidis and M. Tawarmalani, "Semidefinite relaxations for quadratically constrained quadratic programming: a review and comparisons", Mathematical Programming 129 (2011), 129–157. SCIP's quadratic handler separates intersection cuts from maximal quadratic-free sets, the family described in Section 3.3, since version 8.A. Chmiela, G. Muñoz and F. Serrano, "On the implementation and strengthening of intersection cuts for QCQPs", Mathematical Programming 197 (2023), 549–586. Semidefinite bounds themselves are used at the nodes only by the specialized binary quadratic solvers named above and by QuadProgBB.

What parallelizes

For GPU work the quadratic case is where the lifted relaxations are understood well enough to attempt batching them. The RLT linear program has \(O(n^2)\) variables and rows that are outer products of constraint rows. A first-order LP method of the PDHG kind, Section 7.2, needs only the products \(Ax\) and \(A^\top y\), which can be formed from the factors without storing the matrix. First-order semidefinite methods based on the low-rank factorization \(X = R R^\top\) now solve max-cut relaxations with matrices of order \(10^7\) in seconds to minutes on a single GPU. The open problem is a safe dual bound from an inexact low-rank iterate at controlled cost. Only a dual-feasible point gives a valid bound, and a Burer–Monteiro iterate must be converted into one by projecting the dual slack onto the positive semidefinite cone before a tree may prune on it.S. Burer and R. D. C. Monteiro, "A nonlinear programming algorithm for solving semidefinite programs via low-rank factorization", Mathematical Programming 95 (2003), 329–357; N. Boumal, V. Voroninski and A. S. Bandeira, "Deterministic guarantees for Burer–Monteiro factorizations of smooth semidefinite programs", Communications on Pure and Applied Mathematics 73 (2020), 581–608; Q. Han, Z. Lin, H. Liu, C. Chen, Q. Deng, D. Ge and Y. Ye, "Accelerating low-rank factorization-based semidefinite programming algorithms on GPU", arXiv 2407.15049 (2024), the cuLoRADS code, whose max-cut timings on an H100 are the authors' numbers. Section 7.6 returns to GPU semidefinite solvers. The intersection cuts and the semidefinite eigenvector cuts are one eigendecomposition per constraint and one scalar equation per LP ray. They are independent across the violated rows and across the nodes of a frontier, which is the shape a batched kernel wants.

Cones

Several of the strong formulations of this section are not linear programs. The perspective of a quadratic, the envelope of \(x/y\), the hull of \(xy \ge 1\) from Proposition 1.2.5 and the Shor relaxation are conic constraints: the variable is required to lie in a convex cone. This subsection names the cone that covers the portfolio problems of this series, the second-order cone, and lists what can be written with it. It then asks the question a branch and bound has to answer: solve the conic relaxation at every node as a conic program, or approximate the cone by a polyhedron and solve linear programs. The two answers are the lifted polyhedral approximation of Ben-Tal and Nemirovski and the outer approximation with conic certificates of Lubin, Coey and Vielma.

Definition 4.8.1 (second-order cone; MISOCP; dual cone). The second-order cone in \(\mathbb{R}^{k+1}\) is \(\mathcal L^{k+1} = \{ (u, t) \in \mathbb{R}^k \times \mathbb{R} : \|u\|_2 \le t \}\). The rotated second-order cone of Proposition 1.2.5, \(\{ (u, s, t) : \|u\|_2^2 \le 2 s t,\ s, t \ge 0 \}\), is the image of \(\mathcal L^{k+2}\) under a linear map. A second-order cone program, named in the introduction to Section 4, is the minimization of a linear function subject to linear equations and the requirement that affine images of the variables lie in such cones. A mixed-integer second-order cone program (MISOCP) adds integrality on some variables. The dual cone of a closed convex cone \(K\) is \(K^* = \{ \lambda : \lambda^\top s \ge 0 \text{ for all } s \in K \}\). The second-order cone and the semidefinite cone are self-dual, \(K^* = K\). The dual of a conic program is a conic program over the dual cone. Weak duality always holds, and strong duality holds under a Slater condition, which is Theorem 2.2.5 in conic form. Second-order cone programs are solved by interior-point methods in polynomial time, like linear programs.A. Ben-Tal and A. Nemirovski, Lectures on Modern Convex Optimization (SIAM, 2001), Chapters 2 and 3; M. S. Lobo, L. Vandenberghe, S. Boyd and H. Lebret, "Applications of second-order cone programming", Linear Algebra and its Applications 284 (1998), 193–228, which is the catalogue the next proposition condenses.

Proposition 4.8.2 (what the second-order cone can write). Each of the following sets is the projection of a set defined by linear equations and second-order cone constraints.

(a) The hyperbolic set \(\{ (x, s, t) : x^2 \le s t,\ s, t \ge 0 \}\), by the cone identity of Lemma 4.3.2, \(x^2 \le s t,\ s, t \ge 0 \iff \| (2x,\ s - t) \|_2 \le s + t\).

(b) The epigraph of the perspective of a convex quadratic, \(t z \ge x^2\) with \(t, z \ge 0\), which is (a), and hence the hull of the on/off set, Theorem 4.3.4(b).

(c) The convex side of a bilinear inequality, \(xy \ge 1\) with \(x, y \ge 0\). This is (a) with \(x \mapsto 1\): \(\|(2, x - y)\|_2 \le x + y\).

(d) The quadratic-over-linear function \(s^2 / y\) on \(y > 0\), hence the convex envelope of \(x/y\) on a box, whose interior formula is \(s^2 / y\) with \(s\) affine in \(x\) (Section 2.4).

(e) Any convex quadratic constraint \(x^\top P x + q^\top x + r \le 0\) with \(P \succeq 0\). Factor \(P = F^\top F\) and write \(\|F x\|_2^2 \le -q^\top x - r\) as an instance of (a). Norms, sums of norms and the geometric mean of nonnegative variables are representable in the same way.

Parts (a) and (b) were proved in Section 4.3, and the two readings after Lemma 4.3.2, \(t \ge 1\) at \(\lambda = 1\) and \(t \ge 2\) at \(\lambda = \tfrac12\), are the picture to keep. Parts (c) to (e) are the same identity applied to other affine images.The second-order cone form of the \(x/y\) envelope is M. Tawarmalani and N. V. Sahinidis, "Semidefinite relaxations of fractional programs via novel convexification techniques", Journal of Global Optimization 20 (2001), 133–154. SCIP 10.0.0's changelog lists, under nonlinearity, "extended SOC detection to simple bilinear constraints, e.g., x*y >= 1" (CHANGELOG of the scipopt/scip repository, SCIP 10.0.0 released 24 November 2025), which is (c) detected automatically.

The hyperbolic set behind the cone: at a chosen x, the region x² ≤ s t, s, t ≥ 0 is everything above the hyperbola t = x²/s, with the points at s = 1 and s = ½ marked (at x = 1, the two readings of Lemma 4.3.2); it is the single second-order cone constraint ‖(2x, s − t)‖₂ ≤ s + t, and with x ↦ 1 the convex side of x y ≥ 1, case (c) of Proposition 4.8.2.

(The portfolio problems as MISOCPs) The portfolio problems of Section 4.4 live in this class. A mean–variance objective with a factor model is a convex quadratic, case (e). A cardinality limit written with the perspective is a family of rotated cones, case (b). Minimum positions and round lots are linear in the binaries. So the natural exact formulation of the cardinality-constrained portfolio problem is a MISOCP. Vielma, Dunning, Huchette and Lubin showed that writing the convex quadratic in an extended separable form makes the outer approximations below markedly stronger. The extended form uses one cone per factor or per term rather than one cone for the whole quadratic. The reason is the one for which separable extended formulations strengthen outer approximation in Section 3.4.J. P. Vielma, I. Dunning, J. Huchette and M. Lubin, "Extended formulations in mixed integer conic quadratic programming", Mathematical Programming Computation 9 (2017), 369–418. Extended formulations and projection are a subject of their own in integer programming, with lower bounds on the size of any extended formulation of certain polytopes. The lifting of inequalities from a lower-dimensional set to a higher-dimensional one is the general tool behind the hull descriptions of this section.M. Conforti, G. Cornuéjols and G. Zambelli, "Extended formulations in combinatorial optimization", 4OR 8 (2010), 1–48; M. Tawarmalani, J.-P. P. Richard and K. Chung, "Strong valid inequalities for orthogonal disjunctions and bilinear covering sets", Mathematical Programming 124 (2010), 481–512. Disjunctive cuts for conic programs, the conic analogue of Section 3.3's lift-and-project, are F. Kılınç-Karzan, "On minimal valid inequalities for mixed integer conic programs", Mathematics of Operations Research 41 (2016), 477–510, and A. Lodi, M. Tanneau and J. P. Vielma, "Disjunctive cuts in mixed-integer conic optimization", Mathematical Programming 199 (2023), 671–719.

Three ways to relax a cone at a node

(Three node relaxations) A branch and bound over a MISOCP has three choices for the node relaxation. It can solve the continuous conic program at every node by an interior-point method. That gives the exact relaxation value but warm-starts poorly, since interior-point methods restart from the centre. It can keep a convex quadratic relaxation and solve it by an active-set method that warm-starts from the parent. That is what the MIQP engines of CPLEX and Gurobi do when the problem is a quadratic program. Or it can replace every cone by a polyhedron and solve linear programs, warm-started by the dual simplex method, which is the fastest per node and the loosest. MOSEK's mixed-integer conic optimizer exposes the first and third as a switch. Its default solves the conic relaxations by the interior-point method, and one parameter turns on an outer approximation instead.MOSEK Optimizer API for Python 11.2, parameter reference, MSK_IPAR_MIO_CONIC_OUTER_APPROXIMATION: "If this option is turned on outer approximation is used when solving relaxations of conic problems; otherwise interior point is used", default off, docs.mosek.com/latest/pythonapi/parameters.html, read 5 October 2026; the algorithm is described in the same manual's chapter "The optimizer for mixed-integer problems". The polyhedral choice raises the question of how many inequalities a cone costs.

Proposition 4.8.3 (polyhedral approximation of the second-order cone; Ben-Tal and Nemirovski 2001). (a) In the plane, a regular \(k\)-gon circumscribed about the unit circle contains the circle and lies within the factor \(1 / \cos(\pi / k)\) of it. A polyhedral outer approximation of the cone \(\{ \|(x_1, x_2)\| \le t \}\) by the \(k\) corresponding inequalities therefore has relative error \(1 / \cos(\pi/k) - 1\). (b) The same cone admits a lifted polyhedral approximation with \(\nu\) levels of auxiliary variables. Start with \(\xi^0 \ge |x_1|\) and \(\eta^0 \ge |x_2|\), each absolute value written as two linear inequalities. For \(j = 1, \dots, \nu\) let \(\theta_j = \pi / 2^{j+1}\) and add

\[\xi^j = \cos\theta_j\, \xi^{j-1} + \sin\theta_j\, \eta^{j-1}, \qquad \eta^j \ge \big| -\sin\theta_j\, \xi^{j-1} + \cos\theta_j\, \eta^{j-1} \big| ,\]

and finish with \(\xi^\nu \le t\) and \(\eta^\nu \le \tan(\pi / 2^{\nu+1})\, \xi^\nu\). The system has \(2\nu + 2\) auxiliary variables and \(3\nu + 6\) rows. Its projection onto \((x_1, x_2, t)\) is \(\{ \|(x_1, x_2)\| \le t / \cos(\pi / 2^{\nu+1}) \}\), which contains the cone and lies within the factor \(1/\cos(\pi/2^{\nu+1})\) of it. The tower of \(\nu\) rotations is therefore the regular \(2^{\nu+1}\)-gon written with \(O(\nu)\) inequalities instead of \(2^{\nu+1}\). (c) The cone \(\mathcal L^{k+1}\) in \(k + 1\) dimensions is the projection of \(k - 1\) three-dimensional cones arranged in a binary tree. It therefore admits a polyhedral approximation to relative accuracy \(\varepsilon\) with \(O(k \log (1/\varepsilon))\) variables and inequalities.A. Ben-Tal and A. Nemirovski, "On polyhedral approximations of the second-order cone", Mathematics of Operations Research 26 (2001), 193–205. The lifted LP-based branch and bound built on it is J. P. Vielma, S. Ahmed and G. L. Nemhauser, "A lifted linear programming branch-and-bound algorithm for mixed-integer conic quadratic programs", INFORMS Journal on Computing 20 (2008), 438–450.

(How the tower works) The mechanism of (b) is a sequence of foldings. The pair \((\xi^0, \eta^0) = (|x_1|, |x_2|)\) has the norm of \((x_1, x_2)\) and an angle in \([0, \pi/2]\). Level \(j\) rotates the pair by \(\theta_j\) and takes the absolute value of the second coordinate. The norm is unchanged and the angle, which lay in \([0, \pi/2^j]\), now lies in \([0, \pi/2^{j+1}]\). After \(\nu\) levels the point lies in a sector of angle \(\pi/2^{\nu+1}\). The last two rows cut that sector with one edge of the circumscribed \(2^{\nu+1}\)-gon, which is where the factor \(1/\cos(\pi/2^{\nu+1})\) comes from. In the relaxation the \(\eta^j\) are only bounded below, so the norm of the lifted pair can grow from level to level. The two closing rows are what keep the growth within that factor.

The Ben-Tal–Nemirovski tower, every level halves the angle: a point of the unit circle is folded into the first quadrant, then rotated by π/4, π/8, … and folded again, until it lies in a sector of opening π/2^(ν+1); the closing rows cut that sector with one edge of the circumscribed 2^(ν+1)-gon, which is where the factor 1/cos(π/2^(ν+1)) comes from (ν = 4: a 32-gon, excess 4.84·10⁻³; ν = 7, beyond the picture: a 256-gon, excess 7.53·10⁻⁵).
The Ben-Tal--Nemirovski tower: every level halves the angle

  (x1, x2)
     |  absolute values
     v
  (xi^0, eta^0) = (|x1|, |x2|)      angle in [0, pi/2]
     |  rotate by theta_1 = pi/4, then take the absolute
     |  value of the second coordinate
     v
  (xi^1, eta^1)                     angle in [0, pi/4]
     |  rotate by theta_2 = pi/8, absolute value
     v
  (xi^2, eta^2)                     angle in [0, pi/8]
     :
  (xi^nu, eta^nu)                   angle in [0, pi/2^(nu+1)]
     |  xi^nu <= t,  eta^nu <= tan(pi/2^(nu+1)) xi^nu
     v
  one edge of the regular 2^(nu+1)-gon: within the factor
  1/cos(pi/2^(nu+1)) of the cone

  the norm is unchanged at every level; in the relaxation the
  eta^j are only bounded below, and the two closing rows keep the
  growth within that factor
  2 nu + 2 auxiliary variables, 3 nu + 6 rows; nu = 4: a 32-gon,
  excess 4.84e-03; nu = 7: a 256-gon, excess 7.53e-05

(Facets against levels, in numbers) In the plane the arithmetic of (a) is enough to see the problem. One per cent accuracy needs \(k \ge 23\) facets, and the error falls only like \(\pi^2 / (2k^2)\). A cone in fifty dimensions approximated by a direct facet list is therefore out of reach. The tower of (c) reaches one per cent with four levels of two-dimensional rotations per three-dimensional cone and \(10^{-4}\) with seven. The script prints the numbers.

# cone_polygon.py -- outer polyhedral approximations of the second-order
# cone in the plane.
#
# A regular k-gon circumscribed about the unit circle overshoots it by
# the factor 1/cos(pi/k); the Ben-Tal--Nemirovski tower with nu levels
# is the k-gon with k = 2^(nu + 1).

import numpy as np

print(" k    excess 1/cos(pi/k) - 1")
for k in (8, 16, 22, 23, 32, 64, 128):
    print(f"{k:3d}    {100*(1/np.cos(np.pi/k) - 1):6.2f} %")

# the smallest k whose excess is at most one per cent
k = 3
while 1/np.cos(np.pi/k) - 1 > 0.01:
    k += 1
print(f"one percent needs k >= {k}")

print("levels nu    k = 2^(nu+1)   excess")
for nu in (2, 3, 4, 5, 7, 11):
    k = 2**(nu + 1)
    print(f"{nu:6d}       {k:5d}        {1/np.cos(np.pi/k) - 1:.2e}")

It prints \(8.24\), \(1.96\), \(1.03\), \(0.94\), \(0.48\), \(0.12\) and \(0.03\) per cent for \(k = 8, 16, 22, 23, 32, 64, 128\), and the line "one percent needs k >= 23". For the tower it prints \(8.2 \cdot 10^{-2}\), \(2.0 \cdot 10^{-2}\), \(4.8 \cdot 10^{-3}\), \(1.2 \cdot 10^{-3}\), \(7.5 \cdot 10^{-5}\) and \(2.9 \cdot 10^{-7}\) for \(\nu = 2, 3, 4, 5, 7, 11\) levels. The script costs a few dozen evaluations of a cosine, and the evaluations for different \(k\) and \(\nu\) are independent, so the loop is trivially parallel. What it shows is the two rates: the error falls quadratically in the number of facets and exponentially in the number of levels.

A polygon around the circle, and what it costs: a regular k-gon circumscribed about the unit circle overshoots it by the factor 1/cos(π/k), so one per cent needs k ≥ 23 facets and the error falls only like π²/(2k²); the Ben-Tal–Nemirovski tower reaches the same accuracy with 3ν + 6 rows for a 2^(ν+1)-gon, one per cent at ν = 4 and 10⁻⁴ at ν = 7. The numbers are the section's script output.

(Cuts from conic certificates) A fixed polyhedral approximation, however good, leaves the solver with an approximate problem. The alternative is to add linear inequalities to the cone's outer approximation only where the search needs them. That is outer approximation in the sense of Section 3.4 with one refinement: the cuts come from dual solutions of the continuous conic subproblems rather than from gradients.

Algorithm 4.8.4 (outer approximation with conic certificates; Lubin, Yamangil, Bent and Vielma 2018; Coey, Lubin and Vielma 2020).

Algorithm CONIC-OA

Input   min c^T (x, y)  s.t.  A (x, y) + s = b,  s in K,  y integer,
        with K a product of closed convex cones (second-order, rotated,
        exponential, semidefinite), y bounded; tolerance eps.
Output  an eps-optimal solution and a proof, or a proof of
        infeasibility.

 0. Initial cuts: a finite set of inequalities lambda^T s >= 0 with
    lambda in K* (the dual cone), for instance the cuts that make each
    cone's projection onto its axes correct; let P_0 be the polyhedron
    they define.

 1. Master: solve the MILP
       min c^T (x, y)  s.t.  A (x, y) + s = b,  s in P_t,  y integer.
    If infeasible, stop: the problem is infeasible. Its value LB_t is a
    lower bound; let y^t be its integer part.

 2. Subproblem: fix y = y^t and solve the continuous conic program in
    (x, s) with an interior-point method.

       a. If it is optimal with primal-dual pair (x^t, s^t, lambda^t),
          lambda^t in K*: update the incumbent with c^T (x^t, y^t);
          add the K*-cut  lambda^{tT} s >= 0  to P_{t+1}. It is tight
          at s^t and separates every master point that violated the
          cone at the same y, because conic duality makes
          lambda^{tT} s^t = 0 at optimality.

       b. If it is infeasible, the interior-point method returns a dual
          ray lambda in K* with  lambda^T (b - A (x, y^t)) < 0  for
          every x (a certificate of infeasibility); add the cut
          lambda^T s >= 0, which excludes y^t from the master.

 3. Stop when the incumbent is within eps of LB_t; otherwise t := t + 1
    and return to step 1. In the single-tree form the cuts are added as
    lazy constraints (Section 3.4) inside one MILP branch and bound, so
    the master is solved once, by one tree, instead of being restarted
    after every cut.

Invariant
    every cut lambda^T s >= 0 with lambda in K* is valid for s in K, so
    each master is a relaxation; a repeated y^t has a subproblem whose
    cut makes the master's value at y^t equal to the subproblem's
    value, so the gap at that y is zero (the Duran-Grossmann argument
    of Theorem 3.4.5); finitely many y give finite termination.

Cost
    one conic solve per integer assignment visited; the cuts are dense
    in s but few.

Parallel
    subproblems for distinct integer assignments are independent and
    batch; the master is the sequential part. The initial cuts and the
    K*-cuts for every cone of a product cone are independent.
Algorithm CONIC-OA: cuts from conic duality refine the master

     +------------------------------+
     | 1. master MILP over P_t:     |--> infeasible: stop
     |    s in P_t, y integer       |
     +------------------------------+<---------------------+
        |                                                  |
        |  y^t, and the lower bound LB_t                   |
        v                                                  |
     +------------------------------+                      |
     | 2. conic program in (x, s)   |  adds the cut        |
     |    at y = y^t, solved by an  |  lambda^T s >= 0,    |
     |    interior-point method     |  lambda in K*, to    |
     +------------------------------+  P_{t+1}, from       |
        |                              a. the dual         |
        |  a. optimal: update the         solution         |
        |     incumbent with              lambda^t, or     |
        |     c^T (x^t, y^t)           b. a dual ray       |
        |                                 lambda, which    |
        v                                 excludes y^t     |
  3. stop when the incumbent is                            |
     within eps of LB_t;                                   |
     otherwise t := t + 1  --------------------------------+

  single tree: the cuts enter one branch and bound as lazy
  constraints, and the master is solved once, by one tree

(Why the dual ray matters) The distinctive step is 2b. A gradient-based outer approximation has nothing to linearize when the subproblem is infeasible and must solve a feasibility problem instead. The conic dual ray is a certificate that the interior-point solver produces anyway, and the cut built from it is valid for the whole cone. Lubin, Yamangil, Bent and Vielma obtained the extended formulations that make these cuts strong by writing the model in a disciplined form. Every nonlinearity is a named convex function with a known conic representation, as modelling languages such as CVX require, so every nonlinearity becomes its own cone and every cone is separable. Coey, Lubin and Vielma made the certificate step systematic for every standard cone, including the exponential and semidefinite cones, in the Pajarito and Pavito solvers.M. Lubin, E. Yamangil, R. Bent and J. P. Vielma, "Polyhedral approximation in mixed-integer convex optimization", Mathematical Programming 172 (2018), 139–168; C. Coey, M. Lubin and J. P. Vielma, "Outer approximation with conic certificates for mixed-integer convex problems", Mathematical Programming Computation 12 (2020), 249–293.

Where this is used

The MILP vendors named in this chapter, CPLEX, Gurobi and Xpress, accept second-order cone constraints and solve MISOCPs, and the conic interior-point solvers MOSEK and COPT do so natively. SCIP detects second-order cone structure in quadratic constraints and treats it with its own handlers. Since version 10 the detection covers the bilinear form of case (c), as the changelog line quoted in the sidenote to Proposition 4.8.2 says. Mittelmann's MISOCP benchmark of 10 September 2026 lists COPT with 47 of 47 instances solved, SCIP with 39, MOSEK with 37 and Knitro with 33.H. D. Mittelmann, Benchmarks for Optimization Software, "Mixed-integer SOCP Benchmark", page dated 10 September 2026, plato.asu.edu/ftp/misocp.html, read 5 October 2026: 47 instances selected from CBLIB2014 and from Mittelmann's own set; MOSEK 11.2.3, SCIP 10.0.0 with CPLEX as its LP solver, COPT 8.0.0 and Knitro 16.0.0, run in default mode except for a MIP gap of zero, on an Intel i7-11700K (3.6 GHz, 64 GB) with a time limit of one hour; scaled shifted geometric means of run times 6.41, 7.40, 1 and 19.1 in that order. For the nonconvex problems of this series the cone is the convex half of the model. Tracking error, perspectives and the convex sides of bilinear constraints go into cones, and the rest goes into envelopes and branching.

H. D. Mittelmann, Benchmarks for Optimization Software, "Mixed-integer SOCP Benchmark", page dated 10 September 2026, plato.asu.edu/ftp/misocp.html, read 5 October 2026: one page’s snapshot on one day, in the page’s own terms.

What parallelizes

On a GPU the conic relaxation is the place where first-order methods for cones, Section 7.6, meet the tree. The three-way choice above becomes a choice between a batched first-order conic solve with a safe dual bound and a batched linear program over the polyhedral approximation. The conic subproblems of distinct integer assignments, and the certificate cuts of every cone in a product cone, are independent, and they are the batch. The master remains the sequential part.

Reformulation as algorithm

The perspective of Section 4.3 can be turned into an algorithm. Bertsimas, Cory-Wright and Pauphilet took the indicator structure "\(x_i = 0\) unless \(z_i = 1\)", added a ridge regularizer to the objective, and dualized the coupling between \(x\) and \(z\). The result is a problem in the binaries alone. Its objective is a convex function of \(z\), and one convex solve returns the function value together with a subgradient. Outer approximation therefore applies to it directly, and every cut is a perspective cut (Proposition 4.3.5).D. Bertsimas, R. Cory-Wright and J. Pauphilet, "A unified approach to mixed-integer optimization problems with logical constraints", SIAM Journal on Optimization 31 (2021), 2340–2367; arXiv 1907.02109. This subsection derives that function by Fenchel duality, the duality that pairs a convex function with its conjugate, works the two-variable example by hand, states the algorithm, and reports what it has solved. It closes the chapter's run of formulations because it shows a reformulation acting as an algorithm. The regularizer is a modelling choice that buys convexity, and the algorithm is what the convexity makes possible.

The dualized problem

The problem class is

\[\min_{z \in \mathcal Z \subseteq \{0, 1\}^n,\ x \in \mathbb{R}^n}\ d^\top z + c(x) + \Omega(x) \qquad \text{subject to } x_i = 0 \text{ whenever } z_i = 0,\]

with \(d \in \mathbb{R}^n\) a vector of fixed costs, \(c\) a convex loss, possibly encoding convex constraints through the value \(+\infty\), and \(\mathcal Z\) a set of allowed indicator patterns such as \(\{ z : \sum_i z_i \le k \}\). The function \(\Omega\) is a regularizer: either the big-\(M\) indicator, \(\Omega(x) = 0\) if \(\|x\|_\infty \le M\) and \(+\infty\) otherwise, or the ridge \(\Omega(x) = \tfrac{1}{2\gamma} \|x\|_2^2\). Write \(h(z) = \min_x \{ c(x) + \Omega(x) : x_i = 0 \text{ if } z_i = 0 \}\) for the value of the inner problem at a fixed pattern, so that the problem is \(\min_{z \in \mathcal Z} d^\top z + h(z)\). Write \(c^*(\alpha) = \sup_x \alpha^\top x - c(x)\) for the convex conjugate of the loss.

Theorem 4.9.1 (Bertsimas, Cory-Wright and Pauphilet 2021, Theorem 1; the saddle-point reformulation). Assume that for every \(z \in \mathcal Z\) the inner problem is either infeasible or satisfies strong duality. Then

\[h(z) = \max_{\alpha \in \mathbb{R}^n}\ \Big[ -c^*(\alpha) - \sum_{i=1}^n z_i\, \Omega^\star(\alpha_i) \Big], \qquad \Omega^\star(\beta) = \begin{cases} M |\beta| & \text{big-}M, \\[2pt] \tfrac{\gamma}{2} \beta^2 & \text{ridge}, \end{cases}\]

so \(h\) is a pointwise maximum of functions affine in \(z\) and therefore convex on \([0, 1]^n\). If \(\alpha^\star(z)\) attains the maximum, the vector with entries \(-\Omega^\star(\alpha^\star(z)_i)\) is a subgradient of \(h\) at \(z\).

Proof. Encode the indicator constraint nonlinearly: \(h(z) = \min_{x, v} \{ c(v) + \Omega(x) : v = \operatorname{Diag}(z)\, x \}\). This works because \(v_i = z_i x_i\) is \(x_i\) when \(z_i = 1\) and \(0\) when \(z_i = 0\), and \(\Omega\) is minimized by setting the free coordinates of \(x\) to zero. Dualize the coupling equation with a multiplier \(\alpha\):

\[h(z) = \max_\alpha\ \Big[ \min_v \big( c(v) - \alpha^\top v \big) + \min_x \big( \Omega(x) + \alpha^\top \operatorname{Diag}(z)\, x \big) \Big] = \max_\alpha\ \Big[ -c^*(\alpha) + \sum_i \min_{x_i} \big( \Omega_i(x_i) + z_i \alpha_i x_i \big) \Big],\]

using strong duality for the outer equality and the separability of \(\Omega\) for the inner one. Each inner minimum is \(-\Omega_i^*(-z_i \alpha_i)\), and for the two regularizers \(\Omega_i^*(-z_i \alpha_i) = z_i\, \Omega^\star(\alpha_i)\). For the big-\(M\) box \(\Omega_i^*(\beta) = M |\beta|\) and \(z_i \ge 0\) pulls out. For the ridge \(\Omega_i^*(\beta) = \tfrac{\gamma}{2} \beta^2\) and \(z_i^2 = z_i\) on binaries. A maximum of affine functions of \(z\) is convex, and the affine function active at the maximizer gives a subgradient. ∎

Theorem 4.9.1: from the indicator constraint to a convex h(z)

  h(z) = min_x  c(x) + Omega(x)   s.t.  x_i = 0 whenever z_i = 0
           |
           |  encode: v = Diag(z) x, so v_i = x_i if z_i = 1
           |  and v_i = 0 if z_i = 0
           v
  h(z) = min_{x,v}  c(v) + Omega(x)   s.t.  v = Diag(z) x
           |
           |  dualize v = Diag(z) x with a multiplier alpha
           v
  h(z) = max_alpha  -c*(alpha) - sum_i z_i Omega*(alpha_i)
           |
           |  a pointwise maximum of functions affine in z
           v
  h convex on [0,1]^n; at the maximizer alpha*(z) the vector
  with entries -Omega*(alpha*(z)_i) is a subgradient

  Omega*(beta) = M |beta| (big-M),  (gamma/2) beta^2 (ridge)

(What one evaluation costs, and the Boolean relaxation) For the ridge the inner minimizer is \(x_i = -\gamma z_i \alpha_i\). For a least-squares loss \(c(x) = \tfrac12 \|y - Xx\|^2\) the maximizing \(\alpha\) is \(\nabla c\) at the inner minimizer, that is \(\alpha = -X^\top r\) with \(r\) the residual of the ridge regression on the support of \(z\). One function value and one subgradient therefore cost one linear solve, and only the squares \(\alpha_i^2\) enter the subgradient. The dualized function is also the perspective reformulation of Section 4.3 seen from the dual side. Call the problem with \(\mathcal Z\) replaced by \(\operatorname{conv}(\mathcal Z)\), and \(h\) extended to the cube by the formula of the theorem, the Boolean relaxation. It is a convex program.

Proposition 4.9.2 (the ridge case is the perspective reformulation; BCP 2021). With the ridge regularizer the problem equals

\[\min_{z \in \mathcal Z}\ \min_x\ d^\top z + c(x) + \frac{1}{2\gamma} \sum_i \frac{x_i^2}{z_i},\]

with the convention \(x_i^2 / z_i = 0\) when \(x_i = z_i = 0\) and \(+\infty\) when \(x_i \ne 0 = z_i\). The outer-approximation cuts of \(h\) are perspective cuts in the sense of Proposition 4.3.5, the inequalities (4.3.1). The Boolean relaxation \(\min_{z \in \operatorname{conv}(\mathcal Z)} d^\top z + h(z)\) is the perspective relaxation, a second-order cone program when \(c\) is conic.Bertsimas, Cory-Wright and Pauphilet (2021), the theorem on perspective equivalence; the perspective cuts are those of A. Frangioni and C. Gentile, "Perspective cuts for a class of convex 0–1 mixed integer programs", Mathematical Programming 106 (2006), 225–236. The second-order cone form of the Boolean relaxation and a sufficient condition for its tightness are in D. Bertsimas and R. Cory-Wright, "A scalable algorithm for sparse portfolio selection", INFORMS Journal on Computing 34 (2022), 1489–1511.

The Boolean relaxation is therefore the perspective relaxation of Section 4.3 under another name. What is new is the algorithmic use. The function \(h\) is convex on the cube and its evaluation is one convex solve, which is exactly what outer approximation needs.

The two-variable example

(The instance, by hand) Take \(c(x) = \tfrac12 \|x - a\|^2\) with \(a = (1, 0.5)\), the ridge with \(\gamma = 1\), no fixed costs, and at most one nonzero coordinate. The conjugate of \(c\) is \(c^*(\alpha) = \tfrac12 \|\alpha\|^2 + a^\top \alpha\), attained at \(x = a + \alpha\), so the theorem gives

\[h(z) = \max_\alpha \Big[ -\tfrac12 \|\alpha\|^2 - a^\top \alpha - \tfrac12 \sum_i z_i \alpha_i^2 \Big] = \sum_{i=1}^2 \frac{a_i^2}{2 (1 + z_i)},\]

each coordinate maximized at \(\alpha_i = -a_i / (1 + z_i)\). The four patterns give \(h(0, 0) = 0.625\), \(h(1, 0) = 0.375\), \(h(0, 1) = 0.5625\) and \(h(1, 1) = 0.3125\), the last excluded by the cardinality limit. A direct check at \(z = (1, 0)\): with \(x_2 = 0\) the inner problem minimizes \(\tfrac12 (x_1 - 1)^2 + \tfrac12 \cdot 0.25 + \tfrac12 x_1^2\), whose minimizer \(x_1 = \tfrac12\) gives \(0.125 + 0.125 + 0.125 = 0.375\). The subgradient at \((1, 0)\) has entries \(\partial h / \partial z_i = -a_i^2 / (2 (1 + z_i)^2)\), that is \((-\tfrac18, -\tfrac18)\), and the cut

\[\eta \;\ge\; 0.375 - \tfrac18 (z_1 - 1) - \tfrac18 z_2\]

evaluates to \(0.5\) at \((0, 0)\) and to \(0.375\) at \((0, 1)\), both at least the incumbent value \(0.375\). One cut, from one inner solve, prunes every other pattern. The Boolean relaxation is also exact here. Since \(h\) decreases in each coordinate, the relaxation sits on \(z_1 + z_2 = 1\), and the minimum of \(a_1^2 / (2(1 + t)) + a_2^2 / (2(2 - t))\) over \(t \in [0, 1]\) is at \(t = 1\), the integer point. The script checks each of these numbers.

The two-variable example: h on the edge z₁ + z₂ = 1, where the Boolean relaxation sits, and at the four patterns, with (1, 1) excluded by z₁ + z₂ ≤ 1, against the one cut from the inner solve at (1, 0). At the text’s instance, a = (1, 0.5) and γ = 1, the subgradient is (−⅛, −⅛), the cut is η ≥ 0.375 − ⅛(z₁ − 1) − ⅛z₂, and the relaxation’s minimum, 0.3750, is at the integer point (1, 0).
# bcp_dual.py -- the dualized indicator problem of Bertsimas, Cory-Wright
# and Pauphilet by hand.
#
#   min over z in {0,1}^2 with z1 + z2 <= 1, x with x_i = 0 when z_i = 0:
#       c(x) + (1/(2 gamma)) ||x||^2,
#   c(x) = (1/2) ||x - a||^2,  a = (1, 0.5),  gamma = 1.
# Then h(z) = sum_i a_i^2 / (2 (1 + gamma z_i)).

import itertools

import numpy as np

a = np.array([1.0, 0.5])
gamma = 1.0
k = 1

def h(z):
    return np.sum(a**2 / (2 * (1 + gamma * z)))

def grad(z):
    return -gamma * a**2 / (2 * (1 + gamma * z)**2)

def inner(z):
    """Check: the inner minimization over x done directly."""
    # minimizer of (1/2)(x-a)^2 + x^2/(2 gamma) on the support
    x = np.where(z > 0, a / (1 + 1/gamma), 0.0)
    return 0.5 * np.sum((x - a)**2) + np.sum(x**2) / (2 * gamma), x

Z = [np.array(z, float) for z in itertools.product((0, 1), repeat=2)
     if sum(z) <= k]
print("pattern z    h(z)     direct inner minimization   x")
for z in Z:
    val, x = inner(z)
    print(f"{tuple(int(t) for t in z)}     {h(z):.4f}   {val:.4f}"
          f"                     {x}")

z1 = np.array([1.0, 0.0])
g1 = grad(z1)
print(f"\nsubgradient at z = (1, 0): {g1}   (= -a_i^2 / (2 (1 + z_i)^2))")
print("cut  eta >= h(z1) + g1 . (z - z1), evaluated at every pattern:")
for z in Z:
    print(f"   z = {tuple(int(t) for t in z)}: "
          f"{h(z1) + g1 @ (z - z1):.4f}   (true h = {h(z):.4f})")
inc = h(z1)
print(f"incumbent h(1, 0) = {inc:.4f}; every other pattern's cut value is")
print("   >= the incumbent, so one cut proves optimality")

# the Boolean relaxation: minimize h over the simplex-capped box,
# z in [0,1]^2, z1 + z2 <= 1 (h is decreasing, so z1 + z2 = 1)
t = np.linspace(0, 1, 100001)
vals = np.array([h(np.array([s, 1 - s])) for s in t])
i = vals.argmin()
print(f"Boolean (perspective) relaxation: min h over z1 + z2 = 1 is "
      f"{vals[i]:.4f}")
print(f"   at z = ({t[i]:.3f}, {1 - t[i]:.3f}): "
      "equal to the integer optimum")

The script prints \(0.6250\), \(0.5625\) and \(0.3750\) for the three admissible patterns, with the direct inner minimization agreeing at each. It prints the subgradient \((-0.125, -0.125)\), the cut values \(0.5000\), \(0.3750\) and \(0.3750\), and the relaxation value \(0.3750\) at \(z = (1, 0)\). Each evaluation of \(h\) is a closed form here. In general it is one ridge regression on the current support, and the evaluations for different supports are independent.

The algorithm

Algorithm 4.9.3 (outer approximation on the dualized objective; BCP 2021).

Algorithm BCP-OUTER-APPROXIMATION
          (ridge regularizer; the big-M case is the same with
           M|alpha_i| in place of (gamma/2) alpha_i^2)

Input   the pattern set Z (for instance sum_i z_i <= k), fixed costs d,
        a convex loss c with an evaluable conjugate or an inner solver,
        gamma > 0, a tolerance eps.
Output  an optimal pattern z and the corresponding x, with a proof of
        optimality within eps.

 0. Root bound and warm start.

       0a. Solve the Boolean relaxation
              min_{z in conv(Z)} d^T z + h(z)
           by a cutting-plane method (Kelley's method, Proposition 3.7.9:
           minimize the maximum of the subgradient cuts collected so far
           over conv(Z), add the cut at the minimizer, repeat; the in-out
           rule stabilizes it by taking each new cut at a point between
           that minimizer and the best point found so far) or by
           projected subgradient steps: h(z) and
              grad_i h(z) = -(gamma/2) alpha*(z)_i^2
           cost one inner solve. Its value is a lower bound (it equals
           the perspective relaxation, Proposition 4.9.2).

       0b. Round the fractional z* at random (Proposition 4.9.4) and
           improve by swaps; this gives z^1 and an incumbent.

 1. t := 1. Evaluate h(z^1) and a subgradient g^1; if the inner problem
    is infeasible add instead the no-good cut
       sum_i z^1_i (1 - z_i) + sum_i (1 - z^1_i) z_i >= 1,
    which excludes z^1 only.

 2. Master (one branch-and-bound tree, the cuts added as lazy
    constraints, Section 3.4):
       min_{z in Z, eta}  d^T z + eta
       s.t.  eta >= h(z^s) + g^{sT} (z - z^s),  s = 1..t.
    Each time the tree reaches an integer candidate z^{t+1}: evaluate h
    and a subgradient there, add the cut, update the incumbent if
    d^T z^{t+1} + h(z^{t+1}) is better, and set t := t + 1.

 3. Stop when the master's lower bound is within eps of the incumbent.

Invariant
    h is convex on [0,1]^n (Theorem 4.9.1), so every cut underestimates
    h on Z and every master is a valid relaxation; a pattern that recurs
    has a cut tight at it, so the master's value there equals its true
    value, and finiteness of Z gives finite termination (the argument
    of Duran and Grossmann, and Fletcher and Leyffer, Theorem 3.4.5).

Cost
    one inner convex solve per cut. For least squares with ridge, h and
    its subgradient cost one solve with the k x k Gram matrix of the
    selected columns, or with the m x m matrix I + gamma X_S X_S^T; the
    cuts are dense in z.

Parallel
    the inner solves for a batch of candidate patterns are independent
    (batched Cholesky factorizations); the Boolean relaxation's gradient
    is a dense matrix-vector product; the master MILP is the sequential
    part.
Algorithm BCP-OUTER-APPROXIMATION: bounds and cuts

  0a. Boolean relaxation over conv(Z)  -->  a lower bound, and z*
          |
          v
  0b. random rounding of z*, swaps     -->  z^1, an incumbent
          |
          v
   1. inner solve at z^1               -->  h(z^1), g^1: the first
          |                                 cut (or a no-good cut)
          v
   2. master: one tree over z in Z, min d^T z + eta,
      eta >= every cut so far  <-----------------------------+
          |                                                  |
          |  an integer candidate z^{t+1}                    |
          v                                                  |
      inner solve: h(z^{t+1}), g^{t+1}  -->  one more cut ---+
                                             (and the incumbent,
                                             if better)

   3. stop when the master's lower bound is within eps of the
      incumbent

Proposition 4.9.4 (smoothness and rounding; BCP 2021). For \(z, z' \in \operatorname{conv}(\mathcal Z)\) and the ridge regularizer, \(h(z') - h(z) \le \tfrac{\gamma}{2} \sum_i (z_i - z'_i)\, \alpha^\star(z')_i^2\). Suppose \(z^\star\) solves the Boolean relaxation, \(\mathcal R\) is the set of its fractional coordinates and \(|\alpha^\star(\cdot)_i| \le L\). Rounding each fractional coordinate independently to \(1\) with probability \(z^\star_i\) then yields a pattern \(z\) with \(0 \le h(z) - h(z^\star) \le \varepsilon\) with probability at least \(1 - |\mathcal R| \exp(-\varepsilon^2 / \kappa)\), where \(\kappa = \tfrac12 \gamma^2 L^4 |\mathcal R|^2\).Bertsimas, Cory-Wright and Pauphilet (2021), the proposition on smoothness and the theorem on randomized rounding, with the constant \(\kappa\) as printed there for the ridge case (the big-\(M\) case has \(\kappa = 2 M^2 L^2 |\mathcal R|^2\)).

(Why rounding works) The bound is a concentration inequality. Each fractional coordinate is rounded independently, and the smoothness inequality makes \(h\) Lipschitz in each coordinate with constant \(\tfrac{\gamma}{2} L^2\), so the sum of \(|\mathcal R|\) bounded independent changes is unlikely to exceed \(\varepsilon\). This is what makes step 0b more than a heuristic. When the relaxation is nearly integral, as in the two-variable example and in the toy regression below, a rounded pattern is nearly optimal with high probability. The cut it generates is then nearly the final one.

Proposition 4.9.4’s bound: rounding each fractional coordinate of the Boolean relaxation’s solution z* independently to 1 with probability z*ᵢ gives a pattern z with 0 ≤ h(z) − h(z*) ≤ ε with probability at least 1 − |R| exp(−ε²/κ), where κ = ½γ²L⁴|R|²; the values of |R|, γ and L are illustrative.

Where this is used

(What the algorithm has solved) The cardinality-constrained mean–variance problem is Bienstock's problem of 1996, described in Section 4.4. As Bertsimas and Cory-Wright present it, his method branched on subsets of the universe with surrogate constraints and no binaries, and his paper fixed the test problem for the field.D. Bienstock, "Computational study of a family of mixed-integer quadratic programming problems", Mathematical Programming 74 (1996), 121–140. The original could not be read for this series, so the method is described as Bertsimas and Cory-Wright (2022), Section 1.3, present it, and its problem sizes are unverified, as Section 4.4 says. The table of Section 4.4 lists the methods that followed and the largest instances they certified, with 400 assets as the ceiling. The 400 is the figure the BCP abstract names as the previous limit of certifiably optimal methods. Bertsimas and Cory-Wright applied Algorithm 4.9.3 to instances built from the S&P 500, the Russell 1000 and the Wilshire 5000 with up to about 3,200 securities. The master was CPLEX 12.8 and the continuous subproblems were solved by Mosek 9.0. The cardinalities were \(k \in \{10, 50, 100, 200\}\), with two values of the regularization, under a time limit of 600 seconds on one thread. On the 200-to-400-asset instances of Frangioni and Gentile they used \(k \in \{6, 8, 10, 12\}\) and the unconstrained case.Bertsimas and Cory-Wright (2022), Tables 7 to 9 and Appendix B of the arXiv version (1811.00138, v5), read on 5 October 2026; the journal version occupies 23 pages against 47 for the preprint, so the table numbers are those of the preprint. The BCP (2021) abstract gives the comparison "3,200 securities against 400 for earlier certifiably optimal methods". The same framework reports network design with hundreds of nodes and sparse regression with up to 100,000 covariates.The 100,000-covariate scale is named in the abstract of BCP (2021); the experiments at that scale are those of D. Bertsimas, J. Pauphilet and B. Van Parys, "Sparse regression: scalable algorithms and empirical performance", Statistical Science 35 (2020), 555–578, which BCP cite at that point, while the BCP paper's own regression tables use 20,000 covariates and its classification tables 10,000. The exact scalable algorithm for sparse regression with the ridge is D. Bertsimas and B. Van Parys, "Sparse high-dimensional regression: exact scalable algorithms and phase transitions", Annals of Statistics 48 (2020); arXiv 1709.10029.

(Two neighbours, and the caveat on the ridge) Two developments sit next to this one. Hazimeh, Mazumder and Saab's L0BnB solves the same regularized regression problem by a branch and bound whose node relaxation is the perspective relaxation. They solve it not by a conic interior-point method but by coordinate descent with active sets and warm starts, for problems with \(p \sim 10^7\) features. Their abstract reports speed-ups of at least 5000 times over what it calls state-of-the-art exact methods.Hazimeh, Mazumder and Saab (2022), cited in Section 4.3; the heuristic side is H. Hazimeh and R. Mazumder, "Fast best subset selection: coordinate descent and local combinatorial optimization algorithms", Operations Research 68 (2020), 1517–1537. Atamtürk and Gómez's rank-one convexification, Section 4.5, strengthens the perspective relaxation itself when the quadratic is not separable, which is the regime of correlated assets.A. Atamtürk and A. Gómez, "Rank-one convexification for sparse regression", Journal of Machine Learning Research 26 (2025), paper 35, 1–50. One caveat travels with every member of this family: the ridge changes the problem. For weights with \(\|x\|_2 \le 1\), as long-only weights summing to one are, an optimal solution of the regularized problem is a \(1/(2\gamma)\)-optimal solution of the unregularized one. The reason is that the ridge adds at most \(\|x\|_2^2/(2\gamma)\) to any objective value. So \(\gamma\) is a modelling parameter to be chosen or cross-validated, not a solver tolerance.Bertsimas and Cory-Wright (2022), Section 1 and the proposition on sensitivity to the regularization.

The two-variable example as the ridge γ varies: the integer optimum, the Boolean relaxation and the one cut’s value at the competing pattern, with the threshold 1/a₂ − 1 derived from the section’s formula for h. Below, 1/(2γ): for weights with ‖x‖₂ ≤ 1 an optimal solution of the regularized problem is a 1/(2γ)-optimal solution of the unregularized one. The ridge changes the problem, and γ is a modelling parameter to be chosen or cross-validated, not a solver tolerance.

What parallelizes

The expensive parts of Algorithm 4.9.3 are dense linear algebra and sampling: the inner solves for a batch of candidate patterns, the gradient of the Boolean relaxation, the random roundings and the swap searches. The master MILP is the sequential bottleneck. L0BnB's choice of a first-order node relaxation points at the alternative: replace the master by a branch and bound over \(z\) whose node bounds are perspective relaxations solved by batched first-order methods. The kernel such a design runs per candidate pattern is the evaluation of \(h\) and its subgradient. The C++23 listing below writes that kernel for a least-squares loss with a ridge. It uses the Woodbury identity so that a support of size \(k\) costs one \(k \times k\) Cholesky factorization, and it distributes a batch of supports over threads through a shared atomic counter. The loop body is what a CUDA kernel would run per thread, with a batched Cholesky in place of the scalar one.

// bcp_batch.cpp -- evaluate the dualized objective h(z) and its
// subgradient for a batch of supports.
//
// Ridge-regularized least squares: c(x) = (1/2)||y - X x||^2, penalty
// ||x||^2 / (2 gamma), support S = {i : z_i = 1}. Woodbury:
//   r = (I + gamma X_S X_S^T)^{-1} y
//     = y - gamma X_S (I_k + gamma X_S^T X_S)^{-1} X_S^T y,
//   h(z) = (1/2) y^T r,   dh/dz_i = -(gamma / 2) (X_i^T r)^2.
// One k x k Cholesky per support.

#include <atomic>
#include <cmath>
#include <cstdio>
#include <span>
#include <thread>
#include <vector>

// X row-major, m x n
struct Data {
    int m, n;
    std::vector<double> X;
    std::vector<double> y;
    double gamma;
};

// Solve (I + gamma G) w = v for a symmetric positive definite k x k Gram
// matrix G, by Cholesky. A holds I + gamma G and is overwritten by its
// factor L; w holds v and is overwritten by the solution.
static void spd_solve(std::vector<double>& A, std::vector<double>& w,
                      int k) {
    // A := L, lower triangle, in place
    for (int j = 0; j < k; ++j) {
        double d = A[j * k + j];
        for (int p = 0; p < j; ++p)
            d -= A[j * k + p] * A[j * k + p];
        A[j * k + j] = std::sqrt(d);
        for (int i = j + 1; i < k; ++i) {
            double s = A[i * k + j];
            for (int p = 0; p < j; ++p)
                s -= A[i * k + p] * A[j * k + p];
            A[i * k + j] = s / A[j * k + j];
        }
    }

    // forward substitution with L
    for (int i = 0; i < k; ++i) {
        double s = w[i];
        for (int p = 0; p < i; ++p)
            s -= A[i * k + p] * w[p];
        w[i] = s / A[i * k + i];
    }

    // back substitution with L^T
    for (int i = k - 1; i >= 0; --i) {
        double s = w[i];
        for (int p = i + 1; p < k; ++p)
            s -= A[p * k + i] * w[p];
        w[i] = s / A[i * k + i];
    }
}

// h(z) and its subgradient for one support; this body is the kernel a
// GPU would run per thread.
static double evaluate(const Data& D, std::span<const int> S,
                       std::span<double> grad) {
    const int k = int(S.size()), m = D.m, n = D.n;
    std::vector<double> G(k * k, 0.0), v(k, 0.0);

    // the Gram matrix G = X_S^T X_S of the support, stored as
    // I + gamma G, and v = X_S^T y
    for (int a = 0; a < k; ++a) {
        for (int b = 0; b <= a; ++b) {
            double s = 0;
            for (int r = 0; r < m; ++r)
                s += D.X[r * n + S[a]] * D.X[r * n + S[b]];
            G[a * k + b] = G[b * k + a] = D.gamma * s;
        }
        G[a * k + a] += 1.0;
        for (int r = 0; r < m; ++r)
            v[a] += D.X[r * n + S[a]] * D.y[r];
    }

    // v := (I + gamma G)^{-1} X_S^T y
    if (k > 0)
        spd_solve(G, v, k);

    // r = y - gamma X_S v
    std::vector<double> res(D.y);
    for (int r = 0; r < m; ++r)
        for (int a = 0; a < k; ++a)
            res[r] -= D.gamma * D.X[r * n + S[a]] * v[a];

    double h = 0;
    for (int r = 0; r < m; ++r)
        h += D.y[r] * res[r];

    for (int i = 0; i < n; ++i) {
        double t = 0;
        for (int r = 0; r < m; ++r)
            t += D.X[r * n + i] * res[r];
        grad[i] = -0.5 * D.gamma * t * t;
    }
    return 0.5 * h;
}

int main() {
    // The worked example: X = I_2, y = a = (1, 0.5), gamma = 1, so
    // h(z) = sum_i a_i^2 / (2 (1 + z_i)).
    Data D{2, 2, {1, 0, 0, 1}, {1.0, 0.5}, 1.0};
    std::vector<std::vector<int>> supports = {{}, {0}, {1}, {0, 1}};
    std::vector<double> h(supports.size()), grad(supports.size() * D.n);

    // a shared counter hands supports to worker threads
    std::atomic<std::size_t> next{0};
    {
        std::vector<std::jthread> pool;
        for (unsigned t = 0; t < 2; ++t)
            pool.emplace_back([&] {
                for (std::size_t b;
                     (b = next.fetch_add(1)) < supports.size();) {
                    std::span<double> g(grad.data() + b * D.n, D.n);
                    h[b] = evaluate(D, supports[b], g);
                }
            });
    }   // jthreads join here

    for (std::size_t b = 0; b < supports.size(); ++b) {
        std::printf("support {");
        for (int i : supports[b])
            std::printf(" %d", i + 1);
        std::printf(" }:  h = %.4f   subgradient = (%.4f, %.4f)\n",
                    h[b], grad[b * D.n], grad[b * D.n + 1]);
    }
}

On the two-variable example the program prints \(h = 0.6250\) for the empty support, \(0.3750\) for \(\{1\}\) with subgradient \((-0.1250, -0.1250)\), \(0.5625\) for \(\{2\}\) and \(0.3125\) for \(\{1, 2\}\), the values of the Python script. Its cost per support is \(O(k^2 m + k^3)\) for the Gram matrix and its factorization plus \(O(nm)\) for the subgradient. The supports are independent, so a batch of \(B\) of them is \(B\) identical kernels with no communication. On a device the Gram matrices of a batch are one batched matrix product and one batched Cholesky call. That is the shape of the primal side of a GPU solver for this class, and Section 7.6 takes it up.

bcp_batch.cpp: four supports, two threads, one atomic counter

  b                  0          1          2          3
  support           { }        {1}        {2}       {1, 2}
                      \          \        /          /
               next.fetch_add(1) hands out b = 0, 1, 2, 3, one
               at a time; a thread stops when it draws b >= 4
                       /                       \
                 jthread 0                 jthread 1
                       \                       /
               evaluate(D, support b): with G = X_S^T X_S,
               v = (I + gamma G)^{-1} X_S^T y by one k x k
               Cholesky, r = y - gamma X_S v, h = (1/2) y^T r,
               grad_i = -(gamma/2) (X_i^T r)^2
                                   |
                                   v
  h[b]            0.6250     0.3750     0.5625     0.3125
  grad[2b]       -0.5000    -0.1250    -0.5000    -0.1250
  grad[2b + 1]   -0.1250    -0.1250    -0.0312    -0.0312

  every b writes only h[b] and its own two entries of grad

What can be written at all

The chapter has treated formulation as a design. The same set of feasible points can be described by many systems of constraints, and the solver's effort depends on which. This subsection asks the prior question: which sets can be described at all with linear constraints and integer variables, and which with convex constraints and integer variables. It then says what the answer means for the models of this series. The theory is Jeroslow and Lowe's for the linear case and Lubin, Vielma and Zadik's for the convex case. The one proof given in full, the midpoint lemma, is also the lower bound behind the logarithmic formulations of Theorem 4.6.4.

Definition 4.10.1 (MILP- and MICP-representability). Let \(Q \subseteq \mathbb{R}^{n + p + q}\) be a mixed-integer formulation of \(S \subseteq \mathbb{R}^n\) in the sense of Definition 2.1.6, as extended in Section 4.1: a closed convex set with \(x \in S\) if and only if \((x, w, z) \in Q\) for some \((w, z) \in \mathbb{R}^p \times \mathbb{Z}^q\). The set \(S\) is MILP-representable (rational MILP-representable) if some such \(Q\) is a (rational) polyhedron. It is MICP-representable, mixed-integer convex representable, if some such \(Q\) is a closed convex set. The formulation is binary if the integer variables are bounded, which is to say they can be taken in \(\{0, 1\}^q\), and pure if \(p = 0\), with no continuous auxiliaries.R. G. Jeroslow and J. K. Lowe, "Modelling with integer variables", Mathematical Programming Studies 22 (1984), 167–184; R. G. Jeroslow, "Representability in mixed integer programming, I: characterization results", Discrete Applied Mathematics 17 (1987), 223–243; M. Lubin, J. P. Vielma and I. Zadik, "Mixed-integer convex representability", Mathematics of Operations Research 47 (2022), 720–749, Section 2 for the definitions. J. P. Vielma, "Mixed integer linear programming formulation techniques", SIAM Review 57 (2015), 3–57, Section 5, is the expository account of the linear case.

Theorem 4.10.2 (Jeroslow and Lowe, 1984). A set \(S \subseteq \mathbb{R}^n\) is rational MILP-representable if and only if there are rational polytopes \(P_1, \dots, P_k\) and a finite set \(U \subset \mathbb{Z}^n\) with

\[S = \bigcup_{i=1}^k P_i + \operatorname{intcone}(U), \qquad \operatorname{intcone}(U) = \Big\{ \sum_{u \in U} \mu_u u : \mu \in \mathbb{Z}_{\ge 0}^U \Big\}.\]

In particular a bounded set is MILP-representable if and only if it is a finite union of polytopes. A finite union of polyhedra has a binary MILP formulation if and only if the polyhedra share a recession cone.Jeroslow and Lowe (1984), as restated in Lubin, Vielma and Zadik (2022) and in Vielma (2015), Section 5. Rationality matters: with an irrational recession direction the characterization fails, an exercise of M. Conforti, G. Cornuéjols and G. Zambelli, Integer Programming (Springer, 2014), cited by Lubin, Vielma and Zadik.

Proof sketch. Sufficiency for a finite union of polytopes is Balas' theorem (Theorem 4.2.4). The hull formulation with one copy of the variables per polytope, one binary per polytope and the scaled constraints \(A^i \nu^i \le b^i \lambda_i\) is a formulation. A common recession cone is what the hull theorem needs for unbounded polyhedra. The integer cone is added with one general integer variable per generator. Necessity: project the polyhedron \(Q\) onto the integer coordinates and decompose the rational polyhedron into a polytope plus a rational cone. The finitely many integer points of the polytope part give the \(P_i\), and the integral generators of the cone give \(U\). This is the argument that proves Meyer's theorem (Theorem 2.1.4) on the hull of the integer points of a rational polyhedron. ∎

The theorem settles the linear question: everything in a MILP is a finite union of polyhedra, possibly repeated along a lattice. The convex question is subtler, because convex sets can be curved, and a parabola is the example to keep in mind.

Theorem 4.10.2’s shape, illustrated: a polytope P repeated at the nonnegative integer multiples of a generator, P + intcone(U).

Theorem 4.10.3 (Lubin, Vielma and Zadik, 2022). (a) A set \(S\) has a pure binary MICP formulation, with binary integer variables and no continuous auxiliaries, if and only if \(S = \bigcup_{i=1}^d S_i\) for nonempty closed convex sets \(S_i\) with a common recession cone. The set \(Q = \operatorname{conv}\big(\bigcup_i S_i \times \{e_i\}\big) \subseteq \mathbb{R}^{n + d}\), the Cayley embedding, is then a formulation. (b) With continuous auxiliaries the class is strictly larger. The set \(S = \{0\} \cup (-\infty, -1]\), whose two pieces have different recession cones, is binary MICP-representable through \(Q = \{ (x, y, z) : x^2 \le yz,\ x \le -z,\ 0 \le z \le 1,\ y \ge 0 \}\). (c) Midpoint lemma. If \(S\) contains \(w\) points no two of which have their midpoint in \(S\), then every MICP formulation of \(S\) uses at least \(\lceil \log_2 w \rceil\) integer variables. If \(S\) contains infinitely many such points, \(S\) is not MICP-representable. (d) Under a rationality assumption on the formulation, a MICP-representable set is either a finite union of compact convex sets or has a periodic structure along the integer directions.Lubin, Vielma and Zadik (2022): the proposition on pure binary representability, the example of part (b), the midpoint lemma with its corollary, and the structural theorem of the abstract, paraphrased here. The Cayley embedding as the geometric form of Balas' theorem is J. P. Vielma, "Small and strong formulations for unions of convex sets from the Cayley embedding", Mathematical Programming 177 (2019), 21–53.

Proof of (c). Let \(Q \subseteq \mathbb{R}^{n + p + q}\) be a formulation with \(q\) integer variables and let \(x^1, \dots, x^w \in S\) be points no two of whose midpoints lie in \(S\). Each has a lift \((x^i, w^i, z^i) \in Q\) with \(z^i \in \mathbb{Z}^q\). If \(w > 2^q\), two of the vectors \(z^i\) agree coordinatewise modulo \(2\), say \(z^i \equiv z^j \pmod 2\) with \(i \ne j\). Then \(\tfrac12 (z^i + z^j)\) is an integer vector. By convexity of \(Q\) the midpoint \(\tfrac12 \big( (x^i, w^i, z^i) + (x^j, w^j, z^j) \big)\) lies in \(Q\) with an integral last block, so \(\tfrac12 (x^i + x^j) \in S\), a contradiction. Hence \(w \le 2^q\), that is \(q \ge \log_2 w\), and an infinite family of such points rules out every finite \(q\). ∎

(The parity argument, and the point-and-ray check) The proof is a parity argument: a convex set in which two lifts have the same parity pattern contains an integral midpoint. The two-variable check of part (b) is worth doing once. For \(z = 0\) the constraint \(x^2 \le y \cdot 0\) forces \(x = 0\), which is the point. For \(z = 1\) the constraints are \(x \le -1\) and \(y \ge x^2\), which is the ray with a free auxiliary \(y\). The rotated cone \(x^2 \le yz\) is Proposition 4.8.2(a). It glues a point to a ray with a different recession cone, something no finite union of polyhedra with binaries can do by Theorem 4.10.2.

Theorem 4.10.3(b): one binary glues a point to a ray. The slice of Q at the z the reader sets, its shadow on the x-axis, and the set S against what the relaxation projects to. The point and the ray have different recession cones, which no finite union of polyhedra with binaries can glue (Theorem 4.10.2); the rotated cone x² ≤ yz does.

(Two jobs of the midpoint lemma) The midpoint lemma does two jobs in this chapter. First, it is the lower bound behind the logarithmic formulations. The \(K\) midpoints of the pieces of a strictly convex or strictly concave interpolant are \(K\) points no two of whose midpoints lie on the graph, so Theorem 4.6.4's \(\lceil \log_2 K \rceil\) binaries cannot be improved. Second, it separates the convex world from the nonconvex one. Consider the points \((z, z^2)\) for \(z \in \mathbb{Z}\). The midpoint of \((a, a^2)\) and \((b, b^2)\) is \(\big( \tfrac{a + b}{2}, \tfrac{a^2 + b^2}{2} \big)\), and

\[\frac{a^2 + b^2}{2} - \Big( \frac{a + b}{2} \Big)^2 = \frac{(a - b)^2}{4} > 0 \qquad (a \ne b),\]

so no midpoint lies on the parabola. The points \((1, 1)\) and \((3, 9)\) have the midpoint \((2, 5)\), which is not \((2, 4)\), and \((2, 4)\) and \((4, 16)\) have the midpoint \((3, 10)\), which is not \((3, 9)\). There are infinitely many such points, so by part (c) the set \(\{ (z, z^2) : z \in \mathbb{Z} \}\) is not MICP-representable. It is a one-line nonconvex MINLP: \(y = x^2\) with \(x\) integer. Lubin, Vielma and Zadik show by the same lemma that the set of prime numbers and the set of rank-one matrices are not MICP-representable.Lubin, Vielma and Zadik (2022), the examples following the midpoint lemma. The nonconvex side of Rockafellar's watershed, quoted in Section 1.2, is therefore not a modelling inconvenience that a cleverer convex formulation removes. There are sets that only a nonconvex constraint can write, and Jeroslow's undecidability theorem (Theorem 1.5.4) is the price of that expressiveness when the integers are unbounded. With a bound \(|x| \le N\) the parabola's integer points are a finite set and Theorem 4.10.2 applies, with at least \(\lceil \log_2 (2N + 1) \rceil\) binaries by part (c).

Integer points of y = x² and their midpoints (Theorem 4.10.3(c)): for a ≠ b the midpoint of (a, a²) and (b, b²) lies (a − b)²/4 > 0 above the parabola, so no midpoint is in the set. The two cases are the text’s midpoints, (2, 5) of (1, 1) and (3, 9) and (3, 10) of (2, 4) and (4, 16), each 1 above the parabola. With |x| ≤ N the finite set needs at least ⌈log₂(2N + 1)⌉ binaries.

Where this is used

(What the tax model may express) What a tax model may express follows from the two theorems. Every variable in the problem of Section 9 is bounded, so Theorem 4.10.2 says that any finite union of polytopes is available with binaries. The structures of that section are all of that form. A whole-lot constraint is a set of finitely many points. The tax on a lot as a function of the shares sold from it is piecewise linear, with a change of slope on the day the lot has been held for a year. (The holding-period rule, like the wash-sale rule below, is stated in Section 9 with its sources and its unverified flag.) It is a union of segments written with the λ model of Section 4.6. A cardinality limit is a union of coordinate subspaces cut by a box. The wash-sale clause \(b \cdot s = 0\) on a box is the union of the two polyhedra \(\{ b = 0 \}\) and \(\{ s = 0 \}\), whose hull is the triangle of Section 4.2 and which one binary represents exactly. The tracking error is a convex quadratic, Proposition 4.8.2(e). So the whole model is a MISOCP, the class of Section 4.8, and the design choices of this chapter decide how tight its relaxation is. The part of the problem that is genuinely nonconvex in the continuous variables, the bilinear wash-sale product, is representable because its disjunctive form is. A model that kept it as a product would hand a MILP-representable set to a global solver as a nonconvex curve. That is the formulation decision the chapter has been about.

The model’s clause b · s = 0 on a box: the two segments {b = 0} and {s = 0}, a union of two polyhedra, and their hull, the triangle b/B + s/S ≤ 1 of Section 4.2, with a point the reader moves. As the product b · s = 0 it is a nonconvex curve for a global solver; as the disjunction {b = 0} or {s = 0}, one binary represents it exactly. This is the model’s clause (no name bought and sold in one list), not the wash-sale rule itself; Section 9 states the rule with its sources, from memory, and it remains unverified. B and S are illustrative share counts.

What parallelizes

Representability decides what a device receives. A MILP-representable structure reaches a batched bounding kernel as a union of polyhedra with binaries, whose per-piece relaxations are independent, while a set that only a nonconvex constraint can write reaches it as a curve that spatial branching must split. The decision is made in the model, before any kernel runs.

The ladder, in one table

The chapter began with the observation that the relaxation is a design and has since named a dozen designs. The table collects them. It extends the ladder of Section 4.1 with the envelope rungs of Section 2.4 and the lifted rungs of Section 4.7, and it fills in the numbers from this series: what each relaxation replaces, what is solved to evaluate it, when it is the convex hull of the set it relaxes, and what it buys on the running examples. Gaps in the last column are relative to the true optimum, the convention of Section 2.1. The statistic of the relaxation figure of Section 2.1 divides by the bound instead and reads \(10.9\) and \(12.0\) per cent, as Section 4.1 notes. The costs are the class of the problem solved per node, not running times. LP means a linear program solvable by the simplex method with a warm start. SOCP and SDP mean conic programs solved by an interior-point method from a cold start. MILP means a problem that itself needs a tree.

Two conventions for one gap: the overstatement of a bound over the integer optimum 4.95, divided by the bound (the statistic of the relaxation figure of Section 2.1) and divided by the optimum (the table's convention, and Gurobi's and CPLEX's).
relaxationreplacessolved per nodethe hull whennumbers from this series
drop integrality (LP relaxation)\(y \in \mathbb Z\) by \(y \in [l, u]\)LPthe formulation is sharp (integer hull)R1: 5.56 vs 4.95 (12.2%); weak polygon 5.63 (13.7%)
secant / chord of a concave term\(f\) by its chord on \([l, u]\)LPone term alone; exact at the endsR3 root bound 1.3125 vs 2.0; the quartic's band 2.56 deep
αBB underestimator\(f\) by \(f + \sum \alpha_i (x_i-l_i)(x_i-u_i)\)convex NLPnot in general; equal to \(f\) exactly where every coordinate with \(\alpha_i > 0\) is at a bound, in particular at the vertices (Theorem 2.4.19(a))\(x^3\) on \([-1, 1]\): \(-3\) vs \(-1\); exact after one split
McCormick planes\(xy\) by four planesLPone product on a box (exact on the edges)R2: 0.40 vs 0.18; R4: 500 vs 400 (pq = p)
piecewise McCormick, \(k\) piecesthe box by \(k \times k\) cellsMILP (\(k\) or \(\log k\) binaries per axis); LP relaxation = McCormick on the boxunion of the per-cell hulls (as a MILP)R2 MILP optimum 0.233, 0.193, 0.181 for \(k = 2, 3, 16\); LP 0.40; band \(1/(2k^2)\)
big-Ma disjunction by one row per constraintLPtwo polygons in the plane (Proposition 4.2.7), every facet at its exact support constant; not in generalSection 4.2's two polytopes 1.5 vs 1 (50%); fig-disjunction 1.444 vs 0.947 (52.5%)
hull (Balas) formulationa disjunction by copies and weightsLP, \(k\) copiesfor nonempty polyhedra, the closure of the hull; the hull itself when they share a recession cone, in particular for polytopes (Theorem 4.2.4)Section 4.2's two polytopes: 1 at the root
perspective\(z f(x/z)\) for an on/off setSOCP, or LP with perspective cutsthe on/off set; separable sumsfig-perspective at \(x = 0.6\): 1.47 vs 1.86 (big-M 0.58); R5: 11.71% vs 12.45% (big-M 11.12%)
diagonal split + perspective\(\Sigma\) by \((\Sigma - D) + D\); \(D\) uniform, factor or SDP-chosenSOCP (an SDP first if \(D\) is SDP-chosen)not in general; the split decides how closetwo assets, \(\rho = 0.8\), \(D = \delta I\): \(0.9 \to 0.95 \to 1.0\) at \(\delta = \lambda_{\min}\)
rank-one and \(2 \times 2\) hulls (Theorems 4.5.3, 4.5.5)the same, in an extended spaceSOCP or SDP, \(O(n^2)\) variablesnot in general; tighter than the perspective–
full hull (Theorem 4.5.6)the sameSDP, exponentially many verticesexactly (the closed hull), for \(\Sigma \succ 0\)–
λ / SOS2 modela PWL graph by its weightsLP (SOS2 by branching)the graph's own breakpoints (sharp)four breakpoints: \([0.5, 2]\) at \(x = 1\); fig-pwl area 1.250 (\(x^2\)) vs 6.283 (\(\sin\))
RLT level 1\(x_i x_j\) by \(X_{ij}\) plus all pairwise productsLP, \(O(n^2)\) columnslevel \(n\) for 0–1 programs (Sherali–Adams)R2: 0.30; knapsack \(1.5 \to 1.333 \to 1\)
Shor SDP\(X = x x^\top\) by \(M(x, X)\) PSDSDP, \((n+1) \times (n+1)\)exact in value when the only constraint is one strictly feasible quadratic constraint, with no box (S-lemma, Corollary 4.7.9)disc: 0.5 exact; R2: unbounded alone, 0.4136 with diagonal rows
SDP + RLTboth of the aboveSDP\(\operatorname{conv}\{(x, x x^\top)\}\) on \([0,1]^2\) for \(n = 2\)R2: 0.18 exact; R4: 500 (no gain); boxQP \(n = 3\): −2.039 vs −2
Lasserre order \(d\)moments up to degree \(2d\)SDP, \(\binom{n+d}{d} \times \binom{n+d}{d}\)finitely, generically (Nie)R2 order 2: 0.18 (\(6 \times 6\)); R4 order 2: 400.0003 (\(36 \times 36\))
copositive (Burer)the whole QP by one coneconic, NP-hard cone; DNN relaxation tractableexact in value for a QP with linear equalities, \(x \ge 0\) and binaries that these constraints bound in \([0, 1]\) (Theorem 4.7.13); DNN = CP up to order 4–
polyhedral cone approximationa cone by \(O(k \log(1/\varepsilon))\) rowsLPnever; within \(\varepsilon\)1% needs a 23-gon in the plane; 4 tower levels
The ladder of relaxations, with the numbers of this series (R1: the two-variable integer program; R2: max \(xy\) s.t. \(2x + y \le 1.2\); R3: st_e13; R4: Haverly 1; R5: three assets).

(Two qualifications to the table) Two rows need a qualification that the cells cannot hold. In the piecewise McCormick row the per-cell best is the MILP optimum of the partition formulation, which the tree over the selectors reaches. The LP relaxation of that formulation, with the coupling rows outside the disjunction, is the root McCormick bound, by Proposition 2.4.9, restated in Proposition 4.6.8. In the big-M row the hull statement is the one Proposition 4.2.7 proves for the tightest big-M of two polygons in the plane. It requires every facet of both polygons to carry the exact support-function constant of the other polygon, including the facets the other polygon already satisfies, whose constants are then negative. The disjunction figure's smallest valid \(M\) and the two-polytope example of Section 4.2 clamp those constants at zero, which is why both show the gaps in the last column. In three dimensions the tightest big-M of two polytopes is not sharp even without the clamp (Vielma's Example 12).

The ladder on the running examples: each relaxation's gap to the true optimum, the table's convention, for the example chosen (R1: the two-variable integer program; R2: max xy s.t. 2x + y ≤ 1.2; R3: st_e13; R4: Haverly 1; R5: three assets), and the piecewise McCormick MILP optimum on R2 against the number of pieces per axis.

(Three readings of the table) Three readings of the table are the chapter's summary. First, cost and tightness rise together but not in step. The biggest single gain on the running bilinear example, from \(0.30\) to \(0.18\), came from one linear inequality's worth of semidefinite information. The Shor relaxation on its own, at the price of a semidefinite program, was worthless there and on Haverly's problem. Second, "the hull when" is almost always a statement about one structure in isolation. McCormick is the hull of one product, the λ model of one graph, the perspective of one on/off set, the Shor relaxation of one quadratic constraint. The hull of an intersection is smaller than the intersection of hulls. This is why no single relaxation suffices and a tree is still needed. Third, the numbers from the figures are small-instance numbers, and the table's third column, what is solved per node, is the quantity that scales. A relaxation that costs a semidefinite program per node is a root-node tool today. A relaxation that costs a linear program with a warm start is the one every global solver runs at every node.

Where this is used

The relaxation that general-purpose global solvers run at every node is polyhedral: McCormick planes, secants, tangents and RLT rows inside a linear program, in BARON, Couenne, SCIP, ANTIGONE and Gurobi. The perspective enters SCIP as a strengthening of cuts for semicontinuous variables, and the big-\(M\) or hull formulations arrive from the modeller. The conic relaxations run at the nodes of the MISOCP solvers, MOSEK, COPT, Gurobi, CPLEX and Xpress, where the cones are the model's own. The semidefinite and moment relaxations run at the root or in specialized solvers for binary quadratic programs and box-constrained quadratic programs. The piecewise-linear route runs where the vendors put it: as the static approximation Gurobi used by default until version 12, and as the MIP relaxations of gas network models. The dualized formulation of Section 4.9 runs in research codes that drive a commercial MILP master. None of this is chosen by the solver from the model. The formulation decisions of this chapter are made before the solver starts, and the relaxation figure of Section 2.1, two polygons around one set of integers, is the smallest case of the decision.

Where the rungs run in a solver

  before the solver: the modeller's big-M or hull formulations
                   |
                   v
                  root     semidefinite and moment relaxations
                   o       (or in specialized solvers for binary
                  / \      and box-constrained QPs)
                 /   \
                o     o    every node: polyhedral rows in an LP;
               / \   / \   in MISOCP solvers, the model's own cones
              o   o o   o

What parallelizes

Every row of the table has a parallel part and a sequential part, and they are the same two parts each time. The relaxation's rows are independent: the four McCormick planes of every product, the outer-product rows of RLT, the perspective cuts of every indicator, the per-piece error constants of a piecewise-linear model, the certificate cuts of every cone. Building a relaxation is therefore a batch, and so is evaluating it at a frontier of nodes, as Section 6.4 and Section 7.4 describe. The sequential part is the solve that selects among what the rows allow: the simplex method's pivots, the interior-point method's factorizations, the MILP master's tree. First-order methods on GPUs attack exactly that part, for linear programs in Section 7.2 and for cones and semidefinite programs in Section 7.6. They pay for the parallelism with inexact duals, which must be repaired into safe bounds before a tree may prune on them. Which rung of the ladder a GPU solver should stand on is then a quantitative question. It weighs a cheap relaxation evaluated at many nodes against a tight one evaluated at few, and the frontier's size and the pruning lost to batching are the terms of the trade. That question is the subject of Section 7.8.

← Back to all posts