7. The GPU
Draft
The algorithms of Sections 2 to 6 were described as they run on a processor core. One relaxation is solved at a time, by a simplex method that starts from the parent's basis, inside a tree whose nodes are visited one or a few at a time. This section asks what changes when the hardware is a graphics processor. Two facts make the question timely. A first-order method for linear programming touches the constraint matrix only through two matrix–vector products, which a GPU evaluates over every nonzero at once. Since 2021 solvers built on that fact have become competitive with simplex and barrier codes on large instances and have entered the commercial solvers. And a branch-and-bound search is, at any moment, a frontier of many small related problems that differ only in their bounds, which is the shape of work a GPU is built for. The first four subsections treat four things in turn. Section 7.1 sets out the execution model and what fits it. Section 7.2 gives the first-order method and its convergence theory. Section 7.3 proves the theorem that turns an inexact dual vector into a rigorous bound. Section 7.4 works out what happens when first-order solves are placed inside a tree. The remaining subsections cover interior-point and nonlinear solves on the device, quadratic and conic programs, heuristics, and the research agenda.
One search tree on a processor core and on a graphics processor
processor core graphics processor
(Sections 2 to 6) (Section 7)
o o
/ \ / \
o o o o
/ \ / \ / \ / \
* . . . * * * *
one or a few nodes at a time; the frontier: many small
each relaxation by a simplex related problems that differ
method that starts from the only in their bounds
parent's basis
o processed * being solved . open, waiting its turn
What a GPU is good at
A graphics processor is a different machine from a processor core. Which steps of a node's work move to it is decided by the shape of the work, not by its amount. This subsection does four things. It gives the execution model in enough detail to reason about cost. It defines the one quantity that decides whether a kernel is limited by memory or by arithmetic. It recounts the attempts to run the simplex method on a GPU and why they did not take. And it ends with a table of what maps to the device and what does not. The reader who has never written a kernel needs only the three definitions.
The execution model
Definition 7.1.1 (the execution model). A kernel is a function launched on the device and executed by many threads at once. The threads are grouped into warps of 32 on NVIDIA’s devices, and the threads of a warp execute one instruction at a time, in lockstep, on one of the device's streaming multiprocessors. NVIDIA calls this single-instruction, multiple-thread execution. When the threads of a warp reach a branch and take different sides, the warp executes each side in turn with the other threads masked off. This is divergence, and it serializes the work. Each multiprocessor keeps many warps resident and switches among them whenever one waits on memory, which is how the device hides the latency of a memory access without caches of the size a processor core has. A kernel launch costs a fixed latency whatever the kernel does, so a kernel that does little work is dominated by its launch. Data crosses the link between host and device far more slowly than it moves inside device memory.NVIDIA, CUDA C++ Programming Guide, "Hardware Implementation" (the SIMT architecture and hardware multithreading), CUDA 13 documentation, docs.nvidia.com/cuda/cuda-c-programming-guide; the warp size of 32 and the lockstep execution are stated there. The advice to batch small kernels and to keep data on the device is NVIDIA, CUDA C++ Best Practices Guide, "Data transfer between host and device", in the same documentation set. No figure for the launch latency or for the bandwidth ratio is quoted here, because both depend on the device and its generation.
The execution model of Definition 7.1.1
host
^
| the host-device link: data crosses it far more slowly
v than it moves inside device memory
+--------------------------------------------------------------+
| device memory |
+--------------------------------------------------------------+
^ | ^ | ^ |
| v | v | v
+-------------+ +-------------+ +-------------+
| SM | | SM | | SM |
| warp warp | | warp warp | ... | warp warp |
| warp ... | | warp ... | | warp ... |
+-------------+ +-------------+ +-------------+
warp 32 threads, one instruction at a time, in lockstep
SM a streaming multiprocessor: it keeps many warps resident
and switches among them whenever one waits on memory
kernel a function executed by many threads at once; each launch
costs a fixed latency, whatever the kernel does
Three consequences follow and shape every algorithm in this section. Work must be organized in batches large enough to amortize the launch latency. The threads of a warp should execute the same instructions on different data, so data-dependent control flow is expensive. And data that lives on the device should stay there, because a transfer per node costs more than the node. Sparse matrix–vector products, elementwise maps and reductions fit the model. Pivoting, pointer-chasing and branching on the data do not.
Divergence: one warp of 32 threads at a branch on the data
<------- the 32 threads ------->
| before ################################
| side A ##.##...#.###..#.##.#...###.#.##
| side B ..#..###.#...##.#..#.###...#.#..
v
time # executes the instruction . masked off
the threads take different sides of the branch, and the warp runs
each side in turn: divergence serializes the work
Definition 7.1.2 (arithmetic intensity). The arithmetic intensity of a kernel is the number of floating-point operations it performs per byte it moves between memory and the processor. A kernel whose intensity lies below the ratio of the device's peak arithmetic rate to its memory bandwidth is bandwidth-bound: its running time is the time needed to move its data, and additional arithmetic on data already loaded is free. A sparse matrix–vector product does about two floating-point operations per nonzero loaded, one multiplication and one addition, and is bandwidth-bound on every current device. A sparse matrix times a dense matrix with \(b\) columns does \(2b\) operations per nonzero loaded.The two-operations-per-nonzero count and its consequence for batching are the starting point of N. Blin, S. Gualandi, C. Maes, A. Lodi and B. Stellato, "Batched first-order methods for parallel LP solving in MIP", arXiv 2601.21990 (2026). The ratio of peak rate to bandwidth as the dividing line is the roofline model of S. Williams, A. Waterman and D. Patterson, "Roofline: an insightful visual performance model for multicore architectures", Communications of the ACM 52 (2009).
(Why the product is bandwidth-bound) The last sentence of the definition is worth seeing from the kernel's side. Every nonzero of the matrix has to be loaded, its value and its column index, and the kernel multiplies it by one entry of the vector, adds the result into one entry of the output, and never touches it again. There is nothing to amortize the load against. The arithmetic units of any current device can perform far more than two operations in the time those bytes take to arrive, so they wait, and the kernel runs at the speed of the memory. Williams, Waterman and Patterson draw this as a roofline: the attainable rate rises in proportion to the intensity until it meets the device's peak at the ridge, the ratio of peak rate to bandwidth, and is flat beyond it.
(Why batching raises the intensity) The definition carries the whole economic argument for the batching of this section. A first-order LP iteration, an iteration of the method of Section 7.2, which uses only products with the constraint matrix and never factorizes it, is two sparse matrix–vector products. One LP iterated on its own reads its matrix twice per iteration and does almost nothing with the bytes. \(b\) related LPs that share the matrix and differ only in their bounds or objectives read the matrix once for all \(b\) and do \(2b\) operations per nonzero. Their intensity therefore grows with \(b\) until the traffic of the \(b\) iterate vectors takes over. The siblings of a branch-and-bound tree and the \(2n\) problems of optimality-based bound tightening (Section 2.6) are such families.
b LPs that share one matrix: one read of it serves all b
the sparse matrix one LP b LPs
+-------------------+ +---+ +----------------------+
| * . * . . * | | o | | o o ... o |
| . * . . * . | times | o | or | o o ... o |
| * . . * . . | | o | | o o ... o |
| . . * . * * | | o | | o o ... o |
+-------------------+ | o | | o o ... o |
| o | | o o ... o |
+---+ +----------------------+
an iterate b iterate vectors,
vector one column per LP
operations per nonzero 2 2b: a multiplication and
of the matrix loaded an addition for each LP
a first-order iteration is two such products; the intensity grows
with b until the traffic of the b iterate vectors takes over
Definition 7.1.3 (batch). A batch is a family of \(b\) problems that share a structure, such as one constraint matrix with different bounds, right-hand sides or objectives, or one expression graph evaluated on different boxes, so that one kernel can advance all \(b\) at once.
The families named a moment ago are batches in this sense, and it is worth checking each against the definition. The two children of a node of the R1 tree of Section 3.1 are the same LP over the same polygon with one bound changed, \(x_j \le \lfloor \bar x_j \rfloor\) in one child and \(x_j \ge \lfloor \bar x_j \rfloor + 1\) in the other, where \(\bar x\) is the node's relaxed point: one matrix, two bound vectors, \(b = 2\). The open nodes of a whole frontier differ from one another only in their bound vectors too, so they are a batch with \(b\) equal to the frontier's size. The \(2n\) problems of optimality-based bound tightening minimize and maximize each of the \(n\) variables over one relaxation, a batch that shares everything but the objective. And the boxes of a spatial tree, when each is bounded by evaluating the same factorable functions in interval arithmetic, share one expression graph, which is the second shape the definition names.
(The two prices of batching the frontier) The frontier is the open list \(\mathcal L\) of Definition 3.5.1. Bounding it in batches of \(b\) nodes is the frontier batching of Definition 6.4.1, run by Algorithm 6.4.3, and serial best-first search is the case \(b = 1\). The bound of one node does not depend on any other node, so the bounding of a frontier is parallel without any communication. What limits the gain is everything around the bounding. Proposition 6.4.4 bounds the speedup of the whole search by \(1/(s + (1 - s)/g) < 1/s\) when a fraction \(s\) of the work per node stays sequential and the bounding is accelerated by a factor \(g\), and the device-resident designs of Section 6.4 exist to shrink \(s\). That proposition counts only time. A batch also changes the tree. A batch of \(b\) nodes is committed before the incumbents that the batch itself will produce are known, so it bounds nodes that a serial best-first search would have pruned. Section 6.4 measures that second price with the batch figure, and Proposition 6.4.2 says what it consists of. It is small for small \(b\) and large for large \(b\), though not monotone in \(b\).
The two prices of bounding the frontier in batches of b nodes
1. time (Proposition 6.4.4): of the work per node a fraction s
stays sequential, and the bounding, 1 - s, is g times faster
serial [ s ][ 1 - s: the bounding ]
accelerated [ s ][ (1 - s)/g ]
speedup of the whole search <= 1/(s + (1 - s)/g) < 1/s
2. the tree (Proposition 6.4.2): a batch is committed before the
incumbents it will produce are known
serial best-first, b = 1 bound N_1 --> incumbent --> N_2
tested against it: pruned
a batch holding N_1, N_2 bound N_1 and N_2 together -->
incumbent: N_2 already bounded
The simplex method on a GPU
(Why the simplex method stayed on the core) The simplex method was among the first optimization algorithms people tried to move to graphics hardware, and the record is instructive. Bieling, Peschlow and Martini implemented the revised simplex method on a GPU in 2010 and reported a considerable speedup over a widely used CPU implementation. Lalami, Boyer and El Baz reported a maximum speedup of 12.5 on a GTX 260 board in 2011, and a maximum of 24.5 with two Tesla C2050 boards in a multi-GPU version the same year. Ploskas and Samaras gave GPU implementations of several simplex variants in 2015.J. Bieling, P. Peschlow and P. Martini, "An efficient GPU implementation of the revised simplex method", 2010 IEEE International Symposium on Parallel & Distributed Processing, Workshops and PhD Forum (2010); M. E. Lalami, V. Boyer and D. El Baz, "Efficient implementation of the simplex method on a CPU-GPU system", IPDPSW 2011; M. E. Lalami, D. El Baz and V. Boyer, "Multi GPU implementation of the simplex algorithm", HPCC 2011; N. Ploskas and N. Samaras, "Efficient GPU-based implementations of simplex type algorithms", Applied Mathematics and Computation 250 (2015). The two Lalami papers state that their instances are randomly generated and non-sparse. Every one of these codes worked on dense matrices, a tableau or an explicit basis inverse, and the speedups were measured on dense random instances. The simplex method that solvers actually run is the sparse revised method of Section 3.1, and its per-pivot work is of a different kind. The vocabulary is that of the paragraph before Proposition 3.1.12. A basis is a choice of \(m\) columns of the constraint matrix whose solution of the equality system is a vertex. A pivot exchanges one column of the basis for another. Pricing chooses the variable that enters or leaves the basis by its reduced cost. The ratio test chooses the step length that keeps the solution feasible. The dual simplex method keeps the signs of the reduced costs and repairs primal feasibility one pivot at a time, which is why it can start from the parent's basis (Proposition 3.1.12). The work of one pivot of the sparse method is then the following. Two sparse triangular solves are performed with an LU factorization of the basis, which is updated after each pivot and refactorized periodically with Markowitz pivoting, which chooses the pivot that creates the fewest new nonzeros, to preserve sparsity. A pricing step runs over a changing set of candidates. A ratio test follows. Hall and McKinnon showed that on many practical LPs these solves are hyper-sparse, touching a small and irregular set of entries. Huangfu and Hall built the update techniques and the parallel dual simplex of HiGHS on that observation, and obtained a few-fold speedup on a multicore processor and no more.J. A. J. Hall and K. I. M. McKinnon, "Hyper-sparsity in the revised simplex method and how to exploit it", Computational Optimization and Applications 32 (2005); Q. Huangfu and J. A. J. Hall, "Novel update techniques for the revised simplex method", Computational Optimization and Applications 60 (2015); Q. Huangfu and J. A. J. Hall, "Parallelizing the dual revised simplex method", Mathematical Programming Computation 10 (2018). The dense update of order \(mn\) per pivot, which the GPU codes accelerated, is exactly the work the sparse method avoids altogether. The work the sparse method does is small, irregular and strictly sequential from one pivot to the next. The GPU had nothing to offer it. That structural fact, and not any failure of engineering, is why linear programming stayed on the processor core when dense linear algebra left it. It is also why the method that did move, in the next subsection, is one whose iteration is two matrix–vector products over a fixed matrix.
One simplex pivot: the dense tableau and the sparse revised method
dense tableau sparse revised method
(the GPU codes, 2010 to 2015) (what solvers run)
the tableau, m x n the LU factors, m x m
o o o o o o o o o o o o o o . o . . .
o o o o o o o o o o o o o o . . . . .
o o o o o o o o o o o o o o . . . o .
o o o o o o o o o o o o o o o . . . .
o o o o o o o o o o o o o o . . . . .
every entry updated: order mn two sparse triangular solves
per pivot, regular and plentiful touch a small, irregular set
of entries (hyper-sparse)
pivot --> pivot --> pivot pivot --> pivot --> pivot
small, and strictly sequential
from one pivot to the next
o an entry a pivot works on . an entry it does not touch
What maps and what does not
The same test, regular and plentiful work against irregular and sequential work, sorts the rest of a global solver's node pipeline. The table lists the verdicts and the evidence. Each entry is treated at length in its own subsection.
| operation | shape of the work | verdict | where it is used, and the evidence |
|---|---|---|---|
| sparse matrix–vector products | one thread or warp per row; two flops per nonzero | maps | PDHG (next subsection); bandwidth-bound, so batching raises the intensity |
| \(b\) sibling LPs sharing one matrix | one matrix read, \(b\) right-hand sides (sparse \(\times\) dense product) | maps well | batched strong branching and OBBT (Blin et al. 2026); cuOpt 26.02 and 26.04 |
| elementwise projections, clips | one thread per entry | maps | every first-order step |
| reductions (sums, maxima) | a tree of partial sums | maps, with a caveat | norms, step tests, the safe bound; the summation order changes the bits (Propositions 6.2.16 and 6.2.18; 7.8) |
| interval arithmetic on a DAG | same straight-line program on every box; directed rounding | maps | per-operation rounding intrinsics; GPU interval bounders (Zhang et al. 2025) |
| pointwise McCormick relaxations | same program on every box and point | maps | generated CUDA kernels (Gottlieb, Xu and Stuber 2026), with a small operation set |
| local search moves | thousands of candidate moves scored at once | maps | Feasibility Jump, feasibility pump and tabu search on the GPU (cuOpt; CHAP) |
| bound propagation per constraint | one warp per row, bounds merged by atomics, rounds iterated | maps | Sofranac, Gleixner and Pokutta 2022: \(10\times\) to \(20\times\) over one CPU thread |
| sparse factorization, no pivoting | supernodal Cholesky / \(LDL^\top\) | maps in part | cuDSS inside condensed interior-point methods (7.5), at a cost in conditioning |
| simplex pivots | a sequential chain of sparse triangular solves and LU updates | does not | the warm-started dual simplex stays on the CPU in every solver |
| sparse LU with Markowitz pivoting | data-dependent pivot order | does not | the reason the simplex method stayed on the CPU |
| tree bookkeeping, node selection | pointer chasing, priority queues | does not | kept on the host, or on the device in padded tensors (6.4: Gmys; Liu and Lodi) |
| host–device transfer per node | latency per node | does not | Proposition 6.4.4 |
(Four rows of the table explained) Four rows need a word of explanation. The interval and McCormick rows map because evaluating a factorable function on a box is a straight-line program whose control flow is the same for every box. Zhang and coauthors run an interval lower bounder with domain partitioning on the GPU inside MAiNGO, and Gottlieb, Xu and Stuber generate CUDA kernels that evaluate pointwise McCormick relaxations for thousands of boxes at once. The table of Section 6.4 gives the authors' figures for both.H. Zhang, T. Kerkenhoff, N. Kichler, M. Dahmen, A. Mitsos, U. Naumann and D. Bongartz, "Accelerating deterministic global optimization via GPU-parallel interval arithmetic", arXiv 2507.20769 (2025); R. X. Gottlieb, P. Xu and M. D. Stuber, "Automatic source code generation for deterministic global optimization with parallel architectures", Optimization Methods and Software 41 (2026). Both speedups are the authors' numbers on their own test problems, and the supported operation set of the generated kernels was small at the time of reading. Directed rounding on the device is per operation, through the intrinsics named in Algorithm 7.3.7 and its sidenote. The propagation row maps because the tightening a constraint implies for its variables can be computed from the bounds at the start of a round, for every constraint at once, and merged at the end. Sofranac, Gleixner and Pokutta run whole propagation rounds on the device, one warp or block per constraint, with the candidate bounds merged by atomic minimum and maximum, device instructions that read a word, compare and write it back as one uninterruptible step, so that two threads tightening the same bound cannot overwrite each other's result. They report the same fixed points as the sequential code on 893 of 987 MIPLIB 2017 instances, with geometric-mean speedups of 10 to 20 over single-threaded propagation.B. Sofranac, A. Gleixner and S. Pokutta, "Accelerating domain propagation: an efficient GPU-parallel algorithm over sparse matrices", Parallel Computing 109 (2022); the sequential and parallel schedules reach the same fixed point by the order-independence of constraint propagation, which Section 2.6 states and Section 7.8 uses. The local-search row is A. Çördük, P. Sielski, A. Boucher and K. Aatish, "GPU-accelerated primal heuristics for mixed integer programming", arXiv 2510.20499 (2025), and G. K. Tjusila, A. Hoen, N.-C. Kempke, G. Mexi, T. Berthold, A. Gleixner, T. Koch and S. Pokutta, "CHAP: a hybrid GPU-CPU heuristic for MIP", arXiv 2605.05086 (2026). The factorization row names supernodal methods, sparse factorizations that handle the columns of the factor that share one sparsity pattern as a dense block, and Section 7.5 says what they cost on a device. The reduction row carries a caveat that Section 6.2 proved and Section 7.8 takes up again. Floating-point addition is not associative (Proposition 6.2.16), so a parallel sum's result depends on the order in which the partial sums are combined, and the result is reproducible when the kernel fixes that order (Proposition 6.2.18). The vendor libraries' own guarantee is the library-level form of that proposition, and Section 6.2 quotes it.A branch and bound whose pruning decisions depend on a sum must use a fixed-order reduction or accept run-to-run differences in the tree. Proposition 6.2.21 says that the validity of a directed-rounding bound survives a change of order and its reproducibility does not.
One propagation round on the device (Sofranac, Gleixner, Pokutta)
+--> the bounds at the start of the round
| | | | |
| v v v v
| +--------+ +--------+ +--------+ +--------+
| | row 1 | | row 2 | | row 3 | ... | row m |
| +--------+ +--------+ +--------+ +--------+
| | | | |
| v v v v
| candidate bounds for the variables of each row, all at once
| | | | |
| +-------------+------+------+-------------------+
| |
| v
| merged by atomic minimum and maximum
| |
+------ the next round <-----+-----> or stop at a fixed point
row i one warp or block per constraint
the same fixed points as the sequential code on 893 of 987
MIPLIB 2017 instances; geometric-mean speedups of 10 to 20 over
single-threaded propagation
One sum in two orders: the bits can differ (Proposition 6.2.16)
left to right, one addition a tree of partial sums
after another
x_1 x_2 x_1 x_2 x_3 x_4
\ / \ / \ /
(+) x_3 (+) (+)
\ / \ /
(+) x_4 \ /
\ / \ /
(+) (+)
floating-point addition is not associative, so the two results
can differ in their bits; a kernel that fixes the order gets the
same bits in every run (Proposition 6.2.18)
Where this is used
In October 2026 the production picture follows the table exactly. NVIDIA's cuOpt runs its LP first-order method, its barrier factorizations and its primal heuristics on the GPU and its branch and bound on the CPU. Gurobi 13 and COPT 8 run a first-order LP method on the GPU as one of the algorithms of their LP portfolios and keep their MILP and MINLP trees on the processor. MAiNGO has a GPU interval lower bounder in a 2025 preprint, not in its released documentation. No production solver runs a spatial branch and bound on a device. The next three subsections give the first-order method, the bound it needs, and what the tree does to both.
What parallelizes
What parallelizes, at the level of this subsection, is anything that can be written as the same arithmetic on many independent data. Examples are the bounds of a frontier of boxes, the products of a first-order iteration, the moves of a local search, and the propagation of all constraints in one round. What does not is the chain of decisions that makes a search a search.
First-order LP: PDHG
Every relaxation of Sections 2 to 4 is solved, in production, by a simplex method, and the previous subsection explained why that method does not move to a GPU. This subsection presents the method that does. The primal-dual hybrid gradient method, PDHG, solves a linear program by alternating a gradient step on the primal variables with a gradient step on the dual variables. It is a first-order method in the usual sense: it uses only first derivatives, which for a linear program are the vectors \(c\) and \(b\) and the matrix itself, and it never factorizes anything. It touches the constraint matrix only through one product with \(A\) and one with \(A^\top\) per iteration. Its convergence theory is the theory of fixed-point iterations of a nonexpansive operator, a map that never increases the distance between two points, so that the iterates can circle a solution but cannot run away from it. The rate is sublinear in general and linear on a linear program once restarts are added, with a rate constant that can be poor. That theory explains two things a user of these solvers must know: why the method is slow to high accuracy, and why restarts repair this. The engineering that turned the textbook iteration into a solver, PDLP and its GPU descendants, is given as two algorithm boxes. The figure runs the iteration on the two-variable LP. The subsection closes with the measured record of the GPU codes and what the vendors ship. The LP of the figures is a maximization and is treated as one where the figure is discussed. Everywhere else minimization is the default.
The problem and its optimality conditions
Definition 7.2.1 (the LP in two forms). The figures solve, with \(A \in \mathbb{R}^{m \times n}\), \(b \in \mathbb{R}^m\), \(c \in \mathbb{R}^n\) and a box \(X = [l, u]\),
\[\max_x\ c^\top x \quad \text{subject to}\quad A x \le b,\quad l \le x \le u, \tag{P-max}\]whose Lagrangian with multipliers \(y \ge 0\) on the rows is \(L(x, y) = c^\top x - y^\top (A x - b)\). The general form used by PDLP and its descendants is
\[\min_x\ c^\top x \quad \text{subject to}\quad K x \in [q^L, q^U],\quad l \le x \le u, \tag{P-gen}\]with infinite entries allowed in all four bound vectors, so that equalities (\(q^L_i = q^U_i\)), one-sided rows and free variables are special cases. Its Lagrangian is \(L(x, y) = c^\top x - y^\top K x + p(y)\) with \(p(y) = \sum_i \min(q^L_i y_i, q^U_i y_i)\), and the dual variable is confined to the set \(Y\) on which \(p\) is finite: \(y_i \ge 0\) where only \(q^L_i\) is finite, \(y_i \le 0\) where only \(q^U_i\) is finite, and \(y_i\) free for an equality row. Throughout this section the dual vector of a linear program is written \(y\), as the PDLP literature writes it, rather than the \(\lambda\) used for multipliers elsewhere in the monograph. It is the same object. When the LP is in the box form (P-max) or in the form of Theorem 7.3.1, the matrix is written \(A\), and \(K\) is reserved for (P-gen).
Definition 7.2.2 (saddle point, KKT residuals, relative KKT error). A pair \((x^\star, y^\star)\) is a saddle point of \(L\) on \(X \times Y\) if \(L(x^\star, y) \le L(x^\star, y^\star) \le L(x, y^\star)\) for all \(x \in X\) and \(y \in Y\), with the inequalities reversed for a maximization. For (P-gen), write \(r = c - K^\top y\) for the reduced cost of a pair \((x, y)\). Let \(\mathcal R\) be the set of reduced-cost vectors the box can absorb: a component \(\hat r_j > 0\) is allowed only when \(l_j > -\infty\), a component \(\hat r_j < 0\) only when \(u_j < \infty\), so that \(\hat r_j = 0\) is forced for a free variable. With \(\hat r = \Pi_{\mathcal R}(r)\) the projection of the reduced cost onto that set, the three KKT residuals of the pair are
\[\begin{aligned} &\text{primal infeasibility:} && \operatorname{dist}\big(Kx,\ [q^L, q^U]\big), \\ &\text{dual infeasibility:} && \lVert r - \hat r \rVert, \\ &\text{duality gap:} && \Big\lvert\, c^\top x - \Big( p(y) + \sum_{j=1}^n \min(\hat r_j l_j,\ \hat r_j u_j) \Big) \Big\rvert . \end{aligned}\]The quantity in the inner parentheses is the dual objective, finite because \(\hat r \in \mathcal R\). PDLP stops at relative tolerance \(\varepsilon\) when
\[\text{primal infeasibility} \le \varepsilon\,(1 + \lVert q \rVert),\qquad \text{dual infeasibility} \le \varepsilon\,(1 + \lVert c \rVert),\qquad \text{gap} \le \varepsilon\,(1 + \lvert c^\top x \rvert + \lvert \text{dual objective} \rvert),\]and "\(10^{-4}\) accuracy" or "\(10^{-8}\) accuracy" in the GPU LP literature always means this \(\varepsilon\).D. Applegate, M. Díaz, O. Hinder, H. Lu, M. Lubin, B. O'Donoghue and W. Schudy, "Practical large-scale linear programming using primal-dual hybrid gradient", Advances in Neural Information Processing Systems 34 (2021), arXiv 2106.04756; the journal version is "PDLP: a practical first-order method for large-scale linear programming", Mathematical Programming Computation (2026), arXiv 2501.07018. The dual objective is evaluated at the projected reduced cost \(\hat r\) rather than at \(r\) so that it stays finite when a free or one-sided variable's reduced cost has the wrong sign; that wrong-signed part is charged to the dual residual instead. Proposition 7.3.6 identifies the dual objective, evaluated at \(r\), as a safe bound.
(The residuals are the optimality conditions, measured off the solution) The residuals are the optimality conditions of the LP measured at a pair that need not be optimal. By Theorem 2.2.6, applied to (P-gen) with its rows and bounds written as inequalities, a pair \((x^\star, y^\star)\) is a saddle point of \(L\) on \(X \times Y\) exactly when \(x^\star\) is optimal for (P-gen) and \(y^\star\) is optimal for its dual. The saddle value \(L(x^\star, y^\star)\) is then \(z^\star\). The three residuals are the three parts of that theorem: primal feasibility, dual feasibility, which here means that the reduced cost lies in \(\mathcal R\), and a zero gap between the two objectives. PDHG therefore looks for a saddle point, and the residuals measure how far a pair is from being one. This is the sense in which the rest of the subsection is about a fixed-point problem.
(Residuals are not the distance to the value) The residuals measure violations of the optimality conditions, not the distance to the optimal value. A pair with residuals of \(10^{-4}\) can have an objective value whose distance from \(z^\star\) is much larger than \(10^{-4}\lvert z^\star \rvert\), and the pruning rules of a tree care about the value. Section 7.3 is about that distinction.
The iteration
The convergence theory uses a small vocabulary from convex analysis that the monograph has not needed until now.
Definition 7.2.3 (indicator, proximal map, nonexpansive and monotone operators, resolvent, averages). The indicator of a closed convex set \(S\) is \(\iota_S(x) = 0\) for \(x \in S\) and \(+\infty\) otherwise. The proximal map of a proper closed convex function \(g\) with parameter \(\tau > 0\) is
\[\operatorname{prox}_{\tau g}(v) \;=\; \arg\min_x \Big[\, g(x) + \frac{1}{2\tau}\lVert x - v \rVert^2 \Big],\]a single point, because the function minimized is strongly convex. For \(g = \iota_S\) it is the Euclidean projection \(\Pi_S(v)\) whatever \(\tau\) is, and for \(g(x) = c^\top x + \iota_S(x)\) it is \(\Pi_S(v - \tau c)\).
An operator \(T\) on a Euclidean space with inner product \(\langle \cdot, \cdot \rangle_P\) and norm \(\lVert \cdot \rVert_P\) is nonexpansive if \(\lVert Tz - Tz' \rVert_P \le \lVert z - z' \rVert_P\) for all \(z, z'\), and firmly nonexpansive if \(\lVert Tz - Tz' \rVert_P^2 \le \langle Tz - Tz', z - z' \rangle_P\). The second implies the first by the Cauchy–Schwarz inequality. Its fixed-point residual at \(z\) is \(\lVert z - Tz \rVert_P\), and \(\operatorname{Fix} T\) is its set of fixed points. A set-valued map \(F\) is monotone if \(\langle w - w', z - z' \rangle \ge 0\) whenever \(w \in F(z)\) and \(w' \in F(z')\), and maximal monotone if its graph is not properly contained in the graph of another monotone map. The subdifferential of a proper closed convex function is maximal monotone. The resolvent of \(F\) is \(J_F = (I + F)^{-1}\). When \(F\) is maximal monotone, \(J_F\) is single-valued and defined everywhere, its fixed points are the zeros of \(F\), and it is firmly nonexpansive. For an iteration \(z^{k+1} = Tz^k\) the ergodic average is \(\bar z^k = \frac1k \sum_{t=1}^k z^t\). A sequence is Fejér monotone with respect to a set \(S\) if its distance to every point of \(S\) is nonincreasing. Such a sequence is bounded and has at most one cluster point in \(S\), so if every cluster point lies in \(S\) the sequence converges to a point of \(S\).The proximal map is J.-J. Moreau, "Proximité et dualité dans un espace hilbertien", Bulletin de la Société Mathématique de France 93 (1965); maximal monotonicity of the subdifferential is R. T. Rockafellar, "On the maximality of subdifferential mappings", Pacific Journal of Mathematics 33 (1970); the everywhere-defined resolvent is G. J. Minty, "Monotone (nonlinear) operators in Hilbert space", Duke Mathematical Journal 29 (1962). Proofs of all the facts stated here, including the Fejér argument and the firm nonexpansiveness of resolvents in the metric that defines them, are in H. H. Bauschke and P. L. Combettes, Convex Analysis and Monotone Operator Theory in Hilbert Spaces, second edition (Springer, 2017), the chapters on Fejér monotonicity and on resolvents of monotone operators.
Definition 7.2.4 (PDHG and its norm). For a saddle problem \(\min_{x} \max_{y} \langle Kx, y \rangle + g(x) - h(y)\) with proper closed convex \(g, h\) whose proximal maps are computable, one PDHG step with step sizes \(\tau, \sigma > 0\) is
\[x^{k+1} = \operatorname{prox}_{\tau g}\!\left(x^k - \tau K^\top y^k\right),\qquad y^{k+1} = \operatorname{prox}_{\sigma h}\!\left(y^k + \sigma K\,(2 x^{k+1} - x^k)\right). \tag{PDHG}\]The PDHG metric is \(\lVert z \rVert_P^2 = \tau^{-1}\lVert x \rVert^2 + \sigma^{-1}\lVert y \rVert^2 - 2\langle Kx, y \rangle\) for \(z = (x, y)\), which is a norm exactly when \(\tau\sigma\lVert K \rVert^2 < 1\). PDLP writes the two steps with one step size \(\eta\) and a primal weight \(\omega\) as \(\tau = \eta/\omega\) and \(\sigma = \eta\omega\), and the induced norm is \(\lVert z \rVert_\omega^2 = \omega\lVert x \rVert^2 + \lVert y \rVert^2/\omega\).
(For an LP, the proximal maps are projections) For the LP (P-gen) the proximal maps are projections. Take \(g(x) = c^\top x + \iota_{[l,u]}(x)\) and \(h(y) = -p(y) + \iota_Y(y)\), with \(-K\) in the role of the matrix. By Definition 7.2.3 the primal step is \(x^{k+1} = \Pi_{[l,u]}(x^k - \tau(c - K^\top y^k))\), a gradient step on the cost followed by a clip to the box. The dual step for a one-sided row \(K_i x \ge q^L_i\) is \(y_i^{k+1} = \max(0,\ y_i^k + \sigma(q^L_i - K_i(2x^{k+1} - x^k)))\), a gradient step on the constraint violation followed by a clip to the nonnegative half-line. An equality row takes the gradient step without a clip. A two-sided row takes the proximal step of its interval term, which PDLP implements in closed form. The extrapolation \(2x^{k+1} - x^k\) in the dual step is the one feature that distinguishes PDHG from a plain alternation of gradient steps, and the next theorem shows that it is what makes the pair converge.
Theorem 7.2.5 (Chambolle and Pock, 2011). Let \(g, h\) be proper closed convex, let \(K\) be linear with \(\lVert K \rVert = L_K\), and suppose a saddle point exists. If \(\tau\sigma L_K^2 < 1\), the iterates of (PDHG) converge to a saddle point, and for every bounded \(B_1 \times B_2\) the ergodic averages \(\bar x^N = \frac1N\sum_{k=1}^N x^k\) and \(\bar y^N = \frac1N\sum_{k=1}^N y^k\) satisfy
\[\sup_{(x, y) \in B_1 \times B_2}\ \big[L(\bar x^N, y) - L(x, \bar y^N)\big]\ \le\ \frac{1}{N}\ \sup_{(x, y) \in B_1 \times B_2}\ \left[\frac{\lVert x - x^0 \rVert^2}{2\tau} + \frac{\lVert y - y^0 \rVert^2}{2\sigma}\right].\]Proof sketch. The two updates are optimality conditions: \(\tau^{-1}(x^k - x^{k+1}) - K^\top y^k \in \partial g(x^{k+1})\) and \(\sigma^{-1}(y^k - y^{k+1}) + K(2x^{k+1} - x^k) \in \partial h(y^{k+1})\). Subtracting the subgradient inequalities at an arbitrary \((x, y)\) and completing squares gives, for every \(k\),
\[L(x^{k+1}, y) - L(x, y^{k+1}) \;\le\; \tfrac12 \lVert z^k - z \rVert_P^2 - \tfrac12 \lVert z^{k+1} - z \rVert_P^2 - \tfrac12 \lVert z^{k+1} - z^k \rVert_P^2 .\]The cross term \(\langle K(x^{k+1} - x^k), y^{k+1} - y^k \rangle\) produced by the extrapolation is exactly what turns the two separate squared distances into the single \(P\)-norm, and \(\tau\sigma L_K^2 < 1\) is what makes \(P\) positive definite. Summing over \(k\) telescopes the right-hand side. Convexity of \(L\) in \(x\) and concavity in \(y\) move the sum of gaps inside to the averages. Taking \((x, y)\) to be a saddle point shows that \(\lVert z^k - z^\star \rVert_P\) is nonincreasing, so the sequence is Fejér monotone with respect to the saddle set. It also shows that \(\sum_k \lVert z^{k+1} - z^k \rVert_P^2\) is finite, so consecutive iterates come together and every cluster point is a fixed point of the continuous update, hence a saddle point. By Definition 7.2.3 the whole sequence converges. ∎A. Chambolle and T. Pock, "A first-order primal-dual algorithm for convex problems with applications to imaging", Journal of Mathematical Imaging and Vision 40 (2011); the ergodic rates are refined in A. Chambolle and T. Pock, "On the ergodic convergence rates of a first-order primal–dual algorithm", Mathematical Programming 159 (2016).
(Why the plain iteration spirals) In a picture, PDHG is a descent step on \(x\) and an ascent step on \(y\) taken together, and the extrapolation makes the pair behave like a single implicit step. The rate is \(O(1/N)\) for the averages. The last iterate converges, but in general with no rate, and on a linear program it spirals toward the solution rather than approaching it along a line. The spiral has a simple cause. The bilinear term \(\langle Kx, y \rangle\) contributes the field \((-K^\top y,\ Kx)\) to the pair's motion, which is a rotation: it moves \(x\) at right angles to the direction in which \(y\) moves it back, so a step along it circles the saddle point instead of approaching it, and only the proximal damping of the step, which shrinks the circle a little each time, brings the pair in. On a linear program the objective and constraints are linear, so there is nothing else to pull the iterates inward, and the spiral is tight. The figure below shows the spiral. The following proposition is the reason the Halpern and infeasibility theories of this subsection apply to PDHG.
Proposition 7.2.6 (PDHG is a proximal point iteration; He and Yuan, 2012). Let \(F(z) = (\partial g(x) + K^\top y,\ \partial h(y) - Kx)\) be the saddle operator, which is maximal monotone. One PDHG step is \(z^{k+1} = (I + P^{-1}F)^{-1} z^k\), the resolvent of \(P^{-1}F\), that is, the proximal point step for \(F\) in the metric \(P\). Consequently the PDHG operator \(T\) is firmly nonexpansive in \(\lVert \cdot \rVert_P\): \(\lVert Tz - Tz' \rVert_P^2 \le \langle Tz - Tz', z - z' \rangle_P\).
Proof. Rewrite the two optimality conditions of the previous proof as \(P(z^k - z^{k+1}) \in F(z^{k+1})\) with \(P = \begin{pmatrix} I/\tau & -K^\top \\ -K & I/\sigma \end{pmatrix}\). The first block row reads \(\tau^{-1}(x^k - x^{k+1}) - K^\top(y^k - y^{k+1}) \in \partial g(x^{k+1}) + K^\top y^{k+1}\) and the second \(-K(x^k - x^{k+1}) + \sigma^{-1}(y^k - y^{k+1}) \in \partial h(y^{k+1}) - Kx^{k+1}\). The extrapolation term is what produces the off-diagonal blocks. The operator \(P^{-1}F\) is maximal monotone in the inner product \(\langle z, z' \rangle_P = z^\top P z'\), because \(\langle P^{-1}w, z \rangle_P = w^\top z\), and the resolvent of a maximal monotone operator is firmly nonexpansive in the metric that defines it (Definition 7.2.3). ∎B. He and X. Yuan, "Convergence analysis of primal-dual algorithms for a saddle-point problem: from contraction perspective", SIAM Journal on Imaging Sciences 5 (2012). When the objective carries a convex term with an \(L_f\)-Lipschitz gradient, such as the quadratic of a QP relaxation, it can be handled by an explicit gradient step inside the primal update, and the iterates still converge provided \(\tau^{-1} - \sigma\lVert K \rVert^2 > L_f/2\): L. Condat, "A primal–dual splitting method for convex optimization involving Lipschitzian, proximable and linear composite terms", Journal of Optimization Theory and Applications 158 (2013), and B. C. Vũ, "A splitting algorithm for dual monotone inclusions involving cocoercive operators", Advances in Computational Mathematics 38 (2013). Section 7.6 takes up the quadratic case.
Sharpness and restarts
The \(O(1/N)\) rate is the rate for a general convex-concave problem. A linear program has more structure than that, and the structure has a name.
Definition 7.2.7 (sharpness, Hoffman constant, normalized duality gap). A function \(f\) with minimum set \(Z^\star\) is \(\alpha\)-sharp on a set \(S\) if \(f(z) - f^\star \ge \alpha\,\operatorname{dist}(z, Z^\star)\) for all \(z \in S\). The Hoffman constant \(H(G)\) of a polyhedron \(\{z : Gz \ge h\}\) is the smallest constant such that \(\operatorname{dist}(z, \{Gz \ge h\}) \le H(G)\,\lVert (h - Gz)_+ \rVert\) for every \(z\).A. J. Hoffman, "On approximate solutions of systems of linear inequalities", Journal of Research of the National Bureau of Standards 49 (1952). The normalized duality gap of a primal-dual pair \(z = (x, y)\) at radius \(r > 0\) is
\[\rho_r(z) \;=\; \frac{1}{r}\ \max_{\hat z \in W_r(z)}\ \big[L(x, \hat y) - L(\hat x, y)\big],\qquad W_r(z) = \{\hat z \in Z : \lVert \hat z - z \rVert \le r\},\]a localized duality gap that is finite even when the feasible region is unbounded. A primal-dual problem is \(\alpha\)-sharp on \(S\) if \(\rho_r\) is \(\alpha\)-sharp on \(S\) for every \(r \in (0, \operatorname{diam} S]\).
Proposition 7.2.8 (a linear program is sharp; Applegate, Hinder, Lu and Lubin, 2023, Lemma 5). Write the optimality conditions of the LP, primal feasibility, dual feasibility and a zero gap, as one system of linear inequalities \(Gz \ge h\) in the pair \(z = (x, y)\), with Hoffman constant \(H(G)\). For every \(R > 0\) the primal-dual LP is \(\alpha\)-sharp on the ball \(W_R(0)\) with \(\alpha = \big(H(G)\sqrt{1 + 4R^2}\big)^{-1}\).
Proof sketch. The violation of the optimality system at \(z\) is dominated by the localized gap, \(\lVert (h - Gz)_+ \rVert \le \rho_r(z)\sqrt{1 + R^2}\) on \(W_R(0)\) (Lemma 4 of the paper), and Hoffman's inequality bounds the distance to the solution set by that violation. Combine the two. ∎
Theorem 7.2.9 (restarted PDHG converges linearly; Applegate, Hinder, Lu and Lubin, 2023). A primal-dual algorithm here means any method whose ergodic averages \(\bar z^t\) from a start \(z^0\) satisfy the two bounds below, for constants \(C > 0\) and \(q \ge 0\) and every \(t\). The constant \(C\) is the constant of the \(O(1/t)\) ergodic bound of Theorem 7.2.5 after normalization by the radius. The constant \(q\) bounds how far the points averaged can wander from the current iterate, in units of its distance to \(Z^\star\). The two bounds are
\[\rho_{\lVert \bar z^t - z^0 \rVert}(\bar z^t) \;\le\; \frac{2C}{t}\,\lVert \bar z^t - z^0 \rVert, \qquad \lVert \bar z^t - z^0 \rVert \;\le\; (q + 2)\,\operatorname{dist}(z^0, Z^\star).\]PDHG with step \(\eta \le 1/\lVert A \rVert_2\) satisfies both in its own norm \(\lVert \cdot \rVert_P\) with \(C = 1/\eta\) and \(q = 0\), and, for \(\eta < 1/\lVert A \rVert_2\), in the Euclidean norm with \(C = 2/(\eta(1 - \eta\lVert A \rVert_2))\) and \(q = 4(1 + \eta\lVert A \rVert_2)/(1 - \eta\lVert A \rVert_2)\), so that \(C\) is of the order of \(\lVert A \rVert_2\) at the largest admissible step. Fix \(\beta \in (0, 1)\) and put \(t^\star = \lceil 2C(q + 2)/(\alpha\beta) \rceil\).
(i) Fixed-frequency restarts. Restart from the average every \(t^\star\) iterations. If the problem is \(\alpha\)-sharp on \(W_R(z^{0,0})\) with \(R = \frac{q + 2}{1 - \beta}\operatorname{dist}(z^{0,0}, Z^\star)\), then \(\operatorname{dist}(z^{n,0}, Z^\star) \le \beta^n\,\operatorname{dist}(z^{0,0}, Z^\star)\), where \(z^{n,0}\) is the point from which the \(n\)-th restart began.
(ii) Adaptive restarts. Restart from the average the first time the normalized duality gap has decayed by the factor \(\beta\),
\[\rho_{\lVert \bar z^{n,t} - z^{n,0} \rVert}(\bar z^{n,t}) \;\le\; \beta\, \rho_{\lVert z^{n,0} - z^{n-1,0} \rVert}(z^{n,0}).\]If the problem is \(\alpha\)-sharp on a set containing every restart point, then every restart length \(\tau_n\) is at most \(t^\star\), and \(\operatorname{dist}(z^{n,0}, Z^\star) \le \beta^n\,(t^\star/\tau_0)\,\operatorname{dist}(z^{0,0}, Z^\star)\), where \(\tau_0\) is the length of the first restart.
In either case an \(\varepsilon\)-accurate point is reached in \(O\!\left(\frac{C}{\alpha}\log\frac{1}{\varepsilon}\right)\) iterations in all.
Proof sketch. Within one restart, sharpness turns the bound on the normalized gap of the average into \(\operatorname{dist}(\bar z^t, Z^\star) \le \rho(\bar z^t)/\alpha \le \frac{2C(q+2)}{\alpha t}\operatorname{dist}(z^0, Z^\star)\), so after \(t^\star\) inner iterations the distance has contracted by \(\beta\). In case (i) the restart points stay inside \(W_R(z^{0,0})\), where sharpness is assumed, because the restart steps are bounded by \((q+2)\) times a geometric sequence, and induction over the restarts gives the decay. In case (ii) the same estimate shows that the adaptive criterion has triggered by iteration \(t^\star\) at the latest. Chaining the restart conditions then gives the decay, with the factor \(t^\star/\tau_0\) coming from bounding the first restart's gap through its own length rather than through \(t^\star\). ∎D. Applegate, O. Hinder, H. Lu and M. Lubin, "Faster first-order primal-dual methods for linear programming using restarts and sharpness", Mathematical Programming 201 (2023): the two displayed properties are its Property 3, the constants for PDHG are its Propositions 1 and 3 and Corollary 2, and the two restart schemes are its Theorems 1 and 2 with Remark 3. Alternative geometric condition measures that bound the same complexity through the LP's level-set geometry are in Z. Xiong and R. M. Freund, "Computational guarantees for restarted PDHG for LP based on 'limiting error ratios' and LP sharpness", Mathematical Programming (2026), arXiv 2312.14774; linear convergence of the restarted average under a quadratic error bound is proved in O. Fercoq, "Quadratic error bound of the smoothed gap and the restarted averaged primal-dual hybrid gradient", Open Journal of Mathematical Optimization 4 (2023).
Fixed-frequency restarts, Theorem 7.2.9 (i): a restart every t*
iteration 0 t* 2t* 3t*
| | | |
z^{0,0} --> z^{1,0} --> z^{2,0} --> z^{3,0} --> ...
dist to Z* <= d beta d beta^2 d beta^3 d
each arrow: t* PDHG steps from the restart point, then a
restart from the average of those t* iterates
d = dist(z^{0,0}, Z*), t* = ceil( 2C(q + 2) / (alpha beta) ),
beta in (0, 1); the bounds hold if the problem is alpha-sharp
on W_R(z^{0,0}) with R = (q + 2) d / (1 - beta)
Theorem 7.2.10 (optimality of the rate, and the price of not restarting; the same paper). For every \(\alpha\), \(C\) and dimension \(m\) there is a bilinear problem in \(\mathbb{R}^m \times \mathbb{R}^m\), \(\alpha\)-sharp and \(C\)-smooth, on which every primal-dual method whose iterates stay in the span of the previous iterates and gradients satisfies \(\operatorname{dist}(z^t, Z^\star) \ge (1 - \alpha/C)^t\operatorname{dist}(z^0, Z^\star)\) for \(t < m\). Hence \(\Omega\!\left(\frac{C}{\alpha}\log\frac1\varepsilon\right)\) iterations are necessary in the worst case, and restarted PDHG is optimal in this class up to constants. On bilinear problems with \(\kappa = \sigma_{\max}(A)/\sigma^+_{\min}(A)\), PDHG without restarts needs \(\Theta(\kappa^2\log(1/\varepsilon))\) iterations, against \(O(\kappa\log(1/\varepsilon))\) with restarts.
(What the two theorems say together) The two theorems together say what to expect. A linear program is always sharp, so restarted PDHG always converges linearly, and in that sense the method is not sublinear on LP. But the rate constant is a Hoffman constant, a geometric quantity that is large when facets of the feasible polyhedron are nearly parallel or the problem is degenerate, and nothing in the method improves it. The geometry is easy to see in the plane: where two facets meet at a very acute angle, a point that violates both by a small amount can sit far down the thin wedge between them, a long way from the feasible set, so a small residual there says little about the distance to a solution, and that is exactly the ratio the Hoffman constant measures. Sharpness inherits the same constant, which is why the rate can be linear and still slow. The unrestarted iteration is quadratically worse. Plain PDHG spirals toward the solution with a radius that contracts only by a factor \(1 - O(\kappa^{-2})\) per iteration, averaging kills the rotation, and restarting resets the spiral at the much better average.
Halpern variants and reflection
The next three results do the same job for the last iterate instead of the average.
Theorem 7.2.11 (Halpern, 1967; Lieder, 2021). Let \(T\) be nonexpansive with a fixed point and let the Halpern iteration anchored at \(z^0\) be \(z^{k+1} = \frac{k+1}{k+2}\,T(z^k) + \frac{1}{k+2}\,z^0\). Then \(\lVert z^k - Tz^k \rVert \le \frac{2}{k+1}\operatorname{dist}(z^0, \operatorname{Fix} T)\), and the bound is tight.B. Halpern, "Fixed points of nonexpanding maps", Bulletin of the American Mathematical Society 73 (1967), gives the iteration and its convergence; the tight rate is F. Lieder, "On the convergence rate of the Halpern-iteration", Optimization Letters 15 (2021).
The anchor pulls the iterate back toward \(z^0\) by an amount that shrinks exactly fast enough to damp the rotation that slows PDHG. The last iterate then has an \(O(1/k)\) fixed-point residual, where the plain iteration \(z^{k+1} = Tz^k\) has only \(O(1/\sqrt k)\) in general.
Theorem 7.2.12 (restarted Halpern PDHG; Lu and Yang, 2024). Let \(T\) be the PDHG operator of a feasible bounded LP with step size \(\eta \le 1/(2\lVert A \rVert_2)\). For every \(R > 0\) there is a constant \(\alpha_\eta > 0\) with \(\alpha_\eta\operatorname{dist}(z, Z^\star) \le \lVert z - Tz \rVert\) whenever \(\lVert Tz \rVert \le R\): the fixed-point residual is sharp, and \(\alpha_\eta\) plays for it the role that \(\alpha\) plays for the normalized gap in Definition 7.2.7. Apply the Halpern iteration to \(T\) and restart from the current iterate, resetting the anchor, either every \(k^\star = \lceil 2e/\alpha_\eta \rceil\) iterations or adaptively when \(\lVert z^{n,k} - Tz^{n,k} \rVert \le \frac1e\lVert z^{n,0} - Tz^{n,0} \rVert\). Then a restart of length \(\tau_n\) contracts the distance to the solution set by the factor \(2/(\alpha_\eta \tau_n)\), the restart points converge linearly, and an \(\varepsilon\)-accurate point is reached in \(O\!\left(\frac{1}{\alpha_\eta}\log\frac1\varepsilon\right)\) iterations in all, the same order as Theorem 7.2.9 but stated for the last iterate. The method identifies the active set in finite time, after which the rate is governed by a local sharpness constant rather than the global Hoffman constant. On an infeasible or unbounded LP it produces infeasibility certificates at an accelerated linear rate without a nondegeneracy assumption.H. Lu and J. Yang, "Restarted Halpern PDHG for linear programming", arXiv 2407.16144 (2024), Proposition 2 for the sharpness of the residual and Theorems 1 to 5 for the rates; the reflected variant is its Theorem 6. The Halpern iterate differs from the ergodic average on a generic LP because of the projections, so the two are different methods, not two names for one.
Proposition 7.2.13 (reflection halves the constant). If \(T\) is firmly nonexpansive then \(2T - I\) is nonexpansive with the same fixed points, and the Halpern iteration applied to \(2T - I\) gives \(\lVert z^k - Tz^k \rVert \le \frac{1}{k+1}\operatorname{dist}(z^0, \operatorname{Fix} T)\).
Proof. Expanding \(\lVert (2T - I)z - (2T - I)z' \rVert^2\) shows that firm nonexpansiveness of \(T\) is equivalent to nonexpansiveness of \(2T - I\), and the two have the same fixed points. Theorem 7.2.11 applied to \(2T - I\) gives \(\lVert (2T - I)z^k - z^k \rVert = 2\lVert Tz^k - z^k \rVert \le \frac{2}{k+1}\operatorname{dist}(z^0, \operatorname{Fix} T)\). ∎
By Proposition 7.2.6 the PDHG operator is firmly nonexpansive, so the reflected scheme \(z^{k+1} = \frac{k+1}{k+2}\big((1 + \gamma)Tz^k - \gamma z^k\big) + \frac{1}{k+2}z^0\) with \(\gamma = 1\) is admissible. cuPDLPx uses a reflection parameter \(\gamma \in [0, 1]\) and reports that reflection allows longer steps than the plain Halpern update. The same base iteration appears from the splitting side as a Halpern Peaceman–Rachford method, which is the method behind HPR-LP. The Peaceman–Rachford splitting is an operator-splitting iteration that is equivalent to PDHG's base step on this problem. The two codes differ in their restart rules, step and weight control, and engineering, not in the iteration.H. Lu, Z. Peng and J. Yang, "cuPDLPx: a further enhanced GPU-based first-order solver for linear programming", arXiv 2507.14051 (2025); K. Chen, D. Sun, Y. Yuan, G. Zhang and X. Zhao, "HPR-LP: an implementation of an HPR method for solving linear programming", Mathematical Programming Computation 18 (2026), and, for the relation between the two, K. Chen, D. Sun, Y. Yuan, G. Zhang and X. Zhao, "On the relationships among GPU-accelerated first-order methods for solving linear programming", arXiv 2509.23903 (2025). The acceleration also appears from the proximal-point side as D. Kim, "Accelerated proximal point method for maximally monotone operators", Mathematical Programming 190 (2021).
The reflected Halpern step, gamma in [0, 1], anchored at z^0
reflection r = (1 + gamma) T z^k - gamma z^k
= T z^k + gamma (T z^k - z^k)
z^k T z^k r
o-------------------------o-------------------------o
|<----- T z^k - z^k ----->|<- gamma (T z^k - z^k) ->|
anchoring z^{k+1} = (k+1)/(k+2) r + 1/(k+2) z^0
z^0 z^{k+1} r
o---------------------------------------o-----------o
|<------ (k+1)/(k+2) of the way ------->|<--------->|
1/(k+2)
drawn with gamma = 1: then r = (2T - I) z^k, the reflection of
z^k through T z^k (Proposition 7.2.13)
Infeasible nodes
(What an infeasible node has instead of a solution) A branch-and-bound tree produces infeasible nodes constantly, and a method that only converges when a saddle point exists would be useless in one. What an infeasible node has instead of a solution is a certificate. Farkas' lemma, in the form this section needs, says that for a box-constrained system exactly one of two statements holds: either the set \(\{x : Ax \ge b,\ l \le x \le u\}\) is nonempty, or there is a \(y \ge 0\) with \(b^\top y + \sum_j \min\big(-(A^\top y)_j l_j,\ -(A^\top y)_j u_j\big) > 0\). A certificate of infeasibility is such a \(y\). That the two statements cannot hold together is Corollary 7.3.4 below. That one of them always holds is linear programming duality, Theorem 2.2.6, applied to the feasibility problem with zero objective. Its dual function is the positively homogeneous function of \(y\) just displayed, which is finite for every \(y \ge 0\) because the box is compact. If the primal is infeasible the dual is therefore unbounded above, and some \(y\) gives a positive value.J. Farkas, "Theorie der einfachen Ungleichungen", Journal für die reine und angewandte Mathematik 124 (1902). The box form stated here is the general lemma specialized to the rows and the bounds written as inequalities; a textbook statement and proof are in M. Conforti, G. Cornuéjols and G. Zambelli, Integer Programming, Graduate Texts in Mathematics 271 (Springer, 2014), Chapter 3. The last theorem of the theory says what PDHG does when there is nothing to converge to.
Theorem 7.2.14 (infeasibility detection; Applegate, Díaz, Lu and Lubin, 2024). Let \(T\) be the PDHG operator for an LP with \(\tau\sigma\lVert A \rVert_2^2 < 1\) and let \(v\) be the minimum-norm element of \(\operatorname{cl}\operatorname{range}(T - I)\), the infimal displacement vector. (i) The normalized iterates \(z^k/k\) and the normalized averages \(\frac{2}{k+1}(\bar z^k - z^0)\) converge to \(v\) at rate \(O(1/k)\), because \(\operatorname{range}(T - I)\) is closed for PDHG on an LP, \(T - I\) being a composition of affine maps and polyhedral projections. (ii) If the primal and the dual are both feasible, the iterates converge to a saddle point. If the primal (dual) is infeasible, the dual (primal) iterates diverge along a ray and \(v\) is a Farkas certificate of infeasibility. (iii) Under a strict complementarity condition at the limit, the differences \(z^{k+1} - z^k\) converge to \(v\) at an eventual linear rate.D. Applegate, M. Díaz, H. Lu and M. Lubin, "Infeasibility detection with primal-dual hybrid gradient for large-scale linear programming", SIAM Journal on Optimization 34 (2024), Theorems 3, 4 and 6.
On an infeasible LP the operator has no fixed point and the iterates drift at a constant velocity. The drift velocity is the infeasibility certificate. In practice PDLP, cuPDLP and cuOpt detect infeasibility by watching the iterate differences and the normalized averages, which costs nothing beyond the iteration itself. Corollary 7.3.4 shows how the certificate is made rigorous in one further matrix–vector product.
Theorem 7.2.14: PDHG on an infeasible LP, where T has no fixed point
z^0 z^1 z^2 z^3 z^k
o-------->o-------->o-------->o------ . . . ------->o-------->
the iterates drift at a constant velocity v, the minimum-norm
element of cl range(T - I), and v is the certificate:
z^k / k --> v at rate O(1/k) (i)
2/(k+1) (zbar^k - z^0) --> v at rate O(1/k) (i)
z^{k+1} - z^k --> v eventually linearly, under
strict complementarity (iii)
primal (dual) infeasible: the dual (primal) iterates diverge
along a ray, and v is a Farkas certificate of infeasibility (ii)
From iteration to solver
With the theory in hand, here is the method as the figure runs it, in the maximization form (P-max).
Algorithm 7.2.15 (PDHG in the box form).
Algorithm 7.2.15 PDHG for max c'x s.t. A x <= b, l <= x <= u
Input A (m x n), b, c, the box [l, u], step sizes tau, sigma with
tau*sigma*||A||_2^2 < 1, iteration count N
Output iterates (x^k, y^k), k = 0..N, with l <= x^k <= u and y^k >= 0
1. x^0 := any point of the box; y^0 := 0
2. for k = 0, 1, ..., N-1:
3. x^{k+1} := clip( x^k + tau * (c - A' y^k), l, u )
# one product with A'
4. y^{k+1} := max( 0, y^k + sigma * (A (2 x^{k+1} - x^k) - b) )
# one product with A
Invariant
||z^k - z*||_P is nonincreasing for every saddle point z*
(Proposition 7.2.6); the ergodic averages satisfy the O(1/k) gap
bound of Theorem 7.2.5; x^k is always in the box and y^k >= 0.
Cost per step
two sparse matrix-vector products, about 2 nnz(A) flops each, plus
O(m + n) vector work.
Parallel
everything. Each product is one thread or warp per row or column,
the clips are one thread per entry, and nothing inside a step
depends on anything else inside the step.
One PDHG step in the box form: steps 3 and 4 of Algorithm 7.2.15
y^k ---> [ A' y^k ] one product with A'
|
x^k ---------+
v
x^{k+1} := clip( x^k + tau * (c - A' y^k), l, u )
|
x^k ---------+
v
2 x^{k+1} - x^k the extrapolation
|
v
[ A (2 x^{k+1} - x^k) ] one product with A
|
y^k ---------+
v
y^{k+1} := max( 0, y^k + sigma * (A (2 x^{k+1} - x^k) - b) )
A is touched only by the two products; the clip to the box and
the clip at 0 act entry by entry
For the minimization form (P-gen), step 3 becomes \(x^{k+1} = \Pi_{[l,u]}(x^k - \tau(c - K^\top y^k))\) and step 4 becomes \(y^{k+1} = \Pi_Y(y^k + \sigma(q - K(2x^{k+1} - x^k)))\) with the projection appropriate to each row. The textbook iteration needs five additions before it is a solver: a rescaling of the problem, a rule for the step size, a rule for the primal weight, the restarts of Theorem 7.2.9, and a termination test on the original problem.
Definition 7.2.16 (restart, primal weight, preconditioning). A restart replaces the current iterate by a chosen point (the ergodic average, or the Halpern anchor) and begins the inner iteration afresh from it. The primal weight \(\omega\) of Definition 7.2.4 rebalances primal and dual progress. PDLP initializes it at \(\lVert c \rVert_2/\lVert q \rVert_2\) and moves it after each restart toward the ratio of the dual to the primal distance travelled. Diagonal preconditioning replaces \(K\) by \(D_1 K D_2\) with positive diagonal matrices: Ruiz equilibration iterates \(D_1 \leftarrow D_1\operatorname{diag}(\lVert K_{i,\cdot} \rVert_\infty^{-1/2})\) and \(D_2 \leftarrow D_2\operatorname{diag}(\lVert K_{\cdot,j} \rVert_\infty^{-1/2})\), and the Pock–Chambolle preconditioner with parameter \(\alpha\) uses \((D_1)_{ii} = (\sum_j \lvert K_{ij} \rvert^{2-\alpha})^{-1/2}\) and \((D_2)_{jj} = (\sum_i \lvert K_{ij} \rvert^{\alpha})^{-1/2}\). PDLP applies ten Ruiz iterations followed by Pock–Chambolle with \(\alpha = 1\).D. Ruiz, "A scaling algorithm to equilibrate both rows and columns norms in matrices", Technical Report RAL-TR-2001-034, Rutherford Appleton Laboratory (2001), cited as PDLP cites it (the report could not be fetched online on 5 October 2026); T. Pock and A. Chambolle, "Diagonal preconditioning for first order primal-dual algorithms in convex optimization", 2011 IEEE International Conference on Computer Vision (2011).
Algorithm 7.2.17 (PDLP-style restarted PDHG).
Algorithm 7.2.17 PDLP-style restarted PDHG
Input an LP in the form (P-gen); a relative tolerance eps; restart
constants beta_suff, beta_nec, beta_art
Output (x, y) with relative KKT error <= eps (Definition 7.2.2), or an
infeasibility certificate
1. Presolve. Rescale: ten Ruiz iterations, then Pock-Chambolle with
alpha = 1 (Definition 7.2.16).
2. omega := ||c||_2 / ||q||_2 (1 if either is near 0);
eta := 1 / ||K||_inf; z^{0,0} := 0; n := 0; k := 0
3. repeat (outer loop over restarts, counter n):
4. t := 0
5. repeat (inner loop):
6. adaptive step: compute a candidate (x', y') with
tau = eta/omega, sigma = eta*omega;
eta_bar := ||(x' - x, y' - y)||_omega^2
/ ( 2 |(y' - y)' K (x' - x)| );
eta' := min( (1 - (k+1)^(-0.3)) eta_bar,
(1 + (k+1)^(-0.6)) eta );
accept the candidate if eta <= eta_bar, with eta' as
the next eta; else set eta := eta' and retry
7. z^{n,t+1} := (x', y');
update the running average zbar^{n,t+1};
k := k + 1; t := t + 1
8. every 64 iterations: let z_c be the average or the current
iterate, whichever has the smaller progress measure mu (the
normalized duality gap, or the KKT error), and test
(i) mu(z_c) <= beta_suff * mu(z^{n,0}) (sufficient decay)
(ii) mu(z_c) <= beta_nec * mu(z^{n,0}) and mu increased
(necessary decay, no progress)
(iii) t >= beta_art * k (artificial restart)
until one of (i) to (iii) holds
9. z^{n+1,0} := z_c; primal weight update:
omega := exp( theta * log( ||y^{n+1,0} - y^{n,0}||
/ ||x^{n+1,0} - x^{n,0}|| )
+ (1 - theta) log omega ), theta = 0.5
10. termination test on the unscaled problem (Definition 7.2.2);
infeasibility test on the iterate differences and the
normalized averages (Theorem 7.2.14)
11. until converged
Invariant
within an inner loop the P-norm distance to the solution set is
nonincreasing; across restarts it contracts geometrically under
sharpness (Theorem 7.2.9).
Cost per step
two sparse matrix-vector products plus O(m + n); the restart test
adds O(m + n) for the KKT error, or a small trust-region solve for
the normalized gap, which is why the GPU codes use the KKT error.
Parallel
the products and the vector work; the adaptive step needs one
reduction per trial; the restart decision and the primal-weight
update are scalars.
The three loops of Algorithm 7.2.17
1-2 presolve, rescale; omega, eta; z^{0,0} := 0; n := 0; k := 0
|
v
3 outer loop over restarts, counter n <------------------+
| |
4 t := 0 |
| |
5 inner loop <-------------------------------------+ |
| | |
6 adaptive step: candidate (x', y'), eta_bar <-+ | |
eta <= eta_bar? -- no: eta := eta' ----------+ | |
| yes | |
7 z^{n,t+1} := (x', y'); running average | |
zbar^{n,t+1}; k := k + 1; t := t + 1 | |
| | |
8 every 64 iterations: z_c := the average or the | |
iterate, whichever has the smaller mu; | |
(i), (ii) or (iii) holds? ------------- no ------+ |
| yes |
9 z^{n+1,0} := z_c; primal weight update omega |
| |
10 termination test on the unscaled problem; |
infeasibility test |
| |
11 converged? ------------------------------- no ---------+
| yes
v
(x, y) with relative KKT error <= eps, or an
infeasibility certificate
The constants as published: the journal version of PDLP uses \(\beta_{\mathrm{suff}} = 0.1\), \(\beta_{\mathrm{nec}} = 0.9\) and \(\beta_{\mathrm{art}} = 0.5\) with the normalized duality gap as the progress measure. cuPDLP.jl uses the KKT error with \(0.2\), \(0.8\) and \(0.36\).The step-size rule is Algorithm 2 of the NeurIPS paper and the primal-weight update its Section 3.3. The restart constants are those of the Mathematical Programming Computation version; the NeurIPS paper prints the same two constants with their names interchanged (\(\beta_{\mathrm{sufficient}} = 0.9\), \(\beta_{\mathrm{necessary}} = 0.1\)), so the journal values are used here. The cuPDLP.jl constants are from H. Lu and J. Yang, "cuPDLP.jl: a GPU implementation of restarted primal-dual hybrid gradient for linear programming in Julia", Operations Research 73 (2025), arXiv 2311.12180. PDLP's authors state that the primal-weight update has no convergence proof. The adaptive step, step 6, is the one piece of the algorithm that is awkward on a GPU, because each trial needs a reduction and a possible retry before the step can be taken. cuPDLPx removes it.
Algorithm 7.2.18 (reflected restarted Halpern PDHG with a controlled primal weight; cuPDLPx).
Algorithm 7.2.18 Reflected restarted Halpern PDHG
with a PID-controlled primal weight (cuPDLPx)
Input the rescaled LP; a constant step eta := 0.998 / ||A||_2, with
||A||_2 estimated by power iteration (the repeated product
A'A v that converges to the largest singular value);
reflection gamma in [0, 1]; primal weight w := 1
Output the last iterate z with relative KKT error <= eps
1. anchor z^{n,0} := z^0; k := 0
2. repeat:
3. v := PDHG(z^{n,k}) with tau = eta/w, sigma = eta*w
# steps 3 and 4 of Algorithm 7.2.15
4. z^{n,k+1} := (k+1)/(k+2) * ( (1 + gamma) v - gamma z^{n,k} )
+ 1/(k+2) * z^{n,0}
5. r := || z^{n,k+1} - PDHG(z^{n,k+1}) ||_P
(fixed-point residual; evaluated every few iterations)
6. restart if r <= beta_suff * r(z^{n,0}),
or r <= beta_nec * r(z^{n,0}) and r > r_prev,
or k >= beta_art * (total iterations):
z^{n+1,0} := z^{n,k+1};
e := log( sqrt(w) ||x^{n+1,0} - x^{n,0}||
/ ( ||y^{n+1,0} - y^{n,0}|| / sqrt(w) ) );
log w := log w - ( K_P e + K_I sum_i e_i + K_D (e - e_prev) );
k := 0
7. until the KKT test passes on the original instance
Invariant
Halpern's bound ||z^k - T z^k|| <= 2 dist(z^0, Z*)/(k+1) inside an
inner loop, halved by reflection (Proposition 7.2.13); linear
contraction across restarts (Theorem 7.2.12).
Cost per step
two sparse matrix-vector products plus O(m + n); the residual costs
one extra PDHG step when it is evaluated, so it is evaluated
periodically. There is no step-size search and hence no extra
reduction per step.
Parallel
as Algorithm 7.2.15. The constant step is chosen precisely to remove
the sequential trial-and-retry of the adaptive rule.
cuPDLPx is therefore a restarted, reflected Halpern scheme with a constant step size and a proportional–integral–derivative controller on the primal weight, a feedback rule from control engineering that moves the weight by amounts proportional to the current imbalance between primal and dual progress, to the accumulated imbalance, and to its rate of change. Its authors' measurements against cuPDLP.jl on an H100 are in the table. The abstract summarizes them as 2.5 to 5 times on the MIPLIB relaxations and 3 to 6.8 times on Mittelmann's set. The C implementation is about a fifth faster than the Julia prototype.H. Lu, Z. Peng and J. Yang, "cuPDLPx: a further enhanced GPU-based first-order solver for linear programming", arXiv 2507.14051 (2025); code at github.com/MIT-Lu-Lab/cuPDLPx under the Apache 2.0 licence. The speedups are the authors' own, in shifted geometric mean of run times.
| instance set | tolerance | without presolve | with presolve | hard subset |
|---|---|---|---|---|
| 383 MIPLIB 2017 LP relaxations | \(10^{-4}\) | 2.11x to 2.52x | 3.08x | 4.85x |
| 383 MIPLIB 2017 LP relaxations | \(10^{-8}\) | 2.35x to 2.86x | 3.57x | 4.99x |
| Mittelmann's 49-instance LP set | \(10^{-4}\) | about 3x to 4x | 6.8x | – |
One more device belongs to the engineering, because a tree will need it. A first-order primal iterate violates rows by up to the tolerance, and a point that violates rows is not a feasible solution. PDLP repairs it by solving the feasibility problem on its own.
Theorem 7.2.19 (feasibility polishing; Applegate and coauthors, journal version, Appendix C). Let \(X_F = \{x \ge 0 : Ax = b\}\) be nonempty, with Hoffman constant \(H(A)\) in the sense \(H(A)\operatorname{dist}(x, X_F) \le \lVert Ax - b \rVert_2\) for all \(x\), and let \(x^{0,0}\) be an approximately optimal, approximately feasible point for \(\min\{c^\top x : x \in X_F\}\). Run PDHG on the feasibility problem with the objective set to zero, started at \(x^{0,0}\) with dual iterate \(0\) and step \(\eta = \alpha'/\lVert A \rVert_2\) with \(\alpha' \in (0, 1)\), restarting from the primal average every \(\ell\) iterations with
\[\ell \;\ge\; \Big\lceil \frac{2(q + 2)}{\eta\, H(A)} \Big\rceil, \qquad q = 4\,\frac{1 + \alpha'}{1 - \alpha'},\]the constant \(q\) of Theorem 7.2.9 in the Euclidean norm. Then for every \(n \ge 1\) the restart points satisfy
\[\operatorname{dist}(x^{n,0}, X_F) \;\le\; 2^{-n}\operatorname{dist}(x^{0,0}, X_F), \qquad c^\top x^{n,0} - c^\top x^{0,0} \;\le\; 2(q + 2)\,\lVert c \rVert_2\,\frac{\lVert A x^{0,0} - b \rVert_2}{H(A)},\]so the iterates approach the feasible set geometrically and the objective moves by at most a constant times the initial infeasibility.
The practical version alternates primal and dual polishing. It is what let PDLP solve eight of eleven linear programs with 125 million to 6.3 billion nonzeros to a feasibility of \(10^{-8}\) and gaps near one percent within six days. Gurobi's barrier method solved three of the eleven and exceeded a terabyte of memory on the other eight. Without polishing only two of the eleven were solved.D. Applegate, M. Díaz, O. Hinder, H. Lu, M. Lubin, B. O'Donoghue and W. Schudy, "PDLP: a practical first-order method for large-scale linear programming", Mathematical Programming Computation (2026), Appendix C (Lemma 1, Theorem 1 and inequality (22)) and the large-instance experiments; arXiv 2501.07018. Google's research blog of 20 September 2024, "Scaling up linear programming with PDLP", states that PDLP has been in production at Google since May 2023 and that it was co-awarded the Beale–Orchard-Hays Prize at the International Symposium on Mathematical Programming in July 2024; both are the blog's statements, and the prize was not checked against the Mathematical Optimization Society's own announcement. This is the regime in which a first-order method wins outright: one enormous LP, where a factorization does not fit in memory and a billion sparse products do.
The figure and a reference implementation
The figure runs Algorithm 7.2.15 on the two-variable LP of the plane figures, with \(c = (\cos\theta, \sin\theta)\) at \(\theta = 45^\circ\), step sizes \(\tau = \sigma = 0.9/\lVert A \rVert_2\), and restarts to the running average every forty iterations, beside the simplex method and the barrier method on the same problem. In the notation of Definition 7.2.1 the instance is the polygon of R1 (Section 1.6) in the form (P-max): \(m = 4\) rows and \(n = 2\) variables, with
\[A = \begin{pmatrix} 2 & 5 \\ 5 & 2 \\ -3 & 4 \\ 1 & -2 \end{pmatrix},\qquad b = \begin{pmatrix} 24.5 \\ 30.5 \\ 11 \\ 4.2 \end{pmatrix},\qquad l = \begin{pmatrix} 0 \\ 0 \end{pmatrix},\qquad u = \begin{pmatrix} 7 \\ 6 \end{pmatrix},\qquad c = \begin{pmatrix} \cos 45^\circ \\ \sin 45^\circ \end{pmatrix},\]so the dual vector \(y\) has four nonnegative components, one per row, the reduced cost is \(r = c - A^\top y \in \mathbb{R}^2\), and the optimum is \(z^\star = 5.556\) at the vertex \((4.929, 2.929)\), where the first two rows are active. The listing after the figure carries the same data. Its "show" control has four views, described in turn.
(The plane view) In the plane, the simplex method reaches the optimal vertex in 3 pivots, each a sequential basis update that lands exactly on a vertex, while PDHG spirals in through the interior. The scrub, the figure's iteration slider, shows the iterate at \(k = 120\) at \((4.372, 3.099)\) against the vertex \((4.929, 2.929)\). The objective panel plots three curves against the iteration. The primal value \(c^\top x^k\) approaches \(z^\star = 5.556\) from either side. The raw dual value \(b^\top y^k\) sits below \(z^\star\) for most of the run, and for a maximization that means it is not a bound at all. The corrected dual value drawn in orange lies above \(z^\star\) at every iteration. At \(k = 120\) the three read 5.2831, 4.0325 and 6.2088, the last of them 11.75 percent above the optimum. The readout reports that the primal value is within a tenth of a percent after 119 iterations and the corrected bound after 208, and the stats quote the 3 pivots against the 119 iterations.
(The residuals view) The residuals view plots PDLP's three relative residuals in form, with the box kept as a projection rather than dualized, all on a logarithmic scale. They are the primal residual \(\lVert (Ax - b)^+ \rVert/(1 + \lVert b \rVert)\), the positive part of the reduced cost \(\lVert (c - A^\top y)^+ \rVert/(1 + \lVert c \rVert)\), and the gap \(\lvert c^\top x - b^\top y \rvert/(1 + \lvert c^\top x \rvert + \lvert b^\top y \rvert)\) with the raw dual value. At \(k = 120\) the primal residual is \(0\), because the iterate lies inside the polygon, the dual residual is \(1.6 \times 10^{-1}\) and the gap \(1.2 \times 10^{-1}\). With the box dualized, as in Definition 7.2.2, the dual residual is zero by construction and the gap at \(k = 120\) is \(7.4 \times 10^{-2}\) against the corrected value 6.2088. The readout prints both versions. With the fixed restarts to the average, all three residuals first fall below \(10^{-4}\) at \(k = 236\), below \(10^{-6}\) at \(k = 286\) and below \(10^{-8}\) at \(k = 368\), and stay below from \(k = 255\), 328 and 391. With the restarts switched off the same thresholds are reached at \(k = 131\), 205 and 262. A fixed restart period of forty iterations therefore costs iterations on this problem. The iterate walks toward the vertex, the average lags behind the walk, and every restart pulls the iterate back. The adaptive criteria of Algorithm 7.2.17 exist to restart only when the progress measure has decayed, and the figure's take says so.
(The barrier view) The barrier view runs the interior-point method on the same LP. The barrier method, named in Section 1.3 and treated in full in Section 7.5, replaces each inequality by a logarithmic penalty weighted by a parameter \(\mu > 0\) and follows the minimizers of the penalized problem, the central path, as \(\mu\) decreases to zero. At a point of the central path every one of the eight inequalities, four rows and four box bounds, has slack times multiplier equal to \(\mu\), so the gap between the dual value with all eight dualized and the primal value is exactly \(8\mu\). The figure starts the method at \((2, 2)\), divides \(\mu\) by five at each centring, and follows the central path to \((4.928, 2.928)\) at \(\mu = 5\times 10^{-5}\). Its dual values at the central points fall from 34.8 at \(\mu = 4\) through 10.52, 6.51, 5.75 and 5.59 to 5.56.
(All three methods on logarithmic axes) The view with all three methods puts each method's relative distance to the optimum on logarithmic axes. All three end at the same vertex by three different routes, and the simplex route is the shortest in steps and the only one that is sequential in every step.
The orange curve of the objective panel is the subject of the next subsection. For now it is enough to know its formula. For any \(y \ge 0\) and any \(x\) in the box \(l \le x \le u\), with \(r = c - A^\top y\),
\[c^\top x \;\le\; b^\top y \;+\; \sum_{j=1}^{n} \max\big(r_j l_j,\ r_j u_j\big), \tag{SB-max}\]which with \(l = 0\) is \(b^\top y + \sum_j u_j\max(0, r_j)\). Corollary 7.3.2 proves it. The following listing is the reference implementation of restarted PDHG with that bound on the figure's problem. It computes the exact optimum by vertex enumeration and runs plain and restarted PDHG for 400 iterations. It prints the primal value, the raw dual value, the bound (SB-max) evaluated at the running average of the dual iterates, and the largest row violation, and it checks that the bound never falls below the optimum.
# Restarted PDHG with the safe dual bound.
#
# The two-variable LP of the plane figures:
#
# maximize c.x subject to A x <= b, 0 <= x <= u,
# c = (cos 45°, sin 45°).
#
# Safe bound (max form): for any y >= 0 and r = c - A^T y,
#
# c.x <= b.y + sum_j max(r_j l_j, r_j u_j) on the box.
import numpy as np
A = np.array([[2., 5.], [5., 2.], [-3., 4.], [1., -2.]])
b = np.array([24.5, 30.5, 11., 4.2])
l = np.zeros(2)
u = np.array([7., 6.])
th = np.deg2rad(45.0)
c = np.array([np.cos(th), np.sin(th)])
def vertices():
"""The exact optimum by vertex enumeration (the simplex answer)."""
rows = np.vstack([A, np.eye(2), -np.eye(2)])
rhs = np.concatenate([b, u, -l])
best = (-np.inf, None)
for i in range(len(rows)):
for j in range(i + 1, len(rows)):
M = rows[[i, j]]
if abs(np.linalg.det(M)) < 1e-12:
continue
v = np.linalg.solve(M, rhs[[i, j]])
if np.all(rows @ v <= rhs + 1e-9) and c @ v > best[0]:
best = (c @ v, v)
return best
def safe_bound(y):
"""Neumaier-Shcherbina, max form: valid for every y >= 0."""
r = c - A.T @ y
return b @ y + np.sum(np.maximum(r * l, r * u))
def pdhg(N=400, R=40, step=0.9, restart=True):
# tau * sigma * ||A||^2 < 1
tau = sigma = step / np.linalg.norm(A, 2)
x = np.zeros(2)
y = np.zeros(4)
xa = x.copy()
ya = y.copy()
cnt = 1
hist = []
for k in range(1, N + 1):
# one product with A^T
xn = np.clip(x + tau * (c - A.T @ y), l, u)
# one product with A
y = np.maximum(0.0, y + sigma * (A @ (2 * xn - x) - b))
x = xn
# running averages
cnt += 1
xa += (x - xa) / cnt
ya += (y - ya) / cnt
# restart to the average
if restart and cnt > R:
x, y, cnt = xa.copy(), ya.copy(), 1
hist.append((k, c @ x, b @ y, safe_bound(ya),
np.max(A @ x - b, initial=0.0)))
return hist
zstar, xstar = vertices()
print(f"exact optimum (simplex): z* = {zstar:.6f} "
f"at x* = ({xstar[0]:.6f}, {xstar[1]:.6f})")
for restart in (False, True):
H = pdhg(restart=restart)
label = "restarted PDHG (R = 40)" if restart else "plain PDHG"
hit = next((k for k, p, _, _, v in H
if abs(p - zstar) <= 1e-3 * zstar and v <= 1e-3), None)
hits = next((k for k, _, _, s, _ in H
if s - zstar <= 1e-3 * zstar), None)
print(f"\n{label}:\n"
f" primal within 0.1% (row violation <= 1e-3) at k = {hit};\n"
f" safe bound within 0.1% at k = {hits}\n")
print(" k c.x b.y (raw) safe bound max row violation")
for k, p, raw, s, v in H:
if k in (1, 10, 40, 100, 120, 200, 400):
print(f"{k:4d} {p:8.5f} {raw:9.5f} {s:10.5f} {v:9.2e}")
assert all(s >= zstar - 1e-9
for _, _, _, s, _ in H), "safe bound violated"
print("\nevery safe bound of every iteration is >= z*: checked")
exact optimum (simplex): z* = 5.555839 at x* = (4.928571, 2.928571)
plain PDHG:
primal within 0.1% (row violation <= 1e-3) at k = 79;
safe bound within 0.1% at k = None
k c.x b.y (raw) safe bound max row violation
1 0.12504 0.00000 9.19239 0.00e+00
10 1.25036 0.00000 9.19239 0.00e+00
40 4.98164 2.75272 9.12374 1.40e-01
100 5.55584 5.52669 7.07704 2.51e-02
120 5.55584 5.54561 6.82512 3.04e-04
200 5.55584 5.55581 6.31996 7.41e-06
400 5.55584 5.55584 5.93885 2.67e-12
restarted PDHG (R = 40):
primal within 0.1% (row violation <= 1e-3) at k = 119;
safe bound within 0.1% at k = 208
k c.x b.y (raw) safe bound max row violation
1 0.12504 0.00000 9.19239 0.00e+00
10 1.25036 0.00000 9.19239 0.00e+00
40 2.50023 0.08626 9.12374 0.00e+00
100 5.34452 4.16599 6.63979 0.00e+00
120 5.28307 4.03247 6.20878 0.00e+00
200 5.55421 5.49206 5.60844 0.00e+00
400 5.55584 5.55584 5.55584 1.05e-08
every safe bound of every iteration is >= z*: checked
(Reading the listing's output) The exact vertex is the one the figure's simplex reaches in three pivots, and the restarted run's row for \(k = 120\) reproduces the figure's readout to four decimals. The raw dual value is below \(z^\star\) for most of both runs, so it is not a bound. The safe bound is above \(z^\star\) at every one of the 800 iterations checked. Without restarts the last iterate converges quickly, with the primal value within a tenth of a percent at \(k = 79\). But the running average of the duals, from which the bound is read, lags: its bound is still 6.9 percent loose at \(k = 400\). With restarts to the average the iterate itself is slower, at 119 iterations, but the dual average converges and the bound with it. The bound is within a tenth of a percent at \(k = 208\) and exact to five decimals at \(k = 400\). Each iteration costs two products with a \(4 \times 2\) matrix and a handful of vector operations, every one of which is parallel across rows and columns. On this problem the arithmetic is trivial, and the point of the listing is the shape of the iteration, not its speed.
A numpy script run for this post, not reproduced here, extends the listing above with the restart tests of Algorithms 7.2.17 and 7.2.18 and separates the three variants on two problems: the two-variable LP, and a random sparse LP with 300 rows, 200 columns, five percent density and the box \([0, 2]\), constructed to be feasible. It implements the three variants in the maximization form with the cuPDLP.jl restart constants, checks the relative KKT error of Definition 7.2.2 every ten iterations, and gives the same counts from run to run. The table records the first iteration at which that error is below each tolerance. The reader cannot check these counts from the listings of this post, and they are quoted for the shape of the comparison rather than for their exact values.
| plain PDHG, last iterate | restarted average (0.2 / 0.8 / 0.36) | reflected Halpern, restarted | |
|---|---|---|---|
| two-variable LP, 45 degrees | 170 / 220 / 270 | 190 / 230 / 290 | 110 / 140 / 170 |
| random sparse LP, \(300 \times 200\) | 8,950 / – / – | 3,720 / 8,440 / 17,310 | 1,960 / 5,380 / 8,510 |
(Reading the table) On the two-variable problem the three are within a factor of two of one another and all reach \(10^{-8}\) in a few hundred iterations: a linear program in two dimensions defined by a vertex is extremely sharp. On the \(300 \times 200\) problem the plain iteration reaches \(10^{-4}\) after 8,950 iterations and does not reach \(10^{-6}\) within 50,000. Both restarted variants reach \(10^{-8}\). The reflected Halpern variant adds a nearly constant number of iterations per two decades of accuracy, which is the linear convergence of Theorems 7.2.9 and 7.2.12, and the stall of the plain method is the \(\kappa^2\) behaviour of Theorem 7.2.10. The reflected Halpern variant needs between half and two thirds of the iterations of the restarted average at every tolerance; the halving of Proposition 7.2.13 compares it with the unreflected Halpern iteration, which the table does not record.
The measured record
(The codes' own measurements) What the GPU codes report, independently measured where that is possible, is the following. cuPDLP.jl, the first GPU implementation of PDLP, solved on an H100 381 of the 383 MIPLIB 2017 LP relaxations to \(10^{-4}\) with presolve, with a shifted geometric mean time (shift 10 s) of 7.37 s. Gurobi's barrier solved all 383 in 4.65 s and Gurobi's dual simplex 373 in 11.84 s. At \(10^{-8}\) cuPDLP.jl solved 373. Its speedups over CPU PDLP were about 4, 10 and 20 times on small, medium and large instances, with about a second of fixed overhead that no small instance recovers.H. Lu and J. Yang, "cuPDLP.jl: a GPU implementation of restarted primal-dual hybrid gradient for linear programming in Julia", Operations Research 73 (2025), arXiv 2311.12180 (2023). The shifted geometric mean is defined in Section 5.2. cuPDLP-C, the C implementation written with Cardinal Operations, is about half again as fast as the Julia version. At \(10^{-8}\) it solved 369 of the 383 relaxations (shifted mean 18.53 s) against COPT's 383 in 3.11 s, and 41 of the 49 instances of Mittelmann's LP set against COPT's 48.H. Lu, J. Yang, H. Hu, Q. Huangfu, J. Liu, T. Liu, Y. Ye, C. Zhang and D. Ge, "cuPDLP-C: a strengthened implementation of cuPDLP for linear programming by C language", arXiv 2312.14832 (2023). At high accuracy on ordinary instances the CPU barrier code is faster and solves more; the first-order method's case is elsewhere. HPR-LP, the Halpern Peaceman–Rachford code, reports shifted-mean speedups of 2.39 to 5.70 times over PDLP at \(10^{-8}\) with presolve on an A100.K. Chen, D. Sun, Y. Yuan, G. Zhang and X. Zhao, "HPR-LP: an implementation of an HPR method for solving linear programming", Mathematical Programming Computation 18 (2026), arXiv 2408.12179.
(The independent benchmark) The independent record is Mittelmann's LP feasibility benchmark. It asks each code for a primal-dual feasible point on 65 instances, 49 public and 16 undisclosed, with the CPU codes under a 15,000 s limit and the GPU codes under a 1,000 s limit. He showed it at the INFORMS Annual Meeting in Atlanta on 28 October 2025, with the GPU codes on an H100, and his conclusions slide read that cuOpt and COPT's GPU barrier led among the GPU codes and that COPT led among the CPU codes. A rerun on a B200 for a talk in Hong Kong in March 2026 put COPT's GPU barrier first among the GPU codes, with cuOpt second and cuPDLPx third. The ranking among the GPU codes has changed at every update, so any snapshot must carry its date, and the one quoted in full below is the live page's state on 16 September 2026.H. D. Mittelmann, "Latest progress in optimization software", INFORMS Annual Meeting, Atlanta, 28 October 2025, slides 11, 12 and 23, plato.asu.edu/talks/informs2025.pdf. Slide 11 gives, with the CPU codes on an Intel i7-11700K at 3.6 GHz with 64 GB of memory under the 15,000 s limit and the GPU codes on the H100 under the 1,000 s limit, in scaled shifted geometric means with the number solved of 65, COPT 1.19 (65), OptVerse 2.08 (64), MOSEK 3.93 (59), Xpress 6.87 (59), PDLP 20.0 (50), KNITRO 20.3 (49) and HiGHS 26.2 (49) among the CPU codes, and COPT's GPU barrier 1 (64), cuOpt 1.06 (62), cuPDLPx 1.94 (61), cuPDLP 2.22 (56) and HPR-LP 3.03 (59) among the GPU codes. H. D. Mittelmann, "Benchmarking optimization software: a (hi)story", talk at PolyU, Hong Kong, March 2026, slide 32, plato.asu.edu/talks/hongkong26.pdf, with cuOpt 26.02, cuPDLPx 0.2.5, COPT 8.0.3 and HPR-LP-C 0.1.1 on a B200: scaled means 1 for COPT's GPU barrier, 1.19 for cuOpt, 1.74 for cuPDLPx and 2.60 for HPR-LP, with 64, 61, 58 and 54 of the 65 solved. Neither talk says that Gurobi 13 and COPT ship PDHG on the GPU; that is documented by the vendors, below. The October 2025 talk also showed six further instances from Oliver Hinder's collection, three heat-equation and three multicommodity-flow problems, run with a 15,000 s limit for every code. The heat instances have 15,625,000 constraints, 31,628,008 variables and 125 million nonzeros. The flow instances have 1.5 million constraints, 126 million variables and 253 million nonzeros. cuPDLP-C solved all six, in 7,288, 3,696, 3,014, 1,166, 2,260 and 1,461 s. cuOpt ran out of memory on all six. COPT's GPU barrier solved the three flow instances and ran out of memory on the three heat instances. cuPDLPx solved the three flow instances and timed out on the heat instances, HPR-LP solved none within the limit, and a Gurobi 13 beta on the GPU timed out on all six. The conclusions slide records that cuPDLP-C was the only GPU code to solve all the large problems.Slide 12 of the INFORMS talk, which gives the two families' dimensions separately, and slide 23. The instance slide heads its column "cuPDLP" and the conclusions slide names the code cuPDLP-C; the latter, the C implementation of Lu, Yang and coauthors cited above, is used here. The live page is updated as versions change. Its state on 16 September 2026, with the GPU codes on a B200 under a 1,000 s limit and a tolerance of \(10^{-6}\), is the table below.
| code | device | scaled mean | solved of 65 |
|---|---|---|---|
| HPR-LP-C 0.1.3 | GPU | 1.00 | 63 |
| cuOpt 26.08 | GPU | 1.17 | 62 |
| COPT 8 barrier (COPT-G) | GPU | 1.29 | 64 |
| cuPDLPx 0.3.0 | GPU | 2.43 | 57 |
| COPT 8 | CPU | 1.67 | 65 |
| MOSEK | CPU | 5.90 | 56 |
| Xpress (XOPT) | CPU | 9.63 | 59 |
| HiGHS | CPU | 16.9 | 55 |
| KNITRO | CPU | 23.0 | 48 |
| PDLP (OR-Tools) | CPU | 27.8 | 50 |
(Three safe readings and one unsafe one) Three readings are safe and one is not. The best GPU codes are between 1.3 and 1.7 times faster than the best CPU code in shifted geometric mean and solve one to three fewer instances. The ranking among the GPU codes has changed at every update of the page, so any one of them should be quoted with its date. The decisive wins are on the instances with hundreds of millions of nonzeros where the CPU codes do not finish, and there the limit is memory rather than speed. The reading that is not safe is the vendor one. NVIDIA's blog of October 2024 announcing cuOpt's LP solver reports "over 5,000x" on a single instance against a CPU-based solver, which on inspection is a comparison of GPU PDLP with CPU PDLP, not with the simplex and barrier codes of the table.H. D. Mittelmann, "LPfeas benchmark (find PD feasible point) + addendum", page of 16 September 2026, plato.asu.edu/ftp/lpfeas.html; its addendum on Hinder's problems with a 4,800 s limit gives HPR-LP 13 solved (scaled 1), cuPDLPx 8 (3.41), cuOpt 7 (3.95) and COPT's GPU barrier 5 (5.56). The vendor figure is N. Blin, "Accelerate large linear programming problems with NVIDIA cuOpt", NVIDIA Technical Blog, 8 October 2024, a vendor claim.
What the vendors ship
The state of the vendors' products in October 2026 is as follows. Gurobi 13.0.0, released on 11 November 2025, added PDHG to its LP algorithms, selected with Method=6. By default it runs on the CPU. With PDHGGPU=1 it runs on an NVIDIA GPU from a separate download, with its own tolerances PDHGAbsTol, PDHGRelTol and PDHGConvTol, and crossover, the pivoting step that turns an interior point into an optimal basis, is available from the point it produces. Release 13.0.2 of 5 May 2026 states that the GPU PDHG "has matured from the beta state and now is a fully supported feature".Gurobi Optimizer Reference Manual, "Release notes: additions, changes and removals", versions 13.0.0 to 13.0.3, docs.gurobi.com, and the parameter reference for Method, PDHGGPU and Crossover; the release dates are those of gurobipy on PyPI (13.0.0 on 11 November 2025, 13.0.2 on 5 May 2026). Gurobi's own performance figures for version 13 against 12 are vendor claims and are left to Section 5.3. COPT 8.0 selects its first-order method with LpMethod=6 and runs it on the GPU by default when a GPU is present (GPUMode). It offers a GPU barrier for LP as a separate mode and stops the first-order method at PDLPTol, whose default is \(10^{-6}\). Its documentation says the GPU solver supports LP and the root relaxation of MILP.Cardinal Optimizer (COPT) User Guide, version 8.0, "Parameters" (LpMethod, GPUMode, GPUDevice, PDLPTol) and "FAQ: GPU usage related", guide.coap.online/copt/en-doc. HiGHS has shipped cuPDLP-C as its PDLP option since version 1.7.0 (7 March 2024), with a GPU build since 1.10.0 (20 March 2025). Version 1.14.0 (6 April 2026) added a native first-order solver, HiPDLP, which can use NVIDIA GPUs. GAMS 54.1.0 of 15 June 2026 describes the addition and notes that the cuPDLP-C option is to be removed in a future version.GitHub releases of ERGO-Code/HiGHS v1.7.0, v1.10.0 and v1.14.0 (release dates from the GitHub API) and the source files highs/lp_data/HighsOptions.h and docs/src/solvers.md at tag v1.14.0; GAMS, "GAMS 54 release notes", 54.1.0 (15 June 2026), gams.com/54/docs/RN_54.html. NVIDIA's cuOpt solves LP by PDLP, by a GPU barrier method with cuDSS factorizations, or by a CPU dual simplex, by default all three concurrently with the fastest answer returned. Its PDLP stops at a relative accuracy of \(10^{-4}\) by default. Its documentation grades \(10^{-2}\) as low accuracy, \(10^{-4}\) as regular, \(10^{-6}\) as high and \(10^{-8}\) as very high, and release 26.04 of April 2026 added single-precision and mixed-precision PDLP.NVIDIA cuOpt User Guide, release 26.08, "Introduction", "Convex optimization settings", "FAQ" and "Release notes" (the 26.04 entry "Add support for FP32 and mixed precision in PDLP"; GitHub release v26.04.00 of 9 April 2026), docs.nvidia.com/cuopt/user-guide/latest. Two of these facts deserve emphasis. A GPU LP solver is not necessarily a first-order method: COPT's GPU barrier and cuOpt's barrier run a sparse factorization on the device, and Section 7.5 explains what that costs in conditioning. And every one of these codes defaults to a looser tolerance for the first-order method than for the barrier or the simplex, which is the right default for a heuristic and the wrong one for a bound.
Five claims recur in discussions of these solvers, and the theory above disposes of them. The table sets each against the result that answers it.
| claim | what the theory says |
|---|---|
| "a \(10^{-4}\) relative tolerance means a \(10^{-4}\) error in the bound" | the three tolerances bound residuals, not the distance to the optimal value, and the dual residual is relative to \(1 + \lVert c \rVert\); the correction of Section 7.3 can be much larger than \(\varepsilon \lvert z^\star \rvert\) when the box is wide |
| "PDHG converges sublinearly on LP" | the \(O(1/k)\) rate is the general convex-concave rate; restarted PDHG converges linearly on every LP (Theorem 7.2.9), with a Hoffman constant that can be poor |
| "restarts are a heuristic" | fixed-frequency restarts at the right length are the optimal scheme and the adaptive tests trigger no later (Theorem 7.2.9); the primal-weight update is the heuristic part, and PDLP's authors say so |
| "Halpern is just averaging" | the Halpern iterate differs from the ergodic average on a generic LP, has a last-iterate residual bound that averaging does not provide (Theorem 7.2.11), and reflection halves its constant (Proposition 7.2.13) |
| "the GPU makes every LP faster" | it does not: the node LPs of a tree are small and warm-started, and there the dual simplex still wins (Section 7.4) |
Where this is used
Gurobi 13, COPT 8, HiGHS and cuOpt all run PDHG on the GPU today, as one algorithm of a portfolio for LP and, in cuOpt and COPT, for the root relaxation of a MILP. None of the global MINLP solvers of Section 5.3 documents a GPU or first-order path for its relaxations: BARON, SCIP, Couenne, ANTIGONE and MAiNGO all solve their polyhedral relaxations through CPU LP codes.The LP engines as the solvers' manuals describe them (BARON: CLP, CPLEX or Xpress; SCIP: SoPlex by default through its LP interface; Couenne: CLP through Bonmin and the Osi interface; ANTIGONE: CPLEX; MAiNGO: CLP or CPLEX) were not re-fetched for this section, so the statement made in the text is only the negative one: none of the five documents a GPU or first-order option for its relaxations.
What parallelizes
What parallelizes is the whole iteration: the two products are one thread or warp per row or column, the projections one thread per entry, the restart tests a few reductions, and the only sequential dimension is the iteration counter. What the method costs is accuracy per iteration, and a tree that wants to prune on its output has to buy the accuracy or correct the bound. The correction is the next subsection.
Safe bounds from inexact duals
A first-order method returns an approximate dual vector, and an approximate dual value is not a bound. Weak duality (Theorem 2.2.2) says that the dual function at any multiplier bounds the optimum, and Proposition 2.2.4(b) wrote that function in closed form for a linear program over a box. A PDHG iterate is not dual feasible before convergence, and the figure above shows its raw dual value sitting on the wrong side of the optimum for two hundred iterations. A branch and bound prunes on bounds, and a tree that prunes on a number that is not a bound is wrong rather than approximate. This subsection gives the theorem that repairs any dual vector into a valid bound at the cost of one matrix–vector product, and proves it. It shows how directed rounding, the choice of rounding every floating-point operation toward \(-\infty\) or toward \(+\infty\) instead of to the nearest representable number, makes the repaired number rigorous in floating point. It interprets the correction as a Lagrangian dual, extends it to the general LP form and to a frontier of nodes, and gives the two reference implementations the rest of the section uses. The theorem is the bridge between the first-order solvers of the previous subsection and the pruning rules of Section 3. Without it none of what follows in Section 7.4 would be correct.
The theorem
Theorem 7.3.1 (Neumaier and Shcherbina, 2004; the safe lower bound). Consider \(\min\{c^\top x : Ax \ge b,\ l \le x \le u\}\) with finite \(l, u\) and optimal value \(z^\star\). For every \(y \ge 0\), with \(r = c - A^\top y\),
\[z^\star \;\ge\; b^\top y \;+\; \sum_{j=1}^{n} \min\big(r_j l_j,\ r_j u_j\big). \tag{SB}\]The right-hand side of (SB) is the dual function (2.2.1) of Proposition 2.2.4(b). What is new here is the attribution, Corollary 7.3.3 and the batched form of Proposition 7.3.8.
Proof. Let \(x\) be feasible. Then \(c^\top x = r^\top x + y^\top A x \ge r^\top x + y^\top b\), because \(y \ge 0\) and \(Ax \ge b\) componentwise. For each \(j\) the function \(t \mapsto r_j t\) is linear, so on \([l_j, u_j]\) it is minimized at an endpoint, and \(r_j x_j \ge \min(r_j l_j, r_j u_j)\). Sum over \(j\) and take the infimum over feasible \(x\). ∎A. Neumaier and O. Shcherbina, "Safe bounds in linear and mixed-integer linear programming", Mathematical Programming 99 (2004). Neumaier's survey "Complete search in continuous global optimization and constraint satisfaction", Acta Numerica 13 (2004), records the motivating failure, that "even famous state-of-the-art solvers like CPLEX 8.0 (and many other commercial MILP codes) may lose an integral global solution of an innocent-looking mixed integer linear program", and that Jansson extended the certification to unbounded variables: C. Jansson, "Rigorous lower and upper bounds in linear programming", SIAM Journal on Optimization 14 (2004).
(What the correction term charges) The picture is this. Weak duality with the box dropped and only the rows dualized would demand \(r = 0\), which an inexact \(y\) never achieves. Keeping the box as a hard constraint instead, each nonzero reduced cost \(r_j\) is charged the worst it can do over the variable's interval, and the charge is exactly the term \(\min(r_j l_j, r_j u_j)\). When \(y\) is an optimal dual vector and \(r\) the corresponding reduced costs, complementary slackness (Theorem 2.2.6) puts each variable with \(r_j \ne 0\) at the bound that attains the minimum, so (SB) holds with equality and the bound is tight. Nothing is assumed about \(y\) beyond nonnegativity. The vector can come from a simplex code with tolerances, from an interior-point method stopped early, or from a PDHG iterate on a GPU. The price of a poor \(y\) is a weak bound, never an invalid one.
Corollary 7.3.2 (the maximization form). For \(\max\{c^\top x : Ax \le b,\ l \le x \le u\}\) and any \(y \ge 0\), with \(r = c - A^\top y\), \(z^\star \le b^\top y + \sum_j \max(r_j l_j, r_j u_j)\), which is (SB-max). With \(l = 0\) it reads \(z^\star \le b^\top y + \sum_j u_j \max(0, r_j)\).
Proof. Apply the theorem to \(\min\{-c^\top x : -Ax \ge -b,\ l \le x \le u\}\) and change signs. ∎
Corollary 7.3.3 (rigour in floating point). Compute \(r\) in interval arithmetic, so that each \(r_j\) is enclosed in an interval \([\underline r_j, \overline r_j]\) containing its exact value, and evaluate the right-hand side of (SB) with every operation rounded toward \(-\infty\), using \(\min(\underline r_j l_j, \underline r_j u_j, \overline r_j l_j, \overline r_j u_j)\) in place of \(\min(r_j l_j, r_j u_j)\). The number so computed is a rigorous lower bound on \(z^\star\).
Proof. Each replacement can only decrease the value, and rounding toward \(-\infty\) decreases it further. ∎This is the form in which the bound is used by the exact mixed-integer solvers: W. Cook, S. Dash, R. Fukasawa and M. Goycoolea, "Numerically safe Gomory mixed-integer cuts", INFORMS Journal on Computing 21 (2009); D. E. Steffy and K. Wolter, "Valid linear programming bounds for exact mixed-integer programming", INFORMS Journal on Computing 25 (2013), which compares it with two other ways of obtaining valid node bounds; W. Cook, T. Koch, D. E. Steffy and K. Wolter, "A hybrid branch-and-bound approach for exact rational mixed-integer programming", Mathematical Programming Computation 5 (2013); L. Eifler and A. Gleixner, "A computational status update for exact rational mixed integer programming", Mathematical Programming 197 (2023), and "Safe and verified Gomory mixed-integer cuts in a rational mixed-integer program framework", SIAM Journal on Optimization 34 (2024). SCIP 10's exact mode, built on this work, covers the linear case only (Section 5.6).
Corollary 7.3.4 (a rigorous infeasibility certificate). Apply (SB) with \(c = 0\): for any \(y \ge 0\), if \(b^\top y + \sum_j \min(-(A^\top y)_j l_j,\ -(A^\top y)_j u_j) > 0\), then \(\{x : Ax \ge b,\ l \le x \le u\}\) is empty.
Proof. The left-hand side is a lower bound on the minimum of the zero function over the feasible set, which is \(0\) if the set is nonempty. ∎
Any approximate Farkas vector, in particular the infimal displacement vector that PDHG produces on an infeasible node (Theorem 7.2.14), can be checked this way in one further product with directed rounding. Pruning a node as infeasible on a heuristic detection without this check is unsafe. Pruning it on the certificate is as rigorous as pruning on the bound.
Proposition 7.3.5 (the safe bound is a partial Lagrangian dual). Define \(D(y) = b^\top y + \sum_j \min(r_j l_j, r_j u_j)\) with \(r = c - A^\top y\). Then
\[D(y) \;=\; \min_{l \le x \le u}\ \big[\,c^\top x - y^\top(Ax - b)\,\big],\]so \(D\) is the Lagrangian dual function of the LP in which only the rows are dualized and the box is kept. \(D\) is concave and piecewise linear, \(\max_{y \ge 0} D(y) = z^\star\), and \(D\) is Lipschitz in the Euclidean norm with constant at most \(\max_{l \le x \le u} \lVert b - Ax \rVert\), which is at most \(\lVert b \rVert + \lVert A \rVert \max_{l \le x \le u} \lVert x \rVert\).
Proof. The inner minimization separates over coordinates, and each coordinate is a linear function minimized at an endpoint of its interval, which gives the formula. Weak duality is Theorem 7.3.1, and strong duality holds because the box is compact and the problem is linear, so the dual attains \(z^\star\). For the Lipschitz constant, let \(\bar x\) be a minimizer of the inner problem at \(y'\). Then \(D(y) - D(y') \le c^\top \bar x - y^\top(A\bar x - b) - c^\top \bar x + y'^\top(A\bar x - b) = (b - A\bar x)^\top(y - y')\), and \(\lVert b - A\bar x \rVert\) is at most the maximum of \(\lVert b - Ax \rVert\) over the box. Exchanging \(y\) and \(y'\) gives the other direction. ∎
(Why the orange curve behaves as it does) This explains the behaviour of the orange curve in the figure. The correction term is not an ad hoc penalty. It charges each reduced cost that \(y\) leaves the worst it can do over the box. The bound's looseness vanishes at \(y^\star\), and it shrinks at the rate at which \(y^k\) approaches \(y^\star\), not at the rate at which \(b^\top y^k\) approaches \(z^\star\). The reason is the Lipschitz constant of the proposition: since \(D(y^\star) = z^\star\), the looseness \(z^\star - D(y^k)\) is at most that constant times \(\lVert y^k - y^\star \rVert\), whatever the raw value is doing. The raw value \(b^\top y^k\) can lie on either side of \(z^\star\). The corrected value \(D(y^k)\) always lies on the right side. The reading also shows what makes the correction large: the width of the box. On a box of side \(10\) a reduced cost of \(10^{-3}\) costs \(10^{-2}\) per variable. On a box of side \(0.1\) it costs \(10^{-4}\). Every bound tightening of Section 2.6 therefore tightens every subsequent safe bound, before the first-order solver has done anything.
Proposition 7.3.6 (the general form: two-sided rows and infinite bounds). Consider (P-gen) of Definition 7.2.1, \(\min\{c^\top x : Kx \in [q^L, q^U],\ l \le x \le u\}\), where entries of \(q^L\) and \(l\) may be \(-\infty\) and entries of \(q^U\) and \(u\) may be \(+\infty\), with optimal value \(z^\star\). Let \(y \in \mathbb{R}^m\) satisfy \(y_i \ge 0\) whenever \(q^U_i = +\infty\) and \(y_i \le 0\) whenever \(q^L_i = -\infty\), so that \(p(y) = \sum_i \min(q^L_i y_i, q^U_i y_i)\) is finite with the convention \(0 \cdot (\pm\infty) = 0\), and let \(r = c - K^\top y\). Then
\[z^\star \;\ge\; p(y) + \sum_{j=1}^{n} \varphi_j(r_j),\qquad \varphi_j(r_j) = \min(r_j l_j,\ r_j u_j), \tag{SB-gen}\]where \(\varphi_j(r_j) = -\infty\) if \(l_j = -\infty\) and \(r_j > 0\), or \(u_j = +\infty\) and \(r_j < 0\). The right-hand side equals the dual objective of Definition 7.2.2 whenever the dual-infeasibility residual of that definition is zero, because then the projected reduced cost \(\hat r\) coincides with \(r\), and it is \(-\infty\) otherwise. An equality row has a free multiplier and contributes \(q_i y_i\). A free variable contributes \(-\infty\) unless \(r_j = 0\).
Proof. For feasible \(x\), \(c^\top x = r^\top x + y^\top K x\). Row by row, \(t \mapsto y_i t\) is linear on \([q^L_i, q^U_i]\), so \(y_i(Kx)_i \ge \min(y_i q^L_i, y_i q^U_i)\), and the sign condition on \(y_i\) makes the minimum the finite endpoint when the other is infinite. Column by column, \(r_j x_j \ge \varphi_j(r_j)\) for the same reason, with \(\varphi_j = -\infty\) exactly when the linear function is unbounded below on the variable's interval. Sum and take the infimum over feasible \(x\). ∎
(What the proposition says about the solvers) The proposition says two things about the solvers. A PDLP-type code already computes the right-hand side of (SB-gen) in exact arithmetic whenever it evaluates its dual objective at a point with zero dual residual, so what Corollary 7.3.3 adds is only the rounding. At other points PDLP's dual objective, evaluated at \(\hat r\) rather than \(r\), is a finite number that is not a bound. And the bound is useless, not merely weak, for a free or one-sided variable whose reduced cost has the wrong sign. That is one reason the global solvers of Section 5.3 bound every variable before relaxing, and it is why the finite-box case is the one that matters in a tree.
Two examples by hand
Two examples can be checked by hand. The first is the smallest: maximize \(x_1 + x_2\) subject to \(x_1 + x_2 \le 1\) and \(0 \le x \le 1\), whose optimum is \(1\). The single multiplier \(y\) on the row gives \(r = (1 - y, 1 - y)\). At \(y = 0.9\) the raw value \(b^\top y = 0.9\) is below the optimum and is not a bound. The correction adds \(u_j\max(0, r_j) = 0.1\) for each of the two variables, and the safe bound is \(0.9 + 0.1 + 0.1 = 1.1\), valid and loose. At \(y = 1\) the reduced costs vanish and the bound is exactly \(1\). At \(y = 0\) the bound is the box bound \(2\). At \(y = 1.5\), where \(y\) is dual feasible, raw value and safe bound coincide at \(1.5\), as weak duality says they must. The second example is the two-variable LP of the figure at \(45^\circ\), with \(c = (0.70711, 0.70711)\) and optimum \(z^\star = 55/(7\sqrt2) = 5.555839\). Its optimal dual puts multipliers \(c_1/7 = 0.10102\) on the first two rows and zero on the others, since solving \(A_B^\top y = c\) with \(A_B = \begin{pmatrix} 2 & 5 \\ 5 & 2 \end{pmatrix}\) gives \(y_1 = y_2 = c_1/7\). Take instead the inexact dual \(y = (0.1, 0.1, 0, 0)\). Then \(A^\top y = (0.7, 0.7)\), \(r = (0.00711, 0.00711)\) and \(b^\top y = 2.45 + 3.05 = 5.5\), below \(z^\star\) and so not a bound for the maximization. The correction \(\sum_j u_j\max(0, r_j) = 7 \times 0.00711 + 6 \times 0.00711 = 0.0924\) gives the safe bound \(5.5924\), valid and \(0.66\) percent loose. With \(y = 0\) the bound is \(7 \times 0.70711 + 6 \times 0.70711 = 9.1924\), the value of the objective at the far corner of the box, which is where the figure's orange curve starts. The listing checks the arithmetic.
# The safe bound by hand (max form).
#
# z* <= b.y + sum_j max(r_j l_j, r_j u_j) with r = c - A^T y, for any
# y >= 0, evaluated at a few duals y on the two examples of the text.
import numpy as np
def safe_bound(A, b, c, l, u, y):
r = c - A.T @ y
return b @ y + np.sum(np.maximum(r * l, r * u))
# the first example of the text:
# max x1 + x2 s.t. x1 + x2 <= 1, 0 <= x <= 1 (optimum 1)
A = np.array([[1., 1.]])
b = np.array([1.])
c = np.ones(2)
l = np.zeros(2)
u = np.ones(2)
for y in (0.0, 0.9, 1.0, 1.5):
raw = b @ np.array([y])
safe = safe_bound(A, b, c, l, u, np.array([y]))
print(f"first example y = {y:3.1f}: raw b.y = {raw:.2f} "
f"safe bound = {safe:.2f}")
# the two-variable LP of the plane figures at 45 degrees (optimum
# 5.555839); the exact dual is c_1/7 on rows 1 and 2
A = np.array([[2., 5.], [5., 2.], [-3., 4.], [1., -2.]])
b = np.array([24.5, 30.5, 11., 4.2])
l = np.zeros(2)
u = np.array([7., 6.])
c = np.array([np.cos(np.pi / 4), np.sin(np.pi / 4)])
for y in (np.zeros(4),
np.array([0.1, 0.1, 0., 0.]),
np.array([1 / 7, 1 / 7, 0., 0.]) * c[0]):
raw = b @ y
safe = safe_bound(A, b, c, l, u, y)
print(f"R1 y = ({y[0]:.5f}, {y[1]:.5f}, 0, 0): "
f"raw b.y = {raw:.4f} safe bound = {safe:.4f}")
first example y = 0.0: raw b.y = 0.00 safe bound = 2.00
first example y = 0.9: raw b.y = 0.90 safe bound = 1.10
first example y = 1.0: raw b.y = 1.00 safe bound = 1.00
first example y = 1.5: raw b.y = 1.50 safe bound = 1.50
R1 y = (0.00000, 0.00000, 0, 0): raw b.y = 0.0000 safe bound = 9.1924
R1 y = (0.10000, 0.10000, 0, 0): raw b.y = 5.5000 safe bound = 5.5924
R1 y = (0.10102, 0.10102, 0, 0): raw b.y = 5.5558 safe bound = 5.5558
The function is one matrix–vector product, a vector of maxima and two sums. Its cost is negligible next to the solve that produced \(y\), and every line of it is parallel across rows and columns.
R1 at 45 degrees: the safe bound from y = (0.1, 0.1, 0, 0)
y = (0.1, 0.1, 0, 0) -------------------------------+
| |
| one matrix-vector product | b^T y
v v
A^T y = (0.7, 0.7) 2.45 + 3.05 = 5.5,
| below z* = 5.555839:
| r = c - A^T y, not a bound
| c = (0.70711, 0.70711) |
v |
r = (0.00711, 0.00711) |
| |
| u_j max(0, r_j), u = (7, 6) |
v |
7 x 0.00711 + 6 x 0.00711 = 0.0924 |
| |
+-------------------> + <---------------------+
|
v
5.5 + 0.0924 = 5.5924, the safe bound:
valid, and 0.66 percent loose
Rigour in floating point
(Rounding modes on a core and on a device) Rigour needs one more ingredient. The listing above computes (SB) in ordinary floating point, so each operation rounds to nearest and the computed number can exceed the exact bound by a few units in the last place, that is, by a few multiples of the spacing between adjacent floating-point numbers at that magnitude. Corollary 7.3.3 asks instead for every operation to round in the safe direction. On a processor core the rounding direction is a mode of the floating-point unit, set in C++ by std::fesetround(FE_DOWNWARD) or FE_UPWARD under #pragma STDC FENV_ACCESS ON. The mode is thread state, so a batched kernel that spreads nodes over several threads must set it in every worker. A compiler that fuses a product and a sum into a single fused multiply-add changes the operation sequence and must be told not to (-ffp-contract=off in clang). On a GPU there is no cheap per-thread mode switch. Directed rounding is per operation through the intrinsics of the CUDA math library, __dadd_rd, __dadd_ru, __dmul_rd, __dmul_ru, __fma_rd and __fma_ru for double precision, with the same warning about contraction (--fmad=false for the CUDA compiler).NVIDIA, CUDA Math API Reference Manual, "Double precision intrinsics", CUDA 13 documentation; the interval-arithmetic pattern built on them is S. Collange, M. Daumas and D. Defour, "Interval arithmetic in CUDA", in GPU Computing Gems Jade Edition (2012). The rounding modes themselves are those of the IEEE 754 standard, exposed in C and C++ through <cfenv>.
Algorithm 7.3.7 (safe bound from an approximate dual, with directed rounding).
Algorithm 7.3.7 Safe bound from an approximate dual, with directed rounding
(max form; Corollaries 7.3.2 and 7.3.3)
Input A, b, c, a finite box [l, u], any y >= 0 (for instance the PDHG
dual average), rounding intrinsics
Output ub with z* <= ub rigorously
1. for each column j in parallel:
s_lo := sum_i A_ij y_i rounded down;
s_hi := the same sum rounded up # an interval for (A'y)_j
r_lo := c_j - s_hi (rounded down);
r_hi := c_j - s_lo (rounded up) # an interval for r_j
t_j := max( r_lo*l_j, r_lo*u_j, r_hi*l_j, r_hi*u_j ),
each product rounded up
2. ub := ( sum_i b_i y_i rounded up ) + ( sum_j t_j rounded up )
# a tree reduction in a fixed order
Invariant
each interval [s_lo, s_hi] and [r_lo, r_hi] encloses its exact
counterpart; each t_j, each sum and ub is an upper bound on the
exact value.
Cost
one sparse matrix-vector product with directed rounding and two
reductions; negligible next to the solve.
Parallel
the column sums are independent; the final sums are reductions and
must be done in a fixed order, or in integer arithmetic, if
bit-reproducibility is wanted.
CUDA sketch (double)
s_hi = __fma_ru(A_ij, y_i, s_hi); s_lo = __fma_rd(A_ij, y_i, s_lo);
t_j from __dmul_ru(r_hi, u_j) and its three companions;
compile with --fmad=false so that a*b + c is not contracted into a
fused multiply-add.
The enclosures of Algorithm 7.3.7 for one column j (max form)
sum_i A_ij y_i rounded down sum_i A_ij y_i rounded up
| |
v v
s_lo <= (A'y)_j <= s_hi
| |
+--------------+ +--------------+
\ /
\/
/\
/ \
+--------------+ +--------------+
| |
v v
r_lo <= r_j <= r_hi
c_j - s_hi c_j - s_lo
rounded down rounded up
| |
+-----------------+----------------+
|
v
t_j := max( r_lo*l_j, r_lo*u_j, r_hi*l_j, r_hi*u_j ),
each product rounded up >= max(r_j l_j, r_j u_j)
|
v
ub := ( sum_i b_i y_i rounded up ) + ( sum_j t_j rounded up )
>= z*
The minimization form uses the mirror-image rounding, toward \(-\infty\) for everything that must not exceed its exact value and toward \(+\infty\) for the enclosure of \(A^\top y\). The algorithm is written for a single node. In a tree the same computation is performed for every node of the frontier, and the next proposition says what changes between nodes.
A frontier of nodes
Proposition 7.3.8 (node bounds and the batched form). Let the nodes of a branch-and-bound tree share the rows \(Ax \ge b\) and differ in their boxes, so that node \(N\) has relaxation \(\mathrm{LP}_N = \min\{c^\top x : Ax \ge b,\ x \in B_N\}\) with \(B_N = [l^N, u^N] \subseteq B_{\mathrm{parent}}\) and value \(\bar z(N)\), and write \(D_N(y) = b^\top y + \sum_j \min(r_j l^N_j, r_j u^N_j)\) with \(r = c - A^\top y\). For every \(y \ge 0\):
(i) \(D_N(y) \le \bar z(N) \le \min\{f(x) : x \in \mathcal F \cap B_N\}\), so any nonnegative vector, converged or not, is a valid node bound.
(ii) \(D_N(y) \ge D_{\mathrm{parent}}(y)\): the parent's dual gives the child a bound at least as good as the parent's. The improvement in coordinate \(j\) is \(\lvert r_j \rvert\) times the distance by which the bound attaining the minimum moved, so it is zero when \(r_j = 0\).
(iii) If the number computed by Corollary 7.3.3 satisfies \(D_N(y) \ge z_{\mathrm{inc}} - \varepsilon\), then \(N\) contains no feasible point with \(f(x) < z_{\mathrm{inc}} - \varepsilon\) and may be pruned.
(iv) For a frontier of \(b\) nodes with duals \(y^1, \dots, y^b\), the \(b\) bounds are \(b\) independent evaluations of the same formula, each costing one product \(A^\top y^k\) with directed rounding plus \(O(m + n)\) work. With one shared \(y\), for instance the parent's, the reduced-cost interval is computed once and only the box sum differs between nodes, so the \(b\) bounds cost one product plus \(O(nb)\).
Proof. (i) is Theorem 7.3.1 applied to \(\mathrm{LP}_N\), followed by the relaxation inequality of Section 2.1. (ii) A linear function's minimum over a smaller interval is larger, coordinate by coordinate. (iii) follows from (i) and the rigour of Corollary 7.3.3. (iv) is a count. ∎
Proposition 7.3.8 (iv): the bounds of a frontier of b nodes
a dual per node one shared dual
(for instance the parent's)
y^1 y^2 ... y^b y
| | | |
v v v v
A^T y^1 A^T y^2 ... A^T y^b A^T y and the reduced-cost
| | | interval, once
v v v |
D_1 D_2 ... D_b +------+--- ... ---+
| | |
v v v
D_1 D_2 ... D_b
b products with directed one product; only the box
rounding, plus O(m + n) each sum differs: plus O(nb)
(When the parent's dual buys the children nothing) Statement (ii) has a consequence that is easy to miss. A child's box enters the bound only through nonzero reduced costs. If the parent's dual is exact and the branched variable has reduced cost zero at the parent, which is the usual case when the variable is basic and fractional, since a basic variable's reduced cost is zero at a simplex optimum (Section 3.1), then the parent's dual gives every child exactly the parent's bound and nothing more. The children's own duals are needed before the box tightening pays. The listing below shows both cases. It is a C++23 realization of Algorithm 7.3.7 over a frontier of five boxes of the two-variable LP, written in the minimization form \(\min -c^\top x\) subject to \(-Ax \ge -b\). The frontier is the root box \([0, 7] \times [0, 6]\) and the four children of branching on \(x\) at \(4 \mid 5\) and on \(y\) at \(2 \mid 3\). The duals are each node's exact dual perturbed in the seventh digit, as a floating-point solver would return them, and then, in a second call, the root's exact dual for every node. The two directed roundings are set with std::fesetround. A pass toward \(-\infty\) computes the lower enclosure of \(A^\top y\) and of \(b^\top y\). A pass toward \(+\infty\) computes the upper enclosure of \(A^\top y\). A final pass toward \(-\infty\) forms the reduced-cost interval, the worst-case products with the box ends, and the sums. Each sum is a std::transform_reduce, and each node is one element of a std::ranges::for_each over the batch. With a parallel execution policy or a thread pool each worker would have to set its own rounding mode.
The three passes of batched_safe_bound over the five nodes
pass 1 pass 2 pass 3
toward -inf toward +inf toward -inf
+---------------+---------------+-----------------------+
root | s_lo, by_lo | s_hi | r_lo, r_hi -> out[0] |
x<=4 | s_lo, by_lo | s_hi | r_lo, r_hi -> out[1] |
x>=5 | s_lo, by_lo | s_hi | r_lo, r_hi -> out[2] |
y<=2 | s_lo, by_lo | s_hi | r_lo, r_hi -> out[3] |
y>=3 | s_lo, by_lo | s_hi | r_lo, r_hi -> out[4] |
+---------------+---------------+-----------------------+
^ ^
switch to +inf switch to -inf
a row: one node, one element of a for_each, independent of the
other rows (the batch dimension)
s_lo, s_hi: A'y rounded down and up; by_lo: b'y rounded down
out[k] = by_lo + the sum over j of the least of r_lo l, r_lo u,
r_hi l, r_hi u: the safe lower bound of node k
// Safe lower bounds for a frontier of node LPs (Proposition 7.3.8).
//
// Bounds for a frontier of b node LPs that share the rows and differ in
// their boxes, from one approximate dual vector per node; the batch
// size b of the text is K here, since b is the right-hand side. Node k:
//
// min c'x s.t. A x >= b, l_k <= x <= u_k;
//
// for any y_k >= 0 and r = c - A'y_k:
//
// z_k >= b'y_k + sum_j min(r_j l_kj, r_j u_kj).
//
// Two directed roundings make the number rigorous: everything that must
// not exceed its exact value is rounded toward -inf, the enclosure of
// A'y toward +inf.
//
// Compile: clang++ -std=c++23 -ffp-contract=off safe_frontier.cpp
#include <algorithm>
#include <cfenv>
#include <cstdio>
#include <numeric>
#include <ranges>
#include <span>
#include <vector>
#pragma STDC FENV_ACCESS ON
// A row-major m x n; L, U, Y column-major.
struct Frontier {
int m, n, K;
std::vector<double> A, b, c, L, U, Y;
};
// sum_i A_ij y_i in the current rounding mode (one column of A'y);
// the reduction is the kernel's inner loop.
static double aty(const Frontier& f, int j, std::span<const double> y) {
auto rows = std::views::iota(0, f.m);
return std::transform_reduce(
rows.begin(), rows.end(), 0.0, std::plus<>{},
[&](int i) { return f.A[i * f.n + j] * y[i]; });
}
// Safe lower bounds for all K nodes: out[k] <= z_k is guaranteed
// whatever the accuracy of y_k.
static std::vector<double> batched_safe_bound(const Frontier& f) {
const int saved = std::fegetround();
std::vector<double> s_lo(f.n * f.K), s_hi(f.n * f.K);
std::vector<double> by_lo(f.K), out(f.K);
auto cols = std::views::iota(0, f.K);
// pass 1: lower enclosures of A'y and of b'y
std::fesetround(FE_DOWNWARD);
// independent per node: the batch (thread) dimension
std::ranges::for_each(cols, [&](int k) {
std::span<const double> y(&f.Y[k * f.m], f.m);
for (int j = 0; j < f.n; ++j) {
s_lo[k * f.n + j] = aty(f, j, y);
}
auto rows = std::views::iota(0, f.m);
by_lo[k] = std::transform_reduce(
rows.begin(), rows.end(), 0.0, std::plus<>{},
[&](int i) { return f.b[i] * y[i]; });
});
// pass 2: upper enclosure of A'y
std::fesetround(FE_UPWARD);
std::ranges::for_each(cols, [&](int k) {
std::span<const double> y(&f.Y[k * f.m], f.m);
for (int j = 0; j < f.n; ++j) {
s_hi[k * f.n + j] = aty(f, j, y);
}
});
// pass 3: r in [r_lo, r_hi], worst case over the box, sum
std::fesetround(FE_DOWNWARD);
std::ranges::for_each(cols, [&](int k) {
auto idx = std::views::iota(0, f.n);
auto worst_case = [&](int j) {
// <= exact r_j (c - (upper enclosure), rounded down)
const double r_lo = f.c[j] - s_hi[k * f.n + j];
// >= exact r_j (s_lo - c rounded down, then negated)
const double r_hi = -(s_lo[k * f.n + j] - f.c[j]);
const double l = f.L[k * f.n + j];
const double u = f.U[k * f.n + j];
// each product rounded down
return std::min({r_lo * l, r_lo * u, r_hi * l, r_hi * u});
};
const double tail = std::transform_reduce(
idx.begin(), idx.end(), 0.0, std::plus<>{}, worst_case);
out[k] = by_lo[k] + tail;
});
std::fesetround(saved);
return out;
}
int main() {
// The two-variable LP in min form: minimize -(x + y)/sqrt(2) over
// 2x+5y<=24.5, 5x+2y<=30.5, -3x+4y<=11, x-2y<=4.2, written as
// A x >= b with A = -(rows), b = -(rhs).
//
// Frontier: the root box [0,7]x[0,6] and the four children of
// branching on x at 4|5 and on y at 2|3.
//
// Duals: the exact dual of each node's own LP, perturbed in the
// seventh digit as a floating-point solver would return it (node 0
// exact); then the root dual for every node.
const double s = 0.70710678118654752;
Frontier f{4, 2, 5,
{-2, -5, -5, -2, 3, -4, -1, 2},
{-24.5, -30.5, -11, -4.2},
{-s, -s},
{}, {}, {}};
// root x<=4 x>=5 y<=2 y>=3
f.L = {0, 0, 0, 0, 5, 0, 0, 0, 0, 3};
f.U = {7, 6, 4, 6, 7, 6, 7, 2, 7, 6};
const double e = 1e-7;
f.Y = {s / 7, s / 7, 0, 0, // root
s / 5 + e, 0, 0, e, // x<=4
0, s / 2 - e, 0, 0, // x>=5
0, s / 5 + e, e, 0, // y<=2
s / 2 - e, 0, 0, e}; // y>=3
const char* name[] = {"root [0,7]x[0,6]", "x<=4", "x>=5", "y<=2",
"y>=3"};
auto own = batched_safe_bound(f);
for (int k = 0; k < f.K; ++k) {
f.Y[k * f.m + 0] = s / 7;
f.Y[k * f.m + 1] = s / 7;
f.Y[k * f.m + 2] = f.Y[k * f.m + 3] = 0;
}
auto root = batched_safe_bound(f);
std::printf("%-18s %-22s %s\n",
"node", "safe lb, own dual", "safe lb, root dual");
for (int k = 0; k < f.K; ++k) {
std::printf("%-18s %-22.15f %.15f\n", name[k], own[k], root[k]);
}
std::fesetround(FE_DOWNWARD);
volatile double one = 1.0, three = 3.0;
const double d = one / three;
std::fesetround(FE_UPWARD);
const double u = one / three;
std::fesetround(FE_TONEAREST);
std::printf("rounding check: 1/3 down %.17g, up %.17g,\n"
" difference %.3g (one ulp)\n",
d, u, u - d);
}
node safe lb, own dual safe lb, root dual
root [0,7]x[0,6] -5.555838995037160 -5.555838995037160
x<=4 -5.161881172661798 -5.555838995037160
x>=5 -5.480078204195747 -5.555838995037160
y<=2 -5.161882452661798 -5.555838995037160
y>=3 -5.480078324195748 -5.555838995037160
rounding check: 1/3 down 0.33333333333333331, up 0.33333333333333337,
difference 5.55e-17 (one ulp)
(Reading the listing's output) The exact node values, in the maximization form of the figure, are \(5.555838995\) at the root, \(5.161879503\) for the children \(x \le 4\) and \(y \le 2\) (at \((4, 3.3)\) and \((5.3, 2)\)), and \(5.480077554\) for \(x \ge 5\) and \(y \ge 3\) (at \((5, 2.75)\) and \((4.75, 3)\)). Every bound in the first column lies below its node's exact minimum. The root, whose dual is exact, is off by about \(10^{-15}\). The four children are off by \(1.7 \times 10^{-6}\), \(6.5 \times 10^{-7}\), \(3.0 \times 10^{-6}\) and \(7.7 \times 10^{-7}\), the price of the \(10^{-7}\) perturbation of the duals times the box widths. In the second column the root dual gives every child the root bound and nothing more, because the root's reduced costs vanish. This is Proposition 7.3.8 (ii), and it is the reason a batched tree needs a dual per node, or at least a few iterations per node from the parent's dual, before the children's boxes buy anything. The last line checks that the two rounding modes are in force: \(1/3\) rounded down and up differ by one unit in the last place. The kernel does three passes of \(O(\operatorname{nnz}(A) + m + n)\) work per node with two rounding-mode switches in all rather than one per operation. The loop over nodes is the batch dimension and is embarrassingly parallel, the term for tasks that need no communication with one another, one thread or warp per node on a device. Inside a node the sums are reductions whose order must be fixed if the result is to be reproducible bit for bit. A reduction whose result does not depend on the order of its partial sums is available at the price given after Proposition 6.2.19, and Proposition 6.2.18 says when a fixed tree suffices.Section 6.2 discusses what determinism costs a parallel tree, and Section 7.8 lists deterministic reductions among the open problems.
The frontier of the listing: the root box and four children
root [0,7] x [0,6]
5.555838995
|
+----------------+----------------+
| branch on x | branch on y
+-------+-------+ +-------+-------+
| x <= 4 | x >= 5 | y <= 2 | y >= 3
v v v v
[0,4] x [0,6] [5,7] x [0,6] [0,7] x [0,2] [0,7] x [3,6]
at (4, 3.3) at (5, 2.75) at (5.3, 2) at (4.75, 3)
5.161879503 5.480077554 5.161879503 5.480077554
safe bound from the node's own dual (perturbed in the seventh
digit): off by
1.7e-6 6.5e-7 3.0e-6 7.7e-7
safe bound from the root's exact dual: the root's bound and
nothing more, because the root's reduced costs vanish
5.555838995 5.555838995 5.555838995 5.555838995
values in the maximization form of the figure
Where this is used
In exact mixed-integer programming the bound is one of the standard ways of obtaining valid node bounds, and it is what SCIP's exact mode uses alongside exact LP solves. In rigorous global optimization it certifies the LP relaxations of GlobSol and of the COCONUT environment.R. B. Kearfott, "GlobSol user guide", Optimization Methods and Software 24 (2009); H. Schichl and A. Neumaier, "Interval analysis on directed acyclic graphs for global optimization", Journal of Global Optimization 33 (2005); A. Neumaier, O. Shcherbina, W. Huyer and T. Vinkó, "A comparison of complete global optimization solvers", Mathematical Programming 103 (2005). R. B. Kearfott, "Interval computations, rigour and non-rigour in deterministic continuous global optimization", Optimization Methods and Software 26 (2011), discusses what rigour costs. In the GPU codes of the next subsection it appears in two guises. Blin and coauthors evaluate their batched bounds at a dual tolerance of \(10^{-8}\) and subtract a safety margin before using a bound to tighten a variable, which is (SB-gen) of Proposition 7.3.6 with a margin for the residual rather than a directed rounding. The piecewise-linear solver of Guan and coauthors uses the PDHG dual objective directly as a node bound and validates the final answer afterwards rather than correcting each bound, which is faster and not rigorous.N. Blin, S. Gualandi, C. Maes, A. Lodi and B. Stellato, "Batched first-order methods for parallel LP solving in MIP", arXiv 2601.21990 (2026), whose margin is \(\Delta = \varepsilon(1 + \lvert c^\top x \rvert + \lvert \varphi(r) + \varphi(y) \rvert)\) in their notation; Y. Guan, S. Luo, P. Li, T. Chen and K. Xu, "B³-PWL: GPU-batched branch-and-bound for piecewise-linear optimization with SOS2 constraints", arXiv 2608.28988 (2026). And it is the orange curve of the PDHG figure.
What parallelizes
What parallelizes is all of it: one sparse product with directed rounding per node, independent across the frontier, with a fixed-order reduction at the end. What it costs is tightness proportional to the dual error times the box width. The single most effective way to buy that tightness back is the bound tightening of Section 2.6, which is why batched OBBT is among the first problems of Section 7.8 (problem 3), and why problem 1 there prunes on the safe bound throughout. One last caution concerns mixed precision. The iteration may run in single precision, as cuOpt's does when asked. The bound must then be computed in double precision with directed rounding from whatever \(y\) the iteration produced, and its looseness grows with the dual error that the lower precision leaves.
First-order methods inside a tree
A branch-and-bound tree solves thousands of small, similar linear programs. Each child differs from its parent by one bound, and the dual simplex method recovers optimality from the parent's basis in a handful of pivots (Proposition 3.1.12). A first-order method has no basis. It returns a primal point that is slightly infeasible and a dual vector that is slightly suboptimal, and it converges in a number of iterations that depends on the problem's geometry rather than on how close the start was. This subsection works out what each of these differences does to the tree. Warm starts come first. Inexact primal points, for incumbents and for branching, come second. The third topic is the effect of a loosened bound on the number of nodes, with the tolerance figure and two experiments. The fourth is the batching of sibling LPs into one GPU problem, with the safe bound as the pruning test, and the two batched tasks built so far, strong branching and bound tightening. The fifth is the architecture of the one production solver organized around these ideas. The subsection ends with the regime in which a first-order node solve beats the simplex method. The tree is the knapsack tree of Section 3.1 where a tree is needed. The knapsack maximizes, and the propositions say which convention they use.
Warm starts
(Why a warm start buys a first-order method little) The simplex method's advantage in a tree is the warm start, so that comes first. By Theorem 7.2.9 a restarted first-order method needs \(O\big(\frac{C}{\alpha}\log\frac{\operatorname{dist}(z^0, Z^\star)}{\varepsilon}\big)\) iterations. A starting point close to the child's solution therefore helps only through the distance inside the logarithm. Starting ten times closer saves \(\frac{C}{\alpha}\log 10\) iterations, a fixed amount that does not depend on the accuracy asked for and that is a small share of the whole at high accuracy. A dual simplex warm start from the parent's basis, by contrast, typically needs a handful of pivots whatever the accuracy, because the parent's basis is dual feasible for the child (Proposition 3.1.12) and the child's optimum is usually a few pivots away. The published GPU trees reflect this. The piecewise-linear solver of Guan and coauthors does not warm start between parent and child at all, "since PDHG lacks a basis representation" as its authors put it. cuOpt's release 26.02 improved its first-order method's primal and dual warm start, and Gurobi 13 accepts a primal-dual vector as a start for its PDHG, but how much either buys inside a tree is not published.Y. Guan, S. Luo, P. Li, T. Chen and K. Xu, "B³-PWL: GPU-batched branch-and-bound for piecewise-linear optimization with SOS2 constraints", arXiv 2608.28988 (2026); NVIDIA cuOpt User Guide, "Release notes", 26.02; Gurobi Optimizer Reference Manual 13.0, parameter LPWarmStart. Whether a child's sharpness constant is systematically better than its parent's, so that warm starts across siblings pay more than the logarithm suggests, is an open question listed in Section 7.8.
Warm-starting a child LP from its parent
dual simplex restarted first-order method
------------ ----------------------------
carries the parent's optimal carries the parent's iterate
basis, dual feasible for the z^0, at distance dist(z^0, Z*)
child (Proposition 3.1.12) from the child's solutions
| |
v v
typically a handful of pivots, O( (C/alpha) log(dist/eps) )
whatever the accuracy iterations (Theorem 7.2.9)
inside the logarithm: log(dist/eps) = log dist + log(1/eps)
ten times closer: log dist falls by log 10, which saves
(C/alpha) log 10 iterations, the same
amount for every accuracy eps
Inexact primal points
(Repair before use, and guards for branching) A first-order primal iterate at tolerance \(10^{-4}\) violates rows by up to that amount, and a point that violates rows is not a feasible solution. Its objective value may not be used as \(z_{\mathrm{inc}}\) until the point has been repaired, by the feasibility polishing of Theorem 7.2.19, by rounding followed by fix-and-propagate, the heuristic of Section 3.2 that fixes variables one at a time and propagates the bounds after each fixing, or by a problem-specific repair such as the SOS2 repair of the piecewise-linear solver, which restores the condition of Section 4.6 that at most two adjacent weights of a piecewise-linear model are nonzero. The repair costs little and the inaccuracy is harmless for guidance. Kempke and Koch show that first-order LP solutions at \(10^{-4}\) rather than \(10^{-6}\) relative tolerance lose nothing as guidance for fix-and-propagate heuristics. The looser solves are 1.5 times faster to compute on MIPLIB and up to 5 times faster on large instances. With them they solve an energy-system model with 243 million nonzeros and 8 million variables to a gap below 2 percent in under four hours. On the same model, in the words of their abstract, state-of-the-art commercial solvers cannot produce a feasible solution within two days of running time.N.-C. Kempke and T. Koch, "Fix-and-propagate heuristics using low-precision first-order LP solutions for large-scale mixed-integer linear optimization", Mathematical Programming Computation (2026), arXiv 2503.10344. The abstract names no solver; which commercial codes were run is stated in the body of the paper, which was not consulted here. Branching on an inexact point needs the same care as branching on any relaxed point. A variable whose value sits within the tolerance of a bound or of an integer should not be branched on. A cut cannot cut off a point that violates it by less than the LP feasibility tolerance. The branching-point rules of Section 3.5, which pull the split toward the midpoint and clamp it away from the bounds, protect a spatial tree from both. The one difference is that the inexactness is larger and systematic rather than at the level of rounding error, so the guards need to be set at the method's tolerance rather than at \(10^{-9}\).
A first-order primal point in the tree (tolerance 1e-4)
primal iterate: rows violated by up to 1e-4
| |
| for an incumbent | for guidance
v v
not a feasible solution: fix-and-propagate from it:
repair it first, by 1e-4 rather than 1e-6 loses
- feasibility polishing nothing (Kempke and Koch),
(Theorem 7.2.19), and the solves are 1.5x
- rounding, then faster on MIPLIB, up to 5x
fix-and-propagate, or on large instances
- a problem-specific repair,
such as the SOS2 repair
of B3-PWL
|
v
a feasible point: only now
may its value be used as z_inc
Slack in the bound and the size of the tree
The price of the first-order solve is paid in the bound. Section 7.3 made every first-order dual vector a valid bound. The price was looseness proportional to the dual error times the box width, and looseness in the pruning bound changes the number of nodes. The following definitions and propositions quantify how. They use the tree vocabulary of Section 3.1: a node is fathomed when the search disposes of it without branching, because its relaxation is infeasible, because its bound cannot beat the incumbent, or because nothing is left to branch on.
Definition 7.4.1 (bound slack; batched LP family). A node bound has relative slack \(\delta\) if the number used for pruning is \(\bar z(N)(1 + \delta)\) instead of the exact relaxation value \(\bar z(N)\) for a maximization, and \(\bar z(N)(1 - \delta)\) for a minimization. A batched LP family is a set of \(b\) linear programs \(\{(c^k, l^k, u^k, q^{L,k}, q^{U,k})\}_{k=1}^b\) that share the constraint matrix \(A\). The sibling nodes of a tree, which differ in their boxes, and the \(2n\) problems of optimality-based bound tightening, which differ in their objectives, are such families.
(The picture behind the growth of the tree) The picture behind the growth of the tree is this. Suppose every branching improves the bound of each child by the same amount \(g\). A node whose bound exceeds the incumbent by \(G\) needs about \(G/g\) more levels of branching before its descendants' bounds reach the incumbent and are pruned, and each level doubles the number of nodes, so closing a gap \(G\) costs about \(2^{G/g}\) nodes. A relative slack \(\delta\) in every bound asks the search to bring the loosened bound \((1 + \delta)\bar z(N)\), not \(\bar z(N)\), down to the incumbent. That adds about \(\delta\lvert z^\star \rvert\) to the gap every node must close and multiplies the count by \(2^{\delta\lvert z^\star \rvert/g}\). The proposition is the refinement of this picture to unequal gains, in the abstract branching model of Le Bodic and Nemhauser.
The picture behind Proposition 7.4.2 (equal gains g, maximization)
exact bound the tree nodes on
of the nodes the level
z_inc + G o 1
/ \
z_inc + G - g o o 2
/ \ / \
z_inc + G - 2g o o o o 4
: : : :
z_inc o o o o o o o o 2^(G/g) [1]
: : : :
about o o o o o o o o o o 2^(G/g) x [2]
z_inc - delta |z*| 2^(delta |z*| / g)
[1] the exact bounds reach the incumbent: pruned here
[2] the loosened bounds (1 + delta) zbar(N) reach it only here:
about delta |z*| more gap, delta |z*| / g more levels
each level lowers the bound by g and doubles the nodes, so
closing the gap G costs about 2^(G/g) nodes
Proposition 7.4.2 (tree growth under relative slack; the abstract branching model of Le Bodic and Nemhauser). In the abstract branching model, a branching on a variable with gains \((\ell, r)\) improves the bound by \(\ell\) in one child and by \(r\) in the other. The smallest tree that closes a gap \(G\) with a single variable type has size \(t(G) = 1 + t(G - \ell) + t(G - r)\), so \(t(G) = \Theta(\varphi^{G})\), where \(\varphi > 1\) is the root of \(\varphi^{-\ell} + \varphi^{-r} = 1\). If pruning uses a bound with relative slack \(\delta\), the gap that must be closed at every node grows by \(\delta\lvert z^\star \rvert\), and the tree grows by the factor \(\varphi^{\delta\lvert z^\star \rvert}\). For equal gains \(g\) the factor is \(2^{\delta\lvert z^\star \rvert/g}\).
Proof sketch. A node is pruned when its bound reaches the incumbent. With slack the bound must reach \(z^\star/(1 + \delta) \approx z^\star - \delta z^\star\) instead of \(z^\star\), for a maximization with \(z^\star > 0\), which is the extra \(\delta z^\star\) of gap in the picture above. The recurrence \(t(G) = 1 + t(G - \ell) + t(G - r)\) has the growth rate \(\varphi\) of its characteristic equation \(\varphi^{-\ell} + \varphi^{-r} = 1\), which is the content of the paper, and an extra gap of \(\delta z^\star\) multiplies \(t\) by \(\varphi^{\delta z^\star}\). For equal gains \(g\) the equation reads \(2\varphi^{-g} = 1\), so \(\varphi = 2^{1/g}\) and the factor is \(2^{\delta z^\star/g}\), the doubling per level of the picture. ∎P. Le Bodic and G. Nemhauser, "An abstract model for branching and its application to mixed integer programming", Mathematical Programming 166 (2017).
Proposition 7.4.3 (slack bounds the certifiable gap). Consider a maximization with \(z^\star > 0\), bounds used with relative slack \(\delta\), and the gap test "fathom \(N\) when \((1 + \delta)\bar z(N) \le (1 + \varepsilon) z_{\mathrm{inc}}\)". If \(\varepsilon < \delta\) and the incumbent is optimal, the gap test never fathoms a node whose box contains an optimal point. Termination can then come only from fathoming such nodes in some other way, by exhausting a finite tree or by solving them exactly. In particular a spatial branch and bound whose nodes are boxes of positive width, and whose only fathoming rules are the gap test and infeasibility, cannot terminate.
Proof. A node \(N\) whose box contains an optimal point has exact bound \(\bar z(N) \ge z^\star = z_{\mathrm{inc}}\), so \((1 + \delta)\bar z(N) \ge (1 + \delta) z_{\mathrm{inc}} > (1 + \varepsilon) z_{\mathrm{inc}}\), and it is not infeasible. Every branching of such a node produces a child whose box contains the same optimal point, so a node of this kind exists at every depth, and a spatial tree with the two fathoming rules has no way to dispose of it. ∎
Proposition 7.4.3: the boxes that hold an optimal point x*
maximization with z* > 0, eps < delta, and an optimal
incumbent: z_inc = z*
root box, x* in it
/ \
the other child a child with x*
/ \
the other child a child with x*
/ \
the other child a child with x*
:
at every depth
every box N that holds x* has zbar(N) >= z* = z_inc, so
(1 + delta) zbar(N) >= (1 + delta) z_inc > (1 + eps) z_inc:
the gap test never fathoms N, and N is not infeasible; with only
these two fathoming rules a spatial tree cannot terminate
(The tolerance figure's setup) The two propositions say that a bounding method with a fixed relative slack inflates the tree exponentially in the slack and makes any gap tolerance below the slack unreachable without exhaustion. The figure below measures the first effect on the knapsack family of the branch-and-bound figure of Section 3.1. Every node bound is multiplied by \(1 + \delta\) before the search sees it, which is what a first-order solve stopped early and corrected by Section 7.3 would return. The search is otherwise the exact best-bound search of that figure, so the count at \(\delta = 0\) is the same. Both axes are logarithmic, and \(\delta = 0\) sits to the left of an axis break. The default view has 12 items, seed 9 and no starting incumbent, with optimum \(z^\star = 293\) and capacity 171 of a total weight of 342. Its node counts, and those of the two comparison series the figure draws, are in the table.
| \(\delta\) | 0 | 1e-6 | 1e-5 | 1e-4 | 1e-3 | 3e-3 | 1e-2 | 3e-2 | 1e-1 |
|---|---|---|---|---|---|---|---|---|---|
| 12 items, seed 9, no incumbent | 37 | 37 | 37 | 39 | 41 | 43 | 81 | 265 | 1,865 |
| 15 items, seed 9, no incumbent | 67 | 67 | 67 | 67 | 71 | 83 | 177 | 867 | 8,849 |
(Reading the tolerance figure) For the default instance the factors are 1.1 at \(10^{-4}\), 2.2 at \(10^{-2}\) and 50 at \(10^{-1}\), and the last decade of \(\delta\) alone grows the tree 23-fold. For 15 items the factors are 2.6 at \(10^{-2}\) and 132 at \(10^{-1}\). Starting the default instance from the greedy incumbent 290 instead gives 35 nodes at \(\delta = 0\), 81 at \(10^{-2}\) and again 1,865 at \(10^{-1}\). The curve is flat until \(\delta\) approaches \(1/z^\star = 0.0034\), the dotted blue line, because the knapsack's values are integers. A loosening changes a pruning decision only when it carries a bound \(b < z_{\mathrm{inc}} + 1\) across the next integer above the incumbent, which needs \(b(1 + \delta) \ge z_{\mathrm{inc}} + 1\). The figure's take draws the practical conclusion for this instance: a bounding device ten times faster per node comes out ahead only up to \(\delta = 0.03\). Across the 60 instances and both starts the largest tree is 18,083 nodes, below the cap of 20,000 at which the figure would draw a red ring, and for 12 items the dotted grey rule marks the full enumeration tree of 8,191 nodes. The knapsack search terminates at every \(\delta\), including those far above any gap tolerance, because its tree is finite. Once every item is decided a node is a leaf and is fathomed by exhaustion, which is the escape that Proposition 7.4.3 allows and that a spatial tree does not have.
The tolerance curve is flat until delta approaches 1/z* = 0.0034
integer values: a loosening matters only when it carries a
bound b across the next integer above the incumbent
b z_inc + 1
------+-----------------------+--------------------> value
| |
+========> | b (1 + delta) < z_inc + 1:
| | pruned as before
| |
+===========================> b (1 + delta) >= z_inc + 1:
| the decision changes
for b near z* = 293 the loosening delta b reaches one unit at
delta = 1/z* = 0.0034
The same experiment on a correlated knapsack family, with real-valued pruning instead of the integer step, shows the exponential shape of Proposition 7.4.2 without the plateau, and the same slack on the bilinear example of Section 3.5 shows the edge of Proposition 7.4.3. The listing does both. It runs best-first branch and bound with the Dantzig bound, the greedy LP bound of the knapsack from Section 3.1, on ten random instances with 14 items, prunes a node when \((1 + \delta)\,\bar z(N) \le z_{\mathrm{inc}}\), and reports the geometric mean of the node counts. It then runs a spatial branch and bound on \(\max\, xy\) subject to \(2x + y \le 1.2\) on the unit square, with the McCormick bound of each box solved exactly by vertex enumeration and the wider side bisected, stopping when every open box satisfies \((1 + \delta)\bar z(B) \le (1 + \varepsilon) z_{\mathrm{inc}}\), and counts the boxes.
# How a looser node bound grows the tree.
#
# (1) Best-first branch and bound on 0/1 knapsacks with the Dantzig
# bound, pruning a node when (1 + delta) * bound <= incumbent, i.e. with
# a relative slack delta in every bound.
# (2) Spatial branch and bound on the bilinear example
# max x*y s.t. 2x + y <= 1.2 on [0, 1]^2 (Section 3.5),
# with the McCormick LP bound of each box solved exactly by vertex
# enumeration, the wider side bisected, and the search stopped when
# every open box has (1 + delta) * bound <= (1 + eps) * incumbent.
import heapq
import itertools
import numpy as np
def nodes_with_slack(w, v, C, delta):
"""Nodes processed by the knapsack search with slack delta."""
n = len(w)
order = np.argsort(-v / w)
w, v = w[order], v[order]
def bound(i, cap, val):
"""Dantzig: items by value per unit of weight,
fractional last item."""
for j in range(i, n):
if w[j] <= cap:
cap -= w[j]
val += v[j]
else:
return val + v[j] * cap / w[j]
return val
inc = 0.0
nodes = 0
heap = [(-bound(0, C, 0.0), 0, C, 0.0)]
while heap:
nb, i, cap, val = heapq.heappop(heap)
nodes += 1
# prune on the loosened bound
if (1 + delta) * (-nb) <= inc:
continue
if i == n:
inc = max(inc, val)
continue
for take in (1, 0):
if take and w[i] > cap:
continue
cap2, val2 = (cap - w[i], val + v[i]) if take else (cap, val)
# a partial fill is itself feasible
inc = max(inc, val2)
bd = bound(i + 1, cap2, val2)
if (1 + delta) * bd > inc:
heapq.heappush(heap, (-bd, i + 1, cap2, val2))
return nodes
def mccormick_bound(xl, xu, yl, yu):
"""Max w s.t. 2x + y <= 1.2, the box, and the two upper McCormick
planes; returns (w, the vertex (x, y, w)) or (-inf, None)."""
rows = [([2, 1, 0], 1.2),
([-1, 0, 0], -xl), ([1, 0, 0], xu),
([0, -1, 0], -yl), ([0, 1, 0], yu),
# w <= xu*y + yl*x - xu*yl, w <= xl*y + yu*x - xl*yu
([-yl, -xu, 1], -xu * yl), ([-yu, -xl, 1], -xl * yu)]
M = np.array([a for a, _ in rows], float)
r = np.array([t for _, t in rows], float)
best = (-np.inf, None)
# vertex enumeration in (x, y, w)
for i, j, k in itertools.combinations(range(len(rows)), 3):
B = M[[i, j, k]]
if abs(np.linalg.det(B)) < 1e-12:
continue
p = np.linalg.solve(B, r[[i, j, k]])
if np.all(M @ p <= r + 1e-9) and p[2] > best[0]:
best = (p[2], p)
return best
def boxes_with_slack(delta, eps):
"""Boxes processed by the spatial search on the bilinear example."""
inc = 0.0
nodes = 0
bd, p = mccormick_bound(0, 1, 0, 1)
heap = [(-bd, 0.0, 1.0, 0.0, 1.0)]
while heap:
nb, xl, xu, yl, yu = heapq.heappop(heap)
nodes += 1
bd, p = mccormick_bound(xl, xu, yl, yu)
# the relaxed (x, y) is feasible
inc = max(inc, p[0] * p[1])
if (1 + delta) * bd <= (1 + eps) * inc:
continue
if xu - xl >= yu - yl:
xm = 0.5 * (xl + xu)
kids = [(xl, xm, yl, yu), (xm, xu, yl, yu)]
else:
ym = 0.5 * (yl + yu)
kids = [(xl, xu, yl, ym), (xl, xu, ym, yu)]
for kd in kids:
b2, _ = mccormick_bound(*kd)
if (1 + delta) * b2 > (1 + eps) * inc:
heapq.heappush(heap, (-b2, *kd))
return nodes
deltas = [0, 1e-4, 1e-3, 3e-3, 1e-2, 2e-2, 5e-2]
rng = np.random.default_rng(1)
counts = {d: [] for d in deltas}
# ten random instances with 14 items
for s in range(10):
w = rng.integers(10, 100, 14).astype(float)
v = w + rng.integers(5, 25, 14)
C = 0.5 * w.sum()
for d in deltas:
counts[d].append(nodes_with_slack(w, v, C, d))
base = np.exp(np.mean(np.log(counts[0])))
print("knapsack, 14 items, 10 instances:")
print("geometric-mean nodes against the relative bound slack delta")
for d in deltas:
g = np.exp(np.mean(np.log(counts[d])))
print(f" delta = {d:<6g} nodes = {g:8.1f}"
f" ratio to delta = 0: {g / base:6.2f}")
print("bilinear example, spatial branch and bound with McCormick bounds:")
print("boxes against delta (delta <= eps only)")
for eps, ds in ((1e-2, (0, 1e-4, 1e-3, 3e-3, 5e-3, 7e-3, 9e-3, 1e-2)),
(1e-3, (0, 1e-4, 1e-3))):
runs = [f"delta = {d:g}: {boxes_with_slack(d, eps)}" for d in ds]
# three runs to a line
for row in range(0, len(runs), 3):
head = f" eps = {eps:<5g} " if row == 0 else " " * 15
line = "".join(f"{x:<20}" for x in runs[row:row + 3])
print((head + line).rstrip())
knapsack, 14 items, 10 instances:
geometric-mean nodes against the relative bound slack delta
delta = 0 nodes = 57.8 ratio to delta = 0: 1.00
delta = 0.0001 nodes = 60.3 ratio to delta = 0: 1.04
delta = 0.001 nodes = 65.0 ratio to delta = 0: 1.13
delta = 0.003 nodes = 77.2 ratio to delta = 0: 1.34
delta = 0.01 nodes = 153.8 ratio to delta = 0: 2.66
delta = 0.02 nodes = 369.8 ratio to delta = 0: 6.40
delta = 0.05 nodes = 1812.0 ratio to delta = 0: 31.35
bilinear example, spatial branch and bound with McCormick bounds:
boxes against delta (delta <= eps only)
eps = 0.01 delta = 0: 8 delta = 0.0001: 8 delta = 0.001: 8
delta = 0.003: 8 delta = 0.005: 10 delta = 0.007: 11
delta = 0.009: 13 delta = 0.01: 130
eps = 0.001 delta = 0: 13 delta = 0.0001: 13 delta = 0.001: 130
(Reading the listing's output) The quantity \(\ln(N(\delta)/N(0))/\delta\) is 98, 93 and 69 at \(\delta = 0.01\), \(0.02\) and \(0.05\). The growth is roughly exponential in \(\delta\) as the proposition predicts, with \(2^{\delta z^\star/g}\) fitting a per-level gain \(g\) of about one hundredth of \(z^\star\). A slack of one percent costs a factor 2.7 in nodes and five percent a factor 31. On the bilinear example at tolerance \(\varepsilon = 0.01\) the search takes eight boxes for every \(\delta\) up to \(0.003\), then 10, 11 and 13 boxes at \(\delta = 0.005\), \(0.007\) and \(0.009\), and 130 boxes at \(\delta = \varepsilon = 0.01\). At \(\varepsilon = 10^{-3}\) it takes 13 boxes at \(\delta = 0\) and \(10^{-4}\) and 130 at \(\delta = 10^{-3}\). The jump by an order of magnitude as \(\delta\) reaches \(\varepsilon\) is Proposition 7.4.3 at the edge of its hypothesis. At \(\delta = \varepsilon\) the gap test at a box containing the optimum is met only up to rounding, and for \(\delta > \varepsilon\) it is never met, so the listing does not run those cases: the proposition says they cannot stop. The knapsack runs are ten independent instances per \(\delta\) and the McCormick bounds are independent vertex enumerations, so both are parallel across instances and across boxes. Each tree search itself is sequential, one node at a time. The lesson for a solver is the one the figure's take states. A tree that prunes on bounds with slack \(\delta\) must be run with a gap tolerance comfortably above \(\delta\), or the slack must be removed. The safe bound of Section 7.3 removes it, because it has no slack but only a looseness that disappears as the dual converges, and so does solving the few hard nodes to high accuracy. The units matter here. A relative tolerance of \(10^{-4}\) on a bound of $720,000 is $72. The tax problem of Section 9 stops at a gap of a few dollars, so a first-order bound left at the solvers' default tolerance could never certify it, and rescaling the objective does not help, because the slack scales with it.
Batching a frontier
Batching is what makes the first-order node solve worth having despite all this. The following proposition is the formal content of the arithmetic-intensity argument of Section 7.1.
Proposition 7.4.4 (batched PDHG is PDHG on a block problem). Let \(b\) linear programs share the matrix \(A\) and differ in \((c^k, l^k, u^k, q^{L,k}, q^{U,k})\). Stack the iterates as \(X = [x^1 \cdots x^b] \in \mathbb{R}^{n \times b}\) and \(Y = [y^1 \cdots y^b] \in \mathbb{R}^{m \times b}\). Then \(b\) independent PDHG iterations, each with its own step sizes, primal weight and restart schedule, are exactly one PDHG iteration on the block-diagonal problem with matrix \(\operatorname{diag}(A, \dots, A)\), and the two matrix products of that iteration are the sparse-matrix times dense-matrix products \(AX\) and \(A^\top Y\).
Proof. The PDHG operator of the block problem is separable across the blocks, because the matrix is block-diagonal and the projections act columnwise. ∎
Proposition 7.4.4: b LPs that share A, run as one block problem
LP k, k = 1, ..., b: its own c^k, l^k, u^k, q^{L,k}, q^{U,k}, its
own step sizes, primal weight and restart schedule; the same A
the block problem the stacked iterates
+-----+-----+-----+-----+
| A | | | | X = [ x^1 x^2 ... x^b ]
+-----+-----+-----+-----+ n x b
| | A | | |
+-----+-----+-----+-----+ Y = [ y^1 y^2 ... y^b ]
| | | ... | | m x b
+-----+-----+-----+-----+
| | | | A |
+-----+-----+-----+-----+
diag(A, ..., A)
separable across the blocks: one PDHG iteration on it is b
independent iterations, and its two products are A X and A' Y,
sparse matrix times dense matrix; A is read from memory once per
product for all b problems instead of b times
(The block problem on a device) The consequence for a device is that the constraint matrix is read from memory once per product for all \(b\) problems instead of \(b\) times, so the arithmetic intensity of Definition 7.1.2 grows \(b\)-fold until the traffic of \(X\) and \(Y\) takes over. Blin, Gualandi, Maes, Lodi and Stellato implement exactly this. Their batch runs reflected Halpern iterations anchored at the start, with a step size, primal weight and restart schedule per column, the three restart tests of Algorithm 7.2.18, and the constant step \(0.998/\lVert A \rVert_2\). They apply it to the \(2p\) LPs of full strong branching, which differ in one bound each, and to the \(2n\) LPs of bound tightening, which differ in the objective.N. Blin, S. Gualandi, C. Maes, A. Lodi and B. Stellato, "Batched first-order methods for parallel LP solving in MIP", arXiv 2601.21990 (2026). The algorithm below adds to their batched iteration the one thing a tree needs from it: a safe bound per column (Proposition 7.3.8), evaluated periodically, and the retirement of a column as soon as its bound prunes it. The batch then spends its iterations on the nodes that are not yet decided.
Algorithm 7.4.5 (batched PDHG over a frontier of sibling node LPs, with safe bounds and retirement).
Algorithm 7.4.5 Batched PDHG over a frontier of b sibling node LPs
sharing A, with safe bounds and retirement
Input A (shared, in compressed sparse rows);
per-node data c^k, l^k, u^k, q^{L,k}, q^{U,k};
warm starts (x^k_0, y^k_0) from the parent;
node tolerance eps_node; incumbent z_inc; check period T
Output for each node k: a safe bound ub_k (Algorithm 7.3.7),
a flag prune_k, an approximate primal point x^k
1. X := [x^1_0 ... x^b_0], Y := [y^1_0 ... y^b_0];
per-column step eta_k := 0.998/||A||_2,
primal weights w_k := the parent's;
anchors Z^0 := (X, Y); per-column Halpern counters k_j := 0
2. repeat:
3. V := A' Y (one sparse-dense product);
X' := clip( X + (eta/w) (C - V), L, U ) # columnwise clips
4. W := A (2X' - X) (one sparse-dense product);
Y' := proj_rows( Y + (eta w) (W - Q) )
5. Halpern/reflection update per column with the column's own
counter k_j (Algorithm 7.2.18, step 4)
6. every T iterations:
per-column fixed-point residuals, restarts and
primal-weight updates;
per-column safe bound ub_k from the column's PDHG output y
(Algorithm 7.3.7);
mark prune_k if ub_k <= z_inc, and retire the column
(compact the batch)
7. until every column is retired (pruned, or its relative KKT
error <= eps_node)
Invariant
ub_k >= z*_k for every k at every check, whatever the accuracy
(Proposition 7.3.8 (i), in the max form); the columns are
independent PDHG runs, so each inherits the guarantees of
Algorithm 7.2.18.
Cost per step
two sparse-dense products with b right-hand sides, 4 nnz(A) b
flops for two reads of A, plus O((m + n) b).
Parallel
across columns and across the rows of A at once; the batch is
compacted when columns retire so that the products stay dense
in b. Siblings differ in one bound each, so only the bound
matrices differ, and they need not be stored in full (Blin and
coauthors store the differences).
The loop of Algorithm 7.4.5 on a batch of b columns
1. X, Y: one column per node, warm starts from the parents
|
v
3. V := A' Y; X' := clip(...) sparse-dense products, <-----+
4. W := A (2X' - X); Y' := ... b right-hand sides |
5. Halpern/reflection update, per column |
| |
v |
6. every T iterations, for each live column k: |
restarts; safe bound ub_k (Algorithm 7.3.7) |
ub_k <= z_inc --> prune_k, retire k |
relative KKT <= eps_node --> retire k |
compact the batch: retired columns leave X and Y -------------+
|
| every column retired
v
7. stop: for each node k, a safe bound ub_k, a flag prune_k
and an approximate primal point x^k
The listing is the reference implementation on the smallest instance that shows the behaviour. It is a 40-item knapsack LP relaxation with one row, so that the two products are a rank-one outer product and a dot product per column, with 256 nodes, each fixing a random 30 percent of the variables to 0 or 1 through its box. The exact LP value of every node comes from the greedy closed form of Section 3.1. The incumbent is set to 97 percent of the best node's LP value, so that an exact LP would prune 230 of the 256 nodes. Mode A runs plain PDHG in lockstep on all columns for 2,000 iterations and never retires a column. Mode B is Algorithm 7.4.5. It runs the reflected Halpern iteration with a counter per column and the three restart tests on the per-column fixed-point residual. Every ten iterations it computes the safe bound of Corollary 7.3.2 from the PDHG output \(y\), which is nonnegative where the Halpern iterate need not be. A column is retired when its bound is at most the incumbent or its relative KKT error is at most \(10^{-4}\), the node tolerance of the piecewise-linear solver and the default of cuOpt's first-order method. The "alive" column counts the columns entering a check, before that check's retirements.
# Batched PDHG over a frontier of K node LPs that share one constraint
# row, with the safe dual bound per column.
#
# K is the batch size b of the text; b is the knapsack capacity below.
# node k: max c^T x s.t. a^T x <= b, L_k <= x <= U_k
# (L_k, U_k encode the node's fixings)
# safe bound: for any y >= 0 and r = c - a y,
# z_k <= b y + sum_j max(r_j L_kj, r_j U_kj)
# Mode A: plain PDHG in lockstep, no column retired.
# Mode B: Algorithm 7.4.5: reflected Halpern PDHG with a counter per
# column, restarts on the fixed-point residual (constants 0.2 / 0.8 /
# 0.36), a safe bound per column every T iterations, and a column
# retired (compacted out of the batch) when its bound is <= the
# incumbent or its relative KKT error is <= eps_node.
import numpy as np
rng = np.random.default_rng(7)
n, K = 40, 256
c = rng.integers(10, 60, n).astype(float)
a = rng.integers(5, 40, n).astype(float)
b = 0.5 * a.sum()
def greedy_lp(l, u):
"""Exact LP value of a one-row knapsack relaxation with bounds
l <= x <= u."""
cap = b - a @ l
val = c @ l
for j in np.argsort(-c / a):
t = min(u[j] - l[j], max(cap, 0.0) / a[j])
val += c[j] * t
cap -= a[j] * t
return val
# the frontier: each node fixes a random 30% of the variables to 0 or 1
L = np.zeros((n, K))
U = np.ones((n, K))
for k in range(K):
fix = rng.random(n) < 0.3
vals = rng.integers(0, 2, n)
L[fix, k] = vals[fix]
U[fix, k] = vals[fix]
z_exact = np.array([greedy_lp(L[:, k], U[:, k]) for k in range(K)])
inc = 0.97 * z_exact.max()
prunable = int((z_exact <= inc).sum())
# tau * sigma * ||a||^2 < 1
tau = sigma = 0.9 / np.linalg.norm(a)
print(f"instance: n={n} items, K={K} nodes, "
f"30% of the variables fixed per node;")
print(f"incumbent {inc:.1f} = 0.97 * best node LP")
print(f"an exact LP would prune {prunable} of the {K} nodes")
def T_step(X, y, L, U):
"""One PDHG step on every column of the batch (two products
with a)."""
Xn = np.clip(X + tau * (c[:, None] - a[:, None] * y[None, :]), L, U)
return Xn, np.maximum(0.0, y + sigma * (a @ (2 * Xn - X) - b))
def safe(y, L, U):
"""The safe bound of Corollary 7.3.2, one per column."""
r = c[:, None] - a[:, None] * y[None, :]
return (b * y + (np.maximum(r, 0) * U).sum(0)
+ (np.minimum(r, 0) * L).sum(0))
print("\nMode A: plain PDHG in lockstep, nothing retired")
print(" iter pruned-by-safe-bound max(safe-exact)"
" mean primal infeas")
X = np.clip(0.5 * np.ones((n, K)), L, U)
y = np.zeros(K)
for it in range(1, 2001):
X, y = T_step(X, y, L, U)
s = safe(y, L, U)
if it in (10, 50, 100, 200, 500, 1000, 2000):
print(f"{it:5d} {(s <= inc).sum():10d}"
f" {(s - z_exact).max():8.2f}"
f" {np.maximum(a @ X - b, 0).mean():8.3f}")
assert np.all(s >= z_exact - 1e-9), (
"safe bound must never be below the LP value")
print("\nMode B: Algorithm 7.4.5 (reflected Halpern, per-column restarts,")
print(" retire at safe <= incumbent or KKT <= 1e-4)")
print(" iter alive pruned solved max(safe-exact) column-iterations")
print(" " * 36 + "over alive")
# 1e-4: the node tolerance of B3-PWL and cuOpt's PDLP default
Tchk, eps_node, bs, bn, ba = 10, 1e-4, 0.2, 0.8, 0.36
# alive columns only
alive = np.arange(K)
Xa = np.clip(0.5 * np.ones((n, K)), L, U)
ya = np.zeros(K)
X0, y0 = Xa.copy(), ya.copy()
kin = np.zeros(K)
r0 = np.full(K, np.inf)
rprev = np.full(K, np.inf)
S = np.full(K, np.inf)
pruned = np.zeros(K, bool)
solved = np.zeros(K, bool)
work = 0
for it in range(1, 2001):
La, Ua = L[:, alive], U[:, alive]
work += alive.size
# v = T(z)
Xn, yn = T_step(Xa, ya, La, Ua)
# Halpern weight, gamma = 1 (reflection)
lam = (kin + 1) / (kin + 2)
Xa = lam * (2 * Xn - Xa) + (1 - lam) * X0
ya = lam * (2 * yn - ya) + (1 - lam) * y0
kin += 1
if it % Tchk == 0:
# bounds come from the PDHG output yn >= 0
s = safe(yn, La, Ua)
S[alive] = np.minimum(S[alive], s)
prim = np.maximum(a @ Xn - b, 0.0) / (1 + abs(b))
gap = (np.abs(c @ Xn - s)
/ (1 + np.abs(c @ Xn) + np.abs(s)))
kkt = np.maximum(prim, gap)
# fixed-point residual of the column
res = np.sqrt(((Xa - Xn) ** 2).sum(0) / tau
+ (ya - yn) ** 2 / sigma)
r0 = np.where(np.isinf(r0), res, r0)
# the three restart tests
rs = ((res <= bs * r0)
| ((res <= bn * r0) & (res > rprev))
| (kin >= ba * it))
X0[:, rs] = Xn[:, rs]
y0[rs] = yn[rs]
Xa[:, rs] = Xn[:, rs]
ya[rs] = yn[rs]
kin[rs] = 0
r0[rs] = res[rs]
rprev = res
done = (S[alive] <= inc) | (kkt <= eps_node)
pruned[alive[S[alive] <= inc]] = True
solved[alive[(kkt <= eps_node) & ~(S[alive] <= inc)]] = True
if it in (10, 50, 100, 200, 500, 1000, 2000) or done.all():
worst = ((S[alive] - z_exact[alive]).max()
if alive.size else 0.0)
print(f"{it:5d} {alive.size:5d} {pruned.sum():6d}"
f" {solved.sum():6d} {worst:15.3f} {work:17d}")
# compaction: drop retired columns
keep = ~done
alive = alive[keep]
Xa, ya, X0, y0 = Xa[:, keep], ya[keep], X0[:, keep], y0[keep]
kin, r0, rprev = kin[keep], r0[keep], rprev[keep]
if alive.size == 0:
break
assert np.all(S >= z_exact - 1e-9), (
"safe bound must never be below the LP value")
left = alive.size
print(f"{K - left} of {K} columns retired by iteration {it}:")
print(f" {pruned.sum()} pruned by the safe bound"
f" (exact LP prunes {prunable}),")
print(f" {solved.sum()} solved to KKT <= {eps_node:g},"
f" {left} still alive;")
print(f" column-iterations {work} against {K * 2000}")
print(" for Mode A's 2000 lockstep iterations")
instance: n=40 items, K=256 nodes, 30% of the variables fixed per node;
incumbent 1061.5 = 0.97 * best node LP
an exact LP would prune 230 of the 256 nodes
Mode A: plain PDHG in lockstep, nothing retired
iter pruned-by-safe-bound max(safe-exact) mean primal infeas
10 228 23.47 3.989
50 230 2.56 0.337
100 230 1.42 0.098
200 230 0.26 0.030
500 230 0.05 0.001
1000 230 0.01 0.000
2000 230 0.01 0.000
Mode B: Algorithm 7.4.5 (reflected Halpern, per-column restarts,
retire at safe <= incumbent or KKT <= 1e-4)
iter alive pruned solved max(safe-exact) column-iterations
over alive
10 256 229 0 10.687 2560
50 26 230 2 0.323 3610
100 3 230 25 0.038 4150
110 1 230 26 0.038 4160
256 of 256 columns retired by iteration 110:
230 pruned by the safe bound (exact LP prunes 230),
26 solved to KKT <= 0.0001, 0 still alive;
column-iterations 4160 against 512000
for Mode A's 2000 lockstep iterations
(What retirement buys) After ten iterations the safe bound already prunes 228 (Mode A) or 229 (Mode B) of the 230 prunable nodes, with a worst looseness of 23.5 or 10.7 objective units. By iteration 50 both modes prune all 230, although the primal iterates are still infeasible by 0.34 units on average and would be useless as solutions. The bound is nevertheless exact as a certificate, and the assertion at the end of each mode checks that it never fell below the LP value of any node at any check. Retirement is what makes the batch cheap. Mode B's columns need 4,160 column-iterations in all, with every column retired by iteration 110 (230 pruned, 26 solved to \(10^{-4}\)), against 512,000 for Mode A's 2,000 lockstep iterations, a factor of 123, and the compaction, the repacking of the live columns into one contiguous block after each round of retirements, keeps the products dense in the live columns. The nodes that survive pruning are the expensive ones. The 26 unpruned nodes need up to 110 iterations to reach \(10^{-4}\): two had reached it by iteration 50 and twenty-five by iteration 100. A run of the same code to \(10^{-6}\) leaves four of them unconverged after 2,000 iterations under plain PDHG and under the Halpern variant alike, which is the small-sharpness case of Theorem 7.2.10 on nearly degenerate knapsack relaxations: items with nearly equal value-to-weight ratios make the optimal basis nearly non-unique, and that is the geometry that shrinks the sharpness constant of Proposition 7.2.8. A batched node solver therefore stops columns as soon as their bound crosses the incumbent and spends its iterations on the few columns that matter. The rule for when to stop a column that neither prunes nor converges is an open problem (Section 7.8). Per iteration the cost is two products with the shared row over the live columns, \(O(n)\) clips and the Halpern update per column, and every ten iterations one reduced-cost evaluation per column for the bound and the KKT test. Every line is data-parallel across columns and across the entries of a column, and the only sequential dimension is the iteration counter.
Two batched tasks have been built and measured so far, both at the root of a MILP tree. Full strong branching (Section 3.1) solves the \(2p\) LPs of tentatively fixing each fractional variable up and down, a family that differs in one bound per problem. Optimality-based bound tightening solves the \(2n\) LPs of minimizing and maximizing each variable over the relaxation, a family that differs in the objective. Blin and coauthors measure both on a B200 against Gurobi's dual simplex. The table gives their numbers.
| task | problem family | speedup in shifted geometric mean |
|---|---|---|
| full strong branching | combinatorial auctions | 12.0x |
| set covering | 34.6x | |
| maximum independent set | 80.8x | |
| facility location | 488.5x | |
| MIPLIB 2017 | mixed: "dual simplex remains faster on most problems" once warm-started and limited to 200 pivots per LP | |
| bound tightening (OBBT) | average over their set | 25.7x |
The two batched families of the table: what changes per column
full strong branching: 2p LPs, two per fractional variable x_j
+------------------------+------------------------+
| x_j fixed down | x_j fixed up | j = 1..p
+------------------------+------------------------+
A and c shared; the box differs in one bound
bound tightening (OBBT): 2n LPs, two per variable x_j
+------------------------+------------------------+
| min x_j | max x_j | j = 1..n
+------------------------+------------------------+
A and the box shared; the objective differs
They identify batch sizes of 32 to 512 and a large sparse matrix as the regime in which the device wins. For bound tightening they tighten the dual tolerance to \(10^{-8}\) and subtract the safety margin of Section 7.3 before a bound is used. The sequential engineering of OBBT in Section 2.6 was designed for one LP at a time. Candidates whose LP solution already sits at its bound are filtered out. The remaining LPs are ordered so that consecutive ones are similar and the dual simplex warm start is cheap. The duals of each solve are kept as Lagrangian variable bounds. Which of these parts carry over to a batch, and which of the \(2n\) problems to put in one, are open.N. Blin, S. Gualandi, C. Maes, A. Lodi and B. Stellato, arXiv 2601.21990 (2026); A. M. Gleixner, T. Berthold, B. Müller and S. Weltge, "Three enhancements for optimization-based bound tightening", Journal of Global Optimization 67 (2017). The precedent for a batched strong-branching expert on a GPU is the ADMM variant of full strong branching that Nair and coauthors used to generate training data for learned branching: V. Nair et al., "Solving mixed integer programs using neural networks", arXiv 2012.13349 (2020). In cuOpt the two tasks are the parameters CUOPT_MIP_BATCH_PDLP_STRONG_BRANCHING, which uses batched PDLP in place of the dual simplex for strong branching at the root (release 26.02), and CUOPT_MIP_BATCH_PDLP_RELIABILITY_BRANCHING (release 26.04), which does the same inside reliability branching, the rule of Section 3.1 that strong-branches only the variables whose branching history is not yet trusted; both are off by default.
cuOpt
cuOpt is the one production solver organized around the ideas of this subsection, and its development record is the clearest statement of where the field stands. The repository was created on 8 April 2025 under the Apache 2.0 licence. The solver's introduction says that it combines "GPU-accelerated primal heuristics for improving the primal bound with traditional CPU algorithms, including branch and bound, to improve the dual bound", and that "primal heuristics such as local search, feasibility pump, and feasibility jump run on the GPU". It adds that the solver "currently excels at finding high-quality feasible solutions quickly" and that "proving feasible solutions optimal remains under active development".NVIDIA cuOpt User Guide, release 26.08, "Introduction", docs.nvidia.com/cuopt/user-guide/latest/introduction.html; repository github.com/NVIDIA/cuopt. The parameter names are those of "MIP settings" in the same guide. The table collects the release notes. The dates are the release dates on PyPI or GitHub.
| release | date | LP side | MILP side |
|---|---|---|---|
| (open source release) | 8 Apr 2025 (repository created) | PDLP on the GPU | GPU primal heuristics (Feasibility Jump, local search, feasibility pump) with branch and bound on the CPU: the design the introduction describes, already present (the 25.05 notes list only Feasibility Jump fixes) |
| 25.05 | 29 May 2025 | PDLP and dual simplex concurrent; crossover from the PDLP point | first C API for LP and MIP; Feasibility Jump fixes |
| 25.10 | 14 Oct 2025 | barrier method with cuDSS on the GPU; all three LP methods concurrent | Papilo presolve, on by default; parallel best-first and diving threads on the CPU; node presolve |
| 25.12 | 11 Dec 2025 | concurrent root relaxation for MIP (PDLP, barrier, dual simplex) | RINS; factorization reuse in the tree |
| 26.02 | 11 Feb 2026 | improved primal and dual warm start of PDLP | root cuts (Gomory, MIR, knapsack, strong Chvátal–Gomory); parallel reliability branching; batched PDLP for strong branching at the root; experimental deterministic B&B (GPU heuristics excluded) |
| 26.04 | 9 Apr 2026 | FP32 and mixed-precision PDLP | batched PDLP in reliability branching; clique and implied-bound cuts |
| 26.06 | 9 Jun 2026 | convex quadratic and second-order-cone constraints in the barrier | B&B workers with their own heaps that steal nodes from one another (about 3x more nodes explored, vendor figure); symmetry detection; flow-cover cuts |
| 26.08 | 6 Aug 2026 | multi-GPU PDLP with METIS partitioning (2.5x to 8.8x on eight B200s, vendor figure) | zero-half cuts; recursive RINS |
Two corrections to earlier statements about this solver are worth recording. The October 2025 version of this post reported, from a developers' forum reply of 2025 that could not be located again, that the solver had no presolve or cuts. A bound-propagation presolve existed before release 25.10. A Papilo-based presolve, on by default, dates from 25.10 (October 2025), and root cutting planes from 26.02 (February 2026), with seven further cut families and work stealing by 26.08. The remark was true of early 2025 at most and is not true now.NVIDIA cuOpt User Guide, "Release notes" 25.05 to 26.08, docs.nvidia.com/cuopt/user-guide/latest/release-notes.html; release dates from the GitHub releases of NVIDIA/cuopt (v26.02.00 on 11 February 2026, v26.04.00 on 9 April 2026, v26.06.00 on 9 June 2026, v26.08.00 on 6 August 2026) and from the PyPI upload times of cuopt-cu12 (25.05.0 on 29 May 2025, 25.10.0 on 14 October 2025, 25.12.0 on 11 December 2025). The 25.05 and 25.08 notes already mention a "conditional bounds presolve" and a "load balanced bounds presolve". The repository creation date is from the GitHub API. And the division of labour is the one the table of Section 7.1 predicts. The dual bound is produced only by the CPU tree with exact LP solves, the primal bound by anyone, and the first-order method enters the tree only where it can be batched. The developers' own paper on the heuristic side describes PDLP used as an approximate LP solver with a one-second limit and warm starts, a GPU feasibility pump, Feasibility Jump, fix-and-propagate and a probing cache. Its measured results on the MIPLIB 2017 benchmark set are quoted in Section 7.7.A. Çördük, P. Sielski, A. Boucher and K. Aatish, "GPU-accelerated primal heuristics for mixed integer programming", arXiv 2510.20499 (2025).
cuOpt's division of labour between the GPU and the CPU
GPU CPU
+---------------------------+ +---------------------------+
| primal heuristics: local | | branch and bound with |
| search, feasibility pump, | | exact LP solves |
| Feasibility Jump; PDLP as | +---------------------------+
| an approximate LP solver | | |
+---------------------------+ | v
| | the dual bound:
v | the tree only
the primal bound <--------------+
(by anyone)
batched PDLP --> into the tree only where it can be batched:
strong branching at the root (26.02) and
reliability branching (26.04),
both off by default
Beyond cuOpt, the published uses of first-order LP inside a tree are few and structured. The piecewise-linear solver of Guan and coauthors is a GPU-batched branch and bound for problems with SOS2 constraints. It solves batches of node LPs with cuPDLPx at \(10^{-4}\) and runs a batched feasibility pump. It fills each batch to 85 percent of device memory and expands the frontier breadth-first when the queue falls below 90 percent of the batch size. It reports a 9.25 times geometric-mean speedup over cuOpt on 43 synthetic instances, and 10.3 times over SCIP and 5.1 times over HiGHS on valve-point unit-commitment instances, while Gurobi remains about 7 times faster than it on the synthetic set. Sharadga and Mohammadi use PDHG for the LP relaxations of a branch and bound for unit commitment on systems with 4,224 to 6,717 buses. Liu and Lodi process batches of nodes on the device, with padded node tensors, to certify optimal \(k\)-sparse generalized linear models, and report speedups of one to two orders of magnitude over a CPU implementation. Lucas, Meng and Mazumder give a GPU nonlinear branch and bound for sparse regression whose node relaxation is solved by an alternating-direction method. They report runtime improvements over a CPU implementation of their method, over a specialized branch and bound and over commercial MIP solvers, without a factor in the abstract. Both problem classes share a fixed matrix across all nodes. CHAP, the winner of the 2026 computational competition on GPU primal heuristics, coordinates a GPU tabu search and cuPDLPx with CPU fix-and-propagate and feasibility pump through a shared solution pool. Its count of instances with a solution found, against Gurobi and cuOpt, is in Section 7.7, which gives the table. It proves nothing, and does not claim to.Y. Guan, S. Luo, P. Li, T. Chen and K. Xu, arXiv 2608.28988 (2026); H. Sharadga and J. Mohammadi, "GPU-accelerated optimization solver for unit commitment in large-scale power grids", arXiv 2512.06715 (2025); J. Liu and A. Lodi, "From sequential nodes to GPU batches: parallel branch and bound for optimal k-sparse GLMs", arXiv 2605.22188 (2026); R. Lucas, X. Meng and R. Mazumder, "A GPU-accelerated nonlinear branch-and-bound framework for sparse linear models", arXiv 2602.04551 (2026); G. K. Tjusila, A. Hoen, N.-C. Kempke, G. Mexi, T. Berthold, A. Gleixner, T. Koch and S. Pokutta, "CHAP: a hybrid GPU-CPU heuristic for MIP", arXiv 2605.05086 (2026); MIP Workshop 2026, Computational Competition: GPU-Accelerated Primal Heuristics for MIP, mixedinteger.org/2026/competition (results page). All five papers are preprints, and the speedups are the authors' own measurements. No published work reports a complete GPU-bounded branch and bound for general MILP, let alone for nonconvex MINLP, that beats a CPU solver to a proof of optimality, and cuOpt's documentation says as much.
Regimes
The evidence of this subsection and the last two sorts into a short table of regimes.
| situation | ahead in the cited evidence | evidence |
|---|---|---|
| one LP with 1e8 to 1e9 nonzeros, where a factorization does not fit in memory | first order, GPU or CPU | PDLP journal version: 8 of 11 huge LPs against 3 for Gurobi's barrier; Mittelmann's INFORMS talk of 28 Oct 2025: among the GPU codes, cuPDLP-C alone finishing all six of Hinder's instances |
| one large sparse LP to 1e-6 | GPU codes, by a small factor | Mittelmann's LP feasibility page of 16 Sep 2026: the best GPU codes (HPR-LP-C, cuOpt and COPT's GPU barrier; not all first order) 1.3x to 1.7x faster than the best CPU code in shifted geometric mean, 1 to 3 fewer solved |
| one small or medium LP to 1e-8 | CPU barrier, simplex | cuPDLP.jl paper: about 1 s of fixed overhead; cuPDLP-C paper (Lu, Yang et al. 2023; the authors' numbers): COPT 383 of 383 against cuPDLP-C 369, in shifted means of 3.11 s and 18.53 s |
| hundreds of sibling LPs sharing \(A\) (strong branching, OBBT) | batched first order | Blin et al. (authors' numbers): 12x to 489x over Gurobi's dual simplex for strong branching on four structured families; batches of 32 to 512 |
| node LPs warm-started from the parent | dual simplex | Blin et al. on MIPLIB with 200-pivot warm starts; no basis to warm start PDHG from (B3-PWL) |
| a tree whose gap tolerance is below the relative slack of its corrected first-order bounds | exact bounds, or exhaustion of a finite tree; a spatial tree of boxes of positive width, fathomed only by the gap test and infeasibility, cannot terminate | Propositions 7.4.2 and 7.4.3 (a gap tolerance below the slack is unreachable without exhaustion); the tolerance figure |
Where this is used
In October 2026 the first-order node solve is in production inside a tree in exactly one place: cuOpt's batched strong branching at the root and batched reliability branching below it, both off by default. At the root, cuOpt and COPT also run the first-order method concurrently with the barrier and the simplex. Every other use is a preprint on a structured problem class. The research direction this subsection leads to, and Section 7.8 states as open problems with their first experiments, is a nonconvex tree whose node relaxations are McCormick LPs that share their rows except for the box-dependent envelope planes. They would be bounded in batches by Algorithm 7.4.5, with warm starts from the parent, the safe bound as the pruning test, and bound tightening run as a batch before the bounding.
What parallelizes
What parallelizes is the bounding of the whole frontier at once, the batched tightening, and the heuristics. What does not is the chain of branching decisions, the incumbent lag that a batch introduces, and the few hard nodes whose bounds sit within the gap and need near-exact solves. Under every method proposed so far those hard nodes dominate the time.
GPU interior point and NLP
The last three subsections moved the linear relaxation to the device. A nonconvex MINLP solver also solves nonlinear programs. It solves a local NLP for an incumbent at almost every node. It solves sub-NLPs inside the heuristics of Section 3.2. In the convex case, and in the conic corner of Section 4.8, it solves convex NLP relaxations whose values are the bounds. Section 1.3 described the local method that does this work. An interior-point method replaces the inequalities by a barrier and takes Newton steps on the perturbed optimality conditions. Almost all of its time goes into factorizing one symmetric indefinite matrix per step. That factorization needs numerical pivoting. Pivoting is a sequential chain of data-dependent decisions, and the table of Section 7.1 put it on the side of what does not map. This subsection explains the one device that moves it. A block elimination turns the indefinite system into a positive definite one, which a GPU factorizes with a pivot order fixed in advance. The price is paid in conditioning, and the conditioning sets the accuracy such a solve can reach. The reader needs this for two reasons. It is how the production GPU barrier codes of COPT and of NVIDIA's cuOpt work. And it fixes the number that matters most for a MINLP tree: a GPU interior-point solve returns a point accurate to about \(10^{-4}\) to \(10^{-6}\), which is an incumbent and not a bound.
The Newton system and its condensation
Write the problem of Section 1.3 with its inequalities turned into equalities by slacks,
\[\min_{x,\, s}\ f(x) \quad \text{s.t.} \quad g(x) = 0,\qquad h(x) + s = 0,\qquad s \ge 0,\]with \(g : \mathbb{R}^n \to \mathbb{R}^{m_e}\) the equalities and \(h : \mathbb{R}^n \to \mathbb{R}^{m}\) the inequalities. The multipliers are \(\lambda \in \mathbb{R}^{m_e}\) and \(\nu \in \mathbb{R}^{m}\), and the Lagrangian is \(L(x, s, \lambda, \nu) = f(x) + \lambda^\top g(x) + \nu^\top (h(x) + s)\). This subsection follows the lettering of the interior-point literature, with \(\lambda\) on the equalities and \(\nu \ge 0\) on the inequalities, rather than the convention of Section 1.3, where \(\lambda \ge 0\) multiplies the inequalities.
Definition 7.5.1 (barrier subproblem; Newton, augmented and condensed KKT systems). For a barrier parameter \(\mu > 0\) the barrier subproblem is \(\min f(x) - \mu \sum_i \log s_i\) subject to \(g(x) = 0\) and \(h(x) + s = 0\). Its first-order conditions are \(\nabla_x L = 0\), \(g(x) = 0\), \(h(x) + s = 0\) and \(s_i \nu_i = \mu\) for every \(i\). Let \(W = \nabla^2_{xx} L\), \(J_g = \nabla g(x)^\top\), \(J_h = \nabla h(x)^\top\), \(S = \operatorname{diag}(s)\), \(V = \operatorname{diag}(\nu)\) and \(\Sigma = S^{-1} V = \operatorname{diag}(\nu_i / s_i)\). Let \(e\) be the vector of ones and write the four residuals as \(r_d = \nabla_x L\), \(r_s = \nu - \mu S^{-1} e\), \(r_p = h(x) + s\) and \(r_e = g(x)\). The Newton system of the four conditions in the unknowns \((\Delta x, \Delta s, \Delta \nu, \Delta \lambda)\) is
\[\begin{pmatrix} W & 0 & J_h^\top & J_g^\top \\ 0 & \Sigma & I & 0 \\ J_h & I & 0 & 0 \\ J_g & 0 & 0 & 0 \end{pmatrix} \begin{pmatrix} \Delta x \\ \Delta s \\ \Delta \nu \\ \Delta \lambda \end{pmatrix} = - \begin{pmatrix} r_d \\ r_s \\ r_p \\ r_e \end{pmatrix}. \tag{7.5.1}\]Its second block row is the linearization \(\nu_i \Delta s_i + s_i \Delta \nu_i = \mu - s_i \nu_i\) of the complementarity condition, divided by \(s_i\). Solving that row for \(\Delta s = -\Sigma^{-1}(r_s + \Delta \nu)\) and substituting into the third block row gives the augmented KKT system, a symmetric indefinite system in the remaining unknowns,
\[\begin{pmatrix} W & J_h^\top & J_g^\top \\ J_h & -\Sigma^{-1} & 0 \\ J_g & 0 & 0 \end{pmatrix} \begin{pmatrix} \Delta x \\ \Delta \nu \\ \Delta \lambda \end{pmatrix} = - \begin{pmatrix} r_d \\ \tilde r_p \\ r_e \end{pmatrix}, \qquad \tilde r_p = r_p - \Sigma^{-1} r_s = h(x) + s - V^{-1}(S \nu - \mu e). \tag{7.5.2}\]The condensed KKT matrix is \(K_c = W + J_h^\top \Sigma J_h\), the matrix obtained by eliminating \(\Delta \nu\) from (7.5.2) as well.
(Three remarks on the definition) The definition calls for three remarks. The elimination of \(\Delta s\) is what puts \(-\Sigma^{-1}\) in the middle block of (7.5.2) in place of a zero, and (7.5.2) is the system a CPU solver factorizes. The matrix \(W\) is the Hessian of the Lagrangian. For a nonconvex problem it is indefinite, even on the null space of the constraints, and a line-search interior-point method adds a multiple \(\delta_w I\) to it until the step is a descent direction. The test for "until" is the inertia of the matrix, its counts of positive, negative and zero eigenvalues, which Proposition 7.5.3 treats. And \(\Sigma\) is the one object in the system that changes character along the iteration. On the central path its entries range from order \(\mu\) to order \(1/\mu\), and everything in this subsection follows from that.
Proposition 7.5.2 (condensation by block elimination). Let \(\Sigma \succ 0\). Eliminating \(\Delta\nu = \Sigma\,(J_h \Delta x + \tilde r_p)\) from (7.5.2) gives the equivalent system
\[\begin{pmatrix} K_c & J_g^\top \\ J_g & 0 \end{pmatrix} \begin{pmatrix} \Delta x \\ \Delta \lambda \end{pmatrix} = - \begin{pmatrix} r_d + J_h^\top \Sigma \tilde r_p \\ r_e \end{pmatrix}, \qquad K_c = W + J_h^\top \Sigma J_h, \tag{7.5.3}\]which is still indefinite when \(m_e > 0\). Two further eliminations make it positive definite.
(i) (Equality relaxation; LiftedKKT.) Replace each equality \(g_j(x) = 0\) by the pair \(-\tau \le g_j(x) \le \tau\) with a small \(\tau > 0\). Then \(m_e = 0\) and the second block row of (7.5.3) disappears. Let \(\tilde J_h\) stack the Jacobians of the \(m\) inequalities and of the \(2 m_e\) relaxed rows, let \(\tilde\Sigma\) be their barrier diagonal, and let \(\tilde r_d\) and \(\tilde r_p\) be the residuals of the relaxed problem. The system is \(K_\tau \Delta x = -(\tilde r_d + \tilde J_h^\top \tilde\Sigma \tilde r_p)\) with
\[K_\tau = (W + \delta_w I) + \tilde J_h^\top \tilde\Sigma \tilde J_h .\]Let \(\tilde J_A\) be the rows of \(\tilde J_h\) that are active at the solution. If \(W + \delta_w I\) is positive definite on \(\ker \tilde J_A\), then \(K_\tau\) is positive definite whenever the active entries of \(\tilde\Sigma\) are large enough. On the central path they are of order \(1/\mu\), so this is the regime near the solution. For any \(\delta_w > \max\{0, -\lambda_{\min}(W)\}\) the matrix \(K_\tau\) is positive definite outright, which is the fallback the inertia correction uses.
(ii) (Augmented Lagrangian; HyKKT.) For \(\gamma > 0\) add \(\gamma J_g^\top\) times the second block row of (7.5.3) to the first. With \(K_\gamma = K_c + \gamma J_g^\top J_g\) the system becomes
\[K_\gamma \Delta x + J_g^\top \Delta\lambda = b_1 + \gamma J_g^\top b_2,\qquad J_g \Delta x = b_2,\]where \(b_1, b_2\) are the two right-hand sides of (7.5.3). If \(K_c\) is positive definite on \(\ker J_g\), then \(K_\gamma\) is positive definite for all large \(\gamma\). The solution is \(\Delta\lambda = S_\gamma^{-1}\big(J_g K_\gamma^{-1}(b_1 + \gamma J_g^\top b_2) - b_2\big)\) with the Schur complement \(S_\gamma = J_g K_\gamma^{-1} J_g^\top\), followed by \(\Delta x = K_\gamma^{-1}(b_1 + \gamma J_g^\top b_2 - J_g^\top \Delta\lambda)\).
Proof. The elimination of \(\Delta\nu\) is the second block row of (7.5.2) solved for \(\Delta\nu\) and substituted into the first. For (i), after the relaxation every constraint is an inequality with a slack and a multiplier. The same elimination removes all of them and leaves \(K_\tau\). Take \(v \ne 0\). If \(v \in \ker \tilde J_A\), then \(v^\top K_\tau v \ge v^\top (W + \delta_w I) v > 0\), because the inactive rows contribute a nonnegative term. If \(v \notin \ker \tilde J_A\), let \(\sigma_A\) be the smallest active entry of \(\tilde\Sigma\), so that \(v^\top \tilde J_h^\top \tilde\Sigma \tilde J_h v \ge \sigma_A \lVert \tilde J_A v \rVert^2\). Finsler's lemma says that a symmetric matrix positive definite on \(\ker \tilde J_A\) becomes positive definite after adding \(\sigma_0 \tilde J_A^\top \tilde J_A\) for some finite \(\sigma_0\). Hence \(K_\tau \succ 0\) as soon as \(\sigma_A \ge \sigma_0\). If \(\delta_w > \max\{0, -\lambda_{\min}(W)\}\), then \(W + \delta_w I \succ 0\) and \(K_\tau \succ 0\) for every \(\tilde\Sigma \succ 0\). For (ii), the operation on the rows is a nonsingular row operation, so the systems are equivalent. We have \(v^\top K_\gamma v = v^\top K_c v + \gamma \lVert J_g v \rVert^2\). This is positive for \(v \in \ker J_g \setminus \{0\}\) by hypothesis. For \(v \notin \ker J_g\) it is positive once \(\gamma\) is large, since \(v^\top K_c v\) is bounded below on the unit sphere and \(\lVert J_g v \rVert^2\) is bounded away from zero on the unit sphere outside a neighbourhood of \(\ker J_g\). The formulas for \(\Delta\lambda\) and \(\Delta x\) are the two block equations solved in order. ∎
(The two routes as implemented, and the picture) The two routes are the ones implemented on GPUs. The equality relaxation is MadNLP's LiftedKKT. The augmented-Lagrangian route is HyKKT, which Regev and coauthors introduced as a hybrid direct–iterative method. There \(K_\gamma\) is factorized by Cholesky, and the Schur-complement system in \(\Delta\lambda\) is solved by conjugate gradients, the iterative method for a positive definite system that needs one product with the matrix per iteration. That system is small when there are few equalities, and forming its matrix would be expensive, so each conjugate-gradient step costs one pair of triangular solves with the Cholesky factor instead.S. Regev, N.-Y. Chiang, E. Darve, C. G. Petra, M. A. Saunders, K. Świrydowicz and S. Peleš, "HyKKT: a hybrid direct-iterative method for solving KKT linear systems", Optimization Methods and Software 38 (2023), 332–355; F. Pacaud, S. Shin, A. Montoison, M. Schanen and M. Anitescu, "Condensed interior-point methods for scalable nonlinear programming on GPUs", Mathematical Programming Computation (2026), online 10 August 2026, doi 10.1007/s12532-026-00335-0; arXiv 2405.14236 (v3, July 2026), which names the two variants LiftedKKT and HyKKT and analyses both. The picture is this. The indefinite system (7.5.2) is a saddle point: minimize over \(\Delta x\), maximize over the multipliers. Condensation prices the inequalities into the Hessian through \(\Sigma\). That turns the saddle into a bowl in \(\Delta x\), as long as no equalities remain. The two routes remove the equalities in two ways. One turns them into inequalities. The other adds a penalty large enough to make the bowl steep in the directions the equalities constrain.
The script of the next subsection studies the first route in full on a toy problem: minimize \(\tfrac12\lVert x - a \rVert^2\) with \(a = (2, 1, 1)\) subject to \(x_1 + x_2 \le 1\), \(0 \le x \le 3\) and the equality \(x_1 - 2x_2 = 0\), whose optimum is \(x^\star = (2/3, 1/3, 1)\). In the notation of Definition 7.5.1 this has \(n = 3\), one equality \(g(x) = x_1 - 2x_2\), so \(m_e = 1\), and \(m = 7\) inequalities, the row \(x_1 + x_2 - 1 \le 0\) and the six bounds, each with a slack; at the optimum the row and the equality are active and \(x_3 = 1\) touches nothing.
From the Newton system (7.5.1) to a positive definite matrix
(7.5.1) unknowns (dx, ds, dnu, dlam)
[ W 0 J_h' J_g' ]
[ 0 Sig I 0 ]
[ J_h I 0 0 ]
[ J_g 0 0 0 ]
|
| row 2 solved for ds = -Sig^-1 (r_s + dnu)
| and substituted into row 3
v
(7.5.2) augmented, unknowns (dx, dnu, dlam): symmetric
indefinite, the system a CPU solver factorizes (pivoted LDL')
[ W J_h' J_g' ]
[ J_h -Sig^-1 0 ]
[ J_g 0 0 ]
|
| row 2 solved for dnu = Sig (J_h dx + r~_p)
| and substituted into row 1
v
(7.5.3) condensed, unknowns (dx, dlam): still indefinite
when m_e > 0
[ K_c J_g' ] K_c = W + J_h' Sig J_h
[ J_g 0 ]
|
+---------------------------------+
| |
(i) LiftedKKT: (ii) HyKKT: gamma J_g' times
-tau <= g_j(x) <= tau row 2 added to row 1; row 2,
replaces g_j(x) = 0, so m_e = 0 J_g dx = b_2, stays
| |
v v
K_tau = (W + delta_w I) K_gam = K_c + gamma J_g' J_g
+ J~_h' Sig~ J~_h
PD for large active Sig~, and PD for all large gamma;
for every delta_w > Cholesky of K_gam, conjugate
max{0, -lambda_min(W)}; gradients on the Schur
Cholesky, pivot order fixed complement J_g K_gam^-1 J_g'
Sig = diag(nu_i / s_i); r~_p = r_p - Sig^-1 r_s
PD: positive definite, under the hypotheses of Proposition 7.5.2
Why the augmented system needs pivoting and the condensed one does not
A symmetric indefinite factorization \(P^\top K P = L D L^\top\) chooses its pivots by looking at the numbers, which is what makes it sequential. The reason is that an indefinite matrix can present a zero or tiny diagonal entry at any stage of the elimination, so the next pivot, a \(1 \times 1\) or a \(2 \times 2\) block, has to be chosen from the values as they stand after the previous eliminations, and the choice at each stage depends on every choice before it. A positive definite matrix has a positive pivot at every stage in every symmetric ordering, so no choice is needed and the order can be fixed beforehand from the sparsity pattern alone, which is what lets the factorization be scheduled in advance. The condensed route replaces the indefinite factorization by a Cholesky factorization with a pivot order fixed that way. The next proposition says that nothing is lost in the exchange. The quantity the pivoted factorization is used to test is readable from the condensed matrix.
Proposition 7.5.3 (inertia and the condensed matrix). Let \(K = \begin{pmatrix} H & J^\top \\ J & -D \end{pmatrix}\) with \(H\) symmetric \(n \times n\), \(J\) of size \(m \times n\) and \(D \succ 0\) diagonal. Then the inertia of \(K\), the numbers of positive, negative and zero eigenvalues, is \((n, m, 0)\) if and only if the condensed matrix \(H + J^\top D^{-1} J\) is positive definite.
Proof. Block elimination with the invertible block \(-D\) gives the congruence \(\begin{pmatrix} I & J^\top D^{-1} \\ 0 & I \end{pmatrix} K \begin{pmatrix} I & 0 \\ D^{-1} J & I \end{pmatrix} = \begin{pmatrix} H + J^\top D^{-1} J & 0 \\ 0 & -D \end{pmatrix},\) in which the left factor is the transpose of the right one. Congruence preserves inertia (Sylvester's law). The inertia of a block-diagonal matrix is the sum of the inertias of its blocks: \((n, 0, 0)\) for the first block if and only if it is positive definite, and \((0, m, 0)\) for \(-D\). ∎
(The inertia test without pivoting) With \(D = \Sigma^{-1}\) this is exactly the first two blocks of (7.5.2), and the condensed matrix of the proposition is \(K_c\). The interior-point methods of Section 1.3 insist on inertia \((n, m + m_e, 0)\), because that is the condition under which the Newton step is a descent direction for the barrier merit function. When the pivoted factorization reports the wrong inertia they increase \(\delta_w\) and refactorize. That is the inertia correction of the filter line-search method.A. Wächter and L. T. Biegler, "On the implementation of an interior-point filter line-search algorithm for large-scale nonlinear programming", Mathematical Programming 106 (2006), 25–57 (the inertia correction). The inertia argument for block matrices is the additivity formula of E. V. Haynsworth, "Determination of the inertia of a partitioned Hermitian matrix", Linear Algebra and its Applications 1 (1968), 73–81. The proposition says that the same information is available without pivoting. A Cholesky factorization of \(K_c\) either succeeds, in which case the inertia is right, or meets a nonpositive pivot, in which case \(\delta_w\) must grow. That is the test the GPU codes run.
Proposition 7.5.3: the congruence, and two ways to test inertia
[ H J' ] [ H + J' D^-1 J 0 ]
K = [ J -D ] --block elimination--> [ 0 -D ]
(a congruence: Sylvester's law keeps inertia)
inertia(K) = inertia(H + J' D^-1 J) + inertia(-D)
= (n, 0, 0) + (0, m, 0) exactly when
H + J' D^-1 J is PD
with D = Sig^-1 the first block is K_c, and the test runs
on a CPU (Section 1.3) on a GPU
pivoted LDL' of (7.5.2) Cholesky of K_c, pivot
order fixed in advance
| |
v v
inertia (n, m + m_e, 0)? a nonpositive pivot?
| | | |
| yes | no | no | yes
v v v v
solve for increase solve for increase
the step delta_w, the step delta_w,
refactorize refactorize
What condensation costs: conditioning
(How \(\Sigma\) spreads along the central path) On the central path the complementarity products are \(s_i \nu_i = \mu\) for every \(i\). Consider a solution that satisfies strict complementarity (Section 1.3), so that at every constraint exactly one of the slack and the multiplier is zero. Near it, the active inequalities have \(s_i \to 0\) with \(\nu_i\) bounded away from zero, and the inactive ones have \(\nu_i \to 0\) with \(s_i\) bounded away from zero. Hence \(\Sigma_{ii} = \nu_i / s_i\) is of order \(1/\mu\) on the active set and of order \(\mu\) on the inactive set. The condensed matrix \(K_c = W + J_h^\top \Sigma J_h\) therefore contains terms of order \(1/\mu\) along the range of the active Jacobian and terms of order one elsewhere. Pacaud and coauthors make this precise. They also show that the resulting ill-conditioning is harmless in a specific sense.
Theorem 7.5.4 (structured ill-conditioning of the condensed matrix; Pacaud, Shin, Montoison, Schanen and Anitescu, 2026, Theorem 2). Let the current primal-dual iterate satisfy the centrality conditions of the interior-point method near a solution, the conditions that keep the iterate within a fixed neighbourhood of the central path of Section 1.3. Suppose that at the solution the active constraints have linearly independent gradients, strict complementarity holds and the second-order sufficient condition holds. Let \(\Xi = s^\top \nu / m\) be the duality measure. Let \(A\) be the Jacobian of the active constraints (the \(m_a\) active inequalities and the \(m_e\) equalities), \(\ell = m_a + m_e\), and \(s_{\min}\) the smallest active slack. Put \(\underline\sigma = \min\{1/\Xi, \gamma\}\) and \(\bar\sigma = \max\{1/s_{\min}, \gamma\}\). Then, for sufficiently small \(\Xi\), the HyKKT matrix \(K_\gamma\) has (i) \(\ell\) largest eigenvalues that are positive, the largest of order \(\Theta(\bar\sigma)\) and the \(\ell\)-th of order \(\Omega(\underline\sigma)\); (ii) \(n - \ell\) smallest eigenvalues of order \(\Theta(1)\); (iii) condition number \(\kappa_2(K_\gamma) = \Theta(\bar\sigma)\) when \(0 < \ell < n\); and (iv) invariant subspaces within \(O(\underline\sigma^{-1})\) of the range of \(A^\top\) and of its orthogonal complement. The same statements hold for the LiftedKKT matrix \(K_\tau\) with \(m_e = 0\).
(The proof's picture, and why the ill-conditioning is harmless) The proof separates the active constraints from the inactive ones. Index the inactive inequalities by \(N\) and the active ones by \(B\), so that \(H_N\), \(S_N\) and \(V_N\) are the Jacobian, slack and multiplier blocks of the inactive rows and \(S_B\), \(V_B\) those of the active rows. Then \(K_\gamma = W + H_N^\top S_N^{-1} V_N H_N + A^\top D_\gamma A\), where \(D_\gamma\) is the diagonal matrix that carries \(S_B^{-1} V_B\) on the active inequalities and \(\gamma\) on the equalities. The inactive term is \(O(\Xi)\). The eigenvalues are then read off the dominant term \(A^\top D_\gamma A\) and the \(O(1)\) remainder. The details are in the paper.Pacaud, Shin, Montoison, Schanen and Anitescu (2026), Section 4.2 ("Structured ill-conditioning"), Theorem 2 and the remarks after it; the characterization \(s_i = \Theta(\Xi)\), \(\nu_i = \Theta(1)\) for active and \(s_i = \Theta(1)\), \(\nu_i = \Theta(\Xi)\) for inactive constraints under the centrality conditions is their Proposition 1, quoted from S. J. Wright. The paper's error analysis (its Section 4.2.2) concludes that the error in the Newton direction is governed by the \(O(1)\) part of the spectrum rather than by \(\kappa_2(K_\gamma)\), and reports that as \(\gamma\) rises from \(10^4\) to \(10^8\) the conjugate-gradient iterations fall tenfold while the relative residual grows linearly in \(\gamma\). In a picture, the condensed matrix is a bowl that is extremely steep in the \(\ell\) directions the active constraints fix, and gently curved in the other \(n - \ell\). The steep directions are the ones on which the Newton step is nearly determined by the constraints anyway. The condition number is as bad as the steepest direction. But the error that matters, the error in the gently curved directions, is not multiplied by it. That is why Cholesky in double precision survives condition numbers of \(10^{14}\) in these codes. It is also why it survives them only with iterative refinement and only to a limited accuracy. Iterative refinement takes a computed solution, forms the residual of the linear system in working precision, solves for the correction with the same factor and adds it, a fixed number of times. Each pass removes part of the error the factorization's rounding introduced, as long as the factor is accurate enough for the correction to point the right way.
The measurements go with the theorem. Shin, Anitescu and Pacaud report, for the IEEE 118-bus optimal power flow case at tolerance \(10^{-4}\), a condition number of \(9.43 \times 10^{11}\) for the augmented system and \(3.15 \times 10^{14}\) for the condensed one. They write that "reliable convergence is achieved only up to a tolerance of \(10^{-4}\)", and they set their solver's tolerance to \(\epsilon_{\mathrm{mach}}^{1/4} \approx 10^{-4}\) by default when the condensed system is used.S. Shin, M. Anitescu and F. Pacaud, "Accelerating optimal power flow with GPUs: SIMD abstraction of nonlinear programs and condensed-space interior-point methods", Electric Power Systems Research 236 (2024), article 110651 (PSCC 2024); arXiv 2307.16830 (2023). The quotations and the two condition numbers are from its "Known limitations" paragraph and its section on the barrier subproblem. The later paper reaches tolerances from \(10^{-4}\) to \(10^{-8}\) with LiftedKKT and cuDSS, NVIDIA's library of sparse direct factorizations for the device, on a large power-flow instance. It takes 20 to 30 seconds on the GPU against 157 to 302 seconds for Pardiso, a CPU sparse direct solver, on the CPU. It also records the price in its Table 2. The residual norm of the final KKT solve is \(1.2 \times 10^{-2}\) at tolerance \(10^{-4}\) and \(1.2 \times 10^{-6}\) at \(10^{-8}\), about two orders of magnitude worse than the tolerance. The condition number of \(K_\tau\) exceeds \(10^{18}\), and the factorization succeeds only because iterative refinement repairs the loss.Pacaud, Shin, Montoison, Schanen and Anitescu (2026), Table 2 (LiftedKKT: 110 to 200 interior-point iterations and 156.9 to 301.7 s with Pardiso on the CPU, 104 to 114 iterations and 19.9 to 30.4 s with cuDSS on the GPU, accuracy \(1.2 \times 10^{-2}\) to \(1.2 \times 10^{-6}\) as the tolerance goes from \(10^{-4}\) to \(10^{-8}\)) and the paragraph before it, which states that the condition number of \(K_\tau\) can exceed \(10^{18}\).
The script below shows the mechanism on a problem small enough to read. It has three variables, one equality \(x_1 - 2x_2 = 0\), the inequality \(x_1 + x_2 \le 1\) and the bounds \(0 \le x \le 3\), with the objective \(\tfrac12\lVert x - a\rVert^2\) and \(a = (2, 1, 1)\). At the optimum \(x^\star = (2/3, 1/3, 1)\) the inequality and the equality are active, so \(\ell = 2 < n = 3\), and the third coordinate is the null-space direction of Theorem 7.5.4. For each \(\mu\) from \(10^{-1}\) to \(10^{-8}\) the script solves the barrier subproblem by Newton's method on the full system (7.5.1), which puts it on the central path. It then forms the augmented matrix of (7.5.2), the condensed matrix \(K_c\) and the HyKKT matrix \(K_\gamma\) for two values of \(\gamma\), and compares the Newton direction from the pivot-free route against the one from the augmented system.
# condensed_kkt.py: the Newton system of a barrier method three ways,
# along the central path.
#
# Problem: min 1/2 ||x - a||^2 s.t. x1 + x2 <= 1, 0 <= x <= 3,
# x1 - 2 x2 = 0, a = (2, 1, 1).
# Optimum x* = (2/3, 1/3, 1), nu* = 10/9 on the one active inequality,
# lambda* = 2/9 on the equality; x3 touches no active constraint, so the
# active Jacobian has a one-dimensional null space.
import numpy as np
np.set_printoptions(precision=3, suppress=True)
a = np.array([2.0, 1.0, 1.0])
n = 3
W = np.eye(n) # Hessian of f
# g(x) = G x - r <= 0
G = np.vstack([[1, 1, 0], -np.eye(n), np.eye(n)])
r = np.array([1, 0, 0, 0, 3, 3, 3], float)
h = np.array([1.0, -2.0, 0.0]) # h' x = 0
m = len(r)
def newton_full(x, s, nu, lam, mu):
"""One primal-dual Newton step on the full (n + 2m + 1) system.
Returns the steps (dx, ds, dnu, dlam) and the largest residual.
"""
rd = x - a + G.T @ nu + h * lam
rp = G @ x - r + s
re = h @ x
rc = s * nu - mu
N = n + m + 1 + m
K = np.zeros((N, N))
rhs = -np.concatenate([rd, rp, [re], rc])
K[:n, :n] = W
K[:n, n + m + 1:] = G.T
K[:n, n + m] = h
K[n:n + m, :n] = G
K[n:n + m, n:n + m] = np.eye(m)
K[n + m, :n] = h
K[n + m + 1:, n:n + m] = np.diag(nu)
K[n + m + 1:, n + m + 1:] = np.diag(s)
d = np.linalg.solve(K, rhs)
return (d[:n], d[n:n + m], d[n + m + 1:], d[n + m],
max(abs(rd).max(), abs(rp).max(), abs(re), abs(rc).max()))
def inertia(M):
"""(number of positive, number of negative) eigenvalues of M."""
ev = np.linalg.eigvalsh(M)
return int((ev > 0).sum()), int((ev < 0).sum())
x = np.array([0.4, 0.3, 0.5])
s = np.maximum(r - G @ x, 0.5)
nu = 1.0 / s
lam = 0.0
print(" mu Sigma_min Sigma_max 1/s_min kappa(aug) kappa(Kc)")
print(" kappa(K_gam) gam=1e4 gam=1e8"
" |dx_hykkt - dx_aug| / |dx|")
for mu in [1e-1, 1e-2, 1e-3, 1e-4, 1e-5, 1e-6, 1e-7, 1e-8]:
# solve the barrier subproblem for this mu: a point on the central path
for it in range(60):
dx, ds, dnu, dlam, res = newton_full(x, s, nu, lam, mu)
if res < 1e-13:
break
# the fraction-to-the-boundary rule keeps s and nu positive
al = min([1.0]
+ [-0.995 * s[i] / ds[i]
for i in range(m) if ds[i] < 0]
+ [-0.995 * nu[i] / dnu[i]
for i in range(m) if dnu[i] < 0])
x, s, nu, lam = (x + al * dx, s + al * ds,
nu + al * dnu, lam + al * dlam)
Sig = nu / s # the diagonal that condensation introduces
# a nonzero right-hand side
rd = x - a + G.T @ nu + h * lam + 1e-3
rp = G @ x - r + s
re = h @ x
rc = s * nu - 0.5 * mu
# the pivoted route: the augmented system (7.5.2)
K_aug = np.block([[W, G.T, h[:, None]],
[G, -np.diag(1 / Sig), np.zeros((m, 1))],
[h[None, :], np.zeros((1, m)), np.zeros((1, 1))]])
b_aug = np.concatenate([-rd, -rp + rc / nu, [-re]])
dx_aug = np.linalg.solve(K_aug, b_aug)[:n]
# inequalities eliminated
Kc = W + G.T @ (Sig[:, None] * G)
b1 = -rd - G.T @ (Sig * (rp - rc / nu))
b2 = -re
out = []
# HyKKT: K_gam = Kc + gam h h' is positive definite
for gam in (1e4, 1e8):
Kg = Kc + gam * np.outer(h, h)
L = np.linalg.cholesky(Kg) # Cholesky without pivoting
solve = lambda v: np.linalg.solve(L.T, np.linalg.solve(L, v))
t = solve(b1 + gam * h * b2)
u = solve(h)
# Schur complement (1 x 1 here; CG in general)
dlam = (h @ t - b2) / (h @ u)
# the HyKKT direction is t - dlam u
rel = np.linalg.norm(t - dlam * u - dx_aug) / np.linalg.norm(dx_aug)
out.append((np.linalg.cond(Kg), rel))
print(f"{mu:6.0e}{Sig.min():11.2e}{Sig.max():11.2e}"
f"{1/s.min():10.1e}{np.linalg.cond(K_aug):12.2e}"
f"{np.linalg.cond(Kc):11.2e}")
print(f"{'':8}{out[0][0]:20.2e}{out[1][0]:10.2e} "
f"{out[0][1]:.1e} / {out[1][1]:.1e}")
print(f"\nat mu = 1e-8: x = {x},\n"
f" nu = {nu},\n"
f" lambda = {lam:.4f}\n"
f" (x* = (2/3, 1/3, 1), nu* = 10/9, lambda* = 2/9)")
Kc_eq = np.block([[Kc, h[:, None]], [h[None, :], np.zeros((1, 1))]])
print(f"inertia: augmented matrix {inertia(K_aug)} = (n, m + 1);\n"
f" condensed Kc {inertia(Kc)};\n"
f" Kc with the equality row kept {inertia(Kc_eq)}")
print(f"eigenvalues of Kc at mu = 1e-8: {np.linalg.eigvalsh(Kc)}\n"
f" (one of order 1/mu for the active row, two of order 1)")
mu Sigma_min Sigma_max 1/s_min kappa(aug) kappa(Kc)
kappa(K_gam) gam=1e4 gam=1e8 |dx_hykkt - dx_aug| / |dx|
1e-01 1.38e-02 1.77e+01 1.3e+01 1.05e+02 3.31e+01
4.47e+04 4.47e+08 6.3e-14 / 2.3e-10
1e-02 1.40e-03 1.28e+02 1.1e+02 8.01e+02 2.54e+02
4.94e+04 4.94e+08 3.4e-15 / 2.3e-11
1e-03 1.41e-04 1.24e+03 1.1e+03 7.95e+03 2.48e+03
5.02e+04 4.99e+08 1.7e-16 / 1.0e-12
1e-04 1.41e-05 1.24e+04 1.1e+04 7.94e+04 2.47e+04
5.42e+04 5.00e+08 9.6e-17 / 2.2e-15
1e-05 1.41e-06 1.23e+05 1.1e+05 7.94e+05 2.47e+05
2.53e+05 5.00e+08 2.2e-16 / 2.9e-16
1e-06 1.41e-07 1.23e+06 1.1e+06 7.94e+06 2.47e+06
2.47e+06 5.00e+08 4.4e-16 / 4.4e-16
1e-07 1.41e-08 1.23e+07 1.1e+07 7.94e+07 2.47e+07
2.47e+07 5.03e+08 2.8e-16 / 2.8e-16
1e-08 1.41e-09 1.23e+08 1.1e+08 7.94e+08 2.47e+08
2.47e+08 5.42e+08 2.3e-16 / 2.3e-16
at mu = 1e-8: x = [0.667 0.333 1. ],
nu = [1.111 0. 0. 0. 0. 0. 0. ],
lambda = 0.2222
(x* = (2/3, 1/3, 1), nu* = 10/9, lambda* = 2/9)
inertia: augmented matrix (3, 8) = (n, m + 1);
condensed Kc (3, 0);
Kc with the equality row kept (3, 1)
eigenvalues of Kc at mu = 1e-8: [1.000e+00 1.000e+00 2.469e+08]
(one of order 1/mu for the active row, two of order 1)
The script performs eight barrier solves, each a few Newton steps on an \(18 \times 18\) system, and a few eigenvalue decompositions. Nothing in it is parallel. It is a diagnostic, not a solver.
(Five things visible in the table) Five things are visible in the table. First, the diagonal \(\Sigma\) spreads as the theory says. Its smallest entry is \(\Sigma_{\min} \approx 9\mu/64\), from the inactive bound \(x_2 \le 3\), whose slack at the optimum is \(8/3\). Its largest entry is \(\Sigma_{\max} \approx (10/9)^2/\mu\), from the active row, whose multiplier is \(10/9\). So the diagonal spans about \(2\log_{10}(1/\mu) + 1\) orders of magnitude: nine at \(\mu = 10^{-4}\) and seventeen at \(\mu = 10^{-8}\). Second, the augmented matrix has inertia \((3, 8)\) at every point, three positive and eight negative eigenvalues, and the condensed matrix is positive definite at every point, as Proposition 7.5.3 requires. Keeping the equality row leaves one negative eigenvalue, which is why HyKKT adds the penalty. Third, \(K_c\), which lacks the penalty \(\gamma J_g^\top J_g\), has one large eigenvalue, for the active inequality, and two of order one, one of them along the equality direction. In \(K_\gamma\) that direction is lifted to order \(\gamma\), which is Theorem 7.5.4 with \(\ell = 2\) and \(n - \ell = 1\). The condition number of \(K_c\) grows like \(2.5/\mu\). Fourth, the HyKKT condition number is \(\Theta(\max\{1/s_{\min}, \gamma\})\) to within the constant \(\lVert \nabla h \rVert^2 = 5\). With \(\gamma = 10^4\) it sits near \(5 \times 10^4\) until \(1/s_{\min}\) overtakes \(\gamma\) at \(\mu \approx 10^{-4}\), and then it grows with \(1/\mu\). With \(\gamma = 10^8\) it is \(5 \times 10^8\) throughout, as part (iii) of the theorem predicts. Fifth, the Newton direction from the pivot-free route agrees with the pivoted one to \(2.3 \times 10^{-10}\) in the worst case and to machine precision in most rows, in spite of condition numbers of \(10^8\). This is the structured ill-conditioning at work: the direction is determined by the well-conditioned part of the spectrum. On this toy problem the condensed matrix is no worse conditioned than the augmented one. On the 118-bus case the measurement went the other way by a factor of about 300. The theorem is agnostic between the two, since it bounds the condensed matrix alone.
The GPU interior-point step
The derivative evaluations, the assembly of \(K_c\) and the factorization all map to the device once the pivot order is fixed. Two terms from sparse direct solvers are needed to read the box below. A sparse Cholesky factorization is organized by its elimination tree, the dependency order of the factor's columns. A column can be computed as soon as its descendants in the tree are done. The height of the tree therefore bounds the sequential depth of the factorization, and its width bounds the parallelism. A supernodal factorization groups columns with the same sparsity structure into dense blocks and processes each block with dense kernels, which is what a device does well. The algorithm box records one step of the LiftedKKT variant as MadNLP runs it, with ExaModels for derivatives and cuDSS for the factorization.
The elimination tree of a sparse Cholesky factorization (schematic)
-+- o computed last
| / \
height o o
| / \ / \
-+- o o o o computed first
|<- width ->|
o a column of the factor, which can be computed as soon as
its descendants in the tree are done
height: bounds the sequential depth of the factorization
width: bounds the parallelism
Algorithm 7.5.5 One condensed-space interior-point step on a GPU
(LiftedKKT; after Shin, Anitescu and Pacaud 2024 and
Pacaud et al. 2026)
Input current (x, s, nu) with s, nu > 0; barrier parameter mu;
tolerance eps_tol; relaxation tau (about eps_tol);
regularization delta_w; the model as a set of repeated
expression patterns (ExaModels)
Output the next iterate (x+, s+, nu+)
1. (once) replace every equality g_j(x) = 0 by -tau <= g_j(x) <= tau,
so that all constraints are inequalities with slacks
2. evaluate f, h, grad f, the Jacobian J_h and the Lagrangian
Hessian W on the device:
one thread per instance of each expression pattern,
reverse-mode differentiation inside the pattern,
results scattered into fixed sparse structures
3. Sigma := diag(nu_i / s_i);
assemble K_tau := W + delta_w I + J_h' Sigma J_h in a fixed
sparsity pattern (a map over nonzeros)
4. factorize K_tau = L L' (or L D L') by cuDSS, with the symbolic
analysis computed once at the first iteration and reused;
if a pivot is not positive: [Proposition 7.5.3]
delta_w := 10 delta_w (or the first positive value) and
repeat step 4
5. solve for dx by two triangular solves; iterative refinement:
r := b - K_tau dx, dx := dx + K_tau^{-1} r,
a fixed number of times
6. recover ds := -(h(x) + s) - J_h dx and
dnu := Sigma (J_h dx + h(x) + s) - nu + mu / s
(the eliminated rows)
7. step lengths by the fraction-to-the-boundary rule for s and nu,
which caps the step so that every s_i and nu_i stays at least a
fraction (0.995) of its current value away from zero
(two reductions);
filter line search on x (a few evaluations of step 2 without
derivatives);
update mu by the usual rule
Invariant
K_tau is positive definite whenever step 4 succeeds, so dx is a
descent direction for the barrier merit function; the relaxation
tau changes the problem by O(tau), which is below the accuracy
eps_tol to which the inequalities are satisfied anyway.
Cost per step
the factorization (one supernodal Cholesky with a fixed elimination
tree) dominates; derivative evaluation and assembly are maps over
the model's patterns and the matrix's nonzeros.
Parallel
steps 2, 3, 5 (the triangular solves as far as the elimination tree
allows) and 6 in full; step 4 through cuDSS's supernodal kernels;
the line search and the barrier update are scalars. Nothing
branches on data except the pivot test of step 4, which is a single
flag.
One step of Algorithm 7.5.5 on the device, stage by stage
once: 1. every g_j(x) = 0 becomes -tau <= g_j(x) <= tau
4. the symbolic analysis of K_tau, at the first iteration
(x, s, nu), mu
|
v
2. f, h, grad f, J_h, W map: one thread per instance
| of each expression pattern
v
3. Sigma; K_tau := W + delta_w I map over the nonzeros
+ J_h' Sigma J_h
|
v
4. K_tau = L L' by cuDSS <-------+ supernodal kernels
| |
every pivot positive? --no--> delta_w := 10 delta_w (or the
| yes first positive value)
v
5. two triangular solves; then, as far as the elimination
a fixed number of times, tree allows
r := b - K_tau dx,
dx := dx + K_tau^{-1} r
|
v
6. ds, dnu (the eliminated rows) map
|
v
7. fraction to the boundary (0.995) two reductions
filter line search on x a few evaluations of step 2
update mu scalar
|
v
(x+, s+, nu+)
the one branch on data: the pivot test of step 4, a single flag
The numbers reported for this design come with conditions, and the conditions follow each number. Shin, Anitescu and Pacaud evaluate derivatives on the GPU about 500 times faster than AMPL or JuMP on the CPU. They solve their largest instance, case30000_goc, about 4 times faster on a Quadro GV100 than their own solver on a Xeon Gold 6140, and about 10 times faster than Ipopt with MA27, all at tolerance \(10^{-4}\).Shin, Anitescu and Pacaud (2024), abstract and the section on benchmark results: "our method achieves a 4x speedup compared with our solver running on CPUs for the largest instance", and "approximately 10 times faster than state-of-the-art tools (Ipopt, JuMP.jl, and Ma27)". Pacaud and coauthors report that cuDSS factorizes the condensed matrices 4 to 8 times faster than Pardiso in total interior-point time. They report that HyKKT with cuDSS reaches an 8 times speedup over HSL MA27 on the largest PGLib instances and "a remarkable tenfold acceleration" on large optimal power flow, and that a laptop A1000 GPU already beats MA27. On the CUTEst collection the method shows "diminished robustness and limited speedups for edge cases", because a single dense constraint row makes the condensed matrix dense.Pacaud, Shin, Montoison, Schanen and Anitescu (2026), abstract, contribution 3 and Section 6; Table 1 of that paper gives the raw factorization times (cuDSS Cholesky 0.0405 s against CHOLMOD 0.744 s, Pardiso 0.294 s and MA86 0.138 s with eight threads, on one condensed matrix). All comparisons are at the authors' tolerances on the authors' instances. Two follow-ups show where the design is heading. A multi-period optimal power flow with more than ten million variables was solved to \(10^{-4}\) in under ten minutes on a GH200 with 480 GB of unified memory. MadNCL, an augmented-Lagrangian method on the GPU for degenerate problems, solved a 500-bus security-constrained problem with 256 contingencies "fully on the GPU in less than 3 minutes, whereas Knitro takes more than 3 hours".S. Shin, V. Rao, M. Schanen, D. A. Maldonado and M. Anitescu, "Scalable multi-period AC optimal power flow utilizing GPUs with high memory capacities", arXiv 2405.14032 (2024); F. Pacaud, A. Nurkanović, A. Pozharskiy, A. Montoison and S. Shin, "An augmented Lagrangian method on GPU for security-constrained AC optimal power flow", arXiv 2510.13333 (2025); A. Montoison, F. Pacaud, M. Saunders, S. Shin and D. Orban, "MadNCL: a GPU implementation of Algorithm NCL for large-scale, degenerate nonlinear programs", arXiv 2510.05885 (2025). Preprints, authors' numbers.
The cost is as the box says. The factorization is the sequential core even on the device. It is parallel only as far as the elimination tree of the fixed pivot order allows, and everything around it is a map or a reduction. The inner loop a CPU solver cannot parallelize, the data-dependent pivot sequence, has been removed rather than parallelized.
A frontier of small nonlinear programs
A MINLP tree does not solve one large NLP. It solves many small ones: a local NLP per node for an incumbent, one per candidate assignment in the heuristics of Section 3.2, one per start in a multistart. For those the right unit of parallelism is the node, not the matrix. The condensed system has a second virtue here. Its assembly and factorization have the same instruction stream for every node, so a batch of nodes is one kernel. The C++23 listing below is that kernel on a CPU. A frontier of \(K\) nodes shares the Hessian \(W\) and the Jacobian \(J\) of its \(m\) inequalities and differs in the barrier diagonal \(\Sigma_k\) and the right-hand side. Each worker thread takes a strided share of the batch, assembles \(K_k = W + \delta I + J^\top \Sigma_k J\), factorizes it by Cholesky with no pivoting, and solves.
batched_condensed.cpp: the batch in memory and its threads
shared by all nodes: W (n x n), J (m x n), delta
one slice per node k, in three flat arrays:
sig [ Sigma_0 | Sigma_1 | . . . | Sigma_k | . . . ] m each
rhs [ b_0 | b_1 | . . . | b_k | . . . ] n each
sol [ x_0 | x_1 | . . . | x_k | . . . ] n each
thread p takes the nodes k = p, p + P, p + 2P, . . .
node k 0 1 . . . P-1 P P+1 . . . 2P . . .
thread p 0 1 . . . P-1 0 1 . . . 0 . . .
for each of its nodes, in its own Kmat (n x n):
assemble Kmat := W + delta I + J' diag(Sigma_k) J
factorize Kmat = L L', no pivoting --a pivot <= 0--> a failure
solve L y = b_k, then L' x_k = y, into sol
n = 8, m = 12, K = 4096 nodes, P = 8 threads; on a device, one
thread block per node, with the matrix in shared memory
// batched_condensed.cpp: condensed KKT solves for a frontier of K small
// NLP nodes, C++23.
//
// Every node shares the Hessian W and the Jacobian J of its m
// inequalities and differs in the barrier diagonal
// Sigma_k = diag(nu_k / s_k) and its right-hand side. The kernel
// assembles
// K_k = W + delta I + J' Sigma_k J (n x n, positive definite),
// factorizes it by Cholesky with no pivoting (the same symbolic
// structure for every k) and solves. One std::jthread per strided share
// of the batch stands in for one thread block per node on a device.
// Compile check:
// clang++ -std=c++23 -fsyntax-only -Wall -Wextra batched_condensed.cpp
#include <atomic>
#include <chrono>
#include <cmath>
#include <cstdio>
#include <random>
#include <span>
#include <thread>
#include <vector>
constexpr int n = 8, m = 12;
struct Shared {
double W[n][n];
double J[m][n];
double delta;
};
// assemble K = W + delta I + J' diag(sig) J for one node
// (n*n*m multiply-adds, no branches)
static void assemble(const Shared& S, std::span<const double, m> sig,
double K[n][n]) {
for (int i = 0; i < n; ++i)
for (int j = 0; j < n; ++j) {
double v = S.W[i][j] + (i == j ? S.delta : 0.0);
for (int r = 0; r < m; ++r)
v += S.J[r][i] * sig[r] * S.J[r][j];
K[i][j] = v;
}
}
// in-place Cholesky K = L L' (lower triangle) and the two triangular
// solves; returns false if not PD
static bool cholesky_solve(double K[n][n], std::span<double, n> x) {
for (int j = 0; j < n; ++j) {
double d = K[j][j];
for (int k = 0; k < j; ++k)
d -= K[j][k] * K[j][k];
if (d <= 0.0)
return false; // a pivot that is not positive: K is not PD
K[j][j] = std::sqrt(d);
for (int i = j + 1; i < n; ++i) {
double v = K[i][j];
for (int k = 0; k < j; ++k)
v -= K[i][k] * K[j][k];
K[i][j] = v / K[j][j];
}
}
// forward solve with L, then backward solve with L'
for (int i = 0; i < n; ++i) {
double v = x[i];
for (int k = 0; k < i; ++k)
v -= K[i][k] * x[k];
x[i] = v / K[i][i];
}
for (int i = n - 1; i >= 0; --i) {
double v = x[i];
for (int k = i + 1; k < n; ++k)
v -= K[k][i] * x[k];
x[i] = v / K[i][i];
}
return true;
}
int main() {
const int K = 4096, P = 8; // nodes in the frontier, worker threads
Shared S{};
S.delta = 1e-8;
std::mt19937_64 gen(7);
std::uniform_real_distribution<double> U(-1.0, 1.0);
for (int i = 0; i < n; ++i) {
S.W[i][i] = 2.0;
for (int j = 0; j < i; ++j)
S.W[i][j] = S.W[j][i] = 0.1 * U(gen);
}
for (int r = 0; r < m; ++r)
for (int j = 0; j < n; ++j)
S.J[r][j] = U(gen);
std::vector<double> sig(K * m), rhs(K * n), sol(K * n);
// Sigma entries from 1e-4 (inactive) to 1e4 (active)
for (int k = 0; k < K; ++k) {
for (int r = 0; r < m; ++r)
sig[k * m + r] = std::pow(10.0, 4.0 * U(gen));
for (int j = 0; j < n; ++j)
rhs[k * n + j] = U(gen);
}
std::atomic<int> failures{0};
auto t0 = std::chrono::steady_clock::now();
{
std::vector<std::jthread> pool;
for (int p = 0; p < P; ++p)
pool.emplace_back([&, p] {
double Kmat[n][n];
// the batch dimension: independent nodes
for (int k = p; k < K; k += P) {
std::span<const double, m> sg(&sig[k * m], m);
std::span<double, n> x(&sol[k * n], n);
for (int j = 0; j < n; ++j)
x[j] = rhs[k * n + j];
assemble(S, sg, Kmat);
if (!cholesky_solve(Kmat, x))
failures.fetch_add(1, std::memory_order_relaxed);
}
});
}
auto dt = std::chrono::steady_clock::now() - t0;
double ms = std::chrono::duration<double, std::milli>(dt).count();
// check: residual of K x = b, recomputed from scratch
double worst = 0.0;
for (int k = 0; k < K; ++k) {
double Kmat[n][n];
assemble(S, std::span<const double, m>(&sig[k * m], m), Kmat);
for (int i = 0; i < n; ++i) {
double v = -rhs[k * n + i];
for (int j = 0; j < n; ++j)
v += Kmat[i][j] * sol[k * n + j];
worst = std::fmax(worst, std::fabs(v));
}
}
std::printf("K = %d nodes (n = %d, m = %d), %d threads: %.2f ms,\n"
"%d failed factorizations, worst residual %.2e\n",
K, n, m, P, ms, failures.load(), worst);
}
K = 4096 nodes (n = 8, m = 12), 8 threads: 0.79 ms,
0 failed factorizations, worst residual 1.37e-12
(Reading the listing's output) The listing compiles cleanly with clang++ -std=c++23 -fsyntax-only -Wall -Wextra. Built with -O2 -pthread, it factorizes and solves 4,096 condensed systems in about a millisecond on eight cores of a laptop. The wall time varies between runs, from half a millisecond to a few milliseconds when the threads start cold, and the residual does not vary. The worst residual is \(1.4 \times 10^{-12}\) for diagonals spanning eight orders of magnitude. The cost is \(O(n^2 m + n^3/3)\) per node, with no data-dependent control flow except the pivot test, which on a device is a flag per node rather than a branch. What parallelizes is the batch. Inside a node the small Cholesky is sequential in its columns, which is why the published GPU batched kernels assign one thread block to one node and keep the matrix in shared memory. Two published systems measure this pattern. ExaTron is a trust-region Newton solver for bound-constrained problems that runs "fully running on multiple graphics processing units" with no host–device transfer. It solves the many small nonconvex subproblems of a component-wise Lagrangian decomposition of optimal power flow, with linear scaling in the batch size and "more than 35 times speedup on 6 GPUs than on 40 CPU cores". WarpMPC runs ten thousand to over a hundred thousand identically structured quadratic programs per batch inside an alternating-direction method, with an unrolled \(LDL^\top\) factorization.Y. Kim, F. Pacaud, M. Schanen, K. Kim and M. Anitescu, "Leveraging GPU batching for scalable nonlinear programming through massive Lagrangian decomposition", SIAM Journal on Scientific Computing 47 (2025), B1133–B1157 (the 2021 arXiv version 2106.14995 has four authors); H. Hose, S. H. Jeon, C. Khazoom, S. Kim and S. Trimpe, "WarpMPC: large-batch MPC on GPU via ADMM with unrolled \(LDL^\top\) factorization", arXiv 2607.11603 (2026). Section 7.8 returns to the decomposition.
Where the barrier runs on a GPU today
Two production codes and one academic solver run an interior-point method on a device, and none of them is an NLP solver for MINLP. COPT 8 offers a GPU barrier for linear programming (GPUMode = 2, "high-performance GPU for barrier on LP"). In Mittelmann's LP feasibility benchmark it was first among the GPU codes on the H100 run of October 2025, with a scaled shifted geometric mean of \(1\) and 64 of 65 instances solved, and again on the B200 run of March 2026, with \(1\) and 64 solved. On the page of 16 September 2026 it is third behind HPR-LP-C and cuOpt, at \(1.29\) against \(1.00\) and \(1.17\), with 64 solved against 63 and 62. It has solved the most instances of any GPU code at every update.Cardinal Optimizer (COPT) User Guide, version 8.0, "Parameters" (GPUMode), guide.coap.online/copt/en-doc; H. D. Mittelmann, "Latest progress in optimization software", INFORMS Annual Meeting, 28 October 2025, slide 11; the PolyU talk of March 2026, slide 32; and the LPfeas page of 16 September 2026, plato.asu.edu/ftp/lpfeas.html. The GPU codes have a 1,000 s limit against 15,000 s for the CPU codes on those pages. NVIDIA's cuOpt added, in release 25.10, "a primal-dual interior-point method that uses GPU-accelerated sparse Cholesky and LDLT solves via cuDSS" for LP. Release 25.12 extended it to convex quadratic objectives, and release 26.06 to convex quadratic, second-order cone and rotated second-order cone constraints. It defaults the barrier to a relative accuracy of \(10^{-8}\), against \(10^{-4}\) for its first-order method.NVIDIA cuOpt User Guide, release 26.08, "Introduction" (the quoted description of the barrier method), "Convex optimization settings" (the default accuracies) and "Release notes", entries 25.10 (barrier for LP), 25.12 ("QP barrier (beta)"), 26.02 ("QP generally available") and 26.06 ("support for quadratic constraints (convex quadratic, second-order cone, and rotated second-order cone) in the barrier solver"), docs.nvidia.com/cuopt/user-guide/latest; release dates 14 October 2025, 11 December 2025, 11 February 2026 and 9 June 2026 from the PyPI and GitHub records. Section 7.6 refers back to these entries. CuClarabel, the GPU version of the Clarabel conic interior-point solver, is research software from Stanford and Oxford. It factorizes its KKT systems with cuDSS and parallelizes the cone operations by cone type. On quadratic programs its authors report that it is "more than 2 times faster than Gurobi, about 4 times faster than MOSEK and 10x times faster than the existing Rust implementation". They add that a mixed-precision option "can potentially achieve additional acceleration without compromising solution accuracy".Y. Chen, D. Tse, P. Nobel, P. Goulart and S. Boyd, "CuClarabel: GPU acceleration for a conic optimization solver", ACM Transactions on Mathematical Software 52 (2026), article 3; arXiv 2412.19027 (2024), Section 4.2 and the abstract. Authors' measurements on their own test sets, on an RTX 4090. For general nonconvex NLP the GPU stack is MadNLP with ExaModels and cuDSS, HiOp's condensed sparse–dense solver on CUDA and HIP, and ExaGO. None of these is wired into a global MINLP solver, and Gurobi 13's nonlinear barrier runs on the CPU.MadNLP.jl and ExaModels.jl, github.com/MadNLP/MadNLP.jl (MIT licence), with the GPU tutorial of its documentation; S. Peles, K. S. Perumalla, M. Alam, A. J. Mancinelli, R. C. Rutherford, J. Ryan and C. G. Petra, "Porting the nonlinear optimization library HiOp to accelerator-based hardware architectures", arXiv 2605.13736 (2026); S. Abhyankar, S. Peles, T. Becejac, J. Holzer, A. Mancinelli and C. Rutherford, "Exascale Grid Optimization (ExaGO) toolkit", arXiv 2203.10587 (2022); Gurobi Optimizer Reference Manual 13.0, release notes (the nonlinear barrier is a CPU preview with OptimalityTarget=1).
(The tolerance ladder, and its two halves for a MINLP) The tolerance ladder of these codes is the fact to carry forward. A barrier method on a CPU with a pivoted factorization stops at \(10^{-8}\) as a matter of course. The condensed GPU method stops at \(10^{-4}\) by default. It reaches \(10^{-6}\) to \(10^{-8}\) on well-behaved instances, at a cost in residual accuracy. Mittelmann's GPU quadratic-programming table shows the same shape for cuOpt's barrier and first-order codes together: 19 of 21 QPLIB instances solved at \(10^{-4}\), 16 at \(10^{-6}\) and 13 at \(10^{-8}\).H. D. Mittelmann, PolyU talk of March 2026, slide 35: 21 QPLIB QPs on a B200, 18 February 2026; cuOpt scaled means 1, 3.48, 8.29 with 19, 16, 13 solved at \(10^{-4}\), \(10^{-6}\), \(10^{-8}\); HPR-QP 3.05, 8.89, 10.1 with 19, 15, 15 solved. The slide states the device and the date; its time limit was not recorded for this post. For a MINLP the consequence has two halves. A local NLP solve at \(10^{-4}\) returns a point whose objective is an incumbent once the point has been repaired to feasibility, and it returns multipliers that certify nothing. A convex NLP relaxation solved at \(10^{-4}\) returns a value that is not a bound until a correction like Theorem 7.3.1 has been applied to it. For a nonlinear relaxation the correction requires a global minimization of the Lagrangian over the node's box, which is Section 7.8's sixth problem.
Where this is used
In October 2026 the GPU interior-point method is in production for LP (COPT, cuOpt) and for convex QP and conic programs (cuOpt), with CuClarabel as the research code for cones. It is in research use for large sparse nonconvex NLP in power systems and process engineering (MadNLP with ExaModels, HiOp, ExaGO). The global MINLP solvers of Section 5.3 call CPU NLP solvers (Ipopt, CONOPT, Knitro and others) through their NLP interfaces.
What parallelizes
What parallelizes is everything in Algorithm 7.5.5 except the elimination tree of the factorization and the scalar line search, and, across a frontier, the whole solve as a batch of identically shaped kernels. What it costs is the accuracy of the last two digits. For a tree that wants to prune, those two digits decide whether a node closes.
GPU QP, conic and SDP
The relaxations of Section 4 were not all linear programs. The cardinality portfolio of Section 4.4 has a convex quadratic relaxation. The perspective of Section 4.3 turns each indicator into a rotated second-order cone, so the strengthened relaxation is a second-order cone program. Section 4.7 bounded nonconvex quadratics by semidefinite programs. The tax problem of Section 9 is, once the trade directions are fixed, a quadratic program with a factor-model Hessian. Each of these is a convex program with a conic constraint, and each is solved at every node of the corresponding tree. This subsection asks which of them a GPU solves, and how. The answer has the same shape as for linear programming. First-order methods whose iteration is a few matrix–vector products move to the device and stop at \(10^{-4}\) to \(10^{-6}\). Interior-point methods move with the condensation of Section 7.5. The one new element is that a batch of small quadratic programs sharing one Hessian can share one factorization, which makes the batched QP cheaper per problem than the batched LP. The portfolio problems of the monograph are exactly such batches, and the subsection ends with what they look like on a device.
Definition 7.6.1 (the convex QP in two forms). A convex quadratic program is
\[\min_x\ \tfrac12 x^\top Q x + c^\top x \quad \text{s.t.} \quad Ax \in [b^L, b^U],\quad l \le x \le u, \tag{7.6.1}\]with \(Q \succeq 0\). The first-order codes write it in two special forms. PDQP's analysis uses inequality rows \(Ax \le b\) and free \(x\) (the PDQP.jl code also accepts equalities and a box). OSQP uses a single two-sided row block \(l \le Ax \le u\) in which the variable bounds are rows of \(A\). A conic program replaces the box and the rows by \(Ax + s = b\), \(s \in \mathcal K\) for a product \(\mathcal K\) of cones (nonnegative orthant, second-order cones, exponential cones, semidefinite cones). The saddle form that the first-order methods iterate on is
\[\min_{x \in X}\ \max_{y}\ \tfrac12 x^\top Q x + c^\top x - y^\top (Ax - b),\]the Lagrangian of Section 2.2 with the box kept as the primal domain. Its KKT residuals and relative tolerance are those of Definition 7.2.2, with the quadratic term included in the dual residual.
Throughout this subsection the Hessian is \(Q\) and the linear term is \(c\), as in (7.6.1). OSQP's own notation writes them as \(P\) and \(q\). The algorithm box and the listing below keep \(Q\) and \(c\), so that \(q\) remains the dual function of Section 2.2.
First-order quadratic programming
The PDHG iteration of Section 7.2 handles a smooth term by an explicit gradient step, at a cost in the step-size condition.
Theorem 7.6.2 (Condat, 2013; Vũ, 2013). Let \(f\) be convex with an \(L_f\)-Lipschitz gradient and add it to the saddle problem of Definition 7.2.4, so that the primal update becomes \(x^{k+1} = \operatorname{prox}_{\tau g}\!\big(x^k - \tau \nabla f(x^k) - \tau K^\top y^k\big)\) with the dual update unchanged. If
\[\frac{1}{\tau} - \sigma \lVert K \rVert^2 \;>\; \frac{L_f}{2},\]the iterates converge to a saddle point. Proof in the two papers.L. Condat, "A primal–dual splitting method for convex optimization involving Lipschitzian, proximable and linear composite terms", Journal of Optimization Theory and Applications 158 (2013), 460–479, Theorem 3.1; B. C. Vũ, "A splitting algorithm for dual monotone inclusions involving cocoercive operators", Advances in Computational Mathematics 38 (2013), 667–681. In words, the gradient step on \(f\) is an explicit step, and an explicit step on a function with curvature \(L_f\) is stable only when the step is short compared with \(1/L_f\). The coupling term already consumes part of the stability budget through \(\sigma \lVert K \rVert^2\). The two demands add, and the quadratic's share of the budget is \(L_f/2\). With \(f = \tfrac12 x^\top Q x\) and \(L_f = \lVert Q \rVert\) this is PDHG for the QP. Each iteration costs one product with \(Q\), one with \(A\) and one with \(A^\top\), all parallel over nonzeros. On the subspace where \(Q\) vanishes the problem is a linear program, and restarts give the linear rate of Theorem 7.2.9. On the range of \(Q\) the problem is a strongly curved quadratic, and a different idea applies.
(Acceleration, and where it enters PDHG) That idea is acceleration. A plain gradient method on a strongly convex quadratic whose curvature lies between \(\alpha\) and \(L\) needs \(O((L/\alpha)\log(1/\varepsilon))\) steps, because each step shrinks the error by a factor of about \(1 - \alpha/L\). Nesterov's accelerated method evaluates the gradient not at the current iterate but at an extrapolated point, the momentum point, which leans in the direction of the last step. The extrapolation damps the oscillation across the steep directions that limits the plain method, and the step count falls to \(O(\sqrt{L/\alpha}\,\log(1/\varepsilon))\). Lu and Yang apply this to the primal step of PDHG. The gradient of the quadratic is evaluated at a momentum point \(x_{\mathrm{md}}\), a weighted average of the running average and the current iterate. The dual step is extrapolated as in PDHG. The weights \(\beta_k\), \(\theta_k\) and the step sizes \(\eta_k\), \(\tau_k\) grow with the inner iteration count \(k\). Restarts to the running average supply the linear rate, exactly as in Theorem 7.2.9.
Algorithm 7.6.3 Restarted accelerated PDHG for convex QP
(rAPDHG; Lu and Yang)
Input QP in the form min 1/2 x'Qx + c'x s.t. Ax <= b
(x free; dual y >= 0); start (x^{0,0}, y^{0,0});
restart length K; per-inner-step parameters
beta_k = (k+2)/2,
theta_k = k/(k+1),
eta_k = (k+1)/(2(||Q|| + K ||A||)),
tau_k = (k+1)/(2 K ||A||)
Output (x^{n,0}, y^{n,0}) after n restarts
1. for n = 0, 1, 2, ...:
set the running averages (xbar, ybar) := (x^{n,0}, y^{n,0})
and x^{n,-1} := x^{n,0}
2. for k = 0, ..., K - 1:
3. x_md := (1 - 1/beta_k) xbar + (1/beta_k) x^{n,k}
# the momentum point
4. y^{n,k+1} := max(0, y^{n,k}
+ tau_k (A (theta_k (x^{n,k} - x^{n,k-1})
+ x^{n,k})
- b))
# extrapolated dual step
5. x^{n,k+1} := x^{n,k} - eta_k (Q x_md + c + A' y^{n,k+1})
# accelerated gradient step
6. xbar := (1 - 1/beta_k) xbar + (1/beta_k) x^{n,k+1};
ybar likewise # weighted averages
7. restart: (x^{n+1,0}, y^{n+1,0}) := (xbar, ybar)
Invariant
within an inner loop the accelerated PDHG bound on the duality gap
of the average holds; across restarts the distance to the optimal
set contracts by a fixed factor (Theorem 7.6.4).
Cost per step
one product with Q, one with A, one with A' and O(m + n) vector
work.
Parallel
all of it, as for Algorithm 7.2.15; the extrapolation and the
averages are elementwise.
(Two definitions behind the theorem) Two definitions from the paper are needed to state its theorem. Write \(z = (x, y)\) for a primal-dual pair, \(\mathcal Z\) for the set of pairs with \(y \ge 0\), \(\mathcal Z^\star\) for the saddle points, and \(L(x, y) = \tfrac12 x^\top Q x + c^\top x + y^\top(Ax - b)\) for the Lagrangian of the form \(Ax \le b\). For \(\xi \ge 0\) the smoothed duality gap of \(z\) centred at \(\dot z\) is
\[G_\xi(z; \dot z) \;=\; \max_{\hat z \in \mathcal Z}\ \Big\{ L(x, \hat y) - L(\hat x, y) - \frac{\xi}{2}\lVert \hat z - \dot z \rVert^2 \Big\}.\]Without the penalty this is the duality gap of \(z\) measured against every competitor \(\hat z\), which is infinite when \(\mathcal Z\) is unbounded. The penalty keeps the competitors near \(\dot z\) and makes the quantity finite, as the ball does for the normalized duality gap of Definition 7.2.7. The smoothed gap satisfies quadratic growth with constant \(\alpha_\xi > 0\) on a set \(S\) if \(G_\xi(z; z^\star) \ge \alpha_\xi \operatorname{dist}^2(z, \mathcal Z^\star)\) for every \(z \in S\) and every \(z^\star \in \mathcal Z^\star\). In words, the gap grows at least quadratically with the distance from the solution set. Lu and Yang show that a convex QP satisfies this condition on bounded sets, and \(\alpha_\xi\) plays the part of the sharpness constant of Definition 7.2.7.
Theorem 7.6.4 (linear convergence of rAPDHG; Lu and Yang, Theorem 1). Fix \(\xi > 0\) and suppose the smoothed duality gap satisfies quadratic growth with constant \(\alpha_\xi\) on the ball of radius \(R = 3\operatorname{dist}(z^{0,0}, \mathcal Z^\star)/(1 - 1/e)\) around the starting point \(z^{0,0}\). With the parameters of Algorithm 7.6.3 and a restart length
\[K \;\ge\; \max\Big\{ \sqrt{\tfrac{32 e^2 \lVert Q \rVert}{\alpha_\xi}},\ \tfrac{32 e^2 \lVert A \rVert}{\alpha_\xi},\ \sqrt{\tfrac{64 \lVert Q \rVert}{\xi}},\ \tfrac{64 \lVert A \rVert}{\xi},\ \tfrac{\lVert Q \rVert}{\lVert A \rVert} \Big\},\]the restart points satisfy \(\operatorname{dist}(z^{n,0}, \mathcal Z^\star) \le e^{-n} \operatorname{dist}(z^{0,0}, \mathcal Z^\star)\). Hence an \(\varepsilon\)-accurate point is reached in \(O\big(\max\{\sqrt{\lVert Q \rVert/\alpha_\xi},\ \lVert A \rVert/\alpha_\xi,\ \sqrt{\lVert Q \rVert/\xi},\ \lVert A \rVert/\xi,\ \lVert Q \rVert/\lVert A \rVert\}\log(\operatorname{dist}(z^{0,0}, \mathcal Z^\star)/\varepsilon)\big)\) iterations, and this rate is optimal among a wide class of first-order primal-dual methods. Proof in the paper.H. Lu and J. Yang, "A practical and optimal first-order method for large-scale convex quadratic programming", Mathematical Programming 215 (2026), 771–808, online 2 July 2025; arXiv 2311.07710, Definitions 1 and 3 (the smoothed duality gap and quadratic growth), Theorem 1 and Section 4 (the lower-bound instances). The solver is PDQP.jl, which runs on CPU and GPU and is compared with SCS and OSQP in the paper. The square root on \(\lVert Q \rVert\) is the mark of acceleration. A plain gradient method pays the curvature ratio linearly, and an accelerated one pays its square root. The restarted primal-dual hybrid conjugate gradient method of Huang, Zhang, Li, Ge, Liu and Ye goes one step further. It replaces the gradient step on the quadratic by an inexact conjugate-gradient solve of the primal subproblem. The authors show that this keeps the linear rate with an improved constant, and on their large instances it is "approximately 5 times faster than rAPDHG and about 100 times faster than other existing methods" on the GPU.Y. Huang, W. Zhang, H. Li, D. Ge, H. Liu and Y. Ye, "A restarted primal-dual hybrid conjugate gradient method for large-scale quadratic programming", INFORMS Journal on Computing (2025), doi 10.1287/ijoc.2024.0983; arXiv 2405.16160 (2024), abstract and Section 1; authors' measurements. The conic generalization is PDCS, a matrix-free primal-dual method for LP, second-order cone, convex QP and exponential-cone programs with a GPU implementation (cuPDCS). Another member of the family is HPR-QP, the Halpern Peaceman–Rachford method for convex composite quadratic programming that Mittelmann's GPU QP table of Section 7.5 compares with cuOpt.Z. Lin, Z. Xiong, D. Ge and Y. Ye, "A practical GPU-enhanced matrix-free primal-dual method for large-scale conic programs", arXiv 2505.00311 (2025); K. Chen, D. Sun, Y. Yuan, G. Zhang and X. Zhao, "HPR-QP: a dual Halpern Peaceman–Rachford method for solving large-scale convex composite quadratic programming", arXiv 2507.02470 (2025).
The restarts of Algorithm 7.6.3 and Theorem 7.6.4's contraction
z = (x, y) run n = 0 run n = 1
z^{0,0} --K inner steps--> z^{1,0} --K inner steps--> z^{2,0}
then a restart to the averages (xbar, ybar)
distance to the saddle points Z*, at most
d d / e d / e^2
inside a run, step 3 puts the momentum point on a segment:
xbar o-----------------o----------------------------o x^{n,k}
x_md
|<-- 1/beta_k --->| a fraction of the segment,
beta_k = (k+2)/2
d = dist(z^{0,0}, Z*); after n restarts the bound is e^-n d
Operator splitting, and why a batch of QPs shares one factorization
The other first-order family for QP is the alternating direction method of multipliers, ADMM, which minimizes an augmented Lagrangian alternately over two blocks of variables and updates the multipliers after each sweep; its standard implementation is OSQP. It splits the QP into a linear system and a projection, and the linear system has the same matrix at every iteration.
Algorithm 7.6.5 OSQP-style ADMM for
min 1/2 x'Qx + c'x s.t. l <= A x <= u,
and its batched form
Input Q (n x n, PSD), c, A (m x n), bounds l <= u (finite or
infinite); penalty rho > 0, regularization sigma > 0,
relaxation alpha in (1, 2); for a batch: K problems sharing
Q and A, with their own (c_k, l_k, u_k)
Output (x, z, y) with A x = z in [l, u] and the KKT residuals below
tolerance
0. factorize M := Q + sigma I + rho A'A once, or the equivalent
quasi-definite KKT matrix (a symmetric matrix with a positive
definite and a negative definite diagonal block, which has an LDL'
factorization for every symmetric pivot order);
for a batch this is ONE factorization
1. for k = 0, 1, 2, ...:
2. x~ := M^{-1} (sigma x^k - c + A' (rho z^k - y^k))
# the linear system;
for a batch, M^{-1} applied to K right-hand sides
3. z~ := A x~
4. x^{k+1} := alpha x~ + (1 - alpha) x^k
5. z^{k+1} := clip( alpha z~ + (1 - alpha) z^k + y^k / rho, l, u )
# the projection: elementwise
6. y^{k+1} := y^k + rho (alpha z~ + (1 - alpha) z^k - z^{k+1})
# dual update: elementwise
7. stop when ||A x^{k+1} - z^{k+1}|| and
||Q x^{k+1} + c + A' y^{k+1}|| are small (relative to the
problem's scale)
Invariant
z^k is always in [l, u]; for a convex QP the iterates converge to a
primal-dual solution for any rho > 0.
Cost per step
one solve with the cached factor (a dense or sparse triangular
pair; with the Woodbury identity O(n r) for a rank-r plus diagonal
Q), one product with A, O(m + n) elementwise work.
Parallel
the batch dimension is a dense matrix-matrix product (the factor
times K right-hand sides); the clips and dual updates are one
thread per entry per problem; the stopping test is one reduction
per problem.
(Why a batch of QPs shares one factorization) The method is Stellato, Banjac, Goulart, Bemporad and Boyd's. Schubiger, Banjac and Lygeros ported it to CUDA as cuOSQP. They report it "up to two orders of magnitude faster than the CPU implementation" on large QPs, where the linear system is solved by a preconditioned conjugate gradient on the device.B. Stellato, G. Banjac, P. Goulart, A. Bemporad and S. Boyd, "OSQP: an operator splitting solver for quadratic programs", Mathematical Programming Computation 12 (2020), 637–672; M. Schubiger, G. Banjac and J. Lygeros, "GPU acceleration of ADMM for large-scale quadratic programming", Journal of Parallel and Distributed Computing 144 (2020), 55–67. The observation that matters for a tree is step 0. Consider a branch and bound over a convex MIQP, or the direction-fixed subproblems of the tax problem. The nodes differ only in their bounds \(l_k, u_k\) and sometimes in the linear term, so \(M\) is the same for all of them. One factorization serves the whole frontier. The per-iteration work is then a dense product of the cached factor with \(K\) right-hand sides, which is the operation a GPU performs best of all. For the LP there is no matrix to factorize, so the only gain there was the arithmetic intensity of Definition 7.1.2.
The bound that a tree needs from such a batch has a closed form, and it needs no sign conditions on the multipliers when the box is bounded.
Proposition 7.6.6 (a safe bound for a QP from any multiplier vector). Consider the QP (7.6.1) with its box written as rows of \(A\), so that the constraints are \(Ax \in [l, u]\) with finite \(l \le u\), and with \(Q \succ 0\) and optimal value \(z^\star\). For every \(y \in \mathbb{R}^m\), with \(w = c + A^\top y\),
\[z^\star \;\ge\; d(y) \;=\; -\tfrac12\, w^\top Q^{-1} w \;-\; \sum_{i=1}^m \max(y_i l_i,\ y_i u_i). \tag{7.6.2}\]If \(Q\) is only positive semidefinite, the first term is \(\min_x [\tfrac12 x^\top Q x + w^\top x]\), which is \(-\infty\) unless \(w \in \operatorname{range}(Q)\).
Proof. Let \(x\) be feasible and \(z = Ax \in [l, u]\). Then \(\tfrac12 x^\top Q x + c^\top x = \tfrac12 x^\top Q x + w^\top x - y^\top z\), which is at least \(\min_{x'} [\tfrac12 x'^\top Q x' + w^\top x'] - \max_{z' \in [l,u]} y^\top z'\). The first minimum is attained at \(x' = -Q^{-1} w\) with value \(-\tfrac12 w^\top Q^{-1} w\). The second maximum separates over coordinates, and a linear function of \(z'_i\) on \([l_i, u_i]\) is maximized at an endpoint. Take the infimum over feasible \(x\). ∎
(The QP bound as a Lagrangian dual) This is Proposition 7.3.5 for a quadratic objective. The function \(d(y)\) is the Lagrangian dual function with the row intervals treated as the constraint set, it is concave in \(y\), and it equals \(z^\star\) at the optimal multipliers. Any ADMM or PDHG dual iterate may be inserted. The looseness \(z^\star - d(y)\) vanishes at the rate at which \(y\) converges, not at the rate at which the primal residual does, for the reason of Section 7.3: \(d\) is a concave function of \(y\) alone, so its distance below \(z^\star\) is governed by the distance of \(y\) from the optimal multipliers and by nothing the primal iterate does. Evaluating it costs one product with \(A^\top\) and one solve with \(Q\), and for the Hessians of the monograph the solve is cheap.
For a factor-model Hessian \(Q = D + F F^\top\), with \(D \succ 0\) diagonal and \(F \in \mathbb{R}^{n \times r}\) the loadings on \(r\) factors (the \(k\) of Section 4.4), Proposition 4.4.7 gives every product \(Q^{-1} w\) in \(O(nr)\) operations by the Woodbury identity after one \(O(nr^2 + r^3)\) factorization, against \(O(n^2)\) with a dense factor, together with the storage count of Section 4.4: with \(n = 1{,}000\) assets and \(r = 50\) factors the dense covariance has \(500{,}500\) distinct entries and the factor form \(n + r(r+1)/2 + nr = 52{,}275\), the "about 52,000" of Section 4.4. The point here is that the same structure makes the inner solve of Algorithm 7.6.5 and the bound of Proposition 7.6.6 linear in \(n\).The factor-model form \(V = X\Sigma X^\top + D\) and the elimination of the exposure variables are as in N. Moehle, J. Gindi, S. Boyd and M. J. Kochenderfer, "Portfolio construction as linearly constrained separable optimization", Optimization and Engineering 24 (2023), 1667–1687, Appendix A, and N. Moehle, M. J. Kochenderfer, S. Boyd and A. Ang, "Tax-aware portfolio construction via convex optimization", Journal of Optimization Theory and Applications 189 (2021), 364–383, Section 2.
The factor-model Hessian and its Woodbury inverse
Q = D + F F'
+---------+ +---------+ +---+ +---------+
| | | \ | | | | r x n |
| n x n | = | \ | + |n x| +---------+
| | | \ | | r |
+---------+ +---------+ +---+
diagonal, D > 0
Q^-1 w = D^-1 w - D^-1 F (I + F' D^-1 F)^-1 F' D^-1 w
\________________/
r x r, formed once
O(n r) per product after one O(n r^2 + r^3) factorization,
against O(n^2) with a dense factor
storage, dense against factor form:
batched_qp.py, n = 40 assets, r = 4 factors: 820 against 210
Section 4.4, n = 1,000 and r = 50: 500,500 against 52,275
The script below is Algorithm 7.6.5 over a frontier of 256 nodes of a portfolio QP with 40 assets and a four-factor covariance. Each node fixes a random 30 percent of the assets, to zero or to a floor of 4 percent, and leaves the rest free below 25 percent. All nodes share the budget row \(e^\top x = 1\). One inverse of the shared matrix \(M\) is formed. Each ADMM iteration is then one \(40 \times 40\) by \(40 \times 256\) product and elementwise clips. The product is a dense matrix–matrix multiplication, called a GEMM in the vocabulary of the linear-algebra libraries, and it is the operation the device runs at its peak rate. At every check the safe bound (7.6.2) is evaluated for every node from its current dual iterate, with \(Q^{-1}\) applied through the Woodbury identity. It is compared with an incumbent set at the tenth percentile of the node values, so that an exact solve would prune 230 of the 256 nodes.
One ADMM iteration of batched_qp.py, on all 256 nodes at once
Minv, 40 x 40 right-hand sides, 40 x 256,
formed once one column per node
+-----------+ +---+---+---+- . . . -+---+
| | | | | | | |
| | x | 1 | 2 | 3 | . . . |256| one GEMM
| | | | | | | |
+-----------+ +---+---+---+- . . . -+---+
|
v
the clips and updates, entry by entry: column k is clipped to
node k's row intervals [lo, hi] (the budget row and its box)
|
v
at a check: d(y) for every column, one Woodbury product per node
# batched_qp.py: K portfolio QPs that share a factor-model Hessian and
# differ in their boxes, solved together by an OSQP-style ADMM with one
# factorization for the whole batch, with a safe dual bound per node.
#
# node k: min 1/2 x'Qx + c'x s.t. e'x = 1, l_k <= x <= u_k,
# Q = D + F F' (n assets, r factors)
#
# Safe bound: for ANY multiplier vector y on the rows A = [e'; I] with
# row intervals [lo, hi],
# z_k >= d(y) = -1/2 w'Q^{-1}w - sum_i max(y_i lo_i, y_i hi_i),
# w = c + A'y (weak duality; the box is compact)
import numpy as np
rng = np.random.default_rng(3)
n, r, K = 40, 4, 256
# factor loadings, specific variances
F = 0.15 * rng.standard_normal((n, r))
D = (0.10 + 0.20 * rng.random(n)) ** 2
Q = np.diag(D) + F @ F.T
gamma = 5.0
mu_ret = 0.02 + 0.10 * rng.random(n)
c = -mu_ret / gamma
Dinv = 1.0 / D
# Woodbury capacitance (r x r)
Cap = np.linalg.inv(np.eye(r) + F.T @ (Dinv[:, None] * F))
def Qinv(Wm):
"""Q^{-1} applied to the columns of Wm in O(n r) per column."""
t = Dinv[:, None] * Wm
return t - Dinv[:, None] * (F @ (Cap @ (F.T @ t)))
assert np.allclose(Qinv(np.eye(n)), np.linalg.inv(Q))
# the frontier: each node fixes a random 30% of the assets to 0 or to a
# floor, the rest free in [0, 0.25]
L = np.zeros((n, K))
U = np.full((n, K), 0.25)
for k in range(K):
fix = rng.random(n) < 0.3
zero = rng.random(n) < 0.6
L[fix & ~zero, k] = 0.04
U[fix & zero, k] = 0.0
# row intervals of A x
lo = np.vstack([np.ones((1, K)), L])
hi = np.vstack([np.ones((1, K)), U])
A = np.vstack([np.ones((1, n)), np.eye(n)])
sigma, rho, alpha = 1e-6, 0.5, 1.6
# ONE factorization for all K nodes
Minv = np.linalg.inv(Q + sigma * np.eye(n) + rho * (A.T @ A))
def safe_bound(Y):
"""d(y) column by column, vectorized over the batch."""
Wm = c[:, None] + A.T @ Y
return (-0.5 * np.sum(Wm * Qinv(Wm), axis=0)
- np.sum(np.maximum(Y * lo, Y * hi), axis=0))
def admm(iters, check_every=None, inc=None, z_ref=None):
"""Algorithm 7.6.5 on all K nodes; given inc, print the checks."""
X = np.clip(np.full((n, K), 1.0 / n), L, U)
Z = A @ X
Y = np.zeros((n + 1, K))
for it in range(1, iters + 1):
# batched GEMM: n x n by n x K
Xt = Minv @ (sigma * X - c[:, None] + A.T @ (rho * Z - Y))
Xn = alpha * Xt + (1 - alpha) * X
Zt = alpha * (A @ Xt) + (1 - alpha) * Z
# elementwise clips and updates
Zn = np.clip(Zt + Y / rho, lo, hi)
Y = Y + rho * (Zt - Zn)
X, Z = Xn, Zn
if ((check_every and it % check_every == 0)
or it in (10, 50, 100, 200, 500, 1000)):
if inc is not None:
d = safe_bound(Y)
prim = np.abs(A @ X - Z).max(axis=0)
dual = np.abs(Q @ X + c[:, None] + A.T @ Y).max(axis=0)
assert np.all(d <= z_ref + 1e-9), (
"a safe bound exceeded the node value")
print(f"{it:5d} {(d >= inc).sum():6d} "
f"{(z_ref - d).max():9.2e} "
f"{prim.max():8.1e} {dual.max():8.1e}")
return X, Y, Z
# reference solves (20,000 lockstep iterations)
Xr, Yr, Zr = admm(20000)
obj = lambda X: 0.5 * np.sum(X * (Q @ X), axis=0) + c @ X
# feasible points, so z_ref >= z_k >= d(y)
z_ref = obj(np.clip(Zr[1:], L, U) / np.clip(Zr[1:], L, U).sum(axis=0))
kkt = (np.abs(Q @ Xr + c[:, None] + A.T @ Yr).max(),
np.abs(A @ Xr - Zr).max())
# an incumbent that an exact solve would use to prune 90% of the nodes
inc = np.quantile(z_ref, 0.10)
print(f"n = {n} assets, r = {r} factors, K = {K} nodes;\n"
f" reference KKT residuals {kkt[0]:.1e} (dual), "
f"{kkt[1]:.1e} (primal)")
print(f"incumbent {inc:.5f} = 10th percentile of the node values:\n"
f" an exact solve prunes {(z_ref >= inc).sum()} of {K}")
print(f"Woodbury: Q^-1 w costs {2*n*r + n} multiply-adds against {n*n} "
f"for a dense factor;\n"
f" dense factor-model Hessian has {n*(n+1)//2} entries, "
f"factor form {n + r*(r+1)//2 + n*r}")
print("\n iter pruned (safe) max(z - bound) primal res dual res")
admm(1000, inc=inc, z_ref=z_ref)
n = 40 assets, r = 4 factors, K = 256 nodes;
reference KKT residuals 2.4e-15 (dual), 1.5e-14 (primal)
incumbent -0.01914 = 10th percentile of the node values:
an exact solve prunes 230 of 256
Woodbury: Q^-1 w costs 360 multiply-adds against 1600 for a dense factor;
dense factor-model Hessian has 820 entries, factor form 210
iter pruned (safe) max(z - bound) primal res dual res
10 213 1.39e-03 7.9e-02 3.1e-03
50 230 2.96e-07 2.2e-04 1.2e-04
100 230 6.89e-10 2.7e-07 6.1e-06
200 230 1.02e-14 7.3e-10 2.3e-08
500 230 1.04e-17 1.4e-14 2.8e-15
1000 230 1.04e-17 1.6e-14 3.5e-15
(Reading the table) The table reads as follows. After ten iterations the safe bound prunes 213 of the 230 prunable nodes, while the primal iterates still violate their rows by up to \(0.08\). After fifty it prunes all 230 with a worst looseness of \(3 \times 10^{-7}\) in objective units, and the primal residual is still \(2 \times 10^{-4}\). The bound converges faster than the iterate for a structural reason. The dual function \(d\) is a concave quadratic on each orthant of \(y\): the box term is piecewise linear, with kinks at \(y_i = 0\). The ADMM update keeps \(y_i = 0\) exactly on the rows whose clip is inactive, so the iterates stay on the kinks, and the error is quadratic in the dual error of the active rows, through \(Q^{-1}\). The assertion checks at every row of the table that no node's bound ever exceeded that node's reference value, the objective at a feasible point of the node, which is at least the node's value. The cost per iteration is one \(n \times n\) by \(n \times K\) product for the whole frontier, \(O(nK)\) clips and updates, and at a check one Woodbury product per node. No factorization is repeated, and nothing branches on data. A device would run the same lines with the product as a single GEMM and the clips as one kernel over \(n \times K\) entries, which is the pattern the batched MPC solvers below implement at scale.
One row of steps 5 and 6 of Algorithm 7.6.5, and its term in d(y)
step 5: v = alpha z~ + (1 - alpha) z^k + y^k / rho,
z^{k+1} = clip(v, l, u)
step 6: y^{k+1} = y^k + rho (alpha z~ + (1 - alpha) z^k - z^{k+1})
= rho (v - z^{k+1})
v < l_i l_i <= v <= u_i v > u_i
------*----[==================================]----*------> v
l_i u_i
z^{k+1} = l_i z^{k+1} = v z^{k+1} = u_i
y^{k+1} < 0 y^{k+1} = 0 y^{k+1} > 0
the row's term max(y_i l_i, y_i u_i) in (7.6.2):
y_i l_i 0, the kink at y_i = 0 y_i u_i
Interior points, cones and semidefinite programs on the device
The interior-point route of Section 7.5 carries over to cones unchanged. The KKT system of a conic program has the same saddle shape. CuClarabel condenses it, factorizes it with cuDSS and parallelizes the cone operations by cone type, with the quadratic-programming speedups already quoted and second-order cone relaxations of power-flow problems among its tests.Chen, Tse, Nobel, Goulart and Boyd (2026), Sections 4.2 and 4.3. NVIDIA's cuOpt accepts convex quadratic objectives and convex quadratic, second-order cone and rotated second-order cone constraints in its barrier method, in the releases dated in Section 7.5. The perspective relaxation of Section 4.3, whose constraint \(x_i^2 \le \theta_i z_i\) is a rotated cone, can therefore be handed to a GPU solver as it stands. COPT's GPU barrier is for linear programs only. No GPU feature was found in the parts of the MOSEK and CPLEX documentation that could be read for this post.
(Semidefinite programs, and the missing certificate) Semidefinite programs are the hard case. The semidefinite cone's projection is an eigenvalue decomposition, and its interior-point Newton system is dense. The GPU codes avoid both by working with a low-rank factorization \(X = R R^\top\) and an augmented Lagrangian. cuLoRADS solves MaxCut relaxations with matrices of order \(10^7\) in ten seconds to one minute on an H100. cuHALLaR reports speedups of 30 to 140 times on matrix completion, up to 135 times on maximum stable set and 15 to 47 times on phase retrieval, with an instance of eight million rows in 142 seconds.Q. Han, Z. Lin, H. Liu, C. Chen, Q. Deng, D. Ge and Y. Ye, "Accelerating low-rank factorization-based semidefinite programming algorithms on GPU", arXiv 2407.15049 (2024); J. M. Aguirre, D. Cifuentes, V. Guigues, R. D. C. Monteiro, V. H. Nascimento and A. Sujanani, "cuHALLaR: a GPU accelerated low-rank augmented Lagrangian method for large-scale semidefinite programming", arXiv 2505.13719 (2025). Preprints; authors' numbers. Mittelmann's sparse SDP benchmark of 1 October 2026 is the independent measurement. On 75 instances with a 40,000 s limit, cuLoRADS 1.0.0 on an 80 GB H100 has a scaled shifted geometric mean of 3.01. The CPU codes have 1 for COPT, 2.89 for MOSEK, 5.40 for CSDP, 5.32 for SDPT3, 7.99 for SDPA and 29.9 for SeDuMi. It solves 71 of the 75, with 14 at reduced accuracy. The page notes that it "is not applicable to general SDP problems since it requires prior knowledge of a tracebound", that is, of an upper bound on the trace of the optimal matrix.H. D. Mittelmann, "Several SDP codes on sparse and other SDP problems; also on GPUs", page dated 1 October 2026, plato.asu.edu/ftp/sparse_sdp.html. For the monograph the relevance of an SDP solver is as a bound engine. Here the GPU codes have the same gap that the first-order LP codes had before Section 7.3: a low-rank primal iterate is not a dual certificate. A primal iterate, however accurate, is one nearly feasible point, and its value says nothing about how far below it the optimum could lie; only a dual object, a vector of multipliers whose dual value has been computed and made rigorous, bounds the optimum from the other side, and the low-rank methods carry no such object to convergence. Mittelmann's QUBO benchmark of 21 September 2026 shows the consequence. Its exact solvers, where QuBowl and QUPLANE solve 22 of 23 unconstrained binary QPLIB instances and BARON, COPT, SHOT, McSparse, BiqBin and SCIP solve 8 to 13, all run on CPUs and all prove with CPU bounds. The GPU and annealing heuristics for the same problems report solution quality with no bound at all.H. D. Mittelmann, "Nonconvex QUBO-QPLIB benchmark", page dated 21 September 2026, plato.asu.edu/ftp/qubo.html: 23 instances, 12 threads, one hour, on the machine named on the page, which was not recorded for this post; scaled means QuBowl 1 (22 solved), QUPLANE 4.60 (22), BARON 20.8 (13), SHOT with Gurobi 21.3 (12), COPT 25.2 (13), McSparse 33.4 (12), BiqBin 76.3 (9), SCIP with CPLEX 77.3 (8). BiqBin's SDP-bounded branch and bound is N. Gusmeroli, T. Hrga, B. Lužar, J. Povh, M. Siebenhofer and A. Wiegele, "BiqBin: a parallel branch-and-bound solver for binary quadratic problems with linear constraints", ACM Transactions on Mathematical Software 48 (2022). A safe SDP bound from a GPU low-rank solver's dual iterate is Section 7.8's agenda.
The low-rank factorization of the GPU SDP codes
X = R R'
+-----------+ +---+ +-----------+
| | | | | |
| | = | | +-----------+
| | | |
+-----------+ +---+
in the semidefinite a low-rank factor
cone
X: projection onto the cone: an eigenvalue decomposition;
interior-point Newton system: dense
R R': an augmented Lagrangian in R, with neither of the two
Batches of small quadratic programs
The batched QP is older on the GPU than the batched LP, because machine learning and control needed it first. OptNet solves a batch of small quadratic programs by a batched interior-point method as a layer of a neural network, with the batch dimension of the tensor as the problem index. ReLU-QP writes the ADMM iteration of Algorithm 7.6.5 as a weight-tied network of rectified linear units, so that a model-predictive controller, a controller that re-solves a short-horizon QP at every time step, runs on the same hardware as the network it controls. WarpMPC solves ten thousand to over a hundred thousand identically structured QPs per batch, at 8,000 to 250,000 sequential-quadratic-programming iterations per second, by unrolling a sparse \(LDL^\top\) inside ADMM. MPAX implements restarted PDHG for LP and QP in JAX, with batching through vmap, automatic differentiation through the solve, and TPU as well as GPU backends.B. Amos and J. Z. Kolter, "OptNet: differentiable optimization as a layer in neural networks", ICML 2017, arXiv 1703.00443; A. L. Bishop, J. Z. Zhang, S. Gurumurthy, K. Tracy and Z. Manchester, "ReLU-QP: a GPU-accelerated quadratic programming solver for model-predictive control", ICRA 2024, arXiv 2311.18056; H. Hose, S. H. Jeon, C. Khazoom, S. Kim and S. Trimpe, arXiv 2607.11603 (2026); H. Lu, Z. Peng and J. Yang, "MPAX: mathematical programming in JAX", arXiv 2412.09734 (2024). Authors' numbers; MPAX is MIT-licensed. Two facts from this literature transfer to a MINLP tree. Every batched solver fixes the iteration budget per round, so that all problems in the batch execute the same instructions, and it retires or pads the problems that finish early. This is the lane-efficiency question of Section 6.4. And every one of them keeps the problem data on the device between solves, since the problems of one batch are related to those of the next, which is the warm-start question of Section 7.4.
A batch of small QPs under a fixed iteration budget per round
round 1 round 2 round 3
QP 1 ########## | ########## | ####...... |
QP 2 ########## | ###....... | .......... |
QP 3 ########## | ########## | ########## |
. . . | | |
# an iteration: the same instructions for every QP of the batch
. a QP that has finished early: retired or padded
the problem data stays on the device between solves
What the portfolio QP looks like on a device
The portfolio problems of Sections 4.4 and 9 have four properties that place them in the regime above, and the research programme of Section 7.8 is built on them. The separable–affine form, a problem whose objective is a sum of one-variable terms and whose only coupling is a set of affine equalities, and the alternating-direction method below are those of Moehle, Gindi, Boyd and Kochenderfer. The list of four properties is this monograph's reading of their papers.The separable–affine form \(\min \sum_i f_i(x_i)\) s.t. \(Ax = b\) and its ADMM are Moehle, Gindi, Boyd and Kochenderfer (2023), Sections 1 to 3 and 5; the two-stage method and its 744 monthly instances are Moehle, Kochenderfer, Boyd and Ang (2021), Sections 5 and 6. The grouping into four properties is the monograph's, not a list in either paper.
- Accounts are independent. A manager with many accounts solves many unrelated problems of one shape, and the first level of parallelism is a batch of accounts.
- Inside an account, every relaxation and every direction-fixed subproblem is a convex QP with the same factor model and different bounds. The frontier of a tree over trade directions and lot boundaries is therefore the batch of Algorithm 7.6.5 with one factorization, and the bound of Proposition 7.6.6 is its pruning test.
- The nonconvexities are univariate and separable: the tax kink at zero in each asset's cost function, and the whole-lot requirement. In the separable–affine form the proximal step of the alternating-direction method is a closed-form minimization of a piecewise quadratic per asset, one thread per asset per account. The only coupled step is a projection onto \(\{Ax = b\}\) with \(m = k + 1\) rows, the \(k\) factor exposures and the cash row. Moehle, Gindi, Boyd and Kochenderfer report a mean of 251 milliseconds per nonconvex instance and 152 per convexified one on a CPU. Their instances have about a thousand securities and 72 factors, and the KKT matrix is factorized once and cached.
- The duality gap is provably small relative to the account, by Theorem 5.4.7 (Shapley–Folkman) and Theorem 5.4.8 (Udell and Boyd), which Section 9 applies to the tax problem. When a tree is needed it branches on a few directions and lot boundaries, not on \(2^n\) patterns. The two-stage rounding of Moehle, Kochenderfer, Boyd and Ang matched the relaxation bound on 678 of their 744 monthly instances.
(The three shapes on a device) On a device the three steps of the ADMM have three shapes. The proximal step is a map over assets and accounts. The projection is a batched Schur solve, a solve with the small Schur complement of the projection's KKT system, of size \((k+1) \times (k+1)\) per account, preceded and followed by products with the shared \(n \times k\) exposure matrix. The stopping test is a reduction. The cardinality relaxations of Section 4.4 with the perspective are rotated-cone programs that cuOpt's barrier and CuClarabel accept as they are. On the three-asset instance of Section 4.4 the perspective with the uniform diagonal split closes 43 percent of the big-M gap at the default return floor, and the factor-model diagonal closes 72 percent. The table of Section 4.4 gives the figures across the return floor, and Section 4.5 explains why the uniform split is so much smaller. So the diagonal a factor model supplies is the one to extract. What is missing is the tree around the batch, with its bound, its branching on directions and its certificate in dollars. That is Section 9's plan and Section 7.8's first problems.
The separable-affine ADMM over many accounts: three shapes
asset 1 asset 2 . . . asset n
account 1 [ t ] [ t ] . . . [ t ]
account 2 [ t ] [ t ] . . . [ t ]
. . .
+-> proximal step: a map, one thread t per asset per account,
| each a closed-form minimization of a piecewise quadratic
| |
| v
| projection onto {Ax = b}, m = k + 1 rows (k exposures, cash):
| a product with the shared n x k exposure matrix
| -> a batched (k+1) x (k+1) Schur solve, one per account
| -> a product with the exposure matrix again
| |
| v
+-- stopping test, a reduction: until it passes, iterate again
Where this is used
In October 2026 GPU quadratic and conic programming is in production in cuOpt (first-order and barrier, QP and SOCP) and in the GPU first-order codes PDQP, PDHCG, cuPDCS and HPR-QP, with CuClarabel as the research code for conic interior points. The mixed-integer QP and MISOCP solvers of Section 5.3 (CPLEX, Gurobi, Xpress, MOSEK, SCIP) solve their node relaxations on the CPU with warm-started active-set or barrier methods. The GPU semidefinite codes serve as relaxation engines for structured problems, not yet as bound engines for exact solvers.
What parallelizes
What parallelizes is the first-order iteration in full, the batch of a frontier through one shared factorization, the separable proximal steps across assets and accounts, and the low-rank SDP iteration. What does not parallelize is the pivoted factorization that a CPU conic solver uses, and the eigenvalue decomposition that a general SDP projection needs. What is open is the bound. The safe QP bound above is exact arithmetic on a GPU dual, and for semidefinite relaxations no analogue has been written down.
Heuristics on the GPU
Everything in Section 7 so far has been about the dual side of a search: relaxations, bounds and their accuracy. The primal side is where the GPU has actually delivered. The reason is structural. A primal heuristic is either a search over many independent candidates, which is a batch, or a local search whose step is a score for every candidate move, which is a map followed by a reduction. In neither case does a wrong or inexact result cost correctness. An incumbent is checked for feasibility before it is used, and a bad heuristic wastes time and nothing else. This subsection covers the three primal methods that run on devices today, and the measurements behind them. The first is Feasibility Jump and its relatives as bulk computations. The second is the portfolio design in which several heuristics and an approximate first-order LP share a pool of solutions. The third is massively parallel multistart for the continuous part of a MINLP. It ends with the benchmark on which these methods are scored, the primal integral, and with what that benchmark does and does not say.
Feasibility Jump as a bulk computation
Section 3.2 described Feasibility Jump. It is a local search over integer assignments that minimizes the weighted violation \(F^w(x) = \sum_i w_i \max\{0, a_i^\top x - b_i\}\) of the constraints. It moves one variable at a time to the value that reduces the violation most. At a local minimum it raises the weights of the constraints that remain violated, so that the weights act as Lagrange multipliers. Its inner quantity is the jump value.
(The jump value, and its sequential costs) Proposition 3.2.10 defined the one-variable violation function and Section 3.2 gave the jump value as its minimizer, found by a sorted-slope scan, together with the sequential costs: \(O(1)\) to choose a move from a sample of scores, \(O(\eta \log \eta)\) to recompute a jump value, \(O(\eta\mu)\) to update the scores of the neighbours and \(O(m\mu)\) after a weight change, where \(\eta\) is the largest number of rows a variable appears in and \(\mu\) the largest number of variables in a row. The notation of the box below is theirs. With all variables but \(x_j\) fixed at the current point \(\bar x\) and \(d_i = b_i - \sum_{k \ne j} a_{ik} \bar x_k\), the violation as a function of \(x_j = t\) alone is \(G_j(t) = \sum_{i : a_{ij} \ne 0} w_i \max\{0, a_{ij} t - d_i\}\), with critical values \(t_{ij} = d_i / a_{ij}\).B. Luteberget and G. Sartor, "Feasibility Jump: an LP-free Lagrangian MIP heuristic", Mathematical Programming Computation 15 (2023), 365–388, Algorithms 1 and 2; Sections 3.2 and 5.4 give the method's record.
The picture is a convex polygonal function of one variable. Each row that contains \(x_j\) contributes a hinge: zero while the row is satisfied, then rising with slope \(w_i \lvert a_{ij} \rvert\) once the row becomes tight. The sum of the hinges is minimized where its slope changes sign, and that is at a kink or at an end of the interval.
(What a device inverts) The sequential implementation is lazy and incremental. It keeps every variable's jump value and score up to date by touching only the rows the last move changed. That dependence on the last move is what makes it sequential. A device inverts the design. It recomputes everything in bulk, once per round, from the current assignment. The recomputation is a map over constraints and variables, whose cost is paid by parallel lanes rather than by the clock. And it runs many assignments at once. Three terms from GPU programming appear in the box. The constraint matrix is stored in compressed sparse row form, CSR, once by row and once by column: for each row, the list of its nonzero columns and coefficients. A segmented sort sorts many short lists at once, one segment per column. A segmented scan accumulates within each segment, and a segmented argmax picks the best entry of each segment.
Algorithm 7.7.1 One round of Feasibility Jump
over a population of assignments on a device
Input constraints a_i' x <= b_i (CSR by row and by column), bounds;
a population X = [x^1 ... x^P] of assignments with weight
vectors w^1 ... w^P; the shared pool of incumbents
Output the population after one move each (or a weight update), and
any feasible assignment found
1. violations: for every (p, i): # one thread per row per assignment
v_i^p := max(0, a_i' x^p - b_i)
2. feasibility: if sum_i v_i^p = 0 for some p:
write x^p to the pool (a reduction per assignment)
3. jump values: for every (p, j): # one thread per column per assignment
d_i := b_i - a_i' x^p + a_ij x_j^p over the rows of column j;
critical values t_ij;
sort them (a small per-thread sort or a segmented sort over the
column) and scan the slopes to the minimizer v_j^p of G_j;
score s_j^p := G_j(x_j^p) - G_j(v_j^p)
4. local minimum test: if no s_j^p > 0:
w_i^p := w_i^p + 1 on the violated rows of assignment p
(elementwise), and
choose j* among the variables of one violated row
(a segmented argmax)
5. otherwise choose j*^p by sampling a few positive scores and taking
the largest (a sampled reduction per assignment)
6. move: x_{j*}^p := v_{j*}^p
Invariant
as in the sequential method, each move with a positive score
strictly reduces F^w for the current weights, and weights only
grow on constraints violated at a local minimum.
Cost per round
O(nnz(A)) per assignment for steps 1 and 3, plus the per-column
sorts; the whole population at once.
Parallel
everything; the only sequential dependency is between rounds.
Several variables of one assignment may be moved per round when
their rows do not intersect (a conflict-free set), at the price of
stale scores otherwise.
One round of Algorithm 7.7.1: a grid of threads per step
1. violations 3. jump values, scores
thread (p, i) thread (p, j)
x^1, w^1 --> [ v v v ... v ] --> [ s s s ... s ] --> j*^1
x^2, w^2 --> [ v v v ... v ] --> [ s s s ... s ] --> j*^2
... ... ... ...
x^P, w^P --> [ v v v ... v ] --> [ s s s ... s ] --> j*^P
rows i columns j |
| v
v 6. x_{j*}^p := v_{j*}^p
2. if sum_i v_i^p = 0: for every p; then the
x^p to the pool next round
j*^p 4. if no s_j^p > 0 (a local minimum): w_i^p + 1 on the
violated rows of x^p, and j* by a segmented argmax
over one violated row
5. otherwise the variable with the largest of a few
sampled positive scores
Step 3 of Algorithm 7.7.1 on CSR by column: sort, then scan
the rows i of each column j, one segment per column, end to end:
column 1 2 3 ...
rows i [ i i i ] [ i i ] [ i i i i ] ...
for one assignment p, the critical value t_ij of each entry:
[ t t t ] [ t t ] [ t t t t ]
| | |
v segmented sort: every segment on its own
[ t < t < t ] [ t < t ] [ t < t < t < t ]
| | |
v segmented scan: within each segment, the
| slope of G_j accumulated from the left
[ - - + ] [ - + ] [ - - + + ]
| | |
v v v
v_1^p v_2^p v_3^p ...
- / + the slope of G_j just right of that t_ij is < 0 / >= 0;
the minimizer v_j^p is the kink of the first +, where the
slope changes sign, or an end of the interval
This is the design the cuOpt paper describes for its GPU heuristics. It is also the design of the CHAP tabu search, a local search that forbids reversing recent moves, which implements its best-shift step with sort, scan and reduce primitives. The box leaves three questions open: parallel move selection under conflicts, a convergence statement for the weight updates as a dual ascent, and an extension to nonlinear constraints, whose one-dimensional violation function is no longer piecewise linear. They are Section 7.8's fifth problem.A. Çördük, P. Sielski, A. Boucher and K. Aatish, "GPU-accelerated primal heuristics for mixed integer programming", arXiv 2510.20499 (2025); G. K. Tjusila, A. Hoen, N.-C. Kempke, G. Mexi, T. Berthold, A. Gleixner, T. Koch and S. Pokutta, "CHAP: a hybrid GPU-CPU heuristic for MIP", arXiv 2605.05086 (2026), whose GPU side is "a native tabu search featuring a novel best-shift algorithm built on sort, scan, and reduce primitives". On the CPU, Feasibility Jump "now runs by default on FICO Xpress Solver 9.0" (Luteberget and Sartor 2023; the control is FEASIBILITYJUMP, integer, default \(-1\), FICO, Xpress Optimizer Reference Manual) and in HiGHS since version 1.11.0 (6 June 2025, release note: "Added the feasibility jump heuristic … This is on by default").
Portfolios around a shared pool
No GPU heuristic runs alone. The production design, in cuOpt and in the 2026 competition entries, is a portfolio. Several heuristics run on the device. An approximate first-order LP solve supplies a relaxation point. A few CPU heuristics run beside them, and in a complete solver a CPU tree runs as well. All of them read and write one pool of solutions. One component of the pool design needs a name: a probing cache is a store of the bound implications of fixing single variables, computed once and reused to detect infeasible partial assignments early.
Algorithm 7.7.2 Portfolio of GPU heuristics and a CPU tree
sharing a solution pool
(the cuOpt and CHAP design)
Input MILP after presolve; time limit; a pool P of points (feasible
solutions, relaxation points, promising infeasible points)
Output an incumbent, and a proof if the tree finishes
G1 (GPU) PDLP on the root relaxation to low accuracy (1e-4, a
one-second limit in cuOpt's heuristic loop, warm-started);
the relaxation point goes to P
G2 (GPU) Feasibility Jump or tabu search over a population
(Algorithm 7.7.1), seeded from P;
repaired points go to P
G3 (GPU) feasibility pump with batched roundings and a probing cache
for early infeasibility detection;
fix-and-propagate guided by the points in P
C1 (CPU) branch and bound: concurrent root solve, reliability
branching (optionally with batched PDLP), root cuts, node
presolve, workers that steal nodes from one another
C2 (CPU) fix-and-propagate variants, RINS and RENS sub-MIPs
seeded from P
Loop
every component reads the best incumbent from P before pruning or
restarting and writes improvements back;
termination on time, gap or an empty tree
Invariant
the dual bound comes only from C1 (exact LP or safe bound); the
primal bound from anyone; correctness never depends on the GPU
components.
Cost
G2 and G3 are embarrassingly parallel over candidates; C1 is the
sequential core.
Parallel
everything on the device side; across workers on the CPU side; the
pool is the only shared state.
The portfolio of Algorithm 7.7.2: five components, one pool P
device (GPU)
G1 PDLP, root LP G2 Feasibility Jump G3 feasibility pump,
to 1e-4 or tabu search fix-and-propagate
| ^ | ^ |
| relaxation seeds| |repaired |points |
| point | |points |in P |
v | v | v
+--------------------------------------------------------------+
| pool P: feasible solutions, relaxation points, promising |
| infeasible points; the only shared state |
+--------------------------------------------------------------+
^ | ^ |
| | best incumbent, | | seeds
| | before pruning | |
| v | v
C1 branch and bound C2 fix-and-propagate,
RINS and RENS sub-MIPs
host (CPU)
the dual bound comes only from C1; every component reads the
best incumbent from P and writes its improvements back
The measurements of this design are of three kinds. The cuOpt paper reports on the fusion of GPU PDLP as approximate LP solver, a probing cache, feasibility pump, Feasibility Jump and fix-and-propagate, the last three being the heuristics of Section 3.2. In ten-minute runs on an H100 it found 221 feasible solutions with a 22 percent mean primal gap on the presolved MIPLIB 2017 benchmark set.Çördük, Sielski, Boucher and Aatish (2025), abstract and Section 5; a preprint, authors' numbers. The MIP 2026 computational competition on GPU-accelerated primal heuristics was won by CHAP. The table collects its result and the rules under which it was obtained.
| code | where it runs | found a solution | mean primal gap | within 1 % | score (page scale) |
|---|---|---|---|---|---|
| CHAP (winner) | GPU tabu search + cuPDLPx; CPU feasibility pump and fix-and-propagate; shared pool | 47 of 50 | 6.03 % | 19.3 (avg) | 33.84 ± 0.30 |
| Gurobi, default mode | CPU | 44 of 50 | – | – | – |
| cuOpt, heuristics only | GPU | 43 of 50 | – | – | – |
CHAP coordinates its GPU tabu search and cuPDLPx with CPU fix-and-propagate and feasibility pump through a pool that also holds infeasible candidates. The rules required entries to implement GPU heuristics without calling a MIP solver ("pure CPU submissions will not be considered"), and the score combined the primal integral and the final objective.Tjusila et al. (2026), Section 5; MIP Workshop 2026, "Computational competition: GPU-accelerated primal heuristics for MIP" (the Land–Doig competition), rules and results, mixedinteger.org/2026/competition, organizers Tjandraatmadja (chair), Basciftci, Çördük, Gamrath, Hojny, Kronqvist and Lu; honourable mentions cuLocalMIP (Lin, Wang, Yang and Cai) and GPU-MIP (Zhu), whose papers were not read for this post. Which normalization of the primal integral the competition page uses, and whether it is scaled by 100, could not be verified; the score is quoted on the page's own scale. PaNGEA, from the same season, combines relaxation solves with batched local search and parallel sub-problem generation on the device. It reports gap-integral reductions of 8 to 18 percent from moving a single-node heuristic to the GPU, and a further 12 to 19 percent from exploring nodes in parallel, on 283 competition and MIPLIB instances.J. Pauphilet and Y. Wu, "PaNGEA: parallel node generation and exploration algorithm on GPU", arXiv 2610.03090 (2026); a preprint, authors' numbers. Kempke and Koch supply the fact that makes the approximate LP in slot G1 defensible. Fix-and-propagate heuristics guided by a PDLP solution at relative tolerance \(10^{-4}\) find solutions as good as those guided by a \(10^{-6}\) solution, at 1.5 times the speed on MIPLIB and up to 5 times on large instances. With them the authors solve an energy-system model of 243 million nonzeros and 8 million variables to a gap below 2 percent in under four hours, where, in the words of their abstract, state-of-the-art commercial solvers could not produce a feasible solution within two days.N.-C. Kempke and T. Koch, "Fix-and-propagate heuristics using low-precision first-order LP solutions for large-scale mixed-integer linear optimization", Mathematical Programming Computation (2026), doi 10.1007/s12532-026-00312-7; arXiv 2503.10344 (2025). The abstract names no solver, and the body of the paper was not consulted for this post; Section 7.4 quotes the same result. The pattern of these results is the same as that of Section 7.4's table. The device wins where there are many candidates and a loose answer suffices, and the bound is computed on the CPU.
Massively parallel multistart
The continuous part of a nonconvex MINLP needs incumbents too, and the oldest way to find them is to start a local method from many points. The theory is probabilistic and was worked out in the 1980s. Its assumptions, all starts independent and all advanced together, are the assumptions a device satisfies naturally.
(The two facts from Section 1.3) Proposition 1.3.12 gave the two basic facts. With \(N\) independent uniform starts and \(p\) the volume fraction of the global minimizer's basin, the probability that the global minimizer is missed is \((1 - p)^N\), which tends to zero and is never zero. And after \(N\) local searches have found \(w\) distinct minimizers, the posterior expectation of the number \(W\) of local minimizers under the uniform priors is \(\mathbb{E}[W \mid w, N] = w(N - 1)/(N - w - 2)\), with the stopping rule that stops when this estimate falls below \(w + \tfrac12\). One further classical result is needed here.
Theorem 7.7.3 (multi-level single linkage; Rinnooy Kan and Timmer, 1987). Let \(f\) be continuous on a compact \(S \subset \mathbb{R}^n\) with finitely many local minimizers, each with a region of attraction of positive measure under a fixed deterministic local solver. In iteration \(k\) draw \(N\) further uniform sample points and start a local search only from a sample point with no better sample point within a critical distance \(r_k\) that shrinks like \((\log(kN)/(kN))^{1/n}\). Then with probability one every local minimizer is found in finitely many iterations, and only finitely many local searches are ever started. Proof in the two 1987 papers.A. H. G. Rinnooy Kan and G. T. Timmer, "Stochastic global optimization methods, Part I: clustering methods" and "Part II: multi level methods", Mathematical Programming 39 (1987), 27–56 and 57–78. The exact constant in the critical distance was not verified against the paper and is stated qualitatively. The miss probability and the Bayesian estimate are Proposition 1.3.12 of Section 1.3, with their sources.
(What the theorem says about starting, and the estimate about stopping) The theorem says that the sample itself, and not the local solver, decides which starts are worth running: a sample point with a better neighbour within the critical distance is expected to share that neighbour's basin, so its local search is skipped. The estimate of Proposition 1.3.12(ii) says the same about when to stop. For fixed \(w\) the estimate falls toward \(w\) as \(N\) grows, because many starts that find nothing new are evidence that nothing new exists. The rule stops when the expected number of unseen minimizers is below one half. A worked instance of Proposition 1.3.12(i) is Section 1.3's. With basins of 20, 50 and 30 percent of the box and the deepest valley the smallest, one start finds it with probability \(0.2\). Ten starts miss it with probability \(0.8^{10} = 0.107\), and thirty with probability \(0.8^{30} = 0.0012\). On a device the natural form of multistart is the one Proposition 1.3.12 assumes and a CPU never runs: all \(N\) starts advance together.
Algorithm 7.7.4 Lockstep multistart on a device
Input f (and constraints) as a straight-line program on the device;
box S; batch size N; a fixed iteration budget T; a clustering
tolerance; the stopping rule of Proposition 1.3.12(ii)
Output a set of distinct local minimizers with their values; the best
one is the incumbent candidate
1. sample N uniform points of S (one thread per point)
2. for t = 1..T:
every point takes one step of the same local method
(projected gradient, or a fixed-pivot Newton step with the
condensed system of Section 7.5)
with a step rule that needs no per-point branching
(a fixed step, or a fixed number of backtracking trials
evaluated for all points at once)
3. cluster the T-th iterates
(a nearest-neighbour pass; on the device a sort by a hash of the
rounded coordinates followed by a segmented reduction);
w := number of clusters
4. verify each cluster representative: KKT residual and the projected
Hessian's smallest eigenvalue (one small kernel per representative);
discard saddle points and boundary artefacts
5. stop if w (N - 1) / (N - w - 2) < w + 1/2;
otherwise double N and repeat with new samples
Invariant
every representative that passes step 4 is a local minimizer of f
on S; the best of them is a valid incumbent once its feasibility
has been checked to tolerance.
Cost
N T evaluations of f and its gradient, all of the same shape; the
clustering is O(N log N).
Parallel
steps 1, 2 and 4 across points; step 3 across clusters after the
sort; the stopping test is a scalar.
Lockstep multistart (Algorithm 7.7.4): N points, T rounds, a test
1. N uniform points of S; 2. T rounds of the same step for every
point, one thread per point and no per-point branching
point 1 o -------> o -------> o -- ... --> o --+
point 2 o -------> o -------> o -- ... --> o --+
... |
point N o -------> o -------> o -- ... --> o --+
round 1 round 2 round T |
v
3. cluster the T-th iterates: w clusters
|
v
4. verify each representative: KKT residual and
smallest eigenvalue; discard saddle points and
boundary artefacts
|
v
5. w (N - 1) / (N - w - 2) < w + 1/2 ?
| no | yes
v v
double N, new samples, stop: the best verified
and back to 1 minimizer is the
incumbent candidate
The script below runs the algorithm on the six-hump camel function, a standard two-variable test function with six local minimizers, two of them global, on the box \([-2, 2] \times [-1, 1]\), the function and box of the interval-partition script of Section 7.8. It uses a fixed-step projected gradient method with 3,000 rounds for every start. One batch of 102,400 starts is run. It is then sliced into 200 independent runs of each size \(N\). From these the script estimates the probability of finding a global minimizer and compares it with \(1 - (1 - \theta_1)^N\), where \(\theta_1\) is the measure of the two global basins estimated from the whole batch.
# lockstep_multistart.py: multistart as a GPU would run it.
#
# All starts advance together, one fixed-step projected-gradient
# iteration per round (no per-start line search), on the six-hump
# camel function
#
# f(x, y) = (4 - 2.1 x^2 + x^4/3) x^2 + x y + (-4 + 4 y^2) y^2
#
# on [-2, 2] x [-1, 1], whose minimum is -1.031628 at
# (+-0.0898, -+0.7127). Prints the number of distinct minima found
# and the Bayesian estimate w (N-1)/(N-w-2) of their total W
# (Proposition 1.3.12), and the probability of finding a global
# minimizer against 1 - (1 - theta)^N.
import numpy as np
rng = np.random.default_rng(11)
lo, hi = np.array([-2.0, -1.0]), np.array([2.0, 1.0])
def f(p):
"""The six-hump camel function; p[..., 0] is x, p[..., 1] is y."""
x, y = p[..., 0], p[..., 1]
return ((4 - 2.1 * x ** 2 + x ** 4 / 3) * x ** 2
+ x * y
+ (-4 + 4 * y ** 2) * y ** 2)
def grad(p):
"""The gradient of f, at the same points."""
x, y = p[..., 0], p[..., 1]
return np.stack([8 * x - 8.4 * x ** 3 + 2 * x ** 5 + y,
x - 8 * y + 16 * y ** 3], axis=-1)
def descend(S, step=1 / 80, rounds=3000):
"""Every start takes the same instruction stream: SIMD-shaped."""
for _ in range(rounds):
S = np.clip(S - step * grad(S), lo, hi)
return S
def cluster(E, tol=1e-3):
"""Distinct endpoints (the local solver's fixed points)."""
reps = []
for e in E:
if not any(np.linalg.norm(e - m) < tol for m in reps):
reps.append(e)
return reps
seeds, Ns = 200, [4, 8, 16, 32, 64, 128, 256, 512]
# one batch of 102,400 starts
S = descend(lo + (hi - lo) * rng.random((seeds * max(Ns), 2)))
reps = sorted(cluster(S[:20000]), key=f)
# measure of the two global basins
theta = np.mean(f(S) < -1.0316 + 1e-4)
print("endpoints after 3000 lockstep rounds\n"
f"from 102,400 uniform starts: {len(reps)} distinct points")
for m in reps:
print(f" ({m[0]:+.4f}, {m[1]:+.4f}) f = {f(m):+.6f}"
+ (" (global)" if f(m) < -1.0316 + 1e-4 else ""))
print("share of starts that reach a global minimizer: "
f"theta = {theta:.3f}\n")
print(" N P(global found) 1-(1-theta)^N"
" mean w E[W | w, N] max w")
for N in Ns:
runs = S[:seeds * N].reshape(seeds, N, 2)
ws = []
hit = 0
for run in runs:
ws.append(len(cluster(run)))
hit += bool(np.any(f(run) < -1.0316 + 1e-4))
w = np.mean(ws)
est = w * (N - 1) / (N - w - 2) if N > w + 2 else float("inf")
print(f"{N:4d} {hit / seeds:6.3f}"
f" {1 - (1 - theta) ** N:6.3f}"
f" {w:5.2f}"
f" {est:7.2f}"
f" {max(ws)}")
print("\nSection 1.3 check: basins 20 / 50 / 30 %%, deepest first:\n"
" miss probability 0.8^10 = %.3f, 0.8^30 = %.4f"
% (0.8 ** 10, 0.8 ** 30))
endpoints after 3000 lockstep rounds
from 102,400 uniform starts: 6 distinct points
(+0.0898, -0.7127) f = -1.031628 (global)
(-0.0898, +0.7127) f = -1.031628 (global)
(-1.7036, +0.7961) f = -0.215464
(+1.7036, -0.7961) f = -0.215464
(-1.6071, -0.5687) f = +2.104250
(+1.6071, +0.5687) f = +2.104250
share of starts that reach a global minimizer: theta = 0.602
N P(global found) 1-(1-theta)^N mean w E[W | w, N] max w
4 0.985 0.975 2.85 inf 4
8 1.000 0.999 4.07 14.76 6
16 1.000 1.000 5.11 8.62 6
32 1.000 1.000 5.78 7.39 6
64 1.000 1.000 5.98 6.73 6
128 1.000 1.000 6.00 6.35 6
256 1.000 1.000 6.00 6.17 6
512 1.000 1.000 6.00 6.08 6
Section 1.3 check: basins 20 / 50 / 30 %, deepest first:
miss probability 0.8^10 = 0.107, 0.8^30 = 0.0012
(Reading the table) The script evaluates \(102{,}400 \times 3{,}000\) gradients, all of the same shape, and the clustering passes are the only sequential part. Every line of the descent would run as one kernel over the batch on a device. The table reads as follows. The fixed-step method with no per-start decisions reaches the six local minimizers of the function to three decimals from every one of 102,400 starts (the clustering tolerance is \(10^{-3}\)). So the lockstep restriction cost nothing here beyond iterations. The two global basins cover 60 percent of the box, so four starts already find a global minimizer 98.5 percent of the time, against the independence prediction \(1 - 0.398^4 = 0.975\). The Bayesian estimate of Proposition 1.3.12(ii) is the useful column. It is infinite at \(N = 4\), where there are too few starts to say anything. It is \(14.8\) at \(N = 8\), when four minima have been seen on average. It falls through \(8.6\), \(7.4\) and \(6.7\) to \(6.35\) at \(N = 128\), where the stopping rule \(\mathbb{E}[W \mid w, N] < w + \tfrac12 = 6.5\) fires with all six minima in hand, and to \(6.08\) at \(N = 512\). The whole computation is one batch. The cost of 512 starts is the cost of 4 on a device with enough lanes, and the cost of certainty, in the limited sense of Proposition 1.3.12, is the number of rounds and not the number of starts. What the toy leaves out is the hard part of a MINLP incumbent: the integers. The practical form of multistart in a MINLP solver fixes the integer variables from a pool of candidate assignments and runs the continuous local solves in batch. That is the batched local NLP of Section 7.5 with a different source of right-hand sides. BARON runs a randomized multistart local search in its preprocessing (NumLoc). LINDO's global solver combines multistart with its branch and bound. Knitro's multistart runs its local solves on parallel threads and is "deterministic, even when run in parallel using multiple threads".GAMS/BARON manual (NumLoc, default \(-2\): BARON chooses the number of local searches from problem characteristics); Y. Lin and L. Schrage, "The global solver in the LINDO API", Optimization Methods and Software 24 (2009), as cited in Section 5.3; Artelys, Knitro User Guide, "Multi-start" (ms_enable, ms_maxsolves, ms_numthreads), with the quoted sentence and the sentence "Knitro cannot guarantee that multi-start will find the global optimum". The constraint-consensus multistart of SCIP and the OQNLP scatter search of Z. Ugray, L. Lasdon, J. Plummer, F. Glover, J. Kelly and R. Martí, "Scatter search and local NLP solvers: a multistart framework for global optimization", INFORMS Journal on Computing 19 (2007), 328–340, are the CPU designs of Section 3.2. The lane-efficiency model of Section 6.4 says what to design against. A local solver whose work spreads by a factor \(1.65\) between the median and the 84th percentile leaves more than half the lanes idle at \(b = 256\). The fixed iteration budget of Algorithm 7.7.4 is therefore required by the lane model, and the script's fixed round count is not a simplification.
What the benchmark measures
(The primal integral, and the benchmark's scale) Primal heuristics are scored by the primal integral of Definition 2.3.4, \(P(T) = \int_0^T p(t)\,dt\) with \(p(t)\) the primal gap function of Definition 2.3.3, equal to \(1\) while no incumbent exists, and by its time average \(\bar P(T) = P(T)/T \in [0, 1]\). Mittelmann's MIPFEAS benchmark reports the time average, and every \(\bar P\) in the table below is on the benchmark's own scale, which differs from the monograph's.T. Berthold, "Measuring the impact of primal heuristics", Operations Research Letters 41 (2013), 611–614, for \(P(T)\) with the integrand \(1\) while no incumbent exists, which is the convention of Definition 2.3.3. The MIPFEAS pages use a different integrand: \(2\) while no incumbent exists, \(1\) when the incumbent and the optimum differ in sign, and otherwise \(|z(t) - z^\star| / \max(|z(t)|, |z^\star|, 1)\), as printed on H. D. Mittelmann's MIPFEAS page of 16 September 2026, plato.asu.edu/ftp/mipfeas.html, and used by M. Bussieck and S. Dirkse, "Expanding the focus: introducing the MIPFEAS benchmark", GAMS blog, 17 March 2026, updated through 16 September 2026, gams.com/blog/2026/03. Both pages read on 5 October 2026; Section 2.3 records the same reading. MIPFEAS values are therefore not on the scale of Definition 2.3.4. The table collects the MIPFEAS numbers.
| run | code | device | \(\bar P\) | feasible found | proven optimal |
|---|---|---|---|---|---|
| March 2026 | cuOpt 26.02 | H100 | 0.0651 | 225 | 57 |
| March 2026 | cuOpt 26.02 | GB10 | 0.0887 | 223 | 60 |
| March 2026 | HiGHS 1.12 | CPU | 0.0567 | 212 | 97 |
| March 2026 | SCIP 10.0.1 | CPU | 0.0755 | 204 | 77 |
| March 2026 | CBC | CPU | 0.1121 | 192 | 77 |
| March 2026 | virtual mean commercial solver | CPU | 0.0132 | 227 | 180 |
| September 2026 | cuOpt 26.08 | B200 | 0.0286 | 226 | 78 |
| September 2026 | HiGHS 1.15.1 | CPU | 0.0453 | 214 | 107 |
| September 2026 | SCIP 10.0.1 | CPU | 0.0755 | 204 | 77 |
| September 2026 | virtual mean commercial solver | CPU | 0.0105 | 227 | 183 |
(Two readings, and a small example) In March 2026 HiGHS and cuOpt stood at about 4.3 and 4.9 times the commercial mean. Between the March and the September tables cuOpt's mean fell from \(0.0651\) (release 26.02 on an H100) to \(0.0286\) (release 26.08 on a B200), a change that mixes a new release with a new device, while its proven count rose only from 57 to 78.H. D. Mittelmann, PolyU talk of March 2026, slide 36, and the GAMS MIPFEAS page, March and September 2026 versions, read 5 October 2026. The ratios 4.3 and 4.9 are computed from the rounded means on the slide. Two readings follow, and both are needed. cuOpt finds feasible solutions on more instances than any open-source CPU solver, and it proves fewer of them optimal than HiGHS, which is exactly its documentation's self-description. And a \(\bar P(T)\) ranking is a ranking of heuristics, not of solvers: a code that never proves anything can win it. A small example makes the point, computed with the integrand of Definition 2.3.3, so that \(p(t) = 1\) while no incumbent exists. Take \(T = 300\) seconds and \(z^\star = 1000\). A heuristic finds \(1300\) at 2 s, \(1150\) at 5 s, \(1060\) at 20 s and \(1030\) at 120 s, and never proves optimality. Its gaps are \(0.23077\), \(0.13043\), \(0.05660\) and \(0.02913\), so \(P(T) = 2 \cdot 1 + 3 \cdot 0.23077 + 15 \cdot 0.13043 + 100 \cdot 0.05660 + 180 \cdot 0.02913 = 15.6\) s and \(\bar P(T) = 0.052\). An exact solver finds \(1200\) at 40 s, \(1050\) at 90 s and the optimum at 180 s. Its gaps are \(0.16667\) and \(0.04762\), so \(P(T) = 40 + 50 \cdot 0.16667 + 90 \cdot 0.04762 = 52.6\) s and \(\bar P(T) = 0.175\). The heuristic scores 3.4 times better although it never reaches the optimum, because the 40 seconds without any incumbent cost the exact solver \(40/300 = 0.133\) on their own. Under the MIPFEAS integrand those empty seconds count twice, and the two integrals become \(17.6\) s and \(92.6\) s, with \(\bar P(T) = 0.059\) and \(0.309\) and a ratio of \(5.3\). For the research programme of the monograph the number to watch is the dual one, the dual integral of Definition 7.8.1 in Section 7.8.
Where this is used
In October 2026 the GPU primal side is in production in cuOpt, with Feasibility Jump, feasibility pump, local search and fix-and-propagate on the device, a probing cache and an approximate first-order LP. It is in the competition winner CHAP and in PaNGEA, and in the CPU multistart of every global MINLP solver. For nonconvex MINLP no GPU heuristic exists yet, since the moves of Algorithm 7.7.1 assume linear rows and the local solves of Algorithm 7.7.4 assume a continuous problem.
What parallelizes
What parallelizes is the whole of both algorithms, across candidates and across the rows and columns of each candidate's evaluation, with the pool as the only shared state. What does not parallelize is the proof, which every one of these systems leaves to a CPU tree.
The research agenda: ultra-fast nonconvex MINLP on GPUs
This subsection turns the preceding seven into a list of problems. It is written for the reader who intends to work on them. Each problem states what is known, the first experiment that would move it, and the risk that the experiment fails for a reason already visible. The list is ranked by two criteria at once: how much of a nonconvex MINLP solve the problem would move if solved, and how ready its pieces are. The cost model behind the ranking is the one Section 7 has established. On a device, thousands of similar small problems in one batch are cheap. One problem solved to \(10^{-9}\) is expensive. Anything that must be done in a particular order is expensive. A spatial branch and bound spends its time on relaxations, bound tightening, local solves and the tree, in roughly that order. The first three consist of many independent small computations, and the fourth is a sequence of dependent decisions. Ten ranked problems follow, numbered 1 to 10 in the table below, and the text refers to them by these numbers. Three further items close the subsection and are not ranked problems: a table of the problem structures the evidence favours, a measurement protocol, and a note on accuracy on which every other item depends.
| rank | problem |
|---|---|
| 1 | the bound: batched McCormick relaxations with warm starts and safe bounds |
| 2 | evaluation bounds: intervals and pointwise McCormick with directed rounding |
| 3 | tightening: propagation as a Jacobi fixed point, and batched OBBT |
| 4 | the tree: frontier batching, the pruning loss and when to stop a column |
| 5 | the primal side: nonlinear local search and multistart on the device |
| 6 | certified bounds from GPU nonlinear relaxations |
| 7 | learned branching on the device |
| 8 | reproducibility: deterministic reductions and batch schedules |
| 9 | mixed precision and the arithmetic of the bound |
| 10 | decomposition: Lagrangian bounds from batched block solves |
Four things do not exist as of 5 October 2026, and the list is built around their absence.
- No production solver runs a spatial branch and bound for general nonconvex MINLP on a device.
- No published GPU branch and bound beats a CPU solver to a proof of optimality on general MILP, let alone on MINLP.
- No benchmark measures any GPU code on MINLPLib instances. Mittelmann's GPU tracks are LP feasibility, convex continuous QPLIB and sparse SDP.
- No published study measures the scaling of a global MINLP solver across more than one node of a machine.
Each statement carries its date, since a single paper can falsify it.H. D. Mittelmann, "Benchmarks for optimization software", plato.asu.edu/bench.html, entries dated through 1 October 2026 (the three tracks marked "also for GPUs"), and "Mixed integer nonlinear programming benchmark (MINLPLIB)", page dated 26 February 2026, plato.asu.edu/ftp/minlp.html, which has no GPU column; NVIDIA cuOpt User Guide 26.08, "Introduction" ("proving feasible solutions optimal remains under active development"). The absence statements rest on the literature sweeps recorded for this post through the arXiv, Crossref and OpenAlex records and the sources of Sections 6 and 7. On the positive side, every ingredient of a global solver has a GPU prototype in some restricted setting. Sections 7.2 to 7.7 and Section 6.4 described them. They are first-order LP solvers competitive with the best CPU codes, a hybrid MILP design with heuristics on the device and the tree on the host, and batched LP solves inside strong branching and bound tightening. They are also two GPU bounding engines for deterministic global optimization, condensed-space interior points at \(10^{-4}\) to \(10^{-6}\), low-rank SDP solvers for structured problems, and a competition for GPU primal heuristics. The problems below record each prototype under "What is known", with its section.
Decomposition and Lagrangian bounds
(The decomposition theory the problems use) Before the list comes one piece of theory that several of the problems use, stated in Section 3.7. A nonconvex MINLP is often a collection of small nonconvex problems held together by a few constraints. Scenarios of a stochastic program are coupled by the design variables, sites by a shared resource, assets by factor exposures and a budget. The general-purpose solvers of Section 5.3 do not see the shape. They relax every nonconvex term over its box and branch on single variables. Decomposition sees it, and its subproblems are the batch a device wants. What problem 10 needs is the following. Definition 3.7.1 sets up the block-separable problem (3.7.1), with block sets \(S_i\), block objectives \(f_i\) and \(m\) coupling constraints \(\sum_i A_i x_i \le b\), and its Lagrangian relaxation \(q(\lambda) = \sum_i q_i(\lambda) - \lambda^\top b\), whose \(n\) block subproblems \(q_i(\lambda) = \min_{x_i \in S_i} [f_i(x_i) + \lambda^\top A_i x_i]\) are independent. Theorem 3.7.3 and Corollary 3.7.4 say that the dual value \(d^\star = \sup_{\lambda \ge 0} q(\lambda)\) is the value of the convexified problem, in which each block objective is replaced by its convex envelope along the coupling, and that it dominates every relaxation that convexifies the blocks term by term, the McCormick relaxation of each block on its box included. Theorem 3.7.6 (Lagrangian decomposition; Guignard and Kim) strengthens the bound by splitting the coupled variables into copies and pricing their agreement. Theorem 5.4.7 (Shapley–Folkman) and Theorem 5.4.8 (Udell and Boyd) bound the remaining gap by a sum of at most \(m\) block nonconvexities, however many blocks there are. Theorem 3.7.11 (Lagrangian cuts; Karuppiah and Grossmann) turns the block values into cuts for any convex relaxation, and Theorem 3.7.16 (nonconvex generalized Benders decomposition; Li, Tomasgard and Barton) is the other exact scheme. The worked example of Section 3.7, two sites sharing one crew of seven technicians, each site the bilinear running example behind a capacity switch, shows the ladder \(z_{\mathrm{LP}} = -0.3553 < z_{\mathrm{LR}} = -0.2540 < z_{\mathrm{LD}} = -0.2150 < z^\star = -0.1600\), with the Lagrangian dual maximized at the kink \(\lambda^\star = 0.026\), and its script prints the four values. Its tree table closes the instance in five nodes with the Lagrangian bound at the nodes, where a tree on the McCormick relaxation would have to branch on the continuous variables \(x_i\) and \(w_i\) as well. Problem 10 takes the block solves of that example to a device.
The bound: batched McCormick relaxations with warm starts and safe bounds
What is known. In spatial branch and bound the sibling nodes' relaxations share all their rows except the McCormick planes of the branched variable, whose coefficients depend on the new bounds (Theorem 2.4.7). A frontier of nodes is therefore almost, but not exactly, the shared-matrix family of Proposition 7.4.4. Batched PDHG over such a family is Algorithm 7.4.5. The safe bound of Theorem 7.3.1 turns any dual iterate into a valid node bound. Blin and coauthors measured speedups of 12 to 489 times over the dual simplex for strong branching on four structured MILP families at batch sizes of 32 to 512, and 25.7 times for bound tightening. They also measured that the warm-started dual simplex remains faster on most MIPLIB node LPs. The frontier experiment of Section 7.4 showed the safe bound pruning 228 of 230 prunable nodes after ten iterations and all 230 after fifty. With the column-retirement policy of Algorithm 7.4.5 the whole frontier took 4,160 column-iterations against 512,000 in lockstep.Blin, Gualandi, Maes, Lodi and Stellato (2026), cited in full in Section 7.4; the experiment is the one of Section 7.4 (batched_pdhg_frontier.py), whose Mode B retires each column when its safe bound crosses the incumbent or its relative KKT error falls below \(10^{-4}\). Nobody has published any of this for nonconvex relaxations.
The first experiment. Fix a McCormick relaxation skeleton for a factorable MINLP: the auxiliary variables of its DAG, the directed acyclic graph of its expressions from Section 2.4, and the envelope rows, with the rows' coefficients as functions of the box. For a frontier of boxes, generate the coefficients in one kernel. Store the shared rows once and the box-dependent rows as a per-node correction of low rank. Only the envelope rows of the branched variable change, and in each only the coefficients that carry its bounds (Theorem 2.4.7), so the per-node difference has rank at most the number of terms containing that variable. Run Algorithm 7.4.5 with the parent's dual as the warm start, prune on the safe bound, and compare node throughput and tree size against Couenne and SCIP on the fifty smallest nonconvex MINLPLib instances with at most twenty nonconvex variables. Record, for each node, the iteration at which its safe bound first crosses the incumbent and the iteration at which it comes within \(10^{-6}\) of the exact LP value, and fit the two distributions. The column-stopping rule is read off them. Quantify \(\operatorname{dist}(z^\star_{\mathrm{parent}}, Z^\star_{\mathrm{child}})\) against the width of the branched variable, since Theorem 7.2.9 says the warm start helps only through that distance inside a logarithm. Estimate also the child's sharpness constant \(\alpha\) (Definition 7.2.7) against the parent's on the same trees, since a systematically larger \(\alpha\) in the children would make warm starts across siblings pay more than the logarithm of Theorem 7.2.9 suggests.
Problem 1: K sibling relaxations as one batch (node k: A + D_k)
x^1 x^2 . . . x^K the K iterates
+------------------------------------+
A | all rows, with the coefficients | one copy: one
| that the siblings share | sparse matrix-
+------------------------------------+ matrix product
+ +--------+--------+---------+--------+
D_k | D_1 | D_2 | . . . | D_K | one correction
+--------+--------+---------+--------+ per node
D_k nonzero only in the envelope rows of the branched variable,
in the coefficients that carry its bounds (Theorem 2.4.7):
rank at most the number of terms containing that variable;
all K corrections generated from the boxes in one kernel
The risk. Coefficients that depend on the box break the shared-matrix assumption. The sparse matrix–matrix product of Proposition 7.4.4 then becomes a batch of sparse matrix–vector products with different matrices, and the arithmetic-intensity gain of Definition 7.1.2 shrinks. The hard nodes, whose bounds sit within the gap, need near-exact solves and dominate the time under every policy tried so far (Section 7.4). And the inflation of the tree by bound slack (Proposition 7.4.2) can eat the throughput gain, unless the safe bound, which has no slack but only looseness, is used throughout.
Evaluation bounds: intervals and pointwise McCormick with directed rounding
What is known. Two relaxation technologies have actually run on a device for factorable problems, and both are evaluations rather than linear programs. The first uses the mean value form of Section 6.4, the interval \(f(c) + F'(X)^\top (X - c)\) for a box \(X\) with centre \(c\) and an interval enclosure \(F'(X)\) of the gradient over the box, whose excess width shrinks quadratically as the box shrinks, against linearly for the natural interval extension of Section 2.6. Zhang and coauthors partition a node's box into \(m^d\) sub-boxes, evaluate the mean value form on all of them in one kernel inside MAiNGO, and take the minimum as the node bound. They report wall-clock times three orders of magnitude below CPU interval arithmetic without partitioning, and competitiveness with MAiNGO's default McCormick bounder on some case studies. The second technology is pointwise. Gottlieb, Xu and Stuber generate CUDA kernels that evaluate pointwise McCormick relaxations, interval extensions and subgradients for thousands of boxes or points at once, at about 9 nanoseconds per evaluation against 237 on the CPU by their measurement. Their branch and bound, ParBB, is 11 to 22 times faster than EAGO on their test problems, with a small operation set at the time of reading.Zhang et al. (2025) and Gottlieb, Xu and Stuber (2026), cited in full in Sections 6.4 and 7.1, and the SourceCodeMcCormick.jl repository; authors' numbers. Directed rounding per operation is through the CUDA intrinsics __dadd_rd, __dadd_ru, __dmul_rd, __dmul_ru, __fma_rd and __fma_ru of NVIDIA, CUDA Math API Reference Manual, with the implementation pattern of S. Collange, M. Daumas and D. Defour, "Interval arithmetic in CUDA", GPU Computing Gems Jade Edition (2012). The compiler contracts a product and a sum into one fused multiply-add by default (--fmad=true, NVIDIA, CUDA Compiler Driver NVCC, documentation of --fmad and --use_fast_math, read 5 October 2026), so a kernel that writes its two directed roundings with plain operators must be compiled with --fmad=false; a kernel written with the _rd and _ru intrinsics does not depend on the flag. Under a directed rounding mode a fused result is still a directed rounding of an exact quantity, so contraction breaks per-operation error accounting and bitwise reproducibility rather than the enclosure itself. That the intrinsics are never contracted is stated in NVIDIA's programming guide, whose current page could not be re-read for this post.
The script below computes both bounds on the six-hump camel function on \([-2, 2] \times [-1, 1]\), the function of the multistart script of Section 7.7, for partitions of the box into \(m \times m\) sub-boxes with \(m\) from \(1\) to \(64\).
# mean_value_partition.py: the two evaluation bounds of problem 2.
#
# The six-hump camel function
#
# f(x, y) = 4x^2 - 2.1x^4 + x^6/3 + xy - 4y^2 + 4y^4
#
# on [-2, 2] x [-1, 1], minimum -1.0316, bounded on an m x m partition
# of the box. Each sub-box is one independent evaluation (one lane on a
# device); the bound for the box is the minimum over the sub-boxes.
# Interval arithmetic without directed rounding, for illustration.
import numpy as np
def sq(lo, hi):
"""[lo, hi]^2 as an interval."""
if lo >= 0:
return lo * lo, hi * hi
if hi <= 0:
return hi * hi, lo * lo
return 0.0, max(lo * lo, hi * hi)
def mul(a, b):
"""The product of two intervals."""
p = [a[0] * b[0], a[0] * b[1], a[1] * b[0], a[1] * b[1]]
return min(p), max(p)
def add(*ts):
return sum(t[0] for t in ts), sum(t[1] for t in ts)
def scale(c, a):
return (c * a[0], c * a[1]) if c >= 0 else (c * a[1], c * a[0])
def f(x, y):
return 4 * x**2 - 2.1 * x**4 + x**6 / 3 + x * y - 4 * y**2 + 4 * y**4
def F(X, Y):
"""The natural interval extension, term by term."""
X2 = sq(*X)
X4 = sq(*X2)
Y2 = sq(*Y)
Y4 = sq(*Y2)
return add(scale(4, X2), scale(-2.1, X4), scale(1 / 3, mul(X4, X2)),
mul(X, Y), scale(-4, Y2), scale(4, Y4))
def G(X, Y):
"""An interval enclosure of the gradient (odd powers are monotone)."""
X3 = (X[0] ** 3, X[1] ** 3)
X5 = (X[0] ** 5, X[1] ** 5)
Y3 = (Y[0] ** 3, Y[1] ** 3)
return (add(scale(8, X), scale(-8.4, X3), scale(2, X5), Y),
add(X, scale(-8, Y), scale(16, Y3)))
def bounds(m):
"""The minimum over the m x m sub-boxes of the two lower bounds."""
xs = np.linspace(-2, 2, m + 1)
ys = np.linspace(-1, 1, m + 1)
nat = mvf = np.inf
for i in range(m):
for j in range(m):
X, Y = (xs[i], xs[i + 1]), (ys[j], ys[j + 1])
cx, cy = (X[0] + X[1]) / 2, (Y[0] + Y[1]) / 2
nat = min(nat, F(X, Y)[0])
# mean value form: f(c) + gx (X - cx) + gy (Y - cy)
gx, gy = G(X, Y)
mvf = min(mvf, f(cx, cy)
+ mul(gx, (X[0] - cx, X[1] - cx))[0]
+ mul(gy, (Y[0] - cy, Y[1] - cy))[0])
return nat, mvf
gx, gy = np.meshgrid(np.linspace(-2, 2, 2001), np.linspace(-1, 1, 1001))
fstar = f(gx, gy).min()
print(f"true minimum on the box (dense grid): {fstar:.4f}")
print()
print(f"{'':3} {'':9} {'natural':>12} {'mean value':>12}"
f" {'gap of the':>16}")
print(f"{'m':>3} {'sub-boxes':>9} {'extension':>12} {'form':>12}"
f" {'mean value form':>16}")
for m in [1, 2, 4, 8, 16, 32, 64]:
nat, mvf = bounds(m)
print(f"{m:3d} {m * m:9d} {nat:12.3f} {mvf:12.3f}"
f" {fstar - mvf:16.3f}")
true minimum on the box (dense grid): -1.0316
natural mean value gap of the
m sub-boxes extension form mean value form
1 1 -39.600 -322.400 321.368
2 4 -39.600 -88.017 86.985
4 16 -35.017 -38.244 37.212
8 64 -25.538 -13.908 12.876
16 256 -15.431 -4.309 3.277
32 1024 -8.123 -1.117 0.085
64 4096 -3.786 -1.052 0.020
(Reading the partition table) The script's cost is \(m^2\) interval evaluations of \(f\) and of its gradient per row of the table, each independent of the others, which is the batch. The minimum over the sub-boxes is a reduction. The table reads as follows. The mean-value lower bound over \(m \times m\) sub-boxes is \(-322.4\), \(-88.0\), \(-38.2\), \(-13.9\), \(-4.31\), \(-1.117\) and \(-1.052\) for \(m = 1, 2, 4, 8, 16, 32, 64\), against the true minimum \(-1.0316\). The gap to the true minimum falls by a factor of 38 from \(m = 16\) to \(32\), and by a factor near four, the quadratic rate, from \(m = 32\) to \(64\).The quadratic order of the mean value form is the standard result of interval analysis, R. E. Moore, R. B. Kearfott and M. J. Cloud, Introduction to Interval Analysis, SIAM (2009), and A. Neumaier, Interval Methods for Systems of Equations, Cambridge University Press (1990). The natural extension's bound is still \(-3.79\) at \(m = 64\), and it is tighter than the mean value form only on the coarsest partitions, where the gradient enclosure is wide. The \(4{,}096\) evaluations at \(m = 64\) are one small batch on a device and \(4{,}096\) sequential DAG evaluations on a CPU.
Problem 2 on the camel box [-2, 2] x [-1, 1]: m = 2 and m = 4
m = 2: 4 sub-boxes m = 4: 16 sub-boxes
+-------------+-------------+ +-------+-------+-------+-------+
| | | | * | * | * | * |
| * | * | +-------+-------+-------+-------+
| | | | * | * | * | * |
+-------------+-------------+ +-------+-------+-------+-------+
| | | | * | * | * | * |
| * | * | +-------+-------+-------+-------+
| | | | * | * | * | * |
+-------------+-------------+ +-------+-------+-------+-------+
natural extension -39.600 natural extension -35.017
mean value form -88.017 mean value form -38.244
* the centre c of a sub-box X, in f(c) + F'(X)^T (X - c)
one sub-box, one lane; the bound for the box is the minimum over
the sub-boxes, a reduction
m = 64: 4,096 sub-boxes; natural extension -3.786, mean value
form -1.052, against the true minimum -1.0316
The first experiment. Build a GPU-resident interval and McCormick evaluator for factorable DAGs with directed rounding and a complete operation library, the branch-free rules for each univariate function of Section 2.4's catalogue. Bound a frontier of boxes with one linearization point per box and with eight. Compare bound tightness per box and node throughput against the partitioned interval bounder and against MAiNGO's CPU McCormick LP bound at equal wall time, on the same fifty instances as problem 1. A convex relaxation's value and subgradient at one point give a supporting hyperplane. A hyperplane is minimized over a box in closed form, so a rigorous bound needs no LP at all, and several points give a small LP in \(n\) variables.
The risk. Partitioning is exponential in the dimension: \(64^2 = 4{,}096\) sub-boxes become \(64^5 \approx 10^9\). The partitioned bound is therefore confined to nodes with a few nonconvex variables, which is why problem 1's LP bound is still needed. A few linearization points give a weaker bound than the lifted LP, so the tree grows. And completing the operation library with branch-free, directed-rounding rules is necessary before the method can reach general MINLPLib instances at all.
Tightening: propagation as a Jacobi fixed point, and batched OBBT
What is known. Feasibility-based bound tightening (Section 2.6) applies, for each constraint, a map \(F_i\) that shrinks the box using that constraint's activity bounds, and iterates. A sequential solver applies the maps in sequence, each seeing the box the previous one left, which is a Gauss–Seidel schedule. A device applies all of them to the box at the start of the round and merges the results at the end, which is a Jacobi schedule. The two reach the same fixed point.
(Two remarks on the fixed-point theorem) By Theorem 2.6.11 the two schedules reach the same greatest fixed point, provided the contractors, the maps of Section 2.6 that shrink a box without discarding any feasible point, are monotone, contracting and continuous on decreasing chains, which the activity-bound contractors of linear and factorable constraints are. Proposition 2.6.12 says that the Gauss–Seidel round is never behind the Jacobi round after the same number of rounds, so the Jacobi schedule needs at least as many rounds to reach any given box. Two remarks are added here. The first is that the continuity hypothesis of Theorem 2.6.11 is not decorative. On \([0, 10]\) take \(F([a, x]) = [a, g(x)]\) with \(g(x) = x\) on \([0, 1/2]\), \(g(x) = 1/2\) on \((1/2, 1]\) and \(g(x) = (x + 1)/2\) for \(x > 1\). This \(F\) is monotone and contracting, its iteration from \([0, 10]\) converges to \([0, 1]\), which is not a fixed point, and its greatest fixed point is \([0, 1/2]\). The second is that the theorem guarantees the limit and says nothing about the round count. The two-constraint system \(x \le y/2\), \(y \le x/2\) on \([0, 10]^2\) halves the box every round and reaches its fixed point \(\{0\}\) only in the limit, and on a device the round count is the cost. Sofranac, Gleixner and Pokutta implemented the Jacobi schedule on a GPU, one warp or block per constraint, with the candidate bounds merged by atomic minimum and maximum and the rounds iterated to a fixed point with a cap of 100. They report fixed points identical to the sequential code's on 893 of 987 MIPLIB 2017 instances and geometric-mean speedups of 10 to 20 times, up to 180, over single-threaded propagation.B. Sofranac, A. Gleixner and S. Pokutta, "Accelerating domain propagation: an efficient GPU-parallel algorithm over sparse matrices", Parallel Computing 109 (2022), 102874; the same authors' stopping measure is "An algorithm-independent measure of progress for linear constraint propagation", Constraints 27 (2022), 432–455. The fixed-point theorem and the two schedules are Theorem 2.6.11 and Proposition 2.6.12 of Section 2.6, with their sources.
The script below propagates linear rows with both schedules on a chain of twelve constraints, eleven two-variable rows and a bound on the last variable, on the two-cycle above, and on a random sparse system. It counts rounds to a fixed point and checks that the fixed points agree.
# jacobi_fbbt.py: feasibility-based bound tightening with two schedules.
#
# Gauss-Seidel applies the rows in sequence and updates the box at once
# (the sequential solver's loop); Jacobi computes every row's tightening
# from the box at the start of the round and merges at the end (one GPU
# round, Sofranac, Gleixner and Pokutta). Rows are a_i' x <= b_i on a
# box [l, u]; each row tightens its variables from the row's minimum
# activity.
import numpy as np
def tighten_row(ai, bi, l, u):
lo = np.where(ai > 0, ai * l, ai * u)
minact = lo.sum()
nl, nu = l.copy(), u.copy()
for j in np.flatnonzero(ai):
# min activity of the other variables
rest = minact - lo[j]
if ai[j] > 0:
nu[j] = min(nu[j], (bi - rest) / ai[j])
else:
nl[j] = max(nl[j], (bi - rest) / ai[j])
return nl, nu
def propagate(A, b, l, u, schedule, tol=1e-9, cap=500):
"""Run rounds until no bound moves by tol or more (at most cap).
Returns (rounds, l, u); the count includes that final round.
"""
for rnd in range(1, cap + 1):
l0, u0 = l.copy(), u.copy()
if schedule == "gauss-seidel":
for i in range(len(b)):
l, u = tighten_row(A[i], b[i], l, u)
else:
# jacobi: all rows read (l0, u0); merge by max / min
Ls, Us = zip(*(tighten_row(A[i], b[i], l0, u0)
for i in range(len(b))))
l, u = np.max(Ls, axis=0), np.min(Us, axis=0)
if np.abs(l - l0).max() < tol and np.abs(u - u0).max() < tol:
return rnd, l, u
return cap, l, u
def compare(name, A, b, l, u):
"""Both schedules from [l, u]: rounds, agreement, final volume."""
rg, lg, ug = propagate(A, b, l.copy(), u.copy(), "gauss-seidel")
rj, lj, uj = propagate(A, b, l.copy(), u.copy(), "jacobi")
same = np.abs(lg - lj).max() < 1e-7 and np.abs(ug - uj).max() < 1e-7
print(name)
print(f" rounds: Gauss-Seidel {rg:3d}, Jacobi {rj:3d};")
print(f" same fixed point: {same}; "
f"volume {np.prod(uj - lj):.3e} from {np.prod(u - l):.3e}")
n = 12
# (a) a chain x_j <= 0.5 x_{j+1}, x_n <= 1: rows ordered last to first
# (along the chain, good for Gauss-Seidel) and first to last
A = np.zeros((n, n))
b = np.zeros(n)
for j in range(n - 1):
A[j, j], A[j, j + 1] = 1.0, -0.5
A[n - 1, n - 1] = 1.0
b[n - 1] = 1.0
l, u = np.zeros(n), np.full(n, 10.0)
compare("chain of 12, rows ordered last to first", A[::-1], b[::-1], l, u)
print()
compare("chain of 12, rows ordered first to last", A, b, l, u)
print()
# (b) the two-cycle x <= 0.5 y, y <= 0.5 x: converges only in the limit
# (geometric shrinking)
A2 = np.array([[1.0, -0.5], [-0.5, 1.0]])
b2 = np.zeros(2)
compare("two-cycle x <= y/2, y <= x/2 on [0,10]^2 (to 1e-9)", A2, b2,
np.zeros(2), np.full(2, 10.0))
print()
# (c) a random sparse system
rng = np.random.default_rng(5)
m, n = 60, 30
A3 = np.where(rng.random((m, n)) < 0.1, rng.uniform(-1, 1, (m, n)), 0.0)
A3[np.arange(m), rng.integers(0, n, m)] = rng.uniform(0.5, 1, m)
# feasible by construction
x0 = rng.uniform(0, 1, n)
b3 = A3 @ x0 + 0.1 * rng.random(m)
compare("random 60 x 30 system, 10% dense, box [0,10]^30", A3, b3,
np.zeros(n), np.full(n, 10.0))
chain of 12, rows ordered last to first
rounds: Gauss-Seidel 2, Jacobi 13;
same fixed point: True; volume 1.355e-20 from 1.000e+12
chain of 12, rows ordered first to last
rounds: Gauss-Seidel 13, Jacobi 13;
same fixed point: True; volume 1.355e-20 from 1.000e+12
two-cycle x <= y/2, y <= x/2 on [0,10]^2 (to 1e-9)
rounds: Gauss-Seidel 18, Jacobi 34;
same fixed point: True; volume 3.388e-19 from 1.000e+02
random 60 x 30 system, 10% dense, box [0,10]^30
rounds: Gauss-Seidel 15, Jacobi 24;
same fixed point: True; volume 2.549e+01 from 1.000e+30
(Reading the schedules table) The script's cost is one pass over the nonzeros per round and per schedule. In the Jacobi round every row's tightening is independent of the others, which is the parallel part. The Gauss–Seidel round is sequential by construction. The table reads as follows. The fixed points agree in every case, as Theorem 2.6.11 requires. On the chain the sequential schedule needs two rounds when the rows happen to be ordered along the chain and thirteen when they are ordered against it. The Jacobi schedule needs thirteen regardless, because information travels one row per round. On the two-cycle the Jacobi schedule takes about twice the rounds, since Gauss–Seidel halves twice per round. On the random system it takes 24 against 15. Sofranac and coauthors observed about three rounds on average for the sequential schedule on MIPLIB, and more for the parallel one, capped at 100. The Jacobi round costs one pass over the nonzeros for all constraints at once, so on a device the extra rounds are paid for many times over. The open question is a bound on the round count in terms of the constraint graph and the contraction factor of the propagation operator, together with a stopping rule for propagation interleaved with batched bounding on the same device.
Optimality-based bound tightening is the other half. It is \(2n\) linear programs per node that share the whole matrix and differ only in their objectives \(\pm e_j\), the ideal batched family. It is the most expensive component of SCIP's and Couenne's node pipelines. Its engineering, Gleixner, Berthold, Müller and Weltge's filtering, aggressive filtering and Lagrangian variable bounds, was designed for sequential simplex solves.A. M. Gleixner, T. Berthold, B. Müller and S. Weltge, "Three enhancements for optimization-based bound tightening", Journal of Global Optimization 67 (2017), 731–757. Blin et al. (2026) report 25.7 times over sequential Gurobi for MILP bound tightening with a dual tolerance tightened to \(10^{-8}\) and a safety margin \(\Delta = \varepsilon(1 + |c^\top x| + |\phi(r) + \phi(y)|)\) subtracted from each bound before use. The first experiment here is batched OBBT with safe bounds at the root of MINLPLib instances, since each of the \(2n\) bounds is itself a safe bound on a variable by Theorem 7.3.1. Measure box-volume reduction per second against SCIP's OBBT. Then build the batched analogues of the three enhancements, including a rule for which of the \(2n\) problems to admit to a batch. The risk is twofold. OBBT wants the tightest bounds, which is exactly where first-order accuracy is weakest. And a looser variable bound loosens every McCormick plane that uses it (Theorem 2.4.7).
The tree: frontier batching, the pruning loss and when to stop a column
What is known. A batch of \(b\) nodes is committed before the incumbents it will produce are known, so it bounds nodes a serial best-first search would have pruned. The inflation consists of the nodes with bounds in the window between successive incumbents, and of the non-essential nodes a batch reaches when the essential frontier holds fewer than \(b\) (Proposition 6.4.2). It is negligible for small batches and catastrophic for large ones: 1.00 for \(b \le 64\), 1.07 at \(b = 1{,}024\) and 7.9 at \(b = 65{,}536\) on the 132,376-node knapsack tree of Section 6.4. The batch figure's device model has its best wall time at \(b = 128\), with a speedup of 9.30 over sequential against 6.01 at \(b = 32\) and 7.97 at \(b = 256\), where the time is 112.0 units against 893 sequential and 128 of 1,359 nodes are bounded for nothing. The only published batching policy is B³-PWL's: fill the batch to 85 percent of device memory and expand breadth-first when the queue falls below 90 percent of the batch size. Amdahl's bound (Proposition 6.4.4) caps the whole-search speedup at \(1/s\) for a sequential fraction \(s\), which is why every working GPU tree keeps the node data on the device.Guan, Luo, Li, Chen and Xu (2026), cited in full in Section 7.4, for the policy and for the remark that the method does not warm start between parent and child because PDHG "lacks a basis representation"; the batch figure and the inflation numbers are Section 6.4's; the anomaly theorems that make inflation a matter of order rather than of arithmetic are Lai and Sahni (1984), Section 6.2.
The first experiment. Extend the batch simulation of Section 6.4 to spatial branching on the bilinear example and on Haverly's pooling problem (Section 3.5). Add a GPU heuristic running between rounds and an adaptive batch size that tracks the frontier width. Measure inflation against batch size and heuristic strength. Then run the column policy of problem 1 on the same trees, recording the iteration budget spent on columns that end up pruned against columns that end up branched. A model of inflation in terms of the incumbent's lag in rounds, in the style of Proposition 7.4.2's growth model, is what would make the results transfer.
The risk. The results are instance-specific. The best batch size moves during the run, and a fixed one is wrong at both ends of the search. The ramp-up and ramp-down of Section 6.2 reappear as rounds in which the frontier cannot fill the device, so the speedup of a GPU tree is bounded by the tree's width over time, which nobody knows in advance.
The primal side: nonlinear local search and multistart on the device
What is known. Section 7.7 gave it. The GPU heuristics that win competitions are linear (Feasibility Jump, feasibility pump, tabu search on linear violations), and MINLP incumbents still come from CPU local NLP solves. The lockstep multistart of Algorithm 7.7.4 is the natural GPU form, and the Bayesian stopping rule of Proposition 1.3.12 tells it when to stop. Batched small NLP solves exist (ExaTron), and batched QP solves exist at scale (WarpMPC, OptNet). The open items of Feasibility Jump are parallel move selection under conflicts, weight updates as a true dual ascent, and nonlinear constraints. For nonlinear constraints the one-dimensional violation function \(G_j(t)\) is no longer piecewise linear, but it can be evaluated on a grid of candidate values in parallel.
The first experiment. Implement Feasibility Jump moves with nonlinear violations evaluated by the DAG kernels of problem 2, on the pooling problem and on the three-asset cardinality instance of Section 4.4 (expected returns \(0.10, 0.07, 0.04\), volatilities \(0.20, 0.15, 0.10\), correlation \(0.3\)). Measure the feasible-solution rate against SCIP's heuristics in ten seconds. In parallel, run Algorithm 7.7.4 with the integers fixed from a pool of candidate assignments and the continuous local solves batched through the condensed system of Section 7.5, and measure the primal integral against BARON's multistart preprocessing.
The risk. Continuous variables with nonlinear coupling make single-variable jumps ineffective, so a line-search or coordinate-descent move set may be needed. A local NLP solve at \(10^{-4}\) must be repaired to feasibility before its value is an incumbent (Section 7.5). And lane efficiency falls with the spread of iteration counts across the batch (Proposition 6.4.6), so the fixed budget must be chosen against the distribution of local-solve lengths, which is problem-dependent.
Certified bounds from GPU nonlinear relaxations
What is known. For a convex MINLP node, or for a conic relaxation of a nonconvex one, the node relaxation is a convex NLP. Its GPU solve returns a point and multipliers at \(10^{-4}\) to \(10^{-6}\) (Section 7.5), which is not a bound. For linear and quadratic relaxations the correction is exact arithmetic on the dual iterate (Theorem 7.3.1, Proposition 7.6.6). For a general convex relaxation with constraints \(g(x) \le 0\) the analogue is the Lagrangian dual value \(q(\lambda) = \min_{x \in B_N} [f(x) + \lambda^\top g(x)]\) at the returned \(\lambda \ge 0\). It is a valid bound for any \(\lambda\) by weak duality (Section 2.2), provided the inner minimization over the node's box is itself global. For convex \(f\) and \(g\) that inner problem is convex, and in general interval arithmetic bounds it over the box.
The first experiment. For the convex MINLPLib instances, take the multipliers of a condensed-space solve at \(10^{-4}\). Compute an interval lower bound on \(f + \lambda^\top g\) over the node box on the device with the kernels of problem 2. Compare it with the exact node value and with the value the solver reported. Then do the same with one or two subgradient steps on \(q\), each costing one more inner bound.
The risk. Interval bounds on the Lagrangian are loose on large boxes, and the partition numbers of problem 2 show how loose. The correction may therefore be far looser than the \(10^{-4}\) it repairs. Tightening it means partitioning, whose cost is exponential in the number of variables that enter the nonlinear terms.
Learned branching on the device
What is known. The learned-branching literature of Section 5.5 trains a network on the GPU and runs the tree on the CPU. The one line of work that moved part of the search itself to the device is Nair and coauthors' batched alternating-direction strong-branching expert. It produced 1.4 and 12 times more training data in the same time on their two largest sets. The 2026 evidence points away from per-node inference on the device. Sparse models with under 4 percent of a graph network's parameters are faster than both the default solver and the GPU network. Evolved CPU-only branching rules outperform SCIP's and the GPU policies. What shipped for nonconvex MINLP is learned rule selection in Xpress Global, an 8 to 9 percent geometric-mean time reduction and over 10 percent on hard instances.V. Nair et al., "Solving mixed integer programs using neural networks", arXiv 2012.13349 (2020); S. Bayramoğlu, G. L. Nemhauser and N. V. Sahinidis, "Speeding up mixed-integer programming solvers with sparse learning for branching", arXiv 2604.00094 (2026); C. Zhang et al., "BiFE: search-efficient discovery of CPU-only branching policies via LLM-based bi-fidelity evolution", arXiv 2609.36735 (2026); T. Berthold and F. Geis, "Learning to choose branching rules for nonconvex MINLPs", arXiv 2602.09996 (2026). All preprints; authors' numbers.
The first experiment. Compare three ways of choosing the branching variable on the same device budget and the same MINLPLib trees. The first is batched-PDHG strong branching with exact scores from approximate LPs (Algorithm 7.4.5 over all candidates, cuOpt's batch option). The second is pseudocosts, the per-variable averages of past bound improvements that Section 3.1 defines, and the third is a graph network. Measure tree size and time. The learned component worth having on the device is a batched one, scores for thousands of candidates at once. The question is whether learning adds anything once exact strong-branching scores are cheap.
The risk. The exact-but-approximate scores may mislead branching as badly as a network does off its training distribution. And the host–device transfer of the frontier's features per round is the sequential fraction of Proposition 6.4.4 in another form.
Reproducibility: deterministic reductions and batch schedules
What is known. Section 6.2 established the three notions and the two device-specific sources of nondeterminism. A floating-point reduction whose order is decided at run time gives different bits on different runs. Atomics into one word, dynamic scheduling, and library kernels chosen by the number of multiprocessors all decide the order at run time. In Section 6.2's experiment the recursive sum of one 1,000-term vector took 273 different values over a thousand arrival orders, and a fixed tree whose leaves follow the arrival order still took 45. A fixed-schedule reduction over operands in a canonical order is bitwise reproducible (Proposition 6.2.18), as are integer and min/max atomics. A one-bin integer accumulation is order-free (Proposition 6.2.19). Demmel and Nguyen's reproducible summation (Sections 6.2 and 7.3) costs about nine floating-point and three bitwise operations per term for a six-word accumulator. One ulp of relative noise in the bounds moves a 127,505-node knapsack tree to between 128,085 and 128,156 evaluations, a different count in every one of twelve runs, while every run returns the same optimum. Validity under any order comes from directed rounding (Proposition 6.2.21), and reproducibility from a fixed batch schedule (Proposition 6.4.5). The price was measured on the CPU at 2 to 9 percent of time for Xpress Global and a third of the threads idle for Para-B&B. It has not been measured on any GPU.Belotti, Berthold, Gally, Gottwald and Pólik (2025), Table 8, and Zhang et al. (2026, Para-B&B), both cited in full in Section 6.2; cuOpt's deterministic parallel branch and bound is experimental since release 26.02 and excludes the GPU heuristics (release notes).
The first experiment. The question, stated once, is the cost, in wall-clock time and in tree size, of a fully deterministic GPU spatial branch and bound: fixed-schedule reductions, directed rounding, a fixed batch schedule, incumbents applied at round boundaries. The experiment measures whether that cost falls with batch size, as the CPU cost falls with thread count, or rises with it through idle lanes and incumbent lag. Run the deterministic frontier batch of Section 6.4 against its opportunistic variant on the same instances, ten runs each, and report time, idle lanes, node counts and whether the paths coincide. Write the products \(Ax\) and \(A^\top y\) of Algorithm 7.4.5 as gathers with fixed in-warp trees rather than scatters with atomics, and measure the throughput lost.
Problem 8: three ways to sum the terms of one entry (Ax)_i
recursive sum in a fixed tree, leaves a fixed in-warp tree,
arrival order: atomic in arrival order canonical order
adds into one word
(a scatter) (a gather)
t_3 --+ t_3 --+ t_1 --+
t_1 --+ +--+ +--+
t_4 --+--> (Ax)_i t_1 --+ | t_2 --+ |
t_2 --+ +--> (Ax)_i +--> (Ax)_i
t_4 --+ | t_3 --+ |
+--+ +--+
t_2 --+ t_4 --+
in Section 6.2, one 1,000-term vector, a thousand arrival orders:
273 different values 45 different values bitwise reproducible
(Proposition 6.2.18)
t_j = a_ij x_j, the terms of (Ax)_i; arrival order: decided at
run time
The risk. It is small, with two reservations. A deterministic incumbent exchange across heuristic threads needs a synchronous pool, which reintroduces the lag of problem 4. And a logical clock on a device must be a count of rounds and iterations fixed in advance, since the host sees kernel completions and not work, so a work limit in those units is not comparable across devices.
Mixed precision and the arithmetic of the bound
What is known. Three precisions appear in this problem. FP64 and FP32 are the 64-bit and 32-bit IEEE formats. FP16 is the 16-bit floating-point format, and TF32 is a 19-bit format with FP32's exponent range and 10 bits of mantissa. Tensor cores, the matrix-multiply units of recent NVIDIA devices, run at their highest rates in FP16 and TF32. cuOpt 26.04 ships single-precision and mixed-precision PDLP. Kempke and Koch showed that \(10^{-4}\) first-order solutions lose nothing for heuristics. CuClarabel's mixed-precision factorization gives "moderate speed improvements". An interior-point method in single precision exists for differentiable optimization. No paper using tensor cores for linear programming or interval bounding was found.NVIDIA cuOpt release notes, 26.04 ("Add support for FP32 and mixed precision in PDLP", 9 April 2026); Kempke and Koch (2026); Chen, Tse, Nobel, Goulart and Boyd (2026); J. Arrizabalaga, K. Tracy and Z. Manchester, "A differentiable interior-point method in single precision", arXiv 2605.17913 (2026); the negative statement is from the arXiv sweep behind this section. The theorem that makes any of this safe for a tree is Corollary 7.3.3. The bound is computed in double precision with directed rounding from whatever dual the single-precision iteration produced, and its looseness grows with the dual error.
The first experiment. Evaluate Theorem 7.3.1's bound in single precision with the _rd intrinsics on the frontier of Section 7.4 and measure the loss against double. Then run the PDHG iterations in TF32 or FP16 on tensor cores with an FP64 bound evaluation, and measure iterations to a given safe-bound tightness against FP64 iterations.
The risk. Reduced precision may slow PDHG's convergence more than it saves per iteration, since the method's progress is measured in a norm the rounding perturbs. And the safe bound in FP32 is valid but may be too loose to prune anything.
Decomposition: Lagrangian bounds from batched block solves
What is known. The theory is Section 3.7: Theorem 3.7.3 makes every dual value a bound, Theorem 3.7.11 turns block values into cuts, Theorem 3.7.6 is the decomposition bound, and the two-site example of Section 3.7 shows the ladder \(-0.3553 < -0.2540 < -0.2150 < -0.1600\). The one published GPU instance of massive Lagrangian decomposition is ExaTron (Section 7.5). It decomposes optimal power flow by grid components and solves the many small nonconvex subproblems in batch with no host–device transfer. It reports linear scaling in the batch size and more than 35 times speedup on six GPUs against forty CPU cores. Its subproblems are solved locally, so the scheme is a heuristic rather than a global method, and what carries over is the batching pattern, not the guarantee.Kim, Pacaud, Schanen, Kim and Anitescu (2025), abstract; Y. Kim and K. Kim, "Accelerated computation and tracking of AC optimal power flow solutions using GPUs", ICPP Workshops 2022, for the component-wise decomposition with ADMM on a 70,000-bus system. The dual iteration has three parts. The block solves form the first. A reduction forming \(q(\lambda)\) and its supergradient \(\sum_i A_i x_i(\lambda) - b\) forms the second. A multiplier update forms the third, a vector step for the subgradient method or a small quadratic program for a bundle method. The first is a batch of identical small problems, the second a reduction, the third a synchronization point of negligible cost. Every dual value is a valid bound by weak duality (Section 2.2). The multiplier update may therefore be stale, asynchronous or in single precision without invalidating anything, provided the block solves are global and the sum of block values is rounded toward \(-\infty\). The block solves are the problem. A Lagrangian bound needs each \(q_i(\lambda)\) to be a global minimum or a certified lower bound, and no batched global solver for small nonconvex blocks exists.
Problem 10: one dual iteration, in its three parts
+--> lambda
| |
| | 1. block solves: a batch of identical small problems
| +--------------+--------------+-- . . . --+
| v v v v
| q_1(lambda) q_2(lambda) q_3(lambda) q_n(lambda)
| | | | |
| +--------------+--------------+-- . . . --+
| | 2. a reduction
| v
| q(lambda) = sum_i q_i(lambda) - lambda^T b, and the
| supergradient sum_i A_i x_i(lambda) - b
| |
| | 3. the multiplier update, a synchronization point
| v
| a vector step (subgradient method) or a small QP (bundle)
| |
+-------+
each q(lambda) is a valid bound by weak duality if every
q_i(lambda) is a global minimum or a certified lower bound and the
sum is rounded toward -inf; the update may be stale, asynchronous
or in single precision
The first experiment. Scale the two-site example of Section 3.7 to thousands of sites with random costs, capacities and crew requirements, each block the running example with a capacity switch and a closed-form value against which a batched bounder can be checked. Evaluate \(q(\lambda)\) on the device by a batched enumeration kernel over the blocks' integer choices with the continuous part in closed form, and maximize it by a subgradient method with the Polyak step, the step length that uses the best known value of the dual optimum to set how far to move. Then replace the enumeration by the batched interval or McCormick bounder of problem 2 to obtain certified block lower bounds. Measure the dual bound's tightness against the sum of exact block values, and the time per dual iteration against a CPU loop over blocks. Then take Karuppiah and Grossmann's water network, if its data can be obtained, with the scenario subproblems bounded in batch.
The risk. A block solved only to a local minimum gives an invalid bound. Batched global block solves need per-lane spatial branching inside the batch, with the lane-efficiency problem of Section 6.4 inside each block. And the dual ascent for different nodes of a tree proceeds at different rates, so the host must schedule which nodes get another dual iteration, which is problem 4 again.
Which structures benefit
The evidence of Sections 6 and 7 sorts problem structures by how much of their work is regular and shared. The table collects it with its gaps. The first row's evidence is the PDLP paper's journal version, which reports eight of eleven instances with \(10^8\) to \(6 \times 10^9\) nonzeros solved to feasibility \(10^{-8}\) and gaps near one percent within six days, where a barrier method solved three.D. Applegate, M. Díaz, O. Hinder, H. Lu, M. Lubin, B. O'Donoghue and W. Schudy, "PDLP: a practical first-order method for large-scale linear programming", Mathematical Programming Computation (2026), doi 10.1007/s12532-026-00309-2; arXiv 2501.07018 (2025). The other rows cite their sources where they were first quoted: Blin et al. (2026) and cuOpt's release notes in Section 7.4; Moehle et al. (2023) in Section 7.6; Kim et al. (2025), Zhang et al. (2025), Gottlieb, Xu and Stuber (2026), Guan et al. (2026) and Pacaud et al. (2026) in Sections 7.5 and 7.8; Liu and Lodi (arXiv 2605.22188, 2026) and Lucas, Meng and Mazumder (arXiv 2602.04551, 2026) in Section 6.4.
| structure | what batches or vectorizes | evidence (October 2026) | what is missing |
|---|---|---|---|
| one huge sparse LP (\(10^8\) to \(10^9\) nonzeros) | the two products of PDHG | PDLP journal version: 8 of 11 instances with \(10^8\) to \(6 \times 10^9\) nonzeros; Mittelmann's INFORMS talk of 28 Oct 2025: among the GPU codes, cuPDLP-C alone finishing all six of Hinder's instances | nothing for MINLP: such LPs are not node relaxations |
| hundreds of sibling LPs sharing \(A\) (strong branching, OBBT, a frontier) | one matrix read, \(K\) right-hand sides | Blin et al. (authors' numbers): 12× to 489× over Gurobi's dual simplex for strong branching at the root on four MILP families; cuOpt's batch PDLP in strong and reliability branching, off by default | McCormick rows that change with the box (problem 1) |
| sibling QPs sharing \(Q\) and \(A\) (convex MIQP, direction-fixed tax QPs) | one factorization, \(K\) right-hand sides | Section 7.6's batch: 230 of 230 pruned at 50 iterations; OSQP, WarpMPC, OptNet | a tree around the batch with its certificate in dollars (Section 9) |
| factor-model Hessians (\(n\) assets, \(r \ll n\)) | \(O(nr)\) solves and products by Woodbury | Proposition 4.4.7; Moehle et al.: a mean of 251 ms per nonconvex instance on a CPU | batched Schur solves per account on a device (unmeasured) |
| separable nonconvexity (tax kinks, lots, fixed charges) under few coupling rows | closed-form proximal steps, one thread per piece; a duality gap at most the sum of the largest block nonconvexities, one per coupling row that can be active, however many blocks there are, and finite only when every block's domain is an interval (Theorems 5.4.7–8) | the ADMM heuristic and envelope bound of Sections 5.4 and 9 | nothing beyond the CPU heuristic |
| many small independent problems (accounts, scenarios, blocks) | one problem per lane or block | ExaTron: more than 35× on 6 GPUs against 40 CPU cores, with local block solves (a heuristic); batched MPC solvers | batched global solves (problem 10) |
| factorable DAGs on a frontier of boxes | the same straight-line program per box | Zhang et al. (authors' numbers): an interval bounder with wall-clock times three orders of magnitude below CPU interval arithmetic without partitioning; Gottlieb, Xu and Stuber (authors' numbers): pointwise McCormick kernels at about 9 ns per evaluation against 237 on the CPU | a complete operation library with directed rounding (problem 2) |
| piecewise-linear models with many SOS2 sets | fixed rows, box-dependent bounds | B³-PWL (Guan et al., preprint, authors' numbers): 9.25× over cuOpt in geometric mean on 43 synthetic instances, 10.3× over SCIP on valve-point unit-commitment instances; Gurobi about 7× faster than B³-PWL on the synthetic set | general MINLP has no fixed rows |
| sparse regression, \(k\)-sparse GLMs | fixed matrix across all nodes | Liu and Lodi: one to two orders of magnitude over a CPU implementation; Lucas, Meng and Mazumder: runtime improvements over a CPU implementation, a specialized B&B and commercial MIP solvers, no factor in the abstract | transfer to problems without a fixed matrix |
| dense constraint rows | a dense condensed matrix | Pacaud et al.: "diminished robustness and limited speedups" on CUTEst edge cases | a sparse-dense split of the condensed system |
| small warm-startable node LPs | nothing: a few dual simplex pivots | Blin et al. on MIPLIB with 200-pivot warm starts; PDHG "lacks a basis representation" (B³-PWL) | the regime stays on the CPU |
(Reading the table for the tax problem and for MINLPLib) The reading for the tax problem of Section 9 is rows four to six together. A factor model supplies cheap solves and a diagonal to extract. The nonconvexity is separable under \(k + 1\) coupling rows, so the gap is small and the proximal steps are closed-form. Accounts are independent, so the whole of a manager's book is one batch of batches. The reading for general MINLPLib is the negative rows: no fixed matrix, no fixed rows, and dense rows somewhere in most instances. The first experiment on structure is a catalogue. For the 200 instances of Mittelmann's MINLP benchmark, measure three numbers per instance: the fraction of McCormick rows whose coefficients change at a branching, the density of the densest row, and the number of variables that enter nonlinear terms. Sort the instances by the three numbers. The catalogue predicts which of problems 1, 2 and 10 applies to which instance, and it does not exist.
Measurement: the benchmark and the scaling study that do not exist
Section 5.2 gave the methodology of the field's benchmarks and the dates and machines behind every number in the monograph. For the GPU programme two measurements are missing: a GPU code scored on MINLPLib, and a scaling curve of a global MINLP solver across devices. This item defines the measure the two share and fixes the protocol for making them.
The primal integral of Section 2.3 scores the incumbent and says nothing about the proof. The measure for the proof is its mirror image, with the proved bound in place of the incumbent.
Definition 7.8.1 (dual gap function; dual integral and its time average). Let \(z^\star\) be the optimal value of a minimization problem, or the reference value of the benchmark when the optimum is not known, and let \(z_{\mathrm{dual}}(t)\) be the global lower bound a run has proved by time \(t\), with \(z_{\mathrm{dual}}(t) = -\infty\) while no finite bound has been proved. The dual gap function of the run is
\[d(t) \;=\; \begin{cases} 0 & \text{if } z_{\mathrm{dual}}(t) = z^\star, \\ 1 & \text{if } z_{\mathrm{dual}}(t) = -\infty \text{ or } z_{\mathrm{dual}}(t)\, z^\star < 0, \\ \dfrac{|z^\star - z_{\mathrm{dual}}(t)|}{\max\{ |z^\star|, |z_{\mathrm{dual}}(t)| \}} & \text{otherwise}, \end{cases}\]a step function with values in \([0, 1]\) that equals \(1\) before the first finite bound and falls at each improvement of the bound. For a run of length \(T\) the dual integral is \(D(T) = \int_0^T d(t)\,dt\), in units of time, and the time-averaged dual integral is \(\bar D(T) = D(T)/T \in [0, 1]\).
(What the definition changes, and what it keeps apart) The definition is Definitions 2.3.3 and 2.3.4 with the incumbent replaced by the proved bound, and the same conventions apply: the integrand is \(1\) while nothing is known, and steps, nodes or relaxations may replace time. Because \(d\) takes values in \([0, 1]\), \(0 \le \bar D(T) \le 1\), and \(\bar D(T) = 1\) when the run proved no finite bound, which is the value every GPU heuristic of Section 7.7 scores. The primal–dual integral that the SCIP developers use (Section 2.3) integrates the gap between the incumbent and the bound and so mixes the two halves of a solve. The pair \((\bar P(T), \bar D(T))\) keeps them apart. A heuristic that never proves anything can have a small \(\bar P(T)\), as the example of Section 7.7 showed, and it has \(\bar D(T) = 1\).
The protocol fixes everything else a number of this kind depends on. Its eleven steps fix the instances, the reference values, the tolerances, the limits, the measures, the aggregates, the hardware, the scaling arm, the determinism arm, the codes and the reporting. Every convention in it is one that Section 5.2 or Section 6.2 has already stated for CPU solvers.
Algorithm 7.8.2 Scaling and determinism measurement of a global MINLP
solver, on CPUs and on devices
Input the codes to be measured, each with the modes it offers
(deterministic, opportunistic); the benchmark instances with their
reference values; one feasibility tolerance, one relative gap
tolerance and one time limit for all
Output per code and mode: a primal column P-bar(T), a dual column D-bar(T),
a solved count, a scaling curve and a determinism column, each
carrying the conditions it was measured under
1. instances: the 200 instances of Mittelmann's MINLP benchmark, by
name; no instance is dropped after the fact
2. reference values: the MINLPLib primal and dual reference bounds on the
day of the run, recorded with the run; the reference
primal value is z* in Definitions 2.3.3 and 7.8.1
3. tolerances: one feasibility tolerance and one relative gap
tolerance for every code, with the gap formula stated
4. limits: one wall-clock time limit T for every code, on the CPU
and on the device alike; a run that fails or reaches
the limit is charged the limit (Section 5.2)
5. measures: for every run, P-bar(T) of Definition 2.3.4 and
D-bar(T) of Definition 7.8.1, both with the
integrand 1 before the first incumbent or the first
finite bound, and whether the run proved optimality to
the gap tolerance of step 3
6. aggregates: the shifted geometric mean with shift 10 s
(Definition 5.2.1), scaled by the best mean in the
same column; the per-instance files published; no
virtual-best row unless it was computed from them
7. hardware: the machine, its memory and its thread count; the
device, its memory, its driver and library versions;
every solver version; the date of the run
8. scaling arm: every code at 1, 2, 4, 8, 16 and 32 threads, and a
device code at 1, 2, 4, ... devices as well; three
permutations of each instance's rows and columns at
every size (the permutation rule of Section 5.2, with
the count Xpress Global used, Section 6.3); the
speedup of Definition 6.1.1 reported over all solvable
instances and over the instances that need at least
100 nodes
9. determinism arm: deterministic and opportunistic modes as separate
rows; every configuration run twice; the agreement of
node counts and of incumbent sequences between the two
runs reported as a column (Definition 6.2.14)
10. codes: the CPU global solvers of Mittelmann's page
(BARON, SCIP, LINDO, SHOT) at default settings, MAiNGO
with its GPU bounder of problem 2, and every device
code with the mode it ran in; a code that produces no
bound stays in the dual column with D-bar(T) = 1
11. reporting: every number carries the items of steps 1 to 4 and 7
and the per-instance files; the dual column is
reported as it comes out and is never filled with the
primal column's numbers
Invariant
every code in a table faces the same instances, tolerances and limit, so
any two columns are comparable; a number without its list of conditions
is not reported.
The protocol's cost is the number of runs, which for 200 instances, six thread counts, three permutations, two modes and two repetitions is 14,400 runs per code at the time limit of step 4, before any device code is added. The runs are independent of one another, so the measurement is itself a batch and spreads over as many machines as are available. Nothing in the protocol is specific to the GPU programme, and only what it would show for the programme is added here. Its dual track will at first contain only the CPU solvers and MAiNGO, because every GPU heuristic of Section 7.7 produces no bound. That column is to be reported as it stands, with \(\bar D(T) = 1\) for those codes, and not filled with primal numbers.
The runs of Algorithm 7.8.2 for one CPU code (steps 1, 8 and 9)
each of the 200 instances 200
|
+-- at 1, 2, 4, 8, 16 and 32 threads x 6
|
+-- three permutations of rows and columns x 3
|
+-- deterministic and opportunistic mode x 2
|
+-- every configuration run twice x 2
--------
14,400
every run at the time limit T of step 4, independent of the
others: the measurement is itself a batch
The nearest existing evidence is the baseline the protocol would extend. Xpress Global's deterministic parallel tree on its own 1,308-instance set reaches speedups of 1.34, 1.72, 2.19, 2.59 and 2.59 at 2, 4, 8, 16 and 32 threads over all solvable instances. On the instances needing at least 100 nodes the speedups are 1.54, 2.21, 3.13, 3.97 and 4.05, and the solved count falls slightly from 16 to 32 threads. The multi-GPU Chapel codes reach 44 percent efficiency at 1,024 GPUs on N-Queens trees whose bounds cost nanoseconds. Their 20 × 20 flow-shop instances stop scaling beyond 16 nodes, that is 128 GPUs, because the trees are too small. FiberSCIP's MINLP experiments are the one shared-memory MINLP measurement.Belotti, Berthold, Gally, Gottwald and Pólik (2025), Section 3 and Table 7; G. Helbecque, E. Krishnasamy, T. Carneiro, N. Melab and P. Bouvry, "Portable PGAS-based GPU-accelerated branch-and-bound algorithms at scale", Concurrency and Computation: Practice and Experience 37 (2025), e70321 (speedups of 4.02, 8.00, 15.81, 31.12 and 56.30 relative to one 8-GPU node at 8, 16, 32, 64 and 128 nodes on N-Queens, hence 44 percent at 128 nodes); Y. Shinano, S. Heinz, S. Vigerske and M. Winkler, "FiberSCIP — a shared memory parallelization of SCIP", INFORMS Journal on Computing 30 (2018), 11–30. The CPU reference numbers for global MINLP are Mittelmann's page of 26 February 2026: 200 instances, two hours, AMD Ryzen 9 5900X; BARON 158, SCIP 154, LINDO 116 and SHOT 96 solved, and 183 of the 200 (91.5 percent) by at least one of the four, the row "optimal auto settings" of his per-instance table compare.txt, read 5 October 2026. Two caveats travel with the 183: the table's LINDO column is from a run of 6 March 2026 with 119 solved, not the page's 116, so the union mixes two LINDO runs, and excluding the two rows flagged "#" gives 181.
Accounting for accuracy from end to end
The last item is the one every other depends on. A GPU branch and bound as sketched here mixes five kinds of arithmetic. They are an FP32 or mixed-precision first-order iterate, an FP64 safe bound with directed rounding, a relaxed-equality NLP solve at \(10^{-4}\) to \(10^{-6}\), interval bounds with outward rounding, and an exact integer incumbent checked to a feasibility tolerance. Section 8 will give the three tolerances a solver works to. What does not exist is a statement of what the final "gap \(\le \varepsilon\)" of such a solver certifies, and of how the budget \(\varepsilon\) should be split between relaxation error, floating-point error and the gap itself. The exact-MIP literature has the vocabulary: safe bounds, project-and-shift, which turns an approximate dual solution into an exactly feasible one by a projection and a shift toward a known interior point, verified cuts, and a certificate format that an independent checker can replay, all for the linear case.W. Cook, T. Koch, D. E. Steffy and K. Wolter, "A hybrid branch-and-bound approach for exact rational mixed-integer programming", Mathematical Programming Computation 5 (2013), 305–344; L. Eifler and A. Gleixner, "A computational status update for exact rational mixed integer programming", Mathematical Programming 197 (2023), 793–812; K. K. H. Cheung, A. Gleixner and D. E. Steffy, "Verifying integer programming results", IPCO 2017, LNCS 10328, for the certificate format. Section 5.6 gives the numbers. Its transfer to inexact-but-fast GPU bounds has not been written down: a proof that the final gap claim is rigorous, and a certificate a CPU can check after the fact. Rational arithmetic is a poor fit for a device, so the candidates are interval and affine arithmetic in fixed precision, the latter tracking each quantity as an affine form in shared noise symbols so that correlated errors partly cancel. The first experiment is a certificate format for a spatial branch-and-bound tree whose node bounds are safe bounds from first-order duals and interval evaluations, with a sequential checker. The risk is that the checker costs as much as the search.
Accounting for accuracy: five kinds of arithmetic, one gap claim
FP32 or mixed-precision first-order iterate ----+
FP64 safe bound, directed rounding -------------+
relaxed-equality NLP solve, 1e-4 to 1e-6 -------+--> "gap <= eps":
interval bounds, outward rounding --------------+ what does it
exact integer incumbent, feasibility tolerance -+ certify?
eps [ relaxation error | floating-point error | the gap itself ]
how to split the budget: not written down
the first experiment:
the tree --> a certificate --> a sequential checker on a CPU
(node bounds: safe bounds from first-order duals and interval
evaluations)
Where this is used
In October 2026 the device runs, in production, the first-order LP relaxation of four solvers (Gurobi 13, COPT 8, HiGHS and cuOpt), the barrier factorizations of two (COPT and cuOpt), the primal heuristics of one (cuOpt), and no tree. The global MINLP solvers of Section 5.3 use none of it, and the two GPU bounding engines of problem 2 are research prototypes inside MAiNGO and ParBB.
What parallelizes
What parallelizes is the batches: the sibling relaxations of problem 1, the box evaluations of problem 2, the propagation rounds and the \(2n\) bound-tightening problems of problem 3, the local solves of problem 5 and the block solves of problem 10. What does not parallelize is the tree, whose decisions depend on one another, and the hard nodes, whose bounds sit within the gap and need near-exact solves. The agenda above is, in one sentence, to build the batches, to put the safe bound between them and the tree, and to measure the result on MINLPLib with the dual integral reported as it comes out. Section 8 lists what remains of the subject as a whole, and Section 9 applies the programme to the tax problem.
What parallelizes, what does not, and the safe bound between them
what parallelizes: the batches what does not
+----------------------------------+ +--------------------+
| sibling relaxations (1) | | the tree |
| box evaluations (2) | | the hard nodes |
| propagation rounds and the | +--------------------+
| 2n bound-tightening problems (3) | ^
| local solves (5) | |
| block solves (10) | |
+----------------------------------+ |
| |
+--------> the safe bound --------+
(k) problem k