Mathematical Optimization

NN-Generated Lyapunov Metrics for Fast NMPC: What Is Actually Guaranteed?

A critical note on learned Lyapunov terminal costs for NMPC, focusing on Cholesky-structured positive-definite surrogates, horizon compression, and the unresolved role of approximation error.

Positioning: why this problem matters in real systems

Nonlinear model predictive control is attractive because it can handle nonlinear dynamics, constraints, changing references, and multi-step tradeoffs in one optimization problem. That is why NMPC appears naturally in chemical process control, robotic motion, autonomous parking, energy storage operation, hydrogen supply chains, and other systems where the controller must reason ahead while respecting physical limits.

The difficulty is timing. At every sampling instant, NMPC solves a nonlinear optimization problem over a finite horizon. A longer horizon usually gives better foresight: the controller can avoid myopic actions, anticipate delayed constraint effects, and approximate stabilizing behavior more reliably. But a long horizon also increases the online nonlinear programming burden.

A short-horizon NMPC controller is faster, but it can be shortsighted. It may choose an action that looks good over one or two steps but makes the later problem difficult, expensive, or infeasible. The central question of this note is therefore practical: can a learned terminal or continuation cost preserve some of the value of a long horizon while allowing a much shorter online horizon?

Problem setting

The control problem is a parameterized NMPC problem. The state is x, the control is optimized online, and p denotes a parameter vector such as a reference, operating condition, exogenous signal, or steady-state descriptor. The desired state or steady-state target is written as xs.

Long-horizon NMPC implicitly computes a continuation value: the cost-to-go after the first action. If that continuation value were known exactly, a one-step online problem plus the exact continuation value could reproduce the first decision of the long-horizon problem under the same modeling assumptions. This is the Bellman-style horizon-compression intuition behind the method.

The supplied material describes an online horizon of N=1. The missing information is replaced by a learned terminal/continuation surrogate. This reduces the online decision dimension, but it does not remove nonlinear dynamics, nonconvexity, local optima, numerical conditioning, or feasibility issues. It compresses the horizon; it does not turn NMPC into a trivial lookup table.

Prior research gap

Classical stabilizing NMPC usually relies on three terminal ingredients:

  • a terminal set Xf,
  • a terminal cost F,
  • and a local stabilizing controller κf.

These ingredients give a route to recursive feasibility and Lyapunov decrease, but they can be hard to design for strongly nonlinear systems, changing operating points, and high-dimensional parameterized tasks. Long-horizon terminal-constraint-free NMPC can sometimes recover stabilizing behavior by making the finite horizon sufficiently long, but that moves the burden into online computation.

Explicit MPC and neural-network approximate MPC reduce online cost by approximating either the policy or the optimization map. The tradeoff is that approximation often weakens the clean stability and feasibility story. A generic learned terminal cost can approximate a value function, but an arbitrary scalar neural network does not naturally guarantee positive definiteness, nor does it automatically behave like a Lyapunov function.

This is the gap the method is trying to occupy: learn a continuation surrogate that is computationally useful online, but impose enough structure that the terminal cost has at least a Lyapunov-compatible metric form.

Core idea

The core idea is data-driven horizon compression. Offline, the method uses long-horizon NMPC solutions to supervise a terminal/continuation surrogate. Online, the controller solves a much shorter NMPC problem, potentially with N=1, and relies on the learned terminal term to represent the omitted future.

If the exact continuation value V*(x,p) were available, the one-step decomposition would be conceptually clean:

min u ℓ(x,u,p) + V* (x+,p+)

Here, ℓ is the one-step stage cost, x+ is the next state produced by the nonlinear dynamics, and p+ is the updated parameter. The problem is that V* is generally unknown. The paper therefore learns a surrogate, but with a special structure: the neural network generates a positive-definite terminal matrix rather than directly outputting an arbitrary scalar cost.

Mathematical structure: key architecture

The distinctive architecture is this: a feedforward neural network receives p and outputs the lower-triangular entries of a matrix Lθ(p). The terminal matrix is then constructed as

Pθ (p) = Lθ (p) Lθ(p) T + εI.

The term Lθ(p) is the lower-triangular factor generated by the neural network. The matrix Pθ(p) is the learned terminal metric. The scalar ε>0 is a fixed positive regularization constant, and I is the identity matrix.

The learned terminal/continuation surrogate is

V^θ (x,p) = (x-xs) T Pθ (p) (x-xs).

This equation says that the terminal cost is quadratic in the state error x-xs, but the quadratic metric changes with the parameter p. The neural network is therefore not simply “the value function.” It is a generator of a Cholesky-type Lyapunov metric.

The information flow is compact:

parameter p
   |
   v
Feedforward NN f_theta(p)
   |
   v
lower-triangular L_theta(p)
   |
   v
P_theta(p)=L_theta(p)L_theta(p)^T + epsilon I
   |
   v
Vhat_theta(x,p)=(x-xs)^T P_theta(p)(x-xs)

Why epsilon I is added

The product Lθ(p)Lθ(p)T is always positive semidefinite. That is useful, but it is not enough. If Lθ is rank deficient, or if some diagonal entries are zero, the product may have zero eigenvalues. In that case the terminal cost may fail to be strictly positive for nonzero state errors in some directions.

Adding εI shifts every eigenvalue upward by ε. Therefore,

Pθ (p) ≻ 0 for every parameter p.

This is a structural guarantee. It does not depend on the training data being perfect. It follows from the matrix construction itself. Numerically, it also protects the terminal metric from becoming singular or nearly degenerate in directions where the network outputs a weak factor.

Why it can work

The method can work because it uses offline computation to amortize part of the long-horizon NMPC problem. The expensive long-horizon solves teach the terminal surrogate what future cost may look like, while the online controller solves a much smaller nonlinear program.

There is also a useful inductive bias. Many stabilizing control designs use quadratic Lyapunov-like functions near an equilibrium or steady state. A parameter-dependent positive-definite metric can be interpreted as learning how the local geometry of the terminal cost should change across operating regimes. This is more structured than asking an unconstrained network to output any scalar value.

But the evidence and the guarantee must be separated. The supplied material supports the following reading:

  • the positive definiteness of Pθ(p) is guaranteed by construction;
  • the positive definiteness of V^θ with respect to x-xs is guaranteed by construction;
  • exact horizon compression is valid only if the surrogate equals the true continuation value;
  • closed-loop NMPC stability still depends on Lyapunov decrease, feasibility, terminal-set logic, or equivalent assumptions.

This distinction is the heart of the note. The paper’s strongest idea is the neural network generated Cholesky-type metric. Its weakest unresolved point is that this algebraic positive definiteness is weaker than a closed-loop stability guarantee when approximation error is present.

Assumptions and limitations

Positive definiteness does not imply Lyapunov decrease. A function can be strictly positive away from the target and still increase along closed-loop trajectories. Therefore, from

V^θ (x,p) > 0 for x≠xs

one cannot conclude that

V^θ (x+,p+) - V^θ (x,p) ≤ - α (|x-xs|) .

The first statement is a shape property of the terminal cost. The second is a trajectory property of the closed-loop system. They are not equivalent.

Several limitations follow.

  • Sampled decrease penalties do not imply global decrease over all states and parameters.
  • Offline training coverage does not guarantee behavior outside the training distribution.
  • The online optimization problem remains nonlinear when the dynamics are nonlinear.
  • Reducing the NLP horizon reduces dimension, but does not remove local optima or numerical failures.
  • The quadratic-in-state-error form may be too restrictive for strongly nonquadratic value landscapes.
  • Recursive feasibility is not automatically guaranteed by this learned terminal cost.
  • Classical terminal sets and Lyapunov ingredients are not fully removed; much of their role is shifted into offline data generation, supervision, and empirical validation.

This is a fair tradeoff, not a fatal flaw. The method is useful precisely because it buys online speed with offline computation and structural bias. The important point is to state what is bought and what remains unpaid.

Critical assessment: approximation-error critique

The most important missing analysis is approximation-error-aware stability or performance. The clean Bellman-style argument depends on the learned surrogate being exact. In practice, the relevant error is

| V^θ (x,p) - V* (x,p) |.

If this error is nonzero, the one-step controller is no longer solving the exact horizon-compressed problem. A stronger paper would specify what remains true under a bound such as

| V^θ (x,p) - V* (x,p) | ≤ εV .

Here, εV would measure the worst-case value approximation error over a specified domain. Such a result would not automatically give asymptotic stability, but it could support bounded performance degradation if the domain, dynamics, and optimization errors are controlled.

Even more directly, one could ask for a practical Lyapunov decrease statement:

V^θ (x+,p+) - V^θ (x,p) ≤ - α (|x-xs|) + εdec .

The term εdec would represent the residual decrease error. If it is small, the conclusion would likely be practical stability or bounded ultimate behavior, not exact asymptotic convergence. That would be a more honest and useful guarantee for a learned terminal cost.

The supplied material does not establish that such an error-aware theorem is proved. Therefore, the correct reading is conservative: the structure guarantees a positive-definite terminal metric, while closed-loop behavior under approximation error remains a separate analytical question.

Balanced takeaway

This method is best understood as a data-driven horizon-compression method for NMPC. Its main contribution is not that a neural network appears in the controller. The contribution is the NN-generated Cholesky/Lyapunov metric:

NN predicts a matrix factor, not an arbitrary scalar value.
The factor creates a positive-definite terminal metric.
The metric defines a Lyapunov-compatible quadratic terminal cost.

That is a meaningful design choice. It gives the learned terminal cost a control-theoretic shape that generic value approximation lacks.

The limitation is equally clear. Structural positive definiteness is not the same as recursive feasibility, closed-loop Lyapunov decrease, global stability, or exact continuation-value approximation. The method can reduce online computation, but it requires offline long-horizon NMPC solves and inherits the usual risks of function approximation, training distribution mismatch, nonlinear optimization, and feasibility-critical control.

So the right claim is moderate: this is a promising structured surrogate for fast NMPC, especially when long-horizon solutions can be generated offline and the operating domain is well covered. It should not be read as proving that neural networks replace stabilizing MPC design or that global NMPC stability is solved.

References

  • Abdufattokhov, S., Zanon, M., & Bemporad, A. (2024). Learning Lyapunov terminal costs from data for complexity reduction in nonlinear model predictive control. International Journal of Robust and Nonlinear Control, 34(13), 8676-8691.