When trying to uncover tacit knowledge behind complex supply-chain decisions, the obvious data-driven approach is to learn a mapping from situation to decision. Give a black-box model the demand forecast, current inventory, and production conditions, then train it to reproduce the schedule chosen by an experienced planner. Such a model may predict well, but it still leaves the central question unanswered: why did the planner prefer this decision over the other feasible alternatives?
The more interesting approach is to model the objective behind the decision. Perhaps the planner cares far more about avoiding shortages than reducing inventory cost. Perhaps a regular production cycle matters because it protects the plant against uncertain future demand. If these priorities can be inferred from historical decisions, the resulting model does more than imitate an expert. It expresses the expert’s decision logic as an interpretable objective function and generates a decision by solving an optimization problem under known operational constraints. The process becomes inspectable: we can see which trade-offs drive a plan, challenge their interpretation, and identify where the model fails to explain actual practice.
This is why I wanted to introduce the study by Dixit et al. Their Dow production-planning case is specific, but the underlying idea has room to extend to many supply-chain decisions in which constraints are reasonably well understood while preferences remain implicit. Procurement, inventory positioning, production allocation, network operation, and rolling-horizon planning all contain versions of the same problem. The paper does not complete those extensions, but it provides a useful starting architecture: preserve the known mixed-integer model, infer a small set of objective weights from expert plans, and put the learned objective back inside the optimizer.
The forward problem is known; the objective is not
For planning instance , the forward problem is written compactly as
The variables describe production, inventory, campaign starts and ends, cycle lengths, and production gaps. The Dow plant produces eight products on one unit over a 200-period horizon. Product order is fixed by a higher-level planning problem; this MILP chooses campaign lengths while enforcing inventory and scheduling constraints.
The unknown objective is not treated as an unrestricted vector. It is assembled from seven hypothesized cost terms elicited through domain knowledge:
- positive-inventory holding cost;
- lower and upper deviations from desired inventory at campaign starts;
- lower and upper deviations from desired inventory at campaign ends;
- violations of a preferred cycle-length range;
- gaps between production campaigns.
The objective is
Here normalizes the terms and represents the planner’s relative preference. This is already a strong modeling decision. The method does not discover arbitrary human motives; it estimates weights within a basis chosen in advance.
Why direct inverse MILP is awkward
Suppose the historical data contain pairs : the demand and inventory conditions, followed by the plan made by an expert. A direct formulation would choose a cost vector and require each predicted plan to be optimal for a lower-level MILP. It would then minimize the distance between predicted and observed decisions.
This creates a bilevel problem with one mixed-integer lower level per observation. KKT conditions or strong-duality reformulations are not available in the usual way because the lower problems are MILPs. Keeping every predicted schedule and all of its mixed-integer constraints in the inverse master also makes the formulation grow quickly with the number of observations.
The paper instead uses a suboptimality loss. For instance , let
The loss asks how much worse the observed expert plan is than the best feasible plan under the estimated objective:
If , the observed plan is optimal under the estimated objective. A positive value permits a noisy or merely near-optimal expert. The inverse problem minimizes the sum of these gaps plus an penalty that favors a sparse objective. The admissible set for must also rule out the meaningless all-zero objective and fix the otherwise unidentified scale.
This substitution changes what the model tries to reproduce. Decision loss asks for the same schedule. Suboptimality loss asks for an objective under which the expert schedule is hard to beat. Multiple schedules can have similar costs, so the second criterion is weaker. It is also the reason the inverse master no longer needs a predicted mixed-integer schedule for every observation.
Cutting planes turn the universal comparison into a solvable loop
The suboptimality constraint must hold against every feasible alternative:
Enumerating all alternatives is impossible. The algorithm begins with a finite subset of schedules for each training instance. With those schedules fixed, the master problem for the objective weights and suboptimality variables is an LP. It then solves a forward MILP for each observation under the current estimated cost:
If this schedule beats the expert by more than the current , it exposes a violated constraint and is added to the master. The loop alternates between a small LP and independent forward MILPs. It keeps only schedules that currently challenge the inferred objective.
The convergence statement needs a qualifier. To certify that no violated constraint remains, the final cut-generating MILPs must be solved to global optimality. The authors use limited solve times for faster cuts in early iterations and require optimality toward the end. A time-limited separation problem that fails to find a better schedule is not, by itself, a proof that no such schedule exists.
What the Dow case reveals
The dataset contains 70 historical plans: 50 for training and 20 for testing, evaluated over five random splits. A model using inventory holding cost alone predicts inventory levels that are systematically too low. Adding penalties for desired inventory ranges improves the match, and the full seven-term objective improves it further.
The learned seven-term model assigns essentially zero weight to explicit inventory holding cost. Avoiding low inventory, maintaining acceptable cycle lengths, and avoiding production gaps matter more. Interviews with an expert planner support much of this interpretation. This is the paper’s most persuasive result: costs that are easy to measure need not be the costs that govern decisions.
The disagreements are equally informative. Some weight rankings suggested an explanation that the expert rejected. Short production gaps in the historical plans also weakened the inferred gap penalty, even though the planner said intentional idling is normally unacceptable. Unmodeled disruptions may have produced those gaps. Inverse optimization then absorbed a missing constraint or external event into the objective weights.
The authors also let the weights vary across four time buckets and across products. Predictive error decreases. The time-dependent model suggests that shortage avoidance is stronger near the start of the rolling horizon, while regular cycle lengths matter more farther out when demand is uncertain. Product-dependent weights are consistent with different margins and customer priorities, but they can also reflect demand bias or omitted operating conditions. Better fit does not identify which explanation is correct.
What is guaranteed, and what is inferred
The cleanest guarantee is feasibility with respect to the modeled forward problem. A new schedule is still obtained by solving the MILP, so it satisfies the constraints encoded in . A direct neural predictor would need additional machinery to make the same statement. This guarantee is only as good as the constraint model. Maintenance, operator availability, quality events, or customer-specific rules that are missing from the MILP are not covered.
The method does not guarantee recovery of a planner’s true psychological objective. At least five issues remain.
First, the weighted-sum basis is prespecified. A planner may use lexicographic rules—avoid shortages first, then minimize gaps, then reduce inventory—rather than a compensatory weighted sum.
Second, inverse objectives are identifiable only up to positive scale, and highly correlated cost features may admit many nearly equivalent weight vectors. Normalization makes relative weights readable; it does not prove uniqueness.
Third, the study pools plans from multiple planners under one objective. The estimated vector may be a population compromise rather than any individual’s preference.
Fourth, the records are rolling-horizon plans, not fully implemented trajectories. Decisions late in the horizon may be provisional. Time-varying weights may capture this planning convention rather than a genuine change in preference.
Fifth, allowing weights to vary by time and product adds substantial flexibility. The lower test RMSE is useful evidence, but it does not by itself separate real preference heterogeneity from additional fitting capacity.
The right interpretation is therefore modest: the method finds an objective within a chosen feature space that makes historical expert plans approximately optimal under a chosen constraint model. That is weaker than discovering the true objective. It is still valuable. The result is an interpretable, optimization-compatible hypothesis about tacit planning logic, and disagreements with experts become diagnostics for missing objectives, missing constraints, or bad data.
Reference
Dixit, S., Gupta, R., Kelloway, A., Wassick, J., & Zhang, Q. (2026). Uncovering expert objectives in production planning via inverse optimization: An industrial case study. Chemical Engineering Research and Design. https://doi.org/10.1016/j.cherd.2026.07.065. arXiv:2608.07398.