Surrogate Modeling for Process Prediction and Optimization · Week 2

Week 2: Learning Multicomponent Steady-State Responses with MLPs

Train and evaluate a steady-state three-output MLP; follow one gradient update and check concentration consistency.

Week 2: Learning Multicomponent Steady-State Responses with MLPs

More conversion does not tell the whole composition

The desired intermediate is B in A → B → C. At short residence time little B forms; at long residence time more B reacts onward to C.

An engineer needs all three outlet concentrations at each condition. One conversion number does not describe the product composition.

Can one model predict A, B, and C consistently?

A process engineer examines the three reactor products A, B, and C.
Sunwoo Kim · Surrogate Modeling for Process Prediction and Optimization Week 2 · 1 / 21

Three outputs, different errors

All three concentrations have the same units, but their variation differs. An unscaled loss can emphasize one component.

This week learns how an MLP computes these outputs, how one gradient update changes its parameters, and how to evaluate each component.

Sunwoo Kim · Surrogate Modeling for Process Prediction and Optimization Week 2 · 2 / 21

Learn and check the steady-state MLP

  1. Derive steady-state labels for A, B, and C.
  2. Scale the data and follow a forward pass, backpropagation, and one parameter update.
  3. Train the MLP and compare component RMSE and concentration-sum residuals.
Sunwoo Kim · Surrogate Modeling for Process Prediction and Optimization Week 2 · 3 / 21

1. A → B → C in a CSTR

Consider a perfectly mixed, constant-volume, constant-density, isothermal CSTR with consecutive first-order reactions A → B → C. The feed contains A; its B and C concentrations are zero. Temperature is an operating input. With concentrations in mol/L, temperature in K, and time in min, the balances are:

dCAdt=CAf−CAτ−k1CAdCBdt=−CBτ+k1CA−k2CBdCCdt=−CCτ+k2CB
Sunwoo Kim · Surrogate Modeling for Process Prediction and Optimization Week 2 · 4 / 21

Rate constants and operating domain

Both rate constants have units of 1/min. Use the following temperature dependence:

k1(T)=0.2exp[6000(1330−1T)]k2(T)=0.1exp[5000(1330−1T)]
Quantity Value or range
Reference temperature 330 K
Reference k₁; k₂ 0.2; 0.1 per minute
Temperature-sensitivity parameters 6000 K; 5000 K
Temperature T 300–360 K
Residence time τ 1–10 min
Feed A concentration 0.8–1.5 mol/L

The three balances describe A consumption, B formation and consumption, and C formation. Adding them cancels the reaction terms.

Sunwoo Kim · Surrogate Modeling for Process Prediction and Optimization Week 2 · 5 / 21

2. Derive the three-output steady-state map

Set the time derivatives to zero and multiply each balance by τ. Collect the concentration terms:

(1+k1τ)CA=CAf(1+k2τ)CB=k1τCACC=k2τCB
Sunwoo Kim · Surrogate Modeling for Process Prediction and Optimization Week 2 · 6 / 21

Solve for A, B, and C

Solve in the order A, B, C:

CA=CAf1+k1τCB=k1τCAf(1+k1τ)(1+k2τ)CC=k2τCB

At 330 K, τ = 5 min, and feed = 1.2 mol/L, the concentrations are [0.6, 0.4, 0.2] mol/L. Their sum equals the feed concentration. These three outputs are the labels for the steady-state MLP.

Sunwoo Kim · Surrogate Modeling for Process Prediction and Optimization Week 2 · 7 / 21

Why the B response is nonmonotone

At a fixed temperature, the B yield and its maximizing residence time follow from the expression for C_B:

YB=CBCAf=k1τ(1+k1τ)(1+k2τ),τ*=1k1k2
Sunwoo Kim · Surrogate Modeling for Process Prediction and Optimization Week 2 · 8 / 21

Compare the three concentration profiles

At 330 K, τ* is about 7.07 min. Short residence times limit B formation; long residence times allow more B to react to C.

A, B, and C concentration responses to residence time at 330 kelvin, comparing the reference process and MLP.
Three-component response at 330 K. The B concentration has an interior maximum.
Sunwoo Kim · Surrogate Modeling for Process Prediction and Optimization Week 2 · 9 / 21

3. Prepare the supervised dataset

Each row contains inputs [T, τ, C_Af] and outputs [C_A, C_B, C_C]. Input and output arrays both have shape (N, 3). Generate 400 train, 120 validation, and 160 test operating points within the table’s bounds, then calculate their labels with the steady-state map.

The three concentrations have the same units but different variations. In the training data, their standard deviations are approximately [0.283, 0.110, 0.217] mol/L.

Sunwoo Kim · Surrogate Modeling for Process Prediction and Optimization Week 2 · 10 / 21

4. Forward calculation and scaling

Use a 3 → 32 → 32 → 3 MLP, with ReLU hidden layers and a linear output layer:

Each hidden layer first computes a weighted sum and bias. ReLU(a) = max(0, a) retains positive values and sets negative values to zero. The weights and biases are the parameters fitted to the concentration data.

h1=ReLU(W1x̃+b1)h2=ReLU(W2h1+b2)z^=W3h2+b3
Layer Weight shape Bias shape Parameters
Input → hidden 1 32 × 3 32 128
Hidden 1 → hidden 2 32 × 32 32 1056
Hidden 2 → output 3 × 32 3 99
Total     1283
Sunwoo Kim · Surrogate Modeling for Process Prediction and Optimization Week 2 · 11 / 21

Scaling and the training loss

Scale the inputs with the declared lower and upper bounds. Standardize each output with its training mean μ and standard deviation s, and restore physical units after prediction:

x̃j=2xj−ljuj−lj−1zj=yj−μjsj,y^j=μj+sjz^j

The loss averages squared errors across samples and standardized output components:

L=13N∑n=1N∑j=13(z^n,j−zn,j)2

Thus a concentration error is measured relative to that component’s training variation. For example, the smaller B standard deviation gives its raw squared error a larger weight in this loss.

Sunwoo Kim · Surrogate Modeling for Process Prediction and Optimization Week 2 · 12 / 21

5. Backpropagation: one calculation

Take a 3 → 2 → 3 network with zero biases and the following values. For this calculation, use L = ½‖ŷ − y‖² and learning rate 0.1.

x=[10-1],y=[010]W1=[100001],W2=[100111]

The hidden pre-activation a, hidden activation h, prediction, and error are:

a=[1-1],h=[10],y^=[101],e=[1-11]
Sunwoo Kim · Surrogate Modeling for Process Prediction and Optimization Week 2 · 13 / 21

Gradients and the parameter update

The output error propagates through the output weights and the ReLU derivative:

δ2=e,∇W2L=ehT=[10-1010]δ1=(W2Te)⊙[10]=[20]∇W1L=δ1xT=[20-2000]

The bias gradients are δ₁ and δ₂. Update every weight and bias with parameter ← parameter − 0.1 × gradient. The next forward calculation gives [0.26, 0.14, 0.26], and the loss decreases from 1.5 to 0.4374. The notebook checks these gradients using autograd.

Sunwoo Kim · Surrogate Modeling for Process Prediction and Optimization Week 2 · 14 / 21

6. Train and evaluate the steady-state model

Use Adam with learning rate 0.001, batches of 64, and up to 1400 epochs. After each epoch, evaluate validation loss and retain the weights at its minimum. One batch follows these operations:

prediction = model(inputs)
loss = loss_function(prediction, targets)
optimizer.zero_grad()
loss.backward()
optimizer.step()
Training and validation scaled mean squared errors for the steady-state MLP.
Training and validation loss. Prediction metrics below are computed after restoring mol/L units.
Sunwoo Kim · Surrogate Modeling for Process Prediction and Optimization Week 2 · 15 / 21

Evaluate each output component

Test RMSE, mol/L A B C
Output scaling 0.005275 0.002940 0.004423
Unscaled outputs 0.006634 0.003917 0.004162

In this comparison, output scaling reduces A and B RMSE; C RMSE is slightly higher. Evaluate each component rather than combining all three into one error number.

Reference-versus-predicted parity plots for the three concentrations on the test set.
Separate parity plots reveal each component's errors.
Sunwoo Kim · Surrogate Modeling for Process Prediction and Optimization Week 2 · 16 / 21

Check the concentration sum

Also calculate the concentration-sum residual:

r=CA^+CB^+CC^−CAf

Its test RMSE is 0.003781 mol/L. Component prediction errors and the residual describe different aspects of the model.

The three linear outputs are free predictions. Balance-consistent training labels do not force their sum to match the feed concentration.

For feed 1.2 mol/L, predictions [0.62, 0.41, 0.20] mol/L sum to 1.23 mol/L, giving r = 0.03 mol/L. A model predicting mole fractions can likewise return a sum different from one.

A small prediction error does not enforce a required equality.

Sunwoo Kim · Surrogate Modeling for Process Prediction and Optimization Week 2 · 17 / 21

Week 4 preview: put the constraint into the model

We can correct a prediction after training, or include a constraint-respecting transformation in the training calculation.

Approach Where the correction enters
Post-training mapping/projection Train the raw predictor, then correct its outputs.
KKT-hPINN Apply a fixed affine projection using the input and raw prediction before the loss, and at inference.
x→fθ→y^→Πx→y~→loss

KKT-hPINN computes loss on the projected output and backpropagates through the projection. Its coefficients are fixed; the MLP weights learn with that transformation in place.

Week 4 derives the projection and KKT conditions, checks rank assumptions, and follows the gradient. The guarantee concerns the specified linear equalities, up to numerical arithmetic; nonnegativity needs an additional constraint.

Source: Chen et al. (2024), DOI: 10.1016/j.compchemeng.2024.108764.

Sunwoo Kim · Surrogate Modeling for Process Prediction and Optimization Week 2 · 18 / 21

7. Run the required steady-state lab

Download the steady-state notebook.

python -m pip install -r requirements.txt
python week02_lab.py --mode steady --output-dir week02-results

Evaluate the trained predictor by component RMSE and concentration-sum residual.

Sunwoo Kim · Surrogate Modeling for Process Prediction and Optimization Week 2 · 19 / 21

8. Core exercises and submission

  1. Derive C_A, C_B, and C_C and verify their sum. Explain the nonmonotone B curve.
  2. Count the 1283 steady-state MLP parameters and verify the hand-calculated gradient update.
  3. Compare scaled and unscaled outputs by each component’s RMSE.
  4. Inspect concentration-sum residual and minimum predicted concentration on held-out inputs.

Submit the steady-state model, split/scaling details, checked gradient update, and component/consistency evaluation.

Original explanations: PyTorch model construction, autograd, and training.

Sunwoo Kim · Surrogate Modeling for Process Prediction and Optimization Week 2 · 20 / 21

Reading companion

The Week 2 companion and English PDF explain the three-output map, parameter count, scaling, gradient update, training loop, and independent diagnostics.

KKT-hPINN reference: Chen, H., Constante Flores, G. E., & Li, C. (2024). Physics-informed neural networks with hard linear equality constraints. Computers & Chemical Engineering, 189, 108764.

Sunwoo Kim · Surrogate Modeling for Process Prediction and Optimization Week 2 · 21 / 21