Week 2: Learning Multicomponent Steady-State Responses with MLPs
More conversion does not tell the whole composition
The desired intermediate is B in A → B → C. At short residence time little B forms; at long residence time more B reacts onward to C.
An engineer needs all three outlet concentrations at each condition. One conversion number does not describe the product composition.
Can one model predict A, B, and C consistently?

Three outputs, different errors
All three concentrations have the same units, but their variation differs. An unscaled loss can emphasize one component.
This week learns how an MLP computes these outputs, how one gradient update changes its parameters, and how to evaluate each component.
Learn and check the steady-state MLP
- Derive steady-state labels for A, B, and C.
- Scale the data and follow a forward pass, backpropagation, and one parameter update.
- Train the MLP and compare component RMSE and concentration-sum residuals.
1. A → B → C in a CSTR
Consider a perfectly mixed, constant-volume, constant-density, isothermal CSTR with consecutive first-order reactions A → B → C. The feed contains A; its B and C concentrations are zero. Temperature is an operating input. With concentrations in mol/L, temperature in K, and time in min, the balances are:
Rate constants and operating domain
Both rate constants have units of 1/min. Use the following temperature dependence:
| Quantity | Value or range |
|---|---|
| Reference temperature | 330 K |
| Reference k₁; k₂ | 0.2; 0.1 per minute |
| Temperature-sensitivity parameters | 6000 K; 5000 K |
| Temperature T | 300–360 K |
| Residence time τ | 1–10 min |
| Feed A concentration | 0.8–1.5 mol/L |
The three balances describe A consumption, B formation and consumption, and C formation. Adding them cancels the reaction terms.
2. Derive the three-output steady-state map
Set the time derivatives to zero and multiply each balance by τ. Collect the concentration terms:
Solve for A, B, and C
Solve in the order A, B, C:
At 330 K, τ = 5 min, and feed = 1.2 mol/L, the concentrations are [0.6, 0.4, 0.2] mol/L. Their sum equals the feed concentration. These three outputs are the labels for the steady-state MLP.
Why the B response is nonmonotone
At a fixed temperature, the B yield and its maximizing residence time follow from the expression for C_B:
Compare the three concentration profiles
At 330 K, τ* is about 7.07 min. Short residence times limit B formation; long residence times allow more B to react to C.

3. Prepare the supervised dataset
Each row contains inputs [T, τ, C_Af] and outputs [C_A, C_B, C_C]. Input and output arrays both have shape (N, 3). Generate 400 train, 120 validation, and 160 test operating points within the table’s bounds, then calculate their labels with the steady-state map.
The three concentrations have the same units but different variations. In the training data, their standard deviations are approximately [0.283, 0.110, 0.217] mol/L.
4. Forward calculation and scaling
Use a 3 → 32 → 32 → 3 MLP, with ReLU hidden layers and a linear output layer:
Each hidden layer first computes a weighted sum and bias. ReLU(a) = max(0, a) retains positive values and sets negative values to zero. The weights and biases are the parameters fitted to the concentration data.
| Layer | Weight shape | Bias shape | Parameters |
|---|---|---|---|
| Input → hidden 1 | 32 × 3 | 32 | 128 |
| Hidden 1 → hidden 2 | 32 × 32 | 32 | 1056 |
| Hidden 2 → output | 3 × 32 | 3 | 99 |
| Total | 1283 |
Scaling and the training loss
Scale the inputs with the declared lower and upper bounds. Standardize each output with its training mean μ and standard deviation s, and restore physical units after prediction:
The loss averages squared errors across samples and standardized output components:
Thus a concentration error is measured relative to that component’s training variation. For example, the smaller B standard deviation gives its raw squared error a larger weight in this loss.
5. Backpropagation: one calculation
Take a 3 → 2 → 3 network with zero biases and the following values. For this calculation, use L = ½‖ŷ − y‖² and learning rate 0.1.
The hidden pre-activation a, hidden activation h, prediction, and error are:
Gradients and the parameter update
The output error propagates through the output weights and the ReLU derivative:
The bias gradients are δ₁ and δ₂. Update every weight and bias with parameter ← parameter − 0.1 × gradient. The next forward calculation gives [0.26, 0.14, 0.26], and the loss decreases from 1.5 to 0.4374. The notebook checks these gradients using autograd.
6. Train and evaluate the steady-state model
Use Adam with learning rate 0.001, batches of 64, and up to 1400 epochs. After each epoch, evaluate validation loss and retain the weights at its minimum. One batch follows these operations:
prediction = model(inputs)
loss = loss_function(prediction, targets)
optimizer.zero_grad()
loss.backward()
optimizer.step()

Evaluate each output component
| Test RMSE, mol/L | A | B | C |
|---|---|---|---|
| Output scaling | 0.005275 | 0.002940 | 0.004423 |
| Unscaled outputs | 0.006634 | 0.003917 | 0.004162 |
In this comparison, output scaling reduces A and B RMSE; C RMSE is slightly higher. Evaluate each component rather than combining all three into one error number.

Check the concentration sum
Also calculate the concentration-sum residual:
Its test RMSE is 0.003781 mol/L. Component prediction errors and the residual describe different aspects of the model.
The three linear outputs are free predictions. Balance-consistent training labels do not force their sum to match the feed concentration.
For feed 1.2 mol/L, predictions [0.62, 0.41, 0.20] mol/L sum to 1.23 mol/L, giving r = 0.03 mol/L. A model predicting mole fractions can likewise return a sum different from one.
A small prediction error does not enforce a required equality.
Week 4 preview: put the constraint into the model
We can correct a prediction after training, or include a constraint-respecting transformation in the training calculation.
| Approach | Where the correction enters |
|---|---|
| Post-training mapping/projection | Train the raw predictor, then correct its outputs. |
| KKT-hPINN | Apply a fixed affine projection using the input and raw prediction before the loss, and at inference. |
KKT-hPINN computes loss on the projected output and backpropagates through the projection. Its coefficients are fixed; the MLP weights learn with that transformation in place.
Week 4 derives the projection and KKT conditions, checks rank assumptions, and follows the gradient. The guarantee concerns the specified linear equalities, up to numerical arithmetic; nonnegativity needs an additional constraint.
Source: Chen et al. (2024), DOI: 10.1016/j.compchemeng.2024.108764.
7. Run the required steady-state lab
Download the steady-state notebook.
python -m pip install -r requirements.txt
python week02_lab.py --mode steady --output-dir week02-results
Evaluate the trained predictor by component RMSE and concentration-sum residual.
8. Core exercises and submission
- Derive C_A, C_B, and C_C and verify their sum. Explain the nonmonotone B curve.
- Count the 1283 steady-state MLP parameters and verify the hand-calculated gradient update.
- Compare scaled and unscaled outputs by each component’s RMSE.
- Inspect concentration-sum residual and minimum predicted concentration on held-out inputs.
Submit the steady-state model, split/scaling details, checked gradient update, and component/consistency evaluation.
Original explanations: PyTorch model construction, autograd, and training.
Reading companion
The Week 2 companion and English PDF explain the three-output map, parameter count, scaling, gradient update, training loop, and independent diagnostics.
KKT-hPINN reference: Chen, H., Constante Flores, G. E., & Li, C. (2024). Physics-informed neural networks with hard linear equality constraints. Computers & Chemical Engineering, 189, 108764.