Chemical Plants

TalkToAgent: Using Existing LLMs to Explain RL Controllers in Chemical Plants

A critical note on TalkToAgent, an LLM multi-agent framework that reuses existing LLMs to route natural-language questions to XRL tools for chemical process-control policies.

How an LLM can help with controller explanations

The useful role of an LLM in this setting is not to replace process-control analysis. It is to make the analysis easier to ask for. A chemical engineer may not want to decide whether a question requires SHAP, Q-value decomposition, a rollout comparison, or a generated contrastive policy. The LLM can sit between the user’s question and those tools: classify the intent, call the right computation, help assemble executable analysis code when needed, and summarize the resulting figures or numbers in ordinary language.

The LLM helps translate between engineering questions and XRL workflows. The plant-facing evidence still has to come from the controller, the simulator, the reward decomposition, and the contrastive trajectories.

Still, this is not chemical-domain fine-tuning

The first thing to clarify is what this paper is not. This is not a study that fine-tunes an LLM to become a chemical-engineering or chemical-plant specialist. It does not claim that the language model has learned process dynamics, reaction kinetics, plant safety constraints, or controller design from domain-specific training.

The idea is narrower and more practical: take existing LLMs and use them as an interface layer around already defined explainable-RL tools. The LLM is not the source of truth. It routes a user’s natural-language question to an XRL computation, helps generate small pieces of executable analysis code when needed, and turns plots or numerical outputs into a readable explanation. The explanation is grounded in DeepSHAP, Q-value decomposition, simulator rollouts, and contrastive policy comparisons, not in free-form LLM speculation.

That distinction matters for chemical plants. In a safety-critical process-control setting, a fluent explanation is not enough. If the LLM says “the controller increased coolant flow because the reactor temperature was rising,” the useful question is: which attribution, rollout, reward component, or counterfactual simulation supports that sentence?

Why this problem matters for process control

Deep reinforcement learning is attractive in process control because chemical plants are nonlinear, coupled, constrained, and often operated in continuous state-action spaces. A learned policy can sometimes handle dynamics that are awkward for a simple rule-based controller.

The problem is trust. A chemical engineer cannot simply deploy a black-box RL policy because its average simulation performance looks good. The operator needs to ask questions such as: “Why did the agent choose this control action?”, “Which process variables drove the action?”, “What future reward trade-off was the policy expecting?”, “What would have happened under a more conservative behavior?”, and “How would a simple on-off policy compare in this state region?”

Traditional XRL tools answer fragments of these questions. Feature attribution can show which state variables influenced the current action. Reward decomposition can show which future reward terms dominate. A contrastive rollout can compare one action with another. But those tools are scattered, technical, and often require the user to know which method to call.

TalkToAgent tries to reduce that interface burden. The contribution is not a new RL theorem. It is an engineering framework that connects natural-language questions to a structured set of XRL computations for process-control policies.

System structure

The framework has three main roles:

user query
  -> Coordinator: infer intent and select an XRL tool
  -> XRL tool: attribution, reward decomposition, or contrastive simulation
  -> Explainer: translate figures and numerical outputs into natural language

for generated code:
  Coder -> Evaluator -> Debugger -> refined code

The Coordinator maps a question to one of several explanation modes. If the user asks which variables mattered, the system can use a feature-importance tool such as DeepSHAP. If the user asks what the agent expected to gain, it can use Q-value or reward-component decomposition with rollout results. If the user asks “what if the controller behaved differently?”, it can run a contrastive simulation.

The Coder-Evaluator-Debugger loop is used when the explanation requires generated code, especially for reward decomposition or policy-level contrastive explanations. This is a sensible use of LLMs: not to assert a plant explanation directly, but to help assemble executable analysis that can be checked against the simulator.

The XRL tools being orchestrated

The first tool family is feature importance. The paper uses DeepSHAP to attribute an action output to input state variables. In a tank system, for example, the explanation might say that a lower-tank level error contributed strongly to a pump-voltage action. This is useful, but it is still local attribution. It does not prove dynamic causality inside the closed-loop plant.

The second family is expected-outcome explanation. Instead of asking only “which state variable changed the action?”, it asks what future trajectory and reward trade-off the action is associated with. A policy in a quadruple-tank system may accept a short-term imbalance in one level to improve a later tracking objective or reduce control variation. This is closer to the logic of RL because the policy is trained against discounted future return, not only the current error.

In simplified notation, the value being explained is:

Qπ (s,a) = E [ ∑t=0∞ γt rt | s0=s, a0=a, π ] .

The caveat is that this “expectation” is only as good as the learned critic, the reward decomposition, and the simulator rollouts used to approximate it. In stochastic or model-mismatched plants, a representative rollout is not the same thing as a verified expectation over the real plant.

The third family is contrastive explanation. The paper separates this into three levels: “CE-A: action-level contrast”, “CE-B: behavior-level contrast”, and “CE-P: policy-level contrast”.

CE-A changes a specific action and compares the resulting trajectory. This is direct, but it requires the user to specify a numerical alternative action. CE-B accepts qualitative control language such as “more conservative,” “more aggressive,” or “opposite,” then maps that phrase into an action-trajectory transformation. CE-P generates a simple alternative policy, such as an on-off controller, and compares its rollout with the RL policy.

This is the most useful application-facing extension. A chemical engineer is more likely to ask about conservative operation, aggressive cooling, or an on-off fallback policy than to request a specific numerical action perturbation at one time step.

Why the framework is useful

The useful part of TalkToAgent is the interaction design around XRL. It changes the user’s job from “choose the correct explanation algorithm and its arguments” to “ask a process-control question.” That is a real reduction in cognitive load.

It also avoids one obvious failure mode of LLM explanations. The language model is not asked to invent a reason from scratch. It is asked to interpret artifacts produced by predefined tools: SHAP values, reward-component plots, forward simulations, and contrastive trajectories. This does not eliminate hallucination, but it gives the explanation an external anchor.

For chemical plants, that matters because the explanation should be inspectable. A reasonable workflow is:

natural-language question
  -> selected XRL tool
  -> simulator-grounded computation
  -> figure or numerical comparison
  -> natural-language explanation
  -> operator checks whether the explanation matches process intuition

The last step is important. This is not an autonomous safety-certification system. It is an interactive analysis layer for understanding a learned controller.

Experimental setting

The paper evaluates the framework on three process-control benchmarks: a continuous stirred-tank reactor, a quadruple-tank system, and a photobioreactor. The RL controller is trained with SAC. The examples are well chosen because they cover different process-control structures: setpoint tracking, coupled liquid-level dynamics, and yield-oriented bioprocess operation.

The evaluation asks whether the system can map user queries to the correct XRL function and arguments, whether it can generate contrastive policy code reliably enough to run, and whether the resulting multimodal explanations contain domain-relevant information. The reported task-classification accuracy is high across the three benchmarks, with some errors in contrastive queries that are semantically ambiguous. The code-generation loop also reduces repeated failures compared with direct generation, although policy-level contrastive examples are still small in scale.

The cost profile is also worth noting. Simple feature-importance or rollout explanations are relatively cheap. CE-P can consume many more tokens because it involves code generation, evaluation, debugging, and another simulation. In a real deployment, that difference matters. A plant-facing explanation system would need latency, budget, and audit controls.

Critical reading

The strongest limitation is explanation faithfulness. SHAP, reward decomposition, and contrastive rollout are post-hoc tools. They can make a controller more interpretable, but they do not prove that the neural policy internally “reasoned” in the way the explanation describes. A good phrasing is: this action is interpretable through these state attributions and simulated consequences. A bad phrasing is: the agent truly made the decision for this human-like reason.

The second limitation is simulator dependence. Expected-outcome and contrastive explanations depend on forward simulation. If the simulator misses actuator delays, measurement noise, unmodeled constraints, fouling, regime shifts, or plant-model mismatch, the contrastive explanation can be clean and still misleading.

The third limitation is the semantic gap in CE-B. Mapping “conservative” to a smoothing parameter or “aggressive” to an amplified action trajectory is intuitive, but it is a heuristic. In control terms, conservative behavior should ideally be defined through an optimization problem or explicit behavioral constraints, such as smaller action variation, bounded overshoot, or stricter safety margins. Without that, different users may mean different things by the same word.

The fourth limitation is CE-P. Comparing an RL policy against a generated on-off controller is useful, but it should not be oversold as general policy-level explanation. The current form is closer to local counterfactual policy simulation over a chosen interval or condition than to a global explanation of the learned policy.

Finally, the Evaluator is not a formal verifier. It may reduce code-generation errors, but it does not guarantee that a generated policy satisfies the intended process-control specification. For safety-facing use, the generated code should be checked by executable assertions or formal specifications, not only by another LLM agent.

Takeaway

TalkToAgent is best read as a practical orchestration layer for explainable RL in chemical process control. Its value is not that an LLM becomes a chemical plant expert through fine-tuning. Its value is that an existing LLM can help route natural-language engineering questions to grounded XRL tools and then summarize their outputs in a way a process engineer can inspect.

References

Kim, H., Chen, H., Li, C., & Lee, J. M. (2026). TalkToAgent: A multi-agent LLM Framework for natural language explanation of reinforcement learning policies. Computers & Chemical Engineering, 109672.