LLM & Probabilistic Approaches

GraphRAG for Engineering Diagrams: ChatP&ID and P&ID Retrieval

A critical note on ChatP&ID: P&IDs are better treated as structured engineering knowledge graphs than as raw images or raw XML, but the benchmark mainly validates context engineering rather than a new GraphRAG algorithm.

The useful question in this paper is not “Can an LLM read a P&ID?” A multimodal model can already say plausible things about an engineering drawing. The sharper question is this:

What should the LLM be allowed to see?

For Piping and Instrumentation Diagrams, the answer is rarely the raw image. It is also not necessarily the raw smart-P&ID XML. A P&ID is closer to a structured engineering graph: equipment, pipes, valves, instruments, controllers, actuators, line numbers, tags, and process-connectivity relations. ChatP&ID builds on that view. It converts DEXPI smart P&IDs into Neo4j knowledge graphs, creates several abstraction levels, and lets an LLM agent call graph-retrieval tools depending on the question.

That is why the paper fits naturally under LLM & Probabilistic Approaches, even though the object is an engineering diagram. The contribution is not a new language model, a new GNN, or a theorem about graph retrieval. It is a system-level argument about representation: if the input context is cleaner, shorter, and topologically meaningful, then a smaller or cheaper LLM has a better chance of answering process-engineering questions correctly.

The Problem

P&IDs are central documents in plant design, operation, maintenance, management of change, HAZOP, and safety review. They encode which equipment is connected to which line, which valve sits on a path, which instrument measures a variable, and which controller manipulates which actuator.

The difficulty is that industrial P&IDs often exist as PDFs, images, CAD exports, or smart-P&ID files whose structure was designed for engineering software rather than for language-model reasoning. A human can zoom into a pump tag, follow a downstream line, inspect valve symbols, and connect a temperature indicator to a control loop. That process is slow and interpretation-dependent. It scales poorly when many pages, off-page connectors, revisions, and safety questions are involved.

There are three obvious ways to give such information to an LLM.

First, send the image to a multimodal LLM. This has low setup cost, but P&IDs are not ordinary pictures. Small labels, tag numbers, line specifications, instrument bubbles, arrows, and dense symbols are easy to miss or compress away. A model can produce a fluent answer while losing a set pressure, a pump specification, or a path segment.

Second, send the raw DEXPI/XML representation. This gives the model explicit machine-readable information, but the raw file also contains internal IDs, URIs, class hierarchy details, layout information, and other metadata that is not directly useful for process reasoning. The problem becomes “useful engineering data plus semantic noise.” The paper notes that even a single smart-P&ID input can exceed 150,000 tokens.

Third, convert the P&ID into a graph that keeps engineering semantics and topology while removing unnecessary representation noise. This is the path taken by ChatP&ID.

What ChatP&ID Builds

The system starts from DEXPI-compatible smart P&IDs and uses pyDEXPI to transform them into a flowsheet knowledge graph stored in Neo4j. The graph represents objects such as pumps, tanks, heat exchangers, valves, instruments, controllers, and piping components as nodes. Relations such as composition, connection, control, manipulation, and signal flow become edges. Node properties hold tags, design pressures, temperatures, nominal diameters, materials, fail-safe positions, and related attributes when available.

The paper then uses several graph abstraction levels.

Graph level Meaning
Complete graph Keeps pyDEXPI objects nearly one-to-one.
Process graph Compresses lower-level piping composition.
Conceptual graph Keeps mainly equipment, major lines, instruments, and control relations.

This abstraction is not just for visualization. It is context engineering. The complete graph contains more information but also more noise and token cost. The conceptual graph loses detail but may expose the process-level structure more clearly. For many LLM questions, a smaller graph with the right topology is better than a complete graph with many irrelevant fields.

Four Retrieval Modes

The paper compares four GraphRAG tools.

ContextRAG is the simplest. It cleans the graph representation, removes unnecessary metadata, and passes a graph context to the LLM. In topology mode it mainly sends node types and connectivity. In graph mode it also sends attributes such as tags, pressures, temperatures, and specifications. This is closer to semantic compression than to selective retrieval: the whole cleaned graph, or a large part of it, becomes the context.

VectorRAG creates semantic descriptions for nodes, embeds them, embeds the user query, and retrieves the top-k relevant nodes by similarity. Its advantage is context size. The LLM receives only a subset of nodes instead of the whole graph. Its weakness is also clear: semantic closeness is not the same as process relevance. A question about isolation, bypasses, or downstream tracing may require topology that is not captured by a top-k node list.

PathRAG is the most process-engineering-shaped idea in the paper. Engineers do not only search for an equipment item; they follow the line, inspect upstream and downstream neighbors, check valves and instruments, and reason over the resulting path. PathRAG begins from relevant nodes and traverses neighboring graph structure. This is the right direction for questions such as which valves isolate a tank or how a process-stream temperature is controlled.

There is one caveat. The algorithm description must make the neighbor restriction explicit. If the next-hop search is actually an unrestricted vector search at every step, then the method is less a path traversal and more a repeated global retrieval loop. The implementation may still constrain the search locally, but the paper’s pseudocode needs to be read carefully on this point.

CypherRAG asks the LLM to translate natural-language questions into Cypher queries and executes them against Neo4j. This is attractive for structured questions such as listing valves and fail-safe positions. The database can reject malformed syntax. But valid Cypher is not the same as correct engineering interpretation. A query can be executable and still retrieve the wrong subgraph if the LLM misunderstands “upstream isolation valve” or “control loop completeness.”

What Is Guaranteed

The paper is mostly an implementation and benchmark paper, not a formal-methods paper. Its guarantees are operational rather than semantic.

Component What is guaranteed What is not guaranteed
Complete graph Intended mapping from pyDEXPI objects to graph nodes Correctness of the original P&ID
VectorRAG Cosine ranking under the chosen embeddings Factual relevance of the retrieved nodes
PathRAG Bounded search by depth or breadth Inclusion of the correct engineering path
Agent loop Termination through a tool-call limit Correct tool selection
CypherRAG Possible rejection of malformed queries Alignment between valid query and question intent
ContextRAG Context reduction through metadata filtering No information loss after abstraction

This distinction matters. GraphRAG can reduce hallucination risk by giving the model a better-grounded context. It does not prove answer correctness. In engineering settings, that difference is not small. A grounded wrong answer is still wrong.

Why the Approach Works

The main performance mechanism is not mysterious.

First, graph abstraction removes semantic noise. Raw XML contains many tokens that are not useful for answering a process question. Removing those fields improves the ratio between engineering-relevant information and total input tokens.

Second, topology is explicit. In an image-based setting, the vision model must infer lines, arrows, symbols, and connectivity. In a graph setting, connections such as pump to heat exchanger to tank are represented directly.

Third, retrieval can be matched to question type. Attribute questions may work with VectorRAG or CypherRAG. Whole-diagram summaries may work with ContextRAG. Isolation, routing, and control-loop questions need path-aware retrieval. A single document-RAG pattern is not enough because P&ID questions mix attributes, topology, and engineering inference.

Benchmark Reading

The benchmark uses 19 QA pairs across graph queries, path exploration, knowledge inference, and graph summarization. For GPT-5-mini, the reported Table 5 values are:

Method Accuracy Cost/query Time/query
ContextRAG 0.91 $0.0044 24.33 s
VectorRAG 0.82 $0.0023 24.42 s
PathRAG 0.83 $0.0021 54.64 s
CypherRAG 0.86 $0.0016 39.42 s
Multimodal image 0.83 $0.0018 45.55 s
Raw Proteus XML 0.88 $0.0342 52.05 s

ContextRAG has the highest observed accuracy in this table. It is not the lowest absolute cost method. CypherRAG, VectorRAG, PathRAG, and image input are all cheaper per query in the reported numbers. The more defensible conclusion is narrower:

ContextRAG gives the highest observed accuracy and a strong cost-performance tradeoff compared with raw smart-P&ID ingestion.

The larger claim is still useful. When the paper compares conceptual graph context with raw Proteus XML, the graph representation raises accuracy while reducing token cost substantially. That supports the representation argument: the LLM does not need more raw input; it needs better-shaped input.

Evaluation Limits

The evaluation is promising but small. Nineteen QA pairs cannot fully cover industrial P&ID work. The benchmark is weighted toward graph queries, while only one example targets graph summarization. Harder plant questions include bypass-aware isolation, interlock and trip logic, multi-page off-page connector tracing, failure-position reasoning, revision mismatch detection, and inconsistency between a P&ID and a control narrative.

The use of LLM-as-a-judge is also mixed. Semantic similarity and LLM judging are useful for scaling evaluation, but P&ID answers often hinge on exact values, exact tags, and exact path membership. A verbose answer can score well while containing a wrong set pressure or an extra valve. The paper recognizes this issue; it should be treated as a measurement limitation, not a minor detail.

Another methodological weakness is that graph representation and LLM-generated semantic enrichment are not fully separated. VectorRAG and PathRAG embed node descriptions generated by GPT-4o. The performance gain may come from the graph, from the GPT-generated descriptions, from the embedding model, or from their combination. A stronger ablation would compare raw graph attributes, template-generated descriptions, and GPT-generated descriptions.

Novelty

The novelty is not a completely new GraphRAG algorithm. It is the combination:

DEXPI, pyDEXPI, LPG/Neo4j, multi-level graph abstraction, GraphRAG tool benchmarking, and a P&ID chat interface.

The best idea is the framing: P&IDs should be handled as structured engineering knowledge graphs, not as images with labels or raw XML dumps. That framing naturally points toward consistency checking, rule-based auto-correction, revision comparison, control-loop completeness checking, isolation-path generation, HAZOP evidence retrieval, topology-aware process synthesis support, and operator-facing decision support.

The agent part is more limited. Figure 3 is closer to a single ReAct-style LLM agent that calls several retrieval tools than to a multi-agent system where specialized agents coordinate. Multi-agent orchestration is a natural research direction, but it is not the central benchmark object in this paper.

Final Assessment

This paper asks one of the most important questions for LLM systems in engineering diagrams:

What should the model see?

The answer is convincing. Not a raw image. Not raw XML. A graph context in which engineering semantics and topology have already been cleaned and organized.

That direction matters most when the model is small, cost is constrained, or deployment must be auditable. ChatP&ID does not solve P&ID reasoning in full. It shows that the representation layer is not a preprocessing detail. It is where much of the engineering intelligence enters the LLM system.

References

Alimin, A. A.; Schweidtmann, A. M. “GraphRAG for Engineering Diagrams: ChatP&ID Enables LLM Interaction with P&IDs.” arXiv:2603.22528, 2026. https://arxiv.org/abs/2603.22528