Graph Retrieval-Augmented Visual Question Answering for Complex Engineering Drawings
Keywords:
engineering drawings analysis, blueprint interpretation, retrieval-augmented generation, multi-hop reasoning, traceable question answeringAbstract
Complex engineering drawings, such as piping and instrumentation diagrams (P&ID), substation wiring diagrams, and industrial system topology diagrams, contain dense symbols, text annotations, and multi-hop connectivity relations. Existing MLLMs can handle general visual question answering, but they still face two key limitations in engineering drawings: the high context burden of complete drawings or complete structural files, and the lack of verifiable topological evidence for final answers. This paper proposes Neuro-Symbolic Graph Retrieval-Augmented Generation (NSG-RAG), which reformulates engineering drawing question answering as topology-grounded visual question answering under the setting where PID2Graph provides GraphML topology annotations. NSG-RAG does not retrain an end-to-end drawing parser. Instead, it constructs a retrievable evidence space around GraphML. Given a question, the system performs descriptive node grounding, retrieves relevant k-hop subgraphs and multi-hop paths, and verifies whether candidate paths are supported by the underlying graph structure through NetworkX rules. A controlled PID-GraphQA evaluation built from 500 PID2Graph Complete drawings covers symbol recognition, direct connection, negative connection, multi-hop path, and unsupported path claim tasks. Experiments show that, when GraphML is available, graph-structured retrieval and rule verification improve answer accuracy, evidence traceability, and multi-hop topological reasoning. This study is positioned as a proof-of-concept investigation of GraphRAG for complex engineering drawings and provides a reproducible basis for future evaluation with real MLLM backbones and analysis of end-to-end parsing error propagation.
References
Alimin, A.A.; Schweidtmann, A.M. GraphRAG for Engineering Diagrams: ChatP&ID Enables LLM Interaction with P&IDs. arXiv preprint arXiv:2603.22528, 2026. https://doi.org/10.48550/arXiv.2603.22528
Antol, S.; Agrawal, A.; Lu, J.; et al. VQA: Visual Question Answering. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2015; pp. 2425–2433. https://doi.org/10.1109/ICCV.2015.279
Bai, S.; Chen, K.; Liu, X.; et al. Qwen2.5-VL Technical Report. arXiv preprint arXiv:2502.13923, 2025. https://doi.org/10.48550/arXiv.2502.13923
Chen, S.; Yuan, Z.; Zhang, Q.; Hua, W.; Cao, J.; Huang, X. NeuSymEA: Neuro-Symbolic Entity Alignment via Variational Inference. In Proceedings of the Conference on Neural Information Processing Systems (NeurIPS), 2025; Vol. 38. https://doi.org/10.48550/arXiv.2410.04153
Edge, D.; Trinh, H.; Cheng, N.; et al. From Local to Global: A Graph RAG Approach to Query-Focused Summarization. arXiv preprint arXiv:2404.16130, 2024. https://doi.org/10.48550/arXiv.2404.16130
Goyal, Y.; Khot, T.; Agrawal, A.; Summers-Stay, D.; Batra, D.; Parikh, D. Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering. International Journal of Computer Vision 2019, 127(4), 398–414. https://doi.org/10.1007/s11263-018-1116-0
Guo, Z.; Xia, L.; Yu, Y.; Ao, T.; Huang, C. LightRAG: Simple and Fast Retrieval-Augmented Generation. In Findings of the Association for Computational Linguistics: EMNLP 2025, 2025; pp. 10746–10761. https://doi.org/10.18653/v1/2025.findings-emnlp.568
Gupta, M.; Wei, C.; Czerniawski, T.; Eiris, R. PIDQA—Question Answering on Piping and Instrumentation Diagrams. Machine Learning and Knowledge Extraction 2025, 7(2), 39. https://doi.org/10.3390/make7020039
Hsiao, C.-H.; Wang, Y.-C.; Lin, T.-S.; Yeh, Y.-R.; Chen, C.-S. MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation. arXiv preprint arXiv:2512.20626, 2025. https://doi.org/10.48550/arXiv.2512.20626
Hudson, D.A.; Manning, C.D. GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019; pp. 6700–6709. https://doi.org/10.1109/CVPR.2019.00686
Johnson, J.; Hariharan, B.; van der Maaten, L.; Fei-Fei, L.; Zitnick, C.L.; Girshick, R. CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017; pp. 1988–1997. https://doi.org/10.1109/CVPR.2017.215
Karpukhin, V.; Oguz, B.; Min, S.; et al. Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2020; pp. 6769–6781. https://doi.org/10.18653/v1/2020.emnlp-main.550
Kim, H.; Lee, W.; Kim, M.; et al. Deep-Learning-Based Recognition of Symbols and Texts at an Industrially Applicable Level from Images of High-Density Piping and Instrumentation Diagrams. Expert Systems with Applications 2021, 183, 115337. https://doi.org/10.1016/j.eswa.2021.115337
Krishna, R.; Zhu, Y.; Groth, O.; et al. Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations. International Journal of Computer Vision 2017, 123(1), 32–73. https://doi.org/10.1007/s11263-016-0981-7
Lewis, P.; Perez, E.; Piktus, A.; et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Proceedings of the Conference on Neural Information Processing Systems (NeurIPS), 2020; Vol. 33, pp. 9459–9474. https://doi.org/10.48550/arXiv.2005.11401
Li, J.; Li, D.; Savarese, S.; Hoi, S.C.H. BLIP-2: Bootstrapping Language-Image Pre-Training with Frozen Image Encoders and Large Language Models. In Proceedings of the International Conference on Machine Learning (ICML), 2023; Vol. 202, pp. 19730–19742. https://doi.org/10.48550/arXiv.2301.12597
Liu, H.; Li, C.; Wu, Q.; Lee, Y.J. Visual Instruction Tuning. In Proceedings of the Conference on Neural Information Processing Systems (NeurIPS), 2023; Vol. 36, pp. 34892–34916. https://doi.org/10.48550/arXiv.2304.08485
Liu, N.F.; Lin, K.; Hewitt, J.; et al. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics 2024, 12, 157–173. https://doi.org/10.1162/tacl_a_00638
Moon, Y.; Lee, J.; Mun, D.; Lim, S. Deep Learning-Based Method to Recognize Line Objects and Flow Arrows from Image-Format Piping and Instrumentation Diagrams for Digitization. Applied Sciences 2021, 11(21), 10054. https://doi.org/10.3390/app112110054
OpenAI. GPT-4o System Card. arXiv preprint arXiv:2410.21276, 2024. https://doi.org/10.48550/arXiv.2410.21276
Qu, N.; Xue, J.; Gao, S. Automatic Recognition System for P&ID Diagrams in Distributed Control Systems. In Proceedings of the 2025 IEEE International Conference on Computational Intelligence and Smart Application Technology (CISAT), 2025; pp. 674–678. https://doi.org/10.1109/CISAT66811.2025.11181755
Radford, A.; Kim, J.W.; Hallacy, C.; et al. Learning Transferable Visual Models from Natural Language Supervision. In Proceedings of the International Conference on Machine Learning (ICML), 2021; Vol. 139, pp. 8748–8763. https://doi.org/10.48550/arXiv.2103.00020
Schmidt, W.J.; Rincon-Yanez, D.; Kharlamov, E.; Paschke, A. LLMs on the Rise: Neuro-Symbolic AI for Knowledge Graph Construction in Manufacturing: Systematic Literature Review. IEEE Access 2026, 14, 28383–28410. https://doi.org/10.1109/ACCESS.2026.3665999
Stürmer, J.M.; Graumann, M.; Koch, T. From Engineering Diagrams to Graphs: Digitizing P&IDs with Transformers. In Proceedings of the 2025 IEEE 12th International Conference on Data Science and Advanced Analytics (DSAA), 2025; pp. 1–11. https://doi.org/10.1109/DSAA65442.2025.11248012
Stürmer, J.M.; Graumann, M.; Koch, T. PID2Graph. Zenodo, 2025. https://doi.org/10.5281/zenodo.14803338
Tran, D.T.; Tran, T.-K.; Hauswirth, M.; Le Phuoc, D. ReasonVQA: A Multi-Hop Reasoning Benchmark with Structural Knowledge for Visual Question Answering. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025; pp. 18793–18803. https://doi.org/10.1109/ICCV51701.2025.01746
Yu, E.-S.; Cha, J.-M.; Lee, T.; Kim, J.; Mun, D. Features Recognition from Piping and Instrumentation Diagrams in Image Format Using a Deep Learning Network. Energies 2019, 12(23), 4425. https://doi.org/10.3390/en12234425
Yu, J.; Liu, Y.; Gu, J.; Torr, P.; Zhou, D. Can Knowledge-Graph-Based Retrieval Augmented Generation Really Retrieve What You Need? In Proceedings of the Conference on Neural Information Processing Systems (NeurIPS), 2025. https://doi.org/10.48550/arXiv.2510.16582
Yuan, X.; Ning, L.; Ye, Q.; Fan, W.; Li, Q. mKG-RAG: Leveraging Multimodal Knowledge Graphs in Retrieval-Augmented Generation for Knowledge-Intensive VQA. In Proceedings of the ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR), 2026. https://doi.org/10.1145/3805712.3809680
Zhu, J.; Wang, W.; Chen, Z.; et al. InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models. arXiv preprint arXiv:2504.10479, 2025. https://doi.org/10.48550/arXiv.2504.10479
