VLDB 2026 Research / reviewers in the wild / expert
Pan Hu 0001
dblp:96/3740-1
· DBLP profile ↗
20ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0003-1701-9640ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Incremental Maintenance of DatalogMTL MaterialisationsabstractDatalogMTL extends the classical Datalog language with metric temporal logic (MTL), enabling expressive reasoning over temporal data. While existing reasoning approaches, such as materialisation-based and automata-based methods, offer soundness and completeness, they lack support for handling efficient dynamic updates—a crucial requirement for real-world applications that involve frequent data updates. In this work, we propose DRedMTL, an incremental reasoning algorithm for DatalogMTL with bounded intervals. Our algorithm builds upon the classical Delete/Rederive (DRed) algorithm, which incrementally updates the materialisation of a Datalog program. Unlike a Datalog materialisation which is in essence a finite set of facts, a DatalogMTL materialisation has to be represented as a finite set of facts plus periodic intervals indicating how the full materialisation can be constructed through unfolding. To cope with this, our algorithm is equipped with specifically designed operators to efficiently handle such periodic representations of DatalogMTL materialisations. We have implemented this approach and tested it on several publicly available datasets. Experimental results show that DRedMTL often significantly outperforms rematerialisation, sometimes by orders of magnitude. Kaiyue Zhao, Dingqi Chen, Pan Hu 0001 |
AAAI | 4 |
| 2025 | Goal-Driven Reasoning in DatalogMTL with Magic SetsabstractDatalogMTL is a powerful rule-based language for temporal reasoning. Due to its high expressive power and flexible modeling capabilities, it is suitable for a wide range of applications, including tasks from industrial and financial sectors. However, due its high computational complexity, practical reasoning in DatalogMTL is highly challenging. To address this difficulty, we introduce a new reasoning method for DatalogMTL which exploits the magic sets technique—a rewriting approach developed for (non-temporal) Datalog to simulate top-down evaluation with bottom-up reasoning. We have implemented this approach and evaluated it on publicly available benchmarks, showing that the proposed approach significantly and consistently outperformed state-of-the-art reasoning techniques. Kaiyue Zhao, Dongliang Wei, Przemyslaw Andrzej Walega, Dingmin Wang, Hongming Cai 0001, Pan Hu 0001 |
AAAI | 7 |
| 2025 | Knowledge Graph-Based Process Planning for CAD ModelsabstractIn response to the demand for automated analysis and optimized reasoning of part model data faced in the process of ship research, design and manufacturing, this paper combines the knowledge graph with the model design rule base and manufacturing process library to construct a model process generation framework based on the knowledge graph. The paper first standardizes heterogeneous models for each type of CAD model in research design to construct a component hierarchy diagrams for multi-level assembly of models. Then the decomposed models are subjected to feature extraction and association. Then the process analysis based on model features is carried out with the help of knowledge graph. Finally, the automated generation of process plans is realized through the reasoning and consistency verification of process combinations of single components. The project explores the common components in ship research, design and manufacturing, and the case application results show that the project system has high accuracy and good adaptability, and has practical value and reference significance for the interoperability of design and manufacturing process data and the realization of automated process planning. Chenze Li, Pan Hu 0001, Hongming Cai 0001 |
SoMeT | 5 |
| 2025 | DR-RAG: Domain-Rule-based Retrieval-Augmented Generation for aviation digital model design
Xirui Xiong, Hongming Cai 0001, Han Yu 0005, Bingqing Shen, Pan Hu 0001 |
Adv. Eng. Informatics | 5 |
| 2025 | Practical Reasoning in DatalogMTLabstractAbstract DatalogMTL is an extension of Datalog with metric temporal operators that has found an increasing number of applications in recent years. Reasoning in DatalogMTL is, however, of high computational complexity, which makes reasoning in modern data-intensive applications challenging. In this paper we present a practical reasoning algorithm for the full DatalogMTL language, which we have implemented in a system called MeTeoR. Our approach effectively combines an optimised (but generally non-terminating) materialisation (a.k.a. forward chaining) procedure, which provides scalable behaviour, with an automata-based component that guarantees termination and completeness. To ensure favourable scalability of the materialisation component, we propose a novel seminaïve materialisation procedure for DatalogMTL enjoying the non-repetition property, which ensures that each rule instance will be applied at most once throughout its entire execution. Moreover, our materialisation procedure is enhanced with additional optimisations which further reduce the number of redundant computations performed during materialisation by disregarding rules as soon as it is certain that they cannot derive new facts in subsequent materialisation steps. Our extensive evaluation supports the practicality of our approach. Dingmin Wang, Bernardo Cuenca Grau, Przemyslaw Andrzej Walega, Pan Hu 0001 |
Theory Pract. Log. Program. | 4 |
| 2024 | Optimised Storage for Datalog ReasoningabstractMaterialisation facilitates Datalog reasoning by precomputing all consequences of the facts and the rules so that queries can be directly answered over the materialised facts. However, storing all materialised facts may be infeasible in practice, especially when the rules are complex and the given set of facts is large. We observe that for certain combinations of rules, there exist data structures that compactly represent the reasoning result and can be efficiently queried when necessary. In this paper, we present a general framework that allows for the integration of such optimised storage schemes with standard materialisation algorithms. Moreover, we devise optimised storage schemes targeting at transitive rules and union rules, two types of (combination of) rules that commonly occur in practice. Our experimental evaluation shows that our approach significantly improves memory consumption, sometimes by orders of magnitude, while remaining competitive in terms of query answering time. Pan Hu 0001, Yavor Nenov, Ian Horrocks 0001 |
AAAI | 2 |
| 2024 | Parallel Collaborative Reasoning Approaches Based on DatalogMTL in IoT ScenariosabstractAn important task in IoT application scenarios is to perform synergy reasoning on the phenomenal data and events of strong temporal semantics and complex correlations characteristics, with the help of associated knowledge and rules. However, current reasoning methods suffer from difficult rule representation with poor readability, lack of temporal semantics, and high reasoning complexity and inefficiency. To address these problems, this paper proposes parallel collaborative reasoning approaches based on DatalogMTL. Firstly, a series of collaborative access control mechanisms are designed for the concurrent conflict problems. Then, the rule-level parallel and fact-level parallel reasoning methods are presented based on materialization algorithm respectively. In this paper, we take experiments on two relative datasets and verify that our approaches greatly improve the reasoning efficiency and have good scalability in IoT scenarios. Pan Hu 0001, Hongming Cai 0001, Lihong Jiang |
CSCWD | 2 |
| 2024 | CGCI: Cross-granularity Causal Inference framework for engineering Change Propagation Analysis
Yuxiao Wang 0004, Hongming Cai 0001, Bingqing Shen, Pan Hu 0001, Han Yu 0005, Lihong Jiang |
Adv. Eng. Informatics | 4 |
| 2024 | A Cloud-Edge Collaboration Framework for Generating Process Digital TwinabstractTracking the process of remote task execution is critical to timely process analysis by collecting the evidence of correct execution or failure, which generates a process digital twin (DT) for remote supervision. Generally, it will encounter the challenge of constrained communication, high overhead, and high traceability demand, leading to the efficient remote process tracking issue. Existing approaches can address the issue by monitoring or simulating remote task execution. Nevertheless, they do not provide a cost-effective solution, especially when unexpected situation occurs. Thus, we proposed a new cloud-edge collaboration framework for process DT generation. It addresses the efficient remote process tracking issue with a real-virtual collaborative process tracking (RVCPT) approach. The approach contains three patterns of real-virtual collaboration for tracking the entire process of task execution with a coevolution pattern, identifying unexpected situations with a discrimination pattern, and generating a process DT with a real-virtual fusion pattern. This approach can minimize tracking overhead, and meanwhile maintains high traceability, which maximizes the overall cost-effectiveness. With prototype development, case study and experimental evaluation show the applicability and performance advantage of the new cloud-edge collaboration framework in remote supervision. Bingqing Shen, Han Yu 0005, Pan Hu 0001, Hongming Cai 0001, Jingzhi Guo, Boyi Xu, Lihong Jiang |
IEEE Trans. Cloud Comput. | 3 |
| 2024 | Accurate Sampling-Based Cardinality Estimation for Complex Graph QueriesabstractAccurately estimating the cardinality (i.e., the number of answers) of complex queries plays a central role in database systems. This problem is particularly difficult in graph databases, where queries often involve a large number of joins and self-joins. Recently, Park et al. [ 55 ] surveyed seven state-of-the-art cardinality estimation approaches for graph queries. The results of their extensive empirical evaluation show that a sampling method based on theWanderJoinonline aggregation algorithm [ 47 ] consistently offers superior accuracy. We extended the framework by Park et al. [ 55 ] with three additional datasets and repeated their experiments. Our results showed that WanderJoin is indeed very accurate, but it can often take a large number of samples and thus be very slow. Moreover, when queries are complex and data distributions are skewed, it often fails to find valid samples and estimates the cardinality as zero. Finally, complex graph queries often go beyond simple graph matching and involve arbitrary nesting of relational operators such as disjunction, difference, and duplicate elimination. Neither of the methods considered by Park et al. [ 55 ] is applicable to such queries. In this article, we present a novel approach for estimating the cardinality of complex graph queries. Our approach is inspired by WanderJoin, but, unlike all approaches known to us, it can process complex queries with arbitrary operator nesting. Our estimator is strongly consistent, meaning that the average of repeated estimates converges with probability one to the actual cardinality. We present optimisations of the basic algorithm that aim to reduce the chance of producing zero estimates and improve accuracy. We show empirically that our approach is both accurate and quick on complex queries and large datasets. Finally, we discuss how to integrate our approach into a simple dynamic programming query planner, and we confirm empirically that our planner produces high-quality plans that can significantly reduce end-to-end query evaluation times. Pan Hu 0001, Boris Motik |
ACM Trans. Database Syst. | 1 |
| 2023 | Intelligent Manufacturing Collaboration Platform for 3D Curved Plates Based on Graph MatchingabstractThe three-dimensional (3D) curved plate manufacturing is performed by constructing surfaces corresponding to the shape of the curved plate for multi-point forming. However, in the manufacturing process, the rebound restricts the forming accuracy, and the currently adopted rebound control methods cannot predict the rebound amount accurately. Meanwhile, the process involves multi-role collaboration and multiple data conversions and comparisons. These problems lead to a high degree of manual dependence, which affects manufacturing efficiency and accuracy. To address the above problems, this paper proposes a collaborative platform for the intelligent manufacturing of curved plates based on graph matching. Firstly, this paper establishes information models covering the whole process of curved plate manufacturing and forms a unified topology graph model. Then, the intelligent generation method of processing parameters based on graph matching is proposed, which realizes similar case recommendation and case-based processing parameters generation. Finally, we design and develop a collaboration platform based on micro-service architecture to support efficient collaboration among various departments and roles. In this paper, we use sail-shaped curved plates as a case of processing parameters generation and verify that this intelligent method can improve the accuracy of rebound control by comparison with related work, which shows that our method can be effectively applied to curved plate manufacturing. Yanjun Dong, Haoyuan Hu, Pan Hu 0001, Lihong Jiang, Hongming Cai 0001 |
CSCWD | 4 |
| 2023 | Enhancing Datalog Reasoning with Hypertree DecompositionsabstractDatalog reasoning based on the seminaive evaluation strategy evaluates rules using traditional join plans, which often leads to redundancy and inefficiency in practice, especially when the rules are complex. Hypertree decompositions help identify efficient query plans and reduce similar redundancy in query answering. However, it is unclear how this can be applied to materialisation and incremental reasoning with recursive Datalog programs. Moreover, hypertree decompositions require additional data structures and thus introduce nonnegligible overhead in both runtime and memory consumption. In this paper, we provide algorithms that exploit hypertree decompositions for the materialisation and incremental evaluation of Datalog programs. Furthermore, we combine this approach with standard Datalog reasoning algorithms in a modular fashion so that the overhead caused by the decompositions is reduced. Our empirical evaluation shows that, when the program contains complex rules, the combined approach is usually significantly faster than the baseline approach, sometimes by orders of magnitude. Pan Hu 0001, Yavor Nenov, Ian Horrocks 0001 |
IJCAI | 2 |
| 2022 | MeTeoR: Practical Reasoning in Datalog with Metric Temporal OperatorsabstractDatalogMTL is an extension of Datalog with operators from metric temporal logic which has received significant attention in recent years. It is a highly expressive knowledge representation language that is well-suited for applications in temporal ontology-based query answering and stream processing. Reasoning in DatalogMTL is, however, of high computational complexity, making implementation challenging and hindering its adoption in applications. In this paper, we present a novel approach for practical reasoning in DatalogMTL which combines materialisation (a.k.a. forward chaining) with automata-based techniques. We have implemented this approach in a reasoner called MeTeoR and evaluated its performance using a temporal extension of the Lehigh University Benchmark and a benchmark based on real-world meteorological data. Our experiments show that MeTeoR is a scalable system which enables reasoning over complex temporal rules and datasets involving tens of millions of temporal facts. Dingmin Wang, Pan Hu 0001, Przemyslaw Andrzej Walega, Bernardo Cuenca Grau |
AAAI | 2 |
| 2022 | Parallel Construction of Knowledge Graphs from Relational Databases
Jingsheng Yan, Pan Hu 0001, Hongming Cai 0001, Lihong Jiang |
PRICAI (1) | 4 |
| 2022 | Modular materialisation of Datalog programsabstractAnswering queries over large datasets extended with Datalog rules plays a key role in numerous data management applications, and it has been implemented in several highly optimised Datalog systems in both academic and commercial contexts. Many systems implement reasoning via materialisation, which involves precomputing all consequences of the rules and the dataset in a preprocessing step. Some systems also use incremental reasoning algorithms, which can update the materialisation efficiently when the input dataset changes. Such techniques allow queries to be processed without any reference to the rules, so they are often used in applications where the performance of query answering is critical. Existing materialisation and incremental reasoning techniques enumerate all possible ways to apply rules to the data in order to derive all relevant consequences. This, however, can be inefficient because derivations of rules commonly used in practice are redundant; for example, rules axiomatising a binary predicate as symmetric and transitive can have a cubic number of applications, yet they can derive at most a quadratic number of facts. Such redundancy can be a significant source of overhead in practice and can prevent Datalog systems from successfully processing large datasets. To address this issue, in this paper we present a novel framework for modular materialisation and incremental reasoning. Our key idea is that, for certain combinations of rules commonly used in practice, all consequences can be derived using specialised procedures that do not necessarily enumerate all possible rule applications. Thus, our framework supports materialisation and incremental reasoning via a collection of modules. Each module is responsible for deriving consequences of a subset of the program, by using either standard rule application or proprietary algorithms. We prove that such an approach is complete as long as each module satisfies certain properties. Our formalisation of a module is very general, and in fact it allows modules to keep arbitrary auxiliary information. We also show how to realise custom procedures for four types of modules: transitivity, symmetry–transitivity, chain rules, and sequencing elements of a total order. Finally, we demonstrate empirically that using our custom procedures can speed up materialisation and incremental reasoning by several orders of magnitude on several well-known benchmarks. Thus, our technique has the potential to significantly improve the scalability of Datalog reasoners. Pan Hu 0001, Boris Motik, Ian Horrocks 0001 |
Artif. Intell. | 1 |
| 2021 | OWL2Vec*: embedding of OWL ontologiesabstractAbstract Semantic embedding of knowledge graphs has been widely studied and used for prediction and statistical analysis tasks across various domains such as Natural Language Processing and the Semantic Web. However, less attention has been paid to developing robust methods for embedding OWL (Web Ontology Language) ontologies, which contain richer semantic information than plain knowledge graphs, and have been widely adopted in domains such as bioinformatics. In this paper, we propose a random walk and word embedding based ontology embedding method named , which encodes the semantics of an OWL ontology by taking into account its graph structure, lexical information and logical constructors. Our empirical evaluation with three real world datasets suggests that benefits from these three different aspects of an ontology in class membership prediction and class subsumption prediction tasks. Furthermore, often significantly outperforms the state-of-the-art methods in our experiments. Jiaoyan Chen 0001, Pan Hu 0001, Ernesto Jiménez-Ruiz, Ole Magnus Holter, Denvar Antonyrajah, Ian Horrocks 0001 |
Mach. Learn. | 2 |
| 2019 | Modular Materialisation of Datalog ProgramsabstractThe seminaïve algorithm can be used to materialise all consequences of a datalog program, and it also forms the basis for algorithms that incrementally update a materialisation as the input facts change. Certain (combinations of) rules, however, can be handled much more efficiently using custom algorithms. To integrate such algorithms into a general reasoning approach that can handle arbitrary rules, we propose a modular framework for computing and maintaining a materialisation. We split a datalog program into modules that can be handled using specialised algorithms, and we handle the remaining rules using the semina¨ıve algorithm. We also present two algorithms for computing the transitive and the symmetric– transitive closure of a relation that can be used within our framework. Finally, we show empirically that our framework can handle arbitrary datalog programs while outperforming existing approaches, often by orders of magnitude. Pan Hu 0001, Boris Motik, Ian Horrocks 0001 |
AAAI | 1 |
| 2019 | Datalog Reasoning over Compressed RDF Knowledge BasesabstractMaterialisation is often used in RDF systems as a preprocessing step to derive all facts implied by given RDF triples and rules. Although widely used, materialisation considers all possible rule applications and can use a lot of memory for storing the derived facts, which can hinder performance. We present a novel materialisation technique that compresses the RDF triples so that the rules can sometimes be applied to multiple facts at once, and the derived facts can be represented using structure sharing. Our technique can thus require less space, as well as skip certain rule applications. Our experiments show that our technique can be very effective: when the rules are relatively simple, our system is both faster and requires less memory than prominent state-of-the-art RDF systems. Pan Hu 0001, Jacopo Urbani, Boris Motik, Ian Horrocks 0001 |
CIKM | 1 |
| 2018 | Optimised Maintenance of Datalog MaterialisationsabstractTo efficiently answer queries, datalog systems often materialise all consequences of a datalog program, so the materialisation must be updated whenever the input facts change. Several solutions to the materialisation update problem have been proposed. The Delete/Rederive (DRed) and the Backward/Forward (B/F) algorithms solve this problem for general datalog, but both contain steps that evaluate rules "backwards" by matching their heads to a fact and evaluating the partially instantiated rule bodies as queries. We show that this can be a considerable source of overhead even on very small updates. In contrast, the Counting algorithm does not evaluate the rules "backwards," but it can handle only nonrecursive rules. We present two hybrid approaches that combine DRed and B/F with Counting so as to reduce or even eliminate "backward" rule evaluation while still handling arbitrary datalog programs. We show empirically that our hybrid algorithms are usually significantly faster than existing approaches, sometimes by orders of magnitude. Pan Hu 0001, Boris Motik, Ian Horrocks 0001 |
AAAI | 1 |
| 2015 | SLOREV: Using Classical CAD Techniques for 3D Object Extraction from Single Photo
Pan Hu 0001, Hongming Cai 0001, Fenglin Bu |
MMM (2) | 1 |