VLDB 2026 Research / reviewers in the wild / expert
Hao Huang 0014
dblp:04/5616-14
· DBLP profile ↗
7ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0003-3117-0881ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 6 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 5 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Causality with Knowledge Graphs: Semantics and Inference
Hao Huang 0014 |
WSDM | 1 |
| 2026 | Joint Graph Learning for Robust Causal Inference over Knowledge GraphsabstractCausal inference is critical for understanding cause-effect relationships in real-world domains. However, applying it over knowledge graphs (KGs) poses unique challenges due to two key issues: missing attributes caused by the Open-World Assumption and interference effects arising from complex relational dependencies among entities. Existing methods often assume fully observed data or fail to model inter-unit dependencies, leading to biased or unreliable effect estimates. We introduce BaLu, a joint graph learning framework that addresses both challenges through an end-to-end solution. BaLu reformulates the causal inference over KGs as two interconnected tasks: (1) attribute imputation as edge prediction between units (entities) and their attributes, and (2) treatment effect estimation as node prediction that accounts for interference through representation learning. BaLu employs Graph Neural Networks (GNNs) to capture attribute similarity and relational structure, enabling both accurate imputation and interference-aware message passing. Experiments on four benchmark datasets show that BaLu consistently outperforms state-of-the-art baselines—even when enhanced with strong imputation techniques—demonstrating robust performance in incomplete and relationally complex KGs. These results demonstrate that BaLu offers a principled and practical solution for robust causal inference in knowledge-driven domains, empowering data-driven decision-making under real-world conditions of incompleteness and relational complexity. Hao Huang 0014, Maria-Esther Vidal |
WSDM | 1 |
| 2025 | HyKG-CF: A Hybrid Approach for Counterfactual Prediction using Domain Knowledge
Hao Huang 0014, Maria-Esther Vidal |
WSDM | 1 |
| 2025 | Integrating Knowledge Graphs with Symbolic AI: The Path to Interpretable Hybrid AI Systems in MedicineabstractKnowledge Graphs (KGs) are graph-based structures that integrate heterogeneous data, capture domain knowledge, and enable explainable AI through symbolic reasoning. This position paper examines the challenges and research opportunities in integrating KGs with neuro-symbolic AI, highlighting their potential to enhance explainability, scalability, and context-aware reasoning in hybrid AI systems. Using a lung cancer use case, we illustrate how hybrid approaches address tasks such as link prediction—uncovering hidden relationships in medical data—and counterfactual reasoning—analyzing alternative scenarios to understand causal factors. The discussion is framed around TrustKG, which demonstrates how constraint validation, causal reasoning, and user-centric communication can support transparent and reliable decision-making. Additionally, we identify current limitations of KGs, including gaps in knowledge coverage, evolving data integration challenges, and the need for improved usability and impact assessment. These insights are not limited to healthcare but extend to other domains like energy, manufacturing, and mobility, showcasing the broad applicability of KGs. Finally, we propose research directions to unlock their full potential in building robust, transparent, and widely adopted real-world applications. Maria-Esther Vidal, Yashrajsinh Chudasama, Hao Huang 0014, Disha Purohit, Maria Torrente |
J. Web Semant. | 3 |
| 2024 | SemMatch: Semantics-Aware Matching for Causal Inference over Knowledge Graphs
Hao Huang 0014, Maria-Esther Vidal |
WISE (2) | 1 |
| 2022 | Causal Relationship over Knowledge GraphsabstractCausality has been discussed for centuries, and the theory of causal inference over tabular data has been broadly studied and utilized in multiple disciplines. However, only a few works attempt to infer the causality while exploiting the meaning of the data represented in a data structure like knowledge graph. These works offer a glance at the possibilities of causal inference over knowledge graphs, but do not yet consider the metadata, e.g., cardinalities, class subsumption and overlap, and integrity constraints. We propose CareKG, a new formalism to express causal relationships among concepts, i.e., classes and relations, and enable causal queries over knowledge graphs using semantics of metadata. We empirically evaluate the expressiveness of CareKG in a synthetic knowledge graph concerning cardinalities, class subsumption and overlap, integrity constraints. Our initial results indicate that CareKG can represent and measure causal relations with some semantics which are uncovered by state-of-the-art approaches. Hao Huang 0014 |
CIKM | 1 |
| 2021 | Core-Concept-Seeded LDA for Ontology LearningabstractOntologies are powerful semantic models applied for various purposes such as improving system interoperability, information retrieval, question answering, etc. However, building domain ontologies remains a challenging task for humans, especially when the concepts and properties are large or evolving, and also when they are built from large-scale textual data. Machine learning allows to automate the building of ontologies from texts. In particular, clustering techniques have a promising ability on the concept formation task by identifying the cluster of semantically closed terms as a concept. However, current works encounter issues in learning relevant domain-specific clusters or in identifying the relevant concept labels for each cluster. To solve these issues, we propose both to use core concepts from a domain ontology as prior knowledge, and to adapt term clustering with seed knowledge-based LDA models in order to take these core concepts into account. First, each topic is associated with a set of seed terms of a single core concept, then the learning is guided by these seeds to gather in the same topic the terms that refer to its core concept. We evaluate our proposal on two textual corpora and compare it to the baselines (LDA, K-means, and SMBM). The results show that our approach performs significantly better than other methods on the class-balanced dataset and works well on the class-imbalanced dataset with a proper number of topics for each core concept. Hao Huang 0014, Mounira Harzallah, Fabrice Guillet |
KES | 1 |