VLDB 2026 Research / reviewers in the wild / expert
Guang-Yuan Hao
dblp:222/7953
· DBLP profile ↗
9ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0003-4740-6254ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Probabilistic and Bayesian machine learning · 34% Transfer learning and domain adaptation · 20% Knowledge representation and reasoning · 11% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 22 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Transfer learning and domain adaptation
domain adaptation |
1.3 | 2 | 2023 | Taxonomy-Structured Domain Adaptation · ICML 2023 Domain-Indexing Variational Bayes: Interpretable Domain Index for Domain Adaptation · ICLR 2023 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference |
1.0 | 1 | 2026 | BayesAgent: Bayesian Agentic Reasoning Under Uncertainty via Verbalized Probabilistic Graphical Modeling · AAAI 2026 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
1.0 | 1 | 2026 | BayesAgent: Bayesian Agentic Reasoning Under Uncertainty via Verbalized Probabilistic Graphical Modeling · AAAI 2026 |
Natural language and speech › Language models and text generation
LLM agents |
1.0 | 1 | 2026 | BayesAgent: Bayesian Agentic Reasoning Under Uncertainty via Verbalized Probabilistic Graphical Modeling · AAAI 2026 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal discovery |
0.9 | 1 | 2025 | A Conditional Independence Test in the Presence of Discretization · ICLR 2025 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › conditional independence
conditional independence testing |
0.9 | 1 | 2025 | A Conditional Independence Test in the Presence of Discretization · ICLR 2025 |
Machine learning › Efficient and distributed learning
active learning |
0.8 | 1 | 2024 | Composite Active Learning: Towards Multi-Domain Active Learning with Theoretical Guarantees · AAAI 2024 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
causal reasoning |
0.8 | 1 | 2024 | Natural Counterfactuals With Necessary Backtracking · NeurIPS 2024 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › causal reasoning
counterfactual reasoning |
0.8 | 1 | 2024 | Natural Counterfactuals With Necessary Backtracking · NeurIPS 2024 |
Machine learning › Transfer learning and domain adaptation › cross-domain learning
multi-domain learning |
0.8 | 1 | 2024 | Composite Active Learning: Towards Multi-Domain Active Learning with Theoretical Guarantees · AAAI 2024 |
Mathematical optimization
constrained optimization |
0.8 | 1 | 2024 | Natural Counterfactuals With Necessary Backtracking · NeurIPS 2024 |
Machine learning › Transfer learning and domain adaptation › domain adaptation › distribution adaptation
adversarial domain adaptation |
0.7 | 1 | 2023 | Taxonomy-Structured Domain Adaptation · ICML 2023 |
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training |
0.7 | 1 | 2023 | Taxonomy-Structured Domain Adaptation · ICML 2023 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
0.7 | 1 | 2023 | Domain-Indexing Variational Bayes: Interpretable Domain Index for Domain Adaptation · ICLR 2023 |
Natural language and speech › Information extraction and text analysis
sequence labeling |
0.5 | 1 | 2021 | DyLex: Incorporating Dynamic Lexicons into BERT for Sequence Labeling · EMNLP (1) 2021 |
Machine learning › Generative modeling › generative adversarial network
cross-domain image generation |
0.3 | 1 | 2018 | MIXGAN: Learning Concepts from Different Domains for Mixture Generation · IJCAI 2018 |
Machine learning › Generative modeling
generative adversarial network |
0.3 | 1 | 2018 | MIXGAN: Learning Concepts from Different Domains for Mixture Generation · IJCAI 2018 |
Visual content generation and editing › style transfer
image style transfer |
0.3 | 1 | 2018 | MIXGAN: Learning Concepts from Different Domains for Mixture Generation · IJCAI 2018 |
Machine learning › Trustworthy machine learning › calibration
confidence calibration |
0.3 | 1 | 2026 | BayesAgent: Bayesian Agentic Reasoning Under Uncertainty via Verbalized Probabilistic Graphical Modeling · AAAI 2026 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › structure learning › graphical model learning
bayesian network learning |
0.3 | 1 | 2025 | A Conditional Independence Test in the Presence of Discretization · ICLR 2025 |
Natural language and speech › Language models and text generation › pre-trained language model
BERT |
0.1 | 1 | 2021 | DyLex: Incorporating Dynamic Lexicons into BERT for Sequence Labeling · EMNLP (1) 2021 |
Natural language and speech › Language models and text generation
pre-trained language model |
0.1 | 1 | 2021 | DyLex: Incorporating Dynamic Lexicons into BERT for Sequence Labeling · EMNLP (1) 2021 |
Methods — techniques the papers use, named apart from their topics
structural causal model · 1.5backtracking · 1.5verbalized probabilistic graphical modeling · 1.0bayesian inference · 1.0nonparanormal model · 0.9nodewise regression · 0.9bridge equation · 0.9instance-level query strategy · 0.8domain-level budget allocation · 0.8domain indexing · 0.7generative adversarial network · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BayesAgent: Bayesian Agentic Reasoning Under Uncertainty via Verbalized Probabilistic Graphical ModelingabstractHuman cognition excels at transcending sensory input and forming latent representations that structure our understanding of the world. While Large Language Model (LLM) agents demonstrate emergent reasoning and decision-making abilities, they lack a principled framework for capturing latent structures and modeling uncertainty. In this work, we explore for the first time how to bridge LLM agents with probabilistic graphical models (PGMs) to address agentic reasoning under uncertainty. To this end, we introduce Verbalized Probabilistic Graphical Modeling (vPGM), a Bayesian agentic framework that (i) guides LLM agents in following key principles of PGMs through natural language and (ii) refines the resulting posterior distributions via numerical Bayesian inference. Unlike many traditional probabilistic methods requiring substantial domain expertise, vPGM bypasses expert‐driven model design, making it well‐suited for scenarios with limited assumptions. We evaluated our model on several agentic reasoning tasks, both close-ended and open-ended. Our results indicate that the model effectively enhances confidence calibration and text generation quality. Hengguan Huang, Xing Shen 0001, Guang-Yuan Hao, Lingfa Meng, Dianbo Liu, David A. Duchêne, Hao Wang 0014, Samir Bhatt |
AAAI | 3 |
| 2025 | A Conditional Independence Test in the Presence of DiscretizationabstractTesting conditional independence (CI) has many important applications, such as Bayesian network learning and causal discovery. Although several approaches have been developed for learning CI structures for observed variables, those existing methods generally fail to work when the variables of interest can not be directly observed and only discretized values of those variables are available. For example, if $X_1$, $\tilde{X}_2$ and $X_3$ are the observed variables, where $\tilde{X}_2$ is a discretization of the latent variable $X_2$, applying the existing methods to the observations of $X_1$, $\tilde{X}_2$ and $X_3$ would lead to a false conclusion about the underlying CI of variables $X_1$, $X_2$ and $X_3$.
Motivated by this, we propose a CI test specifically designed to accommodate the presence of discretization. To achieve this, a bridge equation and nodewise regression are used to recover the precision coefficients reflecting the conditional dependence of the latent continuous variables under the nonparanormal model. An appropriate test statistic has been proposed, and its asymptotic distribution under the null hypothesis of CI has been derived.
Theoretical analysis, along with empirical validation on various datasets, rigorously demonstrates the effectiveness of our testing methods. Guang-Yuan Hao, Yumou Qiu |
ICLR | 3 |
| 2025 | Permutation-based Rank Test in the Presence of Discretization and Application in Causal Discovery with Mixed DataabstractRecent advances have shown that statistical tests for the rank of cross-covariance matrices play an important role in causal discovery. These rank tests include partial correlation tests as special cases and provide further graphical information about latent variables. Existing rank tests typically assume that all the continuous variables can be perfectly measured, and yet, in practice many variables can only be measured after discretization. For example, in psychometric studies, the continuous level of certain personality dimensions of a person can only be measured after being discretized into order-preserving options such as disagree, neutral, and agree. Motivated by this, we propose Mixed data Permutation-based Rank Test (MPRT), which properly controls the statistical errors even when some or all variables are discretized. Theoretically, we establish the exchangeability and estimate the asymptotic null distribution by permutations; as a consequence, MPRT can effectively control the Type I error in the presence of discretization while previous methods cannot. Empirically, our method is validated by extensive experiments on synthetic data and real-world data to demonstrate its effectiveness as well as applicability in causal discovery (code will be available at https://github.com/dongxinshuai/scm-identify). Xinshuai Dong, Ignavier Ng, Haoyue Dai, Guang-Yuan Hao, Shunxing Fan, Peter Spirtes, Yumou Qiu, Kun Zhang 0001 |
ICML | 5 |
| 2024 | Composite Active Learning: Towards Multi-Domain Active Learning with Theoretical GuaranteesabstractActive learning (AL) aims to improve model performance within a fixed labeling budget by choosing the most informative data points to label. Existing AL focuses on the single-domain setting, where all data come from the same domain (e.g., the same dataset). However, many real-world tasks often involve multiple domains. For example, in visual recognition, it is often desirable to train an image classifier that works across different environments (e.g., different backgrounds), where images from each environment constitute one domain. Such a multi-domain AL setting is challenging for prior methods because they (1) ignore the similarity among different domains when assigning labeling budget and (2) fail to handle distribution shift of data across different domains. In this paper, we propose the first general method, dubbed composite active learning (CAL), for multi-domain AL. Our approach explicitly considers the domain-level and instance-level information in the problem; CAL first assigns domain-level budgets according to domain-level importance, which is estimated by optimizing an upper error bound that we develop; with the domain-level budgets, CAL then leverages a certain instance-level query strategy to select samples to label from each domain. Our theoretical analysis shows that our method achieves a better error bound compared to current AL methods. Our empirical results demonstrate that our approach significantly outperforms the state-of-the-art AL methods on both synthetic and real-world multi-domain datasets. Code is available at https://github.com/Wang-ML-Lab/multi-domain-active-learning. Guang-Yuan Hao, Hengguan Huang, Hao Wang 0014 |
AAAI | 1 |
| 2024 | Natural Counterfactuals With Necessary BacktrackingabstractCounterfactual reasoning is pivotal in human cognition and especially important for providing explanations and making decisions. While Judea Pearl's influential approach is theoretically elegant, its generation of a counterfactual scenario often requires too much deviation from the observed scenarios to be feasible, as we show using simple examples. To mitigate this difficulty, we propose a framework of natural counterfactuals and a method for generating counterfactuals that are more feasible with respect to the actual data distribution. Our methodology incorporates a certain amount of backtracking when needed, allowing changes in causally preceding variables to minimize deviations from realistic scenarios. Specifically, we introduce a novel optimization framework that permits but also controls the extent of backtracking with a "naturalness'' criterion. Empirical experiments demonstrate the effectiveness of our method. The code is available at https://github.com/GuangyuanHao/natural_counterfactuals. Guang-Yuan Hao, Jiji Zhang, Biwei Huang, Hao Wang 0014, Kun Zhang 0001 |
NeurIPS | 1 |
| 2023 | Domain-Indexing Variational Bayes: Interpretable Domain Index for Domain Adaptation
Zihao Xu 0001, Guang-Yuan Hao, Hao He 0011, Hao Wang 0014 |
ICLR | 2 |
| 2023 | Taxonomy-Structured Domain AdaptationabstractDomain adaptation aims to mitigate distribution shifts among different domains. However, traditional formulations are mostly limited to categorical domains, greatly simplifying nuanced domain relationships in the real world. In this work, we tackle a generalization with taxonomy-structured domains, which formalizes domains with nested, hierarchical similarity structures such as animal species and product catalogs. We build on the classic adversarial framework and introduce a novel taxonomist, which competes with the adversarial discriminator to preserve the taxonomy information. The equilibrium recovers the classic adversarial domain adaptation’s solution if given a non-informative domain taxonomy (e.g., a flat taxonomy where all leaf nodes connect to the root node) while yielding non-trivial results with other taxonomies. Empirically, our method achieves state-of-the-art performance on both synthetic and real-world datasets with successful adaptation. Zihao Xu 0001, Hao He 0011, Guang-Yuan Hao, Guang-He Lee, Hao Wang 0014 |
ICML | 4 |
| 2021 | DyLex: Incorporating Dynamic Lexicons into BERT for Sequence LabelingabstractBaojun Wang, Zhao Zhang, Kun Xu, Guang-Yuan Hao, Yuyang Zhang, Lifeng Shang, Linlin Li, Xiao Chen, Xin Jiang, Qun Liu. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Baojun Wang, Guang-Yuan Hao, Lifeng Shang, Linlin Li 0001, Xiao Chen 0012, Xin Jiang 0002, Qun Liu 0001 |
EMNLP (1) | 4 |
| 2018 | MIXGAN: Learning Concepts from Different Domains for Mixture GenerationabstractIn this work, we present an interesting attempt on mixture generation: absorbing different image concepts (e.g., content and style) from different domains and thus generating a new domain with learned concepts. In particular, we propose a mixture generative adversarial network (MIXGAN). MIXGAN learns concepts of content and style from two domains respectively, and thus can join them for mixture generation in a new domain, i.e., generating images with content from one domain and style from another. MIXGAN overcomes the limitation of current GAN-based models which either generate new images in the same domain as they observed in training stage, or require off-the-shelf content templates for transferring or translation. Extensive experimental results demonstrate the effectiveness of MIXGAN as compared to related state-of-the-art GAN-based models. Guang-Yuan Hao, Hong-Xing Yu, Wei-Shi Zheng 0001 |
IJCAI | 1 |