Zihao Wang 0001

dblp:148/9655-1 · DBLP profile ↗
← Back
25ranked-venue papers
5as first author
20since 2021 · last 2026
0000-0002-3919-0396ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 5 first-author · 18 since 2021Databases, data management, data science and information retrieval · 7 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 RPGen: Robust and Differentially Private Synthetic Image Generation
abstract
Differentially private (DP) image synthesis enables the generation of realistic images while bounding privacy leakage, facilitating secure data sharing across organizations. However, the Gaussian noise injected during DP training, such as via DP-SGD, often severely degrades synthesis quality by disrupting model convergence. To address this, we introduce RPGen, a novel framework that enhances diffusion models' parameter robustness to mitigate DP noise effects without compromising privacy guarantees. At its core, RPGen employs adversarial model perturbation (AMP) during public pre-training to build resilience against perturbations, but we identify and tackle the critical issue of robustness transferability across domains. RPGen achieves this through a three-step process: (1) A pre-trained classifier infers labels for private images, aggregated into a class distribution noised with Gaussian mechanism for DP, and public samples are selected to match this privatized distribution for domain alignment; (2) The diffusion model is pre-trained on this curated subset with adversarial model perturbation to foster robustness; (3) The model undergoes fine-tuning on private data using DP-SGD. This synergy of robustness augmentation and transferability optimization yields high-fidelity synthesis. Extensive evaluations on ImageNet for pre-training, with CelebA and CIFAR-10 for synthesis, show RPGen outperforming state-of-the-art baselines across epsilon in 1, 5, 10. On average, it achieves 20.18% lower FID and 5.45% higher classification accuracy. Ablations confirm the efficacy of domain curation and modest perturbations, establishing RPGen as a new benchmark for privacy-utility trade-offs in image generation.
Zihao Wang 0001, Hao Peng 0001, Yuecen Wei, Li Sun 0008, Zhengtao Yu 0001
AAAI1
2026 Activation-Guided Local Editing for Jailbreaking Attacks
abstract
Jiecong Wang, Haoran Li, Hao Peng, Ziqian Zeng, Zihao Wang, Haohua Du, Zhengtao Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiecong Wang, Haoran Li 0003, Hao Peng 0001, Ziqian Zeng, Zihao Wang 0001, Haohua Du, Zhengtao Yu 0001
ACL (1)5
2025 Extending Complex Logical Queries on Uncertain Knowledge Graphs
abstract
The study of machine learning-based logical query-answering enables reasoning with large-scale and incomplete knowledge graphs. This paper further advances this line of research by considering the uncertainty in the knowledge. The uncertain nature of knowledge is widely observed in the real world, but does not align seamlessly with the first-order logic underpinning existing studies. To bridge this gap, we study the setting of soft queries on uncertain knowledge, which is motivated by the establishment of soft constraint programming. We further propose an ML-based approach with both forward inference and backward calibration to answer soft queries on large-scale, incomplete, and uncertain knowledge graphs. Theoretical discussions reveal that our method ensures there are no catastrophic cascading errors in our forward inference algorithm while maintaining the same complexity as state-of-the-art inference algorithms for first-order queries. Empirical results justify the superior performance of our approach against previous ML-based methods with number embedding extensions.
Weizhi Fei, Zihao Wang 0001, Hang Yin 0008, Yang Duan, Yangqiu Song
ACL (1)2
2025 Enhancing Transformers for Generalizable First-Order Logical Entailment
abstract
Tianshi Zheng, Jiazheng Wang, Zihao Wang, Jiaxin Bai, Hang Yin, Zheye Deng, Yangqiu Song, Jianxin Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Tianshi Zheng, Zihao Wang 0001, Jiaxin Bai, Hang Yin 0008, Zheye Deng, Yangqiu Song, Jianxin Li 0002
ACL (1)3
2025 Few-Shot Knowledge Graph Completion via Transfer Knowledge from Similar Tasks
abstract
Knowledge graphs (KGs) are essential in many AI applications but often suffer from incompleteness, limiting their utility. Many relations in KGs have only a few examples, making it challenging to train accurate models. Few-shot learning offers a promising direction by enabling KG completion with only a small number of training triplets. However, most existing approaches treat each relation independently and fail to leverage shared information across tasks. In this paper, we introduce TransNet, a transfer learning method for few-shot KG completion that captures task relationships and reuses knowledge from related tasks. TransNet further incorporates meta-learning to effectively handle unseen relations. Experiments on standard benchmarks demonstrate that TransNet achieves strong performance compared to prior methods. Code and data will be released upon acceptance.
Lihui Liu, Zihao Wang 0001, Dawei Zhou 0003, Ruijie Wang 0004, Sihong He, Hanghang Tong
CIKM2
2025 From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery
abstract
Large Language Models (LLMs) are catalyzing a paradigm shift in scientific discovery, evolving from task-specific automation tools into increasingly autonomous agents and fundamentally redefining research processes and human-AI collaboration.This survey systematically charts this burgeoning field, placing a central focus on the changing roles and escalating capabilities of LLMs in science.Through the lens of the scientific method, we introduce a foundational three-level taxonomy-Tool, Analyst, and Scientist-to delineate their escalating autonomy and evolving responsibilities within the research lifecycle.We further identify pivotal challenges and future research trajectories such as robotic automation, self-improvement, and ethical governance.Overall, this survey provides a conceptual architecture and strategic foresight to navigate and shape the future of AI-driven scientific discovery, fostering both rapid innovation and responsible advancement.
Tianshi Zheng, Zheye Deng, Hong Ting Tsang, Weiqi Wang 0001, Jiaxin Bai, Zihao Wang 0001, Yangqiu Song
EMNLP6
2025 LogiDynamics: Unraveling the Dynamics of Inductive, Abductive and Deductive Logical Inferences in LLM Reasoning
abstract
Tianshi Zheng, Cheng Jiayang, Chunyang Li, Haochen Shi, Zihao Wang, Jiaxin Bai, Yangqiu Song, Ginny Wong, Simon See. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Tianshi Zheng, Cheng Jiayang, Zihao Wang 0001, Jiaxin Bai, Yangqiu Song, Ginny Y. Wong, Simon See
EMNLP5
2025 EFOk-CQA: Towards Knowledge Graph Complex Query Answering beyond Set Operation
abstract
To answer complex queries on knowledge graphs, logical reasoning over incomplete knowledge needs learning-based methods because they are capable of generalizing over unobserved knowledge. Therefore, an appropriate dataset is fundamental to both obtaining and evaluating such methods under this paradigm. In this paper, we propose a comprehensive framework for data generation, model training, and method evaluation that covers the combinatorial space of Existential First-order Queries with multiple variables (EFOk). The combinatorial query space in our framework significantly extends those defined by set operations in the existing literature. Additionally, we construct a dataset, EFOk-CQA, with 741 query types for empirical evaluation, and our benchmark results provide new insights into how query hardness affects the results. Furthermore, we demonstrate that the existing dataset construction process is systematically biased and hinders the appropriate development of query-answering methods, highlighting the importance of our work. Our code and data are provided in https://github.com/HKUST-KnowComp/EFOK-CQA.
Hang Yin 0008, Zihao Wang 0001, Weizhi Fei, Yangqiu Song
KDD (2)2
2024 Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key?
abstract
Recent progress in LLMs discussion suggests that multi-agent discussion improves the reasoning abilities of LLMs.In this work, we reevaluate this claim through systematic experiments, where we propose a novel group discussion framework to enrich the set of discussion mechanisms.Interestingly, our results show that a single-agent LLM with strong prompts can achieve almost the same performance as the best existing discussion approach on a wide range of reasoning tasks and backbone LLMs.We observe that the multi-agent discussion performs better than a single agent only when there is no demonstration in the prompt.Further study reveals the common interaction mechanisms of LLMs during the discussion. 1
Qineng Wang, Zihao Wang 0001, Hanghang Tong, Yangqiu Song
ACL (1)2
2024 LLMs Assist NLP Researchers: Critique Paper (Meta-)Reviewing
abstract
Jiangshu Du, Yibo Wang, Wenting Zhao, Zhongfen Deng, Shuaiqi Liu, Renze Lou, Henry Peng Zou, Pranav Narayanan Venkit, Nan Zhang, Mukund Srinath, Haoran Ranran Zhang, Vipul Gupta, Yinghui Li, Tao Li, Fei Wang, Qin Liu, Tianlin Liu, Pengzhi Gao, Congying Xia, Chen Xing, Cheng Jiayang, Zhaowei Wang, Ying Su, Raj Sanjay Shah, Ruohao Guo, Jing Gu, Haoran Li, Kangda Wei, Zihao Wang, Lu Cheng, Surangika Ranathunga, Meng Fang, Jie Fu, Fei Liu, Ruihong Huang, Eduardo Blanco, Yixin Cao, Rui Zhang, Philip S. Yu, Wenpeng Yin. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Jiangshu Du, Yibo Wang 0001, Wenting Zhao 0006, Zhongfen Deng, Shuaiqi Liu 0002, Renze Lou, Henry Peng Zou, Pranav Venkit, Mukund Srinath, Ranran Haoran Zhang, Tao Li 0039, Fei Wang 0060, Qin Liu 0010, Tianlin Liu, Pengzhi Gao, Congying Xia, Chen Xing, Cheng Jiayang, Zhaowei Wang 0003, Raj Sanjay Shah, Ruohao Guo, Haoran Li 0003, Kangda Wei, Zihao Wang 0001, Lu Cheng 0001, Surangika Ranathunga, Fei Liu 0004, Ruihong Huang, Eduardo Blanco 0002, Yixin Cao 0002, Rui Zhang 0037, Philip S. Yu, Wenpeng Yin 0001
EMNLP29
2024 Generate-on-Graph: Treat LLM as both Agent and KG for Incomplete Knowledge Graph Question Answering
abstract
To address the issues of insufficient knowledge and hallucination in Large Language Models (LLMs), numerous studies have explored integrating LLMs with Knowledge Graphs (KGs).However, these methods are typically evaluated on conventional Knowledge Graph Question Answering (KGQA) with complete KGs, where all factual triples required for each question are entirely covered by the given KG.In such cases, LLMs primarily act as an agent to find answer entities within the KG, rather than effectively integrating the internal knowledge of LLMs and external knowledge sources such as KGs.In fact, KGs are often incomplete to cover all the knowledge required to answer questions.To simulate these real-world scenarios and evaluate the ability of LLMs to integrate internal and external knowledge, we propose leveraging LLMs for QA under Incomplete Knowledge Graph (IKGQA), where the provided KG lacks some of the factual triples for each question, and construct corresponding datasets.To handle IKGQA, we propose a training-free method called Generate-on-Graph (GoG), which can generate new factual triples while exploring KGs.Specifically, GoG performs reasoning through a Thinking-Searching-Generating framework, which treats LLM as both Agent and KG in IKGQA.Experimental results on two datasets demonstrate that our GoG outperforms all previous methods.
Shizhu He, Jiabei Chen, Zihao Wang 0001, Yangqiu Song, Hanghang Tong, Jun Zhao 0001, Kang Liu 0001
EMNLP4
2024 Rethinking Complex Queries on Knowledge Graphs with Neural Link Predictors
abstract
Reasoning on knowledge graphs is a challenging task because it utilizes observed information to predict the missing one. Particularly, answering complex queries based on first-order logic is one of the crucial tasks to verify learning to reason abilities for generalization and composition. Recently, the prevailing method is query embedding which learns the embedding of a set of entities and treats logic operations as set operations and has shown great empirical success. Though there has been much research following the same formulation, many of its claims lack a formal and systematic inspection. In this paper, we rethink this formulation and justify many of the previous claims by characterizing the scope of queries investigated previously and precisely identifying the gap between its formulation and its goal, as well as providing complexity analysis for the currently investigated queries. Moreover, we develop a new dataset containing ten new types of queries with features that have never been considered and therefore can provide a thorough investigation of complex queries. Finally, we propose a new neural-symbolic method, Fuzzy Inference with Truth value (FIT), where we equip the neural link predictors with fuzzy logic theory to support end-to-end learning using complex queries with provable reasoning capability. Empirical results show that our method outperforms previous methods significantly in the new dataset and also surpasses previous methods in the existing dataset at the same time.
Hang Yin 0008, Zihao Wang 0001, Yangqiu Song
ICLR2
2024 Privacy-Preserved Neural Graph Databases
abstract
In the era of large language models (LLMs), efficient and accurate data retrieval has become increasingly crucial for the use of domain-specific or private data in the retrieval augmented generation (RAG). Neural graph databases (NGDBs) have emerged as a powerful paradigm that combines the strengths of graph databases (GDBs) and neural networks to enable efficient storage, retrieval, and analysis of graph-structured data which can be adaptively trained with LLMs. The usage of neural embedding storage and Complex neural logical Query Answering (CQA) provides NGDBs with generalization ability. When the graph is incomplete, by extracting latent patterns and representations, neural graph databases can fill gaps in the graph structure, revealing hidden relationships and enabling accurate query answering. Nevertheless, this capability comes with inherent trade-offs, as it introduces additional privacy risks to the domain-specific or private databases. Malicious attackers can infer more sensitive information in the database using well-designed queries such as from the answer sets of where Turing Award winners born before 1950 and after 1940 lived, the living places of Turing Award winner Hinton are probably exposed, although the living places may have been deleted in the training stage due to the privacy concerns. In this work, we propose a privacy-preserved neural graph database (P-NGDB) framework to alleviate the risks of privacy leakage in NGDBs. We introduce adversarial training techniques in the training stage to enforce the NGDBs to generate indistinguishable answers when queried with private information, enhancing the difficulty of inferring sensitive information through combinations of multiple innocuous queries. Extensive experimental results on three datasets show that our framework can effectively protect private information in the graph database while delivering high-quality public answers responses to queries. The code is available at https://github.com/HKUST-KnowComp/PrivateNGDB.
Haoran Li 0003, Jiaxin Bai, Zihao Wang 0001, Yangqiu Song
KDD4
2024 Optimal Transport Enhanced Cross-City Site Recommendation
abstract
Site recommendation, which aims at predicting the optimal location for brands to open new branches, has demonstrated an important role in assisting decision-making in modern business. In contrast to traditional recommender systems that can benefit from extensive information, site recommendation starkly suffers from extremely limited information and thus leads to unsatisfactory performance. Therefore, existing site recommendation methods primarily focus on several specific name brands and heavily rely on fine-grained human-crafted features to avoid the data sparsity problem. However, such solutions are not able to fulfill the demand for rapid development in modern business. Therefore, we aim to alleviate the data sparsity problem by effectively utilizing data across multiple cities and thereby propose a novel Optimal Transport enhanced Cross-city (OTC) framework for site recommendation. Specifically, OTC leverages optimal transport (OT) on the learned embeddings of brands and regions separately to project the brands and regions from the source city to the target city. Then, the projected embeddings of brands and regions are utilized to obtain the inference recommendation in the target city. By integrating the original recommendation and the inference recommendations from multiple cities, OTC is able to achieve enhanced recommendation results. The experimental results on the real-world OpenSiteRec dataset, encompassing thousands of brands and regions across four metropolises, demonstrate the effectiveness of our proposed OTC in further improving the performance of site recommendation models.
Xinhang Li 0001, Xiangyu Zhao 0001, Zihao Wang 0001, Yang Duan, Yong Zhang 0002, Chunxiao Xing
SIGIR3
2023 Logical Message Passing Networks with One-hop Inference on Atomic Formulas
Zihao Wang 0001, Yangqiu Song, Ginny Y. Wong, Simon See
ICLR1
2023 spred: Solving L1 Penalty with SGD
abstract
We propose to minimize a generic differentiable objective with $L_1$ constraint using a simple reparametrization and straightforward stochastic gradient descent. Our proposal is the direct generalization of previous ideas that the $L_1$ penalty may be equivalent to a differentiable reparametrization with weight decay. We prove that the proposed method, spred, is an exact differentiable solver of $L_1$ and that the reparametrization trick is completely “benign" for a generic nonconvex function. Practically, we demonstrate the usefulness of the method in (1) training sparse neural networks to perform gene selection tasks, which involves finding relevant features in a very high dimensional space, and (2) neural network compression task, to which previous attempts at applying the $L_1$-penalty have been unsuccessful. Conceptually, our result bridges the gap between the sparsity in deep learning and conventional statistical learning.
Liu Ziyin 0001, Zihao Wang 0001
ICML2
2022 Gromov-Wasserstein Guided Representation Learning for Cross-Domain Recommendation
abstract
Cross-Domain Recommendation (CDR) has attracted increasing attention in recent years as a solution to the data sparsity issue. The fundamental paradigm of prior efforts is to train a mapping function based on the overlapping users/items and then apply it to the knowledge transfer. However, due to the commercial privacy policy and the sensitivity of user data, it is unrealistic to explicitly share the user mapping relations and behavior data. Therefore, in this paper, we consider a more practical cross-domain scenario, where there is no explicit overlap between the source and target domains in terms of users/items. Since the user sets of both domains are drawn from the entire population, there may be commonalities between their user characteristics, resulting in comparable user preference distributions. Thus, without the mapping relations at user level, it is feasible to model this distribution-level relation to transfer knowledge between domains. To this end, we propose a novel framework that improves the effect of representation learning on the target domain by aligning the representation distributions between the source and target domains. In addition, GWCDR can be easily integrated with existing single-domain collaborative filtering methods to achieve cross-domain recommendation. Extensive experiments on two pairs of public bidirectional datasets demonstrate the effectiveness of our proposed framework in enhancing the recommendation performance.
Xinhang Li 0001, Zhaopeng Qiu, Xiangyu Zhao 0001, Zihao Wang 0001, Yong Zhang 0002, Chunxiao Xing, Xian Wu 0001
CIKM4
2022 Unsupervised Sentence Textual Similarity with Compositional Phrase Semantics
abstract
Measuring Sentence Textual Similarity (STS) is a classic task that can be applied to many downstream NLP applications such as text generation and retrieval. In this paper, we focus on unsupervised STS that works on various domains but only requires minimal data and computational resources. Theoretically, we propose a light-weighted Expectation-Correction (EC) formulation for STS computation. EC formulation unifies unsupervised STS approaches including the cosine similarity of Additively Composed (AC) sentence embeddings, Optimal Transport (OT), and Tree Kernels (TK). Moreover, we propose the Recursive Optimal Transport Similarity (ROTS) algorithm to capture the compositional phrase semantics by composing multiple recursive EC formulations. ROTS finishes in linear time and is faster than its predecessors. ROTS is empirically more effective and scalable than previous approaches. Extensive experiments on 29 STS tasks under various settings show the clear advantage of ROTS over existing approaches. Detailed ablation studies prove the effectiveness of our approaches.
Zihao Wang 0001, Jiaheng Dou, Yong Zhang 0002
COLING1
2022 Posterior Collapse of a Linear Latent Variable Model
abstract
This work identifies the existence and cause of a type of posterior collapse that frequently occurs in the Bayesian deep learning practice. For a general linear latent variable model that includes linear variational autoencoders as a special case, we precisely identify the nature of posterior collapse to be the competition between the likelihood and the regularization of the mean due to the prior. Our result also suggests that posterior collapse may be a general problem of learning for deeper architectures and deepens our understanding of Bayesian deep learning.
Zihao Wang 0001, Liu Ziyin 0001
NeurIPS1
2021 TC-MIMONet: A Learning-based Transceiver for MIMO Systems with Temporal Correlations
abstract
Data-driven approaches have recently emerged as promising remedies for communication system designs, which leverage deep learning techniques for automated development and optimization. In this paper, we revisit the designs of multi-input multi-output (MIMO) wireless systems and investigate the end-to-end learning for MIMO systems with temporal correlations. Our objective is to develop a MIMO transceiver to improve the communication performance by making fully use of the available temporal information. Although the end-to-end learning framework has been applied to various communication systems, existing designs largely rely on memoryless autoencoders (AEs) and overlook the time dependency. To overcome this issue, we propose a novel learning-based MIMO transceiver, namely, the TC-MIMONet, which extends the conventional memoryless AE-based transceivers by customizing two neural network components with memory. In particular, a long short-term memory (LSTM)-based CSI predictor is adopted at the transmitter, while a two-timescale LSTM-based decoder is developed for the receiver. Simulation results show that TC-MIMONet achieves significant block error rate reduction compared to two baseline schemes without utilizing the available temporal information.
Chunhui Chen 0005, Zihao Wang 0001, Yuyi Mao, Hao Wu 0060, Bo Bai 0001, Gong Zhang 0001
VTC Spring2
2020 A Relaxed Matching Procedure for Unsupervised BLI
abstract
Recently unsupervised Bilingual Lexicon Induction(BLI) without any parallel corpus has attracted much research interest.One of the crucial parts in methods for the BLI task is the matching procedure.Previous works impose a too strong constraint on the matching and lead to many counterintuitive translation pairings.Thus, We propose a relaxed matching procedure to find a more precise matching between two languages.We also find that aligning source and target language embedding space bidirectionally will bring significant improvement.We follow the previous iterative framework to conduct experiments.Results on standard benchmark demonstrate the effectiveness of our proposed method, which substantially outperforms previous unsupervised methods.
Xu Zhao 0007, Zihao Wang 0001, Yong Zhang 0002, Hao Wu 0060
ACL2
2020 Robust Document Distance with Wasserstein-Fisher-Rao metric
abstract
Computing the distance among linguistic objects is an essential problem in natural language processing. The word mover’s distance (WMD) has been successfully applied to measure the document distance by synthesizing the low-level word similarity with the framework of optimal transport (OT). However, due to the global transportation nature of OT, the WMD may overestimate the semantic dissimilarity when documents contain unequal semantic details. In this paper, we propose to address this overestimation issue with a novel Wasserstein-Fisher-Rao (WFR) document distance grounded on unbalanced optimal transport theory. Compared to the WMD, the WFR document distance provides a trade-off between global transportation and local truncation, which leads to a better similarity measure for unequal semantic details. Moreover, an efficient prune strategy is particularly designed for the WFR document distance to facilitate the top-k queries among a large number of documents. Extensive experimental results show that the WFR document distance achieves higher accuracy that WMD and even its supervised variation s-WMD.
Zihao Wang 0001, Datong Zhou, Yong Zhang 0002, Chenglong Rao, Hao Wu 0060
ACML1
2020 Instance Explainable Multi-instance Learning for ROI of Various Data
Xu Zhao 0007, Zihao Wang 0001, Yong Zhang 0002, Chunxiao Xing
DASFAA (2)2
2020 Semi-Supervised Bilingual Lexicon Induction with Two-way Interaction
abstract
Semi-supervision is a promising paradigm for Bilingual Lexicon Induction (BLI) with limited annotations.However, previous semisupervised methods do not fully utilize the knowledge hidden in annotated and nonannotated data, which hinders further improvement of their performance.In this paper, we propose a new semi-supervised BLI framework to encourage the interaction between the supervised signal and unsupervised alignment.We design two message-passing mechanisms to transfer knowledge between annotated and non-annotated data, named prior optimal transport and bi-directional lexicon update respectively.Then, we perform semi-supervised learning based on a cyclic or a parallel parameter feeding routine to update our models.Our framework is a general framework that can incorporate any supervised and unsupervised BLI methods based on optimal transport.Experimental results on MUSE and VecMap datasets show significant improvement of our models.Ablation study also proves that the two-way interaction between the supervised signal and unsupervised alignment accounts for the gain of the overall performance.Results on distant language pairs further illustrate the advantage and robustness of our proposed method.
Xu Zhao 0007, Zihao Wang 0001, Hao Wu 0060, Yong Zhang 0002
EMNLP (1)2
2018 Modeling Patient Visit Using Electronic Medical Records for Cost Profile Estimation
Kangzhi Zhao, Yong Zhang 0002, Zihao Wang 0001, Hongzhi Yin, Xiaofang Zhou 0001, Jin Wang 0007, Chunxiao Xing
DASFAA (2)3