Zhihua Wen

dblp:76/2699 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0002-2224-6197ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Software engineering, systems software and programming languages · 5 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 3 · 2 first-author · 1 since 2021Computer networks · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Knowledge Injection Exists in MoE? Exploring Expert-Aware Contrast Decoding in MoE for Mitigating LLMs' Hallucinations
abstract
Xinyue Fang, Zhiliang Tian, Zhen Huang, Ziyi Pan, Zhihua Wen, Xi Wang, Quntian Fang, Dongsheng Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xinyue Fang, Zhiliang Tian, Zhen Huang 0006, Ziyi Pan, Zhihua Wen, Quntian Fang, Dongsheng Li 0001
ACL (1)5
2026 Lightweight and Accurate Printed Circuit Board Defect Detection Network Based on Cross-scale Feature Fusion
Longxin Zhang, Lvkui Jiang, Zhihua Wen
ICIC (18)4
2026 Rethinking the Hidden Risk of Reranking: Achieving Risk-aware Reranking with Information Gain for RAG with LLMs
abstract
Retrieval-augmented generation (RAG) has become a cornerstone for enhancing large language models (LLMs) with real-time information from the Web, but its performance often heavily depends on the quality of the retrieved documents. Given that RAG systems frequently draw from vast and often noisy Web corpora, ensuring the reliability of retrieved content is paramount. While rerankers improve the factual accuracy of the RAG system by elevating the proportion of ground-truth documents (GD) in high-ranked results, the shifts of document type distributions during reranking remain unclear, hindering the understanding of the reranker's behavior. To bridge this gap, we conduct an empirical study to categorize documents and compare their distribution before and after reranking. We reveal a counterintuitive finding: though rerankers improve the proportion of GD, they also significantly increase the proportion of harmful documents (HD) in top-ranked retrieved documents. It not only narrows the potential context window for ranking the GD higher but also increases the risk of HD misleading the LLMs, potentially leading to the generation and propagation of misinformation across Web platforms. Motivated by this finding, we propose a risk-aware reranking method for RAG with LLMs, which balances the risk and benefit during reranking. Given a query, the RAG framework first retrieves relevant documents. Then, our approach quantifies the potential beneficial and harmful impacts of various documents on the LLMs' generation. To estimate the impacts, we conduct a dual-aspect document impact assessment via information gain, which employs a risk clipping to avoid the numerical fluctuations in the estimation. Finally, we conduct the reranking according to the potential impact of each document, enabling the reranker to significantly reduce the HD proportion. Experiments and analysis across multiple models and datasets, including Wikipedia, web news, and research papers, show the effectiveness of our method. Our code is available at https://github.com/lzz335/hidden_risk_of_reranking.
Zhizhao Liu, Zhihua Wen, Zhiliang Tian, Zhen Huang 0006, Miaorong Zhu, Zimian Wei, Yifu Gao, Liang Ding 0006, Dongsheng Li 0001
WWW2
2025 Zero-resource Hallucination Detection for Text Generation via Graph-based Contextual Knowledge Triples Modeling
abstract
LLMs obtain remarkable performance but suffer from hallucinations. Most research on detecting hallucination focuses on questions with short and concrete correct answers that are easy to check faithfulness. Hallucination detections for text generation with open-ended answers are more hard. Some researchers use external knowledge to detect hallucinations in generated texts, but external resources for specific scenarios are hard to access. Recent studies on detecting hallucinations in long texts without external resources conduct consistency comparison among multiple sampled outputs. To handle long texts, researchers split long texts into multiple facts and individually compare the consistency of each pair of facts. However, these methods (1) hardly achieve alignment among multiple facts; (2) overlook dependencies between multiple contextual facts. In this paper, we propose a graph-based context-aware (GCA) hallucination detection method for text generations, which aligns facts and considers the dependencies between contextual facts in consistency comparison. Particularly, to align multiple facts, we conduct a triple-oriented response segmentation to extract multiple knowledge triples. To model dependencies among contextual triples (facts), we construct contextual triples into a graph and enhance triples’ interactions via message passing and aggregating via RGCN. To avoid the omission of knowledge triples in long texts, we conduct an LLM-based reverse verification by reconstructing the knowledge triples. Experiments show that our model enhances hallucination detection and excels all baselines.
Xinyue Fang, Zhen Huang 0006, Zhiliang Tian, Minghui Fang 0002, Ziyi Pan, Quntian Fang, Zhihua Wen, Hengyue Pan
AAAI7
2025 AGD: Adversarial Game Defense Against Jailbreak Attacks in Large Language Models
abstract
Shilong Pan, Zhiliang Tian, Zhen Huang, Wanlong Yu, Zhihua Wen, Xinwang Liu, Kai Lu, Minlie Huang, Dongsheng Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Shilong Pan, Zhiliang Tian, Zhen Huang 0006, Wanlong Yu, Zhihua Wen, Xinwang Liu 0002, Kai Lu 0001, Minlie Huang, Dongsheng Li 0001
ACL (1)5
2025 Self-Persuasion: A Novel Cognitive Approach to Effective LLM Jailbreaking
Wei Xie 0007, Shuoyoucheng Ma, Zhihua Wen, Enze Wang
CogSci6
2025 A Deep Reinforcement Learning Algorithm with Ordered Action Space for Budget-Aware Workflow Scheduling in Heterogeneous Clouds
Yanfen Zhang, Longxin Zhang, Lili Du, Zhihua Wen, Buqing Cao, Jianguo Chen 0001
ICA3PP (3)4
2025 Scenario-independent Uncertainty Estimation for LLM-based Question Answering via Factor Analysis
abstract
Large language models (LLMs) demonstrate significant potential in various applications; however, they are susceptible to generating hallucinations, which can lead to the spread of online misinformation. Existing studies address hallucination detection by (1) employing reference-based methods that consult external resources for verification or (2) utilizing reference-free methods that mainly estimate answer uncertainty based on LLM's internal states. However, reference-based methods incur significant costs and can be infeasible for obtaining reliable external references. Besides, existing uncertainty estimation (UE) methods often overlook the impact of scenario backgrounds inherited from the query's lexical resources, leading to noise in UE. In almost all real-world applications, users care about the uncertainty concerning semantics or facts instead of the query's scenario information. Therefore, we argue that mitigating scenario-related noise and focusing on semantic information can yield a more desirable UE. In this paper, we introduce a plug-and-play scenario-independent framework to enhance unsupervised UE in LLMs by removing scenario-related noise and focusing on semantic information. This framework is compatible with most existing UE methods, as it leverages only the existing UE methods' outputs. Specifically, we design a scenario-specific sampling to paraphrase queries, maintaining their common semantics while diversifying the scenario distribution. Subsequently, to estimate the contribution of the common semantics, we design a factor analysis (FA) model to disentangle the UE score obtained from the given UE method into a combination of multiple latent factors, which represent the contribution of the common semantics and scenario-related noise. By solving the FA model, we decompose the impact of the most significant factor to approximate the uncertainty caused by the common semantics, thus achieving scenario-independent UE. Extensive experiments and analysis across multiple models and datasets demonstrate the effectiveness of our approach.
Zhihua Wen, Zhizhao Liu, Zhiliang Tian, Shilong Pan, Zhen Huang 0006, Dongsheng Li 0001, Minlie Huang
WWW1
2024 POMP: Probability-driven Meta-graph Prompter for LLMs in Low-resource Unsupervised Neural Machine Translation
abstract
Shilong Pan, Zhiliang Tian, Liang Ding, Haoqi Zheng, Zhen Huang, Zhihua Wen, Dongsheng Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Shilong Pan, Zhiliang Tian, Liang Ding 0006, Haoqi Zheng, Zhen Huang 0006, Zhihua Wen, Dongsheng Li 0001
ACL (1)6
2024 Perception of Knowledge Boundary for Large Language Models through Semi-open-ended Question Answering
abstract
Large Language Models (LLMs) are widely used for knowledge-seeking purposes yet suffer from hallucinations. The knowledge boundary of an LLM limits its factual understanding, beyond which it may begin to hallucinate. Investigating the perception of LLMs' knowledge boundary is crucial for detecting hallucinations and LLMs' reliable generation. Current studies perceive LLMs' knowledge boundary on questions with concrete answers (close-ended questions) while paying limited attention to semi-open-ended questions that correspond to many potential answers. Some researchers achieve it by judging whether the question is answerable or not. However, this paradigm is not so suitable for semi-open-ended questions, which are usually ``partially answerable questions'' containing both answerable answers and ambiguous (unanswerable) answers. Ambiguous answers are essential for knowledge-seeking, but it may go beyond the knowledge boundary of LLMs. In this paper, we perceive the LLMs' knowledge boundary with semi-open-ended questions by discovering more ambiguous answers. First, we apply an LLM-based approach to construct semi-open-ended questions and obtain answers from a target LLM. Unfortunately, the output probabilities of mainstream black-box LLMs are inaccessible to sample more low-probability ambiguous answers. Therefore, we apply an open-sourced auxiliary model to explore ambiguous answers for the target LLM. We calculate the nearest semantic representation for existing answers to estimate their probabilities, with which we reduce the generation probability of high-probability existing answers to achieve a more effective generation. Finally, we compare the results from the RAG-based evaluation and LLM self-evaluation to categorize four types of ambiguous answers that are beyond the knowledge boundary of the target LLM. Following our method, we construct a dataset to perceive the knowledge boundary for GPT-4. We find that GPT-4 performs poorly on semi-open-ended questions and is often unaware of its knowledge boundary. Besides, our auxiliary model, LLaMA-2-13B, is effective in discovering many ambiguous answers, including correct answers neglected by GPT-4 and delusive wrong answers GPT-4 struggles to identify.
Zhihua Wen, Zhiliang Tian, Zexin Jian, Zhen Huang 0006, Pei Ke, Yifu Gao, Minlie Huang, Dongsheng Li 0001
NeurIPS1
2023 Retrieval-Augmented GPT-3.5-Based Text-to-SQL Framework with Sample-Aware Prompting and Dynamic Revision Chain
Chunxi Guo, Zhiliang Tian, Jintao Tang, Shasha Li 0001, Zhihua Wen, Ting Wang 0009
ICONIP (6)5
2023 Prompting GPT-3.5 for Text-to-SQL with De-semanticization and Skeleton Retrieval
Chunxi Guo, Zhiliang Tian, Jintao Tang, Pancheng Wang, Zhihua Wen, Ting Wang 0009
PRICAI (2)5
2022 Emotion-Aware Multimodal Pre-training for Image-Grounded Emotional Response Generation
Zhiliang Tian, Zhihua Wen, Yiping Song, Jintao Tang, Dongsheng Li 0001, Nevin Lianwen Zhang
DASFAA (3)2
2014 The MoJo family: a story about clustering evaluation (invited talk)
abstract
The need to decompose large, complex software systems into smaller, more manageable subsystems has been recognized for more than two decades. Many cluster analysis algorithms have been applied to the software domain, and several algorithms specializing in software clustering have been developed. This in turn has created the need to evaluate and compare clustering results.
Zhihua Wen, Vassilios Tzerpos
ICPC1
2011 Measuring a commercial content delivery network
abstract
Content delivery networks (CDNs) have become a crucial part of the modern Web infrastructure. This paper studies the performance of the leading content delivery provider - Akamai. It measures the performance of the current Akamai platform and considers a key architectural question faced by both CDN designers and their prospective customers: whether the co-location approach to CDN platforms adopted by Akamai, which tries to deploy servers in numerous Internet locations, brings inherent performance benefits over a more consolidated data center approach pursued by other influential CDNs such as Limelight. We believe the methodology we developed for this study will be useful for other researchers in the CDN arena.
Sipat Triukose, Zhihua Wen, Michael Rabinovich
WWW2
2011 Dynamic landmark triangles: A simple and efficient mechanism for inter-host latency estimation
Zhihua Wen, Michael Rabinovich
Comput. Networks1
2008 Network distance estimation with dynamic landmark triangles
abstract
This paper describes an efficient and accurate approach to estimate the network distance between arbitrary Internet hosts. We use three landmark hosts forming a triangle in two-dimensional space to estimate the distance between arbitrary hosts with simple trigonometrical calculations. To improve the accuracy of estimation, we dynamically choose the best triangle for a given pair of hosts using a heuristic algorithm. Our experiments show that this approach achieves both lower computational and network probing cost over the classic landmarks-based approach while producing more accurate estimates.
Zhihua Wen, Michael Rabinovich
SIGMETRICS1
2007 An Analysis of Performance Interference Effects in Virtual Environments
abstract
Virtualization is an essential technology in modern datacenters. Despite advantages such as security isolation, fault isolation, and environment isolation, current virtualization techniques do not provide effective performance isolation between virtual machines (VMs). Specifically, hidden contention for physical resources impacts performance differently in different workload configurations, causing significant variance in observed system throughput. To this end, characterizing workloads that generate performance interference is important in order to maximize overall utility. In this paper, we study the effects of performance interference by looking at system-level workload characteristics. In a physical host, we allocate two VMs, each of which runs a sample application chosen from a wide range of benchmark and real-world workloads. For each combination, we collect performance metrics and runtime characteristics using an instrumented Ken hypervisor. Through subsequent analysis of collected data, we identify clusters of applications that generate certain types of performance interference. Furthermore, we develop mathematical models to predict the performance of a new application from its workload characteristics. Our evaluation shows our techniques were able to predict performance with average error of approximately 5%
Younggyun Koh, Rob C. Knauerhase, Paul Brett, Mic Bowman, Zhihua Wen, Calton Pu
ISPASS5
2007 Facilitating focused internet measurements
abstract
This paper describes our implementation of and initial experiences with DipZoom (for "Deep Internet Performance Zoom"), a novel approach to provide focused, on-demand Internet measurements. Unlike existing approaches that face a difficult challenge of building a measurement platform with sufficiently diverse measurements and measuring hosts, DipZoom implements a matchmaking service instead, which uses P2P concepts to bring together experimenters in need of measurements with external measurement providers. DipZoom offers the following two main contributions. First,since it is just a facilitator for an open community of participants, it promises unprecedented availability of diverse measurements and measuring points. Second, it can be used as a veneer over existing measurement platforms, automating the planning and execution of complex measurements.
Zhihua Wen, Sipat Triukose, Michael Rabinovich
SIGMETRICS1
2006 DipZoom: The Internet Measurements Marketplace
abstract
We describe DipZoom (for "deep Internet performance zoom"), an approach to provide focused, on-demand Internet measurements. Unlike existing approaches that face a difficult challenge of building a measurement platform with sufficiently diverse measurements and measuring hosts, DipZoom offers a matchmaking service instead, which uses P2P concepts to bring together experimenters in need of measurements with external measurement providers. It then harnesses market forces to orchestrate the supply and demand sides in the resulting open eco-system. This paper outlines the overall design of DipZoom, and discusses payment, trust and security issues in the resulting open system.
Michael Rabinovich, Sipat Triukose, Zhihua Wen, Limin Wang 0010
INFOCOM3
2004 Evaluating Similarity Measures for Software Decompositions
abstract
One of the central questions that a similarity measure for software decompositions has to address is whether to consider discrepancies in terms of the nodes of a particular decomposition, or assess similarity based on differences in clustering the edges of the system's dependency graph. We argue that considering nodes or edges in isolation is too one-sided. We outline shortcomings of previous approaches, and introduce the first dissimilarity measure that takes both nodes and edges into account. We also present experiments on real and synthetic data sets that illustrate the differences between various measures.
Zhihua Wen, Vassilios Tzerpos
ICSM1