Xinglin Wang

dblp:02/1010 · DBLP profile ↗
← Back
25ranked-venue papers
7as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 3 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Do Retrieval Augmented Language Models Know When They Don't Know?
abstract
Existing large language models (LLMs) occasionally generate plausible yet factually incorrect responses, known as hallucinations. Two main approaches have been proposed to mitigate hallucinations: retrieval-augmented language models (RALMs) and refusal post-training. However, current research predominantly focuses on their individual effectiveness while overlooking the evaluation of the refusal capability of RALMs. Ideally, if RALMs know when they do not know, they should refuse to answer. In this study, we ask the fundamental question: Do RALMs know when they don’t know? Specifically, we investigate three questions. First, are RALMs well calibrated with respect to different internal and external knowledge states? We examine the influence of various factors. Contrary to expectations, when all retrieved documents are irrelevant, RALMs still tend to refuse questions they could have answered correctly. Next, given the model's pronounced over-refusal behavior, we raise a second question: How does a RALM's refusal ability align with its calibration quality? Our results show that the over-refusal problem can be mitigated through in-context fine-tuning. However, we observe that improved refusal behavior does not necessarily imply better calibration or higher overall accuracy. Finally, we ask: Can we combine refusal-aware RALMs with uncertainty-based answer abstention to mitigate over-refusal? We develop a simple yet effective refusal mechanism for refusal-post-trained RALMs that improves their overall answer quality by balancing refusal and correct answers. Our study provides a more comprehensive understanding of the factors influencing RALM behavior. Meanwhile, we emphasize that uncertainty estimation for RALMs remains an open problem deserving deeper investigation.
Youchao Zhou, Heyan Huang, Xinglin Wang, Shumin Shi, Yang Deng 0002
AAAI5
2026 LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient
abstract
Peiwen Yuan, Shaoxiong Feng, Yiwei Li, Xinglin Wang, Yueqi Zhang, Jiayi Shi, Chuyi Tan, Boyuan Pan, Yao Hu, Kan Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Peiwen Yuan, Shaoxiong Feng, Yiwei Li 0001, Xinglin Wang, Chuyi Tan, Boyuan Pan, Yao Hu 0002, Kan Li 0001
ACL (1)4
2026 MODE+: A benchmark and a probe into multimodal open-domain dialogue evaluation
Hang Yin 0007, Xinglin Wang, Pinren Lu, Bin Sun 0004, Peiwen Yuan, Kan Li 0001
Neurocomputing2
2025 From Sub-Ability Diagnosis to Human-Aligned Generation: Bridging the Gap for Text Length Control via MarkerGen
abstract
Peiwen Yuan, Chuyi Tan, Shaoxiong Feng, Yiwei Li, Xinglin Wang, Yueqi Zhang, Jiayi Shi, Boyuan Pan, Yao Hu, Kan Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Peiwen Yuan, Chuyi Tan, Shaoxiong Feng, Yiwei Li 0001, Xinglin Wang, Boyuan Pan, Yao Hu 0002, Kan Li 0001
ACL (1)5
2025 Beyond One-Size-Fits-All: Tailored Benchmarks for Efficient Evaluation
abstract
Evaluating models on large benchmarks can be very resource-intensive, especially during a period of rapid model evolution. Existing efficient evaluation methods estimate the performance of target models by testing them on a small, static coreset derived from the publicly available evaluation results of source models, which are separate from the target models. However, these approaches rely on the assumption that target models have high prediction consistency with source models, which doesn’t generalize well in practice. To fill this gap, we propose TailoredBench, a method that conducts customized evaluation tailored to each target model. Specifically, a Global-coreset is first constructed as a probe to identify the most consistent source models for each target model with an adaptive source model selection strategy. Afterwards, a scalable K-Medoids clustering algorithm is proposed to extend the Global-coreset to a tailored Native-coreset for each target model. According to the predictions on respective Native-coreset, we estimate the overall performance of target models with a calibrated estimation strategy. Comprehensive experiments on five benchmarks across over 300 models demonstrate that compared to best performing baselines, TailoredBench achieves an average reduction of 31.4% in MAE of accuracy estimates under the same inference budgets, showcasing strong effectiveness and generalizability.
Peiwen Yuan, Shaoxiong Feng, Yiwei Li 0001, Xinglin Wang, Chuyi Tan, Boyuan Pan, Yao Hu 0002, Kan Li 0001
ACL (1)5
2025 UniCBE: An Uniformity-driven Comparing Based Evaluation Framework with Unified Multi-Objective Optimization
abstract
Human preference plays a significant role in measuring large language models and guiding them to align with human values. Unfortunately, current comparing-based evaluation (CBE) methods typically focus on a single optimization objective, failing to effectively utilize scarce yet valuable preference signals. To address this, we delve into key factors that can enhance the accuracy, convergence, and scalability of CBE: suppressing sampling bias, balancing descending process of uncertainty, and mitigating updating uncertainty. Following the derived guidelines, we propose UniCBE, a unified uniformity-driven CBE framework which simultaneously optimize these core objectives by constructing and integrating three decoupled sampling probability matrices, each designed to ensure uniformity in specific aspects. We further ablate the optimal tuple sampling and preference aggregation strategies to achieve efficient CBE. On the AlpacaEval benchmark, UniCBE saves over 17% of evaluation budgets while achieving a Pearson correlation with ground truth exceeding 0.995, demonstrating excellent accuracy and convergence. In scenarios where new models are continuously introduced, UniCBE can even save over 50% of evaluation costs, highlighting its improved scalability.
Peiwen Yuan, Shaoxiong Feng, Yiwei Li 0001, Xinglin Wang, Chuyi Tan, Boyuan Pan, Yao Hu 0002, Kan Li 0001
ICLR4
2025 CogLM: Tracking Cognitive Development of Large Language Models
abstract
Xinglin Wang, Peiwen Yuan, Shaoxiong Feng, Yiwei Li, Boyuan Pan, Heda Wang, Yao Hu, Kan Li. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Xinglin Wang, Peiwen Yuan, Shaoxiong Feng, Yiwei Li 0001, Boyuan Pan, Heda Wang, Yao Hu 0002, Kan Li 0001
NAACL (Long Papers)1
2025 Every Rollout Counts: Optimal Resource Allocation for Efficient Test-Time Scaling
abstract
Test-Time Scaling (TTS) improves the performance of Large Language Models (LLMs) by using additional inference-time computation to explore multiple reasoning paths through search. Yet how to allocate a fixed rollout budget most effectively during search remains underexplored, often resulting in inefficient use of compute at test time. To bridge this gap, we formulate test-time search as a resource allocation problem and derive the optimal allocation strategy that maximizes the probability of obtaining a correct solution under a fixed rollout budget. Within this formulation, we reveal a core limitation of existing search methods: solution-level allocation tends to favor reasoning directions with more candidates, leading to theoretically suboptimal and inefficient use of compute. To address this, we propose Direction-Oriented Resource Allocation (DORA), a provably optimal method that mitigates this bias by decoupling direction quality from candidate count and allocating resources at the direction level. To demonstrate DORA’s effectiveness, we conduct extensive experiments on challenging mathematical reasoning benchmarks including MATH500, AIME2024, and AIME2025. The empirical results show that DORA consistently outperforms strong baselines with comparable computational cost, achieving state-of-the-art accuracy. We hope our findings contribute to a broader understanding of optimal TTS for LLMs.
Xinglin Wang, Yiwei Li 0001, Shaoxiong Feng, Peiwen Yuan, Chuyi Tan, Boyuan Pan, Yao Hu 0002, Kan Li 0001
NeurIPS1
2025 Silencer: From Discovery to Mitigation of Self-Bias in LLM-as-Benchmark-Generator
abstract
LLM-as-Benchmark-Generator methods have been widely studied as a supplement to human annotators for scalable evaluation, while the potential biases within this paradigm remain underexplored. In this work, we systematically define and validate the phenomenon of inflated performance in models evaluated on their self-generated benchmarks, referred to as self-bias, and attribute it to sub-biases arising from question domain, language style, and wrong labels. On this basis, we propose Silencer, a general framework that leverages the heterogeneity between multiple generators at both the sample and benchmark levels to neutralize bias and generate high-quality, self-bias-silenced benchmark. Experimental results across various settings demonstrate that Silencer can suppress self-bias to near zero, significantly improve evaluation effectiveness of the generated benchmark (with an average improvement from 0.655 to 0.833 in Pearson correlation with high-quality human-annotated benchmark), while also exhibiting strong generalizability.
Peiwen Yuan, Yiwei Li 0001, Shaoxiong Feng, Xinglin Wang, Chuyi Tan, Boyuan Pan, Yao Hu 0002, Kan Li 0001
NeurIPS4
2025 Mind the Quote: Enabling Quotation-Aware Dialogue in LLMs via Plug-and-Play Modules
abstract
Human–AI conversation frequently relies on quoting earlier text—“check it with the formula I just highlighted”—yet today’s large language models (LLMs) lack an explicit mechanism for locating and exploiting such spans. We formalize the challenge as span-conditioned generation, decomposing each turn into the dialogue history, a set of token-offset quotation spans, and an intent utterance. Building on this abstraction, we introduce a quotation-centric data pipeline that automatically synthesizes task-specific dialogues, verifies answer correctness through multi-stage consistency checks, and yields both a heterogeneous training corpus and the first benchmark covering five representative scenarios. To meet the benchmark’s zero-overhead and parameter-efficiency requirements, we propose QuAda, a lightweight training-based method that attaches two bottleneck projections to every attention head, dynamically amplifying or suppressing attention to quoted spans at inference time while leaving the prompt unchanged and updating < 2.8% of backbone weights. Experiments across models show that QuAda is suitable for all scenarios and generalizes to unseen topics, offering an effective, plug-and-play solution for quotation-aware dialogue.
Peiwen Yuan, Yiwei Li 0001, Shaoxiong Feng, Xinglin Wang, Chuyi Tan, Boyuan Pan, Yao Hu 0002, Kan Li 0001
NeurIPS5
2024 Turning Dust into Gold: Distilling Complex Reasoning Capabilities from LLMs by Leveraging Negative Data
abstract
Large Language Models (LLMs) have performed well on various reasoning tasks, but their inaccessibility and numerous parameters hinder wide application in practice. One promising way is distilling the reasoning ability from LLMs to small models by the generated chain-of-thought reasoning paths. In some cases, however, LLMs may produce incorrect reasoning chains, especially when facing complex mathematical problems. Previous studies only transfer knowledge from positive samples and drop the synthesized data with wrong answers. In this work, we illustrate the merit of negative data and propose a model specialization framework to distill LLMs with negative samples besides positive ones. The framework consists of three progressive steps, covering from training to inference stages, to absorb knowledge from negative data. We conduct extensive experiments across arithmetic reasoning tasks to demonstrate the role of negative data in distillation from LLM.
Yiwei Li 0001, Peiwen Yuan, Shaoxiong Feng, Boyuan Pan, Bin Sun 0004, Xinglin Wang, Heda Wang, Kan Li 0001
AAAI6
2024 Integrate the Essence and Eliminate the Dross: Fine-Grained Self-Consistency for Free-Form Language Generation
abstract
Xinglin Wang, Yiwei Li, Shaoxiong Feng, Peiwen Yuan, Boyuan Pan, Heda Wang, Yao Hu, Kan Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Xinglin Wang, Yiwei Li 0001, Shaoxiong Feng, Peiwen Yuan, Boyuan Pan, Heda Wang, Yao Hu 0002, Kan Li 0001
ACL (1)1
2024 BatchEval: Towards Human-like Text Evaluation
abstract
Peiwen Yuan, Shaoxiong Feng, Yiwei Li, Xinglin Wang, Boyuan Pan, Heda Wang, Yao Hu, Kan Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Peiwen Yuan, Shaoxiong Feng, Yiwei Li 0001, Xinglin Wang, Boyuan Pan, Heda Wang, Yao Hu 0002, Kan Li 0001
ACL (1)4
2024 Generative Dense Retrieval: Memory Can Be a Burden
abstract
Peiwen Yuan, Xinglin Wang, Shaoxiong Feng, Boyuan Pan, Yiwei Li, Heda Wang, Xupeng Miao, Kan Li. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Peiwen Yuan, Xinglin Wang, Shaoxiong Feng, Boyuan Pan, Yiwei Li 0001, Heda Wang, Xupeng Miao, Kan Li 0001
EACL (1)2
2024 Focused Large Language Models are Stable Many-Shot Learners
abstract
Peiwen Yuan, Shaoxiong Feng, Yiwei Li, Xinglin Wang, Yueqi Zhang, Chuyi Tan, Boyuan Pan, Heda Wang, Yao Hu, Kan Li. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Peiwen Yuan, Shaoxiong Feng, Yiwei Li 0001, Xinglin Wang, Chuyi Tan, Boyuan Pan, Heda Wang, Yao Hu 0002, Kan Li 0001
EMNLP4
2024 Escape Sky-high Cost: Early-stopping Self-Consistency for Multi-step Reasoning
abstract
Self-consistency (SC) has been a widely used decoding strategy for chain-of-thought reasoning. Despite bringing significant performance improvements across a variety of multi-step reasoning tasks, it is a high-cost method that requires multiple sampling with the preset size. In this paper, we propose a simple and scalable sampling process, Early-Stopping Self-Consistency (ESC), to greatly reduce the cost of SC without sacrificing performance. On this basis, one control scheme for ESC is further derivated to dynamically choose the performance-cost balance for different tasks and models. To demonstrate ESC's effectiveness, we conducted extensive experiments on three popular categories of reasoning tasks: arithmetic, commonsense and symbolic reasoning over language models with varying scales. The empirical results show that ESC reduces the average number of sampling of chain-of-thought reasoning by a significant margin on six benchmarks, including MATH (-33.8%), GSM8K (-80.1%), StrategyQA (-76.8%), CommonsenseQA (-78.5%), Coin Flip (-84.2%) and Last Letters (-67.4%), while attaining comparable performances.
Yiwei Li 0001, Peiwen Yuan, Shaoxiong Feng, Boyuan Pan, Xinglin Wang, Bin Sun 0004, Heda Wang, Kan Li 0001
ICLR5
2024 Instruction Embedding: Latent Representations of Instructions Towards Task Identification
abstract
Instruction data is crucial for improving the capability of Large Language Models (LLMs) to align with human-level performance. Recent research LIMA demonstrates that alignment is essentially a process where the model adapts instructions' interaction style or format to solve various tasks, leveraging pre-trained knowledge and skills. Therefore, for instructional data, the most important aspect is the task it represents, rather than the specific semantics and knowledge information. The latent representations of instructions play roles for some instruction-related tasks like data selection and demonstrations retrieval. However, they are always derived from text embeddings, encompass overall semantic information that influences the representation of task categories. In this work, we introduce a new concept, instruction embedding, and construct Instruction Embedding Benchmark (IEB) for its training and evaluation. Then, we propose a baseline Prompt-based Instruction Embedding (PIE) method to make the representations more attention on tasks. The evaluation of PIE, alongside other embedding methods on IEB with two designed tasks, demonstrates its superior performance in accurately identifying task categories. Moreover, the application of instruction embeddings in four downstream tasks showcases its effectiveness and suitability for instruction-related tasks.
Yiwei Li 0001, Shaoxiong Feng, Peiwen Yuan, Xinglin Wang, Boyuan Pan, Heda Wang, Yao Hu 0002, Kan Li 0001
NeurIPS5
2023 Better Correlation and Robustness: A Distribution-Balanced Self-Supervised Learning Framework for Automatic Dialogue Evaluation
abstract
Turn-level dialogue evaluation models (TDEMs), using self-supervised learning (SSL) framework, have achieved state-of-the-art performance in open-domain dialogue evaluation. However, these models inevitably face two potential problems. First, they have low correlations with humans on medium coherence samples as the SSL framework often brings training data with unbalanced coherence distribution. Second, the SSL framework leads TDEM to nonuniform score distribution. There is a danger that the nonuniform score distribution will weaken the robustness of TDEM through our theoretical analysis. To tackle these problems, we propose Better Correlation and Robustness (BCR), a distribution-balanced self-supervised learning framework for TDEM. Given a dialogue dataset, BCR offers an effective training set reconstructing method to provide coherence-balanced training signals and further facilitate balanced evaluating abilities of TDEM. To get a uniform score distribution, a novel loss function is proposed, which can adjust adaptively according to the uniformity of score distribution estimated by kernel density estimation. Comprehensive experiments on 17 benchmark datasets show that vanilla BERT-base using BCR outperforms SOTA methods significantly by 11.3% on average. BCR also demonstrates strong generalization ability as it can lead multiple SOTA methods to attain better correlation and robustness.
Peiwen Yuan, Xinglin Wang, Bin Sun 0004, Yiwei Li 0001
NeurIPS2
2010 Co-Existence Analysis of LTE Micro Cell and LTE Out-Band Backhaul
abstract
In this paper, a stand-alone LTE based out-band backhaul is designed for urban area in NLOS environment, and the interference and compatibility issues relating to co-existence of LTE micro cell and co-located LTE out-band backhaul are investigated by a static system level simulator. Feasibility and recommendation of installing out-band backhaul are analyzed according to the simulation results.
Xinglin Wang, Xiaokun Yang
VTC Fall1
2010 Study on Co-Existence of Macro WCDMA Cell and Micro HSUPA Cell
abstract
In this paper, a static system level simulator is used to investigate the capacity loss in case of co-existence of macro WCDMA and micro HSUPA. The capacity loss under different frequency spacing or ACIRs is obtained. The simulation results show that the micro HSUPA cell has little impact on the macro WCDMA cell with frequency spacing of 5 MHz, while the macro WCDMA cell has larger impact on the micro HSUPA cell. Based on simulation, suggestions of frequency planning and parameter setting are given to avoid large capacity loss due to co-existence.
Xinglin Wang, Xiaokun Yang, Xiaojin Zhang 0003
VTC Spring2
2008 Exact BER Analysis for Signal Code Modulation in Wireless Communication
abstract
Signal code modulation (SCM) is a mixed analog-digital modulation technique that is proposed for relaying a digital communication signal over channels with different Signal-to-Noise Ratios (SNRs) and, in particular, provides an alternative method besides the traditional complete demodulation and re-modulation (Demod/Remod). In this paper, we derive the exact closed-form expressions for the bit error rate (BER) of SCM modulation scheme in additive white Gaussian noise (AWGN) channel and Nakagami fading channel. Furthermore, the diversity gain for SCM is defined.
Shaoqun Fan, Xinglin Wang
VTC Spring2
2007 A Simplified Layered QoS Scheduling Scheme in OFDM Networks
abstract
A simplified layered scheduling scheme is presented, which is based on utility-based scheduling, such as Max-Delay- Utility (MDU) scheduling in this paper. In this scheme, the scheduling is divided into two steps-- macro and micro scheduling. In macro step, the utility functions of traffics are defined, according to which the scheduling order of various services is determined. Then in micro step, the scheduling is among all users of the traffic type which is determined in macro step. By simulation in a multi-user orthogonal frequency division multiplexing (OFDM) network, it is demonstrated that the simplified layered MDU scheduling can effectively handle multiple traffic types with diverse QoS requirements and achieve almost the same performance from the view of utility, delay, throughput and fairness as MDU does. Moreover, the simplified layered MDU scheduling has much lower computational complexity than the MDU scheduling and almost the same utility values.
Zhiqiang He 0001, Weiling Wu, Xinglin Wang
VTC Fall4
2006 An Improved Detection Based on Lattice Reduction in MIMO Systems
abstract
In this paper, list detection based on lattice reduction (LDLR) in MIMO systems is proposed. By generating a list of candidate transmit symbol vectors we can improve the performance of linear detection based on lattice reduction and successive interference cancellation based on lattice reduction significantly. The performance improvement is demonstrated by computer simulations, which also show that a list with two entries is enough for the MMSE based detection to approach the maximum likelihood (ML) detection at high signal-to-noise ratio (SNR) in case of (4,4) MIMO
Xinglin Wang, Zhiqiang He 0001, Kai Niu 0001, Weiling Wu
PIMRC1
2006 List Sphere Decoding Combined with Linear Detection-Based Iterative Soft Interference Cancellation Via Exit Chart
abstract
Iterative list sphere decoding (LSD) can achieve near capacity on a multiple-antenna channel, however, with a rather high complexity. In order to reduce the complexity, in this paper, linear detection-based iterative soft interference cancellation (ISIC) instead of LSD is applied at the second and later iterations. We propose to exploit extrinsic information transfer (EXIT) chart analysis to select a specific ISIC to alleviate large performance degradation. Here MF-based and MMSE-base iterative soft interference cancellations (ISIC) are selected to be combined with the LSD. Simulation shows that the iterative combined detections have little signal-to-noise ratio (SNR) loss and a much lower complexity in comparison with iterative LSD
Xinglin Wang, Kai Niu 0001, Zhiqiang He 0001, Weiling Wu
PIMRC1
2006 Decision Feedback Aided Detection Based on Lattice Reduction in MIMO Systems
abstract
In this paper, a novel MIMO detection named decision feedback aided detection based on lattice reduction (DFDLR) is proposed. Lattice reduction has been proposed to be exploited in signal detection in multiple antenna systems. However, lattice reduction transforms the QAM constellation cube to be a parallelotope, and it is difficult to determine the transformed QAM boundaries, which results in sub-optimal performance with simple quantization. We propose to exploit decision feedback to enhance the QAM boundary control. Simulations show that decision feedback can improve the performance of lattice reduction-based linear detection (LD) and lattice reduction-based successive interference cancellation (SIC) significantly especially with zero forcing (ZF) criterion. It is also shown by simulation that the proposed scheme in conjunction with SIC can approach maximum likelihood (ML) detection with MMSE criterion at high signal-to-noise ratio (SNR)
Xinglin Wang, Kai Niu 0001, Weiling Wu, Martin Weckerle
VTC Spring1