Chenxu Yang

dblp:316/8012 · DBLP profile ↗
← Back
21ranked-venue papers
11as first author
21since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 6 first-author · 11 since 2021Theory of computation · 8 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Test-time Prompt Intervention
abstract
Test-time compute has led to remarkable success in the large language model (LLM) community, particularly for complex tasks, where longer chains of thought (CoTs) are generated to enhance reasoning capabilities. However, growing evidence reveals that such reasoning models often produce CoTs plagued by excessive redundancy, including repetitive verification steps and unnecessary reasoning shifts. The root cause lies in post-training of them that overly rely on outcome reward paradigms, as the data of process reward paradigms, which regulate intermediate reasoning steps, is difficult to construct at scale. To address this, we propose PI, a novel framework for Test-time Prompt Intervention. PI provides an interface to dynamically guide and regulate reasoning paths during inference through timely (When module) and proper (How module) interventions and post-intervention sampling (Which module). This allows human problem-solving expertise and cognitive science principles to be seamlessly integrated into LLMs’ reasoning processes, enhancing controllability and interpretability. Extensive experiments across multiple models and datasets demonstrate that PI significantly shortens CoTs while reducing hallucination, yielding more concise and reliable reasoning.
Chenxu Yang, Qingyi Si, Mz Dai, Dingyu Yao, Mingyu Zheng, Zheng Lin 0001, Weiping Wang 0005
AAAI1
2026 Breaking the Trade-Off Between Faithfulness and Expressiveness for Large Language Models
abstract
Grounding responses in external knowledge represents an effective strategy for mitigating hallucinations in Large Language Models (LLMs). However, current LLMs struggle to seamlessly integrate knowledge while simultaneously maintaining faithfulness (or fidelity) and expressiveness, capabilities that humans naturally possess. This limitation results in outputs that either lack support from external knowledge, thereby compromising faithfulness, or appear overly verbose and unnatural, thus sacrificing expressiveness. In this work, to break the trade-off between faithfulness and expressiveness, we propose Collaborative Decoding (CoDe), a novel approach that dynamically integrates output probabilities generated with and without external knowledge. This integration is guided by distribution divergence and model confidence, enabling the selective activation of relevant and reliable expressions from the model's internal parameters. Furthermore, we introduce a knowledge-aware reranking mechanism that prevents over-reliance on prior parametric knowledge while ensuring proper utilization of provided external information. Through comprehensive experiments, our plug-and-play CoDe framework demonstrates superior performance in enhancing faithfulness without compromising expressiveness across diverse LLMs and evaluation metrics, validating both its effectiveness and generalizability.
Chenxu Yang, Qingyi Si, Lanrui Wang, Zheng Lin 0001
AAAI1
2026 VecInfer: Efficient LLM Inference with Low-Bit KV Cache via Outlier-Suppressed Vector Quantization
abstract
Dingyu Yao, Chenxu Yang, Zhengyang Tong, Zheng Lin, Wei Liu, Jian Luan, Weiping Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Dingyu Yao, Chenxu Yang, Zhengyang Tong, Zheng Lin 0001, Wei Liu 0302, Jian Luan 0001, Weiping Wang 0005
ACL (1)2
2026 Complete graphs without proper subgraphs
Mengya He, Jinxia Liang, Colton Magnant, Chenxu Yang
Discret. Appl. Math.4
2026 Ramsey numbers avoiding properly colored cycles
Jinxia Liang, Colton Magnant, Pouria Salehi Nowbandegani, Meiqin Wei, Chenxu Yang
Discret. Appl. Math.5
2025 Sibyl: Empowering Empathetic Dialogue Generation in Large Language Models via Sensible and Visionary Commonsense Inference
abstract
Recently, there has been a heightened interest in building chatbots based on Large Language Models (LLMs) to emulate human-like qualities in multi-turn conversations. Despite having access to commonsense knowledge to better understand the psychological aspects and causality of dialogue context, even these powerful LLMs struggle to achieve the goals of empathy and emotional support. Current commonsense knowledge derived from dialogue contexts is inherently limited and often fails to adequately anticipate the future course of a dialogue. This lack of foresight can mislead LLMs and hinder their ability to provide effective support. In response to this challenge, we present an innovative framework named Sensible and Visionary Commonsense Knowledge (Sibyl). Designed to concentrate on the immediately succeeding dialogue, this paradigm equips LLMs with the capability to uncover the implicit requirements of the conversation, aiming to elicit more empathetic responses. Experimental results demonstrate that incorporating our paradigm for acquiring commonsense knowledge into LLMs comprehensively enhances the quality of their responses.
Lanrui Wang, Chenxu Yang, Zheng Lin 0001, Hongyin Tang, Yanan Cao 0001, Jingang Wang, Weiping Wang 0005
COLING3
2025 Weights-Rotated Preference Optimization for Large Language Models
abstract
Chenxu Yang, Ruipeng Jia, Mingyu Zheng, Naibin Gu, Zheng Lin, Siyuan Chen, Weichong Yin, Hua Wu, Weiping Wang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Chenxu Yang, Ruipeng Jia, Mingyu Zheng, Naibin Gu, Zheng Lin 0001, Weichong Yin, Hua Wu 0003, Weiping Wang 0005
EMNLP1
2025 Vision-Guided Acoustic Localization with Decoupled Inference for Moving Speakers
Yidi Li 0001, Kairan Zhang, Chenxu Yang, Chongwei Yan, Rongshan Gao, Mingliang Dou
ICIC (19)3
2025 Categorical Attention: Fine-grained Language-guided Noise Filtering Network for Occluded Person Re-Identification
abstract
Person Re-Identification (ReID) aims to match individuals across different camera views, but occlusions in real-world scenarios, such as vehicles or crowds, hinder feature extraction and matching. Current occluded ReID methodologies typically leverage visual augmentation techniques in an attempt to mitigate the disruptive effects of occlusion-induced noise. However, relying solely on visual data fail to effectively filter out occlusion noise. In this paper, we introduce the Fine-grained Language-guided Noise Filtering Network (FLaN-Net) for occluded ReID. FLaN-Net innovatively employs categorical attention mechanism to generate adaptive tokens that capture the following three distinct types of visual information: comprehensive descriptions of individuals, detailed visible attributes, and characteristics of occluding objects. Subsequently, a cross-attention mechanism aligns these prompts with the image, guiding the model to focus on relevant regions. To generate robust and discriminative features for occluded pedestrians, we further introduce a dynamic weighting fusion module that integrates visual, textual, and cross-attention features based on their reliability. Experimental results demonstrate that FLaN-Net outperforms existing methods on occluded ReID benchmarks, offering a robust solution for challenging real-world conditions.
Dayan Wu, Chenxu Yang, Qinghang Su, Zheng Lin 0001
IJCAI3
2025 Towards Realistic Generation: A Multi-Task Agent for Imitating Diverse Character Linguistic Styles
abstract
The advent of large language models (LLMs) has significantly propelled the advancement of Role-Playing Agents (RPAs). However, current Role-Playing Agents predominantly focus on mimicking a character’s fundamental attributes while neglecting the replication of linguistic style, and they are incapable of effectively replicating characters when performing tasks beyond multi-turn dialogues, which results in generated responses that lack authenticity. The reason current RPAs lack this capability is due to the nature of existing character datasets, which lack collections of character quotations and are limited to multi-turn dialogue tasks, constraining the RPA’s performance across other task domains and failing to mimic a character’s linguistic style. To address this gap, we developed a multi-task role-playing dataset named MRstyle, which encompasses a substantial number of real individuals along with their quotations and covers seven different tasks. On this basis, we develop StyleRPA, a Multi-Task Role-Playing Agent (MRPA) that significantly outperforms recent open-source LLMs and RPAs baselines on 7 tasks including Dialogue, Dictionary, Composition, Story Generation, Product Description, Music Commentary, and Open Question Answering.
Qingyi Si, Chenxu Yang, Zheng Lin 0001, Yunzhi Liang, Siyang Tao, Weiping Wang 0005
IJCNN3
2025 S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models
abstract
As Test-Time Scaling emerges as an active research focus in the large language model community, advanced post-training methods increasingly emphasize extending chain-of-thought (CoT) generation length, thereby enhancing reasoning capabilities to approach Deepseek R1-like reasoning models. However, recent studies reveal that reasoning models (even Qwen3) consistently exhibit excessive thought redundancy in CoT generation. This overthinking issue arises from the inherent limitations of conventional outcome-reward reinforcement learning, which systematically overlooks the regulation of intermediate reasoning processes. This paper introduces Serial-Group Decaying-Reward Policy Optimization (S-GRPO), a novel reinforcement learning paradigm that enables models to implicitly evaluate the sufficiency of intermediate reasoning steps, thereby facilitating early exit in CoT generation. Unlike GRPO, which samples multiple possible reasoning paths in parallel (parallel group), S-GRPO only samples one reasoning path and serially selects multiple temporal positions from the path to exit thinking and directly generate answers (serial group). For correct answers within a serial group, rewards gradually decrease based on the exit positions along the reasoning path from front to back. This design encourages the model to produce more accurate and concise thoughts, while also incentivizing early thinking termination when appropriate. Empirical evaluations demonstrate that S-GRPO is compatible with state-of-the-art reasoning models, including Qwen3 and Deepseek-distill. Across diverse benchmarks such as GSM8K, AIME 2024, AMC 2023, MATH-500, and GPQA Diamond, S-GRPO achieves a substantial reduction in sequence length (40.4%~61.1%) while simultaneously improving accuracy (absolute 0.72%~3.92%).
Muzhi Dai, Chenxu Yang, Qingyi Si
NeurIPS2
2025 Fault-tolerance in distance-edge-monitoring sets
Chenxu Yang, Yaping Mao, Ralf Klasing, Yuzhi Xiao
Acta Informatica1
2025 Nordhaus-Gaddum-type results for monitoring edge-geodetic number of graphs
Fanfan Wang, Chenxu Yang
Discret. Appl. Math.3
2025 Constructing disjoint Steiner trees in Sierpiński graphs
abstract
Let $G$ be a graph and $S\subseteq V(G)$ with $|S|\geq 2$. Then the trees $T_1, T_2, \cdots, T_\ell$ in $G$ are \emph{internally disjoint Steiner trees} connecting $S$ (or $S$-Steiner trees) if $E(T_i) \cap E(T_j )=\emptyset$ and $V(T_i)\cap V(T_j)=S$ for every pair of distinct integers $i,j$, $1 \leq i, j \leq \ell$. Similarly, if we only have the condition $E(T_i) \cap E(T_j )=\emptyset$ but without the condition $V(T_i)\cap V(T_j)=S$, then they are \emph{edge-disjoint Steiner trees}. The \emph{generalized $k$-connectivity}, denoted by $κ_k(G)$, of a graph $G$, is defined as $κ_k(G)=\min\{κ_G(S)|S \subseteq V(G) \ \textrm{and} \ |S|=k \}$, where $κ_G(S)$ is the maximum number of internally disjoint $S$-Steiner trees. The \emph{generalized local edge-connectivity} $λ_{G}(S)$ is the maximum number of edge-disjoint Steiner trees connecting $S$ in $G$. The {\it generalized $k$-edge-connectivity} $λ_k(G)$ of $G$ is defined as $λ_k(G)=\min\{λ_{G}(S)\,|\,S\subseteq V(G) \ and \ |S|=k\}$. These measures are generalizations of the concepts of connectivity and edge-connectivity, and they and can be used as measures of vulnerability of networks. It is, in general, difficult to compute these generalized connectivities. However, there are precise results for some special classes of graphs. In this paper, we obtain the exact value of $λ_{k}(S(n,\ell))$ for $3\leq k\leq \ell^n$, and the exact value of $κ_{k}(S(n,\ell))$ for $3\leq k\leq \ell$, where $S(n, \ell)$ is the Sierpiński graphs with order $\ell^n$. As a direct consequence, these graphs provide additional interesting examples when $λ_{k}(S(n,\ell))=κ_{k}(S(n,\ell))$. We also study the some network properties of Sierpiński graphs. Steiner Tree; Generalized Connectivity; Sierpiński Graph
Chenxu Yang, Ping Li 0025, Yaping Mao, Eddie Cheng 0001, Ralf Klasing
Fundam. Informaticae1
2024 On the distance-edge-monitoring numbers of graphs
Chenxu Yang, Ralf Klasing, Yaping Mao, Xingchao Deng
Discret. Appl. Math.1
2024 Perturbation Results for Distance-edge-monitoring Numbers
abstract
Foucaud et al. recently introduced and initiated the study of a new graph-theoretic concept in the area of network monitoring. Given a graph G = ( V( G), E( G)), a set M ⊆ V( G) is a distance-edge-monitoring set if for every edge e ∈ E( G), there is a vertex x ∈ M and a vertex y ∈ V( G) such that the edge e belongs to all shortest paths between x and y. The smallest size of such a set in G is denoted by dem( G). Denoted by G – e (resp. G\ u) the subgraph of G obtained by removing the edge e from G (resp. a vertex u together with all its incident edges from G). In this paper, we first show that dem( G – e) – dem( G) ≤ 2 for any graph G and edge e ∈ E( G). Moreover, the bound is sharp. Next, we construct two graphs G and H to show that dem( G) – dem( G\ u) and dem( H \ v) – dem( H) can be arbitrarily large, where u ∈ V( G) and v ∈ V( H). We also study the relation between dem( H) and dem( G), where H is a subgraph of G. In the end, we give an algorithm to judge whether the distance-edge-monitoring set still remain in the resulting graph when any edge of a graph G is deleted.
Chenxu Yang, Ralf Klasing, Changxiang He, Yaping Mao
Fundam. Informaticae1
2024 Monitoring the edges of a graph using distances with given girth
abstract
International audience
Chenxu Yang, Sun-Yuan Hsieh, Yaping Mao, Ralf Klasing
J. Comput. Syst. Sci.1
2023 Multi-level Adaptive Contrastive Learning for Knowledge Internalization in Dialogue Generation
abstract
Chenxu Yang, Zheng Lin, Lanrui Wang, Chong Tian, Liang Pang, Jiangnan Li, Qirong Ho, Yanan Cao, Weiping Wang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Chenxu Yang, Zheng Lin 0001, Lanrui Wang, Chong Tian, Qirong Ho, Yanan Cao 0001, Weiping Wang 0005
EMNLP1
2023 G&L: An Attention-based Model for Improving Prefetching in Solid-state Drives
abstract
Solid-state drives (SSDs) have become the storage device for many personal computers and enterprise servers. However, the access speed of SSDs is still a bottleneck for computer systems. With prefetching, SSDs are able to predict future requests before moving data to a faster storage unit. Nevertheless, the interleaved access flow of applications poses challenges for existing prefetchers. We propose a model named G&L consisting of two components, namely a generative pre-training (GPT) prefetcher and a logical block address (LBA)-I/O size table. The GPT prefetcher applies a decoder-only framework with a self-attention mechanism to forecast the next LBA. The LBA-I/O size table maps LBAs to their most frequently appeared I/Osize for prediction. To evaluate the entire performance of the G&L model, we propose an algorithm to simulate the buffer when the I/O size changes dynamically. Experiments show that the accuracy, coverage, and f1 score of the GPT prefetcher surpass the baselines in most cases. The prediction accuracy of the LBA-I/O size table exceeds the baseline on all datasets. The G&L model's prefetching accuracy, coverage, and f1 score also outperform the baseline considering variable I/O size.
Chenxu Yang, Xin Man, Jie Shao 0001
IJCNN1
2022 TAKE: Topic-shift Aware Knowledge sElection for Dialogue Generation
abstract
Knowledge-grounded dialogue generation consists of two subtasks: knowledge selection and response generation. The knowledge selector generally constructs a query based on the dialogue context and selects the most appropriate knowledge to help response generation. Recent work finds that realizing who (the user or the agent) holds the initiative and utilizing the role-initiative information to instruct the query construction can help select knowledge. It depends on whether the knowledge connection between two adjacent rounds is smooth to assign the role. However, whereby the user takes the initiative only when there is a strong semantic transition between two rounds, probably leading to initiative misjudgment. Therefore, it is necessary to seek a more sensitive reason beyond the initiative role for knowledge selection. To address the above problem, we propose a Topic-shift Aware Knowledge sElector(TAKE). Specifically, we first annotate the topic shift and topic inheritance labels in multi-round dialogues with distant supervision. Then, we alleviate the noise problem in pseudo labels through curriculum learning and knowledge distillation. Extensive experiments on WoW show that TAKE performs better than strong baselines.
Chenxu Yang, Zheng Lin 0001, Fandong Meng, Weiping Wang 0005, Lanrui Wang, Jie Zhou 0016
COLING1
2022 Transformer-Based Cache Replacement Policy Learning
Chenxu Yang, Jie Shao 0001
WISE2