Wenxiang Chen

dblp:01/8509 · DBLP profile ↗
← Back
24ranked-venue papers
8as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 7 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
abstract
Despite rapid development, large language models (LLMs) still encounter challenges in multi-turn decision-making tasks (i.e., agent tasks) like web shopping and browser navigation, which require making a sequence of intelligent decisions based on environmental feedback. Previous work for LLM agents typically relies on elaborate prompt engineering or fine-tuning with expert trajectories to improve performance. In this work, we take a different perspective: we explore constructing process reward models (PRMs) to evaluate each decision and guide the agent's decision-making process. Unlike LLM reasoning, where each step is scored based on correctness, actions in agent tasks do not have a clear-cut correctness. Instead, they should be evaluated based on their proximity to the goal and the progress they have made. Building on this insight, we propose a re-defined PRM for agent tasks, named AgentPRM, to capture both the interdependence between sequential decisions and their contribution to the final goal. This enables better progress tracking and exploration-exploitation balance. To scalably obtain labeled data for training AgentPRM, we employ a Temporal Difference-based (TD-based) estimation method combined with Generalized Advantage Estimation (GAE), which proves more sample-efficient than prior methods. Extensive experiments across different agentic tasks show that AgentPRM is over 8× more compute-efficient than baselines, and it demonstrates robust improvement when scaling up test-time compute. Moreover, we perform detailed analyses to show how our method works and offer more insights, e.g., applying AgentPRM to the reinforcement learning of LLM agents.
Zhiheng Xi, Chenyang Liao, Zhihao Zhang 0002, Wenxiang Chen, Binghai Wang, Senjie Jin, Yuhao Zhou 0005, Jian Guan 0002, Wei Wu 0014, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
WWW5
2026 Enhancing Multimodal Compositional Understanding of Vision-Language Models With Semantic Decoupling and Feature Coupling
Wenxiang Chen, Housheng Su
IEEE Signal Process. Lett.1
2025 AgentGym: Evaluating and Training Large Language Model-based Agents across Diverse Environments
abstract
Zhiheng Xi, Yiwen Ding, Wenxiang Chen, Boyang Hong, Honglin Guo, Junzhe Wang, Xin Guo, Dingwen Yang, Chenyang Liao, Wei He, Songyang Gao, Lu Chen, Rui Zheng, Yicheng Zou, Tao Gui, Qi Zhang, Xipeng Qiu, Xuanjing Huang, Zuxuan Wu, Yu-Gang Jiang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhiheng Xi, Yiwen Ding, Wenxiang Chen, Boyang Hong, Honglin Guo, Junzhe Wang 0001, Dingwen Yang, Chenyang Liao, Wei He 0024, Songyang Gao, Lu Chen 0001, Yicheng Zou, Tao Gui, Qi Zhang 0001, Xipeng Qiu, Xuanjing Huang 0001, Zuxuan Wu, Yu-Gang Jiang 0001
ACL (1)3
2025 Progressive Training of Transformer for Knowledge Graph Completion Tasks
Wenxiang Chen, Housheng Su
NLPCC (1)1
2025 The rise and potential of large language model based agents: a survey
Zhiheng Xi, Wenxiang Chen, Wei He 0024, Yiwen Ding, Boyang Hong, Ming Zhang 0030, Junzhe Wang 0001, Senjie Jin, Enyu Zhou, Xiaoran Fan, Xiao Wang 0001, Limao Xiong, Yuhao Zhou 0005, Weiran Wang 0003, Changhao Jiang, Yicheng Zou, Zhangyue Yin, Shihan Dou, Rongxiang Weng, Wenjuan Qin, Yongyan Zheng, Xipeng Qiu, Xuanjing Huang 0001, Qi Zhang 0001, Tao Gui
Sci. China Inf. Sci.2
2025 Multi-Label Feature Selection With Missing Features via Implicit Label Replenishment and Positive Correlation Feature Recovery
abstract
Multi-label feature selection can effectively solve the curse of dimensionality problem in multi-label learning. Existing multi-label feature selection methods mostly handle multi-label data without missing features. However, in practical applications, multi-label data with missing features exist widely, and most existing multi-label feature selection methods are not directly applicable. Therefore, we propose a feature selection method for multi-label data with missing features. First, we propose a method to extract implicit label information from the feature space to replenish the binary label information. Second, we learn the positive correlation between features to construct a feature correlation recovery matrix to recover missing features. Finally, we design a sparse model-based multi-label feature selection method for processing multi-label data with missing features and prove the convergence of this method. Comparative experiments with existing feature selection methods demonstrate the effectiveness of our method.
Jianhua Dai 0003, Wenxiang Chen
IEEE Trans. Knowl. Data Eng.2
2025 Instance-Dependent Incomplete Multi-Label Feature Selection by Fuzzy Tolerance Relation and Fuzzy Mutual Implication Granularity
abstract
Multi-label feature selection is an effective approach to mitigate the high-dimensional feature problem in multi-label learning. Most existing multi-label feature selection methods either assume that the data is complete, or that either the features or the labels are incomplete. So far, there are few studies on multi-label data with missing features and labels. In many cases, missing features in instances of multi-label data often lead to missing labels, which is ignored by existing studies. We define this type of data as instance-dependent incomplete multi-label data. In this paper, we propose a feature selection method for instance-dependent incomplete multi-label data. Firstly, we use the positive correlations between features to reconstruct the feature space, thereby recovering missing values and enhancing non-missing values. Secondly, we use fuzzy tolerance relation to guide label recovery, and utilize fuzzy mutual implication granularity to impose structural constraint on the projection matrix. Thirdly, we achieve feature selection by eliminating the impact of incomplete instances and imposing sparse regularization on the projection matrix. Finally, we provide a convergent solution for the proposed feature selection framework. Comparative experiments with existing multi-label feature selection methods show that our method can perform effective feature selection on instance-dependent incomplete multi-label data.
Jianhua Dai 0003, Wenxiang Chen, Witold Pedrycz
IEEE Trans. Knowl. Data Eng.2
2024 ORTicket: Let One Robust BERT Ticket Transfer across Different Tasks
abstract
Pretrained language models can be applied for various downstream tasks but are susceptible to subtle perturbations. Most adversarial defense methods often introduce adversarial training during the fine-tuning phase to enhance empirical robustness. However, the repeated execution of adversarial training hinders training efficiency when transitioning to different tasks. In this paper, we explore the transferability of robustness within subnetworks and leverage this insight to introduce a novel adversarial defense method ORTicket, eliminating the need for separate adversarial training across diverse downstream tasks. Specifically, (i) pruning the full model using the MLM task (the same task employed for BERT pretraining) yields a task-agnostic robust subnetwork(i.e., winning ticket in Lottery Ticket Hypothesis); and (ii) fine-tuning this subnetwork for downstream tasks. Extensive experiments demonstrate that our approach achieves comparable robustness to other defense methods while retaining the efficiency of traditional fine-tuning.This also confirms the significance of selecting MLM task for identifying the transferable robust subnetwork. Furthermore, our method is orthogonal to other adversarial training approaches, indicating the potential for further enhancement of model robustness.
Yuhao Zhou 0005, Wenxiang Chen, Zhiheng Xi, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
LREC/COLING2
2024 Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning
abstract
In this paper, we propose R$^3$: Learning Reasoning through Reverse Curriculum Reinforcement Learning (RL), a novel method that employs only outcome supervision to achieve the benefits of process supervision for large language models. The core challenge in applying RL to complex reasoning is to identify a sequence of actions that result in positive rewards and provide appropriate supervision for optimization. Outcome supervision provides sparse rewards for final results without identifying error locations, whereas process supervision offers step-wise rewards but requires extensive manual annotation. R$^3$ overcomes these limitations by learning from correct demonstrations. Specifically, R$^3$ progressively slides the start state of reasoning from a demonstration’s end to its beginning, facilitating easier model exploration at all stages. Thus, R$^3$ establishes a step-wise curriculum, allowing outcome supervision to offer step-level signals and precisely pinpoint errors. Using Llama2-7B, our method surpasses RL baseline on eight reasoning tasks by $4.1$ points on average. Notably, in program-based reasoning, 7B-scale models perform comparably to larger models or closed-source models with our R$^3$.
Zhiheng Xi, Wenxiang Chen, Boyang Hong, Senjie Jin, Wei He 0024, Yiwen Ding, Shichun Liu, Junzhe Wang 0001, Honglin Guo, Xiaoran Fan, Yuhao Zhou 0005, Shihan Dou, Xiao Wang 0001, Xinbo Zhang, Peng Sun 0006, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ICML2
2024 Feature selection based on neighborhood complementary entropy for heterogeneous data
Wenxiang Chen, Liyun Xia
Inf. Sci.2
2024 A novel multi-label feature selection method based on knowledge consistency-independence index
Xiangbin Liu, Heming Zheng, Wenxiang Chen, Liyun Xia, Jianhua Dai 0003
Inf. Sci.3
2024 Multilabel Feature Selection Based on Fuzzy Mutual Information and Orthogonal Regression
abstract
With the increase of high-dimensional multilabel data, multilabel feature selection (MFS) has received more and more widespread attention. Embedded feature selection methods have been widely studied due to their high efficiency and low computational cost. Fuzzy mutual information, as an effective tool for processing continuous features, is widely used in filter feature selection, which results in many repeated entropy calculations. Most of the existing multilabel embedded feature selection methods are based on least squares regression, which loses a lot of statistical and structural information. To solve the above-mentioned problems, we established an optimization framework based on fuzzy mutual information that considers global correlation to obtain the weight of each feature. Under this framework, many repeated entropy operations are avoided. Then, the weight of each feature is introduced into the orthogonal regression optimization framework as prior knowledge. Finally, two optimization frameworks are comprehensively considered for MFS. Furthermore, considering the characteristics of multilabel data, we extend the proposed method to feature-specific MFS. We conducted sufficient experiments to demonstrate the efficiency of our proposed method.
Jianhua Dai 0003, Wenxiang Chen, Chucai Zhang
IEEE Trans. Fuzzy Syst.3
2023 LDSSNV: A Linkage Disequilibrium-Based Method for the Detection of Somatic Single-Nucleotide Variants
abstract
Single nucleotide variants (SNVs) are very common in human genome and pose a significant effect on cellular proliferation and tumorigenesis in various cancers. Somatic variant and germline variant are the two forms of SNVs. They are the major drivers of inherited diseases and acquired tumors respectively. A reasonable analysis of the next generation sequencing data profiles from cancer genomes could provide crucial information for cancer diagnosis and treatment. Accurate detection of SNVs and distinguishing the two forms are still considered challenging tasks in cancer analysis. Herein, we propose a new approach, LDSSNV, to detect somatic SNVs without matched normal samples. LDSSNV predicts SNVs by training the XGboost classifier on a concise combination of features and distinguishes the two forms based on linkage disequilibrium which is a trait between germline mutations. LDSSNV provides two modes to distinguish the somatic variants from germline variants, the single-mode and multiple-mode by respectively using a single tumor sample and multiple tumor samples. The performance of the proposed method is assessed on both simulation data and real sequencing datasets. The analysis shows that the LDSSNV method outperforms competing methods and can become a robust and reliable tool for analyzing tumor genome variation.
Jingfen Lan, Wenxiang Chen, Ganggang Yin, Haque A. K. Alvi, Kun Xie 0011, Qiang Yu 0003, Xiguo Yuan
IEEE ACM Trans. Comput. Biol. Bioinform.2
2022 CXR Data Annotation and Classification with Pre-trained Language Models
abstract
Clinical data annotation has been one of the major obstacles for applying machine learning approaches in clinical NLP. Open-source tools such as NegBio and CheXpert are usually designed on data from specific institutions, which limit their applications to other institutions due to the differences in writing style, structure, language use as well as label definition. In this paper, we propose a new weak supervision annotation framework with two improvements compared to existing annotation frameworks: 1) we propose to select representative samples for efficient manual annotation; 2) we propose to auto-annotate the remaining samples, both leveraging on a self-trained sentence encoder. This framework also provides a function for identifying inconsistent annotation errors. The utility of our proposed weak supervision annotation framework is applicable to any given data annotation task, and it provides an efficient form of sample selection and data auto-annotation with better classification results for real applications.
Nina Zhou, AiTi Aw, Zhuo Han Liu, Cher Heng Tan, Yonghan Ting, Wenxiang Chen, Jordan Zheng Ting Sim
COLING6
2018 Tunneling between plateaus: improving on a state-of-the-art MAXSAT solver using partition crossover
abstract
There are two important challenges for local search algorithms when applied to Maximal Satisfiability (MAXSAT). 1) Local search spends a great deal of time blindly exploring plateaus in the search space and 2) local search is less effective on application instances. This second problem may be related to local search's inability to exploit problem structure. We propose a genetic recombination operator to address both of these issues. On problems with well defined local optima, partition crossover is able to "tunnel" between local optima to discover new local optima in O(n) time. The PXSAT algorithm combines partition crossover and local search to produce a new way to escape plateaus. Partition crossover locally decomposes the evaluation function for a given instance into independent components, and is guaranteed to find the best solution among an exponential number of candidate solutions in O(n) time. Empirical results on an extensive set of application instances show that the proposed framework substantially improves two of best local search solvers, AdaptG2WSAT and Sparrow, on many application instances. PXSAT combined with AdaptG2WSAT is also able to outperform CCLS, winner of several recent MAXSAT competitions.
Wenxiang Chen, L. Darrell Whitley, Renato Tinós, Francisco Chicano
GECCO1
2017 Selecting Optimal Models Based on Efficiency and Robustness in Multi-valued Biological Networks
abstract
In this paper, we propose an optimization algorithm for literature-derived model and parameter identification in multi-valued biological regulatory networks. Our approach is a multi-objective optimization method where the objectives are inspired from structural Efficiency, dynamical Robustness and biological selectivity of cells in their actions. Given an incomplete model derived from literature and partially instrumented clinical observations, our method identifies the optimal model parameterization by maximizing structural Efficiency, dynamical Robustness and Selectivity. As the parameterization space is super exponential, we implemented our method in a constraint satisfaction framework by defining logical equivalences of the dynamical features. The implemented framework is then solved with a lazy clause solver known as Chuffed. We apply our method on female Hypothalamic-Pituitary-Gonadal axis (HPG) and demonstrate how it is able to identify a model that reproduces the complex menstrual cycle. The algorithm found a structure and parameterization for the 5 node 14 edge (≈ 50% edge density) HPG model with a normalized length cost and robustness of 1.46 and 0.35 respectively in 713 seconds on an Intel core i7 machine.Our method discovered that there are at least 6 more regulatory interactions that must be added to the commonly accepted HPG basic model in order to reproduce the menstrual cycle efficiently and robustly. The discovery of additional interactions suggest that our algorithm provides new insight to the biological model identification by combining the information from literature, clinical measurements and dynamical parameters.
Hooman Sedghamiz, Wenxiang Chen, L. Darrell Whitley, Gordon Broderick
BIBE2
2017 Decomposing SAT Instances with Pseudo Backbones
Wenxiang Chen, L. Darrell Whitley
EvoCOP1
2016 Stochastic Local Search over Minterms on Structured SAT Instances
abstract
We observed that Conjunctive Normal Form (CNF) encodings of structured SAT instances often have a set of consecutive clauses defined over a small number of Boolean variables. To exploit the pattern, we propose a transformation of CNF to an alternative representation, Conjunctive Minterm Canonical Form (CMCF). The transformation is a two-step process: CNF clauses are first partitioned into disjoint subsets such that each subset contains CNF clauses with shared Boolean variables. CNF clauses in each subset are then replaced by Minterm Canonical Form (i.e., partial solutions), which is found by enumeration. We show empirically that a simple Stochastic Local Search (SLS) solver based on CMCF can consistently achieve a higher success rate using fewer evaluations than the SLS solver WalkSAT on two representative classes of structured SAT problems.
Wenxiang Chen, L. Darrell Whitley, Adele E. Howe, Brian W. Goldman
SOCS1
2013 Impact of problem decomposition on Cooperative Coevolution
abstract
Variable Interaction Learning (VIL) is an emerging technique regarding detecting interacting variables so that Cooperative Coevolutionary Evolutionary Algorithms (CCEAs) can decompose problems accordingly and tackle subproblems of smaller sizes. While previous approaches are developed to efficiently perform VIL, no study has been on the actual usefulness of the detected variable interactions in terms of the performance of CCEAs. Since VIL is a computationally expensive task by itself, overly spending time on VIL without notable benefits for CCEAs should be avoided. It is hence critical to study the real impact of problem decomposition on CCEAs. We conduct empirical studies to address three closely related questions: 1) will a better problem decomposition lead to better performance of CCEAs, 2) when will improving problem decomposition benefit CCEAs, and 3) to what extent will improving problem decomposition enhance the performance of CCEAs.
Wenxiang Chen, Ke Tang 0001
IEEE Congress on Evolutionary Computation1
2013 Second order partial derivatives for NK-landscapes
abstract
Local search methods based on explicit neighborhood enumeration require at least $O(n)$ time to identify all possible improving moves. For k-bounded pseudo-Boolean optimization problems, recent approaches have achieved $O(k^2*2^{k})$ runtime cost per move, where $n$ is the number of variables and $k$ is the number of variables per subfunction. Even though the bound is independent of $n$, the complexity per move is still exponential in $k$. In this paper, we propose a second order partial derivatives-based approach that executes first-improvement local search where the runtime cost per move is time polynomial in $k$ and independent of $n$. This method is applied to NK-landscapes, where larger values of $k$ may be of particular interest.
Wenxiang Chen, L. Darrell Whitley, Doug Hains, Adele E. Howe
GECCO1
2013 Hyperplane initialized local search for MAXSAT
abstract
By converting the MAXSAT problem to Walsh polynomials, we can efficiently and exactly compute the hyperplane averages of fixed order k. We use this fact to construct initial solutions based on variable configurations that maximize the sampling of hyperplanes with good average evaluations. The Walsh coefficients can also be used to implement a constant time neighborhood update which is integral to a fast next descent local search for MAXSAT (and for all bounded pseudo-Boolean optimization problems.) We evaluate the effect of initializing local search with hyperplane averages on both the first local optima found by the search and the final solutions found after a fixed number of bit flips. Hyperplane initialization not only provides better evaluations, but also finds local optima closer to the globally optimal solution in fewer bit flips than search initialized with random solutions. A next descent search initialized with hyperplane averages is able to outperform several state-of-the art stochastic local search algorithms on both random and industrial instances of MAXSAT.
Doug Hains, L. Darrell Whitley, Adele E. Howe, Wenxiang Chen
GECCO4
2012 Constant time steepest descent local search with lookahead for NK-landscapes and MAX-kSAT
abstract
A modified form of steepest descent local search is proposed that displays an average complexity of O(1) time per move for NK-Landscape and MAX-kSAT problems. The algorithm uses a Walsh decomposition to identify improving moves. In addition, it is possible to compute a Hamming distance 2 statistical lookahead: if x is the current solution and y is a neighbor of x, it is possible to compute the average evaluation of the neighbors of y. The average over the Hamming distance 2 neighborhood can be used as a surrogate evaluation function to replace f. The same modified steepest descent can be executed in O(1) time using the Hamming distance 2 neighborhood average as the fitness function. In practice, the modifications needed to prove O(1) complexity can be relaxed with little or no impact on runtime performance. Finally, steepest descent local search over the mean of the Hamming distance 2 neighborhood yields superior results compared to using the standard evaluation function for certain types of NK-Landscape problems.
L. Darrell Whitley, Wenxiang Chen
GECCO2
2012 An Empirical Evaluation of O(1) Steepest Descent for NK-Landscapes
L. Darrell Whitley, Wenxiang Chen, Adele E. Howe
PPSN (1)2
2010 Large-Scale Global Optimization Using Cooperative Coevolution with Variable Interaction Learning
Wenxiang Chen, Thomas Weise 0001, Zhenyu Yang 0008, Ke Tang 0001
PPSN (2)1