Yuxian Wang

dblp:65/7833 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
YearPublicationVenuePosition
2026 TinyJudge: Unverifiable Constraint Alignment via Lightweight Specialist Ensembles
abstract
Yirong Zeng, Yufei Liu, Xiao Ding, Yutai Hou, Yuxian Wang, Wu Ning, Haonan Song, Dandan Tu, Qixun Zhang, Yuxiang He, Bibo Cai, Ting Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yirong Zeng, Yutai Hou, Yuxian Wang, Wu Ning, Haonan Song, Dandan Tu, Qixun Zhang, Bibo Cai, Ting Liu 0001
ACL (1)5
2025 iTool: Reinforced Fine-Tuning with Dynamic Deficiency Calibration for Advanced Tool Use
abstract
Yirong Zeng, Xiao Ding, Yuxian Wang, Weiwen Liu, Yutai Hou, Wu Ning, Xu Huang, Duyu Tang, Dandan Tu, Bing Qin, Ting Liu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yirong Zeng, Yuxian Wang, Weiwen Liu, Yutai Hou, Wu Ning, Xu Huang 0008, Duyu Tang, Dandan Tu, Bing Qin 0001, Ting Liu 0001
EMNLP3
2025 ToolACE: Winning the Points of LLM Function Calling
abstract
Function calling significantly extends the application boundary of large language models (LLMs), where high-quality and diverse training data is critical for unlocking this capability. However, collecting and annotating real function-calling data is challenging, while synthetic data from existing pipelines often lack coverage and accuracy. In this paper, we present ToolACE, an automatic agentic pipeline designed to generate accurate, complex, and diverse tool-learning data, specifically tailored to the capabilities of LLMs. ToolACE leverages a novel self-evolution synthesis process to curate a comprehensive API pool of 26,507 diverse APIs. Dialogs are further generated through the interplay among multiple agents, under the guidance of a complexity evaluator. To ensure data accuracy, we implement a dual-layer verification system combining rule-based and model-based checks. We demonstrate that models trained on our synthesized data---even with only 8B parameters---achieve state-of-the-art performance, comparable to the latest GPT-4 models. Our model and a subset of the data are publicly available at https://huggingface.co/Team-ACE.
Weiwen Liu, Xu Huang 0008, Xingshan Zeng, Xinlong Hao, Dexun Li, Shuai Wang 0020, Weinan Gan, Zhengying Liu, Yuanqing Yu, Zezhong Wang 0004, Yuxian Wang, Wu Ning, Yutai Hou, Bin Wang 0004, Chuhan Wu, Yong Liu 0020, Yasheng Wang, Duyu Tang, Dandan Tu, Lifeng Shang, Xin Jiang 0002, Ruiming Tang, Defu Lian, Qun Liu 0001, Enhong Chen
ICLR12
2024 Identifying brain disease genes via integrating brain imaging and molecular network
abstract
Identifying genes associate to brain diseases is crucial for uncovering the related biological mechanisms and for advancing therapeutic interventions. Due to the ongoing advancements in network-based computational approaches, biological networks, especially molecular networks, provide valuable insights for predicting disease genes. However, many methods have ignored brain imaging data when exploring brain disease genes, despite its widely use in neuroscience research. In this paper, we propose a novel framework, Deep Interactive AutoEncoder (DIAE), which integrates brain imaging and molecular-based gene networks to predict brain disease genes. DIAE first constructs a gene association network based on brain imaging and high-resolution whole brain-whole gene expression data. Subsequently, a deep and interactive multi-network integration method is introduced to learn low-dimensional features of genes by combining the brain imaging-based network with other molecular-based gene networks. Finally, these features are utilized to predict brain disease genes using a support vector machine (SVM) model. For performance evaluation, we compare DIAE with three existing state-of-the-art methods in the context of disease gene identification across four brain diseases. The experimental results show the superior performance of DIAE and highlight the effectiveness of the brain imaging-based gene network for predicting brain disease genes.
Yuxian Wang, Jiajie Peng
BIBM2
2023 Gene selection in a single cell gene decision space based on class-consistent technology and fuzzy rough iterative computation model
Guangji Yu, Yuxian Wang
Appl. Intell.4
2023 DFinder: a novel end-to-end graph embedding-based method to identify drug-food interactions
abstract
MOTIVATION: Drug-food interactions (DFIs) occur when some constituents of food affect the bioaccessibility or efficacy of the drug by involving in drug pharmacodynamic and/or pharmacokinetic processes. Many computational methods have achieved remarkable results in link prediction tasks between biological entities, which show the potential of computational methods in discovering novel DFIs. However, there are few computational approaches that pay attention to DFI identification. This is mainly due to the lack of DFI data. In addition, food is generally made up of a variety of chemical substances. The complexity of food makes it difficult to generate accurate feature representations for food. Therefore, it is urgent to develop effective computational approaches for learning the food feature representation and predicting DFIs. RESULTS: In this article, we first collect DFI data from DrugBank and PubMed, respectively, to construct two datasets, named DrugBank-DFI and PubMed-DFI. Based on these two datasets, two DFI networks are constructed. Then, we propose a novel end-to-end graph embedding-based method named DFinder to identify DFIs. DFinder combines node attribute features and topological structure features to learn the representations of drugs and food constituents. In topology space, we adopt a simplified graph convolution network-based method to learn the topological structure features. In feature space, we use a deep neural network to extract attribute features from the original node attributes. The evaluation results indicate that DFinder performs better than other baseline methods. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/23AIBox/23AIBox-DFinder. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tao Wang 0082, Jinjin Yang, Yifu Xiao, Yuxian Wang, Yongtian Wang, Jiajie Peng
Bioinform.5
2023 Unsupervised Bayesian Subpixel Mapping Autoencoder Network for Hyperspectral Images
abstract
Unsupervised subpixel mapping (SPM) of hyperspectral image (HSI) is a challenging task due to the difficulties to integrate different prior information and model constraints into a coherent framework. This paper presents a Bayesian neural network for unsupervised HSI SPM, which has the following characteristics. First, the deep image prior (DIP) achieved by a fully convolutional neural network (FCNN) is used to model the spatial correlation efficiently and adaptively in the subpixel label domain. Second, a discrete spectral mixture model (DSMM) is designed to leverage the forward model for enhanced SPM. Third, an auto-encoder architecture is designed to integrate the FCNN and the DSMM to allow efficient unsupervised representational learning using both data and knowledge. Fourth, an expectation-maximization approach is designed to solve the resulting maximum a posteriori problem, where a purified means approach extracts endmembers, and the gradient descent approach updates FCNN parameters for subpixel label estimation. Comparative experiments on both real and simulated HSIs demonstrate that the proposed method outperforms other state-of-the-art methods in terms of both numerical accuracies and visual subpixel mapping results.
Yuan Fang 0003, Yuxian Wang, Linlin Xu, Yujia Chen 0002, Alexander Wong, David A. Clausi
IEEE Trans. Geosci. Remote. Sens.2
2022 A network-based method for brain disease gene prediction by integrating brain connectome and molecular network
abstract
Brain disease gene identification is critical for revealing the biological mechanism and developing drugs for brain diseases. To enhance the identification of brain disease genes, similarity-based computational methods, especially network-based methods, have been adopted for narrowing down the searching space. However, these network-based methods only use molecular networks, ignoring brain connectome data, which have been widely used in many brain-related studies. In our study, we propose a novel framework, named brainMI, for integrating brain connectome data and molecular-based gene association networks to predict brain disease genes. For the consistent representation of molecular-based network data and brain connectome data, brainMI first constructs a novel gene network, called brain functional connectivity (BFC)-based gene network, based on resting-state functional magnetic resonance imaging data and brain region-specific gene expression data. Then, a multiple network integration method is proposed to learn low-dimensional features of genes by integrating the BFC-based gene network and existing protein-protein interaction networks. Finally, these features are utilized to predict brain disease genes based on a support vector machine-based model. We evaluate brainMI on four brain diseases, including Alzheimer's disease, Parkinson's disease, major depressive disorder and autism. brainMI achieves of 0.761, 0.729, 0.728 and 0.744 using the BFC-based gene network alone and enhances the molecular network-based performance by 6.3% on average. In addition, the results show that brainMI achieves higher performance in predicting brain disease genes compared to the existing three state-of-the-art methods.
Ruijiang Han, Menghan Zhang, Yuxian Wang, Tao Wang 0082, Yongtian Wang, Xuequn Shang 0001, Jiajie Peng
Briefings Bioinform.4
2022 A review and performance evaluation of clustering frameworks for single-cell Hi-C data
abstract
The three-dimensional genome structure plays a key role in cellular function and gene regulation. Single-cell Hi-C (high-resolution chromosome conformation capture) technology can capture genome structure information at the cell level, which provides the opportunity to study how genome structure varies among different cell types. Recently, a few methods are well designed for single-cell Hi-C clustering. In this manuscript, we perform an in-depth benchmark study of available single-cell Hi-C data clustering methods to implement an evaluation system for multiple clustering frameworks based on both human and mouse datasets. We compare eight methods in terms of visualization and clustering performance. Performance is evaluated using four benchmark metrics including adjusted rand index, normalized mutual information, homogeneity and Fowlkes-Mallows index. Furthermore, we also evaluate the eight methods for the task of separating cells at different stages of the cell cycle based on single-cell Hi-C data.
Caiwei Zhen, Yuxian Wang, Jiaquan Geng, Jinghao Peng, Tao Wang 0082, Jianye Hao, Xuequn Shang 0001, Zhongyu Wei, Peican Zhu, Jiajie Peng
Briefings Bioinform.2
2022 BCUN: Bayesian Fully Convolutional Neural Network for Hyperspectral Spectral Unmixing
abstract
Spectral unmixing (SU) plays a fundamental role in hyperspectral image (HSI) processing. Effective SU relies on the accurate and efficient characterization of the noise effect, the endmembers, and the spatial correlation effect in abundances, as well as efficient optimization techniques to estimate these effects. To address these issues, this article presents a Bayesian fully convolutional hyperspectral unmixing network (BCUN) with the following key characteristics. First, a fully convolutional neural network (FCNN)-based deep image prior (DIP) is designed for enhanced characterization and estimation of the spatial context information in abundance maps, leading to more efficient and accurate abundance modeling than the traditional nonnegative least squares (NNLS) approaches. Second, a multivariate Gaussian distribution with an anisotropic covariance matrix is designed to characterize the conditional distribution of the spectral observations, leading to a novel Mahalanobis distance-based loss for FCNN training that is better capable of addressing the noise heterogeneous effect in HSI than the Euclidean distance-based mean squared error (MSE) loss in traditional deep neural networks. Third, the designed conditional distribution of spectral observations also enables the incorporation of the spectral mixture model (SMM) into the FCNN training process for effectively leveraging the knowledge in the forward spectral model. Fourth, the endmembers are modeled and estimated by a “purified means” approach that is capable of better characterizing endmembers. Finally, the above key components are coherently integrated into a Bayesian framework, and the resulting maximuma posteriori(MAP) problem is solved by a designed expectation–maximization (EM) algorithm. Experimental results on both simulated and real HSIs demonstrate that the proposed BCUN approach outperforms the other classical and state-of-the-art methods on both endmember estimation and abundance estimation.
Yuan Fang 0003, Yuxian Wang, Linlin Xu, Rongming Zhuo, Alexander Wong, David A. Clausi
IEEE Trans. Geosci. Remote. Sens.2
2021 Predicting Hepatoma-Related Genes Based on Representation Learning of PPI network and Gene Ontology Annotations
abstract
Hepatoma is the most common type of primary liver cancer with a high mortality rate in the world. The genetic causes of the disease pathology remain largely unknown. Effective discovery of the genes associated with hepatoma has become important in disease prevention, early diagnosis, and therapeutic treatments. With the developments of molecular networks, graph-based methods have been tremendously successful in predicting disease genes based on the hypothesis of guilt-by-association. Network representation learning (NRL) techniques have accelerated disease gene discovery in recent years because of their powerful network feature extraction ability. However, the current network representation learning-based methods for disease gene discovery did not consider the gene features derived from gene ontology annotations, which apriori group genes with similar functions. To fill this gap, here we propose a novel framework to predict hepatoma-related genes based on representation learning from both protein-protein interactions (PPI) network and gene ontology annotations. Our framework has three steps: learning features from PPI network and gene ontologies using NRL techniques, integrating different features based on autoencoder, predicting hepatoma-related genes using machine learning classifiers. Experiments have demonstrated that our framework could accurately predict hepatoma-related genes with AUROC and AUPRC reaching 0.93 and 0.94, respectively. Compared with other methods using only single representation features, our framework also shows superior performance on hepatoma gene prediction.
Tao Wang 0082, Zhiyuan Shao, Yifu Xiao, Xuchao Zhang, Binze Shi, Siyu Chen 0024, Yuxian Wang, Jiajie Peng, Xuequn Shang 0001
BIBM8
2021 An end-to-end heterogeneous graph representation learning-based framework for drug-target interaction prediction
abstract
Accurately identifying potential drug-target interactions (DTIs) is a key step in drug discovery. Although many related experimental studies have been carried out for identifying DTIs in the past few decades, the biological experiment-based DTI identification is still timeconsuming and expensive. Therefore, it is of great significance to develop effective computational methods for identifying DTIs. In this paper, we develop a novel 'end-to-end' learning-based framework based on heterogeneous 'graph' convolutional networks for 'DTI' prediction called end-to-end graph (EEG)-DTI. Given a heterogeneous network containing multiple types of biological entities (i.e. drug, protein, disease, side-effect), EEG-DTI learns the low-dimensional feature representation of drugs and targets using a graph convolutional networks-based model and predicts DTIs based on the learned features. During the training process, EEG-DTI learns the feature representation of nodes in an end-to-end mode. The evaluation test shows that EEG-DTI performs better than existing state-of-art methods. The data and source code are available at: https://github.com/MedicineBiology-AI/EEG-DTI.
Jiajie Peng, Yuxian Wang, Jiaojiao Guan, Ruijiang Han, Jianye Hao, Zhongyu Wei, Xuequn Shang 0001
Briefings Bioinform.2