Xuanbai Ren

dblp:253/6603 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Generative modeling · 34% Optimization for machine learning · 32% Efficient and distributed learning · 18%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 68% Medical and health informatics · 32%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model merging
1.012026
CAML: A Conflict-Aware Molecular Language Model Merging Framework for Multi-Constraint Molecular Generation · ACL (1) 2026
Machine learning › Generative modeling
molecular generation
1.012026
CAML: A Conflict-Aware Molecular Language Model Merging Framework for Multi-Constraint Molecular Generation · ACL (1) 2026
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization
0.912025
Multi-Objective Molecular Design Through Learning Latent Pareto Set · AAAI 2025
Machine learning › Generative modeling
latent space model
0.912025
Multi-Objective Molecular Design Through Learning Latent Pareto Set · AAAI 2025
Machine learning › Graph learning
molecular representation learning
0.912025
Self-supervised Blending Structural Context of Visual Molecules for Robust Drug Interaction Prediction · NeurIPS 2025
Machine learning › Optimization for machine learning › multi-objective optimization
pareto set learning
0.912025
Multi-Objective Molecular Design Through Learning Latent Pareto Set · AAAI 2025
Medical and health informatics › drug safety
drug-drug interaction prediction
0.912025
Self-supervised Blending Structural Context of Visual Molecules for Robust Drug Interaction Prediction · NeurIPS 2025
Bioinformatics and computational biology › molecular informatics
molecular design
0.912025
Multi-Objective Molecular Design Through Learning Latent Pareto Set · AAAI 2025
Bioinformatics and computational biology › gene regulation › regulatory element discovery
enhancer prediction
0.512021
iEnhancer-XG: interpretable sequence-based enhancers and their strength predictor · Bioinform. 2021
Bioinformatics and computational biology › sequence analysis
sequence-based prediction
0.512021
iEnhancer-XG: interpretable sequence-based enhancers and their strength predictor · Bioinform. 2021

Methods — techniques the papers use, named apart from their topics

visual molecule encoding · 1.7surrogate model · 1.7self-supervised pretraining · 1.7local bayesian optimization · 1.7encoder-decoder · 1.7molecular language models · 1.0language model merging · 1.0pseudo dinucleotide composition · 0.5position-specific scoring matrix · 0.5mismatch k-tuple · 0.5k-spectrum profile · 0.5ensemble learning · 0.5
YearPublicationVenuePosition
2026 CAML: A Conflict-Aware Molecular Language Model Merging Framework for Multi-Constraint Molecular Generation
abstract
Xuanbai Ren, Luoda Tan, Pei Liu, Tengfei Ma, Xiangzheng Fu, Longyue Wang, Yiping Liu, Xiangxiang Zeng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xuanbai Ren, Luoda Tan, Pei Liu 0008, Tengfei Ma 0002, Xiangzheng Fu, Longyue Wang, Xiangxiang Zeng
ACL (1)1
2025 Multi-Objective Molecular Design Through Learning Latent Pareto Set
abstract
Molecular design inherently involves the optimization of multiple conflicting objectives, such as enhancing bio-activity and ensuring synthesizability. Evaluating these objectives often requires resource-intensive computations or physical experiments. Current molecular design methodologies typically approximate the Pareto set using a limited number of molecules. In this paper, we present an innovative approach, called Multi-Objective Molecular Design through Learning Latent Pareto Set (MLPS). MLPS initially utilizes an encoder-decoder model to seamlessly transform the discrete chemical space into a continuous latent space. We then employ local Bayesian optimization models to efficiently search for local optimal solutions (i.e., molecules) within predefined trust regions. Using surrogate objective values derived from these local models, we train a global Pareto set learning model to understand the mapping between direction vectors (called “preferences”) in the objective space and the entire Pareto set in the continuous latent space. Both the global Pareto set learning model and local Bayesian optimization models collaborate to discover high-quality solutions and adapt the trust regions dynamically. Our work is an effective endeavor towards learning the Pareto set for multi-objective molecular design, providing decision-makers with the capability to fine-tune their preferences and thoroughly explore the Pareto set. Experimental results demonstrate that MLPS achieves state-of-the-art performance across various multi-objective scenarios, encompassing diverse objective types and varying numbers of objectives. The effectiveness of MLPS was further validated through real-world challenges in discovering antifungal peptides with low toxicity and high activity.
Xuanbai Ren, Yuansheng Liu, Bosheng Song, Xiangxiang Zeng, Hisao Ishibuchi
AAAI3
2025 Self-supervised Blending Structural Context of Visual Molecules for Robust Drug Interaction Prediction
abstract
Identifying drug-drug interactions (DDIs) is critical for ensuring drug safety and advancing drug development, a topic that has garnered significant research interest. While existing methods have made considerable progress, approaches relying solely on known DDIs face a key challenge when applied to drugs with limited data: insufficient exploration of the space of unlabeled pairwise drugs. To address these issues, we innovatively introduce S$^2$VM, a Self-supervised Visual pretraining framework for pair-wise Molecules, to fully fuse structural representations and explore the space of drug pairs for DDI prediction. S$^2$VM incorporates the explicit structure and correlations of visual molecules, such as the positional relationships and connectivity between functional substructures. Specifically, we blend the visual fragments of drug pairs into a unified input for joint encoding and then recover molecule-specific visual information for each drug individually. This approach integrates fine-grained structural representations from unlabeled drug pair data. By using visual fragments as anchors, S$^2$VM effectively captures the spatial information of local molecular components within visual molecules, resulting in more comprehensive embeddings of drug pairs. Experimental results show that S$^2$VM achieves state-of-the-art performance on widely used benchmarks, with Macro-F1 score improvements of 4.21% and 3.31%, respectively. Further extensive results and theoretical analysis demonstrate the effectiveness of S$^2$VM for both few-shot and novel drugs.
Tengfei Ma 0002, Yongsheng Zang, Yujie Chen 0002, Xuanbai Ren, Bosheng Song, Hongxin Xiang, Xiangxiang Zeng
NeurIPS5
2021 iEnhancer-XG: interpretable sequence-based enhancers and their strength predictor
abstract
MOTIVATION: Enhancers are non-coding DNA fragments with high position variability and free scattering. They play an important role in controlling gene expression. As machine learning has become more widely used in identifying enhancers, a number of bioinformatic tools have been developed. Although several models for identifying enhancers and their strengths have been proposed, their accuracy and efficiency have yet to be improved. RESULTS: We propose a two-layer predictor called 'iEnhancer-XG.' It comprises a one-layer predictor (for identifying enhancers) and a second classifier (for their strength) and uses 'XGBoost' as a base classifier and five feature extraction methods, namely, k-Spectrum Profile, Mismatch k-tuple, Subsequence Profile, Position-specific scoring matrix (PSSM) and Pseudo dinucleotide composition (PseDNC). Each method has an independent output. We place the feature vector matrix into the ensemble learning for fusion. This experiment involves the method of 'SHapley Additive explanations' to provide interpretability for the previous black box machine learning methods and improve their credibility. The accuracies of the ensemble learning method are 0.811 (first layer) and 0.657 (second layer). The rigorous 10-fold cross-validation confirms that the proposed method is significantly better than existing technologies. AVAILABILITY AND IMPLEMENTATION: The source code and dataset for the enhancer predictions have been uploaded to https://github.com/jimmyrate/ienhancer-xg. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xuanbai Ren, Xiangzheng Fu, Mingyu Gao 0004, Xiangxiang Zeng
Bioinform.2