Yifei Xing 0001

dblp:44/5901-1 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0002-5206-9515ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Disentangled diffusion model for 3D molecular generation with protein-ligand interaction priors
abstract
MOTIVATION: Structure-based drug design (SBDD) aims to generate ligand molecules that tightly bind to specific protein targets, a critical step in drug discovery. Diffusion models have shown promise for this task, yet existing methods struggle to effectively incorporate protein-ligand interaction priors during generation. Most approaches rely on protein-specific structural priors that remain fixed throughout generation, limiting molecular diversity and failing to capture the dynamic interplay between protein pockets and ligand atoms, which is essential for achieving high binding affinity. RESULTS: We propose DPDiff, a disentangled prior-conditioned diffusion model for protein-specific 3D molecular generation. DPDiff introduces two complementary interaction prior networks that capture geometry-based spatial interactions and sequence-based interactions robust to structural noise. During generation, the model dynamically extracts interaction priors using intermediate diffusion predictions and adaptively fuses them via a time-dependent adapter. A disentangled denoising network balances prior guidance with generative flexibility. Experiments on the CrossDocked2020 dataset demonstrate that DPDiff generates molecules with more realistic 3D structures and state-of-the-art binding affinities, achieving an average Vina Dock score of -8.58 and a high affinity ratio of 69.4%, outperforming existing methods while maintaining favorable drug-likeness and synthetic accessibility. AVAILABILITY AND IMPLEMENTATION: The source code of DPDiff is available at https://github.com/ZerinHwang03/DPDiff.
Zhilin Huang, Ling Yang 0006, Chujun Qin, Yifei Xing 0001, Xiangxin Zhou, Yu Wang 0027, Xin Gao 0001, Wenming Yang
Bioinform.4
2026 DyToS: Budget-aware dynamic token scheduling for efficient multi-modal large language models
Yifei Xing 0001, Ruiping Wang 0001, Dongmei Jiang, Xiangyuan Lan
Neurocomputing1
2026 Learning from easy to hard: Curriculum meta-learning for few-shot node classification
Qilong Yan, Weinan Guan, Yifei Xing 0001, Jingpu Duan, Jian Yin 0001
Inf. Sci.3
2026 AlignMamba-2: Enhancing multimodal fusion and sentiment analysis with modality-aware Mamba
Yan Li 0121, Yifei Xing 0001, Xiangyuan Lan, Xin Li 0034, Dongmei Jiang
Pattern Recognit.2
2026 Decoupled gradient-guided stratification for resource-efficient multi-modal data pruning
Yifei Xing 0001, Ruiping Wang 0001, Xiangyuan Lan, Yaowei Wang 0001
Pattern Recognit. Lett.1
2025 AlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignment
abstract
Cross-Modal alignment is crucial for multimodal representation fusion due to the inherent heterogeneity between modalities. While Transformer-Based methods have shown promising results in modeling inter-modal relationships, their quadratic computational complexity limits their applicability to long-sequence or large-scale data. Although recent Mamba-Based approaches achieve linear complexity, their sequential scanning mechanism poses fundamental challenges in comprehensively modeling cross-modal relationships. To address this limitation, we propose Align-Mamba, an efficient and effective method for multimodal fusion. Specifically, grounded in Optimal Transport, we introduce a local cross-modal alignment module that explicitly learns token-level correspondences between different modalities. Moreover, we propose a global cross-modal alignment loss based on Maximum Mean Discrepancy to implicitly enforce the consistency between different modal distributions. Finally, the unimodal representations after local and global alignment are passed to the Mamba backbone for further cross-modal interaction and multimodal fusion. Extensive experiments on complete and incomplete multimodal fusion tasks demonstrate the effectiveness and efficiency of the proposed method. For instance, on the CMU-MOSI dataset, AlignMamba improves classification accuracy by 0.9%, reduces GPU memory usage by 20.3%, and decreases inference time by 83.3%.
Yan Li 0121, Yifei Xing 0001, Xiangyuan Lan, Xin Li 0034, Dongmei Jiang
CVPR2
2025 Enhancing Visual Understanding in Multimodal Large Language Models with Efficient Feature Alignment and State Space Models
abstract
Multimodal Large Language Models (MLLMs) excel at processing complex tasks involving visual and textual data. However, existing Mamba-based MLLMs face significant challenges in visual feature extraction, resulting in poor cross-modal alignment and compromised performance. To tackle these issues, we present ML-Mamba, a novel architecture built on the Mamba framework to enhance multimodal learning. ML-Mamba features a robust visual encoder, a Mamba-Transformer Projector for optimizing feature alignment and interaction between modalities, and the advanced Mamba large language model. By utilizing techniques such as cluster-based scanning for improved visual feature extraction and a Shared-Specialized feed-forward mechanism, ML-Mamba enhances visual representation quality and overall model efficiency. Extensive benchmarking shows that ML-Mamba outperforms existing models in multimodal tasks, significantly improving inference speed and cross-modal alignment. This work underscores the potential of integrating structured state-space models with advanced transformer components to develop scalable and resource-efficient multimodal models.
Jiakai Pan, Jiahao Tang, Yifei Xing 0001, Zhengzhuo Wang, Shengzhi Shen, Jianguo Hu
ECAI4
2025 EMMA: Empowering Multi-modal Mamba with Structural and Hierarchical Alignment
abstract
Mamba-based architectures have shown to be a promising new direction for deep learning models owing to their competitive performance and sub-quadratic deployment speed. However, current Mamba multi-modal large language models (MLLM) are insufficient in extracting visual features, leading to imbalanced cross-modal alignment between visual and textural latents, negatively impacting performance on multi-modal tasks. In this work, we propose Empowering Multi-modal Mamba with Structural and Hierarchical Alignment (EMMA), which enables the MLLM to extract fine-grained visual information. Specifically, we propose a pixel-wise alignment module to autoregressively optimize the learning and processing of spatial image-level features along with textual tokens, enabling structural alignment at the image level. In addition, to prevent the degradation of visual information during the cross-model alignment process, we propose a multi-scale feature fusion (MFF) module to combine multi-scale visual features from intermediate layers, enabling hierarchical alignment at the feature level. Extensive experiments are conducted across a variety of multi-modal benchmarks. Our model shows lower latency than other Mamba-based MLLMs and is nearly four times faster than transformer-based MLLMs of similar scale during inference. Due to better cross-modal alignment, our model exhibits lower degrees of hallucination and enhanced sensitivity to visual details, which manifests in superior performance across diverse multi-modal benchmarks. Code provided at https://github.com/xingyifei2016/EMMA.
Yifei Xing 0001, Xiangyuan Lan, Ruiping Wang 0001, Dongmei Jiang, Yaowei Wang 0001
ICLR1
2025 Enhanced Motion-aware Latent Diffusion Models for Video Frame Interpolation
abstract
The objective of video frame interpolation (VFI) methods is to enhance video fluency and visual quality by generating intermediate frames between consecutive original frames based on the source video. Recently, diffusion-based VFI methods have made promising progresses, with generated results performing well in perceptual quality. However, these methods have not fully explored how to effectively leverage external motion priors to enhance the model's ability to estimate motion information between adjacent frames, which is crucial for VFI models to avoid generating blurry results due to the motion ambiguity. In this paper, we propose an Enhanced Motion-Aware latent Diffusion model ( EMADiff ) for video frame interpolation. Specifically, we integrate motion priors into the decoder of vector-quantized enhanced motion-aware GAN to guide the information propagation during RGB interpolated frame reconstruction. Furthermore, we propose enhanced motion-aware noising and de-noising procedures. By reducing the discrepancy in attention to motion priors between the forward and reverse processes, our EMADiff effectively utilizes motion priors, alleviates motion ambiguity, and generates realistic content. Comprehensive experiments on benchmark datasets show EMADiff achieves state-of-the-art performance, surpassing existing approaches and producing visually plausible and content-clear results.
Zhilin Huang, Chujun Qin, Yifei Xing 0001, Wenming Yang
ACM Multimedia3
2025 Generic Scene Graph Generation Model with Hierarchical Prompt Learning
Xuhan Zhu, Yifei Xing 0001, Ruiping Wang 0001, Yaowei Wang 0001, Xiangyuan Lan
Int. J. Comput. Vis.2
2024 Hierarchical Prompt Learning for Scene Graph Generation
Xuhan Zhu, Yifei Xing 0001, Ruiping Wang 0001, Yaowei Wang 0001, Xiangyuan Lan
BMVC2
2024 Calibration for Long-tailed Scene Graph Generation
abstract
Miscalibrated models tend to be unreliable and insecure for downstream applications. In this work, we attempt to highlight and remedy miscalibration in current scene graph generation (SGG) models, which has been overlooked by previous works. We discover that obtaining well-calibrated models for SGG is more challenging than conventional calibration settings, as long-tailed SGG training data exacerbates miscalibration with overconfidence in head classes and underconfidence in tail classes. We further analyze which components are explicitly impacted by the long-tailed data during optimization, thereby exacerbating miscalibration and unbalanced learning, including biased parameters, deviated boundaries, and distorted target distribution. To address the above issues, we propose the Compositional Optimization Calibration (COC) method, comprising three modules: i. A parameter calibration module that utilizes a hyperspherical classifier to eliminate the bias introduced by biased parameters. ii. A boundary calibration module that disperses features of majority classes to consolidate the decision boundaries of minority classes and mitigate deviated boundaries. iii. A target distribution calibration module that addresses distorted target distribution, leverages within-triplet prior to guide confidence-aware and label-aware target calibration, and applies curriculum regulation to constrain learning focus from easy to hard classes. Extensive evaluation on popular benchmarks demonstrates the effectiveness of our proposed method in improving model calibration and resolving unbalanced learning for long-tailed SGG. Finally, our proposed method performs best on model calibration compared to different types of calibration methods and achieves state-of-the-art trade-off performance on balanced SGG learning.
Xuhan Zhu, Yifei Xing 0001, Ruiping Wang 0001, Yaowei Wang 0001, Xiangyuan Lan
ACM Multimedia2
2022 Co-domain Symmetry for Complex-Valued Deep Learning
abstract
We study complex-valued scaling as a type of symmetry natural and unique to complex-valued measurements and representations. Deep Complex Networks (DCN) extend real-valued algebra to the complex domain without addressing complex-valued scaling. SurReal extends manifold learning to the complex plane, achieving scaling invariance with manifold distances that discard phase information. Treating complex-valued scaling as a co-domain transformation, we design novel equivariant/invariant layer functions and architectures that exploit co-domain symmetry. We also propose novel complex-valued representations of RGB images, where complex-valued scaling indicates hue shift or correlated changes across color channels. Benchmarked on MSTAR, CIFAR10, CIFAR100, and SVHN, our co-domain symmetric (CDS) classifiers deliver higher accuracy, better generalization, more robustness to co-domain transformations, and lower model bias and variance than DCN and SurReal with far fewer parameters.
Utkarsh Singhal, Yifei Xing 0001, Stella X. Yu
CVPR2
2022 SurReal: Complex-Valued Learning as Principled Transformations on a Scaling and Rotation Manifold
abstract
Complex-valued data are ubiquitous in signal and image processing applications, and complex-valued representations in deep learning have appealing theoretical properties. While these aspects have long been recognized, complex-valued deep learning continues to lag far behind its real-valued counterpart. We propose a principled geometric approach to complex-valued deep learning. Complex-valued data could often be subject to arbitrary complex-valued scaling; as a result, real and imaginary components could covary. Instead of treating complex values as two independent channels of real values, we recognize their underlying geometry: we model the space of complex numbers as a product manifold of nonzero scaling and planar rotations. Arbitrary complex-valued scaling naturally becomes a group of transitive actions on this manifold. We propose to extend the property instead of the form of real-valued functions to the complex domain. We define convolution as the weighted Fréchet mean on the manifold that is equivariant to the group of scaling/rotation actions and define distance transform on the manifold that is invariant to the action group. The manifold perspective also allows us to define nonlinear activation functions, such as tangent ReLU and G -transport, as well as residual connections on the manifold-valued data. We dub our model SurReal, as our experiments on MSTAR and RadioML deliver high performance with only a fractional size of real- and complex-valued baseline models.
Rudrasis Chakraborty, Yifei Xing 0001, Stella X. Yu
IEEE Trans. Neural Networks Learn. Syst.2