Yuxuan Cheng

dblp:289/7947 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Scale Margin Loss for Object Detection
Yuxuan Cheng, Yanjun Zhang 0002, Leo Yu Zhang, Donald Donglong Chen, Yuming Fang 0001
KSEM (2)1
2025 Rethinking Graph Domain Adaptation: A Spectral Contrastive Perspective
Yuxuan Cheng, Wenqi Fan
ECML/PKDD (6)2
2025 CLIP-AdaM: Adapting Multi-view CLIP for Open-set 3D Object Retrieval
abstract
Open-set 3D object retrieval (3DOR) aims to learn discriminative and generalizable embeddings for unseen categories of 3D objects. However, attaining this objective typically requires the costly acquisition of large-scale 3D object datasets and associated resources for model training. Building upon the strong open-world representation capabilities of CLIP, we introduce CLIP-AdaM, which, to our knowledge, represents the first attempt to adapt a CLIP model for open-set 3DOR with minimal effort. We first find that a pretrained CLIP already delivers a surprisingly acceptable performance on multi-view images. To further unleash its potential, we design a customized adapter for learning to aggregate and adapt its pretrained features towards better 3D embeddings. For aggregation, it learns two sets of view scores to weigh the contributions of view images for fusion. One is learned by a tiny view-score network at the instance level, and the other is learned implicitly at the dataset level, aiding generalization to unseen categories. The adaptation component comprises only a basic linear layer yet yields superior results. During training, the adapter with such a small amount of parameters can be efficiently fine-tuned with limited 3D closed-set data, effectively mitigating the overfitting issue while harnessing the prior knowledge from pretrained models. Without bells and whistles, CLIP-AdaM attains state-of-the-art performance on four open-set 3DOR benchmarks. Additionally, it demonstrates strong extensibility to broader scenarios, including zero-shot, few-shot, and seen/unseen 3D representation learning.
Xinwei He 0001, Yuxuan Cheng, Yulong Wang 0002, Yang Zhou 0007, Xiang Bai
SIGIR3
2025 CrossACL: Analytic Continual Learning via Feature Cross for Hyperspectral Image Classification
abstract
Rapidly developing remote sensing technologies expand the volume and variety of hyperspectral images (HSIs). An HSI classification (HSIC) model should be able to adapt to new classes continually while retaining knowledge of previously learned classes to reduce training resources. However, popular HSIC models based on deep neural networks exhibit a significant performance decline in previously learned classes, known as the catastrophic forgetting phenomenon. To efficiently address this issue in HSIC, we propose an analytic continual learning method based on feature cross (CrossACL). CrossACL introduces a novel and training-free feature cross module (FCM) to better adapt to the increasingly complex feature space as the number of HSI classes increases. Furthermore, it utilizes an analytic recursive ridge regression classifier with a closed-form solution. This formulation achieves conditional equivalence between continual learning and joint training on all data seen so far, providing a theoretical guarantee against catastrophic forgetting. In addition, CrossACL introduces a simple but effective oversampling strategy to mitigate classification discrimination due to class-imbalanced HSI samples. Experiments on HSIC datasets demonstrate that CrossACL achieves competitive results compared with state-of-the-art methods at significantly lower computational consumption.
Jianan Ji, Yuxuan Cheng, Peiting Xiong, Huiping Zhuang
IEEE Geosci. Remote. Sens. Lett.3
2024 A4-Unet: Deformable Multi-Scale Attention Network for Brain Tumor Segmentation
abstract
Brain tumor segmentation models have aided diagnosis in recent years. However, they face MRI complexity and variability challenges, including irregular shapes and unclear boundaries, leading to noise, misclassification, and incomplete segmentation, thereby limiting accuracy. To address these issues, we adhere to an outstanding Convolutional Neural Networks (CNNs) design paradigm and propose a novel network named A4-Unet. In A4-Unet, Deformable Large Kernel Attention (DLKA) is incorporated in the encoder, allowing for improved capture of multi-scale tumors. Swin Spatial Pyramid Pooling (SSPP) with cross-channel attention is employed in a bottleneck further to study long-distance dependencies within images and channel relationships. To enhance accuracy, a Combined Attention Module (CAM) with Discrete Cosine Transform (DCT) orthogonality for channel weighting and convolutional element-wise multiplication is introduced for spatial weighting in the decoder. Attention gates (AG) are added in the skip connection to highlight the foreground while suppressing irrelevant background information. The proposed network is evaluated on three authoritative MRI brain tumor benchmarks and a proprietary dataset, and it achieves a 94.4% Dice score on the BraTS 2020 dataset, thereby establishing multiple new state-of-the-art benchmarks. The code is available here: https://github.com/WendyWAAAAANG/A4-Unet.
Ruoxin Wang, Haiming Du, Yuxuan Cheng, Lingjie Yang, Xiaohui Duan, Yunfang Yu, Yu Zhou 0027, Donald Donglong Chen
BIBM4
2024 SimpleFusion: A Simple Fusion Framework for Infrared and Visible Images
Yuxuan Cheng, Xinwei He 0001, Yan Aze, Jinhai Xiang
PRCV (8)2
2022 ByteGNN: Efficient Graph Neural Network Training at Large Scale
abstract
Graph neural networks (GNNs) have shown excellent performance in a wide range of applications such as recommendation, risk control, and drug discovery. With the increase in the volume of graph data, distributed GNN systems become essential to support efficient GNN training. However, existing distributed GNN training systems suffer from various performance issues including high network communication cost, low CPU utilization, and poor end-to-end performance. In this paper, we propose ByteGNN, which addresses the limitations in existing distributed GNN systems with three key designs: (1) an abstraction of mini-batch graph sampling to support high parallelism, (2) a two-level scheduling strategy to improve resource utilization and to reduce the end-to-end GNN training time, and (3) a graph partitioning algorithm tailored for GNN workloads. Our experiments show that ByteGNN outperforms the state-of-the-art distributed GNN systems with up to 3.5--23.8 times faster end-to-end execution, 2--6 times higher CPU utilization, and around half of the network communication cost.
Chenguang Zheng, Yuxuan Cheng, Zhezheng Song, Yifan Wu 0002, Changji Li, James Cheng
Proc. VLDB Endow.3