Rui Xu 0021

dblp:00/4859-21 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
6since 2021 · last 2026
0009-0002-1364-6305ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 DilatedTAD: Enhancing Adaptability to Actions of Varying Durations for Temporal Action Detection
abstract
Temporal Action Detection (TAD) aims to identify action boundaries and their corresponding categories in untrimmed videos, playing a crucial role in long-video understanding. Prior works often struggle to balance the trade-off between capturing long-range dependencies and ensuring computational efficiency. Recently, the state space model Mamba has exhibited impressive capabilities and efficiency in long-term sequence modeling. However, current methods based on Mamba generally lack a unified framework to simultaneously address the redundancy of long-duration actions and the boundary sensitivity of short-duration actions—limitations that largely stem from Mamba’s reliance on limited state representations and its unidirectional modeling. To tackle the aforementioned challenges, we propose DilatedTAD, a novel TAD framework with an expanded receptive field. DilatedTAD leverages the Inter-Parallel DIM component (InterDIM) to integrate multi-scale temporal information, enabling a better trade-off between short-duration and long-duration action detection. InterDIM is built upon our proposed Dilated Mamba (DIM), where multiple DIM branches with different dilation rates are designed to focus on actions of varying durations. Specifically, DIM introduces a novel use of dilation to skip redundant temporal information, thereby enhancing the model’s focus on crucial boundary features. Additionally, a bidirectional modeling design is adopted in DIM to compensate for the lack of future temporal context in the original Mamba architecture. Extensive experiments show that DilatedTAD outperforms state-of-the-art methods on multiple datasets, achieving mAPs of 74.9% (THUMOS14), 42.90% (ActivityNet 1.3), 45.0% (HACS), and 26.3% and 24.3% (EPIC-Kitchens 100). Our code will be publicly available.
Longyang Tang, Bo Zhang 0096, Rui Xu 0021, Junsheng Zhou, Yi Chen 0023
IEEE Trans. Circuits Syst. Video Technol.4
2024 Multi-Attribute Interactions Matter for 3D Visual Grounding
abstract
3D visual grounding aims to localize 3D objects described by free-form language sentences. Following the detection-then-matching paradigm, existing methods mainly focus on embedding object attributes in unimodal feature extraction and multimodal feature fusion, to enhance the discriminability of the proposal feature for accurate grounding. However, most of them ignore the explicit interaction of multiple attributes, causing a bias in unimodal representation and misalignment in multimodal fusion. In this paper, we propose a multi-attribute aware Transformer for 3D visual grounding, learning the multi-attribute interactions to refine the intra-modal and inter-modal grounding cues. Specifically, we first develop an attribute causal analysis module to quantify the causal effect of different attributes for the final prediction, which provides powerful supervision to correct the misleading attributes and adaptively capture other discriminative features. Then, we design an exchanging-based multimodal fusion module, which dynamically replaces tokens with low attribute attention between modalities before directly integrating low-dimensional global features. This ensures an attribute-level multimodal information fusion and helps align the language and vision details more efficiently for fine-grained multimodal features. Extensive experiments show that our method can achieve state-of-the-art performance on ScanRefer and Sr3D/Nr3D datasets. The code is publicly available at https://github.com/volcanoXC/MA2TransVG.
Can Xu 0006, Yuehui Han, Rui Xu 0021, Le Hui, Jin Xie 0001, Jian Yang 0003
CVPR3
2024 Masked Motion Prediction with Semantic Contrast for Point Cloud Sequence Learning
Yuehui Han, Can Xu 0006, Rui Xu 0021, Jianjun Qian, Jin Xie 0001
ECCV (76)3
2023 Transformer-based Point Cloud Generation Network
abstract
Point cloud generation is an important research topic in 3D computer vision, which can provide high-quality datasets for various downstream tasks. However, efficiently capturing the geometry of point clouds remains a challenging problem due to their irregularities. In this paper, we propose a novel transformer-based 3D point cloud generation network to generate realistic point clouds. Specifically, we first develop a transformer-based interpolation module that utilizes k-nearest neighbors at different scales to learn global and local information about point clouds in the feature space. Based on geometric information, we interpolate new point features to upsample the point cloud features. Then, the upsampled features are used to generate a coarse point cloud with spatial coordinate information. We construct a transformer-based refinement module to enhance the upsampled features in feature space with geometric information in coordinate space. Finally, we use a multi-layer perceptron on the upsampled features to generate the final point cloud. Extensive experiments on ShapeNet and ModelNet demonstrate the effectiveness of our proposed method.
Rui Xu 0021, Le Hui, Yuehui Han, Jianjun Qian, Jin Xie 0001
ACM Multimedia1
2023 Scene Graph Masked Variational Autoencoders for 3D Scene Generation
abstract
Generating realistic 3D indoor scenes requires a deep understanding of objects and their spatial relationships. However, existing methods often fail to generate realistic 3D scenes due to the limited understanding of object relationships. To tackle this problem, we propose a Scene Graph Masked Variational Auto-Encoder (SG-MVAE) framework that fully captures the relationships between objects to generate more realistic 3D scenes. Specifically, we first introduce a relationship completion module that adaptively learns the missing relationships between objects in the scene graph. To accurately predict the missing relationships, we employ multi-group attention to capture the correlations between the objects with missing relationships and other objects in the scene. After obtaining the complete scene relationships, we mask the relationships between objects and use a decoder to reconstruct the scene. The reconstruction process enhances the model's understanding of relationships, generating more realistic scenes. Extensive experiments on benchmark datasets show that our model outperforms state-of-the-art methods.
Rui Xu 0021, Le Hui, Yuehui Han, Jianjun Qian, Jin Xie 0001
ACM Multimedia1
2022 Domain Disentangled Generative Adversarial Network for Zero-Shot Sketch-Based 3D Shape Retrieval
abstract
Sketch-based 3D shape retrieval is a challenging task due to the large domain discrepancy between sketches and 3D shapes. Since existing methods are trained and evaluated on the same categories, they cannot effectively recognize the categories that have not been used during training. In this paper, we propose a novel domain disentangled generative adversarial network (DD-GAN) for zero-shot sketch-based 3D retrieval, which can retrieve the unseen categories that are not accessed during training. Specifically, we first generate domain-invariant features and domain-specific features by disentangling the learned features of sketches and 3D shapes, where the domain-invariant features are used to align with the corresponding word embeddings. Then, we develop a generative adversarial network that combines the domain-specific features of the seen categories with the aligned domain-invariant features to synthesize samples, where the synthesized samples of the unseen categories are generated by using the corresponding word embeddings. Finally, we use the synthesized samples of the unseen categories combined with the real samples of the seen categories to train the network for retrieval, so that the unseen categories can be recognized. In order to reduce the domain shift problem, we utilize unlabeled unseen samples to enhance the discrimination ability of the discriminator. With the discriminator distinguishing the generated samples from the unlabeled unseen samples, the generator can generate more realistic unseen samples. Extensive experiments on the SHREC'13 and SHREC'14 datasets show that our method significantly improves the retrieval performance of the unseen categories.
Rui Xu 0021, Zongyan Han, Le Hui, Jianjun Qian, Jin Xie 0001
AAAI1
2020 Progressive Point Cloud Deconvolution Generation Network
Le Hui, Rui Xu 0021, Jin Xie 0001, Jianjun Qian, Jian Yang 0003
ECCV (15)2