Xuesong Chen 0001

dblp:33/205-1 · DBLP profile ↗
← Back
16ranked-venue papers
8as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 8 first-author · 7 since 2021Artificial intelligence and machine learning · 10 · 7 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 M3Net: Multimodal Multi-task Learning for 3D Detection, Segmentation, and Occupancy Prediction in Autonomous Driving
abstract
The perception system for autonomous driving generally requires to handle multiple diverse sub-tasks. However, current algorithms typically tackle individual sub-tasks separately, which leads to low efficiency when aiming at obtaining full-perception results. Some multi-task learning methods try to unify multiple tasks with one model, but do not solve the conflicts in multi-task learning. In this paper, we introduce M3Net, a novel multimodal and multi-task network that simultaneously tackles detection, segmentation, and 3D occupancy prediction for autonomous driving and achieves superior performance than single task model. M3Net takes multimodal data as input and multiple tasks via query-token interactions. To enhance the integration of multi-modal features for multi-task learning, we first propose the Modality-Adaptive Feature Integration (MAFI) module, which enables single-modality features to predict channel-wise attention weights for their high-performing tasks, respectively. Based on integrated features, we then develop task-specific query initialization strategies to accommodate the needs of detection/segmentation and 3D occupancy prediction. Leveraging the properly initialized queries, a shared decoder transforms queries and BEV features layer-wise, facilitating multi-task learning. Furthermore, we propose a Task-oriented Channel Scaling (TCS) module in the decoder to mitigate conflicts between optimizing for different tasks. Additionally, our proposed multi-task querying and TCS module support both Transformer-based decoder and Mamba-based decoder, demonstrating its flexibility to different architectures. M3Net achieves state-of-the-art multi-task learning performance on the nuScenes benchmarks.
Xuesong Chen 0001, Shaoshuai Shi, Tao Ma 0002, Jingqiu Zhou, Simon See, Ka Chun Cheung, Hongsheng Li 0001
AAAI1
2025 GaussianPainter: Painting Point Cloud into 3D Gaussians with Normal Guidance
abstract
In this paper, we present GaussianPainter, the first method to paint a point cloud into 3D Gaussians given a reference image. GaussianPainter introduces an innovative feed-forward approach to overcome the limitations of time-consuming test-time optimization in 3D Gaussian splatting. Our method addresses a critical challenge in the field: the non-uniqueness problem inherent in the large parameter space of 3D Gaussian splatting. This space, encompassing rotation, anisotropic scales, and spherical harmonic coefficients, introduces the challenge of rendering similar images from substantially different Gaussian fields. As a result, feed-forward networks face instability when attempting to directly predict high-quality Gaussian fields, struggling to converge on consistent parameters for a given output. To address this issue, we propose to estimate a surface normal for each point to determine its Gaussian rotation. This strategy enables the network to effectively predict the remaining Gaussian parameters in the constrained space. We further enhance our approach with an appearance injection module, incorporating reference image appearance into Gaussian fields via a multiscale triplane representation. Our method successfully balances efficiency and fidelity in 3D Gaussian generation, achieving high-quality, diverse, and robust 3D content creation from point clouds in a single forward pass. A video is provided in our supplementary material for a more detailed explanation of our method.
Jingqiu Zhou, Lue Fan, Xuesong Chen 0001, Linjiang Huang, Si Liu 0001, Hongsheng Li 0001
AAAI3
2025 SOLVE: Synergy of Language-Vision and End-to-End Networks for Autonomous Driving
abstract
The integration of Vision-Language Models (VLMs) into autonomous driving systems has shown promise in addressing key challenges such as learning complexity, interpretability, and common-sense reasoning. However, existing approaches often struggle with efficient integration and real-time decision-making due to computational demands. In this paper, we introduce SOLVE, an innovative framework that synergizes VLMs with end-to-end (E2E) models to enhance autonomous vehicle planning. Our approach emphasizes knowledge sharing at the feature level through a shared visual encoder, enabling comprehensive interaction between VLM and E2E components. We propose a Trajectory Chain-of-Thought (T-CoT) paradigm, which progressively refines trajectory predictions, reducing uncertainty and improving accuracy. By employing a temporal decoupling strategy, SOLVE achieves efficient cooperation by aligning high-quality VLM outputs with E2E real-time performance. Evaluated on the nuScenes dataset, our method demonstrates significant improvements in trajectory prediction accuracy, paving the way for more robust and reliable autonomous driving systems.
Xuesong Chen 0001, Linjiang Huang, Tao Ma 0002, Rongyao Fang, Shaoshuai Shi, Hongsheng Li 0001
CVPR1
2023 TrajectoryFormer: 3D Object Tracking Transformer with Predictive Trajectory Hypotheses
abstract
3D multi-object tracking (MOT) is vital for many applications including autonomous driving vehicles and service robots. With the commonly used tracking-by-detection paradigm, 3D MOT has made important progress in recent years. However, these methods only use the detection boxes of the current frame to obtain trajectory-box association results, which makes it impossible for the tracker to recover objects missed by the detector. In this paper, we present TrajectoryFormer, a novel point-cloud-based 3D MOT framework. To recover the missed object by detector, we generates multiple trajectory hypotheses with hybrid candidate boxes, including temporally predicted boxes and current-frame detection boxes, for trajectory-box association. The predicted boxes can propagate object’s history trajectory information to the current frame and thus the network can tolerate short-term miss detection of the tracked objects. We combine long-term object motion feature and short-term object appearance feature to create per-hypothesis feature embedding, which reduces the computational overhead for spatial-temporal encoding. Additionally, we introduce a Global-Local Interaction Module to conduct information interaction among all hypotheses and models their spatial relations, leading to accurate estimation of hypotheses. Our TrajectoryFormer achieves state-of-the-art performance on the Waymo 3D MOT benchmarks. Code is available at https://github.com/poodarchu/EFG.
Xuesong Chen 0001, Shaoshuai Shi, Benjin Zhu, Qiang Wang 0023, Ka Chun Cheung, Simon See, Hongsheng Li 0001
ICCV1
2022 MPPNet: Multi-frame Feature Intertwining with Proxy Points for 3D Temporal Object Detection
Xuesong Chen 0001, Shaoshuai Shi, Benjin Zhu, Ka Chun Cheung, Hang Xu 0004, Hongsheng Li 0001
ECCV (8)1
2021 A Unified Multi-Scenario Attacking Network for Visual Object Tracking
abstract
Existing methods of adversarial attacks successfully generate adversarial examples to confuse Deep Neural Networks (DNNs) of image classification and object detection, resulting in wrong predictions. However, these methods are difficult to attack models of video object tracking, because the tracking algorithms could handle sequential information across video frames and the categories of targets tracked are normally unknown in advance. In this paper, we propose a Unified and Effective Network, named UEN, to attack visual object tracking models. There are several appealing characteristics of UEN: (1) UEN could produce various invisible adversarial perturbations according to different attack settings by using only one simple end-to-end network with three ingenious loss function; (2) UEN could generate general visible adversarial patch patterns to attack the advanced trackers in the real-world; (3) Extensive experiments show that UEN is able to attack many state-of-the-art trackers effectively (e.g. SiamRPN-based networks and DiMP) on popular tracking datasets including OTB100, UAV123, and GOT10K, making online real-time attacks possible. The attack results outperform the introduced baseline in terms of attacking ability and attacking efficiency.
Xuesong Chen 0001, Canmiao Fu, Feng Zheng 0001, Yong Zhao 0010, Hongsheng Li 0001, Ping Luo 0002, Guo-Jun Qi
AAAI1
2021 Semantic Scene Completion via Integrating Instances and Scene In-the-Loop
abstract
Semantic Scene Completion aims at reconstructing a complete 3D scene with precise voxel-wise semantics from a single-view depth or RGBD image. It is a crucial but challenging problem for indoor scene understanding. In this work, we present a novel framework named Scene-Instance-Scene Network (SISNet), which takes advantages of both in-stance and scene level semantic information. Our method is capable of inferring fine-grained shape details as well as nearby objects whose semantic categories are easily mixed-up. The key insight is that we decouple the instances from a coarsely completed semantic scene instead of a raw input image to guide the reconstruction of instances and the over-all scene. SISNet conducts iterative scene-to-instance (SI) and instance-to-scene (IS) semantic completion. Specifically, the SI is able to encode objects’ surrounding context for effectively decoupling instances from the scene and each instance could be voxelized into higher resolution to capture finer details. With IS, fine-grained instance information can be integrated back into the 3D scene and thus leads to more accurate semantic scene completion. Utilizing such an iterative mechanism, the scene and instance completion benefits each other to achieve higher completion accuracy. Extensively experiments show that our proposed method consistently outperforms state-of-the-art methods on both real NYU, NYUCAD and synthetic SUNCG-RGBD datasets. The code and the supplementary material will be available at https://github.com/yjcaimeow/SISNet.
Yingjie Cai, Xuesong Chen 0001, Kwan-Yee Lin, Xiaogang Wang 0001, Hongsheng Li 0001
CVPR2
2020 Salience-Guided Cascaded Suppression Network for Person Re-Identification
abstract
Employing attention mechanisms to model both global and local features as a final pedestrian representation has become a trend for person re-identification (Re-ID) algorithms. A potential limitation of these methods is that they focus on the most salient features, but the re-identification of a person may rely on diverse clues masked by the most salient features in different situations, e.g., body, clothes or even shoes. To handle this limitation, we propose a novel Salience-guided Cascaded Suppression Network (SCSN) which enables the model to mine diverse salient features and integrate these features into the final representation by a cascaded manner. Our work makes the following contributions: (i) We observe that the previously learned salient features may hinder the network from learning other important information. To tackle this limitation, we introduce a cascaded suppression strategy, which enables the network to mine diverse potential useful features that be masked by the other salient features stage-by-stage and each stage integrates different feature embedding for the last discriminative pedestrian representation. (ii) We propose a Salient Feature Extraction (SFE) unit, which can suppress the salient features learned in the previous cascaded stage and then adaptively extracts other potential salient feature to obtain different clues of pedestrians. (iii) We develop an efficient feature aggregation strategy that fully increases the network’s capacity for all potential salience features. Finally, experimental results demonstrate that our proposed method outperforms the state-of-the-art methods on four large-scale datasets. Especially, our approach exceeds the current best method by over 7% on the CUHK03 dataset.
Xuesong Chen 0001, Canmiao Fu, Yong Zhao 0010, Feng Zheng 0001, Jingkuan Song, Rongrong Ji, Yi Yang 0001
CVPR1
2020 One-Shot Adversarial Attacks on Visual Tracking With Dual Attention
abstract
Almost all adversarial attacks in computer vision are aimed at pre-known object categories, which could be offline trained for generating perturbations. But as for visual object tracking, the tracked target categories are normally unknown in advance. However, the tracking algorithms also have potential risks of being attacked, which could be maliciously used to fool the surveillance systems. Meanwhile, it is still a challenging task that adversarial attacks on tracking since it has the free-model tracked target. Therefore, to help draw more attention to the potential risks, we study adversarial attacks on tracking algorithms. In this paper, we propose a novel one-shot adversarial attack method to generate adversarial examples for free-model single object tracking, where merely adding slight perturbations on the target patch in the initial frame causes state-of-the-art trackers to lose the target in subsequent frames. Specifically, the optimization objective of the proposed attack consists of two components and leverages the dual attention mechanisms. The first component adopts a targeted attack strategy by optimizing the batch confidence loss with confidence attention while the second one applies a general perturbation strategy by optimizing the feature loss with channel attention. Experimental results show that our approach can significantly lower the accuracy of the most advanced Siamese network-based trackers on three benchmarks.
Xuesong Chen 0001, Xiyu Yan, Feng Zheng 0001, Yong Jiang 0001, Shutao Xia, Yong Zhao 0010, Rongrong Ji
CVPR1
2020 Hijacking Tracker: A Powerful Adversarial Attack on Visual Tracking
abstract
Visual object tracking has made important breakthroughs with the assistance of deep learning models. Unfortunately, recent research has clearly proved that deep learning models are vulnerable to malicious adversarial attacks, which mislead the models making wrong decisions by perturbing the input image. The threat to the models alerts us to pay attention to the model security of deep learning- based tracking algorithms. Therefore, we study the adversarial attacks against advanced trackers based on deep learning to better identify the vulnerability of tracking algorithms. In this paper, we propose to add slight adversarial perturbations to the input image by an inconspicuous but powerful attack strategy-hijacking algorithm. Specifically, the hijacking strategy misleads trackers in two aspects: one is shape hijacking that changes the shape of the model output; the other is position hijacking that gradually pushes the output to any position in the image frame. Besides, we further propose an adaptive optimization approach to integrate two hijacking mechanisms efficiently. Eventually, the hijacking algorithm results in fooling the tracker to track the wrong target gradually. The experimental results demonstrate the powerful attack ability of our method-quickly hijacking state-of-the-art trackers and reducing the accuracy of these models by more than 90% on OTB2015.
Xiyu Yan, Xuesong Chen 0001, Yong Jiang 0001, Shutao Xia, Yong Zhao 0010, Feng Zheng 0001
ICASSP2
2019 Scanet: Spatial-channel Attention Network for 3D Object Detection
abstract
This paper aims to achieve high-accuracy 3D object detection, in which we propose a novel Spatial-Channel Attention Network (SCANet), a two-stage detector that takes both LIDAR point clouds and RGB images as input to generate 3D object estimates. The first stage is a 3D region proposal network (RPN) in which we put forward a new Spatial-Channel Attention (SCA) module and an Extension Spatial Upsample (ESU) module. Using the pyramid pooling structure and global average pooling, the SCA module can not only effectively incorporate multi-scale and global context information, but also produce spatial and channel-wise attention to select discriminative features. The ESU module in the decoder can recover the lost spatial information caused by consecutive pooling operators to generate reliable 3D region proposals. In the second stage, we design a new multi-level fusion scheme for accurate classification and 3D bounding box regression. Experimental results demonstrate that SCANet achieves state-of-the-art performance on the challenging KITTI 3D object detection benchmark.
Haihua Lu, Xuesong Chen 0001, Guiying Zhang, Qiuhao Zhou, Yanbo Ma, Yong Zhao 0010
ICASSP2
2019 Multi-attention Network for Thoracic Disease Classification and Localization
abstract
The chest X-ray is one of the most commonly available radiological examinations for diagnosing lung diseases. This task remains a major challenge due to 1) the shortage of accurate annotations for chest X-ray examinations, 2) the diversity of lesion areas on X-rays from different thoracic disease and 3) the problem of class imbalance in existing chest X-ray databases. In this paper, we propose a new multi-attention convolutional neural network for thoracic disease classification and localization. First, the framework is equipped with squeeze-and-excitation (SE) block as a feature attention module to offer a chance of cross-channel feature recalibration. Second, we propose a novel space attention module to combine global and local information. Third, we present a hard examples attention module to alleviate the class imbalance problem. The comprehensive experiments are performed on the ChestX-ray14 dataset. Quantitative and qualitative results demonstrate that our method outperforms the state-of-the-art algorithm.
Yanbo Ma, Qiuhao Zhou, Xuesong Chen 0001, Haihua Lu, Yong Zhao 0010
ICASSP3
2019 Discriminative Features Reconstruction Network for Semantic Segmentation
abstract
Thanks to the development of convolutional neural networks (CNNs), researchers have proposed lots of effective semantic segmentation models. However, there are still two problems disturbing researchers, one of which is objects misidentification on the image level and another one is poor performance on details, especially the boundary of objects. To tackle these two problems, we propose Discriminative Features Reconstruction Network (DFR) containing two modules: Second-order Pyramid Features Reconstruction Module (SPFR) and Second-order Boundary Attention Module (SBA). Specifically, SPFR fuses different scales features to gain pyramid receptive field. Besides, SPFR extracts second-order statistics data to retrieve more discriminative features. Furthermore, we put forward SBA that is helpful to refine the segmentation results. On SBA, low-level features recover localization details under the high-level feature guidance. Our DFR achieves state-of-the-art performance on PASCAL VOC 2012 dataset with mIoU accuracy 81.1% without pre-training COCO dataset and post-processing.
Qiuhao Zhou, Yanbo Ma, Haihua Lu, Xuesong Chen 0001, Yong Zhao 0010
ICASSP4
2019 Sequentially Refined Spatial and Channel-Wise Feature Aggregation in Encoder-Decoder Network for Single Image Dehazing
abstract
Single image haze removal is a challenging problem due to its inherent ill-posed nature. Several prior-based and learning-based methods have been proposed to solve this problem and they have achieved superior results. However, the performance of prior-based methods is limited by hand-designed features. Meanwhile, the learning-based methods have flaws in spatial correlation. Because they just employ the convolutional neural network (CNN) as an end-to-end mapping module, generating dehazed image directly or learning parameters in atmospheric scattering model, without associating the relationship between feature activation and distribution of haze. Instead of using prior or CNN to simply estimate parameters of atmospheric scattering model, we propose a sequentially refined Spatial and Channel-wise Feature Aggregation (SCFA) dehazing network, called SCFADN. Specifically, our network learns residues between clean images and hazy images through an improved encoder-decoder network which incorporated the proposed sequentially refined SCFA module. This module efficiently learns long-range feature correlation for dehazing residue modeling, which enables the shallow network removing haze effectively. Extensive experiments on synthetic and real datasets results demonstrate that the proposed method achieves superior performance over the state-of-the-art methods.
Xuesong Chen 0001, Haihua Lu, Kaili Cheng, Yanbo Ma, Qiuhao Zhou, Yong Zhao 0010
ICIP1
2019 ACNet: Aggregated Channels Network for Automated Mitosis Detection
Kaili Cheng, Xuesong Chen 0001, Yanbo Ma, Mengjie Bai, Yong Zhao 0010
PAKDD (1)3
2018 CHS-NET: A Cascaded Neural Network with Semi-Focal Loss for Mitosis Detection
abstract
Counting of mitotic figures in hematoxylin and eosin(H&E) stained histological slide is the main indicator of tumor proliferation speed which is an important biomarker indicative of breast cancer patients’ prognosis. It is difficult to detect mitotic cells due to the diversity of the cells and the problem of class imbalance. We propose a new network called CHS-NET which is a cascaded neural network with hard example mining and semi-focal loss to detect mitotic cells in breast cancer. First, we propose a screening network to identify the candidates of mitotic cells preliminary and a refined network to identify mitotic cells from these candidates more accurately. We propose a new feature fusion module in each network to explore complex nonlinear predictors and improve accuracy. Then, we propose a novel loss named semi-focal loss and we use off-line hard example mining to solve the problem of class imbalance and error labeling. Finally, we propose a new training skill of cutting patches in the whole slide image, considering the size and distribution of mitotic cells. Our method achieves 0.68 F1 score which outperforms the best result in Tumor Proliferation Assessment Challenge 2016 held by MICCAI.
Yanbo Ma, Qiuhao Zhou, Kaili Cheng, Xuesong Chen 0001, Yong Zhao 0010
ACML5