EDBT 2026 Demo / reviewers in the wild / expert
Shaoxiang Chen 0001
dblp:04/2928-1
· DBLP profile ↗
19ranked-venue papers
8as first author
13since 2021 · last 2026
0000-0002-7627-7124ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 8 first-author · 10 since 2021Artificial intelligence and machine learning · 14 · 7 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ChatTracker: Enhancing Visual Tracking via LLM-Driven Iterative Description RefinementabstractVisual object tracking focuses on locating a target object within a video sequence based on an initial bounding box. Recently, Vision-Language (VL) trackers have been proposed to utilize additional natural language descriptions to enhance versatility in various applications. Despite this potential, VL trackers still underperform the State-of-the-Art (SoTA) visual trackers in terms of tracking accuracy. We find that this inferiority is primarily due to their heavy reliance on manual textual annotations, which include the frequent provision of ambiguous language descriptions. In this paper, we identify, for the first time, that over 10% of textual annotations in existing VL tracking datasets suffer from inaccuracies through manual evaluation. To address this problem, we propose ChatTracker to leverage the wealth of world knowledge in the Multimodal Large Language Model (MLLM) to generate high-quality language descriptions and enhance tracking performance. To this end, we propose a novel Reflection-based Language Description Refinement Module to iteratively refine the ambiguous and inaccurate descriptions of the target with tracking feedback. To further utilize semantic information produced by MLLM, a simple yet effective VL tracking framework is proposed, which can be easily integrated as a plug-and-play module to boost the performance of both VL and visual trackers. Experimental results show that ChatTracker achieves comparable performance to existing SoTA tracking methods. In addition, language descriptions generated by ChatTracker enhance the performance of various VL trackers and exhibit better text-to-image alignment than annotations in the original dataset. Moreover, our proposed framework can improve the performance of various visual tasks, including Referring Expression Comprehension (REC), Referring Expression Segmentation (RES), and Referring Video Object Segmentation (R-VOS) tasks by providing more accurate language descriptions, which demonstrates the universality of ChatTracker. We release the manual evaluation results and the generated textual descriptions, aiming to drive advancements in VL tracking. Yiming Sun 0006, Mi Zhang 0001, Shaoxiang Chen 0001, Yang Li 0041, Changbo Wang, Jianke Zhu, Steven C. H. Hoi |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Instance-Aware Multi-Camera 3D Object Detection with Structural Priors Mining and Self-Boosting LearningabstractCamera-based bird-eye-view (BEV) perception paradigm has made significant progress in the autonomous driving field. Under such a paradigm, accurate BEV representation construction relies on reliable depth estimation for multi-camera images. However, existing approaches exhaustively predict depths for every pixel without prioritizing objects, which are precisely the entities requiring detection in the 3D space. To this end, we propose IA-BEV, which integrates image-plane instance awareness into the depth estimation process within a BEV-based detector. First, a category-specific structural priors mining approach is proposed for enhancing the efficacy of monocular depth generation. Besides, a self-boosting learning strategy is further proposed to encourage the model to place more emphasis on challenging objects in computation-expensive temporal stereo matching. Together they provide advanced depth estimation results for high-quality BEV features construction, benefiting the ultimate 3D detection. The proposed method achieves state-of-the-art performances on the challenging nuScenes benchmark, and extensive experimental results demonstrate the effectiveness of our designs. Zequn Jie, Shaoxiang Chen 0001, Lechao Cheng, Jingjing Chen 0001, Lin Ma 0002, Yu-Gang Jiang 0001 |
AAAI | 3 |
| 2024 | Making Large Language Models Better Planners with Reasoning-Decision Alignment
Shaoxiang Chen 0001, Sihao Lin, Zequn Jie, Lin Ma 0002, Guangrun Wang, Xiaodan Liang |
ECCV (36) | 3 |
| 2024 | Lumen: Unleashing Versatile Vision-Centric Capabilities of Large Multimodal ModelsabstractLarge Multimodal Model (LMM) is a hot research topic in the computer vision area and has also demonstrated remarkable potential across multiple disciplinary fields. A recent trend is to further extend and enhance the perception capabilities of LMMs. The current methods follow the paradigm of adapting the visual task outputs to the format of the language model, which is the main component of a LMM. This adaptation leads to convenient development of such LMMs with minimal modifications, however, it overlooks the intrinsic characteristics of diverse visual tasks and hinders the learning of perception capabilities. To address this issue, we propose a novel LMM architecture named Lumen, a Large multimodal model with versatile vision-centric capability enhancement. We decouple the LMM's learning of perception capabilities into task-agnostic and task-specific stages. Lumen first promotes fine-grained vision-language concept alignment, which is the fundamental capability for various visual tasks. Thus the output of the task-agnostic stage is a shared representation for all the tasks we address in this paper. Then the task-specific decoding is carried out by flexibly routing the shared representation to lightweight task decoders with negligible training efforts. Comprehensive experimental results on a series of vision-centric and VQA benchmarks indicate that our Lumen model not only achieves or surpasses the performance of existing LMM-based approaches in a range of vision-centric tasks while maintaining general visual understanding and instruction following capabilities. Shaoxiang Chen 0001, Zequn Jie, Jingjing Chen 0001, Lin Ma 0002, Yu-Gang Jiang 0001 |
NeurIPS | 2 |
| 2024 | ChatTracker: Enhancing Visual Tracking Performance via Chatting with Multimodal Large Language ModelabstractVisual object tracking aims to locate a targeted object in a video sequence based on an initial bounding box. Recently, Vision-Language~(VL) trackers have proposed to utilize additional natural language descriptions to enhance versatility in various applications. However, VL trackers are still inferior to State-of-The-Art (SoTA) visual trackers in terms of tracking performance. We found that this inferiority primarily results from their heavy reliance on manual textual annotations, which include the frequent provision of ambiguous language descriptions. In this paper, we propose ChatTracker to leverage the wealth of world knowledge in the Multimodal Large Language Model (MLLM) to generate high-quality language descriptions and enhance tracking performance. To this end, we propose a novel reflection-based prompt optimization module to iteratively refine the ambiguous and inaccurate descriptions of the target with tracking feedback. To further utilize semantic information produced by MLLM, a simple yet effective VL tracking framework is proposed and can be easily integrated as a plug-and-play module to boost the performance of both VL and visual trackers. Experimental results show that our proposed ChatTracker achieves a performance comparable to existing methods. Yiming Sun 0006, Shaoxiang Chen 0001, Junwei Huang, Yang Li 0041, Chenhui Li 0001, Changbo Wang |
NeurIPS | 3 |
| 2023 | MSMDFusion: Fusing LiDAR and Camera at Multiple Scales with Multi-Depth Seeds for 3D Object DetectionabstractFusing LiDAR and camera information is essential for accurate and reliable 3D object detection in autonomous driving systems. This is challenging due to the difficulty of combining multi-granularity geometric and semantic features from two drastically different modalities. Recent approaches aim at exploring the semantic densities of camera features through lifting points in 2D camera images (referred to as “seeds”) into 3D space, and then incorporate 2D semantics via cross-modal interaction or fusion techniques. However, depth information is under-investigated in these approaches when lifting points into 3D space, thus 2D semantics can not be reliably fused with 3D points. Moreover, their multi-modal fusion strategy, which is implemented as concatenation or attention, either can not effectively fuse 2D and 3D information or is unable to perform fine-grained interactions in the voxel space. To this end, we propose a novel framework with better utilization of the depth information and fine-grained cross-modal interaction between LiDAR and camera, which consists of two important components. First, a Multi-Depth Unprojection (MDU) method is used to enhance the depth quality of the lifted points at each interaction level. Second, a Gated Modality-Aware Convolution (GMA-Conv) block is applied to modulate voxels involved with the camera modality in a fine-grained manner and then aggregate multi-modal features into a unified space. Together they provide the detection head with more comprehensive features from LiDAR and camera. On the nuScenes test benchmark, our proposed method, abbreviated as MSMD-Fusion, achieves state-of-the-art results on both 3D object detection and tracking tasks without using test-time-augmentation and ensemble techniques. The code is available at https://github.com/SxJyJay/MSMDFusion. Zequn Jie, Shaoxiang Chen 0001, Jingjing Chen 0001, Lin Ma 0002, Yu-Gang Jiang 0001 |
CVPR | 3 |
| 2023 | Self-Supervised Learning for Semi-Supervised Temporal Language GroundingabstractGiven a text description, Temporal Language Grounding (TLG) aims to localize temporal boundaries of the segments that contain the specified semantics in an untrimmed video. TLG is inherently a challenging task, as it requires comprehensive understanding of both sentence semantics and video contents. Previous works either tackle this task in a fully-supervised setting that requires a large amount of temporal annotations or in a weakly-supervised setting that usually cannot achieve satisfactory performance. Since manual annotations are expensive, to cope with limited annotations, we tackle TLG in a semi-supervised way by incorporating self-supervised learning, and proposeSelf-SupervisedSemi-SupervisedTemporalLanguageGrounding (S$^{4}$TLG). S$^{4}$TLG consists of two parts: (1) A pseudo label generation module that adaptively produces instant pseudo labels for unlabeled samples based on predictions from a teacher model; (2) A self-supervised feature learning module with inter-modal and intra-modal contrastive losses to learn video feature representations under the constraints of video content consistency and video-text alignment. We conduct extensive experiments on the ActivityNet-CD-OOD and Charades-CD-OOD datasets. The results demonstrate that our proposed S$^{4}$TLG can achieve competitive performance compared to fully-supervised state-of-the-art methods while only requiring a small portion of temporal annotations. Fan Luo 0004, Shaoxiang Chen 0001, Jingjing Chen 0001, Zuxuan Wu, Yu-Gang Jiang 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Scene Graph Refinement Network for Visual Question AnsweringabstractVisual Question Answering aims to answer the free-form natural language question based on the visual clues in a given image. It is a difficult problem as it requires understanding the fine-grained structured information of both language and image for compositional reasoning. To establish the compositional reasoning, recent works attempt to introduce the scene graph in VQA. However, as the generated scene graphs are usually quite noisy, it greatly limits the performance of question answering. Therefore, this paper proposes to refine the scene graphs for improving the effectiveness. Specifically, we present a novelSceneGraphRefinement network (SGR), which introduces a transformer-based refinement network to enhance the object and relation features for better classification. Moreover, as the question provides valuable clues for distinguishing whether the$\left\langle \mathit{subject, predicate, object} \right\rangle$triplets are helpful or not, the SGR network exploits the semantic information presented in the questions to select the most relevant relations for question answering. Extensive experiments are conducted on the GQA benchmark demonstrate the effectiveness of our method. Tianwen Qian, Jingjing Chen 0001, Shaoxiang Chen 0001, Bo Wu 0018, Yu-Gang Jiang 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | FT-TDR: Frequency-Guided Transformer and Top-Down Refinement Network for Blind Face InpaintingabstractBlind face inpainting refers to the task of reconstructing visual contents without explicitly indicating the corrupted regions in a face image. Inherently, this task faces two challenges: (1) how to detect various mask patterns of different shapes and contents; (2) how to restore visually plausible and pleasing contents in the masked regions. In this paper, we propose a novel two-stage blind face inpainting method named Frequency-guided Transformer and Top-Down Refinement Network (FT-TDR) to tackle these challenges. Specifically, we first use a transformer-based network to detect the corrupted regions to be inpainted as masks by modeling the relation among different patches. For improved detection results, we also exploit the frequency modality as complementary information and capture the local contextual incoherence to enhance boundary consistency. Then a top-down refinement network is proposed to hierarchically restore features at different levels and generate contents that are semantically consistent with the unmasked face regions. Extensive experiments demonstrate that our method outperforms current state-of-the-art blind and non-blind face inpainting methods qualitatively and quantitatively. Shaoxiang Chen 0001, Zuxuan Wu, Yu-Gang Jiang 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | MORE: Multi-Order RElation Mining for Dense Captioning in 3D Scenes
Shaoxiang Chen 0001, Zequn Jie, Jingjing Chen 0001, Lin Ma 0002, Yu-Gang Jiang 0001 |
ECCV (35) | 2 |
| 2021 | Towards Bridging Event Captioner and Sentence Localizer for Weakly Supervised Dense Event CaptioningabstractDense Event Captioning (DEC) aims to jointly localize and describe multiple events of interest in untrimmed videos, which is an advancement of the conventional video captioning task (generating a single sentence description for a trimmed video). Weakly Supervised Dense Event Captioning (WS-DEC) goes one step further by not relying on human-annotated temporal event boundaries. However, there are few methods trying to tackle this task, and how to connect localization and description remains an open problem. In this paper, we demonstrate that under weak supervision, the event captioning module and localization module should be more closely bridged in order to improve description performance. Different from previous approaches, in our method, the event captioner generates a sentence from a video segment and feeds it to the sentence localizer to reconstruct the segment, and the localizer produces word importance weights as a guidance for the captioner to improve event description. To further bridge the sentence localizer and event captioner, a concept learner is adopted as the basis of the sentence localizer, which can be utilized to construct an induced set of concept features to enhance video features and improve the event captioner. Finally, our proposed method outperforms state-of-the-art WS-DEC methods on the ActivityNet Captions dataset. Shaoxiang Chen 0001, Yu-Gang Jiang 0001 |
CVPR | 1 |
| 2021 | Motion Guided Region Message Passing for Video CaptioningabstractVideo captioning is an important vision task and has been intensively studied in the computer vision community. Existing methods that utilize the fine-grained spatial information have achieved significant improvements, however, they either rely on costly external object detectors or do not sufficiently model the spatial/temporal relations. In this paper, we aim at designing a spatial information extraction and aggregation method for video captioning without the need of external object detectors. For this purpose, we propose a Recurrent Region Attention module to better extract diverse spatial features, and by employing Motion-Guided Cross-frame Message Passing, our model is aware of the temporal structure and able to establish high-order relations among the diverse regions across frames. They jointly encourage information communication and produce compact and powerful video representations. Furthermore, an Adjusted Temporal Graph Decoder is proposed to flexibly update video features and model high-order temporal relations during decoding. Experimental results on three benchmark datasets: MSVD, MSR-VTT, and VATEX demonstrate that our proposed method can outperform state-of-the-art methods. Shaoxiang Chen 0001, Yu-Gang Jiang 0001 |
ICCV | 1 |
| 2021 | Towards Bridging Video and Language by Caption Generation and Sentence LocalizationabstractVarious video understanding tasks (classification, tracking, action detection, etc.) have been extensively studied in the multimedia and computer vision communities over the recent years. While these tasks are important, we think that bridging video and language is a more natural and intuitive way to interact with videos. Caption generation and sentence localization are two representative tasks for connecting video and language, and my research is focused on these two tasks. In this extended abstract, I present approaches for tackling each of these tasks by exploiting fine-grained information in videos, together with ideas about how these two tasks can be connected. So far, my work have demonstrated that these two tasks share a common foundation, and by connecting them to form a cycle, video and language can be more closely bridged. Finally, several challenges and future directions will be discussed. Shaoxiang Chen 0001 |
ACM Multimedia | 1 |
| 2020 | Hierarchical Visual-Textual Graph for Temporal Activity Localization via Language
Shaoxiang Chen 0001, Yu-Gang Jiang 0001 |
ECCV (20) | 1 |
| 2020 | Learning Modality Interaction for Temporal Sentence Localization and Event Captioning in Videos
Shaoxiang Chen 0001, Wei Liu 0005, Yu-Gang Jiang 0001 |
ECCV (4) | 1 |
| 2019 | Motion Guided Spatial Attention for Video CaptioningabstractSequence-to-sequence models incorporated with attention mechanism have shown promising improvements on video captioning. While there is rich information both inside and between frames, spatial attention is rarely explored and motion information is usually handled by 3D-CNNs as just another modality for fusion. On the other hand, researches about human perception suggest that apparent motion can attract attention. Motivated by this, we aim to learn spatial attention on video frames under the guidance of motion information for caption generation. We present a novel video captioning framework by utilizing Motion Guided Spatial Attention (MGSA). The proposed MGSA exploits the motion between video frames by learning spatial attention from stacked optical flow images with a custom CNN. To further relate the spatial attention maps of video frames, we designed a Gated Attention Recurrent Unit (GARU) to adaptively incorporate previous attention maps. The whole framework can be trained in an end-to-end manner. We evaluate our approach on two benchmark datasets, MSVD and MSR-VTT. The experiments show that our designed model can generate better video representation and state of the art results are obtained under popular evaluation metrics such as BLEU@4, CIDEr, and METEOR. Shaoxiang Chen 0001, Yu-Gang Jiang 0001 |
AAAI | 1 |
| 2019 | Semantic Proposal for Activity Localization in Videos via Sentence QueryabstractThis paper presents an efficient algorithm to tackle temporal localization of activities in videos via sentence queries. The task differs from traditional action localization in three aspects: (1) Activities are combinations of various kinds of actions and may span a long period of time. (2) Sentence queries are not limited to a predefined list of classes. (3) The videos usually contain multiple different activity instances. Traditional proposal-based approaches for action localization that only consider the class-agnostic “actionness” of video snippets are insufficient to tackle this task. We propose a novel Semantic Activity Proposal (SAP) which integrates the semantic information of sentence queries into the proposal generation process to get discriminative activity proposals. Visual and semantic information are jointly utilized for proposal ranking and refinement. We evaluate our algorithm on the TACoS dataset and the Charades-STA dataset. Experimental results show that our algorithm outperforms existing methods on both datasets, and at the same time reduces the number of proposals by a factor of at least 10. Shaoxiang Chen 0001, Yu-Gang Jiang 0001 |
AAAI | 1 |
| 2019 | Deep Learning for Video Captioning: A ReviewabstractDeep learning has achieved great successes in solving specific artificial intelligence problems recently. Substantial progresses are made on Computer Vision (CV) and Natural Language Processing (NLP). As a connection between the two worlds of vision and language, video captioning is the task of producing a natural-language utterance (usually a sentence) that describes the visual content of a video. The task is naturally decomposed into two sub-tasks. One is to encode a video via a thorough understanding and learn visual representation. The other is caption generation, which decodes the learned representation into a sequential sentence, word by word. In this survey, we first formulate the problem of video captioning, then review state-of-the-art methods categorized by their emphasis on vision or language, and followed by a summary of standard datasets and representative approaches. Finally, we highlight the challenges which are not yet fully understood in this task and present future research directions. Shaoxiang Chen 0001, Ting Yao 0003, Yu-Gang Jiang 0001 |
IJCAI | 1 |
| 2019 | Black-box Adversarial Attacks on Video Recognition ModelsabstractDeep neural networks (DNNs) are known for their vulnerability to adversarial examples. These are examples that have undergone small, carefully crafted perturbations, and which can easily fool a DNN into making misclassifications at test time. Thus far, the field of adversarial research has mainly focused on image models, under either a white-box setting, where an adversary has full access to model parameters, or a black-box setting where an adversary can only query the target model for probabilities or labels. Whilst several white-box attacks have been proposed for video models, black-box video attacks are still unexplored. To close this gap, we propose the first black-box video attack framework, called V-BAD. V-BAD utilizestentative perturbations transferred from image models andpartition-based rectifications found by the NES to obtain good adversarial gradient estimates with fewer queries to the target model. V-BAD is equivalent to estimating the projection of the adversarial gradient on a selected subspace. Using three benchmark video datasets, we demonstrate that V-BAD can craft both untargeted and targeted attacks to fool two state-of-the-art deep video recognition models. For the targeted attack, it achieves $>$93% success rate using only an average of $3.4 \sim 8.4 \times 10^4$ queries, a similar number of queries to state-of-the-art black-box image attacks. This is despite the fact that videos often have two orders of magnitude higher dimensionality than static images. We believe that V-BAD is a promising new tool to evaluate and improve the robustness of video recognition models to black-box adversarial attacks. Linxi Jiang, Xingjun Ma, Shaoxiang Chen 0001, James Bailey 0001, Yu-Gang Jiang 0001 |
ACM Multimedia | 3 |