Dongsheng Jiang

dblp:85/8729 · DBLP profile ↗
← Back
21ranked-venue papers
0as first author
21since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Visual Bridge: Universal Visual Perception Representations Generating
abstract
Recent advances in diffusion models have achieved remarkable success in isolated computer vision tasks such as text-to-image generation, depth estimation, and optical flow. However, these models are often restricted by a ``single-task-single-model'' paradigm, severely limiting their generalizability and scalability in multi-task scenarios. Motivated by the cross-domain generalization ability of large language models, we propose a universal visual perception framework based on flow matching that can generate diverse visual representations across multiple tasks. Our approach formulates the process as a universal flow-matching problem from image patch tokens to task-specific representations rather than an independent generation or regression problem. By leveraging a strong self-supervised foundation model as the anchor and introducing a multi-scale, circular task embedding mechanism, our method learns a universal velocity field to bridge the gap between heterogeneous tasks, supporting efficient and flexible representation transfer. Extensive experiments on classification, detection, segmentation, depth estimation, and image-text retrieval demonstrate that our model achieves competitive performance in both zero-shot and fine-tuned settings, outperforming prior generalist and several specialist models. Ablation studies further validate the robustness, scalability, and generalization of our framework. Our work marks a significant step towards general-purpose visual perception, providing a solid foundation for future research in universal vision modeling.
Shuguang Dou, Junzhou Li, Zhiheng Yu, Dongsheng Jiang, Shugong Xu
AAAI6
2026 SAM2MOT: A Novel Paradigm of Multi-Object Tracking by Segmentation
abstract
Inspired by Segment Anything 2, which generalizes segmentation from images to videos, we propose SAM2MOT—a novel segmentation-driven paradigm for multi-object tracking that breaks away from the conventional detection-association framework. In contrast to previous approaches that treat segmentation as auxiliary information, SAM2MOT places it at the heart of the tracking process, systematically tackling challenges like false positives and occlusions. Its effectiveness has been thoroughly validated on major MOT benchmarks. Furthermore, SAM2MOT integrates pre-trained detector, pre-trained segmentor with tracking logic into a zero-shot MOT system that requires no fine-tuning. This significantly reduces dependence on labeled data and paves the way for transitioning MOT research from task-specific solutions to general-purpose systems. Experiments on DanceTrack, UAVDT, and BDD100K show state-of-the-art results. Notably, SAM2MOT outperforms existing methods on DanceTrack by +2.1 HOTA and +4.5 IDF1, highlighting its effectiveness in MOT.
Manqi Zhao, Dongsheng Jiang
AAAI5
2025 ProReflow: Progressive Reflow with Decomposed Velocity
abstract
Diffusion models have achieved significant progress in both image and video generation while still suffering from huge computation costs. As an effective solution, rectified flow aims to rectify the diffusion process of diffusion models into a straight line for few-step and even one-step generation. However, in this paper, we suggest that the original training pipeline of reflow is not optimal and introduce two techniques to improve it. Firstly, we introduce progressive reflow, which progressively reflows the diffusion models in local timesteps until the whole diffusion progresses, reducing the difficulty of flow matching. Second, we introduce aligned v-prediction, which highlights the importance of direction matching in flow matching over magnitude matching. Experimental results on SDv1.5 and SDXL demonstrate the effectiveness of our method, for example, conducting on SDv1.5 achieves an FID of 10.70 on MSCOCO2014 validation set with only 4 sampling steps, close to our teacher model (32 DDIM steps, FID = 10.05). Our codes will be released at Github.
Lei Ke, Haohang Xu, Xuefei Ning, Yu Li 0022, Haoling Li, Dongsheng Jiang, Yujiu Yang 0001, Linfeng Zhang 0001
CVPR8
2025 AeroGen: Enhancing Remote Sensing Object Detection with Diffusion-Driven Data Generation
abstract
Remote sensing image object detection (RSIOD) aims to identify and locate specific objects within satellite or aerial imagery. However, there is a scarcity of labeled data in current RSIOD datasets, which significantly limits the performance of current detection algorithms. Although existing techniques, e.g., data augmentation and semi-supervised learning, can mitigate this scarcity issue to some extent, they are heavily dependent on high-quality labeled data and perform worse in rare object classes. To address this issue, this paper proposes a layout-controllable diffusion generative model (i.e. AeroGen) tailored for RSIOD. To our knowledge, AeroGen is the first model to simultaneously support horizontal and rotated bounding box condition generation, thus enabling the generation of high-quality synthetic images that meet specific layout and object category requirements. Additionally, we propose an end-to-end data augmentation framework that integrates a diversity-conditioned generator and a filtering mechanism to enhance both the diversity and quality of generated data. Experimental results demonstrate that the synthetic data produced by our method are of high quality and diversity. Furthermore, the synthetic RSIOD data can significantly improve the detection performance of existing RSIOD models, i.e., the mAP metrics on DIOR, DIOR-R, and HRSC datasets are improved by 3.7%, 4.3%, and 2.43%, respectively. The code is available at here.
Datao Tang, Xiangyong Cao, Jing Yao 0002, Xueru Bai, Dongsheng Jiang, Deyu Meng
CVPR7
2025 FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text
Bingchao Wang, Zhiwei Ning, Jianyu Ding, Xuanang Gao, Dongsheng Jiang
ICCV6
2025 Computation and Memory-Efficient Model Compression with Gradient Reweighting
abstract
Pruning is a commonly employed technique for deep neural networks (DNNs) aiming at compressing the model size to reduce computational and memory costs during inference. In contrast to conventional neural networks, large language models (LLMs) pose a unique challenge regarding pruning efficiency due to their substantial computational and memory demands. Existing methods, particularly optimization-based ones, often require considerable computational resources in gradient estimation because they cannot effectively leverage weight sparsity of the intermediate pruned network to lower compuation and memory costs in each iteration. The fundamental challenge lies in the need to frequently instantiate intermediate pruned sub-models to achieve these savings, a task that becomes infeasible even for moderately sized neural networks. To this end, this paper proposes a novel pruning method for DNNs that is both computationally and memory-efficient. Our key idea is to develop an effective reweighting mechanism that enables us to estimate the gradient of the pruned network in current iteration via reweigting the gradient estimated on an outdated intermediate sub-model instantiated at an earlier stage, thereby significantly reducing model instantiation frequency. We further develop a series of techniques, e.g., clipping and preconditioning matrix, to reduce the variance of gradient estimation and stabilize the optimization process. We conducted extensive experimental validation across various domains. Our approach achieves 50\% sparsity and a 1.58$\times$ speedup in forward pass on Llama2-7B model with only 6 GB of memory usage, outperforming state-of-the-art methods with respect to both perplexity and zero-shot performance. As a by-product, our method is highly suited for distributed sparse training and can achieve a 2 $\times$ speedup over the dense distributed baselines.
Yuesen Liao, Binrui Wu, Yuquan Zhou, Xupeng Shi, Dongsheng Jiang
NeurIPS6
2025 SaFiRe: Saccade-Fixation Reiteration with Mamba for Referring Image Segmentation
abstract
Referring Image Segmentation (RIS) aims to segment the target object in an image given a natural language expression. While recent methods leverage pre-trained vision backbones and more training corpus to achieve impressive results, they predominantly focus on simple expressions—short, clear noun phrases like “red car” or “left girl”. This simplification often reduces RIS to a key word/concept matching problem, limiting the model’s ability to handle referential ambiguity in expressions. In this work, we identify two challenging real-world scenarios: object-distracting expressions, which involve multiple entities with contextual cues, and category-implicit expressions, where the object class is not explicitly stated. To address the challenges, we propose a novel framework, SaFiRe, which mimics the human two-phase cognitive process—first forming a global understanding, then refining it through detail-oriented inspection. This is naturally supported by Mamba’s scan-then-update property, which aligns with our phased design and enables efficient multi-cycle refinement with linear complexity. We further introduce aRefCOCO, a new benchmark designed to evaluate RIS models under ambiguous referring expressions. Extensive experiments on both standard and proposed datasets demonstrate the superiority of SaFiRe over state-of-the-art baselines. Project page: https://zhenjiemao.github.io/SaFiRe/.
Zhenjie Mao, Yuhuan Yang, Chaofan Ma, Dongsheng Jiang, Jiangchao Yao, Ya Zhang 0002, Yanfeng Wang 0001
NeurIPS4
2025 Contrastive Learning via Variational Information Bottleneck
abstract
Recent advances in self-supervised learning have witnessed great achievements, especially with the introduction of contrastive learning, where the goal is to maximize the mutual information between different augmentations of the same image, i.e., positive pairs. However, such optimization does not necessarily correspond to optimal representation due to noisy samples, thus inevitably being over-confident in the relevance between views. As a result, the learned model would capture spurious correlation and retain superfluous information that deteriorates representations. In this paper, we facilitate contrastive learning by reducing superfluous relevance between positive views. To this end, we introduce the representation entropy minimization regularization over the objective of vanilla contrastive learning, which forces representations to retain possibly the least information, thus alleviating superfluous relevance from irrelevant views. Then, we derive the analytical expression of the proposed objective by converting it to an information bottleneck problem and solving via variation approximation, which leads to a novel contrastive learning framework, termed as CLIMB, short for Contrastive Learning via variational InforMation Bottleneck. Experiments over multiple benchmarks demonstrate that CLIMB brings consistent improvement. Notably, using DINO as an instantiation, CLIMB achieves 4.5% and 3.5% gain under the k-NN classification metric with EfficientNet-B0 and ResNet-50 as backbones, respectively.
Jin Li 0057, Xiaopeng Zhang 0008, Dongsheng Jiang, Wenrui Dai, Hongkai Xiong, Qi Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Improving Image Restoration Through Removing Degradations in Textual Representations
abstract
In this paper, we introduce a new perspective for improving image restoration by removing degradation in the textual representations of a given degraded image. Intuitively, restoration is much easier on text modality than image one. For example, it can be easily conducted by removing degradation-related words while keeping the content-aware words. Hence, we combine the advantages of images in detail description and ones of text in degradation removal to perform restoration. To address the cross-modal assistance, we propose to map the degraded images into textual representations for removing the degradations, and then convert the restored textual representations into a guidance image for assisting image restoration. In particular, We ingeniously embed an image-to-text mapper and text restoration module into CLIP-equipped text-to-image models to generate the guidance. Then, we adopt a simple coarse-to-fine approach to dynamically inject multi-scale information from guidance to image restoration networks. Extensive experiments are conducted on various image restoration tasks, including deblurring, dehazing, deraining, and denoising, and all-in-one image restoration. The results showcase that our method outperforms state-of-the-art ones across all these tasks. The codes and models are available at https://github.com/mrluin/TextualDegRemoval.
Jingbo Lin, Zhilu Zhang 0001, Yuxiang Wei 0001, Dongwei Ren, Dongsheng Jiang, Qi Tian 0001, Wangmeng Zuo
CVPR5
2024 Domain-Adaptive Semantic Segmentation Emerges From Vision-Language Supervised Domain-Debiased Self-Training
abstract
Unsupervised domain adaptive semantic segmentation leverages synthetic data to train a segmentation model and transfers it to unlabeled real images. Due to the style difference, the transferred model suffers from the domain gap. Even worse, some classes exhibit the extreme domain gap, where the feature distributions undergo a complete shift between the two domains. To alleviate it, we propose a domain-debiased self-training strategy with CLIP to distill its domain-agnostic knowledge. Specifically, we enforce the consistency between the feature maps from our segmentation model and the image encoder of CLIP. Meanwhile, the text embeddings from the text encoder for each class serve as a domain-agnostic classifier to support a domain-debiased feature learning condition. Experimental results under standard UDA settings demonstrate that our proposed strategy consistently improves the UDA segmentation performance based on different backbones and with different large pre-trained models.
Huayu Wang, Zekun Jiang, Lingxi Xie, Dongsheng Jiang, Wei Shen 0002, Qi Tian 0001
ICASSP4
2024 ControlVideo: Training-free Controllable Text-to-video Generation
abstract
Text-driven diffusion models have unlocked unprecedented abilities in image generation, whereas their video counterpart lags behind due to the excessive training cost. To avert the training burden, we propose a training-free ControlVideo to produce high-quality videos based on the provided text prompts and motion sequences. Specifically, ControlVideo adapts a pre-trained text-to-image model (i.e., ControlNet) for controllable text-to-video generation. To generate continuous videos without flicker effect, we propose an interleaved-frame smoother to smooth the intermediate frames. In particular, interleaved-frame smoother splits the whole videos with successive three-frame clips, and stabilizes each clip by updating the middle frame with the interpolation among other two frames in latent space. Furthermore, a fully cross-frame interaction mechanism have been exploited to further enhance the frame consistency, while a hierarchical sampler is employed to produce long videos efficiently. Extensive experiments demonstrate that our ControlVideo outperforms the state-of-the-arts both quantitatively and qualitatively. It is worthy noting that, thanks to the efficient designs, ControlVideo could generate both short and long videos within several minutes using one NVIDIA 2080Ti. Code and videos are available at [this link](https://github.com/YBYBZhang/ControlVideo).
Yabo Zhang, Yuxiang Wei 0001, Dongsheng Jiang, Xiaopeng Zhang 0008, Wangmeng Zuo, Qi Tian 0001
ICLR3
2024 Adaptive B-spline curve fitting with minimal control points using an improved sparrow search algorithm for geometric modeling of aero-engine blades
Suihao Lu, Dongsheng Jiang
Multim. Syst.4
2024 Consensus Synergizes With Memory: A Simple Approach for Anomaly Segmentation in Urban Scenes
abstract
Anomaly segmentation is a critical task for safety-critical applications, such as autonomous driving in urban environments. Its objective is to detect out-of-distribution (OOD) samples with unseen categories, given a pre-trained segmentation model. The core challenge of this task is how to distinguish hard in-distribution samples from OOD samples, which has not been explicitly discussed in previous research. In this paper, we propose a simple yet effective approach named CosMe (Consensus Synergizes with Memory) to address this challenge. CosMe consists of two key components: 1) building a memory bank comprising seen prototypes extracted from multiple layers of the given segmentation model, and 2) training an auxiliary model that mimics the behavior of the given model and using the consensus of their mid-level features as complementary cues that synergize with the memory bank. The former serves as a baseline that can detect all potential outliers, including both OOD and hard in-distribution samples; the latter assists in distinguishing between these two types of outliers. Experimental results on several urban scene anomaly segmentation datasets demonstrate that CosMe outperforms previous approaches by a significant margin.
Jiazhong Cen, Zekun Jiang, Lingxi Xie, Dongsheng Jiang, Wei Shen 0002, Qi Tian 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 SDPT: Semantic-Aware Dimension-Pooling Transformer for Image Segmentation
abstract
Image segmentation plays a critical role in autonomous driving by providing vehicles with a detailed and accurate understanding of their surroundings. Transformers have recently shown encouraging results in image segmentation. However, transformer-based models are challenging to strike a better balance between performance and efficiency. The computational complexity of the transformer-based models is quadratic with the number of inputs, which severely hinders their application in dense prediction tasks. In this paper, we present the semantic-aware dimension-pooling transformer (SDPT) to mitigate the conflict between accuracy and efficiency. The proposed model comprises an efficient transformer encoder for generating hierarchical features and a semantic-balanced decoder for predicting semantic masks. In the encoder, a dimension-pooling mechanism is used in the multi-head self-attention (MHSA) to reduce the computational cost, and a parallel depth-wise convolution is used to capture local semantics. Simultaneously, we further apply this dimension-pooling attention (DPA) to the decoder as a refinement module to integrate multi-level features. With such a simple yet powerful encoder-decoder framework, we empirically demonstrate that the proposed SDPT achieves excellent performance and efficiency on various popular benchmarks, including ADE20K, Cityscapes, and COCO-Stuff. For example, our SDPT achieves 48.6$\%$mIOU on the ADE20K dataset, which outperforms the current methods with fewer computational costs. The codes can be found at https://github.com/HuCaoFighting/SDPT.
Hu Cao, Guang Chen 0001, Hengshuang Zhao, Dongsheng Jiang, Xiaopeng Zhang 0008, Qi Tian 0001, Alois C. Knoll
IEEE Trans. Intell. Transp. Syst.4
2023 USAGE: A Unified Seed Area Generation Paradigm for Weakly Supervised Semantic Segmentation
abstract
Seed area generation is usually the starting point of weakly supervised semantic segmentation (WSSS). Computing the Class Activation Map (CAM) from a multi-label classification network is the de facto paradigm for seed area generation, but CAMs generated from Convolutional Neural Networks (CNNs) and Transformers are prone to be under- and over-activated, respectively, which makes the strategies to refine CAMs for CNNs usually inappropriate for Transformers, and vice versa. In this paper, we propose a Unified optimization paradigm for Seed Area GEneration (USAGE) for both types of networks, in which the objective function to be optimized consists of two terms: One is a generation loss, which controls the shape of seed areas by a temperature parameter following a deterministic principle for different types of networks; The other is a regularization loss, which ensures the consistency between the seed areas that are generated by self-adaptive network adjustment from different views, to overturn false activation in seed areas. Experimental results show that USAGE consistently improves seed area generation for both CNNs and Transformers by large margins, e.g., outperforming state-of-the-art methods by a mIoU of 4.1% on PASCAL VOC. Moreover, based on the USAGE-generated seed areas on Transformers, we achieve state-of-the-art WSSS results on both PASCAL VOC and MS COCO.
Zelin Peng, Guanchun Wang, Lingxi Xie, Dongsheng Jiang, Wei Shen 0002, Qi Tian 0001
ICCV4
2023 Progressively Compressed Auto-Encoder for Self-supervised Representation Learning
Jin Li 0057, Xiaopeng Zhang 0008, Yabo Chen, Dongsheng Jiang, Wenrui Dai, Hongkai Xiong, Qi Tian 0001
ICLR5
2023 Segment Anything in 3D with NeRFs
abstract
Recently, the Segment Anything Model (SAM) emerged as a powerful vision foundation model which is capable to segment anything in 2D images. This paper aims to generalize SAM to segment 3D objects. Rather than replicating the data acquisition and annotation procedure which is costly in 3D, we design an efficient solution, leveraging the Neural Radiance Field (NeRF) as a cheap and off-the-shelf prior that connects multi-view 2D images to the 3D space. We refer to the proposed solution as SA3D, for Segment Anything in 3D. It is only required to provide a manual segmentation prompt (e.g., rough points) for the target object in a single view, which is used to generate its 2D mask in this view with SAM. Next, SA3D alternately performs mask inverse rendering and cross-view self-prompting across various views to iteratively complete the 3D mask of the target object constructed with voxel grids. The former projects the 2D mask obtained by SAM in the current view onto 3D mask with guidance of the density distribution learned by the NeRF; The latter extracts reliable prompts automatically as the input to SAM from the NeRF-rendered 2D mask in another view. We show in experiments that SA3D adapts to various scenes and achieves 3D segmentation within minutes. Our research offers a generic and efficient methodology to lift a 2D vision foundation model to 3D, as long as the 2D model can steadily address promptable segmentation across multiple views.
Jiazhong Cen, Zanwei Zhou, Jiemin Fang, Chen Yang 0023, Wei Shen 0002, Lingxi Xie, Dongsheng Jiang, Xiaopeng Zhang 0008, Qi Tian 0001
NeurIPS7
2023 AiluRus: A Scalable ViT Framework for Dense Prediction
abstract
Vision transformers (ViTs) have emerged as a prevalent architecture for vision tasks owing to their impressive performance. However, their complexity dramatically increases when handling long token sequences, particularly for dense prediction tasks that require high-resolution input. Notably, dense prediction tasks, such as semantic segmentation or object detection, emphasize more on the contours or shapes of objects, while the texture inside objects is less informative. Motivated by this observation, we propose to apply adaptive resolution for different regions in the image according to their importance. Specifically, at the intermediate layer of the ViT, we select anchors from the token sequence using the proposed spatial-aware density-based clustering algorithm. Tokens that are adjacent to anchors are merged to form low-resolution regions, while others are preserved independently as high-resolution. This strategy could significantly reduce the number of tokens, and the following layers only handle the reduced token sequence for acceleration. At the output end, the resolution of the feature map is recovered by unfolding merged tokens for task prediction. Consequently, we can considerably accelerate ViTs for dense prediction tasks. The proposed method is evaluated across three different datasets and demonstrates promising performance. For instance, "Segmenter ViT-L" can be accelerated by 48\% FPS without fine-tuning, while maintaining the performance. Moreover, our method can also be applied to accelerate fine-tuning. Experiments indicate that we can save 52\% training time while accelerating 2.46$\times$ FPS with only a 0.09\% performance drop.
Jin Li 0057, Xiaopeng Zhang 0008, Bowen Shi 0003, Dongsheng Jiang, Wenrui Dai, Hongkai Xiong, Qi Tian 0001
NeurIPS5
2023 A Survey on Label-Efficient Deep Image Segmentation: Bridging the Gap Between Weak Supervision and Dense Prediction
abstract
The rapid development of deep learning has made a great progress in image segmentation, one of the fundamental tasks of computer vision. However, the current segmentation algorithms mostly rely on the availability of pixel-level annotations, which are often expensive, tedious, and laborious. To alleviate this burden, the past years have witnessed an increasing attention in building label-efficient, deep-learning-based image segmentation algorithms. This paper offers a comprehensive review on label-efficient image segmentation methods. To this end, we first develop a taxonomy to organize these methods according to the supervision provided by different types of weak labels (including no supervision, inexact supervision, incomplete supervision and inaccurate supervision) and supplemented by the types of segmentation problems (including semantic segmentation, instance segmentation and panoptic segmentation). Next, we summarize the existing label-efficient image segmentation methods from a unified perspective that discusses an important question: how to bridge the gap between weak supervision and dense prediction - the current methods are mostly based on heuristic priors, such as cross-pixel similarity, cross-label constraint, cross-view consistency, and cross-image relation. Finally, we share our opinions about the future research directions for label-efficient deep image segmentation.
Wei Shen 0002, Zelin Peng, Huayu Wang, Jiazhong Cen, Dongsheng Jiang, Lingxi Xie, Xiaokang Yang 0001, Qi Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2022 SdAE: Self-distillated Masked Autoencoder
Yabo Chen, Yuchen Liu 0006, Dongsheng Jiang, Xiaopeng Zhang 0008, Wenrui Dai, Hongkai Xiong, Qi Tian 0001
ECCV (30)3
2022 A Transformer-Based Decoder for Semantic Segmentation with Multi-level Context Mining
Bowen Shi 0003, Dongsheng Jiang, Xiaopeng Zhang 0008, Wenrui Dai, Junni Zou, Hongkai Xiong, Qi Tian 0001
ECCV (28)2