EDBT 2026 Demo / reviewers in the wild / expert
Chen-Lin Zhang
dblp:177/9120
· DBLP profile ↗
18ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0002-3168-1852ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 8 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Single Image Rolling Shutter Removal with Diffusion ModelsabstractWe present RS-Diffusion, the first Diffusion Models-based method for single-frame Rolling Shutter (RS) correction. RS artifacts compromise visual quality of frames due to the row-wise exposure of CMOS sensors. Most previous methods have focused on multi-frame approaches, using temporal information from consecutive frames for the motion rectification. However, few approaches address the more challenging but important single frame RS correction. In this work, we present an ``image-to-motion" framework via diffusion techniques, with a designed patch-attention module. In addition, we present the RS-Real dataset, comprised of captured RS frames alongside their corresponding Global Shutter (GS) ground-truth pairs. The GS frames are corrected from the RS ones, guided by the corresponding Inertial Measurement Unit (IMU) gyroscope data acquired during capture. Experiments show that RS-Diffusion surpasses previous single-frame RS methods, demonstrates the potential of diffusion-based approaches, and provides a valuable dataset for further research. Zhanglei Yang, Haipeng Li 0001, Mingbo Hong, Chen-Lin Zhang, Shuaicheng Liu |
AAAI | 4 |
| 2024 | An Asymmetric Augmented Self-Supervised Learning Method for Unsupervised Fine-Grained Image HashingabstractUnsupervised fine-grained image hashing aims to learn compact binary hash codes in unsupervised settings, addressing challenges posed by large-scale datasets and dependence on supervision. In this paper, we first identify a granularity gap between generic and fine-grained datasets for unsupervised hashing methods, highlighting the inadequacy of conventional self-supervised learning for fine-grained visual objects. To bridge this gap, we propose the Asymmetric Augmented Self-Supervised Learning (A2-SSL) method, comprising three modules. The asymmetric augmented SSL module employs suitable augmentation strategies for positive/negative views, preventing fine-grained category confusion inherent in conventional SSL. Part-oriented dense contrastive learning utilizes the Fisher Vector framework to capture and model fine- grained object parts, enhancing unsupervised representations through part-level dense contrastive learning. Self-consistent hash code learning introduces a reconstruction task aligned with the self-consistency principle, guiding the model to emphasize comprehensive features, particularly fine-grained patterns. Experimental results on five benchmark datasets demonstrate the superiority of A2-SSL over existing methods, affirming its efficacy in unsupervised fine-grained image hashing. Feiran Hu, Chen-Lin Zhang, Jiangliang Guo, Xiu-Shen Wei, Lin Zhao 0003, Lingyan Gao |
CVPR | 2 |
| 2024 | End-to-End Temporal Action Detection with 1B Parameters Across 1000 FramesabstractRecently, temporal action detection (TAD) has seen significant performance improvement with end-to-end training. However, due to the memory bottleneck, only models with limited scales and limited data volumes can afford end-to-end training, which inevitably restricts TAD performance. In this paper, we reduce the memory consumption for end-to-end training, and manage to scale up the TAD backbone to 1 billion parameters and the input video to 1,536 frames, leading to significant detection performance. The key to our approach lies in our proposed temporal-informative adapter (TIA), which is a novel lightweight module that reduces training memory. Using TIA, we free the humongous backbone from learning to adapt to the TAD task by only updating the parameters in TIA. TIA also leads to better TAD representation by temporally aggregating context from adjacent frames throughout the backbone. We evaluate our model across four representative datasets. Owing to our efficient design, we are able to train end-to-end on VideoMAEv2-giant and achieve 75.4% mAP on THUMOS14, being the first end-to-end model to outper-form the best feature-based methods. Code is available at https://github.com/sming256/AdaTAD. Shuming Liu 0001, Chen-Lin Zhang, Chen Zhao 0002, Bernard Ghanem |
CVPR | 2 |
| 2024 | RecDiffusion: Rectangling for Image Stitching with Diffusion ModelsabstractImage stitching from different captures often results in non-rectangular boundaries, which is often considered un-appealing. To solve non-rectangular boundaries, current solutions involve cropping, which discards image content, inpainting, which can introduce unrelated content, or warping, which can distort non-linear features and introduce artifacts. To overcome these issues, we introduce a novel diffusion-based learning framework, RecDiffusion, for image stitching rectangling. This framework combines Motion Diffusion Models (MDM) to generate motion fields, ef-fectively transitioning from the stitched image's irregular borders to a geometrically corrected intermediary. Fol-lowed by Content Diffusion Models (CDM) for image de-tail refinement. Notably, our sampling process utilizes a weighted map to identify regions needing correction during each iteration of CDM. Our RecDiffusion ensures geomet-ric accuracy and overall visual appeal, surpassing all pre-vious methods in both quantitative and qualitative measures when evaluated on public benchmarks. Code is released at https://github.com/haippp/RecDiffusion. Tianhao Zhou, Haipeng Li 0001, Ao Luo, Chen-Lin Zhang, Bing Zeng 0001, Shuaicheng Liu |
CVPR | 5 |
| 2024 | A Unified Image Compression Method for Human Perception and Multiple Vision Tasks
Sha Guo, Lin Sui, Chen-Lin Zhang, Zhuo Chen 0006, Wenhan Yang, Ling-Yu Duan |
ECCV (71) | 3 |
| 2023 | A Simple and Efficient Pipeline to Build an End-to-End Spatial-Temporal Action DetectorabstractSpatial-temporal action detection is a vital part of video understanding. Current spatial-temporal action detection methods mostly use an object detector to obtain person candidates and classify these person candidates into different action categories. So-called two-stage methods are heavy and hard to apply in real-world applications. Some existing methods build one-stage pipelines, But a large performance drop exists with the vanilla one-stage pipeline and extra classification modules are needed to achieve comparable performance. In this paper, we explore a simple and effective pipeline to build a strong one-stage spatial-temporal action detector. The pipeline is composed by two parts: one is a simple end-to-end spatial-temporal action detector. The proposed end-to-end detector has minor architecture changes to current proposal-based detectors and does not add extra action classification modules. The other part is a novel labeling strategy to utilize unlabeled frames in sparse annotated data. We named our model as SE-STAD. The proposed SE-STAD achieves around 2% mAP boost and around 80% FLOPs reduction. Our code will be released at https://github.com/4paradigm-CV/SE-STAD Lin Sui, Chen-Lin Zhang, Lixin Gu |
WACV | 2 |
| 2023 | Salvage of Supervision in Weakly Supervised Object Detection and SegmentationabstractWeakly supervised vision tasks, including detection and segmentation, have attracted much attention in the vision community recently. However, the lack of detailed and precise annotations in the weakly supervised case leads to a large accuracy gap between weakly- and fully-supervised methods. In this article, we propose a new framework, Salvage of Supervision (SoS), with the key idea being to effectively harness every potentially useful supervisory signal in weakly supervised vision tasks. Starting with weakly supervised object detection (WSOD), we propose SoS-WSOD to shrink the technology gap between WSOD and FSOD, which utilizes the weak image-level labels, the pseudo-labels, and the power of semi-supervised object detection for WSOD. Moreover, SoS-WSOD removes restrictions in traditional WSOD methods, including the reliance on ImageNet pretraining and inability to use modern backbones. The SoS framework also extends to weakly supervised semantic segmentation and instance segmentation. On several weakly supervised vision benchmarks, SoS achieves significant performance boost and generalization ability. Lin Sui, Chen-Lin Zhang, Jianxin Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Weakly supervised foreground learning for weakly supervised localization and detection
Chen-Lin Zhang, Yin Li 0003, Jianxin Wu 0001 |
Pattern Recognit. | 1 |
| 2022 | Salvage of Supervision in Weakly Supervised Object DetectionabstractWeakly supervised object detection (WSOD) has recently attracted much attention. However, the lack of bounding-box supervision makes its accuracy much lower than fully supervised object detection (FSOD), and currently modern FSOD techniques cannot be applied to WSOD. To bridge the performance and technical gaps between WSOD and FSOD, this paper proposes a new framework, Salvage of Supervision (SoS), with the key idea being to harness every potentially useful supervisory signal in WSOD: the weak image-level labels, the pseudo-labels, and the power of semi-supervised object detection. This paper proposes new approaches to utilize these weak and noisy signals effectively, and shows that each type of supervisory signal brings in notable improvements, outperforms existing WSOD methods (which mainly use only the weak labels) by large margins. The proposed SoS- WSOD method also has the ability to freely use modern FSOD techniques. SoS-WSOD achieves 64.4 mAP50on VOC2007, 61.9 mAP50on VOC2012 and 16.6 mAP50:95on MS-COCO, and also has fast inference speed. Ablations and visualization further verify the effectiveness of SoS. Lin Sui, Chen-Lin Zhang, Jianxin Wu 0001 |
CVPR | 2 |
| 2022 | ActionFormer: Localizing Moments of Actions with Transformers
Chen-Lin Zhang, Jianxin Wu 0001, Yin Li 0003 |
ECCV (4) | 1 |
| 2020 | Rethinking the Route Towards Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) aims to localize objects with only image-level labels. Previous methods often try to utilize feature maps and classification weights to localize objects using image level annotations indirectly. In this paper, we demonstrate that weakly supervised object localization should be divided into two parts: class-agnostic object localization and object classification. For class-agnostic object localization, we should use class-agnostic methods to generate noisy pseudo annotations and then perform bounding box regression on them without class labels. We propose the pseudo supervised object localization (PSOL) method as a new way to solve WSOL. Our PSOL models have good transferability across different datasets without fine-tuning. With generated pseudo bounding boxes, we achieve 58.00% localization accuracy on ImageNet and 74.74% localization accuracy on CUB-200, which have a large edge over previous models. Chen-Lin Zhang, Yun-Hao Cao, Jianxin Wu 0001 |
CVPR | 1 |
| 2020 | ApproxDet: content and contention-aware approximate object detection for mobilesabstractAdvanced video analytic systems, including scene classification and object detection, have seen widespread success in various domains such as smart cities and autonomous systems. With an evolution of heterogeneous client devices, there is incentive to move these heavy video analytics workloads from the cloud to mobile devices for low latency and real-time processing and to preserve user privacy. However, most video analytic systems are heavyweight and are trained offline with some pre-defined latency or accuracy requirements. This makes them unable to adapt at runtime in the face of three types of dynamism --- the input video characteristics change, the amount of compute resources available on the node changes due to co-located applications, and the user's latency-accuracy requirements change. In this paper we introduce ApproxDet, an adaptive video object detection framework for mobile devices to meet accuracy-latency requirements in the face of changing content and resource contention scenarios. To achieve this, we introduce a multi-branch object detection kernel, which incorporates a data-driven modeling approach on the performance metrics, and a latency SLA-driven scheduler to pick the best execution branch at runtime. We evaluate ApproxDet on a large benchmark video dataset and compare quantitatively to AdaScale and YOLOv3. We find that ApproxDet is able to adapt to a wide variety of contention and content characteristics and outshines all baselines, e.g., it achieves 52% lower latency and 11.1% higher accuracy over YOLOv3. Our software is open-sourced at https://github.com/purdue-dcsl/ApproxDet. Ran Xu 0003, Chen-Lin Zhang, Pengcheng Wang 0001, Jayoung Lee, Subrata Mitra, Somali Chaterji, Yin Li 0003, Saurabh Bagchi |
SenSys | 2 |
| 2019 | Noise-Aware Network Embedding for Multiplex NetworkabstractNetwork embedding aims at learning the latent representations of nodes while preserving the complex structure of the underlying graph. Real-world networks are usually related with each other via common nodes, the so-called multiplex network. To make the data mining work on the multiplex network more actionable, it become urgent and essential to transform it into low-dimension vector space. Recently, several works have been proposed to leverage the complementary information for embedding. However, they suffer from sacrificing distinct properties of the counterparts in different layers, as they preserve much noise information into embedding vectors. In this paper, we propose a Noise-Aware Network Embedding approach for Multiplex Network, namely NANE. Unlike previous works, NANE considers the roles of an identical node in different layers, and adopts a more robust and flexible strategy to rationally integrate the cross-layer information while keeping the unique characteristic of each layer. We perform extensive evaluations on several real-world datasets. The experimental results demonstrate that our NANE can achieve better performance on link prediction task and significantly outperform previous methods especially in noisy multiplex network scenarios. Xiaokai Chu, Xinxin Fan, Di Yao 0001, Chen-Lin Zhang, Jingping Bi |
IJCNN | 4 |
| 2019 | Unsupervised object discovery and co-localization by deep descriptor transformation
Xiu-Shen Wei, Chen-Lin Zhang, Jianxin Wu 0001, Chunhua Shen, Zhi-Hua Zhou |
Pattern Recognit. | 2 |
| 2019 | Improving CNN linear layers with power mean non-linearity
Chen-Lin Zhang, Jianxin Wu 0001 |
Pattern Recognit. | 1 |
| 2018 | Coarse-to-Fine: A RNN-Based Hierarchical Attention Model for Vehicle Re-identification
Xiu-Shen Wei, Chen-Lin Zhang, Lingqiao Liu, Chunhua Shen, Jianxin Wu 0001 |
ACCV (2) | 2 |
| 2018 | Deep Bimodal Regression of Apparent Personality Traits from Short Video SequencesabstractApparent personality analysis (APA) is an important problem of personality computing, and furthermore, automatic APA becomes a hot and challenging topic in computer vision and multimedia. In this paper, we propose a deep learning solution to APA from short video sequences. In order to capture rich information from both the visual and audio modality of videos, we tackle these tasks with our Deep Bimodal Regression (DBR) framework. In DBR, for the visual modality, we modify the traditional convolutional neural networks for exploiting important visual cues. In addition, taking into account the model efficiency, we extract audio representations and build a linear regressor for the audio modality. For combining the complementary information from the two modalities, we ensemble these predicted regression scores by both early fusion and late fusion. Finally, based on the proposed framework, we come up with a solution for the Apparent Personality Analysis competition track in the ChaLearn Looking at People challenge in association with ECCV 2016. Our DBR is the winner (first place) of this challenge with 86 registered participants. Beyond the competition, we further investigate the performance of different loss functions in our visual models, and prove non-convex loss functions for regression are optimal on the human-labeled video data. Xiu-Shen Wei, Chen-Lin Zhang, Hao Zhang 0038, Jianxin Wu 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2017 | Deep Descriptor Transforming for Image Co-LocalizationabstractReusable model design becomes desirable with the rapid expansion of machine learning applications. In this paper, we focus on the reusability of pre-trained deep convolutional models. Specifically, different from treating pre-trained models as feature extractors, we reveal more treasures beneath convolutional layers, i.e., the convolutional activations could act as a detector for the common object in the image co-localization problem. We propose a simple but effective method, named Deep Descriptor Transforming (DDT), for evaluating the correlations of descriptors and then obtaining the category-consistent regions, which can accurately locate the common object in a set of images. Empirical studies validate the effectiveness of the proposed DDT method. On benchmark image co-localization datasets, DDT consistently outperforms existing state-of-the-art methods by a large margin. Moreover, DDT also demonstrates good generalization ability for unseen categories and robustness for dealing with noisy data. Xiu-Shen Wei, Chen-Lin Zhang, Yao Li 0003, Chen-Wei Xie, Jianxin Wu 0001, Chunhua Shen, Zhi-Hua Zhou |
IJCAI | 2 |