VLDB 2026 Research / reviewers in the wild / expert
Guanchen Ding
dblp:283/8505
· DBLP profile ↗
9ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0002-9523-1850ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Efficient Neural Rate Control for JPEG-AIabstractRate control (RC) is a critical component in learned image compression (LIC), particularly in the emerging JPEG-AI standard, which enables adaptive bitrate achievement to meet diverse bandwidth constraints. JPEG-AI default RC employs an iterative optimization process, wherein a pre-trained RC model is selected and the (generated) latent representations are adjusted based on the mismatch between actual and target bitrates. Despite satisfactory results, such a trial-and-error paradigm necessitates multiple processing cycles, resulting in inevitable computational overhead. We propose an efficient neural rate control framework for JPEG-AI to address this limitation. Our idea is to train a ResNet-based neural control (NRC) to learn the mapping from the input images and target bitrates to the optimal coding parameters. The trained NRC can then be applied to predict the coding parameters based on the new input images and target bitrates directly. Experimental results on DIV2K and MSCOCO datasets show that our NRC achieves comparable rate-distortion performance while reducing encoding time by about 5× compared to JPEG-AI default RC. Guanchen Ding, Zhenzhong Chen 0001, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Counting Beyond Domains: Toward Alignment in Unsupervised Domain Adaptation in Remote Sensing Object Counting
Guanchen Ding, Daiqin Yang, Zhenzhong Chen 0001, Chang Wen Chen |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Towards Omniscient Feature Alignment for Video RescalingabstractVideo super-resolution often reconstructs high-resolution (HR) video from low-resolution (LR) video that has been downsampled using predefined methods, which is an ill-posedness problem. Recent video rescaling algorithms alleviate this problem by jointly training the downsampling and upsampling processes. However, they primarily exploit the shallow temporal correlations among video frames, overlooking the intricate, long-term sequential depth dependencies within the video. In this paper, we propose an omniscient feature alignment to leverage the bidirectional deep temporal information for video rescaling, namely OFA-VRN. In the downsampling phase, the proposed method separates the input HR video into LR frames and high-frequency components using haar wavelet transform and explicitly embeds the high-frequency components into the LR frames. In this way, detailed information is stored in the frame and maintains visual perception quality in downsampled videos. During the upsampling phase, we use an advanced bidirectional propagation paradigm to enhance temporal information aggregation capabilities. By incorporating the proposed omniscient feature alignment, the network is capable of leveraging multi-frame feature information from the triplet dimension to further alleviate misalignment issues, thereby enhancing its capacity for deep temporal information utilization. The experiments on Vid4 and Vimeo90K-T demonstrate that our model achieves competitive performance compared to the state-of-the-art methods. Guanchen Ding, Chang Wen Chen |
ICASSP | 1 |
| 2024 | Domain-Agnostic Crowd Counting via Uncertainty-Guided Style Diversity AugmentationabstractDomain shift significantly hinders crowd counting performance in unseen domains. Domain adaptation methods tackle this issue using target domain images but falter when acquiring these images is difficult. Moreover, they demand additional training time for fine-tuning. To address this issue, we propose an Uncertainty-Guided Style Diversity Augmentation (UGSDA) method, enabling the models to be trained solely on the source domain and directly generalized to various target domains. It is achieved by generating sufficiently diverse and realistic samples during the training process. Specifically, our UGSDA method incorporates three tailor-designed components: the Global Styling Elements Extraction (GSEE) module, the Local Uncertainty Perturbations (LUP) module, and the Density Distribution Consistency (DDC) loss. The GSEE extracts global style elements from the feature space of the whole source domain. The LUP aims to obtain uncertainty perturbations from the batch-level input to form style distributions beyond the source domain, which used to generate diversified stylized samples together with global style elements. To regulate the extent of perturbations, the DDC loss imposes constraints between the source samples and the stylized samples, ensuring the stylized samples maintain a higher degree of realism and reliability. Comprehensive experiments validate the superiority of our approach, demonstrating its strong generalization capabilities across various datasets and models. Code is available at https://github.com/gcding/UGSDA-pytorch. Guanchen Ding, Lingbo Liu, Zhenzhong Chen 0001, Chang Wen Chen |
ACM Multimedia | 1 |
| 2024 | DOPNet: Dense Object Prediction Network for Multiclass Object Counting and Localization in Remote Sensing ImagesabstractObject counting and localization for remote sensing images are effective means to solve large-scale object analysis problems. Nowadays, most counting methods obtain the number of objects by employing convolutional neural network (CNN) to regress a density map of objects. Even if these leading methods have achieved impressive performances, they simply focus on estimating the number of single-class objects, without providing location information and cannot support multiclass objects. To tackle these problems, a point-based network named Dense Object Prediction Network (DOPNet) is proposed for multiclass object counting and localization for remote sensing images. DOPNet differs from the conventional approach of predicting multiple density maps by incorporating category attributes into the predicted objects, enabling the accurate counting and localization of multiclass objects. Specifically, DOPNet adopts a multiscale architecture (MS) to provide dense predictions of object proposals. A scale adaptive feature enhancement module (SAFEM) is designed to predict scales of objects for the suppression of duplicate proposals. Given only point level annotations for training, a pseudo-box generation algorithm is designed to find the most suitable pseudo-box of each annotated object for the supervision of scale learning. Comprehensive experiments prove that DOPNet can achieve preferable performance on challenging benchmarks of counting while providing object locations. Code and pre-trained models are available athttps://github.com/Ceoilmp/DOPNet. Mingpeng Cui, Guanchen Ding, Daiqin Yang, Zhenzhong Chen 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Crowd Counting via Unsupervised Cross-Domain Feature AdaptationabstractGiven an image, crowd counting aims to estimate the amount of target objects in the image. With un-predictable installation situations of surveillance systems (or other equipments), crowd counting images from different data sets may exhibit severe discrepancies in viewing angle, scale, lighting condition, etc. As it is usually expensive and time-consuming to annotate each data set for model training, it has been an essential issue in crowd counting to transfer a well-trained model on a labeled data set (source domain) to a new data set (target domain). To tackle this problem, we propose a cross-domain learning network to learn the domain gaps in an unsupervised learning manner. The proposed network comprises of a Multi-granularity Feature-aware Discriminator (MFD) module, a Domain-invariant Feature Adaptation (DFA) module, and a Cross-domain Vanishing Bridge (CVB) module to remove domain-specific information from the extracted features and promote the mapping performances of the network. Unlike most existing methods that use only Global Feature Discriminator (GFD) to align features at image level, an additional Local Feature Discriminator (LFD) is inserted and together with GFD form the MFD module. As a complement to MFD, LFD refines features at pixel level and has the ability to align local features. The DFA module explicitly measures the distances between the source domain features and the target domain features and aligns the marginal distribution of their features with Maximum Mean Discrepancy (MMD). Finally, the CVB module provides an incremental capability of removing the impact of interfering part of the extracted features. Several well-known networks are adopted as the backbone of our algorithm to prove the effectiveness of the proposed adaptation structure. Comprehensive experiments demonstrate that our model achieves competitive performance to the state-of-the-art methods. Guanchen Ding, Daiqin Yang |
IEEE Trans. Multim. | 1 |
| 2022 | The First Challenge on Moving Object Detection and Tracking in Satellite Videos: Methods and ResultsabstractIn this paper, we briefly summarize the first challenge on moving object detection and tracking in satellite videos (SatVideoDT). This challenge has three tracks related to satellite video analysis, including moving object detection (Track 1), single object tracking (Track 2), and multiple-object tracking (Track 3). 123, 89, and 70 participants successfully registered, while 37, 42, and 29 teams submitted their final results on the test datasets for Tracks 1-3, respectively. The top-performing methods and their results in each track are described with details. This challenge establishes a new benchmark for satellite video analysis. Yulan Guo, Qingyong Hu, Feng Zhang 0046, Ye Zhang 0037, Hanyun Wang, Chenguang Dai, Weilong Guo, Xiyu Qi, Kelong Tu, Shudan Zhu, Lai Chen, Bin Lin 0013, Chaocan Xue, Jinlei Zheng, Limei Qin, Ying Li 0017, Manqi Zhao, Lu Ruan 0003, Mingpeng Cui, Guanchen Ding, Guangwei Jiang, Zhenzhong Chen 0001, Kaiyang Cao, Lingyu Kong, Shaodong Chen, Zhicheng Zhao 0001, Qin Shen, Lei Liu 0049, Chenglong Li 0002, Yun Xiao 0003 |
ICPR | 25 |
| 2022 | Object Counting for Remote-Sensing Images via Adaptive Density Map-Assisted LearningabstractObject counting has attracted a lot of attention in remote sensing image analysis. In density map based object counting algorithms, the ground truth density maps generated by fix-sized Gaussian kernels ignore the spatial features of the objects. In this paper, an Adaptive Density Map Assisted Learning algorithm (ADMAL) is proposed, which taps into spatial features of the objects from the beginning phase of ground truth density map generation. ADMAL consists of two networks: a Contexture Aware Density Map Generation (CADMG) network and a Transformer-based Density Map Estimation (TDME) network. The CADMG network is designed to generate a ground truth density map from each annotated point map. Comparing with Gaussian convolved density maps, the ground truth density maps generated by CADMG will be tailored according to the texture and neighborhood relationship among objects, which can promote the learning effect of the TDME network. TDME is the core network for object counting. The backbone of the TDME network adopts a Swin transformer structure, the self-attention mechanism of which possesses a larger receptive field for effective feature extraction in remote sensing images. Comprehensive experiments prove that the ground truth density map generated by CADMG can help various density map estimation networks achieve better training effects, among which TDME achieves the best performances. Moreover, the ADMAL algorithm can achieve preferable object counting performances for both satellite-based image and drone-based image. Code and pre-trained models are available at https://github.com/gcding/ADMAL-pytorch. Guanchen Ding, Mingpeng Cui, Daiqin Yang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Drone-Based Car Counting via Density Map LearningabstractCar counting on drone-based images is a challenging task in computer vision. Most advanced methods for counting are based on density maps. Usually, density maps are first generated by convolving ground truth point maps with a Gaussian kernel for later model learning (generation). Then, the counting network learns to predict density maps from input images (estimation). Most studies focus on the estimation problem while overlooking the generation problem. In this paper, a training framework is proposed to generate density maps by learning and train generation and estimation subnetworks jointly. Experiments demonstrate that our method outperforms other density map-based methods and shows the best performance on drone-based car counting. Jingxian Huang, Guanchen Ding, Yujia Guo, Daiqin Yang |
VCIP | 2 |