EDBT 2026 Demo / reviewers in the wild / expert
Manlin Zhang
dblp:69/7814
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0003-2319-2749ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DiffusionEngine: Diffusion model is scalable data engine for object detection
Manlin Zhang, Jie Wu 0032, Yuxi Ren, Ming Li 0010, Andy Jinhua Ma |
Pattern Recognit. | 1 |
| 2024 | Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image SynthesisabstractDiffusion model is a promising approach to image generation and has been employed for Pose-Guided Person Image Synthesis (PGPIS) with competitive performance. While existing methods simply align the person appearance to the target pose, they are prone to overfitting due to the lack of a high-level semantic understanding on the source person image. In this paper, we propose a novel Coarse-to-Fine Latent Diffusion (CFLD) method for PGPIS. In the absence of image-caption pairs and textual prompts, we de-velop a novel training paradigm purely based on images to control the generation process of a pre-trained text-to-image diffusion model. A perception-refined decoder is designed to progressively refine a set of learnable queries and extract semantic understanding of person images as a coarse-grained prompt. This allows for the decoupling of fine-grained appearance and pose information controls at different stages, and thus circumventing the potential over-fitting problem. To generate more realistic texture details, a hybrid- granularity attention module is proposed to encode multi-scale fine-grained appearance features as bias terms to augment the coarse-grained prompt. Both quantitative and qualitative experimental results on the DeepFashion benchmark demonstrate the superiority of our method over the state of the arts for PGPIS. Code is available at https://github.com/YanzuoLu/CFLD. Yanzuo Lu, Manlin Zhang, Andy Jinhua Ma, Xiaohua Xie, Jian-Huang Lai |
CVPR | 2 |
| 2024 | TreeReward: Improve Diffusion Model via Tree-Structured Feedback LearningabstractRecently, there has been significant progress in leveraging human feedback to enhance diffusion-based image generation, garnering considerable interest and attention. However, existing methods fail to achieve a fine-grained performance boost for the following challenges: i) insufficient amount of fine-grained feedback data; ii) lack of effective fine-grained feedback learning framework; To tackle these challenges, we present TreeReward to facilitate the fine-grained feedback optimization for diffusion models. Specifically, to address the limitation of the fine-grained feedback data, we first design a novel "AI + Expert" feedback data construction pipeline, yielding about 2.2M high-quality feedback dataset encompassing six fine-grained dimensions at a relatively low cost. Built upon this dataset, we introduce a tree-structure reward model to exploit the fine-grained feedback data efficiently and provide tailored optimization during feedback learning. We validate the feedback learning performance of our method across different fine-grained dimensions and various downstream tasks. Extensive experiments on both Stable Diffusion v1.5 (SD1.5) and Stable Diffusion XL (SDXL) demonstrate the effectiveness of our method in enhancing the general and fine-grained generation and downstream tasks generalization. Jie Wu 0030, Huafeng Kuang, Haiming Zhang 0001, Yuxi Ren, Manlin Zhang, Xuefeng Xiao 0001, Guanbin Li |
ACM Multimedia | 7 |
| 2023 | UGC: Unified GAN Compression for Efficient Image-to-Image TranslationabstractRecent years have witnessed the prevailing progress of Generative Adversarial Networks (GANs) in image-to-image translation. However, the success of these GAN models hinges on ponderous computational costs and labor-expensive training data. Current efficient GAN learning techniques often fall into two orthogonal aspects: i) model slimming via reduced calculation costs; ii) data/label-efficient learning with fewer training data/labels. To combine the best of both worlds, we propose a new learning paradigm, Unified GAN Compression (UGC), with a unified optimization objective to seamlessly prompt the synergy of model-efficient and label-efficient learning. UGC sets up semi-supervised-driven network architecture search and adaptive online semi-supervised distillation stages sequentially, which formulates a heterogeneous mutual learning scheme to obtain an architecture-flexible, label-efficient, and performance-excellent model. Extensive experiments demonstrate that UGC obtains state-of-the-art lightweight models even with less than 50% labels. UGC that compresses 40× MACs can achieve 21.43 FID on edges→shoes with 25% labels, which even outperforms the original model with 100% labels by 2.75 FID. Yuxi Ren, Jie Wu 0032, Manlin Zhang, Xuefeng Xiao 0001, Rui Wang 0089 |
ICCV | 4 |
| 2023 | Patch Shuffle and Pixel Contrast: Dual Consistency Learning for Semi-supervised Lung Tumor Segmentation
Chenyu Cai, Manlin Zhang, Yanxu Hu, Andy Jinhua Ma |
PRCV (5) | 3 |
| 2022 | Suppressing Static Visual Cues via Normalizing Flows for Self-Supervised Video Representation LearningabstractDespite the great progress in video understanding made by deep convolutional neural networks, feature representation learned by existing methods may be biased to static visual cues. To address this issue, we propose a novel method to suppress static visual cues (SSVC) based on probabilistic analysis for self-supervised video representation learning. In our method, video frames are first encoded to obtain latent variables under standard normal distribution via normalizing flows. By modelling static factors in a video as a random variable, the conditional distribution of each latent variable becomes shifted and scaled normal. Then, the less-varying latent variables along time are selected as static cues and suppressed to generate motion-preserved videos. Finally, positive pairs are constructed by motion-preserved videos for contrastive learning to alleviate the problem of representation bias to static cues. The less-biased video representation can be better generalized to various downstream tasks. Extensive experiments on publicly available benchmarks demonstrate that the proposed method outperforms the state of the art when only single RGB modality is used for pre-training. Manlin Zhang, Andy Jinhua Ma |
AAAI | 1 |
| 2022 | Improving Pre-trained Masked Autoencoder via Locality Enhancement for Person Re-identification
Yanzuo Lu, Manlin Zhang, Andy Jinhua Ma, Xiaohua Xie, Jian-Huang Lai |
PRCV (2) | 2 |
| 2022 | Multi-Level Temporal Dilated Dense Prediction for Action Recognitionabstract3D convolutional neural networks have achieved great success for action recognition. However, large variations of temporal dynamics have not been properly processed and low-level features have not been fully exploited in most existing works. To solve these two problems, we present a general and flexible framework, namely multi-level temporal dilated dense prediction network, which can incorporate with most of existing methods as backbone to improve the temporal modeling capacity. In the proposed method, a novel temporal dilated dense prediction block is designed to fully utilize temporal features with various temporal dilated rates for dense prediction while maintaining relatively low computational cost. To fuse information from low to high levels, our method combines the predictions from multiple such blocks inserted at different stages of the backbone network. In-depth analysis is given to show that short- to long-term temporal dependencies can be captured and multi-level spatio-temporal features are effectively fused for video action recognition by the proposed method. Experimental results demonstrate that our method achieves impressive performance improvement on four publicly available action recognition benchmarks including Charades, Kinetics, Something-Something-V1 and HMDB51. Manlin Zhang, Andy Jinhua Ma |
IEEE Trans. Multim. | 3 |
| 2021 | Learning Spatio-temporal Representation by Channel Aliasing Video PerceptionabstractIn this paper, we propose a novel pretext task namely Channel Aliasing Video Perception (CAVP) for self-supervised video representation learning. The main idea of our approach is to generate channel aliasing videos, which carry different motion cues simultaneously by assembling distinct channels from different videos. With the generated channel aliasing videos, we propose to recognize the number of different motion flows within a channel aliasing video for perception of discriminative motion cues. As a plug-and-play method, the proposed pretext task can be integrated into a co-training framework with other self-supervised learning methods to further improve the performance. Experimental results on publicly available action recognition benchmarks verify the effectiveness of our method for spatio-temporal representation learning. Manlin Zhang, Andy Jinhua Ma |
ACM Multimedia | 3 |