VLDB 2026 Research / reviewers in the wild / expert
Lele Cheng
dblp:169/3288
· DBLP profile ↗
10ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0002-4267-6063ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Paragraph-to-Image Generation with Information-Enriched Diffusion Model
Weijia Wu 0001, Zhuang Li 0002, Yefei He, Zheng Shou 0001, Chunhua Shen, Lele Cheng, Tingting Gao |
Int. J. Comput. Vis. | 6 |
| 2024 | FashionERN: Enhance-and-Refine Network for Composed Fashion Image RetrievalabstractThe goal of composed fashion image retrieval is to locate a target image based on a reference image and modified text. Recent methods utilize symmetric encoders (e.g., CLIP) pre-trained on large-scale non-fashion datasets. However, the input for this task exhibits an asymmetric nature, where the reference image contains rich content while the modified text is often brief. Therefore, methods employing symmetric encoders encounter a severe phenomenon: retrieval results dominated by reference images, leading to the oversight of modified text. We propose a Fashion Enhance-and-Refine Network (FashionERN) centered around two aspects: enhancing the text encoder and refining visual semantics. We introduce a Triple-branch Modifier Enhancement model, which injects relevant information from the reference image and aligns the modified text modality with the target image modality. Furthermore, we propose a Dual-guided Vision Refinement model that retains critical visual information through text-guided refinement and self-guided refinement processes. The combination of these two models significantly mitigates the reference dominance phenomenon, ensuring accurate fulfillment of modifier requirements. Comprehensive experiments demonstrate our approach's state-of-the-art performance on four commonly used datasets. Yanzhe Chen, Huasong Zhong, Xiangteng He, Yuxin Peng 0001, Jiahuan Zhou, Lele Cheng |
AAAI | 6 |
| 2024 | Decouple Content and Motion for Conditional Image-to-Video GenerationabstractThe goal of conditional image-to-video (cI2V) generation is to create a believable new video by beginning with the condition, i.e., one image and text. The previous cI2V generation methods conventionally perform in RGB pixel space, with limitations in modeling motion consistency and visual continuity. Additionally, the efficiency of generating videos in pixel space is quite low. In this paper, we propose a novel approach to address these challenges by disentangling the target RGB pixels into two distinct components: spatial content and temporal motions. Specifically, we predict temporal motions which include motion vector and residual based on a 3D-UNet diffusion model. By explicitly modeling temporal motions and warping them to the starting image, we improve the temporal consistency of generated videos. This results in a reduction of spatial redundancy, emphasizing temporal details. Our proposed method achieves performance improvements by disentangling content and motion, all without introducing new structural complexities to the model. Extensive experiments on various datasets confirm our approach's superior performance over the majority of state-of-the-art methods in both effectiveness and efficiency. Cuifeng Shen, Yulu Gan, Xiongwei Zhu, Lele Cheng, Tingting Gao, Jinzhi Wang |
AAAI | 5 |
| 2024 | Show Me a Video: A Large-Scale Narrated Video Dataset for Coherent Story IllustrationabstractIllustrating a multi-sentence story with visual content is a significant challenge in multimedia research. While previous works have focused on sequential story-to-visual representations at the image level or representing a single sentence with a video clip, illustrating a long multi-sentence story with coherent videos remains an under-explored area. In this paper, we propose the task of video-based story illustration that focuses on the goal of visually illustrating a story with retrieved video clips. To support this task, we first create a large-scale dataset of coherent video stories in each sample, consisting of 85K narrative stories with 60 pairs of consistent clips and texts. We then propose the Story Context-Enhanced Model, which leverages local and global contextual information within the story, inspired by sequence modeling in language understanding. Through comprehensive quantitative experiments, we demonstrate the effectiveness of our baseline model. In addition, qualitative results and detailed user studies reveal that our method can retrieve coherent video sequences from stories. The dataset and code will be made publicly athttps://nfy-dot.github.io/CVSV-dataset/. Yu Lu 0019, Feiyue Ni, Linchao Zhu, Zongxin Yang, Ruihua Song, Lele Cheng, Yi Yang 0001 |
IEEE Trans. Multim. | 8 |
| 2023 | Real20M: A Large-scale E-commerce Dataset for Cross-domain RetrievalabstractIn e-commerce, products and micro-videos serve as two primary carriers. Introducing cross-domain retrieval between these carriers can establish associations, thereby leading to the advancement of specific scenarios, such as retrieving products based on micro-videos or recommending relevant videos based on products. However, existing datasets only focus on retrieval within the product domain while neglecting the micro-video domain and often ignore the multi-modal characteristics of the product domain. Additionally, these datasets strictly limit their data scale through content alignment and use a content-based data organization format that hinders the inclusion of user retrieval intentions. To address these limitations, we propose the PKU Real20M dataset, a large-scale e-commerce dataset designed for cross-domain retrieval. We adopt a query-driven approach to efficiently gather over 20 million e-commerce products and micro-videos, including multimodal information. Additionally, we design a three-level entity prompt learning framework to align inter-modality information from coarse to fine. Moreover, we introduce the Query-driven Cross-Domain retrieval framework (QCD), which leverages user queries to facilitate efficient alignment between the product and micro-video domains. Extensive experiments on two downstream tasks validate the effectiveness of our proposed approaches. The dataset and source code are available at https://github.com/PKU-ICST-MIPL/Real20M_ACMMM2023. Yanzhe Chen, Huasong Zhong, Xiangteng He, Yuxin Peng 0001, Lele Cheng |
ACM Multimedia | 5 |
| 2023 | MV-Diffusion: Motion-aware Video Diffusion ModelabstractIn this paper, we present a Motion-aware Video Diffusion Model (MV-Diffusion) for enhancing the temporal consistency of generated videos using autoregressive diffusion models. Despite the success of diffusion models in various vision generation tasks, generating high-quality and realistic videos with coherent temporal structure remains a challenging problem. Current methods have primarily focused on capturing implicit motion features within a restricted window of RGB frames, rather than explicitly modeling the motion. To address this, we focus on improving the temporal modeling ability of the current autoregressive video diffusion approach by leveraging rich temporal trajectory information in a global context and explicitly modeling local motion trends. The main contributions of this research include: (1) a Trajectory Modeling (TM) block that enhances the model's conditioning by incorporating global motion trajectory information, (2) a Motion Trend Attention (MTA) block that utilizes a cross-attention mechanism to explicitly infer motion trends from the optical flow rather than implicitly learning from RGB input. Experimental results on three video generation tasks using four datasets show the effectiveness of our proposed MV-Diffusion, outperforming existing state-of-the-art approaches. The code is available at https://github.com/PKU-ICST-MIPL/MV-Diffusion_ACMMM2023. Zijun Deng, Xiangteng He, Yuxin Peng 0001, Xiongwei Zhu, Lele Cheng |
ACM Multimedia | 5 |
| 2021 | Learning From Large-Scale Noisy Web Data With Ubiquitous Reweighting for Image ClassificationabstractMany important advances of deep learning techniques have originated from the efforts of addressing the image classification task on large-scale datasets. However, the construction of clean datasets is costly and time-consuming since the Internet is overwhelmed by noisy images with inadequate and inaccurate tags. In this paper, we propose a Ubiquitous Reweighting Network (URNet) that can learn an image classification model from noisy web data. By observing the web data, we find that there are five key challenges, i.e., imbalanced class sizes, high intra-classes diversity and inter-class similarity, imprecise instances, insufficient representative instances, and ambiguous class labels. With these challenges in mind, we assume every training instance has the potential to contribute positively by alleviating the data bias and noise via reweighting the influence of each instance according to different class sizes, large instance clusters, its confidence, small instance bags, and the labels. In this manner, the influence of bias and noise in the data can be gradually alleviated, leading to the steadily improving performance of URNet. Experimental results in the WebVision 2018 challenge with 16 million noisy training images from 5000 classes show that our approach outperforms state-of-the-art models and ranks first place in the image classification task. Jia Li 0003, Yafei Song 0002, Lele Cheng, Pengcheng Yuan, Shumin Han |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Weakly Supervised Learning with Side Information for Noisy Labeled Images
Lele Cheng, Xiangzeng Zhou, Dangwei Li, Hong Shang |
ECCV (30) | 1 |
| 2015 | Facial landmark detection via cascade multi-channel convolutional neural networkabstractThis paper presents a novel cascade multi-channel convolutional neural networks(CMC-CNN) approach for face alignment. Several CNN are jointly used for the finally output. In our method, each stage CNN takes the local region around the landmarks as input, and each local patches does convolution separately, which can lead network to learn local high-level features. Then a fully connected layer is put to learn global information from these local features. Our methods has achieves the state-of-the-art results when tested on the 300 Face in-the-Wild(300-W) dataset. Qiqi Hou, Jinjun Wang, Lele Cheng, Yihong Gong |
ICIP | 3 |
| 2015 | Robust Deep Auto-encoder for Occluded Face RecognitionabstractOcclusions by sunglasses, scarf, hats, beard, shadow etc, can significantly reduce the performance of face recognition systems. Although there exists a rich literature of researches focusing on face recognition with illuminations, poses and facial expression variations, there is very limited work reported for occlusion robust face recognition. In this paper, we present a method to restore occluded facial regions using deep learning technique to improve face recognition performance. Inspired by SSDA for facial occlusion removal with known occlusion type and explicit occlusion location detection from a preprocessing step, this paper further introduces Double Channel SSDA (DC-SSDA) which requires no prior knowledge of the types and the locations of occlusions. Experimental results based on CMU-PIE face database have showed that, the proposed method is robust to a variety of occlusion types and locations, and the restored faces could yield significant recognition performance improvements over occluded ones. Lele Cheng, Jinjun Wang, Yihong Gong, Qiqi Hou |
ACM Multimedia | 1 |