Mengli Cheng

dblp:220/3324 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
5since 2021 · last 2025
0009-0007-7603-9787ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Generative modeling · 42% Image recognition and object detection · 14% Information extraction and text analysis · 11%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 56% Recommender systems · 44%

Topics — the 16 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
0.912025
EasyAnimate: High-Performance Video Generation Framework with Hybrid Windows Attention and Reward Backpropagation · ACM Multimedia 2025
Machine learning › Generative modeling › diffusion model
diffusion transformer
0.912025
EasyAnimate: High-Performance Video Generation Framework with Hybrid Windows Attention and Reward Backpropagation · ACM Multimedia 2025
Machine learning › Generative modeling › diffusion model
video diffusion model
0.912025
EasyAnimate: High-Performance Video Generation Framework with Hybrid Windows Attention and Reward Backpropagation · ACM Multimedia 2025
Machine learning › Representation and self-supervised learning › pre-training
multimodal pretraining
0.812024
MuLTI: Efficient Video-and-Language Understanding with Text-Guided MultiWay-Sampler and Multiple Choice Modeling · AAAI 2024
Computer vision › Vision and language
video-language understanding
0.812024
MuLTI: Efficient Video-and-Language Understanding with Text-Guided MultiWay-Sampler and Multiple Choice Modeling · AAAI 2024
Machine learning and data management › automated machine learning
hyperparameter optimization
0.712023
EasyRec: An Easy-to-Use, Extendable and Efficient Framework for Building Industrial Recommendation Systems · AAAI 2023
Recommender systems
industrial recommendation
0.712023
EasyRec: An Easy-to-Use, Extendable and Efficient Framework for Building Industrial Recommendation Systems · AAAI 2023
Machine learning › Efficient and distributed learning
distributed training
0.512021
EasyASR: A Distributed Machine Learning Platform for End-to-end Automatic Speech Recognition · AAAI 2021
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
end-to-end speech recognition
0.512021
EasyASR: A Distributed Machine Learning Platform for End-to-end Automatic Speech Recognition · AAAI 2021
Natural language and speech › Information extraction and text analysis › document analysis
document structure extraction
0.412020
One-shot Text Field labeling using Attention and Belief Propagation for Structure Information Extraction · ACM Multimedia 2020
Machine learning › Generative modeling
generative adversarial network
0.412019
Scene Text Recognition with Auto-Aligned Feature Generator · ICDM 2019
Computer vision › Image recognition and object detection
scene text recognition
0.412019
Scene Text Recognition with Auto-Aligned Feature Generator · ICDM 2019
Computer vision › Image recognition and object detection › scene text detection
multi-oriented scene text detection
0.312018
IncepText: A New Inception-Text Module with Deformable PSROI Pooling for Multi-Oriented Scene Text Detection · IJCAI 2018
Computer vision › Image recognition and object detection
scene text detection
0.312018
IncepText: A New Inception-Text Module with Deformable PSROI Pooling for Multi-Oriented Scene Text Detection · IJCAI 2018
Machine learning › Generative modeling › video generation
efficient video generation
0.312025
EasyAnimate: High-Performance Video Generation Framework with Hybrid Windows Attention and Reward Backpropagation · ACM Multimedia 2025
Machine learning and data management
online learning
0.212023
EasyRec: An Easy-to-Use, Extendable and Efficient Framework for Building Industrial Recommendation Systems · AAAI 2023

Methods — techniques the papers use, named apart from their topics

end-to-end neural ASR · 1.0distributed training · 1.0reward backpropagation · 0.9multimodal large language model · 0.9hybrid window attention · 0.9self-attention · 0.8multiple choice modeling · 0.8adapt-pooling residual mapping · 0.8online learning · 0.7hyperparameter optimization · 0.7feature selection · 0.7belief propagation · 0.4attention mechanism · 0.4
YearPublicationVenuePosition
2025 EasyAnimate: High-Performance Video Generation Framework with Hybrid Windows Attention and Reward Backpropagation
abstract
This paper introduces EasyAnimate, an efficient and high quality video generation framework that leverages diffusion transformers to achieve high-quality video production, encompassing data processing, model training, and end-to-end inference. Despite substantial advancements achieved by video diffusion models, existing video generation models still struggles with slow generation speeds and less-than-ideal video quality. To improve training and inference efficiency without compromising performance, we propose Hybrid Window Attention. We design the multidirectional sliding window attention in Hybrid Window Attention, which provides stronger receptive capabilities in 3D dimensions compared to naive one, while reducing the model's computational complexity as the video sequence length increases. To enhance video generation quality, we optimize EasyAnimate using reward backpropagation to better align with human preferences. As a post-training method, it greatly enhances the model's performance while ensuring efficiency. In addition to the aforementioned improvements, EasyAnimate integrates a series of further refinements that significantly improve both computational efficiency and model performance. We introduce a new training strategy called Training with Token Length to resolve uneven GPU utilization in training videos of varying resolutions and lengths, thereby enhancing efficiency. Additionally, we use a multimodal large language model as the text encoder to improve text comprehension of the model. Experiments demonstrate significant enhancements resulting from the above improvements. The EasyAnimate achieves state-of-the-art performance on both the VBench leaderboard and human evaluation. Code and pre-trained models are available at https://github.com/aigc-apps/EasyAnimate.
Kunzhe Huang, Xinyi Zou, Yunkuo Chen, Mengli Cheng, Jun Huang 0007
ACM Multimedia6
2024 MuLTI: Efficient Video-and-Language Understanding with Text-Guided MultiWay-Sampler and Multiple Choice Modeling
abstract
Video-and-language understanding has a variety of applications in the industry, such as video question answering, text-video retrieval, and multi-label classification. Existing video-and-language understanding methods generally adopt heavy multi-modal encoders and feature fusion modules, which consume high computational costs. Specially, they have difficulty dealing with dense video frames or long text prevalent in industrial applications. This paper proposes MuLTI, a highly accurate and efficient video-and-language understanding model that achieves efficient and effective feature fusion and rapid adaptation to downstream tasks. Specifically, we design a Text-Guided MultiWay-Sampler based on adapt-pooling residual mapping and self-attention modules to sample long sequences and fuse multi-modal features, which reduces the computational costs and addresses performance degradation caused by previous samplers. Therefore, MuLTI can handle longer sequences with limited computational costs. Then, to further enhance the model's performance and fill in the lack of pretraining tasks in the video question answering, we propose a new pretraining task named Multiple Choice Modeling. This task bridges the gap between pretraining and downstream tasks and improves the model's ability to align video and text features. Benefiting from the efficient feature fusion module and the new pretraining task, MuLTI achieves state-of-the-art performance on multiple datasets. Implementation and pretrained models will be released.
Yunkuo Chen, Mengli Cheng
AAAI4
2023 EasyRec: An Easy-to-Use, Extendable and Efficient Framework for Building Industrial Recommendation Systems
abstract
We present EasyRec, an easy-to-use, extendable and efficient recommendation framework for building industrial recommendation systems. Our EasyRec framework is superior in the following aspects:first, EasyRec adopts a modular and pluggable design pattern to reduce the efforts to build custom models; second, EasyRec implements hyper-parameter optimization and feature selection algorithms to improve model performance automatically; third, EasyRec applies online learning to adapt to the ever-changing data distribution. The code is released: https://github.com/alibaba/EasyRec.
Mengli Cheng, Hongsheng Jin
AAAI1
2021 EasyASR: A Distributed Machine Learning Platform for End-to-end Automatic Speech Recognition
abstract
We present EasyASR, a distributed machine learning platform for training and serving large-scale Automatic Speech Recognition (ASR) models, as well as collecting and processing audio data at scale. Our platform is built upon the Machine Learning Platform for AI of Alibaba Cloud. Its main functionality is to support efficient learning and inference for end-to-end ASR models on distributed GPU clusters. It allows users to learn ASR models with either pre-defined or user-customized network architectures via simple user interface. On EasyASR, we have produced state-of-the-art results over several public datasets for Mandarin speech recognition.
Chengyu Wang 0001, Mengli Cheng, Jun Huang 0007
AAAI2
2021 Weakly Supervised Construction of ASR Systems from Massive Video Data
abstract
Building large-scale Automatic Speech Recognition (ASR) systems from scratch is significantly challenging, mostly due to the time-consuming and financially-expensive process of annotating a large amount of audio data with transcripts.Although several unsupervised pre-training models have been proposed, applying such models directly might be sub-optimal if more labeled, training data could be obtained without a large cost.In this paper, we present a weakly supervised framework for constructing ASR systems with massive video data.As videos often contain human-speech audio aligned with subtitles, we consider videos as an important knowledge source, and propose an effective approach to extract high-quality audio aligned with transcripts from videos based on text detection and Optical Character Recognition.The underlying ASR model can be fine-tuned to fit any domain-specific target training datasets after weakly supervised pre-training.Extensive experiments show that our framework can easily produce state-of-the-art results on six public datasets for Mandarin speech recognition.
Mengli Cheng, Chengyu Wang 0001, Jun Huang 0007
Interspeech1
2020 One-shot Text Field labeling using Attention and Belief Propagation for Structure Information Extraction
abstract
Structured information extraction from document images usually consists of three steps: text detection, text recognition, and text field labeling. While text detection and text recognition have been heavily studied and improved a lot in literature, text field labeling is less explored and still faces many challenges. Existing learning based methods for text labeling task usually require a large amount of labeled examples to train a specific model for each type of document. However, collecting large amounts of document images and labeling them is difficult and sometimes impossible due to privacy issues. Deploying separate models for each type of document also consumes a lot of resources. Facing these challenges, we explore one-shot learning for the text field labeling task. Existing one-shot learning methods for the task are mostly rule-based and have difficulty in labeling fields in crowded regions with few landmarks and fields consisting of multiple separate text regions. To alleviate these problems, we proposed a novel deep end-to-end trainable approach for one-shot text field labeling, which makes use of attention mechanism to transfer the layout information between document images. We further applied conditional random field on the transferred layout information for the refinement of field labeling. We collected and annotated a real-world one-shot field labeling dataset with a large variety of document types and conducted extensive experiments to examine the effectiveness of the proposed model. To stimulate research in this direction, the collected dataset and the one-shot model will be released (https://github.com/AlibabaPAI/one_shot_text_labeling).
Mengli Cheng, Minghui Qiu, Jun Huang 0007, Wei Lin 0016
ACM Multimedia1
2019 Scene Text Recognition with Auto-Aligned Feature Generator
abstract
Scene text recognition has attracted increasing attention in computer vision due to its various applications. Most of the existing scene text recognition methods are under the encoder-decoder framework. In order to improve text feature learning of these methods, Generative Adversarial Networks (GANs) are recently integrated to generate clean text images without distorted letters. However, the existing GANs assume the input images are spatially aligned, while the words in natural images are often in irregular shapes. The misalignment brings a big problem for both image generation and text recognition. In this paper, we present a novel text feature alignment network to solve this problem. Our method can handle both horizontal and vertical images with irregular texts. Our proposed framework is end-to-end trainable, and extensive experiments on several public benchmarks demonstrate its superiority in terms of both effectiveness and efficiency.
Qiangpeng Yang, Hongsheng Jin, Mengli Cheng, Wenmeng Zhou, Jun Huang 0007, Wei Lin 0016
ICDM3
2018 IncepText: A New Inception-Text Module with Deformable PSROI Pooling for Multi-Oriented Scene Text Detection
abstract
Incidental scene text detection, especially for multi-oriented text regions, is one of the most challenging tasks in many computer vision applications.Different from the common object detection task, scene text often suffers from a large variance of aspect ratio, scale, and orientation. To solve this problem, we propose a novel end-to-end scene text detector IncepText from an instance-aware segmentation perspective. We design a novel Inception-Text module and introduce deformable PSROI pooling to deal with multi-oriented text detection. Extensive experiments on ICDAR2015, RCTW-17, and MSRA-TD500 datasets demonstrate our method's superiority in terms of both effectiveness and efficiency. Our proposed method achieves 1st place result on ICDAR2015 challenge and the state-of-the-art performance on other datasets. Moreover, we have released our implementation as an OCR product which is available for public access.
Qiangpeng Yang, Mengli Cheng, Wenmeng Zhou, Minghui Qiu, Wei Lin 0016
IJCAI2