Mingyu Ouyang

dblp:354/8141 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Image and video processing · 50% Image and video coding · 50%
Artificial intelligence
2 papers
Language models and text generation · 90% Multi-agent systems · 10%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › controllable text generation
personalized text generation
1.012026
SlideTailor: Personalized Presentation Slide Generation for Scientific Papers · AAAI 2026
Natural language and speech › Language models and text generation › text summarization › document summarization
presentation slide generation
1.012026
SlideTailor: Personalized Presentation Slide Generation for Scientific Papers · AAAI 2026
Image and video coding › transform coding
DCT coefficient recovery
0.812024
JPEG Quantized Coefficient Recovery via DCT Domain Spatial-Frequential Transformer · IEEE Trans. Image Process. 2024
Image and video processing
image restoration
0.812024
JPEG Quantized Coefficient Recovery via DCT Domain Spatial-Frequential Transformer · IEEE Trans. Image Process. 2024
Image and video processing › image restoration › compression artifact removal
JPEG artifact removal
0.812024
JPEG Quantized Coefficient Recovery via DCT Domain Spatial-Frequential Transformer · IEEE Trans. Image Process. 2024
Image and video coding
transform coding
0.812024
JPEG Quantized Coefficient Recovery via DCT Domain Spatial-Frequential Transformer · IEEE Trans. Image Process. 2024
Human-AI interaction › automation
GUI automation
0.812024
AssistGUI: Task-Oriented PC Graphical User Interface Automation · CVPR 2024
Knowledge, reasoning and agents › Multi-agent systems
multi-agent collaboration
0.212024
AssistGUI: Task-Oriented PC Graphical User Interface Automation · CVPR 2024

Methods — techniques the papers use, named apart from their topics

multi-agent collaboration · 1.5large language model · 1.5chain-of-speech · 1.0agentic framework · 1.0transformer · 0.8deep neural network · 0.8
YearPublicationVenuePosition
2026 SlideTailor: Personalized Presentation Slide Generation for Scientific Papers
abstract
Automatic presentation slide generation can greatly streamline content creation. However, since preferences of each user may vary, existing under-specified formulations often lead to suboptimal results that fail to align with individual user needs. We introduce a novel task that conditions paper-to-slides generation on user-specified preferences. We propose a human behavior-inspired agentic framework, SlideTailor, that progressively generates editable slides in a user-aligned manner. Instead of requiring users to write their preferences in detailed textual form, our system only asks for a paper-slides example pair and a visual template—natural and easy-to-provide artifacts that implicitly encode rich user preferences across content and visual style. Despite the implicit and unlabeled nature of these inputs, our framework effectively distills and generalizes the preferences to guide customized slide generation. We also introduce a novel chain-of-speech mechanism to align slide content with planned oral narration. Such a design significantly enhances the quality of generated slides and enables downstream applications like video presentations. To support this new task, we construct a benchmark dataset that captures diverse user preferences, with carefully designed interpretable metrics for robust evaluation. Extensive experiments demonstrate the effectiveness of our framework.
Wenzheng Zeng, Mingyu Ouyang, Langyuan Cui, Hwee Tou Ng
AAAI2
2024 AssistGUI: Task-Oriented PC Graphical User Interface Automation
abstract
Graphical User Interface (GUI) automation holds significant promise for assisting users with complex tasks, thereby boosting human productivity. Existing works leveraging Large Language Model (LLM) or LLM-based AI agents have shown capabilities in automating tasks on Android and Web platforms. However, these tasks are primarily aimed at simple device usage and entertainment operations. This paper presents a novel benchmark, Assistgui, to evaluate whether models are capable of manipulating the mouse and keyboard on the Windows platform in response to user-requested tasks. We carefully collected a set of 100 tasks from nine widely-used software applications, such as, After Effects and MS Word, each accompanied by the necessary project files for better evaluation. Moreover, we propose a multi-agent collaboration framework, which incorporates four agents to perform task decomposition, GUI parsing, action generation, and reflection. Our experimental results reveal that our multi-agent collaboration mechanism outshines existing methods in performance. Nevertheless, the potential remains substantial, with the best model attaining only a 46% success rate on our benchmark. We conclude with a thorough analysis of the current methods' limitations, setting the stage for future breakthroughs in this domain.
Difei Gao, Lei Ji 0001, Zechen Bai, Mingyu Ouyang, Dongxing Mao, Qinchen Wu, Peiyi Wang, Xiangwu Guo, Hengxu Wang, Luowei Zhou, Zheng Shou 0001
CVPR4
2024 JPEG Quantized Coefficient Recovery via DCT Domain Spatial-Frequential Transformer
abstract
JPEG compression adopts the quantization of Discrete Cosine Transform (DCT) coefficients for effective bit-rate reduction, whilst the quantization could lead to a significant loss of important image details. Recovering compressed JPEG images in the frequency domain has recently garnered increasing interest, complementing the multitude of restoration techniques established in the pixel domain. However, existing DCT domain methods typically suffer from limited effectiveness in handling a wide range of compression quality factors or fall short in recovering sparse quantized coefficients and the components across different colorspaces. To address these challenges, we propose a DCT domain spatial-frequential Transformer, namely DCTransformer, for JPEG quantized coefficient recovery. Specifically, a dual-branch architecture is designed to capture both spatial and frequential correlations within the collocated DCT coefficients. Moreover, we incorporate the operation of quantization matrix embedding, which effectively allows our single model to handle a wide range of quality factors, and a luminance-chrominance alignment head that produces a unified feature map to align different-sized luminance and chrominance components. Our proposed DCTransformer outperforms the current state-of-the-art JPEG artifact removal techniques, as demonstrated by our extensive experiments.
Mingyu Ouyang, Zhenzhong Chen 0001
IEEE Trans. Image Process.1