VLDB 2026 Research / reviewers in the wild / expert
Mingyu Ouyang
dblp:354/8141
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
1 paper |
Image and video processing · 50% Image and video coding · 50% | |
| Artificial intelligence
2 papers |
Language models and text generation · 90% Multi-agent systems · 10% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-AI interaction · 100% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › controllable text generation
personalized text generation |
1.0 | 1 | 2026 | SlideTailor: Personalized Presentation Slide Generation for Scientific Papers · AAAI 2026 |
Natural language and speech › Language models and text generation › text summarization › document summarization
presentation slide generation |
1.0 | 1 | 2026 | SlideTailor: Personalized Presentation Slide Generation for Scientific Papers · AAAI 2026 |
Image and video coding › transform coding
DCT coefficient recovery |
0.8 | 1 | 2024 | JPEG Quantized Coefficient Recovery via DCT Domain Spatial-Frequential Transformer · IEEE Trans. Image Process. 2024 |
Image and video processing
image restoration |
0.8 | 1 | 2024 | JPEG Quantized Coefficient Recovery via DCT Domain Spatial-Frequential Transformer · IEEE Trans. Image Process. 2024 |
Image and video processing › image restoration › compression artifact removal
JPEG artifact removal |
0.8 | 1 | 2024 | JPEG Quantized Coefficient Recovery via DCT Domain Spatial-Frequential Transformer · IEEE Trans. Image Process. 2024 |
Image and video coding
transform coding |
0.8 | 1 | 2024 | JPEG Quantized Coefficient Recovery via DCT Domain Spatial-Frequential Transformer · IEEE Trans. Image Process. 2024 |
Human-AI interaction › automation
GUI automation |
0.8 | 1 | 2024 | AssistGUI: Task-Oriented PC Graphical User Interface Automation · CVPR 2024 |
Knowledge, reasoning and agents › Multi-agent systems
multi-agent collaboration |
0.2 | 1 | 2024 | AssistGUI: Task-Oriented PC Graphical User Interface Automation · CVPR 2024 |
Methods — techniques the papers use, named apart from their topics
multi-agent collaboration · 1.5large language model · 1.5chain-of-speech · 1.0agentic framework · 1.0transformer · 0.8deep neural network · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SlideTailor: Personalized Presentation Slide Generation for Scientific PapersabstractAutomatic presentation slide generation can greatly streamline content creation. However, since preferences of each user may vary, existing under-specified formulations often lead to suboptimal results that fail to align with individual user needs. We introduce a novel task that conditions paper-to-slides generation on user-specified preferences. We propose a human behavior-inspired agentic framework, SlideTailor, that progressively generates editable slides in a user-aligned manner. Instead of requiring users to write their preferences in detailed textual form, our system only asks for a paper-slides example pair and a visual template—natural and easy-to-provide artifacts that implicitly encode rich user preferences across content and visual style. Despite the implicit and unlabeled nature of these inputs, our framework effectively distills and generalizes the preferences to guide customized slide generation. We also introduce a novel chain-of-speech mechanism to align slide content with planned oral narration. Such a design significantly enhances the quality of generated slides and enables downstream applications like video presentations. To support this new task, we construct a benchmark dataset that captures diverse user preferences, with carefully designed interpretable metrics for robust evaluation. Extensive experiments demonstrate the effectiveness of our framework. Wenzheng Zeng, Mingyu Ouyang, Langyuan Cui, Hwee Tou Ng |
AAAI | 2 |
| 2024 | AssistGUI: Task-Oriented PC Graphical User Interface AutomationabstractGraphical User Interface (GUI) automation holds significant promise for assisting users with complex tasks, thereby boosting human productivity. Existing works leveraging Large Language Model (LLM) or LLM-based AI agents have shown capabilities in automating tasks on Android and Web platforms. However, these tasks are primarily aimed at simple device usage and entertainment operations. This paper presents a novel benchmark, Assistgui, to evaluate whether models are capable of manipulating the mouse and keyboard on the Windows platform in response to user-requested tasks. We carefully collected a set of 100 tasks from nine widely-used software applications, such as, After Effects and MS Word, each accompanied by the necessary project files for better evaluation. Moreover, we propose a multi-agent collaboration framework, which incorporates four agents to perform task decomposition, GUI parsing, action generation, and reflection. Our experimental results reveal that our multi-agent collaboration mechanism outshines existing methods in performance. Nevertheless, the potential remains substantial, with the best model attaining only a 46% success rate on our benchmark. We conclude with a thorough analysis of the current methods' limitations, setting the stage for future breakthroughs in this domain. Difei Gao, Lei Ji 0001, Zechen Bai, Mingyu Ouyang, Dongxing Mao, Qinchen Wu, Peiyi Wang, Xiangwu Guo, Hengxu Wang, Luowei Zhou, Zheng Shou 0001 |
CVPR | 4 |
| 2024 | JPEG Quantized Coefficient Recovery via DCT Domain Spatial-Frequential TransformerabstractJPEG compression adopts the quantization of Discrete Cosine Transform (DCT) coefficients for effective bit-rate reduction, whilst the quantization could lead to a significant loss of important image details. Recovering compressed JPEG images in the frequency domain has recently garnered increasing interest, complementing the multitude of restoration techniques established in the pixel domain. However, existing DCT domain methods typically suffer from limited effectiveness in handling a wide range of compression quality factors or fall short in recovering sparse quantized coefficients and the components across different colorspaces. To address these challenges, we propose a DCT domain spatial-frequential Transformer, namely DCTransformer, for JPEG quantized coefficient recovery. Specifically, a dual-branch architecture is designed to capture both spatial and frequential correlations within the collocated DCT coefficients. Moreover, we incorporate the operation of quantization matrix embedding, which effectively allows our single model to handle a wide range of quality factors, and a luminance-chrominance alignment head that produces a unified feature map to align different-sized luminance and chrominance components. Our proposed DCTransformer outperforms the current state-of-the-art JPEG artifact removal techniques, as demonstrated by our extensive experiments. Mingyu Ouyang, Zhenzhong Chen 0001 |
IEEE Trans. Image Process. | 1 |