Pengyu Yan

dblp:41/11473 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 PosterVerse: A Full-Workflow Framework for Commercial-Grade Poster Generation with HTML-Based Scalable Typography
abstract
Commercial-grade poster design demands the seamless integration of aesthetic appeal with precise, informative content delivery. Current automated poster generation systems face significant limitations, including incomplete design workflows, poor text rendering accuracy, and insufficient flexibility for commercial applications. To address these challenges, we propose PosterVerse, a full-workflow, commercial-grade poster generation method that seamlessly automates the entire design process while delivering high-density and scalable text rendering. PosterVerse replicates professional design through three key stages: (1) blueprint creation using fine-tuned LLMs to extract key design elements from user requirements, (2) graphical background generation via customized diffusion models to create visually appealing imagery, and (3) unified layout-text rendering with an MLLM-powered HTML engine to guarantee high text accuracy and flexible customization. In addition, we introduce PosterDNA, a commercial-grade, HTML-based dataset tailored for training and validating poster design models. To the best of our knowledge, PosterDNA is the first Chinese poster generation dataset to introduce HTML typography files, enabling scalable text rendering and fundamentally solving the challenges of rendering small and high-density text. Experimental results demonstrate that PosterVerse consistently produces commercial-grade posters with appealing visuals, accurate text alignment, and customizable layouts, making it a promising solution for automating commercial poster design.
Junle Liu, Peirong Zhang 0001, Yuyi Zhang 0002, Pengyu Yan, Xinyue Zhou, Fengjun Guo
AAAI4
2025 Reviving Cultural Heritage: A Novel Approach for Comprehensive Historical Document Restoration
abstract
Yuyi Zhang, Peirong Zhang, Zhenhua Yang, Pengyu Yan, Yongxin Shi, Pengwei Liu, Fengjun Guo, Lianwen Jin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yuyi Zhang 0002, Peirong Zhang 0001, Zhenhua Yang, Pengyu Yan, Yongxin Shi, Pengwei Liu, Fengjun Guo
ACL (1)4
2025 FormerPose: An efficient multi-scale fusion Transformer network based on RGB-D for 6D pose estimation
Pihong Hou, Yongfang Zhang, Pengyu Yan
J. Vis. Commun. Image Represent.4
2024 ChartReformer: Natural Language-Driven Chart Image Editing
Pengyu Yan, Mahesh Bhosale, Jay Lal, Bikhyat Adhikari, David S. Doermann
ICDAR (1)1
2024 Artemis: Towards Referential Understanding in Complex Videos
abstract
Videos carry rich visual information including object description, action, interaction, etc., but the existing multimodal large language models (MLLMs) fell short in referential understanding scenarios such as video-based referring. In this paper, we present Artemis, an MLLM that pushes video-based referential understanding to a finer level. Given a video, Artemis receives a natural-language question with a bounding box in any video frame and describes the referred target in the entire video. The key to achieving this goal lies in extracting compact, target-specific video features, where we set a solid baseline by tracking and selecting spatiotemporal features from the video. We train Artemis on the newly established ViderRef45K dataset with 45K video-QA pairs and design a computationally efficient, three-stage training procedure. Results are promising both quantitatively and qualitatively. Additionally, we show that Artemis can be integrated with video grounding and text summarization tools to understand more complex scenarios. Code and data are available at https://github.com/NeurIPS24Artemis/Artemis.
Jihao Qiu, Lingxi Xie, Tianren Ma, Pengyu Yan, David S. Doermann, Qixiang Ye, Yunjie Tian
NeurIPS6
2024 Two-stage approximation allocation approach for real-time parking reservations considering stochastic requests and reusable resources
Mingyan Bai, Pengyu Yan, Xiaoqiang Cai, Xiang T. R. Kong
Adv. Eng. Informatics2
2024 Parcel consolidation approach and routing algorithm for last-mile delivery by unmanned aerial vehicles
Pengyu Yan, Kaize Yu, Peifan Li
Expert Syst. Appl.2
2023 SpaDen: Sparse and Dense Keypoint Estimation for Real-World Chart Understanding
Saleem Ahmed, Pengyu Yan, David S. Doermann, Srirangaraj Setlur, Venu Govindaraju
ICDAR (2)2
2023 Context-Aware Chart Element Detection
Pengyu Yan, Saleem Ahmed, David S. Doermann
ICDAR (1)1
2018 Unsupervised Saliency Detection in 3-D-Video Based on Multiscale Segmentation and Refinement
abstract
In this letter, we propose an unsupervised salient object detection method in three-dimensional videos. Both temporal and depth information are efficiently considered, and multiscale architecture and graph-based refinement are built to improve accuracy and robustness. First, the input video frame is segmented into nonoverlapping superpixels by combining both appearance and depth information at the input. A multiscale architecture is also deployed after the segmentation with different segmentation parameters. Second, the initial saliency score of each segmented superpixel in each scale is calculated via global contrast, which is defined by appearance, depth, and motion cues from two consecutive frames. Third, the initial saliency in each scale is refined by smoothing over graphs built by three spatial–temporal feature priors—color, depth, and motion. Finally, the result is obtained by fusing three refined saliency maps in three scales. The experiments on two widely used datasets illustrate that our method outperforms state-of-the-art algorithms in terms of accuracy, robustness, and reliability.
Ping Zhang 0023, Pengyu Yan, Fengcan Shen
IEEE Signal Process. Lett.2
2018 A Heuristic for Inserting Randomly Arriving Jobs Into an Existing Hoist Schedule
abstract
Hoist scheduling in automated electroplating lines has been extensively studied in a static environment. However, practical electroplating lines are subject to diversified unforeseen disruptions that require frequent rescheduling to maintain or optimize system performance. This paper addresses a hoist scheduling problem, where randomly arriving jobs need to be inserted into an existing schedule without changing the sequence of hoist moves already scheduled. The objective is to minimize the total completion time of all the jobs in the existing schedule and a newly inserted job. We develop a polynomial-time heuristic that adjusts the starting times of the existing hoist moves to a limited extent but does not bring about a severe disturbance of the existing hoist moves. We compare our algorithm with two existing approaches with different rescheduling policies (i.e., partial and zero adjustment of the existing schedule). We empirically analyze the productivity and the stability of the schedules generated by the three approaches. Computational results demonstrate that our algorithm can generate more productive and stable schedules than the two existing approaches.
Pengyu Yan, Ada Che, Eugene Levner
IEEE Trans Autom. Sci. Eng.1