EDBT 2026 Demo / reviewers in the wild / expert
Yupeng Zhou
dblp:209/8098
· DBLP profile ↗
29ranked-venue papers
11as first author
24since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 8 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dimension-Aware Active Annotation for Aesthetic Perception via Multi-Agent Human-AI CollaborationabstractTo cultivate students' aesthetic development, teachers must objectively interpret and evaluate the artistic qualities and emotional resonance within their paintings—a process known as aesthetic perception. This evaluation process is labor-intensive and susceptible to biases due to variations among individual teachers. Advances in artificial intelligence (AI) motivate the use of AI-driven models to automate and enhance this aesthetic perception task. However, building effective AI-driven aesthetic perception models requires extensive datasets, which are typically labor-intensive and costly to gather. To address this, we propose a novel framework that selectively identifies the most challenging dimensions of aesthetic perception for expert annotation, using AI-generated pseudo-annotations to reduce cost and improve model performance. Our framework integrates a multi-agent active learning strategy to systematically annotate scores across multiple dimensions of aesthetic perception. Initially, we train an aesthetic perception model using a small, manually annotated dataset, establishing primary annotation capabilities. Then, this trained model generates pseudo-annotations for unlabeled data across various aesthetic dimensions (e.g., humor, happiness). To ensure annotation quality and relevance, a multi-agent system evaluates these pseudo-annotations, identifying dimensions requiring expert human input based on metrics such as model estimation confidence. Human experts provide targeted annotations selectively, refining the dataset and guiding an iterative improvement cycle. Through repeated refinement, the model progressively enhances both its predictive accuracy and its automated annotation proficiency. Our optimization approach dynamically balances accuracy, annotation relevance, and human effort. Extensive experiments conducted on two real-world datasets demonstrate the effectiveness of our framework. Ye Zhang 0014, Dongjie Wang 0001, Yupeng Zhou, Minghao Yin |
AAAI | 4 |
| 2026 | An exact algorithm with a new upper bound and reductions for maximum edge weighted clique in massive sparse graphs
Shuli Hu, Yupeng Zhou, Minghao Yin |
Frontiers Comput. Sci. | 4 |
| 2026 | SRFormerV2: Taking a Closer Look at Permuted Self-Attention for Image Super-ResolutionabstractPrevious works have shown that increasing the window size for Transformer-based image super-resolution models (e.g., SwinIR) can significantly improve the model performance. Still, the computation overhead is also considerable when the window size gradually increases. In this paper, we present SRFormer, a simple but novel method that can enjoy the benefit of large window self-attention but introduces even less computational burden. The core of our SRFormer is the permuted self-attention (PSA), which strikes an appropriate balance between the channel and spatial information for self-attention. Without any bells and whistles, we show that our SRFormer achieves a 33.86 dB PSNR score on the Urban100 dataset, which is 0.46 dB higher than that of SwinIR but uses fewer parameters and computations. In addition, we also attempt to scale up the model by further enlarging the window size and channel numbers to explore the potential of Transformer-based models. Experiments show that our scaled model, named SRFormerV2, can further improve the results and achieves state-of-the-art. We hope our simple and effective approach could be useful for future research in super-resolution model design. Yupeng Zhou, Zhen Li 0031, Chunle Guo, Li Liu 0002, Ming-Ming Cheng, Qibin Hou |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | LLM4AD: Large Language Models for Autonomous Driving - Concept, Review, Benchmark, Experiments, and Future TrendsabstractWith the broader adoption and highly successful development of large language models (LLMs), there has been growing interest and demand for applying LLMs to autonomous driving technology. Driven by their natural language (NL) understanding and reasoning capabilities, LLMs have the potential to enhance various aspects of autonomous driving systems, from perception and scene understanding to interactive decision-making. This article first introduces the novel concept of designing LLMs for autonomous driving (LLM4AD), followed by a review of existing LLM4AD studies. Then, a comprehensive benchmark is proposed for evaluating the instruction-following and reasoning abilities of LLM4AD systems, which includes LaMPilot-Bench, CARLA Leaderboard 1.0 Benchmark in simulation and NuPlanQA for multiview visual question answering (VQA). Furthermore, extensive real-world experiments are conducted on autonomous vehicle platforms, examining both on-cloud and on-edge LLM deployment for personalized decision-making and motion control. Next, the future trends of integrating language diffusion models into autonomous driving are explored, exemplified by the proposed vision-language diffusion (ViLaD) framework. Finally, the main challenges of LLM4AD are discussed, including latency, deployment, security and privacy, safety, trust and transparency, and personalization. Can Cui 0009, Yunsheng Ma, Sungyeon Park 0001, Zichong Yang, Yupeng Zhou, Peiran Liu 0003, Juanwu Lu, Juntong Peng, Jiaru Zhang, Ruqi Zhang, Lingxi Li 0001, Yaobin Chen, Jitesh H. Panchal, Amr Abdelraouf, Kyungtae Han, Ziran Wang |
Proc. IEEE | 5 |
| 2026 | A Hierarchical Test Platform for Vision Language Model (VLM)-Integrated Real-World Autonomous DrivingabstractVision-Language Models (VLMs) have demonstrated significant promise for autonomous driving due to their powerful multimodal reasoning capabilities. However, adapting VLMs from generic data to safety-critical driving contexts introduces a notable challenge known as domain shift. Existing simulation-based and dataset-driven evaluation approaches struggle to accurately replicate real-world complexities, lacking repeatable closed-loop evaluation and flexible scenario manipulation. Furthermore, current real-world testing platforms typically focus on isolated modules and do not support comprehensive interaction with VLM-based systems. Consequently, there is a critical need for a holistic testing architecture capable of integrating perception, planning, and control modules, accommodating VLM-based systems, and supporting configurable real-world testing scenarios. In this article, we address this critical gap by proposing a hierarchical real-world test platform specialized in the rigorous evaluation of VLM-integrated autonomous driving systems. Specifically, our platform features have: a lightweight, structured, and low-latency middleware pipeline specialized for seamless VLM integration; a hierarchical modular architecture enabling flexible substitution between conventional and VLM-based autonomy components, providing exceptional deployment flexibility for rapid experimentation; and sophisticated closed-loop scenario-based testing capabilities on a controlled test track, facilitating comprehensive evaluation of the entire full-stack VLM-integrated autonomous driving pipeline, from perception, reasoning, decision-making, and planning to final vehicle maneuvers. Through an extensive real-world case study, we demonstrate the effectiveness of our platform in evaluating the performance and robustness of VLM-integrated autonomous driving under diverse realistic conditions. Project page and codes: https://github.com/YupengZhouPurdue/VLMTest . Yupeng Zhou, Can Cui 0009, Juntong Peng, Zichong Yang, Juanwu Lu, Jitesh H. Panchal, Bin Yao 0001, Ziran Wang |
ACM Trans. Internet Things | 1 |
| 2025 | Multi-type MOOCs Recommendation: Leveraging Deep Multi-Relational Representation and Hierarchical ReasoningabstractMassive open online courses (MOOCs) recommendation provides online courses tailored to learners' individual preferences. Existing literature is limited by: 1) Ignoring the interrelations among courses, knowledge concepts, and videos, which leads to suboptimal recommendation performance; 2) Neglecting the hierarchical interactions between learners and components like courses, knowledge concepts, and videos, which makes it difficult to capture learners' intentions accurately. To address them, we propose a novel multi-type MOOCs recommendation framework, which enables multi-type educational content recommendations. This framework includes two important components: multi-relational representation and hierarchical reasoning. Regarding multi-relational representation, we first create two static course-relational and knowledge concept-relational graphs based on domain knowledge and construct a dynamic video-relational graph using learners' browsing historical sequences. Then, we capture the interactions among different components by learning the corresponding embeddings via graph neural networks. Regarding hierarchical reasoning, we implement a hierarchical beam search strategy to narrow down the candidate courses, knowledge concepts, and videos by calculating joint probability. Finally, we introduce an optional layer to increase the diversity and reasonableness of video recommendations by estimating learners' intentions. Extensive experiments are conducted to show the effectiveness, robustness, and interpretability of our method. Ye Zhang 0014, Yanqi Gao, Dongjie Wang 0001, Yupeng Zhou, Zhaoyang Sun, Minghao Yin |
AAAI | 4 |
| 2025 | AR-1-to-3: Single Image to Consistent 3D Object via Next-View Prediction
Xuying Zhang, Yupeng Zhou, Kai Wang 0001, Zhen Li 0031, Shaohui Jiao, Daquan Zhou, Qibin Hou, Ming-Ming Cheng |
ICCV | 2 |
| 2025 | Structure-Aware Handwritten Text Recognition via Graph-Enhanced Cross-Modal Mutual LearningabstractExisting handwriting recognition methods only focus on learning visual patterns by modeling low-level relationships of adjacent pixels, while overlooking the intrinsic geometric structures of characters. In this paper, we propose a novel graph-enhanced cross-modal mutual learning network GCM to fully process handwritten text images alongside their corresponding geometric graphs, which consists of one shared cross-modal encoder and two parallel inverse decoders. Specifically, the encoder simultaneously extracts visual and geometric information from the cross-modal inputs, and the decoders fuse the multi-modal features for prediction under the guidance of cross-modal fusion. Moreover, two parallel decoders sequentially aggregate cross-modal features in inverse orders (V→G and G→V) but are enhanced through mutual distillation at each time-step, which involves one-to-one knowledge transfer and fully leverages complementary cross-modal information from both directions. Notably, only one branch of GCM is activated in inference, thus avoiding the increase of the model parameters and computation costs for testing. Experiments show that our method outperforms previous state-of-the-art methods on public benchmarks such as IAM, RIMES, and ICDAR-2013 when no extra training data is utilized. Ji Gan, Yupeng Zhou, Jiaxu Leng, Xinbo Gao 0001 |
IJCAI | 2 |
| 2025 | On-Board Vision-Language Models (VLMs) for Personalized Motion Control of Autonomous VehiclesabstractPersonalized driving refers to an autonomous vehicle’s ability to adapt its driving behavior or control strategies to match individual users’ preferences and driving styles while maintaining safety and comfort standards. However, existing works either fail to capture every individual’s preference precisely or become computationally inefficient as the user base expands. Vision-Language Models (VLMs) offer promising solutions to this front through their natural language understanding and scene reasoning capabilities. In this work, we propose a lightweight yet effective on-board VLM framework that provides low-latency personalized driving performance while maintaining strong reasoning capabilities. Our solution incorporates a Retrieval-Augmented Generation (RAG)-based memory module that enables continuous learning of individual driving preferences through human feedback. Through comprehensive real-world vehicle experiments, our system has demonstrated the ability to provide safe, comfortable, and personalized driving experiences across various scenarios and significantly reduce takeover rates by up to 76.9%. To the best of our knowledge, this work represents the first personalized VLM motion control system in real-world autonomous vehicles. The demo video can be watched at https://tinyurl.com/4xsnz79n. Can Cui 0009, Zichong Yang, Yupeng Zhou, Juntong Peng, Sungyeon Park 0001, Yunsheng Ma, Wenqian Ye, Yiheng Feng, Jitesh H. Panchal, Lingxi Li 0001, Yaobin Chen, Ziran Wang |
IROS | 3 |
| 2025 | Accelerating influence through communities: A scalable approach for maximizing budgeted influence in large-scale networks
Xingjian Ji, Hanhui Liu, Qinglong Hou, Shuli Hu, Minghao Yin, Yupeng Zhou |
Expert Syst. Appl. | 6 |
| 2025 | MaskDiffusion: Boosting Text-to-Image Consistency with Conditional Mask
Yupeng Zhou, Daquan Zhou, Yaxing Wang, Jiashi Feng, Qibin Hou |
Int. J. Comput. Vis. | 1 |
| 2025 | An incremental algorithm for dynamic graph coloring based on graph reduction and adaptive recoloring strategies
Yupeng Zhou, Hanhui Liu, Shuli Hu, Minghao Yin |
J. Supercomput. | 1 |
| 2024 | MRMLREC: A Two-Stage Approach for Addressing Data Sparsity in MOOC Video Recommendation (Student Abstract)abstractWith the abundance of learning resources available on massive open online courses (MOOCs) platforms, the issue of interactive data sparsity has emerged as a significant challenge.This paper introduces MRMLREC, an efficient MOOC video recommendation which consists of two main stages: multi-relational representation and multi-level recommendation, aiming to solve the problem of data sparsity. In the multi-relational representation stage, MRMLREC adopts a tripartite approach, constructing relational graphs based on temporal sequences, courses-videos relation, and knowledge concepts-video relation. These graphs are processed by a Graph Convolution Network (GCN) and two variant Graph Attention Networks (GAT) to derive representations. A variant of the Long Short-Term Memory Network (LSTM) then integrates these multi-dimensional data to enhance the overall representation. The multi-level recommendation stage introduces three prediction tasks at varying levels—courses, knowledge concepts, and videos—to mitigate data sparsity and improve the interpretability of video recommendations. Beam search (BS) is employed to identify top-β items at each level, refining the subsequent level's search space and enhancing recommendation efficiency. Additionally, an optional layer offers both personalization and diversification modes, ensuring variety in recommended videos and maintaining learner engagement. Comprehensive experiments demonstrate the effectiveness of MRMLREC on two real-world instances from Xuetang X. Ye Zhang 0014, Yanqi Gao, Yupeng Zhou, Minghao Yin |
AAAI | 3 |
| 2024 | StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video GenerationabstractFor recent diffusion-based generative models, maintaining consistent content across a series of generated images, especially those containing subjects and complex details, presents a significant challenge. In this paper, we propose a simple but effective self-attention mechanism, termed Consistent Self-Attention, that boosts the consistency between the generated images. It can be used to augment pre-trained diffusion-based text-to-image models in a zero-shot manner. Based on the images with consistent content, we further show that our method can be extended to long range video generation by introducing a semantic space temporal motion prediction module, named Semantic Motion Predictor. It is trained to estimate the motion conditions between two provided images in the semantic spaces. This module converts the generated sequence of images into videos with smooth transitions and consistent subjects that are more stable than the modules based on latent spaces only, especially in the context of long video generation. By merging these two novel components, our framework, referred to as StoryDiffusion, can describe a text-based story with consistent images or videos encompassing a rich variety of contents. The proposed StoryDiffusion encompasses pioneering explorations in visual story generation with the presentation of images and videos, which we hope could inspire more research from the aspect of architectural modifications. Yupeng Zhou, Daquan Zhou, Ming-Ming Cheng, Jiashi Feng, Qibin Hou |
NeurIPS | 1 |
| 2024 | A three-phase sheep optimization algorithm for numerical and engineering optimization problems
Dapeng Qu, Shilin Peng, Zeyu Wen, Changjiu Yu, Yupeng Zhou |
Expert Syst. Appl. | 8 |
| 2023 | SRFormer: Permuted Self-Attention for Single Image Super-ResolutionabstractPrevious works have shown that increasing the window size for Transformer-based image super-resolution models (e.g., SwinIR) can significantly improve the model performance but the computation overhead is also considerable. In this paper, we present SRFormer, a simple but novel method that can enjoy the benefit of large window self-attention but introduces even less computational burden. The core of our SRFormer is the permuted self-attention (PSA), which strikes an appropriate balance between the channel and spatial information for self-attention. Our PSA is simple and can be easily applied to existing super-resolution networks based on window self-attention. Without any bells and whistles, we show that our SRFormer achieves a 33.86dB PSNR score on the Urban100 dataset, which is 0.46dB higher than that of SwinIR but uses fewer parameters and computations. We hope our simple and effective approach can serve as a useful tool for future research in super-resolution model design. Our code is available at https://github.com/HVision-NKU/SRFormer. Yupeng Zhou, Zhen Li 0031, Chunle Guo, Song Bai 0001, Ming-Ming Cheng, Qibin Hou |
ICCV | 1 |
| 2023 | A master-apprentice evolutionary algorithm for maximum weighted set K-covering problem
Yupeng Zhou, Mingjie Fan, Yiyuan Wang 0002, Minghao Yin |
Appl. Intell. | 1 |
| 2023 | A new local search algorithm with greedy crossover restart for the dominating tree problem
Dangdang Niu, Bin Liu 0023, Minghao Yin, Yupeng Zhou |
Expert Syst. Appl. | 4 |
| 2023 | OLFWA: A novel fireworks algorithm with new explosion operator and two stages information utilization
Mingjie Fan, Yupeng Zhou, Mingzhang Han, Xinchao Zhao, Lingjuan Ye |
Inf. Sci. | 2 |
| 2022 | NukCP: An Improved Local Search Algorithm for Maximum k-Club ProblemabstractThe maximum k-club problem (MkCP) is an important clique relaxation problem with wide applications. Previous MkCP algorithms only work on small-scale instances and are not applicable for large-scale instances. For solving instances with different scales, this paper develops an efficient local search algorithm named NukCP for the MkCP which mainly includes two novel ideas. First, we propose a dynamic reduction strategy, which makes a good balance between the time efficiency and the precision effectiveness of the upper bound calculation. Second, a stratified threshold configuration checking strategy is designed by giving different priorities for the neighborhood in the different levels. Experiments on a broad range of different scale instances show that NukCP significantly outperforms the state-of-the-art MkCP algorithms on most instances. Jiejiang Chen, Yiyuan Wang 0002, Shaowei Cai 0001, Minghao Yin, Yupeng Zhou, Jieyu Wu |
AAAI | 5 |
| 2022 | HEA-D: A Hybrid Evolutionary Algorithm for Diversified Top-k Weight Clique Search ProblemabstractThe diversified top-k weight clique (DTKWC) search problem is an important generalization of the diversified top-k clique (DTKC) search problem with extensive applications, which extends the DTKC search problem by taking into account the weight of vertices. In this paper, we formulate DTKWC search problem using mixed integer linear program constraints and propose an efficient hybrid evolutionary algorithm (HEA-D) that combines a clique-based crossover operator and an effective simulated annealing-based local optimization procedure to find high-quality local optima. The experimental results show that HEA-D performs much better than the existing methods on two representative real-world benchmarks. Jun Wu 0020, Chu Min Li 0001, Yupeng Zhou, Minghao Yin, Dangdang Niu |
IJCAI | 3 |
| 2022 | Solving multi-objective constrained minimum weighted bipartite assignment problem: a case study on energy-aware radio broadcast scheduling
Yupeng Zhou, Mingjie Fan, Feifei Ma, Minghao Yin |
Sci. China Inf. Sci. | 1 |
| 2022 | Combining max-min ant system with effective local search for solving the maximum set k-covering problem
Yupeng Zhou, Shuli Hu, Yiyuan Wang 0002, Minghao Yin |
Knowl. Based Syst. | 1 |
| 2021 | Solving diversified top-k weight clique search problem
Junping Zhou, Chu Min Li 0001, Yupeng Zhou, Lili Liang |
Sci. China Inf. Sci. | 3 |
| 2020 | Contention-Aware Mapping and Scheduling Optimization for NoC-Based MPSoCs (Student Abstract)abstractWe consider spacial and temporal aspects of communication to avoid contention in Network-on-Chip (NoC) architectures. A constraint model is constructed such that the design concerns can be evaluated, and an efficient evolutionary algorithm with various heuristics is proposed to search for better solutions. Experimentations from random benchmarks demonstrate the efficiency of our method in multi-objective optimization and the effectiveness of our techniques in avoiding network contention. Yupeng Zhou, Rongjie Yan, Anyu Cai, Yige Yan, Minghao Yin |
AAAI | 1 |
| 2018 | An efficient local search for partial vertex cover problem
Yupeng Zhou, Yiyuan Wang 0002, Jian Gao 0007 |
Neural Comput. Appl. | 1 |
| 2017 | A Hybrid Multi-objective Evolutionary Algorithm for Energy-Aware Allocation and Scheduling Optimization of MPSoCsabstractMPSoCs are increasingly being adopted in the design of emerging complex embedded systems. Resource limitations require designers to find optimizations among various design considerations. Task mapping and scheduling become one of the key issues in designing such systems. To meet the requirements of makespan minimization and workload balance for energy-aware MPSoCs, the paper presents a unified formulation to find satisfied task mapping and scheduling solutions. The model considers both computation and communication cost, and enables applying dynamic power management (DPM) for energy optimization. To efficiently approximate the Pareto front of the optimization problem, we propose a multi-objective hybrid algorithm (MOHA) by integrating a Pareto local search into an evolutionary process, with a problem-specific initialization. Experimental results from realistic benchmarks demonstrate that the proposed techniques are able to generate high-quality solutions of realistic applications on the target architecture, compared with state-of-the-art methods. Rongjie Yan, Yupeng Zhou, Yige Yan, Minghao Yin, Min Yu 0006, Feifei Ma, Kai Huang 0002 |
ICTAI | 2 |
| 2017 | GRASP for connected dominating set problems
Shuli Hu, Jian Gao 0007, Yupeng Zhou, Yiyuan Wang 0002, Minghao Yin |
Neural Comput. Appl. | 4 |
| 2017 | A path cost-based GRASP for minimum independent dominating set problem
Yiyuan Wang 0002, Yupeng Zhou, Minghao Yin |
Neural Comput. Appl. | 3 |