VLDB 2026 Research / reviewers in the wild / expert
Shuai Peng
dblp:115/6470
· DBLP profile ↗
24ranked-venue papers
11as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 5 since 2021Computer networks · 6 · 5 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Security and privacy · 3 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automated building outline extraction from digital surface models and orthophoto: a novel contour-based approachabstractBuilding outlines have many applications. However, owing to the diversity of buildings and the complexity of the surrounding environment, automatic extraction of building outlines from remote sensing data remains challenging. This paper presents a novel approach for extracting building outlines from digital surface models (DSM) and orthophotographs. The DSM provides initial contour lines, while the orthophotograph indicates where vegetation is obstructing the building outline. The approach introduces two key algorithms: Distance-Constrained Clustering (DCC), to cluster contour lines, and Gradient-based Optimal Contour Selection (G-OCS), to select building outlines. Vegetation information is used to recover obstructed building outlines and improve outline accuracy and completeness. Experimental results, using the International Society for Photogrammetry and Remote Sensing (ISPRS) Vaihingen benchmark dataset, demonstrate the method’s performance (quality metric: 85.0% for individual regions, 73.6% for individual objects, and 99.1% for objects >50 m). Validation using a dataset from Shandong Province (China) confirmed the method’s robustness and applicability for complex urban environments. The approach effectively handles challenges such as interference from vegetation and irregular building structures, outperforming techniques such as WHUZ, CNN/8F+, and HD-Net. This novel method automates building outline extraction and provides useful building information, with applications in urban planning, disaster management, and smart city development. Fangyuqing Jin, Xing Li 0022, Yihu Zhu, Zirui Ou, Yaoyao Ren, Shuai Peng, Wei Liu 0095, Erzhu Li, Lianpeng Zhang |
Int. J. Geogr. Inf. Sci. | 6 |
| 2025 | LogicPro: Improving Complex Logical Reasoning via Program-Guided LearningabstractJin Jiang, Yuchen Yan, Yang Liu, Jianing Wang, Shuai Peng, Xunliang Cai, Yixin Cao, Mengdi Zhang, Liangcai Gao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shuai Peng, Liangcai Gao |
ACL (1) | 5 |
| 2025 | SketchRef: a Multi-Task Evaluation Benchmark for Sketch SynthesisabstractSketching is a powerful artistic technique for capturing essential visual information about real-world objects and has increasingly attracted attention in image synthesis research. However, the field lacks a unified benchmark to evaluate the performance of various synthesis methods. To address this, we propose SketchRef, the first comprehensive multi-task evaluation benchmark for sketch synthesis. SketchRef fully leverages the shared characteristics between sketches and reference photos. It introduces two primary tasks: category prediction and structural consistency estimation, the latter being largely overlooked in previous studies. These tasks are further divided into five subtasks across four domains: animals, common things, human body, and faces. Recognizing the inherent trade-off between recognizability and simplicity in sketches, we are the first to quantify this balance by introducing a recognizability calculation method constrained by simplicity, mRS, ensuring fair and meaningful evaluations. To validate our approach, we collected 7,920 responses from art enthusiasts, confirming the effectiveness of our proposed evaluation metrics. Additionally, we evaluate the performance of existing sketch synthesis methods on our benchmark, highlighting their strengths and weaknesses. We hope this study establishes a standardized benchmark and offers valuable insights for advancing sketch synthesis algorithms. Xingyue Lin, Xingjian Hu, Shuai Peng, Liangcai Gao |
ICME | 3 |
| 2025 | Vote & Mix: Plug-and-Play Token Reduction for Efficient Vision TransformerabstractDespite the remarkable success of Vision Transformers (ViTs) in various visual tasks, they are often hindered by substantial computational cost. In this work, we introduce Vote&Mix (VoMix), a plug-and-play and parameter-free token reduction method, which can be readily applied to off-the-shelf ViT models without any training. VoMix tackles the computational redundancy of ViTs by identifying tokens with high homogeneity through a layer-wise token similarity voting mechanism. Subsequently, the selected tokens are mixed into the retained set, thereby preserving visual information. Experiments demonstrate VoMix significantly improves the speed-accuracy tradeoff of ViTs on both images and videos. Without any training, VoMix achieves a 2× increase in throughput of existing ViT-H on ImageNet-1K and a 2.4× increase in throughput of existing ViT-L on Kinetics-400 video dataset, with a mere 0.3% drop in top-1 accuracy. Shuai Peng, Di Fu, Baole Wei, Liangcai Gao, Zhi Tang 0001 |
ICME | 1 |
| 2025 | Uni-MuMER: Unified Multi-Task Fine-Tuning of Vision-Language Model for Handwritten Mathematical Expression RecognitionabstractHandwritten Mathematical Expression Recognition (HMER) remains a persistent challenge in Optical Character Recognition (OCR) due to the inherent freedom of symbol layouts and variability in handwriting styles.
Prior methods have faced performance bottlenecks by proposing isolated architectural modifications, making them difficult to integrate coherently into a unified framework.
Meanwhile, recent advances in pretrained vision-language models (VLMs) have demonstrated strong cross-task generalization, offering a promising foundation for developing unified solutions.
In this paper, we introduce Uni-MuMER, which fully fine-tunes a VLM for the HMER task without modifying its architecture, effectively injecting domain-specific knowledge into a generalist framework.
Our method integrates three data-driven tasks: Tree-Aware Chain-of-Thought (Tree-CoT) for structured spatial reasoning, Error-Driven Learning (EDL) for reducing confusion among visually similar
characters, and Symbol Counting (SC) for improving recognition consistency in long expressions.
Experiments on the CROHME and HME100K datasets show that Uni-MuMER achieves super state-of-the-art performance,
outperforming the best lightweight specialized model SSAN by 16.31\% and the top-performing VLM Gemini2.5-flash by 24.42\% under zero-shot setting.
Our datasets, models, and code are open-sourced at: https://github.com/BFlameSwift/Uni-MuMER Shuai Peng, Baole Wei, Liangcai Gao |
NeurIPS | 4 |
| 2025 | Ghost of the Navigator: Spoofing Attack Against Direction-of-Arrival EstimationabstractDirection of Arrival (DOA) estimation has been widely studied and applied in real-life systems. However, little attention has been given to the risk posed by attackers trying to hide real targets or create fake ones. This poses a serious security and safety concern, as such attacks could lead to system failures or accidents. In this paper, we proposeDOA spoofing, the first physical layer attack that exploits the semantic gap between DOA estimation outputs and the existing physical location of targets to spoof the perceived sources’ location. We consider two scenarios:i)the attacker wants to create ghost targets at desired locations while being undetected, andii)the attacker wants to change the DOA estimated at the receiver for a target victim. To evaluate the effectiveness of DOA spoofing, we test it against several widely used and well-established DOA estimation algorithms. Via extensive numerical simulations, we demonstrate that in both scenarios the attacker can spoof DOA estimates at the receiver while not being detected among the targets. To the best of our knowledge, this paper presents the first study to examine the impact of DOA attacks on the behavior of radar perception models. Therefore, although not providing a real-life testbed due to the challenges of implementing a non-commercial multi-antenna system, our contribution exposes novel threats thanks to the use of algorithms and models commonly used in real-life DOA estimation. Saiqin Xu, Alessandro Brighente, Baixiao Chen, Mauro Conti, Shuai Peng |
IEEE Internet Things J. | 5 |
| 2025 | Gridless 2D DOA estimation for sparse planar arrays via 2-level Toeplitz reconstruction
Shuai Peng, Baixiao Chen, Saiqin Xu |
Signal Process. | 1 |
| 2024 | Efficient Exploration on Worst-Case Delay Performance of Networked Industrial Control Systems via Network Calculus and Deep LearningabstractAs computer networking technologies such as Ethernet gain momentum in modern industrial control systems (ICSs), deterministic delay performance, which refers to the worst-case latencies provisioned by the network infrastructure, has become a critical property required by real-time control, automation, and operations. Although analytical methods that identify exact worst-case bounds on flow-specific communication delays have been proposed, there remain two hurdles discouraging further explorations of network-wide deterministic delay performance: (i) methods based on mixed integer linear programming (MILP) have high computational complexities; and (ii) existing methods are flow-specific. This paper proposes a deep-learning-assisted approach to understanding the deterministic delay performance of networked industrial control systems. Transforming flow-specific worst-case delay bounding into network-wide deterministic delay analysis, our approach facilitates the incorporation of application-specific hard-real-time performance constraints. By incorporating graph neural networks (GNNs), our approach significantly reduces the runtime costs and provides sufficiently tight approximations to the optimal MILP solutions. Through a combination of numerical analyses and simulations, we demon-strate that our approach enables the timely exploration of various worst-case scenarios, thereby paving the way for agile service provisioning in time-critical and/or delay-sensitive ICSs as well as model-based design of self-healing networked ICSs. Zhiqi Liang, Jiajie Zeng, Shuai Peng, Xiaoguang Ma |
SRDS | 3 |
| 2024 | CoLive: Edge-Assisted Clustered Learning Framework for Viewport Prediction in 360$^{\circ }$ Live StreamingabstractThe exceptionally high bandwidth requirement for delivering high-quality live 360$^\circ$video poses a significant challenge to current network capacity. Mitigating such bandwidth starvation necessitates accurate field-of-view (FoV) prediction to focus limited resources on the viewer's area of interest. However, FoV prediction for live 360$^\circ$streaming can be complex due to the time-sensitive nature of live content and the limited knowledge available for model training. Our paper introduces a novel framework,CoLive, for predicting the FoV in 360$^\circ$live streaming.CoLiveaccelerates FoV prediction by offloading model training from viewers to the edge and migrating saliency feature detection to the server side. Observations on user clustering of viewing behaviors further motivate us to propose a novel dynamic clustered learning algorithm. The algorithm dynamically groups users according to their model update gradients and enables them to train a shared model that better suits their viewing preferences. We conduct extensive experiments on the public 360$^\circ$video datasets and demonstrate thatCoLiveoutperforms state-of-the-art solutions in terms of prediction performance and bandwidth savings. Xingyan Chen, Shuai Peng, Yu Zhao 0019, Mingwei Xu 0001, Changqiao Xu |
IEEE Trans. Multim. | 4 |
| 2023 | Spherical Convolution-based Saliency Detection for FoV Prediction in 360-degree Video StreamingabstractField of view (FoV) prediction is a crucial issue in 360° video streaming, which is the basis for selectively transmitting panoramic videos to reduce bandwidth. The saliency feature is a very important part of FoV prediction. The saliency area identifies a user’s region of interest (RoI) and reflects the user’s viewing behavior preference. The regular convolutional neural network (CNN) cannot effectively extract the spatial representation of panoramic video content because significant geometric distortion will be introduced after panoramic video projection, especially in polar regions. In this paper, we propose a depth neural network model based on spherical convolution, which can learn the spatial features of the 360° videos by encoding the distortion invariance into the architecture of CNNs. A series of experiments on the public 360° video saliency dataset show the proposed model outperforms the existing saliency models. Finally, we embed the proposed saliency network into a popular FoV prediction framework and propose a complete FoV prediction framework for 360° video streaming. Shuai Peng, Jialu Hu, Changqiao Xu |
IWCMC | 1 |
| 2022 | CoLive: An Edge-Assisted Online Learning Framework for Viewport Prediction in 360° Live StreamingabstractThe ever-increasing demand for bandwidth resources when delivering premium quality 360° video challenges the current network capacity. To alleviate such bandwidth pressure, it is imperative to predict the viewport via observing the content visual feature and historical viewing behaviors, which thereby allows the system to concentrate the limited resource on viewer's region of interest in 360° content. However, enabling accurate viewport prediction for 360° live streaming is non-trivial given the time-sensitive of live content and shortage of pre-knowledge on the visual features and viewing behaviors. In this paper, we propose CoLive, an edge-assisted online viewport prediction framework. CoLive incorporates edge computing to offload the prediction model training from viewers and migrates the saliency feature detection to the server side for reducing the processing delay. Viewers can also collaboratively train a central predicting model via sharing their loss gradients. This central model, together with the saliency feature detection, further prompts accuracy prediction and learning acceleration, especially for new incoming viewers. A series of experiments on the public 360° video dataset show how our solution achieves better performance compared with state-of-the-art solutions. Shuai Peng, Xingyan Chen, Yu Zhao 0019, Mingwei Xu 0001, Changqiao Xu |
ICME | 2 |
| 2022 | Compute Like Humans: Interpretable Step-by-step Symbolic Computation with Deep Neural NetworkabstractNeural network capability in symbolic computation has emerged in much recent work. However, symbolic computation is always treated as an end-to-end blackbox prediction task, where human-like symbolic deductive logic is missing. In this paper, we argue that any complex symbolic computation can be broken down to a sequence of finite Fundamental Computation Transformations (FCT), which are grounded as certain mathematical expression computation transformations. The entire computation sequence represents a full human understandable symbolic deduction process. Instead of studying on different end-to-end neural network applications, this paper focuses on approximating FCT which further build up symbolic deductive logic. To better mimic symbolic computations with math expression transformations, we propose a novel tree representation learning architecture GATE (Graph Aggregation Transformer Encoder) for math expressions. We generate a large-scale math expression transformation dataset for training purpose and collect a real-world dataset for validation. Experiments demonstrate the feasibility of producing step-by-step human-like symbolic deduction sequences with the proposed approach, which outperforms other neural network approaches and heuristic approaches. Shuai Peng, Di Fu, Yijun Liang, Gu Xu, Liangcai Gao, Zhi Tang 0001 |
KDD | 1 |
| 2021 | Image to LaTeX with Graph Neural Network for Mathematical Formula Recognition
Shuai Peng, Liangcai Gao, Zhi Tang 0001 |
ICDAR (2) | 1 |
| 2021 | Handwritten Mathematical Expression Recognition with Bidirectionally Trained Transformer
Wenqi Zhao, Liangcai Gao, Zuoyu Yan, Shuai Peng, Ziyin Zhang |
ICDAR (2) | 4 |
| 2021 | YOLO-face: a real-time face detector
Weijun Chen 0006, Hongbo Huang, Shuai Peng, Changsheng Zhou, Cuiping Zhang |
Vis. Comput. | 3 |
| 2020 | More trainable inception-ResNet for face recognition
Shuai Peng, Hongbo Huang, Weijun Chen 0006, Weiwei Fang |
Neurocomputing | 1 |
| 2018 | XPA: An Open Soruce IDE for XACML Policies (S)abstractThis paper presents XPA (XACML Policy Analyzer), an open source IDE (Integrated Development Environment) for testing, debugging, and mutating XACML 3.0 policies.XACML is an OASIS standard for specifying attributebased access control policies.XPA provides a variety of new techniques for generating test cases from policies, localizing bugs in faulty policies, and repairing faulty policy elements.XPA has been applied to numerous XACML policies from the literature and real-world applications.These policies have been used to quantitatively evaluate the effectiveness of various testing and debugging methods.For system developers and administrators, XPA is a practical IDE for developing dependable XACML policies.For access control researchers, XPA offers a versatile toolkit for studying and evaluating new testing, debugging, and verification techniques. Roshan Shrestha, Shuai Peng, Turner Lehmbecker, Dianxiang Xu |
SEKE | 2 |
| 2016 | Towards automatic repair of access control policiesabstractAccess control policies written in the XACML standard language tend to be complex due to the great variety of attribute types and operations for fine-grained access control. The complexity not only increases the likelihood of having authorization faults in access control policies, but also makes it challenging to find and fix these faults. This paper presents an approach for automating the process of debugging XACML policies. It consists of two main techniques: fault localization and mutation-based policy repair. Fault localization aims to find the most suspicious policy elements according to the correlation between the execution information of policy elements and the test execution results. Mutation-based policy repair aims to modify the suspicious policy elements by using well-defined mutation operators. Our empirical studies have used a large number of faulty policies with one or two seeded faults. Our approach was able to repair all of them automatically. We have also compared several scoring methods for ranking suspicious policy elements. The results show that Naish2 and CBI-Inc are very efficient for automatic repair. Dianxiang Xu, Shuai Peng |
PST | 2 |
| 2016 | Automated Fault Localization of XACML PoliciesabstractAccess control policies in distributed systems, particularly implemented in the XACML standard language, are increasingly complex. Faults may exist in complex policies for various reasons such as misunderstanding of the access control requirements, omissions, and coding errors. These faults, if not removed before deployment, may lead to unauthorized accesses or denial of service. Manual localization of these faults, however, can be a challenging task. Inspired by spectrum-based fault localization for software debugging, this paper presents an approach for automatically localizing the fault(s) in a given XACML policy by exploring test coverage information of the policy elements. We investigate two test coverage criteria (i.e., reachability and firing) of policy elements and 14 scoring methods for ranking policy elements to determine the fault location(s). To evaluate the fault localization methods, we have used real-world policy files with different levels of complexity and a large number of policy mutants with one or two seeded faults. The experiment results show that the firing-based Naish2 and CBI-Inc methods are effective in fault localization of XACML policies. Dianxiang Xu, Shuai Peng |
SACMAT | 3 |
| 2015 | Energy neutral clustering for energy harvesting wireless sensors networks
Shuai Peng, Ting Wang 0017, Chor Ping Low |
Ad Hoc Networks | 1 |
| 2015 | Energy neutral directed diffusion for energy harvesting wireless sensor networks
Shuai Peng, Chor Ping Low |
Comput. Commun. | 1 |
| 2014 | Prediction free energy neutral power management for energy harvesting wireless sensor nodes
Shuai Peng, Chor Ping Low |
Ad Hoc Networks | 1 |
| 2013 | Energy Neutral Routing for energy harvesting Wireless Sensor NetworksabstractIn recent years, energy harvesting techniques have made it possible to overcome the energy resource limitation suffered by traditional Wireless Sensor Networks. In particular, it is now possible for the network to operate perpetually if sufficient energy is harvested to replenish the energy that is being consumed. Hence the need to prolong network lifetime is no longer the primary objective for the design of network routing protocols. Instead, the aim now is to maintain the network in a Energy Neutral state under which a certain performance level can be maintained perpetually. In this paper we propose a routing protocol, which we refer to as the Energy Neutral Routing (ENR) Protocol, that is able to achieve this objective. ENR also provides fair distribution of traffic loads and it is able to guarantee a linear bound on the total number of communication control messages exchanged, which in turn improves the overall distinct packet delivery ratio. Empirical studies are carried out to evaluate and compare the performance of ENR against other routing protocols. Shuai Peng, Chor Ping Low |
WCNC | 1 |
| 2012 | Throughput optimal energy neutral management for energy harvesting wireless sensor networksabstractEnergy harvesting techniques have enabled the provisioning of alternative energy sources beside the conventional one which is typically provided by the batteries. Using such harvesting techniques together with proper energy management mechanisms, a Energy Neutral state can be achieved so that desired performance level can be supported perpetually. Existing energy neutral management mechanisms rely on the prediction of the amount of energy that can be harvested in the future. Such mechanisms suffer from the time consuming prediction processes and great fluctuations in the actual implementations. Hence, in this paper we propose real time adaptive energy management policies that enable energy neutral management based solely on observed information in the past instead of predicting the amount of energy that can be harvested in the future. We will show that these policies are throughput optimal in the sense that we are able to ensure that the data queue of each sensor node remain stable when the largest possible data rate is applied under different channel capacity scenarios. We also propose an energy harvest-store(use) method to reduce the energy losses caused by storing energy in non ideal energy buffers. Simulation studies show that our proposed throughput optimal policies and harvest-store(use) method are indeed able to improve the overall throughput as compared to existing schemes. Shuai Peng, Chor Ping Low |
WCNC | 1 |