VLDB 2026 Research / reviewers in the wild / expert
Xiaohui Peng 0002
dblp:06/8456-2
· DBLP profile ↗
23ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0003-2706-6379ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 2 first-author · 11 since 2021Computer networks · 4 · 3 since 2021Artificial intelligence and machine learning · 2Software engineering, systems software and programming languages · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enabling High-Utilization and Low-Contention FaaS: A Request-Level Resource Provisioning ApproachabstractFunction-as-a-Service offers cost efficiency but often suffers from resource underutilization. This underutilization stems from the instance-level resource provisioning pattern, an issue that existing optimizations have failed to resolve fundamentally. The core problem is that static coarse-grained instance-level resource allocation cannot match the millisecond-level burstiness of dynamic requests. Consequently, it is difficult for current systems to achieve high resource utilization while maintaining high quality of service (QoS) guarantees. To address the problem, this paper advocates a shift to request-level resource provisioning, which redefines the individual request as the atomic unit for scheduling and resource management. We implement this approach in RRP, a scalable FaaS platform that enables efficient per-request resource allocation and release. RRP unifies instance placement and request routing with low-overhead, millisecond-level global visibility. Our evaluation shows that RRP significantly outperforms state-of-the-art instance-level platforms and algorithms. By matching resources to each request’s needs and isolating them from contention, RRP achieves low latency and high utilization. Specifically, on real-world Azure traces, RRP achieves speedups of 1.33 × –30.15 × for average end-to-end latency and 1.37 × –61.46 × for P99 latency, and raises CPU utilization from 44.80%–56.32% to 72.49% under bursty loads. Runfu Li, Zishu Yu, Yifan Wang 0005, Xiaohui Peng 0002, Ninghui Sun, Zhiwei Xu 0002 |
HPDC | 4 |
| 2026 | Group-TopK: Optimizing Distributed Training on Edge Devices via Communication Compression
Yifan Wang 0005, Xiaohui Peng 0002, Haohao Ma, Hui Sun 0002, Deke Guo, Boyu Diao |
HPDC | 2 |
| 2026 | LASS: Reducing Cold Startup Latency in Serverless Through Loaded Library SharingabstractIn serverless scenario, function invocation runs in an individual container. Lightweight container technology has significantly reduced the startup latency of container. The library loading process now becomes a critical performance bottleneck of serverless function cold startup. The state-of-the-art approaches leverage the process fork operation to reduce the cold startup latency in serverless computing by reusing the loaded libraries. However, the fork operation can only share libraries between parent process and forked process. For security, the libraries loaded by the parent process should be a subset of those required by the forked process, which limits opportunities to eliminate library loading overhead. To address this problem, we propose theLASSsystem, which enables multiple processes to share initialized libraries in a composable and efficient manner.LASSallows a process to securely reuse libraries loaded by multiple processes, thereby reducing library loading latency to the millisecond level. Compared to the state-of-the-art approaches,LASScan improve average library loading speed by more than 10.3×, and reduce 99thpercentile end-to-end latency by 34%–57%. Zishu Yu, Runfu Li, Yifan Wang 0005, Xiaohui Peng 0002, Zhiwei Xu 0002 |
IEEE Trans. Computers | 5 |
| 2026 | Breaking Cloud Dependencies: A Distributed Ledger Approach to IoT Device Usufruct Management
Xiaohui Peng 0002, Yifan Wang 0005, Deke Guo |
IEEE Trans. Mob. Comput. | 2 |
| 2026 | SEPP-FLBC: A Secure and Efficient Privacy Protection Scheme Using Federate Learning and Blockchain for Edge-End-Cloud DevicesabstractThe convergence of federated learning (FL) and blockchain in edge-end-cloud systems offers promising opportunities for privacy-preserving collaborative intelligence. However, existing blockchain-enhanced FL (BFL) approaches remain vulnerable to malicious participants and lack robust protection for model updates. To address these issues, we propose SEPP-FLBC, a Secure and Efficient Privacy Protection framework based on Federated Learning and Blockchain Committees. SEPP-FLBC introduces a novel blockchain committee consensus mechanism to validate model updates and defend against unreliable nodes. It further employs a refined multi-party communication paradigm to facilitate indirect and secure data interactions, reducing the risk of information leakage. Additionally, differential privacy noise is applied to model updates to enhance resistance to inference attacks. A formal convergence analysis is conducted to ensure model stability and minimize overhead. Extensive experiments on benchmark datasets demonstrate that SEPP-FLBC achieves superior accuracy while maintaining strong privacy guarantees and communication efficiency, outperforming state-of-the-art BFL methods in both security and performance. Libo Feng, Junwei Guo, Fake Fang, Zhenli He, Yimin Yu, Shaowen Yao 0001, Xiaohui Peng 0002 |
IEEE Trans. Serv. Comput. | 7 |
| 2025 | Gensor: A Graph-Based Construction Tensor Compilation Method for Deep LearningabstractHigh-performance deep learning depends on efficient tensor programs. In recent years, automatic tensor program optimization, also known as tensor compilation, has emerged as the primary approach to generating efficient tensor programs. However, how to generate kernels with higher performance in a shorter time is still the key challenge. In this paper, we present Gensor, a graph-based construction tensor compilation method for deep learning, to further improve the performance of construction tensor compilation. Unlike existing tree-based methods, Gensor abstracts construction space into a graph structure. Gensor then explores the construction space with Markov analysis. Gensor takes tensor programs as states and models scheduling primitives as transition actions between these states. Therefore, the process of tensor program construction optimization is abstracted as a graph traversal process. This approach expands the optimization space, improving operator performance while ensuring rapid optimization. Extensive experiments with typical operators demonstrate that Gensor significantly outperforms the state-of-the-art methods on GPUs for both cloud servers and edge devices. As a result, Gensor can generate operator kernels in seconds, with performance increasing by 18 % on average, reaching a maximum of 30 %. It also achieves high speedup for end-to-end models like ResNet50 and GPT-2, with an average acceleration of 20 %. Hangda Liu, Boyu Diao, Xiaohui Peng 0002, Yongjun Xu 0001 |
IPDPS | 5 |
| 2025 | Tide: A Distributed Runtime Management Framework for Things-Edge-Cloud Computing ContinuumabstractThe increasing number of connected IoT devices produces massive amounts of sensed data at the network edge. A new computing paradigm, called Things-Edge-Cloud (TEC) collaboration, has been proposed to meet real-time and high-throughput requirements. Most of the existing work focuses on workload scheduling across the computing nodes in TEC with ad-hoc implementations using runtime and management frameworks designed for the cloud. In this paper, we propose RSEP to model the entities and their relationship in the TEC computing continuum (TEC3). We then design and implement Tide—a distributed runtime management framework for TEC3based on the RSEP model, which enables elastic resource allocation and seamless computation offloading. It employs runtime environment isolation and physical resource binding to enforce strong isolation without incurring performance penalties. To decouple runtime and framework, Tide provides a set of portable application interfaces that allow the management of variety runtimes. We implement Tide from scratch and compare its latency and throughput with KubeEdge, Ray, and bare-metal implementations. Experimental results show that Tide improves throughput by 2x and reduces average latency, 95th percentile latency, and latency standard deviation by${4 5. 8 2 \%, 4 8. 3 6 \%}$, and${2 8. 9 6 \%}$, respectively. Specifically, Tide achieves${8 7. 3 \%}$of the ideal goodput, exceeding other platforms more than 10x. Xiaohui Peng 0002, Wenkai Yan, Yifan Wang 0005, Shoujian Zheng, Zhiwei Xu 0002 |
IPDPS | 1 |
| 2025 | A Case Study on Benchmarking Distributed AI SystemsabstractThe rapid growth of artificial intelligence (AI), particularly in computer vision (CV), necessitates distributed computing for efficient model training. Existing benchmarks often lack adaptability to emerging scenarios or focus on limited applications. To address these gaps, this paper studies a case on a comprehensive benchmark suite for distributed AI training systems. We classifies AI tasks into four categories, LargeScale, Moderate Complexity, High Load, and High-Performance, based on single-load computation and load concurrency, with representative models evaluated on Ray and DeepSpeed across diverse hardware. The experiments reveal fragmented framework performance. DeepSpeed excels in stability and efficiency for Large-Scale and Moderate Complexity tasks, leveraging advanced memory optimization. Ray outperforms in High Load and High-Performance tasks due to its dynamic resource scheduling but shows greater variability. These results highlight the need for task-specific framework selection tailored to hardware and performance requirements. We provides valuable insights for optimizing distributed AI training and bridges limitations in current benchmarks. Future work aims to expand task categories and framework support to align with the evolving demands of distributed AI systems. Jianwei Gao, Xiaohui Peng 0002, Jiamu Liu, Yifan Wang 0005, Deke Guo |
IWQoS | 2 |
| 2025 | upTSA: A DIMM-Based Near Data Processing Accelerator for Time Series Analysis
Shunchen Shi, Fan Yang 0096, Qijia Yang, Xiaohui Peng 0002, Xueqi Li 0001, Ninghui Sun |
NPC (1) | 4 |
| 2025 | EdgeInferFlow: A Distributed Inference Acceleration Method for Deep Learning Chained Structure Models for Edge Devices
Hanfeng Zhai, Yifan Wang 0005, Xiaohui Peng 0002, Xueqi Li 0001 |
NPC (1) | 3 |
| 2025 | DSparse: A Distributed Training Method for Edge Clusters Based on Sparse Update
Xiaohui Peng 0002, Yixuan Sun, Zhenghui Zhang, Yifan Wang 0005 |
J. Comput. Sci. Technol. | 1 |
| 2024 | Snapipeline: Accelerating Snapshot Startup for FaaS ContainersabstractDue to the frequent starts and stops of numerous services in FaaS (Function as a Service), reducing cold start overhead is a core issue in improving the performance of container-based FaaS services. Snapshot and restore-based mechanisms effectively reduce the cold start time of containers by transforming container initialization overhead into restoration overhead. Consequently, this mechanism has become a research hotspot in accelerating the cold start of FaaS containers. Researchers introduce snapshot compression and decompress the snapshots to reduce the storage cost before starting instances. However, existing works have the following shortcomings: (1) File-mapped memory pages are not processed during snapshot compression, resulting in a significant amount of redundant data in memory; (2) The serial execution of snapshot decompression and instance restoration leads to high instance startup latency. Yuqiao Lan, Xiaohui Peng 0002, Yifan Wang 0005 |
SoCC | 2 |
| 2024 | CMS: A Computility Resource Status Management and Storage Framework
Xiaohui Peng 0002, Kuo Chang, Yifan Wang 0005 |
NPC (1) | 1 |
| 2024 | Hawk: An Efficient NALM System for Accurate Low-Power Appliance RecognitionabstractNon-intrusive Appliance Load Monitoring (NALM) aims to recognize individual appliance usage from the main meter without indoor sensors. However, existing systems struggle to balance dataset construction efficiency and event/state recognition accuracy, especially for low-power appliance recognition. This paper introduces Hawk, an efficient and accurate NALM system that operates in two stages: dataset construction and event recognition. In the data construction stage, we efficiently collect a balanced and diverse dataset, HawkDATA, based on balanced Gray code and enable automatic data annotations via a sampling synchronization strategy called shared perceptible time. During the event recognition stage, our algorithm pipeline integrates steady-state differential pre-processing and voting-based post-processing for accurate event recognition from the aggregate current. Experimental results show that HawkDATA takes only 1/71.5 of the collection time to collect 6.34x more appliance state combinations than the baseline. In HawkDATA and a widely used dataset, Hawk achieves an average F1 score of 93.94% for state recognition and 97.07% for event recognition, which is a 47.98% and 11.57% increase over SOTA algorithms. Furthermore, selected appliance subsets and the model trained from HawkDATA are deployed in two real-world scenarios with many unknown background appliances. The average F1 scores of event recognition are 96.02% and 94.76%. Hawk's source code and HawkDATA are accessible at https://github.com/WZiJ/SenSys24-Hawk. Xingzhou Zhang, Yifan Wang 0005, Xiaohui Peng 0002, Zhiwei Xu 0002 |
SenSys | 4 |
| 2022 | Flet-Edge: A Full Life-cycle Evaluation Tool for deep learning framework on the EdgeabstractDeep learning frameworks, such as TensorFlow, PyTorch, MXNet, and Paddle Paddle are widely used and studied by industry. At the same time, AIoT (Artificial Intelligence and Internet of Things) and edge computing have provided more deep learning scenarios on the edge. In order to develop and deploy AIoT applications, we need to evaluate deep learning frameworks from ease-of-use and performance. To describe the full life-cycle performance of deep learning frameworks on the edge, this paper proposed a metric set, PDR, includes three comprehensive submetrics: Programming complexity, Deployment complexity, and Runtime performance. Based on the PDR, this paper designed and implemented a full life-cycle evaluation tool, Flet-Edge, which can automatically collect and present the PDR’s metrics, visually. Finally, to verify the availability of the Flet-Edge, this paper built a heterogeneous edge device cluster and carried out three case studies. With only one configuration file as input, the FletEdge will collect the twelve metrics of training or inference tasks and output them in text or chart. By observing the hierarchical roofline diagram provided by the Flet-Edge, this paper shows that the Flet-edge has the ability to optimize software and hardware of deep learning. Xiaoyang Jiang, Xingzhou Zhang, Xiaohui Peng 0002 |
ICPADS | 3 |
| 2022 | Hebo System: Trusted Copyright Authorization in Computer NetworksabstractThe DRM (Digital Rights Management) systems protect owners’ copyrights by controlling consumers’ access to digital works. However, they fail to provide authorization evidence if customers use digital works on other platforms outside the DRM systems. There is no such evidence that can trustworthily be disseminated in computer networks, which results in many copyright lawsuits. To address the problem, we propose the Copyrights Authorization Model (CAM) and the Hebo system to ensure the consensus on copyright authorization in computer networks. The CAM proves that participants agree on copyright authorizations if they are traceable, integrated, and non-repudiated. Based on the CAM, we design the ledger of trusted authorization forest that keeps the three properties and independent zones that maintain the ledger. The Hebo system is composed of these zones. It can provide authorization evidence for consumers to avoid copyright disputes. Besides, it has the advantages of a beneficial locality. The TPS (transactions per second) can increase with the number of independent zones, nodes can flexibly choose ledgers according to their capability, the system reduces redundant storage, and zones allow asynchrony. Finally, we evaluate the system on five different platforms, and the average time costs of operations are less than 60 milliseconds. The TPS of a single node depends on the configuration of the hardware, which is an average of 138, 19, 24, 21, and 98 in Server, MacBook, J-Nano, Pi4B, and J-TX2, respectively. Yixuan Sun, Xiaohui Peng 0002 |
ICPADS | 4 |
| 2020 | Fengyi: Trusted Data Sharing in VANETs with BlockchainabstractSuperior to traditional vehicles, intelligent vehicles (IV) can share data in Vehicular Ad-Hoc Networks (VANETs) to provide a more comfortable and safer driving experience, based on the assumption that the data sharing is accountable and reliable. However, trusted data sharing in VANETs is always a paramount concern. We define that the trusted data sharing includes the three properties, namely data sharing accountability, privacy preservation, and transmission confidentiality. To address the problem, we propose a comprehensive solution including the trusted ledger model (TLM) and implement the Fengyi system to verify it. The TLM is a model that ensures the consistency of multiple data resources in a low trust distributed computing environment. Then, distributed Fengyi ledgers based on the TLM are proposed to keep data sharing accountable and private in VANETs. The Fengyi system is designed and implemented to provide authentication and encrypted communication services with the ledgers. Finally, we deploy the Fengyi system on three different platforms, checking the effectiveness and efficiency of the Fengyi system. The results show that the system can ensure trusted data sharing in VANETs, and the time cost for the verification of data sharing on-road is average 253.33μs, 38% lower than that in recent research. Yifan Wang 0005, Xiaohui Peng 0002 |
PRDC | 4 |
| 2019 | T-REST: An Open-Enabled Architectural Style for the Internet of ThingsabstractComputing offloading is a key challenge of new rising computing paradigms of the Internet of Things (IoT) like edge computing, which shifts computations to data sources as near as possible to gain the benefits, such as low latency and energy efficiency. However, the fragmentation problem of IoT devices results in a heterogeneous and disordered ecosystem, hindering the interoperating demands of computing offloading. What we need is an open-enabled ecosystem which allows third-party developers to create and update functions of deployed devices dynamically. We propose Things-representational state transfer (T-REST) which is an extension of the representational state transfer (REST) architectural style to address this problem. It integrates contents and computations together to inherit the uniform interface principle from REST. Three architectural constraints are added to REST: 1) reusable remote evaluation; 2) dynamic time series representation; and 3) computational hypertext. A novel event triggering mechanism is designed to decouple the tight coupling of front-end content accesses and back-end computations for resources. A reference prototype, named T-REST engine, is implemented to verify the proposed architecture with the open-enabled style, distributed semantics, and computing offloading features. Discussions show that T-REST preserves the benefits of REST. In addition, it achieves extremely lightweight footprints and can perform computing offloading through open-enabled architectures. Zhiwei Xu 0002, Lu Chao, Xiaohui Peng 0002 |
IEEE Internet Things J. | 3 |
| 2019 | Ecosystem of Things: Hardware, Software, and ArchitectureabstractEdge computing is a continuum that includes the computing resources from cloud to things. Ecosystem of things (EoT) is a subsystem of the ecosystem of edge computing, which potentially contains trillions of devices of things and directly interacts with the physical world. This paper surveys the state of the art of EoT by focusing on the computing infrastructure aspect with a forward-looking perspective. We point out a trend of smart edge computing with four types of smartness and intelligence. We address three fundamental questions. 1) What capabilities and how much energy efficiency are the hardware providing? What is the future growth potential? 2) What abstractions are provided by the system software? Are they adequate to support smart edge computing? 3) What ecosystem architectures have been proposed for the coordination of things, the edge, and the cloud? Are they meeting the needs to encourage innovation but avoid unnecessary ecosystem fragmentation? We examine advances from both industry and academia, including research results, visions, and project concepts. We also point out future research directions. Lu Chao, Xiaohui Peng 0002, Zhiwei Xu 0002, Lei Zhang 0008 |
Proc. IEEE | 2 |
| 2019 | Deep learning for sensor-based activity recognition: A survey
Jindong Wang 0001, Yiqiang Chen 0001, Shuji Hao, Xiaohui Peng 0002, Lisha Hu |
Pattern Recognit. Lett. | 4 |
| 2019 | A Novel Feature Incremental Learning Method for Sensor-Based Activity RecognitionabstractRecognizing activities of daily living is an important research topic for health monitoring and elderly care. However, most existing activity recognition models only work with static and pre-defined sensor configurations. Enabling an existing activity recognition model to adapt to the emergence of new sensors in a dynamic environment is a significant challenge. In this paper, we propose a novel feature incremental learning method, namely the Feature Incremental Random Forest (FIRF), to improve the performance of an existing model with a small amount of data on newly appeared features. It consists of two important components - 1) a mutual information based diversity generation strategy (MIDGS) and 2) a feature incremental tree growing mechanism (FITGM). MIDGS enhances the internal diversity of random forests, while FITGM improves the accuracy of individual decision trees. To evaluate the performance of FIRF, we conduct extensive experiments on three well-known public datasets for activity recognition. Experimental results demonstrate that FIRF is significantly more accurate and efficient compared with other state-of-the-art methods. It has the potential to allow the dynamic exploitation of new sensors in changing environments. Chunyu Hu 0001, Yiqiang Chen 0001, Xiaohui Peng 0002, Han Yu 0001, Chenlong Gao, Lisha Hu |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2018 | Stratified Transfer Learning for Cross-domain Activity RecognitionabstractIn activity recognition, it is often expensive and time-consuming to acquire sufficient activity labels. To solve this problem, transfer learning leverages the labeled samples from the source domain to annotate the target domain which has few or none labels. Existing approaches typically consider learning a global domain shift while ignoring the intra-affinity between classes, which will hinder the performance of the algorithms. In this paper, we propose a novel and general cross-domain learning framework that can exploit the intra-affinity of classes to perform intra-class knowledge transfer. The proposed framework, referred to as Stratified Transfer Learning (STL), can dramatically improve the classification accuracy for cross-domain activity recognition. Specifically, STL first obtains pseudo labels for the target domain via majority voting technique. Then, it performs intra-class knowledge transfer iteratively to transform both domains into the same subspaces. Finally, the labels of target domain are obtained via the second annotation. To evaluate the performance of STL, we conduct comprehensive experiments on three large public activity recognition datasets (i.e., OPPORTUNITY, PAMAP2, and UCI DSADS), which demonstrates that STL significantly outperforms other state-of-the-art methods w.r.t. classification accuracy (improvement of 7.68%). Furthermore, we extensively investigate the performance of STL across different degrees of similarities and activity levels between domains. And we also discuss the potential of STL in other pervasive computing applications to provide empirical experience for future research. Jindong Wang 0001, Yiqiang Chen 0001, Lisha Hu, Xiaohui Peng 0002, Philip S. Yu |
PerCom | 4 |
| 2018 | A novel random forests based class incremental learning method for activity recognition
Chunyu Hu 0001, Yiqiang Chen 0001, Lisha Hu, Xiaohui Peng 0002 |
Pattern Recognit. | 4 |