VLDB 2026 Research / reviewers in the wild / expert
Xiaohui Wei 0002
dblp:25/4202-2
· DBLP profile ↗
84ranked-venue papers
38as first author
53since 2021 · last 2026
0000-0001-5597-3625ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 45 · 21 first-author · 32 since 2021Computer networks · 15 · 7 first-author · 6 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 8 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Co-designing architecture and feature guidance for efficient video understanding
Xingwang Wang 0003, Xiaohui Wei 0002, Kun Yang 0001 |
Comput. Vis. Image Underst. | 3 |
| 2026 | Deeply understanding features to achieve efficient remote sensing image classification
Xingwang Wang 0003, Xiaohui Wei 0002, Yafeng Sun, Kun Yang 0001 |
Expert Syst. Appl. | 3 |
| 2026 | PromptGuard: Safeguarding large vision-language models via adversarial prompt tuning
Changbao Zhou, Hengshan Yue, Ming Yan 0007, Xiaohui Wei 0002 |
Knowl. Based Syst. | 4 |
| 2026 | Feature-based optimization enables 2D CNNs for efficient spatio-temporal perception
Xingwang Wang 0003, Xiaohui Wei 0002, Yafeng Sun, Kun Yang 0001 |
Pattern Recognit. | 3 |
| 2026 | Sift: Channel-Wise Historical Embedding for High Efficiency Distributed Graph Neural Network Training with Accuracy GuaranteeabstractDistributed Graph Neural Network (DGNN) is a powerful tool in large-scale graph representation learning. However, high data-transfer overhead among workers in a DGNN training job confines its scalability and thus the overall performance. Vertex-wise historical embedding methods have demonstrated high potential to alleviate the problems, but still suffer from severe accuracy loss and limited performance scalability, which has been attributed to the information loss of critical channels in historical vertices and redundant information in local channels. This article explores the optimization of channel level and construct a quantitative accuracy model for channel-wise historical embedding. We propose Sift, a novel DGNN training framework, supporting channel-wise partial historical embedding with accuracy guarantee. Sift has three components: a historical embedding evaluator with channel-wise quantitative accuracy model, a sawtooth-like matrix rearrangement for accelerating message passing, and a hybrid parallel framework for overlapping communication overhead. Comprehensive experimental results show that Sift achieves near-linear parallel convergence speedup, outperforming the state-of-the-art baselines by up to 72% in total training performance and up to 21% in convergence speed. Zhewen Xu, Hongliang Li 0003, Junze Han, Hengshan Yue, Hairui Zhao 0002, Dongyuan Tian, Zijian Li 0007, Xiaohui Wei 0002 |
ACM Trans. Archit. Code Optim. | 8 |
| 2026 | Voltage Channel: Exploiting GPU Voltage Noise for Covert and Side Channel Attacks
Zhanyuntian Li, Jingweijia Tan, Kaige Yan, Haixiao Xu, Xiaohui Wei 0002 |
IEEE Trans. Computers | 5 |
| 2026 | FR2eRAM: A Fault-Resilient Graph Processing Paradigm in Realistic ReRAMsabstractWith the explosive growth of modern graph data, ReRAM-based graph processing paradigms are emerging as promising solutions to the “memory wall” bottleneck. However, existing paradigms often lean on overly idealized ReRAM architectures, overlooking the critical influence of various hardware faults due to the analog nature and immature fabrication processes of realistic ReRAMs. These hardware faults can lead to convergence anomalies or unacceptable output deviations (i.e., Severe Errors), undermining the reliability of ReRAM-based graph processing. While some studies enhance the reliability of ReRAM-based computation through fault-aware remapping or robust algorithm design, the unique graph execution characteristics make these efforts challenging to migrate effectively. In this work, we first develop ReGFI, a microarchitecture-level Fault Injection framework for ReRAM-based Graph processing. Unlike traditional fault injection methods that only introduce random algorithm-level faults, ReGFI precisely maps hardware faults into the microarchitectural graph execution flow to effectively characterize their effects on the execution correctness. Based on ReGFI, we propose FReRAM, a Fault-Resilient graph processing paradigm in realistic ReRAMs. Firstly, observing the fault robustness of graph vertices compared to graph edges, we reverse-map the vertex values to the non-ideal crossbar while using the edge values as inputs, for proactive SE avoidance. Then, leveraging the cell idleness in ReRAM crossbars and bit-wise reliability discrepancies of graph data, we recycle idle crossbar cells and squeeze out approximable Least Significant Bits to robust-encode the fault-sensitive bits, for further SE elimination. Experimental results exhibit that FReRAM achieves 88.84% SE reduction while incurring negligible overhead for fault-resilient ReRAM-based graph processing. Furthermore, we evaluate the effectiveness and performance of FReRAM under different hardware configurations, graph datasets, and data formats. Hengshan Yue, Nan Jiang 0013, Zongdian Li, Jiaguo Deng, Yu Huang 0013, Meikang Qiu, Xiaohui Wei 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2025 | GraphFI: An Efficient Fault Injection Framework for Graph Processing on GPGPUsabstractAs graph tasks become pervasive in real-time and safetycritical domains (e.g., financial fraud detection and electrical power systems), it is also essential to guarantee their reliable execution beyond pursuing extraordinary performance. However, due to the neglect of consideration for graph-specific execution paradigm, existing Fault Injection (FI) reliability analysis methods typically incur inaccurate system error resilience characterization, making it challenging to provide helpful guidance for efficient and reliable graph processing paradigm design. This paper proposes GraphFI, an efficient Graph Fault Injection framework on the universal parallel tasks acceleration platform (i.e., GPGPUs). Our key insight is progressively excavating the graph-specific error propagation and effect mechanisms, thereby avoiding blind FI trials. Firstly, observing that iterations with similar active vertex set exhibit similar error behavior, we propose iteration-driven GraphFI (ID-GraphFI) to solely select representative iterations for fast error resilience profile assessment. Secondly, by detecting resilience-similarity communities in graph topology, we propose topology-driven GraphFI (TD-GraphFI) that only selects representative vertices for community overall reliability evaluation. Thirdly, by exploring the graph-specific fault monotonic property, we propose the monotonicity-driven GraphFI (MD-GraphFI) to granularly draw system severe error boundaries for predictable/unnecessary fault injection avoidance. Merging them all, GraphFI can reduce system fault site space by up to two orders of magnitude, which achieves $2.1 \sim 15.2 \times$ speedup compared to SOTA methods while providing better reliability assessment accuracy. Nan Jiang 0013, Hengshan Yue, Jingweijia Tan, Mengting Zhou, Wenda Wei, Meikang Qiu, Xiaohui Wei 0002 |
DAC | 9 |
| 2025 | Accelerating Graph Sampling in GNN Systems Through Accuracy-Aware Data Reuse
Nan Jiang 0013, Hengshan Yue, Xiaohui Wei 0002 |
ICA3PP (3) | 5 |
| 2025 | A Mapping Strategy Optimization Framework for Systolic Array Accelerators
Hengshan Yue, Haixiao Xu, Xiaohui Wei 0002 |
ICA3PP (1) | 6 |
| 2025 | GraphFT: A Lightweight Fault-tolerant Framework for Iterative Graph Processing
Xiaohui Wei 0002, Mengting Zhou, Nan Jiang 0013, Xiang Li 0197, Hengshan Yue |
WASA (3) | 1 |
| 2025 | Harnessing dynamic graph differential operators for efficient data-driven wind prediction
Xiaohui Wei 0002, Zhewen Xu, Hongliang Li 0003, Jieyun Hao, Hengshan Yue, Changzheng Liu |
GeoInformatica | 1 |
| 2025 | A distinct classification of attention mechanisms in video understanding
Xingwang Wang 0003, Yafeng Sun, Kun Yang 0001, Xiaohui Wei 0002 |
Inf. Sci. | 5 |
| 2025 | ResCheckpointer: Building Program Error Resilience-Aware Checkpointing Mechanism for HPC Systems
Xiaohui Wei 0002, Shiyu Tong, Zhongao Sun, Xiang Li 0197, Hengshan Yue |
J. Comput. Sci. Technol. | 1 |
| 2025 | EAAR: Efficient and Accurate Action Recognition model with enhanced spatio-temporal perception
Xingwang Wang 0003, Yafeng Sun, Kun Yang 0001, Xiaohui Wei 0002 |
Neural Networks | 5 |
| 2025 | Evaluating GPU's Instruction-Level Error Characteristics Under Low Supply VoltagesabstractSupply voltage underscaling has been an effective approach to improve the energy-efficiency of modern high-performance processors, such as GPUs. However, energy efficiency and reliability are two sides of a trade-off. Undervolting will inevitably undermine reliability, since it reduces chip manufacturers’ voltage guardbands that is designed to ensure correct operations under worst-case scenarios. To achieve optimal energy efficiency while maintaining enough reliability, it is necessary to deeply understand the error characteristics caused by undervolting. Unlike previous works which focus mostly on program level, we perform the first comprehensive instruction-level voltage margin and error characteristics evaluation for GPU architectures. We systematically measure the error probability and patterns of GPU instructions during undervolting. Then, we also analyze the impact of locations (SMs, threads, and bits) and operand data values on the error characteristics. Based on our observations, we reduce the voltage to the minimum safe limit for different instructions which achieves 18.37% energy saving, and we further propose an error detection strategy which reduces the performance and energy overhead by 14.8% with negligible 0.01% degradation for error detection rate. Jingweijia Tan, Jiashuo Wang, Kaige Yan, Xiaohui Wei 0002, Xin Fu 0001 |
IEEE Trans. Computers | 4 |
| 2025 | A Survey of Change Point Detection in Dynamic GraphsabstractChange point detection is crucial for identifying state transitions and anomalies in dynamic systems, with applications in network security, health care, and social network analysis. Dynamic systems are represented by dynamic graphs with spatial and temporal dimensions. As objects and their relations in a dynamic graph change over time, detecting these changes is essential. Numerous methods for change point detection in dynamic graphs have been developed, but no systematic review exists. This paper addresses this gap by introducing change point detection tasks in dynamic graphs, discussing two tasks based on input data types: detection in graph snapshot series (focusing on graph topology changes) and time series on graphs (focusing on changes in graph entities with temporal dynamics). We then present related challenges and applications, provide a comprehensive taxonomy of surveyed methods, including datasets and evaluation metrics, and discuss promising research directions. Shang Gao 0005, Dandan Guo, Xiaohui Wei 0002, Jon G. Rokne, Hui Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2025 | GEREM: Fast and Precise Error Resilience Assessment for GPU MicroarchitecturesabstractGPUs are widely used hardware acceleration platforms in many areas due to their great computational throughput. In the meanwhile, GPUs are vulnerable to transient hardware faults in the post-Moore era. Analyzing the error resilience of GPUs are critical for both hardware and software. Statistical fault injection approaches are commonly used for error resilience analysis, which are highly accurate but very time consuming. In this work, we propose GEREM, a first framework to speed up fault injection process so as to estimate the error resilience of GPU microarchitectures swiftly and precisely. We find early fault behaviors can be used to accurately predict the final outcomes of program execution. Based on this observation, we categorize the early behaviors of hardware faults into GPU Early Fault Manifestation models (EFMs). For data structures, EFMs are early propagation characteristics of faults, while for pipeline instructions, EFMs are heuristic properties of several instruction contexts. We further observe that EFMs are determined by static microarchitecture states, so we can capture them without actually simulating the program execution process under fault injections. Leveraging these observations, our GEREM framework first profiles the microarchitectural states related for EFMs at one time. It then injects faults into the profiled traces to immediately generate EFMs. For data storage structures, EFMs are directly used to predict final fault outcomes, while for pipeline instructions, machine learning is used for prediction. Evaluation results show GEREM precisely assesses the error resilience of GPU microarchitecture structures with$237\times$speedup on average comparing with traditional fault injections. Jingweijia Tan, An Zhong, Kaige Yan, Xiaohui Wei 0002, Guanpeng Li |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2024 | PGSampler: Accelerating GPU-Based Graph Sampling in GNN Systems via Workload FusionabstractGraph Neural Networks (GNNs) have demonstrated remarkable performance across various domains. Sample-based training, a practical strategy for training on large-scale graphs, often faces time-consuming graph sampling challenges. To address this, GPU-based graph sampling has been introduced, while there is still room for further efficiency improvements. Though several prior works have been proposed to accelerate the computation or memory access for GPU-based graph sampling, we show that the performance bottlenecks induced by small workload cannot be ignored. In this paper, we propose PGSampler, an efficient system for accelerating GPU-based graph sampling. First, PGSampler leverages a barrier-free execution mode to fuse workload, significantly improving the resource utilization. By altering the sampling execution mode, PGSampler also reduces the preprocessing time before kernel execution, thus accelerating the whole sampling process. Next, based on the new sampling execution mode, considering the dynamically generated nature of sampling tasks, PGSampler adopts a persistent kernel design and uses the task queue to assign tasks, achieving dynamic load balancing. Evaluations with diverse parameter settings show that PGSampler can achieve up to 2.22 × performance speedup over the state-of-the-art GNN system DGL. Xiaohui Wei 0002, Weikai Tang, Hao Qi 0004, Hengshan Yue |
CLUSTER | 1 |
| 2024 | HAp-FT: A Hybrid Approximate Fault Tolerance Framework for DNN AcceleratorabstractNowadays, as Deep Neural Networks (DNNs) become ubiquitous in mission-critical domains (e.g., automatic driving systems), ensuring the reliable execution of domain-specific DNN accelerators in the presence of hardware faults is increasingly essential. However, recent fault tolerance attempts either suffer from expensive performance overhead or are limited in scalability by specific fault modes. This paper proposes HAp-FT, a Hybrid Approximate fault Tolerance framework that integrates both proactive error alleviation and efficient error detection design philosophies. First, considering the vertical accumulation execution paradigm of the systolic array, HAp-FT proactively transfers the reliability risk of vulnerable weight to neighboring Processing Elements (PEs) for advance fault effect alleviation. Then, leveraging the inter-filter similarity characteristics, HAp-FT remaps similar filters to adjacent columns of PEs to achieve real-time intragroup approximate error detection. Experimental results exhibit that HAp-FT can recover 98.38%95.74% accuracy degradation incurred by transient/permanent hardware faults while only introducing 0.18% performance overhead, 2.69% extra area, and 2.38% energy overhead. Moreover, to satisfy the reliability requirements of various application scenarios, HAp-FT is also portable to systolic arrays under all dataflow strategies and adaptable quantized models. Xiaohui Wei 0002, Zeyu Guan, Fengyi Li, Hengshan Yue |
ICCD | 1 |
| 2024 | EFNAS: Efficient Federated Neural Architecture Search Across AIoT DevicesabstractFederated neural architecture search tailors deep learning models to accommodate varied client data in Federated Learning (FL) scenarios. However, the simultaneous optimization of multiple subnetworks leads to substantial GPU memory overhead in differentiable Neural Architecture Search (NAS) methods. Additionally, after each client searches local architecture, the conventional weighted averaging approach may result in a loss of architectural feature information and failure to capture architectural diversity, limiting model performance and expressiveness. To address these challenges, we propose EFNAS, a novel computation-efficient and aggregation-effective Federated NAS framework. Specifically, we propose Single Path Local Search (SPLS) to automatically search for the optimal network architecture with minimal complexity. SPLS uses Gumbel Softmax to continually reparameterize the probability distribution to reduce complexity. Furthermore, we propose Client-Centric Architecture Aggregation (CCAA), considering the aggregated architecture as a graph with subgraphs representing client architectures. CCAA leverages probabilistic distributions extracted from subgraphs that frequently appear across multiple clients to derive a global architecture. Guided by this global architecture, each client compares the accuracy of its locally searched architecture with the global architecture, selecting the architecture best suited for its final model architecture. Comprehensive experiments on various datasets demonstrate that EFNAS achieves excellent performance while guaranteeing efficiency during searching compared to other methods. Xiaohui Wei 0002, Guanhua Chen 0003, Hairui Zhao 0002, Hengshan Yue |
IJCNN | 1 |
| 2024 | Improving INT Inference Resilience of CNNs by Vulnerability-Aware Mixed-Precision QuantizationabstractThe application of INT quantization allows CNN model size and computation to be reduced exponentially, making extreme quantization a hot research topic. However, firstly the decrease in bitwidth increases the probability of error in critical bits, which is unacceptable for safety-critical scenarios that require high resilience of CNNs. Secondly, extreme quantization often requires retraining with high overhead. Considering the higher tolerance of errors in high-bitwidth INTs and the local resilience differences of CNNs, we use mixed precision quantization (MPQ) to trade off resilience and compression effects. Existing MPQ methods mostly focus on compression as the first goal and lack indicators of resilience or vulnerability. Based on a large number of vulnerability analyses, we propose to use the maximum bit average offset and the approximation parameter average quantization error to measure layer vulnerability. And we measure the vulnerability of the entire CNN in terms of the accumulation of layer offsets. Accordingly, we design a complete vulnerability-aware MPQ (VulnA-MPQ) framework. From the perspective of minimizing vulnerability and compression ratio, VulnA-MPQ is performed on multiple CNNs using genetic algorithm (GA) with low overhead. The CNNs only need to be fine-tuned to restore accuracy after VulnA-MPQ. Experiments demonstrate that VulnA-MPQ not only maintains accuracy and compression ratio better but also improves CNNs fault tolerance. Within the 2% accuracy threshold, the random error data tolerated by VGG16 weights and activations are improved by 33%. The number of random errors tolerated by MobileNetV3-large weights is improved by 8.87 times, and that of activations is improved by 4.64 times. Xiaohui Wei 0002, Zeyu Guan, Hengshan Yue |
ISPA | 1 |
| 2024 | CraftRGP: A Comprehensive Reliability Analysis Framework Towards ReRAM-Based Graph ProcessingabstractThe burgeoning proliferation of graph data has rendered graph processing within conventional von Neumann computing architectures increasingly inefficient. Resistive Random Access Memory (ReRAM) emerges as a promising solution to circumvent the "memory wall" bottleneck encountered in graph processing. Nevertheless, the inherent non-idealities of ReRAM devices raise reliability concerns, the ramifications of which for graph processing necessitate a comprehensive exploration. In this work, we propose CraftRGP, a Comprehensive Reliability Analysis Framework Towards ReRAM-Based Graph Processing. CraftRGP simulates the operation of graph processing on ReRAM devices at the underlying hardware level, and employs software fault injection to simulate three primary fault models on ReRAM devices: permanent faults, transient faults, and sensing faults. Based on CraftRGP, we systematically evaluate the fault resilience of ReRAM-based graph processing along the dimensions of graph applications, fault models, and hardware architectures. Finally, utilizing the observed reliability characteristics with CraftRGP, we exemplify two use cases to demonstrate that CraftRGP can guide system designers to facilitate better reliability designs. Xiaohui Wei 0002, Jiaguo Deng, Zongdian Li, Nan Jiang 0013, Hengshan Yue |
ITC-Asia | 1 |
| 2024 | HiRM: Hierarchical resource management for earth system models on many-core clusters
Zhewen Xu, Xiaohui Wei 0002, Jieyun Hao, Hongliang Li 0003, Zhaohui Ding |
CCF Trans. High Perform. Comput. | 2 |
| 2024 | DUAL-C: Building a "soft error efficient" on-the-fly compression mechanism for raw video data at edge devices
Xiaohui Wei 0002, Hengshan Yue, Nan Jiang 0013, Jianpeng Zhao 0001, Meikang Qiu |
Future Gener. Comput. Syst. | 1 |
| 2024 | HSAS: Efficient task scheduling for large scale heterogeneous systolic array accelerator cluster
Kaige Yan, Yanshuang Song, Jingweijia Tan, Xiaohui Wei 0002, Xin Fu 0001 |
Future Gener. Comput. Syst. | 5 |
| 2024 | DGFormer: a physics-guided station level weather forecasting model with dynamic spatial-temporal graph neural network
Zhewen Xu, Xiaohui Wei 0002, Jieyun Hao, Junze Han, Hongliang Li 0003, Changzheng Liu, Zijian Li 0007, Dongyuan Tian, Nong Zhang |
GeoInformatica | 2 |
| 2024 | ALERT: A lightweight defense mechanism for enhancing DNN robustness against T-BFA
Xiaohui Wei 0002, Yumin Yan, Nan Jiang 0013, Hengshan Yue |
J. Syst. Archit. | 1 |
| 2024 | SAR: Sharpness-Aware minimization for enhancing DNNs' Robustness against bit-flip errors
Changbao Zhou, Jiawei Du 0002, Ming Yan 0007, Hengshan Yue, Xiaohui Wei 0002, Joey Tianyi Zhou |
J. Syst. Archit. | 5 |
| 2024 | ReIPE: Recycling Idle PEs in CNN Accelerator for Vulnerable Filters Soft-Error DetectionabstractTo satisfy prohibitively massive computational requirements of current deep Convolutional Neural Networks (CNNs), CNN-specific accelerators are widely deployed in large-scale systems. Caused by high-energy neutrons and α-particle strikes, soft error may lead to catastrophic failures when CNN is deployed on high integration density accelerators. As CNNs become ubiquitous in mission-critical domains, ensuring the reliable execution of CNN accelerators in the presence of soft errors is increasingly essential. In this article, we propose to Re cycle I dle P rocessing E lements (PEs) in the CNN accelerator for vulnerable filters soft error detection (ReIPE). Considering the error-sensitivity of filters, ReIPE first carries out a filter-level gradient analysis process to replace fault injection for fast filter-wise error resilience estimation. Then, to achieve maximal reliability benefits, combining the hardware-level systolic array idleness and software-level CNN filter-wise error resilience profile, ReIPE preferentially duplicated loads the most vulnerable filters onto systolic array to recycle idle-column PEs for opportunistically redundant execution (error detection). Exploiting the data reuse properties of accelerators, ReIPE incorporates the error detection process into the original computation flow of accelerators to perform real-time error detection. Once the error is detected, ReIPE will trigger a correction round to rectify the erroneous output. Experimental results performed on LeNet-5, Cifar-10-CNN, AlexNet, ResNet-20, VGG-16, and ResNet-50 exhibit that ReIPE can cover 96.40% of errors while reducing 75.06% performance degradation and 67.79% energy consumption of baseline dual modular redundancy on average. Moreover, to satisfy the reliability requirements of various application scenarios, ReIPE is also applicable for pruned, quantized, and Transformer-based models, as well as portable to other accelerator architectures. Xiaohui Wei 0002, Hengshan Yue, Jingweijia Tan, Zeyu Guan, Nan Jiang 0013, Xinyang Zheng, Jianpeng Zhao 0001, Meikang Qiu |
ACM Trans. Archit. Code Optim. | 1 |
| 2024 | ApproxDup: Developing an Approximate Instruction Duplication Mechanism for Efficient SDC Detection in GPGPUsabstractNowadays, selective instruction duplication (SelDup) is the typical approach to detect silent data corruption (SDC) in GPGPU. However, owing to the up-to-billions fault sites of parallel GPGPU kernel functions, it usually introduces tremendous overhead to perform fault injections (FIs) for obtaining the duplication-candidate instruction set (although can be conducted in parallel). Moreover, current SelDup typically considers all SDCs severe and tends to duplicate more instructions. The nontrivial duplication overhead seriously restricts the deployment of current SelDup on resource-constrained systems (e.g., embedded GPGPUs). To address the above challenges, this article proposes an approximate instruction duplication (ApproxDup) mechanism for efficient SDC detection in GPGPUs. First, to replace the expensive FI-based duplication-candidate instructions identified method, we drive out a machine learning (ML)-based model (SDC-predictor) for instructionwise SDC proneness and severity estimation. Our key insight is that instruction type/functionality and instruction dependency set can efficaciously characterize the instructionwise SDC proneness in GPGPUs. In contrast, the instruction’s original data magnitude, fault propagation range, and error detected features can distinguish its SDC severity. Second, incorporating the concept of approximate computing, we propose ApproxDup that preferentially duplicates severe-SDC-prone instructions while relaxing the detection of minor/detectable SDCs for traditional SelDup overhead reduction. Experimental results exhibit that ApproxDup can cover 92.51% of severe SDCs while merely increasing 38% of dynamic instructions, which achieves a better tradeoff between reliability and performance compared with the state-of-the-art SelDup. Furthermore, we discuss the effectiveness of the proposed method on different ML models/applications/GPGPU architectures. Xiaohui Wei 0002, Nan Jiang 0013, Hengshan Yue, Jianpeng Zhao 0001, Guangli Li, Meikang Qiu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | Exploiting Complex Network-Based Clustering for Personalization-Enhanced Hierarchical Federated Edge LearningabstractFederated Learning (FL) has been extensively applied in urban environmental prediction tasks of mobile edge computing by training a global machine learning model without data sharing. However, the training of FL faces the challenges such as the poor generalization capability of a single global model over heterogeneous data and hefty communication overhead caused by the frequent model exchange between massive edge servers and remote cloud servers. To address such issues, we propose HPFL-CN, a novel communication-efficient Hierarchical Personalized Federated edge Learning framework with Complex Network clustering. HPFL-CN introduces Privacy-preserving Feature Clustering (PFC) to extract privacy-preserving low-dimensional feature representations of each edge server via mapping the environmental data to different complex network domains for clustering similar edge servers accurately. Based on the clustering results of PFC, anedge-mediator-cloudhierarchical architecture is proposed to realize personalization at the cluster level by Effective Hierarchical Scheduling (EHS). Furthermore, to adapt to dynamic scenarios of new edge servers joining and streaming data generation, we further extend HPFL-CN to Adaptive personalized federated learning with dynamic grouping (Ada-HPFL-CN), which can flexibly re-group edge servers and adjust mixed model weights and the model aggregation frequency adaptively. Our extensive experiments on real-world datasets demonstrate the efficacy of our framework, which outperforms state-of-the-art FL methods regarding personalization and communication efficiency performance. Zijian Li 0007, Zihan Chen 0001, Xiaohui Wei 0002, Shang Gao 0005, Hengshan Yue, Zhewen Xu, Tony Q. S. Quek |
IEEE Trans. Mob. Comput. | 3 |
| 2023 | Detecting SDCs in GPGPUs Through Efficient Partial Thread Redundancy
Xiaohui Wei 0002, Nan Jiang 0013, Hengshan Yue |
ICA3PP (7) | 1 |
| 2023 | GLAM-SERP: Building a Graph Learning-Assisted Model for Soft Error Resilience Prediction in GPGPUs
Xiaohui Wei 0002, Jianpeng Zhao 0001, Nan Jiang 0013, Hengshan Yue |
ICA3PP (4) | 1 |
| 2023 | CFPA: Cognitive Federated Partial Adaptation for Effective PersonalizationabstractFederated Learning (FL) is a promising machine learning paradigm to train a global model from multiple clients while ensuring privacy. A key challenge in FL is the data heterogeneity that impairs the generalization of the single global model on each client. As a prevalent solution, the existing personalized FL primarily concentrates on the client-level personalization process, which may pose an overfitting problem since the local data of a single device is usually limited. Additionally, these methods lack the cognitive ability for collaborative learning among similar clients, which also restricts the performance of personalized models. To address this, we propose CFPA, a novel communication-efficient, and computation-efficient Cognitive Federated Partial Adaptation framework to boost personalization performance via the Soft-grouping Weighted Aggregation (SWA) strategy. Specifically, CFPA decouples the model into the body and the head and requires clients collaboratively train a well-performing federated pre-personalized model. Then, CFPA leverages the distribution knowledge extracted from the output feature maps of the convolutional layers in the federated pre-personalized model to identify the similarity among clients efficiently. Guided by the similarity matrix, CFPA further performs weighted federated adaption on the head of each local model, ultimately generating a personalized local model for each client. Comprehensive experiments on three benchmark datasets with various heterogeneous settings demonstrate that CFPA outperforms other state-of-the-art FL approaches. Xiaohui Wei 0002, Didi Jiao, Shiyu Tong, Zijian Li 0007, Chenghao Ren, Hengshan Yue |
MSN | 1 |
| 2023 | FASS-pruner: customizing a fine-grained CNN accelerator-aware pruning framework via intra-filter splitting and inter-filter shuffling
Xiaohui Wei 0002, Xinyang Zheng, Guangli Li, Hengshan Yue |
CCF Trans. High Perform. Comput. | 1 |
| 2023 | Saca-FI: A microarchitecture-level fault injection framework for reliability analysis of systolic array based CNN accelerator
Jingweijia Tan, Qixiang Wang, Kaige Yan, Xiaohui Wei 0002, Xin Fu 0001 |
Future Gener. Comput. Syst. | 4 |
| 2023 | HSM-SMCS: Task Assignment Based on Hybrid Sensing Modes in Sparse Mobile CrowdsensingabstractSparse mobile crowdsensing (Sparse MCS) is an emerging paradigm for urban-scale sensing applications, which recruits suitable participants to complete sensing tasks in only a few selected cells and then infers data of unsensed cells for saving sensing costs and obtaining high-quality sensing maps. In Sparse MCS, one crucial issue is task assignment, in which the platform selects cells whose sensing data can reduce inferred sensing maps errors (i.e., cell selection) and recruits the participant set with the maximum contribution for performing tasks (i.e., participant recruitment). The research on participant recruitment mainly focuses on single participatory-based or single opportunistic-based sensing mode. Due to the complementarity of two sensing modes, recruiting participants by only one sensing mode would result in wasting sensing resources and compromising the quality of task completion. Thus, combining the advantages of two sensing modes, we propose a task assignment framework based on hybrid sensing modes in Sparse MCS (HSM-SMCS) for achieving a good tradeoff between sensing quality and cost. Specifically, we propose a heuristic two-stage search strategy that simultaneously recruits opportunistic and participatory participants to perform tasks in significant cells within the constraint of total costs, considering their contributions to sensing map inference. Thereinto, for opportunistic participants, mobility prediction greatly affects task assignment effectiveness. However, existing prediction algorithms lead to unsatisfactory outcomes when the historical trajectory data of opportunistic participants are scarce. To effectively improve the predictive accuracy, we design a mobility prediction model based on transfer learning. The experimental evaluation on real trajectory data sets and sensor data sets of corresponding areas demonstrates that our framework outperforms state-of-the-art methods with higher quality reconstructed sensing maps. Xiaohui Wei 0002, Zijian Li 0007, Chenghao Ren, Shang Gao 0005 |
IEEE Internet Things J. | 1 |
| 2023 | TC-SEPM: Characterizing soft error resilience of CNNs on Tensor Cores from program and microarchitecture perspectives
Xiaohui Wei 0002, Changbao Zhou, Hengshan Yue, Joey Tianyi Zhou |
J. Syst. Archit. | 1 |
| 2023 | MCM-GPU Voltage Noise Characterization and Architecture-Level MitigationabstractDue to manufacturing process and yield constraints, scaling GPU performance via increasing chip area becomes difficult. In the meanwhile, the demand for high computational throughput is increasing for high performance computing applications. As an alternative, multichip module GPU (MCM-GPU) achieves performance scalability via integrating multiple GPU chip modules (GPMs) on the same package. However, large MCM-GPU systems are susceptible to voltage noise effects, which cause voltage instability during program execution and result in energy inefficiency. In this work, we first model and analyze the voltage noise of MCM-GPUs at architecture level in detail. We characterize the voltage noise distributions of MCM-GPUs at different levels and under various design parameters. We further propose two architecture level voltage noise mitigation approaches, including GPM aware mitigation (GAM) and droop magnitude aware smoothing (DMAS), that leverage the voltage noise characteristics of MCM-GPUs. Evaluation shows both techniques are effective in reducing the voltage droop magnitudes and achieve good energy savings with negligible performance degradation. Jingweijia Tan, Weiren Wang, Kaige Yan, Xiaohui Wei 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | Location-and-Preference Joint Prediction for Task Assignment in Spatial CrowdsourcingabstractWith the rapid development of mobile networks and the ubiquity of mobile devices, spatial crowdsourcing (SC), which refers to assigning spatial–temporal tasks to moving workers, has drawn increasing attention. Thus, many researchers aim at various task assignment methods in SC. However, existing works generally consider workers’ location and preference categories separately. Ignorance of the correlation between them can often lead to poor assignment results. In this article, we propose a location-and-preference joint prediction model (JPM) to predict workers’ locations and preference categories jointly at each sample timestamp. Based on the predictive location probability distribution and preference probability distribution, we elaborately design a greedy multiattribute joint task assignment algorithm (MAJA) to maximize the average number of completed tasks under constraints. Then, an overall procedure incorporating the JPM and MAJA, called the location-and-preference joint prediction-based task assignment (LPJTA), is implemented to focus on assigning tasks to workers who are near the task location and willing to perform the task based on predicting locations and preference categories. We theoretically analyze the time complexity and approximation ratio of the proposed methods and construct extensive experiments on three real datasets to empirically verify their effectiveness, comparing with the state-of-the-art baselines. Xiaohui Wei 0002, Bingyi Sun, Jiaxu Cui, Meikang Qiu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | Improving the Performance of CNN Accelerator Architecture under the Impact of Process VariationsabstractConvolutional neural network (CNN) accelerators are popular specialized platforms for efficient CNN processing. As semiconductor manufacturing technology scales down to nano scale, process variation dramatically affects the chip’s quality. Process variation causes delay variation within the chip due to transistor parameter differences. CNN accelerators adopt a large number of processing elements (PEs) for parallel computing, which are highly susceptible to process variation effects. Fast CNN processing desires consistent performance among PEs; otherwise the processing speed is limited by the slowest PE within the chip. In this work, we first quantitatively model and analyze the impact of process variation on CNN accelerators’ operating frequency. We further analyze the utilization of CNN accelerators and the characteristics of CNN models. We then leverage the PE underutilization to propose a sub-matrix reformation mechanism and leverage the pixel similarity of images to propose a weight transfer technique. Both techniques are able to tolerate the low-frequency PEs and achieve performance improvement at chip level. Furthermore, a novel resilience-aware mapping technique that exploits the diversity in the importance of weights is also proposed to improve the performance. Evaluation results show that our techniques are able to achieve significant processing speed improvement with negligible accuracy loss. Jingweijia Tan, Weiren Wang, Maodi Ma, Xiaohui Wei 0002, Kaige Yan |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2022 | HRCache: Edge-End Collaboration for Mobile Deep Vision Based on H.264 and Approximated ReuseabstractTo accomplish computationally intensive visual tasks on mobile devices with limited memory and computation capability, a common solution is to offload the tasks to the edge with more powerful computing resources. Nevertheless, due to the extra communication, even edge computing is challenged by the growing demands for real-time interactive tasks (e.g., VR and AR), which are generally performed by deep CNNs with heavy computation. In this paper, we design a novel system called HRCache to reduce the end-to-end latency by effectively recompressing video data and reusing cached inference results on the edge. As similar image regions have been indicated in the offloading video coding, the edge can approximately infer these regions with results stored on the edge to save the time for calculating these regions. Moreover, since the data in these regions do not require calculation, the transmitted data can be further simplified to reduce transmission latency. Furthermore, HRCache can quickly and continuously adjust the coding parameters to adapt the accuracy loss and overall latency with application requirements. Compared with the original offloading schemes, HRCache significantly reduces the average latency, about 13.60% to 18.83%, at little accuracy loss of 1.25% in Top-1 accuracy for classification and 0.135 in IoU for object detection. Xiaohui Wei 0002, Xiukun Wei, Xingwang Wang 0003, Yundi Wang, Yan Niu |
IPCCC | 1 |
| 2022 | MSSA-FL: High-Performance Multi-stage Semi-asynchronous Federated Learning with Non-IID Data
Xiaohui Wei 0002, Mingkai Hou, Chenghao Ren, Xiang Li 0197, Hengshan Yue |
KSEM (2) | 1 |
| 2022 | Approximation Algorithms for Reliability-Aware Maximum VoI on AUV-Aided Data Collections
Xiaohui Wei 0002, Xingwang Wang 0003, Chenghao Ren, Meikang Qiu |
NPC | 2 |
| 2022 | HPFL-CN: Communication-Efficient Hierarchical Personalized Federated Edge Learning via Complex Network Feature ClusteringabstractFederated Learning (FL), a promising privacy-preserving distributed learning paradigm, has been extensively applied in urban environmental prediction tasks of Mobile Edge Computing (MEC) by training a global machine learning model without data sharing. However, it is hard for the shared global model to be well generalized among local edge servers, due to the statistical data heterogeneity, especially in real-world urban environmental data. Besides, the existing FL approaches may result in excessive communication and computation overhead due to the frequent transmission and aggregation of model parameters between massive edge servers and remote cloud servers. To address the above issues, we propose HPFL-CN, a novel communication-efficient Hierarchical Personalized Federated edge Learning framework via Complex Network feature clustering, aiming to cluster edge servers with similar environmental data distributions and then high-efficiently train personalized models for each cluster via hierarchical architecture. Specifically, HPFL-CN introduces Privacy-preserving Feature Clustering (PFC) to extract privacy-preserving low-dimensional feature representations of each edge server via mapping the environmental data to different complex network domains for clustering similar edge servers accurately. According to the clustering results of PFC, HPFL-CN further introduces an edge-mediator-cloud architecture for hierarchical model aggregation by Effective Hierarchical Scheduling (EHS), in which every mediator coordinates the training of edge servers within each cluster and periodically uploads model to cloud server for global model aggregation. Meanwhile, each mediator server would find a trade-off between cloud and edge models to realize personalization within clusters. Our extensive experiments on real-world datasets demonstrate the effectiveness and generalization of HPFL-CN, which outperforms other state-of-the-art FL methods regarding personalization performance and communication efficiency. Zijian Li 0007, Zihan Chen 0001, Xiaohui Wei 0002, Shang Gao 0005, Chenghao Ren, Tony Q. S. Quek |
SECON | 3 |
| 2022 | Cooperative task assignment in spatial crowdsourcing via multi-agent deep reinforcement learning
Xiang Li 0197, Shang Gao 0005, Xiaohui Wei 0002 |
J. Syst. Archit. | 4 |
| 2022 | Dynamic Graph-Level Neural Network for SAR Image Change DetectionabstractThe graph neural network (GNN) has been widely applied to image analysis and recognition. Recently, a semisupervised graph convolutional network (ssGCN) method has been proposed to change detection and obtains promising performance on very-high-resolution remote sensing images. However, a synthetic aperture radar (SAR) image is subject to speckle noise, and there is no explicit structure. In this letter, an end-to-end dynamic graph-level neural network (DGLNN) is proposed to exploit the local structure of each pixel neighborhood block at a graph level and learn a more discriminative graph for change detection. Moreover, in the training of DGLNN, a$K$-nearest neighborhood is employed to reconstruct edges between nodes instead of the fixed edges between two nodes so that each node exploits the features from different neighbor nodes. The proposed method is verified by cross-domain SAR image change detection on four sets of SAR images and compared with five state-of-the-art deep-learning-based SAR image change detection methods. The overall experimental results show that the proposed DGLNN obtains outstanding performance. Rongfang Wang, Liang Wang 0043, Xiaohui Wei 0002, Jiawei Chen 0001, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Eff-ECC: Protecting GPGPUs Register File With a Unified Energy-Efficient ECC MechanismabstractGraphics processing units (GPUs) are widely used in general-purpose high-performance computing applications (i.e., GPGPUs), which require reliable execution in the presence of soft errors. To support massive thread-level parallelism, a sizeable register file is adopted in GPUs, which is highly vulnerable to soft errors. Although modern commercial GPUs provide single-error-correction double-error-detection (SEC-DED) error correction code (ECC) for the register file, it consumes a considerable amount of energy due to frequent register accesses and leakage power of ECC storage. In this article, we propose to leverage the error sensitivity of instructions, the duplicate characteristics of the same-named registers, and the error sensitivity of data bits to build a unified energy-efficient ECC mechanism for a GPGPUs register file (Eff-ECC), which consists of instruction-aware ECC (IA-ECC), duplication-aware ECC (DA-ECC), and bit-aware ECC (BA-ECC). Considering the error sensitivity of instructions, IA-ECC merely implements ECCs for the write registers of critical instructions. Observing the same-named registers across threads usually keeps the same data, DA-ECC avoids unnecessary ECC generation and verification for duplicate register values. Leveraging the inherent error-tolerance features of the program, BA-ECC merely protects significant bits of registers to combat the crucial error. Experimental results demonstrate that Eff-ECC tremendously reduces 86.46% energy consumption of traditional SEC-DED ECC. Hengshan Yue, Xiaohui Wei 0002, Jingweijia Tan, Nan Jiang 0013, Meikang Qiu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | Detecting SDCs in GPGPUs Through an Efficient Instruction Duplication Mechanism
Xiaohui Wei 0002, Nan Jiang 0013, Hengshan Yue |
KSEM | 1 |
| 2021 | G-SEPM: building an accurate and efficient soft error prediction model for GPGPUsabstractAs GPUs become ubiquitous in large-scale general purpose HPC systems (GPGPUs), ensuring the reliable execution of such systems in the presence of soft errors is increasingly essential. To provide insights into how resilient GPU programs are toward soft errors, researchers typically rely on random Fault Injection (FI) to evaluate the tolerance of programs. However, it is expensive to obtain a statistically significant resilience profile and not suitable to identify all the error-critical fault sites of GPU programs. Hengshan Yue, Xiaohui Wei 0002, Guangli Li, Jianpeng Zhao 0001, Nan Jiang 0013, Jingweijia Tan |
SC | 2 |
| 2021 | Coordinated process scheduling algorithms for coupled earth system modelsabstractAbstract It is becoming increasingly significant for humans to predict and understand future climate changes using coupled climate system models. Although the performance and scalability of individual physical components have improved over the past few years, coupled climate systems still suffer from low efficiency. This paper focuses on the process scheduling problem for the widely applied coupled earth system model (CESM). The proposed resource allocation strategies allow components to execute on a compromised suboptimal setup and still maintain approximately the best parallel speedup. With this flexible resource allocation strategy, we further propose a coordinated process scheduling algorithm (CPSA). More notably, we propose an upgraded version called CPSA‐B, which makes efficient resource sharing configurations, including resource allocation and process layout of components. We integrate CPSA and CPSA‐B as pre‐arrangement tools into the CESM program and deploy them on the Huawei Kunpeng platform. The speedup curves of the CESM components are prepared in advance, based on sampling tests. Experimental data show that CPSA‐B reduces up to 58% of the execution time compared with the CESM default strategy. The algorithm has low complexity and can efficiently find solutions for large input sizes. Xiaohui Wei 0002, Zhewen Xu, Hongliang Li 0003, Zhaohui Ding |
Concurr. Comput. Pract. Exp. | 1 |
| 2021 | STAC: a spatio-temporal approximate method in data collection applications
Xiaohui Wei 0002, Sijie Yan, Xingwang Wang 0003, Mohsen Guizani, Xiaojiang Du |
Pervasive Mob. Comput. | 1 |
| 2020 | LAD-ECC: Energy-Efficient ECC Mechanism for GPGPUs Register FileabstractGraphics Processing Units (GPUs) are widely used in general-purpose high-performance computing applications (i.e., GPGPUs), which require reliable execution in the presence of soft errors. To support massive thread level parallelism, a sizeable register file is adopted in GPUs, which is highly vulnerable to soft errors. Although modern commercial GPUs provide single-error-correction double-error-detection (SEC-DED) ECC for the register file, it consumes a considerable amount of energy due to frequent register accesses and leakage power of ECC storage.In this paper, we propose to Leverage Approximation and Duplication characteristics of register values to build an energy-efficient ECC mechanism (LAD-ECC) in GPGPUs, which consists of APproximation-aware ECC (AP-ECC) and Duplication-Aware ECC (DA-ECC). Leveraging the inherent error tolerance features, AP-ECC merely protects significant bits of registers to combat the critical error. Observing same-named registers across threads usually keep the same data, DA-ECC avoids unnecessary ECC generation and verification for duplicate register values. Experimental results demonstrate that our LAD-ECC tremendously reduces 69.72% energy consumption of traditional SEC-DED ECC. Xiaohui Wei 0002, Hengshan Yue, Jingweijia Tan |
DATE | 1 |
| 2020 | A Novel Clustering-Based Filter Pruning Method for Efficient Deep Neural Networks
Xiaohui Wei 0002, Xiaoxian Shen, Changbao Zhou, Hengshan Yue |
ICA3PP (2) | 1 |
| 2020 | CPSA: A Coordinated Process Scheduling Algorithm for Coupled Earth System ModelabstractCoupled climate system models are important tools for climatologists to predict and understand future climate. These models are usually resource-consuming due to the large number of processors required and long execution time. Although the performance and scalability of individual physical system model have been improved over the past years, coupled climate systems still suffer from low efficiency when sharing resource across models. This paper focuses on the process scheduling strategy of Coupled Earth System Model (CESM), a widely applied coupled system model. Instead of pursuing best speedup efficiency for individual component, the proposed resource allocation strategy allows components to execute on compromised sub-optimal setup and still maintains relatively high parallel speedup. With this flexible resource allocation strategy, we further propose a Coordinated Process Scheduling Algorithm (CPSA) to make efficient resource sharing configurations, including resource allocation and process layout of components. We integrate CPSA as a tool into CESM program, and deploy it on Huawei Kunpeng Platform. Speedup curves of CESM components are prepared in advance based on sampling tests. Experimental data show that our algorithm reduces up to 52.6% of execution time compared with CESM default strategy. We also present simulation data to show that our algorithm is efficient for the platforms with up to a million cores. Hongliang Li 0003, Zhewen Xu, Fangyu Tang, Xiaohui Wei 0002, Zhaohui Ding |
ICCCN | 4 |
| 2020 | Reducing Fault-tolerant Overhead for Distributed Stream Processing with Approximate BackupabstractThe stream processing model continuously processes online data in an on-pass fashion that can be more vulnerable to failures than other offline-data processing schemes. Checkpoint-based fault-tolerant methods have been widely used to enhance the reliability of stream processing systems. To ensure exact data recoveries upon failures, full-backup mechanisms are used to store a complete copy of data, which introduces substantial runtime overhead and increases output latency. In the meantime, a wide range of online processing applications prefer quick-and-dirty results with a slight degradation inaccuracy to delayed exact results. This paper introduces a novel approximate fault-tolerant problem (OAFP) with the objective of reducing the failure-free fault-tolerant overhead and ensuring user-defiled output accuracy requirement upon failure at the same time. We present an approximate fault-tolerant scheme based on sampling backup mechanism and study the trade-off between fault-tolerant overhead and output accuracy in stream processing systems. We proposed two algorithms to compute backup plans for both single-node failure and correlated failure scenarios. Extensive experiments with different types of stream topologies are conducted on our simulator to verify the correctness and effectiveness of our approach. We prove our solution guarantees the output accuracy requirement with minimum FT latency for directed acyclic graph (DAG) stream topologies with single-node failures. Yuan Zhuang 0003, Xiaohui Wei 0002, Hongliang Li 0003, Mingkai Hou, Yundi Wang |
ICCCN | 2 |
| 2020 | G-SEAP: Analyzing and characterizing soft-error aware approximation in GPGPUs
Xiaohui Wei 0002, Hengshan Yue, Shang Gao 0005, Ruyu Zhang, Jingweijia Tan |
Future Gener. Comput. Syst. | 1 |
| 2020 | Efficient receiver-based flooding in mobile ad hoc networks
Xin Bai 0004, Xiaohui Wei 0002, Sen Bai |
Wirel. Networks | 2 |
| 2019 | Process Variation Mitigation on Convolutional Neural Network Accelerator ArchitectureabstractConvolutional Neural Network (CNN) accelerators are popular specialized platforms for efficient CNN processing. As semiconductor manufacturing technology scales down to nano scale, process variation dramatically affects the chip's quality. Process variation causes delay variation within the chip due to transistor parameter differences. CNN accelerators adopt a large number of Processing Elements (PEs) for parallel computing, which are highly susceptible to process variation effects. Fast CNN processing desires consistent performance among PEs, otherwise the processing speed is limited by the slowest PE within the chip. In this work, we first quantitatively model and analyze the impact of process variation on CNN accelerator's operating frequency. We further analyze the utilization of CNN accelerator and the characteristics of CNN models. We then leverage the PE underutilization to propose a sub-matrix reformation mechanism and leverage the pixel similarity of images to propose a weight transfer technique. Both techniques are able to tolerate the low-frequency PEs, and achieve performance improvement at chip level. Evaluation results show our techniques are able to achieve significant processing speed improvement with negligible accuracy loss. Maodi Ma, Jingweijia Tan, Xiaohui Wei 0002, Kaige Yan |
ICCD | 3 |
| 2019 | An optimal checkpointing model with online OCI adjustment for stream processing applicationsabstractSummary Checkpoint‐based fault‐tolerant (FT) methods have been widely used to enhance the reliability of stream processing systems, but a checkpointing process usually introduces considerable overhead. It is a critical issue to choose the optimal checkpoint interval (OCI) that maximizes the processing efficiency. Traditional OCI models consider the recovery time equals to the execution time from the last checkpoint to the failure moment. However, for stream processing jobs, the recovery time is related to reprocessing workloads, depending on the real‐time input data before a failure. A new model is needed to choose the OCI for stream processing applications. Moreover, the input data rate of a stream processing job fluctuates over time. To solve these problems, we present a novel DSPS OCI (DOCI) model in this paper. We prove that it maximizes the processing efficiency for a given time. We propose an approach to dynamically adjust the OCI for an application to accommodate the workload fluctuations. We conduct simulation experiments to verify the effectiveness of our DOCI model and the efficiency of the online OCI adjustment algorithm. Experimental results with a real‐world dataset show that DOCI achieves an improvement on system efficiency by up to 32%, compared with existing FT approaches. Yuan Zhuang 0003, Xiaohui Wei 0002, Hongliang Li 0003, Xubin He |
Concurr. Comput. Pract. Exp. | 2 |
| 2019 | A survey on quality-assurance approximate stream processing and applications
Xiaohui Wei 0002, Yuanyuan Liu 0007, Xingwang Wang 0003, Bingyi Sun, Shang Gao 0005, Jon G. Rokne |
Future Gener. Comput. Syst. | 1 |
| 2019 | Pec: Proactive Elastic Collaborative Resource Scheduling in Data Stream ProcessingabstractIn the Distributed Parallel Stream Processing Systems (DPSPS), elastic resource allocation allows applications to dynamically response to workload fluctuations. However, resource provisioning can be particularly challenging, due to the unpredictability of the workload. In addition, unlike CPU resources, bandwidth resources are often ignored in resource allocation. Moreover, resource allocation and resource placement are considered separately. In this paper, we investigate the proactive elastic resource scheduling problem for computation-intensive and communication-intensive applications, which aims at meeting the latency requirement with the minimal energy cost, and propose a dynamic collaborative strategy from the systemic perspective. Specifically, we first model a collaborative workload prediction pattern to accurately predict the upcoming workload, and construct a latency estimation model to estimate the latency of the application. Then, we design an energy-efficient resource pre-allocation method, in which the CPU frequency adjustment and the stability of resource reconfigurations are both considered. Finally, we present a communication-aware resource placement approach. Simulation results show that, compared with the reactive strategies, our strategy achieves an obviously better latency performance, and effectively avoids unnecessary resource adjustments. Meanwhile, the energy consumption is about saved by 50 percent on average, and the communication cost is maintained at a very low level of 4 percent. Xiaohui Wei 0002, Xiang Li 0197, Xingwang Wang 0003, Shang Gao 0005, Hongliang Li 0003 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2018 | An Optimal Checkpointing Model with Online OCI Adjustment for Stream Processing ApplicationsabstractCheckpoint-based fault tolerant method has been widely used to enhance the reliability of Distributed Stream Processing Engines (DSPEs), but a checkpointing process usually introduces considerable overhead. It is a critical issue to choose the Optimal Checkpoint Interval (OCI) that maximizes the processing efficiency. Traditional OCI models consider the recovery time only related to the execution time from the last checkpoint to the moment of the failure. They are not suitable for stream processing jobs because the recovery time is related to the reprocessing workload, which depends on the realtime input data before a failure. A new model is needed to choose the OCI for stream processing applications. Moreover, the input data rate of an stream processing job fluctuates over time. The OCI of an application should also be adjusted dynamically according to the input workload. To solve these problems, we present a novel DSPS Optimal Checkpoint Interval (DOCI) model in this paper. We prove that it maximizes the processing efficiency for a given time period. We propose an approach to dynamically adjust the OCI for an application to accommodate the realtime workload fluctuations. We conduct simulation experiments to verify the effectiveness of DOCI model and the efficiency of the online OCI adjustment algorithm. Experimental results with a real-world dataset show DOCI achieves an improvement on system efficiency by up to 40%, comparing with existing fault-tolerant approaches. Yuan Zhuang 0003, Xiaohui Wei 0002, Hongliang Li 0003, Xubin He |
ICCCN | 2 |
| 2018 | Energy-Aware Allocation of Approximate Query Processing Over Data Streams with Error GuaranteeabstractWith increasing real-time and resource-intensive requirements, approximate computing is widely adopted to improve the performance of query processing over data streams. However, existing works concentrate on simple queries with single-step operations, such as point or join queries. There are a large number of nested queries with selection or filtering operations before aggregation. In this poster, we focus on approximate nested stream queries. We first propose a novel approximate model, SCM-sketches, that makes two-stage approximation for nested query answering with guaranteed errors. In the first stage for nested filtering operations, we use the sampling method to compress the arriving data. Then in the second stage, a sketch is used for further aggregation or join operations. We also theoretically analyze the effect of error propagation on approximate errors. Compared with existing sketch-based methods, experiment results with real-life datasets verify the effectiveness of SCM-sketches. Xiaohui Wei 0002, Yuanyuan Liu 0007, Shang Gao 0005, Xingwang Wang 0003 |
IWQoS | 1 |
| 2018 | An Online Approximate Stream Processing Framework with Customized Error ControlabstractIn online approximate stream processing, customers generally submit their requests with some specific quality requirements (e.g. maximum error). This raises a critical problem that online quality control is necessary to meet customized requirements. Since continuous arriving data needs to be processed immediately, it brings the difficulty of acquiring knowledge which significantly affects the efficiency of sampling. Hence, it's more challenging to ensure a prescribed level of quality without knowledge about data. In this paper, we present an adaptive approximate processing framework for online stream applications to address the challenges mentioned above. Specially, we first design a new data knowledge learning scheme to stratify the arriving stream data. Then, based on the online learning results, we propose a dynamic sampling strategy with the consideration of the stream rate. Finally, we further present a double-check error control mechanism to manage the output quality. Experiments with real world datasets show that the proposed approximate framework is not only applicable to different data distributions, but also provides a customized error control. Xiaohui Wei 0002, Yuanyuan Liu 0007, Xingwang Wang 0003, Shang Gao 0005 |
IWQoS | 1 |
| 2018 | Data Quality Aware Task Allocation Under a Feasible Budget in Mobile CrowdsensingabstractSatisfying spatial-temporal coverage requirement in the interested regions while considering the quality of the sensing data with budget limitation is a major research challenge in mobile crowdsensing. Most existing research in this field focus on the number of sensor readings collected in each covered subarea and do not consider individual differences of participants for contributing to data quality improvement. In this paper, we propose a novel coverage metric, quality coverage, which considers both the spatial coverage and the quality of sensing data and then use task allocation approaches to achieve highly diverse and spatial quality coverage level within a limited budget for different application scenarios. Xiaohui Wei 0002, Shang Gao 0005 |
IWQoS | 1 |
| 2018 | On pricing approximate queries
Xingwang Wang 0003, Xiaohui Wei 0002, Yuanyuan Liu 0007, Shang Gao 0005 |
Inf. Sci. | 2 |
| 2018 | A Task Allocation Method for Stream Processing with Recovery Latency Constraint
Hongliang Li 0003, Jie Wu 0001, Xiang Li 0197, Xiaohui Wei 0002 |
J. Comput. Sci. Technol. | 5 |
| 2017 | Task Allocation for Stream Processing with Recovery Latency GuaranteeabstractStream processing applications continuously process large amounts of online streaming data in real-time or near real-time. They have strict latency constraints, but they are also vulnerable to failures. Failure recoveries may slow down the entire processing pipeline and break latency constraints. Upstream backup is one of the most widely applied fault-tolerant schemes for stream processing systems. It introduces complex backup dependencies to tasks, and increases the difficulty of controlling recovery latencies. Moreover, when dependent tasks are located on the same processor, they fail at the same time in processor-level failures, bringing extra recovery latencies that increase the impacts of failures. This paper presents a correlated failure effect model to describe the recovery latency of a stream topology in processor-level failures for an allocation plan. We introduce a Recovery-latency-aware Task Allocation Problem (RTAP) that seeks task allocation plans for stream topologies that will achieve guaranteed recovery latencies. We present a heuristic algorithm with a computational complexity of O(nlog^2n) to solve the problem. Extensive experiments were conducted to verify the correctness and effectiveness of our approach. Hongliang Li 0003, Jie Wu 0001, Xiang Li 0197, Xiaohui Wei 0002 |
CLUSTER | 5 |
| 2017 | Parallel Regional Growth Marching Cubes: Efficient Surface Reconstruction Based on MPI
Xiaohui Wei 0002, Xinyan Bao, Xiaoli Pang, Haolong Cui |
ICIG (2) | 1 |
| 2017 | Integrated recovery and task allocation for stream processingabstractStream processing applications continueously process large-scale data streams online. The throughput of a stream processing application must match the input rate to avoid loss of data. Failures affect throughput because a task failure can suspend itself from producing new data and can even cause an application-level halt. The key motivation of this work is to mitigate the performance degradation caused by task-level failures. We introduce a novel Integrated Recovery Model (IRM) that allows resource sharing among both failure-free tasks and recovering tasks on a processor. The failure-free tasks slow down to accelerate a task recovery rather than suspending their actions and waiting for the recovery to finish; waiting causes a complete halt of the application. In this way, the recovery is seamless and does not suspend the entire system. The performance slowdown is related to both the failure-free processing cost and recovery cost on each processor. Moreover, the recovery cost of a task is related to the Fault-Tolerant Configuration (FTC) of the stream application. This paper introduces a novel task allocation problem that, given an FTC, can constrain processing performance during recoveries (i.e. throughput slowdown ratio) while minimizing the amount of resource occupied. We propose both a greedy algorithm and a heuristic algorithm with computational complexities of O(n log n) and O(n log2n), respectively, to solve the problem. Extensive experiments verify the correctness and effectiveness of our approach. Our approach enables continuous processing results and seamless failure recoveries with a constrained slowdown ratio. Hongliang Li 0003, Jie Wu 0001, Xiang Li 0197, Xiaohui Wei 0002, Yuan Zhuang 0003 |
IPCCC | 5 |
| 2017 | An indoor positioning approach using sibling signal patterns in enterprise WiFi infrastructureabstractThe indoor positioning technology plays an important role in the application scenarios requiring indoor location. In this paper, the WiFi signals under modern enterprise WiFi infrastructure and signal patterns between coexisting access points (APs) are investigated. Sibling signal patterns are defined and processed to generate Beacon APs that have higher confidence for positioning. Then a positioning approach using Beacon APs is proposed and shows improved positioning accuracy. The proposed schemes are fully designed, implemented and evaluated in a real-world environment, revealing its effectiveness and efficiency. Kun Yang 0001, Xiaohui Wei 0002 |
IWCMC | 4 |
| 2017 | Minimum Backups for Stream Processing With Recovery Latency GuaranteesabstractThe stream processing model continuously processes online data in an on-pass fashion that can be more vulnerable to failures than other big-data processing schemes. Existing fault-tolerant (FT) approaches have been presented to enhance the reliability of stream processing systems. However, the fundamental tradeoff between recovery latency and FT overhead is still unclear, so these scheme cannot provide recovery latency guarantees. This paper introduces the FT Configuration (FTC) problem and presents a solution for guaranteed recovery latency with minimum backups. A failure effect model is presented to describe the relationship between recovery latency and FTC (the amount and locations of backups). With this model, we design an algorithm to compute FTCs for different types of stream topologies according to recovery latency requirements. Extensive experiments are conducted to verify the correctness and effectiveness of our approach. We prove that our algorithm guarantees recovery latencies for all directed acyclic graph (DAG) stream topologies. For line(s) and tree topologies, our algorithm solves the FTC problem with a time complexity of O(N). For a general DAG topology, a heuristic function is used to generate FTCs. This causes fewer than 10% more backups on average compared to the optimal solution with a time complexity of O(N2). Hongliang Li 0003, Jie Wu 0001, Xiang Li 0197, Xiaohui Wei 0002 |
IEEE Trans. Reliab. | 5 |
| 2016 | Mining Electronic Physical Records, a trialabstractIn China, the rapid development of medical informatization and the increasing public health awareness greatly stimulate the growth of Electronic Physical Records (EPR). This paper focuses on analyzing and mining electronic physical records to provide constructive suggestions for people's lifestyles using data of thirty thousand individual records in 2009 within the nine-month period from April to December, at China-Japan Union Hospital of Jilin University. During the analyzing process, numerous challenges were faced such as data storage, data noise and data incompletion etc. Thus we designed a framework as the resolution to these problems, and utilized data mining methods including association rules to find the characteristics of all ages in Changchun and gained some satisfactory results such as the abnormal BMI increasing with age. We also utilized the community discovery algorithm to explore the relationship between abnormal physical indicators. However, the result is not satisfactory and the module cant even reach 0.1. To solve the problem of information isolation in China, we also implement a data platform. Xiaohui Wei 0002, Wenyang Zou, Shang Gao 0005 |
ASONAM | 1 |
| 2016 | Maximal Independent Sets in Heterogeneous Wireless Ad Hoc NetworksabstractIn ad hoc wireless networks, a Connected Dominating Set (CDS) has been extensively used as a Virtual Backbone (VB) for routing. The majority of approximation algorithms for constructing a small CDS in wireless ad hoc networks follow a general two-phased approach. The first phase is to construct a Dominating Set (DS) and the second phase is to connect the nodes in it. Generally, in the first phase, a Maximum Independent Set (MIS) is used as the DS. The relation between the size of a Maximum Independent Set and a Minimum Connected Dominating Set (MCDS) plays the key role in the performance analyses of these two-phased algorithms. In homogeneous wireless ad hoc networks modeled as Unit Disk Graphs (UDG) and Unit Ball Graphs (UBG), the relation between them has been well studied. However, in heterogeneous wireless networks which generally modeled as Disk Graphs with Bidirectional links (DGB) and Ball Graphs with Bidirectional links (BGB), upper bounds for the size of MISs have seldom been studied. In this paper, we give tighter upper bounds for the size of MISs in heterogeneous wireless ad hoc networks. When the maximum and minimum transmission range are relatively close, our result is much better. In DGB, when the transmission range ratio is (1,1.152], (1.152,1.307], (1.307,1.407], (1.407,1.462], (1.462,1.515], (1.515,1.618], (1.618,1.932], we prove that the size of any Maximal Independent Set (MIS) is upper bounded by 6opt + 1, 7opt + 1, 8opt + 1, 9opt + 1, 10opt + 1, 11opt + 1, 16.7778opt + 1.2222, where opt denotes the size of an optimal solution of the CDS problem. Sen Bai, Xiangjiu Che, Xin Bai 0004, Xiaohui Wei 0002 |
IEEE Trans. Mob. Comput. | 4 |
| 2016 | A Heuristic Clustering-Based Task Deployment Approach for Load Balancing Using Bayes Theorem in Cloud EnvironmentabstractAiming at the current problems that most physical hosts in the cloud data center are so overloaded that it makes the whole cloud data center'load imbalanced and that existing load balancing approaches have relatively high complexity, this paper has focused on the selection problem of physical hosts for deploying requested tasks and proposed a novel heuristic approach called Load Balancing based on Bayes and Clustering (LB-BC). Most previous works, generally, utilize a series of algorithms through optimizing the candidate target hosts within an algorithm cycle and then picking out the optimal target hosts to achieve the immediate load balancing effect. However, the immediate effect doesn't guarantee high execution efficiency for the next task although it has abilities in achieving high resource utilization. Based on this argument, LB-BC introduces the concept of achieving the overall load balancing in a long-term process in contrast to the immediate load balancing approaches in the current literature. LB-BC makes a limited constraint about all physical hosts aiming to achieve a task deployment approach with global search capability in terms of the performance function of computing resource. The Bayes theorem is combined with the clustering process to obtain the optimal clustering set of physical hosts finally. Simulation results show that compared with the existing works, the proposed approach has reduced the failure number of task deployment events obviously, improved the throughput, and optimized the external services performance of cloud data centers. Jia Zhao 0003, Kun Yang 0001, Xiaohui Wei 0002, Yan Ding 0001, Liang Hu 0001, Gaochao Xu |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2014 | MapReduce delay scheduling with deadline constraintabstractSUMMARY MapReduce programming paradigm has been widely applied to solve large‐scale data‐intensive problems. Intensive studies of MapReduce scheduling have been carried out to improve MapReduce system performance. Delay scheduling is a common way to achieve high data locality and system performance. However, inappropriate delays can lead to low system throughput and potentially break the original job priority constraints. This paper proposes a deadline‐enabled delay (DLD) scheduling algorithm that optimizes job delay decisions according to real‐time resource availability and resource competition, while still meets job deadline constraints. Experimental results illustrate that the resource availability estimation method of DLD is accurate (92%). Compared with other approaches, DLD reduces job turnaround time by 22% in average while keeping a high locality rate (88%).Copyright © 2013 John Wiley & Sons, Ltd. Hongliang Li 0003, Xiaohui Wei 0002, Qingwu Fu |
Concurr. Comput. Pract. Exp. | 2 |
| 2014 | Topology-Aware Partial Virtual Cluster Mapping Algorithm on Shared Distributed InfrastructuresabstractNovel virtualized HPC centers provide isolated and configurable Virtual Clusters (VC) on shared distributed infrastructures as execution environments for parallel and distributed applications. These VCs are usually customized and deployed per job in runtime. Allocating physical resources for VC is known as Virtual Cluster Mapping (VCM) problem, which is a critical issue that affects both performance of the VC and resource utilization of the system. Most previous works treat all Virtual Machines (VMs) in a VC request equally. However, because sub-jobs in a parallel job usually perform different roles, the corresponding VMs in a VC that execute these sub-jobs respectively should have different levels of importance. Based on this argument, this paper introduces the concept of partial VC mapping in contrast to the full mapping methodology in the current literatures. To fulfill partial mapping, the important backbone communication structure of parallel job called Communication Skeleton (CS) is derived based on the network topology among virtual nodes. To generate the CS of a job, mechanisms for evaluating the importance of nodes are proposed. Eventually, a Topology-aware Partial Virtual Cluster Mapping algorithm (TOP-VCM) is proposed which is based on sub-graph isomorphism detection. TOP-VCM can fully satisfy the nodes/links requirements in CS to ensure the execution performance with only slight degradation of other trivial nodes/links to significantly reduce the mapping difficulty. Simulation results have shown that TOP-VCM has significantly improved the total revenue, the utilization of physical resources and the performance of mapping algorithm while satisfying the VC requirements. Xiaohui Wei 0002, Hongliang Li 0003, Kun Yang 0001, Lei Zou 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2008 | Implement the Grid Workflow Scheduling for Data Intensive Applications with CSF4abstractGrid computing technology is able to integrate and share large-scale distributed computation and data resource to facilitate the scientific researches. Recently, the grid workflow support and large-scale distributed data management are becoming two main requirements of scientists and researchers in many fields, such as bioinformatics, high-energy physics etc. In this paper, we proposed to support grid workflow for data intensive applications using CSF4 scheduling plug-ins. The grid workflow scheduling and data aware scheduling policies are implemented in two scheduling plug-ins, grid workflow plug-in and grid data aware plug-in, respectively. The two scheduling plug-ins can work together smoothly. The data aware plug-in will automatically dispatch the workflow tasks to the grid sites which are close to data replicas. At last, the experiment results are given to show the improvement of system performance and optimization of scheduling. Zhaohui Ding, Xiaohui Wei 0002, Yaoguang Yuan, Wilfred W. Li, Osamu Tatebe |
eScience | 2 |
| 2006 | Building Cyberinfrastructure for Bioinformatics Using Service Oriented Architecture
Wilfred W. Li, Sriram Krishnan, Kurt Mueller, Kohei Ichikawa, Susumu Date, Sargis Dallakyan, Michel F. Sanner, Chris Misleh, Zhaohui Ding, Xiaohui Wei 0002, Osamu Tatebe, Peter W. Arzberger |
CCGRID | 10 |
| 2005 | Integrating Local Job Scheduler - LSFTM with GfarmTM
Xiaohui Wei 0002, Wilfred W. Li, Osamu Tatebe, Gaochao Xu, Liang Hu 0001, Jiubin Ju |
ISPA | 1 |
| 2000 | SFT: A Consistent Checkpointing Algorithm with Short Freezing Time
Xiaohui Wei 0002, Jiubin Ju |
J. Comput. Sci. Technol. | 1 |
| 2000 | SCR Algorithm: Saving/Restoring States of File Systems
Xiaohui Wei 0002, Jiubin Ju |
J. Comput. Sci. Technol. | 1 |