EDBT 2026 Demo / reviewers in the wild / expert
Xiaojun Zhai
dblp:70/9455
· DBLP profile ↗
46ranked-venue papers
5as first author
31since 2021 · last 2026
0000-0002-1030-8311ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 3 first-author · 15 since 2021Computer networks · 10 · 9 since 2021Software engineering, systems software and programming languages · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Security and privacy · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Energy-Efficient and Reliable Task Mapping and Offloading for Multicore Edge Devices With DVFSabstractMulticore platforms based on NoC are promising architectures for safety-critical applications. Application execution performance is determined by task mapping, with reliable execution, real-time response, and energy efficiency as requirements. We can perform task duplication, DVFS, and multipath routing to meet these requirements during task mapping. Furthermore, the computation platforms have limited computation capacity and energy supply in several application domains. Some complex tasks can be offloaded from the edge device to the cloud for execution. However, such task offloading influences task mapping on the edge device. Existing approaches seldom consider the correlation of task offloading to the cloud and task mapping on the edge device. To address this limitation, we jointly consider task mapping inside the NoC-based multicore edge device and task offloading to the cloud to optimize energy consumption while satisfying reliability and real-time constraints. This problem is formulated as a mixed-integer nonlinear programming and linearized to find the optimal solution. We propose a novel three-step heuristic with a feedback mechanism to enhance task schedulability and reduce computation time. We evaluate the behavior of our approaches through exhaustive simulations. The results show that our approaches outperform existing methods in terms of energy efficiency, task reliability, and schedulability. Lei Mo, Tamim M. Al-Hasan, Angeliki Kritikakou, Xiaojun Zhai, Olivier Sentieys, Shibo He |
IEEE Internet Things J. | 5 |
| 2026 | Mitigating Scalability Challenges in LUT-Based Neural Networks via Pruning OptimisationsabstractModern deep neural networks heavily rely on a large number of multiply-accumulate operations, which constitute the predominant computational cost. To address this, Look-Up Table (LUT)-based matrix multiplications have emerged as a promising alternative for reducing the computational cost and time of the multiply-accumulate operations in a neural network. However, the LUT-based neural network still faces the scalability challenge due to the inherent limitations of LUT-based matrix multiplication. To mitigate these scalability limitations, this paper proposes a scalable and energy-efficient LUT-based approximate matrix multiplication unit (LUT-MU) constituting the basic component of the neural networks by integrating a pruning strategy on the MADDNESS algorithm, a LUT-based matrix multiplication methodology. With increasing problem size and precision demands in matrix multiplication, our proposed LUT-MU architecture effectively constrains resource expansion. The case study shows that deploying our LUT-MU in neural network architectures, including fully connected layers (MNIST) and ResNets (CIFAR-10, ImageNet)—on XCZU7EV and XCZU19EG FPGAs, produces up to 1.6× throughput improvement and 4.2× energy efficiency gains over mainstream CUDA-based network implementations, and 1.8× energy efficiency compared to leading quantised neural network implementations, with moderate impact on accuracy. Compared to original MADDNESS-based neural networks, our LUT-MU shows 1.3 to 2.6× resource savings based on various resolution configuration settings of MADDNESS. Xuqi Zhu, Huaizhi Zhang, Chandrajit Pal, Sangeet Saha, Klaus D. McDonald-Maier, Xiaojun Zhai |
IEEE Trans. Computers | 8 |
| 2026 | QoS-Aware Approximate Task Mapping on Heterogeneous Multicore Platforms with DVFS and Task MigrationabstractHeterogeneous Multicore Platforms (HMPs) have been widely adopted to execute tasks across a range of applications. Under limited system resources and diverse application requirements, allocating and executing dependent Approximate Computing (AC) tasks on these platforms to achieve high Quality-of-Service (QoS) is challenging. Dynamic Voltage and Frequency Scaling (DVFS) and task migration have proven effective for improving QoS while balancing time and energy consumption. However, existing approaches often overlook the migration overhead and the resulting dynamic changes in task dependencies, which can adversely affect mapping outcomes. To address these issues, this article presents a novel AC task mapping method that maximizes system QoS under multiple constraints on HMPs, accounting for task migration overhead, DVFS, and changes in Directed Acyclic Graph (DAG) topology. We first formulate this joint design problem as a complex nonlinear programming problem. Next, we linearize the nonlinear terms without performance loss by introducing auxiliary variables and additional constraints. Building on this formulation, we propose an optimal (OPT) and a low-complexity Heuristic Algorithm (HEU), derived from problem decomposition and a greedy strategy, which divides the Mixed-Integer Non-Linear Programming (MINLP) problem into two smaller subproblems with fewer variables and constraints, solving them sequentially. The simulation results show that the proposed OPT method achieves higher QoS performance, measured at about 2.389 times on average and up to 4.115 times, while its feasibility is increased to about 3.263 times on average and up to 9.667 times, compared to other state-of-the-art methods. In addition, the average QoS of the proposed HEU method is about 0.577 times that of the proposed method, but its computation time is over a thousand times shorter. Hengyan Song, Lei Mo, Tamim M. Al-Hasan, Angeliki Kritikakou, Xiaojun Zhai, Shibo He, Olivier Sentieys |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2025 | Late Breaking Results: Approximated LUT-Based Neural Networks for FPGA Accelerated InferenceabstractThis work presents LUT-MU, an approximated LUT-based Matrix Multiplication (MM) architecture designed for FPGA-based Neural Network (NN) inference across. The proposed architecture maximises the utilisation of on-chip memory bandwidth through dedicated memory distribution and pipeline design, addressing performance limitations inherent to LUT-based MM. Experimental evaluation demonstrates that LUT-MU achieves a four-fold improvement in NN inference throughput whilst reducing hardware resource consumption by 80% with only a 5% decline in accuracy. These results validate that our optimisation approach successfully resolves the performance constraints caused by the limited arithmetic intensity and memory bandwidth, enabling the LUT-MU to serve as a foundation for efficient NN acceleration systems. Xuqi Zhu, Huaizhi Zhang, Tamim M. Al-Hasan, Klaus D. McDonald-Maier, Xiaojun Zhai |
DATE | 6 |
| 2025 | Wireless Single-Camera Markerless Motion Capture System for Healthcare ApplicationsabstractSingle-camera markerless systems have emerged as a robust methodology for human motion capture and rehabilitation applications. Traditional methodologies typically necessitate multiple strategically positioned cameras or special equipment, including sensors to capture patient ambulatory motion, requiring preliminary calibration and synchronization procedures, which may incur significant costs. This paper presents a wireless single-camera markerless framework for rehabilitation applications that leverages advanced deep learning (DL) architectures to estimate and extract three-dimensional skeletal coordinates from monocular camera views of ambulatory patients. The extracted skeletal representation is subsequently transmitted across wireless communication channels. Then, the rendering technique has been applied for displaying virtual movement of the patient for privacy enhancement. Simulations demonstrate the effectiveness of the framework while maintaining motion assessment capabilities, presenting opportunities for deployment in remote healthcare monitoring scenarios. Areej Athama, Apoorva Srivastava, Shengyang Huang, Kezhi Wang, Yongmin Li 0001, Xiaojun Zhai |
HPCC | 6 |
| 2025 | Computationally Efficient FPGA-based Large Language Model Inference for Real-Time Decision-Making in Robotic SystemsabstractIntegrating Large Language Models (LLMs) into modern robotic systems presents significant computational and energy constraint challenges, particularly for human-centered robotic applications. This paper presents a novel hardware optimization technique for deploying LLMs on resource-constrained embedded devices, achieving an up to 77% reduction in computational latency through an FPGA implementation in comparison to other popular embedded computing devices (e.g., CPU and GPUs). Additionally, we demonstrate our methodology by deploying a LLaMA 2-7B model on a Unitree Go2 robotic dog integrated with the proposed FPGA platform. The proposed optimization framework preserves real-time interaction capabilities while significantly reducing computational and energy overhead, facilitating efficient natural language processing for human-robot interaction in safety-critical and dynamic environments. Experimental results demonstrate that the FPGA-based LLaMA 2-7B implementation achieves up to 6.06-fold and 1.95-fold higher throughput compared to baseline CPU and GPU implementations while maintaining comparable inference accuracy. Furthermore, the proposed FPGA design surpasses existing state-of-the-art FPGA implementations, delivering a 30% improvement in computational efficiency. Huaizhi Zhang, Tamim M. Al-Hasan, Xuqi Zhu, Weiyong Si, Klaus D. McDonald-Maier, Xiaojun Zhai |
IROS | 7 |
| 2025 | A Privacy-aware Quantilisation Approach for Efficient Edge Deep Learning AcceleratorabstractData privacy is one of the key concerns in machine learning model applications at the edge, especially in sensitive domains such as the healthcare sector. Here adversaries may exploit Membership Inference Attacks (MIAs) to determine if particular data points were used as part of datasets from the model’s training set, potentially leading to further data leakage issues. Although privacy preservation techniques like differential privacy (DP) can mitigate such risks during the training phase, this often results in degradation of model accuracy, making them less suitable for cloud training and edge deployment paradigms. For AI edge applications, existing research works for designing neural network accelerators primarily prioritize computational performance and power efficiency as their design target. In this paper, we introduce a novel privacy-aware quantilisation approach for deep learning accelerators and analyse the tradeoff between computational efficiency and privacy protection. The proposed system allows to adjust privacy constraints through tunable parameters, enabling flexible deployment on edge devices while meeting privacy and performance constraints. We have evaluated the proposed design on an AMD VCK190 board using a range of hypothetical MIA benchmarks. The results demonstrate that the proposed approach can effectively reduce the success rate of MIA attacks across multiple performance metrics. Huaizhi Zhang, Xuqi Zhu, Klaus D. McDonald-Maier, Xiaojun Zhai |
ISCAS | 5 |
| 2025 | Reliable and Energy Optimized Task Mapping for Heterogeneous Multicore NoC Based on Partial Task Duplication and Multipath RoutingabstractThe increasing integration of heterogeneous processors on a chip presents significant challenges for efficient management in Multi-Processor System-on-Chip (MPSoC) platforms. Network-on-Chip (NoC) architectures offer a flexible and scalable interconnection paradigm through router-based communication. However, mapping dependent, real-time tasks in NoC environments critically affects data processing and transmission efficiency. An optimized task mapping scheme must address constraints such as real-time deadlines, energy consumption, and reliability, which are key metrics for modern NoCs. Existing approaches often overlook the complex interplay between communication paths and their associated energy costs, resulting in suboptimal resource utilization. This paper proposes a comprehensive task mapping framework that jointly optimizes energy efficiency and reliability by integrating Dynamic Voltage and Frequency Scaling (DVFS), multi-path data routing, task allocation, scheduling, and partial task duplication. We formulate the problem as a complex combinatorial optimization task and transform it into a solvable form with reduced computational complexity. Simulation results demonstrate that the proposed method achieves superior energy efficiency by reducing energy consumption by up to 39.7%, reducing computation time, and improving task schedulability compared to existing state-of-the-art approaches. Lei Mo, Tamim M. Al-Hasan, Minyu Cui, Xiaojun Zhai, Qing Gao 0001, Shibo He |
IEEE Internet Things J. | 5 |
| 2025 | RENOWNED: A Real-Time Anomaly Detection and Mitigation Framework in Edge-Enabled IoVabstractThe rapid adoption of smart vehicles and their interconnection through the Internet of Vehicles (IoV) has increased the use of electronic control units (ECUs) in cars. These ECUs, while enabling advanced features, also present a larger target for cyberattacks, which can disrupt critical functions and jeopardize safety. The time-sensitive nature of automotive systems necessitates swift responses, making the protection of ECUs crucial. The imprecise computation (IC) task model can mitigate the risk of task completion failures by generating acceptable approximation results within deadlines when achieving absolute accuracy becomes difficult within fixed deadlines and energy budgets. This article introduces RENOWNED, a solution that ensures the normal functioning of these controller area networks (CAN) controlled ECUs even in the face of anomalies. It combines anomaly detection and mitigation through the HEALING module to maintain the desired performance. The anomaly detection module uses graph attention networks (GAT) to identify unusual processor behavior. If an anomaly is detected, the HEALING module takes over, reallocating tasks based on the available resources to guarantee that deadlines are met and energy constraints are not exceeded. Experiments have shown that RENOWNED delivers a Quality of Service (QoS) of 25% to 64% when system utilisation is varied in the range from 40% to 90%. It exhibits an excelling performance in detecting anomalies, achieving a 97.6% accuracy even when the magnitude mixed anomaly signals are very minute. Thus, our proposed RENOWNED offers a robust way to enhance the reliability and energy efficiency of safety-critical automotive applications prevalent in IoV. Chandrajit Pal, Sangeet Saha, Xiaojun Zhai, Klaus D. McDonald-Maier |
IEEE Internet Things J. | 3 |
| 2025 | Modeling and Risk Analysis of Cooperative Adaptive Cruise Control Systems Based on Petri Nets and Distributed Edge IntelligenceabstractFueled by advancements in intelligent transportation systems, the Internet of Vehicles (IoV) seeks to connect smart vehicles, road infrastructure, and users into a unified network, enhancing traffic efficiency and reducing accident risks. Centralized cloud data collection raises concerns about privacy and communication overhead. To address these, distributed edge intelligence (DEI) reduces transmission costs and improves privacy by implementing machine learning at the network edge. In this context, cooperative adaptive cruise control (CACC) systems, combined with DEI in the IoV framework, enhance transportation system intelligence through real-time data processing and decentralized decision making. This article proposes a modeling and analysis method for CACC systems based on Petri nets. The datasets are automatically generated using tools developed by our team, and machine-learning methods are utilized to perform risk prediction analysis on the CACC model. From the perspective of Petri nets synchronization, we propose risk mitigation strategies from a design standpoint. The research results show that the proposed method significantly reduces signal accumulation and enhances synchronization in CACC systems. This improvement provides new theoretical support and technical guidance for the design and implementation of CACC systems, ultimately enhancing their safety and reliability. Wangyang Yu 0001, Yumeng Cheng, Xianwen Fang, Xiaojun Zhai, Hongyuan Jing |
IEEE Internet Things J. | 4 |
| 2025 | Formal Modeling of Hybrid System Based on Semi-continuous Colored Petri Net: A Case Study of Adaptive Cruise Control SystemabstractMany Next-Generation consumer electronic devices would be distributed hybrid electronic systems, such as UAVs (Unmanned Aerial Vehicles) and smart electronic cars. The safety and risk control are the key issues for the sustainability of such consumer electronic systems. The modeling of hybrid electronic systems is difficult to be abstracted by traditional Petri Nets. This also makes the reachable marking graph unable to be applied to Petri Nets of the hybrid electronic systems. This paper proposes a novel Petri Net to model and analyze the hybrid electronic systems. We name it a Semi-continuous Colored Petri Net (SCPN) that inherits the excellent modeling capabilities and analysis methods of Petri Nets, and can formally depict hybrid quantities. In addition, we propose the construction algorithm for an SCPN reachable marking graph and prove its finiteness. Finally, we model and analyze an Adaptive Cruise Control (ACC) system of smart electronic cars as an example to prove the validity of SCPN. We use the proposed SCPN to model and analyze the running process of an ACC system under the continuous deceleration scenario of the front vehicle. The application study shows that the ACC system has logic flaws under the constant headway strategy when the front vehicle continues to decelerate. Based on this analysis, improvements to the SCPN of the ACC system are made, effectively enhancing its safety and logical correctness. Wangyang Yu 0001, Yumeng Cheng, Lu Liu 0001, Fei Hao 0001, Xiaojun Zhai, Minsi Chen |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2025 | Contention and Reliability-Aware Energy Efficiency Task Mapping on NoC-Based MPSoCsabstractRecently, network-on-chip (NoC)-based multiprocessor system-on-chips (MPSoCs) have become popular computing platforms for real-time applications due to high communication performance and energy efficiency over traditional bus-based MPSoCs. Due to the nature of network structures, network congestion along with transient faults, can significantly affect communication efficiency and system reliability. Most existing works have rarely focused on the concurrent optimization of network contention, reliability, and energy consumption. Here, we study the problem of contention and reliability-aware task mapping under real-time constraints for dynamic voltage and frequency scaling-enabled NoC. The problem entails optimizing voltage/frequency on cores and links to reduce energy consumption and ensure system reliability, while task mapping and slack time are adopted to alleviate network contention and reduce latency. We aim to minimize computation and communication energy and balance workload. This problem is formulated as a mixed-integer nonlinear programming, and we present an effective linearization scheme that equivalently transforms it into a mixed-integer linear programming to find the optimal solution. To reduce computation time, we propose a three-step heuristic, including task allocation, frequency scaling and edge scheduling, and communication contention management. Finally, we perform extensive simulations to evaluate the proposed method. The results show we can achieve 31.6% and 21.7% energy savings, with 95.5% and 98.6% less contention than the existing methods. Lei Mo, Xinmei Li, Angeliki Kritikakou, Xiaojun Zhai |
IEEE Trans. Reliab. | 4 |
| 2025 | APPARENT: AI-Powered Platform Anomaly Detection in Edge ComputingabstractEmbedded systems serving as IoT nodes are often vulnerable to malicious and unknown runtime software that could compromise the system, steal sensitive data, and cause undesirable system behaviour. Commercially available embedded systems used in automation, medical equipment, and automotive industries, are especially exposed to this vulnerability since they lack the resources to incorporate conventional safety features and are challenging to mitigate through conventional approaches. We propose a novel system design coined as APPARENT which can identify program characteristics by monitoring and counting the maximum possible low-level hardware events from Hardware Performance Counters (HPCs) that occur during the program's execution and analyse the correlation among the counts of various monitored events. To further utilise these captured events as features we propose a self-supervised machine learning algorithm that combines a Graph Attention Network GAT and a Generative Topographic Mapping GTM to detect unusual program behaviour as anomalies to enhance the system security. Our proposed methodology takes advantage of attributes like program counter, cycles per instruction, and physical and virtual timers at various exception levels of the embedded processor to identify abnormal activity. APPARENT identifies unknown program behaviours not present in the training phase with an accuracy of over 98.46% on Autobench EEMBC benchmarks. Chandrajit Pal, Sangeet Saha, Xiaojun Zhai, Gareth Howells 0001, Klaus D. McDonald-Maier |
IEEE Trans. Sustain. Comput. | 3 |
| 2025 | A Distributed Data-Driven and Machine Learning Method for High-Level Causal Analysis in Sustainable IoT SystemsabstractA causal relationship forms when one event triggers another's change or occurrence. Causality helps to understand connections among events, explain phenomena, and facilitate better decision-making. In IoT systems, massive consumption of energy may lead to specific types of air pollution. There are causal relationships among air pollutants. Analyzing their interactions allows for targeted adjustments in energy use, like shifting to cleaner energy and cutting high-emission sources. This reduces air pollution and boosts energy sustainability, aiding sustainable development. This paper introduces a distributed data-driven machine learning method for high-level causal analysis (DMHC), which extracts general and high-level Complex Event Processing (CEP) rules from unlabeled data. CEP rules can capture the interactions among events and represent the causal relationships among them. DMHC deploys a two-layer LSTM attention mechanism model and decision tree algorithm to filter and label data, extracting general CEP rules. Afterward, it proceeds to generate event logs based on general rules with heuristic mining (HM), extracting high-level CEP rules that pertain to causal relationships. These high-level rules complement the extracted general rules and reflect the causal relationships among the general rules. The proposed high-level methodology is validated using a real air quality dataset. Wangyang Yu 0001, Jing Zhang 0024, Lu Liu 0001, Xiaojun Zhai, Ruhul Kabir Howlader |
IEEE Trans. Sustain. Comput. | 5 |
| 2024 | Evaluating Lightweight GAN- and Adapted CTGAN-Based Data Synthesis for Predictive Maintenance in High-Radiation EnvironmentsabstractThis paper presents a comparative analysis of two developed Generative Adversarial Network (GAN) architectures for synthesizing sensor data in predictive maintenance (PdM) applications within high-radiation environments. The study ad-dresses the challenge of data scarcity in such settings, where experimental runs are constrained by the risk of device failure and economic considerations. The two GAN models: GAN-1 uses the Conditional Tabular GAN (CTGAN) architecture, and GAN-2 employs a custom network. These models generated synthetic datasets that were used to train and evaluate three machine learning algorithms: Random Forest, k-Nearest Neighbours, and eXtreme Gradient Boosting. The performance of these PdM models trained on synthetic data was compared against models trained on the original limited dataset. Results demonstrate that GAN-1 produced synthetic data closely mirroring the characteristics of the original dataset, enabling PdM models to achieve comparable performance levels. This study highlights the potential of GAN-based data synthesis in enhancing PdM model development for high-radiation environments, offering a viable solution to the challenges of limited data availability in such harsh settings. The findings have significant implications for improving operational reliability and safety in nuclear and other extreme environments where electronic systems are deployed. Tamim M. Al-Hasan, Xiaojun Zhai, Klaus D. McDonald-Maier, Faycal Bensaali, Alice Cryer |
BDCAT | 2 |
| 2024 | Enhancing security in e-business processes: Utilizing dynamic slicing of Colored Petri Nets for logical vulnerability detection
Wangyang Yu 0001, Lu Liu 0001, Xiaojun Zhai, Yumeng Cheng |
Future Gener. Comput. Syst. | 4 |
| 2024 | Energy Minimization of RIS-Assisted Cooperative UAV-USV MEC NetworkabstractUnmanned surface vehicles (USVs) are becoming increasingly significant in fulfilling integrated sensing, computing, and communication with the emergence of bidirectional computation tasks. However, Quality-of-Service provisioning is still challenging since USVs are restricted with limited onboard resources and direct links between them and shore-based terrestrial base stations (TBSs) are frequently blocked. This article proposes a novel reconfigurable intelligent surface (RIS)-assisted cooperative unmanned aerial vehicle (UAV)–USV mobile-edge computing (MEC) network architecture, where RIS-mounted tethered UAV (TUAV) and rotary-wing UAVs (RUAVs) are collaboratively utilized to serve USVs. RUAVs energy minimization is formulated by jointly considering TUAV hovering altitude, RIS phase-shift vector, RUAV service selection indicator, and RUAVs turning points. A heuristic solution is proposed to tackle the formulated problem, where the original problem is first decoupled into three subproblems, e.g., the joint optimization of RIS phase-shift vector and TUAV hovering altitude subproblem, RUAVs service selection indicator subproblem, and RUAVs turning points subproblem, each of which is solved by the proposed modified alternative direction method of multiplier (ADMM) algorithm, the proposed enhanced simulated annealing (ESA) algorithm and the proposed successive convex approximation (SCA)-based algorithm. In this way, the challenging problem can be efficiently solved iteratively. The results show that the proposed solution can decrease RUAVs energy consumption by nearly 29% compared to numerous selected advanced algorithms. Moreover, the performance of the proposed solution regarding typical penalty coefficients and number of RIS reflecting elements is investigated. Yangzhe Liao, Yuanyan Song, Si-Yu Xia, Yi Han 0007, Ning Xu 0006, Xiaojun Zhai |
IEEE Internet Things J. | 6 |
| 2024 | Low-Latency Data Computation of Inland Waterway USVs for RIS-Assisted UAV MEC NetworkabstractUnmanned Surface Vehicles (USVs) in inland waterways have drawn increasing attention for their excellent capability to serve maritime time-consuming missions such as autonomous navigation and intelligent monitoring. However, USVs struggle to accomplish emerging computation-intensive tasks (e.g., sensor, telemetry, etc) timely due to the limited on-board resources. This paper proposes a novel reconfigurable intelligent surface (RIS)-assisted unmanned aerial vehicle (UAV) multi-access edge computing (MEC) network architecture to support low-latency USVs data computation with time window. Aiming to enhance USVs task processing efficiency, the minimization of USVs task processing time is formulated by jointly considering UAVs flight route selection, USVs execution mode selection, UAVs hovering coordinates and RIS phase shift vector. A heuristic solution is proposed to tackle the formulated challenging problem iteratively. The original problem is decoupled into three subproblems: an enhanced deferred acceptance algorithm is proposed to solve UAVs flight route selection subproblem; an enhanced Lagrangian relaxation method is proposed to solve USVs execution mode selection subproblem; a joint alternating direction method of multipliers (ADMM)-successive convex approximation (SCA)-based algorithm is proposed to solve UAVs hovering coordinates subproblem. Experiment results demonstrate that the proposed solution can decrease task processing time by approximately 54% compared with numerous selected advanced algorithms. Moreover, the performance of the proposed solution under typical UAVs caching capability and the number of UAVs has been investigated. Yangzhe Liao, Yuanyan Song, Yi Han 0007, Ning Xu 0006, Xiaojun Zhai, Zhenhui Yuan |
IEEE Internet Things J. | 6 |
| 2024 | Modeling and Analysis of ETC Control System with Colored Petri Net and Dynamic SlicingabstractNowadays, Electronic Toll Collection (ETC) control systems have been widely adopted to smoothen traffic flow on highways. However, as it is a complex business interaction system, there are inevitably flaws in its control logic process, such as the problem of vehicle fee evasion. We find that there is more than one way for vehicles to evade fees. This shows that it is difficult to ensure the completeness of its design. Therefore, it is necessary to adopt a novel formal method to model and analyze its design, detect flaws, and modify it. In this article, a Colored Petri net (CPN) is introduced to establish its model. To analyze and modify the system model more efficiently, a dynamic slicing method of CPN is proposed. First, a static slice is obtained from the static slicing criterion by backtracking. Second, considering all binding elements that can be enabled under the initial marking, a forward slice is obtained from the dynamic slicing criterion by traversing. Third, the dynamic slicing of CPN is obtained by taking the intersection of both slices. The proposed dynamic slicing method of CPN can be used to formalize and verify the behavior properties of an ETC control system, and the flaws can be detected effectively. As a case study, the flaw about a vehicle that has not completed the payment following the previous vehicle to pass the railing is detected by the proposed method. Wangyang Yu 0001, Jinming Kong, Zhijun Ding, Xiaojun Zhai, Zhiqiang Li 0003 |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2023 | Bayesian Optimization for Efficient Heterogeneous MPSoC Based DNN Accelerator Runtime TuningabstractWith the explosive growth of Internet of Things (IoT) devices and applications, deploying Deep Neural Networks (DNNs) on resource-constrained embedded edge devices has become a popular research trend. Because such systems have limited resources, they need to rely on optimising resource utilisation to meet performance requirements. However, for scenarios where the DNN application and workloads are dynamically changing, the offline system optimisation technique cannot achieve optimal runtime performance in practical environments. Hence, in this PhD project, we propose a Bayesian Optimisation (BO)-based runtime tuning scheme for improving energy efficiency of heterogeneous MPSoC-based DNN accelerator in the context of DNN applications. By seeking suitable hardware configurations of the accelerator for dynamic DNN inference workloads ranging from 200 M to 600 M FLOPs (floating-point operations) at runtime, the recommended configuration can averagely save up to 15.33% energy consumption from a random configuration setting. Xuqi Zhu, Sangeet Saha, Xiaojun Zhai, Klaus D. McDonald-Maier |
FPL | 4 |
| 2023 | Grading and Calculation of Synchronic Distance in Petri Nets for Trustworthy Modeling and analyzingabstractSynchronization plays a crucial role in computer systems, providing support for system security, data consistency, and coordination. It contributes to the establishment and application of trust, security, and dependability in distributed systems and concurrent computing to a significant extent. This article makes innovative contributions in the field of synchronic distance in Petri net. We provide refined definitions for the hierarchical classification of synchronic levels in Petri net, proposing the concepts of absolute synchronization, strong synchronization, and extended synchronization based on different conditions. Furthermore, we propose an innovative method for calculating synchronic distance. This method can automate the calculation of synchronic distance between any two transitions using computer computation, resulting in improved accuracy and reduced errors. This novel approach provides an effective tool for system security and trustworthy modeling, as accurate synchronic distance calculations allow for better evaluation of synchronic distance between different transitions, leading to the identification of potential security vulnerabilities and design flaws, thereby enhancing the credibility of decision-making and promoting the reliability of models and analysis results. To validate the proposed method, we introduce a specific example of a Petri net with concurrency, demonstrate the practicality and effectiveness of the proposed method and algorithm through analysis of this example. Our work extends the research on Petri net synchronic distance, further advancing the understanding and exploration of this field. Yumeng Cheng, Wangyang Yu 0001, Xiaojun Zhai, Fei Hao 0001 |
TrustCom | 3 |
| 2022 | Energy Minimization for IRS-assisted UAV-empowered Wireless CommunicationsabstractNon-terrestrial wireless communications have evolved into a technology enabler for seamless connectivity and ubiquitous computing services in the beyond fifth-generation (B5G) and sixth-generation (6G) networks, aiming to provision reliable and energy efficient communications among aerial platforms and ground mobile users. This paper considers intelligent reflecting surface (IRS)-assisted unmanned aerial vehicle (UAV)-empowered wireless communication, which exploits both the high mobility of UAV and passive beamforming gain brought by IRS. The energy minimization of rotary-wing UAV is formulated by jointly considering numerous quality of service (QoS) constraints with intricately coupled variables. To tackle the formulated challenging problem, a heuristic algorithm is proposed. First, we decouple it into several subproblems. Moreover, we jointly investigate offloading decisions of Internet of Thing (IoT) devices by the proposed enhanced differential evolution algorithm. Then, minorization-maximization algorithm (MMA) is utilized to solve the optimization of IRS phase shift-vector. Moreover, ant colony optimization (ACO) algorithm is proposed to optimize UAV flight route indicator matrix. Numerical results validate the effectiveness of the proposed algorithm. The results show that the proposed solution can remarkably decrease UAV flight distance while improving the network energy efficiency in comparison with numerous advanced algorithms. Yangzhe Liao, Jiaying Liu 0011, Yi Han 0007, Qingsong Ai, Quan Liu 0001, Xiaojun Zhai |
MSN | 7 |
| 2022 | UAV Swarm Trajectory and Cooperative Beamforming Design in Double-IRS Assisted Wireless CommunicationsabstractNon-terrestrial communications have emerged as a technological enabler for seamless connectivity and ubiquitous computation services in the upcoming beyond fifth generation (B5G) and sixth generation (6G) networks. However, there exist numerous practical technical limitations, such as high deployment cost, massive energy consumption, high probability of information transmission blockage and dynamic propagation environments and so forth. Thanks to the rapid developments of meta-materials, the cost-effective and energy-efficiency intelligent reconfigurable surface (IRS) has been globally recognized as a revolutionized technology to construct smart radio environments. In this paper, a novel double-IRS assisted unmanned aerial vehicles (UAV)-swarm-enabled communication network architecture is proposed, where two UAV swarms are integrated with the main IRS reflector and subreflector, respectively. The energy minimization problem of UAV swarm carried main IRS is formulated, subject to a list of quality of service (QoS) constraints. To tackle the formulated challenging problem, we first decouple the original problem into two subproblems. Then, a heuristic algorithm is proposed, where the enhanced differential evolution (DE) algorithm is proposed to optimize the UAV swarm trajectory and the alternate optimization algorithm is utilized to optimize the cooperative reflect beamforming vector. Numerical results validate that the proposed algorithm outperforms several selected advanced algorithms regarding UAV swarm energy consumption. Moreover, the network performance under the different number of IRS elements is investigated. Yangzhe Liao, Xiaojun Zhai |
MSN | 4 |
| 2022 | Benchmark Tool for Detecting Anomalous Program Behaviour on Embedded DevicesabstractThis paper presents an open-source benchmark tool for anomaly detection in program behaviour, using program counter (PC) and instruction type information. It is introducing anomalies in artificial way, allowing for fine-grained evaluation with adjustable sliding window sizes and preprocessing configuration. The usage of the benchmark, including demonstrated data collection, does not require any additional hardware other than a standard computer. The benchmark uses the output of llvm-objdump program to focus on non-library code which allows for rapid evaluation of various detection methods with different configurations. The proposed tool extracts features derived from processor’s PC and instruction type information and then utilizes the features to identify abnormal behavior using 4 different anomaly detection algorithms. New detection methods can be easily incorporated into the benchmark, which provides a solid foundation for evaluating novel, previously unseen methods against methods we selected for our experiment. Michal Borowski, Sangeet Saha, Xiaojun Zhai, Klaus D. McDonald-Maier |
TrustCom | 3 |
| 2022 | InterpolatedXY: a two-step strategy to normalize DNA methylation microarray data avoiding sex biasabstractMOTIVATION: Data normalization is an essential step to reduce technical variation within and between arrays. Due to the different karyotypes and the effects of X chromosome inactivation, females and males exhibit distinct methylation patterns on sex chromosomes; thus, it poses a significant challenge to normalize sex chromosome data without introducing bias. Currently, existing methods do not provide unbiased solutions to normalize sex chromosome data, usually, they just process autosomal and sex chromosomes indiscriminately. RESULTS: Here, we demonstrate that ignoring this sex difference will lead to introducing artificial sex bias, especially for thousands of autosomal CpGs. We present a novel two-step strategy (interpolatedXY) to address this issue, which is applicable to all quantile-based normalization methods. By this new strategy, the autosomal CpGs are first normalized independently by conventional methods, such as funnorm or dasen; then the corrected methylation values of sex chromosome-linked CpGs are estimated as the weighted average of their nearest neighbors on autosomes. The proposed two-step strategy can also be applied to other non-quantile-based normalization methods, as well as other array-based data types. Moreover, we propose a useful concept: the sex explained fraction of variance, to quantitatively measure the normalization effect. AVAILABILITY AND IMPLEMENTATION: The proposed methods are available by calling the function 'adjustedDasen' or 'adjustedFunnorm' in the latest wateRmelon package (https://github.com/schalkwyk/wateRmelon), with methods compatible with all the major workflows, including minfi. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yucheng Wang 0008, Tyler J. Gorrie-Stone, Olivia A. Grant, Alexandria D. Andrayas, Xiaojun Zhai, Klaus D. McDonald-Maier, Leonard C. Schalkwyk |
Bioinform. | 5 |
| 2022 | ACCURATE: Accuracy Maximization for Real-Time Multicore Systems With Energy-Efficient Way-Sharing CachesabstractImproving result accuracy in approximate computing (AC)-based real-time applications without violating deadlines has recently become an active research domain. Execution time of AC real-time tasks can individually be separated into: execution of the mandatory part to obtain a result of acceptable quality, followed by a partial/complete execution of the optional part to improve the result accuracy of the initial result within a given deadline. However, obtaining higher result accuracy at the cost of enhanced execution time may lead to deadline violation, along with higher energy usage. We present ACCURATE, a novel hybrid offline–online approximate real-time scheduling approach that first schedules AC-based tasks on multicore with an objective to maximize result accuracy and determines operational processing speeds for each task constrained by system-wide power limit, deadline, and task dependency. At runtime, by employing a way-sharing technique (WH_LLC) at the last level cache (LLC), ACCURATE improves performance, which is further leveraged, to enhance result accuracy by executing more from the optional part and to improve the energy efficiency of the cache by turning off a controlled number of cache ways. ACCURATE also exploits the slacks either to improve the result accuracy of the tasks or to enhance the energy efficiency of the underlying system, or both. ACCURATE achieves 85% QoS with 36% average reduction in cache leakage consumption with a 24% average gain in energy-delay product (EDP) for a 4-core-based chip multiprocessor (CMP) with 6.4% average improvement in performance. Sangeet Saha, Shounak Chakraborty 0001, Xiaojun Zhai, Shoaib Ehsan, Klaus D. McDonald-Maier |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | RASA: Reliability-Aware Scheduling Approach for FPGA-Based Resilient Embedded Systems in Extreme EnvironmentsabstractField-programmable gate arrays (FPGAs) offer the flexibility of general-purpose processors along with the performance efficiency of dedicated hardware that essentially renders it as a platform of choice for modern-day robotic systems for achieving real-time performance. Such robotic systems when deployed in harsh environments often get plagued by faults due to extreme conditions. Consequently, the real-time applications running on FPGA become susceptible to errors which call for a reliability-aware task scheduling approach, the focus of this article. We attempt to address this challenge using a hybrid offline-online approach. Given a set of periodic real-time tasks that require to be executed, the offline component generates a feasible preemptive schedule with specific preemption points. At runtime, these preemption events are utilized for fault detection. Upon detecting any faulty execution at such distinct points, the reliability-aware scheduling approach, RASA, orchestrates the recovery mechanism to remediate the scenario without jeopardizing the predefined schedule. Effectiveness of the proposed strategy has been verified through simulation-based experiments and we observed that the RASA is able to achieve 72% of task acceptance rate even under 70% of system workloads with high fault occurrence rates. Sangeet Saha, Xiaojun Zhai, Shoaib Ehsan, Shakaiba Majeed, Klaus D. McDonald-Maier |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2021 | EnSuRe: Energy & Accuracy Aware Fault-tolerant Scheduling on Real-time Heterogeneous SystemsabstractThis paper proposes an energy efficient real-time scheduling strategy called EnSuRe, which (i) executes real-time tasks on low power consuming primary processors to enhance the system accuracy by maintaining the deadline and (ii) provides reliability against a fixed number of transient faults by selectively executing backup tasks on high power consuming backup processor. Simulation results reveal that EnSuRe consumes nearly 25% less energy, compared to existing techniques, while satisfying the fault tolerance requirements. EnSuRe is also able to achieve 75% system accuracy with 50% system utilisation. Further, the obtained simulation outcomes are validated on benchmark tasks via a fault injection framework on Xilinx ZYNQ APSoC heterogeneous dual core platform. Sangeet Saha, Adewale Adetomi, Xiaojun Zhai, Server Kasap, Shoaib Ehsan, Tughrul Arslan, Klaus D. McDonald-Maier |
IOLTS | 3 |
| 2021 | Design and Implementation of a RISC V Processor on FPGAabstractThe RISC-V ISA is becoming one of the leading instruction sets for the Internet-of-Things and System-on-Chip applications. Due to its strong security features and open-source nature, it is becoming a competitor to the popular ARM architecture. This paper describes the design of a light weight, open-source implementation of a RISCV processor using modern hardware design teclmiques, the implementation of the design onto a Field Programmable Gate Array (FPGA), and its testing. We wanted to create a RISC-V processor that is easy for beginners to learn from and lightweight enough to be implemented on even small FPGAs. While there are existing opensource implementations of RISC-V processors, none are intuitive enough for a beginner to follow. For this reason, in this paper we have minimised the use of conventions and components in modern processors that are not strictly necessary for a barebones implementation. For example, the processor does not include pipelining and uses a simple Harvard architecture. The barebones nature of the design allows for a lot of potential for upgradability. The implementation of each component, and the corresponding test benches, are written in concise and conventional System Verilog. The project produced a RISC-V processor with files for targeting Basys 3 Artix-7 FPGA. Performance was tested using the Dhyrstone benchmark and achieved a strong 2276 DMIPs/MHz, even outperforming the ARM Cortex-A9, while maintaining very low resource utilization on the FPGA. Ludovico Poli, Sangeet Saha, Xiaojun Zhai, Klaus D. McDonald-Maier |
MSN | 3 |
| 2021 | Editorial for FGCS special issue: Intelligent IoT systems for healthcare and rehabilitation
Qingsong Ai, Wei Meng 0003, Faycal Bensaali, Xiaojun Zhai, Lu Liu 0001, Nasser Alaraje |
Future Gener. Comput. Syst. | 4 |
| 2021 | A Deep Segmentation Network of Multi-Scale Feature Fusion Based on Attention Mechanism for IVOCT Lumen ContourabstractRecently, coronary heart disease has attracted more and more attention, where segmentation and analysis for vascular lumen contour are helpful for treatment. And intravascular optical coherence tomography (IVOCT) images are used to display lumen shapes in clinic. Thus, an automatic segmentation method for IVOCT lumen contour is necessary to reduce the doctors' workload while ensuring diagnostic accuracy. In this paper, we proposed a deep residual segmentation network of multi-scale feature fusion based on attention mechanism (RSM-Network, Residual Squeezed Multi-Scale Network) to segment the lumen contour in IVOCT images. Firstly, three different data augmentation methods including mirror level turnover, rotation and vertical flip are considered to expand the training set. Then in the proposed RSM-Network, U-Net is contained as the main body, considering its characteristic of accepting input images with any sizes. Meanwhile, the combination of residual network and attention mechanism is applied to improve the ability of global feature extraction and solve the vanishing gradient problem. Moreover, the pyramid feature extraction structure is introduced to enhance the learning ability for multi-scale features. Finally, in order to increase the matching degree between the actual output and expected output, the cross entropy loss function is also used. A series of metrics are presented to evaluate the performance of our proposed network and the experimental results demonstrate that the proposed RSM-Network can learn the contour details better, contributing to strong robustness and accuracy for IVOCT lumen contour segmentation. Chenxi Huang 0001, Yisha Lan, Gaowei Xu, Xiaojun Zhai, Jipeng Wu, Fan Lin, Nianyin Zeng, Qingqi Hong, E. Y. K. Ng, Yonghong Peng |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2020 | A self-scrubbing scheme for embedded systems in radiation environmentsabstractAs one of the most important components in the embedded systems, the SRAM are sensitive to radiation effects. When the embedded systems working in the extreme radiation environments, the bit flips could occur frequently and decrease the reliability of the systems significantly. In this paper, the self-scrubbing RAM scheme is proposed for light wight embedded systems in the extreme radiation environments. In the scheme, both scrubbing and ECC are used to mitigate the large number of the errors in the RAMs. The separately scrubber is designed to scrub the RAM separately. Therefore it is can be able to operating the scrubbing, when the CPUs are busy. In addition, the scrubber is a portable modules and the hardware costs do not grow with the size of the available RAM. The results of the real world radiation experiments show that it can correct most errors in the RAM under neutron radiation where the errors rates in unhardened RAMs is approximately 1.2bit/(KB·h). The results of the 6 hours radiation experiments show that the error rates of in the conventional ECC RAM is approximately 4.3×10-4bit/(KB·h), while the self-scrubbing RAMs is less than 8.7×10-5bit/(KB·h). Yufan Lu, Xiaojun Zhai, Sangeet Saha, Shoaib Ehsan, Klaus D. McDonald-Maier |
IOLTS | 2 |
| 2020 | A Framework and Protocol for Dynamic Management of Fault Tolerant Systems in Harsh EnvironmentsabstractRobots can be used to deal with hazardous materials like nuclear waste. Unfortunately, electronic components are also susceptible to radiation effects. Current proposals to tackle this issue solve only parts of the problems for the specific scenarios and the specific types of radiation. At the same time, current computational devices should provide run-time capabilities to monitor and adapt to different situations. In this paper, we target a possible solution presenting a framework which provides the flexibility to employ fault-tolerant techniques on distributed systems. As proof of concept, we target a fault-tolerant technique to extend the operating time of systems in harsh environments. Results show a very low overhead, of few microseconds, to execute a majority voter with replicated tasks. Eduardo Wächter, Server Kasap, Xiaojun Zhai, Shoaib Ehsan, Klaus D. McDonald-Maier |
IOLTS | 3 |
| 2020 | Latency-Based Analytic Approach to Forecast Cloud Workload Trend for Sustainable DatacentersabstractCloud datacenters are turning out to be massive energy consumers and environment polluters, which necessitate the need for promoting sustainable computing approaches for achieving environment-friendly datacentre execution. Direct causes of excess energy consumption of the datacentre include running servers at low level of workloads and over-provisioning of server resources to the arriving workloads during execution. To this end, predicting the future workload demands and their respective behaviors at the datacenters are being the focus of recent researches in the context of sustainable datacenters. But prediction analytics of cloud workloads suffer various limitations imposed by the dynamic and unclear characteristics of Cloud workloads. This paper proposes a novel forecasting model named K-means based Rand Variable Learning Rate Backpropagation Neural Network (K-RVLBPNN) for predicting the future workload arrival trend, by exploiting the latency sensitivity characteristics of Cloud workloads, based on a combination of improved K-means clustering algorithm and Backpropagation Neural Network (BPNN) algorithm. Experiments conducted on real-world Cloud datasets shows that the proposed model shows better prediction accuracy, outperforming the traditional Hidden Markov Model, Naïve Bayes Classifier, and our earlier RVLBPNN model, respectively. Yao Lu 0021, Lu Liu 0001, John Panneerselvam, Xiaojun Zhai, Nick Antonopoulos |
IEEE Trans. Sustain. Comput. | 4 |
| 2019 | Hemelb Acceleration and Visualization for Cerebral AneurysmsabstractA weakness in the wall of a cerebral artery causing a dilation or ballooning of the blood vessel is known as a cerebral aneurysm. Optimal treatment requires fast and accurate diagnosis of the aneurysm. HemeLB is a fluid dynamics solver for complex geometries developed to provide neurosurgeons with information related to the flow of blood in and around aneurysms. On a cost efficient platform, HemeLB could be employed in hospitals to provide surgeons with the simulation results in real-time. In this work, we developed an improved version of HemeLB for GPU implementation and result visualization. A visualization platform for smooth interaction with end users is also presented. Finally, a comprehensive evaluation of this implementation is reported. The results demonstrate that the proposed implementation achieves a maximum performance of 15,168,964 site updates per second, and is capable of speeding up HemeLB for deployment in hospitals and clinical investigations. Sahar Soheilian Esfahani, Peter V. Coveney, Xiaojun Zhai, Minsi Chen, Abbes Amira, Faycal Bensaali, Julien Abinahed, Sarada Dakua, Georges Younes 0003, Robin A. Richardson |
ICIP | 3 |
| 2019 | Zynq SoC based acceleration of the lattice Boltzmann methodabstractSummary Cerebral aneurysm is a life‐threatening condition. It is a weakness in a blood vessel that may enlarge and bleed into the surrounding area. In order to understand the surrounding environmental conditions during the interventions or surgical procedures, a simulation of blood flow in cerebral arteries is needed. One of the effective simulation approaches is to use the lattice Boltzmann (LB) method. Due to the computational complexity of the algorithm, the simulation is usually performed on high performance computers. In this paper, efficient hardware architectures of the LB method on a Zynq system‐on‐chip (SoC) are designed and implemented. The proposed architectures have first been simulated in Vivado HLS environment and later implemented on a ZedBoard using the software‐defined SoC (SDSoC) development environment. In addition, a set of evaluations of different hardware architectures of the LB implementation is discussed in this paper. The experimental results show that the proposed implementation is able to accelerate the processing speed by a factor of 52 compared to a dual‐core ARM processor‐based software implementation. Xiaojun Zhai, Abbes Amira, Faycal Bensaali, AlMaha Al-Shibani, Asma Al-Nassr, Asmaa El-Sayed, Mohammad Eslami, Sarada Dakua, Julien Abinahed |
Concurr. Comput. Pract. Exp. | 1 |
| 2019 | Energy-efficient Static Task Scheduling on VFI-based NoC-HMPSoCs for Intelligent Edge Devices in Cyber-physical SystemsabstractThe interlinked processing units in modern Cyber-Physical Systems (CPS) creates a large network of connected computing embedded systems. Network-on-Chip (NoC)-based Multiprocessor System-on-Chip (MPSoC) architecture is becoming a de facto computing platform for real-time applications due to its higher performance and Quality-of-Service (QoS). The number of processors has increased significantly on the multiprocessor systems in CPS; therefore, Voltage Frequency Island (VFI) has been recently adopted for effective energy management mechanism in the large-scale multiprocessor chip designs. In this article, we investigated energy-efficient and contention-aware static scheduling for tasks with precedence and deadline constraints on intelligent edge devices deploying heterogeneous VFI-based NoC-MPSoCs (VFI-NoC-HMPSoC) with DVFS-enabled processors. Unlike the existing population-based optimization algorithms, we proposed a novel population-based algorithm called ARSH-FATI that can dynamically switch between explorative and exploitative search modes at run-time. Our static scheduler ARHS-FATI collectively performs task mapping, scheduling, and voltage scaling. Consequently, its performance is superior to the existing state-of-the-art approach proposed for homogeneous VFI-based NoC-MPSoCs. We also developed a communication contention-aware Earliest Edge Consistent Deadline First (EECDF) scheduling algorithm and gradient descent--inspired voltage scaling algorithm called Energy Gradient Decent (EGD). We introduced a notion of Energy Gradient (EG) that guides EGD in its search for island voltage settings and minimize the total energy consumption. We conducted the experiments on eight real benchmarks adopted from Embedded Systems Synthesis Benchmarks (E3S). Our static scheduling approach ARSH-FATI outperformed state-of-the-art technique and achieved an average energy-efficiency of ∼24% and ∼30% over CA-TMES-Search and CA-TMES-Quick, respectively. Umair Ullah Tariq, Haider Ali 0001, Lu Liu 0001, John Panneerselvam, Xiaojun Zhai |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2018 | Fast Algorithm for HEVC Intra Prediction Based on Adaptive Mode Decision and Early Termination of CU PartitionabstractHigh Efficiency Video Coding (HEVC) introduces 35 intra prediction modes and a flexible quad-tree coding structure, which remarkably improve the coding efficiency. However, in intra prediction, the cost computation in Rough Mode Decision (RMD) process and Rate Distortion Optimized (RDO) process suffers from a pretty high complexity compared with H.264. To deal with this problem, a modified RMD process is proposed, in which all 35 prediction modes are divided into groups according to their phase angle to reduce the candidate modes. Xiaojun Zhai, Zhi Liu 0008, Changzhi An |
DCC | 2 |
| 2018 | Service-Oriented System Engineering
Nik Bessis, Xiaojun Zhai, Stelios Sotiriadis |
Future Gener. Comput. Syst. | 2 |
| 2018 | Guest Editorial Special Issue on Real-Time Data Processing for Internet of ThingsabstractWith the development of the Internet of Things (IoT), various large-scale real-time data processing applications for handling real-time sensor data are becoming one of the important applications in cloud computing. The academia, the industry, and even the government institutions have already begun to pay close attention to how to efficiently process large amounts of sensor data in real-time using cloud computing technology. Although cloud computing technology has attracted much attention with high-performance, there are strong needs for improving data processing efficiency of large-scale real-time data for IoT-based applications. In addition to this, currently the IoT paradigm is facing increasing difficulty to handle the data generated from IoT applications. As a result of this, it is challenging to ensure low latency and network bandwidth consumption, optimal utilization of computational recourses, scalability, security, and energy efficiency of IoT devices while moving all data to the cloud. Therefore, this centralized computing model is starting to shift to a decentralized model termed as edge computing, that allows data to be handled from the cloud to local devices such as smartphones, smart gateways or routers, local PCs or sensor nodes on a smaller scale in real-time. Faycal Bensaali, Xiaojun Zhai, Abbes Amira, Lu Liu 0001 |
IEEE Internet Things J. | 2 |
| 2017 | ECG encryption and identification based security solution on the Zynq SoC for connected health systems
Xiaojun Zhai, Amine Ait Si Ali, Abbes Amira, Faycal Bensaali |
J. Parallel Distributed Comput. | 1 |
| 2016 | Heterogeneous Implementation of ECG Encryption and Identification on the Zynq SoCabstractThis paper presents an innovative and safe connected health solution for human identification. The system consists of the encryption and decryption of ECG signals using the advanced encryption standard (AES) as well as the recognition of individuals based on ECG biometrics. Heterogeneous and efficient implementation of the proposed system has been performed on a Xilinx ZC702 Zynq based prototyping board. Various IP-cores have been created based on the high level synthesis (HLS) implementation of the AES cipher, AES decipher and ECG identification blocks. The proposed hardware implementation has shown promising results since it met the real-time requirements and outclassed current field programmable gate array (FPGA) based systems in multiple key metrics including power consumption, processing time and hardware resources usage. The implemented system needs 10.71 ms to process one ECG sample and consumes 107mW while using only 30% of all available on-chip resources. Amine Ait Si Ali, Xiaojun Zhai, Abbes Amira, Faycal Bensaali, Naeem Ramzan |
FCCM | 2 |
| 2015 | Exploring ICMetrics to detect abnormal program behaviour on embedded devices
Xiaojun Zhai, Kofi Appiah, Shoaib Ehsan, Gareth Howells 0001, Huosheng Hu, Dongbing Gu, Klaus D. McDonald-Maier |
J. Syst. Archit. | 1 |
| 2015 | A Method for Detecting Abnormal Program Behavior on Embedded DevicesabstractA potential threat to embedded systems is the execution of unknown or malicious software capable of triggering harmful system behavior, aimed at theft of sensitive data or causing damage to the system. Commercial off-the-shelf embedded devices, such as embedded medical equipment, are more vulnerable as these type of products cannot be amended conventionally or have limited resources to implement protection mechanisms. In this paper, we present a self-organizing map (SOM)-based approach to enhance embedded system security by detecting abnormal program behavior. The proposed method extracts features derived from processor's program counter and cycles per instruction, and then utilises the features to identify abnormal behavior using the SOM. Results achieved in our experiment show that the proposed method can identify unknown program behaviors not included in the training set with over 98.4% accuracy. Xiaojun Zhai, Kofi Appiah, Shoaib Ehsan, Gareth Howells 0001, Huosheng Hu, Dongbing Gu, Klaus D. McDonald-Maier |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2013 | Automatic number plate recognition system on an ARM-DSP and FPGA heterogeneous SoC platforms
Zoe Jeffrey, Xiaojun Zhai, Faycal Bensaali, Reza Sotudeh, Aladdin M. Ariyaeeinia |
Hot Chips Symposium | 2 |
| 2010 | License plate localisation based on morphological operationsabstractAutomatic Number Plate Recognition (ANPR) systems allow users to track, identify and monitor moving vehicles by automatically extracting their number plates. This paper presents an improved method to locate car plates in an ANPR system. The proposed method is based on morphological open and close operations where different Structuring Elements (SE) are used to maximally eliminate non-plate region and enhance plate region. This method has been tested using a database of UK number plates and results achieved have shown significant improvements in terms of the detection rate compare to other existing plate localisation systems. Xiaojun Zhai, Faycal Bensaali, Soodamani Ramalingam |
ICARCV | 1 |