Zhiyi Huang 0001

dblp:13/6225-1 · DBLP profile ↗
← Back
89ranked-venue papers
12as first author
28since 2021 · last 2025
0000-0001-8561-2556ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 47 · 8 first-author · 14 since 2021Artificial intelligence and machine learning · 9 · 7 since 2021Computer networks · 8 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 ROCKET: An RNS-based Photonic Accelerator for High-Precision and Energy-Efficient DNN Training
abstract
In recent years, the rapid development of Deep Neural Networks (DNNs) has posed significant challenges in terms of training duration and costs. High-frequency, low-power photonic computing has emerged as a highly promising solution. However, the substantial cost of data conversion and the limitations introduced by noise in photonic devices continue to hinder the realization of high-precision and energy-efficient DNN training. To address this challenge, we propose a novel photonic accelerator, ROCKET, based on the Residue Number System (RNS). RNS is based on modular arithmetic and enables support for high-precision computation through parallel multi-path low-precision operations. First, we leverage specialized lookup tables to enable high-throughput, low-latency conversions between high-precision and low-precision numerical representations. Next, we design a low-power photonic accelerator architecture utilizing intensity modulators, which minimizes the number of computational components while maximizing data reuse. Subsequently, we propose a hybrid photonic-electronic pipelined dataflow to maximize parallelism within the photonic-electronic computation path. Finally, we develop a high-frequency (4.096 GHz) hybrid photonic-electronic prototype using FPGA, Radio Frequency (RF), and photonic components to validate the feasibility of the ROCKET. Our large-scale simulations on seven mainstream DNN models show that, compared to the A100 GPU, TPU v4, and the state-of-the-art photonic accelerator Mirage, ROCKET achieves speedups of 33×, 243×, and 198×, respectively, while saving energy by factors of 64×, 204×, and 142×.
Hao Zhang 0058, Haibo Zhang 0001, Chengpeng Xia, Zhiyi Huang 0001, Yawen Chen 0001, Amanda S. Barnard
ICS4
2025 Gated parallel feature fusion multi-task learning for motor imagery EEG classification
Xianheng Wang, Veronica Liesaputra, Zhiyi Huang 0001
Expert Syst. Appl.3
2025 ChipAI: A scalable chiplet-based accelerator for efficient DNN inference using silicon photonics
abstract
To enhance the precision of inference, deep neural network (DNN) models have been progressively growing in scale and complexity, leading to increased latency and computational resource demands. This growth necessitates scalable architectures, such as chiplet-based accelerators, to accommodate the substantial volume of deep learning inference tasks. However, the efficiency, energy consumption, and scalability of existing accelerators are severely constrained by metallic interconnects. Photonic interconnects, on the contrary, offer a promising alternative, with their advantages of low latency, high bandwidth, high energy efficiency, and simplified communication processes. In this paper, we propose ChipAI, an accelerator designed based on photonic interconnects for accelerating DNN inference tasks. ChipAI implements an efficient hybrid optical network that supports effective inter-chiplet and intra-chiplet data sharing, thereby enhancing parallel processing capabilities. Additionally, we propose a flexible dataflow leveraging the ChipAI architecture and the characteristics of DNN models, facilitating efficient architectural mapping of DNN layers. Simulation on various DNN models demonstrates that, compared to the state-of-the-art chiplet-based DNN accelerator with photonic interconnects, ChipAI can reduce the DNN inference time and energy consumption by up to 82% and 79%, respectively.
Hao Zhang 0058, Haibo Zhang 0001, Zhiyi Huang 0001, Yawen Chen 0001
J. Syst. Archit.3
2025 A Non-Invasive Blood Glucose Detection System Based on Photoplethysmogram With Multiple Near-Infrared Sensors
abstract
Recent advancements in non-invasive blood glucose detection have seen progress in both photoplethysmogram and multiple near-infrared methods. While the former shows better predictability of baseline glucose levels, it lacks sensitivity to daily fluctuations. Near-infrared methods respond well to short-term changes but face challenges due to individual and environmental factors. To address this, we developed a novel fingertip blood glucose detection system combining both methods. Using multiple light sensors and a lightweight deep learning model, our system achieved promising results in oral glucose tolerance tests. A total of 10 participants were involved in the study, each providing approximately 700 data segments of about 10 seconds each. With a root mean squared error of 0.242 mmol/L and 100% accuracy in zone A of the Parkes error grid, our approach demonstrates the potential of multiple near-infrared sensors for non-invasive glucose detection.
Zhiyi Huang 0001, Houbing Song, Yuan-Ting Zhang, Yuan Zhang 0007, Zhen Mei 0002
IEEE J. Biomed. Health Informatics3
2024 CoDe : A Cooperative and Decentralized Collision Avoidance Algorithm for Small-Scale UAV Swarms Considering Energy Efficiency
abstract
This paper introduces a cooperative and decentralized collision avoidance algorithm (CoDe) for small-scale UAV swarms consisting of up to three UAVs. CoDe improves energy efficiency of UAVs by achieving effective cooperation among UAVs. Moreover, CoDe is specifically tailored for UAV’s operations by addressing the challenges faced by existing schemes, such as ineffectiveness in selecting actions from continuous action spaces and high computational complexity. CoDe is based on Multi-Agent Reinforcement Learning (MARL), and finds cooperative policies by incorporating a novel credit assignment scheme. The novel credit assignment scheme estimates the contribution of an individual by subtracting a baseline from the joint action value for the swarm. The credit assignment scheme in CoDe outperforms other benchmarks as the baseline takes into account not only the importance of a UAV’s action but also the interrelation between UAVs. Furthermore, extensive experiments are conducted against existing MARL-based and conventional heuristic-based algorithms to demonstrate the advantages of the proposed algorithm.
Shuangyao Huang, Haibo Zhang 0001, Zhiyi Huang 0001
IROS3
2024 SQPMF: successive point of interest recommendation system based on probability matrix factorization
Zhiyi Huang 0001
Appl. Intell.2
2024 An in-depth survey on Deep Learning-based Motor Imagery Electroencephalogram (EEG) classification
abstract
Electroencephalogram (EEG)-based Brain-Computer Interfaces (BCIs) build a communication path between human brain and external devices. Among EEG-based BCI paradigms, the most commonly used one is motor imagery (MI). As a hot research topic, MI EEG-based BCI has largely contributed to medical fields and smart home industry. However, because of the low signal-to-noise ratio (SNR) and the non-stationary characteristic of EEG data, it is difficult to correctly classify different types of MI-EEG signals. Recently, the advances in Deep Learning (DL) significantly facilitate the development of MI EEG-based BCIs. In this paper, we provide a systematic survey of DL-based MI-EEG classification methods. Specifically, we first comprehensively discuss several important aspects of DL-based MI-EEG classification, covering input formulations, network architectures, public datasets, etc. Then, we summarize problems in model performance comparison and give guidelines to future studies for fair performance comparison. Next, we fairly evaluate the representative DL-based models using source code released by the authors and meticulously analyse the evaluation results. By performing ablation study on the network architecture, we found that (1) effective feature fusion is indispensable for multi-stream CNN-based models. (2) LSTM should be combined with spatial feature extraction techniques to obtain good classification performance. (3) the use of dropout contributes little to improving the model performance, and that (4) adding fully connected layers to the models significantly increases their parameters but it might not improve their performance. Finally, we raise several open issues in MI-EEG classification and provide possible future research directions.
Xianheng Wang, Veronica Liesaputra, Yi Wang 0126, Zhiyi Huang 0001
Artif. Intell. Medicine5
2024 OBhunter: An ensemble spectral-angular based transformer network for occlusion detection
Jiangnan Zhang, Kewen Xia, Zhiyi Huang 0001, Romoke Grace Akindele
Expert Syst. Appl.3
2024 Power-Adaptive Communication With Channel-Aware Transmission Scheduling in WBANs
abstract
Radio links in Wireless Body Area Networks (WBANs) are highly subject to short and long-term attenuation due to the unstable network topology and frequent body blockage. This instability makes it challenging to achieve reliable and energy-efficient communication, but on the other hand, provides a great potential for the sending nodes to dynamically schedule the transmissions at the time with the best-expected channel quality. Motivated by this, we propose IGE (Improved Gilbert-Elliott Markov chain model), a memory-efficient Markov chain model to monitor channel fluctuations and provide a long-term channel prediction. We then design ATPS (Adaptive Transmission Power Selection), a deadline-constrained channel scheduling scheme that enables a sending node to buffer the packets when the channel is bad and schedule them to be transmitted when the channel is expected to be good within a deadline. ATPS can self-learn the pattern of channel changes without imposing a significant computation or memory overhead on the sending node. We evaluate the performance of ATPS through experiments using TelosB motes under different scenarios with different body postures and packet rates. We further compare ATPS with several state-of-the-art schemes including the optimal scheduling policy in which the optimal transmission time for each packet is calculated based on the collected RSSI (Received Signal Strength Indicator) samples in an off-line manner. The experimental results reveal that ATPS performs almost as efficiently as the optimal scheme in high-date-rate scenarios and has a similar trend on power level usage.
Abbas Arghavani, Haibo Zhang 0001, Zhiyi Huang 0001, Yawen Chen 0001
IEEE Internet Things J.3
2024 Feature specific progressive improvement for salient object detection
Xianheng Wang, Veronica Liesaputra, Zhiyi Huang 0001
Pattern Recognit.4
2024 E2CoPre: Energy Efficient and Cooperative Collision Avoidance for UAV Swarms With Trajectory Prediction
abstract
This paper presents a novel solution to address the challenges in achieving energy efficiency and cooperation for collision avoidance in UAV swarms. The proposed method combines Artificial Potential Field (APF) and Particle Swarm Optimization (PSO) techniques. APF provides environmental awareness and implicit coordination to UAVs, while PSO searches for collision-free and energy-efficient trajectories for each UAV in a decentralized manner under the implicit coordination. This decentralized approach is achieved by minimizing a novel cost function that leverages the advantages of the active contour model from image processing. Additionally, future trajectories are predicted by approximating the minima of the novel cost function using calculus of variation, which enables proactive actions and defines the initial conditions for PSO. We propose a two-branch trajectory planning framework that ensures UAVs only change altitudes when necessary for energy considerations. Extensive experiments are conducted to evaluate the effectiveness and efficiency of our method in various situations.
Shuangyao Huang, Haibo Zhang 0001, Zhiyi Huang 0001
IEEE Trans. Intell. Transp. Syst.3
2023 Performance Comparison of Distributed DNN Training on Optical Versus Electrical Interconnect Systems
Yawen Chen 0001, Zhiyi Huang 0001, Haibo Zhang 0001, Hui Tian 0001
ICA3PP (1)3
2023 OpTree: An Efficient Algorithm for All-gather Operation in Optical Interconnect Systems
abstract
All-gather collective communication is one of the most important communication primitives in parallel and distributed computation, which plays an essential role in many high performance computing (HPC) applications such as distributed Deep Learning (DL) with model and hybrid parallelisms. To solve the communication bottleneck of All-gather, optical interconnection network can provide unprecedented high bandwidth and reliability for data transfer among the distributed nodes. However, most traditional All-gather algorithms are designed for electrical interconnection, which cannot fit well for optical interconnect systems, resulting in poor performance. This paper proposes an efficient scheme, called OpTree, for All-gather operation on optical interconnect systems. OpTree derives an optimal m-ary tree corresponding to the optimal number of communication stages, which achieves the minimum communication time. We further analyze and compare the communication steps of OpTree with existing All-gather algorithms. Theoretical results exhibit that OpTree requires much less number of communication steps than existing All-gather algorithms on optical interconnect systems. Simulation results show that OpTree can reduce communication time by 72.21 %, 94.30%, and 88.58% compared to three existing All-gather schemes Wrht, Ring, and NE, respectively.
Yawen Chen 0001, Zhiyi Huang 0001, Haibo Zhang 0001
ICC3
2023 Wrht: Efficient All-reduce for Distributed DNN Training in Optical Interconnect Systems
abstract
Communication efficiency is crucial for accelerating distributed deep neural network (DNN) training. All-reduce, a vital communication primitive, is responsible for reducing model parameters in distributed DNN training. However, most existing All-reduce algorithms, designed for traditional electrical interconnect systems, fall short due to bandwidth limitations. Optical interconnects, with superior bandwidth, low transmission delay, and less power consumption, emerge as viable alternatives. We propose Wrht (Wavelength Reused Hierarchical Tree), an efficient scheme for implementing the All-reduce operation in optical interconnect systems. Wrht leverages wavelength-division multiplexing (WDM) to minimize the communication time in distributed data-parallel DNN training. We calculate the required wavelengths, minimum communication steps, and optimal communication time, considering optical communication constraints. Simulations with real-world DNN models indicate that Wrht notably reduces communication time. On average, compared with three conventional All-reduce algorithms, Wrht achieves reductions of 65.23%, 43.81%, and 82.22% respectively in optical interconnect systems, and 61.23% and 55.51% compared with two algorithms in electrical systems. This highlights Wrht’s potential to enhance communication efficiency in DNN training using optical interconnects.
Yawen Chen 0001, Zhiyi Huang 0001, Haibo Zhang 0001
ICPP3
2023 SEECHIP: A Scalable and Energy-Efficient Chiplet-based GPU Architecture Using Photonic Links
abstract
The continuous increase in GPU performance benefits a wide range of high-performance computing (HPC) applications. Slower growth of transistor density and limited size of chip die are now posing significant challenges to scale GPUs. The chiplet technology provides a potential solution to surpass these limitations. However, the performance of these chiplet-based GPUs is often constrained by the metallic-based interconnects between the chiplets. Emerging technologies such as photonic interconnect can overcome the limitations of metallic interconnects, offering several superior properties, such as high bandwidth density and low energy consumption. In this paper, we propose SEECHIP: a Scalable and Energy-Efficient CHIPlet-based GPU architecture using photonic links. SEECHIP introduces a novel photonic inter-chiplet network that supports both unicast and broadcast communication, providing the same transmission bandwidth at both the sending and receiving ends. In addition, we propose a tailored hierarchical memory architecture, which is more suitable for the parallelization of general-purpose HPC applications. Simulation results using 14 benchmarks show that SEECHIP can achieve and reduction in execution time and energy consumption, respectively, as compared to other GPUs with metallic or photonic interconnects. Simulation results also show that SEECHIP has good scalability compared with the other GPUs.
Hao Zhang 0058, Yawen Chen 0001, Zhiyi Huang 0001, Haibo Zhang 0001
ICPP3
2023 Efficient All-Reduce for Distributed DNN Training in Optical Interconnect Systems
abstract
All-reduce is the crucial communication primitive to reduce model parameters in distributed Deep Neural Networks (DNN) training. Most existing all-reduce algorithms are designed for traditional electrical interconnect systems, which cannot meet the communication requirements for distributed training of large DNNs due to the low data bandwidth of the electrical interconnect systems. One of the promising alternatives for electrical interconnect is optical interconnect, which can provide high bandwidth, low transmission delay, and low power cost. We propose an efficient scheme called WRHT (Wavelength Reused Hierarchical Tree) for implementing all-reduce operation in optical interconnect systems. WRHT can take advantage of WDM (Wavelength Division Multiplexing) to reduce the communication time of distributed data-parallel DNN training. Simulations using real DNN models show that, compared to all-reduce algorithms in the electrical and optical network systems, our approach reduces communication time by 75.76% and 91.86%, respectively.
Yawen Chen 0001, Zhiyi Huang 0001, Haibo Zhang 0001, Fangfang Zhang 0002
PPoPP3
2023 Decentralized piggybacking-based dissemination of Cooperative Awareness Messages in vehicular ad-hoc networks
Guangbing Xiao, Haibo Zhang 0001, Zhiyi Huang 0001, Yawen Chen 0001
Comput. Networks3
2023 ETAM: Ensemble transformer with attention modules for detection of small objects
Jiangnan Zhang, Kewen Xia, Zhiyi Huang 0001, Romoke Grace Akindele
Expert Syst. Appl.3
2023 ESA: An efficient sequence alignment algorithm for biological database search on Sunway TaihuLight
Hao Zhang 0058, Zhiyi Huang 0001, Yawen Chen 0001, Jianguo Liang, Xiran Gao
Parallel Comput.2
2023 Routing and Wavelength Assignment for Multiple Multicasts in Optical Network-on-Chip (ONoC)
abstract
Optical network-on-chip (ONoC) is an emerging chip-scale optical interconnection technology to realize high-performance and power-efficient intercore communication for many-core processors. Multicast communication is popularly used in parallel applications on chip. However, existing researches for multicast in ONoC mainly focus on the optimization of one multicast. This limits the practical applications of the research outcomes because we often face the dynamic formation of multiple multicast groups in real network systems. In this article, we define the problem of routing and wavelength assignment for multiple multicasts in ONoC with the objective of minimizing the number of wavelengths required. To solve the problem, we first formulate it as an integer programming model for general topologies. Then we design routing policies for special instances that optimally use only one wavelength on mesh topology. For general instances, we design a group-partitioning routing algorithm for multiple multicasts (GPRMM). GPRMM decouples a group of multicasts into a number of subgroups, each of which matching one of the special instances. Theoretical results show that the number of wavelengths required by GPRMM is no more than the Destination Density$\sigma _{d}$, i.e., the maximum number of multicasts with destinations in the same row or column. Moreover, we find the upper bound and the lower bound on the number of wavelengths required for GPRMM. The wavelength requirement is also upper bounded by the network size$n$for an$n\times n$mesh network. Simulation results show that GPRMM can reduce the number of wavelengths by 26.7% compared with previous methods. GPRMM has the advantages of low routing complexity, low wavelength requirement, low power consumption, and good scalability.
Yawen Chen 0001, Zhiyi Huang 0001, Haibo Zhang 0001, Huaxi Gu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 Comparing the performance of multi-layer perceptron training on electrical and optical network-on-chips
Yawen Chen 0001, Zhiyi Huang 0001, Haibo Zhang 0001, Hao Zhang 0058, Chengpeng Xia
J. Supercomput.3
2023 Tuatara: Location-Driven Power-Adaptive Communication for Wireless Body Area Networks
abstract
Radio links in wireless body area networks (WBANs) suffer from both short-term and long-term variations due to the dynamic network topology and frequent blockage caused by body movements, making it challenging to achieve reliable, energy-efficient and real-time data communication. Through experiments with TelosB motes, we observe a strong positive relationship between the channel quality and the location of the sensor node relative to the gateway. Motivated by this observation, we design Tuatara, a novel power-aware communication protocol that allows each sensor node to dynamically adjust its transmission power based on the channel status inferred from its instant location, aiming to save energy, reduce interference, and improve communication reliability. Combining the orientations measured by motion sensors with the anatomical constraints of body movements, each sensor node can locally estimate its instant location relative to the gateway. Based on a probabilistic model, power level selection is converted to calculate the optimal probability of selecting each power level at a given location, with the objective of minimizing the transmission cost. A learning scheme is designed to adaptively update the power level selection probabilities, making Tuatara self-adaptable to changes in the signal propagation environment. Experimental results demonstrate that Tuatara outperforms the state-of-the-art protocols in various scenarios, with performance close to that of the optimal power selection solution even in scenarios where the packet rate is very low.
Abbas Arghavani, Haibo Zhang 0001, Zhiyi Huang 0001, Yawen Chen 0001
IEEE Trans. Mob. Comput.3
2022 Astraea: towards QoS-aware and resource-efficient multi-stage GPU services
abstract
Multi-stage user-facing applications on GPUs are widely-used nowa- days, and are often implemented to be microservices. Prior re- search works are not applicable to ensuring QoS of GPU-based microservices due to the different communication patterns and shared resource contentions. We propose Astraea to manage GPU microservices considering the above factors. In Astraea, a microser- vice deployment policy is used to maximize the supported peak service load while ensuring the required QoS. To adaptively switch the communication methods between microservices according to different deployments, we propose an auto-scaling GPU communi- cation framework. The framework automatically scales based on the currently used hardware topology and microservice location, and adopts global memory-based techniques to reduce intra-GPU communication. Astraea increases the supported peak load by up to 82.3% while achieving the desired 99%-ile latency target compared with state-of-the-art solutions.
Wei Zhang 0149, Quan Chen 0002, Kaihua Fu, Ningxin Zheng, Zhiyi Huang 0001, Jingwen Leng, Minyi Guo
ASPLOS5
2022 Special issue on programming models and applications for multicores and manycores 2019-2020
Min Si, Quan Chen 0002, Zhiyi Huang 0001
Concurr. Comput. Pract. Exp.3
2022 Special Issue on Programming Models and Applications for Multicores and Manycores 2020
abstract
Special Issue on Programming Models
Min Si, Quan Chen 0002, Zhiyi Huang 0001
Concurr. Comput. Pract. Exp.3
2022 A fast lasso-based method for inferring higher-order interactions
abstract
Large-scale genotype-phenotype screens provide a wealth of data for identifying molecular alterations associated with a phenotype. Epistatic effects play an important role in such association studies. For example, siRNA perturbation screens can be used to identify combinatorial gene-silencing effects. In bacteria, epistasis has practical consequences in determining antimicrobial resistance as the genetic background of a strain plays an important role in determining resistance. Recently developed tools scale to human exome-wide screens for pairwise interactions, but none to date have included the possibility of three-way interactions. Expanding upon recent state-of-the-art methods, we make a number of improvements to the performance on large-scale data, making consideration of three-way interactions possible. We demonstrate our proposed method, Pint, on both simulated and real data sets, including antibiotic resistance testing and siRNA perturbation screens. Pint outperforms known methods in simulated data, and identifies a number of biologically plausible gene effects in both the antibiotic and siRNA models. For example, we have identified a combination of known tumour suppressor genes that is predicted (using Pint) to cause a significant increase in cell proliferation.
Kieran Elmes, Astra Heywood, Zhiyi Huang 0001, Alex Gavryushkin
PLoS Comput. Biol.3
2021 Performance Comparison of Multi-layer Perceptron Training on Electrical and Optical Network-on-Chips
Yawen Chen 0001, Zhiyi Huang 0001, Haibo Zhang 0001
PDCAT3
2021 I-Scheduler: Iterative scheduling for distributed stream processing systems
Leila Eskandari, Jason Mair, Zhiyi Huang 0001, David M. Eyers
Future Gener. Comput. Syst.3
2019 AnxietyDecoder: An EEG-based Anxiety Predictor using a 3-D Convolutional Neural Network
abstract
In this paper, we propose and implement an EEG-based three-dimensional Convolutional Neural Network architecture, `AnxietyDecoder', to predict anxious personality and decode its potential biomarkers from the participants. Since Goal-Conflict-Specific-Rhythmicity (GCSR) in the EEG is a sign of an anxiety-related system working, we first propose a two-dimensional Conflict-focused CNN (2-D CNN). It simulates the GCSR extraction process but with the advantages of automatic frequency band selection and functional contrast calculation optimization, thus providing more comprehensive trait anxiety predictions. Then, to generate more targeted hierarchical features from local spatio-temporal scale to global, we propose a three-dimensional Conflict-focused CNN (3-D CNN), which simultaneously integrates information in the temporal and brain-topology-related spatial dimensions. In addition, we embed Layer-wise Relevance Propagation (LRP) into our model to reveal the essential brain areas that are correlated to anxious personality. The experimental results show that the percentage variance accounted for by our three-dimensional Conflict-focused CNN is 33%, which is almost four times higher than the previous theoretically derived GCSR contrast (7%). Meanwhile, it also outperforms the 2-D model (26%) and the t-test difference between the 3-D and 2-D models is significant (t(4) = 5.4962, p = 0.0053). What's more, the reverse engineering results provide an interpretable way to understand the prediction decision-making and participants' anxiety personality. Our proposed AnxietyDecoder not only sets a new benchmark for EEG-based anxiety prediction but also reveals essential EEG components that contribute to the decision-making, and thus sheds some light on the anxiety biomarker research.
Yi Wang 0126, Brendan McCane, Neil McNaughton, Zhiyi Huang 0001, Shabah M. Shadli, Phoebe Neo
IJCNN4
2019 TrustZone for Supervised Asymmetric Multiprocessing Systems
abstract
Many modern forms of asymmetric multiprocessing (AMP) architecture use hypervisors to increase software security by isolating the system software in virtual machines. However, efficient virtualisation depends on hardware support that is not available across all products. Within modern ARM architectures, the aforementioned software isolation can also be implemented using ARM TrustZone technology. This paper presents a TrustZone-based AMP architecture (TZ-AMP) that can consolidate multiple system software environments securely on devices that lack hardware virtualisation support. We evaluate our prototype on the ARMv7-A architecture, and demonstrate TrustZone-based context-switching performance in the order of microseconds, confirming that TZ-AMP maintains high performance while also achieving hardware-backed software security.
Mahdi Amiri-Kordestani, David M. Eyers, Zhiyi Huang 0001, Morteza Biglari-Abhari
PDCAT3
2019 Manila: Using a densely populated PMC-space for power modelling within large-scale systems
Jason Mair, Zhiyi Huang 0001, David M. Eyers
Parallel Comput.2
2019 International workshop on programming models and applications for multicores and manycores (PMAM 2018)
Min Si, Zhiyi Huang 0001, Pavan Balaji
Parallel Comput.2
2019 Chimp: A Learning-based Power-aware Communication Protocol for Wireless Body Area Networks
abstract
Radio links in wireless body area networks (WBANs) commonly experience highly time-varying attenuation due to the dynamic network topology and frequent occlusions caused by body movements, making it challenging to design a reliable, energy-efficient, and real-time communication protocol for WBANs. In this article, we present Chimp, a learning-based power-aware communication protocol in which each sending node can self-learn the channel quality and choose the best transmission power level to reduce energy consumption and interference range while still guaranteeing high communication reliability. Chimp is designed based on learning automata that uses only the acknowledgment packets and motion data from a local gyroscope sensor to infer the real-time channel status. We design a new cost function that takes into account the energy consumption, communication reliability and interference and develop a new learning function that can guarantee to select the optimal transmission power level to minimize the cost function for any given channel quality. For highly dynamic postures such as walking and running, we exploit the correlation between channel quality and motion data generated by a gyroscope sensor to fastly estimate channel quality, eliminating the need to use expensive channel sampling procedures. We evaluate the performance of Chimp through experiments using TelosB motes equipped with the MPU-9250 motion sensor chip and compare it with the state-of-the-art protocols in different body postures. Experimental results demonstrate that Chimp outperforms existing schemes and works efficiently in most common body postures. In high-date-rate scenarios, it achieves almost the same performance as the optimal power assignment scheme in which the optimal power level for each transmission is calculated based on the collected channel measurements in an off-line manner.
Abbas Arghavani, Haibo Zhang 0001, Zhiyi Huang 0001, Yawen Chen 0001
ACM Trans. Embed. Comput. Syst.3
2019 Wavelength-Reused Hierarchical Optical Network on Chip Architecture for Manycore Processors
abstract
Manycore processor is becoming the mainstream platform for cloud computing applications. However, the design of high-performance and sustainable inter-core communication network is still a challenging problem. Optical Network on Chip (ONoC) is an emerging chip-scale optical communication technology with high bandwidth capacity and energy efficiency. In this paper, we present a Wavelength Reused Hierarchical ONoC architecture, WRH-ONoC. It leverages the nonblocking wavelength-routed λ-router and hierarchical networking to reuse the limited number of wavelengths. In WRH-ONoC, all the cores are grouped into multiple subsystems, and the cores in the same subsystem are directly interconnected using a λ-router for nonblocking communication. For inter-subsystem communication, all subsystems are further connected through multiple λ-routers and gateways in a hierarchical manner. Thus, the available wavelengths can be reused in different λ-routers. Furthermore, WRHm-ONoC, an efficient extension with multicast ability is also proposed. Given the numbers of cores and available wavelengths, we derive the minimum hardware requirement, the expected end-to-end delay, and the maximum data rate. Theoretical analysis and simulation results indicate WRH-ONoC achieves prominent improvement on the communication performance and sustainability, e.g., 46.0 percent of reduction on zero-load delay and 72.7 percent of improvement on throughput for 400 cores with the modest hardware/energy costs.
Haibo Zhang 0001, Yawen Chen 0001, Zhiyi Huang 0001, Huaxi Gu
IEEE Trans. Sustain. Comput.4
2018 EmotioNet: A 3-D Convolutional Neural Network for EEG-based Emotion Recognition
abstract
In this paper, an emotional EEG-specific three-dimensional Convolutional Neural Network, EmotioNet, is proposed and implemented to accurately recognize emotion states. For the first time, raw data in the benchmark emotional EEG database, i.e., DEAP, are used as the input to a CNN architecture. In order to investigate the spatio-temporal character of emotional features, the effectiveness of 2-D and 3-D convolution kernels, which extract spatial and temporal features separately and simultaneously, are compared in detail. Furthermore, two major problems of EEG-based emotion recognition, namely, covariance shift and the unreliability of emotional ground truth, are described, and the effectiveness of batch normalization and dense prediction, which alleviate these problems respectively, are also investigated. Experimental results show that 3-D kernels, batch normalization, and dense prediction are all essential techniques for the emotional EEG-specific CNN architecture. The proposed EmotioNet, namely, a 3-D covariance shift adaptation-based CNN with a dense prediction layer, achieves classification rates of 73.3% and 72.1% for arousal and valence, equivalent to the best performance of several previous studies. Importantly, our results are based on automatic feature extraction, which is in contrast to previous handcrafted features. Therefore, EmotioNet provides a new method for EEG-based emotion recognition.
Yi Wang 0126, Zhiyi Huang 0001, Brendan McCane, Phoebe Neo
IJCNN2
2018 T3-Scheduler: A topology and Traffic aware two-level Scheduler for stream processing systems in a heterogeneous cluster
Leila Eskandari, Jason Mair, Zhiyi Huang 0001, David M. Eyers
Future Gener. Comput. Syst.3
2018 Distributed sparse bundle adjustment algorithm based on three-dimensional point partition and asynchronous communication
abstract
Sparse bundle adjustment (SBA) is a key but time- and memory-consuming step in three-dimensional (3D) reconstruction. In this paper, we propose a 3D point-based distributed SBA algorithm (DSBA) to improve the speed and scalability of SBA. The algorithm uses an asynchronously distributed sparse bundle adjustment (A-DSBA) to overlap data communication with equation computation. Compared with the synchronous DSBA mechanism (SDSBA), A-DSBA reduces the running time by 46%. The experimental results on several 3D reconstruction datasets reveal that our distributed algorithm running on eight nodes is up to five times faster than that of the stand-alone parallel SBA. Furthermore, the speedup of the proposed algorithm (running on eight nodes with 48 cores) is up to 41 times that of the serial SBA (running on a single node).
Xiaolong Shen, Yong Dou, Steven Mills, David M. Eyers, Huan Feng, Zhiyi Huang 0001
Frontiers Inf. Technol. Electron. Eng.6
2018 Principal Component Analysis Based Filtering for Scalable, High Precision k-NN Search
abstract
Approximate$k$Nearest Neighbours (A$k$NN) search is widely used in domains such as computer vision and machine learning. However, A$k$NN search in high-dimensional datasets does not scale well on multicore platforms, due to its large memory footprint. Parallel A$k$NN search using space subdivision for filtering helps reduce the memory footprint, but its loss of precision is unstable. In this paper, we propose a new data filtering method—PCAF—for parallel A$k$NN search based on principal component analysis. PCAF improves on previous methods, demonstrating sustained, high scalability for a wide range of high-dimensional datasets on both Intel and AMD multicore platforms. Moreover, PCAF maintains highly precise A$k$NN search results.
Huan Feng, David M. Eyers, Steven Mills, Yongwei Wu 0001, Zhiyi Huang 0001
IEEE Trans. Computers5
2017 ATPS: Adaptive Transmission Power Selection for Communication in Wireless Body Area Networks
abstract
Since radio links in wireless body area networks (WBANs) commonly experience highly time-varying attenuation due to topology instability, communication protocols with fixed transmission power cannot produce a very good performance in terms of energy consumption, interference range, and communication reliability. We explain that how channel behaviourcan be modelled using Markov Chain. Then, a power-adaptive communication protocol for WBANs is developed in which each sensor node can self-learn its channel and dynamically adjust itstransmission power. We evaluate our scheme through implementing the idea using the TelosB motes. The results demonstrate that our scheme can self-learn the channel behaviours, and reduce energy consumption and interference.
Abbas Arghavani, Haibo Zhang 0001, Zhiyi Huang 0001, Yawen Chen 0001
LCN3
2017 CRAFT reducing the effort for indoor localisation
abstract
Indoor localisation systems have slowly become more and more accurate. Each localisation system needs tuning to affect reasonable performance. In this paper we propose CRAFT, a crowd sourced approach to constructing a WiFi fingerprint database. The method uses a temporarily deployment of a small number of anchor nodes to roughly locate the position of the WiFi sample. Through thorough experiments in a real-world building, CRAFT's error is 2.2 m a decrease of 25% when compare to other published results.
Paul Crane, Zhiyi Huang 0001, Haibo Zhang 0001
PIMRC2
2017 A cooperative offloading game on data recovery for reliable broadcast in VANET
abstract
Summary The rapidly growing demand for accident‐free driving in intelligent transportation makes reliable broadcast a critical factor for vehicularad hocnetworks. Existing solutions always try to improve the broadcast reliability by retransmitting lost packets. However, the excessive retransmissions can easily cause unpredictable time delay and even broadcast storms, rendering the reliable broadcast problem unsolved. In this paper, a novel reliable broadcast scheme is proposed by exploring the advantages of lost data piggybacking. Our scheme allows all the vehicles to piggyback some received packets cooperatively to help other vehicles to recover the lost packets. We formulate the cooperative piggybacking problem as a cooperative offloading game and present a decentralized solution to compute the optimal data piggybacking solutions based on only partial network information. A reward‐penalty scheme is designed for the offloading process to impel all the vehicles' decisions that converge to the Nash equilibrium, which is proved to be the global optimal solution to the decentralized offloading scheme. Simulation results show that the proposed cooperative offloading scheme can achieve much higher broadcast reliability and lower propagation delay, in comparison with existing solutions. In a small vehicle network, all lost cooperative awareness messages can be successfully recovered within 25 ms after the initial broadcast by using the data traces generated by GEMV2. Copyright © 2016 John Wiley & Sons, Ltd.
Guangbing Xiao, Haibo Zhang 0001, Houcine Hassan, Yawen Chen 0001, Zhiyi Huang 0001, Ning Sun 0007
Concurr. Comput. Pract. Exp.5
2016 PCAF: Scalable, High Precision k-NN Search Using Principal Component Analysis Based Filtering
abstract
Approximate k Nearest Neighbours (AkNN) search is widely used in domains such as computer vision and machine learning. However, AkNN search in high dimensional datasets does not work well on multicore platforms. It scales poorly due to its large memory footprint. Current parallel AkNN search using space subdivision for filtering helps reduce the memory footprint, but leads to loss of precision. We propose a new data filtering method -- PCAF -- for parallel AkNN search based on principal components analysis. PCAF improves on previous methods by demonstrating sustained, high scalability for a wide range of high dimensional datasets on both Intel and AMD multicore platforms. Moreover, PCAF maintains high precision in terms of the AkNN search results.
Huan Feng, David M. Eyers, Steven Mills, Yongwei Wu 0001, Zhiyi Huang 0001
ICPP5
2016 Emender: Signal filter for trilateration based indoor localisation
abstract
Various applications of indoor localisation (e.g. tracking firemen in a burning building, or navigation for the blind) require accurate location estimates. A common indoor localisation approach using commodity mobile phones is to perform trilateration with distance estimates derived from the strength of radio signals, however, they can vary wildly especially indoors. We propose a simple filtering technique to exclude measurements which adversely effect localisation accuracy. Through experimentation in a real building, across state of the art geometries and filtering techniques, our proposed filter shows an increase in accuracy by at least 30% and decrease the time taken to estimate the location by an order of magnitude.
Paul Crane, Zhiyi Huang 0001, Haibo Zhang 0001
PIMRC2
2016 Decentralized Cooperative Piggybacking for Reliable Broadcast in the VANET
abstract
Reliably broadcasting safety information to neighboring vehicles is a big challenge in vehicular ad-hoc networks (VANETs), due to the dynamic network topology and the unreliable wireless channels. In this paper we present two decentralized cooperative schemes to enhance broadcast reliability by exploiting the advantage of message piggybacking. The key idea is to let each vehicle optimally piggyback some messages it has received when broadcasting with the expectation that the neighboring vehicles can recover its lost messages through the piggybacked messages. We first present greedy piggybacking, in which each vehicle announces its lost messages to neighboring vehicles and makes piggybacking decisions based on message losses in its neighbors. We observed that some lost messages still cannot be successfully recovered in greedy piggybacking due to the asymmetric wireless communications, and further proposed a mutual learning based scheme to overcome the drawback of greedy piggybacking. We evaluated the performance of the two schemes through trace-driven simulations, and results show that both schemes can achieve significant improvement on broadcast reliability in VANETs in comparison with the existing solutions.
Guangbing Xiao, Haibo Zhang 0001, Zhiyi Huang 0001, Yawen Chen 0001
VTC Spring3
2016 Programming models and applications for multicores and manycores
abstract
Rapid advancements in multicore and manycore chips have been a revolution within chip manufacturing, almost eradicating single-core processors. From high-end servers to mobile phones, multicore and manycore chips are steadily entering every single aspect of information technology. However, programming multicore and manycore architectures remains challenging today. To fully utilize these chips, parallel programming models that allow sequential programs and programs utilizing limited parallelism to transition to architectures with massive parallelism, while maintaining good performance and productive development, are urgently needed. This special issue contains six articles selected from the 2014 International Workshop on Programming Models and Applications for Multicores and Manycores (PMAM 2014). These articles cover issues in parallel programming languages and models, work-stealing in parallel runtime environments, heterogeneous computing, and their implications on computational patterns. In ‘Adaptive Demand-aware Work-Stealing in Multi-programmed Multi-core Architectures’ 1, Chen et al. discuss how work-stealing models can be extended to support environments where multiple parallel programs might be concurrently executing. They propose a new work-stealing algorithm called demand-aware work-stealing, which alleviates this issue by allowing programs to donate and steal idle cores as needed. In ‘Palirria: Accurate On-line Parallelism Estimation for Adaptive Work-Stealing’ 2, Varisteas and Brorsson present a self-adapting work-stealing scheduling method for nested fork/join parallelism called ‘Palirria’. The proposed approach can be used to estimate the number of utilizable workers and self-adapt accordingly. In ‘Efficient CPU-GPU cooperative computing for solving the subset-sum problem’ 3, Wan et al. study the impact of heterogeneous computing with CPUs and Graphics Processing Unit (GPUs) for solving the subset-sum problem. The authors observe that while heterogeneous computing is prominent, not much study has been performed in the simultaneous usage of both CPU and GPU resources during computation. This paper proposes an efficient CPU–GPU cooperative computing scheme for solving the subset-sum problem, which enables the full utilization of all the computing power of both CPUs and GPUs. In ‘Dynamic Partitioning-based JPEG Decompression on Heterogeneous Multicore Architectures’ 4, Sodsong et al. introduce a novel JPEG decoding scheme for heterogeneous architectures consisting of a CPU and a general-purpose GPU. They employ an offline profiling step to determine the performance of a system's CPU and GPU with respect to JPEG decoding. Then the runtime partitioning and scheduling scheme exploits task, data, and pipeline parallelism by scheduling the non-parallelizable entropy decoding task on the CPU, whereas inverse discrete cosine transformations, color conversions, and upsampling are conducted on both the CPU and the GPU. In ‘Compiler Transformation of Nested Loops for GPGPUs’ 5, Tian et al. present their experiences in creating an open-source OpenACC compiler in an industrial framework (OpenUH as a branch of Open64). They also discuss in detail the techniques that they developed for loop-scheduling reduction operations on General Purpose Graphics Processing Unit (GPGPUs). In ‘Vectorizing Unstructured Mesh Computations for Many-core Architectures’ 6, Reguly et al. present results on achieving high performance through vectorization on CPUs and the Xeon-Phi on a key class of irregular applications: unstructured mesh computations. Using Single Instruction Multiple Threads (SIMT) and Single Instruction Multiple Data (SIMD) programming models, they show how unstructured mesh computations map to OpenCL or vector intrinsics through the use of code generation techniques in the OP2 Domain Specific Library and explore how irregular memory accesses and race conditions can be organized on different hardware. We hope that the articles in this special issue will provide readers with relevant insights into the emerging parallel programming models for multicore and manycore systems. Pavan Balaji holds appointments as a Computer Scientist and Group Lead at the Argonne National Laboratory, as an Institute Fellow of the Northwestern-Argonne Institute of Science and Engineering at Northwestern University, and as a Research Fellow of the Computation Institute at the University of Chicago. He leads the Programming Models and Runtime Systems group at Argonne. His research interests include parallel programming models and runtime systems for communication and I/O on extreme-scale supercomputing systems, modern system architecture, cloud computing systems, data-intensive computing, and big-data sciences. He has nearly 150 publications in these areas and has delivered nearly 150 talks and tutorials at various conferences and research institutes. Dr Balaji is a recipient of several awards including the U.S. Department of Energy Early Career award in 2012, TEDxMidwest Emerging Leader award in 2013, Crain's Chicago 40 under 40 award in 2012, Los Alamos National Laboratory Director's Technical Achievement award in 2005, Ohio State University Outstanding Researcher award in 2005, six best paper awards, one best paper finalist, and one best poster finalist. He has served as a chair or editor for nearly 50 journals, conferences, and workshops and as a technical program committee member in numerous conferences and workshops. He is a senior member of the IEEE and a professional member of the ACM. More details about Dr Balaji are available at http://www.mcs.anl.gov/balaji. Contact him at [email protected] Zhiyi Huang is an Associate Professor at the University of Otago, New Zealand. His research interests include parallel and distributed computing, multicore architectures, parallel programming models and environments, task scheduling, operating systems, green computing, and computer networks. More details about Prof. Zhiyi Huang are available at http://www.cs.otago.ac.nz/staffpriv/hzy. Contact him at [email protected]
Pavan Balaji, Zhiyi Huang 0001
Concurr. Comput. Pract. Exp.2
2015 Quantifying the Energy Efficiency Challenges of Achieving Exascale Computing
abstract
Power and performance are two potentially opposing objectives in the design of a supercomputer, where increases in performance often come at the cost of increased power consumption and vice versa. The task of simultaneously maximising both objectives is becoming an increasingly prominent challenge in the development of future exascale supercomputers. To gain some perspective on the scale of the challenge, we analyse the power and performance trends for the Top500 and Green500 supercomputer lists. We then present the PαPW metric, which we use to evaluate the scalability of power efficiency, projecting the development of an exascale system. From this analysis, we found that when both power and performance are considered, the projected date of achieving an exascale system falls far beyond the current target of 2020.
Jason Mair, Zhiyi Huang 0001, David M. Eyers, Yawen Chen 0001
CCGRID2
2015 WRH-ONoC: A wavelength-reused hierarchical architecture for optical Network on Chips
abstract
Optical Network on Chip (ONoC) is a promising technology for the next-generation many-core chip multiprocessors owing to its tremendous advantages in low power consumption, low communication delay, and high bandwidth. In this paper we present WRH-ONoC, a novel wavelength-reused hierarchical architecture that is capable of interconnecting thousands of cores using a limited number of wavelengths while providing extremely high-throughput data communication between connected cores. In WRH-ONoC, the cores are divided into small subsystems that are interconnected using multiple λ-routers and gateways in a hierarchical manner. Each λ-router can provide non-blocking parallel communication among the directly connected cores or gateways, and all λ-routers can reuse the limited number of available wavelengths. Communications between cores in different subsystems are routed via gateways in which optical signals can change their wavelengths via optical-electrical signal conversions. For a given number of cores, we give the minimum number of levels, λ-routers, and gateways required to interconnect these cores, and derive the expected end-to-end data communication delay under the Uniform-Poisson traffic pattern. Both theoretical analysis and simulation results demonstrate that WRH-ONoC can achieve significant improvement on performance and reduction on hardware cost in comparison with the existing solutions.
Haibo Zhang 0001, Yawen Chen 0001, Zhiyi Huang 0001, Huaxi Gu
INFOCOM4
2015 Efficient Selection Algorithm for Fast k-NN Search on GPUs
abstract
k Nearest Neighbours (k-NN) search is a fundamental problem in many computer vision and machine learning tasks. These tasks frequently involve a large number of high-dimensional vectors, which require intensive computations. Recent research work has shown that the Graphics Processing Unit (GPU) is a promising platform for solving k-NN search. However, these search algorithms often meet a serious bottleneck on GPUs due to a selection procedure, called k-selection, which is the final stage of k-NN and significantly affects the overall performance. In this paper, we propose new data structures and optimization techniques to accelerate k-selection on GPUs. Three key techniques are proposed: Merge Queue, Buffered Search and Hierarchical Partition. Compared with previous works, the proposed techniques can significantly improve the computing efficiency of k-selection on GPUs. Experimental results show that our techniques can achieve an up to 4:2× performance improvement over the state-of-the-art methods.
Xiaoxin Tang, Zhiyi Huang 0001, David M. Eyers, Steven Mills, Minyi Guo
IPDPS2
2015 Scalable Multicore k-NN Search via Subspace Clustering for Filtering
abstract
k Nearest Neighbors (k-NN) search is a widely used category of algorithms with applications in domains such as computer vision and machine learning. Despite the desire to process increasing amounts of high-dimensional data within these domains, k-NN algorithms scale poorly on multicore systems because they hit a memory wall. In this paper, we propose a novel data filtering strategy for k-NN search algorithms on multicore platforms. By excluding unlikely features during the k-NN search process, this strategy can reduce the amount of computation required as well as the memory footprint. It is complementary to the data selection strategies used in other state-of-the-art k-NN algorithms. A Subspace Clustering for Filtering (SCF) method is proposed to implement the data filtering strategy. Experimental results on four k-NN algorithms show that SCF can significantly improve their performance on three modern multicore platforms with only a small loss of search precision.
Xiaoxin Tang, Zhiyi Huang 0001, David M. Eyers, Steven Mills, Minyi Guo
IEEE Trans. Parallel Distributed Syst.2
2014 Data filtering for scalable high-dimensional k-NN search on multicore systems
abstract
K Nearest Neighbors (k-NN) search is a widely used category of algorithms with applications in domains such as computer vision and machine learning. With the rapidly increasing amount of data available, and their high dimensionality, k-NN algorithms scale poorly on multicore systems because they hit a memory wall. In this paper, we propose a novel data filtering strategy, named Subspace Clustering for Filtering (SCF), for k-NN search algorithms on multicore platforms. By excluding unlikely features in k-NN search, this strategy can reduce memory footprint as well as computation. Experimental results on four k-NN algorithms show that SCF can improve their performance on two modern multicore platforms with insignificant loss of search precision.
Xiaoxin Tang, Steven Mills, David M. Eyers, Kai-Cheung Leung, Zhiyi Huang 0001, Minyi Guo
HPDC5
2014 BWS: Beacon-driven wake-up scheme for train localization using wireless sensor networks
abstract
Real-time train localization using wireless sensor networks (WSNs) offers huge benefits in terms of cost reduction and safety enhancement in railway environments. A challenging problem in WSN-based train localization is how to guarantee timely communication between the anchor sensors deployed along the track and the gateway deployed on the train with minimum energy consumption. This paper presents an energy-efficient scheme for timely communication between the gateway and the anchor sensors, in which each anchor sensor runs an asynchronous duty-cycling protocol to conserve energy and wakes up only when it goes into the communication range of the gateway. A beacon-driven wake-up scheme is designed, and we establish the upper bound on the amount of time that an anchor sensor can sleep in one duty cycle to guarantee timely wake-up once a train approaches. We also give a thorough theoretical analysis for the energy efficiency of our scheme and give the optimal amount of time that an anchor sensor should sleep in terms of minimizing the total energy consumption at each anchor sensor. We evaluate the performance of our scheme through simulations, and results show that our scheme can wake up anchor sensors timely at a very low cost on energy consumption.
Adeel Javed, Haibo Zhang 0001, Zhiyi Huang 0001, Jeremiah D. Deng
ICC3
2014 Optimal link scheduling for delay-constrained periodic traffic over unreliable wireless links
abstract
This paper investigates the problem of scheduling delay-constrained traffic in a single-hop wireless industrial network in which different source devices have different data rates. We aim to maximize the packet delivery reliability while meeting the deadline for each packet. The transmission scheduling problem is decomposed into two sub-problems: subperiod-based slot allocation and slot-based transmission scheduling. The former sub-problem is formulated as a nonlinear integer programming problem, and we present a solution with polynomial-time complexity by converting it to a linear integer programming problem. For the latter sub-problem, we demonstrate that the existence of a feasible optimal schedule depends on the order of the elements in the slot allocation vector produced by solving the former subproblem. An algorithm is designed to compute a feasible slot allocation that sustains a realizable schedule. Simulation results demonstrate that our scheme ensures each device has almost the same packet delivery rate in different report periods, which is important for maintaining the stability of control systems.
Haibo Zhang 0001, Zhiyi Huang 0001, Michael Albert 0001
INFOCOM3
2014 SIB: noise reduction in fingerprint-based indoor localisation using multiple transmission powers
abstract
Research efforts into indoor localisation have focused on improving the accuracy of location estimates. In this paper, we propose a novel approach called SIB that uses RSSI values from low-power transmissions to exclude the noisy measurements from usual high-power RSSI measurements. SIB can effectively reduce the effect of noise in fingerprint-based localisation according to our analysis on the function of power loss ratio to transmission distance. Our results, based on evaluation in a real-world environment with noisy data, show a decrease in the geometric error of 85% in our indoor localisation.
Paul Crane, Zhiyi Huang 0001, Haibo Zhang 0001
MUM2
2014 Special issue on programming models and applications for multicores and manycores - Guest Editors' Introduction
Pavan Balaji, Zhiyi Huang 0001
Parallel Comput.2
2013 A Particle Filter Based Train Localization Scheme Using Wireless Sensor Networks
abstract
Real-time train localization is essential to ensure the safety of modern railway transportation. This paper investigates the feasibility to achieve real-time and accurate train localization using wireless sensor networks. We carry out on-site experiments in a railway environment and demonstrate that Received Signal Strength Indicator (RSSI) is a good estimator for train localization. By combining the advantages of RSSI-based distance estimation and particle filtering techniques, we design a particle filter based train localization scheme and propose a novel Weighted RSSI Likelihood Function (WRLF) for updating the weights of particles. The proposed scheme is evaluated through simulations using the data obtained from the on-site measurements. Simulation results demonstrate that our scheme can achieve high localization accuracy, and is robust to changes in train speed and the deployment density of anchor sensors.
Jothi V. N. Vijayakumar, Haibo Zhang 0001, Zhiyi Huang 0001, Adeel Javed
DASC3
2013 Performance Tuning on Multicore Systems for Feature Matching within Image Collections
abstract
Parallel programming is the mainstream for today's HPC applications. Programmers need to parallelize their programs to achieve better performance on multicore systems. However, due to a lack of good understanding of parallelism in algorithms, scheduling policy in runtime systems, and multicore architectures, programmers usually find it very hard to write high-performance, scalable programs on these parallel platforms. Although using a parallelized library written by experts can reduce the amount of work for coding, it does not automatically guarantee good performance according to our study. A better understanding of parallelism in algorithms, the OS/runtime systems, and hardware architectures is necessary if programmers wish to further improve performance. In this paper, we use SIFT-based feature matching within large-scale image collections to show the importance of three factors-the level of parallelism, scheduling policy, and memory architecture-that affect the performance of large-scale feature matching on multicore systems. We demonstrate experimental results using programs based on OpenCV and OpenMP, which are executed on both 16-core and 64-core machines. From our experimental results, we find that images with a large number of features achieve poor scalability on the 64-core machine due to a poor cache utilization. To address this issue of cache performance, we propose a Divide-and-Merge algorithm that divides the feature space into several small sub-spaces so that they fit within the cache. Our experiments show that the performance tuning addressing all of the three factors improves the speedup of feature matching from 10.6× to 21.5× on the 64-core machine. While the speedup is improved by 103%, the scalability of the feature matching algorithm is improved by up to 6.45 times on the 64-core machine with our performance tuning. Our study indicates that performance tuning on multicore systems is very challenging even for a simple image processing algorithm.
Xiaoxin Tang, Steven Mills, David M. Eyers, Zhiyi Huang 0001, Kai-Cheung Leung, Minyi Guo
ICPP4
2013 Performance evaluation of View-Oriented Transactional Memory
Zhiyi Huang 0001, Kai-Cheung Leung
Parallel Comput.1
2013 Restricted admission control in view-oriented transactional memory
Kai-Cheung Leung, Yawen Chen 0001, Zhiyi Huang 0001
J. Supercomput.3
2013 Adaptive Cache Aware Bitier Work-Stealing in Multisocket Multicore Architectures
abstract
Modern multicore computers often adopt a multisocket multicore architecture with shared caches in each socket. However, traditional work-stealing schedulers tend to pollute the shared cache and incur more cache misses due to their random stealing. To relieve this problem, this paper proposes an Adaptive Cache-Aware Bi-tier work-stealing (A-CAB) scheduler. A-CAB improves the performance of memory-bound applications by reducing memory footprint and cache misses of tasks running inside the same CPU socket. A-CAB adaptively uses a DAG partitioner to divide an execution Directed Acyclic Graph (DAG) into the intersocket tier and the intrasocket tier. Tasks in the intersocket tier are scheduled across sockets while tasks in the intrasocket tier are scheduled within the same socket. Experimental results tell us that A-CAB can improve the performance of memory-bound applications up to 74.4 percent compared with the traditional work-stealing.
Quan Chen 0002, Minyi Guo, Zhiyi Huang 0001
IEEE Trans. Parallel Distributed Syst.3
2012 MOS-based Handover Protocol for Next Generation Wireless Networks
abstract
Next Generation Wireless Networks (NGWNs) are expected to provide high data rate and optimized quality of service to multimedia and real-time applications over the Internet Protocol (IP) networks. To achieve these goals, handover plays a very critical role in maintaining the seamless connectivity when mobile terminals move across different cells or networks. In this paper, we propose a novel scheme compliant with the IEEE 802.21 standard for handover in an integrated scenario with UMTS and WiMAX networks. We use the call quality, measured using Mean Opinion Score (MOS), as the major metric for handover optimization. We compare the proposed MOS-based handover scheme with the traditional RSS-based handover scheme. The numerical results demonstrate that our proposed scheme can maintain high call quality and reduce the probabilities for both handover dropping and call dropping.
Sheetal Jadhav, Haibo Zhang 0001, Zhiyi Huang 0001
AINA3
2012 CATS: cache aware task-stealing based on online profiling in multi-socket multi-core architectures
abstract
Multi-socket Multi-core architectures with shared caches in each socket have become mainstream when a single multi-core chip cannot provide enough computing capacity for high performance computing. However, traditional task-stealing schedulers tend to pollute the shared cache and incur severe cache misses due to their randomness in stealing. To address the problem, this paper proposes a Cache Aware Task-Stealing (CATS) scheduler, which uses the shared cache efficiently with an online profiling method and schedules tasks with shared data to the same socket. CATS adopts an online DAG partitioner based on the profiling information to ensure tasks with shared data can efficiently utilize the shared cache. One outstanding novelty of CATS is that it does not require any extra user-provided information. Experimental results show that CATS can improve the performance of memory-bound programs up to 74.4% compared with the traditional task-stealing scheduler.
Quan Chen 0002, Minyi Guo, Zhiyi Huang 0001
ICS3
2012 WATS: Workload-Aware Task Scheduling in Asymmetric Multi-core Architectures
abstract
Asymmetric Multi-Core (AMC) architectures have shown high performance as well as power efficiency. However, current parallel programming environments do not perform well on AMC due to their assumption that all cores are symmetric and provide equal performance. Their random task scheduling policies, such as task-stealing, can result in unbalanced workloads in AMC and severely degrade the performance of parallel applications. To balance the workloads of parallel applications in AMC, this paper proposes a Workload-Aware Task Scheduling (WATS) scheme that adopts history-based task allocation and preference-based task stealing. The history-based task allocation is based on a near-optimal, static task allocation using the historical statistics collected during the execution of a parallel application. The preference-based task stealing, which steals tasks based on a preference list, can dynamically adjust the workloads in AMC if the task allocation is less optimal due to approximation in the history-based task allocation. Experimental results show that WATS can improve the performance of CPU-bound applications up to 82.7% compared with the random task scheduling policies.
Quan Chen 0002, Yawen Chen 0001, Zhiyi Huang 0001, Minyi Guo
IPDPS3
2011 CAB: Cache Aware Bi-tier Task-Stealing in Multi-socket Multi-core Architecture
abstract
Modern multi-core computers often adopt a multi-socket multi-core architecture with shared caches in each socket. However, traditional task-stealing schedulers tend to pollute the shared cache and incur more cache misses due to their random stealing. To relieve this problem, this paper proposes a Cache Aware Bi-tier (CAB) task-stealing scheduler, which improves the performance of memory-bound applications by reducing memory footprint and cache misses of tasks running inside the same CPU socket. CAB uses an automatic partitioning method to divide an execution Directed Acyclic Graph (DAG) into the inter-socket tier and the intra-socket tier. Tasks generated in the inter-socket tier are scheduled across sockets, while tasks generated in the intra-socket tier are scheduled within the same socket. Experimental results show that CAB can improve the performance of memory-bound applications up to 68.7% compared with the traditional task-stealing.
Quan Chen 0002, Zhiyi Huang 0001, Minyi Guo, Jingyu Zhou
ICPP2
2011 Handover Delay in Mobile WiMAX: A Simulation Study
abstract
Worldwide Interoperability for Microwave Access (WiMAX) deployment is growing at a rapid pace. Since Mobile WiMAX has the key advantage of serving large coverage areas per base station, it has become a popular emerging technology for handling mobile clients. However, serving a large number of Mobile Stations (MS) in practice requires an efficient handover scheme. Currently, mobile WiMAX has a long handover delay that contributes to the overall end-to-end communication delay. Recent research is focusing on increasing the efficiency of hand over schemes. In this paper, we analyse the performance of the two standardised handover schemes, namely the Mobile IP and the ASN-based Network Mobility (ABNM), in mobile WiMAX using simulation. Our results clearly indicate that ABNM is more efficient for handover in terms of handover delay and throughput.
Bhaskar Ashoka, David M. Eyers, Zhiyi Huang 0001
PDCAT3
2011 Performance Evaluation of Quality of VoIP in WiMAX and UMTS
abstract
Next Generation Wireless Networks (NGWNs) focus on convergence of different Radio Access Technologies (RATs) providing good Quality of Service (QoS) for applications such as Voice over IP traffic (VoIP) and video streaming. The voice applications over IP networks are growing rapidly due to their increasing popularity and cost. To meet the demand of providing high-quality of VoIP at anytime and from anywhere, it is imperative to design suitable QoS model. In this paper we conduct simulation study to evaluate the QoS performance of WiMAX and UMTS for supporting VoIP. We designed simulation modules in OPNET for WiMAX and UMTS, and carried out extensive simulations to evaluate and analyze several important performance metrics such as Mean Opinion Score (MOS), end-to-end delay, jitter and packet delay variation. Simulation results show that WiMAX outscores the UMTS with a sufficient margin, and is the better technology to support VoIP applications compared with UMTS.
Sheetal Jadhav, Haibo Zhang 0001, Zhiyi Huang 0001
PDCAT3
2010 Using memory mapping to support cactus stacks in work-stealing runtime systems
abstract
Many multithreaded concurrency platforms that use a work-stealing runtime system incorporate a "cactus stack," wherein a function's accesses to stack variables properly respect the function's calling ancestry, even when many of the functions operate in parallel. Unfortunately, such existing concurrency platforms fail to satisfy at least one of the following three desirable criteria:
I-Ting Angelina Lee, Silas Boyd-Wickizer, Zhiyi Huang 0001, Charles E. Leiserson
PACT3
2010 Preface
Zhiyi Huang 0001, John H. Hine, Laurent Lefèvre, Tony McGregor
J. Supercomput.1
2010 Data race: tame the beast
K. Leung, Zhiyi Huang 0001, Paul Werstein
J. Supercomput.2
2009 Maotai 2.0: Data Race Prevention in View-Oriented Parallel Programming
abstract
This paper proposes a data race prevention scheme, which can prevent data races in the View-Oriented Parallel Programming (VOPP) model. VOPP is a novel shared-memory data-centric parallel programming model, which uses views to bundle mutual exclusion with data access. We have implemented the data race prevention scheme with a memory protection mechanism. Experimental results show that the extra overhead of memory protection is trivial in our applications. We also present a new VOPP implementation-Maotai 2.0, which has advanced features such as deadlock avoidance, producer/consumer view and system queues, in addition to the data race prevention scheme. The performance of Maotai 2.0 is evaluated and compared with modern programming models such as OpenMP and Cilk.
K. Leung, Zhiyi Huang 0001, Paul Werstein
PDCAT2
2008 Cross Layer Protocol Support for Live Streaming Media
abstract
Delivering live streaming content over the Internet requires low delay and smooth packet transmission rate. TCP introduces rate oscillations and requires more buffering and bandwidth to sustain uninterrupted playback. In this paper we propose a new framework which facilitates streaming flows. Our solution provides a smoother rate control than TCP and improves streaming performance based on cross layer feedback between the transport protocol and streaming server. We present our experimental results through simulation.
Syed Hasan, Laurent Lefèvre, Zhiyi Huang 0001, Paul Werstein
AINA3
2008 Maotai: View-Oriented Parallel Programming on CMT Processors
abstract
View-oriented parallel programming (VOPP) is a novel parallel programming model which uses views for communication between multiple processes. With the introduction of views, mutual exclusion and shared data access are bundled together, which offers both convenience and high performance to parallel programming. This paper presents the implementation of VOPP on chip-multi threading processors, e.g. UltraSPARC T1. We demonstrate that our implementation of VOPP on multi-core platforms (namely Maotai) shows significantly better performance than directly applying the original DSM implementation of VOPP (namely VODCA) on our platform. Besides, we compare the performance of VOPP with MPI and OpenMP. The experimental results demonstrate that VOPP has better scalability than both MPI and OpenMP on our platform.
Zhiyi Huang 0001
ICPP2
2008 GPU as a General Purpose Computing Resource
abstract
In the last few years, GPUs(Graphics Processing Units) have made rapid development. Their ever-increasing computing power and decreasing cost have attracted attention from both industry and academia. In addition to graphics applications, researchers are interested in using them for general purpose computing. Recently, NVIDIA released a new computing architecture, CUDA (compute united device architecture), for its GeForce 8 series, Quadro FX, and Tesla GPU products. This new architecture can change fundamentally the way in which GPUs are used. In this paper, we study the programmability of CUDA and its GeForce 8 GPU and compare its performance with general purpose processors, in order to investigate its suitability for general purpose computation.
Zhiyi Huang 0001, Paul Werstein, Martin K. Purvis
PDCAT2
2008 Virtual Aggregated Processor in Multi-core Computers
abstract
Parallel computing has been in the spotlight with the advent of multi-core computers. The popular multithreading model does not scale very well when there are hundreds or thousands of cores, since it can only help exploit coarse-grained parallelism. There exist a lot of fine-grained parallelism to be exploited in I/O tasks and memory accesses during execution of a thread. Our Counter-Amdahl's Law tells us that it is more effective to parallelize the serial fraction of a parallel algorithm rather than the parallelized fraction in order to maximize the speedup. In this paper, we have proposed a Virtual Aggregated Processor that is aiming at speeding up execution of a thread through exploiting the fine-grained parallelism in I/O tasks and memory accesses. We have proposed and implemented two techniques, helper thread and I/O specialization, to demonstrate the potential effectiveness of the Virtual Aggregated Processor technology.
Zhiyi Huang 0001, Andrew Trotman, Xiangfei Jia, Mariusz Nowostawski, Nathan Rountree, Paul Werstein
PDCAT1
2008 Application-Specific Disk I/O Optimisation for a Search Engine
abstract
Operating systems only provide general-purpose I/O optimisation since they have to service various types of applications. However, application level I/O optimisation can achieve better performance since an application has a better knowledge of how to optimise disk I/O for the application. In this paper we provide a solution for application-specific I/O for optimising a search engine. It shows a 28% improvement when compared to the general-purpose I/O optimisation of Linux. Our result also shows a 11% improvement when the Linux I/O optimisation is bypassed.
Xiangfei Jia, Andrew Trotman, Richard A. O'Keefe, Zhiyi Huang 0001
PDCAT4
2007 Revisit of View-Oriented Parallel Programming
abstract
Traditional parallel programming styles have many problems which hinder the development of parallel applications. The message passing style can be too complex for many programmers. While shared memory based parallel programming is relatively easy, it requires programmers to guarantee there is no data race in programs by using mutually exclusive locks. Data race conditions are generally difficult to debug and difficult to prevent as well. The view-oriented parallel programming (VOPP) is a novel shared-memory-based programming style. It removes the burden of guaranteeing data race free from the programmers. With the VOPP approach, shared data objects in a parallel program are divided into views according to the memory access pattern of the parallel algorithm. Data race is not an issue in VOPP, since mutual exclusion is automatically done by the underlying system when a view is accessed. The programmer only needs to synchronize the access of views using synchronization primitives like barriers. By removing data races of view access, VOPP makes it easier to code and less difficult to debug programs. It provides potential performance advantages on multi-core systems as well as cluster computers. It will also provide useful information for efficient implementation of transactional memory.
Zhiyi Huang 0001
CCGRID1
2007 Performance Evaluation of View-Oriented Parallel Programming on Cluster of Computers
Haifeng Shang, Zhiyi Huang 0001
HPCC5
2007 A Remote Memory Swapping System for Cluster Computers
abstract
This paper describes the use of remote memory for virtual memory swapping in a cluster computer. Our design uses a lightweight kernel-to-kernel communications channel for fast, efficient data transfer. Performance tests are made to compare our system to normal hard disk swapping. The tests show significantly improved performance when data access is random.
Paul Werstein, Xiangfei Jia, Zhiyi Huang 0001
PDCAT3
2006 VODCA: View-Oriented, Distributed, Cluster-Based Approach to Parallel Computing
Zhiyi Huang 0001, Martin K. Purvis
CCGRID1
2006 Load Balancing in a Cluster Computer
abstract
This paper proposes a load balancing algorithm for distributed use of a cluster computer. It uses load information including CPU queue length, CPU utilisation, memory utilisation and network traffic to decide the load of each node. This algorithm is compared to an algorithm using only the CPU queue length. The performance evaluation results show that the proposed algorithm performs well
Paul Werstein, Hailing Situ, Zhiyi Huang 0001
PDCAT3
2005 View-oriented update protocol with integrated diff for view-based consistency
abstract
This paper proposes a view-oriented update protocol with integrated diff for efficient implementation of a view-based consistency model which supports a novel view-oriented parallel programming style based on distributed shared memory. View-oriented parallel programming requires the programmer to divide the shared data into views according to the nature of the parallel algorithm and its memory access pattern. The advantage of this programming style is that it offers the potential for the underlying distributed shared memory system to optimize consistency maintenance. The View-oriented update protocol with integrated diff is proposed to exploit this performance potential. This protocol is compared with a traditional diff-based protocol and an existing home-based protocol. Experimental results demonstrate that the performance of the proposed protocol is significantly better than the diff-based protocol and the home-based protocol.
Zhiyi Huang 0001, Martin K. Purvis, Paul Werstein
CCGRID1
2005 Performance Evaluation of View-Oriented Parallel Programming
abstract
This paper evaluates the performance of a novel view-oriented parallel programming style for parallel programming on cluster computers. View-oriented parallel programming is based on distributed shared memory which is friendly and easy for programmers to use. It requires the programmer to divide shared data into views according to the memory access pattern of the parallel algorithm. One of the advantages of this programming style is that it offers the performance potential for the underlying distributed shared memory system to optimize consistency maintenance. Also it allows the programmer to participate in performance optimization of a program through wise partitioning of the shared data into views. Experimental results demonstrate a significant performance gain of the programs based on the view-oriented parallel programming style.
Zhiyi Huang 0001, Martin K. Purvis, Paul Werstein
ICPP1
2005 Performance Comparison between VOPP and MPI
abstract
View-Oriented Parallel Programming is based on Distributed Shared Memory which is friendly and easy for programmers to use. It requires the programmer to divide shared data into views according to the memory access pattern of the parallel algorithm. One of the advantages of this programming style is that it offers the performance potential for the underlying Distributed Shared Memory system to optimize consistency maintenance. Also it allows the programmer to participate in performance optimization of a program through wise partitioning of the shared data into views. In this paper, we compare the performance of View- Oriented Parallel Programming against Message Passing Interface. Our experimental results demonstrate a performance gap between View-Oriented Parallel Programming and Message Passing Interface. The contributing overheads behind the performance gap are discussed and analyzed, which sheds much light on further performance improvement of View-Oriented Parallel Programming. Key Words: Distributed Shared Memory, View-based Consistency, View-Oriented Parallel Programming, Cluster Computing, Message Passing Interface
Zhiyi Huang 0001, Martin K. Purvis, Paul Werstein
PDCAT1
2004 View-Oriented Parallel Programming and View-Based Consistency
Zhiyi Huang 0001, Martin K. Purvis, Paul Werstein
PDCAT1
2004 Locabus: A Kernel to Kernel Communication Channel for Cluster Computing
Paul Werstein, Mark Pethick, Zhiyi Huang 0001
PDCAT3
2001 View-Based Consistency and Its Implementation
abstract
The paper proposes a novel view based consistency model for distributed shared memory. A view is a set of ordinary, data objects that a processor has the right to access in a data-race-free program. The view based consistency model only requires that the data objects of a view are updated before a processor accesses them. Compared with other memory consistency models, the view based consistency model can achieve data selection without user annotation and can reduce much false-sharing effect. This model has been implemented based on TreadMarks. Performance results have shown that for all our applications, the view based consistency model outperforms the lazy release consistency model.
Zhiyi Huang 0001, Stephen Cranefield, Martin K. Purvis, Chengzheng Sun
CCGRID1
2000 Handling side-effects and cuts with selective recomputation in parallel Prolog
Zhiyi Huang 0001, Chengzheng Sun, Abdul Sattar 0001
Future Gener. Comput. Syst.1
1998 Toward Transparent Selective Sequential Consistency in Distributed Shared Memory Systems
abstract
This paper proposes a transparent selective sequential consistency approach to distributed shared memory (DSM) systems. First, three basic techniques-time selection, processor selection, and data selection-are analyzed for improving the performance of strictly sequential consistency DSM systems, and a transparent approach to achieving these selections is proposed. Then, this paper focuses on the protocols and techniques devised to achieve transparent data selection, including a novel selective lazy/eager updates propagation protocol for propagating updates on shared data objects, and the critical region updated pages set scheme to automatically detect the associations between shared data objects and synchronization objects. The proposed approach is able to offer the same potential performance advantages as the entry consistency model or the scope consistency model, but it imposes no extra burden to programmers and never fails to execute programs correctly. The devised protocols and techniques have been implemented and experimented with in the context of the TreadMarks DSM system. Performance results have shown that for many applications, our transparent data selection approach outperforms the lazy release consistency model using a lazy or eager updates propagation protocol.
Chengzheng Sun, Zhiyi Huang 0001, Wan-Ju Lei, Abdul Sattar 0001
ICDCS2
1997 Handling Side-effects with Selective Recomputation in AND/OR Parallel Execution Models
Zhiyi Huang 0001, Chengzheng Sun, Abdul Sattar 0001
ICLP1
1993 Parallel execution of prolog on shared-memory multiprocessors
Yaoqing Gao, Dingxing Wang, Meiming Shen, Zhiyi Huang 0001, Shouren Hu, Giorgio Levi
J. Comput. Sci. Technol.5