Yawen Chen 0001

dblp:30/2325 · DBLP profile ↗
← Back
58ranked-venue papers
8as first author
27since 2021 · last 2026
0000-0001-7006-2459ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 29 · 5 first-author · 14 since 2021Computer networks · 16 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Security and privacy · 1Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CD-DPC: Centrifugal degree based density peaks clustering algorithm
Linlin Ma, Xincheng Liu, Huihui Chu, Yuzhen Zhao, Yawen Chen 0001, Wenke Zang
Pattern Recognit.7
2025 BITLUME: Precision-Flexible Photonic Computing for Ultra-Fast and Energy-Efficient DNN Acceleration
abstract
As deep learning expands across emerging domains, computational demands are pushing traditional electronic accelerators to their limits. Silicon photonics has emerged as a promising technology for accelerating deep learning workloads, but precision remains a challenge due to noise and non-idealities. In this paper, we present BITLUME, a novel photonic computing unit that enables multiplications beyond 8-bit precision through a precision-flexible scheme. We further propose an optimized round-truncation algorithm and data mapping strategy for BITLUME to reduce optoelectronic conversions, enhance data reuse, and maintain computational accuracy. A hybrid optoelectronic architecture integrating BITLUME is developed and validated using a prototype built with FPGA, RF, and photonic components, achieving 3.7× lower end-to-end latency than the A100 GPU in dot product. Simulations of training seven DNN models at FP32 show that BITLUME achieves up to 3.35× and 10.78× speedup, and 1.53× and 4.12× energy savings, compared to the state-of-the-art photonic accelerator and A100 GPU, respectively.
Chengpeng Xia, Haibo Zhang 0001, Hao Zhang 0058, Yawen Chen 0001, Amanda S. Barnard
ICCAD4
2025 ROCKET: An RNS-based Photonic Accelerator for High-Precision and Energy-Efficient DNN Training
abstract
In recent years, the rapid development of Deep Neural Networks (DNNs) has posed significant challenges in terms of training duration and costs. High-frequency, low-power photonic computing has emerged as a highly promising solution. However, the substantial cost of data conversion and the limitations introduced by noise in photonic devices continue to hinder the realization of high-precision and energy-efficient DNN training. To address this challenge, we propose a novel photonic accelerator, ROCKET, based on the Residue Number System (RNS). RNS is based on modular arithmetic and enables support for high-precision computation through parallel multi-path low-precision operations. First, we leverage specialized lookup tables to enable high-throughput, low-latency conversions between high-precision and low-precision numerical representations. Next, we design a low-power photonic accelerator architecture utilizing intensity modulators, which minimizes the number of computational components while maximizing data reuse. Subsequently, we propose a hybrid photonic-electronic pipelined dataflow to maximize parallelism within the photonic-electronic computation path. Finally, we develop a high-frequency (4.096 GHz) hybrid photonic-electronic prototype using FPGA, Radio Frequency (RF), and photonic components to validate the feasibility of the ROCKET. Our large-scale simulations on seven mainstream DNN models show that, compared to the A100 GPU, TPU v4, and the state-of-the-art photonic accelerator Mirage, ROCKET achieves speedups of 33×, 243×, and 198×, respectively, while saving energy by factors of 64×, 204×, and 142×.
Hao Zhang 0058, Haibo Zhang 0001, Chengpeng Xia, Zhiyi Huang 0001, Yawen Chen 0001, Amanda S. Barnard
ICS5
2025 TextCG: A text classification framework based on information fusion via cross-graph attention network and gated recurrent unit
Wenke Zang, Wenbin Ma, Yuzhen Zhao, Yuanhua Wang, Xiyu Liu 0001, Yawen Chen 0001
Inf. Sci.7
2025 ChipAI: A scalable chiplet-based accelerator for efficient DNN inference using silicon photonics
abstract
To enhance the precision of inference, deep neural network (DNN) models have been progressively growing in scale and complexity, leading to increased latency and computational resource demands. This growth necessitates scalable architectures, such as chiplet-based accelerators, to accommodate the substantial volume of deep learning inference tasks. However, the efficiency, energy consumption, and scalability of existing accelerators are severely constrained by metallic interconnects. Photonic interconnects, on the contrary, offer a promising alternative, with their advantages of low latency, high bandwidth, high energy efficiency, and simplified communication processes. In this paper, we propose ChipAI, an accelerator designed based on photonic interconnects for accelerating DNN inference tasks. ChipAI implements an efficient hybrid optical network that supports effective inter-chiplet and intra-chiplet data sharing, thereby enhancing parallel processing capabilities. Additionally, we propose a flexible dataflow leveraging the ChipAI architecture and the characteristics of DNN models, facilitating efficient architectural mapping of DNN layers. Simulation on various DNN models demonstrates that, compared to the state-of-the-art chiplet-based DNN accelerator with photonic interconnects, ChipAI can reduce the DNN inference time and energy consumption by up to 82% and 79%, respectively.
Hao Zhang 0058, Haibo Zhang 0001, Zhiyi Huang 0001, Yawen Chen 0001
J. Syst. Archit.4
2025 Virtual regularized bipartite graph learning for multi-view subspace clustering
Linlin Ma, Wenke Zang, Xincheng Liu, Yuzhen Zhao, Xiyu Liu 0001, Zhenni Jiang, Baoqiang Yan, Yawen Chen 0001
Knowl. Based Syst.8
2024 Power-Adaptive Communication With Channel-Aware Transmission Scheduling in WBANs
abstract
Radio links in Wireless Body Area Networks (WBANs) are highly subject to short and long-term attenuation due to the unstable network topology and frequent body blockage. This instability makes it challenging to achieve reliable and energy-efficient communication, but on the other hand, provides a great potential for the sending nodes to dynamically schedule the transmissions at the time with the best-expected channel quality. Motivated by this, we propose IGE (Improved Gilbert-Elliott Markov chain model), a memory-efficient Markov chain model to monitor channel fluctuations and provide a long-term channel prediction. We then design ATPS (Adaptive Transmission Power Selection), a deadline-constrained channel scheduling scheme that enables a sending node to buffer the packets when the channel is bad and schedule them to be transmitted when the channel is expected to be good within a deadline. ATPS can self-learn the pattern of channel changes without imposing a significant computation or memory overhead on the sending node. We evaluate the performance of ATPS through experiments using TelosB motes under different scenarios with different body postures and packet rates. We further compare ATPS with several state-of-the-art schemes including the optimal scheduling policy in which the optimal transmission time for each packet is calculated based on the collected RSSI (Received Signal Strength Indicator) samples in an off-line manner. The experimental results reveal that ATPS performs almost as efficiently as the optimal scheme in high-date-rate scenarios and has a similar trend on power level usage.
Abbas Arghavani, Haibo Zhang 0001, Zhiyi Huang 0001, Yawen Chen 0001
IEEE Internet Things J.4
2024 Density peaks clustering based on density voting and neighborhood diffusion
Wenke Zang, Jing Che, Linlin Ma, Xincheng Liu, Aoyu Song, Jingwen Xiong, Yuzhen Zhao, Xiyu Liu 0001, Yawen Chen 0001
Inf. Sci.9
2024 Routing and Wavelength Assignment Algorithm for Mesh-based Multiple Multicasts in Optical Network-on-chip
Cui Yu, Yawen Chen 0001, Boyong Gao
Theory Comput. Syst.3
2023 Performance Comparison of Distributed DNN Training on Optical Versus Electrical Interconnect Systems
Yawen Chen 0001, Zhiyi Huang 0001, Haibo Zhang 0001, Hui Tian 0001
ICA3PP (1)2
2023 OpTree: An Efficient Algorithm for All-gather Operation in Optical Interconnect Systems
abstract
All-gather collective communication is one of the most important communication primitives in parallel and distributed computation, which plays an essential role in many high performance computing (HPC) applications such as distributed Deep Learning (DL) with model and hybrid parallelisms. To solve the communication bottleneck of All-gather, optical interconnection network can provide unprecedented high bandwidth and reliability for data transfer among the distributed nodes. However, most traditional All-gather algorithms are designed for electrical interconnection, which cannot fit well for optical interconnect systems, resulting in poor performance. This paper proposes an efficient scheme, called OpTree, for All-gather operation on optical interconnect systems. OpTree derives an optimal m-ary tree corresponding to the optimal number of communication stages, which achieves the minimum communication time. We further analyze and compare the communication steps of OpTree with existing All-gather algorithms. Theoretical results exhibit that OpTree requires much less number of communication steps than existing All-gather algorithms on optical interconnect systems. Simulation results show that OpTree can reduce communication time by 72.21 %, 94.30%, and 88.58% compared to three existing All-gather schemes Wrht, Ring, and NE, respectively.
Yawen Chen 0001, Zhiyi Huang 0001, Haibo Zhang 0001
ICC2
2023 Wrht: Efficient All-reduce for Distributed DNN Training in Optical Interconnect Systems
abstract
Communication efficiency is crucial for accelerating distributed deep neural network (DNN) training. All-reduce, a vital communication primitive, is responsible for reducing model parameters in distributed DNN training. However, most existing All-reduce algorithms, designed for traditional electrical interconnect systems, fall short due to bandwidth limitations. Optical interconnects, with superior bandwidth, low transmission delay, and less power consumption, emerge as viable alternatives. We propose Wrht (Wavelength Reused Hierarchical Tree), an efficient scheme for implementing the All-reduce operation in optical interconnect systems. Wrht leverages wavelength-division multiplexing (WDM) to minimize the communication time in distributed data-parallel DNN training. We calculate the required wavelengths, minimum communication steps, and optimal communication time, considering optical communication constraints. Simulations with real-world DNN models indicate that Wrht notably reduces communication time. On average, compared with three conventional All-reduce algorithms, Wrht achieves reductions of 65.23%, 43.81%, and 82.22% respectively in optical interconnect systems, and 61.23% and 55.51% compared with two algorithms in electrical systems. This highlights Wrht’s potential to enhance communication efficiency in DNN training using optical interconnects.
Yawen Chen 0001, Zhiyi Huang 0001, Haibo Zhang 0001
ICPP2
2023 SEECHIP: A Scalable and Energy-Efficient Chiplet-based GPU Architecture Using Photonic Links
abstract
The continuous increase in GPU performance benefits a wide range of high-performance computing (HPC) applications. Slower growth of transistor density and limited size of chip die are now posing significant challenges to scale GPUs. The chiplet technology provides a potential solution to surpass these limitations. However, the performance of these chiplet-based GPUs is often constrained by the metallic-based interconnects between the chiplets. Emerging technologies such as photonic interconnect can overcome the limitations of metallic interconnects, offering several superior properties, such as high bandwidth density and low energy consumption. In this paper, we propose SEECHIP: a Scalable and Energy-Efficient CHIPlet-based GPU architecture using photonic links. SEECHIP introduces a novel photonic inter-chiplet network that supports both unicast and broadcast communication, providing the same transmission bandwidth at both the sending and receiving ends. In addition, we propose a tailored hierarchical memory architecture, which is more suitable for the parallelization of general-purpose HPC applications. Simulation results using 14 benchmarks show that SEECHIP can achieve and reduction in execution time and energy consumption, respectively, as compared to other GPUs with metallic or photonic interconnects. Simulation results also show that SEECHIP has good scalability compared with the other GPUs.
Hao Zhang 0058, Yawen Chen 0001, Zhiyi Huang 0001, Haibo Zhang 0001
ICPP2
2023 Efficient All-Reduce for Distributed DNN Training in Optical Interconnect Systems
abstract
All-reduce is the crucial communication primitive to reduce model parameters in distributed Deep Neural Networks (DNN) training. Most existing all-reduce algorithms are designed for traditional electrical interconnect systems, which cannot meet the communication requirements for distributed training of large DNNs due to the low data bandwidth of the electrical interconnect systems. One of the promising alternatives for electrical interconnect is optical interconnect, which can provide high bandwidth, low transmission delay, and low power cost. We propose an efficient scheme called WRHT (Wavelength Reused Hierarchical Tree) for implementing all-reduce operation in optical interconnect systems. WRHT can take advantage of WDM (Wavelength Division Multiplexing) to reduce the communication time of distributed data-parallel DNN training. Simulations using real DNN models show that, compared to all-reduce algorithms in the electrical and optical network systems, our approach reduces communication time by 75.76% and 91.86%, respectively.
Yawen Chen 0001, Zhiyi Huang 0001, Haibo Zhang 0001, Fangfang Zhang 0002
PPoPP2
2023 Decentralized piggybacking-based dissemination of Cooperative Awareness Messages in vehicular ad-hoc networks
Guangbing Xiao, Haibo Zhang 0001, Zhiyi Huang 0001, Yawen Chen 0001
Comput. Networks4
2023 ESA: An efficient sequence alignment algorithm for biological database search on Sunway TaihuLight
Hao Zhang 0058, Zhiyi Huang 0001, Yawen Chen 0001, Jianguo Liang, Xiran Gao
Parallel Comput.3
2023 Routing and Wavelength Assignment for Multiple Multicasts in Optical Network-on-Chip (ONoC)
abstract
Optical network-on-chip (ONoC) is an emerging chip-scale optical interconnection technology to realize high-performance and power-efficient intercore communication for many-core processors. Multicast communication is popularly used in parallel applications on chip. However, existing researches for multicast in ONoC mainly focus on the optimization of one multicast. This limits the practical applications of the research outcomes because we often face the dynamic formation of multiple multicast groups in real network systems. In this article, we define the problem of routing and wavelength assignment for multiple multicasts in ONoC with the objective of minimizing the number of wavelengths required. To solve the problem, we first formulate it as an integer programming model for general topologies. Then we design routing policies for special instances that optimally use only one wavelength on mesh topology. For general instances, we design a group-partitioning routing algorithm for multiple multicasts (GPRMM). GPRMM decouples a group of multicasts into a number of subgroups, each of which matching one of the special instances. Theoretical results show that the number of wavelengths required by GPRMM is no more than the Destination Density$\sigma _{d}$, i.e., the maximum number of multicasts with destinations in the same row or column. Moreover, we find the upper bound and the lower bound on the number of wavelengths required for GPRMM. The wavelength requirement is also upper bounded by the network size$n$for an$n\times n$mesh network. Simulation results show that GPRMM can reduce the number of wavelengths by 26.7% compared with previous methods. GPRMM has the advantages of low routing complexity, low wavelength requirement, low power consumption, and good scalability.
Yawen Chen 0001, Zhiyi Huang 0001, Haibo Zhang 0001, Huaxi Gu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 STADIA: Photonic Stochastic Gradient Descent for Neural Network Accelerators
abstract
Deep Neural Networks (DNNs) have demonstrated great success in many fields such as image recognition and text analysis. However, the ever-increasing sizes of both DNN models and training datasets make deep leaning extremely computation- and memory-intensive. Recently, photonic computing has emerged as a promising technology for accelerating DNNs. While the design of photonic accelerators for DNN inference and forward propagation of DNN training has been widely investigated, the architectural acceleration for equally important backpropagation of DNN training has not been well studied. In this paper, we propose a novel silicon photonic-based backpropagation accelerator for high performance DNN training. Specifically, a general-purpose photonic gradient descent unit named STADIA is designed to implement the multiplication, accumulation, and subtraction operations required for computing gradients using mature optical devices including Mach-Zehnder Interferometer (MZI) and Mircoring Resonator (MRR), which can significantly reduce the training latency and improve the energy efficiency of backpropagation. To demonstrate efficient parallel computing, we propose a STADIA-based backpropagation acceleration architecture and design a dataflow by using wavelength-division multiplexing (WDM). We analyze the precision of STADIA by quantifying the precision limitations imposed by losses and noises. Furthermore, we evaluate STADIA with different element sizes by analyzing the power, area and time delay for photonic accelerators based on DNN models such as AlexNet, VGG19 and ResNet. Simulation results show that the proposed architecture STADIA can achieve significant improvement by 9.7× in time efficiency and 147.2× in energy efficiency, compared with the most advanced optical-memristor based backpropagation accelerator.
Chengpeng Xia, Yawen Chen 0001, Haibo Zhang 0001, Jigang Wu
ACM Trans. Embed. Comput. Syst.2
2023 Comparing the performance of multi-layer perceptron training on electrical and optical network-on-chips
Yawen Chen 0001, Zhiyi Huang 0001, Haibo Zhang 0001, Hao Zhang 0058, Chengpeng Xia
J. Supercomput.2
2023 h-Restricted H-structure connectivity and h-restricted H-substructure connectivity of hypercube
Cui Yu, Boyong Gao, Yawen Chen 0001
J. Supercomput.4
2023 Tuatara: Location-Driven Power-Adaptive Communication for Wireless Body Area Networks
abstract
Radio links in wireless body area networks (WBANs) suffer from both short-term and long-term variations due to the dynamic network topology and frequent blockage caused by body movements, making it challenging to achieve reliable, energy-efficient and real-time data communication. Through experiments with TelosB motes, we observe a strong positive relationship between the channel quality and the location of the sensor node relative to the gateway. Motivated by this observation, we design Tuatara, a novel power-aware communication protocol that allows each sensor node to dynamically adjust its transmission power based on the channel status inferred from its instant location, aiming to save energy, reduce interference, and improve communication reliability. Combining the orientations measured by motion sensors with the anatomical constraints of body movements, each sensor node can locally estimate its instant location relative to the gateway. Based on a probabilistic model, power level selection is converted to calculate the optimal probability of selecting each power level at a given location, with the objective of minimizing the transmission cost. A learning scheme is designed to adaptively update the power level selection probabilities, making Tuatara self-adaptable to changes in the signal propagation environment. Experimental results demonstrate that Tuatara outperforms the state-of-the-art protocols in various scenarios, with performance close to that of the optimal power selection solution even in scenarios where the packet rate is very low.
Abbas Arghavani, Haibo Zhang 0001, Zhiyi Huang 0001, Yawen Chen 0001
IEEE Trans. Mob. Comput.4
2022 Three-stage auction scheme for computation offloading on mobile blockchain with edge computing
abstract
Summary Blockchain has been applied in wide range of fields to guarantee security. However, it has been very challenging for blockchain to flourish in mobile environment with limited resources. Existing studies mainly assume that single mobile user can buy the whole resources from edge servers in mobile blockchain. This paper formulates the problem of maximizing the social welfare for computation offloading in mobile blockchain. A three‐stage auction scheme with approximation ratio of based on group‐buying mechanism is proposed to allocate edge server resources for mobile blockchain applications. In the first stage, the miners are divided into groups, and a Vickrey–Clarke–Groves based auction is proposed to determine the bid of each group for each edge server. In the second stage, a matching algorithm is proposed to match edge servers and Access Points for maximizing the profit of edge servers. In the third stage, the edge server resources are allocated to mobile users for mining base on the results in the above stages. We prove that our auction scheme guarantees truthfulness, individual rationality and budget balance. Simulation results show that, the social welfare of our scheme is improved by 33.78%, 21.84%, 19.69%, and 6.69% for 1000 miners, compared with the existing works.
Chengpeng Xia, Yalan Wu, Long Chen 0006, Yawen Chen 0001, Jigang Wu
Concurr. Comput. Pract. Exp.4
2022 Energy-Efficient Non-Orthogonal Multiple Access for Downlink Communication in Mobile Edge Computing Systems
abstract
Downlink mobile edge computing (MEC) networks are requiblack to serve increasing large number of Internet of Things (IoT) devices with limited battery capacity. In order to serve massive user equipments with low power consumption requirements, in this paper, we propose an energy-efficient multi-carrier non-orthogonal multiple access (MC-NOMA) design which allows more than two IoT devices to multiplex and access the same subcarrier band. With the aim to minimize the total transmit energy while meeting the demands of each IoT device such as the low latency, in our design, we first derive the optimal successive interference cancellation (SIC) policy and minimum power allocated to every IoT device. Then we propose an optimal greedy algorithm to allocate the frequency blocks, and formulate the optimization of the computational resource allocation as a min-max problem. Subsequently, we characterize the MC-NOMA network with the potential game model, and present a scheduling scheme to manage massive IoT devices. Simulation results demonstrate that our proposed scheme can consume 3-10 dB less energy in a MEC network deployed with 256 IoT devices compablack with the conventional orthogonal multiple access (OMA) scheme and non-orthogonal multiple access (NOMA) scheme.
Lin Zhang 0023, Furong Fang, Guixun Huang, Yawen Chen 0001, Haibo Zhang 0001, Yuan Jiang 0008, Weibin Ma
IEEE Trans. Mob. Comput.4
2022 Low RF-Complexity Digital Transmit Beamforming for Large-Scale Millimeter Wave MIMO Systems
abstract
Digital beamforming (DBF) with full Radio-Frequency (RF) complexity requires a separate RF chain per transmit antenna, which is not only difficult to realize for large-scale MIMO systems but also vulnerable to hardware imperfection. In this paper, we show that DBF architectures don’t necessarily require the number of RF chains to be the same as the number of transmit antennas. We propose a low RF-complexity transmit DBF architecture that enables to use a small number of RF chains to serve a large number of transmit antennas. In our architecture, the transmit antennas are divided into groups. All antennas in the same group share the same RF chain in a time-multiplexed manner to preserve the signal from baseband till the antenna aperture. A novel antenna grouping algorithm is proposed to dynamically group the transmit antennas in a way that each antenna group forms a close-to-rank-one channel matrix with the receive antennas. Under ideal hardware conditions, we show that our DBF architecture can achieve nearly the same performance as that of the conventional full RF-complexity DBF in terms of bandwidth and spectral efficiency. However, when hardware imperfection is considered, our architecture outperforms full RF-complexity DBF in both spectral and energy efficiency because it is more robust to inter RF-chain cross-talk effects due to the reduction on the number of required RF chains.
Haibo Zhang 0001, Yawen Chen 0001, Naveed Iqbal 0003
IEEE Trans. Wirel. Commun.3
2021 Performance Comparison of Multi-layer Perceptron Training on Electrical and Optical Network-on-Chips
Yawen Chen 0001, Zhiyi Huang 0001, Haibo Zhang 0001
PDCAT2
2021 Photonic Computing and Communication for Neural Network Accelerators
Chengpeng Xia, Yawen Chen 0001, Haibo Zhang 0001, Hao Zhang 0058, Jigang Wu
PDCAT2
2021 An efficient shortest path algorithm for content-based routing on 2-D mesh accelerator networks
Huaxi Gu, Wenting Wei, Yawen Chen 0001
Future Gener. Comput. Syst.5
2020 Full Digital Transmit Beamforming with Low RF Complexity for Large-scale mmWave MIMO system
abstract
Conventional full digital transmit beamforming requires a separate Radio-Frequency (RF) chain per transmit antenna, which is not easy to realize for large-scale antenna systems due to the high hardware cost, system complexity and power consumption. In this paper, we propose a new scheme that enables full digital transmit beamforming with low RF complexity. We consider a multiple-input-multiple-output (MIMO) system with N transmit antennas and M receive antennas where N > M, and transmit antennas are connected to L RF chains, where L <; N. In our scheme, transmit antennas are dynamically divided into L groups based on the multipath channel spatial correlation profile. The beamforming gain is achieved by letting highly correlated antennas in the same group. Each group of antennas is connected to one out of L RF chains and each group transmits an independent data stream. RF chain is designed in such a way that it independently preserves baseband signals of a group of antennas in a time-division-multiplexing (TDM) way. Our scheme is evaluated in simulations by comparing with full RF complexity digital transmit beamforming. Simulation results demonstrate that, our proposal is suitable for doubly massive MIMO systems, particularly in line-of-sight (LOS) scenarios.
Haibo Zhang 0001, Yawen Chen 0001, Naveed Iqbal 0003
ICC3
2019 Chimp: A Learning-based Power-aware Communication Protocol for Wireless Body Area Networks
abstract
Radio links in wireless body area networks (WBANs) commonly experience highly time-varying attenuation due to the dynamic network topology and frequent occlusions caused by body movements, making it challenging to design a reliable, energy-efficient, and real-time communication protocol for WBANs. In this article, we present Chimp, a learning-based power-aware communication protocol in which each sending node can self-learn the channel quality and choose the best transmission power level to reduce energy consumption and interference range while still guaranteeing high communication reliability. Chimp is designed based on learning automata that uses only the acknowledgment packets and motion data from a local gyroscope sensor to infer the real-time channel status. We design a new cost function that takes into account the energy consumption, communication reliability and interference and develop a new learning function that can guarantee to select the optimal transmission power level to minimize the cost function for any given channel quality. For highly dynamic postures such as walking and running, we exploit the correlation between channel quality and motion data generated by a gyroscope sensor to fastly estimate channel quality, eliminating the need to use expensive channel sampling procedures. We evaluate the performance of Chimp through experiments using TelosB motes equipped with the MPU-9250 motion sensor chip and compare it with the state-of-the-art protocols in different body postures. Experimental results demonstrate that Chimp outperforms existing schemes and works efficiently in most common body postures. In high-date-rate scenarios, it achieves almost the same performance as the optimal power assignment scheme in which the optimal power level for each transmission is calculated based on the collected channel measurements in an off-line manner.
Abbas Arghavani, Haibo Zhang 0001, Zhiyi Huang 0001, Yawen Chen 0001
ACM Trans. Embed. Comput. Syst.4
2019 Wavelength-Reused Hierarchical Optical Network on Chip Architecture for Manycore Processors
abstract
Manycore processor is becoming the mainstream platform for cloud computing applications. However, the design of high-performance and sustainable inter-core communication network is still a challenging problem. Optical Network on Chip (ONoC) is an emerging chip-scale optical communication technology with high bandwidth capacity and energy efficiency. In this paper, we present a Wavelength Reused Hierarchical ONoC architecture, WRH-ONoC. It leverages the nonblocking wavelength-routed λ-router and hierarchical networking to reuse the limited number of wavelengths. In WRH-ONoC, all the cores are grouped into multiple subsystems, and the cores in the same subsystem are directly interconnected using a λ-router for nonblocking communication. For inter-subsystem communication, all subsystems are further connected through multiple λ-routers and gateways in a hierarchical manner. Thus, the available wavelengths can be reused in different λ-routers. Furthermore, WRHm-ONoC, an efficient extension with multicast ability is also proposed. Given the numbers of cores and available wavelengths, we derive the minimum hardware requirement, the expected end-to-end delay, and the maximum data rate. Theoretical analysis and simulation results indicate WRH-ONoC achieves prominent improvement on the communication performance and sustainability, e.g., 46.0 percent of reduction on zero-load delay and 72.7 percent of improvement on throughput for 400 cores with the modest hardware/energy costs.
Haibo Zhang 0001, Yawen Chen 0001, Zhiyi Huang 0001, Huaxi Gu
IEEE Trans. Sustain. Comput.3
2018 User Localization Using Random Access Channel Signals in LTE Networks with Massive MIMO
abstract
Recent studies show that real-time precise user localization enables to deliver accurate beamforming in MIMO systems without the need for channel estimation. This paper presents new solutions for accurate user localization in massive MIMO LTE systems. A key novelty of the developed schemes is the ability to locate users during LTE's random access channel synchronization procedure before they are connected to the network, by which the obtained location information can be immediately used to optimize the allocation of radio resource and perform accurate beamforming. To achieve this, the developed solutions leverage the advantages of spherical wave propagation since it allows simultaneously estimating the angle of arrival and the propagation distance from the user equipment to each antenna element in the base station. We design solutions for both single-path line-of-sight communication and multi-path propagation environments. The developed schemes were evaluated through both simulations and proof-of-concept experiments. Simulation results show that both algorithms can achieve decimeter-level localization accuracy using 64 and more antenna elements for the distances up to 300 meters. The proof-of-concept experiment justifies the feasibility of user localization based on the estimation of the shape of the incoming wavefront.
Aleksei Fedorov, Haibo Zhang 0001, Yawen Chen 0001
ICCCN3
2018 A joint optimization method for NoC topology generation
Kun Wang 0001, Huaxi Gu, Yintang Yang, Yawen Chen 0001, Haibo Zhang 0001
J. Supercomput.6
2017 Testudo: A Low Latency and High-Efficient Memory-Centric Network Using Optical Interconnect
abstract
With the continuing-scaling of future multicore processors, the performance requirements on memory access has been put forward much higher. Memory- centric network is deemed as a promising communication paradigm for core-to-memory interconnect in future multicore processors. However, the traditional electrical interconnect has the drawbacks of limited capacity, high communication delay, poor scalability and low energy efficiency, which further limits the performance improvement of the system. To support high-performance communication for memory access, we propose the Testudo architecture, an optically connected memory-centric network (MCN), which utilizes the emerging optical interconnect technology and 3D-stacking memory technology to achieve high bandwidth, low power consumption and high scalability. Testudo is designed based on multiple optical crossbar organized in a torus- like topology. Each optical crossbar is in multiple-write- multiple-read construction, which provides high connectivity for IP cores. By employing an all optical, token-based arbitration scheme with low complexity, the memory access communication is contention-free. Simulation results show that Testudo improves the performance significantly compared to the electrical mesh topology.
Shixiong Qi, Huaxi Gu, Haibo Zhang 0001, Yawen Chen 0001
GLOBECOM4
2017 Geometry-based modeling and simulation of 3D multipath propagation channel with realistic spatial characteristics
abstract
For high-resolution massive MIMO and very large antenna arrays, wireless channel models have to scrutinize the detailed space features of the surrounding environment. Existing models such as WINNER and 3GPP are not appropriate for validating and evaluating new concepts for 4G/5G as they do not consider the spatial characteristics of the real environment. The simplified 3D shapes of simulated objects in geometry-based channel models, which are constructed using vertical and horizontal planes, may cause significant difference from the real channel. In this paper, we present an approach to model the specular reflection of a signal from an arbitrary inclined surface by taking into account the signal's polarization and a spatial distribution of massive MIMO antenna elements. The approach was validated through simulating LTE uplink transmissions in an environment modeled based on Google Maps. Results showed the importance of considering detailed 3D characteristics of the surroundings in simulations. We observed that even slightly inclined walls can have significant influence on channels in comparison with models with only vertical and horizontal surfaces due to different propagation paths, different angles of reflection, and different changes of polarizations.
Aleksei Fedorov, Haibo Zhang 0001, Yawen Chen 0001
ICC3
2017 ATPS: Adaptive Transmission Power Selection for Communication in Wireless Body Area Networks
abstract
Since radio links in wireless body area networks (WBANs) commonly experience highly time-varying attenuation due to topology instability, communication protocols with fixed transmission power cannot produce a very good performance in terms of energy consumption, interference range, and communication reliability. We explain that how channel behaviourcan be modelled using Markov Chain. Then, a power-adaptive communication protocol for WBANs is developed in which each sensor node can self-learn its channel and dynamically adjust itstransmission power. We evaluate our scheme through implementing the idea using the TelosB motes. The results demonstrate that our scheme can self-learn the channel behaviours, and reduce energy consumption and interference.
Abbas Arghavani, Haibo Zhang 0001, Zhiyi Huang 0001, Yawen Chen 0001
LCN4
2017 A cooperative offloading game on data recovery for reliable broadcast in VANET
abstract
Summary The rapidly growing demand for accident‐free driving in intelligent transportation makes reliable broadcast a critical factor for vehicularad hocnetworks. Existing solutions always try to improve the broadcast reliability by retransmitting lost packets. However, the excessive retransmissions can easily cause unpredictable time delay and even broadcast storms, rendering the reliable broadcast problem unsolved. In this paper, a novel reliable broadcast scheme is proposed by exploring the advantages of lost data piggybacking. Our scheme allows all the vehicles to piggyback some received packets cooperatively to help other vehicles to recover the lost packets. We formulate the cooperative piggybacking problem as a cooperative offloading game and present a decentralized solution to compute the optimal data piggybacking solutions based on only partial network information. A reward‐penalty scheme is designed for the offloading process to impel all the vehicles' decisions that converge to the Nash equilibrium, which is proved to be the global optimal solution to the decentralized offloading scheme. Simulation results show that the proposed cooperative offloading scheme can achieve much higher broadcast reliability and lower propagation delay, in comparison with existing solutions. In a small vehicle network, all lost cooperative awareness messages can be successfully recovered within 25 ms after the initial broadcast by using the data traces generated by GEMV2. Copyright © 2016 John Wiley & Sons, Ltd.
Guangbing Xiao, Haibo Zhang 0001, Houcine Hassan, Yawen Chen 0001, Zhiyi Huang 0001, Ning Sun 0007
Concurr. Comput. Pract. Exp.4
2017 3D network-on-chip design for embedded ubiquitous computing systems
Huaxi Gu, Yawen Chen 0001, Yintang Yang, Kun Wang 0001
J. Syst. Archit.3
2017 Adaptive Message Routing and Replication in Mobile Opportunistic Networks for Connected Communities
abstract
Mobile opportunistic networking is a promising technology that can supplement existing cellular and WiFi networks to provide desirable services for smart and connected communities. Message routing is the most compelling challenge in mobile opportunistic networks due to the lack of contemporaneous end-to-end paths and the resource constraints at mobile devices. To improve the probability of successful message delivery, most existing routing schemes use the past contact history to predict future contacts for message forwarding, and exploit message replication and redundancy for multicopy routing. However, most existing prediction-based routing schemes simply use the average pairwise contact probability as the routing metric and neglect the benefits of exploring fine-grained contact information such as pairwise repeated contact patterns to improve the accuracy of predicting future contacts. Moreover, there is no efficient mechanism that can adaptively control message replication in a decentralized manner to achieve both high probability of successful message delivery and low message overhead. To address these problems, we present FGAR, a routing protocol designed for mobile opportunistic networks by leveraging fine-grained contact characterization and adaptive message replication. In FGAR, contact history is characterized in a fine-grained manner with timing information using a sliding window mechanism, and future contacts are predicted based on the fine-grained contact information, thereby improving the accuracy of contact prediction. We further design an efficient message replication scheme in which message replication is controlled in a fully decentralized manner by taking into account the expected message delivery probability, the replication history, and the quality of the encountered device. A replica is generated only when it is necessary to fulfill the expected message delivery probability. We evaluate our scheme through trace-driven simulations, and the simulation results show that FGAR outperforms existing schemes. In comparison with PRoPHET, FGAR can achieve more than 20% improvement on average on successful message delivery, whereas the message overhead has been reduced by a factor up to 15.
Haibo Zhang 0001, Luming Wan, Yawen Chen 0001, Laurence T. Yang, Lizhi Peng
ACM Trans. Internet Techn.3
2016 Decentralized Cooperative Piggybacking for Reliable Broadcast in the VANET
abstract
Reliably broadcasting safety information to neighboring vehicles is a big challenge in vehicular ad-hoc networks (VANETs), due to the dynamic network topology and the unreliable wireless channels. In this paper we present two decentralized cooperative schemes to enhance broadcast reliability by exploiting the advantage of message piggybacking. The key idea is to let each vehicle optimally piggyback some messages it has received when broadcasting with the expectation that the neighboring vehicles can recover its lost messages through the piggybacked messages. We first present greedy piggybacking, in which each vehicle announces its lost messages to neighboring vehicles and makes piggybacking decisions based on message losses in its neighbors. We observed that some lost messages still cannot be successfully recovered in greedy piggybacking due to the asymmetric wireless communications, and further proposed a mutual learning based scheme to overcome the drawback of greedy piggybacking. We evaluated the performance of the two schemes through trace-driven simulations, and results show that both schemes can achieve significant improvement on broadcast reliability in VANETs in comparison with the existing solutions.
Guangbing Xiao, Haibo Zhang 0001, Zhiyi Huang 0001, Yawen Chen 0001
VTC Spring4
2016 Routing in Delay Tolerant Networks with fine-grained contact characterisation and dynamic message replication
abstract
Pairwise contacts in Delay-Tolerant Networks (DTNs) for applications such as bus or smartphone based social networking commonly show some regular repeating patterns. Most existing routing protocols only implicitly exploit these patterns to predict future contacts. To enhance message delivery rate, most of the schemes allow messages to be replicated and forwarded to encountered nodes. However, there is no efficient mechanism for dynamically controlling message replication to achieve high message delivery rate with very low message overhead. In this paper, we present FGDR, a routing protocol designed for DTNs by leveraging fine-grained contact characterisation and dynamic message replication. In FGDR, the history contact is characterised in a fine-grained manner using a sliding window mechanism, and an up-to-date future contact prediction can be made based on the most recent history data. We design an efficient message replication scheme, in which replication is controlled in a fully decentralised manner by taking into account the expected message delivery rate, the replication history, and the quality of the encountered node. A replica can be generated only when it is necessary to fulfill the expected message delivery rate. We evaluate our scheme through trace-driven simulations, and results show FGDR can achieve much higher message delivery rate with lower message overhead in comparison with existing schemes.
Luming Wan, Haibo Zhang 0001, Yawen Chen 0001
WoWMoM4
2016 Note on Edge-Colored Graphs for Networks with Homogeneous Faults
abstract
The failure on all homogeneous devices due to the same reason is called homogeneous fault in networks. In contrast, heterogeneous platforms deployed simultaneously in the network are more robust against homogeneous faults. One of the challenging problems is how to design survivable networks that against homogeneous faults. This paper utilizes edge-colored graphs to investigate the network topology with homogeneous faults, in order to guarantee network connectivity using minimum number of links. Two types of network topologies are proposed on the edge-colored graph. One type of networks is characterized by the fact that all the edges of the same color form a Hamiltonian path or a Hamiltonian cycle. An upper bound on the number of colors used in the proposed network topologies is obtained. The network topologies of the second type have edges colored with at most five colors. Additionally, the subnetworks induced by the edges of two colors contain a Hamiltonian path, or a Hamilton cycle in some cases.
Rui Hou 0006, Jigang Wu, Yawen Chen 0001, Haibo Zhang 0001
Comput. J.3
2015 Quantifying the Energy Efficiency Challenges of Achieving Exascale Computing
abstract
Power and performance are two potentially opposing objectives in the design of a supercomputer, where increases in performance often come at the cost of increased power consumption and vice versa. The task of simultaneously maximising both objectives is becoming an increasingly prominent challenge in the development of future exascale supercomputers. To gain some perspective on the scale of the challenge, we analyse the power and performance trends for the Top500 and Green500 supercomputer lists. We then present the PαPW metric, which we use to evaluate the scalability of power efficiency, projecting the development of an exascale system. From this analysis, we found that when both power and performance are considered, the projected date of achieving an exascale system falls far beyond the current target of 2020.
Jason Mair, Zhiyi Huang 0001, David M. Eyers, Yawen Chen 0001
CCGRID4
2015 WRH-ONoC: A wavelength-reused hierarchical architecture for optical Network on Chips
abstract
Optical Network on Chip (ONoC) is a promising technology for the next-generation many-core chip multiprocessors owing to its tremendous advantages in low power consumption, low communication delay, and high bandwidth. In this paper we present WRH-ONoC, a novel wavelength-reused hierarchical architecture that is capable of interconnecting thousands of cores using a limited number of wavelengths while providing extremely high-throughput data communication between connected cores. In WRH-ONoC, the cores are divided into small subsystems that are interconnected using multiple λ-routers and gateways in a hierarchical manner. Each λ-router can provide non-blocking parallel communication among the directly connected cores or gateways, and all λ-routers can reuse the limited number of available wavelengths. Communications between cores in different subsystems are routed via gateways in which optical signals can change their wavelengths via optical-electrical signal conversions. For a given number of cores, we give the minimum number of levels, λ-routers, and gateways required to interconnect these cores, and derive the expected end-to-end data communication delay under the Uniform-Poisson traffic pattern. Both theoretical analysis and simulation results demonstrate that WRH-ONoC can achieve significant improvement on performance and reduction on hardware cost in comparison with the existing solutions.
Haibo Zhang 0001, Yawen Chen 0001, Zhiyi Huang 0001, Huaxi Gu
INFOCOM3
2015 Constructing Edge-Colored Graph for Heterogeneous Networks
Rui Hou 0006, Jigang Wu, Yawen Chen 0001, Haibo Zhang 0001, Xiufeng Sui
J. Comput. Sci. Technol.3
2014 TB-SnW: Trust-based Spray-and-Wait routing for delay-tolerant networks
Aysha Al Hinai, Haibo Zhang 0001, Yawen Chen 0001, Yidong Li
J. Supercomput.3
2013 Analyzing Packet-Level Routing in Data Centers
abstract
Data centers host diverse applications with stringent QoS requirements. The key issue is to eliminate network congestions which severely degrade application performance. One effective solution is to balance the traffic load in the datacenter regular topologies. Many previous strategies focused on optimized flow routing, and these solutions can hardly achieve ideal load balance while guaranteeing QoS of different traffic flows due to the limitations in practical. In this paper, we discuss packet-level routing and analyze its merit for fine-grained load balance in data centers. Though packet-level routing interacts poorly with TCP in traditional network settings, we prove that it can be adapted to datacenter environment. Motived by the work done by Dixit [4] [5], we assert that packet-level routing is the right choice for data centers. Our simulation results demonstrate that packet-level routing better fulfills datacenter requirements.
Ruoyan Liu, Huaxi Gu, Yawen Chen 0001, Haibo Zhang 0001
DASC3
2013 Restricted admission control in view-oriented transactional memory
Kai-Cheung Leung, Yawen Chen 0001, Zhiyi Huang 0001
J. Supercomput.2
2012 WATS: Workload-Aware Task Scheduling in Asymmetric Multi-core Architectures
abstract
Asymmetric Multi-Core (AMC) architectures have shown high performance as well as power efficiency. However, current parallel programming environments do not perform well on AMC due to their assumption that all cores are symmetric and provide equal performance. Their random task scheduling policies, such as task-stealing, can result in unbalanced workloads in AMC and severely degrade the performance of parallel applications. To balance the workloads of parallel applications in AMC, this paper proposes a Workload-Aware Task Scheduling (WATS) scheme that adopts history-based task allocation and preference-based task stealing. The history-based task allocation is based on a near-optimal, static task allocation using the historical statistics collected during the execution of a parallel application. The preference-based task stealing, which steals tasks based on a preference list, can dynamically adjust the workloads in AMC if the task allocation is less optimal due to approximation in the history-based task allocation. Experimental results show that WATS can improve the performance of CPU-bound applications up to 82.7% compared with the random task scheduling policies.
Quan Chen 0002, Yawen Chen 0001, Zhiyi Huang 0001, Minyi Guo
IPDPS2
2012 Mitigating Blackhole Attacks in Delay Tolerant Networks
abstract
Unlike the conventional routing techniques in the Internet where routing privileges are given to trustworthy and fully authenticated nodes, Delay Tolerant Networks (DTNs) allow any node to participate in routing due to the lack of consistent infrastructure and central administration. This creates new security challenges as even authorized nodes in DTNs could inject several malicious threats against the network. This paper investigates novel solutions based on the Spray-and-Wait (SnW) routing protocol for mitigating black hole attacks in DTNs. A new knowledge-based routing scheme, called Trust-Based Spray- and-Wait protocol (TB-SnW), is proposed. The routing decisions in TB-SnW protocol are made based on the trust levels that are computed at each node using its historic routing records. Simulation results show that the TB-SnW protocol can achieve better performance in terms of mitigating Byzantine attacks and reducing message delivery delay compared with the Spray-and-Wait protocol.
Aysha Al Hinai, Haibo Zhang 0001, Yawen Chen 0001
PDCAT3
2011 Routing and wavelength assignment for hypercube communications embedded on optical chordal ring networks of degrees 3 and 4
Yawen Chen 0001, Hong Shen 0001, Haibo Zhang 0001
Comput. Commun.1
2011 Embedding Meshes and Tori on Double-Loop Networks of the Same Size
abstract
Double-loop networks are extensions of ring networks and are widely used in the design and implementation of local area networks and parallel processing architectures. However, embedding of other types of networks on double-loop networks has not been well studied due to the topological complexity of double-loop networks. The traditional L-shape [CHECK END OF SENTENCE], designed to compute the diameter of double-loop networks, is not effective to solve the embedding problem. We propose a novel tessellation approach to partition the geometric plane of double-loop networks into a set of parallelogram tiles, called P-shape. Based on the characteristics of P-shape, we design a simple embedding scheme, namely, P-shape embedding, that embeds meshes and tori on double-loop networks in a systematic way. Under P-shape embedding, we evaluate the embedding metrics of dilation, average dilation, and congestion, which depend heavily on the parameters of P-shape. A main merit of P-shape embedding is that a large fraction of embedded mesh/torus edges have edge dilation 1, resulting in a low average dilation. These are the first results, to our knowledge, for embedding meshes and tori on double-loop networks which is of great significance due to the popularity of these architectures. Our P-shape construction bridges between regular graphs and double-loop networks, and provides a powerful tool for studying double-loop networks.
Yawen Chen 0001, Hong Shen 0001
IEEE Trans. Computers1
2010 Routing and wavelength assignment for hypercube in array-based WDM optical networks
Yawen Chen 0001, Hong Shen 0001
J. Parallel Distributed Comput.1
2008 Balancing energy consumption for uniform data gathering wireless sensor networks
abstract
No abstract available.
Haibo Zhang 0001, Hong Shen 0001, Yawen Chen 0001, Zonghua Zhang
PODC3
2007 Wavelength Assignment for Directional Hypercube Communications on a Class of WDM Optical Networks
abstract
Hypercube communication is one of the most versatile and efficient communication patterns shared by a large number of computational problems. In this paper, we study routing and wavelength assignment for realizing hypercube communications on WDM optical networks including linear arrays and rings with the consideration of communication directions. Specifically, we consider this problem for both bidirectional and unidirectional hypercube communications. For each case, we identify a lower bound on the number of wavelengths required, and present a simple embedding scheme and wavelength assignment algorithm that uses a provably near-optimal number of wavelengths. By realizing hypercube computations in optical networks, the hypercube computation speed can be significantly improved compared with the traditional electronic networks.
Yawen Chen 0001, Hong Shen 0001
ICPP1
2006 Embedding Hypercube Communications on Optical Chordal Ring Networks
abstract
Hypercube communication is one of the most versatile and efficient communication patterns for parallel computation. Routing and wavelength assignments for realizing hypercube communications on WDM linear arrays, rings, meshes and tori have been discussed in our past researches. In this paper, we study routing and wavelength assignment for realizing hypercube communications on WDM chordal ring networks of degree 3. We design embedding scheme and derive the number of wavelengths required for different chord length. Based on embedding scheme of double cycle embedding, we also provide the analysis of chord length with optimal number of wavelengths to realize hypercube communications on 3-degree chordal rings. Results show that the wavelength requirement for realizing hypercube communications on optical networks has been further reduced on optical 3-degree chordal ring networks compared with some topologies discussed before. Our results have both theoretical and practical significance as WDM optical networks have an increasing popularity
Yawen Chen 0001, Hong Shen 0001, Haibo Zhang 0001
LCN1
2006 Wavelength Assignment for Realizing Parallel FFT on Regular Optical Networks
Yawen Chen 0001, Hong Shen 0001, Fang'ai Liu
J. Supercomput.1
2005 An Improved Scheme of Wavelength Assignment for Parallel FFT Communication Pattern on a Class of Regular Optical Networks
Yawen Chen 0001, Hong Shen 0001
NPC1
2005 Wavelength Assignment for Parallel FFT Communication Pattern on Linear Arrays by Lattice Embedding
abstract
Fast Fourier Transform(FFT) represents a common communication pattern shared by a large class of scientific and engineering problems and wavelength assignment is a key issue to increase efficiency and reduce cost in Wavelength Division Multiplexing (WDM) optical networks. In this paper, we propose a new scheme for the wavelength assignment of parallel FFT communication pattern on WDM linear arrays. By lattice embedding, the number of wavelengths required to realize parallel FFT communication pattern on WDM linear arrays significantly improves the known result. Our proposed embedding method also provides a new approach to the hypercube layout problem considering connections dimension by dimension rather than all connections as in the traditional approach.
Yawen Chen 0001, Hong Shen 0001
PDCAT1