Jiansong Zhang 0001

dblp:38/6831-1 · DBLP profile ↗
← Back
39ranked-venue papers
7as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 30 · 6 first-author · 1 since 2021Systems, architecture and hardware · 8 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Fast and Scalable Selective Retransmission for RDMA
Peihao Huang, Guo Chen 0001, Xin Zhang 0117, Huijun Shen, Ying Bian, Yuanwei Lu, Zhenyuan Ruan, Bojie Li, Jiansong Zhang 0001, Yongfeng Liu, Zhigang Chen 0001
INFOCOM11
2025 SAFE: A Scalable Homomorphic Encryption Accelerator for Vertical Federated Learning
abstract
Privacy preservation has become a critical concern for governments, hospitals, and large corporations. Homomorphic encryption (HE) enables a ciphertext-based computation paradigm with strong security guarantees. In emerging cross-agency data cooperation scenarios like vertical federated learning (VFL), HE protects the data interaction from exposure to counterparts. However, computation on ciphertext has significant performance challenges due to increased data size and substantial overhead. Related work has been proposed to accelerate HE using parallel hardware, such as GPUs, FPGAs, and ASICs. However, many existing hardware accelerators target specific HE operations, such as number theoretic transform (NTT) and key switching, providing limited performance improvement for end-to-end applications. Others support bootstrapping, which requires quite a large ASIC design. To better support existing VFL training applications, we propose SAFE, an HE accelerator for scalable homomorphic matrix-vector products (HMVPs), which is the performance bottleneck. SAFE adopts a coefficient-wise encoded HMVP algorithm, despite a vanilla mode, we further explore the compressed and concatenated modes, which can fully utilize the polynomial encoding slots. The proposed hardware architecture, customized for HMVP dataflow, supports spatial and temporal parallelization of function units. The most costly polynomial function, NTT, is implemented with a low-area constant geometry unit which improves efficiency by$2.43\times $. SAFE is implemented as a CPU-FPGA heterogeneous acceleration system, unleashing the multithread potential. The evaluation demonstrates an up to$36\times $speed-up in end-to-end federated logistic regression training.
Yanheng Lu, Xuanle Ren, Ruiguang Zhong, Jiansong Zhang 0001, Hanghang Wu, Xiaofu Zheng, Tingqiang Chu, Cheng Hong 0001, Changzheng Wei, Dimin Niu, Yuan Xie 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2023 CHAM: A Customized Homomorphic Encryption Accelerator for Fast Matrix-Vector Product
abstract
Homomorphic encryption (HE) is a promising technique for privacy-preserving computing because it allows computation on encrypted data without decryption. HE, however, suffers from poor performance due to enlarged data size and exploded amount of computation. Related work has been proposed to accelerate HE using GPUs, FPGAs, and ASICs. The existing work, however, aims at specific HE schemes and fails to consider the fast-evolving algorithms. For example, HE algorithms that combine different HE schemes have demonstrated capability of supporting more types of HE operations and ciphertexts. Moreover, some existing hardware accelerators target small HE operations (such as number theoretic transform and key-switch), which however provides limited or even neglected performance improvement for end-to-end applications. To better support existing privacy-preserving applications (e.g., logistic regression and neural network inference), we propose CHAM, an HE accelerator, for high-performance matrix-vector product, which can be easily extended to 2-D and 3-D convolutions. Motivated by the evolution of algorithms, CHAM supports not only traditional HE operations, but also different types of ciphertexts and the conversion between them. We implement CHAM with Xilinx FPGAs. The evaluation demonstrates 1800× speed-up for matrix-vector product, 36× speed-up for logistic regression, and 144× speed-up for Beaver triple generation compared to the existing work.
Xuanle Ren, Yanheng Lu, Ruiguang Zhong, Jiansong Zhang 0001, Hanghang Wu, Xiaofu Zheng, Tingqiang Chu, Cheng Hong 0001, Changzheng Wei, Dimin Niu, Yuan Xie 0001
DAC7
2022 An Efficient Hardware Design for Accelerating Sparse CNNs With NAS-Based Models
abstract
Deep convolutional neural networks (CNNs) have achieved remarkable performance at the cost of huge computation. As the CNN models become more complex and deeper, compressing CNNs to sparse by pruning the redundant connection in the networks has emerged as an attractive approach to reduce the amount of computation and memory requirement. On the other hand, FPGAs have been demonstrated to be an effective hardware platform to accelerate CNN inference. However, most existing FPGA accelerators focus on dense CNN models, which are inefficient when executing sparse models as most of the arithmetic operations involve addition and multiplication with zero operands. In this work, we propose an accelerator with software–hardware co-design for sparse CNNs on FPGAs. To efficiently deal with the irregular connections in the sparse convolutional layers, we propose a weight-oriented dataflow that exploits element–matrix multiplication as the key operation. Each weight is processed individually, which yields low decoding overhead. Then, we design an FPGA accelerator that features a tile look-up table (TLUT) and a channel multiplexer (CMUX). The TLUT is designed to match the index between sparse weights and input pixels. Using TLUT, the runtime decoding overhead is mitigated by using an efficient indexing operation. Moreover, we propose a weight layout to enable efficient on-chip memory access without conflicts. To cooperate with the weight layout, a CMUX is inserted to locate the address. Finally, we build a neural architecture search (NAS) engine that leverages the reconfigurability of FPGAs to generate an efficient CNN model and choose the optimal hardware design parameters. The experiments demonstrate that our accelerator can achieve 223.4-309.0 GOP/s for the modern CNNs on Xilinx ZCU102, which provides a$2.4\times $–$12.9\times $speedup over previous dense CNN accelerators on FPGAs. Our FPGA-aware NAS approach shows$2\times $speedup over MobileNetV2 with 1.5% accuracy loss.
Yun Liang 0001, Liqiang Lu, Yicheng Jin, Jiaming Xie, Ruirui Huang, Jiansong Zhang 0001, Wei Lin 0016
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2021 FleetRec: Large-Scale Recommendation Inference on Hybrid GPU-FPGA Clusters
abstract
We present FleetRec, a high-performance and scalable recommendation inference system within tight latency constraints. FleetRec takes advantage of heterogeneous hardware including GPUs and the latest FPGAs equipped with high-bandwidth memory. By disaggregating computation and memory to different types of hardware and bridging their connections by high-speed network, FleetRec gains the best of both worlds, and can naturally scale out by adding nodes to the cluster. Experiments on three production models up to 114 GB show that FleetRec outperforms optimized CPU baseline by more than one order of magnitude in terms of throughput while achieving significantly lower latency.
Wenqi Jiang 0001, Zhenhao He, Shuai Zhang 0007, Kai Zeng 0002, Jiansong Zhang 0001, Tongxuan Liu, Yong Li 0020, Jingren Zhou 0001, Ce Zhang 0001, Gustavo Alonso
KDD6
2020 CEFS: compute-efficient flow scheduling for iterative synchronous applications
abstract
Iterative Synchronous Applications (ISApps) are popular in today's data centers, represented by distributed deep learning (DL) training. In ISApps, multiple nodes carry out the computing task iteratively, with globally synchronizing the results in each iteration. To increase the scaling efficiency of ISApps, in this paper we propose a new flow scheduling approach, called CEFS. CEFS saves the waiting time of computing nodes from two aspects. For a single node, flows with data which can trigger earlier computation at the node are assigned with higher priority; among nodes, flows towards slower nodes are assigned with higher priority.
Shuai Wang 0028, Dan Li 0001, Jiansong Zhang 0001, Wei Lin 0016
CoNEXT3
2019 PAI-FCNN: FPGA Based Inference System for Complex CNN Models
abstract
Convolutional Neural Network (CNN) models are becoming complex with advanced OPs and structures, which introduces design challenges for FPGA-based system. In this paper, we present the design of an FPGA-based CNN inference system, PAI-FCNN, to support modern complex CNN models. PAI-FCNN consists of scalable hardware design and a model reconstruction flow in software compiler. In this way, advanced OPs like Deconv, Conv with upsampling, Dilated Conv, Concatenation can be processed by PAI-FCNN with high performance and hardware efficiency. PAI-FCNN also incorporates reduced precision to boost computing capacity, and the emerging CNN-RNN (Recurrent Neural Network) hybrid models are supported. Our experiments on both PC and embedded FPGA platforms show that the system consistently performs in an efficient manner. PAI-FCNN achieves better throughput and power efficiency than GPU solutions.
Lixue Xia, Lansong Diao, Zhao Jiang, Hao Liang 0003, Kai Chen 0008, Shunli Dou, Zibin Su, Jiansong Zhang 0001, Wei Lin 0016
ASAP10
2019 An Efficient Hardware Accelerator for Sparse Convolutional Neural Networks on FPGAs
abstract
Deep convolutional neural networks (CNN) have achieved remarkable performance with the cost of huge computation. As the CNN model becomes more complex and deeper, compressing CNN to sparse by pruning the redundant connection in networks has emerged as an attractive approach to reduce the amount of computation and memory requirement. In recent years, FPGAs have been demonstrated to be an effective hardware platform to accelerate CNN inference. However, most existing FPGA architectures focus on dense CNN models. The architecture designed for dense CNN models are inefficient when executing sparse models as most of the arithmetic operations involve addition and multiplication with zero operands. On the other hand, recent sparse FPGA accelerators only focus on FC layers. In this work, we aim to develop an FPGA accelerator for sparse CNNs. To efficiently deal with the irregular connection in the sparse convolutional layer, we propose a weight-oriented dataflow that processes each weight individually. Then we design an FPGA architecture which can handle input-weight connection and weight-output connection efficiently. For input-weight connection, we design a tile look-up table to eliminate the runtime indexing match of compressed weights. Moreover, we develop a weight layout to enable high on-chip memory access. To cooperate with the weight layout, a channel multiplexer is inserted to locate the address which can ensure no data access conflict. Experiments demonstrate that our accelerator can achieve 223.4-309.0 GOP/s for the modern CNNs on Xilinx ZCU102, which provides a 3.6x-12.9x speedup over previous dense CNN FPGA accelerators.
Liqiang Lu, Jiaming Xie, Ruirui Huang, Jiansong Zhang 0001, Wei Lin 0016, Yun Liang 0001
FCCM4
2019 PAI-FCNN: FPGA Based CNN Inference System
abstract
We describe the FPGA subsystem of the Platform of Artificial Intelligence (PAI) in Alibaba Group, called PAI-FCNN. PAI-FCNN plays the role of a heterogeneous back-end for CNN inference, together with other CPU, GPU and ASIC subsystems in PAI. Driven by various business needs, we built PAI-FCNN from scratch since two years ago. We present our experience from FPGA/compiler design and implementation, to system evaluation and deployment. In particular, in order to address three practical challenges: (1) Efficient processing for diverse operators and model structure such as Deconv, Dilated Conv, Up-sampling, PReLu and Concatenation. (2) Serving multiple highly-different models on single FPGA hardware. (3) Competitive performance with alternative GPU or ASIC solutions, we extensively perform joint software & hardware design to optimize system efficiency across multiple CNN models, which includes model reconstruction in compiler software and flexible data access in data-flow CNN processor. We also incorporate reduced precision and model retraining to boost system capacity. Using U-net as an example, on Xilinx KU115 chip, with the help of 74.9% efficiency on Int16-precision hardware (with 3.226TOPS capacity) and 72.9% efficiency on mixed-int8/int3-precision hardware (with 14.746TOPS capacity), we achieve slightly better throughput and 2X higher power efficiency than P4.
Lansong Diao, Zhao Jiang, Hao Liang 0003, Chang'an Ye, Kai Chen 0008, Shunli Dou, Lixue Xia, Jiansong Zhang 0001, Wei Lin 0016
FPGA10
2019 Speedy: An Accelerator for Sparse Convolutional Neural Networks on FPGAs
abstract
Deep convolutional neural networks (CNNs) have achieved remarkable performance with the cost of huge computation. Moreover, the current trend of CNNs is towards more complex and deeper topology. Compressing CNNs to sparse have emerged as the most attractive approach to reduce the amount of computation and memory requirement. This compression is achieved by pruning the redundant connection in networks. FPGAs have been an effective solution to accelerate CNN inference for its high parallel computing, flexibility and energy-efficiency. Although existing FPGA architectures are able to excellently process dense CNN models, they cannot benefit from the computation reduction when accelerating the sparse CNN models. Because most of the arithmetic operations involve addition and multiplication with zero operands, meanwhile accelerating sparse CNN models incurs significant data encoding and decoding overhead. In this paper, we propose a FPGA accelerator Speedy that can efficiently exploit sparsity in CNN models. We first investigate the dataflow design space to explore the available performance with different parallelization strategies. The result of exploration is Speedy dataflow which provides enough parallel multiplications and maximizes the weight reuse. Then, we propose a novel data representation combined with memory partition technique to increase the on-chip bandwidth. Finally, we propose Speedy FPGA architecture in which we apply line buffer design and high-throughput PE. In the experiments, we evaluate Speedy on contemporary neural networks. Speedy provides flexible parameters for different FPGA scale. First, we evaluate the resources utilization and hardware efficiency with different design configurations. Then we compare our design with previous FPGA implementations. Overall, Speedy achieves 11.3x-20.8x and 1.5x-6.8x speed up for Alexnet and VGGnet with 90% weight sparsity.
Liqiang Lu, Yun Liang 0001, Ruirui Huang, Wei Lin 0016, Xiaoyuan Cui, Jiansong Zhang 0001
FPGA6
2019 Ouroboros: An Inference Engine for Deep Learning Based TTS on Embedded Devices
abstract
This article consists of a collection of slides from the author's conference presentation.
Jiansong Zhang 0001, Lixue Xia, Zhao Jiang, Hao Liang 0003, Shouda Liu, Wei Lin 0016, Yuan Xie 0001
Hot Chips Symposium1
2019 BeamRaster: A Practical Fast Massive MU-MIMO System With Pre-Computed Precoders
abstract
In order to achieve more dramatic spatial multiplexing gains, both industry and academia have pushed towards the massive Multi-User Multi-Input and Multi-Output (MU-MIMO) systems. However, traditional linear precoding techniques do not scale up well with the number of antennas, i.e., they either have high implementation difficulties (zero-forcing) or sacrifice wireless capacity as a price (conjugate or codebook-based precoding). In this paper, we present a novel precoding scheme, BeamRaster, which is a fast and high efficient scheme for massive MU-MIMO system. Inspired from the codebook-based precoding, BeamRaster pre-computes a set of angle-domain beam filters that divide the channel into directional subspaces. Unlike previous work, BeamRaster carefully manages the cross-interference using (1) a grating table to track the correlation among beams in real-time, (2) an interference-aware user-beam selection, and (3) a pre-distortion method to cancel the residual interference because of side-lobes. We implement and evaluate the BeamRaster using FPGA and software defined radio platform. On one hand, BeamRaster is easy to implement in hardware, i.e., it can realize the precoding for a 64-antenna MU-MIMO system in real time with a single Altera Stratix V FPGA. On the other hand, both the experiments with medium-scale antennas and simulations with large-scale antennas show that BeamRaster can achieve high capacity gain.
Wencong Xiao, Yuechen Tao, Jiansong Zhang 0001, Wenjie Wang 0001
IEEE Trans. Mob. Comput.6
2019 Polarization-Based Visible Light Positioning
abstract
Visible Light Positioning (VLP) provides a promising means to achieve indoor localization with sub-meter accuracy. We observe that the Visible Light Communication (VLC) methods in existing VLP systems rely on intensity-based modulation, and thus they require a high pulse rate to prevent flickering. However, the high pulse rate adds an unnecessary and heavy burden for receiving with cameras. To eliminate this burden, we propose the polarization-based modulation, which is flicker-free, to enable a low pulse rate VLC. In this way, we make VLP light-weight enough even for devices with a camera. This paper presents the VLP system PIXEL, which realizes our idea. In PIXEL, we developed 1) a novel color-based modulation scheme to handle user's mobility and 2) a sensor-assisted positioning algorithm to ease user's burden in capturing multiple location anchors. Our experiments based on the prototype show that PIXEL can provide 10 cm-level real-time VLP for wearables and smartphones with camera resolution as coarse as 60 pixel × 80 pixel and CPU frequency as low as 300 MHz.
Zhice Yang, Zeyu Wang 0001, Jiansong Zhang 0001, Qian Zhang 0001
IEEE Trans. Mob. Comput.3
2019 MP-RDMA: Enabling RDMA With Multi-Path Transport in Datacenters
abstract
RDMA is becoming prevalent because of its low latency, high throughput and low CPU overhead. However, in current datacenters, RDMA remains a single path transport which is prone to failures and falls short to utilize the rich parallel network paths. Unlike previous multi-path approaches, which mainly focus on TCP, this paper presents a multi-path transport for RDMA, i.e. MP-RDMA, which efficiently utilizes the rich network paths in datacenters. MP-RDMA employs three novel techniques to address the challenge of limited RDMA NICs on-chip memory size: 1) a multi-path ACK-clocking mechanism to distribute traffic in a congestion-aware manner without incurring per-path states; 2) an out-of-order aware path selection mechanism to control the level of out-of-order delivered packets, thus minimizes the meta data required to them; 3) a synchronise mechanism to ensure in-order memory update whenever needed. With all these techniques, MP-RDMA only adds 66B to each connection state compared to single-path RDMA. Our evaluation with an FPGA-based prototype demonstrates that compared with single-path RDMA, MP-RDMA can significantly improve the robustness under failures ( $2\times \sim 4\times $ higher throughput under 0.5%~10% link loss ratio) and improve the overall network utilization by up to 47%.
Guo Chen 0001, Yuanwei Lu, Bojie Li, Kun Tan 0002, Yongqiang Xiong, Peng Cheng 0005, Jiansong Zhang 0001, Thomas Moscibroda
IEEE/ACM Trans. Netw.7
2018 Cutting the Cord: Designing a High-quality Untethered VR System with Low Latency Remote Rendering
abstract
This paper introduces an end-to-end untethered VR system design and open platform that can meet virtual reality latency and quality requirements at 4K resolution over a wireless link. High-quality VR systems generate graphics data at a data rate much higher than those supported by existing wireless-communication products such as Wi-Fi and 60GHz wireless communication. The necessary image encoding, makes it challenging to maintain the stringent VR latency requirements. To achieve the required latency, our system employs a Parallel Rendering and Streaming mechanism to reduce the add-on streaming latency, by pipelining the rendering, encoding, transmission and decoding procedures. Furthermore, we introduce a Remote VSync Driven Rendering technique to minimize display latency. To evaluate the system, we implement an end-to-end remote rendering platform on commodity hardware over a 60Ghz wireless network. Results show that the system can support current 2160x1200 VR resolution at 90Hz with less than 16ms end-to-end latency, and 4K resolution with 20ms latency, while keeping a visually lossless image quality to the user.
Ruiguang Zhong, Wuyang Zhang, Yunxin Liu 0001, Jiansong Zhang 0001, Marco Gruteser
MobiSys5
2018 Multi-Path Transport for RDMA in Datacenters
Yuanwei Lu, Guo Chen 0001, Bojie Li, Kun Tan 0002, Yongqiang Xiong, Peng Cheng 0005, Jiansong Zhang 0001, Enhong Chen, Thomas Moscibroda
NSDI7
2018 DCAP: Improving the Capacity of WiFi Networks with Distributed Cooperative Access Points
abstract
This paper presents the Distributed Cooperative Access Points (DCAP) system that can simultaneously serve multiple clients using cooperative beamforming to increase the capacity of WiFi-type wireless networks. The distributed APs are connected by Ethernet and driven by independent low-cost local oscillators. To facilitate cooperative beamforming, we address three major challenges: the phase synchronization, the channel state information (CSI) measurement, and the user selection. Specifically, we develop 1) a cooperative tracking scheme to track signal phase drifts at symbol level without adding extra hardware complexity; 2) an incremental CSI estimation mechanism that removes the per-frame CSI measurement overhead of previous approaches; and 3) a simple random user selection algorithm that scales the network capacity linearly and delivers over 70 percent performance compared to the optimal but complex greedy algorithm. We implement DCAP on the Sora software radio platform and evaluate it in a wireless network with nine nodes. Experimental results show that the cooperative beamforming is feasible in practice, and our cooperative phase tracking can ensure strict phase alignment (≤ 0.03 radian) among APs during the entire beamforming period (1.2 ms). Otherwise, without tracking, phases may drift by 0.3 radian over merely 600 μs, causing that the symbol SNR decreases as large as 20 dB.
Taotao Wang, Qing Yang 0006, Jiansong Zhang 0001, Soung Chang Liew, Shengli Zhang 0001
IEEE Trans. Mob. Comput.4
2017 Memory Efficient Loss Recovery for Hardware-based Transport in Datacenter
abstract
Limited by the small on-chip memory, hardware-based transport typically implements go-back-N loss recovery mechanism, which costs very few memory but is well-known to perform inferior even under small packet loss ratio. We present MELO, an efficient selective retransmission mechanism for hardware-based transport, which consumes only a constant small memory regardless of the number of concurrent connections. Specifically, MELO employs an architectural separation between data and meta data storage and uses a shared bits pool allocation mechanism to reduce meta data on-chip memory footprint. By only adding in average 23B extra on-chip states for each connection, MELO achieves up to 14.02x throughput while reduces 99% tail FCT by 3.11x compared with go-back-N under certain loss ratio.
Yuanwei Lu, Guo Chen 0001, Zhenyuan Ruan, Wencong Xiao, Bojie Li, Jiansong Zhang 0001, Yongqiang Xiong, Peng Cheng 0005, Enhong Chen
APNet6
2017 BikeLoc: a Real-time High-Precision Bicycle Localization System Using Synthetic Aperture Radar
abstract
In recent years we have witnessed the rapid development of smart bicycles. For example, Mobike1 is able to interact with smartphones. As we all known, accurate bicycle localization system is one of the most critical technologies for the development of smart bicycles. However, GPS's error is at meter-level and it performs poorly under skyscrapers and in tunnels.
Hongjiang Lyu, Linghe Kong, Chengzhang Li, Yunxin Liu 0001, Jiansong Zhang 0001, Guihai Chen
APNet5
2017 Proximity based IoT device authentication
abstract
Internet of Things (IoT) devices are largely embedded devices which lack a sophisticated user interface, e.g., touch screen, keyboard, etc. As a consequence, traditional Pre-Shared Key (PSK) based authentication for mobile devices becomes difficult to apply. For example, according to our study on home automation devices which leverage smartphone for PSK input, the current process does not protect against active impersonating attack and also leaks the Wi-Fi password to eavesdroppers, i.e., currently these IoT devices can be exploited to enter into critical infrastructures, e.g., home networks. Motivated by this real-world security vulnerability, in this paper we propose a novel proximity-based mechanism for IoT device authentication, called Move2Auth, for the purpose of enhancing IoT device security. In Move2Auth, we require user to hold smartphone and perform one of two hand-gestures (moving towards and away, and rotating) in front of IoT device. By combining (1) large RSS-variation and (2) matching between RSS-trace and smartphone sensor-trace, Move2Auth can reliably detect proximity and authenticate IoT device accordingly. Based on our implementation on Samsung Galaxy smartphone and commodity Wi-Fi adapter, we prove Move2Auth can protect against powerful active attack, i.e., the false-positive rate is consistently lower than 0.5%.
Jiansong Zhang 0001, Zeyu Wang 0001, Zhice Yang, Qian Zhang 0001
INFOCOM1
2017 Latency-based WiFi congestion control in the air for dense WiFi networks
abstract
WiFi has become the primary method to access the Internet. However, the WiFi-hop latency, particularly in dense-WiFi environments, is far from satisfactory [1], to support delay-sensitive applications such as Web browsing and VoIP. The WiFi latency mainly comes from two kinds of queues: the host queue and the distributed queue, which is caused by CSMA/CA mechanism when multiple nodes contend for the channel. While the host queue can be easily bypassed using priority scheduling at end-host, the distributed queue is not. Previously, IEEE 802.11e tries to provide priorities in this distributed queue by adjusting the MAC layer parameters, but it does not scale when there are increasing number of delay-sensitive flows. In this paper, we propose and design QAir, a practical solution to reduce WiFi latency of delay-sensitive flows in dense WiFi networks. QAir takes a different approach to transfer this distributed queue to host queue. Consequently, the delay-sensitive flows can bypass the entire queue and their latency can be greatly reduced. QAir works in a distributed manner with no centralized scheduler. We have implemented QAir on commodity WiFi devices. Experimental results show that, compared to the 802.11 DCF baseline, QAir can reduce the average WiFi-hop latency of delay-sensitive flows by 50-75%.
Changhua Pei, Youjian Zhao, Yunxin Liu 0001, Kun Tan 0002, Jiansong Zhang 0001, Yuan Meng 0002, Dan Pei
IWQoS5
2017 Lightweight Display-to-device Communication using Electromagnetic Radiation and FM Radio
abstract
This paper presents Shadow, a novel display-to-device communication system working in radio frequency. It leverages the Electromagnetic Radiation (EMR) emanated from display to transmit information. We modulate the signal transmitted to display to make the corresponding EMR signal fall into the FM band. In this way. when users' devices with FM ability are approaching to the display, it can automatically receive information from the display. Since Shadow does not rely on camera, wearable devices can also obtain information from display. In addition to the lightweight communication ability, Shadow keeps the communication truly invisible by only transmitting in the period that will not be shown on the display panel. Our system requires no modification on hardware, and we implement it with commodity display system and mobile devices. Results show it can achieve 1.5 kbps at distances of up to 20cm from the display panel.
Zhice Yang, Jiansong Zhang 0001, Zeyu Wang 0001, Qian Zhang 0001
MobiHoc2
2015 SpaceHub: A Smart Relay System for Smart Home
abstract
With the proliferation of smart wireless devices in our homes, the cross-technology interference increasingly becomes an important issue. This paper presents a novel smart relay system, called SpaceHub, which leverages an multi-antenna relay node to mitigate cross-technology interference for all communicating devices which may only have single antenna. In SpaceHub, the relay node overhears wireless communications in the air, separates the collided signals, and forwards the separated (cleaned) signals to their intended receivers without a prior knowledge of the wireless signal structures. The core component of SpaceHub is a blind signal separator that constructs spatial filters using the angle-of-arrival information of collided signals. We have implemented SpaceHub on a software radio platform and our evaluation shows SpaceHub signal separator can suppress the interference up to 23dB, and is robust against the power or relative locations of interfering signals.
Lizhao You, Jiansong Zhang 0001, Wenjie Wang 0001
HotNets4
2015 Enabling TDMA for today's wireless LANs
abstract
Today's WLANs are struggling to provide desirable features like high efficiency, fairness and QoS because of the use of Distributed Coordination Function (DCF). In this paper we present OpenTDMF, an architecture to enable TDMA on commodity WLAN devices. Our hope is to provide the desirable features without entirely rebuilding the WLAN infrastructure. OpenTDMF is inspired by and architecturally similar to Software Defined Networking (SDN). Specifically, we leverage the backhaul of WLAN to coordinate all the stations for channel access. This fine-grained coordination is performed in a decoupled control plane which includes a central controller and programmable APs. To realize OpenTDMF on commodity WLAN devices, we develop several novel techniques to achieve μs-level time synchronization among all the APs. We also enable AP-triggered uplink transmission so that all the transmissions in the WLAN can be determined. We implemented a prototype of OpenTDMF based on commodity WLAN devices. Empirical results validate the OpenTDMF design and demonstrate its benefits.
Zhice Yang, Jiansong Zhang 0001, Kun Tan 0001, Qian Zhang 0001, Yongguang Zhang
INFOCOM2
2015 Turning Waste into Wealth: Enabling Communication in Guardband Whitespace
abstract
Similar to TV bands, the guardband frequencies are not occupied therefore are whitespace that potentially allows additional communication activities. Considering the difference to TV whitespace, we propose independent communication for guardband whitespace. In this paper, we present the Pilotfish system which realizes independent communication and turns guardband whitespace into new communication channels. To address the big challenges of interference mitigation, we employ novel PHY design which includes specially customized FBMC and an Nulled Decoding technique to null the strong background signal in guardbands. We implemented Pilotfish using software radio system. Empirical evaluation results validate the Pilotfish design in both PHY and MAC.
Jiansong Zhang 0001, Jin Zhang 0001, Kun Tan 0001, Lin Yang 0009, Qian Zhang 0001, Yongguang Zhang
MobiHoc1
2015 Video: Lightweight Visible Light Communication for Indoor Positioning
abstract
Visible light positioning (VLP) is an emerging positioning technique that utilizes indoor light sources to broadcast archer locations through visible light communication (VLC). Benefited by the densely deployed light lamps, VLP holds the promise for more accurate positioning accuracy than RF based approaches, and thus enables interesting applications such as retail navigation and shelf-level advertising in supermarkets and shopping malls.
Zeyu Wang 0001, Zhice Yang, Jiansong Zhang 0001, Qian Zhang 0001
MobiSys3
2015 Demo: Lightweight Visible Light Communication forIndoor Positioning
abstract
Visible light positioning (VLP) is an emerging positioning technique that utilizes indoor light sources to broadcast archer locations through visible light communication (VLC). Benefited by the densely deployed light lamps, VLP holds the promise for more accurate positioning accuracy than RF based approaches, and thus enables interesting applications such as retail navigation and shelf-level advertising in supermarkets and shopping malls.
Zeyu Wang 0001, Zhice Yang, Jiansong Zhang 0001, Qian Zhang 0001
MobiSys3
2015 Wearables Can Afford: Light-weight Indoor Positioning with Visible Light
abstract
Visible Light Positioning (VLP) provides a promising means to achieve indoor localization with sub-meter accuracy. We observe that the Visible Light Communication (VLC) methods in existing VLP systems rely on intensity-based modulation, and thus they require a high pulse rate to prevent flickering. However, the high pulse rate adds an unnecessary and heavy burden to receiving devices. To eliminate this burden, we propose the polarization-based modulation, which is flicker-free, to enable a low pulse rate VLC. In this way, we make VLP light-weight enough even for resource-constrained wearable devices, e.g. smart glasses. Moreover, the polarization-based VLC can be applied to any illuminating light sources, thereby eliminating the dependency on LED.
Zhice Yang, Zeyu Wang 0001, Jiansong Zhang 0001, Qian Zhang 0001
MobiSys3
2013 Flexible array of inexpensive radios
abstract
In this demo, we propose to use multiple inexpensive off-the-shelf radios to build the FAIR system that can be flexibly configured to realize (1) non-contiguous spectrum access (2) MIMO and beamforming (3) constructing wider-band radio. While non-contiguous spectrum access and MIMO/beamforming are naturally supported by the FAIR system, we further develop radio bonding technique for constructing wider-band radio. Radio bonding provides a cost effective alternative to proprietory radio development that it can realize a non-existing wider-band radio using multiple commodity narrower-band radios. We demonstrate FAIR and radio bonding based on Sora 2.0.
Jiansong Zhang 0001
MobiCom1
2013 BigStation: enabling scalable real-time signal processingin large mu-mimo systems
abstract
Multi-user multiple-input multiple-output (MU-MIMO) is the latest communication technology that promises to linearly increase the wireless capacity by deploying more antennas on access points (APs). However, the large number of MIMO antennas will generate a huge amount of digital signal samples in real time. This imposes a grand challenge on the AP design by multiplying the computation and the I/O requirements to process the digital samples. This paper presents BigStation, a scalable architecture that enables realtime signal processing in large-scale MIMO systems which may have tens or hundreds of antennas. Our strategy to scale is to extensively parallelize the MU-MIMO processing on many simple and low-cost commodity computing devices. Our design can incrementally support more antennas by proportionally adding more computing devices. To reduce the overall processing latency, which is a critical constraint for wireless communication, we parallelize the MU-MIMO processing with a distributed pipeline based on its computation and communication patterns. At each stage of the pipeline, we further use data partitioning and computation partitioning to increase the processing speed. As a proof of concept, we have built a BigStation prototype based on commodity PC servers and standard Ethernet switches. Our prototype employs 15 PC servers and can support real-time processing of 12 software radio antennas. Our results show that the BigStation architecture is able to scale to tens to hundreds of antennas. With 12 antennas, our BigStation prototype can increase wireless capacity by 6.8x with a low mean processing delay of 860μs. While this latency is not yet low enough for the 802.11 MAC, it already satisfies the real-time requirements of many existing wireless standards, e.g., LTE and WCDMA.
Qing Yang 0006, Hongyi Yao, Ji Fang, Jiansong Zhang 0001, Yongguang Zhang
SIGCOMM7
2013 Fine-Grained Channel Access in Wireless LAN
abstract
With the increasing of physical-layer (PHY) data rate in modern wireless local area networks (WLANs) (e.g., 802.11n), the overhead of media access control (MAC) progressively degrades data throughput efficiency. This trend reflects a fundamental aspect of the current MAC protocol, which allocates the channel as a single resource at a time. This paper argues that, in a high data rate WLAN, the channel should be divided into separate subchannels whose width is commensurate with the PHY data rate and typical frame size. Multiple stations can then contend for and use subchannels simultaneously according to their traffic demands, thereby increasing overall efficiency. We introduce FICA, a fine-grained channel access method that embodies this approach to media access using two novel techniques. First, it proposes a new PHY architecture based on orthogonal frequency division multiplexing (OFDM) that retains orthogonality among subchannels while relying solely on the coordination mechanisms in existing WLAN, carrier sensing and broadcasting. Second, FICA employs a frequency-domain contention method that uses physical-layer Request to Send/Clear to Send (RTS/CTS) signaling and frequency domain backoff to efficiently coordinate subchannel access. We have implemented FICA, both MAC and PHY layers, using a software radio platform, and our experiments demonstrate the feasibility of the FICA design. Furthermore, our simulation results show FICA can improve the efficiency of WLANs from a few percent to 600% compared to existing 802.11.
Ji Fang, Yuanyang Zhang, Shouyuan Chen, Lixin Shi, Jiansong Zhang 0001, Yongguang Zhang, Zhenhui Tan
IEEE/ACM Trans. Netw.6
2012 Frame retransmissions considered harmful: improving spectrum efficiency using Micro-ACKs
abstract
Retransmissions reduce the efficiency of data communication in wireless networks because of: (i) per-retransmission packet headers, (ii) contention overhead on every retransmission, and (iii) redundant bits in every retransmission. In fact, every retransmission nearly doubles the time to successfully deliver the packet. To improve spectrum efficiency in a lossy environment, we propose a new in-frame retransmission scheme using uACKs. Instead of waiting for the entire transmission to end before sending the ACK, the receiver sends smaller uACKs for every few symbols, on a separate narrow feedback channel. Based on these uACKs, the sender only retransmits the lost symbols after the last data symbol in the frame, thereby adaptively changing the frame size to ensure it is successfully delivered. We have implemented uACK on the Sora platform. Experiments with our prototype validate the feasibility of symbol-level uACK . By significantly reducing the retransmistion overhead, the sender is able to aggressively use higher data rate for a lossy link. Both improve the overall network efficiency. Our experimental results from a controlled environment and an 9-node software radio testbed show that uACK can have up to 140% throughput gain over 802.11g and up to 60% gain over the best known retransmission scheme.
Jiansong Zhang 0001, Haichen Shen, Kun Tan 0001, Ranveer Chandra, Yongguang Zhang, Qian Zhang 0001
MobiCom1
2010 MPAP: virtualization architecture for heterogenous wireless APs
abstract
This demonstration shows a novel virtualization architecture, called Multi-Purpose Access Point (MPAP), which can virtualize multiple heterogenous wireless standards based on software radio. The basic idea is to deploy a wide-band radio front-end to receive wireless signals from all wireless standards sharing the same spectrum band, and use separate software base-bands to demodulate information stream for each wireless standard. Based on software radio, MPAP consolidates multiple wireless devices into single hardware platform, and allows them to share the same general-purpose computing resource. Different software base-bands can easily communicate and coordinate with one another. Thus, it also provides better coexistence among heterogenous wireless standards. As an example, we demonstrate to use non-contiguous OFDM in 802.11g PHY to avoid the mutual interference with narrow-band ZigBee communication.
Ji Fang, Jiansong Zhang 0001, Haichen Shen, Yongguang Zhang
SIGCOMM3
2010 Fine-grained channel access in wireless LAN
abstract
Modern communication technologies are steadily advancing the physical layer (PHY) data rate in wireless LANs, from hundreds of Mbps in current 802.11n to over Gbps in the near future. As PHY data rates increase, however, the overhead of media access control (MAC) progressively degrades data throughput efficiency. This trend reflects a fundamental aspect of the current MAC protocol, which allocates the channel as a single resource at a time.
Ji Fang, Yuanyang Zhang, Shouyuan Chen, Lixin Shi, Jiansong Zhang 0001, Yongguang Zhang
SIGCOMM6
2010 Experimenting software radio with the Sora platform
abstract
Sora is a fully programmable, high performance software radio platform based on commodity general-purpose PC. In this demonstration, we illustrate the main features of the Sora platform that provide researchers flexible and powerful means to conduct wireless experiments at different levels with various goals. Specifically, the demonstrator will show four useful applications for wireless research that are built based on the Sora platform: 1) A capture tool that allows one to take a snapshot on a wireless channel; 2) a signal generation tool that allows one to transmit arbitrary baseband wave-form over the air, from a monophonic tone to a complex modulated frame; 3) an on-line real-time receiving application that uses the Sora User-Mode Extension; and 4) a fully featured Software radio WiFi driver (SoftWiFi) that can seamlessly inter-operate with commercial WiFi cards.
Jiansong Zhang 0001, Sen Xiang, Qiufeng Yin, Ji Fang, Yongguang Zhang
SIGCOMM1
2009 SAM: enabling practical spatial multiple access in wireless LAN
abstract
Spatial multiple access holds the promise to boost the capacity of wireless networks when an access point has multiple antennas. Due to the asynchronous and uncontrolled nature of wireless LANs, conventional MIMO technology does not work efficiently when concurrent transmissions from multiple stations are uncoordinated. In this paper, we present the design and implementation of a crosslayer system, called SAM, that addresses the challenges of enabling spatial multiple access for multiple devices in a random access network like WLAN. SAM uses a chain-decoding technique to reliably recover the channel parameters for each device, and iteratively decode concurrent frames with misaligned symbol timings and frequency offsets. We propose a new MAC protocol, called CCMA, to enable concurrent transmissions by different mobile stations while remaining backward compatible with 802.11. Finally, we implement the PHY and MAC layer of SAM using the Sora high-performance software radio platform. Our evaluation results under real wireless conditions show that SAM can improve network uplink throughput by 70% with two antennas over 802.11.
Ji Fang, Wei Wang 0002, Jiansong Zhang 0001, Mi Chen, Geoffrey M. Voelker
MobiCom5
2009 Sora: High Performance Software Radio Using General Purpose Multi-core Processors
Jiansong Zhang 0001, Ji Fang, Yusheng Ye, Yongguang Zhang, Wei Wang 0002, Geoffrey M. Voelker
NSDI2
2009 XOR Rescue: Exploiting Network Coding in Lossy Wireless Networks
abstract
It is well-known that wireless links are error-prone and require retransmissions for recovering frames from errors and losses. Network coding (NC) has been proposed for more efficient MAC-layer retransmissions in WLANs. However, existing schemes employed the reception report mechanism, which is both inefficient and expensive. Furthermore, they considered neither fairness nor the effects of time-varying heterogeneous wireless networks. These issues are critical for achieving full benefit of network coding. Without addressing them, these schemes may even impair system performance. In this paper, a novel MAC-layer retransmission scheme, namely XOR rescue (XORR) is proposed. It estimates the reception status without extra overheads and devises a new coding metric, which accommodates the effects of the frames size and the channel condition. Finally, XORR employs NC-aware fair opportunistic scheduling, which is theoretically proven to be fair, i.e. not only the service time is evenly allocated, but also it always improves the expected goodput for every wireless station. It is further verified by theoretic analyses, extensive simulations and testbed experiments. Our results show that XORR outperforms the non-coding fair opportunistic scheduling and 802.11 by 25% and 40%, respectively.
Fang-Chun Kuo, Xiang-Yang Li 0001, Jiansong Zhang 0001, Xiaoming Fu 0001
SECON4
2008 A Practical SNR-Guided Rate Adaptation
abstract
Rate adaptation is critical to the system performance of wireless networks. Typically, rate adaptation is considered as a MAC layer mechanism in IEEE 802.11. Most previous work relies only on frame losses to infer channel quality, but performs poorly if frame losses are mainly caused by interference. Recently SNR- based rate adaptation schemes have been proposed, but most of them have not been studied in a real environment. In this paper, we first conduct a systematic measurement-based study to confirm that in general SNR is a good prediction tool for channel quality, and identify two key challenges for this to be used in practice: (1) The SNR measures in hardware are often uncalibrated, and thus the SNR thresholds are hardware dependent. (2) The direct prediction from SNR to frame delivery ratio (FDR) is often over optimistic under interference conditions. Based on these observations, we present a novel practical SNR- Guided Rate Adaptation (SGRA) scheme. We implement and evaluate SGRA in a real test-bed and compare it with other three algorithms: ARF, RRAA and HRC. Our results show that SGRA outperforms the other three algorithms in all cases we have tested.
Jiansong Zhang 0001, Yongguang Zhang
INFOCOM1