EDBT 2026 Demo / reviewers in the wild / expert
Bojie Li
dblp:122/5219 · also Bojie Lv
· DBLP profile ↗
48ranked-venue papers
11as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 33 · 10 first-author · 15 since 2021Systems, architecture and hardware · 6 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Advancing RDMA Scalability With High PerformanceabstractDue to its superior performance, Remote Direct Memory Access (RDMA) has been widely deployed in data center networks. It provides applications with ultra-high throughput, ultra-low latency, and far lower CPU utilization than TCP/IP software network stack. However, the connection states that must be stored on the RDMA NIC (RNIC) and the small NIC memory result in poor scalability. The performance drops significantly when the RNIC needs to maintain a large number of concurrent connections. We propose StaR (Stateless RDMA), which solves the scalability problem of RDMA by transferring states to the other communication end in a trusted network. Leveraging the asymmetric communication pattern in data center applications, StaRlets the communication end with low NIC memory usage to save states for the other end with high NIC memory usage, thus making the RNIC on the bottleneck side stateless. We implemented StaR on an FPGA board with a 10Gbps network port and NS-3, evaluating its performance on a testbed with 9 machines, each equipped with StaR NICs, and verified its scalability stability by conducting a larger-scale simulation with 200 fully connected nodes using a 100Gbps link. The experimental results show that in high concurrency scenarios, the throughput of StaR can reach up to 4.13x and 1.35x of the original RNIC and the latest software-based solution, respectively. Xijin Yin, Guo Chen 0001, Xizheng Wang, Huichen Dai, Bojie Li, Binzhang Fu, Kun Tan 0002 |
IEEE Trans. Netw. | 6 |
| 2026 | mmAlert: A Simultaneous Device Localization and Target Tracking System via Cooperative Passive SensingabstractIn this paper, a cooperative passive sensing system in millimeter-wave (mmWave) band for simultaneous device localization and target tracking, namely mmAlert, is proposed. Specifically, in uplink communication with at least two transmitters, the receiver receives the line-of-sight (LoS) signals and the scattered signals off a moving target, respectively. Based on the received signals of the sensing time intervals, when a passive target moves along one or multiple unknown trajectories, mmAlert could measure the angles-of-arrival (AoAs) and bistatic Doppler frequencies of the echoes from the sensing target, and then jointly estimate the locations of the transmitters and the trajectories of the target. Specifically, the transmitters’ locations and the moving target’s trajectories can be searched by minimizing the weighted mean squared error of the AoA and Doppler measurements. The optimal solution of the minimization problem is prohibitive due to the large number of variables. Hence, a low-complexity algorithm based on the alternating optimization is proposed, where the extended Kalman filter (EKF) is introduced to quickly shape the trajectories. The mmAlert is implemented in a 60GHz communication testbed. The experiment shows with the received signal spanning a single trajectory, the average localization error of the transmitters and average trajectory reconstruction error are 0.76 m and 0.29 m, respectively. The average errors are suppressed to 0.07 m and 0.2 m respectively, if the received signal spanning 50 trajectories is used. This justifies the benefit of trajectory diversity in localization and tracking. Chao Yu 0007, Bojie Li, Chunxi Chen, Rui Wang 0007 |
IEEE Trans. Wirel. Commun. | 2 |
| 2025 | Fast and Scalable Selective Retransmission for RDMA
Peihao Huang, Guo Chen 0001, Xin Zhang 0117, Huijun Shen, Ying Bian, Yuanwei Lu, Zhenyuan Ruan, Bojie Li, Jiansong Zhang 0001, Yongfeng Liu, Zhigang Chen 0001 |
INFOCOM | 10 |
| 2025 | ReinWiFi: Application-Layer QoS Optimization of WiFi Networks with Reinforcement LearningabstractThe enhanced distributed channel access (EDCA) mechanism is used in current wireless fidelity (WiFi) networks to support priority requirements of heterogeneous applications. However, the EDCA mechanism can not adapt to particular quality-of-service (QoS) objective, network topology, and interference level. In this paper, a novel reinforcement-learningbased scheduling framework is proposed and implemented to optimize the application-layer quality-of-service (QoS) of a WiFi network with commercial adapters and unknown interference. Particularly, application-layer tasks of file delivery and delaysensitive communication are jointly scheduled by adjusting the contention window sizes and application-layer throughput limitation, such that the throughput of the former and the round trip time of the latter can be optimized. Due to the unknown interference and vendor-dependent implementation of the WiFi adapters, the relation between the scheduling policy and the system QoS is unknown. Hence, a reinforcement learning method is proposed, in which a novel Q-network is trained to map from the historical scheduling parameters and QoS observations to the current scheduling action. It is demonstrated on a testbed that the proposed framework can achieve a significantly better performance than the EDCA mechanism. Qianren Li, Bojie Li, Yuncong Hong, Rui Wang 0007 |
VTC2025-Spring | 2 |
| 2025 | A Dynamic Improvement Framework for Vehicular Task OffloadingabstractIn this paper, the task offloading from vehicles with random velocities is optimized via a novel dynamic improvement framework. Particularly, in a vehicular network with multiple vehicles and base stations (BSs), computing tasks of vehicles are offloaded via BSs to an edge server. Due to the random velocities, the exact trajectories of vehicles cannot be predicted in advance. Hence, instead of deterministic optimization, the cell association, uplink time and throughput allocation of multiple vehicles in a period of task offloading are formulated as a finite-horizon Markov decision process. In the proposed solution framework, we first obtain a reference scheduling scheme of cell association, uplink time and throughput allocation via deterministic optimization at the very beginning. The reference scheduling scheme is then used to approximate the value functions of the Bellman's equations, and the actual scheduling action is determined in each time slot according to the current system state and approximate value functions. Thus, the intensive computation for value iteration in the conventional solution is eliminated. Moreover, a nontrivial average cost upper bound is provided for the proposed solution framework. In the simulation, the random trajectories of vehicles are generated from a high-fidelity traffic simulator. It is shown that the performance gain of the proposed scheduling framework over the baselines is significant. Qianren Li, Yuncong Hong, Bojie Li, Rui Wang 0007 |
WCNC | 3 |
| 2025 | A Dynamic Programming Framework for Vehicular Task Offloading With Successive Action ImprovementabstractIn this paper, task offloading from vehicles with random velocities is optimized via a novel dynamic programming framework. Particularly, in a vehicular network with multiple vehicles and base stations (BSs), computing tasks of vehicles are offloaded via BSs to an edge server. Due to the random velocities, the exact locations of vehicles versus time, namely trajectories, cannot be determined in advance. Hence, instead of deterministic optimization, the cell association, uplink time, and throughput allocation of multiple vehicles during a period of task offloading are formulated as a finite-horizon Markov decision process. In order to derive a low-complexity solution algorithm, a two-time-scale framework is proposed. The scheduling period is divided into super slots, each super slot is further divided into a number of time slots. At the beginning of each super slot, we first obtain a reference scheduling scheme of cell association, uplink time and throughput allocation via deterministic optimization, yielding an approximation of the optimal value function. Within the super slot, the actual scheduling action of each time slot is determined by making improvement to the approximate value function according to the system state. Due to the successive improvement framework, a non-trivial average cost upper bound could be derived. In the simulation, the random trajectories of vehicles are generated from a high-fidelity traffic simulator. It is shown that the performance gain of the proposed scheduling framework over the baselines is significant. Qianren Li, Yuncong Hong, Bojie Li, Rui Wang 0007 |
IEEE Trans. Commun. | 3 |
| 2024 | Sensing-Assisted Adaptive Channel Contention for Mobile Delay-Sensitive CommunicationsabstractThis paper proposes an adaptive channel contention mechanism to optimize the queuing performance of a distributed millimeter wave (mmWave) uplink system with the capability of environment and mobility sensing. The mobile agents determine their back-off timer parameters according to their local knowledge of the uplink queue lengths, channel quality, and future channel statistics, where the channel prediction relies on the environment and mobility sensing. The optimization of queuing performance with this adaptive channel contention mechanism is formulated as a decentralized multi-agent Markov decision process (MDP). Although the channel contention actions are determined locally at the mobile agents, the optimization of local channel contention policies of all mobile agents is conducted in a centralized manner according to the system statistics before the scheduling. In the solution, the local policies are approximated by analytical models, and the optimization of their parameters becomes a stochastic optimization problem along an adaptive Markov chain. An unbiased gradient estimation is proposed so that the local policies can be optimized efficiently via the stochastic gradient descent method. It is demonstrated by simulation that the proposed gradient estimation is significantly more efficient in optimization than the existing methods, e.g., simultaneous perturbation stochastic approximation (SPSA). Bojie Li, Qianren Li, Rui Wang 0007 |
GLOBECOM | 1 |
| 2024 | Delay-Aware Two-Time-Scale Scheduling for mmWave Systems With Mobility and Environment KnowledgeabstractA two-time-scale approximate Markov decision process (MDP) is proposed to optimize the uplink queuing performance of a millimeter wave (mmWave) communication system. Exploiting the wireless sensing techniques, the locations of signal reflectors, blockers, and mobile agents, as well as the mobility pattern of agents, become available knowledge to facilitate predictive scheduling. Notice that the state variation of the wireless channel fading and transmission queue is much faster than that of mobile agents’ locations. The joint optimization of uplink power adaptation, time allocation, and analog beamforming of all the frames is formulated as a two-time-scale MDP with the average uplink energy, queuing length, and buffer overflow rate in the minimization objective. The scheduling in the larger time scale with an infinite horizon, i.e., the analog transceiver beamforming, adapts with the random motion of mobile agents; whereas the scheduling in the smaller time scale with a finite horizon, i.e., uplink power and time allocation, adapts with both the small-scale channel fading, queue dynamics and large-scale random motion. The optimal solution of two-time-scale MDP requests iterative optimization between Bellman’s equations of both time scales, whose computation complexity is prohibitive. A novel low-complexity solution framework is then proposed to obtain the optimal larger-time-scale and sub-optimal smaller-time-scale policies, where the performance of the proposed scheme is bounded analytically. Benefiting from the motion sensing and blockage prediction, the proposed scheduling scheme outperforms existing benchmarks in the numerical simulations. Bojie Li, Rui Wang 0007 |
IEEE Trans. Commun. | 1 |
| 2024 | COSMO: Dynamic Uploading Scheduling in mmWave-Based Sensor Networks with Mobile BlockersabstractWireless sensor networks (WSNs) leveraging millimeter wave (mmWave) communication for bandwidth-demanding applications is considered in this article. Despite the large bandwidth, the delivery of delay-sensitive information collected by sensors may still face significant latency due to the vulnerability to intermittent link blockage. Hence, the guarantee of low age of information (AoI) in mmWave WSNs is not straightforward. In this article, the wireless sensing and dynamic programming techniques are jointly exploited to relieve the above issue. The former tracks the human blockers and predicts the chance of link blockage; the latter optimizes the transmission of multiple sensors based on the prediction. Particularly, the long-term optimization of sampling, uplink time and power allocation policies in a sensor network can be formulated as an infinite-horizon Markov decision process (MDP) with discounted cost, where the state transition probabilities can be predicted via wireless sensing. A novel low-complexity solution framework, namely COSMO, with a guaranteed performance in the worst case, is proposed. Simulations show that compared with heuristic benchmarks, benefiting from the prediction of the link blockage, COSMO can significantly suppress the average system cost, which consists of both AoI and energy consumption. Yifei Sun 0003, Bojie Li, Haisheng Tan, Rui Wang 0007, Francis C. M. Lau 0001 |
ACM Trans. Sens. Networks | 2 |
| 2024 | Predictive Delay-Aware Scheduling With Receiver Rotation Detection and mmWave Channel LearningabstractIn this paper, the joint downlink delay-aware scheduling in a large time span, where the rotation of User Equipments (UEs) may lead to significant channel variation, is investigated via a novel approximate Markov Decision Process (MDP) method. Specifically, we consider the joint downlink power allocation and receiving UE selection of a number of successive frames in a millimeter Wave (mmWave) system with quasi-static scattering clusters in the channel and rotating UEs. The propagation statistics of scattering clusters can be tracked via a learning method. Since the rotation of UEs can be detected, future channel statistics can be forecast via embedded motion sensors. Hence, the overall scheduling is formulated as a finite-horizon MDP with non-stationary predictable state transition probabilities, where the average queuing delay and probability of transmission buffer overflow are considered in the objective of scheduling optimization. A novel low-complexity solution framework with an analytical performance bound is proposed to save the efforts of value iteration. Benefiting from the forecast of system statistics, superior performance to the benchmarks is shown by numerical simulations, particularly in the suppression of buffer overflow rate. Preliminary experiments via an mmWave testbed are conducted to demonstrate the feasibility of the sensor-assisted mmWave beam alignment. Yifei Sun 0003, Bojie Li, Rui Wang 0007, Haisheng Tan, Francis C. M. Lau 0001 |
IEEE Trans. Wirel. Commun. | 2 |
| 2024 | Rate-Splitting With Hybrid Messages: DoF Analysis of the Two-User MIMO Broadcast Channel With Imperfect CSITabstractMost of the existing research on degrees-of-freedom (DoF) with imperfect channel state information at the transmitter (CSIT) assume the messages are private, which may not reflect reality as the two receivers can request the same content. To overcome this limitation, we therefore consider the hybrid unicast and multicast messages. In particular, we characterize the optimal DoF region for the two-user multiple-input multiple-output (MIMO) broadcast channel (BC) with imperfect CSIT and hybrid messages. For the converse, we establish a three-step procedure to exploit the utmost possible relaxation. For the achievability, since the DoF region is with specific three-dimensional structure regarding antenna configurations and CSIT qualities, we verify the existence or non-existence of corner point candidates via the feature of antenna configurations and CSIT qualities categorization, and provide a hybrid message-aware rate-splitting scheme. Besides, we show that to achieve the strictly positive corner points, it is unnecessary to split the unicast messages into private and common parts. This implies adding a multicast message may mitigate the rate-splitting complexity. Tong Zhang 0026, Yufan Zhuang, Gaojie Chen 0001, Shuai Wang 0004, Bojie Li, Rui Wang 0007, Pei Xiao 0001 |
IEEE Trans. Wirel. Commun. | 5 |
| 2023 | FastWake: Revisiting Host Network Stack for Interrupt-mode RDMAabstractPolling and interrupt has long been a trade-off in RDMA systems. Polling has lower latency but each CPU core can only run one thread. Interrupt enables time sharing among multiple threads but has higher latency. Many applications such as databases have hundreds of threads, which is much larger than the number of cores. So, they have to use interrupt mode to share cores among threads, and the resulting RDMA latency is much higher than the hardware limits. In this paper, we analyze the root cause of high costs in RDMA interrupt delivery, and present FastWake, a practical redesign of interrupt-mode RDMA host network stack using commodity RDMA hardware, Linux OS, and unmodified applications. Our first approach to fast thread wake-up completely removes interrupts. We design a per-core dispatcher thread to poll all the completion queues of the application threads on the same core, and utilize a kernel fast path to context switch to the thread with an incoming completion event. The approach above would keep CPUs running at 100% utilization, so we design an interrupt-based approach for scenarios with power constraints. Observing that waking up a thread on the same core as the interrupt is much faster than threads on other cores, we dynamically adjust RDMA event queue mappings to improve interrupt core affinity. In addition, we revisit the kernel path of thread wake-up, and remove the overheads in virtual file system (VFS), locking, and process scheduling. Experiments show that FastWake can reduce RDMA latency by 80% on x86 and 77% on ARM at the cost of < 30% higher power utilization than traditional interrupts, and the latency is only 0.3 ∼ 0.4 μ s higher than the limits of underlying hardware. When power saving is desired, our interrupt-based approach can still reduce interrupt-mode RDMA latency by 59% on x86 and 52% on ARM. Bojie Li, Zihao Xiang, Xiaoliang Wang 0001, Han Ruan, Jingbin Zhou, Kun Tan 0002 |
APNet | 1 |
| 2023 | Predictive Resource Allocation in mmWave Systems with Rotation DetectionabstractMillimeter wave (MmWave) has been regarded as a promising technology to support high-capacity communications in 5G era. However, its high-layer performance such as latency and packet drop rate in the long term highly depends on resource allocation because mmWave channel suffers significant fluctuation with rotating users due to mmWave sparse channel property and limited field-of-view (FoV) of antenna arrays. In this paper, downlink transmission scheduling considering rotation of user equipments (UE) and limited antenna FoV in an mmWave system is optimized via a novel approximate Markov decision process (MDP) method. Specifically, we consider the joint downlink UE selection and power allocation in a number of frames where future orientations of rotating UEs can be predicted via embedded motion sensors. The problem is formulated as a finite-horizon MDP with non-stationary state transition probabilities. A novel low-complexity solution framework is proposed via one iteration step over a base policy whose average future cost can be predicted with analytical expressions. It is demonstrated by simulations that compared with existing benchmarks, the proposed scheme can schedule the downlink transmission and suppress the packet drop rate efficiently in non-stationary mmWave links. Yifei Sun 0003, Bojie Li, Rui Wang 0007, Haisheng Tan, Francis C. M. Lau 0001 |
ICC | 2 |
| 2023 | Dynamic Uploading Scheduling in mmWave-Based Sensor Networks via Mobile Blocker DetectionabstractThe freshness of information, measured as Age of Information (AoI), is critical for many applications in next-generation wireless sensor networks (WSNs). Due to its high bandwidth, millimeter wave (mmWave) communication is seen to be frequently exploited in WSNs to facilitate the deployment of bandwidth-demanding applications. However, the vulnerability of mmWave to user mobility typically results in link blockage and thus postponed real-time communications. In this paper, joint sampling and uploading scheduling in an AoI-oriented WSN working in mmWave band is considered, where a single human blocker is moving randomly and signal propagation paths may be blocked. The locations of signal reflectors and the real-time position of the blocker can be detected via wireless sensing technologies. With the knowledge of blocker motion pattern, the statistics of future wireless channels can be predicted. As a result, the AoI degradation arising from link blockage can be forecast and mitigated. Specifically, we formulate the long-term sampling, uplink transmission time and power allocation as an infinite-horizon Markov decision process (MDP) with discounted cost. Due to the curse of dimensionality, the optimal solution is infeasible. A novel low-complexity solution framework with guaranteed performance in the worst case is proposed where the forecast of link blockage is exploited in a value function approximation. Simulations show that compared with several heuristic benchmarks, our proposed policy, benefiting from the awareness of link blockage, can reduce average cost up to 49.6%. Yifei Sun 0003, Bojie Li, Rui Wang 0007, Haisheng Tan, Francis C. M. Lau 0001 |
ICPADS | 2 |
| 2023 | Joint Beamforming Design for Integrated Passive Sensing and Communications in V2I NetworksabstractThis letter investigates joint beamforming and power optimization in a vehicle-to-infrastructure (V2I) system for integrated passive sensing and communications techniques. Specifically, the base station (BS) with multiple antennas delivers downlink data to multiple vehicles, respectively, and a sensing receiver simultaneously utilizes the downlink data signal to sense the motion of one vehicle. The sensing receiver is deployed separately. It collects the line-of-sight (LoS) signals from the BS and the scattered signals via the target vehicle, such that the distance and velocity of the target vehicle can be detected in a passive manner. As a result, the joint downlink beamforming and power optimization at the BS would be able to maximize the weighted summation of downlink data rates, subject to constraints on the signal-to-interference-plus-noise ratios (SINRs) of the two signals for passive sensing. In order to solve the above non-convex optimization problem, we first derive the optimal receiving beams for the data and sensing receivers given arbitrary transmission beam design, and then propose a low-complexity iterative algorithm to find a sub-optimal design of transmission beams based on the successive convex approximation (SCA) method. Simulations demonstrate the good performance, fast convergence and useful design insights. Bojie Li, Tong Zhang 0026, Rui Wang 0007, Pak-Chung Ching |
IEEE Signal Process. Lett. | 2 |
| 2023 | Modeling the Interplay between Loop Tiling and Fusion in Optimizing Compilers Using Affine RelationsabstractLoop tiling and fusion are two essential transformations in optimizing compilers to enhance the data locality of programs. Existing heuristics either perform loop tiling and fusion in a particular order, missing some of their profitable compositions, or execute ad-hoc implementations for domain-specific applications, calling for a generalized and systematic solution in optimizing compilers. In this article, we present a so-called basteln (an abbreviation for backward slicing of tiled loop nests) strategy in polyhedral compilation to better model the interplay between loop tiling and fusion. The basteln strategy first groups loop nests by preserving their parallelism/tilability and next performs rectangular/parallelogram tiling to the output groups that produce data consumed outside the considered program fragment. The memory footprints required by each tile are then computed, from which the upward exposed data are extracted to determine the tile shapes of the remaining fusion groups. Such a tiling mechanism can construct complex tile shapes imposed by the dependences between these groups, which are further merged by a post-tiling fusion algorithm for enhancing data locality without losing the parallelism/tilability of the output groups. The basteln strategy also takes into account the amount of redundant computations and the fusion of independent groups, exhibiting a general applicability. We integrate the basteln strategy into two optimizing compilers, with one a general-purpose optimizer and the other a domain-specific compiler for deploying deep learning models. The experiments are conducted on CPU, GPU, and a deep learning accelerator to demonstrate the effectiveness of the approach for a wide class of application domains, including deep learning, image processing, sparse matrix computation, and linear algebra. In particular, the basteln strategy achieves a mean speedup of 1.8× over cuBLAS/cuDNN and 1.1× over TVM on GPU when used to optimize deep learning models; it also outperforms PPCG and TVM by 11% and 20%, respectively, when generating code for the deep learning accelerator. Jie Zhao 0002, Jinchen Xu, Peng Di, Wang Nie, Yanzhi Yi, Zhen Geng, Renwei Zhang, Bojie Li, Zhiliang Gan, Xuefeng Jin 0004 |
ACM Trans. Comput. Syst. | 10 |
| 2022 | Distributed Job Dispatching in Edge Computing Networks With Random Transmission Latency: A Low-Complexity POMDP ApproachabstractJob dispatching is a fundamental problem in edge computing for load balancing among multiple edge servers. When implementing an edge computing system with distributed job dispatchers in a sizable network, such as a metropolitan area network (MAN), the highly dynamic transmission latency is nonnegligible, which could lead to outdated information being shared. Moreover, the fully observed system state is beyond reach as the reception of any broadcast is time consuming. In this article, we investigate the online distributed job dispatching problem in edge computing, where multiple access points (APs) collect jobs and then dispatch each job to an edge server. The distributed dispatcher on each AP would receive partially and outdated information exchanged via periodic broadcast. Hence, we formulate the distributed job dispatching problem by leveraging the partially observable Markov decision process (POMDP) and propose a novel approximate Markov decision process (MDP) solution framework, calledDecMDP, that bypasses the huge time complexity of conventional POMDP solutions. Both analytical and semi-analytical performance lower bounds are derived for the approximate MDP solution. Furthermore, we extendDecMDPto handle a more general scenario wherea prioriknowledge of the system is absent. Finally, extensive simulations based on the Google Cluster traces show that our policy can achieve the best performance when compared with heuristic baselines, e.g., achieving 20.67% reduction in average job response time, and consistently performs well under various parameter settings. Yuncong Hong, Bojie Li, Rui Wang 0007, Haisheng Tan, Zhenhua Han, Francis C. M. Lau 0001 |
IEEE Internet Things J. | 2 |
| 2021 | StaR: Breaking the Scalability Limit for RDMAabstractDue to its superior performance, Remote Direct Memory Access (RDMA) has been widely deployed in data center networks. It provides applications with ultra-high throughput, ultra-low latency, and far lower CPU utilization than TCP/IP software network stack. However, the connection states that must be stored on the RDMA NIC (RNIC) and the small NIC memory result in poor scalability. The performance drops significantly when the RNIC needs to maintain a large number of concurrent connections.We propose StaR (Stateless RDMA), which solves the scalability problem of RDMA by transferring states to the other communication end. Leveraging the asymmetric communication pattern in data center applications, StaR lets the communication end with low concurrency save states for the other end with high concurrency, thus making the RNIC on the bottleneck side to be stateless. We have implemented StaR on an FPGA board with 10Gbps network port and evaluated its performance on a testbed with 9 machines all equipped with StaR NICs. The experimental results show that in high concurrency scenarios, the throughput of StaR can reach up to 4.13x and 1.35x of the original RNIC and the latest software-based solution, respectively. Xizheng Wang, Guo Chen 0001, Xijin Yin, Huichen Dai, Bojie Li, Binzhang Fu, Kun Tan 0002 |
ICNP | 5 |
| 2021 | AKG: automatic kernel generation for neural processing units using polyhedral transformationsabstractExisting tensor compilers have proven their effectiveness in deploying deep neural networks on general-purpose hardware like CPU and GPU, but optimizing for neural processing units (NPUs) is still challenging due to the heterogeneous compute units and complicated memory hierarchy. Jie Zhao 0002, Bojie Li, Wang Nie, Zhen Geng, Renwei Zhang, Xiong Gao, Zheng Li 0035, Peng Di, Xuefeng Jin 0004 |
PLDI | 2 |
| 2021 | 1Pipe: scalable total order communication in data center networksabstractThis paper proposes 1Pipe, a novel communication abstraction that enables different receivers to process messages from senders in a consistent total order. More precisely, 1Pipe provides both unicast and scattering (i.e., a group of messages to different destinations) in a causally and totally ordered manner. 1Pipe provides a best effort service that delivers each message at most once, as well as a reliable service that guarantees delivery and provides restricted atomic delivery for each scattering. 1Pipe can simplify and accelerate many distributed applications, e.g., transactional key-value stores, log replication, and distributed data structures. Bojie Li, Gefei Zuo, Wei Bai 0001 |
SIGCOMM | 1 |
| 2020 | Adaptive Video Streaming for Massive MIMO Networks via Novel Approximate MDPabstractThe scheduling of downlink video streaming in a massive multiple-input-multiple-output (MIMO) network is considered in this paper, where active users arrive randomly to request video contents of a finite playback duration via their service base stations. Each video consists of a sequence of segments, which can be transmitted to the requesting users with variable video bitrates. To facilitate adaptive video streaming, a number of physical-layer frames are grouped as a super frame. We formulate the adaptation of transmitted segment number, frame allocation and segment bitrate in all the super frames as an infinite-horizon Markov decision process (MDP), whose objective is a discounted measurement of the average Quality-of-Experience (QoE). A novel approximate MDP method is proposed to obtain a low-complexity scheduling policy. Specifically, a baseline policy is introduced and its asymptotic value function is derived analytically. The low-complexity scheduling policy will be obtained from one-step iteration based on the analytical expression, which becomes a performance lower bound on the derived policy. It is shown by simulations that the proposed low-complexity scheduling policy has significant performance gain over the baseline policy. Qiao Lan, Bojie Li, Rui Wang 0007, Yi Gong 0001, Kaibin Huang |
ICC | 2 |
| 2020 | Online Distributed Job Dispatching with Outdated and Partially-Observable InformationabstractIn this paper, we investigate online distributed job dispatching in an edge computing system residing in a Metropolitan Area Network (MAN). Specifically, job dispatchers are implemented on access points (APs) which collect jobs from mobile users and distribute each job to a server at the edge or the cloud. A signaling mechanism with periodic broadcast is introduced to facilitate cooperation among APs. The transmission latency is non-negligible in MAN, which leads to outdated information sharing among APs. Moreover, the fully-observed system state is discouraged as reception of all broadcast is time consuming. Therefore, we formulate the distributed optimization of job dispatching strategies among the APs as a Markov decision process with partial and outdated system state, i.e., partially observable Markov Decision Process (POMDP). The conventional solution for POMDP is impractical due to huge time complexity. We propose a novel low-complexity solution framework for distributed job dispatching, based on which the optimization of job dispatching policy can be decoupled via an alternative policy iteration algorithm, so that the distributed policy iteration of each AP can be made according to partial and outdated observation. A theoretical performance lower bound is proved for our approximate MDP solution. Furthermore, we conduct extensive simulations based on the Google Cluster trace. The evaluation results show that our policy can achieve as high as 20.67% reduction in average job response time compared with heuristic baselines, and our algorithm consistently performs well under various parameter settings. Yuncong Hong, Bojie Li, Rui Wang 0007, Haisheng Tan, Zhenhua Han, Hao Zhou 0001, Francis C. M. Lau 0001 |
MSN | 2 |
| 2020 | Joint Optimization of File Placement and Delivery in Cache-Assisted Wireless Networks With Limited Lifetime and Cache SpaceabstractIn this paper, the scheduling of downlink file transmission in one cell with the assistance of cache nodes with finite cache space is studied. Specifically, requesting users arrive randomly and the base station (BS) reactively multicasts files to the requesting users and selected cache nodes. The latter can offload the traffic in their coverage areas from the BS. We consider the joint optimization of the abovementioned file placement and delivery within a finite lifetime subject to the cache space constraint. Within the lifetime, the allocation of multicast power and symbol number for each file transmission at the BS is formulated as a dynamic programming problem with a random stage number. Note that there are no existing solutions to this problem. We develop an asymptotically optimal solution framework by transforming the original problem to an equivalent finite-horizon Markov decision process (MDP) with a fixed stage number. A novel approximation approach is then proposed to address the curse of dimensionality, where the analytical expressions of approximate value functions are provided. We also derive analytical bounds on the exact value function and approximation error. The approximate value functions depend on some system statistics, e.g., requesting users’ distribution. One reinforcement learning algorithm is proposed for the scenario where these statistics are unknown. Bojie Li, Rui Wang 0007, Ying Cui 0001, Yi Gong 0001, Haisheng Tan |
IEEE Trans. Commun. | 1 |
| 2020 | Online Deadline-Aware Task Dispatching and Scheduling in Edge ComputingabstractIn this article, we study online deadline-aware task dispatching and scheduling in edge computing. We jointly considerthe management of the networking and computing resources to meet the maximum number of deadlines. We propose an online algorithm, named Dedas, which greedily schedules newly arriving tasks and considers whether to replace some existing tasks in order to make the new deadlines satisfied. We derive a non-trivial competitive ratio of Dedas theoretically, and our analysis is asymptotically tight. Besides, we implement a distributed approximation D - Dedas with a better scalability and less than 10 percent performance loss compared with the centralized algorithm Dedas. We then build DeEdge, an edge computing testbed installed with typical latency-sensitive applications such as IoT sensor monitoring and face matching. We adopt a real-world data trace from the Google cluster for large-scale emulations. Extensive testbed experiments and simulations demonstrate that the deadline miss ratio of Dedas is stable for online tasks, which is reduced by up to 60 percent compared with state-of-the-art methods. Moreover, Dedas performs well in minimizing the average task completion time. Jiaying Meng, Haisheng Tan, Xiang-Yang Li 0001, Zhenhua Han, Bojie Li |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2020 | Adaptive Video Streaming for Massive MIMO Networks via Approximate MDP and Reinforcement LearningabstractThe scheduling of downlink video streaming in a massive multiple-input multiple-output (MIMO) network is considered in this paper, where active users arrive randomly to request video contents of a finite playback duration via their service base stations (BSs). Each video content consisting of a sequence of segments can be transmitted to the requesting users with variable video bitrates. We formulate the joint control of transmitted segment number, frame allocation and segment bitrate in all the super frames (each comprising multiple frames) as an infinite-horizon Markov decision process (MDP). The maximization objective is a discounted measurement of the average Quality of Experience (QoE). Since there is no efficient method for scheduling design with random user arrivals and departures in the existing literature, a novel approximate MDP method is proposed to obtain a low-complexity scheduling policy, where a lower bound on its performance is derived. Specifically, we first introduce a baseline policy and derive its asymptotic value function. One-step policy iteration is then applied to improve this value function, yielding the mentioned low-complexity policy. Finally, we propose a novel and efficient reinforcement learning (RL) algorithm to evaluate the value function when the prior knowledge on user arrival intensity is absent. Qiao Lan, Bojie Li, Rui Wang 0007, Kaibin Huang, Yi Gong 0001 |
IEEE Trans. Wirel. Commun. | 2 |
| 2019 | Towards Stateless RNIC for Data Center NetworksabstractBecause of small NIC on-chip memory, the massive connection states maintained on Remote Direct Memory Access (RDMA) NIC (RNIC) significantly limit its scalability. When the number of concurrent connections grows, RNICs have to frequently fetch connection states from host memory, leading to dramatic performance degradation. In this paper, we propose StaR, which fundamentally solves this scalability issue by making RNIC stateless. Leveraging the asymmetric communication pattern in data center applications, the StaR RNIC stores zero connection-related states by moving all the connection states to the other end. Through careful design, StaR RNICs can maintain unchanged RDMA semantics and avoid security issues even when processing traffic statelessly. Preliminary simulation results show that StaR can improve the aggregate throughput by more than 160x (stress test) and 4x (application) compared to original RNICs. Pulin Pan, Guo Chen 0001, Xizheng Wang, Huichen Dai, Bojie Li, Binzhang Fu, Kun Tan 0002 |
APNet | 5 |
| 2019 | MDP-Based Scheduling Design for Mobile-Edge Computing Systems with Random User ArrivalabstractIn this paper, we investigate the scheduling design of a mobile-edge computing (MEC) system, where the random arrival of mobile devices with computation tasks in both spatial and temporal domains is considered. The binary computation offloading model is adopted. Every task is indivisible and can be computed at either the mobile device or the MEC server. We formulate the optimization of task offloading decision, uplink transmission device selection and power allocation in all the frames as an infinite-horizon Markov decision process (MDP). Due to the uncertainty in device number and location, conventional approximate MDP approaches to addressing the curse of dimensionality cannot be applied. A novel low- complexity sub-optimal solution framework is then proposed. We first introduce a baseline scheduling policy, whose value function can be derived analytically. Then, one-step policy iteration is adopted to obtain a sub-optimal scheduling policy whose performance can be bounded analytically. Simulation results show that the gain of the sub-optimal policy over various benchmarks is significant. Shanfeng Huang, Bojie Li, Rui Wang 0007 |
GLOBECOM | 2 |
| 2019 | Joint Heterogeneous Server Placement and Application Configuration in Edge ComputingabstractThe rapid development of the Internet of Things (IoT) has brought profound changes in the cloud computing paradigm. One promising computing model in IoT-related applications is edge computing, which can decrease the request response time by deploying edge servers close to IoT devices. Two fundamental problems in edge computing include how to place a limited number of edge servers on the candidate locations (e.g., the Access Points) and how to configure the applications on each server. In this paper, we jointly study the edge server placement and the application configuration to minimize the weighted sum of the service cost and edge server opening cost. We propose a local-search based algorithm, named SPAC, to solve the problem efficiently with its approximation ratio analyzed. Extensive simulations on Google data traces demonstrate SPAC reduces the total cost by up to 60% compared with state-of-the-art methods, and outperforms the baselines consistently in different parameter settings. Jiaying Meng, Chaoliang Zeng, Haisheng Tan, Zining Li, Bojie Li, Xiang-Yang Li 0001 |
ICPADS | 5 |
| 2019 | Dedas: Online Task Dispatching and Scheduling with Bandwidth Constraint in Edge ComputingabstractIn this paper, we study online deadline-aware task dispatching and scheduling in edge computing. We jointly consider management of the networking bandwidth and computing resources to meet the maximum number of deadlines. We propose an online algorithm Dedas, which greedily schedules newly arriving tasks and considers whether to replace some existing tasks in order to make the new deadlines satisfied. We derive a non-trivial competitive ratio theoretically, and our analysis is asymptotically tight. We then build DeEdge, an edge computing testbed installed with typical latency-sensitive applications such as IoT sensor monitoring and face matching. Besides, we adopt a real-world data trace from the Google cluster for large-scale emulations. Extensive testbed experiments and simulations demonstrate that the deadline miss ratio of Dedas is stable for online tasks, which is reduced by up to 60% compared with state-of-the-art methods. Moreover, Dedas performs well in minimizing the average task completion time. Jiaying Meng, Haisheng Tan, Wanli Cao, Liuyan Liu, Bojie Li |
INFOCOM | 6 |
| 2019 | Socksdirect: datacenter sockets can be fast and compatibleabstractCommunication intensive applications in hosts with multi-core CPU and high speed networking hardware often put considerable stress on the native socket system in an OS. Existing socket replacements often leave significant performance on the table, as well have limitations on compatibility and isolation. Bojie Li, Tianyi Cui, Wei Bai 0001 |
SIGCOMM | 1 |
| 2019 | Joint Downlink Scheduling for File Placement and Delivery in Cache-Assisted Wireless Networks With Finite File LifetimeabstractIn this paper, downlink transmission scheduling of popular files is optimized with the assistance of wireless cache nodes. Specifically, the requests of each file, which is further divided into a number of segments, are modeled as a Poisson point process within its finite lifetime. Two downlink transmission modes are considered: 1) the base station reactively multicasts the file segments to the requesting users and selected cache nodes and (2) the base station proactively multicasts some file segments to the selected cache nodes without requests. The cache nodes with decoded file segments can help to offload the traffic via other spectrum. Without the proactive multicast, we formulate the downlink transmission resource minimization as a dynamic programming problem with random stage number, which can be approximated via a finite-horizon Markov decision process (MDP) with fixed stage number. To address the prohibitively huge state space, we propose a low-complexity scheduling policy by linearly approximating the value functions of the MDP, where the bound on the approximation error is derived. Moreover, we propose a learning-based algorithm to evaluate the approximated value functions for unknown geographical distribution of requesting users. Finally, given the above reactive multicast policy, a proactive multicast policy is introduced to exploit the temporal diversity of shadowing effect. It is shown by simulation that the proposed low-complexity reactive multicast policy can significantly reduce the resource consumption at the base station, and the proactive multicast policy can further improve the performance. Bojie Li, Lexiang Huang, Rui Wang 0007 |
IEEE Trans. Commun. | 1 |
| 2019 | MP-RDMA: Enabling RDMA With Multi-Path Transport in DatacentersabstractRDMA is becoming prevalent because of its low latency, high throughput and low CPU overhead. However, in current datacenters, RDMA remains a single path transport which is prone to failures and falls short to utilize the rich parallel network paths. Unlike previous multi-path approaches, which mainly focus on TCP, this paper presents a multi-path transport for RDMA, i.e. MP-RDMA, which efficiently utilizes the rich network paths in datacenters. MP-RDMA employs three novel techniques to address the challenge of limited RDMA NICs on-chip memory size: 1) a multi-path ACK-clocking mechanism to distribute traffic in a congestion-aware manner without incurring per-path states; 2) an out-of-order aware path selection mechanism to control the level of out-of-order delivered packets, thus minimizes the meta data required to them; 3) a synchronise mechanism to ensure in-order memory update whenever needed. With all these techniques, MP-RDMA only adds 66B to each connection state compared to single-path RDMA. Our evaluation with an FPGA-based prototype demonstrates that compared with single-path RDMA, MP-RDMA can significantly improve the robustness under failures ( $2\times \sim 4\times $ higher throughput under 0.5%~10% link loss ratio) and improve the overall network utilization by up to 47%. Guo Chen 0001, Yuanwei Lu, Bojie Li, Kun Tan 0002, Yongqiang Xiong, Peng Cheng 0005, Jiansong Zhang 0001, Thomas Moscibroda |
IEEE/ACM Trans. Netw. | 3 |
| 2018 | ST-Accel: A High-Level Programming Platform for Streaming Applications on FPGAabstractIn recent years we have witnessed the emergence of the FPGA in many high-performance systems. This is due to FPGA's high reconfigurability and improved user-friendly programming environment. OpenCL, supported by major FPGA vendors, is a high-level programming platform that liberates hardware developers from having to deal with the complex and error-prone HDL development. While OpenCL exposes a GPU-like programming model, which is well-suited for compute-intensive tasks, in many state-of-art systems that deploy FPGA, we observe that the workloads are streaming-like, which is communication-intensive. This mismatch leads to low throughput and high end-to-end latency. In this paper, we propose ST-Accel, a new high-level programming platform for streaming applications on FPGA. It has the following advantages: (i) ST-Accel adopts the multiprocessing programming model to capture the inherent pipeline-level parallelism of streaming applications while reducing the end-to-end latency. (ii) A message-passing-based host/FPGA communication model is used to avoid the coherency issue of shared memory, thus enabling host/FPGA communication during kernel execution. (iii) ST-Accel provides a high-level abstraction for I/O devices to support direct I/O device access that eliminates the overhead of host CPU and reduces the I/O latency. (iv) ST-Accel enables the decoupled access/execute architecture to maximize the utilization of I/O devices. (v) The host/FPGA communication interface is redesigned to cater to the demands of both latency-critical and throughput-critical scenarios. The experimental results on the Amazon AWS cloud and local machine show that ST-Accel can achieve 1.6X-166X throughput and 1/3 latency for typical streaming workloads when compared to OpenCL. Zhenyuan Ruan, Bojie Li, Peipei Zhou 0001, Jason Cong |
FCCM | 3 |
| 2018 | Joint Optimization of File Placement and Delivery in Cache-Assisted Wireless NetworksabstractIn this paper, the downlink file transmission in one cell with the assistance of cache nodes is studied. Specifically, the base station (BS) reactively delivers files to cache nodes and a requesting user in a multicast manner. Therefore, one file transmission may lead to the cache status update, which further affects the future file transmissions. We consider the joint optimization of file placement and delivery. In particular, we first formulate the optimization of transmission power and time in one finite file lifetime as a Markov Decision Process (MDP) with a random number of stages, where the objective is to minimize the transmission resource at the BS. It is shown that the optimal solution can be obtained via a revised Bellman's equation. Due to the curse of dimensionality, a novel approximation approach is proposed, where the value functions of the Bellman's equation can be calculated from analytical expressions. Hence, iterative algorithms, which appear in the general approximate MDP solutions, can be avoided. Moreover, an bound on the approximation error is also provided. Bojie Li, Rui Wang 0007, Ying Cui 0001, Haisheng Tan |
GLOBECOM | 1 |
| 2018 | Cellular Offloading via Downlink Cache PlacementabstractIn this paper, the downlink file transmission within a finite lifetime is optimized with the assistance of wireless cache nodes. Specifically, the number of requests within the lifetime of one file is modeled as a Poisson point process. The base station multicasts files to downlink users and the selected the cache nodes, so that the cache nodes can help to forward the files in the next file request. Thus we formulate the downlink transmission as a Markov decision process (MDP) with random number of stages, where transmission power and time on each transmission are the control policy. Due to random number of file transmissions, we first proposed a revised Bellman's equation, where the optimal control policy can be derived. In order to address the prohibitively huge state space, we also introduce a low-complexity sub-optimal solution based on a linear approximation of the value function. The approximated value function can be calculated analytically, so that conventional numerical value iteration can be eliminated. Moreover, the gap between the approximated value function and the real value function is bounded analytically. It is shown by simulation that, with the approximated MDP approach, the proposed algorithm can significantly reduce the resource consumption at the base station. Bojie Li, Lexiang Huang, Rui Wang 0007 |
ICC | 1 |
| 2018 | Multi-Path Transport for RDMA in Datacenters
Yuanwei Lu, Guo Chen 0001, Bojie Li, Kun Tan 0002, Yongqiang Xiong, Peng Cheng 0005, Jiansong Zhang 0001, Enhong Chen, Thomas Moscibroda |
NSDI | 3 |
| 2018 | FUSO: Fast Multi-Path Loss Recovery for Data Center Networks
Guo Chen 0001, Yuanwei Lu, Yuan Meng 0002, Bojie Li, Kun Tan 0002, Dan Pei, Peng Cheng 0005, Layong Luo, Yongqiang Xiong, Xiaoliang Wang 0001, Youjian Zhao |
IEEE/ACM Trans. Netw. | 4 |
| 2017 | Memory Efficient Loss Recovery for Hardware-based Transport in DatacenterabstractLimited by the small on-chip memory, hardware-based transport typically implements go-back-N loss recovery mechanism, which costs very few memory but is well-known to perform inferior even under small packet loss ratio. We present MELO, an efficient selective retransmission mechanism for hardware-based transport, which consumes only a constant small memory regardless of the number of concurrent connections. Specifically, MELO employs an architectural separation between data and meta data storage and uses a shared bits pool allocation mechanism to reduce meta data on-chip memory footprint. By only adding in average 23B extra on-chip states for each connection, MELO achieves up to 14.02x throughput while reduces 99% tail FCT by 3.11x compared with go-back-N under certain loss ratio. Yuanwei Lu, Guo Chen 0001, Zhenyuan Ruan, Wencong Xiao, Bojie Li, Jiansong Zhang 0001, Yongqiang Xiong, Peng Cheng 0005, Enhong Chen |
APNet | 5 |
| 2017 | Ultra-high-throughput massive MIMO field-trial over radio computing architecture with peak spectrum efficiency of 79.82 bps/HzabstractMassive multiple-input multiple-output (MIMO) has been considered as one of the key technologies in 5G communication, owing to its promising potential to increase the spectrum efficiency significantly. The capacity grows rapidly with the increasing number of antennas. However, in real-time environment, the more antennas are employed, the more computing resources are required, and this increase is even in a nonlinear progression. In this work, radio computing architecture (RCA) with multiple parallel general purpose processors (GPPs) and automatic distribution system (ADS) are utilized to overcome the strict real-time constraints and establish a highly flexible prototype platform for massive MIMO research. The feasibility and performance of the GPP-based prototype has been investigated through field trial. A collocated 64-RF-channel base station (BS) with each RF channel driving 3 antennas can simultaneously serve twelve 8-antenna user equipments (UEs). The maximum overall bandwidth is 200 MHz, and the maximum flow number is pre-defined as 24. In this field trial, a peak total user throughput of 11.29 Gbps and a cell spectrum efficiency of 79.82 bps/Hz are achieved. As far as we know, this is the world's highest throughput ever achieved in the sub 6 GHz band. Furthermore, apart from multi-user MIMO, single-user MIMO is also tested in this field trial. Wenliang Liang, Yuanquan Wang 0003, Bojie Li, Jie Sheng, Yuchao Han, Haihua Shen, Liang Gu, Yuya Saito, Anass Benjebbour, Yoshihisa Kishiyama, Xin Wang 0073, Xiaolin Hou, Huiling Jiang |
PIMRC | 3 |
| 2017 | Field trial on TDD massive MIMO system with polar codeabstractA large scale field trial is conducted to investigate the integrated performance of massive multiple-input multiple-output (MIMO) and polar code, which are two key technologies for 5th generation (5G) mobile system. A more simple and effective algorithm named polarization weight (PW) is used in polar construction. Based on uplink and downlink channel reciprocity in the time-division duplex (TDD) mode, the practical performance of single user (SU) massive MIMO prototype with 64 independent RF channels and 200 MHz bandwidth is investigated under different user equipment (UE) deployments, different layer numbers and moving speeds. Compared to massive MIMO with turbo code, it is shown that significant performance gains can be obtained in TDD massive MIMO system with polar code. For 3 layers scheduled in massive MIMO system with polar code, the maximum user throughput of 1.64 Gbps and spectrum efficiency of 11.64 bit/s/Hz can be achieved. Based on these observations, the feasibility and performance of TDD massive MIMO system with polar code are verified. Wenliang Liang, Bojie Li, Liang Gu, Jie Sheng, Pengcheng Qiu, Jian Wang 0001, Yuanquan Wang 0003 |
PIMRC | 3 |
| 2017 | KV-Direct: High-Performance In-Memory Key-Value Store with Programmable NICabstractPerformance of in-memory key-value store (KVS) continues to be of great importance as modern KVS goes beyond the traditional object-caching workload and becomes a key infrastructure to support distributed main-memory computation in data centers. Recent years have witnessed a rapid increase of network bandwidth in data centers, shifting the bottleneck of most KVS from the network to the CPU. RDMA-capable NIC partly alleviates the problem, but the primitives provided by RDMA abstraction are rather limited. Meanwhile, programmable NICs become available in data centers, enabling in-network processing. In this paper, we present KV-Direct, a high performance KVS that leverages programmable NIC to extend RDMA primitives and enable remote direct key-value access to the main host memory. Bojie Li, Zhenyuan Ruan, Wencong Xiao, Yuanwei Lu, Yongqiang Xiong, Andrew Putnam, Enhong Chen |
SOSP | 1 |
| 2017 | Trial Results of 5G New Air Interface TechnologiesabstractThis paper presents field trial results of innovative air interface technologies: sparse code multiple access (SCMA), filtered orthogonal frequency division multiplexing (f-OFDM), and massive multiple-input multiple-output (MIMO). To confirm the improvement of the promising technologies, the counterparts in long term evolution (LTE) are selected as the comparison baseline. For each technology, detailed parameter configurations and deployment scenarios are addressed at first, then precise test procedures are described, finally, extensive analysis are provided to testify the high spectral efficiency, flexibility and system capacity introduced by the new technologies. According to the trial results, SCMA can achieve 3 times connection number and throughput than SC-FDMA in uplink, and 90% throughput gains over OFDMA in downlink. Waveform f-OFDM is validated to support flexible resources allocation thus avoiding asynchronous transmission in both uplink and downlink. The large number of antennas in Massive MIMO accommodates more user and offers around 10 times throughput than single user scenario in downlink. Hongyi Zhao, Dageng Chen, Bijun Zhang, Liang Gu, Bojie Li, Huilian Yang, Tingjian Tian, Zhendong Luo, Kejun Wei |
VTC Spring | 9 |
| 2017 | Field Trial Investigation of Wired and Wireless Calibration Schemes for Real-Time Massive MIMO PrototypeabstractRadio frequency (RF) channel reciprocity calibration is one of the most important technologies in massive multiple-input multiple-output (MIMO) system. With the increasing of the antenna size, the calibration procedure becomes much more challenging. In this paper, wired- and wireless-calibration schemes have been proposed and investigated through field-trial test based on a real-time massive MIMO prototype with 64 independent RF channels. The prototype is capable of supporting 12 200-MHz user equipments (UEs) simultaneously. The field trial test results validate both schemes can significantly improve the overall system throughput. For single-user (SU) MIMO scenario, the throughput can be enhanced from 0.71 Gbps to 1.46 Gbps and 1.54 Gbps with wireless- and wired-calibration schemes, respectively. For multi-user (MU) MIMO scenario, the corresponding enhancement can be from 0.26 Gbps to 9.18 Gbps and 11.29 Gbps, respectively. In terms of calibration accuracy, the wired-calibration scheme has better performance, but the wireless one has lower system complexity and cost. To the best of our knowledge, this is the first field trial investigation of wireless-calibration scheme and comparison with wired calibration in a real-time high-throughput massive MIMO prototype. Wenliang Liang, Yuanquan Wang 0003, Keyu Song, Bojie Li, Haihua Shen, Shenfei Zhang, Liang Gu, Yuuya Saito, Anass Benjebbour, Yoshihisa Kishiyama |
VTC Fall | 4 |
| 2017 | Large Scale Field Experimental Trial of Downlink TDD Massive MIMO at the 4.5 GHz BandabstractThe 5th generation of mobile communications system (5G) is nowadays expected to support very diverse applications, devices, and services such as enhanced mobile broadband (eMBB) and the Internet of things (IoTs). In order to meet the future requirements, new 5G candidate radio access technologies have been proposed and studied. For example, Massive multiple-input multiple-output (MIMO) is an attractive technology to further improve the spectrum efficiency and to boost system capacity. This paper presents a large scale field experimental trial of downlink time division duplex (TDD) Massive MIMO using the developed 5G test bed at the 4.5 GHz band. Specifically, the system performance of multi-user (MU) Massive MIMO with different numbers of sets of user equipment (UEs) and spatial layers is investigated, including the impact of different inter-UE separations and UE deployments, which have not been well investigated compared to the previous studies. The experimental results show that the throughput performance for a wide UE separation linearly increases as the number of layers and UEs increases, while the performance improvement is diminished when inter-UE separation narrows. In addition, the maximum total user throughput of 11.29 Gbps, which corresponds to a spectrum efficiency of 79.82 bps/Hz/cell, was achieved with 24 layers. The results indicate that TDD Massive MIMO using the developed 5G test bed greatly improves the spectrum efficiency for large scale MU-MIMO. Yuya Saito, Anass Benjebbour, Yoshihisa Kishiyama, Xin Wang 0073, Xiaolin Hou, Huiling Jiang, Wenliang Liang, Bojie Li, Liang Gu, Tsuyoshi Kashima |
VTC Spring | 9 |
| 2017 | Introduction on IMT-2020 5G Trials in ChinaabstractUltrahigh data rate, massive connectivity, ultralow latency, and high reliability, as well as ultraflexible air interface design to support diversified usage scenarios, are the major targets of 5G radio access network design. To meet these targets, advanced new radio transmission technologies, including new waveform, new channel coding, non-orthogonal multiple access, as well as massive antenna techniques, have been proposed and studied worldwide in both academic and industries. Most of the gains claimed in the literature are from simulations or by small scale lab testing. Instead, this paper elaborates on the first hand results and analysis of the large-scale field trials carried out in China for a selection of key 5G technology components as well as some of their combinations, including filtered-orthogonal frequency division multiplexing (OFDM) (f-OFDM), polar code, sparse code multiple access, and massive multi-input multi-output (MIMO) for both uplink and downlink cellular networks. Testing results verify the feasibility and benefit of each technology contributing toward the diversified 5G targets, as well as the feasibility for these technologies to work jointly for higher spectrum efficiency, larger connectivity, and lower latency. Hongyi Zhao, Yan Chen 0010, Dageng Chen, Bijun Zhang, Liang Gu, Bojie Li, Huilian Yang, Tingjian Tian, Zhendong Luo, Kejun Wei |
IEEE J. Sel. Areas Commun. | 10 |
| 2016 | ClickNP: Highly flexible and High-performance Network Processing with Reconfigurable HardwareabstractHighly flexible software network functions (NFs) are crucial components to enable multi-tenancy in the clouds. However, software packet processing on a commodity server has limited capacity and induces high latency. While software NFs could scale out using more servers, doing so adds significant cost. This paper focuses on accelerating NFs with programmable hardware, i.e., FPGA, which is now a mature technology and inexpensive for datacenters. However, FPGA is predominately programmed using low-level hardware description languages (HDLs), which are hard to code and difficult to debug. More importantly, HDLs are almost inaccessible for most software programmers. This paper presents ClickNP, a FPGA-accelerated platform for highly flexible and high-performance NFs with commodity servers. ClickNP is highly flexible as it is completely programmable using high-level C-like languages, and exposes a modular programming abstraction that resembles Click Modular Router. ClickNP is also high performance. Our prototype NFs show that they can process traffic at up to 200 million packets per second with ultra-low latency ($< 2\mu$s). Compared to existing software counterparts, with FPGA, ClickNP improves throughput by 10x, while reducing latency by 10x. To the best of our knowledge, ClickNP is the first FPGA-accelerated platform for NFs, written completely in high-level language and achieving 40 Gbps line rate at any packet size. Bojie Li, Kun Tan 0002, Layong Luo, Yanqing Peng, Renqian Luo, Ningyi Xu, Yongqiang Xiong, Peng Cheng 0005 |
SIGCOMM | 1 |
| 2016 | Fast and Cautious: Leveraging Multi-path Diversity for Transport Loss Recovery in Data Centers
Guo Chen 0001, Yuanwei Lu, Yuan Meng 0002, Bojie Li, Kun Tan 0002, Dan Pei, Peng Cheng 0005, Layong Luo, Yongqiang Xiong, Xiaoliang Wang 0001, Youjian Zhao |
USENIX ATC | 4 |
| 2012 | A multimedia service migration protocol for single user multiple devicesabstractThis paper describes a new protocol SMP, which supports multimedia transfer for single-user, multiple-device scenarios. Through its novel naming and control/data plane designs, SMP is able to retain the current client and server protocol operations while placing new functions at the proxy. Our initial evaluation has confirmed its viability. Chi-Yu Li 0001, Ioannis Pefkianakis, Bojie Li, Chenghui Peng, Songwu Lu |
ICC | 3 |