VLDB 2026 Research / reviewers in the wild / expert
Bingyang Wu
dblp:06/8847
· DBLP profile ↗
19ranked-venue papers
9as first author
9since 2021 · last 2026
0000-0001-7221-8007ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 10 · 4 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FastServe: Iteration-Level Preemptive Scheduling for Large Language Model Inference
Bingyang Wu, Yinmin Zhong, Fangyue Liu, Yuanhang Sun, Gang Huang 0001, Xuanzhe Liu, Xin Jin 0008 |
NSDI | 1 |
| 2026 | Epiphron: Resource-Efficient Distributed Key-Value StorageabstractIn-memory key-value storage necessitates a substantial quantity of computation and storage resources for both performance and scalability, thereby diminishing the resources available for user applications. The emergence of programmable network hardware, including SmartNICs and programmable switches, provides the opportunity to offload operations from server CPUs. We present Epiphron, a novel distributed in-memory key-value store architecture that co-designs with off-path SmartNICs and programmable switches. Facing the limited performance of off-path SmartNICs, Epiphron successfully achieves high resource efficiency while keeping load balancing and fault tolerance by$(i)$hybridizing erasure coding with replication in storage management,$(ii)$accelerating read operations with a new data plane design (conflict detection and RDMA-compatible forwarding) on programmable switches,$(iii)$employing a network protocol extended from one-sided RDMA. We evaluate Epiphron on Barefoot Tofino switches, NVIDIA BlueField-2 SmartNICs, and commodity servers. The experimental results demonstrate that compared to existing solutions, Epiphron improves throughput by up to 2.2× and consumes 47% less memory while completely bypassing server CPUs. Ruidong Zhu, Bingyang Wu, Xin Yao 0008, Renhai Chen, Gong Zhang 0001, Xuanzhe Liu, Xin Jin 0008 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | Optimizing RLHF Training for Large Language Models with Stage Fusion
Yinmin Zhong, Bingyang Wu, Changyi Wan, Hanpeng Hu, Ranchen Ming, Yibo Zhu 0001, Xin Jin 0008 |
NSDI | 3 |
| 2024 | dLoRA: Dynamically Orchestrating Requests and Adapters for LoRA LLM Serving
Bingyang Wu, Ruidong Zhu, Peng Sun 0006, Xuanzhe Liu, Xin Jin 0008 |
OSDI | 1 |
| 2024 | LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence ParallelismabstractThe context window of large language models (LLMs) is rapidly increasing, leading to a huge variance in resource usage between different requests as well as between different phases of the same request. Restricted by static parallelism strategies, existing LLM serving systems cannot efficiently utilize the underlying resources to serve variable-length requests in different phases. To address this problem, we propose a new parallelism paradigm, elastic sequence parallelism (ESP), to elastically adapt to the variance across different requests and phases. Based on ESP, we design and build LoongServe, an LLM serving system that (1) improves computation efficiency by elastically adjusting the degree of parallelism in real-time, (2) improves communication efficiency by reducing key-value cache migration overhead and overlapping partial decoding communication with computation, and (3) improves GPU memory efficiency by reducing key-value cache fragmentation across instances. Our evaluation under diverse real-world datasets shows that LoongServe improves the throughput by up to 3.85× compared to chunked prefill and 5.81× compared to prefill-decoding disaggregation. Bingyang Wu, Yinmin Zhong, Peng Sun 0006, Xuanzhe Liu, Xin Jin 0008 |
SOSP | 1 |
| 2023 | Transparent GPU Sharing in Container Clouds for Deep Learning Workloads
Bingyang Wu, Zhihao Bai, Xuanzhe Liu, Xin Jin 0008 |
NSDI | 1 |
| 2023 | XRON: A Hybrid Elastic Cloud Overlay Network for Video Conferencing at Planetary ScaleabstractQuality and cost are two key considerations for video conferencing services. Service providers face a dilemma when selecting network tiers to build their infrastructure---relying on Internet links has poor quality, while using premium links brings excessive cost. Bingyang Wu, Kun Qian 0021, Bo Li 0061, Dennis Cai, Ennan Zhai, Xuanzhe Liu, Xin Jin 0008 |
SIGCOMM | 1 |
| 2022 | AMOS: enabling automatic mapping for tensor computations on spatial accelerators with hardware abstractionabstractHardware specialization is a promising trend to sustain performance growth. Spatial hardware accelerators that employ specialized and hierarchical computation and memory resources have recently shown high performance gains for tensor applications such as deep learning, scientific computing, and data mining. To harness the power of these hardware accelerators, programmers have to use specialized instructions with certain hardware constraints. However, these hardware accelerators and instructions are quite new and there is a lack of understanding of the hardware abstraction, performance optimization space, and automatic methodologies to explore the space. Existing compilers use hand-tuned computation implementations and optimization templates, resulting in sub-optimal performance and heavy development costs. Size Zheng 0001, Renze Chen, Anjiang Wei, Yicheng Jin, Qin Han, Liqiang Lu, Bingyang Wu, Shengen Yan, Yun Liang 0001 |
ISCA | 7 |
| 2022 | NeoFlow: A Flexible Framework for Enabling Efficient Compilation for High Performance DNN TrainingabstractDeep neural networks (DNNs) are increasingly deployed in various image recognition and natural language processing applications. The continuous demand for accuracy and high performance has led to innovations in DNN design and a proliferation of new operators. However, existing DNN training frameworks such as PyTorch and TensorFlow only support a limited range of operators and rely on hand-optimized libraries to provide efficient implementations for these operators. To evaluate novel neural networks with new operators, the programmers have to either replace the holistic new operators with existing operators or provide low-level implementations manually. Therefore, a critical requirement for DNN training frameworks is to provide high-performance implementations for the neural networks containing new operators automatically in the absence of efficient library support. In this paper, we introduce NeoFlow, which is a flexible framework for enabling efficient compilation for high-performance DNN training. NeoFlow allows the programmers to directly write customized expressions as new operators to be mapped to graph representation and low-level implementations automatically, providing both high programming productivity and high performance. First, NeoFlow provides expression-based automatic differentiation to support customized model definitions with new operators. Then, NeoFlow proposes an efficient compilation system that partitions the neural network graph into subgraphs, explores optimized schedules, and generates high-performance libraries for subgraphs automatically. Finally, NeoFlow develops an efficient runtime system to combine the compilation and training as a whole by overlapping their execution. In the experiments, we examine the numerical accuracy and performance of NeoFlow. The results show that NeoFlow can achieve similar or even better performance at the operator and whole graph level for DNNs compared to deep learning frameworks. Especially, for novel networks training, the geometric mean speedups of NeoFlow to PyTorch, TensorFlow, and CuDNN are 3.16X, 2.43X, and 1.92X, respectively. Size Zheng 0001, Renze Chen, Yicheng Jin, Anjiang Wei, Bingyang Wu, Shengen Yan, Yun Liang 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2018 | Performance Analysis of OMP-Based Channel Estimations in Mobile OFDM SystemsabstractTo analyze the performance of orthogonal matching pursuit (OMP)-based compressed channel estimation (CCE) with deterministic pilot patterns, we propose a mathematical framework by defining four normalized mean square errors (NMSEs): the total NMSE (NMSET), the NMSE on dominant channel components (NMSED), the NMSE caused by “lost errors” (NMSEL), and the NMSE caused by “false alarms” (NMSEF). Then, we derive a formula with a closed form for evaluating the upper bound of NMSED in the ideal case (NMSED,UB). Using the proposed analytical framework, the main findings include: 1) the NMSED,UB is determined by the following four parameters: the deterministic pilot pattern, the maximum Doppler shift, the number of dominant multipath components, and the SNR; 2) the NMSED,UB can be viewed as an approximation of practical NMSET in the case that the probability of the successes of OMP exceeds a certain threshold, in which both NMSEL and NMSEF are neglectable; and 3) using linear regression models, the practical bit-error-rate performance also can be predicted well based on the proposed NMSED,UB. We believe that the proposed framework provides a useful tool for adaptively optimizing pilot parameters according to rapidly time-varying channel conditions when using OMP-based CCEs in mobile OFDM systems. Guoping Tan, Bingyang Wu, Thorsten Herfet |
IEEE Trans. Wirel. Commun. | 2 |
| 2017 | Correlation-driven optimized Taylor expansion precoding for massive MIMO systems with correlated channelsabstractHardware-efficient low-complexity precoding is very important in the downlink of Massive MIMO systems for mitigating interference and optimizing performance. In this paper, we propose a correlation-driven optimized Taylor expansion (CD-OTE) precoding scheme to simplify linear minimum mean square error (MMSE) precoding. In order to simplify the hardware-expensive matrix inversion involved in the linear MMSE pre-coder, a Taylor expansion with optimized polynomial coefficients and selection of the most relevant correlation coefficients is proposed. We take into consideration the correlation between different users' channels and develop a general design criterion. Both convergence and complexity analyses are carried out. Simulation results show that the proposed CD-OTE precoder is significantly better than previously reported techniques, while requiring a similar cost. Wence Zhang, Rodrigo C. de Lamare, Cunhua Pan, Ming Chen 0001, Bingyang Wu, Xu Bao 0001 |
ICC | 5 |
| 2017 | Widely Linear Precoding for Large-Scale MIMO with IQI: Algorithms and Performance AnalysisabstractIn this paper, we study widely linear precoding techniques to mitigate in-phase/quadrature-phase (IQ) imbalance (IQI) in the downlink of large-scale multiple-input multiple-output (MIMO) systems. We adopt a real-valued signal model, which considers the IQI at the transmitter, and then develop widely linear zero-forcing (WL-ZF), widely linear matched filter, widely linear minimum mean-squared error, and widely linear block-diagonalization (WL-BD) type precoding algorithms for both single- and multiple-antenna users. We also present a performance analysis of WL-ZF and WL-BD. It is proved that without IQI, WL-ZF has exactly the same multiplexing gain and power offset as ZF, while when IQI exists, WL-ZF achieves the same multiplexing gain as ZF with ideal IQ branches, but with a minor power loss, which is related to the system scale and the IQ parameters. We also compare the performance of WL-BD with BD. The analysis shows that with ideal IQ branches, WL-BD has the same data rate as BD, while when IQI exists, WL-BD achieves the same multiplexing gain as BD without IQ imbalance. Numerical results verify the analysis and show that the proposed widely linear type precoding methods significantly outperform their conventional counterparts with IQI and approach those with ideal IQ branches. Wence Zhang, Rodrigo C. de Lamare, Cunhua Pan, Ming Chen 0001, Jianxin Dai, Bingyang Wu, Xu Bao 0001 |
IEEE Trans. Wirel. Commun. | 6 |
| 2016 | Power Control in D2D Underlay Massive MIMO Systems with Pilot ReuseabstractThis paper studies pilot reuse and data transmit power control in a D2D underlay massive MIMO system over fading channels. In order to reduce the length of pilots, we propose to reuse a set of orthogonal pilots among DUEs, and the graph coloring based pilot allocation (GCPA) algorithm is utilized to allocate pilots to DUEs. Linear minimum mean square error (LMMSE) filters are used for signal detection. We then derive the lower bound of D2D links' average signal-to-interference-plus-noise ratio (SINR), and formulate a power control problem to minimize D2D links' data transmit power under the target SINR constraints. An iterative method converging to the unique optimal solution is proposed. Simulation results show that the analytical lower bound of the average SINR closely matches the simulated average SINR. What's more, pilot resources can be saved greatly by pilot reuse, and the effect of pilot contamination to the system can be almost neglected by allocating proper number of pilots to DUEs and applying GCPA algorithm. Hao Xu 0003, Zhaohui Yang 0001, Bingyang Wu, Jianfeng Shi 0001, Ming Chen 0001 |
VTC Spring | 3 |
| 2014 | Pricing-based distributed power control for weighted sum energy-efficiency maximization in ad hoc networksabstractWe consider the problem of maximizing the weighted sum energy efficiency (WS-EE) in ad hoc networks. To solve this problem in a distributed manner, one novel distributed adaptive-pricing algorithm is developed based on limited information exchange among the nodes. Specifically, each node updates its current interference information and broadcasts it to the other nodes. Having collected all this information, each node can adjust its transmit power accordingly with simple arithmetical operations. Then iterate these two steps. This algorithm is strictly proven to be convergent and can attain the KKT optimality conditions of the problem. Moreover, an alternative centralized algorithm based on gradient projection method is proposed to serve as the performance benchmark. Simulation results show that the proposed distributed algorithm converges rapidly. Furthermore, this distributed algorithm performs as well as the centralized one and significantly outperforms the existing algorithm in terms of the WS-EE. Cunhua Pan, Bingyang Wu, Nuo Huang, Hong Ren, Ming Chen 0001 |
GLOBECOM | 2 |
| 2011 | Comments on "Partial Channel Feedback Schemes Maximizing Overall Efficiency in Wireless Networks"abstractThe proof of Proposition 2 in the above-mentioned paper was incorrect due to some flaws in the prerequisite analysis of the proposed hybrid feedback scheme. We fix the flaws and correct the proof in this comment. From the corrected proof, we arrive at a new observation that the hybrid feedback scheme has exactly the same feedback overhead as the pure opportunistic feedback scheme if the threshold is adaptively set for a target utilization. Simulation results demonstrate the correctness of our analysis. Chunjuan Diao, Ming Chen 0001, Bingyang Wu |
IEEE Trans. Wirel. Commun. | 3 |
| 2005 | Iterative channel estimation and signal detection in clipped OFDMabstractHigh peak-to-average power ratio (PAPR) is an inherit drawback of orthogonal frequency-division multiplexing (OFDM) systems, and it limits the use of OFDM technology. In all the PAPR reduction methods, digital clipping is probably the simplest one and it is very effective in the solution of the PAPR problem in OFDM systems. However, clipping introduces additional distortion, which degrades the performance of channel estimation and signal detection. In this paper, the performance of channel estimation using the simple least square (LS) method in OFDM systems with a clipping operation is analyzed, a decision-aided channel estimation method to mitigate the clipping distortion is developed, and the potential gain in channel estimation of the decision-aided method is derived theoretically. Simulation results show that, in clipped OFDM, the theoretical potential gain in channel estimation can be obtained by a few iterations of an iterative decision-aided channel estimation and signal detection process, and, simultaneously, the performance of signal detection is effectively improved Bingyang Wu, Shixin Cheng, Haifeng Wang 0002 |
GLOBECOM | 1 |
| 2005 | Trellis factor search PTS for PAPR reduction in OFDMabstractOFDM symbols exhibit a property of high peak-to-average power ratios (PAPR), which degrades system performance. Partial transmit sequence (PTS) is a kind of methods to depress the PAPR effectively. To obtain good trade-offs between complexity and performance on PAPR reduction using PTS schemes, a trellis structure based PTS factor search method is proposed and investigated in this paper. The trellis search is with a variant constraint length L/sub C/, 1/spl les/L/sub C//spl les/V-1, where V is the number of PTS sub-blocks. The method is to decide a PTS factor by searching all possible paths obtained by varying L/sub C/ consecutive factors. Using different constraint length, trellis factor search PTS exhibits different PAPR reduction performance with different computation burden. The trellis search can be viewed as a general PTS factor search model including the full search and the iterative search appeared in previous papers. Bingyang Wu, Shixin Cheng, Haifeng Wang 0002 |
PIMRC | 1 |
| 2004 | SER evaluation of OFDM with Nyquist rate and oversampled clipping using estimated channel informationabstractClipping is a popular method to overcome large peak-to-mean envelope power ratios (PMEPRs) in orthogonal frequency-division multiplexing (OFDM) systems, but it results in nonlinear distortion, which degrades signal detection performance. In this paper, clipping effects on signal detection with a simple channel estimation in OFDM are investigated, signal-to-noise and distortion ratio (SNDR) results of detected signals are derived, and symbol error rates (SERs) of clipped OFDM in frequency selective Rayleigh fading channels are evaluated. The evaluated SER results are supported by simulations. Bingyang Wu, Shixin Cheng, Haifeng Wang 0002, Ming Chen 0001 |
PIMRC | 1 |
| 2003 | Clipping effects on channel estimation and signal detection in OFDMabstractClipping is a popular and simple method to overcome large peak to mean envelope power ratio (PMEPR) in OKDM system, but it results in a nonlinear distortion, which degrade the performance of channel estimation and signal detection. In this paper we estimated OFDM channel using a simple LS method based on a scheme that only part of the OFDM tones arc used for pilot, in which the clipping distortion is unknown for the receiver, and then we analyze the clipping effects on channel estimation. According to the estimated channel, we detected the OFDM signal and analyze the clipping effects on signal detection. We derived the BER solution of clipping OFDM in Rayleigh fading channel and the numerical solution is very close to simulation results. Bingyang Wu, Shixin Cheng, Haifeng Wang 0002 |
PIMRC | 1 |