Wenquan Xu

dblp:64/3114 · DBLP profile ↗
← Back
23ranked-venue papers
10as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 17 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 4 · 4 first-author · 2 since 2021Systems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Turbo: Efficiently Serving Long-Context Large Language Models with In-Network Aggregation
abstract
LLM supporting long contexts faces a critical memory bottleneck due to the linear growth of KV cache. Distributing the storage across multiple GPUs alleviates this burden but introduces significant communication overhead or traffic incast, especially during the decoding phase. We propose Turbo, a first-of-its-kind in-network aggregation system that accelerates long-context inference by offloading query broadcast and attention aggregation to switches. We address three key challenges to map complex attention mechanisms onto restricted switch hardware: (i) To bypass the switch's inability to buffer global states or perform complex operations, we devise online table-based aggregation, which decomposes global reduction into pairwise operations and approximates nonlinear functions via lookup tables. (ii) To circumvent the restriction on retroactive state access in RMT pipelines, we introduce a rolling forward scheme that propagates states to enable cross-stage updates. (iii) To mitigate aggregation stragglers caused by topology-induced load imbalance, we construct a load-aware aggregation tree that optimizes workload distribution. Evaluations on a Tofino2-based testbed show that Turbo reduces end-to-end inference latency by up to 37%. Large-scale simulations on NS-3 demonstrate that Turbo significantly outperforms state-of-the-art baselines in both inference latency and network traffic reduction with negligible accuracy loss.
Ying Wan 0001, Yuchen Xu 0003, Chuwen Zhang, Yingsheng Huang, Wenquan Xu, Jialin Li 0001, Mingwei Xu 0001, Wenfei Wu, Congcong Miao
SIGCOMM6
2025 DIR-EC: intra-region enhancement and inter-region collaboration network for facial expression recognition
Fuji Ren, Wenquan Xu
Vis. Comput.6
2024 OptimusPrime: Unleash Dataplane Programmability through a Transformable Architecture
abstract
Network dataplane calls for better programmability. Current programmable network processing chips are based on either pipeline or multi-core Run-To-Completion (RTC) architecture with various trade-offs in flexibility, performance, and cost. The existing attempts to amalgamate the strengths of the two are stilted and inflexible. In this paper, we challenge the status quo by introducing a more fluid and organic programmable chip architecture, OptimusPrime, built from identical hardware blocks. Unlike the conventional static hybrid architecture, OptimusPrime allows each block to be transformed into either a pipeline stage processor or a multi-core RTC processor through software-defined configuration, enabling versatile data plane programming tailored to a wide range of applications (e.g., stateful packet processing and in-network computing). We integrate the C and P4 languages for application programming and develop algorithms to map a user program to the optimal distribution of pipeline stages and RTC cores. We demonstrate the viability of OptimusPrime through practical use cases such as in-network aggregation, in-network caching, and network function integration. We developed an FPGA-based prototype and a software-based ASIC simulator to validate the feasibility of OptimusPrime, which can be used by switches and smartNICs to enhance their programmability to a new level with high performance and low cost.
Zhikang Chen, Haoyu Song 0001, Hanyi Zhou, Tong Yun, Wenquan Xu, Tian Pan 0001, Bin Liu 0001
SIGCOMM7
2024 Stability Analysis of Deep Belief Network: Based SD-AR Model for Nonlinear Time Series
abstract
Abstract As for nonlinear time series prediction, many different kinds of varying-coefficient models have been proposed and analysised in recent years. A kind of varying functional-coefficient autoregressive model, called the deep belief network-based state-dependent autoregressive (DBN-AR) model is considered in this paper. The stability conditions and existing conditions of limit cycle of the DBN-AR model are also studied. An especial designed parameter estimation method is used to identify the DBN-AR model. The DBN-AR model is used to predict the famous Canadian lynx data and Henon chaotic series, the prediction capability of the DBN-AR model is compared with other prediction models, the experimental results show that the DBN-AR model obtains better prediction accuracy.
Wenquan Xu
Neural Process. Lett.1
2024 OBMA: Scalable Route Lookups With Fast and Zero-Interrupt Updates
abstract
Software-based IP route lookup is a key component for packet forwarding in Software Defined Networks. Running lookup algorithms on commodity CPUs is flexible and scalable, which shows advantages on cost and power consumption over the hardware-based forwarding engines. However, dynamic network functions and services make route updates more frequent than ever. Existing algorithms often fall short of the incremental update requirements. In this paper, we propose the Overlay BitMap Algorithm (OBMA), which contains several variations, to support extraordinary update performance while maintaining the highest-in-class lookup speed and storage efficiency. Starting from the basic OBMA_B, we develop two variations with different tradeoffs for different application scenarios. OBMA_L supports faster lookups than OBMA_B at a small cost of update speed. OBMA_S achieves better storage efficiency than OBMA_B at a small cost of lookup throughput. We run our algorithms on a commodity CPU and evaluate them with real-world route tables and traces. The experiments show that OBMA achieves the lowest memory footprint, the highest update speed, and over 200 Mpps lookup throughput. Specifically, OBMA_S reduces the memory footprint to 3.98 bytes/prefix which is 25.33% smaller that of the state-of-the-art Poptrie; OBMA_L supports 252.02 Mpps lookup throughput with a single thread, and more than 600 Mpps with multiple parallel threads in a single CPU, significantly outperforming the state-of-the-art Poptrie and SAIL; OBMA_B supports updates at a rate of 14.58M updates/s which is 15 times faster than Poptrie. The tests show that the update process has little interference with the lookup process for OBMA, and achieves zero-interrupt to lookups with multiple threads.
Chuwen Zhang, Haoyu Song 0001, Ying Wan 0003, Wenquan Xu, Bin Liu 0001
IEEE/ACM Trans. Netw.5
2023 ISAC: In-Switch Approximate Cache for IoT Object Detection and Recognition
Wenquan Xu, Haoyu Song 0001, Bin Liu 0001
INFOCOM1
2023 ClickINC: In-network Computing as a Service in Heterogeneous Programmable Data-center Networks
abstract
In-Network Computing (INC) has found many applications for performance boosts or cost reduction. However, given heterogeneous devices, diverse applications, and multi-path network typologies, it is cumbersome and error-prone for application developers to effectively utilize the available network resources and gain predictable benefits without impeding normal network functions. Previous work is oriented to network operators more than application developers. We develop ClickINC to streamline the INC programming and deployment using a unified and automated workflow. Click-INC provides INC developers a modular programming abstractions, without concerning to the states of the devices and the network topology. We describe the ClickINC framework, model, language, workflow, and corresponding algorithms. Experiments on both an emulator and a prototype system demonstrate its feasibility and benefits.
Wenquan Xu, Haoyu Song 0001, Zhikang Chen, Wenfei Wu, Guyue Liu, Yinchao Zhang, Zerui Tian, Bin Liu 0001
SIGCOMM1
2022 TSN-Peeper: an Efficient Traffic Monitor in Time-Sensitive Networking
abstract
Time-Sensitive Networking (TSN) is proposed in recent years to satisfy the strict performance requirements of time-sensitive traffic in a growing number of emerging applications. Even though several traffic scheduling algorithms have been standardized for TSN to pursue this goal, time-sensitive flows may not be forwarded as planned and thus fail to achieve the expected performance in real networks. The fundamental cause lies in the fact that static offline planning cannot adapt to the intrinsic dynamic factors in TSN (e.g., time-synchronization error) at runtime. Hence, next-generation TSN will benefit from a closed-loop design where a performance monitoring system provides feedback of real-time packet-forwarding information. In our research, TSN-Peeper, a light-weight, fast-response and full-coverage TSN performance monitoring system, is designed and evaluated. This paper describes its architecture design and data collection mechanisms that enable timely identification and collection of packet-forwarding misbehavior at low-cost in TSN. TSN-Peeper offloads the misbehavior identification in the switch to relieve the burden on the controller and network bandwidth. To reduce the interruption frequency to the controller, it uses probe packets to collect misbehavior information in aggregation with optimized path planning. To realize controllable reporting delays, it optimizes the sending moments of probe packets according to the flow settings. Experimental results verify that TSN-Peeper offers fast response with low cost while providing full coverage and being scalable.
Chuwen Zhang, Zerui Tian, Liang Cheng 0001, Yuxi Liu 0017, Ying Wan 0001, Wenquan Xu, Tian Pan 0001, Yang Xu 0010, Yi Wang 0004, Hailong Zhu, Bin Liu 0001
ICNP9
2022 Enabling In-situ Programmability in Network Data Plane: From Architecture to Language
Zhikang Chen, Haoyu Song 0001, Wenquan Xu, Tong Yun, Bin Liu 0001
NSDI4
2022 Effective capacity optimization for UDN
abstract
Abstract For the ultra‐dense distributed network (UDN), the intensive deployment leads to the collision problem of the unlicensed spectrum access due to the random contention access mechanism of the unlicensed spectrum, which is difficult to improve the user quality of service (QoS). Here, the wireless access process of the licensed assisted access (LAA) network is modelled as a four‐state semi‐Markov model. Then, the effective capacity (EC) of the UDN based on the random access is derived, and the network capacity limit under the QoS requirements in the unlicensed spectrum is given. Based on the obtained EC, two power control algorithms are proposed to maximize the EC and effective energy efficiency (EEE) respectively under the user QoS requirements. Finally, numerical results verify the accuracy of the proposed theory, and the capacity‐delay domain of the UDN is given. The simulation results show that the proposed algorithms improve the EC and the EEE compared with existing classical schemes.
Shaoguo Xie, Liefu Ai, Wenquan Xu
IET Commun.3
2022 A Hybrid Modeling Method Based on Linear AR and Nonlinear DBN-AR Model for Time Series Forecasting
Wenquan Xu, Hui Peng 0001, Xiaoyong Zeng, Feng Zhou 0005, Xiaoying Tian
Neural Process. Lett.1
2021 In-situ Programmable Switching using rP4: Towards Runtime Data Plane Programmability
abstract
The existing chip architecture and programming language are incapable of supporting in-service updates by loading or offloading on-demand protocols and functions at runtime. We examine the fundamental reasons for the inflexibility and design a new In-situ Programmable Switch Architecture (IPSA) as a fix. We further design rP4, a P4 extension, for programming IPSA-based devices. To manifest the in-situ programming feasibility, we develop an rP4 compiler and demonstrate several use cases on both a software switch, ipbm, and an FPGA-based prototype. Our preliminary experiments and analysis show that, compared to PISA, IPSA provides higher flexibility in enabling runtime functional update with limited performance and gate-count penalty. The in-situ programming capability enabled by IPSA and rP4 opens a promising design space for programmable networks.
Haoyu Song 0001, Zhikang Chen, Wenquan Xu, Bin Liu 0001
HotNets5
2021 SODA: Similar 3D Object Detection Accelerator at Network Edge for Autonomous Driving
abstract
Offloading the 3D object detection from autonomous vehicles to MEC is appealing because of the gains on quality, latency, and energy. However, detection requests lead to repetitive computations since the multitudinous requests share approximate detection results. It is crucial to reduce such fuzzy redundancy by reusing the previous results. A key challenge is that the requests mapping to the reusable result are only similar but not identical. An efficient method for similarity matching is needed to justify the use case. To this end, by taking advantage of TCAM's ap-proximate matching capability and NMC's computing efficiency, we design SODA, a first-of-its-kind hardware accelerator which sits in the mobile base stations between autonomous vehicles and MEC servers. We design efficient feature encoding and partition algorithms for SODA to ensure the quality of the similarity matching and result reuse. Our evaluation shows that SODA significantly improves the system performance and the detection results exceed the accuracy requirements on the subject matter, qualifying SODA as a practical domain-specific solution.
Wenquan Xu, Haoyu Song 0001, Linyang Hou, Xinggong Zhang, Chuwen Zhang, Wei Hu 0003, Yi Wang 0004, Bin Liu 0001
INFOCOM1
2021 Scalable Hardware Content Router: Architecture, Modeling and Performance
abstract
Current Internet is evolving with the gradual shift from the traditional host-to-host communication model to the new host-to-content paradigm, which will eventually lead to a network of caches. The novel Named Data Networking (NDN) has been proposed as a future Internet architecture to embrace this paradigmatic shift, where caching becomes an ubiquitous functionality available at each router.A router with the functionality of content caching, running on NDN mechanisms, is termed as an NDN-based content router. Previous researchers focused on software content routers (SCR), which leverage a commercial off-the-shelf computer to execute content caching/accessing and named-based packet forwarding. SCR can only achieve limited throughput, which is far below the speed requirements of modern routers. Facing this situation, in this paper, we propose a hardware-based content router (HCR), aiming at purchasing wire-speed processing. We design a physically concise architecture for decoupling the packet buffers in line cards from the content caches attached to storage cards, enabling separate management and optimization while facilitating a modular structure for smooth capacity upgrade in response to increasing storage utilization. For lowering the operating complexity and reducing the storage management cost, we choose to employ distributed caches working in a cooperated manner by using consistent hashing. We model several candidate storage organizing schemes and carry out theoretical analyses for comparison. Analytical and synthetic workload-driven results show that the consistent hashing scheme achieves high cache performance and low cost simultaneously.
Bin Liu 0001, Huichen Dai, Wenquan Xu, Tong Yun, Ji Miao
IWQoS3
2021 PQR: Prediction-supported Quality-aware Routing for Uninterrupted Vehicle Communication
abstract
Vehicle to Vehicle (V2V) communication opens a new way to make vehicles directly communicate with each other, providing faster responses for time-sensitive tasks than cellular networks. Effective V2V routing protocols are essential yet challenging, as the high dynamic road environment makes communication easy to break. Many prediction methods proposed in the existing protocols to address this issue are either flawed or have a poor effect. In this paper, to cope with the two aspects of the problems that cause communication interrupt, i.e., link breaks and route quality degradation, we design an acceleration-based trajectory prediction algorithm to estimate the link lifetime, and a machine learning model to predict route quality. Based on the prediction algorithms, we propose PQR, a Prediction-supported Quality-aware Routing protocol, which can proactively switch to a better route before the current link breaks or the route quality degrades. Especially, considering the limitations of the current routing protocols, we elaborate a new hybrid routing protocol that integrates the topology-based method and location-based method to achieve instant communication. Simulation results show that PQR outperforms the existing protocols in Packet Delivery Ratio (PDR), Roundtrip Time (RTT), and Normalized Routing Overhead (NRO). Specifically, we have also implemented a vehicular testbed to demonstrate PQR’s real-world performance, and results show that PQR achieves almost no packet loss with latency less than 10ms during route handoff for topology change.
Wenquan Xu, Xuefeng Ji, Chuwen Zhang, Beichuan Zhang 0001, Yu Wang 0003, Xiaojun Wang 0001, Yunsheng Wang 0001, Jianping Wang 0001, Bin Liu 0001
IWQoS1
2020 GlobalInsight: An LSTM Based Model for Multi-Vehicle Trajectory Prediction
abstract
Intelligent Transport System (ITS) raises the increasing demand on accurate vehicle trajectory prediction for navigation efficiency. The rapidly developing 5G networks provides communications with high transmission bandwidth and super-low latency, paving the way for Mobile Edge Computing (MEC) to calculate more accurate trajectory prediction for vehicles, as the MEC server holds more comprehensive vehicular information. However, the current methods for trajectory prediction are not efficient due to the dynamical environment. To address this issue, we propose GlobalInsight, a Long Short-Term Memory (LSTM) based model, which runs on the MEC to perform accurate trajectory prediction for multiple vehicles no matter how scenario changes. In particular, we use three auxiliary layers to respectively capture the principal component of vehicle features, social interaction of adjacent vehicles, and the cross-vehicle correlation of similar vehicles. We further integrate the above information into LSTM in the main layer to enhance the trajectory learning and prediction. We evaluate our model under the NGSIM dataset, and experimental results exhibit that our model outperforms the state-of-the-art approaches.
Wenquan Xu, Zhikang Chen, Chuwen Zhang, Xuefeng Ji, Yunsheng Wang 0001, Bin Liu 0001
ICC1
2020 PBC: Effective Prefix Caching for Fast Name Lookups
Chuwen Zhang, Haoyu Song 0001, Beichuan Zhang 0001, Yi Wang 0004, Ying Wan 0001, Wenquan Xu, Bin Liu 0001
Networking7
2020 A Three-level Routing Hierarchy in improved SDN-MEC-VANET Architecture
abstract
Existing routing algorithms that based on traditional Vehicular Ad-Hoc NETwork (VANET) architectures cannot provide fast and diverse routing services due to dynamic and unstable environment. To address this issue, we propose a three-level routing hierarchy in improved Software-Defined VANET architecture based on Mobile Edge Computing (SDN-MEC-VANET) to improve routing performance and enrich the data transmission mode for the VANET. Moreover, it can be applied to almost all VANET protocols, enabling protocol-independent forwarding. Besides, this improved architecture can coordinate different edge devices to timely adjust the service delivery strategy under the predictive correction from controllers, providing high-bandwidth and low-delay transmission for Internet of Vehicles (IoV). Meanwhile, MEC technology is introduced to perform local control, leveraging the storage and computing capabilities of edge devices to reduce the processing pressure of the controller. Simulation results show that our routing algorithm in improved network architecture can achieve a higher packet delivery ratio within a reasonable delay than other approaches under different scenarios of network scale, communication frequency and vehicular velocity.
Xuefeng Ji, Wenquan Xu, Chuwen Zhang, Bin Liu 0001
WCNC2
2020 NIHR: Name/ID Hybrid Routing in Information-centric VANET
abstract
Vehicular Ad hoc network (VANET) has received great attention in recent research, but many challenges still lie in innovating efficient routing protocols to support the highly dynamic environment. Existing ID-based routing protocols cannot fundamentally tackle the dynamic topology problem in VANET. The recent emerging Information-Centric Networking (ICN) makes routing decisions based on data itself instead of a particular host, seeming to have the potential to handle the dynamic topology, but problems (e.g., severe flooding overhead) still remain. Therefore, inspired by the idea of ICN, we propose a name/ID hybrid routing (NIHR) protocol that combines the data-namebased routing and host-ID-based routing to address the above two issues simultaneously. In particular, we develop an announce strategy to improve the efficiency of the in-network cache, and we design a bloom filter based structure to achieve fast content lookup. Simulation results show NIHR's high performance in terms of Packet Delivery Ratio (PDR), Roundtrip time (RTT) and roundtrip hop count. Especially, to verify NIHR's performance in real-world scenarios, we have implemented a vehicular real-time video conference system based on MK5 OBU [1] (On-Board Unit).
Wenquan Xu, Xuefeng Ji, Chuwen Zhang, Bin Liu 0001
WCNC1
2020 DBN based SD-ARX model for nonlinear time series prediction and analysis
Wenquan Xu, Hui Peng 0001, Xiaoying Tian
Appl. Intell.1
2019 P3R: Realizing Robust Routing for VANET Using Trajectory Prediction and Crossroad Recognition
abstract
High topology dynamics and intermittent connectivity in Vehicular Ad hoc Network (VANET) bring huge challenges to end-to-end communication. Existing routing protocols for MANET such as AODV and OLSR work fine under modest mobility, but have a difficult time to handle frequent topology changes in VANET. This paper proposes Peeking at the Past and Present Routing (P3R), a routing protocol that will calculate next-hops when the past forwarding is considered invalid. The next-hop calculation is based on the predicted locations of forwarder's neighbors and the packet's destination node, overcoming the inaccuracy caused by stale location information. Furthermore, we differentiate vehicles on crossroads as they have high connectivity in actual urban streets. In this way, P3R is able to deal with link breakages quickly and exploit new links. Simulation results show that P3R outperforms state-of-the-art alternatives in terms of packet delivery ratio, delay and cost, while maintaining strong scalability and robustness. We also implement P3R in a real vehicular testbed and the results reveal it has high connectivity on real streets.
Chuwen Zhang, Huichen Dai, Yang Li 0062, Wenquan Xu, Xuefeng Ji, Ying Wan 0001, Gong Zhang 0001, Bin Liu 0001
ICPADS5
2019 A hybrid modelling method for time series forecasting based on a linear regression model and deep learning
Wenquan Xu, Hui Peng 0001, Xiaoyong Zeng, Feng Zhou 0005, Xiaoying Tian
Appl. Intell.1
2018 OBMA: Minimizing Bitmap Data Structure with Fast and Uninterrupted Update Processing
abstract
Software-based IP route lookup is one of the key components in Software Defined Networks. To address challenges on density, power and cost, Commodity CPU is preferred over other platforms to run lookup algorithms. As network functions become richer and more dynamic, route updates are more frequent. Unfortunately, previous works put less effort on fast incremental updates. On the other hand, The cache in CPU could be a performance limiter due to its small size, which requires algorithm designers to give high priority on storage efficiency in addition to time complexity. In this paper, we propose a new route lookup algorithm, OBMA, which improves update performance and storage efficiency while maintaining high lookup speed. The extensive experiments over real-word traces show that OBMA reduces the memory footprint to just 4.52 bytes/prefix, supports update speed up to 7.2 M/s which is 12.5 times faster than the state-of-the-art algorithm Poptrie. Besides, OBMA achieves up to 195.87 Mpps lookup speed with a single thread. Tests on comprehensive performance of lookup and update show that OBMA can sustain high lookup speed with update speed increasing.
Chuwen Zhang, Haoyu Song 0001, Ying Wan 0001, Wenquan Xu, Huichen Dai, Yang Li 0062, Bin Liu 0001
IWQoS5