EDBT 2026 Demo / reviewers in the wild / expert
Chuwen Zhang
dblp:170/9703
· DBLP profile ↗
29ranked-venue papers
6as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 19 · 5 first-author · 11 since 2021Systems, architecture and hardware · 4 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Composable Emulation Framework for Whitebox Switches
Congcong Miao, Xianneng Zou, Chuwen Zhang, Qihang Liu, Zhijie Yan, Yanke Zhang, Yong Jiang 0001, Qiao Xiang, Xin Jin 0008, Zili Meng, Ang Chen 0001 |
NSDI | 3 |
| 2026 | PAED: Physics-Anchored and Emotion Disentangled Memory Diffusion for Talking Face Generation
Xingchu Zhang, Chuwen Zhang, Zehao Deng |
PAKDD (1) | 2 |
| 2026 | Turbo: Efficiently Serving Long-Context Large Language Models with In-Network AggregationabstractLLM supporting long contexts faces a critical memory bottleneck due to the linear growth of KV cache. Distributing the storage across multiple GPUs alleviates this burden but introduces significant communication overhead or traffic incast, especially during the decoding phase. We propose Turbo, a first-of-its-kind in-network aggregation system that accelerates long-context inference by offloading query broadcast and attention aggregation to switches. We address three key challenges to map complex attention mechanisms onto restricted switch hardware: (i) To bypass the switch's inability to buffer global states or perform complex operations, we devise online table-based aggregation, which decomposes global reduction into pairwise operations and approximates nonlinear functions via lookup tables. (ii) To circumvent the restriction on retroactive state access in RMT pipelines, we introduce a rolling forward scheme that propagates states to enable cross-stage updates. (iii) To mitigate aggregation stragglers caused by topology-induced load imbalance, we construct a load-aware aggregation tree that optimizes workload distribution. Evaluations on a Tofino2-based testbed show that Turbo reduces end-to-end inference latency by up to 37%. Large-scale simulations on NS-3 demonstrate that Turbo significantly outperforms state-of-the-art baselines in both inference latency and network traffic reduction with negligible accuracy loss. Ying Wan 0001, Yuchen Xu 0003, Chuwen Zhang, Yingsheng Huang, Wenquan Xu, Jialin Li 0001, Mingwei Xu 0001, Wenfei Wu, Congcong Miao |
SIGCOMM | 3 |
| 2025 | An Enhanced Alternating Direction Method of Multipliers-Based Interior Point Method for Linear and Conic OptimizationabstractThe alternating-direction-method-of-multipliers-based (ADMM-based) interior point method, or ABIP method, is a hybrid algorithm that effectively combines interior point method (IPM) and first-order methods to achieve a performance boost in large-scale linear optimization. Different from traditional IPM that relies on computationally intensive Newton steps, the ABIP method applies ADMM to approximately solve the barrier penalized problem. However, similar to other first-order methods, this technique remains sensitive to condition number and inverse precision. In this paper, we provide an enhanced ABIP method with multiple improvements. First, we develop an ABIP method to solve the general linear conic optimization and establish the associated iteration complexity. Second, inspired by some existing methods, we develop different implementation strategies for the ABIP method, which substantially improve its performance in linear optimization. Finally, we conduct extensive numerical experiments in both synthetic and real-world data sets to demonstrate the empirical advantage of our developments. In particular, the enhanced ABIP method achieves a 5.8× reduction in the geometric mean of run time on 105 selected linear optimization instances from Netlib, and it exhibits advantages in certain structured problems, such as support vector machine and PageRank. However, the enhanced ABIP method still falls behind commercial solvers in many benchmarks, especially when high accuracy is desired. We posit that it can serve as a complementary tool alongside well-established solvers. History: Accepted by Antonio Frangioni, Area Editor for Design & Analysis of Algorithms—Continuous. Funding: This research was supported by the National Natural Science Foundation of China [Grants 72394360, 72394364, 72394365, 72225009, 72171141, and 72150001] and by the Program for Innovative Research Team of Shanghai University of Finance and Economics. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2023.0017 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2023.0017 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ . Wenzhi Gao, Dongdong Ge, Bo Jiang 0007, Yuntian Jiang, Jingsong Liu, Chenyu Xue 0001, Yinyu Ye 0001, Chuwen Zhang |
INFORMS J. Comput. | 11 |
| 2024 | Trust Region Methods for Nonconvex Stochastic Optimization beyond Lipschitz SmoothnessabstractIn many important machine learning applications, the standard assumption of having a globally Lipschitz continuous gradient may fail to hold. This paper delves into a more general (L0, L1)-smoothness setting, which gains particular significance within the realms of deep neural networks and distributionally robust optimization (DRO). We demonstrate the significant advantage of trust region methods for stochastic nonconvex optimization under such generalized smoothness assumption. We show that first-order trust region methods can recover the normalized and clipped stochastic gradient as special cases and then provide a unified analysis to show their convergence to first-order stationary conditions. Motivated by the important application of DRO, we propose a generalized high-order smoothness condition, under which second-order trust region methods can achieve a complexity of O(epsilon(-3.5)) for convergence to second-order stationary points. By incorporating variance reduction, the second-order trust region method obtains an even better complexity of O(epsilon(-3)), matching the optimal bound for standard smooth optimization. To our best knowledge, this is the first work to show convergence beyond the first-order stationary condition for generalized smooth optimization. Preliminary experiments show that our proposed algorithms perform favorably compared with existing methods. Chenghan Xie, Chuwen Zhang, Dongdong Ge, Yinyu Ye 0001 |
AAAI | 3 |
| 2024 | Empower Programmable Pipeline for Advanced Stateful Packet Processing
Zhikang Chen, Haoyu Song 0001, Yinchao Zhang, Hanyi Zhou, Ruoyu Sun 0009, Wenkuo Dong, Chuwen Zhang, Yang Xu 0010, Bin Liu 0001 |
NSDI | 10 |
| 2024 | A Homogenization Approach for Gradient-Dominated Stochastic OptimizationabstractGradient dominance property is a condition weaker than strong convexity, yet sufficiently ensures global convergence even in non-convex optimization. This property finds wide applications in machine learning, reinforcement learning (RL), and operations management. In this paper, we propose the stochastic homogeneous second-order descent method (SHSODM) for stochastic functions enjoying gradient dominance property based on a recently proposed homogenization approach. Theoretically, we provide its sample complexity analysis, and further present an enhanced result by incorporating variance reduction techniques. Our findings show that SHSODM matches the best-known sample complexity achieved by other second-order methods for gradient-dominated stochastic optimization but without cubic regularization. Empirically, since the homogenization approach only relies on solving extremal eigenvector problem at each iteration instead of Newton-type system, our methods gain the advantage of cheaper computational cost and robustness in ill-conditioned problems. Numerical experiments on several RL tasks demonstrate the better performance of SHSODM compared to other off-the-shelf methods. Jiyuan Tan, Chenyu Xue 0001, Chuwen Zhang, Dongdong Ge, Yinyu Ye 0001 |
UAI | 3 |
| 2024 | A Customized Augmented Lagrangian Method for Block-Structured Integer ProgrammingabstractInteger programming with block structures has received considerable attention recently and is widely used in many practical applications such as train timetabling and vehicle routing problems. It is known to be NP-hard due to the presence of integer variables. We define a novel augmented Lagrangian function by directly penalizing the inequality constraints and establish the strong duality between the primal problem and the augmented Lagrangian dual problem. Then, a customized augmented Lagrangian method is proposed to address the block-structures. In particular, the minimization of the augmented Lagrangian function is decomposed into multiple subproblems by decoupling the linking constraints and these subproblems can be efficiently solved using the block coordinate descent method. We also establish the convergence property of the proposed method. To make the algorithm more practical, we further introduce several refinement techniques to identify high-quality feasible solutions. Numerical experiments on a few interesting scenarios show that our proposed algorithm often achieves a satisfactory solution and is quite effective. Rui Wang 0060, Chuwen Zhang, Shanwen Pu, Jianjun Gao 0001, Zaiwen Wen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | OBMA: Scalable Route Lookups With Fast and Zero-Interrupt UpdatesabstractSoftware-based IP route lookup is a key component for packet forwarding in Software Defined Networks. Running lookup algorithms on commodity CPUs is flexible and scalable, which shows advantages on cost and power consumption over the hardware-based forwarding engines. However, dynamic network functions and services make route updates more frequent than ever. Existing algorithms often fall short of the incremental update requirements. In this paper, we propose the Overlay BitMap Algorithm (OBMA), which contains several variations, to support extraordinary update performance while maintaining the highest-in-class lookup speed and storage efficiency. Starting from the basic OBMA_B, we develop two variations with different tradeoffs for different application scenarios. OBMA_L supports faster lookups than OBMA_B at a small cost of update speed. OBMA_S achieves better storage efficiency than OBMA_B at a small cost of lookup throughput. We run our algorithms on a commodity CPU and evaluate them with real-world route tables and traces. The experiments show that OBMA achieves the lowest memory footprint, the highest update speed, and over 200 Mpps lookup throughput. Specifically, OBMA_S reduces the memory footprint to 3.98 bytes/prefix which is 25.33% smaller that of the state-of-the-art Poptrie; OBMA_L supports 252.02 Mpps lookup throughput with a single thread, and more than 600 Mpps with multiple parallel threads in a single CPU, significantly outperforming the state-of-the-art Poptrie and SAIL; OBMA_B supports updates at a rate of 14.58M updates/s which is 15 times faster than Poptrie. The tests show that the update process has little interference with the lookup process for OBMA, and achieves zero-interrupt to lookups with multiple threads. Chuwen Zhang, Haoyu Song 0001, Ying Wan 0003, Wenquan Xu, Bin Liu 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2023 | FASTeller: A Hardware Partial Aggregator for Accurate Flow Counting in Cloud NetworksabstractAccurate per-flow counting is beyond the capability of network switches due to the sheer flow number. The conventional divide-and-conquer method by distributing the traffic to multiple servers for software processing is costly. The solution therefore quests for a combination of hardware and software where the hardware with limited resources aims to undertake a part of the job and reduce the workload of software, achieving a desirable balance of cost and performance. To this end we design FASTeller to be deployed on SmartNICs. It is tuned to maximize the counting aggregation level in hardware, leaving the server a much lower workload for accurate per-flow counting and sparing the server capacity for post-counting functions such as network intrusion detection. The novelty lies in the multi-tier hardware caching data structure which is tailored for the flow distribution properties of real traffic. We build an FPGA-based prototype and evaluate the performance of FASTeller. The low-cost implementation achieves the highest performance among the methods in comparison and can easily sustain the accurate perflow counting for 100Gbps traffic with the least software load. Tong Yun, Yinxin Kuang, Zhuang Ling, Haoyu Song 0001, Peilong Wang, Chuwen Zhang, Mao Miao, Zhaogeng Li, Donghua Huang, Bin Liu 0001 |
ICNP | 6 |
| 2022 | TSN-Peeper: an Efficient Traffic Monitor in Time-Sensitive NetworkingabstractTime-Sensitive Networking (TSN) is proposed in recent years to satisfy the strict performance requirements of time-sensitive traffic in a growing number of emerging applications. Even though several traffic scheduling algorithms have been standardized for TSN to pursue this goal, time-sensitive flows may not be forwarded as planned and thus fail to achieve the expected performance in real networks. The fundamental cause lies in the fact that static offline planning cannot adapt to the intrinsic dynamic factors in TSN (e.g., time-synchronization error) at runtime. Hence, next-generation TSN will benefit from a closed-loop design where a performance monitoring system provides feedback of real-time packet-forwarding information. In our research, TSN-Peeper, a light-weight, fast-response and full-coverage TSN performance monitoring system, is designed and evaluated. This paper describes its architecture design and data collection mechanisms that enable timely identification and collection of packet-forwarding misbehavior at low-cost in TSN. TSN-Peeper offloads the misbehavior identification in the switch to relieve the burden on the controller and network bandwidth. To reduce the interruption frequency to the controller, it uses probe packets to collect misbehavior information in aggregation with optimized path planning. To realize controllable reporting delays, it optimizes the sending moments of probe packets according to the flow settings. Experimental results verify that TSN-Peeper offers fast response with low cost while providing full coverage and being scalable. Chuwen Zhang, Zerui Tian, Liang Cheng 0001, Yuxi Liu 0017, Ying Wan 0001, Wenquan Xu, Tian Pan 0001, Yang Xu 0010, Yi Wang 0004, Hailong Zhu, Bin Liu 0001 |
ICNP | 1 |
| 2021 | FastUp: Fast TCAM Update for SDN Switches in Datacenter NetworksabstractTCAM is widely used for flow table lookup in Software-Defined Networking (SDN) switches for datacenter and enterprise networks. While its lookup throughput is unparalleled, TCAM updating, particularly for new rule insertions, can impair the overall system performance. A rule insertion entails two steps: 1) Computing the rule moving operations; and 2) Interrupting the TCAM lookups to apply the operations. In previous work, the performance gain on one step is always at the expense of the performance loss on the other. However, update throughput and latency depend on both. In this paper, we present a faster and more balanced TCAM update scheme, which not only achieves the shortest interrupt time so far but also significantly reduces the computation time. By using a novel sequential stack, FastUp reduces the time and space complexity of the state-of-the-art schemes from$O(m^{2})$and$O(m)$to$O(m\log h)$and$O(h)$, respectively, where$h << m$. Evaluations show that FastUp shortens the computation time and the interrupt time by$100\times$and$1.6\times$, respectively, which is equivalent to update delay${15\times}$reduction and$\mathbf{10\times}$update throughput gain against the state-of-the-art schemes. Moreover, we debunk a common mistake and show the dynamic programming based algorithm cannot be used to solve the reorder problem, and instead we use a bidirectional rule moving method to address the problem. In addition, we propose a practical method to find the theoretical lower bound of interrupt time in relatively large TCAM, which can be used to evaluate the optimality degree of TCAM update schemes. Evaluations show that FastUp achieves 90 % optimality. Ying Wan 0001, Haoyu Song 0001, Hao Che, Yang Xu 0010, Yi Wang 0004, Chuwen Zhang, Zhijun Wang 0001, Tian Pan 0001, Hao Li 0011, Hong Jiang 0001, Chengchen Hu, Bin Liu 0001 |
ICDCS | 6 |
| 2021 | PIPO: Efficient Programmable Scheduling for Time Sensitive NetworkingabstractTime Sensitive Networking (TSN) is an emerging Ethernet technology for real-time systems. To address different Quality-of-Service (QoS) requirements of applications, IEEE 802.1 TSN Task Group has standardized several packet scheduling and shaping algorithms. The software implementation of these algorithms is hard to meet the performance requirements, while the hardware implementation in Application-Specific Integrated Circuit (ASIC) is inflexible. A hardware-programmable scheduler is necessary to deal with this dilemma. Among the existing primitives, the most expressive one is Push-In-Extract-Out (PIEO), but its complexity makes the implementation very expensive. A relatively lower-cost implementation of PIEO cannot guarantee the scheduling correctness for the most critical Time-Triggered (TT) traffic in TSN. As a remedy, in this paper we propose a new Push-In-Pick-Out (PIPO) primitive under a TSN programmable scheduling framework. Composed of simple priority queues, PIPO can express all existing TSN scheduling and shaping algorithms, and is flexible enough to support future ones. Our PIPO implementation guarantees the TT traffic scheduling correctness. The simulation results corroborate the theoretical analysis that the low-cost PIPO can closely approximate PIEO and sustain a high bandwidth utilization. The prototype on Xilinx FPGA shows that, with 2,048 inputs, the PIPO-based scheduler achieves a throughput of 70 Mpps, which is 1.64x higher than the PIEO-based one, but using only 14.7% Look-Up Tables (LUTs) and 40.5% Block RAMs of the latter. Chuwen Zhang, Zhikang Chen, Haoyu Song 0001, Ruyi Yao, Yang Xu 0010, Yi Wang 0004, Ji Miao, Bin Liu 0001 |
ICNP | 1 |
| 2021 | Adaptive Batch Update in TCAM: How Collective Optimization Beats Individual OnesabstractRule update in TCAM has long been identified as a key technical challenge due to the rule order constraint. Existing algorithms take each rule update as an independent task. However, emerging applications produce batch rule update requests. Processing the updates individually causes high aggregated cost which can strain the processor and/or incur excessive TCAM lookup interrupts. This paper presents the first true batch update algorithm, ABUT. Unlike the other alleged batch update algorithms, ABUT collectively evaluates and optimizes the TCAM placement for whole batches throughout. By applying the topology grouping and maintaining the group order invariance in TCAM, ABUT achieves substantial computing time reduction yet still yields the best-in-class placement cost. Our evaluations show that ABUT is ideal for low-latency and high-throughput batch TCAM updates in modern high-performance switches. Ying Wan 0001, Haoyu Song 0001, Yang Xu 0010, Chuwen Zhang, Yi Wang 0004, Bin Liu 0001 |
INFOCOM | 4 |
| 2021 | SODA: Similar 3D Object Detection Accelerator at Network Edge for Autonomous DrivingabstractOffloading the 3D object detection from autonomous vehicles to MEC is appealing because of the gains on quality, latency, and energy. However, detection requests lead to repetitive computations since the multitudinous requests share approximate detection results. It is crucial to reduce such fuzzy redundancy by reusing the previous results. A key challenge is that the requests mapping to the reusable result are only similar but not identical. An efficient method for similarity matching is needed to justify the use case. To this end, by taking advantage of TCAM's ap-proximate matching capability and NMC's computing efficiency, we design SODA, a first-of-its-kind hardware accelerator which sits in the mobile base stations between autonomous vehicles and MEC servers. We design efficient feature encoding and partition algorithms for SODA to ensure the quality of the similarity matching and result reuse. Our evaluation shows that SODA significantly improves the system performance and the detection results exceed the accuracy requirements on the subject matter, qualifying SODA as a practical domain-specific solution. Wenquan Xu, Haoyu Song 0001, Linyang Hou, Xinggong Zhang, Chuwen Zhang, Wei Hu 0003, Yi Wang 0004, Bin Liu 0001 |
INFOCOM | 6 |
| 2021 | PQR: Prediction-supported Quality-aware Routing for Uninterrupted Vehicle CommunicationabstractVehicle to Vehicle (V2V) communication opens a new way to make vehicles directly communicate with each other, providing faster responses for time-sensitive tasks than cellular networks. Effective V2V routing protocols are essential yet challenging, as the high dynamic road environment makes communication easy to break. Many prediction methods proposed in the existing protocols to address this issue are either flawed or have a poor effect. In this paper, to cope with the two aspects of the problems that cause communication interrupt, i.e., link breaks and route quality degradation, we design an acceleration-based trajectory prediction algorithm to estimate the link lifetime, and a machine learning model to predict route quality. Based on the prediction algorithms, we propose PQR, a Prediction-supported Quality-aware Routing protocol, which can proactively switch to a better route before the current link breaks or the route quality degrades. Especially, considering the limitations of the current routing protocols, we elaborate a new hybrid routing protocol that integrates the topology-based method and location-based method to achieve instant communication. Simulation results show that PQR outperforms the existing protocols in Packet Delivery Ratio (PDR), Roundtrip Time (RTT), and Normalized Routing Overhead (NRO). Specifically, we have also implemented a vehicular testbed to demonstrate PQR’s real-world performance, and results show that PQR achieves almost no packet loss with latency less than 10ms during route handoff for topology change. Wenquan Xu, Xuefeng Ji, Chuwen Zhang, Beichuan Zhang 0001, Yu Wang 0003, Xiaojun Wang 0001, Yunsheng Wang 0001, Jianping Wang 0001, Bin Liu 0001 |
IWQoS | 3 |
| 2021 | T-Cache: Efficient Policy-Based Forwarding Using Small TCAMabstractTernary Content Addressable Memory (TCAM) is widely used by modern routers and switches to support policy-based forwarding due to its incomparable lookup speed and flexible matching patterns. However, the limited TCAM capacity does not scale with the ever-increasing rule table size due to the high hardware cost and high power consumption. At present, using TCAM just as a rule cache is an appealing solution, but one must resolve several tricky issues including the rule dependency and the associated TCAM updates. In this paper, we propose a new approach which can generate dependency-free rules to cache. By removing the rule dependency, the complex TCAM update problem also disappears. We provide the complete T-cache system design including slow path processing and cache replacement, and implement a T-cache prototype on Barefoot Tofino switches. We conduct comprehensive software simulations and hardware experiments based on real-world and synthesized rule tables and packet traces to show that T-cache is efficient and robust for network traffic in various scenarios. Ying Wan 0001, Haoyu Song 0001, Yang Xu 0010, Tian Pan 0001, Chuwen Zhang, Yi Wang 0004, Bin Liu 0001 |
IEEE/ACM Trans. Netw. | 6 |
| 2020 | GlobalInsight: An LSTM Based Model for Multi-Vehicle Trajectory PredictionabstractIntelligent Transport System (ITS) raises the increasing demand on accurate vehicle trajectory prediction for navigation efficiency. The rapidly developing 5G networks provides communications with high transmission bandwidth and super-low latency, paving the way for Mobile Edge Computing (MEC) to calculate more accurate trajectory prediction for vehicles, as the MEC server holds more comprehensive vehicular information. However, the current methods for trajectory prediction are not efficient due to the dynamical environment. To address this issue, we propose GlobalInsight, a Long Short-Term Memory (LSTM) based model, which runs on the MEC to perform accurate trajectory prediction for multiple vehicles no matter how scenario changes. In particular, we use three auxiliary layers to respectively capture the principal component of vehicle features, social interaction of adjacent vehicles, and the cross-vehicle correlation of similar vehicles. We further integrate the above information into LSTM in the main layer to enhance the trajectory learning and prediction. We evaluate our model under the NGSIM dataset, and experimental results exhibit that our model outperforms the state-of-the-art approaches. Wenquan Xu, Zhikang Chen, Chuwen Zhang, Xuefeng Ji, Yunsheng Wang 0001, Bin Liu 0001 |
ICC | 3 |
| 2020 | FastUp: Compute a Better TCAM Update Scheme in Less Time for SDN SwitchesabstractWhile widely used for flow tables in SDN switches, TCAM faces challenges for rule updates. Both the computation time and interrupt time need to be short. We propose FastUp, a new TCAM update algorithm, which improves the previous dynamic programming-based algorithms. Evaluations show that FastUp shortens the computation time by 40~100× and the interrupt time by 1.2~2.5×. In addition, we are the first to prove the NP-hardness of the optimal TCAM update problem, and provide a practical method to evaluate an algorithm's degree of optimality. Experiments show that FastUp's optimality reaches 90%. Ying Wan 0001, Haoyu Song 0001, Hao Che, Yang Xu 0010, Yi Wang 0004, Chuwen Zhang, Zhijun Wang 0001, Tian Pan 0001, Hao Li 0011, Hong Jiang 0001, Chengchen Hu, Zhikang Chen, Bin Liu 0001 |
ICDCS | 6 |
| 2020 | HOLNET: A Holistic Traffic Control Framework for Datacenter NetworksabstractIn this paper, we put forward a HOListic traffic control framework for datacenter NETworks (HOLNET). HOLNET reformulates the network utility maximization (NUM) framework into a HOLNET NUM framework that fully harnesses the potential of the existing NUM-based solutions to allow large families of traffic control protocols of various degrees of sophistication to be developed, i.e., host-based, single or multiple Class-of-Service (CoS) enabled, single or multi-path congestion control, with or without in-network load balancing. Unlike the existing solutions that are largely empirical and point by design, HOLNET is a principled, systematic framework. All the protocols in a family developed under HOLNET share a common, user-defined global optimization objective and fairness criterion. As a result, the protocols in a family can be fairly compared and carefully selected to fully explore the performance, scalability and design complexity tradeoffs. Case studies, based on both a single and a multi-path host-based solutions, demonstrate the viability and flexibility in HOLNET design space exploration. To further test the backward compatibility and performance with respect to some existing lightweight solutions, we develop HOLNET-UTA, an integrated congestion control and load balancing protocol, achieving TCP-fair resource allocation. HOLNET-UTA is found by simulation to improve the average flow completion time (FCT) by more than 20%, compared to DRILL with DCTCP. Zhijun Wang 0001, Akshit Singhal, Yunxiang Wu, Chuwen Zhang, Hao Che, Hong Jiang 0001, Bin Liu 0001, Constantino M. Lagoa |
ICNP | 4 |
| 2020 | T-cache: Dependency-free Ternary Rule Cache for Policy-based ForwardingabstractTernary Content Addressable Memory (TCAM) is widely used by modern routers and switches to support policy-based forwarding. However, the limited TCAM capacity does not scale with the ever-increasing rule table size. Using TCAM just as a rule cache is a plausible solution, but one must resolve several tricky issues including the rule dependency and the associated TCAM updates. In this paper, we propose a new approach which can generate dependency-free rules to cache. By removing the rule dependency, the TCAM update problem also disappears. We provide the complete T-cache system design including slow path processing and cache replacement. Evaluations based on real-world and synthesized rule tables and traces show that T-cache is efficient and robust for network traffic in various scenarios. Ying Wan 0001, Haoyu Song 0001, Yang Xu 0010, Tian Pan 0001, Chuwen Zhang, Bin Liu 0001 |
INFOCOM | 6 |
| 2020 | PBC: Effective Prefix Caching for Fast Name Lookups
Chuwen Zhang, Haoyu Song 0001, Beichuan Zhang 0001, Yi Wang 0004, Ying Wan 0001, Wenquan Xu, Bin Liu 0001 |
Networking | 1 |
| 2020 | A Three-level Routing Hierarchy in improved SDN-MEC-VANET ArchitectureabstractExisting routing algorithms that based on traditional Vehicular Ad-Hoc NETwork (VANET) architectures cannot provide fast and diverse routing services due to dynamic and unstable environment. To address this issue, we propose a three-level routing hierarchy in improved Software-Defined VANET architecture based on Mobile Edge Computing (SDN-MEC-VANET) to improve routing performance and enrich the data transmission mode for the VANET. Moreover, it can be applied to almost all VANET protocols, enabling protocol-independent forwarding. Besides, this improved architecture can coordinate different edge devices to timely adjust the service delivery strategy under the predictive correction from controllers, providing high-bandwidth and low-delay transmission for Internet of Vehicles (IoV). Meanwhile, MEC technology is introduced to perform local control, leveraging the storage and computing capabilities of edge devices to reduce the processing pressure of the controller. Simulation results show that our routing algorithm in improved network architecture can achieve a higher packet delivery ratio within a reasonable delay than other approaches under different scenarios of network scale, communication frequency and vehicular velocity. Xuefeng Ji, Wenquan Xu, Chuwen Zhang, Bin Liu 0001 |
WCNC | 3 |
| 2020 | NIHR: Name/ID Hybrid Routing in Information-centric VANETabstractVehicular Ad hoc network (VANET) has received great attention in recent research, but many challenges still lie in innovating efficient routing protocols to support the highly dynamic environment. Existing ID-based routing protocols cannot fundamentally tackle the dynamic topology problem in VANET. The recent emerging Information-Centric Networking (ICN) makes routing decisions based on data itself instead of a particular host, seeming to have the potential to handle the dynamic topology, but problems (e.g., severe flooding overhead) still remain. Therefore, inspired by the idea of ICN, we propose a name/ID hybrid routing (NIHR) protocol that combines the data-namebased routing and host-ID-based routing to address the above two issues simultaneously. In particular, we develop an announce strategy to improve the efficiency of the in-network cache, and we design a bloom filter based structure to achieve fast content lookup. Simulation results show NIHR's high performance in terms of Packet Delivery Ratio (PDR), Roundtrip time (RTT) and roundtrip hop count. Especially, to verify NIHR's performance in real-world scenarios, we have implemented a vehicular real-time video conference system based on MK5 OBU [1] (On-Board Unit). Wenquan Xu, Xuefeng Ji, Chuwen Zhang, Bin Liu 0001 |
WCNC | 3 |
| 2019 | P3R: Realizing Robust Routing for VANET Using Trajectory Prediction and Crossroad RecognitionabstractHigh topology dynamics and intermittent connectivity in Vehicular Ad hoc Network (VANET) bring huge challenges to end-to-end communication. Existing routing protocols for MANET such as AODV and OLSR work fine under modest mobility, but have a difficult time to handle frequent topology changes in VANET. This paper proposes Peeking at the Past and Present Routing (P3R), a routing protocol that will calculate next-hops when the past forwarding is considered invalid. The next-hop calculation is based on the predicted locations of forwarder's neighbors and the packet's destination node, overcoming the inaccuracy caused by stale location information. Furthermore, we differentiate vehicles on crossroads as they have high connectivity in actual urban streets. In this way, P3R is able to deal with link breakages quickly and exploit new links. Simulation results show that P3R outperforms state-of-the-art alternatives in terms of packet delivery ratio, delay and cost, while maintaining strong scalability and robustness. We also implement P3R in a real vehicular testbed and the results reveal it has high connectivity on real streets. Chuwen Zhang, Huichen Dai, Yang Li 0062, Wenquan Xu, Xuefeng Ji, Ying Wan 0001, Gong Zhang 0001, Bin Liu 0001 |
ICPADS | 1 |
| 2019 | Ultra-Fast Bloom Filters using SIMD TechniquesabstractThe network link speed is growing at an ever-increasing rate, which requires all network functions on routers/switches to keep pace. Bloom filter is a widely-used membership check data structure in networking applications. Correspondingly, it also faces the urgent demand of improving the performance in membership check speed. To this end, this paper proposes a new Bloom filter variant called Ultra-Fast Bloom Filters (UFBF), by leveraging the Single Instruction Multiple Data (SIMD) techniques. We make three improvements for UFBF to accelerate the membership check speed. First, we develop a novel hash computation algorithm which can compute multiple hash functions in parallel with the use of SIMD instructions. Second, we elaborate a Bloom filter's bit-test process from sequential to parallel, enabling more bit-tests per unit time. Third, we improve the cache efficiency of membership check by encoding an element's information to a small block so that it can fit into a cache-line. We further generalize UFBF, called c-UFBF, to make UFBF supporting large number of hash functions. Both theoretical analysis and extensive evaluations show that the UFBF greatly outperforms the state-of-the-art Bloom filter variants on membership check speed. Jianyuan Lu, Ying Wan 0001, Yang Li 0062, Chuwen Zhang, Huichen Dai, Yi Wang 0004, Gong Zhang 0001, Bin Liu 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2018 | OBMA: Minimizing Bitmap Data Structure with Fast and Uninterrupted Update ProcessingabstractSoftware-based IP route lookup is one of the key components in Software Defined Networks. To address challenges on density, power and cost, Commodity CPU is preferred over other platforms to run lookup algorithms. As network functions become richer and more dynamic, route updates are more frequent. Unfortunately, previous works put less effort on fast incremental updates. On the other hand, The cache in CPU could be a performance limiter due to its small size, which requires algorithm designers to give high priority on storage efficiency in addition to time complexity. In this paper, we propose a new route lookup algorithm, OBMA, which improves update performance and storage efficiency while maintaining high lookup speed. The extensive experiments over real-word traces show that OBMA reduces the memory footprint to just 4.52 bytes/prefix, supports update speed up to 7.2 M/s which is 12.5 times faster than the state-of-the-art algorithm Poptrie. Besides, OBMA achieves up to 195.87 Mpps lookup speed with a single thread. Tests on comprehensive performance of lookup and update show that OBMA can sustain high lookup speed with update speed increasing. Chuwen Zhang, Haoyu Song 0001, Ying Wan 0001, Wenquan Xu, Huichen Dai, Yang Li 0062, Bin Liu 0001 |
IWQoS | 1 |
| 2017 | Ultra-Fast Bloom Filters using SIMD techniquesabstractThe network link speed is increasing at an alarming rate, which requires all network functions on routers/switches to keep pace. Bloom filter is a widely-used membership check data structure in network applications. It also faces the urgent demand of improving the performance in membership check speed. To this end, this paper proposes a new Bloom filter variant called Ultra-Fast Bloom Filters, by leveraging the SIMD techniques. We make three improvements for the UFBF to accelerate the membership check speed. First, we develop a novel hash computation algorithm which can compute multiple hash functions in parallel with the use of SIMD instructions. Second, we change a Bloom filter's bit-test process from sequential to parallel. Third, we increase the cache efficiency of membership check by encoding an element's information to a small block which can easily fit into a cache-line. Both theoretical analysis and extensive simulations show that the UFBF greatly exceeds the state-of-the-art Bloom filter variants on membership check speed. Jianyuan Lu, Ying Wan 0001, Yang Li 0062, Chuwen Zhang, Huichen Dai, Yi Wang 0004, Gong Zhang 0001, Bin Liu 0001 |
IWQoS | 4 |
| 2015 | Hyperspectral and multispectral image fusion using CNMF with minimum endmember simplex volume and abundance sparsity constraintsabstractHyperspectral (HS) remote sensing image with finer spectral information has great advantages in feature identification and classification. However, the spatial resolution of HS image is usually low due to practical limitations. In this paper, the low-spatial-resolution HS image is fused with the high-spatial-resolution multispectral (MS) image of the same observation scene to improve its spatial resolution. A novel spectral unmixing based HS and MS image fusion approach (VSC-CNMF) is proposed, in which CNMF with minimum endmember simplex volume and abundance sparsity constraints is employed for coupled unmixing of HS and MS images. Simulative experiments are employed for verification and comparison. The experimental results illustrate that the newly proposed VSC-CNMF based HS and MS fusion algorithm outperforms several state-of-the-art unmixing based fusion approaches in cases with moderate number of endmembers. Yifan Zhang 0006, Chuwen Zhang, Mingyi He, Shaohui Mei |
IGARSS | 4 |