VLDB 2026 Research / reviewers in the wild / expert
Wenwen Fu
dblp:170/9943
· DBLP profile ↗
15ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 1 first-author · 6 since 2021Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Deadlock-Free Bridge Module for Inter-Chiplet Cache-Coherent Communication in an Open Chiplet EcosystemabstractThe envisioned open chiplet ecosystem promises significant reductions in chip design complexity and cost by enabling designers to rapidly assemble standard chiplets from diverse vendors. Constructing such an open chiplet ecosystem requires support from Network-on-Chip (NoC) routing algorithms, as integrating multiple chiplets onto an interposer can potentially lead to inter-chiplet deadlock. Prior work avoids deadlock through methods like turn restrictions, virtual channel isolation, or packet injection control, or recovers from deadlock using mechanisms such as escape channels or bubble flow control. These approaches achieve a favorable balance regarding modularity, performance, and cost. However, they still rely on designers possessing detailed knowledge of the internal NoC architecture within each chiplet. This requirement impedes the development of a truly open ecosystem, as it constrains chiplet interoperability and vendor independence. Addressing this limitation, we propose a Deadlock-Free Bridge Module (DFBM) designed to resolve interchiplet deadlock without relying on the specifics of individual chiplet NoC implementations. The DFBM infers the transmission behavior of inter-chiplet packets by analyzing the dependency relationships among coherence protocol transaction flows. It then employs a packet injection control mechanism to isolate inter- and intra-chiplet packets, thereby preventing deadlock. DFBMs can be seamlessly interconnected between arbitrary chiplets to achieve deadlock-freedom, eliminating the need for modifications to their internal NoC architectures. Experimental results demonstrate that DFBM incurs only 2.5% area overhead, while achieving a performance improvement ranging from 1% to 7%. Zhiqiang Chen 0006, Wenwen Fu, Yongwen Wang |
HPCA | 2 |
| 2026 | PDE-TSN: Enable TSN Autonomous Self-healing under Link Faults
Wenwen Fu, Xuyan Jiang, Wei Quan 0004, Tao Li 0008, Zhigang Sun 0002 |
SECON | 2 |
| 2026 | DP4C: A SoC Architecture for NN-Driven Network Functions With the Intelligent PlaneabstractNeural-network-driven (NN-driven) network functions and their implementation on the data plane are emerging topics due to demonstrated accuracy and high performance. Meanwhile, we argue that deploying NN-driven network functions should satisfy two design goals: the generality to support various NN models, and the flexibility to operate various network functions. Unfortunately, existing work cannot satisfy both goals simultaneously. In this paper, we introduce the concept of the Intelligent Plane for NN-driven network functions, and propose DP4C, a cross-plane SoC architecture that integrates the intelligent, control, and data planes within a single chip. DP4C comprises the programmable NN inference engine that iteratively executes inference to ensure model generality in the intelligent plane, a multi-core RISC-V CPU that parses inference results into diverse network functions through its architectural flexibility in the control plane, and the switch fabric in the data plane. To further eliminate the performance bottleneck, we propose (i) the direct register access mechanism coupled with custom instructions to reduce the overhead of cross-plane data migration; and (ii) the multi-core pipelining with adaptive batch-processing for multi-core CPU. DP4C SoC is fabricated using 130 nm technology and has already been deployed in industrial IoT environments. We also design three distinct test cases to evaluate DP4C, fully demonstrating the model generality and operational flexibility. Dong Wen 0004, Tao Li 0008, Wenwen Fu, Chenglong Li 0007, Zhuochen Fan, Chao Zhuo, Zhiting Xiong, Junnan Li 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | Megabits Down to Kilobits: Memory-Efficient Time-Aware Shaping for TSNabstractTime-Sensitive Networking (TSN) provides bounded latency and low jitter for cyber-physical systems, such as industrial control. As a key component of TSN, the Time-Aware Shaper (TAS) applies gate control rules to control the transmission time of frames in critical flows. TAS stores the gate control rules for each frame in the gate control table. However, in typical industrial setups, the memory usage of the table could reach over tens of megabits and even exceed the total memory capacity of TSN switches.To address this issue, we propose a memory-efficient TAS design named METAS. It transitions from a per-frame to a per-flow approach. METAS stores one persistent rule for a flow and dynamically generates a temporary rule for a frame only when the frame arrives. We prototyped METAS on an FPGA, and experimental results show that METAS reduces memory usage from 14.34 Mbits to 288 Kbits when supporting 1,024 flows, using just 1.56% of the FPGA’s logic resources while maintaining microsecondlevel latency and nanosecond-level jitter for critical flows. Xuyan Jiang, Wenwen Fu, Xiangrui Yang 0002, Wenfei Wu |
DAC | 2 |
| 2025 | Facial Expression Generation from Text with FaceCLIP
Wenwen Fu, Wenjuan Gong, Chen-Yang Yu, Wei Wang 0115, Jordi Gonzàlez 0001 |
J. Comput. Sci. Technol. | 1 |
| 2025 | FooDog: Empower TSN for Efficient PolicingabstractTime-Sensitive Networking (TSN) is an emerging real-time Ethernet technology that provides deterministic communication for time-sensitive (TS) traffic. At its core, TSN utilizes Per-Stream Filtering and Policing (PSFP) gates to mitigate the disruption of unavoidable frame drift. However, as first identified in this work, the naive PSFP gate design results in heavy memory usage, which hinders normal switching functions. This work proposes an efficient PSFP gate design called FooDog. FooDog employs a two-stage structure and a dual-engine policing mechanism to realize memory-efficient, logic-compact, and fast policing while maintaining minimal latency and jitter for TS traffic. Results on FPGA prototypes show that FooDog consumes only hundreds of kilobits of memory, reducing on-chip memory overheads by more than 90% compared to the unoptimized PSFP gate design. Additionally, it maintains end-to-end latency in the microsecond range and jitter below 150 nanoseconds under abnormal traffic conditions, comparable to typical TSN performance without anomalies. Xuyan Jiang, Xiangrui Yang 0002, Tongqing Zhou, Wenfei Wu, Wenwen Fu, Wei Quan 0004, Yingwen Chen 0001, Yihao Jiao, Zhigang Sun 0002 |
IEEE Trans. Netw. | 5 |
| 2024 | Hebo: FPGA-based Transfer Time Planning for Volatile Traffic in TSNabstractTime-Sensitive Networking (TSN) is an advanced technology designed for real-time Ethernet communications, providing extremely low latency, minimal jitter, and lossless data transfer for time-sensitive critical traffic. Despite its benefits, TSN faces challenges with volatile traffic, where the time between frames constantly changes, leading to potential network performance issues. To tackle this issue, this paper introduces Hebo, a novel solution designed for zero frame loss and minimal latency of volatile traffic. Hebo employs a centralized controller, built with field-programmable gate array (FPGA), to dynamically plan the timing of volatile traffic in real-time. This approach ensures real-time data transmission across the network by efficiently allocating network resources. Our real-world and simulation experiments show that Hebo could significantly improve network performance, achieving less than 100 microseconds in end-to-end delay and eliminating nearly all frame loss (reducing it from approximately 80% to zero) in industrial automation scenarios. Xuyan Jiang, Zitong Wang 0002, Xiangrui Yang 0002, Yihao Jiao, Tianci Yu, Wenwen Fu, Yinhan Sun, Zhigang Sun 0002 |
IWQoS | 7 |
| 2023 | Poster Abstract: A Network-on-Chip Router Architecture for Industrial Internet-of-Thing GatewaysabstractMore processors are integrated into Industrial Internet-of-Thing gateways to perform increasing emerging applications. Network-on-chip (NoC) offers a scalable, high-throughput, and energy-efficient communicate infrastructure. However, existing NoC routers cannot guarantee differentiated quality-of-service (QoS) for diversified applications. Hence, we propose a novel NoC router architecture with the gate control mechanism to provide customized QoS. Chenglong Li 0007, Cunlu Li, Wenwen Fu, Tao Li 0008 |
IPSN | 3 |
| 2023 | DRA: Ultra-Low Latency Network I/O for TSN Embedded End-SystemsabstractTime-Sensitive Networking (TSN) is a promising open-source technique for hard real-time embedded fields such as industrial automation and autonomous driving. The embedded end-system deployed in TSN must guarantee deterministic latency and jitter for network I/O in above scenarios. Recent researchers focus on eliminating jitter but neglect to reduce latency. The network I/O latency is still too high to satisfy the dozen-microsecond requirements of latency-sensitive applications. We observed that (1) the data path from registers to external storage is actually the bottleneck for latency reduction, and (2) the serial processing of CPU and NIC can be further optimized. Therefore, we proposed DRA (Direct Register Access), a novel network I/O mechanism to achieve microsecond-level latency. DRA delivers whole packet data from the NIC directly into extended registers inside the CPU, avoiding the waste of time to move data between internal registers and external storage. Moreover, DRA promotes parallelization of CPU processing and NIC transfer to reduce latency. Considering the increasing popularity of RISC-V ISA in embedded systems, we prototype DRA using an open-source RISC-V core on FPGA and evaluate it under real-life application scenarios. Compared with existing mechanisms, experimental results demonstrate that DRA reduces the network I/O latency and jitter by at least 60% and 30%, achieving microsecond-level network I/O processing. Chenglong Li 0007, Tao Li 0008, Junnan Li 0002, Wenwen Fu |
IWQoS | 4 |
| 2023 | Fenglin-I: An Open-Source Time-Sensitive Networking Chip Enabling Agile CustomizationabstractTime-Sensitive Networking (TSN) technology is experiencing diverse application requirements and forming a complicated standard system. It is extremely difficult to design a one-fits-all chip for all TSN applications. Therefore, application-driven TSN chip customization is inevitable. Generally, chip customization starts from a “clean-slate”. For complicated ASIC chips, that results in significant development overhead. Inspired by RISC-V chips, an open-source template will significantly reduce the customization complexity. Along this road, we propose an open-source TSN chip named Fenglin-I. Fenglin-I includes a high-level abstraction to build a relationship between application requirements and chip implementation, source code of a real chip named FastTSN to provide reference code for chip implementation, and software tools to facilitate chip verification. Based on Fenglin-I, we further propose a TSN chip customization method that provides step-by-step guidance about customizing TSN chips agilely. To verify the effectiveness of Fenglin-I and the proposed customization method, we use FPGA arrays to prototype and verify FastTSN. The results show that FastTSN achieves microsecond-level transmission jitter for unicast and multicast time-critical traffic. Additionally, we demonstrate two domain-specific TSN chip customization cases in which the customized chips reuse at least 84$\%$of FastTSN code while meeting their requirements. Wenwen Fu, Wei Quan 0004, Jinli Yan, Zhigang Sun 0002 |
IEEE Trans. Computers | 1 |
| 2022 | TASP: Enabling Time-Triggered Task Scheduling in TSN-Based Mixed-Criticality SystemsabstractDistributed mixed-criticality system (DMCS) has been widely used in various critical domains such as self-driving cars and space crafts. To guarantee the end-to-end QoS (i.e., deadline/jitter requirements) of sensing-controlling-actuating control loops (CL) applications, DMCS adopts Time-Sensitive Networking (TSN), an emerging real-time Ethernet technology, for communication between end systems (ES). TSN provides a synchronized global clock and guarantees bounded delay for time-critical traffic in CLs, making it possible for DMCS to collaboratively schedule the computation (on ES) and communication (in TSN) to meet the Quality of Service (QoS) requirement. However, as modern DMCS tends to use fully-fledged Linux distributions (rather than a custom real-time OS) on ES to enjoy Linux’s mature ecosystem, it is challenging for DMCS to realize TSN-based QoS guarantees because the event-triggered scheduling of Linux on ES is incompatible with TSN.This paper proposes TAsk Scheduling Puppeteer (TASP), a mechanism that schedules CL tasks based on TSN without modifying the Linux OS. The key idea of TASP is to manipulate task scheduling by controlling the timing of CL packet submissions at the interface between TSN and ES. Specifically, TASP extracts two parameters: Fore Guardband (ForeGB) and Back Guardband (BackGB). During the ForeGB period before a CL packet’s submission, TASP forbids any packets’ submission; while during the BackGB period after a CL packet’s submission, TASP forbids any other CL packets’ submission. ForeGB and BackGB can ensure that there is at most one schedulable task on the ES at any time, and thus Linux has no choice but to schedule the only task, making the Linux scheduler a puppet. We have deployed TASP and evaluated it in real-world TSN switches based on an open-source TSN project, OpenTSN. The TASP-enabled ES can achieve task scheduling precisely based on TSN’s global clock, which outperforms the original ES by reducing end-to-end jitter from milliseconds to microseconds. Xuyan Jiang, Yiming Zhang 0003, Wenwen Fu, Xiangrui Yang 0002, Yinhan Sun, Zhigang Sun 0002 |
IWQoS | 3 |
| 2020 | TSN-Builder: Enabling Rapid Customization of Resource-Efficient Switches for Time-Sensitive NetworkingabstractTime-Sensitive Networking (TSN) emerges as a promising technique empowering deterministic forwarding on standard Ethernet without sacrificing compatibility. There are some commercial off-the-shelf (COTS) switches that support TSN recently. However, the resource partitioning on these switches is normally inefficient for the on-chip memory resource in many specific application scenarios. We observe that the critical requirements (e.g., topology, flow features) of these scenarios are pre-determined. Thus, developing a TSN switch in a Top-down approach is feasible and urgently needed.In this paper, we propose TSN-Builder, a template-based developing model for customizing resource-efficient TSN switches rapidly with targeted application-dependent requirements. TSN-Builder decomposes the integrated TSN switching function into multiple function templates. With a fine-grained resource abstraction, TSN-Builder provides platform-independent customization interfaces for developers to customize the resource parameters. We prototype TSN switches on FPGA to evaluate the resource consumption and performance under different application scenarios. Experimental results show that TSN-Builder reduces the on-chip memory by up to 80.53% under the same Quality-of-Service, compared to the resource configuration in the COTS switch. Jinli Yan, Wei Quan 0004, Xiangrui Yang 0002, Wenwen Fu, Zhigang Sun 0002 |
DAC | 4 |
| 2019 | STRIDE: Single-Trip-Time Based Reliable Data Transport Protocol for the Reconfigurable CloudabstractIn a recent development, reconfigurable clouds become a viable solution to overcome practical problems in clouds, such as scalability, delay, etc., by offloading computation tasks to reconfigurable hardware, FPGA. Several existing techniques, such as TCP/IP Offload Engine (TOE) and Lightweight Transport Layer (LTL), are still hard to be implemented in real-world deployment due to large overhead or stringent dependency of the underlying network. In this paper, we propose STRIDE, a novel inter-FPGA data communication protocol, to provide reliable end-to-end communication which addresses practical problems in deployment. In our design, STRIDE leverages FPGA's abilities through programming, such as precise timestamping, to make more accurate measurement on end-to-end delay and deliver more precise control in managing traffic in cloud. We implement STRIDE on a FPGA-based network experimental platform and demonstrate that STRIDE reduces various hardware resources consumption by 36% to 49% compared to TOE. Additionally, it also improves flow completion time in comparison to TOE by 2.2X. We further demonstrate STRIDE outperforms DCTCP and TCP-Vegas in OMNET simulator by up to 1.8X and 2.3X on average and 99th percentile respectively in large scale setting. Wenwen Fu, Tao Li 0008, Jialun Yang, Junnan Li 0002, Zhigang Sun 0002 |
ICC | 1 |
| 2018 | FAS: Using FPGA to Accelerate and Secure SDN Software SwitchesabstractSoftware-Defined Networking (SDN) promises the vision of more flexible and manageable networks but requires certain level of programmability in the data plane to accommodate different forwarding abstractions. SDN software switches running on commodity multicore platforms are programmable and are with low deployment cost. However, the performance of SDN software switches is not satisfactory due to the complex forwarding operations on packets. Moreover, this may hinder the performance of real-time security on software switch. In this paper, we analyze the forwarding procedure and identify the performance bottleneck of SDN software switches. An FPGA-based mechanism for accelerating and securing SDN switches, named FAS (FPGA-Accelerated SDN software switch), is proposed to take advantage of the reconfigurability and high-performance advantages of FPGA. FAS improves the performance as well as the capacity against malicious traffic attacks of SDN software switches by offloading some functional modules. We validate FAS on an FPGA-based network processing platform. Experiment results demonstrate that the forwarding rate of FAS can be 44% higher than the original SDN software switch. In addition, FAS provides new opportunity to enhance the security of SDN software switches by allowing the deployment of bump-in-the-wire security modules (such as packet detectors and filters) in FPGA. Wenwen Fu, Tao Li 0008, Zhigang Sun 0002 |
Secur. Commun. Networks | 1 |
| 2015 | Satellite derived sea surface salinity validation in South China Sea areaabstractAfter almost 5 years of SMOS launched, accuracy of satellite SSS measurements is evaluated/validated in most areas. But In South-China Sea area (4 ° N-25 ° N, 105 ° E-125 ° E), few calibration/validation efforts is made in this area. In this paper we will validate the satellite (SMOS/Aquarius) derived SSS measurements based on moored buoys and ARGO in-situ measurements. For SMOS SSS measurements, 3-day, 10-day and monthly averaged SSS products with space resolution of 0.25°×0.25° and 1°×1° are validated in the period of year 2011 to year 2013 in South China Sea area. Aquarius SSS products of daily, 7-day and monthly averaged data with 1°×1° space solution are also validated in the period of year 2011 to year 2014. An overall mean difference of -0.3psu (0.6psu for RMSE) for Aquarius measurements compared to -0.6psu (1.6psu for RMSE) for SMOS is given in our validation. Hongping Li, Wenwen Fu, Haihua Chen 0004, Changjun Li |
IGARSS | 2 |