VLDB 2026 Research / reviewers in the wild / expert
Tao Li 0008
dblp:75/4601-8
· DBLP profile ↗
33ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0001-7168-3628ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 13 · 1 first-author · 5 since 2021Systems, architecture and hardware · 9 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Security and privacy · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EvalMuse-40K: A Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Alignment EvaluationabstractText-to-Image (T2I) generation models have achieved significant advancements. Correspondingly, many automated methods emerge to evaluate the image-text alignment capabilities of generative models. However, the performance comparison among these automated methods is constrained by the limited scale of existing datasets. Additionally, existing datasets lack the capacity to assess the performance of automated methods at a fine-grained level. In this study, we contribute an EvalMuse-40K dataset, gathering 40K image-text pairs with fine-grained human annotations for image-text alignment-related tasks. In the construction process, we employ various strategies such as balanced prompt sampling and data re-annotation to ensure the diversity and reliability of our dataset. This allows us to comprehensively evaluate the performance of image-text alignment methods for T2I models. Based on this dataset, we introduce an efficient automated evaluation method termed FGA-BLIP2, which enables Fine-Grained Alignment evaluation solely by inputting images and text leveraging BLIP2, without visual question answering for each fine-grained element. Experimental results show the proposed FGA-BLIP2 efficiently achieves good performance on multiple image-text alignment datasets. Meanwhile, benefiting from the high efficiency and fine-grained evaluation capability of FGA-BLIP2, we apply it as a reward model to improve text-to-image models, which effectively enhances the image-text alignment ability of text-to-image models. Shuhao Han, Haotian Fan, Jiachen Fu, Tao Li 0008, Junhui Cui, Yunqiu Wang, Yang Tai, Chunle Guo, Chongyi Li |
AAAI | 5 |
| 2026 | PDE-TSN: Enable TSN Autonomous Self-healing under Link Faults
Wenwen Fu, Xuyan Jiang, Wei Quan 0004, Tao Li 0008, Zhigang Sun 0002 |
SECON | 7 |
| 2026 | KaleidoScope: A Co-Processor for Neural-Network-Driven Intelligent Data Plane
Dong Wen 0004, Zhongpei Liu, Tong Yang 0003, Tianyun Li, Yanshu Wang, Tao Li 0008, Zhuochen Fan, Qing Li 0006, Zhigang Sun 0002 |
IEEE Trans. Computers | 6 |
| 2026 | DP4C: A SoC Architecture for NN-Driven Network Functions With the Intelligent PlaneabstractNeural-network-driven (NN-driven) network functions and their implementation on the data plane are emerging topics due to demonstrated accuracy and high performance. Meanwhile, we argue that deploying NN-driven network functions should satisfy two design goals: the generality to support various NN models, and the flexibility to operate various network functions. Unfortunately, existing work cannot satisfy both goals simultaneously. In this paper, we introduce the concept of the Intelligent Plane for NN-driven network functions, and propose DP4C, a cross-plane SoC architecture that integrates the intelligent, control, and data planes within a single chip. DP4C comprises the programmable NN inference engine that iteratively executes inference to ensure model generality in the intelligent plane, a multi-core RISC-V CPU that parses inference results into diverse network functions through its architectural flexibility in the control plane, and the switch fabric in the data plane. To further eliminate the performance bottleneck, we propose (i) the direct register access mechanism coupled with custom instructions to reduce the overhead of cross-plane data migration; and (ii) the multi-core pipelining with adaptive batch-processing for multi-core CPU. DP4C SoC is fabricated using 130 nm technology and has already been deployed in industrial IoT environments. We also design three distinct test cases to evaluate DP4C, fully demonstrating the model generality and operational flexibility. Dong Wen 0004, Tao Li 0008, Wenwen Fu, Chenglong Li 0007, Zhuochen Fan, Chao Zhuo, Zhiting Xiong, Junnan Li 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2026 | UTFormer: An Ultra-Lightweight Transformer for Traffic Classification
Dong Wen 0004, Tianyun Li, Zhuochen Fan, Qing Li 0006, Fa Zhu, Chenglong Li 0007, Athanasios V. Vasilakos, Tao Li 0008 |
IEEE Trans. Inf. Forensics Secur. | 9 |
| 2026 | Adaptive Affinity Memorization With Layer Mutation for Multimodal Deepfake Continual Detection
Jianbin Ye, Bo Liu 0014, Zijian Gao, Wuyang Chen 0002, Tao Li 0008, Huaimin Wang 0001, Kele Xu |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | Approaching 100% Confidence in Stream Summary through ReliableSketchabstractTo approximate sums of values in key-value data streams, sketches are widely used in databases and networking systems.They offer high-confidence approximations for any given key while ensuring low time and space overhead.While existing sketches are proficient in estimating individual keys, they struggle to maintain this high confidence across all keys collectively, an objective that is critically important in both algorithm theory and its practical applications.We propose ReliableSketch, the first to control the error of all keys to less than Λ with a small failure probability Δ, requiring only 𝑂 (1 + Δ ln ln( 𝑁 Λ )) amortized time and 𝑂 ( 𝑁 Λ + ln( 1 Δ )) space.Furthermore, its simplicity makes it hardware-friendly, and we implement it on CPU servers, FPGAs, and programmable switches.Our experiments show that under the same small space, ReliableSketch not only keeps all keys' errors below Λ but also delivers competitive throughput among accuracy-oriented baselines, outperforming * Both authors contributed equally to this research. Yuhan Wu 0001, Hanbo Wu, Xilai Liu, Yuxuan Tian 0001, Yikai Zhao 0001, Tong Yang 0003, Kaicheng Yang 0001, Tao Li 0008, Lihua Miao, Gaogang Xie |
IMC | 10 |
| 2025 | EasyViT: An Adaptive Collaborative Edge Computing Framework for Vision TransformerabstractDeploying Vision Transformers (ViTs) in edge computing environments presents significant challenges due to their high computing demands and the resource constraints of edge devices. While collaborative edge computing and dynamic token dropping offer potential solutions, existing approaches suffer from rigid strategies that fail to adapt to diverse conditions of network and computing resources at the edge. This paper introduces EasyViT, an adaptive framework that optimizes ViT deployment through the joint coordination of collaborative edge computing and dynamic token dropping. Key innovations include: (1) A token dropping model that integrates dynamic token dropping and collaborative edge environments, formulating an integer linear programming (ILP) optimization problem. (2) An Approximate Stochastic Gradient Descent (ASGD) method with atomic gradient calculation, which transforms the NP-hard ILP problem into a continuous space for rapid near-optimal solution generation. Extensive evaluations on a real-world edge testbed with multiple Raspberry Pi nodes demonstrate that EasyViT achieves 1.06–5.06× speedup over baseline methods under 20 configurations of edge environments, while maintaining model accuracy within 2.8% degradation. The proposed framework exhibits the adaptability across diverse ViT architectures, network bandwidths, and computing resources. Dong Wen 0004, Guanping Liang, Tianyun Li, Junnan Li 0002, Tao Li 0008 |
IEEE Internet Things J. | 6 |
| 2025 | A Performance-Balanced Scheduling Algorithm for Diverse Real-World TSN ScenariosabstractTime-Sensitive Networking (TSN) achieves low-delay and low-jitter traffic transmission through different traffic scheduling mechanisms. However, despite numerous algorithms developed based on these mechanisms, most fail to concurrently support multipath, hybrid, and multicast traffic, which are prevalent in real-world scenarios. Moreover, balancing performance metrics such as success rate, bandwidth utilization, and computation overhead remains challenging for these algorithms, significantly limiting their application in diverse TSN scenarios. To solve this problem, this paper proposes a universal ultra-low-delay and zero-jitter traffic scheduling model. Based on this model, this paper further designs a performance-balanced algorithm. The algorithm improves traffic scheduling success rate through joint routing and scheduling, increases network bandwidth utilization through hybrid traffic scheduling, and achieves low computation overhead through policy-based searching. Finally, extensive experiments demonstrate that the algorithm effectively balances performance metrics across diverse real-world scenarios. It achieves high scheduling success rate under real-world traffic loads ($\gt $20% improvement over non-joint routing), increased bandwidth utilization in the presence of hybrid traffic (18.3% enhancement over non-hybrid traffic scheduling), and low computation overhead ($\lt $2 minutes). Xuyan Jiang, Rulin Liu, Tao Li 0008, Wei Quan 0004, Zhigang Sun 0002 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2024 | Dissect Black Box: Interpreting for Rule-Based Explanations in Unsupervised Anomaly DetectionabstractIn high-stakes sectors such as network security, IoT security, accurately distinguishing between normal and anomalous data is critical due to the significant implications for operational success and safety in decision-making. The complexity is exacerbated by the presence of unlabeled data and the opaque nature of black-box anomaly detection models, which obscure the rationale behind their predictions. In this paper, we present a novel method to interpret the decision-making processes of these models, which are essential for detecting malicious activities without labeled attack data. We put forward the Segmentation Clustering Decision Tree (SCD-Tree), designed to dissect and understand the structure of normal data distributions. The SCD-Tree integrates predictions from the anomaly detection model into its splitting criteria, enhancing the clustering process with the model's insights into anomalies. To further refine these segments, the Gaussian Boundary Delineation (GBD) algorithm is employed to define boundaries within each segmented distribution, effectively delineating normal from anomalous data points. At this point, this approach addresses the curse of dimensionality by segmenting high-dimensional data and ensures resilience to data drift and perturbations through flexible boundary fitting. We transform the intricate operations of anomaly detection into an interpretable rule's format, constructing a comprehensive set of rules for understanding. Our method's evaluation on diverse datasets and models demonstrates superior explanation accuracy, fidelity, and robustness over existing method, proving its efficacy in environments where interpretability is paramount. Ruoyu Li 0003, Nengwu Wu, Qing Li 0006, Xinhan Lin, Tao Li 0008, Yong Jiang 0001 |
NeurIPS | 7 |
| 2024 | 2FA Sketch: Two-Factor Armor Sketch for Accurate and Efficient Heavy Hitter Detection in Data Streams
Xilai Liu, Xinyi Zhang 0004, Tao Li 0008, Tong Yang 0003, Gaogang Xie |
NPC (2) | 4 |
| 2023 | Poster Abstract: A Network-on-Chip Router Architecture for Industrial Internet-of-Thing GatewaysabstractMore processors are integrated into Industrial Internet-of-Thing gateways to perform increasing emerging applications. Network-on-chip (NoC) offers a scalable, high-throughput, and energy-efficient communicate infrastructure. However, existing NoC routers cannot guarantee differentiated quality-of-service (QoS) for diversified applications. Hence, we propose a novel NoC router architecture with the gate control mechanism to provide customized QoS. Chenglong Li 0007, Cunlu Li, Wenwen Fu, Tao Li 0008 |
IPSN | 4 |
| 2023 | DRA: Ultra-Low Latency Network I/O for TSN Embedded End-SystemsabstractTime-Sensitive Networking (TSN) is a promising open-source technique for hard real-time embedded fields such as industrial automation and autonomous driving. The embedded end-system deployed in TSN must guarantee deterministic latency and jitter for network I/O in above scenarios. Recent researchers focus on eliminating jitter but neglect to reduce latency. The network I/O latency is still too high to satisfy the dozen-microsecond requirements of latency-sensitive applications. We observed that (1) the data path from registers to external storage is actually the bottleneck for latency reduction, and (2) the serial processing of CPU and NIC can be further optimized. Therefore, we proposed DRA (Direct Register Access), a novel network I/O mechanism to achieve microsecond-level latency. DRA delivers whole packet data from the NIC directly into extended registers inside the CPU, avoiding the waste of time to move data between internal registers and external storage. Moreover, DRA promotes parallelization of CPU processing and NIC transfer to reduce latency. Considering the increasing popularity of RISC-V ISA in embedded systems, we prototype DRA using an open-source RISC-V core on FPGA and evaluate it under real-life application scenarios. Compared with existing mechanisms, experimental results demonstrate that DRA reduces the network I/O latency and jitter by at least 60% and 30%, achieving microsecond-level network I/O processing. Chenglong Li 0007, Tao Li 0008, Junnan Li 0002, Wenwen Fu |
IWQoS | 2 |
| 2023 | TreeSensing: Linearly Compressing Sketches with FlexibilityabstractA Sketch is an excellent probabilistic data structure, which records the approximate statistics of data streams. Linear additivity is an important property of sketches. This paper studies how to keep the linear property after sketch compression. Most existing compression methods do not keep the linear property. We propose TreeSensing, an accurate, efficient, and flexible framework to linearly compress sketches. In TreeSensing, we first separate a sketch into two parts according to counter values. For the sketch with small counters, we propose a technique called TreeEncoding to compress it into a hierarchical structure. For the sketch with large counters, we propose a technique called SketchSensing to compress it using compressive sensing. We theoretically analyze the accuracy of TreeSensing. We use TreeSensing to compress 7 sketches and conduct two end-to-end experiments: distributed measurement and distributed machine learning. Experimental results show that TreeSensing outperforms prior art on both accuracy and efficiency, which achieves up to 100× smaller error and 5.1× higher speed than state-of-the-art Cluster-Reduce. All related codes are open-sourced. Zirui Liu 0002, Yixin Zhang 0002, Yifan Zhu 0011, Ruwen Zhang, Tong Yang 0003, Kun Xie 0001, Tao Li 0008, Bin Cui 0001 |
Proc. ACM Manag. Data | 8 |
| 2023 | A Deterministic Embedded End-System Tightly Coupled With TSN ScheduleabstractDistributed real-time systems (DRTSs) composed of many embedded end-systems have been widely adopted in the industrial fields. Time-sensitive networking (TSN), as a promising communication infrastructure for DRTS, has shown great potential in industry and academia. TSN assumes that an end-system can release critical tasks to process critical packets strictly according to the prescheduled time. Unfortunately, two factors currently damage this assumption: 1) the jitter caused by system architecture during task release, task execution, and packet transmission; and 2) a TSN schedule result may exceed the execution capability of the end-system and cause conflicts. This article proposed deterministic chip (DetChip), a system-on-chip capable of deterministically implementing the TSN schedule result. DetChip supports time-triggered task release, time-predictable task execution and precise network transmission. Based on DetChip, this article first formalizes the execution capability of the end-system as end-system constraints (ECs). Existing TSN scheduling algorithms integrating ECs can solve conflicts by co-scheduling end-systems and the TSN network. Compared with previous works, DetChip only introduces few clock cycles jitter for critical task execution according to the TSN schedule result. The proposed ECs obtain$2\times$–$8\times$more conflict-free solutions for advanced scheduling algorithms with a linear increase in time overhead. Compared with the general end-system, DetChip can reduce$6\times$–$20\times$processing jitter to achieve better clock synchronization. Chenglong Li 0007, Zonghui Li, Tao Li 0008, Cunlu Li |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | MapEmbed: Perfect Hashing with High Load Factor and Fast UpdateabstractPerfect hashing is a hash function that maps a set of distinct keys to a set of continuous integers without collision. However,most existing perfect hash schemes are static, which means that they cannot support incremental updates, while most datasets in practice are dynamic. To address this issue, we propose a novel hashing scheme, namely MapEmbed Hashing. Inspired by divide-and-conquer and map-and-reduce, our key idea is named map-and-embed and includes two phases: 1) Map all keys into many small virtual tables; 2) Embed all small tables into a large table by circular move. Our experimental results show that under the same experimental setting, the state-of-the-art perfect hashing (dynamic perfect hashing) can achieve around 15% load factor, around 0.3 Mops update speed, while our MapEmbed achieves around 90% ~ 95% load factor, and around 8.0 Mops update speed per thread. All codes of ours and other algorithms are open-sourced at GitHub. Yuhan Wu 0001, Zirui Liu 0002, Jie Gui, Haochen Gan, Yuhao Han, Tao Li 0008, Ori Rottenstreich, Tong Yang 0003 |
KDD | 7 |
| 2020 | Update Latency Optimization of Packet Classification for SDN Switch on FPGAabstractFPGA is widely used in real-time network processing such as packet classification in SDN switches due to high performance and programmability. BV-based approaches on FPGA provide a performance guarantee for multi-field packet classification, but no update latency guarantee. We thus present SplitBV for the efficient update by splitting the ruleset into subrulesets that can be performed in parallel. Results show that our approach can reduce 73% and 36% update latency on average for synthetic 5-tuple rules and OpenFlow1.0 rules respectively. Chenglong Li 0007, Tao Li 0008, Junnan Li 0002, Zilin Shi |
FCCM | 2 |
| 2019 | TabTree: A TSS-assisted Bit-selecting Tree Scheme for Packet Classification with Balanced Rule MappingabstractTo support fast rule updates in SDN, the Open vSwitch implements Priority Sorting Tuple Space Search (PSTSS) for its packet classifications. Although it has good performance on rule updates, it has a performance concern on table lookups. In contrast, decision tree methods are being actively investigated for high throughput, but they are not able to support fast updates because of rule replications. CutSplit, the state-of-the-art decision tree scheme, provides a novel rule update mechanism by avoiding tree reconstructions. However, its average update time is still two orders of magnitude larger than PSTSS. Meanwhile, existing decision trees are not only unbalanced but also depth unbounded, making them difficult to be optimized on FPGA. In this paper, we present a new decision tree scheme called TabTree, which achieves high performance on both lookups and updates. By mapping rules into tree nodes dynamically, a very limited number of balanced trees with bounded depths can be generated without the trouble of rule replications. Experimental results show that, TabTree has comparable update performance to PSTSS, but it outperforms PSTSS significantly in terms of number of memory accesses for packet classification. Additionally, TabTree is more practical for implementations on FPGA. Wenjun Li 0004, Tong Yang 0003, Yeim-Kuan Chang, Tao Li 0008, Hui Li 0022 |
ANCS | 4 |
| 2019 | STRIDE: Single-Trip-Time Based Reliable Data Transport Protocol for the Reconfigurable CloudabstractIn a recent development, reconfigurable clouds become a viable solution to overcome practical problems in clouds, such as scalability, delay, etc., by offloading computation tasks to reconfigurable hardware, FPGA. Several existing techniques, such as TCP/IP Offload Engine (TOE) and Lightweight Transport Layer (LTL), are still hard to be implemented in real-world deployment due to large overhead or stringent dependency of the underlying network. In this paper, we propose STRIDE, a novel inter-FPGA data communication protocol, to provide reliable end-to-end communication which addresses practical problems in deployment. In our design, STRIDE leverages FPGA's abilities through programming, such as precise timestamping, to make more accurate measurement on end-to-end delay and deliver more precise control in managing traffic in cloud. We implement STRIDE on a FPGA-based network experimental platform and demonstrate that STRIDE reduces various hardware resources consumption by 36% to 49% compared to TOE. Additionally, it also improves flow completion time in comparison to TOE by 2.2X. We further demonstrate STRIDE outperforms DCTCP and TCP-Vegas in OMNET simulator by up to 1.8X and 2.3X on average and 99th percentile respectively in large scale setting. Wenwen Fu, Tao Li 0008, Jialun Yang, Junnan Li 0002, Zhigang Sun 0002 |
ICC | 2 |
| 2019 | A Heterogeneous Parallel Packet Processing Architecture for NFV AccelerationabstractNetwork function virtualization (NFV) offers a new way to design, deploy and manage networking services. It is of vital importance to exploit heterogeneous parallelism between hardware and software, in order to improve virtulization performance and quality of virtualized network services. In this poster, we propose a novel heterogeneous parallel architecture that highly exploits the parallelism inside packet processing, and implementation efficacy with hardware processing engines and software threads. We present two packet processing pipelines with three implemented VNF instances to better demonstrate the efficiency of heterogeneous parallelism in accelerating NFV. We show the performance of our proposed architecture with various virtualized requirements and traffics in a well-deployed network environment. Experimental results reveal that it can achieve accelerated NFV performance, as well as provide a wide class of VNFs to improve the quality of virtualized network services. Jinshu Su, Biao Han 0003, Gaofeng Lv, Tao Li 0008, Zhigang Sun 0002 |
ICNP | 4 |
| 2019 | FAST: enabling fast software/hardware prototype for network experimentationabstractThe evolution of new technologies in network community is getting ever faster. Yet it remains the case that prototyping those novel mechanisms on a real-world system (i.e. CPU-FPGA platforms) is both time and labor consuming, which has a serious impact on the research timeliness. In order to bring researchers out of trivial process in prototype development, this paper proposed FAST, a software hardware co-design framework for fast network prototyping. With the programming abstraction of FAST, researchers are able to prototype (using C, verilog or both) a wide spectrum of network boxes rapidly based on all kinds of CPU-FPGA platforms. FAST framework takes care of managing DMA, PCIe and Linux Kernel while providing a unified API for researchers so they can focus only on the packet processing functions. We demonstrate FAST framework's easy to use features with a number of prototypes and show we can get over 10x gains in performance or 1000x better accuracy in clock synchronization compared with their software versions. Xiangrui Yang 0002, Zhigang Sun 0002, Junnan Li 0002, Jinli Yan, Tao Li 0008, Wei Quan 0004, Donglai Xu, Gianni Antichi |
IWQoS | 5 |
| 2019 | A Memory Optimized Architecture for Multi-Field Packet Classification (Brief Announcement)abstractThe high-performance hardware architectures for multi-field packet classification have been studied over the past decade. Although many FPGA-based solutions can achieve very high throughput, the limited FPGA resources severely hinders the scalability of the rulesets or matching fields. To address this issue, we present a parallel architecture named Wildcard-removed Two-dimensional Pipeline (WeeTP) to save memory usage of wildcards and reduce logic resources. WeeTP uses the Maximum Wildcard Overlap (MWO) algorithm to maximize the compression percentage by rearranging the ruleset. We implement and evaluate WeeTP on an Intel STRATIX V FPGA. Experimental results show that our approach can save 37% and 41% memory consumption on average for real 5-tuple rules and OpenFlow rules, respectively. Chenglong Li 0007, Tao Li 0008, Junnan Li 0002 |
SPAA | 2 |
| 2018 | Demonstration of Path-Based Packet Batcher for Accelerating Vectorized Packet ProcessingabstractRecently, a major challenge on generic multi-core network processing platforms is how to improve packet processing performance. Vector packet processor (VPP) is a modularized and high- performance software framework for building network dataplane applications. The key idea of VPP is to reduce instruction cache (i-cache) misses with vectorized packet processing. However, the packets in a vector may traverse different processing paths in some scenarios. In such case, the vector is split into several smaller vectors, and the per- packet overhead would increase. In this paper, we propose a Path-based Packet Batcher (PPB) to accelerate VPP. PPB is transparent to VPP, and it requires no modification to VPP. Before VPP processes packets, PPB batches the packets based on the processing paths they will traverse. We build a prototype based on FPGA to evaluate the performance optimizations to VPP with PPB. Experiment results show that the reduction of i-cache misses can be up to 57.6% when the batch size is 128. Jinli Yan, Tao Li 0008, Gaofeng Lv, Zhigang Sun 0002 |
SECON | 2 |
| 2018 | FAS: Using FPGA to Accelerate and Secure SDN Software SwitchesabstractSoftware-Defined Networking (SDN) promises the vision of more flexible and manageable networks but requires certain level of programmability in the data plane to accommodate different forwarding abstractions. SDN software switches running on commodity multicore platforms are programmable and are with low deployment cost. However, the performance of SDN software switches is not satisfactory due to the complex forwarding operations on packets. Moreover, this may hinder the performance of real-time security on software switch. In this paper, we analyze the forwarding procedure and identify the performance bottleneck of SDN software switches. An FPGA-based mechanism for accelerating and securing SDN switches, named FAS (FPGA-Accelerated SDN software switch), is proposed to take advantage of the reconfigurability and high-performance advantages of FPGA. FAS improves the performance as well as the capacity against malicious traffic attacks of SDN software switches by offloading some functional modules. We validate FAS on an FPGA-based network processing platform. Experiment results demonstrate that the forwarding rate of FAS can be 44% higher than the original SDN software switch. In addition, FAS provides new opportunity to enhance the security of SDN software switches by allowing the deployment of bump-in-the-wire security modules (such as packet detectors and filters) in FPGA. Wenwen Fu, Tao Li 0008, Zhigang Sun 0002 |
Secur. Commun. Networks | 2 |
| 2016 | Self-described buffer: A novel mechanism to improve packet I/O efficiency in LinuxabstractSocket buffer (SKB) is the standard data structure for exchanging packets and their control information between NIC driver and protocol stack. The overhead of dynamic SKB management has been considered as the significant bottleneck in packet I/O. Some novel non-SKB mechanisms, such as DPDK, were thus proposed to solve the problem. However, these mechanisms usually cannot be widely adopted in the data path of most packet forwarding applications, due to their incompatibility with SKB. In this paper, a new SKB-compatible mechanism, namely Self-described buffer (SDB), is proposed to improve the efficiency of packet I/O. SDB eliminates SKB allocation/deallocation overhead by offloading SKB management into NIC hardware. It also reduces the overhead of dynamic binding/unbinding operations existed in SKB management by statically binding related information in advance using the free space of Databuf. To evaluate the proposed approach, a SDB-enabled NIC and its driver has been designed and implemented based on FPGA. Experimental results show that the proposed SDB achieves 2× throughput compared with a traditional SKB mechanism in raw packet forwarding, and 34.75% improvement for typical network forwarding applications (e.g. IP forwarding, Bridge forwarding and SDN forwarding) on average. Jinli Yan, Zhigang Sun 0002, Tao Li 0008, Donglai Xu |
IWQoS | 4 |
| 2016 | Design and implementation of Software Defined Hardware Counters for SDN
Tao Li 0008, Biao Han 0003, Zhigang Sun 0002 |
Comput. Networks | 2 |
| 2015 | Towards high-performance packet processing on commodity multi-cores: current issues and future directions
Jinli Yan, Zhigang Sun 0002, Tao Li 0008, Minxuan Zhang |
Sci. China Inf. Sci. | 4 |
| 2014 | Demostration of Self-Described Buffer for Accelerating Packet Forwarding on Multi-core ServersabstractNetwork processing platform based on the multi-core CPU becomes more and more prevailing in nowadays. Buffer allocation/deallocation operations consume a large number of CPU cycles in packet I/O process. The problem becomes even worse in the scenario of packet forwarding, as buffer allocation/deallocation operations are more frequent than the host-based network applications. We thus propose a novel data structure for packet buffer management on multi-cores, named Self-Described Buffer (SDB), which merges the separated descriptor and metadata into packet buffer. SDB management overhead can be greatly reduced by utilizing the compact data structure, and zero-overhead buffer management can be further achieved by offloading SDB allocation/deallocation operations to NIC. We have prototyped SDB enabled NIC, named BcNIC, on NetFPGA-10G. In the demo, we will illustrate the advantages of the SDB scheme by comparing the performance of BcNIC with the traditional NIC on multi-core platforms. Zhigang Sun 0002, Tao Li 0008, Biao Han 0003, Gaofeng Lv |
CloudCom | 3 |
| 2014 | The Demonstration of Hyper Software Defined Hardware CountersabstractSoftware Defined Networking (SDN) provides efficient network and traffic management for data center network. As underlying devices in SDN, SDN switches must maintain a large number of hardware counters. Implementation of these counters faces serious challenges for SDN switches, i.e., High memory consumption and inflexibility. Thus, we previously proposed Software Defined Hardware Counters (SDHC), which decouples definition and implementation of counters to overcome these challenges. However, like traditional hardware counters, SDHC only supports passive statistical mode (i.e., The values of the counters can be only read passively by the controller). Based on the passive mode, most of applications need to send request messages at some frequency to obtain statistics, which causes some critical problems for SDN: i) low statistical accuracy, ii) high network bandwidth consumption. Hyper Software Defined Hardware Counters (Hyper SDHC) is thus proposed by extending SDHC. Through introducing the timer-triggering and updating-triggering statistics-reporting mechanisms, Hyper SDHC can naturally support active statistical mode, i.e., Counters actively report their values according to triggering condition. It can greatly enhance the statistical accuracy and reduce network bandwidth consumption between controller and switch. The demo of Hyper SDHC is implemented based on Net Magic platform. The demo will exhibit how Hyper SDHC works and how it supports a typical video quality monitor application. Tao Li 0008, Biao Han 0003, Zhigang Sun 0002 |
CloudCom | 2 |
| 2014 | Design of Software Defined hardware counters for SDNabstractImplementation of counters is a critical challenge for switches in today's Software-Defined Networking (SDN). In this paper, we address the current challenges in implementing SDN counters: high memory consumption, low utilization, and inflexibility. We introduce the concept of software defined hardware counters (SDHCs) for SDN. Our main idea is to make the switch-local CPU flexibly allocate memory space to each counter required by controllers. The ASIC of SDN switches transmits event records to the CPU, which contain updating information of the counters. Furthermore, the ASIC provides non-semantic counter memory space to be allocated by the CPU. Based on the proposed SDHC, an SDN controller can flexibly apply/release various counters for each counter category (e.g., each flow entry, each port) through the south-bound interface. It is shown that SDHC achieves high flexibility while reducing the memory space on ASIC. It also improves the update performance through alleviating the CPU overhead. Finally, we evaluate the performance of SDHC through comprehensive simulation study. Tao Li 0008, Biao Han 0003, Zhigang Sun 0002 |
LANMAN | 2 |
| 2011 | Using NetMagic to observe fine-grained per-flow latency measurementsabstractWe introduce NetMagic to demonstrate the efficacy of RLI architecture RLI for the fine-grained per-flow latency measurements. In this demo, the main function of RLI is implemented in NetMagic, which is the key component of our experimental network comprising several computers and switches. We are going to show how NetMagic can provide rapid implementation and evaluation of RLI architecture that is difficult with commercial switch or router platforms. In the demo, the estimated fine-grained per-flow latency by RLI is monitored and dynamically presented. Further, the true latency with a resolution of 8ns is also provided by NetMagic for the evaluation. The efficacy of RLI architecture can be observed in a real-time fashion by the difference between estimated latencies and true ones. Tao Li 0008, Zhigang Sun 0002, Chunbo Jia, Myungjin Lee |
SIGCOMM | 1 |
| 2010 | Selecting profitable custom instructions for reconfigurable processors
Tao Li 0008, Jigang Wu, Siew-Kei Lam, Thambipillai Srikanthan, Xicheng Lu |
J. Syst. Archit. | 1 |
| 2009 | Fast enumeration of maximal valid subgraphs for custom-instruction identificationabstractExtensible processors are increasingly becoming popular as they allow for incorporating custom instructions to meet design constraints. However, identifying custom instructions under architectural input/output ports constraint is a time consuming process particularly when large applications are considered. To rapidly identify the most profitable custom instructions with large inputs and outputs, this paper proposes a novel identification algorithm for enumerating maximal convex subgraphs containing no invalid node (i.e., maximal valid subgraphs). The proposed enumerating strategy is based on divide-and-conquer with a top-down manner, rather than the bottom-up manner utilized in the state-of-the-art. The division operation only considers invalid inner nodes of the given DFG, rather than taking all the invalid nodes into account, and thus accelerates enumeration of the maximal valid subgraphs. Experimental results show that, the improvement over the latest work is more than 90% for 60% DFG instances of the acknowledged benchmarks. Tao Li 0008, Zhigang Sun 0002, Jigang Wu, Xicheng Lu |
CASES | 1 |