VLDB 2026 Research / reviewers in the wild / expert
Ning Weng
dblp:24/3415
· DBLP profile ↗
32ranked-venue papers
4as first author
2since 2021 · last 2024
0000-0003-4869-8383ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 14 · 2 first-authorSystems, architecture and hardware · 7 · 1 first-authorSecurity and privacy · 5 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorSoftware engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | ProGen: Projection-Based Adversarial Attack Generation Against Network Intrusion DetectionabstractAdversarial attacks, widely recognized as significant threats to machine learning (ML) models in computer vision and natural language processing, can have more severe consequences when targeting ML-based Network Intrusion Detection Systems (NIDS). These attacks, characterized by data manipulation, necessitate a focused investigation grounded in the unique attributes of the data and practical constraints inherent to the target scenario, as opposed to indiscriminately applying methodologies borrowed from other domains. Since network traffic is complex unstructured data, ML models are commonly used in existing studies to explore how perturbations can defeat ML-based IDS. However, two challenges persist in the realm of traffic-space adversarial attack generation. First, raw traffic data cannot be directly input into ML models. Second, determining the appropriate perturbation scale and direction is challenging, particularly in the case of multi-class NIDS. In this work, we propose a projection-based adversarial attack generation framework, ProGen, to address these two challenges. ProGen is inspired by two observed characteristics of the NIDS scenario: flexible representation and clear objective. ProGen uses a basic feature sequence (BFS) space to represent network traffic in a way that aligns with realistic requirements. To achieve a clear objective, ProGen utilizes a traffic space generative adversarial network (GAN) to approximate distribution mapping between malicious traffic and benign traffic. To better apply the generative model for adversarial attacks, we further design constraints to preserve the functions of the adversarial traffic. We’ve successfully demonstrated the effectiveness of ProGen on six common ML models using the CSE-CIC-IDS2018, CIC-IDS-2017, and UNSW-NB15 datasets; however, we’re yet to validate these findings in real network environments. We visualize the generated distributions of the BFS elements to illustrate the projecting effect under the designed realistic constraints. The results of attack effectiveness tests show that attacks generated from ProGen can significantly reduce the detection performance across different ML models. Minxiao Wang, Ning Yang 0009, Nicolas J. Forcade-Perkins, Ning Weng |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | K-GetNID: Knowledge-Guided Graphs for Early and Transferable Network Intrusion DetectionabstractDeveloping early and transferable Network Intrusion Detection Systems (NIDSs) is essential for robust network security. Early detection prevents further damage, while transferable NIDSs enables reuse across diverse networks. For Machine Learning (ML) and Deep Learning (DL)-based NIDSs, transferability significantly reduces data collection and annotation costs for timely attack mitigation. Current DL-based early intrusion detection studies often focus on identifying attacks from the first few packets, neglecting the crucial aspect of adjustable early detection. Additionally, most DL-based NIDS methods overlook transferability during both design and evaluation phases. To address these limitations, we propose K-GetNID, a knowledge-guided graph learning-based NIDS that excels in both early and transferability. We introduce a Heterogeneous Temporal Graph (HTGraph) to represent the dynamic feature series of network flows, providing enough information for early detection. Additionally, we construct this HTGraph format based on prior knowledge about feature types and correlations to assist the neural network in learning general and transferable knowledge for intrusion detection. We develop a corresponding Heterogeneous Temporal Graph Neural Network (HTGNN) model to learn from the HTGraph format. Furthermore, an Adjustable Early Detection Decoder is designed to enhance the generalization of the proposed model to the input distribution shifts caused by early detection. Experiments on CIC-IDS-2017 and UNSW-NB15 datasets show that K-GetNID matches the performance of deep learning methods, excelling in adjustable early intrusion detection and transferability. Minxiao Wang, Ning Yang 0009, Ning Weng |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2020 | Multibit Tries Packet Classification with Deep Reinforcement LearningabstractHigh performance packet classification is a key component to support scalable network applications like firewalls, intrusion detection, and differentiated services. With ever increasing in the line-rate in core networks, it becomes a great challenge to design a scalable and high performance packet classification solution using hand-tuned heuristics approaches. In this paper, we present a scalable learning-based packet classification engine and its performance evaluation. By exploiting the sparsity of ruleset, our algorithm uses a few effective bits (EBs) to extract a large number of candidate rules with just a few of memory access. These effective bits are learned with deep reinforcement learning and they are used to create a bitmap to filter out the majority of rules which do not need to be full-matched to improve the online system performance. Moreover, our EBs learning-based selection method is independent of the ruleset, which can be applied to varying rulesets. Our multibit tries classification engine outperforms lookup time both in worst and average case by 55% and reduce memory footprint, compared to traditional decision tree without EBs. Hasibul Jamil, Ning Weng |
HPSR | 2 |
| 2019 | Scalable Many-Field Packet Classification for Traffic Steering in SDN SwitchesabstractPacket classification is a key function for software-defined networking switches. For example, OpenFlow Switch examines up to 15 required fields, against thousands of rules. With the proliferation of new header fields in a ruleset, it becomes a great challenge to design a high performance packet classification solution. In this paper, we present a scalable many-field packet classification algorithm and its prototype implementation on a graphics processing unit (GPU). By exploiting the sparsity of ruleset, our algorithm uses a few effective bits (EB) to divide a large ruleset into multiple subsets with low rule replication for economic memory usage at the offline stage. These EB are chosen based on our selection metrics: wildcard ratio, independence index, and diversity index. Using these EB, our algorithm can quickly filter out the majority rules which do not need to be full-matched to improve the online system performance. Moreover, the choice of EB is adjustable to meet the implementation requirements of varying environments for good performance scalability. Our prototype on a single NVIDIA K20C GPU achieves more than 160 MPPS throughput for 100K 15-field synthetic ruleset. Cheng-Liang Hsieh, Ning Weng, Wei Wei 0020 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2017 | NF-switch: VNFs-enabled SDN switches for high performance service function chainingabstractCurrent Software-Defined Networking (SDN) switch examines many packet headers to support the flow-based packet forwarding. To support the application-based forwarding for a service function chaining (SFC), SDN switch requires packet header modifications to identify the processing status and multiple packet matchings to steer the network traffics to different VMs in a specific order. Both challenges require significant computation resources in a system and result in severe system performance degradation. To improve the system performance and keep the flexibility, this paper proposes NF-Switch to eliminate the requirement of packet modifications and to reduce the number of matchings for the application-based forwarding. Compared to the native implementation, our experimental results show that NF-Switch reduces the system processing latency to one quarter and increases the system throughput about 3 times for a SFC with 10 network functions (NFs). Moreover, the proposed solution maintains the content switching time to update the implemented SFC for a better system scalability regarding to the number of NFs in a SFC. Cheng-Liang Hsieh, Ning Weng |
ICNP | 2 |
| 2016 | Many-Field Packet Classification for Software-Defined Networking SwitchesabstractPacket classification is a core problem for OpenFlow-based software-defined networking switches, which required 38 packet header fields per flow to be examined against thousands of rules in a ruleset. With the trend of continue growing number of fields in a rule and the number of rules in rule set, it will be a great challenge to design a high performance packet classification solution with the capability to easy update new rule and fields. In this paper, we present a scalable many-field packet classification algorithm with varying rulesets and its prototype implementation on a graphics processing unit. The proposed algorithm constructs multiple lookup tables and merges partial lookup results for a small ruleset to accelerate the overall packet classification process by using effective bit positions in a ruleset with three selecting metrics: wildcard ratio, independence index, and diversity index. Those lookup tables made with effective bit positions are flat with a low rule replication ratio. Besides, they are adjustable to meet different implementation environments for a good performance scalability between different ruleset sizes. Our prototype on a single NVIDIA K20C GPU achieves 198 MPPS, 186 MPPS, 163 MPPS throughput for 1K, 32K, and 100K 15-field ruleset. Cheng-Liang Hsieh, Ning Weng |
ANCS | 2 |
| 2016 | Virtual Network Functions Instantiation on SDN Switches for Policy-Aware Traffic SteeringabstractSoftware-Defined Networking (SDN) provides the capability to steer traffic in a network to lower the management cost. Network Function Virtualization (NFV) gives the chance to implement network functions at the right time and the right place to increase operation flexibility. Together SDN and NFV show the potential to create an agile system with a low operations cost and a high customer satisfaction. However, the combination of SDN with NFV results in the redundant packet forwarding traffic inside SDN in order to forward packets based on the deployed network functions for a service chain. Besides, it also increase the computation requirement of a controller for possible packet header modifications and flow states management. In this paper, we propose to implement network functions on SDN switches to lower the traffic inside of SDN and the computation requirement of a SDN controller. We create network function modules for open virtual switches and make those functions to be managed by a controller with an algorithm to streamline the implemented service chains. Our results show that the proposed system can reduce about 2/3 of current network traffic compared to the current solutions without the modification of forwarding tables and packets. Cheng-Liang Hsieh, Ning Weng |
ANCS | 2 |
| 2016 | A high-throughput DPI engine on GPU via algorithm/implementation co-optimization
Cheng-Liang Hsieh, Lucas Vespa, Ning Weng |
J. Parallel Distributed Comput. | 3 |
| 2015 | Scalable Many-Field Packet Classification using Multidimensional-Cutting Via Selective Bit-ConcatenationabstractOpenFlow Switch in Software-Defined Networking (SDN) has changed packet classification from standard 5-tuple to arbitrary many-field. The growing number of fields in a rule and the increasing number of rules in a ruleset poses great challenges for packet classification in terms of performance, storage, and update cost. In this paper, we design a two-stage packet classification system to address those issues by exploiting ruleset sparsity and rule fields independence. A ruleset is examined offline with proposed matrices to find representative bits from different field in a rule. We leverage those representative bits and concatenate them as sample values to divide a ruleset into several subsets in sample spaces. Each subset is given a unique address for each sample space. A ruleset update only affects those related addresses. The proposed pre-filtering stage comes out only highly related rules by intersecting candidate rules from different sample spaces for full match process. Out system throughput is 356 MPPS for 1K 15-field rules and 213 MPPS for 100K 15-field rules when using a single NVIDIA K20C GPU card. Cheng-Liang Hsieh, Ning Weng |
ANCS | 2 |
| 2014 | High performance multi-field packet classification using bucket filtering and GPU processingabstractThe literature review shows a trend to arbitrary number of multi-field packet classification is evolved from standard 5-tuple matching to support new applications like OpenFlow switch which processes upto 15 fields \cite{Yun:ISCAHPP13}. However, arbitrary number of multi-field packet classification becomes a great challenge regarding to performance, memory requirement, and update cost. In this paper, a high performance multi-field packet classification system is designed and implemented using bucket filtering and GPU processing. Cheng-Liang Hsieh, Ning Weng |
ANCS | 2 |
| 2012 | Quality-of-information modeling and adapting for delay-sensitive sensor network applicationsabstractAcceptable Quality-of-Information (QoI) is essential for sensor network applications such as infrastructure health monitoring because it directly impacts public safety. However, it is a challenging problem to attain good QoI of sensor applications due to unpredictable environment noise, unreliable network communication, and varying requirements for wide variety of sensor applications. We believe the first step to addressing this challenge is to develop an application-independent QoI model. In this paper, we present a fundamental quality-of-information model based on signal-to-noise ratio. Our model addresses information quality by considering sensor measurement quality, network quality and sensor application requirements by end users. Furthermore, we develop a quality-aware scheduling framework which exploits an analytical queue model to calculate and adapt sensor node sampling rate and base station scheduling priority in order to optimize overall quality. Our results show that sensor measurement quality in terms of sampling rate, network quality in terms of loss rate and delay, all play significant roles in impacting overall quality. A QoI-aware scheduler thus is an effective approach to quantify information quality and adapt for unpredictable sensor networks. Mini Mathew, Ning Weng, Lucas Vespa |
IPCCC | 2 |
| 2011 | MS-DFA: Multiple-Stride Pattern Matching for Scalable Deep Packet InspectionabstractAs network speeds continue to increase, so does the need for scalable pattern matching for deep packet scanning applications such as signature-based network intrusion detection. Multiple-stride deterministic finite automaton (DFA) increases the performance of pattern matching, because they allow multiple bytes of a packet to be scanned simultaneously. However, traditional multiple-stride DFA either rely on specific hardware for parallel comparison or have a huge memory requirements due to state explosion. In this paper, we present a high throughput, multiple-stride pattern-matching architecture that requires a small storage cost and no specific hardware. The basic idea is to group DFA states/transitions into three coarse-grained and variable-size blocks, so that each individual block can employ different-specific methods to optimize storage requirements and performance. The blocks are naturally identified based on basic observations of DFA characteristics: prefix, linear trie and state dependencies. The performance evaluation is done using the Snort pattern sets. We show that multi-byte striding DFA achieves multi Gb/s pure content inspection in software, while utilizing <3 bytes per pattern character. Lucas Vespa, Ning Weng, Ramaswamy Ramaswamy |
Comput. J. | 2 |
| 2011 | Deep packet pre-filtering and finite state encoding for adaptive intrusion detection system
Ning Weng, Lucas Vespa, Benfano Soewito |
Comput. Networks | 1 |
| 2011 | Information quality model and optimization for 802.15.4-based wireless sensor networks
Ning Weng, I-Hung Li, Lucas Vespa |
J. Netw. Comput. Appl. | 1 |
| 2011 | Hybrid pattern matching for trusted intrusion detectionabstractAbstract Intrusion Detection Systems (IDSs) rely on pattern matching to detect and thwart a network attack by comparing packets with a database of known attack patterns. The key requirements of trusted intrusion detection are accurate pattern matching, adaptive, and reliable reconfiguration for new patterns. To address these requirements, this paper presents a trusted intrusion detection by utilizing hybrid pattern matching engines: FPGA‐based and multicore‐based pattern matching engine. To achieve synchronization of these two pattern matching engines, methodologies including multi‐threading DFA and clustered state coding have been developed. These hybrid pattern matching engines increases the reliability and trustworthy of intrusion detection systems. Copyright © 2009 John Wiley & Sons, Ltd. Benfano Soewito, Lucas Vespa, Ning Weng, Haibo Wang 0005 |
Secur. Commun. Networks | 3 |
| 2011 | Deterministic finite automata characterization and optimization for scalable pattern matchingabstractMemory-based Deterministic Finite Automata (DFA) are ideal for pattern matching in network intrusion detection systems due to their deterministic performance and ease of update of new patterns, however severe DFA memory requirements make it impractical to implement thousands of patterns. This article aims to understand the basic relationship between DFA characteristics and memory requirements, and to design a practical memory-based pattern matching engine. We present a methodology that consists of theoretical DFA characterization, encoding optimization, and implementation architecture. Results show the validity of the characterization metrics, effectiveness of the encoding techniques, and efficiency of the memory-based pattern engines. Lucas Vespa, Ning Weng |
ACM Trans. Archit. Code Optim. | 2 |
| 2010 | Quality-aware scheduling metrics for adaptive sensor networksabstractQuality of service in sensor networks is a difficult problem due to unpredictable environment noise, unreliable network communication and varying requirements for wide varieties of applications. In this paper, we present fundamental quality of information metrics using signal-to-noise ratio. These metrics address information quality under varying sensing environments, noise and network bandwidth, and are completely application independent. We use these metrics to develop a quality-aware scheduling system (QSS) which exploits cross-layer control of sensors to effectively schedule data sensing and forwarding. Particularly, we develop and evaluate several QSS scheduling mechanisms: passive, reactive and perceptive. These mechanisms can adapt to environment noise and bandwidth variation by dynamically changing sensor rates. Our results indicate that our QSS is a novel and effective approach to improve the QoS for sensor networks. Lucas Vespa, Ning Weng |
LCN | 2 |
| 2009 | Testbed for evaluating worm containment systemsabstractDangerous worms like CodeRed or Slammer can spread millions of probe packets in just seconds which can result in thousands of infected hosts and large losses. Fast and effective containment strategies are crucially important to protect the Internet Infrastructure. Toward this goal of fast and effective worm containment, different techniques have been presented such as address blacklisting and content filtering [3], anomaly detection [6] and signature-based detection [5]. Meanwhile recently developed worm models [1] enable us to develop a testbed to accurately and quickly evaluate the efficiency of these defense mechanisms. In this paper, we present a testbed which utilizes software agents to allow large scale simulation with individual host functionality. We utilize this testbed to evaluate our containment systems in terms of security and performance tradeoff. Ritam Chakrovorty, Lucas Vespa, Ning Weng |
ANCS | 3 |
| 2009 | Theoretic analysis of finite automata for memory-based pattern matchingabstractIn the midst of vastly numbered and quickly growing internet security threats, Network Intrusion Detection System (NIDS) [2] becomes more important to network security every day. Vital to effective NIDS is a multi-pattern matching engine which requires deterministic performance and adaptability to new threats [3]. Memory-based Deterministic Finite Automata (DFA) are ideal for pattern matching but have severe memory requirements [1] that make them difficult to implement. Many previous heuristic techniques have been proposed to reduce memory requirements, however in this paper, we aim to effectively understand the basic relationship between DFA characteristics and memory, in order to create minimal memory DFA implementations. We show what DFA characteristics either cause or reduce memory requirements, as well as how to optimize DFA to exploit those characteristics. Specifically, we introduce the concepts of State Independence and State Irregularity, which are DFA characteristics that can reduce memory waste and allow for memory reuse. Furthermore, we introduce DFA normalization which optimizes DFA to fully exploit these characteristics. Altogether this work serves as a source for how to extract and utilize DFA characteristics to create minimal memory implementations. Lucas Vespa, Ning Weng |
ANCS | 2 |
| 2009 | P3FSM: Portable Predictive Pattern Matching Finite State MachineabstractSignature-based network intrusion detection requires fast and reconfigurable pattern matching for deep packet inspection. In our previous work we address this problem with a hardware based pattern matching engine that utilizes a novel state encoding scheme to allow memory efficient use of Deterministic Finite Automata. In this work we expand on these concepts to create a completely software based system, P3FSM, which combines the properties of hardware based systems with the portability and programmability of software. Specifically we introduce two methods, character aware and SDFA, for encoding predictive state codes which can forecast the next states of our FSM. The result is software based pattern matching which is fast, reconfigurable, memory-efficient and portable. Lucas Vespa, Mini Mathew, Ning Weng |
ASAP | 3 |
| 2009 | Predictive Pattern Matching for Scalable Network Intrusion Detection
Lucas Vespa, Mini Mathew, Ning Weng |
ICICS | 3 |
| 2009 | Deterministic Finite Automata Characterization for Memory-Based Pattern Matching
Lucas Vespa, Ning Weng |
ICICS | 2 |
| 2009 | Concurrent workload mapping for multicore security systemsabstractAbstract Multicore based network processors are promising components to build real‐time and scalable security systems to protect the networks and systems. The parallel nature of the processing system makes it challenging for application developers to concurrently program security systems for high performance. In this paper we present an automatic programming methodology that considers application complexity, traffic variation, and attack signatures update. In particular, our mapping algorithm concurrently takes advantage of parallelism in the level of tasks, applications, and packets to achieve optimal performance. We present results that show the effectiveness of the analysis, mapping, and the performance of the model methodology. Copyright © 2009 John Wiley & Sons, Ltd. Benfano Soewito, Ning Weng |
Concurr. Comput. Pract. Exp. | 2 |
| 2009 | Analysis of network processing workloads
Ramaswamy Ramaswamy, Ning Weng, Tilman Wolf |
J. Syst. Archit. | 2 |
| 2009 | Analytic modeling of network processors for parallel workload mappingabstractNetwork processors are heterogeneous system-on-chip multiprocessors that are optimized to perform packet forwarding and processing tasks at Gigabit data rates. To meet the performance demands of increasing link speeds and complex network applications, network processors are implemented with several dozen embedded processor cores and hardware accelerators that run multiple packet processing applications in parallel. The parallel nature of the processing system makes it increasingly difficult for application developers to understand and manage resources and map processing tasks to the hardware. To address this problem, we present a methodology for profiling and analyzing network processor applications, mapping processing tasks to a generalized network processor architecture, and analytically determining the expected throughput performance. The key novelty of this work is not only the adaptation of application analysis and mapping algorithms to heterogeneous network processors, but also that the entire process can be automated and hidden from the application developer. Starting with the analysis of a uniprocessor implementation of the application, the process yields a mapping of the partitioned application that shows best performance for a given network processor system. The simplicity of the proposed randomized mapping algorithm allows the use of this methodology in network processor runtime systems where dynamic reallocation of tasks is necessary but processing power is limited. We present results that show the effectiveness of the analysis and mapping methodology as well as its application to design space exploration. Ning Weng, Tilman Wolf |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2008 | Methodology for evaluating string matching algorithms on multiprocessorabstractThe Internet is suffering caused by the lacking of security. One of the most promising ways to provide security is Intrusion Detection Systems (IDSs). The heart of almost every IDSs is a string matching algorithm, which is a very computational intensive task. Network Processors (NPs), a specialized multiprocessor, can provide flexibility and high performance for string matching. This paper evaluates several key string matching algorithms using a comprehensive simulation framework. Starting from a uniprocessor profiling, the framework constructs task graphs for string matching algorithms. Then task graphs are mapped onto NPs together with other network applications. The system throughput is determined by the analytical performance model. With this framework, we can evaluate the performance of different string matching algorithms on NPs. Our results show that shift table based algorithms (SFKSearch and Wu-Manber) and finite automaton based Aho-Corasick are complementary: SFKSearch and Wu-Manber do better job in NPs for good packet and larger pattern length due to better inter-task parallelism and shifting; Aho-Corasick does not depend on minimal pattern length and shows relative small processing cost variation between bad and good packets. Benfano Soewito, Ning Weng |
AICCSA | 2 |
| 2008 | Mapping task graphs onto Network Processors using genetic algorithmabstractNetwork Processors (NPs) are embedded system-on-a-chip multiprocessors that are optimized to perform simple packet processing tasks at data rates of several Gigabytes per second. They are the key components to build a performance-scalable and function-flexible network systems. To meet the performance demands of increasing link speeds and more complex network applications, NPs are implemented with several dozen of processor cores and run multiple packet processing applications in parallel. This trend makes it increasingly difficult for application developers to program NPs for high performance. This paper presents an automated task scheduling technique to address this parallel programming complexity. Our proposed technique is based on GA. By incorporating tasks dependency into scheduling list and encoding task scheduling list as a chromosome, GA can quickly remove the invalid mappings and evolve to the high quality solutions. This technique takes advantage of task-level and application-level parallelism to maximize system performance for a given NPs architecture. The simulation results show that this proposed technique can generate high quality mapping comparing to other heuristics by mapping some sample network applications. This work will also enable researchers and engineers to systematically evaluate and quantitatively understand the NPs system issues including application partitioning, architecture organizing, workload mapping and run-time operating. Ning Weng, Nandeesh Kumar, Satish Dechu, Benfano Soewito |
AICCSA | 1 |
| 2008 | Implementing high-speed string matching hardware for network intrusion detection systemsabstractThis paper presents a string matching hardware on FPGA for network intrusion detection systems. The proposed architecture, consisting of packet classifiers and strings matching verifiers, achieves superb throughput by using several mechanisms. First, based on incoming packet contents, the packet classifiers can dramatically reduce the number of strings to be matched for each packet and, accordingly, feed the packet to a proper verifier to conduct matching. Second, a novel multi-threading finite state machine (FSM) is proposed, which improves FSM clock frequency and allows multiple packets to be examined by a single FSM simultaneously. Design techniques for high-speed interconnect and interface circuits are also presented. Experimental results are presented to explore the trade-offs between system performance, strings partition granularity and hardware resource cost Atul Mahajan, Benfano Soewito, Sai K. Parsi, Ning Weng, Haibo Wang 0005 |
FPGA | 4 |
| 2007 | Methodology for Evaluating DNA Pattern Searching Algorithms on MultiprocessorabstractPattern matching has been one of the major operations in modern bioengineering especially in Bioinformatics. Prior work on this area have focus on either pursuing mathematically efficient matching algorithms or hardwired approach. As multicore processor are becoming mainstream, developers need to determine how to take advantage of multicore technology for pattern matching. In this paper, we propose a methodology to evaluate pattern search algorithms for DNA on Multiprocessor. Our evaluation methodology is an automatic simulation framework. Starting from a uniprocessor profiling, the framework constructs task graphs for string matching algorithms. Then task graphs are mapped onto multiprocessor. The system's performance is determined by the analytical performance model. With this framework, we can evaluate the performance of different algorithms on multiprocessor. Our case studies show that finite automaton based (Aho-Corasick) is more efficient than shift table based algorithms (SFKSearch and Wu-Manber) on uniprocessor, however, Wu-Manber is 3 times efficient than Aho-Corasick on multiprocessor due to its inherent parallelism. Benfano Soewito, Ning Weng |
BIBE | 2 |
| 2005 | Design considerations for network processor operating systemsabstractNetwork processors (NPs) promise a flexible, programmable packet processing infrastructure for network systems. To make full use of the capabilities of network processors, it is imperative to provide the ability to dynamically adapt to changing traffic patterns and to provide run-time support in the form of a network processor operating system. The differences to existing operating systems and the main challenges lie in the multiprocessor nature of NPs, their on-chip resources constraints, and the real-time processing requirements. In this paper, we explore the key design tradeoffs that need to be considered when designing a network processor operating system. In particular, we explore the performance impact of (1) application analysis for partitioning, (2) network traffic characterization, (3) workload mapping, and (4) run-time adaptation. We present and discuss qualitative and quantitative results in the context of a particular application analysis and mapping framework, but the observations and conclusions are generally applicable to any run-time environment for network processors. Tilman Wolf, Ning Weng, Chia-Hui Tai |
ANCS | 2 |
| 2005 | Analysis of Network Processing WorkloadsabstractNetwork processing is becoming an increasingly important paradigm as the Internet moves towards an architecture with more complex functionality inside the network. Modern routers not only forward packets, but also process headers and payloads to implement a variety of functions related to security, performance, and customization. It is important to get a detailed understanding of the workloads associated with this processing in order to be able to develop efficient network processing engines. We present a tool called PacketBench, which provides a framework for implementing network processing applications and obtaining an extensive set of workload characteristics. PacketBench provides the support functions to handle various packet traces and manage packet memory. For statistics collection, PacketBench provides the ability to derive a number of microarchitectural and networking related metrics. The understanding of workload details of network processing has many practical applications. As network processing systems move towards highly parallel embedded systems, it is becoming increasingly important to explore the processing requirements of individual packets rather than averaged statistics. We show a range of workload results that focus on individual packets and the variation between them Ramaswamy Ramaswamy, Ning Weng, Tilman Wolf |
ISPASS | 2 |
| 2004 | Characterizing network processing delayabstractComputer networks have progressed from a simple store-and-forward medium to a complex communication infrastructure. Routers in the network need to implement a variety of functions ranging from simple packet classification for forwarding and firewalling to complex payload modifications for encryption and content adaptation. As these functions increase in number and complexity, more processing time is required, and packets experience a significant processing delay. In most network simulations, this delay has not been addressed because it was considered negligible. However, we show that this network processing delay can reach the magnitude of long-distance propagation delay and thus becomes a significant contributor to the overall packet delay. We evaluate different network applications and develop a model that characterizes packet processing cost with only a few parameters that can easily be derived from our simulations. To validate our simulation and our model, we compare them to actual network measurements. The contributions of this work can be used to increase the accuracy of network simulations and improve network performance estimations. Ramaswamy Ramaswamy, Ning Weng, Tilman Wolf |
GLOBECOM | 2 |