John W. Lockwood

dblp:31/1714 · DBLP profile ↗
← Back
54ranked-venue papers
7as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 35 · 5 first-authorComputer networks · 18 · 2 first-authorSecurity and privacy · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
9 papers
Routing and switching · 55% Internet architecture and protocols · 19% Software-defined and programmable networks · 15%
Computer architecture, parallel and distributed computing, and storage systems
11 papers
Reconfigurable computing and FPGAs · 74% Processor architecture and microarchitecture · 12% GPUs and heterogeneous computing · 9%
Network and information security
3 papers
Network security · 100%

Topics — the 27 heaviest of 32, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Routing and switching
IP lookup
0.132005
Shape Shifting Tries for Faster IP Route Lookup · ICNP 2005
Scalable IP lookup for Internet routers · IEEE J. Sel. Areas Commun. 2003
Scalable IP Lookup for Programmable Routers · INFOCOM 2002
Reconfigurable computing and FPGAs › FPGA-based network processing
FPGA-based pattern matching
0.122006
Fast and Scalable Pattern Matching for Network Intrusion Detection Systems · IEEE J. Sel. Areas Commun. 2006
Context-free-grammar based token tagger in reconfigurable devices · FPGA 2006
Network security › intrusion detection and prevention
intrusion detection
0.122006
Fast and Scalable Pattern Matching for Network Intrusion Detection Systems · IEEE J. Sel. Areas Commun. 2006
Efficient packet classification for network intrusion detection using FPGA · FPGA 2005
Reconfigurable computing and FPGAs
FPGA-based network processing
0.122005
Efficient packet classification for network intrusion detection using FPGA · FPGA 2005
A framework for rule processing in reconfigurable network systems (abstract only) · FPGA 2005
Internet architecture and protocols
overlay networks
0.112007
Supercharging planetlab: a high performance, multi-application, overlay network platform · SIGCOMM 2007
Processor architecture and microarchitecture › special-purpose processor
network processor
0.112007
Supercharging planetlab: a high performance, multi-application, overlay network platform · SIGCOMM 2007
Network security › intrusion detection and prevention › intrusion detection › pattern matching
multi-pattern matching
0.112006
Fast and Scalable Pattern Matching for Network Intrusion Detection Systems · IEEE J. Sel. Areas Commun. 2006
Reconfigurable computing and FPGAs
reconfigurable computing
0.112006
Fast and Scalable Pattern Matching for Network Intrusion Detection Systems · IEEE J. Sel. Areas Commun. 2006
Software-defined and programmable networks › programmable data plane
FPGA-based packet processing
0.122001
Reprogrammable network packet processing on the field programmable port extender (FPX) · FPGA 2001
Field programmable port extender (FPX) for distributed routing and queuing · FPGA 2000
Software-defined and programmable networks
programmable data plane
0.122001
Reprogrammable network packet processing on the field programmable port extender (FPX) · FPGA 2001
Field programmable port extender (FPX) for distributed routing and queuing · FPGA 2000
Routing and switching › IP lookup
hash-based lookup
0.112005
Fast hash table lookup using extended bloom filter: an aid to network processing · SIGCOMM 2005
Routing and switching › IP lookup
IPv6 lookup
0.112005
Shape Shifting Tries for Faster IP Route Lookup · ICNP 2005
Internet architecture and protocols › packet processing
packet classification
0.112005
Fast hash table lookup using extended bloom filter: an aid to network processing · SIGCOMM 2005
Routing and switching › data plane
router data plane
0.112005
Fast hash table lookup using extended bloom filter: an aid to network processing · SIGCOMM 2005
Internet of things and sensor networks › wireless sensor network
sensor fusion
0.112005
Sensor fusion and correlation · SenSys 2005
Network security
intrusion detection and prevention
0.112005
A framework for rule processing in reconfigurable network systems (abstract only) · FPGA 2005
GPUs and heterogeneous computing › packet processing
packet classification
0.112005
Efficient packet classification for network intrusion detection using FPGA · FPGA 2005
Routing and switching › packet switch
router
0.012003
Scalable IP lookup for Internet routers · IEEE J. Sel. Areas Commun. 2003
Routing and switching
routing
0.012000
Field programmable port extender (FPX) for distributed routing and queuing · FPGA 2000
Automata and formal languages › formal grammars
context-free grammar
0.012006
Context-free-grammar based token tagger in reconfigurable devices · FPGA 2006
Routing and switching
input-queued switch
0.011997
A High-Performance OC-12/OC-48 Queue Design Prototype for Input-buffered ATM Switches · INFOCOM 1997
Internet architecture and protocols
quality of service
0.011997
A High-Performance OC-12/OC-48 Queue Design Prototype for Input-buffered ATM Switches · INFOCOM 1997
Routing and switching
switching
0.011997
A High-Performance OC-12/OC-48 Queue Design Prototype for Input-buffered ATM Switches · INFOCOM 1997
Network measurement and analytics
flow state management
0.012005
Fast hash table lookup using extended bloom filter: an aid to network processing · SIGCOMM 2005
Routing and switching › data plane › router data plane
high-speed packet processing
0.012005
Shape Shifting Tries for Faster IP Route Lookup · ICNP 2005
Hardware accelerators and domain-specific architectures › network function acceleration
TCAM-based classification
0.012005
Efficient packet classification for network intrusion detection using FPGA · FPGA 2005
Reconfigurable computing and FPGAs
FPGA prototyping
0.011997
A High-Performance OC-12/OC-48 Queue Design Prototype for Input-buffered ATM Switches · INFOCOM 1997

Methods — techniques the papers use, named apart from their topics

on-chip embedded memory · 0.1multihashing · 0.1bloom filter · 0.1tree-bitmap · 0.1string matching · 0.1regular expression matching · 0.1bit vector algorithm · 0.1TCAM · 0.1trie partitioning · 0.1multiple hashing · 0.1extended bloom filter · 0.1dynamic programming · 0.1tree bitmap algorithm · 0.0partial bitstream generation · 0.0bitfile transformation · 0.0SRAM · 0.0CAM · 0.0dynamic module loading · 0.0
YearPublicationVenuePosition
2015 Scalable Key/Value Search in Datacenters
abstract
Key/Value Store (KVS) is a fundamental service used widely in modern data centers to associate keys with data values. KVS systems, such as Redis, Memcached, and Dynamo DB have traditionally been implemented with software and run on clusters of microprocessor-based servers. In this work an alternate approach is taken that performs KVS with gate ware in Field Programmable Gate Array (FPGA) logic. We leverage an efficient, open-standard, binary message format to transfer keys and values over Ethernet. Results of three different implementations of this KVS were compared -- software running on a Linux server with network data sent over UDP/IP sockets, kernel bypass using Intel's Data Plane Development Kit (DPDK), and with pure FPGA logic implemented in gate ware. We characterize the three implementations in terms of throughput, latency, and power.
John W. Lockwood
FCCM1
2015 Comparison of Key/Value Store (KVS) in software and programmable hardware
John W. Lockwood
Hot Chips Symposium1
2014 Guest Editorial Deep Packet Inspection: Algorithms, Hardware, and Applications
abstract
The thirteen articles in this special section explore the technology of deep packet inspection (DPI). DPI examines the content in packet payloads to search for signatures of network applications, signs of malicious activities, and leaks of sensitive information, rather than just examine packet headers for information such as IP addresses and port numbers. The inspection provides network devices with rich information of application protocol messages in packet payloads, and enables them to make intelligent decisions in packet processing based on the information. The papers are organized into the following four sections: (1) Scalable Algorithms and Architectures for DPI, (2) Network Traffic Analysis with DPI, (3) Network Protocol Identification with DPI, and (4) Network Security Analysis with DPI.
Ying-Dar Lin, Po-Ching Lin, Viktor Prasanna 0001, H. Jonathan Chao, John W. Lockwood
IEEE J. Sel. Areas Commun.5
2009 A Packet Generator on the NetFPGA Platform
abstract
A packet generator and network traffic capture system has been implemented on the NetFPGA. The NetFPGA is an open networking platform accelerator that enables rapid development of hardware-accelerated packet processing applications. The packet generator application allows Internet packets to be transmitted at line rate on up to four gigabit Ethernet ports simultaneously. Data transmitted is specified in a standard PCAP file, transferred to local memory on the NetFPGA card, then sent on the gigabit links using a precise data rate, inter-packet delay, and number of iterations specified by the user. The hardware circuit also simultaneously operates as a packet capture system, allowing traffic to be captured from up to all four of the gigabit Ethernet ports. Timestamps are recorded and traffic can be transferred back to the host and stored using the same PCAP format. The project has been implemented as a fully open-source project and serves as an exemplar project on how to build and distribute NetFPGA applications. All of the code (Verilog hardware, system software, verification scripts, make files, and support tools) can be freely downloaded from the NetFPGA.org Website. Benchmarks comparing this hardware-accelerated application to the fastest available PC with a PCIe NIC shows that the FPGA-based hardware-accelerator far exceeds the performance possible using TCP-reply software.
G. Adam Covington, Glen Gibb, John W. Lockwood, Nick McKeown
FCCM3
2008 Reconfigurable content-based router using hardware-accelerated language parser
abstract
This article presents a dense logic design for matching multiple regular expressions with a field programmable gate array (FPGA) at 10+ Gbps. It leverages on the design techniques that enforce the shortest critical path on most FPGA architectures while optimizing the circuit size. The architecture is capable of supporting a maximum throughput of 12.90 Gbps on a Xilinx Virtex 4 LX200 and its performance is linearly scalable with size. Additionally, this article presents techniques for parsing data streams to provide semantic information for patterns found within a data stream. We illustrate how a content-based router can be implemented with our parsing techniques using an XML parser as an example. The content-based router presented was designed, implemented, and tested in a Xilinx Virtex XCV2000E FPGA on the FPX platform. It is capable of processing 32-bits of data per clock cycle and runs at 100 MHz. This allows the system to process and route XML messages at 3.2 Gbps.
James Moscola, John W. Lockwood, Young H. Cho
ACM Trans. Design Autom. Electr. Syst.2
2007 Changing Output Quality for Thermal Management
abstract
A growing number of embedded computing systems are used outside of environmentally controlled locations. In locations such as remote parts of deserts, deep ocean floors, and outer space, it is not only difficult to predict environmental effects on a system, they also allow very limited accessibility once a system is deployed. Therefore, it is often necessary for system parameters to be over provisioned to guarantee correct functionality under worst case environmental conditions. This often leads to an end system that is suboptimal for typical conditions.
Phillip H. Jones, James Moscola, Young H. Cho, John W. Lockwood
FCCM4
2007 Adaptive Thermoregulation for Applications on Reconfigurable Devices
abstract
A biological organism's ability to sense and adapt to its environment is essential to its survival. Likewise, environmentally aware computing systems avail themselves to a longer operational life and a wider range of applications than traditional systems. In this paper, we propose a novel circuit design methodology that allows parameterizable hardware to self-regulate its temperature. We apply this methodology to an image recognition system on an Xilinx Virtex 4 FX100 field programmable gate array (FPGA). The image recognition system sustains a safe operational temperature by automatically adjusting its frequency and output quality. The circuit sacrifices output performance and quality to lower its internal temperature as the ambient temperature increases, and can leverage cooler temperatures by increasing output performance and quality. Furthermore, the circuit will shutdown if the ambient temperature becomes too hot for the device to function properly. A performance evaluation of our adaptive circuit under various thermal conditions shows up to a 4× factor increase in performance and a 2× factor increase in quality over a system without dynamic thermal control.
Phillip H. Jones, James Moscola, Young H. Cho, John W. Lockwood
FPL4
2007 Supercharging planetlab: a high performance, multi-application, overlay network platform
abstract
In recent years, overlay networks have become an important vehicle for delivering Internet applications. Overlay network nodes are typically implemented using general purpose servers or clusters. We investigate the performance benefits of more integrated architectures, combining general-purpose servers with high performance Network Processor (NP) subsystems. We focus on PlanetLab as our experimental context and report on the design and evaluation of an experimental PlanetLab platform capable of much higher levels of performance than typical system configurations. To make it easier for users to port applications, the system supports a fast path/slow path application structure that facilitates the mapping of the most performance-critical parts of an application onto an NP subsystem, while allowing the more complex control and exception-handling to be implemented within the programmer-friendly environment provided by conventional servers. We report on implementations of two sample applications, an IPv4 router, and a forwarding application for the Internet Indirection Infrastructure. We demonstrate an 80x improvement in packet processing rates and comparable reductions in latency.
Jonathan S. Turner, Patrick Crowley, John D. DeHart, Amy Freestone, Brandon Heller, Fred Kuhns, Sailesh Kumar, John W. Lockwood, Michael Wilson 0001, Charlie Wiseman, David Zar
SIGCOMM8
2006 Fast packet classification using bloom filters
abstract
Ternary Content Addressable Memory (TCAM), although widely used for general packet classification, is an expensive and high power-consuming device. Algorithmic solutions which rely on commodity memory chips are relatively inexpensive and power-efficient but have not been able to match the generality and performance of TCAMs. Therefore, the development of fast and power-efficient algorithmic packet classification techniques continues to be a research subject.In this paper we propose a new approach to packet classification which combines architectural and algorithmic techniques. Our starting point is the well-known crossproduct algorithm which is fast but has significant memory overhead due to the extra rules needed to represent the crossproducts. We show how to modify the crossproduct method in a way that drastically reduces the memory requirement without compromising on performance. Unnecessary accesses to the off-chip memory are avoided by filtering them through on-chip Bloom filters. For packets that match p rules in a rule set, our algorithm requires just 4 + p + ε independent memory accesses to return all matching rules, where ε << 1 is a small constant that depends on the false positive rate of the Bloom filters. Using two commodity SRAM chips, a throughput of 38 Million packets per second can be achieved. For rule set sizes ranging from a few hundred to several thousand filters, the average rule set expansion factor attributable to the algorithm is just 1.2 to 1.4. The average memory consumption per rule is 32 to 45 bytes.
Sarang Dharmapurikar, Haoyu Song 0001, Jonathan S. Turner, John W. Lockwood
ANCS4
2006 A Scalable Hybrid Regular Expression Pattern Matcher
abstract
In this paper, the authors present a reconfigurable hardware architecture for searching for regular expression patterns in streaming data. This new architecture is created by combining two popular pattern matching techniques: a pipelined character grid architecture (Baker, 2004), and a regular expression NFA architecture (Cho, 2006). The resulting hybrid architecture can scale the number of input characters while still maintaining the ability to scan for regular expression patterns
James Moscola, Young H. Cho, John W. Lockwood
FCCM3
2006 Hierarchical Clustering using Reconfigurable Devices
abstract
Non-hierarchical k-means algorithms have been implemented in hardware, most frequently for image clustering. Here, we focus on hierarchical clustering of text documents based on document similarity. To our knowledge, this is the first work to present a hierarchical clustering algorithm designed for hardware implementation and ours is the first hardware-accelerated implementation
Shobana Padmanabhan, Moshe Looks, Dan Legorreta, Young H. Cho, John W. Lockwood
FCCM5
2006 Context-free-grammar based token tagger in reconfigurable devices
abstract
We present a high performance reconfigurable hardware architecture for detecting patterns as well as their contextual meaning. By analyzing for both semantic content and structure, the accuracy of content-level processing systems can be improved. Our system is built using semantics defined by context-free-grammar (CFG) to tag the streaming data. Unlike the traditional table look up and stack based engines used in CFG parsers, we explore a new method that maps the grammar structure on to the Field Programmable Gate Arrays (FPGA) hardware. The structure is a direct translation of the grammar which enables the meaning of the patterns to be determined based on the location of its detection. The parallel pattern detection engines are instantiated using FPGA resources. Our implementation scans for the regular expression patterns and determines their semantics as defined by a grammar. This highly parallel and fine grained pipelined engine with 8 bit input bus can operate in bandwidth above 2 Gbps. For a simple XML grammar example, our engine can detect and tag the patterns at 1.57 Gbps on Xilinx VirtexE FPGA and 4.26 Gbps on the new Virtex 4 devices.
Young H. Cho, James Moscola, John W. Lockwood
FPGA3
2006 High Speed Document Clustering in Reconfigurable Hardware
abstract
High-performance document clustering systems enable similar documents to be automatically organized into groups. In the past, the large amount of computational time needed to cluster documents prevented practical use of such systems with a large number of documents. A full hardware implementation of the K-means clustering algorithm has been designed and implemented in reconfigurable hardware that clusters 512k documents rapidly. This implementation, uses four parallel cosine distance metrics to cluster document vectors that each have 4000 dimensions. The synthesized hardware runs on the field programmable port extender (FPX) platform at a clock rate of 80 MHz. Although the clock rate on the Xilinx VirtexE 2000 is slower than a CPU, the implementation runs 26 times faster than an algorithmically equivalent software algorithm running on an Intel 3.60 GHz Xeon. The same architecture was used to synthesize a faster and larger design for the Xilinx Virtex4 LX200. This larger implementation can contain up to 25 parallel cosine distance metrics. The implementation synthesized with a clock rate of 250 MHz and outperforms the equivalent software by a factor of 328
G. Adam Covington, Charles L. G. Comstock, Andrew A. Levine, John W. Lockwood, Young H. Cho
FPL4
2006 A Thermal Management and Profiling Method for Reconfigurable Hardware Applications
abstract
Given large circuit sizes, high clock frequencies, and possibly extreme operating environments, Field Programmable Gate Arrays (FPGAs) are capable of heating beyond their designed thermal limits. As new circuits are developed for FPGAs and deployed remotely, engineers are challenged to determine in advance if the device will operate within recommended thermal ranges. The amount of power consumed by the circuit depends on how an algorithm is compiled into hardware, how the circuit is placed and routed, and the patterns of data that pass through the system. The amount of heat that can be dissipated depends on the thermal transfer characteristics of the package, the air flow that passes over the package, and the ambient temperature of the remote systems. Rather than designing a system to handle unreasonable worst-case situations, we have implemented a thermal management system that continuously monitors the temperature of the FPGA and reprograms the device if the temperate approaches the outer limits of safe operating conditions. Our system measures the junction temperature of a Xilinx Virtex FPGA using a built-in thermal diode. Using the temperature monitoring mechanism, we have studied the steady-state and transient conditions of multiple benchmark circuits implemented in an FPGA logic on the Field-programmable Port Extender (FPX) development platform. We observed properties of these benchmark circuits that enable us to predict power and thermal characteristics for real applications. We propose a Dynamic Thermal Management (DTM) strategy for FPGAs based on temperature feedback.
Phillip H. Jones, John W. Lockwood, Young H. Cho
FPL2
2006 Implementation of Network Application Layer Parser for Multiple TCP/IP Flows in Reconfigurable Devices
abstract
This paper presents an implementation of a high-performance network application layer parser in FPGAs. At the core of the architecture resides a pattern matcher and a parser. The pattern matcher scans for patterns in high-speed streaming TCP data streams. The parser core augments each pattern found with semantic information determined from the patterns location within the data stream. The packet payload parser can provide a higher level of understanding of a data stream for many network applications. Such applications include high performance XML parsers, content-based/aware routers, and others. Additionally, a TCP processor allows stateful packet payload parsing of up to 8 million simultaneous TCP flows. The payload parser has been implemented in a Xilinx Virtex E 2000 FPGA on the Field-Programmable Port Extender platform. The parsing module runs at 200 MHz and parse raw data at 6.4 Gbps. The payload parser, integrated with the TCP processor, runs at 100 MHz for a throughput of 3.2 Gbps
James Moscola, Young H. Cho, John W. Lockwood
FPL3
2006 An adaptive frequency control method using thermal feedback for reconfigurable hardware applications
abstract
Reconfigurable circuits running in field programmable gate arrays (FPGAs) can be dynamically optimized for power based on computational requirements and thermal conditions of the environment. In the past, FPGA circuits were typically small and operated at a low frequency. Few users were concerned about high-power consumption and the heat generated by FPGA devices. The current generation of FPGAs, however, use extensive pipelining techniques to achieve high data processing rates and dense layouts that can generate significant amounts of heat. FPGA circuits can be synthesized that can generate more heat than the package can dissipate. For FPGAs that operate in controlled environments, heatsinks and fans can be mounted to the device to extract heat from the device. When FPGA devices do not operate in a controlled environment, however, changes to ambient temperature due to factors such as the failure of a fan or a reconfiguration of bitfile running on the device can drastically change the operating conditions. A protection mechanism is needed to ensure the proper operation of the FPGA circuits when such a change occurs. To address these issues, we have devised a reconfigurable temperature monitoring system that gives feedback to the FPGA circuit using the measured junction temperature of the device. Using this feedback, we designed a novel dual frequency switching system that allows the FPGA circuits to maintain the highest level of performance for a given maximum junction temperature. Our working system has been implemented and deployed on the field programmable port extender (FPX) platform at Washington University in St. Louis. Our experimental results with a scalable image correlation circuit show up to a 2.4times factor increase in performance as compared to a system without thermal feedback. Our circuit ensures that the device performs the maximum required computation while always operating within a safe temperature range
Phillip H. Jones, Young H. Cho, John W. Lockwood
FPT3
2006 Vision for liquid architecture
abstract
In the liquid architecture project, we are exploring ways in which architectural flexibility can be exploited to improve the execution properties of individual applications. Here, we report on successes we have had to date in this area, and present our vision of where this research should proceed into the future.
Roger D. Chamberlain, Ron Cytron, Jason E. Fritts, John W. Lockwood
IPDPS4
2006 Reconfigurable context-free grammar based data processing hardware with error recovery
abstract
This paper presents an architecture for context-free grammar (CFG) based data processing hardware for re-configurable devices. Our system leverages on CFGs to tokenize and parse data streams into a sequence of words with corresponding semantics. Such a tokenizing and parsing engine is sufficient for processing grammatically correct input data. However, most pattern recognition applications must consider data sets that do not always conform to the predefined grammar. Therefore, we augment our system to detect and recover from grammatical errors while extracting useful information. Unlike the table look up method used in traditional CFG parsers, we map the structure of the grammar rules directly onto the field programmable gate array (FPGA). Since every part of the grammar is mapped onto independent logic, the resulting design is an efficient parallel data processing engine. To evaluate our design, we implement several XML parsers in an FPGA. Our XML parsers are able to process the full content of the packets up to 3.59 Gbps on Xilinx Virtex 4 devices
James Moscola, Young H. Cho, John W. Lockwood
IPDPS3
2006 Automatic application-specific microarchitecture reconfiguration
abstract
Applications for constrained embedded systems are subject to strict time constraints and restrictive resource utilization. With soft core processors, application developers can customize the processor for their application, constrained by resources but aimed at high application performance. With such freedom in the design space of the processor, however, comes complexity. We present here an automatic optimization technique that helps the developers with the processor microarchitecture customization. A naive approach exploring all possible configurations is exponential with the number of parameters and hence is clearly infeasible, even with only tens of reconfigurable parameters. Instead, our approach runs in time that is linear with the number of parameter values, based on an assumption of parameter independence. This makes the approach feasible and scalable. For the dimensions that we customize, namely application runtime and hardware resources, we formulate their costs as a constrained binary integer nonlinear optimization program. Though the results are not guaranteed to be optimal, we find they are near-optimal in practice. Our technique itself is general and can be applied to other design-space exploration problems
Shobana Padmanabhan, Ron Cytron, Roger D. Chamberlain, John W. Lockwood
IPDPS4
2006 Rethinking Hardware Support for Network Analysis and Intrusion Prevention
Vern Paxson, Krste Asanovic, Sarang Dharmapurikar, John W. Lockwood, Ruoming Pang, Robin Sommer, Nicholas Weaver
HotSec4
2006 Fast and Scalable Pattern Matching for Network Intrusion Detection Systems
abstract
High-speed packet content inspection and filtering devices rely on a fast multipattern matching algorithm which is used to detect predefined keywords or signatures in the packets. Multipattern matching is known to require intensive memory accesses and is often a performance bottleneck. Hence, specialized hardware-accelerated algorithms are required for line-speed packet processing. We present hardware-implementable pattern matching algorithm for content filtering applications, which is scalable in terms of speed, the number of patterns and the pattern length. Our algorithm is based on a memory efficient multihashing data structure called Bloom filter. We use embedded on-chip memory blocks in field programmable gate array/very large scale integration chips to construct Bloom filters which can suppress a large fraction of memory accesses and speed up string matching. Based on this concept, we first present a simple algorithm which can scan for several thousand short (up to 16 bytes) patterns at multigigabit per second speeds with a moderately small amount of embedded memory and a few mega bytes of external memory. Furthermore, we modify this algorithm to be able to handle arbitrarily large strings at the cost of a little more on-chip memory. We demonstrate the merit of our algorithm through theoretical analysis and simulations performed on Snort's string set
Sarang Dharmapurikar, John W. Lockwood
IEEE J. Sel. Areas Commun.2
2005 Fast and scalable pattern matching for content filtering
abstract
High-speed packet content inspection and filtering devices rely on a fast multi-pattern matching algorithm which is used to detect predefined keywords or signatures in the packets. Multi-pattern matching is known to require intensive memory accesses and is often a performance bottleneck. Hence specialized hardware-accelerated algorithms are being developed for line-speed packet processing. While several pattern matching algorithms have already been developed for such applications, we find that most of them suffer from scalability issues. To support a large number of patterns, the throughput is compromised or vice versa.We present a hardware-implementable pattern matching algorithm for content filtering applications, which is scalable in terms of speed, the number of patterns and the pattern length. We modify the classic Aho-Corasick algorithm to consider multiple characters at a time for higher throughput. Furthermore, we suppress a large fraction of memory accesses by using Bloom filters implemented with a small amount of on-chip memory. The resulting algorithm can support matching of several thousands of patterns at more than 10 Gbps with the help of a less than 50 KBytes of embedded memory and a few megabytes of external SRAM. We demonstrate the merit of our algorithm through theoretical analysis and simulations performed on Snort's string set.
Sarang Dharmapurikar, John W. Lockwood
ANCS2
2005 A Framework for Rule Processing in Reconfigurable Network Systems
abstract
High-performance rule processing systems are needed by network administrators in order to protect Internet systems from attack. Researchers have been working to implement components of intrusion detection systems (IDS), such as the highly popular Snort system, in reconfigurable hardware. While considerable progress has been made in the areas of string matching and header processing, complete systems have not yet been demonstrated that effectively combine all of the functionality necessary to perform rule processing for network systems. In this paper, a framework for implementing a rule processing system in reconfigurable hardware is presented. The framework integrates the functionality to scan dataflows for regular expressions, fixed strings, and header values. It also allows modules to be added to perform extended functionality to support all features found in Snort rules. Reconfigurability and flexibility are key components of the framework that enable it to adapt to protect Internet systems from threats including malicious worms, computer viruses, and network intruders. To prove the framework viable, a system has been built that scans all bytes of transmission control protocol/Internet protocol (TCP/IP) traffic entering and leaving a network's gateway at multi-gigabit rates. Using Xilinx FPGA hardware on the field programmable port extender (FPX) platform, the framework can process 32,768 complex rules at data rates of 2.5 Gbps. Systems to handle data at 10 Gbps rates can be built today using the same framework in the latest reconfigurable hardware devices such as the Virtex 4.
Michael Attig, John W. Lockwood
FCCM2
2005 A framework for rule processing in reconfigurable network systems (abstract only)
abstract
High-performance rule processing systems are needed by network administrators in order to protect Internet systems from attack. Researchers have been working to implement components of Intrusion Detection Systems, such as the highly popular Snort system, in reconfigurable hardware. While considerable progress has been made in the areas of string matching and header processing, complete systems have yet been demonstrated that effectively combine all of the functionality necessary to perform intrusion detection and prevention for real network systems.In this paper, we discuss a framework for implementing a rule processing system in reconfigurable hardware. The framework integrates the functionality to scan data flows for regular expressions, fixed strings, and header values. It also allows plug-in modules to be added to the system in order to perform the extended functionality to support the remaining set of Snort rules in modular components. Reconfigurability and flexibility are key components of the system that enable it adapt to protect Internet systems from threats including malicious worms, computer viruses, and network intruders.To prove the system viable, a system has been built that scans all bytes of Transmission Control protocol/Internet Protocol (TCP/IP) traffic entering and leaving a network's gateway at multi-gigabit rates and performs deep-packet inspection. Using Xilinx FPGA hardware on the Field programmable Port Extender (FPX), the platform can process 32,768 complex rules at data rates of 2.5 Gbps. Systems to handle data at 10 Gbps rates can be built today using the same framework in the most recent generation of reconfigurable hardware devices such as the Virtex4.
Michael Attig, John W. Lockwood
FPGA2
2005 Efficient packet classification for network intrusion detection using FPGA
abstract
Using FPGA technology for real-time network intrusion detection has gained many research efforts recently. In this paper, a novel packet classification architecture called BV-TCAM is presented, which is implemented for an FPGA-based Network Intrusion Detection System (NIDS). The classifier can report multiple matches at gigabit per second network link rates. The BV-TCAM architecture combines the Ternary Content Addressable Memory (TCAM) and the Bit Vector (BV) algorithm to effectively compress the data representations and boost throughput. A tree-bitmap implementation of the BV algorithm is used for source and destination port lookup while a TCAM performs the lookup of the other header fields, which can be represented as a prefix or exact value. The architecture eliminates the requirement for prefix expansion of port ranges. With the aid of a small embedded TCAM, packet classification can be implemented in a relatively small part of the available logic of an FPGA. The design is prototyped and evaluated in a Xilinx FPGA XCV2000E on the FPX platform. Even with the most difficult set of rules and packet inputs, the circuit is fast enough to sustain OC48 traffic throughput. Using larger and faster FPGAs, the system can work at speeds greater than OC192.
Haoyu Song 0001, John W. Lockwood
FPGA2
2005 HAIL: A Hardware-Accelerated Algorithm for Language Identification
abstract
A hardware-accelerated algorithm has been designed to automatically identify the primary languages used in documents transferred over the Internet. The algorithm has been implemented in hardware on the field programmable port extender (FPX) platform. This system, referred to as the hardware-accelerated identification of languages (HAIL) project, identifies the primary languages used in content transferred over transmission control protocol (TCP)/Internet protocol (IP) networks that operate at rates exceeding 2.4 Gigabits/second. We demonstrate that this hardware accelerated circuit, operating on a Xilinx XCV2000E-8 FPGA, far outperforms software algorithms running on modern personal computers while maintaining extremely high levels of accuracy.
Charles M. Kastner, G. Adam Covington, Andrew A. Levine, John W. Lockwood
FPL4
2005 Snort Offloader: A Reconfigurable Hardware NIDS Filter
abstract
Software-based network intrusion detection systems (NIDS) often fail to keep up with high-speed network links. In this paper an FPGA-based pre-filter is presented that reduces the amount of traffic sent to a software-based NIDS for inspection. Simulations using real network traces and the Snort rule set show that a pre-filter can reduce up to 90% of network traffic that would have otherwise been processed by Snort software. The projected performance enables a computer to perform real-time intrusion detection of malicious content passing over a 10 Gbps network using FPGA hardware that operates with 10 Gbps of throughput and software that needs only to operate with 1 Gbps of throughput.
Haoyu Song 0001, Todd Sproull, Michael Attig, John W. Lockwood
FPL4
2005 Optimizing memory bandwidth of a multi-channel packet buffer
abstract
Backbone routers typically require large buffers to hold packets during congestion. A thumb rule is to provide a buffer at every link, equal to the product of the round trip time and the link capacity. This translates into Gigabytes of buffers operating at line rate at every link. Such a size and rate necessitates the use of SDRAM with bandwidth of, for example, 80 Gbps for link speed of 40 Gbps. With speedup in the switch fabrics used in most routers, the bandwidth requirement of the buffer increases further. While multiple SDRAM devices can be used in parallel to achieve high bandwidth and storage capacity, a wide logical data bus composed of these devices results in suboptimal performance for arbitrarily sized packets. An alternative is to divide the wide logical data bus into multiple logical channels and store packets into them independently. However, in such an organization, the cumulative pin count grows due to additional address buses which might offset the performance gained. We find that due to several existing memory technologies and their characteristics and with Internet traffic composed of particular sized packets, a judiciously architected data channel can greatly enhance the performance per pin. In this paper, we derive an expression for the effective memory bandwidth of a parallel channel packet buffer and show how it can be optimized for a given number of I/O pins available for interfacing to memory. We believe that our model can greatly aid packet buffer designers to achieve the best performance.
Sarang Dharmapurikar, Sailesh Kumar, John W. Lockwood, Patrick Crowley
GLOBECOM3
2005 Advances for networks & internet
John W. Lockwood, George Kesidis
GLOBECOM1
2005 Multi-pattern signature matching for hardware network intrusion detection systems
abstract
Network intrusion detection system (NIDS) performs deep inspections on the packet payload to identify, deter and contain the malicious attacks over the Internet. It needs to perform exact matching on multi-pattern signatures in real time. In this paper we introduce an efficient data structure called extended Bloom filter (EBF) and the corresponding algorithm to perform the multi-pattern signature matching. We also present a technique to support long signature matching so that we need only to maintain a limited number of supported signature lengths for the EBFs. We show that at reasonable hardware cost we can achieve very fast and almost time-deterministic exact matching for thousands of signatures. The architecture takes the advantages of embedded multi-port memories in FPGAs and can be used to build a full-featured hardware-based NIDS.
Haoyu Song 0001, John W. Lockwood
GLOBECOM2
2005 Shape Shifting Tries for Faster IP Route Lookup
abstract
Some of the fastest practical algorithms for IP route lookup are based on space-efficient encodings of multi-bit tries (M. Degermark, et al., 1997, W. Eatherton, 1999). Unfortunately, the time required by these algorithms grows in proportion to the address length, making them less attractive for IPv6. This paper describes and evaluates a new data structure called a shape-shifting trie, in which the data structure nodes correspond to arbitrarily shaped subtrees of the underlying binary trie for a given set of address prefixes. The ability to adapt the node shape to the trie reduces the number of nodes that must be accessed to perform a lookup, especially for tries with large sparse regions. We give a fast algorithm for optimally dividing a trie into nodes so as to minimize the maximum lookup depth. We show that seven data structure accesses are sufficient for route tables with more than 150,000 IPv6 prefixes. This makes it possible to achieve wire-speed processing for OC192 link using a single QDRII SRAM chip.
Haoyu Song 0001, Jonathan S. Turner, John W. Lockwood
ICNP3
2005 Sensor fusion and correlation
abstract
No abstract available.
Todd Sproull, Richard Hough, John W. Lockwood, Christopher K. Zuver, Kent English, John Meier
SenSys3
2005 Fast hash table lookup using extended bloom filter: an aid to network processing
abstract
Hash tables are fundamental components of several network processing algorithms and applications, including route lookup, packet classification, per-flow state management and network monitoring. These applications, which typically occur in the data-path of high-speed routers, must process and forward packets with little or no buffer, making it important to maintain wire-speed throughout. A poorly designed hash table can critically affect the worst-case throughput of an application, since the number of memory accesses required for each lookup can vary. Hence, high throughput applications require hash tables with more predictable worst-case lookup performance. While published papers often assume that hash table lookups take constant time, there is significant variation in the number of items that must be accessed in a typical hash table search, leading to search times that vary by a factor of four or more.We present a novel hash table data structure and lookup algorithm which improves the performance over a naive hash table by reducing the number of memory accesses needed for the most time-consuming lookups. This allows designers to achieve higher lookup performance for a given memory bandwidth, without requiring large amounts of buffering in front of the lookup engine. Our algorithm extends the multiple-hashing Bloom Filter data structure to support exact matches and exploits recent advances in embedded memory technology. Through a combination of analysis and simulations we show that our algorithm is significantly faster than a naive hash table using the same amount of memory, hence it can support better throughput for router applications that use hash tables.
Haoyu Song 0001, Sarang Dharmapurikar, Jonathan S. Turner, John W. Lockwood
SIGCOMM4
2004 Implementation Results of Bloom Filters for String Matching
abstract
Network intrusion detection and prevention systems (IDPS) use string matching to scan Internet packets for malicious content. Bloom filters offer a mechanism to search for a large number of strings efficiently and concurrently when implemented with field programmable gate array (FPGA) technology. A string matching circuit has been implemented within the FPX platform using Bloom filters. Using 155 block RAMs on a single Xilinx VirtexE 2000 FPGA, the circuit scans for 35,475 unique signatures.
Michael Attig, Sarang Dharmapurikar, John W. Lockwood
FCCM3
2004 Secure Remote Control of Field-programmable Network Devices
abstract
A circuit and an associated lightweight protocol have been developed to secure communication between a control console and remote programmable network devices. The circuit provides encryption, data integrity checking and sequence number verification to ensure confidentiality, integrity and authentication of control messages sent over the public Internet. All of these functions are performed directly in FPGA hardware to provide high throughput and near-zero latency. The circuit has been used to control and configure remote firewalls and intrusion detection systems. The circuit could also be used to control and configure other distributed network applications.
Haoyu Song 0001, John W. Lockwood, James Moscola
FCCM3
2004 Automated Method to Generate Bitstream Intellectual Property Cores for Virtex FPGAs
Edson Lemos Horta, John W. Lockwood
FPL2
2004 A Modular System for FPGA-Based TCP Flow Processing in High-Speed Networks
David V. Schuehler, John W. Lockwood
FPL2
2003 Implementation of a Content-Scanning Module for an Internet Firewall
abstract
A module has been implemented in Field Programmable Gate Array (FPGA) hardware that scans the content of Internet packets at Gigabits/second rates. All of the packet processing operations are performed using reconfigurable hardware within a single Xilinx Virtex XCV2000E FPGA. A set of layered protocol wrappers is used to parse the headers and payloads of packets for Internet protocol data. A content matching server automatically generates the Finite State Machines (FSMs) to search for regular expressions. The complete system is operated on the Field-programmable Port Extender (FPX) platform.
James Moscola, John W. Lockwood, Ronald Prescott Loui, Michael Pachos
FCCM2
2003 An Extensible, System-On-Programmable-Chip, Content-Aware Internet Firewall
John W. Lockwood, Christopher E. Neely, Christopher K. Zuver, James Moscola, Sarang Dharmapurikar, David Lim
FPL1
2003 A TCP/IP Based Multi-device Programming Circuit
David V. Schuehler, Harvey Ku, John W. Lockwood
FPL3
2003 Beyond performance: secure and fair memory management for multiple systems on a chip
abstract
Developments in VLSI technologies create the possibility of hosting several independent (sub) systems in a single chip. There is a need to share a number of resources, especially off-chip resources, which creates new constraints in the design process. Although performance is still a key constraint, sharing implies that secure access to those resources and QoS guarantees are needed. In this paper, an architecture is presented that achieves the goals listed above. The Embedded Hardware Manager acts as a middleware between the applications and the resources, taking the role of resource manager and security agent. The results show that it can prevent resource misuse and undue information peeking or even altering while maintaining individual QoS guarantees. At the same time, high performance is still achieved.
Carlos Macián, Sarang Dharmapurikar, John W. Lockwood
FPT3
2003 System-on-chip packet processor for an experimental network services platform
abstract
As the focus of networking research shifts from raw performance to the delivery of advanced network services, there is a growing need for open-platform systems for extensible networking research. The Applied Research Laboratory at Washington University in Saint Louis has developed a flexible network services platform (NSP) to meet this need. The NSP provides an extensible platform for prototyping next-generation network services and applications. The paper describes the design of a system-on-chip packet processor for the NSP which performs all core packet processing functions, including segmentation and reassembly, packet classification, route lookup, and queue management. Targeted to a commercial configurable logic device, the system is designed to support gigabit links and switch fabrics with a 2:1 speed advantage. We provide resource consumption results for each component of the packet processor design.
David E. Taylor, Alex Chandra, Sarang Dharmapurikar, John W. Lockwood, Wenjing Tang, Jonathan S. Turner
GLOBECOM5
2003 Scalable IP lookup for Internet routers
abstract
Internet protocol (IP) address lookup is a central processing function of Internet routers. While a wide range of solutions to this problem have been devised, very few simultaneously achieve high lookup rates, good update performance, high memory efficiency, and low hardware cost. High performance solutions using content addressable memory devices are a popular but high-cost solution, particularly when applied to large databases. We present an efficient hardware implementation of a previously unpublished IP address lookup architecture, invented by Eatherton and Dittia (see M.S. thesis, Washington Univ., St. Louis, MO, 1998). Our experimental implementation uses a single commodity synchronous random access memory chip and less than 10% of the logic resources of a commercial configurable logic device, operating at 100 MHz. With these quite modest resources, it can perform over 9 million lookups/s, while simultaneously processing thousands of updates/s, on databases with over 100000 entries. The lookup structure requires 6.3 bytes per address prefix: less than half that required by other methods. The architecture allows performance to be scaled up by using parallel fast IP lookup (FIPL) engines, which interleave accesses to a common memory interface. This architecture allows performance to scale up directly with available memory bandwidth. We describe the tree bitmap algorithm, our implementation of it in a dynamically extensible gigabit router being developed at Washington University in Saint Louis, and the results of performance experiments designed to assess its performance under realistic operating conditions.
David E. Taylor, Jonathan S. Turner, John W. Lockwood, Todd Sproull, David B. Parlour
IEEE J. Sel. Areas Commun.3
2002 Dynamic hardware plugins in an FPGA with partial run-time reconfiguration
abstract
Tools and a design methodology have been developed to support partial run-time reconfiguration of FPGA logic on the Field Programmable Port Extender. High-speed Internet packet processing circuits on this platform are implemented as Dynamic Hardware Plugin (DHP) modules that fit within a specific region of an FPGA device. The PARBIT tool has been developed to transform and restructure bitfiles created by standard computer aided design tools into partial bitsteams that program DHPs. The methodology allows the platform to hot-swap application-specific DHP modules without disturbing the operation of the rest of the system.
Edson Lemos Horta, John W. Lockwood, David E. Taylor, David B. Parlour
DAC2
2002 Control and Configuration Software for a Reconfigurable Networking Hardware Platform
abstract
A suite of tools called NCHARGE (Networked Configurable Hardware Administrator for Reconfiguration and Governing via End-systems) has been developed to simplify the co-design of hardware and software components that process packets within a network of Field Programmable Gate Arrays (FPGAs). A key feature of NCHARGE is that it provides a high-performance packet interface to hardware and standard Application Programming Interface (API) between software and reprogrammable hardware modules. Using this API, multiple software processes can communicate to one or more hardware modules using standard TCP/IP sockets. NCHARGE also provides a Web-Based User Interface to simplify the configuration and control of an entire network switch that contains several software and hardware modules.
Todd Sproull, John W. Lockwood, David E. Taylor
FCCM2
2002 Using PARBIT to Implement Partial Run-Time Reconfigurable Systems
Edson Lemos Horta, John W. Lockwood, Sergio Takeo Kofuji
FPL2
2002 Scalable IP Lookup for Programmable Routers
abstract
Continuing growth in optical link speeds places increasing demands on the performance of Internet routers, while deployment of embedded and distributed network services imposes new demands for flexibility and programmability. IP address lookup has become a significant performance bottleneck for the highest performance routers. Amid the vast array of academic and commercial solutions to the problem, few achieve a favorable balance of performance, efficiency, and cost. New commercial products utilize content addressable memory (CAM) devices to achieve high lookup speeds at an exorbitantly high hardware cost with limited flexibility. In contrast, this paper describes an efficient, scalable lookup engine design, able to achieve high performance with the use of a small portion of a reconfigurable logic device and a commodity random access memory (RAM) device. The Fast Internet Protocol Lookup (FIPL) engine is an implementation of Eatherton and Dittia's previously unpublished Tree Bitmap algorithm (1998) targeted to an open-platform research router. FIPL can be scaled to achieve guaranteed worst-case performance of over 9 million lookups per second with a single SRAM operating at the fairly modest clock speed of 100 MHz. Experimental evaluation of FIPL throughput, latency, and update performance is provided using a sample routing table from Mae West.
David E. Taylor, John W. Lockwood, Todd Sproull, Jonathan S. Turner, David B. Parlour
INFOCOM2
2002 Dynamic hardware plugins: exploiting reconfigurable hardware for high-performance programmable routers
David E. Taylor, Jonathan S. Turner, John W. Lockwood, Edson Lemos Horta
Comput. Networks3
2001 Reprogrammable network packet processing on the field programmable port extender (FPX)
abstract
A prototype platform has been developed that allows processing of packets at the edge of a multi-gigabit-per-second network switch. This system, the Field Programmable Port Extender (FPX), enables packet processing functions to be implemented as modular components in reprogrammable hardware. All logic on the on the FPX is implemented in two Field Programmable Gate Arrays (FPGAs). Packet processing functions in the system are implemented as dynamically-loadable modules.
John W. Lockwood, Naji Naufel, Jonathan S. Turner, David E. Taylor
FPGA1
2001 Reconfigurable Router Modules Using Network Protocol Wrappers
Florian Braun, John W. Lockwood, Marcel Waldvogel
FPL2
2000 Field programmable port extender (FPX) for distributed routing and queuing
abstract
Field Programmable Gate Arrays (FPGAs) are being used to provide fast Internet Protocol (IP) packet routing and advanced queuing in a highly scalable network switch. A new module, called the Field-programmable Port Extender (FPX), is being built to augment the Washington University Gigabit Switch (WUGS) with reprogrammable logic.
John W. Lockwood, Jonathan S. Turner, David E. Taylor
FPGA1
2000 Compensation modeling for QoS support on a wireless network
abstract
This paper presents the design of a new error compensation model for providing QoS support on a wireless network. The compensation model uses different compensation strategies for the different priority classes in conjunction with a multiclass priority fair queuing (MPFQ) algorithm. Criteria for fair error compensation and a classification of error compensation models are established in order to show the properties of different models, and to select the best error compensation for each priority class. Simulation results show that the new MPFQ compensation model meets the long-term fairness guarantees and provides an improved flow separation.
Stefan Bucheli, Jay R. Moorman, John W. Lockwood
GLOBECOM3
1999 Implementation of campus-wide wireless network services using ATM, virtual LANs, and wireless basestations
abstract
A campus-wide wireless LAN has been implemented that combines the scalability of virtual LANs and the flexibility of switched Ethernet with the mobility of wireless LAN access. A suite of general purpose software tools has been developed and deployed to monitor the status of the network, measure the bandwidth utilization, control access to the network, and authenticate users.
John W. Lockwood
WCNC1
1997 A High-Performance OC-12/OC-48 Queue Design Prototype for Input-buffered ATM Switches
abstract
This paper presents the design and prototype of an intelligent, 3-dimensional-queue (3DQ) for high-performance, scalable, input-buffered ATM switches. The 3DQ uses pointers and linked lists to organize ATM cells into multiple virtual queues according to priority, destination, and virtual connection. It enforces per-virtual connection quality-of-service (QoS) and eliminates head-of-line (HOL) blocking. Using field-programmable-gate-array (FPGA) devices, our prototype hardware can process ATM cells at 622 Mb/s (OC-12). Using more aggressive technology (multi-chip-module (MCM) and fast GaAs logic), the same 3DQ can process cells at 2.5 Gb/s (OC-48). Using the 3DQ and matrix-unit-cell-scheduler (MUCS) as essential components, an input-buffered ATM switch system has been designed, which can achieve near-100% link bandwidth utilization.
John W. Lockwood, J. D. Will
INFOCOM2