EDBT 2026 Demo / reviewers in the wild / expert
John W. Lockwood
dblp:31/1714
· DBLP profile ↗
54ranked-venue papers
7as first author
0since 2021 · last 2015
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 35 · 5 first-authorComputer networks · 18 · 2 first-authorSecurity and privacy · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer networks
9 papers |
Routing and switching · 55% Internet architecture and protocols · 19% Software-defined and programmable networks · 15% | |
| Computer architecture, parallel and distributed computing, and storage systems
11 papers |
Reconfigurable computing and FPGAs · 74% Processor architecture and microarchitecture · 12% GPUs and heterogeneous computing · 9% | |
| Network and information security
3 papers |
Network security · 100% |
Topics — the 27 heaviest of 32, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Routing and switching
IP lookup |
0.1 | 3 | 2005 | Shape Shifting Tries for Faster IP Route Lookup · ICNP 2005 Scalable IP lookup for Internet routers · IEEE J. Sel. Areas Commun. 2003 Scalable IP Lookup for Programmable Routers · INFOCOM 2002 |
Reconfigurable computing and FPGAs › FPGA-based network processing
FPGA-based pattern matching |
0.1 | 2 | 2006 | Fast and Scalable Pattern Matching for Network Intrusion Detection Systems · IEEE J. Sel. Areas Commun. 2006 Context-free-grammar based token tagger in reconfigurable devices · FPGA 2006 |
Network security › intrusion detection and prevention
intrusion detection |
0.1 | 2 | 2006 | Fast and Scalable Pattern Matching for Network Intrusion Detection Systems · IEEE J. Sel. Areas Commun. 2006 Efficient packet classification for network intrusion detection using FPGA · FPGA 2005 |
Reconfigurable computing and FPGAs
FPGA-based network processing |
0.1 | 2 | 2005 | Efficient packet classification for network intrusion detection using FPGA · FPGA 2005 A framework for rule processing in reconfigurable network systems (abstract only) · FPGA 2005 |
Internet architecture and protocols
overlay networks |
0.1 | 1 | 2007 | Supercharging planetlab: a high performance, multi-application, overlay network platform · SIGCOMM 2007 |
Processor architecture and microarchitecture › special-purpose processor
network processor |
0.1 | 1 | 2007 | Supercharging planetlab: a high performance, multi-application, overlay network platform · SIGCOMM 2007 |
Network security › intrusion detection and prevention › intrusion detection › pattern matching
multi-pattern matching |
0.1 | 1 | 2006 | Fast and Scalable Pattern Matching for Network Intrusion Detection Systems · IEEE J. Sel. Areas Commun. 2006 |
Reconfigurable computing and FPGAs
reconfigurable computing |
0.1 | 1 | 2006 | Fast and Scalable Pattern Matching for Network Intrusion Detection Systems · IEEE J. Sel. Areas Commun. 2006 |
Software-defined and programmable networks › programmable data plane
FPGA-based packet processing |
0.1 | 2 | 2001 | Reprogrammable network packet processing on the field programmable port extender (FPX) · FPGA 2001 Field programmable port extender (FPX) for distributed routing and queuing · FPGA 2000 |
Software-defined and programmable networks
programmable data plane |
0.1 | 2 | 2001 | Reprogrammable network packet processing on the field programmable port extender (FPX) · FPGA 2001 Field programmable port extender (FPX) for distributed routing and queuing · FPGA 2000 |
Routing and switching › IP lookup
hash-based lookup |
0.1 | 1 | 2005 | Fast hash table lookup using extended bloom filter: an aid to network processing · SIGCOMM 2005 |
Routing and switching › IP lookup
IPv6 lookup |
0.1 | 1 | 2005 | Shape Shifting Tries for Faster IP Route Lookup · ICNP 2005 |
Internet architecture and protocols › packet processing
packet classification |
0.1 | 1 | 2005 | Fast hash table lookup using extended bloom filter: an aid to network processing · SIGCOMM 2005 |
Routing and switching › data plane
router data plane |
0.1 | 1 | 2005 | Fast hash table lookup using extended bloom filter: an aid to network processing · SIGCOMM 2005 |
Internet of things and sensor networks › wireless sensor network
sensor fusion |
0.1 | 1 | 2005 | Sensor fusion and correlation · SenSys 2005 |
Network security
intrusion detection and prevention |
0.1 | 1 | 2005 | A framework for rule processing in reconfigurable network systems (abstract only) · FPGA 2005 |
GPUs and heterogeneous computing › packet processing
packet classification |
0.1 | 1 | 2005 | Efficient packet classification for network intrusion detection using FPGA · FPGA 2005 |
Routing and switching › packet switch
router |
0.0 | 1 | 2003 | Scalable IP lookup for Internet routers · IEEE J. Sel. Areas Commun. 2003 |
Routing and switching
routing |
0.0 | 1 | 2000 | Field programmable port extender (FPX) for distributed routing and queuing · FPGA 2000 |
Automata and formal languages › formal grammars
context-free grammar |
0.0 | 1 | 2006 | Context-free-grammar based token tagger in reconfigurable devices · FPGA 2006 |
Routing and switching
input-queued switch |
0.0 | 1 | 1997 | A High-Performance OC-12/OC-48 Queue Design Prototype for Input-buffered ATM Switches · INFOCOM 1997 |
Internet architecture and protocols
quality of service |
0.0 | 1 | 1997 | A High-Performance OC-12/OC-48 Queue Design Prototype for Input-buffered ATM Switches · INFOCOM 1997 |
Routing and switching
switching |
0.0 | 1 | 1997 | A High-Performance OC-12/OC-48 Queue Design Prototype for Input-buffered ATM Switches · INFOCOM 1997 |
Network measurement and analytics
flow state management |
0.0 | 1 | 2005 | Fast hash table lookup using extended bloom filter: an aid to network processing · SIGCOMM 2005 |
Routing and switching › data plane › router data plane
high-speed packet processing |
0.0 | 1 | 2005 | Shape Shifting Tries for Faster IP Route Lookup · ICNP 2005 |
Hardware accelerators and domain-specific architectures › network function acceleration
TCAM-based classification |
0.0 | 1 | 2005 | Efficient packet classification for network intrusion detection using FPGA · FPGA 2005 |
Reconfigurable computing and FPGAs
FPGA prototyping |
0.0 | 1 | 1997 | A High-Performance OC-12/OC-48 Queue Design Prototype for Input-buffered ATM Switches · INFOCOM 1997 |
Methods — techniques the papers use, named apart from their topics
on-chip embedded memory · 0.1multihashing · 0.1bloom filter · 0.1tree-bitmap · 0.1string matching · 0.1regular expression matching · 0.1bit vector algorithm · 0.1TCAM · 0.1trie partitioning · 0.1multiple hashing · 0.1extended bloom filter · 0.1dynamic programming · 0.1tree bitmap algorithm · 0.0partial bitstream generation · 0.0bitfile transformation · 0.0SRAM · 0.0CAM · 0.0dynamic module loading · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | Scalable Key/Value Search in DatacentersabstractKey/Value Store (KVS) is a fundamental service used widely in modern data centers to associate keys with data values. KVS systems, such as Redis, Memcached, and Dynamo DB have traditionally been implemented with software and run on clusters of microprocessor-based servers. In this work an alternate approach is taken that performs KVS with gate ware in Field Programmable Gate Array (FPGA) logic. We leverage an efficient, open-standard, binary message format to transfer keys and values over Ethernet. Results of three different implementations of this KVS were compared -- software running on a Linux server with network data sent over UDP/IP sockets, kernel bypass using Intel's Data Plane Development Kit (DPDK), and with pure FPGA logic implemented in gate ware. We characterize the three implementations in terms of throughput, latency, and power. John W. Lockwood |
FCCM | 1 |
| 2015 | Comparison of Key/Value Store (KVS) in software and programmable hardware
John W. Lockwood |
Hot Chips Symposium | 1 |
| 2014 | Guest Editorial Deep Packet Inspection: Algorithms, Hardware, and ApplicationsabstractThe thirteen articles in this special section explore the technology of deep packet inspection (DPI). DPI examines the content in packet payloads to search for signatures of network applications, signs of malicious activities, and leaks of sensitive information, rather than just examine packet headers for information such as IP addresses and port numbers. The inspection provides network devices with rich information of application protocol messages in packet payloads, and enables them to make intelligent decisions in packet processing based on the information. The papers are organized into the following four sections: (1) Scalable Algorithms and Architectures for DPI, (2) Network Traffic Analysis with DPI, (3) Network Protocol Identification with DPI, and (4) Network Security Analysis with DPI. Ying-Dar Lin, Po-Ching Lin, Viktor Prasanna 0001, H. Jonathan Chao, John W. Lockwood |
IEEE J. Sel. Areas Commun. | 5 |
| 2009 | A Packet Generator on the NetFPGA PlatformabstractA packet generator and network traffic capture system has been implemented on the NetFPGA. The NetFPGA is an open networking platform accelerator that enables rapid development of hardware-accelerated packet processing applications. The packet generator application allows Internet packets to be transmitted at line rate on up to four gigabit Ethernet ports simultaneously. Data transmitted is specified in a standard PCAP file, transferred to local memory on the NetFPGA card, then sent on the gigabit links using a precise data rate, inter-packet delay, and number of iterations specified by the user. The hardware circuit also simultaneously operates as a packet capture system, allowing traffic to be captured from up to all four of the gigabit Ethernet ports. Timestamps are recorded and traffic can be transferred back to the host and stored using the same PCAP format. The project has been implemented as a fully open-source project and serves as an exemplar project on how to build and distribute NetFPGA applications. All of the code (Verilog hardware, system software, verification scripts, make files, and support tools) can be freely downloaded from the NetFPGA.org Website. Benchmarks comparing this hardware-accelerated application to the fastest available PC with a PCIe NIC shows that the FPGA-based hardware-accelerator far exceeds the performance possible using TCP-reply software. G. Adam Covington, Glen Gibb, John W. Lockwood, Nick McKeown |
FCCM | 3 |
| 2008 | Reconfigurable content-based router using hardware-accelerated language parserabstractThis article presents a dense logic design for matching multiple regular expressions with a field programmable gate array (FPGA) at 10+ Gbps. It leverages on the design techniques that enforce the shortest critical path on most FPGA architectures while optimizing the circuit size. The architecture is capable of supporting a maximum throughput of 12.90 Gbps on a Xilinx Virtex 4 LX200 and its performance is linearly scalable with size. Additionally, this article presents techniques for parsing data streams to provide semantic information for patterns found within a data stream. We illustrate how a content-based router can be implemented with our parsing techniques using an XML parser as an example. The content-based router presented was designed, implemented, and tested in a Xilinx Virtex XCV2000E FPGA on the FPX platform. It is capable of processing 32-bits of data per clock cycle and runs at 100 MHz. This allows the system to process and route XML messages at 3.2 Gbps. James Moscola, John W. Lockwood, Young H. Cho |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2007 | Changing Output Quality for Thermal ManagementabstractA growing number of embedded computing systems are used outside of environmentally controlled locations. In locations such as remote parts of deserts, deep ocean floors, and outer space, it is not only difficult to predict environmental effects on a system, they also allow very limited accessibility once a system is deployed. Therefore, it is often necessary for system parameters to be over provisioned to guarantee correct functionality under worst case environmental conditions. This often leads to an end system that is suboptimal for typical conditions. Phillip H. Jones, James Moscola, Young H. Cho, John W. Lockwood |
FCCM | 4 |
| 2007 | Adaptive Thermoregulation for Applications on Reconfigurable DevicesabstractA biological organism's ability to sense and adapt to its environment is essential to its survival. Likewise, environmentally aware computing systems avail themselves to a longer operational life and a wider range of applications than traditional systems. In this paper, we propose a novel circuit design methodology that allows parameterizable hardware to self-regulate its temperature. We apply this methodology to an image recognition system on an Xilinx Virtex 4 FX100 field programmable gate array (FPGA). The image recognition system sustains a safe operational temperature by automatically adjusting its frequency and output quality. The circuit sacrifices output performance and quality to lower its internal temperature as the ambient temperature increases, and can leverage cooler temperatures by increasing output performance and quality. Furthermore, the circuit will shutdown if the ambient temperature becomes too hot for the device to function properly. A performance evaluation of our adaptive circuit under various thermal conditions shows up to a 4× factor increase in performance and a 2× factor increase in quality over a system without dynamic thermal control. Phillip H. Jones, James Moscola, Young H. Cho, John W. Lockwood |
FPL | 4 |
| 2007 | Supercharging planetlab: a high performance, multi-application, overlay network platformabstractIn recent years, overlay networks have become an important vehicle for delivering Internet applications. Overlay network nodes are typically implemented using general purpose servers or clusters. We investigate the performance benefits of more integrated architectures, combining general-purpose servers with high performance Network Processor (NP) subsystems. We focus on PlanetLab as our experimental context and report on the design and evaluation of an experimental PlanetLab platform capable of much higher levels of performance than typical system configurations. To make it easier for users to port applications, the system supports a fast path/slow path application structure that facilitates the mapping of the most performance-critical parts of an application onto an NP subsystem, while allowing the more complex control and exception-handling to be implemented within the programmer-friendly environment provided by conventional servers. We report on implementations of two sample applications, an IPv4 router, and a forwarding application for the Internet Indirection Infrastructure. We demonstrate an 80x improvement in packet processing rates and comparable reductions in latency. Jonathan S. Turner, Patrick Crowley, John D. DeHart, Amy Freestone, Brandon Heller, Fred Kuhns, Sailesh Kumar, John W. Lockwood, Michael Wilson 0001, Charlie Wiseman, David Zar |
SIGCOMM | 8 |
| 2006 | Fast packet classification using bloom filtersabstractTernary Content Addressable Memory (TCAM), although widely used for general packet classification, is an expensive and high power-consuming device. Algorithmic solutions which rely on commodity memory chips are relatively inexpensive and power-efficient but have not been able to match the generality and performance of TCAMs. Therefore, the development of fast and power-efficient algorithmic packet classification techniques continues to be a research subject.In this paper we propose a new approach to packet classification which combines architectural and algorithmic techniques. Our starting point is the well-known crossproduct algorithm which is fast but has significant memory overhead due to the extra rules needed to represent the crossproducts. We show how to modify the crossproduct method in a way that drastically reduces the memory requirement without compromising on performance. Unnecessary accesses to the off-chip memory are avoided by filtering them through on-chip Bloom filters. For packets that match p rules in a rule set, our algorithm requires just 4 + p + ε independent memory accesses to return all matching rules, where ε << 1 is a small constant that depends on the false positive rate of the Bloom filters. Using two commodity SRAM chips, a throughput of 38 Million packets per second can be achieved. For rule set sizes ranging from a few hundred to several thousand filters, the average rule set expansion factor attributable to the algorithm is just 1.2 to 1.4. The average memory consumption per rule is 32 to 45 bytes. Sarang Dharmapurikar, Haoyu Song 0001, Jonathan S. Turner, John W. Lockwood |
ANCS | 4 |
| 2006 | A Scalable Hybrid Regular Expression Pattern MatcherabstractIn this paper, the authors present a reconfigurable hardware architecture for searching for regular expression patterns in streaming data. This new architecture is created by combining two popular pattern matching techniques: a pipelined character grid architecture (Baker, 2004), and a regular expression NFA architecture (Cho, 2006). The resulting hybrid architecture can scale the number of input characters while still maintaining the ability to scan for regular expression patterns James Moscola, Young H. Cho, John W. Lockwood |
FCCM | 3 |
| 2006 | Hierarchical Clustering using Reconfigurable DevicesabstractNon-hierarchical k-means algorithms have been implemented in hardware, most frequently for image clustering. Here, we focus on hierarchical clustering of text documents based on document similarity. To our knowledge, this is the first work to present a hierarchical clustering algorithm designed for hardware implementation and ours is the first hardware-accelerated implementation Shobana Padmanabhan, Moshe Looks, Dan Legorreta, Young H. Cho, John W. Lockwood |
FCCM | 5 |
| 2006 | Context-free-grammar based token tagger in reconfigurable devicesabstractWe present a high performance reconfigurable hardware architecture for detecting patterns as well as their contextual meaning. By analyzing for both semantic content and structure, the accuracy of content-level processing systems can be improved. Our system is built using semantics defined by context-free-grammar (CFG) to tag the streaming data. Unlike the traditional table look up and stack based engines used in CFG parsers, we explore a new method that maps the grammar structure on to the Field Programmable Gate Arrays (FPGA) hardware. The structure is a direct translation of the grammar which enables the meaning of the patterns to be determined based on the location of its detection. The parallel pattern detection engines are instantiated using FPGA resources. Our implementation scans for the regular expression patterns and determines their semantics as defined by a grammar. This highly parallel and fine grained pipelined engine with 8 bit input bus can operate in bandwidth above 2 Gbps. For a simple XML grammar example, our engine can detect and tag the patterns at 1.57 Gbps on Xilinx VirtexE FPGA and 4.26 Gbps on the new Virtex 4 devices. Young H. Cho, James Moscola, John W. Lockwood |
FPGA | 3 |
| 2006 | High Speed Document Clustering in Reconfigurable HardwareabstractHigh-performance document clustering systems enable similar documents to be automatically organized into groups. In the past, the large amount of computational time needed to cluster documents prevented practical use of such systems with a large number of documents. A full hardware implementation of the K-means clustering algorithm has been designed and implemented in reconfigurable hardware that clusters 512k documents rapidly. This implementation, uses four parallel cosine distance metrics to cluster document vectors that each have 4000 dimensions. The synthesized hardware runs on the field programmable port extender (FPX) platform at a clock rate of 80 MHz. Although the clock rate on the Xilinx VirtexE 2000 is slower than a CPU, the implementation runs 26 times faster than an algorithmically equivalent software algorithm running on an Intel 3.60 GHz Xeon. The same architecture was used to synthesize a faster and larger design for the Xilinx Virtex4 LX200. This larger implementation can contain up to 25 parallel cosine distance metrics. The implementation synthesized with a clock rate of 250 MHz and outperforms the equivalent software by a factor of 328 G. Adam Covington, Charles L. G. Comstock, Andrew A. Levine, John W. Lockwood, Young H. Cho |
FPL | 4 |
| 2006 | A Thermal Management and Profiling Method for Reconfigurable Hardware ApplicationsabstractGiven large circuit sizes, high clock frequencies, and possibly extreme operating environments, Field Programmable Gate Arrays (FPGAs) are capable of heating beyond their designed thermal limits. As new circuits are developed for FPGAs and deployed remotely, engineers are challenged to determine in advance if the device will operate within recommended thermal ranges. The amount of power consumed by the circuit depends on how an algorithm is compiled into hardware, how the circuit is placed and routed, and the patterns of data that pass through the system. The amount of heat that can be dissipated depends on the thermal transfer characteristics of the package, the air flow that passes over the package, and the ambient temperature of the remote systems. Rather than designing a system to handle unreasonable worst-case situations, we have implemented a thermal management system that continuously monitors the temperature of the FPGA and reprograms the device if the temperate approaches the outer limits of safe operating conditions. Our system measures the junction temperature of a Xilinx Virtex FPGA using a built-in thermal diode. Using the temperature monitoring mechanism, we have studied the steady-state and transient conditions of multiple benchmark circuits implemented in an FPGA logic on the Field-programmable Port Extender (FPX) development platform. We observed properties of these benchmark circuits that enable us to predict power and thermal characteristics for real applications. We propose a Dynamic Thermal Management (DTM) strategy for FPGAs based on temperature feedback. Phillip H. Jones, John W. Lockwood, Young H. Cho |
FPL | 2 |
| 2006 | Implementation of Network Application Layer Parser for Multiple TCP/IP Flows in Reconfigurable DevicesabstractThis paper presents an implementation of a high-performance network application layer parser in FPGAs. At the core of the architecture resides a pattern matcher and a parser. The pattern matcher scans for patterns in high-speed streaming TCP data streams. The parser core augments each pattern found with semantic information determined from the patterns location within the data stream. The packet payload parser can provide a higher level of understanding of a data stream for many network applications. Such applications include high performance XML parsers, content-based/aware routers, and others. Additionally, a TCP processor allows stateful packet payload parsing of up to 8 million simultaneous TCP flows. The payload parser has been implemented in a Xilinx Virtex E 2000 FPGA on the Field-Programmable Port Extender platform. The parsing module runs at 200 MHz and parse raw data at 6.4 Gbps. The payload parser, integrated with the TCP processor, runs at 100 MHz for a throughput of 3.2 Gbps James Moscola, Young H. Cho, John W. Lockwood |
FPL | 3 |
| 2006 | An adaptive frequency control method using thermal feedback for reconfigurable hardware applicationsabstractReconfigurable circuits running in field programmable gate arrays (FPGAs) can be dynamically optimized for power based on computational requirements and thermal conditions of the environment. In the past, FPGA circuits were typically small and operated at a low frequency. Few users were concerned about high-power consumption and the heat generated by FPGA devices. The current generation of FPGAs, however, use extensive pipelining techniques to achieve high data processing rates and dense layouts that can generate significant amounts of heat. FPGA circuits can be synthesized that can generate more heat than the package can dissipate. For FPGAs that operate in controlled environments, heatsinks and fans can be mounted to the device to extract heat from the device. When FPGA devices do not operate in a controlled environment, however, changes to ambient temperature due to factors such as the failure of a fan or a reconfiguration of bitfile running on the device can drastically change the operating conditions. A protection mechanism is needed to ensure the proper operation of the FPGA circuits when such a change occurs. To address these issues, we have devised a reconfigurable temperature monitoring system that gives feedback to the FPGA circuit using the measured junction temperature of the device. Using this feedback, we designed a novel dual frequency switching system that allows the FPGA circuits to maintain the highest level of performance for a given maximum junction temperature. Our working system has been implemented and deployed on the field programmable port extender (FPX) platform at Washington University in St. Louis. Our experimental results with a scalable image correlation circuit show up to a 2.4times factor increase in performance as compared to a system without thermal feedback. Our circuit ensures that the device performs the maximum required computation while always operating within a safe temperature range Phillip H. Jones, Young H. Cho, John W. Lockwood |
FPT | 3 |
| 2006 | Vision for liquid architectureabstractIn the liquid architecture project, we are exploring ways in which architectural flexibility can be exploited to improve the execution properties of individual applications. Here, we report on successes we have had to date in this area, and present our vision of where this research should proceed into the future. Roger D. Chamberlain, Ron Cytron, Jason E. Fritts, John W. Lockwood |
IPDPS | 4 |
| 2006 | Reconfigurable context-free grammar based data processing hardware with error recoveryabstractThis paper presents an architecture for context-free grammar (CFG) based data processing hardware for re-configurable devices. Our system leverages on CFGs to tokenize and parse data streams into a sequence of words with corresponding semantics. Such a tokenizing and parsing engine is sufficient for processing grammatically correct input data. However, most pattern recognition applications must consider data sets that do not always conform to the predefined grammar. Therefore, we augment our system to detect and recover from grammatical errors while extracting useful information. Unlike the table look up method used in traditional CFG parsers, we map the structure of the grammar rules directly onto the field programmable gate array (FPGA). Since every part of the grammar is mapped onto independent logic, the resulting design is an efficient parallel data processing engine. To evaluate our design, we implement several XML parsers in an FPGA. Our XML parsers are able to process the full content of the packets up to 3.59 Gbps on Xilinx Virtex 4 devices James Moscola, Young H. Cho, John W. Lockwood |
IPDPS | 3 |
| 2006 | Automatic application-specific microarchitecture reconfigurationabstractApplications for constrained embedded systems are subject to strict time constraints and restrictive resource utilization. With soft core processors, application developers can customize the processor for their application, constrained by resources but aimed at high application performance. With such freedom in the design space of the processor, however, comes complexity. We present here an automatic optimization technique that helps the developers with the processor microarchitecture customization. A naive approach exploring all possible configurations is exponential with the number of parameters and hence is clearly infeasible, even with only tens of reconfigurable parameters. Instead, our approach runs in time that is linear with the number of parameter values, based on an assumption of parameter independence. This makes the approach feasible and scalable. For the dimensions that we customize, namely application runtime and hardware resources, we formulate their costs as a constrained binary integer nonlinear optimization program. Though the results are not guaranteed to be optimal, we find they are near-optimal in practice. Our technique itself is general and can be applied to other design-space exploration problems Shobana Padmanabhan, Ron Cytron, Roger D. Chamberlain, John W. Lockwood |
IPDPS | 4 |
| 2006 | Rethinking Hardware Support for Network Analysis and Intrusion Prevention
Vern Paxson, Krste Asanovic, Sarang Dharmapurikar, John W. Lockwood, Ruoming Pang, Robin Sommer, Nicholas Weaver |
HotSec | 4 |
| 2006 | Fast and Scalable Pattern Matching for Network Intrusion Detection SystemsabstractHigh-speed packet content inspection and filtering devices rely on a fast multipattern matching algorithm which is used to detect predefined keywords or signatures in the packets. Multipattern matching is known to require intensive memory accesses and is often a performance bottleneck. Hence, specialized hardware-accelerated algorithms are required for line-speed packet processing. We present hardware-implementable pattern matching algorithm for content filtering applications, which is scalable in terms of speed, the number of patterns and the pattern length. Our algorithm is based on a memory efficient multihashing data structure called Bloom filter. We use embedded on-chip memory blocks in field programmable gate array/very large scale integration chips to construct Bloom filters which can suppress a large fraction of memory accesses and speed up string matching. Based on this concept, we first present a simple algorithm which can scan for several thousand short (up to 16 bytes) patterns at multigigabit per second speeds with a moderately small amount of embedded memory and a few mega bytes of external memory. Furthermore, we modify this algorithm to be able to handle arbitrarily large strings at the cost of a little more on-chip memory. We demonstrate the merit of our algorithm through theoretical analysis and simulations performed on Snort's string set Sarang Dharmapurikar, John W. Lockwood |
IEEE J. Sel. Areas Commun. | 2 |
| 2005 | Fast and scalable pattern matching for content filteringabstractHigh-speed packet content inspection and filtering devices rely on a fast multi-pattern matching algorithm which is used to detect predefined keywords or signatures in the packets. Multi-pattern matching is known to require intensive memory accesses and is often a performance bottleneck. Hence specialized hardware-accelerated algorithms are being developed for line-speed packet processing. While several pattern matching algorithms have already been developed for such applications, we find that most of them suffer from scalability issues. To support a large number of patterns, the throughput is compromised or vice versa.We present a hardware-implementable pattern matching algorithm for content filtering applications, which is scalable in terms of speed, the number of patterns and the pattern length. We modify the classic Aho-Corasick algorithm to consider multiple characters at a time for higher throughput. Furthermore, we suppress a large fraction of memory accesses by using Bloom filters implemented with a small amount of on-chip memory. The resulting algorithm can support matching of several thousands of patterns at more than 10 Gbps with the help of a less than 50 KBytes of embedded memory and a few megabytes of external SRAM. We demonstrate the merit of our algorithm through theoretical analysis and simulations performed on Snort's string set. Sarang Dharmapurikar, John W. Lockwood |
ANCS | 2 |
| 2005 | A Framework for Rule Processing in Reconfigurable Network SystemsabstractHigh-performance rule processing systems are needed by network administrators in order to protect Internet systems from attack. Researchers have been working to implement components of intrusion detection systems (IDS), such as the highly popular Snort system, in reconfigurable hardware. While considerable progress has been made in the areas of string matching and header processing, complete systems have not yet been demonstrated that effectively combine all of the functionality necessary to perform rule processing for network systems. In this paper, a framework for implementing a rule processing system in reconfigurable hardware is presented. The framework integrates the functionality to scan dataflows for regular expressions, fixed strings, and header values. It also allows modules to be added to perform extended functionality to support all features found in Snort rules. Reconfigurability and flexibility are key components of the framework that enable it to adapt to protect Internet systems from threats including malicious worms, computer viruses, and network intruders. To prove the framework viable, a system has been built that scans all bytes of transmission control protocol/Internet protocol (TCP/IP) traffic entering and leaving a network's gateway at multi-gigabit rates. Using Xilinx FPGA hardware on the field programmable port extender (FPX) platform, the framework can process 32,768 complex rules at data rates of 2.5 Gbps. Systems to handle data at 10 Gbps rates can be built today using the same framework in the latest reconfigurable hardware devices such as the Virtex 4. Michael Attig, John W. Lockwood |
FCCM | 2 |
| 2005 | A framework for rule processing in reconfigurable network systems (abstract only)abstractHigh-performance rule processing systems are needed by network administrators in order to protect Internet systems from attack. Researchers have been working to implement components of Intrusion Detection Systems, such as the highly popular Snort system, in reconfigurable hardware. While considerable progress has been made in the areas of string matching and header processing, complete systems have yet been demonstrated that effectively combine all of the functionality necessary to perform intrusion detection and prevention for real network systems.In this paper, we discuss a framework for implementing a rule processing system in reconfigurable hardware. The framework integrates the functionality to scan data flows for regular expressions, fixed strings, and header values. It also allows plug-in modules to be added to the system in order to perform the extended functionality to support the remaining set of Snort rules in modular components. Reconfigurability and flexibility are key components of the system that enable it adapt to protect Internet systems from threats including malicious worms, computer viruses, and network intruders.To prove the system viable, a system has been built that scans all bytes of Transmission Control protocol/Internet Protocol (TCP/IP) traffic entering and leaving a network's gateway at multi-gigabit rates and performs deep-packet inspection. Using Xilinx FPGA hardware on the Field programmable Port Extender (FPX), the platform can process 32,768 complex rules at data rates of 2.5 Gbps. Systems to handle data at 10 Gbps rates can be built today using the same framework in the most recent generation of reconfigurable hardware devices such as the Virtex4. Michael Attig, John W. Lockwood |
FPGA | 2 |
| 2005 | Efficient packet classification for network intrusion detection using FPGAabstractUsing FPGA technology for real-time network intrusion detection has gained many research efforts recently. In this paper, a novel packet classification architecture called BV-TCAM is presented, which is implemented for an FPGA-based Network Intrusion Detection System (NIDS). The classifier can report multiple matches at gigabit per second network link rates. The BV-TCAM architecture combines the Ternary Content Addressable Memory (TCAM) and the Bit Vector (BV) algorithm to effectively compress the data representations and boost throughput. A tree-bitmap implementation of the BV algorithm is used for source and destination port lookup while a TCAM performs the lookup of the other header fields, which can be represented as a prefix or exact value. The architecture eliminates the requirement for prefix expansion of port ranges. With the aid of a small embedded TCAM, packet classification can be implemented in a relatively small part of the available logic of an FPGA. The design is prototyped and evaluated in a Xilinx FPGA XCV2000E on the FPX platform. Even with the most difficult set of rules and packet inputs, the circuit is fast enough to sustain OC48 traffic throughput. Using larger and faster FPGAs, the system can work at speeds greater than OC192. Haoyu Song 0001, John W. Lockwood |
FPGA | 2 |
| 2005 | HAIL: A Hardware-Accelerated Algorithm for Language IdentificationabstractA hardware-accelerated algorithm has been designed to automatically identify the primary languages used in documents transferred over the Internet. The algorithm has been implemented in hardware on the field programmable port extender (FPX) platform. This system, referred to as the hardware-accelerated identification of languages (HAIL) project, identifies the primary languages used in content transferred over transmission control protocol (TCP)/Internet protocol (IP) networks that operate at rates exceeding 2.4 Gigabits/second. We demonstrate that this hardware accelerated circuit, operating on a Xilinx XCV2000E-8 FPGA, far outperforms software algorithms running on modern personal computers while maintaining extremely high levels of accuracy. Charles M. Kastner, G. Adam Covington, Andrew A. Levine, John W. Lockwood |
FPL | 4 |
| 2005 | Snort Offloader: A Reconfigurable Hardware NIDS FilterabstractSoftware-based network intrusion detection systems (NIDS) often fail to keep up with high-speed network links. In this paper an FPGA-based pre-filter is presented that reduces the amount of traffic sent to a software-based NIDS for inspection. Simulations using real network traces and the Snort rule set show that a pre-filter can reduce up to 90% of network traffic that would have otherwise been processed by Snort software. The projected performance enables a computer to perform real-time intrusion detection of malicious content passing over a 10 Gbps network using FPGA hardware that operates with 10 Gbps of throughput and software that needs only to operate with 1 Gbps of throughput. Haoyu Song 0001, Todd Sproull, Michael Attig, John W. Lockwood |
FPL | 4 |
| 2005 | Optimizing memory bandwidth of a multi-channel packet bufferabstractBackbone routers typically require large buffers to hold packets during congestion. A thumb rule is to provide a buffer at every link, equal to the product of the round trip time and the link capacity. This translates into Gigabytes of buffers operating at line rate at every link. Such a size and rate necessitates the use of SDRAM with bandwidth of, for example, 80 Gbps for link speed of 40 Gbps. With speedup in the switch fabrics used in most routers, the bandwidth requirement of the buffer increases further. While multiple SDRAM devices can be used in parallel to achieve high bandwidth and storage capacity, a wide logical data bus composed of these devices results in suboptimal performance for arbitrarily sized packets. An alternative is to divide the wide logical data bus into multiple logical channels and store packets into them independently. However, in such an organization, the cumulative pin count grows due to additional address buses which might offset the performance gained. We find that due to several existing memory technologies and their characteristics and with Internet traffic composed of particular sized packets, a judiciously architected data channel can greatly enhance the performance per pin. In this paper, we derive an expression for the effective memory bandwidth of a parallel channel packet buffer and show how it can be optimized for a given number of I/O pins available for interfacing to memory. We believe that our model can greatly aid packet buffer designers to achieve the best performance. Sarang Dharmapurikar, Sailesh Kumar, John W. Lockwood, Patrick Crowley |
GLOBECOM | 3 |
| 2005 | Advances for networks & internet
John W. Lockwood, George Kesidis |
GLOBECOM | 1 |
| 2005 | Multi-pattern signature matching for hardware network intrusion detection systemsabstractNetwork intrusion detection system (NIDS) performs deep inspections on the packet payload to identify, deter and contain the malicious attacks over the Internet. It needs to perform exact matching on multi-pattern signatures in real time. In this paper we introduce an efficient data structure called extended Bloom filter (EBF) and the corresponding algorithm to perform the multi-pattern signature matching. We also present a technique to support long signature matching so that we need only to maintain a limited number of supported signature lengths for the EBFs. We show that at reasonable hardware cost we can achieve very fast and almost time-deterministic exact matching for thousands of signatures. The architecture takes the advantages of embedded multi-port memories in FPGAs and can be used to build a full-featured hardware-based NIDS. Haoyu Song 0001, John W. Lockwood |
GLOBECOM | 2 |
| 2005 | Shape Shifting Tries for Faster IP Route LookupabstractSome of the fastest practical algorithms for IP route lookup are based on space-efficient encodings of multi-bit tries (M. Degermark, et al., 1997, W. Eatherton, 1999). Unfortunately, the time required by these algorithms grows in proportion to the address length, making them less attractive for IPv6. This paper describes and evaluates a new data structure called a shape-shifting trie, in which the data structure nodes correspond to arbitrarily shaped subtrees of the underlying binary trie for a given set of address prefixes. The ability to adapt the node shape to the trie reduces the number of nodes that must be accessed to perform a lookup, especially for tries with large sparse regions. We give a fast algorithm for optimally dividing a trie into nodes so as to minimize the maximum lookup depth. We show that seven data structure accesses are sufficient for route tables with more than 150,000 IPv6 prefixes. This makes it possible to achieve wire-speed processing for OC192 link using a single QDRII SRAM chip. Haoyu Song 0001, Jonathan S. Turner, John W. Lockwood |
ICNP | 3 |
| 2005 | Sensor fusion and correlationabstractNo abstract available. Todd Sproull, Richard Hough, John W. Lockwood, Christopher K. Zuver, Kent English, John Meier |
SenSys | 3 |
| 2005 | Fast hash table lookup using extended bloom filter: an aid to network processingabstractHash tables are fundamental components of several network processing algorithms and applications, including route lookup, packet classification, per-flow state management and network monitoring. These applications, which typically occur in the data-path of high-speed routers, must process and forward packets with little or no buffer, making it important to maintain wire-speed throughout. A poorly designed hash table can critically affect the worst-case throughput of an application, since the number of memory accesses required for each lookup can vary. Hence, high throughput applications require hash tables with more predictable worst-case lookup performance. While published papers often assume that hash table lookups take constant time, there is significant variation in the number of items that must be accessed in a typical hash table search, leading to search times that vary by a factor of four or more.We present a novel hash table data structure and lookup algorithm which improves the performance over a naive hash table by reducing the number of memory accesses needed for the most time-consuming lookups. This allows designers to achieve higher lookup performance for a given memory bandwidth, without requiring large amounts of buffering in front of the lookup engine. Our algorithm extends the multiple-hashing Bloom Filter data structure to support exact matches and exploits recent advances in embedded memory technology. Through a combination of analysis and simulations we show that our algorithm is significantly faster than a naive hash table using the same amount of memory, hence it can support better throughput for router applications that use hash tables. Haoyu Song 0001, Sarang Dharmapurikar, Jonathan S. Turner, John W. Lockwood |
SIGCOMM | 4 |
| 2004 | Implementation Results of Bloom Filters for String MatchingabstractNetwork intrusion detection and prevention systems (IDPS) use string matching to scan Internet packets for malicious content. Bloom filters offer a mechanism to search for a large number of strings efficiently and concurrently when implemented with field programmable gate array (FPGA) technology. A string matching circuit has been implemented within the FPX platform using Bloom filters. Using 155 block RAMs on a single Xilinx VirtexE 2000 FPGA, the circuit scans for 35,475 unique signatures. Michael Attig, Sarang Dharmapurikar, John W. Lockwood |
FCCM | 3 |
| 2004 | Secure Remote Control of Field-programmable Network DevicesabstractA circuit and an associated lightweight protocol have been developed to secure communication between a control console and remote programmable network devices. The circuit provides encryption, data integrity checking and sequence number verification to ensure confidentiality, integrity and authentication of control messages sent over the public Internet. All of these functions are performed directly in FPGA hardware to provide high throughput and near-zero latency. The circuit has been used to control and configure remote firewalls and intrusion detection systems. The circuit could also be used to control and configure other distributed network applications. Haoyu Song 0001, John W. Lockwood, James Moscola |
FCCM | 3 |
| 2004 | Automated Method to Generate Bitstream Intellectual Property Cores for Virtex FPGAs
Edson Lemos Horta, John W. Lockwood |
FPL | 2 |
| 2004 | A Modular System for FPGA-Based TCP Flow Processing in High-Speed Networks
David V. Schuehler, John W. Lockwood |
FPL | 2 |
| 2003 | Implementation of a Content-Scanning Module for an Internet FirewallabstractA module has been implemented in Field Programmable Gate Array (FPGA) hardware that scans the content of Internet packets at Gigabits/second rates. All of the packet processing operations are performed using reconfigurable hardware within a single Xilinx Virtex XCV2000E FPGA. A set of layered protocol wrappers is used to parse the headers and payloads of packets for Internet protocol data. A content matching server automatically generates the Finite State Machines (FSMs) to search for regular expressions. The complete system is operated on the Field-programmable Port Extender (FPX) platform. James Moscola, John W. Lockwood, Ronald Prescott Loui, Michael Pachos |
FCCM | 2 |
| 2003 | An Extensible, System-On-Programmable-Chip, Content-Aware Internet Firewall
John W. Lockwood, Christopher E. Neely, Christopher K. Zuver, James Moscola, Sarang Dharmapurikar, David Lim |
FPL | 1 |
| 2003 | A TCP/IP Based Multi-device Programming Circuit
David V. Schuehler, Harvey Ku, John W. Lockwood |
FPL | 3 |
| 2003 | Beyond performance: secure and fair memory management for multiple systems on a chipabstractDevelopments in VLSI technologies create the possibility of hosting several independent (sub) systems in a single chip. There is a need to share a number of resources, especially off-chip resources, which creates new constraints in the design process. Although performance is still a key constraint, sharing implies that secure access to those resources and QoS guarantees are needed. In this paper, an architecture is presented that achieves the goals listed above. The Embedded Hardware Manager acts as a middleware between the applications and the resources, taking the role of resource manager and security agent. The results show that it can prevent resource misuse and undue information peeking or even altering while maintaining individual QoS guarantees. At the same time, high performance is still achieved. Carlos Macián, Sarang Dharmapurikar, John W. Lockwood |
FPT | 3 |
| 2003 | System-on-chip packet processor for an experimental network services platformabstractAs the focus of networking research shifts from raw performance to the delivery of advanced network services, there is a growing need for open-platform systems for extensible networking research. The Applied Research Laboratory at Washington University in Saint Louis has developed a flexible network services platform (NSP) to meet this need. The NSP provides an extensible platform for prototyping next-generation network services and applications. The paper describes the design of a system-on-chip packet processor for the NSP which performs all core packet processing functions, including segmentation and reassembly, packet classification, route lookup, and queue management. Targeted to a commercial configurable logic device, the system is designed to support gigabit links and switch fabrics with a 2:1 speed advantage. We provide resource consumption results for each component of the packet processor design. David E. Taylor, Alex Chandra, Sarang Dharmapurikar, John W. Lockwood, Wenjing Tang, Jonathan S. Turner |
GLOBECOM | 5 |
| 2003 | Scalable IP lookup for Internet routersabstractInternet protocol (IP) address lookup is a central processing function of Internet routers. While a wide range of solutions to this problem have been devised, very few simultaneously achieve high lookup rates, good update performance, high memory efficiency, and low hardware cost. High performance solutions using content addressable memory devices are a popular but high-cost solution, particularly when applied to large databases. We present an efficient hardware implementation of a previously unpublished IP address lookup architecture, invented by Eatherton and Dittia (see M.S. thesis, Washington Univ., St. Louis, MO, 1998). Our experimental implementation uses a single commodity synchronous random access memory chip and less than 10% of the logic resources of a commercial configurable logic device, operating at 100 MHz. With these quite modest resources, it can perform over 9 million lookups/s, while simultaneously processing thousands of updates/s, on databases with over 100000 entries. The lookup structure requires 6.3 bytes per address prefix: less than half that required by other methods. The architecture allows performance to be scaled up by using parallel fast IP lookup (FIPL) engines, which interleave accesses to a common memory interface. This architecture allows performance to scale up directly with available memory bandwidth. We describe the tree bitmap algorithm, our implementation of it in a dynamically extensible gigabit router being developed at Washington University in Saint Louis, and the results of performance experiments designed to assess its performance under realistic operating conditions. David E. Taylor, Jonathan S. Turner, John W. Lockwood, Todd Sproull, David B. Parlour |
IEEE J. Sel. Areas Commun. | 3 |
| 2002 | Dynamic hardware plugins in an FPGA with partial run-time reconfigurationabstractTools and a design methodology have been developed to support partial run-time reconfiguration of FPGA logic on the Field Programmable Port Extender. High-speed Internet packet processing circuits on this platform are implemented as Dynamic Hardware Plugin (DHP) modules that fit within a specific region of an FPGA device. The PARBIT tool has been developed to transform and restructure bitfiles created by standard computer aided design tools into partial bitsteams that program DHPs. The methodology allows the platform to hot-swap application-specific DHP modules without disturbing the operation of the rest of the system. Edson Lemos Horta, John W. Lockwood, David E. Taylor, David B. Parlour |
DAC | 2 |
| 2002 | Control and Configuration Software for a Reconfigurable Networking Hardware PlatformabstractA suite of tools called NCHARGE (Networked Configurable Hardware Administrator for Reconfiguration and Governing via End-systems) has been developed to simplify the co-design of hardware and software components that process packets within a network of Field Programmable Gate Arrays (FPGAs). A key feature of NCHARGE is that it provides a high-performance packet interface to hardware and standard Application Programming Interface (API) between software and reprogrammable hardware modules. Using this API, multiple software processes can communicate to one or more hardware modules using standard TCP/IP sockets. NCHARGE also provides a Web-Based User Interface to simplify the configuration and control of an entire network switch that contains several software and hardware modules. Todd Sproull, John W. Lockwood, David E. Taylor |
FCCM | 2 |
| 2002 | Using PARBIT to Implement Partial Run-Time Reconfigurable Systems
Edson Lemos Horta, John W. Lockwood, Sergio Takeo Kofuji |
FPL | 2 |
| 2002 | Scalable IP Lookup for Programmable RoutersabstractContinuing growth in optical link speeds places increasing demands on the performance of Internet routers, while deployment of embedded and distributed network services imposes new demands for flexibility and programmability. IP address lookup has become a significant performance bottleneck for the highest performance routers. Amid the vast array of academic and commercial solutions to the problem, few achieve a favorable balance of performance, efficiency, and cost. New commercial products utilize content addressable memory (CAM) devices to achieve high lookup speeds at an exorbitantly high hardware cost with limited flexibility. In contrast, this paper describes an efficient, scalable lookup engine design, able to achieve high performance with the use of a small portion of a reconfigurable logic device and a commodity random access memory (RAM) device. The Fast Internet Protocol Lookup (FIPL) engine is an implementation of Eatherton and Dittia's previously unpublished Tree Bitmap algorithm (1998) targeted to an open-platform research router. FIPL can be scaled to achieve guaranteed worst-case performance of over 9 million lookups per second with a single SRAM operating at the fairly modest clock speed of 100 MHz. Experimental evaluation of FIPL throughput, latency, and update performance is provided using a sample routing table from Mae West. David E. Taylor, John W. Lockwood, Todd Sproull, Jonathan S. Turner, David B. Parlour |
INFOCOM | 2 |
| 2002 | Dynamic hardware plugins: exploiting reconfigurable hardware for high-performance programmable routers
David E. Taylor, Jonathan S. Turner, John W. Lockwood, Edson Lemos Horta |
Comput. Networks | 3 |
| 2001 | Reprogrammable network packet processing on the field programmable port extender (FPX)abstractA prototype platform has been developed that allows processing of packets at the edge of a multi-gigabit-per-second network switch. This system, the Field Programmable Port Extender (FPX), enables packet processing functions to be implemented as modular components in reprogrammable hardware. All logic on the on the FPX is implemented in two Field Programmable Gate Arrays (FPGAs). Packet processing functions in the system are implemented as dynamically-loadable modules. John W. Lockwood, Naji Naufel, Jonathan S. Turner, David E. Taylor |
FPGA | 1 |
| 2001 | Reconfigurable Router Modules Using Network Protocol Wrappers
Florian Braun, John W. Lockwood, Marcel Waldvogel |
FPL | 2 |
| 2000 | Field programmable port extender (FPX) for distributed routing and queuingabstractField Programmable Gate Arrays (FPGAs) are being used to provide fast Internet Protocol (IP) packet routing and advanced queuing in a highly scalable network switch. A new module, called the Field-programmable Port Extender (FPX), is being built to augment the Washington University Gigabit Switch (WUGS) with reprogrammable logic. John W. Lockwood, Jonathan S. Turner, David E. Taylor |
FPGA | 1 |
| 2000 | Compensation modeling for QoS support on a wireless networkabstractThis paper presents the design of a new error compensation model for providing QoS support on a wireless network. The compensation model uses different compensation strategies for the different priority classes in conjunction with a multiclass priority fair queuing (MPFQ) algorithm. Criteria for fair error compensation and a classification of error compensation models are established in order to show the properties of different models, and to select the best error compensation for each priority class. Simulation results show that the new MPFQ compensation model meets the long-term fairness guarantees and provides an improved flow separation. Stefan Bucheli, Jay R. Moorman, John W. Lockwood |
GLOBECOM | 3 |
| 1999 | Implementation of campus-wide wireless network services using ATM, virtual LANs, and wireless basestationsabstractA campus-wide wireless LAN has been implemented that combines the scalability of virtual LANs and the flexibility of switched Ethernet with the mobility of wireless LAN access. A suite of general purpose software tools has been developed and deployed to monitor the status of the network, measure the bandwidth utilization, control access to the network, and authenticate users. John W. Lockwood |
WCNC | 1 |
| 1997 | A High-Performance OC-12/OC-48 Queue Design Prototype for Input-buffered ATM SwitchesabstractThis paper presents the design and prototype of an intelligent, 3-dimensional-queue (3DQ) for high-performance, scalable, input-buffered ATM switches. The 3DQ uses pointers and linked lists to organize ATM cells into multiple virtual queues according to priority, destination, and virtual connection. It enforces per-virtual connection quality-of-service (QoS) and eliminates head-of-line (HOL) blocking. Using field-programmable-gate-array (FPGA) devices, our prototype hardware can process ATM cells at 622 Mb/s (OC-12). Using more aggressive technology (multi-chip-module (MCM) and fast GaAs logic), the same 3DQ can process cells at 2.5 Gb/s (OC-48). Using the 3DQ and matrix-unit-cell-scheduler (MUCS) as essential components, an input-buffered ATM switch system has been designed, which can achieve near-100% link bandwidth utilization. John W. Lockwood, J. D. Will |
INFOCOM | 2 |