EDBT 2026 Demo / reviewers in the wild / expert
Reza Rahimi
dblp:143/7528
· DBLP profile ↗
10ranked-venue papers
2as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Memory systems · 42% Hardware accelerators and domain-specific architectures · 34% Reconfigurable computing and FPGAs · 24% | |
| Artificial intelligence
1 paper |
Autonomous driving · 44% Reinforcement learning · 44% Robot navigation and mapping · 13% | |
| Network and information security
1 paper |
Network security · 100% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems › in-memory computing
in-memory automata accelerator |
1.3 | 3 | 2021 | Sunder: Enabling Low-Overhead and Scalable Near-Data Pattern Matching Acceleration · MICRO 2021 Impala: Algorithm/Architecture Co-Design for In-Memory Multi-Stride Pattern Matching · HPCA 2020 eAP: A Scalable and Efficient In-Memory Accelerator for Automata Processing · MICRO 2019 |
Hardware accelerators and domain-specific architectures › pattern matching accelerator
automata processor |
1.2 | 3 | 2020 | Impala: Algorithm/Architecture Co-Design for In-Memory Multi-Stride Pattern Matching · HPCA 2020 FlexAmata: A Universal and Efficient Adaption of Applications to Spatial Automata Processing Accelerators · ASPLOS 2020 A Scalable Solution for Rule-Based Part-of-Speech Tagging on Novel Hardware Accelerators · KDD 2018 |
Reconfigurable computing and FPGAs
automata processing |
0.6 | 3 | 2021 | ASPEN: A Scalable In-SRAM Architecture for Pushdown Automata · MICRO 2018 Sunder: Enabling Low-Overhead and Scalable Near-Data Pattern Matching Acceleration · MICRO 2021 eAP: A Scalable and Efficient In-Memory Accelerator for Automata Processing · MICRO 2019 |
Reconfigurable computing and FPGAs
FPGA routing architecture |
0.4 | 1 | 2019 | eAP: A Scalable and Efficient In-Memory Accelerator for Automata Processing · MICRO 2019 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.3 | 1 | 2018 | Navigating Occluded Intersections with Autonomous Vehicles Using Deep Reinforcement Learning · ICRA 2018 |
Robotics › Autonomous driving › autonomous vehicle navigation
intersection navigation |
0.3 | 1 | 2018 | Navigating Occluded Intersections with Autonomous Vehicles Using Deep Reinforcement Learning · ICRA 2018 |
Memory systems
in-memory computing |
0.3 | 1 | 2018 | ASPEN: A Scalable In-SRAM Architecture for Pushdown Automata · MICRO 2018 |
Hardware accelerators and domain-specific architectures › pattern matching accelerator
regular expression matching accelerator |
0.3 | 1 | 2018 | A Scalable Solution for Rule-Based Part-of-Speech Tagging on Novel Hardware Accelerators · KDD 2018 |
Memory systems
processing-in-memory |
0.3 | 2 | 2021 | Sunder: Enabling Low-Overhead and Scalable Near-Data Pattern Matching Acceleration · MICRO 2021 eAP: A Scalable and Efficient In-Memory Accelerator for Automata Processing · MICRO 2019 |
Network security › intrusion detection and prevention
intrusion detection |
0.1 | 1 | 2020 | Impala: Algorithm/Architecture Co-Design for In-Memory Multi-Stride Pattern Matching · HPCA 2020 |
Network security › intrusion detection and prevention › intrusion detection
pattern matching |
0.1 | 1 | 2020 | Impala: Algorithm/Architecture Co-Design for In-Memory Multi-Stride Pattern Matching · HPCA 2020 |
Robotics › Robot navigation and mapping › active perception
active sensing |
0.1 | 1 | 2018 | Navigating Occluded Intersections with Autonomous Vehicles Using Deep Reinforcement Learning · ICRA 2018 |
Data mining › pattern mining › tree mining
frequent subtree mining |
0.1 | 1 | 2018 | ASPEN: A Scalable In-SRAM Architecture for Pushdown Automata · MICRO 2018 |
Methods — techniques the papers use, named apart from their topics
algorithm-architecture co-design · 0.9in-memory computing · 0.5automata processing · 0.5finite automata mapping · 0.4place and route · 0.4finite automata · 0.4regular expression matching · 0.3deep reinforcement learning · 0.3automata processor · 0.3FPGA · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Sunder: Enabling Low-Overhead and Scalable Near-Data Pattern Matching AccelerationabstractAutomata processing is an efficient computation model for regular expressions and other forms of sophisticated pattern matching. The demand for high-throughput and real-time pattern matching in many applications, including network intrusion detection and spam filters, has motivated several in-memory architectures for automata processing. Existing in-memory architectures focus on accelerating the pattern-matching kernel, but either fail to support a practical reporting solution or optimistically assume that the reporting stage is not the performance bottleneck. However, gathering and processing the reports can be the major bottleneck, especially when the reporting frequency is high. Moreover, all the existing in-memory architectures work with a fixed processing rate (mostly 8-bit/cycle), and they do not adjust the input consumption rate based on the properties of the applications, which can lead to throughput and capacity loss. Elaheh Sadredini, Reza Rahimi, Mohsen Imani, Kevin Skadron |
MICRO | 2 |
| 2020 | FlexAmata: A Universal and Efficient Adaption of Applications to Spatial Automata Processing AcceleratorsabstractPattern matching, especially for complex patterns with many variations, is an important task in many big-data applications and maps well to finite automata. Recently, a variety of research has focused on hardware acceleration of automata processing, especially via spatial architectures that directly map the patterns to massively parallel hardware elements, such as in FPGAs and in-memory solutions. We observed that all existing automata-acceleration architectures are designed based on fixed, 8-bit symbol processing, derived from ASCII processing. However, the alphabet size in pattern-matching applications varies from just a few up to billions of unique symbols. This makes it difficult to provide a universal and efficient mapping of this wide variety of automata applications to existing automata accelerators. Elaheh Sadredini, Reza Rahimi, Marzieh Lenjani, Mircea R. Stan, Kevin Skadron |
ASPLOS | 2 |
| 2020 | Grapefruit: An Open-Source, Full-Stack, and Customizable Automata Processing on FPGAsabstractRegular expressions have been widely used in various application domains such as network security, machine learning, and natural language processing. Increasing demand for accelerated regular expressions, or equivalently finite automata, has motivated many efforts in designing FPGA accelerators. However, there is no framework that is publicly available, comprehensive, parameterizable, general, full-stack, and easy-touse, all in one, for design space exploration for a wide range of growing pattern matching applications on FPGAs. In this paper, we present Grapefruit, the first open-source, full-stack, efficient, scalable, and extendable automata processing framework on FPGAs. Grapefruit is equipped with an integrated compiler with many parameters for automata simulation, verification, minimization, transformation, and optimizations. Our modular and standard design allows researchers to add capabilities and explore various features for a target application. Our experimental results show that the hardware generated by Grapefruit performs 9%80% better than prior work that is not fully end-to-end and has 3.4 × higher throughput in a multi-stride solution than a single-stride solution. Reza Rahimi, Elaheh Sadredini, Mircea R. Stan, Kevin Skadron |
FCCM | 1 |
| 2020 | Impala: Algorithm/Architecture Co-Design for In-Memory Multi-Stride Pattern MatchingabstractHigh-throughput and concurrent processing of thousands of patterns on each byte of an input stream is critical for many applications with real-time processing needs, such as network intrusion detection, spam filters, virus scanners, and many more. The demand for accelerated pattern matching has motivated several recent in-memory accelerator architectures for automata processing, which is an efficient computation model for pattern matching. Our key observations are: (1) all these architectures are based on 8-bit symbol processing (derived from ASCII), and our analysis on a large set of real-world automata benchmarks reveals that the 8-bit processing dramatically under-utilizes hardware resources, and (2) multi-stride symbol processing, a major source of throughput growth, is not explored in the existing in-memory solutions. This paper presents Impala, a multi-stride in-memory automata processing architecture by leveraging our observations. The key insight of our work is that transforming 8-bit processing to 4-bit processing exponentially reduces hardware resources for state-matching and improves resource utilization. This, in turn, brings the opportunity to have a denser design, and be able to utilize more memory columns to process multiple symbols per cycle with a linear increase in state-matching resources. Impala thus introduces threefold area, throughput, and energy benefits at the expense of increased offline compilation time. Our empirical evaluations on a wide range of automata benchmarks reveal that Impala has on average 2.7× (up to 3.7×) higher throughput per unit area and 1.22× lower power consumption than Cache Automaton, which is the best performing prior work. Elaheh Sadredini, Reza Rahimi, Marzieh Lenjani, Mircea R. Stan, Kevin Skadron |
HPCA | 2 |
| 2019 | eAP: A Scalable and Efficient In-Memory Accelerator for Automata ProcessingabstractAccelerating finite automata processing benefits regular-expression workloads and a wide range of other applications that do not map obviously to regular expressions, including pattern mining, bioinformatics, and machine learning. Existing in-memory automata processing accelerators suffer from inefficient routing architectures. They are either incapable of efficiently place-and-route a highly connected automaton or require an excessive amount of hardware resources. Elaheh Sadredini, Reza Rahimi, Vaibhav Verma, Mircea R. Stan, Kevin Skadron |
MICRO | 2 |
| 2018 | Navigating Occluded Intersections with Autonomous Vehicles Using Deep Reinforcement LearningabstractProviding an efficient strategy to navigate safely through unsignaled intersections is a difficult task that requires determining the intent of other drivers. We explore the effectiveness of Deep Reinforcement Learning to handle intersection problems. Using recent advances in Deep RL, we are able to learn policies that surpass the performance of a commonly-used heuristic approach in several metrics including task completion time and goal success rate and have limited ability to generalize. We then explore a system's ability to learn active sensing behaviors to enable navigating safely in the case of occlusions. Our analysis, provides insight into the intersection handling problem, the solutions learned by the network point out several shortcomings of current rule-based methods, and the failures of our current deep reinforcement learning system point to future research directions. David Isele, Reza Rahimi, Akansel Cosgun, Kaushik Subramanian, Kikuo Fujimura |
ICRA | 2 |
| 2018 | A Scalable Solution for Rule-Based Part-of-Speech Tagging on Novel Hardware AcceleratorsabstractPart-of-speech (POS) tagging is the foundation of many natural language processing applications. Rule-based POS tagging is a wellknown solution, which assigns tags to the words using a set of predefined rules. Many researchers favor statistical-based approaches over rule-based methods for better empirical accuracy. However, until now, the computational cost of rule-based POS tagging has made it difficult to study whether more complex rules or larger rulesets could lead to accuracy competitive with statistical approaches. In this paper, we leverage two hardware accelerators, the Automata Processor (AP) and Field Programmable Gate Arrays (FPGA), to accelerate rule-based POS tagging by converting rules to regular expressions and exploiting the highly-parallel regular-expressionmatching ability of these accelerators. We study the relationship between rule set size and accuracy, and observe that adding more rules only poses minimal overhead on the AP and FPGA. This allows a substantial increase in the number and complexity of rules, leading to accuracy improvement. Our experiments on Treebank and Brown corpora achieve up to 2,600X and 1,914X speedups on the AP and on the FPGA respectively over rule-based methods on the CPU in the rule-matching stage, up to 58× speedup over the Perceptron POS tagger on the CPU in total testing time, and up to 253× speedup over the LSTM tagger on the GPU in total testing time, while showing a competitive accuracy compared to neural-network and statistical solutions. Elaheh Sadredini, Deyuan Guo, Chunkun Bo, Reza Rahimi, Kevin Skadron, Hongning Wang |
KDD | 4 |
| 2018 | ASPEN: A Scalable In-SRAM Architecture for Pushdown AutomataabstractMany applications process some form of tree-structured or recursively-nested data, such as parsing XML or JSON web content as well as various data mining tasks. Typical CPU processing solutions are hindered by branch misprediction penalties while attempting to reconstruct nested structures and also by irregular memory access patterns. Recent work has demonstrated improved performance for many data processing applications through memory-centric automata processing engines. Unfortunately, these architectures do not support a computational model rich enough for tasks such as XML parsing. In this paper, we present ASPEN, a general-purpose, scalable, and reconfigurable memory-centric architecture for processing of tree-like data. We take inspiration from previous automata processing architectures, but support the richer deterministic pushdown automata computational model. We propose a custom datapath capable of performing the state matching, stack manipulation, and transition routing operations of pushdown automata, all efficiently stored and computed in memory arrays. Further, we present compilation algorithms for transforming large classes of existing grammars to pushdown automata executable on ASPEN, and demonstrate their effectiveness on four different languages: Cool (object oriented programming), DOT (graph visualization), JSON, and XML. Finally, we present an empirical evaluation of two application scenarios for ASPEN: XML parsing, and frequent subtree mining. The proposed architecture achieves an average 704.5 ns per KB parsing XML compared to 9983 ns per KB in a state-of-the-art XML parser across 23 benchmarks. We also demonstrate a 37.2x and 6x better end-to-end speedup over CPU and GPU implementations of subtree mining. Kevin Angstadt, Arun Subramaniyan 0001, Elaheh Sadredini, Reza Rahimi, Kevin Skadron, Westley Weimer, Reetuparna Das |
MICRO | 4 |
| 2017 | Frequent subtree mining on the automata processor: challenges and opportunitiesabstractFrequency counting of complex patterns such as subtrees is more challenging than for simple itemsets and sequences, as the number of possible candidate patterns in a tree is much higher than one-dimensional data structures, with dramatically higher processing times. In this paper, we propose a new and scalable solution for frequent subtree mining (FTM) on the Automata Processor (AP), a new and highly parallel accelerator architecture. We present a multi-stage pruning framework on the AP, called AP-FTM, to reduce the search space of FTM candidates. This achieves up to 353X speedup at the cost of a small reduction in accuracy, on four real-world and synthetic datasets, when compared with PatternMatcher, a practical and exact CPU solution. To provide a fully accurate and still scalable solution, we propose a hybrid method to combine AP-FTM with a CPU exact-matching approach, and achieve up to 262X speedup over PatternMatcher on a challenging database. We also develop a GPU algorithm for FTM, but show that the AP also outperforms this. The results on a synthetic database show the AP advantage grows further with larger datasets. Elaheh Sadredini, Reza Rahimi, Ke Wang 0011, Kevin Skadron |
ICS | 2 |
| 2016 | A high-performance OpenFlow software switchabstractSoftware switches offer flexibility to service providers but potentially suffer from low performance. A software switch called Lagopus was implemented using Intel's Data Plane Development Kit (DPDK), which offers libraries for high-performance packet handling. Prior work on software switches focused on characterizing packet forwarding throughput. In this work, we evaluated the impact of certain parameters and settings in Lagopus on application performance and studied packet drop rates. The importance of receive-thread packet classification for load balancing and to send delay-sensitive flows to a different worker thread from high-throughput flows was first demonstrated. Next, we showed that a loop-count variable used to control packet batching should be kept small in case link utilization is low. Finally, we showed that packet drop rate could be non-zero when the OpenFlow table size is large and packet arrival rate is high, and interestingly, the packet drop rate was higher with four worker threads than with a single worker thread. This implies a need for careful calibration and planning of the parameters of parallelization. Reza Rahimi, Malathi Veeraraghavan, Yoshihiro Nakajima, Hirokazu Takahashi, Yusuke Nakajima, Satoru Okamoto, Naoaki Yamanaka |
HPSR | 1 |