EDBT 2026 Demo / reviewers in the wild / expert
Marziyeh Nourian
dblp:183/4737
· DBLP profile ↗
9ranked-venue papers
4as first author
3since 2021 · last 2026
0009-0007-5124-8283ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
High-performance computing · 47% Hardware accelerators and domain-specific architectures · 18% Electronic design automation · 18% | |
| Theoretical computer science
1 paper |
Graph algorithms and graph theory · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing
supercomputer architecture |
1.0 | 1 | 2026 | UpDown: A Supercomputer Co-Designed for Scalable Graph Processing · IEEE Trans. Parallel Distributed Syst. 2026 |
Hardware accelerators and domain-specific architectures › pattern matching accelerator
automata processor |
0.4 | 1 | 2019 | Evaluating High Performance Pattern Matching on the Automata Processor · IEEE Trans. Computers 2019 |
Electronic design automation › hardware verification and test
pattern matching |
0.4 | 1 | 2019 | Evaluating High Performance Pattern Matching on the Automata Processor · IEEE Trans. Computers 2019 |
Reconfigurable computing and FPGAs
reconfigurable computing |
0.4 | 1 | 2019 | Evaluating High Performance Pattern Matching on the Automata Processor · IEEE Trans. Computers 2019 |
Graph algorithms and graph theory
graph processing |
0.3 | 1 | 2026 | UpDown: A Supercomputer Co-Designed for Scalable Graph Processing · IEEE Trans. Parallel Distributed Syst. 2026 |
Methods — techniques the papers use, named apart from their topics
co-design · 2.0NFA emulation · 0.4FPGA implementation · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UpDown: Efficient Manycore based on Many Threading and Scalable Memory ParallelismabstractManycore architectures are a promising direction for single-chip performance. They typically use in-order cores and caches as building blocks and can produce good performance on regular applications. However, on irregular applications, they have low core utilization due to data-dependent control-flow and memory access. Andronicus Rajasukumar, Ruiqi Xu 0001, Tianchi Zhang 0005, Yuqing Wang 0011, Tianshuo Su, Marziyeh Nourian, Jianru Ding, Jiya Su, Rajat Khandelwal, Alexander Fell, David F. Gleich, Yanjing Li, Henry Hoffmann, Andrew A. Chien |
ICS | 6 |
| 2026 | UpDown: A Supercomputer Co-Designed for Scalable Graph ProcessingabstractTraditional supercomputers have focused on dense computation performance as exemplified by HPL. Graph processing applications differ with extreme irregularity ($10^{9}$imbalance in skewed, real-world graphs) that produces unpredictable work, parallelism, memory access, and communication. Together, these make scalable performance and programming difficult. We describe the UpDown system architecture, co-designed for irregular graph computations. UpDown provides efficient fine-grained thread invocations ($\sim$10 instructions), direct messaging (no network interface card) for scalable local and global messaging, and split-transaction memory operations that enable extremely high memory bandwidth. Combined with architectural support for global addressing and an aggressive network design, these UpDown features enable direct exploitation of edge and vertex parallelism, using it to deliver breakthrough graph processing performance and programmability. We evaluate the performance of the UpDown system using a challenging suite of graph applications (Pagerank, Breadth-first Search, Triangle Counting, Partial Match, etc). For a single-node, results show 100-fold performance advantage over multicore CPUs. Compared to today's fastest scalable parallel computers UpDown achieves 1000-fold performance increases. UpDown delivers these levels of performance with high-level programmability, these programs directly express vertex-edge parallelism which UpDown exploits directly in hardware. Andrew A. Chien, Charles Colley, Jianru Ding, Alexander Fell, David F. Gleich, Henry Hoffmann, Moubarak Jeje, Rajat Khandelwal, Yanjing Li, Jose M. Monsalve Diaz, Marziyeh Nourian, Andronicus Rajasukumar, Jiya Su, Tianshuo Su, Yuqing Wang 0011, Ruiqi Xu 0001, Tianchi Zhang 0005 |
IEEE Trans. Parallel Distributed Syst. | 11 |
| 2022 | Data Transformation Acceleration using Deterministic Finite-State TransducersabstractData transformation tasks are a critical and costly part of many data processing and analytics applications. A simple computing model that can efficiently represent data transformation and be mapped to different platforms can provide programmers with the flexibility o f u sing different data representations and allow for exploiting different platforms, including general-purpose processors and accelerators.We propose extended Deterministic Finite State Transducers (DFST+s), a computing model that enables the compact expression of data transformations (a significantly terser expression compared to the DFSTs model, a traditional computational abstraction for data transformation), aiding their correct and efficient implementation. We define the TF ORM language to facilitate expressing the DFST+, and the TFORM virtual machine to enable a further compact expression, leading to a high performance and portable implementation. We propose two TFORM VM execution models and evaluate them using a variety of data transformations (from Apache Parquet file format and sparse matrices). Our results show both effective portability across CPU and a hardware accelerator, and performance increases of 1.7× and 11.7× geometric mean, respectively, over a custom CPU implementation of the same transformations. Marziyeh Nourian, Tri Nguyen 0002, Andrew A. Chien, Michela Becchi |
IEEE Big Data | 1 |
| 2020 | Optimizing Complex OpenCL Code for FPGA: A Case Study on Finite Automata TraversalabstractWhile FPGAs have been traditionally considered hard to program, recently there have been efforts aimed to allow the use of high-level programming models and libraries intended for multi-core CPUs and GPUs to program FPGAs. For example, both Intel and Xilinx are now providing toolchains to deploy OpenCL code onto FPGA. However, because the nature of the parallelism offered by GPU and FPGA devices is fundamentally different, OpenCL code optimized for GPU can prove very inefficient on FPGA, in terms of both performance and hardware resource utilization. This paper explores this problem on finite automata traversal. In particular, we consider an OpenCL NFA traversal kernel optimized for GPU but exhibiting FPGA-friendly characteristics, namely: limited memory requirements, lack of synchronization, and SIMD execution. We explore a set of structural code changes, custom and best-practice optimizations to retarget this code to FPGA. We showcase the effect of these optimizations on an Intel Stratix V FPGA board using various NFA topologies from different application domains. Our evaluation shows that, while the resource requirements of the original code exceed the capacity of the FPGA in use, our optimizations lead to significant resource savings and allow the transformed code to fit the FPGA for all considered NFA topologies. In addition, our optimizations lead to speedups up to 4x over an already optimized code-variant aimed to fit the NFA traversal kernel on FPGA. Some of the proposed optimizations can be generalized for other applications and introduced in OpenCL-to-FPGA compiler. Marziyeh Nourian, Mostafa Eghbali Zarch, Michela Becchi |
ICPADS | 1 |
| 2019 | Evaluating High Performance Pattern Matching on the Automata ProcessorabstractIn this paper, we study the acceleration of applications that identify all the occurrences of thousands of string-patterns in an input data-stream using the Automata Processor (AP). For this evaluation, we use two applications from two fields, namely, cybersecurity and bioinformatics. The first application, called Fast-SNAP, scans network data for 4312 signatures of intrusion derived from the popular open-source Snort database. Using the resources of a single AP-board, Fast-SNAP can scan for all these signatures at 1 Gbps. The second application, called PROTOMATA, looks for all the occurrences of 1,309 motifs from the PROSITE database in protein sequences. PROTOMATA is up to 68 times faster than the state-of-the-art CPU implementation. As a comparison, we emulate the execution of the same NFAs by programming FPGAs using state-of-the-art techniques. We find that the performance derived by using the resources of a single AP-board, which houses 32 AP-chips, is comparable to that of the resources of five to six large FPGAs. The design techniques used in this paper are generic and may be applicable to the development of similar applications on the AP. Indranil Roy, Ankit Srivastava, Matt Grimm, Marziyeh Nourian, Michela Becchi, Srinivas Aluru |
IEEE Trans. Computers | 4 |
| 2018 | A Compiler Framework for Fixed-Topology Non-Deterministic Finite Automata on SIMD PlatformsabstractAutomata processing is at the core of many pattern matching applications. Previous efforts have proposed designs to accelerate automata traversal on various parallel platforms. Many existing acceleration methods store the finite automata states and transitions in memory. As a result, for these designs memory size and bandwidth are the main limiting factors to performance and power efficiency. Many applications, however, require processing several fixed-topology automata that differ only in the symbols associated to the transitions. This property enables the design of alternative, memory-efficient solutions. In this work, we target fixed-topology non-deterministic finite automata (NFAs) and propose a memory-efficient design that embeds the automata topology in code and stores only the transition symbols in memory. Our solution is suitable to SIMD architectures. In particular, we propose a compiler framework that, given a set of NFAs with fixed topology and the hardware configuration of the target SIMD platform, deploys the NFAs and the input streams to be processed onto the target device so as to exploit the available parallelism, maximize hardware utilization and optimize the memory access patterns. Our compiler framework performs a combination of platform-agnostic and platform-specific design decisions and optimizations. We showcase our compiler framework on GPU, Intel Xeon Phi and Intel Xeon Skylake platforms. Marziyeh Nourian, Hancheng Wu, Michela Becchi |
ICPADS | 1 |
| 2017 | A Memory-Efficient GPU Method for Hamming and Levenshtein Distance SimilarityabstractSeveral applications from computational linguistic and genomics perform similarity analysis between sequences of characters based on Hamming and Levenshtein distance measures. Hamming and Levenshtein distance-based matching maps well onto non-deterministic finite automata (NFAs), which have been accelerated on GPUs. However, designed with the flexibility to support generic topologies, existing NFA engines have inefficiencies when processing fixed-topology NFAs. In this work we target this problem and propose two methods to improve the preprocessing and traversal performance of Levenshtein and Hamming distance NFAs. Our methods are based on the following observation: for these fixed-topologies, the transitions do not need to be stored in device memory, but they can be inferred from the reference string (i.e., the string to be matched against) and the NFA topology alone. Our first, basic implementation (implicit-active-sq) minimizes preprocessing by bypassing NFA construction and packing, but exhibits several traversal inefficiencies. Our optimized method (implicit-rearranged-sv) includes space and time optimizations for global memory and radically different shared memory access patterns within the kernel, while at the same time incurring only modest preprocessing overhead over the basic implementation. Our experimental evaluation shows that, on large NFAs consisting of several millions states, implicit-rearranged-sv outperforms traditional GPU engines both in terms of traversal throughput (3-22x speedup) and preprocessing time (856-12,237x speedup). Andrew Todd, Marziyeh Nourian, Michela Becchi |
HiPC | 2 |
| 2017 | Demystifying automata processing: GPUs, FPGAs or Micron's AP?abstractMany established and emerging applications perform at their core some form of pattern matching, a computation that maps naturally onto finite automata abstractions. As a consequence, in recent years there has been a substantial amount of work on high-speed automata processing, which has led to a number of implementations targeting a variety of parallel platforms: CPUs, GPUs, FPGAs, ASICs, and Network Processors. More recently, Micron has announced its Automata Processor (AP), a DRAM-based accelerator of non-deterministic finite automata (NFA). Despite the abundance of work in this domain, the advantages and disadvantages of different automata processing accelerators and the innovation space in this area are still unclear. Marziyeh Nourian, Xiaodong Yu 0001, Wu-chun Feng, Michela Becchi |
ICS | 1 |
| 2016 | High Performance Pattern Matching Using the Automata ProcessorabstractIn this paper, we study the acceleration of applications that require searching for all occurrences of thousands of string-patterns in an input data-stream, using the Automata Processor (AP). For this purpose, we use two applications from two fields, namely, network security and bioinformatics. The first application, called Fast-SNAP (for Fast-SNort using AP), scans network data for 4312 signatures of intrusion derived from the popular open-source Snort database. Using the resources of a single AP board, Fast-SNAP can scan for all these signatures at 10.3 Gbps. The second application, called PROTOMATA (for PROTein autOMATA), looks for all occurrences of 1308 protein motifs from the PROSITE database in protein sequences. PROTOMATA is up to half a million times faster than its single-CPU-based counterpart. The techniques developed to program these applications may be useful in the design and development of similar applications using this new hardware accelerator. Indranil Roy, Ankit Srivastava, Marziyeh Nourian, Michela Becchi, Srinivas Aluru |
IPDPS | 3 |