EDBT 2026 Demo / reviewers in the wild / expert
Nariman Eskandari
dblp:163/7155
· DBLP profile ↗
3ranked-venue papers
2as first author
0since 2021 · last 2019
0000-0003-3038-5186ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorTheory of computation · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Cloud and datacenter computing · 44% Reconfigurable computing and FPGAs · 44% High-performance computing · 13% |
Topics — the 1 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing › cluster computing
cluster deployment |
0.1 | 1 | 2019 | A Modular Heterogeneous Stack for Deploying FPGAs and CPUs in the Data Center · FPGA 2019 |
Methods — techniques the papers use, named apart from their topics
high-level synthesis · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | A Modular Heterogeneous Stack for Deploying FPGAs and CPUs in the Data CenterabstractIn this work we present a heterogeneous deployment stack, calledGalapagos, that includes the abstraction of individual nodes (FPGAsand CPUs), the communication protocols between nodes and theorchestration and connection of these nodes into clusters. The stackwe create is also highly modular, allowing users to explore a designspace in the implementation of their cluster such as different net-work protocols or communication layers. The communication layerwe have currently implemented within our hardware stack, calledHUMboldt, handles heterogeneous communication between multi-ple FPGAs and CPUs. We implementHUMboldtusing High-LevelSynthesis (HLS) to ensure functional portability of communicatingkernels, allowing us to prototype hardware kernels in software. Ourresults have shown that our modular approach to this heterogeneousdeployment stack has introduced very little area and latency over-head in the FPGAs and can still perform at line-rate, bottleneckedsolely by the network links connecting the nodes. Our results alsohighlight the scalability of our design as our performance remainslimited by the network links when the cluster size increases. Nariman Eskandari, Naif Tarafdar, Daniel Ly-Ma, Paul Chow |
FPGA | 1 |
| 2017 | Heterogeneous virtualized network function framework for the data centerabstractWe present a framework for creating heterogeneous virtualized network function (VNF) service chains from cloud data center resources. Traditionally, these functions are packaged in software images within a catalog of networking applications that can be loaded onto a virtual machine CPU, and can be offered to users as a service. Our framework combines the best of both software and hardware by allowing users to chain traditional software-based VNFs with hardware-based VNFs that the user provides as an IP to generate a bitstream or a pre-generated VNF as part of a library. To accomplish this, our framework first creates the hardware bitstreams and programs the FPGA VNFs, loads any software VNFs requested, and programs the network to daisy chain the VNFs together. Furthermore, this enables an incremental design flow where the user can start by implementing a chain of VNFs in software and incrementally substitute software VNFs for their hardware counterparts. Our paper investigates two case studies to show the ability to switch between hardware and software VNFs in our framework and to demonstrate the benefit of using hardware VNFs. The first study is signature matching at fixed offsets, similar to matching packet headers. In this case study, the CPU can keep up at line-rate using specialized networking drivers. The second case study involves string matching within a packet, which requires scanning through the entire frame. In this case, the CPU performance drops to approximately 20 percent of the input rate, whereas the FPGA can continue to keep up at line-rate. Naif Tarafdar, Thomas Lin, Nariman Eskandari, David Lion, Alberto Leon-Garcia, Paul Chow |
FPL | 3 |
| 2014 | A fast emulator for ARM-based embedded systemsabstractThis paper presents a high-performance implementation for an Intel 8080 emulator on a Raspberry Pi device. The problem was defined as a software contest in MEMOCODE 2014 and this implementation took the second place in this contest. We deployed several optimization techniques and employed best programming practices to increase the performance of the naïve reference implementation. Improving data structure usage and modifying function calls are the techniques that resulted in higher performance of this implementation. Our implementation has about 2.5 times speedup over the reference code of the contest. Nariman Eskandari, Hatef Madani, Armin Ahmadzadeh, Mohsen Mahmoudi Aznaveh, Saeid Gorgin 0001 |
MEMOCODE | 1 |