EDBT 2026 Demo / reviewers in the wild / expert
Amir Momeni
dblp:22/8791
· DBLP profile ↗
5ranked-venue papers
3as first author
0since 2021 · last 2016
0000-0002-6941-8179ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 first-authorSoftware engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Emerging computing paradigms · 36% GPUs and heterogeneous computing · 23% Integrated circuit design · 18% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% |
Topics — the 6 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Emerging computing paradigms › approximate computing › approximate circuit design
approximate arithmetic circuits |
0.2 | 1 | 2015 | Design and Analysis of Approximate Compressors for Multiplication · IEEE Trans. Computers 2015 |
Emerging computing paradigms
approximate computing |
0.2 | 1 | 2015 | Design and Analysis of Approximate Compressors for Multiplication · IEEE Trans. Computers 2015 |
Processor architecture and microarchitecture
computer arithmetic |
0.2 | 1 | 2015 | Design and Analysis of Approximate Compressors for Multiplication · IEEE Trans. Computers 2015 |
Integrated circuit design
digital circuit design |
0.2 | 1 | 2015 | Design and Analysis of Approximate Compressors for Multiplication · IEEE Trans. Computers 2015 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.1 | 1 | 2015 | Bridging Architecture and Programming for Throughput-Oriented Vision Processing (Abstract Only) · FPGA 2015 |
GPUs and heterogeneous computing
vision processing |
0.1 | 1 | 2015 | Bridging Architecture and Programming for Throughput-Oriented Vision Processing (Abstract Only) · FPGA 2015 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.4parallelism granularity classification · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | Hardware thread reordering to boost OpenCL throughput on FPGAsabstractAvailability of OpenCL for FPGAs has raised new questions about the efficiency of massive thread-level parallelism on FPGAs. The general trend is toward creating deep pipelining and in-order execution of many OpenCL threads across a shared data-path. While this can be a very effective approach for regular kernels, its efficiency significantly diminishes for irregular kernels with runtime-dependent control flow. We need to look for new approaches to improve execution efficiency of FPGAs when targeting irregular OpenCL kernels. This paper proposes a novel solution, called Hardware Thread Reordering (HTR), to boost the throughput of the FPGAs when executing irregular kernels possessing non-deterministic runtime control flow. The key insight of HRT is out-of-order OpenCL thread execution over a shared data-path to achieve significantly higher throughput. The thread reordering is performed at a basic-block level granularity. The synthesized basic-blocks are extended with independent pipeline control signals and context registers to bypass the live values of reordered threads. We demonstrate the efficiency of our proposed solution on three parallel irregular kernels. For the experiments, we utilize the LegUp tool to compare the baseline (in-order) data-path with HTR-enhanced data-path. Our RTL simulation results demonstrate that HTR-enhanced data-path achieves up to 11× increase in kernels throughput at a very low overhead (less than 2× increase in FPGA resources). Amir Momeni, Hamed Tabkhi, Gunar Schirner, David R. Kaeli |
ICCD | 1 |
| 2015 | Bridging Architecture and Programming for Throughput-Oriented Vision Processing (Abstract Only)abstractWith the expansion of OpenCL support across many heterogeneous devices (including FPGAs, GPUs and CPUs), the programmability of these systems has been significantly increased. At the same time, new questions arise about which device should be targeted for each OpenCL software kernel. Once we select a device, then we are left to customize the application, selecting the right granularity of parallelism and frequency of host-to-device communication. In this paper, we study the impact of source-level decisions on the overall execution time when developing OpenCL program across different heterogeneous devices. We focus on two mainstream architecture classes (GPUs and FPGAs), and consider throughput-oriented advanced vision processing. To guide this exploration, we propose a new vertical classification for selecting the grain of parallelism for advanced vision processing applications. To carry out this study we have selected the Mean-shift object tracking algorithm as a representative candidate of advanced vision algorithms. Overall, our evaluation demonstrates that fine-grained parallelism can greatly benefit FPGA execution (up to a 4X speed-up), while a combination of coarse-grained and fine-grained parallelism achieves the best performance on a GPU (up to a 6X speed-up). Also, there can be a large benefit if we can execute both the parallel and serial parts of the program on a FPGA (up to a 21X speed-up). Amir Momeni, Hamed Tabkhi, Gunar Schirner, David R. Kaeli |
FPGA | 1 |
| 2015 | NUPAR: A Benchmark Suite for Modern GPU ArchitecturesabstractHeterogeneous systems consisting of multi-core CPUs, Graphics Processing Units (GPUs) and many-core accelerators have gained widespread use by application developers and data-center platform developers. Modern day heterogeneous systems have evolved to include advanced hardware and software features to support a spectrum of application patterns. Heterogeneous programming frameworks such as CUDA, OpenCL, and OpenACC have all introduced new interfaces to enable developers to utilize new features on these platforms. In emerging applications, performance optimization is not only limited to effectively exploiting data-level parallelism, but includes leveraging new degrees of concurrency and parallelism to accelerate the entire application. Yash Ukidave, Fanny Nina Paravecino, Leiming Yu, Charu Kalra, Amir Momeni, Zhongliang Chen, Nick Materise, Brett Daley, Perhaad Mistry, David R. Kaeli |
ICPE | 5 |
| 2015 | Design and Analysis of Approximate Compressors for MultiplicationabstractInexact (or approximate) computing is an attractive paradigm for digital processing at nanometric scales. Inexact computing is particularly interesting for computer arithmetic designs. This paper deals with the analysis and design of two new approximate 4-2 compressors for utilization in a multiplier. These designs rely on different features of compression, such that imprecision in computation (as measured by the error rate and the so-called normalized error distance) can meet with respect to circuit-based figures of merit of a design (number of transistors, delay and power consumption). Four different schemes for utilizing the proposed approximate compressors are proposed and analyzed for a Dadda multiplier. Extensive simulation results are provided and an application of the approximate multipliers to image processing is presented. The results show that the proposed designs accomplish significant reductions in power dissipation, delay and transistor count compared to an exact design; moreover, two of the proposed multiplier designs provide excellent capabilities for image multiplication with respect to average normalized error distance and peak signal-to-noise ratio (more than 50 dB for the considered image examples). Amir Momeni, Jie Han 0001, Paolo Montuschi, Fabrizio Lombardi |
IEEE Trans. Computers | 1 |
| 2011 | Hardware design of a new genetic based disk scheduling method
Hossein Rahmani 0001, Mohammad Reza Bonyadi, Amir Momeni, Mohsen Ebrahimi Moghaddam, Maghsoud Abbaspour |
Real Time Syst. | 3 |