EDBT 2026 Demo / reviewers in the wild / expert
Aykut Dengi
dblp:74/3715
· DBLP profile ↗
4ranked-venue papers
0as first author
0since 2021 · last 2019
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Hardware accelerators and domain-specific architectures · 28% Reconfigurable computing and FPGAs · 28% Integrated circuit design · 22% | |
| Artificial intelligence
1 paper |
Deep learning architectures and training · 100% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Reconfigurable computing and FPGAs
FPGA accelerator |
0.4 | 1 | 2019 | An Energy-Efficient FPGA Implementation of an LSTM Network Using Approximate Computing · FPGA 2019 |
Reconfigurable computing and FPGAs
FPGA power reduction |
0.4 | 1 | 2019 | An Energy-Efficient FPGA Implementation of an LSTM Network Using Approximate Computing · FPGA 2019 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
LSTM accelerator |
0.4 | 1 | 2019 | An Energy-Efficient FPGA Implementation of an LSTM Network Using Approximate Computing · FPGA 2019 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.4 | 1 | 2019 | An Energy-Efficient FPGA Implementation of an LSTM Network Using Approximate Computing · FPGA 2019 |
Electronic design automation › physical design
clock network synthesis |
0.3 | 1 | 2017 | A Clock Skewing Strategy to Reduce Power and Area of ASIC Circuits · DAC 2017 |
Integrated circuit design › clocking
clock skew |
0.3 | 1 | 2017 | A Clock Skewing Strategy to Reduce Power and Area of ASIC Circuits · DAC 2017 |
Integrated circuit design
low-power circuit design |
0.3 | 1 | 2017 | A Clock Skewing Strategy to Reduce Power and Area of ASIC Circuits · DAC 2017 |
Electronic design automation
physical design |
0.3 | 1 | 2017 | A Clock Skewing Strategy to Reduce Power and Area of ASIC Circuits · DAC 2017 |
Machine learning › Deep learning architectures and training › recurrent neural network
LSTM |
0.1 | 1 | 2019 | An Energy-Efficient FPGA Implementation of an LSTM Network Using Approximate Computing · FPGA 2019 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.1 | 1 | 2019 | An Energy-Efficient FPGA Implementation of an LSTM Network Using Approximate Computing · FPGA 2019 |
Methods — techniques the papers use, named apart from their topics
quantization · 0.8approximate computing · 0.8source-target identification algorithm · 0.3differential flip-flop design · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | An Energy-Efficient FPGA Implementation of an LSTM Network Using Approximate ComputingabstractLong Short-Term Memory (LSTM) Recurrent Neural network (RNN) is known for its capability in modeling temporal aspects of data and has been shown to produce promising results in sequence learning tasks such as language modeling. However, due to the large number of model parameters and compute-intensive operations, existing FPGA implementations of LSTM cells are not sufficiently energy-efficient as they require large area and exhibit high power consumption. This work describes a substantially different hardware implementation of an LSTM which includes several architectural innovations to achieve high throughput and energy-efficiency. This paper includes extensive exploration of the design trade-offs and demonstrates the advantages for one common application - language modeling. Implementation of the design on a Xilinx Zynq XC7Z030 FPGA for language modeling shows significant improvements in throughput and energy-efficiency as compared to the state-of-the-art designs. It is worth mentioning that the proposed LSTM hardware architecture is also applicable to other applications that use LSTM as part of the neural network model (e.g., CNN-RNN models) or in whole (e.g., RNN models). Elham Azari, Aykut Dengi, Sarma B. K. Vrudhula |
FPGA | 2 |
| 2018 | FPGAs with Reconfigurable Threshold Logic Gates for Improved Performance, Power and AreaabstractThis paper proposes an alternative FPGA tile structure that consists of three traditional LUTs combined with a new reconfigurable threshold logic cell (TLC). The TLC requires only 7 SRAM cells and can be configured to implement one of several threshold functions. The proposed architecture is implemented in a 28nm FDSOI process, and is evaluated on standard benchmark circuits and several large complex function blocks. The results demonstrate an average reduction of 8.9% in register count, 15.4% in multiplexer count, 7% average reduction in Basic Logic Element (BLE) area, and 8.2% average reduction in BLE power, with a maximum decrease in register count up to 64%, BLE multiplexer count up to 68%, BLE Area up to 51.6% and BLE power up to 61.6% without loss in performance. We also show a reduction of 21% in the area of a tile. Ankit Wagle, Aykut Dengi, Sarma B. K. Vrudhula |
FPL | 3 |
| 2018 | Design Considerations for Energy-Efficient and Variation-Tolerant Nonvolatile LogicabstractSystems powered by harvested energy must consume very low power and withstand frequent interruptions in power. Nonvolatile logic (NVL) addresses the latter by saving the system state in flipflops enhanced with spin-transfer torque magnetic tunnel junctions (STT-MTJs) as the nonvolatile storage devices. Manufacturing variations in the STT-MTJs and in CMOS transistors significantly reduce yield, leading to over-design and high-energy consumption. A detailed analysis of the design tradeoffs in the driver circuitry for performing backup and restore, and a novel method to design the energy optimal driver for a given yield is presented. Next, efficient designs of two nonvolatile flip-flop (NVFF) circuits are presented, in which the backup time is determined on a per-chip basis, resulting in minimizing the energy wastage and satisfying the yield constraint. To achieve a yield of 98%, the conventional approach would have to expend nearly 5× more energy than the minimum required, whereas the proposed tunable approach expends only 26% more energy than the minimum. Also included are the energy consumption of the proposed NVFF designs when used in two larger function blocks. Experimental results were based on a commercial 40-nm process design kit, and HSPICE simulations with foundry supplied statistical models and data. Aykut Dengi, Sarma B. K. Vrudhula |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | A Clock Skewing Strategy to Reduce Power and Area of ASIC CircuitsabstractA new method for reducing power and area of standard cell ASICs is described. The method is based on deliberately introducing clock skew without the use of extra buffers in the clock network. This is done by having some flipflops, called sources, generate clock signals for other flipflops, called targets. The method involves two key features: (1) the design of new differential flipflop, referred to as KVFF, that is functionally identical to a master-slave edge-triggered D flipflop, but in addition, produces an completion signal that is a skewed version of its input clock, which is used to clock other flipflops; and (2) an efficient algorithm that identifies the sources and targets involved in the new clocking scheme, with the objective of reducing area and power. These are reduced because deliberate skew introduces extra slack on the logic cones that feed the target flipflops, which is exploited by synthesis tools to reduce area and power. In addition, the overhead of conventional methods of introducing skew, e.g. buffers, is eliminated. Using commercial tools, significant improvements in power and area are shown on placed and routed netlists of several circuits. Niranjan Kulkarni, Aykut Dengi, Sarma B. K. Vrudhula |
DAC | 2 |