EDBT 2026 Demo / reviewers in the wild / expert
Dilip P. Vasudevan
dblp:20/3373 · also Dilip Vasudevan
· DBLP profile ↗
11ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0003-0931-309XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Emerging computing paradigms · 47% Integrated circuit design · 25% Hardware accelerators and domain-specific architectures · 22% | |
| Theoretical computer science
1 paper |
Logic in computer science · 100% |
Topics — the 2 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Integrated circuit design
superconducting logic |
0.4 | 1 | 2020 | A Computational Temporal Logic for Superconducting Accelerators · ASPLOS 2020 |
Emerging computing paradigms › analog computing
race logic |
0.4 | 1 | 2019 | Boosted Race Trees for Low Energy Classification · ASPLOS 2019 |
Methods — techniques the papers use, named apart from their topics
formal analysis · 0.9analog circuit modeling · 0.9race logic · 0.4ensemble learning · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | InFormer: A High-throughput, Ultra-efficient In-memory Compute-based Floating-point Arithmetic Accelerator for Transformers
Hasita Veluri, Dilip P. Vasudevan |
ACM Great Lakes Symposium on VLSI | 2 |
| 2023 | Benefits of Optimistic Parallel Discrete Event Simulation for Network-on-Chip SimulationabstractThe end of Moore's law has placed a two-fold demand on hardware simulation. Firstly, efficient co-design requires fast simulation of hardware systems in order to vet proposed designs. Secondly, modern simulator platforms need to become increasingly concurrent as well. To address these challenges, we develop an optimistic time warp-based parallelization for the Structural Simulation Toolkit (SST). Our optimistic PDES engine hides synchronization costs by speculatively executing tasks, leading to better compute resource utilization on modern multicore architectures. Given the significant engineering effort to make custom existing hardware models reversible, we also develop a new SST component, called escher. The escher workflow instruments arbitrary SST applications and replays event traces through the different SST parallelizations. This is a useful tool for hardware engineers who want to understand the benefits of optimistic parallelization before undertaking the significant software engineering effort to make their hardware models reversible. We demonstrate this workflow by generating traces for a tiled mesh-NOC architecture and show a 2.1x to 3.7x speed-up using our optimistic SST parallelization versus the current conservative SST parallelization. Maximilian H. Bremer, Nirmalendu Bikash Patra, Tan Nguyen 0001, Dilip P. Vasudevan, Cy P. Chan |
DS-RT | 4 |
| 2021 | SRNoC: A Statically-Scheduled Circuit-Switched Superconducting Race Logic NoCabstractTemporal encoding has been shown to be a natural fit for single flux quantum (SFQ) superconducting computing since SFQ already encodes information with the presence or absence of voltage pulses. However, past work in SFQ has focused on binary-encoded networks on chip (NoCs). In this paper, we propose superconducting rotary NoC (SRNoC), a NoC where both data and control paths operate in the temporal domain following the race logic (RL) convention. Therefore, SFQ chips with temporal compute or memory can use SRNoC to avoid converting between the temporal and binary domains that would result from using a binary-encoded NoC. Using RL also enables SRNoC to be area-efficient, mitigating SFQ technology's low device density. SRNoC treats pulses as independent packets and delivers them to outputs without changing their value, i.e. preserving the RL convention. SRNoC operates on a fixed, rotating connection schedule between inputs and outputs. In each connection window, multiple pulses (packets) can be transmitted sequentially. SRNoC provides 13.1x higher throughput per port per Josephson junction (JJ) compared to the best-performing of three demonstrated NoCs. George Michelogiannakis, Darren Lyles, Patricia Gonzalez-Guerrero, Meriam Gay Bautista, Dilip P. Vasudevan, Anastasiia Butko |
IPDPS | 5 |
| 2020 | A Computational Temporal Logic for Superconducting AcceleratorsabstractSuperconducting logic offers the potential to perform computation at tremendous speeds and energy savings. However, a "semantic gap" lies between the level-driven logic that traditional hardware designs accept as a foundation and the pulse-driven logic that is naturally supported by the most compelling superconducting technologies. A pulse, unlike a level signal, will fire through a channel for only an instant. Arranging the network of superconducting components so that input pulses always arrive simultaneously to "logic gates'' to maintain the illusion of Boolean-only evaluation is a significant engineering hurdle. In this paper, we explore computing in a new and more native tongue for superconducting logic: time of arrival. Building on recent work in delay-based computations we show that superconducting logic can naturally compute directly over temporal relationships between pulse arrivals, that the computational relationships between those pulse arrivals can be formalized through a functional extension to a temporal predicate logic used in the verification community, and that the resulting architectures can operate asynchronously and describe real and useful computations. We verify our hypothesis through a combination of detailed analog circuit models, a formal analysis of our abstractions, and an evaluation in the context of several superconducting accelerators. Georgios Tzimpragos, Dilip P. Vasudevan, Nestan Tsiskaridze, George Michelogiannakis, Advait Madhavan, Jennifer Volk, John Shalf, Timothy Sherwood |
ASPLOS | 2 |
| 2020 | Language Support for Navigating Architecture Design in Closed FormabstractAs computer architecture continues to expand beyond software-agnostic microarchitecture to specialized and heterogeneous logic or even radically different emerging computing models (e.g., quantum cores, DNA storage units), detailed cycle-level simulation is no longer presupposed. Exploring designs under such complex interacting relationships (e.g., performance, energy, thermal, frequency) calls for a more integrative but higher-level approach. We propose Charm, a modeling language supporting closed-form high-level architecture modeling. Charm enables mathematical representations of mutually dependent architectural relationships to be specified, composed, checked, evaluated, reused, and shared. The language is interpreted through a combination of automatic symbolic evaluation, scalable graph transformation, and efficient compiler techniques, generating executable DAGs and optimized analysis procedures. Charm also exploits the advancements in satisfiability modulo theory solvers to automatically search the design space to help architects explore multiple design knobs simultaneously (e.g., different CNN tiling configurations). Through two case studies, we demonstrate that Charm allows one to define high-level architecture models in a clean and concise format, maximize reusability and shareability, capture unreasonable assumptions, and significantly ease design space exploration at a high level. Weilong Cui, Georgios Tzimpragos, Bill Tao, Joseph McMahan, Deeksha Dangwal, Nestan Tsiskaridze, George Michelogiannakis, Dilip P. Vasudevan, Timothy Sherwood |
ACM J. Emerg. Technol. Comput. Syst. | 8 |
| 2019 | Boosted Race Trees for Low Energy ClassificationabstractWhen extremely low-energy processing is required, the choice of data representation makes a tremendous difference. Each representation (e.g. frequency domain, residue coded, log-scale) comes with a unique set of trade-offs --- some operations are easier in that domain while others are harder. We demonstrate that race logic, in which temporally coded signals are getting processed in a dataflow fashion, provides interesting new capabilities for in-sensor processing applications. Specifically, with an extended set of race logic operations, we show that tree-based classifiers can be naturally encoded, and that common classification tasks can be implemented efficiently as a programmable accelerator in this class of logic. To verify this hypothesis, we design several race logic implementations of ensemble learners, compare them against state-of-the-art classifiers, and conduct an architectural design space exploration. Our proof-of-concept architecture, consisting of 1,000 reconfigurable Race Trees of depth 6, will process 15.2M frames/s, dissipating 613mW in 14nm CMOS. Georgios Tzimpragos, Advait Madhavan, Dilip P. Vasudevan, Dmitri B. Strukov, Timothy Sherwood |
ASPLOS | 3 |
| 2019 | PARADISE - Post-Moore Architecture and Accelerator Design Space Exploration Using Device Level Simulation and ExperimentsabstractAn increasing number of technologies are being proposed to preserve digital computing performance scaling as lithographic scaling slows. These technologies include new devices, specialized architectures, memories, and 3D integration. Currently, no end-to-end tool flow is available to rapidly perform architectural-level evaluation using device-level models and for a variety of emerging technologies at once. We propose PARADISE: An open-source comprehensive methodology to evaluate emerging technologies with a vertical simulation flow from the individual device level all the way up to the architec-turallevel. To demonstrate its effectiveness, we use PARADISE to perform end-to-end simulation and analysis of heterogeneous architectures using CNFETs, TFETs, and NCFETs, along with multiple hardware designs. To demonstrate its accuracy, we show that PARADISE has only a 6% mean deviation for delay and 9% for power compared to previous studies using commercial synthesis tools. Dilip P. Vasudevan, George Michelogiannakis, David Donofrio, John Shalf |
ISPASS | 1 |
| 2015 | The Bit-Nibble-Byte MicroEngine (BnB) for Efficient Computing on Short DataabstractEnergy is a critical challenge in computing performance. Due to "word size creep" from modern CPUs are inefficient for short-data element processing. We propose and evaluate a new microarchitecture called "Bit-Nibble-Byte"(BnB). We describe our design which includes both long fixed point vectors and as well as novel variable length instructions. Together, these features provide energy and performance benefits on a wide range of applications. We evaluate BnB with a detailed design of 5 vector sizes (128,256,512,1024,2048) mapped into 32nm and 7nm transistor technologies, and in combination with a variety of memory systems (DDR3 and HMC). The evaluation is based on both handwritten and compiled code with a custom compiler built for BnB. Our results include significant performance (19x-252x) and energy benefits (5.6x-140.7x) for short bit-field operations typically assumed to require hardwired accelerators and large-scale applications with compiled code. Dilip P. Vasudevan, Andrew A. Chien |
ACM Great Lakes Symposium on VLSI | 1 |
| 2011 | Design of a Low Power, Sub-Threshold, Asynchronous Arithmetic Logic Unit Using a Bidirectional AdderabstractA novel asynchronous bidirectional arithmetic Logic Unit (ALU) is introduced in this paper. The adder in the proposed design is a ripple carry adder with the bidirectional characteristic. The ALU is designed with asynchronous dual rail circuit style. Several ALUs with sizes ranging from 4bits to 32 bits were built. Their power and performance metrics were compared with the conventional ALUs built with the fast adders designed with dynamic logic style. Significant power reduction with the sub-threshold operating voltage is achieved. Also the design is compared with the ALU design proposed for reversible quantum computers in the CMOS context to show the logic efficiency of the proposed design around 30 % in area. Power reduction of 9-26% was achieved for the addition operation and and 19.5 - 75.1% for the logical operation on the proposed 32 bit ALU, compared to the conventional dynamic logic based ALU operated over the voltage range 0.2-0.3V. Jiaoyan Chen 0002, Dilip P. Vasudevan, Emanuel M. Popovici, Michel P. Schellekens |
DSD | 2 |
| 2010 | Static Average Case Power Estimation Technique for Block CiphersabstractIn this paper a new static average case dynamic power estimation technique is introduced based on the property of randomness preservation for digital circuits. The proposed technique is validated by estimating the average case power for a block cipher, DES with a lower estimation error percentage of 0.9481 % and lesser simulation time with a pattern reduction of (2" × 2"!)-(2" × 2" × 2) for n bit design. The same technique can be extended to any block cipher, including the AES and IDEA-NXT. Tingcong Ye, Dilip P. Vasudevan, Jiaoyan Chen 0002, Emanuel M. Popovici, Michel P. Schellekens |
DSD | 2 |
| 2004 | A Novel Approach for On-line Testable Reversible Logic Circuit DesigabstractTwo testable reversible logic gates are proposed in this paper. These gates can be used to implement reversible digital circuits with various levels of complexity. The major feature of these gates is that they provide online-testability for circuits implemented using these gates. The application of these gates in testable ripple carry, carry-skip adders and MCNC benchmark circuits have been illustrated. Dilip P. Vasudevan, Parag K. Lala, James Patrick Parkerson |
Asian Test Symposium | 1 |