VLDB 2026 Research / reviewers in the wild / expert
Advait Madhavan
dblp:123/2368
· DBLP profile ↗
11ranked-venue papers
6as first author
2since 2021 · last 2024
0000-0002-4121-1336ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 6 first-author · 2 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Emerging computing paradigms · 51% Hardware accelerators and domain-specific architectures · 30% Integrated circuit design · 16% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% |
Topics — the 9 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Emerging computing paradigms
approximate computing |
0.8 | 1 | 2024 | Energy Efficient Convolutions with Temporal Arithmetic · ASPLOS (2) 2024 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator › convolution acceleration
convolution accelerator |
0.8 | 1 | 2024 | Energy Efficient Convolutions with Temporal Arithmetic · ASPLOS (2) 2024 |
Emerging computing paradigms › analog computing
race logic |
0.6 | 2 | 2019 | Boosted Race Trees for Low Energy Classification · ASPLOS 2019 Race Logic: A hardware acceleration for dynamic programming algorithms · ISCA 2014 |
Integrated circuit design
superconducting logic |
0.4 | 1 | 2020 | A Computational Temporal Logic for Superconducting Accelerators · ASPLOS 2020 |
Emerging computing paradigms
approximate and stochastic computing |
0.2 | 1 | 2016 | Energy efficient computation with asynchronous races · DAC 2016 |
Integrated circuit design
low-power circuit design |
0.2 | 1 | 2016 | Energy efficient computation with asynchronous races · DAC 2016 |
Emerging computing paradigms › approximate and stochastic computing
stochastic computing |
0.2 | 1 | 2024 | Energy Efficient Convolutions with Temporal Arithmetic · ASPLOS (2) 2024 |
Bioinformatics and computational biology
sequence alignment |
0.1 | 2 | 2016 | Energy efficient computation with asynchronous races · DAC 2016 Race Logic: A hardware acceleration for dynamic programming algorithms · ISCA 2014 |
Bioinformatics and computational biology › sequence alignment
DNA sequence alignment |
0.1 | 1 | 2014 | Race Logic: A hardware acceleration for dynamic programming algorithms · ISCA 2014 |
Methods — techniques the papers use, named apart from their topics
formal analysis · 0.9analog circuit modeling · 0.9temporal encoding · 0.8multiply-accumulate · 0.8delay encoding · 0.5current starved inverters · 0.5race logic · 0.4ensemble learning · 0.4timing-based computation · 0.4systolic array · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Energy Efficient Convolutions with Temporal ArithmeticabstractConvolution is an important operation at the heart of many applications, including image processing, object detection, and neural networks. While data movement and coordination operations continue to be important areas for optimization in general-purpose architectures, for computation fused with sensor operation, the underlying multiply-accumulate (MAC) operations dominate power consumption. Non-traditional data encoding has been shown to reduce the energy consumption of this arithmetic, with options including everything from reduced-precision floating point to fully stochastic operation, but all of these approaches start with the assumption that a complete analog-to-digital conversion (ADC) has already been done for each pixel. While analog-to-time converters have been shown to use less energy, arithmetically manipulating temporally encoded signals beyond simple min, max, and delay operations has not previously been possible, meaning operations such as convolution have been out of reach. In this paper we show that arithmetic manipulation of temporally encoded signals is possible, practical to implement, and extremely energy efficient. Rhys Gretsch, Peiyang Song 0002, Advait Madhavan, Jeremy Lau, Timothy Sherwood |
ASPLOS (2) | 3 |
| 2021 | Temporal State Machines: Using Temporal Memory to Stitch Time-based Graph Computationsabstractmappings of algorithms into hardware rely on researcher ingenuity and result in custom architectures that are difficult to systematize. We propose to associate race logic with the mathematical field of tropical algebra, enabling a more methodical approach toward building temporal circuits. This association between the mathematical primitives of tropical algebra and generalized race logic computations guides the design of temporally coded tropical circuits. It also serves as a framework for expressing high-level timing-based algorithms. This abstraction, when combined with temporal memory, allows for the systematic exploration of race logic-based temporal architectures by making it possible to partition feed-forward computations into stages and organize them into a state machine. We leverage analog memristor-based temporal memories to design such a state machine that operates purely on time-coded wavefronts. We implement a version of Dijkstra's algorithm to evaluate this temporal state machine. This demonstration shows the promise of expanding the expressibility of temporal computing to enable it to deliver significant energy and throughput advantages. Advait Madhavan, Matthew W. Daniels, Mark D. Stiles |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2020 | A Computational Temporal Logic for Superconducting AcceleratorsabstractSuperconducting logic offers the potential to perform computation at tremendous speeds and energy savings. However, a "semantic gap" lies between the level-driven logic that traditional hardware designs accept as a foundation and the pulse-driven logic that is naturally supported by the most compelling superconducting technologies. A pulse, unlike a level signal, will fire through a channel for only an instant. Arranging the network of superconducting components so that input pulses always arrive simultaneously to "logic gates'' to maintain the illusion of Boolean-only evaluation is a significant engineering hurdle. In this paper, we explore computing in a new and more native tongue for superconducting logic: time of arrival. Building on recent work in delay-based computations we show that superconducting logic can naturally compute directly over temporal relationships between pulse arrivals, that the computational relationships between those pulse arrivals can be formalized through a functional extension to a temporal predicate logic used in the verification community, and that the resulting architectures can operate asynchronously and describe real and useful computations. We verify our hypothesis through a combination of detailed analog circuit models, a formal analysis of our abstractions, and an evaluation in the context of several superconducting accelerators. Georgios Tzimpragos, Dilip P. Vasudevan, Nestan Tsiskaridze, George Michelogiannakis, Advait Madhavan, Jennifer Volk, John Shalf, Timothy Sherwood |
ASPLOS | 5 |
| 2020 | Lessons Learned the Hard Wayabstract“Fail often to succeed sooner” is a common mantra that we are told is the secret to success. When reporting research results, however, scholars rarely write about their failed attempts and only focus on the successful ones. Perhaps the source of this disconnect between what we preach and what we do can be found in the underlying assumption that published work is meant to move the field forward and failed attempts supposedly do not. The goal of the confessions presented in this paper is to show that even failed attempts are genuine and valuable contributions to our field provided that we learn from our mistakes and correct them. The 27 confessions span from planning oversights, digital and analog design errors, misunderstanding of devices, overlooked parasitics, LVS errors, and troubles in testing. Tobi Delbruck, Ibrahim M. Elfadel, Shahzad Muzaffar, Germain Haessig, Bo Wang 0012, Amine Bermak, Rui Graca, Luis A. Camuñas-Mesa, Bathiya Senevirathna, Pamela Abshire, Bernabé Linares-Barranco, Saeed Afshar, Shih-Chii Liu, Runchun Wang, Piotr Dudek, Stephen J. Carey, José M. de la Rosa 0001, Marc Dandin, Sheung Lu, Vincent Frick, Teresa Serrano-Gotarredona, Paula López Martinez 0001, Melika Payvand, Advait Madhavan, Eric R. Fossum, Juan Camilo Vasquez Tieck, Yan Liu 0016, Timothy G. Constandinou, Alexander Serb, Ricardo Carmona-Galán, Robert Nawrocki, Walter D. Leon-Salas |
ISCAS | 24 |
| 2020 | Storing and Retrieving Wavefronts with Resistive Temporal MemoryabstractWe extend the reach of temporal computing schemes by developing a memory for multi-channel temporal patterns or “wavefronts.” This temporal memory re-purposes conventional one-transistor-one-resistor (1T1R) memristor crossbars for use in an arrival-time coded, single-event-per-wire temporal computing environment. The memristor resistances and the associated circuit capacitances provide the necessary time constants, enabling the memory array to store and retrieve wavefronts. The retrieval operation of such a memory is naturally in the temporal domain and the resulting wavefronts can be used to trigger time-domain computations. While recording the wavefronts can be done using standard digital techniques, that approach has substantial translation costs between temporal and digital domains. To avoid these costs, we propose a spike timing dependent plasticity (STDP) inspired wavefront recording scheme to capture incoming wavefronts. We simulate these designs with experimentally validated memristor models and analyze the effects of memristor non-idealities on the operation of such a memory. Advait Madhavan, Mark D. Stiles |
ISCAS | 1 |
| 2019 | Boosted Race Trees for Low Energy ClassificationabstractWhen extremely low-energy processing is required, the choice of data representation makes a tremendous difference. Each representation (e.g. frequency domain, residue coded, log-scale) comes with a unique set of trade-offs --- some operations are easier in that domain while others are harder. We demonstrate that race logic, in which temporally coded signals are getting processed in a dataflow fashion, provides interesting new capabilities for in-sensor processing applications. Specifically, with an extended set of race logic operations, we show that tree-based classifiers can be naturally encoded, and that common classification tasks can be implemented efficiently as a programmable accelerator in this class of logic. To verify this hypothesis, we design several race logic implementations of ensemble learners, compare them against state-of-the-art classifiers, and conduct an architectural design space exploration. Our proof-of-concept architecture, consisting of 1,000 reconfigurable Race Trees of depth 6, will process 15.2M frames/s, dissipating 613mW in 14nm CMOS. Georgios Tzimpragos, Advait Madhavan, Dilip P. Vasudevan, Dmitri B. Strukov, Timothy Sherwood |
ASPLOS | 2 |
| 2018 | High-Throughput Pattern Matching With CMOL FPGA Circuits: Case for Logic-in-Memory ComputingabstractIn this paper, we propose a novel CMOS+ MOLecular (CMOL) field-programmable gate array (FPGA) circuit architecture to perform massively parallel, high-throughput computations, which is especially useful for pattern matching tasks and multidimensional associative searches. In the new architecture, patterns are stored as resistive states of emerging nonvolatile memory nanodevices, while the analyzed data are streamed via CMOS subsystem. The main improvements over prior work offered by the proposed circuits are increased nanodevice utilization and, as a result, substantially higher throughput, which is demonstrated by a detailed analysis of the implementation of pattern matching task on the new architecture. For example, our estimates show that the proposed CMOL FPGA circuits based on the 22-nm CMOS technology and one crossbar layer with 22-nm nanowire half-pitch allows up to 12.5% average nanodevice utilization, i.e., the fraction of the devices turned to the high conductive state, as compared to a typical ~0.1% of the original CMOL FPGA circuits. This in turn enables throughput close to 7.1 × 1016bits/s/cm2at ~ 1 fJ/bit energy efficiency, for matching of ~ 107250-bit patterns stored locally on a 1 cm2chip. These numbers are at least 2 orders of magnitude better throughput as compared to that of other state-of-the-art FPGA methods, and begin to approach ternary content-addressable memory -like performance at similar CMOS technology nodes. More generally, we argue that the proposed concept combines the versatility of reconfigurable architectures and density of the associative memories. It can be viewed as a very tight symbiotic integration of memory and logic functions for high-performance logic-in-memory computing. Advait Madhavan, Timothy Sherwood, Dmitri B. Strukov |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2016 | Energy efficient computation with asynchronous racesabstractBy encoding information as digital signal propagation delay, rather than conventional logic levels, some basic processing operations become exceedingly energy efficient to implement. The result of such a computation can then be observed by relative timing differences between injected signals. We demonstrate the embodiment of such an approach utilizing current starved inverters as delay elements and characterize application-level artifacts of circuit-level variance. Specifically we chose the well-studied DNA sequence alignment problem for comparison and we show that, for the synthesized design, asynchronous races are 10× more energy efficient and 4× denser at comparable speeds as compared to prior approaches. Advait Madhavan, Timothy Sherwood, Dmitri B. Strukov |
DAC | 1 |
| 2015 | A configurable CMOS memory platform for 3D-integrated memristorsabstractMemristors are emerging as powerful nanoscale devices for diverse applications, such as high-density memories and neuromorphic applications. However, this nascent technology requires considerable advancement before this vision is realized. We present a highly configurable CMOS interface chip which enables the characterization of on-chip memristors, especially for memory applications. The chip was fabricated in On-Semi 3M2P 0.5 μm occupying 2×2 mm2. The chip design allows for post-CMOS fabrication of memristors. The interface between the memristor and the CMOS circuitry was provided via a top metal contact. The chip was designed to support an area-distributed interface decoupling CMOS pitch and memristor pitch, enabling high-density memristor integration. Measurement results on post-CMOS fabricated Ag/SiO2/Pt memristive devices are reported. Though we have shown the results from one memristive material stack, thorough chip characterization demonstrates the versatility of the chip enabling its use with a wide variety of materials stacks. Melika Payvand, Advait Madhavan, Miguel Angel Lastras-Montaño, Amirali Ghofrani, Justin Rofeh, Kwang-Ting Cheng, Dmitri B. Strukov, Luke Theogarajan |
ISCAS | 2 |
| 2014 | Race Logic: A hardware acceleration for dynamic programming algorithmsabstractWe propose a novel computing approach, dubbed “Race Logic”, in which information, instead of being represented as logic levels, as is done in conventional logic, is represented as a timing delay. Under this new information representation, computations can be performed by observing the relative propagation times of signals injected into the circuit (i.e. the outcome of races). Race Logic is especially suited for solving problems related to the traversal of directed acyclic graphs commonly used in dynamic programming algorithms. The main advantage of this novel approach is that information processing (min-max and addition operations) can be very efficiently expressed through the manipulation of the natural delay chaining inherent to digital designs, which then results in superior latency, throughput, and energy efficiency. To verify this hypothesis, we designed several Race Logic implementations of a DNA global sequence alignment engine and compared it to the state-of-the-art conventional systolic array implementation. Our synthesized design shows that synchronous Race Logic is up to 4× faster when both approaches are mapped to a 0.5μm CMOS standard cell technology. At the same time the throughput for sequence matching per circuit area is about 3× higher at 5× lower power density for 20-long-symbol DNA sequences. Advait Madhavan, Timothy Sherwood, Dmitri B. Strukov |
ISCA | 1 |
| 2012 | Mapping of image and network processing tasks on high-throughput CMOL FPGA circuits
Advait Madhavan, Dmitri B. Strukov |
VLSI-SoC | 1 |