EDBT 2026 Demo / reviewers in the wild / expert
Ashwin Sanjay Lele
dblp:261/2888
· DBLP profile ↗
10ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0002-2440-905XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SATA: Sparsity-Aware Scheduling for Selective Token AttentionabstractTransformers have become the foundation of numerous state-of-the-art AI models across diverse domains, thanks to their powerful attention mechanism for modeling long-range dependencies. However, the quadratic scaling complexity of attention poses significant challenges for efficient hardware implementation. While techniques such as quantization and pruning help mitigate this issue, selective token attention offers a promising alternative by narrowing the attention scope to only the most relevant tokens, reducing computation and filtering out noise.In this work, we propose SATA, a locality-centric dynamic scheduling scheme that proactively manages sparsely distributed access patterns from selective Query-Key operations. By reordering operand flow and exploiting data locality, our approach enables early fetch and retirement of intermediate Query/Key vectors, improving system utilization. We implement and evaluate our token management strategy in a control and compute system, using runtime traces from selective-attention-based models. Experimental results show that our method improves system throughput by up to 1.76× and boosts energy efficiency by 2.94×, while incurring minimal scheduling overhead. Zhenkun Fan, Zishen Wan, Che-Kai Liu, Ashwin Sanjay Lele, Win-San Khwa, Meng-Fan Chang, Arijit Raychowdhury |
DATE | 4 |
| 2025 | Finding the Pareto Frontier of Low-Precision Data Formats and MAC Architecture for LLM InferenceabstractTo accelerate AI applications, numerous data formats and physical implementations of matrix multiplication have been proposed, creating a complex design space. This paper studies the efficient MAC implementation of the integer, floating-point, posit, and logarithmic number system (LNS) data formats and Microscaling (MX) and VectorScaled Quantization (VSQ) block data formats. We evaluate the area, power, and numerical accuracy (evaluated as signal-to-quantization noise ratio) of $\mathbf{3 5, 0 0 0}$ MAC designs spanning each data format and several key design parameters such as the inner product size and accumulation width. We find that for the same numerical accuracy, pareto optimal MAC designs with emerging data formats (LNS16, MXINT8, VSQINT4) achieve $1.8 \times 2.2 \times$, and $1.9 \times$ TOPs/W improvement compared to FP16, FP8, and FP4 dot product implementations. Brian Crafton, Xiaochen Peng, Xiaoyu Sun 0001, Ashwin Sanjay Lele, Win-San Khwa, Kerem Akarvardar |
DAC | 4 |
| 2023 | Neuromorphic Swarm on RRAM Compute-in-Memory Processor for Solving QUBO ProblemabstractCombinatorial optimization problems prevail in engineering and industry. Some are NP-hard and thus become difficult to solve on edge devices due to limited power and computing resources. Quadratic Unconstrained Binary Optimization (QUBO) problem is a valuable emerging model that can formulate numerous combinatorial problems, such as Max-Cut, traveling salesman problems, and graphic coloring. QUBO model also reconciles with two emerging computation models, quantum computing and neuromorphic computing, which can potentially boost the speed and energy efficiency in solving combinatorial problems. In this work, we design a neuromorphic QUBO solver composed of a swarm of spiking neural networks (SNN) that conduct a population-based meta-heuristic search for solutions. The proposed model can achieve about x20 40 speedup on large QUBO problems in terms of time steps compared to a traditional neural network solver. As a codesign, we evaluate the neuromorphic swarm solver on a 40nm 25mW Resistive RAM (RRAM) Compute-in-Memory (CIM) SoC with a 2.25MB RRAM-based accelerator and an embedded Cortex M3 core. The collaborative SNN swarm can fully exploit the specialty of CIM accelerator in matrix and vector multiplications. Compared to previous works, such an algorithm-hardware synergized solver exhibits advantageous speed and energy efficiency for edge devices. Ashwin Sanjay Lele, Muya Chang, Samuel Spetalnick, Brian Crafton, Arijit Raychowdhury, Yan Fang 0002 |
DAC | 1 |
| 2023 | Live Demonstration: Hybrid RRAM and SRAM SoC for Fused Frame and Event Target TrackingabstractEvent and frame cameras capture the complemen-tary spatial and temporal details of a scene providing an accuracy vs. latency trade-off. Fusing these processing modalities using convolutional (CNN) and spiking neural networks (SNN) respectively has been shown for target tracking. We present our heterogeneous RRAM compute-in-memory (CIM) and SRAM compute-near-memory (CNM) SoC for simultaneous processing of CNN and SNN. We will show the advantage of using fused vision over frame-only vision and demonstrate python programmable data streaming. The visitors will be able to see the processing-dependent dynamic power gating of non-volatile RRAM and in-memory error correction capability. Ashwin Sanjay Lele, Muya Chang, Samuel Spetalnick, Yan Fang 0002, Brian Crafton, Shota Konno, Arijit Raychowdhury |
ISCAS | 1 |
| 2022 | Circuit and System Technologies for Energy-Efficient Edge Robotics: (Invited Paper)abstractAs we march towards the age of ubiquitous intelligence, we note that AI and intelligence are progressively moving from the cloud to the edge. The success of Edge-AI is pivoted on innovative circuits and hardware that can enable inference and limited learning in resource-constrained edge autonomous systems. This paper introduces a series of ultra-low-power accelerator and system designs on enabling the intelligence in edge robotic platforms, including reinforcement learning neuro-morphic control, swarm intelligence, and simultaneous mapping and localization. We put an emphasis on the impact of the mixed-signal circuit, neuro-inspired computing system, benchmarking and software infrastructure, as well as algorithm-hardware co-design to realize the most energy-efficient Edge-AI ASICs for the next-generation intelligent and autonomous systems. Zishen Wan, Ashwin Sanjay Lele, Arijit Raychowdhury |
ASP-DAC | 2 |
| 2022 | Fusing Frame and Event Vision for High-speed Optical Flow for Edge ApplicationabstractOptical flow computation with frame-based cameras provides high accuracy but the speed is limited either by the model size of the algorithm or by the frame rate of the camera. This makes it inadequate for high-speed applications. Event cameras provide continuous asynchronous event streams overcoming the frame-rate limitation. However, the algorithms for processing the data either borrow frame like setup limiting the speed or suffer from lower accuracy. We fuse the complementary accuracy and speed advantages of the frame and event-based pipelines to provide high-speed optical flow while maintaining a low error rate. Our bio-mimetic network is validated with the MVSEC dataset showing 19% error degradation at 4$\times$ speed up. We then demonstrate the system with a high-speed drone flight scenario where a high-speed event camera computes the flow even before the optical camera sees the drone making it suited for applications like tracking and segmentation. This work shows the fundamental trade-offs in frame-based processing may be overcome by fusing data from other modalities. Ashwin Sanjay Lele, Arijit Raychowdhury |
ISCAS | 1 |
| 2020 | Bio-inspired Gait Imitation of Hexapod Robot Using Event-Based Vision Sensor and Spiking Neural NetworkabstractLearning how to walk is a sophisticated neurological task for most animals. In order to walk, the brain must synthesize multiple cortices, neural circuits, and diverse sensory inputs. Some animals, like humans, imitate surrounding individuals to speed up their learning. When humans watch their peers, visual data is processed through a visual cortex in the brain. This complex problem of imitation-based learning forms associations between visual data and muscle actuation through Central Pattern Generation (CPG). Reproducing this imitation phenomenon on low power, energy-constrained robots that are learning to walk remains challenging and unexplored. We propose a bio-inspired feed-forward approach based on neuromorphic computing and event-based vision to address the gait imitation problem. The proposed method trains a "student" hexapod to walk by watching an "expert" hexapod moving its legs. The student processes the flow of Dynamic Vision Sensor (DVS) data with a one-layer Spiking Neural Network (SNN). The SNN of the student successfully imitates the expert within a small convergence time of ten iterations and exhibits energy efficiency at the sub-microjoule level. Justin Ting, Yan Fang 0002, Ashwin Sanjay Lele, Arijit Raychowdhury |
IJCNN | 3 |
| 2020 | Circuit Cost Reduction for Online STDP using NIPIN Selector as Timekeeping Device in RRAM SynapseabstractOn-chip implementation of spike-time dependent plasticity in spiking neural networks using RRAM synapses requires pulse shaping circuits (PSC) to drive RRAMs. PSCs convert the temporal separation between pre and post neuron spikes to appropriate voltages that get applied across the synapse. The speculation of PSCs consuming the majority of circuit resources in the neuron circuits calls for methods simplifying the PSC. A recently demonstrated NIPIN timekeeping device based selector facilitates this, showing learning with square pulses using its inherent hole storage physics. However, a quantitative advantage achieved by utilizing a timekeeping device to evaluate its necessity is unavailable in the literature. Also, a model is required to carry out large scale circuit simulations for crossbar arrays using this device as selector. In this work, we design and compare the PSCs for different selector devices proposed in the literature to show 133× reduction in energy per spike and 8× reduction in the area of neuron circuit using NIPIN as the selector device compared to previously shown diode selector. We also present an experimentally calibrated model for the device for future explorations. Our results show that the small fraction energy and area occupied by the leaky-integrate and fire part of the circuit makes optimization of PSCs a priority. Thus, our work highlights the importance of mimicking biology by the use of simple spikes from neurons and performing time-keeping at the synapse in implementations of learning circuits. Ashwin Sanjay Lele, Anand Naik, Lakshya Bandhu, Bhaskar Das, Udayan Ganguly |
ISCAS | 1 |
| 2020 | Online Reward-Based Training of Spiking Central Pattern Generator for Hexapod LocomotionabstractOnline learning in legged robot under stringent performance and energy constraints thwarts the application of conventional reinforcement learning and optimization algorithms. The integration of complex sensors and data pre-processing required in using these algorithms makes this more challenging. Spiking neural networks allow local learning and low computing power opening new possibilities neuromorphic paradigm to such tasks. Central pattern generation based learning to walk in hexapod robots perfectly matches the temporal learning in SNNs allowing end-to-end learning. We propose a stochastic reinforcement-based algorithm allowing the hexapod to learn using the reward generated by the gyro sensors and camera-based visual inputs. The system is implemented on a Raspberry pi to demonstrate convergence to bio-observed gait patterns. Ashwin Sanjay Lele, Yan Fang 0002, Justin Ting, Arijit Raychowdhury |
VLSI-SOC | 1 |
| 2013 | Relational algorithms for multi-bulk-synchronous processorsabstractRelational databases remain an important application infrastructure for organizing and analyzing massive volumes of data. At the same time, processor architectures are increasingly gravitating towards Multi-Bulk-Synchronous processor (Multi-BSP) architectures employing throughput-optimized memory systems, lightweight multi-threading, and Single-Instruction Multiple-Data (SIMD) core organizations. This paper explores the mapping of primitive relational algebra operations onto such architectures to improve the throughput of data warehousing applications built on relational databases. Gregory Frederick Diamos, Haicheng Wu, Jin Wang 0010, Ashwin Sanjay Lele, Sudhakar Yalamanchili |
PPoPP | 4 |