Udari De Alwis

dblp:295/7114 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
0000-0002-0824-7724ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Smart Imager with Object Detection Exploiting Edge-Frame-Base Processing and Bounding Box Extraction for μW Power Purely-Harvested Sensor Nodes
abstract
Battery-less and cost-sensitive vision nodes are becoming essential in IoT-scale sensor networks, where in/near-sensor AI enables local recognition while minimizing data transmission. However, achieving multi-class object detection under available peak power budgets (<10 µW) and low-cost fabrication remains a major challenge. Existing smart imagers either lack on-chip intelligence or exceed such power budgets due to costly sensing and computing. This paper presents a fully-integrated smart imager performing multi-class object detection at 8.51 μW (equivalent to the power from a 7mm × 6mm harvester at 300 lux) in standard 180nm CMOS. The system processes 1-bit edge-extracted frames, applies tile-level novelty detection for bounding-box ROI extraction, and computes CENTRIST features over cropped regions. A low-power approximate linear SVM classifies detected objects at 130 pW/pixel power. Unlike prior architectures, the proposed system maintains full image readout, supports flexible learning-based inference, and avoids custom optics and CIS processing. This makes it the first battery-less smart imager capable of flexible, multi-object detection in low-cost standard CMOS technology.
Hayate Okuhara, Udari De Alwis, Liu Yue, Karim Ali 0007, Massimo Alioto
DATE2
2026 Characterizing Machine Learning Force Fields as Emerging Molecular Dynamics Workloads on Graphics Processing Units
abstract
Molecular dynamics (MD) simulates the time evolution of atomic systems governed by interatomic forces, and the fidelity of these simulations depends critically on the underlying force model. Classical force fields (CFFs) rely on fixed functional forms fitted to experimental or theoretical data, offering computational efficiency and broad applicability but limited accuracy in chemically diverse or reactive environments. In contrast, machine learning force fields (MLFFs) deliver near-quantum-chemical accuracy at molecular-mechanics cost by learning interatomic interactions directly from high-level electronic-structure data.While MLFFs offer improved accuracy at a fraction of the cost of quantum methods, they introduce significant computational overhead, particularly in descriptor evaluation and neural network inference. These operations pose challenges for parallel hardware due to irregular memory access, minimum data reuse and inefficient kernel execution.This work investigates the hardware performance of such models using poly-alanine chains, a novel benchmark molecule system(s) with controllable input size, which used as performance evaluation test cases highlighting the computational bottlenecks of the graphical processor units when scaling out MLFF simulations. The analysis identifies key bottlenecks in descriptor and force computation, memory handling, highlighting the opportunities for improvements in the emerging area of MLFF based MD in drug discovery, that has received limited attention from a computer architecture perspective.
Udari De Alwis, Benjamin E. Mayer, Tom J. Ashby, Maria Barrera, Timon Evenblij, Joyjit Kundu
ISPASS1