Uday Kumar Reddy Vengalam

dblp:223/7813 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
5since 2021 · last 2023
0000-0001-6748-0625ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2023 Supporting Energy-based Learning with an Ising Machine substrate: a Case Study on RBM
abstract
Nature apparently does a lot of computation constantly. If we can harness some of that computation at an appropriate level, we can potentially perform certain type of computation (much) faster and more efficiently than we can do with a von Neumann computer. Indeed, many powerful algorithms are inspired by nature and are thus prime candidates for nature-based computation. One particular branch of this effort that has seen some recent rapid advances is Ising machines. Some Ising machines are already showing better performance and energy efficiency for optimization problems. Through design iterations and co-evolution between hardware and algorithm, we expect more benefits from nature-based computing systems in the future. In this paper, we make a case for an augmented Ising machine suitable for both training and inference using an energy-based machine learning algorithm. We show that with a small change, the Ising substrate accelerates key parts of the algorithm and achieves non-trivial speedup and efficiency gain. With a more substantial change, we can turn the machine into a self-sufficient gradient follower to virtually complete training entirely in hardware. This can bring about 29x speedup and about 1000x reduction in energy compared to a Tensor Processing Unit (TPU) host.
Uday Kumar Reddy Vengalam, Yongchao Liu 0003, Tong Geng, Hui Wu 0007, Michael C. Huang 0001
MICRO1
2022 QuBRIM: A CMOS Compatible Resistively-Coupled Ising Machine with Quantized Nodal Interactions
abstract
Physical Ising machines have been shown to solve combinatoric optimization problems with orders-of-magnitude improvements in speed and energy efficiency o ver v on N eumann systems. However, building such a system is still in its infancy and a scalable, robust implementation remains challenging. CMOS-compatible electronic Ising machines (e.g., [1]) are promising as the mature technology helps bring scale, speed, and energy efficiency to the dynamical system. However, subtle issues can arise when using voltage-controlled transistors to act as programmable resistive coupling. In this paper, we propose a version of resistively-coupled Ising machine using quantized nodal interactions (QuBRIM), which significantly i mproved the predictability of the coupling resistor. The functionality of QuBRIM is demonstrated by solving the well-known Max-Cut problem using both behavioral and circuit level simulations in 45 nm CMOS technology node. We show that the dynamical system naturally seeks local minima in the objective function's energy landscape and that by applying spin-fix a nnealing, t he system reaches a global minimum with a high probability.
Yiqiao Zhang, Uday Kumar Reddy Vengalam, Anshujit Sharma, Michael C. Huang 0001, Zeljko Ignjatovic
ICCAD2
2022 A CMOS Compatible Bistable Resistively-coupled Ising Machine-BRIM
abstract
Ising machines and other nature-based computing platforms have recently become attractive due to their potential of outperforming conventional computers when solving problems that involve a large number of competing alternatives, such as combinatorial optimizations. In this paper, a newly proposed resistively-coupled Ising machine with bistable nodes (BRIM) is designed and simulated in a 45nm CMOS process. The performance of the proposed machine is evaluated based on its capability of solving the Max-cut graph problem. A spin-fix annealing technique is applied to help escape local minima and improve the solution quality. Simulation result shows that this technique effectively increases the probability of finding the Max-cut solution by 50.5% and reduces the solution error to 1.73 on average.
Yiqiao Zhang, Richard Afoakwa, Uday Kumar Reddy Vengalam, Michael C. Huang 0001, Zeljko Ignjatovic
ISCAS3
2022 LoopIn: A Loop-Based Simulation Sampling Mechanism
abstract
Understanding program behavior is at the heart of general-purpose architecture design. Whether we are testing a new design offline or making a design adapt to changing behavior online, a central assumption is that the test cases represent real workload in steady state. Typical computer programs have been known to exhibit patterns of runtime behavior that repeat during the course of their execution. Simulation and adaptation strategies all exploit this repetition to some extent. In this paper, we introduce a simple mechanism that is more explicit in identifying and exploiting behavior repetition at the granularity of (broadly defined) loops. The result is that a typical benchmark will be categorized into tens of loops. In terms of architectural simulations, this strategy will create a moderate number (on the orders of 100) of relatively short (tens of thousands of instructions) segments. There are two major benefits in our view. The first and more quantifiable benefit is that, the strategy requires less simulation and obtains increased accuracy compared to the commonly used SimPoint approach. Second, instead of depicting average statistics of an entire program, we can accurately describe intra-program behavior variation, which simple sampling strategies cannot. LoopIn produces many small simulation segments. In certain usage scenarios, microarchitectural state warm-up may be costly. In these cases, an existing tool BLRL can help create efficient warm-up arrangements.
Uday Kumar Reddy Vengalam, Anshujit Sharma, Michael C. Huang 0001
ISPASS1
2021 BRIM: Bistable Resistively-Coupled Ising Machine
abstract
Physical Ising machines rely on nature to guide a dynamical system towards an optimal state which can be read out as a heuristical solution to a combinatorial optimization problem. Such designs that use nature as a computing mechanism can lead to higher performance and/or lower operation costs. Quantum annealers are a prominent example of such efforts. However, existing Ising machines are generally bulky and energy intensive. Such disadvantages may be acceptable if these designs provide some significant intrinsic advantages at a much larger scale in the future, which remains to be seen. But for now, integrated electronic designs of Ising machines allow more immediate applications. We propose one such design that uses bistable nodes, coupled with programmable and variable strengths. The design is fully CMOS compatible for on-chip applications and demonstrates competitive solution quality and significantly superior execution time and energy.
Richard Afoakwa, Yiqiao Zhang, Uday Kumar Reddy Vengalam, Zeljko Ignjatovic, Michael C. Huang 0001
HPCA3
2018 Enabling Scientific Computing on Memristive Accelerators
abstract
Linear algebra is ubiquitous across virtually every field of science and engineering, from climate modeling to macroeconomics. This ubiquity makes linear algebra a prime candidate for hardware acceleration, which can improve both the run time and the energy efficiency of a wide range of scientific applications. Recent work on memristive hardware accelerators shows significant potential to speed up matrix-vector multiplication (MVM), a critical linear algebra kernel at the heart of neural network inference tasks. Regrettably, the proposed hardware is constrained to a narrow range of workloads: although the eight-to 16-bit computations afforded by memristive MVM accelerators are acceptable for machine learning, they are insufficient for scientific computing where high-precision floating point is the norm. This paper presents the first proposal to enable scientific computing on memristive crossbars. Three techniques are explored — reducing overheads by exploiting exponent range locality, early termination of fixed-point computation, and static operation scheduling — that together enable a fixed-point memristive accelerator to perform high-precision floating point without the exorbitant cost of naïve floating-point emulation on fixed-point hardware. A heterogeneous collection of crossbars with varying sizes is proposed to efficiently handle sparse matrices, and an algorithm for mapping the dense subblocks of a sparse matrix to an appropriate set of crossbars is investigated. The accelerator can be combined with existing GPU-based systems to handle datasets that cannot be efficiently handled by the memristive accelerator alone. The proposed optimizations permit the memristive MVM concept to be applied to a wide range of problem domains, respectively improving the execution time and energy dissipation of sparse linear solvers by 10.3x and 10.9x over a purely GPU-based system.
Ben Feinberg, Uday Kumar Reddy Vengalam, Nathan Whitehair, Engin Ipek
ISCA2