Muya Chang

dblp:237/0955 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0002-3035-1106ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2026 CryoBoost: A 40nm Cryogenic-CMOS Matrix Multiplication Accelerator for Energy Efficient Computing
abstract
This paper proposes a cryogenic Matrix Multiplication (MATMUL) accelerator chip to address the exponential increase in energy consumption in AI model training. The accelerator, designed in a 40nm CMOS process, leverages a Liquid Nitrogen-based cooling system, operating from 300K to 77K. The design comprises 10 processing elements (PEs) operating on 4x4 matrices at INT8 precision, interconnected by a core data ring. The PEs use a load-store architecture with a 128-bit very long instruction word (VLIW)-based controller. The paper presents a detailed characterization of the supply voltage versus frequency, performance, and power across different temperatures. The results indicate a significant reduction in power at cryogenic temperatures, with up to 45.7% reduction in power at 77K compared to 300K at iso performance. The maximum energy efficiency increases from 1.975GHz/W at 300K to 2.497GHz/W at 77K, yielding a 26.4% gain. This translates to up to 20% reduction in training energy for Large Language Models.
Rakshith Saligram, Samuel Spetalnick, Brian Crafton, Muya Chang, Alec Nordlund, Joshua Gess, Ruslan Nagimov, Arijit Raychowdhury
ACM Great Lakes Symposium on VLSI4
2024 A 24/48V to 0.8V-1.2V All-Digital Synchronous Buck Converter with Package-Integrated GaN power FETs and 180nm Silicon Controller IC
abstract
This paper presents a 24V/48V input, 0.8V- 1.2V output, two-phase, single-stage point-of-load (PoL) synchronous buck converter with enhanced-mode Gallium Nitride (GAN) N-FET based output stage and 180nm HV BCD silicon based all-digital control. The GaN devices, silicon controller chip and bootstrapping (BST) capacitors are heterogeneously integrated on an organic package substrate, thus providing a System in Package (SiP) solution, to enable high efficiency (>76% for 48:1, >86% for 24:1) while delivering 10.52W at 1V at a switching frequency of 5MHz, with a net Figure of Merit (FoM) of 11,520 MHz•V - a 13% improvement over the state-of-the art (SoA).
Kaushik Bhattacharyya, Minxiang Gong, Muya Chang, Xin Zhang 0025, Arijit Raychowdhury
ISCAS3
2024 E-Gaze: Gaze Estimation With Event Camera
abstract
Near-eye gaze estimation is a task that maps the recording of an eye captured by an adjacent camera to the direction of a person's gaze in space. In contrast to frame-based cameras, event cameras are characterized by high sensing rates, low latency, sparse asynchronous data outputs, and high dynamic range, which are well suited for recording the fast eye movements. However, algorithms and system designs that operate on frame-based cameras are not applicable to event-based data, due to the natural differences in the data characteristics. In this work, we study the pattern of near-eye event-based data streams and extract eye features to estimate gaze. First, by analyzing eye parts and movements, and harnessing the polar, spatial, and temporal distribution of the events, we introduce a real-time pipeline to extract pupil features. Second, we present a recurrent neural network with a proposed coordinate-to-angle loss function to accurately estimate gaze from pupil feature sequence. We demonstrated that our system achieves accurate real-time estimation with angular accuracy of 0.46° and update rates of 950 Hz, thus opening up avenues for novel applications. To our knowledge, this is the first system that operates only on event-based data to perform gaze estimation.
Nealson Li, Muya Chang, Arijit Raychowdhury
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 In-Situ Privacy via Mixed-Signal Perturbation and Hardware-Secure Data Reversibility
abstract
The swift proliferation of edge intelligence and ubiquitous data generation have heightened privacy into a pressing societal need. State-of-the-art reversible privacy protection requires significant hardware resources at the edge with distinct architecture for sensors and security, leading to a rise in hardware overhead and expanded attack surfaces. To address these challenges, we propose a time-domain mixed-signal (TD-MS) circuit architecture facilitating in-situ privacy (ISP) with hardware-secured data reversibility. The proposed TD-MS ISP unites data acquisition, data conversion, key generation, and protection while providing authorized device-specific unclonable data recovery for forensic purposes. At the system level, we demonstrate the attack resilience and privacy-preserving computation performance by implementing a custom embedded system applied to real-world surveillance scenarios. At the circuit level, we showcase custom TD-MS circuits, evaluating their energy and area efficiency against a digital baseline implemented in 65nm technology. With full-stack SPICE simulations for both the baseline digital and proposed TD-MS circuits, we measured a$670\times$energy/frame savings against the embedded system,$3\times$area reduction and$3.2\times$energy TD-MS gains over digital.
Steven Davis, Boyang Cheng, Muya Chang, Ningyuan Cao
IEEE Trans. Circuits Syst. I Regul. Pap.4
2023 Neuromorphic Swarm on RRAM Compute-in-Memory Processor for Solving QUBO Problem
abstract
Combinatorial optimization problems prevail in engineering and industry. Some are NP-hard and thus become difficult to solve on edge devices due to limited power and computing resources. Quadratic Unconstrained Binary Optimization (QUBO) problem is a valuable emerging model that can formulate numerous combinatorial problems, such as Max-Cut, traveling salesman problems, and graphic coloring. QUBO model also reconciles with two emerging computation models, quantum computing and neuromorphic computing, which can potentially boost the speed and energy efficiency in solving combinatorial problems. In this work, we design a neuromorphic QUBO solver composed of a swarm of spiking neural networks (SNN) that conduct a population-based meta-heuristic search for solutions. The proposed model can achieve about x20 40 speedup on large QUBO problems in terms of time steps compared to a traditional neural network solver. As a codesign, we evaluate the neuromorphic swarm solver on a 40nm 25mW Resistive RAM (RRAM) Compute-in-Memory (CIM) SoC with a 2.25MB RRAM-based accelerator and an embedded Cortex M3 core. The collaborative SNN swarm can fully exploit the specialty of CIM accelerator in matrix and vector multiplications. Compared to previous works, such an algorithm-hardware synergized solver exhibits advantageous speed and energy efficiency for edge devices.
Ashwin Sanjay Lele, Muya Chang, Samuel Spetalnick, Brian Crafton, Arijit Raychowdhury, Yan Fang 0002
DAC2
2023 Privacy-by-Sensing with Time-domain Differentially-Private Compressed Sensing
abstract
With the ubiquitous IoT sensors and enormous real-time data generation, data privacy is becoming a critical societal concern. State-of-the-art privacy protection methods all demand significant hardware overhead due to computation-insensitive algorithms and divided sensor/security architecture. In this paper, we propose a generic time-domain circuit architecture that protects raw data by enabling a differentially-private compressed sensing (DP-CS) algorithm secured by physical unclonable functions (PUF). To address privacy concerns and hardware overhead at the same time, a robust unified PUF and time-domain mixed-signal (TD-MS) module are designed, where PUF enables private and secure entropy generation. To evaluate the proposed design against a digital baseline, we performed experiments based on synthesized circuits and SPICE simulation and measured a 2.9x area reduction and 3.2x energy gains. We also measured high-quality PUF generation with TD-MS circuit with a inter-die Hamming distance of 52% and a low intra-die Hamming distance of 2.8%. Furthermore, we performed attack and algorithm performance measurements demonstrating the proposed design preserves data privacy even under attack, and the machine learning performance has minimal degradation (within 2%) compared to the digital baseline.
Boyang Cheng, Pengyu Zeng, Steven Davis, Muya Chang, Ningyuan Cao
DATE5
2023 Live Demonstration: Hybrid RRAM and SRAM SoC for Fused Frame and Event Target Tracking
abstract
Event and frame cameras capture the complemen-tary spatial and temporal details of a scene providing an accuracy vs. latency trade-off. Fusing these processing modalities using convolutional (CNN) and spiking neural networks (SNN) respectively has been shown for target tracking. We present our heterogeneous RRAM compute-in-memory (CIM) and SRAM compute-near-memory (CNM) SoC for simultaneous processing of CNN and SNN. We will show the advantage of using fused vision over frame-only vision and demonstrate python programmable data streaming. The visitors will be able to see the processing-dependent dynamic power gating of non-volatile RRAM and in-memory error correction capability.
Ashwin Sanjay Lele, Muya Chang, Samuel Spetalnick, Yan Fang 0002, Brian Crafton, Shota Konno, Arijit Raychowdhury
ISCAS2
2022 Stochastic Mixed-Signal Circuit Design for In-Sensor Privacy
abstract
The ubiquitous data acquisition and extensive data exchange of sensors pose severe security and privacy concerns for the end-users and the public. To enable real-time protection of raw data, it is demanding to facilitate privacy-preserving algorithms at data generation, or in-sensory privacy. However, due to the severe sensor resource constraints and intensive computation/security cost, it remains an open question of how to enable data protection algorithms with efficient circuit techniques. To answer this question, this paper discusses the potential of a stochastic mixed-signal (SMS) circuit for ultra-low-power, small-foot-print data security. In particular, this paper discusses digitally-controlled-oscillators (DCO) and their advantages in (1) seamless analog interface, (2) stochastic computation efficiency, and (3) unified entropy generation over conventional digital circuit baselines. With DCO as an illustrative case, we target (1) SMS privacy-preserving architecture definition and systematic SMS analysis on its performance gains across various hardware/software configurations, and (2) revisit analog/mixed-signal voltage/transistor scaling in the context of entropy-based data protection.
Ningyuan Cao, Boyang Cheng, Muya Chang
ICCAD4
2019 Efficient Signal Reconstruction via Distributed Least Square Optimization on a Systolic FPGA Architecture
abstract
Optimization problems form the basis of a wide gamut of computationally challenging tasks in signal processing, machine learning, resource planning and so on. Out of these, convex optimization, and in particular least square optimization, covers a vast majority; and recent advances in iterative algorithms to solve such problems of large dimensions have gained traction. Multi-core designs with systolic or semi-systolic architectures can be a key enabler for implementing discrete dynamical systems and realize massively scalable architectures to solve such optimization algorithms. In this paper, we present a platform architecture implemented in programmable FPGA hardware to solve a template problem in distributed optimization, namely signal reconstruction from non-uniform sampling. This is a quintessential problem with wide-spread applications in signal processing, computational imaging etc. We expect such an architectural exploration to open up promising opportunities to solve distributed optimizations that are becoming increasingly important in real-world applications. The complete system design, mapping and optimization into an FPGA architecture as well as analysis of convergence and scalability have been presented.
Muya Chang, Samantak Gangopadhyay, Tomer Hamam, Justin K. Romberg, Arijit Raychowdhury
ICASSP1