EDBT 2026 Demo / reviewers in the wild / expert
Malte J. Rasch
dblp:68/5535
· DBLP profile ↗
11ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0002-7988-4624ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 since 2021Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Assessing the Performance of Analog Training for Transfer LearningabstractAnalog in-memory computing is a next-generation computing paradigm that promises fast, parallel, and energy-efficient deep learning training and transfer learning (TL). However, achieving this promise has remained elusive due to a lack of suitable training algorithms. Analog memory devices exhibit asymmetric and non-linear switching behavior in addition to device-to-device variation, meaning that most, if not all, of the current off-the-shelf training algorithms cannot achieve good training outcomes. Also, recently introduced algorithms have enjoyed limited attention, as they require bi-directionally switching devices of unrealistically high symmetry and precision and are highly sensitive. A new algorithm chopped TTv2 (c-TTv2), has been introduced, which leverages the chopped technique to address many of the challenges mentioned above. In this paper, we assess the performance of the c-TTv2 algorithm for analog TL using a Swin-ViT model on a subset of the CIFAR100 dataset. We also investigate the robustness of our algorithm to changes in some device specifications, including weight transfer noise, symmetry point skew, and symmetry point variability. Omobayode Fagbohungbe, Corey Lammie, Malte J. Rasch, Takashi Ando, Tayfun Gokmen, Vijay Narayanan |
ISCAS | 3 |
| 2024 | Analog AI as a Service: A Cloud Platform for In-Memory ComputingabstractThis paper introduces the Analog AI Cloud Composer platform, a service that allows users to access Analog In-Memory Computing (AIMC) simulation and computing resources over the cloud. We introduce the concept of an Analog AI as a Service (AAaaS). AIMC offers a novel approach for decreasing both the latency and energy usage associated with Deep Neural Network (DNN) inference and training. This platform democratizes access to AIMC computing, making it available to a broader audience, including researchers, developers, and businesses. Emphasizing a user-friendly, no-code approach, AAaaS integrates the Analog Hardware Acceleration Kit (AIHWKit) simulation platform within a fully managed cloud environment. We discuss the architecture of the Analog AI Cloud Composer (AAICC), focusing on its key services such as inference, training, and AIMC hardware access. The platform's design, grounded in cloud services and guidelines, ensures a secure, data-centric user experience with robust control and validation mechanisms. Kaoutar El Maghraoui, Kim Tran, Kurtis Ruby, Borja Godoy, Jordan Murray, Manuel Le Gallo-Bourdeau, Todd Deshane, Pablo Gonzalez, Diego Moreda, Hadjer Benmeziane, Corey Lammie, Julian Büchel, Malte J. Rasch, Abu Sebastian, Vijay Narayanan |
SSE | 13 |
| 2024 | Improving the Accuracy of Analog-Based In-Memory Computing Accelerators Post-TrainingabstractAnalog-Based In-Memory Computing (AIMC) inference accelerators can be used to efficiently execute Deep Neural Network (DNN) inference workloads. However, to mitigate accuracy losses, due to circuit and device non-idealities, Hardware-Aware (HWA) training methodologies must be employed. These typically require significant information about the underlying hardware. In this paper, we propose two Post-Training (PT) optimization methods to improve accuracy after training is performed. For each crossbar, the first optimizes the conductance range of each column, and the second optimizes the input, i.e, Digital-to-Analog Converter (DAC), range. It is demonstrated that, when these methods are employed, the complexity during training, and the amount of information about the underlying hardware can be reduced, with no notable change in accuracy (≤0.1%) when finetuning the pretrained RoBERTa transformer model for all General Language Understanding Evaluation (GLUE) benchmark tasks. Additionally, it is demonstrated that further optimizing learned parameters PT improves accuracy. Corey Lammie, Athanasios Vasilopoulos, Julian Büchel, Giacomo Camposampiero, Manuel Le Gallo, Malte J. Rasch, Abu Sebastian |
ISCAS | 6 |
| 2024 | Towards Exact Gradient-based Training on Analog In-memory ComputingabstractGiven the high economic and environmental costs of using large vision or language models, analog in-memory accelerators present a promising solution for energy-efficient AI. While inference on analog accelerators has been studied recently, the training perspective is underexplored. Recent studies have shown that the "workhorse" of digital AI training - stochastic gradient descent (SGD) algorithm converges inexactly when applied to model training on non-ideal devices. This paper puts forth a theoretical foundation for gradient-based training on analog devices. We begin by characterizing the non-convergent issue of SGD, which is caused by the asymmetric updates on the analog devices. We then provide a lower bound of the asymptotic error to show that there is a fundamental performance limit of SGD-based analog training rather than an artifact of our analysis.
To address this issue, we study a heuristic analog algorithm called Tiki-Taka that has recently exhibited superior empirical performance compared to SGD. We rigorously show its ability to converge to a critical point exactly and hence eliminate the asymptotic error. The simulations verify the correctness of the analyses. Zhaoxian Wu, Tayfun Gokmen, Malte J. Rasch, Tianyi Chen 0002 |
NeurIPS | 3 |
| 2023 | Architectures and Circuits for Analog-memory-based Hardware Accelerators for Deep Neural Networks (Invited)abstractAnalog non-volatile memory (NVM)-based accelerators for Deep Neural Networks (DNNs) can achieve high-throughput and energy-efficient multiply-accumulate (MAC) operations by taking advantage of massively parallelized analog compute, implemented with Ohm's law and Kirchhoff's current law on arrays of resistive memory devices. Competitive end-to-end DNN accuracies can be obtained, provided that weights are accurately programmed onto NVM devices and MAC operations are sufficiently linear. In this paper, we report architectural and circuit advances for such Analog NVM-based accelerators. We describe a highly heterogeneous and programmable accelerator architecture for DNN inference that combines analog NVM memory-array “Tiles” for weight-stationary, energy-efficient MAC operations, together with heterogeneous special-function Compute-Cores for auxiliary digital computation. Massively parallel vectors of neuron-activation data are exchanged over short distances using a dense and efficient circuit-switched 2D mesh, enabling a wide range of DNN workloads, including CNNs, LSTMs, and Transformers. We also show a 14-nm inference chip consisting of multiple$\mathbf{512}\times \mathbf{512}$arrays of Phase Change Memory (PCM) devices which implements multiple DNN benchmarks using such a circuit-switched 2D mesh. Hsinyu Tsai, Pritish Narayanan, Shubham Jain 0004, Stefano Ambrogio, Kohji Hosokawa, Masatoshi Ishii, Charles Mackin, Ching-Tzu Chen, Atsuya Okazaki, Akiyo Nomura, Irem Boybat, Ramachandran Muralidhar, Martin M. Frank, Takeo Yasuda, Alexander M. Friz, Yasuteru Kohda, An Chen 0002, Andrea Fasoli, Malte J. Rasch, Stanislaw Wozniak, Jose Luquin, Vijay Narayanan, Geoffrey W. Burr |
ISCAS | 19 |
| 2022 | Analog-memory-based 14nm Hardware Accelerator for Dense Deep Neural Networks including TransformersabstractAnalog non-volatile memory (NVM)-based accelerators for deep neural networks perform high-throughput and energy-efficient multiply-accumulate (MAC) operations (e.g., high TeraOPS/W) by taking advantage of massively parallelized analog MAC operations, implemented with Ohm’s law and Kirchhoff’s current law on array-matrices of resistive devices. While the wide-integer and floating-point operations offered by conventional digital CMOS computing are much more suitable than analog computing for conventional applications that require high accuracy and true reproducibility, deep neural networks can still provide competitive end-to-end results even with modest (e.g., 4-bit) precision in synaptic operations. In this paper, we describe a 14-nm inference chip, comprising multiple 512$\times$ 512 arrays of Phase Change Memory (PCM) devices, which can deliver software-equivalent inference accuracy for MNIST handwritten-digit recognition and recurrent LSTM benchmarks, by using compensation techniques to finesse analog-memory challenges such as conductance drift and noise. We also project accuracy for Natural Language Processing (NLP) tasks performed with a state-of-art large Transformer-based model, BERT, when mapped onto an extended version of this same fundamental chip architecture. Atsuya Okazaki, Pritish Narayanan, Stefano Ambrogio, Kohji Hosokawa, Hsinyu Tsai, Akiyo Nomura, Takeo Yasuda, Charles Mackin, Alexander M. Friz, Masatoshi Ishii, Yasuteru Kohda, Katie Spoon, An Chen 0002, Andrea Fasoli, Malte J. Rasch, Geoffrey W. Burr |
ISCAS | 15 |
| 2019 | Training Large-Scale Spiking Neural Networks on Multi-core Neuromorphic System Using Backpropagation
Megumi Ito, Malte J. Rasch, Masatoshi Ishii, Atsuya Okazaki, SangBum Kim, Junka Okazawa, Akiyo Nomura, Kohji Hosokawa, Wilfried Haensch |
ICONIP (3) | 2 |
| 2012 | A Kernel Two-Sample Test
Arthur Gretton, Karsten M. Borgwardt, Malte J. Rasch, Bernhard Schölkopf, Alexander J. Smola |
J. Mach. Learn. Res. | 3 |
| 2011 | Learning Variance Statistics of Natural Images
Libo Ma, Malte J. Rasch, Si Wu 0001 |
ISNN (2) | 2 |
| 2007 | A Kernel Approach to Comparing Distributions
Arthur Gretton, Karsten M. Borgwardt, Malte J. Rasch, Bernhard Schölkopf, Alexander J. Smola |
AAAI | 3 |
| 2006 | A Kernel Method for the Two-Sample-ProblemabstractWe propose two statistical tests to determine if two samples are from different dis- tributions. Our test statistic is in both cases the distance between the means of the two samples mapped into a reproducing kernel Hilbert space (RKHS). The first test is based on a large deviation bound for the test statistic, while the second is based on the asymptotic distribution of this statistic. The test statistic can be com- puted in O(m2) time. We apply our approach to a variety of problems, including attribute matching for databases using the Hungarian marriage method, where our test performs strongly. We also demonstrate excellent performance when compar- ing distributions over graphs, for which no alternative tests currently exist. Arthur Gretton, Karsten M. Borgwardt, Malte J. Rasch, Bernhard Schölkopf, Alexander J. Smola |
NIPS | 3 |