EDBT 2026 Demo / reviewers in the wild / expert
Jack Cai
dblp:312/6051
· DBLP profile ↗
5ranked-venue papers
4as first author
5since 2021 · last 2025
0009-0008-1892-1875ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Attention-Level SpeculationabstractAs Large Language Models (LLMs) grow in size and context length, efficient inference strategies are essential to maintain low-latency token generation. Unfortunately, conventional tensor and data parallelism face diminishing returns when scaling across multiple devices. We propose a novel form—attention-level speculative parallelism (ALSpec)—that predicts self-attention outputs to execute subsequent operations early on separate devices. Our approach overlaps attention and non-attention computations, reducing the attention latency overhead at 128K context length by up to 5x and improving end-to-end decode latency by up to 1.65x, all without sacrificing quality. We establish the fundamental pillars for speculative execution and provide an execution paradigm that simplifies implementation. We show that existing attention-approximation methods perform well on simple information retrieval tasks, but they fail in advanced reasoning and math. Combined with speculative execution, we can approximate up to 90% of self-attention without harming model correctness. Demonstrated on Tenstorrent’s NPU devices, we scale up LLM inference beyond current techniques, paving the way for faster inference in transformer models. Jack Cai, Ammar Vora, Randolph Zhang, Mark O'Connor, Mark C. Jeffrey |
ICML | 1 |
| 2024 | In-Memory Transformer Self-Attention Mechanism Using Passive Memristor CrossbarabstractTransformers have emerged as the state-of-the-art architecture for natural language processing (NLP) and computer vision. However, they are inefficient in both conventional and in-memory computing architectures as doubling their sequence length quadruples their time and memory complexity due to their self-attention mechanism. Traditional methods optimize self-attention using memory-efficient algorithms or approximate methods, such as locality-sensitive hashing (LSH) attention that reduces time and memory complexity from O(L2) to O(L log L). In this work, we propose a hardware-level solution that further improves the computational efficiency of LSH attention by utilizing in-memory computing with semi-passive memristor arrays. We demonstrate that LSH can be performed with low-resolution, energy-efficient 0T1R arrays performing stochastic memristive vector-matrix multiplication (VMM). Using circuit-level simulation, we show our proposed method is feasible as a drop-in approximation in Large Language Models (LLMs) with no degradation in evaluation metrics. Our results set the foundation for future works on computing the entire transformer architecture in-memory. Jack Cai, Muhammad Ahsan Kaleem, Roman Genov, Mostafa Rahimi Azghadi, Amirali Amirsoleimani |
ISCAS | 1 |
| 2023 | HESSPROP: Mitigating Memristive DNN Weight Mapping Errors with Hessian BackpropagationabstractA universal objective function to minimize mem-ristive crossbar deep neural network weight mapping errors through Hessian backpropagation (HessProp) is presented. Hes-sProp minimizes the$L_{2}$norm of the neural network gradient to achieve a flat minima in a neural network's weight space. We hypothesize that this leads to robustness against small perturbations of weights. The stochastic weight mapping phe-nomenon on memristor crossbars is simulated, and the proposed method was evaluated on image classification tasks using the MNIST dataset. The result demonstrates on average 40.81% and 41.45% groundbreaking accuracy increase for distilled and large memristive convolutional neural networks in worst-case scenarios. Jack Cai, Muhammad Ahsan Kaleem, Amirali Amirsoleimani, Roman Genov |
ISCAS | 1 |
| 2023 | A Survey of Ensemble Methods for Mitigating Memristive Neural Network Non-idealitiesabstractIn this work, ensemble methods are presented and tested as universal ways to improve the performance of Memristive Deep Neural Networks (MDNNs) with non-idealities. The Generalized Ensemble Method and Weighted Voting ensemble methods improve the accuracy of classification on the MNIST dataset by 6.5% and 6.6% respectively, thus showing that they are more effective than basic Ensemble Averaging which has been investigated before, as well as other methods such as Voting. Different weighting schemes for Weighted Voting were tested, and we present Algorithm 1 and 2, which are the theoretically and experimentally optimal weighting schemes respectively. Our work serves as a guideline for choosing ensemble methods for MDNNs. Muhammad Ahsan Kaleem, Jack Cai, Amirali Amirsoleimani, Roman Genov |
ISCAS | 2 |
| 2022 | HYPERLOCK: In-Memory Hyperdimensional Encryption in Memristor Crossbar ArrayabstractWe present a novel cryptography architecture based on memristor crossbar array, binary hypervectors, and neural network. Utilizing the stochastic and unclonable nature of memristor crossbar and error tolerance of binary hypervectors and neural network, implementation of the algorithm on memristor crossbar simulation is made possible. We demonstrate that with an increasing dimension of the binary hypervectors, the nonidealities in the memristor circuit can be effectively controlled. At the fine level of controlled crossbar non-ideality, noise from memristor circuit can be used to encrypt data while being sufficiently interpretable by neural network for decryption. We applied our algorithm on image cryptography for proof of concept, and to text en/decryption with 100% decryption accuracy despite crossbar noises. Our work shows the potential and feasibility of using memristor crossbars as an unclonable stochastic encoder unit of cryptography on top of their existing functionality as a vectormatrix multiplication acceleration device. Jack Cai, Amirali Amirsoleimani, Roman Genov |
ISCAS | 1 |