VLDB 2026 Research / reviewers in the wild / expert
Kareem Ibrahim
dblp:373/5480
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2024
0000-0002-3951-0708ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Hardware accelerators and domain-specific architectures · 67% Memory systems · 33% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
computational neuroscience |
0.8 | 1 | 2024 | Marple: Scalable Spike Sorting for Untethered Brain-Machine Interfacing · ASPLOS (2) 2024 |
Bioinformatics and computational biology › neuroscience › neuroinformatics › neural data analysis
spike sorting |
0.8 | 1 | 2024 | Marple: Scalable Spike Sorting for Untethered Brain-Machine Interfacing · ASPLOS (2) 2024 |
Hardware accelerators and domain-specific architectures › bioinformatics accelerator
biomedical accelerator |
0.8 | 1 | 2024 | Marple: Scalable Spike Sorting for Untethered Brain-Machine Interfacing · ASPLOS (2) 2024 |
Memory systems
memory compression |
0.8 | 1 | 2024 | Atalanta: A Bit is Worth a "Thousand" Tensor Values · ASPLOS (2) 2024 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator |
0.8 | 1 | 2024 | Atalanta: A Bit is Worth a "Thousand" Tensor Values · ASPLOS (2) 2024 |
Information retrieval › pattern matching
template matching |
0.2 | 1 | 2024 | Marple: Scalable Spike Sorting for Untethered Brain-Machine Interfacing · ASPLOS (2) 2024 |
Methods — techniques the papers use, named apart from their topics
machine learning · 2.3hardware-software co-design · 0.8arithmetic coding · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Atalanta: A Bit is Worth a "Thousand" Tensor ValuesabstractAtalanta is a lossless, hardware/software co-designed compression technique for the tensors of fixed-point quantized deep neural networks. Atalanta increases effective memory capacity, reduces off-die traffic, and/or helps to achieve the desired performance/energy targets while using smaller off-die memories during inference. Atalanta is architected to deliver nearly identical coding efficiency compared to Arithmetic Coding while avoiding its complexity, overhead, and bandwidth limitations. Indicatively, the Atalanta decoder and encoder units each use less than 50B of internal storage. In hardware, Atalanta is implemented as an assist over any machine learning accelerator transparently compressing/decompressing tensors just before the off-die memory controller. This work shows the performance and energy efficiency of Atalanta when implemented in a 65nm technology node. Atalanta reduces data footprint of weights and activations to 60% and 48% respectively on average over a wide set of 8-bit quantized models and complements a wide range of quantization methods. Integrated with a Tensorcore-based accelerator, Atalanta boosts the speedup and energy efficiency to 1.44× and 1.37×, respectively. Atalanta is effective at compressing the stashed activations during training for fixed-point inference. Alberto Delmas Lascorz, Mostafa Mahmoud, Ali Hadi Zadeh, Milos Nikolic 0002, Kareem Ibrahim, Christina Giannoula, Ameer Abdelhadi, Andreas Moshovos |
ASPLOS (2) | 5 |
| 2024 | Marple: Scalable Spike Sorting for Untethered Brain-Machine InterfacingabstractSpike sorting is the process of parsing electrophysiological signals from neurons to identify if, when, and which particular neurons fire. Spike sorting is a particularly difficult task in computational neuroscience due to the growing scale of recording technologies and complexity in traditional spike sorting algorithms. Previous spike sorters can be divided into software-based and hardware-based solutions. Software solutions are highly accurate but operate on recordings after-the-fact, and often require utilization of high-power GPUs to process in a timely fashion, and they cannot be used in portable applications. Hardware solutions suffer in terms of accuracy due to the simplification of mechanisms for implementation's sake and process only up to 128 inputs. This work answers the question: "How much computation power and memory storage is needed to sort spikes from 1000s of channels to keep up with advances in probe technology?" We analyze the computational and memory requirements for modern software spike sorters to identify their potential bottlenecks - namely in the template memory storage. We architect Marple, a highly optimized hardware pipeline for spike sorting which incorporates a novel mechanism to reduce the template memory storage from 8 - 11x. Marple is scalable, uses a flexible vector-based back-end to perform neuron identification, and a fixed-function front-end to filter the incoming streams into areas of interest. The implementation is projected to use just 79mW in 7nm, when spike sorting 10K channels at peak activity. We further demonstrate, for the first time, a machine learning replacement for the template matching stage. Eugene Sha, Andy Wei Liu, Kareem Ibrahim, Mostafa Mahmoud, Christina Giannoula, Ameer Abdelhadi, Andreas Moshovos |
ASPLOS (2) | 3 |