VLDB 2026 Research / reviewers in the wild / expert
Sayeh Sharify
dblp:188/6204
· DBLP profile ↗
9ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Efficient and distributed learning · 74% Language models and text generation · 13% Optimization for machine learning · 13% | |
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Hardware accelerators and domain-specific architectures · 78% Memory systems · 12% Processor architecture and microarchitecture · 7% |
Topics — the 15 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
model compression |
1.8 | 3 | 2025 | ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals · ICML 2025 Mixed-Precision Quantization for Deep Vision Models with Integer Quadratic Programming · DAC 2025 Loom: exploiting weight and activation precisions to accelerate convolutional neural networks · DAC 2018 |
Machine learning › Efficient and distributed learning › model compression › quantization
mixed-precision quantization |
1.7 | 2 | 2025 | ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals · ICML 2025 Mixed-Precision Quantization for Deep Vision Models with Integer Quadratic Programming · DAC 2025 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
1.1 | 3 | 2020 | Late Breaking Results: Building an On-Chip Deep Learning Memory Hierarchy Brick by Brick · DAC 2020 ShapeShifter: Enabling Fine-Grain Data Width Adaptation in Deep Learning · MICRO 2019 Loom: exploiting weight and activation precisions to accelerate convolutional neural networks · DAC 2018 |
Natural language and speech › Language models and text generation
large language model inference |
0.9 | 1 | 2025 | ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals · ICML 2025 |
Machine learning › Efficient and distributed learning › model compression › quantization
post-training quantization |
0.9 | 1 | 2025 | ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals · ICML 2025 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator |
0.8 | 2 | 2019 | Laconic deep learning inference acceleration · ISCA 2019 Bit-Tactical: A Software/Hardware Approach to Exploiting Value and Bit Sparsity in Neural Networks · ASPLOS 2019 |
Memory systems
memory compression |
0.5 | 2 | 2020 | ShapeShifter: Enabling Fine-Grain Data Width Adaptation in Deep Learning · MICRO 2019 Late Breaking Results: Building an On-Chip Deep Learning Memory Hierarchy Brick by Brick · DAC 2020 |
Machine learning › Efficient and distributed learning
inference acceleration |
0.4 | 1 | 2019 | Laconic deep learning inference acceleration · ISCA 2019 |
Hardware accelerators and domain-specific architectures › sparsity exploitation
bit-level sparsity exploitation |
0.4 | 1 | 2019 | Bit-Tactical: A Software/Hardware Approach to Exploiting Value and Bit Sparsity in Neural Networks · ASPLOS 2019 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › CNN accelerator
CNN inference accelerator |
0.3 | 1 | 2018 | Loom: exploiting weight and activation precisions to accelerate convolutional neural networks · DAC 2018 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › DNN accelerator
precision-scalable accelerator |
0.3 | 1 | 2018 | Loom: exploiting weight and activation precisions to accelerate convolutional neural networks · DAC 2018 |
Processor architecture and microarchitecture
data-parallel architecture |
0.3 | 1 | 2017 | Bit-pragmatic deep neural network computing · MICRO 2017 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN accelerator |
0.3 | 1 | 2017 | Bit-pragmatic deep neural network computing · MICRO 2017 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.2 | 2 | 2019 | ShapeShifter: Enabling Fine-Grain Data Width Adaptation in Deep Learning · MICRO 2019 Loom: exploiting weight and activation precisions to accelerate convolutional neural networks · DAC 2018 |
Energy-efficient computing › energy-efficient architecture
energy-efficient accelerator |
0.1 | 1 | 2017 | Bit-pragmatic deep neural network computing · MICRO 2017 |
Methods — techniques the papers use, named apart from their topics
sensitivity analysis · 0.9random rotation · 0.9principal component analysis · 0.9low-rank residual · 0.9integer quadratic programming · 0.9fine-grain data width encoding · 0.8dynamic width selection · 0.8per-layer precision · 0.7bit-parallel multiply-accumulate · 0.7lossless compression · 0.4fixed-point quantization · 0.4static scheduling · 0.4sparse shuffling network · 0.4shift-and-add multiplication · 0.3serial-parallel multiplication · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mixed-Precision Quantization for Deep Vision Models with Integer Quadratic ProgrammingabstractQuantization is a widely used technique to compress neural networks. Assigning uniform bit-widths across all layers can result in significant accuracy degradation at low precision and inefficiency at high precision. Mixed-precision quantization (MPQ) addresses this by assigning varied bit-widths to layers, optimizing the accuracy-efficiency trade-off. Existing sensitivity-based methods for MPQ assume that quantization errors across layers are independent, which leads to suboptimal choices. We introduce CLADO, a practical sensitivity-based MPQ algorithm that captures cross-layer dependency of quantization error. CLADO approximates pairwise cross-layer errors using linear equations on a small data subset. Layerwise bit-widths are assigned by optimizing a new MPQ formulation based on cross-layer quantization errors using an Integer Quadratic Program. Experiments with CNN and vision transformer models on ImageNet demonstrate that CLADO achieves state-of-the-art mixed-precision quantization performance. Code repository available here.11https://github.com/JamesTuna/CLADOMPQ Sayeh Sharify, Michael Orshansky |
DAC | 2 |
| 2025 | ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank ResidualsabstractPost-training quantization (PTQ) of large language models (LLMs) holds the promise in reducing the prohibitive computational cost at inference time. Quantization of all weight, activation and key-value (KV) cache tensors to 4-bit without significantly degrading generalizability is challenging, due to the high quantization error caused by extreme outliers in activations. To tackle this problem, we propose ResQ, a PTQ method that pushes further the state-of-the-art. By means of principal component analysis (PCA), it identifies a low-rank subspace (in practice 1/8 of the hidden dimension) in which activation variances are highest, and keep the coefficients within this subspace in high precision, e.g. 8-bit, while quantizing the rest to 4-bit. Within each subspace, invariant random rotation is applied to further suppress outliers. We show that this is a provably optimal mixed precision quantization scheme that minimizes error. With the Llama and Qwen2.5 families of models, we demonstrate that ResQ outperforms recent uniform and mixed precision PTQ methods on a variety of benchmarks, achieving up to 33% lower perplexity on Wikitext than the next best method SpinQuant, and upto 3X speedup over 16-bit baseline. Anonymous code repository available at https://anonymous.4open.science/r/project-resq-2142. Utkarsh Saxena, Sayeh Sharify, Kaushik Roy 0001 |
ICML | 2 |
| 2025 | Understanding the Difficulty of Low-Precision Post-Training Quantization for LLMsabstractLarge language models of high parameter counts are computationally expensive, yet can be made much more efficient by compressing their weights to very low numerical precision. This can be achieved either through post-training quantization by minimizing local, layer-wise quantization errors, or through quantization-aware fine-tuning by minimizing the global loss function. In this study, we discovered that, under the same data constraint, the former approach nearly always fared worse than the latter, a phenomenon particularly prominent when the numerical precision is very low. We further showed that this difficulty of post-training quantization arose from stark misalignment between optimization of the local and global objective functions. Our findings suggested limited utility in minimization of local quantization error and the importance of direct quantization-aware fine-tuning, in the regime of large models at very low precision. Zifei Xu, Sayeh Sharify, Wanzin Yazar, Tristan James Webb |
IJCNN | 2 |
| 2020 | Late Breaking Results: Building an On-Chip Deep Learning Memory Hierarchy Brick by BrickabstractData accesses between on- and off-chip memories account for a large fraction of overall energy consumption during inference with deep learning networks. We present Boveda, a lossless on-chip memory compression technique for neural networks operating on fixed-point values. Boveda reduces the datawidth used per block of values to be only as long as necessary: since most values are of small magnitude Boveda drastically reduces their footprint. Boveda can be used to increase the effective on-chip capacity, to reduce off-chip traffic, or to reduce the on-chip memory capacity needed to achieve a performance/energy target. Boveda reduces total model footprint to 53%. Isak Edo Vivancos, Sayeh Sharify, Milos Nikolic 0002, Ciaran Bannon, Mostafa Mahmoud, Alberto Delmas Lascorz, Andreas Moshovos |
DAC | 2 |
| 2019 | Bit-Tactical: A Software/Hardware Approach to Exploiting Value and Bit Sparsity in Neural NetworksabstractWeight and activation sparsity can be leveraged in hardware to boost the performance and energy efficiency of Deep Neural Networks during inference. Fully capitalizing on sparsity requires re-scheduling and mapping the execution stream to deliver non-zero weight/activation pairs to multiplier units for maximal utilization and reuse. However, permitting arbitrary value re-scheduling in memory space and in time places a considerable burden on hardware to perform dynamic at-runtime routing and matching of values, and incurs significant energy inefficiencies. Bit-Tactical (TCL) is a neural network accelerator where the responsibility for exploiting weight sparsity is shared between a novel static scheduling middleware, and a co-designed hardware front-end with a lightweight sparse shuffling network comprising two (2- to 8-input) multiplexers per activation input. We empirically motivate two back-end designs chosen to target bit-sparsity in activations, rather than value-sparsity, with two benefits: a) we avoid handling the dynamically sparse whole-value activation stream, and b) we uncover more ineffectual work. TCL outperforms other state-of-the-art accelerators that target sparsity for weights and activations, the dynamic precision requirements of activations, or their bit-level sparsity for a variety of neural networks. Alberto Delmas Lascorz, Patrick Judd, Dylan Malone Stuart, Zissis Poulos, Mostafa Mahmoud, Sayeh Sharify, Milos Nikolic 0002, Kevin Siu, Andreas Moshovos |
ASPLOS | 6 |
| 2019 | Laconic deep learning inference accelerationabstractWe present a method for transparently identifying ineffectual computations during inference with Deep Learning models. Specifically, by decomposing multiplications down to the bit level, the amount of work needed by multiplications during inference can be potentially reduced by at least 40× across a wide selection of neural networks (8b and 16b). This method produces numerically identical results and does not affect overall accuracy. We present Laconic, a hardware accelerator that implements this approach to boost energy efficiency for inference with Deep Learning Networks. Laconic judiciously gives up some of the work reduction potential to yield a low-cost, simple, and energy efficient design that outperforms other state-of-the-art accelerators: an optimized DaDianNao-like design [13], Eyeriss [15], SCNN [71], Pragmatic [3], and BitFusion [83]. We study 16b, 8b, and 1b/2b fixed-point quantized models. Sayeh Sharify, Alberto Delmas Lascorz, Mostafa Mahmoud, Milos Nikolic 0002, Kevin Siu, Dylan Malone Stuart, Zissis Poulos, Andreas Moshovos |
ISCA | 1 |
| 2019 | ShapeShifter: Enabling Fine-Grain Data Width Adaptation in Deep LearningabstractWe show that selecting a data width for all values in Deep Neural Networks, quantized or not and even if that width is different per layer, amounts to worst-case design. Much shorter data widths can be used if we target the common case by adjusting the data type width at a much finer granularity. We propose ShapeShifter, where we group weights and activations and encode them using a width specific to each group and where typical group sizes vary from 16 to 256 values. The per group widths are selected statically for the weights and dynamically by hardware for the activations. We present two applications of ShapeShifter. In the first, that is applicable to any system, ShapeShifter reduces off- and on-chip storage and communication. This ShapeShifter-based memory compression is simple and low cost yet reduces off-chip traffic to 33% and 36% for 8-bit and 16-bit models respectively. This makes it possible to sustain higher performance for a given off-chip memory interface while also boosting energy efficiency. In the second application, we show how ShapeShifter can be implemented as a surgical extension over designs that exploit variable precision in time. Alberto Delmas Lascorz, Sayeh Sharify, Isak Edo Vivancos, Dylan Malone Stuart, Omar Mohamed Awad, Patrick Judd, Mostafa Mahmoud, Milos Nikolic 0002, Kevin Siu, Zissis Poulos, Andreas Moshovos |
MICRO | 2 |
| 2018 | Loom: exploiting weight and activation precisions to accelerate convolutional neural networksabstractLoom (LM), a hardware inference accelerator for Convolutional Neural Networks (CNNs) is presented. In LM every bit of data precision that can be saved translates to proportional performance gains. For both weights and activations LM exploits profile-derived per layer precisions. However, at runtime LM further trims activation precisions at a much smaller than a layer granularity. On average, across several image classification CNNs and for a configuration that can perform the equivalent of 128 16b × 16b multiply-accumulate operations per cycle LM outperforms a state-of-the-art bit-parallel accelerator [3] by 3.19× without any loss in accuracy while being 2.59× more energy efficient. LM can trade-off accuracy for additional improvements in execution performance and energy efficiency and compares favorably to an accelerator that targeted only activation precisions. Sayeh Sharify, Alberto Delmas Lascorz, Kevin Siu, Patrick Judd, Andreas Moshovos |
DAC | 1 |
| 2017 | Bit-pragmatic deep neural network computingabstractDeep Neural Networks expose a high degree of parallelism, making them amenable to highly data parallel architectures. However, data-parallel architectures often accept inefficiency in individual computations for the sake of overall efficiency. We show that on average, activation values of convolutional layers during inference in modern Deep Convolutional Neural Networks (CNNs) contain 92% zero bits. Processing these zero bits entails ineffectual computations that could be skipped. We propose Pragmatic (PRA), a massively data-parallel architecture that eliminates most of the ineffectual computations on-the-fly, improving performance and energy efficiency compared to state-of-the-art high-performance accelerators [5]. The idea behind PRA is deceptively simple: use serial-parallel shift-and-add multiplication while skipping the zero bits of the serial input. However, a straightforward implementation based on shift-and-add multiplication yields unacceptable area, power and memory access overheads compared to a conventional bit-parallel design. PRA incorporates a set of design decisions to yield a practical, area and energy efficient design. Jorge Albericio, Alberto Delmas Lascorz, Patrick Judd, Sayeh Sharify, Gerard O'Leary, Roman Genov, Andreas Moshovos |
MICRO | 4 |