Amir H. Abdi

dblp:200/9533 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0002-3169-4477ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Efficient and distributed learning · 37% Language models and text generation · 21% Trustworthy machine learning · 18%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 100%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › large language model inference
long-context inference
1.722025
MMInference: Accelerating Pre-filling for Long-Context Visual Language Models via Modality-Aware Permutation Sparse Attention · ICML 2025
SCBench: A KV Cache-Centric Analysis of Long-Context Methods · ICLR 2025
Machine learning › Efficient and distributed learning
inference acceleration
1.622025
MMInference: Accelerating Pre-filling for Long-Context Visual Language Models via Modality-Aware Permutation Sparse Attention · ICML 2025
MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention · NeurIPS 2024
Machine learning › Efficient and distributed learning › inference acceleration
pre-filling acceleration
1.622025
MMInference: Accelerating Pre-filling for Long-Context Visual Language Models via Modality-Aware Permutation Sparse Attention · ICML 2025
MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention · NeurIPS 2024
Machine learning › Deep learning architectures and training › attention mechanism
sparse attention
1.622025
MMInference: Accelerating Pre-filling for Long-Context Visual Language Models via Modality-Aware Permutation Sparse Attention · ICML 2025
MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention · NeurIPS 2024
Machine learning › Trustworthy machine learning
interpretability
0.912025
Llama See, Llama Do: A Mechanistic Perspective on Contextual Entrainment and Distraction in LLMs · ACL (1) 2025
Machine learning › Efficient and distributed learning › KV cache management
KV cache compression
0.912025
SCBench: A KV Cache-Centric Analysis of Long-Context Methods · ICLR 2025
Machine learning › Efficient and distributed learning
KV cache management
0.912025
SCBench: A KV Cache-Centric Analysis of Long-Context Methods · ICLR 2025
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability
0.912025
Llama See, Llama Do: A Mechanistic Perspective on Contextual Entrainment and Distraction in LLMs · ACL (1) 2025
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization
long-context modeling
0.812024
MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention · NeurIPS 2024
Machine learning › Generative modeling
iterative refinement
0.712023
Scaleformer: Iterative Multi-scale Refining Transformers for Time Series Forecasting · ICLR 2023
Machine learning › Trustworthy machine learning › uncertainty estimation
selective classification
0.712023
Towards Better Selective Classification · ICLR 2023
Machine learning › Deep learning architectures and training
transformer
0.712023
Scaleformer: Iterative Multi-scale Refining Transformers for Time Series Forecasting · ICLR 2023
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization › long-context modeling
long-context language model
0.312025
SCBench: A KV Cache-Centric Analysis of Long-Context Methods · ICLR 2025
Computer vision › Vision and language
vision-language model
0.312025
MMInference: Accelerating Pre-filling for Long-Context Visual Language Models via Modality-Aware Permutation Sparse Attention · ICML 2025
Performance modeling and evaluation
benchmarking
0.312025
SCBench: A KV Cache-Centric Analysis of Long-Context Methods · ICLR 2025
Performance modeling and evaluation › benchmarking › machine learning benchmarking
long-context benchmark
0.312025
SCBench: A KV Cache-Centric Analysis of Long-Context Methods · ICLR 2025

Methods — techniques the papers use, named apart from their topics

sparse attention · 1.7quantization · 1.7prompt compression · 1.7dynamic sparse attention · 1.6permutation-based sparsity · 0.9differentiable masking · 0.9counterfactual prompting · 0.9GPU kernel optimization · 0.9pattern identification · 0.8GPU kernels · 0.8
YearPublicationVenuePosition
2025 Llama See, Llama Do: A Mechanistic Perspective on Contextual Entrainment and Distraction in LLMs
abstract
We observe a novel phenomenon, contextual entrainment, across a wide range of language models (LMs) and prompt settings, providing a new mechanistic perspective on how LMs become distracted by “irrelevant” contextual information in the input prompt. Specifically, LMs assign significantly higher logits (or probabilities) to any tokens that have previously appeared in the context prompt, even for random tokens. This suggests that contextual entrainment is a mechanistic phenomenon, occurring independently of the relevance or semantic relation of the tokens to the question or the rest of the sentence. We find statistically significant evidence that the magnitude of contextual entrainment is influenced by semantic factors. Counterfactual prompts have a greater effect compared to factual ones, suggesting that while contextual entrainment is a mechanistic phenomenon, it is modulated by semantic factors.We hypothesise that there is a circuit of attention heads — the entrainment heads — that corresponds to the contextual entrainment phenomenon. Using a novel entrainment head discovery method based on differentiable masking, we identify these heads across various settings. When we “turn off” these heads, i.e., set their outputs to zero, the effect of contextual entrainment is significantly attenuated, causing the model to generate output that capitulates to what it would produce if no distracting context were provided. Our discovery of contextual entrainment, along with our investigation into LM distraction via the entrainment heads, marks a key step towards the mechanistic analysis and mitigation of the distraction problem.
Jingcheng Niu, Xingdi Yuan, Hamidreza Saghir, Amir H. Abdi
ACL (1)5
2025 SCBench: A KV Cache-Centric Analysis of Long-Context Methods
abstract
Long-context Large Language Models (LLMs) have enabled numerous downstream applications but also introduced significant challenges related to computational and memory efficiency. To address these challenges, optimizations for long-context inference have been developed, centered around the KV cache. However, existing benchmarks often evaluate in single-request, neglecting the full lifecycle of the KV cache in real-world use. This oversight is particularly critical, as KV cache reuse has become widely adopted in LLMs inference frameworks, such as vLLM and SGLang, as well as by LLM providers, including OpenAI, Microsoft, Google, and Anthropic. To address this gap, we introduce SCBENCH (SharedContextBENCH), a comprehensive benchmark for evaluating long-context methods from a KV cache centric perspective: 1) KV cache generation, 2) KV cache compression, 3) KV cache retrieval, and 4) KV cache loading. Specifically, SCBench uses test examples with shared context, ranging 12 tasks with two shared context modes, covering four categories of long-context capabilities: string retrieval, semantic retrieval, global information, and multi-task. With SCBench, we provide an extensive KV cache-centric analysis of eight categories long-context solutions, including Gated Linear RNNs (Codestal-Mamba), Mamba-Attention hybrids (Jamba-1.5-Mini), and efficient methods such as sparse attention, KV cache dropping, quantization, retrieval, loading, and prompt compression. The evaluation is conducted on six Transformer-based long-context LLMs: Llama-3.1-8B/70B, Qwen2.5-72B/32B, Llama-3-8B-262K, and GLM-4-9B. Our findings show that sub-O(n) memory methods suffer in multi-turn scenarios, while sparse encoding with O(n) memory and sub-O(n^2) pre-filling computation perform robustly. Dynamic sparsity yields more expressive KV caches than static patterns, and layer-level sparsity in hybrid architectures reduces memory usage with strong performance. Additionally, we identify attention distribution shift issues in long-generation scenarios.
Huiqiang Jiang, Qianhui Wu, Xufang Luo, Surin Ahn, Chengruidong Zhang, Amir H. Abdi, Dongsheng Li 0002, Jianfeng Gao 0001, Yuqing Yang 0001, Lili Qiu
ICLR7
2025 MMInference: Accelerating Pre-filling for Long-Context Visual Language Models via Modality-Aware Permutation Sparse Attention
abstract
The integration of long-context capabilities with visual understanding unlocks unprecedented potential for Vision Language Models (VLMs). However, the quadratic attention complexity during the pre-filling phase remains a significant obstacle to real-world deployment. To overcome this limitation, we introduce MMInference (Multimodality Million tokens Inference), a dynamic sparse attention method that accelerates the prefilling stage for long-context multi-modal inputs. First, our analysis reveals that the temporal and spatial locality of video input leads to a unique sparse pattern, the Grid pattern. Simultaneously, VLMs exhibit markedly different sparse distributions across different modalities. We introduce a permutation-based method to leverage the unique Grid pattern and handle modality boundary issues. By offline search the optimal sparse patterns for each head, MMInference constructs the sparse distribution dynamically based on the input. We also provide optimized GPU kernels for efficient sparse computations. Notably, MMInference integrates seamlessly into existing VLM pipelines without any model modifications or fine-tuning. Experiments on multi-modal benchmarks-including Video QA, Captioning, VisionNIAH, and Mixed-Modality NIAH-with state-of-the-art long-context VLMs (LongVila, LlavaVideo, VideoChat-Flash, Qwen2.5-VL) show that MMInference accelerates the pre-filling stage by up to 8.3x at 1M tokens while maintaining accuracy. Our code is available at https://ama.ms/MMInference.
Huiqiang Jiang, Chengruidong Zhang, Qianhui Wu, Xufang Luo, Surin Ahn, Amir H. Abdi, Dongsheng Li 0002, Jianfeng Gao 0001, Yuqing Yang 0001, Lili Qiu
ICML7
2024 MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
abstract
The computational challenges of Large Language Model (LLM) inference remain a significant barrier to their widespread deployment, especially as prompt lengths continue to increase. Due to the quadratic complexity of the attention computation, it takes 30 minutes for an 8B LLM to process a prompt of 1M tokens (i.e., the pre-filling stage) on a single A100 GPU. Existing methods for speeding up prefilling often fail to maintain acceptable accuracy or efficiency when applied to long-context LLMs. To address this gap, we introduce MInference (Milliontokens Inference), a sparse calculation method designed to accelerate pre-filling of long-sequence processing. Specifically, we identify three unique patterns in long-context attention matrices-the A-shape, Vertical-Slash, and Block-Sparse-that can be leveraged for efficient sparse computation on GPUs. We determine the optimal pattern for each attention head offline and dynamically build sparse indices based on the assigned pattern during inference. With the pattern and sparse indices, we perform efficient sparse attention calculations via our optimized GPU kernels to significantly reduce the latency in the pre-filling stage of longcontext LLMs. Our proposed technique can be directly applied to existing LLMs without any modifications to the pre-training setup or additional fine-tuning. By evaluating on a wide range of downstream tasks, including InfiniteBench, RULER, PG-19, and Needle In A Haystack, and models including LLaMA-3-1M, GLM-4-1M, Yi-200K, Phi-3-128K, and Qwen2-128K, we demonstrate that MInference effectively reduces inference latency by up to 10x for pre-filling on an A100, while maintaining accuracy. Our code is available at https://aka.ms/MInference.
Huiqiang Jiang, Chengruidong Zhang, Qianhui Wu, Xufang Luo, Surin Ahn, Zhenhua Han, Amir H. Abdi, Dongsheng Li 0002, Chin-Yew Lin, Yuqing Yang 0001, Lili Qiu
NeurIPS8
2023 Towards Better Selective Classification
Leo Feng, Mohamed Osama Ahmed, Hossein Hajimirsadeghi, Amir H. Abdi
ICLR4
2023 Scaleformer: Iterative Multi-scale Refining Transformers for Time Series Forecasting
Mohammad Amin Shabani, Amir H. Abdi, Lili Meng, Tristan Sylvain
ICLR2
2022 TD-GEN: Graph Generation Using Tree Decomposition
abstract
We propose TD-GEN, a graph generation framework based on tree decomposition, and introduce a reduced upper bound on the maximum number of decisions needed for graph generation. The framework includes a permutation invariant tree generation model which forms the backbone of graph generation. Tree nodes are supernodes, each representing a cluster of nodes in the graph. Graph nodes and edges are incrementally generated inside the clusters by traversing the tree supernodes, respecting the structure of the tree decomposition, and following node sharing decisions between the clusters. Further, we discuss the shortcomings of the standard evaluation criteria based on statistical properties of the generated graphs. We propose to compare the generalizability of models based on expected likelihood. Empirical results on a variety of standard graph generation datasets demonstrate the superior performance of our method.
Hamed Shirzad, Hossein Hajimirsadeghi, Amir H. Abdi, Greg Mori
AISTATS3
2020 On Modelling Label Uncertainty in Deep Neural Networks: Automatic Estimation of Intra- Observer Variability in 2D Echocardiography Quality Assessment
abstract
Uncertainty of labels in clinical data resulting from intra-observer variability can have direct impact on the reliability of assessments made by deep neural networks. In this paper, we propose a method for modelling such uncertainty in the context of 2D echocardiography (echo), which is a routine procedure for detecting cardiovascular disease at point-of-care. Echo imaging quality and acquisition time is highly dependent on the operator's experience level. Recent developments have shown the possibility of automating echo image quality quantification by mapping an expert's assessment of quality to the echo image via deep learning techniques. Nevertheless, the observer variability in the expert's assessment can impact the quality quantification accuracy. Here, we aim to model the intra-observer variability in echo quality assessment as an aleatoric uncertainty modelling regression problem with the introduction of a novel method that handles the regression problem with categorical labels. A key feature of our design is that only a single forward pass is sufficient to estimate the level of uncertainty for the network output. Compared to the 0.11 ± 0.09 absolute error (in a scale from 0 to 1) archived by the conventional regression method, the proposed method brings the error down to 0.09 ± 0.08, where the improvement is statistically significant and equivalents to 5.7% test accuracy improvement. The simplicity of the proposed approach means that it could be generalized to other applications of deep learning in medical imaging, where there is often uncertainty in clinical labels.
Zhibin Liao, Hani Girgis, Amir H. Abdi, Hooman Vaseli, Jorden Hetherington, Robert Rohling, Ken Gin, Teresa Tsang, Purang Abolmaesumi
IEEE Trans. Medical Imaging3
2019 Variational Shape Completion for Virtual Planning of Jaw Reconstructive Surgery
Amir H. Abdi, Mehran Pesteie, Eitan Prisman, Purang Abolmaesumi, Sidney S. Fels
MICCAI (5)1
2019 Cardiac Phase Detection in Echocardiograms With Densely Gated Recurrent Neural Networks and Global Extrema Loss
abstract
Accurate detection of end-systolic (ES) and end-diastolic (ED) frames in an echocardiographic cine series can be difficult but necessary pre-processing step for the development of automatic systems to measure cardiac parameters. The detection task is challenging due to variations in cardiac anatomy and heart rate often associated with pathological conditions. We formulate this problem as a regression problem and propose several deep learning-based architectures that minimize a novel global extrema structured loss function to localize the ED and ES frames. The proposed architectures integrate convolution neural networks (CNNs)-based image feature extraction model and recurrent neural networks (RNNs) to model temporal dependencies between each frame in a sequence. We explore two CNN architectures: DenseNet and ResNet, and four RNN architectures: long short-term memory, bi-directional LSTM, gated recurrent unit (GRU), and Bi-GRU, and compare the performance of these models. The optimal deep learning model consists of a DenseNet and GRU trained with the proposed loss function. On average, we achieved 0.20 and 1.43 frame mismatch for the ED and ES frames, respectively, which are within reported inter-observer variability for the manual detection of these frames.
Fatemeh Taheri Dezaki, Zhibin Liao, Christina Luong 0001, Hani Girgis, Neeraj Dhungel, Amir H. Abdi, Delaram Behnami, Ken Gin, Robert Rohling, Purang Abolmaesumi, Teresa Tsang
IEEE Trans. Medical Imaging6
2017 Quality Assessment of Echocardiographic Cine Using Recurrent Neural Networks: Feasibility on Five Standard View Planes
Amir H. Abdi, Christina Luong 0001, Teresa Tsang, John Jue, Ken Gin, Darwin Yeung, Dale Hawley, Robert Rohling, Purang Abolmaesumi
MICCAI (3)1
2017 Automatic Quality Assessment of Echocardiograms Using Convolutional Neural Networks: Feasibility on the Apical Four-Chamber View
abstract
Echocardiography (echo) is a skilled technical procedure that depends on the experience of the operator. The aim of this paper is to reduce user variability in data acquisition by automatically computing a score of echo quality for operator feedback. To do this, a deep convolutional neural network model, trained on a large set of samples, was developed for scoring apical four-chamber (A4C) echo. In this paper, 6,916 end-systolic echo images were manually studied by an expert cardiologist and were assigned a score between 0 (not acceptable) and 5 (excellent). The images were divided into two independent training-validation and test sets. The network architecture and its parameters were based on the stochastic approach of the particle swarm optimization on the training-validation data. The mean absolute error between the scores from the ultimately trained model and the expert's manual scores was 0.71 ± 0.58. The reported error was comparable to the measured intra-rater reliability. The learned features of the network were visually interpretable and could be mapped to the anatomy of the heart in the A4C echo, giving confidence in the training result. The computation time for the proposed network architecture, running on a graphics processing unit, was less than 10 ms per frame, sufficient for real-time deployment. The proposed approach has the potential to facilitate the widespread use of echo at the point-of-care and enable early and timely diagnosis and treatment. Finally, the approach did not use any specific assumptions about the A4C echo, so it could be generalizable to other standard echo views.
Amir H. Abdi, Christina Luong 0001, Teresa Tsang, Gregory Allan, Saman Nouranian, John Jue, Dale Hawley, Sarah Fleming, Ken Gin, Jody Swift, Robert Rohling, Purang Abolmaesumi
IEEE Trans. Medical Imaging1
2017 Correction to "Automatic Quality Assessment of Echocardiograms Using Convolutional Neural Networks: Feasibility on the Apical Four-Chamber View"
abstract
In the above-title paper [ibid., vol. 36, no. 6, pp. 1221-1230, Jun. 2017], the first footnote should have indicated the following information: A. H. Abdi and C. Luong are joint first authors.
Amir H. Abdi, Christina Luong 0001, Teresa Tsang, Gregory Allan, Saman Nouranian, John Jue, Dale Hawley, Sarah Fleming, Ken Gin, Jody Swift, Robert Rohling, Purang Abolmaesumi
IEEE Trans. Medical Imaging1