Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Mahdi Nikdan

dblp:298/2929 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Efficient and distributed learning · 73% Language models and text generation · 25% Optimization for machine learning · 2%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 82% Processor architecture and microarchitecture · 18%

Topics — the 16 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model training
1.722025
Quartet: Native FP4 Training Can Be Optimal for Large Language Models · NeurIPS 2025
QuEST: Stable Training of LLMs with 1-Bit Weights and Activations · ICML 2025
Machine learning › Efficient and distributed learning
low-precision training
1.722025
Quartet: Native FP4 Training Can Be Optimal for Large Language Models · NeurIPS 2025
QuEST: Stable Training of LLMs with 1-Bit Weights and Activations · ICML 2025
Machine learning › Efficient and distributed learning › model compression › quantization
quantization-aware training
1.722025
HALO: Hadamard-Assisted Lower-Precision Optimization for LLMs · NeurIPS 2025
QuEST: Stable Training of LLMs with 1-Bit Weights and Activations · ICML 2025
Machine learning › Efficient and distributed learning
model compression
1.622025
QuEST: Stable Training of LLMs with 1-Bit Weights and Activations · ICML 2025
RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation · ICML 2024
Natural language and speech › Language models and text generation
large language model fine-tuning
1.122025
Efficient Data Selection at Scale via Influence Distillation · NeurIPS 2025
RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation · ICML 2024
Machine learning › Efficient and distributed learning
data selection
0.912025
Efficient Data Selection at Scale via Influence Distillation · NeurIPS 2025
Machine learning › Efficient and distributed learning
distributed training
0.912025
HALO: Hadamard-Assisted Lower-Precision Optimization for LLMs · NeurIPS 2025
Machine learning › Efficient and distributed learning › distributed training › data parallel training
fully sharded data parallel
0.912025
HALO: Hadamard-Assisted Lower-Precision Optimization for LLMs · NeurIPS 2025
Machine learning › Efficient and distributed learning › data selection
influence-based data selection
0.912025
Efficient Data Selection at Scale via Influence Distillation · NeurIPS 2025
Natural language and speech › Language models and text generation › instruction tuning
instruction data selection
0.912025
Efficient Data Selection at Scale via Influence Distillation · NeurIPS 2025
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.812024
RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation · ICML 2024
Machine learning › Efficient and distributed learning › sparse computation
sparse backpropagation
0.712023
SparseProp: Efficient Sparse Backpropagation for Faster Training of Neural Networks at the Edge · ICML 2023
Machine learning › Efficient and distributed learning › model compression
sparse training
0.712023
SparseProp: Efficient Sparse Backpropagation for Faster Training of Neural Networks at the Edge · ICML 2023
Hardware accelerators and domain-specific architectures › machine learning accelerator › training accelerator
neural network training accelerator
0.712023
SparseProp: Efficient Sparse Backpropagation for Faster Training of Neural Networks at the Edge · ICML 2023
Machine learning › Optimization for machine learning
second-order optimization
0.312025
Efficient Data Selection at Scale via Influence Distillation · NeurIPS 2025
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.312025
Quartet: Native FP4 Training Can Be Optimal for Large Language Models · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

low-precision scaling law · 1.7FP4 training · 1.7CUDA kernel · 1.7trust gradient estimator · 0.9hadamard rotation · 0.9hadamard normalization · 0.9adam · 0.9PEFT · 0.9MSE-optimal fitting · 0.9FSDP · 0.9vectorized CPU implementation · 0.7sparse backpropagation · 0.7
YearPublicationVenuePosition
2025 QuEST: Stable Training of LLMs with 1-Bit Weights and Activations
abstract
One approach to reducing the massive costs of large language models (LLMs) is the use of quantized or sparse representations for training or deployment. While post-training compression methods are very popular, the question of obtaining even more accurate compressed models by directly training over such representations, i.e., Quantization-Aware Training (QAT), is still open: for example, a recent study put the "optimal" bit-width at which models can be trained using QAT, while staying accuracy-competitive with standard FP16/BF16 precision, at 8-bits weights and activations. We advance this state-of-the-art via a new method called QuEST, for which we demonstrate optimality at 4-bits and stable convergence as low as 1-bit weights and activations. QuEST achieves this by improving two key aspects of QAT methods: (1) accurate and fast quantization of the (continuous) distributions of weights and activations via Hadamard normalization and MSE-optimal fitting; (2) a new trust gradient estimator based on the idea of explicitly minimizing the error between the noisy gradient computed over quantized states and the "true" (but unknown) full-precision gradient. Experiments on Llama-type architectures show that QuEST induces stable scaling laws across the entire range of hardware-supported precisions, and can be extended to sparse representations. We provide GPU kernel support showing that models produced by QuEST can be executed efficiently. Our code is available at https://github.com/IST-DASLab/QuEST.
Andrei Panferov, Jiale Chen 0004, Soroush Tabesh, Mahdi Nikdan, Dan Alistarh
ICML4
2025 HALO: Hadamard-Assisted Lower-Precision Optimization for LLMs
abstract
Quantized training of Large Language Models (LLMs) remains an open challenge, as maintaining accuracy while performing all matrix multiplications in low precision has proven difficult. This is particularly the case when fine-tuning pre-trained models, which can have large weight, activation, and error (output gradient) outlier values that make lower-precision optimization difficult. To address this, we present HALO, a new quantization-aware training approach for Transformers that enables accurate and efficient low-precision training by combining 1) strategic placement of Hadamard rotations in both forward and backward passes, which mitigate outliers, 2) high-performance kernel support, and 3) FSDP integration for low-precision communication. Our approach ensures that all large matrix multiplications during the forward and backward passes are executed in lower precision. Applied to LLaMa models, HALO achieves near-full-precision-equivalent results during fine-tuning on various tasks, while delivering up to 1.41x end-to-end speedup for full fine-tuning on RTX 4090 GPUs. HALO efficiently supports both standard and parameter-efficient fine-tuning (PEFT). Our results demonstrate the first practical approach to fully quantized LLM fine-tuning that maintains accuracy in INT8 and FP6 precision, while delivering performance benefits.
Saleh Ashkboos, Mahdi Nikdan, Rush Tabesh, Roberto L. Castro, Torsten Hoefler, Dan Alistarh
NeurIPS2
2025 Quartet: Native FP4 Training Can Be Optimal for Large Language Models
abstract
Training large language models (LLMs) models directly in low-precision offers a way to address computational costs by improving both throughput and energy efficiency. For those purposes, NVIDIA's recent Blackwell architecture facilitates very low-precision operations using FP4 variants. Yet, current algorithms for training LLMs in FP4 precision face significant accuracy degradation and often rely on mixed-precision fallbacks. In this paper, we investigate hardware-supported FP4 training and introduce a new approach for accurate, end-to-end FP4 training with all the major computations (i.e., linear layers) in low precision. Through extensive evaluations on Llama-type models, we reveal a new low-precision scaling law that quantifies performance trade-offs across bit-widths and training setups. Guided by this investigation, we design an "optimal" technique in terms of accuracy-vs-computation, called Quartet. We implement Quartet using optimized CUDA kernels tailored for Blackwell, demonstrating that fully FP4-based training is a competitive alternative to FP16 half-precision and to FP8 training. Our code is available at https://github.com/IST-DASLab/Quartet .
Roberto L. Castro, Andrei Panferov, Rush Tabesh, Oliver Sieberling, Jiale Chen 0004, Mahdi Nikdan, Saleh Ashkboos, Dan Alistarh
NeurIPS6
2025 Efficient Data Selection at Scale via Influence Distillation
abstract
Effective data selection is critical for efficient training of modern Large Language Models (LLMs). This paper introduces Influence Distillation, a novel, mathematically-justified framework for data selection that employs second-order information to optimally weight training samples. By distilling each sample's influence on a target distribution, our method assigns model-specific weights that are used to select training data for LLM fine-tuning, guiding it toward strong performance on the target domain. We derive these optimal weights for both Gradient Descent and Adam optimizers. To ensure scalability and reduce computational cost, we propose a $\textit{landmark-based approximation}$: influence is precisely computed for a small subset of "landmark" samples and then efficiently propagated to all other samples to determine their weights. We validate Influence Distillation by applying it to instruction tuning on the Tulu V2 dataset, targeting a range of tasks including GSM8k, SQuAD, and MMLU, across several models from the Llama and Qwen families. Experiments show that Influence Distillation matches or outperforms state-of-the-art performance while achieving up to $3.5\times$ faster selection.
Mahdi Nikdan, Vincent Cohen-Addad, Dan Alistarh, Vahab S. Mirrokni
NeurIPS1
2024 RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation
abstract
We investigate parameter-efficient fine-tuning (PEFT) methods that can provide good accuracy under limited computational and memory budgets in the context of large language models (LLMs). We present a new PEFT method called Robust Adaptation (RoSA) inspired by robust principal component analysis that jointly trains $\textit{low-rank}$ and highly-sparse components on top of a set of fixed pretrained weights to efficiently approximate the performance of a full-fine-tuning (FFT) solution. Across a series of challenging generative tasks such as grade-school math and SQL query generation, which require fine-tuning for good performance, we show that RoSA outperforms LoRA, pure sparse fine-tuning, and alternative hybrid methods at the same parameter budget, and can even recover the performance of FFT on some tasks. We provide system support for RoSA to complement the training algorithm, specifically in the form of sparse GPU kernels which enable memory- and computationally-efficient training, and show that it is also compatible with low-precision base weights, resulting in the first joint representation combining quantization, low-rank and sparse approximations. Our code is available at https://github.com/IST-DASLab/RoSA.
Mahdi Nikdan, Soroush Tabesh, Elvir Crncevic, Dan Alistarh
ICML1
2023 SparseProp: Efficient Sparse Backpropagation for Faster Training of Neural Networks at the Edge
abstract
We provide an efficient implementation of the backpropagation algorithm, specialized to the case where the weights of the neural network being trained are sparse. Our algorithm is general, as it applies to arbitrary (unstructured) sparsity and common layer types (e.g., convolutional or linear). We provide a fast vectorized implementation on commodity CPUs, and show that it can yield speedups in end-to-end runtime experiments, both in transfer learning using already-sparsified networks, and in training sparse networks from scratch. Thus, our results provide the first support for sparse training on commodity hardware.
Mahdi Nikdan, Tommaso Pegolotti, Eugenia Iofinova, Eldar Kurtic, Dan Alistarh
ICML1
2021 3D Image Segmentation With Sparse Annotation by Self-Training and Internal Registration
abstract
Anatomical image segmentation is one of the foundations for medical planning. Recently, convolutional neural networks (CNN) have achieved much success in segmenting volumetric (3D) images when a large number of fully annotated 3D samples are available. However, rarely a volumetric medical image dataset containing a sufficient number of segmented 3D images is accessible since providing manual segmentation masks is monotonous and time-consuming. Thus, to alleviate the burden of manual annotation, we attempt to effectively train a 3D CNN using a sparse annotation where ground truth on just one 2D slice of the axial axis of each training 3D image is available. To tackle this problem, we propose a self-training framework that alternates between two steps consisting of assigning pseudo annotations to unlabeled voxels and updating the 3D segmentation network by employing both the labeled and pseudo labeled voxels. To produce pseudo labels more accurately, we benefit from both propagation of labels (or pseudo-labels) between adjacent slices and 3D processing of voxels. More precisely, a 2D registration-based method is proposed to gradually propagate labels between consecutive 2D slices and a 3D U-Net is employed to utilize volumetric information. Ablation studies on benchmarks show that cooperation between the 2D registration and the 3D segmentation provides accurate pseudo-labels that enable the segmentation network to be trained effectively when for each training sample only even one segmented slice by an expert is available. Our method is assessed on the CHAOS and Visceral datasets to segment abdominal organs. Results demonstrate that despite utilizing just one segmented slice for each 3D image (that is weaker supervision in comparison with the compared weakly supervised methods) can result in higher performance and also achieve closer results to the fully supervised manner.
Adeleh Bitarafan, Mahdi Nikdan, Mahdieh Soleymani Baghshah
IEEE J. Biomed. Health Informatics2