EDBT 2026 Demo / reviewers in the wild / expert
Vahid Partovi Nia
dblp:178/0912
· DBLP profile ↗
22ranked-venue papers
0as first author
16since 2021 · last 2026
0000-0001-6673-4224ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 11 since 2021Artificial intelligence and machine learning · 11 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enterprise Resource Planning Using Multi-Type Transformers in Ferro-Titanium Industry
Samira Yazdanpourmoghadam, Mahan Balal Pour, Vahid Partovi Nia |
ICPRAM | 3 |
| 2025 | OAC: Output-adaptive Calibration for Accurate Post-training QuantizationabstractDeployment of Large Language Models (LLMs) has major computational costs, due to their rapidly expanding size. Compression of LLMs reduces the memory footprint, latency, and energy required for their inference. Post-training Quantization (PTQ) techniques have been developed to compress LLMs while avoiding expensive re-training. Most PTQ approaches formulate the quantization error based on a layer-wise Euclidean loss, ignoring the model output. Then, each layer is calibrated using its layer-wise Hessian to update the weights towards minimizing the quantization error. The Hessian is also used for detecting the most salient weights to quantization. Such PTQ approaches are prone to accuracy drop in low-precision quantization. We propose Output-adaptive Calibration (OAC) to incorporate the model output in the calibration process. We formulate the quantization error based on the distortion of the output cross-entropy loss. OAC approximates the output-adaptive Hessian for each layer under reasonable assumptions to reduce the computational complexity. The output-adaptive Hessians are used to update the weight matrices and detect the salient weights towards maintaining the model output. Our proposed method outperforms the state-of-the-art baselines such as SpQR and BiLLM, especially, at extreme low-precision (2-bit and binary) quantization. Ali Edalati, Alireza Ghaffari, Mahsa Ghazvini Nejad, Boxing Chen, Masoud Asgharian, Vahid Partovi Nia |
AAAI | 7 |
| 2025 | PoT-PTQ: Two-Step Power-of-Two Post-Training for LLMsabstractLarge Language Models (LLMs) have demonstrated remarkable performance across various natural language processing (NLP) tasks. However, their deployment is challenging due to the substantial computational resources required. Power-of-two (PoT) quantization is a general tool to counteract this difficulty. Albeit previous works on PoT quantization can be efficiently dequantized on CPUs using fixed-point addition, it showed less effectiveness on GPUs. The reason is entanglement of the sign bit and sequential bit manipulations needed for dequantization. We propose a novel POT quantization framework for LLM weights that (i) outperforms state-of-the-art accuracy in extremely low-precision number formats, and (ii) enables faster inference through more efficient dequantization. To maintain the accuracy of the quantized model, we introduce a two-step post-training algorithm: (i) initialize the quantization scales with a robust starting point, and (ii) refine these scales using a minimal calibration set. The performance of our PoT post-training algorithm surpasses the current state-of-the-art in integer quantization, particularly at low precisions such as 2- and 3-bit formats. Our PoT quantization accelerates the dequantization step required for the floating point inference and leads to 3.67× speed up on a NVIDIA V100, and 1.63× on a NVIDIA RTX 4090, compared to uniform integer dequantization. Xinyu Wang 0061, Vahid Partovi Nia, Peng Lu 0006, Jerry Huang, Xiao-Wen Chang, Boxing Chen, Yufei Cui |
ECAI | 2 |
| 2025 | Zeroth Order Optimization for Pretraining Language ModelsabstractABSTRACT: The physical memory for training Large Language Models (LLMs) grow with the model size, and are limited to the GPU memory. In particular, back-propagation that requires the computation of the first-order derivatives adds to this memory overhead. Training extremely large language models with memory-efficient algorithms is still a challenge with theoretical and practical implications. Back-propagation-free training algorithms, also known as zeroth-order methods, are recently examined to address this challenge. Their usefulness has been proven in fine-tuning of language models. However, so far, there has been no study for language model pretraining using zeroth-order optimization, where the memory constraint is manifested more severely. We build the connection between the second order, the first order, and the zeroth order theoretically. Then, we apply the zeroth order optimization to pre-training light-weight language models, and discuss why they cannot be readily applied. We show in p articular that the curse of dimensionality is the main obstacle, and pave the way towards modifications of zeroth order methods for pre-training such models. Nathan Allaire, Mahsa Ghazvini Nejad, Sébastien Le Digabel, Vahid Partovi Nia |
ICPRAM | 4 |
| 2025 | Rethinking Post-Training Quantization: Introducing a Statistical Pre-Calibration Approach
Alireza Ghaffari, Sharareh Younesian, Boxing Chen, Vahid Partovi Nia, Masoud Asgharian |
ICPRAM | 4 |
| 2024 | Mitigating Outlier Activations in Low-Precision Fine-Tuning of Language Models
Alireza Ghaffari, Justin Yu, Mahsa Ghazvini Nejad, Masoud Asgharian, Boxing Chen, Vahid Partovi Nia |
ICPRAM | 6 |
| 2023 | DenseShift : Towards Accurate and Efficient Low-Bit Power-of-Two QuantizationabstractEfficiently deploying deep neural networks on low-resource edge devices is challenging due to their ever-increasing resource requirements. To address this issue, researchers have proposed multiplication-free neural networks, such as Power-of-Two quantization, or also known as Shift networks, which aim to reduce memory usage and simplify computation. However, existing low-bit Shift networks are not as accurate as their full-precision counterparts, typically suffering from limited weight range encoding schemes and quantization loss. In this paper, we propose the DenseShift network, which significantly improves the accuracy of Shift networks, achieving competitive performance to full-precision networks for vision and speech applications. In addition, we introduce a method to deploy an efficient DenseShift network using non-quantized floating-point activations, while obtaining 1.6× speed-up over existing methods. To achieve this, we demonstrate that zero-weight values in low-bit Shift networks do not contribute to model capacity and negatively impact inference computation. To address this issue, we propose a zero-free shifting mechanism that simplifies inference and increases model capacity. We further propose a sign-scale decomposition design to enhance training efficiency and a low-variance random initialization strategy to improve the model’s transfer learning performance. Our extensive experiments on various computer vision and speech tasks demonstrate that DenseShift outperforms existing low-bit multiplication-free networks and achieves competitive performance compared to full-precision networks. Furthermore, our proposed approach exhibits strong transfer learning performance without a drop in accuracy. Our code was released on GitHub. Xinlin Li 0001, Bang Liu 0003, Rui Heng Yang, Vanessa Courville, Vahid Partovi Nia |
ICCV | 6 |
| 2023 | On the Convergence of Stochastic Gradient Descent in Low-Precision Number FormatsabstractDeep learning models are dominating almost all artificial intelligence tasks such as vision, text, and speech processing.Stochastic Gradient Descent (SGD) is the main tool for training such models, where the computations are usually performed in single-precision floating-point number format.The convergence of single-precision SGD is normally aligned with the theoretical results of real numbers since they exhibit negligible error.However, the numerical error increases when the computations are performed in low-precision number formats.This provides compelling reasons to study the SGD convergence adapted for low-precision computations.We present both deterministic and stochastic analysis of the SGD algorithm, obtaining bounds that show the effect of number format.Such bounds can provide guidelines as to how SGD convergence is affected when constraints render the possibility of performing high-precision computations remote. Matteo Cacciola, Antonio Frangioni, Masoud Asgharian, Alireza Ghaffari, Vahid Partovi Nia |
ICPRAM | 5 |
| 2023 | Understanding Neural Network Binarization with Forward and Backward Proximal QuantizersabstractIn neural network binarization, BinaryConnect (BC) and its variants are considered the standard. These methods apply the sign function in their forward pass and their respective gradients are backpropagated to update the weights. However, the derivative of the sign function is zero whenever defined, which consequently freezes training. Therefore, implementations of BC (e.g., BNN) usually replace the derivative of sign in the backward computation with identity or other approximate gradient alternatives. Although such practice works well empirically, it is largely a heuristic or ``training trick.'' We aim at shedding some light on these training tricks from the optimization perspective. Building from existing theory on ProxConnect (PC, a generalization of BC), we (1) equip PC with different forward-backward quantizers and obtain ProxConnect++ (PC++) that includes existing binarization techniques as special cases; (2) derive a principled way to synthesize forward-backward quantizers with automatic theoretical guarantees; (3) illustrate our theory by proposing an enhanced binarization algorithm BNN++; (4) conduct image classification experiments on CNNs and vision transformers, and empirically verify that BNN++ generally achieves competitive results on binarizing these models. Yiwei Lu 0001, Yaoliang Yu, Xinlin Li 0001, Vahid Partovi Nia |
NeurIPS | 4 |
| 2022 | Convolutional Neural Network Compression through Generalized Kronecker Product DecompositionabstractModern Convolutional Neural Network (CNN) architectures, despite their superiority in solving various problems, are generally too large to be deployed on resource constrained edge devices. In this paper, we reduce memory usage and floating-point operations required by convolutional layers in CNNs. We compress these layers by generalizing the Kronecker Product Decomposition to apply to multidimensional tensors, leading to the Generalized Kronecker Product Decomposition (GKPD). Our approach yields a plug-and-play module that can be used as a drop-in replacement for any convolutional layer. Experimental results for image classification on CIFAR-10 and ImageNet datasets using ResNet, MobileNetv2 and SeNet architectures substantiate the effectiveness of our proposed approach. We find that GKPD outperforms state-of-the-art decomposition methods including Tensor-Train and Tensor-Ring as well as other relevant compression methods such as pruning and knowledge distillation. Marawan Gamal Abdel Hameed, Marzieh S. Tahaei, Ali Mosleh 0003, Vahid Partovi Nia |
AAAI | 4 |
| 2022 | EuclidNets: Combining Hardware and Architecture Design for Efficient Training and Inference
Mariana Oliveira Prazeres, Xinlin Li 0001, Adam M. Oberman, Vahid Partovi Nia |
ICPRAM | 4 |
| 2022 | iRNN: Integer-only Recurrent Neural NetworkabstractRecurrent neural networks (RNN) are used in many real-world text and speech applications. They include complex modules such as recurrence, exponential-based activation, gate interaction, unfoldable normalization, bi-directional dependence, and attention. The interaction between these elements prevents running them on integer-only operations without a significant performance drop. Deploying RNNs that include layer normalization and attention on integer-only arithmetic is still an open problem. We present a quantization-aware training method for obtaining a highly accurate integer-only recurrent neural network (iRNN). Our approach supports layer normalization, attention, and an adaptive piecewise linear approximation of activations (PWL), to serve a wide range of RNNs on various applications. The proposed method is proven to work on RNN-based language models and challenging automatic speech recognition, enabling AI applications on the edge. Our iRNN maintains similar performance as its full-precision counterpart, their deployment on smartphones improves the runtime performance by $2\times$, and reduces the model size by $4\times$. Eyyüb Sari, Vanessa Courville, Vahid Partovi Nia |
ICPRAM | 3 |
| 2022 | KroneckerBERT: Significant Compression of Pre-trained Language Models Through Kronecker Decomposition and Knowledge DistillationabstractMarzieh Tahaei, Ella Charlaix, Vahid Nia, Ali Ghodsi, Mehdi Rezagholizadeh. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Marzieh S. Tahaei, Ella Charlaix, Vahid Partovi Nia, Ali Ghodsi 0001, Mehdi Rezagholizadeh |
NAACL-HLT | 3 |
| 2022 | Is Integer Arithmetic Enough for Deep Learning Training?abstractThe ever-increasing computational complexity of deep learning models makes their training and deployment difficult on various cloud and edge platforms. Replacing floating-point arithmetic with low-bit integer arithmetic is a promising approach to save energy, memory footprint, and latency of deep learning models. As such, quantization has attracted the attention of researchers in recent years. However, using integer numbers to form a fully functional integer training pipeline including forward pass, back-propagation, and stochastic gradient descent is not studied in detail. Our empirical and mathematical results reveal that integer arithmetic seems to be enough to train deep learning models. Unlike recent proposals, instead of quantization, we directly switch the number representation of computations. Our novel training method forms a fully integer training pipeline that does not change the trajectory of the loss and accuracy compared to floating-point, nor does it need any special hyper-parameter tuning, distribution adjustment, or gradient clipping. Our experimental results show that our proposed method is effective in a wide variety of tasks such as classification (including vision transformers), object detection, and semantic segmentation. Alireza Ghaffari, Marzieh S. Tahaei, Mohammadreza Tayaranian, Masoud Asgharian, Vahid Partovi Nia |
NeurIPS | 5 |
| 2021 | Demystifying and Generalizing BinaryConnectabstractBinaryConnect (BC) and its many variations have become the de facto standard for neural network quantization. However, our understanding of the inner workings of BC is still quite limited. We attempt to close this gap in four different aspects: (a) we show that existing quantization algorithms, including post-training quantization, are surprisingly similar to each other; (b) we argue for proximal maps as a natural family of quantizers that is both easy to design and analyze; (c) we refine the observation that BC is a special case of dual averaging, which itself is a special case of the generalized conditional gradient algorithm; (d) consequently, we propose ProxConnect (PC) as a generalization of BC and we prove its convergence properties by exploiting the established connections. We conduct experiments on CIFAR-10 and ImageNet, and verify that PC achieves competitive performance. Tim Dockhorn, Yaoliang Yu, Eyyüb Sari, Mahdi Zolnouri, Vahid Partovi Nia |
NeurIPS | 5 |
| 2021 | S$^3$: Sign-Sparse-Shift Reparametrization for Effective Training of Low-bit Shift NetworksabstractShift neural networks reduce computation complexity by removing expensive multiplication operations and quantizing continuous weights into low-bit discrete values, which are fast and energy-efficient compared to conventional neural networks. However, existing shift networks are sensitive to the weight initialization and yield a degraded performance caused by vanishing gradient and weight sign freezing problem. To address these issues, we propose S$^3$ re-parameterization, a novel technique for training low-bit shift networks. Our method decomposes a discrete parameter in a sign-sparse-shift 3-fold manner. This way, it efficiently learns a low-bit network with weight dynamics similar to full-precision networks and insensitive to weight initialization. Our proposed training method pushes the boundaries of shift neural networks and shows 3-bit shift networks compete with their full-precision counterparts in terms of top-1 accuracy on ImageNet. Xinlin Li 0001, Bang Liu 0003, Yaoliang Yu, Wulong Liu, Chunjing Xu, Vahid Partovi Nia |
NeurIPS | 6 |
| 2020 | Activation Adaptation in Neural NetworksabstractMany neural network architectures rely on the choice of the activation function for each hidden layer. Given the activation function, the neural network is trained over the bias and the weight parameters. The bias catches the center of the activation, and the weights capture the scale. Here we propose to train the network over a shape parameter as well. This view allows each neuron to tune its own activation function and adapt the neuron curvature towards a better prediction. This modification only adds one further equation to the back-propagation for each neuron. Re-formalizing activation functions as CDF generalizes the class of activation function extensively. We aimed at generalizing an extensive class of activation functions to study: i) skewness and ii) smoothness of activation functions. Here we introduce adaptive Gumbel activation function as a bridge between Gumbel and sigmoid. A similar approach is used to invent a smooth version of ReLU. Our comparison with common activation functions suggests different data representation especially in early neural network layers. This adaptation also provides prediction improvement. Farnoush Farhadi, Vahid Partovi Nia, Andrea Lodi 0001 |
ICPRAM | 2 |
| 2019 | Active Learning for High-Dimensional Binary FeaturesabstractErbium-doped fiber amplifier (EDFA) is an optical amplifier/repeater device used to boost the intensity of optical signals being carried through fiber optic communication networks. A highly accurate EDFA model - to predict the signal gain for each channel - is required because of its crucial role in optical network management and optimization. EDFA channel inputs (i.e. features) either carry signal or are idle, therefore they can be treated as binary features. However, channel outputs (and the corresponding signal gains) are continuous values. Labeled training data is very expensive to collect for EDFA devices, therefore we devise an active learning strategy suitable for binary features to overcome this issue. We propose to take advantage of sparse linear models to simplify the predictive model. This approach improves signal gain prediction and accelerates active learning query generation. We show the performance of our proposed active learning strategies on simulated data and real EDFA data. Ali Vahdat, Mouloud Belbahri, Vahid Partovi Nia |
CNSM | 3 |
| 2019 | Multiomics modeling of the immunome, transcriptome, microbiome, proteome and metabolome adaptations during human pregnancyabstractMotivation: Multiple biological clocks govern a healthy pregnancy. These biological mechanisms produce immunologic, metabolomic, proteomic, genomic and microbiomic adaptations during the course of pregnancy. Modeling the chronology of these adaptations during full-term pregnancy provides the frameworks for future studies examining deviations implicated in pregnancy-related pathologies including preterm birth and preeclampsia. Results: We performed a multiomics analysis of 51 samples from 17 pregnant women, delivering at term. The datasets included measurements from the immunome, transcriptome, microbiome, proteome and metabolome of samples obtained simultaneously from the same patients. Multivariate predictive modeling using the Elastic Net (EN) algorithm was used to measure the ability of each dataset to predict gestational age. Using stacked generalization, these datasets were combined into a single model. This model not only significantly increased predictive power by combining all datasets, but also revealed novel interactions between different biological modalities. Future work includes expansion of the cohort to preterm-enriched populations and in vivo analysis of immune-modulating interventions based on the mechanisms identified. Availability and implementation: Datasets and scripts for reproduction of results are available through: https://nalab.stanford.edu/multiomics-pregnancy/. Supplementary information: Supplementary data are available at Bioinformatics online. Mohammad Sajjad Ghaemi, Daniel B. DiGiulio, Kévin Contrepois, Benjamin J. Callahan, Thuy T. M. Ngo, Brittany Lee-McMullen, Benoit Lehallier, Anna Robaczewska, David Mcilwain, Yael Rosenberg-Hasson, Ronald J. Wong, Cecele Quaintance, Anthony Culos, Natalie Stanley, Athena Tanada, Amy Tsai, Dyani Gaudilliere, Edward Ganio, Xiaoyuan Han, Kazuo Ando, Leslie McNeil, Martha Tingle, Paul H. Wise, Ivana Maric, Marina Sirota, Tony Wyss-Coray, Virginia D. Winn, Maurice L. Druzin, Ronald Gibbs, Gary L. Darmstadt, David B. Lewis, Vahid Partovi Nia, Bruno Agard, Robert Tibshirani, Garry P. Nolan, Michael Snyder 0001, David A. Relman, Stephen R. Quake, Gary M. Shaw, David K. Stevenson, Martin S. Angst, Brice Gaudilliere, Nima Aghaeepour |
Bioinform. | 32 |
| 2018 | Causal Inference and Mechanism Clustering of A Mixture of Additive Noise ModelsabstractThe inference of the causal relationship between a pair of observed variables is a fundamental problem in science, and most existing approaches are based on one single causal model. In practice, however, observations are often collected from multiple sources with heterogeneous causal models due to certain uncontrollable factors, which renders causal analysis results obtained by a single model skeptical. In this paper, we generalize the Additive Noise Model (ANM) to a mixture model, which consists of a finite number of ANMs, and provide the condition of its causal identifiability. To conduct model estimation, we propose Gaussian Process Partially Observable Model (GPPOM), and incorporate independence enforcement into it to learn latent parameter associated with each observation. Causal inference and clustering according to the underlying generating mechanisms of the mixture model are addressed in this work. Experiments on synthetic and real data demonstrate the effectiveness of our proposed approach. Shoubo Hu, Zhitang Chen, Vahid Partovi Nia, Lai-Wan Chan, Yanhui Geng |
NeurIPS | 3 |
| 2016 | Statistical Measurement Validation with Application to Electronic Nose TechnologyabstractAn artificial olfaction called electronic nose (e-nose) relies on an array of gas sensors with the capability of mimicking the human sense of smell. Applying an appropriate pattern recognition on the sensor’s output returns odor concentration and odor classification. Odor concentration plays a key role in analyzing odors. Assuring the validity of measurements in each stage of sampling is a critical issue in sampling odors. An accurate prediction for odor concentration demands for careful monitoring of the gas sensor array measurements through time. The existing e-noses capture all odor changes in its environment with possibly varying range of error. Consequently, some measurements may distort the pattern recognition results. We explore e-nose data and provide a statistical algorithm to assess the data validity. Our online algorithm is computationally efficient and treats data as being sampled. Mina Mirshahi, Vahid Partovi Nia, Luc Adjengue |
ICPRAM | 2 |
| 2016 | Flight deck crew reserve: From data to forecasting
Amir-Hosein Homaie-Shandizi, Vahid Partovi Nia, Michel Gamache, Bruno Agard |
Eng. Appl. Artif. Intell. | 2 |