EDBT 2026 Demo / reviewers in the wild / expert
Soroush Abbasi Koohpayegani
dblp:277/5486
· DBLP profile ↗
16ranked-venue papers
4as first author
15since 2021 · last 2025
0000-0001-8023-1519ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MCNC: Manifold-Constrained Reparameterization for Neural CompressionabstractThe outstanding performance of large foundational models across diverse tasks,
from computer vision to speech and natural language processing, has significantly
increased their demand. However, storing and transmitting these models poses
significant challenges due to their massive size (e.g., 750GB for Llama 3.1 405B).
Recent literature has focused on compressing the original weights or reducing the
number of parameters required for fine-tuning these models. These compression
methods generally constrain the parameter space, for example, through low-rank
reparametrization (e.g., LoRA), pruning, or quantization (e.g., QLoRA) during
or after the model training. In this paper, we present a novel model compres-
sion method, which we term Manifold-Constrained Neural Compression (MCNC).
This method constrains the parameter space to low-dimensional pre-defined and
frozen nonlinear manifolds, which effectively cover this space. Given the preva-
lence of good solutions in over-parameterized deep neural networks, we show that
by constraining the parameter space to our proposed manifold, we can identify
high-quality solutions while achieving unprecedented compression rates across
a wide variety of tasks and architectures. Through extensive experiments in
computer vision and natural language processing tasks, we demonstrate that our
method significantly outperforms state-of-the-art baselines in terms of compres-
sion, accuracy, and/or model reconstruction time. Our code is publicly available at
https://github.com/mint-vu/MCNC. Chayne Thrash, Reed Andreas, Ali Abbasi 0008, Parsa Nooralinejad, Soroush Abbasi Koohpayegani, Hamed Pirsiavash, Soheil Kolouri |
ICLR | 5 |
| 2024 | SlowFormer: Adversarial Attack on Compute and Energy Consumption of Efficient Vision TransformersabstractRecently, there has been a lot of progress in reducing the computation of deep models at inference time. These methods can reduce both the computational needs and power usage of deep models. Some of these approaches adaptively scale the compute based on the input instance. We show that such models can be vulnerable to a universal adversarial patch attack, where the attacker optimizes for a patch that when pasted on any image, can increase the compute and power consumption of the model. We run experiments with three different efficient vision transformer methods showing that in some cases, the attacker can increase the computation to the maximum possible level by simply pasting a patch that occupies only 8% of the image area. We also show that a standard adversarial training defense method can reduce some of the attack's success. We believe adaptive efficient methods will be necessary in the future to lower the power usage of expensive deep models, so we hope our paper encourages the community to study the robustness of these methods and develop better defense methods for the proposed attack. Code is available at: https://github.com/UCDvision/SlowFormer Navaneet K. L., Soroush Abbasi Koohpayegani, Essam Sleiman, Hamed Pirsiavash |
CVPR | 2 |
| 2024 | CompGS: Smaller and Faster Gaussian Splatting with Vector Quantization
Navaneet K. L., Kossar Pourahmadi, Soroush Abbasi Koohpayegani, Hamed Pirsiavash |
ECCV (32) | 3 |
| 2024 | NOLA: Compressing LoRA using Linear Combination of Random BasisabstractFine-tuning Large Language Models (LLMs) and storing them for each downstream task or domain is impractical because of the massive model size (e.g., 350GB in GPT-3).
Current literature, such as LoRA, showcases the potential of low-rank modifications to the original weights of an LLM, enabling efficient adaptation and storage for task-specific models. These methods can reduce the number of parameters needed to fine-tune an LLM by several orders of magnitude. Yet, these methods face two primary limitations: (1) the parameter count is lower-bounded by the rank one decomposition, and (2) the extent of reduction is heavily influenced by both the model architecture and the chosen rank. We introduce NOLA, which overcomes the rank one lower bound present in LoRA. It achieves this by re-parameterizing the low-rank matrices in LoRA using linear combinations of randomly generated matrices (basis) and optimizing the linear mixture coefficients only. This approach allows us to decouple the number of trainable parameters from both the choice of rank and the network architecture. We present adaptation results using GPT-2, LLaMA-2, and ViT in natural language and computer vision tasks. NOLA performs as well as LoRA models with much fewer number of parameters compared to LoRA with rank one, the best compression LoRA can archive. Particularly, on LLaMA-2 70B, our method is almost 20 times more compact than the most compressed LoRA without degradation in accuracy. Our code is available here: https://github.com/UCDvision/NOLA Soroush Abbasi Koohpayegani, Navaneet K. L., Parsa Nooralinejad, Soheil Kolouri, Hamed Pirsiavash |
ICLR | 1 |
| 2024 | SimA: Simple Softmax-free Attention for Vision TransformersabstractRecently, vision transformers have become very popular. However, deploying them in many applications is computationally expensive partly due to the Softmax layer in the attention block. We introduce a simple yet effective, Softmaxfree attention block, SimA, which normalizes query and key matrices with simple ℓ1-norm instead of using Softmax layer. Then, the attention block in SimA is a simple multiplication of three matrices, so SimA can dynamically change the ordering of the computation at the test time to achieve linear computation on the number of tokens or the number of channels. We empirically show that SimA applied to three SOTA variations of transformers, DeiT, XCiT, and CvT, results in on-par accuracy compared to the SOTA models, without any need for Softmax layer. Interestingly, changing SimA from multi-head to single-head has only a small effect on the accuracy, which further simplifies the attention block. Moreover, we show that SimA is much faster on small edge devices, e.g., Raspberry Pi, which we believe is due to higher complexity of Softmax layer on those devices. The code is available here: https://github.com/UCDvision/sima Soroush Abbasi Koohpayegani, Hamed Pirsiavash |
WACV | 1 |
| 2024 | A Closer Look at Robustness of Vision Transformers to Backdoor AttacksabstractTransformer architectures are based on self-attention mechanism that processes images as a sequence of patches. As their design is quite different compared to CNNs, it is important to take a closer look at their vulnerability to back-door attacks and how different transformer architectures affect robustness. Backdoor attacks happen when an attacker poisons a small part of the training images with a specific trigger or backdoor which will be activated later. The model performance is good on clean test images, but the attacker can manipulate the decision of the model by showing the trigger on an image at test time. In this paper, we compare state-of-the-art architectures through the lens of backdoor attacks, specifically how attention mechanisms affect robustness. We observe that the well known vision transformer architecture (ViT) is the least robust architecture and ResMLP, which belongs to a class called Feed Forward Networks (FFN), is most robust to backdoor attacks among state-of-the-art architectures. We also find an intriguing difference between transformers and CNNs - interpretation algorithms effectively highlight the trigger on test images for transformers but not for CNNs. Based on this observation, we find that a test-time image blocking defense reduces the attack success rate by a large margin for transformers. We also show that such blocking mechanisms can be incorporated during the training process to improve robustness even further. We believe our experimental findings will encourage the community to understand the building block components in developing novel architectures robust to back-door attacks. Code is available here: https://github.com/UCDvision/backdoor_transformer.git Akshayvarun Subramanya, Soroush Abbasi Koohpayegani, Aniruddha Saha, Ajinkya Tejankar, Hamed Pirsiavash |
WACV | 2 |
| 2023 | PRANC: Pseudo RAndom Networks for Compacting deep modelsabstractWe demonstrate that a deep model can be reparametrized as a linear combination of several randomly initialized and frozen deep models in the weight space. During training, we seek local minima that reside within the subspace spanned by these random models (i.e., ‘basis’ networks). Our framework, PRANC, enables significant compaction of a deep model. The model can be reconstructed using a single scalar ‘seed,’ employed to generate the pseudo-random ‘basis’ networks, together with the learned linear mixture coefficients. In practical applications, PRANC addresses the challenge of efficiently storing and communicating deep models, a common bottleneck in several scenarios, including multi-agent learning, continual learners, federated systems, and edge devices, among others. In this study, we employ PRANC to condense image classification models and compress images by compacting their associated implicit neural networks. PRANC outperforms baselines with a large margin on image classification when compressing a deep model almost 100 times. Moreover, we show that PRANC enables memory-efficient inference by generating layer-wise weights on the fly. The source code of PRANC is here: https://github.com/UCDvision/PRANC Parsa Nooralinejad, Ali Abbasi 0008, Soroush Abbasi Koohpayegani, Kossar Pourahmadi, Rana Muhammad Shahroz Khan, Soheil Kolouri, Hamed Pirsiavash |
ICCV | 3 |
| 2023 | Multi-Agent Lifelong Implicit Neural LearningabstractImplicit neural representations (INRs) have emerged as powerful tools for the continuous representation of signals, finding applications in imaging, computer graphics, and signal compression. Additionally, decentralized multi-agent systems are crucial in various applications, frequently leading to enhanced reliability and efficiencies in computation and communication. In this paper, we explore using multi-agent Lifelong Learning (LL) systems for learning INRs. We propose a rigorous problem setup and evaluation plan to investigate the efficacy of such systems compared to single-agent and multi-task learning baselines. Our research, conducted across varied dimensions, demonstrates promising results, thereby contributing a novel perspective to the realm of continual learning. Soheil Kolouri, Ali Abbasi 0008, Soroush Abbasi Koohpayegani, Parsa Nooralinejad, Hamed Pirsiavash |
IEEE Signal Process. Lett. | 3 |
| 2022 | Consistent Explanations by Contrastive LearningabstractPost-hoc explanation methods, e.g., Grad-CAM, enable humans to inspect the spatial regions responsible for a particular network decision. However, it is shown that such explanations are not always consistent with human priors, such as consistency across image transformations. Given an interpretation algorithm, e.g., Grad-CAM, we introduce a novel training method to train the model to produce more consistent explanations. Since obtaining the ground truth for a desired model interpretation is not a well-defined task, we adopt ideas from contrastive self-supervised learning, and apply them to the interpretations of the model rather than its embeddings. We show that our method, Contrastive Grad-CAM Consistency (CGC), results in Grad-CAM interpretation heatmaps that are more consistent with human annotations while still achieving comparable classification accuracy. Moreover, our method acts as a regularizer and improves the accuracy on limited-data, fine-grained classification settings. In addition, because our method does not rely on annotations, it allows for the incorporation of unlabeled data into training, which enables better generalization of the model. The code is available here: https://github.com/UCDvision/CGC Vipin Pillai, Soroush Abbasi Koohpayegani, Ashley Ouligian, Dennis Fong, Hamed Pirsiavash |
CVPR | 2 |
| 2022 | Backdoor Attacks on Self-Supervised LearningabstractLarge-scale unlabeled data has spurred recent progress in self-supervised learning methods that learn rich vi-sual representations. State-of-the-art self-supervised methods for learning representations from images (e.g., MoCo, BYOL, MSF) use an inductive bias that random augmentations (e.g., random crops) of an image should produce similar embeddings. We show that such methods are vulnerable to backdoor attacks - where an attacker poisons a small part of the unlabeled data by adding a trigger (image patch chosen by the attacker) to the images. The model performance is good on clean test images, but the attacker can manipulate the decision of the model by showing the trigger at test time. Backdoor attacks have been studied extensively in supervised learning and to the best of our knowledge, we are the first to study them for self-supervised learning. Backdoor attacks are more practical in self-supervised learning, since the use of large unlabeled data makes data inspection to remove poisons prohibitive. We show that in our targeted attack, the attacker can produce many false positives for the target category by using the trigger at test time. We also propose a defense method based on knowledge distillation that succeeds in neutralizing the attack. Our code is available here: https://github.com/UMBCvisionISSL-Backdoor Aniruddha Saha, Ajinkya Tejankar, Soroush Abbasi Koohpayegani, Hamed Pirsiavash |
CVPR | 3 |
| 2022 | Adaptive Token Sampling for Efficient Vision Transformers
Mohsen Fayyaz, Soroush Abbasi Koohpayegani, Farnoush Rezaei Jafari, Sunando Sengupta, Hamid Reza Vaezi Joze, Eric Sommerlade, Hamed Pirsiavash, Juergen Gall |
ECCV (11) | 2 |
| 2022 | Constrained Mean Shift Using Distant yet Related Neighbors for Representation Learning
Navaneet K. L., Soroush Abbasi Koohpayegani, Ajinkya Tejankar, Kossar Pourahmadi, Akshayvarun Subramanya, Hamed Pirsiavash |
ECCV (31) | 2 |
| 2021 | Regression as a Simple Yet Effective Tool for Self-supervised Knowledge Distillation
Navaneet K. L., Soroush Abbasi Koohpayegani, Ajinkya Tejankar, Hamed Pirsiavash |
BMVC | 2 |
| 2021 | Mean Shift for Self-Supervised LearningabstractMost recent self-supervised learning (SSL) algorithms learn features by contrasting between instances of images or by clustering the images and then contrasting between the image clusters. We introduce a simple mean-shift algorithm that learns representations by grouping images together without contrasting between them or adopting much of prior on the structure or number of the clusters. We simply "shift" the embedding of each image to be close to the "mean" of the neighbors of its augmentation. Since the closest neighbor is always another augmentation of the same image, our model will be identical to BYOL when using only one nearest neighbor instead of 5 used in our experiments. Our model achieves 72.4% on ImageNet linear evaluation with ResNet50 at 200 epochs outperforming BYOL. Also, our method outperforms the SOTA by a large margin when using weak augmentations only, facilitating adoption of SSL for other modalities. Our code is available here: https://github.com/UMBCvision/MSF Soroush Abbasi Koohpayegani, Ajinkya Tejankar, Hamed Pirsiavash |
ICCV | 1 |
| 2021 | ISD: Self-Supervised Learning by Iterative Similarity DistillationabstractRecently, contrastive learning has achieved great results in self-supervised learning, where the main idea is to pull two augmentations of an image (positive pairs) closer compared to other random images (negative pairs). We argue that not all negative images are equally negative. Hence, we introduce a self-supervised learning algorithm where we use a soft similarity for the negative images rather than a binary distinction between positive and negative pairs. We iteratively distill a slowly evolving teacher model to the student model by capturing the similarity of a query image to some random images and transferring that knowledge to the student. Specifically, our method should handle unbalanced and unlabeled data better than existing contrastive learning methods, because the randomly chosen negative set might include many samples that are semantically similar to the query image. In this case, our method labels them as highly similar while standard contrastive methods label them as negatives. Our method achieves comparable results to the state-of-the-art models. Our code is available here: https://github.com/UMBCvision/ISD. Ajinkya Tejankar, Soroush Abbasi Koohpayegani, Vipin Pillai, Paolo Favaro, Hamed Pirsiavash |
ICCV | 2 |
| 2020 | CompRess: Self-Supervised Learning by Compressing RepresentationsabstractSelf-supervised learning aims to learn good representations with unlabeled data. Recent works have shown that larger models benefit more from self-supervised learning than smaller models. As a result, the gap between supervised and self-supervised learning has been greatly reduced for larger models. In this work, instead of designing a new pseudo task for self-supervised learning, we develop a model compression method to compress an already learned, deep self-supervised model (teacher) to a smaller one (student). We train the student model so that it mimics the relative similarity between the datapoints in the teacher's embedding space. For AlexNet, our method outperforms all previous methods including the fully supervised model on ImageNet linear evaluation (59.0% compared to 56.5%) and on nearest neighbor evaluation (50.7% compared to 41.4%). To the best of our knowledge, this is the first time a self-supervised AlexNet has outperformed supervised one on ImageNet classification. Our code is available here: https://github.com/UMBCvision/CompRess Soroush Abbasi Koohpayegani, Ajinkya Tejankar, Hamed Pirsiavash |
NeurIPS | 1 |