VLDB 2026 Research / reviewers in the wild / expert
Hamed Pirsiavash
dblp:07/6340
· DBLP profile ↗
59ranked-venue papers
8as first author
22since 2021 · last 2025
0000-0002-2528-3305ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 45 · 6 first-author · 18 since 2021Artificial intelligence and machine learning · 44 · 8 first-author · 15 since 2021Systems, architecture and hardware · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MCNC: Manifold-Constrained Reparameterization for Neural CompressionabstractThe outstanding performance of large foundational models across diverse tasks,
from computer vision to speech and natural language processing, has significantly
increased their demand. However, storing and transmitting these models poses
significant challenges due to their massive size (e.g., 750GB for Llama 3.1 405B).
Recent literature has focused on compressing the original weights or reducing the
number of parameters required for fine-tuning these models. These compression
methods generally constrain the parameter space, for example, through low-rank
reparametrization (e.g., LoRA), pruning, or quantization (e.g., QLoRA) during
or after the model training. In this paper, we present a novel model compres-
sion method, which we term Manifold-Constrained Neural Compression (MCNC).
This method constrains the parameter space to low-dimensional pre-defined and
frozen nonlinear manifolds, which effectively cover this space. Given the preva-
lence of good solutions in over-parameterized deep neural networks, we show that
by constraining the parameter space to our proposed manifold, we can identify
high-quality solutions while achieving unprecedented compression rates across
a wide variety of tasks and architectures. Through extensive experiments in
computer vision and natural language processing tasks, we demonstrate that our
method significantly outperforms state-of-the-art baselines in terms of compres-
sion, accuracy, and/or model reconstruction time. Our code is publicly available at
https://github.com/mint-vu/MCNC. Chayne Thrash, Reed Andreas, Ali Abbasi 0008, Parsa Nooralinejad, Soroush Abbasi Koohpayegani, Hamed Pirsiavash, Soheil Kolouri |
ICLR | 6 |
| 2025 | FAMOUS: Fault Attack Mitigation via Exploiting Invariances in Deep Neural NetworksabstractImplementing Deep Neural Networks (DNNs) in hardware is essential due to rising Power-Performance-Area (PPA) demands and the limitations of GPUs in meeting them. However, such accelerators are vulnerable to Fault Injection Attacks (FIAs), such as those induced by laser illumination or Rowhammer. FAMOUS protects against FIAs by exploiting invariances in DNNs—particularly permutation invariance—by dynamically swapping convolutional channels and linear layer connections during runtime. This misleads attackers aiming to corrupt critical weights that significantly impact model output. We evaluate FAMOUS on transformer models (ViT-tiny and ViT-small) across multiple datasets. Even with 100 faults injected into essential weights, accuracy drops are minimal (≈4.7 and 0.02 points on ImageNet-1k), compared to severe drops (59 and 70 points) without protection. CNNs also benefit from FAMOUS, though to a lesser extent. Javad Bahrami, Parsa Nooralinejad, Hamed Pirsiavash, Naghmeh Karimi |
ITC | 3 |
| 2024 | BrainWash: A Poisoning Attack to Forget in Continual LearningabstractContinual learning has gained substantial attention within the deep learning community, offering promising solutions to the challenging problem of sequential learning. Yet, a largely unexplored facet of this paradigm is its susceptibility to adversarial attacks, especially with the aim of inducing forgetting. In this paper, we introduce “Brain-Wash,” a novel data poisoning method tailored to impose forgetting on a continual learner. By adding the Brain-Wash noise to a variety of baselines, we demonstrate how a trained continual learner can be induced to forget its previously learned tasks catastrophically, even when using these continual learning baselines. An important feature of our approach is that the attacker requires no access to previous tasks' data and is armed merely with the model's current parameters and the data belonging to the most recent task. Our extensive experiments highlight the efficacy of Brain Wash, showcasing degradation in performance across various regularization and memory replay-based continual learning methods. Our code is available here: https://github.com/mint-vuIBrainwash Ali Abbasi 0008, Parsa Nooralinejad, Hamed Pirsiavash, Soheil Kolouri |
CVPR | 3 |
| 2024 | SlowFormer: Adversarial Attack on Compute and Energy Consumption of Efficient Vision TransformersabstractRecently, there has been a lot of progress in reducing the computation of deep models at inference time. These methods can reduce both the computational needs and power usage of deep models. Some of these approaches adaptively scale the compute based on the input instance. We show that such models can be vulnerable to a universal adversarial patch attack, where the attacker optimizes for a patch that when pasted on any image, can increase the compute and power consumption of the model. We run experiments with three different efficient vision transformer methods showing that in some cases, the attacker can increase the computation to the maximum possible level by simply pasting a patch that occupies only 8% of the image area. We also show that a standard adversarial training defense method can reduce some of the attack's success. We believe adaptive efficient methods will be necessary in the future to lower the power usage of expensive deep models, so we hope our paper encourages the community to study the robustness of these methods and develop better defense methods for the proposed attack. Code is available at: https://github.com/UCDvision/SlowFormer Navaneet K. L., Soroush Abbasi Koohpayegani, Essam Sleiman, Hamed Pirsiavash |
CVPR | 4 |
| 2024 | CompGS: Smaller and Faster Gaussian Splatting with Vector Quantization
Navaneet K. L., Kossar Pourahmadi, Soroush Abbasi Koohpayegani, Hamed Pirsiavash |
ECCV (32) | 4 |
| 2024 | NOLA: Compressing LoRA using Linear Combination of Random BasisabstractFine-tuning Large Language Models (LLMs) and storing them for each downstream task or domain is impractical because of the massive model size (e.g., 350GB in GPT-3).
Current literature, such as LoRA, showcases the potential of low-rank modifications to the original weights of an LLM, enabling efficient adaptation and storage for task-specific models. These methods can reduce the number of parameters needed to fine-tune an LLM by several orders of magnitude. Yet, these methods face two primary limitations: (1) the parameter count is lower-bounded by the rank one decomposition, and (2) the extent of reduction is heavily influenced by both the model architecture and the chosen rank. We introduce NOLA, which overcomes the rank one lower bound present in LoRA. It achieves this by re-parameterizing the low-rank matrices in LoRA using linear combinations of randomly generated matrices (basis) and optimizing the linear mixture coefficients only. This approach allows us to decouple the number of trainable parameters from both the choice of rank and the network architecture. We present adaptation results using GPT-2, LLaMA-2, and ViT in natural language and computer vision tasks. NOLA performs as well as LoRA models with much fewer number of parameters compared to LoRA with rank one, the best compression LoRA can archive. Particularly, on LLaMA-2 70B, our method is almost 20 times more compact than the most compressed LoRA without degradation in accuracy. Our code is available here: https://github.com/UCDvision/NOLA Soroush Abbasi Koohpayegani, Navaneet K. L., Parsa Nooralinejad, Soheil Kolouri, Hamed Pirsiavash |
ICLR | 5 |
| 2024 | FAT-RABBIT: Fault-Aware Training towards Robustness AgainstBit-flip Based Attacks in Deep Neural NetworksabstractMachine learning and in particular deep learning is used in a broad range of crucial applications. Implementing such models in custom hardware can be highly beneficial thanks to their low power and computation latency compared to GPUs. However, an error in their output can lead to disastrous outcomes. An adversary may force misclassification in the model’s outcome by inducing a number of bit-flips in the targeted locations; thus declining the accuracy. To fill the gap, this paper presents FAT-RABBIT, a cost-effective mechanism designed to mitigate such threats by training the model such that there would be few weights that can be highly impactful in the outcome; thus reducing the sensitivity of the model to the fault injection attacks. Moreover, to increase robustness against bit-wise large perturbations, we propose an optimization scheme so-called M-SAM. We then augment FAT-RABBIT with the M-SAM optimizer to further bolster model accuracy against bit-flipping fault attacks. Notably, these approaches incur no additional hardware overhead. Our experimental results demonstrate the robustness of FAT-RABBIT and its augmented version, called Augmented FAT-RABBIT, against such attacks. Hossein Pourmehrani, Javad Bahrami, Parsa Nooralinejad, Hamed Pirsiavash, Naghmeh Karimi |
ITC | 4 |
| 2024 | SimA: Simple Softmax-free Attention for Vision TransformersabstractRecently, vision transformers have become very popular. However, deploying them in many applications is computationally expensive partly due to the Softmax layer in the attention block. We introduce a simple yet effective, Softmaxfree attention block, SimA, which normalizes query and key matrices with simple ℓ1-norm instead of using Softmax layer. Then, the attention block in SimA is a simple multiplication of three matrices, so SimA can dynamically change the ordering of the computation at the test time to achieve linear computation on the number of tokens or the number of channels. We empirically show that SimA applied to three SOTA variations of transformers, DeiT, XCiT, and CvT, results in on-par accuracy compared to the SOTA models, without any need for Softmax layer. Interestingly, changing SimA from multi-head to single-head has only a small effect on the accuracy, which further simplifies the attention block. Moreover, we show that SimA is much faster on small edge devices, e.g., Raspberry Pi, which we believe is due to higher complexity of Softmax layer on those devices. The code is available here: https://github.com/UCDvision/sima Soroush Abbasi Koohpayegani, Hamed Pirsiavash |
WACV | 2 |
| 2024 | A Closer Look at Robustness of Vision Transformers to Backdoor AttacksabstractTransformer architectures are based on self-attention mechanism that processes images as a sequence of patches. As their design is quite different compared to CNNs, it is important to take a closer look at their vulnerability to back-door attacks and how different transformer architectures affect robustness. Backdoor attacks happen when an attacker poisons a small part of the training images with a specific trigger or backdoor which will be activated later. The model performance is good on clean test images, but the attacker can manipulate the decision of the model by showing the trigger on an image at test time. In this paper, we compare state-of-the-art architectures through the lens of backdoor attacks, specifically how attention mechanisms affect robustness. We observe that the well known vision transformer architecture (ViT) is the least robust architecture and ResMLP, which belongs to a class called Feed Forward Networks (FFN), is most robust to backdoor attacks among state-of-the-art architectures. We also find an intriguing difference between transformers and CNNs - interpretation algorithms effectively highlight the trigger on test images for transformers but not for CNNs. Based on this observation, we find that a test-time image blocking defense reduces the attack success rate by a large margin for transformers. We also show that such blocking mechanisms can be incorporated during the training process to improve robustness even further. We believe our experimental findings will encourage the community to understand the building block components in developing novel architectures robust to back-door attacks. Code is available here: https://github.com/UCDvision/backdoor_transformer.git Akshayvarun Subramanya, Soroush Abbasi Koohpayegani, Aniruddha Saha, Ajinkya Tejankar, Hamed Pirsiavash |
WACV | 5 |
| 2023 | Defending Against Patch-based Backdoor Attacks on Self-Supervised LearningabstractRecently, self-supervised learning (SSL) was shown to be vulnerable to patch-based data poisoning backdoor attacks. It was shown that an adversary can poison a small part of the unlabeled data so that when a victim trains an SSL model on it, the final model will have a back-door that the adversary can exploit. This work aims to defend self-supervised learning against such attacks. We use a three-step defense pipeline, where we first train a model on the poisoned data. In the second step, our proposed defense algorithm (PatchSearch) uses the trained model to search the training data for poisoned samples and removes them from the training set. In the third step, a final model is trained on the cleaned-up training set. Our results show that PatchSearch is an effective defense. As an example, it improves a model's accuracy on images containing the trigger from 38.2% to 63.7% which is very close to the clean model's accuracy, 64.6%. More-over, we show that PatchSearch outperforms baselines and state-of-the-art defense approaches including those using additional clean, trusted data. Our code is available at https://github.com/UCDvision/PatchSearch Ajinkya Tejankar, Maziar Sanjabi, Qifan Wang 0001, Sinong Wang, Hamed Firooz, Hamed Pirsiavash, Liang Tan 0005 |
CVPR | 6 |
| 2023 | Is Multi-Task Learning an Upper Bound for Continual Learning?abstractContinual learning and multi-task learning are commonly used machine learning techniques for learning from multiple tasks. However, existing literature assumes multi-task learning as a reasonable performance upper bound for various continual learning algorithms, without rigorous justification. Additionally, in a multi-task setting, a small subset of tasks may behave as adversarial tasks, negatively impacting overall learning performance. On the other hand, continual learning approaches can avoid the negative impact of adversarial tasks and maintain performance on the remaining tasks, resulting in better performance than multi-task learning. This paper introduces a novel continual self-supervised learning approach, where each task involves learning an invariant representation for a specific class of data augmentations. We demonstrate that this approach results in naturally contradicting tasks and that, in this setting, continual learning often outperforms multi-task learning on benchmark datasets, including MNIST, CIFAR-10, and CIFAR-100. Zihao Wu 0003, Hamed Pirsiavash, Soheil Kolouri |
ICASSP | 3 |
| 2023 | PRANC: Pseudo RAndom Networks for Compacting deep modelsabstractWe demonstrate that a deep model can be reparametrized as a linear combination of several randomly initialized and frozen deep models in the weight space. During training, we seek local minima that reside within the subspace spanned by these random models (i.e., ‘basis’ networks). Our framework, PRANC, enables significant compaction of a deep model. The model can be reconstructed using a single scalar ‘seed,’ employed to generate the pseudo-random ‘basis’ networks, together with the learned linear mixture coefficients. In practical applications, PRANC addresses the challenge of efficiently storing and communicating deep models, a common bottleneck in several scenarios, including multi-agent learning, continual learners, federated systems, and edge devices, among others. In this study, we employ PRANC to condense image classification models and compress images by compacting their associated implicit neural networks. PRANC outperforms baselines with a large margin on image classification when compressing a deep model almost 100 times. Moreover, we show that PRANC enables memory-efficient inference by generating layer-wise weights on the fly. The source code of PRANC is here: https://github.com/UCDvision/PRANC Parsa Nooralinejad, Ali Abbasi 0008, Soroush Abbasi Koohpayegani, Kossar Pourahmadi, Rana Muhammad Shahroz Khan, Soheil Kolouri, Hamed Pirsiavash |
ICCV | 7 |
| 2023 | MASTAF: A Model-Agnostic Spatio-Temporal Attention Fusion Network for Few-shot Video ClassificationabstractWe propose MASTAF, a Model-Agnostic Spatio-Temporal Attention Fusion network for few-shot video classification. MASTAF takes input from a general video spatial and temporal representation,e.g., using 2D CNN, 3D CNN, and Video Transformer. Then, to make the most of such representations, we use self- and cross-attention models to highlight the critical spatio-temporal region to increase the inter-class variations and decrease the intra-class variations. Last, MASTAF applies a lightweight fusion network and a nearest neighbor classifier to classify each query video. We demonstrate that MASTAF improves the state-of-the-art performance on three few-shot video classification benchmarks(UCF101, HMDB51, and Something-Something-V2), e.g., by up to 91.6%, 69.5%, and 60.7% for five-way one-shot video classification, respectively. Huanle Zhang, Hamed Pirsiavash, Xin Liu 0002 |
WACV | 2 |
| 2023 | Multi-Agent Lifelong Implicit Neural LearningabstractImplicit neural representations (INRs) have emerged as powerful tools for the continuous representation of signals, finding applications in imaging, computer graphics, and signal compression. Additionally, decentralized multi-agent systems are crucial in various applications, frequently leading to enhanced reliability and efficiencies in computation and communication. In this paper, we explore using multi-agent Lifelong Learning (LL) systems for learning INRs. We propose a rigorous problem setup and evaluation plan to investigate the efficacy of such systems compared to single-agent and multi-task learning baselines. Our research, conducted across varied dimensions, demonstrates promising results, thereby contributing a novel perspective to the realm of continual learning. Soheil Kolouri, Ali Abbasi 0008, Soroush Abbasi Koohpayegani, Parsa Nooralinejad, Hamed Pirsiavash |
IEEE Signal Process. Lett. | 5 |
| 2022 | Consistent Explanations by Contrastive LearningabstractPost-hoc explanation methods, e.g., Grad-CAM, enable humans to inspect the spatial regions responsible for a particular network decision. However, it is shown that such explanations are not always consistent with human priors, such as consistency across image transformations. Given an interpretation algorithm, e.g., Grad-CAM, we introduce a novel training method to train the model to produce more consistent explanations. Since obtaining the ground truth for a desired model interpretation is not a well-defined task, we adopt ideas from contrastive self-supervised learning, and apply them to the interpretations of the model rather than its embeddings. We show that our method, Contrastive Grad-CAM Consistency (CGC), results in Grad-CAM interpretation heatmaps that are more consistent with human annotations while still achieving comparable classification accuracy. Moreover, our method acts as a regularizer and improves the accuracy on limited-data, fine-grained classification settings. In addition, because our method does not rely on annotations, it allows for the incorporation of unlabeled data into training, which enables better generalization of the model. The code is available here: https://github.com/UCDvision/CGC Vipin Pillai, Soroush Abbasi Koohpayegani, Ashley Ouligian, Dennis Fong, Hamed Pirsiavash |
CVPR | 5 |
| 2022 | Backdoor Attacks on Self-Supervised LearningabstractLarge-scale unlabeled data has spurred recent progress in self-supervised learning methods that learn rich vi-sual representations. State-of-the-art self-supervised methods for learning representations from images (e.g., MoCo, BYOL, MSF) use an inductive bias that random augmentations (e.g., random crops) of an image should produce similar embeddings. We show that such methods are vulnerable to backdoor attacks - where an attacker poisons a small part of the unlabeled data by adding a trigger (image patch chosen by the attacker) to the images. The model performance is good on clean test images, but the attacker can manipulate the decision of the model by showing the trigger at test time. Backdoor attacks have been studied extensively in supervised learning and to the best of our knowledge, we are the first to study them for self-supervised learning. Backdoor attacks are more practical in self-supervised learning, since the use of large unlabeled data makes data inspection to remove poisons prohibitive. We show that in our targeted attack, the attacker can produce many false positives for the target category by using the trigger at test time. We also propose a defense method based on knowledge distillation that succeeds in neutralizing the attack. Our code is available here: https://github.com/UMBCvisionISSL-Backdoor Aniruddha Saha, Ajinkya Tejankar, Soroush Abbasi Koohpayegani, Hamed Pirsiavash |
CVPR | 4 |
| 2022 | Adaptive Token Sampling for Efficient Vision Transformers
Mohsen Fayyaz, Soroush Abbasi Koohpayegani, Farnoush Rezaei Jafari, Sunando Sengupta, Hamid Reza Vaezi Joze, Eric Sommerlade, Hamed Pirsiavash, Juergen Gall |
ECCV (11) | 7 |
| 2022 | Constrained Mean Shift Using Distant yet Related Neighbors for Representation Learning
Navaneet K. L., Soroush Abbasi Koohpayegani, Ajinkya Tejankar, Kossar Pourahmadi, Akshayvarun Subramanya, Hamed Pirsiavash |
ECCV (31) | 6 |
| 2021 | Explainable Models with Consistent InterpretationsabstractGiven the widespread deployment of black box deep neural networks in computer vision applications, the interpretability aspect of these black box systems has recently gained traction. Various methods have been proposed to explain the results of such deep neural networks. However, some recent works have shown that such explanation methods are biased and do not produce consistent interpretations. Hence, rather than introducing a novel explanation method, we learn models that are encouraged to be interpretable given an explanation method. We use Grad-CAM as the explanation algorithm and encourage the network to learn consistent interpretations along with maximizing the log-likelihood of the correct class. We show that our method outperforms the baseline on the pointing game evaluation on ImageNet and MS-COCO datasets respectively. We also introduce new evaluation metrics that penalize the saliency map if it lies outside the ground truth bounding box or segmentation mask, and show that our method outperforms the baseline on these metrics as well. Moreover, our model trained with interpretation consistency generalizes to other explanation algorithms on all the evaluation metrics. The code and models are publicly available. Vipin Pillai, Hamed Pirsiavash |
AAAI | 2 |
| 2021 | Regression as a Simple Yet Effective Tool for Self-supervised Knowledge Distillation
Navaneet K. L., Soroush Abbasi Koohpayegani, Ajinkya Tejankar, Hamed Pirsiavash |
BMVC | 4 |
| 2021 | Mean Shift for Self-Supervised LearningabstractMost recent self-supervised learning (SSL) algorithms learn features by contrasting between instances of images or by clustering the images and then contrasting between the image clusters. We introduce a simple mean-shift algorithm that learns representations by grouping images together without contrasting between them or adopting much of prior on the structure or number of the clusters. We simply "shift" the embedding of each image to be close to the "mean" of the neighbors of its augmentation. Since the closest neighbor is always another augmentation of the same image, our model will be identical to BYOL when using only one nearest neighbor instead of 5 used in our experiments. Our model achieves 72.4% on ImageNet linear evaluation with ResNet50 at 200 epochs outperforming BYOL. Also, our method outperforms the SOTA by a large margin when using weak augmentations only, facilitating adoption of SSL for other modalities. Our code is available here: https://github.com/UMBCvision/MSF Soroush Abbasi Koohpayegani, Ajinkya Tejankar, Hamed Pirsiavash |
ICCV | 3 |
| 2021 | ISD: Self-Supervised Learning by Iterative Similarity DistillationabstractRecently, contrastive learning has achieved great results in self-supervised learning, where the main idea is to pull two augmentations of an image (positive pairs) closer compared to other random images (negative pairs). We argue that not all negative images are equally negative. Hence, we introduce a self-supervised learning algorithm where we use a soft similarity for the negative images rather than a binary distinction between positive and negative pairs. We iteratively distill a slowly evolving teacher model to the student model by capturing the similarity of a query image to some random images and transferring that knowledge to the student. Specifically, our method should handle unbalanced and unlabeled data better than existing contrastive learning methods, because the randomly chosen negative set might include many samples that are semantically similar to the query image. In this case, our method labels them as highly similar while standard contrastive methods label them as negatives. Our method achieves comparable results to the state-of-the-art models. Our code is available here: https://github.com/UMBCvision/ISD. Ajinkya Tejankar, Soroush Abbasi Koohpayegani, Vipin Pillai, Paolo Favaro, Hamed Pirsiavash |
ICCV | 5 |
| 2020 | Hidden Trigger Backdoor AttacksabstractWith the success of deep learning algorithms in various domains, studying adversarial attacks to secure deep models in real world applications has become an important research topic. Backdoor attacks are a form of adversarial attacks on deep networks where the attacker provides poisoned data to the victim to train the model with, and then activates the attack by showing a specific small trigger pattern at the test time. Most state-of-the-art backdoor attacks either provide mislabeled poisoning data that is possible to identify by visual inspection, reveal the trigger in the poisoned data, or use noise to hide the trigger. We propose a novel form of backdoor attack where poisoned data look natural with correct labels and also more importantly, the attacker hides the trigger in the poisoned data and keeps the trigger secret until the test time. We perform an extensive study on various image classification settings and show that our attack can fool the model by pasting the trigger at random locations on unseen images although the model performs well on clean data. We also show that our proposed attack cannot be easily defended using a state-of-the-art defense algorithm for backdoor attacks. Aniruddha Saha, Akshayvarun Subramanya, Hamed Pirsiavash |
AAAI | 3 |
| 2020 | Universal Litmus Patterns: Revealing Backdoor Attacks in CNNsabstractThe unprecedented success of deep neural networks in many applications has made these networks a prime target for adversarial exploitation. In this paper, we introduce a benchmark technique for detecting backdoor attacks (aka Trojan attacks) on deep convolutional neural networks (CNNs). We introduce the concept of Universal Litmus Patterns (ULPs), which enable one to reveal backdoor attacks by feeding these universal patterns to the network and analyzing the output (i.e., classifying the network as `clean' or `corrupted'). This detection is fast because it requires only a few forward passes through a CNN. We demonstrate the effectiveness of ULPs for detecting backdoor attacks on thousands of networks with different architectures trained on four benchmark datasets, namely the German Traffic Sign Recognition Benchmark (GTSRB), MNIST, CIFAR10, and Tiny-ImageNet. The codes and train/test models for this paper can be found here: https://umbcvision.github.io/Universal-Litmus-Patterns/. Soheil Kolouri, Aniruddha Saha, Hamed Pirsiavash, Heiko Hoffmann |
CVPR | 3 |
| 2020 | On-Chip Voltage and Temperature Digital Sensor for Security, Reliability, and PortabilityabstractThe integrated circuits can be exposed to various stresses during run-time due to unexpected environmental conditions or attacks. Ensuring that a circuit is not working out-of-specification via sensing its operating conditions, e.g., temperature and voltage, is highly useful in detecting anomalies. Analog sensors have been used to monitor the operating conditions for a long time, however, weaknesses including lack of portability to thin technology nodes, costly & complex calibration process, and low attack resistance make such sensors inefficient. Digital sensors, via considering the temperature and voltage effects altogether instead of treating each separately, have been demonstrated as a qualified replacement. In this paper, we develop an integrated framework for continuous monitoring of the operating voltage and temperature of each chip. The framework includes an embedded on-chip sensor circuitry along with a Neural Network model that quantifies the temperature and voltage values via processing the data collected by this sensor. The experimental results confirm the high accuracy of the proposed framework in tracking on-chip voltage and temperature variations, i.e., with the average error of 0.014V in a range of 0.65V to 1.4V, and the average error of 3.9°C in a range of -10°C to 150°C, respectively. Md Toufiq Hasan Anik, Mohammad Ebrahimabadi, Hamed Pirsiavash, Jean-Luc Danger, Sylvain Guilley, Naghmeh Karimi |
ICCD | 3 |
| 2020 | COOT: Cooperative Hierarchical Transformer for Video-Text Representation LearningabstractMany real-world video-text tasks involve different levels of granularity, such as frames and words, clip and sentences or videos and paragraphs, each with distinct semantics. In this paper, we propose a Cooperative hierarchical Transformer (COOT) to leverage this hierarchy information and model the interactions between different levels of granularity and different modalities. The method consists of three major components: an attention-aware feature aggregation layer, which leverages the local temporal context (intra-level, e.g., within a clip), a contextual transformer to learn the interactions between low-level and high-level semantics (inter-level, e.g. clip-video, sentence-paragraph), and a cross-modal cycle-consistency loss to connect video and text. The resulting method compares favorably to the state of the art on several benchmarks while having few parameters. Simon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash, Thomas Brox |
NeurIPS | 3 |
| 2020 | CompRess: Self-Supervised Learning by Compressing RepresentationsabstractSelf-supervised learning aims to learn good representations with unlabeled data. Recent works have shown that larger models benefit more from self-supervised learning than smaller models. As a result, the gap between supervised and self-supervised learning has been greatly reduced for larger models. In this work, instead of designing a new pseudo task for self-supervised learning, we develop a model compression method to compress an already learned, deep self-supervised model (teacher) to a smaller one (student). We train the student model so that it mimics the relative similarity between the datapoints in the teacher's embedding space. For AlexNet, our method outperforms all previous methods including the fully supervised model on ImageNet linear evaluation (59.0% compared to 56.5%) and on nearest neighbor evaluation (50.7% compared to 41.4%). To the best of our knowledge, this is the first time a self-supervised AlexNet has outperformed supervised one on ImageNet classification. Our code is available here: https://github.com/UMBCvision/CompRess Soroush Abbasi Koohpayegani, Ajinkya Tejankar, Hamed Pirsiavash |
NeurIPS | 3 |
| 2019 | Fooling Network Interpretation in Image ClassificationabstractDeep neural networks have been shown to be fooled rather easily using adversarial attack algorithms. Practical methods such as adversarial patches have been shown to be extremely effective in causing misclassification. However, these patches are highlighted using standard network interpretation algorithms, thus revealing the identity of the adversary. We show that it is possible to create adversarial patches which not only fool the prediction, but also change what we interpret regarding the cause of the prediction. Moreover, we introduce our attack as a controlled setting to measure the accuracy of interpretation algorithms. We show this using extensive experiments for Grad-CAM interpretation that transfers to occluding patch interpretation as well. We believe our algorithms can facilitate developing more robust network interpretation tools that truly explain the network's underlying decision making process. Akshayvarun Subramanya, Vipin Pillai, Hamed Pirsiavash |
ICCV | 3 |
| 2018 | Boosting Self-Supervised Learning via Knowledge TransferabstractIn self-supervised learning, one trains a model to solve a so-called pretext task on a dataset without the need for human annotation. The main objective, however, is to transfer this model to a target domain and task. Currently, the most effective transfer strategy is fine-tuning, which restricts one to use the same model or parts thereof for both pretext and target tasks. In this paper, we present a novel framework for self-supervised learning that overcomes limitations in designing and comparing different tasks, models, and data domains. In particular, our framework decouples the structure of the self-supervised model from the final task-specific fine-tuned model. This allows us to: 1) quantitatively assess previously incompatible models including handcrafted features; 2) show that deeper neural network models can learn better representations from the same pretext task; 3) transfer knowledge learned with a deep model to a shallower one and thus boost its learning. We use this framework to design a novel self-supervised task, which achieves state-of-the-art performance on the common benchmarks in PASCAL VOC 2007, ILSVRC12 and Places by a significant margin. Our learned features shrink the mAP gap between models trained via self-supervised learning and supervised learning from 5.9% to 2.6% in object detection on PASCAL VOC 2007. Mehdi Noroozi, Ananth Vinjimoor, Paolo Favaro, Hamed Pirsiavash |
CVPR | 4 |
| 2018 | Cross-Modal Scene NetworksabstractPeople can recognize scenes across many different modalities beyond natural images. In this paper, we investigate how to learn cross-modal scene representations that transfer across modalities. To study this problem, we introduce a new cross-modal scene dataset. While convolutional neural networks can categorize scenes well, they also learn an intermediate representation not aligned across modalities, which is undesirable for cross-modal transfer applications. We present methods to regularize cross-modal convolutional neural networks so that they have a shared representation that is agnostic of the modality. Our experiments suggest that our scene representation can help transfer representations across modalities for retrieval. Moreover, our visualizations suggest that units emerge in the shared representation that tend to activate on consistent concepts independently of the modality. Yusuf Aytar, Lluís Castrejón, Carl Vondrick, Hamed Pirsiavash, Antonio Torralba 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | Weakly Supervised Cascaded Convolutional NetworksabstractObject detection is a challenging task in visual understanding domain, and even more so if the supervision is to be weak. Recently, few efforts to handle the task without expensive human annotations is established by promising deep neural network. A new architecture of cascaded networks is proposed to learn a convolutional neural network (CNN) under such conditions. We introduce two such architectures, with either two cascade stages or three which are trained in an end-to-end pipeline. The first stage of both architectures extracts best candidate of class specific region proposals by training a fully convolutional network. In the case of the three stage architecture, the middle stage provides object segmentation, using the output of the activation maps of first stage. The final stage of both architectures is a part of a convolutional neural network that performs multiple instance learning on proposals extracted in the previous stage(s). Our experiments on the PASCAL VOC 2007, 2010, 2012 and large scale object datasets, ILSVRC 2013, 2014 datasets show improvements in the areas of weakly-supervised object detection, classification and localization. Ali Diba, Vivek Sharma 0001, Ali Mohammad Pazandeh, Hamed Pirsiavash, Luc Van Gool |
CVPR | 4 |
| 2017 | Representation Learning by Learning to CountabstractWe introduce a novel method for representation learning that uses an artificial supervision signal based on counting visual primitives. This supervision signal is obtained from an equivariance relation, which does not require any manual annotation. We relate transformations of images to transformations of the representations. More specifically, we look for the representation that satisfies such relation rather than the transformations that match a given representation. In this paper, we use two image transformations in the context of counting: scaling and tiling. The first transformation exploits the fact that the number of visual primitives should be invariant to scale. The second transformation allows us to equate the total number of visual primitives in each tile to that in the whole image. These two transformations are combined in one constraint and used to train a neural network with a contrastive loss. The proposed task produces representations that perform on par or exceed the state of the art in transfer learning benchmarks. Mehdi Noroozi, Hamed Pirsiavash, Paolo Favaro |
ICCV | 2 |
| 2017 | Automated Detection of Substance Use-Related Social Media Posts Based on Image and Text AnalysisabstractNowadays, teens and young adults spend a significant amount of time on social media. According to the national survey of American attitudes on substance abuse, American teens who spend time on social media sites are at increased risk of smoking, drinking and illicit drug use. Reducing teens' exposure to substance use-related social media posts may help minimize their risk of future substance use and addiction. In this paper, we present a method for automated detection of substance userelated social media posts. With this technology, substance userelated content can be automatically filtered out from social media. To detect substance use related social media posts, we employ the state-of-the-art social media analytics that combines Neural Network-based image and text processing technologies. Our evaluation results demonstrate that image features derived using Convolutional Neural Network and textual features derived using neural document embedding are effective in identifying substance use-related social media posts. Arpita Roy, Anamika Paul, Hamed Pirsiavash, Shimei Pan |
ICTAI | 3 |
| 2016 | Joint Semantic Segmentation and Depth Estimation with Deep Convolutional NetworksabstractMulti-scale deep CNNs have been used successfully for problems mapping each pixel to a label, such as depth estimation and semantic segmentation. It has also been shown that such architectures are reusable and can be used for multiple tasks. These networks are typically trained independently for each task by varying the output layer(s) and training objective. In this work we present a new model for simultaneous depth estimation and semantic segmentation from a single RGB image. Our approach demonstrates the feasibility of training parts of the model for each task and then fine tuning the full, combined model on both tasks simultaneously using a single loss function. Furthermore we couple the deep CNN with fully connected CRF, which captures the contextual relationships and interactions between the semantic and depth cues improving the accuracy of the final results. The proposed model is trained and evaluated on NYUDepth V2 dataset [23] outperforming the state of the art methods on semantic segmentation and achieving comparable results on the task of depth estimation. Arsalan Mousavian, Hamed Pirsiavash, Jana Kosecka |
3DV | 2 |
| 2016 | Learning Aligned Cross-Modal Representations from Weakly Aligned DataabstractPeople can recognize scenes across many different modalities beyond natural images. In this paper, we investigate how to learn cross-modal scene representations that transfer across modalities. To study this problem, we introduce a new cross-modal scene dataset. While convolutional neural networks can categorize cross-modal scenes well, they also learn an intermediate representation not aligned across modalities, which is undesirable for crossmodal transfer applications. We present methods to regularize cross-modal convolutional neural networks so that they have a shared representation that is agnostic of the modality. Our experiments suggest that our scene representation can help transfer representations across modalities for retrieval. Moreover, our visualizations suggest that units emerge in the shared representation that tend to activate on consistent concepts independently of the modality. Lluís Castrejón, Yusuf Aytar, Carl Vondrick, Hamed Pirsiavash, Antonio Torralba 0001 |
CVPR | 4 |
| 2016 | DeepCAMP: Deep Convolutional Action & Attribute Mid-Level PatternsabstractThe recognition of human actions and the determination of human attributes are two tasks that call for fine-grained classification. Indeed, often rather small and inconspicuous objects and features have to be detected to tell their classes apart. In order to deal with this challenge, we propose a novel convolutional neural network that mines mid-level image patches that are sufficiently dedicated to resolve the corresponding subtleties. In particular, we train a newly designed CNN (DeepPattern) that learns discriminative patch groups. There are two innovative aspects to this. On the one hand we pay attention to contextual information in an original fashion. On the other hand, we let an iteration of feature learning and patch clustering purify the set of dedicated patches that we use. We validate our method for action classification on two challenging datasets: PASCAL VOC 2012 Action and Stanford 40 Actions, and for attribute recognition we use the Berkeley Attributes of People dataset. Our discriminative mid-level mining CNN obtains state-of-the-art results on these datasets, without a need for annotations about parts and poses. Ali Diba, Ali Mohammad Pazandeh, Hamed Pirsiavash, Luc Van Gool |
CVPR | 3 |
| 2016 | Predicting Motivations of Actions by Leveraging TextabstractUnderstanding human actions is a key problem in computer vision. However, recognizing actions is only the first step of understanding what a person is doing. In this paper, we introduce the problem of predicting why a person has performed an action in images. This problem has many applications in human activity understanding, such as anticipating or explaining an action. To study this problem, we introduce a new dataset of people performing actions annotated with likely motivations. However, the information in an image alone may not be sufficient to automatically solve this task. Since humans can rely on their lifetime of experiences to infer motivation, we propose to give computer vision systems access to some of these experiences by using recently developed natural language models to mine knowledge stored in massive amounts of text. While we are still far away from fully understanding motivation, our results suggest that transferring knowledge from language into vision can help machines understand why people in images might be performing an action. Carl Vondrick, Deniz Oktay, Hamed Pirsiavash, Antonio Torralba 0001 |
CVPR | 3 |
| 2016 | Anticipating Visual Representations from Unlabeled VideoabstractAnticipating actions and objects before they start or appear is a difficult problem in computer vision with several real-world applications. This task is challenging partly because it requires leveraging extensive knowledge of the world that is difficult to write down. We believe that a promising resource for efficiently learning this knowledge is through readily available unlabeled video. We present a framework that capitalizes on temporal structure in unlabeled video to learn to anticipate human actions and objects. The key idea behind our approach is that we can train deep networks to predict the visual representation of images in the future. Visual representations are a promising prediction target because they encode images at a higher semantic level than pixels yet are automatic to compute. We then apply recognition algorithms on our predicted representation to anticipate objects and actions. We experimentally validate this idea on two datasets, anticipating actions one second in the future and objects five seconds in the future. Carl Vondrick, Hamed Pirsiavash, Antonio Torralba 0001 |
CVPR | 2 |
| 2016 | Generating Videos with Scene DynamicsabstractWe capitalize on large amounts of unlabeled video in order to learn a model of scene dynamics for both video recognition tasks (e.g. action classification) and video generation tasks (e.g. future prediction). We propose a generative adversarial network for video with a spatio-temporal convolutional architecture that untangles the scene's foreground from the background. Experiments suggest this model can generate tiny videos up to a second at full frame rate better than simple baselines, and we show its utility at predicting plausible futures of static images. Moreover, experiments and visualizations show the model internally learns useful features for recognizing actions with minimal supervision, suggesting scene dynamics are a promising signal for representation learning. We believe generative video models can impact many applications in video understanding and simulation. Carl Vondrick, Hamed Pirsiavash, Antonio Torralba 0001 |
NIPS | 2 |
| 2016 | Visualizing Object Detection Features
Carl Vondrick, Aditya Khosla, Hamed Pirsiavash, Tomasz Malisiewicz, Antonio Torralba 0001 |
Int. J. Comput. Vis. | 3 |
| 2015 | Learning visual biases from human imaginationabstractAlthough the human visual system can recognize many concepts under challengingconditions, it still has some biases. In this paper, we investigate whether wecan extract these biases and transfer them into a machine recognition system.We introduce a novel method that, inspired by well-known tools in humanpsychophysics, estimates the biases that the human visual system might use forrecognition, but in computer vision feature spaces. Our experiments aresurprising, and suggest that classifiers from the human visual system can betransferred into a machine with some success. Since these classifiers seem tocapture favorable biases in the human visual system, we further present an SVMformulation that constrains the orientation of the SVM hyperplane to agree withthe bias from human visual system. Our results suggest that transferring thishuman bias into machines may help object recognition systems generalize acrossdatasets and perform better when very little training data is available. Carl Vondrick, Hamed Pirsiavash, Aude Oliva, Antonio Torralba 0001 |
NIPS | 2 |
| 2014 | Parsing Videos of Actions with Segmental GrammarsabstractReal-world videos of human activities exhibit temporal structure at various scales, long videos are typically composed out of multiple action instances, where each instance is itself composed of sub-actions with variable durations and orderings. Temporal grammars can presumably model such hierarchical structure, but are computationally difficult to apply for long video streams. We describe simple grammars that capture hierarchical temporal structure while admitting inference with a finite-state-machine. This makes parsing linear time, constant storage, and naturally online. We train grammar parameters using a latent structural SVM, where latent subactions are learned automatically. We illustrate the effectiveness of our approach over common baselines on a new half-million frame dataset of continuous YouTube videos. Hamed Pirsiavash, Deva Ramanan |
CVPR | 1 |
| 2014 | Assessing the Quality of Actions
Hamed Pirsiavash, Carl Vondrick, Antonio Torralba 0001 |
ECCV (6) | 1 |
| 2013 | Parsing IKEA Objects: Fine Pose EstimationabstractWe address the problem of localizing and estimating the fine-pose of objects in the image with exact 3D models. Our main focus is to unify contributions from the 1970s with recent advances in object detection: use local keypoint detectors to find candidate poses and score global alignment of each candidate pose to the image. Moreover, we also provide a new dataset containing fine-aligned objects with their exactly matched 3D models, and a set of models for widely used objects. We also evaluate our algorithm both on object detection and fine pose estimation, and show that our method outperforms state-of-the art algorithms. Joseph J. Lim, Hamed Pirsiavash, Antonio Torralba 0001 |
ICCV | 2 |
| 2012 | Detecting activities of daily living in first-person camera viewsabstractWe present a novel dataset and novel algorithms for the problem of detecting activities of daily living (ADL) in firstperson camera views. We have collected a dataset of 1 million frames of dozens of people performing unscripted, everyday activities. The dataset is annotated with activities, object tracks, hand positions, and interaction events. ADLs differ from typical actions in that they can involve long-scale temporal structure (making tea can take a few minutes) and complex object interactions (a fridge looks different when its door is open). We develop novel representations including (1) temporal pyramids, which generalize the well-known spatial pyramid to approximate temporal correspondence when scoring a model and (2) composite object models that exploit the fact that objects look different when being interacted with. We perform an extensive empirical evaluation and demonstrate that our novel representations produce a two-fold improvement over traditional approaches. Our analysis suggests that real-world ADL recognition is “all about the objects,” and in particular, “all about the objects being interacted with.” Hamed Pirsiavash, Deva Ramanan |
CVPR | 1 |
| 2012 | Steerable part modelsabstractWe describe a method for learning steerable deformable part models. Our models exploit the fact that part templates can be written as linear filter banks. We demonstrate that one can enforce steerability and separability during learning by applying rank constraints. These constraints are enforced with a coordinate descent learning algorithm, where each step can be solved with an off-the-shelf structured SVM solver. The resulting models are orders of magnitude smaller than their counterparts, greatly simplifying learning and reducing run-time computation. Limiting the degrees of freedom also reduces overfitting, which is useful for learning large part vocabularies from limited training data. We learn steerable variants of several state-of-the-art models for object detection, human pose estimation, and facial landmark estimation. Our steerable models are smaller, faster, and often improve performance. Hamed Pirsiavash, Deva Ramanan |
CVPR | 1 |
| 2011 | AVSS 2011 demo session: A large-scale benchmark dataset for event recognition in surveillance videoabstractSummary form only given. We present a concept for automatic construction site monitoring by taking into account 4D information (3D over time), that is acquired from highly-overlapping digital aerial images. On the one hand today's maturity of flying micro aerial vehicles (MAVs) enables a low-cost and an efficient image acquisition of high-quality data that maps construction sites entirely from many varying viewpoints. On the other hand, due to low-noise sensors and high redundancy in the image data, recent developments in 3D reconstruction workflows have benefited the automatic computation of accurate and dense 3D scene information. Having both an inexpensive high-quality image acquisition and an efficient 3D analysis workflow enables monitoring, documentation and visualization of observed sites over time with short intervals. Relating acquired 4D site observations, composed of color, texture, geometry over time, largely supports automated methods toward full scene understanding, the acquisition of both the change and the construction site's progress. Sangmin Oh, Anthony Hoogs, A. G. Amitha Perera, Naresh P. Cuntoor, Chia-Chih Chen, Jong Taek Lee, Saurajit Mukherjee, Jake K. Aggarwal, Hyungtae Lee, Larry Davis 0001, Eran Swears, Xiaoyang Wang 0001, Kishore K. Reddy, Mubarak Shah, Carl Vondrick, Hamed Pirsiavash, Deva Ramanan, Jenny Yuen, Antonio Torralba 0001, Bi Song, Anesco Fong, Amit K. Roy-Chowdhury, Mita Desai |
AVSS | 17 |
| 2011 | A large-scale benchmark dataset for event recognition in surveillance videoabstractWe introduce a new large-scale video dataset designed to assess the performance of diverse visual event recognition algorithms with a focus on continuous visual event recognition (CVER) in outdoor areas with wide coverage. Previous datasets for action recognition are unrealistic for real-world surveillance because they consist of short clips showing one action by one individual [15, 8]. Datasets have been developed for movies [11] and sports [12], but, these actions and scene conditions do not apply effectively to surveillance videos. Our dataset consists of many outdoor scenes with actions occurring naturally by non-actors in continuously captured videos of the real world. The dataset includes large numbers of instances for 23 event types distributed throughout 29 hours of video. This data is accompanied by detailed annotations which include both moving object tracks and event examples, which will provide solid basis for large-scale evaluation. Additionally, we propose different types of evaluation modes for visual recognition tasks and evaluation metrics along with our preliminary experimental results. We believe that this dataset will stimulate diverse aspects of computer vision research and help us to advance the CVER tasks in the years ahead. Sangmin Oh, Anthony Hoogs, A. G. Amitha Perera, Naresh P. Cuntoor, Chia-Chih Chen, Jong Taek Lee, Saurajit Mukherjee, Jake K. Aggarwal, Hyungtae Lee, Larry Davis 0001, Eran Swears, Xiaoyang Wang 0001, Kishore K. Reddy, Mubarak Shah, Carl Vondrick, Hamed Pirsiavash, Deva Ramanan, Jenny Yuen, Antonio Torralba 0001, Bi Song, Anesco Fong, Amit K. Roy-Chowdhury, Mita Desai |
CVPR | 17 |
| 2011 | Globally-optimal greedy algorithms for tracking a variable number of objectsabstractWe analyze the computational problem of multi-object tracking in video sequences. We formulate the problem using a cost function that requires estimating the number of tracks, as well as their birth and death states. We show that the global solution can be obtained with a greedy algorithm that sequentially instantiates tracks using shortest path computations on a flow network. Greedy algorithms allow one to embed pre-processing steps, such as nonmax suppression, within the tracking algorithm. Furthermore, we give a near-optimal algorithm based on dynamic programming which runs in time linear in the number of objects and linear in the sequence length. Our algorithms are fast, simple, and scalable, allowing us to process dense input data. This results in state-of-the-art performance. Hamed Pirsiavash, Deva Ramanan, Charless C. Fowlkes |
CVPR | 1 |
| 2010 | TalkMiner: a lecture webcast search engineabstractThe design and implementation of a search engine for lecture webcasts is described. A searchable text index is created allowing users to locate material within lecture videos found on a variety of websites such as YouTube and Berkeley webcasts. The index is created from words on the presentation slides appearing in the video along with any associated metadata such as the title and abstract when available. John Adcock, Matthew Cooper 0002, Laurent Denoue, Hamed Pirsiavash, Lawrence A. Rowe |
ACM Multimedia | 4 |
| 2010 | TalkMiner: a search engine for online lecture videoabstractTalkMiner is a search engine for lecture webcasts. Lecture videos are processed to recover a set of distinct slide images and OCR is used to generate a list of indexable terms from the slides. On our prototype system, users can search and browse lists of lectures, slides in a specific lecture, and play the lecture video. Over 10,000 lecture videos have been indexed from a variety of sources. A public website will be published in mid 2010 that allows users to experiment with the search engine. John Adcock, Matthew Cooper 0002, Laurent Denoue, Hamed Pirsiavash, Lawrence A. Rowe |
ACM Multimedia | 4 |
| 2009 | Personal photo album summarizationabstractPhoto album summarization is the process of selecting a subset of photos from a larger collection which best preserves the information in the entire set and is semantically coherent. In this paper we propose a system which uses heterogeneous information sources associated with digital photos and generates a summary. Our algorithm adapts itself based on the type of event it is summarizing (Yearbook, Week or Single Day Event) We model the summarization problem as a retrieval problem based on different types of queries. We propose some evaluation metrics for the summary. We use an intuitive web based interface to present the results so that users can further explore the summary in an interactive way. This system is our submission to the CeWe Challenge for the Next Generation of Tangible Multimedia Products. Pinaki Sinha, Hamed Pirsiavash, Ramesh Jain 0001 |
ACM Multimedia | 2 |
| 2009 | Bilinear classifiers for visual recognitionabstractWe describe an algorithm for learning bilinear SVMs. Bilinear classifiers are a discriminative variant of bilinear models, which capture the dependence of data on multiple factors. Such models are particularly appropriate for visual data that is better represented as a matrix or tensor, rather than a vector. Matrix encodings allow for more natural regularization through rank restriction. For example, a rank-one scanning-window classifier yields a separable filter. Low-rank models have fewer parameters and so are easier to regularize and faster to score at run-time. We learn low-rank models with bilinear classifiers. We also use bilinear classifiers for transfer learning by sharing linear factors between different classification tasks. Bilinear classifiers are trained with biconvex programs. Such programs are optimized with coordinate descent, where each coordinate step requires solving a convex program - in our case, we use a standard off-the-shelf SVM solver. We demonstrate bilinear SVMs on difficult problems of people detection in video sequences and action classification of video sequences, achieving state-of-the-art results in both. Hamed Pirsiavash, Deva Ramanan, Charless C. Fowlkes |
NIPS | 1 |
| 2009 | Towards Environment-to-Environment (E2E) multimedia communication systems
Vivek K. Singh 0001, Hamed Pirsiavash, Ish Rishabh, Ramesh Jain 0001 |
Multim. Tools Appl. | 2 |
| 2009 | Opti-Acoustic Stereo Imaging: On System Calibration and 3-D Target ReconstructionabstractUtilization of an acoustic camera for range measurements is a key advantage for 3-D shape recovery of underwater targets by opti-acoustic stereo imaging, where the associated epipolar geometry of optical and acoustic image correspondences can be described in terms of conic sections. In this paper, we propose methods for system calibration and 3-D scene reconstruction by maximum likelihood estimation from noisy image measurements. The recursive 3-D reconstruction method utilized as initial condition a closed-form solution that integrates the advantages of two other closed-form solutions, referred to as the range and azimuth solutions. Synthetic data tests are given to provide insight into the merits of the new target imaging and 3-D reconstruction paradigm, while experiments with real data confirm the findings based on computer simulations, and demonstrate the merits of this novel 3-D reconstruction paradigm. Shahriar Negahdaripour, Hicham Sekkati, Hamed Pirsiavash |
IEEE Trans. Image Process. | 3 |
| 2007 | Integration of Motion Cues in Optical and Sonar Videos for 3-D PositioningabstractTarget-based positioning and 3-D target reconstruction are critical capabilities in deploying submersible platforms for a range of underwater applications, e.g., search and inspection missions. While optical cameras provide high-resolution and target details, they are constrained by limited visibility range. In highly turbid waters, target at up to distances of 10 s of meters can be recorded by high-frequency (MHz) 2-D sonar imaging systems that have become introduced to the commercial market in years. Because of lower resolution and SNR level and inferior target details compared to optical camera in favorable visibility conditions, the integration of both sensing modalities can enable operation in a wider range of conditions with generally better performance compared to deploying either system alone. In this paper, estimate of the 3-D motion of the integrated system and the 3-D reconstruction of scene features are addressed. We do not require establishing matches between optical and sonar features, referred to as opti-acoustic correspondences, but rather matches in either the sonar or optical motion sequences. In addition to improving the motion estimation accuracy, advantages of the system comprise overcoming certain inherent ambiguities of monocular vision, e.g., the scale-factor ambiguity, and dual interpretation of planar scenes. We discuss how the proposed solution provides an effective strategy to address the rather complex opti-acoustic stereo matching problem. Experiment with real data demonstrate our technical contribution. Shahriar Negahdaripour, Hamed Pirsiavash, Hicham Sekkati |
CVPR | 2 |
| 2007 | Opti-Acoustic Stereo Imaging, System Calibration and 3-D ReconstructionabstractUtilization of an acoustic camera for range measurements is a key advantage for 3-D shape recovery of underwater targets by opti-acoustic stereo imaging, where the associated epipolar geometry of optical and acoustic image correspondences can be described in terms of conic sections. In this paper, we propose methods for system calibration and 3-D scene reconstruction by maximum likelihood estimation from noisy image measurements. The recursive 3-D reconstruction method utilized as initial condition a closed-form solution that integrates the advantages of so-called range and azimuth solutions. Synthetic data tests are given to provide insight into the merits of the new target imaging and 3-D reconstruction paradigm, while experiments with real data confirm the findings based on computer simulations, and demonstrate the merits of this novel 3-D reconstruction paradigm. Shahriar Negahdaripour, Hicham Sekkati, Hamed Pirsiavash |
CVPR | 3 |
| 2005 | An Iterative Approach for Reconstruction of Arbitrary Sparsely Sampled Magnetic Resonance ImagesabstractIn many fast MR imaging techniques, K-space is sampled sparsely in order to gain a fast traverse of K-space. These techniques use non-Cartesian sampling trajectories like radial, zigzag, and spiral. In the reconstruction procedure, usually interpolation methods are used to obtain missing samples on a regular grid. In this paper, we propose an iterative method for image reconstruction which uses the black marginal area of the image. The proposed iterative solution offers a great enhancement in the quality of the reconstructed image in comparison with conventional algorithms like zero filling and neural network. This method is applied on MRI data and its improved performance over other methods is demonstrated. Hamed Pirsiavash, Mohammad Soleymani 0003, Gholam-Ali Hossein-Zadeh |
CBMS | 1 |
| 2005 | A Robust Free Size OCR for Omni-Font Persian/Arabic Printed Document Using Combined MLP/SVM
Hamed Pirsiavash, Ramin Mehran, Farbod Razzazi |
CIARP | 1 |