EDBT 2026 Demo / reviewers in the wild / expert
Muhammad Salman Asif
dblp:21/1910 · also M. Salman Asif
· DBLP profile ↗
46ranked-venue papers
5as first author
32since 2021 · last 2025
0000-0001-5993-3903ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 35 · 5 first-author · 22 since 2021Artificial intelligence and machine learning · 27 · 22 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MMP: Towards Robust Multi-Modal Learning with Masked Modality Projection
Niki Nezakati, Md Kaykobad Reza, Ameya Patil 0001, Mashhour Solh, Muhammad Salman Asif |
IEEE Big Data | 5 |
| 2025 | RENO: Real-Time Neural Compression for 3D LiDAR Point CloudsabstractDespite the substantial advancements demonstrated by learning-based neural models in the LiDAR Point Cloud Compression (LPCC) task, realizing real-time compression—an indispensable criterion for numerous industrial applications—remains a formidable challenge. This paper proposes RENO, the first real-time neural codec for 3D LiDAR point clouds, achieving superior performance with a lightweight model. RENO skips the octree construction and directly builds upon the multiscale sparse tensor representation. Instead of the multi-stage inferring, RENO devises sparse occupancy codes, which exploit cross-scale correlation and derive voxels’ occupancy in a one-shot manner, greatly saving processing time. Experimental results demonstrate that the proposed RENO achieves real-time coding speed, 10 fps at 14-bit depth on a desktop platform (e.g., one RTX 3090 GPU) for both encoding and decoding processes, while providing 12.25% and 48.34% bit-rate savings compared to G-PCCv23 and Draco, respectively, at a similar quality. RENO model size is merely 1MB, making it attractive for practical applications. The source code is available at https://github.com/NJUVISION/RENO. Kang You, Tong Chen 0004, Dandan Ding, Muhammad Salman Asif, Zhan Ma 0001 |
CVPR | 4 |
| 2025 | VOccl3D: A Video Benchmark Dataset for 3D Human Pose and Shape Estimation Under Real Occlusions
Yash Garg, Saketh Bachu, Arindam Dutta, Rohit Lal, Sarosij Bose, Calvin-Khang Ta, Muhammad Salman Asif, Amit K. Roy-Chowdhury |
ICCV | 7 |
| 2025 | Targeted Unlearning with Single Layer Unlearning GradientabstractMachine unlearning methods aim to remove sensitive or unwanted content from trained models, but typically demand extensive model updates at significant computational cost while potentially degrading model performance on both related and unrelated tasks. We propose Single Layer Unlearning Gradient (SLUG) as an efficient method to unlearn targeted information by updating a single critical layer using a one-time gradient computation. SLUG uses layer importance and gradient alignment metrics to identify the optimal layer for targeted information removal while preserving the model utility. We demonstrate the effectiveness of SLUG for CLIP, Stable Diffusion, and vision-language models (VLMs) in removing concrete (e.g., identities and objects) and abstract concepts (e.g., artistic styles). On the UnlearnCanvas benchmark, SLUG achieves comparable unlearning performance to existing methods while requiring significantly less computational resources. Our proposed approach offers a practical solution for targeted unlearning that is computationally efficient and precise. Our code is available at https://github.com/CSIPlab/SLUG Zikui Cai, Yaoteng Tan, Muhammad Salman Asif |
ICML | 3 |
| 2025 | STRIDE: Single-Video Based Temporally Continuous Occlusion-Robust 3D Pose EstimationabstractAccurately estimating 3D human poses is crucial for fields like action recognition, gait recognition, and virtual/augmented reality. However, predicting human poses under severe occlusion remains a persistent and significant challenge. Existing image-based estimators struggle with heavy occlusions due to a lack of temporal context, resulting in inconsistent predictions, while video-based models, despite benefiting from temporal data, face limitations with prolonged occlusions over multiple frames. Additionally, existing algorithms often struggle to generalize unseen videos. Addressing these challenges, we propose STRIDE (Single-video based TempoRally contInuous Occlusion-Robust 3D Pose Estimation), a novel Test-Time Training (TTT) approach to fit a human motion prior for estimating 3D human poses for each video. Our proposed approach handles occlusions not encountered during the model's training by refining a sequence of noisy initial pose estimates into accurate, temporally coherent poses at test time, effectively overcoming the limitations of existing methods. Our flexible, model-agnostic framework allows us to use any off-the-shelf 3D pose estimation method to improve robustness and temporal consistency. We validate STRIDE's efficacy through comprehensive experiments on multiple challenging datasets where it not only outperforms existing single-image and video-based pose estimation models but also showcases superior handling of substantial occlusions, achieving fast, robust, accurate, and temporally consistent 3D pose estimates. Code is made publicly available at https://github.com/take2rohit/stride Rohit Lal, Saketh Bachu, Yash Garg, Arindam Dutta, Calvin-Khang Ta, Hannah Dela Cruz, Dripta S. Raychaudhuri, Muhammad Salman Asif, Amit K. Roy-Chowdhury |
WACV | 8 |
| 2025 | Robust Multimodal Learning With Missing Modalities via Parameter-Efficient AdaptationabstractMultimodal learning seeks to utilize data from multiple sources to improve the overall performance of downstream tasks. It is desirable for redundancies in the data to make multimodal systems robust to missing or corrupted observations in some correlated modalities. However, we observe that the performance of several existing multimodal networks significantly deteriorates if one or multiple modalities are absent at test time. To enable robustness to missing modalities, we propose a simple and parameter-efficient adaptation procedure for pretrained multimodal networks. In particular, we exploit modulation of intermediate features to compensate for the missing modalities. We demonstrate that such adaptation can partially bridge performance drop due to missing modalities and outperform independent, dedicated networks trained for the available modality combinations in some cases. The proposed adaptation requires extremely small number of parameters (e.g., fewer than 1% of the total parameters) and applicable to a wide range of modality combinations and tasks. We conduct a series of experiments to highlight the missing modality robustness of our proposed method on five different multimodal tasks across seven datasets. Our proposed method demonstrates versatility across various tasks and datasets, and outperforms existing methods for robust multimodal learning with missing modalities. Md Kaykobad Reza, Ashley Prater-Bennette, Muhammad Salman Asif |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Disguise without Disruption: Utility-Preserving Face De-identificationabstractWith the rise of cameras and smart sensors, humanity generates an exponential amount of data. This valuable information, including underrepresented cases like AI in medical settings, can fuel new deep-learning tools. However, data scientists must prioritize ensuring privacy for individuals in these untapped datasets, especially for images or videos with faces, which are prime targets for identification methods. Proposed solutions to de-identify such images often compromise non-identifying facial attributes relevant to downstream tasks. In this paper, we introduce Disguise, a novel algorithm that seamlessly de-identifies facial images while ensuring the usability of the modified data. Unlike previous approaches, our solution is firmly grounded in the domains of differential privacy and ensemble-learning research. Our method involves extracting and substituting depicted identities with synthetic ones, generated using variational mechanisms to maximize obfuscation and non-invertibility. Additionally, we leverage supervision from a mixture-of-experts to disentangle and preserve other utility attributes. We extensively evaluate our method using multiple datasets, demonstrating a higher de-identification rate and superior consistency compared to prior approaches in various downstream tasks. Zikui Cai, Zhongpai Gao, Benjamin Planche, Meng Zheng 0002, Terrence Chen, Muhammad Salman Asif, Ziyan Wu 0001 |
AAAI | 6 |
| 2024 | PNeRV: Enhancing Spatial Consistency via Pyramidal Neural Representation for VideosabstractThe primary focus of Neural Representation for Videos (NeRV) is to effectively model its spatiotemporal consis-tency. However, current NeRV systems often face a signif-icant issue of spatial inconsistency, leading to decreased perceptual quality. To address this issue, we introduce the Pyramidal Neural Representation for Videos (PNeRV), which is built on a multi-scale information connection and comprises a lightweight rescaling operator, Kronecker Fully-connected layer (KFc), and a Benign Selective Mem-ory (BSM) mechanism. The KFc, inspired by the tensor de-composition of the vanilla Fully-connected layer, facilitates low-cost rescaling and global correlation modeling. BSM merges high-level features with granular ones adaptively. Furthermore, we provide an analysis based on the Univer-sal Approximation Theory of the NeRV system and vali-date the effectiveness of the proposed PNeRV. We conducted comprehensive experiments to demonstrate that PNeRV sur-passes the performance of contemporary NeRV models, achieving the best results in video regression on UVG and DAVIS under various metrics (PSNR, SSIM, LPIPS, and FVD). Compared to vanilla N eRV, P N eRV achieves$a+4.49$dB gain in PSNR and a 231% increase in FVD on UVG, along with$a+3.28$dB PSNR and 634% FVD increase on DAVIS. Muhammad Salman Asif, Zhan Ma 0001 |
CVPR | 2 |
| 2024 | EDformer: Transformer-Based Event Denoising Across Varied Noise Levels
Bin Jiang 0018, Bohan Qu, Muhammad Salman Asif, Zhan Ma 0001 |
ECCV (27) | 4 |
| 2024 | Token-Based Spatiotemporal Representation of the EventsabstractThe event camera’s low power consumption and ability to capture microsecond brightness changes make it attractive for various computer vision tasks. Existing event representation methods typically convert events into frames, voxel grids, or spikes for deep neural networks (DNNs). However, these approaches often sacrifice temporal granularity or require specialized devices for processing. This work introduces a novel token-based event representation, where each event is considered a fundamental processing unit termed an event-token. This approach preserves the sequence’s intricate spatiotemporal attributes at the event level. Moreover, we propose a Three-way Attention mechanism in the Event Transformer Block (ETB) to collaboratively construct temporal and spatial correlations between events. We compare our proposed token-based event representation extensively with other prevalent methods for object classification and optical flow estimation. The experimental results showcase its competitive performance while demanding minimal computational resources on standard devices. Bin Jiang 0018, Muhammad Salman Asif, Xun Cao, Zhan Ma 0001 |
ICASSP | 3 |
| 2024 | Parameter-Efficient Adaptation for Computational ImagingabstractDeep learning-based methods provide remarkable performance in a number of computational imaging problems. Examples include end-to-end trained networks that map measurements to unknown signals, plug-and-play (PnP) methods that use pretrained denoisers as image prior, and model-based unrolled networks that train artifact removal blocks. Many of these methods lack robustness and fail to generalize with distribution shifts in data, measurements, and noise. In this paper, we present a simple framework to perform domain adaptation as data and measurement distribution shifts. Our method learns a small number of factors to add in a pretrained model to bridge the gap in performance. We present a number of experiments on accelerated magnetic resonance imaging (MRI) reconstruction and image deblurring to demonstrate that our method requires a small amount of memory and parameter overhead to adapt to new domains. Nebiyou Yismaw, Ulugbek Kamilov, Muhammad Salman Asif |
ICASSP | 3 |
| 2024 | Prior Mismatch and Adaptation in PnP-ADMM with a Nonconvex Convergence AnalysisabstractPlug-and-Play (PnP) priors is a widely-used family of methods for solving imaging inverse problems by integrating physical measurement models with image priors specified using image denoisers. PnP methods have been shown to achieve state-of-the-art performance when the prior is obtained using powerful deep denoisers. Despite extensive work on PnP, the topic of distribution mismatch between the training and testing data has often been overlooked in the PnP literature. This paper presents a set of new theoretical and numerical results on the topic of prior distribution mismatch and domain adaptation for the alternating direction method of multipliers (ADMM) variant of PnP. Our theoretical result provides an explicit error bound for PnP-ADMM due to the mismatch between the desired denoiser and the one used for inference. Our analysis contributes to the work in the area by considering the mismatch under nonconvex data-fidelity terms and expansive denoisers. Our first set of numerical results quantifies the impact of the prior distribution mismatch on the performance of PnP-ADMM on the problem of image super-resolution. Our second set of numerical results considers a simple and effective domain adaption strategy that closes the performance gap due to the use of mismatched denoisers. Our results suggest the relative robustness of PnP-ADMM to prior distribution mismatch, while also showing that the performance gap can be significantly reduced with only a few training samples from the desired distribution. Shirin Shoushtari, Jiaming Liu 0001, Edward P. Chandler, Muhammad Salman Asif, Ulugbek Kamilov |
ICML | 4 |
| 2024 | Efficient Visual Computing With Camera RAW SnapshotsabstractConventional cameras capture image irradiance (RAW) on a sensor and convert it to RGB images using an image signal processor (ISP). The images can then be used for photography or visual computing tasks in a variety of applications, such as public safety surveillance and autonomous driving. One can argue that since RAW images contain all the captured information, the conversion of RAW to RGB using an ISP is not necessary for visual computing. In this paper, we propose a novel ρ-Vision framework to perform high-level semantic understanding and low-level compression using RAW images without the ISP subsystem used for decades. Considering the scarcity of available RAW image datasets, we first develop an unpaired CycleR2R network based on unsupervised CycleGAN to train modular unrolled ISP and inverse ISP (invISP) models using unpaired RAW and RGB images. We can then flexibly generate simulated RAW images (simRAW) using any existing RGB image dataset and finetune different models originally trained in the RGB domain to process real-world camera RAW images. We demonstrate object detection and image compression capabilities in RAW-domain using RAW-domain YOLOv3 and RAW image compressor (RIC) on camera snapshots. Quantitative results reveal that RAW-domain task inference provides better detection accuracy and compression efficiency compared to that in the RGB domain. Furthermore, the proposed ρ-Vision generalizes across various camera sensors and different task-specific models. An added benefit of employing the ρ-Vision is the elimination of the need for ISP, leading to potential reductions in computations and processing times. Ming Lu 0003, Xu Zhang 0006, Xin Feng 0007, Muhammad Salman Asif, Zhan Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Ensemble-based Blackbox Attacks on Dense PredictionabstractWe propose an approach for adversarial attacks on dense prediction models (such as object detectors and segmentation). It is well known that the attacks generated by a single surrogate model do not transfer to arbitrary (blackbox) victim models. Furthermore, targeted attacks are often more challenging than the untargeted attacks. In this paper, we show that a carefully designed ensemble can create effective attacks for a number of victim models. In particular, we show that normalization of the weights for individual models plays a critical role in the success of the attacks. We then demonstrate that by adjusting the weights of the ensemble according to the victim model can further improve the performance of the attacks. We performed a number of experiments for object detectors and segmentation to highlight the significance of the our proposed methods. Our proposed ensemble-based method outperforms existing blackbox attack methods for object detection and segmentation. Finally we show that our proposed method can also generate a single perturbation that can fool multiple blackbox detection and segmentation models simultaneously. Code is available at https://github.com/CSIPlab/EBAD Zikui Cai, Yaoteng Tan, Muhammad Salman Asif |
CVPR | 3 |
| 2023 | DNeRV: Modeling Inherent Dynamics via Difference Neural Representation for VideosabstractExisting implicit neural representation (INR) methods do not fully exploit spatiotemporal redundancies in videos. Index-based INRs ignore the content-specific spatial features and hybrid INRs ignore the contextual dependency on adjacent frames, leading to poor modeling capability for scenes with large motion or dynamics. We analyze this limitation from the perspective of function fitting and reveal the importance of frame difference. To use explicit motion information, we propose Difference Neural Representation for Videos (DNeRV), which consists of two streams for content and frame difference. We also introduce a collaborative content unit for effective feature fusion. We test DNeRV for video compression, inpainting, and interpolation. D-NeRVachieves competitive results against the state-of-the-art neural compression approaches and outperforms existing implicit methods on downstream inpainting and interpolation for$960\times 1920$videos. Muhammad Salman Asif, Zhan Ma 0001 |
CVPR | 2 |
| 2023 | Privacy Preserving Face Recognition with Lensless CameraabstractThe widespread adoption of facial recognition technology is a global phenomenon. Facial recognition systems leverage upon image data containing faces. This poses serious threats to user privacy as the data is exposed to potential data breaches. In this paper, we propose a face recognition system that works without compromising user privacy. It utilizes data captured by FlatCam - a lensless camera. FlatCam captures the scene as a sensor measurement that is visually unintelligible. The proposed system preserves user privacy since it works directly on FlatCam’s sensor measurements without the need of FlatCam camera parameters which are required for pixel reconstruction. We propose a frequency domain deep learning solution that computes the DCT of the sensor field at multiple resolutions and organizes it into sub-bands before training a classification network with attention. The multi-resolution DCT subband representation leads to huge performance gains when compared to using the sensor measurement directly for training. Our proposed system was trained and tested on a real lensless camera dataset - the Flat-Cam Face dataset. Privacy of user is preserved during both training and testing. Experimental results demonstrate the effectiveness of our method. Chris Henry, Muhammad Salman Asif, Zhu Li 0001 |
ICASSP | 2 |
| 2023 | Compressive Sensing with Tensorized AutoencoderabstractDeep networks can be trained to map images into a low-dimensional latent space. In many cases, different images in a collection are articulated versions of one another; for example, same object with different lighting, background, or pose. Furthermore, in many cases, parts of images can be corrupted by noise or missing entries. In this paper, our goal is to recover images without access to the ground-truth (clean) images using the articulations as structural prior of the data. Such recovery problems fall under the domain of compressive sensing. We propose to learn autoencoder with tensor ring factorization on the the embedding space to impose structural constraints on the data. In particular, we use a tensor ring structure in the bottleneck layer of the autoencoder that utilizes the soft labels of the structured dataset. We empirically demonstrate the effectiveness of the proposed approach for inpainting and denoising applications. The resulting method achieves better reconstruction quality compared to other generative prior-based self-supervised recovery approaches for compressive sensing. Rakib Hyder, Muhammad Salman Asif |
ICASSP | 2 |
| 2023 | Leveraging Local Patch Differences in Multi-Object Scenes for Generative Adversarial AttacksabstractState-of-the-art generative model-based attacks against image classifiers overwhelmingly focus on single-object(i.e., single dominant object) images. Different from such settings, we tackle a more practical problem of generating adversarial perturbations using multi-object (i.e., multiple dominant objects) images as they are representative of most real-world scenes. Our goal is to design an attack strategy that can learn from such natural scenes by leveraging the local patch differences that occur inherently in such images (e.g. difference between the local patch on the object ‘person’ and the object ‘bike’ in a traffic scene). Our key idea is to misclassify an adversarial multi-object image by confusing the victim classifier for each local patch in the image. Based on this, we propose a novel generative attack (called Local Patch Difference or LPD-Attack) where a novel contrastive loss function uses the aforesaid local differences in feature space of multi-object scenes to optimize the perturbation generator. Through various experiments across diverse victim convolutional neural networks, we show that our approach outperforms baseline generative attacks with highly transferable perturbations when evaluated under different white-box and black-box settings. Abhishek Aich, Shasha Li 0002, Chengyu Song, Muhammad Salman Asif, Srikanth V. Krishnamurthy, Amit K. Roy-Chowdhury |
WACV | 4 |
| 2023 | GCNDepth: Self-supervised monocular depth estimation based on graph convolutional networkabstractDepth estimation is a challenging task of 3D reconstruction to enhance the accuracy sensing of environment awareness. This work brings a new solution with improvements, which increases the quantitative and qualitative understanding of depth maps compared to existing methods. Recently, convolutional neural networks (CNN) have demonstrated their extraordinary ability to estimate depth maps from monocular videos. However, traditional CNN does not support a topological structure, and they can work only on regular image regions with determined sizes and weights. On the other hand, graph convolutional networks (GCN) can handle the convolution of non-Euclidean data, and they can be applied to irregular image regions within a topological structure. Therefore, to preserve object geometric appearances and objects locations in the scene, in this work, we aim to exploit GCN for a self-supervised monocular depth estimation model. Our model consists of two parallel auto-encoder networks: the first is an auto-encoder that will depend on ResNet-50 and extract the feature from the input image and on multi-scale GCN to estimate the depth map. In turn, the second network will be used to estimate the ego-motion vector (i.e., 3D pose) between two consecutive frames based on ResNet-18. The estimated 3D pose and depth map will be used to construct the target image. A combination of loss functions related to photometric, reprojection, and smoothness is used to cope with bad depth prediction and preserve the discontinuities of the objects. Our method and performance are improved quantitatively and qualitatively. In particular, our method provided comparable and promising results with a high prediction accuracy of 89% on the publicly available KITTI dataset. Our method also offers 40% reduction in the number of trainable parameters compared to the state of the art solutions.In addition, we tested our trained model with Make3D dataset to evaluate the trained model on a new dataset with low resolution images. The source code is publicly available at (https://github.com/ArminMasoumian/GCNDepth.git) Armin Masoumian, Hatem A. Rashwan, Saddam Abdulwahab, Julián Cristiano, Muhammad Salman Asif, Domenec Puig |
Neurocomputing | 5 |
| 2022 | Context-Aware Transfer Attacks for Object DetectionabstractBlackbox transfer attacks for image classifiers have been extensively studied in recent years. In contrast, little progress has been made on transfer attacks for object detectors. Object detectors take a holistic view of the image and the detection of one object (or lack thereof) often depends on other objects in the scene. This makes such detectors inherently context-aware and adversarial attacks in this space are more challenging than those targeting image classifiers. In this paper, we present a new approach to generate context-aware attacks for object detectors. We show that by using co-occurrence of objects and their relative locations and sizes as context information, we can successfully generate targeted mis-categorization attacks that achieve higher transfer success rates on blackbox object detectors than the state-of-the-art. We test our approach on a variety of object detectors with images from PASCAL VOC and MS COCO datasets and demonstrate up to 20 percentage points improvement in performance compared to the other state-of-the-art methods. Zikui Cai, Xinxin Xie, Shasha Li 0001, Mingjun Yin, Chengyu Song, Srikanth V. Krishnamurthy, Amit K. Roy-Chowdhury, Muhammad Salman Asif |
AAAI | 8 |
| 2022 | Zero-Query Transfer Attacks on Context-Aware Object DetectorsabstractAdversarial attacks perturb images such that a deep neural network produces incorrect classification results. A promising approach to defend against adversarial attacks on natural multi-object scenes is to impose a context-consistency check, wherein, if the detected objects are not consistent with an appropriately defined context, then an attack is suspected. Stronger attacks are needed to fool such context-aware detectors. We present the first approach for generating context-consistent adversarial attacks that can evade the context-consistency check of black-box object detectors operating on complex, natural scenes. Unlike many black-box attacks that perform repeated attempts and open themselves to detection, we assume a “zero-query” setting, where the attacker has no knowledge of the classification decisions of the victim system. First, we derive multiple attack plans that assign incorrect labels to victim objects in a context-consistent manner. Then we design and use a novel data structure that we call the perturbation success probability matrix, which enables us to filter the attack plans and choose the one most likely to succeed. This final attack plan is implemented using a perturbation-bounded adversarial attack algorithm. We compare our zero-query attack against a few-query scheme that repeatedly checks if the victim system is fooled. We also compare against state-of-the-art context-agnostic attacks. Against a context-aware defense, the fooling rate of our zero-query approach is significantly higher than context-agnostic approaches and higher than that achievable with up to three rounds of the fewquery scheme. Zikui Cai, Shantanu Rane, Alejandro E. Brito, Chengyu Song, Srikanth V. Krishnamurthy, Amit K. Roy-Chowdhury, Muhammad Salman Asif |
CVPR | 7 |
| 2022 | Incremental Task Learning with Incremental Rank Updates
Rakib Hyder, Ken Shao, Boyu Hou, Panos P. Markopoulos, Ashley Prater-Bennette, Muhammad Salman Asif |
ECCV (23) | 6 |
| 2022 | GAMA: Generative Adversarial Multi-Object Scene AttacksabstractThe majority of methods for crafting adversarial attacks have focused on scenes with a single dominant object (e.g., images from ImageNet). On the other hand, natural scenes include multiple dominant objects that are semantically related. Thus, it is crucial to explore designing attack strategies that look beyond learning on single-object scenes or attack single-object victim classifiers. Due to their inherent property of strong transferability of perturbations to unknown models, this paper presents the first approach of using generative models for adversarial attacks on multi-object scenes. In order to represent the relationships between different objects in the input scene, we leverage upon the open-sourced pre-trained vision-language model CLIP (Contrastive Language-Image Pre-training), with the motivation to exploit the encoded semantics in the language space along with the visual space. We call this attack approach Generative Adversarial Multi-object Attacks (GAMA). GAMA demonstrates the utility of the CLIP model as an attacker's tool to train formidable perturbation generators for multi-object scenes. Using the joint image-text features to train the generator, we show that GAMA can craft potent transferable perturbations in order to fool victim classifiers in various attack settings. For example, GAMA triggers ~16% more misclassification than state-of-the-art generative approaches in black-box settings where both the classifier architecture and data distribution of the attacker are different from the victim. Our code is available here: https://abhishekaich27.github.io/gama.html Abhishek Aich, Calvin-Khang Ta, Akash Gupta 0001, Chengyu Song, Srikanth V. Krishnamurthy, Muhammad Salman Asif, Amit K. Roy-Chowdhury |
NeurIPS | 6 |
| 2022 | ADC: Adversarial attacks against object Detection that evade Context consistency checksabstractDeep Neural Networks (DNNs) have been shown to be vulnerable to adversarial examples, which are slightly perturbed input images which lead DNNs to make wrong predictions. To protect from such examples, various defense strategies have been proposed. A very recent defense strategy for detecting adversarial examples, that has been shown to be robust to current attacks, is to check for intrinsic context consistencies in the input data, where context refers to various relationships (e.g., object-to-object co-occurrence relationships) in images. In this paper, we show that even context consistency checks can be brittle to properly crafted adversarial examples and to the best of our knowledge, we are the first to do so. Specifically, we propose an adaptive framework to generate examples that subvert such defenses, namely, Adversarial attacks against object Detection that evade Context consistency checks (ADC). In ADC, we formulate a joint optimization problem which has two attack goals, viz., (i) fooling the object detector and (ii) evading the context consistency check system, at the same time. Experiments on both PASCAL VOC and MS COCO datasets show that examples generated with ADC fool the object detector with a success rate of over 85% in most cases, and at the same time evade the recently proposed context consistency checks, with a "bypassing" rate of over 80% in most cases. Our results suggest that "how to robustly model con- text and check its consistency," is still an open problem. Mingjun Yin, Shasha Li 0001, Chengyu Song, Muhammad Salman Asif, Amit K. Roy-Chowdhury, Srikanth V. Krishnamurthy |
WACV | 4 |
| 2022 | Spatial Temporal Video Enhancement Using Alternating ExposuresabstractHigh-speed video acquisition under poor illumination conditions is a challenging task. Imaging using long exposure can ensure brightness and suppress noise. However, the captured images may be blurry due to fast object movements or camera shakes. Imaging with short exposure can record sharp textures, but the high camera gain may cause noticeable noise. To alleviate this dilemma, we design a camera system using alternating exposures, where frames expose cyclically in a short-long way. The system consists of restoration and interpolation modules to reconstruct sharp, noise-reduced, high-frame-rate frames from low-frame-rate alternate-exposed input images. We design an optical-flow-based alternate-complementary alignment architecture for spatial enhancement, which effectively aligns the short-exposed and long-exposed images in a two-stage progressive way. Moreover, it explores complementary information from short-exposed and long-exposed inputs to ensure consistency between outputs. We propose a flow-enhanced frame interpolation module for temporal enhancement, which refines the intermediate flows and reconstructs the intermediate images based on the restored images of the alignment network and warped input neighboring frames. The whole network with two modules is end-to-end jointly learnable. We first evaluate the algorithm on simulation data. To demonstrate practicality, we then test it on real data by setting up a prototype camera. We propose an effective spatial degradation regularization strategy to reduce the domain gap between simulation and real data. Besides, we extend our method by integrating multi-frame exposure fusion technology to reduce overexposure areas in real scenarios. Experimental results show that our method performs favorably against state-of-the-art methods on both synthetic data and real-world data. Wang Shen, Guo Lu, Guangtao Zhai, Li Chen 0021, Muhammad Salman Asif |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | Exploiting Multi-Object Relationships for Detecting Adversarial Attacks in Complex ScenesabstractVision systems that deploy Deep Neural Networks (DNNs) are known to be vulnerable to adversarial examples. Recent research has shown that checking the intrinsic consistencies in the input data is a promising way to detect adversarial attacks (e.g., by checking the object co-occurrence relationships in complex scenes). However, existing approaches are tied to specific models and do not offer generalizability. Motivated by the observation that language descriptions of natural scene images have already captured the object co-occurrence relationships that can be learned by a language model, we develop a novel approach to perform context consistency checks using such language models. The distinguishing aspect of our approach is that it is independent of the deployed object detector and yet offers very high accuracy in terms of detecting adversarial examples in practical scenes with multiple objects. Experiments on the PASCAL VOC and MS COCO datasets show that our method can outperform state-of-the-art methods in detecting adversarial attacks. Mingjun Yin, Shasha Li 0001, Zikui Cai, Chengyu Song, Muhammad Salman Asif, Amit K. Roy-Chowdhury, Srikanth V. Krishnamurthy |
ICCV | 5 |
| 2021 | A Simple Framework for 3D Lensless Imaging with Programmable MasksabstractLensless cameras provide a framework to build thin imaging systems by replacing the lens in a conventional camera with an amplitude or phase mask near the sensor. Existing methods for lensless imaging can recover the depth and intensity of the scene, but they require solving computationally-expensive inverse problems. Furthermore, existing methods struggle to recover dense scenes with large depth variations. In this paper, we propose a lensless imaging system that captures a small number of measurements using different patterns on a programmable mask. In this context, we make three contributions. First, we present a fast recovery algorithm to recover textures on a fixed number of depth planes in the scene. Second, we consider the mask design problem, for programmable lensless cameras, and provide a design template for optimizing the mask patterns with the goal of improving depth estimation. Third, we use a refinement network as a post-processing step to identify and remove artifacts in the reconstruction. These modifications are evaluated extensively with experimental results on a lensless camera prototype to showcase the performance benefits of the optimized masks and recovery algorithms over the state of the art. Yucheng Zheng, Aswin C. Sankaranarayanan, Muhammad Salman Asif |
ICCV | 4 |
| 2021 | Solving Fourier Phase Retrieval with a Reference Image as a Sequence of Linear Inverse ProblemsabstractFourier phase retrieval problem is equivalent to the recovery of a two-dimensional image from its autocorrelation measurements. This problem is generally nonlinear and nonconvex. Good initialization and prior information about the support or sparsity of the target image are often critical for a robust recovery. In this paper, we show that the presence of a known reference image can help us solve the nonlinear phase retrieval problem as a sequence of small linear inverse problems. Instead of recovering the entire image at once, our sequential method recovers a small number of rows or columns by solving a linear deconvolution problem at every step. Existing methods for the reference-based (holographic) phase retrieval either assume that the reference and target images are sufficiently separated so that the recovery problem is linear or recover the image via nonlinear optimization. In contrast, our proposed method does not require the separation condition. We performed an extensive set of simulations to demonstrate that our proposed method can successfully recover images from autocorrelation data under different settings of reference placement and noise. Fahimeh Arab, Muhammad Salman Asif |
ICIP | 2 |
| 2021 | Data-Driven Illumination Patterns For Coded Diffraction ImagingabstractSignal recovery from nonlinear measurements involves solving an iterative optimization problem. In this paper, we present a framework to optimize the sensing parameters to improve the quality of the signal recovered by the given iterative method. In particular, we learn illumination patterns to recover signals from coded diffraction patterns using a fixed-cost alternating minimization-based phase retrieval method. Coded diffraction phase retrieval is a physically realistic system in which the signal is first modulated by a sequence of codes before the sensor records its Fourier amplitude. We represent the phase retrieval method as an unrolled network with a fixed number of layers and minimize the recovery error by optimizing over the measurement parameters. Since the number of iterations/layers are fixed, the recovery runs under a fixed cost. We present extensive simulation results on a variety of datasets under different conditions and a comparison with existing methods. Our results demonstrate that the proposed method provides near-perfect reconstruction using patterns learned with a small number of training images. Our proposed method provides significant improvements over existing methods both in terms of accuracy and speed. Zikui Cai, Rakib Hyder, Muhammad Salman Asif |
ICIP | 3 |
| 2021 | Adversarial Attacks on Black Box Video Classifiers: Leveraging the Power of Geometric TransformationsabstractWhen compared to the image classification models, black-box adversarial attacks against video classification models have been largely understudied. This could be possible because, with video, the temporal dimension poses significant additional challenges in gradient estimation. Query-efficient black-box attacks rely on effectively estimated gradients towards maximizing the probability of misclassifying the target video. In this work, we demonstrate that such effective gradients can be searched for by parameterizing the temporal structure of the search space with geometric transformations. Specifically, we design a novel iterative algorithm GEOmetric TRAnsformed Perturbations (GEO-TRAP), for attacking video classification models. GEO-TRAP employs standard geometric transformation operations to reduce the search space for effective gradients into searching for a small group of parameters that define these operations. This group of parameters describes the geometric progression of gradients, resulting in a reduced and structured search space. Our algorithm inherently leads to successful perturbations with surprisingly few queries. For example, adversarial examples generated from GEO-TRAP have better attack success rates with ~73.55% fewer queries compared to the state-of-the-art method for video adversarial attacks on the widely used Jester dataset. Overall, our algorithm exposes vulnerabilities of diverse video classification models and achieves new state-of-the-art results under black-box settings on two large datasets. Shasha Li 0001, Abhishek Aich, Shitong Zhu, Muhammad Salman Asif, Chengyu Song, Amit K. Roy-Chowdhury, Srikanth V. Krishnamurthy |
NeurIPS | 4 |
| 2021 | Recovery Analysis for Plug-and-Play Priors using the Restricted Eigenvalue ConditionabstractThe plug-and-play priors (PnP) and regularization by denoising (RED) methods have become widely used for solving inverse problems by leveraging pre-trained deep denoisers as image priors. While the empirical imaging performance and the theoretical convergence properties of these algorithms have been widely investigated, their recovery properties have not previously been theoretically analyzed. We address this gap by showing how to establish theoretical recovery guarantees for PnP/RED by assuming that the solution of these methods lies near the fixed-points of a deep neural network. We also present numerical results comparing the recovery performance of PnP/RED in compressive sensing against that of recent compressive sensing algorithms based on generative models. Our numerical results suggest that PnP with a pre-trained artifact removal network provides significantly better results compared to the existing state-of-the-art methods. Jiaming Liu 0001, Muhammad Salman Asif, Brendt Wohlberg, Ulugbek Kamilov |
NeurIPS | 2 |
| 2021 | A Dual Camera System for High Spatiotemporal Resolution Video AcquisitionabstractThis paper presents a dual camera system for high spatiotemporal resolution (HSTR) video acquisition, where one camera shoots a video with high spatial resolution and low frame rate (HSR-LFR) and another one captures a low spatial resolution and high frame rate (LSR-HFR) video. Our main goal is to combine videos from LSR-HFR and HSR-LFR cameras to create an HSTR video. We propose an end-to-end learning framework, AWnet, mainly consisting of a FlowNet and a FusionNet that learn an adaptive weighting function in pixel domain to combine inputs in a frame recurrent fashion. To improve the reconstruction quality for cameras used in reality, we also introduce noise regularization under the same framework. Our method has demonstrated noticeable performance gains in terms of both objective PSNR measurement in simulation with different publicly available video and light-field datasets and subjective evaluation with real data captured by dual iPhone 7 and Grasshopper3 cameras. Ablation studies are further conducted to investigate and explore various aspects, such as reference structure, camera parallax, exposure time, etc) of our system to fully understand its capability for potential applications. Zhan Ma 0001, Muhammad Salman Asif, Yiling Xu, Wenbo Bao, Jun Sun 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Robust High Dynamic Range (HDR) Imaging with Complex Motion and Parallax
Zhiyuan Pu, Peiyao Guo, Muhammad Salman Asif, Zhan Ma 0001 |
ACCV (2) | 3 |
| 2020 | Non-Adversarial Video Synthesis with Learned PriorsabstractMost of the existing works in video synthesis focus on generating videos using adversarial learning. Despite their success, these methods often require input reference frame or fail to generate diverse videos from the given data distribution, with little to no uniformity in the quality of videos that can be generated. Different from these methods, we focus on the problem of generating videos from latent noise vectors, without any reference input frames. To this end, we develop a novel approach that jointly optimizes the input latent space, the weights of a recurrent neural network and a generator through non-adversarial learning. Optimizing for the input latent space along with the network weights allows us to generate videos in a controlled environment, i.e., we can faithfully generate all videos the model has seen during the learning process as well as new unseen videos. Extensive experiments on three challenging and diverse datasets well demonstrate that our proposed approach generates superior quality videos compared to the existing state-of-the-art methods. Abhishek Aich, Akash Gupta 0001, Rameswar Panda, Rakib Hyder, Muhammad Salman Asif, Amit K. Roy-Chowdhury |
CVPR | 5 |
| 2020 | Solving Phase Retrieval with a Learned Reference
Rakib Hyder, Zikui Cai, Muhammad Salman Asif |
ECCV (30) | 3 |
| 2020 | Fourier Phase Retrieval with Arbitrary Reference SignalabstractFourier phase retrieval problem aims at recovering a signal from its Fourier amplitude measurements. A good initialization and prior information about the sparsity or support of the target signal is critical for robust recovery. Holographic phase retrieval is a related problem in which the presence of a reference signal makes the signal recovery problem linear. The existing methods, however, only work if the support of the reference and target signals are sufficiently separated. In this paper, we present a Fourier phase retrieval algorithm in the presence of a known (reference) signal at arbitrary location in the scene. We assume that a small part of the target signal is known without any other assumption about the support and separation of the known and unknown parts of the signal. Our recovery algorithm is based on alternating minimization and gradient descent. We demonstrate that our proposed method significantly improves the Fourier phase retrieval for natural images and synthetic images with multiple objects. Fahimeh Arab, Muhammad Salman Asif |
ICASSP | 2 |
| 2020 | Low-Rank Tensor Ring Model for Completing Missing Visual DataabstractLow rank tensor factorization can be viewed as a higher order generalization of low-rank matrix factorization, both of which have been used for image and video representation and reconstruction from compressive measurements. In this paper, we present an algorithm for recovering low-rank tensors from massively under-sampled or missing data. We use low-rank tensor ring (TR) factorization to model images and videos. We observed that TR factorization models are robust to random missing entries but they fail in the cases when large blocks or slices of data are missing. We developed the following two types of algorithms to fill large missing blocks: An algorithm that incrementally updates the tensor rank and implicitly enforces correlations among different modes of the tensor. A framework to incorporate information about similarities between different modes of the tensor to enforce explicit similarity constraints between the missing and known parts of the tensors. We present simulation experiments on YaleB dataset to demonstrate the performance of our methods. Muhammad Salman Asif, Ashley Prater-Bennette |
ICASSP | 1 |
| 2020 | Coded Illumination and Multiplexing for Lensless ImagingabstractMask-based lensless cameras offer an alternative option to conventional cameras. Compared to conventional cameras, lensless cameras can be extremely thin, flexible, and lightweight. Despite these advantages, the quality of images recovered from the lensless cameras is often poor because of the ill-conditioning of the underlying linear system. In this paper, we propose a new method to address the problem of illconditioning by combining coded illumination patterns with the mask-based lensless imaging. We assume that the object is illuminated with multiple binary patterns and the camera acquires a sequence of images for different illumination patterns. We propose a low-complexity, recursive algorithm that avoids storing all the images or creating a large system matrix. We present simulation results on standard test images under various extreme conditions and demonstrate that the quality of the image improves significantly with a small number of illumination patterns. Yucheng Zheng, Rongjia Zhang, Muhammad Salman Asif |
ICASSP | 3 |
| 2020 | SweepCam - Depth-Aware Lensless Imaging Using Programmable MasksabstractLensless cameras, while extremely useful for imaging in constrained scenarios, struggle with resolving scenes with large depth variations. To resolve this, we propose imaging with a set of mask patterns displayed on a programmable mask, and introduce a computational focusing operator that helps to resolve the depth of scene points. As a result, the proposed imager can resolve dense scenes with large depth variations, allowing for more practical applications of lensless cameras. We also present a fast reconstruction algorithm for scene at multiple depths that reduces reconstruction time by two orders of magnitude. Finally, we build a prototype to show the proposed method improves both image quality and depth resolution of lensless cameras. Shigeki Nakamura, Muhammad Salman Asif, Aswin C. Sankaranarayanan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Alternating Phase Projected Gradient Descent with Generative Priors for Solving Compressive Phase RetrievalabstractThe classical problem of phase retrieval arises in various signal acquisition systems. Due to the ill-posed nature of the problem, the solution requires assumptions on the structure of the signal. In the last several years, sparsity and support-based priors have been leveraged successfully to solve this problem. In this work, we propose replacing the sparsity/support priors with generative priors and propose two algorithms to solve the phase retrieval problem. Our proposed algorithms combine the ideas from AltMin approach for non-convex sparse phase retrieval and projected gradient descent approach for solving linear inverse problems using generative priors. We empirically show that the performance of our method with projected gradient descent is superior to the existing approach for solving phase retrieval under generative priors. We support our method with an analysis of sample complexity with Gaussian measurements. Rakib Hyder, Viraj Shah, Chinmay Hegde, Muhammad Salman Asif |
ICASSP | 4 |
| 2019 | Exactly Decoding a Vector through Relu ActivationabstractWe consider learning a d-dimensional parameter w through nonlinear input/output relation governed by ReLU activation. We study a supervised learning setup in which we want to decode w from input/output pairs (x, y). We consider an additive model with nonlinear ReLU activation that can be represented as y = Σk=1dReLU(w[k] + x[k]). Such a model appears in representation learning and recommendation systems where w corresponds to an unknown embedding of a user or item and the x correspond to embedding of known probe vectors. In this paper, we show that a gradient descent algorithm linearly converges with O(d) samples and quickly finds the true parameter w under mild assumptions. Our assumptions are in terms of the input distribution that captures the fundamentals of the problems. We also demonstrate the performance of our algorithm with numerical simulations. Samet Oymak, Muhammad Salman Asif |
ICASSP | 2 |
| 2019 | Multilinear Compressive Sensing With Tensor Ring FactorizationabstractTensor factorization has become a powerful tool for representation and analysis of multi-dimensional data. Low rank tensor factorization can be viewed as a higher order generalization of low-rank matrix factorization, both of which have been used for image and video representation and reconstruction from compressive measurements. In this paper, we present an algorithm for reconstructing images and videos from compressive measurements using tensor ring factorization model. We use a projected gradient descent approach that alternates between gradient descent over a loss function and projection onto a low-rank tensor ring structure. We also present a computationally efficient initialization step for the special case of multilinear compressive sensing. We present simulation results to demonstrate the performance of our algorithm on real images and videos. Muhammad Salman Asif, Ashley Prater-Bennette |
ICIP | 1 |
| 2018 | Lensless 3D Imaging Using Mask-Based CamerasabstractRecently, coded masks have been used to demonstrate a thin form-factor lensless camera, FlatCam, in which a mask is placed immediately on top of a bare image sensor. In this paper, we present an imaging model and algorithm to jointly estimate depth and intensity information in the scene from a single or multiple FlatCams. We use a light field representation to model the mapping of 3D scene onto the sensor in which light rays from different depths yield different modulation patterns. We present a greedy depth pursuit algorithm to search the 3D volume and estimate the depth and intensity of each pixel within the camera field-of-view. We present simulation results to analyze the performance of our proposed model and algorithm with different FlatCam settings. Muhammad Salman Asif |
ICASSP | 1 |
| 2015 | FPA-CS: Focal plane array-based compressive imaging in short-wave infraredabstractCameras for imaging in short and mid-wave infrared spectra are significantly more expensive than their counterparts in visible imaging. As a result, high-resolution imaging in those spectrum remains beyond the reach of most consumers. Over the last decade, compressive sensing (CS) has emerged as a potential means to realize inexpensive short-wave infrared cameras. One approach for doing this is the single-pixel camera (SPC) where a single detector acquires coded measurements of a high-resolution image. A computational reconstruction algorithm is then used to recover the image from these coded measurements. Unfortunately, the measurement rate of a SPC is insufficient to enable imaging at high spatial and temporal resolutions. We present a focal plane array-based compressive sensing (FPA-CS) architecture that achieves high spatial and temporal resolutions. The idea is to use an array of SPCs that sense in parallel to increase the measurement rate, and consequently, the achievable spatio-temporal resolution of the camera. We develop a proof-of-concept prototype in the short-wave infrared using a sensor with 64× 64 pixels; the prototype provides a 4096× increase in the measurement rate compared to the SPC and achieves a megapixel resolution at video rate using CS techniques. Huaijin G. Chen, Muhammad Salman Asif, Aswin C. Sankaranarayanan, Ashok Veeraraghavan |
CVPR | 2 |
| 2011 | Estimation and dynamic updating of time-varying signals with sparse variationsabstractThis paper presents an algorithm for an ℓ1-regularized Kalman filter. Given observations of a discrete-time linear dynamical system with sparse errors in the state evolution, we estimate the state sequence by solving an optimization algorithm that balances fidelity to the measurements (measured by the standard ℓ2norm) against the sparsity of the innovations (measured using the ℓ1norm). We also derive an efficient algorithm for updating the estimate as the system evolves. This dynamic updating algorithm uses a homotopy scheme that tracks the solution as new measurements are slowly worked into the system and old measurements are slowly removed. The effective cost of adding new measurements is a number of low-rank updates to the solution of a linear system of equations that is roughly proportional to the joint sparsity of all the innovations in the time interval of interest. Muhammad Salman Asif, Adam S. Charles, Justin K. Romberg, Christopher J. Rozell |
ICASSP | 1 |
| 2010 | Streaming Compressive Sensing for high-speed periodic videosabstractThe ability of Compressive Sensing (CS) to recover sparse signals from limited measurements has been recently exploited in computational imaging to acquire high-speed periodic and near-periodic videos using only a low-speed camera with coded exposure and intensive off-line processing. Each low-speed frame integrates a coded sequence of high-speed frames during its exposure time. The high-speed video can be reconstructed from the low-speed coded frames using a sparse recovery algorithm. This paper presents a new streaming CS algorithm specifically tailored to this application. Our streaming approach allows causal on-line acquisition and reconstruction of the video, with a small, controllable, and guaranteed buffer delay and low computational cost. The algorithm adapts to changes in the signal structure and, thus, outperforms the off-line algorithm in realistic signals. Muhammad Salman Asif, Dikpal Reddy, Petros Boufounos, Ashok Veeraraghavan |
ICIP | 1 |