VLDB 2026 Research / reviewers in the wild / expert
Muhammad A. Shah
dblp:142/5481
· DBLP profile ↗
10ranked-venue papers
9as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Revisiting Acoustic Features for Robust ASRabstractAutomatic Speech Recognition (ASR) systems must be robust to the myriad types of noises present in real-world environments including environmental noise, room impulse response, special effects as well as attacks by malicious actors (adversarial attacks). Recent works seek to improve accuracy and robustness by developing novel Deep Neural Networks (DNNs) and curating diverse training datasets for them, while using relatively simple acoustic features. While this approach improves robustness to the types of noise present in the training data, it confers limited robustness against unseen noises and negligible robustness to adversarial attacks. In this paper, we revisit the approach of earlier works that developed acoustic features inspired by biological auditory perception that could be used to perform accurate and robust ASR. In contrast, Specifically, we evaluate the ASR accuracy and robustness of several biologically inspired acoustic features. In addition to several features from prior works, such as gammatone filterbank features (GammSpec), we also propose two new acoustic features called frequency masked spectrogram (FreqMask) and difference of gammatones spectrogram (DoGSpec) to simulate the neuro-psychological phenomena of frequency masking and lateral suppression. Experiments on diverse models and datasets show that (1) DoGSpec achieves significantly better robustness than the highly popular log mel spectrogram (LogMelSpec) with minimal accuracy degradation, and (2) GammSpec achieves better accuracy and robustness to non-adversarial noises from the Speech Robust Bench benchmark, but it is outperformed by DoGSpec against adversarial attacks. Muhammad A. Shah, Bhiksha Raj |
ICASSP | 1 |
| 2025 | Speech Robust Bench: A Robustness Benchmark For Speech RecognitionabstractAs Automatic Speech Recognition (ASR) models become ever more pervasive, it is important to ensure that they make reliable predictions under corruptions present in the physical and digital world. We propose Speech Robust Bench (SRB), a comprehensive benchmark for evaluating the robustness of ASR models to diverse corruptions. SRB is composed of 114 input perturbations which simulate an heterogeneous range of corruptions that ASR models may encounter when deployed in the wild. We use SRB to evaluate the robustness of several state-of-the-art ASR models and observe that model size and certain modeling choices such as the use of discrete representations, or self-training appear to be conducive to robustness. We extend this analysis to measure the robustness of ASR models on data from various demographic subgroups, namely English and Spanish speakers, and males and females. Our results revealed noticeable disparities in the model's robustness across subgroups. We believe that SRB will significantly facilitate future research towards robust ASR models, by making it easier to conduct comprehensive and comparable robustness evaluations. Muhammad A. Shah, David Solans Noguero, Mikko A. Heikkilä, Bhiksha Raj, Nicolas Kourtellis |
ICLR | 1 |
| 2024 | Fixed Inter-Neuron Covariability Induces Adversarial RobustnessabstractVulnerability to adversarial perturbations is a major flaw of Deep Neural Networks (DNNs) that brings their reliability in real-world scenarios into question. However, biological perception, which DNNs emulate, is highly robust to such perturbations. This suggests that DNNs diverge from biological perception and develop vulnerability to adversarial perturbations. One point of divergence is that the activity of biological neurons is correlated and the structure of this correlation tends to be static across stimuli and over time, even if it hampers performance and learning. Conversely, we observe that small perturbations of the inputs can significantly change the inter-neuron correlations. We hypothesize that constraining the DNN’s neurons to respect a fixed inter-neuron correlation structure would improve its robustness. To test this hypothesis, we develop a Self-Consistent Activation (SCA) layer, which consists of neurons whose activations are optimized to conform to a fixed, but learned, correlation pattern. We train models with our SCA layer on image and audio recognition tasks and evaluate their accuracy under AutoAttack, an ensemble of white and black box adversarial attacks. Models with an SCA layer achieved up to 6% (abs) higher accuracy than traditional MLPs, without being trained on adversarially perturbed data. The code is available at https://github.com/ahmedshah1494/SCA-Layer. Muhammad A. Shah, Bhiksha Raj |
ICASSP | 1 |
| 2023 | Training on Foveated Images Improves Robustness to Adversarial AttacksabstractDeep neural networks (DNNs) have been shown to be vulnerable to adversarial attacks
-- subtle, perceptually indistinguishable perturbations of inputs that change the response of the model. In the context of vision, we hypothesize that an important contributor to the robustness of human visual perception is constant exposure to low-fidelity visual stimuli in our peripheral vision. To investigate this hypothesis, we develop RBlur, an image transform that simulates the loss in fidelity of peripheral vision by blurring the image and reducing its color saturation based on the distance from a given fixation point. We show that compared to DNNs trained on the original images, DNNs trained on images transformed by RBlur are substantially more robust to adversarial attacks, as well as other, non-adversarial, corruptions, achieving up to 25% higher accuracy on perturbed data. Muhammad A. Shah, Aqsa Kashaf, Bhiksha Raj |
NeurIPS | 1 |
| 2021 | High-Frequency Adversarial Defense for Speech and AudioabstractRecent work suggests that adversarial examples are enabled by high-frequency components in the dataset. In the speech domain where spectrograms are used extensively, masking those components seems like a sound direction for defenses against attacks. We explore a smoothing approach based on additive noise masking in priority high frequencies. We show that this approach is much more robust than the naive noise filtering approach, and a promising research direction. We successfully apply our defense on a Librispeech speaker identification task, and on the UrbanSound8K audio classification dataset. Raphaël Olivier, Bhiksha Raj, Muhammad A. Shah |
ICASSP | 3 |
| 2021 | Towards Adversarial Robustness Via Compact Feature RepresentationsabstractDeep Neural Networks (DNNs), while providing state-of-the-art performance in a wide variety of tasks, have been shown to be vulnerable to adversarial attacks. Recent studies have posited that this vulnerability arises because DNNs operate over a grossly overspecified input space with very sparse human supervision due to which they tend to learn spurious features that humans would ignore. These spurious features provide an attack vector for the adversary because perturbing these features would not alter the human’s decision but may alter the model’s prediction. In this paper we explore hypothesis that reducing the size of the model’s feature representation while maintaining its generalizability would discard spurious features while retaining perceptually relevant ones. We find that after the size of the feature representation has been reduced the models exhibit increased adversarial robustness, while suffering only a minimal loss in accuracy. In addition to being more robust, models with compact feature representations have the benefit of being more resource efficient. Muhammad A. Shah, Raphaël Olivier, Bhiksha Raj |
ICASSP | 1 |
| 2021 | Evaluating the Vulnerability of End-to-End Automatic Speech Recognition Models to Membership Inference Attacks
Muhammad A. Shah, Joseph Szurley, Athanasios Mouchtaris, Jasha Droppo |
Interspeech | 1 |
| 2020 | Optimal Strategies For Comparing Covariates To Solve Matching ProblemsabstractMany machine learning tasks can be posed as matching problems in which we are given a “probe” entry that we expect matches some of the entries in our “gallery”. The general solution to these problems is to retrieve matching entries based on statistical dependencies between the probe and the gallery data that are learned using complex models. Often, however, there are other common covariates to the probe and gallery data which might be easily inferred and may explain some of the statistical dependencies between the two. In this paper we present a probabilistic framework to derive optimal matching strategies based only on covariate features for three broad tasks, namely N-way classification, pairwise verification and ranking. We use canonical metrics to determine the maximum performance that can be expected if only covariate features are used and determine the marginal gain of using complex models. We find that covariate matching achieves an EER within 10% of a CNN in the verification task, and an MAP within 22% of the a DNN based model in the ranking task. Muhammad A. Shah, Raphaël Olivier, Bhiksha Raj |
ICPR | 1 |
| 2018 | Hitting Three Birds with One System: A Voice-Based CAPTCHA for the Modern UserabstractCAPTCHA challenges are used all over the Internet to prevent automated scripts from spamming web services. However, recent technological developments have rendered the conventional CAPTCHA insecure and inconvenient to use. In this paper, we propose vCAPTCHA, a voice-based CAPTCHA system that would: (1) enable more secure human authentication, (2) more conveniently integrate with modern devices accessing web services, and (3) help collect vast amounts of annotated speech data for different languages, accents, and dialects that are under-represented in the current speech corpora, thus making speech technologies accessible to more people around the world. vCAPTCHA requires users to speak their responses, in order to unlock or use different web services, instead of typing them. These user responses are analyzed to determine if they were indeed naturally produced, and transcribed to ensure that they contain the challenge sentence. We build a prototype for vCAPTCHA in order to assess its performance and practicality. Our preliminary results show that we are able to achieve an attack success rate as low as 2.3% while maintaining a human success rate comparable to current CAPTCHAs, on ASVspoof datasets. Muhammad A. Shah, Khaled A. Harras |
ICWS | 1 |
| 2014 | ChromoHub V2: cancer genomicsabstractSUMMARY: Cancer genomics data produced by next-generation sequencing support the notion that epigenetic mechanisms play a central role in cancer. We have previously developed Chromohub, an open access online interface where users can map chemical, structural and biological data from public repositories on phylogenetic trees of protein families involved in chromatin mediated-signaling. Here, we describe a cancer genomics interface that was recently added to Chromohub; the frequency of mutation, amplification and change in expression of chromatin factors across large cohorts of cancer patients is regularly extracted from The Cancer Genome Atlas and the International Cancer Genome Consortium and can now be mapped on phylogenetic trees of epigenetic protein families. Explorators of chromatin signaling can now easily navigate the cancer genomics landscape of writers, readers and erasers of histone marks, chromatin remodeling complexes, histones and their chaperones. AVAILABILITY AND IMPLEMENTATION: http://www.thesgc.org/chromohub/. Muhammad A. Shah, Remi Denton, Matthieu Schapira |
Bioinform. | 1 |