Hafiz Malik

dblp:74/3256 · DBLP profile ↗
← Back
48ranked-venue papers
14as first author
16since 2021 · last 2026
0000-0001-6006-3888ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 10 first-author · 3 since 2021Security and privacy · 14 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 11 · 7 since 2021Systems, architecture and hardware · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Lightweight deep learning framework with feature fusion for multiclass classification of chronic wounds
Mustafa Alhababi, Muteb Aljasem, Zaid A. Aldoulah, Gregory Auner, Hafiz Malik, Lubna K. Alazzawi
Eng. Appl. Artif. Intell.5
2025 Multilingual Dataset Integration Strategies for Robust Audio Deepfake Detection: A SAFE Challenge System
abstract
The SAFE Challenge evaluates synthetic speech detection across three tasks: unmodified audio, processed audio with compression artifacts, and laundered audio designed to evade detection. We systematically explore self-supervised learning (SSL) front-ends, training data compositions, and audio length configurations for robust deepfake detection. Our AASIST-based approach incorporates WavLM large frontend with RawBoost augmentation, trained on a multilingual dataset of 256,600 samples spanning 9 languages and over 70 TTS systems from Codec-Fake, MLAAD v5, SpoofCeleb, Famous Figures, and MAILABS. Through extensive experimentation with different SSL frontends, three training data versions, and two audio lengths, we achieved second place in both Task 1 (unmodified audio detection) and Task 3 (laundered audio detection), demonstrating strong generalization and robustness.
Hashim Ali 0003, Surya Subramani, Nithin Sai Adupa, Lekha Bollinani, Sali El-Loh, Hafiz Malik
ASRU6
2025 Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges
abstract
Recent advances in speech synthesis have introduced unprecedented challenges in maintaining voice authenticity, particularly concerning public figures who are frequent targets of impersonation attacks. This paper presents a comprehensive methodology for collecting, curating, and generating synthetic speech data for political figures and a detailed analysis of challenges encountered. We introduce a systematic approach incorporating an automated pipeline for collecting high-quality bonafide speech samples, featuring transcription-based segmentation that significantly improves synthetic speech quality. We experimented with various synthesis approaches; from single-speaker to zero-shot synthesis, and documented the evolution of our methodology. The resulting dataset comprises bonafide and synthetic speech samples from ten public figures, demonstrating superior quality with a NISQA-TTS naturalness score of 3.69 and the highest human misclassification rate of 61.9\%.
Hashim Ali 0003, Surya Subramani, Raksha Varahamurthy, Nithin Sai Adupa, Lekha Bollinani, Hafiz Malik
INTERSPEECH6
2025 Sniper: Countering Locker Ransomware Attacks Through Natural Language Processing
abstract
Mobile systems have evolved into versatile devices that end users depend on for carrying out their daily tasks. Unfortunately, the mobile sector has recently fallen prey to a series of ransomware campaigns designed to lock users out of their devices and extort them for payment. In response to these challenges, we propose a novel runtime system that dynamically restores device access by undoing the effects of locker ransomware. A key observation made by this work is that attackers rely on the display of a ransom note on the victim’s device to demand payment. Based on this observation, we develop a solution that combines the monitoring of mobile app activity with a natural language processing (NLP) unit that harnesses transformers to detect the appearance of ransom notes. We extensively validate the robustness of our solution against more than five thousand ransomware samples and show that our solution reliably recovers from all the malicious samples that we tested, including overlay screen and change pin-code ransomware. Finally, an evaluation of our proof-of-concept implementation shows minimal performance impact while running a mix of mobile benchmark applications.
Abdulrahman Abu Elkhail, Anys Bacha, Hafiz Malik
IEEE Trans. Dependable Secur. Comput.3
2024 Is Audio Spoof Detection Robust to Laundering Attacks?
abstract
Voice-cloning (VC) systems have seen an exceptional increase in the realism of synthesized speech in recent years. The high quality of synthesized speech and the availability of low-cost VC services have given rise to many potential abuses of this technology. Several detection methodologies have been proposed over the years that can detect voice spoofs with reasonably good accuracy. However, these methodologies are mostly evaluated on clean audio databases, such as ASVSpoof 2019. This paper evaluates SOTA Audio Spoof Detection approaches in the presence of laundering attacks. In that regard, a new laundering attack database, called the ASVSpoof Laundering Database, is created. This database is based on the ASVSpoof 2019 (LA) eval database comprising a total of 1388.22 hours of audio recordings. Seven SOTA audio spoof detection approaches are evaluated on this laundered database. The results indicate that SOTA systems perform poorly in the presence of aggressive laundering attacks, especially reverberation and additive noise attacks. This suggests the need for robust audio spoof detection.
Hashim Ali 0003, Surya Subramani, Shefali Sudhir, Raksha Varahamurthy, Hafiz Malik
IH&MMSec5
2023 Generation of Time-Varying Impedance Attacks Against Haptic Shared Control Steering Systems
abstract
The safety-critical nature of vehicle steering is one of the main motivations for exploring the space of possible cyber-physical attacks against the steering systems of modern vehicles. This paper investigates the adversarial capabilities for destabilizing the interaction dynamics between human drivers and vehicle haptic shared control (HSC) steering systems. In contrast to the conventional robotics literature, where the main objective is to render the human-automation interaction dynamics stable by ensuring passivity, this paper takes the exact opposite route. In particular, to investigate the damaging capabilities of a successful cyber-physical attack, this paper demonstrates that an attacker who targets the HSC steering system can destabilize the interaction dynamics between the human driver and the vehicle HSC steering system through synthesis of time-varying impedance profiles. Specifically, it is shown that the adversary can utilize a properly designed non-passive and time-varying adversarial impedance target dynamics, which are fed with a linear combination of the human driver and the steering column torques. Using these target dynamics, it is possible for the adversary to generate in realtime a reference angular command for the driver input device and the directional control steering assembly of the vehicle. Furthermore, it is shown that the adversary can make the steering wheel and the vehicle steering column angular positions to follow the reference command generated by the time-varying impedance target dynamics using proper adaptive control strategies. Numerical simulations demonstrate the effectiveness of such time-varying impedance attacks, which result in a non-passive and inherently unstable interaction between the driver and the HSC steering system.
Alireza Mohammadi 0001, Hafiz Malik
IROS2
2023 Deepfakes generation and detection: state-of-the-art, open challenges, countermeasures, and way forward
Momina Masood, Marriam Nawaz, Khalid Mahmood 0003, Ali Javed, Aun Irtaza, Hafiz Malik
Appl. Intell.6
2023 Seamlessly Safeguarding Data Against Ransomware Attacks
abstract
Encryption has become an indispensable technology for preserving confidentiality. Unfortunately, cybercriminals have re-purposed this technology to deny users access to their data. This trend has sparked an onslaught of ransomware attacks, that resulted in several victims being extorted to pay ransoms in return for restoring their maliciously encrypted data. In response to these challenges, we propose a novel runtime solution that seamlessly defends against cryptographic ransomware. A key observation made by this work is that maliciously encrypted data is initially buffered in the OS's page cache before it is flushed to the underlying storage device. Based on this observation, we develop a solution that efficiently manages data synchronization between the memory and storage subsystems to prevent maliciously encrypted data from being permanently committed to the underlying storage. We extensively validate the robustness of this approach against more than one thousand ransomware samples and show that our design reliably restores all encrypted files. Furthermore, our solution is resilient to ransomware that employ techniques including master boot record infection and multi-threaded attacks. Finally, an evaluation of our proof-of-concept implementation shows minimal performance impact while running a mix of compute and I/O bound applications.
Abdulrahman Abu Elkhail, Nada Lachtar, Duha Ibdah, Rustam Aslam, Anys Bacha, Hafiz Malik
IEEE Trans. Dependable Secur. Comput.7
2023 Vehicle Lateral Motion Dynamics Under Braking/ABS Cyber-Physical Attacks
abstract
In face of an increasing number of automotive cyber-physical threat scenarios, the issue of adversarial destabilization of the lateral motion of target vehicles through direct attacks on their steering systems has been extensively studied. A more subtle question is whether a cyberattacker can destabilize the target vehicle lateral motion through improper engagement of the vehicle brakes and/or anti-lock braking systems (ABS). Motivated by such a question, this paper investigates the impact of cyber-physical attacks that exploit the braking/ABS systems to adversely affect the lateral motion stability of the targeted vehicles. Using a hybrid physical/dynamic tire-road friction model, it is shown that if a braking system/ABS attacker manages to continuously vary the longitudinal slips of the wheels, they can violate the necessary conditions for asymptotic stability of the underlying linear time-varying (LTV) dynamics of the lateral motion. Furthermore, the minimal perturbations of the wheel longitudinal slips that result in lateral motion instability under fixed slip values are derived. Finally, a real-time algorithm for monitoring the lateral motion dynamics of vehicles against braking/ABS cyber-physical attacks is devised. This algorithm, which can be efficiently computed using the modest computational resources of automotive embedded processors, can be utilized along with other intrusion detection techniques to infer whether a vehicle braking system/ABS is experiencing a cyber-physical attack. Numerical simulations in the presence of realistic CAN bus delays, destabilizing slip value perturbations obtained from solving quadratic programs on an embedded ARM Cortex-M3 emulator, and side-wind gusts demonstrate the effectiveness of the proposed methodology.
Alireza Mohammadi 0001, Hafiz Malik, Masoud Abbaszadeh
IEEE Trans. Inf. Forensics Secur.2
2022 Distilling Facial Knowledge with Teacher-Tasks: Semantic-Segmentation-Features For Pose-Invariant Face-Recognition
abstract
This paper demonstrates a novel approach to improve face-recognition pose-invariance using semantic-segmentation features. The proposed Seg-Distilled-ID network jointly learns identification and semantic-segmentation tasks, where the segmentation task is then "distilled" (MobileNet en-coder). Performance is benchmarked against three state-of-the-art encoders on a publicly available data-set emphasizing head-pose variations. Experimental evaluations show the Seg-Distilled-ID network shows notable robustness benefits, achieving 99.9% test-accuracy in comparison to 81.6% on ResNet-101, 96.1% on VGG-19 and 96.3% on InceptionV3. This is achieved using approximately one-tenth of the top encoder’s inference parameters. These results demonstrate distilling semantic-segmentation features can efficiently ad-dress face-recognition pose-invariance.
Ali Hassani 0005, Zaid A. El-Shair, Rafi Ud Daula Refat, Hafiz Malik
ICIP4
2022 Voice spoofing detector: A unified anti-spoofing framework
Ali Javed, Khalid Mahmood 0003, Hafiz Malik, Aun Irtaza
Expert Syst. Appl.3
2022 Defocus blur detection using novel local directional mean patterns (LDMP) and segmentation via KNN matting
Awais Khan 0007, Aun Irtaza, Ali Javed, Tahira Nazir, Hafiz Malik, Khalid Mahmood 0003, Muhammad Ammar Khan
Frontiers Comput. Sci.5
2022 Graph-Based Intrusion Detection System for Controller Area Networks
abstract
The controller area network (CAN) is the most widely used intra-vehicular communication network in the automotive industry. Because of its simplicity in design, it lacks most of the requirements needed for a security-proven communication protocol. However, a safe and secured environment is imperative for autonomous as well as connected vehicles. Therefore CAN security is considered one of the important topics in the automotive research community. In this article, we propose a four-stage intrusion detection system that uses the chi-squared method and can detect any kind of strong and weak cyber attacks in a CAN. This work is the first-ever graph-based defense system proposed for the CAN. Our experimental results show that we have a very low 5.26% misclassification for denial of service (DoS) attack, 10% misclassification for fuzzy attack, 4.76% misclassification for replay attack, and no misclassification for spoofing attack. In addition, the proposed methodology exhibits up to 13.73% better accuracy compared to existing ID sequence-based methods.
Riadul Islam, Rafi Ud Daula Refat, Sai Manikanta Yerram, Hafiz Malik
IEEE Trans. Intell. Transp. Syst.4
2021 An Application Agnostic Defense Against the Dark Arts of Cryptojacking
abstract
The popularity of cryptocurrencies has garnered interest from cybercriminals, spurring an onslaught of cryptojacking campaigns that aim to hijack computational resources for the purpose of mining cryptocurrencies. In this paper, we present a cross-stack cryptojacking defense system that spans the hardware and OS layers. Unlike prior work that is confined to detecting cryptojacking behavior within web browsers, our solution is application agnostic. We show that tracking instructions that are frequently used in cryptographic hash functions serve as reliable signatures for fingerprinting cryptojacking activity. We demonstrate that our solution is resilient to multi-threaded and throttling evasion techniques that are commonly employed by cryptojacking malware. We characterize the robustness of our solution by extensively testing a diverse set of workloads that include real consumer applications. Finally, an evaluation of our proof-of-concept implementation shows minimal performance impact while running a mix of benchmark applications.
Nada Lachtar, Abdulrahman Abu Elkhail, Anys Bacha, Hafiz Malik
DSN4
2021 Voice spoofing detection corpus for single and multi-order audio replays
Roland Baumann, Khalid Mahmood 0003, Ali Javed, Andersen Ball, Brandon Kujawa, Hafiz Malik
Comput. Speech Lang.6
2021 Secure Automatic Speaker Verification (SASV) System Through sm-ALTP Features and Asymmetric Bagging
abstract
The growing number of voice-enabled devices and applications consider automatic speaker verification (ASV) a fundamental component. However, maximum outreach for ASV in critical domains e.g., financial services and health care, is not possible unless we overcome security breaches caused by voice cloning algorithms and replayed audios. Therefore, to overcome these vulnerabilities, a secure ASV (SASV) system based on the novel sign modified acoustic local ternary pattern (sm-ALTP) features and asymmetric bagging-based classifier-ensemble with enhanced attack vector is presented. The proposed audio representation approach clusters the high and low frequency components in audio frames by normally distributing frequency components against a convex function. Then, the neighborhood statistics are applied to capture the user specific vocal tract information. The proposed SASV system simultaneously verifies the bonafide speakers and detects the voice cloning attack, cloning algorithm used to synthesize cloned audio (in the defined settings), and voice-replay attacks over the ASVspoof 2019 dataset. In addition, the proposed method detects the voice replay and cloned voice replay attacks over the VSDC dataset. Both the voice cloning algorithm detection and cloned-replay attack detection are novel concepts introduced in this paper. The voice cloning algorithm detection module determines the voice cloning algorithm used to generate the fake audios. Whereas, the cloned voice replay attack detection is performed to determine the SASV behavior when audio samples are simultaneously contemplated with cloning and replay artifacts.
Muteb Aljasem, Aun Irtaza, Hafiz Malik, Noushin Saba, Ali Javed, Khalid Mahmood 0003, Mohammad Meharmohammadi
IEEE Trans. Inf. Forensics Secur.3
2020 Dark Firmware: A Systematic Approach to Exploring Application Security Risks in the Presence of Untrusted Firmware
Duha Ibdah, Nada Lachtar, Abdulrahman Abu Elkhail, Anys Bacha, Hafiz Malik
RAID5
2020 Sensor Data Integrity Verification for Autonomous Vehicles Using Spread 3D Dither QIM
abstract
An autonomous vehicle is a modern-day cyber-physical system with multiple sensors connected over its internal network. Many crucial features of the autonomy depend on the data from these sensors that get transmitted over the vehicle network to reach the data processing central modules. Protecting and verifying the sensor data integrity as it traverses the vehicle network is essential for the proper functioning of the autonomous vehicle. We propose a novel approach called spread dither 3D quantization index modulation (QIM) to verify the integrity of LiDAR sensor data used in autonomous vehicle applications. We propose a framework to verify the sensor data integrity in time-critical autonomous vehicle applications based on a data hiding technique called dither modulation. The sensor data is embedded with a watermark using the proposed spread 3D dither modulation at the sensor domain. The embedded sensor data can be directly used by the data processing algorithms due to very low embedded induced distortion. The sensor data integrity can be verified by a parallel process that can decode the embedded sensor data to detect and localize the tampering. Our proposed countermeasure framework is verified with real-life LiDAR data frames against the simulated transmission layer insider attacks on sensor data such as object insertion and deletion. The dither QIM method's optimum parameter values are also deduced through multiple experiments.
Raghu Changalvala, Hafiz Malik
VTC Fall2
2020 A decision tree framework for shot classification of field sports videos
Ali Javed, Khalid Mahmood 0003, Aun Irtaza, Hafiz Malik
J. Supercomput.4
2019 Replay and key-events detection for sports video summarization using confined elliptical local ternary patterns and extreme learning machine
Ali Javed, Aun Irtaza, Yasmeen Khaliq, Hafiz Malik, Muhammad Tariq Mahmood
Appl. Intell.4
2019 Multimodal framework based on audio-visual features for summarisation of cricket videos
abstract
Sports broadcasters generate an enormous amount of video content on the cyberspace due to massive viewership all over the world. Analysis and consumption of this huge repository urges the broadcasters to apply video summarisation to extract the exciting segments from the entire video to capture user's interest and reap the storage and transmission benefits. Therefore, in this study an automatic method for key‐events detection and summarisation based on audio‐visual features is presented for cricket videos. Acoustic local binary pattern features are used to capture excitement level in the audio stream, which is used to train a binary support vector machine (SVM) classifier. Trained SVM classifier is used to label audio frame as an excited or non‐excited frame. Excited audio frames are used to select candidate key‐video frames. A decision tree‐based classifier is trained to detect key‐events in the input cricket videos that are then used for video summarisation. Performance of the proposed framework has been evaluated on a diverse dataset of cricket videos belonging to different tournaments and broadcasters. Experimental results indicate that the proposed method achieves an average accuracy of 95.5%, which signifies its effectiveness.
Ali Javed, Aun Irtaza, Hafiz Malik, Muhammad Tariq Mahmood, Syed Muhammad Adnan Shah
IET Image Process.3
2018 Digital multimedia audio forensics: past, present and future
Mohammed Zakariah, Muhammad Khurram Khan, Hafiz Malik
Multim. Tools Appl.3
2017 Autonomous Decentralized Privacy-Enabled Data Preparation Architecture for Multicenter Clinical Observational Research
abstract
Tailoring treatment and clinical decision making to a person's unique characteristics is the next milestone for healthcare informatics, but for it to be accomplished, big data analytics for identifying risk factors and other hidden patterns among patients become paramount. In future these analytics will take the form of multicenter observational research, for which data preparation is vital. Specifically, quality data must be obtained in a timely manner while protecting the privacy of patients in the health records shared among researchers. Furthermore, the coordination and cooperation of a fluctuating number of medical data sources containing these records for clinical data distribution is an additional requirement in multicenter studies. Thus, we propose an autonomous decentralized, privacy-enabled data preparation architecture and novel SEDTM algorithm to meet these requirements, censuring sensitive information via filtration, and extracting relevant clinical data with a fully automated approach. Our evaluation demonstrates a 40% - 60% increase in the retrieval of quality patient data, compared to traditional semantic similarity, for our proposed SEDTM algorithm.
Khalid Mahmood 0003, Varun Sathyan, Hisham Kanaan, Ghaus M. Malik, Hafiz Malik
ISADS5
2017 Audio splicing detection and localization using environmental signature
Yifan Chen 0001, Rui Wang 0007, Hafiz Malik
Multim. Tools Appl.4
2016 Joint-channel modeling to attack QIM steganography
Hafiz Malik, K. P. Subbalakshmi, Rajarathnam Chandramouli
Multim. Tools Appl.1
2016 An Efficient Framework for Automatic Highlights Generation from Sports Videos
abstract
This letter presents a framework for replay detection in sports videos to generate highlights. For replay detection, the proposed work exploits the following facts: 1) broadcasters introduce gradual transition (GT) effect both at the start and at the end of a replay segment (RS), and 2) the absence of score captions (SCs) in an RS. The dual-threshold-based method is used to detect GT frames from the input video. A pair of successive GT frames is used to extract the candidate RSs. All frames in the selected segment are processed to detect SC. To this end, temporal running average is used to filter out temporal variations. First- and second-order statistics are used to binarize the running average image, which is fed to optical character recognition stage for character recognition. The absence/presence of SC is used for replay/live frame labeling. The SC detection stage complements the GT detection process, therefore, a combination of both is expected to result in superior computational complexity and detection accuracy. The performance of the proposed system is evaluated on 22 videos of four different sports (e.g., Cricket, tennis, baseball, and basketball). Experimental results indicate that the proposed method can achieve average detection accuracy ≥ 94.7%.
Ali Javed, Khalid Bashir Bajwa, Hafiz Malik, Aun Irtaza
IEEE Signal Process. Lett.3
2016 Anti-Forensics of Environmental-Signature-Based Audio Splicing Detection and Its Countermeasure via Rich-Features Classification
abstract
Numerous methods for detecting audio splicing have been proposed. Environmental-signature-based methods are considered to be the most effective forgery detection methods. The performance of existing audio forensic analysis methods is generally measured in the absence of any anti-forensic attack. Effectiveness of these methods in the presence of anti-forensic attacks is therefore unknown. In this paper, we propose an effective anti-forensic attack for environmental-signature-based splicing detection method and countermeasures to detect the presence of the anti-forensic attack. For anti-forensic attack, dereverberation-based processing is proposed. Three dereverberation methods are considered to tamper with the acoustic environment signature. Experimental results indicate that the proposed dereverberation-based anti-forensic attack significantly degrades the performance of the selected splicing detection method. The proposed countermeasures exploit artifacts introduced by the anti-forensic processing. To detect the presence of potential anti-forensic processing, a machine learning-based framework is proposed. In particular, the proposed anti-forensic detection method uses a rich-feature model consisting of Fourier coefficients, spectral properties, high-order statistics of musical noise residuals, and modulation spectral coefficients to capture traces of dereverberation attacks. The performance of the proposed framework is evaluated on both synthetic data and real-world speech recordings. The experimental results show that the proposed rich-feature model can detect the presence of anti-forensic processing with an average accuracy of 95%.
Yifan Chen 0001, Rui Wang 0007, Hafiz Malik
IEEE Trans. Inf. Forensics Secur.4
2016 Channel selection for simultaneous move game in cognitive radio ad hoc networks
Qurratul-Ain Minhas, Hasan Mahmood, Hafiz Malik
Wirel. Networks3
2014 Audio source authentication and splicing detection using acoustic environmental signature
abstract
Audio splicing is one of the most common manipulation techniques in the audio forensic world. In this paper, the magnitudes of acoustic channel impulse response and ambient noise are considered as the environmental signature and used to authenticate the integrity of query audio and identify the spliced audio segments. The proposed scheme firstly extracts the magnitudes of channel impulse response and ambient noise by applying the spectrum classification technique to each suspected frame. Then, correlation between the magnitudes of query frame and reference frame is calculated. An optimal threshold determined according to the statistical distribution of similarities is used to identify the spliced frames. Furthermore, a refining step using the relationship between adjacent frames is adopted to reduce the false positive rate and false negative rate. Effectiveness of the proposed method is tested on two data sets consisting of speech recordings of human speakers. Performance of the proposed method is evaluated for various experimental settings. Experimental results show that the proposed method not only detects the presence of spliced frames, but also localizes the forgery segments. Comparison results with previous work illustrate the superiority of the proposed scheme.
Yifan Chen 0001, Rui Wang 0007, Hafiz Malik
IH&MMSec4
2014 A new investment strategy based on data mining and Neural Networks
abstract
In this paper, we present a new investment strategy for optimal gains on investments in the stock market. Neural Network (NN)-based framework is used for trading prediction and forecasting. To this end, statistical measures based on return and volatility are used to filter out low performing sectors in the stock market. A simple but effective method based on price Simple Moving Averages (SMAs) is used to measure volatility for a given stock. The proposed NN-based system uses the strongest performing indices for stock market forecasting. In addition to predicting investment decisions such as Buy or Sell, the proposed framework also aims at maximizing investment gains (or returns). The proposed NN-based framework rely on historical data and provides investors investing strategies for optimal trading. Training data is extracted extracted from historical weekly data (from the Yahoo Finance). Simulation results indicate that the proposed framework can help investors making investment decisions and increasing their trading profitability.
Hafiz Malik
IJCNN2
2013 Unsupervised multimodal VAD using sequential hierarchy
abstract
In speech processing systems, the performance of the Voice Activity Detector (VAD) is a bottleneck to the whole system. Traditional VADs are solely based on acoustic features. Additional modality in form of visual information is used to make robust VADs. In this paper, we propose a multimodal VAD based on decision fusion between two modalities. Visual VAD (VVAD) decision vectors are interpolated so that logical operators can be applied to both modalities. In order to avoid this interpolation, we suggest a sequential arrangement of both subsystems to achieve a multimodal VAD. The proposed method considerably reduces false alarm rates when compared with performance of standalone audio VAD (AVAD).
Rameez Ahmad, Syed Paymaan Raza, Hafiz Malik
CIDM3
2013 Visual Speech Detection Using an Unsupervised Learning Framework
abstract
This paper presents an unsupervised learning framework for visual speech detection. Bimodal GMM is used to model visual features, i.e., mouth region intensity, which varies during speech. Variation in the mouth region intensity is used for visual speech and non-speech classification. The GMM parameters are estimated using the EM algorithm. Performance of the proposed algorithm is evaluated using a dataset consisting of 14 video clips containing almost 20, 000 frames. Performance of the proposed algorithm is also compared with existing state-of-the-art. Experimental results show that the proposed method achieves high detection and low false alarm rates.
Rameez Ahmad, Syed Paymaan Raza, Hafiz Malik
ICMLA (2)3
2013 Acoustic Environment Identification and Its Applications to Audio Forensics
abstract
An audio recording is subject to a number of possible distortions and artifacts. Consider, for example, artifacts due to acoustic reverberation and background noise. The acoustic reverberation depends on the shape and the composition of a room, and it causes temporal and spectral smearing of the recorded sound. The background noise, on the other hand, depends on the secondary audio source activities present in the evidentiary recording. Extraction of acoustic cues from an audio recording is an important but challenging task. Temporal changes in the estimated reverberation and background noise can be used for dynamic acoustic environment identification (AEI), audio forensics, and ballistic settings. We describe a statistical technique to model and estimate the amount of reverberation and background noise variance in an audio recording. An energy-based voice activity detection method is proposed for automatic decaying-tail-selection from an audio recording. Effectiveness of the proposed method is tested using a data set consisting of speech recordings. The performance of the proposed method is also evaluated for both speaker-dependent and speaker-independent scenarios.
Hafiz Malik
IEEE Trans. Inf. Forensics Secur.1
2013 Audio Recording Location Identification Using Acoustic Environment Signature
abstract
An audio recording is subject to a number of possible distortions and artifacts. Consider, for example, artifacts due to acoustic reverberation and background noise. The acoustic reverberation depends on the shape and the composition of the room, and it causes temporal and spectral smearing of the recorded sound. The background noise, on the other hand, depends on the secondary audio source activities present in the evidentiary recording. Extraction of acoustic cues from an audio recording is an important but challenging task. Temporal changes in the estimated reverberation and background noise can be used for dynamic acoustic environment identification (AEI), audio forensics, and ballistic settings. We describe a statistical technique based on spectral subtraction to estimate the amount of reverberation and nonlinear filtering based on particle filtering to estimate the background noise. The effectiveness of the proposed method is tested using a data set consisting of speech recordings of two human speakers (one male and one female) made in eight acoustic environments using four commercial grade microphones. Performance of the proposed method is evaluated for various experimental settings such as microphone independent, semi- and full-blind AEI, and robustness to MP3 compression. Performance of the proposed framework is also evaluated using Temporal Derivative-based Spectrum and Mel-Cepstrum (TDSM)-based features. Experimental results show that the proposed method improves AEI performance compared with the direct method (i.e., feature vector is extracted from the audio recording directly). In addition, experimental results also show that the proposed scheme is robust to MP3 compression attack.
Hafiz Malik
IEEE Trans. Inf. Forensics Secur.2
2012 Recording environment identification using acoustic reverberation
abstract
Acoustic environment leaves its fingerprints in the audio recording captured in it. Acoustic reverberation and background noise are generally used to characterize an acoustic environment. Acoustic reverberation depends on the shape and the composition of a room, therefore, differences in the estimated reverberation can be used in a forensic and ballistic settings and acoustic environment identification (AEI). We describe a framework that uses acoustic reverberation to characterize recording environment and use it for AEI. Inverse filtering is used to estimate the reverberation component from audio recording. A 48-dimensional feature vector consisting of Mel-frequency Cepstral Coefficients and Logarithmic Mel-spectral Coefficients is used to capture traces of reverberation. A multi-class support vector machine (SVM) classifier is used for AEI. Experimental results show that the proposed system can successfully identify a recording environment for regular as well as blind AEI.
Hafiz Malik
ICASSP1
2012 Nonparametric Steganalysis of QIM Steganography Using Approximate Entropy
abstract
This paper proposes an active steganalysis method for quantization index modulation (QIM)-based steganography. The proposed nonparametric steganalysis method uses irregularity (or randomness) in the test image to distinguish between the cover image and the stego image. We have shown that plain quantization (quantization without message embedding) induces regularity in the resulting quantized object, whereas message embedding using QIM increases irregularity in the resulting QIM-stego. Approximate entropy, an algorithmic entropy measure, is used to quantify irregularity in the test image. The QIM-stego image is then analyzed to estimate secret message length. To this end, the QIM codebook is estimated from the QIM-stego image using first-order statistics of the image coefficients in the embedding domain. The estimated codebook is then used to estimate secret message. Simulation results show that the proposed scheme can successfully estimate the hidden message from the QIM-stego with very low decoding error probability. For a given cover object the decoding error probability depends on embedding rate and decreases monotonically, approaching zero as the embedding rate approaches one.
Hafiz Malik, K. P. Subbalakshmi, Ramamurti Chandramouli
IEEE Trans. Inf. Forensics Secur.1
2010 Audio forensics from acoustic reverberation
abstract
An audio recording is subject to a number of possible distortions and artifacts. For example, the persistence of sound, due to multiple reflections from various surfaces in a room, causes temporal and spectral smearing of the recorded sound. This distortion is referred to as audio reverberation time. We describe a technique to model and estimate the amount of reverberation in an audio recording. Because reverberation depends on the shape and composition of a room, differences in the estimated reverberation can be used in a forensic and ballistic setting.
Hafiz Malik, Hany Farid
ICASSP1
2010 Digital audio forensics using background noise
abstract
This paper presents a new audio forensics method based on background noise in the audio signals. The traditional speech enhancement algorithms improve the quality of speech signals, however, existing methods leave traces of speech in the removed noise. Estimated noise using these existing methods contains traces of speech signal, also known as leakage signal. Although this speech leakage signal has low SNR, yet it can be perceived easily by listening to the estimated noise signal, it therefore cannot be used for audio forensics applications. For reliable audio authentication, a better noise estimation method is desirable. To achieve this goal, a two-step framework is proposed to estimate the background noise with minimal speech leakage signal. A correlation based similarity measure is then applied to determine the integrity of speech signal. The proposed method has been evaluated for different speech signals recorded in various environments. The results show that it performs better than the existing speech enhancement algorithms with significant improvement in terms of SNR value.
Sohaib Ikram, Hafiz Malik
ICME2
2010 Statistical modeling of footprints of QIM steganography
abstract
In this paper, a new model is proposed to characterize distortion due to message embedding. The proposed statistical model is used to develop a parametric steganalysis technique to attack quantization index modulation (QIM) steganography. We have shown that quantization with message embedding (a.k.a. QIM) introduces relatively stronger disturbance in the local-correlation in the test-image than quantization without message embedding. Presented steganalysis technique exploits rich spatial/temporal correlation in the natural images to estimate local-randomness in the test-image. The local-randomness estimated from the test-image is modeled using generalized Gamma distribution (GGD). A binary hypothesis test, based on generalized likelihood ratio test (GLRT), is used to detect the QIM-stego image. Simulation results show that the proposed method can successfully distinguish between the quantized-cover and the QIM-stego with very low false alarm rates.
Hafiz Malik
ICME1
2008 Commentary Paper 2 on "A Probabilistic Bayesian Framework for Model-Based Object Tracking Using Undecimated Wavelet Packet Descriptors"
abstract
This uses combination of corner-based model and coefficients of undecimated wavelet packet transform (UWPT) for the proposed probabilistic Bayesian framework for object tracking. The UWPT coefficients are calculated for patch around each corner. The proposed scheme uses local descriptors e.g. UWPT coefficients, to improve global representation of object shape model. The proposed scheme then estimates global position of the object using voting based on coherency among the model corners.
Hafiz Malik
AVSS1
2008 Commentary Paper 3 on Visual Players Detection and Tracking in Soccer Matches
abstract
This manuscript presents object detection and tracking in video clips of soccer matches. Background subtraction based framework on pixel energy is proposed to detect objects in the input video clip with varying light conditions, high frame rates, and real-time processing. Unsupervised clustering is then used classify detected objects into various classes. A stochastic approach based on maximum a posteriori probability (MAP) is proposed for object tracking.
Hafiz Malik
AVSS1
2008 Robust audio watermarking using frequency-selective spread spectrum
abstract
A novel audio watermarking scheme based on frequency-selective spread spectrum (FSSS) technique is presented. Unlike most of the existing spread spectrum (SS) watermarking schemes that use the entire audible frequency range for watermark embedding, the proposed scheme randomly selects subband(s) signal(s) of the host audio signal for watermark embedding. The proposed FSSS scheme provides a natural mechanism to exploit the band-dependent frequency-masking characteristics of the human auditory system to ensure the fidelity of the host audio signal and the robustness of the embedded information. Key attributes of the proposed scheme include reduced host interference in watermark detection, better fidelity, secure embedding and improved multiple watermark embedding capability. To detect the embedded watermark, two blind watermark detection methods are examined, one based on normalised correlation and the other based on estimation correlation. Extensive simulation results are presented to analyse the performance of the proposed scheme for various signal manipulations and standard benchmark attacks. A comparison with the existing full-band SS-based schemes is also provided to show the improved performance of the proposed scheme.
Hafiz Malik, Rashid Ansari, Ashfaq Khokhar 0001
IET Inf. Secur.1
2006 Blind Detection for Additive Embedding Using Underdetermined ICA
abstract
This paper presents an efficient blind watermark detection scheme for additive embedding (AE) based on underdetermined independent component analysis (ICA) framework. The proposed detector assumes that the host signal and the watermark obey non-Gaussian distributions and watermark embedding follows AE model. The proposed blind watermark detector employs blind source separation (BSS) for underdetermined mixtures for watermark estimation. Simulation results are presented showing that the proposed detector performs significantly better than existing correlation based blind detectors operating without suppressing the host signal interference at the detector.
Hafiz Malik, Ashfaq Khokhar 0001, Rashid Ansari, Marco Salvemini
ISM1
2005 Improved watermark detection for spread-spectrum based watermarking using independent component analysis
abstract
This paper presents an efficient blind watermark detection/decoding scheme for spread spectrum (SS) based watermarking, exploiting the fact that in SS-based embedding schemes the embedded watermark and the host signal are mutually independent and obey non-Gaussian distribution. The proposed scheme employs the theory of independent component analysis (ICA) and posed the watermark detection as a blind source separation problem. The proposed ICA-based blind detection/decoding scheme has been simulated using real-world audio clips. The simulation results show that the ICA-based detector can detect and decode watermark with extremely low decoding bit error probability (less than 0.01) against common watermarking attacks and benchmark degradations.
Hafiz Malik, Ashfaq Khokhar 0001, Rashid Ansari
Digital Rights Management Workshop1
2004 Data-hiding in audio using frequency-selective phase alteration
abstract
A novel perception-based data hiding technique for digital audio is proposed. It exploits the lower sensitivity of the human auditory system (HAS) to phase distortion in audio compared with magnitude distortion. Audio is decomposed into subband signals, some of which are selected for embedding data with a controlled alteration of phase using suitable allpass digital filters. The proposed scheme is robust to standard data manipulations yielding less than 2% error probability against compression, re-sampling, re-quantization, random chopping and noise addition. The proposed method is also robust to desynchronization attacks.
Rashid Ansari, Hafiz Malik, Ashfaq Khokhar 0001
ICASSP (5)2
2004 Robust audio watermarking using frequency selective spread spectrum theory
abstract
A new method is proposed for robust audio watermarking using direct-sequence spread spectrum in combination with the subband decomposition of the audio signal. The method exploits the frequency masking characteristics of the human auditory system (HAS) and inserts the watermark into a randomly selected frequency band of the input audio signal. Performance of the proposed system is evaluated for robustness to signal manipulations such as contamination with additive noise, resampling, compression, filtering, multiple watermark insertion, and random chopping. Experimental results show that the capacity of the proposed watermarking scheme is relatively high compared with existing spread spectrum based audio watermarking schemes.
Hafiz Malik, Ashfaq Khokhar 0001, Rashid Ansari
ICASSP (5)1
2004 Robust data-hiding in audio
abstract
A novel high capacity data hiding technique for digital audio is proposed. Imperceptibility of the embedded data is ensured based on the masking property of the human auditory system (HAS). Audio signal is decomposed into subband signals, some of which are selected for embedding data using finite-length impulse response approximations to allpass digital filters. Data detection is based on finding the filter pole-zero locations, which is achieved by power spectrum estimation of the data embedded audio signal. Performance of the proposed scheme is evaluated for different data encoding strategies. The proposed method is robust to desynchronization attacks as well as other standard data manipulation attacks.
Hafiz Malik, Ashfaq Khokhar 0001, Rashid Ansari
ICME1
2002 Predominant pitch contour extraction from audio signals
abstract
This paper describes a computationally efficient method for estimating the predominant pitch in audio recordings. The proposed method is intended for building a content-based indexing and retrieval system that can search in a audio database using the melody line of a complex input audio sample. Available pitch estimation methods are effective primarily when dealing with recordings of human voice that is either unaccompanied or accompanied with one or two musical instruments. These methods perform poorly when applied to pitch estimation in complex music signals due to their reliance on directly estimating the fundamental frequency (F/sub 0/), a task that is affected by the overlapping presence in frequency of instrumental sounds such as those of guitar, piano, etc. In our method we exploit the higher harmonic structure of the human voice to develop a low-complexity system for estimating predominant pitch. Experimental results show that this computationally efficient method provides a robust estimate of predominant pitch in real-world audio signals with 85% success rate.
Hafiz Malik, Ashfaq Khokhar 0001, Rashid Ansari, Bruno Cappe de Baillon
ICME (2)1