Muhammad Shahid Muneer

dblp:333/8776 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0007-6093-8251ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 MOSAIV: Multi-Agent LLM Swarms for Automated Multimedia News Verification: Fake News Detection
abstract
Verifying social media content from active conflict zones requires rapid geolocation, source attribution, forensic analysis, and multi-platform verification—tasks that overwhelm individual analysts at scale. We present Multi-agent OSINT Swarm for Automated Information Verification (MOSAIV), a three-stage agentic swarm built on Large Language Models (LLMs) for automated multimedia news verification. MOSAIV operates in three sequential phases: (1) a Prime Agent use 50 labeled training samples via few-shot in-context learning to produce a shared context document and a reusable 7-step verification skill specification; (2) multiple parallel Verification Agents each process social media posts using the primed skill, performing live web searches, OSINT analysis, and structured report generation; and (3) a dedicated Localization Agent independently verifies GPS coordinates and produces bounding-box-annotated evidence images, dual-panel OpenStreetMap location cards, and live source evidence thumbnails, multiple evidence artifacts per run in total. We evaluated 10 conflict-zone validation cases provided by the MV2026 challenge. MOSAIV achieves 10/10 precise GPS coordinates, 100% forensic analysis coverage, and an average of 8.1 independent sources per report, demonstrating that few-shot priming and agent swarms are critical for high-quality automated OSINT verification. We further conduct the first systematic model-capability scaling study across six Claude variants, identifying GPS precision, forensic completeness, and propaganda context detection.
Muhammad Shahid Muneer, Khoa Van Tran, Simon S. Woo
ICMR1
2026 ICR-Net: Robust Deepfake Detection Under Temporal Corruption
Hyeongjun Choi, Muhammad Shahid Muneer, Binh Minh Le, Simon S. Woo
PAKDD (3)3
2026 AEON: Adaptive Embedding Optimized Noise for Robust Watermarking in Diffusion Models
abstract
The widespread use of synthetic image generation models and the challenges associated with authenticity preservation have fueled the demand for robust watermarking methods to safeguard authenticity and protect the copyright of synthetic images. Existing watermarking methods embed. Invisible signatures in synthetic images often compromise image quality and remain susceptible to multiple watermark removal attacks, including reconstruction and forgery methods. To overcome this issue, we propose a novel watermarking approach, AEON, which seamlessly integrates the watermark into the latent diffusion process and ensures the watermark aligns with scene semantics in the final image. Unlike existing invisible in-diffusion watermarking and traditional hash-based methods, our approach adapts the neural synthesized hash-based watermark to the semantics of the generated image during the intermediate diffusion process instead of embedding traditional hashes with the initial noise. Our proposed approach a) modulates the noise sampling in each diffusion denoising iteration through a learnable watermark embedding, b) optimizes consistency, reconstruction, and similarity loss, enforcing local and global alignment between the watermark structure and the underlying image content, and c) generates a strong watermark by allowing late embedding of the watermark in the diffusion process. Empirical results demonstrate the effectiveness of the proposed approach in retaining quality and its robustness against cumulative adversarial attacks.
Muhammad Shahid Muneer, Simon S. Woo
WACV1
2025 Beyond Masking: Landmark-based Representation Learning and Knowledge-Distillation for Audio-Visual Deepfake Detection
abstract
Audio-visual deepfake detection methods demonstrate strong performance on academic datasets but fail significantly when applied to real-world. To address the shortcomings of previous approaches, we utilize landmarks dynamic information. First, we propose Landmark-based Distillation (LBD), motivated by I-JEPA's representation learning approach. LBD utilizes KL-divergence to align facial landmark predictions from visual and audio encoders, enforcing focus on geometric facial features rather than spurious background information. Second, we introduce Multimodal Temporal Information Alignment (MTIA), which employs contrastive learning to enhance temporal consistency between audio and visual representations. We conduct experiments on academic datasets and web-based deepfakes collected from diverse social media platforms, serving as real-world examples. Our proposed landmark-guided distillation framework achieves computational efficiency while improving multimodal video deepfake detection performance across a diverse range of deepfakes compared to existing methods. The code is available at https://github.com/Ckck12/Beyond-Masking.
Muhammad Shahid Muneer, Simon S. Woo
CIKM2
2024 UGAD: Universal Generative AI Detector utilizing Frequency Fingerprints
abstract
In the wake of a fabricated explosion image at the Pentagon, an ability to discern real images from fake counterparts has never been more critical. Our study introduces a novel multi-modal approach to detect AI-generated images amidst the proliferation of new generation methods such as Diffusion models. Our method, UGAD, encompasses three key detection steps: First, we transform the RGB images into YCbCr channels and apply an Integral Radial Operation to emphasize salient radial features. Secondly, the Spatial Fourier Extraction operation is used for a spatial shift, utilizing a pre-trained deep learning network for optimal feature extraction. Finally, the deep neural network classification stage processes the data through dense layers using softmax for classification. Our approach significantly enhances the accuracy of differentiating between real and AI-generated images, as evidenced by a 12.64% increase in accuracy and 28.43% increase in AUC compared to existing state-of-the-art methods.
Inzamamul Alam, Muhammad Shahid Muneer, Simon S. Woo
CIKM2