VLDB 2026 Research / reviewers in the wild / expert
Bikash Dutta
dblp:371/6340
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ACID Test: A Benchmark for Cultural Safety and Alignment in LALMsabstractLarge Audio Language Models (LALMs) are transforming AI by processing and generating human language directly from audio. As these models proliferate in real-world applications, it becomes critical to evaluate their performance to ensure equitable and safe use across diverse linguistic and cultural contexts. We present the first comprehensive study of cultural bias in LALMs, extending text-based harm frameworks to the audio modality to analyze how linguistic diversity influences model behavior and uncover challenges in interpreting audio nuances. To address this, we introduce the Audio Cultural Intelligence Dataset (ACID), a multilingual audio–text benchmark spanning 1,315 hours across diverse languages and cultural contexts, and we conduct a systematic evaluation of 10 open-source and two closed-source models. Our results reveal substantial performance disparities across languages and cultural settings and show that biases manifest distinctly when models process audio inputs. These findings highlight the need to evaluate LALMs not only for technical accuracy but also for fair and culturally sensitive behavior, motivating the development of inclusive datasets and culturally aware training practices for safer and more equitable audio language models. Bikash Dutta, Adit Jain, Rishabh Ranjan, Mayank Vatsa, Richa Singh 0001 |
AAAI | 1 |
| 2025 | Can RAG-Driven Enhancements Amplify Audio LLMs for Low-Resource Languages?abstractThe proliferation of Large Language Models (LLMs) has transformed Natural Language Processing (NLP), yet their development has largely overlooked low-resource languages. This paper addresses this disparity by evaluating three prominent Large Audio Language Models (LALMs) – LTU-AS, GAMA, and Pengi – across tasks like Automatic Speech Recognition (ASR), Audio Question Answering (AQA), and audio classification tasks in Hindi and code-mixed Hindi-English (aka Hinglish). We also explore the potential of Retrieval-Augmented Generation (RAG) to boost LALM performance in these low-resource settings. Our findings highlight significant performance discrepancies, with LALMs performing well in audio classification but struggling with ASR and AQA. While RAG shows potential, especially for audio classification, its impact is inconsistent across tasks. This work offers critical insights into the challenges of using LALMs for low-resource languages and provides a foundation for developing more inclusive and adaptable AI systems for complex multilingual tasks. Bikash Dutta, Rishabh Ranjan, Akshat Jain, Richa Singh 0001, Mayank Vatsa |
ICASSP | 1 |
| 2025 | Can Quantized Audio Language Models Perform Zero-Shot Spoofing Detection?
Bikash Dutta, Rishabh Ranjan, Shyam Sathvik, Mayank Vatsa, Richa Singh 0001 |
INTERSPEECH | 1 |
| 2025 | Non-invasive TB Detection Using Acoustic and Semantic Features from Cough Sounds
Yasmeena Akhter, Rishabh Ranjan, Bikash Dutta, Mayank Vatsa, Richa Singh 0001 |
MICCAI (1) | 3 |
| 2024 | BirdCollect: A Comprehensive Benchmark for Analyzing Dense Bird Flock AttributesabstractAutomatic recognition of bird behavior from long-term, un controlled outdoor imagery can contribute to conservation efforts by enabling large-scale monitoring of bird populations. Current techniques in AI-based wildlife monitoring have focused on short-term tracking and monitoring birds individually rather than in species-rich flocks. We present Bird-Collect, a comprehensive benchmark dataset for monitoring dense bird flock attributes. It includes a unique collection of more than 6,000 high-resolution images of Demoiselle Cranes (Anthropoides virgo) feeding and nesting in the vicinity of Khichan region of Rajasthan. Particularly, each image contains an average of 190 individual birds, illustrating the complex dynamics of densely populated bird flocks on a scale that has not previously been studied. In addition, a total of 433 distinct pictures captured at Keoladeo National Park, Bharatpur provide a comprehensive representation of 34 distinct bird species belonging to various taxonomic groups. These images offer details into the diversity and the behaviour of birds in vital natural ecosystem along the migratory flyways. Additionally, we provide a set of 2,500 point-annotated samples which serve as ground truth for benchmarking various computer vision tasks like crowd counting, density estimation, segmentation, and species classification. The benchmark performance for these tasks highlight the need for tailored approaches for specific wildlife applications, which include varied conditions including views, illumination, and resolutions. With around 46.2 GBs in size encompassing data collected from two distinct nesting ground sets, it is the largest birds dataset containing detailed annotations, showcasing a substantial leap in bird research possibilities. We intend to publicly release the dataset to the research community. The database is available at: https://iab-rubric.org/resources/wildlife-dataset/birdcollect Kshitiz, Sonu Sreshtha, Bikash Dutta, Muskan Dosi, Mayank Vatsa, Richa Singh 0001, Saket Anand, Sudeep Sarkar, Sevaram Mali Parihar |
AAAI | 3 |
| 2024 | Faking Fluent: Unveiling the Achilles' Heel of Multilingual Deepfake DetectionabstractWith the rapid advancement of deep learning techniques, the generation of audio deepfakes has achieved remarkable realism across various languages and accents. However, the effectiveness of audio deepfake detection models in diverse linguistic environments remains a crucial area of investigation. This paper presents the first empirical study on the robustness of current audio deepfake detection algorithms across different languages and accents. We evaluate whether these models maintain their effectiveness across varied linguistic domains or perform better in specific language contexts. Our comprehensive analysis examines state-of-the-art audio deepfake detection models trained on the ASVspoof 2019 and BhashaBluff datasets, assessing their performance across four diverse datasets: three representing similar-language variations (Speech Accent Archive, Svarah, and the UK English Accent Dataset) and one representing a different language (Vaani). Our results and supporting analysis indicate that while current models perform well on benchmark datasets, their ability to generalize across diverse linguistic conditions is limited. We identify potential vulnerabilities in existing models when faced with unfamiliar languages or accents, highlighting the need for more inclusive and adaptable detection systems. Our results highlight the need to enhance the robustness of audio deepfake detection across the global linguistic spectrum and emphasize the importance of developing models capable of effectively identifying synthetic speech, regardless of language or accent. Rishabh Ranjan, Bikash Dutta, Mayank Vatsa, Richa Singh 0001 |
IJCB | 2 |