EDBT 2026 Demo / reviewers in the wild / expert
Rishabh Ranjan
dblp:148/9922
· DBLP profile ↗
33ranked-venue papers
13as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 10 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 9 first-author · 15 since 2021Security and privacy · 7 · 4 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ACID Test: A Benchmark for Cultural Safety and Alignment in LALMsabstractLarge Audio Language Models (LALMs) are transforming AI by processing and generating human language directly from audio. As these models proliferate in real-world applications, it becomes critical to evaluate their performance to ensure equitable and safe use across diverse linguistic and cultural contexts. We present the first comprehensive study of cultural bias in LALMs, extending text-based harm frameworks to the audio modality to analyze how linguistic diversity influences model behavior and uncover challenges in interpreting audio nuances. To address this, we introduce the Audio Cultural Intelligence Dataset (ACID), a multilingual audio–text benchmark spanning 1,315 hours across diverse languages and cultural contexts, and we conduct a systematic evaluation of 10 open-source and two closed-source models. Our results reveal substantial performance disparities across languages and cultural settings and show that biases manifest distinctly when models process audio inputs. These findings highlight the need to evaluate LALMs not only for technical accuracy but also for fair and culturally sensitive behavior, motivating the development of inclusive datasets and culturally aware training practices for safer and more equitable audio language models. Bikash Dutta, Adit Jain, Rishabh Ranjan, Mayank Vatsa, Richa Singh 0001 |
AAAI | 3 |
| 2026 | Just-in-Time-OPRFs and a Modular Framework for Fast Private Set Intersection
Mihir Bellare, Rishabh Ranjan, Doreen Riepel |
CRYPTO (8) | 2 |
| 2026 | UCX is All You Need: A Universal Transform for Committing Authenticated Encryption
Mihir Bellare, Rishabh Ranjan, Nujud Senan, Basel Alomair |
CRYPTO (6) | 2 |
| 2025 | Can RAG-Driven Enhancements Amplify Audio LLMs for Low-Resource Languages?abstractThe proliferation of Large Language Models (LLMs) has transformed Natural Language Processing (NLP), yet their development has largely overlooked low-resource languages. This paper addresses this disparity by evaluating three prominent Large Audio Language Models (LALMs) – LTU-AS, GAMA, and Pengi – across tasks like Automatic Speech Recognition (ASR), Audio Question Answering (AQA), and audio classification tasks in Hindi and code-mixed Hindi-English (aka Hinglish). We also explore the potential of Retrieval-Augmented Generation (RAG) to boost LALM performance in these low-resource settings. Our findings highlight significant performance discrepancies, with LALMs performing well in audio classification but struggling with ASR and AQA. While RAG shows potential, especially for audio classification, its impact is inconsistent across tasks. This work offers critical insights into the challenges of using LALMs for low-resource languages and provides a foundation for developing more inclusive and adaptable AI systems for complex multilingual tasks. Bikash Dutta, Rishabh Ranjan, Akshat Jain, Richa Singh 0001, Mayank Vatsa |
ICASSP | 2 |
| 2025 | ILLUSION: Unveiling Truth with a Comprehensive Multi-Modal, Multi-Lingual Deepfake DatasetabstractThe proliferation of deepfakes and AI-generated content has led to a surge in media forgeries and misinformation, necessitating robust detection systems. However, current datasets lack diversity across modalities, languages, and real-world scenarios. To address this gap, we present ILLUSION (Integration of Life-Like Unique Synthetic Identities and Objects from Neural Networks), a large-scale, multi-modal
deepfake dataset comprising 1.3 million samples spanning audio-visual forgeries, 26 languages, challenging noisy environments, and various manipulation protocols. Generated using 28 state-of-the-art generative techniques, ILLUSION includes
faceswaps, audio spoofing, synchronized audio-video manipulations, and synthetic media while ensuring a balanced representation of gender and skin tone for unbiased evaluation. Using Jaccard Index and UpSet plot analysis, we demonstrate ILLUSION’s distinctiveness and minimal overlap with existing datasets, emphasizing its novel generative coverage. We benchmarked image, audio, video, and multi-modal detection models, revealing key challenges such as performance degradation in multilingual and multi-modal contexts, vulnerability to real-world distortions, and limited generalization to zero-day attacks. By bridging synthetic and real-world complexities, ILLUSION provides a challenging yet essential platform for advancing deepfake detection research. The dataset is publicly available at https://www.iab-rubric.org/illusion-database. Kartik Thakral, Rishabh Ranjan, Akshat Jain, Mayank Vatsa, Richa Singh 0001 |
ICLR | 2 |
| 2025 | SHIELD: A Self-supervised, Silicosis-focused Hierarchical Imaging Framework for Occupational Lung Disease DiagnosisabstractSilicosis is an irreversible lung disease caused by silica dust exposure in industrial settings. Early detection is crucial, but automatic diagnostic methods are hindered by limited data availability. We propose SHIELD - a self-supervised, Silicosis-focused Hierarchical Imaging framework for early occupational Lung disease Diagnosis. Our method leverages a multi-resolution jigsaw puzzle pretext task on CXR images to extract and preserve features for lung region analysis. By employing a pyramidal strategy to generate pretrained models at various resolutions, followed by fine-tuning and a two-level ensembling across diverse deep learning architectures, SHIELD achieves enhanced diagnostic accuracy. We validate our approach on a publicly collected CXR dataset of 3044 samples from public health centers in India. SHIELD achieves 72% accuracy, demonstrating up to 20% improvement over baseline approaches. This work advances medical image analysis and supports UN Sustainable Development Goal 3 by providing cost-effective early screening in resource-limited settings. Yasmeena Akhter, Rishabh Ranjan, Richa Singh 0001, Mayank Vatsa |
IJCAI | 2 |
| 2025 | Can Quantized Audio Language Models Perform Zero-Shot Spoofing Detection?
Bikash Dutta, Rishabh Ranjan, Shyam Sathvik, Mayank Vatsa, Richa Singh 0001 |
INTERSPEECH | 2 |
| 2025 | LitMAS: A Lightweight and Generalized Multi-Modal Anti-Spoofing Framework for Biometric Security
Nidheesh Gorthi, Kartik Thakral, Rishabh Ranjan, Richa Singh 0001, Mayank Vatsa |
INTERSPEECH | 3 |
| 2025 | Multimodal Zero-Shot Framework for Deepfake Hate Speech Detection in Low-Resource Languages
Rishabh Ranjan, Ayinala Likhith, Mayank Vatsa, Richa Singh 0001 |
INTERSPEECH | 1 |
| 2025 | SynHate: Detecting Hate Speech in Synthetic Deepfake Audio
Rishabh Ranjan, Kishan Pipariya, Mayank Vatsa, Richa Singh 0001 |
INTERSPEECH | 1 |
| 2025 | Non-invasive TB Detection Using Acoustic and Semantic Features from Cough Sounds
Yasmeena Akhter, Rishabh Ranjan, Bikash Dutta, Mayank Vatsa, Richa Singh 0001 |
MICCAI (1) | 2 |
| 2024 | The Concrete Security of Two-Party Computation: Simple Definitions, and Tight Proofs for PSI and OPRFs
Mihir Bellare, Rishabh Ranjan, Doreen Riepel, Ali Aldakheel |
ASIACRYPT (6) | 2 |
| 2024 | Faking Fluent: Unveiling the Achilles' Heel of Multilingual Deepfake DetectionabstractWith the rapid advancement of deep learning techniques, the generation of audio deepfakes has achieved remarkable realism across various languages and accents. However, the effectiveness of audio deepfake detection models in diverse linguistic environments remains a crucial area of investigation. This paper presents the first empirical study on the robustness of current audio deepfake detection algorithms across different languages and accents. We evaluate whether these models maintain their effectiveness across varied linguistic domains or perform better in specific language contexts. Our comprehensive analysis examines state-of-the-art audio deepfake detection models trained on the ASVspoof 2019 and BhashaBluff datasets, assessing their performance across four diverse datasets: three representing similar-language variations (Speech Accent Archive, Svarah, and the UK English Accent Dataset) and one representing a different language (Vaani). Our results and supporting analysis indicate that while current models perform well on benchmark datasets, their ability to generalize across diverse linguistic conditions is limited. We identify potential vulnerabilities in existing models when faced with unfamiliar languages or accents, highlighting the need for more inclusive and adaptable detection systems. Our results highlight the need to enhance the robustness of audio deepfake detection across the global linguistic spectrum and emphasize the importance of developing models capable of effectively identifying synthetic speech, regardless of language or accent. Rishabh Ranjan, Bikash Dutta, Mayank Vatsa, Richa Singh 0001 |
IJCB | 1 |
| 2024 | Context Encoded Multi-Modal Attention Network for Detecting Audio SpoofingabstractHumans interpret speech through both acoustic and textual signals. Recognizing the importance of incorporating contextual information into speech processing, especially for spoofing detection, we introduce a multimodal representation framework that integrates text with audio signals to enhance spoof detection effectiveness. This research advances in two significant ways: firstly, through the creation of a new multimodal spoof detection model called Raw-BERT, and secondly, by developing an innovative multi-headed multimodal attention network that synergistically merges text and audio data for enhanced detection capabilities. We have extensively evaluated the proposed framework across diverse datasets in various languages, demonstrating that the addition of textual context significantly boosts model performance. Overall, the proposed model demonstrates superior results, consistently outperforming existing spoof detection algorithms in multiple evaluations on benchmark audio spoofing datasets. Rishabh Ranjan, Mayank Vatsa, Richa Singh 0001 |
IJCB | 1 |
| 2024 | Position: Relational Deep Learning - Graph Representation Learning on Relational DatabasesabstractMuch of the world’s most valued data is stored in relational databases and data warehouses, where the data is organized into tables connected by primary-foreign key relations. However, building machine learning models using this data is both challenging and time consuming because no ML algorithm can directly learn from multiple connected tables. Current approaches can only learn from a single table, so data must first be manually joined and aggregated into this format, the laborious process known as feature engineering. Feature engineering is slow, error prone and leads to suboptimal models. Here we introduce Relational Deep Learning (RDL), a blueprint for end-to-end learning on relational databases. The key is to represent relational databases as a temporal, heterogeneous graphs, with a node for each row in each table, and edges specified by primary-foreign key links. Graph Neural Networks then learn representations that leverage all input data, without any manual feature engineering. We also introduce RelBench, and benchmark and testing suite, demonstrating strong initial results. Overall, we define a new research area that generalizes graph machine learning and broadens its applicability. Matthias Fey, Weihua Hu, Jan Eric Lenssen, Rishabh Ranjan, Joshua Robinson 0001, Rex Ying, Jiaxuan You, Jure Leskovec |
ICML | 5 |
| 2024 | SelfVC: Voice Conversion With Iterative Refinement using Self TransformationsabstractWe propose SelfVC, a training strategy to iteratively improve a voice conversion model with self-synthesized examples. Previous efforts on voice conversion focus on factorizing speech into explicitly disentangled representations that separately encode speaker characteristics and linguistic content. However, disentangling speech representations to capture such attributes using task-specific loss terms can lead to information loss. In this work, instead of explicitly disentangling attributes with loss terms, we present a framework to train a controllable voice conversion model on entangled speech representations derived from self-supervised learning (SSL) and speaker verification models. First, we develop techniques to derive prosodic information from the audio signal and SSL representations to train predictive submodules in the synthesis model. Next, we propose a training strategy to iteratively improve the synthesis model for voice conversion, by creating a challenging training objective using self-synthesized examples. We demonstrate that incorporating such self-synthesized examples during training improves the speaker similarity of generated speech as compared to a baseline voice conversion model trained solely on heuristically perturbed inputs. Our framework is trained without any text and achieves state-of-the-art results in zero-shot voice conversion on metrics evaluating naturalness, speaker similarity, and intelligibility of synthesized audio. Paarth Neekhara, Shehzeen Hussain, Rafael Valle, Boris Ginsburg, Rishabh Ranjan, Shlomo Dubnov, Farinaz Koushanfar, Julian J. McAuley |
ICML | 5 |
| 2024 | RelBench: A Benchmark for Deep Learning on Relational DatabasesabstractWe present RelBench, a public benchmark for solving predictive tasks in relational databases with deep learning. RelBench provides databases and tasks spanning diverse domains, scales, and database dimensions, and is intended to be a foundational infrastructure for future research in this direction. We use RelBench to conduct the first comprehensive empirical study of graph neural network (GNN) based predictive models on relational data, as recently proposed by Fey et al. 2024. End-to-end learned GNNs are capable fully exploiting the predictive signal encoded in links between entities, marking a significant shift away from the dominant paradigm of manual feature engineering combined with tabular machine learning. To thoroughly evaluate GNNs against the prior gold-standard we conduct a user study, where an experienced data scientist manually engineers features for each task. In this study, GNNs learn better models whilst reducing human work needed by more than an order of magnitude. This result demonstrates the power of GNNs for solving predictive tasks in relational databases, opening up new research opportunities. Joshua Robinson 0001, Rishabh Ranjan, Weihua Hu, Jiaqi Han 0001, Alejandro Dobles, Matthias Fey, Jan Eric Lenssen, Yiwen Yuan, Zecheng Zhang, Jure Leskovec |
NeurIPS | 2 |
| 2024 | Post-Hoc Reversal: Are We Selecting Models Prematurely?abstractTrained models are often composed with post-hoc transforms such as temperature scaling (TS), ensembling and stochastic weight averaging (SWA) to improve performance, robustness, uncertainty estimation, etc. However, such transforms are typically applied only after the base models have already been finalized by standard means. In this paper, we challenge this practice with an extensive empirical study. In particular, we demonstrate a phenomenon that we call post-hoc reversal, where performance trends are reversed after applying post-hoc transforms. This phenomenon is especially prominent in high-noise settings. For example, while base models overfit badly early in training, both ensembling and SWA favor base models trained for more epochs. Post-hoc reversal can also prevent the appearance of double descent and mitigate mismatches between test loss and test error seen in base models. Preliminary analyses suggest that these transforms induce reversal by suppressing the influence of mislabeled examples, exploiting differences in their learning dynamics from those of clean examples. Based on our findings, we propose post-hoc selection, a simple technique whereby post-hoc metrics inform model development decisions such as early stopping, checkpointing, and broader hyperparameter choices. Our experiments span real-world vision, language, tabular and graph datasets. On an LLM instruction tuning dataset, post-hoc selection results in >1.5x MMLU improvement compared to naive selection. Rishabh Ranjan, Mrigank Raman, Carlos Guestrin, Zachary C. Lipton |
NeurIPS | 1 |
| 2023 | SV-DeiT: Speaker Verification with DeiTCap Spoofing DetectionabstractAs advancements in automatic speech generation continue to progress, the ability to distinguish between real and fake samples has diminished. In addition, current spoofing detection algorithms struggle to perform well on new and unseen test distributions. To address these challenges, this paper presents two contributions. First, inspired by the success of transformer and capsule networks in high representation capabilities, we propose the DeiTCap spoof detection network on spectrogram audio features. This framework utilizes multi-head attention, sub-entities (capsules) in the audio domain and a modified routing algorithm to identify capsule agreement. The proposed spoof detection algorithm is integrated into the spoofing aware speaker recognition framework SV-DeiT. Second, we introduce a novel text-to-speech dataset TRADIF created with cutting-edge transformers and diffusion models to evaluate the generalizability of countermeasure systems. Our proposed DeiT-Cap achieves an EER of 1.08% on the evaluation set of the ASVSpoof2019 LA dataset. Moreover, the proposed network demonstrates strength in cross-domain training-testing with two different datasets, highlighting its robustness and versatility. Rishabh Ranjan, Mayank Vatsa, Richa Singh 0001 |
IJCB | 1 |
| 2023 | On AI-Assisted Pneumoconiosis Detection from Chest X-raysabstractAccording to theWorld Health Organization, Pneumoconiosis affects millions of workers globally, with an estimated 260,000 deaths annually. The burden of Pneumoconiosis is particularly high in low-income countries, where occupational safety standards are often inadequate, and the prevalence of the disease is increasing rapidly. The reduced availability of expert medical care in rural areas, where these diseases are more prevalent, further adds to the delayed screening and unfavourable outcomes of the disease. This paper aims to highlight the urgent need for early screening and detection of Pneumoconiosis, given its significant impact on affected individuals, their families, and societies as a whole. With the help of low-cost machine learning models, early screening, detection, and prevention of Pneumoconiosis can help reduce healthcare costs, particularly in low-income countries. In this direction, this research focuses on designing AI solutions for detecting different kinds of Pneumoconiosis from chest X-ray data. This will contribute to the Sustainable Development Goal 3 of ensuring healthy lives and promoting well-being for all at all ages, and present the framework for data collection and algorithm for detecting Pneumoconiosis for early screening. The baseline results show that the existing algorithms are unable to address this challenge. Therefore, it is our assertion that this research will improve state-of-the-art algorithms of segmentation, semantic segmentation, and classification not only for this disease but in general medical image analysis literature. Yasmeena Akhter, Rishabh Ranjan, Richa Singh 0001, Mayank Vatsa, Santanu Chaudhury |
IJCAI | 2 |
| 2023 | Uncovering the Deceptions: An Analysis on Audio Spoofing Detection and Future ProspectsabstractAudio has become an increasingly crucial biometric modality due to its ability to provide an intuitive way for humans to interact with machines. It is currently being used for a range of applications including person authentication to banking to virtual assistants. Research has shown that these systems are also susceptible to spoofing and attacks. Therefore, protecting audio processing systems against fraudulent activities such as identity theft, financial fraud, and spreading misinformation, is of paramount importance. This paper reviews the current state-of-the-art techniques for detecting audio spoofing and discusses the current challenges along with open research problems. The paper further highlights the importance of considering the ethical and privacy implications of audio spoofing detection systems. Lastly, the work aims to accentuate the need for building more robust and generalizable methods, the integration of automatic speaker verification and countermeasure systems, and better evaluation protocols. Rishabh Ranjan, Mayank Vatsa, Richa Singh 0001 |
IJCAI | 1 |
| 2022 | STATNet: Spectral and Temporal features based Multi-Task Network for Audio Spoofing DetectionabstractWith the rise in mobile phone users and VoIP, voice has emerged as an easy and accessible biometric modality for identification or verification tasks. Given the increasing usage of voice biometrics, the security of these systems is also of paramount importance. Researchers have demon-strated that Automatic Speaker Verification (ASV) systems are prone to spoofing attacks like synthetic speech or fake speech, which can be used maliciously for a variety of tasks such as impersonation, fake news spreading, and opinion formation. This research proposes a deep convolution-based multi-task network which performs both spoof detection and source identification for synthetic speech. The pro-posed model is evaluated on three datasets ASVspoof2019 LA, FOR-Norm and In-the- Wild Audio Deepfake dataset. The results demonstrate the EER of 2.456%, 0.814%, and 0.199% on the ASVspoof2019 LA, FOR-Norm, and In-the-Wild Audio Deepfake datasets. In addition, we have also demonstrated results for cross-dataset evaluation and speech source identification. Rishabh Ranjan, Mayank Vatsa, Richa Singh 0001 |
IJCB | 1 |
| 2022 | Exploiting Epochs and Symmetries in Analysing MPI ProgramsabstractCommunication nondeterminism is one of the main reasons for the intractability of verification of message passing concurrency. In many practical message passing programs, the non-deterministic communication structure is symmetric and decomposed into epochs to obtain efficiency. Thus, symmetries and epoch structure can be exploited to reduce verification complexity. In this paper, we present a dynamic-symbolic runtime verification technique for single-path MPI programs, which (i) exploits communication symmetries by way of specifying symmetry breaking predicates (SBP) and (ii) performs compositional verification based on epochs. On the one hand, SBPs prevent the symbolic decision procedure from exploring isomorphic parts of the search space, and on the other hand, epochs restrict the size of a program needed to be analyzed at a point in time. We show that our analysis is sound and complete for single-path MPI programs on a given input. Using our prototype tool SIMIAN, we further demonstrate that our approach leads to (i) a significant reduction in verification times and (ii) scaling up to larger benchmark sizes compared to prior trace verifiers. Rishabh Ranjan, Ishita Agrawal, Subodh Sharma 0001 |
ASE | 1 |
| 2022 | A Solver-free Framework for Scalable Learning in Neural ILP ArchitecturesabstractThere is a recent focus on designing architectures that have an Integer Linear Programming (ILP) layer within a neural model (referred to as \emph{Neural ILP} in this paper). Neural ILP architectures are suitable for pure reasoning tasks that require data-driven constraint learning or for tasks requiring both perception (neural) and reasoning (ILP). A recent SOTA approach for end-to-end training of Neural ILP explicitly defines gradients through the ILP black box [Paulus et al. [2021]] – this trains extremely slowly, owing to a call to the underlying ILP solver for every training data point in a minibatch. In response, we present an alternative training strategy that is \emph{solver-free}, i.e., does not call the ILP solver at all at training time. Neural ILP has a set of trainable hyperplanes (for cost and constraints in ILP), together representing a polyhedron. Our key idea is that the training loss should impose that the final polyhedron separates the positives (all constraints satisfied) from the negatives (at least one violated constraint or a suboptimal cost value), via a soft-margin formulation. While positive example(s) are provided as part of the training data, we devise novel techniques for generating negative samples. Our solution is flexible enough to handle equality as well as inequality constraints. Experiments on several problems, both perceptual as well as symbolic, which require learning the constraints of an ILP, show that our approach has superior performance and scales much better compared to purely neural baselines and other state-of-the-art models that require solver-based training. In particular, we are able to obtain excellent performance in 9 x 9 symbolic and visual Sudoku, to which the other Neural ILP solver is not able to scale. Yatin Nandwani, Rishabh Ranjan, Mausam, Parag Singla |
NeurIPS | 2 |
| 2022 | GREED: A Neural Framework for Learning Graph Distance FunctionsabstractSimilarity search in graph databases is one of the most fundamental operations in graph analytics. Among various distance functions, graph and subgraph edit distances (GED and SED respectively) are two of the most popular and expressive measures. Unfortunately, exact computations for both are NP-hard. To overcome this computational bottleneck, neural approaches to learn and predict edit distance in polynomial time have received much interest. While considerable progress has been made, there exist limitations that need to be addressed. First, the efficacy of an approximate distance function lies not only in its approximation accuracy, but also in the preservation of its properties. To elaborate, although GED is a metric, its neural approximations do not provide such a guarantee. This prohibits their usage in higher order tasks that rely on metric distance functions, such as clustering or indexing. Second, several existing frameworks for GED do not extend to SED due to SED being asymmetric. In this work, we design a novel siamese graph neural network called Greed, which through a carefully crafted inductive bias, learns GED and SED in a property-preserving manner. Through extensive experiments across $10$ real graph datasets containing up to $7$ million edges, we establish that Greed is not only more accurate than the state of the art, but also up to $3$ orders of magnitude faster. Even more significantly, due to preserving the triangle inequality, the generated embeddings are indexable and consequently, even in a CPU-only environment, Greed is up to $50$ times faster than GPU-powered computations of the closest baseline. Rishabh Ranjan, Siddharth Grover, Sourav Medya, Venkatesan T. Chakaravarthy, Yogish Sabharwal, Sayan Ranu |
NeurIPS | 1 |
| 2021 | Multi-Scale Residual Network for Covid-19 Diagnosis Using Ct-ScansabstractTo mitigate the outbreak of highly contagious COVID-19, we need a sensitive, robust automated diagnostic tool. This paper proposes a three-level approach to separate the cases of COVID-19, pneumonia from normal patients using chest CT scans. At the first level, we fine tune a multi-scale ResNet50 model for feature extraction from all the slices of CT scan for each patient. By using multi-scale residual network, we can learn different sizes of infection, thereby making the detection possible at early stages too. These extracted features are used to train a patient-level classifier, at the second level. Four different classifiers are trained at this stage. Finally, predictions of patient level classifiers are combined by training an ensemble classifier. We test the proposed method on three sets of data released by ICASSP, COVID-19 Signal Processing Grand Challenge (SPGC). The proposed method has been successful in classifying the three classes with a validation accuracy of 94.9% and testing accuracy of 88.89%. Pratyush Garg, Rishabh Ranjan, Kamini Upadhyay, Monika Agrawal 0002, Desh Deepak |
ICASSP | 2 |
| 2020 | Robust Source Counting and DOA Estimation Using Spatial Pseudo-Spectrum and Convolutional Neural NetworkabstractMany signal processing-based methods for sound source direction-of-arrival estimation produce a spatial pseudo-spectrum of which the local maxima strongly indicate the source directions. Due to different levels of noise, reverberation and different number of overlapping sources, the spatial pseudo-spectra are noisy even after smoothing. In addition, the number of sources is often unknown. As a result, selecting the peaks from these spectra is susceptible to error. Convolutional neural network has been successfully applied to many image processing problems in general and direction-of-arrival estimation in particular. In addition, deep learning-based methods for direction-of-arrival estimation show good generalization to different environments. We propose to use a 2D convolutional neural network with multi-task learning to robustly estimate the number of sources and the directions-of-arrival from short-time spatial pseudo-spectra, which have useful directional information from audio input signals. This approach reduces the tendency of the neural network to learn unwanted association between sound classes and directional information, and helps the network generalize to unseen sound classes. The simulation and experimental results show that the proposed methods outperform other directional-of-arrival estimation methods in different levels of noise and reverberation, and different number of sources. Thi Ngoc Tho Nguyen, Woon-Seng Gan, Rishabh Ranjan, Douglas L. Jones |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Parametric Hear through Equalization for Augmented Reality AudioabstractAugmented Reality (AR) audio applications require headphones to be acoustically transparent so that real sounds can pass through unaltered for natural fusion with virtual sounds. In this paper, we consider a multiple source scenario for hear through (HT) equalization (EQ) using closed-back circumaural headsets. AR headset prototype (described in our previous study) is used to capture real sounds from external microphones and compute the directional HT filters using adaptive filtering. This method is best suited for single source scenarios as one best filter corresponding to the estimated source direction is optimally used for HT filtering. In this paper, we propose parametric HT EQ for multiple-source scenarios in time-frequency domain by estimating a sub-band Direction of Arrival (DOA) using neural networks (NN) and selecting the corresponding HT filters from a pre-computed database. Objective analysis using spectral difference (SD) is used to evaluate the performance of different HT EQ filters with open ear scenario used as a reference. Using dummy head measurements with bandlimited pink noise and real source signals, it was found that the proposed integrated system significantly improves the performance over the conventional HT system in multiple source scenarios. Rishabh Ranjan, Jianjun He 0001, Woon-Seng Gan |
ICASSP | 2 |
| 2017 | Fast HRFT measurement system with unconstrained head movements for 3D audio in virtual and augmented reality applicationsabstractBinaural audio plays an indispensable role in virtual reality (VR) and augmented reality (AR). Binaural audio recreates the sensation of the three dimensional auditory experience using Head- Related Transfer Functions (HRTFs). HRTFs are as unique as our fingerprint. To achieve an immersive audio experience, HRTFs measured from every particular user is required. Nowadays, the conventional methods for HRTF measurements requires a wellcontrolled environment, hardly any movement of the user, and projecting to the user a high level of unpleasant sound in a rather long duration. Such difficulties have greatly limited the use of individually measurement HRTFs and hinder the authenticity of immersive audio. To solve these problems, we proposed a fast and convenient HRTF measurement system that is an order of magnitude faster and more importantly, it does not place any constraints on the user's movement. With the help of a head-tracker and advanced adaptive signal processing algorithms, this system is able to achieve satisfactory HRTF measurement accuracy. In this demonstration, we will present a fast real-time HRTF acquisition system and show how the individualized HRTFs improve the audio experience in VR/AR applications. Nguyen Duy Hai, Nitesh Kumar Chaudhary, Santi Peksi, Rishabh Ranjan, Jianjun He 0001, Woon-Seng Gan |
ICASSP | 4 |
| 2016 | Fast continuous HRTF acquisition with unconstrained movements of human subjectsabstractHead related transfer function (HRTF) is widely used in 3D audio reproduction, especially over headphones. Conventionally, HRTF database is acquired at discrete directions and the acquisition process is time-consuming. Recent works have been proposed to improve HRTF acquisition efficiency via continuous acquisition. However, these HRTF acquisition techniques still require subject to sit still (with limited head movement) in a rotating chair. In this paper, we further relax the head movement constraint during acquisition by using a head tracker. The proposed continuous HRTF acquisition technique relies on the activation based normalized least-mean-square (ANLMS) algorithm to extract HRTF on the fly. Experimental results validated the accuracy of the proposed technique, when compared with the standard static acquisition technique. Jianjun He 0001, Rishabh Ranjan, Woon-Seng Gan |
ICASSP | 2 |
| 2015 | A hybrid speaker array-headphone system for immersive 3D audio reproductionabstractSpatial sound systems aim at rendering realistic sound experience to the listeners with uniform sound fields in the entire listening area. Today with the advancement of multichannel surround sound techniques, such systems are being practically realized, especially, at theatres, lecture halls, auditoriums, etc. Current practices, which are most widely used as home theatre systems, are based on multichannel stereophony, like 5.1, 10.2 and higher surround channel system. These systems require multiple loudspeakers to be placed in fixed configuration but often constrained by the room size. Sound reproduction systems like wave field synthesis (WFS) based on principle of natural propagation of sound waves, can create replica of true sound field uniformly over an extended listening area. However, WFS based systems too require hundreds of densely spaced loudspeakers enclosing the listener area and thus, difficult to realize in homes. In this paper, we introduce a new hybrid system by combining the WFS and binaural synthesis over headphones (based on active noise control techniques) to reduce the need of installing loudspeakers everywhere in a living room. Rishabh Ranjan, Woon-Seng Gan |
ICASSP | 1 |
| 2015 | Natural Listening over Headphones in Augmented Reality Using Adaptive Filtering TechniquesabstractAugmented reality (AR), which composes of virtual and real world environments, is becoming one of the major topics of research interest due to the advent of wearable devices. Today, AR is commonly used as assistive display to enhance the perception of reality in education, gaming, navigation, sports, entertainment, simulators, etc. However, most of the past works have mainly concentrated on the visual aspects of AR. Auditory events are one of the essential components in human perceptions in daily life but the augmented reality solutions have been lacking in this regard till now compared to visual aspects. Therefore, there is a need of natural listening in AR systems to give a holistic experience to the user. A new headphones configuration is presented in this work with two pairs of binaural microphones attached to headphones (one internal and one external microphone on each side). This paper focuses on enabling natural listening using open headphones employing adaptive filtering techniques to equalize the headset such that virtual sources are perceived as close as possible to sounds emanating from the physical sources. This would also require a superposition of virtual sources with the physical sound sources, as well as ambience. Modified versions of the filtered-x normalized least mean square algorithm (FxNLMS) are proposed in the paper to converge faster to the optimum solution as compared to the conventional FxNLMS. Measurements are carried out with open structure type headphones to evaluate their performance. Subjective test was conducted using individualized binaural room impulse responses (BRIRs) to evaluate the perceptual similarity between real and virtual sounds. Rishabh Ranjan, Woon-Seng Gan |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2014 | Fast and efficient real-time GPU based implementation of wave field synthesisabstractWave Field Synthesis (WFS) aims to replicate true sound field in an extended listening area with the help of loudspeaker arrays. WFS practical setups are heavily computational, as they need to drive many loudspeakers to accurately render multiple virtual sources. Thus, performance bottleneck occurs due to the sequential implementation on PCs with few cores. In addition, real-time spatial audio reproduction systems like WFS are subjected to hard real-time constraints, limiting system throughput and require cascading of several PCs to improve performance. In this paper, a fast and efficient graphics processing unit (GPU) based implementation of WFS is proposed to enhance the system throughput by extracting maximum data parallelism in the algorithm. The proposed method, implemented on NVidia C2075 GPU, uses block based partitioning approach to achieve peak system throughput of 1,400 Msamples per second, while rendering up to 200 real-time sound sources. Rishabh Ranjan, Woon-Seng Gan |
ICASSP | 1 |