VLDB 2026 Research / reviewers in the wild / expert
Saurav Pahuja
dblp:348/1765
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0003-1170-5199ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ATGnet: Adaptive Temporal Graph Network for EEG-enabled Sound Source Tracking in Cocktail Party ScenariosabstractDecoding selective auditory attention from electroencephalography (EEG) signals has gained considerable interest. However, few studies have looked into tracking the dynamic trajectory of moving sound source in complex auditory environments, e.g. with multiple moving speakers. We propose a novel model, namely Adaptive Temporal Graph Network (ATGnet), to continuously track the sound source trajectory using spatial-temporal EEG representations. ATGnet incorporates an adaptive graph topology to extract spatial features, and a graph-convolutional long short-term memory (GC-LSTM) network to capture spatial-temporal dependency. We evaluated ATGnet by performing within-subject leave-one-trial-out cross-validation on EEG signals from 10 participants. Experiment results indicate that ATGnet effectively overcomes the variation of signals across trials and subjects. They further confirm that ATGnet robustly tracks both attended and unattended sound sources, and significantly outperforms traditional methods. ATGnet offers a promising solution to continuous sound source tracking in dynamic conditions, with potential applications in neuro-steered hearing devices. Saurav Pahuja, Gabriel Ivucic, Siqi Cai 0002, Dashanka De Silva, Tanja Schultz, Haizhou Li 0001 |
ICASSP | 1 |
| 2025 | Selective Auditory Attention Decoding in Naturalistic Conversations Using EEG-Based Speech Envelope Tracking in Multi-Speaker Environments
Gabriel Ivucic, Saurav Pahuja, Dashanka De Silva, Tanja Schultz |
INTERSPEECH | 2 |
| 2025 | GTAnet: Geometry-Guided Temporal Attention for EEG-Based Sound Source Tracking in Cocktail Party Scenarios
Saurav Pahuja, Gabriel Ivucic, Siqi Cai 0002, Dashanka De Silva, Haizhou Li 0001, Tanja Schultz |
INTERSPEECH | 1 |
| 2025 | NeuroSpex+: Dual-Task Training of Neuro-Guided Speaker Extraction with Speech Envelope and Waveform
Dashanka De Silva, Siqi Cai 0002, Saurav Pahuja, Tanja Schultz, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2024 | Leveraging Graphic and Convolutional Neural Networks for Auditory Attention Detection with EEG
Saurav Pahuja, Gabriel Ivucic, Pascal Himmelmann, Siqi Cai 0002, Tanja Schultz, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2024 | Neurospex: Neuro-Guided Speaker Extraction With Cross-Modal FusionabstractIn the study of auditory attention, it has been revealed that there exists a robust correlation between attended speech and elicited neural responses, measurable through electroencephalography (EEG). Therefore, it is possible to use the attention information available within EEG signals to guide the extraction of the target speaker in a cocktail party computationally. In this paper, we present a neuro-guided speaker extraction model, i.e. NeuroSpex, using the EEG response of the listener as the sole auxiliary reference cue to extract attended speech from monaural speech mixtures. We propose a novel EEG signal encoder that captures the attention information. Additionally, we propose a cross-attention (CA) mechanism to enhance the speech feature representations, generating a speaker extraction mask. Experimental results on a publicly available dataset demonstrate that our proposed model outperforms two baseline models across various evaluation metrics. Dashanka De Silva, Siqi Cai 0002, Saurav Pahuja, Tanja Schultz, Haizhou Li 0001 |
SLT | 3 |
| 2023 | Multi-Head Attention and GRU for Improved Match-Mismatch Classification of Speech Stimulus and EEG ResponseabstractThis work is based on the participation by the HyperAttention team in the Auditory EEG Decoding Challenge, 2023 (ICASSP 2023 Signal Processing Grand Challenge) task 1, which deals with the match-mismatch classification of speech stimuli and EEG responses of human listeners. We demonstrate the benefits of using mel-spectrograms instead of speech envelopes as input features as well as the effectiveness of Multi-Head Attention and GRU for EEG and speech processing. With a total score of 79.05 %, we reach the second place in the challenge. Marvin Borsdorf, Saurav Pahuja, Gabriel Ivucic, Siqi Cai 0002, Haizhou Li 0001, Tanja Schultz |
ICASSP | 2 |
| 2023 | Enhancing Subject-Independent EEG-Based Auditory Attention Decoding with WGAN and Pearson Correlation CoefficientabstractElectroencephalography (EEG) related research faces a significant challenge of subject independence due to the variation in brain signals and responses among individuals. While deep learning models hold promise in addressing this challenge, their effectiveness depends on large datasets for training and generalization across participants. To overcome this limitation, we propose a solution to the above limitation by increasing the size and quality of training data for subject-independent auditory attention decoding (AAD) using EEG with deep learning. Specifically, our method employs a Wasserstein Generative Adversarial Network (WGAN) to generate synthetic data, with Pearson correlation filtering the most realistic samples. We evaluated this method on a publicly available dataset of selective auditory attention experiments and showed superior performance in subject-independent AAD performance. The mixed training set, consisting of both real and artificial data generated by the WGAN+Pearson Correlation Coefficient, demonstrated approximately 4% improvement in AAD accuracy for a 1-second window. These results demonstrate that deep learning remains a viable approach to overcoming data scarcity in subject-independent AAD tasks based on EEG. Moreover, the proposed method has the potential to improve the generalization and reliability of EEG classification tasks. Saurav Pahuja, Gabriel Ivucic, Felix Putze, Siqi Cai 0002, Haizhou Li 0001, Tanja Schultz |
SMC | 1 |