Mandar Gogate

dblp:187/5503 · DBLP profile ↗
← Back
25ranked-venue papers
6as first author
20since 2021 · last 2026
0000-0003-1712-9014ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 8 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Mamba- SBRNet : Real-Time Lightweight Student Behaviour Object Detection Model
abstract
ABSTRACT Detecting student behaviour objects in classroom environments is crucial for assessing educational progress, optimizing teaching strategies and improving student learning outcomes. With the ongoing advancement of educational informatization, analysing classroom behaviour has become an important tool for enhancing teaching quality and personalized learning. However, current student behaviour object detection models based on CNN and Transformer architectures face challenges such as large parameter sizes and high inference delays when deployed on edge devices in classrooms, limiting their practical application. To address these issues, this study proposes a lightweight student behaviour detection framework based on the Mamba architecture, aimed at balancing computational efficiency and detection accuracy. First, the framework based on the state‐space model (SSM) efficiently captures global dependencies, using local convolutions to enhance detection accuracy and scene understanding while maintaining real‐time performance. Second, the C2CGA module increases attention diversity through feature splitting, self‐attention, cascading and projected concatenation, deepening the network while reducing computational overhead. Finally, the A2CMoCA module aggregates multi‐scale features, improving the learning of small objects and occluded behaviours. Experiments on a self‐built classroom behaviour dataset (containing eight typical teaching behaviours) show that the proposed method achieves 91.5% detection accuracy while maintaining a lightweight design. Compared to the baseline model, its computational efficiency (5.9G FLOPs) is reduced by 56.6%, the parameter size is compressed to 3.65 M (a 39% reduction) and the inference speed is 3.2 ms, meeting the real‐time monitoring requirements in classroom teaching scenarios.
Le Zou, Yuanhang Xia, Fengling Jiang, Yimin Wu, Kia Dashtipour, Mandar Gogate, Amir Hussain 0001, Xiaofeng Wang 0009
Expert Syst. J. Knowl. Eng.7
2026 Fourier fusion and dual-path attention enhancement network for medical image segmentation
Le Zou, Xiangxu Bu, Zhize Wu, Fengling Jiang, Lingma Sun, Kia Dashtipour, Mandar Gogate, Xiaofeng Wang 0009, Amir Hussain 0001
Multim. Syst.7
2026 Multimodal Cognitive Load Estimation With Radio Frequency Sensing and Pupillometry in Complex Auditory Environments
abstract
The detection of listening effort or cognitive load (CL) has been a major research challenge in recent years. Most conventional techniques utilise physiological or audio-visual sensors and are privacy-invasive and computationally complex. The challenges of synchronization, data alignment and accessibility limitations potentially increase the noise and error probability, compromising the accuracy of CL estimates. This innovative work presents a multi-modal, non-invasive and privacy-preserving approach that combines Radio Frequency (RF) and pupillometry sensing to address these challenges. Custom RF sensors are first designed and developed to capture blood flow changes in specific brain regions with high spatial resolution. Next, multi-modal fusion with pupillometry sensing is proposed and shown to offer a robust assessment of cognitive and listening effort through pupil size and pupil dilation. Our novel approach evaluates RF sensing to estimate CL from cerebral blood flow variations utilizing pupillometry as a baseline. A first-of-its-kind, multi-modal dataset is collected as a new benchmark resource in a controlled environment with participants to comprehend target speech with varying background noise levels. The framework is statistically evaluated using intraclass correlation for pupillometry data (average ICC> 0.95). The correlation between pupillometry and RF data is established through Pearson's correlation (average PCC> 0.79). Further, CL is classified into high and low categories based on RF data using K-means clustering. Future work involves integrating RF sensors with glasses to estimate listening effort for hearing-aid users and utilising RF measurements to optimize speech enhancement based on individual's listening effort and complexity of acoustic environment.
Usman Anwar, Adeel Hussain, Mandar Gogate, Kia Dashtipour, Tughrul Arslan, Amir Hussain 0001, Peter Lomax
IEEE J. Biomed. Health Informatics3
2025 Towards Personalised Audio Visual Speech Enhancement
Mandar Gogate, Kia Dashtipour, Amir Hussain 0001
INTERSPEECH1
2025 Investigating Gender Bias in Text-to-Audio Generation Models
Aarish Shah Mohsin, Mohammad Nadeem, Shahab Saquib Sohail, Tughrul Arslan, Mandar Gogate, Nasir Saleem, Amir Hussain 0001
INTERSPEECH5
2025 Transfer Learning-Based Automatic Sentiment Annotation of a Twitter-Based Arabic Mental Illness (AMI) Dataset
abstract
ABSTRACT Sentiment analysis, crucial for discerning emotional tones in text, relies on manual annotation to train machine learning models and is considered the gold standard for creating annotated corpora. However, this process is time‐consuming, labour‐intensive, and prone to biases. This paper proposes an automatic annotation approach for the Twitter‐based Arabic Mental Illness (AMI) dataset, which encompasses both Modern Standard Arabic and Dialectal Arabic. The approach leverages transfer learning with existing manually annotated datasets and three advanced Arabic language models to automate annotation, thereby enriching Arabic as a low‐resource language with labelled sentiment data. Validation was conducted by comparing the automatically generated annotations to manual annotation on the same dataset, achieving strong inter‐annotator agreement with a Cohen's Kappa statistic of k = 0.8457. Additionally, various baseline models were evaluated on the AMI dataset, identifying AraBERT as the top performer with the highest F1 score and accuracy.
Arwa Diwali, Kawther Saeedi, Kia Dashtipour, Mandar Gogate, Zain U. Hussain, Adam Howard, Amir Hussain 0001
Expert Syst. J. Knowl. Eng.4
2025 A Novel Approach to Fire Detection With Enhanced Target Localisation and Recognition
abstract
ABSTRACT Real‐time monitoring of fires is crucial for safeguarding lives and property. However, current fire detection methods still suffer from issues such as redundant feature information, poor network generalisation capabilities and low perception of target location information. To address these challenges, a novel fire detection method called YOLO‐FDI has been proposed. This method utilises partial convolution and coordinate convolution with attention mechanisms and Alpha loss at different stages. Specifically, to enhance target localisation accuracy, an attention mechanism is integrated into the model to autonomously focus on fire‐affected areas. In terms of feature extraction, partial convolution is employed to reduce computational redundancy and memory access, improving performance and effectively extracting spatial features. During the feature fusion stage, coordinate convolution embeds feature information into coordinate data, further enhancing the coordinate perception capabilities of pixels on the feature map, thereby improving adaptability and accuracy in detecting fire targets. Additionally, the model utilises Alpha loss to enhance flexibility and robustness in fire object detection and recognition. Experimental results demonstrate the effectiveness of the proposed model based on three self‐constructed datasets. Compared to the baseline YOLOv7 model, its mAP has improved by 4.5 percentage points, 1.7 percentage points and 2.6 percentage points, respectively. This method demonstrates the capability to accurately represent fire targets and exhibits better stability and reliability in fire target detection, effectively reducing false positives and missed detections.
Le Zou, Fengling Jiang, Zhize Wu, Lingma Sun, Mandar Gogate, Kia Dashtipour, Amir Hussain 0001
Expert Syst. J. Knowl. Eng.7
2024 Edged based audio-visual speech enhancement demonstrator
Song Chen 0005, Mandar Gogate, Kia Dashtipour, Jasper Kirton-Wingate, Adeel Hussain, Faiyaz Doctor, Tughrul Arslan, Amir Hussain 0001
INTERSPEECH2
2024 Real-Time Gaze-directed speech enhancement for audio-visual hearing-aids
Arif Reza Anway, Bryony Buck, Mandar Gogate, Kia Dashtipour, Michael A. Akeroyd, Amir Hussain 0001
INTERSPEECH3
2024 DDformer: Dimension decomposition transformer with semi-supervised learning for underwater image enhancement
Zhi Gao 0005, Jing Yang 0041, Fengling Jiang, Xixiang Jiao, Kia Dashtipour, Mandar Gogate, Amir Hussain 0001
Knowl. Based Syst.6
2024 Sentiment Analysis Meets Explainable Artificial Intelligence: A Survey on Explainable Sentiment Analysis
abstract
Sentiment analysis can be used to derive knowledge that is connected to emotions and opinions from textual data generated by people. As computer power has grown, and the availability of benchmark datasets has increased, deep learning models based on deep neural networks have emerged as the dominant approach for sentiment analysis. While these models offer significant advantages, their lack of interpretability poses a major challenge in comprehending the rationale behind their reasoning and prediction processes, leading to complications in the models' explainability. Further, only limited research has been carried out into developing deep learning models that describe their internal functionality and behaviors. In this timely study, we carry out a first of its kind overview of key sentiment analysis techniques and eXplainable artificial intelligence (XAI) methodologies that are currently in use. Furthermore, we provide a comprehensive review of sentiment analysis explainability.
Arwa Diwali, Kawther Saeedi, Kia Dashtipour, Mandar Gogate, Erik Cambria, Amir Hussain 0001
IEEE Trans. Affect. Comput.4
2024 Context-Aware Audio-Visual Speech Enhancement Based on Neuro-Fuzzy Modeling and User Preference Learning
abstract
It is estimated that by 2050 approximately one in ten individuals globally will experience disabling hearing impairment. In the presence of everyday reverberant noise, a substantial proportion of individual users encounter challenges in speech comprehension. This study introduces a novel application of neuro-fuzzy modeling that synergizes and fuses audio-visual speech enhancement (AV SE) with an initial user preference learning based framework. Specifically, our approach uniquely integrates multimodal AV speech data with innovative SE methods and fuzzy inferencing techniques. This integration is further enriched by incorporating a user-preference learning model that adapts to environmental and user-specific contexts, including signal-to-noise ratios, sound power, and the quality of visual information. The proposed framework facilitates the incorporation of clinical measures such as user cognitive load (or listening effort) with real-world uncertainty to steer the system outputs. We employ an adaptive fuzzy neural network to derive the most effective Sugeno fuzzy inference model, employing particle swarm optimization to ensure optimal SE by considering sound power, ambient noise levels, and visual quality. Experimental results utilize our new benchmark AV multitalker challenge dataset to demonstrate the superiority of our user preference-informed, context-aware AV SE approach in enhancing speech intelligibility and quality in challenging noisy conditions, marking a significant advancement over conventional methods while reducing energy consumption. The conclusion supports the ecological scalability of our approach and its potential for real-world applications, setting a new benchmark in AV SE research, paving the way for future assistive hearing and communication technologies.
Song Chen 0005, Jasper Kirton-Wingate, Faiyaz Doctor, Usama Arshad, Kia Dashtipour, Mandar Gogate, Zahid Halim, Ahmed Yassin Al-Dubai, Tughrul Arslan, Amir Hussain 0001
IEEE Trans. Fuzzy Syst.6
2023 5G-IoT Cloud based Demonstration of Real-Time Audio-Visual Speech Enhancement for Multimodal Hearing-aids
Ankit Gupta 0008, Abhijeet Bishnu, Mandar Gogate, Kia Dashtipour, Tughrul Arslan, Ahsan Adeel, Amir Hussain 0001, Tharmalingam Ratnarajah, Mathini Sellathurai
INTERSPEECH3
2023 Application for Real-time Audio-Visual Speech Enhancement
Mandar Gogate, Kia Dashtipour, Amir Hussain 0001
INTERSPEECH1
2023 Live Demonstration: Cloud-based Audio-Visual Speech Enhancement in Multimodal Hearing-aids
abstract
Hearing loss is among the most serious public health problems, affecting as much as 20% of the worldwide population. Even cutting-edge multi-channel audio-only speech enhancement (SE) algorithms used in modern hearing aids face significant hurdles since they typically magnify noises while failing to boost speech understanding in crowded social environments. Recently, for the first time we proposed a novel integration of 5G cloud-radio access network, internet of things (IoT), and strong privacy algorithms to develop 5G IoT enabled hearing aid (HA) [1]. In this demonstration, we show the first-ever transceiver (PHY layer) model for cloud-based audio-visual (AV) SE, which meets the requirements for high data rate and low latency of forthcoming multi-modal HAs (such as Google glasses with integrated HAs). Even in highly noisy conditions like cafés, clubs, conferences, meetings, etc., the transceiver [2] transmits raw AV information from a hearing aid system to a cloud-based platform and obtains a clear signal. In Fig. 1, we illustrate an example of our cloud-based AV SE hearing aid demonstration. Herein, the left-side computer and Universal Software Radio Peripheral (USRP) x310 function as IoT systems (hearing aids), the right-side USRP serves as an access point or base station, and the right-side computer serves as a cloud-server for operating NN-based SE models. Please take note that the channel between the HA device and the cloud is defined as an uplink channel, whereas the channel between the access point (cloud) and the HA device is defined as a downlink channel. Given the time-varying sensitivity of the data received at HA devices, the uplink channel can therefore handle a variety of data rates. As a result, a customized long-term evolution (LTE)-based frame structure is developed for uplink transmission of data. It provides error-correction codes in the 1.4 MHz and 3 MHz bandwidths with a variety of modulations and code rates. Furthermore, the cloud access point simply supports a limited transmission rate because it only transmits audio data to the HA equipment. In order to support real-time AV SE, a modified frame structure for LTE with 1.4 MHz of bandwidth is developed. The AV SE algorithm receives cropped lip images of the target speaker and a noisy speech spectrogram, and it produces an ideal binary mask that lessens the noise-dominant regions while improving the speech-dominant areas. We use the depth-wise separable convolutions, reduced STFT window size of 32 ms, smaller STFT window shift of 8 ms, and 64 convolutions in the audio feature extraction layers of our Cochlea-Net [3] multi-modal AV SE neural network architecture to reduce processing latency. Furthermore, the visual feature extraction framework is employed. Our proposed architecture can handle streaming data frame-by-frame. Thus, the users will experience for the first time the real-world development of a physical layer transceiver that can perform AV SE in real-time under strict latency and data rate requirements. For this demonstration, we will bring two computers and two USRP x310 devices.
Abhijeet Bishnu, Ankit Gupta 0008, Mandar Gogate, Kia Dashtipour, Tughrul Arslan, Ahsan Adeel, Amir Hussain 0001, Mathini Sellathurai, Tharmalingam Ratnarajah
ISCAS3
2023 Live Demonstration: Real-time Multi-modal Hearing Assistive Technology Prototype
abstract
Hearing loss affects at least 1.5 billion people globally. The WHO estimates 83% of people who could benefit from hearing aids do not use them. Barriers to HA uptake are multifaceted but include ineffectiveness of current HA technology in noisy environments with multiple competing noise sources where human performance is known to be dependent upon input from both the aural and visual senses.
Mandar Gogate, Adeel Hussain, Kia Dashtipour, Amir Hussain 0001
ISCAS1
2022 A Novel Frame Structure for Cloud-Based Audio-Visual Speech Enhancement in Multimodal Hearing-aids
abstract
In this paper, we design a first of its kind transceiver (PHY layer) prototype for cloud-based audio-visual (AV) speech enhancement (SE) complying with high data rate and low latency requirements of future multimodal hearing assistive technology. The innovative design needs to meet multiple challenging constraints including up/down link communications, delay of transmission and signal processing, and real-time AV SE models processing. The transceiver includes device detection, frame detection, frequency offset estimation, and channel estimation capabilities. We develop both uplink (hearing aid to the cloud) and downlink (cloud to hearing aid) frame structures based on the data rate and latency requirements. Due to the varying nature of uplink information (audio and lip-reading), the uplink channel supports multiple data rate frame structure, while the downlink channel has a fixed data rate frame structure. In addition, we evaluate the latency of different PHY layer blocks of the transceiver for developed frame structures using LabVIEW NXG. This can be used with software defined radio (such as Universal Software Radio Peripheral) for real-time demonstration scenarios.
Abhijeet Bishnu, Ankit Gupta 0008, Mandar Gogate, Kia Dashtipour, Ahsan Adeel, Amir Hussain 0001, Mathini Sellathurai, Tharmalingam Ratnarajah
HealthCom3
2022 AVSE Challenge: Audio-Visual Speech Enhancement Challenge
abstract
Audio-visual speech enhancement is the task of improving the quality of a speech signal when video of the speaker is available. It opens-up the opportunity of improving speech intelligibility in adverse listening scenarios that are currently too challenging for audio-only speech enhancement models. The Audio-Visual Speech Enhancement (AVSE) challenge aims to set the first benchmark in this area. We provide participants with datasets and scripts to test their audio-visual speech enhancement models under a common framework for both training and evaluation. The data is derived from real-world videos, and comprises noisy mixes, in which audio from target speaker is mixed with either a competing speaker or a noise signal. The submitted systems are evaluated by conducting AV intelligibility tests involving human participants. We expect this challenge to be a platform for advancing the field of audio-visual speech-enhancement and to provide further insight about the scope and limitations of current AV speech enhancement approaches.
Andrea Lorena Aldana Blanco, Cassia Valentini-Botinhao, Ondrej Klejch, Mandar Gogate, Kia Dashtipour, Amir Hussain 0001, Peter Bell 0001
SLT4
2021 A novel context-aware multimodal framework for persian sentiment analysis
Kia Dashtipour, Mandar Gogate, Erik Cambria, Amir Hussain 0001
Neurocomputing2
2021 Lane-DeepLab: Lane semantic segmentation in automatic driving scenarios for high-definition maps
Fengling Jiang, Jing Yang 0041, Mandar Gogate, Kia Dashtipour, Amir Hussain 0001
Neurocomputing5
2020 Deep Neural Network Driven Binaural Audio Visual Speech Separation
abstract
The central auditory pathway exploits the auditory signals and visual information sent by both ears and eyes to segregate speech from multiple competing noise sources and help disambiguate phonological ambiguity. In this study, inspired from this unique human ability, we present a deep neural network (DNN) that ingest the binaural sounds received at the two ears as well as the visual frames to selectively suppress the competing noise sources individually at both ears. The model exploits the noisy binaural cues and noise robust visual cues to improve speech intelligibility. The comparative simulation results in terms of objective metrics such as PESQ, STOI, SI-SDR and DBSTOI demonstrate significant performance improvement of the proposed audio-visual (AV) DNN as compared to the audio-only (A-only) variant of the proposed model. Finally, subjective listening tests with the real noisy AV ASPIRE corpus shows the superiority of the proposed AV DNN as compared to state-of-the-art approaches.
Mandar Gogate, Kia Dashtipour, Peter Bell 0001, Amir Hussain 0001
IJCNN1
2020 Visual Speech In Real Noisy Environments (VISION): A Novel Benchmark Dataset and Deep Learning-Based Baseline System
abstract
In this paper, we present VIsual Speech In real nOisy eNvironments (VISION), a first of its kind audio-visual (AV) corpus comprising 2500 utterances from 209 speakers, recorded in real noisy environments including social gatherings, streets, cafeterias and restaurants. While a number of speech enhancement frameworks have been proposed in the literature that exploit AV cues, there are no visual speech corpora recorded in real environments with a sufficient variety of speakers, to enable evaluation of AV frameworks' generalisation capability in a wide range of background visual and acoustic noises. The main purpose of our AV corpus is to foster research in the area of AV signal processing and to provide a benchmark corpus that can be used for reliable evaluation of AV speech enhancement systems in everyday noisy settings. In addition, we present a baseline deep neural network (DNN) based spectral mask estimation model for speech enhancement. Comparative simulation results with subjective listening tests demonstrate significant performance improvement of the baseline DNN compared to state-of-the-art speech enhancement approaches.
Mandar Gogate, Kia Dashtipour, Amir Hussain 0001
INTERSPEECH1
2020 A hybrid Persian sentiment analysis framework: Integrating dependency grammar based rules and deep neural networks
Kia Dashtipour, Mandar Gogate, Jingpeng Li 0001, Fengling Jiang, Amir Hussain 0001
Neurocomputing2
2019 A novel deep learning driven, low-cost mobility prediction approach for 5G cellular networks: The case of the Control/Data Separation Architecture (CDSA)
abstract
One of the fundamental goals of mobile networks is to enable uninterrupted access to wireless services without compromising the expected quality of service (QoS). This paper reports a number of significant contributions. First, a novel analytical model is proposed for holistic handover (HO) cost evaluation, that integrates signaling overhead, latency, call dropping, and radio resource wastage. The developed mathematical model is applicable to several cellular architectures, but the focus here is on the Control/Data Separation Architecture (CDSA). Second, data-driven HO prediction is proposed and evaluated as part of the holistic cost, for the first time, through novel application of a recurrent deep learning architecture, specifically, a stacked long-short-term memory (LSTM) model. Finally, simulation results and preliminary analysis reveal different cases where non-predictive and predictive deep neural networks can be effectively utilized, based on HO management requirements. Both analytical and machine learning models are evaluated with a benchmark, real-world dataset measuring human behaviors and interactions. Numerical and comparative simulation results demonstrate the potential of our proposed deep learning-driven HO management framework, as a future benchmark for the mobile networking and machine learning communities.
Metin Öztürk, Mandar Gogate, Oluwakayode Onireti, Ahsan Adeel, Amir Hussain 0001, Muhammad Ali Imran 0001
Neurocomputing2
2018 DNN Driven Speaker Independent Audio-Visual Mask Estimation for Speech Separation
abstract
Human auditory cortex excels at selectively suppressing background noise to focus on a target speaker. The process of selective attention in the brain is known to contextually exploit the available audio and visual cues to better focus on target speaker while filtering out other noises. In this study, we propose a novel deep neural network (DNN) based audiovisual (AV) mask estimation model. The proposed AV mask estimation model contextually integrates the temporal dynamics of both audio and noise-immune visual features for improved mask estimation and speech separation. For optimal AV features extraction and ideal binary mask (IBM) estimation, a hybrid DNN architecture is exploited to leverages the complementary strengths of a stacked long short term memory (LSTM) and convolution LSTM network. The comparative simulation results in terms of speech quality and intelligibility demonstrate significant performance improvement of our proposed AV mask estimation model as compared to audio-only and visual-only mask estimation approaches for both speaker dependent and independent scenarios.
Mandar Gogate, Ahsan Adeel, Ricard Marxer, Jon Barker, Amir Hussain 0001
INTERSPEECH1