VLDB 2026 Research / reviewers in the wild / expert
Aditya Parikh
dblp:192/5221 · also Aditya Kamlesh Parikh
· DBLP profile ↗
11ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 5 first-author · 10 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Theory of computation · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generating High Quality Synthetic Data for Dutch Medical ConversationsabstractContains fulltext : 334342.pdf (Publisher’s version ) (Closed access) Contains fulltext : 334342.pdf (Author’s version preprint ) (Open Access) Cecilia Kuan, Aditya Parikh, Henk van den Heuvel |
LREC | 2 |
| 2026 | Rubric-Guided Fine-tuning of SpeechLLMs for Multi-Aspect, Multi-Rater L2 Reading-Speech AssessmentabstractContains fulltext : 331543.pdf (Publisher’s version ) (Closed access) Aditya Parikh, Cristian Tejedor García, Catia Cucchiarini, Helmer Strik |
LREC | 1 |
| 2025 | Leveraging Transfer Learning with Class-Specific Decoders for Laparoscopic Segmentationabstract7089 Priya Tomar, Aditya Parikh, Christian Bauckhage, Rafet Sifa |
IEEE Big Data | 2 |
| 2025 | First Investigation of Deep Learning for Intraoperative Gauze Segmentation in Minimally Invasive Abdominal SurgeryabstractSurgical gauze is an essential part of surgical procedures, which is primarily used for controlling bleeding and absorbing bodily fluids. The post-surgical retention of gauze can lead to serious complications in the patient's health and necessitate additional surgery for gauze removal. In the wake of data scarcity, the research on gauze segmentation on the real-world surgical data remains underexplored. In this work, we investigate the use of deep learning methods for gauze segmentation in robotassisted minimally invasive abdominal surgeries, utilizing an inhouse surgical dataset prepared at a university hospital. The training data reflects a realistic surgical setting and extensive diversity in spatial, morphological, and visual attributes of three different gauze categories. We have investigated prevalently used segmentation architectures, including CNN-based, transformer-based, and hybrid architectures, to provide a proof-of-concept for gauze segmentation in a realistic setting. Besides, we investigate the influence of additional sub-optimally annotated, auto-tracked segmentation masks to address the bottleneck of data scarcity and performance optimization. Our results demonstrate the efficacy of real-world data to counter the main challenge reported by prior works - the trade-off between blood presence and gauze detection. The incorporation of auto-track annotations enables performance enhancements, particularly in generic cases. The integration of effective segmentation approaches will benefit robotguided surgical procedures and various downstream applications by providing a precise delineation of foreign objects, enhancing patient safety and surgical outcomes. Priya Tomar, Maximilian Broß, Philipp Feodorovici, Jan Arensmeyer, Philipp Leifels, Aditya Parikh, Hanno Matthaei, Christian Bauckhage, Helen Schneider, Rafet Sifa |
DSAA | 6 |
| 2025 | Evaluating Logit-Based GOP Scores for Mispronunciation DetectionabstractPronunciation assessment relies on goodness of pronunciation (GOP) scores, traditionally derived from softmax-based posterior probabilities. However, posterior probabilities may suffer from overconfidence and poor phoneme separation, limiting their effectiveness. This study compares logit-based GOP scores with probability-based GOP scores for mispronunciation detection. We conducted our experiment on two L2 English speech datasets spoken by Dutch and Mandarin speakers, assessing classification performance and correlation with human ratings. Logit-based methods outperform probability-based GOP in classification, but their effectiveness depends on dataset characteristics. The maximum logit GOP shows the strongest alignment with human perception, while a combination of different GOP scores balances probability and logit features. The findings suggest that hybrid GOP methods incorporating uncertainty modeling and phoneme-specific weighting improve pronunciation assessment. Aditya Parikh, Cristian Tejedor García, Catia Cucchiarini, Helmer Strik |
INTERSPEECH | 1 |
| 2025 | Enhancing GOP in CTC-Based Mispronunciation Detection with Phonological KnowledgeabstractComputer-Assisted Pronunciation Training (CAPT) systems employ automatic measures of pronunciation quality, such as the goodness of pronunciation (GOP) metric. GOP relies on forced alignments, which are prone to labeling and segmentation errors due to acoustic variability. While alignment-free methods address these challenges, they are computationally expensive and scale poorly with phoneme sequence length and inventory size. To enhance efficiency, we introduce a substitution-aware alignment-free GOP that restricts phoneme substitutions based on phoneme clusters and common learner errors. We evaluated our GOP on two L2 English speech datasets, one with child speech, My Pronunciation Coach (MPC), and SpeechOcean762, which includes child and adult speech. We compared RPS (restricted phoneme substitutions) and UPS (unrestricted phoneme substitutions) setups within alignment-free methods, which outperformed the baseline. We discuss our results and outline avenues for future research. Aditya Parikh, Cristian Tejedor García, Catia Cucchiarini, Helmer Strik |
INTERSPEECH | 1 |
| 2024 | Ensembles of Hybrid and End-to-End Speech RecognitionabstractWe propose a method to combine the hybrid Kaldi-based Automatic Speech Recognition (ASR) system with the end-to-end wav2vec 2.0 XLS-R ASR using confidence measures. Our research is focused on the low-resource Irish language. Given the limited available open-source resources, neither the standalone hybrid ASR nor the end-to-end ASR system can achieve optimal performance. By applying the Recognizer Output Voting Error Reduction (ROVER) technique, we illustrate how ensemble learning could facilitate mutual error correction between both ASR systems. This paper outlines the strategies for merging the hybrid Kaldi ASR model and the end-to-end XLS-R model with the help of confidence scores. Although contemporary state-of-the-art end-to-end ASR models face challenges related to prediction overconfidence, we utilize Renyi’s entropy-based confidence approach, tuned with temperature scaling, to align it with the Kaldi ASR confidence. Although there was no significant difference in the Word Error Rate (WER) between the hybrid and end-to-end ASR, we could achieve a notable reduction in WER after ensembling through ROVER. This resulted in an almost 14% Word Error Rate Reduction (WERR) on our primary test set and an approximately 20% WERR on other noisy and imbalanced test data. Aditya Parikh, Louis ten Bosch, Henk van den Heuvel |
LREC/COLING | 1 |
| 2024 | SignON - a Co-creative Machine Translation for Sign and Spoken Languages (end-of-project results, contributions and lessons learned)abstractSignON, a 3-year Horizon 20202 project addressing the lack of technology and services for MT between sign languages (SLs) and spoken languages (SpLs) ended in December 2023. SignON was unprecedented. Not only it addressed the wider complexity of the aforementioned problem – from research and development of recognition, translation and synthesis, through development of easy-to-use mobile applications and a cloud-based framework to do the “heavy lifting” as well as to establishing ethical, privacy and inclusivenesspolicies and operation guidelines – but also engaged with the deaf and hard of hearing communities in an effective co-creation approach where these main stakeholders drove the development in the right direction and had the final say.Currently we are witnessing advances in natural language processing for SLs, including MT. SignON was one of the largest projects that contributed to this surge with 17 partners and more than 60 consortium members, working in parallel with other international and European initiatives, such as project EASIER and others. Dimitar Sht. Shterionov, Vincent Vandeghinste, Mirella De Sisto, Aoife Brady, Mathieu De Coster, Lorraine Leeson, Andy Way, Josep Blat, Frankie Picron, Davy Van Landuyt, Marcello Paolo Scipioni, Aditya Parikh, Louis ten Bosch, John J. O'Flaherty, Joni Dambre, Caro Brosens, Jorn Rijckaert, Víctor Ubieto Nogales, Bram Vanroy, Santiago Egea Gómez, Ineke Schuurman, Gorka Labaka, Adrián Núñez-Marcos, Irene Murtagh, Euan McGill, Horacio Saggion |
EAMT (2) | 12 |
| 2023 | SignON: Sign Language Translation. Progress and challengesabstractSignON (https://signon-project.eu/) is a Horizon 2020 project, running from 2021 until the end of 2023, which addresses the lack of technology and services for the automatic translation between sign languages (SLs) and spoken languages, through an inclusive, human-centric solution, hence contributing to the repertoire of communication media for deaf, hard of hearing (DHH) and hearing individuals. In this paper, we present an update of the status of the project, describing the approaches developed to address the challenges and peculiarities of SL machine translation (SLMT). Vincent Vandeghinste, Dimitar Sht. Shterionov, Mirella De Sisto, Aoife Brady, Mathieu De Coster, Lorraine Leeson, Josep Blat, Frankie Picron, Marcello Paolo Scipioni, Aditya Parikh, Louis ten Bosch, John J. O'Flaherty, Joni Dambre, Jorn Rijckaert, Bram Vanroy, Víctor Ubieto Nogales, Santiago Egea Gómez, Ineke Schuurman, Gorka Labaka, Adrián Núñez-Marcos, Irene Murtagh, Euan McGill, Horacio Saggion |
EAMT | 10 |
| 2022 | Sign Language Translation: Ongoing Development, Challenges and Innovations in the SignON ProjectabstractThe SignON project (www.signon-project.eu) focuses on the research and development of a Sign Language (SL) translation mobile application and an open communications framework. SignON rectifies the lack of technology and services for the automatic translation between signed and spoken languages, through an inclusive, humancentric solution which facilitates communication between deaf, hard of hearing (DHH) and hearing individuals. We present an overview of the current status of the project, describing the milestones reached to date and the approaches that are being developed to address the challenges and peculiarities of Sign Language Machine Translation (SLMT). Dimitar Sht. Shterionov, Mirella De Sisto, Vincent Vandeghinste, Aoife Brady, Mathieu De Coster, Lorraine Leeson, Josep Blat, Frankie Picron, Marcello Paolo Scipioni, Aditya Parikh, Louis ten Bosch, John J. O'Flaherty, Joni Dambre, Jorn Rijckaert |
EAMT | 10 |
| 2016 | Disease Detection and Severity Estimation in Cotton Plant from Unconstrained ImagesabstractThe primary focus of this paper is to detect disease and estimate its stage for a cotton plant using images. Most disease symptoms are reflected on the cotton leaf. Unlike earlier approaches, the novelty of the proposal lies in processing images captured under uncontrolled conditions in the field using normal or a mobile phone camera by an untrained person. Such field images have a cluttered background making leaf segmentation very challenging. The proposed work use two cascaded classifiers. Using local statistical features, first classifier segments leaf from the background. Then using hue and luminance from HSV colour space another classifier is trained to detect disease and find its stage. The developed algorithm is a generalised as it can be applied for any disease. However as a showcase, we detect Grey Mildew, widely prevalent fungal disease in North Gujarat, India. Aditya Parikh, Mehul S. Raval, Chandrasinh Parmar, Sanjay Chaudhary |
DSAA | 1 |