VLDB 2026 Research / reviewers in the wild / expert
Ambuj Mehrish
dblp:179/0740
· DBLP profile ↗
16ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0003-4240-9915ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DialogXpert: Driving Intelligent and Emotion-Aware Conversations Through Online Value-Based Reinforcement Learning with LLM PriorsabstractLarge-language-model (LLM) agents excel at reactive dialogue but struggle with proactive, goal-driven interactions due to myopic decoding and costly planning. We introduce DialogXpert, which leverages a frozen LLM to propose a small, high-quality set of candidate actions per turn and employs a compact Q-network over fixed BERT embeddings trained via temporal-difference learning to select optimal moves within this reduced space. By tracking the user's emotions DialogXpert tailors each decision to advance the task while nurturing a genuine, empathetic connection. Across negotiation, emotional support, and tutoring benchmarks, DialogXpert drives conversations to under 3 turns with success rates exceeding 94% and, with a larger LLM prior, pushes success above 97% while markedly improving negotiation outcomes. This framework delivers real-time, strategic, and emotionally intelligent dialogue planning at scale. Tazeek Bin Abdur Rakib, Ambuj Mehrish, Lay-Ki Soon, Wern Han Lim, Soujanya Poria |
AAAI | 2 |
| 2025 | Reward-Guided Tree Search for Inference Time Alignment of Large Language ModelsabstractChia-Yu Hung, Navonil Majumder, Ambuj Mehrish, Soujanya Poria. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Chia-Yu Hung, Navonil Majumder, Ambuj Mehrish, Soujanya Poria |
NAACL (Long Papers) | 3 |
| 2024 | HYPERTTS: Parameter Efficient Adaptation in Text to Speech Using HypernetworksabstractNeural speech synthesis, or text-to-speech (TTS), aims to transform a signal from the text domain to the speech domain. While developing TTS architectures that train and test on the same set of speakers has seen significant improvements, out-of-domain speaker performance still faces enormous limitations. Domain adaptation on a new set of speakers can be achieved by fine-tuning the whole model for each new domain, thus making it parameter-inefficient. This problem can be solved by Adapters that provide a parameter-efficient alternative to domain adaptation. Although famous in NLP, speech synthesis has not seen much improvement from Adapters. In this work, we present HyperTTS, which comprises a small learnable network, “hypernetwork”, that generates parameters of the Adapter blocks, allowing us to condition Adapters on speaker representations and making them dynamic. Extensive evaluations of two domain adaptation settings demonstrate its effectiveness in achieving state-of-the-art performance in the parameter-efficient regime. We also compare different variants of , comparing them with baselines in different studies. Promising results on the dynamic adaptation of adapter parameters using hypernetworks open up new avenues for domain-generic multi-speaker TTS systems. The audio samples and code are available at https://github.com/declare-lab/HyperTTS. Yingting Li, Rishabh Bhardwaj, Ambuj Mehrish, Bo Cheng 0001, Soujanya Poria |
LREC/COLING | 3 |
| 2024 | Accented Text-to-Speech Synthesis with a Conditional Variational AutoencoderabstractAccent plays a significant role in speech communication, influencing one's capability to understand as well as conveying a person's identity. This paper introduces a novel and efficient framework for accented Text-to-Speech (TTS) synthesis based on a Conditional Variational Autoencoder. It has the ability to synthesize a selected speaker's voice, and convert this to any desired target accent. Our thorough experiments validate the effectiveness of the proposed framework using both objective and subjective evaluations. The results also show remarkable performance in terms of the model's ability to manipulate accents in the synthesized speech. Overall, our proposed framework presents a promising avenue for future accented TTS research. Jan Melechovský, Ambuj Mehrish, Berrak Sisman, Dorien Herremans |
TENCON | 2 |
| 2024 | Accent Conversion in Text-to-Speech Using Multi-Level VAE and Adversarial TrainingabstractWith rapid globalization, the need to build inclu-sive and representative speech technology cannot be overstated. Accent is an important aspect of speech that needs to be taken into consideration while building inclusive speech synthesizers. Inclusive speech technology aims to erase any biases towards specific groups, such as people of certain accent. We note that state-of-the-art Text-to-Speech (TTS) systems may currently not be suitable for all people, regardless of their background, as they are designed to generate high-quality voices without focusing on accent. In this paper, we propose a TTS model that utilizes a Multi-Level Variational Autoencoder with adversarial learning to address accented speech synthesis and conversion in TTS, with a vision for more inclusive systems in the future. We evaluate the performance through both objective metrics and subjective listening tests. The results show an improvement in accent conversion ability compared to the baseline. Jan Melechovský, Ambuj Mehrish, Berrak Sisman, Dorien Herremans |
TENCON | 2 |
| 2023 | Evaluating Parameter-Efficient Transfer Learning Approaches on SURE Benchmark for Speech UnderstandingabstractFine-tuning is widely used as the default algorithm for transfer learning from pre-trained models. Parameter inefficiency can however arise when, during transfer learning, all the parameters of a large pre-trained model need to be updated for individual downstream tasks. As the number of parameters grows, fine-tuning is prone to overfitting and catastrophic forgetting. In addition, full fine-tuning can become prohibitively expensive when the model is used for many tasks. To mitigate this issue, parameter-efficient transfer learning algorithms, such as adapters and prefix tuning, have been proposed as a way to introduce a few trainable parameters that can be plugged into large pre-trained language models such as BERT, HuBERT. In this paper, we introduce the Speech UndeRstanding Evaluation (SURE) benchmark for parameter-efficient learning for various speech processing tasks. Additionally, we introduce a new adapter, ConvAdapter, based on 1D convolution. We show that ConvAdapter outperforms the standard adapters while showing comparable performance against prefix tuning and Low-Rank Adaptation with only 0.94% of trainable parameters. Yingting Li, Ambuj Mehrish, Rishabh Bhardwaj, Navonil Majumder, Bo Cheng 0001, Shuai Zhao 0001, Amir Zadeh 0001, Rada Mihalcea, Soujanya Poria |
ICASSP | 2 |
| 2023 | ADAPTERMIX: Exploring the Efficacy of Mixture of Adapters for Low-Resource TTS AdaptationabstractThere are significant challenges for speaker adaptation in textto-speech for languages that are not widely spoken or for speakers with accents or dialects that are not well-represented in the training data. To address this issue, we propose the use of the ”mixture of adapters” method. This approach involves adding multiple adapters within a backbone-model layer to learn the unique characteristics of different speakers. Our approach outperforms the baseline, with a noticeable improvement of 5% observed in speaker preference tests when using only one minute of data for each new speaker. Moreover, following the adapter paradigm, we fine-tune only the adapter parameters (11% of the total model parameters). This is a significant achievement in parameter-efficient speaker adaptation, and one of the first models of its kind. Overall, our proposed approach offers a promising solution to the speech synthesis techniques, particularly for adapting to speakers from diverse backgrounds. Ambuj Mehrish, Abhinav Ramesh Kashyap, Yingting Li, Navonil Majumder, Soujanya Poria |
INTERSPEECH | 1 |
| 2023 | Text-to-Audio Generation using Instruction Guided Latent Diffusion ModelabstractThe immense scale of the recent large language models (LLM) allows many interesting properties, such as, instruction- and chain-of-thought-based fine-tuning, that has significantly improved zero- and few-shot performance in many natural language processing (NLP) tasks. Inspired by such successes, we adopt such an instruction-tuned LLM Flan-T5 as the text encoder for text-to-audio (TTA) generation-a task where the goal is to generate an audio from its textual description. The prior works on TTA either pre-trained a joint text-audio encoder or used a non-instruction-tuned model, such as, T5. Consequently, our latent diffusion model (LDM)-based approach (Tango) outperforms the state-of-the-art AudioLDM on most metrics and stays comparable on the rest on AudioCaps test set, despite training the LDM on a 63 times smaller dataset and keeping the text encoder frozen. This improvement might also be attributed to the adoption of audio pressure level-based sound mixing for training set augmentation, whereas the prior methods take a random mix. Deepanway Ghosal, Navonil Majumder, Ambuj Mehrish, Soujanya Poria |
ACM Multimedia | 3 |
| 2023 | Towards lifelong human assisted speaker diarization
Meysam Shamsi, Anthony Larcher, Loïc Barrault, Sylvain Meignier, Yevhenii Prokopalo, Marie Tahon, Ambuj Mehrish, Simon Petit-Renaud, Olivier Galibert, Samuel Gaist, André Anjos, Sébastien Marcel, Marta R. Costa-jussà |
Comput. Speech Lang. | 7 |
| 2022 | Learning Accent Representation with Multi-Level VAE Towards Controllable Speech SynthesisabstractAccent is a crucial aspect of speech that helps define one's identity. We note that the state-of-the-art Text-to-Speech (TTS) systems can achieve high-quality generated voice, but still lack in terms of versatility and customizability. Moreover, they generally do not take into account accent, which is an important feature of speaking style. In this work, we utilize the concept of Multi-level VAE (ML-VAE) to build a control mechanism that aims to disentangle accent from a reference accented speaker; and to synthesize voices in different accents such as English, American, Irish, and Scottish. The proposed framework can also achieve high-quality accented voice generation for multi-speaker setup, which we believe is remarkable. We investigate the performance through objective metrics and conduct listening experiments for a subjective performance assessment. We showed that the proposed method achieves good performance for naturalness, speaker similarity, and accent similarity. Jan Melechovský, Ambuj Mehrish, Dorien Herremans, Berrak Sisman |
SLT | 2 |
| 2021 | Speaker Embeddings for Diarization of Broadcast Data In The Allies ChallengeabstractDiarization consists in the segmentation of speech signals and the clustering of homogeneous speaker segments. State-of-the-art systems typically operate upon speaker embeddings, such as i-vectors or neural x-vectors, extracted from mel cepstral coefficients (MFCCs) or spectrograms. The recent SincNet architecture extracts x-vectors directly from raw speech signals. The work reported in this paper compares the performance of different embeddings extracted from MFCCs or the raw signal for speaker diarization and broadcast media treated with compression and sub-sampling, operations which typically degrade performance. Experiments are performed with the new ALLIES database that was designed to complement existing, publicly available French corpora of broadcast radio and TV shows. Results show that, in adverse conditions, with compression and sampling mismatch, SincNet x-vectors outperform i-vectors and x-vectors by relative DERs of 43% and 73% respectively. Additionally we found that SincNet x-vectors are not the absolute best embeddings but are more robust to data mismatch than others. Anthony Larcher, Ambuj Mehrish, Marie Tahon, Sylvain Meignier, Jean Carrive, David Doukhan, Olivier Galibert, Nicholas W. D. Evans |
ICASSP | 2 |
| 2021 | The LIUM Human Active Correction Platform for Speaker Diarization
Alexandre Flucha, Anthony Larcher, Ambuj Mehrish, Sylvain Meignier, Florian Plaut, Nicolas Poupon, Yevhenii Prokopalo, Adrien Puertolas, Meysam Shamsi, Marie Tahon |
Interspeech | 3 |
| 2020 | Egocentric Analysis of Dash-Cam Videos for Vehicle ForensicsabstractVideo acquisition using dashboard-mounted cameras has recently achieved massive popularity around the world. One of the major developments following the dash-cam's popularity is that videos captured by them can be used as testimony during scenarios, like traffic violations and accidents. The widespread deployment of dash-cams brings new problems ranging from the compromise of privacy by uploading these videos on public websites using videos captured from other cars for making fraudulent claims. Therefore, there is a compelling need to address the problems associated with the usage of dash-cam videos. In this paper, we discuss and highlight the importance of the emerging area of multimedia vehicle forensics. We propose an algorithm for linking a dash-cam video to a specific car. The proposed algorithm is useful for various applications, for example, insurance companies can authenticate the origin of video before processing the claim. In a different scenario of illegitimate video upload on the Web, the video can be traced back to the car it originated from. To this end, we make use of motion blur extracted from dash-cam videos for generating a discriminative feature. We observe that the subtle motion pattern of every vehicle can serve as its unique signature. We extract motion blur from dash-cam videos and use random forest trees for classifying the vehicle correctly. The experimental results on thousands of frames obtained from dash-cam videos of several cars show the effectiveness of our approach. We further investigate the process of forging the signature of a car and propose a counter forensics method to detect such forgery. Also, we discuss the application of our technique to other potential platforms where the camera can be mounted, for example, on the chest of a person. We believe that ours is the first work that describes this new area of research. Ambuj Mehrish, Puneet Jain, A. Venkata Subramanyam, Mohan Kankanhalli |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Robust PRNU estimation from probabilistic raw measurements
Ambuj Mehrish, A. Venkata Subramanyam, Sabu Emmanuel |
Signal Process. Image Commun. | 1 |
| 2017 | Multimedia signatures for vehicle forensicsabstractThe use of dashboard-mounted video cameras is rapidly spreading in many countries around the world. Widespread usage of dash-cams brings new problems, for example, dash-cam videos are uploaded on public websites which contain footage of other cars with the number-plates visible. This can potentially compromise privacy. Further, dash-cam videos can be used as evidence in case of accidents. There have been even cases of usage of dash-cam videos for insurance claims. Not only genuine claims can be made but fraudulent claims using some other cars footage can be used. The use as well as misuse of dash-cam videos is going to be widely prevalent in the near future. In this paper, we present a solution to problem of identifying the vehicle in which the dashboard camera is mounted. Our technique can be used by insurance companies for authenticating the source of origin (the vehicle on which the camera is mounted) of video before processing the insurance claim. We make use of features extracted from motion blur as a feature which is generated due to the unique motion of a vehicle. We have found that the subtle motion pattern of every vehicle acts as unique signature. To the best of our knowledge, ours is the first work that describes this new area of research. Ambuj Mehrish, A. Venkata Subramanyam, Mohan Kankanhalli |
ICME | 1 |
| 2016 | Sensor Pattern Noise Estimation Using Probabilistically Estimated RAW ValuesabstractPhoto response nonuniformity (PRNU) is consider as reliable camera fingerprint for identifying source of a digital images. Digital cameras use various image processing operations to map linear color measurements (raw data) into nonlinear narrow gamut image. This nonlinear transformation affects estimation of PRNU. To undo the effect of nonlinear transformation, in this letter, we propose to estimate PRNU from probabilistically obtained raw values. Since not all cameras provide raw values as their output, we propose to compute estimate of raw values from the JPEG images using probabilistic color derendering procedure. The estimated raw values are modeled as a Poisson process and then maximum likelihood estimation (MLE) is used for PRNU estimation. The experimental results show that, the digital camera identification using our proposed PRNU estimate is better than using other popular PRNU estimate. Ambuj Mehrish, A. Venkata Subramanyam, Sabu Emmanuel |
IEEE Signal Process. Lett. | 1 |