VLDB 2026 Research / reviewers in the wild / expert
Rejwanul Haque
dblp:50/2942
· DBLP profile ↗
24ranked-venue papers
8as first author
11since 2021 · last 2026
0000-0003-1680-0099ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 8 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Terminology-Aware Retrieval-Augmented Knowledge Distillation for Biomedical Neural Machine TranslationabstractKnowledge distillation (KD) compresses large teacher models into smaller student models by transferring soft labels or intermediate activations. While effective in general domains, KD alone falls short in specialised machine translation (MT) settings, such as biomedical translation. The student inherits only the teacher’s compressed knowledge and lacks access to external domain information. Moreover, standard KD typically relies on abundant parallel data, which is often unavailable in domain-specific scenarios. To address these limitations, we combine KD with retrieval-augmented generation (RAG) in a few-shot setting. We propose a retrieval-augmented enhanced few-shot KD framework for French-to-English biomedical translation task. The student learns to retrieve relevant in-domain knowledge from an external database, complementing the teacher’s supervision. We design and compare several retrieval strategies to enhance student capacity. Experiments show that with our terminology-aware retrieval-based methods, the student achieves performance comparable to or better than the teacher, while preserving translation quality and efficiency. Maria Zafar, Souhail Bakkali, Rejwanul Haque |
EAMT (1) | 3 |
| 2023 | Adaptive Machine Translation with Large Language ModelsabstractConsistency is a key requirement of high-quality translation. It is especially important to adhere to pre-approved terminology and adapt to corrected translations in domain-specific projects. Machine translation (MT) has achieved significant progress in the area of domain adaptation. However, real-time adaptation remains challenging. Large-scale language models (LLMs) have recently shown interesting capabilities of in-context learning, where they learn to replicate certain input-output text generation patterns, without further fine-tuning. By feeding an LLM at inference time with a prompt that consists of a list of translation pairs, it can then simulate the domain and style characteristics. This work aims to investigate how we can utilize in-context learning to improve real-time adaptive MT. Our extensive experiments show promising results at translation time. For example, GPT-3.5 can adapt to a set of in-domain sentence pairs and/or terminology while translating a new sentence. We observe that the translation quality with few-shot in-context learning can surpass that of strong encoder-decoder MT systems, especially for high-resource languages. Moreover, we investigate whether we can combine MT from strong encoder-decoder models with fuzzy matches, which can further improve translation quality, especially for less supported languages. We conduct our experiments across five diverse language pairs, namely English-to-Arabic (EN-AR), English-to-Chinese (EN-ZH), English-to-French (EN-FR), English-to-Kinyarwanda (EN-RW), and English-to-Spanish (EN-ES). Yasmin Moslem, Rejwanul Haque, John D. Kelleher, Andy Way |
EAMT | 2 |
| 2023 | Instance-Based Domain Adaptation for Improving Terminology TranslationabstractTerms are essential indicators of a domain, and domain term translation is dealt with priority in any translation workflow. Translation service providers who use machine translation (MT) expect term translation to be unambiguous and consistent with the context and domain in question. Although current state-of-the-art neural MT (NMT) models are able to produce high-quality translations for many languages, they are still not at the level required when it comes to translating domain-specific terms. This study presents a terminology-aware instance- based adaptation method for improving terminology translation in NMT. We conducted our experiments for French-to-English and found that our proposed approach achieves a statistically significant improvement over the baseline NMT system in translating domain-specific terms. Specifically, the translation of multi-word terms is improved by 6.7% compared to the strong baseline. Prashanth Nayak, John D. Kelleher, Rejwanul Haque, Andy Way |
MTSummit (1) | 3 |
| 2022 | Identifying Fake News in Brazilian Portuguese
Marcelo Fischer, Rejwanul Haque, Paul Stynes, Pramod Pathak |
NLDB | 2 |
| 2021 | Not All Contexts are Important: The Impact of Effective Context in Conversational Neural Machine TranslationabstractMultilingual chat systems are the need of the hour for organizations who render online conversational services to their customers. This can effectively be facilitated using the robust machine translation (MT) systems. Translating user-generated contents is regarded as one of the challenging tasks for MT. As for translating conversations, it is more challenging since the meaning of any particular utterance in a conversation usually depends on its context. Moreover, chats in conversational systems are usually informal, contain code-mixed (mix of more than one language), and other grammatical inconsistencies. In this paper, we use state-of-the-art Transformer models to build our MT systems and to translate conversations between the customer service agents and customers. We propose a novel method which effectively selects contextual information from the source text of conversation to be translated. We also employ a terminology-based pseudo in-domain corpus mining strategy for fine-tuning our translation model. We evaluate our methods on the German-English WMT20 Shared Task on Chat Translation dataset, and obtain 61.1 and 63.9 BLEU points on evaluation test sets for English-to-German and German-to-English, respectively, surpassing the present state-of-the-art MT systems by 0.7 BLEU and 1.5 BLEU points, respectively. In this paper, we use state-of-the-art Transformer models to build our MT systems and to translate conversations between the customer service agents and customers. We propose a novel method which effectively selects contextual information from the source text of conversation to be translated. We also employ a terminology-based pseudo in-domain corpus mining strategy for fine-tuning our translation model. We evaluate our methods on the German-English WMT20 Shared Task on Chat Translation dataset, and obtain 61.1 and 63.9 BLEU points on evaluation test sets for English-to-German and German-to-English, respectively, surpassing the present state-of-the-art MT systems by 0.7 BLEU and 1.5 BLEU points, respectively. Baban Gain, Rejwanul Haque, Asif Ekbal |
IJCNN | 2 |
| 2021 | Augmenting Training Data for Low-Resource Neural Machine Translation via Bilingual Word Embeddings and BERT Language Modelling
Akshai Ramesh, Haque Usuf Uhana, Venkatesh Balavadhani Parthasarathy, Rejwanul Haque, Andy Way |
IJCNN | 4 |
| 2021 | Investigating Active Learning in Interactive Neural Machine TranslationabstractInteractive-predictive translation is a collaborative iterative process and where human translators produce translations with the help of machine translation (MT) systems interactively. Various sampling techniques in active learning (AL) exist to update the neural MT (NMT) model in the interactive-predictive scenario. In this paper and we explore term based (named entity count (NEC)) and quality based (quality estimation (QE) and sentence similarity (Sim)) sampling techniques – which are used to find the ideal candidates from the incoming data – for human supervision and MT model’s weight updation. We carried out experiments with three language pairs and viz. German-English and Spanish-English and Hindi-English. Our proposed sampling technique yields 1.82 and 0.77 and 0.81 BLEU points improvements for German-English and Spanish-English and Hindi-English and respectively and over random sampling based baseline. It also improves the present state-of-the-art by 0.35 and 0.12 BLEU points for German-English and Spanish-English and respectively. Human editing effort in terms of number-of-words-changed also improves by 5 and 4 points for German-English and Spanish-English and respectively and compared to the state-of-the-art. Kamal Kumar Gupta, Dhanvanth Boppana, Rejwanul Haque, Asif Ekbal, Pushpak Bhattacharyya |
MTSummit (1) | 3 |
| 2021 | Augmenting training data with syntactic phrasal-segments in low-resource neural machine translation
Kamal Kumar Gupta, Sukanta Sen, Rejwanul Haque, Asif Ekbal, Pushpak Bhattacharyya, Andy Way |
Mach. Transl. | 3 |
| 2021 | Recent advances of low-resource neural machine translation
Rejwanul Haque, Chao-Hong Liu, Andy Way |
Mach. Transl. | 1 |
| 2021 | Philipp Koehn: Neural Machine TranslationabstractAbstract Neural machine translation (NMT) is an approach to machine translation (MT) that uses deep learning techniques, a broad area of machine learning based on deep artificial neural networks (NNs). The book Neural Machine Translation by Philipp Koehn targets a broad range of readers including researchers, scientists, academics, advanced undergraduate or postgraduate students, and users of MT, covering wider topics including fundamental and advanced neural network-based learning techniques and methodologies used to develop NMT systems. The book demonstrates different linguistic and computational aspects in terms of NMT with the latest practices and standards and investigates problems relating to NMT. Having read this book, the reader should be able to formulate, design, implement, critically assess and evaluate some of the fundamental and advanced deep learning techniques and methods used for MT. Koehn himself notes that he was somewhat overtaken by events, as originally this book was envisaged only as a chapter in a revised, extended version of his 2009 book Statistical Machine Translation . However, in the interim, NMT completely overtook this previously dominant paradigm, and this new book is likely to serve as the reference of note for the field for some time to come, despite the fact that new techniques are coming onstream all the time. Wandri Jooste, Rejwanul Haque, Andy Way |
Mach. Transl. | 2 |
| 2021 | Reinforced NMT for Sentiment and Content Preservation in Low-resource ScenarioabstractThe preservation of domain knowledge from source to the target is crucial in any translation workflows. Hence, translation service providers that use machine translation (MT) in production could reasonably expect that the translation process should transfer both the underlying pragmatics and the semantics of the source-side sentences into the target language. However, recent studies suggest that the MT systems often fail to preserve such crucial information (e.g., sentiment, emotion, gender traits) embedded in the source text in the target. In this context, the raw automatic translations are often directly fed to other natural language processing (NLP) applications (e.g., sentiment classifier) in a cross-lingual platform. Hence, the loss of such crucial information during the translation could negatively affect the performance of such downstream NLP tasks that heavily rely on the output of the MT systems. In our current research, we carefully balance both the sides (i.e., sentiment and semantics) during translation, by controlling a global-attention-based neural MT (NMT), to generate translations that encode the underlying sentiment of a source sentence while preserving its non-opinionated semantic content. Toward this, we use a state-of-the-art reinforcement learning method, namely, actor-critic , that includes a novel reward combination module, to fine-tune the NMT system so that it learns to generate translations that are best suited for a downstream task, viz. sentiment classification while ensuring the source-side semantics is intact in the process. Experimental results for Hindi–English language pair show that our proposed method significantly improves the performance of the sentiment classifier and alongside results in an improved NMT system. Divya Kumari, Asif Ekbal, Rejwanul Haque, Pushpak Bhattacharyya, Andy Way |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2020 | Modelling Source- and Target- Language Syntactic Information as Conditional Context in Interactive Neural Machine TranslationabstractIn interactive machine translation (MT), human translators correct errors in automatic translations in collaboration with the MT systems, which is seen as an effective way to improve the productivity gain in translation. In this study, we model source-language syntactic constituency parse and target-language syntactic descriptions in the form of supertags as conditional context for interactive prediction in neural MT (NMT). We found that the supertags significantly improve productivity gain in translation in interactive-predictive NMT (INMT), while syntactic parsing somewhat found to be effective in reducing human effort in translation. Furthermore, when we model this source- and target-language syntactic information together as the conditional context, both types complement each other and our fully syntax-informed INMT model statistically significantly reduces human efforts in a French–to–English translation task, achieving 4.30 points absolute (corresponding to 9.18% relative) improvement in terms of word prediction accuracy (WPA) and 4.84 points absolute (corresponding to 9.01% relative) reduction in terms of word stroke ratio (WSR) over the baseline. Kamal Kumar Gupta, Rejwanul Haque, Asif Ekbal, Pushpak Bhattacharyya, Andy Way |
EAMT | 2 |
| 2020 | Syntax-Informed Interactive Neural Machine TranslationabstractIn interactive machine translation (MT), human translators correct errors in automatic translations in collaboration with the MT systems, and this is an effective way to improve productivity gain in translation. Phrase-based statistical MT (PB-SMT) has been the mainstream approach to MT for the past 30 years, both in academia and industry. Neural MT (NMT), an end-to-end learning approach to MT, represents the current state-of-the-art in MT research. The recent studies on interactive MT have indicated that NMT can significantly outperform PB-SMT. In this work, first we investigate the possibility of integrating lexical syntactic descriptions in the form of supertags into the state-of-the-art NMT model, Transformer. Then, we explore whether integration of supertags into Transformer could indeed reduce human efforts in translation in an interactive-predictive platform. From our investigation we found that our syntax-aware interactive NMT (INMT) framework significantly reduces simulated human efforts in the French-to-English and Hindi- to-English translation tasks, achieving a 2.65 point absolute corresponding to 5.65% relative improvement and a 6.55 point absolute corresponding to 19.1% relative improvement, respectively, in terms of word prediction accuracy (WPA) over the respective baselines. Kamal Kumar Gupta, Rejwanul Haque, Asif Ekbal, Pushpak Bhattacharyya, Andy Way |
IJCNN | 2 |
| 2020 | Investigating Query Expansion and Coreference Resolution in Question Answering on BERT
Santanu Bhattacharjee, Rejwanul Haque, Gideon Maillette de Buy Wenniger, Andy Way |
NLDB | 2 |
| 2020 | Analysing terminology translation errors in statistical and neural machine translation
Rejwanul Haque, Mohammed Hasanuzzaman, Andy Way |
Mach. Transl. | 1 |
| 2019 | Evaluating Terminology Translation in MT
Rejwanul Haque, Mohammed Hasanuzzaman, Andy Way |
CICLing (1) | 1 |
| 2019 | Ruslan Mitkov, Johanna Monti, Gloria Corpas Pastor, and Violeta Seretan (eds): Multiword units in machine translation and translation technology - Current Issues in Linguistic Theory, Volume 341, John Benjamin Publishing Company, Amsterdam & Philadelphia, 2018, ix+259 pp, ISBN 978-90-272-0060-0 (HB), ISBN 978-90-272-6420-6 (e-book)
Rejwanul Haque, Mohammed Hasanuzzaman, Andy Way |
Mach. Transl. | 1 |
| 2011 | Integrating source-language context into phrase-based statistical machine translation
Rejwanul Haque, Sudip Kumar Naskar, Antal van den Bosch, Andy Way |
Mach. Transl. | 1 |
| 2009 | Using Supertags as Source Language Context in SMT
Rejwanul Haque, Sudip Kumar Naskar, Yanjun Ma, Andy Way |
EAMT | 1 |
| 2009 | Dependency Relations as Source Context in Phrase-Based SMT
Rejwanul Haque, Sudip Kumar Naskar, Antal van den Bosch, Andy Way |
PACLIC | 1 |
| 2009 | Experiments on Domain Adaptation for English--Hindi SMT
Rejwanul Haque, Sudip Kumar Naskar, Josef van Genabith, Andy Way |
PACLIC | 1 |
| 2008 | Bengali, Hindi and Telugu to English Ad-hoc Bilingual Task
Sivaji Bandyopadhyay, Tapabrata Mondal, Sudip Kumar Naskar, Asif Ekbal, Rejwanul Haque, Srinivasa Rao Godhavarthy |
IJCNLP | 5 |
| 2008 | Named Entity Recognition in Bengali: A Conditional Random Field Approach
Asif Ekbal, Rejwanul Haque, Sivaji Bandyopadhyay |
IJCNLP | 2 |
| 2008 | Language Independent Named Entity Recognition in Indian Languages
Asif Ekbal, Rejwanul Haque, Amitava Das 0001, Venkateswarlu Poka, Sivaji Bandyopadhyay |
IJCNLP | 2 |