Shamil Chollampatt

dblp:182/2351 · DBLP profile ↗
← Back
12ranked-venue papers
8as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 8 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Machine translation · 46% Language models and text generation · 31% Efficient and distributed learning · 12%

Topics — the 22 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › text generation
grammatical error correction
1.552019
Cross-Sentence Grammatical Error Correction · ACL (1) 2019
Neural Quality Estimation of Grammatical Error Correction · EMNLP 2018
A Multilayer Convolutional Encoder-Decoder Neural Network for Grammatical Error Correction · AAAI 2018
Natural language and speech › Machine translation
neural machine translation
0.922023
CLAD-ST: Contrastive Learning with Adversarial Data for Robust Speech Translation · EMNLP 2023
Neural Network Translation Models for Grammatical Error Correction · IJCAI 2016
Natural language and speech › Machine translation › speech translation
cascaded speech translation
0.712023
CLAD-ST: Contrastive Learning with Adversarial Data for Robust Speech Translation · EMNLP 2023
Natural language and speech › Language models and text generation › text summarization
dialogue summarization
0.712023
Select, Prompt, Filter: Distilling Large Language Models for Summarizing Conversations · EMNLP 2023
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.712023
Select, Prompt, Filter: Distilling Large Language Models for Summarizing Conversations · EMNLP 2023
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
LLM distillation
0.712023
Select, Prompt, Filter: Distilling Large Language Models for Summarizing Conversations · EMNLP 2023
Natural language and speech › Machine translation
robust machine translation
0.712023
CLAD-ST: Contrastive Learning with Adversarial Data for Robust Speech Translation · EMNLP 2023
Natural language and speech › Machine translation
speech translation
0.712023
CLAD-ST: Contrastive Learning with Adversarial Data for Robust Speech Translation · EMNLP 2023
Natural language and speech › Language models and text generation
text summarization
0.712023
Select, Prompt, Filter: Distilling Large Language Models for Summarizing Conversations · EMNLP 2023
Natural language and speech › Machine translation › computer-assisted translation
automatic post-editing
0.412020
Can Automatic Post-Editing Improve NMT? · EMNLP (1) 2020
Natural language and speech › Machine translation
constrained machine translation
0.412020
Lexically Constrained Neural Machine Translation with Levenshtein Transformer · ACL 2020
Natural language and speech › Machine translation › constrained machine translation
lexically constrained translation
0.412020
Lexically Constrained Neural Machine Translation with Levenshtein Transformer · ACL 2020
Machine learning › Deep learning architectures and training › encoder-decoder architecture
convolutional encoder-decoder
0.312018
A Multilayer Convolutional Encoder-Decoder Neural Network for Grammatical Error Correction · AAAI 2018
Machine learning › Deep learning architectures and training
encoder-decoder architecture
0.312018
A Multilayer Convolutional Encoder-Decoder Neural Network for Grammatical Error Correction · AAAI 2018
Natural language and speech › Machine translation › machine translation evaluation
translation quality estimation
0.312018
Neural Quality Estimation of Grammatical Error Correction · EMNLP 2018
Machine learning › Representation and self-supervised learning › contrastive learning › robust contrastive learning
adversarial contrastive learning
0.212023
CLAD-ST: Contrastive Learning with Adversarial Data for Robust Speech Translation · EMNLP 2023
Machine learning › Representation and self-supervised learning
contrastive learning
0.212023
CLAD-ST: Contrastive Learning with Adversarial Data for Robust Speech Translation · EMNLP 2023
Natural language and speech › Language models and text generation › prompting
prompt engineering
0.212023
Select, Prompt, Filter: Distilling Large Language Models for Summarizing Conversations · EMNLP 2023
Natural language and speech › Machine translation
statistical machine translation
0.222018
A Multilayer Convolutional Encoder-Decoder Neural Network for Grammatical Error Correction · AAAI 2018
Adapting Grammatical Error Correction Based on the Native Language of Writers with Neural Network Joint Models · EMNLP 2016
Machine learning › Deep learning architectures and training › sequence modeling › sequence generation
non-autoregressive generation
0.112020
Lexically Constrained Neural Machine Translation with Levenshtein Transformer · ACL 2020
Natural language and speech › Language models and text generation
text generation
0.112020
Lexically Constrained Neural Machine Translation with Levenshtein Transformer · ACL 2020
Natural language and speech › Language models and text generation
reranking
0.112018
Neural Quality Estimation of Grammatical Error Correction · EMNLP 2018

Methods — techniques the papers use, named apart from their topics

attention · 0.7semantic similarity sampling · 0.7prompt retrieval · 0.7knowledge distillation · 0.7curriculum learning · 0.7contrastive learning · 0.7adversarial examples · 0.7neural post-editing model · 0.4levenshtein transformer · 0.4beam search · 0.4
YearPublicationVenuePosition
2025 Cross-lingual Evaluation of Multilingual Text Generation
abstract
Scaling automatic evaluation of multilingual text generation of LLMs to new tasks, domains, and languages remains a challenge. Traditional evaluation on benchmark datasets carries the risk of reference data leakage in LLM training or involves additional human annotation effort. The alternative strategy of using another LLM as a scorer also faces uncertainty about the ability of this LLM itself to score non-English text. To address these issues, we propose an annotation-free cross-lingual evaluation protocol for multilingual text generation. Given an LLM candidate to be evaluated and a set of non-English inputs for a particular text generation task, our method first generates English references from the translation of the non-English inputs into English. This is done by an LLM that excels in the equivalent English text generation task. The non-English text generated by the LLM candidate is compared against the generated English references using a cross-lingual evaluation metric to assess the ability of the candidate LLM on multilingual text generation. Our protocol shows a high correlation to the reference-based ROUGE metric in four languages on news text summarization. We also evaluate a diverse set of LLMs in over 90 languages with different prompting strategies to study their multilingual generative abilities.
Shamil Chollampatt, Minh-Quang Pham, Sathish Reddy Indurthi, Marco Turchi
COLING1
2023 CLAD-ST: Contrastive Learning with Adversarial Data for Robust Speech Translation
abstract
The cascaded approach continues to be the most popular choice for speech translation (ST).This approach consists of an automatic speech recognition (ASR) model and a machine translation (MT) model that are used in a pipeline to translate speech in one language to text in another language.MT models are often trained on well-formed text and therefore lack robustness while translating noisy ASR outputs in the cascaded approach, degrading the overall translation quality significantly.We address this robustness problem in downstream MT models by forcing the MT encoder to bring the representations of a noisy input closer to its clean version in the semantic space.This is achieved by introducing a contrastive learning method that leverages adversarial examples in the form of ASR outputs paired with their corresponding human transcripts to optimize the network parameters.In addition, a curriculum learning strategy is then used to stabilize the training by alternating the standard MT log-likelihood loss and the contrastive losses.Our approach achieves significant gains of up to 3 BLEU scores in English-German and English-French speech translation without hurting the translation quality on clean text.
Sathish Indurthi, Shamil Chollampatt, Ravi Agrawal, Marco Turchi
EMNLP2
2023 Select, Prompt, Filter: Distilling Large Language Models for Summarizing Conversations
abstract
Large language models (LLMs) like ChatGPT can be expensive to train, deploy, and use for specific natural language generation tasks such as text summarization and for certain domains.A promising alternative is to fine-tune relatively smaller language models (LMs) on a particular task using high-quality, in-domain datasets.However, it can be prohibitively expensive to get such high-quality training data.This issue has been mitigated by generating weakly supervised data via knowledge distillation (KD) of LLMs.We propose a three-step approach to distill ChatGPT and fine-tune smaller LMs for summarizing forum conversations.More specifically, we design a method to selectively sample a large unannotated corpus of forum conversation using a semantic similarity metric.Then, we use the same metric to retrieve suitable prompts for ChatGPT from a small annotated validation set in the same domain.The generated dataset is then filtered to remove lowquality instances.Our proposed select-promptfilter KD approach leads to significant improvements of up to 6.6 ROUGE-2 score by leveraging sufficient in-domain pseudo-labelled data, over a standard KD approach given the same size of training data.
Minh-Quang Pham, Sathish Indurthi, Shamil Chollampatt, Marco Turchi
EMNLP3
2020 Lexically Constrained Neural Machine Translation with Levenshtein Transformer
abstract
This paper proposes a simple and effective algorithm for incorporating lexical constraints in neural machine translation.Previous work either required re-training existing models with the lexical constraints or incorporating them during beam search decoding with significantly higher computational overheads.Leveraging the flexibility and speed of a recently proposed Levenshtein Transformer model (Gu et al., 2019), our method injects terminology constraints at inference time without any impact on decoding speed.Our method does not require any modification to the training procedure and can be easily applied at runtime with custom dictionaries.Experiments on English-German WMT datasets show that our approach improves an unconstrained baseline and previous approaches.
Raymond Hendy Susanto, Shamil Chollampatt, Liling Tan
ACL2
2020 Can Automatic Post-Editing Improve NMT?
abstract
Automatic post-editing (APE) aims to improve machine translations, thereby reducing human post-editing effort.APE has had notable success when used with statistical machine translation (SMT) systems but has not been as successful over neural machine translation (NMT) systems.This has raised questions on the relevance of APE task in the current scenario.However, the training of APE models has been heavily reliant on large-scale artificial corpora combined with only limited human post-edited data.We hypothesize that APE models have been underperforming in improving NMT translations due to the lack of adequate supervision.To ascertain our hypothesis, we compile a larger corpus of human post-edits of English to German NMT.We empirically show that a state-of-art neural APE model trained on this corpus can significantly improve a strong in-domain NMT system, challenging the current understanding in the field.We further investigate the effects of varying training data sizes, using artificial training data, and domain specificity for the APE task.
Shamil Chollampatt, Raymond Hendy Susanto, Liling Tan, Ewa Szymanska
EMNLP (1)1
2019 Cross-Sentence Grammatical Error Correction
abstract
Automatic grammatical error correction (GEC) research has made remarkable progress in the past decade.However, all existing approaches to GEC correct errors by considering a single sentence alone and ignoring crucial cross-sentence context.Some errors can only be corrected reliably using cross-sentence context and models can also benefit from the additional contextual information in correcting other errors.In this paper, we address this serious limitation of existing approaches and improve strong neural encoder-decoder models by appropriately modeling wider contexts.We employ an auxiliary encoder that encodes previous sentences and incorporate the encoding in the decoder via attention and gating mechanisms.Our approach results in statistically significant improvements in overall GEC performance over strong baselines across multiple test sets.Analysis of our cross-sentence GEC model on a synthetic dataset shows high performance in verb tense corrections that require cross-sentence context.
Shamil Chollampatt, Hwee Tou Ng
ACL (1)1
2018 A Multilayer Convolutional Encoder-Decoder Neural Network for Grammatical Error Correction
abstract
We improve automatic correction of grammatical, orthographic, and collocation errors in text using a multilayer convolutional encoder-decoder neural network. The network is initialized with embeddings that make use of character N-gram information to better suit this task. When evaluated on common benchmark test data sets (CoNLL-2014 and JFLEG), our model substantially outperforms all prior neural approaches on this task as well as strong statistical machine translation-based systems with neural and task-specific features trained on the same data. Our analysis shows the superiority of convolutional neural networks over recurrent neural networks such as long short-term memory (LSTM) networks in capturing the local context via attention, and thereby improving the coverage in correcting grammatical errors. By ensembling multiple models, and incorporating an N-gram language model and edit features via rescoring, our novel method becomes the first neural approach to outperform the current state-of-the-art statistical machine translation-based approach, both in terms of grammaticality and fluency.
Shamil Chollampatt, Hwee Tou Ng
AAAI1
2018 A Reassessment of Reference-Based Grammatical Error Correction Metrics
abstract
Several metrics have been proposed for evaluating grammatical error correction (GEC) systems based on grammaticality, fluency, and adequacy of the output sentences. Previous studies of the correlation of these metrics with human quality judgments were inconclusive, due to the lack of appropriate significance tests, discrepancies in the methods, and choice of datasets used. In this paper, we re-evaluate reference-based GEC metrics by measuring the system-level correlations with humans on a large dataset of human judgments of GEC outputs, and by properly conducting statistical significance tests. Our results show no significant advantage of GLEU over MaxMatch (M2), contradicting previous studies that claim GLEU to be superior. For a finer-grained analysis, we additionally evaluate these metrics for their agreement with human judgments at the sentence level. Our sentence-level analysis indicates that comparing GLEU and M2, one metric may be more useful than the other depending on the scenario. We further qualitatively analyze these metrics and our findings show that apart from being less interpretable and non-deterministic, GLEU also produces counter-intuitive scores in commonly occurring test examples.
Shamil Chollampatt, Hwee Tou Ng
COLING1
2018 Neural Quality Estimation of Grammatical Error Correction
abstract
Grammatical error correction (GEC) systems deployed in language learning environments are expected to accurately correct errors in learners' writing.However, in practice, they often produce spurious corrections and fail to correct many errors, thereby misleading learners.This necessitates the estimation of the quality of output sentences produced by GEC systems so that instructors can selectively intervene and re-correct the sentences which are poorly corrected by the system and ensure that learners get accurate feedback.We propose the first neural approach to automatic quality estimation of GEC output sentences that does not employ any hand-crafted features.Our system is trained in a supervised manner on learner sentences and corresponding GEC system outputs with quality score labels computed using human-annotated references.Our neural quality estimation models for GEC show significant improvements over a strong feature-based baseline.We also show that a state-of-the-art GEC system can be improved when quality scores are used as features for reranking the N-best candidates.
Shamil Chollampatt, Hwee Tou Ng
EMNLP1
2016 Adapting Grammatical Error Correction Based on the Native Language of Writers with Neural Network Joint Models
abstract
An important aspect for the task of grammatical error correction (GEC) that has not yet been adequately explored is adaptation based on the native language (L1) of writers, despite the marked influences of L1 on second language (L2) writing.In this paper, we adapt a neural network joint model (NNJM) using L1-specific learner text and integrate it into a statistical machine translation (SMT) based GEC system.Specifically, we train an NNJM on general learner text (not L1-specific) and subsequently train on L1-specific data using a Kullback-Leibler divergence regularized objective function in order to preserve generalization of the model.We incorporate this adapted NNJM as a feature in an SMT-based English GEC system and show that adaptation achieves significant F 0.5 score gains on English texts written by L1 Chinese, Russian, and Spanish writers.
Shamil Chollampatt, Duc Tam Hoang, Hwee Tou Ng
EMNLP1
2016 Neural Network Translation Models for Grammatical Error Correction
Shamil Chollampatt, Kaveh Taghipour, Hwee Tou Ng
IJCAI1
2016 Exploiting N-Best Hypotheses to Improve an SMT Approach to Grammatical Error Correction
Duc Tam Hoang, Shamil Chollampatt, Hwee Tou Ng
IJCAI2