Igor Shalyminov

dblp:205/8962 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0001-9664-1774ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Faithful, Unfaithful or Ambiguous? Multi-Agent Debate with Initial Stance for Summary Evaluation
abstract
Mahnaz Koupaee, Jake W. Vincent, Saab Mansour, Igor Shalyminov, Han He, Hwanjun Song, Raphael Shu, Jianfeng He, Yi Nian, Amy Wing-mei Wong, Kyu J. Han, Hang Su. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Mahnaz Koupaee, Jake W. Vincent, Saab Mansour, Igor Shalyminov, Han He, Hwanjun Song, Raphael Shu, Yi Nian, Amy Wing-mei Wong, Kyu J. Han
NAACL (Long Papers)4
2024 FineSurE: Fine-grained Summarization Evaluation using LLMs
abstract
Automated evaluation is crucial for streamlining text summarization benchmarking and model development, given the costly and timeconsuming nature of human evaluation.Traditional methods like ROUGE do not correlate well with human judgment, while recently proposed LLM-based metrics provide only summary-level assessment using Likertscale scores.This limits deeper model analysis, e.g., we can only assign one hallucination score at the summary level, while at the sentence level, we can count sentences containing hallucinations.To remedy those limitations, we propose FineSurE, a fine-grained evaluator specifically tailored for the summarization task using large language models (LLMs).It also employs completeness and conciseness criteria, in addition to faithfulness, enabling multi-dimensional assessment.We compare various open-source and proprietary LLMs as backbones for FineSurE.In addition, we conduct extensive benchmarking of FineSurE against SOTA methods including NLI-, QA-, and LLM-based methods, showing improved performance especially on the completeness and conciseness dimensions.The code is available at https://github.com/ DISL-Lab/FineSurE-ACL24.
Hwanjun Song, Igor Shalyminov, Jason Cai, Saab Mansour
ACL (1)3
2024 Can Your Model Tell a Negation from an Implicature? Unravelling Challenges With Intent Encoders
abstract
Yuwei Zhang, Siffi Singh, Sailik Sengupta, Igor Shalyminov, Hang Su, Hwanjun Song, Saab Mansour. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Siffi Singh, Sailik Sengupta, Igor Shalyminov, Hwanjun Song, Saab Mansour
ACL (1)4
2024 Controllable Contextualized Image Captioning: Directing the Visual Narrative Through User-Defined Highlights
Shunqi Mao, Chaoyi Zhang, Hwanjun Song, Igor Shalyminov, Tom Weidong Cai
ECCV (50)5
2024 MAGID: An Automated Pipeline for Generating Synthetic Multi-modal Datasets
abstract
Hossein Aboutalebi, Hwanjun Song, Yusheng Xie, Arshit Gupta, Lijia Sun, Hang Su, Igor Shalyminov, Nikolaos Pappas, Siffi Singh, Saab Mansour. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Hossein Aboutalebi, Hwanjun Song, Yusheng Xie, Arshit Gupta, Lijia Sun, Igor Shalyminov, Nikolaos Pappas 0004, Siffi Singh, Saab Mansour
NAACL-HLT7
2024 CERET: Cost-Effective Extrinsic Refinement for Text Generation
abstract
Jason Cai, Hang Su, Monica Sunkara, Igor Shalyminov, Saab Mansour. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jason Cai, Monica Sunkara, Igor Shalyminov, Saab Mansour
NAACL-HLT4
2024 Semi-Supervised Dialogue Abstractive Summarization via High-Quality Pseudolabel Selection
abstract
Jianfeng He, Hang Su, Jason Cai, Igor Shalyminov, Hwanjun Song, Saab Mansour. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jason Cai, Igor Shalyminov, Hwanjun Song, Saab Mansour
NAACL-HLT4
2024 TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization
abstract
Liyan Tang, Igor Shalyminov, Amy Wong, Jon Burnsky, Jake Vincent, Yu’an Yang, Siffi Singh, Song Feng, Hwanjun Song, Hang Su, Lijia Sun, Yi Zhang, Saab Mansour, Kathleen McKeown. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Liyan Tang, Igor Shalyminov, Amy Wing-mei Wong, Jon Burnsky, Jake W. Vincent, Siffi Singh, Song Feng 0001, Hwanjun Song, Lijia Sun, Yi Zhang 0053, Saab Mansour, Kathy McKeown
NAACL-HLT2
2021 GRTr: Generative-Retrieval Transformers for Data-Efficient Dialogue Domain Adaptation
abstract
Domain adaptation has recently become a key problem in dialogue systems research. Deep learning, while being the preferred technique for modeling such systems, works best given massive training data. However, in real-world scenarios, such resources are rarely available for new domains, and the ability to train with a few dialogue examples can be considered essential. Pre-training on large data sources and adapting to the target data has become the standard method for few-shot problems within the deep learning framework. In this paper, we presentgrtr, a hybrid generative-retrieval model based on the large-scale general-purpose language model GPT[2] fine-tuned to the multi-domainmetalwoz dataset. In addition to robust and diverse response generation provided by the GPT[2], our model is able to estimate generation confidence, and is equipped with retrieval logic as a fallback for the cases when the estimate is low.grtr is the winning entry at the fast domain adaptation task of DSTC-8 in human evaluation ($>$4% improvement over the 2nd place system). It also attains superior performance to a series of baselines on automated metrics onmetalwoz andmultiwoz, a multi-domain dataset of goal-oriented dialogues. In this paper, we also conduct a study ofgrtr's performance in the setup of limited adaptation data, evaluating the model's overall response prediction performance onmetalwoz and goal-oriented performance onmultiwoz.
Igor Shalyminov, Alessandro Sordoni, Adam Atkinson, Hannes Schulz
IEEE ACM Trans. Audio Speech Lang. Process.1
2020 Fast Domain Adaptation for Goal-Oriented Dialogue Using a Hybrid Generative-Retrieval Transformer
abstract
Goal-oriented dialogue systems are now widely adopted in industry, where practical aspects of using them becomes of key importance. As such, it is expected from such systems to fit into a rapid prototyping cycle for new products and domains. For data-driven dialogue systems (especially those based on deep learning) that amounts to maintaining production-level performance having been provided with a few `seed' dialogue examples, normally referred to as data efficiency.With extremely data-dependent deep learning methods, the most promising way to achieve practical data efficiency is transfer learning-i.e., leveraging a greater, highly represented data source for training a base model, then fine-tuning it to available in-domain data.In this paper, we present a hybrid generative-retrieval model that can be trained using transfer learning. By using GPT-2 as the base model and fine-tuning it to the multidomain MetaLWOz dataset, we obtain a robust dialogue model able to perform both response generation and ranking1. Combining both, it outperforms several competitive generative-only and retrieval-only baselines, measured by language modeling quality on MetaLWOz as well as in goal- oriented metrics (Intent/Slot Fl-scores) on the MultiWoz corpus.
Igor Shalyminov, Alessandro Sordoni, Adam Atkinson, Hannes Schulz
ICASSP1
2019 Data-Efficient Goal-Oriented Conversation with Dialogue Knowledge Transfer Networks
abstract
Igor Shalyminov, Sungjin Lee, Arash Eshghi, Oliver Lemon. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Igor Shalyminov, Arash Eshghi, Oliver Lemon
EMNLP/IJCNLP (1)1
2019 Contextual Out-of-domain Utterance Handling with Counterfeit Data Augmentation
abstract
Neural dialog models often lack robustness to anomalous user input and produce inappropriate responses which leads to frustrating user experience. Although there are a set of prior approaches to out-of-domain (OOD) utterance detection, they share a few restrictions: they rely on OOD data or multiple sub-domains, and their OOD detection is context-independent which leads to suboptimal performance in a dialog. The goal of this paper is to propose a novel OOD detection method that does not require OOD data by utilizing counterfeit OOD turns in the context of a dialog. For the sake of fostering further research, we also release new dialog datasets which are 3 publicly available dialog corpora augmented with OOD turns in a controllable way. Our method outperforms state-of-the-art dialog models equipped with a conventional OOD detection mechanism by a large margin in the presence of OOD utterances.
Igor Shalyminov
ICASSP2
2019 Few-Shot Dialogue Generation Without Annotated Data: A Transfer Learning Approach
abstract
Learning with minimal data is one of the key challenges in the development of practical, production-ready goal-oriented dialogue systems.In a real-world enterprise setting where dialogue systems are developed rapidly and are expected to work robustly for an evergrowing variety of domains, products, and scenarios, efficient learning from a limited number of examples becomes indispensable.In this paper, we introduce a technique to achieve state-of-the-art dialogue generation performance in a few-shot setup, without using any annotated data.We do this by leveraging background knowledge from a larger, more highly represented dialogue sourcenamely, the MetaLWOz dataset.We evaluate our model on the Stanford Multi-Domain Dialogue Dataset, consisting of human-human goal-oriented dialogues in in-car navigation, appointment scheduling, and weather information domains.We show that our few-shot approach achieves state-of-the art results on that dataset by consistently outperforming the previous best model in terms of BLEU and Entity F1 scores, while being more data-efficient by not requiring any data annotation.
Igor Shalyminov, Arash Eshghi, Oliver Lemon
SIGdial1
2017 Bootstrapping incremental dialogue systems from minimal data: the generalisation power of dialogue grammars
abstract
We investigate an end-to-end method for automatically inducing task-based dialogue systems from small amounts of unannotated dialogue data.It combines an incremental semantic grammar -Dynamic Syntax and Type Theory with Records (DS-TTR) -with Reinforcement Learning (RL), where language generation and dialogue management are a joint decision problem.The systems thus produced are incremental: dialogues are processed word-by-word, shown previously to be essential in supporting natural, spontaneous dialogue.We hypothesised that the rich linguistic knowledge within the grammar should enable a combinatorially large number of dialogue variations to be processed, even when trained on very few dialogues.Our experiments show that our model can process 74% of the Facebook AI bAbI dataset even when trained on only 0.13% of the data (5 dialogues).It can in addition process 65% of bAbI+, a corpus 1 we created by systematically adding incremental dialogue phenomena such as restarts and self-corrections to bAbI.We compare our model with a state-of-theart retrieval model, memn2n (Bordes et al., 2017).We find that, in terms of semantic accuracy, memn2n shows very poor robustness to the bAbI+ transformations even when trained on the full bAbI dataset.
Arash Eshghi, Igor Shalyminov, Oliver Lemon
EMNLP2