VLDB 2026 Research / reviewers in the wild / expert
Igor Shalyminov
dblp:205/8962
· DBLP profile ↗
14ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0001-9664-1774ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Faithful, Unfaithful or Ambiguous? Multi-Agent Debate with Initial Stance for Summary EvaluationabstractMahnaz Koupaee, Jake W. Vincent, Saab Mansour, Igor Shalyminov, Han He, Hwanjun Song, Raphael Shu, Jianfeng He, Yi Nian, Amy Wing-mei Wong, Kyu J. Han, Hang Su. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Mahnaz Koupaee, Jake W. Vincent, Saab Mansour, Igor Shalyminov, Han He, Hwanjun Song, Raphael Shu, Yi Nian, Amy Wing-mei Wong, Kyu J. Han |
NAACL (Long Papers) | 4 |
| 2024 | FineSurE: Fine-grained Summarization Evaluation using LLMsabstractAutomated evaluation is crucial for streamlining text summarization benchmarking and model development, given the costly and timeconsuming nature of human evaluation.Traditional methods like ROUGE do not correlate well with human judgment, while recently proposed LLM-based metrics provide only summary-level assessment using Likertscale scores.This limits deeper model analysis, e.g., we can only assign one hallucination score at the summary level, while at the sentence level, we can count sentences containing hallucinations.To remedy those limitations, we propose FineSurE, a fine-grained evaluator specifically tailored for the summarization task using large language models (LLMs).It also employs completeness and conciseness criteria, in addition to faithfulness, enabling multi-dimensional assessment.We compare various open-source and proprietary LLMs as backbones for FineSurE.In addition, we conduct extensive benchmarking of FineSurE against SOTA methods including NLI-, QA-, and LLM-based methods, showing improved performance especially on the completeness and conciseness dimensions.The code is available at https://github.com/ DISL-Lab/FineSurE-ACL24. Hwanjun Song, Igor Shalyminov, Jason Cai, Saab Mansour |
ACL (1) | 3 |
| 2024 | Can Your Model Tell a Negation from an Implicature? Unravelling Challenges With Intent EncodersabstractYuwei Zhang, Siffi Singh, Sailik Sengupta, Igor Shalyminov, Hang Su, Hwanjun Song, Saab Mansour. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Siffi Singh, Sailik Sengupta, Igor Shalyminov, Hwanjun Song, Saab Mansour |
ACL (1) | 4 |
| 2024 | Controllable Contextualized Image Captioning: Directing the Visual Narrative Through User-Defined Highlights
Shunqi Mao, Chaoyi Zhang, Hwanjun Song, Igor Shalyminov, Tom Weidong Cai |
ECCV (50) | 5 |
| 2024 | MAGID: An Automated Pipeline for Generating Synthetic Multi-modal DatasetsabstractHossein Aboutalebi, Hwanjun Song, Yusheng Xie, Arshit Gupta, Lijia Sun, Hang Su, Igor Shalyminov, Nikolaos Pappas, Siffi Singh, Saab Mansour. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Hossein Aboutalebi, Hwanjun Song, Yusheng Xie, Arshit Gupta, Lijia Sun, Igor Shalyminov, Nikolaos Pappas 0004, Siffi Singh, Saab Mansour |
NAACL-HLT | 7 |
| 2024 | CERET: Cost-Effective Extrinsic Refinement for Text GenerationabstractJason Cai, Hang Su, Monica Sunkara, Igor Shalyminov, Saab Mansour. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jason Cai, Monica Sunkara, Igor Shalyminov, Saab Mansour |
NAACL-HLT | 4 |
| 2024 | Semi-Supervised Dialogue Abstractive Summarization via High-Quality Pseudolabel SelectionabstractJianfeng He, Hang Su, Jason Cai, Igor Shalyminov, Hwanjun Song, Saab Mansour. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jason Cai, Igor Shalyminov, Hwanjun Song, Saab Mansour |
NAACL-HLT | 4 |
| 2024 | TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue SummarizationabstractLiyan Tang, Igor Shalyminov, Amy Wong, Jon Burnsky, Jake Vincent, Yu’an Yang, Siffi Singh, Song Feng, Hwanjun Song, Hang Su, Lijia Sun, Yi Zhang, Saab Mansour, Kathleen McKeown. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Liyan Tang, Igor Shalyminov, Amy Wing-mei Wong, Jon Burnsky, Jake W. Vincent, Siffi Singh, Song Feng 0001, Hwanjun Song, Lijia Sun, Yi Zhang 0053, Saab Mansour, Kathy McKeown |
NAACL-HLT | 2 |
| 2021 | GRTr: Generative-Retrieval Transformers for Data-Efficient Dialogue Domain AdaptationabstractDomain adaptation has recently become a key problem in dialogue systems research. Deep learning, while being the preferred technique for modeling such systems, works best given massive training data. However, in real-world scenarios, such resources are rarely available for new domains, and the ability to train with a few dialogue examples can be considered essential. Pre-training on large data sources and adapting to the target data has become the standard method for few-shot problems within the deep learning framework. In this paper, we presentgrtr, a hybrid generative-retrieval model based on the large-scale general-purpose language model GPT[2] fine-tuned to the multi-domainmetalwoz dataset. In addition to robust and diverse response generation provided by the GPT[2], our model is able to estimate generation confidence, and is equipped with retrieval logic as a fallback for the cases when the estimate is low.grtr is the winning entry at the fast domain adaptation task of DSTC-8 in human evaluation ($>$4% improvement over the 2nd place system). It also attains superior performance to a series of baselines on automated metrics onmetalwoz andmultiwoz, a multi-domain dataset of goal-oriented dialogues. In this paper, we also conduct a study ofgrtr's performance in the setup of limited adaptation data, evaluating the model's overall response prediction performance onmetalwoz and goal-oriented performance onmultiwoz. Igor Shalyminov, Alessandro Sordoni, Adam Atkinson, Hannes Schulz |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | Fast Domain Adaptation for Goal-Oriented Dialogue Using a Hybrid Generative-Retrieval TransformerabstractGoal-oriented dialogue systems are now widely adopted in industry, where practical aspects of using them becomes of key importance. As such, it is expected from such systems to fit into a rapid prototyping cycle for new products and domains. For data-driven dialogue systems (especially those based on deep learning) that amounts to maintaining production-level performance having been provided with a few `seed' dialogue examples, normally referred to as data efficiency.With extremely data-dependent deep learning methods, the most promising way to achieve practical data efficiency is transfer learning-i.e., leveraging a greater, highly represented data source for training a base model, then fine-tuning it to available in-domain data.In this paper, we present a hybrid generative-retrieval model that can be trained using transfer learning. By using GPT-2 as the base model and fine-tuning it to the multidomain MetaLWOz dataset, we obtain a robust dialogue model able to perform both response generation and ranking1. Combining both, it outperforms several competitive generative-only and retrieval-only baselines, measured by language modeling quality on MetaLWOz as well as in goal- oriented metrics (Intent/Slot Fl-scores) on the MultiWoz corpus. Igor Shalyminov, Alessandro Sordoni, Adam Atkinson, Hannes Schulz |
ICASSP | 1 |
| 2019 | Data-Efficient Goal-Oriented Conversation with Dialogue Knowledge Transfer NetworksabstractIgor Shalyminov, Sungjin Lee, Arash Eshghi, Oliver Lemon. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Igor Shalyminov, Arash Eshghi, Oliver Lemon |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Contextual Out-of-domain Utterance Handling with Counterfeit Data AugmentationabstractNeural dialog models often lack robustness to anomalous user input and produce inappropriate responses which leads to frustrating user experience. Although there are a set of prior approaches to out-of-domain (OOD) utterance detection, they share a few restrictions: they rely on OOD data or multiple sub-domains, and their OOD detection is context-independent which leads to suboptimal performance in a dialog. The goal of this paper is to propose a novel OOD detection method that does not require OOD data by utilizing counterfeit OOD turns in the context of a dialog. For the sake of fostering further research, we also release new dialog datasets which are 3 publicly available dialog corpora augmented with OOD turns in a controllable way. Our method outperforms state-of-the-art dialog models equipped with a conventional OOD detection mechanism by a large margin in the presence of OOD utterances. Igor Shalyminov |
ICASSP | 2 |
| 2019 | Few-Shot Dialogue Generation Without Annotated Data: A Transfer Learning ApproachabstractLearning with minimal data is one of the key challenges in the development of practical, production-ready goal-oriented dialogue systems.In a real-world enterprise setting where dialogue systems are developed rapidly and are expected to work robustly for an evergrowing variety of domains, products, and scenarios, efficient learning from a limited number of examples becomes indispensable.In this paper, we introduce a technique to achieve state-of-the-art dialogue generation performance in a few-shot setup, without using any annotated data.We do this by leveraging background knowledge from a larger, more highly represented dialogue sourcenamely, the MetaLWOz dataset.We evaluate our model on the Stanford Multi-Domain Dialogue Dataset, consisting of human-human goal-oriented dialogues in in-car navigation, appointment scheduling, and weather information domains.We show that our few-shot approach achieves state-of-the art results on that dataset by consistently outperforming the previous best model in terms of BLEU and Entity F1 scores, while being more data-efficient by not requiring any data annotation. Igor Shalyminov, Arash Eshghi, Oliver Lemon |
SIGdial | 1 |
| 2017 | Bootstrapping incremental dialogue systems from minimal data: the generalisation power of dialogue grammarsabstractWe investigate an end-to-end method for automatically inducing task-based dialogue systems from small amounts of unannotated dialogue data.It combines an incremental semantic grammar -Dynamic Syntax and Type Theory with Records (DS-TTR) -with Reinforcement Learning (RL), where language generation and dialogue management are a joint decision problem.The systems thus produced are incremental: dialogues are processed word-by-word, shown previously to be essential in supporting natural, spontaneous dialogue.We hypothesised that the rich linguistic knowledge within the grammar should enable a combinatorially large number of dialogue variations to be processed, even when trained on very few dialogues.Our experiments show that our model can process 74% of the Facebook AI bAbI dataset even when trained on only 0.13% of the data (5 dialogues).It can in addition process 65% of bAbI+, a corpus 1 we created by systematically adding incremental dialogue phenomena such as restarts and self-corrections to bAbI.We compare our model with a state-of-theart retrieval model, memn2n (Bordes et al., 2017).We find that, in terms of semantic accuracy, memn2n shows very poor robustness to the bAbI+ transformations even when trained on the full bAbI dataset. Arash Eshghi, Igor Shalyminov, Oliver Lemon |
EMNLP | 2 |