EDBT 2026 Demo / reviewers in the wild / expert
Kayo Yin
dblp:262/3579
· DBLP profile ↗
14ranked-venue papers
9as first author
13since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 9 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Which Attention Heads Matter for In-Context Learning?abstractLarge language models (LLMs) exhibit impressive in-context learning (ICL) capability, enabling them to generate relevant responses from a handful of task demonstrations in the prompt.
Prior studies have suggested two different explanations for the mechanisms behind ICL:
induction heads that find and copy relevant tokens, and function vector (FV) heads whose activations compute a latent encoding of the ICL task.
To better understand which of the two distinct mechanisms drives ICL, we study and compare induction heads and FV heads in 12 language models. Through detailed ablations, we find that few-shot ICL is driven primarily by FV heads, especially in larger models. We also find that FV and induction heads are connected: many FV heads
start as induction heads during training before transitioning to the FV mechanism. This leads us to speculate that induction facilitates learning the more complex FV mechanism for ICL. Kayo Yin, Jacob Steinhardt |
ICML | 1 |
| 2025 | SHADES: Towards a Multilingual Assessment of Stereotypes in Large Language ModelsabstractMargaret Mitchell, Giuseppe Attanasio, Ioana Baldini, Miruna Clinciu, Jordan Clive, Pieter Delobelle, Manan Dey, Sil Hamilton, Timm Dill, Jad Doughman, Ritam Dutt, Avijit Ghosh, Jessica Zosa Forde, Carolin Holtermann, Lucie-Aimée Kaffee, Tanmay Laud, Anne Lauscher, Roberto L Lopez-Davila, Maraim Masoud, Nikita Nangia, Anaelia Ovalle, Giada Pistilli, Dragomir Radev, Beatrice Savoldi, Vipul Raheja, Jeremy Qin, Esther Ploeger, Arjun Subramonian, Kaustubh Dhole, Kaiser Sun, Amirbek Djanibekov, Jonibek Mansurov, Kayo Yin, Emilio Villa Cueva, Sagnik Mukherjee, Jerry Huang, Xudong Shen, Jay Gala, Hamdan Al-Ali, Tair Djanibekov, Nurdaulet Mukhituly, Shangrui Nie, Shanya Sharma, Karolina Stanczak, Eliza Szczechla, Tiago Timponi Torrent, Deepak Tunuguntla, Marcelo Viridiano, Oskar Van Der Wal, Adina Yakefu, Aurélie Névéol, Mike Zhang, Sydney Zink, Zeerak Talat. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Margaret Mitchell, Giuseppe Attanasio, Ioana Baldini, Miruna-Adriana Clinciu, Jordan Clive, Pieter Delobelle, Manan Dey, Sil Hamilton, Timm Dill, Jad Doughman, Ritam Dutt, Avijit Ghosh, Jessica Zosa Forde, Carolin Holtermann, Lucie-Aimée Kaffee, Tanmay Laud, Anne Lauscher, Roberto L. Lopez-Davila, Maraim Masoud, Nikita Nangia, Anaelia Ovalle, Giada Pistilli, Dragomir R. Radev, Beatrice Savoldi, Vipul Raheja, Jeremy Qin, Esther Ploeger, Arjun Subramonian, Kaustubh D. Dhole, Kaiser Sun, Amirbek Djanibekov, Jonibek Mansurov, Kayo Yin, Emilio Villa Cueva, Sagnik Mukherjee, Jerry Huang, Jay Gala, Hamdan Al-Ali, Tair Djanibekov, Nurdaulet Mukhituly, Shangrui Nie, Shanya Sharma, Karolina Stanczak, Eliza Szczechla, Tiago Timponi Torrent, Deepak Tunuguntla, Marcelo Viridiano, Oskar Van Der Wal, Adina Yakefu, Aurélie Névéol, Mike Zhang, Sydney Zink, Zeerak Talat |
NAACL (Long Papers) | 33 |
| 2024 | American Sign Language Handshapes Reflect Pressures for Communicative EfficiencyabstractCommunicative efficiency is a key topic in linguistics and cognitive psychology, with many studies demonstrating how the pressure to communicate with minimal effort guides the form of natural language. However, this phenomenon is rarely explored in signed languages. This paper shows how handshapes in American Sign Language (ASL) reflect these efficiency pressures and provides new evidence of communicative efficiency in the visual-gestural modality.We focus on hand configurations in native ASL signs and signs borrowed from English to compare efficiency pressures from both ASL and English usage. First, we develop new methodologies to quantify the articulatory effort needed to produce handshapes and the perceptual effort required to recognize them. Then, we analyze correlations between communicative effort and usage statistics in ASL or English. Our findings reveal that frequent ASL handshapes are easier to produce and that pressures for communicative efficiency mostly come from ASL usage, rather than from English lexical borrowing. Kayo Yin, Terry Regier, Daniel Klein 0001 |
ACL (1) | 1 |
| 2024 | Using Language Models to Disambiguate Lexical Choices in TranslationabstractIn translation, a concept represented by a single word in a source language can have multiple variations in a target language.The task of lexical selection requires using context to identify which variation is most appropriate for a source text.We work with native speakers of nine languages to create DTAiLS, a dataset of 1,377 sentence pairs that exhibit cross-lingual concept variation when translating from English.We evaluate recent LLMs and neural machine translation systems on DTAiLS, with the bestperforming model, GPT-4, achieving from 67 to 85% accuracy across languages.Finally, we use language models to generate English rules describing target-language concept variations.Providing weaker models with high-quality lexical rules improves accuracy substantially, in some cases reaching or outperforming GPT-4. Concept: date (fruit)Variations and Generated Rules Khorma refers to the fruit of the date palm when it is fully ripe and dried.It is commonly consumed as a sweet, chewy snack or used in various dishes, particularly desserts.Rotab refers to fresh, soft dates that are partially ripe.These dates are moister and sweeter than fully ripe, dried dates (khorma).Rotab is often eaten as a fresh fruit or used in cooking where a softer, sweeter texture is desired.Kharak refers to dates that are unripe and are in a semi-dried state.They are less sweet compared to rotab and khorma and are often used in cooking or further processed into other forms. Josh Barua, Sanjay Subramanian, Kayo Yin, Alane Suhr |
EMNLP | 3 |
| 2024 | ASL STEM Wiki: Dataset and Benchmark for Interpreting STEM ArticlesabstractKayo Yin, Chinmay Singh, Fyodor O Minakov, Vanessa Milan, Hal Daumé Iii, Cyril Zhang, Alex Xijie Lu, Danielle Bragg. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Kayo Yin, Chinmay Singh, Fyodor O. Minakov, Vanessa Milan, Hal Daumé III, Cyril Zhang, Alex Lu 0002, Danielle Bragg |
EMNLP | 1 |
| 2023 | When Does Translation Require Context? A Data-driven, Multilingual ExplorationabstractAlthough proper handling of discourse significantly contributes to the quality of machine translation (MT), these improvements are not adequately measured in common translation quality metrics.Recent works in context-aware MT attempt to target a small set of discourse phenomena during evaluation, however not in a fully systematic way.In this paper, we develop the Multilingual Discourse-Aware (MUDA) benchmark, a series of taggers that identify and evaluate model performance on discourse phenomena in any given dataset.The choice of phenomena is inspired by a novel methodology to systematically identify translations requiring context.We confirm the difficulty of previously studied phenomena while uncovering others that were previously unaddressed.We find that common context-aware MT models make only marginal improvements over context-agnostic models, which suggests these models do not handle these ambiguities effectively.We release code and data for 14 language pairs to encourage the MT community to focus on accurately capturing discourse phenomena.1 Patrick Fernandes, Kayo Yin, Emmy Liu, André F. T. Martins, Graham Neubig |
ACL (1) | 2 |
| 2022 | Interpreting Language Models with Contrastive ExplanationsabstractModel interpretability methods are often used to explain NLP model decisions on tasks such as text classification, where the output space is relatively small.However, when applied to language generation, where the output space often consists of tens of thousands of tokens, these methods are unable to provide informative explanations.Language models must consider various features to predict a token, such as its part of speech, number, tense, or semantics.Existing explanation methods conflate evidence for all these features into a single explanation, which is less interpretable for human understanding.To disentangle the different decisions in language modeling, we focus on explaining language models contrastively: we look for salient input tokens that explain why the model predicted one token instead of another.We demonstrate that contrastive explanations are quantifiably better than non-contrastive explanations in verifying major grammatical phenomena, and that they significantly improve contrastive model simulatability for human observers.We also identify groups of contrastive decisions where the model uses similar evidence, and we are able to characterize what input tokens models use during various language generation decisions.1 Kayo Yin, Graham Neubig |
EMNLP | 1 |
| 2022 | Including Signed Languages in Natural Language Processing (Extended Abstract)abstractSigned languages are the primary means of communication for many deaf and hard of hearing individuals. Since signed languages exhibit all the fundamental linguistic properties of natural language, we believe that tools and theories of Natural Language Processing (NLP) are crucial towards its modeling. However, existing research in Sign Language Processing (SLP) seldom attempt to explore and leverage the linguistic organization of signed languages. This position paper calls on the NLP community to include signed languages as a research area with high social and scientific impact. We first discuss the linguistic properties of signed languages to consider during their modeling. Then, we review the limitations of current SLP models and identify the open challenges to extend NLP to signed languages. Finally, we urge (1) the adoption of an efficient tokenization method; (2) the development of linguistically-informed models; (3) the collection of real-world signed language data; (4) the inclusion of local signed language communities as an active and leading voice in research. Kayo Yin, Malihe Alikhani |
IJCAI | 1 |
| 2021 | Measuring and Increasing Context Usage in Context-Aware Machine TranslationabstractPatrick Fernandes, Kayo Yin, Graham Neubig, André F. T. Martins. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Patrick Fernandes, Kayo Yin, Graham Neubig, André F. T. Martins |
ACL/IJCNLP (1) | 2 |
| 2021 | Do Context-Aware Translation Models Pay the Right Attention?abstractKayo Yin, Patrick Fernandes, Danish Pruthi, Aditi Chaudhary, André F. T. Martins, Graham Neubig. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Kayo Yin, Patrick Fernandes, Danish Pruthi, Aditi Chaudhary, André F. T. Martins, Graham Neubig |
ACL/IJCNLP (1) | 1 |
| 2021 | Including Signed Languages in Natural Language ProcessingabstractKayo Yin, Amit Moryossef, Julie Hochgesang, Yoav Goldberg, Malihe Alikhani. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Kayo Yin, Amit Moryossef, Julie Hochgesang, Yoav Goldberg, Malihe Alikhani |
ACL/IJCNLP (1) | 1 |
| 2021 | When is Wall a Pared and when a Muro?: Extracting Rules Governing Lexical SelectionabstractLearning fine-grained distinctions between vocabulary items is a key challenge in learning a new language.For example, the noun "wall" has different lexical manifestations in Spanish -"pared" refers to an indoor wall while "muro" refers to an outside wall.However, this variety of lexical distinction may not be obvious to non-native learners unless the distinction is explained in such a way.In this work, we present a method for automatically identifying fine-grained lexical distinctions, and extracting concise descriptions explaining these distinctions in a human-and machine-readable format.We confirm the quality of these extracted descriptions in a language learning setup for two languages, Spanish and Greek, where we use them to teach non-native speakers when to translate a given ambiguous word into its different possible translations.Code and data are publicly released here.1 Aditi Chaudhary, Kayo Yin, Antonios Anastasopoulos, Graham Neubig |
EMNLP (1) | 2 |
| 2021 | Signed Coreference ResolutionabstractCoreference resolution is key to many natural language processing tasks and yet has been relatively unexplored in Sign Language Processing.In signed languages, space is primarily used to establish reference.Solving coreference resolution for signed languages would not only enable higher-level Sign Language Processing systems, but also enhance our understanding of language in different modalities and of situated references, which are key problems in studying grounded language.In this paper, we: (1) introduce Signed Coreference Resolution (SCR), a new challenge for coreference modeling and Sign Language Processing; (2) collect an annotated corpus of German Sign Language with gold labels for coreference together with an annotation software for the task; (3) explore features of hand gesture, iconicity, and spatial situated properties and move forward to propose a set of linguistically informed heuristics and unsupervised models for the task; (4) put forward several proposals about ways to address the complexities of this challenge effectively 1 . Kayo Yin, Kenneth DeHaan, Malihe Alikhani |
EMNLP (1) | 1 |
| 2020 | Better Sign Language Translation with STMC-TransformerabstractSign Language Translation (SLT) first uses a Sign Language Recognition (SLR) system to extract sign language glosses from videos.Then, a translation system generates spoken language translations from the sign language glosses.This paper focuses on the translation system and introduces the STMC-Transformer which improves on the current state-of-the-art by over 5 and 7 BLEU respectively on gloss-to-text and video-to-text translation of the PHOENIX-Weather 2014T dataset.On the ASLG-PC12 corpus, we report an increase of over 16 BLEU.We also demonstrate the problem in current methods that rely on gloss supervision.The videoto-text translation of our STMC-Transformer outperforms translation of GT glosses.This contradicts previous claims that GT gloss translation acts as an upper bound for SLT performance and reveals that glosses are an inefficient representation of sign language.For future SLT research, we therefore suggest an end-to-end training of the recognition and translation models, or using a different sign language annotation scheme. Kayo Yin, Jesse Read |
COLING | 1 |