Shane Storks

dblp:239/4098 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0002-5826-4426ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Language models and text generation · 24% Knowledge representation and reasoning · 21% Vision and language · 14%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computing education · 100%
Software engineering, system software, and programming languages
1 paper
Empirical software engineering · 100%

Topics — the 19 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
in-context learning
1.422024
Eliciting In-Context Learning in Vision-Language Models for Videos Through Curated Data Distributional Properties · EMNLP 2024
In-Context Analogical Reasoning with Pre-Trained Language Models · ACL (1) 2023
Knowledge, reasoning and agents › Multi-agent systems › emergent communication
compositional language emergence
1.012026
Discovering Properties of Inflectional Morphology in Neural Emergent Communication · ACL (1) 2026
Knowledge, reasoning and agents › Multi-agent systems
emergent communication
1.012026
Discovering Properties of Inflectional Morphology in Neural Emergent Communication · ACL (1) 2026
Machine learning › Trustworthy machine learning
interpretability
1.012026
Sparse Feature Coactivation Reveals Causal Semantic Modules in Large Language Models · ACL (1) 2026
Natural language and speech › Information extraction and text analysis › computational morphology
morphological inflection
1.012026
Discovering Properties of Inflectional Morphology in Neural Emergent Communication · ACL (1) 2026
Knowledge, reasoning and agents › Knowledge representation and reasoning
semantic representation
1.012026
Sparse Feature Coactivation Reveals Causal Semantic Modules in Large Language Models · ACL (1) 2026
Natural language and speech › Language models and text generation › natural language understanding
linguistic knowledge in language models
0.912025
Mind the Gap: How BabyLMs Learn Filler-Gap Dependencies · EMNLP 2025
Computer vision › Vision and language
video-language model
0.812024
Eliciting In-Context Learning in Vision-Language Models for Videos Through Curated Data Distributional Properties · EMNLP 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning
analogical reasoning
0.712023
In-Context Analogical Reasoning with Pre-Trained Language Models · ACL (1) 2023
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.712023
From Heuristic to Analytic: Cognitively Motivated Strategies for Coherent Physical Commonsense Reasoning · EMNLP 2023
Natural language and speech › Language models and text generation
large language model reasoning
0.712023
From Heuristic to Analytic: Cognitively Motivated Strategies for Coherent Physical Commonsense Reasoning · EMNLP 2023
Knowledge, reasoning and agents › Knowledge representation and reasoning › commonsense reasoning
physical commonsense reasoning
0.712023
From Heuristic to Analytic: Cognitively Motivated Strategies for Coherent Physical Commonsense Reasoning · EMNLP 2023
Computing education › AI education
NLP education
0.712023
NLP Reproducibility For All: Understanding Experiences of Beginners · ACL (1) 2023
Empirical software engineering
reproducibility
0.712023
NLP Reproducibility For All: Understanding Experiences of Beginners · ACL (1) 2023
Robotics › Robot navigation and mapping
embodied instruction following
0.612022
DANLI: Deliberative Agent for Following Natural Language Instructions · EMNLP 2022
Natural language and speech › Language models and text generation › instruction following
instruction-following agents
0.612022
DANLI: Deliberative Agent for Following Natural Language Instructions · EMNLP 2022
Computer vision › Vision and language
vision-language model
0.312025
Transparent and Coherent Procedural Mistake Detection · EMNLP 2025
Computer vision › Video understanding and tracking › multimodal video understanding
video narration
0.212024
Eliciting In-Context Learning in Vision-Language Models for Videos Through Curated Data Distributional Properties · EMNLP 2024
Computer vision › Vision and language
visual reasoning
0.212023
In-Context Analogical Reasoning with Pre-Trained Language Models · ACL (1) 2023

Methods — techniques the papers use, named apart from their topics

fine-tuning · 1.5user study · 1.3survey · 1.3sparse autoencoder · 1.0feature coactivation · 1.0deep neural network · 1.0wh-licensing scores · 0.9natural language inference · 0.9grammaticality contrasts · 0.9flip tests · 0.9training paradigm · 0.8data distribution properties · 0.8
YearPublicationVenuePosition
2026 Sparse Feature Coactivation Reveals Causal Semantic Modules in Large Language Models
abstract
Ruixuan Deng, Xiaoyang Hu, Miles Gilberti, Shane Storks, Aman Taxali, Mike Angstadt, Chandra Sripada, Joyce Chai. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ruixuan Deng, Xiaoyang Hu, Miles Gilberti, Shane Storks, Aman Taxali, Michael Angstadt, Chandra Sekhar Sripada, Joyce Y. Chai
ACL (1)4
2026 Discovering Properties of Inflectional Morphology in Neural Emergent Communication
abstract
Emergent communication (EmCom) with deep neural network-based agents promises to yield insights into the nature of human language, but remains focused primarily on a few subfieldspecific goals and metrics that prioritize communication schemes which represent attributes with unique characters one-to-one and compose them syntactically.We thus reinterpret a common EmCom setting, the attribute-value reconstruction game, by imposing a smallvocabulary constraint to simulate double articulation, and formulating a novel setting analogous to naturalistic inflectional morphology (enabling meaningful comparison to natural language communication schemes).We develop new metrics and explore variations of this game motivated by real properties of inflectional morphology: concatenativity and fusion.Through our experiments, we discover that simulated phonological constraints encourage concatenative morphology, and emergent languages replicate the tendency of natural languages to fuse grammatical attributes.
Miles Gilberti, Shane Storks, Huteng Dai
ACL (1)2
2025 Mind the Gap: How BabyLMs Learn Filler-Gap Dependencies
abstract
Humans acquire syntactic constructions like filler-gap dependencies from limited and often noisy input.Can neural language models do the same?We investigate this question by evaluating GPT-2 models trained on childoriented input from the BabyLM Challenge.Our experiments focus on whether these "baby" language models acquire filler-gap dependencies, generalize across constructions, and respect structural constraints such as island effects.We apply a suite of syntactic constructions to four models trained on child language, including two base models (trained on 10M and 100M tokens) and two well-performing models from the BabyLM Challenge (ConcreteGPT and BabbleGPT).We evaluate model behavior using wh-licensing scores, flip tests, and grammaticality contrasts across four constructions.Results show that BabyLM-scale models partially acquire filler-gap dependencies but often fail to generalize or fully capture island constraints.Our code and datasets are available at https://github.com/um-cap-lab/ EMNLP-2025-submission.
Chi-Yun Chang, Xueyang Huang, Humaira Nasir, Shane Storks, Olawale Akingbade, Huteng Dai
EMNLP4
2025 Transparent and Coherent Procedural Mistake Detection
abstract
Procedural mistake detection (PMD) is a challenging problem of classifying whether a human user (observed through egocentric video) has successfully executed a task (specified by a procedural text).Despite significant recent efforts, machine performance in the wild remains nonviable, and the reasoning processes underlying this performance are opaque.As such, we extend PMD to require generating visual self-dialog rationales to inform decisions.Given the impressive, mature image understanding capabilities observed in recent visionand-language models (VLMs), we curate a suitable benchmark dataset for PMD based on individual frames.As our reformulation enables unprecedented transparency, we leverage a natural language inference (NLI) model to formulate two automated metrics for the coherence of generated rationales.We establish baselines for this reframed task, showing that VLMs struggle off-the-shelf, but with some trade-offs, their accuracy, coherence, and efficiency can be improved by incorporating these metrics into common inference and finetuning methods.Lastly, our multi-faceted metrics visualize common outcomes, highlighting areas for further improvement.
Shane Storks, Itamar Bar-Yossef, Yayuan Li, Jason J. Corso, Joyce Y. Chai
EMNLP1
2024 Eliciting In-Context Learning in Vision-Language Models for Videos Through Curated Data Distributional Properties
abstract
A major reason behind the recent success of large language models (LLMs) is their incontext learning capability, which makes it possible to rapidly adapt them to downstream textbased tasks by prompting them with a small number of relevant demonstrations.While large vision-language models (VLMs) have recently been developed for tasks requiring both text and images, they largely lack in-context learning over visual information, especially in understanding and generating text about videos.In this work, we implement Emergent In-context Learning on Videos (EILeV), a novel training paradigm that induces in-context learning over video and text by capturing key properties of pre-training data found by prior work to be essential for in-context learning in transformers.In our experiments, we show that EILeV-trained models outperform other off-the-shelf VLMs in few-shot video narration for novel, rare actions.Furthermore, we demonstrate that these key properties of bursty distributions, skewed marginal distributions, and dynamic meaning each contribute to varying degrees to VLMs' in-context learning capability in narrating procedural videos.Our results, analysis, and EILeV-trained models yield numerous insights about the emergence of in-context learning over video and text, creating a foundation for future work to optimize and scale VLMs for open-domain video understanding and reasoning.1
Keunwoo Peter Yu, Fengyuan Hu, Shane Storks, Joyce Y. Chai
EMNLP4
2023 In-Context Analogical Reasoning with Pre-Trained Language Models
abstract
Analogical reasoning is a fundamental capacity of human cognition that allows us to reason abstractly about novel situations by relating them to past experiences.While it is thought to be essential for robust reasoning in AI systems, conventional approaches require significant training and/or hard-coding of domain knowledge to be applied to benchmark tasks.Inspired by cognitive science research that has found connections between human language and analogy-making, we explore the use of intuitive language-based abstractions to support analogy in AI systems.Specifically, we apply large pre-trained language models (PLMs) to visual Raven's Progressive Matrices (RPM), a common relational reasoning test.By simply encoding the perceptual features of the problem into language form, we find that PLMs exhibit a striking capacity for zero-shot relational reasoning, exceeding human performance and nearing supervised vision-based methods.We explore different encodings that vary the level of abstraction over task features, finding that higherlevel abstractions further strengthen PLMs' analogical reasoning.Our detailed analysis reveals insights on the role of model complexity, incontext learning, and prior knowledge in solving RPM tasks.
Xiaoyang Hu, Shane Storks, Richard L. Lewis, Joyce Y. Chai
ACL (1)2
2023 NLP Reproducibility For All: Understanding Experiences of Beginners
abstract
As natural language processing (NLP) has recently seen an unprecedented level of excitement, and more people are eager to enter the field, it is unclear whether current research reproducibility efforts are sufficient for this group of beginners to apply the latest developments.To understand their needs, we conducted a study with 93 students in an introductory NLP course, where students reproduced the results of recent NLP papers.Surprisingly, we find that their programming skill and comprehension of research papers have a limited impact on their effort spent completing the exercise.Instead, we find accessibility efforts by research authors to be the key to success, including complete documentation, better coding practice, and easier access to data files.Going forward, we recommend that NLP researchers pay close attention to these simple aspects of open-sourcing their work, and use insights from beginners' feedback to provide actionable ideas on how to better support them.
Shane Storks, Keunwoo Peter Yu, Ziqiao Ma 0001, Joyce Y. Chai
ACL (1)1
2023 From Heuristic to Analytic: Cognitively Motivated Strategies for Coherent Physical Commonsense Reasoning
abstract
Pre-trained language models (PLMs) have shown impressive performance in various language tasks.However, they are prone to spurious correlations, and often generate illusory information.In real-world applications, PLMs should justify decisions with formalized, coherent reasoning chains, but this challenge remains under-explored.Cognitive psychology theorizes that humans are capable of utilizing fast and intuitive heuristic thinking to make decisions based on past experience, then rationalizing the decisions through slower and deliberative analytic reasoning.We incorporate these interlinked dual processes in fine-tuning and in-context learning with PLMs, applying them to two language understanding tasks that require coherent physical commonsense reasoning.We show that our proposed Heuristic-Analytic Reasoning (HAR) strategies drastically improve the coherence of rationalizations for model decisions, yielding state-of-the-art results on Tiered Reasoning for Intuitive Physics (TRIP).We also find that this improved coherence is a direct result of more faithful attention to relevant language context in each step of reasoning.Our findings suggest that human-like reasoning strategies can effectively improve the coherence and reliability of PLM reasoning.
Shane Storks, Fengyuan Hu, Sungryull Sohn, Moontae Lee, Honglak Lee, Joyce Y. Chai
EMNLP2
2022 DANLI: Deliberative Agent for Following Natural Language Instructions
abstract
Yichi Zhang, Jianing Yang, Jiayi Pan, Shane Storks, Nikhil Devraj, Ziqiao Ma, Keunwoo Yu, Yuwei Bao, Joyce Chai. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Yichi Zhang 0001, Jiayi Pan 0002, Shane Storks, Nikhil Devraj, Ziqiao Ma 0001, Keunwoo Peter Yu, Yuwei Bao, Joyce Y. Chai
EMNLP4