Stephanie M. Lukin

dblp:63/11299 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
7since 2021 · last 2025
0000-0001-8761-167XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 8 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 7 · 3 first-author · 4 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 From Camera-Eye to AI: Exploring the Interplay of Cinematography and Computational Visual Storytelling
Brett A. Halperin, Stephanie M. Lukin
CHI2
2025 Aha! - Predicting What Matters Next: Online Highlight Detection Without Looking Ahead
abstract
Real-time understanding of continuous video streams is essential for intelligent agents operating in high-stakes environments, including autonomous vehicles, surveillance drones, and disaster response robots. Yet, most existing video understanding and highlight detection methods assume access to the entire video during inference, making them unsuitable for online or streaming scenarios. In particular, current models optimize for offline summarization, failing to support step-by-step reasoning needed for real-time decision-making. We introduce Aha, an autoregressive highlight detection framework that predicts the relevance of each video frame against a task described in natural language. Without accessing future video frames, Aha utilizes a multimodal vision-language model and lightweight, decoupled heads trained on a large, curated dataset of human-centric video labels. To enable scalability, we introduce the Dynamic SinkCache mechanism that achieves constant memory usage across infinite-length streams without degrading performance on standard benchmarks. This encourages the hidden representation to capture high-level task objectives, enabling effective frame-level rankings for informativeness, relevance, and uncertainty with respect to the natural language task. Aha achieves state-of-the-art (SOTA) performance on highlight detection benchmarks, surpassing even prior offline, full-context approaches and video-language models by +5.9\% on TVSum and +8.3\% on Mr.Hisum in mAP (mean Average Precision). We explore Aha’s potential for real-world robotics applications given a task-oriented natural language input and a continuous, robot-centric video. Both experiments demonstrate Aha's potential effectiveness as a real-time reasoning module for downstream planning and long-horizon understanding.
Aiden Chang, Celso de Melo, Stephanie M. Lukin
NeurIPS3
2024 Artificial Dreams: Surreal Visual Storytelling as Inquiry Into AI 'Hallucination'
abstract
What does it mean for stochastic artificial intelligence (AI) to “hallucinate” when performing a literary task as open-ended as creative visual storytelling? In this paper, we investigate AI “hallucination” by stress-testing a visual storytelling algorithm with different visual and textual inputs designed to probe dream logic inspired by cinematic surrealism. Following a close reading of 100 visual stories that we deem artificial dreams, we describe how AI “hallucination” in computational visual storytelling is the opposite of groundedness: literary expression that is ungrounded in the visual or textual inputs. We find that this lack of grounding can be a source of either creativity or harm entangled with bias and illusion. In turn, we disentangle these obscurities and discuss steps toward addressing the perils while harnessing the potentials for innocuous cases of AI “hallucination” to enhance the creativity of visual storytelling.
Brett A. Halperin, Stephanie M. Lukin
Conference on Designing Interactive Systems2
2024 SCOUT: A Situated and Multi-Modal Human-Robot Dialogue Corpus
abstract
We introduce the Situated Corpus Of Understanding Transactions (SCOUT), a multi-modal collection of human-robot dialogue in the task domain of collaborative exploration. The corpus was constructed from multiple Wizard-of-Oz experiments where human participants gave verbal instructions to a remotely-located robot to move and gather information about its surroundings. SCOUT contains 89,056 utterances and 310,095 words from 278 dialogues averaging 320 utterances per dialogue. The dialogues are aligned with the multi-modal data streams available during the experiments: 5,785 images and 30 maps. The corpus has been annotated with Abstract Meaning Representation and Dialogue-AMR to identify the speaker’s intent and meaning within an utterance, and with Transactional Units and Relations to track relationships between utterances to reveal patterns of the Dialogue Structure. We describe how the corpus and its annotations have been used to develop autonomous human-robot systems and enable research in open questions of how humans speak to robots. We release this corpus to accelerate progress in autonomous, situated, human-robot dialogue, especially in the context of navigation tasks where details about the environment need to be discovered.
Stephanie M. Lukin, Claire Bonial, Matthew Marge, Taylor Hudson, Cory J. Hayes, Kimberly A. Pollard, Anthony Baker, Ashley Foots, Ron Artstein, Felix Gervits, Mitchell Abrams, Cassidy Henry, Lucia Donatelli, Anton Leuski, Susan G. Hill, David R. Traum, Clare R. Voss
LREC/COLING1
2023 Envisioning Narrative Intelligence: A Creative Visual Storytelling Anthology
abstract
In this paper, we collect an anthology of 100 visual stories from authors who participated in our systematic creative process of improvised story-building based on image sequences. Following close reading and thematic analysis of our anthology, we present five themes that characterize the variations found in this creative visual storytelling process: (1) Narrating What is in Vision vs. Envisioning; (2) Dynamically Characterizing Entities/Objects; (3) Sensing Experiential Information About the Scenery; (4) Modulating the Mood; (5) Encoding Narrative Biases. In understanding the varied ways that people derive stories from images, we offer considerations for collecting story-driven training data to inform automatic story generation. In correspondence with each theme, we envision narrative intelligence criteria for computational visual storytelling as: creative, reliable, expressive, grounded, and responsible. From these criteria, we discuss how to foreground creative expression, account for biases, and operate in the bounds of visual storyworlds.
Brett A. Halperin, Stephanie M. Lukin
CHI2
2023 Navigating to Success in Multi-Modal Human-Robot Collaboration: Analysis and Corpus Release
abstract
Human-guided robotic exploration is a useful approach to gathering information at remote locations, especially those that might be too risky, inhospitable, or inaccessible for humans. Maintaining common ground between the remotely-located partners is a challenge, one that can be facilitated by multi-modal communication. In this paper, we explore how participants utilized multiple modalities to investigate a remote location with the help of a robotic partner. Participants issued spoken natural language instructions and received from the robot: text-based feedback, continuous 2D LIDAR mapping, and upon-request static photographs. We noticed that different strategies were adopted in terms of use of the modalities, and hypothesize that these differences may be correlated with success at several exploration sub-tasks. We found that requesting photos may have improved the identification and counting of some key entities (doorways in particular) and that this strategy did not hinder the amount of overall area exploration. Future work with larger samples may reveal the effects of more nuanced photo and dialogue strategies, which can inform the training of robotic agents. Additionally, we announce the release of our unique multi-modal corpus of human-robot communication in an exploration context: SCOUT, the Situated Corpus on Understanding Transactions.
Stephanie M. Lukin, Kimberly A. Pollard, Claire Bonial, Taylor Hudson, Ron Artstein, Clare R. Voss, David R. Traum
RO-MAN1
2022 The Search for Agreement on Logical Fallacy Annotation of an Infodemic
abstract
We evaluate an annotation schema for labeling logical fallacy types, originally developed for a crowd-sourcing annotation paradigm, now using an annotation paradigm of two trained linguist annotators. We apply the schema to a variety of different genres of text relating to the COVID-19 pandemic. Our linguist (as opposed to crowd-sourced) annotation of logical fallacies allows us to evaluate whether the annotation schema category labels are sufficiently clear and non-overlapping for both manual and, later, system assignment. We report inter-annotator agreement results over two annotation phases as well as a preliminary assessment of the corpus for training and testing a machine learning algorithm (Pattern-Exploiting Training) for fallacy detection and recognition. The agreement results and system performance underscore the challenging nature of this annotation task and suggest that the annotation schema and paradigm must be iteratively evaluated and refined in order to arrive at a set of annotation labels that can be reproduced by human annotators and, in turn, provide reliable training data for automatic detection and recognition systems.
Claire Bonial, Austin Blodgett, Taylor Hudson, Stephanie M. Lukin, Jeffrey Micher, Douglas Summers-Stay, Peter Sutor Jr., Clare R. Voss
LREC4
2020 Dialogue-AMR: Abstract Meaning Representation for Dialogue
abstract
This paper describes a schema that enriches Abstract Meaning Representation (AMR) in order to provide a semantic representation for facilitating Natural Language Understanding (NLU) in dialogue systems. AMR offers a valuable level of abstraction of the propositional content of an utterance; however, it does not capture the illocutionary force or speaker’s intended contribution in the broader dialogue context (e.g., make a request or ask a question), nor does it capture tense or aspect. We explore dialogue in the domain of human-robot interaction, where a conversational robot is engaged in search and navigation tasks with a human partner. To address the limitations of standard AMR, we develop an inventory of speech acts suitable for our domain, and present “Dialogue-AMR”, an enhanced AMR that represents not only the content of an utterance, but the illocutionary force behind it, as well as tense and aspect. To showcase the coverage of the schema, we use both manual and automatic methods to construct the “DialAMR” corpus—a corpus of human-robot dialogue annotated with standard AMR and our enriched Dialogue-AMR schema. Our automated methods can be used to incorporate AMR into a larger NLU pipeline supporting human-robot dialogue.
Claire Bonial, Lucia Donatelli, Mitchell Abrams, Stephanie M. Lukin, Stephen Tratz, Matthew Marge, Ron Artstein, David R. Traum, Clare R. Voss
LREC4
2018 Dialogue Structure Annotation for Multi-Floor Interaction
David R. Traum, Cassidy Henry, Stephanie M. Lukin, Ron Artstein, Felix Gervits, Kimberly A. Pollard, Claire Bonial, Su Lei, Clare R. Voss, Matthew Marge, Cory J. Hayes, Susan G. Hill
LREC3
2018 Consequences and Factors of Stylistic Differences in Human-Robot Dialogue
abstract
This paper identifies stylistic differences in instruction-giving observed in a corpus of human-robot dialogue.Differences in verbosity and structure (i.e., single-intent vs. multi-intent instructions) arose naturally without restrictions or prior guidance on how users should speak with the robot.Different styles were found to produce different rates of miscommunication, and correlations were found between style differences and individual user variation, trust, and interaction experience with the robot.Understanding potential consequences and factors that influence style can inform design of dialogue systems that are robust to natural variation from human users.
Stephanie M. Lukin, Kimberly A. Pollard, Claire Bonial, Matthew Marge, Cassidy Henry, Ron Artstein, David R. Traum, Clare R. Voss
SIGDIAL Conference1
2018 Controlling Personality-Based Stylistic Variation with Neural Natural Language Generators
abstract
Natural language generators for taskoriented dialogue must effectively realize system dialogue actions and their associated semantics.In many applications, it is also desirable for generators to control the style of an utterance.To date, work on task-oriented neural generation has primarily focused on semantic fidelity rather than achieving stylistic goals, while work on style has been done in contexts where it is difficult to measure content preservation.Here we present three different sequence-to-sequence models and carefully test how well they disentangle content and style.We use a statistical generator, PERSONAGE, to synthesize a new corpus of over 88,000 restaurant domain utterances whose style varies according to models of personality, giving us total control over both the semantic content and the stylistic variation in the training data.We then vary the amount of explicit stylistic supervision given to the three models.We show that our most explicit model can simultaneously achieve high fidelity to both semantic and stylistic goals: this model adds a context vector of 36 stylistic parameters as input to the hidden state of the encoder at each time step, showing the benefits of explicit stylistic supervision, even when the amount of training data is large.
Shereen Oraby, Lena Reed, Shubhangi Tandon, Sharath T. S., Stephanie M. Lukin, Marilyn A. Walker
SIGDIAL Conference5
2017 Argument Strength is in the Eye of the Beholder: Audience Effects in Persuasion
abstract
Stephanie Lukin, Pranav Anand, Marilyn Walker, Steve Whittaker. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017.
Stephanie M. Lukin, Pranav Anand, Marilyn A. Walker, Steve Whittaker 0001
EACL (1)1
2016 PersonaBank: A Corpus of Personal Narratives and Their Story Intention Graphs
Stephanie M. Lukin, Kevin Bowden, Casey Barackman, Marilyn A. Walker
LREC1
2015 Narrative Variations in a Virtual Storyteller
Stephanie M. Lukin, Marilyn A. Walker
IVA1
2015 Generating Sentence Planning Variations for Story Telling
abstract
There has been a recent explosion in applications for dialogue interaction ranging from direction-giving and tourist information to interactive story systems. Yet the natural language generation (NLG) component for many of these systems remains largely handcrafted. This limitation greatly restricts the range of applications; it also means that it is impossible to take advantage of recent work in expressive and statistical language generation that can dynamically and automatically produce a large number of variations of given content. We propose that a solution to this problem lies in new methods for developing language generation resources. We describe the ES-TRANSLATOR, a computational language generator that has previously been applied only to fables, and quantitatively evaluate the domain independence of the EST by applying it to personal narratives from weblogs. We then take advantage of recent work on language generation to create a parameterized sentence planner for story generation that provides aggregation operations, variations in discourse and in point of view. Finally, we present a user evaluation of different personal narrative retellings.
Stephanie M. Lukin, Lena Reed, Marilyn A. Walker
SIGDIAL Conference1
2014 Building Community and Commitment with a Virtual Coach in Mobile Wellness Programs
Stephanie M. Lukin, G. Michael Youngblood, Honglu Du, Marilyn A. Walker
IVA1
2014 Getting Reliable Annotations for Sarcasm in Online Dialogues
Reid Swanson, Stephanie M. Lukin, Luke Eisenberg, Thomas Chase Corcoran, Marilyn A. Walker
LREC2
2014 Extracting relevant knowledge for the detection of sarcasm and nastiness in the social web
Raquel Justo, Thomas Chase Corcoran, Stephanie M. Lukin, Marilyn A. Walker, M. Inés Torres
Knowl. Based Syst.3
2013 Generating Different Story Tellings from Semantic Representations of Narrative
Elena Rishes, Stephanie M. Lukin, David K. Elson, Marilyn A. Walker
ICIDS2
2011 A Machine Learning Approach to End-to-End RTT Estimation and its Application to TCP
abstract
In this paper, we explore a novel approach to end-to-end round-trip time (RTT) estimation using a machine-learning technique known as the Experts Framework. In our proposal, each of several "experts" guesses a fixed value. The weighted average of these guesses estimates the RTT, with the weights updated after every RTT measurement based on the difference between the estimated and actual RTT. Through extensive simulations we show that the proposed machine-learning algorithm adapts very quickly to changes in the RTT. Our results show a considerable reduction in the number of retransmitted packets and a increase in goodput, in particular on more heavily congested scenarios. We corroborate our results through "live" experiments using an implementation of the proposed algorithm in the Linux kernel. These experiments confirm the higher accuracy of the machine learning approach with more than 40% improvement, not only over the standard TCP, but also over the well known Eifel RTT estimator.
Bruno Astuto A. Nunes, Kerry Veenstra, William Ballenthin, Stephanie M. Lukin, Katia Obraczka
ICCCN4