EDBT 2026 Demo / reviewers in the wild / expert
Ethan Fast
dblp:39/8285
· DBLP profile ↗
13ranked-venue papers
11as first author
0since 2021 · last 2020
0009-0006-9093-1299ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 8 · 7 first-authorArtificial intelligence and machine learning · 5 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Information extraction and text analysis · 49% Reinforcement learning · 18% Question answering and dialogue systems · 14% | |
| Human-computer interaction and pervasive computing
3 papers |
Learning and educational technologies · 74% Ubiquitous computing and smart environments · 26% | |
| Software engineering, system software, and programming languages
3 papers |
Programming languages and type systems · 47% Software maintenance and evolution · 32% Empirical software engineering · 20% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Computational social science and digital humanities · 100% | |
| Theoretical computer science
1 paper |
Automated reasoning and model checking · 100% |
Topics — the 16 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems
conversational agents |
0.3 | 1 | 2018 | Iris: A Conversational Agent for Complex Tasks · CHI 2018 |
Natural language and speech › Information extraction and text analysis › lexical resources › lexical resource construction
lexicon induction |
0.3 | 1 | 2017 | Lexicons on Demand: Neural Word Embeddings for Large-Scale Text Analysis · IJCAI 2017 |
Natural language and speech › Information extraction and text analysis
sentiment analysis |
0.3 | 1 | 2017 | Long-Term Trends in the Public Perception of Artificial Intelligence · AAAI 2017 |
Machine learning › Representation and self-supervised learning › word representation
word embedding |
0.3 | 1 | 2017 | Lexicons on Demand: Neural Word Embeddings for Large-Scale Text Analysis · IJCAI 2017 |
Ubiquitous computing and smart environments › context recognition
activity recognition |
0.2 | 1 | 2016 | Augur: Mining Human Behaviors from Fiction to Power Interactive Systems · CHI 2016 |
Programming languages and type systems › language design
language extension |
0.2 | 1 | 2016 | Meta: Enabling Programming Languages to Learn from the Crowd · UIST 2016 |
Empirical software engineering
mining software repositories |
0.2 | 1 | 2014 | Emergent, crowd-scale programming practice in the IDE · CHI 2014 |
Learning and educational technologies › educational assessment
automated grading |
0.2 | 1 | 2013 | Crowd-scale interactive formal reasoning and analytics · UIST 2013 |
Automated reasoning and model checking › theorem proving
proof checking |
0.2 | 1 | 2013 | Crowd-scale interactive formal reasoning and analytics · UIST 2013 |
Learning and educational technologies
online learning |
0.1 | 1 | 2020 | Reinforcement Learning for the Adaptive Scheduling of Educational Activities · CHI 2020 |
Programming languages and type systems
domain-specific languages |
0.1 | 1 | 2018 | Iris: A Conversational Agent for Complex Tasks · CHI 2018 |
Programming languages and type systems
type systems |
0.1 | 1 | 2018 | Iris: A Conversational Agent for Complex Tasks · CHI 2018 |
Natural language and speech › Information extraction and text analysis
topic model |
0.1 | 1 | 2017 | Long-Term Trends in the Public Perception of Artificial Intelligence · AAAI 2017 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › commonsense reasoning
commonsense knowledge base |
0.1 | 1 | 2016 | Augur: Mining Human Behaviors from Fiction to Power Interactive Systems · CHI 2016 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › ontology
lexical knowledge base |
0.1 | 1 | 2016 | Empath: Understanding Topic Signals in Large-Scale Text · CHI 2016 |
Computational social science and digital humanities
social media analysis |
0.1 | 1 | 2016 | Identifying Dogmatism in Social Media: Signals and Models · EMNLP 2016 |
Methods — techniques the papers use, named apart from their topics
crowdsourcing · 1.8reinforcement learning · 0.9controlled experiment · 0.9active learning · 0.9user study · 0.7text corpus analysis · 0.6sentiment analysis · 0.6statistical modeling · 0.5feature engineering · 0.5proof cache · 0.3neural word embeddings · 0.3vector model · 0.2text mining · 0.2runtime instrumentation · 0.2statistical analysis · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Reinforcement Learning for the Adaptive Scheduling of Educational ActivitiesabstractAdaptive instruction for online education can increase learning gains and decrease the work required of learners, instructors, and course designers. Reinforcement Learning (RL) is a promising tool for developing instructional policies, as RL models can learn complex relationships between course activities, learner actions, and educational outcomes. This paper demonstrates the first RL model to schedule educational activities in real time for a large online course through active learning. Our model learns to assign a sequence of course activities while maximizing learning gains and minimizing the number of items assigned. Using a controlled experiment with over 1,000 learners, we investigate how this scheduling policy affects learning gains, dropout rates, and qualitative learner feedback. We show that our model produces better learning gains using fewer educational activities than a linear assignment condition, and produces similar learning gains to a self-directed condition using fewer educational activities and with lower dropout rates. Jonathan Bassen, Bharathan Balaji, Michael Schaarschmidt, Candace Thille, Jay Painter, Dawn Zimmaro, Alex Games, Ethan Fast, John C. Mitchell |
CHI | 8 |
| 2018 | Iris: A Conversational Agent for Complex TasksabstractToday, most conversational agents are limited to simple tasks supported by standalone commands, such as getting directions or scheduling an appointment. To support more complex tasks, agents must be able to generalize from and combine the commands they already understand. This paper presents a new approach to designing conversational agents inspired by linguistic theory, where agents can execute complex requests interactively by combining commands through nested conversations. We demonstrate this approach in Iris, an agent that can perform open-ended data science tasks such as lexical analysis and predictive modeling. To power Iris, we have created a domain-specific language that transforms Python functions into combinable automata and regulates their combinations through a type system. Running a user study to examine the strengths and limitations of our approach, we find that data scientists completed a modeling task 2.6 times faster with Iris than with Jupyter Notebook. Ethan Fast, Binbin Chen 0002, Julia Mendelsohn, Jonathan Bassen, Michael S. Bernstein |
CHI | 1 |
| 2018 | OARS: exploring instructor analytics for online learningabstractLearning analytics systems have the potential to bring enormous value to online education. Unfortunately, many instructors and platforms do not adequately leverage learning analytics in their courses today. In this paper, we report on the value of these systems from the perspective of course instructors. We study these ideas through OARS, a modular and real-time learning analytics system that we deployed across more than ten online courses with tens of thousands of learners. We leverage this system as a starting point for semi-structured interviews with a diverse set of instructors. Our study suggests new design goals for learning analytics systems, the importance of real-time analytics to many instructors, and the value of flexibility in data selection and aggregation for an instructor when working with an analytics system. Jonathan Bassen, Iris Howley, Ethan Fast, John C. Mitchell, Candace Thille |
L@S | 3 |
| 2017 | Long-Term Trends in the Public Perception of Artificial IntelligenceabstractAnalyses of text corpora over time can reveal trends in beliefs, interest, and sentiment about a topic. We focus on views expressed about artificial intelligence (AI) in the New York Times over a 30-year period. General interest, awareness, and discussion about AI has waxed and waned since the field was founded in 1956. We present a set of measures that captures levels of engagement, measures of pessimism and optimism, the prevalence of specific hopes and concerns, and topics that are linked to discussions about AI over decades. We find that discussion of AI has increased sharply since 2009, and that these discussions have been consistently more optimistic than pessimistic. However, when we examine specific concerns, we find that worries of loss of control of AI, ethical concerns for AI, and the negative impact of AI on work have grown in recent years. We also find that hopes for AI in healthcare and education have increased over time. Ethan Fast, Eric Horvitz |
AAAI | 1 |
| 2017 | Lexicons on Demand: Neural Word Embeddings for Large-Scale Text AnalysisabstractHuman language is colored by a broad range of topics, but existing text analysis tools only focus on a small number of them. We present Empath, a tool that can generate and validate new lexical categories on demand from a small set of seed terms (like "bleed" and "punch" to generate the category violence). Empath draws connotations between words and phrases by learning a neural embedding across billions of words on the web. Given a small set of seed words that characterize a category, Empath uses its neural embedding to discover new related terms, then validates the category with a crowd-powered filter. Empath also analyzes text across 200 built-in, pre-validated categories we have generated such as neglect, government, and social media. We show that Empath's data-driven, human validated categories are highly correlated (r=0.906) with similar categories in LIWC. Ethan Fast, Binbin Chen 0002, Michael S. Bernstein |
IJCAI | 1 |
| 2016 | Empath: Understanding Topic Signals in Large-Scale TextabstractHuman language is colored by a broad range of topics, but existing text analysis tools only focus on a small number of them. We present Empath, a tool that can generate and validate new lexical categories on demand from a small set of seed terms (like "bleed" and "punch" to generate the category violence). Empath draws connotations between words and phrases by deep learning a neural embedding across more than 1.8 billion words of modern fiction. Given a small set of seed words that characterize a category, Empath uses its neural embedding to discover new related terms, then validates the category with a crowd-powered filter. Empath also analyzes text across 200 built-in, pre-validated categories we have generated from common topics in our web dataset, like neglect, government, and social media. We show that Empath's data-driven, human validated categories are highly correlated (r=0.906) with similar categories in LIWC. Ethan Fast, Binbin Chen 0002, Michael S. Bernstein |
CHI | 1 |
| 2016 | Augur: Mining Human Behaviors from Fiction to Power Interactive SystemsabstractFrom smart homes that prepare coffee when we wake, to phones that know not to interrupt us during important conversations, our collective visions of HCI imagine a future in which computers understand a broad range of human behaviors. Today our systems fall short of these visions, however, because this range of behaviors is too large for designers or programmers to capture manually. In this paper, we instead demonstrate it is possible to mine a broad knowledge base of human behavior by analyzing more than one billion words of modern fiction. Our resulting knowledge base, Augur, trains vector models that can predict many thousands of user activities from surrounding objects in modern contexts: for example, whether a user may be eating food, meeting with a friend, or taking a selfie. Augur uses these predictions to identify actions that people commonly take on objects in the world and estimate a user's future activities given their current situation. We demonstrate Augur-powered, activity-based systems such as a phone that silences itself when the odds of you answering it are low, and a dynamic music player that adjusts to your present activity. A field deployment of an Augur-powered wearable camera resulted in 96% recall and 71% precision on its unsupervised predictions of common daily activities. A second evaluation where human judges rated the system's predictions over a broad set of input images found that 94% were rated sensible. Ethan Fast, William McGrath, Pranav Rajpurkar, Michael S. Bernstein |
CHI | 1 |
| 2016 | Identifying Dogmatism in Social Media: Signals and ModelsabstractWe explore linguistic and behavioral features of dogmatism in social media and construct statistical models that can identify dogmatic comments.Our model is based on a corpus of Reddit posts, collected across a diverse set of conversational topics and annotated via paid crowdsourcing.We operationalize key aspects of dogmatism described by existing psychology theories (such as over-confidence), finding they have predictive power.We also find evidence for new signals of dogmatism, such as the tendency of dogmatic posts to refrain from signaling cognitive processes.When we use our predictive model to analyze millions of other Reddit posts, we find evidence that suggests dogmatism is a deeper personality trait, present for dogmatic users across many different domains, and that users who engage on dogmatic comments tend to show increases in dogmatic posts themselves. Ethan Fast, Eric Horvitz |
EMNLP | 1 |
| 2016 | Shirtless and Dangerous: Quantifying Linguistic Signals of Gender Bias in an Online Fiction Writing Community
Ethan Fast, Tina Vachovsky, Michael S. Bernstein |
ICWSM | 1 |
| 2016 | Meta: Enabling Programming Languages to Learn from the CrowdabstractCollectively authored programming resources such as Q&A sites and open-source libraries provide a limited window into how programs are constructed, debugged, and run. To address these limitations, we introduce Meta: a language extension for Python that allows programmers to share functions and track how they are used by a crowd of other programmers. Meta functions are shareable via URL and instrumented to record runtime data. Combining thousands of Meta functions with their collective runtime data, we demonstrate tools including an optimizer that replaces your function with a more efficient version written by someone else, an auto-patcher that saves your program from crashing by finding equivalent functions in the community, and a proactive linter that warns you when a function fails elsewhere in the community. We find that professional programmers are able to use Meta for complex tasks (creating new Meta functions that, for example, cross-validate a logistic regression), and that Meta is able to find 44 optimizations (for a 1.45 times average speedup) and 5 bug fixes across the crowd. Ethan Fast, Michael S. Bernstein |
UIST | 1 |
| 2014 | Emergent, crowd-scale programming practice in the IDEabstractWhile emergent behaviors are uncodified across many domains such as programming and writing, interfaces need explicit rules to support users. We hypothesize that by codifying emergent programming behavior, software engineering interfaces can support a far broader set of developer needs. To explore this idea, we built Codex, a knowledge base that records common practice for the Ruby programming language by indexing over three million lines of popular code. Codex enables new data-driven interfaces for programming systems: statistical linting, identifying code that is unlikely to occur in practice and may constitute a bug; pattern annotation, automatically discovering common programming idioms and annotating them with metadata using expert crowdsourcing; and library generation, constructing a utility package that encapsulates and reflects emergent software practice. We evaluate these applications to find Codex captures a broad swatch of programming practice, statistical linting detects problematic code snippets, and pattern annotation discovers nontrivial idioms such as basic HTTP authentication and database migration templates. Our work suggests that operationalizing practice-driven knowledge in structured domains such as programming can enable a new class of user interfaces. Ethan Fast, Daniel Steffee, Lucy Wang, Joel Brandt, Michael S. Bernstein |
CHI | 1 |
| 2013 | Crowd-scale interactive formal reasoning and analyticsabstractLarge online courses often assign problems that are easy to grade because they have a fixed set of solutions (such as multiple choice), but grading and guiding students is more difficult in problem domains that have an unbounded number of correct answers. One such domain is derivations: sequences of logical steps commonly used in assignments for technical, mathematical and scientific subjects. We present DeduceIt, a system for creating, grading, and analyzing derivation assignments in any formal domain. DeduceIt supports assignments in any logical formalism, provides students with incremental feedback, and aggregates student paths through each proof to produce instructor analytics. DeduceIt benefits from checking thousands of derivations on the web: it introduces a proof cache, a novel data structure which leverages a crowd of students to decrease the cost of checking derivations and providing real-time, constructive feedback. We evaluate DeduceIt with 990 students in an online compilers course, finding students take advantage of its incremental feedback and instructors benefit from its structured insights into course topics. Our work suggests that automated reasoning can extend online assignments and large-scale education to many new domains. Ethan Fast, Colleen Lee, Alex Aiken, Michael S. Bernstein, Daphne Koller |
UIST | 1 |
| 2010 | Designing better fitness functions for automated program repairabstractEvolutionary methods have been used to repair programs automatically, with promising results. However, the fitness function used to achieve these results was based on a few simple test cases and is likely too simplistic for larger programs and more complex bugs. We focus here on two aspects of fitness evaluation: efficiency and precision. Efficiency is an issue because many programs have hundreds of test cases, and it is costly to run each test on every individual in the population. Moreover, the precision of fitness functions based on test cases is limited by the fact that a program either passes a test case, or does not, which leads to a fitness function that can take on only a few distinct values. This paper investigates two approaches to enhancing fitness functions for program repair, incorporating (1) test suite selection to improve efficiency and (2) formal specifications to improve precision. We evaluate test suite selection on 10 programs, improving running time for automated repair by 81%. We evaluate program invariants using the Fitness Distance Correlation (FDC) metric, demonstrating significant improvements and smoother evolution of repairs Ethan Fast, Claire Le Goues, Stephanie Forrest, Westley Weimer |
GECCO | 1 |