VLDB 2026 Research / reviewers in the wild / expert
Amber Shinsel
dblp:45/8771
· DBLP profile ↗
4ranked-venue papers
1as first author
0since 2021 · last 2014
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
1 paper |
Software testing · 100% |
Topics — the 2 heaviest of 2, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Software testing › regression testing
test selection |
0.2 | 1 | 2014 | You Are the Only Possible Oracle: Effective Test Selection for End Users of Interactive Machine Learning Systems · IEEE Trans. Software Eng. 2014 |
Software testing
test oracle |
0.1 | 1 | 2014 | You Are the Only Possible Oracle: Effective Test Selection for End Users of Interactive Machine Learning Systems · IEEE Trans. Software Eng. 2014 |
Methods — techniques the papers use, named apart from their topics
user study · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | You Are the Only Possible Oracle: Effective Test Selection for End Users of Interactive Machine Learning SystemsabstractHow do you test a program when only a single user, with no expertise in software testing, is able to determine if the program is performing correctly? Such programs are common today in the form of machine-learned classifiers. We consider the problem of testing this common kind of machine-generated program when the only oracle is an end user: e.g., only you can determine if your email is properly filed. We present test selection methods that provide very good failure rates even for small test suites, and show that these methods work in both large-scale random experiments using a “gold standard” and in studies with real users. Our methods are inexpensive and largely algorithm-independent. Key to our methods is an exploitation of properties of classifiers that is not possible in traditional software testing. Our results suggest that it is plausible for time-pressured end users to interactively detect failures-even very hard-to-find failures-without wading through a large number of successful (and thus less useful) tests. We additionally show that some methods are able to find the arguably most difficult-to-detect faults of classifiers: cases where machine learning algorithms have high confidence in an incorrect result. Alex Groce, Todd Kulesza, Chaoqiang Zhang, Shalini Shamasunder, Margaret M. Burnett, Weng-Keen Wong, Simone Stumpf, Shubhomoy Das, Amber Shinsel, Forrest Bice, Kevin McIntosh |
IEEE Trans. Software Eng. | 9 |
| 2011 | Mini-crowdsourcing end-user assessment of intelligent assistants: A cost-benefit studyabstractIntelligent assistants sometimes handle tasks too important to be trusted implicitly. End users can establish trust via systematic assessment, but such assessment is costly. This paper investigates whether, when, and how bringing a small crowd of end users to bear on the assessment of an intelligent assistant is useful from a cost/benefit perspective. Our results show that a mini-crowd of testers supplied many more benefits than the obvious decrease in workload, but these benefits did not scale linearly as mini-crowd size increased - there was a point of diminishing returns where the cost-benefit ratio became less attractive. Amber Shinsel, Todd Kulesza, Margaret M. Burnett, William Curran, Alex Groce, Simone Stumpf, Weng-Keen Wong |
VL/HCC | 1 |
| 2010 | Does My Model Work? Evaluation Abstractions of Cognitive ModelersabstractAre the abstractions that scientific modelers use to build their models in a modeling language the same abstractions they use to evaluate the correctness of their models? The extent to which such differences exist seems likely to correspond to additional effort of modelers in determining whether their models work as intended. In this paper, we therefore investigate the distinction between "programming abstractions" and "evaluation abstractions". As the basis of our investigation, we conducted a case study on cognitive modeling. We report modelers' evaluation abstractions, and the lengths they went to in evaluating their models. From these results, we derive design implications for several categories of persistent, first-class evaluation abstractions in future debugging tools for modelers. Christopher Bogart, Margaret M. Burnett, Scott Douglass, David Piorkowski, Amber Shinsel |
VL/HCC | 5 |
| 2010 | Explanatory Debugging: Supporting End-User Debugging of Machine-Learned ProgramsabstractMany machine-learning algorithms learn rules of behavior from individual end users, such as task-oriented desktop organizers and handwriting recognizers. These rules form a “program” that tells the computer what to do when future inputs arrive. Little research has explored how an end user can debug these programs when they make mistakes. We present our progress toward enabling end users to debug these learned programs via a Natural Programming methodology. We began with a formative study exploring how users reason about and correct a text-classification program. From the results, we derived and prototyped a concept based on “explanatory debugging”, then empirically evaluated it. Our results contribute methods for exposing a learned program's logic to end users and for eliciting user corrections to improve the program's predictions. Todd Kulesza, Simone Stumpf, Margaret M. Burnett, Weng-Keen Wong, Yann Riche, Travis Moore, Ian Oberst, Amber Shinsel, Kevin McIntosh |
VL/HCC | 8 |