Ivan Stelmakh

dblp:222/3200 · DBLP profile ↗
← Back
11ranked-venue papers
8as first author
9since 2021 · last 2024
0000-0002-9237-7379ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 7 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Question answering and dialogue systems · 36% Trustworthy machine learning · 24% Language models and text generation · 18%
Interdisciplinary, comprehensive, and emerging computing
5 papers
Computational social science and digital humanities · 100%
Theoretical computer science
4 papers
Algorithmic game theory and mechanism design · 57% Mathematical optimization · 33% Graph algorithms and graph theory · 10%
Human-computer interaction and pervasive computing
3 papers
Collaborative and social computing · 50% Human-AI interaction · 38% Usability and user experience research · 12%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Software engineering, system software, and programming languages
1 paper
Software testing · 50% Program analysis · 50%

Topics — the 17 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computational social science and digital humanities › science of science
peer review
1.032021
A Novice-Reviewer Experiment to Address Scarcity of Qualified Reviewers in Large Conferences · AAAI 2021
On Testing for Biases in Peer Review · NeurIPS 2019
Catch Me if I Can: Detecting Strategic Behaviour in Peer Assessment · AAAI 2021
Machine learning › Trustworthy machine learning
fairness
0.812024
Debiasing Evaluations That Are Biased by Evaluations · J. Mach. Learn. Res. 2024
Natural language and speech › Question answering and dialogue systems › robust question answering
ambiguous question answering
0.612022
ASQA: Factoid Questions Meet Long-Form Answers · EMNLP 2022
Natural language and speech › Language models and text generation › natural language understanding › question answering
factoid question answering
0.612022
ASQA: Factoid Questions Meet Long-Form Answers · EMNLP 2022
Natural language and speech › Question answering and dialogue systems
long-form question answering
0.612022
ASQA: Factoid Questions Meet Long-Form Answers · EMNLP 2022
Information retrieval
evaluation
0.612022
ASQA: Factoid Questions Meet Long-Form Answers · EMNLP 2022
Information retrieval › evaluation › task-based evaluation
question answering evaluation
0.612022
ASQA: Factoid Questions Meet Long-Form Answers · EMNLP 2022
Human-AI interaction
human-AI collaboration
0.512021
Towards Fair, Equitable, and Efficient Peer Review · AAAI 2021
Collaborative and social computing › information sharing › scholarly communication
peer review
0.512021
Towards Fair, Equitable, and Efficient Peer Review · AAAI 2021
Mathematical optimization › regularization
regularized optimization
0.512021
Debiasing Evaluations That Are Biased by Evaluations · AAAI 2021
Program analysis
false alarm reduction
0.412019
On Testing for Biases in Peer Review · NeurIPS 2019
Software testing
statistical testing
0.412019
On Testing for Biases in Peer Review · NeurIPS 2019
Algorithmic game theory and mechanism design › market design › matching markets
reviewer assignment
0.412019
On Testing for Biases in Peer Review · NeurIPS 2019
Machine learning › Learning theory › model selection
cross-validation
0.212024
Debiasing Evaluations That Are Biased by Evaluations · J. Mach. Learn. Res. 2024
Computational social science and digital humanities
science of science
0.112021
Towards Fair, Equitable, and Efficient Peer Review · AAAI 2021
Collaborative and social computing
human computation
0.112021
A Novice-Reviewer Experiment to Address Scarcity of Qualified Reviewers in Large Conferences · AAAI 2021
Graph algorithms and graph theory › graph algorithms › network flow
maximum flow
0.112021
PeerReview4All: Fair and Accurate Reviewer Assignment in Peer Review · J. Mach. Learn. Res. 2021

Methods — techniques the papers use, named apart from their topics

regularized optimization · 2.3cross-validation · 2.3summarization · 1.1automated metric · 1.1hypothesis testing · 1.1subjective-score model · 1.0statistical hypothesis testing · 1.0minimax analysis · 1.0incremental max-flow · 1.0
YearPublicationVenuePosition
2024 Debiasing Evaluations That Are Biased by Evaluations
abstract
It is common to evaluate a set of items by soliciting people to rate them. For example, universities ask students to rate the teaching quality of their instructors, and conference organizers ask authors of submissions to evaluate the quality of the reviews. However, in these applications, students often give a higher rating to a course if they receive higher grades in a course, and authors often give a higher rating to the reviews if their papers are accepted to the conference. In this work, we call these external factors the "outcome" experienced by people, and consider the problem of mitigating these outcome-induced biases in the given ratings when some information about the outcome is available. We formulate the information about the outcome as a known partial ordering on the bias. We propose a debiasing method by solving a regularized optimization problem under this ordering constraint, and also provide a carefully designed cross-validation method that adaptively chooses the appropriate amount of regularization. We provide theoretical guarantees on the performance of our algorithm, as well as experimental evaluations.
Jingyan Wang 0001, Ivan Stelmakh, Yuting Wei 0001, Nihar B. Shah
J. Mach. Learn. Res.2
2022 ASQA: Factoid Questions Meet Long-Form Answers
abstract
An abundance of datasets and availability of reliable evaluation metrics have resulted in strong progress in factoid question answering (QA).This progress, however, does not easily transfer to the task of long-form QA, where the goal is to answer questions that require in-depth explanations.The hurdles include (i) a lack of high-quality data, and (ii) the absence of a well-defined notion of the answer's quality.In this work, we address these problems by (i) releasing a novel dataset and a task that we call ASQA (Answer Summaries for Questions which are Ambiguous); and (ii) proposing a reliable metric for measuring performance on ASQA.Our task focuses on factoid questions that are ambiguous, that is, have different correct answers depending on interpretation.Answers to ambiguous questions should synthesize factual information from multiple sources into a long-form summary that resolves the ambiguity.In contrast to existing long-form QA tasks (such as ELI5), ASQA admits a clear notion of correctness: a user faced with a good summary should be able to answer different interpretations of the original ambiguous question.We use this notion of correctness to define an automated metric of performance for ASQA.Our analysis demonstrates an agreement between this metric and human judgments, and reveals a considerable gap between human performance and strong baselines.
Ivan Stelmakh, Yi Luan, Bhuwan Dhingra, Ming-Wei Chang
EMNLP1
2022 No Rose for MLE: Inadmissibility of MLE for Evaluation Aggregation Under Levels of Expertise
abstract
A number of applications including crowd-sourced labeling and peer review require aggregation of labels or evaluations sourced from multiple evaluators. There is often additional information available pertaining to the evaluators’ expertise. A natural approach for aggregation is to consider the widely studied Dawid-Skene model (or its extensions incorporating evaluators’ expertise), and employ the standard maximum likelihood estimator (MLE). While MLE is in general widely used in practice and enjoys a number of appealing theoretical guarantees, in this work we provide a negative result for the MLE. Specifically, we prove that the MLE is asymptotically inadmissible for a special case of evaluation aggregation with expertise level information. We show this by constructing an alternative estimator that we show is significantly better than the MLE in certain parameter regimes and at least as good elsewhere. Finally, simulations reveal that our findings may hold in more general conditions than what we theoretically analyze.
Charvi Rastogi, Ivan Stelmakh, Nihar B. Shah, Sivaraman Balakrishnan
ISIT2
2021 Debiasing Evaluations That Are Biased by Evaluations
abstract
It is common to evaluate a set of items by soliciting people to rate them. For example, universities ask students to rate the teaching quality of their instructors, and conference organizers ask authors of submissions to evaluate the quality of the reviews. However, in these applications, students often give a higher rating to a course if they receive higher grades in a course, and authors often give a higher rating to the reviews if their papers are accepted to the conference. In this work, we call these external factors the "outcome" experienced by people, and consider the problem of mitigating these outcome-induced biases in the given ratings when some information about the outcome is available. We formulate the information about the outcome as a known partial ordering on the bias. We propose a debiasing method by solving a regularized optimization problem under this ordering constraint, and also provide a carefully designed cross-validation method that adaptively chooses the appropriate amount of regularization. We provide theoretical guarantees on the performance of our algorithm, as well as experimental evaluations.
Jingyan Wang 0001, Ivan Stelmakh, Yuting Wei 0001, Nihar B. Shah
AAAI2
2021 Towards Fair, Equitable, and Efficient Peer Review
abstract
Peer review is the backbone of academia. The rapid growth of the number of submissions to leading publication venues has identified a need for automation of some parts of the peer-review pipeline and nowadays human referees are required to interact with various interfaces and technologies in this process. However, there exists evidence that if such interactions are not carefully designed, they can exacerbate various problems related to fairness and efficiency of the process. In my research, I aim to design a Human-AI collaboration pipeline in peer review to mitigate these issues and ensure that science progresses in a fair, equitable, and efficient manner.
Ivan Stelmakh
AAAI1
2021 Catch Me if I Can: Detecting Strategic Behaviour in Peer Assessment
abstract
We consider the issue of strategic behaviour in various peer-assessment tasks, including peer grading of exams or homeworks and peer review in hiring or promotions. When a peer-assessment task is competitive (e.g., when students are graded on a curve), agents may be incentivized to misreport evaluations in order to improve their own final standing. Our focus is on designing methods for detection of such manipulations. Specifically, we consider a setting in which agents evaluate a subset of their peers and output rankings that are later aggregated to form a final ordering. In this paper, we investigate a statistical framework for this problem and design a principled test for detecting strategic behaviour. We prove that our test has strong false alarm guarantees and evaluate its detection ability in practical settings. For this, we design and conduct an experiment that elicits strategic behaviour from subjects and release a dataset of patterns of strategic behaviour that may be of independent interest. We use this data to run a series of real and semi-synthetic evaluations that reveal a strong detection power of our test.
Ivan Stelmakh, Nihar B. Shah, Aarti Singh
AAAI1
2021 A Novice-Reviewer Experiment to Address Scarcity of Qualified Reviewers in Large Conferences
abstract
Conference peer review constitutes a human-computation process whose importance cannot be overstated: not only it identifies the best submissions for acceptance, but, ultimately, it impacts the future of the whole research area by promoting some ideas and restraining others. A surge in the number of submissions received by leading AI conferences has challenged the sustainability of the review process by increasing the burden on the pool of qualified reviewers which is growing at a much slower rate. In this work, we consider the problem of reviewer recruiting with a focus on the scarcity of qualified reviewers in large conferences. Specifically, we design a procedure for (i) recruiting reviewers from the population not typically covered by major conferences and (ii) guiding them through the reviewing pipeline. In conjunction with the ICML 2020 --- a large, top-tier machine learning conference --- we recruit a small set of reviewers through our procedure and compare their performance with the general population of ICML reviewers. Our experiment reveals that a combination of the recruiting and guiding mechanisms allows for a principled enhancement of the reviewer pool and results in reviews of superior quality compared to the conventional pool of reviews as evaluated by senior members of the program committee (meta-reviewers).
Ivan Stelmakh, Nihar B. Shah, Aarti Singh, Hal Daumé III
AAAI1
2021 PeerReview4All: Fair and Accurate Reviewer Assignment in Peer Review
abstract
We consider the problem of automated assignment of papers to reviewers in conference peer review, with a focus on fairness and statistical accuracy. Our fairness objective is to maximize the review quality of the most disadvantaged paper, in contrast to the commonly used objective of maximizing the total quality over all papers. We design an assignment algorithm based on an incremental max-flow procedure that we prove is near-optimally fair. Our statistical accuracy objective is to ensure correct recovery of the papers that should be accepted. We provide a sharp minimax analysis of the accuracy of the peer-review process for a popular objective-score model as well as for a novel subjective-score model that we propose in the paper. Our analysis proves that our proposed assignment algorithm also leads to a near-optimal statistical accuracy. Finally, we design a novel experiment that allows for an objective comparison of various assignment algorithms, and overcomes the inherent difficulty posed by the absence of a ground truth in experiments on peer-review. The results of this experiment as well as of other experiments on synthetic and real data corroborate the theoretical guarantees of our algorithm.
Ivan Stelmakh, Nihar B. Shah, Aarti Singh
J. Mach. Learn. Res.1
2021 Prior and Prejudice: The Novice Reviewers' Bias against Resubmissions in Conference Peer Review
abstract
Modern machine learning and computer science conferences are experiencing a surge in the number of submissions that challenges the quality of peer review as the number of competent reviewers is growing at a much slower rate. To curb this trend and reduce the burden on reviewers, several conferences have started encouraging or even requiring authors to declare the previous submission history of their papers. Such initiatives have been met with skepticism among authors, who raise the concern about a potential bias in reviewers' recommendations induced by this information. In this work, we investigate whether reviewers exhibit a bias caused by the knowledge that the submission under review was previously rejected at a similar venue, focusing on a population of novice reviewers who constitute a large fraction of the reviewer pool in leading machine learning and computer science conferences. We design and conduct a randomized controlled trial closely replicating the relevant components of the peer-review pipeline with $133$ reviewers (master's, junior PhD students, and recent graduates of top US universities) writing reviews for $19$ papers. The analysis reveals that reviewers indeed become negatively biased when they receive a signal about paper being a resubmission, giving almost 1 point lower overall score on a 10-point Likert item (Δ = -0.78, 95% CI = [-1.30, -0.24]) than reviewers who do not receive such a signal. Looking at specific criteria scores (originality, quality, clarity and significance), we observe that novice reviewers tend to underrate quality the most.
Ivan Stelmakh, Nihar B. Shah, Aarti Singh, Hal Daumé III
Proc. ACM Hum. Comput. Interact.1
2019 PeerReview4All: Fair and Accurate Reviewer Assignment in Peer Review
abstract
We consider the problem of automated assignment of papers to reviewers in conference peer review, with a focus on fairness and statistical accuracy. Our fairness objective is to maximize the review quality of the most disadvantaged paper, in contrast to the popular objective of maximizing the total quality over all papers. We design an assignment algorithm based on an incremental max-flow procedure that we prove is near-optimally fair. Our statistical accuracy objective is to ensure correct recovery of the papers that should be accepted. With a sharp minimax analysis we also prove that our algorithm leads to assignments with strong statistical guarantees both in an objective-score model as well as a novel subjective-score model that we propose in this paper.
Ivan Stelmakh, Nihar B. Shah, Aarti Singh
ALT1
2019 On Testing for Biases in Peer Review
abstract
We consider the issue of biases in scholarly research, specifically, in peer review. There is a long standing debate on whether exposing author identities to reviewers induces biases against certain groups, and our focus is on designing tests to detect the presence of such biases. Our starting point is a remarkable recent work by Tomkins, Zhang and Heavlin which conducted a controlled, large-scale experiment to investigate existence of biases in the peer reviewing of the WSDM conference. We present two sets of results in this paper. The first set of results is negative, and pertains to the statistical tests and the experimental setup used in the work of Tomkins et al. We show that the test employed therein does not guarantee control over false alarm probability and under correlations between relevant variables, coupled with any of the following conditions, with high probability can declare a presence of bias when it is in fact absent: (a) measurement error, (b) model mismatch, (c) reviewer calibration. Moreover, we show that the setup of their experiment may itself inflate false alarm probability if (d) bidding is performed in non-blind manner or (e) popular reviewer assignment procedure is employed. Our second set of results is positive, in that we present a general framework for testing for biases in (single vs. double blind) peer review. We then present a hypothesis test with guaranteed control over false alarm probability and non-trivial power even under conditions (a)--(c). Conditions (d) and (e) are more fundamental problems that are tied to the experimental setup and not necessarily related to the test.
Ivan Stelmakh, Nihar B. Shah, Aarti Singh
NeurIPS1