VLDB 2026 Research / reviewers in the wild / expert
Ivan Stelmakh
dblp:222/3200
· DBLP profile ↗
11ranked-venue papers
8as first author
9since 2021 · last 2024
0000-0002-9237-7379ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 7 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Question answering and dialogue systems · 36% Trustworthy machine learning · 24% Language models and text generation · 18% | |
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Computational social science and digital humanities · 100% | |
| Theoretical computer science
4 papers |
Algorithmic game theory and mechanism design · 57% Mathematical optimization · 33% Graph algorithms and graph theory · 10% | |
| Human-computer interaction and pervasive computing
3 papers |
Collaborative and social computing · 50% Human-AI interaction · 38% Usability and user experience research · 12% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Software testing · 50% Program analysis · 50% |
Topics — the 17 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computational social science and digital humanities › science of science
peer review |
1.0 | 3 | 2021 | A Novice-Reviewer Experiment to Address Scarcity of Qualified Reviewers in Large Conferences · AAAI 2021 On Testing for Biases in Peer Review · NeurIPS 2019 Catch Me if I Can: Detecting Strategic Behaviour in Peer Assessment · AAAI 2021 |
Machine learning › Trustworthy machine learning
fairness |
0.8 | 1 | 2024 | Debiasing Evaluations That Are Biased by Evaluations · J. Mach. Learn. Res. 2024 |
Natural language and speech › Question answering and dialogue systems › robust question answering
ambiguous question answering |
0.6 | 1 | 2022 | ASQA: Factoid Questions Meet Long-Form Answers · EMNLP 2022 |
Natural language and speech › Language models and text generation › natural language understanding › question answering
factoid question answering |
0.6 | 1 | 2022 | ASQA: Factoid Questions Meet Long-Form Answers · EMNLP 2022 |
Natural language and speech › Question answering and dialogue systems
long-form question answering |
0.6 | 1 | 2022 | ASQA: Factoid Questions Meet Long-Form Answers · EMNLP 2022 |
Information retrieval
evaluation |
0.6 | 1 | 2022 | ASQA: Factoid Questions Meet Long-Form Answers · EMNLP 2022 |
Information retrieval › evaluation › task-based evaluation
question answering evaluation |
0.6 | 1 | 2022 | ASQA: Factoid Questions Meet Long-Form Answers · EMNLP 2022 |
Human-AI interaction
human-AI collaboration |
0.5 | 1 | 2021 | Towards Fair, Equitable, and Efficient Peer Review · AAAI 2021 |
Collaborative and social computing › information sharing › scholarly communication
peer review |
0.5 | 1 | 2021 | Towards Fair, Equitable, and Efficient Peer Review · AAAI 2021 |
Mathematical optimization › regularization
regularized optimization |
0.5 | 1 | 2021 | Debiasing Evaluations That Are Biased by Evaluations · AAAI 2021 |
Program analysis
false alarm reduction |
0.4 | 1 | 2019 | On Testing for Biases in Peer Review · NeurIPS 2019 |
Software testing
statistical testing |
0.4 | 1 | 2019 | On Testing for Biases in Peer Review · NeurIPS 2019 |
Algorithmic game theory and mechanism design › market design › matching markets
reviewer assignment |
0.4 | 1 | 2019 | On Testing for Biases in Peer Review · NeurIPS 2019 |
Machine learning › Learning theory › model selection
cross-validation |
0.2 | 1 | 2024 | Debiasing Evaluations That Are Biased by Evaluations · J. Mach. Learn. Res. 2024 |
Computational social science and digital humanities
science of science |
0.1 | 1 | 2021 | Towards Fair, Equitable, and Efficient Peer Review · AAAI 2021 |
Collaborative and social computing
human computation |
0.1 | 1 | 2021 | A Novice-Reviewer Experiment to Address Scarcity of Qualified Reviewers in Large Conferences · AAAI 2021 |
Graph algorithms and graph theory › graph algorithms › network flow
maximum flow |
0.1 | 1 | 2021 | PeerReview4All: Fair and Accurate Reviewer Assignment in Peer Review · J. Mach. Learn. Res. 2021 |
Methods — techniques the papers use, named apart from their topics
regularized optimization · 2.3cross-validation · 2.3summarization · 1.1automated metric · 1.1hypothesis testing · 1.1subjective-score model · 1.0statistical hypothesis testing · 1.0minimax analysis · 1.0incremental max-flow · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Debiasing Evaluations That Are Biased by EvaluationsabstractIt is common to evaluate a set of items by soliciting people to rate them. For example, universities ask students to rate the teaching quality of their instructors, and conference organizers ask authors of submissions to evaluate the quality of the reviews. However, in these applications, students often give a higher rating to a course if they receive higher grades in a course, and authors often give a higher rating to the reviews if their papers are accepted to the conference. In this work, we call these external factors the "outcome" experienced by people, and consider the problem of mitigating these outcome-induced biases in the given ratings when some information about the outcome is available. We formulate the information about the outcome as a known partial ordering on the bias. We propose a debiasing method by solving a regularized optimization problem under this ordering constraint, and also provide a carefully designed cross-validation method that adaptively chooses the appropriate amount of regularization. We provide theoretical guarantees on the performance of our algorithm, as well as experimental evaluations. Jingyan Wang 0001, Ivan Stelmakh, Yuting Wei 0001, Nihar B. Shah |
J. Mach. Learn. Res. | 2 |
| 2022 | ASQA: Factoid Questions Meet Long-Form AnswersabstractAn abundance of datasets and availability of reliable evaluation metrics have resulted in strong progress in factoid question answering (QA).This progress, however, does not easily transfer to the task of long-form QA, where the goal is to answer questions that require in-depth explanations.The hurdles include (i) a lack of high-quality data, and (ii) the absence of a well-defined notion of the answer's quality.In this work, we address these problems by (i) releasing a novel dataset and a task that we call ASQA (Answer Summaries for Questions which are Ambiguous); and (ii) proposing a reliable metric for measuring performance on ASQA.Our task focuses on factoid questions that are ambiguous, that is, have different correct answers depending on interpretation.Answers to ambiguous questions should synthesize factual information from multiple sources into a long-form summary that resolves the ambiguity.In contrast to existing long-form QA tasks (such as ELI5), ASQA admits a clear notion of correctness: a user faced with a good summary should be able to answer different interpretations of the original ambiguous question.We use this notion of correctness to define an automated metric of performance for ASQA.Our analysis demonstrates an agreement between this metric and human judgments, and reveals a considerable gap between human performance and strong baselines. Ivan Stelmakh, Yi Luan, Bhuwan Dhingra, Ming-Wei Chang |
EMNLP | 1 |
| 2022 | No Rose for MLE: Inadmissibility of MLE for Evaluation Aggregation Under Levels of ExpertiseabstractA number of applications including crowd-sourced labeling and peer review require aggregation of labels or evaluations sourced from multiple evaluators. There is often additional information available pertaining to the evaluators’ expertise. A natural approach for aggregation is to consider the widely studied Dawid-Skene model (or its extensions incorporating evaluators’ expertise), and employ the standard maximum likelihood estimator (MLE). While MLE is in general widely used in practice and enjoys a number of appealing theoretical guarantees, in this work we provide a negative result for the MLE. Specifically, we prove that the MLE is asymptotically inadmissible for a special case of evaluation aggregation with expertise level information. We show this by constructing an alternative estimator that we show is significantly better than the MLE in certain parameter regimes and at least as good elsewhere. Finally, simulations reveal that our findings may hold in more general conditions than what we theoretically analyze. Charvi Rastogi, Ivan Stelmakh, Nihar B. Shah, Sivaraman Balakrishnan |
ISIT | 2 |
| 2021 | Debiasing Evaluations That Are Biased by EvaluationsabstractIt is common to evaluate a set of items by soliciting people to rate them. For example, universities ask students to rate the teaching quality of their instructors, and conference organizers ask authors of submissions to evaluate the quality of the reviews. However, in these applications, students often give a higher rating to a course if they receive higher grades in a course, and authors often give a higher rating to the reviews if their papers are accepted to the conference. In this work, we call these external factors the "outcome" experienced by people, and consider the problem of mitigating these outcome-induced biases in the given ratings when some information about the outcome is available. We formulate the information about the outcome as a known partial ordering on the bias. We propose a debiasing method by solving a regularized optimization problem under this ordering constraint, and also provide a carefully designed cross-validation method that adaptively chooses the appropriate amount of regularization. We provide theoretical guarantees on the performance of our algorithm, as well as experimental evaluations. Jingyan Wang 0001, Ivan Stelmakh, Yuting Wei 0001, Nihar B. Shah |
AAAI | 2 |
| 2021 | Towards Fair, Equitable, and Efficient Peer ReviewabstractPeer review is the backbone of academia. The rapid growth of the number of submissions to leading publication venues has identified a need for automation of some parts of the peer-review pipeline and nowadays human referees are required to interact with various interfaces and technologies in this process. However, there exists evidence that if such interactions are not carefully designed, they can exacerbate various problems related to fairness and efficiency of the process. In my research, I aim to design a Human-AI collaboration pipeline in peer review to mitigate these issues and ensure that science progresses in a fair, equitable, and efficient manner. Ivan Stelmakh |
AAAI | 1 |
| 2021 | Catch Me if I Can: Detecting Strategic Behaviour in Peer AssessmentabstractWe consider the issue of strategic behaviour in various peer-assessment tasks, including peer grading of exams or homeworks and peer review in hiring or promotions. When a peer-assessment task is competitive (e.g., when students are graded on a curve), agents may be incentivized to misreport evaluations in order to improve their own final standing. Our focus is on designing methods for detection of such manipulations. Specifically, we consider a setting in which agents evaluate a subset of their peers and output rankings that are later aggregated to form a final ordering. In this paper, we investigate a statistical framework for this problem and design a principled test for detecting strategic behaviour. We prove that our test has strong false alarm guarantees and evaluate its detection ability in practical settings. For this, we design and conduct an experiment that elicits strategic behaviour from subjects and release a dataset of patterns of strategic behaviour that may be of independent interest. We use this data to run a series of real and semi-synthetic evaluations that reveal a strong detection power of our test. Ivan Stelmakh, Nihar B. Shah, Aarti Singh |
AAAI | 1 |
| 2021 | A Novice-Reviewer Experiment to Address Scarcity of Qualified Reviewers in Large ConferencesabstractConference peer review constitutes a human-computation process whose importance cannot be overstated: not only it identifies the best submissions for acceptance, but, ultimately, it impacts the future of the whole research area by promoting some ideas and restraining others. A surge in the number of submissions received by leading AI conferences has challenged the sustainability of the review process by increasing the burden on the pool of qualified reviewers which is growing at a much slower rate. In this work, we consider the problem of reviewer recruiting with a focus on the scarcity of qualified reviewers in large conferences. Specifically, we design a procedure for (i) recruiting reviewers from the population not typically covered by major conferences and (ii) guiding them through the reviewing pipeline. In conjunction with the ICML 2020 --- a large, top-tier machine learning conference --- we recruit a small set of reviewers through our procedure and compare their performance with the general population of ICML reviewers. Our experiment reveals that a combination of the recruiting and guiding mechanisms allows for a principled enhancement of the reviewer pool and results in reviews of superior quality compared to the conventional pool of reviews as evaluated by senior members of the program committee (meta-reviewers). Ivan Stelmakh, Nihar B. Shah, Aarti Singh, Hal Daumé III |
AAAI | 1 |
| 2021 | PeerReview4All: Fair and Accurate Reviewer Assignment in Peer ReviewabstractWe consider the problem of automated assignment of papers to reviewers in conference peer review, with a focus on fairness and statistical accuracy. Our fairness objective is to maximize the review quality of the most disadvantaged paper, in contrast to the commonly used objective of maximizing the total quality over all papers. We design an assignment algorithm based on an incremental max-flow procedure that we prove is near-optimally fair. Our statistical accuracy objective is to ensure correct recovery of the papers that should be accepted. We provide a sharp minimax analysis of the accuracy of the peer-review process for a popular objective-score model as well as for a novel subjective-score model that we propose in the paper. Our analysis proves that our proposed assignment algorithm also leads to a near-optimal statistical accuracy. Finally, we design a novel experiment that allows for an objective comparison of various assignment algorithms, and overcomes the inherent difficulty posed by the absence of a ground truth in experiments on peer-review. The results of this experiment as well as of other experiments on synthetic and real data corroborate the theoretical guarantees of our algorithm. Ivan Stelmakh, Nihar B. Shah, Aarti Singh |
J. Mach. Learn. Res. | 1 |
| 2021 | Prior and Prejudice: The Novice Reviewers' Bias against Resubmissions in Conference Peer ReviewabstractModern machine learning and computer science conferences are experiencing a surge in the number of submissions that challenges the quality of peer review as the number of competent reviewers is growing at a much slower rate. To curb this trend and reduce the burden on reviewers, several conferences have started encouraging or even requiring authors to declare the previous submission history of their papers. Such initiatives have been met with skepticism among authors, who raise the concern about a potential bias in reviewers' recommendations induced by this information. In this work, we investigate whether reviewers exhibit a bias caused by the knowledge that the submission under review was previously rejected at a similar venue, focusing on a population of novice reviewers who constitute a large fraction of the reviewer pool in leading machine learning and computer science conferences. We design and conduct a randomized controlled trial closely replicating the relevant components of the peer-review pipeline with $133$ reviewers (master's, junior PhD students, and recent graduates of top US universities) writing reviews for $19$ papers. The analysis reveals that reviewers indeed become negatively biased when they receive a signal about paper being a resubmission, giving almost 1 point lower overall score on a 10-point Likert item (Δ = -0.78, 95% CI = [-1.30, -0.24]) than reviewers who do not receive such a signal. Looking at specific criteria scores (originality, quality, clarity and significance), we observe that novice reviewers tend to underrate quality the most. Ivan Stelmakh, Nihar B. Shah, Aarti Singh, Hal Daumé III |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2019 | PeerReview4All: Fair and Accurate Reviewer Assignment in Peer ReviewabstractWe consider the problem of automated assignment of papers to reviewers in conference peer review, with a focus on fairness and statistical accuracy. Our fairness objective is to maximize the review quality of the most disadvantaged paper, in contrast to the popular objective of maximizing the total quality over all papers. We design an assignment algorithm based on an incremental max-flow procedure that we prove is near-optimally fair. Our statistical accuracy objective is to ensure correct recovery of the papers that should be accepted. With a sharp minimax analysis we also prove that our algorithm leads to assignments with strong statistical guarantees both in an objective-score model as well as a novel subjective-score model that we propose in this paper. Ivan Stelmakh, Nihar B. Shah, Aarti Singh |
ALT | 1 |
| 2019 | On Testing for Biases in Peer ReviewabstractWe consider the issue of biases in scholarly research, specifically, in peer review. There is a long standing debate on whether exposing author identities to reviewers induces biases against certain groups, and our focus is on designing tests to detect the presence of such biases. Our starting point is a remarkable recent work by Tomkins, Zhang and Heavlin which conducted a controlled, large-scale experiment to investigate existence of biases in the peer reviewing of the WSDM conference. We present two sets of results in this paper. The first set of results is negative, and pertains to the statistical tests and the experimental setup used in the work of Tomkins et al. We show that the test employed therein does not guarantee control over false alarm probability and under correlations between relevant variables, coupled with any of the following conditions, with high probability can declare a presence of bias when it is in fact absent: (a) measurement error, (b) model mismatch, (c) reviewer calibration. Moreover, we show that the setup of their experiment may itself inflate false alarm probability if (d) bidding is performed in non-blind manner or (e) popular reviewer assignment procedure is employed. Our second set of results is positive, in that we present a general framework for testing for biases in (single vs. double blind) peer review. We then present a hypothesis test with guaranteed control over false alarm probability and non-trivial power even under conditions (a)--(c). Conditions (d) and (e) are more fundamental problems that are tied to the experimental setup and not necessarily related to the test. Ivan Stelmakh, Nihar B. Shah, Aarti Singh |
NeurIPS | 1 |