VLDB 2026 Research / reviewers in the wild / expert
Mark Lindsey
dblp:57/692
· DBLP profile ↗
7ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0003-3071-4090ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-author · 4 since 2021Computer networks · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Machine learning and data management · 82% Information retrieval · 18% | |
| Artificial intelligence
2 papers |
Speech recognition and synthesis · 100% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 100% |
Topics — the 8 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning and data management
active learning |
1.0 | 1 | 2026 | The Value of Corrective Feedback in the Online Active Learning Paradigm · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Machine learning and data management › online learning
online active learning |
1.0 | 1 | 2026 | The Value of Corrective Feedback in the Online Active Learning Paradigm · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Natural language and speech › Speech recognition and synthesis › spoken document processing
speech summarization |
0.8 | 1 | 2024 | Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization? · ACL (1) 2024 |
Information retrieval
evaluation |
0.2 | 1 | 2024 | Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization? · ACL (1) 2024 |
Information retrieval › retrieval evaluation
reference-free evaluation |
0.2 | 1 | 2024 | Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization? · ACL (1) 2024 |
Network measurement and analytics
wireless network measurement |
0.0 | 1 | 2004 | Analysis of wireless information locality and association patterns in a campus · INFOCOM 2004 |
Wireless networking
WLAN |
0.0 | 1 | 2004 | Analysis of wireless information locality and association patterns in a campus · INFOCOM 2004 |
Content delivery and video streaming
caching |
0.0 | 1 | 2004 | Analysis of wireless information locality and association patterns in a campus · INFOCOM 2004 |
Methods — techniques the papers use, named apart from their topics
human evaluation · 2.3automatic metrics · 2.3LLM-based evaluation · 2.3corrective feedback · 2.0active learning · 2.0measurement study · 0.0markov chain modeling · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Value of Corrective Feedback in the Online Active Learning ParadigmabstractOnline Active Learning (OAL) is a powerful tool for classifying evolving data streams using limited annotations from a human operator who is a domain expert. The objective of the OAL learning paradigm is to minimize jointly the classification error rate and the annotation cost across the data stream by posing periodic Active Learning (AL) queries. In this paper, this objective is extended to include identification of classifier errors by the expert during the typical workflow. To this end, Corrective Feedback (CF) is introduced as a second channel of interaction between the expert and the learning algorithm, complementary to the AL channel, that allows the algorithm to obtain additional training labels without disrupting the expert's workflow. Online Active Learning with Corrective Feedback (OAL-CF) is formally defined as a paradigm, and its efficacy is proven through experimental application to two binary classification tasks, Spoken Language Verification and Voice-Type Discrimination. Finally, the effects of adding CF to the OAL paradigm are analyzed in terms of classification performance, annotation cost, trends over time, and class balance of the collected training data. Overall, the addition of CF results in a 53% relative reduction in cost compared to OAL without CF. Mark Lindsey, Francis Kubala, Richard M. Stern |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Iterative Feedback in the Online Active Learning ParadigmabstractOnline Active Learning with Corrective Feedback (OALCF) is a new machine learning paradigm that learns to detect events of interest in streaming audio by adapting to feedback from an operational user, who is a domain expert. The machine learns from each batch of the data stream by posing active learning queries to the expert and updating its model before making predictions. The expert reviews the predictions and indicates which of them are incorrect. This feedback is used to update the model for the next batch. In this paper, we introduce iterative feedback, where active learning and corrective feedback are performed multiple times per batch. We validate this approach on a large Spoken Language Verification task with 35 low-prevalence languages. The evaluation metric used (IMLM) accounts for the total cost of the method (i.e., error and feedback cost combined). The iterative algorithm achieves a 47.7% relative reduction in total cost compared to the original OAL-CF algorithm. Mark Lindsey, Francis Kubala, Richard M. Stern |
ASRU | 1 |
| 2025 | A Unified Metric for Simultaneous Evaluation of Error Rate and Annotation CostabstractPattern classification systems have traditionally been trained using a set of labeled training data and subsequently evaluated using different testing data. The cost of labeling the training data is typically substantial. Online Human-In-The-Loop (HITL) algorithms present an alternate approach that enables useful classification for many real-world applications using much less labeled data. These classifiers begin with a very small amount of training data and iteratively improve their performance by labeling a selected small number of utterances manually. Unfortunately, there is no unified evaluation metric that considers both classifier performance and annotation cost, which makes it difficult to evaluate these algorithms objectively. Furthermore, the lack of such a metric restricts the evaluation of online learning algorithms to prequential evaluation (before the classifier is adapted to the newly-labeled evaluation data), which does not realistically reflect the algorithm’s ability to adapt to the data stream in real time. This paper introduces the Interactive Machine Learning Metric (IMLM), a new unified evaluation metric that makes the combination of performance and annotation cost for binary classification tasks far less arbitrary. This metric is well suited for the evaluation of online HITL algorithms and also allows for fair comparison of different algorithms after adapting to the evaluation data. The value and appropriateness of IMLM is demonstrated by evaluating a series of Online Active Learning algorithms on a Spoken Language Verification task. Mark Lindsey, Francis Kubala, Richard M. Stern |
ICASSP | 1 |
| 2024 | Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?abstractReference summaries for abstractive speech summarization require human annotation, which can be performed by listening to an audio recording or by reading textual transcripts of the recording.In this paper, we examine whether summaries based on annotators listening to the recordings differ from those based on annotators reading transcripts.Using existing intrinsic evaluation based on human evaluation, automatic metrics, LLM-based evaluation, and a retrieval-based reference-free method.We find that summaries are indeed different based on the source modality, and that speechbased summaries are more factually consistent and information-selective than transcript-based summaries.Meanwhile, transcript-based summaries are impacted by recognition errors in the source, and expert-written summaries are more informative and reliable.We make all the collected data and analysis code public 1 to facilitate the reproduction of our work and advance research in this area. Roshan S. Sharma, Suwon Shon, Mark Lindsey, Hira Dhamyal, Bhiksha Raj |
ACL (1) | 3 |
| 2023 | Reducing the Cost of Spoof Detection Labeling using Mixed-Strategy Active Learning and Pretrained ModelsabstractActive learning is a powerful method for reducing the amount of labeled training data needed for a machine learning model to learn a task without degrading performance. This is accomplished by iteratively selecting the most informative samples from an unlabeled dataset to be labeled by an oracle (i.e., a human annotator) using an active learning sampling strategy. Pretrained models have been used in recent years as frontends for active learning neural networks to increase efficiency. This work applies active learning with pretrained models to the spoof detection task with the following two goals: 1) the identification of which pretrained speech models and active learning strategies are most effective for the spoof detection task, and 2) the development of an active learning method that selects the optimal sampling strategy from a list of available strategies at each step of the active learning process. This mixed strategy is shown to outperform all individual strategies for the task. Mark Lindsey, Nathaniel R. Robinson, Francis Kubala, Richard M. Stern |
ASRU | 1 |
| 2023 | Unsupervised Voice Type Discrimination Score Adaptation Using X-Vector ClustersabstractVoice type discrimination (VTD) is the task of automatically detecting speech produced in the same room as a recording device ("live speech") among other speech and non-speech noises, such as traffic noises or radio broadcasts ("distractor audio"). Existing work has described methods for performing the VTD task. This paper presents a method for adapting the output of these existing methods in an unsupervised manner via x-vector clustering and correlation. This adaptation method can be applied to the output of any VTD algorithm, requires no additional training data, and has been shown to yield a relative decrease in decision cost function (DCF) score of up to 47% on a standardized database collected for the task. Mark Lindsey, Tyler Vuong, Richard M. Stern |
ICASSP | 1 |
| 2004 | Analysis of wireless information locality and association patterns in a campusabstractOur goal is to explore characteristics of the environment that provide opportunities for caching, prefetching, coverage planning, and resource reservation. We conduct a one-month measurement study of locality phenomena among wireless Web users and their association patterns on a major university campus using the IEEE 802.11 wireless infrastructure. We evaluate the performance of different caching paradigms, such as single user cache, cache attached to an access point (AP), and peer-to-peer caching. In several settings such caching mechanisms could be beneficial. Unlike other measurement studies in wired networks in which 25% to 40% of documents draw 70% of Web access, our traces indicate that 13% of unique URLS draws this number of Web accesses. In addition, the overall ideal hit ratio of the user cache, cache attached to an access point, and peer-to-peer caching paradigms (where peers are coresident within an AP) are 51%, 55%, and 23%, respectively. We distinguish wireless clients based on their inter-building mobility, their visits to APs, their continuous walks in the wireless infrastructure, and their wireless information access during these periods. We model the associations as a Markov chain using as state information the most recent AP visits. We can predict with high probability (86%) the next AP with which a wireless client will associate. Also, there are APs with a high percentage of user revisits. Such measurements can benefit protocols and algorithms that aim to improve the performance of the wireless infrastructures by load balancing, admission control, and resource reservation across APs. Francisco Chinchilla, Mark Lindsey, Maria Papadopouli |
INFOCOM | 2 |