Dayou Yu

dblp:319/4611 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Trustworthy machine learning · 38% Efficient and distributed learning · 30% Learning paradigms · 19%
Databases, data mining, and information retrieval
2 papers
Data mining · 54% Machine learning and data management · 46%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
active learning
1.832023
Discover-Then-Rank Unlabeled Support Vectors in the Dual Space for Multi-Class Active Learning · ICML 2023
STARS: Spatial-Temporal Active Re-sampling for Label-Efficient Learning from Noisy Annotations · AAAI 2023
A Gaussian Process-Bayesian Bernoulli Mixture Model for Multi-Label Active Learning · NeurIPS 2021
Machine learning › Learning paradigms
multi-label classification
1.322024
Evidential Mixture Machines: Deciphering Multi-Label Correlations for Active Learning Sensitivity · NeurIPS 2024
A Gaussian Process-Bayesian Bernoulli Mixture Model for Multi-Label Active Learning · NeurIPS 2021
Machine learning › Trustworthy machine learning › uncertainty estimation › neural network uncertainty
evidential deep learning
0.812024
Evidential Mixture Machines: Deciphering Multi-Label Correlations for Active Learning Sensitivity · NeurIPS 2024
Machine learning › Efficient and distributed learning › active learning › active learning for classification
multi-label active learning
0.812024
Evidential Mixture Machines: Deciphering Multi-Label Correlations for Active Learning Sensitivity · NeurIPS 2024
Machine learning › Trustworthy machine learning › uncertainty estimation › uncertainty-aware learning
uncertainty-aware classification
0.812024
Evidential Mixture Machines: Deciphering Multi-Label Correlations for Active Learning Sensitivity · NeurIPS 2024
Data mining › data reduction
data pruning
0.812024
Balancing Feature Similarity and Label Variability for Optimal Size-Aware One-shot Subset Selection · ICML 2024
Machine learning › Trustworthy machine learning › model validation
active testing
0.712023
Actively Testing Your Model While It Learns: Realizing Label-Efficient Learning in Practice · NeurIPS 2023
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels
0.712023
STARS: Spatial-Temporal Active Re-sampling for Label-Efficient Learning from Noisy Annotations · AAAI 2023
Machine learning › Trustworthy machine learning
robustness
0.712023
STARS: Spatial-Temporal Active Re-sampling for Label-Efficient Learning from Noisy Annotations · AAAI 2023
Machine learning and data management
active learning
0.712023
Actively Testing Your Model While It Learns: Realizing Label-Efficient Learning in Practice · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.512021
A Gaussian Process-Bayesian Bernoulli Mixture Model for Multi-Label Active Learning · NeurIPS 2021
Machine learning › Learning paradigms › multi-label classification
label correlation modeling
0.512021
A Gaussian Process-Bayesian Bernoulli Mixture Model for Multi-Label Active Learning · NeurIPS 2021
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.512021
A Gaussian Process-Bayesian Bernoulli Mixture Model for Multi-Label Active Learning · NeurIPS 2021
Machine learning › Efficient and distributed learning › data selection
coreset selection
0.212024
Balancing Feature Similarity and Label Variability for Optimal Size-Aware One-shot Subset Selection · ICML 2024

Methods — techniques the papers use, named apart from their topics

core-set loss bound · 1.5beta-scoring importance function · 1.5uncertainty estimation · 0.8mixture of bernoulli · 0.8evidential learning · 0.8maximum-margin classifier · 0.7max-margin classifier · 0.7kernel methods · 0.7early stopping · 0.7dual space analysis · 0.7active testing · 0.7active resampling · 0.7active learning · 0.7active feedback · 0.7
YearPublicationVenuePosition
2024 Balancing Feature Similarity and Label Variability for Optimal Size-Aware One-shot Subset Selection
abstract
Subset or core-set selection offers a data-efficient way for training deep learning models. One-shot subset selection poses additional challenges as subset selection is only performed once and full set data become unavailable after the selection. However, most existing methods tend to choose either diverse or difficult data samples, which fail to faithfully represent the joint data distribution that is comprised of both feature and label information. The selection is also performed independently from the subset size, which plays an essential role in choosing what types of samples. To address this critical gap, we propose to conduct Feature similarity and Label variability Balanced One-shot Subset Selection (BOSS), aiming to construct an optimal size-aware subset for data-efficient deep learning. We show that a novel balanced core-set loss bound theoretically justifies the need to simultaneously consider both diversity and difficulty to form an optimal subset. It also reveals how the subset size influences the bound. We further connect the inaccessible bound to a practical surrogate target which is tailored to subset sizes and varying levels of overall difficulty. We design a novel Beta-scoring importance function to delicately control the optimal balance of diversity and difficulty. Comprehensive experiments conducted on both synthetic and real data justify the important theoretical properties and demonstrate the superior performance of BOSS as compared with the competitive baselines.
Abhinab Acharya, Dayou Yu, Qi Yu 0001, Xumin Liu
ICML2
2024 Evidential Mixture Machines: Deciphering Multi-Label Correlations for Active Learning Sensitivity
abstract
Multi-label active learning is a crucial yet challenging area in contemporary machine learning, often complicated by a large and sparse label space. This challenge is further exacerbated in active learning scenarios where labeling resources are constrained. Drawing inspiration from existing mixture of Bernoulli models, which efficiently compress the label space into a more manageable weight coefficient space by learning correlated Bernoulli components, we propose a novel model called Evidential Mixture Machines (EMM). Our model leverages mixture components derived from unsupervised learning in the label space and improves prediction accuracy by predicting weight coefficients following the evidential learning paradigm. These coefficients are aggregated as proxy pseudo counts to enhance component offset predictions. The evidential learning approach provides an uncertainty-aware connection between input features and the predicted coefficients and components. Additionally, our method combines evidential uncertainty with predicted label embedding covariances for active sample selection, creating a richer, multi-source uncertainty metric beyond traditional uncertainty scores. Experiments on synthetic datasets show the effectiveness of evidential uncertainty prediction and EMM's capability to capture label correlations through predicted components. Further testing on real-world datasets demonstrates improved performance compared to existing multi-label active learning methods.
Dayou Yu, Weishi Shi, Qi Yu 0001
NeurIPS1
2023 STARS: Spatial-Temporal Active Re-sampling for Label-Efficient Learning from Noisy Annotations
abstract
Active learning (AL) aims to sample the most informative data instances for labeling, which makes the model fitting data efficient while significantly reducing the annotation cost. However, most existing AL models make a strong assumption that the annotated data instances are always assigned correct labels, which may not hold true in many practical settings. In this paper, we develop a theoretical framework to formally analyze the impact of noisy annotations and show that systematically re-sampling guarantees to reduce the noise rate, which can lead to improved generalization capability. More importantly, the theoretical framework demonstrates the key benefit of conducting active re-sampling on label-efficient learning, which is critical for AL. The theoretical results also suggest essential properties of an active re-sampling function with a fast convergence speed and guaranteed error reduction. This inspires us to design a novel spatial-temporal active re-sampling function by leveraging the important spatial and temporal properties of maximum-margin classifiers. Extensive experiments conducted on both synthetic and real-world data clearly demonstrate the effectiveness of the proposed active re-sampling function.
Dayou Yu, Weishi Shi, Qi Yu 0001
AAAI1
2023 Discover-Then-Rank Unlabeled Support Vectors in the Dual Space for Multi-Class Active Learning
abstract
We propose to approach active learning (AL) from a novel perspective of discovering and then ranking potential support vectors by leveraging the key properties of the dual space of a sparse kernel max-margin predictor. We theoretically analyze the change of a hinge loss in the dual form and provide both the upper and lower bounds that are deeply connected to the key geometric properties induced by the dual space, which then help us identify various types of important data samples for AL. These bounds inform the design of a novel sampling strategy that leverages class-wise evidence as a key vehicle, formed through an affine combination of dual variables and kernel evaluation. We construct two distinct types of sampling functions, including discovery and ranking. The former focuses on samples with low total evidence from all classes, which signifies their potential to support exploration; the latter exploits the current decision boundary to identify the most conflicting regions for sampling, aiming to further refine the decision boundary. These two functions, which are complementary to each other, are automatically arranged into a two-phase active sampling process that starts with the discovery and then transitions to the ranking of data points to most effectively balance exploration and exploitation. Experiments on various real-world data demonstrate the state-of-the-art AL performance achieved by our model.
Dayou Yu, Weishi Shi, Qi Yu 0001
ICML1
2023 Actively Testing Your Model While It Learns: Realizing Label-Efficient Learning in Practice
abstract
In active learning (AL), we focus on reducing the data annotation cost from the model training perspective. However, "testing'', which often refers to the model evaluation process of using empirical risk to estimate the intractable true generalization risk, also requires data annotations. The annotation cost for "testing'' (model evaluation) is under-explored. Even in works that study active model evaluation or active testing (AT), the learning and testing ends are disconnected. In this paper, we propose a novel active testing while learning (ATL) framework that integrates active learning with active testing. ATL provides an unbiased sample-efficient estimation of the model risk during active learning. It leverages test samples annotated from different periods of a dynamic active learning process to achieve fair model evaluations based on a theoretically guaranteed optimal integration of different test samples. Periodic testing also enables effective early-stopping to further save the total annotation cost. ATL further integrates an "active feedback'' mechanism, which is inspired by human learning, where the teacher (active tester) provides immediate guidance given by the prior performance of the student (active learner). Our theoretical result reveals that active feedback maintains the label complexity of the integrated learning-testing objective, while improving the model's generalization capability. We study the realistic setting where we maximize the performance gain from choosing "testing'' samples for feedback without sacrificing the risk estimation accuracy. An agnostic-style analysis and empirical evaluations on real-world datasets demonstrate that the ATL framework can effectively improve the annotation efficiency of both active learning and evaluation tasks.
Dayou Yu, Weishi Shi, Qi Yu 0001
NeurIPS1
2021 A Gaussian Process-Bayesian Bernoulli Mixture Model for Multi-Label Active Learning
abstract
Multi-label classification (MLC) allows complex dependencies among labels, making it more suitable to model many real-world problems. However, data annotation for training MLC models becomes much more labor-intensive due to the correlated (hence non-exclusive) labels and a potential large and sparse label space. We propose to conduct multi-label active learning (ML-AL) through a novel integrated Gaussian Process-Bayesian Bernoulli Mixture model (GP-B$^2$M) to accurately quantify a data sample's overall contribution to a correlated label space and choose the most informative samples for cost-effective annotation. In particular, the B$^2$M encodes label correlations using a Bayesian Bernoulli mixture of label clusters, where each mixture component corresponds to a global pattern of label correlations. To tackle highly sparse labels under AL, the B$^2$M is further integrated with a predictive GP to connect data features as an effective inductive bias and achieve a feature-component-label mapping. The GP predicts coefficients of mixture components that help to recover the final set of labels of a data sample. A novel auxiliary variable based variational inference algorithm is developed to tackle the non-conjugacy introduced along with the mapping process for efficient end-to-end posterior inference. The model also outputs a predictive distribution that provides both the label prediction and their correlations in the form of a label covariance matrix. A principled sampling function is designed accordingly to naturally capture both the feature uncertainty (through GP) and label covariance (through B$^2$M) for effective data sampling. Experiments on real-world multi-label datasets demonstrate the state-of-the-art AL performance of the proposed GP-B$^2$M model.
Weishi Shi, Dayou Yu, Qi Yu 0001
NeurIPS2