Abhinab Acharya

dblp:384/4310 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Data mining · 100%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 2 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › data reduction
data pruning
0.812024
Balancing Feature Similarity and Label Variability for Optimal Size-Aware One-shot Subset Selection · ICML 2024
Machine learning › Efficient and distributed learning › data selection
coreset selection
0.212024
Balancing Feature Similarity and Label Variability for Optimal Size-Aware One-shot Subset Selection · ICML 2024

Methods — techniques the papers use, named apart from their topics

core-set loss bound · 1.5beta-scoring importance function · 1.5
YearPublicationVenuePosition
2024 Balancing Feature Similarity and Label Variability for Optimal Size-Aware One-shot Subset Selection
abstract
Subset or core-set selection offers a data-efficient way for training deep learning models. One-shot subset selection poses additional challenges as subset selection is only performed once and full set data become unavailable after the selection. However, most existing methods tend to choose either diverse or difficult data samples, which fail to faithfully represent the joint data distribution that is comprised of both feature and label information. The selection is also performed independently from the subset size, which plays an essential role in choosing what types of samples. To address this critical gap, we propose to conduct Feature similarity and Label variability Balanced One-shot Subset Selection (BOSS), aiming to construct an optimal size-aware subset for data-efficient deep learning. We show that a novel balanced core-set loss bound theoretically justifies the need to simultaneously consider both diversity and difficulty to form an optimal subset. It also reveals how the subset size influences the bound. We further connect the inaccessible bound to a practical surrogate target which is tailored to subset sizes and varying levels of overall difficulty. We design a novel Beta-scoring importance function to delicately control the optimal balance of diversity and difficulty. Comprehensive experiments conducted on both synthetic and real data justify the important theoretical properties and demonstrate the superior performance of BOSS as compared with the competitive baselines.
Abhinab Acharya, Dayou Yu, Qi Yu 0001, Xumin Liu
ICML1