Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Sakshi Gandhi

dblp:237/9679 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Information extraction and text analysis · 44% Efficient and distributed learning · 44% Image recognition and object detection · 13%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
auto labeling
0.412020
GOGGLES: Automatic Image Labeling with Affinity Coding · SIGMOD Conference 2020
Natural language and speech › Information extraction and text analysis
data annotation
0.412020
GOGGLES: Automatic Image Labeling with Affinity Coding · SIGMOD Conference 2020
Computer vision › Image recognition and object detection
image classification
0.112020
GOGGLES: Automatic Image Labeling with Affinity Coding · SIGMOD Conference 2020
Machine learning and data management › weak supervision
data programming
0.112020
GOGGLES: Automatic Image Labeling with Affinity Coding · SIGMOD Conference 2020

Methods — techniques the papers use, named apart from their topics

hierarchical generative model · 0.9few-shot learning · 0.9affinity coding · 0.9
YearPublicationVenuePosition
2020 GOGGLES: Automatic Image Labeling with Affinity Coding
abstract
Generating large labeled training data is becoming the biggest bottleneck in building and deploying supervised machine learning models. Recently, the data programming paradigm has been proposed to reduce the human cost in labeling training data. However, data programming relies on designing labeling functions which still requires significant domain expertise. Also, it is prohibitively difficult to write labeling functions for image datasets as it is hard to express domain knowledge using raw features for images (pixels). We propose affinity coding, a new domain-agnostic paradigm for automated training data labeling. The core premise of affinity coding is that the affinity scores of instance pairs belonging to the same class on average should be higher than those of pairs belonging to different classes, according to some affinity functions. We build the GOGGLES system that implements affinity coding for labeling image datasets by designing a novel set of reusable affinity functions for images, and propose a novel hierarchical generative model for class inference using a small development set. We compare GOGGLES with existing data programming systems on 5 image labeling tasks from diverse domains. GOGGLES achieves labeling accuracies ranging from a minimum of 71% to a maximum of 98% without requiring any extensive human annotation. In terms of end-to-end performance, GOGGLES outperforms the state-of-the-art data programming system Snuba by 21% and a state-of-the-art few-shot learning technique by 5%, and is only 7% away from the fully supervised upper bound.
Nilaksh Das, Sanya Chaba, Renzhi Wu, Sakshi Gandhi, Polo Chau, Xu Chu 0002
SIGMOD Conference4