Shweta Garg 0004

dblp:305/6200 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2022
0000-0003-1737-0406ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Information extraction and text analysis · 84% Knowledge representation and reasoning · 8% Graph learning · 8%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
text classification
0.612022
MotifClass: Weakly Supervised Text Classification with Higher-order Metadata Information · WSDM 2022
Natural language and speech › Information extraction and text analysis › text classification
weakly supervised text classification
0.612022
MotifClass: Weakly Supervised Text Classification with Higher-order Metadata Information · WSDM 2022
Natural language and speech › Information extraction and text analysis
named entity recognition
0.512021
ChemNER: Fine-Grained Chemistry Named Entity Recognition with Ontology-Guided Distant Supervision · EMNLP (1) 2021
Machine learning › Graph learning
heterogeneous graph
0.212022
MotifClass: Weakly Supervised Text Classification with Higher-order Metadata Information · WSDM 2022
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph
0.212022
MotifClass: Weakly Supervised Text Classification with Higher-order Metadata Information · WSDM 2022
Natural language and speech › Information extraction and text analysis › relation extraction
distant supervision
0.112021
ChemNER: Fine-Grained Chemistry Named Entity Recognition with Ontology-Guided Distant Supervision · EMNLP (1) 2021

Methods — techniques the papers use, named apart from their topics

pseudo-labeling · 0.6motif mining · 0.6sequence labeling · 0.5ontology-guided distant supervision · 0.5
YearPublicationVenuePosition
2022 MotifClass: Weakly Supervised Text Classification with Higher-order Metadata Information
abstract
We study the problem of weakly supervised text classification, which aims to classify text documents into a set of pre-defined categories with category surface names only and without any annotated training document provided. Most existing classifiers leverage textual information in each document. However, in many domains, documents are accompanied by various types of metadata (e.g., authors, venue, and year of a research paper). These metadata and their combinations may serve as strong category indicators in addition to textual contents. In this paper, we explore the potential of using metadata to help weakly supervised text classification. To be specific, we model the relationships between documents and metadata via a heterogeneous information network. To effectively capture higher-order structures in the network, we use motifs to describe metadata combinations. We propose a novel framework, named MotifClass , which (1) selects category-indicative motif instances, (2) retrieves and generates pseudo-labeled training samples based on category names and indicative motif instances, and (3) trains a text classifier using the pseudo training data. Extensive experiments on real-world datasets demonstrate the superior performance of MotifClass to existing weakly supervised text classification approaches. Further analysis shows the benefit of considering higher-order metadata information in our framework.
Yu Zhang 0044, Shweta Garg 0004, Yu Meng 0001, Xiusi Chen, Jiawei Han 0001
WSDM2
2021 ChemNER: Fine-Grained Chemistry Named Entity Recognition with Ontology-Guided Distant Supervision
abstract
Scientific literature analysis needs fine-grained named entity recognition (NER) to provide a wide range of information for scientific discovery.For example, chemistry research needs to study dozens to hundreds of distinct, fine-grained entity types, making consistent and accurate annotation difficult even for crowds of domain experts.On the other hand, domain-specific ontologies and knowledge bases (KBs) can be easily accessed, constructed, or integrated, which makes distant supervision realistic for fine-grained chemistry NER.In distant supervision, training labels are generated by matching mentions in a document with the concepts in the knowledge bases (KBs).However, this kind of KB-matching suffers from two major challenges: incomplete annotation and noisy annotation.We propose CHEMNER, an ontologyguided, distantly-supervised method for finegrained chemistry NER to tackle these challenges.It leverages the chemistry type ontology structure to generate distant labels with novel methods of flexible KB-matching and ontology-guided multi-type disambiguation.It significantly improves the distant label generation for the subsequent sequence labeling model training.We also provide an expertlabeled, chemistry NER dataset with 62 finegrained chemistry types (e.g., chemical compounds and chemical reactions).Experimental results show that CHEMNER is highly effective, outperforming substantially the stateof-the-art NER methods (with .25 absolute F1 score improvement).
Xuan Wang 0008, Vivian Hu, Xiangchen Song, Shweta Garg 0004, Jinfeng Xiao, Jiawei Han 0001
EMNLP (1)4