VLDB 2026 Research / reviewers in the wild / expert
Jaehyo Yoo
dblp:249/2815
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2023
0000-0002-3600-6362ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Information extraction and text analysis · 51% Generative modeling · 33% Question answering and dialogue systems · 16% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling › synthetic data generation
dataset generation |
1.2 | 2 | 2023 | Automatic Creation of Named Entity Recognition Datasets by Querying Phrase Representations · ACL (1) 2023 Simple Questions Generate Named Entity Recognition Datasets · EMNLP 2022 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
1.2 | 2 | 2023 | Automatic Creation of Named Entity Recognition Datasets by Querying Phrase Representations · ACL (1) 2023 Simple Questions Generate Named Entity Recognition Datasets · EMNLP 2022 |
Natural language and speech › Information extraction and text analysis › named entity recognition
weakly supervised named entity recognition |
0.7 | 1 | 2023 | Automatic Creation of Named Entity Recognition Datasets by Querying Phrase Representations · ACL (1) 2023 |
Natural language and speech › Question answering and dialogue systems
question generation |
0.6 | 1 | 2022 | Simple Questions Generate Named Entity Recognition Datasets · EMNLP 2022 |
Information retrieval › document retrieval › text search
phrase-based retrieval |
0.2 | 1 | 2023 | Automatic Creation of Named Entity Recognition Datasets by Querying Phrase Representations · ACL (1) 2023 |
Methods — techniques the papers use, named apart from their topics
phrase embedding search · 1.3embedding distance verification · 1.3open-domain question answering · 0.6ask-to-generate · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Automatic Creation of Named Entity Recognition Datasets by Querying Phrase RepresentationsabstractMost weakly supervised named entity recognition (NER) models rely on domain-specific dictionaries provided by experts.This approach is infeasible in many domains where dictionaries do not exist.While a phrase retrieval model was used to construct pseudo-dictionaries with entities retrieved from Wikipedia automatically in a recent study, these dictionaries often have limited coverage because the retriever is likely to retrieve popular entities rather than rare ones.In this study, we present a novel framework, HighGEN, that generates NER datasets with high-coverage pseudo-dictionaries.Specifically, we create entity-rich dictionaries with a novel search method, called phrase embedding search, which encourages the retriever to search a space densely populated with various entities.In addition, we use a new verification process based on the embedding distance between candidate entity mentions and entity types to reduce the false-positive noise in weak labels generated by high-coverage dictionaries.We demonstrate that HighGEN outperforms the previous best model by an average F1 score of 4.7 across five NER benchmark datasets. Hyunjae Kim, Jaehyo Yoo, Seunghyun Yoon 0002, Jaewoo Kang |
ACL (1) | 2 |
| 2022 | Simple Questions Generate Named Entity Recognition DatasetsabstractRecent named entity recognition (NER) models often rely on human-annotated datasets, requiring the significant engagement of professional knowledge on the target domain and entities.This research introduces an ask-to-generate approach that automatically generates NER datasets by asking questions in simple natural language to an open-domain question answering system (e.g., "Which disease?").Despite using fewer in-domain resources, our models, solely trained on the generated datasets, largely outperform strong low-resource models by an average F1 score of 19.4 for six popular NER benchmarks.Furthermore, our models provide competitive performance with rich-resource models that additionally leverage in-domain dictionaries provided by domain experts.In few-shot NER, we outperform the previous best model by an F1 score of 5.2 on three benchmarks and achieve new state-of-the-art performance.The code and datasets are available at https://github.com/dmis-lab/GeNER. Hyunjae Kim, Jaehyo Yoo, Seunghyun Yoon 0002, Jinhyuk Lee, Jaewoo Kang |
EMNLP | 2 |