VLDB 2026 Research / reviewers in the wild / expert
Ji-Ung Lee
dblp:238/0875
· DBLP profile ↗
7ranked-venue papers
4as first author
5since 2021 · last 2023
0000-0002-8428-2003ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Information extraction and text analysis · 100% | |
| Human-computer interaction and pervasive computing
1 paper |
Learning and educational technologies · 77% Human-AI interaction · 23% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computing education · 100% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis
sentiment analysis |
0.5 | 1 | 2021 | Investigating label suggestions for opinion mining in German Covid-19 social media · ACL/IJCNLP (1) 2021 |
Learning and educational technologies
active learning |
0.4 | 1 | 2020 | Empowering Active Learning to Jointly Optimize System and User Demands · ACL 2020 |
Natural language and speech › Information extraction and text analysis
data annotation |
0.1 | 1 | 2021 | Investigating label suggestions for opinion mining in German Covid-19 social media · ACL/IJCNLP (1) 2021 |
Human-AI interaction › human-centered AI
human-centered machine learning |
0.1 | 1 | 2020 | Empowering Active Learning to Jointly Optimize System and User Demands · ACL 2020 |
Methods — techniques the papers use, named apart from their topics
active learning · 0.9label suggestion · 0.5user study · 0.4corpus-based experiment · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Lessons Learned from a Citizen Science Project for Natural Language ProcessingabstractJan-Christoph Klie, Ji-Ung Lee, Kevin Stowe, Gözde Şahin, Nafise Sadat Moosavi, Luke Bates, Dominic Petrak, Richard Eckart De Castilho, Iryna Gurevych. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Jan-Christoph Klie, Ji-Ung Lee, Kevin Stowe, Gözde Gül Sahin, Nafise Sadat Moosavi, Luke Bates, Dominic Petrak, Richard Eckart de Castilho, Iryna Gurevych |
EACL | 2 |
| 2023 | Efficient Methods for Natural Language Processing: A SurveyabstractAbstract Recent work in natural language processing (NLP) has yielded appealing results from scaling model parameters and training data; however, using only scale to improve performance means that resource consumption also grows. Such resources include data, time, storage, or energy, all of which are naturally limited and unevenly distributed. This motivates research into efficient methods that require fewer resources to achieve similar results. This survey synthesizes and relates current methods and findings in efficient NLP. We aim to provide both guidance for conducting NLP under limited resources, and point towards promising research directions for developing more efficient methods. Marcos V. Treviso, Ji-Ung Lee, Tianchu Ji, Betty van Aken, Manuel R. Ciosici, Michael Hassid, Kenneth Heafield, Sara Hooker, Colin Raffel, Pedro Henrique Martins, André F. T. Martins, Jessica Zosa Forde, Peter A. Milder, Edwin Simpson, Noam Slonim, Jesse Dodge, Emma Strubell, Niranjan Balasubramanian, Leon Derczynski, Iryna Gurevych, Roy Schwartz 0001 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2022 | Annotation Curricula to Implicitly Train Non-Expert AnnotatorsabstractAbstract Annotation studies often require annotators to familiarize themselves with the task, its annotation scheme, and the data domain. This can be overwhelming in the beginning, mentally taxing, and induce errors into the resulting annotations; especially in citizen science or crowdsourcing scenarios where domain expertise is not required. To alleviate these issues, this work proposes annotation curricula, a novel approach to implicitly train annotators. The goal is to gradually introduce annotators into the task by ordering instances to be annotated according to a learning curriculum. To do so, this work formalizes annotation curricula for sentence- and paragraph-level annotation tasks, defines an ordering strategy, and identifies well-performing heuristics and interactively trained models on three existing English datasets. Finally, we provide a proof of concept for annotation curricula in a carefully designed user study with 40 voluntary participants who are asked to identify the most fitting misconception for English tweets about the Covid-19 pandemic. The results indicate that using a simple heuristic to order instances can already significantly reduce the total annotation time while preserving a high annotation quality. Annotation curricula thus can be a promising research direction to improve data collection. To facilitate future research—for instance, to adapt annotation curricula to specific tasks and expert annotation scenarios—all code and data from the user study consisting of 2,400 annotations is made available.1 Ji-Ung Lee, Jan-Christoph Klie, Iryna Gurevych |
Comput. Linguistics | 1 |
| 2022 | Erratum: Annotation Curricula to Implicitly Train Non-Expert AnnotatorsabstractAbstract The authors of this work (“Annotation Curricula to Implicitly Train Non-Expert Annotators” by Ji-Ung Lee, Jan-Christoph Klie, and Iryna Gurevych in Computational Linguistics 48:2 https://doi.org/10.1162/coli_a_00436) discovered an incorrect inequality symbol in section 5.3 (page 360). The paper stated that the differences in the annotation times for the control instances result in a p-value of 0.200 which is smaller than 0.05 (p = 0.200 < 0.05). As 0.200 is of course larger than 0.05, the correct inequality symbol is p = 0.200 > 0.05, which is in line with the conclusion that follows in the text. The paper has been updated accordingly. Ji-Ung Lee, Jan-Christoph Klie, Iryna Gurevych |
Comput. Linguistics | 1 |
| 2021 | Investigating label suggestions for opinion mining in German Covid-19 social mediaabstractTilman Beck, Ji-Ung Lee, Christina Viehmann, Marcus Maurer, Oliver Quiring, Iryna Gurevych. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Tilman Beck, Ji-Ung Lee, Christina Viehmann, Marcus Maurer, Oliver Quiring, Iryna Gurevych |
ACL/IJCNLP (1) | 2 |
| 2020 | Empowering Active Learning to Jointly Optimize System and User DemandsabstractExisting approaches to active learning maximize the system performance by sampling unlabeled instances for annotation that yield the most efficient training.However, when active learning is integrated with an end-user application, this can lead to frustration for participating users, as they spend time labeling instances that they would not otherwise be interested in reading.In this paper, we propose a new active learning approach that jointly optimizes the seemingly counteracting objectives of the active learning system (training efficiently) and the user (receiving useful instances).We study our approach in an educational application, which particularly benefits from this technique as the system needs to rapidly learn to predict the appropriateness of an exercise to a particular user, while the users should receive only exercises that match their skills.We evaluate multiple learning strategies and user types with data from real users and find that our joint approach better satisfies both objectives when alternative methods lead to many unsuitable exercises for end users.1 Ji-Ung Lee, Christian M. Meyer, Iryna Gurevych |
ACL | 1 |
| 2019 | Manipulating the Difficulty of C-TestsabstractWe propose two novel manipulation strategies for increasing and decreasing the difficulty of C-tests automatically. This is a crucial step towards generating learner-adaptive exercises for self-directed language learning and preparing language assessment tests. To reach the desired difficulty level, we manipulate the size and the distribution of gaps based on absolute and relative gap difficulty predictions. We evaluate our approach in corpus-based experiments and in a user study with 60 participants. We find that both strategies are able to generate C-tests with the desired difficulty level. Ji-Ung Lee, Erik Schwan, Christian M. Meyer |
ACL (1) | 1 |