EDBT 2026 Demo / reviewers in the wild / expert
Gjorgjina Cenikj
dblp:264/2904
· DBLP profile ↗
3ranked-venue papers in the field
3as first author
2since 2021 · last 2022
0000-0002-2723-0821ORCID · verified
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 3 (3 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | SciFoodNER: Food Named Entity Recognition for Scientific TextabstractNamed Entity Recognition (NER) and Named Entity Linking (NEL) are key tasks in Information Extraction, addressing the identification and normalization of entity mentions from raw text. In the domain of food and nutrition, there have been several NER methods already developed, however, when applied to scientific text, they fail to generalize and produce large performance degradation. This introduces the need for new food NER and NEL models, developed specifically for extracting food entities from scientific text. In this paper, we present a scientific food NER and NEL model, SciFoodNER, obtained by fine-tuning transformer models on a corpus of scientific abstracts annotated with food entities. The models can identify mentions of food entites from raw text, and link the food entities to the Hansard Taxonomy, the FoodOn ontology and the Systematised Nomenclature of Medicine Clinical Terms (SNOMEDCT). Out of the evaluated models, the BioBERT model achieves the best results, reaching a median macro-averaged F1 score of 0.90 for the NER task, 0.66 for the NEL task linking to the Hansard Taxonomy, 0.43 for the NEL task linking to the FoodOn ontology and 0.58 for the NEL task linking to the SNOMEDCT ontology. Gjorgjina Cenikj, Gasper Petelin, Barbara Korousic-Seljak, Tome Eftimov |
IEEE Big Data | 1 |
| 2021 | Skills Named-Entity Recognition for Creating a Skill Inventory of Today's WorkplaceabstractTo trace the skills of the employees and to link them to their responsibilities, organizations need support from methods that automatically support such kind of activities. Even more, the COVID-19 pandemic completely transformed organizations’ workplace and culture, so project managers should take care of the skills of their team members in order to successfully realize a project. For this purpose, we have introduced rule- based named-entity recognition methods that are able to extract soft and technical skills using textual data such as job postings, curricula vitae (CVs), shout-outs, performance reviews, etc. The best model using a fusion of several skills-related dictionaries provides an F1 score of 59%, which can be further used as a base for creating an annotated silver corpus consisting of textual data and all skills found there. Further, we demonstrated the application of the developed model to link the required skills for a certain job title. Being able to trace the skills in an automatic way, we can further link them to tasks and actions to fully understand their economic value. Gjorgjina Cenikj, Bozhanka Vitanova, Tome Eftimov |
IEEE BigData | 1 |
| 2020 | BuTTER: BidirecTional LSTM for Food Named-Entity RecognitionabstractIn the modern era of big data, one of the biggest challenges is to find an efficient way of extracting information from unstructured data and structuring it in a form that can be interpreted and utilized by both humans and computers. In this paper, we focus on the domain of food and nutrition by introducing a Machine Learning (ML) based Named Entity Recognition (NER) method, which is a crucial step in extracting information from unstructured textual data. To the best of our knowledge, this is the first corpus-based food NER method that has been enabled by the recently published FoodBase corpus. The method is based on Bidirectional Long Short-Term Memory (BiLSTM) in conjunction with Conditional Random Fields (CRF) and Representation Learning (RL). Our experiments show that, despite the relatively small amount of annotated data, BuTTER is able to successfully identify food entities from raw text, with the best of the proposed models achieving an average macro F1 score of 0.946. Gjorgjina Cenikj, Gorjan Popovski, Riste Stojanov, Barbara Korousic-Seljak, Tome Eftimov |
IEEE BigData | 1 |