Milica Ikonic Nesic

dblp:306/8896 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Integrating TEI, NER/NEL, Textometry, and Linked Data for a Semantically Enriched Interview Corpus
abstract
This paper presents a pipeline that converts unstructured interview transcripts into a semantically enriched, queryable knowledge resource. The texts from the Digitalne Ikone 20+ interview collection were first encoded in TEI XML (Text Encoding Initiative), marking interview boundaries, paragraph breaks, speaker turns with identifiers, dates, and topics. This structural encoding underpins downstream NLP and enables structured querying (e.g., by speaker). We then applied Named Entity Recognition to identify persons, places, organizations, and events, and embedded the results directly in TEI. In the third stage, Named Entity Linking mapped entity mentions to canonical Wikidata identifiers via context-aware disambiguation; missing entries were added to Wikidata when necessary. The resulting TEI+NER/NEL corpus, serialized as linked data, follows the NIF (NLP Interchange Framework). The pipeline also supports retrieval-augmented summarization that retrieves evidence passages and prompts LLMs (implemented with DSPy) to produce faithful interview summaries. We discuss design choices (TXM for textometry with JeRTeh resources; TESLA models for NER/NEL), report qualitative gains in interpretability through semantic links, and outline future work on domain-adapted NER/NEL, graph-based completion, and more expressive RAG architectures. The approach is replicable for other oral-history or media corpora and advances practical, evidence-grounded access to cultural archives and beyond.
Ranka Stankovic, Tamara Vucenovic, Biljana Rujevic, Milica Ikonic Nesic, Mihailo Skoric
LREC4
2024 Topic Modeling of the SrpELTeC Corpus: A Comparison of NMF, LDA, and BERTopic
abstract
Topic modeling is an effective way to gain insight into large amounts of data.Some of the most widely used topic models are Latent Dirichlet allocation (LDA) and Nonnegative Matrix Factorization (NMF).However, new ways to mine topics have emerged with the rise of self-attention models and pretrained language models.BERTopic represents the current stateof-the-art when it comes to modeling topics.In this paper, we compared LDA, NMF, and BERTopic performance on literary texts in the Serbian language, both quantitatively by measuring Topic Coherency (TC) and Topic Diversity (TD), and by conducting a qualitative evaluation of the obtained topics.Additionally, for BERTopic, we compared multilingual sentence transformer embeddings with the Jerteh-355 monolingual embeddings for Serbian.NMF yielded the best Topic Coherency results, while BERTopic with Jerteh-355 embeddings gave the best Topic Diveristy.The monolingual Serbian Jerteh-355 embeddings also outperformed sentence transformer embeddings in both TC and TD.
Teodora Mihajlov, Milica Ikonic Nesic, Ranka Stankovic, Olivera Kitanovic
FedCSIS2
2024 SrpCNNeL: Serbian Model for Named Entity Linking
abstract
This paper presents the development of a Named Entity Linking (NEL) model to the Wikidata knowledge base for the Serbian language, named SrpCNNeL.The model was trained to recognize and link seven different named entity types (persons, locations, organizations, professions, events, demonyms, and works of art) on a dataset containing sentences from novels, legal documents, as well as sentences generated from the Wikidata knowledge base and the Leximirka lexical database.The resulting model demonstrated robust performance, achieving an F1 score of 0.8 on the test set.Considering that the dataset contains the highest number of locations linked to the knowledge base, an evaluation was conducted on an independent dataset and compared to the baseline Spacy Entity Linker for locations only.
Milica Ikonic Nesic, Sasa Petalinkar, Ranka Stankovic, Milos Utvic, Olivera Kitanovic
FedCSIS1
2022 Distant Reading in Digital Humanities: Case Study on the Serbian Part of the ELTeC Collection
abstract
In this paper we present the Serbian part of the ELTeC multilingual corpus of novels written in the time period 1840-1920. The corpus is being built in order to test various distant reading methods and tools with the aim of re-thinking the European literary history. We present the various steps that led to the production of the Serbian sub-collection: the novel selection and retrieval, text preparation, structural annotation, POS-tagging, lemmatization and named entity recognition. The Serbian sub-collection was published on different platforms in order to make it freely available to various users. Several use examples show that this sub-collection is usefull for both close and distant reading approaches.
Ranka Stankovic, Cvetana Krstev, Branislava Sandrih, Dusko Vitas, Mihailo Skoric, Milica Ikonic Nesic
LREC6