Diana Nurbakova

dblp:130/5835 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0002-6620-7771ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Theory of computation · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 IILAP+: Exploring AI-Assisted Dataset Creation for Critical Thinking Assessment
abstract
As LLM-based chatbots reshape traditional information-seeking behaviours, the need for AI literacy becomes increasingly urgent. Fluent, confident AI outputs can mask errors and biases, making users' ability to question, verify, and contextualize information crucial to preventing misinformation. In this paper, we introduce IILAP+, which extends the Interactive Information Literacy Assessment Platform (IILAP) with an AI assistant to support the creation of assessment datasets for critical reading of AI-generated content. Unlike the original IILAP, which relies on teacher-curated datasets, IILAP+ uses retrieval-augmented generation grounded in tiered verification sources, decomposes answers into atomic claims verified via SPARQL queries and natural language inference, then injects controlled errors. An educator-in-the-loop interface allows instructors to configure, review, edit, and publish generated items. We describe the architecture, the error taxonomy with mappings to the ACRL Information Literacy Framework, and a comparative evaluation of two LLM backends (DeepSeek V3, GPT-4.1-nano) across two seed question sources, analysing answer quality, error injection reliability, and cost tradeoffs. Results show that AI-assisted dataset creation can reduce authoring effort while preserving key assessment affordances, provided human oversight remains essential, positioning AI as a complementary tool rather than a replacement for expert curation.
Diana Nurbakova, Liana Ermakova
L@S1
2023 HUMMUS: A Linked, Healthiness-Aware, User-centered and Argument-Enabling Recipe Data Set for Recommendation
abstract
The overweight and obesity rate is increasing for decades worldwide. Healthy nutrition is, besides education and physical activity, one of the various keys to tackle this issue. In an effort to increase the availability of digital, healthy recommendations, the scientific area of food recommendation extends its focus from the accuracy of the recommendations to beyond-accuracy goals like transparency and healthiness. To address this issue a data basis is required, which in the ideal case encompasses user-item interactions like ratings and reviews, food-related information such as recipe details, nutritional data, and in the best case additional data which describes the food items and their relations semantically. Though several recipe recommendation data sets exist, to the best of our knowledge, a holistic large-scale healthiness-aware and connected data sets have not been made available yet. The lack of such data could partially explain the poor popularity of the topic of healthy food recommendation when compared to the domain of movie recommendation. In this paper, we show that taking into account only user-item interactions is not sufficient for a recommendation. To close this gap, we propose a connected data set called HUMMUS (Health-aware User-centered recoMMendation and argUment-enabling data Set) collected from Food.com containing multiple features including rich nutrient information, text reviews, and ratings, enriched by the authors with extra features such as Nutri-scores and connections to semantic data like the FoodKG and the FoodOn ontology. We hope that these data will contribute to the healthy food recommendation domain.
Felix Bölz, Diana Nurbakova, Sylvie Calabretto, Armin Gerl, Lionel Brunie, Harald Kosch
RecSys2
2022 Automatic Simplification of Scientific Texts: SimpleText Lab at CLEF-2022
Liana Ermakova, Patrice Bellot, Jaap Kamps, Diana Nurbakova, Irina Ovchinnikova, Eric SanJuan, Élise Mathurin, Sílvia Araújo, Radia Hannachi, Stéphane Huet, Nicolas Poinsu
ECIR (2)4
2021 Anytime Subgroup Discovery in High Dimensional Numerical Data
abstract
Subgroup discovery (SD) enables one to elicit patterns that strongly discriminate a class label. When it comes to numerical data, most of the existing SD approaches perform data discretizations and thus suffer from information loss. A few algorithms avoid such a loss by considering the search space of every interval pattern built on the dataset numerical values and provide an “anytime” property: at any moment, they are able to provide a result that improves over time. Given a sufficient time/memory budget, they may eventually complete an exhaustive search. However, such approaches are often intractable when dealing with high-dimensional numerical data, for instance, when extracting features from real-life multivariate time series. To overcome such limitations, we propose MonteCloPi, an approach based on a bottom-up exploration of numerical patterns with a Monte Carlo Tree Search. It enables to have a better exploration-exploitation trade-off between exploration and exploitation when sampling huge search spaces. Our extensive set of experiments proves the efficiency of MonteCloPi on high-dimensional data with hundreds of attributes. We finally discuss the actionability of discovered subgroups when looking for skill analysis from Rocket League action logs.
Romain Mathonat, Diana Nurbakova, Jean-François Boulicaut, Mehdi Kaytoue-Uberall
DSAA2
2021 Text Simplification for Scientific Information Access - CLEF 2021 SimpleText Workshop
Liana Ermakova, Patrice Bellot, Pavel Braslavski 0001, Jaap Kamps, Josiane Mothe, Diana Nurbakova, Irina Ovchinnikova, Eric SanJuan
ECIR (2)6
2021 Anytime mining of sequential discriminative patterns in labeled sequences
Romain Mathonat, Diana Nurbakova, Jean-François Boulicaut, Mehdi Kaytoue-Uberall
Knowl. Inf. Syst.2
2020 Enhancing Recommendation Diversity using Determinantal Point Processes on Knowledge Graphs
abstract
Top-N recommendations are widely applied in various real life domains and keep attracting intense attention from researchers and industry due to available multi-type information, new advances in AI models and deeper understanding of user satisfaction. Whileaccuracy has been the prevailing issue of the recommendation problem for the last decades, other facets of the problem, namelydiversity andexplainability, have received much less attention. In this paper, we focus on enhancing diversity of top-N recommendation, while ensuring the trade-off between accuracy and diversity. Thus, we propose an effective framework DivKG leveraging knowledge graph embedding and determinantal point processes (DPP). First, we capture different kinds of relations among users, items and additional entities through a knowledge graph structure. Then, we represent both entities and relations as k-dimensional vectors by optimizing a margin-based loss with all kinds of historical interactions. We use these representations to construct kernel matrices of DPP in order to make top-N diversified predictions. We evaluate our framework on MovieLens datasets coupled with IMDb dataset. Our empirical results show substantial improvement over the state-of-the-art regarding both accuracy and diversity metrics.
Lu Gan 0004, Diana Nurbakova, Léa Laporte, Sylvie Calabretto
SIGIR2
2019 SeqScout: Using a Bandit Model to Discover Interesting Subgroups in Labeled Sequences
abstract
It is extremely useful to exploit labeled datasets not only to learn models but also to improve our understanding of a domain and its available targeted classes. The so-called subgroup discovery task has been considered for a long time. It concerns the discovery of patterns or descriptions, the set of supporting objects of which have interesting properties, e.g., they characterize or discriminate a given target class. Though many subgroup discovery algorithms have been proposed for transactional data, discovering subgroups within labeled sequential data and thus searching for descriptions as sequential patterns has been much less studied. In that context, exhaustive exploration strategies can not be used for real-life applications and we have to look for heuristic approaches. We propose the algorithm SeqScout to discover interesting subgroups (w.r.t. a chosen quality measure) from labeled sequences of itemsets. This is a new sampling algorithm that mines discriminant sequential patterns using a multi-armed bandit model. It is an anytime algorithm that, for a given budget, finds a collection of local optima in the search space of descriptions and thus subgroups. It requires a light configuration and it is independent from the quality measure used for pattern scoring. Furthermore, it is fairly simple to implement. We provide qualitative and quantitative experiments on several datasets to illustrate its added-value.
Romain Mathonat, Diana Nurbakova, Jean-François Boulicaut, Mehdi Kaytoue-Uberall
DSAA2
2018 Recommendation of Activity Sequences during Distributed Events
abstract
Leisure activities constitute an important part of our life. Nowadays, the offer of activities to undertake is constantly growing. This can be easily seen not only by the increasing number of social events created and promoted on Facebook, Couchsurfing, etc., but also by the appearance of specialised online services and event-based social networks, such as Meetup, Eventbrite, etc. Moreover, multi-day events (e.g. conventions, festivals, cruise trips, exhibitions), to which we refer to as distributed events, attract thousands of participants. Their attendees are often overwhelmed with the amount of available options. Recommender systems appear as a common solution in such a context. In this project, we formulate the problem of recommendation of activity sequences and aim at providing an integrated support for users to create a personalised itinerary of activities in order to facilitate their decision making process which events to join. Such assistance is expected to bring a positive impact on well-being and satisfaction with life of individuals.
Diana Nurbakova
UMAP1