VLDB 2026 Research / reviewers in the wild / expert
Mirela Popa
dblp:165/6065
· DBLP profile ↗
14ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0002-6449-1158ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 since 2021Artificial intelligence and machine learning · 5 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Little More Like This: Text-to-Image Retrieval with Vision-Language Models Using Relevance FeedbackabstractLarge vision-language models (VLMs) enable intuitive visual search using natural language queries. However, improving their performance often requires fine-tuning and scaling to larger model variants. In this work, we propose a mechanism inspired by traditional text-based search to improve retrieval performance at inference time: relevance feedback. While relevance feedback can serve as an alternative to fine-tuning, its model-agnostic design also enables use with fine-tuned VLMs. Specifically, we introduce and evaluate four feedback strategies for VLM-based retrieval. First, we revise classical pseudo-relevance feedback (PRF), which refines query embeddings based on top-ranked results. To address its limitations, we propose generative relevance feedback (GRF), which uses synthetic captions for query refinement. Furthermore, we introduce an attentive feedback summarizer (AFS), a custom transformer-based model that integrates multimodal fine-grained features from relevant items. Finally, we simulate explicit feedback using ground-truth captions as an upper-bound baseline. Experiments on Flickr30k and COCO with the VLM backbones show that GRF, AFS, and explicit feedback improve retrieval performance by 3–5% in MRR@5 for smaller VLMs, and 1–3% for larger ones, compared to retrieval with no feedback. Moreover, AFS, similarly to explicit feedback, mitigates query drift and is more robust than GRF in iterative, multi-turn retrieval settings. Our findings demonstrate that relevance feedback can consistently enhance retrieval across VLMs and open up opportunities for interactive and adaptive visual search. Bulat Khaertdinov, Mirela Popa, Nava Tintarev |
WACV | 2 |
| 2025 | VisualReF: Interactive Image Search Prototype with Visual Relevance FeedbackabstractIn the absence of interaction history, image recommendations often depend on content-based approaches. Prompted by user queries in natural language, such systems rank items based on the similarity between textual and visual features. However, these approaches typically rely on static queries and do not offer alternative feedback mechanisms. In this paper, we present VisualReF: an interactive image retrieval prototype that introduces visual relevance feedback through fine-grained user annotations. Built on vision-language models (VLMs) for retrieval, our system allows users to label relevant and irrelevant regions in retrieved images. These regions are captioned using a generative vision-language model to refine the query vector. Our work bridges the gap between conventional static image retrieval and interactive, user-guided search by introducing visual relevance feedback. Finally, our prototype contributes to the field of visual recommendation by empowering researchers with practical tools for: (i) collecting region-level visual relevance signals from users, (ii) supporting integration of human feedback into interactive search pipelines, and (iii) explaining how the relevance feedback model perceives user input. Bulat Khaertdinov, Mirela Popa, Nava Tintarev |
RecSys | 2 |
| 2025 | Proactive robot task sequencing through real-time hand motion prediction in human-robot collaborationabstractHuman–robot collaboration (HRC) is essential for improving productivity and safety across various industries. While reactive motion re-planning strategies are useful, there is a growing demand for proactive methods that predict human intentions to enable more efficient collaboration. This study addresses this need by introducing a framework that combines deep learning-based human hand trajectory forecasting with heuristic optimization for robotic task sequencing. The deep learning model advances real-time hand position forecasting using a multi-task learning loss to account for both hand positions and contact delay regression, achieving state-of-the-art performance on the Ego4D Future Hand Prediction benchmark. By integrating hand trajectory predictions into task planning, the framework offers a cohesive solution for HRC. To optimize task sequencing, the framework incorporates a Dynamic Variable Neighborhood Search (DynamicVNS) heuristic algorithm, which allows robots to pre-plan task sequences and avoid potential collisions with human hand positions. DynamicVNS provides significant computational advantages over the generalized VNS method. The framework was validated on a UR10e robot performing a visual inspection task in a HRC scenario, where the robot effectively anticipated and responded to human hand movements in a shared workspace. Experimental results highlight the system’s effectiveness and potential to enhance HRC in industrial settings by combining predictive accuracy and task planning efficiency. • We enhance hand position forecasting with a novel loss, achieving state-of-the-art results. • Our method unifies forecasting and task planning using the Dynamic TSP with Time Windows. • We enable real-time hand motion prediction and seamless integration into robot planning. Shyngyskhan Abilkassov, Michael Gentner, Almas Shintemirov, Eckehard G. Steinbach, Mirela Popa |
Image Vis. Comput. | 5 |
| 2024 | XPCA Gen: Extended PCA Based Tabular Data Generation ModelabstractThe proposed method XPCA Gen, introduces a novel approach for synthetic tabular data generation by utilising relevant patterns present in the data. This is performed using principle components obtained through XPCA (probabilistic interpretation of standard PCA) decomposition of original data. Since new data points are obtained by synthesizing the principle components, the generated data is an accurate and noise redundant representation of original data with a good diversity of data points. The experimental results obtained on benchmark datasets (e.g. CMC, PID) demonstrate performance in ML utility metrics (accuracy, precision, recall), showing its ability to capture inherent patterns in the dataset. Along with ML utility metrics, high Hausdorff distance indicates diversity in generated data without compromising statistical properties. Moreover, this is not a data hungry method like other complex neural networks. Overall, XPCA Gen emerges as a promising solution for data privacy preservation and robust model training with diverse samples. Sreekala Kallidil Padinjarekkara, Jessica Alecci, Mirela Popa |
ICPRAM | 3 |
| 2024 | XPCA Gen: Extended PCA Based Tabular Data Generation Model with Regularisation
Sreekala Kallidil Padinjarekkara, Jessica Alecci, Mirela Popa |
ICPRAM | 3 |
| 2024 | Tab-VAE: A Novel VAE for Generating Synthetic Tabular DataabstractVariational Autoencoders (VAEs) suffer from a well-known problem of overpruning or posterior collapse due to strong regularization while working in a sufficiently high-dimensional latent space. When VAEs are used to generate tabular data, categorical one-hot encoded data expand the dimensionality of the feature space dramatically, making modeling multi-class categorical data challenging. In this paper, we propose Tab-VAE, a novel VAE-based approach to generate synthetic tabular data that tackles this challenge by introducing a sampling technique at inference for categorical variables. A detailed review of the current state-of-theart models shows that most of the tabular data generation approaches draw methodologies from Generative Adversarial Networks (GANs) while a simpler more stable VAE method is ignored. Our extensive evaluation of the Tab-VAE with other leading generative models shows Tab-VAE improves the state-of-the-art VAEs significantly. It also shows that Tab-VAE outperforms the best GAN-based tabular data generators, paving the way for a powerful and less computationally expensive tabular data generation model. Syed Mahir Tazwar, Max Knobbout, Enrique Hortal, Mirela Popa |
ICPRAM | 4 |
| 2023 | Analysis of the Impact of Data Augmentation on the Performance of Deep Learning Models in Multispectral Food Authenticity IdentificationabstractFood authenticity is a significant concern in the meat industry, demanding effective detection methods.This study explores the use of multispectral imaging (MSI) and deep learning for meat adulteration detection.We evaluate different deep learning models using transfer learning and preprocessing techniques in a multi-level adulteration classification task.In addition, we propose a novel approach called one-band mixed augmentation for band selection in MSI data, which outperforms traditional reflectance-based feature selection and enhances model robustness.Furthermore, employing the ninecrop approach for dataset augmentation improved the accuracy from 0.63 to 0.74 for DenseNet201 model without transfer learning.This research contributes to advancing food safety assessment practices and provides insights into the application of deep learning for preventing food adulteration.The proposed one-band mixed augmentation approach offers a novel strategy for handling band selection challenges in MSI data analysis. Yaru Zhang, Arif Yilmaz, Mirela Popa, Christopher Brewster |
FedCSIS | 3 |
| 2022 | The Effect of Spatial and Temporal Occlusion on Word Level Sign Language RecognitionabstractDriven by the appeal of real-world applicable models, we investigate how temporal and spatial occlusion affect sign language recognition. Utilizing only a crop of the hands and pose flow, we maintain accuracies comparable to an I3D baseline for the WLASL dataset using a video transformer model (VTN), implying that hand crops might contain enough information for accurate prediction. Moreover, we find that a crop of only the right hand provides enough data to train an accurate model, achieving results of 0.2% less than the baseline for AUTSL and 4.7% less across all WLASL datasets. Sampling a video every fifth frame achieves comparative results to baseline, with 8 frame sequences performing better for AUTSL (0.4% less than baseline) and 16 frames performing better for WLASL (0.2% for WLASL 100 and 300). Our results indicate the feasibility of utilizing less information for sign language recognition, however more research is necessary to apply these findings in real-world scenarios. Ajkel Mino, Mirela Popa, Alexia Briassouli |
ICIP | 2 |
| 2020 | A hierarchical autoencoder learning model for path prediction and abnormality detection
Dario Dotti, Mirela Popa, Stylianos Asteriadis |
Pattern Recognit. Lett. | 2 |
| 2020 | Being the Center of Attention: A Person-Context CNN Framework for Personality RecognitionabstractThis article proposes a novel study on personality recognition using video data from different scenarios. Our goal is to jointly model nonverbal behavioral cues with contextual information for a robust, multi-scenario, personality recognition system. Therefore, we build a novel multi-stream Convolutional Neural Network (CNN) framework, which considers multiple sources of information. From a given scenario, we extract spatio-temporal motion descriptors from every individual in the scene, spatio-temporal motion descriptors encoding social group dynamics, and proxemics descriptors to encode the interaction with the surrounding context. All the proposed descriptors are mapped to the same feature space facilitating the overall learning effort. Experiments on two public datasets demonstrate the effectiveness of jointly modeling the mutual Person-Context information, outperforming the state-of-the art-results for personality recognition in two different scenarios. Last, we present CNN class activation maps for each personality trait, shedding light on behavioral patterns linked with personality attributes. Dario Dotti, Mirela Popa, Stylianos Asteriadis |
ACM Trans. Interact. Intell. Syst. | 2 |
| 2019 | Multimodal and Temporal Perception of Audio-visual Cues for Emotion RecognitionabstractIn Audio-Video Emotion Recognition (AVER), the idea is to have a human-level understanding of emotions from video clips. There is a need to bring these two modalities into a unified framework, to effectively learn multimodal fusion for AVER. In addition, literature studies lack in-depth analysis and utilization of how emotions vary as a function of time. Psychological and neurological studies show that negative and positive emotions are not recognized at the same speed. In this paper, we propose a novel multimodal temporal deep network framework that embeds video clips using their audio-visual content, onto a metric space, where their gap is reduced and their complementary and supplementary information is explored. We address two research questions, (1) how audio-visual cues contribute to emotion recognition and (2) how temporal information impacts the recognition rate and speed of emotions. The proposed method is evaluated on two datasets, CREMA-D and RAVDESS. The study findings are promising, achieving the state-of-the-art performance on both datasets, and showing a significant impact of multimodal and temporal emotion perception. Esam Ghaleb, Mirela Popa, Stylianos Asteriadis |
ACII | 2 |
| 2018 | Towards Affect Recognition through Interactions with Learning MaterialsabstractAffective state recognition has recently attracted a notable amount of attention in the research community, as it can be directly linked to a student's performance during learning. Consequently, being able to retrieve the affect of a student can lead to more personalized education, targeting higher degrees of engagement and, thus, optimizing the learning experience and its outcomes. In this paper, we apply Machine Learning (ML) and present a novel approach for affect recognition in Technology-Enhanced Learning (TEL) by understanding learners' experience through tracking their interactions with a serious game as a learning platform. We utilize a variety of interaction parameters to examine their potential to be used as an indicator of the learner's affective state. Driven by the Theory of Flow model, we investigate the correspondence between the prediction of users' self-reported affective states and the interaction features. Cross-subject evaluation using Support Vector Machines (SVMs) on a dataset of 32 participants interacting with the platform demonstrated that the proposed framework could achieve a significant precision in affect recognition. The subject-based evaluation highlighted the benefits of an adaptive personalized learning experience, contributing to achieving optimized levels of engagement. Esam Ghaleb, Mirela Popa, Enrique Hortal, Stylianos Asteriadis, Gerhard Weiss 0001 |
ICMLA | 2 |
| 2017 | Multimodal monitoring of Parkinson's and Alzheimer's patients using the ICT4LIFE platformabstractThe analysis of multimodal data collected by innovative imaging sensors, Internet of Things (IoT) devices and user interactions, can provide smart and automatic distant monitoring of patients and reveal valuable insights for early detection and/or prevention of events related to their health situation. In this paper, we present a platform called ICT4LIFE which starting from low-level data capturing and performing multimodal fusion to extract relevant features, can perform high-level reasoning to provide relevant data on monitoring and evolution of the patient, and trigger proper actions for improving the quality of life of the patient. Federico Alvarez, Mirela Popa, Nicholas Vretos, Alberto Belmonte-Hernandez, Stylianos Asteriadis, Vassilios Solachidis, Triana Mariscal, Dario Dotti, Petros Daras |
AVSS | 2 |
| 2017 | ICT4Life open source libraries supporting multimodal analysis of different diseasesabstractThe ICT4Life Open Source framework contains libraries for acquiring and processing data from different sensors, machine learning algorithms for activity recognition, as well as fusion methods of multiple modalities either at an early or at a late stage. The main purpose of the introduced system is to enable an easy customization of patients' monitoring using different types of sensors. Furthermore, by allowing an easy integration of new sensors or types of activities, the proposed subsystem supports the development of new solutions for different diseases, than the ones considered in the ICT4Life p roject. Thomas Theodoridis, Vassilios Solachidis, Nicholas Vretos, Petros Daras, Dario Dotti, Mirela Popa, Gustavo Hernández, Federico Alvarez, Alejandro Gonzalez Paton, Angel Lopez |
AVSS | 6 |