Adrian Cosma

dblp:232/3326 · DBLP profile ↗
← Back
23ranked-venue papers
9as first author
19since 2021 · last 2026
0000-0003-0307-2520ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 6 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 On Model and Data Scaling for Skeleton-based Self-Supervised Gait Recognition
abstract
Gait recognition from video streams is a challenging problem in computer vision biometrics due to the subtle differences between gaits and numerous confounding factors. Recent advancements in self-supervised pretraining have led to the development of robust gait recognition models that are invariant to walking covariates. While neural scaling laws have transformed model development in other domains by linking performance to data, model size, and compute, their applicability to gait remains unexplored. In this work, we conduct the first empirical study scaling on skeleton-based self-supervised gait recognition to quantify the effect of data quantity, model size and compute on downstream gait recognition performance. We pretrain multiple variants of GaitPT -- a transformer-based architecture -- on a dataset of 2.7 million walking sequences collected in the wild. We evaluate zero-shot performance across four benchmark datasets to derive scaling laws for data, model size, and compute. Our findings demonstrate predictable power-law improvements in performance with increased scale and confirm that data and compute scaling significantly influence downstream accuracy. We further isolate architectural contributions by comparing GaitPT with GaitFormer under controlled compute budgets. These results provide practical insights into resource allocation and performance estimation for real-world gait recognition systems.
Adrian Cosma, Andy Catruna, Emilian Radoi
AAAI1
2026 Database-Agnostic Gait Enrollment using SetTransformers
abstract
Gait recognition has emerged as a powerful tool for unobtrusive and long-range identity analysis, with growing relevance in surveillance and monitoring applications. Although recent advances in deep learning and large-scale datasets have enabled highly accurate recognition under closed-set conditions, real-world deployment demands open-set gait enrollment, which means determining whether a new gait sample corresponds to a known identity or represents a previously unseen individual. In this work, we introduce a transformer-based framework for open-set gait enrollment that is both dataset-agnostic and recognition-architecture-agnostic. Our method leverages a SetTransformer to make enrollment decisions based on the embedding of a probe sample and a context set drawn from the gallery, without requiring task-specific thresholds or retraining for new environments. By decoupling enrollment from the main recognition pipeline, our model is generalized across different datasets, gallery sizes, and identity distributions. We propose an evaluation protocol that uses existing datasets in different ratios of identities and walks per identity. We instantiate our method using skeleton-based gait representations and evaluate it on two benchmark datasets (CASIA-B and PsyMo), using embeddings from three state-of-the-art recognition models (GaitGraph, GaitFormer, and GaitPT). We show that our method is flexible, is able to accurately perform enrollment in different scenarios, and scales better with data compared to traditional approaches. We will make the code and dataset scenarios publicly available.
Nicoleta-Nina Basoc, Adrian Cosma, Andy Catruna, Emilian Radoi
FG2
2026 MoME: Estimating Psychological Traits from Gait with Multi-Stage Mixture of Movement Experts
abstract
Gait encodes rich biometric and behavioural information, yet leveraging the manner of walking to infer psychological traits remains a challenging and underexplored problem. We introduce a hierarchical Multi-Stage Mixture of Movement Experts (MoME) architecture for multi-task prediction of psychological attributes from gait sequences represented as 2D poses. MoME processes the walking cycle in four stages of movement complexity, employing lightweight expert models to extract spatio-temporal features and task-specific gating modules to adaptively weight experts across traits and stages. Evaluated on the PsyMo benchmark covering 17 psychological traits, our method outperforms state-of-the-art gait analysis models, achieving a 37.47% weighted F1 score at the run level and 44.6% at the subject level. Our experiments show that integrating auxiliary tasks such as identity recognition, gender prediction, and BMI estimation further improves psychological trait estimation. Our findings demonstrate the viability of multi-task gait-based learning for psychological trait estimation and provide a foundation for future research on movement-informed psychological inference.
Andy Catruna, Adrian Cosma, Emilian Radoi
FG2
2026 Simplifying Spatio-Temporal Graphs in Skeleton-Based Gait Analysis
Adrian Cosma, Emilian Radoi
FG1
2025 The Strawberry Problem: Emergence of Character-level Understanding in Tokenized Language Models
abstract
Despite their remarkable progress across diverse domains, Large Language Models (LLMs) consistently fail at simple characterlevel tasks, such as counting letters in words, due to a fundamental limitation: tokenization.In this work, we frame this limitation as a problem of low mutual information and analyze it in terms of concept emergence.Using a suite of 19 synthetic tasks that isolate character-level reasoning in a controlled setting, we show that such capabilities emerge suddenly and only late in training.We find that percolation-based models of concept emergence explain these patterns, suggesting that learning character composition is not fundamentally different from learning commonsense knowledge.To address this bottleneck, we propose a lightweight architectural modification that significantly improves character-level reasoning while preserving the inductive advantages of subword models.Together, our results bridge low-level perceptual gaps in tokenized LMs and provide a principled framework for understanding and mitigating their structural blind spots.We make our code publicly available.
Adrian Cosma, Stefan Ruseti, Emilian Radoi, Mihai Dascalu
EMNLP1
2025 The Expression of Happiness in Social Media of Individuals Reporting Depression
abstract
Depression has long been studied in the NLP field, with most works focusing on individuals’ negative emotions. People with depression experience happiness, but this has not been extensively studied. Previous works have shown that sentiment or emotion classification approaches are unsuitable for extracting happy moments because they may not be expressed only in positive words. In this work, we conduct a large-scale study of happy moments from social media texts of individuals mentioning a depression diagnosis. We develop an extensive deep learning-based framework to extract happy moments from text, and annotate them with semantic topics, gender labels, and agency and sociality measures. We analyze over 400,000 happy moments and show significant differences in topics, agency, and sociality of users in the depression and control groups, varying by gender. We found that male and female users in the depression group expressed more sociality in their happy moments than control users. Furthermore, male users’ agency was not impaired in depression, while female users in the depression group expressed fewer happy moments with agency than the control group. Our research can inform psychology interventions, which can foster feelings of longer-lasting happiness and represent a promising path of collaboration between computational linguistics and psychology.
Ana-Maria Bucur, Berta Chulvi, Adrian Cosma, Paolo Rosso
IEEE Trans. Affect. Comput.3
2024 RoCode: A Dataset for Measuring Code Intelligence from Problem Definitions in Romanian
abstract
Recently, large language models (LLMs) have become increasingly powerful and have become capable of solving a plethora of tasks through proper instructions in natural language. However, the vast majority of testing suites assume that the instructions are written in English, the de facto prompting language. Code intelligence and problem solving still remain a difficult task, even for the most advanced LLMs. Currently, there are no datasets to measure the generalization power for code-generation models in a language other than English. In this work, we present RoCode, a competitive programming dataset, consisting of 2,642 problems written in Romanian, 11k solutions in C, C++ and Python and comprehensive testing suites for each problem. The purpose of RoCode is to provide a benchmark for evaluating the code intelligence of language models trained on Romanian / multilingual text as well as a fine-tuning set for pretrained Romanian models. Through our results and review of related works, we argue for the need to develop code models for languages other than English.
Adrian Cosma, Ioan-Bogdan Iordache, Paolo Rosso
LREC/COLING1
2024 Reading Between the Frames: Multi-modal Depression Detection in Videos from Non-verbal Cues
David Gimeno-Gómez, Ana-Maria Bucur, Adrian Cosma, Carlos D. Martínez-Hinarejos, Paolo Rosso
ECIR (1)3
2024 How Hard is this Test Set? NLI Characterization by Exploiting Training Dynamics
abstract
Natural Language Inference (NLI) evaluation is crucial for assessing language understanding models; however, popular datasets suffer from systematic spurious correlations that artificially inflate actual model performance.To address this, we propose a method for the automated creation of a challenging test set without relying on the manual construction of artificial and unrealistic examples.We categorize the test set of popular NLI datasets into three difficulty levels by leveraging methods that exploit training dynamics.This categorization significantly reduces spurious correlation measures, with examples labeled as having the highest difficulty showing markedly decreased performance and encompassing more realistic and diverse linguistic phenomena.When our characterization method is applied to the training set, models trained with only a fraction of the data achieve comparable performance to those trained on the full dataset, surpassing other dataset characterization techniques.Our research addresses limitations in NLI dataset construction, providing a more authentic evaluation of model performance with implications for diverse NLU applications.
Adrian Cosma, Stefan Ruseti, Mihai Dascalu, Cornelia Caragea
EMNLP1
2024 The Paradox of Motion: Evidence for Spurious Correlations in Skeleton-Based Gait Recognition Models
abstract
Gait, an unobtrusive biometric, is valued for its capability to identify individuals at a distance, across external outfits and environmental conditions. This study challenges the prevailing assumption that vision-based gait recognition, in particular skeleton-based gait recognition, relies primarily on motion patterns, revealing a significant role of the implicit anthropometric information encoded in the walking sequence. We show through a comparative analysis that removing height information leads to notable performance degradation across three models and two benchmarks (CASIA-B and GREW). Furthermore, we propose a spatial transformer model processing individual poses, disregarding any temporal information, which achieves unreasonably good accuracy, emphasizing the bias towards appearance information and indicating spurious correlations in existing benchmarks. These findings underscore the need for a nuanced understanding of the interplay between motion and appearance in vision-based gait recognition, prompting a reevaluation of the methodological assumptions in this field. Our experiments indicate that “in-the-wild” datasets are less prone to spurious correlations, prompting the need for more diverse and large scale datasets for advancing the field.
Andy Catruna, Adrian Cosma, Emilian Radoi
FG2
2024 CrossGaze: A Strong Method for 3D Gaze Estimation in the Wild
abstract
Gaze estimation, the task of predicting where an individual is looking, is a critical task with direct applications in areas such as human-computer interaction and virtual reality. Estimating the direction of looking in unconstrained environments is difficult, due to the many factors that can obscure the face and eye regions. In this work we propose CrossGaze, a strong baseline for gaze estimation, that leverages recent developments in computer vision architectures and attention-based modules. Unlike previous approaches, our method does not require a specialized architecture, utilizing already established models that we integrate in our architecture and adapt for the task of 3D gaze estimation. This approach allows for seamless updates to the architecture as any module can be replaced with more powerful feature extractors. On the Gaze360 benchmark, our model surpasses several state-of-the-art methods, achieving a mean angular error of 9.94°. Our proposed model serves as a strong foundation for future research and development in gaze estimation, paving the way for practical and accurate gaze prediction in real-world scenarios. The code is available at: https://github.com/AndyCatruna/CrossGaze.
Andy Catruna, Adrian Cosma, Emilian Radoi
FG2
2024 GaitPT: Skeletons are All You Need for Gait Recognition
abstract
The analysis of patterns of walking is an important area of research that has numerous applications in security, healthcare, sports and human-computer interaction. Lately, walking patterns have been regarded as a unique fingerprinting method for automatic person identification at a distance. In this work, we propose a novel gait recognition architecture called Gait Pyramid Transformer (GaitPT) that leverages pose estimation skeletons to capture unique walking patterns, without relying on appearance information. GaitPT adopts a hierarchical transformer architecture that effectively extracts both spatial and temporal features of movement in an anatomically consistent manner, guided by the structure of the human skeleton. Our results show that GaitPT achieves state-of-the-art performance compared to other skeleton-based gait recognition works, in both controlled and in-the-wild scenarios. GaitPT obtains 82.6% average accuracy on CASIA-B, surpassing other works by a margin of 6%. Moreover, it obtains 52.16% Rank-1 accuracy on GREW, outperforming both skeleton-based and appearance-based approaches. The code is available at: https://github.com/AndyCatruna/GaitPT.
Andy Catruna, Adrian Cosma, Emilian Radoi
FG2
2024 Aligning Actions and Walking to LLM-Generated Textual Descriptions
abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities in various domains, including data augmentation and synthetic data generation. This work explores the use of LLMs to generate rich textual descriptions for motion sequences, encompassing both actions and walking patterns. We leverage the expressive power of LLMs to align motion representations with high-level linguistic cues, addressing two distinct tasks: action recognition and retrieval of walking sequences based on appearance attributes. For action recognition, we employ LLMs to generate textual descriptions of actions in the BABEL-60 dataset, facilitating the alignment of motion sequences with linguistic representations. In the domain of gait analysis, we investigate the impact of appearance attributes on walking patterns by generating textual descriptions of motion sequences from the DenseGait dataset using LLMs. These descriptions capture subtle variations in walking styles influenced by factors such as clothing choices and footwear. Our approach demonstrates the potential of LLMs in aug-menting structured motion attributes and aligning multimodal representations. The findings contribute to the advancement of comprehensive motion understanding and open up new av-enues for leveraging LLMs in multimodal alignment and data augmentation for motion analysis. We make the code publicly available at https://github.com/Radu1999/WalkAndText
Radu Chivereanu, Adrian Cosma, Andy Catruna, Razvan Rughinis, Emilian Radoi
FG2
2024 Gait Recognition from Highly Compressed Videos
abstract
Surveillance footage represents a valuable resource and opportunities for conducting gait analysis. However, the typical low quality and high noise levels in such footage can severely impact the accuracy of pose estimation algorithms, which are foundational for reliable gait analysis. Existing literature suggests a direct correlation between the efficacy of pose estimation and the subsequent gait analysis results. A common mitigation strategy involves fine-tuning pose estimation models on noisy data to improve robustness. However, this approach may degrade the downstream model's performance on the original high-quality data, leading to a trade-off that is undesirable in practice. We propose a processing pipeline that incorporates a task-targeted artifact correction model specifically designed to pre-process and enhance surveillance footage before pose estimation. Our artifact correction model is optimized to work alongside a state-of-the-art pose estimation network, HRNet, without requiring repeated fine-tuning of the pose estimation model. Furthermore, we propose a simple and robust method for obtaining low quality videos that are annotated with poses in an automatic manner with the purpose of training the artifact correction model. We systematically evaluate the performance of our artifact correction model against a range of noisy surveillance data and demonstrate that our approach not only achieves improved pose estimation on low-quality surveillance footage, but also preserves the integrity of the pose estimation on high resolution footage. Our experiments show a clear enhancement in gait analysis performance, supporting the viability of the proposed method as a superior alternative to direct fine-tuning strategies. Our contributions pave the way for more reliable gait analysis using surveillance data in real-world applications, regardless of data quality.
Andrei Niculae, Andy Catruna, Adrian Cosma, Daniel Rosner, Emilian Radoi
FG3
2024 PsyMo: A Dataset for Estimating Self-Reported Psychological Traits from Gait
abstract
Psychological trait estimation from external factors such as movement and appearance is a challenging and longstanding problem in psychology, and is principally based on the psychological theory of embodiment. To date, attempts to tackle this problem have utilized private small-scale datasets with intrusive body-attached sensors. Potential applications of an automated system for psychological trait estimation include estimation of occupational fatigue and psychology, and marketing and advertisement. In this work, we propose PsyMo (Psychological traits from Motion), a novel, multi-purpose and multi-modal dataset for exploring psychological cues manifested in walking patterns. We gathered walking sequences from 312 subjects in 7 different walking variations and 6 camera angles. In conjunction with walking sequences, participants filled in 6 psychological questionnaires, totaling 17 psychometric attributes related to personality, self-esteem, fatigue, aggressiveness and mental health. We propose two evaluation protocols for psychological trait estimation. Alongside the estimation of self-reported psychological traits from gait, the dataset can be used as a drop-in replacement to benchmark methods for gait recognition. We anonymize all cues related to the identity of the subjects and publicly release only silhouettes, 2D / 3D human skeletons and 3D SMPL human meshes.
Adrian Cosma, Emilian Radoi
WACV1
2023 It's Just a Matter of Time: Detecting Depression with Time-Enriched Multimodal Transformers
Ana-Maria Bucur, Adrian Cosma, Paolo Rosso, Liviu P. Dinu
ECIR (1)2
2023 GaitMorph: Transforming Gait by Optimally Transporting Discrete Codes
abstract
Gait, the manner of walking, has been proven to be a reliable biometric with uses in surveillance, marketing and security. A promising new direction for the field is training gait recognition systems without explicit human annotations, through self-supervised learning approaches. Such methods are heavily reliant on strong augmentations for the same walking sequence to induce more data variability and to simulate additional walking variations. Current data augmentation schemes are heuristic and cannot provide the necessary data variation as they are only able to provide simple temporal and spatial distortions. In this work, we propose GaitMorph, a novel method to modify the walking variation for an input gait sequence. Our method entails the training of a high-compression model for gait skeleton sequences that leverages unlabelled data to construct a discrete and interpretable latent space, which preserves identity-related features. Furthermore, we propose a method based on optimal transport theory to learn latent transport maps on the discrete codebook that morph gait sequences between variations. We perform extensive experiments and show that our method is suitable to synthesize additional views for an input sequence.
Adrian Cosma, Emilian Radoi
IJCB1
2022 Life is not Always Depressing: Exploring the Happy Moments of People Diagnosed with Depression
abstract
In this work, we explore the relationship between depression and manifestations of happiness in social media. While the majority of works surrounding depression focus on symptoms, psychological research shows that there is a strong link between seeking happiness and being diagnosed with depression. We make use of Positive-Unlabeled learning paradigm to automatically extract happy moments from social media posts of both controls and users diagnosed with depression, and qualitatively analyze them with linguistic tools such as LIWC and keyness information. We show that the life of depressed individuals is not always bleak, with positive events related to friends and family being more noteworthy to their lives compared to the more mundane happy events reported by control users.
Ana-Maria Bucur, Adrian Cosma, Liviu P. Dinu
LREC2
2021 From Face to Gait: Weakly-Supervised Learning of Gender Information from Walking Patterns
abstract
Obtaining demographics information from video is valuable for a range of real-world applications. While approaches that leverage facial features for gender inference are very successful in restrained environments, they do not work in most real-world scenarios when the subject is not facing the camera, has the face obstructed or the face is not clear due to distance from the camera or poor resolution. We propose a weakly-supervised method for learning gender information of people based on their manner of walking. We make use of state-of-the art facial analysis models to automatically annotate front- view walking sequences and generalise to unseen angles by leveraging gait-based label propagation. Our results show on par or higher performance with facial analysis models with an F1 score of 91 % and the ability to successfully generalise to scenarios in which facial analysis is unfeasible due to subjects not facing the camera or having the face obstructed.
Andy Catruna, Adrian Cosma, Emilian Radoi
FG2
2020 Self-supervised Representation Learning on Document Images
Adrian Cosma, Mihai Ghidoveanu, Michael Panaitescu-Liess, Marius Popescu
DAS1
2020 Black-Box Ripper: Copying black-box models using generative evolutionary algorithms
abstract
We study the task of replicating the functionality of black-box neural models, for which we only know the output class probabilities provided for a set of input images. We assume back-propagation through the black-box model is not possible and its training images are not available, e.g. the model could be exposed only through an API. In this context, we present a teacher-student framework that can distill the black-box (teacher) model into a student model with minimal accuracy loss. To generate useful data samples for training the student, our framework (i) learns to generate images on a proxy data set (with images and classes different from those used to train the black-box) and (ii) applies an evolutionary strategy to make sure that each generated data sample exhibits a high response for a specific class when given as input to the black box. Our framework is compared with several baseline and state-of-the-art methods on three benchmark data sets. The empirical evidence indicates that our model is superior to the considered baselines. Although our method does not back-propagate through the black-box network, it generally surpasses state-of-the-art methods that regard the teacher as a glass-box model. Our code is available at: https://github.com/antoniobarbalau/black-box-ripper.
Antonio Barbalau, Adrian Cosma, Radu Tudor Ionescu, Marius Popescu
NeurIPS2
2020 A Generic and Model-Agnostic Exemplar Synthetization Framework for Explainable AI
Antonio Barbalau, Adrian Cosma, Radu Tudor Ionescu, Marius Popescu
ECML/PKDD (2)2
2019 CamLoc: Pedestrian Location Estimation through Body Pose Estimation on Smart Cameras
abstract
Advances in hardware and algorithms are driving the exponential growth of Internet of Things (IoT), with increasingly more pervasive computations being performed near the data generation sources. With this wave of technology, a range of intelligent devices can perform local inferences (activity recognition, fitness monitoring, etc.), which have obvious advantages: reduced inference latency for interactive (real-time) applications and better data privacy by processing user data locally. Video processing can benefit many applications and data labelling systems, although performing this efficiently at the edge of the Internet is not trivial. In this paper, we show that accurate pedestrian location estimation is achievable using deep neural networks on fixed cameras with limited computing resources. Our approach, CamLoc, uses pose estimation from key body points detection to extend pedestrian skeleton when the entire body is not in view (occluded by obstacles or partially outside the frame). Our evaluation dataset contains over 2100 frames from surveillance cameras (including two cameras simultaneously pointing at the same scene from different angles), in 42 different scenarios of activity and occlusion. We make this dataset available together with annotations indicating the exact 2D position of person in frame as ground-truth information. CamLoc achieves good location estimation accuracy in these complex scenarios with high levels of occlusion, matching the performance of state-of-the-art solutions, but using less computing resources and attaining a higher inference throughput.
Adrian Cosma, Emilian Radoi, Valentin Radu
IPIN1