VLDB 2026 Research / reviewers in the wild / expert
Tiberiu Sosea
dblp:278/8001
· DBLP profile ↗
11ranked-venue papers
6as first author
10since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On the Calibration of Image Semi-Supervised Learning ModelsabstractSemi-supervised learning (SSL) has demonstrated high performance in image classification tasks by effectively utilizing both labeled and unlabeled data. However, existing SSL methods often suffer from poor calibration, with models yielding overconfident predictions that misrepresent actual prediction likelihoods. Recently, neural networks trained with mixup that linearly interpolates random examples from the training set have shown better calibration in supervised settings. However, calibration of neural models remains under-explored in SSL settings. Although effective in supervised model calibration, random mixup of pseudolabels in SSL presents challenges due to the overconfidence and unreliability of pseudolabels. In this work, we introduce CalibrateMix, a targeted mixup-based approach that aims to improve the calibration of SSL models while maintaining or even improving their classification accuracy. Our method leverages training dynamics of labeled and unlabeled samples to identify ''easy-to-learn'' and ''hard-to-learn'' samples, which in turn are utilized in a targeted mixup of easy and hard samples. Experimental results across several benchmark datasets show that our method achieves lower expected calibration error (ECE) and superior accuracy compared to existing SSL approaches. Mehrab Mustafy Rahman, Jayanth Mohan, Tiberiu Sosea, Cornelia Caragea |
AAAI | 3 |
| 2024 | GunStance: Stance Detection for Gun Control and Gun RegulationabstractNikesh Gyawali, Iustin Sirbu, Tiberiu Sosea, Sarthak Khanal, Doina Caragea, Traian Rebedea, Cornelia Caragea. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Nikesh Gyawali, Iustin Sirbu, Tiberiu Sosea, Sarthak Khanal, Doina Caragea, Traian Rebedea, Cornelia Caragea |
ACL (1) | 3 |
| 2024 | Sarcasm Detection in a Disaster ContextabstractDuring natural disasters, people often use social media platforms such as Twitter to ask for help, to provide information about the disaster situation, or to express contempt about the unfolding event or public policies and guidelines. This contempt is in some cases expressed as sarcasm or irony. Understanding this form of speech in a disaster-centric context is essential to improving natural language understanding of disaster-related tweets. In this paper, we introduce HurricaneSARC, a dataset of 15,000 tweets annotated for intended sarcasm, and provide a comprehensive investigation of sarcasm detection using pre-trained language models. Our best model is able to obtain as much as 0.70 F1 on our dataset. We also demonstrate that the performance on HurricaneSARC can be improved by leveraging intermediate task transfer learning Tiberiu Sosea, Junyi Jessy Li, Cornelia Caragea |
LREC/COLING | 1 |
| 2023 | Label Smoothing for Emotion Detection (Student Abstract)abstractAutomatically detecting emotions from text has countless applications, ranging from large scale opinion mining to social robots in healthcare and education. However, emotions are subjective in nature and are often expressed in ambiguous ways. At the same time, detecting emotions can also require implicit reasoning, which may not be available as surface- level, lexical information. In this work, we conjecture that the overconfidence of pre-trained language models such as BERT is a critical problem in emotion detection and show that alleviating this problem can considerably improve the generalization performance. We carry out comprehensive experiments on four emotion detection benchmark datasets and show that calibrating our model predictions leads to an average improvement of 1.35% in weighted F1 score. George Maratos, Tiberiu Sosea, Cornelia Caragea |
AAAI | 2 |
| 2023 | Unsupervised Extractive Summarization of Emotion TriggersabstractUnderstanding what leads to emotions during large-scale crises is important as it can provide groundings for expressed emotions and subsequently improve the understanding of ongoing disasters.Recent approaches (Zhan et al., 2022) trained supervised models to both detect emotions and explain emotion triggers (events and appraisals) via abstractive summarization.However, obtaining timely and qualitative abstractive summaries is expensive and extremely time-consuming, requiring highlytrained expert annotators.In time-sensitive, high-stake contexts, this can block necessary responses.We instead pursue unsupervised systems that extract triggers from text.First, we introduce COVIDET-EXT, augmenting (Zhan et al., 2022)'s abstractive dataset (in the context of the COVID-19 crisis) with extractive triggers.Second, we develop new unsupervised learning models that can jointly detect emotions and summarize their triggers.Our best approach, entitled Emotion-Aware Pagerank, incorporates emotion information from external sources combined with a language understanding module, and outperforms strong baselines.We release our data and code at Tiberiu Sosea, Hongli Zhan, Junyi Jessy Li, Cornelia Caragea |
ACL (1) | 1 |
| 2023 | MarginMatch: Improving Semi-Supervised Learning with Pseudo-MarginsabstractWe introduce MarginMatch, a new SSL approach combining consistency regularization and pseudo-labeling, with its main novelty arising from the use of unlabeled data training dynamics to measure pseudo-label quality. Instead of using only the model's confidence on an unlabeled example at an arbitrary iteration to decide if the example should be masked or not, MarginMatch also analyzes the behavior of the model on the pseudo-labeled examples as the training progresses, to ensure low quality predictions are masked out. MarginMatch brings substantial improvements on four vision benchmarks in low data regimes and on two large-scale datasets, emphasizing the importance of enforcing high-quality pseudo-labels. Notably, we obtain an improvement in error rate over the state-of-the-art of 3.25% on CIFAR-100 with only 25 labels per class and of 3.78% on STL-10 using as few as 4 labels per class. We make our code available at https://github.com/tsosea2/MarginMatch. Tiberiu Sosea, Cornelia Caragea |
CVPR | 1 |
| 2022 | Multimodal Semi-supervised Learning for Disaster Tweet ClassificationabstractDuring natural disasters, people often use social media platforms, such as Twitter, to post information about casualties and damage produced by disasters. This information can help relief authorities gain situational awareness in nearly real time, and enable them to quickly distribute resources where most needed. However, annotating data for this purpose can be burdensome, subjective and expensive. In this paper, we investigate how to leverage the copious amounts of unlabeled data generated on social media by disaster eyewitnesses and affected individuals during disaster events. To this end, we propose a semi-supervised learning approach to improve the performance of neural models on several multimodal disaster tweet classification tasks. Our approach shows significant improvements, obtaining up to 7.7% improvements in F-1 in low-data regimes and 1.9% when using the entire training data. We make our code and data publicly available at https://github.com/iustinsirbu13/multimodal-ssl-for-disaster-tweet-classification. Iustin Sirbu, Tiberiu Sosea, Cornelia Caragea, Doina Caragea, Traian Rebedea |
COLING | 2 |
| 2022 | Why Do You Feel This Way? Summarizing Triggers of Emotions in Social Media PostsabstractCrises such as the COVID-19 pandemic continuously threaten our world and emotionally affect billions of people worldwide in distinct ways.Understanding the triggers leading to people's emotions is of crucial importance.Social media posts can be a good source of such analysis, yet these texts tend to be charged with multiple emotions, with triggers scattering across multiple sentences.This paper takes a novel angle, namely, emotion detection and trigger summarization, aiming to both detect perceived emotions in text, and summarize events and their appraisals that trigger each emotion.To support this goal, we introduce COVIDET (Emotions and their Triggers during Covid-19), a dataset of ~1, 900 English Reddit posts related to COVID-19, which contains manual annotations of perceived emotions and abstractive summaries of their triggers described in the post.We develop strong baselines to jointly detect emotions and summarize emotion triggers.Our analyses show that COVIDET presents new challenges in emotion-specific summarization, as well as multi-emotion detection in long social media posts.* Hongli Zhan and Tiberiu Sosea contributed equally.Reddit Post 1: My sibling is 19 and she constantly goes places with her friends and to there houses and its honestly stressing me out.2: Our grandfather lives with us and he has dementia along with other health issues and my mom has diabetes and heart problems and I have autoimmune diseases & chronic health issues.3: She also has asthma.4: Its stressing me out because despite this she seems to not care about how badly it would affect all of us if we were to get the virus.5: And sadly I feel like its not much I can do she literally doesn't respect my mom and though I'm older she doesn't respect me either.6: Its so frustrating. Emotions and Abstractive Summaries of TriggersEmotion: anger Abstractive Summary of Trigger: My sister having absolutely no regard for any of our family's health coupled with the fact that I can't do anything about it is so aggravating to me.Emotion: fear Abstractive Summary of Trigger: My sibling, who, in spite of our family's myriad of issues that all make us high-risk people, continuously goes out and about, which makes her likely to get infected.I am scared for all of us right now. Hongli Zhan, Tiberiu Sosea, Cornelia Caragea, Junyi Jessy Li |
EMNLP | 2 |
| 2022 | EnsyNet: A Dataset for Encouragement and Sympathy DetectionabstractMore and more people turn to Online Health Communities to seek social support during their illnesses. By interacting with peers with similar medical conditions, users feel emotionally and socially supported, which in turn leads to better adherence to therapy. Current studies in Online Health Communities focus only on the presence or absence of emotional support, while the available datasets are scarce or limited in terms of size. To enable development on emotional support detection, we introduce EnsyNet, a dataset of 6,500 sentences annotated with two types of support: encouragement and sympathy. We train BERT-based classifiers on this dataset, and apply our best BERT model in two large scale experiments. The results of these experiments show that receiving encouragements or sympathy improves users’ emotional state, while the lack of emotional support negatively impacts patients’ emotional state. Tiberiu Sosea, Cornelia Caragea |
LREC | 1 |
| 2022 | Emotion analysis and detection during COVID-19abstractUnderstanding emotions that people express during large-scale crises helps inform policy makers and first responders about the emotional states of the population as well as provide emotional support to those who need such support. We present CovidEmo, a dataset of ~3,000 English tweets labeled with emotions and temporally distributed across 18 months. Our analyses reveal the emotional toll caused by COVID-19, and changes of the social narrative and associated emotions over time. Motivated by the time-sensitive nature of crises and the cost of large-scale annotation efforts, we examine how well large pre-trained language models generalize across domains and timeline in the task of perceived emotion prediction in the context of COVID-19. Our analyses suggest that cross-domain information transfers occur, yet there are still significant gaps. We propose semi-supervised learning as a way to bridge this gap, obtaining significantly better performance using unlabeled data from the target domain. Tiberiu Sosea, Chau Pham 0003, Alexander Tekle, Cornelia Caragea, Junyi Jessy Li |
LREC | 1 |
| 2020 | CancerEmo: A Dataset for Fine-Grained Emotion DetectionabstractEmotions are an important element of human nature, often affecting the overall wellbeing of a person.Therefore, it is no surprise that the health domain is a valuable area of interest for emotion detection, as it can provide medical staff or caregivers with essential information about patients.However, progress on this task has been hampered by the absence of large labeled datasets.To this end, we introduce CANCEREMO , an emotion dataset created from an online health community and annotated with eight fine-grained emotions.We perform a comprehensive analysis of these emotions and develop deep learning models on the newly created dataset.Our best BERT model achieves an average F1 of 71%, which we improve further using domain-specific pretraining. Tiberiu Sosea, Cornelia Caragea |
EMNLP (1) | 1 |