EDBT 2026 Demo / reviewers in the wild / expert
Eliseo Bao Souto
dblp:359/3295 · also Eliseo Bao
· DBLP profile ↗
4ranked-venue papers
4as first author
4since 2021 · last 2026
0009-0000-8457-1115ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Evidence of Depression Symptoms via Prompt InductionabstractDepression places substantial pressure on mental health services, and many people describe their experiences outside clinical settings in high-volume user-generated text (e.g., online forums and social media). Automatically identifying clinical symptom evidence in such text can therefore complement limited clinical capacity and scale to large populations. We address this need through sentence-level classification of 21 depression symptoms from the BDI-II questionnaire, using BDI-Sen, a dataset annotated for symptom relevance. This task is fine-grained and highly imbalanced, and we find that common LLM approaches (zero-shot, in-context learning, and fine-tuning) struggle to apply consistent relevance criteria for most symptoms. We propose Symptom Induction (SI), a novel approach which compresses labeled examples into short, interpretable guidelines that specify what counts as evidence for each symptom and uses these guidelines to condition classification. Across four LLM families and eight models, SI achieves the best overall weighted F1 on BDI-Sen, with especially large gains for infrequent symptoms. Cross-domain evaluation on an external dataset further shows that induced guidelines generalize across other diseases shared symptomatology (bipolar and eating disorders). Eliseo Bao Souto, Anxo Pérez, David Otero 0001, Javier Parapar |
SIGIR | 1 |
| 2025 | ReDSM5: A Reddit Dataset for DSM-5 Depression DetectionabstractDepression is a pervasive mental health condition that affects hundreds of millions of individuals worldwide, yet many cases remain undiagnosed due to barriers in traditional clinical access and pervasive stigma. Social media platforms, and Reddit in particular, offer rich, user-generated narratives that can reveal early signs of depressive symptomatology. However, existing computational approaches often label entire posts simply as depressed or not depressed, without linking language to specific criteria from the DSM-5, the standard clinical framework for diagnosing depression. This limits both clinical relevance and interpretability. To address this gap, we introduce ReDSM5, a novel Reddit corpus comprising 1484 long-form posts, each exhaustively annotated at the sentence level by a licensed psychologist for the nine DSM-5 depression symptoms. For each label, the annotator also provides a concise clinical rationale grounded in DSM-5 methodology. We conduct an exploratory analysis of the collection, examining lexical, syntactic, and emotional patterns that characterize symptom expression in social media narratives. Compared to prior resources, ReDSM5 uniquely combines symptom-specific supervision with expert explanations, facilitating the development of models that not only detect depression but also generate human-interpretable reasoning. We establish baseline benchmarks for both multi-label symptom classification and explanation generation, providing reference results for future research on detection and interpretability. Eliseo Bao Souto, Anxo Pérez, Javier Parapar |
CIKM | 1 |
| 2025 | MindWell: A Conversational Agent for Professional Depression Screening on Social Media
Eliseo Bao Souto, Anxo Pérez, Javier Parapar |
ECIR (5) | 1 |
| 2025 | Towards Explainable and Safe Systems for Health DataabstractThe integration of artificial intelligence into healthcare holds immense potential to transform medical services through personalized care, early diagnosis, and more informed decision-making. However, despite significant advances in digital health technologies, the adoption of AI-driven systems remains limited, particularly in sensitive domains such as mental health, due to concerns around transparency, safety, and user trust. These concerns are compounded by the complexity of clinical language, the high stakes of decision-making, and the need for accountability in automated systems. This research addresses these challenges by investigating how Large Language Models (LLMs) and related techniques can be adapted to create explainable, interpretable, and safe health information systems. Positioned at the intersection of health informatics, Natural Language Processing (NLP), and Information Retrieval (IR), this work explores both the potential and limitations of LLMs in clinical contexts. While LLMs are effective at processing large textual datasets and producing fluent responses, their black-box nature, tendency to hallucinate, and associated privacy risks make them unsuitable for direct application in healthcare without significant adaptation. Eliseo Bao Souto |
SIGIR | 1 |