EDBT 2026 Demo / reviewers in the wild / expert
Bojian Hou
dblp:305/4170
· DBLP profile ↗
3ranked-venue papers in the field
0as first author
3since 2021 · last 2025
0000-0002-3894-4547ORCID · corroborated
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 2Data Mining & Knowledge Discovery · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | IRIS: Interpretable Risk Clustering Intelligence for Survival AnalysisabstractSurvival analysis models have evolved significantly with deep learning approaches, yet often lack interpretability and meaningful risk stratification capabilities. We present Interpretable Risk Clustering Intelligence for Survival Analysis (IRIS), a novel framework that addresses the critical task of risk clustering while enhancing both input-level and model-body interpretability. Unlike traditional survival models that perform post-hoc risk clustering, IRIS learns to cluster patients into meaningful risk groups directly from data while providing transparent feature importance estimation through feature contribution functions. We validate IRIS on several benchmark datasets, a real-world Alzheimer's disease dataset, and an electronic health record dataset, showing superior performance in risk clustering and predictive reliability with only a modest decrease in time-to-event prediction accuracy compared to state-of-the-art methods. Our results show that IRIS successfully balances the trade-off between interpretability and prediction performance in risk-based survival analysis, offering clinicians actionable insights for treatment planning and resource allocation. Kazi Noshin, Bojian Hou, Mary Regina Boland, Zixuan Wen, Boning Tong, Li Shen 0001, Aidong Zhang 0001 |
IEEE Big Data | 2 |
| 2025 | MentalChat16K: A Benchmark Dataset for Conversational Mental Health AssistanceabstractWe introduce MentalChat16K, an English benchmark dataset combining a synthetic mental health counseling dataset and a dataset of anonymized transcripts from interventions between Behavioral Health Coaches and Caregivers of patients in palliative or hospice care. Covering a diverse range of conditions like depression, anxiety, and grief, this curated dataset is designed to facilitate the development and evaluation of large language models for conversational mental health assistance. By providing a high-quality resource tailored to this critical domain, MentalChat16K aims to advance research on empathetic, personalized AI solutions to improve access to mental health support services. The dataset prioritizes patient privacy, ethical considerations, and responsible data usage. MentalChat16K presents a valuable opportunity for the research community to innovate AI technologies that can positively impact mental well-being. The dataset is available at https://huggingface.co/datasets/ShenLab/MentalChat16K and the code and documentation are hosted on GitHub at https://github.com/PennShenLab/MentalChat16K. Tianyi Wei, Bojian Hou, Patryk Orzechowski, Shu Yang 0009, Ruochen Jin, Rachael Paulbeck, Joost B. Wagenaar, George Demiris, Li Shen 0001 |
KDD (2) | 3 |
| 2024 | SEFD: Semantic-Enhanced Framework for Detecting LLM-Generated Textabstracttechniques that often evade existing detection methods. To address this challenge, we present a novel semantic-enhanced framework for detecting LLM-generated text (SEFD) that leverages a retrieval-based mechanism to fully utilize text semantics. Our framework improves upon existing detection methods by systematically integrating retrieval-based techniques with traditional detectors, employing a carefully curated retrieval mechanism that strikes a balance between comprehensive coverage and computational efficiency. We showcase the effectiveness of our approach in sequential text scenarios common in real-world applications, such as online forums and Q&A platforms. Through comprehensive experiments across various LLM-generated texts and detection methods, we demonstrate that our framework substantially enhances detection accuracy in paraphrasing scenarios while maintaining robustness for standard LLM-generated content. This work contributes significantly to ongoing efforts to safeguard information integrity in an era where AI-generated content is increasingly prevalent. Weiqing He, Bojian Hou, Tianqi Shang, D. Ataee Tarzanagh, Qi Long, Li Shen 0001 |
IEEE Big Data | 2 |