Lika Lomidze

dblp:436/9756 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Trustworthy machine learning · 100%
Human-computer interaction and pervasive computing
1 paper
Health and well-being technologies · 100%

Topics — the 2 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
safety evaluation
1.012026
Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations · ACL (1) 2026
Health and well-being technologies
mental health technology
0.312026
Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

persona simulation · 2.0emotion modeling · 2.0LLM-assisted classification · 2.0
YearPublicationVenuePosition
2026 Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations
abstract
There are growing concerns about the risks posed by AI companion applications designed for emotional engagement.Existing safety evaluations often rely on self-reported user data or interviews, offering limited insights into real-time dynamics.We present the first end-to-end scalable framework for controlled simulation and safety evaluation of multi-turn interactions with AI companion applications.Our framework integrates four key components: persona construction with clinical and psychometric validation, persona-specific scenario generation, scenario-driven multi-turn simulation with a dialogue refinement module that preserves persona fidelity, and harm evaluation.We apply this framework to evaluate how Replika, a widely used AI companion app, responds to high-risk user groups.We construct 9 personas representing individuals with depression, anxiety, PTSD, eating disorders, and incel identity, and collect 1,674 dialogue pairs across 25 high-risk scenarios.We combine emotion modeling and LLM-assisted utterance-and harm-level classification to analyze these exchanges.Results show that Replika exhibits a narrow emotional range dominated by curiosity and care, while frequently mirroring or normalizing unsafe content such as self-harm, disordered eating, and violentfantasy narratives.These findings highlight how controlled persona simulations can serve as a scalable testbed for evaluating safety risks in AI companions. 1 Content Warning: This paper includes examples of dialogues involving self-harm, disordered eating, and misogynistic language.
Prerna Juneja, Lika Lomidze
ACL (1)2