EDBT 2026 Demo / reviewers in the wild / expert
Kilichbek Haydarov
dblp:259/1409
· DBLP profile ↗
7ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0002-3062-2228ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Vision and language · 41% Generative modeling · 15% Language models and text generation · 10% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% |
Topics — the 11 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
image captioning |
1.1 | 2 | 2022 | It is Okay to Not Be Okay: Overcoming Emotional Bias in Affective Image Captioning by Contrastive Data Collection · CVPR 2022 ArtEmis: Affective Language for Visual Art · CVPR 2021 |
Computer vision › Vision and language
affective reasoning |
0.8 | 1 | 2024 | Affective Visual Dialog: A Large-Scale Benchmark for Emotional Reasoning Based on Visually Grounded Conversations · ECCV (75) 2024 |
Computer vision › Vision and language › image captioning
cross-lingual image captioning |
0.8 | 1 | 2024 | No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 Languages · EMNLP 2024 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.8 | 1 | 2024 | Adversarial Text to Continuous Image Generation · CVPR 2024 |
Natural language and speech › Question answering and dialogue systems
visual dialog |
0.8 | 1 | 2024 | Affective Visual Dialog: A Large-Scale Benchmark for Emotional Reasoning Based on Visually Grounded Conversations · ECCV (75) 2024 |
Robotics › Motion planning and robot control › robot control › contact control › contact task control › robot force control
adaptive impedance control |
0.6 | 1 | 2022 | It is Okay to Not Be Okay: Overcoming Emotional Bias in Affective Image Captioning by Contrastive Data Collection · CVPR 2022 |
Machine learning › Trustworthy machine learning
fairness |
0.6 | 1 | 2022 | It is Okay to Not Be Okay: Overcoming Emotional Bias in Affective Image Captioning by Contrastive Data Collection · CVPR 2022 |
Natural language and speech › Information extraction and text analysis
emotion attribution |
0.5 | 1 | 2021 | ArtEmis: Affective Language for Visual Art · CVPR 2021 |
Natural language and speech › Language models and text generation › natural language understanding
emotion understanding |
0.5 | 1 | 2021 | ArtEmis: Affective Language for Visual Art · CVPR 2021 |
Computer vision › Vision and language
art image understanding |
0.4 | 2 | 2024 | No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 Languages · EMNLP 2024 ArtEmis: Affective Language for Visual Art · CVPR 2021 |
Natural language and speech › Information extraction and text analysis
emotion recognition |
0.2 | 1 | 2024 | Affective Visual Dialog: A Large-Scale Benchmark for Emotional Reasoning Based on Visually Grounded Conversations · ECCV (75) 2024 |
Methods — techniques the papers use, named apart from their topics
multimodal emotion recognition · 1.7large language model · 1.7emotional dissonance computation · 1.7word-level attention · 0.8multilingual annotation · 0.8hypernetwork weight modulation · 0.8benchmark construction · 0.8contrastive data collection · 0.6METEOR · 0.6CIDEr · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards AI-Assisted Psychotherapy: Emotion-Guided Generative InterventionsabstractLarge language models (LLMs) hold promise for therapeutic interventions, yet most existing datasets rely solely on text, overlooking nonverbal emotional cues essential to real-world therapy.To address this, we introduce a multimodal dataset of 1,441 publicly sourced therapy session videos containing both dialogue and non-verbal signals such as facial expressions and vocal tone.Inspired by Hochschild's concept of emotional labor, we propose a computational formulation of emotional dissonance-the mismatch between facial and vocal emotion-and use it to guide emotionally aware prompting.Our experiments show that integrating multimodal cues, especially dissonance, improves the quality of generated interventions.We also find that LLM-based evaluators misalign with expert assessments in this domain, highlighting the need for humancentered evaluation.Project page link: https: //kilichbek.github.io/webpage/mental/ Kilichbek Haydarov, Youssef Mohamed, Emilio Goldenhersch, Paul OCallaghan, Li-jia Li |
EMNLP | 1 |
| 2024 | Adversarial Text to Continuous Image GenerationabstractExisting GAN-based text-to-image models treat images as 2D pixel arrays. In this paper, we approach the text-to-image task from a different perspective, where a 2D image is represented as an implicit neural representation (INR). We show that straightforward conditioning of the unconditional INR-based GAN method on text inputs is not enough to achieve good performance. We propose a word-level attention-based weight modulation operator that controls the generation process of INR-GAN based on hypernetworks. Our experiments on benchmark datasets show that HyperCGAN achieves competitive performance to existing pixel-based methods and retains the properties of continuous generative models. Project page link: https://kilichbek.github.io/webpagelhypercgan. Kilichbek Haydarov, Aashiq Muhamed, Xiaoqian Shen, Jovana Lazarevic, Ivan Skorokhodov, Chamuditha Jayanga Galappaththige, Mohamed Elhoseiny 0001 |
CVPR | 1 |
| 2024 | Affective Visual Dialog: A Large-Scale Benchmark for Emotional Reasoning Based on Visually Grounded Conversations
Kilichbek Haydarov, Xiaoqian Shen, Avinash Madasu, Mahmoud Salem, Gamaleldin Elsayed, Mohamed Elhoseiny 0001 |
ECCV (75) | 1 |
| 2024 | No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 LanguagesabstractYoussef Mohamed, Runjia Li, Ibrahim Said Ahmad, Kilichbek Haydarov, Philip Torr, Kenneth Church, Mohamed Elhoseiny. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Youssef Mohamed, Runjia Li, Ibrahim Said Ahmad, Kilichbek Haydarov, Philip Torr 0001, Kenneth Church 0001, Mohamed Elhoseiny 0001 |
EMNLP | 4 |
| 2022 | It is Okay to Not Be Okay: Overcoming Emotional Bias in Affective Image Captioning by Contrastive Data CollectionabstractDatasets that capture the connection between vision, language, and affection are limited, causing a lack of understanding of the emotional aspect of human intelligence. As a step in this direction, the ArtEmis dataset was recently introduced as a large-scale dataset of emotional reactions to images along with language explanations of these chosen emotions. We observed a significant emotional bias towards instance-rich emotions, making trained neural speakers less accurate in describing under-represented emotions. We show that collecting new data, in the same way, is not effective in mitigating this emotional bias. To remedy this problem, we propose a contrastive data collection approach to balance ArtEmis with a new complementary dataset such that a pair of similar images have contrasting emotions (one positive and one negative). We collected 260,533 instances using the proposed method, we combine them with ArtEmis, creating a second iteration of the dataset. The new combined dataset, dubbed ArtEmis v2.0, has a balanced distribution of emotions with explanations revealing more fine details in the associated painting. Our experiments show that neural speakers trained on the new dataset improve CIDEr and METEOR evaluation metrics by 20% and 7%, respectively, compared to the biased dataset. Finally, we also show that the performance per emotion of neural speakers is improved across all the emotion categories, significantly on under-represented emotions. The collected dataset and code are available at https://artemisdataset-v2.org. Youssef Mohamed, Faizan Farooq Khan, Kilichbek Haydarov, Mohamed Elhoseiny 0001 |
CVPR | 3 |
| 2021 | ArtEmis: Affective Language for Visual ArtabstractWe present a novel large-scale dataset and accompanying machine learning models aimed at providing a detailed understanding of the interplay between visual content, its emotional effect, and explanations for the latter in language. In contrast to most existing annotation datasets in computer vision, we focus on the affective experience triggered by visual artworks and ask the annotators to indicate the dominant emotion they feel for a given image and, crucially, to also provide a grounded verbal explanation for their emotion choice. As we demonstrate below, this leads to a rich set of signals for both the objective content and the affective impact of an image, creating associations with abstract concepts (e.g., "freedom" or "love"), or references that go beyond what is directly visible, including visual similes and metaphors, or subjective references to personal experiences. We focus on visual art (e.g., paintings, artistic photographs) as it is a prime example of imagery created to elicit emotional responses from its viewers. Our dataset, termed ArtEmis, contains 455K emotion attributions and explanations from humans, on 80K artworks from WikiArt. Building on this data, we train and demonstrate a series of captioning systems capable of expressing and explaining emotions from visual stimuli. Remarkably, the captions produced by these systems often succeed in reflecting the semantic and abstract content of the image, going well beyond systems trained on existing datasets. The collected dataset and developed methods are available at https://artemisdataset.org. Panos Achlioptas, Maks Ovsjanikov, Kilichbek Haydarov, Mohamed Elhoseiny 0001, Leonidas J. Guibas |
CVPR | 3 |
| 2020 | One-Shot Learning for Surveillance Anomaly Recognition using Siamese 3D CNNabstractOne-shot image recognition has been explored for many applications in computer vision community. However, its applications in video analytics is not deeply investigated yet. For instance, surveillance anomaly recognition is an open challenging problem and one of its hurdles is the lack of accurate temporally annotated data. This paper addresses the lack of data issue using one-shot learning strategy and proposes an anomaly recognition framework which exploits a 3D CNN siamese network that yields the similarity between two anomaly sequences. This paper also investigates the existing 3D CNNs for this task and then proposes a lightweight 3D CNN model that efficiently handles one-shot anomaly recognition. Once our network is trained, then we can use the powerful discriminative 3D CNN features to predict anomalies not only for the new data but also for entirely new classes. The proposed model is trained using temporally annotated test set of UCF Crime dataset. Finally, the trained model is used to recognize the anomalies and produce temporal automatic labels for the video level weakly annotated training set of the dataset. Amin Ullah, Khan Muhammad 0001, Kilichbek Haydarov, Ijaz Ul Haq, Mi Young Lee, Sung Wook Baik |
IJCNN | 3 |