EDBT 2026 Demo / reviewers in the wild / expert
Vicky Kalogeiton
dblp:157/8338 · also Vicky S. Kalogeiton
· DBLP profile ↗
30ranked-venue papers
6as first author
21since 2021 · last 2026
0000-0002-7368-6993ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 3 first-author · 16 since 2021Artificial intelligence and machine learning · 20 · 4 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorSecurity and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Audio-Visual Generation
Tae-Hyun Oh, Vicky Kalogeiton, Stavros Petridis, Sergey Tulyakov, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 2 |
| 2025 | Around the World in 80 Timesteps: A Generative Approach to Global Visual GeolocationabstractGlobal visual geolocation consists in predicting where an image was captured anywhere on Earth. Since not all images can be localized with the same precision, this task inherently involves a degree of ambiguity. However, existing approaches are deterministic and overlook this aspect. In this paper, we propose the first generative approach for visual geolocation based on diffusion and flow matching, and an extension to Riemannian flow matching, where the denoising process operates directly on the Earth's surface. Our model achieves state-of-the-art performance on three visual geolocation benchmarks: OpenStreetView-5M, YFCC100M, and iNat21. In addition, we introduce the task of probabilistic visual geolocation, where the model predicts a probability distribution over all possible locations instead of a single point. We implement new metrics and baselines for this task, demonstrating the advantages of our generative approach. Codes and models are available here. Nicolas Dufour, Vicky Kalogeiton, David Picard, Loïc Landrieu |
CVPR | 2 |
| 2025 | AKiRa: Augmentation Kit on Rays for Optical Video GenerationabstractRecent advances in text-conditioned video diffusion have greatly improved video quality. However, these methods offer limited or sometimes no control to users on camera aspects, including dynamic camera motion, zoom, distorted lens and focus shifts. These motion and optical aspects are crucial for adding controllability and cinematic elements to generation frameworks, ultimately resulting in visual content that draws focus, enhances mood, and guides emotions according to filmmakers’ controls. In this paper, we aim to close the gap between controllable video generation and camera optics. To achieve this, we propose AKiRa (Augmentation Kit on Rays), a novel augmentation framework that builds and trains a camera adapter with a complex camera model over an existing video generation backbone. It enables fine-tuned control over camera motion as well as complex optical parameters (focal length, distortion, aperture) to achieve cinematic effects such as zoom, fish-eye effect, and bokeh. Extensive experiments demonstrate AKiRa’s effectiveness in combining and composing camera optics while outperforming all state-of-the-art methods. This work sets a new landmark in controlled and optically enhanced video generation, paving the way for future optical video generation methods. Xi Wang 0024, Robin Courant, Marc Christie, Vicky Kalogeiton |
CVPR | 4 |
| 2025 | Di[M]O: Distilling Masked Diffusion Models Into One-Step Generator
Yuanzhi Zhu 0001, Xi Wang 0024, Stéphane Lathuilière, Vicky Kalogeiton |
ICCV | 4 |
| 2025 | T-REGS: Minimum Spanning Tree Regularization for Self-Supervised LearningabstractSelf-supervised learning (SSL) has emerged as a powerful paradigm for learning representations without labeled data, often by enforcing invariance to input transformations such as rotations or blurring.
Recent studies have highlighted two pivotal properties for effective representations: (i) avoiding dimensional collapse-where the learned features occupy only a low-dimensional subspace, and (ii) enhancing uniformity of the induced distribution.
In this work, we introduce T-REGS, a simple regularization framework for SSL based on the length of the Minimum Spanning Tree (MST) over the learned representation.
We provide theoretical analysis demonstrating that T-REGS simultaneously mitigates dimensional collapse and promotes distribution uniformity on arbitrary compact Riemannian manifolds.
Several experiments on synthetic data and on classical SSL benchmarks validate the effectiveness of our approach at enhancing representation quality. Julie Mordacq, David Loiseaux, Vicky Kalogeiton, Steve Oudot |
NeurIPS | 3 |
| 2025 | LEAD: Latent Realignment for Human Motion DiffusionabstractAbstract Our goal is to generate realistic human motion from natural language. Modern methods often face a trade‐off between model expressiveness and text‐to‐motion (T2M) alignment. Some align text and motion latent spaces but sacrifice expressiveness; others rely on diffusion models producing impressive motions but lacking semantic meaning in their latent space. This may compromise realism, diversity and applicability. Here, we address this by combining latent diffusion with a realignment mechanism, producing a novel, semantically structured space that encodes the semantics of language. Leveraging this capability, we introduce the task of textual motion inversion to capture novel motion concepts from a few examples. For motion synthesis, we evaluate LEAD on HumanML3D and KIT‐ML and show comparable performance to the state‐of‐the‐art in terms of realism, diversity and text‐motion consistency. Our qualitative analysis and user study reveal that our synthesised motions are sharper, more human‐like and comply better with the text compared to modern methods. For motion textual inversion (MTI), our method demonstrates improvements in capturing out‐of‐distribution characteristics in comparison to traditional VAEs. Nefeli Andreou, Xi Wang 0024, Victoria Fernández Abrevaya, Marie-Paule Cani, Yiorgos Chrysanthou, Vicky Kalogeiton |
Comput. Graph. Forum | 6 |
| 2024 | Collaborating Foundation Models for Domain Generalized Semantic SegmentationabstractDomain Generalized Semantic Segmentation (DGSS) deals with training a model on a labeled source domain with the aim of generalizing to unseen domains during inference. Existing DGSS methods typically effectuate robust features by means of Domain Randomization (DR). Such an approach is often limited as it can only account for style diversification and not content. In this work, we take an orthogonal approach to DGSS and propose to use an assembly of CoLlaborative FOUndation models for Domain Generalized Semantic Segmentation (CLOUDS). In detail, CLOUDS is a framework that integrates Foundation Models of various kinds: (i) CLIP backbone for its robust feature representation, (ii) Diffusion Model to diversify the content, thereby covering various modes of the possible target distribution, and (iii) Segment Anything Model (SAM) for iteratively refining the predictions of the segmentation model. Extensive experiments show that our CLOUDS excels in adapting from synthetic to real DGSS benchmarks and under varying weather conditions, notably outperforming prior methods by 5.6% and 6.7% on averaged mIoU, respectively. Our code is available at https://github.com/yasserben/CLOUDS Yasser Benigmim, Subhankar Roy, Slim Essid, Vicky Kalogeiton, Stéphane Lathuilière |
CVPR | 4 |
| 2024 | Don't Drop Your Samples! Coherence-Aware Training Benefits Conditional DiffusionabstractConditional diffusion models are powerful generative models that can leverage various types of conditional information, such as class labels, segmentation masks, or text captions. However, in many real-world scenarios, conditional infor-mation may be noisy or unreliable due to human annotation errors or weak alignment. In this paper, we propose the Coherence-Aware Diffusion (CAD), a novel method that in-tegrates coherence in conditional information into diffusion models, allowing them to learn from noisy annotations with-out discarding data. We assume that each data point has an associated coherence score that reflects the quality of the conditional information. We then condition the diffusion model on both the conditional information and the coherence score. In this way, the model learns to ignore or discount the conditioning when the coherence is low. We show that CAD is theoretically sound and empirically effective on various conditional generation tasks. Moreover, we show that lever-aging coherence generates realistic and diverse samples that respect conditional information better than models trained on cleaned datasets where samples with low coherence have been discarded. Code and weights here. Nicolas Dufour, Victor Besnier, Vicky Kalogeiton, David Picard |
CVPR | 3 |
| 2024 | E.T. the Exceptional Trajectories: Text-to-Camera-Trajectory Generation with Character Awareness
Robin Courant, Nicolas Dufour, Xi Wang 0024, Marc Christie, Vicky Kalogeiton |
ECCV (4) | 5 |
| 2024 | Learning the What and How of Annotation in Video Object SegmentationabstractVideo Object Segmentation (VOS) is crucial for several applications, from video editing to video data generation. Training a VOS model requires an abundance of manually labeled training videos. The de-facto traditional way of annotating objects requires humans to draw detailed segmentation masks on the target objects at each video frame. This annotation process, however, is tedious and time-consuming. To reduce this annotation cost, in this paper, we propose EVA-VOS, a human-in-the-loop annotation framework for video object segmentation. Unlike the traditional approach, we introduce an agent that predicts iteratively both which frame ("What") to annotate and which annotation type ("How") to use. Then, the annotator annotates only the selected frame that is used to update a VOS module, leading to significant gains in annotation time. We conduct experiments on the MOSE and the DAVIS datasets and we show that: (a) EVA-VOS leads to masks with accuracy close to the human agreement 3.5× faster than the standard way of annotating videos; (b) our frame selection achieves state-of-the-art performance; (c) EVA-VOS yields significant performance gains in terms of annotation time compared to all other methods and baselines. Thanos Delatolas, Vicky Kalogeiton, Dim P. Papadopoulos |
WACV | 2 |
| 2024 | FunnyNet-W: Multimodal Learning of Funny Moments in Videos in the WildabstractAbstract Automatically understanding funny moments (i.e., the moments that make people laugh) when watching comedy is challenging, as they relate to various features, such as body language, dialogues and culture. In this paper, we propose FunnyNet-W, a model that relies on cross- and self-attention for visual, audio and text data to predict funny moments in videos. Unlike most methods that rely on ground truth data in the form of subtitles, in this work we exploit modalities that come naturally with videos: (a) video frames as they contain visual information indispensable for scene understanding, (b) audio as it contains higher-level cues associated with funny moments, such as intonation, pitch and pauses and (c) text automatically extracted with a speech-to-text model as it can provide rich information when processed by a Large Language Model. To acquire labels for training, we propose an unsupervised approach that spots and labels funny audio moments. We provide experiments on five datasets: the sitcoms TBBT, MHD, MUStARD, Friends, and the TED talk UR-Funny. Extensive experiments and analysis show that FunnyNet-W successfully exploits visual, auditory and textual cues to identify funny moments, while our findings reveal FunnyNet-W’s ability to predict funny moments in the wild. FunnyNet-W sets the new state of the art for funny moment detection with multimodal cues on all datasets with and without using ground truth information. Robin Courant, Vicky Kalogeiton |
Int. J. Comput. Vis. | 3 |
| 2023 | Reward Function Design for Crowd Simulation via Reinforcement LearningabstractCrowd simulation is important for video-games design, since it enables to populate virtual worlds with autonomous avatars that navigate in a human-like manner. Reinforcement learning has shown great potential in simulating virtual crowds, but the design of the reward function is critical to achieving effective and efficient results. In this work, we explore the design of reward functions for reinforcement learning-based crowd simulation. We provide theoretical insights on the validity of certain reward functions according to their analytical properties, and evaluate them empirically using a range of scenarios, using the energy efficiency as the metric. Our experiments show that directly minimizing the energy usage is a viable strategy as long as it is paired with an appropriately scaled guiding potential, and enable us to study the impact of the different reward components on the behavior of the simulated crowd. Our findings can inform the development of new crowd simulation techniques, and contribute to the wider study of human-like navigation. Ariel Kwiatkowski, Vicky Kalogeiton, Julien Pettré, Marie-Paule Cani |
MIG | 2 |
| 2023 | Understanding reinforcement learned crowds
Ariel Kwiatkowski, Vicky Kalogeiton, Julien Pettré, Marie-Paule Cani |
Comput. Graph. | 2 |
| 2022 | FunnyNet: Audiovisual Learning of Funny Moments in Videos
Robin Courant, Vicky Kalogeiton |
ACCV (4) | 3 |
| 2022 | SCAM! Transferring Humans Between Images with Semantic Cross Attention Modulation
Nicolas Dufour, David Picard, Vicky Kalogeiton |
ECCV (14) | 3 |
| 2022 | Contrastive Masked Transformers for Forecasting Renal Transplant Function
Léo Milecki, Vicky Kalogeiton, Sylvain Bodard, Dany Anglicheau, Jean-Michel Correas, Marc-Olivier Timsit, Maria Vakalopoulou |
MICCAI (8) | 2 |
| 2022 | A Survey on Reinforcement Learning Methods in Character AnimationabstractAbstract Reinforcement Learning is an area of Machine Learning focused on how agents can be trained to make sequential decisions, and achieve a particular goal within an arbitrary environment. While learning, they repeatedly take actions based on their observation of the environment, and receive appropriate rewards which define the objective. This experience is then used to progressively improve the policy controlling the agent's behavior, typically represented by a neural network. This trained module can then be reused for similar problems, which makes this approach promising for the animation of autonomous, yet reactive characters in simulators, video games or virtual reality environments. This paper surveys the modern Deep Reinforcement Learning methods and discusses their possible applications in Character Animation, from skeletal control of a single, physically‐based character to navigation controllers for individual agents and virtual crowds. It also describes the practical side of training DRL systems, comparing the different frameworks available to build such agents. Ariel Kwiatkowski, Eduardo Alvarado, Vicky Kalogeiton, C. Karen Liu, Julien Pettré, Michiel van de Panne, Marie-Paule Cani |
Comput. Graph. Forum | 3 |
| 2022 | LAEO-Net++: Revisiting People Looking at Each Other in VideosabstractCapturing the 'mutual gaze' of people is essential for understanding and interpreting the social interactions between them. To this end, this paper addresses the problem of detecting people Looking At Each Other (LAEO) in video sequences. For this purpose, we propose LAEO-Net++, a new deep CNN for determining LAEO in videos. In contrast to previous works, LAEO-Net++ takes spatio-temporal tracks as input and reasons about the whole track. It consists of three branches, one for each character's tracked head and one for their relative position. Moreover, we introduce two new LAEO datasets: UCO-LAEO and AVA-LAEO. A thorough experimental evaluation demonstrates the ability of LAEO-Net++ to successfully determine if two people are LAEO and the temporal window where it happens. Our model achieves state-of-the-art results on the existing TVHID-LAEO video dataset, significantly outperforming previous approaches. Finally, we apply LAEO-Net++ to a social network, where we automatically infer the social relationship between pairs of people based on the frequency and duration that they LAEO, and show that LAEO can be a useful tool for guided search of human interactions in videos. Manuel J. Marín-Jiménez, Vicky Kalogeiton, Pablo Medina-Suarez, Andrew Zisserman |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Multimodal Gait Recognition Under Missing ModalitiesabstractMultimodal systems for gait recognition have gained a lot of attention. However, there is a clear gap in the study of missing modalities, which represents real-life scenarios where sensors fail or data get corrupted. Here, we investigate how to handle missing modalities for gait recognition. We propose a single and flexible framework that uses a variable number of input modalities. For each modality, it consists of a branch and a binary unit indicating whether the modality is available; these are gated and merged together. Finally, it generates a single and compact ‘multimodal’ gait signature that encodes biometric information of the input. Our framework outperforms the state of the art on TUM-GAID and extensive experiments reveal its effectiveness for handling missing modalities even in the multiview setup of CASIA-B. The code is available online: https://github.com/avagait/gaitmiss. Rubén Delgado-Escaño, Francisco M. Castro, Nicolás Guil, Vicky Kalogeiton, Manuel J. Marín-Jiménez |
ICIP | 4 |
| 2021 | Multiple Style Transfer Via Variational AutoencoderabstractModern works on style transfer focus on transferring style from a single image. Recently, some approaches study multiple style transfer; these, however, are either too slow or fail to mix multiple styles. We propose ST-VAE, a Variational AutoEncoder for latent space-based style transfer. It performs multiple style transfer by projecting nonlinear styles to a linear latent space, enabling to merge styles via linear interpolation before transferring the new style to the content image. To evaluate ST-VAE, we experiment on COCO for single and multiple style transfer. We also present a case study revealing that ST-VAE outperforms other methods while being faster, flexible, and setting a new path for multiple style transfer. Vicky Kalogeiton, Marie-Paule Cani |
ICIP | 2 |
| 2021 | UGaitNet: Multimodal Gait Recognition With Missing Input ModalitiesabstractGait recognition systems typically rely solely on silhouettes for extracting gait signatures. Nevertheless, these approaches struggle with changes in body shape and dynamic backgrounds; a problem that can be alleviated by learning from multiple modalities. However, in many real-life systems some modalities can be missing, and therefore most existing multimodal frameworks fail to cope with missing modalities. To tackle this problem, in this work, we propose UGaitNet, a unifying framework for gait recognition, robust to missing modalities. UGaitNet handles and mingles various types and combinations of input modalities, i.e. pixel gray value, optical flow, depth maps, and silhouettes, while being camera agnostic. We evaluate UGaitNet on two public datasets for gait recognition: CASIA-B and TUM-GAID, and show that it obtains compact and state-of-the-art gait descriptors when leveraging multiple or missing modalities. Finally, we show that UGaitNet with optical flow and grayscale inputs achieves almost perfect (98.9%) recognition accuracy on CASIA-B (same-view “normal”) and 100% on TUM-GAID (“ellapsed time”). Code will be available. Manuel J. Marín-Jiménez, Francisco M. Castro, Rubén Delgado-Escaño, Vicky Kalogeiton, Nicolás Guil |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2020 | Constrained Video Face Clustering using1NN Relations
Vicky Kalogeiton, Andrew Zisserman |
BMVC | 1 |
| 2020 | Smooth-AP: Smoothing the Path Towards Large-Scale Image Retrieval
Andrew Brown 0006, Weidi Xie, Vicky Kalogeiton, Andrew Zisserman |
ECCV (9) | 3 |
| 2019 | LAEO-Net: Revisiting People Looking at Each Other in VideosabstractCapturing the ‘mutual gaze’ of people is essential for understanding and interpreting the social interactions between them. To this end, this paper addresses the problem of detecting people Looking At Each Other (LAEO) in video sequences. For this purpose, we propose LAEO-Net, a new deep CNN for determining LAEO in videos. In contrast to previous works, LAEO-Net takes spatio-temporal tracks as input and reasons about the whole track. It consists of three branches, one for each character’s tracked head and one for their relative position. Moreover, we introduce two new LAEO datasets: UCO-LAEO and AVA-LAEO. A thorough experimental evaluation demonstrates the ability of LAEO-Net to successfully determine if two people are LAEO and the temporal window where it happens. Our model achieves state-of-the-art results on the existing TVHID-LAEO video dataset, significantly outperforming previous approaches. Manuel J. Marín-Jiménez, Vicky Kalogeiton, Pablo Medina-Suarez, Andrew Zisserman |
CVPR | 2 |
| 2019 | Real-Time Active SLAM and Obstacle Avoidance for an Autonomous Robot Based on Stereo VisionabstractIn this article, the problem of real-time robot exploration and map building (active SLAM) is considered. A single stereo vision camera is exploited by a fully autonomous robot to navigate, localize itself, define its surroundings, and avoid any possible obstacle in the aim of maximizing the mapped region following the optimal route. A modified version of the so-called cognitive-based adaptive optimization algorithm is introduced for the robot to successfully complete its tasks in real time and avoid any local minima entrapment. The method’s effectiveness and performance were tested under various simulation environments as well as real unknown areas with the use of properly equipped robots. Vicky Kalogeiton, Konstantinos Ioannidis, Georgios Ch. Sirakoulis, Elias B. Kosmatopoulos |
Cybern. Syst. | 1 |
| 2017 | Joint Learning of Object and Action DetectorsabstractWhile most existing approaches for detection in videos focus on objects or human actions separately, we aim at jointly detecting objects performing actions, such as cat eating or dog jumping. We introduce an end-to-end multitask objective that jointly learns object-action relationships. We compare it with different training objectives, validate its effectiveness for detecting objects-actions in videos, and show that both tasks of object and action detection benefit from this joint learning. Moreover, the proposed architecture can be used for zero-shot learning of actions: our multitask objective leverages the commonalities of an action performed by different objects, e.g. dog and cat jumping, enabling to detect actions of an object without training with these object-actions pairs. In experiments on the A2D dataset [50], we obtain state-of-the-art results on segmentation of object-action pairs. We finally apply our multitask architecture to detect visual relationships between objects in images of the VRD dataset [24]. Vicky Kalogeiton, Philippe Weinzaepfel, Vittorio Ferrari, Cordelia Schmid |
ICCV | 1 |
| 2017 | Action Tubelet Detector for Spatio-Temporal Action LocalizationabstractCurrent state-of-the-art approaches for spatio-temporal action localization rely on detections at the frame level that are then linked or tracked across time. In this paper, we leverage the temporal continuity of videos instead of operating at the frame level. We propose the ACtion Tubelet detector (ACT-detector) that takes as input a sequence of frames and outputs tubelets, i.e., sequences of bounding boxes with associated scores. The same way state-of-the-art object detectors rely on anchor boxes, our ACT-detector is based on anchor cuboids. We build upon the SSD framework [19]. Convolutional features are extracted for each frame, while scores and regressions are based on the temporal stacking of these features, thus exploiting information from a sequence. Our experimental results show that leveraging sequences offrantes significantly improves detection performance over using individual frames. The gain of our tubelet detector can be explained by both more accurate scores and more precise localization. Our ACT-detector outperforms the state-of-the-art methods for frame-mAP and video-mAP on the J-HMDB [12] and UCF-101 [31] datasets, in particular at high overlap thresholds. Vicky Kalogeiton, Philippe Weinzaepfel, Vittorio Ferrari, Cordelia Schmid |
ICCV | 1 |
| 2017 | Programmable Crossbar Quantum-Dot Cellular Automata CircuitsabstractQuantum-dot fabrication and characterization is a well-established technology, which is used in photonics, quantum optics, and nanoelectronics. Four quantum-dots placed at the corners of a square form a unit cell, which can hold a bit of information and serve as a basis for quantum-dot cellular automata (QCA) nanoelectronic circuits. Although several basic QCA circuits have been designed, fabricated, and tested, proving that quantum-dots can form functional, fast and low-power nanoelectronic circuits, QCA nanoelectronics still remain at its infancy. One of the reasons for this is the lack of design automation tools, which will facilitate the systematic design of large QCA circuits that contemporary applications demand. Here we present novel, programmable QCA circuits, which are based on crossbar architecture. These circuits can be programmed to implement any Boolean function in analogy to CMOS field-programmable gate arrays and open the road that will lead to full design automation of QCA nanoelectronic circuits. Using this architecture we designed and simulated QCA circuits that proved to be area efficient, stable, and reliable. Vicky Kalogeiton, Dim P. Papadopoulos, Orestis Liolis, Vasilios A. Mardiris, Georgios Ch. Sirakoulis, Ioannis Karafyllidis |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2016 | Analysing Domain Shift Factors between Videos and Images for Object DetectionabstractObject detection is one of the most important challenges in computer vision. Object detectors are usually trained on bounding-boxes from still images. Recently, video has been used as an alternative source of data. Yet, for a given test domain (image or video), the performance of the detector depends on the domain it was trained on. In this paper, we examine the reasons behind this performance gap. We define and evaluate different domain shift factors: spatial location accuracy, appearance diversity, image quality and aspect distribution. We examine the impact of these factors by comparing performance before and after factoring them out. The results show that all four factors affect the performance of the detectors and their combined effect explains nearly the whole performance gap. Vicky Kalogeiton, Vittorio Ferrari, Cordelia Schmid |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2013 | Automatic summarization and annotation of videos with lack of metadata information
Dim P. Papadopoulos, Vicky Kalogeiton, Savvas A. Chatzichristofis, Nikos Papamarkos |
Expert Syst. Appl. | 2 |