Evlampios Apostolidis

dblp:138/0884 · also Evlambios E. Apostolidis · DBLP profile ↗
← Back
28ranked-venue papers
12as first author
15since 2021 · last 2025
0000-0001-5376-7158ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 27 · 11 first-author · 14 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 An Experimental Study on Generating Plausible Textual Explanations for Video Summarization
abstract
In this paper, we present our experimental study on generating plausible textual explanations for the outcomes of video summarization. For the needs of this study, we extend an existing framework for multigranular explanation of video summarization by integrating a SOTA Large Multimodal Model (LLaVA-OneVision) and prompting it to produce natural language descriptions of the obtained visual explanations. Following, we focus on one of the most desired characteristics for explainable AI, the plausibility of the obtained explanations that relates with their alignment with the humans' reasoning and expectations. Using the extended framework, we propose an approach for evaluating the plausibility of visual explanations by quantifying the semantic overlap between their textual descriptions and the textual descriptions of the corresponding video summaries, with the help of two methods for creating sentence embed dings (SBERT, SimCSE). Based on the extended framework and the proposed plausibility evaluation approach, we conduct an exper-imental study using a SOTA method (CA-SUM) and two datasets (SumMe, TVSum) for video summarization, to examine whether the more faithful explanations are also the more plausible ones, and identify the most appropriate approach for generating plausible textual explanations for video summarization.
Thomas Eleftheriadis, Evlampios Apostolidis, Vasileios Mezaris
CBMI2
2025 TSalV360: A Method and Dataset for Text-driven Saliency Detection in 360-Degrees Videos
abstract
In this paper, we deal with the task of text-driven saliency detection in 360° videos. For this, we introduce the TSV360 dataset which includes 16,000 triplets of ERP frames, textual descriptions of salient objects/events in these frames, and the associated ground-truth saliency maps. Following, we extend and adapt a SOTA visual-based approach for 360°video saliency detection, and develop the TSalV360 method that takes into account a user-provided text description of the desired objects and/or events. This method leverages a SOTA vision-language model for data representation and integrates a similarity estimation module and a viewport spatio-temporal cross-attention mechanism, to discover dependencies between the different data modalities. Quantitative and qualitative evalu-ations using the TSV360 dataset, showed the competitiveness of TSalV360 compared to a SOTA visual-based approach and documented its competency to perform customized text-driven saliency detection in 360°videos.
Ioannis Kontostathis, Evlampios Apostolidis, Vasileios Mezaris
CBMI2
2025 SD-VSum: A Method and Dataset for Script-Driven Video Summarization
abstract
In this work, we introduce the task of script-driven video summarization, which aims to produce a summary of the full-length video by selecting the parts that are most relevant to a user-provided script outlining the visual content of the desired summary. Following, we extend a recently-introduced large-scale dataset for generic video summarization (VideoXum) by producing natural language descriptions of the different human-annotated summaries that are available per video. In this way we make it compatible with the introduced task, since the available triplets of ''video, summary and summary description'' can be used for training a method that is able to produce different summaries for a given video, driven by the provided script about the content of each summary. Finally, we develop a new network architecture for script-driven video summarization (SD-VSum), that employs a cross-modal attention mechanism for aligning and fusing information from the visual and text modalities. Our experimental evaluations demonstrate the advanced performance of SD-VSum against SOTA approaches for query-driven and generic (unimodal and multimodal) summarization from the literature, and document its capacity to produce video summaries that are adapted to each user's needs about their content.
Manolis Mylonas, Evlampios Apostolidis, Vasileios Mezaris
ACM Multimedia2
2025 Enhancing User Control in AI-Based Video Summarization for Social Media
Ioannis Kontostathis, Evlampios Apostolidis, Konstantinos Apostolidis, Vasileios Mezaris
MMM (5)2
2024 Facilitating the Production of Well-Tailored Video Summaries for Sharing on Social Media
Evlampios Apostolidis, Konstantinos Apostolidis, Vasileios Mezaris
MMM (4)1
2024 An Integrated System for Spatio-temporal Summarization of 360-Degrees Videos
Ioannis Kontostathis, Evlampios Apostolidis, Vasileios Mezaris
MMM (4)2
2024 AI and data-driven media analysis of TV content for optimised digital content marketing
abstract
Abstract To optimise digital content marketing for broadcasters, the Horizon 2020 funded ReTV project developed an end-to-end process termed “Trans-Vector Publishing” and made it accessible through a Web-based tool termed “Content Wizard”. This paper presents this tool with a focus on each of the innovations in data and AI-driven media analysis to address each key step in the digital content marketing workflow: topic selection, content search and video summarisation. First, we use predictive analytics over online data to identify topics the target audience will give the most attention to at a future time. Second, we use neural networks and embeddings to find the video asset closest in content to the identified topic. Third, we use a GAN to create an optimally summarised form of that video for publication, e.g. on social networks. The result is a new and innovative digital content marketing workflow which meets the needs of media organisations in this age of interactive online media where content is transient, malleable and ubiquitous.
Lyndon J. B. Nixon, Konstantinos Apostolidis, Evlampios Apostolidis, Damianos Galanopoulos, Vasileios Mezaris, Basil Philipp, Rasa Bocyte
Multim. Syst.3
2023 Selecting A Diverse Set Of Aesthetically-Pleasing and Representative Video Thumbnails Using Reinforcement Learning
abstract
This paper presents a new reinforcement-based method for video thumbnail selection (called RL-DiVTS), that relies on estimates of the aesthetic quality, representativeness and visual diversity of a small set of selected frames, made with the help of tailored reward functions. The proposed method integrates a novel diversity-aware Frame Picking mechanism that performs a sequential frame selection and applies a reweighting process to demote frames that are visually-similar to the already selected ones. Experiments on two benchmark datasets (OVP and YouTube), using the top-3 matching evaluation protocol, show the competitiveness of RL-DiVTS against other SoA video thumbnail selection and summarization approaches from the literature.
Evlampios Apostolidis, Georgios Balaouras, Vasileios Mezaris, Ioannis Patras
ICIP1
2022 Explaining video summarization based on the focus of attention
abstract
In this paper we propose a method for explaining video summarization. We start by formulating the problem as the creation of an explanation mask which indicates the parts of the video that influenced the most the estimates of a video summarization network, about the frames’ importance. Then, we explain how the typical analysis pipeline of attention-based networks for video summarization can be used to define explanation signals, and we examine various attention-based signals that have been studied as explanations in the NLP domain. We evaluate the performance of these signals by investigating the video summarization network’s input-output relationship according to different replacement functions, and utilizing measures that quantify the capability of explanations to spot the most and least influential parts of a video. We run experiments using an attention-based network (CA-SUM) and two datasets (SumMe and TVSum) for video summarization. Our evaluations indicate the advanced performance of explanations formed using the inherent attention weights, and demonstrate the ability of our method to explain the video summarization results using clues about the focus of the attention mechanism.
Evlampios Apostolidis, Georgios Balaouras, Vasileios Mezaris, Ioannis Patras
ISM1
2022 Summarizing Videos using Concentrated Attention and Considering the Uniqueness and Diversity of the Video Frames
abstract
In this work, we describe a new method for unsupervised video summarization. To overcome limitations of existing unsupervised video summarization approaches, that relate to the unstable training of Generator-Discriminator architectures, the use of RNNs for modeling long-range frames' dependencies and the ability to parallelize the training process of RNN-based network architectures, the developed method relies solely on the use of a self-attention mechanism to estimate the importance of video frames. Instead of simply modeling the frames' dependencies based on global attention, our method integrates a concentrated attention mechanism that is able to focus on non-overlapping blocks in the main diagonal of the attention matrix, and to enrich the existing information by extracting and exploiting knowledge about the uniqueness and diversity of the associated frames of the video. In this way, our method makes better estimates about the significance of different parts of the video, and drastically reduces the number of learnable parameters. Experimental evaluations using two benchmarking datasets (SumMe and TVSum) show the competitiveness of the proposed method against other state-of-the-art unsupervised summarization approaches, and demonstrate its ability to produce video summaries that are very close to the human preferences. An ablation study that focuses on the introduced components, namely the use of concentrated attention in combination with attention-based estimates about the frames' uniqueness and diversity, shows their relative contributions to the overall summarization performance.
Evlampios Apostolidis, Georgios Balaouras, Vasileios Mezaris, Ioannis Patras
ICMR1
2021 Combining Global and Local Attention with Positional Encoding for Video Summarization
abstract
This paper presents a new method for supervised video summarization. To overcome drawbacks of existing RNN-based summarization architectures, that relate to the modeling of long-range frames’ dependencies and the ability to parallelize the training process, the developed model re-lies on the use of self-attention mechanisms to estimate the importance of video frames. Contrary to previous attention-based summarization approaches that model the frames’ dependencies by observing the entire frame sequence, our method combines global and local multi-head attention mechanisms to discover different modelings of the frames’ dependencies at different levels of granularity. Moreover, the utilized attention mechanisms integrate a component that encodes the temporal position of video frames - this is of major importance when producing a video summary. Experiments on two datasets (SumMe and TVSum) demonstrate the effectiveness of the proposed model compared to existing attention-based methods, and its competitiveness against other state-of-the-art supervised summarization approaches. An ablation study that focuses on our main proposed components, namely the use of global and local multi-head attention mechanisms in collaboration with an absolute positional encoding component, shows their relative contributions to the overall summarization performance.
Evlampios Apostolidis, Georgios Balaouras, Vasileios Mezaris, Ioannis Patras
ISM1
2021 Combining Adversarial and Reinforcement Learning for Video Thumbnail Selection
abstract
This paper presents a new method for unsupervised video thumbnail selection. The developed network architecture selects video thumbnails based on two criteria: the representativeness and the aesthetic quality of their visual content. Training relies on a combination of adversarial and reinforcement learning. The former is used to train a discriminator, whose goal is to distinguish the original from a reconstructed version of the video based on a small set of candidate thumbnails. The discriminator's feedback is a measure of the representativeness of the selected thumbnails. This measure is combined with estimates about the aesthetic quality of the thumbnails (made using a SoA Fully Convolutional Network) to form a reward and train the thumbnail selector via reinforcement learning. Experiments on two datasets (OVP and Youtube) show the competitiveness of the proposed method against other SoA approaches. An ablation study with respect to the adopted thumbnail selection criteria documents the importance of considering the aesthetics, and the contribution of this information when used in combination with measures about the representativeness of the visual content.
Evlampios Apostolidis, Eleni Adamantidou, Vasileios Mezaris, Ioannis Patras
ICMR1
2021 Content Wizard: demo of a trans-vector digital video publication tool
abstract
In order to optimise the distribution of video assets online, media organizations need tailor their offerings for specific digital channels and better understand the interests of their audiences at particular points in time, which are often influenced by contemporary new stories and trends on social media. For this purpose, the research project ReTV has developed a Web-based tool termed ’Content Wizard’ which demonstrates an end-to-end, semi-automated workflow for video content creation, adaptation and distribution across digital channels. Digital assets can be selected based on predicted future trending topics, re-purposed according to the different digital channels they will be published upon and scheduled for the optimal future publication date. The result is an innovative video publication workflow that meets the marketing needs of media organisations in this age of transient online media spread across multiple channels.
Lyndon J. B. Nixon, Konstantinos Apostolidis, Evlampios Apostolidis, Damianos Galanopoulos, Vasileios Mezaris, Basil Philipp, Rasa Bocyte
IMX3
2021 Video Summarization Using Deep Neural Networks: A Survey
abstract
Video summarization technologies aim to create a concise and complete synopsis by selecting the most informative parts of the video content. Several approaches have been developed over the last couple of decades, and the current state of the art is represented by methods that rely on modern deep neural network architectures. This work focuses on the recent advances in the area and provides a comprehensive survey of the existing deep-learning-based methods for generic video summarization. After presenting the motivation behind the development of technologies for video summarization, we formulate the video summarization task and discuss the main characteristics of a typical deep-learning-based analysis pipeline. Then, we suggest a taxonomy of the existing algorithms and provide a systematic review of the relevant literature that shows the evolution of the deep-learning-based video summarization technologies and leads to suggestions for future developments. We then report on protocols for the objective evaluation of video summarization algorithms, and we compare the performance of several deep-learning-based approaches. Based on the outcomes of these comparisons, as well as some documented considerations about the amount of annotated data and the suitability of evaluation protocols, we indicate potential future research directions.
Evlampios Apostolidis, Eleni Adamantidou, Alexandros I. Metsai, Vasileios Mezaris, Ioannis Patras
Proc. IEEE1
2021 AC-SUM-GAN: Connecting Actor-Critic and Generative Adversarial Networks for Unsupervised Video Summarization
abstract
This paper presents a new method for unsupervised video summarization. The proposed architecture embeds an Actor-Critic model into a Generative Adversarial Network and formulates the selection of important video fragments (that will be used to form the summary) as a sequence generation task. The Actor and the Critic take part in a game that incrementally leads to the selection of the video key-fragments, and their choices at each step of the game result in a set of rewards from the Discriminator. The designed training workflow allows the Actor and Critic to discover a space of actions and automatically learn a policy for key-fragment selection. Moreover, the introduced criterion for choosing the best model after the training ends, enables the automatic selection of proper values for parameters of the training process that are not learned from the data (such as the regularization factor σ). Experimental evaluation on two benchmark datasets (SumMe and TVSum) demonstrates that the proposed AC-SUM-GAN model performs consistently well and gives SoA results in comparison to unsupervised methods, that are also competitive with respect to supervised methods.
Evlampios Apostolidis, Eleni Adamantidou, Alexandros I. Metsai, Vasileios Mezaris, Ioannis Patras
IEEE Trans. Circuits Syst. Video Technol.1
2020 Performance over Random: A Robust Evaluation Protocol for Video Summarization Methods
abstract
This paper proposes a new evaluation approach for video summarization algorithms. We start by studying the currently established evaluation protocol; this protocol, defined over the ground-truth annotations of the SumMe and TVSum datasets, quantifies the agreement between the user-defined and the automatically-created summaries with F-Score, and reports the average performance on a few different training/testing splits of the used dataset. We evaluate five publicly-available summarization algorithms under a large-scale experimental setting with 50 randomly-created data splits. We show that the results reported in the papers are not always congruent with their performance on the large-scale experiment, and that the F-Score cannot be used for comparing algorithms evaluated on different splits. We also show that the above shortcomings of the established evaluation protocol are due to the significantly varying levels of difficulty among the utilized splits, that affect the outcomes of the evaluations. Further analysis of these findings indicates a noticeable performance correlation among all algorithms and a random summarizer. To mitigate these shortcomings we propose an evaluation protocol that makes estimates about the difficulty of each used data split and utilizes this information during the evaluation process. Experiments involving different evaluation settings demonstrate the increased representativeness of performance results when using the proposed evaluation approach, and the increased reliability of comparisons when the examined methods have been evaluated on different data splits.
Evlampios Apostolidis, Eleni Adamantidou, Alexandros I. Metsai, Vasileios Mezaris, Ioannis Patras
ACM Multimedia1
2020 Unsupervised Video Summarization via Attention-Driven Adversarial Learning
Evlampios Apostolidis, Eleni Adamantidou, Alexandros I. Metsai, Vasileios Mezaris, Ioannis Patras
MMM (1)1
2020 A Web Service for Video Summarization
abstract
This paper presents a Web service that supports the automatic generation of video summaries for user-submitted videos. The developed Web application decomposes the video into segments, evaluates the fitness of each segment to be included in the video summary and selects appropriate segments until a pre-defined time budget is filled. The integrated deep-learning-based video analysis and summarization technologies exhibit state-of-the-art performance and, by exploiting the processing capabilities of modern GPUs, offer faster than real-time processing. Configurations for generating video summaries that fulfill the specifications for posting on the most common video sharing platforms and social networks are available in the user interface of this application, enabling the one-click generation of distribution-channel-specific summaries.
Chrysa Collyda, Konstantinos Apostolidis, Evlampios Apostolidis, Eleni Adamantidou, Alexandros I. Metsai, Vasileios Mezaris
IMX3
2019 Multimodal Video Annotation for Retrieval and Discovery of Newsworthy Video in a News Verification Scenario
Lyndon J. B. Nixon, Evlampios Apostolidis, Fotini Markatopoulou, Ioannis Patras, Vasileios Mezaris
MMM (1)2
2019 Detecting Tampered Videos with Multimedia Forensics and Deep Learning
Markos Zampoglou, Fotini Markatopoulou, Grégoire Mercier, Despoina Touska, Evlampios Apostolidis, Symeon Papadopoulos, Roger Cozien, Ioannis Patras, Vasileios Mezaris, Ioannis Kompatsiaris
MMM (1)5
2018 A Motion-Driven Approach for Fine-Grained Temporal Segmentation of User-Generated Videos
Konstantinos Apostolidis, Evlampios Apostolidis, Vasileios Mezaris
MMM (1)2
2017 VideoAnalysis4ALL: An On-line Tool for the Automatic Fragmentation and Concept-based Annotation, and the Interactive Exploration of Videos
abstract
This paper presents the VideoAnalysis4ALL tool that supports the automatic fragmentation and concept-based annotation of videos, and the exploration of the annotated video fragments through an interactive user interface. The developed web application decomposes the video into two different granularities, namely shots and scenes, and annotates each fragment by evaluating the existence of a number (several hundreds) of high-level visual concepts in the keyframes extracted from these fragments. Through the analysis the tool enables the identification and labeling of semantically coherent video fragments, while its user interfaces allow the discovery of these fragments with the help of human-interpretable concepts. The integrated state-of-the-art video analysis technologies perform very well and, by exploiting the processing capabilities of multi-thread / multi-core architectures, reduce the time required for analysis to approximately one third of the video's duration, thus making the analysis three times faster than real-time processing.
Chrysa Collyda, Evlampios Apostolidis, Alexandros Pournaras, Fotini Markatopoulou, Vasileios Mezaris, Ioannis Patras
ICMR2
2017 A web-based tool for fast instance-level labeling of videos and the creation of spatiotemporal media fragments
Anastasia Ioannidou, Evlampios Apostolidis, Chrysa Collyda, Vasileios Mezaris
Multim. Tools Appl.2
2016 VERGE: A Multimodal Interactive Search Engine for Video Browsing and Retrieval
Anastasia Moumtzidou, Theodoros Mironidis, Evlampios Apostolidis, Fotini Markatopoulou, Anastasia Ioannidou, Ilias Gialampoukidis, Konstantinos Avgerinakis, Stefanos Vrochidis, Vasileios Mezaris, Ioannis Kompatsiaris, Ioannis Patras
MMM (2)3
2015 VERGE: A Multimodal Interactive Video Search Engine
Anastasia Moumtzidou, Konstantinos Avgerinakis, Evlampios Apostolidis, Fotini Markatopoulou, Konstantinos Apostolidis, Theodoros Mironidis, Stefanos Vrochidis, Vasileios Mezaris, Ioannis Kompatsiaris, Ioannis Patras
MMM (2)3
2014 Fast shot segmentation combining global and local visual descriptors
abstract
This paper introduces an algorithm for fast temporal segmentation of videos into shots. The proposed method detects abrupt and gradual transitions, based on the visual similarity of neighboring frames of the video. The descriptive efficiency of both local (SURF) and global (HSV histograms) descriptors is exploited for assessing frame similarity, while GPU-based processing is used for accelerating the analysis. Specifically, abrupt transitions are initially detected between successive video frames where there is a sharp change in the visual content, which is expressed by a very low similarity score. Then, the calculated scores are further analysed for the identification of frame-sequences where a progressive change of the visual content takes place and, in this way gradual transitions are detected. Finally, a post-processing step is performed aiming to identify outliers due to object/camera movement and flash-lights. The experiments show that the proposed algorithm achieves high accuracy while being capable of faster-than-real-time analysis.
Evlampios Apostolidis, Vasileios Mezaris
ICASSP1
2014 Automatic fine-grained hyperlinking of videos within a closed collection using scene segmentation
abstract
This paper introduces a framework for establishing links between related media fragments within a collection of videos. A set of analysis techniques is applied for extracting information from different types of data. Visual-based shot and scene segmentation is performed for defining media fragments at different granularity levels, while visual cues are detected from keyframes of the video via concept detection and optical character recognition (OCR). Keyword extraction is applied on textual data such as the output of OCR, subtitles and metadata. This set of results is used for the automatic identification and linking of related media fragments. The proposed framework exhibited competitive performance in the Video Hyperlinking sub-task of MediaEval 2013, indicating that video scene segmentation can provide more meaningful segments, compared to other decomposition methods, for hyperlinking purposes.
Evlampios Apostolidis, Vasileios Mezaris, Mathilde Sahuguet, Benoit Huet, Barbora Cervenková, Daniel Stein, Stefan Eickeler, José Luis Redondo García, Raphaël Troncy, Lukás Pikora
ACM Multimedia1
2014 VERGE: An Interactive Search Engine for Browsing Video Collections
Anastasia Moumtzidou, Konstantinos Avgerinakis, Evlampios Apostolidis, Vera Aleksic, Fotini Markatopoulou, Christina Papagiannopoulou, Stefanos Vrochidis, Vasileios Mezaris, Reinhard Busch, Ioannis Kompatsiaris
MMM (2)3