Vasileios Mezaris

dblp:80/6216 · DBLP profile ↗
← Back
142ranked-venue papers
13as first author
36since 2021 · last 2026
0000-0002-0121-4364ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 129 · 12 first-author · 34 since 2021Artificial intelligence and machine learning · 17 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 12 · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 1 since 2021Computer networks · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 VERGE in VBS 2026
Nick Pantelidis, Eleni Kosmidou, Damianos Galanopoulos, Dimitris Georgalis, Stefanos Pasios, Konstantinos Apostolidis, Andreas Goulas, Maria Pegia, Georgios Tsionkis, Konstantinos Gkountakos, Grigorios Kouvrakis, Anastasia Moumtzidou, Ilias Gialampoukidis, Stefanos Vrochidis, Vasileios Mezaris, Ioannis Kompatsiaris
MMM (4)15
2025 An Experimental Study on Generating Plausible Textual Explanations for Video Summarization
abstract
In this paper, we present our experimental study on generating plausible textual explanations for the outcomes of video summarization. For the needs of this study, we extend an existing framework for multigranular explanation of video summarization by integrating a SOTA Large Multimodal Model (LLaVA-OneVision) and prompting it to produce natural language descriptions of the obtained visual explanations. Following, we focus on one of the most desired characteristics for explainable AI, the plausibility of the obtained explanations that relates with their alignment with the humans' reasoning and expectations. Using the extended framework, we propose an approach for evaluating the plausibility of visual explanations by quantifying the semantic overlap between their textual descriptions and the textual descriptions of the corresponding video summaries, with the help of two methods for creating sentence embed dings (SBERT, SimCSE). Based on the extended framework and the proposed plausibility evaluation approach, we conduct an exper-imental study using a SOTA method (CA-SUM) and two datasets (SumMe, TVSum) for video summarization, to examine whether the more faithful explanations are also the more plausible ones, and identify the most appropriate approach for generating plausible textual explanations for video summarization.
Thomas Eleftheriadis, Evlampios Apostolidis, Vasileios Mezaris
CBMI3
2025 TSalV360: A Method and Dataset for Text-driven Saliency Detection in 360-Degrees Videos
abstract
In this paper, we deal with the task of text-driven saliency detection in 360° videos. For this, we introduce the TSV360 dataset which includes 16,000 triplets of ERP frames, textual descriptions of salient objects/events in these frames, and the associated ground-truth saliency maps. Following, we extend and adapt a SOTA visual-based approach for 360°video saliency detection, and develop the TSalV360 method that takes into account a user-provided text description of the desired objects and/or events. This method leverages a SOTA vision-language model for data representation and integrates a similarity estimation module and a viewport spatio-temporal cross-attention mechanism, to discover dependencies between the different data modalities. Quantitative and qualitative evalu-ations using the TSV360 dataset, showed the competitiveness of TSalV360 compared to a SOTA visual-based approach and documented its competency to perform customized text-driven saliency detection in 360°videos.
Ioannis Kontostathis, Evlampios Apostolidis, Vasileios Mezaris
CBMI3
2025 VidCtx: Context-aware Video Question Answering with Image Models
abstract
To address computational and memory limitations of Large Multimodal Models in the Video Question-Answering task, several recent methods extract textual representations per frame (e.g., by captioning) and feed them to a Large Language Model (LLM) that processes them to produce the final response. However, in this way, the LLM does not have access to visual information and often has to process repetitive textual descriptions of nearby frames. To address those shortcomings, in this paper, we introduce VidCtx, a novel training-free VideoQA framework which integrates both modalities, i.e. both visual information from input frames and textual descriptions of others frames that give the appropriate context. More specifically, in the proposed framework a pre-trained Large Multimodal Model (LMM) is prompted to extract at regular intervals, question-aware textual descriptions (captions) of video frames. Those will be used as context when the same LMM will be prompted to answer the question at hand given as input a) a certain frame, b) the question and c) the context/caption of an appropriate frame. To avoid redundant information, we chose as context the descriptions of distant frames. Finally, a simple yet effective max pooling mechanism is used to aggregate the frame-level decisions. This methodology enables the model to focus on the relevant segments of the video and scale to a high number of frames. Experiments show that VidCtx achieves competitive performance among approaches that rely on open models on three public Video QA benchmarks, NExT-QA, IntentQA and STAR. Our code is available at https://github.com/IDT-ITI/VidCtx.
Andreas Goulas, Vasileios Mezaris, Ioannis Patras
ICME2
2025 SD-VSum: A Method and Dataset for Script-Driven Video Summarization
abstract
In this work, we introduce the task of script-driven video summarization, which aims to produce a summary of the full-length video by selecting the parts that are most relevant to a user-provided script outlining the visual content of the desired summary. Following, we extend a recently-introduced large-scale dataset for generic video summarization (VideoXum) by producing natural language descriptions of the different human-annotated summaries that are available per video. In this way we make it compatible with the introduced task, since the available triplets of ''video, summary and summary description'' can be used for training a method that is able to produce different summaries for a given video, driven by the provided script about the content of each summary. Finally, we develop a new network architecture for script-driven video summarization (SD-VSum), that employs a cross-modal attention mechanism for aligning and fusing information from the visual and text modalities. Our experimental evaluations demonstrate the advanced performance of SD-VSum against SOTA approaches for query-driven and generic (unimodal and multimodal) summarization from the literature, and document its capacity to produce video summaries that are adapted to each user's needs about their content.
Manolis Mylonas, Evlampios Apostolidis, Vasileios Mezaris
ACM Multimedia3
2025 Enhancing User Control in AI-Based Video Summarization for Social Media
Ioannis Kontostathis, Evlampios Apostolidis, Konstantinos Apostolidis, Vasileios Mezaris
MMM (5)4
2025 VERGE in VBS 2025
Nick Pantelidis, Dimitris Georgalis, Maria Pegia, Damianos Galanopoulos, Konstantinos Apostolidis, Klearchos Stavrothanasopoulos, Anastasia Moumtzidou, Konstantinos Gkountakos, Ilias Gialampoukidis, Stefanos Vrochidis, Vasileios Mezaris, Ioannis Kompatsiaris
MMM (5)11
2024 Finding Video Shots for Immersive Journalism Through Text-to-Video Search
abstract
Video assets from archives or online platforms can provide relevant content for embedding into immersive scenes or for generation of 3D objects or scenes. However, XR content creators lack tools to find relevant video segments for their chosen topic. In this paper, we explore the use case of journalists creating immersive experiences for news stories and their need to find related video material to create and populate a 3D scene. An innovative approach creates text and video embeddings and matches textual input queries to relevant video shots. This is provided via a Web dashboard for search and retrieval across video collections, with selected shots forming the input to content creation tools to generate and populate an immersive scene, meaning journalists do not need specialist knowledge to communicate stories via XR.
Lyndon J. B. Nixon, Damianos Galanopoulos, Vasileios Mezaris
CBMI3
2024 Video Shot Discovery Through Text2Video Embeddings in a News Analytics Dashboard
abstract
This demonstration will show how video shot discovery through joint text-video embedding has been integrated into a journalistic workflow through a news monitoring dashboard with the purpose of identifying suitable video material for the creation of immersive scenes around a chosen news story or topic.
Lyndon J. B. Nixon, Damianos Galanopoulos, Vasileios Mezaris, Alexander Hubmann-Haidvogel, Daniel Fischl, Arno Scharl
CBMI3
2024 Verge: Simplifying Video Search for Novice Users
abstract
This paper presents an updated iteration of the VERGE interactive video retrieval system. It offers various search options like free text and concept-based text search, color similarity, people and face detection, and visual and semantic similarity search. The system is designed to handle large amounts of data efficiently using advanced indexing techniques and state-of-the-art AI technology for visual content analysis. This paper describes enhancements made to improve usability for non-expert users, particularly through changes to the search and browsing interface.
Nick Pantelidis, Maria Pegia, Damianos Galanopoulos, Konstantinos Apostolidis, Dimitris Georgalis, Klearchos Stavrothanasopoulos, Anastasia Moumtzidou, Konstantinos Gkountakos, Ilias Gialampoukidis, Stefanos Vrochidis, Vasileios Mezaris, Ioannis Kompatsiaris
CBMI11
2024 Exploiting LMM-Based Knowledge for Image Classification Tasks
Maria Tzelepi, Vasileios Mezaris
EANN2
2024 Online Anchor-Based Training For Image Classification Tasks
abstract
In this paper, we aim to improve the performance of a deep learning model towards image classification tasks, proposing a novel anchor-based training methodology, named Online Anchor-based Training (OAT). The OAT method, guided by the insights provided in the anchor-based object detection methodologies, instead of learning directly the class labels, proposes to train a model to learn percentage changes of the class labels with respect to defined anchors. We define as anchors the batch centers at the output of the model. Then, during the test phase, the predictions are converted back to the original class label space, and the performance is evaluated. The effectiveness of the OAT method is validated on four datasets.
Maria Tzelepi, Vasileios Mezaris
ICIP2
2024 LMM-Regularized CLIP Embeddings for Image Classification
abstract
In this paper we deal with image classification tasks using the powerful CLIP vision-language model. Our goal is to advance the classification performance using the CLIP’s image encoder, by proposing a novel Large Multimodal Model (LMM) based regularization method. The proposed method uses an LMM to extract semantic descriptions for the images of the dataset. Then, it uses the CLIP’s text encoder, frozen, in order to obtain the corresponding text embeddings and compute the mean semantic class descriptions. Subsequently, we adapt the CLIP’s image encoder by adding a classification head, and we train it along with the image encoder output, apart from the main classification objective, with an additional auxiliary objective. The additional objective forces the embeddings at the image encoder’s output to become similar to their corresponding LMM-generated mean semantic class descriptions. In this way, it produces embeddings with enhanced discrimination ability, leading to improved classification performance. The effectiveness of the proposed regularization method is validated through extensive experiments on three image classification datasets.
Maria Tzelepi, Vasileios Mezaris
ISM2
2024 Facilitating the Production of Well-Tailored Video Summaries for Sharing on Social Media
Evlampios Apostolidis, Konstantinos Apostolidis, Vasileios Mezaris
MMM (4)3
2024 An Integrated System for Spatio-temporal Summarization of 360-Degrees Videos
Ioannis Kontostathis, Evlampios Apostolidis, Vasileios Mezaris
MMM (4)3
2024 VERGE in VBS 2024
Nick Pantelidis, Maria Pegia, Damianos Galanopoulos, Konstantinos Apostolidis, Klearchos Stavrothanasopoulos, Anastasia Moumtzidou, Konstantinos Gkountakos, Ilias Gialampoukidis, Stefanos Vrochidis, Vasileios Mezaris, Ioannis Kompatsiaris, Björn Þór Jónsson 0001
MMM (4)10
2024 Exploring Multi-modal Fusion for Image Manipulation Detection and Localization
Konstantinos Triaridis, Vasileios Mezaris
MMM (3)2
2024 AI and data-driven media analysis of TV content for optimised digital content marketing
abstract
Abstract To optimise digital content marketing for broadcasters, the Horizon 2020 funded ReTV project developed an end-to-end process termed “Trans-Vector Publishing” and made it accessible through a Web-based tool termed “Content Wizard”. This paper presents this tool with a focus on each of the innovations in data and AI-driven media analysis to address each key step in the digital content marketing workflow: topic selection, content search and video summarisation. First, we use predictive analytics over online data to identify topics the target audience will give the most attention to at a future time. Second, we use neural networks and embeddings to find the video asset closest in content to the identified topic. Third, we use a GAN to create an optimally summarised form of that video for publication, e.g. on social networks. The result is a new and innovative digital content marketing workflow which meets the needs of media organisations in this age of interactive online media where content is transient, malleable and ubiquitous.
Lyndon J. B. Nixon, Konstantinos Apostolidis, Evlampios Apostolidis, Damianos Galanopoulos, Vasileios Mezaris, Basil Philipp, Rasa Bocyte
Multim. Syst.5
2023 Selecting A Diverse Set Of Aesthetically-Pleasing and Representative Video Thumbnails Using Reinforcement Learning
abstract
This paper presents a new reinforcement-based method for video thumbnail selection (called RL-DiVTS), that relies on estimates of the aesthetic quality, representativeness and visual diversity of a small set of selected frames, made with the help of tailored reward functions. The proposed method integrates a novel diversity-aware Frame Picking mechanism that performs a sequential frame selection and applies a reweighting process to demote frames that are visually-similar to the already selected ones. Experiments on two benchmark datasets (OVP and YouTube), using the top-3 matching evaluation protocol, show the competitiveness of RL-DiVTS against other SoA video thumbnail selection and summarization approaches from the literature.
Evlampios Apostolidis, Georgios Balaouras, Vasileios Mezaris, Ioannis Patras
ICIP3
2023 Masked Feature Modelling for the unsupervised pre-training of a Graph Attention Network block for bottom-up video event recognition
abstract
In this paper, we introduce Masked Feature Modelling (MFM), a novel approach for the unsupervised pretraining of a Graph Attention Network (GAT) block. MFM utilizes a pretrained Visual Tokenizer to reconstruct masked features of objects within a video, leveraging the MiniKinetics dataset. We then incorporate the pre-trained GAT block into a state-of-the-art bottom-up supervised video-event recognition architecture, ViGAT, to improve the model’s starting point and overall accuracy. Experimental evaluations on the YLI-MED dataset demonstrate the effectiveness of MFM in improving event recognition performance.
Dimitrios Daskalakis, Nikolaos Gkalelis, Vasileios Mezaris
ISM3
2023 VERGE in VBS 2023
Nick Pantelidis, Stelios Andreadis, Maria Pegia, Anastasia Moumtzidou, Damianos Galanopoulos, Konstantinos Apostolidis, Despoina Touska, Konstantinos Gkountakos, Ilias Gialampoukidis, Stefanos Vrochidis, Vasileios Mezaris, Ioannis Kompatsiaris
MMM (1)11
2022 Explaining video summarization based on the focus of attention
abstract
In this paper we propose a method for explaining video summarization. We start by formulating the problem as the creation of an explanation mask which indicates the parts of the video that influenced the most the estimates of a video summarization network, about the frames’ importance. Then, we explain how the typical analysis pipeline of attention-based networks for video summarization can be used to define explanation signals, and we examine various attention-based signals that have been studied as explanations in the NLP domain. We evaluate the performance of these signals by investigating the video summarization network’s input-output relationship according to different replacement functions, and utilizing measures that quantify the capability of explanations to spot the most and least influential parts of a video. We run experiments using an attention-based network (CA-SUM) and two datasets (SumMe and TVSum) for video summarization. Our evaluations indicate the advanced performance of explanations formed using the inherent attention weights, and demonstrate the ability of our method to explain the video summarization results using clues about the focus of the attention mechanism.
Evlampios Apostolidis, Georgios Balaouras, Vasileios Mezaris, Ioannis Patras
ISM3
2022 Gated-ViGAT: Efficient Bottom-Up Event Recognition and Explanation Using a New Frame Selection Policy and Gating Mechanism
abstract
In this paper, Gated-ViGAT, an efficient approach for video event recognition, utilizing bottom-up (object) information, a new frame sampling policy and a gating mechanism is proposed. Specifically, the frame sampling policy uses weighted in-degrees (WiDs), derived from the adjacency matrices of graph attention networks (GATs), and a dissimilarity measure to select the most salient and at the same time diverse frames representing the event in the video. Additionally, the proposed gating mechanism fetches the selected frames sequentially, and commits early-exiting when an adequately confident decision is achieved. In this way, only a few frames are processed by the computationally expensive branch of our network that is responsible for the bottom-up information extraction. The experimental evaluation on two large, publicly available video datasets (MiniKinetics, ActivityNet) demonstrates that Gated-ViGAT provides a large computational complexity reduction in comparison to our previous approach (ViGAT), while maintaining the excellent event recognition and explainability performance1.
Nikolaos Gkalelis, Dimitrios Daskalakis, Vasileios Mezaris
ISM3
2022 TAME: Attention Mechanism Based Feature Fusion for Generating Explanation Maps of Convolutional Neural Networks
abstract
The apparent "black box" nature of neural networks is a barrier to adoption in applications where explainability is essential. This paper presents TAME (Trainable Attention Mechanism for Explanations)1, a method for generating explanation maps with a multi-branch hierarchical attention mechanism. TAME combines a target model’s feature maps from multiple layers using an attention mechanism, transforming them into an explanation map. TAME can easily be applied to any convolutional neural network (CNN) by streamlining the optimization of the attention mechanism’s training method and the selection of target model’s feature maps. After training, explanation maps can be computed in a single forward pass. We apply TAME to two widely used models, i.e. VGG-16 and ResNet-50, trained on ImageNet and show improvements over previous top-performing methods. We also provide a comprehensive ablation study comparing the performance of different variations of TAME’s architecture.
Mariano Ntrougkas, Nikolaos Gkalelis, Vasileios Mezaris
ISM3
2022 Summarizing Videos using Concentrated Attention and Considering the Uniqueness and Diversity of the Video Frames
abstract
In this work, we describe a new method for unsupervised video summarization. To overcome limitations of existing unsupervised video summarization approaches, that relate to the unstable training of Generator-Discriminator architectures, the use of RNNs for modeling long-range frames' dependencies and the ability to parallelize the training process of RNN-based network architectures, the developed method relies solely on the use of a self-attention mechanism to estimate the importance of video frames. Instead of simply modeling the frames' dependencies based on global attention, our method integrates a concentrated attention mechanism that is able to focus on non-overlapping blocks in the main diagonal of the attention matrix, and to enrich the existing information by extracting and exploiting knowledge about the uniqueness and diversity of the associated frames of the video. In this way, our method makes better estimates about the significance of different parts of the video, and drastically reduces the number of learnable parameters. Experimental evaluations using two benchmarking datasets (SumMe and TVSum) show the competitiveness of the proposed method against other state-of-the-art unsupervised summarization approaches, and demonstrate its ability to produce video summaries that are very close to the human preferences. An ablation study that focuses on the introduced components, namely the use of concentrated attention in combination with attention-based estimates about the frames' uniqueness and diversity, shows their relative contributions to the overall summarization performance.
Evlampios Apostolidis, Georgios Balaouras, Vasileios Mezaris, Ioannis Patras
ICMR3
2022 VERGE in VBS 2022
Stelios Andreadis, Anastasia Moumtzidou, Damianos Galanopoulos, Nick Pantelidis, Konstantinos Apostolidis, Despoina Touska, Konstantinos Gkountakos, Maria Pegia, Ilias Gialampoukidis, Stefanos Vrochidis, Vasileios Mezaris, Ioannis Kompatsiaris
MMM (2)11
2022 Data-driven personalisation of television content: a survey
Lyndon J. B. Nixon, Jeremy D. Foss, Konstantinos Apostolidis, Vasileios Mezaris
Multim. Syst.4
2022 Special issue on data-driven personalisation of television content
Lyndon J. B. Nixon, Jeremy D. Foss, Vasileios Mezaris
Multim. Syst.3
2021 A Fast Smart-Cropping Method and Dataset for Video Retargeting
abstract
In this paper a method that re-targets a video to a different aspect ratio using cropping is presented. We argue that cropping methods are more suitable for video aspect ratio transformation when the minimization of semantic distortions is a prerequisite. For our method, we utilize visual saliency to find the image regions of attention, and we employ a filtering-through-clustering technique to select the main region of focus. We additionally introduce the first publicly available benchmark dataset for video cropping, annotated by 6 human subjects. Experimental evaluation on the introduced dataset shows the competitiveness of our method.
Konstantinos Apostolidis, Vasileios Mezaris
ICIP2
2021 Combining Global and Local Attention with Positional Encoding for Video Summarization
abstract
This paper presents a new method for supervised video summarization. To overcome drawbacks of existing RNN-based summarization architectures, that relate to the modeling of long-range frames’ dependencies and the ability to parallelize the training process, the developed model re-lies on the use of self-attention mechanisms to estimate the importance of video frames. Contrary to previous attention-based summarization approaches that model the frames’ dependencies by observing the entire frame sequence, our method combines global and local multi-head attention mechanisms to discover different modelings of the frames’ dependencies at different levels of granularity. Moreover, the utilized attention mechanisms integrate a component that encodes the temporal position of video frames - this is of major importance when producing a video summary. Experiments on two datasets (SumMe and TVSum) demonstrate the effectiveness of the proposed model compared to existing attention-based methods, and its competitiveness against other state-of-the-art supervised summarization approaches. An ablation study that focuses on our main proposed components, namely the use of global and local multi-head attention mechanisms in collaboration with an absolute positional encoding component, shows their relative contributions to the overall summarization performance.
Evlampios Apostolidis, Georgios Balaouras, Vasileios Mezaris, Ioannis Patras
ISM3
2021 A Web Service for Video Smart-Cropping
abstract
This paper presents a Web service that supports the automatic transformation of a video's aspect ratio. We employ a modified smart-cropping technique from the literature that aims to minimize the loss of semantically important visual content. We integrate this method in an easy-to-use publicly- accessible Web service where a video can be uploaded and automatically transformed to the desired aspect ratio. We also demonstrate that the algorithmic modifications we introduced in the process of building our Web service offer performance gains, when compared to the original method of the literature.
Konstantinos Apostolidis, Vasileios Mezaris
ISM2
2021 Combining Adversarial and Reinforcement Learning for Video Thumbnail Selection
abstract
This paper presents a new method for unsupervised video thumbnail selection. The developed network architecture selects video thumbnails based on two criteria: the representativeness and the aesthetic quality of their visual content. Training relies on a combination of adversarial and reinforcement learning. The former is used to train a discriminator, whose goal is to distinguish the original from a reconstructed version of the video based on a small set of candidate thumbnails. The discriminator's feedback is a measure of the representativeness of the selected thumbnails. This measure is combined with estimates about the aesthetic quality of the thumbnails (made using a SoA Fully Convolutional Network) to form a reward and train the thumbnail selector via reinforcement learning. Experiments on two datasets (OVP and Youtube) show the competitiveness of the proposed method against other SoA approaches. An ablation study with respect to the adopted thumbnail selection criteria documents the importance of considering the aesthetics, and the contribution of this information when used in combination with measures about the representativeness of the visual content.
Evlampios Apostolidis, Eleni Adamantidou, Vasileios Mezaris, Ioannis Patras
ICMR3
2021 VERGE in VBS 2021
Stelios Andreadis, Anastasia Moumtzidou, Konstantinos Gkountakos, Nick Pantelidis, Konstantinos Apostolidis, Damianos Galanopoulos, Ilias Gialampoukidis, Stefanos Vrochidis, Vasileios Mezaris, Ioannis Kompatsiaris
MMM (2)9
2021 Content Wizard: demo of a trans-vector digital video publication tool
abstract
In order to optimise the distribution of video assets online, media organizations need tailor their offerings for specific digital channels and better understand the interests of their audiences at particular points in time, which are often influenced by contemporary new stories and trends on social media. For this purpose, the research project ReTV has developed a Web-based tool termed ’Content Wizard’ which demonstrates an end-to-end, semi-automated workflow for video content creation, adaptation and distribution across digital channels. Digital assets can be selected based on predicted future trending topics, re-purposed according to the different digital channels they will be published upon and scheduled for the optimal future publication date. The result is an innovative video publication workflow that meets the marketing needs of media organisations in this age of transient online media spread across multiple channels.
Lyndon J. B. Nixon, Konstantinos Apostolidis, Evlampios Apostolidis, Damianos Galanopoulos, Vasileios Mezaris, Basil Philipp, Rasa Bocyte
IMX5
2021 Video Summarization Using Deep Neural Networks: A Survey
abstract
Video summarization technologies aim to create a concise and complete synopsis by selecting the most informative parts of the video content. Several approaches have been developed over the last couple of decades, and the current state of the art is represented by methods that rely on modern deep neural network architectures. This work focuses on the recent advances in the area and provides a comprehensive survey of the existing deep-learning-based methods for generic video summarization. After presenting the motivation behind the development of technologies for video summarization, we formulate the video summarization task and discuss the main characteristics of a typical deep-learning-based analysis pipeline. Then, we suggest a taxonomy of the existing algorithms and provide a systematic review of the relevant literature that shows the evolution of the deep-learning-based video summarization technologies and leads to suggestions for future developments. We then report on protocols for the objective evaluation of video summarization algorithms, and we compare the performance of several deep-learning-based approaches. Based on the outcomes of these comparisons, as well as some documented considerations about the amount of annotated data and the suitability of evaluation protocols, we indicate potential future research directions.
Evlampios Apostolidis, Eleni Adamantidou, Alexandros I. Metsai, Vasileios Mezaris, Ioannis Patras
Proc. IEEE4
2021 AC-SUM-GAN: Connecting Actor-Critic and Generative Adversarial Networks for Unsupervised Video Summarization
abstract
This paper presents a new method for unsupervised video summarization. The proposed architecture embeds an Actor-Critic model into a Generative Adversarial Network and formulates the selection of important video fragments (that will be used to form the summary) as a sequence generation task. The Actor and the Critic take part in a game that incrementally leads to the selection of the video key-fragments, and their choices at each step of the game result in a set of rewards from the Discriminator. The designed training workflow allows the Actor and Critic to discover a space of actions and automatically learn a policy for key-fragment selection. Moreover, the introduced criterion for choosing the best model after the training ends, enables the automatic selection of proper values for parameters of the training process that are not learned from the data (such as the regularization factor σ). Experimental evaluation on two benchmark datasets (SumMe and TVSum) demonstrates that the proposed AC-SUM-GAN model performs consistently well and gives SoA results in comparison to unsupervised methods, that are also competitive with respect to supervised methods.
Evlampios Apostolidis, Eleni Adamantidou, Alexandros I. Metsai, Vasileios Mezaris, Ioannis Patras
IEEE Trans. Circuits Syst. Video Technol.4
2020 Structured Pruning of LSTMs via Eigenanalysis and Geometric Median for Mobile Multimedia and Deep Learning Applications
abstract
In this paper, a novel structured pruning approach for learning efficient long short-term memory (LSTM) network architectures is proposed. More specifically, the eigenvalues of the covariance matrix associated with the responses of each LSTM layer are computed and utilized to quantify the layers' redundancy and automatically obtain an individual pruning rate for each layer. Subsequently, a Geometric Median based (GM-based) criterion is used to identify and prune in a structured way the most redundant LSTM units, realizing the pruning rates derived in the previous step. The experimental evaluation on the Penn Treebank text corpus and the large-scale YouTube-8M audio-video dataset for the tasks of word-level prediction and visual concept detection, respectively, shows the efficacy of the proposed approach1.
Nikolaos Gkalelis, Vasileios Mezaris
ISM2
2020 Attention Mechanisms, Signal Encodings and Fusion Strategies for Improved Ad-hoc Video Search with Dual Encoding Networks
abstract
In this paper, the problem of unlabeled video retrieval using textual queries is addressed. We present an extended dual encoding network which makes use of more than one encodings of the visual and textual content, as well as two different attention mechanisms. The latter serve the purpose of highlighting temporal locations in every modality that can contribute more to effective retrieval. The different encodings of the visual and textual inputs, along with early/late fusion strategies, are examined for further improving performance. Experimental evaluations and comparisons with state-of-the-art methods document the merit of the proposed network.
Damianos Galanopoulos, Vasileios Mezaris
ICMR2
2020 Performance over Random: A Robust Evaluation Protocol for Video Summarization Methods
abstract
This paper proposes a new evaluation approach for video summarization algorithms. We start by studying the currently established evaluation protocol; this protocol, defined over the ground-truth annotations of the SumMe and TVSum datasets, quantifies the agreement between the user-defined and the automatically-created summaries with F-Score, and reports the average performance on a few different training/testing splits of the used dataset. We evaluate five publicly-available summarization algorithms under a large-scale experimental setting with 50 randomly-created data splits. We show that the results reported in the papers are not always congruent with their performance on the large-scale experiment, and that the F-Score cannot be used for comparing algorithms evaluated on different splits. We also show that the above shortcomings of the established evaluation protocol are due to the significantly varying levels of difficulty among the utilized splits, that affect the outcomes of the evaluations. Further analysis of these findings indicates a noticeable performance correlation among all algorithms and a random summarizer. To mitigate these shortcomings we propose an evaluation protocol that makes estimates about the difficulty of each used data split and utilizes this information during the evaluation process. Experiments involving different evaluation settings demonstrate the increased representativeness of performance results when using the proposed evaluation approach, and the increased reliability of comparisons when the examined methods have been evaluated on different data splits.
Evlampios Apostolidis, Eleni Adamantidou, Alexandros I. Metsai, Vasileios Mezaris, Ioannis Patras
ACM Multimedia4
2020 AI4TV 2020: 2nd International Workshop on AI for Smart TV Content Production, Access and Delivery
abstract
Technological developments in comprehensive video understanding - detecting and identifying visual elements of a scene, combined with audio understanding (music, speech), as well as aligned with textual information such as captions, subtitles, etc. and background knowledge - have been undergoing a significant revolution during recent years. The workshop brings together experts from academia and industry in order to discuss the latest progress in artificial intelligence research in topics related to multimodal information analysis, and in particular, semantic analysis of video, audio, and textual information for smart digital TV content production, access and delivery.
Raphaël Troncy, Jorma Laaksonen, Hamed Rezazadegan Tavakoli, Lyndon J. B. Nixon, Vasileios Mezaris, Mohammad Hosseini 0002
ACM Multimedia5
2020 VERGE in VBS 2020
Stelios Andreadis, Anastasia Moumtzidou, Konstantinos Apostolidis, Konstantinos Gkountakos, Damianos Galanopoulos, Emmanouil Michail, Ilias Gialampoukidis, Stefanos Vrochidis, Vasileios Mezaris, Ioannis Kompatsiaris
MMM (2)9
2020 Unsupervised Video Summarization via Attention-Driven Adversarial Learning
Evlampios Apostolidis, Eleni Adamantidou, Alexandros I. Metsai, Vasileios Mezaris, Ioannis Patras
MMM (1)4
2020 Subclass Deep Neural Networks: Re-enabling Neglected Classes in Deep Network Training for Multimedia Classification
Nikolaos Gkalelis, Vasileios Mezaris
MMM (1)2
2020 A Web Service for Video Summarization
abstract
This paper presents a Web service that supports the automatic generation of video summaries for user-submitted videos. The developed Web application decomposes the video into segments, evaluates the fitness of each segment to be included in the video summary and selects appropriate segments until a pre-defined time budget is filled. The integrated deep-learning-based video analysis and summarization technologies exhibit state-of-the-art performance and, by exploiting the processing capabilities of modern GPUs, offer faster than real-time processing. Configurations for generating video summaries that fulfill the specifications for posting on the most common video sharing platforms and social networks are available in the user interface of this application, enabling the one-click generation of distribution-channel-specific summaries.
Chrysa Collyda, Konstantinos Apostolidis, Evlampios Apostolidis, Eleni Adamantidou, Alexandros I. Metsai, Vasileios Mezaris
IMX6
2019 AI4TV 2019: 1st International Workshop on AI for Smart TV Content Production, Access and Delivery
abstract
Technological developments in comprehensive video understanding - detecting and identifying visual elements of a scene, combined with audio understanding (music, speech), as well as aligned with textual information such as captions, subtitles, etc. and background knowledge - have been undergoing a significant revolution during recent years. The workshop brings together experts from academia and industry in order to discuss the latest progress in artificial intelligence research in topics related to multimodal information analysis, and in particular, semantic analysis of video, audio, and textual information for smart digital TV content production, access and delivery.
Raphaël Troncy, Jorma Laaksonen, Hamed Rezazadegan Tavakoli, Lyndon J. B. Nixon, Vasileios Mezaris
ACM Multimedia5
2019 VERGE in VBS 2019
Stelios Andreadis, Anastasia Moumtzidou, Damianos Galanopoulos, Fotini Markatopoulou, Konstantinos Apostolidis, Thanassis Mavropoulos, Ilias Gialampoukidis, Stefanos Vrochidis, Vasileios Mezaris, Ioannis Kompatsiaris, Ioannis Patras
MMM (2)9
2019 Image Aesthetics Assessment Using Fully Convolutional Neural Networks
Konstantinos Apostolidis, Vasileios Mezaris
MMM (1)2
2019 Temporal Lecture Video Fragmentation Using Word Embeddings
Damianos Galanopoulos, Vasileios Mezaris
MMM (2)2
2019 Multimodal Video Annotation for Retrieval and Discovery of Newsworthy Video in a News Verification Scenario
Lyndon J. B. Nixon, Evlampios Apostolidis, Fotini Markatopoulou, Ioannis Patras, Vasileios Mezaris
MMM (1)5
2019 Training Researchers with the MOVING Platform
Iacopo Vagliano, Angela Fessl, Franziska Günther, Thomas Köhler 0001, Vasileios Mezaris, Ahmed Saleh 0003, Ansgar Scherp, Ilija Simic
MMM (2)5
2019 Detecting Tampered Videos with Multimedia Forensics and Deep Learning
Markos Zampoglou, Fotini Markatopoulou, Grégoire Mercier, Despoina Touska, Evlampios Apostolidis, Symeon Papadopoulos, Roger Cozien, Ioannis Patras, Vasileios Mezaris, Ioannis Kompatsiaris
MMM (1)9
2019 A deep generic to specific recognition model for group membership analysis using non-verbal cues
abstract
Automatic understanding and analysis of groups has attracted increasing attention in the vision and multimedia communities in recent years. However, little attention has been paid to the automatic analysis of the non-verbal behaviors and how this can be utilized for analysis of group membership, i.e., recognizing which group each individual is part of. This paper presents a novel Support Vector Machine (SVM) based Deep Specific Recognition Model (DeepSRM) that is learned based on a generic recognition model . The generic recognition model refers to the model trained with data across different conditions, i.e., when people are watching movies of different types. Although the generic recognition model can provide a baseline for the recognition model trained for each specific condition, the different behaviors people exhibit in different conditions limit the recognition performance of the generic model. Therefore, the specific recognition model is proposed for each condition separately and built on top of the generic recognition model . A number of experiments are conducted using a database aiming to study group analysis while each group (i.e., four participants together) were watching a number of long movie segments. Our experimental results show that the proposed deep specific recognition model (44%) outperforms the generic recognition model (26%). The recognition of group membership also indicates that the non-verbal behaviors of individuals within a group share commonalities.
Wenxuan Mou, Christos Tzelepis, Vasileios Mezaris, Hatice Gunes, Ioannis Patras
Image Vis. Comput.3
2019 Implicit and Explicit Concept Relations in Deep Neural Networks for Multi-Label Video/Image Annotation
abstract
In this paper, we propose a deep convolutional neural network (DCNN) architecture that addresses the problem of video/image concept annotation by exploiting concept relations at two different levels. At the first level, we build on ideas from multi-task learning, and propose an approach to learn concept-specific representations that are sparse, linear combinations of representations of latent concepts. By enforcing the sharing of the latent concept representations, we exploit the implicit relations between the target concepts. At a second level, we build on ideas from structured output learning and propose the introduction, at training time, of a new cost term that explicitly models the correlations between the concepts. By doing so, we explicitly model the structure in the output space (i.e., the concept labels). Both of the above are implemented using standard convolutional layers and are incorporated in a single DCNN architecture that can then be trained end-to-end with standard back-propagation. Experiments on four large-scale video and image data sets show that the proposed DCNN improves concept annotation accuracy and outperforms the related state-of-the-art methods.
Fotini Markatopoulou, Vasileios Mezaris, Ioannis Patras
IEEE Trans. Circuits Syst. Video Technol.2
2018 A Motion-Driven Approach for Fine-Grained Temporal Segmentation of User-Generated Videos
Konstantinos Apostolidis, Evlampios Apostolidis, Vasileios Mezaris
MMM (1)3
2018 VERGE in VBS 2018
Anastasia Moumtzidou, Stelios Andreadis, Fotini Markatopoulou, Damianos Galanopoulos, Ilias Gialampoukidis, Stefanos Vrochidis, Vasileios Mezaris, Ioannis Kompatsiaris, Ioannis Patras
MMM (2)7
2018 Linear Maximum Margin Classifier for Learning from Uncertain Data
abstract
In this paper, we propose a maximum margin classifier that deals with uncertainty in data input. More specifically, we reformulate the SVM framework such that each training example can be modeled by a multi-dimensional Gaussian distribution described by its mean vector and its covariance matrix-the latter modeling the uncertainty. We address the classification problem and define a cost function that is the expected value of the classical SVM cost when data samples are drawn from the multi-dimensional Gaussian distributions that form the set of the training examples. Our formulation approximates the classical SVM formulation when the training examples are isotropic Gaussians with variance tending to zero. We arrive at a convex optimization problem, which we solve efficiently in the primal form using a stochastic gradient descent approach. The resulting classifier, which we name SVM with Gaussian Sample Uncertainty (SVM-GSU), is tested on synthetic data and five publicly available and popular datasets; namely, the MNIST, WDBC, DEAP, TV News Channel Commercial Detection, and TRECVID MED datasets. Experimental results verify the effectiveness of the proposed method.
Christos Tzelepis, Vasileios Mezaris, Ioannis Patras
IEEE Trans. Pattern Anal. Mach. Intell.2
2017 Generic to Specific Recognition Models for Membership Analysis in Group Videos
abstract
Automatic understanding and analysis of groups has attracted increasing attention in the vision and multimedia communities in recent years. However, little attention has been paid to the automatic analysis of group membership - i.e., recognizing which group the individual in question is part of. This paper presents a novel two-phase Support Vector Machine (SVM) based specific recognition model that is learned using an optimized generic recognition model. We conduct a set of experiments using a database collected to study group analysis from multimodal cues while each group (i.e., four participants together) were watching a number of long movie segments. Our experimental results show that the proposed specific recognition model (52%) outperforms the generic recognition model trained across all different videos (35%) and the independent recognition model trained directly on each specific video (33%) using linear SVM.
Wenxuan Mou, Christos Tzelepis, Vasileios Mezaris, Hatice Gunes, Ioannis Patras
FG3
2017 Expo: An Expectation-oriented System for Selecting Important Photos from Personal Collections
abstract
The diffusion of digital photography lets people take hundreds of photos during personal events, such as trips and ceremonies. Many methods have been developed for summarizing such large personal photo collections. However, they usually emphasize the coverage of the original collection, without considering which photos users would select, i.e. their expectations. In this paper we present Expo, a system that aims at selecting which photos users perceive as most important and would have selected, thus meeting their expectations. It does not rely on any manually provided annotation, thus keeping the effort of users low. Photos are processed by applying a wide set of image processing techniques and a subset of a required size is selected. Users can review and modify the automatic selection. The system can also be used to gather training data by letting users select their preferred photos from the imported collections.
Andrea Ceroni, Vassilios Solachidis, Claudia Niederée, Olga Papadopoulou, Vasileios Mezaris
ICMR5
2017 VideoAnalysis4ALL: An On-line Tool for the Automatic Fragmentation and Concept-based Annotation, and the Interactive Exploration of Videos
abstract
This paper presents the VideoAnalysis4ALL tool that supports the automatic fragmentation and concept-based annotation of videos, and the exploration of the annotated video fragments through an interactive user interface. The developed web application decomposes the video into two different granularities, namely shots and scenes, and annotates each fragment by evaluating the existence of a number (several hundreds) of high-level visual concepts in the keyframes extracted from these fragments. Through the analysis the tool enables the identification and labeling of semantically coherent video fragments, while its user interfaces allow the discovery of these fragments with the help of human-interpretable concepts. The integrated state-of-the-art video analysis technologies perform very well and, by exploiting the processing capabilities of multi-thread / multi-core architectures, reduce the time required for analysis to approximately one third of the video's duration, thus making the analysis three times faster than real-time processing.
Chrysa Collyda, Evlampios Apostolidis, Alexandros Pournaras, Fotini Markatopoulou, Vasileios Mezaris, Ioannis Patras
ICMR5
2017 Concept Language Models and Event-based Concept Number Selection for Zero-example Event Detection
abstract
Zero-example event detection is a problem where, given an event query as input but no example videos for training a detector, the system retrieves the most closely related videos. In this paper we present a fully-automatic zero-example event detection method that is based on translating the event description to a predefined set of concepts for which previously trained visual concept detectors are available. We adopt the use of Concept Language Models (CLMs), which is a method of augmenting semantic concept definition, and we propose a new concept-selection method for deciding on the appropriate number of the concepts needed to describe an event query. The proposed system achieves state-of-the-art performance in automatic zero-example event detection.
Damianos Galanopoulos, Fotini Markatopoulou, Vasileios Mezaris, Ioannis Patras
ICMR3
2017 Query and Keyframe Representations for Ad-hoc Video Search
abstract
This paper presents a fully-automatic method that combines video concept detection and textual query analysis in order to solve the problem of ad-hoc video search. We present a set of NLP steps that cleverly analyse different parts of the query in order to convert it to related semantic concepts, we propose a new method for transforming concept-based keyframe and query representations into a common semantic embedding space, and we show that our proposed combination of concept-based representations with their corresponding semantic embeddings results to improved video search accuracy. Our experiments in the TRECVID AVS 2016 and the Video Search 2008 datasets show the effectiveness of the proposed method compared to other similar approaches.
Fotini Markatopoulou, Damianos Galanopoulos, Vasileios Mezaris, Ioannis Patras
ICMR3
2017 Incremental Accelerated Kernel Discriminant Analysis
abstract
In this paper a novel incremental dimensionality reduction (DR) technique called incremental accelerated kernel discriminant analysis (IAKDA) is proposed. Consisting of the eigenvalue decomposition of a relatively small-size matrix and the recursive block Cholesky factorization of the kernel matrix, a nonlinear DR transformation is efficiently computed at each incremental step. Moreover, employing factorization techniques of excellent numerical stability, IAKDA effectively removes data nonlinearities in the low dimensional subspace. Experimental evaluation on various multimedia tasks and datasets confirms that the proposed approach combined with linear support vector machines (LSVMs) offers improved mean average precision (MAP) and provides an impressive training time speedup over batch KDA and also over traditional LSVM and kernel SVM (KSVM).
Nikolaos Gkalelis, Vasileios Mezaris
ACM Multimedia2
2017 MuVer'17: First International Workshop on Multimedia Verification
abstract
This paper gives an overview of the First International Workshop on Multimedia Verification, organized as part of the 2017 ACM Multimedia Conference. The paper outlines the current verification scene and needs, discusses the goals of the workshop, and presents the workshop's program, consisting of two invited keynote talks and three presentations of full papers that have been accepted at the workshop.
Vasileios Mezaris, Lyndon J. B. Nixon, Symeon Papadopoulos, Jochen Spangenberg
ACM Multimedia1
2017 MultiEdTech 2017: 1st International Workshop on Multimedia-based Educational and Knowledge Technologies for Personalized and Social Online Training
abstract
Educational and Knowledge Technologies (EdTech), especially in connection to multimedia content and the vision of mobile and personalized learning, is a hot topic in both academia and the business start-ups ecosystem. The driver and enabler of this is on the one side the development and widespread availability of multimedia materials and MOOCs, which represent multimedia content produced specifically for supporting e-learning; and, on the other side, the ever increasing availability of all sorts on information on the Internet and in social media channels (e. g. lectures, research papers, user-generated videos, news items), which, despite not directly targeting e-learning, can prove to be valuable complements to the more targeted learning materials. Although the availability of such content is not a problem these days, finding the right content and associating different relevant pieces of multimedia so as to enable a comprehensive learning experience on a chosen subject is by no means a trivial task. This workshop provides research in areas related to multimedia-based educational and knowledge technologies and particularly on the use of multimedia search and retrieval, analysis and understanding, browsing, summarization, recommendation, and visualization technologies on multimedia content available in specialized learning platforms, the Web, mobile devices and/or social networks for supporting personalized and adaptive e-learning and training.
Ansgar Scherp, Vasileios Mezaris, Thomas Köhler 0001, Alex Hauptmann 0001
ACM Multimedia2
2017 VERGE in VBS 2017
Anastasia Moumtzidou, Theodoros Mironidis, Fotini Markatopoulou, Stelios Andreadis, Ilias Gialampoukidis, Damianos Galanopoulos, Anastasia Ioannidou, Stefanos Vrochidis, Vasileios Mezaris, Ioannis Kompatsiaris, Ioannis Patras
MMM (2)9
2017 Comparison of Fine-Tuning and Extension Strategies for Deep Convolutional Neural Networks
Nikiforos Pittaras, Fotini Markatopoulou, Vasileios Mezaris, Ioannis Patras
MMM (1)3
2017 A web-based tool for fast instance-level labeling of videos and the creation of spatiotemporal media fragments
Anastasia Ioannidou, Evlampios Apostolidis, Chrysa Collyda, Vasileios Mezaris
Multim. Tools Appl.4
2017 Automatic Synchronization of Multi-user Photo Galleries
abstract
In this paper we address the issue of photo galleries synchronization, where pictures related to the same event are collected by different users. Existing solutions to address the problem are usually based on unrealistic assumptions, like time consistency across photo galleries, and often heavily rely on heuristics, therefore limiting the applicability to real-world scenarios. We propose a solution that achieves better generalization performance for the synchronization task compared to the available literature. The method is characterized by three stages: at first, deep convolutional neural network features are used to assess the visual similarity among the photos; then, pairs of similar photos are detected across different galleries and used to construct a graph; eventually, a probabilistic graphical model is used to estimate the temporal offset of each pair of galleries, by traversing the minimum spanning tree extracted from this graph. The experimental evaluation is conducted on four publicly available datasets covering different types of events, demonstrating the strength of our proposed method. A thorough discussion of the obtained results is provided for a critical assessment of the quality in synchronization.
Emanuele Sansone, Konstantinos Apostolidis, Nicola Conci, Giulia Boato, Vasileios Mezaris, Francesco G. B. De Natale
IEEE Trans. Multim.5
2017 Introduction to Special Issue on Deep Learning for Mobile Multimedia
abstract
No abstract available.
Kaoru Ota, Minh-Son Dao, Vasileios Mezaris, Francesco G. B. De Natale
ACM Trans. Multim. Comput. Commun. Appl.3
2017 Deep Learning for Mobile Multimedia: A Survey
abstract
Deep Learning (DL) has become a crucial technology for multimedia computing. It offers a powerful instrument to automatically produce high-level abstractions of complex multimedia data, which can be exploited in a number of applications, including object detection and recognition, speech-to- text, media retrieval, multimodal data analysis, and so on. The availability of affordable large-scale parallel processing architectures, and the sharing of effective open-source codes implementing the basic learning algorithms, caused a rapid diffusion of DL methodologies, bringing a number of new technologies and applications that outperform, in most cases, traditional machine learning technologies. In recent years, the possibility of implementing DL technologies on mobile devices has attracted significant attention. Thanks to this technology, portable devices may become smart objects capable of learning and acting. The path toward these exciting future scenarios, however, entangles a number of important research challenges. DL architectures and algorithms are hardly adapted to the storage and computation resources of a mobile device. Therefore, there is a need for new generations of mobile processors and chipsets, small footprint learning and inference algorithms, new models of collaborative and distributed processing, and a number of other fundamental building blocks. This survey reports the state of the art in this exciting research area, looking back to the evolution of neural networks, and arriving to the most recent results in terms of methodologies, technologies, and applications for mobile environments.
Kaoru Ota, Minh-Son Dao, Vasileios Mezaris, Francesco G. B. De Natale
ACM Trans. Multim. Comput. Commun. Appl.3
2016 Online multi-task learning for semantic concept detection in video
abstract
In this paper we propose an online multi-task learning algorithm for video concept detection. In particular, we extend the Efficient Lifelong Learning Algorithm (ELLA) in the following ways: (a) we solve the objective function of ELLA using quadratic programming instead of solving the Lasso problem, (b) we add a new label-based constraint that considers concept correlations, (c) we use linear SVMs as base learners instead of logistic regression. Experimental results show improvement over both the single-task learning methods typically used in this problem and the original ELLA algorithm.
Fotini Markatopoulou, Vasileios Mezaris, Ioannis Patras
ICIP2
2016 Video aesthetic quality assessment using kernel Support Vector Machine with isotropic Gaussian sample uncertainty (KSVM-IGSU)
abstract
In this paper we propose a video aesthetic quality assessment method that combines the representation of each video according to a set of photographic and cinematographic rules, with the use of a learning method that takes the video representation's uncertainty into consideration. Specifically, our method exploits the information derived from both low- and high-level analysis of video layout, leading to a photo- and motion-based video representation scheme. Subsequently, a kernel Support Vector Machine (SVM) extension, the KSVM-iGSU, is trained to classify the videos and retrieve those of high aesthetic value. Experimental results on our large dataset verify the effectiveness of the proposed method. We also make publicly available our dataset, in order to facilitate research in the area of video aesthetic quality assessment.
Christos Tzelepis, Eftichia Mavridaki, Vasileios Mezaris, Ioannis Patras
ICIP3
2016 AKSDA-MSVM: A GPU-accelerated Multiclass Learning Framework for Multimedia
abstract
In this paper, a combined nonlinear dimensionality reduction and multiclass classification framework is proposed. Specifically, a novel discriminant analysis (DA) technique, called accelerated kernel subclass discriminant analysis (AKSDA), derives a discriminant subspace, and a linear multiclass support vector machine (MSVM) computes a set of separating hyperplanes in the derived subspace. Moreover, within this framework an approach for accelerating the computation of multiple Gram matrices and an associated late fusion scheme are presented. Experimental evaluation in five multimedia datasets, on tasks such as video event detection and news document classification, shows that the proposed framework achieves excellent results in terms of both training time and generalization performance.
Stavros Arestis-Chartampilas, Nikolaos Gkalelis, Vasileios Mezaris
ACM Multimedia3
2016 Deep Multi-task Learning with Label Correlation Constraint for Video Concept Detection
abstract
In this work we propose a method that integrates multi-task learning (MTL) and deep learning. Our method appends a MTL-like loss to a deep convolutional neural network, in order to learn the relations between tasks together at the same time, and also incorporates the label correlations between pairs of tasks. We apply the proposed method on a transfer learning scenario, where our objective is to fine-tune the parameters of a network that has been originally trained on a large-scale image dataset for concept detection, so that it be applied on a target video dataset and a corresponding new set of target concepts. We evaluate the proposed method for the video concept detection problem on the TRECVID 2013 Semantic Indexing dataset. Our results show that the proposed algorithm leads to better concept-based video annotation than existing state-of-the-art methods.
Fotini Markatopoulou, Vasileios Mezaris, Ioannis Patras
ACM Multimedia2
2016 Ordering of Visual Descriptors in a Classifier Cascade Towards Improved Video Concept Detection
Fotini Markatopoulou, Vasileios Mezaris, Ioannis Patras
MMM (1)2
2016 VERGE: A Multimodal Interactive Search Engine for Video Browsing and Retrieval
Anastasia Moumtzidou, Theodoros Mironidis, Evlampios Apostolidis, Fotini Markatopoulou, Anastasia Ioannidou, Ilias Gialampoukidis, Konstantinos Avgerinakis, Stefanos Vrochidis, Vasileios Mezaris, Ioannis Kompatsiaris, Ioannis Patras
MMM (2)9
2016 Video Event Detection Using Kernel Support Vector Machine with Isotropic Gaussian Sample Uncertainty (KSVM-iGSU)
Christos Tzelepis, Vasileios Mezaris, Ioannis Patras
MMM (1)2
2016 Learning to detect video events from zero or very few video examples
Christos Tzelepis, Damianos Galanopoulos, Vasileios Mezaris, Ioannis Patras
Image Vis. Comput.3
2016 Event-based media processing and analysis: A survey of the literature
abstract
Research on event-based processing and analysis of media is receiving an increasing attention from the scientific community due to its relevance for an abundance of applications, from consumer video management and video surveillance to lifelogging and social media.Events have the ability to semantically encode relationships of different informational modalities, such as visual-audio-text, time, involved agents and objects, with the spatio-temporal component of events being a key feature for contextual analysis.This unveils an enormous potential for exploiting new information sources and opening new research directions.In this paper, we survey the existing literature in this field.We extensively review the employed conceptualization of the notion of event in multimedia, the techniques for event representation and modeling, the feature representation and event inference approaches for the problems of event detection in audio, visual, and textual content.Furthermore, we review some key event-based multimedia applications, and various benchmarking activities that provide solid frameworks for measuring the performance of different event processing and analysis systems.We provide an in-depth discussion of the insights obtained from reviewing the literature and identify future directions and challenges.
Christos Tzelepis, Zhigang Ma, Vasileios Mezaris, Bogdan Ionescu, Ioannis Kompatsiaris, Giulia Boato, Nicu Sebe, Shuicheng Yan
Image Vis. Comput.3
2016 Special issue on multimedia in ecology
Concetto Spampinato, Vasileios Mezaris, Jacco van Ossenbruggen
Multim. Syst.2
2016 Special issue on "Fine-grained categorization in ecological multimedia"
Concetto Spampinato, Vasileios Mezaris, Marco Cristani
Pattern Recognit. Lett.2
2015 Concept Detection in Multimedia Web Resources About Home Made Explosives
abstract
This work investigates the effectiveness of a state-of-the-art concept detection framework for the automatic classification of multimedia content, namely images and videos, embedded in publicly available Web resources containing recipes for the synthesis of Home Made Explosives (HMEs), to a set of predefined semantic concepts relevant to the HME domain. The concept detection framework employs advanced methods for video (shot) segmentation, visual feature extraction (using SIFT, SURF, and their variations), and classification based on machine learning techniques (logistic regression). The evaluation experiments are performed using an annotated collection of multimedia HME content discovered on the Web, and a set of concepts, which emerged both from an empirical study, and were also provided by domain experts and interested stakeholders, including Law Enforcement Agencies personnel. The experiments demonstrate the satisfactory performance of our framework, which in turn indicates the significant potential of the adopted approaches on the HME domain.
George Kalpakis, Theodora Tsikrika, Fotini Markatopoulou, Nikiforos Pittaras, Stefanos Vrochidis, Vasileios Mezaris, Ioannis Patras, Ioannis Kompatsiaris
ARES6
2015 Cascade of classifiers based on binary, non-binary and deep convolutional network descriptors for video concept detection
abstract
In this paper we propose a cascade architecture that can be used to train and combine different visual descriptors (local binary, local non-binary and Deep Convolutional Neural Network-based) for video concept detection. The proposed architecture is computationally more efficient than typical state-of-the-art video concept detection systems, without affecting the detection accuracy. In addition, this work presents a detailed study on combining descriptors based on Deep Convolutional Neural Networks with other popular local descriptors, both within a cascade and when using different late-fusion schemes. We evaluate our methods on the extensive video dataset of the 2013 TRECVID Semantic Indexing Task.
Fotini Markatopoulou, Vasileios Mezaris, Ioannis Patras
ICIP2
2015 A comprehensive aesthetic quality assessment method for natural images using basic rules of photography
abstract
In this paper we propose a comprehensive photo aesthetic assessment method that represents each photo according to a set of photographic rules. Specifically, our method exploits the information derived from both low- and high-level analysis of photo layout, not only for the photo as a whole, but also for specific spatial regions of it. Five feature vectors are introduced for describing the photo's simplicity, colorfulness, sharpness, pattern and composition. Subsequently, they are concatenated in a final feature vector where a Support Vector Machine (SVM) classifier is applied in order to perform the aesthetic quality evaluation. The experimental results and comparisons show that our approach achieves consistently more accurate quality assessment than the relevant literature methods and also that the proposed features can be combined with other generic image features so as to enhance the performance of previous methods.
Eftichia Mavridaki, Vasileios Mezaris
ICIP2
2015 To Keep or not to Keep: An Expectation-oriented Photo Selection Method for Personal Photo Collections
abstract
When selecting important photos from a personal photo collection - e.g. for creating an enjoyable sub-collection for revisiting or preservation - photos are not considered in isolation. Therefore, collection-level criteria are also taken into account by automated photo selection methods. However, the typical two-step process of first clustering and subsequently picking from the clusters seems to overstress coverage as a criterion when applied to the task of selecting the photos most important to a user. We, therefore, propose a novel expectation-oriented photo selection method, which combines a variety of collection-level and image-level selection criteria in a flexible way. In our evaluation, which is based on large real-world personal photo collections with overall more than 18,000 images, we show that our method outperforms state-of-the-art photo selection methods. In addition, the proposed method does not rely on any manual annotations, making it applicable in realistic settings of personal photo collections.
Andrea Ceroni, Vassilios Solachidis, Claudia Niederée, Olga Papadopoulou, Nattiya Kanhabua, Vasileios Mezaris
ICMR6
2015 Exploiting Multiple Web Resources towards Collecting Positive Training Samples for Visual Concept Learning
abstract
The number of images uploaded to the web is enormous and is rapidly increasing. The purpose of our work is to use these for acquiring positive training data for visual concept learning. Manually creating training data for visual concept classifiers is an expensive and time consuming task. We propose an approach which automatically collects positive training samples from the Web by constructing a multitude of text queries and retaining for each query only very few top-ranked images returned by each one of the different web image search engines (Google, Flickr and Bing). In this way, we sift the burden of false positive rejection to the Web search engines and directly assemble a rich set of high-quality positive training samples. Experiments on forty concepts, evaluated on the ImageNet dataset, show the merit of the proposed approach.
Olga Papadopoulou, Vasileios Mezaris
ICMR2
2015 GPU Accelerated Generalised Subclass Discriminant Analysis for Event and Concept Detection in Video
abstract
In this paper a discriminant analysis (DA) technique called accelerated generalised subclass discriminant analysis (AGSDA) and its GPU implementation are presented. This method identifies a discriminant subspace of the input space in three steps: a) Gram matrix computation, b) eigenvalue decomposition of the between subclass factor matrix, and c) computation of the solution of a linear matrix system with symmetric positive semidefinite (SPSD) matrix of coefficients. Based on the fact that the computationally intensive parts of AGSDA, i.e. Gram matrix computation and identification of the SPSD linear matrix system solution, are highly parallelisable, a GPU implementation of AGSDA is proposed. Experimental results on large-scale datasets of TRECVID for event and concept detection show that our GPU-AGSDA method combined with LSVM outperforms LSVM alone in training time, memory consumption, and detection accuracy.
Stavros Arestis-Chartampilas, Nikolaos Gkalelis, Vasileios Mezaris
ACM Multimedia3
2015 About Events, Objects, and their Relationships: Human-centered Event Understanding from Multimedia
abstract
HuEvent'15 is a continuation of previous year's successful workshop on events in multimedia. It focuses on the human-centered aspects of understanding events from multimedia content. This includes the notion of objects and their relation to events. The workshop brings together researchers from the different areas in multimedia and beyond that are interested in understanding the concept of events.
Ansgar Scherp, Vasileios Mezaris, Bogdan Ionescu, Francesco G. B. De Natale
ACM Multimedia2
2015 A Study on the Use of a Binary Local Descriptor and Color Extensions of Local Descriptors for Video Concept Detection
Fotini Markatopoulou, Nikiforos Pittaras, Olga Papadopoulou, Vasileios Mezaris, Ioannis Patras
MMM (1)4
2015 VERGE: A Multimodal Interactive Video Search Engine
Anastasia Moumtzidou, Konstantinos Avgerinakis, Evlampios Apostolidis, Fotini Markatopoulou, Konstantinos Apostolidis, Theodoros Mironidis, Stefanos Vrochidis, Vasileios Mezaris, Ioannis Kompatsiaris, Ioannis Patras
MMM (2)8
2014 Fast shot segmentation combining global and local visual descriptors
abstract
This paper introduces an algorithm for fast temporal segmentation of videos into shots. The proposed method detects abrupt and gradual transitions, based on the visual similarity of neighboring frames of the video. The descriptive efficiency of both local (SURF) and global (HSV histograms) descriptors is exploited for assessing frame similarity, while GPU-based processing is used for accelerating the analysis. Specifically, abrupt transitions are initially detected between successive video frames where there is a sharp change in the visual content, which is expressed by a very low similarity score. Then, the calculated scores are further analysed for the identification of frame-sequences where a progressive change of the visual content takes place and, in this way gradual transitions are detected. Finally, a post-processing step is performed aiming to identify outliers due to object/camera movement and flash-lights. The experiments show that the proposed algorithm achieves high accuracy while being capable of faster-than-real-time analysis.
Evlampios Apostolidis, Vasileios Mezaris
ICASSP2
2014 No-reference blur assessment in natural images using Fourier transform and spatial pyramids
abstract
In this paper we propose a no-reference image blur assessment model that performs partial blur detection in the frequency domain. Specifically, our method exploits the information derived from the power spectrum of the Fourier transform. The latter is computed for both the entire image and several patches of it, in order to estimate the distribution of low and high frequencies, and is appropriately encoded so as to preserve some information about the spatial arrangement of the frequency distribution in the image. Finally, a Support Vector Machine (SVM) classifier is applied to the above features, serving as the image blur quality evaluator. For a proper training and evaluation of the proposed method, we proceeded with creating and using a large image dataset consisting of more than 2400 digital photographs, which we make publicly available. The results show the efficiency of our method in assessing not only artificially-distorted images but also naturally-blurred ones.
Eftichia Mavridaki, Vasileios Mezaris
ICIP2
2014 Video event detection using generalized subclass discriminant analysis and linear support vector machines
abstract
In this paper, a two-phase approach to event detection in video is proposed. This combines a novel nonlinear Discriminant Analysis (DA) method called Generalized Subclass DA (GSDA), to identify a discriminant subspace, and a Linear Support Vector Machine (LSVM), to efficiently learn the event in the derived subspace. The proposed GSDA-LSVM framework is used as an alternative to the Kernel Support Vector Machine (KSVM) approach, which despite its excellent classification accuracy requires significant computational resources for learning the events (i.e., for identifying the kernel parameters and KSVM penalty term) in large-scale video collections. In contrary, using the GSDA-LSVM approach the SVM penalty term can be rapidly identified in the lower dimensional subspace. Moreover, an additional speed up in deriving this lower-dimensional space is achieved by using the proposed GSDA method instead of conventional nonlinear subclass DA methods such as KSDA or KMSDA. This is made possible by GSDA exploiting the special structure of the inter-between-subclass scatter matrix to reformulate the original KSDA eigenvalue problem to one involving matrices of much smaller dimension. The proposed GSDA-LSVM approach leads to more accurate event detection and to computational efficiency gains, as shown by experimental results on the extensive TRECVID MED 2010 and 2012 datasets.
Nikolaos Gkalelis, Vasileios Mezaris
ICMR2
2014 Automatic fine-grained hyperlinking of videos within a closed collection using scene segmentation
abstract
This paper introduces a framework for establishing links between related media fragments within a collection of videos. A set of analysis techniques is applied for extracting information from different types of data. Visual-based shot and scene segmentation is performed for defining media fragments at different granularity levels, while visual cues are detected from keyframes of the video via concept detection and optical character recognition (OCR). Keyword extraction is applied on textual data such as the output of OCR, subtitles and metadata. This set of results is used for the automatic identification and linking of related media fragments. The proposed framework exhibited competitive performance in the Video Hyperlinking sub-task of MediaEval 2013, indicating that video scene segmentation can provide more meaningful segments, compared to other decomposition methods, for hyperlinking purposes.
Evlampios Apostolidis, Vasileios Mezaris, Mathilde Sahuguet, Benoit Huet, Barbora Cervenková, Daniel Stein, Stefan Eickeler, José Luis Redondo García, Raphaël Troncy, Lukás Pikora
ACM Multimedia2
2014 Video hyperlinking
abstract
This is the abstract for the ``Video Hyperlinking'' tutorial, presented as part of the 2014 ACM Multimedia Conference. Video hyperlinking is the introduction of links that originate from pieces of video material and point to other relevant content, be it video or any other form of digital content. The tutorial presents the state of the art in video hyperlinking approaches and in relevant enabling technologies, such as video analysis and multimedia indexing and retrieval. Several alternative strategies, based on text, visual and/or audio information are introduced, evaluated and discussed, providing the audience with details on what works and what doesn't on real broadcast material.
Vasileios Mezaris, Benoit Huet
ACM Multimedia1
2014 HuEvent'14: 2014 workshop on human-centered event understanding from multimedia
abstract
This workshop focuses on the human-centered aspects of understanding events from multimedia content. This includes the notion of objects and their relation to events. The workshop brings together researchers from the different areas in multimedia and beyond that are interested in understanding the concept of events.
Ansgar Scherp, Vasileios Mezaris, Bogdan Ionescu, Francesco G. B. De Natale
ACM Multimedia2
2014 Summary Abstract for the 3rd ACM International Workshop on Multimedia Analysis for Ecological Data
abstract
The 3rd ACM International Workshop on Multimedia Anal- ysis for Ecological Data (MAED'14) is held as part of ACM Multimedia 2014.
Concetto Spampinato, Vasileios Mezaris, Marco Cristani
ACM Multimedia2
2014 A Comparative Study on the Use of Multi-label Classification Techniques for Concept-Based Video Indexing and Annotation
Fotini Markatopoulou, Vasileios Mezaris, Ioannis Kompatsiaris
MMM (1)2
2014 VERGE: An Interactive Search Engine for Browsing Video Collections
Anastasia Moumtzidou, Konstantinos Avgerinakis, Evlampios Apostolidis, Vera Aleksic, Fotini Markatopoulou, Christina Papagiannopoulou, Stefanos Vrochidis, Vasileios Mezaris, Reinhard Busch, Ioannis Kompatsiaris
MMM (2)8
2014 ReSEED: social event dEtection dataset
abstract
Nowadays, digital cameras are very popular among people and quite every mobile phone has a build-in camera. Social events have a prominent role in people's life. Thus, people take pictures of events they take part in and more and more of them upload these to well-known online photo community sites like Flickr. The number of pictures uploaded to these sites is still proliferating and there is a great interest in automatizing the process of event clustering so that every incoming (picture) document can be assigned to the corresponding event without the need of human interaction. These social events are defined as events that are planned by people, attended by people and for which the social multimedia are also captured by people. There is an urgent need to develop algorithms which are capable of grouping media by the social events they depict or are related to. In order to train, test, and evaluate such algorithms and frameworks, we present a dataset that consists of about 430,000 photos from Flickr together with the underlying ground truth consisting of about 21,000 social events. All the photos are accompanied by their textual metadata. The ground truth for the event groupings has been derived from event calendars on the Web that have been created collaboratively by people. The dataset has been used in the Social Event Detection (SED) task that was part of the MediaEval Benchmark for Multimedia Evaluation 2013. This task required participants to discover social events and organize the related media items in event-specific clusters within a collection of Web multimedia documents. In this paper we describe how the dataset has been collected and the creation of the ground truth together with a proposed evaluation methodology and a brief description of the corresponding task challenge as applied in the context of the Social Event Detection task.
Timo Reuter, Symeon Papadopoulos, Vasileios Mezaris, Philipp Cimiano
MMSys3
2014 Real-life events in multimedia: detection, representation, retrieval, and applications
Vasileios Mezaris, Ansgar Scherp, Ramesh Jain 0001, Mohan Kankanhalli
Multim. Tools Appl.1
2014 Survey on modeling and indexing events in multimedia
Ansgar Scherp, Vasileios Mezaris
Multim. Tools Appl.2
2014 Video Tomographs and a Base Detector Selection Strategy for Improving Large-Scale Video Concept Detection
abstract
In this paper, we deal with the problem of video concept detection to use the concept detection results toward a more effective concept-based video retrieval. The key novelties of this paper are as follows: 1) the use of spatio-temporal video slices (tomographs) in the same way that visual keyframes are typically used in video concept detection schemes. These spatio-temporal slices capture in a compact way motion patterns that are useful for detecting semantic concepts and are used for training a number of base detectors. The latter augment the set of keyframe-based base detectors that can be trained using different frame representations. 2) The introduction of a generic methodology, built upon a genetic algorithm, for controlling which subset of the available base detectors (consequently, which subset of the possible shot representations) should be combined for developing an optimal detector for each specific concept. This methodology is directly applicable to the learning of hundreds of diverse concepts, while diverging from the one-size-fits-all approach that is typically used in problems of this size. The proposed techniques are evaluated on the datasets of the 2011 and 2012 Semantic Indexing Task of TRECVID, each comprising several hundred hours of heterogeneous video clips and ground-truth annotations for tens of concepts that exhibit significant variation in terms of generality, complexity, and human participation. The experimental results manifest the merit of the proposed techniques.
Panagiotis Sidiropoulos, Vasileios Mezaris, Ioannis Kompatsiaris
IEEE Trans. Circuits Syst. Video Technol.2
2013 Video event recounting using mixture subclass discriminant analysis
abstract
In this paper, a new feature selection method is used, in combination with a semantic model vector video representation, in order to enumerate the key semantic evidences of an event in a video signal. In particular, a set of semantic concept detectors is firstly used for estimating a model vector for each video signal, where each element of the model vector denotes the degree of confidence that the respective concept is depicted in the video. Then, a novel feature selection method is learned for each event of interest. This method is based on exploiting the first two eigenvectors derived using the eigenvalue formulation of the mixture subclass discriminant analysis. Subsequently, given a video-event pair, the proposed method jointly evaluates the significance of each concept for the detection of the given event and the degree of confidence with which this concept is detected in the given video, in order to decide which concepts provide the strongest evidence in support of the provided video-event link. Experimental results using a video collection of TRECVID demonstrate the effectiveness of the proposed video event recounting method.
Nikolaos Gkalelis, Vasileios Mezaris, Ioannis Kompatsiaris, Tania Stathaki
ICIP2
2013 Enhancing video concept detection with the use of tomographs
abstract
In this work we deal with the problem of video concept detection, for the purpose of using the detection results towards more effective concept-based video retrieval. In order to handle this task, we propose using spatio-temporal video slices, called video tomographs, in the same way that visual keyframes are typically used in traditional keyframe-based video concept detection schemes. Video tomographs capture in a compact way motion patterns that are present in the video, and are used in this work for training a number of base detectors. The latter augment the set of keyframe-based base detectors that can be trained on different image representations. Combining the keyframe-based and tomograph-based detectors, improved concept detection accuracy can be achieved. The proposed approach is evaluated on a dataset that is extensive both in terms of video duration and concept variation. The experimental results manifest the merit of the proposed approach.
Panagiotis Sidiropoulos, Vasileios Mezaris, Ioannis Kompatsiaris
ICIP2
2013 Video event detection using a subclass recoding error-correcting output codes framework
abstract
In this paper, complex video events are learned and detected using a novel subclass recoding error-correcting outputs (SRECOC) design. In particular, a set of pre-trained concept detectors along different low-level visual feature types are used to provide a model vector representation of video signals. Subsequently, a subclass partitioning algorithm is used to divide only the target event class to several subclasses and learn one subclass detector for each event subclass. The pool of the subclass detectors is then combined under a SRECOC framework to provide a single event detector. This is achieved by first exploiting the properties of the linear loss-weighted decoding measure in order to derive a probability estimate along the different event subclass detectors, and then utilizing the sum probability rule along event subclasses to retrieve a single degree of confidence for the presence of the target event in a particular test video. Experimental results on the large-scale video collections of the TRECVID Multimedia Event Detection (MED) task verify the effectiveness of the proposed method. Moreover, the effect of weak or strong concept detectors on the accuracy of the resulting event detectors is examined.
Nikolaos Gkalelis, Vasileios Mezaris, Michail Dimopoulos, Ioannis Kompatsiaris, Tania Stathaki
ICME2
2013 Summary abstract for the 2nd ACM international workshop on multimedia analysis for ecological data
abstract
The 2nd ACM International Workshop on Multimedia Analysis for Ecological Data (MAED'13) is held as part of ACM Multimedia 2013. MAED'13, following the first workshop of the MAED series (MAED'12) that was held as part of ACM Multimedia 2012, is concerned with the processing, interpretation, and visualization of ecology-related multimedia content with the aim to support biologists in their investigations for analyzing and monitoring natural environments.
Concetto Spampinato, Vasileios Mezaris, Jacco van Ossenbruggen
ACM Multimedia2
2013 Improving event detection using related videos and relevance degree support vector machines
abstract
In this paper, a new method that exploits related videos for the problem of event detection is proposed, where related videos are videos that are closely but not fully associated with the event of interest. In particular, the Weighted Margin SVM formulation is modified so that related class observations can be effectively incorporated in the optimization problem. The resulting Relevance Degree SVM is especially useful in problems where only a limited number of training observations is provided, e.g., for the EK10Ex subtask of TRECVID MED, where only ten positive and ten related samples are provided for the training of a complex event detector. Experimental results on the TRECVID MED 2011 dataset verify the effectiveness of the proposed method.
Christos Tzelepis, Nikolaos Gkalelis, Vasileios Mezaris, Ioannis Kompatsiaris
ACM Multimedia3
2013 The 2012 social event detection dataset
abstract
This paper presents the 2012 Social Event Detection dataset (SED2012). The dataset constitutes a challenging benchmark for methods that detect social events in large collections of multimedia items. More specifically, the dataset comprises more than 160 thousands of Flickr photos and their accompanying metadata, as well as a list of 149 manually selected and annotated target events, each of which is defined as a set of relevant photos. This paper discusses the challenges defined as part of SED 2012, the data collection process, the dataset and its basic statistics, the ground truth creation and the suggested evaluation methodology.
Symeon Papadopoulos, Emmanouil Schinas, Vasileios Mezaris, Raphaël Troncy, Ioannis Kompatsiaris
MMSys3
2013 Mixture Subclass Discriminant Analysis Link to Restricted Gaussian Model and Other Generalizations
abstract
In this paper, a theoretical link between mixture subclass discriminant analysis (MSDA) and a restricted Gaussian model is first presented. Then, two further discriminant analysis (DA) methods, i.e., fractional step MSDA (FSMSDA) and kernel MSDA (KMSDA) are proposed. Linking MSDA to an appropriate Gaussian model allows the derivation of a new DA method under the expectation maximization (EM) framework (EM-MSDA), which simultaneously derives the discriminant subspace and the maximum likelihood estimates. The two other proposed methods generalize MSDA in order to solve problems inherited from conventional DA. FSMSDA solves the subclass separation problem, that is, the situation in which the dimensionality of the discriminant subspace is strictly smaller than the rank of the inter-between-subclass scatter matrix. This is done by an appropriate weighting scheme and the utilization of an iterative algorithm for preserving useful discriminant directions. On the other hand, KMSDA uses the kernel trick to separate data with nonlinearly separable subclass structure. Extensive experimentation shows that the proposed methods outperform conventional MSDA and other linear discriminant analysis variants.
Nikolaos Gkalelis, Vasileios Mezaris, Ioannis Kompatsiaris, Tania Stathaki
IEEE Trans. Neural Networks Learn. Syst.2
2012 Multimedia analysis for ecological data
abstract
The ACM International Workshop on Multimedia Analysis for Ecological Data (MAED'12) is held as part of ACM Multimedia 2012. MAED'12 is concerned with the processing, interpretation, and visualization of ecology-related multimedia content with the aim to support biologists in their investigations for analyzing and monitoring natural environments, with particular attention to living organisms and pollution effects.
Concetto Spampinato, Vasileios Mezaris, Jacco van Ossenbruggen
ACM Multimedia2
2012 Linear Subclass Support Vector Machines
abstract
In this letter, linear subclass support vector machines (LSSVMs) are proposed that can efficiently learn a piecewise linear decision function for binary classification problems. This is achieved using a nongaussianity criterion to derive the subclass structure of the data, and a new formulation of the optimization problem that exploits the subclass information. LSSVMs provide low computation cost during training and evaluation, and offer competitive recognition performance in comparison to other popular SVM-based algorithms. Experimental results on various datasets confirm the advantages of LSSVMs.
Nikolaos Gkalelis, Vasileios Mezaris, Ioannis Kompatsiaris, Tania Stathaki
IEEE Signal Process. Lett.2
2012 Differential Edit Distance: A Metric for Scene Segmentation Evaluation
abstract
In this paper, a novel approach to evaluating video temporal decomposition algorithms is presented. The evaluation measures typically used to this end are nonlinear combinations of precision-recall or coverage-overflow, which are not metrics and additionally possess undesirable properties, such as nonsymmetricity. To alleviate these drawbacks, we introduce a novel unidimensional measure that is proven to be metric and satisfies a number of qualitative prerequisites that previous measures do not. This measure is named differential edit distance (DED), since it can be seen as a variation of the well-known edit distance. After defining DED, we further introduce an algorithm that computes it in less than cubic time. Finally, DED is extensively compared with state-of-the-art measures, namely, the harmonic means (F-score) of precision-recall and coverage-overflow. The experiments include comparisons of qualitative properties, the time required for optimizing the parameters of scene segmentation algorithms with the help of these measures, and a user study gauging the agreement of these measures with the users' assessment of the segmentation results. The results confirm that the proposed measure is a unidimensional metric that is effective in evaluating scene segmentation techniques and in helping to optimize their parameters.
Panagiotis Sidiropoulos, Vasileios Mezaris, Ioannis Kompatsiaris, Josef Kittler
IEEE Trans. Circuits Syst. Video Technol.2
2011 High-level event detection system based on discriminant visual concepts
abstract
This paper demonstrates a new approach to detecting high-level events that may be depicted in images or video frames. Given a non-annotated content item, a large number of previously trained visual concept detectors are applied to it and their responses are used for representing the content item with a model vector in a high-dimensional concept space. Subsequently, an improved subclass discriminant analysis method is used for identifying a concept subspace within the aforementioned concept space, that is most appropriate for detecting and recognizing the target high-level events. In this subspace, the nearest neighbor rule is used for comparing the non-annotated content item with a few known example instances of the target events. The high-level events used as target events in the present version of the system are those defined for the TRECVID 2010 Multimedia Event Detection (MED) task.
Ioannis Tsampoulatidis, Nikolaos Gkalelis, Anastasios Dimou, Vasileios Mezaris, Ioannis Kompatsiaris
ICMR4
2011 Modeling and representing events in multimedia
abstract
This paper presents an overview of the Joint Workshop on Modeling and Representing Events (JMRE), which is held as part of ACM Multimedia 2011. JMRE is concerned with the understanding of events from multimedia, and with using events in order to better organize and consume multimedia.
Vasileios Mezaris, Ansgar Scherp, Ramesh Jain 0001, Mohan Kankanhalli, Huiyu Zhou 0001, Jianguo Zhang 0001, Liang Wang 0001, Zhengyou Zhang
ACM Multimedia1
2011 A comparative study of object-level spatial context techniques for semantic image analysis
Georgios Th. Papadopoulos, Carsten Saathoff, Hugo Jair Escalante, Vasileios Mezaris, Ioannis Kompatsiaris, Michael G. Strintzis
Comput. Vis. Image Underst.4
2011 Mixture Subclass Discriminant Analysis
abstract
In this letter, mixture subclass discriminant analysis (MSDA) that alleviates two shortcomings of subclass discriminant analysis (SDA) is proposed. In particular, it is shown that for data with Gaussian homoscedastic subclass structure a) SDA does not guarantee to provide the discriminant subspace that minimizes the Bayes error, and, b) the sample covariance matrix can not be used as the minimization metric of the discriminant analysis stability criterion (DSC). Based on this analysis MSDA modifies the objective function of SDA and utilizes a novel partitioning procedure to aid discrimination of data with Gaussian homoscedastic subclass structure. Experimental results confirm the improved classification performance of MSDA.
Nikolaos Gkalelis, Vasileios Mezaris, Ioannis Kompatsiaris
IEEE Signal Process. Lett.2
2011 Temporal Video Segmentation to Scenes Using High-Level Audiovisual Features
abstract
In this paper, a novel approach to video temporal decomposition into semantic units, termed scenes, is presented. In contrast to previous temporal segmentation approaches that employ mostly low-level visual or audiovisual features, we introduce a technique that jointly exploits low-level and high-level features automatically extracted from the visual and the auditory channel. This technique is built upon the well-known method of the scene transition graph (STG), first by introducing a new STG approximation that features reduced computational cost, and then by extending the unimodal STG-based temporal segmentation technique to a method for multimodal scene segmentation. The latter exploits, among others, the results of a large number of TRECVID-type trained visual concept detectors and audio event detectors, and is based on a probabilistic merging process that combines multiple individual STGs while at the same time diminishing the need for selecting and fine-tuning several STG construction parameters. The proposed approach is evaluated on three test datasets, comprising TRECVID documentary films, movies, and news-related videos, respectively. The experimental results demonstrate the improved performance of the proposed approach in comparison to other unimodal and multimodal techniques of the relevant literature and highlight the contribution of high-level audiovisual features toward improved video segmentation to scenes.
Panagiotis Sidiropoulos, Vasileios Mezaris, Ioannis Kompatsiaris, Hugo Meinedo, Miguel M. F. Bugalho, Isabel Trancoso
IEEE Trans. Circuits Syst. Video Technol.2
2010 On the use of feature tracks for dynamic concept detection in video
abstract
This paper proposes the use of feature tracks for the detection of concepts in video, particularly dynamic concepts. Feature tracks are defined as sets of local interest points found in different frames of a video shot that exhibit spatio-temporal and visual continuity, defining a trajectory in the 2D+Time space. The extraction of feature tracks and the selection and representation of an appropriate subset of them allow the generation of a Bag-of-Spatiotemporal-Words model for the shot, which facilitates capturing the dynamics of video content. The experimental evaluation of the proposed approach highlights how the selection of such feature tracks for the definition of the Bag-of-Spatiotemporal-Words model enhances the results of traditional keyframe-based concept detection techniques.
Vasileios Mezaris, Anastasios Dimou, Ioannis Kompatsiaris
ICIP1
2010 Probabilistic combination of spatial context with visual and co-occurrence information for semantic image analysis
abstract
In this paper, a probabilistic approach to combining spatial context with visual and co-occurrence information for semantic image analysis is presented. Overall, the examined image is segmented and subsequently an initial classification of the resulting image regions to semantic concepts is performed based solely on visual information. Then, a Genetic Algorithm (GA) is introduced for deciding on the optimal semantic image interpretation, realizing image analysis as a global optimization problem. The fundamental novelty of this work is that the GA incorporates in its evolutionary procedure a set of Bayesian Networks (BNs), which probabilistically learn the impact of the available spatial, visual and co-occurrence information on the final outcome for every possible pair of semantic concepts. Experimental results on two publicly available datasets demonstrate the efficiency of the proposed approach.
Georgios Th. Papadopoulos, Vasileios Mezaris, Ioannis Kompatsiaris, Michael G. Strintzis
ICIP2
2010 A Statistical Learning Approach to Spatial Context Exploitation for Semantic Image Analysis
abstract
In this paper, a statistical learning approach to spatial context exploitation for semantic image analysis is presented. The proposed method constitutes an extension of the key parts of the authors' previous work on spatial context utilization, where a Genetic Algorithm (GA) was introduced for exploiting fuzzy directional relations after performing an initial classification of image regions to semantic concepts using solely visual information. In the extensions reported in this work, a more elaborate approach is followed during the spatial knowledge acquisition and modeling process. Additionally, the impact of every resulting spatial constraint on the final outcome is adaptively adjusted. Experimental results as well as comparative evaluation on three datasets of varying complexity in terms of the total number of supported semantic concepts demonstrate the efficiency of the proposed method.
Georgios Th. Papadopoulos, Vasileios Mezaris, Ioannis Kompatsiaris, Michael G. Strintzis
ICPR2
2010 Modeling, detecting, and processing events in multimedia
abstract
No abstract available.
Ansgar Scherp, Ramesh Jain 0001, Mohan Kankanhalli, Vasileios Mezaris
ACM Multimedia4
2009 Combining multimodal and temporal contextual information for semantic video analysis
abstract
In this paper, a graphical modeling-based approach to semantic video analysis is presented for jointly realizing modality fusion and temporal context exploitation. Overall, the examined video sequence is initially segmented into shots and for every resulting shot appropriate color, motion and audio features are extracted. Then, Hidden Markov Models (HMMs) are employed for performing an initial association of each shot with the semantic classes that are of interest separately for every modality. Subsequently, an integrated Bayesian Network (BN) is introduced for simultaneously performing information fusion and temporal contextual knowledge exploitation, contrary to the usual practice of performing each task separately. The final outcome of the overall video analysis approach is the association of a semantic class with every shot. Experimental results as well as comparative evaluation from the application of the proposed approach in the domain of news broadcast video are presented.
Georgios Th. Papadopoulos, Vasileios Mezaris, Ioannis Kompatsiaris, Michael G. Strintzis
ICIP2
2009 Multi-modal scene segmentation using scene transition graphs
abstract
In this work the problem of automatic decomposition of video into elementary semantic units, known in the literature as scenes, is addressed. Two multi-modal automatic scene segmentation techniques are proposed, both building upon the Scene Transition Graph (STG). In the first of the proposed approaches, speaker diarization results are used for introducing a post-processing step to the STG construction algorithm, with the objective of discarding scene boundaries erroneously identified according to visual-only dissimilarity. In the second approach, speaker diarization and additional audio analysis results are employed and a separate audio-based STG is constructed, in parallel to the original STG based on visual information. The two STGs are subsequently combined. Preliminary results from the application of the proposed techniques to broadcast videos reveal their improved performance over previous approaches.
Panagiotis Sidiropoulos, Vasileios Mezaris, Ioannis Kompatsiaris, Hugo Meinedo, Isabel Trancoso
ACM Multimedia2
2009 Integrating Image Segmentation and Classification for Fuzzy Knowledge-Based Multimedia Indexing
Thanos Athanasiadis, Nikos Simou, Georgios Th. Papadopoulos, Rachid Benmokhtar, Krishna Chandramouli, Vassilis Tzouvaras, Vasileios Mezaris, Marios Phinikettos, Yannis Avrithis, Ioannis Kompatsiaris, Benoit Huet, Ebroul Izquierdo
MMM7
2009 Statistical Motion Information Extraction and Representation for Semantic Video Analysis
abstract
In this paper, an approach to semantic video analysis that is based on the statistical processing and representation of the motion signal is presented. Overall, the examined video is temporally segmented into shots and for every resulting shot appropriate motion features are extracted; using these, hidden Markov models (HMMs) are employed for performing the association of each shot with one of the semantic classes that are of interest. The novel contributions of this paper lie in the areas of motion information processing and representation. Regarding the motion information processing, the kurtosis of the optical flow motion estimates is calculated for identifying which motion values originate from true motion rather than measurement noise. Additionally, unlike the majority of the approaches of the relevant literature that are mainly limited to global- or camera-level motion representations, a new representation for providing local-level motion information to HMMs is also presented. It focuses only on the pixels where true motion is observed. For the selected pixels, energy distribution-related information, as well as a complementary set of features that highlight particular spatial attributes of the motion signal, are extracted. Experimental results, as well as comparative evaluation, from the application of the proposed approach in the domains ofTennis,NewsandVolleyballbroadcast video, andHuman Actionvideo demonstrate the efficiency of the proposed method.
Georgios Th. Papadopoulos, Alexia Briassouli, Vasileios Mezaris, Ioannis Kompatsiaris, Michael G. Strintzis
IEEE Trans. Circuits Syst. Video Technol.3
2008 Estimation and representation of accumulated motion characteristics for semantic event detection
abstract
In this paper, a motion-based approach for detecting high-level semantic events in video sequences is presented. Its main characteristic is its generic nature, i.e. it can be directly applied to any possible domain of concern without the need for domain-specific algorithmic modifications or adaptations. For realizing event detection, the examined video sequence is initially segmented into shots and for every resulting shot appropriate motion features are extracted. Then, Hidden Markov Models (HMMs) are employed for performing the association of each shot with one of the high-level semantic events that are of interest in any given domain. Regarding the motion feature extraction procedure, a new representation for providing local-level motion information to HMMs is presented, while motion characteristics from previous frames are also exploited. Experimental results as well as comparative evaluation from the application of the proposed approach in the domain of news broadcast video are presented.
Georgios Th. Papadopoulos, Vasileios Mezaris, Ioannis Kompatsiaris, Michael G. Strintzis
ICIP2
2008 Gradual transition detection using color coherence and other criteria in a video shot meta-segmentation framework
abstract
Shot segmentation provides the basis for almost all high-level video content analysis approaches, validating it as one of the major prerequisites for efficient video semantic analysis, indexing and retrieval. The successful detection of both gradual and abrupt transitions is necessary to this end. In this paper a new gradual transition detection algorithm is proposed, that is based on novel criteria such as color coherence change that exhibit less sensitivity to local or global motion than previously proposed ones. These criteria, each of which could serve as a standalone gradual transition detection approach, are then combined using a machine learning technique, to result in a meta-segmentation scheme. Besides significantly improved performance, advantage of the proposed scheme is that there is no need for threshold selection, as opposed to what would be the case if any of the proposed features were used by themselves and as is typically the case in the relevant literature. Performance evaluation and comparison with four other popular algorithms reveals the effectiveness of the proposed technique.
Efthymia Tsamoura, Vasileios Mezaris, Ioannis Kompatsiaris
ICIP2
2007 Automated IVUS Contour Detection Using Intesity Features and Radial Basis Function Approximation
abstract
Intravascular ultrasound (IVUS) constitutes a valuable technique for the diagnosis and management of coronary atherosclerosis. The detection of lumen and media-adventitia borders in IVUS images represents a necessary step in the utilization of the IVUS data for the reliable quantitative assessment of atherosclerotic lesions. In this paper, a fully automated technique for the detection of lumen and media-adventitia boundaries is developed. This comprises two different steps for contour initialization, one for each corresponding contour of interest, and a procedure for the refinement of the detected contours based on Radial Basis Function approximation. The proposed approach is shown to be capable of performing quick and reliable automated segmentation of IVUS images.
Maria Papadogiorgaki, Vasileios Mezaris, Yiannis S. Chatzizisis, George D. Giannoglou, Ioannis Kompatsiaris
CBMS2
2007 Video Segmentation and Semantics Extraction from the Fusion of Motion and Color Information
abstract
In recent years, digital multimedia technologies have evolved significantly, and are finding numerous applications, over the internet, and even over mobile networks. Thus, the video processing community has started focusing more intensively on the extraction of higher level information from multimedia data. This paper proposes a novel two-stage video processing system that aims to segment and extract semantically meaningful information, which can help achieve higher level interpretation of video. The flow fields present in the video are accumulated over several frames and their statistics are processed to derive an "activity area", that is characteristic of the type of events taking place. The color information complements the motion data, and is used for the accurate segmentation of the moving entities in each frame. The joint use of the activity area and accurate segmentation can serve as a first step to the further semantic interpretation of the video, including the recognition and accurate localization of moving objects of interest. We present experiments that demonstrate the effectiveness of our method for real videos.
Alexia Briassouli, Vasileios Mezaris, Ioannis Kompatsiaris
ICIP (3)2
2007 Joint Motion and Color Statistical Video Processing for Motion Segmentation
abstract
Vast amounts of digital multimedia data are being produced and distributed today, so methods for the efficient and reliable extraction of information from video data are becoming necessary. We present a novel motion segmentation algorithm, which accurately extracts moving objects from a video, and also provides a likelihood map, for each object pixel assignment. The flow is estimated, and accumulated over several frames, to give action masks. Color segmentation clusters regions of similar color in each frame. A novel, likelihood ratio-based method for the statistical comparison of color layers in the regions of activity and the background is presented and compared with an Earth mover's distance-based approach. Our method also gives the likelihood with which each pixel is assigned to a moving object in each frame. Experiments with real sequences illustrate the advantages of our method, namely that it gives overall more reliable results, and also provides the likelihood map for the segmented object.
Alexia Briassouli, Vasileios Mezaris, Ioannis Kompatsiaris
ICME2
2006 Knowledge-Assisted Image Analysis Based on Context and Spatial Optimization
abstract
In this article, an approach to semantic image analysis is presented. Under the proposed approach, ontologies are used to capture general, spatial, and contextual knowledge of a domain, and a genetic algorithm is applied to realize the final annotation. The employed domain knowledge considers high-level information in terms of the concepts of interest of the examined domain, contextual information in the form of fuzzy ontological relations, as well as low-level information in terms of prototypical low-level visual descriptors. To account for the inherent ambiguity in visual information, uncertainty has been introduced in the spatial relations definition. First, an initial hypothesis set of graded annotations is produced for each image region, and then context is exploited to update appropriately the estimated degrees of confidence. Finally, a genetic algorithm is applied to decide the most plausible annotation by utilizing the visual and the spatial concepts definitions included in the domain ontology. Experiments with a collection of photographs belonging to two different domains demonstrate the performance of the proposed approach.
Georgios Th. Papadopoulos, Phivos Mylonas, Vasileios Mezaris, Yannis Avrithis, Ioannis Kompatsiaris
Int. J. Semantic Web Inf. Syst.3
2006 Object-based MPEG-2 video indexing and retrieval in a collaborative environment
Vasileios Mezaris, Ioannis Kompatsiaris, Michael G. Strintzis
Multim. Tools Appl.1
2005 A genetic algorithm-based approach to knowledge-assisted video analysis
abstract
Efficient video content management and exploitation requires extraction of the underlying semantics, a non-trivial task associating low-level features of the image domain and high-level semantic descriptions. In this paper, a knowledge-assisted approach for extracting semantics of domain-specific video content is presented. Domain knowledge considers both low-level features (color, motion, shape) and spatial behavior (topological and directional information). During the preprocessing step, a set of over-segmented homogenous atom-regions is generated and their low-level and spatial descriptions are extracted. A genetic algorithm is then applied in order to find the optimal interpretation according to a specific domain conceptualization. The proposed approach was tested on the formula one, tennis and beach vacations domains showing promising results.
Nikola Voisine, Stamatia Dasiopoulou, Frédéric Precioso, Vasileios Mezaris, Ioannis Kompatsiaris, Michael G. Strintzis
ICIP (3)4
2005 Knowledge-assisted semantic video object detection
abstract
An approach to knowledge-assisted semantic video object detection based on a multimedia ontology infrastructure is presented. Semantic concepts in the context of the examined domain are defined in an ontology, enriched with qualitative attributes (e.g., color homogeneity), low-level features (e.g., color model components distribution), object spatial relations, and multimedia processing methods (e.g., color clustering). Semantic Web technologies are used for knowledge representation in the RDF(S) metadata standard. Rules in F-logic are defined to describe how tools for multimedia analysis should be applied, depending on concept attributes and low-level features, for the detection of video objects corresponding to the semantic concepts defined in the ontology. This supports flexible and managed execution of various application and domain independent multimedia analysis tasks. Furthermore, this semantic analysis approach can be used in semantic annotation and transcoding systems, which take into consideration the users environment including preferences, devices used, available network bandwidth and content identity. The proposed approach was tested for the detection of semantic objects on video data of three different domains.
Stamatia Dasiopoulou, Vasileios Mezaris, Ioannis Kompatsiaris, Vasileios Papastathis, Michael G. Strintzis
IEEE Trans. Circuits Syst. Video Technol.2
2004 A knowledge-based approach to domain-specific compressed video analysis
abstract
A novel approach to domain-specific video analysis is proposed. The proposed approach is based on exploiting domain-specific knowledge in the form of an ontology to detect video objects corresponding to the semantic concepts defined in the ontology. The association between the visual objects and the defined semantic concepts is performed by taking into account both qualitative attributes of the semantic objects (e.g., color homogeneity), indicating necessary preprocessing methods (color clustering, respectively), and numerical data generated via training (e.g., color models, also defined in the ontology). To enable fast and efficient processing, this methodology is applied to MPEG-2 video, requiring only its partial decoding. The proposed approach is demonstrated in the domain of Formula-1 racing video and shows promising results.
Vasileios Mezaris, Ioannis Kompatsiaris, Michael G. Strintzis
ICIP1
2004 Combining Multiple Segmentation Algorithms and the MPEG-7 eXperimentation Model in the Schema Reference System
abstract
The SCHEMA Network of Excellence aims to bring together a critical mass of universities, research centers, industrial partners and end users, in order to design a reference system for content-based semantic scene analysis, interpretation and understanding. In this paper, advances in the development of the SCHEMA reference system are reported, focusing on the application of region-based image retrieval using automatic segmentation. More specifically, the integration of four segmentation algorithms and the MPEG-7 eXperimentation Model with the reference system are discussed, along with the motivation behind these and various other choices that were made during the development of the reference system. Experimental results for this system, as well as results for an earlier version of it employing proprietary descriptors, are shown using a common collection of images. Comparative evaluation of these versions, both in terms of retrieval accuracy and in terms of time-efficiency, allows the evaluation of the reference system as a whole as well as the evaluation of the usability of different components integrated with it, such as the MPEG-7 eXperimentation Model. These results illustrate the efficiency of the proposed system, as well as its suitability in serving as a test-bed for evaluating and comparing different algorithms and approaches pertaining to the content-based and semantic manipulation of visual information.
Vasileios Mezaris, Charalambos Doulaverakis, Raúl Medina Beltrán de Otálora, Stephan Herrmann 0002, Ioannis Kompatsiaris, Michael G. Strintzis
IV1
2004 Still Image Segmentation Tools For Object-Based Multimedia Applications
abstract
In this paper, a color image segmentation algorithm and an approach to large-format image segmentation are presented, both focused on breaking down images to semantic objects for object-based multimedia applications. The proposed color image segmentation algorithm performs the segmentation in the combined intensity–texture–position feature space in order to produce connected regions that correspond to the real-life objects shown in the image. A preprocessing stage of conditional image filtering and a modified K-Means-with-connectivity-constraint pixel classification algorithm are used to allow for seamless integration of the different pixel features. Unsupervised operation of the segmentation algorithm is enabled by means of an initial clustering procedure. The large-format image segmentation scheme employs the aforementioned segmentation algorithm, providing an elegant framework for the fast segmentation of relatively large images. In this framework, the segmentation algorithm is applied to reduced versions of the original images, in order to speed-up the completion of the segmentation, resulting in a coarse-grained segmentation mask. The final fine-grained segmentation mask is produced with partial reclassification of the pixels of the original image to the already formed regions, using a Bayes classifier. As shown by experimental evaluation, this novel scheme provides fast segmentation with high perceptual segmentation quality.
Vasileios Mezaris, Ioannis Kompatsiaris, Michael G. Strintzis
Int. J. Pattern Recognit. Artif. Intell.1
2004 Real-time compressed-domain spatiotemporal segmentation and ontologies for video indexing and retrieval
abstract
In this paper, a novel algorithm is presented for the real-time, compressed-domain, unsupervised segmentation of image sequences and is applied to video indexing and retrieval. The segmentation algorithm uses motion and color information directly extracted from the MPEG-2 compressed stream. An iterative rejection scheme based on the bilinear motion model is used to effect foreground/background segmentation. Following that, meaningful foreground spatiotemporal objects are formed by initially examining the temporal consistency of the output of iterative rejection, clustering the resulting foreground macroblocks to connected regions and finally performing region tracking. Background segmentation to spatiotemporal objects is additionally performed. MPEG-7 compliant low-level descriptors describing the color, shape, position, and motion of the resulting spatiotemporal objects are extracted and are automatically mapped to appropriate intermediate-level descriptors forming a simple vocabulary termed object ontology. This, combined with a relevance feedback mechanism, allows the qualitative definition of the high-level concepts the user queries for (semantic objects, each represented by a keyword) and the retrieval of relevant video segments. Desired spatial and temporal relationships between the objects in multiple-keyword queries can also be expressed, using the shot ontology. Experimental results of the application of the segmentation algorithm to known sequences demonstrate the efficiency of the proposed segmentation approach. Sample queries reveal the potential of employing this segmentation algorithm as part of an object-based video indexing and retrieval scheme.
Vasileios Mezaris, Ioannis Kompatsiaris, Nikolaos V. Boulgouris, Michael G. Strintzis
IEEE Trans. Circuits Syst. Video Technol.1
2004 Video object segmentation using Bayes-based temporal tracking and trajectory-based region merging
abstract
A novel unsupervised video object segmentation algorithm is presented, aiming to segment a video sequence to objects: spatiotemporal regions representing a meaningful part of the sequence. The proposed algorithm consists of three stages: initial segmentation of the first frame using color, motion, and position information, based on a variant of the K-means-with-connectivity-constraint algorithm; a temporal tracking algorithm, using a Bayes classifier and rule-based processing to reassign changed pixels to existing regions and to efficiently handle the introduction of new regions; and a trajectory-based region merging procedure that employs the long-term trajectory of regions, rather than the motion at the frame level, so as to group them to objects with different motion. As shown by experimental evaluation, this scheme can efficiently segment video sequences with fast moving or newly appearing objects. A comparison with other methods shows segmentation results corresponding more accurately to the real objects appearing on the image sequence.
Vasileios Mezaris, Ioannis Kompatsiaris, Michael G. Strintzis
IEEE Trans. Circuits Syst. Video Technol.1
2003 An ontology approach to object-based image retrieval
abstract
In this paper, an image retrieval methodology suited for search in large collections of heterogeneous images is presented. The proposed approach employs a fully unsupervised segmentation algorithm to divide images into regions. Low-level features describing the color, position, size and shape of the resulting regions are extracted and are automatically mapped to appropriate intermediate-level descriptors forming a simple vocabulary termed object ontology. The object ontology is used to allow the qualitative definition of the high-level concepts the user queries for (semantic objects, each represented by a keyword) in a human-centered fashion. When querying, clearly irrelevant image regions are rejected using the intermediate-level descriptors; following that, a relevance feedback mechanism employing the low-level features is invoked to produce the final query results. The proposed approach bridges the gap between keyword-based approaches, which assume the existence of rich image captions or require manual evaluation and annotation of every image of the collection, and query-by-example approaches, which assume that the user queries for images similar to one that already is at his disposal.
Vasileios Mezaris, Ioannis Kompatsiaris, Michael G. Strintzis
ICIP (2)1
2002 A framework for the efficient segmentation of large-format color images
abstract
A novel approach to large-format image segmentation is presented, focused on usage in content-based multimedia applications. The proposed framework aims at facilitating the time-efficient segmentation of large-format images while maintaining the high perceptual quality of the segmentation result. For this to be achieved, the employed segmentation algorithm is applied to reduced versions of the large-format images, in order to speed-up its execution, resulting in a coarse-grained segmentation mask. The final fine-grained segmentation mask is produced by an enhancement stage that involves partial reclassification of the pixels of the original image using a Bayes classifier. As shown by experimental evaluation, this novel scheme provides fast segmentation with high perceptual segmentation quality.
Vasileios Mezaris, Ioannis Kompatsiaris, Michael G. Strintzis
ICIP (1)1