EDBT 2026 Demo / reviewers in the wild / expert
Nikos Nikolaidis 0001
dblp:n/NikosNikolaidis · also Nikolaos Nikolaidis 0001
· DBLP profile ↗
123ranked-venue papers
10as first author
17since 2021 · last 2026
0000-0003-1515-7986ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 95 · 10 first-author · 9 since 2021Artificial intelligence and machine learning · 25 · 8 since 2021Systems, architecture and hardware · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | REGEN: Real-Time Photorealism Enhancement in Games via a Dual-Stage Generative Network FrameworkabstractPhotorealism is an important aspect of modern video games since it can shape player experience and impact immersion, narrative engagement, and visual fidelity. To achieve photorealism, beyond traditional rendering pipelines, generative models have been increasingly adopted as an effective approach for bridging the gap between the visual realism of synthetic and real worlds. However, under real-time constraints of video games, existing generative approaches continue to face a tradeoff between visual quality and runtime efficiency. In this work, we present a framework for enhancing the photorealism of rendered game frames using generative networks. We propose REGEN, which first employs a robust unpaired image-to-image translation model to generate semantically consistent photorealistic frames. These generated frames are then used to create a paired dataset, which transforms the problem to a simpler unpaired image-to-image translation. This enables training with a lightweight method, achieving real-time inference without compromising visual quality. We evaluate REGEN on Unreal Engine, showing, by employing the CMMD metric, that it achieves comparable or slightly improved visual quality compared to the robust method, while improving the frame rate by 12x. Additional experiments also validate that REGEN adheres to the semantic preservation of the initial robust image-to-image translation method and maintains temporal consistency. Code, pre-trained models, and demos for this work are available at: https://github.com/stefanos50/REGEN Stefanos Pasios, Nikos Nikolaidis 0001 |
IEEE Trans. Games | 2 |
| 2025 | Efficient deterministic renewable energy forecasting guided by multiple-location weather data
Charalampos Symeonidis, Nikos Nikolaidis 0001 |
Neural Comput. Appl. | 2 |
| 2025 | CARLA2Real: A Tool for Reducing the Sim2real Appearance Gap in CARLA SimulatorabstractSimulators are indispensable for research in autonomous systems such as self-driving cars, autonomous robots, and drones. Despite significant progress in various simulation aspects, such as graphical realism, an evident gap persists between the virtual and real-world environments. Since the ultimate goal is to deploy the autonomous systems in the real world, reducing the sim2real gap is of utmost importance. In this paper, we employ a state-of-the-art approach to enhance the photorealism of simulated data, aligning them with the visual characteristics of real-world datasets. Based on this, we developed CARLA2Real, an easy-to-use, publicly available tool (plug-in) for the widely used and open-source CARLA simulator. This tool enhances the output of CARLA in near real-time, achieving a frame rate of 13 FPS, translating it to the visual style and realism of real-world datasets such as Cityscapes, KITTI, and Mapillary Vistas. By employing the proposed tool, we generated synthetic datasets from both the simulator and the enhancement model outputs, including their corresponding ground truth annotations for tasks related to autonomous driving. Then, we performed a number of experiments to evaluate the impact of the proposed approach on feature extraction and semantic segmentation methods when trained on the enhanced synthetic data. The results demonstrate that the sim2real appearance gap is significant and can indeed be reduced by the introduced approach. Comparisons with a state-of-the-art image-to-image translation approach are also provided. The tool, pre-trained models, and associated data for this work are available for download at:https://github.com/stefanos50/CARLA2Real Stefanos Pasios, Nikos Nikolaidis 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | Enhancing visual object tracking robustness through a lightweight denoising moduleabstractAbstract Visual object tracking is crucial for numerous applications ranging from smartphones to autonomous vehicles. However, the impact of input noise on tracking performance remains underexplored. This paper presents a lightweight neural network module designed to enhance the robustness of 2D tracking methods against various types of noise. By performing image-to-image translation, the proposed robust tracking module (RTM) standardizes the operational space of tracking algorithms, thereby improving their resilience. Experimental results on benchmark datasets demonstrate the effectiveness of RTM in mitigating performance degradation caused by noise. Additionally, we introduce an evaluation toolkit that facilitates the assessment of tracking robustness against common noise types. The source code of the proposed method is available at https://github.com/iason1907/RTM . Iason Karakostas, Vasileios Mygdalis, Nikos Nikolaidis 0001, Ioannis Pitas |
Vis. Comput. | 3 |
| 2024 | A synthetic human-centric dataset generation pipeline for active robotic vision
Charalampos Georgiadis, Nikolaos Passalis, Nikos Nikolaidis 0001 |
Pattern Recognit. Lett. | 3 |
| 2023 | Wind Energy Prediction Guided by Multiple-Location Weather Forecasts
Charalampos Symeonidis, Nikos Nikolaidis 0001 |
EANN | 2 |
| 2023 | Efficient Feature Extraction for Non-Maximum Suppression in Visual Person DetectionabstractNon-Maximum Suppression (NMS) is a post-processing step in almost every visual object detector, tasked with rapidly pruning the number of overlapping detected candidate rectangular Regions-of-Interest (RoIs) and replacing them with a single, more spatially accurate detection (in pixel coordinates). The common Greedy NMS algorithm suffers from drawbacks, due to the need for careful manual tuning. In visual person detection, most NMS methods typically suffer when analyzing crowded scenes with high levels of in-between occlusions. This paper proposes a modification on a deep neural architecture for NMS, suitable for such cases and capable of efficiently cooperating with recent neural object detectors. The method approaches the NMS problem as a rescoring task, aiming to ideally assign precisely one detection per object. The proposed modification exploits the extraction of RoI representations, semantically capturing the region’s visual appearance, from information-rich feature maps computed by the detector’s intermediate layers. Experimental evaluation on two common public person detection datasets shows improved accuracy against competing methods, with acceptable inference speed. Charalampos Symeonidis, Ioannis Mademlis, Ioannis Pitas, Nikos Nikolaidis 0001 |
ICASSP | 4 |
| 2023 | ActiveFace: A Synthetic Active Perception Dataset for Face RecognitionabstractActive vision aims to enhance the efficiency of computer vision methods by enabling the capturing sensor, usually placed on a robot or, more generally, an autonomous system, to dynamically adjust its viewpoint, position or parameters in real-time. This capability allows for more precise decision-making by the model. However, training and evaluating an active vision model often necessitates a substantial number of annotated images, which must be captured under various sensor and environmental settings. These diverse images enable the model to learn the underlying dynamics of the active perception process. Unfortunately, collecting and annotating such datasets is a challenging and expensive task. It involves not only providing hand-crafted ground truth annotations but also ensuring that actions, such as moving around / towards / away from a person, are properly “imitated” to enable active vision approaches to model the corresponding active perception dynamics. To address these limitations, in this paper we propose a synthetic facial image generation pipeline specifically designed to support active face recognition, developed using a highly realistic simulation framework based on Unity. The developed pipeline allows for the generation of facial images for a set of persons at various view angles, distances, illumination conditions, and backgrounds. We demonstrate the effectiveness of our approach by training and evaluating a recently proposed embedding-based active face recognizer, as well as extending it to perform 2 axis control, leveraging the additional information provided by the generated dataset. To facilitate replication and encourage the use of the generated dataset for training and evaluating other active vision approaches, we also provide the associated assets and the developed dataset generation pipeline. Charalampos Georgiadis, Nikolaos Passalis, Nikos Nikolaidis 0001 |
MMSP | 3 |
| 2023 | A multiple-UAV architecture for autonomous media production
Ioannis Mademlis, Arturo Torres-González, Jesús Capitán, Maurizio Montagnuolo, Alberto Messina, Fulvio Negro, Cédric Le Barz, Rita Cunha, Bruno J. Guerreiro, Fan Zhang 0017, Stephen Boyle, Gregoire Guerout, Anastasios Tefas, Nikos Nikolaidis 0001, David Bull 0001, Ioannis Pitas |
Multim. Tools Appl. | 15 |
| 2023 | Using synthesized facial views for active face recognition
Efstratios Kakaletsis, Nikos Nikolaidis 0001 |
Mach. Vis. Appl. | 2 |
| 2023 | Neural Attention-Driven Non-Maximum Suppression for Person DetectionabstractNon-maximum suppression (NMS) is a post-processing step in almost every visual object detector. NMS aims to prune the number of overlapping detected candidate regions-of-interest (RoIs) on an image, in order to assign a single and spatially accurate detection to each object. The default NMS algorithm (GreedyNMS) is fairly simple and suffers from severe drawbacks, due to its need for manual tuning. A typical case of failure with high application relevance is pedestrian/person detection in the presence of occlusions, where GreedyNMS doesn't provide accurate results. This paper proposes an efficient deep neural architecture for NMS in the person detection scenario, by capturing relations of neighboring RoIs and aiming to ideally assign precisely one detection per person. The presented Seq2Seq-NMS architecture assumes a sequence-to-sequence formulation of the NMS problem, exploits the Multihead Scale-Dot Product Attention mechanism and jointly processes both geometric and visual properties of the input candidate RoIs. Thorough experimental evaluation on three public person detection datasets shows favourable results against competing methods, with acceptable inference runtime requirements. Charalampos Symeonidis, Ioannis Mademlis, Ioannis Pitas, Nikos Nikolaidis 0001 |
IEEE Trans. Image Process. | 4 |
| 2022 | Auth-Persons: A Dataset for Detecting Humans in Crowds from Aerial ViewsabstractRecent advances in artificial intelligence, control and sensing technologies have facilitated the development of autonomous Unmanned Aerial Vehicles (UAVs). Detecting humans from video input captured on-the-fly from UAVs is a critical task for ensuring flight safety, mostly handled with lightweight Deep Neural Networks (DNNs). However the detection of individual people in the case of dense crowds and/or distribution shifts (i.e., significant visual differences between the training and the test sets) is still very challenging. This paper presents AUTH-Persons, a new, annotated, publicly available video dataset, that consists of both real and synthetic footage, suitable for training and evaluating aerial-view person detection algorithms. The synthetic data were collected from 8 visually distinct photorealistic outdoor environments and they mostly contain scenes with crowded areas, where heavy occlusions and high person densities pose challenges to common detectors. This dataset is employed to evaluate the generalization performance of various state-of-the-art detection frameworks, by testing them on environments that are visually distinct from those they have been trained on. Finally, given that Non-Maximum Suppression (NMS) methods at the end of person detection pipelines typically suffer in crowded scenes, the performance of various NMS algorithms is also compared in AUTH-Persons. Charalampos Symeonidis, Ioannis Mademlis, Ioannis Pitas, Nikos Nikolaidis 0001 |
ICIP | 4 |
| 2022 | Multilayer Online Self-Acquired Knowledge DistillationabstractOnline knowledge distillation has been proposed as an auspicious approach for circumventing the flaws of the conventional offline distillation (i.e., complex, and computationally and memory demanding process). In this work, a novel online self-distillation method, named Multilayer Online Self-Acquired Knowledge Distillation (MOSAKD), is proposed, aiming to develop fast-to-execute and effective models that can comply with applications with memory and computational restrictions, e.g., robotics applications. The MOSAKD method is able to mine additional knowledge both from the intermediate and the output layers of a deep neural model in an online fashion. To achieve this goal, k-nn non-parametric density estimation for estimating the unknown probability distributions of the data samples in the feature space generated by any neural layer is used. This enables us to compute the soft labels that explicitly express the similarities of the data with the classes, by directly estimating the posterior class probabilities of the data samples. The experimental evaluation on four datasets, including a dataset of synthetic images, indicates the effectiveness of the MOSAKD method and the superiority over existing online distillation methods. Maria Tzelepi, Charalampos Symeonidis, Nikos Nikolaidis 0001, Anastasios Tefas |
ICPR | 3 |
| 2022 | OpenDR: An Open Toolkit for Enabling High Performance, Low Footprint Deep Learning for RoboticsabstractExisting Deep Learning (DL) frameworks typically do not provide ready-to-use solutions for robotics, where very specific learning, reasoning, and embodiment problems exist. Their relatively steep learning curve and the different methodologies employed by DL compared to traditional approaches, along with the high complexity of DL models, which often leads to the need of employing specialized hardware accelerators, further increase the effort and cost needed to employ DL models in robotics. Also, most of the existing DL methods follow a static inference paradigm, as inherited by the traditional computer vision pipelines, ignoring active perception, which can be employed to actively interact with the environment in order to increase perception accuracy. In this paper, we present the Open Deep Learning Toolkit for Robotics (OpenDR). OpenDR aims at developing an open, non-proprietary, efficient, and modular toolkit that can be easily used by robotics companies and research institutions to efficiently develop and deploy AI and cognition technologies to robotics applications, providing a solid step towards addressing the aforementioned challenges. We also detail the design choices, along with an abstract interface that was created to overcome these challenges. This interface can describe various robotic tasks, spanning beyond traditional DL cognition and inference, as known by existing frameworks, incorporating openness, homogeneity and robotics-oriented perception e.g., through active perception, as its core design principles. Nikolaos Passalis, S. Pedrazzi, Robert Babuska, Wolfram Burgard, D. Dias, F. Ferro, Moncef Gabbouj, Ole Green, Alexandros Iosifidis, Erdal Kayacan, Jens Kober, O. Michel, Nikos Nikolaidis 0001, Paraskevi Nousi, Roel Pieters, Maria Tzelepi, Abhinav Valada, Anastasios Tefas |
IROS | 13 |
| 2022 | A UAV Video Data Generation Framework for Improved Robustness of UAV Detection MethodsabstractRecent advances have facilitated the development and popularization of Unmanned Aerial Vehicles (UAVs) that can operate semi or fully autonomously. The real-time, accurate visual detection of UAVs is crucial for various tasks and applications including surveillance (e.g., detecting UAVs flying over restricted areas such as airports) or multi-robot systems (e.g., a swarm of UAVs that need to cooperate and avoid collisions between swarm members in GPS-denied environments). The small target-to-image ratio and large similarity with other flying objects makes the visual detection of UAVs a challenging task. In addition, data distribution shifts can have a major negative impact to UAV detection frameworks, often trained on a wide variety of datasets to achieve an adequate level of robustness. As an attempt to mitigate the effect of these issues, we present a method that can generate realistic annotated video data depicting flying UAVs, using as input real background videos and 3D UAV models. The conducted experimental evaluation showed that the synthetic data are both challenging and realistic and that detectors trained on a combination of real-world and synthetic data, exhibit an improved generalization performance, achieving better precision rates when evaluated on real datasets that are visually distinct from the corresponding real training data. Charalampos Symeonidis, Charalampos Anastasiadis, Nikos Nikolaidis 0001 |
MMSP | 3 |
| 2021 | Efficient Realistic Data Generation Framework Leveraging Deep Learning-Based Human Digitization
Charalampos Symeonidis, Paraskevi Nousi, Pavlos Tosidis, Konstantinos Tsampazis, Nikolaos Passalis, Anastasios Tefas, Nikos Nikolaidis 0001 |
EANN | 7 |
| 2021 | Multiview vision-based human crowd localization for UAV fleet flight safety
Efstratios Kakaletsis, Ioannis Mademlis, Nikos Nikolaidis 0001, Ioannis Pitas |
Signal Process. Image Commun. | 3 |
| 2020 | Shot type constraints in UAV cinematography for autonomous target tracking
Iason Karakostas, Ioannis Mademlis, Nikos Nikolaidis 0001, Ioannis Pitas |
Inf. Sci. | 3 |
| 2019 | Shot Type Feasibility in Autonomous UAV CinematographyabstractAerial cinematography relying on camera-equipped umanned aerial vehicles (UAVs), or drones, has revolutionized media production during the past years. Autonomous UAV function-alities are already being employed to a degree, in a manner structured mainly around visual target tracking. From a cinematographic point of view, the desired shot type (i.e., Close-Up, Long Shot, etc.) is the most important factor affecting the artistic result. Achieving a specific shot type depends on the target-to-camera distance and the camera focal length. However, the interaction between UAV/camera motion trajectory (e.g., Orbit, Chase, etc.) and the visual tracker requirements constrains the range of feasible shot types at each time instance. In this paper, which extends previous work, these constraints are explored for a number of standard UAV/camera motion types, UAV shot types are classified and rules regarding shot feasibility over time are analytically derived. The proposed rules are evaluated in a realistic UAV simulation environment and achieve high performance, indicating possible benefits from their integration into an intelligent shooting system. Iason Karakostas, Ioannis Mademlis, Nikos Nikolaidis 0001, Ioannis Pitas |
ICASSP | 3 |
| 2019 | Semantic Map Annotation Through UAV Video Analysis Using Deep Learning Models in ROS
Efstratios Kakaletsis, Maria Tzelepi, Pantelis I. Kaplanoglou, Charalampos Symeonidis, Nikos Nikolaidis 0001, Anastasios Tefas, Ioannis Pitas |
MMM (2) | 5 |
| 2019 | Action recognition by fusing depth video and skeletal data information
Ioannis Kapsouras, Nikos Nikolaidis 0001 |
Multim. Tools Appl. | 2 |
| 2019 | Neurons With Paraboloid Decision Boundaries for Improved Neural Network Classification PerformanceabstractIn mathematical terms, an artificial neuron computes the inner product of a d-dimensional input vector x with its weight vector w, compares it with a bias value w0and fires based on the result of this comparison. Therefore, its decision boundary is given by the equation wTx + w0= 0. In this paper, we propose replacing the linear hyperplane decision boundary of a neuron with a curved, paraboloid decision boundary. Thus, the decision boundary of the proposed paraboloid neuron is given by the equation (hTx + h0)2- ||x - p||22= 0, where h and h0denote the parameters of the directrix and p denotes the coordinates of the focus. Such paraboloid neural networks are proven to have superior recognition accuracy in a number of applications. Nikolaos Tsapanos, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | Label Propagation on Facial Images Using Similarity and Dissimilarity Labelling ConstraintsabstractIn this paper, a novel multimedia data (specifically facial images) label propagation method is presented that is based on the inclusion of labelling constraints in the objective function of the MLPP-CLP state of the art algorithm. The proposed method can incorporate pairwise facial image similarity and dissimilarity constraints into the objective function of the aforementioned method. Experiments which have been conducted on facial image labelling in three stereoscopic movies, confirm the increased labelling accuracy of the proposed method. Efstratios Kakaletsis, Olga Zoidi, Ioannis Tsingalis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
ICIP | 5 |
| 2018 | UAV Cinematography Constraints Imposed by Visual Target TrackingabstractCamera-equipped drones have recently revolutionized aerial cinematography, allowing easy acquisition of impressive footage. Although they are currently manually operated, autonomous functionalities based on machine learning and computer vision are becoming popular. However, the emerging area of autonomous UAV filming has to face several challenges, especially when visually tracking fast and unpredictably moving targets. In the latter case, an important issue is how to determine the shot types that are achievable without risking failure of the 2D visual tracker. This paper studies the constraints imposed to cinematography decision-making during autonomous UAV shooting. It focuses on formalizing and geometrically modelling common target-following UAV motion types, in order to analytically determine the maximum permissible camera focal length (therefore, the range of feasible shot types) for avoiding visual target tracking failure. Iason Karakostas, Ioannis Mademlis, Nikos Nikolaidis 0001, Ioannis Pitas |
ICIP | 3 |
| 2018 | Challenges in Autonomous UAV Cinematography: An OverviewabstractAutonomous UAV cinematography is an active research field with exciting potential for the media industry. It bears the promise of greatly facilitating UAV shooting for various applications, while significantly reducing the costs compared to manual shooting. However, the general problem has not been clearly defined and the challenges arising from current legislation and technology restrictions have not been fully charted. A complete overview of issues related to autonomous UAV cinematography is needed, pertaining to the current situation in the field, so as to guide immediate-future research. The purpose of this paper is to lay exactly this groundwork, with the expectation of providing a global perspective to multiple domain-specific research communities. The outlined issues are partitioned into challenges deriving from ethical/legal/safety considerations and from operational/production requirements. A brief survey of current technological solutions, including their limitations, is also provided for each issue. Ioannis Mademlis, Vasileios Mygdalis, Nikos Nikolaidis 0001, Ioannis Pitas |
ICME | 3 |
| 2018 | Fast constrained person identity label propagation in stereo videos using a pruned similarity matrix
Efstratios Kakaletsis, Olga Zoidi, Ioannis Tsingalis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
Signal Process. Image Commun. | 5 |
| 2018 | Positive and Negative Label PropagationsabstractThis paper extends the state-of-the-art label propagation (LP) framework in the propagation of negative labels. More specifically, the state-of-the-art LP methods propagate information of the form “the sample i should be assigned the label k.” The proposed method extends the state-of-theart framework by considering additional information of the form “the sample i should not be assigned the label k.” A theoretical analysis is presented in order to include negative LP in the problem formulation. Moreover, a method for selecting the negative labels in cases when they are not inherent from the data structure is presented. Furthermore, the incorporation of negative label information in two multigraph LP methods is presented. Finally, a discussion on the proposed algorithm extension to out of sample data, as well as scalability issues, is presented. Experimental results in various scenarios showed that the incorporation of negative label information increases, in all cases, the classification accuracy of the state of the art. Olga Zoidi, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | Summarization of human activity videos via low-rank approximationabstractSummarization of videos depicting human activities is a timely problem with important applications, e.g., in the domains of surveillance or film/TV production, that steadily becomes more relevant. Research on video summarization has mainly relied on global clustering or local (frame-by-frame) saliency methods to provide automated algorithmic solutions for key-frame extraction. This work presents a method based on selecting as key-frames video frames able to optimally reconstruct the entire video. The novelty lies in modelling the reconstruction algebraically as a Column Subset Selection Problem (CSSP), resulting in extracting key-frames that correspond to elementary visual building blocks. The problem is formulated under an optimization framework and approximately solved via a genetic algorithm. The proposed video summarization method is being evaluated using a publicly available annotated dataset and an objective evaluation metric. According to the quantitative results, it clearly outperforms the typical clustering approach. Ioannis Mademlis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
ICASSP | 3 |
| 2017 | Multimodal speaker clustering in full length movies
Ioannis Kapsouras, Anastasios Tefas, Nikos Nikolaidis 0001, Geoffroy Peeters, Elie-Laurent Benaroya, Ioannis Pitas |
Multim. Tools Appl. | 3 |
| 2017 | Automatic Detection of 3D Quality Defects in Stereoscopic Videos Using Binocular DisparityabstractThe 3D video quality issues that may disturb the human visual system and negatively impact the 3D viewing experience are well known and become more relevant as the availability of 3D video content increases, primarily through 3D cinema, but also through 3D television. In this paper, we propose four algorithms that exploit available stereo disparity information, in order to detect disturbing stereoscopic effects, namely, stereoscopic window violations, bent window effects, uncomfortable fusion object objects, and depth jump cuts on stereo videos. After detecting such issues, the proposed algorithms characterize them, based on the stress they cause to the viewer's visual system. Qualitative representative examples, quantitative experimental results on a custom-made video data set, a parameter sensitivity study, and comments on the computational complexity of the algorithms are provided, in order to assess the accuracy and the performance of stereoscopic quality defect detection. Sotirios Delis, Ioannis Mademlis, Nikos Nikolaidis 0001, Ioannis Pitas |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | Movie shot selection preserving narrative propertiesabstractAutomatic shot selection is an important aspect of movie summarization that is helpful both to producers and to audiences, e.g., for market promotion or browsing purposes. However, most of the related research has focused on shot selection based on low-level video content, which disregards semantic information, or on narrative properties extracted from text, which requires the movie script to be available. In this work, semantic shot selection based on the narrative prominence of movie characters in both the visual and the audio modalities is investigated, without the need for additional data such as a script. The output is a movie summary that only contains video frames from selected movie shots. Selection is controlled by a user-provided shot retention parameter, that removes key-frames/key-segments from the skim based on actor face appearances and speech instances. This novel process (Multimodal Shot Pruning, or MSP) is algebraically modelled as a multimodal matrix Column Subset Selection Problem, which is solved using an evolutionary computing approach. Ioannis Mademlis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
MMSP | 3 |
| 2016 | Exploiting stereoscopic disparity for augmenting human activity recognition performance
Ioannis Mademlis, Alexandros Iosifidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
Multim. Tools Appl. | 4 |
| 2016 | Big Data Analysis for Media ProductionabstractA typical high-end film production generates several terabytes of data per day, either as footage from multiple cameras or as background information regarding the set (laser scans, spherical captures, etc). This paper presents solutions to improve the integration of the multiple data sources, and understand their quality and content, which are useful both to support creative decisions on-set (or near it) and enhance the postproduction process. The main cinema specific contributions, tested on a multisource production dataset made publicly available for research purposes, are the monitoring and quality assurance of multicamera set-ups, multisource registration and acceleration of 3-D reconstruction, anthropocentric visual analysis techniques for semantic content annotation, and integrated 2-D–3-D web visualization tools. We discuss as well improvements carried out in basic techniques for acceleration, clustering and visualization, which were necessary to deal with the very large multisource data, and can be applied to other big data problems in diverse application fields. Josep Blat, Alun Evans, Hansung Kim 0001, Evren Imre, Lukás Polok, Viorela Ila, Nikos Nikolaidis 0001, Pavel Zemcík, Anastasios Tefas, Pavel Smrz, Adrian Hilton 0001, Ioannis Pitas |
Proc. IEEE | 7 |
| 2016 | Multimodal Stereoscopic Movie Summarization Conforming to Narrative CharacteristicsabstractVideo summarization is a timely and rapidly developing research field with broad commercial interest, due to the increasing availability of massive video data. Relevant algorithms face the challenge of needing to achieve a careful balance between summary compactness, enjoyability, and content coverage. The specific case of stereoscopic 3D theatrical films has become more important over the past years, but not received corresponding research attention. In this paper, a multi-stage, multimodal summarization process for such stereoscopic movies is proposed, that is able to extract a short, representative video skim conforming to narrative characteristics from a 3D film. At the initial stage, a novel, low-level video frame description method is introduced (frame moments descriptor) that compactly captures informative image statistics from luminance, color, optical flow, and stereoscopic disparity video data, both in a global and in a local scale. Thus, scene texture, illumination, motion, and geometry properties may succinctly be contained within a single frame feature descriptor, which can subsequently be employed as a building block in any key-frame extraction scheme, e.g., for intra-shot frame clustering. The computed key-frames are then used to construct a movie summary in the form of a video skim, which is post-processed in a manner that also considers the audio modality. The next stage of the proposed summarization pipeline essentially performs shot pruning, controlled by a user-provided shot retention parameter, that removes segments from the skim based on the narrative prominence of movie characters in both the visual and the audio modalities. This novel process (multimodal shot pruning) is algebraically modeled as a multimodal matrix column subset selection problem, which is solved using an evolutionary computing approach. Subsequently, disorienting editing effects induced by summarization are dealt with, through manipulation of the video skim. At the last step, the skim is suitably post-processed in order to reduce stereoscopic video defects that may cause visual fatigue. Ioannis Mademlis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
IEEE Trans. Image Process. | 3 |
| 2016 | Visual Voice Activity Detection in the WildabstractThe visual voice activity detection (V-VAD) problem in unconstrained environments is investigated in this paper. A novel method for V-VAD in the wild, exploiting local shape and motion information appearing at spatiotemporal locations of interest for facial video segment description and the bag of words model for facial video segment representation, is proposed. Facial video segment classification is subsequently performed using the state-of-the-art classification algorithms. Experimental results on one publicly available V-VAD dataset denote the effectiveness of the proposed method, since it achieves better generalization performance in unseen users, when compared to the recently proposed state-of-the-art methods. Additional results on a new unconstrained dataset provide evidence that the proposed method can be effective even in such cases in which any other existing method fails. Fotini Patrona, Alexandros Iosifidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
IEEE Trans. Multim. | 4 |
| 2015 | Multimodal Speaker Diarization Utilizing Face Clustering Information
Ioannis Kapsouras, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
ICIG (2) | 3 |
| 2015 | Visual voice activity detection based on spatiotemporal information and bag of wordsabstractA novel method for Visual Voice Activity Detection (V-VAD) that exploits local shape and motion information appearing at spatiotemporal locations of interest for facial region video description and the Bag of Words (BoW) model for facial region video representation is proposed in this paper. Facial region video classification is subsequently performed based on Single-hidden Layer Feedforward Neural (SLFN) network trained by applying the recently proposed kernel Extreme Learning Machine (kELM) algorithm on training facial videos depicting talking and non-talking persons. Experimental results on two publicly available V-VAD data sets, denote the effectiveness of the proposed method, since better generalization performance in unseen users is achieved, compared to recently proposed state-of-the-art methods. Fotini Patrona, Alexandros Iosifidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
ICIP | 4 |
| 2015 | Kernel matrix trimming for improved Kernel K-means clusteringabstractThe Kernel k-Means algorithm for clustering extends the classic k-Means clustering algorithm. It uses the kernel trick to implicitly calculate distances on a higher dimensional space, thus overcoming the classic algorithm's inability to handle data that are not linearly separable. Given a set of n elements to cluster, the n × n kernel matrix is calculated, which contains the dot products in the higher dimensional space of every possible combination of two elements. This matrix is then referenced to calculate the distance between an element and a cluster center, as per classic k-Means. In this paper, we propose a novel algorithm for zeroing elements of the kernel matrix, thus trimming the matrix, which results in reduced memory complexity and improved clustering performance. Nikolaos Tsapanos, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
ICIP | 3 |
| 2015 | Object motion analysis description in stereo video content
Theodoris Theodoridis, Konstantinos Papachristou, Nikos Nikolaidis 0001, Ioannis Pitas |
Comput. Vis. Image Underst. | 3 |
| 2015 | Person identity recognition on motion capture data using multiple actions
Ioannis Kapsouras, Nikos Nikolaidis 0001 |
Mach. Vis. Appl. | 2 |
| 2015 | A distributed framework for trimmed Kernel k-Means clustering
Nikolaos Tsapanos, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
Pattern Recognit. | 3 |
| 2015 | Facial image clustering in stereoscopic videos using double spectral analysis
Georgios Orfanidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
Signal Process. Image Commun. | 3 |
| 2014 | Facial image clustering in stereo videos using local binary patterns and double spectral analysisabstractIn this work we propose the use of local binary patterns in combination with double spectral analysis for facial image clustering applied to 3D (stereoscopic) videos. Double spectral clustering involves the fusion of two well known algorithms: Normalized cuts and spectral clustering in order to improve the clustering performance. The use of local binary patterns upon selected fiducial points on the facial images proved to be a good choice for describing images. The framework is applied on 3D videos and makes use of the additional information deriving from the existence of two channels, left and right for further improving the clustering results. Georgios Orfanidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
CIDM | 3 |
| 2014 | Stereoscopic video description for human action recognitionabstractIn this paper, a stereoscopic video description method is proposed that indirectly incorporates scene geometry information derived from stereo disparity, through the manipulation of video interest points. This approach is flexible and able to cooperate with any monocular low-level feature descriptor. The method is evaluated on the problem of recognizing complex human actions in natural settings, using a publicly available action recognition database of unconstrained stereoscopic 3D videos, coming from Hollywood movies. It is compared both against competing depth-aware approaches and a state-of-the-art monocular algorithm. Experimental results denote that the proposed approach outperforms them and achieves state-of-the-art performance. Ioannis Mademlis, Alexandros Iosifidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
CIMSIVP | 4 |
| 2014 | Semi-supervised dimensionality reduction on data with multiple representations for label propagation on facial imagesabstractIn this paper a novel method is introduced for semi-supervised dimensionality reduction on facial images extracted from stereo videos. It operates on image data with multiple representations and calculates a projection matrix that preserves locality information and a priori pairwise information, in the form of must-link and cannot-link constraints between the various data representations, as well as label information for a percentage of the data. The final data representation is a linear combination of the projections of all data representations. The performance of the proposed Semi-supervised Multiple Locality Preserving Projections method was evaluated in person identity label propagation on facial images extracted from stereo movies. Experimental results showed that the proposed method outperforms state of the art methods. Olga Zoidi, Nikos Nikolaidis 0001, Ioannis Pitas |
ICASSP | 2 |
| 2014 | A correspondence based method for activity recognition in human skeleton motion sequencesabstractIn this paper we present an algorithm for efficient activity recognition operating upon human skeleton motion sequences, derived through motion capture systems or by analyzing the output of RGB-D sensors. Our approach is driven from the assumption that, if two such sequences describe similar activities, then, consecutive frames (poses) of one sequence are expected to be similar to consecutive frames of the other. The proposed method adopts a quaternion based distance metric to calculate the similarity between poses and an intuitive method for estimating a similarity score between two skeleton motion sequences, based on the structure of a pose correspondence matrix. Our method achieved 99.5% correct activity recognition, when applied on motion capture data, in a classification task consisting of 18 classes of activities. Eftychia Fotiadou, Nikos Nikolaidis 0001 |
ICIP | 2 |
| 2014 | Stereoscopic video shot clustering into semantic concepts based on visual and disparity informationabstractIn this paper, we propose a framework for clustering shots from stereoscopic videos into clusters that correspond to semantic concepts exploiting visual and disparity information. Various color, disparity and texture descriptors are applied to shot key frames for obtaining low-level representations. Self Organizing Maps are subsequently employed upon various combinations of these representations in order to determine a lattice of representative semantic concepts. Experimental results on performances and football stereoscopic videos show that the use of disparity information leads to better clustering compared to using visual information only. Konstantinos Papachristou, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
ICIP | 3 |
| 2014 | Label propagation on data with multiple representations through multi-graph locality preserving projectionsabstractIn this paper a novel method is introduced for propagating label information on data with multiple representations. The method performs dimensionality reduction of the data by calculating a projection matrix that preserves locality information and a priori pairwise information, in the form of must-link and cannot-link constraints between the various data representations. The final data representations are then fused, in order to perform label propagation. The performance of the proposed method was evaluated on facial images extracted from stereo movies and on the UCF11 action recognition database. Experimental results showed that the proposed method outperforms state of the art methods. Olga Zoidi, Nikos Nikolaidis 0001, Ioannis Pitas |
ICIP | 2 |
| 2014 | Action Recognition in Motion Capture Data Using a Bag of Postures ApproachabstractIn this paper we introduce a novel method for movement recognition in motion capture data. A movement is regarded as a combination of basic movement patterns, the so-called dynemes. Initially a K-means variant that takes into account the periodic nature of angular data is applied on training data to discover the most discriminative dynemes. Each frame is then assigned to one of these dynemes and a histogram that describes the frequency of occurrence of these dynemes for each movement is constructed. SVM classification and sparse representation based classification are used for movement recognition on the test data. The effectiveness and robustness of this method is shown through experimental results on a standard dataset of motion capture data. Ioannis Kapsouras, Nikos Nikolaidis 0001 |
ICPR | 2 |
| 2014 | Efficient automatic detection of 3D video artifactsabstractThis paper summarizes some common artifacts in stereo video content. These artifacts lead to poor even uncomfortable 3D viewing experience. Efficient approaches for detecting three typical artifacts, sharpness mismatch, synchronization mismatch and stereoscopic window violation, are presented in detail. Sharpness mismatch is estimated by measuring the width deviations of edge pairs in depth planes. Synchronization mismatch is detected based on the motion inconsistencies of feature points between the stereoscopic channels in a short time frame. Stereoscopic window violation is detected, using connected component analysis, when objects hit the vertical frame boundaries while being in front of the virtual screen. For experiments, test sequences were created in a professional studio environment and state-of-the-art metrics were used for evaluating the proposed approaches. The experimental results show that our algorithms have considerable robustness in detecting 3D defects. Mohan Liu, Ioannis Mademlis, Patrick Ndjiki-Nya, Jean-Charles Le Quintrec, Nikos Nikolaidis 0001, Ioannis Pitas |
MMSP | 5 |
| 2014 | 2D/3D AudioVisual content analysis & descriptionabstractIn this paper, we propose a way of using the Audio-Visual Description Profile (AVDP) of the MPEG-7 standard for 2D or stereo video and multichannel audio content description. Our aim is to provide means of using AVDP in such a way, that 3D video and audio content can be correctly and consistently described. Since AVDP semantics do not include ways for dealing with 3D audiovisual content, a new semantic framework within AVDP is proposed and examples of using AVDP to describe the results of analysis algorithms on stereo video and multichannel audio content are presented. Ioannis Pitas, Konstantinos Papachristou, Nikos Nikolaidis 0001, Marco Liuni, Elie-Laurent Benaroya, Geoffroy Peeters, Axel Röbel, Antje Linnemann, Mohan Liu, Sebastian Gerke |
MMSP | 3 |
| 2014 | Shot type characterization in 2D and 3D video contentabstractDue to the enormous increase of video and image content on the web in the last decades, automatic video annotation became a necessity. The successful annotation of video and image content facilitate a successful indexing and retrieval in search databases. In this work we study a variety of possible shot type characterizations that can be assigned in a single video frame or still image. Possible ways to propagate these characterizations to a video segment (or to an entire shot) are also discussed. A method for the detection of Over-the-Shoulder shots in 3D (stereo) video is also proposed. Ioannis Tsingalis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
MMSP | 3 |
| 2014 | Action recognition on motion capture data using a dynemes and forward differences representation
Ioannis Kapsouras, Nikos Nikolaidis 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2014 | Activity-based methods for person recognition in motion capture sequences
Eftychia Fotiadou, Nikos Nikolaidis 0001 |
Pattern Recognit. Lett. | 2 |
| 2014 | Stereo object tracking with fusion of texture, color and disparity information
Olga Zoidi, Nikos Nikolaidis 0001, Anastasios Tefas, Ioannis Pitas |
Signal Process. Image Commun. | 2 |
| 2014 | Person Identity Label Propagation in Stereo VideosabstractIn this paper a novel method is introduced for propagating person identity labels on facial images extracted from stereo videos. It operates on image data with multiple representations and calculates a projection matrix that preserves locality information and a priori pairwise information, in the form of must-link and cannot-link constraints between the various data representations. The final data representation is a linear combination of the projections of all data representations. Moreover, the proposed method takes into account information obtained through data clustering. This information is exploited during the data propagation step in two ways: to regulate the similarity strength between the projected data and to indicate which samples should be selected for label propagation initialization. The performance of the proposed Multiple Locality Preserving Projections with Cluster-based Label Propagation (MLPP-CLP) method was evaluated on facial images extracted from stereo movies. Experimental results showed that the proposed method outperforms state of the art methods. Olga Zoidi, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
IEEE Trans. Multim. | 3 |
| 2013 | Feature Comparison and Feature Fusion for Traditional Dances Recognition
Ioannis Kapsouras, Stylianos Karanikolos, Nikos Nikolaidis 0001, Anastasios Tefas |
EANN (1) | 3 |
| 2013 | Variational Bayesian inference for stereo object trackingabstractIn this paper, we deal with object tracking in stereo video sequences. We introduce a Bayesian framework for utilizing the results of any conventional single channel object tracker, in order to accomplish the refinement of the tracking accuracy in the left/right video channel. In this Bayesian framework, a variational Bayesian algorithm is employed to this end, where a priori information about the object displacement (movement) over time is incorporated by means of a prior distribution. This a priori information is obtained in a pre-processing step, in which the object displacement over time is estimated. Experiments demonstrate the efficiency of the proposed post-processing methodology in terms of tracking accuracy. Giannis K. Chantas, Nikos Nikolaidis 0001, Ioannis Pitas |
ICASSP | 2 |
| 2013 | Appearance based object tracking in stereo sequencesabstractA novel algorithm is proposed, that performs tracking of rigid objects in 3D videos, without knowledge of the camera calibration parameters, by exploiting only visual information obtained from the left and right video channels, namely luminance and disparity information. The proposed algorithm exploits noisy disparity maps that have been extracted by a real-time disparity estimation algorithm. The algorithm employs two appearance-based representation methods for describing the object texture. The first one combines luminance with disparity information and the second one employs Local Steering Kernel (LSK) descriptors. Olga Zoidi, Nikos Nikolaidis 0001, Ioannis Pitas |
ICASSP | 2 |
| 2013 | Automatic 3D defects identification in stereoscopic videosabstract3DTV and 3D cinema have become quite popular during the last few years. It is now well understood that certain 3D video quality issues may have a negative effect in the 3D viewing experience. In this paper, we propose two novel algorithms that exploit available disparity information, in order to detect two disturbing stereoscopic issues, namely Stereoscopic Window Violations (SWV) and bent window effects. The algorithms' performance is tested on a number of examples. The proposed algorithms can be used for assessing the overall quality of stereoscopic video content or in order to enable fixing the detected issues in a post-production stage. Sotirios Delis, Nikos Nikolaidis 0001, Ioannis Pitas |
ICIP | 2 |
| 2012 | Variational Bayesian inference for forward-backward visual tracking in stereo sequencesabstractIn this paper we propose a Bayesian framework for accurate object tracking in stereoscopic sequences. Object detection and forward tracking are first combined according to predefined rules to get a first set of tracked regions candidates. Backward tracking is then applied to provide another set of possible object localizations. Moreover, this strategy is applied herein in stereoscopic video. We introduce a Bayesian inference algorithm which is used to merge the information of both forward and backward tracking in order to refine the tracked region localization results. Experiments, performed on face tracking, show that the proposed method provides higher tracking accuracy than a forward tracker. Giannis K. Chantas, Nikos Nikolaidis 0001, Ioannis Pitas |
ICIP | 2 |
| 2012 | Multi-view human movement recognition based on fuzzy distances and linear discriminant analysis
Alexandros Iosifidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
Comput. Vis. Image Underst. | 3 |
| 2012 | Multiplicative update rules for incremental training of multiclass support vector machines
Symeon Nikitidis, Nikos Nikolaidis 0001, Ioannis Pitas |
Pattern Recognit. | 2 |
| 2012 | Subclass discriminant Nonnegative Matrix Factorization for facial image analysis
Symeon Nikitidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
Pattern Recognit. | 3 |
| 2012 | Shape matching using a binary search tree structure of weak classifiers
Nikolaos Tsapanos, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
Pattern Recognit. | 3 |
| 2012 | Video fingerprinting using Latent Dirichlet Allocation and facial images
Nicholas Vretos, Nikos Nikolaidis 0001, Ioannis Pitas |
Pattern Recognit. | 2 |
| 2011 | Facial expression recognition using clustering discriminant Non-negative Matrix FactorizationabstractNon-negative Matrix Factorization (NMF) is among the most popular subspace methods widely used in a variety of image processing problems. Recently, a discriminant NMF method that incorporates Linear Discriminant Analysis criteria and achieves an efficient decomposition of the provided data to its discriminant parts has been proposed. However, this approach poses several limitations since it assumes that the underline data distribution forms compact sets which is often unrealistic. To remedy this limitation we regard that data inside each class form various number of clusters and apply a Clustering based Discriminant Analysis. The proposed method combines appropriate discriminant constraints in the NMF decomposition cost function in order to address the problem of finding discriminant projections that enhance class separability in the reduced dimensional projection space. Experimental results performed on the Cohn-Kanade database verified the effectiveness of the proposed method in the facial expression recognition task. Symeon Nikitidis, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
ICIP | 3 |
| 2011 | 3D facial expression recognition using Zernike moments on depth imagesabstractIn this paper we propose a new method for 3D facial expression recognition. We make use of the Zernike moments, which are calculated in the depth image of a 3D facial point cloud. Combining, the Zernike moments along with the 3D point clouds and the depth images, we succeed in tackling problems arising in facial expression recognition due to affine transformations of the data, such as translation, rotation and scaling which, in other approaches are considered very harmful in the overall accuracy of a facial expression recognition algorithm. Support vector machines are used in order to classify the previously extracted features. Results are drawn in two publicly available databases for 3D facial expression recognition. Nicholas Vretos, Nikos Nikolaidis 0001, Ioannis Pitas |
ICIP | 2 |
| 2010 | Multi-view object and human body part detection utilizing 3D scene informationabstractThe aim of this paper is to present a new method for multiview object or human body (or body part) detection. The basic idea consists of using a single view detector in every view of a scene captured by multiple cameras and then combining the results using the 3D information of the scene. The method can improve the results of the single view detector, while also localizing the object/human in the 3D space. This results in a robust way for rejecting the false detections, amending the missed detections and associating the results of the single view detector across views. Georgios Sfiris, Nikos Nikolaidis 0001, Ioannis Pitas |
ICIP | 2 |
| 2010 | Video replica detection utilizing R-trees and frame-based votingabstractA novel color-based two-step, coarse-to-fine video replica detection system is proposed in this paper. The first step uses an R-tree in order to perform a coarse selection of the database (original) videos that potentially match the query video. A training procedure that utilizes attacked versions of the database videos and aims at achieving robustness to attacks is being used. A frame-based voting procedure is also involved. A refinement step that processes the set of videos returned by the first step in order to select the final matching video (if any) follows. The performance of the system has been evaluated on a database of short videos with good results. Dimitrios Zotos, Nikos Nikolaidis 0001, Ioannis Pitas |
ICME | 2 |
| 2010 | Incremental Training of Multiclass Support Vector MachinesabstractWe present a new method for the incremental training of multiclass Support Vector Machines that provides computational efficiency for training problems in the case where the training data collection is sequentially enriched and dynamic adaptation of the classifier is required. An auxiliary function that incorporates some desired characteristics in order to provide an upper bound of the objective function which summarizes the multiclass classification task has been designed and the global minimizer for the enriched dataset is found using a warm start algorithm, since faster convergence is expected when starting from the previous global minimum. Experimental evidence on two data collections verified that our method is faster than retraining the classifier from scratch, while the achieved classification accuracy is maintained at the same level. Symeon Nikitidis, Nikos Nikolaidis 0001, Ioannis Pitas |
ICPR | 2 |
| 2010 | Movement recognition exploiting multi-view informationabstractIn this paper a novel view-invariant movement recognition method is presented. A multi-camera setup is used to capture the movement from different observation angles. Identification of the position of each camera with respect to the subject's body is achieved by a procedure based on morphological operations and the proportions of the human body. Binary body masks from frames of all cameras, consistently arranged through the previous procedure, are concatenated to produce the so-called multi-view binary mask. These masks are rescaled and vectorized to create feature vectors in the input space. Fuzzy vector quantization is performed to associate input feature vectors with movement representations and linear discriminant analysis is used to map movements in a low dimensionality discriminant feature space. Experimental results show that the method can achieve very satisfactory recognition rates. Alexandros Iosifidis, Nikos Nikolaidis 0001, Ioannis Pitas |
MMSP | 2 |
| 2010 | Image replica detection system utilizing R-trees and linear discriminant analysis
Spiros Nikolopoulos, Stefanos Zafeiriou, Nikos Nikolaidis 0001, Ioannis Pitas |
Pattern Recognit. | 3 |
| 2009 | A model-based facial expression recognition algorithm using Principal Components AnalysisabstractIn this paper, we propose a new method for facial expression recognition. We utilize the Candide facial grid and apply principal components analysis (PCA) to find the two eigenvectors of the model vertices. These eigenvectors along with the barycenter of the vertices are used to define a new coordinate system where vertices are mapped. Support vector machines (SVMs) are then used for the facial expression classification task. The method is invariant to in-plane translation and rotation as well as scaling of the face and achieves very satisfactory results. Nicholas Vretos, Nikos Nikolaidis 0001, Ioannis Pitas |
ICIP | 2 |
| 2009 | View indepedent human movement recognition from multi-view video exploiting a circular invariant posture representationabstractIn this paper a novel method for view independent human movement representation and recognition, exploiting the rich information contained in multi-view videos, is proposed. The binary masks of a multi-view posture image are first vectorized, concatenated and the view correspondence problem between train and test samples is solved using the circular shift invariance property of the discrete Fourier transform (DFT) magnitudes. Then, using fuzzy vector quantization (FVQ) and linear discriminant analysis (LDA), different movements are represented and classified. This method allows view independent movement recognition, without the use of calibrated cameras, a-priori view correspondence information or 3D model reconstruction. A multi-view video database has been constructed for the assessment of the proposed algorithm. Evaluation of this algorithm on the new database, shows that it is particularly efficient and robust, and can achieve good recognition performance. Nikolaos Gkalelis, Nikos Nikolaidis 0001, Ioannis Pitas |
ICME | 2 |
| 2009 | Frontal view recognition in multiview video sequencesabstractIn this paper, a novel method is proposed as a solution to the problem of frontal view recognition from multiview image sequences. Our aim is to correctly identify the view that corresponds to the camera placed in front of a person, or the camera whose view is closer to a frontal one. By doing so, frontal face images of the person can be acquired, in order to be used in face or facial expression recognition techniques that require frontal faces to achieve a satisfactory result. The proposed method firstly employs the Discriminant Non-Negative Matrix Factorization (DNMF) algorithm on the input images acquired from every camera. The output of the algorithm is then used as an input to a support vector machines (SVMs) system that classifies the head poses acquired from the cameras to two classes that correspond to the frontal or non frontal pose. Experiments conducted on the IDIAP database demonstrate that the proposed method achieves an accuracy of 98.6% in frontal view recognition. Irene Kotsia, Nikos Nikolaidis 0001, Ioannis Pitas |
ICME | 2 |
| 2009 | A perceptual hashing algorithm using latent dirichlet allocationabstractThis paper investigates the possibility of extracting latent aspects of a video, using visual information about humans (e.g. actors' faces), in order to develop a fingerprinting (replica detection) framework. We employ a generative probabilistic model, namely Latent Dirichlet Allocation (LDA), so as to capture latent aspects of a video, using facial semantic information derived from the video. We use the bag-of-words concept, (bag-of-faces in our case) in order to ensure exchangeability of the latent variables (e.g. topics). The video topics are modeled as a mixture of distributions of faces in each video. This generative probabilistic model has already been used in the case of text modeling with good results. Experimental results provide evidence that the proposed method performs very efficiently for video fingerprinting. Nicholas Vretos, Nikos Nikolaidis 0001, Ioannis Pitas |
ICME | 2 |
| 2009 | 3D head pose estimation in monocular video sequences by sequential camera self-calibrationabstractThis paper presents a novel approach for estimating 3D head pose in single-view video sequences acquired by an uncalibrated camera. Following the initialization by a face detector, a tracking technique localizes the faces in each frame in the video sequence. Head pose estimation is performed by using a structure from motion and self-calibration technique in a sequential way. The proposed method was applied to the IDIAP database that contains head pose ground truth data. The obtained results demonstrate that the method can estimate the head pose with satisfying accuracy. Ioannis Marras, Nikos Nikolaidis 0001, Ioannis Pitas |
MMSP | 2 |
| 2009 | Facial feature detection using distance vector fields
Stylianos Asteriadis, Nikos Nikolaidis 0001, Ioannis Pitas |
Pattern Recognit. | 2 |
| 2009 | Semantic video fingerprinting and retrieval using face information
Costas I. Cotsaces, Nikos Nikolaidis 0001, Ioannis Pitas |
Signal Process. Image Commun. | 2 |
| 2009 | 3-D Head Pose Estimation in Monocular Video Sequences Using Deformable Surfaces and Radial Basis FunctionsabstractThis paper presents a novel approach for estimating 3-D head pose in single-view video sequences. Following initialization by a face detector, a tracking technique that utilizes a 3-D deformable surface model to approximate the facial image intensity is used to track the face in the video sequence. Head pose estimation is performed by using a feature vector which is a byproduct of the equations that govern the deformation of the surface model used in the tracking. The afore-mentioned vector is used as input in a radial basis function interpolation network in order to estimate the 3-D head pose. The proposed method was applied to IDIAP head pose estimation database. The obtained results show that the method can estimate the head direction vector with very good accuracy. Michail Krinidis, Nikos Nikolaidis 0001, Ioannis Pitas |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | Visual Lip Activity Detection and Speaker Detection Using Mouth Region IntensitiesabstractIn this letter, we introduce a novel approach for lip activity detection and speaker detection, using solely visual information. The main idea in this work is to apply signal detection algorithms to a simple and easily extracted feature from the mouth region. We argue that the increased average value and standard deviation of the number of pixels with low intensities that the mouth region of a speaking person demonstrates can be used as visual cues for detecting visual speech. We then proceed in deriving a statistical algorithm that utilizes this fact for the efficient characterization of visual speech and silence in video sequences. Furthermore, we employ the lip activity detection method in order to determine the active speaker(s) in a multi-person environment. Spyridon Siatras, Nikos Nikolaidis 0001, Michail Krinidis, Ioannis Pitas |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Texture and Shape Information Fusion for Facial Action Unit RecognitionabstractA novel method that fuses texture and shape information to achieve Facial Action Unit (FAU) recognition from video sequences is proposed. In order to extract the texture information, a subspace method based on Discriminant Non- negative Matrix Factorization (DNMF) is applied on the difference images of the video sequence, calculated taking under consideration the neutral and the most expressive frame, to extract the desired classification label. The shape information consists of the deformed Candide facial grid (more specifically the grid node displacements between the neutral and the most expressive facial expression frame) that corresponds to the facial expression depicted in the video sequence. The shape information is afterwards classified using a two-class Support Vector Machine (SVM) system. The fusion of texture and shape information is performed using Median Radial Basis Functions (MRBFs) Neural Networks (NNs) in order to detect the set of present FAUs. The accuracy achieved in the Cohn-Kanade database is equal to 92.1% when recognizing the 17 FAUs that are responsible for facial expression development. Irene Kotsia, Stefanos Zafeiriou, Nikos Nikolaidis 0001, Ioannis Pitas |
ACHI | 3 |
| 2008 | Motivating class-specific nonlinear projections for single and multiple view face verificationabstractIn this paper we motivate the use of class-specific nonlinear subspace methods for face verification. The problem of face verification is considered as a two-class problem (genuine versus impostor class). The typical Fisher's linear discriminant analysis (FLDA) gives only one or two projections in a two-class problem. This is a very strict limitation to the search of discriminant dimensions. As for the FLDA for N class problems (N > 2) the transformation is not person specific. In order to remedy these limitations of FLDA, exploit the individuality of human faces and take into consideration the fact that the distribution of facial images, under different viewpoints, illumination variations and facial expression is highly complex and non-linear, novel kernel discriminant algorithms are used. The new method was tested in the face verification problem using single and multiple view datasets and found to outperform other commonly used kernel approaches. Georgios Goudelis, Stefanos Zafeiriou, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
ICIP | 4 |
| 2008 | Face-Based Digital Signatures for Video RetrievalabstractThe characterization of a video segment by a digital signature is a fundamental task in video processing. It is necessary for video indexing and retrieval, copyright protection, and other tasks. Semantic video signatures are those that are based on high-level content information rather than on low-level features of the video stream. The major advantage of such signatures is that they are highly invariant to nearly all types of distortion. A major semantic feature of a video is the appearance of specific persons in specific video frames. Because of the great amount of research that has been performed on the subject of face detection and recognition, the extraction of such information is generally tractable, or will be in the near future. We have developed a method that uses the pre-extracted output of face detection and recognition to perform fast semantic query-by-example retrieval of video segments. We also give the results of the experimental evaluation of our method on a database of real video. One advantage of our approach is that the evaluation of similarity is convolution-based, and is thus resistant to perturbations in the signature and independent of the exact boundaries of the query segment. Costas I. Cotsaces, Nikos Nikolaidis 0001, Ioannis Pitas |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Nonnegative Matrix Factorization in Polynomial Feature SpaceabstractPlenty of methods have been proposed in order to discover latent variables (features) in data sets. Such approaches include the principal component analysis (PCA), independent component analysis (ICA), factor analysis (FA), etc., to mention only a few. A recently investigated approach to decompose a data set with a given dimensionality into a lower dimensional space is the so-called nonnegative matrix factorization (NMF). Its only requirement is that both decomposition factors are nonnegative. To approximate the original data, the minimization of the NMF objective function is performed in the Euclidean space, where the difference between the original data and the factors can be minimized by employing L(2)-norm. In this paper, we propose a generalization of the NMF algorithm by translating the objective function into a Hilbert space (also called feature space) under nonnegativity constraints. With the help of kernel functions, we developed an approach that allows high-order dependencies between the basis images while keeping the nonnegativity constraints on both basis images and coefficients. Two practical applications, namely, facial expression and face recognition, show the potential of the proposed approach. Ioan Buciu, Nikos Nikolaidis 0001, Ioannis Pitas |
IEEE Trans. Neural Networks | 2 |
| 2007 | Facial Expression Recognition in Videos using a Novel Multi-Class Support Vector Machines VariantabstractIn this paper, a novel class of support vector machines (SVM) is introduced to deal with facial expression recognition. The proposed classifier incorporates statistic information about the classes under examination into the classical SVM. The developed system performs facial expression recognition in facial videos. The grid tracking and deformation algorithm used tracks the Candide grid over time as the facial expression evolves, until the frame that corresponds to the greatest facial expression intensity. The geometrical displacement of Candide nodes is used as an input to the bank of novel SVM classifiers, that are utilized to recognize the six basic facial expressions. The experiments on the Cohn-Kanade database show a recognition accuracy of 98.2%. Irene Kotsia, Nikos Nikolaidis 0001, Ioannis Pitas |
ICASSP (2) | 2 |
| 2007 | Using Deformable Surface Models to Derive a DCT-Like 2D TransformabstractThis paper introduces a 2D discrete, non-separable transform for image processing, which can be regarded as a combination of the well known discrete cosine transform (DCT) with an analytically derived quantization table that includes a compression ratio selection parameter. A 3D deformable surface model is used to approximate the image intensity and the introduced discrete transform is an intermediate step of the explicit surface deformation governing equations. The proposed transform is applied to lossy image compression and the obtained results are compared to those of a DCT-based compression scheme. Michail Krinidis, Nikos Nikolaidis 0001, Ioannis Pitas |
ICME | 2 |
| 2007 | The discrete modal transform and its application to lossy image compression
Michail Krinidis, Nikos Nikolaidis 0001, Ioannis Pitas |
Signal Process. Image Commun. | 2 |
| 2007 | 2-D Feature-Point Selection and Tracking Using 3-D Physics-Based Deformable SurfacesabstractThis paper presents a novel approach for selecting and tracking feature points in video sequences. In this approach, the image intensity is represented by a 3-D deformable surface model. The proposed approach relies on selecting and tracking feature points by exploiting the so-called generalized displacement vector that appears in the explicit surface deformation governing equations. This vector is proven to be a combination of the output of various line- and edge-detection masks, thus leading to distinct, robust features. The proposed method was compared, in terms of tracking accuracy and robustness, with a well-known tracking algorithm, Kanade-Lucas-Tomasi (KLT), and a tracking algorithm based on scale-invariant feature transform (SIFT) features. The proposed method was experimentally shown to be more precise and robust than both KLT and SIFT tracking. Moreover, the feature-point selection scheme was tested against the SIFT and Harris feature points, and it was demonstrated to provide superior results. Michail Krinidis, Nikos Nikolaidis 0001, Ioannis Pitas |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | An Optimal Detector Structure for the Fourier Descriptors Domain Watermarking of 2D Vector GraphicsabstractAbstract-Polygonal lines constitute a key graphical primitive in 2D vector graphics data. Thus, the ability to apply a digital watermark to such an entity would enable the watermarking of cartoons, drawings, and Geographical Information Systems (GIS) data in vector graphics format. This paper builds on and extends an existing algorithm that achieves polygonal line watermarking by modifying the Fourier descriptors magnitude in an imperceptible way. Watermarks embedded by this technique can be detected in rotated, translated, scaled, or reflected polygonal lines. The detection of such watermarks had been previously carried out through a correlator detector. In this paper, analysis of the statistics of the Fourier descriptors is exploited to devise an optimal blind detector. Furthermore, the problem of watermarking multiple lines, as well as other implementation issues are being addressed. Experimental results verify the imperceptibility and robustness of the proposed method. Víctor Rodríguez-Doncel, Nikos Nikolaidis 0001, Ioannis Pitas |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2006 | Temporal Video Segmentation by Graph PartitioningabstractA novel temporal video segmentation method that, in addition to abrupt cuts, can detect with very high accuracy gradual transitions such as dissolves, fades and wipes is proposed. The method relies on evaluating mutual information between multiple pairs of frames within a certain temporal frame window. This way we create a graph where the frames are nodes and the measures of similarity correspond to the weights of the edges. By finding and disconnecting the weak connections between nodes we separate the graph to subgraphs ideally corresponding to the shots. Experiments on TRECVID2004 video test set containing different types of shot transitions and significant object and camera motion inside the shots prove that the method is very efficient. Zuzana Cernekova, Nikos Nikolaidis 0001, Ioannis Pitas |
ICASSP (2) | 2 |
| 2006 | Video Indexing by Face Occurrence-Based SignaturesabstractThe extraction of a digital signature from a video segment in order to uniquely identify it, is often a necessary prerequisite for video indexing, copyright protection and other tasks. Semantic video signatures are those that are based on high-level content information rather than on low-level features of the video stream, their major advantage being that they are invariant to nearly all types of distortion. Since a major semantic feature of a video is the appearance of specific people in specific frames, we have developed a method that uses the pre-extracted output of face detection and recognition to perform fast semantic indexing and retrieval of video segments. We give the results of the experimental evaluation of our method on an artificial database created using a probabilistic model of the creation of video Costas I. Cotsaces, Nikos Nikolaidis 0001, Ioannis Pitas |
ICASSP (2) | 2 |
| 2006 | Fusion of Geometrical and Texture Information for Facial Expression RecognitionabstractA novel method based on geometrical and texture information is proposed for facial expression recognition from video sequences. The discriminant non-negative matrix factorization (DNMF) algorithm is applied at the image of the last frame of the video sequence, corresponding to the greatest intensity of the facial expression, thus extracting the texture information. A support vector machines (SVMs) system is used for the classification of the geometrical information derived from tracking the Candide grid over the video sequence. The geometrical information consists of the differences of the node coordinates between the neutral (first) and the fully expressed facial expression (last) video frame. The fusion of texture and geometrical information obtained is performed using SVMs. The accuracy achieved is 98,7% when recognizing the six basic facial expressions. Irene Kotsia, Nikos Nikolaidis 0001, Ioannis Pitas |
ICIP | 2 |
| 2006 | A Novel Replica Detection System using Binary Classifiers, R-Trees, and PCAabstractReplica detection is a prerequisite for the discovery of copyright infringement and detection of illicit content. For this purpose, content-based systems can be an efficient alternative to watermarking. Rather than imperceptibly embedding a signal, content-based systems rely-on image similarity. Certain content-based systems use adaptive classifiers to detect replicas. In such systems, a suspect image is tested against every original, which can become computationally prohibitive as the number of original images grows. In this paper, we propose using R-tree indexing to decrease the necessary number of comparisons and rapidly select the most likely originals. Experimental results show that the proposed system performs very satisfactorily and that up to 99.3% of the originals can be discarded before applying the binary classifiers. Yannick Maret, Spiros Nikolopoulos, Frédéric Dufaux, Touradj Ebrahimi, Nikos Nikolaidis 0001 |
ICIP | 5 |
| 2006 | Virtual Dental Patient: a System for Virtual Teeth DrillingabstractThis paper introduces, a virtual teeth drilling system named virtual dental patient designed to aid dentists in getting acquainted with the teeth anatomy, the handling of drilling instruments and the challenges associated with the drilling procedure. The basic aim of the system is to be used for the training of dental students. The application features a 3D model of the face and the oral cavity that can be adapted to the characteristics of a specific person and animated. Drilling using a haptic device is performed on realistic teeth models (constructed from real data), within the oral cavity. Results and intermediate steps of the drilling procedure can be saved for future use Ioannis Marras, Leontios Papaleontiou, Nikos Nikolaidis 0001, Kleoniki Lyroudia, Ioannis Pitas |
ICME | 3 |
| 2006 | Image Replica Detection using R-Trees and Linear Discriminant AnalysisabstractIn this paper a novel system for image replica detection is presented. The system uses color-based descriptors in order to extract robust features for image representation. These features are used for indexing the images in a database using an R-tree. When a query about whether a test image is a replica of an image in the database is submitted, the R-tree is traversed and a set of candidate images is retrieved. Then, in order to obtain a single result and at the same time reduce the number of decision errors the system is enhanced with linear discriminant analysis (LDA). The conducted experiments show that the proposed approach is very promising Spiros Nikolopoulos, Stefanos Zafeiriou, Panagiotis Sidiropoulos, Nikos Nikolaidis 0001, Ioannis Pitas |
ICME | 4 |
| 2006 | On the initialization of the DNMF algorithmabstractA subspace supervised learning algorithm named discriminant non-negative matrix factorization (DNMF) has been recently proposed for classifying human facial expressions. It decomposes images into a set of basis images and corresponding coefficients. Usually, the algorithm starts with random basis image and coefficient initialization. Then, at each iteration, both basis images and coefficients are updated to minimize the underlying cost function. The algorithm may need several thousands of iterations to obtain cost function minimization. We provide a way to significantly improve the speed of the algorithm convergence by constructing initial basis images that meet the sparseness and orthogonality requirements and approximate the final minimization solution. To experimentally evaluate the new approach, we have applied DNMF using the random and the proposed initialization procedure to recognize six basic facial expressions. While fewer iteration steps are needed with the proposed initialization, the recognition accuracy remains within satisfactory levels. Ioan Buciu, Nikos Nikolaidis 0001, Ioannis Pitas |
ISCAS | 2 |
| 2006 | Digital image processing techniques for the detection and removal of cracks in digitized paintingsabstractAn integrated methodology for the detection and removal of cracks on digitized paintings is presented in this paper. The cracks are detected by thresholding the output of the morphological top-hat transform. Afterward, the thin dark brush strokes which have been misidentified as cracks are removed using either a median radial basis function neural network on hue and saturation data or a semi-automatic procedure based on region growing. Finally, crack filling using order statistics filters or controlled anisotropic diffusion is performed. The methodology has been shown to perform very well on digitized paintings suffering from cracks. I. Giakoumis, Nikos Nikolaidis 0001, Ioannis Pitas |
IEEE Trans. Image Process. | 2 |
| 2005 | An Audio-Visual Database For Evaluating Person Tracking AlgorithmsabstractThis paper presents an audio-visual database that can be used as a reference database for testing and evaluation of video, audio or joint audio-visual person tracking algorithms, as well as speaker localization methods. Additional possible uses include the testing of face detection and pose estimation algorithms. A number of different scenes are included in the database, ranging from simple to complex scenes that can challenge existing algorithms. They include different subjects, with appearances that can cause problems to video tracking algorithms, (e.g. facial features such as beards, glasses, etc.), optimal and artificially created sub-optimal lighting conditions, subject movement based on simple as well as random motion trajectories, different distances from the camera/microphones and occlusion. The database incorporates ground truth data (3D position in time) originating from a commercially available 4-camera infrared (IR) tracking system. Examples of how the database can be used to evaluate video and audio tracking algorithms are also provided. Michail Krinidis, Georgios N. Stamou, Heinz Teutsch, Sascha Spors, Nikos Nikolaidis 0001, Rudolf Rabenstein |
ICASSP (2) | 5 |
| 2005 | Enhanced transform-domain correlation-based audio watermarkingabstractVarious watermarking techniques have been proposed so far, aiming at the copyright protection of audio signals. Little effort has been made, however, in taking under consideration the spectrum of the watermark sequence itself and exploiting its frequency properties. An enhanced audio watermarking technique, based on correlation detection, is introduced in this paper, where high-frequency chaotic watermarks are multiplicatively embedded in the low frequencies of the DFT domain. A series of experiments have been conducted to demonstrate both detection reliability and robustness against attacks. Anastasios Tefas, Alexia Giannoula, Nikos Nikolaidis 0001, Ioannis Pitas |
ICASSP (2) | 3 |
| 2005 | Object tracking based on morphological elastic graph matchingabstractThis paper presents a novel method for real-time tracking of objects in video sequences. Tracking is performed using the so-called morphological elastic graph matching algorithm. When applied to faces, initialization of the tracking algorithm is performed by means of a novel face detection and facial feature extraction step. The obtained results show good performance in scenes with complex background. Comparison with an existing feature-based tracking method using measures based on ground truth data proves the superiority of the proposed method. Georgios N. Stamou, Nikos Nikolaidis 0001, Ioannis Pitas |
ICIP (1) | 2 |
| 2004 | Anatomically-based 3D face and oral cavity model for creating virtual medical patientsabstractThis work presents a new hierarchical, modular and scalable model of a human face and oral cavity based on anatomical data taken from the Visible Human Project (National Institute of Health). The described model can further be adapted to any particular face and oral cavity of a human head by means of a finite elements method (FEM). The final aim will be to construct functional virtual medical patients. Georgios Moschos, Nikos Nikolaidis 0001, Ioannis Pitas |
ICME | 2 |
| 2003 | Improving the detection reliability of correlation-based watermarking techniquesabstractThe performance of watermarking schemes based on correlation detection is closely related to the frequency characteristics of the watermark sequence. In order to improve both detection reliability and robustness against attacks, embedding of watermarks with high-frequency spectrum, in the low frequencies of the DFT domain, is introduced in this paper and theoretical analysis of correlation based watermarking techniques with multiplicative embedding is performed. The proposed watermarking framework is successfully applied to audio signals, demonstrating its superiority with respect to both robustness and inaudibility. Experiments are conducted, in order to verify the validity of the theoretical analysis results. Alexia Giannoula, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
ICME | 3 |
| 2002 | Watermarking of sets of polygonal lines using fusion techniquesabstractA blind watermarking method for the copyright protection of sets of polygonal lines in vector graphics images and GIS data (elevation contour maps) is presented. The paper focuses mainly on the use of simple fusion rules for combining the detector outputs from each polygonal line in order to come up with a global detection result. Experimental comparison of the various fusion methods using both synthetic and real data (elevation maps) is provided. Alexia Giannoula, Nikos Nikolaidis 0001, Ioannis Pitas |
ICME (2) | 2 |
| 2002 | Watermark detection: benchmarking perspectivesabstractBenchmarking of watermarking algorithms is a complicated task that requires examination of a set of mutually dependent performance factors (algorithm complexity, decoding/detection performance, and perceptual quality). This paper will focus on detection/decoding performance evaluation and try to summarize its basic principles. A methodology for deriving the corresponding performance metrics will also be provided. Nikos Nikolaidis 0001, Vassilios Solachidis, Anastasios Tefas, Ioannis Pitas |
ICME (2) | 1 |
| 2001 | Bernoulli shift generated chaotic watermarks: theoretic investigationabstractThe paper statistically analyzes the behaviour of chaotic watermark signals generated by n-way Bernoulli shift maps. For this purpose, a simple blind copyright protection watermarking system is considered. The analysis involves theoretical evaluation of the system detection reliability, when a correlator detector is used. The aim of the paper is twofold: (i) to introduce the n-way Bernoulli shift generated chaotic watermarks and theoretically contemplate their properties with respect to detection reliability and (ii) to establish theoretically their potential superiority against the widely used pseudorandom watermarks. Experimental verification of the theoretical analysis results is also performed. Sofia Tsekeridou, Vassilios Solachidis, Nikos Nikolaidis 0001, Athanasios Nikolaidis, Anastasios Tefas, Ioannis Pitas |
ICASSP | 3 |
| 2001 | Digital image processing in painting restoration and archivingabstractDigital image processing and analysis can be an important tool for the restoration of works of art. This paper presents three applications of image processing in this field: a method for digital crack restoration of paintings, a technique for color restoration of old paintings and a method for mosaicing of partial images of works of art painted on curved surfaces. A digital archiving system for works of arts is also described. Nikos Nikolaidis 0001, Ioannis Pitas |
ICIP (1) | 1 |
| 2001 | A benchmarking protocol for watermarking methodsabstractA benchmarking system for watermarking algorithms is described. The proposed benchmarking system can be used to evaluate the performance of watermarking methods used for copyright protection, authentication, fingerprinting, etc. Although the system described is used for image watermarking, the general framework can be used, by introducing a different set of attacks, for benchmarking of video and audio data. Nikos Nikolaidis 0001, Sofia Tsekeridou, Anastasios Tefas, Vassilios Solachidis, Athanasios Nikolaidis, Ioannis Pitas |
ICIP (3) | 1 |
| 2001 | Statistical analysis of a watermarking system based on Bernoulli chaotic sequences
Sofia Tsekeridou, Vassilios Solachidis, Nikos Nikolaidis 0001, Athanasios Nikolaidis, Anastasios Tefas, Ioannis Pitas |
Signal Process. | 3 |
| 2001 | Robust audio watermarking in the time domainabstractThe audio watermarking method proposed in this paper offers copyright protection to an audio signal by time domain processing. The strength of audio signal modifications is limited by the necessity to produce an output signal that is perceptually similar to the original one. The watermarking method presented here does not require the use of the original signal for watermark detection. The watermark signal is generated using a key, i.e., a single number known only to the copyright owner. Watermark embedding depends on the audio signal amplitude and frequency in a way that minimizes the audibility of the watermark signal. The embedded watermark is robust to common audio signal manipulations like MPEG audio coding, cropping, time shifting, filtering, resampling, and requantization. P. Bassia, Ioannis Pitas, Nikos Nikolaidis 0001 |
IEEE Trans. Multim. | 3 |
| 2000 | Watermarking polygonal lines using Fourier descriptorsabstractA method for watermarking of polygonal lines is proposed. The watermark is embedded in the Fourier descriptors of the polygonal line causing minor distortions to the coordinates of the vertices of the polygonal line. Watermarks generated by this technique can be successfully detected even after rotation, translation, scaling and reflection of the host polygonal line. Vassilios Solachidis, Nikos Nikolaidis 0001, Ioannis Pitas |
ICASSP | 2 |
| 2000 | Fourier Descriptors Watermarking of Vector Graphics ImagesabstractA blind method for watermarking of vector graphics images through polygonal line modification is proposed. The watermark is embedded in the Fourier descriptors of the polygonal lines, causing invisible distortions to the vertices coordinates. The properties of the Fourier descriptors ensure that watermarks generated by this technique withstand rotation, translation, scaling, reflection, change of traversal starting point/direction and smoothing. Vassilios Solachidis, Nikos Nikolaidis 0001, Ioannis Pitas |
ICIP | 2 |
| 2000 | Copyright Protection of Still Images Using Self-Similar Chaotic WatermarksabstractA spatial domain watermarking algorithm for still images is proposed in this paper. The watermarks are generated by combining scaled versions of 2-D chaotic signals, in order to attain a self-similar structure. In addition, the 2-D chaotic signal generation procedure contains operations which ensure that lowpass watermarks are obtained. Such characteristics (lowpass spectrum, self-similarity) enable robustness against compression, lowpass filtering or, generally, distortions of lowpass nature, as well as, cropping and scaling. An additional advantage of the proposed algorithm is that the detection procedure does not require the original image. Sofia Tsekeridou, Nikos Nikolaidis 0001, Nicholas D. Sidiropoulos, Ioannis Pitas |
ICIP | 2 |
| 1998 | Robust image watermarking in the spatial domain
Nikos Nikolaidis 0001, Ioannis Pitas |
Signal Process. | 1 |
| 1996 | Copyright protection of images using robust digital signaturesabstractA method for copyright protection of digital images is presented. Copyright protection is achieved by embedding an "invisible" signal, known as digital signature, in the digital image. Signature casting is performed in the spatial domain by slightly modifying the intensity level of randomly selected image pixels. Signature detection is done by comparing the mean intensity value of the marked pixels against that of the not marked pixels. Statistical hypothesis testing is used for this purpose. The signature can be designed in such a way that it is resistant to JPEG compression and lowpass filtering. This is done by minimizing the energy content of the signature signal in higher frequencies. Experiments on real image data verify the effectiveness of the method. Nikos Nikolaidis 0001, Ioannis Pitas |
ICASSP | 1 |
| 1996 | Multichannel L filters based on reduced orderingabstractNonlinear multichannel signal processing is an emerging research topic with numerous applications. In this paper we use the so-called reduced ordering (R-ordering) principle to introduce a new family of L filters for vector-valued observations. The coefficients of the proposed filters can be deduced so that the filters are optimal with respect to the output mean squared error. Expressions for the unconstrained, unbiased and location invariant optimal filter coefficients are derived. The calculation of moments of the R-ordered vectors that are involved in these expressions is also discussed. Experiments with noisy two-channel vector fields and noisy color images are presented in order to demonstrate the superiority of the proposed filters over other multichannel filters. Nikos Nikolaidis 0001, Ioannis Pitas |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1996 | Order statistics learning vector quantizerabstractWe propose a novel class of learning vector quantizers (LVQs) based on multivariate data ordering principles. A special case of the novel LVQ class is the median LVQ, which uses either the marginal median or the vector median as a multivariate estimator of location. The performance of the proposed marginal median LVQ in color image quantization is demonstrated by experiments. Ioannis Pitas, Constantine Kotropoulos, Nikos Nikolaidis 0001, Ruikang Yang, Moncef Gabbouj |
IEEE Trans. Image Process. | 3 |
| 1995 | Edge detection operators for angular dataabstractPhysical quantities referring to angles, like vector direction, color hue etc., are periodic in nature. Due to this periodicity edge detectors proposed for data on the line cannot be used to detect edges in angle-related signals. In this paper we use estimators of circular dispersion to introduce edge detectors for angular signals and discuss their application in edge detection on hue images. Extensions of the notion of quasi-range to circular data are also proposed. These "circular" quasi-ranges have good and user-controlled properties as edge detectors on noisy angular signals. The performance of the proposed edge operators is evaluated on angular edges, using both qualitative and quantitative criteria. Nikos Nikolaidis 0001, Ioannis Pitas |
ICIP | 1 |
| 1994 | Application of directional statistics in vector direction estimationabstractVector direction estimation can be very important in applications like hue color component filtering or motion direction estimation from noisy motion vector fields. Vector representation and manipulation in polar coordinates greatly facilitates the accomplishment of the previous task. Based on angular estimators of location, the authors introduce a number of "circular" filters, i.e. filters for angular input data. These filters include circular mean, median and a-trimmed mean filters. Emphasis is given to a circular median filter for which the approximate output pdf as well as other interesting properties are derived. The effectiveness of angular estimators of location in noise filtering is studied in simulations involving color images. Simulation results clearly indicate that circular filters can be used effectively to remove noise when the estimation of color hue is of primary importance.> Nikos Nikolaidis 0001, Ioannis Pitas |
ICASSP (5) | 1 |
| 1994 | A Class of Order Statistics Learning Vector QuantizersabstractA novel class of Learning Vector Quantizers (LVQs) based on multivariate order statistics is proposed in order to overcome the drawback that the estimators for obtaining the reference vectors in LVQ do not have robustness either against erroneous choices for the winner vector or against the outliers that may exist in vector-valued observations. The performance of the proposed variants of LVQ is demonstrated by experiments. In the case of marginal median LVQ, its asymptotic properties are derived as well.> Ioannis Pitas, Constantine Kotropoulos, Nikos Nikolaidis 0001, Ruikang Yang, Moncef Gabbouj |
ISCAS | 3 |
| 1994 | Directional statistics in nonlinear vector field filtering
Nikos Nikolaidis 0001, Ioannis Pitas |
Signal Process. | 1 |
| 1993 | Combined Evaluation of Motion and Disparity Vector Fields for Stereoscopic Sequence Coding
Nikos Nikolaidis 0001, Ioannis Pitas, Michael G. Strintzis |
CAIP | 1 |