EDBT 2026 Demo / reviewers in the wild / expert
Ebroul Izquierdo
dblp:31/893
· DBLP profile ↗
206ranked-venue papers
29as first author
12since 2021 · last 2024
0000-0002-7142-3970ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 186 · 25 first-author · 10 since 2021Artificial intelligence and machine learning · 18 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 11 · 5 first-authorDatabases, data management, data science and information retrieval · 5Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorComputer networks · 2Security and privacy · 2Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Soft biometrics: a surveyabstractAbstract The field of biometrics research encompasses the need to associate an identity to an individual based on the persons physiological or behaviour traits. While the use of intrusive techniques such as retina scans and finger print identification has resulted in highly accurate systems, the scalability of such systems in real-world applications such as surveillance and border security has been limited. As a branch of biometrics research, the origin of soft biometrics could be traced back to need for non-intrusive solutions for extracting physiological traits of a person. Following high number of research outcomes reported in the literature on soft biometrics, this paper aims to consolidate the scope of soft biometrics research across four thematic schemes (i) a detailed review of soft biometrics research data sets, their annotation strategies and building a largest novel collection of soft traits; (ii) the assessment of metrics that affect the performance of soft biometrics system; (iii) a comparative analysis on feature and modality level fusion reported in the literature for enhancing the system performance; and (iv) a performance analysis of hybrid soft biometrics recognition system using multi-scale criterion. The paper also presents a detailed analysis on the global traits associated to person identity such as gender, age and ethnicity. The contribution of the paper is to provide a comprehensive review of scientific literature, identify open challenges and offer insights on new research directions in the filed. Bilal Hassan, Ebroul Izquierdo, Tomas Piatrik |
Multim. Tools Appl. | 2 |
| 2024 | SSFE-M: A Self-Supervised Feature Extraction Model for Enhanced Camera CalibrationabstractPredicting the corresponding homographies from broadcast frames is an important task in camera calibration. However, the prediction accuracy of many previous methods is unsatisfactory due to low-quality semantic segmentation and feature extraction. In this paper, an effective camera calibration model is proposed to address these issues. The proposed model consists of four processing stages. First, a pix2pix conditional Generative Adversarial Network (cGAN) is employed to transform the broadcast frame into the segmented image. This network can perform robustly even with limited training data. Second, a self-supervised feature extraction network is proposed to represent the segmented image as a 128- dimensional vector, which distinctively captures critical features of the segmented image. Third, the extracted feature vector is sent to a homography database to compute the best-matching homography using the K-D tree algorithm. Finally, the predicted homography is refined by running an Enhanced Correlation Coefficient (ECC) technique. The proposed model is evaluated on the 2014 World Cup and National Basketball datasets. The achieved results are compared to the state-of-the-art approaches, as well as several variants based on the U-net, VGG-16, and ResNet-50. Moreover, the difference between area-based and linebased segmentation is compared and analyzed. The experimental results demonstrate that the robustness and effectiveness of the proposed model are very competitive in diverse sports environments. Ebroul Izquierdo |
IEEE Signal Process. Lett. | 2 |
| 2023 | Efficient Convolution and Transformer-based Network for Video Frame InterpolationabstractVideo frame interpolation is an increasingly important research task with several key industrial applications in the video coding, broadcast and production sectors. Recently, transformers have been introduced to the field resulting in substantial performance gains. However, this comes at a cost of greatly increased memory usage, training and inference time. In this paper, a novel method integrating a transformer encoder and convolutional features is proposed. This network reduces the memory burden by close to 50% and runs up to four times faster during inference time compared to existing transformer-based interpolation methods. A dual-encoder architecture is introduced which combines the strength of convolutions in modelling local correlations with those of the transformer for long-range dependencies. Quantitative evaluations are conducted on various benchmarks with complex motion to showcase the robustness of the proposed method, achieving competitive performance compared to state-of-the-art interpolation networks. Issa Khalifeh, Luka Murn, Marta Mrak, Ebroul Izquierdo |
ICIP | 4 |
| 2023 | QCBA: improving rule classifiers learned from quantitative data by recovering information lost by discretisationabstractAbstract A prediscretisation of numerical attributes which is required by some rule learning algorithms is a source of inefficiencies. This paper describes new rule tuning steps that aim to recover lost information in the discretisation and new pruning techniques that may further reduce the size of rule models and improve their accuracy. The proposed QCBA method was initially developed to postprocess quantitative attributes in models generated by Classification based on associations (CBA) algorithm, but it can also be applied to the results of other rule learning approaches. We demonstrate the effectiveness on the postprocessing of models generated by five association rule classification algorithms (CBA, CMAR, CPAR, IDS, SBRL) and two first-order logic rule learners (FOIL2 and PRM). Benchmarks on 22 datasets from the UCI repository show smaller size and the overall best predictive performance for FOIL2+QCBA compared to all seven baselines. Postoptimised CBA models have a better predictive performance compared to the state-of-the-art rule learner CORELS in this benchmark. The article contains an ablation study for the individual postprocessing steps and a scalability analysis on the KDD’99 Anomaly detection dataset. Tomás Kliegr, Ebroul Izquierdo |
Appl. Intell. | 2 |
| 2023 | A Four-Point Camera Calibration Method for Sport VideosabstractThe main objective of camera calibration is to estimate the homography between a sports field template and the corresponding field in a video frame. State-of-the-art approaches exploit deep-learning technologies to achieve this aim. However, the lack of large training datasets leads to unreliable results. Search-based approaches deliver better accuracy but at the expense of a very high computational cost. In this paper, a four-point camera calibration method is proposed to address these challenges. It consists of four stages: first, a conditional generative adversarial network (cGAN) is used to generate meaningfully segmented video frames. This cGAN is a general image-to-image translation technique which has robust performance even for small training datasets. The aim is to eliminate foreground objects that greatly impede accurate estimation of four suitable points for the subsequent stage. Then, a regression network which can estimate four points from a single segmented video frame is introduced. In this stage we aim at keeping the low computational cost. In the next processing stage, the estimated four points and the original selected four points are used to calculate the homography. In this third stage a Direct Linear Transformation (DLT) algorithm is applied. In the last processing stage, the accuracy of the estimated homography is refined by running an Enhanced Correlation Coefficient (ECC) module. The proposed method is extensively evaluated using a standard 2014 World Cup dataset. The obtained results are compared to the state-of-the-art approaches and its variations based on U-net, VGG-16, ResNet-50. The obtained results show that the proposed method is superior in terms of accuracy and computational efficiency. To demonstrate the effectiveness and robustness of the proposed method, it is also validated on National Basketball, 2018 World Cup and cross dataset. In these diverse environments, the proposed method still maintains the competitive performance. Ebroul Izquierdo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Ecological Impact Assessment Framework for areas affected by Natural DisastersabstractThe forest's biodiversity consists of relations between trees, animals, the environment, and surrounding communities. Their existence required a certain balance both in number and composition. The diversity of the element itself creates a chain that connects each of the living things. Consistently, those mutual relationships are sometimes disturbed by pressures, whether man-made pressures or natural pressures. As a consequence of that event, the biodiversity loses its balance and becomes vulnerable to disaster. The fact that forest fire cases damage every living thing in the forest is becoming a massive issue in forest management. In some instances, the balance of forest biodiversity assembles an ecological resilience essential to the forest condition in combating disturbance. This paper reviews the biodiversity elements and their relationship to the extent to which elements will support ecological resilience. This is a review of 58 studies related to biodiversity balance and ecological resilience. The review discovered evidence that biodiversity components are connected and support each other. However, not every relation contributes to ecological resilience. As a result, we assess several biodiversity elements that might be useful in supporting ecological resilience, which are tree, environment, animal, and community. We also provide 2 case examples case to get the value of some biodiversity elements using a deep learning method. Arief Setyanto, Kusrini Kusrini, Gardyas Bidari Adninda, Renindya Kartikakirana, Rhisa Aidilla Suprapto, Arif Laksito, I Made Artha Agastya, Krishna Chandramouli, Andrea Majlingova, Yvonne Brodrechtová, Konstantinos P. Demestichas, Ebroul Izquierdo |
CBMI | 12 |
| 2022 | Complexity Reduction of Learned In-Loop Filtering in Video CodingabstractIn video coding, in-loop filters are applied on reconstructed video frames to enhance their perceptual quality, before storing the frames for output. Conventional in-loop filters are obtained by hand-crafted methods. Recently, learned filters based on convolutional neural networks that utilize attention mechanisms have been shown to improve upon traditional techniques. However, these solutions are typically significantly more computationally expensive, limiting their potential for practical applications. The proposed method uses a novel combination of sparsity and structured pruning for complexity reduction of learned in-loop filters. This is done through a three-step training process of magnitude-guided weight pruning, insignificant neuron identification and removal, and fine-tuning. Through initial tests we find that network parameters can be significantly reduced with a minimal impact on network performance. Woody Bayliss, Luka Murn, Ebroul Izquierdo, Qianni Zhang, Marta Mrak |
ISCAS | 3 |
| 2022 | A Fast and Effective Framework for Camera Calibration in Sport VideosabstractComputing the relative homography between the sports field template and the corresponding field in a video frame is an important task in camera calibration. In this paper, a fast and effective framework is proposed for addressing this task. The proposed framework has three processing modules. First, a semantic segmentation network is presented to obtain the segmented video frames. Second, a regression network is developed and combined with the direct linear transformation (DLT) algorithm to compute the homography. Third, the enhanced correlation coefficient (ECC) technique is leveraged to refine the estimated homography. The proposed framework is evaluated on 2014 World Cup dataset. The experimental results are compared to the state-of-the-art approaches. The experimental results demonstrate that the accuracy in the proposed framework is superior and the computation speed is competitive. Ebroul Izquierdo |
VCIP | 2 |
| 2021 | ApparelNet: Person Verification Encompassing Auxiliary Attachments VariationabstractAuxiliary attachments change is more frequent than essential clothing in human beings. Usually, auxiliary attachments include coat, jumper, hat and bag etc. It is one of the hardest recognition task in machine vision. It becomes more difficult if a specific person is reappearing after a longer time period while the other influential factors are angle variation and walking speed etc. One of the key application areas for person verification is border control where auxiliary attachments variation is more common. It is usually reflection of ethnicity or fashion. In machine vision, availability of such datasets is very limited, in particular, having reappearance after longer time period i.e., more than weeks or months. To overcome limited dataset problem, transfer learning is a leading solution for improved verification. In this paper, we proposed an aggregated deep learning model called, ApparelNet, more specifically for person verification in border control environment. We used Front-View Gait (FVG) to evaluate the performance of our aggregated model. The FVG is a pedestrian dataset of people encompassing auxiliary attachments variation, having three different angles from the camera and three different walking speeds. Our ApparelNet acquires single image based detection confidence using OpenPose and later additional layers of pre-trained EfficientNetB0 are trained on custom FVG dataset, including fine tuning of the overall EfficientNetB0. The EfficientNetB0 is highly efficient and scalable transfer learning model from the family of Deep-CNN. Overall, our ApparelNet reported training and validation accuracy of 98%, while looking at border control scenario, model verification is performed by selecting random images of 12 different individuals and prediction probability is computed which accumulates to 96%. In our opinion, the model has strong candidature for person reidentification where goal is one-to-many recognition. It may become an ancillary component of any biometrics system too. Bilal Hassan, Ebroul Izquierdo |
MMSP | 2 |
| 2021 | A High Accuracy Camera Calibration Method for Sport VideosabstractCamera calibration for sport videos enables precise and natural delivery of graphics on video footage and several other special effects. This in turns substantially improves the visual experience in the audience and facilitates sports analysis within or after the live show. In this paper, we propose a high accuracy camera calibration method for sport videos. First, we generate a homography database by uniformly sampling camera parameters. This database includes more than 91 thousand different homography matrices. Then, we use the conditional generative adversarial network (cGAN) to achieve semantic segmentation splitting the broadcast frames into four classes. In a subsequent processing step, we build an effective feature extraction network to extract the feature of semantic segmented images. After that, we search for the feature in the database to find the best matching homography. Finally, we refine the homography by image alignment. In a comprehensive evaluation using the 2014 World Cup dataset, our method outperforms other state-of-the-art techniques. Ebroul Izquierdo |
VCIP | 2 |
| 2021 | Large-scale multimedia signal processing for security and digital forensics [SI 1163]
Ebroul Izquierdo, Krishna Chandramouli, Anthony Tung Shuen Ho, Hyoung Joong Kim |
Multim. Tools Appl. | 1 |
| 2021 | Modeling Acceleration Properties for Flexible INTRA HEVC Complexity ControlabstractIt is a very well-known fact, that the high complexity of the High Efficiency Video Coding standard (HEVC) is the main hurdle for its wide deployment and use. To tackle this problem, a number of recent research outcomes exploit heuristic algorithms and machine learning, including deep learning, to reduce the coding complexity. However, in most cases, each encoder module, i.e., encoding process, is first accelerated individually, and then different acceleration algorithms are manually combined. Without a holistic strategy, the acceleration potential of multi-module combination is not exploited and the Rate-Distortion (RD) loss is generally not well controlled. To tackle these shortcomings, this paper exploits the acceleration properties of different modules, i.e., the numerical representation of potential time saving and possible RD loss, from which a heuristic model is explored. Then a Heuristic Model Oriented Framework (HMOF) is proposed which adapts the properties of modules to underlying acceleration algorithms. In the framework, two advanced acceleration algorithms, including Border Considered CNN (BC-CNN)-based Coding Unit (CU) partition and Naive Bayes-based Prediction Unit (PU) partition, are proposed for the CU and PU modules, respectively. Further, by leveraging the heuristic model as the guidance to combine the proposed acceleration algorithms, HMOF is globally optimized, where different time saving budgets are wisely allocated to different modules and a theoretically minimal RD loss is achieved. According to the experimental results, through fusing a suitable deep learning technique and a Bayes-Based prediction, the proposed acceleration framework HMOF enable multiple acceleration choices. Here the proposed joint optimization strategy help to make a choice leading to the best cost-performance. Furthermore, within the proposed framework, intra coding time can be precisely controlled with negligible Bjøntegaard delta bit-rate (BDBR) loss. In this context, as a complexity control method, HMOF outperforms the state-of-the-art complexity reduction algorithms under a similar complexity reduction ratio. These results partially demonstrate the superiority of the proposed technique. Yan Huang 0033, Li Song 0001, Rong Xie 0004, Ebroul Izquierdo, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Adaptive Aggregated Tracklet Linking for Multi-Face TrackingabstractThis paper addresses the problem of multi-face tracking. The proposed framework first assembles discrete detections into short reliable sequences (code-named `tracklets') and second, links them via hierarchical clustering. A key challenge in face tracking in real world scenarios `on the wild' or under adverse conditions is the drastic variations across scenes (e.g: variations in illumination, pose, occlusion, motion blur, etc.). To this end, the paper presents a simple hierarchical clustering approach which exploits the intra-scene compactness while minimizing the impact of inter-scene variations. The algorithm adapts to different domains by dynamically learning domain-specific parameters through density dispersion analysis. The proposed framework is evaluated on three databases, which demonstrate competitive performance with state-of the-art clustering and tracklet linking approaches. In addition, the paper presents a challenging database which is publicly available, on which the end-to-end tracking performance is reported. Samadhi Wickrama Arachchilage, Ebroul Izquierdo |
ICIP | 2 |
| 2020 | Clusterface: Joint Clustering and Classification for Set-Based Face RecognitionabstractDeep learning technology has enabled successful modeling of complex facial features when high quality images are available. Nonetheless, accurate modeling and recognition of human faces in real world scenarios `on the wild' or under adverse conditions remains an open problem. When unconstrained faces are mapped into deep features, variations such as illumination, pose, occlusion, etc., can create inconsistencies in the resultant feature space. Hence, deriving conclusions based on direct associations could lead to degraded performance. This rises the requirement for a basic feature space analysis prior to face recognition. This paper devises a joint clustering and classification scheme which learns deep face associations in an easy-to-hard way. Our method is based on hierarchical clustering where the early iterations tend to preserve high reliability. The rationale of our method is that a reliable clustering result can provide insights on the distribution of the feature space, that can guide the classification that follows. Experimental evaluations on three tasks, face verification, face identification and rank-order search, demonstrates better or competitive performance compared to the state-of-the-art, on all three experiments. Samadhi Wickrama Arachchilage, Ebroul Izquierdo |
ICPR | 2 |
| 2020 | SSDL: Self-Supervised Domain Learning for Improved Face RecognitionabstractFace recognition in unconstrained environments is challenging due to variations in illumination, quality of sensing, motion blur and etc. An individual's face appearance can vary drastically under different conditions creating a gap between train (source) and varying test (target) data. The domain gap could cause decreased performance levels in direct knowledge transfer from source to target. Despite fine-tuning with domain specific data could be an effective solution, collecting and annotating data for all domains is extremely expensive. To this end, we propose a self-supervised domain learning (SSDL) scheme that trains on triplets mined from unlabelled data. A key factor in effective discriminative learning, is selecting informative triplets. Building on most confident predictions, we follow an “easy-to-hard” scheme of alternate triplet mining and self-learning. Comprehensive experiments on four different benchmarks show that SSDL generalizes well on different domains. Samadhi Wickrama Arachchilage, Ebroul Izquierdo |
ICPR | 2 |
| 2020 | Deep Homography-Based Video StabilizationabstractVideo stabilization is fundamental for providing good Quality of Experience for viewers and generating suitable content for video applications. In this scenario, Digital Video Stabilization (DVS) is convenient and economical for casual or amateur recording because it neither requires specific equipment nor demands knowledge of the device used for recording. Although DVS has been a research topic for decades, with a number of proposals from industry and academia, traditional methods tend to fail in a number of scenarios, e.g. with occlusion, textureless areas, parallax, dark areas, amongst others. On the other hand, defining a smooth camera path is a hard task in Deep Learning scenarios. This paper proposes a video stabilization system based on traditional and Deep Learning methods. First, we leverage Spatial Transformer Networks (STNs) to learn transformation parameters between image pairs, then utilize this knowledge to stabilize videos: we obtain the motion parameters between frame pairs and then smooth the camera path using moving averages. Our approach aims at combining the strengths of both Deep Learning and traditional methods: the ability of STNs to estimate motion parameters between two frames and the effectiveness of moving averages to smooth camera paths. Experimental results show that our system outperforms state-of- the-art proposals and a commercial solution. Maria Silvia Ito, Ebroul Izquierdo |
ICPR | 2 |
| 2020 | One Shot Logo Recognition Based on Siamese Neural NetworksabstractThis work presents an approach for one-shot logo recognition that relies on a Siamese neural network (SNN) embedded with a pre-trained model that is fine-tuned on a challenging logo dataset. Although the model is fine-tuned using logo images, the training and testing datasets do not have overlapped categories; meaning that, all the classes used for testing the one-shot recognition framework remain unseen during the fine-tuning process. The recognition process follows the standard SNN approach in which a pair of input images are encoded by each sister network. The encoded outputs for each image are afterwards compared using a trained metric and thresholded to define matches and mismatches. The proposed approach achieves an accuracy of 77.07% under the one-shot constraints in the QMUL-OpenLogo dataset. Code is available at https://github.com/cjvargasc/oneshot_siamese/. Camilo Vargas, Qianni Zhang, Ebroul Izquierdo |
ICMR | 3 |
| 2019 | Fast Inter-prediction Based on Decision Trees for AV1 EncodingabstractThe AOMedia Video 1 (AV1) standard can achieve considerable compression efficiency thanks to the usage of many advanced tools and improvements, such as advanced inter-prediction modes. However, these come at the cost of high computational complexity of encoder, which may limit the benefits of the standard in practical applications. This paper shows that not all sequences benefit from using all such modes, which indicates that a number of encoder optimisations can be introduced to speed up AV1 encoding. A method based on decision trees is proposed to selectively decide whether to test all inter modes. Appropriate features are extracted and used to perform the decision for each block. Experimental results show that the proposed method can reduce the encoding time on average by 43.4% with limited impact on the coding efficiency. Jieon Kim, Saverio G. Blasi, André Seixas Dias, Marta Mrak, Ebroul Izquierdo |
ICASSP | 5 |
| 2019 | CNN Accelerated Intra Video Coding, Where Is the Upper Bound?abstractThe very high complexity of the High Efficiency Video Coding standard (HEVC) is the main hurdle for its wide deployment and use. To tackle this problem, a number of recent research outcomes exploit Convolutional Neural Network (CNN) in each HEVC module for reducing the coding complexity. In this paper an effective method to analyse the potential of CNN techniques to reduce the computational cost of HEVC is proposed. A theoretical upper bound for the effectiveness of this approach in common HEVC modules is investigated. The theoretical maximum of learning-based complexity reduction in HEVC and possible reasons for Rate-Distortion (RD) loss are investigated. On the basis of this analysis, an Intra Video Coding Acceleration (IVCA) scheme is proposed, where Border Considered CNN (BC-CNN) based Coding Unit (CU) partition and heuristic Prediction Unit (PU) partition are seamlessly integrated. According to the experimental results, 66.7% of intra coding time can be saved with negligible 1.71% Bjøntegaard delta bit-rate (BDBR) loss. These results partially demonstrate the superiority of the proposed technique against other state-of-the-art approaches aiming at reducing HEVC complexity in intra mode. Yan Huang 0033, Li Song 0001, Ebroul Izquierdo |
PCS | 3 |
| 2019 | A Framework for Real-Time Face-RecognitionabstractThe advent and wide use of deep-learning technology has enabled tremendous advancements in the accuracy of face recognition under favourable conditions. Nonetheless, the reported near-perfect performance on classic benchmarks like lfw, does not include complications in unconstrained application. The research reported in this paper addresses some of the critical challenges of face recognition under adverse conditions. In this context, we introduce an end-to-end framework for real-time video-based face recognition. This system detects, tracks and recognizes individuals from live video feed. The proposed system addresses three key challenges of video-based face recognition systems: end-to-end computational complexity, in the wild recognition and multi-person recognition. We exploit sophisticated deep neural networks for face detection and facial feature extraction, while minimizing the computational overhead from the rest of the modules in the recognition pipeline. A comprehensive evaluation shows that the proposed system can effectively recognize faces under unconstrained conditions, at elevated frames per second rates. Samadhi Wickrama Arachchilage, Ebroul Izquierdo |
VCIP | 2 |
| 2019 | A Dataset and Evaluation Framework for Deep Learning Based Video Stabilization SystemsabstractAlthough traditional methods for video stabilization are time consuming and prone to failure, more promising Deep Learning based solutions have not been thoroughly studied yet. This is mostly due to the lack of suitable training and testing datasets. To address this problem, this paper introduces a comprehensive dataset for training and assessing techniques for video stabilization. It consists of many shaky video sequences, their stable videos and the respective motion parameters that map each frame of the stable video into the corresponding frame in the unstable video. An important aspect of the dataset is the availability of motion parameters. This critical feature enables better assessment of any video stabilization technique, since it allows for additional comparison of the estimated motion parameters with the provided Ground Truth motion parameters. Using this dataset, an evaluation framework for video stabilization technology is also introduced. To demonstrate the practical use of the introduced dataset and the evaluation framework, we compare the performance of two state of the art techniques for video stabilization. The results of this extensive evaluation are presented and both the database and the evaluation framework are also provided in this paper. Maria Silvia Ito, Ebroul Izquierdo |
VCIP | 2 |
| 2019 | Real-Time Multi-Target Multi-Camera Tracking with Spatial-Temporal InformationabstractVideo security monitoring has always been an important mission for safety reason. In an entire surveillance system, there are usually several cameras distributed sparsely to cover a wide range of public areas (e.g., school, shopping mall or infrastructure). Tracking person through this cameras network is challenging due to different camera perspectives, illumination changes and pose variations. Several algorithms for Multi-Target Multi-Camera tracking (MTMCT) have been proposed in offline method which has delay in getting result. Addressing the need for real-time computation of people tracks through multi camera, the paper proposes an online tracking algorithm. The contributions include (1) online real-time framework which can be used in practical application, (2) extend a single camera multi object tracking (MOT) algorithm to be suitable for multi-camera tracking and (3) use spatial-temporal information to strengthen cross camera person recall performance. The proposed algorithm has been benchmarked against the literature review of MTMC algorithms. Ebroul Izquierdo |
VCIP | 2 |
| 2019 | Advanced Super-Resolution Using Lossless Pooling Convolutional NetworksabstractIn this paper, we present a novel deep learning-based approach for still image super-resolution, that unlike the mainstream models does not rely solely on the input low resolution image for high quality upsampling, and takes advantage of a set of artificially created auxiliary self-replicas of the input image that are incorporated in the neural network to create an enhanced and accurate upscaling scheme. Inclusion of the proposed lossless pooling layers, and the fusion of the input self-replicas enable the model to exploit the high correlation between multiple instances of the same content, and eventually result in significant improvements in the quality of the super-resolution, which is confirmed by extensive evaluations. Farzad Toutounchi, Ebroul Izquierdo |
WACV | 2 |
| 2019 | Time-Constrained Video Delivery Using Adaptive Coding ParametersabstractIn some applications, video content needs to be encoded and uploaded to a remote destination within a pre-defined amount of time. In order to guarantee that the overall processing time does not exceed certain time constraints, a system that performs joint encoding and uploading time control is needed. Such a system requires the flexibility to control both the time spent in the encoding process and the bit-rate. This is because the latter significantly influences the transmission time, especially for transmissions under low-bandwidth constraints. This paper proposes a novel approach to address this challenge by adapting the quantization parameters (QP) of a video encoder in order to meet overall processing time requirements, including both encoding and uploading time. The proposed QP adaptation approach relies on mechanisms to accurately predict the encoding time and bit-rate during the encoding process for the incoming group of pictures. This in turn allows adequate QP selections that result in accurately meeting the overall time constraints. A comprehensive experimental evaluation shows that the proposed QP adaptation approach can accurately meet the overall time constraints by efficiently adapting the encoding process to different target times and bandwidth conditions. André Seixas Dias, Shenglan Huang, Saverio G. Blasi, Marta Mrak, Ebroul Izquierdo |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2018 | An Effective Strategy for Early Skip Mode Decision in HEVCabstractHigh Efficiency Video Coding (HEVC) can provide up to twice better compression efficiency than its predecessor, H.264/Advanced Video Coding (AVC) but its high encoding complexity is a main obstacle for practical applications. In this paper, an early SKIP mode decision scheme is proposed to reduce the complexity of inter mode decision in HEVC. The proposed method firstly employs the merge_flag information of spatial neighbouring candidates for the SKIP mode. In case the best spatial candidate of the SKIP mode is encoded with merge_flag equal to one, the rate-distortion (RD) cost value of 2N×2N Merge or inter 2N×2N PU mode is further exploited to determine the early SKIP mode for the current CU. Experimental results show that the proposed method can provide on average 39.07% encoder speed-up for an average 0.53% bitrate increase in Random Access (RA) main configuration compared to the HEVC test model (HM) 16.13 reference software. Jieon Kim, Ebroul Izquierdo |
ICIP | 2 |
| 2018 | Complexity-Constrained Video Encoding and Delivery using Configuration Transfer MatrixabstractMany applications require video content to be encoded and uploaded under specific complexity constraints. While many speed-ups are available in practical video encoder implementations, it is difficult to predict the impact of such techniques on the actual content being encoded and therefore select the best configuration to meet the given constraints. A method is proposed in this paper to automatically select the encoder configuration in order to meet complexity constraints in terms of encoding and uploading time, using a pre-trained encoder configuration transfer matrix. The algorithm ensures that the content is processed within the specified targets, as presented in the experimental evaluation, where it is shown that the encoder can accurately meet specific constraints under a variety of conditions. Saverio G. Blasi, André Seixas Dias, Marta Mrak, Shenglan Huang, Ebroul Izquierdo |
PCS | 5 |
| 2018 | Estimation of Rate Control Parameters for Video Coding Using CNNabstractRate-control is essential to ensure efficient video delivery. Typical rate-control algorithms rely on bit allocation strategies, to appropriately distribute bits among frames. As reference frames are essential for exploiting temporal redundancies, intra frames are usually assigned a larger portion of the available bits. In this paper, an accurate method to estimate number of bits and quality of intra frames is proposed, which can be used for bit allocation in a rate-control scheme. The algorithm is based on deep learning, where networks are trained using the original frames as inputs, while distortions and sizes of compressed frames after encoding are used as ground truths. Two approaches are proposed where either local or global distortions are predicted. María Santamaría 0001, Ebroul Izquierdo, Saverio G. Blasi, Marta Mrak |
VCIP | 2 |
| 2018 | Foreground Segmentation with Tree-Structured Sparse RPCAabstractBackground subtraction is a fundamental video analysis technique that consists of creation of a background model that allows distinguishing foreground pixels. We present a new method in which the image sequence is assumed to be made up of the sum of a low-rank background matrix and a dynamic tree-structured sparse matrix. The decomposition task is then solved using our approximated Robust Principal Component Analysis (ARPCA) method which is an extension to the RPCA that can handle camera motion and noise. Our model dynamically estimates the support of the foreground regions via a superpixel generation step, so that spatial coherence can be imposed on these regions. Unlike conventional smoothness constraints such as MRF, our method is able to obtain crisp and meaningful foreground regions, and in general, handles large dynamic background motion better. To reduce the dimensionality and the curse of scale that is persistent in the RPCA-based methods, we model the background via Column Subset Selection Problem, that reduces the order of complexity and hence decreases computation time. Comprehensive evaluation on four benchmark datasets demonstrate the effectiveness of our method in outperforming state-of-the-art alternatives. Salehe Erfanian Ebadi, Ebroul Izquierdo |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Deep Co-Space: Sample Mining Across Feature Transformation for Semi-Supervised LearningabstractAiming at improving the performance of visual classification in a cost-effective manner, this paper proposes an incremental semi-supervised learning paradigm called deep co-space (DCS). Unlike many conventional semi-supervised learning methods usually performed within a fixed feature space, our DCS gradually propagates information from labeled samples to unlabeled ones along with deep feature learning. We regard deep feature learning as a series of steps pursuing feature transformation, i.e., projecting the samples from a previous space into a new one, which tends to select the reliable unlabeled samples with respect to this setting. Specifically, for each unlabeled image instance, we measure its reliability by calculating the category variations of feature transformation from two different neighborhood variation perspectives and merged them into a unified sample mining criterion deriving from Hellinger distance. Then, those samples keeping stable correlation to their neighboring samples (i.e., having small category variation in distribution) across the successive feature space transformation are automatically received labels and incorporated into the model for incrementally training in terms of classification. Our extensive experiments on standard image classification benchmarks (e.g., Caltech-256 and SUN-397) demonstrate that the proposed framework is capable of effectively mining from large-scale unlabeled images, which boosts image classification performance and achieves promising results compared with other semi-supervised learning methods. Ziliang Chen 0001, Keze Wang, Xiao Wang 0014, Ebroul Izquierdo, Liang Lin 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2018 | ADORE: An Adaptive Holons Representation Framework for Human Pose EstimationabstractIn this paper, the problem of human pose estimation in a 2D still image is addressed. A framework called adaptive holons representation (ADORE) that takes advantage of local and global cues is proposed to improve the pose estimation accuracy. In particular, ADORE is made up of two components: 1) the holons part, independent losses pose nets (ILPNs) is designed to first infer joints location on the global level; and 2) the adaptive part, convolutional local detectors (CLDs) is proposed to subsequently detect the joints in the potential regions generated by ILPN. Pose estimation is formulated as a classification problem toward body joints in ILPN, which consists of two independent loss layers that, respectively, instruct the learning of x and y coordinates of a joint. Experimental results on two challenging benchmark tasks demonstrate that our proposed framework is more efficient than other deep models and has desirable performance. Xiuyuan Chen, Qianni Zhang, Ebroul Izquierdo |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2018 | CUNet: A Compact Unsupervised Network For Image ClassificationabstractIn this paper, we propose a compact network called compact unsupervised network (CUNet) to address the image classification challenge. Contrasting the usual learning approach of convolutional neural networks, learning is achieved by the simple K-means on diverse image patches. This approach performs well even with scarcely labeled training images, greatly reducing the computational cost, while maintaining high discriminative power. Furthermore, we propose a new weighted pooling method in which different weighting values of adjacent neurons are considered. This strategy leads to improved classification since the network becomes more robust against small image distortions. In the output layer, CUNet integrates feature maps obtained in the last hidden layer, and straightforwardly computes histograms in nonoverlapped blocks. To reduce feature redundancy, we also implement the max-pooling operation on adjacent blocks to select the most competitive features. Comprehensive experiments on well-established databases are conducted to validate the classification performances of the introduced CUNet approach. Mengdie Mao, Gaipeng Kong, Xi Wu 0004, Qianni Zhang, Xiaochun Cao, Ebroul Izquierdo |
IEEE Trans. Multim. | 8 |
| 2017 | Multiview 3D sensing and analysis for high quality point cloud reconstructionabstractMultiview 3D reconstruction techniques enable digital reconstruction of 3D objects from the real world by fusing different viewpoints of the same object into a single 3D representation. This process is by no means trivial and the acquisition of high quality point cloud representations of dynamic 3D objects is still an open problem. In this paper, an approach for high fidelity 3D point cloud generation using low cost 3D sensing hardware is presented. The proposed approach runs in an efficient low-cost hardware setting based on several Kinect v2 scanners connected to a single PC. It performs autocalibration and runs in real-time exploiting an efficient composition of several filtering methods including Radius Outlier Removal (ROR), Weighted Median filter (WM) and Weighted Inter-Frame Average filtering (WIFA). The performance of the proposed method has been demonstrated through efficient acquisition of dense 3D point clouds of moving objects. Andrej Satnik, Ebroul Izquierdo, Richard Orjesek |
ICMV | 2 |
| 2017 | An efficient video super-resolution approach based on sparse representationabstractBlock-wise super-resolution methods, and in particular sparse representation based approaches, often focus on spatial upsampling of still images. Applying such models to videos is an extremely time consuming process due to the expensive sparse coding process for every block of each frame, and the conventional exhaustive overlapping blocks processing for reducing the blocking artifacts. In this paper, we introduce an approach that enables us to skip the sparse representation process for the static portions of a video by predicting the high resolution blocks from previously super-resolved frames, exploiting the significant amount of correlation between the adjacent frames in a video. Our approach is also enhanced by adaptive in-loop filters for removing the blocking and pixel-wise artifacts, replacing the overlapping blocks structure of super-resolution. Our method provides comparable results in terms of image quality with respect to the state of the art, and can reduce the processing time of the super-resolution methods by 63% in average. It enables the block based super-resolution methods to be applied in upsampling of videos to very high resolutions with a reasonable processing time. Farzad Toutounchi, Valia Guerra-Ones, Ebroul Izquierdo |
MMSP | 3 |
| 2017 | Adaptive packet scheduling for scalable video streaming with network coding
Shenglan Huang, Ebroul Izquierdo, Pengwei Hao |
J. Vis. Commun. Image Represent. | 2 |
| 2017 | Content-Adaptive Sketch Portrait Generation by Decompositional Representation LearningabstractSketch portrait generation benefits a wide range of applications such as digital entertainment and law enforcement. Although plenty of efforts have been dedicated to this task, several issues still remain unsolved for generating vivid and detail-preserving personal sketch portraits. For example, quite a few artifacts may exist in synthesizing hairpins and glasses, and textural details may be lost in the regions of hair or mustache. Moreover, the generalization ability of current systems is somewhat limited since they usually require elaborately collecting a dictionary of examples or carefully tuning features/components. In this paper, we present a novel representation learning framework that generates an end-to-end photo-sketch mapping through structure and texture decomposition. In the training stage, we first decompose the input face photo into different components according to their representational contents (i.e., structural and textural parts) by using a pre-trained convolutional neural network (CNN). Then, we utilize a branched fully CNN for learning structural and textural representations, respectively. In addition, we design a sorted matching mean square error metric to measure texture patterns in the loss function. In the stage of sketch rendering, our approach automatically generates structural and textural representations for the input photo and produces the final result via a probabilistic fusion scheme. Extensive experiments on several challenging benchmarks suggest that our approach outperforms example-based synthesis algorithms in terms of both perceptual and objective metrics. In addition, the proposed method also has better generalization ability across data set without additional training. Dongyu Zhang 0002, Liang Lin 0004, Tianshui Chen, Xian Wu 0007, Wenwei Tan, Ebroul Izquierdo |
IEEE Trans. Image Process. | 6 |
| 2016 | Boosting Zero-Shot Image Classification via Pairwise Relationship Learning
Hefeng Wu, Shujin Lin, Ebroul Izquierdo |
ACCV (1) | 6 |
| 2016 | Foreground Segmentation via Dynamic Tree-Structured Sparse RPCA
Salehe Erfanian Ebadi, Ebroul Izquierdo |
ECCV (1) | 2 |
| 2016 | Adaptive packet scheduling for scalable video streaming with network codingabstractIn advanced Peer-to-Peer delivery systems, each user downloads a video stream and at the same time uploads the same stream to other users. In live streaming, scalable video streaming is also proposed to allow partial decoding of the video streaming at a reduced resolution, frame-rate or quality, adapting to different display requirements and receptions conditions of heterogeneous receivers. In live streaming applications, Network Coding (NC) has the potential to bring substantially higher throughput while reducing transmission delays. However, the bandwidth efficiency of NC transmission is still a problem to live media streaming due to the lack of synchronization among peers. To solve this problem, we introduce a scalable streaming system, which includes a layer selection algorithm and a distributed packet scheduling algorithm. The proposed system is tested by streaming a scalable video over a peer-to-peer network. The experimental results confirm that the proposed algorithm achieves better video quality, better delivery ratio, and lower transmission redundancy. Shenglan Huang, Ebroul Izquierdo, Pengwei Hao |
ICC | 2 |
| 2016 | Dynamic tree-structured sparse RPCA via column subset selection for background modeling and foreground detectionabstractVideo analysis often begins with background subtraction, which consists of creation of a background model that allows distinguishing foreground pixels. Recent evaluation of background subtraction techniques demonstrated that there are still considerable challenges facing these methods. Processing per-pixel basis from the background is not only time-consuming but also can dramatically affect foreground region detection, if region cohesion and contiguity is not considered in the model. We present a new method in which we regard the image sequence to be made up of the sum of a low-rank background matrix and a dynamic tree-structured sparse matrix, and solve the decomposition using our approximated Robust Principal Component Analysis method extended to handle camera motion. Furthermore, to reduce the curse of dimensionality and scale, we introduce a low-rank background modeling via Column Subset Selection that reduces the order of complexity, decreases computation time, and eliminates the huge storage need for large videos. Salehe Erfanian Ebadi, Valia Guerra-Ones, Ebroul Izquierdo |
ICIP | 3 |
| 2016 | Steadiness analysis for optimal GOP size selection in HEVCabstractHigh Efficiency Video Coding is the latest and most advanced video coding standard. It supports various Group Of Pictures (GOP) sizes and types such as low delay and random access. The size of the GOP substantially influences the temporal coding process. Therefore, a suitable GOP selection strategy can have a significant impact in the compression efficiency. In this paper, a strategy for GOP selection is proposed. It is derived from a new, low complexity measure of the temporal “steadiness” in the video content. Contrasting conventional approaches the proposed technique does not relies on previous estimation of motion vectors or motion information in the video sequence. As such, it can be used for automated encoding parameter optimization at the start of the coding process. The proposed technique leads to improved encoding efficiency at a negligible computational cost, when compared with the standard coding settings. A comprehensive experimental evaluation confirm that on average -6.69%, BD-rate gain and 14.18% time savings can be achieved. Vigneswaran Poobalasingam, Ebroul Izquierdo |
ICIP | 2 |
| 2016 | Formal representation of events in a surveillance domain ontologyabstractFollowing the exponential deployment of surveillance systems across a wide-spread region of geographic locations, detection and representation of events has become a critical element in automated surveillance systems. In this paper, we present an extensive ontology framework for representing complex semantic events. The proposed ontology builds on DOLCE ontology and relies on the linguistic and cognitive modelling of philosophical knowledge to achieve interoperability between proprietary surveillance systems. The explicit definition of event vocabulary presented in the paper is aimed at aiding forensic analysts to objectively identify and represent complex events. The expressiveness of the proposed ontology framework is described in the context of London Riots which took place in 2011. Faranak Sobhani, Krishna Chandramouli, Qianni Zhang, Ebroul Izquierdo |
ICIP | 4 |
| 2016 | HEVC encoder optimisations using adaptive coding unit visiting orderabstractThe flexible partitioning scheme and increased number of prediction modes in the High Efficiency Video Coding (HEVC) standard are largely responsible for both its high compression efficiency and computational complexity. Each frame in HEVC is partitioned in Coding Tree Units (CTUs) of fixed size, which are then recursively partitioned in Coding Units (CUs). In typical implementations, CUs in a CTU are visited from top to bottom at each level of recursion. In this paper, a different approach is used in which CUs in a CTU can be adaptively visited also in reverse order from bottom to top. Three novel algorithms to reduce complexity of HEVC depth selection, mode decision and inter-prediction are presented, based on this adaptive visiting order. Experimental results show that the proposed encoder achieves on average 38.2% speed-ups compared to fast reference HEVC implementations with pre-built speed-ups enabled, for very limited efficiency losses. Ivan Zupancic, Saverio G. Blasi, Eduardo Peixoto, Ebroul Izquierdo |
ICIP | 4 |
| 2016 | Optimised selection of structure of pictures for video codingabstractEncoders based on the High Efficiency Video Coding (HEVC) standard consider an input sequence as a succession of slices grouped in Structures of Pictures (SOP). The SOP used while encoding specifies many parameters, such as the coding order of frames, or the reference frames used during inter-prediction. Reference encoders typically make use of a fixed SOP structure of a given size, which is periodically repeated throughout the whole sequence. In this paper, the usage of unconventional SOP structures is first analysed, showing that most sequences benefit from usage of larger SOPs, and that the selection of the optimal SOP is highly content dependent. As a result, an algorithm is proposed to automatically select the optimal SOP size based on a low-complexity texture analysis of neighbouring frames. The algorithm is capable of adaptively changing SOP size during the encoding. Extensive evaluation shows that consistent bit-rate reductions are reported at the same objective quality as an effect of using the proposed algorithm. Vigneswaran Poobalasingam, Ebroul Izquierdo, Saverio G. Blasi, Marta Mrak |
MMSP | 2 |
| 2016 | Improved combined intra prediction for higher video compression efficiencyabstractIntra prediction is an important component of video compression. The High Efficiency Video Coding (HEVC) standard supports advanced Intra prediction tools to achieve remarkable Intra coding performance. However, higher compression efficiency performance may be achieved using Combined Intra Prediction (CIP). CIP consists in combining conventional Intra reference samples, with samples extracted from within the current block being encoded. In this paper, an improvement to the original CIP approach is proposed to achieve higher compression efficiency. Due to the fact the encoder does not have access to reconstruction samples while compressing a block, a controlled drift occurs at the block level between the Intra prediction blocks generated at the encoder and decoder when using CIP. The proposed approach aims at reducing such a drift, hence improving CIP performance. This Improved CIP is able to achieve up to 1.4 % BD-rate savings with respect to the HEVC reference software. The paper also proposes the combination of Improved CIP with an additional tool for enhancing the accuracy of Intra prediction. Up to 1.9 % BD-rate savings can be achieved using this approach. André Seixas Dias, Saverio G. Blasi, Marta Mrak, Ebroul Izquierdo |
PCS | 4 |
| 2016 | Fast motion estimation based on neighbouring cost similarityabstractMany video coding standards, including the state-of-the-art High Efficiency Video Coding (HEVC) standard, employ Motion Estimation (ME) at sub-pel precision level to improve the prediction accuracy. However, interpolation and additional Motion Vector (MV) search associated with sub-pel ME result in increased computational complexity. In this paper, an algorithm for fast ME is proposed based on neighbouring sample cost similarity. The costs corresponding to the neighbouring MVs around the optimal MV at different precision levels are analyzed and used to interpolate an error surface. The estimated location of the surface minimum is then used to either skip the remaining ME process, or to limit the number of points to test. Experimental evaluation on high resolution sequences shows that the proposed method can speed up the ME module by 55.5% for 1.5% BD rate losses. Furthermore, the proposed fast quarter-pel ME reduces the quarter-pel ME time by 79.6% while achieving only 17.6% of theoretical encoding losses. Ivan Zupancic, Ebroul Izquierdo |
PCS | 2 |
| 2016 | Two-pass rate control for UHDTV delivery with HEVCabstractRate control has been regarded as an indispensable video coding tool for virtually any application involving video transmission. With the advent of many flexible tools introduced in the current state-of-the-art High Efficiency Video Coding (HEVC) standard, previous Rate-Distortion (RD) models used for rate control become insufficiently accurate. To overcome this issue, a new RD model has been recently proposed based on a robust correspondence between the rate and Lagrange multiplier λ. However, existing methods based on this model tend to perform sub-optimally after the scene change. In this paper, a two-pass rate control method is proposed, targeting Ultra High Definition Television (UHDTV) applications. In the first pass, a fast encoder with a reduced set of coding tools is used to obtain the data used for rate allocation and model parameter initialisation utilised during the second pass. Multiple encoding steps required to derive this information are avoided with the proposed variable quantization parameter framework. Experimental evaluation showed that the proposed two-pass rate control method achieves on average 4.4% BD-rate loss, compared with variable bit-rate encoding. That significantly outperforms the state-of-the-art HEVC rate control method with an average BD-rate loss of 8.8%. Ivan Zupancic, Ebroul Izquierdo, Matteo Naccari, Marta Mrak |
PCS | 2 |
| 2016 | Multi-generation packet scheduling for live streaming with network codingabstractOver the last decade, the emergence of new multimedia devices has motivated the research on efficient media streaming mechanisms that can adapt to dynamic network conditions and heterogeneous devices' capabilities. Network coding (NC) as a rateless code has been applied to collaborative media streaming applications and brings substantial improvements regarding throughput and delay in collaborative media streaming applications. However, little attention has been given to the recoverability of encoded data, especially for the streaming with a strict deadline. When a receiver obtains insufficient packets, it is impossible for receivers to recover any original packets. This in turn leads to severe quality of experience. In this paper, we solve the unrecoverable transmission by determining a scalable layer subscription and scheduling strategy. This scalable layer subscription algorithm is treated as a video quality maximization problem in multiple generations. Then this problem is solved using a dynamic programming algorithm. Experimental results confirm that the proposed algorithm brings better data recoverability and a better quality of service in terms of better video quality, delivery ratio, lower redundancy rate, and significantly lower superfluous video packet rate under different network sizes than conventional random-push schemes. Shenglan Huang, Ebroul Izquierdo, Pengwei Hao |
VCIP | 2 |
| 2016 | Decomposition and matching: Towards efficient automatic Chinese character stroke extractionabstractIn this paper, given images of Chinese characters, we present an automatic stroke extraction system, which consists of character decomposition, shape matching, cross area extraction and stroke segment combination. First, a character is decomposed into isolated stroke structures according to the connection. Then, we extract shape contexts of stroke structures and find the matched counterparts in a standard database by shape matching. Cross points are computed, and cross areas are extracted according to a proposed adaptive cross area extraction method based on point-to-boundary orientation distance. For a shape structure with the matched structure, we optimize cross point set and combine its stroke segments according to the correct cross point set and combination way of its matched counterpart. For those without matched structures, we propose an angle based stroke segment combination method to combine the segments into a complete stroke. Experimental results indicate that the proposed system achieves high accuracy and demonstrates prominent augmentation on efficiency. Zongyi Xu, Qianni Zhang, Ebroul Izquierdo |
VCIP | 5 |
| 2016 | Rethinking random Hough Forests for video database indexing and pattern searchabstractHough Forests have demonstrated effective performance in object detection tasks, which has potential to translate to exciting opportunities in pattern search. However, current systems are incompatible with the scalability and performance requirements of an interactive visual search. In this paper, we pursue this potential by rethinking the method of Hough Forests training to devise a system that is synonymous with a database search index that can yield pattern search results in near real time. The system performs well on simple pattern detection, demonstrating the concept is sound. However, detection of patterns in complex and crowded street-scenes is more challenging. Some success is demonstrated in such videos, and we describe future work that will address some of the key questions arising from our work to date. Craig Henderson, Ebroul Izquierdo |
Comput. Vis. Media | 2 |
| 2016 | Symmetric stability of low level feature detectors
Craig Henderson, Ebroul Izquierdo |
Pattern Recognit. Lett. | 2 |
| 2016 | A Hierarchical Distributed Processing Framework for Big Image DataabstractThis paper introduces an effective processing framework nominated Image Cloud Processing (ICP) to powerfully cope with the data explosion in image processing field. While most previous researches focus on optimizing the image processing algorithms to gain higher efficiency, our work dedicates to providing a general framework for those image processing algorithms, which can be implemented in parallel so as to achieve a boost in time efficiency without compromising the results performance along with the increasing image scale. The proposed ICP framework consists of two mechanisms, i.e., Static ICP (SICP) and Dynamic ICP (DICP). Specifically, SICP is aimed at processing the big image data pre-stored in the distributed system, while DICP is proposed for dynamic input. To accomplish SICP, two novel data representations named P-Image and Big-Image are designed to cooperate with MapReduce to achieve more optimized configuration and higher efficiency. DICP is implemented through a parallel processing procedure working with the traditional processing mechanism of the distributed system. Representative results of comprehensive experiments on the challenging ImageNet dataset are selected to validate the capacity of our proposed ICP framework over the traditional state-of-the-art methods, both in time efficiency and quality of results. Xiaochun Cao, Ebroul Izquierdo |
IEEE Trans. Big Data | 8 |
| 2016 | Robust Feature Matching in Long-Running Poor-Quality VideosabstractWe describe a methodology that is designed to match key point and region-based features in real-world images, acquired from long-running security cameras with no control over the environment. We detect frame duplication and images from static scenes that have no activity to prevent processing saliently identical images, and describe a novel blur-sensitive feature detection method, a combinatorial feature descriptor, and a distance calculation that efficiently unites texture and color attributes to discriminate feature correspondence in low-quality images. Our methods are tested by performing key point matching on real-world security images such as outdoor closed-circuit television videos that are low quality and acquired in uncontrolled conditions with visual distortions caused by weather, crowded scenes, emergency lighting, or the high angle of the camera mounting. We demonstrate an improvement in accuracy of matching key points between images compared with state-of-the-art feature descriptors. We use key point features from a Harris corner detector, scale-invariant feature transform, speeded-up robust features, binary robust invariant scalable keypoints, and features from accelerated segment test as well as MSER and MSCR region detectors to provide a comprehensive analysis of our generic method. We demonstrate feature matching using a 138D descriptor that improves the matching performance of a state-of-the-art 384D color descriptor with just 36% of the storage requirements. Craig Henderson, Ebroul Izquierdo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Holons Visual Representation for Image RetrievalabstractAlong with the enlargement of image scale, convolutional local features, such as SIFT, are ineffective for representing or indexing and more compact visual representations are required. Due to the intrinsic mechanism, the state-of-the-art vector of locally aggregated descriptors (VLAD) has a few limits. Based on this, we propose a new descriptor named holons visual representation (HVR). The proposed HVR is a derivative mutational self-contained combination of global and local information. It exploits both global characteristics and the statistic information of local descriptors in the image dataset. It also takes advantages of local features of each image and computes their distribution with respect to the entire local descriptor space. Accordingly, the HVR is computed by a two-layer hierarchical scheme, which splits the local feature space and obtains raw partitions, as well as the corresponding refined partitions. Then, according to the distances from the centroids of partition spaces to local features and their spatial correlation, we assign the local features into their nearest raw partitions and refined partitions to obtain the global description of an image. Compared with VLAD, HVR holds critical structure information and enhances the discriminative power of individual representation with a small amount of computation cost, while using the same memory overhead. Extensive experiments on several benchmark datasets demonstrate that the proposed HVR outperforms conventional approaches in terms of scalability as well as retrieval accuracy for images with similar intra local information. Gaipeng Kong, Qianni Zhang, Xiaochun Cao, Ebroul Izquierdo |
IEEE Trans. Multim. | 6 |
| 2016 | Bandwidth-Efficient Packet Scheduling for Live Streaming With Network CodingabstractNetwork coding (NC) brings substantial improvements in terms of throughput and delay in collaborative media streaming applications. A key aspect of NC-driven live peer-to-peer streaming is the packet scheduling policy. Indeed, lack of synchronization among peers usually results in significantly redundant packet transmission, which in turn leads to severe bandwidth inefficiencies. In this paper, we address the problem of finding a suitable asynchronous packet scheduling policy that greatly helps to overcome this critical redundant transmission problem. We propose a bandwidth cost minimization technique under a full video packet recovery constraint. In order to add scalability and improved performance, we also further derive a distributed packet scheduling algorithm. Both implementation and analytical considerations of the proposed approaches are described in this paper. Experimental results confirm that the proposed algorithms deliver higher bandwidth efficiency with reduced redundancy and communication overhead rate and, consequently, better quality-of-service in terms of improved video quality and delivery ratio. Shenglan Huang, Ebroul Izquierdo, Pengwei Hao |
IEEE Trans. Multim. | 2 |
| 2016 | Content-Based Guided Image Filtering, Weighted Semi-Global Optimization, and Efficient Disparity Refinement for Fast and Accurate Disparity EstimationabstractThis paper presents a novel approach, which relies on content-based guided image filtering and weighted semi-global optimization for fast and accurate disparity estimation. The approach uses a pixel-based cost term that combines gradient, Gabor-Feature, and color information. The pixel-based matching costs are filtered by applying guided image filtering, which relies on rectangular support windows of two different sizes. In this way, two filtered costs are estimated for each pixel. Among the two filtered costs, the one that will be finally assigned to each pixel depends on the local image content around this pixel. The filtered cost volume is further refined by exploiting weighted semi-global optimization, which improves the disparity estimation accuracy. Finally, the disparity refinement in outlier regions relies on a straightforward and time-efficient outliers handling scheme and on a simple approach which deals with the disparity outliers at depth discontinuities. Experimental results on the Middlebury online stereo evaluation benchmark and 27 additional Middlebury stereo pairs prove that our method is able to generate disparity maps with high accuracy while keeping the computational cost low. Georgios Kordelas, Dimitrios S. Alexiadis, Petros Daras, Ebroul Izquierdo |
IEEE Trans. Multim. | 4 |
| 2016 | Inter-Prediction Optimizations for Video Coding Using Adaptive Coding Unit Visiting OrderabstractThe flexible partitioning scheme and increased number of prediction modes in the high efficiency video coding (HEVC) standard are largely responsible for both its high compression efficiency and computational complexity. In typical HEVC encoder implementations, coding units (CUs) in a coding tree unit (CTU) are visited from top to bottom at each level of recursion to select the optimal coding configuration. In this paper, a novel approach is presented in which CUs in a CTU can be adaptively visited also ina reverse, bottom to top visiting order. This reverse CU (RCU) visiting order allows for different algorithmic optimizations for further complexity reduction of many HEVC encoding steps, especially under challenging conditions, such as highly textured or fast moving content. In particular, algorithms to reduce complexity of HEVC depth selection, mode decision, and inter-prediction are presented here based on the coding information obtained from higher depths when using the RCU visiting order. Experimental results show that enabling different stages of the proposed algorithm can achieve average speedups from 16.3% to 36.6% compared to fast reference HEVC implementation with pre-built speed-ups enabled (up to 51.2% in some cases), for 0.3% to 2.2% BD-rate penalty. Ivan Zupancic, Saverio G. Blasi, Eduardo Peixoto, Ebroul Izquierdo |
IEEE Trans. Multim. | 4 |
| 2015 | Action Recognition based on Subdivision-Fusion ModelabstractThis paper proposes a novel Subdivision-Fusion Model (SFM) to recognize human actions. In most action recognition tasks, overlapping feature distribution is a common problem leading to overfitting. In the subdivision stage of the proposed SFM, samples in each category are clustered. Then, such samples are grouped into multiple more concentrated subcategories. Boundaries for the subcategories are easier to find and as consequence overfitting is avoided. In the subsequent fusion stage, the multi-subcategories classification results are converted back to the original category recognition problem. Two methods to determine the number of clusters are provided. The proposed model has been thoroughly tested with four popular datasets. In the Hollywood2 dataset, an accuracy of 79.4% is achieved, outperforming the state-of-the-art accuracy of 64.3%. The performance on the YouTube Action dataset has been improved from 75.8% to 82.5%, while considerably improvements are also observed on the KTH and UCF50 datasets. Zong Bo Hao, Linlin Lu, Qianni Zhang, Ebroul Izquierdo, Juanyu Yang |
BMVC | 5 |
| 2015 | Horizontal flip-invariant sketch recognition via local patch hashingabstractThis paper introduces a flip aware patch matching frame-work that facilitates scalable sketch recognition. An overlapping spatial grid is utilized to generate an ensemble of patches for each sketch. We rank similarities between freely drawn sketches via a spatial voting process where similar patches in terms of shape and structure arbitrate for the result. Patch similarity is efficiently estimated via the min-hash algorithm. A novel spatial aware reverse index structure ensures the scalability of our scheme. We show the benefits of horizontal flip invariance and structural information in sketch recognition and demonstrate state-of-the-art results in two challenging sketch datasets. Konstantinos Bozas, Ebroul Izquierdo |
ICASSP | 2 |
| 2015 | Multiple Early Termination for fast HEVC coding of UHD contentabstractThe recently ratified High Efficiency Video Coding (HEVC) standard is significantly outperforming previous video coding standards in terms of compression efficiency. However, this comes at the cost of very high computational complexity, which may limit its real-time usage, particularly when targeting Ultra High Definition (UHD) applications. In this paper, an analysis of HEVC coding on UHD content is presented, showing that on average more than 18% of the total encoding time is spent performing uni-directional Motion Estimation (ME) even when using fast algorithms such as Enhanced Predictive Zonal Search (EPZS). In order to speed up the ME process, a novel approach for fast inter prediction is proposed in this paper based on a Multiple Early Termination (MET) decision process. EPZS is only performed in blocks in which it is needed based on local features of the encoded content, or it is skipped otherwise. Experimental results show that the algorithm achieves on average 9.3% speed-ups over conventional HEVC, at the cost of very small BD-rate losses. Ivan Zupancic, Saverio G. Blasi, Ebroul Izquierdo |
ICASSP | 3 |
| 2015 | Context adaptive mode sorting for fast HEVC mode decisionabstractTypical H.265/High Efficiency Video Coding (HEVC) encoder implementations test a variety of prediction modes and select the optimal configuration for each block in terms of Rate-Distortion (RD) cost. A fast HEVC mode decision algorithm is proposed here referred to as Context Adaptive Mode Sorting (CAMS). The frequency of selection of modes and their RD costs are collected while encoding the training frames based on local parameters (the context). This information is then used to sort and restrict the prediction modes to test for each context, while the optimal mode found using CAMS on each CU is validated based on the RD cost distributions found during the training. Experimental results show that the method reduces total encoding time of fast HEVC implementations on average by 29.3%, at modest efficiency losses. Saverio G. Blasi, Eduardo Peixoto, Bruno Macchiavello, Edson M. Hung, Ivan Zupancic, Ebroul Izquierdo |
ICIP | 6 |
| 2015 | Efficient background subtraction with low-rank and sparse matrix decompositionabstractDecomposition of a video scene into background and foreground is an old problem, for which novel approaches in the last years have been proposed. The robust subspace approach based on a low-rank plus sparse matrix decomposition has shown a great ability to identify static parts from moving objects in video sequences. However, those models are still insufficient in realistic environments. In this paper, we propose a modified approximated robust PCA algorithm that can handle moving cameras and takes advantage of the block sparse structure of the pixels corresponding to the moving objects. Additionally, we propose a novel SVD-free algorithm for the case of rank-1 background that outperforms current state-of-the-art methods in computation cost/time as well as performance. Finally, experiments and numerical results evaluating the proposed methods are demonstrated. Salehe Erfanian Ebadi, Valia Guerra-Ones, Ebroul Izquierdo |
ICIP | 3 |
| 2015 | Distributed Bandwidth-efficient Packet Scheduling for Live Streaming with Network CodingabstractThis paper proposed a distributed packet scheduling algorithm for live peer-to-peer streaming system, where network coding is extended to improve the efficiency in bandwidth utilization. A problem of superfluous packet transmission due to the lack of synchronization among peers is identified. This problem often leads to bandwidth inefficiencies. We solve the problem of finding a suitable asynchronous packet scheduling policy by posing and solving a bandwidth allocation problem. This proposed scheduling policy can reduce the superfluous packet transmission, thereby achieving a more efficient bandwidth usage and an improved quality of service in live streaming applications. Experimental results confirm that the proposed scheme demonstrates significantly better video quality, delivery ratio under different network size and different loss rate compared with other push-based schemes. Shenglan Huang, Ebroul Izquierdo, Pengwei Hao |
ACM Multimedia | 2 |
| 2015 | Fast HEVC coding using reverse CU visitingabstractThe High Efficiency Video Coding (HEVC) standard makes use of flexible partitioning to achieve high compression ratios. Each frame is divided in Coding Tree Units (CTUs) of fixed size, which are further partitioned into Coding Units (CUs) following a recursive quadtree structure. Typically, CUs at each level of recursion are tested to select the optimal coding configuration, hence this process is extremely demanding in terms of computational complexity. In this paper, a method to reduce complexity of HEVC quadtree configuration selection and mode decision is presented, based on a reverse bottom-to-top visiting order of CUs in the quadtree. By visiting smallest CUs first, information can be extracted to make decisions on larger CUs. The encoder adaptively selects whether a CTU is encoded using the reverse CU visiting, allowing for considerably faster encoding under all conditions. Experimental results show that the algorithm achieves on average 21% speedups over previous state-of-the-art fast HEVC algorithms, and up to 36% for some sequences, at very limited efficiency losses. Saverio G. Blasi, Ivan Zupancic, Ebroul Izquierdo, Eduardo Peixoto |
PCS | 3 |
| 2015 | Adaptive precision motion estimation for HEVC codingabstractMost video coding standards, including the state-of-the-art High Efficiency Video Coding (HEVC), make use of sub-pixel Motion Estimation (ME) with Motion Vectors (MV) at fractional precisions to achieve high compression ratios. Unfortunately, sub-pixel ME comes at very high computational costs due to the interpolation step and additional motion searches. In this paper, a fast sub-pixel ME algorithm is proposed. The MV precision is adaptively selected on each block to skip the half or quarter precision steps when not needed. The algorithm bases the decision on local features, such as the behaviour of the residual error samples, and global features, such as the amount of edges in the pictures. Experimental results show that the method reduces total encoding time by up to 17.6% compared to conventional HEVC, at modest efficiency losses. Saverio G. Blasi, Ivan Zupancic, Ebroul Izquierdo, Eduardo Peixoto |
PCS | 3 |
| 2015 | HEVC coding optimisation for Ultra High Definition television servicesabstractUltra High Definition TV (UHDTV) services are being trialled while UHD streaming services have already seen commercial débuts. The amount of data associated with these new services is very high thus extremely efficient video compression tools are required for delivery to the end user. The recently published High Efficiency Video Coding (HEVC) standard promises a new level of compression efficiency, up to 50% better than its predecessor, Advanced Video Coding (AVC). The greater efficiency in HEVC is obtained at much greater computational cost compared to AVC. A practical encoder must optimise the choice of coding tools and devise strategies to reduce the complexity without affecting the compression efficiency. This paper describes the results of a study aimed at optimising HEVC encoding for UHDTV content. The study first reviews the available HEVC coding tools to identify the best configuration before developing three new algorithms to further reduce the computational cost. The proposed optimisations can provide an additional 11.5% encoder speed-up for an average 3.1% bitrate increase on top of the best encoder configuration. Matteo Naccari, Andrea Gabriellini, Marta Mrak, Saverio G. Blasi, Ivan Zupancic, Ebroul Izquierdo |
PCS | 6 |
| 2015 | Navigation in REVERIE's virtual environmentsabstractThis work presents a novel navigation system for social collaborative virtual environments populated with multiple characters. The navigation system ensures collision free movement of avatars and agents. It supports direct user manipulation, automated path planning, positioning to get seated, and follow-me behaviour for groups. In follow-me mode, the socially aware system manages the mise en place of individuals within a group. A use case centred around on an educational virtual trip to the European Parliament created for the REVERIE FP7 project, also serves as an example to bring forward aspects of such navigational requirements. Fiona M. Rivera, Fons Fuijk, Ebroul Izquierdo |
VR | 3 |
| 2015 | Enhanced disparity estimation in stereo images
Georgios Kordelas, Dimitrios S. Alexiadis, Petros Daras, Ebroul Izquierdo |
Image Vis. Comput. | 4 |
| 2015 | Frequency-Domain Intra Prediction Analysis and Processing for High-Quality Video CodingabstractMost of the advances in video coding technology focus on applications that require low bitrates, for example, for content distribution on a mass scale. For these applications, the performance of conventional coding methods is typically sufficient. Such schemes inevitably introduce large losses to the signal, which are unacceptable for numerous other professional applications such as capture, production, and archiving. To boost the performance of video codecs for high-quality content, better techniques are needed especially in the context of the prediction module. An analysis of conventional intra prediction methods used in the state-of-the-art High Efficiency Video Coding (HEVC) standard is reported in this paper, in terms of the prediction performance of such methods in the frequency domain. Appropriately modified encoder and decoder schemes are presented and used for this paper. The analysis shows that conventional intra prediction methods can be improved, especially for high frequency components of the signal which are typically difficult to predict. A novel approach to improve the efficiency of high-quality video coding is also presented in this paper based on such analysis. The modified encoder scheme allows for an additional stage of processing performed on the transformed prediction to replace selected frequency components of the signal with specifically defined synthetic content. The content is introduced in the signal using feature-dependent lookup tables. The approach is shown to achieve consistent gains against conventional HEVC with up to -5.2% coding gains in terms of bitrate savings. Saverio G. Blasi, Marta Mrak, Ebroul Izquierdo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2014 | Masking of transformed intra-predicted blocks for high quality image and video codingabstractMany professional applications for image and video coding require very high levels of quality of the decoded signal. Under these conditions, even state-of-the-art standards such as High Efficiency Video Coding (HEVC) may not provide sufficient compression efficiency. Large values of the residual samples, especially at high frequency components, are difficult to encode and result in high bitrates of the coded signal. A novel scheme for image and video coding is presented in this paper to enhance compression at high quality levels, based on separate transform to the frequency domain of original and prediction signals. Selected frequency components of the prediction signal are discarded by means of appropriate masking patterns prior to the residual computation. The optimal masking pattern is selected for each transformed block and signalled in the bitstream. The approach is shown achieving gains against conventional HEVC under high quality constraints when coding both still images and video sequences. Saverio G. Blasi, Marta Mrak, Ebroul Izquierdo |
ICIP | 3 |
| 2014 | Temporal face embedding and propagation in photo collectionsabstractWe present a two-step approach for modeling facial variations and class likelihoods over time. Unlike traditional approaches, we explicitly model the temporal domain that is often available, for example, in consumer photos or surveillance systems. Our combined approach draws upon the concepts of manifold transformation and semi-supervised graph-based propagation to simultaneously recognize faces across entire photo collections. Experiments on two datasets demonstrate improved face recognition accuracy. Markus Brenner, Tamar Avraham, Michael Lindenbaum, Ebroul Izquierdo |
ICIP | 4 |
| 2014 | Optimized scalable video transmission over P2P network with hierarchical network codingabstractThis paper proposes a new push-based peer-to-peer (P2P) communication method to optimally transmit scalable video packets over lossy networks. In our scheme, a rate allocation optimization is performed at the sender node. Different from previous sender/receiver driven schemes, our scheme does not need to update buffer maps or request packets periodically, which reduces the amount of redundant packets and yields to less traffic of video data over the Internet. The proposed optimized rate allocation algorithm calculates in advance the number of needed packets according to the cumulative uplink rate between senders and the receiver, and the loss rate of the link. After calculation, each sender sends the hierarchical network coded packets to receivers based on the estimated bandwidth and resource share allocation. With this method, the waste of bandwidth, and the delay caused by the communications among peers can both be reduced. Shenglan Huang, Michele Sanna, Ebroul Izquierdo, Pengwei Hao |
ICIP | 3 |
| 2014 | Revisiting guided image filter based stereo matching and scanline optimization for improved disparity estimationabstractIn this paper the scanline optimization used for stereo matching, is revisited. In order to improve the performance of this semi-global technique, a new criterion to check depth discontinuity, is introduced. This criterion is defined according to the mean-shift-based image segmentation result. Additionally, this work proposes the employment of a pixel dissimilarity metric for the computation of the cost term, which is then provided to the guided image filter approach to estimate the initial cost volume. The algorithm is tested on the four images of the online Middlebury stereo evaluation benchmark. Moreover, it is tested on 27 additional Middlebury stereo pairs for assessing thoroughly its performance. The extended comparison verifies the efficiency of this work. Georgios Kordelas, Dimitrios S. Alexiadis, Petros Daras, Ebroul Izquierdo |
ICIP | 4 |
| 2014 | REVERIE: Natural human interaction in virtual immersive environmentsabstractREVERIE (REal and Virtual Engagement in Realistic Immersive Environments [1]) targets novel research to address the demanding challenges involved with developing state-of-the-art technologies for online human interaction. The REVERIE framework enables users to meet, socialise and share experiences online by integrating cutting-edge technologies for 3D data acquisition and processing, networking, autonomy and real-time rendering. In this paper, we describe the innovative research that is showcased through the REVERIE integrated framework through richly defined use-cases which demonstrate the validity and potential for natural interaction in a virtual immersive and safe environment. Previews of the REVERIE demo and its key research components can be viewed at www.youtube.com/user/REVERIEFP7. Julie A. Wall, Ebroul Izquierdo, Lemonia Argyriou, David S. Monaghan, Noel E. O'Connor, Steven Poulakos, Aljoscha Smolic, Rufael Mekuria |
ICIP | 2 |
| 2014 | Joint People Recognition across Photo Collections Using Sparse Markov Random Fields
Markus Brenner, Ebroul Izquierdo |
MMM (1) | 2 |
| 2014 | Tools for User Interaction in Immersive Environments
Noel E. O'Connor, Dimitrios S. Alexiadis, Konstantinos C. Apostolakis, Petros Daras, Ebroul Izquierdo, Yingbo Li, David S. Monaghan, Fiona M. Rivera, C. Stevens, Sigurd Van Broeck, Julie A. Wall, Haolin Wei |
MMM (2) | 5 |
| 2014 | Accurate stereo 3D point cloud generation suitable for multi-view stereo reconstructionabstractThis paper proposes a novel methodology for generating 3D point clouds of good accuracy from stereo pairs. Initially, the methodology defines some conditions for the proper selection of image pairs. Then, the selected stereo images are used to estimate dense correspondences using the Daisy descriptor. An efficient two-phase strategy to remove outliers is then introduced. Finally, the 3D point cloud is refined by combining sub-pixel accuracy correspondences estimation and the moving least squares algorithm. The proposed methodology can be exploited by multi-view stereo algorithms due to its good accuracy and its fast computation. Georgios Kordelas, Petros Daras, Patrycia Klavdianos, Ebroul Izquierdo, Qianni Zhang |
VCIP | 4 |
| 2014 | Fast intra mode decision for HEVC based on texture characteristic from RMD and MPMabstractIn this paper, we proposed a fast intra mode decision algorithm for HEVC to further reduce the candidate modes of RDOQ or even skip RDOQ for a PU, which exploited not only the texture consistency of neighbouring PUs reflected by the relation between the optimal RMD mode and MPM, but also the texture characteristic in a PU reflected by the best two RMD modes. Experimental results show that our proposed algorithm can reduce 31.3% intra coding time with 0.56% BD-rate loss on average. Youwei Chen, Ebroul Izquierdo |
VCIP | 3 |
| 2014 | H.264/AVC to HEVC Video Transcoder Based on Dynamic Thresholding and Content ModelingabstractThe new video coding standard, High Efficiency Video Coding (HEVC), was developed to succeed the current standard, H.264/AVC, as the state of the art in video compression. However, there is a lot of legacy content encoded with H.264/AVC. This paper proposes and evaluates several transcoding algorithms from the H.264/AVC to the HEVC format. In particular, a novel transcoding architecture, in which the first frames of the sequence are used to compute the parameters so that the transcoder can learn the mapping for that particular sequence, is proposed. Then, two types of mode mapping algorithms are proposed. In the first solution, a single H.264/AVC coding parameter is used to determine the outgoing HEVC partitions using dynamic thresholding. The second solution uses linear discriminant functions to map the incoming H.264/AVC coding parameters to the outgoing HEVC partitions. This paper contains experiments designed to study the impact of the number of frames used for training in the transcoder. Comparisons with existing transcoding solutions reveal that the proposed work results in lower rate-distortion loss at a competitive complexity performance. Eduardo Peixoto, Tamer Shanableh, Ebroul Izquierdo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2014 | Enabling Geometry-Based 3-D Tele-Immersion With Fast Mesh Compression and Linear Rateless Codingabstract3-D tele-immersion (3DTI) enables participants in remote locations to share, in real time, an activity. It offers users interactive and immersive experiences, but it challenges current media-streaming solutions. Work in the past has mainly focused on the efficient delivery of image-based 3-D videos and on realistic rendering and reconstruction of geometry-based 3-D objects. The contribution of this paper is a real-time streaming component for 3DTI with dynamic reconstructed geometry. This component includes both a novel fast compression method and a rateless packet protection scheme specifically designed towards the requirements imposed by real time transmission of live-reconstructed mesh geometry. Tests on a large dataset show an encoding speed-up up to ten times at comparable compression ratio and quality, when compared with the high-end MPEG-4 SC3DMC mesh encoders. The implemented rateless code ensures complete packet loss protection of the triangle mesh object and a delivery delay within interactive bounds. Contrary to most linear fountain codes, the designed codec enables real-time progressive decoding allowing partial decoding each time a packet is received. This approach is compared with transmission over TCP in packet loss rates and latencies, typical in managed WAN and MAN networks, and heavily outperforms it in terms of end-to-end delay. The streaming component has been integrated into a larger 3DTI environment that includes state of the art 3-D reconstruction and rendering modules. This resulted in a prototype that can capture, compress transmit, and render triangle mesh geometry in real-time in realistic internet conditions as shown in experiments. Compared with alternative methods, lower interactive end-to-end delay and frame rates over three times higher are achieved. Rufael Mekuria, Michele Sanna, Ebroul Izquierdo, Dick C. A. Bulterman, Pablo César |
IEEE Trans. Multim. | 3 |
| 2013 | Knowledge modeling for privacy-by-design in smart surveillance solutionabstractAs new information and communications systems are being equipped with more aggressive capabilities to enable smart surveillance, individuals' private and ethical data is more exposed to potential threats. Consequently, the attention of researchers and policy makers has become increasingly focused on controlling the emerging threats to privacy. In order to ensure that a surveillance system framework complies with the legal, ethical and privacy requirements of the law, in this paper we present a Surveillance Ontology extending the SKOS foundational ontology. The fundamental principles of privacy-by-design (PbD) demand that the surveillance framework consider data minimization, user control, accountability and data separation. Hence, the objective of this ontology is to translate the high-level linguistic rules into the information that can be processed and used to assess the compliance of the video analysis module with the rules defined. Krishna Chandramouli, Virginia Fernandez Arguedas, Ebroul Izquierdo |
AVSS | 3 |
| 2013 | Enhanced inter-prediction using Merge Prediction Transformation in the HEVC codecabstractMerge prediction is a novel technique introduced in the HEVC standard to improve inter-prediction exploiting redundancy of themotion information. We propose in this paper a new approach to enhance the Merge mode in a typical HEVC encoder using parametric transformations of the Merge prediction candidates. An Enhanced Inter-Prediction module is implemented in HEVC using Merge Prediction Transformation (MPT), integrated with the HEVC new features such as the large coding units (CU) and the recursive prediction unit partitioning. The MPT parameters are quantised according to the CU depth and the current QP. The optimal quantization steps are derived via statistical analysis as illustrated in the paper. Results show consistent improvements over conventional HEVC encoding in terms of rate-distortion performance, with a small impact on the encoding complexity and negligible impact on the decoding complexity. Saverio G. Blasi, Eduardo Peixoto, Ebroul Izquierdo |
ICASSP | 3 |
| 2013 | Mode decision with enhanced inter-prediction in HEVCabstractThe HEVC standard makes use of inter-prediction to exploit temporal redundancy in order to obtain efficient compression. A novel approach is proposed in this paper where the inter-prediction for a coding unit (CU) is further enhanced by means of simple parametric transformations. The Enhanced Inter-Prediction (EIP) module is embedded within the mode-decision module at a CU level, to obtain a more accurate inter-prediction and hence possibly reducing the residuals before transform and quantisation. The EIP parameters are encoded in the CU header, and the exact rate-distortion (RD) cost after reconstruction is computed to make sure that the EIP is only used when it is effective. The approach is also improved by means of efficient quantisation of the EIP parameters. The optimal quantisation steps for each CU are found following from analytical considerations as illustrated in the paper. Results show consistent improvements over conventional HEVC encoding in terms of PSNR and bitrate performance. Saverio G. Blasi, Eduardo Peixoto, Ebroul Izquierdo |
ICIP | 3 |
| 2013 | An H.264/AVC to HEVC video transcoder based on mode mappingabstractThe emerging video coding standard, HEVC, was developed to replace the current standard, H.264/AVC. However, in order to promote inter-operability with existing systems using the H.264/AVC, transcoding from H.264/AVC to the HEVC codec is highly needed. This paper presents a transcoding solution that uses machine learning techniques in order to map H.264/AVC macroblocks into HEVC coding units (CUs). Two alternatives to build the machine learning model are evaluated. The first uses a static training, where the model is built offline and used to transcode any video sequence. The other uses a dynamic training, with two well-defined stages: a training stage and a transcoding stage. In the training stage, full re-encoding is performed while the H.264/AVC and the HEVC information are gathered. This information is then used to build a model, which is used in the transcoding stage to classify the HEVC CU partitioning. Both solutions are tested with well-known video sequences and evaluated in terms of rate-distortion (RD) and complexity. The proposed method is on average 2.26 times faster than the trivial transcoder using fast motion estimation, while yielding a RD loss of only 3.6% in terms of bitrate. Eduardo Peixoto, Bruno Macchiavello, Edson M. Hung, Alexandre Zaghetto, Tamer Shanableh, Ebroul Izquierdo |
ICIP | 6 |
| 2013 | People recognition in ambiguously labeled Photo CollectionsabstractWe show how to recognize people based on their faces in Consumer Photo Collections while also incorporating context in the form of ambiguous labels. Such labels can be assigned to single photos (depicting multiple people) as well as to entire sets of photos (e.g. relating to events). To achieve this, we devise a unified framework that has a graphical model along a distance-based face description method at its core. We evaluate our probabilistic approach by performing experiments on two datasets, one of which includes around 5000 face appearances spanning nearly ten years. Markus Brenner, Ebroul Izquierdo |
ICME | 2 |
| 2013 | Human action recognition by fast dense trajectoriesabstractIn this paper, we propose the fast dense trajectories algorithm for human action recognition. Dense trajectories are robust to fast irregular motions and outperform other state-of-the-art descriptors such as KLT tracker or SIFT descriptors. However, the use of dense trajectories is time consuming. To improve the efficiency, we extract feature trajectories in the ROI rather than in the whole frames, and we use the temporal pyramids to achieve adaptable mechanism for different action speed. We evaluate the method on the dataset of Huawei/3DLife -- 3D human reconstruction and action recognition Grand Challenge in ACM Multimedia 2013. Experimental results show a significant improvement over the dense trajectories descriptor in real-time, and adaptable to different speed. Zong Bo Hao, Qianni Zhang, Ebroul Izquierdo, Nan Sang |
ACM Multimedia | 3 |
| 2013 | Mining People's Appearances to Improve Recognition in Photo Collections
Markus Brenner, Ebroul Izquierdo |
MMM (1) | 2 |
| 2013 | Gender-aided people recognition in photo collectionsabstractWe show how to recognize people based on their faces in Consumer Photo Collections while also incorporating context in the form of gender information. We devise and explore a unified framework that has a graphical model along a distance-based face description method at its core. We jointly recognize people across an entire photo collection to also consider the specifics of photos that depict multiple people. Experiments on two datasets demonstrate and validate the effectiveness of our probabilistic approach compared to traditional methods that do not consider gender information. Markus Brenner, Ebroul Izquierdo |
MMSP | 2 |
| 2013 | A 3D tele-immersion system based on live captured mesh geometryabstract3D Tele-immersion enables participants in remote locations to share, in real-time, an activity. It offers users natural interactivity and immersive experiences, but it challenges current networking solutions. Work in the past has mainly focused on the efficient delivery of image-based 3D videos and on the realistic rendering and reconstruction of geometry-based 3D objects. The contribution of this paper is a complete media pipeline that allows for geometry-based 3D tele-immersion. Unlike previous approaches, that stream videos or video plus depth estimate, our streaming module can transmit the live-reconstructed 3D representations (triangle meshes). Based on a set of comparative experiments, this paper details the architecture and describes a novel component that can efficiently stream geometry in real-time. This component includes both a novel fast local compression algorithm and a rateless packet protection scheme geared towards the requirements imposed by real-time transmission of live-capture mesh geometry. Tests on a large dataset show an encoding and decoding speed-up of over 10 times at similar compression and quality rates, when compared to the high-end MPEG-4 SC3DMC mesh encoder. The implemented rateless code ensures complete packet loss protection of the triangle mesh object and avoids delay introduced by retransmissions. This approach is compared to a streaming mechanism over TCP and outperforms it at packet loss rates over 2% and/or latencies over 9 ms in terms of end-to-end transmission delay. As reported in this paper, the component has been successfully integrated into a larger tele-immersive environment that includes beyond state of the art 3D reconstruction and rendering modules. This resulted in a prototype that can capture, compress transmit and render triangle mesh geometry in real-time over the internet. Rufael Mekuria, Michele Sanna, Stefano Asioli, Ebroul Izquierdo, Dick C. A. Bulterman, Pablo César |
MMSys | 4 |
| 2013 | Improving inter prediction in HEVC with residual DPCM for lossless screen content codingabstractVideo content containing computer generated objects is usually denoted as screen content and is becoming popular in applications such as desktop sharing, wireless displays, etc. Screen content images and videos are characterized by high frequency details such as sharp edges and high contrast image areas. On these areas classical lossy encoding tools - spatial transform plus quantization - may significantly compromise their quality and intelligibility. Therefore, lossless coding is used instead and improved coding tools should be specifically devised for screen content. In this context this paper proposes a residual differential pulse code modulation (RDPCM) applied to inter predicted residuals and tested in the context of the HEVC range extension development. The proposed method exploits the spatial correlation present in blocks containing edges or text areas which are poorly predicted by motion compensation. In addition to the baseline inter RDCPM, two improvements to the compression efficiency and the overall throughput are presented and assessed. When compared to HEVC lossless coding as specified in Version 1 of the standard, the proposed algorithm achieves up to 8% average bitrate reduction while not increasing the overall decoding complexity. Matteo Naccari, Saverio G. Blasi, Marta Mrak, Ebroul Izquierdo |
PCS | 4 |
| 2013 | Enhanced Inter-Prediction Via Shifting Transformation in the H.264/AVCabstractInter-prediction based on block-based motion estimation (ME) is used in most video codecs. The closer the prediction to the target block, the lower the residual, and thus more efficient compression can be achieved. In this paper, a new technique called enhanced inter-prediction (EIP) is proposed to improve the prediction candidates using an additional transformation acting while performing ME. A parametric transformation acts within the coding loop of each block to modify the prediction for each motion vector candidate. The EIP is validated in the particular case of a single-parameter shifting transformation. This paper presents an efficient algorithm to compute the best shift for each prediction candidate and a model to select the optimal prediction based on minimum cost integrating the approach with existing rate-distortion optimization techniques in the H.264/AVC video codec. Results show significant improvements with an average of 6% bit-rate reduction compared to the original H.264/AVC. Saverio G. Blasi, Eduardo Peixoto, Ebroul Izquierdo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2013 | MPEG-2 to HEVC Video Transcoding With Content-Based ModelingabstractThis paper proposes an efficient MPEG-2 to High Efficiency Video Coding (HEVC) video transcoder. The objective of the transcoder is to migrate the abundant MPEG-2 video content to the emerging HEVC video coding standard. The transcoder introduces a content-based machine learning solution to predict the depth of the HEVC coding units. The proposed transcoder utilizes full re-encoding to find a mapping between the incoming MPEG-2 coding information and the outgoing HEVC depths of the coding units. Once the model is built, a switch to transcoding mode occurs. Hence, the model is content based and varies from one video sequence to another. The transcoder is compared against full re-encoding using the default HEVC fast motion estimation. Using HEVC test sequences, it is shown that a speedup factor of up to 3 is achieved, while reducing the bitrate of the incoming video by around 50%. In comparison to full re-encoding, an average of 3.9% excessive bitrate is encountered with an average PSNR drop of 0.1 dB. Since this is the first work to report on MPEG-2 to HEVC video transcoding, the reported results can be used as a benchmark for future transcoding research. Tamer Shanableh, Eduardo Peixoto, Ebroul Izquierdo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2013 | Histology Image Retrieval in Optimized Multifeature SpacesabstractContent based histology image retrieval systems have shown great potential in supporting decision making in clinical activities, teaching, and biological research. In content based image retrieval, feature combination plays a key role. It aims at enhancing the descriptive power of visual features corresponding to semantically meaningful queries. It is particularly valuable in histology image analysis where intelligent mechanisms are needed for interpreting varying tissue composition and architecture into histological concepts. This paper presents an approach to automatically combine heterogeneous visual features for histology image retrieval. The aim is to obtain the most representative fusion model for a particular keyword that is associated to multiple query images. The core of this approach is a multi-objective learning method, which aims to understand an optimal visual-semantic matching function by jointly considering the different preferences of the group of query images. The task is posed as an optimisation problem, and a multi-objective optimisation strategy is employed in order to handle potential contradictions in the query images associated to the same keyword. Experiments were performed on two different collections of histology images. The results show that it is possible to improve a system for content based histology image retrieval by using an appropriately defined multi-feature fusion model, which takes careful consideration of the structure and distribution of visual features. Qianni Zhang, Ebroul Izquierdo |
IEEE J. Biomed. Health Informatics | 2 |
| 2013 | Multifeature analysis and semantic context learning for image classificationabstractThis article introduces an image classification approach in which the semantic context of images and multiple low-level visual features are jointly exploited. The context consists of a set of semantic terms defining the classes to be associated to unclassified images. Initially, a multiobjective optimization technique is used to define a multifeature fusion model for each semantic class. Then, a Bayesian learning procedure is applied to derive a context model representing relationships among semantic classes. Finally, this context model is used to infer object classes within images. Selected results from a comprehensive experimental evaluation are reported to show the effectiveness of the proposed approaches. Qianni Zhang, Ebroul Izquierdo |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2012 | A game theoretic framework for optimal resource allocation in P2P scalable video streamingabstractIn this paper we describe a game theoretic framework for scalable video streaming over a peer-to-peer network. The proposed system integrates optimal resource allocation functionalities with an incentive provision mechanism for data sharing. First of all, we introduce an algorithm for packet scheduling that allows users to download a specific sub-set of the original scalable bit-stream, depending on the current network conditions. Furthermore, we present an algorithm that aims both at identifying free-riders and minimising transmission delay. Uncooperative peers are cut out of this system, while users upload more data to those which have less to share, in order to fully exploit the resources of all the peers. Experimental evaluation shows that this model can effectively cope with free-riders and minimise transmission delay for scalable video streaming. Stefano Asioli, Naeem Ramzan, Ebroul Izquierdo |
ICASSP | 3 |
| 2012 | Residual error curvature estimation and adaptive classification for selective sub-pel precision motion estimationabstractWe present a novel approach for adaptive precision motion estimation based on a classification of the residual error curvature. A fast algorithm is proposed to estimate the curvature of the interpolated residual surface using the error samples after integer precision motion estimation. We also propose an original technique to compute and successively update a set of thresholds using the information from previously coded frames. The optimal motion vector precision is then selected for each block according to the current thresholds. The approach is compared in terms of PSNR of the motion compensated reconstruction against conventional state of the art sub-pel motion estimation algorithms, and it is shown to efficiently reduce complexity and coding times of a typical video encoder with negligible effects on the prediction accuracy. Saverio G. Blasi, Ebroul Izquierdo |
ICASSP | 2 |
| 2012 | Exploiting social relationships for free-riders detection in minimum-delay P2P scalable video streamingabstractIn this paper we describe a game theoretic framework for scalable video streaming over a peer-to-peer network that exploits social relationships. The proposed system integrates optimal resource allocation functionalities with an incentive provision mechanism for data sharing. First of all, we introduce an algorithm for packet scheduling that allows users to download a specific sub-set of the original scalable bit-stream, depending on the current network conditions. Furthermore, we present an algorithm that aims both at identifying free-riders and minimising transmission delay by exploiting social relationships among peers. Non social/uncooperative peers are cut out of this system, while users upload more data to those which have less to share, in order to fully exploit the resources of all the users. Stefano Asioli, Naeem Ramzan, Ebroul Izquierdo |
ICIP | 3 |
| 2012 | A complexity-scalable transcoder from H.264/AVC to the new HEVC codecabstractThe emerging video coding standard, HEVC, is currently approaching the final stage of development prior to standardization. However, the current H.264/AVC standard is very successful, and it has been widely adopted for many applications. Thus, transcoding between these codecs will be highly needed once the HEVC codec is finalised. This paper studies the performance of one of the most common techniques for heterogeneous transcoding, motion vector (MV) reuse, in a H.264/AVC to HEVC transcoder. Furthermore, it proposes a new transcoder that is capable of complexity scalability, trading off rate-distortion performance for complexity reduction. The proposed transcoder is based on a new metric to compute the similarity of the H.264/AVC MVs, which is used to decide which HEVC partitions are tested on the transcoder. Eduardo Peixoto, Ebroul Izquierdo |
ICIP | 2 |
| 2012 | Stage-based 3D scene reconstruction from single image
Pengwei Hao, Ebroul Izquierdo |
ICPR | 3 |
| 2012 | Social event detection and retrieval in collaborative photo collectionsabstractIn this paper, we present an approach to detect social events and retrieve associated photos in collaboratively annotated photo collections. We combine data of various modalities such as time, location, and textual and visual features within a framework that has a classification model at its core. Compared to traditional approaches that mainly consider the photos only as a source of information, we also incorporate external information from datasets and online web services to further improve the performance. Experiments based on the MediaEval Social Event Detection Dataset demonstrate the effectiveness of our approach. Markus Brenner, Ebroul Izquierdo |
ICMR | 2 |
| 2012 | The privacy challenges of in-depth video analyticsabstractThe increasing need for both automated and privacy-respecting CCTV systems adds many challenges to the tasks of Video Analytics. The growing capabilities of automated surveillance systems lead to the automatic extraction and processing of complex and potentially personal data related to individuals who may not even be aware of it. This paper discusses the issues related to the processing of potentially sensitive information extracted from various computer vision algorithms and techniques such as person tracking, soft biometrics, behaviour analysis or multi-modal person identification, and delivers insights regarding technical solutions and guidelines for achieving a higher level of privacy compliance. Tomas Piatrik, Virginia Fernandez Arguedas, Ebroul Izquierdo |
MMSP | 3 |
| 2012 | A game theoretic approach to minimum-delay scalable video transmission over P2P
Stefano Asioli, Naeem Ramzan, Ebroul Izquierdo |
Signal Process. Image Commun. | 3 |
| 2012 | Special issue on advances in 2D/3D Video Streaming Over P2P Networks
Naeem Ramzan, Ebroul Izquierdo, Hyunggon Park, Aggelos K. Katsaggelos, Johan A. Pouwelse |
Signal Process. Image Commun. | 2 |
| 2012 | Video streaming over P2P networks: Challenges and opportunities
Naeem Ramzan, Hyunggon Park, Ebroul Izquierdo |
Signal Process. Image Commun. | 3 |
| 2012 | Transcoding From Hybrid Nonscalable to Wavelet-Based Scalable Video CodecsabstractScalable video coding (SVC) enables low complexity adaptation of compressed video, providing an efficient solution for content delivery through heterogeneous networks and to diverse displays. However, legacy video and most commercially available content capturing devices use conventional nonscalable coding, e.g., H.264/AVC. This paper proposes an efficient transcoder from H.264/AVC to a wavelet-based SVC. It aims at exploiting the advantages offered by fine granularity SVC technology when dealing with conventional coders and legacy video. The proposed transcoder was developed to cope with important functionalities of H.264/AVC, such as flexible reference frame (RF) selection. It is able to work with different coding configurations of H.264/AVC, including IPP or IBBP with multiple RFs. Moreover, many of the techniques presented in this paper are generic in the sense that they can be used for transcoding with many popular wavelet-based and hybrid-based video coding architectures. To reduce the transcoder's complexity, motion information and residual data extracted from a compressed H.264/AVC stream are exploited. Experimental results show a very good performance of the proposed transcoder in terms of decoded video quality and system complexity. Eduardo Peixoto, Toni Zgaljic, Ebroul Izquierdo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2012 | Reading Users' Minds From Their Eyes: A Method for Implicit Image AnnotationabstractThis paper explores the possible solutions for image annotation and retrieval by implicitly monitoring user attention via eye-tracking. Features are extracted from the gaze trajectory of users examining sets of images to provide implicit information on the target template that guides visual attention. Our Gaze Inference System (GIS) is a fuzzy logic based framework that analyzes the gaze-movement features to assign a user interest level (UIL) from 0 to 1 to every image that appeared on the screen. Because some properties of the gaze features are unique for every user, our user adaptive framework builds a new processing system for every new user to achieve higher accuracy. The generated UILs can be used for image annotation purposes; however, the output of our system is not limited as it can be used also for retrieval or other scenarios. The developed framework produces promising and reliable UILs where approximately 53% of target images in the users' minds can be identified by the machine with an error of less than 20% and the top 10% of them with no error. We show in this paper that the existing information in gaze patterns can be employed to improve the machine's judgement of image content by assessment of human interest and attention to the objects inside virtual environments. Seyed Navid Haji Mirza, Michael J. Proulx, Ebroul Izquierdo |
IEEE Trans. Multim. | 3 |
| 2011 | Augmented reality mirror for virtual facial alterationsabstractWe present a system for virtual mirror experience that performs attentive facial geometric alterations in augmented reality. The virtual mirror is simulated using commonly available PC with webcam that capture, process and display video in real-time. High realism is obtained by considerate 3D-aware warping of the 2D captured video. A Kalman-based real-time face tracker is used for 3D head pose estimation and accurate facial features localization. The 3D face model used is adapted to the person in front of the mirror by utilizing active shape models for facial landmarks detection, followed by z-depth progressive refining. Geometric adjustments are performed on 3D face vertices while the 2D warping is calculated utilizing the location of back-projected-to-2D face model vertices. The evaluation of our system shows that realistic facial modifications can be rendered in scenarios that correspond to typical usage of a real mirror. Vlado Kitanovski, Ebroul Izquierdo |
ICIP | 2 |
| 2011 | Enhanced visualisation of dance performance from automatically synchronised multimodal recordingsabstractThe Huawei/3DLife Grand Challenge Dataset provides multimodal recordings of Salsa dancing, consisting of audiovisual streams along with depth maps and inertial measurements. In this paper, we propose a system for augmented reality-based evaluations of Salsa dancer performances. An essential step for such a system is the automatic temporal synchronisation of the multiple modalities captured from different sensors, for which we propose efficient solutions. Furthermore, we contribute modules for the automatic analysis of dance performances and present an original software application, specifically designed for the evaluation scenario considered, which enables an enhanced dance visualisation experience, through the augmentation of the original media with the results of our automatic analyses. Marc Gowing, Philip Kelly, Noel E. O'Connor, Cyril Concolato, Slim Essid, Jean Le Feuvre, Robin Tournemenne, Ebroul Izquierdo, Vlado Kitanovski, Qianni Zhang |
ACM Multimedia | 8 |
| 2010 | A Multi-Pattern Search Algorithm for Block Motion Estimation in Video CodingabstractIn this paper, we propose a novel multi-pattern based search technique, TCon search, for fast block matching motion estimation. It starts with small cross shaped and small triangular shaped patterns. Afterwards, based on the previous step optimal motion vector, the search pattern for next step is selected. Except first and last step, each search step considers only three points thus reducing the number of search points significantly. Experimental results demonstrate that the proposed TCon search algorithm performs better than the well-known diamond search (DS) and cross-diamond-hexagonal search (CDHS) algorithms. Compared with the DS algorithm, the proposed TCon search performs up to 2.67 times faster in terms of search point computation and up to 1.67 times faster than CDHS algorithm while comparable quality of reconstructed sequence is maintained. Muhammad Akram 0003, Ebroul Izquierdo |
APWeb | 2 |
| 2010 | An O-FDP Framework in 3D Model Based ReconstructionabstractThree dimensional scene synthesis and analysis has drawn considerable attention in manufacturing industries due to its multifaceted applications ranging from real time recognition, verification, vehicle guidance etc. Nowadays, there is a growing surge in FDP (Focused disparity map) of given objects of interest of multiview stereo images. In order to design customized specific application such as industrial part verification, we propose to use O-FDP (Optimal Focused Disparity Map) that spins off from the fact that only basic geometric information is available from DMU (Digital Mock Up) models such as CATIA in industrial installations. Instead of using the whole image information, we use only the experiential information that is really necessary for application. The proposed framework unifies LIFE(Local Invariant Feature Extraction) techniques, edge information, epipolar geometry and object silhouette information. The framework results are presented and compared with state of the art work. Karthikeyan Vaiapury, Ebroul Izquierdo |
APWeb | 2 |
| 2010 | Detection and enhancement of moving objects in surveillance centric codingabstractThe coexistence of multiple cameras, especially in wireless video surveillance systems imposes severe constraints in terms of computational resources, power supply and bandwidth. These limitations hamper coding and transmission of high-quality video streams. In this paper, we propose a new approach for video coding and transmission of surveillance video. It integrates the coding sub-system and the motion detection module to enhance the quality of moving objects in low-bit rate streams. At the sender side, videos are downsampled and compressed. At the decoder, the video is upsampled and the foreground quality is enhanced by detecting meaningful edges of moving objects via the Hough transform. The quality of the background is also enhanced through progressive update of the background model. Nicola Conci, Ebroul Izquierdo |
ICASSP | 2 |
| 2010 | An Interactive Game for Semi-Automatic Image Annotation
Lasantha Seneviratne, Ebroul Izquierdo |
ICASSP | 2 |
| 2010 | Fast multiframe motion estimation for surveillance videosabstractWe propose a fast multiple reference frames based motion estimation technique for surveillance videos. In the very first reference frame of each motion vector search, successive elimination algorithm is used to find the best motion vector in the search window. For the remaining reference frames, a difference between current and previous reference frames is performed to indentify candidate matching blocks in the current reference frame. Different block matching strategies are proposed to find the optimum motion vector. Experimental evaluation shows that significant reduction in computational complexity can be achieved by applying the proposed strategy. Muhammad Akram 0003, Ebroul Izquierdo |
ICIP | 2 |
| 2010 | Transcoding from H.264/AVC to awavelet-based scalable video codecabstractScalable Video Coding (SVC) enables low complexity adaptation according to transmission and display requirements, providing an efficient solution for video content delivery through heterogeneous networks. However, legacy video and most commercially available content capturing devices use conventional non-scalable coding, e.g., H.264/AVC, to compress and store video streams. As a consequence and in order to fully exploit the advantages of SVC technology, efficient transcoding from conventionally coded to scalable content is urgently needed. In this paper an efficient transcoder from H.264/AVC to a wavelet-based SVC is proposed. The complexity of the transcoder is kept very low by using information extracted directly from the decoded H.264/AVC bitstream, such as motion vectors and the presence of residual data. The proposed approach has been tested with well known benchmarking sequences, showing a good performance in terms of decoded video quality and system complexity. Eduardo Peixoto, Toni Zgaljic, Ebroul Izquierdo |
ICIP | 3 |
| 2010 | Pyramidal Model for Image Semantic SegmentationabstractWe present a new hierarchical model applied to the problem of image semantic segmentation, that is, the association to each pixel in an image with a category label (e.g. tree, cow, building, ...). This problem is usually addressed with a combination of an appearance-based pixel classification and a pixel context model. In our proposal, the images are initially over-segmented in dense patches. The proposed pyramidal model naturally embeds the compositional nature of a scene to achieve a multi-scale contextualisation of patches. This is obtained by imposing an order on the patches aggregation operations towards the final scene. The nodes of the pyramid (that is, a dendrogram) thus represent patch clusters, or super-patches. The probabilistic model favours the homogeneous labelling of super-patches that are likely to contain a single object instance, modelling the uncertainty in identifying such super-patches. The proposed model has several advantages, including the computational efficiency, as well as the expandability. Initial results place the model in line with other works in the recent literature. Giuseppe Passino, Ioannis Patras, Ebroul Izquierdo |
ICPR | 3 |
| 2010 | 3DLife: Bringing the Media Internet to Life
Ebroul Izquierdo, Tomas Piatrik, Qianni Zhang |
ISoLA (2) | 1 |
| 2010 | ACM workshop on surreal media and virtual cloningabstractThis paper gives an overview of ACM Multimedia 2010 Workshop on Surreal Media and Virtual Cloning, including research work towards the creation of surreal media and realistic 3D virtual environments where virtual humans and objects can interact remotely. The primary objective is to discuss key research issues related to the generation of surreal media and 3D cooperative virtual worlds. We expect that the one-day program will bring together research groups from related fields and explore research problems, potential applications and collaborative opportunities. Ebroul Izquierdo, Yang Cai 0002, Qianni Zhang, Manuel García-Herranz |
ACM Multimedia | 1 |
| 2010 | Subjective evaluation of scalable video coding for content distributionabstractThis paper investigates the influence of the combination of the scalability parameters in scalable video coding (SVC) schemes on the subjective visual quality. We aim at providing guidelines for an adaptation strategy of SVC that can select the optimal scalability options for resource-constrained networks. Extensive subjective tests are conducted by using two different scalable video codecs and high definition contents. The results are analyzed with respect to five dimensions, namely, codec, content, spatial resolution, temporal resolution, and frame quality. Jong-Seok Lee, Francesca De Simone, Naeem Ramzan, Zhijie Zhao, Engin Kurutepe, Thomas Sikora, Jörn Ostermann, Ebroul Izquierdo, Touradj Ebrahimi |
ACM Multimedia | 8 |
| 2010 | H.264/AVC to wavelet-based scalable video transcoding supporting multiple coding configurationsabstractScalable Video Coding (SVC) enables low complexity adaptation of the compressed video, providing an efficient solution for video content delivery through heterogeneous networks and to different displays. However, legacy video and most commercially available content capturing devices use conventional non-scalable coding, e.g., H.264/AVC. This paper proposes an efficient transcoder from H.264/AVC to a wavelet-based SVC to exploit the advantages offerend by the SVC technology. The proposed transcoder is able to cope with different coding configurations in H.264/AVC, such as IPP or IBBP with multiple reference frames. To reduce the transcoder's complexity, motion information and presence of the residual data extracted from the decoded H.264/AVC video are exploited. Experimental results show a good performance of the proposed transcoder in terms of decoded video quality and system complexity. Eduardo Peixoto, Toni Zgaljic, Ebroul Izquierdo |
PCS | 3 |
| 2010 | RUSHES - an annotation and retrieval engine for multimedia semantic units
Oliver Schreer, Ingo Feldmann, Isabel Alonso Mediavilla, Pedro Concejero, Abdul Hamid Sadka, Mohammad Rafiq Swash, Sergio Benini, Riccardo Leonardi, Tijana Janjusevic, Ebroul Izquierdo |
Multim. Tools Appl. | 10 |
| 2010 | A Probabilistic Approach for Vision-Based Fire Detection in VideosabstractAutomated fire detection is an active research topic in computer vision. In this paper, we propose and analyze a new method for identifying fire in videos. Computer vision-based fire detection algorithms are usually applied in closed-circuit television surveillance scenarios with controlled background. In contrast, the proposed method can be applied not only to surveillance but also to automatic video classification for retrieval of fire catastrophes in databases of newscast content. In the latter case, there are large variations in fire and background characteristics depending on the video instance. The proposed method analyzes the frame-to-frame changes of specific low-level features describing potential fire regions. These features are color, area size, surface coarseness, boundary roughness, and skewness within estimated fire regions. Because of flickering and random characteristics of fire, these features are powerful discriminants. The behavioral change of each one of these features is evaluated, and the results are then combined according to the Bayes classifier for robust fire recognition. In addition, a priori knowledge of fire events captured in videos is used to significantly improve the classification results. For edited newscast videos, the fire region is usually located in the center of the frames. This fact is used to model the probability of occurrence of fire as a function of the position. Experiments illustrated the applicability of the method. Paulo Vinicius Koerich Borges, Ebroul Izquierdo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | Latent Semantics Local Distribution for CRF-based Image Semantic SegmentationabstractSemantic image segmentation is the task of assigning a semantic label to every pixel of an image. This task is posed as a supervised learning problem in which the appearance of areas that correspond to a number of semantic categories are learned from a dataset of manually labelled images. This paper proposes a method that combines a region-based probabilistic graphical model that builds on the recent success of Conditional Random Fields (CRFs) in the problem of semantic segmentation, with a salient-points-based bagsof-words paradigm. In a first stage, the image is oversegmented into patches. Then, in a CRF-based formulation we learn both the appearance for each semantic category and the neighbouring relations between patches. In addition to patch features, we also consider information extracted on salient points that are detected in the patch’s vicinity. A visual word is associated to each salient point. Two different types of information are used. First, we consider the local weighted distribution of visual words. Using local (i.e. centred at each patch) word histograms enriches the classical global bags-of-word representation with positional information on word distributions. Second, we consider the un-normalised local distribution of a set of latent topics that are obtained by probabilistic Latent Semantic Analysis (pLSA). This distribution is obtained by the weighted accumulation of the latent topic distributions that are associated to the visual words in the area. The advantage of this second approach lays in the separate representation of the semantic content for each visual word. This allows us to consider the word contributions as independent in the CRF formulation without introducing too strong simplification assumptions. Tests on a publicly available dataset demonstrate the validity of the proposed salient point integration strategies. The results obtained with different configurations show an advance compared to other leading works in the area. Giuseppe Passino, Ioannis Patras, Ebroul Izquierdo |
BMVC | 3 |
| 2009 | Subspace clustering of images using Ant colony OptimisationabstractContent-based image retrieval can be dramatically improved by providing a good initial clustering of visual data. The problem of image clustering is that most current algorithms are not able to identify individual clusters that exist in different feature subspaces. In this paper, we propose a novel approach for subspace clustering based on Ant Colony Optimisation and its learning mechanism. The proposed algorithm breaks the assumption that all of the clusters in a dataset are found in the same set of dimensions by assigning weights to features according to the local correlations of data along each dimension. Experiment results on real image datasets show the need for feature selection in clustering and the benefits of selecting features locally. Tomas Piatrik, Ebroul Izquierdo |
ICIP | 2 |
| 2009 | Visualising the Query Space of the Image CollectionabstractIn this paper, we propose an information visualisation solution for multimedia retrieval based on semantic concepts defined in an image database. The proposed visualisation approach utilises concept map, Venn diagram and Fisheye distortion to enable effective and efficient visualisation of image database. In addition, the proposed approach enables displaying the local and global views of the collection subset, selected as relevant by the user. The proposed solution is evaluated on Corel 700 dataset with 10 semantic concepts for the following user actions: exploratory browsing, querying and detecting patterns. Tijana Janjusevic, Ebroul Izquierdo |
IV | 2 |
| 2009 | Integrating Image Segmentation and Classification for Fuzzy Knowledge-Based Multimedia Indexing
Thanos Athanasiadis, Nikos Simou, Georgios Th. Papadopoulos, Rachid Benmokhtar, Krishna Chandramouli, Vassilis Tzouvaras, Vasileios Mezaris, Marios Phinikettos, Yannis Avrithis, Ioannis Kompatsiaris, Benoit Huet, Ebroul Izquierdo |
MMM | 12 |
| 2009 | Special issue on scalable coded media beyond compression
G. Charith K. Abhayaratne, Ebroul Izquierdo, Marta Mrak, Stefano Tubaro |
Signal Process. Image Commun. | 2 |
| 2009 | An efficient optimisation scheme for scalable surveillance centric video communications
Naeem Ramzan, Toni Zgaljic, Ebroul Izquierdo |
Signal Process. Image Commun. | 3 |
| 2009 | Lagrange multiplier selection in wavelet-based scalable video coding for quality scalability
Shuai Wan, Fuzheng Yang 0001, Ebroul Izquierdo |
Signal Process. Image Commun. | 3 |
| 2008 | A probabilistic model for flood detection in video sequencesabstractIn this paper we propose a new image event detection method for identifying flood in videos. Traditional image based flood detection is often used in remote sensing and satellite imaging applications. In contrast, the proposed method is applied for retrieval of flood catastrophes in newscast content, which present great variation in flood and background characteristics, depending on the video instance. Different flood regions in different images share some common features which are reasonably invariant to lightness, camera angle or background scene. These features are texture, relation among color channels and saturation characteristics. The method analyses the frame-to-frame change in these features and the results are combined according to the Bayes classifier to achieve a decision (i.e. flood happens, flood does not happen). In addition, because the flooded region is usually located around the lower and middle parts of an image, a model for the probability of occurrence of flood as a function of the vertical position is proposed, significantly improving the classification performance. Experiments illustrated the applicability of the method and the improved performance in comparison to other techniques. Paulo Vinicius Koerich Borges, Joceli Mayer, Ebroul Izquierdo |
ICIP | 3 |
| 2008 | Global-to-local oriented perception on blurry visual informationabstractA system for perception on blurry visual information is described and validated. Essential capture in global pathway and saliency highlight in local pathway are integrated for low-quality visual information applications. The system can differentiate scenes from various semantic meanings using a spatial layout of context information, which capture the “essential” of the scene. The system can even further discriminate the contents contained in the scene via local highlight. Distinct from previous frameworks, the system presents the entire scheme of being biologically plausible and application efficiency, thus offering a straightforward platform for rapid analysis and interpretation on blurry visual information. Ebroul Izquierdo |
ICIP | 2 |
| 2008 | On the role of structure in part-based object detectionabstractPart-based approaches in image analysis aim at exploiting the considerable discriminative power embedded in relations among image parts. Nonetheless, learning structural information is not always possible without the availability of a training set of classified parts, and taking into account this additional information can even degrade the performance of the system. In this paper, a discriminative graphical model for object detection is introduced and used in order to analyse and report results on the role of structural information in image classification tasks. Giuseppe Passino, Ioannis Patras, Ebroul Izquierdo |
ICIP | 3 |
| 2008 | The sparse image representation for automated image retrievalabstractWe describe a novel sparse image representation for full automated content-based image retrieval using the latent semantic indexing (LSI) approach and also a novel statistical-based model for the efficient dimensional reduction of sparse data. Although images can be represented sparsely for instance by the discrete cosine transform (DCT) coefficients, this sparsity character is destroyed during the LSI-based dimension reduction process. In our approach, we keep the memory limit of the decomposed data by a statistical model of the sparse data. The aim is to find a small but "important" sub-set of coefficients, which represent semantics of images efficiently. The effectiveness of our novel approach is demonstrated by the large scale image similarity task of the NIST TrecVid 2007 benchmark. Pavel Praks, Radek Kucera, Ebroul Izquierdo |
ICIP | 3 |
| 2008 | Efficient scalable video transmission based on two-dimensional error protection schemeabstractAn efficient approach for transmission of scalable video over erroneous channel is proposed. The proposed approach jointly optimises the bit allocation between a wavelet-based scalable video coding framework and a forward error correction codes. The forward error correction is based on the two-dimensional error protection scheme which applies unequal error protection in temporal-spatial layers as well as in quality layers. The scheme minimizes the reconstructed video distortion at the decoder subject to a constraint on the overall transmission bit-rate budget with limited complexity. This minimization is achieved by exploiting the combined scalability in temporal-spatial and quality domain. Experimental results clearly demonstrate the superiority of the proposed approach over conventional forward error correction techniques. It also significantly improves the performance of end to end scalable video transmission at all channel bit rates. Naeem Ramzan, Ebroul Izquierdo |
ICIP | 2 |
| 2008 | Layout Methods for Intuitive Partitioning of Visualization SpaceabstractIn this paper we address two relevant tasks in image visualisation research: layout methods for presenting content and relations within image databases; and optimal solutions for efficient use of the entire display space. We introduce a novel approach to enable users searching on large image archives to distinguish heterogeneous sets of images. Thus, helping them to navigate or browse image databases according to relevant query directions. Two methods for mapping similarity relations between images and cognitive partitioning of the display space are presented. Tijana Janjusevic, Ebroul Izquierdo |
IV | 2 |
| 2008 | Document Image Processing for Paper Side CommunicationsabstractThis paper proposes the use of higher order statistical moments in document image processing to improve the performance of systems which transmit side information through the print and scan channel. Examples of such systems are multilevel 2-D bar codes and certification via text luminance modulation. These systems print symbols with different luminances, according to the target side information. In previous works, the detection of a received symbol is usually performed by evaluating the average luminance or spectral characteristics of the received signal. This paper points out that, whenever halftoning algorithms are used in the printing process, detection can be improved by observing that third and fourth order statistical moments of the transmitted symbol also change, depending on the luminance level. This work provides a thorough analysis for those moments used as detection metrics. A print and scan channel model is exploited to derive the relationship between the modulated luminance level and the higher order moments of a halftone image. This work employs a strategy to merge the different moments into a single metric to achieve a reduced detection error rate. A transmission protocol for printed documents is proposed which takes advantage of the resulting higher robustness achieved with the combined detection metrics. The applicability of the introduced document image analysis approach is validated by comprehensive computer simulations. Paulo Vinicius Koerich Borges, Joceli Mayer, Ebroul Izquierdo |
IEEE Trans. Multim. | 3 |
| 2008 | Robust and Transparent Color Modulation for Text Data HidingabstractThis paper improves the use of text color modulation (TCM) as a reliable text document data hiding method. Using TCM, the characters in a document have their color components modified (possibly unperceptually) according to a side message to be embedded. This work presents a detection metric and an analysis determining the detection error rate in TCM, considering an assumed print and scan (PS) channel model. In addition, a perceptual impact model is employed to evaluate the perceptual difference between a modified and a non-modified character. Combining this perceptual model and the results from the detection error analysis it is possible to determine the optimum color modulation values. The proposed detection metric also exploits the orientation characteristics of color halftoning to reduce the error rate. In particular, because color halftoning algorithms use different screen orientation angles for each color channel, this is used as an effective feature to detect the embedded message. Experiments illustrate the validity of the analysis and the applicability of the method. Paulo Vinicius Koerich Borges, Joceli Mayer, Ebroul Izquierdo |
IEEE Trans. Multim. | 3 |
| 2007 | Efficient side information encoding for text hardcopy documentsabstractThis paper proposes a new coding method that increases significantly the signal-to-watermark ratio in document watermarking algorithms. A possible approach to text document watermarking is to consider text characters as a data structure consisting of several modifiable features such as size, shape, position, luminance, among others. In existing algorithms, these features can be modified sequentially according to bit values to be embedded. In contrast, the solution proposed here uses a positional information coding approach to embed information. Using this approach, the information is related to the position of modified characters, and not to the bit embedded on each character. This coding is based on combinatorial analysis and it can embed more bits in comparison to the usual methods, given a distortion constraint. An analysis showing the superior performance of positional coding for this type of application is presented. Experiments validate the analysis and the applicability of the method. Paulo Vinicius Koerich Borges, Ebroul Izquierdo, Joceli Mayer |
AVSS | 2 |
| 2007 | Wavelet and Eigen-Space Feature Extraction for Classification of Metallography Images
Pavel Praks, Marcin Grzegorzek, Rudolf Moravec, Ladislav Válek, Ebroul Izquierdo |
EJC | 5 |
| 2007 | Performance Analysis of Text Halftone ModulationabstractThis paper analyzes the use of text halftone modulation (THM) as a text hardcopy watermarking method. Using THM, text characters in a document have their luminances modified from the standard black to a gray level generated with a given halftone screen, according to a message to be transmitted. The application of THM has been discussed in ((R. Villan et al., 2006), (K. Matsuit and K. Tanaka, 1994)). In this paper, a spectral metric is proposed to detect the embedded message. Based on this metric, an error rate analysis of halftone modulation is presented considering the effects of the print and scan channel. Experiments validate the analysis and the applicability of the method. Paulo Vinicius Koerich Borges, Joceli Mayer, Ebroul Izquierdo |
ICIP (3) | 3 |
| 2007 | A Knowledge Structuring Technique for Image ClassificationabstractA system for image analysis and classification based on a knowledge structuring technique is presented. The knowledge structuring technique automatically creates a relevance map from salient areas of natural images. It also derives a set of well-structured representations from low-level description to drive the final classification. The backbone of the knowledge structuring technique is a distribution mapping strategy involving two basic modules: structured low-level feature extraction using convolution neural network and a topology representation module based on a growing cell structure network. Classification is achieved by simulating high-level top-down visual information perception and classifying using an incremental Bayesian parameter estimation method. The proposed modular system architecture offers straightforward expansion to include user relevance feedback, contextual input, and multimodal information if available. Ebroul Izquierdo |
ICIP (6) | 2 |
| 2007 | Spatially Adaptive Wavelet Transform for Video Coding with Multi-Scale Motion CompensationabstractIn this paper a technique that enables efficient synthesis of the prediction signal for application in multi-scale motion compensation is presented. The technique targets prediction of high-pass spatial subbands for motion compensation at higher scales. Since in the targeted framework these subbands are obtained by high-pass filtering of prediction signal in pixel domain, an adaptive approach for filtering is proposed to support decomposition of differently predicted frame areas. In this way an efficient application of different prediction modes at all scales used for compensation is enabled. Experimental results show that for fast sequences where such an application of different prediction modes is crucial, the proposed adaptive transform introduces significant objective and visual improvements. Marta Mrak, Ebroul Izquierdo |
ICIP (2) | 2 |
| 2007 | Error Robustness Scheme for Scalable Video Based on the Concatenation of LDPC and Turbo CodesabstractIn this paper, a novel approach for transmission of scalable video over wireless channel is proposed. The proposed approach jointly optimises the bit allocation between a wavelet-based scalable video coding framework and a forward error correction codes. The forward error correction codes is based on the serial concatenation of LDPC codes and turbo codes. Turbo codes shows good performance at high error rates region but LDPC outperforms turbo codes at low error rates. So the concatenation of LDPC and TC enhances the performance at both low and high signal to noise ratios. The scheme minimizes the reconstructed video distortion at the decoder subject to a constraint on the overall transmission bitrate budget. The minimization is achieved by exploiting the source rate distortion characteristics and the statistics of the available codes. Furthermore, an efficient decoding algorithm is proposed. Experimental results clearly demonstrate the superiority of the proposed approach over conventional forward error correction techniques. Naeem Ramzan, Shuai Wan, Ebroul Izquierdo |
ICIP (6) | 3 |
| 2007 | Lagrange Multiplier Selection for 3-D Wavelet Based Scalable Video CodingabstractIn this paper a thorough analysis on the theoretical rate distortion model and the rate distortion performance in an open-loop structure is conducted. A Lagrange multiplier selection for 3D wavelet based scalable video coding is then derived. The proposed Lagrange multiplier is adaptive with respect to the characteristics of video content. Furthermore, it is especially suitable for 3-D wavelet based scalable video coding where quantisation steps are unavailable. Extensive experimental results have demonstrated the effectiveness of the proposed Lagrange multiplier selection. Fuzheng Yang 0001, Shuai Wan, Ebroul Izquierdo |
ICIP (2) | 3 |
| 2007 | Optimised Compression Strategy in Wavelet-Based Video Coding using Improved Context ModelsabstractAccurate probability estimation is a key to efficient compression in entropy coding phase of state-of-the-art video coding systems. Probability estimation can be enhanced if contexts in which symbols occur are used during the probability estimation phase. However, these contexts have to be carefully designed in order to avoid negative effects. Methods that use tree structures to model contexts of various syntax elements have been proven efficient in image and video coding. In this paper we use such structure to build optimised contexts for application in scalable wavelet-based video coding. With the proposed approach context are designed separately for intra-coded frames and motion-compensated frames considering varying statistics across different spatio-temporal subbands. Moreover, contexts are separately designed for different bit-planes. Comparison with compression using fixed contexts from embedded ZeroBlock coding (EZBC) has been performed showing improvements when context modelling on tree structures is applied. Toni Zgaljic, Marta Mrak, Ebroul Izquierdo |
ICIP (3) | 3 |
| 2007 | An Efficient Joint Source-Channel Coding for Wavelet Based Scalable VideoabstractA robust and efficient approach for scalable video transmission over wireless channels is presented. The proposed approach jointly optimizes source and channel coding in order to minimize the overall end-to-end distortion. In particular, the forward error correction method based on turbo codes is considered. Aiming at improving the overall performance of the underlying joint source-channel coding, the combination of the channel coding rate, interleaver and packet size are optimized for turbo codes subject to a constraint on the overall transmission bitrate budget. Experimental results show that the proposed approach outperforms conventional forward error correction techniques at all bit error rates, even in very adverse conditions. Naeem Ramzan, Shuai Wan, Ebroul Izquierdo |
ISCAS | 3 |
| 2007 | Segmentation of Document Images Using Higher Order StatisticsabstractThis work presents an efficient post-segmentation method for separating text from the background in document images. For this task, this paper proposes the use of textured patterns to represent text in documents, instead of the standard black. It is shown that, in poor quality documents, text segmentation is more efficient when the characters in the document are represented in a halftoned gray level prior to printing. This occurs because the halftoning process induces statistical characteristics that help the text to be distinguished from noise or background. A typical case are noisy printed and scanned documents. Experiments validate the analysis and the applicability of the segmentation method. An important application for the method is in the postal service, where letters have their addresses segmented for automatic sorting. Paulo Vinicius Koerich Borges, Joceli Mayer, Ebroul Izquierdo |
MMSP | 3 |
| 2007 | Image classification with the use of radial basis function neural networks and the minimization of the localized generalization error
Wing W. Y. Ng, Andrés Dorado, Daniel S. Yeung, Witold Pedrycz, Ebroul Izquierdo |
Pattern Recognit. | 5 |
| 2007 | Guest Editorial
Luigi Atzori, Ebroul Izquierdo, Pascal Frossard, Özgür B. Akan |
Signal Process. Image Commun. | 2 |
| 2007 | Logotype detection to support semantic-based video annotation
Julián Ramos Cózar, Nicolás Guil, José María González-Linares, Emilio L. Zapata, Ebroul Izquierdo |
Signal Process. Image Commun. | 5 |
| 2007 | Signal processing: Image communication, special issue on content-based multimedia indexing and retrieval
Ebroul Izquierdo, Jenny Benois-Pineau, Régine André-Obrecht |
Signal Process. Image Commun. | 1 |
| 2007 | Perceptually adaptive joint deringing-deblocking filtering for scalable video transmission over wireless networks
Shuai Wan, Marta Mrak, Naeem Ramzan, Ebroul Izquierdo |
Signal Process. Image Commun. | 4 |
| 2007 | Bit-stream allocation methods for scalable video coding supporting wireless communications
Toni Zgaljic, Nikola Sprljan, Ebroul Izquierdo |
Signal Process. Image Commun. | 3 |
| 2007 | Adaptive salient block-based image retrieval in multi-feature space
Qianni Zhang, Ebroul Izquierdo |
Signal Process. Image Commun. | 2 |
| 2007 | An Object- and User-Driven System for Semantic-Based Image Annotation and RetrievalabstractIn this paper, a system for object-based semi-automatic indexing and retrieval of natural images is introduced. Three important concepts underpin the proposed system: a new strategy to fuse different low-level content descriptions; a learning technique involving user relevance feedback; and a novel object based model to link semantic terms and visual objects. To achieve high accuracy in the retrieval and subsequent annotation processes several low-level image primitives are combined in a suitable multifeatures space. This space is modelled in a structured way exploiting both low-level features and spatial contextual relations of image blocks. Support vector machines are used to learn from gathered information through relevance feedback. An adaptive convolution kernel is defined to handle the proposed structured multifeature space. The positive definite property of the introduced kernel is proven, as essential condition for uniqueness and optimality of the convex optimization in support vector machines. The proposed system has been thoroughly evaluated and selected results are reported in this paper Divna Djordjevic, Ebroul Izquierdo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | A Biologically Inspired System for Classification of Natural ImagesabstractA system for visual information analysis and classification based on a biologically inspired visual selective attention model with knowledge structuring is presented. The system is derived from well-known analogous processes in the visual system of primates and inference procedures of the human brain. It consists of three main units: biologically inspired visual selective attention, knowledge structuring, and clustering of visual information. The biologically inspired visual selective attention unit closely follow the mechanisms of the visual what pathway and where pathway in the primates' brain. It uses a bottom-up approach to generate a salient area based on low-level features extracted from natural images. The scale selection to determine suitable size of salient areas uses a maximum entropy approach. This unit also contains a low-level top-down selective attention module that performs decisions on interesting objects by human interaction. In this module, a reinforcement/inhibition mechanism is exploited. The knowledge structuring unit automatically creates a relevance map from salient image areas generated by the biologically inspired unit. It also derives a set of well-structured representations from low-level descriptions to drive the final classification. The knowledge structuring unit relys on human knowledge to produce suitable links between low-level descriptions and high-level representation on a limited training set. The backbone of this unit is a distribution mapping strategy involving two basic modules: structured low-level feature extraction using convolution neural network and a topology representation module based on a growing cell structure network. The third unit of the system classification is achieved by simulating high-level top-down visual information perception and clustering using an incremental Bayesian parameter estimation method. The proposed modular system architecture offers straightforward expansion to include user relevance feedback, contextual input, and multimodal information if available Ebroul Izquierdo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | Editorial: Knowledge Engineering, Semantics, and Signal Processing in Audio-Visual Information RetrievalabstractThe thirteen papers in this special issue focus on knowledge engineering, semantics, and signal processing in audio-visual information retrieval. The selected papers are briefly summarized. Ebroul Izquierdo |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | Rate-Distortion Optimized Motion-Compensated Prediction for Packet Loss Resilient Video CodingabstractA rate-distortion optimized motion-compensated prediction method for robust video coding is proposed. Contrasting methods from the conventional literature, the proposed approach uses the expected reconstructed distortion after transmission, instead of the displaced frame difference in motion estimation. Initially, the end-to-end reconstructed distortion is estimated through arecursive per-pixel estimation algorithm. Then the total bit rate for motion-compensated encoding is predicted using a suitable rate distortion model. The results are fed into the Lagrangian optimization at the encoder to perform motion estimation. Here, the encoder automatically finds an optimized motion compensated prediction by estimating the best tradeoff between coding efficiency and end-to-end distortion. Finally, rate-distortion optimization is applied again to estimate the macroblock mode. This process uses previously selected optimized motion vectors and their corresponding reference frames. It also considers intraprediction. Extensive computer simulations in lossy channel environments were conducted to assess the performance of the proposed method. Selected results for both single and multiple reference frames settings are described. A comparative evaluation using other conventional techniques from the literature was also conducted. Furthermore, the effects of mismatches between the actual channel packet loss rate and the one assumed at the encoder side have been evaluated and reported in this paper. Shuai Wan, Ebroul Izquierdo |
IEEE Trans. Image Process. | 2 |
| 2006 | Mpeg2 Watermarking Channel Protection Using Duo-Binary Turbo CodesabstractThis paper describes scheme for protection of watermarking channel and its capacity enhancement using state-of-the-art error correction technique - turbo coding. Duo-binary codes were used for protection since they are perform better then classical turbo coders in terms of better convergence for iterative decoding, a large minimum distance and also computational expensiveness. A spread spectrum watermarking technique is used to insert watermark. Proposed pseudo-random watermarking bits spreading, amplitude adjustment in DCT domain based on block classification and bit-rate preserving increased signal to noise ratio of the watermarking channel. However, it was essential to introduce an error correction technique in order to achieve high capacity and robustness. In addition, experimental results on robustness to transcoding are presented Ivan Damnjanovic 0001, Naeem Ramzan, Ebroul Izquierdo |
ICASSP (2) | 3 |
| 2006 | An Entropy Coding Scheme for Multi-Component Scalable Motion InformationabstractFully scalable video bit-stream requires layered structure of most of its components. For that reason few methods targeting scalability on motion information have been proposed over the last decade. However, layered representation requires new entropy coding strategies able to efficiently handling of redundancies between different layers. In this paper three methods for entropy coding of layered motion information are proposed. The influence of these schemes on the reconstructed video quality has been also studied Toni Zgaljic, Marta Mrak, Nikola Sprljan, Ebroul Izquierdo |
ICASSP (2) | 4 |
| 2006 | Optimizing Metrics Combining Low-Level Visual Descriptors for Image Annotation and RetrievalabstractAn object oriented approach for key-word based image annotation and classification is presented. It considers combinations of low-level descriptors and suitable metrics to represent and measure similarity between semantically meaningful objects. The objective is to obtain "optimal" metrics based on a linear combination of single metrics and descriptors in a multi-feature space. The proposed approach estimates an optimal linear combination of predefined metrics by applying a Multi-Objective Optimization technique based on a Pareto Archived Evolution Strategy. The proposed approach has been evaluated and tested for annotation of objects in images. Qianni Zhang, Ebroul Izquierdo |
ICASSP (2) | 2 |
| 2006 | Image Classification using Chaotic Particle Swarm OptimizationabstractParticle swarm optimization is one of several meta-heuristic algorithms inspired by biological systems. The chaotic modeling of particle swarm optimization is presented in this paper with application to image classification. The performance of this modified particle swarm optimization algorithm is compared with standard particle swarm optimization. Numerical results of this comparative study are performed on binary classes of images from the Corel dataset. Krishna Chandramouli, Ebroul Izquierdo |
ICIP | 2 |
| 2006 | Knowledge-Based Image Processing for Classification and Recognition in Surveillance ApplicationsabstractThis short paper serves as preface for the ICIP special session on knowledge-based image processing for classification in surveillance applications. This special session presents work on integrative research aimed at automatic classification and semantic-based recognition of scenes and events for surveillance applications. The focus is integrative research on low-level multimedia analysis, knowledge extraction and semantic analysis. The eight papers selected for the special session target convergence of these fields by integrating, for a purpose, what can be disparate disciplines. The integration covers image processing, knowledge representation, information retrieval and semantic analysis. The targeted application scenario involves the processing of input data captured by video-cameras where security is of interest. The work presented in this special session has originated in two large international cooperative projects funded by the European commission under the sixth framework programme of the Information Society Technology (1ST). The mandate and scope of these two projects is very generic embracing several applications that rely on technology for bridging the gap between low-level content descriptions that can be computed automatically by a machine and the richness and subjectivity of semantics in high-level human interpretations of audiovisual media (i.e. the semantic gap). However, the section presented in these conference proceedings is restricted to research work undertaken under more limited scenarios for specific video-based surveillance. Ebroul Izquierdo |
ICIP | 1 |
| 2006 | Evaluation of Techniques for Modeling of Layered Motion StructureabstractMotion information scalability is important for scalable bit-stream adaptation on low bit-rates, when motion rate occupies a significant portion of the total bit-rate. This type of scalability can be achieved by layered representation of motion block partitioning and predictive coding of associated motion vectors across these layers. So far, several approaches for creating layered motion structure targeting quality scalability have been proposed and in this paper their accuracy is evaluated. For that purpose optimal motion models have been found. It has been shown that simple evaluation of reconstruction error at the encoder side improves suboptimal modeling techniques. Marta Mrak, Nikola Sprljan, Ebroul Izquierdo |
ICIP | 3 |
| 2006 | Scalable Video Transmission Using Double Binary Turbo CodeabstractIn this paper, we propose a novel efficient scheme for robust video transmission over wireless channels. The schema consists of motion compensated spatio-temporal wavelet decomposition scalable coder and double binary turbo code. The scalable bitstream is adapted with unequal error protection according to the relevance of video layers in the scalable bitstream. Cyclic redundancy check bits are also added to check error free delivery of scalable bitstream. Double binary turbo codes were used for protection since they perform better than classical turbo coders in terms of better convergence for iterative decoding and computational cost. Selected results of experimental evaluation are reported in the paper. Naeem Ramzan, Ebroul Izquierdo |
ICIP | 2 |
| 2006 | End-to-End Rate-Distortion Optimized Motion EstimationabstractAn end-to-end rate-distortion optimized motion estimation method for robust video coding in lossy networks is proposed. In this method the expected reconstructed distortion after transmission and the total bit rate for displaced frame difference are estimated at the encoder. The results are fed into the Lagrangian optimization at the encoder to perform motion estimation. Here the encoder automatically finds an optimized motion compensated prediction by estimating the best trade off between coding efficiency and end-to-end distortion. Computer simulations in lossy channel environments were conducted to assess the performance of the proposed method. A comparative evaluation using other conventional techniques from the literature was also conducted. Shuai Wan, Ebroul Izquierdo, Fuzheng Yang 0001, Yilin Chang |
ICIP | 2 |
| 2006 | Capacity Enhancement of Compressed Domain Watermarking Channel Using Duo-binary Coding
Ivan Damnjanovic 0001, Ebroul Izquierdo |
IWDW | 2 |
| 2006 | Improving the Efficiency of a Least Median of Squares Schema for the Estimation of the Fundamental MatrixabstractA robust and efficient approach to estimate the fundamental matrix is proposed. The main goal is to reduce the computational cost involved in the estimation when robust schemas are applied. The backbone of the proposed technique is the conventional Least Median of Squares (LMedS) technique. It is well known that the LMedS is one of the most robust regressors for highly contaminated data and unstable models. Unfortunately, its computational complexity renders it useless for practical applications. To overcome this problem, a small number of low-dimensionality least-square problems are solved using well-selected subsets from the input data. The results of this initial approach are fed into the LMedS schema, which is applied to recover the final estimation of the Fundamental matrix. The complexity is substantially reduced by applying a selection process based on an effective statistical analysis of the inherent correlation of the input data. This analysis is used to define a suitable clustering of the data and to drive the subset selection aiming at the reduction of the search space in the LMedS schema. It is shown that avoiding redundancies better estimates can be obtained while keeping the computational cost low. Selected results of computer experiments were conducted to assess the performance of the proposed technique. María Trujillo, Ebroul Izquierdo |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2005 | A Resolution Adaptive Interpolation Technique for Enhanced Decoding of Scalable Coded VideoabstractSubpixel accurate motion compensated temporal filtering introduces a significant coding gain in scalable 3D wavelet video codecs. The influence of the chosen subpixel interpolation technique has not yet been fully analysed in the context of resolution scalability. That problem is addressed in this paper. It is shown that support for increased accuracy and resolution adaptive spatial interpolation needs to be featured in a scalable video decoder, when low resolution sequences are targeted. Using the proposed resolution adaptive filters based on sinc kernels leads to improved decoding performance at low resolution in the sense of achieving higher quality while reducing the complexity of the system. Marta Mrak, Nikola Sprljan, Ebroul Izquierdo |
ICASSP (2) | 3 |
| 2005 | A fast error protection scheme for transmission of embedded coded images over unreliable channels and fixed packet sizeabstractJoint source-channel coding enables efficient transmission of embedded bitstreams over unreliable channels. We address channels with fixed packetisation and decoding without or with minimal delays. The computation of an optimal protection scheme for such bitstreams is generally an exponential complexity problem and hence not applicable in a straightforward implementation. Using the rate-distortion characteristics of the source bitstream and the dynamic programming approach we construct an efficient unequal error protection scheme for the predefined channels. Our algorithm is of linear complexity and thus applicable in real time scenarios. Nikola Sprljan, Marta Mrak, Ebroul Izquierdo |
ICASSP (3) | 3 |
| 2004 | Bringing User Satisfaction to Media Access: The IST BUSMAN ProjectabstractInteractive and seamless access to video content is facilitated by technology that is able to annotate and retrieve media data efficiently and automatically with the minimal human interventions. Media should be accessible by intended users quickly and independent of the database size, regardless of the homogeneous or heterogeneous nature of the delivery channels or the user computing platform. These requirements are driving research and development efforts in the European project BUSMAN; in which a system for efficient annotation, seamless delivery of and access to video components across heterogeneous channels is being developed. The BUSMAN system architecture is based on the needs and expectations of two types of users: content providers and end users. In this paper an overview of the BUSMAN system is presented and its main functionalities are described. Ebroul Izquierdo, Ivan Damnjanovic 0001, Paulo Villegas, Li-Qun Xu, Stephan Herrmann 0002 |
IV | 1 |
| 2004 | Dimensionality Reduction for Content-Based Image ClassificationabstractEffective ways of organizing image descriptors is a critical design step of content-based image classification systems. Suitable descriptors are selected according to the problem domain for generating the feature space. Using several descriptors improves accuracy of representation but risen some challenges such as non linear combination, expensive computation and the curse of dimensionality. In This work an approach using a non parametric statistical test for effective dimensionality reduction is presented. The proposed method facilitates feature discrimination and keeps relevant information. Edyta Mrówka, Andrés Dorado, Witold Pedrycz, Ebroul Izquierdo |
IV | 4 |
| 2004 | A rule-based video annotation systemabstractA generic system for automatic annotation of videos is introduced. The proposed approach is based on the premise that the rules needed to infer a set of high-level concepts from low-level descriptors cannot be defined a priori. Rather, knowledge embedded in the database and interaction with an expert user is exploited to enable system learning. Underpinning the system at the implementation level is preannotated data that dynamically creates signification links between a set of low-level features extracted directly from the video dataset and high-level semantic concepts defined in the lexicon. The lexicon may consist of words, icons, or any set of symbols that convey the meaning to the user. Thus, the lexicon is contingent on the user, application, time, and the entire context of the annotation process. The main system modules use fuzzy logic and rule mining techniques to approximate human-like reasoning. A rule-knowledge base is created on a small sample selected by the expert user during the learning phase. Using this rule-knowledge base, the system automatically assigns keywords from the lexicon to nonannotated video clips in the database. Using common low-level video representations, the system performance was assessed on a database containing hundreds of broadcasting videos. The experimental evaluation showed robust and high annotation accuracy. The system architecture offers straightforward expansion to relevance feedback and autonomous learning capabilities. Andrés Dorado, Janko Calic, Ebroul Izquierdo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2004 | Introduction to the Special Issue on Audio and Video Analysis for Multimedia Interactive Services
Ebroul Izquierdo, Aggelos K. Katsaggelos, Michael G. Strintzis |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2004 | Combining a fuzzy rule-based classifier and illumination Invariance for improved building detectionabstractThe problem of edge-based classification of natural video sequences containing buildings and captured under changing lighting conditions is addressed in this letter. The introduced approach is derived from two empiric observations: In static regions the likelihood of finding features that match the patterns of "buildings" is high because buildings are rigid static objects; and misclassification can be reduced by filtering out image regions changing or deforming in time. These regions may contain objects semantically different to buildings but with a highly similar edge distribution, e.g., high frequency of vertical and horizontal edges. Using these observations a strategy is devised in which a fuzzy rule-based classification technique is combined with a method for changing region detection in outdoor scenes. The proposed approach has been implemented and tested with sequences showing changes in the lighting conditions. Selected results from the experimental evaluation are reported. Vesna Zeljkovic, Andrés Dorado, Ebroul Izquierdo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2003 | Semantic labeling of images combining color, texture and keywordsabstractContent-based image retrieval systems combine perceptual features such as color, texture and shape with semantic concepts for improving the quality of the query's results. In this paper, an annotation technique that combines color and texture with keywords is presented. A method based on color similarity along with a keyword mining technique is used to propagate keywords extracted from a sub-set of annotated images into a large-scale database. A method based on texture properties is applied to link keywords with regions within the images. Finally, an approach for semantic labeling of images is described. In this approach, accuracy of the annotations is estimated and the relationships among keywords are identified. The presented annotation technique is useful for labeling images with keywords construing the underlying semantic content. Andrés Dorado, Ebroul Izquierdo |
ICIP (3) | 2 |
| 2003 | Scalable object-based image retrievalabstractDigital visual libraries have currently available huge amounts of content in unstructured, non-indexed form. Since these collections keep growing fast, retrieving specific images is becoming extremely difficult. It is too slow to linearly search all the stored feature vectors to find those that satisfy the query criteria. Scalability is crucial for an image retrieval system to be practical and realistic. In this paper a simple hierarchical object descriptor scheme, which is compact, flexible, and inherently suited for hierarchical search, is described. By integrating a suitable segmentation algorithm into the descriptor generation schema, the proposed approach becomes object oriented. Basically, features used for the extraction of image regions belonging to single physical objects are used in the definition of object descriptors. The resulting technique generates compact scalable descriptions for each object in the database. Experimental results show the performance of the presented schema in terms of accuracy and scalability. Tsz Ying Lui, Ebroul Izquierdo |
ICIP (3) | 2 |
| 2003 | Semi-Automatic Image Annotation Using Frequent Keyword MiningabstractResearch in content-based image retrieval is an expanding discipline with an accelerated growth in the last ten years. Advances in telecommunications and the huge demand of visual information on Internet and mobile devices is occupying the attention of the researchers in developing efficient systems to ease the task of useful visual information retrieval by the users. We present a semiautomatic image annotation process using the low-level image descriptor fuzzy color signature to extract the most similar images from an annotated database and frequent pattern mining to select the candidates keywords for annotating the new image. The idea is aimed at establishing a bridge between visual data and their interpretation using a weak semantic approach. Andrés Dorado, Ebroul Izquierdo |
IV | 2 |
| 2003 | An ill-posed operator for secure image authenticationabstractMany problems in science and technology are modeled by ill-posed operators and the main difficulty in obtaining accurate solutions is caused by the high instability of such operators. The paper introduces a new schema for secure image authentication. It is based on a rather unconventional approach - the extremely high sensitivity of ill-posed operators to any change in the input data is turned into a tool to achieve fragile watermarking for secure image verification. The ill-posed operator of concern is based on a highly ill-conditioned matrix interrelating the watermark and the original image. Authentication is achieved by solving the least squares problem associated with the underlying linear operator. Regarding general requirements for secure watermark-based authentication, analytical and practical aspects of the introduced technique are discussed. It is shown that the proposed authentication schema is highly secure while providing excellent tamper localization. Several experiments were conducted to demonstrate the effectiveness of the technique when it is subjected to different attacks, including vector quantization counterfeiting and cropping. Ebroul Izquierdo, Valia Guerra-Ones |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2003 | Estimating the essential matrix by efficient linear techniquesabstractIn the problem of recovering the 3D structure of a scene from its 2D projections, a fundamental low-level computer vision task is the estimation of the epipolar geometry. Accurate estimation of the epipolar geometry uses computationally expensive iteration schemes based on nonlinear algebraic constraints to deal with the ill-posedness of the problem. Linear techniques are computationally efficient but extremely unstable. Theoretical and practical aspects of linear methods are analyzed and fundamental results are derived from the study. Two main causes of instability are considered. The first one refers to the lack of homogeneity in the input data. To deal with this problem, a highly efficient scaling approach is introduced. The optimality of the technique is proven theoretically and heuristically. It is shown that a second source of instability arises from the linear dependency between rows of the matrix of the linear system. The effect of this problem in the estimation of the essential matrix is analyzed. An additional strategy is introduced to overcome this difficulty. This strategy improves the stability and accuracy of the linear approach even further while reducing the computational cost. Numerical experiments to evaluate the effectiveness of the proposed techniques are reported. Ebroul Izquierdo, Valia Guerra-Ones |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2003 | Introduction to the special issue on authentication, copyright protection, and information hiding
Ebroul Izquierdo, Hyoung Joong Kim, Benoît Macq |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2003 | Efficient and accurate image based camera registrationabstractA technique for efficient and accurate camera registration based on stereo image analysis is presented. Initially, few correspondences are estimated with high accuracy using a probabilistic relaxation technique. Accuracy is achieved by considering the continuous approximations of selected image areas using second order polynomials and a relaxation rule defined according to the likelihood that estimates obey stereoscopic constraints. The extrinsic camera parameters are then obtained using a novel efficient and robust approach derived from the classic eight point algorithm. Efficiency is achieved by solving a parametric linear optimization problem rather than a nonlinear one as more conventional methods attempt to do. Robustness is obtained by applying two novel strategies: normalization of the initial data via a simple but efficient diagonal scaling approach, and regularization of the underlying linear parametric optimization problem using meaningful constraints. The performance of the presented methods is assessed in several computer experiments using natural video data. Ebroul Izquierdo |
IEEE Trans. Multim. | 1 |
| 2002 | Temporal video segmentation for real-time key frame extractionabstractThe extensive amount of media coverage today, generates difficulties in identifying and selecting desired information. Browsing and retrieval systems become more and more necessary in order to support users with powerful and easy-to-use tools for searching, browsing and summarization of information content. The starting point for these tasks in video browsing and retrieval systems is the low level analysis of video content, especially the segmentation of video content into shots. This paper presents a fast and efficient way to detect shot changes using only the temporal distribution of macroblock types in MPEG compressed video. The notion of a dominant reference frame is introduced here. A dominant frame denotes the reference frame (I or P) used as prediction reference for most of the macroblocks from a subsequent B frame. Janko Calic, Sorin Sav, Ebroul Izquierdo, Seán Marlow, Noel Murphy, Noel E. O'Connor |
ICASSP | 3 |
| 2002 | A generic video analysis and segmentation systemabstractA generic video analysis system for supervised and unsupervised segmentation is described. The idea behind the presented concept is to integrate different advanced segmentation techniques to obtain a robust, efficient and modular segmentation system for natural video and still images. The system entails several independent modules. Each one of these modules encapsulates a complete video processing technique The intermediate results obtained from each single module are merged and further processed by a set of intelligent rules to achieve a highly accurate final segmentation. The modular structure of the system allows it to be extended continuously and with ease by adding new independent modules. The intermediate segmentation results of newly added modules are linked to the other system results via the rule processor. A user friendly graphical interface (GUI) is also provided. The functionality of the GUI is twofold: it serves as input interface to pass processing parameters to the system and as semi-automatic segmentation tool for user interaction and manually refinement of automatically generated segmentation masks. Selected results obtained with the current version of the video analysis system are reported. Ebroul Izquierdo, Jianhui Xia, Roland Mech |
ICASSP | 1 |
| 2002 | A multiresolution technique for video indexing and retrievalabstractThis paper presents a novel approach to multiresolution analysis and scalability in video indexing and retrieval. A scalable algorithm for video parsing and key-frame extraction is introduced. The technique is based on real-time analysis of MPEG motion variables and scalable metrics simplification by discrete contour evolution. Furthermore, a hierarchical key-frame retrieval method using scalable colour histogram analysis is presented. It offers customisable levels of detail in the descriptor space, where the relevance order is determined by degradation of the image, and not by degradation of the image histogram. To assess the performance of the approach, several experiments have been conducted. Selected results are reported. Janko Calic, Ebroul Izquierdo |
ICIP (1) | 2 |
| 2002 | Fuzzy color signaturesabstractWith the large and increasing amount of visual information available in digital libraries and the Web, efficient and robust systems for image retrieval are urgently needed. A compact color descriptor scheme and an efficient metric to compare and retrieve images is presented. An image adaptive color clustering method, called fuzzy color signature, is proposed. The original image colors are mapped into a small number of representative colors using a peaks detection function derived from the color distribution. Fuzzy color signatures are then used as image descriptors. To compare image descriptors the earth mover's distance is used. Several experiments have been conducted to assess the performance of the proposed technique. Andrés Dorado, Ebroul Izquierdo |
ICIP (1) | 2 |
| 2002 | Robust Local Watermarking on Salient Image Areas
Ebroul Izquierdo |
IWDW | 2 |
| 2002 | Key components for an advanced segmentation systemabstractAn advanced image and video segmentation system is proposed. The system builds on existing work, but extends it to achieve efficiency and robustness, which are the two major shortcomings of segmentation methods developed so far. Six different schemes containing several approaches tailored for diverse applications constitute the core of the system. The first two focus on very-low complexity image segmentation addressing real-time applications under specific assumptions. The third scheme is a highly efficient implementation of the powerful nonlinear diffusion model. The other three schemes address the more complex task of physical object segmentation using information about the scene structure or motion. These techniques are based on an extended diffusion model and morphology. The main objective of this work has been to develop a robust and efficient segmentation system for natural video and still images. This goal has been achieved by advancing the state-of-art in terms of pushing forward the frontiers of current methods to meet the challenges of the segmentation task in different situations under reasonable computational cost. Consequently, more efficient methods and novel strategies to issues for which current approaches fail are developed. The performance of the presented segmentation schemes has been assessed by processing several video sequences. Qualitative and quantitative result of this assessment are also reported. Ebroul Izquierdo, Mohammed Ghanbari 0001 |
IEEE Trans. Multim. | 1 |
| 2001 | A novel linear technique to estimate the epipolar geometryabstractThe accurate reconstruction of the 3D scene structure from two different projections and the estimation of the camera scene geometry is of paramount importance in many computer vision tasks. Most of the information about the camera-scene geometry is encapsulated in the fundamental matrix. Estimating the fundamental matrix has been an object of research for many years and continues to be a challenging task in current computer vision systems. While nonlinear iterative approaches have been successful in dealing with the high instability of the underlying problem, their inherent large workload makes these approaches inappropriate for real-time applications. Practical aspects of highly efficient linear methods are studied and a novel low-cost and accurate linear algorithm is introduced. The performance of the proposed approach is assessed by several experiments on real images. Ebroul Izquierdo, Valia Guerra-Ones |
ICASSP | 1 |
| 2001 | A hierarchical approach for low-access latency image indexing and retrievalabstractA major problem to be faced when efficient schemas for image indexing and retrieval are envisaged is the large workload and high complexity of underlying image-processing algorithms. The research leading to this paper focused on this problem, i.e. efficiency and scalability. A novel representation of color histograms is presented. The introduced model is based on hierarchical image adaptive color histograms. One important aspect of the envisaged hierarchic definition is that details at various levels of scale can be obtained. For fast search of similar features in large databases large sets of non-similar images are discriminated at very low computational cost using low detailed descriptions. The search is then progressively refined until only a few very similar objects or images are found and ranked using higher levels of detail. To assess the performance of the approach several experiments have been conducted. Selected results are reported. Ebroul Izquierdo |
ICIP (1) | 1 |
| 2001 | Linear Programming Concept VisualizationabstractA visualization scheme as a tool to find the solution of any 3D-linear programming problem is introduced. The presented approach is highly suitable to draw and visualize interactively the feasible region of a given 3D-linear programming problem. It can be used for a better understanding of the solution process when different methods devoted to solve the underlying problem are applied, e.g., the simplex method. The proposed technique comprises sensitive analysis during the solution process and interactive visualization of the feasible region. The analysis leading to the introduced schema also shows that the design of an appropriate and simple vertex representation is crucial to manage any order of degeneracy. To deal with this paradigm and to some extent to formalize it, the concept of adjacency invariance is introduced. Several experiments have been conducted to test and assess the performance of the introduced concepts and techniques. Jean Pierre Charalambos, Ebroul Izquierdo |
IV | 2 |
| 2001 | Image Based Rendering Using Rational FiltersabstractImage based rendering is a powerful technique for 3D object visualization and representation. Using this approach, arbitrary object or scene views are generated automatically from a small set of real reference images. An algorithm for image based rendering using stereo image pairs is presented. The main goal is to produce a realistic 3D 'continuous look-around' effect along the stereo baseline. To minimize the distortion of the reconstructed pixels due to undefined disparity values, a nonlinear interpolator is proposed. It uses the sparse available disparity map to generate a dense field that partially reconstructs original disparity edge information, producing a sharper intermediate view. The identification of occluded and non-occluded areas is also used to aid the view synthesis process. A special treatment for occluded image areas is also considered in the proposed technique. Several computer experiments have been conducted to assess the performance of the presented method. Ebroul Izquierdo |
IV | 1 |
| 2000 | A Highly Robust Regressor and its Application in Computer Vision
Ebroul Izquierdo |
BMVC | 1 |
| 2000 | A low cost hybrid diffusion technique for object segmentationabstractIn this paper a low-complexity nonlinear filtering technique to smooth textures preserving object contours is presented. The approach is based on a hybrid combination of both isotropic and anisotropic recursive filtering. Using only intensity information the segment borders obtained by applying nonlinear filtering do not necessarily coincide with physical object contours, especially in the case of textured objects. To segment images into regions with physical meaning additional information extracted from disparity or motion is used to weight the filter coefficients. The presented technique has been successfully tested in the context of object segmentation of natural scenes and object-based disparity estimation for stereoscopic applications. Ebroul Izquierdo, Mohammed Ghanbari 0001 |
ICASSP | 1 |
| 2000 | Shape Matching Using a Curvature Based Polygonal Approximation in Scale-SpaceabstractThe emerging MPEG-7 standard demands shape description and shape retrieval techniques. Polygonal approximations of the shape contours give attractive solutions in this domain, because of the description simplicity. This paper introduces a shape matching technique based on the turning function comparison of the shape contour polygonal approximations. António M. G. Pinheiro, Ebroul Izquierdo, Mohammed Ghanbari 0001 |
ICIP | 2 |
| 2000 | Data-Driven Nonlinear Diffusion for Object SegmentationabstractWe propose a novel technique for effective object segmentation, which is based on the combination of image simplification via data-driven nonlinear diffusion and subsequent efficient segmentation of the simplified image. In particular, the data, taking the form of a disparity field from a stereo analysis, has been used to modulate the diffusion process. The strength of this strategy consists of an ability to smooth considerably the details of the imaging scene within the objects' boundaries while inhibiting the diffusion across the boundaries, preserving and even enhancing the object borders. As such, from the simplified image, a simple but efficient histogram-based thresholding and labeling technique can be used to extract precisely an object boundary in its entirety. Li-Qun Xu, Ebroul Izquierdo |
ICIP | 2 |
| 2000 | A robust and efficient scale-space based metric for the evaluation of MPEG-4 VOPsabstractNew MPEG-4 functionalities require the segmentation of input video into different layers (VOPs). Usually, these layers contain arbitrarily shaped objects representing meaningful content of the video stream. With the introduction of the new functionalities in MPEG4, the need of objective and subjective assessment of segmented image quality has emerged. In this paper we introduce an efficient and reliable metric to evaluate segmentation results by comparing them with a given ground truth. Beyond this application the proposed technique can be used for real-time shape description and retrieval in the context of the emerging MPEG7, as well as in general pattern recognition tasks. Selected results obtained by using this metric within these application areas are reported. Ebroul Izquierdo, António M. G. Pinheiro, Mohammed Ghanbari 0001 |
ISCAS | 1 |
| 2000 | Video-Based Camera Registration for Augmented RealityabstractA video-based approach for camera scene registration in augmented reality systems is presented. The presented technique relies on the definition of a model, which is derived from an appropriate parametric linear optimization problem. The optimal parameters are sought in the solution space defined by physically meaningful constraints. Solving the underlying regularized linear problem, we expect to overcome the major shortcoming observed in image-based augmented reality and telepresence systems: the extreme lack of robustness due to the ill-posed nature of the calibration problem. Several computer experiments have been conducted in order to assess the performance of the introduced technique. Ebroul Izquierdo, Valia Guerra-Ones |
IV | 1 |
| 2000 | Image-based rendering and 3D modeling: A complete framework
Ebroul Izquierdo, Jens-Rainer Ohm |
Signal Process. Image Commun. | 1 |
| 1999 | Motion-driven object segmentation in scale-spaceabstractIn this paper we present a method for motion segmentation, in which accurate grouping of pixels undergoing the same motion is targeted. In the presented technique true object edges are first obtained by combining anisotropic diffusion of the original image with edge detection and contour reconstruction in the inherent scale-space. Contours are then matched according to the distance given by a metric defined on their polygonal approximations and the shape of the one-dimensional intensity function along the contour. Masks of objects are obtained by merging image areas inside of edges having the same motion. The performance of the presented technique has been evaluated by computer simulations. Ebroul Izquierdo, Mohammed Ghanbari 0001 |
ICASSP | 1 |
| 1999 | Video Composition by Spatiotemporal Object Segmentation, 3D-Structure and TrackingabstractA stereo vision based system for composition of natural and computer generated images is presented. The system focuses on the solution of four essential tasks in computer vision: disparity estimation, object segmentation, modeling and tracking. These tasks are performed in direct interaction with each other using novel and available multiview analysis techniques and standard computer graphics algorithms. The system is assessed by processing natural video sequences. Selected results are reported. Ebroul Izquierdo, Mohammed Ghanbari 0001 |
IV | 1 |
| 1999 | Disparity/segmentation analysis: matching with an adaptive window and depth-driven segmentationabstractMost of the emerging content-based multimedia technologies are based on efficient methods to solve machine early vision tasks. Among other tasks, object segmentation is perhaps the most important problem in single image processing, whereas pixel-correspondence estimation is the crucial task in multiview image analysis. The solution of these two problems is the key for the development of the majority of leading-edge interactive video-communication technologies and telepresence systems. In this paper, we present a robust framework comprised of joined pixel-correspondence estimation and image segmentation in video sequences taken simultaneously from different perspectives. An improved concept for stereo-image analysis based on block matching with a local adaptive window is introduced. The size and shape of the reference window is calculated adaptively according to the degree of reliability of disparities estimated previously. Considerable improvements are obtained just within object borders or image areas that become occluded by applying the proposed block-matching model. An initial object segmentation is obtained by merging neighboring sampling positions with disparity vectors of similar size and direction. Starting from this initial segmentation, true object borders are detected using a contour-matching algorithm. In this process, the contour of the initial segmentation is taken as a reference pattern, and the edges extracted from the original images, by applying a multiscale algorithm, are the candidates for the true object contour. The performance of the introduced methods has been verified. Ebroul Izquierdo |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1999 | Modeling arbitrary objects based on geometric surface conformityabstractWe address the problem of efficient and flexible modeling of arbitrary three-dimensional (3-D) objects and the accurate tracking of the generated model. These goals are reached by combining available multiview image analysis tools with a straightforward 3-D modeling method, which exploit well-established techniques from both computer vision and computer graphics, improving and combining them with new strategies. The basic idea of the technique presented is to use feature points and relevant edges in the images as nodes and edges of an initial two-dimensional wire grid. The method is adaptive in the sense that an initial rough surface approximation is progressively refined at the locations where the triangular patches do not approximate the surface accurately. The approximation error is measured according to the distance of the model to the object surface, taking into account the reliability of the depth estimated from the stereo image analysis. Once the initial wireframe is available, it is deformed and updated from frame to frame according to the motion of the object points chosen to be nodes. At the end of this process we obtain a temporally consistent 3-D model, which accurately approximates the visible object surface and reflects the physical characteristics of the surface with as few planar patches as possible. The performance of the presented methods is confirmed by several computer experiments. Ebroul Izquierdo |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1998 | Virtual 3D-view generation from stereoscopic video dataabstractMulti viewpoint synthesis from stereoscopic video data is a key technology in most emerging content based multimedia systems. The different techniques presented in the article are based on the analysis of 2D images acquired by two or more cameras. It is shown, how suitable analysis tools for disparity estimation and segmentation can benefit from each other. For the generation of virtual images, two methods with different levels of trade-off between complexity and degree of freedom are described. The first approach is disparity compensated view interpolation, which is capable of generating intermediate views along the interocular axis. The second is a more complex approach, which in a first step generates a 3D model of the object. This model can be rotated to any orientation, the original texture is mapped onto the surface, and arbitrary views can be generated by rendering the obtained surface. Ebroul Izquierdo, Mohammed Ghanbari 0001 |
SMC | 1 |
| 1998 | Image Analysis for 3D Modeling, Rendering, and Virtual View Generation
Ebroul Izquierdo, Silko Kruse |
Comput. Vis. Image Underst. | 1 |
| 1998 | A realtime hardware system for stereoscopic videoconferencing with viewpoint adaptation
Jens-Rainer Ohm, Karsten Grüneberg, Emile A. Hendriks, Ebroul Izquierdo, Dimitris Kalivas, Michael Karl 0001, Dionysis Papadimatos, André Redert |
Signal Process. Image Commun. | 4 |
| 1997 | Stereo matching for enhanced telepresence in three-dimensional videocommunicationsabstractA robust approach for joint motion and disparity estimation in stereo sequences to synthesize arbitrary intermediate views is presented. The improved concept for stereo image analysis is based on a modified block matching algorithm. In which a cost function consisting of area-based correlation together with an appropriately weighted temporal smoothness term is applied. A confidence measure to evaluate the reliability of estimated correspondences is introduced. In occluded image areas and image points with unreliable motion or disparity assignments, considerable improvements are obtained applying an edge-assisted vector interpolation strategy. Two different image synthesis concepts are presented as well. The reported approach is verified by processing a set of sequences taken with stereo cameras having large interaxial distances. Computer simulations show that telepresence illusion with continuous motion parallax and good image quality can be obtained using the methods presented. Ebroul Izquierdo |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1997 | An object-based system for stereoscopic viewpoint synthesisabstractThis paper describes algorithms that were developed for a real-time stereoscopic videoconferencing systems with viewpoint adaptation. The goal is a real telepresence illusion, which is achieved by synthesis of intermediate views from a stereoscopic camera shot with a rather large baseline. The actual viewpoint will be adapted according to the head position of the viewer, such that the impression of motion parallax is produced. The object-based system first identifies foreground and background regions and applies disparity estimation to the foreground object. A hierarchical block matching algorithm is employed for this purpose which takes into account the position of high-activity feature points and the object/background border positions. Using the disparity estimator's output, it is possible to generate arbitrary intermediate views by projections from the left- and right-view images. For this purpose, we have also developed an object-based interpolation algorithm, taking into account a very simple convex-surface model of a person's face and body. Though the algorithms had to be held rather simple under the constraint of hardware feasibility, we obtain a good quality of the intermediate-view images. Finally, we describe the hardware concept for the disparity estimator, which is the most complicated part of the algorithm. Jens-Rainer Ohm, Ebroul Izquierdo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |