VLDB 2026 Research / reviewers in the wild / expert
A. Aydin Alatan
dblp:08/4639
· DBLP profile ↗
109ranked-venue papers
13as first author
15since 2021 · last 2024
0000-0001-5556-7301ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 91 · 12 first-author · 12 since 2021Artificial intelligence and machine learning · 9 · 5 since 2021Databases, data management, data science and information retrieval · 6 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-authorSecurity and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Deep Metric Learning with Chance ConstraintsabstractDeep metric learning (DML) aims to minimize empirical expected loss of the pairwise intra-/inter- class proximity violations in the embedding space. We relate DML to feasibility problem of finite chance constraints. We show that minimizer of proxy-based DML satisfies certain chance constraints, and that the worst case generalization performance of the proxy-based methods can be characterized by the radius of the smallest ball around a class proxy to cover the entire domain of the corresponding class samples, suggesting multiple proxies per class helps performance. To provide a scalable algorithm as well as exploiting more proxies, we consider the chance constraints implied by the minimizers of proxy-based DML instances and reformulate DML as finding a feasible point in intersection of such constraints, resulting in a problem to be approximately solved by iterative projections. Simply put, we repeatedly train a regularized proxy-based loss and re-initialize the proxies with the embeddings of the deliberately selected new samples. We applied our method with 4 well-accepted DML losses and show the effectiveness with extensive evaluations on 4 popular DML benchmarks. Code is available at: https://github.com/yetigurbuz/ccp-dml Yeti Ziya Gurbuz, Ogul Can, A. Aydin Alatan |
WACV | 3 |
| 2023 | Knowledge Distillation Layer that Lets the Student Decide
Ada Gorgun, Yeti Ziya Gurbuz, A. Aydin Alatan |
BMVC | 3 |
| 2023 | Generalized Sum Pooling for Metric LearningabstractA common architectural choice for deep metric learning is a convolutional neural network followed by global average pooling (GAP). Albeit simple, GAP is a highly effective way to aggregate information. One possible explanation for the effectiveness of GAP is considering each feature vector as representing a different semantic entity and GAP as a convex combination of them. Following this perspective, we generalize GAP and propose a learnable generalized sum pooling method (GSP). GSP improves GAP with two distinct abilities: i) the ability to choose a subset of semantic entities, effectively learning to ignore nuisance information, and ii) learning the weights corresponding to the importance of each entity. Formally, we propose an entropy-smoothed optimal transport problem and show that it is a strict generalization of GAP, i.e., a specific realization of the problem gives back GAP. We show that this optimization problem enjoys analytical gradients enabling us to use it as a direct learnable replacement for GAP. We further propose a zero-shot loss to ease the learning of GSP. We show the effectiveness of our method with extensive evaluations on 4 popular metric learning benchmarks. Code is available at: GSP-DML Framework Yeti Ziya Gurbuz, Ozan Sener, A. Aydin Alatan |
ICCV | 3 |
| 2023 | Generalizable Embeddings with Cross-Batch Metric LearningabstractGlobal average pooling (GAP) is a popular component in deep metric learning (DML) for aggregating features. Its effectiveness is often attributed to treating each feature vector as a distinct semantic entity and GAP as a combination of them. Albeit substantiated, such an explanation’s algorithmic implications to learn generalizable entities to represent unseen classes, a crucial DML goal, remain unclear. To address this, we formulate GAP as a convex combination of learnable prototypes. We then show that the prototype learning can be expressed as a recursive process fitting a linear predictor to a batch of samples. Building on that perspective, we consider two batches of disjoint classes at each iteration and regularize the learning by expressing the samples of a batch with the prototypes that are fitted to the other batch. We validate our approach on 4 popular DML benchmarks. Yeti Ziya Gurbuz, A. Aydin Alatan |
ICIP | 2 |
| 2023 | E-VFIA: Event-Based Video Frame Interpolation with AttentionabstractVideo frame interpolation (VFI) is a fundamental vision task that aims to synthesize several frames between two consecutive original video images. Most algorithms aim to accomplish VFI by using only keyframes, which is an ill-posed problem since the keyframes usually do not yield any accurate precision about the trajectories of the objects in the scene. On the other hand, event-based cameras provide more precise information between the keyframes of a video. Some recent state-of-the-art event-based methods approach this problem by utilizing event data for better optical flow estimation to interpolate for video frame by warping. Nonetheless, those methods heavily suffer from the ghosting effect. On the other hand, some of kernel-based VFI methods that only use frames as input, have shown that deformable convolutions, when backed up with transformers, can be a reliable way of dealing with long-range dependencies. We propose event-based video frame interpolation with attention (E-VFIA), as a lightweight kernelbased method. E-VFIA fuses event information with standard video frames by deformable convolutions to generate high quality interpolated frames. The proposed method represents events with high temporal resolution and uses a multi-head selfattention mechanism to better encode event-based information, while being less vulnerable to blurring and ghosting artifacts; thus, generating crispier frames. The simulation results show that the proposed technique outperforms current state-of-the-art methods (both frame and event-based) with a significantly smaller model size. Multimedia material: The code is available at https://github.com/ahmetakman/E-VFIA Onur Selim Kiliç, Ahmet Akman, A. Aydin Alatan |
ICRA | 3 |
| 2023 | TMO-Det: Deep tone-mapping optimized with and for object detection
Ismail Hakki Kocdemir, Alper Koz, Ahmet Oguz Akyüz, Alan Chalmers, A. Aydin Alatan, Sinan Kalkan |
Pattern Recognit. Lett. | 5 |
| 2022 | Feature Embedding by Template Matching as a ResNet Block
Ada Gorgun, Yeti Ziya Gurbuz, A. Aydin Alatan |
BMVC | 3 |
| 2022 | Depth is all you Need: Single-Stage Weakly Supervised Semantic Segmentation From Image-Level SupervisionabstractThe costly process of obtaining semantic segmentation labels has driven research towards to weakly supervised semantic segmentation (WSSS) methods, with only image-level labels available for training. The lack of dense semantic scene representation requires methods to increase complexity to obtain additional semantic information (i.e. object/stuff extent and boundary) about the scene. This is often done though increased model complexity and sophisticated multi-stage training/refinement procedures. However, the lack of 3D geometric structure of a single image makes these efforts desperate at a certain point. In this work, we propose to harness (inverse) depth maps estimated from one single image via a monocular depth estimation model to integrate the 3D geometric structure of the scene into the segmentation model. In light of this proposal, we develop an end-to-end segmentation-based network model and a self-supervised training process to train for semantic masks from only image-level annotations in a single stage. Our experiments show that our one-stage method achieves comparable segmentation performance (val: 64.32, test: 64.91) on Pascal VOC when compared with those significantly more complex pipelines and outperforms SOTA single-stage methods. Mustafa Ergul, A. Aydin Alatan |
ICIP | 2 |
| 2022 | Improved Hard Example Mining Approach for Single Shot Object DetectorsabstractHard example mining methods generally improve the performance of the object detectors, which suffer from imbalanced training sets. In this work, two existing hard example mining approaches (LRM and focal loss, FL) are adapted and combined in a state-of-the-art real-time object detector, YOLOv5. The effectiveness of the proposed approach for improving the performance on hard examples is extensively evaluated. The proposed method increases mAP by 3% compared to using the original loss function and around 1-2% compared to using the hard-mining methods (LRM or FL) individually on 2021 Anti-UAV Challenge Dataset. Aybora Koksal, Önder Tuzcuoglu, Kutalmis Gokalp Ince, Yoldas Ataseven, A. Aydin Alatan |
ICIP | 5 |
| 2022 | Object Detection for Autonomous Driving: High-Dynamic Range vs. Low-Dynamic Range ImagesabstractAn important problem in autonomous driving is to perceive objects even under challenging illumination conditions. Despite this problem, existing solutions use low-dynamic range (LDR) images for object detection for autonomous driving. In this paper, we provide a novel analysis on whether high-dynamic range (HDR) images can provide better performance for object detection for autonomous driving. To this end, we choose a seminal deep object detector and systematically evaluate its performance when trained with (i) LDR images, (ii) HDR images, and (iii) tone-mapped LDR images for scenes with different illuminations. We show that a detector with HDR images pre-processed with normalization and gamma correction can only marginally perform better than a detector with LDR or tone-mapped LDR images. Our analysis of this unexpected finding reveals that a detector with HDR images requires significantly more samples as the space of HDR images is significantly larger than that of LDR images. Ismail Hakki Kocdemir, Ahmet Oguz Akyüz, Alper Koz, Alan Chalmers, A. Aydin Alatan, Sinan Kalkan |
MMSP | 5 |
| 2022 | Vision-based estimation of the number of occupants using video cameras
Ipek Gursel Dino, M. Esat Kalfaoglu, Orcun Koral Iseri, Bilge Erdogan, Sinan Kalkan, A. Aydin Alatan |
Adv. Eng. Informatics | 6 |
| 2021 | Blind Deinterleaving of Signals in Time Series with Self-Attention Based Soft Min-Cost Flow LearningabstractWe propose an end-to-end learning approach to address deinterleaving of patterns in time series, in particular, radar signals. We link signal clustering problem to min-cost flow as an equivalent problem once the proper costs exist. We formulate a bi-level optimization problem involving min-cost flow as a sub-problem to learn such costs from the supervised training data. We then approximate the lower level optimization problem by self-attention based neural networks and provide a trainable framework that clusters the patterns in the input as the distinct flows. We evaluate our method with extensive experiments on a large dataset with several challenging scenarios to show the efficiency. Ogul Can, Yeti Ziya Gurbuz, Berkin Yildirim, A. Aydin Alatan |
ICASSP | 4 |
| 2021 | Deep Metric Learning With Alternating Projections Onto Feasible SetsabstractMinimizers of the typical distance metric learning loss functions can be considered as “feasible points” satisfying a set of constraints imposed by the training data. We reformulate distance metric learning problem as finding a feasible point of a constraint set where the embedding vectors of the training data satisfy desired intra-class and inter-class proximity. The feasible set induced by the constraint set is expressed as the intersection of the relaxed feasible sets which enforce the proximity constraints only for particular samples (a sample from each class) of the training data. Then, the feasible point problem is to be approximately solved by performing alternating projections onto those feasible sets. Such an approach introduces a regularization term and results in minimizing a typical loss function with a systematic batch set construction where these batches are constrained to contain the same sample from each class for a certain number of iterations. The proposed technique is applied with the well-accepted losses and evaluated on three popular benchmark datasets for image retrieval and clustering. Outperforming state-of-the-art, the proposed approach consistently improves the performance of the integrated loss functions with no additional computational cost. Ogul Can, Yeti Ziya Gurbuz, A. Aydin Alatan |
ICIP | 3 |
| 2021 | Metu Loss: Metric Learning With Entangled Triplet Unified LossabstractMetric learning aims to define a distance that measures the semantic difference between the instances in a dataset. In this paper, we analyze several well-known triplet loss functions and argue that the gradients of these triplet loss functions do not move the instances in each triplet in the desired direction with the right magnitude. Hence, in order to determine precise interacting forces, we establish a weak analogy with the phenomena of the electromagnetic forces affecting a charged body in free space. Since gradients of the loss function (matching up to the potential energy) with respect to the anchor, positive and negative instances (corresponding to the point charges) of any valid triplet give forces, a loss function can be obtained in a reverse manner, i.e., starting from the desired interacting forces. Based on this idea, we propose a novel triplet loss function, namely, metric learning with entangled triplet unified (METU) loss, that considers the distances between instances in a triplet as well. In order to present only the effect of the loss function, no mining (including the effect of hinge function) is utilized during the experiments. Based on the results on the used fine-grained dataset, CUB-200-2011, it can be concluded that the proposed loss function outperforms the scores of the state-of-the-art methods in the same category. Kaan Karaman, A. Aydin Alatan |
ICIP | 2 |
| 2021 | HDR Image Construction from Trifocal Multiexposure ImagesabstractWith the progress of autonomous vehicles, the sensing of the environment in more detail with higher dynamic ranges has become more important to classify surrounding objects and obstacles. While stereo HDR images for this purpose can provide advantages compared to the conventional LDR images, they suffer from limited dynamic ranges and spike-like noises due to the inaccuracies in disparity estimation. In this paper, we formulate the HDR image construction problem from trifocal multi-exposure images and develop a method which improves disparity estimation for better HDR image construction. Given the symmetric geometry of the trifocal setup, the proposed method uses the equivalence of disparities from middle to left and middle to right images to determine the reliable regions. The HDR radiance for the pixels in these reliable regions are estimated by using the weighted average of the warped images in different exposures and the middle image, whereas the radiance values outside the reliable regions are estimated by using only the middle image. The experiments with different exposure combinations for left, middle and right images reveal better performances of the proposed method compared to the stereo HDR imaging. It is also observed that the improvements are more apparent for larger disparities between the cameras. Alper Koz, Baris Demirkiliç, Yunus Bilge Kurt, Ahmet Oguz Akyüz, Sinan Kalkan, A. Aydin Alatan, Alan Chalmers |
MMSP | 6 |
| 2019 | A Novel BoVW Mimicking End-To-End Trainable CNN Classification Framework Using Optimal Transport TheoryabstractAn end-to-end trainable convolutional neural network (CNN) framework which mimics bag of visual words (BoVW) is proposed for image classification. To this end, a new paradigm for histogram-like image representation is introduced and optimal transport (OT) distance is utilized for the similarity assessment. Any patch of an image is considered as a unique visual word and the image is represented as the uniform histogram of the visual words with the histogram bins associated to embedding vectors according to the semantic meanings of the corresponding visual words. Thus, in the CNN framework, the output of the last convolutional block is considered as the global representation of the image and the embeddings are inherently learned within the classification framework. With the proposed formulation, undesired quantization for the BoVW representation is no more required; moreover, the learned CNN features are naturally interpretable. The experiments on CIFAR-10, CIFAR-100 and SVHN datasets show that the replacement of the global pooling and fully connected layers with the proposed representation together with OT distance improves the baseline CNN framework. Yeti Ziya Gurbuz, A. Aydin Alatan |
ICIP | 2 |
| 2019 | Quadruplet Selection Methods for Deep Embedding LearningabstractRecognition of objects with subtle differences has been used in many practical applications, such as car model recognition and maritime vessel identification. For discrimination of the objects in fine-grained detail, we focus on deep embedding learning by using a multi-task learning framework, in which the hierarchical labels (coarse and fine labels) of the samples are utilized both for classification and a quadruplet-based loss function. In order to improve the recognition strength of the learned features, we present a novel feature selection method specifically designed for four training samples of a quadruplet. By experiments, it is observed that the selection of very hard negative samples with relatively easy positive ones from the same coarse and fine classes significantly increases some performance metrics in a fine-grained dataset when compared to selecting the quadruplet samples randomly. The feature embedding learned by the proposed method achieves favorable performance against its state-of-the-art counterparts. Kaan Karaman, Erhan Gundogdu, Aykut Koç, A. Aydin Alatan |
ICIP | 4 |
| 2019 | Single Image Noise Level Estimation Using Dark Channel PriorabstractNoise level is required as an input parameter in various image processing applications. In this work, we use the dark channel prior (DCP) to estimate the noise level of an image degraded by additive white Gaussian noise. We develop an approximate model of the probability density function of the dark channel of the noisy image. Using this model, the noise level is determined with the maximum likelihood estimation method from the dark channel intensity values of the noisy image. The results show that our method is faster than the state-of-the-art methods by about two orders of magnitude while providing slightly inferior estimation performance. Aziz Berkay Yesilyurt, Aybüke Erol, Fatih Kamisli, A. Aydin Alatan |
ICIP | 4 |
| 2018 | Extended Object Tracking and Shape ClassificationabstractRecent extended target tracking algorithms provide reliable shape estimates while tracking objects. The estimated extent of the objects can also be used for online classification. In this work, we propose to use a Bayesian classifier to identify different objects based on their contour estimates during tracking. The proposed method uses the uncertainty information provided by the estimation covariance of the tracker. Barkin Tuncer, Murat Kumru, Emre Özkan, A. Aydin Alatan |
FUSION | 4 |
| 2018 | Object Localization Without Bounding Box Information Using Generative Adversarial Reinforcement LearningabstractObject localization can be defined as the task of finding the bounding boxes of objects in a scene. Most of the state-of-the-art approaches utilize meticulously handcrafted training datasets. In this work, we are aiming to create a generative adversarial reinforcement learning framework, which can work without having any explicit bounding box information. Instead of relying on bounding boxes, our framework uses tightly cropped object images as training data. Our image localization framework consists of two parts: a reinforcement learning agent (RL agent) and a discriminator. The RL agent takes input scenes and crops them with the objective of creating a tightly cropped object image. The discriminator tries to distinguish whether the image is generated by the RL agent or it comes from a tightly cropped object database. Experiments indicate that it is possible to achieve a promising localization performance without having explicit bounding box data. It can be concluded that generative adversarial reinforcement learning is an important tool in dealing with other learning problems where explicit input/output paired data is not available. Eren Halici, A. Aydin Alatan |
ICIP | 2 |
| 2018 | Improving Proposal-Based Object Detection Using Convolutional Context FeaturesabstractA novel extension to proposal-based detection is proposed in order to learn convolutional context features for determining boundaries of objects better. Objects and their context are aimed to be learned through parallel convolutional stages. The resulting object and context feature maps are combined in such a way that they preserve their spatial relationship. The proposed algorithm is trained and evaluated on PASCAL VOC 2007 detection benchmark dataset and yielded improvements in performance over baseline, for all classes, especially the ones with distinctive context. Emre Can Kaya, A. Aydin Alatan |
ICIP | 2 |
| 2018 | Fine-grained recognition of maritime vessels and land vehicles by deep feature embeddingabstractRecent advances in large‐scale image and video analysis have empowered the potential capabilities of visual surveillance systems. In particular, deep learning‐based approaches bring in substantial benefits in solving certain computer vision problems such as fine‐grained object recognition. Here, the authors mainly concentrate on classification and identification of maritime vessels and land vehicles, which are the key constituents of visual surveillance systems. Employing publicly available data sets for maritime vessels and land vehicles, the authors aim to improve visual recognition. Specifically, the authors focus on five tasks regarding visual recognition; coarse‐grained classification, fine‐grained classification, coarse‐grained retrieval, fine‐grained retrieval, and verification. To increase the performance in these tasks, the authors utilise a multi‐task learning framework and present a novel loss function which simultaneously considers deep feature learning and classification by exploiting the available hierarchical labels of individual samples and the global statistics of distances between the data pairs. The authors observe that the proposed multi‐task learning model improves the fine‐grained recognition performance on MARVEL and Stanford Cars data sets, compared to training of a model targeting a single recognition task. Berkan Solmaz, Erhan Gundogdu, Veysel Yücesoy, Aykut Koç, A. Aydin Alatan |
IET Comput. Vis. | 5 |
| 2018 | Analysis of Airborne LiDAR Point Clouds With Spectral Graph FilteringabstractSeparation of ground and nonground measurements is an essential task in the analysis of light detection and ranging (LiDAR) point clouds; however, it is challenge to implement a LiDAR filtering algorithm that integrates the mathematical definition of various landforms. In this letter, we propose a novel LiDAR filtering algorithm that adapts to the irregular structure and 3-D geometry of LiDAR point clouds. We exploit weighted graph representations to analyze the 3-D point cloud on its original domain. Then, we consider airborne LiDAR data as an irregular elevation signal residing on graph vertices. Based on a spectral graph approach, we introduce a new filtering algorithm that distinguishes ground and nonground points in terms of their spectral characteristics. Our complete filtering framework consists of outlier removal, iterative graph signal filtering, and erosion steps. Experimental results indicate that the proposed framework achieves a good accuracy on the scenes with data gaps and classifies the nonground points on bridges and complex shapes satisfactorily, while those are usually not handled well by the state-of-the-art filtering methods. Eda Bayram, Pascal Frossard, Elif Vural, A. Aydin Alatan |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2018 | Good Features to Correlate for Visual TrackingabstractDuring the recent years, correlation filters have shown dominant and spectacular results for visual object tracking. The types of the features that are employed in these family of trackers significantly affect the performance of visual tracking. The ultimate goal is to utilize robust features invariant to any kind of appearance change of the object, while predicting the object location as properly as in the case of no appearance change. As the deep learning based methods have emerged, the study of learning features for specific tasks has accelerated. For instance, discriminative visual tracking methods based on deep architectures have been studied with promising performance. Nevertheless, correlation filter based (CFB) trackers confine themselves to use the pre-trained networks which are trained for object classification problem. To this end, in this manuscript the problem of learning deep fully convolutional features for the CFB visual tracking is formulated. In order to learn the proposed model, a novel and efficient backpropagation algorithm is presented based on the loss function of the network. The proposed learning framework enables the network model to be flexible for a custom design. Moreover, it alleviates the dependency on the network trained for classification. Extensive performance analysis shows the efficacy of the proposed custom design in the CFB tracking framework. By fine-tuning the convolutional parts of a state-of-the-art network and integrating this model to a CFB tracker, which is the top performing one of VOT2016, 18% increase is achieved in terms of expected average overlap, and tracking failures are decreased by 25%, while maintaining the superiority over the state-of-the-art methods in OTB-2013 and OTB-2015 tracking datasets. Erhan Gundogdu, A. Aydin Alatan |
IEEE Trans. Image Process. | 2 |
| 2017 | Roadesic distance: Flow-aware tracklet association cost for wide area surveillanceabstractLong-term multi-target tracking via tracklet merging in wide area surveillance has crucial importance to improve tracker performances and operational requirements. Min-cost network flow formulation for multi-target tracking is adopted for the tracklet merging problem. In order to improve the continuity of the computed flows by the min-cost network flow framework, a novel tracklet association cost is proposed to be utilized in this network. The proposed cost is based on connecting two tracklets by considering the traffic flow which is estimated from the precomputed tracklets. Such an approach enforces spatial consistencies between tracks by imposing these relations into the association cost. Hence, without violating the min-cost network flow formulation, a constraint to enforce spatial consistency can be implicitly obtained. The proposed cost function can be further exploited to interpolate gaps between the merged tracklets for postprocessing. The experimental results show that proposed association cost improves baseline framework that uses costs considering only two tracklets at a time, as well as some other tracklet merge algorithms from the literature. Yeti Ziya Gurbuz, Ogul Can, A. Aydin Alatan |
ICIP | 3 |
| 2017 | Extending Correlation Filter-Based Visual Tracking by Tree-Structured Ensemble and Spatial WindowingabstractCorrelation filters have been successfully used in visual tracking due to their modeling power and computational efficiency. However, the state-of-the-art correlation filter-based (CFB) tracking algorithms tend to quickly discard the previous poses of the target, since they consider only a single filter in their models. On the contrary, our approach is to register multiple CFB trackers for previous poses and exploit the registered knowledge when an appearance change occurs. To this end, we propose a novel tracking algorithm [of complexity O(D)] based on a large ensemble of CFB trackers. The ensemble [of size O(2D)] is organized over a binary tree (depth D), and learns the target appearance subspaces such that each constituent tracker becomes an expert of a certain appearance. During tracking, the proposed algorithm combines only the appearance-aware relevant experts to produce boosted tracking decisions. Additionally, we propose a versatile spatial windowing technique to enhance the individual expert trackers. For this purpose, spatial windows are learned for target objects as well as the correlation filters and then the windowed regions are processed for more robust correlations. In our extensive experiments on benchmark datasets, we achieve a substantial performance increase by using the proposed tracking algorithm together with the spatial windowing. Erhan Gundogdu, Huseyin Ozkan, A. Aydin Alatan |
IEEE Trans. Image Process. | 3 |
| 2016 | Ensemble Of adaptive correlation filters for robust visual trackingabstractCorrelation filters have recently been popular due to their success in short-term single-object tracking as well as their computational efficiency. Nevertheless, the appearance model of a single correlation filter based tracking algorithm quickly forgets the past poses of the target object due to the updates over time. To overcome this undesired forgetting, our approach is to run trackers with separate models simultaneously. Hence, we propose a novel tracker relying on an ensemble of correlation filters, where the ensemble is obtained via a decision tree partitioning in the object appearance space. Our technique efficiently searches among the ensemble trackers and activates the ones which are most specialized on the current object appearance. Our tracking method is capable of switching frequently in the ensemble. Thus, an inherently adaptive and non-linear learning rate is achieved. Moreover, we demonstrate the superior performance of our method in benchmark video sequences. Erhan Gundogdu, Huseyin Ozkan, A. Aydin Alatan |
AVSS | 3 |
| 2016 | Fisher-selective search for object detectionabstractAn enhancement to one of the existing visual object detection approaches is proposed for generating candidate windows that improves detection accuracy at no additional computational cost. Hypothesis windows for object detection are obtained based on Fisher Vector representations over initially obtained superpixels. In order to obtain new window hypotheses, hierarchical merging of superpixel regions are applied, depending upon improvements on some objectiveness measures with no additional cost due to additivity of Fisher Vectors. The proposed technique is further improved by concatenating these representations with that of deep networks. Based on the results of the simulations on typical data sets, it can be argued that the approach is quite promising for its use of handcrafted features left to dust due to the rise of deep learning. Ilker Buzcu, A. Aydin Alatan |
ICIP | 2 |
| 2016 | Spatial windowing for correlation filter based visual trackingabstractCorrelation filters have been extensively studied to address online visual object tracking task, while achieving favourable performance against the-state-of-the-art methods in various benchmark datasets. Nevertheless, undesired conditions, i.e. partial occlusions or abrupt deformations of the object appearance, severely degrade the performance of correlation filter based tracking methods. To this end, we propose a method for estimating a spatial window for the object observation such that the correlation output of the correlation filter and the windowed observation (i.e. element-wise multiplication of the window and the observation) is improved, especially in these adverse conditions. This approach leads to a performance uplift in the tracking result compared to the classical windowing operation. Moreover, the estimated spatial window of the object patch indicates the object regions that are useful for correlation. We observe a considerable amount of performance increase in the benchmark video sequences by using the proposed visual tracking method. Erhan Gundogdu, A. Aydin Alatan |
ICIP | 2 |
| 2016 | Object classification in infrared images using deep representationsabstractIn this study, we address the problem of infrared (IR) object classification that divides the object appearance space hierarchically with a binary decision tree structure. Binary decisions are made by using the special features of the object appearances. These features are extracted using a fully connected deep neural network learnt by training samples. At each node of the tree, we train individual deep CNNs such that each node specializes in its corresponding subspace. The proposed classification algorithm is evaluated in our generated dataset, which consists of IR targets collected from different video records obtained from different IR sensors (both midwave and longwave) and taken from real world field. The generated dataset consists of four different class labels as ship/boat, tank, plane and helicopter containing a total of 16K samples. Using the proposed tree-based classifier, we observe a favourable performance increase in our dataset against a single deep CNN classifier. Erhan Gundogdu, Aykut Koç, A. Aydin Alatan |
ICIP | 3 |
| 2016 | Superpixel based hyperspectral target detectionabstractUsing the spectral signature of a target by means of matching the signature with the pixels of an acquired hyperspectral image has been proven as an effective way of classifying hyperspectral pixels in most of the proposed methods in hyperspectral image analysis. A disadvantage of these methods is however to use only the spectral characteristics of pixels for detection while ignoring the spatial relations between the neighbouring pixels. In this paper, we propose a hyperspectral target detection method which uses also the spatial neigboorhood information as well as the spectral characteristics of hyperspectral pixels. To this end, we first utilize superpixelization method [1] to describe the neigborhood relation between the hyperspectral pixels, which has been previously developed and proved to be better compared to a pioneer state-of-the-art superpixel algorithm, SLIC [2]. Second, we investigate the best representatives for superpixels among different alternatives, such as centroids, medoid and mean, and modify the well-known hyperspectral target detection algorithm using orthogonal subspace projection, DTDCA [3], appropriately for superpixels. The improvements of the proposed approach over DTDCA in terms of the detection and false detection rates are verified on real hyperspectral images taken from wheat and corn fields with a VNIR camera. Akin Caliskan, Emrecan Bati, Alper Koz, A. Aydin Alatan |
IGARSS | 4 |
| 2016 | Automatic road detection from gray-level images in Wide Area SurveillanceabstractWide Area Surveillance (WAS) systems are capable of providing continuous surveillance of critical areas as wide as city center (approximately 20 km square), mostly as a gray-level video. Utilization of road information for WAS systems increases moving vehicle tracking performance, while reducing the false alarm rates that might occur due to tall buildings, shadows or terrain. Two different novel approaches for automatic road detection from gray values images are presented in this paper. In the first approach, a probabilistic model for the road pixels of a gray-scale WAS image is obtained by utilizing parallel line detection and tubularity estimation. In the second approach, these road probabilities are converted into a graph representation for local areas. These graphs are solved by using graph cut formulation which exploits min-cut, max-flow algorithm. As a result of this solution, the road mask is extracted by applying a hierarchical model that results with a transition from local to global representation. Although the methods in literature mostly utilize multi-spectral images that result wih a smaller resolution yielding surveillance of a limited region, the proposed method uses gray-scale images which enable surveillance of much wider areas. The proposed method was tested some high resolution WAS images and resulted with promising results. Ogul Can, Yeti Ziya Gurbuz, A. Aydin Alatan |
IGARSS | 3 |
| 2016 | A local extrema based method on 2D brightness temperature maps for detection of archaeological artifactsabstractArchaeological studies using computer vision based analysis methods on thermal imageries mainly lack an important stage of pointwise detection of artifact positions, which is needed for the automation of the system in a generic application. In this paper, we propose a pointwise detection method working in the thermal range of hyperspectral band for archaeological artifacts. The proposed method first optimally converts a given 3D hyperspectral image of the searched scene into a 2D brightness-temperature map by minimizing the mean square error (MSE) between the spectral radiance of a pixel and the Planck curves generated at different temperatures. The local maxima and minima are then found on the resulting 2D map as the candidate points. Finally, a score assignment is performed on the candidate points by using their temperature difference with respect to their neighborhood. The results on the thermal images taken from a test scene have indicated a good correlation between the extrema points and artifact positions. Alper Koz, Hilal Soydan, H. Sebnem Düzgün, A. Aydin Alatan |
IGARSS | 4 |
| 2016 | Comparison of 3D local and global descriptors for similarity retrieval of range data
Neslihan Bayramoglu, A. Aydin Alatan |
Neurocomputing | 2 |
| 2015 | RGBD data based pose estimation: Why sensor fusion?
Osman Serdar Gedik, A. Aydin Alatan |
FUSION | 2 |
| 2015 | Sparse recursive filtering for O(1) stereo matchingabstractRecursive edge-aware filters have been proved to be one of the most efficient approaches for cost aggregation in stereo matching. However, disparity search space dependency, as a result of full search, is the bottle-neck of these local techniques that prevent further reduction in computation. In this paper, the cost aggregation and correspondence search problems are re-formulated to enable adaptive search for each pixel during recursive operations that provides significant reduction in computational complexity. In that manner, fixed number of disparity candidates are tested for each pixel, regardless of the search space, that are aggregated through sparse recursive filtering. Hierarchical approach is exploited to pick disparity candidates for each pixel. The experimental results show that the proposed approach has linear complexity with the image size and in practice it speeds up the recursive approaches almost four times with a marginal decrease in matching accuracy. Compared to the state-of-the-art techniques, hierarchical sparse recursive aggregation is possibly the fastest approach with a competitive accuracy based on Middlebury benchmarking. Yeti Ziya Gurbuz, A. Aydin Alatan, Cevahir Çigla |
ICIP | 2 |
| 2015 | LASP: Local adaptive super-pixelsabstractIn this study, a novel gradient ascent approach is proposed for super-pixel extraction in which spectral statistics and super-pixel geometry are utilized to obtain an optimal Bayesian classifier for pixel to super-pixel label assignment. Utilization of the spectral variances and super-pixel areas reduces the dependency on user selected global parameters, while increasing robustness and adaptability. Proposed Local Adaptive Super-Pixels (LASP) approach exploits hexagonal tiling, while achieving some refinement during initialization in order to improve computation time and accuracy. The experiments conducted on Berkeley segmentation database show that LASP outperforms the existing methods in terms of boundary recall and computation time. Moreover, the proposed method provides lower bleeding error performance compared to the existing gradient ascent techniques. Kutalmis Gokalp Ince, Cevahir Çigla, A. Aydin Alatan |
ICIP | 3 |
| 2015 | Hyperspectral superpixel extraction using boundary updates based on optimal spectral similarity metricabstractThe high spectral resolution of hyperspectral images (HSI) requires a heavy processing load. Assigning each pixel to a group in the image, which is called superpixel, and processing the superpixels instead of the pixels is resorted as a means to overcome this challenge in the hyperspectral literature. In this paper, we propose an algorithm to segment a hyperspectral image into superpixels by means of iteratively updating the boundary pixels of superpixels. We first explore the optimal similarity metric for the boundary pixel updates with the contraint of keeping the superpixel boundaries aligned with the object boundaries in the image. We investigate two approaches for similarity detection between pixels during this update, first comparing the hyperspectral pixels individually, and second, comparing the pixels by using also their neigborhood. The spectral similarity metrics used for investigation are selected as spectral angle mapping (SAM) [1], spectral information divergence (SID) [2] and spatial coherence distance [3] due to their common usage. The proposed approach is compared with a pioneer state-of-the-art superpixel algorithm, SLIC [4], and its superiority is verified in terms of the superpixelization performance metrics, namely boundary recall and undersegmentation error [5]. Akin Cahskan, Alper Koz, A. Aydin Alatan |
IGARSS | 3 |
| 2015 | Joint utilization of local appearance and geometric invariants for 3D object recognition
Medeni Soysal, A. Aydin Alatan |
Multim. Tools Appl. | 2 |
| 2015 | Convexity constrained efficient superpixel and supervoxel extraction
H. Emrah Tasli, Cevahir Çigla, A. Aydin Alatan |
Signal Process. Image Commun. | 3 |
| 2014 | Occlusion-aware 3D multiple object tracker with two cameras for visual surveillanceabstractAn occlusion-aware multiple deformable object tracker for visual surveillance from two cameras is presented. Each object is tracked by a separate particle filter tracker, which is initiated upon detection of a new person and terminated when s/he leaves the scene. Objects are considered as 3D points at their centre of masses as if their mass density is uniform. Point objects and corresponding silhouette centroids in two views together with the epipolar geometry they satisfy resulted in a practical tracking methodology. An occlusion filter is described, that provides the tracker filters conditional occlusion probabilities of the objects, given their estimated positions. Advances over the previous work; in the computation of conditional occlusion probabilities, in incorporation of these probabilities in the particle filter, and in maintaining tracking of separating objects after long periods of moving close-by, are presented on PETS 2006, PETS 2009 and EPFL datasets. Osman Topcu, A. Aydin Alatan, Ali Ozer Ercan |
AVSS | 2 |
| 2014 | Uncertainty modeling for efficient visual odometry via inertial sensors on mobile devicesabstractMost of the mobile applications require efficient and precise computation of the device pose, and almost every mobile device has inertial sensors already equipped together with a camera. This fact makes sensor fusion quite attractive for increasing efficiency during pose tracking. However, the state-of-the-art fusion algorithms have a major shortcoming: lack of well-defined uncertainty introduced to the system during the prediction stage of the fusion filters. Such a drawback results in determining covariances heuristically, and hence, requirement for data-dependent tuning to achieve high performance or even convergence of these filters. In this paper, we propose an inertially-aided visual odometry system that requires neither heuristics nor parameter tuning; computation of the required uncertainties on all the estimated variables are obtained after minimum number of assumptions. Moreover, the proposed system simultaneously estimates the metric scale of the pose computed from a monocular image stream. The experimental results indicate that the proposed scale estimation outperforms the state-of-the-art methods, whereas the pose estimation step yields quite acceptable results in real-time on resource constrained systems. Yagiz Aksoy, A. Aydin Alatan |
ICIP | 2 |
| 2014 | Occlusion-aware HMM-based tracking by learningabstractRecently, an emerging class of methods, namely tracking by detection, achieved quite promising results on challenging tracking data sets. These techniques train a classifier in an online manner to separate the object from its background. These methods only take input location of the object and a random feature pool; then, a classifier bootstraps itself by using the current tracker state and extracted positive and negative samples. Following these approaches, a novel tracking system is proposed. A feature selection method is introduced to increase the discriminative power of the classifier. During tracking, a Hidden Markov Model (HMM) is utilized to filter the features that improve the performance. Moreover, a state of the proposed HMM is allocated to handle occlusions. The proposed tracker is tested on publicly available challenging video sequences and superior tracking results are achieved in real-time. Tughan Marpuc, A. Aydin Alatan |
ICIP | 2 |
| 2014 | MRF-based planar co-segmentation for depth compressionabstractAn energy based planar depth representation is proposed to obtain an efficient depth compression tool for 3DV applications. The proposed segmentation-based depth compression approach is designed by reflecting the rate-distortion tradeoff into the energy terms. A PEARL based algorithm is developed to obtain the planar approximations of depth images. Lastly depth reconstruction and novel view rendering results of the proposal compared with the state-of-the-art methods. The experiments show that the planar approach performs superior rendering results than JPEG 2000 and HEVC standards. Burak Özkalayci, A. Aydin Alatan |
ICIP | 2 |
| 2014 | Geometry-constrained spatial pyramid adaptation for image classificationabstractThis paper proposes a geometry-constrained spatial pyramid adaptation approach for the image classification task. Scene geometry is used as an input parameter for generating the spatial pyramid definitions. The resulting region adaptation is performed in accordance with the predefined geometric guidelines and underlying image characteristics. Using an approximate global geometric correspondence, exploits the idea that images of the same category share a spatial similarity. This assumption is evaluated and justified in an object classification framework, in which generated region segments are used as an enhancement to the widely utilized “spatial pyramid” method. Fixed region pyramids are replaced by the proposed locally coherent geometrically consistent region segments. Performance of the proposed method on object classification framework is evaluated on the 20 class Pascal VOC 2007 dataset. The proposed method shows consistent increase in the mean average precision (MAP) score for different experimental scenarios. H. Emrah Tasli, Ronan Sicre, Theo Gevers, A. Aydin Alatan |
ICIP | 4 |
| 2014 | E3D-D2D: Embedding in 3D, detection in 2D through projective invariantsabstractA novel watermarking method is presented in which the data embedded into a 3D model is extracted from an arbitrary 2D view by using a perspective projective invariant. The data is embedded into 3D positions of selected interest points on a 3D mesh. Determining the interest point modification vectors for ensuring watermark detection constitutes an important part of the proposed method. Different watermark embedding schemes based on optimization of the watermark function are implemented and evaluated. Another important contribution of the proposed method is selection of the interest points to ensure that they remain detectable after data embedding and rendering. A novel method to identify such repeatable interest points is also presented. Simulations are performed on random 3D point sets as well as realistic 3D models. The results indicate that data embedding in 3D and detection in 2D promises a new direction in watermarking research. Yagiz Yasaroglu, A. Aydin Alatan |
ICIP | 2 |
| 2014 | Multimodal concept detection in broadcast media: KavTan
Medeni Soysal, K. Berker Logoglu, Mashar Tekin, Ersin Esen, Ahmet Saracoglu, Banu Oskay Acar, Ezgi C. Ozan, Tugrul K. Ates, Hakan Sevimli, Ayça Müge Sevinç, Ilkay Atil, Savas Özkan, Mehmet Ali Arabaci, Seda Tankiz, Talha Karadeniz 0002, Duygu Oskay Önür, Sezin Selçuk, A. Aydin Alatan, Tolga Çiloglu |
Multim. Tools Appl. | 18 |
| 2014 | An efficient recursive edge-aware filter
Cevahir Çigla, A. Aydin Alatan |
Signal Process. Image Commun. | 2 |
| 2014 | 3D Planar Representation of Stereo Depth Images for 3DTV ApplicationsabstractThe depth modality of the multiview video plus depth (MVD) format is an active research area, whose main objective is to develop depth image based rendering friendly efficient compression methods. As a part of this research, a novel 3D planar-based depth representation is proposed. The planar approximation of multiple depth images are formulated as an energy-based co-segmentation problem by a Markov random field model. The energy terms of this problem are designed to mimic the rate-distortion tradeoff for a depth compression application. A novel algorithm is developed for practical utilization of the proposed planar approximations in stereo depth compression. The co-segmented regions are also represented as layered planar structures forming a novel single-reference MVD format. The ability of the proposed layered planar MVD representation in decoupling the texture and geometric distortions make it a promising approach. Proposed 3D planar depth compression approaches are compared against the state-of-the-art image/video coding standards by objective and visual evaluation and yielded competitive performance. Burak Özkalayci, A. Aydin Alatan |
IEEE Trans. Image Process. | 2 |
| 2014 | Efficient MRF Energy Propagation for Video Segmentation via Bilateral FiltersabstractSegmentation of an object from a video is a challenging task in multimedia applications. Depending on the application, automatic or interactive methods are desired; however, regardless of the application type, efficient computation of video object segmentation is crucial for time-critical applications; specifically, mobile and interactive applications require near real-time efficiencies. In this paper, we address the problem of video segmentation from the perspective of efficiency. We initially redefine the problem of video object segmentation as the propagation of MRF energies along the temporal domain. For this purpose, a novel and efficient method is proposed to propagate MRF energies throughout the frames via bilateral filters without using any global texture, color or shape model. Recently presented bi-exponential filter is utilized for efficiency, whereas a novel technique is also developed to dynamically solve graph-cuts for varying, non-lattice graphs in general linear filtering scenario. These improvements are experimented for both automatic and interactive video segmentation scenarios. Moreover, in addition to the efficiency, segmentation quality is also tested both quantitatively and qualitatively. Indeed, for some challenging examples, significant time efficiency is observed without loss of segmentation quality. Ozan Sener, Kemal Ugur, A. Aydin Alatan |
IEEE Trans. Multim. | 3 |
| 2013 | Fusing 2D and 3D clues for 3D tracking using visual and range data
Osman Serdar Gedik, A. Aydin Alatan |
FUSION | 2 |
| 2013 | Loosely coupled Kalman filtering for fusion of Visual Odometry and inertial navigation
Salim Sirtkaya, Burak Seymen, A. Aydin Alatan |
FUSION | 3 |
| 2013 | Recognition of 3D objects from unconstrained 2D images by using local appearance and affine geometryabstractThis paper introduces a novel method, which utilizes local appearance descriptions in a more efficient way, for 3D object recognition. Geometrically consistent local features are identified using affine 3D and 2D geometric invariants, without any reliance on partial or global planarity. Geometric invariants replace the traditional, highly constrained 2D affine transform estimation/verification step, and provides the ability to directly verify 3D geometric consistency. The accuracy and robustness of the method in highly cluttered scenes are presented in the experiments. Medeni Soysal, A. Aydin Alatan |
ICME | 2 |
| 2013 | Super pixel extraction via convexity induced boundary adaptationabstractThis study presents an efficient super-pixel extraction algorithm with major contributions to the state-of-the-art in terms of accuracy and computational complexity. Segmentation accuracy is improved through convexity constrained geodesic distance utilization; while computational efficiency is achieved by replacing complete region processing with boundary adaptation idea. Starting from the uniformly distributed rectangular equal-sized super-pixels, region boundaries are adapted to intensity edges iteratively by assigning boundary pixels to the most similar neighboring super-pixels. At each iteration, super-pixel regions are updated and hence progressively converging to compact pixel groups. Experimental results with state-of-the-art comparisons, validate the performance of the proposed technique in terms of both accuracy and speed. H. Emrah Tasli, Cevahir Çigla, Theo Gevers, A. Aydin Alatan |
ICME | 4 |
| 2013 | Information permeability for stereo matching
Cevahir Çigla, A. Aydin Alatan |
Signal Process. Image Commun. | 2 |
| 2013 | User assisted disparity remapping for stereo images
H. Emrah Tasli, A. Aydin Alatan |
Signal Process. Image Commun. | 2 |
| 2013 | 3-D Rigid Body Tracking Using Vision and Depth SensorsabstractIn robotics and augmented reality applications, model-based 3-D tracking of rigid objects is generally required. With the help of accurate pose estimates, it is required to increase reliability and decrease jitter in total. Among many solutions of pose estimation in the literature, pure vision-based 3-D trackers require either manual initializations or offline training stages. On the other hand, trackers relying on pure depth sensors are not suitable for AR applications. An automated 3-D tracking algorithm, which is based on fusion of vision and depth sensors via extended Kalman filter, is proposed in this paper. A novel measurement-tracking scheme, which is based on estimation of optical flow using intensity and shape index map data of 3-D point cloud, increases 2-D, as well as 3-D, tracking performance significantly. The proposed method requires neither manual initialization of pose nor offline training, while enabling highly accurate 3-D tracking. The accuracy of the proposed method is tested against a number of conventional techniques, and a superior performance is clearly observed in terms of both objectively via error metrics and subjectively for the rendered scenes. Osman Serdar Gedik, A. Aydin Alatan |
IEEE Trans. Cybern. | 2 |
| 2012 | Interactive 2D-3D image conversion for mobile devicesabstractWe propose a complete still image based 2D-3D mobile conversion system for touch screen use. The system consists of interactive segmentation followed by 3D rendering. The interactive segmentation is conducted dynamically by color Gaussian mixture model updates and dynamic-iterative graph-cut. A coloring gesture is used to guide the way and entertain the user during the process. Output of the image segmentation is then fed to the 3D rendering stage of the system. For rendering stage, two novel improvements are proposed to handle holes resulting from depth image based rendering process. These improvements are also expected to enhance the 3D perception. These two methods are subjectively tested and their results are presented. Yagiz Aksoy, Ozan Sener, A. Aydin Alatan, Kemal Ugur |
ICIP | 3 |
| 2012 | Optimal Data Embedding in 3D Models for Extraction from 2D Views Using Perspective Invariants
Yagiz Yasaroglu, A. Aydin Alatan |
IWDW | 2 |
| 2012 | Dominant sets based movie scene detection
Ufuk Sakarya, Ziya Telatar, A. Aydin Alatan |
Signal Process. | 3 |
| 2011 | Joint Utilization of Appearance, Geometry and Chance for Scene Logo RetrievalabstractA novel approach involving the comparison of appearance and geometrical similarity of local patterns via a combined description is presented. Candidate groups of interest points are identified based on unlikeliness of being matched by chance. For each of the keypoints in these groups, a novel description is proposed. This description utilizes quantized appearance descriptors of interest points to avoid the necessity of matching each test descriptor to each template descriptor. Additionally, one-to-many matching is possible in contrast to its counterparts in the literature. Geometrical descriptions are based on multiple small groups of points, namely quads, in barycentric coordinates, instead of a single large group that is susceptible to partial transformations. These advantages render the proposed algorithm robust to significant appearance changes, especially due to affine transformations, while being resistant to random false matches through simultaneous utilization of geometrical part of the descriptor. This generic, robust template matching technique is evaluated in an application of scene logo retrieval. Medeni Soysal, A. Aydin Alatan |
Comput. J. | 2 |
| 2011 | Robust Video Data Hiding Using Forbidden Zone Data Hiding and Selective EmbeddingabstractVideo data hiding is still an important research topic due to the design complexities involved. We propose a new video data hiding method that makes use of erasure correction capability of repeat accumulate codes and superiority of forbidden zone data hiding. Selective embedding is utilized in the proposed method to determine host signal samples suitable for data hiding. This method also contains a temporal synchronization scheme in order to withstand frame drop and insert attacks. The proposed framework is tested by typical broadcast material against MPEG-2, H.264 compression, frame-rate conversion attacks, as well as other well-known video data hiding methods. The decoding error values are reported for typical system parameters. The simulation results indicate that the framework can be successfully utilized in video data hiding applications. Ersin Esen, A. Aydin Alatan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | A novel shadow restoration algorithm based on atmospheric effects for aerial imagesabstractIn aerial images, the performance of the segmentation and object recognition algorithms could suffer due to shadows in the scene. This effort describes a novel shadow restoration algorithm based on atmospheric effects and characteristics of sun light for aerial images. Firstly, shadow regions are detected exploiting the Rayleigh scattering phenomena and the well-known fact related to the low illumination intensity in the shadow regions. After detection, shadow restoration is achieved by first restoring partially occluded shadow areas, as a result of modeling these transition regions with a continuous function that considers shadow formations. Next, fully occluded shadow regions are restored by first segmenting the image into multiple uniformly illuminated regions, then multiplying the intensity values in these regions with a constant, which is determined by the ratio of intensities between each segment and its non-shadow neighborhood. The simulation results indicate improvements over similar work from the literature. Çaglar Aytekin, A. Aydin Alatan |
ICIP | 2 |
| 2010 | Efficient graph-based image segmentation via speeded-up turbo pixelsabstractAn efficient graph based image segmentation algorithm exploiting a novel and fast turbo pixel extraction method is introduced. The images are modeled as weighted graphs whose nodes correspond to super pixels; and normalized cuts are utilized to obtain final segmentation. Utilizing super pixels provides an efficient and compact representation; the graph complexity decreases by hundreds in terms of node number. Connected K-means with convexity constraint is the key tool for the proposed super pixel extraction. Once the pixels are grouped into super pixels, iterative bi-partitioning of the weighted graph, as introduced in normalized cuts, is performed to obtain segmentation map. Supported by various experiments, the proposed two stage segmentation scheme can be considered to be one of the most efficient graph based segmentation algorithms providing high quality results. Cevahir Çigla, A. Aydin Alatan |
ICIP | 2 |
| 2010 | Frame-rate conversion for multiview video exploiting 3D motion modelsabstractA frame-rate conversion (FRC) scheme for increasing the frame-rate of multiview video for reduction of motion blur in hold-type displays is proposed. In order to obtain high quality inter-frames, the proposed method utilizes 3D motion models relying on the 3D scene information extractable from multiview video. First of all, independently moving objects (IMOs) are segmented by using a depth-based object segmentation method. Then, interest points on IMOs are obtained via scale invariant feature transform (SIFT). Estimating 3D rotation and translation matrices of moving rigid objects by the help of SIFT features between successive frames; 3D scene is reconstructed at any desired time instant. Inter-frames for the desired view of the multiview set are rendered by projecting the reconstructed 3D scene on this view. Simulation results reveal that, by utilizing true 3D object trajectory, the proposed method interpolates frames with increased quality compared to some conventional techniques. Osman Serdar Gedik, A. Aydin Alatan |
ICIP | 2 |
| 2010 | Multi-resolution motion estimation for motion compensated frame interpolationabstractA multi-resolution motion estimation scheme is proposed for tracking of the true 2D motion in video sequences for motion compensated image interpolation. The proposed algorithm utilizes frames with different resolutions and adaptive block dimensions for efficient representation of motion. Firstly, motion vectors for each block are assigned as a result of predictive search in each pass. Then, the outlier motion vectors are detected and corrected at the end of each pass. Simulation results with respect to different quality metrics reveal that the motion fields generated by the proposed algorithm is of high quality, when compared with state-of-the-art motion estimation algorithms in literature. Bertan Günyel, A. Aydin Alatan |
ICIP | 2 |
| 2010 | Shape Index SIFT: Range Image Recognition Using Local FeaturesabstractRange image recognition gains importance in the recent years due to the developments in acquiring, displaying, and storing such data. In this paper, we present a novel method for matching range surfaces. Our method utilizes local surface properties and represents the geometry of local regions efficiently. Integrating the Scale Invariant Feature Transform (SIFT) with the shape index (SI) representation of the range images allows matching of surfaces with different scales and orientations. We apply the method for scaled, rotated, and occluded range images and demonstrate the effectiveness it by comparing the previous studies. Neslihan Bayramoglu, A. Aydin Alatan |
ICPR | 2 |
| 2010 | Watermarking of Free-view VideoabstractWith the advances in image based rendering (IBR) in recent years, generation of a realistic arbitrary view of a scene from a number of original views has become cheaper and faster. One of the main applications of this progress has emerged as free-view TV(FTV), where TV-viewers select freely the viewing position and angle via IBR on the transmitted multiview video. Noting that the TV-viewer might record a personal video for this arbitrarily selected view and misuse this content, it is apparent that copyright and copy protection problems also exist and should be solved for FTV. In this paper, we focus on this newly emerged problem by proposing a watermarking method for free-view video. The watermark is embedded into every frame of multiple views by exploiting the spatial masking properties of the human visual system. Assuming that the position and rotation of the virtual camera is known, the proposed method extracts the watermark successfully from an arbitrarily generated virtual image. In order to extend the method for the case of an unknown virtual camera position and rotation, the transformations on the watermark pattern due to image based rendering operations are analyzed. Based upon this analysis, camera position and homography estimation methods are proposed for the virtual camera. The encouraging simulation results promise not only a novel method, but also a new direction for watermarking research. Alper Koz, Cevahir Çigla, A. Aydin Alatan |
IEEE Trans. Image Process. | 3 |
| 2009 | Watermarking for depth-image-based renderingabstractIn this paper, a novel watermarking method for depth-image-based rendering methods employed in free viewpoint television systems is proposed. The proposed method uses a correlation-based approach; where a watermark pattern is warped for each different view of the multi-view source video, and embedded to the texture maps of those views in spatial domain. First, the projection matrix of the 2D rendered image is estimated; then, using this matrix, a rendered watermark is formed by warping the original watermark pattern. Finally, symmetric phase only matched filtering is used to find the correlation between the rendered image and the rendered watermark. Extended simulations with various attacks show the applicability of the proposed solution. Eren Halici, A. Aydin Alatan |
ICIP | 2 |
| 2009 | Region-based motion-compensated frame rate up-conversion by homography parameter interpolationabstractA new region-based frame interpolation algorithm is proposed based on the segmented motion layers with planar perspective motion models. It is shown that performing the motion model interpolation in the homography parameter space is equivalent to interpolating the parameters of the real camera motion, which requires decomposition of the homography matrix, under several practically reasonable assumptions. Based on this reasoning, backward and forward motion models from the interpolation frame(s) to the original frames of the sequence are estimated for each motion layer. Layer support maps at the point(s) of interpolation in time are generated by using these interpolated motion models and the layer maps of the neighboring original frames. Finally, pixel intensities of the interpolation frame are determined by a series of conditions on the layer occlusion relations and the intensities at the transformed locations of the original frames. Experimental results show that the proposed algorithm achieves visually pleasing results without blur or halo effects on dynamic scenes with complex motion. Engin Türetken, A. Aydin Alatan |
ICIP | 2 |
| 2009 | Special issue on advances in three-dimensional television and video: Guest editorial
Ugur Güdükbay, A. Aydin Alatan |
Signal Process. Image Commun. | 2 |
| 2009 | Rate-Distortion Efficient Piecewise Planar 3-D Scene Representation From 2-D ImagesabstractIn any practical application of the 2-D-to-3-D conversion that involves storage and transmission, representation efficiency has an undisputable importance that is not reflected in the attention the topic received. In order to address this problem, a novel algorithm, which yields efficient 3-D representations in the rate distortion sense, is proposed. The algorithm utilizes two views of a scene to build a mesh-based representation incrementally, via adding new vertices, while minimizing a distortion measure. The experimental results indicate that, in scenes that can be approximated by planes, the proposed algorithm is superior to the dense depth map and, in some practical situations, to the block motion vector-based representations in the rate-distortion sense. Evren Imre, A. Aydin Alatan, Ugur Güdükbay |
IEEE Trans. Image Process. | 2 |
| 2008 | Region-based image segmentation via graph cutsabstractA graph theoretic color image segmentation algorithm is proposed, in which the popular normalized cuts image segmentation method is improved with modifications on its graph structure. The image is represented by a weighted undirected graph, whose nodes correspond to over-segmented regions, instead of pixels, that decreases the complexity of the overall algorithm. In addition, the link weights between the nodes are calculated through the intensity similarities of the neighboring regions. The irregular distribution of the nodes, as a result of such a modification, causes a bias towards combining regions with high number of links. This bias is removed by limiting the number of links for each node. Finally, segmentation is achieved by bipartitioning the graph recursively according to the minimization of the normalized cut measure. The simulation results indicate that the proposed segmentation scheme performs quite faster than the traditional normalized cut methods, as well as yielding better segmentation results due to its region-based representation. Cevahir Çigla, A. Aydin Alatan |
ICIP | 2 |
| 2008 | Segmentation in multi-view video via color, depth and motion cuesabstractIn the light of dense depth map estimation, motion estimation and object segmentation, the research on multi-view video (MW) content has becoming increasingly popular due to its wide application areas in the near future. In this work, object segmentation problem is studied by additional cues due to depth and motion fields. Segmentation is achieved by modeling images as graphical models and performing popular normalized cuts method with some modifications. In the graphical models, each node is represented by a group of pixels, instead of individual pixels, which are obtained as a result of over-segmentation of the images. These over-segmented regions are also utilized in the dense depth map estimation step; in which 3D planar models are assigned for each of these sub-regions. Moreover, optical flow is estimated based on afline motion assumption for these regions. The links of the graphical models are weighted according to the depth, motion and color similarities of the pixel groups due to these regions. Once the links are obtained, segmentation is achieved by recursively bi-partitioning the graph via removing the weak links. Experiments indicate that the proposed framework achieves precise segmentation results for MVV sequences. Cevahir Çigla, A. Aydin Alatan |
ICIP | 2 |
| 2008 | Watermarking for Image Based Rendering via homography-based virtual camera location estimationabstractThe recent advances in Image Based Rendering (IBR) have pioneered freely determining the viewing position and angle in a scene from multi-view video. Noting that a person could also record a personal video for this arbitrarily selected view and misuse this content, apparently, copyright and copy protection problems also exist and should be solved for IBR applications. In our recent work [1], we have proposed a watermarking method, which embeds the watermark pattern into every frame of multi-view video and extracts this watermark from a rendered image, generated by the nearest-interpolation based light-field rendering (LFR). This paper presents a novel solution for the challenging problem of watermark detection in bilinear interpolation, namely the most attractive and promising interpolation method for LFR-based applications. Moreover, the location of the virtual camera could be completely arbitrary in this new formulation. The results show that the watermark could be extracted successfully for LFR via bilinear interpolation for any imagery camera location and rotation, as long as the visual quality of the rendered image is preserved. Alper Koz, Cevahir Çigla, A. Aydin Alatan |
ICIP | 3 |
| 2008 | Error detection and concealment for video transmission using information hiding
Ayhan Yilmaz, A. Aydin Alatan |
Signal Process. Image Commun. | 2 |
| 2008 | Oblivious Spatio-Temporal Watermarking of Digital Video by Exploiting the Human Visual SystemabstractImperceptibility requirement in video watermarking is more challenging compared with its image counterpart due to the additional dimension existing in video. The embedding system should not only yield spatially invisible watermarks for each frame of the video, but it should also take the temporal dimension into account in order to avoid any flicker distortion between frames. While some of the methods in the literature approach this problem by only allowing arbitrarily small modifications within frames in different transform domains, some others simply use implicit spatial properties of the human visual system (HVS), such as luminance masking, spatial masking, and contrast masking. In addition, some approaches exploit explicitly the spatial thresholds of HVS to determine the location and strength of the watermark. However, none of the former approaches have focused on guaranteeing temporal invisibility and achieving maximum watermark strength along the temporal direction. In this paper, temporal dimension is exploited for video watermarking by means of utilizing temporal sensitivity of the HVS. The proposed method utilizes the temporal contrast thresholds of HVS to determine the maximum strength of watermark, which still gives imperceptible distortion after watermark insertion. Compared with some recognized methods in the literature, the proposed method avoids the typical visual degradations in the watermarked video, while still giving much better robustness against common video distortions, such as additive Gaussian noise, video coding, frame rate conversions, and temporal shifts, in terms of bit error rate. Alper Koz, A. Aydin Alatan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | Region-Based Dense Depth Extraction from Multi-View VideoabstractA novel multi-view region-based dense depth map estimation problem is presented, based on a modified plane-sweeping strategy. In this approach, the whole scene is assumed to be region-wise planar. These planar regions are defined by back-projections of the over-segmented homogenous color regions on the images and the plane parameters are determined by angle-sweeping at different depth levels. The position and rotation of the plane patches are estimated robustly by minimizing a segment-based cost function, which considers occlusions, as well. The quality of depth map estimates is measured via reconstruction quality of the conjugate views, after warping segments into these views by the resulting homographies. Finally, a greedy-search algorithm is applied to refine the reconstruction quality and update the plane equations with visibility constraint. Based on the simulation results, it is observed that the proposed algorithm handles large un-textured regions, depth discontinuities at object boundaries, slanted surfaces, as well as occlusions. Cevahir Çigla, Xenophon Zabulis, A. Aydin Alatan |
ICIP (5) | 3 |
| 2007 | Rate-Distortion Based Piecewise Planar 3D Scene Geometry RepresentationabstractThis paper proposes a novel 3D piecewise planar reconstruction algorithm, to build a 3D scene representation that minimizes the intensity error between a particular frame and its prediction. 3D scene geometry is exploited to remove the visual redundancy between frame pairs for any predictive coding scheme. This approach associates the rate increase with the quality of representation, and is shown to be rate-distortion efficient by the experiments. Evren Imre, A. Aydin Alatan, Ugur Güdükbay |
ICIP (5) | 2 |
| 2007 | Towards 3-D scene reconstruction from broadcast video
Evren Imre, Sebastian Knorr, Burak Özkalayci, Ugur Topay, A. Aydin Alatan, Thomas Sikora |
Signal Process. Image Commun. | 5 |
| 2007 | Scene Representation Technologies for 3DTV - A Surveyabstract3-D scene representation is utilized during scene extraction, modeling, transmission and display stages of a 3DTV framework. To this end, different representation technologies are proposed to fulfill the requirements of 3DTV paradigm. Dense point-based methods are appropriate for free-view 3DTV applications, since they can generate novel views easily. As surface representations, polygonal meshes are quite popular due to their generality and current hardware support. Unfortunately, there is no inherent smoothness in their description and the resulting renderings may contain unrealistic artifacts. NURBS surfaces have embedded smoothness and efficient tools for editing and animation, but they are more suitable for synthetic content. Smooth subdivision surfaces, which offer a good compromise between polygonal meshes and NURBS surfaces, require sophisticated geometry modeling tools and are usually difficult to obtain. One recent trend in surface representation is point-based modeling which can meet most of the requirements of 3DTV, however the relevant state-of-the-art is not yet mature enough. On the other hand, volumetric representations encapsulate neighborhood information that is useful for the reconstruction of surfaces with their parallel implementations for multiview stereo algorithms. Apart from the representation of 3-D structure by different primitives, texturing of scenes is also essential for a realistic scene rendering. Image-based rendering techniques directly render novel views of a scene from the acquired images, since they do not require any explicit geometry or texture representation. 3-D human face and body modeling facilitate the realistic animation and rendering of human figures that is quite crucial for 3DTV that might demand real-time animation of human bodies. Physically based modeling and animation techniques produce impressive results, thus have potential for use in a 3DTV framework for modeling and animating dynamic scenes. As a concluding remark, it can be argued that 3-D scene and texture representation techniques are mature enough to serve and fulfill the requirements of 3-D extraction, transmission and display sides in a 3DTV scenario. A. Aydin Alatan, Yücel Yemez, Ugur Güdükbay, Xenophon Zabulis, Karsten Müller 0001, Çigdem Eroglu Erdem, C. Weigel, Aljoscha Smolic |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | 3-D Time-Varying Scene Capture Technologies - A SurveyabstractAdvances in image sensors and evolution of digital computation is a strong stimulus for development and implementation of sophisticated methods for capturing, processing and analysis of 3D data from dynamic scenes. Research on perspective time-varying 3D scene capture technologies is important for the upcoming 3DTV displays. Methods such as shape-from-texture, shape-from-shading, shape-from-focus, and shape-from-motion extraction can restore 3D shape information from a single camera data. The existing techniques for 3D extraction from single-camera video sequences are especially useful for conversion of the already available vast mono-view content to the 3DTV systems. Scene-oriented single-camera methods such as human face reconstruction and facial motion analysis, body modeling and body motion tracking, and motion recognition solve efficiently a variety of tasks. 3D multicamera dynamic acquisition and reconstruction, their hardware specifics including calibration and synchronization and software demands form another area of intensive research. Different classes of multiview stereo algorithms such as those based on cost function computing and optimization, fusing of multiple views, and feature-point reconstruction are possible candidates for dynamic 3D reconstruction. High-resolution digital holography and pattern projection techniques such as coded light or fringe projection for real-time extraction of 3D object positions and color information could manifest themselves as an alternative to traditional camera-based methods. Apart from all of these approaches, there also are some active imaging devices capable of 3D extraction such as the 3D time-of-flight camera, which provides 3D image data of its environment by means of a modulated infrared light source. Elena Stoykova, A. Aydin Alatan, Philip W. Benzie, Nikolaos Grammalidis, Sotiris Malassiotis, Jörn Ostermann, S. Piekh, Ventseslav Sainov, Christian Theobalt, T. Thevar, Xenophon Zabulis |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2006 | Forbidden Zone Data HidingabstractA new blind data hiding method is proposed based on a novel concept of forbidden zone, where no alteration is allowed in a host signal during message embedding step. Depending on the desired probability of error, the range of the forbidden zone varies, as a compromise between robustness and embedding distortion. Hence, the proposed method makes use of this zone through a single control parameter in conjunction with modulating quantizers for message embedding. The superiority of the proposed scheme over QIM is presented theoretically, as well empirically via simulations. The method is further compared with DC-QIM and its weaknesses and strengths are stated. Ersin Esen, A. Aydin Alatan |
ICIP | 2 |
| 2006 | Prioritized Sequential 3D Reconstruction in Video Sequences with Multiple MotionsabstractIn this study, an algorithm is proposed to solve the multi-frame structure from motion (MFSfM) problem for monocular video sequences in dynamic scenes. The algorithm uses the epipolar criterion to segment the features belonging to independently moving objects. Once the features are segmented, corresponding objects are reconstructed individually by using a sequential algorithm, which is also capable of prioritizing the frame pairs with respect to their reliability and information content, thus achieving a fast and accurate reconstruction through efficient processing of the available data. A tracker is utilized to increase the baseline distance between views and to improve the F-matrix estimation, which is beneficial to both the segmentation and the 3D structure estimation processes. The experimental results demonstrate that our approach has the potential to effectively deal with the multi-body MFSfM problem in a generic video sequence. Evren Imre, Sebastian Knorr, A. Aydin Alatan, Thomas Sikora |
ICIP | 3 |
| 2006 | Free-View Watermarking for Free-View TelevisionabstractThe recent advances in Image Based Rendering (IBR) has pioneered a new technology, free-view television, in which TV-viewers select freely the viewing position and angle by the application of IBR on the transmitted multi-view video. Noting that the TV-viewer might also record a personal video for this arbitrarily selected view and misuse this content, it is apparent that copyright and copy protection problems also exist and should be solved for free-view TV. In this paper, we focus on this problem by proposing a watermarking method for free-view video. The watermark is embedded into every frame of multiple views by exploiting the spatial masking properties of the Human Visual System (HVS). Assuming that the position and rotation for the imagery view is known, the proposed method extracts the watermark successfully from an arbitrarily generated image. In order to extend the method for the case of an unknown imagery camera position and rotation, the modifications on the watermark pattern due to image based rendering operations are also analyzed. Based on this analysis, a camera position and homography estimation method is proposed considering the operations in image based rendering. The results show that the watermark detection is achieved successfully for the cases in which the imagery camera is arbitrarily located on the camera plane. Alper Koz, Cevahir Çigla, A. Aydin Alatan |
ICIP | 3 |
| 2006 | Efficient Bayesian Track-Before-DetectabstractThis paper presents a novel Bayesian recursive track-before-detect (TBD) algorithm for detection and tracking of dim targets in optical image sequences. The algorithm eliminates the need for storing past observations by recursively incorporating new data acquired through sensor to the existing information. It calculates the likelihood ratio for optimal detection and estimates target state simultaneously. The technique does not require velocity-matched filtering and hence, it is capable of detecting any target moving in any direction. The algorithm is tested with both synthetic and real video sequences, and is shown to be capable of performing sufficiently well for very low signal-to-noise ratio situations. Serhat Tekinalp, A. Aydin Alatan |
ICIP | 2 |
| 2005 | Oblivious video watermarking using temporal sensitivity of HVSabstractA novel oblivious video watermarking technique based on temporal sensitivity of human visual system (HVS) is proposed. The method exploits the temporal contrast thresholds of HVS to determine the maximum strength of watermark, which still gives imperceptible distortion after watermark insertion. Compared to the other methods in the literature, which do not use any HVS properties explicitly [F. Deguillaume et al 1999] or exploits only spatial properties of HVS [F. Hartung et al, 1998], the proposed method guarantees to avoid flickering problem in the watermarked video and gives better robustness results to video distortions, such as additive Gaussian noise, ITU H.263+ coding at medium bit rates and frame averaging in terms of bit error rate. Alper Koz, A. Aydin Alatan |
ICIP (1) | 2 |
| 2005 | Video Adaptation for Transmission Channels by Utility ModelingabstractThe satisfaction a user gets from watching a video in a resource limited device, can be formulated by Utility Theory. The resulting video adaptation is optimal in the sense that the adapted video maximizes the user satisfaction, which is modeled through subjective tests comprising of 3 independent utility components : crispness, motion smoothness and content visibility. These components are maximized in terms of coding parameters by obtaining a Pareto optimal set. In this manuscript, inclusion of transmission channel capacity into the subjective utility model of user satisfaction is addressed. It is proposed that using the maximum channel capacity as a restriction metric, certain members of the Pareto optimal solution set can be eliminated such that the remaining members are suitable for transmission through the given channel. Once the reduced solution set is obtained, an additional figure of merit can be used to pick a single solution from this set, depending on the application scenario. Özgür Deniz Önür, A. Aydin Alatan |
ICME | 2 |
| 2004 | Data hiding using trellis coded quantizationabstractInformation theoretic tools lead to the design and analysis of new blind data hiding methods. A novel quantization-based blind method, which uses trellis coded quantization, is proposed in this manuscript. The redundancy in initial state selection during trellis coded quantization is exploited to hide information as the index of this initial state. This index is recovered at the receiver by Viterbi decoding after comparison with all initial states. The performance of the proposed method is compared against other well-known approaches via simulations and promising results are obtained. Based on these results, the proposed method can be preferred in certain applications with high distortion attacks. Ersin Esen, A. Aydin Alatan |
ICIP | 2 |
| 2003 | Utilization of texture, contrast and color homogeneity for detecting and recognizing text from video framesabstractIt is possible to index and manage large video archives in a more efficient manner by detecting and recognizing text within video frames. There are some inherent properties of videotext, such as distinguishing texture, higher contrast against background, and uniform color, making it detectable. By employing these properties, it is possible to detect text regions and binarize the image for character recognition, in this paper, a complete framework for detection and recognition of videotext is presented. The results from Gabor-based texture analysis, contrast-based segmentation and color homogeneity are merged to obtain minimum number of candidate regions before binarization. The performance of the system is tested for its recognition rate for various combinations and it is observed that the results give recognition rates, reasonable for most practical purposes. Serhat Tekinalp, A. Aydin Alatan |
ICIP (3) | 2 |
| 2003 | Error concealment of video sequences by data hidingabstractA complete error resilient video transmission codec is proposed, utilizing imperceptible embedded information for combined detecting, resynchronization and reconstruction of the errors and lost data. Utilization of data hiding for this problem provides a reserve information about the video to the receiver while unchanging the transmitted bit-stream syntax; hence, improves the reconstruction video quality without significant extra channel utilization. A spatial domain error recovery technique, which hides edge orientation information of a block, and a resynchronization technique, which embeds bit-length of a block into other blocks are combined, as well as some parity information about the hidden data, to conceal channel errors on intra-coded frames of a video sequence. The inter-coded frames are basically recovered by hiding motion vector information into the next frames. The simulation results show that the proposed approach performs superior to conventional approaches for concealing the errors in binary symmetric channels, especially for higher bit-rates and error-rates. A. Aydin Alatan |
ICIP (2) | 2 |
| 2002 | Foveated image watermarkingabstractThe spatial resolution of the human visual system (HVS) decreases rapidly away from the point of fixation (foveation point). By exploiting this fact, we propose a watermarking approach that embeds the watermark energy into the image periphery according to foveation-based HVS contrast thresholds. Compared to other HVS-based watermarking methods, the simulation results demonstrate an improvement in the robustness of the proposed approach against image degradations, such as JPEG compression, cropping and additive Gaussian noise, in terms of subjective measures, based on foveation. In addition, the method proposed for still images is adapted for video and the robustness of the adapted method is tested against ITU H.263+ coding. Alper Koz, A. Aydin Alatan |
ICIP (3) | 2 |
| 2001 | Automatic multi-modal dialogue scene indexingabstractAn automatic algorithm for indexing dialogue scenes in multimedia content is proposed. The content is segmented into dialogue scenes using the state transitions of a hidden Markov model (HMM). Each shot is classified using both audio and visual information to determine the state/scene transitions for this model. Face detection and silence/speech/music classification are the basic tools which are utilized to index the scenes. While face information is extracted after applying some heuristics to skin-colored regions, audio analysis is achieved by examining signal energy, periodicity and zero crossing rate (ZCR) of the audio waveform. The simulation results show the possibility of automatically indexing the dialogues using the proposed algorithm. A. Aydin Alatan |
ICIP (3) | 1 |
| 2001 | Multi-Modal Dialog Scene Detection Using Hidden Markov Models for Content-Based Multimedia Indexing
A. Aydin Alatan, Ali N. Akansu, Marilyn Wolf |
Multim. Tools Appl. | 1 |
| 2000 | A New Method for Optimal Rate Allocation for Progressive Image Transmission over Noisy ChannelsabstractWe present a new method for optimal rate allocation between a progressive image coder and a channel coder for noisy channels. A mathematical model for the embedded bit streams is developed and used as a new metric for such a joint source-channel coding optimization problem. It is further shown that maximization of such a new metric is equivalent to the maximization of a conventional PSNR measure. PSNR improvements up to 0.3 dB over very noisy channels at low bit rate are achieved by using the proposed method without any additional overheads. Furthermore, the PSNR performance obtained by using the proposed method is upper bounded by those of the conventional Monte Carlo method. Computational results and numerical comparisons consistently verify the merit of the proposed technique. Minyi Zhao, A. Aydin Alatan, Ali N. Akansu |
Data Compression Conference | 2 |
| 2000 | Comparative analysis of hidden Markov models for multi-modal dialogue scene indexingabstractA class of audio-visual content is segmented into dialogue scenes using the state transitions of a novel hidden Markov model (HMM). Each shot is classified using both the audio track and the visual content to determine the state/scene transitions of the model. After simulations with circular and left-to-right HMM topologies, it is observed that both performing very well with multi-modal inputs. Moreover, for the circular topology, the comparisons between different training and observation sets show that audio and face information together gives the most consistent results among different observation sets. A. Aydin Alatan, Ali N. Akansu, Marilyn Wolf |
ICASSP | 1 |
| 2000 | Dynamic UEP of embedded image bit streams over noisy channelsabstractWe present a dynamic unequal error protection (UEP) framework for the embedded image bit streams over noisy channels. A general source model for M-rate UEP schemes is derived and an optimization algorithm for 2-rate dynamic UEP schemes is developed. To facilitate the design and implementation of UEP schemes, we also derived a necessary condition and an upper bound for UEP gains. Simulation results demonstrate about 0.3 dB improvements over equal error protection (EEP) schemes at the price of negligible overheads while the progressiveness of the original bit streams is still kept. Minyi Zhao, A. Aydin Alatan, Ali N. Akansu |
ICASSP | 2 |
| 2000 | Compressed Domain MPEG-2 Video Editing with VBV RequirementabstractA novel method is proposed to achieve efficient MPEG-2 video editing in compressed domain while preserving video buffer verifier (VBV) requirements. Different cases are determined, according to the VBV modes of the bit-streams to be concatenated. For each case, the minimum number of zero-stuffing bits or shortest waiting time between two streams is determined analytically, so that the resulting bit-stream is still VBV-compliant. The simulation results show that the proposed method is applicable to any MPEG-2 bit-stream independent of its encoder. Ren Egawa, A. Aydin Alatan, Ali N. Akansu |
ICIP | 2 |
| 2000 | Unequal error protection of SPIHT encoded image bit streamsabstractA derivative of the set partitioning into hierarchical trees (SPIHT) image coding method, which generates substreams with different error-resilience properties, is proposed. By dividing the image bit stream into three classes, substreams with different immunity properties are obtained. The unequal protection of these substreams with different channel coding rates improves the overall performance of the method against channel errors. Simulation results show the superiority of the proposed method over some of the state-of-the-art methods. A. Aydin Alatan, Minyi Zhao, Ali N. Akansu |
IEEE J. Sel. Areas Commun. | 1 |
| 1999 | On the choice of transforms for data hiding in compressed videoabstractWe present an information-theoretic approach to obtain an estimate of the number of bits that can be hidden in compressed image sequences. We show how addition of the message signal in a suitable transform domain rather than the spatial domain can significantly increase the data hiding capacity. We compare the data hiding capacities achievable with different block transforms and show that the choice of the transform should depend on the robustness needed. While it is better to choose transforms with good energy compaction property (like DCT, wavelet etc.) when the robustness required is low, transforms with poorer energy compaction property (like the Hadamard or Hartley transform) are preferable choices for higher robustness requirements. Mahalingam Ramkumar, Ali N. Akansu, A. Aydin Alatan |
ICASSP | 3 |
| 1999 | A Robust Data Hiding Scheme for Images Using DFTabstractWe present a data hiding scheme for still images, in which only the magnitude of the DFT coefficients are altered to embed the hidden information bits. We begin by identifying various components of a typical data hiding scheme, like the decomposition used, signature design, self-noise suppression and signaling. The final choice of components for the proposed data hiding scheme are well tailored by utilizing theoretical reasonings and experimental observations. Performance results show good robustness of the data hiding scheme for both JPEG and SPIHT image compression. Mahalingam Ramkumar, Ali N. Akansu, A. Aydin Alatan |
ICIP (2) | 3 |
| 1998 | Joint Utilization of Fixed and Variable-Length Codes for Improving Synchronization Immunity for Image TransmissionabstractRobust transmission of images is achieved by using fixed and variable-length coding together without much loss in compression efficiency. The probability distribution function of a DCT coefficient can be divided into two regions using a threshold, so that one portion contains roughly equiprobable transform coefficients. While fixed-length coding, which is a powerful solution to the synchronization problem, is used in this inner equiprobable region without sacrificing compression, the outer (saturating) region is reserved for variable-length codes. The proposed image coder first encodes the bit allocated DCT coefficients using a fixed-rate quantizer bank, then the saturated values for these coefficients are encoded using an entropy constrained scalar quantizer, followed by an arithmetic encoder. Our simulations show that the proposed encoder is appropriate for applications in which an acceptable quality must always be maintained in any channel condition. A. Aydin Alatan, John W. Woods |
ICIP (1) | 1 |
| 1998 | 3-D motion estimation of rigid objects for video coding applications using an improved iterative version of the E-matrix methodabstractAs an alternative to current two-dimensional (2-D) motion models, a robust three-dimensional (3-D) motion estimation method is proposed to be utilized in object-based video coding applications. Since the popular E-matrix method is well known for its susceptibility to input errors, a performance indicator, which tests the validity of the estimated 3-D motion parameters both explicitly and implicitly, is defined. This indicator is utilized within the RANSAC method to obtain a robust set of 2-D motion correspondences which leads to better 3-D motion parameters for each object. The experimental results support the superiority of the proposed method over direct application of the E-matrix method. A. Aydin Alatan, Levent Onural |
IEEE Signal Process. Lett. | 1 |
| 1998 | Image sequence analysis for emerging interactive multimedia services-the European COST 211 frameworkabstractFlexibility and efficiency of coding, content extraction, and content-based search are key research topics in the field of interactive multimedia. Ongoing ISO MPEG-4 and MPEG-7 activities are targeting standardization to facilitate such services. European COST Telecommunications activities provide a framework for research collaboration. At present a significant effort of the COST 211/sup ter/ group activities is dedicated toward image and video sequence analysis and segmentation-an important technological aspect for the success of emerging object-based MPEG-4 and MPEG-7 multimedia applications. The current work of COST 211 is centered around the test model, called the analysis model (AM). The essential feature of the AM is its ability to fuse information from different sources to achieve a high-quality object segmentation. The current information sources are the intermediate results from frame-based (still) color segmentation, motion vector based segmentation, and change-detection-based segmentation. Motion vectors, which form the basis for the motion vector based intermediate segmentation, are estimated from consecutive frames. A recursive shortest spanning tree (RSST) algorithm is used to obtain intermediate color and motion vector based segmentation results. A rule-based region processor fuses the intermediate results; a postprocessor further refines the final segmentation output. The results of the current AM are satisfactory. A. Aydin Alatan, Levent Onural, Michael Wollborn, Roland Mech, Ertem Tuncel, Thomas Sikora |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1998 | Estimation of depth fields suitable for video compression based on 3-D structure and motion of objectsabstractIntensity prediction along motion trajectories removes temporal redundancy considerably in video compression algorithms. In three-dimensional (3-D) object-based video coding, both 3-D motion and depth values are required for temporal prediction. The required 3-D motion parameters for each object are found by the correspondence-based E-matrix method. The estimation of the correspondences-two-dimensional (2-D) motion field-between the frames and segmentation of the scene into objects are achieved simultaneously by minimizing a Gibbs energy. The depth field is estimated by jointly minimizing a defined distortion and bit-rate criterion using the 3-D motion parameters. The resulting depth field is efficient in the rate-distortion sense. Bit-rate values corresponding to the lossless encoding of the resultant depth fields are obtained using predictive coding; prediction errors are encoded by a Lempel-Ziv algorithm. The results are satisfactory for real-life video scenes. A. Aydin Alatan, Levent Onural |
IEEE Trans. Image Process. | 1 |
| 1997 | A Rule-Based Method for Object Segmentation in Video SequencesabstractObject segmentation and tracking are problems within the scope of MPEG-4 and MPEG-7 standardization activities. A novel algorithm for both object segmentation and tracking is presented. The algorithm fuses motion, color, and accumulated previous segmentation data at 'region level', in contrast to conventional 'pixel level' approaches. The information fusion is achieved by a rule-based region processing unit which intelligently utilizes the motion information to locate the objects in the scene, the color information to extract the true boundaries, and the segmentation result of the previous frame for tracking the objects. The algorithm is generic in the sense that the modules prior to the rule-based region processor can independently be replaced by alternative units which can achieve the same tasks. In the proposed algorithm, while the recursive-shortest-spanning-tree (RSST) algorithm is used for segmentation purposes, hierarchical-block-matching (HBM) is utilized for estimating motion between frames. The simulation results are very promising for this novel object segmentation approach. A. Aydin Alatan, Ertem Tuncel, Levent Onural |
ICIP (2) | 1 |
| 1996 | Joint estimation and optimum encoding of depth field for 3-D object-based video codingabstract3-D motion models can be used to remove temporal redundancy between image frames. For efficient encoding using 3-D motion information, apart from the 3-D motion parameters, a dense depth field must also be encoded to achieve 2-D motion compensation on the image plane. Inspired by rate-distortion theory, a novel method is proposed to optimally encode the dense depth fields of the moving objects in the scene. Using two intensity frames and 3-D motion parameters as inputs, an encoded depth field can be obtained by jointly minimizing a distortion criteria and a bit-rate measure. Since the method gives directly an encoded field as an output, it does not require an estimate of the field to be encoded. By efficiently encoding the depth field during the experiments, it is shown that the 3-D motion models can be used in object based video compression algorithms. A. Aydin Alatan, Levent Onural |
ICIP (2) | 1 |
| 1995 | Object based 3-D motion and structure estimationabstractMotion analysis is the most crucial part of object-based coding. Motion in a 3-D environment can be analyzed better by using a 3-D motion model compared to its 2-D counterpart and hence may improve coding efficiency. Gibbs formulated joint segmentation and estimation of 2-D motion not only improves the performance, but also generates robust point correspondences which are necessary for linear 3-D motion estimation algorithms. Estimated 3-D motion parameters are used to find the structure of the previously segmented objects by minimizing another Gibbs energy. Such an approach achieves error immunity compared to linear algorithms. Experimental results are promising and hence the proposed motion and structure analysis method is a candidate to be used in object-based (or even knowledge-based) video coding schemes. A. Aydin Alatan, Levent Onural |
ICIP | 1 |
| 1994 | Gibbs Random Field Model Based 3-D Motion Estimation by Weakened Rigidityabstract3-D motion estimation from a video sequence remains a challenging problem. Modelling the local interactions between the 3-D motion parameters is possible by using Gibbs random fields. An energy function which gives the joint probability distribution of the motion vectors, is constructed. The most probable motion vector set is found by maximizing the probability, represented by this distribution. Since the 3-D motion estimation problem is ill-posed, the regularization is achieved by an initial rigidity assumption. Afterwards, the rigidity is weakened hierarchically, until the finest level is reached. At the finest level, each point has its own motion vector and the "weak-connection" between these vectors are described by the energy function. The high computational cost Q decreased considerably by the multiprecision approach. The simulation results support all our discussions.> A. Aydin Alatan, Levent Onural |
ICIP (2) | 1 |