EDBT 2026 Demo / reviewers in the wild / expert
Fatih Porikli
dblp:p/FatihMuratPorikli · also Fatih Murat Porikli
· DBLP profile ↗
265ranked-venue papers
25as first author
73since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 189 · 23 first-author · 43 since 2021Artificial intelligence and machine learning · 167 · 5 first-author · 62 since 2021Systems, architecture and hardware · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3Computer networks · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HyperNet Fields: Efficiently Training Hypernetworks without Ground Truth by Learning Weight TrajectoriesabstractTo efficiently adapt large models or to train generative models of neural representations, Hypernetworks have drawn interest. While hypernetworks work well, training them is cumbersome, and often requires ground truth optimized weights for each sample. However, obtaining each of these weights is a training problem of its own—one needs to train, e.g., adaptation weights or even an entire neural field for hypernetworks to regress to. In this work, we propose a method to train hypernetworks, without the need for any per-sample ground truth. Our key idea is to learn a Hypernetwork ‘Field’ and estimate the entire trajectory of network weight training instead of simply its converged state. In other words, we introduce an additional input to the Hypernetwork, the convergence state, which then makes it act as a neural field that models the entire convergence pathway of a task network. A critical benefit in doing so is that the gradient of the estimated weights at any convergence state must then match the gradients of the original task—this constraint alone is sufficient to train the Hypernetwork Field. We demonstrate the effectiveness of our method through the task of personalized image generation and 3D shape reconstruction from images and point clouds, demonstrating competitive results without any per-sample ground truth. Eric Hedlin, Munawar Hayat, Fatih Porikli, Kwang Moo Yi, Shweta Mahajan |
CVPR | 3 |
| 2025 | Distilling Multi-modal Large Language Models for Autonomous DrivingabstractAutonomous driving demands safe motion planning, especially in critical "long-tail" scenarios. Recent end-to-end autonomous driving systems leverage large language models (LLMs) as planners to improve generalizability to rare events. However, using LLMs at test time introduces high computational costs. To address this, we propose DiMA, an end-to-end autonomous driving system that maintains the efficiency of an LLM-free (or vision-based) planner while leveraging the world knowledge of an LLM. DiMA distills the information from a multi-modal LLM to a vision-based end-to-end planner through a set of specially designed surrogate tasks. Under a joint training strategy, a scene encoder common to both networks produces structured representations that are semantically grounded as well as aligned to the final planning objective. Notably, the LLM is optional at inference, enabling robust planning without compromising on efficiency. Training with DiMA results in a 37% reduction in the L2 trajectory error and an 80% reduction in the collision rate of the vision-based planner, as well as a 44% trajectory error reduction in long-tail scenarios. DiMA also achieves state-of-the-art performance on the nuScenes planning benchmark. Deepti Hegde, Rajeev Yasarla, Shizhong Han, Apratim Bhattacharyya, Shweta Mahajan, Litian Liu, Risheek Garrepalli, Vishal M. Patel, Fatih Porikli |
CVPR | 10 |
| 2025 | CustomKD: Customizing Large Vision Foundation for Edge Model Improvement via Knowledge DistillationabstractWe propose a novel knowledge distillation approach, CustomKD, that effectively leverages large vision foundation models (LVFMs) to enhance the performance of edge models (e.g., MobileNetV3). Despite recent advancements in LVFMs, such as DINOv2 and CLIP, their potential in knowledge distillation for enhancing edge models remains underexplored. While knowledge distillation is a promising approach for improving the performance of edge models, the discrepancy in model capacities and heterogeneous architectures between LVFMs and edge models poses a significant challenge. Our observation indicates that although utilizing larger backbones (e.g., ViT-S to ViT-L) in teacher models improves their downstream task performances, the knowledge distillation from the large teacher models fails to bring as much performance gain for student models as for teacher models due to the large model discrepancy. Our simple yet effective CustomKD customizes the well-generalized features inherent in LVFMs to a given student model in order to reduce model discrepancies. Specifically, beyond providing well-generalized original knowledge from teachers, CustomKD aligns the features of teachers to those of students, making it easy for students to understand and overcome the large model discrepancy overall. CustomKD significantly improves the performances of edge models in scenarios with unlabeled data such as unsupervised domain adaptation (e.g., OfficeHome and DomainNet) and semi-supervised learning (e.g., CIFAR-100 and ImageNet), achieving the new state-of-the-art performances. Jungsoo Lee, Debasmit Das, Munawar Hayat, Sungha Choi, Kyuwoong Hwang, Fatih Porikli |
CVPR | 6 |
| 2025 | ConsNoTrainLoRA: Data-driven Weight Initialization of Low-Rank Adapters Using Constraints
Debasmit Das, Hyoungwoo Park, Munawar Hayat, Seokeon Choi, Sungrack Yun, Fatih Porikli |
ICCV | 6 |
| 2025 | Understanding Personal Concept in Open-Vocabulary Semantic Segmentation
Sunghyun Park 0005, Jungsoo Lee, Shubhankar Borse, Munawar Hayat, Sungha Choi, Kyuwoong Hwang, Fatih Porikli |
ICCV | 7 |
| 2025 | DuoLoRA: Cycle-Consistent and Rank-Disentangled Content-Style Personalization
Aniket Roy, Shubhankar Borse, Shreya Kadambi, Debasmit Das, Shweta Mahajan, Risheek Garrepalli, Hyojin Park 0004, Ankita Nayak, Rama Chellappa, Munawar Hayat, Fatih Porikli |
ICCV | 11 |
| 2025 | LoRA-X: Bridging Foundation Models with Training-Free Cross-Model AdaptationabstractThe rising popularity of large foundation models has led to a heightened demand for parameter-efficient fine-tuning methods, such as Low-Rank Adaptation (LoRA), which offer performance comparable to full model fine-tuning while requiring only a few additional parameters tailored to the specific base model. When such base models are deprecated and replaced, all associated LoRA modules must be retrained, requiring access to either the original training data or a substantial amount of synthetic data that mirrors the original distribution. However, the original data is often inaccessible due to privacy or licensing issues, and generating synthetic data may be impractical and insufficiently representative. These factors complicate the fine-tuning process considerably. To address this challenge, we introduce a new adapter, Cross-Model Low-Rank Adaptation (LoRA-X), which enables the training-free transfer of LoRA parameters across source and target models, eliminating the need for original or synthetic training data. Our approach imposes the adapter to operate within the subspace of the source base model. This constraint is necessary because our prior knowledge of the target model is limited to its weights, and the criteria for ensuring the adapter’s transferability are restricted to the target base model’s weights and subspace. To facilitate the transfer of LoRA parameters of the source model to a target model, we employ the adapter only in the layers of the target model that exhibit an acceptable level of subspace similarity. Our extensive experiments demonstrate the effectiveness of LoRA-X for text-to-image generation, including Stable Diffusion v1.5 and Stable Diffusion XL. Farzad Farhadzadeh, Debasmit Das, Shubhankar Borse, Fatih Porikli |
ICLR | 4 |
| 2025 | Sort-free Gaussian Splatting via Weighted Sum RenderingabstractRecently, 3D Gaussian Splatting (3DGS) has emerged as a significant advancement in 3D scene reconstruction, attracting considerable attention due to its ability to recover high-fidelity details while maintaining low complexity. Despite the promising results achieved by 3DGS, its rendering performance is constrained by its dependence on costly non-commutative alpha-blending operations. These operations mandate complex view dependent sorting operations that introduce computational overhead, especially on the resource-constrained platforms such as mobile phones. In this paper, we propose Weighted Sum Rendering, which approximates alpha blending with weighted sums, thereby removing the need for sorting. This simplifies implementation, delivers superior performance, and eliminates the ``popping'' artifacts caused by sorting. Experimental results show that optimizing a generalized Gaussian splatting formulation to the new differentiable rendering yields competitive image quality. The method was implemented and tested in a mobile device GPU, achieving on average $1.23\times$ faster rendering. Qiqi Hou, Randall Rauwendaal, Hoang Le, Farzad Farhadzadeh, Fatih Porikli, Alex Bourd, Amir Said |
ICLR | 6 |
| 2025 | PADRe: A Unifying Polynomial Attention Drop-in Replacement for Efficient Vision TransformerabstractWe present Polynomial Attention Drop-in Replacement (PADRe), a novel and unifying framework designed to replace the conventional self-attention mechanism in transformer models. Notably, several recent alternative attention mechanisms, including Hyena, Mamba, SimA, Conv2Former, and Castling-ViT, can be viewed as specific instances of our PADRe framework. PADRe leverages polynomial functions and draws upon established results from approximation theory, enhancing computational efficiency without compromising accuracy. PADRe's key components include multiplicative nonlinearities, which we implement using straightforward, hardware-friendly operations such as Hadamard products, incurring only linear computational and memory costs. PADRe further avoids the need for using complex functions such as Softmax, yet it maintains comparable or superior accuracy compared to traditional self-attention. We assess the effectiveness of PADRe as a drop-in replacement for self-attention across diverse computer vision tasks. These tasks include image classification, image-based 2D object detection, and 3D point cloud object detection. Empirical results demonstrate that PADRe runs significantly faster than the conventional self-attention (11x~43x faster on server GPU and mobile NPU) while maintaining similar accuracy when substituting self-attention in the transformer models. Pierre-David Létourneau, Manish Kumar Singh 0002, Hsin-Pai Cheng, Shizhong Han, Yunxiao Shi, Dalton Jones, Harper Langston, Fatih Porikli |
ICLR | 9 |
| 2025 | Zero-Shot Adaptation of Parameter-Efficient Fine-Tuning in Diffusion ModelsabstractWe introduce ProLoRA, enabling zero-shot adaptation of parameter-efficient fine-tuning in text-to-image diffusion models. ProLoRA transfers pre-trained low-rank adjustments (e.g., LoRA) from a source to a target model without additional training data. This overcomes the limitations of traditional methods that require retraining when switching base models, often challenging due to data constraints. ProLoRA achieves this via projection of source adjustments into the target model’s weight space, leveraging subspace and null space similarities and selectively targeting aligned layers. Evaluations on established text-to-image models demonstrate successful knowledge transfer and comparable performance without retraining. Farzad Farhadzadeh, Debasmit Das, Shubhankar Borse, Fatih Porikli |
ICML | 4 |
| 2025 | H3O: Hyper-Efficient 3D Occupancy Prediction with Heterogeneous Supervisionabstract3D occupancy prediction has recently emerged as a new paradigm for holistic 3D scene understanding and provides valuable information for downstream planning in autonomous driving. Most existing methods, however, are computationally expensive, requiring costly attention-based 2D- 3D transformation and 3D feature processing. In this paper, we present a novel 3D occupancy prediction approach, H30, which features highly efficient architecture designs that incur a significantly lower computational cost as compared to the current state-of-the-art methods. In addition, to compensate for the ambiguity in ground-truth 3D occupancy labels, we advocate leveraging auxiliary tasks to complement the direct 3D supervision. In particular, we integrate multi-camera depth estimation, semantic segmentation, and surface normal estimation via differentiable volume rendering, supervised by corresponding 2D labels that introduces rich and heterogeneous supervision signals. We conduct extensive experiments on the Occ3D-nuScenes and SemanticKITTI benchmarks that demonstrate the superiority of our proposed H30. Yunxiao Shi, Amin Ansari, Fatih Porikli |
ICRA | 4 |
| 2025 | MultiHuman-Testbench: Benchmarking Image Generation for Multiple HumansabstractGeneration of images containing multiple humans, performing complex actions, while preserving their facial identities, is a significant challenge. A major factor contributing to this is the lack of a a dedicated benchmark. To address this, we introduce MultiHuman-Testbench, a novel benchmark for rigorously evaluating generative models for multi-human generation. The benchmark comprises 1800 samples, including carefully curated text prompts, describing a range of simple to complex human actions. These prompts are matched with a total of 5,550 unique human face images, sampled uniformly to ensure diversity across age, ethnic background, and gender. Alongside captions, we provide human-selected pose conditioning images which accurately match the prompt. We propose a multi-faceted evaluation suite employing four key metrics to quantify face count, ID similarity, prompt alignment, and action detection. We conduct a thorough evaluation of a diverse set of models, including zero-shot approaches and training-based methods, with and without regional priors. We also propose novel techniques to incorporate image and region isolation using human segmentation and Hungarian matching, significantly improving ID similarity. Our proposed benchmark and key findings provide valuable insights and a standardized tool for advancing research in multi-human image generation. Shubhankar Borse, Seokeon Choi, Sunghyun Park 0005, Jeongho Kim 0007, Shreya Kadambi, Risheek Garrepalli, Sungrack Yun, Durga Malladi, Fatih Porikli |
NeurIPS | 9 |
| 2025 | Generalized Contrastive Learning for Universal Multimodal RetrievalabstractDespite their consistent performance improvements, cross-modal retrieval models (e.g., CLIP) show degraded performances with retrieving keys composed of fused image-text modality (e.g., Wikipedia pages with both images and text). To address this critical challenge, multimodal retrieval has been recently explored to develop a unified single retrieval model capable of retrieving keys across diverse modality combinations. A common approach involves constructing new composed sets of image-text triplets (e.g., retrieving a pair of image and text given a query image). However, such an approach requires careful curation to ensure the dataset quality and fails to generalize to unseen modality combinations. To overcome these limitations, this paper proposes Generalized Contrastive Learning (GCL), a novel loss formulation that improves multimodal retrieval performance without the burdensome need for new dataset curation. Specifically, GCL operates by enforcing contrastive learning across all modalities within a mini-batch, utilizing existing image-caption paired datasets to learn a unified representation space. We demonstrate the effectiveness of GCL by showing consistent performance improvements on off-the-shelf multimodal retrieval models (e.g., VISTA, CLIP, and TinyCLIP) using the M-BEIR, MMEB, and CoVR benchmarks. Jungsoo Lee, Janghoon Cho, Hyojin Park 0004, Durga Malladi, Kyuwoong Hwang, Fatih Porikli, Sungha Choi |
NeurIPS | 6 |
| 2025 | ODG: Occupancy Prediction Using Dual GaussiansabstractOccupancy prediction infers fine-grained 3D geometry and semantics from camera images of the surrounding environment, making it a critical perception task for autonomous driving. Existing methods either adopt dense grids as scene representation which is difficult to scale to high resolution, or learn the entire scene using a single set of sparse queries, which is insufficient to handle the various object characteristics. In this paper, we present ODG, a hierarchical dual sparse Gaussian representation to effectively capture complex scene dynamics. Building upon the observation that driving scenes can be universally decomposed into static and dynamic counterparts, we define dual Gaussian queries to better model the diverse scene objects. We utilize a hierarchical Gaussian transformer to predict the occupied voxel centers and semantic classes along with the Gaussian parameters. Leveraging the real-time rendering capability of 3D Gaussian Splatting, we also impose rendering supervision with available depth and semantic map annotations injecting pixel-level alignment to boost occupancy learning. Extensive experiments on the Occ3D-nuScenes and Occ3D-Waymo benchmarks demonstrate our proposed method sets new state-of-the-art results while maintaining low inference cost. Yunxiao Shi, Yinhao Zhu, Herbert Cai, Shizhong Han, Jisoo Jeong, Amin Ansari, Fatih Porikli |
NeurIPS | 7 |
| 2025 | HexaGen3D: StableDiffusion is One Step Away from Fast and Diverse Text-to-3D GenerationabstractDespite the latest remarkable advances in generative modeling, efficient generation of high-quality 3D objects from textual prompts remains a difficult task. A key chal-lenge lies in data scarcity: the most extensive 3D datasets encompass merely millions of samples, while their 2D coun-terparts contain billions of text-image pairs. To address this, we propose a novel approach which harnesses the power of large, pretrained 2D diffusion models. More specifically, our approach, HexaGen3D, fine-tunes a pre-trained text-to-image model to jointly predict 6 orthographic projections and the corresponding 3D latent. We then decode these latents to generate a textured mesh. Hex-aGen3D does not require per-sample optimization, and can infer high-quality and diverse objects from textual prompts in 7 seconds, offering significantly better quality-to-latency trade-offs than existing approaches. Furthermore, Hexa-Gen3D demonstrates strong generalization to new objects or compositions. Antoine Mercier 0005, Ramin Nakhli, Mahesh Reddy, Rajeev Yasarla, Fatih Porikli, Guillaume Berger |
WACV | 6 |
| 2025 | Planar Gaussian SplattingabstractThis paper presents Planar Gaussian Splatting (PGS), a novel neural rendering approach to learn the 3D geometry and parse the 3D planes of a scene, directly from multiple RGB images. The PGS leverages Gaussian primitives to model the scene and employ a hierarchical Gaussian mixture approach to group them. Similar Gaussians are progressively merged probabilistically in the tree-structured Gaussian mixtures to identify distinct 3D plane instances and form the overall 3D scene geometry. In order to enable the grouping, the Gaussian primitives contain additional parameters, such as plane descriptors derived by lifting 2D masks from a general 2D segmentation model and surface normals. Experiments show that the proposed PGS achieves state-of-the-art performance in 3D planar reconstruction without requiring either 3D plane labels or depth supervision. In contrast to existing supervised methods that have limited generalizability and struggle under domain shift, PGS maintains its performance across datasets thanks to its neural rendering and scene-specific optimization mechanism, while also being significantly faster than existing optimization-based approaches. Farhad G. Zanjani, Hanno Ackermann, Leyla Mirvakhabova, Fatih Porikli |
WACV | 5 |
| 2024 | Clockwork Diffusion: Efficient Generation With Model-Step DistillationabstractThis work aims to improve the efficiency of text-to-image diffusion models. While diffusion models use computationallyexpensive UNet-based denoising operations in ev-ery generation step, we identify that not all operations are equally relevant for the final output quality. In par-ticular, we observe that UNet layers operating on high-res feature maps are relatively sensitive to small pertur-bations. In contrast, low-res feature maps influence the semantic layout of the final image and can often be per-turbed with no noticeable change in the output. Based on this observation, we propose Clockwork Diffusion, a method that periodically reuses computation from preceding denoising steps to approximate low-res feature maps at one or more subsequent steps. For multiple base-lines, and for both text-to-image generation and image editing, we demonstrate that Clockwork leads to compa-rable or improved perceptual scores with drastically re-duced computational complexity. As an example, for Sta-ble Diffusion vI.5 with 8 DPM++ steps we save 32% of FLOPs with negligible FID and CLIP change. We re-lease code at https://github.com/Qualcomm-AI-research/clockwork-diffusion AmirHossein Habibian, Amir Ghodrati, Noor Fathima, Guillaume Sautière, Risheek Garrepalli, Fatih Porikli, Jens Petersen |
CVPR | 6 |
| 2024 | OCAI: Improving Optical Flow Estimation by Occlusion and Consistency Aware InterpolationabstractThe scarcity of ground-truth labels poses one major challenge in developing optical flow estimation models that are both generalizable and robust. While current methods rely on data augmentation, they have yet to fully exploit the rich information available in labeled video sequences. We propose OCAI, a method that supports robust frame interpolation by generating intermediate video frames along-side optical flows in between. Utilizing a forward warping approach, OCAI employs occlusion awareness to resolve ambiguities in pixel values and fills in missing values by leveraging the forward-backward consistency of optical flows. Additionally, we introduce a teacher-student style semi-supervised learning method on top of the interpolated frames. Using a pair of unlabeled frames and the teacher model's predicted optical flow, we generate interpolated frames and flows to train a student model. The teacher's weights are maintained using Exponential Moving Averaging of the student. Our evaluations demonstrate perceptually superior interpolation quality and enhanced optical flow accuracy on established benchmarks such as Sintel and KITTI. Jisoo Jeong, Risheek Garrepalli, Jamie Menjay Lin, Munawar Hayat, Fatih Porikli |
CVPR | 6 |
| 2024 | DeCoTR: Enhancing Depth Completion with 2D and 3D AttentionsabstractIn this paper, we introduce a novel approach that har-nesses both 2D and 3D attentions to enable highly accurate depth completion without requiring iterative spatial propa-gations. Specifically, we first enhance a baseline convolutional depth completion model by applying attention to 2D features in the bottleneck and skip connections. This effectively improves the performance of this simple network and sets it on par with the latest, complex transformer-based models. Leveraging the initial depths and features from this network, we uplift the 2D features to form a 3D point cloud and construct a 3D point transformer to process it, allowing the model to explicitly learn and exploit 3D geometric features. In addition, we propose normalization techniques to process the point cloud, which improves learning and leads to better accuracy than directly using point transformers off the shelf. Furthermore, we incorporate global attention on downsampled point cloud features, which enables long-range context while still being computationally feasible. We evaluate our method, DeCoTr, on established depth Completion benchmarks, including NYU Depth V2 and KITTI, showcasing that it sets new state-of-the-art performance. We further conduct zero-shot evaluations on ScanNet and DDAD benchmarks and demonstrate that DeCoTR has su-perior generalizability compared to existing approaches. Yunxiao Shi, Manish Kumar Singh 0002, Fatih Porikli |
CVPR | 4 |
| 2024 | Neural Graphics Texture Compression Supporting Random Access
Farzad Farhadzadeh, Qiqi Hou, Hoang Le, Amir Said, Randall Rauwendaal, Alex Bourd, Fatih Porikli |
ECCV (37) | 7 |
| 2024 | Object-Centric Diffusion for Efficient Video Editing
Kumara Kahatapitiya, Adil Karjauv, Davide Abati, Fatih Porikli, Yuki Markus Asano, AmirHossein Habibian |
ECCV (57) | 4 |
| 2024 | FutureDepth: Learning to Predict the Future Improves Video Depth Estimation
Rajeev Yasarla, Manish Kumar Singh 0002, Yunxiao Shi, Jisoo Jeong, Yinhao Zhu, Shizhong Han, Risheek Garrepalli, Fatih Porikli |
ECCV (5) | 9 |
| 2024 | Neural Mesh Fusion: Unsupervised 3D Planar Surface UnderstandingabstractThis paper presents Neural Mesh Fusion (NMF), an efficient approach for joint optimization of polygon mesh from multi-view image observations and unsupervised 3D planar-surface parsing of the scene. In contrast to implicit neural representations, NMF directly learns to deform surface triangle mesh and generate an embedding for unsupervised 3D planar segmentation through gradient-based optimization directly on the surface mesh. The conducted experiments show that NMF obtains competitive results compared to state-of-the-art multi-view planar reconstruction, while not requiring any ground-truth 3D or planar supervision. Moreover, NMF is significantly more computationally efficient compared to implicit neural rendering-based scene reconstruction approaches. Farhad G. Zanjani, Yinhao Zhu, Leyla Mirvakhabova, Fatih Porikli |
ICIP | 5 |
| 2024 | Skip-Attention: Improving Vision Transformers by Paying Less AttentionabstractThis work aims to improve the efficiency of vision transformers (ViTs). While ViTs use computationally expensive self-attention operations in every layer, we identify that these operations are highly correlated across layers -- a key redundancy that causes unnecessary computations. Based on this observation, we propose SkipAT a method to reuse self-attention computation from preceding layers to approximate attention at one or more subsequent layers. To ensure that reusing self-attention blocks across layers does not degrade the performance, we introduce a simple parametric function, which outperforms the baseline transformer's performance while running computationally faster. We show that SkipAT is agnostic to transformer architecture and is effective in image classification, semantic segmentation on ADE20K, image denoising on SIDD, and video denoising on DAVIS. We achieve improved throughput at the same-or-higher accuracy levels in all these tasks. Shashanka Venkataramanan, Amir Ghodrati, Yuki Markus Asano, Fatih Porikli, AmirHossein Habibian |
ICLR | 4 |
| 2024 | FouRA: Fourier Low-Rank AdaptationabstractWhile Low-Rank Adaptation (LoRA) has proven beneficial for efficiently fine-tuning large models, LoRA fine-tuned text-to-image diffusion models lack diversity in the generated images, as the model tends to copy data from the observed training samples. This effect becomes more pronounced at higher values of adapter strength and for adapters with higher ranks which are fine-tuned on smaller datasets. To address these challenges, we present FouRA, a novel low-rank method that learns projections in the Fourier domain along with learning a flexible input-dependent adapter rank selection strategy. Through extensive experiments and analysis, we show that FouRA successfully solves the problems related to data copying and distribution collapse while significantly improving the generated image quality. We demonstrate that FouRA enhances the generalization of fine-tuned models thanks to its adaptive rank selection. We further show that the learned projections in the frequency domain are decorrelated and prove effective when merging multiple adapters. While FouRA is motivated for vision tasks, we also demonstrate its merits for language tasks on commonsense reasoning and GLUE benchmarks. Shubhankar Borse, Shreya Kadambi, Nilesh Prasad Pandey, Kartikeya Bhardwaj, Viswanath Ganapathy, Sweta Priyadarshi, Risheek Garrepalli, Rafael Esteves 0002, Munawar Hayat, Fatih Porikli |
NeurIPS | 10 |
| 2024 | Hollowed Net for On-Device Personalization of Text-to-Image Diffusion ModelsabstractRecent advancements in text-to-image diffusion models have enabled the personalization of these models to generate custom images from textual prompts. This paper presents an efficient LoRA-based personalization approach for on-device subject-driven generation, where pre-trained diffusion models are fine-tuned with user-specific data on resource-constrained devices. Our method, termed Hollowed Net, enhances memory efficiency during fine-tuning by modifying the architecture of a diffusion U-Net to temporarily remove a fraction of its deep layers, creating a hollowed structure. This approach directly addresses on-device memory constraints and substantially reduces GPU memory requirements for training, in contrast to previous methods that primarily focus on minimizing training steps and reducing the number of parameters to update. Additionally, the personalized Hollowed Net can be transferred back into the original U-Net, enabling inference without additional memory overhead. Quantitative and qualitative analyses demonstrate that our approach not only reduces training memory to levels as low as those required for inference but also maintains or improves personalization performance compared to existing methods. Wonguk Cho, Seokeon Choi, Debasmit Das, Matthias Reisser, Taesup Kim, Sungrack Yun, Fatih Porikli |
NeurIPS | 7 |
| 2024 | Guidance Through Surrogate: Toward a Generic Diagnostic AttackabstractAdversarial training (AT) is an effective approach to making deep neural networks robust against adversarial attacks. Recently, different AT defenses are proposed that not only maintain a high clean accuracy but also show significant robustness against popular and well-studied adversarial attacks, such as projected gradient descent (PGD). High adversarial robustness can also arise if an attack fails to find adversarial gradient directions, a phenomenon known as "gradient masking." In this work, we analyze the effect of label smoothing on AT as one of the potential causes of gradient masking. We then develop a guided mechanism to avoid local minima during attack optimization, leading to a novel attack dubbed guided projected gradient attack (G-PGA). Our attack approach is based on a "match and deceive" loss that finds optimal adversarial directions through guidance from a surrogate model. Our modified attack does not require random restarts a large number of attack iterations or a search for optimal step size. Furthermore, our proposed G-PGA is generic, thus it can be combined with an ensemble attack strategy as we demonstrate in the case of auto-attack, leading to efficiency and convergence speed improvements. More than an effective attack, G-PGA can be used as a diagnostic tool to reveal elusive robustness due to gradient masking in adversarial defenses. Muzammal Naseer, Salman Khan 0001, Fatih Porikli, Fahad Shahbaz Khan |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | DejaVu: Conditional Regenerative Learning to Enhance Dense PredictionabstractWe present DejaVu, a novel framework which leverages conditional image regeneration as additional supervision during training to improve deep networks for dense prediction tasks such as segmentation, depth estimation, and surface normal prediction. First, we apply redaction to the input image, which removes certain structural information by sparse sampling or selective frequency removal. Next, we use a conditional regenerator, which takes the redacted image and the dense predictions as inputs, and reconstructs the original image by filling in the missing structural information. In the redacted image, structural attributes like boundaries are broken while semantic context is largely preserved. In order to make the regeneration feasible, the conditional generator will then require the structure information from the other input source, i.e., the dense predictions. As such, by including this conditional regeneration objective during training, DejaVu encourages the base network to learn to embed accurate scene structure in its dense prediction. This leads to more accurate predictions with clearer boundaries and better spatial consistency. When it is feasible to leverage additional computation, DejaVu can be extended to incorporate an attention-based regeneration module within the dense prediction network, which further improves accuracy. Through extensive experiments on multiple dense prediction benchmarks such as Cityscapes, COCO, ADE20K, NYUD-v2, and KITTI, we demonstrate the efficacy of employing DejaVu during training, as it out-performs SOTA methods at no added computation cost. Shubhankar Borse, Debasmit Das, Hyojin Park 0004, Risheek Garrepalli, Fatih Porikli |
CVPR | 6 |
| 2023 | DistractFlow: Improving Optical Flow Estimation via Realistic Distractions and Pseudo-LabelingabstractWe propose a novel data augmentation approach, DistractFlow, for training optical flow estimation models by introducing realistic distractions to the input frames. Based on a mixing ratio, we combine one of the frames in the pair with a distractor image depicting a similar domain, which allows for inducing visual perturbations congruent with natural objects and scenes. We refer to such pairs as distracted pairs. Our intuition is that using semantically meaningful distractors enables the model to learn related variations and attain robustness against challenging deviations, compared to conventional augmentation schemes focusing only on low-level aspects and modifications. More specifically, in addition to the supervised loss computed between the estimated flow for the original pair and its ground-truth flow, we include a second supervised loss defined between the distracted pair's flow and the original pair's ground-truth flow, weighted with the same mixing ratio. Furthermore, when unlabeled data is available, we extend our augmentation approach to self-supervised settings through pseudo-labeling and cross-consistency regularization. Given an original pair and its distracted version, we enforce the estimated flow on the distracted pair to agree with the flow of the original pair. Our approach allows increasing the number of available training pairs significantly without requiring additional annotations. It is agnostic to the model architecture and can be applied to training any optical flow estimation models. Our extensive evaluations on multiple benchmarks, including Sintel, KITTI, and SlowFlow, show that DistractFlow improves existing models consistently, outperforming the latest state of the art. Jisoo Jeong, Risheek Garrepalli, Fatih Porikli |
CVPR | 4 |
| 2023 | X3KD: Knowledge Distillation Across Modalities, Tasks and Stages for Multi-Camera 3D Object DetectionabstractRecent advances in 3D object detection (3DOD) have obtained remarkably strong results for LiDAR-based models. In contrast, surround-view 3DOD models based on multiple camera images underperform due to the necessary view transformation of features from perspective view (PV) to a 3D world representation which is ambiguous due to missing depth information. This paper introduces X3KD, a comprehensive knowledge distillation framework across different modalities, tasks, and stages for multi-camera 3DOD. Specifically, we propose cross-task distillation from an instance segmentation teacher (X-IS) in the PV feature extraction stage providing supervision without ambiguous error backpropagation through the view transformation. After the transformation, we apply cross-modal feature distillation (X-FD) and adversarial training (X-AT) to improve the 3D world representation of multi-camera features through the information contained in a LiDAR-based 3DOD teacher. Finally, we also employ this teacher for cross-modal output distillation (X-OD), providing dense supervision at the prediction stage. We perform extensive ablations of knowledge distillation at different stages of multi-camera 3DOD. Our final X3KD model outperforms previous state-of-the-art approaches on the nuScenes and Waymo datasets and generalizes to RADAR-based 3DOD. Qualitative results video at https://youtu.be/1do9DPFmr38. Marvin Klingner, Shubhankar Borse, Varun Ravi Kumar, Behnaz Rezaei, Venkatraman Narayanan, Senthil Kumar Yogamani, Fatih Porikli |
CVPR | 7 |
| 2023 | PartSLIP: Low-Shot Part Segmentation for 3D Point Clouds via Pretrained Image-Language ModelsabstractGeneralizable 3D part segmentation is important but challenging in vision and robotics. Training deep models via conventional supervised methods requires large-scale 3D datasets with fine-grained part annotations, which are costly to collect. This paper explores an alternative way for low-shot part segmentation of 3D point clouds by leveraging a pretrained image-language model, GLIP. which achieves superior performance on open-vocabulary 2D detection. We transfer the rich knowledge from 2D to 3D through GLIP-based part detection on point cloud rendering and a novel 2D-to-3D label lifting algorithm. We also utilize multi-view 3D priors and few-shot prompt tuning to boost performance significantly. Extensive evaluation on PartNet and PartNet-Mobility datasets shows that our method enables excellent zero-shot 3D part segmentation. Our few-shot version not only outperforms existing few-shot approaches by a large margin but also achieves highly competitive results compared to the fully supervised counterpart. Furthermore, we demonstrate that our method can be directly applied to iPhone-scanned point clouds without significant domain gaps. Minghua Liu, Yinhao Zhu, Shizhong Han, Zhan Ling, Fatih Porikli, Hao Su 0001 |
CVPR | 6 |
| 2023 | Neural 5G Indoor Localization with IMU SupervisionabstractRadio signals are well suited for user localization because they are ubiquitous, can operate in the dark and maintain privacy. Many prior works learn mappings between channel state information (CSI) and position fully-supervised. However, that approach relies on position labels which are very expensive to acquire. In this work, this requirement is relaxed by using pseudo-labels during deployment, which are calculated from an inertial measurement unit (IMU). We propose practical algorithms for IMU double integration and training of the localization system. We show decimeter-level accuracy on simulated and challenging real data of 5G measurements. Our IMU-supervised method performs similarly to fully-supervised, but requires much less effort to deploy. Aleksandr Ermolov, Shreya Kadambi, Maximilian Arnold, Mohammed Hirzallah, Roohollah Amiri, Deepak Singh Mahendar Singh, Srinivas Yerramalli, Daniel Dijkman, Fatih Porikli, Taesang Yoo, Bence Major |
GLOBECOM | 9 |
| 2023 | Transadapt: A Transformative Framework for Online Test Time Adaptive Semantic SegmentationabstractTest-time adaptive (TTA) semantic segmentation adapts a source pre-trained image semantic segmentation model to unlabeled batches of target domain test images, different from real-world, where samples arrive one-by-one in an online fashion. To tackle online settings, we propose TransAdapt, a framework that uses transformer and input transformations to improve segmentation performance. Specifically, we pre-train a transformer-based module on a segmentation network that transforms unsupervised segmentation output to a more reliable supervised output, without requiring test-time online training. To also facilitate test-time adaptation, we propose an unsupervised loss based on the transformed input that enforces the model to be invariant and equivariant to photometric and geometric perturbations, respectively. Overall, our framework produces higher quality segmentation masks with up to 17.6% and 2.8% mIOU improvement over no-adaptation and competitive baselines, respectively. Debasmit Das, Shubhankar Borse, Hyojin Park 0004, Kambiz Azarian, Risheek Garrepalli, Fatih Porikli |
ICASSP | 7 |
| 2023 | Efficient neural supersampling on a novel gaming datasetabstractReal-time rendering for video games has become increasingly challenging due to the need for higher resolutions, framerates and photorealism. Supersampling has emerged as an effective solution to address this challenge. Our work introduces a novel neural algorithm for super-sampling rendered content that is 4× more efficient than existing methods while maintaining the same level of accuracy. Additionally, we introduce a new dataset which provides auxiliary modalities such as motion vectors and depth generated using graphics rendering features like viewport jittering and mipmap biasing at different resolutions. We believe that this dataset fills a gap in the current dataset landscape and can serve as a valuable resource to help measure progress in the field and advance the state-of-the-art in super-resolution techniques for gaming content. Antoine Mercier 0005, Ruan Erasmus, Yashesh Savani, Manik Dhingra, Fatih Porikli, Guillaume Berger |
ICCV | 5 |
| 2023 | Factorized Inverse Path Tracing for Efficient and Accurate Material-Lighting EstimationabstractInverse path tracing has recently been applied to joint material and lighting estimation, given geometry and multi-view HDR observations of an indoor scene. However, it has two major limitations: path tracing is expensive to compute, and ambiguities exist between reflection and emission. Our Factorized Inverse Path Tracing (FIPT) addresses these challenges by using a factored light transport formulation and finds emitters driven by rendering errors. Our algorithm enables accurate material and lighting optimization faster than previous work, and is more effective at resolving ambiguities. The exhaustive experiments on synthetic scenes show that our method (1) outperforms state-of-the-art indoor inverse rendering and relighting methods particularly in the presence of complex illumination effects; (2) speeds up inverse path tracing optimization to less than an hour. We further demonstrate robustness to noisy inputs through material and lighting estimates that allow plausible relighting in a real scene. The source code is available at: https://github.com/lwwu2/fipt Liwen Wu, Rui Zhu 0026, Mustafa B. Yaldiz, Yinhao Zhu, Janarbek Matai, Fatih Porikli, Tzu-Mao Li, Manmohan Krishna Chandraker, Ravi Ramamoorthi |
ICCV | 7 |
| 2023 | MAMo: Leveraging Memory and Attention for Monocular Video Depth EstimationabstractWe propose MAMo, a novel memory and attention framework for monocular video depth estimation. MAMo can augment and improve any single-image depth estimation networks into video depth estimation models, enabling them to take advantage of the temporal information to predict more accurate depth. In MAMo, we augment model with memory which aids the depth prediction as the model streams through the video. Specifically, the memory stores learned visual and displacement tokens of the previous time instances. This allows the depth network to cross-reference relevant features from the past when predicting depth on the current frame. We introduce a novel scheme to continuously update the memory, optimizing it to keep tokens that correspond with both the past and the present visual information. We adopt attention-based approach to process memory features where we first learn the spatiotemporal relation among the resultant visual and displacement memory tokens using self-attention module. Further, the output features of self-attention are aggregated with the current visual features through cross-attention. The cross-attended features are finally given to a decoder to predict depth on the current frame. Through extensive experiments on several benchmarks, including KITTI, NYU-Depth V2, and DDAD, we show that MAMo consistently improves monocular depth estimation networks and sets new state-of-the-art (SOTA) accuracy. Notably, our MAMo video depth estimation provides higher accuracy with lower latency, when comparing to SOTA cost-volume-based video depth models. Rajeev Yasarla, Jisoo Jeong, Yunxiao Shi, Risheek Garrepalli, Fatih Porikli |
ICCV | 6 |
| 2023 | 4D Panoptic Segmentation as Invariant and Equivariant Field PredictionabstractIn this paper, we develop rotation-equivariant neural networks for 4D panoptic segmentation. 4D panoptic segmentation is a benchmark task for autonomous driving that requires recognizing semantic classes and object instances on the road based on LiDAR scans, as well as assigning temporally consistent IDs to instances across time. We observe that the driving scenario is symmetric to rotations on the ground plane. Therefore, rotation-equivariance could provide better generalization and more robust feature learning. Specifically, we review the object instance clustering strategies and restate the centerness-based approach and the offset-based approach as the prediction of invariant scalar fields and equivariant vector fields. Other subtasks are also unified from this perspective, and different invariant and equivariant layers are designed to facilitate their predictions. Through evaluation on the standard 4D panoptic segmentation benchmark of SemanticKITTI, we show that our equivariant models achieve higher accuracy with lower computational costs compared to their non-equivariant counterparts. Moreover, our method sets the new state-of-the-art performance and achieves 1st place on the SemanticKITTI 4D Panoptic Segmentation leaderboard. Minghan Zhu, Shizhong Han, Maani Ghaffari Jadidi, Fatih Porikli, Shubhankar Borse |
ICCV | 5 |
| 2023 | Self-Supervised Geometric Correspondence for Category-Level 6D Object Pose Estimation in the Wild
Shubhankar Borse, Fatih Porikli, Xiaolong Wang 0004 |
ICLR | 5 |
| 2023 | OpenShape: Scaling Up 3D Shape Representation Towards Open-World UnderstandingabstractWe introduce OpenShape, a method for learning multi-modal joint representations of text, image, and point clouds. We adopt the commonly used multi-modal contrastive learning framework for representation alignment, but with a specific focus on scaling up 3D representations to enable open-world 3D shape understanding. To achieve this, we scale up training data by ensembling multiple 3D datasets and propose several strategies to automatically filter and enrich noisy text descriptions. We also explore and compare strategies for scaling 3D backbone networks and introduce a novel hard negative mining module for more efficient training. We evaluate OpenShape on zero-shot 3D classification benchmarks and demonstrate its superior capabilities for open-world recognition. Specifically, OpenShape achieves a zero-shot accuracy of 46.8% on the 1,156-category Objaverse-LVIS benchmark, compared to less than 10% for existing methods. OpenShape also achieves an accuracy of 85.3% on ModelNet40, outperforming previous zero-shot baseline methods by 20% and performing on par with some fully-supervised methods. Furthermore, we show that our learned embeddings encode a wide range of visual and semantic concepts (e.g., subcategories, color, shape, style) and facilitate fine-grained text-3D and image-3D interactions. Due to their alignment with CLIP embeddings, our learned shape representations can also be integrated with off-the-shelf CLIP-based models for various applications, such as point cloud captioning and point cloud-conditioned image generation. Minghua Liu, Ruoxi Shi, Kaiming Kuang, Yinhao Zhu, Shizhong Han, Fatih Porikli, Hao Su 0001 |
NeurIPS | 8 |
| 2023 | X-Align: Cross-Modal Cross-View Alignment for Bird's-Eye-View SegmentationabstractBird’s-eye-view (BEV) grid is a typical representation of the perception of road components, e.g., drivable area, in autonomous driving. Most existing approaches rely on cameras only to perform segmentation in BEV space, which is fundamentally constrained by the absence of reliable depth information. The latest works leverage both camera and LiDAR modalities but suboptimally fuse their features using simple, concatenation-based mechanisms.In this paper, we address these problems by enhancing the alignment of the unimodal features in order to aid feature fusion, as well as enhancing the alignment between the cameras’ perspective view (PV) and BEV representations. We propose X-Align, a novel end-to-end cross-modal and cross-view learning framework for BEV segmentation consisting of the following components: (i) a novel CrossModal Feature Alignment (X-FA) loss, (ii) an attentionbased Cross-Modal Feature Fusion (X-FF) module to align multi-modal BEV features implicitly, and (iii) an auxiliary PV segmentation branch with Cross-View Segmentation Alignment (X-SA) losses to improve the PV-to-BEV transformation. We evaluate our proposed method across two commonly used benchmark datasets, i.e., nuScenes and KITTI-360. Notably, X-Align significantly outperforms the state-of-the-art by 3 absolute mIoU points on nuScenes. We also provide extensive ablation studies to demonstrate the effectiveness of the individual components. Shubhankar Borse, Marvin Klingner, Varun Ravi Kumar, Abdulaziz Almuzairee, Senthil Kumar Yogamani, Fatih Porikli |
WACV | 7 |
| 2023 | X-Align++: cross-modal cross-view alignment for Bird's-eye-view segmentation
Shubhankar Borse, Marvin Klingner, Varun Ravi Kumar, Abdulaziz Almuzairee, Senthil Kumar Yogamani, Fatih Porikli |
Mach. Vis. Appl. | 7 |
| 2023 | Consistency and Diversity Induced Human Motion SegmentationabstractSubspace clustering is a classical technique that has been widely used for human motion segmentation and other related tasks. However, existing segmentation methods often cluster data without guidance from prior knowledge, resulting in unsatisfactory segmentation results. To this end, we propose a novel Consistency and Diversity induced human Motion Segmentation (CDMS) algorithm. Specifically, our model factorizes the source and target data into distinct multi-layer feature spaces, in which transfer subspace learning is conducted on different layers to capture multi-level information. A multi-mutual consistency learning strategy is carried out to reduce the domain gap between the source and target data. In this way, the domain-specific knowledge and domain-invariant properties can be explored simultaneously. Besides, a novel constraint based on the Hilbert Schmidt Independence Criterion (HSIC) is introduced to ensure the diversity of multi-level subspace representations, which enables the complementarity of multi-level representations to be explored to boost the transfer learning performance. Moreover, to preserve the temporal correlations, an enhanced graph regularizer is imposed on the learned representation coefficients and the multi-level representations of the source data. The proposed model can be efficiently solved using the Alternating Direction Method of Multipliers (ADMM) algorithm. Extensive experimental results on public human motion datasets demonstrate the effectiveness of our method against several state-of-the-art approaches. Tao Zhou 0002, Huazhu Fu, Chen Gong 0002, Ling Shao 0001, Fatih Porikli, Haibin Ling, Jianbing Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Adaptive Siamese Tracking With a Compact Latent NetworkabstractIn this article, we provide an intuitive viewing to simplify the Siamese-based trackers by converting the tracking task to a classification. Under this viewing, we perform an in-depth analysis for them through visual simulations and real tracking examples, and find that the failure cases in some challenging situations can be regarded as the issue of missing decisive samples in offline training. Since the samples in the initial (first) frame contain rich sequence-specific information, we can regard them as the decisive samples to represent the whole sequence. To quickly adapt the base model to new scenes, a compact latent network is presented via fully using these decisive samples. Specifically, we present a statistics-based compact latent feature for fast adjustment by efficiently extracting the sequence-specific information. Furthermore, a new diverse sample mining strategy is designed for training to further improve the discrimination ability of the proposed compact latent network. Finally, a conditional updating strategy is proposed to efficiently update the basic models to handle scene variation during the tracking phase. To evaluate the generalization ability and effectiveness and of our method, we apply it to adjust three classical Siamese-based trackers, namely SiamRPN++, SiamFC, and SiamBAN. Extensive experimental results on six recent datasets demonstrate that all three adjusted trackers obtain the superior performance in terms of the accuracy, while having high running speed. Xingping Dong, Jianbing Shen, Fatih Porikli, Jiebo Luo 0001, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Stylized Adversarial DefenseabstractDeep Convolution Neural Networks (CNNs) can easily be fooled by subtle, imperceptible changes to the input images. To address this vulnerability, adversarial training creates perturbation patterns and includes them in the training set to robustify the model. In contrast to existing adversarial training methods that only use class-boundary information (e.g., using a cross-entropy loss), we propose to exploit additional information from the feature space to craft stronger adversaries that are in turn used to learn a robust model. Specifically, we use the style and content information of the target sample from another class, alongside its class-boundary information to create adversarial perturbations. We apply our proposed multi-task objective in a deeply supervised manner, extracting multi-scale feature knowledge to create maximally separating adversaries. Subsequently, we propose a max-margin adversarial training approach that minimizes the distance between source image and its adversary and maximizes the distance between the adversary and the target image. Our adversarial training approach demonstrates strong robustness compared to state-of-the-art defenses, generalizes well to naturally occurring corruptions and data distributional shifts, and retains the model's accuracy on clean examples. Muzammal Naseer, Salman Khan 0001, Munawar Hayat, Fahad Shahbaz Khan, Fatih Porikli |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | A Survey on Deep Learning Technique for Video SegmentationabstractVideo segmentation-partitioning video frames into multiple segments or objects-plays a critical role in a broad range of practical applications, from enhancing visual effects in movie, to understanding scenes in autonomous driving, to creating virtual background in video conferencing. Recently, with the renaissance of connectionism in computer vision, there has been an influx of deep learning based approaches for video segmentation that have delivered compelling performance. In this survey, we comprehensively review two basic lines of research - generic object segmentation (of unknown categories) in videos, and video semantic segmentation - by introducing their respective task settings, background concepts, perceived need, development history, and main challenges. We also offer a detailed overview of representative literature on both methods and datasets. We further benchmark the reviewed methods on several well-known datasets. Finally, we point out open issues in this field, and suggest opportunities for further research. We also provide a public website to continuously track developments in this fast advancing field: https://github.com/tfzhou/VS-Survey. Tianfei Zhou, Fatih Porikli, David Crandall, Luc Van Gool, Wenguan Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Panoptic, Instance and Semantic Relations: A Relational Context Encoder to Enhance Panoptic SegmentationabstractThis paper presents a novel framework to integrate both semantic and instance contexts for panoptic segmentation. In existing works, it is common to use a shared backbone to extract features for both things (countable classes such as vehicles) and stuff (uncountable classes such as roads). This, however, fails to capture the rich relations among them, which can be utilized to enhance visual understanding and segmentation performance. To address this short-coming, we propose a novel Panoptic, Instance, and Semantic Relations (PISR) module to exploit such contexts. First, we generate panoptic encodings to summarize key features of the semantic classes and predicted instances. A Panoptic Relational Attention (PRA) module is then applied to the encodings and the global feature map from the backbone. It produces a feature map that captures 1) the relations across semantic classes and instances and 2) the relations between these panoptic categories and spatial features. PISR also automatically learns to focus on the more important instances, making it robust to the number of instances used in the relational attention module. Moreover, PISR is a general module that can be applied to any existing panoptic segmentation architecture. Through extensive evaluations on panoptic segmentation benchmarks like Cityscapes, COCO, and ADE20K, we show that PISR attains considerable improvements over existing approaches. Shubhankar Borse, Hyojin Park 0004, Debasmit Das, Risheek Garrepalli, Fatih Porikli |
CVPR | 6 |
| 2022 | Imposing Consistency for Optical Flow EstimationabstractImposing consistency through proxy tasks has been shown to enhance data-driven learning and enable self-supervision in various tasks. This paper introduces novel and effective consistency strategies for optical flow estimation, a problem where labels from real-world data are very challenging to derive. More specifically, we propose occlusion consistency and zero forcing in the forms of self-supervised learning and transformation consistency in the form of semi-supervised learning. We apply these consistency techniques in a way that the network model learns to describe pixel-level motions better while requiring no additional annotations. We demonstrate that our consistency strategies applied to a strong baseline network model using the original datasets and labels provide further improvements, attaining the state-of-the-art results on the KITTI-2015 scene flow benchmark in the non-stereo category. Our method achieves the best foreground accuracy (4.33% in Fl-all) over both the stereo and non-stereo categories, even though using only monocular image inputs. Jisoo Jeong, Jamie Menjay Lin, Fatih Porikli, Nojun Kwak |
CVPR | 3 |
| 2022 | Real-Time, Accurate, and Consistent Video Semantic Segmentation via Unsupervised Adaptation and Cross-Unit Deployment on Mobile DeviceabstractThis demonstration showcases our innovations on efficient, accurate, and temporally consistent video semantic segmentation on mobile device. We employ our test-time unsupervised scheme, AuxAdapt, to enable the segmentation model to adapt to a given video in an online manner. More specifically, we leverage a small auxiliary network to perform weight updates and keep the large, main segmen-tation network frozen. This significantly reduces the computational cost of adaptation when compared to previous methods (e.g., Tent, DVP), and at the same time, prevents catastrophic forgetting. By running AuxAdapt, we can considerably improve the temporal consistency of video segmentation while maintaining the accuracy. We demonstrate how to efficiently deploy our adaptive video segmentation algorithm on a smartphone powered by a Snapdragon® Mobile Platform11Snapdragon is a product of Qualcomm Technologies, Inc. and/or its subsidiaries., Rather than simply running the entire algorithm on the GPU, we adopt a crossunit deployment strategy. The main network, which will be frozen during test time, will perform inferences on a highly optimized AI accelerator unit, while the small auxiliary net-work, which will be updated on the fly, will run forward passes and back-propagations on the GPU. Such a deployment scheme best utilizes the available processing power on the smartphone and enables real-time operation of our adaptive video segmentation algorithm. We provide example videos in supplementary material. Hyojin Park 0004, Alan Yessenbayev, Tushar Singhal, Navin Kumar Adhikari, Yizhe Zhang 0001, Shubhankar Borse, Frank Mayer, Balaji Calidas, Nilesh Prasad Pandey, Fatih Porikli |
CVPR | 12 |
| 2022 | IRISformer: Dense Vision Transformers for Single-Image Inverse Rendering in Indoor ScenesabstractIndoor scenes exhibit significant appearance variations due to myriad interactions between arbitrarily diverse object shapes, spatially-changing materials, and complex lighting. Shadows, highlights, and inter-reflections caused by visible and invisible light sources require reasoning about long-range interactions for inverse rendering, which seeks to recover the components of image formation, namely, shape, material, and lighting. In this work, our intuition is that the long-range attention learned by transformer architectures is ideally suited to solve longstanding challenges in single-image inverse rendering. We demonstrate with a specific instantiation of a dense vision transformer, IRISformer, that excels at both single-task and multi-task reasoning required for inverse rendering. Specifically, we propose a transformer architecture to simultaneously estimate depths, normals, spatially-varying albedo, roughness and lighting from a single image of an indoor scene. Our extensive evaluations on benchmark datasets demonstrate state-of-the-art results on each of the above tasks, enabling applications like object insertion and material editing in a single unconstrained real image, with greater photorealism than prior works. Code and data are publicly released.11https://github.com/ViLab-UCSD/IRISformer Rui Zhu 0026, Zhengqin Li, Janarbek Matai, Fatih Porikli, Manmohan Krishna Chandraker |
CVPR | 4 |
| 2022 | SALISA: Saliency-Based Input Sampling for Efficient Video Object Detection
Babak Ehteshami Bejnordi, AmirHossein Habibian, Fatih Porikli, Amir Ghodrati |
ECCV (10) | 3 |
| 2022 | Delta Distillation for Efficient Video Processing
AmirHossein Habibian, Haitam Ben Yahia, Davide Abati, Efstratios Gavves, Fatih Porikli |
ECCV (35) | 5 |
| 2022 | Learning Implicit Feature Alignment Function for Semantic Segmentation
Hanzhe Hu, Yinbo Chen, Shubhankar Borse, Fatih Porikli, Xiaolong Wang 0004 |
ECCV (29) | 6 |
| 2022 | Online Adaptive Personalization for Face Anti-SpoofingabstractFace authentication systems require a robust anti-spoofing module as they can be deceived by fabricating spoof images of authorized users. Most recent face anti-spoofing methods rely on optimized architectures and training objectives to alleviate the distribution shift between train and test users. However, in real online scenarios, past data from a user contains valuable information that could be used to alleviate the distribution shift. We thus introduce OAP (Online Adaptive Personalization): a lightweight solution which can adapt the model online using unlabeled data. OAP can be applied on top of most anti-spoofing methods without the need to store original biometric images. Through experimental evaluation on the SiW dataset, we show that OAP improves recognition performance of existing methods on both single video setting and continual setting, where spoof videos are interleaved with live ones to simulate spoofing attacks. We also conduct ablation studies to confirm the design choices for our solution. Davide Belli, Debasmit Das, Bence Major, Fatih Porikli |
ICIP | 4 |
| 2022 | ConFeSS: A Framework for Single Source Cross-Domain Few-Shot Learning
Debasmit Das, Sungrack Yun, Fatih Porikli |
ICLR | 3 |
| 2022 | On Improving Adversarial Transferability of Vision Transformers
Muzammal Naseer, Kanchana Ranasinghe, Salman Khan 0001, Fahad Shahbaz Khan, Fatih Porikli |
ICLR | 5 |
| 2022 | Dynamic Iterative Refinement for Efficient 3D Hand Pose EstimationabstractWhile hand pose estimation is a critical component of most interactive extended reality and gesture recognition systems, contemporary approaches are not optimized for computational and memory efficiency. In this paper, we propose a tiny deep neural network of which partial layers are recursively exploited for refining its previous estimations. During its iterative refinements, we employ learned gating criteria to decide whether to exit from the weight-sharing loop, allowing per-sample adaptation in our model. Our network is trained to be aware of the uncertainty in its current predictions to efficiently gate at each iteration, estimating variances after each loop for its keypoint estimates. Additionally, we investigate the effectiveness of end-to-end and progressive training protocols for our recursive structure on maximizing the model capacity. With the proposed setting, our method consistently outperforms state-of-the-art 2D/3D hand pose estimation approaches in terms of both accuracy and efficiency for widely used benchmarks. John Yang 0001, Yash Bhalgat, Simyung Chang, Fatih Porikli, Nojun Kwak |
WACV | 4 |
| 2022 | AuxAdapt: Stable and Efficient Test-Time Adaptation for Temporally Consistent Video Semantic SegmentationabstractIn video segmentation, generating temporally consistent results across frames is as important as achieving frame-wise accuracy. This paper presents an efficient, intuitive, and unsupervised online adaptation method, AuxAdapt, for improving the temporal consistency of most neural network models. It does not require optical flow and only takes one pass of the video. Since inconsistency mainly arises from the model’s uncertainty in its output, we propose an adaptation scheme where the model learns from its own segmentation decisions as it streams a video, which allows producing more confident and temporally consistent labeling for similarly-looking pixels across frames. For stability and efficiency, we leverage a small auxiliary segmentation network (AuxNet) to assist with this adaptation. More specifically, AuxNet readjusts the decision of the original segmentation network (Main-Net) by adding its own estimations to that of MainNet. At every frame, only AuxNet is updated via back-propagation while keeping MainNet fixed. We extensively evaluate our test-time adaptation approach on standard video benchmarks, including Cityscapes, CamVid, and KITTI. The results demonstrate that our approach provides label-wise accurate, temporally consistent, and computationally efficient adaptation. Yizhe Zhang 0001, Shubhankar Borse, Fatih Porikli |
WACV | 4 |
| 2022 | Perceptual Consistency in Video SegmentationabstractIn this paper, we present a novel perceptual consistency perspective on video semantic segmentation, which can capture both temporal consistency and pixel-wise correctness. Given two nearby video frames, perceptual consistency measures how much the segmentation decisions agree with the pixel correspondences obtained via matching general perceptual features. More specifically, for each pixel in one frame, we find the most perceptually correlated pixel in the other frame. Our intuition is that such a pair of pixels are highly likely to belong to the same class. Next, we assess how much the segmentation agrees with such perceptual correspondences, based on which we derive the perceptual consistency of the segmentation maps across these two frames. Utilizing perceptual consistency, we can evaluate the temporal consistency of video segmentation by measuring the perceptual consistency over consecutive pairs of segmentation maps in a video. Furthermore, given a sparsely labeled test video, perceptual consistency can be utilized to aid with predicting the pixel-wise correctness of the segmentation on an unlabeled frame. More specifically, by measuring the perceptual consistency between the predicted segmentation and the available ground truth on a nearby frame and combining it with the segmentation confidence, we can accurately assess the classification correctness on each pixel. Our experiments show that the proposed perceptual consistency can more accurately evaluate the temporal consistency of video segmentation as compared to flow-based measures. Furthermore, it can help more confidently predict segmentation accuracy on unlabeled test frames, as compared to using classification confidence alone. Finally, our proposed measure can be used as a regularizer during the training of segmentation models, which leads to more temporally consistent video segmentation while maintaining accuracy. Yizhe Zhang 0001, Shubhankar Borse, Ying Wang 0051, Ning Bi, Xiaoyun Jiang, Fatih Porikli |
WACV | 7 |
| 2022 | Editorial: Human visual saliency and artificial neural attention in deep learning
Wenguan Wang, Ming-Ming Cheng, Haibin Ling, Fatih Porikli |
Neurocomputing | 4 |
| 2022 | Image Segmentation Using Deep Learning: A SurveyabstractImage segmentation is a key task in computer vision and image processing with important applications such as scene understanding, medical image analysis, robotic perception, video surveillance, augmented reality, and image compression, among others, and numerous segmentation algorithms are found in the literature. Against this backdrop, the broad success of deep learning (DL) has prompted the development of new image segmentation approaches leveraging DL models. We provide a comprehensive review of this recent literature, covering the spectrum of pioneering efforts in semantic and instance segmentation, including convolutional pixel-labeling networks, encoder-decoder architectures, multiscale and pyramid-based approaches, recurrent networks, visual attention models, and generative models in adversarial settings. We investigate the relationships, strengths, and challenges of these DL-based segmentation models, examine the widely used datasets, compare performances, and discuss promising research directions. Shervin Minaee, Yuri Boykov, Fatih Porikli, Antonio Plaza, Nasser Kehtarnavaz, Demetri Terzopoulos |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Underwater Image Enhancement With Hyper-Laplacian Reflectance PriorsabstractUnderwater image enhancement aims at improving the visibility and eliminating color distortions of underwater images degraded by light absorption and scattering in water. Recently, retinex variational models show remarkable capacity of enhancing images by estimating reflectance and illumination in a retinex decomposition course. However, ambiguous details and unnatural color still challenge the performance of retinex variational models on underwater image enhancement. To overcome these limitations, we propose a hyper-laplacian reflectance priors inspired retinex variational model to enhance underwater images. Specifically, the hyper-laplacian reflectance priors are established with thel1/2-norm penalty on first-order and second-order gradients of the reflectance. Such priors exploit sparsity-promoting and complete-comprehensive reflectance that is used to enhance both salient structures and fine-scale details and recover the naturalness of authentic colors. Besides, thel2norm is found to be suitable for accurately estimating the illumination. As a result, we turn a complex underwater image enhancement issue into simple subproblems that separately and simultaneously estimate the reflection and the illumination that are harnessed to enhance underwater images in a retinex variational model. We mathematically analyze and solve the optimal solution of each subproblem. In the optimization course, we develop an alternating minimization algorithm that is efficient on element-wise operations and independent of additional prior knowledge of underwater conditions. Extensive experiments demonstrate the superiority of the proposed method in both subjective results and objective assessments over existing methods. Peixian Zhuang, Fatih Porikli, Chongyi Li |
IEEE Trans. Image Process. | 3 |
| 2021 | HS3: Learning with Proper Task Complexity in Hierarchically Supervised Semantic Segmentation
Shubhankar Borse, Yizhe Zhang 0001, Fatih Porikli |
BMVC | 4 |
| 2021 | X-Distill: Improving Self-Supervised Monocular Depth via Cross-Task Distillation
Janarbek Matai, Shubhankar Borse, Yizhe Zhang 0001, Amin Ansari, Fatih Porikli |
BMVC | 6 |
| 2021 | Conditional Model Selection for Efficient Video Understanding
Mihir Jain, Haitam Ben Yahia, Amir Ghodrati, AmirHossein Habibian, Fatih Porikli |
BMVC | 5 |
| 2021 | InverseForm: A Loss Function for Structured Boundary-Aware SegmentationabstractWe present a novel boundary-aware loss term for semantic segmentation using an inverse-transformation network, which efficiently learns the degree of parametric transformations between estimated and target boundaries. This plug-in loss term complements the cross-entropy loss in capturing boundary transformations and allows consistent and significant performance improvement on segmentation backbone models without increasing their size and computational complexity. We analyze the quantitative and qualitative effects of our loss function on three indoor and outdoor segmentation benchmarks, including Cityscapes, NYU-Depth-v2, and PASCAL, integrating it into the training phase of several backbone networks in both single-task and multi-task settings. Our extensive experiments show that the proposed method consistently outperforms base-lines, and even sets the new state-of-the-art on two datasets. Shubhankar Borse, Ying Wang 0051, Yizhe Zhang 0001, Fatih Porikli |
CVPR | 4 |
| 2021 | Efficient Action Recognition via Dynamic Knowledge PropagationabstractEfficient action recognition has become crucial to extend the success of action recognition to many real-world applications. Contrary to most existing methods, which mainly focus on selecting salient frames to reduce the computation cost, we focus more on making the most of the selected frames. To this end, we employ two networks of different capabilities that operate in tandem to efficiently recognize actions. Given a video, the lighter network processes more frames while the heavier one only processes a few. In order to enable the effective interaction between the two, we propose dynamic knowledge propagation based on a cross-attention mechanism. This is the main component of our framework that is essentially a student-teacher architecture, but as the teacher model continues to interact with the student model during inference, we call it a dynamic student-teacher framework. Through extensive experiments, we demonstrate the effectiveness of each component of our framework. Our method outperforms competing state-of-the-art methods on two video datasets: ActivityNet-v1.3 and Mini-Kinetics. Hanul Kim 0001, Mihir Jain, Juntae Lee, Sungrack Yun, Fatih Porikli |
ICCV | 5 |
| 2021 | On Generating Transferable Targeted PerturbationsabstractWhile the untargeted black-box transferability of adversarial perturbations has been extensively studied before, changing an unseen model’s decisions to a specific ‘targeted’ class remains a challenging feat. In this paper, we propose a new generative approach for highly transferable targeted perturbations (TTP). We note that the existing methods are less suitable for this task due to their reliance on class-boundary information that changes from one model to another, thus reducing transferability. In contrast, our approach matches the perturbed image ‘distribution’ with that of the target class, leading to high targeted transferability rates. To this end, we propose a new objective function that not only aligns the global distributions of source and target images, but also matches the local neighbourhood structure between the two domains. Based on the proposed objective, we train a generator function that can adaptively synthesize perturbations specific to a given input. Our generative approach is in-dependent of the source or target domain labels, while consistently performs well against state-of-the-art methods on a wide range of attack settings. As an example, we achieve 32.63% target transferability from (an adversarially weak) VGG19BNto (a strong) WideResNet on ImageNet val. set, which is 4× higher than the previous best generative attack and 16× better than instance-specific iterative attack. Code is available at: https://github.com/Muzammal-Naseer/TTP. Muzammal Naseer, Salman Khan 0001, Munawar Hayat, Fahad Shahbaz Khan, Fatih Porikli |
ICCV | 5 |
| 2021 | Modality-Agnostic Topology Aware LocalizationabstractThis work presents a data-driven approach for the indoor localization of an observer on a 2D topological map of the environment. State-of-the-art techniques may yield accurate estimates only when they are tailor-made for a specific data modality like camera-based system that prevents their applicability to broader domains. Here, we establish a modality-agnostic framework (called OT-Isomap) and formulate the localization problem in the context of parametric manifold learning while leveraging optimal transportation. This framework allows jointly learning a low-dimensional embedding as well as correspondences with a topological map. We examine the generalizability of the proposed algorithm by applying it to data from diverse modalities such as image sequences and radio frequency signals. The experimental results demonstrate decimeter-level accuracy for localization using different sensory inputs. Farhad G. Zanjani, Ilia Karmanov, Hanno Ackermann, Daniel Dijkman, Simone Merlin, Max Welling, Fatih Porikli |
NeurIPS | 7 |
| 2021 | Sketch-specific data augmentation for freehand sketch recognition
Ying Zheng 0009, Hongxun Yao, Xiaoshuai Sun, Shengping Zhang, Sicheng Zhao, Fatih Porikli |
Neurocomputing | 6 |
| 2021 | Deblur and deep depth from single defocus image
Saeed Anwar, Zeeshan Hayder, Fatih Porikli |
Mach. Vis. Appl. | 3 |
| 2021 | Dynamical Hyperparameter Optimization via Deep Reinforcement Learning in TrackingabstractHyperparameters are numerical pre-sets whose values are assigned prior to the commencement of a learning process. Selecting appropriate hyperparameters is often critical for achieving satisfactory performance in many vision problems, such as deep learning-based visual object tracking. However, it is often difficult to determine their optimal values, especially if they are specific to each video input. Most hyperparameter optimization algorithms tend to search a generic range and are imposed blindly on all sequences. In this paper, we propose a novel dynamical hyperparameter optimization method that adaptively optimizes hyperparameters for a given sequence using an action-prediction network leveraged on continuous deep Q-learning. Since the observation space for object tracking is significantly more complex than those in traditional control problems, existing continuous deep Q-learning algorithms cannot be directly applied. To overcome this challenge, we introduce an efficient heuristic strategy to handle high dimensional state space, while also accelerating the convergence behavior. The proposed algorithm is applied to improve two representative trackers, a Siamese-based one and a correlation-filter-based one, to evaluate its generalizability. Their superior performances on several popular benchmarks are clearly demonstrated. Our source code is available at https://github.com/shenjianbing/dqltracking. Xingping Dong, Jianbing Shen, Wenguan Wang, Ling Shao 0001, Haibin Ling, Fatih Porikli |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2021 | Deep ancient Roman Republican coin classification via feature fusion and attention
Hafeez Anwar, Saeed Anwar, Sebastian Zambanini, Fatih Porikli |
Pattern Recognit. | 4 |
| 2021 | Knowledge memorization and generation for action recognition in still images
Wankou Yang, Yazhou Yao, Fatih Porikli |
Pattern Recognit. | 4 |
| 2020 | Multi-Mutual Consistency Induced Transfer Subspace Learning for Human Motion SegmentationabstractHuman motion segmentation based on transfer subspace learning is a rising interest in action-related tasks. Although progress has been made, there are still several issues within the existing methods. First, existing methods transfer knowledge from source data to target tasks by learning domain-invariant features, but they ignore to preserve domain-specific knowledge. Second, the transfer subspace learning is employed in either low-level or high-level feature spaces, but few methods consider fusing multi-level features for subspace learning. To this end, we propose a novel multi-mutual consistency induced transfer subspace learning framework for human motion segmentation. Specifically, our model factorizes the source and target data into distinct multi-layer feature spaces and reduces the distribution gap between them through a multi-mutual consistency learning strategy. In this way, the domain-specific knowledge and domain-invariant properties can be explored simultaneously. Our model also conducts the transfer subspace learning on different layers to capture multi-level structural information. Further, to preserve the temporal correlations, we project the learned representations into a block-like space. The proposed model is efficiently optimized by using the Augmented Lagrange Multiplier (ALM) algorithm. Experimental results on four human motion datasets demonstrate the effectiveness of our method over other state-of-the-art approaches. Tao Zhou 0002, Huazhu Fu, Chen Gong 0002, Jianbing Shen, Ling Shao 0001, Fatih Porikli |
CVPR | 6 |
| 2020 | A Self-supervised Approach for Adversarial RobustnessabstractAdversarial examples can cause catastrophic mistakes in Deep Neural Network (DNNs) based vision systems e.g., for classification, segmentation and object detection. The vulnerability of DNNs against such attacks can prove a major roadblock towards their real-world deployment. Transferability of adversarial examples demand generalizable defenses that can provide cross-task protection. Adversarial training that enhances robustness by modifying target model's parameters lacks such generalizability. On the other hand, different input processing based defenses fall short in the face of continuously evolving attacks. In this paper, we take the first step to combine the benefits of both approaches and propose a self-supervised adversarial training mechanism in the input space. By design, our defense is a generalizable approach and provides significant robustness against the unseen adversarial attacks (e.g. by reducing the success rate of translation-invariant ensemble attack from 82.6% to 31.9% in comparison to previous stateof-the-art). It can be deployed as a plug-and-play solution to protect a variety of vision systems, as we demonstrate for the case of classification, segmentation and detection. Code is available at: https://github.com/ Muzammal-Naseer/NRP. Muzammal Naseer, Salman Khan 0001, Munawar Hayat, Fahad Shahbaz Khan, Fatih Porikli |
CVPR | 5 |
| 2020 | CLNet: A Compact Latent Network for Fast Adjusting Siamese Trackers
Xingping Dong, Jianbing Shen, Ling Shao 0001, Fatih Porikli |
ECCV (20) | 4 |
| 2020 | Structured Convolutions for Efficient Neural Network DesignabstractIn this work, we tackle model efficiency by exploiting redundancy in the implicit structure of the building blocks of convolutional neural networks. We start our analysis by introducing a general definition of Composite Kernel structures that enable the execution of convolution operations in the form of efficient, scaled, sum-pooling components. As its special case, we propose Structured Convolutions and show that these allow decomposition of the convolution operation into a sum-pooling operation followed by a convolution with significantly lower complexity and fewer weights. We show how this decomposition can be applied to 2D and 3D kernels as well as the fully-connected layers. Furthermore, we present a Structural Regularization loss that promotes neural network layers to leverage on this desired structure in a way that, after training, they can be decomposed with negligible performance loss. By applying our method to a wide range of CNN architectures, we demonstrate 'structured' versions of the ResNets that are up to 2x smaller and a new Structured-MobileNetV2 that is more efficient while staying within an accuracy loss of 1% on ImageNet and CIFAR-10 datasets. We also show similar structured versions of EfficientNet on ImageNet and HRNet architecture for semantic segmentation on the Cityscapes dataset. Our method performs equally well or superior in terms of the complexity reduction in comparison to the existing tensor decomposition and channel pruning methods. Yash Bhalgat, Yizhe Zhang 0001, Jamie Menjay Lin, Fatih Porikli |
NeurIPS | 4 |
| 2020 | Component Attention Guided Face Super-Resolution Network: CAGFaceabstractTo make the best use of the underlying structure of faces, the collective information through face datasets and the intermediate estimates during the upsampling process, here we introduce a fully convolutional multi-stage neural network for 4× super-resolution for face images. We implicitly impose facial component-wise attention maps using a segmentation network to allow our network to focus on face-inherent patterns. Each stage of our network is composed of a stem layer, a residual backbone, and spatial upsampling layers. We recurrently apply stages to reconstruct an intermediate image, and then reuse its space-to-depth converted versions to bootstrap and enhance image quality progressively. Our experiments show that our face super-resolution method achieves quantitatively superior and perceptually pleasing results in comparison to state of the art. Ratheesh Kalarot, Fatih Porikli |
WACV | 3 |
| 2020 | Zero-Shot Object Detection: Joint Recognition and Localization of Novel Concepts
Shafin Rahman, Salman Khan 0001, Fatih Porikli |
Int. J. Comput. Vis. | 3 |
| 2020 | Hallucinating Unaligned Face Images by Multiscale Transformative Discriminative Networks
Xin Yu 0002, Fatih Porikli, Basura Fernando, Richard I. Hartley |
Int. J. Comput. Vis. | 2 |
| 2020 | Robust visual tracking with channel attention and focal loss
Dongdong Li 0004, GongJian Wen, Yangliu Kuai, Lingxiao Zhu, Fatih Porikli |
Neurocomputing | 5 |
| 2020 | Cost-sensitive joint feature and dictionary learning for face recognition
Guoqing Zhang 0002, Fatih Porikli, Huaijiang Sun, Quan-Sen Sun, Guiyu Xia, Yuhui Zheng |
Neurocomputing | 2 |
| 2020 | Semantic Face Hallucination: Super-Resolving Very Low-Resolution Face Images with Supplementary AttributesabstractGiven a tiny face image, existing face hallucination methods aim at super-resolving its high-resolution (HR) counterpart by learning a mapping from an exemplary dataset. Since a low-resolution (LR) input patch may correspond to many HR candidate patches, this ambiguity may lead to distorted HR facial details and wrong attributes such as gender reversal and rejuvenation. An LR input contains low-frequency facial components of its HR version while its residual face image, defined as the difference between the HR ground-truth and interpolated LR images, contains the missing high-frequency facial details. We demonstrate that supplementing residual images or feature maps with additional facial attribute information can significantly reduce the ambiguity in face super-resolution. To explore this idea, we develop an attribute-embedded upsampling network, which consists of an upsampling network and a discriminative network. The upsampling network is composed of an autoencoder with skip-connections, which incorporates facial attribute vectors into the residual features of LR inputs at the bottleneck of the autoencoder, and deconvolutional layers used for upsampling. The discriminative network is designed to examine whether super-resolved faces contain the desired attributes or not and then its loss is used for updating the upsampling network. In this manner, we can super-resolve tiny (16×16 pixels) unaligned face images with a large upscaling factor of 8× while reducing the uncertainty of one-to-many mappings remarkably. By conducting extensive evaluations on a large-scale dataset, we demonstrate that our method achieves superior face hallucination results and outperforms the state-of-the-art. Xin Yu 0002, Basura Fernando, Richard I. Hartley, Fatih Porikli |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Can We See More? Joint Frontalization and Hallucination of Unaligned Tiny FacesabstractIn popular TV programs (such as CSI), a very low-resolution face image of a person, who is not even looking at the camera in many cases, is digitally super-resolved to a degree that suddenly the person's identity is made visible and recognizable. Of course, we suspect that this is merely a cinematographic special effect and such a magical transformation of a single image is not technically possible. Or, is it? In this paper, we push the boundaries of super-resolving (hallucinating to be more accurate) a tiny, non-frontal face image to understand how much of this is possible by leveraging the availability of large datasets and deep networks. To this end, we introduce a novel Transformative Adversarial Neural Network (TANN) to jointly frontalize very-low resolution (i.e., 16 × 16 pixels) out-of-plane rotated face images (including profile views) and aggressively super-resolve them (8×), regardless of their original poses and without using any 3D information. TANN is composed of two components: a transformative upsampling network which embodies encoding, spatial transformation and deconvolutional layers, and a discriminative network that enforces the generated high-resolution frontal faces to lie on the same manifold as real frontal face images. We evaluate our method on a large set of synthesized non-frontal face images to assess its reconstruction performance. Extensive experiments demonstrate that TANN generates both qualitatively and quantitatively superior results achieving over 4 dB improvement over the state-of-the-art. Xin Yu 0002, Fatemeh Shiri, Bernard Ghanem, Fatih Porikli |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Underwater scene prior inspired deep underwater image and video enhancement
Chongyi Li, Saeed Anwar, Fatih Porikli |
Pattern Recognit. | 3 |
| 2020 | Feature mask network for person re-identification
Guodong Ding, Salman Khan 0001, Zhenmin Tang, Fatih Porikli |
Pattern Recognit. Lett. | 4 |
| 2020 | When Correlation Filters Meet Siamese Networks for Real-Time Complementary TrackingabstractDiscriminative correlation filter (DCF)-based trackers have recently exhibited high efficiency and impressive robustness to challenging factors, such as illumination change and partial occlusion. However, in cases with fast motion and full occlusion, these trackers drift off soon and can hardly re-detect the target from the restricted search region due to the boundary effect. On the contrary, recent work using a fully convolutional Siamese network (Siamfc) locates the exemplar image within a large search image but suffers from coarse location and distractors. In this paper, we propose a real-time complementary tracker (RCT) by integrating DCF and Siamfc into a two-stage tracking framework where DCF and Siamfc share mutual advantages and complement each other. In the first stage of this framework, RCT locates the target coarsely but robustly with Siamfc. In the second stage, the derived coarse location is refined by DCF for higher accuracy. For efficiency reasons, Siamfc in the first stage is activated occasionally based on the tracking status inferred from the correlation response map of DCF in the second stage. Comprehensive experiments are performed on three popular benchmark datasets: OTB2013, OTB2015, and VOT2016. On OTB2013, RCT runs with over 40 f/s and achieves an absolute gain of 4.8% and 5.2% in mean overlap precision compared with two base trackers (Staple and Siamfc). On VOT2016, RCT makes a good balance between performance and efficiency, ranking fifth in EAO and first in EFO compared with the top five trackers. Dongdong Li 0004, Fatih Porikli, GongJian Wen, Yangliu Kuai |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Scale-Aware Crowd Counting via Depth-Embedded Convolutional Neural NetworksabstractScale variation of pedestrians in a crowd image presents a significant challenge for vision-based people counting systems. Such variations are mainly caused by perspective-related distortions due to the camera pose relative to the ground plane. Following the density-based counting paradigm, we postulate that generating density values adaptive to object scales plays a critical role in the accuracy of the final counting results. Motivated by this, we distill the underlying information from depth cues to obtain scale-aware representations that can respond to object scales considering the fact that the scale is inversely proportional to the object depth. Specifically, we propose a depth embedding module as add-ons into existing networks. This module exploits essential depth cues to spatially re-calibrate the magnitude of the original features. In this way, the objects, although in the same class, will attain distinct representations according to their scales, which directly benefits the estimation of scale-aware density values. We conduct a comprehensive analysis of the effects of the depth embedding module and validate that exploiting depth cues to perceive object scale variations in convolutional neural networks improves crowd counting performances. Our experiments demonstrate the effectiveness of the proposed approach on four popular benchmark datasets. Muming Zhao, Jian Zhang 0002, Fatih Porikli, Bingbing Ni, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Joint Stereo Video Deblurring, Scene Flow Estimation and Moving Object SegmentationabstractStereo videos for the dynamic scenes often show unpleasant blurred effects due to the camera motion and the multiple moving objects with large depth variations. Given consecutive blurred stereo video frames, we aim to recover the latent clean images, estimate the 3D scene flow and segment the multiple moving objects. These three tasks have been previously addressed separately, which fail to exploit the internal connections among these tasks and cannot achieve optimality. In this paper, we propose to jointly solve these three tasks in a unified framework by exploiting their intrinsic connections. To this end, we represent the dynamic scenes with the piece-wise planar model, which exploits the local structure of the scene and expresses various dynamic scenes. Under our model, these three tasks are naturally connected and expressed as the parameter estimation of 3D scene structure and camera motion (structure and motion for the dynamic scenes). By exploiting the blur model constraint, the moving objects and the 3D scene structure, we reach an energy minimization formulation for joint deblurring, scene flow and segmentation. We evaluate our approach extensively on both synthetic datasets and publicly available real datasets with fast-moving objects, camera motion, uncontrolled lighting conditions and shadows. Experimental results demonstrate that our method can achieve significant improvement in stereo video deblurring, scene flow estimation and moving object segmentation, over state-of-the-art methods. Liyuan Pan, Yuchao Dai, Miaomiao Liu 0001, Fatih Porikli, Quan Pan 0001 |
IEEE Trans. Image Process. | 4 |
| 2019 | See More, Know More: Unsupervised Video Object Segmentation With Co-Attention Siamese NetworksabstractWe introduce a novel network, called as CO-attention Siamese Network (COSNet), to address the unsupervised video object segmentation task from a holistic view. We emphasize the importance of inherent correlation among video frames and incorporate a global co-attention mechanism to improve further the state-of-the-art deep learning based solutions that primarily focus on learning discriminative foreground representations over appearance and motion in short-term temporal segments. The co-attention layers in our network provide efficient and competent stages for capturing global correlations and scene context by jointly computing and appending co-attention responses into a joint feature space. We train COSNet with pairs of video frames, which naturally augments training data and allows increased learning capacity. During the segmentation stage, the co-attention model encodes useful information by processing multiple reference frames together, which is leveraged to infer the frequently reappearing and salient foreground objects better. We propose a unified and end-to-end trainable framework where different co-attention variants can be derived for mining the rich context within videos. Our extensive experiments over three large benchmarks manifest that COSNet outperforms the current alternatives by a large margin. We will publicly release our implementation and models. Xiankai Lu, Wenguan Wang, Chao Ma 0004, Jianbing Shen, Ling Shao 0001, Fatih Porikli |
CVPR | 6 |
| 2019 | Cross-Domain Transferability of Adversarial PerturbationsabstractAdversarial examples reveal the blind spots of deep neural networks (DNNs) and represent a major concern for security-critical applications. The transferability of adversarial examples makes real-world attacks possible in black-box settings, where the attacker is forbidden to access the internal parameters of the model. The underlying assumption in most adversary generation methods, whether learning an instance-specific or an instance-agnostic perturbation, is the direct or indirect reliance on the original domain-specific data distribution. In this work, for the first time, we demonstrate the existence of domain-invariant adversaries, thereby showing common adversarial space among different datasets and models. To this end, we propose a framework capable of launching highly transferable attacks that crafts adversarial patterns to mislead networks trained on wholly different domains. For instance, an adversarial function learned on Paintings, Cartoons or Medical images can successfully perturb ImageNet samples to fool the classifier, with success rates as high as $\sim$99\% ($\ell_{\infty} \le 10$). The core of our proposed adversarial function is a generative network that is trained using a relativistic supervisory signal that enables domain-invariant perturbations. Our approach sets the new state-of-the-art for fooling rates, both under the white-box and black-box scenarios. Furthermore, despite being an instance-agnostic perturbation function, our attack outperforms the conventionally much stronger instance-specific attack methods. Muzammal Naseer, Salman Khan 0001, Muhammad Haris Khan, Fahad Shahbaz Khan, Fatih Porikli |
NeurIPS | 5 |
| 2019 | Local Gradients Smoothing: Defense Against Localized Adversarial AttacksabstractDeep neural networks (DNNs) have shown vulnerability to adversarial attacks, i.e., carefully perturbed inputs designed to mislead the network at inference time. Recently introduced localized attacks, Localized and Visible Adversarial Noise (LaVAN) and Adversarial patch, pose a new challenge to deep learning security by adding adversarial noise only within a specific region without affecting the salient objects in an image. Driven by the observation that such attacks introduce concentrated high-frequency changes at a particular image location, we have developed an effective method to estimate noise location in gradient domain and transform those high activation regions caused by adversarial noise in image domain while having minimal effect on the salient object that is important for correct classification. Our proposed Local Gradients Smoothing (LGS) scheme achieves this by regularizing gradients in the estimated noisy region before feeding the image to DNN for inference. We have shown the effectiveness of our method in comparison to other defense methods including Digital Watermarking, JPEG compression, Total Variance Minimization (TVM) and Feature squeezing on ImageNet dataset. In addition, we systematically study the robustness of the proposed defense mechanism against Back Pass Differentiable Approximation (BPDA), a state of the art attack recently developed to break defenses that transform an input sample to minimize the adversarial effect. Compared to other defense mechanisms, LGS is by far the most resistant to BPDA in localized adversarial attack setting. Muzammal Naseer, Salman Khan 0001, Fatih Porikli |
WACV | 3 |
| 2019 | Recovering Faces From Portraits with Auxiliary Facial AttributesabstractRecovering a photorealistic face from an artistic portrait is a challenging task since crucial facial details are often distorted or completely lost in artistic compositions. To handle this loss, we propose an Attribute-guided Face Recovery from Portraits (AFRP) that utilizes a Face Recovery Network (FRN) and a Discriminative Network (DN). FRN consists of an autoencoder with residual block-embedded skip-connections and incorporates facial attribute vectors into the feature maps of input portraits at the bottleneck of the autoencoder. DN has multiple convolutional and fully-connected layers, and its role is to enforce FRN to generate authentic face images with corresponding facial attributes dictated by the input attribute vectors. For the preservation of identities, we impose the recovered and ground-truth faces to share similar visual features. Specifically, DN determines whether the recovered image looks like a real face and checks if the facial attributes extracted from the recovered image are consistent with given attributes. Our method can recover photorealistic identity-preserving faces with desired attributes from unseen stylized portraits, artistic paintings, and hand-drawn sketches. On large-scale synthesized and sketch datasets, we demonstrate that our face recovery method achieves state-of-the-art results. Fatemeh Shiri, Xin Yu 0002, Fatih Porikli, Richard I. Hartley, Piotr Koniusz |
WACV | 3 |
| 2019 | Identity-Preserving Face Recovery from Stylized Portraits
Fatemeh Shiri, Xin Yu 0002, Fatih Porikli, Richard I. Hartley, Piotr Koniusz |
Int. J. Comput. Vis. | 3 |
| 2019 | Learning target-aware correlation filters for visual tracking
Dongdong Li 0004, GongJian Wen, Yangliu Kuai, Fatih Porikli |
J. Vis. Commun. Image Represent. | 5 |
| 2019 | Beyond feature integration: a coarse-to-fine framework for cascade correlation tracking
Dongdong Li 0004, GongJian Wen, Yangliu Kuai, Fatih Porikli |
Mach. Vis. Appl. | 4 |
| 2019 | Regularization of deep neural networks with spectral dropout
Salman Khan 0001, Munawar Hayat, Fatih Porikli |
Neural Networks | 3 |
| 2019 | Image Deblurring with a Class-Specific PriorabstractA fundamental problem in image deblurring is to recover reliably distinct spatial frequencies that have been suppressed by the blur kernel. To tackle this issue, existing image deblurring techniques often rely on generic image priors such as the sparsity of salient features including image gradients and edges. However, these priors only help recover part of the frequency spectrum, such as the frequencies near the high-end. To this end, we pose the following specific questions: (i) Does any image class information offer an advantage over existing generic priors for image quality restoration? (ii) If a class-specific prior exists, how should it be encoded into a deblurring framework to recover attenuated image frequencies? Throughout this work, we devise a class-specific prior based on the band-pass filter responses and incorporate it into a deblurring strategy. More specifically, we show that the subspace of band-pass filtered images and their intensity distributions serve as useful priors for recovering image frequencies that are difficult to recover by generic image priors. We demonstrate that our image deblurring framework, when equipped with the above priors, significantly outperforms many state-of-the-art methods using generic image priors or class-specific exemplars. Saeed Anwar, Cong Phuoc Huynh, Fatih Porikli |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Semi-Supervised Video Object Segmentation with Super-TrajectoriesabstractWe introduce a semi-supervised video segmentation approach based on an efficient video representation, called as "super-trajectory". A super-trajectory corresponds to a group of compact point trajectories that exhibit consistent motion patterns, similar appearances, and close spatiotemporal relationships. We generate the compact trajectories using a probabilistic model, which enables handling of occlusions and drifts effectively. To reliably group point trajectories, we adopt a modified version of the density peaks based clustering algorithm that allows capturing rich spatiotemporal relations among trajectories in the clustering process. We incorporate two intuitive mechanisms for segmentation, called as reverse-tracking and object re-occurrence, for robustness and boosting the performance. Building on the proposed video representation, our segmentation method is discriminative enough to accurately propagate the initial annotations in the first frame onto the remaining frames. Our extensive experimental analyses on three challenging benchmarks demonstrate that, given the annotation in the first frame, our method is capable of extracting the target objects from complex backgrounds, and even reidentifying them after prolonged occlusions, producing high-quality video object segments. Wenguan Wang, Jianbing Shen, Fatih Porikli, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Real-Time Deep Tracking via Corrective Domain AdaptationabstractVisual tracking is one of the fundamental problems in computer vision. Recently, some deep-learning-based tracking algorithms have been illustrating record-breaking performances. However, due to the high complexity of neural networks, most deep trackers suffer from low tracking speed and are, thus, impractical in many real-world applications. Some recently proposed deep trackers with smaller network structure achieve high efficiency while at the cost of significant decrease in precision. In this paper, we propose to transfer the deep feature, which is learned originally for image classification to the visual tracking domain. The domain adaptation is achieved via some “grafted” auxiliary networks, which are trained by regressing the object location in tracking frames. This adaptation improves the tracking performance significantly both on accuracy and efficiency. The yielded deep tracker is real time and also illustrates the state-of-the-art accuracies in the experiment involving two well-adopted benchmarks with more than 100 test videos. Furthermore, the adaptation is also naturally used for introducing the objectness concept into visual tracking. This removes a long-standing target ambiguity in visual tracking tasks, and we illustrate the empirical superiority of the more well-defined task. Xinyu Wang 0010, Fumin Shen, Yi Li 0025, Fatih Porikli, Mingwen Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2019 | Introduction to the Special Section on Deep Learning for Visual SurveillanceabstractWe are now living in an era of visual information where data is unceasingly generated and pushed into consumption at astounding rates. A remarkable portion of this sensory input comes in the form of videos streaming from large-scale surveillance infrastructures as well as consumer-grade monitoring systems. The sheer amount of ground-based, aerial and mobile video surveillance data demands fittingly competent, accurate, effective techniques to extract useful cues and provide assistance for detection, prevention, and intervention tasks in traffic, safety, security, defense, forensic, health, biology, ethology, and retail space management applications. Fatih Porikli, Larry Davis 0001, Qi Wang 0009, Yi Li 0025, Carlo S. Regazzoni |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Video Saliency Detection via Sparsity-Based Reconstruction and PropagationabstractVideo saliency detection aims to continuously discover the motion-related salient objects from the video sequences. Since it needs to consider the spatial and temporal constraints jointly, video saliency detection is more challenging than image saliency detection. In this paper, we propose a new method to detect the salient objects in video based on sparse reconstruction and propagation. With the assistance of novel static and motion priors, a single-frame saliency model is first designed to represent the spatial saliency in each individual frame via the sparsity-based reconstruction. Then, through a progressive sparsity-based propagation, the sequential correspondence in the temporal space is captured to produce the inter-frame saliency map. Finally, these two maps are incorporated into a global optimization model to achieve spatio-temporal smoothness and global consistency of the salient object in the whole video. The experiments on three large-scale video saliency datasets demonstrate that the proposed method outperforms the state-of-the-art algorithms both qualitatively and quantitatively. Runmin Cong, Jianjun Lei 0001, Huazhu Fu, Fatih Porikli, Qingming Huang, Chunping Hou |
IEEE Trans. Image Process. | 4 |
| 2019 | Quadruplet Network With One-Shot Learning for Fast Visual Object TrackingabstractIn the same vein of discriminative one-shot learning, Siamese networks allow recognizing an object from a single exemplar with the same class label. However, they do not take advantage of the underlying structure of the data and the relationship among the multitude of samples as they only rely on the pairs of instances for training. In this paper, we propose a new quadruplet deep network to examine the potential connections among the training instances, aiming to achieve a more powerful representation. We design a shared network with four branches that receive a multi-tuple of instances as inputs and are connected by a novel loss function consisting of pair loss and triplet loss. According to the similarity metric, we select the most similar and the most dissimilar instances as the positive and negative inputs of triplet loss from each multi-tuple. We show that this scheme improves the training performance. Furthermore, we introduce a new weight layer to automatically select suitable combination weights, which will avoid the conflict between triplet and pair loss leading to worse performance. We evaluate our quadruplet framework by model-free tracking-by-detection of objects from a single initial exemplar in several visual object tracking benchmarks. Our extensive experimental analysis demonstrates that our tracker achieves superior performance with a real-time processing speed of 78 frames/s. Our source code is available. Xingping Dong, Jianbing Shen, Dongming Wu 0005, Kan Guo, Xiaogang Jin 0001, Fatih Porikli |
IEEE Trans. Image Process. | 6 |
| 2019 | Feature Affinity-Based Pseudo Labeling for Semi-Supervised Person Re-IdentificationabstractVision-based person re-identification aims to match a person's identity across multiple images, which is a fundamental task in multimedia content analysis and retrieval. Deep neural networks have recently manifested great potential in this task. However, a major bottleneck of existing supervised deep networks is their reliance on a large amount of annotated training data. Manual labeling for person identities in large-scale surveillance camera systems is quite challenging and incurs significant costs. Some recent studies adopt generative model outputs as training data augmentation. To more effectively use these synthetic data for an improved feature learning and re-identification performance, this paper proposes a novel feature affinity-based pseudo labeling method with two possible label encodings. To the best of our knowledge, this is the first study that employs pseudo-labeling by measuring the affinity of unlabeled samples with the underlying clusters of labeled data samples using the intermediate feature representations from deep networks. We propose training the network with the joint supervision of cross-entropy loss together with a center regularization term, which not only ensures discriminative feature representation learning but also simultaneously predicts pseudo-labels for unlabeled data. We show that both label encodings can be learned in a unified manner and help improve the overall performance. Our extensive experiments on three person re-identification datasets: Market-1501, DukeMTMC-reID, and CUHK03, demonstrate significant performance boost over the state-of-the-art person re-identification approaches. Guodong Ding, Shanshan Zhang 0001, Salman Khan 0001, Zhenmin Tang, Jian Zhang 0002, Fatih Porikli |
IEEE Trans. Multim. | 6 |
| 2019 | Robust Object Tracking Using Manifold Regularized Convolutional Neural NetworksabstractIn visual tracking, usually only a small number of samples are labeled, and most existing deep learning based trackers ignore abundant unlabeled samples that could provide additional information for deep trackers to boost their tracking performance. An intuitive way to explain unlabeled data is to incorporate manifold regularization into the common classification loss functions, but the high computational cost may prohibit those deep trackers from practical applications. To overcome this issue, we propose a two-stage approach to a deep tracker that takes into account both labeled and unlabeled samples. The annotation of unlabeled samples is propagated from its labeled neighbors first by exploring the manifold space that these samples are assumed to lie in. Then, we refine it by training a deep convolutional neural network using both labeled and unlabeled data in a supervised manner. Online visual tracking is further carried out under the framework of particle filters with the presented manifold regularized deep model being updated every few frames. Experimental results on different tracking datasets demonstrate that our tracker outperforms most existing tracking approaches. The source code and results are available at: https://github.com/shenjianbing/MRCNNTracking. Hongwei Hu, Bo Ma 0001, Jianbing Shen, Hanqiu Sun, Ling Shao 0001, Fatih Porikli |
IEEE Trans. Multim. | 6 |
| 2019 | Saliency Integration: An Arbitrator ModelabstractSaliency integration has attracted much attention on unifying saliency maps from multiple saliency models. Previous offline integration methods usually face two challenges: 1) if most of the candidate saliency models misjudge the saliency on an image, the integration result will lean heavily on those inferior candidate models; and 2) an unawareness of the ground truth saliency labels brings difficulty in estimating the expertise of each candidate model. To address these problems, in this paper, we propose an arbitrator model (AM) for saliency integration. First, we incorporate the consensus of multiple saliency models and the external knowledge into a reference map to effectively rectify the misleading by candidate models. Second, our quest for ways of estimating the expertise of the saliency models without ground truth labels gives rise to two distinct online model-expertise estimation methods. Finally, we derive a Bayesian integration framework to reconcile the saliency models of varying expertise and the reference map. To extensively evaluate the proposed AM model, we test 27 state-of-the-art saliency models, covering both traditional and deep learning ones, on various combinations over four datasets. The evaluation results show that the AM model improves the performance substantially compared to the existing state-of-the-art integration methods, regardless of the chosen candidate saliency models. Yingyue Xu, Xiaopeng Hong, Fatih Porikli, Xin Liu 0012, Jie Chen 0001, Guoying Zhao 0001 |
IEEE Trans. Multim. | 3 |
| 2019 | Submodular Function Optimization for Motion Clustering and Image SegmentationabstractIn this paper, we propose a framework of maximizing quadratic submodular energy with a knapsack constraint approximately, to solve certain computer vision problems. The proposed submodular maximization problem can be viewed as a generalization of the classic 0/1 knapsack problem. Importantly, maximization of our knapsack constrained submodular energy function can be solved via dynamic programing. We further introduce a range-reduction step prior to dynamic programing as a two-stage procedure for more efficient maximization. In order to demonstrate the effectiveness of the proposed energy function and its maximization algorithm, we apply it to two representative computer vision tasks: image segmentation and motion trajectory clustering. Experimental results of image segmentation demonstrate that our method outperforms the classic segmentation algorithms of graph cuts and random walks. Moreover, our framework achieves better performance than state-of-the-art methods on the motion trajectory clustering task. Jianbing Shen, Xingping Dong, Jianteng Peng, Xiaogang Jin 0001, Ling Shao 0001, Fatih Porikli |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2018 | Zero-Shot Object Detection: Learning to Simultaneously Recognize and Localize Novel Concepts
Shafin Rahman, Salman Khan 0001, Fatih Porikli |
ACCV (1) | 3 |
| 2018 | Hyperparameter Optimization for Tracking With Continuous Deep Q-LearningabstractHyperparameters are numerical presets whose values are assigned prior to the commencement of the learning process. Selecting appropriate hyperparameters is critical for the accuracy of tracking algorithms, yet it is difficult to determine their optimal values, in particular, adaptive ones for each specific video sequence. Most hyperparameter optimization algorithms depend on searching a generic range and they are imposed blindly on all sequences. Here, we propose a novel hyperparameter optimization method that can find optimal hyperparameters for a given sequence using an action-prediction network leveraged on Continuous Deep Q-Learning. Since the common state-spaces for object tracking tasks are significantly more complex than the ones in traditional control problems, existing Continuous Deep Q-Learning algorithms cannot be directly applied. To overcome this challenge, we introduce an efficient heuristic to accelerate the convergence behavior. We evaluate our method on several tracking benchmarks and demonstrate its superior performance1. Xingping Dong, Jianbing Shen, Wenguan Wang, Yu Liu 0074, Ling Shao 0001, Fatih Porikli |
CVPR | 6 |
| 2018 | A Deeper Look at Power NormalizationsabstractPower Normalizations (PN) are very useful non-linear operators in the context of Bag-of-Words data representations as they tackle problems such as feature imbalance. In this paper, we reconsider these operators in the deep learning setup by introducing a novel layer that implements PN for non-linear pooling of feature maps. Specifically, by using a kernel formulation, our layer combines the feature vectors and their respective spatial locations in the feature maps produced by the last convolutional layer of CNN. Linearization of such a kernel results in a positive definite matrix capturing the second-order statistics of the feature vectors, to which PN operators are applied. We study two types of PN functions, namely (i) MaxExp and (ii) Gamma, addressing their role and meaning in the context of nonlinear pooling. We also provide a probabilistic interpretation of these operators and derive their surrogates with well-behaved gradients for end-to-end CNN learning. We apply our theory to practice by implementing the PN layer on a ResNet-50 model and showcase experiments on four benchmarks for fine-grained recognition, scene recognition, and material classification. Our results demonstrate state-of-the-part performance across all these tasks. Piotr Koniusz, Fatih Porikli |
CVPR | 3 |
| 2018 | Video Representation Learning Using Discriminative PoolingabstractPopular deep models for action recognition in videos generate independent predictions for short clips, which are then pooled heuristically to assign an action label to the full video segment. As not all frames may characterize the underlying action-indeed, many are common across multiple actions-pooling schemes that impose equal importance on all frames might be unfavorable. In an attempt to tackle this problem, we propose discriminative pooling, based on the notion that among the deep features generated on all short clips, there is at least one that characterizes the action. To this end, we learn a (nonlinear) hyperplane that separates this unknown, yet discriminative, feature from the rest. Applying multiple instance learning in a large-margin setup, we use the parameters of this separating hyperplane as a descriptor for the full video segment. Since these parameters are directly related to the support vectors in a max-margin framework, they serve as robust representations for pooling of the features. We formulate a joint objective and an efficient solver that learns these hyperplanes per video and the corresponding action classifiers over the hyperplanes. Our pooling scheme is end-to-end trainable within a deep framework. We report results from experiments on three benchmark datasets spanning a variety of challenges and demonstrate state-of-the-art performance across these tasks. Jue Wang 0010, Anoop Cherian, Fatih Porikli, Stephen Gould |
CVPR | 3 |
| 2018 | One-Shot Action Localization by Learning Sequence Matching NetworkabstractLearning based temporal action localization methods require vast amounts of training data. However, such large-scale video datasets, which are expected to capture the dynamics of every action category, are not only very expensive to acquire but are also not practical simply because there exists an uncountable number of action classes. This poses a critical restriction to the current methods when the training samples are few and rare (e.g. when the target action classes are not present in the current publicly available datasets). To address this challenge, we conceptualize a new example-based action detection problem where only a few examples are provided, and the goal is to find the occurrences of these examples in an untrimmed video sequence. Towards this objective, we introduce a novel one-shot action localization method that alleviates the need for large amounts of training samples. Our solution adopts the one-shot learning technique of Matching Network and utilizes correlations to mine and localize actions of previously unseen classes. We evaluate our one-shot action localization method on the THUMOS14 and ActivityNet datasets, of which we modified the configuration to fit our one-shot problem setup. Xuming He 0001, Fatih Porikli |
CVPR | 3 |
| 2018 | Super-Resolving Very Low-Resolution Face Images With Supplementary AttributesabstractGiven a tiny face image, existing face hallucination methods aim at super-resolving its high-resolution (HR) counterpart by learning a mapping from an exemplar dataset. Since a low-resolution (LR) input patch may correspond to many HR candidate patches, this ambiguity may lead to distorted HR facial details and wrong attributes such as gender reversal. An LR input contains low-frequency facial components of its HR version while its residual face image, defined as the difference between the HR ground-truth and interpolated LR images, contains the missing high-frequency facial details. We demonstrate that supplementing residual images or feature maps with additional facial attribute information can significantly reduce the ambiguity in face super-resolution. To explore this idea, we develop an attribute-embedded upsampling network, which consists of an upsampling network and a discriminative network. The upsampling network is composed of an autoencoder with skip-connections, which incorporates facial attribute vectors into the residual features of LR inputs at the bottleneck of the autoencoder and deconvolutional layers used for upsampling. The discriminative network is designed to examine whether super-resolved faces contain the desired attributes or not and then its loss is used for updating the upsampling network. In this manner, we can super-resolve tiny (16×16 pixels) unaligned face images with a large upscaling factor of 8× while reducing the uncertainty of one-to-many mappings remarkably. By conducting extensive evaluations on a large-scale dataset, we demonstrate that our method achieves superior face hallucination results and outperforms the state-of-the-art. Xin Yu 0002, Basura Fernando, Richard I. Hartley, Fatih Porikli |
CVPR | 4 |
| 2018 | Museum Exhibit Identification Challenge for the Supervised Domain Adaptation and Beyond
Piotr Koniusz, Yusuf Tas, Mehrtash Harandi, Fatih Porikli |
ECCV (16) | 5 |
| 2018 | Face Super-Resolution Guided by Facial Component Heatmaps
Xin Yu 0002, Basura Fernando, Bernard Ghanem, Fatih Porikli, Richard I. Hartley |
ECCV (9) | 4 |
| 2018 | Depth Map Completion by Jointly Exploiting Blurry Color Images and Sparse Depth MapsabstractWe aim at predicting a complete and high-resolution depth map from incomplete, sparse and noisy depth measurements. Existing methods handle this problem either by exploiting various regularizations on the depth maps directly or resorting to learning based methods. When the corresponding color images are available, the correlation between the depth maps and the color images are used to improve the completion performance, assuming the color images are clean and sharp. However, in real world dynamic scenes, color images are often blurry due to the camera motion and the moving objects in the scene. In this paper, we propose to tackle the problem of depth map completion by jointly exploiting the blurry color image sequences and the sparse depth map measurements, and present an energy minimization based formulation to simultaneously complete the depth maps, estimate the scene flow and deblur the color images. Our experimental evaluations on both outdoor and indoor scenarios demonstrate the state-of-the-art performance of our approach. Liyuan Pan, Yuchao Dai, Miaomiao Liu 0001, Fatih Porikli |
WACV | 4 |
| 2018 | Identity-Preserving Face Recovery from PortraitsabstractRecovering the latent photorealistic faces from their artistic portraits aids human perception and facial analysis. However, a recovery process that can preserve identity is challenging because the fine details of real faces can be distorted or lost in stylized images. In this paper, we present a new Identity-preserving Face Recovery from Portraits (IFRP) to recover latent photorealistic faces from unaligned stylized portraits. Our IFRP method consists of two components: Style Removal Network (SRN) and Discriminative Network (DN). The SRN is designed to transfer feature maps of stylized images to the feature maps of the corresponding photorealistic faces. By embedding spatial transformer networks into the SRN, our method can compensate for misalignments of stylized faces automatically and output aligned realistic face images. The role of the DN is to enforce recovered faces to be similar to authentic faces. To ensure the identity preservation, we promote the recovered and ground-truth faces to share similar visual features via a distance measure which compares features of recovered and ground-truth faces extracted from a pre-trained VGG network. We evaluate our method on a large-scale synthesized dataset of real and stylized face pairs and attain state of the art results. In addition, our method can recover photorealistic faces from previously unseen stylized portraits, original paintings and human-drawn sketches. Fatemeh Shiri, Fatih Porikli, Richard I. Hartley, Piotr Koniusz |
WACV | 2 |
| 2018 | Instance-Aware Detailed Action Labeling in VideosabstractWe address the problem of detailed sequence labeling of complex activities in videos, which aims to assign an action label to every frame. Previous work typically focus on predicting action class labels for each frame in a sequence without reasoning action instances. However, such category-level labeling is inefficient in encoding the global constraints at the action instance level and tends to produce inconsistent results. In this work we consider a fusion approach that exploits the synergy between action detection and sequence labeling for complex activities. To this end, we propose an instance-aware sequence labeling method that utilizes the cues from action instance detection. In particular, we design an LSTM-based fusion network that integrates framewise action labeling and action instance prediction to produce a final consistent labeling. To evaluate our method, we create a large-scale RGBD video dataset on gym activities for sequence labeling and action detection called GADD. The experimental results on GADD dataset show that our method outperforms all the state-of-the-art methods consistently in terms of labeling accuracy. Xuming He 0001, Fatih Porikli |
WACV | 3 |
| 2018 | Cascade residuals guided nonlinear dictionary learning
Tong Zhang 0023, Fatih Porikli |
Comput. Vis. Image Underst. | 2 |
| 2018 | Video anomaly detection and localization by local motion based joint video representation and OCELM
Siqi Wang 0001, En Zhu, Jianping Yin, Fatih Porikli |
Neurocomputing | 4 |
| 2018 | Machine learning for big visual analysis
Jun Yu 0002, Xue Mei, Fatih Porikli, Jason J. Corso |
Mach. Vis. Appl. | 3 |
| 2018 | Saliency-Aware Video Object SegmentationabstractVideo saliency, aiming for estimation of a single dominant object in a sequence, offers strong object-level cues for unsupervised video object segmentation. In this paper, we present a geodesic distance based technique that provides reliable and temporally consistent saliency measurement of superpixels as a prior for pixel-wise labeling. Using undirected intra-frame and inter-frame graphs constructed from spatiotemporal edges or appearance and motion, and a skeleton abstraction step to further enhance saliency estimates, our method formulates the pixel-wise segmentation task as an energy minimization problem on a function that consists of unary terms of global foreground and background models, dynamic location models, and pairwise terms of label smoothness potentials. We perform extensive quantitative and qualitative experiments on benchmark datasets. Our method achieves superior performance in comparison to the current state-of-the-art in terms of accuracy and speed. Wenguan Wang, Jianbing Shen, Ruigang Yang, Fatih Porikli |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2018 | Hyperparameter selection of one-class support vector machine by self-adaptive data shifting
Siqi Wang 0001, Qiang Liu 0004, En Zhu, Fatih Porikli, Jianping Yin |
Pattern Recognit. | 4 |
| 2018 | LightenNet: A Convolutional Neural Network for weakly illuminated image enhancement
Chongyi Li, Jichang Guo, Fatih Porikli, Yanwei Pang |
Pattern Recognit. Lett. | 3 |
| 2018 | Distinctive action sketch for human action recognition
Ying Zheng 0009, Hongxun Yao, Xiaoshuai Sun, Sicheng Zhao, Fatih Porikli |
Signal Process. | 5 |
| 2018 | End-to-End Feature Integration for Correlation Filter Tracking With Channel AttentionabstractRecently, the performance advancement of discriminative correlation filter (DCF) based trackers is predominantly driven by the use of deep convolutional features. As convolutional features from multiple layers capture different target information, existing works integrate hierarchical convolutional features to enhance target representation. However, these works separate feature integration from DCF learning and hardly benefit from end-to-end training. In this letter, we incorporates feature integration and DCF learning in a unified convolutional neural network. This network reformulates feature integration as a differential module that concatenates features from the shallow and deep layers. A channel attention mechanism is introduced to adaptively impose channel-wise weight on the integrated features. Experimental results on OTB100 and UAV123 demonstrate that our method achieves significant performance improvement while running in real-time. Dongdong Li 0004, GongJian Wen, Yangliu Kuai, Fatih Porikli |
IEEE Signal Process. Lett. | 4 |
| 2018 | Pushing the Limits of Deep CNNs for Pedestrian DetectionabstractCompared with other applications in computer vision, convolutional neural networks (CNNs) have underperformed on pedestrian detection. A breakthrough was made very recently using sophisticated deep CNN (DCNN) models, with a number of handcrafted features or explicit occlusion handling mechanism. In this paper, we show that by reusing the convolutional feature maps of a DCNN model as image features to train an ensemble of boosted decision models, we are able to achieve the best reported accuracy without using specially designed learning algorithms. We empirically identify and disclose important implementation details. We also show that pixel labeling may be simply combined with a detector to boost the detection performance. By adding complementary handcrafted features such as optical flow, the DCNN-based detector can be further improved. We advance the state-of-the-art results by lowering the log-average miss rate from 11.7% to 8.9% on the Caltech data set and from 11.2% to 8.6% on the Inria data set. We also achieve a comparable result to state-of-the-art approaches on the KITTI data set. Qichang Hu, Peng Wang 0015, Chunhua Shen, Anton van den Hengel, Fatih Porikli |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2018 | Not All Negatives Are Equal: Learning to Track With Multiple Background ClustersabstractConventional tracking-by-detection approaches for visual object tracking often assume that the task at hand is a binary foreground-versus-background classification problem, in which the background is a single, generic, and all-inclusive class. In contrast, here we argue that the background appearance, for the most part, possesses a more complicated structure that would benefit from further partitioning into multiple contextual clusters. Our observation is that, although the background class is contemplated to contain a vast intra-class variation, during the tracking process, only a small portion of this diversity is present at the current frame around the foreground object. This observation motivates us to build multiple fine-grained foreground-versus-contextual-cluster models that provide more discriminative classifications, and consequently more robust and accurate foreground object tracking. For each cluster, we employ a structured output support vector machine (SSVM), and in an online manner, we combine the responses of multiple classifiers. To this end, we apply a top-level SSVM that models the tracked foreground object. We show that our refined modeling of the background is better than naïvely growing the complexity of a single foreground-background classifier, i.e., increasing the number of support vectors that existing approaches rely on, which cause overfitting issues. Our extensive evaluations on large benchmark data sets demonstrate that our tracker consistently outperforms the current state-of-the-art while having comparable computational requirements. Gao Zhu, Fatih Porikli, Hongdong Li |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Guest Editorial Introduction to the Special Issue on Large Scale and Nonlinear Similarity Learning for Intelligent Video AnalysisabstractLearning similarity and distance measures has become increasingly important for the analysis, matching, retrieval, recognition, and categorization of video and multimedia data. With the ubiquitous use of digital imaging devices, mobile terminals and social networks, there are massive volumes of heterogeneous and homogeneous video and multimedia data from multiple sources, views, and domains, e.g., news media websites, microblog, mobile phone, social networking, etc. Similarity and distance-based constraints can also be extended and incorporated to boost classification and relationship learning. Moreover, the spatio-temporal coherence among video data can also be utilized for self-supervised learning of similarity and distance metrics. This trend has brought several challenging issues for developing similarity and metric learning methods for large scale and weakly annotated data, where outliers and incorrectly annotated data are inevitable. Recently, scalability has been investigated to cope with lightweight and large scale metric learning, while nonlinear similarity models have shown their great potentials in learning invariant representation and nonlinear measures of video and multimedia data. Wangmeng Zuo, Liang Lin 0004, Alan L. Yuille, Horst Bischof, Lei Zhang 0006, Fatih Porikli |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2018 | A Unified Approach for Conventional Zero-Shot, Generalized Zero-Shot, and Few-Shot LearningabstractPrevalent techniques in zero-shot learning do not generalize well to other related problem scenarios. Here, we present a unified approach for conventional zero-shot, generalized zero-shot and few-shot learning problems. Our approach is based on a novel Class Adapting Principal Directions (CAPD) concept that allows multiple embeddings of image features into a semantic space. Given an image, our method produces one principal direction for each seen class. Then, it learns how to combine these directions to obtain the principal direction for each unseen class such that the CAPD of the test image is aligned with the semantic embedding of the true class, and opposite to the other classes. This allows efficient and class-adaptive information transfer from seen to unseen classes. In addition, we propose an automatic process for selection of the most useful seen classes for each unseen class to achieve robustness in zero-shot learning. Our method can update the unseen CAPD taking the advantages of few unseen images to work in a few-shot learning scenario. Furthermore, our method can generalize the seen CAPDs by estimating seen-unseen diversity that significantly improves the performance of generalized zero-shot learning. Our extensive evaluations demonstrate that the proposed approach consistently achieves superior performance in zero-shot, generalized zero-shot and few/one-shot learning problems. Shafin Rahman, Salman Khan 0001, Fatih Porikli |
IEEE Trans. Image Process. | 3 |
| 2018 | Imagining the Unimaginable Faces by Deconvolutional NetworksabstractWe tackle the challenge of constructing 64 pixels for each individual pixel of a thumbnail face image. We show that such an aggressive super-resolution objective can be attained by taking advantage of the global context and making the best use of the prior information portrayed by the image class. Our input image is so small (e.g., pixels) that it can be considered as a patch of itself. Thus, conventional patch-matching-based super-resolution solutions are unsuitable. In order to enhance the resolution while enforcing the global context, we incorporate a pixel-wise appearance similarity objective into a deconvolutional neural network, which allows efficient learning of mappings between low-resolution input images and their high-resolution counterparts in the training data set. Furthermore, the deconvolutional network blends the learned high-resolution constituent parts in an authentic manner, where the face structure is naturally imposed and the global context is preserved. To account for the possible artifacts in upsampled feature maps, we employ a sub-network composed of additional convolutional layers. During training, we use roughly aligned images (only eye locations), yet demonstrate that our network has the capacity to super-resolve face images regardless of pose and facial expression variations. This significantly reduces the requirement of precisely face alignments in the data set. Owing to the network topology we apply, our method is robust to translational misalignments. In addition, our method is able to upsample rotational unaligned faces with data augmentation. Our extensive experimental analysis manifests that our method achieves more appealing and superior results than the state of the art. Xin Yu 0002, Fatih Porikli |
IEEE Trans. Image Process. | 2 |
| 2018 | Large-Scale Metric Learning: A Voyage From Shallow to DeepabstractDespite its attractive properties, the performance of the recently introduced Keep It Simple and Straightforward MEtric learning (KISSME) method is greatly dependent on principal component analysis as a preprocessing step. This dependence can lead to difficulties, e.g., when the dimensionality is not meticulously set. To address this issue, we devise a unified formulation for joint dimensionality reduction and metric learning based on the KISSME algorithm. Our joint formulation is expressed as an optimization problem on the Grassmann manifold, and hence enjoys the properties of Riemannian optimization techniques. Following the success of deep learning in recent years, we also devise end-to-end learning of a generic deep network for metric learning using our derivation. Masoud Faraki, Mehrtash Harandi, Fatih Porikli |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | A Comprehensive Look at Coding Techniques on Riemannian ManifoldsabstractCore to many learning pipelines is visual recognition such as image and video classification. In such applications, having a compact yet rich and informative representation plays a pivotal role. An underlying assumption in traditional coding schemes [e.g., sparse coding (SC)] is that the data geometrically comply with the Euclidean space. In other words, the data are presented to the algorithm in vector form and Euclidean axioms are fulfilled. This is of course restrictive in machine learning, computer vision, and signal processing, as shown by a large number of recent studies. This paper takes a further step and provides a comprehensive mathematical framework to perform coding in curved and non-Euclidean spaces, i.e., Riemannian manifolds. To this end, we start by the simplest form of coding, namely, bag of words. Then, inspired by the success of vector of locally aggregated descriptors in addressing computer vision problems, we will introduce its Riemannian extensions. Finally, we study Riemannian form of SC, locality-constrained linear coding, and collaborative coding. Through rigorous tests, we demonstrate the superior performance of our Riemannian coding schemes against the state-of-the-art methods on several visual classification tasks, including head pose classification, video-based face recognition, and dynamic scene recognition. Masoud Faraki, Mehrtash Harandi, Fatih Porikli |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | Robust Object Tracking by Nonlinear LearningabstractWe propose a method that obtains a discriminative visual dictionary and a nonlinear classifier for visual tracking tasks in a sparse coding manner based on the globally linear approximation for a nonlinear learning theory. Traditional discriminative tracking methods based on sparse representation learn a dictionary in an unsupervised way and then train a classifier, which may not generate both descriptive and discriminative models for targets by treating dictionary learning and classifier learning separately. In contrast, the proposed tracking approach can construct a dictionary that fully reflects the intrinsic manifold structure of visual data and introduces more discriminative ability in a unified learning framework. Finally, an iterative optimization approach, which computes the optimal dictionary, the associated sparse coding, and a classifier, is introduced. Experiments on two benchmarks show that our tracker achieves a better performance compared with some popular tracking algorithms. Bo Ma 0001, Hongwei Hu, Jianbing Shen, Ling Shao 0001, Fatih Porikli |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2017 | Face Hallucination with Tiny Unaligned Images by Transformative Discriminative Neural NetworksabstractConventional face hallucination methods rely heavily on accurate alignment of low-resolution (LR) faces before upsampling them. Misalignment often leads to deficient results and unnatural artifacts for large upscaling factors. However, due to the diverse range of poses and different facial expressions, aligning an LR input image, in particular when it is tiny, is severely difficult. To overcome this challenge, here we present an end-to-end transformative discriminative neural network (TDN) devised for super-resolving unaligned and very small face images with an extreme upscaling factor of 8. Our method employs an upsampling network where we embed spatial transformation layers to allow local receptive fields to line-up with similar spatial supports. Furthermore, we incorporate a class-specific loss in our objective through a successive discriminative network to improve the alignment and upsampling performance with semantic information. Extensive experiments on large face datasets show that the proposed method significantly outperforms the state-of-the-art. Xin Yu 0002, Fatih Porikli |
AAAI | 2 |
| 2017 | Combined Internal and External Category-Specific Image Denoising
Saeed Anwar, Cong Phuoc Huynh, Fatih Porikli |
BMVC | 3 |
| 2017 | Depth Estimation and Blur Removal from a Single Out-of-focus Image
Saeed Anwar, Zeeshan Hayder, Fatih Porikli |
BMVC | 3 |
| 2017 | Joint Discriminative Bayesian Dictionary and Classifier LearningabstractWe propose to jointly learn a Discriminative Bayesian dictionary along a linear classifier using coupled Beta-Bernoulli Processes. Our representation model uses separate base measures for the dictionary and the classifier, but associates them to the class-specific training data using the same Bernoulli distributions. The Bernoulli distributions control the frequency with which the factors (e.g. dictionary atoms) are used in data representations, and they are inferred while accounting for the class labels in our approach. To further encourage discrimination in the dictionary, our model uses separate (sets of) Bernoulli distributions to represent data from different classes. Our approach adaptively learns the association between the dictionary atoms and the class labels while tailoring the classifier to this relation with a joint inference over the dictionary and the classifier. Once a test sample is represented over the dictionary, its representation is accurately labelled by the classifier due to the strong coupling between the dictionary and the classifier. We derive the Gibbs Sampling equations for our joint representation model and test our approach for face, object, scene and action recognition to establish its effectiveness. Naveed Akhtar, Ajmal Mian, Fatih Porikli |
CVPR | 3 |
| 2017 | Learning an Invariant Hilbert Space for Domain AdaptationabstractThis paper introduces a learning scheme to construct a Hilbert space (i.e., a vector space along its inner product) to address both unsupervised and semi-supervised domain adaptation problems. This is achieved by learning projections from each domain to a latent space along the Mahalanobis metric of the latent space to simultaneously minimizing a notion of domain variance while maximizing a measure of discriminatory power. In particular, we make use of the Riemannian optimization techniques to match statistical properties (e.g., first and second order statistics) between samples projected into the latent space from different domains. Upon availability of class labels, we further deem samples sharing the same label to form more compact clusters while pulling away samples coming from different classes. We extensively evaluate and contrast our proposal against state-of-the-art methods for the task of visual domain adaptation using both handcrafted and deep-net features. Our experiments show that even with a simple nearest neighbor classifier, the proposed method can outperform several state-of-the-art methods benefitting from more involved classification schemes. Samitha Herath, Mehrtash Harandi, Fatih Porikli |
CVPR | 3 |
| 2017 | Domain Adaptation by Mixture of Alignments of Second-or Higher-Order Scatter TensorsabstractIn this paper, we propose an approach to the domain adaptation, dubbed Second-or Higher-order Transfer of Knowledge (So-HoT), based on the mixture of alignments of second-or higher-order scatter statistics between the source and target domains. The human ability to learn from few labeled samples is a recurring motivation in the literature for domain adaptation. Towards this end, we investigate the supervised target scenario for which few labeled target training samples per category exist. Specifically, we utilize two CNN streams: the source and target networks fused at the classifier level. Features from the fully connected layers fc7 of each network are used to compute second-or even higher-order scatter tensors, one per network stream per class. As the source and target distributions are somewhat different despite being related, we align the scatters of the two network streams of the same class (within-class scatters) to a desired degree with our bespoke loss while maintaining good separation of the between-class scatters. We train the entire network in end-to-end fashion. We provide evaluations on the standard Office benchmark (visual domains) and RGB-D combined with Caltech256 (depth-to-rgb transfer). We attain state-of-the-art results. Piotr Koniusz, Yusuf Tas, Fatih Porikli |
CVPR | 3 |
| 2017 | Simultaneous Stereo Video Deblurring and Scene Flow EstimationabstractVideos for outdoor scene often show unpleasant blur effects due to the large relative motion between the camera and the dynamic objects and large depth variations. Existing works typically focus monocular video deblurring. In this paper, we propose a novel approach to deblurring from stereo videos. In particular, we exploit the piece-wise planar assumption about the scene and leverage the scene flow information to deblur the image. Unlike the existing approach [31] which used a pre-computed scene flow, we propose a single framework to jointly estimate the scene flow and deblur the image, where the motion cues from scene flow estimation and blur information could reinforce each other, and produce superior results than the conventional scene flow estimation or stereo deblurring methods. We evaluate our method extensively on two available datasets and achieve significant improvement in flow estimation and removing the blur effect over the state-of-the-art methods. Liyuan Pan, Yuchao Dai, Miaomiao Liu 0001, Fatih Porikli |
CVPR | 4 |
| 2017 | Hallucinating Very Low-Resolution Unaligned and Noisy Face Images by Transformative Discriminative AutoencodersabstractMost of the conventional face hallucination methods assume the input image is sufficiently large and aligned, and all require the input image to be noise-free. Their performance degrades drastically if the input image is tiny, unaligned, and contaminated by noise. In this paper, we introduce a novel transformative discriminative autoencoder to 8X super-resolve unaligned noisy and tiny (16X16) low-resolution face images. In contrast to encoder-decoder based autoencoders, our method uses decoder-encoder-decoder networks. We first employ a transformative discriminative decoder network to upsample and denoise simultaneously. Then we use a transformative encoder network to project the intermediate HR faces to aligned and noise-free LR faces. Finally, we use the second decoder to generate hallucinated HR images. Our extensive evaluations on a very large face dataset show that our method achieves superior hallucination results and outperforms the state-of-the-art by a large margin of 1.82dB PSNR. Xin Yu 0002, Fatih Porikli |
CVPR | 2 |
| 2017 | Scene Categorization with Spectral FeaturesabstractSpectral signatures of natural scenes were earlier found to be distinctive for different scene types with varying spatial envelope properties such as openness, naturalness, ruggedness, and symmetry. Recently, such handcrafted features have been outclassed by deep learning based representations. This paper proposes a novel spectral description of convolution features, implemented efficiently as a unitary transformation within deep network architectures. To the best of our knowledge, this is the first attempt to use deep learning based spectral features explicitly for image classification task. We show that the spectral transformation decorrelates convolutional activations, which reduces co-adaptation between feature detections, thus acts as an effective regularizer. Our approach achieves significant improvements on three large-scale scene-centric datasets (MIT-67, SUN-397, and Places-205). Furthermore, we evaluated the proposed approach on the attribute detection task where its superior performance manifests its relevance to semantically meaningful characteristics of natural scenes. Salman Khan 0001, Munawar Hayat, Fatih Porikli |
ICCV | 3 |
| 2017 | Super-Trajectory for Video SegmentationabstractWe introduce a novel semi-supervised video segmentation approach based on an efficient video representation, called as “super-trajectory”. Each super-trajectory corresponds to a group of compact trajectories that exhibit consistent motion patterns, similar appearance and close spatiotemporal relationships. We generate trajectories using a probabilistic model, which handles occlusions and drifts in a robust and natural way. To reliably group trajectories, we adopt a modified version of the density peaks based clustering algorithm that allows capturing rich spatiotemporal relations among trajectories in the clustering process. The presented video representation is discriminative enough to accurately propagate the initial annotations in the first frame onto the remaining video frames. Extensive experimental analysis on challenging benchmarks demonstrate our method is capable of distinguishing the target objects from complex backgrounds and even reidentifying them after occlusions. Wenguan Wang, Jianbing Shen, Jianwen Xie, Fatih Porikli |
ICCV | 4 |
| 2017 | Deep tracking with objectnessabstractVisual tracking is a fundamental problem in computer vision. However, due to the (sometimes) ambiguous target information given at the first frame, it has also been criticized as less well-posed compared with other tasks with clearly-defined targets, such as object detection and semantic segmentation. In this paper, we try to evaluate the importance of object category in visual tracking by tracking objects with known object types. The proposed algorithm, termed Deep-Track with Objectness (DTO), naturally combines the state-of-the-art deep-learning-based detectors and trackers, which essentially share a large part of the network. In DTO, a deep tracker, which is scale-fixed and sensitive to small translations tracks the object in a relative short lifespan. A deep detector, which is scale-changeable and robust to pose or illumination changes guides the deep tracker in a longer lifespan. As the deep tracker and detector share the main part of their networks, no much extra computation is imposed while the performance gain is significant. We test the proposed algorithm on two well-accepted benchmarks and on both of them, the proposed method increases the tracking accuracies remarkably compared with state-of-the-art visual trackers. Xinyu Wang 0010, Yi Li 0025, Fatih Porikli, Mingwen Wang 0001 |
ICIP | 4 |
| 2017 | Integrated deep and shallow networks for salient object detectionabstractDeep convolutional neural network (CNN) based salient object detection methods have achieved state-of-the-art performance and outperform those unsupervised methods with a wide margin. In this paper, we propose to integrate deep and unsupervised saliency for salient object detection under a unified framework. Specifically, our method takes results of unsupervised saliency (Robust Background Detection, RBD) and normalized color images as inputs, and directly learns an end-to-end mapping between inputs and the corresponding saliency maps. The color images are fed into a Fully Convolutional Neural Networks (FCNN) adapted from semantic segmentation to exploit high-level semantic cues for salient object detection. Then the results from deep FCNN and RBD are concatenated to feed into a shallow network to map the concatenated feature maps to saliency maps. Finally, to obtain a spatially consistent saliency map with sharp object boundaries, we fuse superpixel level saliency map at multi-scale. Extensive experimental results on 8 benchmark datasets demonstrate that the proposed method outperforms the state-of-the-art approaches with a margin. Jing Zhang 0052, Bo Li 0090, Yuchao Dai, Fatih Porikli, Mingyi He |
ICIP | 4 |
| 2017 | Robust and real-time deep tracking via multi-scale domain adaptationabstractVisual tracking is a fundamental problem in computer vision. Recently, some deep-learning-based tracking algorithms have been achieving record-breaking performances. However, due to the high complexity of deep learning, most deep trackers suffer from low tracking speed, and thus are impractical in many real-world applications. Some new deep trackers with smaller network structure achieve high efficiency while at the cost of significant decrease on precision. In this paper, we propose to transfer the feature for image classification to the visual tracking domain via convolutional channel reductions. The channel reduction could be simply viewed as an additional convolutional layer with the specific task. It not only extracts useful information for object tracking but also significantly increases the tracking speed. To better accommodate the useful feature of the target in different scales, the adaptation filters are designed with different sizes. The yielded visual tracker is real-time and also illustrates the state-of-the-art accuracies in the experiment involving two well-adopted benchmarks with more than 100 test videos. Xinyu Wang 0010, Yi Li 0025, Fumin Shen, Fatih Porikli |
ICME | 5 |
| 2017 | Learning a perspective-embedded deconvolution network for crowd countingabstractWe present a novel deep learning framework for crowd counting by learning a perspective-embedded deconvolution network. Perspective is an inherent property of most surveillance scenes. Unlike the traditional approaches that exploit the perspective as a separate normalization, we propose to fuse the perspective into a deconvolution network, aiming to obtain a robust, accurate and consistent crowd density map. Through layer-wise fusion, we merge perspective maps at different resolutions into the deconvolution network. With the injection of perspective, our network is driven to learn to combine the underlying scene geometric constraints adaptively, thus enabling an accurate interpretation from high-level feature maps to the pixel-wise crowd density map. In addition, our network allows generating density map for arbitrary-sized input in an end-to-end fashion. The proposed method achieves competitive result on the WorldExpo2010 crowd dataset. Muming Zhao, Jian Zhang 0002, Fatih Porikli, Wenjun Zhang 0001 |
ICME | 3 |
| 2017 | Learning deep structured network for weakly supervised change detectionabstractConventional change detection methods require a large number of images to learn background models or depend on tedious pixel-level labeling by humans. In this paper, we present a weakly supervised approach that needs only image-level labels to simultaneously detect and localize changes in a pair of images. To this end, we employ a deep neural network with DAG topology to learn patterns of change from image-level labeled training data. On top of the initial CNN activations, we define a CRF model to incorporate the local differences and context with the dense connections between individual pixels. We apply a constrained mean-field algorithm to estimate the pixel-level labels, and use the estimated labels to update the parameters of the CNN in an iterative EM framework. This enables imposing global constraints on the observed foreground probability mass function. Our evaluations on four benchmark datasets demonstrate superior detection and localization performance. Salman Khan 0001, Xuming He 0001, Fatih Porikli, Mohammed Bennamoun, Ferdous Sohel, Roberto Togneri |
IJCAI | 3 |
| 2017 | Ordered Pooling of Optical Flow Sequences for Action RecognitionabstractTraining of Convolutional Neural Networks (CNNs) on long video sequences is computationally expensive due to the substantial memory requirements and the massive number of parameters that deep architectures demand. Early fusion of video frames is thus a standard technique, in which several consecutive frames are first agglomerated into a compact representation, and then fed into the CNN as an input sample. For this purpose, a summarization approach that represents a set of consecutive RGB frames by a single dynamic image to capture pixel dynamics is proposed recently. In this paper, we introduce a novel ordered representation of consecutive optical flow frames as an alternative and argue that this representation captures the action dynamics more efficiently than RGB frames. We provide intuitions on why such a representation is better for action recognition. We validate our claims on standard benchmark datasets and demonstrate that using summaries of flow images lead to significant improvements over RGB frames while achieving accuracy comparable to the stateof-the-art on UCF101 and HMDB datasets. Jue Wang 0010, Anoop Cherian, Fatih Porikli |
WACV | 3 |
| 2017 | Deep Salient Object Detection by Integrating Multi-level CuesabstractA key problem in salient object detection is how to effectively exploit the multi-level saliency cues in a unified and data-driven manner. In this paper, building upon the recent success of deep neural networks, we propose a fully convolutional neural network based approach empowered with multi-level fusion to salient object detection. By integrating saliency cues at different levels through fully convolutional neural networks and multi-level fusion, our approach could effectively exploit both learned semantic cues and higher-order region statistics for edge-accurate salient object detection. First, we fine-tune a fully convolutional neural network for semantic segmentation to adapt it to salient object detection to learn a suitable yet coarse perpixel saliency prediction map. This map is often smeared across salient object boundaries since the local receptive fields in the convolutional network apply naturally on both sides of such boundaries. Second, to enhance the resolution of the learned saliency prediction and to incorporate higher-order cues that are omitted by the neural network, we propose a multi-level fusion approach where super-pixel level coherency in saliency is exploited. Our extensive experimental results on various benchmark datasets demonstrate that the proposed method outperforms the state-of the-art approaches. Jing Zhang 0052, Yuchao Dai, Fatih Porikli |
WACV | 3 |
| 2017 | Learning Spatial Transforms for Refining Object Segment ProposalsabstractWe address the problem of object segment proposal generation, which is a critical step in many instance-level semantic segmentation and scene understanding pipelines. In contrast to prior works that predict binary segment masks from images, we take an alternative refinement approach to improve the quality of a given segment candidate pool. In particular, we propose an efficient deep network that learns 2D spatial transforms to warp an initial object mask towards nearby object region. We formulate this segment refinement task as a regression problem and design a novel feature pooling strategy in our deep network to predict an affine transformation for each object mask. We evaluate our method extensively on two challenging public benchmarks and apply our refinement network to three different initial segment proposal settings. Our results show sizable improvements in average recall across all the settings, achieving the state-of-the-art performances. Xuming He 0001, Fatih Porikli |
WACV | 3 |
| 2017 | Going deeper into action recognition: A survey
Samitha Herath, Mehrtash Harandi, Fatih Porikli |
Image Vis. Comput. | 3 |
| 2017 | Breaking video into pieces for action recognition
Ying Zheng 0009, Hongxun Yao, Xiaoshuai Sun, Xuesong Jiang, Fatih Porikli |
Multim. Tools Appl. | 5 |
| 2017 | No fuss metric learning, a Hilbert space scenario
Masoud Faraki, Mehrtash Harandi, Fatih Porikli |
Pattern Recognit. Lett. | 3 |
| 2017 | Forest Change Detection in Incomplete Satellite Images With Deep Neural NetworksabstractLand cover change monitoring is an important task from the perspective of regional resource monitoring, disaster management, land development, and environmental planning. In this paper, we analyze imagery data from remote sensing satellites to detect forest cover changes over a period of 29 years (1987-2015). Since the original data are severely incomplete and contaminated with artifacts, we first devise a spatiotemporal inpainting mechanism to recover the missing surface reflectance information. The spatial filling process makes use of the available data of the nearby temporal instances followed by a sparse encoding-based reconstruction. We formulate the change detection task as a region classification problem. We build a multiresolution profile (MRP) of the target area and generate a candidate set of bounding-box proposals that enclose potential change regions. In contrast to existing methods that use handcrafted features, we automatically learn region representations using a deep neural network in a data-driven fashion. Based on these highly discriminative representations, we determine forest changes and predict their onset and offset timings by labeling the candidate set of proposals. Our approach achieves the state-of-the-art average patch classification rate of 91.6% (an improvement of ~16%) and the mean onset/offset prediction error of 4.9 months (an error reduction of five months) compared with a strong baseline. We also qualitatively analyze the detected changes in the unlabeled image regions, which demonstrate that the proposed forest change detection approach is scalable to new regions. Salman Khan 0001, Xuming He 0001, Fatih Porikli, Mohammed Bennamoun |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2017 | Category-Specific Object Image DenoisingabstractWe present a novel image denoising algorithm that uses external, category specific image database. In contrast to existing noisy image restoration algorithms that search patches either from a generic database or noisy image itself, our method first selects clean images similar to the noisy image from a database that consists of images of the same class. Then, within the spatial locality of each noisy patch, it assembles a set of "support patches" from the selected images. These noisy-free support samples resemble the noisy patch and correspond principally to the identical part of the depicted object. In addition, we employ a content adaptive distribution model for each patch, where we derive the parameters of the distribution from the support patches. We formulate noise removal task as an optimization problem in the transform domain. Our objective function composed of a Gaussian fidelity term that imposes category specific information, and a low-rank term that encourages the similarity between the noisy and the support patches in a robust manner. The denoising process is driven by an iterative selection of support patches and optimization of the objective function. Our extensive experiments on five different object categories confirm the benefit of incorporating category-specific information to noise removal and demonstrate the superior performance of our method over the state-of-the-art alternatives. Saeed Anwar, Fatih Porikli, Cong Phuoc Huynh |
IEEE Trans. Image Process. | 2 |
| 2017 | Visual Tracking by Sampling in Part SpaceabstractIn this paper, we present a novel part-based visual tracking method from the perspective of probability sampling. Specifically, we represent the target by a part space with two online learned probabilities to capture the structure of the target. The proposal distribution memorizes the historical performance of different parts, and it is used for the first round of part selection. The acceptance probability validates the specific tracking stability of each part in a frame, and it determines whether to accept its vote or to reject it. By doing this, we transform the complex online part selection problem into a probability learning one, which is easier to tackle. The observation model of each part is constructed by an improved supervised descent method and is learned in an incremental manner. Experimental results on two benchmarks demonstrate the competitive performance of our tracker against state-of-the-art methods. Lianghua Huang, Bo Ma 0001, Jianbing Shen, Ling Shao 0001, Fatih Porikli |
IEEE Trans. Image Process. | 6 |
| 2017 | Higher Order Energies for Image SegmentationabstractA novel energy minimization method for general higher order binary energy functions is proposed in this paper. We first relax a discrete higher order function to a continuous one, and use the Taylor expansion to obtain an approximate lower order function, which is optimized by the quadratic pseudo-Boolean optimization or other discrete optimizers. The minimum solution of this lower order function is then used as a new local point, where we expand the original higher order energy function again. Our algorithm does not restrict to any specific form of the higher order binary function or bring in extra auxiliary variables. For concreteness, we show an application of segmentation with the appearance entropy, which is efficiently solved by our method. Experimental results demonstrate that our method outperforms the state-of-the-art methods. Jianbing Shen, Jianteng Peng, Xingping Dong, Ling Shao 0001, Fatih Porikli |
IEEE Trans. Image Process. | 5 |
| 2017 | Selective Video Object CutoutabstractConventional video segmentation approaches rely heavily on appearance models. Such methods often use appearance descriptors that have limited discriminative power under complex scenarios. To improve the segmentation performance, this paper presents a pyramid histogram-based confidence map that incorporates structure information into appearance statistics. It also combines geodesic distance-based dynamic models. Then, it employs an efficient measure of uncertainty propagation using local classifiers to determine the image regions, where the object labels might be ambiguous. The final foreground cutout is obtained by refining on the uncertain regions. Additionally, to reduce manual labeling, our method determines the frames to be labeled by the human operator in a principled manner, which further boosts the segmentation performance and minimizes the labeling effort. Our extensive experimental analyses on two big benchmarks demonstrate that our solution achieves superior performance, favorable computational efficiency, and reduced manual labeling in comparison to the state of the art. Wenguan Wang, Jianbing Shen, Fatih Porikli |
IEEE Trans. Image Process. | 3 |
| 2017 | Optimal Couple Projections for Domain Adaptive Sparse Representation-Based ClassificationabstractIn recent years, sparse representation-based classification (SRC) is one of the most successful methods and has been shown impressive performance in various classification tasks. However, when the training data have a different distribution than the testing data, the learned sparse representation may not be optimal, and the performance of SRC will be degraded significantly. To address this problem, in this paper, we propose an optimal couple projections for domain-adaptive SRC (OCPD-SRC) method, in which the discriminative features of data in the two domains are simultaneously learned with the dictionary that can succinctly represent the training and testing data in the projected space. OCPD-SRC is designed based on the decision rule of SRC, with the objective to learn coupled projection matrices and a common discriminative dictionary such that the between-class sparse reconstruction residuals of data from both domains are maximized, and the within-class sparse reconstruction residuals of data are minimized in the projected low-dimensional space. Thus, the resulting representations can well fit SRC and simultaneously have a better discriminant ability. In addition, our method can be easily extended to multiple domains and can be kernelized to deal with the nonlinear structure of data. The optimal solution for the proposed method can be efficiently obtained following the alternative optimization method. Extensive experimental results on a series of benchmark databases show that our method is better or comparable to many state-of-the-art methods. Guoqing Zhang 0002, Huaijiang Sun, Fatih Porikli, Yazhou Liu, Quan-Sen Sun |
IEEE Trans. Image Process. | 3 |
| 2017 | Automatic Refinement Strategies for Manual Initialization of Object TrackersabstractTracking objects across multiple frames is a well-investigated problem in computer vision. The majority of the existing algorithms that assume an accurate initialization is readily available. However, in many real-life settings, in particular for applications where the video is streaming in real time, the initialization has to be provided by a human operator. This limitation raises an inevitable uncertainty issue. Here, we first collect a large and new data set of inputs that consists of more than 20 K human initialization clicks, by several subjects under three practical user interface scenarios for the popular TB50 tracking benchmark. We analyze the factors and mechanisms of human input, derive statistical models, and show that human input always contains deviations, which exacerbate further when the relative object-camera motion becomes large. We also design and evaluate alternative refinement schemes, and propose a strategy that refits an object window on the most probable target region after a single click. To compensate for the human initialization errors, our method generates window proposals using objectness cues extracted from color and motion attributes, accumulates them into a likelihood map that is weighted by the initial click position and visual saliency scores, and assigns the final window by the maximum likelihood estimate. Our experiments demonstrate that the presented refinement strategy effectively reduces human input errors. Hao Zhu 0002, Fatih Porikli |
IEEE Trans. Image Process. | 2 |
| 2016 | Object-Aware Dictionary Learning with Deep Features
Yurui Xie, Fatih Porikli, Xuming He 0001 |
ACCV (2) | 2 |
| 2016 | Learning to Generate Object Segment Proposals with Multi-modal Cues
Xuming He 0001, Fatih Porikli |
ACCV (1) | 3 |
| 2016 | Sparse Coding on Cascaded Residuals
Tong Zhang 0023, Fatih Porikli |
ACCV (4) | 2 |
| 2016 | Model-Free Multiple Object Tracking with Shared Proposals
Gao Zhu, Fatih Porikli, Hongdong Li |
ACCV (2) | 2 |
| 2016 | When VLAD Met HilbertabstractIn many challenging visual recognition tasks where training data is limited, Vectors of Locally Aggregated Descriptors (VLAD) have emerged as powerful image/video representations that compete with or outperform state-of the-art approaches. In this paper, we address two fundamental limitations of VLAD: its requirement for the local descriptors to have vector form and its restriction to linear classifiers due to its high-dimensionality. To this end, we introduce a kernelized version of VLAD. This not only lets us inherently exploit more sophisticated classification schemes, but also enables us to efficiently aggregate nonvector descriptors (e.g., manifold-valued data) in the VLAD framework. Furthermore, we propose an approximate formulation that allows us to accelerate the coding process while still benefiting from the properties of kernel VLAD. Our experiments demonstrate the effectiveness of our approach at handling manifold-valued data, such as covariance descriptors, on several classification tasks. Our results also evidence the benefits of our nonlinear VLAD descriptors against the linear ones in Euclidean space using several standard benchmark datasets. Mehrtash Harandi, Mathieu Salzmann, Fatih Porikli |
CVPR | 3 |
| 2016 | Beyond Local Search: Tracking Objects Everywhere with Instance-Specific ProposalsabstractMost tracking-by-detection methods employ a local search window around the predicted object location in the current frame assuming the previous location is accurate, the trajectory is smooth, and the computational capacity permits a search radius that can accommodate the maximum speed yet small enough to reduce mismatches. These, however, may not be valid always, in particular for fast and irregularly moving objects. Here, we present an object tracker that is not limited to a local search window and has ability to probe efficiently the entire frame. Our method generates a small number of "high-quality" proposals by a novel instance-specific objectness measure and evaluates them against the object model that can be adopted from an existing tracking-by-detection approach as a core tracker. During the tracking process, we update the object model concentrating on hard false-positives supplied by the proposals, which help suppressing distractors caused by difficult background clutters, and learn how to re-rank proposals according to the object model. Since we reduce significantly the number of hypotheses the core tracker evaluates, we can use richer object descriptors and stronger detector. Our method outperforms most recent state-of-the-art trackers on popular tracking benchmarks, and provides improved robustness for fast moving objects as well as for ultra lowframerate videos. Gao Zhu, Fatih Porikli, Hongdong Li |
CVPR | 2 |
| 2016 | Tensor Representations via Kernel Linearization for Action Recognition from 3D Skeletons
Piotr Koniusz, Anoop Cherian, Fatih Porikli |
ECCV (4) | 3 |
| 2016 | Ultra-Resolving Face Images by Discriminative Generative Networks
Xin Yu 0002, Fatih Porikli |
ECCV (5) | 2 |
| 2016 | Less Is More: Towards Compact CNNs
Fatih Porikli |
ECCV (4) | 3 |
| 2016 | Semantic context and depth-aware object proposal generationabstractThis paper presents a context-aware object proposal generation method for stereo images. Unlike existing methods which mostly rely on image-based or depth features to generate object candidates, we propose to incorporate additional geometric and high-level semantic context information into the proposal generation. Our method starts from an initial object proposal set, and encode objectness for each proposal using three types of features , including a CNN feature, a geometric feature computed from dense depth map, and a semantic context feature from pixel-wise scene labeling. We then train an efficient random forest classifier to re-rank the initial proposals and a set of linear regressors to fine-tune the location of each proposal. Experiments on the KITTI dataset show our approach significantly improves the quality of the initial proposals and achieves the state-of-the-art performance using only a fraction of original object candidates. Xuming He 0001, Fatih Porikli, Laurent Kneip |
ICIP | 3 |
| 2016 | Detection and characterization of Intrinsic symmetry of 3D shapesabstractA comprehensive framework for detection and characterization of partial intrinsic symmetry over 3D shapes is proposed. To identify prominent symmetric regions which overlap in space and vary in form, the proposed framework is decoupled into a Correspondence Space Voting (CSV) procedure followed by a Transformation Space Mapping (TSM) procedure. In the CSV procedure, significant symmetries are first detected by identifying surface point pairs on the input shape that exhibit local similarity in terms of their intrinsic geometry while simultaneously maintaining an intrinsic distance structure at a global level. To allow detection of potentially overlapping symmetric shape regions, a global intrinsic distance-based voting scheme is employed to ensure the inclusion of only those point pairs that exhibit significant intrinsic symmetry. In the TSM procedure, the Functional Map framework is employed to generate the final map of symmetries between point pairs. The TSM procedure ensures the retrieval of the underlying dense correspondence map throughout the 3D shape that follows a particular symmetry. The TSM procedure is also shown to result in the formulation of a metric symmetry space where each point in the space represents a specific symmetry transformation and the distance between points represents the complexity between the corresponding transformations. Experimental results show that the proposed framework can successfully analyze complex 3D shapes that possess rich symmetries. Anirban Mukhopadhyay 0003, Suchendra M. Bhandarkar, Fatih Porikli |
ICPR | 3 |
| 2016 | Finetuning Convolutional Neural Networks for visual aestheticsabstractInferring the aesthetic quality of images is a challenging computer vision task due to its subjective and conceptual nature. Most image aesthetics evaluation approaches focused on designing handcrafted features, and only a few adopted learning of relevant and imperative characteristics in a data-driven manner. In this paper, we propose to attune Convolutional Neural Networks (CNNs) for image aesthetics. Unlike previous deep learning based techniques, we employ pretrained models, namely AlexNet [12] and the 16-layer VGGNet [20], and calibrate them to estimate visual aesthetic quality. This enables exploiting automatically the inherent information from much larger scale and more diversified image datasets. We tested our methods on AVA and CUHKPQ image aesthetics datasets on two different training-testing partitions, and compared the performance using both local and contextual information. Experimental results suggest that our strategy is robust, effective and superior to the state-of-the-art approaches. Yeqing Wang, Yi Li 0025, Fatih Porikli |
ICPR | 3 |
| 2016 | Anomaly detection in crowded scenes by SL-HOF descriptor and foreground classificationabstractWith the widespread use of surveillance cameras, massive video data analysis has become an extremely labor-intensive work. In this paper, we propose an efficient approach to detect video anomaly in crowded scenes based on Spatially Localized Histogram of Optical Flow (SL-HOF) descriptor and foreground classification. For motion description, the new SL-HOF descriptor can not only preserve classic HOF descriptor's favorable capability of characterizing the motion velocity and direction of foreground in crowded scene, but also depicts the spatial distribution of optical flow, which implicitly encodes the structure and local motion information of foreground objects in videos. SL-HOF is shown to significantly outperform other classic video descriptors. To further boost the performance of anomaly localization, we then introduce Robust PCA based foreground classification to discriminate anomalous foreground texture. Instead of computationally expensive approaches like l1-norm Sparse Coding, we adopt classic one-class SVM (OCSVM) to model normal video events and detect outliers (anomaly). Our experiments on the challenging UCSD datasets show our approach can achieve state-of-the-art results when compared to existing video anomaly detection methods. Siqi Wang 0001, En Zhu, Jianping Yin, Fatih Porikli |
ICPR | 4 |
| 2016 | Image set classification by symmetric positive semi-definite matricesabstractRepresenting images and videos by covariance descriptors and leveraging the inherent manifold structure of Symmetric Positive Definite (SPD) matrices leads to enhanced performances in various visual recognition tasks. However, when covariance descriptors are used to represent image sets, the result is often rank-deficient. Thus, most existing approaches adhere to blind perturbation with predefined regularizers just to be able to employ inference tools. To overcome this problem, we introduce novel similarity measures specifically designed for rank-deficient covariance descriptors, i.e., symmetric positive semi-definite matrices. In particular, we derive positive definite kernels that can be decomposed into the kernels on the cone of SPD matrices and kernels on the Grassmann manifolds. Our experiments evidence that, our method achieves superior results for image set classification on various recognition tasks including hand gesture classification, face recognition from video sequences, and dynamic scene categorization. Masoud Faraki, Mehrtash Harandi, Fatih Porikli |
WACV | 3 |
| 2016 | Convolutional neural net bagging for online visual tracking
Yi Li 0025, Fatih Porikli |
Comput. Vis. Image Underst. | 3 |
| 2016 | A Novel Performance Evaluation Methodology for Single-Target TrackersabstractThis paper addresses the problem of single-target tracker performance evaluation. We consider the performance measures, the dataset and the evaluation system to be the most important components of tracker evaluation and propose requirements for each of them. The requirements are the basis of a new evaluation methodology that aims at a simple and easily interpretable tracker comparison. The ranking-based methodology addresses tracker equivalence in terms of statistical significance and practical differences. A fully-annotated dataset with per-frame annotations with several visual attributes is introduced. The diversity of its visual properties is maximized in a novel way by clustering a large number of videos according to their visual attributes. This makes it the most sophistically constructed and annotated dataset to date. A multi-platform evaluation system allowing easy integration of third-party trackers is presented as well. The proposed evaluation methodology was tested on the VOT2014 challenge on the new dataset and 38 trackers, making it the largest benchmark to date. Most of the tested trackers are indeed state-of-the-art since they outperform the standard baselines, resulting in a highly-challenging benchmark. An exhaustive analysis of the dataset from the perspective of tracking difficulty is carried out. To facilitate tracker comparison a new performance visualization technique is proposed. Matej Kristan, Jiri Matas, Ales Leonardis, Tomás Vojír, Roman P. Pflugfelder, Gustavo Fernández, Georg Nebehay, Fatih Porikli, Luka Cehovin |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2016 | DeepTrack: Learning Discriminative Feature Representations Online for Robust Visual TrackingabstractDeep neural networks, albeit their great success on feature learning in various computer vision tasks, are usually considered as impractical for online visual tracking, because they require very long training time and a large number of training samples. In this paper, we present an efficient and very robust tracking algorithm using a single convolutional neural network (CNN) for learning effective feature representations of the target object in a purely online manner. Our contributions are multifold. First, we introduce a novel truncated structural loss function that maintains as many training samples as possible and reduces the risk of tracking error accumulation. Second, we enhance the ordinary stochastic gradient descent approach in CNN training with a robust sample selection mechanism. The sampling mechanism randomly generates positive and negative samples from different temporal distributions, which are generated by taking the temporal relations and label noise into account. Finally, a lazy yet effective updating scheme is designed for CNN training. Equipped with this novel updating algorithm, the CNN model is robust to some long-existing difficulties in visual tracking, such as occlusion or incorrect detections, without loss of the effective adaption for significant appearance changes. In the experiment, our CNN tracker outperforms all compared state-of-the-art methods on two recently proposed benchmarks, which in total involve over 60 video sequences. The remarkable performance improvement over the existing trackers illustrates the superiority of the feature representations, which are learned purely online via the proposed deep learning framework. Yi Li 0025, Fatih Porikli |
IEEE Trans. Image Process. | 3 |
| 2016 | Visual Tracking Under Motion BlurabstractMost existing tracking algorithms do not explicitly consider the motion blur contained in video sequences, which degrades their performance in real-world applications where motion blur often occurs. In this paper, we propose to solve the motion blur problem in visual tracking in a unified framework. Specifically, a joint blur state estimation and multi-task reverse sparse learning framework are presented, where the closed-form solution of blur kernel and sparse code matrix is obtained simultaneously. The reverse process considers the blurry candidates as dictionary elements, and sparsely represents blurred templates with the candidates. By utilizing the information contained in the sparse code matrix, an efficient likelihood model is further developed, which quickly excludes irrelevant candidates and narrows the particle scale down. Experimental results on the challenging benchmarks show that our method performs well against the state-of-the-art trackers. Bo Ma 0001, Lianghua Huang, Jianbing Shen, Ling Shao 0001, Ming-Hsuan Yang 0001, Fatih Porikli |
IEEE Trans. Image Process. | 6 |
| 2016 | Correspondence Driven Saliency TransferabstractIn this paper, we show that large annotated data sets have great potential to provide strong priors for saliency estimation rather than merely serving for benchmark evaluations. To this end, we present a novel image saliency detection method called saliency transfer. Given an input image, we first retrieve a support set of best matches from the large database of saliency annotated images. Then, we assign the transitional saliency scores by warping the support set annotations onto the input image according to computed dense correspondences. To incorporate context, we employ two complementary correspondence strategies: a global matching scheme based on scene-level analysis and a local matching scheme based on patch-level inference. We then introduce two refinement measures to further refine the saliency maps and apply the random-walk-with-restart by exploring the global saliency structure to estimate the affinity between foreground and background assignments. Extensive experimental results on four publicly available benchmark data sets demonstrate that the proposed saliency algorithm consistently outperforms the current state-of-the-art methods. Wenguan Wang, Jianbing Shen, Ling Shao 0001, Fatih Porikli |
IEEE Trans. Image Process. | 4 |
| 2016 | Fast Detection of Multiple Objects in Traffic Scenes With a Common Detection FrameworkabstractTraffic scene perception (TSP) aims to extract accurate real-time on-road environment information, which involves three phases: detection of objects of interest, recognition of detected objects, and tracking of objects in motion. Since recognition and tracking often rely on the results from detection, the ability to detect objects of interest effectively plays a crucial role in TSP. In this paper, we focus on three important classes of objects: traffic signs, cars, and cyclists. We propose to detect all the three important objects in a single learning-based detection framework. The proposed framework consists of a dense feature extractor and detectors of three important classes. Once the dense features have been extracted, these features are shared with all detectors. The advantage of using one common framework is that the detection speed is much faster, since all dense features need only to be evaluated once in the testing phase. In contrast, most previous works have designed specific detectors using different features for each of these three classes. To enhance the feature robustness to noises and image deformations, we introduce spatially pooled features as a part of aggregated channel features. In order to further improve the generalization performance, we propose an object subcategorization method as a means of capturing the intraclass variation of objects. We experimentally demonstrate the effectiveness and efficiency of the proposed framework in three detection applications: traffic sign detection, car detection, and cyclist detection. The proposed framework achieves the competitive performance with state-of-the-art approaches on several benchmark data sets. Qichang Hu, Sakrapee Paisitkriangkrai, Chunhua Shen, Anton van den Hengel, Fatih Porikli |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2015 | Enforcing Point-wise Priors on Binary Segmentation
Feng Li 0005, Fatih Porikli |
BMVC | 2 |
| 2015 | More about VLAD: A leap from Euclidean to Riemannian manifoldsabstractThis paper takes a step forward in image and video coding by extending the well-known Vector of Locally Aggregated Descriptors (VLAD) onto an extensive space of curved Riemannian manifolds. We provide a comprehensive mathematical framework that formulates the aggregation problem of such manifold data into an elegant solution. In particular, we consider structured descriptors from visual data, namely Region Covariance Descriptors and linear subspaces that reside on the manifold of Symmetric Positive Definite matrices and the Grassmannian manifolds, respectively. Through rigorous experimental validation, we demonstrate the superior performance of this novel Riemannian VLAD descriptor on several visual classification tasks including video-based face recognition, dynamic scene recognition, and head pose classification. Masoud Faraki, Mehrtash Harandi, Fatih Porikli |
CVPR | 3 |
| 2015 | Saliency-aware geodesic video object segmentationabstractWe introduce an unsupervised, geodesic distance based, salient video object segmentation method. Unlike traditional methods, our method incorporates saliency as prior for object via the computation of robust geodesic measurement. We consider two discriminative visual features: spatial edges and temporal motion boundaries as indicators of foreground object locations. We first generate framewise spatiotemporal saliency maps using geodesic distance from these indicators. Building on the observation that foreground areas are surrounded by the regions with high spatiotemporal edge values, geodesic distance provides an initial estimation for foreground and background. Then, high-quality saliency results are produced via the geodesic distances to background regions in the subsequent frames. Through the resulting saliency maps, we build global appearance models for foreground and background. By imposing motion continuity, we establish a dynamic location model for each frame. Finally, the spatiotemporal saliency maps, appearance models and dynamic location models are combined into an energy minimization framework to attain both spatially and temporally coherent object segmentation. Extensive quantitative and qualitative experiments on benchmark video dataset demonstrate the superiority of the proposed method over the state-of-the-art algorithms. Wenguan Wang, Jianbing Shen, Fatih Porikli |
CVPR | 3 |
| 2015 | Approximate infinite-dimensional Region Covariance Descriptors for image classificationabstractWe introduce methods to estimate infinite-dimensional Region Covariance Descriptors (RCovDs) by exploiting two feature mappings, namely random Fourier features and the Nyström method. In general, infinite-dimensional RCovDs offer better discriminatory power over their low-dimensional counterparts. However, the underlying Riemannian structure, i.e., the manifold of Symmetric Positive Definite (SPD) matrices, is out of reach to great extent for infinite-dimensional RCovDs. To overcome this difficulty, we propose to approximate the infinite-dimensional RCovDs by making use of the aforementioned explicit mappings. We will empirically show that the proposed finite-dimensional approximations of infinite-dimensional RCovDs consistently outperform the low-dimensional RCovDs for image classification task, while enjoying the Riemannian structure of the SPD manifolds. Moreover, our methods achieve the state-of-the-art performance on three different image classification tasks. Masoud Faraki, Mehrtash Harandi, Fatih Porikli |
ICASSP | 3 |
| 2015 | Class-Specific Image DeblurringabstractIn image deblurring, a fundamental problem is that the blur kernel suppresses a number of spatial frequencies that are difficult to recover reliably. In this paper, we explore the potential of a class-specific image prior for recovering spatial frequencies attenuated by the blurring process. Specifically, we devise a prior based on the class-specific subspace of image intensity responses to band-pass filters. We learn that the aggregation of these subspaces across all frequency bands serves as a good class-specific prior for the restoration of frequencies that cannot be recovered with generic image priors. In an extensive validation, our method, equipped with the above prior, yields greater image quality than many state-of-the-art methods by up to 5 dB in terms of image PSNR, across various image categories including portraits, cars, cats, pedestrians and household objects. Saeed Anwar, Cong Phuoc Huynh, Fatih Porikli |
ICCV | 3 |
| 2015 | Linearization to Nonlinear Learning for Visual TrackingabstractDue to unavoidable appearance variations caused by occlusion, deformation, and other factors, classifiers for visual tracking are nonlinear as a necessity. Building on the theory of globally linear approximations to nonlinear functions, we introduce an elegant method that jointly learns a nonlinear classifier and a visual dictionary for tracking objects in a semi-supervised sparse coding fashion. This establishes an obvious distinction from conventional sparse coding based discriminative tracking algorithms that usually maintain two-stage learning strategies, i.e., learning a dictionary in an unsupervised way then followed by training a classifier. However, the treating dictionary learning and classifier training as separate stages may not produce both descriptive and discriminative models for objects. By contrast, our method is capable of constructing a dictionary that not only fully reflects the intrinsic manifold structure of the data, but also possesses discriminative power. This paper presents an optimization method to obtain such an optimal dictionary, associated sparse coding, and a classifier in an iterative process. Our experiments on a benchmark show our tracker attains outstanding performance compared with the state-of-the-art algorithms. Bo Ma 0001, Hongwei Hu, Jianbing Shen, Fatih Porikli |
ICCV | 5 |
| 2015 | Material Classification on Symmetric Positive Definite ManifoldsabstractThis paper tackles the problem of categorizing materials and textures by exploiting the second order statistics. To this end, we introduce the Extrinsic Vector of Locally Aggregated Descriptors (E-VLAD), a method to combine local and structured descriptors into a unified vector representation where each local descriptor is a Covariance Descriptor (CovD). In doing so, we make use of an accelerated method of obtaining a visual codebook where each atom is itself a CovD. We will then introduce an efficient way of aggregating local CovDs into a vector representation. Our method could be understood as an extrinsic extension of the highly acclaimed method of Vector of Locally Aggregated Descriptors [17] (or VLAD) to CovDs. We will show that the proposed method is extremely powerful in classifying materials/ textures and can outperform complex machineries even with simple classifiers. Masoud Faraki, Mehrtash Harandi, Fatih Porikli |
WACV | 3 |
| 2015 | Lie-Struck: Affine Tracking on Lie Groups Using Structured SVMabstractThis paper presents a novel and reliable tracking-by detection method for image regions that undergo affine transformations such as translation, rotation, scale, dilatation and shear deformations, which span the six degrees of freedom of motion. Our method takes advantage of the intrinsic Lie group structure of the 2D affine motion matrices and imposes this motion structure on a kernelized structured output SVM classifier that provides an appearance based prediction function to directly estimate the object transformation between frames using geodesic distances on manifolds unlike the existing methods proceeding by linearizing the motion. We demonstrate that these combined motion and appearance model structures greatly improve the tracking performance while an incorporated particle filter on the motion hypothesis space keeps the computational load feasible. Experimentally, we show that our algorithm is able to outperform state-of-the-art affine trackers in various scenarios. Gao Zhu, Fatih Porikli, Yansheng Ming, Hongdong Li |
WACV | 2 |
| 2015 | Discriminative feature learning from big data for visual recognition
Zhuolin Jiang, Zhe Lin 0001, Haibin Ling, Fatih Porikli, Ling Shao 0001, Pavan Turaga |
Pattern Recognit. | 4 |
| 2015 | Robust Video Object CosegmentationabstractWith ever-increasing volumes of video data, automatic extraction of salient object regions became even more significant for visual analytic solutions. This surge has also opened up opportunities for taking advantage of collective cues encapsulated in multiple videos in a cooperative manner. However, it also brings up major challenges, such as handling of drastic appearance, motion pattern, and pose variations, of foreground objects as well as indiscriminate backgrounds. Here, we present a cosegmentation framework to discover and segment out common object regions across multiple frames and multiple videos in a joint fashion. We incorporate three types of cues, i.e., intraframe saliency, interframe consistency, and across-video similarity into an energy optimization framework that does not make restrictive assumptions on foreground appearance and motion model, and does not require objects to be visible in all frames. We also introduce a spatio-temporal scale-invariant feature transform (SIFT) flow descriptor to integrate across-video correspondence from the conventional SIFT-flow into interframe motion flow from optical flow. This novel spatio-temporal SIFT flow generates reliable estimations of common foregrounds over the entire video data set. Experimental results show that our method outperforms the state-of-the-art on a new extensive data set (ViCoSeg). Wenguan Wang, Jianbing Shen, Xuelong Li 0001, Fatih Porikli |
IEEE Trans. Image Process. | 4 |
| 2014 | Robust Online Visual Tracking with a Single Convolutional Neural Network
Yi Li 0025, Fatih Porikli |
ACCV (5) | 3 |
| 2014 | Counting people by clustering person detector outputsabstractWe present a people counting system that estimates the number of people in a scene by employing a clustering scheme based on Dirichlet Process Mixture Models (DP-MMs) which takes outputs of a person detector system as input. For each frame, we run a person detector on the frame, take its output as a set of detection areas and define a set of features based on spatial, color and temporal information for each detection. Then using these features, we cluster the detections using DPMMs and Gibbs sampling while having no restriction on the number of clusters, thus can estimate an arbitrary number of people or groups of people. We finally define a measure to calculate the actual number of people within each cluster to infer the final estimation of the number of people in the scene. Ibrahim Saygin Topkaya, Hakan Erdogan, Fatih Porikli |
AVSS | 3 |
| 2014 | DeepTrack: Learning Discriminative Feature Representations by Convolutional Neural Networks for Visual Tracking
Yi Li 0025, Fatih Porikli |
BMVC | 3 |
| 2014 | Bregman Divergences for Infinite Dimensional Covariance MatricesabstractWe introduce an approach to computing and comparing Covariance Descriptors (CovDs) in infinite-dimensional spaces. CovDs have become increasingly popular to address classification problems in computer vision. While CovDs offer some robustness to measurement variations, they also throw away part of the information contained in the original data by only retaining the second-order statistics over the measurements. Here, we propose to overcome this limitation by first mapping the original data to a high-dimensional Hilbert space, and only then compute the CovDs. We show that several Bregman divergences can be computed between the resulting CovDs in Hilbert space via the use of kernels. We then exploit these divergences for classification purpose. Our experiments demonstrate the benefits of our approach on several tasks, such as material and texture recognition, person re-identification, and action recognition from motion capture data. Mehrtash Harandi, Mathieu Salzmann, Fatih Porikli |
CVPR | 3 |
| 2014 | Robust Orthonormal Subspace Learning: Efficient Recovery of Corrupted Low-Rank MatricesabstractLow-rank matrix recovery from a corrupted observation has many applications in computer vision. Conventional methods address this problem by iterating between nuclear norm minimization and sparsity minimization. However, iterative nuclear norm minimization is computationally prohibitive for large-scale data (e.g., video) analysis. In this paper, we propose a Robust Orthogonal Subspace Learning (ROSL) method to achieve efficient low-rank recovery. Our intuition is a novel rank measure on the low-rank matrix that imposes the group sparsity of its coefficients under orthonormal subspace. We present an efficient sparse coding algorithm to minimize this rank measure and recover the low-rank matrix at quadratic complexity of the matrix size. We give theoretical proof to validate that this rank measure is lower bounded by nuclear norm and it has the same global minimum as the latter. To further accelerate ROSL to linear complexity, we also describe a faster version (ROSL+) empowered by random sampling. Our extensive experiments demonstrate that both ROSL and ROSL+ provide superior efficiency against the state-of-the-art methods at the same level of recovery accuracy. Xianbiao Shu, Fatih Porikli, Narendra Ahuja |
CVPR | 2 |
| 2014 | Super-resolving Noisy ImagesabstractOur goal is to obtain a noise-free, high resolution (HR) image, from an observed, noisy, low resolution (LR) image. The conventional approach of preprocessing the image with a denoising algorithm, followed by applying a super-resolution (SR) algorithm, has an important limitation: Along with noise, some high frequency content of the image (particularly textural detail) is invariably lost during the denoising step. This 'denoising loss' restricts the performance of the subsequent SR step, wherein the challenge is to synthesize such textural details. In this paper, we show that high frequency content in the noisy image (which is ordinarily removed by denoising algorithms) can be effectively used to obtain the missing textural details in the HR domain. To do so, we first obtain HR versions of both the noisy and the denoised images, using a patch-similarity based SR algorithm. We then show that by taking a convex combination of orientation and frequency selective bands of the noisy and the denoised HR images, we can obtain a desired HR image where (i) some of the textural signal lost in the denoising step is effectively recovered in the HR domain, and (ii) additional textures can be easily synthesized by appropriately constraining the parameters of the convex combination. We show that this part-recovery and part-synthesis of textures through our algorithm yields HR images that are visually more pleasing than those obtained using the conventional processing pipeline. Furthermore, our results show a consistent improvement in numerical metrics, further corroborating the ability of our algorithm to recover lost signal. Abhishek Singh 0002, Fatih Porikli, Narendra Ahuja |
CVPR | 2 |
| 2014 | Sparse Dictionaries for Semantic Segmentation
Lingling Tao, Fatih Porikli, René Vidal |
ECCV (5) | 2 |
| 2014 | Recycled linear classifiers for multiclass classificationabstractMany machine learning applications employ a multiclass classification stage that uses multiple binary linear classifiers as building blocks. Among these, commonly used strategies such as one-vs-one classification can require learning a large number of hyperplanes, even when the number of classes to be discriminated among is modest. Further, when the data being classified is inherently high-dimensional, the storage and computational complexity associated with the application of multiple linear classifiers can ignite critical resource management issues. This work describes a novel multiclass classification method based on efficient use of a single “recycled” linear classifier (or ReLiC), which addresses these storage and implementation complexity issues. The proposed approach amounts to constraining the entire collection of hyperplanes to be circularly-shifted versions of each other, enabling classification procedures that may be implemented with efficient operations, such as circular convolution (which can be efficiently computed using transform domain techniques), and simple sampling/thresholding operations. We show that the optimization task associated with our proposed approach can be formulated as a quadratic program, and we introduce an efficient distributed procedure for its solution based on an alternating direction method of multipliers. Simulation results demonstrate that the performance of the proposed approach is comparable with the more complex, traditional multiclass linear classification strategies, suggesting the proposed approach is a viable alternative in large-scale data classification tasks. Akshay Soni, Jarvis D. Haupt, Fatih Porikli |
ICASSP | 3 |
| 2014 | Detecting 3D geometric boundaries of indoor scenes under varying lightingabstractThe goal of this research is to identify 3D geometric boundaries in a set of 2D photographs of a static indoor scene under unknown, changing lighting conditions. A 3D geometric boundary is a contour located at a 3D depth discontinuity or a discontinuity in the surface normal. These boundaries can be used effectively for reasoning about the 3D layout of a scene. To distinguish 3D geometric boundaries from 2D texture edges, we analyze the illumination subspace of local appearance at each image location. In indoor time-lapse photography and surveillance video, we frequently see images that are lit by unknown combinations of uncalibrated light sources. We introduce an algorithm for semi-binary nonnegative matrix factorization (SBNMF) to decompose such images into a set of lighting basis images, each of which shows the scene lit by a single light source. These basis images provide a natural, succinct representation of the scene, enabling tasks such as scene editing (e.g., relighting) and shadow edge identification. Jie Ni, Tim K. Marks, Oncel Tuzel, Fatih Porikli |
WACV | 4 |
| 2014 | Special issue on car navigation and vehicle systems
Fatih Porikli, Luc Van Gool |
Mach. Vis. Appl. | 1 |
| 2014 | Classification and Boosting with Multiple Collaborative RepresentationsabstractRecent advances have shown a great potential to explore collaborative representations of test samples in a dictionary composed of training samples from all classes in multi-class recognition including sparse representations. In this paper, we present two multi-class classification algorithms that make use of multiple collaborative representations in their formulations, and demonstrate performance gain of exploring this extra degree of freedom. We first present the Collaborative Representation Optimized Classifier (CROC), which strikes a balance between the nearest-subspace classifier, which assigns a test sample to the class that minimizes the distance between the sample and its principal projection in the selected class, and a Collaborative Representation based Classifier (CRC), which assigns a test sample to the class that minimizes the distance between the sample and its collaborative components. Several well-known classifiers become special cases of CROC under different regularization parameters. We show classification performance can be improved by optimally tuning the regularization parameter through cross validation. We then propose the Collaborative Representation based Boosting (CRBoosting) algorithm, which generalizes the CROC to incorporate multiple collaborative representations. Extensive numerical examples are provided with performance comparisons of different choices of collaborative representations, in particular when the test sample is available via compressive measurements. Yuejie Chi, Fatih Porikli |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | A Novel Video Dataset for Change Detection BenchmarkingabstractChange detection is one of the most commonly encountered low-level tasks in computer vision and video processing. A plethora of algorithms have been developed to date, yet no widely accepted, realistic, large-scale video data set exists for benchmarking different methods. Presented here is a unique change detection video data set consisting of nearly 90 000 frames in 31 video sequences representing six categories selected to cover a wide range of challenges in two modalities (color and thermal infrared). A distinguishing characteristic of this benchmark video data set is that each frame is meticulously annotated by hand for ground-truth foreground, background, and shadow area boundaries-an effort that goes much beyond a simple binary label denoting the presence of change. This enables objective and precise quantitative comparison and ranking of video-based change detection algorithms. This paper discusses various aspects of the new data set, quantitative performance metrics used, and comparative results for over two dozen change detection algorithms. It draws important conclusions on solved and remaining issues in change detection, and describes future challenges for the scientific community. The data set, evaluation tools, and algorithm rankings are available to the public on a website and will be updated with feedback from academia and industry in the future. Nil Goyette, Pierre-Marc Jodoin, Fatih Porikli, Janusz Konrad, Prakash Ishwar |
IEEE Trans. Image Process. | 3 |
| 2013 | Harmonic Variance: A Novel Measure for In-focus SegmentationabstractWe introduce an efficient measure for estimating the degree of in-focus within an image region. This measure, harmonic mean of variances, is computed from the statistical properties of the image in its bandpass filtered versions. We incorporate the harmonic variance into a graph Laplacian spectrum based segmentation framework in order to accurately align the in-focus measure responses on the underlying image structure. Our results demonstrate the effectiveness of this novel measure for determining and segmenting in-focus regions in low depth-of-field images. Feng Li 0005, Fatih Porikli |
BMVC | 2 |
| 2013 | Unconstrained 1D range and 2D image based human detectionabstractAn accurate and computationally very fast multimodal human detector is presented. This 1D+2D detector fuses 1D range scan and 2D image information via an effective geometric descriptor and a silhouette based visual representation within a radial basis function kernel support vector machine learning framework. Unlike the existing approaches, the proposed 1D+2D detector does not make any restrictive assumptions on the range scan positions, thus it is applicable to a wide range of real-life detection tasks. To analyze the discriminative power of the geometric descriptor, a range scan only version, 1D+, is also evaluated. Extensive experiments demonstrate that the 1D+2D detector works robustly under challenging imaging conditions and achieves several orders of magnitude performance improvement while reducing the computational load drastically. In addition, a new multi-modal (LIDAR, depth image, optical image) dataset, DontHitMe, is introduced. This dataset contains 40,000 registered frames and 3,600 manually annotated human objects. It depicts challenging illumination conditions in indoors and outdoors environments and is publicly available to our community. Mehmet Kemal Kocamaz, Fatih Porikli |
IROS | 2 |
| 2013 | Support Vector Shape: A Classifier-Based Shape RepresentationabstractWe introduce a novel implicit representation for 2D and 3D shapes based on Support Vector Machine (SVM) theory. Each shape is represented by an analytic decision function obtained by training SVM, with a Radial Basis Function (RBF) kernel so that the interior shape points are given higher values. This empowers support vector shape (SVS) with multifold advantages. First, the representation uses a sparse subset of feature points determined by the support vectors, which significantly improves the discriminative power against noise, fragmentation, and other artifacts that often come with the data. Second, the use of the RBF kernel provides scale, rotation, and translation invariant features, and allows any shape to be represented accurately regardless of its complexity. Finally, the decision function can be used to select reliable feature points. These features are described using gradients computed from highly consistent decision functions instead from conventional edges. Our experiments demonstrate promising results. Hien Van Nguyen, Fatih Porikli |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | A comparison of methods for non-rigid 3D shape retrieval
Zhouhui Lian, Afzal Godil, Benjamin Bustos, Mohamed Daoudi, Jeroen Hermans, Shun Kawamura, Yukinori Kurita, Guillaume Lavoué, Hien Van Nguyen, Ryutarou Ohbuchi, Yuki Ohkita, Yuya Ohishi, Fatih Porikli, Martin Reuter 0001, Ivan Sipiran, Dirk Smeets, Paul Suetens, Hedi Tabia, Dirk Vandermeulen |
Pattern Recognit. | 13 |
| 2012 | Connecting the dots in multi-class classification: From nearest subspace to collaborative representationabstractWe present a novel multi-class classifier that strikes a balance between the nearest-subspace classifier, which assigns a test sample to the class that minimizes the distance between the test sample and its principal projection in the selected class, and a collaborative representation based classifier, which classifies a sample to the class that minimizes the distance between the collaborative components of the test sample by using all training samples from all classes as the dictionary and its projection in the selected class. In our formulation, the sparse representation based classifier [1] and nearest subspace classifier become special cases under different regularization parameters. We show that the classification performance can be improved by optimally tuning the regularization parameter, which can be done at almost no extra computational cost. We give extensive numerical examples for digit identification and face recognition with performance comparisons of different choices of collaborative representations, in particular when only a partial observation of the test sample is available via compressive sensing measurements. Yuejie Chi, Fatih Porikli |
CVPR | 2 |
| 2012 | A clustering approach to optimize online dictionary learningabstractDictionary learning has emerged as a powerful tool for low level image processing tasks such as denoising and inpainting, as well as sparse coding and representation of images. While there has been extensive work on the development of online and offline dictionary learning algorithms to perform the aforementioned tasks, the problem of choosing an appropriate dictionary size is not as widely addressed. In this paper, we introduce a new scheme to reduce and optimize dictionary size in an online setting by synthesizing new atoms from multiple previous ones. We show that this method performs as well as existing offline and online dictionary learning algorithms in terms of representation accuracy while achieving significant speedup in dictionary reconstruction and image encoding times. Our method not only helps in choosing smaller and more representative dictionaries, but also enables learning of more incoherent ones. Nikhil Rao 0001, Fatih Porikli |
ICASSP | 2 |
| 2012 | Additive noise removal by sparse reconstruction on image affinity netsabstractThis paper presents a new image denoising method based on sparse reconstruction by dictionary learning and collaborative filtering. First, we form an affinity net, in which a node represents an image patch, for the given image by clustering similar patches. For each cluster, we learn an undercomplete dictionary and represent clusters nodes by imposing sparsity inducing norm as a combination of few atoms. Depending on its affinity to other nodes, a single node could be present in multiple clusters making the clusters overlapping. This enables a single global estimation for each filtered pixel to be obtained by collaboratively aggregating its reconstructed patches in the corresponding clusters. Extensive experimental results demonstrate superior performance for additive noise removal without requiring the correct noise variance. Rajagopalan Sundaresan, Fatih Porikli |
ICASSP | 2 |
| 2012 | Multiple dictionary learning for blocking artifacts reductionabstractWe present a structured dictionary learning method to remove blocking artifacts without blurring edges or making any assumption over image gradients. Instead of a single overcomplete dictionary, we build multiple subspaces and impose sparsity on nonzero reconstruction coefficients when we project a given texture sample on each subspace separately. In case the texture matches to the dataset with which the subspace is trained, the corresponding response will be stronger and that subspace will be chosen to represent the texture. In this manner we compute the representations of all patches in the image and aggregate these to obtain the final image. Since the block artifacts are small in magnitude in comparison to actual image edges, aggregation efficiently removes the artifacts but keep the image gradients. We discuss the choices of subspace parameterizations and adaptation to given data. Our results on a large dataset of benchmark images demonstrate that the presented method provides superior results in terms of pixel-wise (PSNR) and perceptual (SSIM) measures. Yi Wang 0009, Fatih Porikli |
ICASSP | 2 |
| 2012 | Compressive Clustering of High-Dimensional DataabstractIn this paper we focus on realistic clustering problems where the input data is high-dimensional and the clusters have complex, multimodal distribution. In this challenging setting the conventional methods, such as k-centers family, hierarchical clustering or those based on model fitting, are inefficient and typically converge far from the globally optimal solution. As an alternative, we propose a novel unsupervised learning approach which is based on the compressive sensing paradigm. The key idea underlying our algorithm is to monitor the distance between the test sample and its principal projection in each cluster, and continue re-assigning it to the cluster yielding the smallest residual. As a result, we obtain an iterative procedure which, under the compressive assumptions, minimizes the total reconstruction error of all samples from their nearest clusters. To evaluate the proposed approach, we have conducted a series of experiments involving various image collections where the task was to automatically group similar objects. Comparison of the obtained results with those yielded by the state-of-the-art clustering methods provides evidence for high discriminative power of our algorithm. Andrzej Ruta, Fatih Porikli |
ICMLA (1) | 2 |
| 2012 | Coverage optimized active learning for k - NN classifiersabstractFast image recognition and classification is extremely important in various robotics applications such as exploration, rescue, localization, etc. k-nearest neighbor (kNN) classifiers are popular tools used in classification since they involve no explicit training phase, and are simple to implement. However, they often require large amounts of training data to work well in practice. In this paper, we propose a batch-mode active learning algorithm for efficient training of kNN classifiers, that substantially reduces the amount of training required. As opposed to much previous work on iterative single-sample active selection, the proposed system selects samples in batches. We propose a coverage formulation that enforces selected samples to be distributed such that all data points have labeled samples at a bounded maximum distance, given the training budget, so that there are labeled neighbors in a small neighborhood of each point. Using submodular function optimization, the proposed algorithm presents a near-optimal selection strategy for an otherwise intractable problem. Further we employ uncertainty sampling along with coverage to incorporate model information and improve classification. Finally, we use locality sensitive hashing for fast retrieval of nearest neighbors during active selection as well as classification, which provides 1-2 orders of magnitude speedups thus allowing real-time classification with large datasets. Ajay J. Joshi, Fatih Porikli, Nikolaos Papanikolopoulos |
ICRA | 2 |
| 2012 | Scalable Active Learning for Multiclass Image ClassificationabstractMachine learning techniques for computer vision applications like object recognition, scene classification, etc., require a large number of training samples for satisfactory performance. Especially when classification is to be performed over many categories, providing enough training samples for each category is infeasible. This paper describes new ideas in multiclass active learning to deal with the training bottleneck, making it easier to train large multiclass image classification systems. First, we propose a new interaction modality for training which requires only yes-no type binary feedback instead of a precise category label. The modality is especially powerful in the presence of hundreds of categories. For the proposed modality, we develop a Value-of-Information (VOI) algorithm that chooses informative queries while also considering user annotation cost. Second, we propose an active selection measure that works with many categories and is extremely fast to compute. This measure is employed to perform a fast seed search before computing VOI, resulting in an algorithm that scales linearly with dataset size. Third, we use locality sensitive hashing to provide a very fast approximation to active learning, which gives sublinear time scaling, allowing application to very large datasets. The approximation provides up to two orders of magnitude speedups with little loss in accuracy. Thorough empirical evaluation of classification accuracy, noise sensitivity, imbalanced data, and computational performance on a diverse set of image datasets demonstrates the strengths of the proposed algorithms. Ajay J. Joshi, Fatih Porikli, Nikolaos Papanikolopoulos |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Data driven frequency mapping for computationally scalable object detectionabstractNonlinear kernel Support Vector Machines achieve better generalizations, yet their training and evaluation speeds are prohibitively slow for real-time object detection tasks where the number of data points in training and the number of hypotheses to be tested in evaluation are in the order of millions. To accelerate the training and particularly testing of such nonlinear kernel machines, we map the input data onto a low-dimensional spectral (Fourier) feature space using a cosine transform, design a kernel that approximates the classification objective in a supervised setting, and apply a fast linear classifier instead of the conventional radial basis functions. We present a data driven hypotheses generation technique and a LogistBoost feature selection. Our experimental results demonstrate the computational improvements 20~100× while maintaining a high classification accuracy in comparison to SVM linear and radial kernel basis function classifiers. Fatih Porikli, Huseyin Ozkan |
AVSS | 1 |
| 2011 | CrossTrack: Robust 3D tracking from two cross-sectional viewsabstractOne of the challenges in radiotherapy of moving tumors is to determine the location of the tumor accurately. Existing solutions to the problem are either invasive or inaccurate. We introduce a non-invasive solution to the problem by tracking the tumor in 3D using bi-plane ultrasound image sequences. We present CrossTrack, a novel tracking algorithm in this framework. We pose the problem as recursive inference of 3D location and tumor boundary segmentation in the two ultrasound views using the tumor 3D model as a prior. For the segmentation task, a robust graph-based approach is deployed as follows: First, robust segmentation priors are obtained through the tumor 3D model. Second, a unified graph combining information across time and multiple views is constructed with a robust weighting function. For the tracking task, an effective mechanism for recovery from respiration-induced occlusion is introduced. Our experiments show the robustness of CrossTrack in handling challenging tumor shapes and disappearance scenarios, with sub-voxel accuracy, and almost 100% precision and recall, significantly outperforming baseline solutions. Mohamed E. Hussein 0001, Fatih Porikli, Rui Li 0053, Suayb S. Arslan |
CVPR | 2 |
| 2011 | Concentric ring signature descriptor for 3D objectsabstractWe present a 3D feature descriptor that represents local topologies within a set of folded concentric rings by distances from local points to a projection plane. This feature, called as Concentric Ring Signature (CORS), possesses similar computational advantages to point signatures yet provides more accurate matches. It produces more compact and discriminative descriptors than shape context. It robust to noise and occlusions. As opposed to spin images, CORS does not require the point normal estimations, therefore it is directly applicable to sparse point clouds where the point densities are insufficiently low. Under the same settings, we demonstrate that the discriminative power of CORS is superior to conventional approaches producing twice as good estimates with the percentage of correct match scores improving from 39% to 88%. Hien Van Nguyen, Fatih Porikli |
ICIP | 2 |
| 2011 | Special issue on dynamic textures in video
A. Enis Çetin, Fatih Porikli |
Mach. Vis. Appl. | 2 |
| 2011 | In-vehicle camera traffic sign detection and recognition
Andrzej Ruta, Fatih Porikli, Shintaro Watanabe, Yongmin Li 0001 |
Mach. Vis. Appl. | 2 |
| 2010 | CycleStack: Inferring periodic behavior via temporal sequence visualization in ultrasound videoabstractA range of well-known treatment methods for destroying tumor and similar harmful growth in human body utilizes the coherence between the inherently periodic movement of the affected body part and periodic respiratory signal of the patient, with the objective of minimizing damage to surrounding normal tissues. Such methods require constant monitoring by an operator who observes the 3D body motion via its 2D projection onto an ultrasound imaging plane and studies the synchronism of this motion with the respiratory signal. Keeping an attentive eye on the respiratory signal as well as the ultrasound video for the entire treatment period is often inconvenient and burdensome. In this paper, we propose a video visualization technique called CycleStack Plot which reduces this cognitive overhead by blending the video and the signal together in a stack-like layout. This visualization reveals the inherent synchronism between the target's movement and the respiratory signal, visually highlights significant phase shifts of either of the two cyclic phenomena, with the hope of arresting the operator's attention. Our proposed visualization also provides a visual overview for the post-treatment analysis which enables educated users to quickly and effectively skim through the excessively long process. This paper demonstrates the utility of CycleStack Plot with a case study using real ultrasound videos. In addition, a user study has been performed to evaluate the merits and limitations of the proposed method with respect to the conventional way of watching a video and a signal side-by-side. Even though the motivation of the proposed visualization is improvement of medical applications that use ultrasound, the core techniques discussed here have potential to be extended to other application domains requiring analysis of cyclic patterns from videos. Teng-Yok Lee, Abon Chaudhuri, Fatih Porikli, Han-Wei Shen |
PacificVis | 3 |
| 2010 | Breaking the interactive bottleneck in multi-class classification with active selection and binary feedbackabstractMulti-class classification schemes typically require human input in the form of precise category names or numbers for each example to be annotated - providing this can be impractical for the user when a large (and possibly unknown) number of categories are present. In this paper, we propose a multi-class active learning model that requires only binary (yes/no type) feedback from the user. For instance, given two images the user only has to say whether they belong to the same class or not. We first show the interactive benefits of such a scheme with user experiments. We then propose a Value of Information (VOI)-based active selection algorithm in the binary feedback model. The algorithm iteratively selects image pairs for annotation so as to maximize accuracy, while also minimizing user annotation effort. To our knowledge, this is the first multi-class active learning approach that requires only yes/no inputs. Experiments show that the proposed method can substantially minimize user supervision compared to the traditional training model, on problems with as many as 100 classes. We also demonstrate that the system is robust to real-world issues such as class population imbalance and labeling noise. Ajay J. Joshi, Fatih Porikli, Nikolaos Papanikolopoulos |
CVPR | 2 |
| 2010 | Scene-Adaptive Human Detection with Incremental Active LearningabstractIn many computer vision tasks, scene changes hinder the generalization ability of trained classifiers. For instance, a human detector trained with one set of images is unlikely to perform well in different scene conditions. In this paper, we propose an incremental learning method for human detection that can take generic training data and build a new classifier adapted to the new deployment scene. Two operation modes are proposed: i) a completely autonomous mode wherein first few empty frames of video are used for adaptation, and ii) an active learning approach with user in the loop, for more challenging scenarios including situations where empty initialization frames may not exist. Results show the strength of the proposed methods for quick adaptation. Ajay J. Joshi, Fatih Porikli |
ICPR | 2 |
| 2010 | Human State Classification and Predication for Critical Care Monitoring by Real-Time Bio-signal AnalysisabstractTo address the challenges in critical care monitoring, we present a multi-modality bio-signal modeling and analysis modeling framework for real-time human state classification and predication. The novel bioinformatic framework is developed to solve the human state classification and predication issues from two aspects: a) achieve 1:1 mapping between the bio-signal and the human state via discriminant feature analysis and selection by using probabilistic principle component analysis (PPCA); b) avoid time-consuming data analysis and extensive integration resources by using Dynamic Bayesian Network (DBN). In addition, intelligent and automatic selection of the most suitable sensors from the bio-sensor array is also integrated in the proposed DBN. Xiaokun Li, Fatih Porikli |
ICPR | 2 |
| 2010 | Multi-class batch-mode active learning for image classificationabstractAccurate image classification is crucial in many robotics and surveillance applications - for example, a vision system on a robot needs to accurately recognize the objects seen by its camera. Object recognition systems typically need a large amount of training data for satisfactory performance. The problem is particularly acute when many object categories are present. In this paper we present a batch-mode active learning framework for multi-class image classification systems. In active learning, images are to be chosen for interactive labeling, instead of passively accepting training data. Our framework addresses two important issues: i) it handles redundancy between different images which is crucial when batch-mode selection is performed; and ii) we pose batch-selection as a submodular function optimization problem that makes an inherently intractable problem efficient to solve, while having approximation guarantees. We show results on image classification data in which our approach substantially reduces the amount of training required over the baseline. Ajay J. Joshi, Fatih Porikli, Nikolaos Papanikolopoulos |
ICRA | 2 |
| 2010 | Compressed Domain Video Object SegmentationabstractWe present a compressed domain video object segmentation method for the MPEG encoded video sequences. For a fraction of the raw domain analysis, compressed domain segmentation provides the essentiala prioriinformation to many vision tasks from surveillance to transcoding that require fast processing of large volumes of data where pixel-resolution boundary extraction is not required. Our method generates accurate segmentation maps in block resolution at hierarchically varying object levels, which empowers application to determine the most pertinent partition of images. It exploits the block structure of the compressed video to minimize the amount of data to be processed. All the available motion flow within a group of pictures is projected onto a single layer, which also consists of the frequency decomposition of color pattern. Then, by starting from the blocks where the spatial energy is small, it expands homogeneous regions while automatically adapting local similarity criteria. We also formulate an alternative solution that applies a kernel-based clustering where separate spatial, transform, and motion kernels are used to establish the affinity. We show that both region expansion and mean shift produce similar results as the computationally expensive raw domain segmentation. Finally, a binary clustering iteratively merges the most similar regions to generate a hierarchical partition tree. Fatih Porikli, Faisal I. Bashir, Huifang Sun |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2009 | Regressed Importance Sampling on Manifolds for Efficient Object TrackingabstractIn this paper, a new integrated particle filter is proposed for video object tracking. After particles are generated by importance sampling, each particle is regressed on the transformation space where the mapping function is learned offline by regression on pose manifold using Lie algebra, leading to a more effective allocation of particles. Experimental results on synthetic and real sequences clearly demonstrate the improved pose (affine) tracking performance of the proposed method compared with the original regression tracker and particle filters. Fatih Porikli |
AVSS | 1 |
| 2009 | Multi-class active learning for image classificationabstractOne of the principal bottlenecks in applying learning techniques to classification problems is the large amount of labeled training data required. Especially for images and video, providing training data is very expensive in terms of human time and effort. In this paper we propose an active learning approach to tackle the problem. Instead of passively accepting random training examples, the active learning algorithm iteratively selects unlabeled examples for the user to label, so that human effort is focused on labeling the most “useful” examples. Our method relies on the idea of uncertainty sampling, in which the algorithm selects unlabeled examples that it finds hardest to classify. Specifically, we propose an uncertainty measure that generalizes margin-based uncertainty to the multi-class case and is easy to compute, so that active learning can handle a large number of classes and large data sizes efficiently. We demonstrate results for letter and digit recognition on datasets from the UCI repository, object recognition results on the Caltech-101 dataset, and scene categorization results on a dataset of 13 natural scene categories. The proposed method gives large reductions in the number of training examples required over random selection to achieve similar classification accuracy, with little computational overhead. Ajay J. Joshi, Fatih Porikli, Nikolaos Papanikolopoulos |
CVPR | 2 |
| 2009 | A new method for tracking performance evaluation based on a reflective model and perturbation analysisabstractIn this paper, we present a novel methodology for tracking performance evaluation. Considering the continuity of the image sequences in a video, we define a new measurement called tracking difficulty which incorporates the local sequence information among a small image sequence centered at each frame. We subsequently use a reflective model to formulate tracking difficulty. Tracking difficulty curves can not only illustrate at which parts of the video one tracking algorithm performs well or poor, but also provide a way to compare the performance of different tracking algorithms. We further add perturbation analysis to the reflective model to examine how sensitive the tracking algorithm is to noise. Results on data sets are presented to show the effectiveness of our evaluation method. Fatih Porikli, Dan Schonfeld |
ICASSP | 2 |
| 2009 | Kernel methods for weakly supervised mean shift clusteringabstractMean shift clustering is a powerful unsupervised data analysis technique which does not require prior knowledge of the number of clusters, and does not constrain the shape of the clusters. The data association criteria is based on the underlying probability distribution of the data points which is defined in advance via the employed distance metric. In many problem domains, the initially designed distance metric fails to resolve the ambiguities in the clustering process. We present a novel semi-supervised kernel mean shift algorithm where the inherent structure of the data points is learned with a few user supplied constraints in addition to the original metric. The constraints we consider are the pairs of points that should be clustered together. The data points are implicitly mapped to a higher dimensional space induced by the kernel function where the constraints can be effectively enforced. The mode seeking is then performed on the embedded space and the approach preserves all the advantages of the original mean shift algorithm. Experiments on challenging synthetic and real data clearly demonstrate that significant improvements in clustering accuracy can be achieved by employing only a few constraints. Oncel Tuzel, Fatih Porikli, Peter Meer |
ICCV | 2 |
| 2009 | Object detection via boosted deformable featuresabstractIt is a common practice to model an object for detection tasks as a boosted ensemble of many models built on features of the object. In this context, features are defined as subregions with fixed relative locations and extents with respect to the object's image window. We introduce using deformable features with boosted ensembles. A deformable features adapts its location depending on the visual evidence in order to match the corresponding physical feature. Therefore, deformable features can better handle deformable objects. We empirically show that boosted ensembles of deformable features perform significantly better than boosted ensembles of fixed features for human detection. Mohamed A. Hussein 0001, Fatih Porikli, Larry Davis 0001 |
ICIP | 2 |
| 2009 | A Comprehensive Evaluation Framework and a Comparative Study for Human DetectorsabstractWe introduce a framework for evaluating human detectors that considers the practical application of a detector on a full image using multisize sliding-window scanning. We produce detection error tradeoff (DET) curves relating the miss detection rate and the false-alarm rate computed by deploying the detector on cropped windows and whole images, using, in the latter, either image resize or feature resize. Plots for cascade classifiers are generated based on confidence scores instead of on variation of the number of layers. To assess a method's overall performance on a given test, we use the average log miss rate (ALMR) as an aggregate performance score. To analyze the significance of the obtained results, we conduct 10-fold cross-validation experiments. We applied our evaluation framework to two state-of-the-art cascade-based detectors on the standard INRIA person dataset and a local dataset of near-infrared images. We used our evaluation framework to study the differences between the two detectors on the two datasets with different evaluation methods. Our results show the utility of our framework. They also suggest that the descriptors used to represent features and the training window size are more important in predicting the detection performance than the nature of the imaging process, and that the choice between resizing images or features can have serious consequences. Mohamed E. Hussein 0001, Fatih Porikli, Larry Davis 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2008 | Kernel integral images: A framework for fast non-uniform filteringabstractIntegral images are commonly used in computer vision and computer graphics applications. Evaluation of box filters via integral images can be performed in constant time, regardless of the filter size. Although Heckbert (1986) extended the integral image approach for more complex filters, its usage has been very limited, in practice. In this paper, we present an extension to integral images that allows for application of a wide class of non-uniform filters. Our approach is superior to Heckbertpsilas in terms of precision requirements and suitability for parallelization. We explain the theoretical basis of the approach and instantiate two concrete examples: filtering with bilinear interpolation, and filtering with approximated Gaussian weighting. Our experiments show the significant speedups we achieve, and the higher accuracy of our approach compared to Heckbertpsilas. Mohamed E. Hussein 0001, Fatih Porikli, Larry Davis 0001 |
CVPR | 2 |
| 2008 | Boosting adaptive linear weak classifiers for online learning and trackingabstractOnline boosting methods have recently been used successfully for tracking, background subtraction etc. Conventional online boosting algorithms emphasize on interchanging new weak classifiers/features to adapt with the change over time. We are proposing a new online boosting algorithm where the form of the weak classifiers themselves are modified to cope with scene changes. Instead of replacement, the parameters of the weak classifiers are altered in accordance with the new data subset presented to the online boosting process at each time step. Thus we may avoid altogether the issue of how many weak classifiers to be replaced to capture the change in the data or which efficient search algorithm to use for a fast retrieval of weak classifiers. A computationally efficient method has been used in this paper for the adaptation of linear weak classifiers. The proposed algorithm has been implemented to be used both as an online learning and a tracking method. We show quantitative and qualitative results on both UCI datasets and several video sequences to demonstrate improved performance of our algorithm. Toufiq Parag, Fatih Porikli, Ahmed M. Elgammal |
CVPR | 2 |
| 2008 | Constant time O(1) bilateral filteringabstractThis paper presents three novel methods that enable bilateral filtering in constant time O(1) without sampling. Constant time means that the computation time of the filtering remains same even if the filter size becomes very large. Our first method takes advantage of the integral histograms to avoid the redundant operations for bilateral filters with box spatial and arbitrary range kernels. For bilateral filters constructed by polynomial range and arbitrary spatial filters, our second method provides a direct formulation by using linear filters of image powers without any approximation. Lastly, we show that Gaussian range and arbitrary spatial bilateral filters can be expressed by Taylor series as linear filter decompositions without any noticeable degradation of filter response. All these methods drastically decrease the computation time by cutting it down constant times (e.g. to 0.06 seconds per 1MB image) while achieving very high PSNRpsilas over 45 dB. In addition to the computational advantages, our methods are straightforward to implement. Fatih Porikli |
CVPR | 1 |
| 2008 | Learning on lie groups for invariant detection and trackingabstractThis paper presents a novel learning based tracking model combined with object detection. The existing techniques proceed by linearizing the motion, which makes an implicit Euclidean space assumption. Most of the transformations used in computer vision have matrix Lie group structure. We learn the motion model on the Lie algebra and show that the formulation minimizes a first order approximation to the geodesic error. The learning model is extended to train a class specific tracking function, which is then integrated to an existing pose dependent object detector to build a pose invariant object detection algorithm. The proposed model can accurately detect objects in various poses, where the size of the search space is only a fraction compared to the existing object detection methods. The detection rate of the original detector is improved by more than 90% for large transformations. Oncel Tuzel, Fatih Porikli, Peter Meer |
CVPR | 2 |
| 2008 | Joint tracking and video registration by factorial Hidden Markov modelsabstractTracking moving objects from image sequences obtained by a moving camera is a difficult problem since there exists apparent motion of the static background. It becomes more difficult when the camera motion between the consecutive frames is very large. Traditionally, registration is applied before tracking to compensate for the camera motion using parametric motion models. At the same time, the tracking result highly depends on the performance of registration. This raises problems when there are big moving objects in the scene and the registration algorithm is prone to fail, since the tracker easily drifts away when poor registration results occur. In this paper, we tackle this problem by registering the frames and tracking the moving objects simultaneously within the factorial hidden Markov model framework using particle filters. Under this framework, tracking and registration are not working separately, but mutually benefit each other by interacting. Particles are drawn to provide the candidate geometric transformation parameters and moving object parameters. Background is registered according to the geometric transformation parameters by maximizing a joint gradient function. A state-of-the-art covariance tracker is used to track the moving object. The tracking score is obtained by incorporating both background and foreground information. By using knowledge of the position of the moving objects, we avoid blindly registering the image pairs without taking the moving object regions into account. We apply our algorithm to moving object tracking on numerous image sequences with camera motion and show the robustness and effectiveness of our method. Xue Mei, Fatih Porikli |
ICASSP | 2 |
| 2008 | Likelihood Map Fusion for Visual Object TrackingabstractVisual object tracking can be considered as a figure-ground classification task. In this paper, different features are used to generate a set of likelihood maps for each pixel indicating the probability of that pixel belonging to foreground object or scene background. For example, intensity, texture, motion, saliency and template matching can all be used to generate likelihood maps. We propose a generic likelihood map fusion framework to combine these heterogeneous features into a fused soft segmentation suitable for mean-shift tracking. All the component likelihood maps contribute to the segmentation based on their classification confidence scores (weights) learned from the previous frame. The evidence combination framework dynamically updates the weights such that, in the fused likelihood map, discriminative foreground/background information is preserved while ambiguous information is suppressed. The framework is applied here to track ground vehicles from thermal airborne video, and is also compared to other state-of-the-art algorithms. Zhaozheng Yin, Fatih Porikli, Robert T. Collins |
WACV | 2 |
| 2008 | Pedestrian Detection via Classification on Riemannian ManifoldsabstractWe present a new algorithm to detect pedestrian in still images utilizing covariance matrices as object descriptors. Since the descriptors do not form a vector space, well known machine learning techniques are not well suited to learn the classifiers. The space of d-dimensional nonsingular covariance matrices can be represented as a connected Riemannian manifold. The main contribution of the paper is a novel approach for classifying points lying on a connected Riemannian manifold using the geometry of the space. The algorithm is tested on INRIA and DaimlerChrysler pedestrian datasets where superior detection rates are observed over the previous approaches. Oncel Tuzel, Fatih Porikli, Peter Meer |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Detection of temporarily static regions by processing video at different frame ratesabstractThis paper presents an abandoned item and illegally parked vehicle detection method for single static camera video surveillance applications. By processing the input video at different frame rates, two backgrounds are constructed; one for short-term and another for long-term. Each of these backgrounds is defined as a mixture of Gaussian models, which are adapted using online Bayesian update. Two binary foreground maps are estimated by comparing the current frame with the backgrounds, and motion statistics are aggregated in a likelihood image by applying a set of heuristics to the foreground maps. Likelihood image is then used to differentiate between the pixels that belong to moving objects, temporarily static regions and scene background. Depending on the application, the temporary static regions indicate abandoned items, illegally parked vehicles, objects removed from the scene, etc. The presented pixel-wise method does not require object tracking, thus its performance is not upper-bounded to error prone detection and correspondence tasks that usually fail for crowded scenes. It accurately segments objects even if they are fully occluded. It can also be effectively implemented on a parallel processing architecture. Fatih Porikli |
AVSS | 1 |
| 2007 | Human Detection via Classification on Riemannian ManifoldsabstractWe present a new algorithm to detect humans in still images utilizing covariance matrices as object descriptors. Since these descriptors do not lie on a vector space, well known machine learning techniques are not adequate to learn the classifiers. The space of d-dimensional nonsingular covariance matrices can be represented as a connected Riemannian manifold. We present a novel approach for classifying points lying on a Riemannian manifold by incorporating the a priori information about the geometry of the space. The algorithm is tested on INRIA human database where superior detection rates are observed over the previous approaches. Oncel Tuzel, Fatih Porikli, Peter Meer |
CVPR | 2 |
| 2007 | Probabilistic Visual Tracking via Robust Template Matching and Incremental Subspace UpdateabstractIn this paper, we present a probabilistic algorithm for visual tracking that incorporates robust template matching and incremental sub-space update. There are two template matching methods used in the tracker: one is robust to small perturbation and the other to background clutter. Each method yields a probability of matching. Further, the templates are modeled using mixed probabilities and updated once the templates in the library cannot capture the variation of object appearance. We also model the tracking history using a nonlinear subspace that is described by probabilistic kernel principal components analysis, which provides a third probability. The most-recent tracking result is added to the nonlinear subspace incrementally. This update is performed efficiently by augmenting the kernel Gram matrix with one row and one column. The product of the three probabilities is defined as the observation likelihood used in a particle filter to derive the tracking result. Experimental results demonstrate the efficiency and effectiveness of the proposed algorithm. Xue Mei, Shaohua Kevin Zhou, Fatih Porikli |
ICME | 3 |
| 2007 | Collaborative tracking of objects in EPTZ camerasabstractThis paper addresses the issue of multi-source collaborative object tracking in high-definition (HD) video sequences. Specifically, we propose a new joint tracking paradigm for the multiple stream electronic pan-tilt-zoom (EPTZ) cameras. These cameras are capable of transmitting a low resolution thumbnail (LRT) image of the whole field of view as well as a high-resolution cropped (HRC) image for the target region. We exploit this functionality to perform joint tracking in both low resolution image of the whole field of view as well as high resolution image of the moving target. Our system detects objects of interest in the LRT image by background subtraction and tracks them using iterative coupled refinement in both LRT and HRC images. We compared the performance of our joint tracking system with that of tracking only in the HD mode. The results of our experiments show improved performance in terms of higher frame rates and better localization. Faisal I. Bashir, Fatih Porikli |
VCIP | 2 |
| 2006 | Robust License Plate Detection Using Covariance Descriptor in a Neural Network FrameworkabstractWe present a license plate detection algorithm that employs a novel image descriptor. Instead of using conventional gradient filters and intensity histograms, we compute a covariance matrix of low-level pixel-wise features within a given image window. Unlike the existing approaches, this matrix effectively captures both statistical and spatial properties within the window. We normalize the covariance matrix using local variance scores and restructure the unique coefficients into a feature vector form. Then, we feed these coefficients into a multi-layer neural network. Since no explicit similarity or distance computation is required in this framework, we are able to keep the computational load of the detection process low. To further accelerate the covariance matrix extraction process, we adapt an integral image based data propagation technique. Our extensive analysis shows that the detection process is robust against noise, illumination distortions, and rotation. In addition, the presented method does not require careful fine tuning of the decision boundaries. Fatih Porikli, Tekin Kocak |
AVSS | 1 |
| 2006 | Covariance Tracking using Model Update Based on Lie AlgebraabstractWe propose a simple and elegant algorithm to track nonrigid objects using a covariance based object description and a Lie algebra based update mechanism. We represent an object window as the covariance matrix of features, therefore we manage to capture the spatial and statistical properties as well as their correlation within the same representation. The covariance matrix enables efficient fusion of different types of features and modalities, and its dimensionality is small. We incorporated a model update algorithm using the Lie group structure of the positive definite matrices. The update mechanism effectively adapts to the undergoing object deformations and appearance changes. The covariance tracking method does not make any assumption on the measurement noise and the motion of the tracked objects, and provides the global optimal solution. We show that it is capable of accurately detecting the nonrigid, moving objects in non-stationary camera sequences while achieving a promising detection rate of 97.4 percent. Fatih Porikli, Oncel Tuzel, Peter Meer |
CVPR (1) | 1 |
| 2006 | Region Covariance: A Fast Descriptor for Detection and Classification
Oncel Tuzel, Fatih Porikli, Peter Meer |
ECCV (2) | 2 |
| 2006 | Fast Construction of Covariance Matrices for Arbitrary Size Image WindowsabstractWe propose an integral image based algorithm to extract feature covariance matrices of all possible rectangular regions within a given image. Covariance is an essential indicator of how much the deviation of two or more variables match. In our case, these variables correspond to point-wise features, e.g. coordinates, color values, gradients, edge magnitude and orientation, local histograms, filter responses, etc. We significantly improve the speed of the covariance computation by taking advantage of the spatial arrangement of image points using integral images, which are intermediate representations used for calculation of region sums. Each point of the integral image corresponds to the summation of all point values inside the feature image rectangle bounded by the upper left corner and the point of interest. Using this representation, any rectangular region sum can be computed in constant time. We follow a similar idea for fast calculation of region covariance. We construct integral images for all separate features as well as integral images of the multiplication of any two feature combinations. Using these set of integral images and region corner point coordinates, we directly extract the covariance matrix coefficients. We show that the proposed method reduces the computational load to quadratic time. Fatih Porikli, Oncel Tuzel |
ICIP | 1 |
| 2006 | Shape-Regulated Particle Filtering for Tracking Non-Rigid ObjectsabstractThis paper presents an active contour based algorithm for tracking non-rigid objects in heavily cluttered scenes. We decompose the non-rigid contour tracking problem into three subproblems: 2D motion estimation, deformation detection, and shape regulation. First, we employ a particle filter to estimate the affine transform parameters between successive frames. Second, by using a dynamic object model, we generate a probabilistic map of deformation to reshape its contour. Finally, we project the updated model onto a trained shape subspace to constrain deformations to be within possible object appearances. Our experiments show that the proposed algorithm significantly improves the performance of the tracker. Jie Shao 0007, Rama Chellappa, Fatih Porikli |
ICIP | 3 |
| 2005 | Performance evaluation of event detection solutions: the CREDS experienceabstractIn video surveillance projects, automatic and real-time event detection solutions are required to guarantee an efficient and cost-effective use of the infrastructure. Many solutions have been proposed to automatically detect a variety of events of interest. However, not all solutions and technologies may satisfy all the requirements of the surveillance scenario. For this reason, performance evaluation of existing event detection solutions becomes an important step in the deployment of video surveillance projects. In this paper, we propose a practical approach that aims at minimizing the ground truth generation problem and the expertise required to evaluate and compare the results by introducing specific requirements of specific event detection scenarios. This approach is believed to be applicable for an initial evaluation of candidate solutions to a specific surveillance scenario before more exhaustive tests in an integrated environment. The proposed method is under evaluation in the framework of the challenge of real-time event detection solutions (CREDS). Francesco Ziliani, Sergio A. Velastin, Fatih Porikli, Lucio Marcenaro, Timothy P. Kelliher, Andrea Cavallaro, Philippe Bruneaut |
AVSS | 3 |
| 2005 | Integral Histogram: A Fast Way To Extract Histograms in Cartesian SpacesabstractWe present a novel method, which we refer as an integral histogram, to compute the histograms of all possible target regions in a Cartesian data space. Our method has three distinct advantages: 1) It is computationally superior to the conventional approach. The integral histogram method makes it possible to employ even an exhaustive search process in real-time, which was impractical before. 2) It can be extended to higher data dimensions, uniform and nonuniform bin formations, and multiple target scales without sacrificing its computational advantages. 3) It enables the description of higher level histogram features. We exploit the spatial arrangement of data points, and recursively propagate an aggregated histogram by starting from the origin and traversing through the remaining points along either a scan-line or a wave-front. At each step, we update a single bin using the values of integral histogram at the previously visited neighboring data points. After the integral histogram is propagated, histogram of any target region can be computed easily by using simple arithmetic operations. Fatih Porikli |
CVPR (1) | 1 |
| 2005 | Shadow Flow: A Recursive Method to Learn Moving Cast ShadowsabstractWe present a novel algorithm to detect and remove cast shadows in a video sequence by taking advantage of the statistical prevalence of the shadowed regions over the object regions. We model shadows using multivariate Gaussians. We apply a weak classifier as a pre-filter. We project shadow models into a quantized color space to update a shadow flow function. We use shadow flow, background models, and current frame to determine the shadow and object regions. This method has several advantages: It does not require a color space transformation. We pose the problem in the RGB color space, and we can carry out the same analysis in other Cartesian spaces as well. It is data-driven and adapts to the changing shadow conditions. In other words, accuracy of our method is not limited by the preset values. Furthermore, it does not assume any 3D models for the target objects or tracking of the cast shadows between frames. Our results show that the detection performance is superior than the benchmark method. Fatih Porikli, Jay Thornton |
ICCV | 1 |
| 2005 | Multi-Kernel Object TrackingabstractIn this paper, we present an object tracking algorithm for the low-frame-rate video in which objects have fast motion. The conventional mean-shift tracking fails in case the relocation of an object is large and its regions between the consecutive frames do not overlap. We provide a solution to this problem by using multiple kernels centered at the high motion areas. In addition, we improve the convergence properties of the mean-shift by integrating two likelihood terms, background and template similarities, in the iterative update mechanism. Our simulations prove the effectiveness of the proposed method Fatih Porikli, Oncel Tuzel |
ICME | 1 |
| 2004 | A hidden markov model framework for traffic event detection using video featuresabstractA novel approach for highway traffic event detection in video is presented. The proposed algorithm extracts event features directly from compressed video and detects traffic event using a Gaussian mixture hidden Markov model (GMHMM). First, an invariant feature vector is extracted from discrete cosine transform (DCT) domain and macro-block vectors after MPEG video stream is parsed. The extracted feature vector accurately describes the change of traffic state and is robust towards different camera setups and illumination situations, such as sunny, cloud, and night. Six traffic patterns are studied and a GMHMM is trained to model these patterns in offline stage. Then, Viterbi algorithm is used to determine the most likely traffic condition. The proposed algorithm is efficient both in terms of computational complexity and memory requirement. The experimental results prove the system has a high detection rate. The presented model based system can be easily extended for detection of similar traffic events. Xiaokun Li, Fatih Porikli |
ICIP | 2 |
| 2004 | Nonlinear warping function recovery by scan-line search using dynamic programming
Fatih Porikli |
ICIP | 1 |
| 2004 | Learning object trajectory patterns by spectral clusteringabstractWe develop a trajectory pattern learning method that has two significant advantages over past work. First, we represent trajectories in the HMM parameter space, thus we overcome the normalization problems of existing methods. Second, we determine common trajectory paths by analyzing the optimal cluster number rather than using a predefined number of clusters. We compute affinity matrices and apply eigenvector decomposition to find clusters. We prove that the number of clusters governs the number of eigenvectors used to span the feature affinity space. We are thus able to determine automatically the optimal number of patterns. We show that the proposed algorithm accurately detects common paths for various camera setups Fatih Porikli |
ICME | 1 |
| 2003 | Inter-camera color calibration by correlation model functionabstractA novel solution to the inter-camera color calibration problem, which is very important for multicamera systems is presented. We propose a distance metric and a model function to evaluate the inter-camera radiometric properties. Instead of depending on the shape assumptions of brightness transfer function to find separate radiometric responses, we derive a nonparametric function to model color distortion for pair-wise camera combinations. Our method is based on correlation matrix analysis and dynamic programming. The correlation matrix is computed from three 1-D color histograms, and the model function is obtained from a minimum cost path traced within the matrix. The model function enables accurate compensation of color mismatches, which cannot be done with conventional distance metrics. Furthermore, we show that our metric can be reduced to other commonly used metrics with suitable simplification. Our simulations prove the effectiveness of the proposed method even for severe color distortions. Fatih Porikli |
ICIP (2) | 1 |
| 2003 | Multi-camera calibration, object tracking and query generationabstractAn automatic object tracking and video summarization method for multi-camera systems with a large number of non-overlapping field-of-view cameras is explained. In this framework, video sequences are stored for each object as opposed to storing a sequence for each camera. Object-based representation enables annotation of video segments, and extraction of content semantics for further analysis. We also present a novel solution to the inter-camera color calibration problem. The transitive model function enables effective compensation for lighting changes and radiometric distortions for large-scale systems. After initial calibration, objects are tracked at each camera by background subtraction and mean-shift analysis. The correspondence of objects between different cameras is established by using a Bayesian belief network. This framework empowers the user to get a concise response to queries such as "which locations did an object visit on Monday and what did it do there?". Fatih Porikli, Ajay Divakaran |
ICME | 1 |
| 2002 | Automatic threshold determination of centroid-linkage region growing by MPEG-7 dominant color descriptorsabstractThe emerging MPEG-7 standard embodies a visual descriptor that is associated with the dominant colors of an image. A threshold adaptation method for region based image and video segmentation that takes advantage of the MPEG-7 dominant color descriptor is presented. This method enables assignment of region growing parameters without any low-level processing. In the standard, it is proposed that the dominant colors be extracted by clustering of color histograms. This property is used to determine color homogeneity that is formulated into the Lorentzian-based color distance norm and corresponding thresholds. The proposed algorithm is compared with other region growing algorithms, and results show that the threshold adaptation performs faster and more robustly. Fatih Porikli |
ICIP (1) | 1 |
| 2002 | Constrained video object segmentation by color masks and MPEG-7 descriptorsabstractWe present an automatic and computationally conservative boundary extraction method using available priori information in a framework consists of change detection mask, region growing, and trajectory motion. Instead of segmenting a entire video frame, only the regions belong to a target object specified by a set of rules are detected. One example of such rules is skin color features of a human body part. The processing domain is limited to the pixels that satisfy color or geometric rules. These rules are represented as a detection mask and implemented as a look-up table. As a result, significant computational reduction and real-time performance are achieved. The framework utilizes color consistency within a centroid-linkage growing technique to grow initial regions. Region seeds are selected among the pixels in the detection mask. The similarity thresholds are adapted from the MPEG-7 dominant color descriptors. The segmentation results of a frame diffused to the next frame and region statistics such as trajectory, percentage of the changed pixels, etc., are registered to determine the moving regions. A computational load comparison of the constrained region growing and regular region growing shows significant reduction in the complexity. Fatih Porikli, Yao Wang 0001 |
ICME (1) | 1 |
| 2001 | Object Segmentation of Color Video Sequences
Fatih Porikli |
CAIP | 1 |
| 2001 | Dynamic bandwidth allocation with optimal number of renegotiations in ATM networksabstractThis paper introduces a scheme for online dynamic bandwidth allocation for variable bit rate (VBR) traffic over ATM networks. The presented method determines the optimum bandwidth renegotiation time and bandwidth amount to allocate to a VBR traffic source by minimizing predefined cost functions. A traffic rate predictor designed by wavelets is provided as feedback to the system. The results show that the introduced scheme minimizes both under-utilization of the available capacity and queuing delays. The proposed method can also be deployed in weighted fair queuing disciplines to update dynamically a weight coefficient assigned to an application. Fatih Porikli, Zafer Sahinoglu |
ICCCN | 1 |
| 2001 | Accurate detection of edge orientation for color and multi-spectral imageryabstractThe frequency domain properties of an image are used for precise detection of edge orientation in color and multi-spectral imagery. The orientation estimation is established as a minimization problem, formulated as a tensor method, and simplified by solving its dual in terms of spatial partial derivatives of the image. First, the spectral density distribution around each pixel is obtained. The edge orientation is determined by fitting a straight line to this distribution. A matching error is devised in tensor form, and minimized by rotating the frequency domain principal axes. The orientation is computed from the spatial derivatives by transposing frequency domain operations to the spatial domain. The estimated edge orientations and magnitudes for different bands are converted to vectors and summed in the vector domain. A comparison of this method with the widely used estimators shows that the adapted tensor method improves the estimation precision even in the presence of extreme noise. Fatih Porikli |
ICIP (1) | 1 |
| 2001 | An unsupervised multi-resolution object extraction algorithm using video-cubeabstractWe propose a fast video object segmentation method that detects object boundaries accurately, and does not require any user assistance. Video streams are considered as 3D data, called video-cubes, to take advantage of 3D signal processing techniques. After a video sequence is filtered, marker nodes are selected from the color gradient. A volume around each marker is grown by using color/texture distance criteria. Then volumes that have similar characteristics are merged. Self-descriptors for each volume, mutual descriptors for each pair of volumes are computed. These descriptors capture motion and spatial information of volumes. In the clustering stage, volumes are classified into objects in a fine-to-coarse hierarchy. While applying and relaxing descriptor based adaptive, similarity scores are estimated for each possible pair-wise combination of volumes. The pair that gives the maximum score is clustered iteratively. Finally, an object-based multi-resolution representation tree is assembled. Fatih Porikli, Yao Wang 0001 |
ICIP (2) | 1 |
| 1997 | Adaptive stripe based patch matching for depth estimationabstractA novel stereo matching technique for depth estimation in stereoscopic image pairs is presented. The input image pair is preprocessed in the intensity domain and together with edge maps an adaptive mesh in which individual elements approximate linearly modeled regions are obtained. Then, an iterative stripe based, quadrilateral patch matching technique is employed to estimate the depth map from the image pair in a hierarchical manner. Finally, the resultant map is postprocessed to smooth the depth map at the patch borders. The quality of the test results demonstrates the effectiveness of the technique. Fatih Porikli, Yao Wang 0001, Cassandra T. Swain |
ICASSP | 1 |
| 1997 | Stripe mesh based disparity estimation by using 3-D Hough transformabstractThis correspondence presents a novel matching technique for disparity estimation. The technique combines advantages of an adaptive stripe-based mesh structure and Hough transform. First, a mesh that is composed of triangular patches and fits to the depth changes is generated by using edge maps. In each patch, the depth of the scene is approximated by a surface. A search space is built by difference maps that are obtained by subtracting the left image from the shifted versions of the right image. From the search space, 3-D Hough spaces are produced such that a point of the Hough space represents a surface in the search space. Then, a matching algorithm finds the combination of surfaces that gives minimum matching error in a neighborhood around patches by using Hough spaces. Continuity and smoothness constraints are formulated as probability density functions that modify the Hough spaces of the following patches by using the estimated surfaces for the previous ones. Our experiments demonstrate the accuracy of the proposed method. Fatih Porikli |
ICIP (3) | 1 |