VLDB 2026 Research / reviewers in the wild / expert
Hakan Bilen
dblp:97/2993
· DBLP profile ↗
64ranked-venue papers
14as first author
38since 2021 · last 2025
0000-0002-6947-6918ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 53 · 12 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 35 · 7 first-author · 21 since 2021Systems, architecture and hardware · 2 · 2 first-authorComputer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Odd-One-Out: Anomaly Detection by Comparing with NeighborsabstractThis paper introduces a novel anomaly detection (AD) problem aimed at identifying ‘odd-looking’ objects within a scene by comparing them to other objects present. Unlike traditional AD benchmarks with fixed anomaly criteria, our task detects anomalies specific to each scene by inferring a reference group of regular objects. To address occlusions, we use multiple views of each scene as input, construct 3D object-centric models for each instance from 2D views, enhancing these models with geometrically consistent part-aware representations. Anomalous objects are then detected through cross-instance comparison. We also introduce two new benchmarks, ToysAD-8K and PartsAD-15K as testbeds for future research in this task. We provide a comprehensive analysis of our method quantitatively and qualitatively on these benchmarks. Ankan Bhunia, Changjian Li 0001, Hakan Bilen |
CVPR | 3 |
| 2025 | DepthCues: Evaluating Monocular Depth Perception in Large Vision ModelsabstractLarge-scale pre-trained vision models are becoming increasingly prevalent, offering expressive and generalizable visual representations that benefit various downstream tasks. Recent studies on the emergent properties of these models have revealed their high-level geometric understanding, in particular in the context of depth perception. However, it remains unclear how depth perception arises in these models without explicit depth supervision provided during pre-training. To investigate this, we examine whether the monocular depth cues, similar to those used by the human visual system, emerge in these models. We introduce a new benchmark, DepthCues, designed to evaluate depth cue understanding, and present findings across 20 diverse and representative pre-trained vision models. Our analysis shows that human-like depth cues emerge in more recent larger models. We also explore enhancing depth perception in large vision models by fine-tuning on DepthCues, and find that even without dense depth supervision, this improves depth estimation. To support further research, our benchmark and evaluation code will be made publicly available for studying depth perception in vision models. Duolikun Danier, Mehmet Aygun, Changjian Li 0001, Hakan Bilen, Oisin Mac Aodha |
CVPR | 4 |
| 2025 | Interactive Anomaly Detection for Articulated Objects via Motion AnticipationabstractThis paper presents a novel problem, interactive anomaly detection (AD) for articulated objects, and introduces a tailored solution that detects functional anomalies by integrating vision, interaction, and anticipation. Unlike traditional AD methods that rely on passive visual observations, our approach actively manipulates objects to reveal anomalies that would otherwise remain hidden. Our method learns to generate a sequence of actions to interact exclusively with normal objects and to anticipate the resulting normal motion. During inference, the model applies predicted actions to the object and compares the observed motion with the anticipated motion to detect anomalies. Additionally, we introduce a new benchmark, PartNet-IAD, for interactive AD, which includes articulated objects with realistic functional anomalies. Experiments show strong generalization to detect anomalies in both seen and unseen object categories. Code and dataset will be released. Ankan Bhunia, Changjian Li 0001, Hakan Bilen |
NeurIPS | 3 |
| 2025 | Jamais Vu: Exposing the Generalization Gap in Supervised Semantic CorrespondenceabstractSemantic correspondence (SC) aims to establish semantically meaningful matches across different instances of an object category. We illustrate how recent supervised SC methods remain limited in their ability to generalize beyond sparsely annotated training keypoints, effectively acting as keypoint detectors. To address this, we propose a novel approach for learning dense correspondences by lifting 2D keypoints into a canonical 3D space using monocular depth estimation. Our method constructs a continuous canonical manifold that captures object geometry without requiring explicit 3D supervision or camera annotations. Additionally, we introduce SPair-U, an extension of SPair-71k with novel keypoint annotations, to better assess generalization. Experiments not only demonstrate that our model significantly outperforms supervised baselines on unseen keypoints, highlighting its effectiveness in learning robust correspondences, but that unsupervised baselines outperform supervised counterparts when generalized across different datasets. Octave Mariotti, Zhipeng Du, Yash Bhalgat, Oisin Mac Aodha, Hakan Bilen |
NeurIPS | 5 |
| 2025 | Spatially-Adaptive Hash Encodings for Neural Surface ReconstructionabstractPositional encodings are a common component of neural scene reconstruction methods, and provide a way to bias the learning of neural fields towards coarser or finer representations. Current neural surface reconstruction methods use a “one-size-fits-all” approach to encoding, choosing a fixed set of encoding functions, and therefore bias, across all scenes. Current state-of-the-art surface reconstruction approaches leverage grid-based multi-resolution hash encoding in order to recover high-detail geometry. We propose a learned approach which allows the network to choose its encoding basis as a function of space, by masking the contribution of features stored at separate grid resolutions. The resulting spatially adaptive approach allows the network to fit a wider range of frequencies without introducing noise. We test our approach on standard benchmark surface reconstruction datasets and achieve state-of-the-art performance on two benchmark datasets. Thomas Walker, Octave Mariotti, Amir Vaxman, Hakan Bilen |
WACV | 4 |
| 2024 | Looking 3D: Anomaly Detection with 2D-3D AlignmentabstractAutomatic anomaly detection based on visual cues holds practical significance in various domains, such as manu-facturing and product quality assessment. This paper introduces a new conditional anomaly detection problem, which involves identifying anomalies in a query image by comparing it to a reference shape. To address this challenge, we have created a large dataset, BrokenChairs-180K, consisting of around 180K images, with diverse anomalies, geometries, and textures paired with 8,143 reference 3D shapes. To tackle this task, we have proposed a novel transformer-based approach that explicitly learns the correspondence between the query image and reference 3D shape via feature alignment and leverages a customized attention mechanism for anomaly detection. Our approach has been rigorously evaluated through comprehensive experiments, serving as a benchmark for future research in this domain. Ankan Bhunia, Changjian Li 0001, Hakan Bilen |
CVPR | 3 |
| 2024 | Improving Semantic Correspondence with Viewpoint-Guided Spherical MapsabstractRecent self-supervised models produce visual features that are not only effective at encoding image-level, but also pixel-level, semantics. They have been reported to obtain impressive results for dense visual semantic correspondence estimation, even outperforming fully-supervised methods. Nevertheless, these models still fail in the pres-ence of challenging image characteristics such as symme-tries and repeated parts. To address these limitations, we propose a new semantic correspondence estimation method that supplements state-of-the-art self-supervised features with 3D understanding via a weak geometric spherical prior. Compared to more involved 3D pipelines, our model provides a simple and effective way of injecting informative geometric priors into the learned representation while requiring only weak viewpoint information. We also propose a new evaluation metric that better accounts for re-peated part and symmetry-induced mistakes. We show that our method succeeds in distinguishing between symmetric views and repeated parts across many object categories in the challenging SPair-71 k dataset and also in generalizing to previously unseen classes in the AwA dataset. Octave Mariotti, Oisin Mac Aodha, Hakan Bilen |
CVPR | 3 |
| 2024 | RECANTFormer: Referring Expression Comprehension with Varying Numbers of TargetsabstractThe Generalized Referring Expression Comprehension (GREC) task extends classic REC by generating image bounding boxes for objects referred to in natural language expressions, which may indicate zero, one, or multiple targets. This generalization enhances the practicality of REC models for diverse real-world applications. However, the presence of varying numbers of targets in samples makes GREC a more complex task, both in terms of training supervision and final prediction selection strategy. Addressing these challenges, we introduce RECANTFormer, a one-stage method for GREC that combines a decoder-free (encoder-only) transformer architecture with DETR-like Hungarian matching. Our approach consistently outperforms baselines by significant margins in three GREC datasets. Bhathiya Hemanthage, Hakan Bilen, Phil J. Bartie, Christian Dondrup, Oliver Lemon |
EMNLP | 2 |
| 2024 | Multi-task Learning with 3D-Aware RegularizationabstractDeep neural networks have become the standard solution for designing models that can perform multiple dense computer vision tasks such as depth estimation and semantic segmentation thanks to their ability to capture complex correlations in high dimensional feature space across tasks. However, the cross-task correlations that are learned in the unstructured feature space can be extremely noisy and susceptible to overfitting, consequently hurting performance. We propose to address this problem by introducing a structured 3D-aware regularizer which interfaces multiple tasks through the projection of features extracted from an image encoder to a shared 3D feature space and decodes them into their task output space through differentiable rendering. We show that the proposed method is architecture agnostic and can be plugged into various prior multi-task backbones to improve their performance; as we evidence using standard benchmarks NYUv2 and PASCAL-Context. Wei-Hong Li 0001, Steven McDonagh 0001, Ales Leonardis, Hakan Bilen |
ICLR | 4 |
| 2024 | Articulate your NeRF: Unsupervised articulated object modeling via conditional view synthesisabstractWe propose a novel unsupervised method to learn pose and part-segmentation of articulated objects with rigid parts.
Given two observations of an object in different articulation states, our method learns the geometry and appearance of object parts by using an implicit model from the first observation, distills the part segmentation and articulation from the second observation while rendering the latter observation.
Additionally, to tackle the complexities in the joint optimization of part segmentation and articulation, we propose a voxel grid based initialization strategy and a decoupled optimization procedure.
Compared to the prior unsupervised work, our model obtains significantly better performance, generalizes to objects with multiple parts while it can be efficiently from few views for the latter observation. Jianning Deng, Kartic Subr, Hakan Bilen |
NeurIPS | 3 |
| 2024 | Divide and Conquer: Rethinking Ambiguous Candidate Identification in Multimodal Dialogues with Pseudo-LabellingabstractAmbiguous Candidate Identification (ACI) in multimodal dialogue is the task of identifying all potential objects that a user's utterance could be referring to in a visual scene, in cases where the reference cannot be uniquely determined.End-to-end models are the dominant approach for this task, but have limited real-world applicability due to unrealistic inference-time assumptions such as requiring predefined catalogues of items.Focusing on a more generalized and realistic ACI setup, we demonstrate that a modular approach, which first emphasizes language-only reasoning over dialogue context before performing vision-language fusion, significantly outperforms end-to-end trained baselines.To mitigate the lack of annotations for training the language-only module (student), we propose a pseudo-labelling strategy with a prompted Large Language Model (LLM) as the teacher. Bhathiya Hemanthage, Christian Dondrup, Hakan Bilen, Oliver Lemon |
SIGDIAL | 3 |
| 2024 | Universal Representations: A Unified Look at Multiple Task and Domain LearningabstractAbstract We propose a unified look at jointly learning multiple vision tasks and visual domains through universal representations, a single deep neural network. Learning multiple problems simultaneously involves minimizing a weighted sum of multiple loss functions with different magnitudes and characteristics and thus results in unbalanced state of one loss dominating the optimization and poor results compared to learning a separate model for each problem. To this end, we propose distilling knowledge of multiple task/domain-specific networks into a single deep neural network after aligning its representations with the task/domain-specific ones through small capacity adapters. We rigorously show that universal representations achieve state-of-the-art performances in learning of multiple dense prediction problems in NYU-v2 and Cityscapes, multiple image classification problems from diverse domains in Visual Decathlon Dataset and cross-domain few-shot learning in MetaDataset. Finally we also conduct multiple analysis through ablation and qualitative studies. Wei-Hong Li 0001, Xialei Liu, Hakan Bilen |
Int. J. Comput. Vis. | 3 |
| 2023 | RenderDiffusion: Image Diffusion for 3D Reconstruction, Inpainting and GenerationabstractDiffusion models currently achieve state-of-the-art performance for both conditional and unconditional image generation. However, so far, image diffusion models do not support tasks required for 3D understanding, such as view-consistent 3D generation or single-view object reconstruction. In this paper, we present RenderDiffusion, the first diffusion model for 3D generation and inference, trained using only monocular 2D supervision. Central to our method is a novel image denoising architecture that generates and renders an intermediate three-dimensional representation of a scene in each denoising step. This enforces a strong inductive structure within the diffusion process, providing a 3D consistent representation while only requiring 2D supervision. The resulting 3D representation can be rendered from any view. We evaluate RenderDiffusion on FFHQ, AFHQ, ShapeNet and CLEVR datasets, showing competitive performance for generation of 3D scenes and inference of 3D scenes from 2D images. Additionally, our diffusion-based approach allows us to use 2D inpainting to edit 3D scenes. Titas Anciukevicius, Zexiang Xu, Matthew Fisher, Paul Henderson, Hakan Bilen, Niloy J. Mitra, Paul Guerrero 0001 |
CVPR | 5 |
| 2023 | Learning Action Changes by Measuring Verb-Adverb Textual RelationshipsabstractThe goal of this work is to understand the way actions are performed in videos. That is, given a video, we aim to predict an adverb indicating a modification applied to the action (e.g. cut “finely”). We cast this problem as a regression task. We measure textual relationships between verbs and adverbs to generate a regression target representing the action change we aim to learn. We test our approach on a range of datasets and achieve state-of-the-art results on both adverb prediction and antonym classification. Furthermore, we outperform previous work when we lift two commonly assumed conditions: the availability of action labels during testing and the pairing of adverbs as antonyms. Existing datasets for adverb recognition are either noisy, which makes learning difficult, or contain actions whose appearance is not influenced by adverbs, which makes evaluation less reliable. To address this, we collect a new high quality dataset: Adverbs in Recipes (AIR). We focus on instructional recipes videos, curating a set of actions that exhibit meaningful visual changes when performed differently. Videos in AIR are more tightly trimmed and were manually reviewed by multiple annotators to ensure high labelling quality. Results show that models learn better from AIR given its cleaner videos. At the same time, adverb prediction on AIR is challenging, demonstrating that there is considerable room for improvement. Davide Moltisanti, Frank Keller, Hakan Bilen, Laura Sevilla-Lara |
CVPR | 3 |
| 2023 | Semi-supervised multimodal coreference resolution in image narrationsabstractIn this paper, we study multimodal coreference resolution, specifically where a longer descriptive text, i.e., a narration is paired with an image.This poses significant challenges due to fine-grained image-text alignment, inherent ambiguity present in narrative language, and unavailability of large annotated training sets.To tackle these challenges, we present a data efficient semi-supervised approach that utilizes image-narration pairs to resolve coreferences and narrative grounding in a multimodal context.Our approach incorporates losses for both labeled and unlabeled data within a crossmodal framework.Our evaluation shows that the proposed approach outperforms strong baselines both quantitatively and qualitatively, for the tasks of coreference resolution and narrative grounding. Arushi Goel, Basura Fernando, Frank Keller, Hakan Bilen |
EMNLP | 4 |
| 2023 | Who are you referring to? Coreference resolution in image narrationsabstractCoreference resolution aims to identify words and phrases which refer to the same entity in a text, a core task in natural language processing. In this paper, we extend this task to resolving coreferences in long-form narrations of visual scenes. First, we introduce a new dataset with annotated coreference chains and their bounding boxes, as most existing image-text datasets only contain short sentences without coreferring expressions or labeled chains. We propose a new technique that learns to identify coref-erence chains using weak supervision, only from image-text pairs and a regularization using prior linguistic knowledge. Our model yields large performance gains over several strong baselines in resolving coreferences. We also show that coreference resolution helps improve grounding narratives in images. Arushi Goel, Basura Fernando, Frank Keller, Hakan Bilen |
ICCV | 4 |
| 2023 | Accelerating Self-Supervised Learning via Efficient Training StrategiesabstractRecently the focus of the computer vision community has shifted from expensive supervised learning towards self-supervised learning of visual representations. While the performance gap between supervised and self-supervised has been narrowing, the time for training self-supervised deep networks remains an order of magnitude larger than its supervised counterparts, which hinders progress, imposes carbon cost, and limits societal benefits to institutions with substantial resources. Motivated by these issues, this paper investigates reducing the training time of recent self-supervised methods by various model-agnostic strategies that have not been used for this problem. In particular, we study three strategies: an extendable cyclic learning rate schedule, a matching progressive augmentation magnitude and image resolutions schedule, and a hard positive mining strategy based on augmentation difficulty. We show that all three methods combined lead up to 2.7 times speed-up in the training time of several self-supervised methods while retaining comparable performance to the standard self-supervised learning setting. Mustafa Taha Koçyigit, Timothy M. Hospedales, Hakan Bilen |
WACV | 3 |
| 2023 | Dataset Condensation with Distribution MatchingabstractComputational cost of training state-of-the-art deep models in many learning problems is rapidly increasing due to more sophisticated models and larger datasets. A recent promising direction for reducing training cost is dataset condensation that aims to replace the original large training set with a significantly smaller learned synthetic set while preserving the original information. While training deep models on the small set of condensed images can be extremely fast, their synthesis remains computationally expensive due to the complex bi-level optimization and secondorder derivative computation. In this work, we propose a simple yet effective method that synthesizes condensed images by matching feature distributions of the synthetic and original training images in many sampled embedding spaces. Our method significantly reduces the synthesis cost while achieving comparable or better performance. Thanks to its efficiency, we apply our method to more realistic and larger datasets with sophisticated neural architectures and obtain a significant performance boost1. We also show promising practical benefits of our method in continual learning and neural architecture search. Bo Zhao 0038, Hakan Bilen |
WACV | 2 |
| 2022 | ViewNeRF: Unsupervised Viewpoint Estimation Using Category-Level Neural Radiance Fields
Octave Mariotti, Oisin Mac Aodha, Hakan Bilen |
BMVC | 3 |
| 2022 | Not All Relations are Equal: Mining Informative Labels for Scene Graph GenerationabstractScene graph generation (SGG) aims to capture a wide variety of interactions between pairs of objects, which is essential for full scene understanding. Existing SGG methods trained on the entire set of relations fail to acquire complex reasoning about visual and textual correlations due to various biases in training data. Learning on trivial relations that indicate generic spatial configuration like ‘on’ instead of informative relations such as ‘parked on’ does not enforce this complex reasoning, harming generalization. To address this problem, we propose a novel framework for SGG training that exploits relation labels based on their informativeness. Our model-agnostic training procedure imputes missing informative relations for less informative samples in the training data and trains a SGG model on the imputed labels along with existing annotations. We show that this approach can successfully be used in conjunction with state-of-the-art SGG methods and improves their performance significantly in multiple metrics on the standard Visual Genome benchmark. Furthermore, we obtain considerable improvements for unseen triplets in a more challenging zero-shot setting. Arushi Goel, Basura Fernando, Frank Keller, Hakan Bilen |
CVPR | 4 |
| 2022 | Cross-domain Few-shot Learning with Task-specific AdaptersabstractIn this paper, we look at the problem of cross-domain few-shot classification that aims to learn a classifier from previously unseen classes and domains withfew labeled samples. Recent approaches broadly solve this problem by pa-rameterizing their few-shot classifiers with task-agnostic and task-specific weights where the former is typically learned on a large training set and the latter is dynamically predicted through an auxiliary network conditioned on a small support set. In this work, we focus on the estimation of the latter, and propose to learn task-specific weights from scratch directly on a small support set, in contrast to dynamically estimating them. In particular, through systematic analysis, we show that task-specific weights through parametric adapters in matrix form with residual connections to multiple intermediate layers of a backbone network significantly improves the per-formance of the state-of-the-art models in the Meta-Dataset benchmark with minor additional cost. Wei-Hong Li 0001, Xialei Liu, Hakan Bilen |
CVPR | 3 |
| 2022 | Learning Multiple Dense Prediction Tasks from Partially Annotated DataabstractDespite the recent advances in multi-task learning of dense prediction problems, most methods rely on expensive labelled datasets. In this paper, we present a label efficient approach and look at jointly learning of multiple dense prediction tasks on partially annotated data (i.e. not all the task labels are available for each image), which we call multi-task partially-supervised learning. We propose a multi-task training procedure that successfully leverages task relations to supervise its multi-task learning when data is partially annotated. In particular, we learn to map each task pair to a joint pairwise task-space which enables sharing information between them in a computationally efficient way through another network conditioned on task pairs, and avoids learning trivial cross-task relations by retaining high-level information about the input image. We rigorously demonstrate that our proposed method effectively exploits the images with unlabelled tasks and outperforms existing semi-supervised learning approaches and related methods on three standard benchmarks. Wei-Hong Li 0001, Xialei Liu, Hakan Bilen |
CVPR | 3 |
| 2022 | CAFE: Learning to Condense Dataset by Aligning FeaturesabstractDataset condensation aims at reducing the network training effort through condensing a cumbersome training set into a compact synthetic one. State-of-the-art approaches largely rely on learning the synthetic data by matching the gradients between the real and synthetic data batches. Despite the intuitive motivation and promising results, such gradient-based methods, by nature, easily overfit to a biased set of samples that produce dominant gradients, and thus lack a global supervision of data distribution. In this paper, we propose a novel scheme to Condense dataset by Aligning FEatures (CAFE), which explicitly attempts to preserve the real-feature distribution as well as the discriminant power of the resulting synthetic set, lending itself to strong generalization capability to various architectures. At the heart of our approach is an effective strategy to align features from the real and synthetic data across various scales, while accounting for the classification of real samples. Our scheme is further backed up by a novel dynamic bi-level optimization, which adaptively adjusts parameter updates to prevent over-/under-fitting. We validate the proposed CAFE across various datasets, and demonstrate that it generally outperforms the state of the art: on the SVHN dataset, for example, the performance gain is up to 11%. Extensive experiments and analysis verify the effectiveness and necessity of proposed designs. Kai Wang 0036, Bo Zhao 0038, Shuo Yang 0006, Shuo Wang 0001, Guan Huang 0003, Hakan Bilen, Xinchao Wang, Yang You 0001 |
CVPR | 8 |
| 2022 | 3D Equivariant Graph Implicit Functions
Yunlu Chen, Basura Fernando, Hakan Bilen, Matthias Nießner, Efstratios Gavves |
ECCV (3) | 3 |
| 2022 | Visual Representation Learning over Latent Domains
Lucas Deecke, Timothy M. Hospedales, Hakan Bilen |
ICLR | 3 |
| 2022 | Learning to Annotate Part Segmentation with Gradient Matching
Yu Yang 0011, Xiaotian Cheng, Hakan Bilen, Xiangyang Ji |
ICLR | 3 |
| 2022 | Learning to Predict Keypoints and Structure of Articulated Objects without SupervisionabstractReasoning about the structure and motion of novel object classes is a core ability in human cognition, crucial for manipulating objects and predicting their possible motion. We present a method that learns to infer the skeleton structure of a novel articulated object from a single image, in terms of joints and rigid links connecting them. The model learns without supervision from a dataset of objects having diverse structures, in different poses and states of articulation. To achieve this, it is trained to explain the differences between pairs of images in terms of a latent skeleton that defines how to transform one into the other. Experiments on several datasets show that our model predicts joint locations significantly more accurately than prior works on unsupervised keypoint discovery; moreover, unlike existing methods, it can predict varying numbers of joints depending on the observed object. It also successfully predicts the connections between joints, even for structures not seen during training. Titas Anciukevicius, Paul Henderson, Hakan Bilen |
ICPR | 3 |
| 2022 | Distilling Representations from GAN Generator via Squeeze and SpanabstractIn recent years, generative adversarial networks (GANs) have been an actively studied topic and shown to successfully produce high-quality realistic images in various domains. The controllable synthesis ability of GAN generators suggests that they maintain informative, disentangled, and explainable image representations, but leveraging and transferring their representations to downstream tasks is largely unexplored. In this paper, we propose to distill knowledge from GAN generators by squeezing and spanning their representations. We \emph{squeeze} the generator features into representations that are invariant to semantic-preserving transformations through a network before they are distilled into the student network. We \emph{span} the distilled representation of the synthetic domain to the real domain by also using real training data to remedy the mode collapse of GANs and boost the student network performance in a real domain. Experiments justify the efficacy of our method and reveal its great significance in self-supervised representation learning. Code is available at https://github.com/yangyu12/squeeze-and-span. Yu Yang 0011, Xiaotian Cheng, Chang Liu 0030, Hakan Bilen, Xiangyang Ji |
NeurIPS | 4 |
| 2022 | CartaGenie: Context-Driven Synthesis of City-Scale Mobile Network Traffic SnapshotsabstractMobile network traffic data offers unprecedented opportunities for innovative studies within and beyond networking. However, progress is hindered by the very limited access that the research community at large has to the real-world mobile network data that is needed to develop and dependably test mobile traffic data-driven solutions. As a contribution to overcome this barrier, we propose CartaGenie, a generator of realistic mobile traffic snapshots at city scale. Taking a deep generative modeling approach and through a tailored conditional generator design, CartaGenie can synthesize high-fidelity and artifact-free spatial traffic snapshots using only contextual information about the target geographical region that is easily found in public repositories. Hence, CartaGenie allows researchers to create their own realistic datasets of spatial traffic from open data about their region of interest. Experiments with real-world mobile traffic measurements collected in multiple metropolitan areas show that CartaGenie can produce dependable network traffic loads for areas where no prior traffic information is available, significantly outperforming a comprehensive set of benchmarks. Moreover, tests with practical case studies demonstrate that the synthetic data generated by CartaGenie is as good as real data in supporting diverse research-oriented mobile traffic data-driven applications. Kai Xu 0014, Rajkarn Singh, Hakan Bilen, Marco Fiore 0001, Mahesh K. Marina, Yue Wang 0008 |
PerCom | 3 |
| 2022 | Learning Foreground-Background Segmentation from Improved Layered GANsabstractDeep learning approaches heavily rely on high-quality human supervision which is nonetheless expensive, time-consuming, and error-prone, especially for image segmentation task. In this paper, we propose a method to automatically synthesize paired photo-realistic images and segmentation masks for the use of training a foreground-background segmentation network. In particular, we learn a generative adversarial network that decomposes an image into foreground and background layers, and avoid trivial decompositions by maximizing mutual information between generated images and latent variables. The improved layered GANs can synthesize higher quality datasets from which segmentation networks of higher performance can be learned. Moreover, the segmentation networks are employed to stabilize the training of layered GANs in return, which are further alternately trained with Layered GANs. Experiments on a variety of single-object datasets show that our method achieves competitive generation quality and segmentation performance compared to related methods. Yu Yang 0011, Hakan Bilen, Qiran Zou, Wing Yin Cheung, Xiangyang Ji |
WACV | 2 |
| 2021 | SpectraGAN: spectrum based generation of city scale spatiotemporal mobile network traffic dataabstractCity-scale spatiotemporal mobile network traffic data can support numerous applications in and beyond networking. However, operators are very reluctant to share their data, which is curbing innovation and research reproducibility. To remedy this status quo, we propose SpectraGAN, a novel deep generative model that, upon training with real-world network traffic measurements, can produce high-fidelity synthetic mobile traffic data for new, arbitrary sized geographical regions over long periods. To this end, the model only requires publicly available context information about the target region, such as population census data. SpectraGAN is an original conditional GAN design with the defining feature of generating spectra of mobile traffic at all locations of the target region based on their contextual features. Evaluations with mobile traffic measurement datasets collected by different operators in 13 cities across two European countries demonstrate that SpectraGAN can synthesize more dependable traffic than a range of representative baselines from the literature. We also show that synthetic data generated with SpectraGAN yield similar results to that with real data when used in applications like radio access network infrastructure power savings and resource allocation, or dynamic population mapping. Kai Xu 0014, Rajkarn Singh, Marco Fiore 0001, Mahesh K. Marina, Hakan Bilen, Howard Benn, Cezary Ziemlicki |
CoNEXT | 5 |
| 2021 | Universal Representation Learning from Multiple Domains for Few-shot ClassificationabstractIn this paper, we look at the problem of few-shot image classification that aims to learn a classifier for previously unseen classes and domains from few labeled samples. Recent methods use various adaptation strategies for aligning their visual representations to new domains or select the relevant ones from multiple domain-specific feature extractors. In this work, we present URL, which learns a single set of universal visual representations by distilling knowledge of multiple domain-specific networks after co-aligning their features with the help of adapters and centered kernel alignment. We show that the universal representations can be further refined for previously unseen domains by an efficient adaptation step in a similar spirit to distance learning methods. We rigorously evaluate our model in the recent Meta-Dataset benchmark and demonstrate that it significantly outperforms the previous methods while being more efficient. Wei-Hong Li 0001, Xialei Liu, Hakan Bilen |
ICCV | 3 |
| 2021 | ViewNet: Unsupervised Viewpoint Estimation from Conditional GenerationabstractUnderstanding the 3D world without supervision is currently a major challenge in computer vision as the annotations required to supervise deep networks for tasks in this domain are expensive to obtain on a large scale. In this paper, we address the problem of unsupervised viewpoint estimation. We formulate this as a self-supervised learning task, where image reconstruction provides the supervision needed to predict the camera viewpoint. Specifically, we make use of pairs of images of the same object at training time, from unknown viewpoints, to self-supervise training by combining the viewpoint information from one image with the appearance information from the other. We demonstrate that using a perspective spatial transformer allows efficient viewpoint learning, outperforming existing unsupervised approaches on synthetic data, and obtains competitive results on the challenging PASCAL3D+ dataset. Octave Mariotti, Oisin Mac Aodha, Hakan Bilen |
ICCV | 3 |
| 2021 | Dataset Condensation with Gradient Matching
Bo Zhao 0038, Konda Reddy Mopuri, Hakan Bilen |
ICLR | 3 |
| 2021 | Neural Feature Matching in Implicit 3D RepresentationsabstractRecently, neural implicit functions have achieved impressive results for encoding 3D shapes. Conditioning on low-dimensional latent codes generalises a single implicit function to learn shared representation space for a variety of shapes, with the advantage of smooth interpolation. While the benefits from the global latent space do not correspond to explicit points at local level, we propose to track the continuous point trajectory by matching implicit features with the latent code interpolating between shapes, from which we corroborate the hierarchical functionality of the deep implicit functions, where early layers map the latent code to fitting the coarse shape structure, and deeper layers further refine the shape details. Furthermore, the structured representation space of implicit functions enables to apply feature matching for shape deformation, with the benefits to handle topology and semantics inconsistency, such as from an armchair to a chair with no arms, without explicit flow functions or manual annotations. Yunlu Chen, Basura Fernando, Hakan Bilen, Thomas Mensink, Efstratios Gavves |
ICML | 3 |
| 2021 | Transfer-Based Semantic Anomaly DetectionabstractDetecting semantic anomalies is challenging due to the countless ways in which they may appear in real-world data. While enhancing the robustness of networks may be sufficient for modeling simplistic anomalies, there is no good known way of preparing models for all potential and unseen anomalies that can potentially occur, such as the appearance of new object classes. In this paper, we show that a previously overlooked strategy for anomaly detection (AD) is to introduce an explicit inductive bias toward representations transferred over from some large and varied semantic task. We rigorously verify our hypothesis in controlled trials that utilize intervention, and show that it gives rise to surprisingly effective auxiliary objectives that outperform previous AD paradigms. Lucas Deecke, Lukas Ruff, Robert A. Vandermeulen, Hakan Bilen |
ICML | 4 |
| 2021 | Dataset Condensation with Differentiable Siamese AugmentationabstractIn many machine learning problems, large-scale datasets have become the de-facto standard to train state-of-the-art deep networks at the price of heavy computation load. In this paper, we focus on condensing large training sets into significantly smaller synthetic sets which can be used to train deep neural networks from scratch with minimum drop in performance. Inspired from the recent training set synthesis methods, we propose Differentiable Siamese Augmentation that enables effective use of data augmentation to synthesize more informative synthetic images and thus achieves better performance when training networks with augmentations. Experiments on multiple image classification benchmarks demonstrate that the proposed method obtains substantial gains over the state-of-the-art, 7% improvements on CIFAR10 and CIFAR100 datasets. We show with only less than 1% data that our method achieves 99.6%, 94.9%, 88.5%, 71.5% relative performance on MNIST, FashionMNIST, SVHN, CIFAR10 respectively. We also explore the use of our method in continual learning and neural architecture search, and show promising results. Bo Zhao 0038, Hakan Bilen |
ICML | 2 |
| 2021 | Continual Representation Learning for Biometric IdentificationabstractWith the explosion of digital data in recent years, continuously learning new tasks from a stream of data without forgetting previously acquired knowledge has become increasingly important. In this paper, we propose a new continual learning (CL) setting, namely "continual representation learning", which focuses on learning better representation in a continuous way. We also provide two large-scale multi-step benchmarks for biometric identification, where the visual appearance of different classes are highly relevant. In contrast to requiring the model to recognize more learned classes, we aim to learn feature representation that can be better generalized to not only previously unseen images but also unseen classes/identities. For the new setting, we propose a novel approach that performs the knowledge distillation over a large number of identities by applying the neighbourhood selection and consistency relaxation strategies to improve scalability and flexibility of the continual learning model. We demonstrate that existing CL methods can improve the representation in the new setting, and our method achieves better results than the competitors. Bo Zhao 0038, Shixiang Tang, Dapeng Chen, Hakan Bilen, Rui Zhao 0001 |
WACV | 4 |
| 2020 | Self-Supervised Learning of Interpretable Keypoints From Unlabelled VideosabstractWe propose a new method for recognizing the pose of objects from a single image that for learning uses only unlabelled videos and a weak empirical prior on the object poses. Video frames differ primarily in the pose of the objects they contain, so our method distils the pose information by analyzing the differences between frames. The distillation uses a new dual representation of the geometry of objects as a set of 2D keypoints, and as a pictorial representation, i.e. a skeleton image. This has three benefits: (1) it provides a tight 'geometric bottleneck' which disentangles pose from appearance, (2) it can leverage powerful image-to-image translation networks to map between photometry and geometry, and (3) it allows to incorporate empirical pose priors in the learning process. The pose priors are obtained from unpaired data, such as from a different dataset or modality such as mocap, such that no annotated image is ever used in learning the pose recognition network. In standard benchmarks for pose recognition for humans and faces, our method achieves state-of-the-art performance among methods that do not require any labelled images for training. Project page: http://www.robots.ox.ac.uk/~vgg/research/unsupervised_pose/. Tomas Jakab, Ankush Gupta 0001, Hakan Bilen, Andrea Vedaldi |
CVPR | 3 |
| 2020 | Weakly Supervised Gaussian Networks for Action DetectionabstractDetecting temporal extents of human actions in videos is a challenging computer vision problem that requires detailed manual supervision including frame-level labels. This expensive annotation process limits deploying action detectors to a limited number of categories. We propose a novel method, called WSGN, that learns to detect actions from weak supervision, using only video-level labels. WSGN learns to exploit both video-specific and dataset-wide statistics to predict relevance of each frame to an action category. This strategy leads to significant gains in action detection for two standard benchmarks THU-MOS14 and Charades. Our method obtains excellent results compared to state-of-the-art methods that uses similar features and loss functions on THUMOS14 dataset. Similarly, our weakly supervised method is only 0.3% mAP behind a state-of-the-art supervised method on challenging Charades dataset for action localization. Basura Fernando, Cheston Tan, Hakan Bilen |
WACV | 3 |
| 2019 | Unsupervised Learning of Landmarks by Descriptor Vector ExchangeabstractEquivariance to random image transformations is an effective method to learn landmarks of object categories, such as the eyes and the nose in faces, without manual supervision. However, this method does not explicitly guarantee that the learned landmarks are consistent with changes between different instances of the same object, such as different facial identities. In this paper, we develop a new perspective on the equivariance approach by noting that dense landmark detectors can be interpreted as local image descriptors equipped with invariance to intra-category variations. We then propose a direct method to enforce such an invariance in the standard equivariant loss. We do so by exchanging descriptor vectors between images of different object instances prior to matching them geometrically. In this manner, the same vectors must work regardless of the specific object identity considered. We use this approach to learn vectors that can simultaneously be interpreted as local descriptors and dense landmarks, combining the advantages of both. Experiments on standard benchmarks show that this approach can match, and in some cases surpass state-of-the-art performance amongst existing methods that learn landmarks without supervision. Code is available at www.robots.ox.ac.uk/~vgg/research/DVE/. James Thewlis, Samuel Albanie, Hakan Bilen, Andrea Vedaldi |
ICCV | 3 |
| 2019 | Mode Normalization
Lucas Deecke, Iain Murray 0001, Hakan Bilen |
ICLR (Poster) | 3 |
| 2018 | Efficient Parametrization of Multi-Domain Deep Neural NetworksabstractA practical limitation of deep neural networks is their high degree of specialization to a single task and visual domain. Recently, inspired by the successes of transfer learning, several authors have proposed to learn instead universal feature extractors that, used as the first stage of any deep network, work well for several tasks and domains simultaneously. Nevertheless, such universal features are still somewhat inferior to specialized networks. To overcome this limitation, in this paper we propose to consider instead universal parametric families of neural networks, which still contain specialized problem-specific models, but differing only by a small number of parameters. We study different designs for such parametrizations, including series and parallel residual adapters, joint adapter compression, and parameter allocations, and empirically identify the ones that yield the highest compression. We show that, in order to maximize performance, it is necessary to adapt both shallow and deep layers of a deep network, but the required changes are very small. We also show that these universal parametrization are very effective for transfer learning, where they outperform traditional fine-tuning techniques. Sylvestre-Alvise Rebuffi, Hakan Bilen, Andrea Vedaldi |
CVPR | 2 |
| 2018 | Unsupervised Learning of Object Landmarks through Conditional Image GenerationabstractWe propose a method for learning landmark detectors for visual objects (such as the eyes and the nose in a face) without any manual supervision. We cast this as the problem of generating images that combine the appearance of the object as seen in a first example image with the geometry of the object as seen in a second example image, where the two examples differ by a viewpoint change and/or an object deformation. In order to factorize appearance and geometry, we introduce a tight bottleneck in the geometry-extraction process that selects and distils geometry-related features. Compared to standard image generation problems, which often use generative adversarial networks, our generation task is conditioned on both appearance and geometry and thus is significantly less ambiguous, to the point that adopting a simple perceptual loss formulation is sufficient. We demonstrate that our approach can learn object landmarks from synthetic image deformations or videos, all without manual supervision, while outperforming state-of-the-art unsupervised landmark detectors. We further show that our method is applicable to a large variety of datasets - faces, people, 3D objects, and digits - without any modifications. Tomas Jakab, Ankush Gupta 0001, Hakan Bilen, Andrea Vedaldi |
NeurIPS | 3 |
| 2018 | Modelling and unsupervised learning of symmetric deformable object categoriesabstractWe propose a new approach to model and learn, without manual supervision, the symmetries of natural objects, such as faces or flowers, given only images as input. It is well known that objects that have a symmetric structure do not usually result in symmetric images due to articulation and perspective effects. This is often tackled by seeking the intrinsic symmetries of the underlying 3D shape, which is very difficult to do when the latter cannot be recovered reliably from data. We show that, if only raw images are given, it is possible to look instead for symmetries in the space of object deformations. We can then learn symmetries from an unstructured collection of images of the object as an extension of the recently-introduced object frame representation, modified so that object symmetries reduce to the obvious symmetry groups in the normalized space. We also show that our formulation provides an explanation of the ambiguities that arise in recovering the pose of symmetric objects from their shape or images and we provide a way of discounting such ambiguities in learning. James Thewlis, Hakan Bilen, Andrea Vedaldi |
NeurIPS | 2 |
| 2018 | Few-shot Learning of Homogeneous Human Locomotion StylesabstractAbstract Using neural networks for learning motion controllers from motion capture data is becoming popular due to the natural and smooth motions they can produce, the wide range of movements they can learn and their compactness once they are trained. Despite these advantages, these systems require large amounts of motion capture data for each new character or style of motion to be generated, and systems have to undergo lengthy retraining, and often reengineering, to get acceptable results. This can make the use of these systems impractical for animators and designers and solving this issue is an open and rather unexplored problem in computer graphics. In this paper we propose a transfer learning approach for adapting a learned neural network to characters that move in different styles from those on which the original neural network is trained. Given a pretrained character controller in the form of a Phase‐Functioned Neural Network for locomotion, our system can quickly adapt the locomotion to novel styles using only a short motion clip as an example. We introduce a canonical polyadic tensor decomposition to reduce the amount of parameters required for learning from each new style, which both reduces the memory burden at runtime and facilitates learning from smaller quantities of data. We show that our system is suitable for learning stylized motions with few clips of motion data and synthesizing smooth motions in real‐time. Ian Mason, Sebastian Starke, Hakan Bilen, Taku Komura |
Comput. Graph. Forum | 4 |
| 2018 | Action Recognition with Dynamic Image NetworksabstractWe introduce the concept of dynamic image, a novel compact representation of videos useful for video analysis, particularly in combination with convolutional neural networks (CNNs). A dynamic image encodes temporal data such as RGB or optical flow videos by using the concept of 'rank pooling'. The idea is to learn a ranking machine that captures the temporal evolution of the data and to use the parameters of the latter as a representation. We call the resulting representation dynamic image because it summarizes the video dynamics in addition to appearance. This powerful idea allows to convert any video to an image so that existing CNN models pre-trained with still images can be immediately extended to videos. We also present an efficient approximate rank pooling operator that runs two orders of magnitude faster than the standard ones with any loss in ranking performance and can be formulated as a CNN layer. To demonstrate the power of the representation, we introduce a novel four stream CNN architecture which can learn from RGB and optical flow frames as well as from their dynamic image representations. We show that the proposed network achieves state-of-the-art performance, 95.5 and 72.5 percent accuracy, in the UCF101 and HMDB51, respectively. Hakan Bilen, Basura Fernando, Efstratios Gavves, Andrea Vedaldi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2017 | Self-Supervised Video Representation Learning with Odd-One-Out NetworksabstractWe propose a new self-supervised CNN pre-training technique based on a novel auxiliary task called odd-one-out learning. In this task, the machine is asked to identify the unrelated or odd element from a set of otherwise related elements. We apply this technique to self-supervised video representation learning where we sample subsequences from videos and ask the network to learn to predict the odd video subsequence. The odd video subsequence is sampled such that it has wrong temporal order of frames while the even ones have the correct temporal order. Therefore, to generate a odd-one-out question no manual annotation is required. Our learning machine is implemented as multi-stream convolutional neural network, which is learned end-to-end. Using odd-one-out networks, we learn temporal representations for videos that generalizes to other related tasks such as action recognition. On action classification, our method obtains 60.3% on the UCF101 dataset using only UCF101 data for training which is approximately 10% better than current state-of-the-art self-supervised learning methods. Similarly, on HMDB51 dataset we outperform self-supervised state-of-the art methods by 12.7% on action classification task. Basura Fernando, Hakan Bilen, Efstratios Gavves, Stephen Gould |
CVPR | 2 |
| 2017 | Unsupervised Learning of Object Landmarks by Factorized Spatial EmbeddingsabstractLearning automatically the structure of object categories remains an important open problem in computer vision. In this paper, we propose a novel unsupervised approach that can discover and learn landmarks in object categories, thus characterizing their structure. Our approach is based on factorizing image deformations, as induced by a viewpoint change or an object deformation, by learning a deep neural network that detects landmarks consistently with such visual effects. Furthermore, we show that the learned landmarks establish meaningful correspondences between different object instances in a category without having to impose this requirement explicitly. We assess the method qualitatively on a variety of object types, natural and man-made. We also show that our unsupervised landmarks are highly predictive of manually-annotated landmarks in face benchmark datasets, and can be used to regress these with a high degree of accuracy. James Thewlis, Hakan Bilen, Andrea Vedaldi |
ICCV | 2 |
| 2017 | Learning multiple visual domains with residual adaptersabstractThere is a growing interest in learning data representations that work well for many different types of problems and data. In this paper, we look in particular at the task of learning a single visual representation that can be successfully utilized in the analysis of very different types of images, from dog breeds to stop signs and digits. Inspired by recent work on learning networks that predict the parameters of another, we develop a tunable deep network architecture that, by means of adapter residual modules, can be steered on the fly to diverse visual domains. Our method achieves a high degree of parameter sharing while maintaining or even improving the accuracy of domain-specific representations. We also introduce the Visual Decathlon Challenge, a benchmark that evaluates the ability of representations to capture simultaneously ten very different visual domains and measures their ability to recognize well uniformly. Sylvestre-Alvise Rebuffi, Hakan Bilen, Andrea Vedaldi |
NIPS | 2 |
| 2017 | Unsupervised learning of object frames by dense equivariant image labellingabstractOne of the key challenges of visual perception is to extract abstract models of 3D objects and object categories from visual measurements, which are affected by complex nuisance factors such as viewpoint, occlusion, motion, and deformations. Starting from the recent idea of viewpoint factorization, we propose a new approach that, given a large number of images of an object and no other supervision, can extract a dense object-centric coordinate frame. This coordinate frame is invariant to deformations of the images and comes with a dense equivariant labelling neural network that can map image pixels to their corresponding object coordinates. We demonstrate the applicability of this method to simple articulated objects and deformable objects such as human faces, learning embeddings from random synthetic transformations or optical flow correspondences, all without any manual supervision. James Thewlis, Hakan Bilen, Andrea Vedaldi |
NIPS | 2 |
| 2016 | Dynamic Image Networks for Action RecognitionabstractWe introduce the concept of dynamic image, a novel compact representation of videos useful for video analysis especially when convolutional neural networks (CNNs) are used. The dynamic image is based on the rank pooling concept and is obtained through the parameters of a ranking machine that encodes the temporal evolution of the frames of the video. Dynamic images are obtained by directly applying rank pooling on the raw image pixels of a video producing a single RGB image per video. This idea is simple but powerful as it enables the use of existing CNN models directly on video data with fine-tuning. We present an efficient and effective approximate rank pooling operator, speeding it up orders of magnitude compared to rank pooling. Our new approximate rank pooling CNN layer allows us to generalize dynamic images to dynamic feature maps and we demonstrate the power of our new representations on standard benchmarks in action recognition achieving state-of-the-art performance. Hakan Bilen, Basura Fernando, Efstratios Gavves, Andrea Vedaldi, Stephen Gould |
CVPR | 1 |
| 2016 | Weakly Supervised Deep Detection NetworksabstractWeakly supervised learning of object detection is an important problem in image understanding that still does not have a satisfactory solution. In this paper, we address this problem by exploiting the power of deep convolutional neural networks pre-trained on large-scale image-level classification tasks. We propose a weakly supervised deep detection architecture that modifies one such network to operate at the level of image regions, performing simultaneously region selection and classification. Trained as an image classifier, the architecture implicitly learns object detectors that are better than alternative weakly supervised detection systems on the PASCAL VOC data. The model, which is a simple and elegant end-to-end architecture, outperforms standard data augmentation and fine-tuning techniques for the task of image-level classification as well. Hakan Bilen, Andrea Vedaldi |
CVPR | 1 |
| 2016 | Integrated perception with recurrent multi-task neural networksabstractModern discriminative predictors have been shown to match natural intelligences in specific perceptual tasks in image classification, object and part detection, boundary extraction, etc. However, a major advantage that natural intelligences still have is that they work well for all perceptual problems together, solving them efficiently and coherently in an integrated manner. In order to capture some of these advantages in machine perception, we ask two questions: whether deep neural networks can learn universal image representations, useful not only for a single task but for all of them, and how the solutions to the different tasks can be integrated in this framework. We answer by proposing a new architecture, which we call multinet, in which not only deep image features are shared between tasks, but where tasks can interact in a recurrent manner by encoding the results of their analysis in a common shared representation of the data. In this manner, we show that the performance of individual tasks in standard benchmarks can be improved first by sharing features between them and then, more significantly, by integrating their solutions in the common representation. Hakan Bilen, Andrea Vedaldi |
NIPS | 1 |
| 2015 | Weakly supervised object detection with convex clusteringabstractWeakly supervised object detection, is a challenging task, where the training procedure involves learning at the same time both, the model appearance and the object location in each image. The classical approach to solve this problem is to consider the location of the object of interest in each image as a latent variable and minimize the loss generated by such latent variable during learning. However, as learning appearance and localization are two interconnected tasks, the optimization is not convex and the procedure can easily get stuck in a poor local minimum, i.e. the algorithm “misses” the object in some images. In this paper, we help the optimization to get close to the global minimum by enforcing a “soft” similarity between each possible location in the image and a reduced set of “exemplars”, or clusters, learned with a convex formulation in the training images. The help is effective because it comes from a different and smooth source of information that is not directly connected with the main task. Results show that our method improves a strong baseline based on convolutional neural network features by more than 4 points without any additional features or extra computation at testing time but only adding a small increment of the training time due to the convex clustering. Hakan Bilen, Marco Pedersoli, Tinne Tuytelaars |
CVPR | 1 |
| 2014 | Weakly Supervised Detection with Posterior Regularization
Hakan Bilen, Marco Pedersoli, Tinne Tuytelaars |
BMVC | 1 |
| 2014 | Object Classification with Adaptable RegionsabstractIn classification of objects substantial work has gone into improving the low level representation of an image by considering various aspects such as different features, a number of feature pooling and coding techniques and considering different kernels. Unlike these works, in this paper, we propose to enhance the semantic representation of an image. We aim to learn the most important visual components of an image and how they interact in order to classify the objects correctly. To achieve our objective, we propose a new latent SVM model for category level object classification. Starting from image-level annotations, we jointly learn the object class and its context in terms of spatial location (where) and appearance (what). Furthermore, to regularize the complexity of the model we learn the spatial and co-occurrence relations between adjacent regions, such that unlikely configurations are penalized. Experimental results demonstrate that the proposed method can consistently enhance results on the challenging Pascal VOC dataset in terms of classification and weakly supervised detection. We also show how semantic representation can be exploited for finding similar content. Hakan Bilen, Marco Pedersoli, Vinay P. Namboodiri, Tinne Tuytelaars, Luc Van Gool |
CVPR | 1 |
| 2014 | Object and Action Classification with Latent Window Parameters
Hakan Bilen, Vinay P. Namboodiri, Luc Van Gool |
Int. J. Comput. Vis. | 1 |
| 2012 | Developing robust vision modules for microsystems applications
Hakan Bilen, Muhammet A. Hocaoglu, Mustafa Unel, Asif Sabanovic |
Mach. Vis. Appl. | 1 |
| 2011 | Object and Action Classification with Latent VariablesabstractIn this paper we propose a generic framework to incorporate unobserved auxiliary information for classifying objects and actions. This framework allows us to explicitly account for localisation and alignment of representations for generic object and action classes as latent variables. We approach this problem in the discriminative setting as learning a max-margin classifier that infers the class label along with the latent variables. Through this paper we make the following contributions a) We provide a method for incorporating latent variables into object and action classification b) We specifically account for the presence of an explicit class related subregion which can include foreground and/or background. c) We explore a way to learn a better classifier by iterative expansion of the latent parameter space. We demonstrate the performance of our approach by rigorous experimental evaluation on a number of standard object and action recognition datasets. © 2011. The copyright of this document resides with its authors. Hakan Bilen, Vinay P. Namboodiri, Luc Van Gool |
BMVC | 1 |
| 2011 | Action recognition: A region based approachabstractWe address the problem of recognizing actions in reallife videos. Space-time interest point-based approaches have been widely prevalent towards solving this problem. In contrast, more spatially extended features such as regions have not been so popular. The reason is, any local region based approach requires the motion flow information for a specific region to be collated temporally. This is challenging as the local regions are deformable and not well delineated from the surroundings. In this paper we address this issue by using robust tracking of regions and we show that it is possible to obtain region descriptors for classification of actions. This paper lays the groundwork for further investigation into region based approaches. Through this paper we make the following contributions a) We advocate identification of salient regions based on motion segmentation b) We adopt a state-of-the art tracker for robust tracking of the identified regions rather than using isolated space-time blocks c) We propose optical flow based region descriptors to encode the extracted trajectories in piece-wise blocks. We demonstrate the performance of our system on real-world data sets. Hakan Bilen, Vinay P. Namboodiri, Luc Van Gool |
WACV | 1 |
| 2009 | Novel parameter estimation schemes in microsystemsabstractThis paper presents two novel estimation methods that are designed to enhance our ability of observing, positioning, and physically transforming the objects and/or biological structures in micromanipulation tasks. In order to effectively monitor and position the microobjects, an online calibration method with submicron precision via a recursive least square solution is presented. To provide the adequate information to manipulate the biological structures without damaging the cell or tissue during an injection, a nonlinear spring-mass-damper model is introduced and mechanical properties of a zebrafish embryo are obtained. These two methods are validated on a microassembly workstation and the results are evaluated quantitatively. Hakan Bilen, Muhammet A. Hocaoglu, Eray A. Baran, Mustafa Unel, Devrim Gozuacik |
ICRA | 1 |
| 2008 | Micromanipulation Using a Microassembly Workstation with Vision and Force Sensing
Hakan Bilen, Mustafa Unel |
ICIC (1) | 1 |
| 2007 | A comparative study of conventional visual servoing schemes in microsystem applicationsabstractThis paper presents an experimental comparison of conventional (calibrated and uncalibrated) image based visual servoing methods in various microsystem applications. Both visual servoing techniques were tested on a microassembly workstation, and their regulation and tracking performances are evaluated. Calibrated visual servoing demands the optical system calibration for the image Jacobian estimation and if a precise optical system calibration is done, it ensures a better accuracy, precision and settling time compared with the uncalibrated approach. On the other hand, in the uncalibrated approach, optical system calibration is not required and since the Jacobian is estimated dynamically, it is more flexible. Hakan Bilen, Muhammet A. Hocaoglu, Erol Ozgur, Mustafa Unel, Asif Sabanovic |
IROS | 1 |