EDBT 2026 Demo / reviewers in the wild / expert
Ziad Al-Halah
dblp:147/2698
· DBLP profile ↗
32ranked-venue papers
9as first author
17since 2021 · last 2026
0000-0001-6887-0385ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 7 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 8 first-author · 11 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sample-efficient Audio-Visual Learning of Scene AcousticsabstractAbstract An environment acoustic model represents how sound is transformed by the physical characteristics of an indoor environment, for any given source/receiver location. Whereas traditional methods for constructing such models assume dense geometry and/or sound measurements throughout the environment, we explore how to infer room impulse responses (RIRs) based on a sparse set of images and echoes observed in the space, as well as how to choose where to collect these audio-visual observations. Towards that goal, we first introduce a transformer-based method that uses self-attention to build a rich acoustic context, then infers the RIRs of arbitrary query source-receiver locations through cross-attention. Then, motivated by real-world physical constraints in collecting these observations, we further introduce active acoustic sampling , a new task in which a mobile agent jointly constructs the environment acoustic model and spatial occupancy map on-the-fly from sparse audio-visual observations. We train a reinforcement learning (RL) policy that guides agent navigation toward optimal acoustic data sampling positions, rewarding information gain for the full environment model. Evaluating on diverse unseen 3D indoor environments, our method outperforms the state-of-the-art and—in a major departure from traditional methods—generalizes to novel environments in a few-shot manner. Furthermore, when augmented with our active sampling policy, it successfully guides an embodied agent to acoustically informative positions given real-world exploration constraints, outperforming both traditional navigation agents and prior acoustic rendering methods. Project: http://vision.cs.utexas.edu/projects/fewShot-RIR . Arjun Somayazulu, Sagnik Majumder, Changan Chen, Ziad Al-Halah, Kristen Grauman |
Int. J. Comput. Vis. | 4 |
| 2025 | Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional VideosabstractGiven a multi-view video, which viewpoint is most informative for a human observer? Existing methods rely on heuristics or expensive "best-view" supervision to answer this question, limiting their applicability. We propose a weakly supervised approach that leverages language accompanying an instructional multi-view video as a means to recover its most informative viewpoint(s). Our key hypothesis is that the more accurately an individual view can predict a view-agnostic text summary, the more informative it is. To put this into action, we propose LANGVIEW, a framework that uses the relative accuracy of view-dependent caption predictions as a proxy for best view pseudo-labels. Then, those pseudo-labels are used to train a view selector, together with an auxiliary camera pose predictor that enhances view-sensitivity. During inference, our model takes as input only a multi-view video—no language or camera poses—and returns the best viewpoint to watch at each timestep. On two challenging datasets comprised of diverse multi-camera setups and how-to activities, our model consistently outperforms state-of-the-art baselines, both with quantitative metrics and human evaluation. Project: https://vision.cs.utexas.edu/projects/which-view-shows-it-best. Sagnik Majumder, Tushar Nagarajan, Ziad Al-Halah, Reina Pradhan, Kristen Grauman |
CVPR | 3 |
| 2025 | Switch-a-View: View Selection Learned from Unlabeled In-the-Wild Videos
Sagnik Majumder, Tushar Nagarajan, Ziad Al-Halah, Kristen Grauman |
ICCV | 3 |
| 2025 | How Would it Sound? Material-Controlled Multimodal Acoustic Profile Generation for Indoor Scenes
Mahnoor Fatima Saad, Ziad Al-Halah |
ICCV | 2 |
| 2024 | Learning Spatial Features from Audio-Visual Correspondence in Egocentric VideosabstractWe propose a self-supervised method for learning repre- sentations based on spatial audio-visual correspondences in egocentric videos. Our method uses a masked auto- encoding framework to synthesize masked binaural audio through the synergy of audio and vision, thereby learning useful spatial relationships between the two modalities We use our pretrained features to tackle two downstream video tasks requiring spatial understanding in social scenar- ios: active speaker detection and spatial audio denoising. Through extensive experiments, we show that our features are generic enough to improve over multiple state-of-the- art baselines on both tasks on two challenging egocentric video datasets that offer binaural audio, EgoCom and Easy-Com. Project: http://vision.cs.utexas.edu/projects/ego_av_corr. Sagnik Majumder, Ziad Al-Halah, Kristen Grauman |
CVPR | 2 |
| 2023 | NaQ: Leveraging Narrations as Queries to Supervise Episodic MemoryabstractSearching long egocentric videos with natural language queries (NLQ) has compelling applications in augmented reality and robotics, where a fluid index into everything that a person (agent) has seen before could augment human memory and surface relevant information on demand. However, the structured nature of the learning problem (free-form text query inputs, localized video temporal window outputs) and its needle-in-a-haystack nature makes it both technically challenging and expensive to supervise. We introduce Narrations-as-Queries (NaQ), a data augmentation strategy that transforms standard video-text narrations into training data for a video query localization model. Validating our idea on the Ego4D benchmark, we find it has tremendous impact in practice. NaQ improves multiple top models by substantial margins (even doubling their accuracy), and yields the very best results to date on the Ego4D NLQ challenge, soundly outperforming all challenge winners in the CVPR and ECCV 2022 competitions and topping the current public leaderboard. Beyond achieving the state-of-the-art for NLQ, we also demonstrate unique properties of our approach such as the ability to perform zero-shot and fewshot NLQ, and improved performance on queries about long-tail object categories. Code and models: http://vision.cs.utexas.edu/projects/naq. Santhosh K. Ramakrishnan, Ziad Al-Halah, Kristen Grauman |
CVPR | 2 |
| 2023 | SpotEM: Efficient Video Search for Episodic MemoryabstractThe goal in episodic memory (EM) is to search a long egocentric video to answer a natural language query (e.g., “where did I leave my purse?”). Existing EM methods exhaustively extract expensive fixed-length clip features to look everywhere in the video for the answer, which is infeasible for long wearable-camera videos that span hours or even days. We propose SpotEM, an approach to achieve efficiency for a given EM method while maintaining good accuracy. SpotEM consists of three key ideas: 1) a novel clip selector that learns to identify promising video regions to search conditioned on the language query; 2) a set of low-cost semantic indexing features that capture the context of rooms, objects, and interactions that suggest where to look; and 3) distillation losses that address the optimization issues arising from end-to-end joint training of the clip selector and EM model. Our experiments on 200+ hours of video from the Ego4D EM Natural Language Queries benchmark and three different EM models demonstrate the effectiveness of our approach: computing only 10% – 25% of the clip features, we preserve 84% – 97% of the original EM model’s accuracy. Project page: https://vision.cs.utexas.edu/projects/spotem Santhosh K. Ramakrishnan, Ziad Al-Halah, Kristen Grauman |
ICML | 2 |
| 2023 | A domain-agnostic approach for characterization of lifelong learning systems
Megan M. Baker, Alexander New, Mario Aguilar-Simon, Ziad Al-Halah, Sébastien M. R. Arnold, Eseoghene Benjamin, Andrew P. Brna, Ethan Brooks, Ryan C. Brown, Zachary A. Daniels, Anurag Reddy Daram, Fabien Delattre, Ryan Dellana, Eric Eaton, Haotian Fu, Kristen Grauman, Jesse Hostetler, Shariq Iqbal, Cassandra Kent, Nicholas Ketz, Soheil Kolouri, George Dimitri Konidaris, Dhireesha Kudithipudi, Erik G. Learned-Miller, Michael L. Littman, Sandeep Madireddy, Jorge A. Mendez, Eric Q. Nguyen, Christine D. Piatko, Praveen K. Pilly, Aswin Raghavan, Abrar Rahman, Santhosh K. Ramakrishnan, Neale Ratzlaff, Andrea Soltoggio, Peter Stone 0001, Indranil Sur, Zhipeng Tang, Saket Tiwari, Kyle Vedder, Felix Wang, Zifan Xu, Angel Yanguas-Gil, Harel Yedidsion, Shangqun Yu, Gautam K. Vallabha |
Neural Networks | 4 |
| 2022 | Zero Experience Required: Plug & Play Modular Transfer Learning for Semantic Visual NavigationabstractIn reinforcement learning for visual navigation, it is common to develop a model for each new task, and train that model from scratch with task-specific interactions in 3D environments. However, this process is expensive; mas-sive amounts of interactions are needed for the model to generalize well. Moreover, this process is repeated when-ever there is a change in the task type or the goal modality. We present a unified approach to visual navigation using a novel modular transfer learning model. Our model can ef-fectively leverage its experience from one source task and apply it to multiple target tasks (e.g., ObjectNav, Room-Nav, Vi ewNav) with various goal modalities (e.g., image, sketch, audio, label). Furthermore, our model enables zero-shot experience learning, whereby it can solve the target tasks without receiving any task-specific interactive training. Our experiments on multiple photorealistic datasets and challenging tasks show that our approach learns faster, generalizes better, and outperforms SoTA models by a sig-nificant margin. Project page: https://vision.cs.utexas.edu/projects/zsel/ Ziad Al-Halah, Santhosh K. Ramakrishnan, Kristen Grauman |
CVPR | 1 |
| 2022 | PONI: Potential Functions for ObjectGoal Navigation with Interaction-free LearningabstractState-of-the-art approaches to ObjectGoal navigation (ObjectNav) rely on reinforcement learning and typically require significant computational resources and time for learning. We propose Potential functions for ObjectGoal Navigation with Interaction-free learning (PONI), a modular approach that disentangles the skills of 'where to look?’ for an object and 'how to navigate to$(x,\ y)$?’. Our key insight is that 'where to look?’ can be treated purely as a perception problem, and learned without environment interactions. To address this, we propose a network that predicts two complementary potential functions conditioned on a semantic map and uses them to decide where to look for an unseen object. We train the potential function network using supervised learning on a passive dataset of top-down semantic maps, and integrate it into a modular framework to perform ObjectNav. Experiments on Gibson and Matterport3D demonstrate that our method achieves the stateof-the-art for ObjectNav while incurring up to$1,600\times less$computational cost for training. Code and pre-trained models are available.11Website: https://vision.cs.utexas.edu/projects/poni/ Santhosh K. Ramakrishnan, Devendra Singh Chaplot, Ziad Al-Halah, Jitendra Malik, Kristen Grauman |
CVPR | 3 |
| 2022 | Environment Predictive Coding for Visual Navigation
Santhosh K. Ramakrishnan, Tushar Nagarajan, Ziad Al-Halah, Kristen Grauman |
ICLR | 3 |
| 2022 | Few-Shot Audio-Visual Learning of Environment AcousticsabstractRoom impulse response (RIR) functions capture how the surrounding physical environment transforms the sounds heard by a listener, with implications for various applications in AR, VR, and robotics. Whereas traditional methods to estimate RIRs assume dense geometry and/or sound measurements throughout the environment, we explore how to infer RIRs based on a sparse set of images and echoes observed in the space. Towards that goal, we introduce a transformer-based method that uses self-attention to build a rich acoustic context, then predicts RIRs of arbitrary query source-receiver locations through cross-attention. Additionally, we design a novel training objective that improves the match in the acoustic signature between the RIR predictions and the targets. In experiments using a state-of-the-art audio-visual simulator for 3D environments, we demonstrate that our method successfully generates arbitrary RIRs, outperforming state-of-the-art methods and---in a major departure from traditional methods---generalizing to novel environments in a few-shot manner. Project: http://vision.cs.utexas.edu/projects/fs_rir Sagnik Majumder, Changan Chen, Ziad Al-Halah, Kristen Grauman |
NeurIPS | 3 |
| 2021 | Semantic Audio-Visual NavigationabstractRecent work on audio-visual navigation assumes a constantly-sounding target and restricts the role of audio to signaling the target’s position. We introduce semantic audio-visual navigation, where objects in the environment make sounds consistent with their semantic meaning (e.g., toilet flushing, door creaking) and acoustic events are sporadic or short in duration. We propose a transformer-based model to tackle this new semantic AudioGoal task, incorporating an inferred goal descriptor that captures both spatial and semantic properties of the target. Our model’s persistent multimodal memory enables it to reach the goal even long after the acoustic event stops. In support of the new task, we also expand the SoundSpaces audio simulations to provide semantically grounded sounds for an array of objects in Matterport3D. Our method strongly outperforms existing audio-visual navigation methods by learning to associate semantic, acoustic, and visual cues.1 Changan Chen, Ziad Al-Halah, Kristen Grauman |
CVPR | 2 |
| 2021 | Fashion IQ: A New Dataset Towards Retrieving Images by Natural Language FeedbackabstractConversational interfaces for the detail-oriented retail fashion domain are more natural, expressive, and user friendly than classical keyword-based search interfaces. In this paper, we introduce the Fashion IQ dataset to support and advance research on interactive fashion image retrieval. Fashion IQ is the first fashion dataset to provide human-generated captions that distinguish similar pairs of garment images together with side-information consisting of real-world product descriptions and derived visual attribute labels for these images. We provide a detailed analysis of the characteristics of the Fashion IQ data, and present a transformer-based user simulator and interactive image retriever that can seamlessly integrate visual attributes with image features, user feedback, and dialog history, leading to improved performance over the state of the art in dialog-based image retrieval. We believe that our dataset will encourage further work on developing more natural and real-world applicable conversational shopping assistants.1 Hui Wu 0009, Yupeng Gao, Ziad Al-Halah, Steven Rennie, Kristen Grauman, Rogério Feris |
CVPR | 4 |
| 2021 | Move2Hear: Active Audio-Visual Source SeparationabstractWe introduce the active audio-visual source separation problem, where an agent must move intelligently in order to better isolate the sounds coming from an object of interest in its environment. The agent hears multiple audio sources simultaneously (e.g., a person speaking down the hall in a noisy household) and it must use its eyes and ears to automatically separate out the sounds originating from a target object within a limited time budget. Towards this goal, we introduce a reinforcement learning approach that trains movement policies controlling the agent’s camera and microphone placement over time, guided by the improvement in predicted audio separation quality. We demonstrate our approach in scenarios motivated by both augmented reality (system is already co-located with the target object) and mobile robotics (agent begins arbitrarily far from the target object). Using state-of-the-art realistic audio-visual simulations in 3D environments, we demonstrate our model’s ability to find minimal movement sequences with maximal payoff for audio source separation. Project: http://vision.cs.utexas.edu/projects/move2hear. Sagnik Majumder, Ziad Al-Halah, Kristen Grauman |
ICCV | 2 |
| 2021 | Learning to Set Waypoints for Audio-Visual Navigation
Changan Chen, Sagnik Majumder, Ziad Al-Halah, Ruohan Gao, Santhosh K. Ramakrishnan, Kristen Grauman |
ICLR | 3 |
| 2021 | Modeling Fashion Influence From PhotosabstractThe evolution of clothing styles and their migration across the world is intriguing, yet difficult to describe quantitatively. We propose to discover and quantify fashion influences from catalog and social media photos. We explore fashion influence along two channels: geolocation and fashion brands. We introduce an approach that detects which of these entities influence which other entities in terms of propagating their styles. We then leverage the discovered influence patterns to inform a novel forecasting model that predicts the future popularity of any given style within any given city or brand. To demonstrate our idea, we leverage public large-scale datasets of 7.7M Instagram photos from 44 major world cities (where styles are worn with variable frequency) as well as 41K Amazon product photos (where styles are purchased with variable frequency). Our model learns directly from the image data how styles move between locations and how certain brands affect each other's designs in a predictable way. The discovered influence relationships reveal how both cities and brands exert and receive fashion influence for an array of visual styles inferred from the images. Furthermore, the proposed forecasting model achieves state-of-the-art results for challenging style forecasting tasks. Our results indicate the advantage of grounding visual style evolution both spatially and temporally, and for the first time, they quantify the propagation of inter-brand and inter-city influences. Project page:https://www.cs.utexas.edu/~ziad/influence_from_photos.html Ziad Al-Halah, Kristen Grauman |
IEEE Trans. Multim. | 1 |
| 2020 | From Paris to Berlin: Discovering Fashion Style Influences Around the WorldabstractThe evolution of clothing styles and their migration across the world is intriguing, yet difficult to describe quantitatively. We propose to discover and quantify fashion influences from everyday images of people wearing clothes. We introduce an approach that detects which cities influence which other cities in terms of propagating their styles. We then leverage the discovered influence patterns to inform a forecasting model that predicts the popularity of any given style at any given city into the future. Demonstrating our idea with GeoStyle-a large-scale dataset of 7.7M images covering 44 major world cities, we present the discovered influence relationships, revealing how cities exert and receive fashion influence for an array of 50 observed visual styles. Furthermore, the proposed forecasting model achieves state-of-the-art results for a challenging style forecasting task, showing the advantage of grounding visual style evolution both spatially and temporally. Ziad Al-Halah, Kristen Grauman |
CVPR | 1 |
| 2020 | SoundSpaces: Audio-Visual Navigation in 3D Environments
Changan Chen, Unnat Jain, Carl Schissler, Sebastià Vicenc Amengual Garí, Ziad Al-Halah, Vamsi K. Ithapu, Philip W. Robinson, Kristen Grauman |
ECCV (6) | 5 |
| 2020 | VisualEchoes: Spatial Image Representation Learning Through Echolocation
Ruohan Gao, Changan Chen, Ziad Al-Halah, Carl Schissler, Kristen Grauman |
ECCV (9) | 3 |
| 2020 | Occupancy Anticipation for Efficient Exploration and Navigation
Santhosh K. Ramakrishnan, Ziad Al-Halah, Kristen Grauman |
ECCV (5) | 2 |
| 2019 | SPaSe - Multi-Label Page Segmentation for Presentation SlidesabstractWe introduce the first benchmark dataset for slide-page segmentation. Presentation slides are one of the most prominent document types used to exchange ideas across the web, educational institutes and businesses. This document format is marked with a complex layout which contains a rich variety of graphical (e.g. diagram, logo), textual (e.g. heading, affiliation) and structural components (e.g. enumeration, legend). This vast and popular knowledge source is still unattainable by modern machine learning technique due to lack of annotated data. To tackle this issue, we introduce SPaSe (Slide Page Segmentation), a novel dataset containing in total 2000 slides with dense, pixel-wise annotations of 25 classes. We show that slide segmentation reveals some interesting properties that characterize this task. Unlike the common image segmentation problem, disjoint classes tend to have a high overlap of regions, thus posing this segmentation task as a multi-label problem. Furthermore, many of the frequently encountered classes in slides are location sensitive (e.g. title, footnote). Hence, we believe our dataset represents a challenging and interesting benchmark for novel segmentation models. Finally, we evaluate state-of-the-art deep segmentation models on our dataset and show that it is suitable for developing deep learning models without any need of pre-training. Our dataset will be released to the public to foster further research on this interesting task. Monica-Laura Haurilet, Ziad Al-Halah, Rainer Stiefelhagen |
WACV | 2 |
| 2018 | Informed Democracy: Voting-based Novelty Detection for Action Recognition
Alina Roitberg, Ziad Al-Halah, Rainer Stiefelhagen |
BMVC | 2 |
| 2017 | Automatic Discovery, Association Estimation and Learning of Semantic Attributes for a Thousand CategoriesabstractAttribute-based recognition models, due to their impressive performance and their ability to generalize well on novel categories, have been widely adopted for many computer vision applications. However, usually both the attribute vocabulary and the class-attribute associations have to be provided manually by domain experts or large number of annotators. This is very costly and not necessarily optimal regarding recognition performance, and most importantly, it limits the applicability of attribute-based models to large scale data sets. To tackle this problem, we propose an end-to-end unsupervised attribute learning approach. We utilize online text corpora to automatically discover a salient and discriminative vocabulary that correlates well with the human concept of semantic attributes. Moreover, we propose a deep convolutional model to optimize class-attribute associations with a linguistic prior that accounts for noise and missing data in text. In a thorough evaluation on ImageNet, we demonstrate that our model is able to efficiently discover and learn semantic attributes at a large scale. Furthermore, we demonstrate that our model outperforms the state-of-the-art in zero-shot learning on three data sets: ImageNet, Animals with Attributes and aPascal/aYahoo. Finally, we enable attribute-based learning on ImageNet and will share the attributes and associations for future research. Ziad Al-Halah, Rainer Stiefelhagen |
CVPR | 1 |
| 2017 | Fashion Forward: Forecasting Visual Style in FashionabstractWhat is the future of fashion? Tackling this question from a data-driven vision perspective, we propose to forecast visual style trends before they occur. We introduce the first approach to predict the future popularity of styles discovered from fashion images in an unsupervised manner. Using these styles as a basis, we train a forecasting model to represent their trends over time. The resulting model can hypothesize new mixtures of styles that will become popular in the future, discover style dynamics (trendy vs. classic), and name the key visual attributes that will dominate tomorrow's fashion. We demonstrate our idea applied to three datasets encapsulating 80,000 fashion products sold across six years on Amazon. Results indicate that fashion forecasting benefits greatly from visual analysis, much more than textual or meta-data cues surrounding products. Ziad Al-Halah, Rainer Stiefelhagen, Kristen Grauman |
ICCV | 1 |
| 2016 | Recovering the Missing Link: Predicting Class-Attribute Associations for Unsupervised Zero-Shot LearningabstractCollecting training images for all visual categories is not only expensive but also impractical. Zero-shot learning (ZSL), especially using attributes, offers a pragmatic solution to this problem. However, at test time most attribute-based methods require a full description of attribute associations for each unseen class. Providing these associations is time consuming and often requires domain specific knowledge. In this work, we aim to carry out attribute-based zero-shot classification in an unsupervised manner. We propose an approach to learn relations that couples class embeddings with their corresponding attributes. Given only the name of an unseen class, the learned relationship model is used to automatically predict the class-attribute associations. Furthermore, our model facilitates transferring attributes across data sets without additional effort. Integrating knowledge from multiple sources results in a significant additional improvement in performance. We evaluate on two public data sets: Animals with Attributes and aPascal/aYahoo. Our approach outperforms state-of the-art methods in both predicting class-attribute associations and unsupervised ZSL by a large margin. Ziad Al-Halah, Makarand Tapaswi, Rainer Stiefelhagen |
CVPR | 1 |
| 2016 | Naming TV characters by watching and analyzing dialogsabstractPerson identification in TV series has been a popular research topic over the last decade. In this area, most approaches either use manually annotated data or extract character supervision from a combination of subtitles and transcripts. However, both approaches have key drawbacks that hinder application of these methods at a large scale - manual annotation is expensive and transcripts are often hard to obtain. We investigate the topic of automatically labeling all character appearances in TV series using information obtained solely from subtitles. This task is extremely difficult as the dialogs between characters provide very sparse and weakly supervised data. We address these challenges by exploiting recent advances in face descriptors and Multiple Instance Learning methods. We propose methods to create MIL bags and evaluate and discuss several MIL techniques. The best combination achieves an average precision over 80% on three diverse TV series. We demonstrate that only using subtitles provides good results on identifying characters in TV series and wish to encourage the community towards this problem. Monica-Laura Haurilet, Makarand Tapaswi, Ziad Al-Halah, Rainer Stiefelhagen |
WACV | 3 |
| 2016 | Transfer metric learning for action similarity using high-level semantics
Ziad Al-Halah, Lukas Rybok, Rainer Stiefelhagen |
Pattern Recognit. Lett. | 1 |
| 2015 | Accio: A Data Set for Face Track Retrieval in Movies Across AgeabstractVideo face recognition is a very popular task and has come a long way. The primary challenges such as illumination, resolution and pose are well studied through multiple data sets. However there are no video-based data sets dedicated to study the effects of aging on facial appearance. We present a challenging face track data set, Harry Potter Movies Aging Data set (Accio1), to study and develop age invariant face recognition methods for videos. Our data set not only has strong challenges of pose, illumination and distractors, but also spans a period of ten years providing substantial variation in facial appearance. We propose two primary tasks: within and across movie face track retrieval; and two protocols which differ in their freedom to use external data. We present baseline results for the retrieval performance using a state-of-the-art face track descriptor. Our experiments show clear trends of reduction in performance as the age gap between the query and database increases. We will make the data set publicly available for further exploration in age-invariant video face recognition. Esam Ghaleb, Makarand Tapaswi, Ziad Al-Halah, Hazim Kemal Ekenel, Rainer Stiefelhagen |
ICMR | 3 |
| 2015 | How to Transfer? Zero-Shot Object Recognition via Hierarchical Transfer of Semantic AttributesabstractAttribute based knowledge transfer has proven very successful in visual object analysis and learning previously unseen classes. However, the common approach learns and transfers attributes without taking into consideration the embedded structure between the categories in the source set. Such information provides important cues on the intraattribute variations. We propose to capture these variations in a hierarchical model that expands the knowledge source with additional abstraction levels of attributes. We also provide a novel transfer approach that can choose the appropriate attributes to be shared with an unseen class. We evaluate our approach on three public datasets: a Pascal, Animals with Attributes and CUB-200-2011 Birds. The experiments demonstrate the effectiveness of our model with significant improvement over state-of-the-art. Ziad Al-Halah, Rainer Stiefelhagen |
WACV | 1 |
| 2014 | What to Transfer? High-Level Semantics in Transfer Metric Learning for Action SimilarityabstractLearning from few examples is considered a very challenging task where transfer learning proved to be beneficial. Such a learning framework exploits previous experiences and knowledge to compensate for the lack of training data in a novel domain. Knowledge representation plays a vital role in the type and performance of transfer learning approaches, as well as its robustness against negative transfer effect. This aspect is usually not considered in most of the proposed transfer learning methodologies, where the focus is either on the transfer type or on the representation. In this work, we study the use of various high-level semantics in transfer metric learning. We propose a generic transfer metric learning framework, and analyze the effect of different semantic similarity spaces on transfer type and efficiency against negative transfer. Furthermore, we introduce a hierarchical knowledge representation model based on the embedded structure in the attribute semantic space. The evaluation of the framework on challenging transfer settings in the context of action similarity demonstrates the effectiveness of our approach. Ziad Al-Halah, Lukas Rybok, Rainer Stiefelhagen |
ICPR | 1 |
| 2014 | "Important stuff, everywhere!" Activity recognition with salient proto-objects as contextabstractObject information is an important cue to discriminate between activities that draw part of their meaning from context. Most of current work either ignores this information or relies on specific object detectors. However, such object detectors require a significant amount of training data and complicate the transfer of the action recognition framework to novel domains with different objects and object-action relationships. Motivated by recent advances in saliency detection, we propose to employ salient proto-objects for unsupervised discovery of object- and object-part candidates and use them as a contextual cue for activity recognition. Our experimental evaluation on three publicly available data sets shows that the integration of proto-objects and simple motion features substantially improves recognition performance, outperforming the state-of-the-art. Lukas Rybok, Boris Schauerte, Ziad Al-Halah, Rainer Stiefelhagen |
WACV | 3 |