EDBT 2026 Demo / reviewers in the wild / expert
Raffay Hamid
dblp:38/51
· DBLP profile ↗
30ranked-venue papers
9as first author
12since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 7 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 6 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | M-LLM Based Video Frame Selection for Efficient Video UnderstandingabstractRecent advances in Multi-Modal Large Language Models (M-LLMs) show promising results in video reasoning. Popular Multi-Modal Large Language Model (M-LLM) frameworks usually apply naive uniform sampling to reduce the number of video frames that are fed into an M-LLM, particularly for long context videos. However, it could lose crucial context in certain periods of a video, so that the downstream M-LLM may not have sufficient visual information to answer a question. To attack this pain point, we propose a light-weight M-LLM-based frame selection method that adaptively select frames that are more relevant to users’ queries. In order to train the proposed frame selector, we introduce two supervision signals (i) Spatial signal, where single frame importance score by prompting an M-LLM; (ii) Temporal signal, in which multiple frames selection by prompting Large Language Model (LLM) using the captions of all frame candidates. The selected frames are then digested by a frozen downstream video M-LLM for visual reasoning and question answering. Empirical results show that the proposed M-LLM video frame selector improves the performances various downstream video Large Language Model (video-LLM) across medium (ActivityNet, NExT-QA) and long (EgoSchema, LongVideoBench) context video question answering benchmarks. Kai Hu 0010, Xiaohan Nie, Son Tran, Tal Neiman, Lingyun Wang 0005, Mubarak Shah, Raffay Hamid, Trishul Chilimbi |
CVPR | 9 |
| 2025 | CoLLM: A Large Language Model for Composed Image RetrievalabstractComposed Image Retrieval (CIR) is a complex task that aims to retrieve images based on a multimodal query. Typical training data consists of triplets containing a reference image, a textual description of desired modifications, and the target image, which are expensive and time-consuming to acquire. The scarcity of CIR datasets has led to zero-shot approaches utilizing synthetic triplets or leveraging vision-language models (VLMs) with ubiquitous web-crawled image-caption pairs. However, these methods have significant limitations: synthetic triplets suffer from limited scale, lack of diversity, and unnatural modification text, while image-caption pairs hinder joint embedding learning of the multimodal query due to the absence of triplet data. Moreover, existing approaches struggle with complex and nuanced modification texts that demand sophisticated fusion and understanding of vision and language modalities. We present CoLLM, a one-stop framework that effectively addresses these limitations. Our approach generates triplets on-the-fly from image-caption pairs, enabling supervised training without manual annotation. We leverage Large Language Models (LLMs) to generate joint embeddings of reference images and modification texts, facilitating deeper multimodal fusion. Additionally, we introduce Multi-Text CIR (MTCIR), a large-scale dataset comprising 3.4M samples, and refine existing CIR benchmarks (CIRR and Fashion-IQ) to enhance evaluation reliability. Experimental results demonstrate that CoLLM achieves state-of-the-art performance across multiple CIR benchmarks and settings. MTCIR yields competitive results, with up to 15% performance improvement. Our refined benchmarks provide more reliable evaluation metrics for CIR models, contributing to the advancement of this important field. Project page is at collm-cvpr25.github.io. Chuong Huynh, Ashish Tawari, Mubarak Shah, Son Tran, Raffay Hamid, Trishul Chilimbi, Abhinav Shrivastava |
CVPR | 6 |
| 2023 | Movies2Scenes: Using Movie Metadata to Learn Scene RepresentationabstractUnderstanding scenes in movies is crucial for a variety of applications such as video moderation, search, and recommendation. However, labeling individual scenes is a time-consuming process. In contrast, movie level metadata (e.g., genre, synopsis, etc.) regularly gets produced as part of the film production process, and is therefore significantly more commonly available. In this work, we propose a novel contrastive learning approach that uses movie metadata to learn a general-purpose scene representation. Specifically, we use movie metadata to define a measure of movie similarity, and use it during contrastive learning to limit our search for positive scene-pairs to only the movies that are considered similar to each other. Our learned scene representation consistently outperforms existing state-of-the-art methods on a diverse set of tasks evaluated using multiple benchmark datasets. Notably, our learned representation offers an average improvement of 7.9% on the seven classification tasks and 9.7% improvement on the two regression tasks in LVU dataset. Furthermore, using a newly collected movie dataset, we present comparative results of our scene representation on a set of video moderation tasks to demonstrate its generalizability on previously less explored tasks. Shixing Chen, Chun-Hao Liu, Xiaohan Nie, Maxim Arap, Raffay Hamid |
CVPR | 6 |
| 2023 | LEMaRT: Label-Efficient Masked Region Transform for Image HarmonizationabstractWe present a simple yet effective self-supervised pre-training method for image harmonization which can leverage large-scale unannotated image datasets. To achieve this goal, we first generate pre-training data online with our Label-Efficient Masked Region Transform (LEMaRT) pipeline. Given an image, LEMaRT generates a foreground mask and then applies a set of transformations to perturb various visual attributes, e.g., defocus blur, contrast, saturation, of the region specified by the generated mask. We then pre-train image harmonization models by recovering the original image from the perturbed image. Secondly, we introduce an image harmonization model, namely SwinIH, by retrofitting the Swin Transformer [27] with a combination of local and global self-attention mechanisms. Pre-training SwinIH with LEMaRT results in a new state of the art for image harmonization, while being label-efficient, i.e., consuming less annotated data for fine-tuning than existing methods. Notably, on iHarmony4 dataset [8], SwinIH outperforms the state of the art, i.e., SCS-Co [16] by a margin of 0.4 dB when it is fine-tuned on only 50% of the training data, and by 1.0 dB when it is trained on the full training dataset. Cong Phuoc Huynh, Maxim Arap, Raffay Hamid |
CVPR | 5 |
| 2023 | Selective Structured State-Spaces for Long-Form Video UnderstandingabstractEffective modeling of complex spatiotemporal dependencies in long-form videos remains an open problem. The recently proposed Structured State-Space Sequence ($S4$) model with its linear complexity offers a promising direction in this space. However, we demonstrate that treating all imagetokens equally as done by$S4$model can adversely affect its efficiency and accuracy. To address this limitation, we present a novel Selective$S4$(i.e.,$S5)$model that employs a lightweight mask generator to adaptively select informative image tokens resulting in more efficient and accurate modeling of long-term spatiotemporal dependencies in videos. Unlike previous mask-based token reduction methods used in transformers, our$S5$model avoids the dense self-attention calculation by making use of the guidance of the momentum-updated$S4$model. This enables our model to efficiently discard less informative tokens and adapt to various long-form video understanding tasks more effectively. However, as is the case for most token reduction methods, the informative image tokens could be dropped incorrectly. To improve the robustness and the temporal horizon of our model, we propose a novel long-short masked contrastive learning (LSMCL) approach that enables our model to predict longer temporal context using shorter input videos. We present extensive comparative results using three challenging long-form video understanding datasets (LVU, COIN and Breakfast), demonstrating that our approach consistently outperforms the previous state-of-the-art S4 model by up to 9.6% accuracy while reducing its memory footprint by 23%. Pichao Wang, Linda Liu, Mohamed Omar, Raffay Hamid |
CVPR | 7 |
| 2022 | Robust Cross-Modal Representation Learning with Progressive Self-DistillationabstractThe learning objective of vision-language approach of CLIP [63] does not effectively account for the noisy many-to-many correspondences found in web-harvested image captioning datasets, which contributes to its compute and data inefficiency. To address this challenge, we introduce a novel training framework based on cross-modal contrastive learning that uses progressive self-distillation and soft image-text alignments to more efficiently learn robust representations from noisy data. Our model distills its own knowledge to dynamically generate soft-alignment targets for a subset of images and captions in every minibatch, which are then used to update its parameters. Extensive evaluation across 14 benchmark datasets shows that our method consistently outperforms its CLIP counterpart in multiple settings, including: (a) zero-shot classification, (b) linear probe transfer, and (c) image-text retrieval, without incurring extra computational cost. Analysis using an ImageNet-based robustness test-bed [70] reveals that our method offers better effective robustness to natural distribution shifts compared to both ImageNet-trained models and CLIP itself. Lastly, pretraining with datasets spanning two orders of magnitude in size shows that our improvements over CLIP tend to scale with number of training examples. Alex Andonian, Shixing Chen, Raffay Hamid |
CVPR | 3 |
| 2022 | Depth-Guided Sparse Structure-from-Motion for Movies and TV ShowsabstractExisting approaches for Structure from Motion (SfM) produce impressive 3-D reconstruction results especially when using imagery captured with large parallax. However, to create engaging video-content in movies and TV shows, the amount by which a camera can be moved while filming a particular shot is often limited. The resulting small-motion parallax between video frames makes standard geometry-based SfM approaches not as effective for movies and TV shows. To address this challenge, we propose a simple yet effective approach that uses single-frame depth-prior obtained from a pretrained network to significantly improve geometry-based SfM for our small-parallax setting. To this end, we first use the depth-estimates of the detected keypoints to reconstruct the point cloud and camera-pose for initial two-view reconstruction. We then perform depth-regularized optimization to register new images and triangulate the new points during incremental reconstruction. To comprehensively evaluate our approach, we introduce a new dataset (StudioSfM) consisting of 130 shots with 21K frames from 15 studio-produced videos that are manually annotated by a professional CG studio. We demonstrate that our approach: (a) significantly improves the quality of 3-D reconstruction for our small-parallax setting, (b) does not cause any degradation for data with large-parallax, and (c) maintains the generalizability and scalability of geometry-based sparse SfM. Our dataset can be obtained at https://github.com/amazon-researchlsmall-baseline-camera-tracking. Sheng Liu 0017, Xiaohan Nie, Raffay Hamid |
CVPR | 3 |
| 2022 | CNN-based Audio Event Recognition for Automated Violence Classification and Rating for Prime Video Content
Kenny Qiu, Raffay Hamid |
INTERSPEECH | 5 |
| 2022 | A comprehensive empirical review of modern voice activity detection approaches for movies and TV shows
Sandeep Joshi, Tamojit Chatterjee, Raffay Hamid |
Neurocomputing | 4 |
| 2021 | Shot Contrastive Self-Supervised Learning for Scene Boundary DetectionabstractScenes play a crucial role in breaking the storyline of movies and TV episodes into semantically cohesive parts. However, given their complex temporal structure, finding scene boundaries can be a challenging task requiring large amounts of labeled training data. To address this challenge, we present a self-supervised shot contrastive learning approach (ShotCoL) to learn a shot representation that maximizes the similarity between nearby shots compared to randomly selected shots. We show how to apply our learned shot representation for the task of scene boundary detection to offer state-of-the-art performance on the MovieNet [33] dataset while requiring only ~25% of the training labels, using 9× fewer model parameters and offering 7× faster runtime. To assess the effectiveness of ShotCoL on novel applications of scene boundary detection, we take on the problem of finding timestamps in movies and TV episodes where video-ads can be inserted while offering a minimally disruptive viewing experience. To this end, we collected a new dataset called AdCuepoints with 3, 975 movies and TV episodes, 2.2 million shots and 19, 119 minimally disruptive ad cue-point labels. We present a thorough empirical analysis on this dataset demonstrating the effectiveness of ShotCoL for ad cue-points detection. Shixing Chen, Xiaohan Nie, David Fan 0001, Dongqing Zhang, Vimal Bhat, Raffay Hamid |
CVPR | 6 |
| 2021 | Intro and Recap Detection for Movies and TV SeriesabstractModern video streaming service companies offer millions of video-titles for its customers. A lot of these titles have repetitive introductory and recap parts in the beginning that customers have to manually skip in order to achieve an uninterrupted viewing experience. To avoid this unnecessary friction, some of the services have recently added "skip-intro" and "skip-recap" buttons to their video players before the intro and recap parts start. To efficiently scale this experience to their entire catalogs, it is important to automate the process of finding the intro and recap portions of titles. In this work, we pose intro and recap detection as a supervised sequence labeling problem and propose a novel end-to-end deep learning framework to this end. Specifically, we use CNNs to extract both visual and audio features from videos, and fuse these features using a B-LSTM in order to capture the various long and short term dependencies among different frame-features over time. Finally, we use a CRF to jointly optimize the sequence labeling for the intro and recap parts of the titles. We present a thorough empirical analysis of our model compared to several other deep learning based architectures and demonstrate the superior performance of our approach. Kripa Chettiar, Ben Cheung, Vernon Germano, Raffay Hamid |
WACV | 5 |
| 2021 | A Robust and Efficient Framework for Sports-Field RegistrationabstractWe propose a novel framework to register sports-fields as they appear in broadcast sports videos. Unlike previous approaches, we particularly address the challenge of field- registration when: (a) there are not enough distinguishable features on the field, and (b) no prior knowledge is available about the camera. To this end, we detect a grid of key- points distributed uniformly on the entire field instead of using only sparse local corners and line intersections, thereby extending the keypoint coverage to the texture-less parts of the field as well. To further improve keypoint based homography estimate, we differentialbly warp and align it with a set of dense field-features defined as normalized distance- map of pixels to their nearest lines and key-regions. We predict the keypoints and dense field-features simultaneously using a multi-task deep network to achieve computational efficiency. To have a comprehensive evaluation, we have compiled a new dataset called SportsFields which is collected from 192 video-clips from 5 different sports covering large environmental and camera variations. We empirically demonstrate that our algorithm not only achieves state of the art field-registration accuracy but also runs in real-time for HD resolution videos using commodity hardware. Xiaohan Nie, Shixing Chen, Raffay Hamid |
WACV | 3 |
| 2016 | Geospatial Correspondences for Multimodal RegistrationabstractThe growing availability of very high resolution (<;1 m/pixel) satellite and aerial images has opened up unprecedented opportunities to monitor and analyze the evolution of land-cover and land-use across the world. To do so, images of the same geographical areas acquired at different times and, potentially, with different sensors must be efficiently parsed to update maps and detect land-cover changes. However, a naϊve transfer of ground truth labels from one location in the source image to the corresponding location in the target image is generally not feasible, as these images are often only loosely registered (with up to ± 50m of non-uniform errors). Furthermore, land-cover changes in an area over time must be taken into account for an accurate ground truth transfer. To tackle these challenges, we propose a mid-level sensor-invariant representation that encodes image regions in terms of the spatial distribution of their spectral neighbors. We incorporate this representation in a Markov Random Field to simultaneously account for nonlinear mis-registrations and enforce locality priors to find matches between multi-sensor images. We show how our approach can be used to assist in several multimodal land-cover update and change detection problems. Diego Marcos, Raffay Hamid, Devis Tuia |
CVPR | 2 |
| 2016 | Toward a Generalizable Image Representation for Large-Scale Change Detection: Application to Generic Damage AnalysisabstractEach year, multiple catastrophic events impact vulnerable populations around the planet. Assessing the damage caused by these events in a timely and accurate manner is crucial for efficient execution of relief efforts to help the victims of these calamities. Given the low accessibility of the damaged areas, high-resolution optical satellite imagery has emerged as a valuable source of information to quickly asses the extent of damage by manually analyzing the pre- and postevent imagery of the region. To make this analysis more efficient, multiple learning techniques using a variety of image representations have been proposed. However, most of these representations are prone to variabilities in capture angle, sun location, and seasonal variations. To evaluate these representations in the context of damage detection, we present a benchmark of 86 pre- and postevent image pairs with respective reference data derived from United Nation Operational Satellite Applications Programme (UNOSAT) assessment maps, spanning a total area of 4665 km2from 11 different locations around the world. The technical contribution of our work is a novel image representation based on shape distributions of image patches encoded with locality-constrained linear coding. We empirically demonstrate that our proposed representation provides an improvement of at least 5%, in equal error rate, over alternate approaches. Finally, we present a thorough robustness analysis of the considered representational schemes, with respect to capture-angle variabilities and multiple sensor combinations. Lionel Gueguen, Raffay Hamid |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2015 | Large-scale damage detection using satellite imageryabstractSatellite imagery is a valuable source of information for assessing damages in distressed areas undergoing a calamity, such as an earthquake or an armed conflict. However, the sheer amount of data required to be inspected for this assessment makes it impractical to do it manually. To address this problem, we present a semi-supervised learning framework for large-scale damage detection in satellite imagery. We present a comparative evaluation of our framework using over 88 million images collected from 4, 665 KM2from 12 different locations around the world. To enable accurate and efficient damage detection, we introduce a novel use of hierarchical shape features in the bags-of-visual words setting. We analyze how practical factors such as sun, sensor-resolution, satellite-angle, and registration differences impact the effectiveness our proposed representation, and compare it to five alternative features in multiple learning settings. Finally, we demonstrate through a user-study that our semi-supervised framework results in a ten-fold reduction in human annotation time at a minimal loss in detection accuracy compared to manual inspection. Lionel Gueguen, Raffay Hamid |
CVPR | 2 |
| 2015 | Hardware compliant approximate image codesabstractIn recent years, several feature encoding schemes for the bags-of-visual-words model have been proposed. While most of these schemes produce impressive results, they all share an important limitation: their high computational complexity makes it challenging to use them for large-scale problems. In this work, we propose an approximate locality-constrained encoding scheme that offers significantly better computational efficiency (~ 40×) than its exact counterpart, with comparable classification accuracy. Using the perturbation analysis of least-squares problems, we present a formal approximation error analysis of our approach, which helps distill the intuition behind the robustness of our method. We present a thorough set of empirical analyses on multiple standard data-sets, to assess the capability of our encoding scheme for its representational as well as discriminative accuracy. Da Kuang, Alex Gittens, Raffay Hamid |
CVPR | 3 |
| 2015 | Fast Approximate Matching of Videos from Hand-Held Cameras for Robust Background SubtractionabstractWe identify a novel instance of the background subtraction problem that focuses on extracting near-field foreground objects captured using handheld cameras. Given two user-generated videos of a scene, one with and the other without the foreground object (s), our goal is to efficiently generate an output video with only the foreground object (s) present in it. We cast this challenge as a spatio-temporal frame matching problem, and propose an efficient solution for it that exploits the temporal smoothness of the video sequences. We present theoretical analyses for the error bounds of our approach, and validate our findings using a detailed set of simulation experiments. Finally, we present the results of our approach tested on multiple real videos captured using handheld cameras, and compare them to several alternate foreground extraction approaches. Raffay Hamid, Atish Das Sarma, Dennis DeCoste, Neel Sundaresan |
WACV | 1 |
| 2014 | Compact Random Feature MapsabstractKernel approximation using randomized feature maps has recently gained a lot of interest. In this work, we identify that previous approaches for polynomial kernel approximation create maps that are rank deficient, and therefore do not utilize the capacity of the projected feature space effectively. To address this challenge, we propose compact random feature maps (CRAFTMaps) to approximate polynomial kernels more concisely and accurately. We prove the error bounds of CRAFTMaps demonstrating their superior kernel reconstruction performance compared to the previous approximation schemes. We show how structured random matrices can be used to efficiently generate CRAFTMaps, and present a single-pass algorithm using CRAFTMaps to learn non-linear multi-class classifiers. We present experiments on multiple standard data-sets with performance competitive with state-of-the-art results. Raffay Hamid, Ying Xiao 0003, Alex Gittens, Dennis DeCoste |
ICML | 1 |
| 2014 | What makes an image popular?abstractHundreds of thousands of photographs are uploaded to the internet every minute through various social networking and photo sharing platforms. While some images get millions of views, others are completely ignored. Even from the same users, different photographs receive different number of views. This begs the question: What makes a photograph popular? Can we predict the number of views a photograph will receive even before it is uploaded? These are some of the questions we address in this work. We investigate two key components of an image that affect its popularity, namely the image content and social context. Using a dataset of about 2.3 million images from Flickr, we demonstrate that we can reliably predict the normalized view count of images with a rank correlation of 0.81 using both image content and social cues. In this paper, we show the importance of image cues such as color, gradients, deep learning features and the set of objects present, as well as the importance of various social cues such as number of friends or number of photos uploaded that lead to high or low popularity of images. Aditya Khosla, Atish Das Sarma, Raffay Hamid |
WWW | 3 |
| 2014 | A visualization framework for team sports captured using multiple static cameras
Raffay Hamid, Ramkrishan K. Kumar, Jessica K. Hodgins, Irfan A. Essa |
Comput. Vis. Image Underst. | 1 |
| 2013 | Dense Non-rigid Point-Matching Using Random ProjectionsabstractWe present a robust and efficient technique for matching dense sets of points undergoing non-rigid spatial transformations. Our main intuition is that the subset of points that can be matched with high confidence should be used to guide the matching procedure for the rest. We propose a novel algorithm that incorporates these high-confidence matches as a spatial prior to learn a discriminative subspace that simultaneously encodes both the feature similarity as well as their spatial arrangement. Conventional subspace learning usually requires spectral decomposition of the pair-wise distance matrix across the point-sets, which can become inefficient even for moderately sized problems. To this end, we propose the use of random projections for approximate subspace learning, which can provide significant time improvements at the cost of minimal precision loss. This efficiency gain allows us to iteratively find and remove high-confidence matches from the point sets, resulting in high recall. To show the effectiveness of our approach, we present a systematic set of experiments and results for the problem of dense non-rigid image-feature matching. Raffay Hamid, Dennis DeCoste, Chih-Jen Lin |
CVPR | 1 |
| 2013 | Large-Scale Video Summarization Using Web-Image PriorsabstractGiven the enormous growth in user-generated videos, it is becoming increasingly important to be able to navigate them efficiently. As these videos are generally of poor quality, summarization methods designed for well-produced videos do not generalize to them. To address this challenge, we propose to use web-images as a prior to facilitate summarization of user-generated videos. Our main intuition is that people tend to take pictures of objects to capture them in a maximally informative way. Such images could therefore be used as prior information to summarize videos containing a similar set of objects. In this work, we apply our novel insight to develop a summarization algorithm that uses the web-image based prior information in an unsupervised manner. Moreover, to automatically evaluate summarization algorithms on a large scale, we propose a framework that relies on multiple summaries obtained through crowdsourcing. We demonstrate the effectiveness of our evaluation framework by comparing its performance to that of multiple human evaluators. Finally, we present results for our framework tested on hundreds of user-generated videos. Aditya Khosla, Raffay Hamid, Chih-Jen Lin, Neel Sundaresan |
CVPR | 2 |
| 2013 | Palette power: enabling visual search through colorsabstractWith the explosion of mobile devices with cameras, online search has moved beyond text to other modalities like images, voice, and writing. For many applications like Fashion, image-based search offers a compelling interface as compared to text forms by better capturing the visual attributes. In this paper we present a simple and fast search algorithm that uses color as the main feature for building visual search. We show that low level cues such as color can be used to quantify image similarity and also to discriminate among products with different visual appearances. We demonstrate the effectiveness of our approach through a mobile shopping application\footnote{eBay Fashion App available at https://itunes.apple.com/us/app/ebay-fashion/id378358380?mt=8 and eBay image swatch is the feature indexing millions of real world fashion images}. Our approach outperforms several other state-of-the-art image retrieval algorithms for large scale image data. Anurag Bhardwaj, Atish Das Sarma, Wei Di, Raffay Hamid, Robinson Piramuthu, Neel Sundaresan |
KDD | 4 |
| 2010 | Player localization using multiple static cameras for sports visualizationabstractWe present a novel approach for robust localization of multiple people observed using multiple cameras. We use this location information to generate sports visualizations, which include displaying a virtual offside line in soccer games, and showing players' positions and motion patterns. Our main contribution is the modeling and analysis for the problem of fusing corresponding players' positional information as finding minimum weight K-length cycles in complete K-partite graphs. To this end, we use a dynamic programming based approach that varies over a continuum of being maximally to minimally greedy in terms of the number of paths explored at each iteration. We present an end-to-end sports visualization framework that employs our proposed algorithm-class. We demonstrate the robustness of our framework by testing it on 60,000 frames of soccer footage captured over 5 different illumination conditions, play types, and team attire. Raffay Hamid, Ramkrishan K. Kumar, Matthias Grundmann 0002, Irfan A. Essa, Jessica K. Hodgins |
CVPR | 1 |
| 2009 | A novel sequence representation for unsupervised analysis of human activities
Raffay Hamid, Siddhartha Maddi, Amos Y. Johnson, Aaron F. Bobick, Irfan A. Essa, Charles L. Isbell Jr. |
Artif. Intell. | 1 |
| 2008 | Taylor expansion based classifier adaptation: Application to person detectionabstractBecause of the large variation across different environments, a generic classifier trained on extensive data-sets may perform sub-optimally in a particular test environment. In this paper, we present a general framework for classifier adaptation, which improves an existing generic classifier in the new test environment. Viewing classifier learning as a cost minimization problem, we perform classifier adaptation by combining the cost function on the old data-sets with the cost function on the data-set collected from the new environment. The former term is further approximated with its second order Taylor expansion to reduce the amount of information that needs to be saved for adaptation. Unlike traditional approaches that are often designed for a specific application or classifier, our scheme is applicable to various types of classifiers and user labels. We demonstrate this property on two popular classifiers (logistic regression and boosting), while using two types of user labels (direct labels and similarity labels). Extensive experiments conducted for the task of person detection in conference-room environments show that significant performance improvement can be achieved with our proposed method. Cha Zhang, Raffay Hamid, Zhengyou Zhang |
CVPR | 2 |
| 2007 | Structure from Statistics - Unsupervised Activity Analysis using Suffix TreesabstractModels of activity structure for unconstrained environments are generally not available a priori. Recent representational approaches to this end are limited by their computational complexity, and ability to capture activity structure only up to some fixed temporal scale. In this work, we propose Suffix Trees as an activity representation to efficiently extract structure of activities by analyzing their constituent event-subsequences over multiple temporal scales. We empirically compare Suffix Trees with some of the previous approaches in terms of feature cardinality, discriminative prowess, noise sensitivity and activity-class discovery. Finally, exploiting properties of Suffix Trees, we present a novel perspective on anomalous subsequences of activities, and propose an algorithm to detect them in linear-time. We present comparative results over experimental data, collected from a kitchen environment to demonstrate the competence of our proposed framework. Raffay Hamid, Siddhartha Maddi, Aaron F. Bobick, Irfan A. Essa |
ICCV | 1 |
| 2005 | Detection and Explanation of Anomalous Activities: Representing Activities as Bags of Event n-GramsabstractWe present a novel representation and method for detecting and explaining anomalous activities in a video stream. Drawing from natural language processing, we introduce a representation of activities as bags of event n-grams, where we analyze the global structural information of activities using their local event statistics. We demonstrate how maximal cliques in an undirected edge-weighted graph of activities, can be used in an unsupervised manner, to discover regular sub-classes of an activity class. Based on these discovered sub-classes, we formulate a definition of anomalous activities and present a way to detect them. Finally, we characterize each discovered sub-class in terms of its "most representative member" and present an information-theoretic method to explain the detected anomalies in a human-interpretable form. Raffay Hamid, Amos Y. Johnson, Samir Batta, Aaron F. Bobick, Charles L. Isbell Jr., Graham Coleman |
CVPR (1) | 1 |
| 2004 | a CAPpella: programming by demonstration of context-aware applicationsabstractContext-aware applications are applications that implicitly take their context of use into account by adapting to changes in a user's activities and environments. No one has more intimate knowledge about these activities and environments than end-users themselves. Currently there is no support for end-users to build context-aware applications for these dynamic settings. To address this issue, we present a CAPpella, a programming by demonstration Context-Aware Prototyping environment intended for end-users. Users "program" their desired context-aware behavior (situation and associated action) in situ, without writing any code, by demonstrating it to a CAPpella and by annotating the relevant portions of the demonstration. Using a meeting and medicine-taking scenario, we illustrate how a user can demonstrate different behaviors to a CAPpella. We describe a CAPpella's underlying system to explain how it supports users in building behaviors and present a study of 14 end-users to illustrate its feasibility and usability. Anind K. Dey, Raffay Hamid, Chris Beckmann, Ian Li, Daniel Hsu 0001 |
CHI | 2 |
| 2004 | Audio-visual flow -a variational approach to multi-modal flow estimationabstractJust as a motion field is associated to a moving object, an audio field can be associated to an object that can behave as a sound source. The flow field of such a sound source which moves over time would not only have an optical component, but also an audio component; something we call audio-visual flow. In this paper we present a common structure tensor based variational framework for dense audiovisual flow-field estimation. The proposed scheme improves the rank of the local structure tensor by incorporating an audio information channel which is substantially uncorrelated from the complementing visual information channel. The scheme allows ascribing weights to individual sensor modalities based on the confidence in their corresponding measurements. Results are presented to demonstrate how combining multiple modalities in our proposed framework can provide a possible solution to temporary full visual occlusions. Raffay Hamid, Aaron F. Bobick, Anthony J. Yezzi |
ICIP | 1 |