EDBT 2026 Demo / reviewers in the wild / expert
Anthony Hoogs
dblp:15/5671
· DBLP profile ↗
62ranked-venue papers
9as first author
17since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 49 · 8 first-author · 11 since 2021Artificial intelligence and machine learning · 35 · 6 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Aligning Machiavellian Agents: Behavior Steering via Test-Time Policy ShapingabstractThe deployment of decision-making AI agents presents a critical challenge in maintaining alignment with human values or guidelines while operating in complex, dynamic environments. Agents trained solely to achieve their objectives may adopt harmful behavior, exposing a key trade-off between maximizing the reward function and maintaining alignment. For pre-trained agents, ensuring alignment is particularly challenging, as retraining can be a costly and slow process. This is further complicated by the diverse and potentially conflicting attributes representing the ethical values for alignment. To address these challenges, we propose a test-time alignment technique based on model-guided policy shaping. Our method allows precise control over individual behavioral attributes, generalizes across diverse reinforcement learning (RL) environments, and facilitates a principled trade-off between ethical alignment and reward maximization without requiring agent retraining. We evaluate our approach using the MACHIAVELLI benchmark, which comprises 134 text-based game environments and thousands of annotated scenarios involving ethical decisions. The RL agents are first trained to maximize the reward in their respective games. At test time, we apply policy shaping via scenario-action attribute classifiers to ensure decision alignment with ethical attributes. We compare our approach against prior training-time methods and general-purpose agents, as well as study several types of ethical violations and power-seeking behavior. Our results demonstrate that test-time policy shaping provides an effective and scalable solution for mitigating unethical behavior across diverse environments and alignment attributes. Dena Mujtaba, Brian Hu 0001, Anthony Hoogs, Arslan Basharat |
AAAI | 3 |
| 2025 | A Semantically Impactful Image Manipulation Dataset: Characterizing Image Manipulations Using Semantic SignificanceabstractWe investigate how to characterize semantic significance (SS) in detecting image manipulations (IMD) for media forensics. We introduce the Characterization of Seman-tic Impact for IMD (CSI-IMD) dataset, which focuses on localizing and evaluating the semantic impact of image manipulations to counter advanced generative techniques. Our evaluation of 10 state-of-the-art IMD and localization methods on CSI-IMD reveals key insights. Unlike existing datasets, CSI-IMD provides detailed semantic annotations beyond traditional manipulation masks, aiding in the development of new defensive strategies. The dataset features manipulations from advanced generation methods, offering various levels of semantic significance. It is divided into two parts: a gold-standard set of 1,000 manu-ally annotated manipulations with high-quality control, and an extended set of 500,000 automated manipulations for large-scale training and analysis. We also propose a new SS-focused task to assess the impact of semantically targeted manipulations. Our experiments show that current IMD methods struggle with manipulations created using stable diffusion, with TruFor and Cat-Net performing the best among those tested. The CSI-IMD dataset will become available at https://github.com/csiimd/csiimd. Ming-Ching Chang, Matthias Kirchner, Zhenfei Zhang, Xin Li 0005, Arslan Basharat, Anthony Hoogs |
WACV | 7 |
| 2024 | Defending Against Social Engineering Attacks in the Age of LLMsabstractLin Ai, Tharindu Sandaruwan Kumarage, Amrita Bhattacharjee, Zizhou Liu, Zheng Hui, Michael S. Davinroy, James Cook, Laura Cassani, Kirill Trapeznikov, Matthias Kirchner, Arslan Basharat, Anthony Hoogs, Joshua Garland, Huan Liu, Julia Hirschberg. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Lin Ai, Tharindu Kumarage, Amrita Bhattacharjee, Zizhou Liu, Zheng Hui, Michael Davinroy, James Cook, Laura Cassani, Kirill Trapeznikov, Matthias Kirchner, Arslan Basharat, Anthony Hoogs, Joshua Garland, Huan Liu 0001, Julia Hirschberg |
EMNLP | 12 |
| 2024 | Geowatch for Detecting Heavy Construction in Heterogeneous Time Series of Satellite ImagesabstractLearning from multiple sensors is challenging due to spatio-temporal misalignment and differences in resolution and captured spectra. To that end, we introduce GeoWATCH, a flexible framework for training models on long sequences of satellite images sourced from multiple sensor platforms, which is designed to handle image classification, activity recognition, object detection, or object tracking tasks. Our system includes a novel partial weight loading mechanism based on sub-graph isomorphism which allows for continually training and modifying a network over many training cycles. This has allowed us to train a lineage of models over a long period of time, which we have observed has improved performance as we adjust configurations while maintaining a core backbone. Jon Crall, Connor Greenwell, David Joy, Matthew Leotta, Aashish Chaudhary, Anthony Hoogs |
IGARSS | 6 |
| 2024 | Uncovering Bias in Building Damage Assessment from Satellite ImageryabstractWe identify a bias in a commonly used dataset for building damage detection, evaluate its effects on existing deep learning models, and devise mitigation strategies to overcome it. We find that the data contains significantly more groups of damaged buildings than single ones leading to skewed machine learning evaluations. Consequently, deep learning models heavily rely on surrounding context rather than individual building damage when classifying supporting our claim. Specifically, the dataset includes extraneous damage surrounding buildings such as debris, fallen trees, and other damaged buildings which results in deep neural networks overfitting to these features. We analyze the top-5 solutions of the xView2 challenge, which focuses on building damage classification using satellite imagery as provided by the xBD dataset. Our experiments reveal that these models struggle to accurately identify isolated damaged buildings, potentially causing oversights in critical disaster scenarios and delaying humanitarian aid. Finally, we devise a new augmentation strategy to reduce this bias in disaster datasets and show it improves real-world outcomes. Dennis Melamed, Cameron Johnson, Isaac Gerg, Russell Blue, Anthony Hoogs, Brian Clipp, Philip Morrone |
IGARSS | 6 |
| 2024 | FishTrack23: An Ensemble Underwater Dataset for Multi-Object TrackingabstractTracking and classifying fish in optical underwater imagery presents several challenges which are encountered less frequently in terrestrial domains. Video may contain large schools comprised of many individuals, dynamic natural backgrounds, highly variable target scales, volatile collection conditions, and non-fish moving confusers including debris, marine snow, and other organisms. Additionally, there is a lack of large public datasets for algorithm evaluation available in this domain. The contributions of this paper is three fold. First, we present the FishTrack23 dataset which provides a large quantity of expert-annotated fish groundtruth tracks, in imagery and video collected across a range of different backgrounds, locations, collection conditions, and organizations. Approximately 850k bounding boxes across 26k tracks are included in the release of the ensemble, with potential for future growth in later releases. Second, we evaluate improvements upon baseline object detectors, trackers and classifiers on the dataset. Lastly, we integrate these methods into web and desktop interfaces to expedite annotation generation on new datasets. Matthew Dawkins, Jack Prior, Bryon Lewis, Robin Faillettaz, Thompson Banez, Mary Salvi, Audrey K. Rollo, Julien Simon, Matthew D. Campbell, Matthew Lucero, Aashish Chaudhary, Benjamin L. Richards, Anthony Hoogs |
WACV | 13 |
| 2023 | Xaitk-Saliency: An Open Source Explainable AI Toolkit for SaliencyabstractAdvances in artificial intelligence (AI) using techniques such as deep learning have fueled the recent progress in fields such as computer vision. However, these algorithms are still often viewed as "black boxes", which cannot easily explain how they arrived at their final output decisions. Saliency maps are one commonly used form of explainable AI (XAI), which indicate the input features an algorithm paid attention to during its decision process. Here, we introduce the open source xaitk-saliency package, an XAI framework and toolkit for saliency. We demonstrate its modular and flexible nature by highlighting two example use cases for saliency maps: (1) object detection model comparison and (2) doppelganger saliency for person re-identification. We also show how the xaitk-saliency package can be paired with visualization tools to support the interactive exploration of saliency maps. Our results suggest that saliency maps may play a critical role in the verification and validation of AI models, ensuring their trusted use and deployment. The code is publicly available at: https://github.com/xaitk/xaitk-saliency. Brian Hu 0001, Paul Tunison, Brandon Richard Webster, Anthony Hoogs |
AAAI | 4 |
| 2023 | Open Set Action Recognition via Multi-Label Evidential LearningabstractExisting methods for open set action recognition focus on novelty detection that assumes video clips show a single action, which is unrealistic in the real world. We propose a new method for open set action recognition and novelty detection via MUlti-Label Evidential learning (MULE), that goes beyond previous novel action detection methods by addressing the more general problems of single or multiple actors in the same scene, with simultaneous action(s) by any actor. Our Beta Evidential Neural Network estimates multi-action uncertainty with Beta densities based on actor-context-object relation representations. An evidence debiasing constraint is added to the objective function for optimization to reduce the static bias of video representations, which can incorrectly correlate predictions and static cues. We develop a primal-dual average scheme update-based learning algorithm to optimize the proposed problem and provide corresponding theoretical analysis. Besides, uncertainty and belief-based novelty estimation mechanisms are formulated to detect novel actions. Extensive experiments on two real-world video datasets show that our proposed approach achieves promising performance in single/multi-actor, single/multi-action settings. Our code and models are released at https://github.com/charliezhaoyinpeng/mule. Chen Zhao 0010, Dawei Du, Anthony Hoogs, Christopher Funk |
CVPR | 3 |
| 2023 | Novel Object Detection in Remote Sensing ImageryabstractNovel object detection in remote sensing is challenging due to small objects, background clutter and open-set recognition. To discover and identify objects that are not within the set of known classes in training, we developed the extreme value theory-based novel object detection framework. Specifically, we first employed the state-of-the-art object detector to extract object detections. Then, the extreme value model (EVM) is trained based on the features of extracted detections. Thus our method can characterize the distribution of outliers in the known classes-based distributions to classify novel objects. If the novelty score is larger than the pre-set threshold, we assign this sample to novel classes; otherwise the sample is classified by the original object detector. To adapt to our task, we hold out a novel set of 18 of overall 60 classes in the xView dataset in satellite imagery. The experimental results on the xView dataset show the effectiveness of our proposed approach over traditional softmax thresholding. Dawei Du, Christopher Funk, Katarina Doctor, Anthony Hoogs |
IGARSS | 4 |
| 2023 | MEVID: Multi-view Extended Videos with Identities for Video Person Re-IdentificationabstractIn this paper, we present the Multi-view Extended Videos with Identities (MEVID) dataset for large-scale, video person re-identification (ReID) in the wild. To our knowledge, MEVID represents the most-varied video person ReID dataset, spanning an extensive indoor and outdoor environment across nine unique dates in a 73-day window, various camera viewpoints, and entity clothing changes. Specifically, we label the identities of 158 unique people wearing 598 outfits taken from 8, 092 tracklets, average length of about 590 frames, seen in 33 camera views from the very-large-scale MEVA person activities dataset. While other datasets have more unique identities, MEVID emphasizes a richer set of information about each individual, such as: 4 outfits/identity vs. 2 outfits/identity in CCVID, 33 viewpoints across 17 locations vs. 6 in 5 simulated locations for MTA, and 10 million frames vs. 3 million for LS-VID. Being based on the MEVA video dataset, we also inherit data that is intentionally demographically balanced to the continental United States. To accelerate the annotation process, we developed a semi-automatic annotation framework and GUI that combines state-of-the-art real-time models for object detection, pose estimation, person ReID, and multi-object tracking. We evaluate several state-of-the-art methods on MEVID challenge problems and comprehensively quantify their robustness in terms of changes of outfit, scale, and background location. Our quantitative analysis on the realistic, unique aspects of MEVID shows that there are significant remaining challenges in video person ReID and indicates important directions for future research. Daniel Davila, Dawei Du, Bryon Lewis, Christopher Funk, Joseph VanPelt, Roderic Collins, Kellie Corona, Matt S. Brown, Scott McCloskey, Anthony Hoogs, Brian Clipp |
WACV | 10 |
| 2023 | Reconstructing Humpty Dumpty: Multi-feature Graph Autoencoder for Open Set Action RecognitionabstractMost action recognition datasets and algorithms assume a closed world, where all test samples are instances of the known classes. In open set problems, test samples may be drawn from either known or unknown classes. Existing open set action recognition methods are typically based on extending closed set methods by adding post hoc analysis of classification scores or feature distances and do not capture the relations among all the video clip elements. Our approach uses the reconstruction error to determine the novelty of the video since unknown classes are harder to put back together and thus have a higher reconstruction error than videos from known classes. We refer to our solution to the open set action recognition problem as "Humpty Dumpty", due to its reconstruction abilities. Humpty Dumpty is a novel graph-based autoencoder that accounts for contextual and semantic relations among the clip pieces for improved reconstruction. A larger reconstruction error leads to an increased likelihood that the action can not be reconstructed, i.e., can not put Humpty Dumpty back together again, indicating that the action has never been seen before and is novel/unknown. Extensive experiments are performed on two publicly available action recognition datasets including HMDB-51 and UCF-101, showing the state-of-the-art performance for open set action recognition. Dawei Du, Ameya Shringi, Anthony Hoogs, Christopher Funk |
WACV | 3 |
| 2022 | Cascade Transformers for End-to-End Person SearchabstractThe goal of person search is to localize a target person from a gallery set of scene images, which is extremely challenging due to large scale variations, pose/viewpoint changes, and occlusions. In this paper, we propose the Cascade Occluded Attention Transformer (COAT) for end-to-end person search. Our three-stage cascade design focuses on detecting people in the first stage, while later stages simultaneously and progressively refine the representation for person detection and re-identification. At each stage the occluded attention transformer applies tighter intersection over union thresholds, forcing the network to learn coarse-to-fine pose/scale invariant features. Meanwhile, we calculate each detection's occluded attention to differentiate a person's tokens from other people or the background. In this way, we simulate the effect of other objects occluding a person of interest at the token-level. Through comprehensive experiments, we demonstrate the benefits of our method by achieving state-of-the-art performance on two benchmark datasets. Rui Yu 0002, Dawei Du, Rodney LaLonde, Daniel Davila, Christopher Funk, Anthony Hoogs, Brian Clipp |
CVPR | 6 |
| 2022 | Discover and Mitigate Unknown Biases with Debiasing Alternate Networks
Zhiheng Li 0002, Anthony Hoogs, Chenliang Xu |
ECCV (13) | 2 |
| 2022 | Novelty Detection in Remote Sensing ImageryabstractObject detection and classification in remote sensing imagery have been studied for decades, and has had a resurgence recently with significant improvements from deep learning. Most approaches follow the standard target recognition paradigm by assuming a fixed set of known object classes. The detector/classifier is trained on these, and attempts to disregard everything else. However, the real-world is complicated and unpredictable; often, there are new, interesting objects that are similar to known classes, but sufficiently different such that the system will (correctly) ignore them. The goal of novelty detection is to detect instances of new object types rather than misclassifying them as known types or background, while continuing to correctly classify instances of known object types. The primary challenge in novelty detection is determining how different a new image should be in order to be novel, vs. a new condition or variant of a known class. To address this, our method performs novelty detection in imagery using extreme value theory (EVT) operating in a CNN-based feature space. EVT characterizes the distribution of outliers in long-tailed distributions to identify novelties. We conducted experiments on the xView dataset for object detection and classification in satellite imagery, reducing it to a classification dataset by using its annotated bounding boxes on objects and holding out a set of 18 of its 60 classes as novelties. Our results indicate that EVT is effective at distinguishing novel from known object classes, even when novel classes are similar to known ones. Dawei Du, Christopher Funk, Anthony Hoogs |
IGARSS | 3 |
| 2022 | 1st ACM SIGKDD Workshop on Ethical Artificial Intelligence: Methods and Applications (EAI-KDD22)abstractEthical AI has become increasingly important and it has been attracting attention from academia and industry, due to its increased popularity in real-world applications with fairness concerns. It also places fundamental importance on ethical considerations in determining legitimate and illegitimate uses of AI. Organizations that apply ethical AI have clearly stated well-defined review processes to ensure adherence to legal guidelines. Therefore, the wave of research at the intersection of ethical AI in data mining and machine learning has also influenced other fields of science, including computer vision, natural language processing, reinforcement learning, and social science. Despite these successes, ethical AI still faces many challenges. Consequently, there is an urgent need to bring experts and researchers together at prestigious venues to discuss ethical AI, which has been rarely seen in previous KDD conferences. This workshop will provide a premium platform for both research and industry from different backgrounds to exchange ideas on opportunities, challenges, and cutting-edge techniques in ethical AI. Chen Zhao 0010, Feng Chen 0001, Xintao Wu, Christopher Funk, Anthony Hoogs |
KDD | 5 |
| 2022 | X-MIR: EXplainable Medical Image RetrievalabstractDespite significant progress in the past few years, machine learning systems are still often viewed as "black boxes," which lack the ability to explain their output decisions. In high-stakes situations such as healthcare, there is a need for explainable AI (XAI) tools that can help open up this black box. In contrast to approaches which largely tackle classification problems in the medical imaging domain, we address the less-studied problem of explainable image retrieval. We test our approach on a COVID-19 chest X-ray dataset and the ISIC 2017 skin lesion dataset, showing that saliency maps help reveal the image features used by models to determine image similarity. We evaluated three different saliency algorithms, which were either occlusion-based, attention-based, or relied on a form of activation mapping. We also develop quantitative evaluation metrics that allow us to go beyond simple qualitative comparisons of the different saliency algorithms. Our results have the potential to aid clinicians when viewing medical images and addresses an urgent need for interventional tools in response to COVID-19. The source code is publicly available at: https://gitlab.kitware.com/brianhhu/x-mir. Brian Hu 0001, Bhavan Vasu, Anthony Hoogs |
WACV | 3 |
| 2021 | MEVA: A Large-Scale Multiview, Multimodal Video Dataset for Activity DetectionabstractWe present the Multiview Extended Video with Activities (MEVA) dataset[6], a new and very-large-scale dataset for human activity recognition. Existing security datasets either focus on activity counts by aggregating public video disseminated due to its content, which typically excludes same-scene background video, or they achieve persistence by observing public areas and thus cannot control for activity content. Our dataset is over 9300 hours of untrimmed, continuous video, scripted to include diverse, simultaneous activities, along with spontaneous background activity. We have annotated 144 hours for 37 activity types, marking bounding boxes of actors and props. Our collection observed approximately 100 actors performing scripted scenarios and spontaneous background activity over a three- week period at access-controled venue, collecting in multiple modalities with overlapping and non-overlapping in-door and outdoor viewpoints. The resulting data includes video from 38 RGB and thermal IR cameras, 42 hours of UAV footage, as well as GPS locations for the actors. 122 hours of annotation are sequestered in support of the NIST Activity in Extended Video (ActEV) challenge; the other 22 hours of annotation and the corresponding video are available on our website, along with an additional 306 hours of ground camera data, 4.6 hours of UAV data, and 9.6 hours of GPS logs. Additional derived data includes camera models geo-registering the outdoor cameras and a dense 3D point cloud model of the outdoor scene. The data was collected with IRB oversight and approval and released under a CC-BY-4.0 license. Kellie Corona, Katie Osterdahl, Roderic Collins, Anthony Hoogs |
WACV | 4 |
| 2020 | DOA-GAN: Dual-Order Attentive Generative Adversarial Network for Image Copy-Move Forgery Detection and LocalizationabstractImages can be manipulated for nefarious purposes to hide content or to duplicate certain objects through copy-move operations. Discovering a well-crafted copy-move forgery in images can be very challenging for both humans and machines; for example, an object on a uniform background can be replaced by an image patch of the same background. In this paper, we propose a Generative Adversarial Network with a dual-order attention model to detect and localize copy-move forgeries. In the generator, the first-order attention is designed to capture copy-move location information, and the second-order attention exploits more discriminative features for the patch co-occurrence. Both attention maps are extracted from the affinity matrix and are used to fuse location-aware and co-occurrence features for the final detection and localization branches of the network. The discriminator network is designed to further ensure more accurate localization results. To the best of our knowledge, we are the first to propose such a network architecture with the 1st-order attention mechanism from the affinity matrix. We have performed extensive experimental validation and our state-of-the-art results strongly demonstrate the efficacy of the proposed approach. Ashraful Islam, Chengjiang Long, Arslan Basharat, Anthony Hoogs |
CVPR | 4 |
| 2019 | Multi-Modal Detection Fusion on a Mobile UGV for Wide-Area, Long-Range SurveillanceabstractWe introduce a self-contained, mobile surveillance system designed to remotely detect and track people in real time, at long ranges, and over a wide field of view in cluttered urban and natural settings. The system is integrated with an unmanned ground vehicle, which hosts an array of four IR and four high-resolution RGB cameras, navigational sensors, and onboard processing computers. High-confidence, low-false-alarm-rate person tracks are produced by fusing motion detections and single-frame CNN person detections between co-registered RGB and IR video streams. Processing speeds are increased by using semantic scene segmentation and a tiered inference scheme to focus processing on the most salient regions of the 43° × 7.8° composite field of view. The system autonomously produces alerts of human presence and movement within the field of view, which are disseminated over a radio network and remotely viewed on a tablet computer. We present an ablation study quantifying the benefits that multi-sensor, multi-detector fusion brings to the problem of detecting people in challenging outdoor environments with shadows, occlusions, clutter, and variable weather conditions. Matt Brown, Keith Fieldhouse, Eran Swears, Paul Tunison, Adam Romlein, Anthony Hoogs |
WACV | 6 |
| 2019 | Deep Neural Networks in Fully Connected CRF for Image Labeling with Social Network MetadataabstractWe propose a novel method for predicting image labels by fusing image content descriptors with the social media context of each image. An image uploaded to a social media site such as Flickr often has meaningful, associated information, such as comments and other images the user has uploaded, that is complementary to pixel content and helpful in predicting labels. Prediction challenges such as ImageNet [6]and MSCOCO [19] use only pixels, while other methods make predictions purely from social media context [21]. Our method is based on a novel fully connected Conditional Random Field (CRF) framework, where each node is an image, and consists of two deep Convolutional Neural Networks (CNN) and one Recurrent Neural Network (RNN) that model both textual and visual node/image information. The edge weights of the CRF graph represent textual similarity and link-based metadata such as user sets and image groups. We model the CRF as an RNN for both learning and inference, and incorporate the weighted ranking loss and cross entropy loss into the CRF parameter optimization to handle the training data imbalance issue. Our proposed approach is evaluated on the MIR-9K dataset and experimentally outperforms current state-of-the-art approaches. Chengjiang Long, Roddy Collins, Eran Swears, Anthony Hoogs |
WACV | 4 |
| 2017 | Compact image representation by binary component analysisabstractWe propose binary component analysis (BCA) for compact image representation and apply it to fractional dimension reduction and Binaryface representation. BCA is similar to the widely used principal component analysis (PCA), but with the restriction of the base vectors taking binary values of +1 and -1, instead of real values. In a finite set of all binary vectors, the projection of the correlated data onto the binary components (BC) have the largest possible variances. BCA leads to fractional dimension reduction, where binary components successively reduce data variance. The top few BC's capture the largest variances, and each additional BC further reduces the variance by a fraction of the amount reduced in one dimension by PCA. BCA is applied to compact face representation as Binaryface and used for face classification. Zhaohui Sun, Anthony Hoogs |
ICIP | 2 |
| 2017 | Reflection correspondence for exposing photograph manipulationabstractModern photo editing software enables increasingly realistic image manipulations, through splicing, region copy and paste, and content-aware fill. One potential flaw in manipulations is generating realistic object reflections, such as in bodies of water or glass surfaces. Image reflections involve complicated interactions between lighting, surface materials, and geometry, and they can be very hard to fake. Any inconsistency between the directly observed scene and its corresponding reflection is a strong indication of content manipulation. We propose a novel, physical-level image forensic approach to verify the integrity of reflections through identifying mismatches between objects and their reflections. The proposed algorithm consists of reflection-invariant feature detection and matching, geometric transform estimation, and robust change detection between a scene region and its warped reflection region. This approach complements the large body of digital-level forensic methods, with the advantage of being robust to digital-level image transformations such as compression, blurring, and added noise. Experimental results on authentic and manipulated images demonstrate its efficacy. Eric Wengrowski, Zhaohui Sun, Anthony Hoogs |
ICIP | 3 |
| 2017 | An Open-Source Platform for Underwater Image and Video AnalyticsabstractGlobal fisheries and the future of sustainable seafood are predicated on healthy populations of various species of fish and shellfish. Recent developments in the collection of large-volume optical data by autonomous underwater vehicles (AUVs), stationary camera arrays, and towed vehicles has made it possible for fishery scientists to generate species-specific, size-structured abundance estimates for different species of marine organisms via imagery. The immense volume of data collected by such devices quickly exceeds manual processing capacity and creates a strong need for automatic image analysis. This paper presents an open-source computer vision software platform designed to integrate common image and video analytics, such as stereo calibration, object detection and object classification, into a sequential data processing pipeline that is easy to program, multi-threaded, and generic. The system provides a cross-language common interface for each of these components, multiple implementations of each, as well as unified methods for evaluating and visualizing the results of different methods for accomplishing the same task. Matthew Dawkins, Linus Sherrill, Keith Fieldhouse, Anthony Hoogs, Benjamin L. Richards, Lakshman Prasad, Kresimir Williams, Nathan Lauffenburger, Gaoang Wang |
WACV | 4 |
| 2016 | Tattoo detection and localization using region-based deep learningabstractTattoos have been increasingly used as a discriminative soft biometric for people identification, such as criminal and victim identification in forensics investigation and law enforcement. However, automatic detection of tattoo images and accurate localization of the regions of interest are challenged by the large variations in artistic composition, color, shape, texture, location on the body, local geometric shape (e.g. neck and finger), imaging conditions, and image quality. In this paper, we train a tattoo detector from the Tatt-C and PASCAL VOC 2007 image datasets using region-based deep learning. The detector can effectively determine if an image contains tattoos and the locations of tattoo regions. We carry out a comprehensive evaluation of our tattoo image classification and detection localization. The detector improves upon the state-of-the-art algorithms in the Tatt-C challenge, achieving a better detection error trade-off curve. It yields low confidence scores on randomly sampled non-tattoo images from 397 scene categories in the MIT-SUN dataset. In addition, the same detector is also validated on the NTU Tattoo Image Dataset with 10000 images. Zhaohui Sun, Jeff Baumes, Paul Tunison, Matthew W. Turek, Anthony Hoogs |
ICPR | 5 |
| 2016 | Image-oriented economic perspective on user behavior in multimedia social forums: An analysis on supply, consumption, and saliency
Sangmin Oh, Megha Pandey, Ilseo Kim, Anthony Hoogs |
Pattern Recognit. Lett. | 4 |
| 2015 | An end-to-end system for content-based video retrieval using behavior, actions, and appearance with interactive query refinementabstractWe describe a system for content-based retrieval from large surveillance video archives, using behavior, action and appearance of objects. Objects are detected, tracked, and classified into broad categories. Their behavior and appearance are characterized by action detectors and descriptors, which are indexed in an archive. Queries can be posed as video exemplars, and the results can be refined through relevance feedback. The contributions of our system include the fusion of behavior and action detectors with appearance for matching; the improvement of query results through interactive query refinement (IQR), which learns a discriminative classifier online based on user feedback; and reasonable performance on low resolution, poor quality video. The system operates on video from ground cameras and aerial platforms, both RGB and IR. Performance is evaluated on publicly-available surveillance datasets, showing that subtle actions can be detected under difficult conditions, with reasonable improvement from IQR. Anthony Hoogs, A. G. Amitha Perera, Roderic Collins, Arslan Basharat, Keith Fieldhouse, Chuck Atkins, Linus Sherrill, Benjamin Boeckel, Russell Blue, Matthew Woehlke, C. Greco, Zhaohui Sun, Eran Swears, Naresh P. Cuntoor, J. Luck, B. Drew, D. Hanson, D. Rowley, J. Kopaz, T. Rude, D. Keefe, Amit Srivastava, Saurabh Khanwalkar, Chia-Chih Chen, Jake K. Aggarwal, Larry Davis 0001, Yaser Yacoob, Dong Liu 0001, Shih-Fu Chang, Bi Song, Amit K. Roy-Chowdhury, Kenneth Sullivan, Jelena Tesic, Shivkumar Chandrasekaran, B. S. Manjunath, K. Reddy, Mubarak Shah, K. Chang, Tsuhan Chen, Mita Desai |
AVSS | 1 |
| 2014 | Real-time heads-up display detection in videoabstractVideo from surveillance cameras, aerial sensors, video games, and other sources may occasionally contain text, heads-up displays (HUDs), lens debris, or other artifacts superimposed on top of some scene. In standard video processing pipelines, the early detection and filtering of these image-plane aligned obstructions can be helpful for improving the accuracy of later operations, such as video stabilization, tracking, or object recognition. This paper presents one such technique to automatically accomplish this, which first extracts various pixel-level features which jointly take into account local spatiotemporal variations around each pixel. Features extracted from multiple frames are then utilized by a novel classification system to determine if any obstructions are present and, if possible, to categorize them into known types. Experimental results show promising performance on a variety of different categories of HUD, in addition to other types of on-screen display. Matthew Dawkins, A. G. Amitha Perera, Anthony Hoogs |
AVSS | 3 |
| 2014 | Complex algorithm optimization through probabilistic search of its configuration treeabstractWe present a novel algorithm to automatically configure complex video processing systems when adapting them to new datasets or scenarios. This has the main benefit of significantly reducing the time spent by a system expert on parameter tuning. Our approach has two main components: (1) a configuration tree structure that organizes system parameters and procedures in a systematic manner and facilitates identifying viable system configurations; and (2) a probabilistic sampling approach on the configuration tree that efficiently searches for the optimal configuration. We have successfully applied the proposed approach to optimize the configurations for two different video processing modules: a motion detection & tracking pipeline and a streaming video segmentation algorithm. The automatically discovered configurations produced better performance than the best manual configurations, while requiring significantly less human effort and domain-specific expertise. Yiliang Xu, Arslan Basharat, Jacob Becker, Anthony Hoogs |
AVSS | 4 |
| 2014 | An Efficient Online Hierarchical Supervoxel Segmentation Algorithm for Time-critical Applications
Yiliang Xu, Dezhen Song, Anthony Hoogs |
BMVC | 3 |
| 2014 | Complex Activity Recognition Using Granger Constrained DBN (GCDBN) in Sports and Surveillance VideoabstractModeling interactions of multiple co-occurring objects in a complex activity is becoming increasingly popular in the video domain. The Dynamic Bayesian Network (DBN) has been applied to this problem in the past due to its natural ability to statistically capture complex temporal dependencies. However, standard DBN structure learning algorithms are generatively learned, require manual structure definitions, and/or are computationally complex or restrictive. We propose a novel structure learning solution that fuses the Granger Causality statistic, a direct measure of temporal dependence, with the Adaboost feature selection algorithm to automatically constrain the temporal links of a DBN in a discriminative manner. This approach enables us to completely define the DBN structure prior to parameter learning, which reduces computational complexity in addition to providing a more descriptive structure. We refer to this modeling approach as the Granger Constraints DBN (GCDBN). Our experiments show how the GCDBN outperforms two of the most relevant state-of-the-art graphical models in complex activity classification on handball video data, surveillance data, and synthetic data. Eran Swears, Anthony Hoogs, Kim Boyer |
CVPR | 2 |
| 2014 | Context aided video-to-text information fusion
Erik Blasch, James G. Nagy, Alex Aved, Eric K. Jones, William M. Pottenger, Arslan Basharat, Anthony Hoogs, Riad I. Hammoud, Genshe Chen, Dan Shen 0004, Haibin Ling |
FUSION | 7 |
| 2014 | Personalized Economy of Images in Social Forums: An Analysis on Supply, Consumption, and SaliencyabstractIn this work, we focus on the novel problem of analyzing individual user's behavioral patterns regarding images shared on social forums. In particular, we view diverse user activities on social multimedia services as an economy, where the first activity mode of sharing or posting is interpreted as supply, and another mode of activity such as commenting on images is interpreted as consumption. To characterize user profiles in these two behavioral modes, we propose an approach to characterize users' supply and consumption profiles based on the image content types with which they engage. We then present various statistical analyses, which confirm that there is an unexpected significant difference between these two behavioral modes. In addition, we introduce a statistical approach to identify users with salient profiles, which can be useful for social multimedia services for blocking users with undesirable behavior or viral content promotion. We showcase the benefits of the proposed saliency detection approach and its extension to detect significant key images from a complex dataset, which exhibits the inherent multi-modal nature of user bases of social multimedia services. Sangmin Oh, Megha Pandey, Ilseo Kim, Anthony Hoogs, Jeff Baumes |
ICPR | 4 |
| 2014 | Real-time multi-target tracking at 210 megapixels/second in Wide Area Motion ImageryabstractWe present a real-time, full-frame, multi-target Wide Area Motion Imagery (WAMI) tracking system that utilizes distributed processing to handle high data rates while maintaining high track quality. The proposed architecture processes the WAMI data as a series of geospatial tiles and implements both process- and thread-level parallelism across multiple compute nodes. Each tile is processed independently, from decoding the image through generating tracks that are finally merged across all tiles by an inter-tile linker (ITL) module. A high performance PostgreSQL database with GIS extensions is used to control the flow of intermediate data between each tracking process. High quality tracks are produced efficiently due to robust, effective algorithmic modules including: multi-frame moving object detection and track initialization; tracking based on the fusion of motion and appearance with a goal of very pure tracks; and online track linking based on multiple features. In addition, we have configured a high-performance compute cluster using high density blade servers, Infiniband networking, and an HPC filesystem. The compute cluster enables full-frame, state-of-the-art tracking of vehicles or dismounts at the WAMI sensor's native 1.25Hz frame-rate, while only taking 7u of rack space and providing 210 megapixels/second throughput. Arslan Basharat, Matthew W. Turek, Yiliang Xu, Chuck Atkins, David Stoup, Keith Fieldhouse, Paul Tunison, Anthony Hoogs |
WACV | 8 |
| 2013 | System and algorithms on detection of objects embedded in perspective geometry using monocular camerasabstractIn this work, we present a framework to detect objects embedded in complex perspective geometry. Our goal is to accurately identify objects such as people standing in balconies or windows on building facades of surrounding buildings. Compared to traditional computer vision work focused on activity analysis from a horizontal view, our framework provides a solution for the application domain of mobile surveillance in urban areas. A novel solution for a monocular camera is formulated by tightly coupling various computational modules including geometric analysis, segmentation, scale estimation, and object detection. In particular, our proposed approach alleviates the effect of the perspective geometry and corresponding distortion in object appearance effectively, and provides accurate scale priors to eliminate unlikely object detection hypotheses. The experimental results on collected video dataset show that the proposed approach is more accurate than traditional detection approaches based on brute-force scanning windows. Yiliang Xu, Sangmin Oh, Fan Yang 0016, Zhuolin Jiang, Naresh P. Cuntoor, Anthony Hoogs, Larry Davis 0001 |
AVSS | 6 |
| 2013 | A Minimum Error Vanishing Point Detection Approach for Uncalibrated Monocular Images of Man-Made EnvironmentsabstractWe present a novel vanishing point detection algorithm for uncalibrated monocular images of man-made environments. We advance the state-of-the-art by a new model of measurement error in the line segment extraction and minimizing its impact on the vanishing point estimation. Our contribution is twofold: 1) Beyond existing hand-crafted models, we formally derive a novel consistency measure, which captures the stochastic nature of the correlation between line segments and vanishing points due to the measurement error, and use this new consistency measure to improve the line segment clustering. 2) We propose a novel minimum error vanishing point estimation approach by optimally weighing the contribution of each line segment pair in the cluster towards the vanishing point estimation. Unlike existing works, our algorithm provides an optimal solution that minimizes the uncertainty of the vanishing point in terms of the trace of its covariance, in a closed-form. We test our algorithm and compare it with the state-of-the-art on two public datasets: York Urban Dataset and Eurasian Cities Dataset. The experiments show that our approach outperforms the state-of-the-art. Yiliang Xu, Sangmin Oh, Anthony Hoogs |
CVPR | 3 |
| 2013 | Pyramid Coding for Functional Scene Element Recognition in Video ScenesabstractRecognizing functional scene elements in video scenes based on the behaviors of moving objects that interact with them is an emerging problem of interest. Existing approaches have a limited ability to characterize elements such as cross-walks, intersections, and buildings that have low activity, are multi-modal, or have indirect evidence. Our approach recognizes the low activity and multi-model elements (crosswalks/intersections) by introducing a hierarchy of descriptive clusters to form a pyramid of codebooks that is sparse in the number of clusters and dense in content. The incorporation of local behavioral context such as person-enter-building and vehicle-parking nearby enables the detection of elements that do not have direct motion-based evidence, e.g. buildings. These two contributions significantly improve scene element recognition when compared against three state-of-the-art approaches. Results are shown on typical ground level surveillance video and for the first time on the more complex Wide Area Motion Imagery. Eran Swears, Anthony Hoogs, Kim Boyer |
ICCV | 2 |
| 2012 | Human Action Recognition in Large-Scale Datasets Using Histogram of Spatiotemporal GradientsabstractResearch in human action recognition has advanced along multiple fronts in recent years to address various types of actions including simple, isolated actions in staged data (e.g., KTH dataset), complex actions (e.g., Hollywood dataset) and naturally occurring actions in surveillance videos (e.g, VIRAT dataset). Several techniques including those based on gradient, flow and interest-points have been developed for their recognition. Most perform very well in standard action recognition datasets, but fail to produce similar results in more complex, large-scale datasets. Here we analyze the reasons for this less than successful generalization by considering a state-of-the-art technique, histogram of oriented gradients in spatiotemporal volumes as an example. This analysis may prove useful in developing robust and effective techniques for action recognition. Kishore K. Reddy, Naresh P. Cuntoor, A. G. Amitha Perera, Anthony Hoogs |
AVSS | 4 |
| 2012 | Human-robot teamwork using activity recognition and human instructionabstractWe address the problem of robot participation as a team member in group activities. In this human-robot interactive setting, the goal is to enable a robot to detect objects around it using on-board video cameras, and determine which activity they are performing. It then determines how it should move to participate in the activity based on its intended role, derived from domain knowledge, and human instruction. A human controller can initialize, correct and refine the robot's behavior and recognition of surrounding activities. A probabilistic framework is used for joint estimation of the robot's desired motion and the activity surrounding it while incorporating human instruction. Group activities consisting of multiple people moving about the robot with a common goal are used to demonstrate the proposed approach. We focus on activities with geometric structures (line, wedge, echelon, etc.) with an arbitrary number of participants, and handle the noisy observations from video detection and tracking. Results using both simulated and real video data are used to demonstrate activity recognition and incorporating human instruction. Naresh P. Cuntoor, Roderic Collins, Anthony Hoogs |
IROS | 3 |
| 2012 | Learning and recognizing complex multi-agent activities with applications to american football playsabstractWe are interested in modeling and recognizing complex behaviors in video, where multiple agents are interacting in a time-varying manner and in a spatially-localized domain such as American football. Our approach pushes the model complexity onto the observations by using a multi-variate kernel density while maintaining a simple HMM model. The temporal interactions of objects are captured by coupling the kernel observation distributions with a time-varying state-transition matrix, producing a Non-Stationary Kernel HMM (NSK-HMM). This modeling philosophy specifically addresses several issues that plague the more complex stationary models with simple observations, i.e. Dynamic Multi-Linked HMM (DML-HMM) and the Time-Delayed Probabilistic Graphical Model (TDPGM). These include: smaller training datasets, sensitivity to intra class variability and/or dense uninformative clutter tracks. Experiments are performed in the American football video domain, where the offensive plays are the activities. Comparisons are made to the DML-HMM and an extension of the TDPGM to DBNs (TDDBN). The NSK-HMM achieves a 57.7% classification accuracy across seven activities, while the DML-HMM is 26.7% and the TDDBN is 21.3%. When tested on four activities the NSK-HMM achieves a 76.0% accuracy. Eran Swears, Anthony Hoogs |
WACV | 2 |
| 2011 | AVSS 2011 demo session: A large-scale benchmark dataset for event recognition in surveillance videoabstractSummary form only given. We present a concept for automatic construction site monitoring by taking into account 4D information (3D over time), that is acquired from highly-overlapping digital aerial images. On the one hand today's maturity of flying micro aerial vehicles (MAVs) enables a low-cost and an efficient image acquisition of high-quality data that maps construction sites entirely from many varying viewpoints. On the other hand, due to low-noise sensors and high redundancy in the image data, recent developments in 3D reconstruction workflows have benefited the automatic computation of accurate and dense 3D scene information. Having both an inexpensive high-quality image acquisition and an efficient 3D analysis workflow enables monitoring, documentation and visualization of observed sites over time with short intervals. Relating acquired 4D site observations, composed of color, texture, geometry over time, largely supports automated methods toward full scene understanding, the acquisition of both the change and the construction site's progress. Sangmin Oh, Anthony Hoogs, A. G. Amitha Perera, Naresh P. Cuntoor, Chia-Chih Chen, Jong Taek Lee, Saurajit Mukherjee, Jake K. Aggarwal, Hyungtae Lee, Larry Davis 0001, Eran Swears, Xiaoyang Wang 0001, Kishore K. Reddy, Mubarak Shah, Carl Vondrick, Hamed Pirsiavash, Deva Ramanan, Jenny Yuen, Antonio Torralba 0001, Bi Song, Anesco Fong, Amit K. Roy-Chowdhury, Mita Desai |
AVSS | 2 |
| 2011 | A large-scale benchmark dataset for event recognition in surveillance videoabstractWe introduce a new large-scale video dataset designed to assess the performance of diverse visual event recognition algorithms with a focus on continuous visual event recognition (CVER) in outdoor areas with wide coverage. Previous datasets for action recognition are unrealistic for real-world surveillance because they consist of short clips showing one action by one individual [15, 8]. Datasets have been developed for movies [11] and sports [12], but, these actions and scene conditions do not apply effectively to surveillance videos. Our dataset consists of many outdoor scenes with actions occurring naturally by non-actors in continuously captured videos of the real world. The dataset includes large numbers of instances for 23 event types distributed throughout 29 hours of video. This data is accompanied by detailed annotations which include both moving object tracks and event examples, which will provide solid basis for large-scale evaluation. Additionally, we propose different types of evaluation modes for visual recognition tasks and evaluation metrics along with our preliminary experimental results. We believe that this dataset will stimulate diverse aspects of computer vision research and help us to advance the CVER tasks in the years ahead. Sangmin Oh, Anthony Hoogs, A. G. Amitha Perera, Naresh P. Cuntoor, Chia-Chih Chen, Jong Taek Lee, Saurajit Mukherjee, Jake K. Aggarwal, Hyungtae Lee, Larry Davis 0001, Eran Swears, Xiaoyang Wang 0001, Kishore K. Reddy, Mubarak Shah, Carl Vondrick, Hamed Pirsiavash, Deva Ramanan, Jenny Yuen, Antonio Torralba 0001, Bi Song, Anesco Fong, Amit K. Roy-Chowdhury, Mita Desai |
CVPR | 2 |
| 2010 | Content-Based Retrieval of Functional Objects in Video Using Scene Context
Sangmin Oh, Anthony Hoogs, Matthew W. Turek, Roderic Collins |
ECCV (1) | 2 |
| 2010 | Unsupervised Learning of Functional Categories in Video Scenes
Matthew W. Turek, Anthony Hoogs, Roderic Collins |
ECCV (2) | 2 |
| 2010 | Track Initialization in Low Frame Rate and Low Resolution VideosabstractThe problem of object detection and tracking has received relatively less attention in low frame rate and low resolution videos. Here we focus on motion segmentation in videos where objects appear small (less than 30-pixel tall people) and have low frame rate (less than 5 Hz). We study challenging cases where some of the, otherwise successful, approaches may break down. We investigate a number of popular techniques in computer vision that have been shown to be useful for discriminating various spatio-temporal signatures. These include: Histogram of oriented Gradients (HOG), Histogram of oriented optical Flow (HOF) and Haar-features (Viola and Jones). We use these feature to classify the motion segmentations into person vs. other and vehicle vs. other. We rely on aligned motion history images to create a more consistent object representation across frames. We present results on these features using webcam data and wide-area aerial video sequences. Naresh P. Cuntoor, Arslan Basharat, A. G. Amitha Perera, Anthony Hoogs |
ICPR | 4 |
| 2010 | Unsupervised Learning of Activities in Video Using Scene ContextabstractUnsupervised learning of semantic activities from video collected over time is an important problem for visual surveillance and video scene understanding. Our goal is to cluster tracks into semantically interpretable activity models that are independent of scene locations; most previous work in video scene understanding is focused on learning location-specific normalcy models. Location-independent models can be used to detect instances of the same activity anywhere in the scene, or even across multiple scenes. Our insight for this unsupervised activity learning problem is to incorporate scene context to characterize the behavior of every track. By scene context, we mean local scene structures, such as building entrances, parking spots and roads, that moving objects frequently interact with. Each track is attributed with large number of potentially useful features that capture the relationships and interactions with a set of existing scene context elements. Once feature vectors are obtained, tracks are grouped in this feature space using state-of-the-art clustering techniques, without considering scene location. Experiments are conducted on webcam video of a complex scene, with many interacting objects and very noisy tracks resulting from low frame rates and poor image quality. Our results demonstrate that location-independent and semantically interpretable groupings can be successfully obtained using unsupervised clustering methods, and that the models are superior to standard location-dependent clustering. Sangmin Oh, Anthony Hoogs |
ICPR | 2 |
| 2010 | Image Comparison by Compound Disjoint Information with Applications to Perceptual Visual Quality Assessment, Image Registration and Tracking
Zhaohui Sun, Anthony Hoogs |
Int. J. Comput. Vis. | 2 |
| 2008 | Video Activity Recognition in the Real World
Anthony Hoogs, A. G. Amitha Perera |
AAAI | 1 |
| 2008 | Evaluation of Localized Semantics: Data, Methodology, and Experiments
Kobus Barnard, Quanfu Fan, Ranjini Swaminathan, Anthony Hoogs, Roderic Collins, Pascale Rondot, John P. Kaufhold |
Int. J. Comput. Vis. | 4 |
| 2006 | Object Boundary Detection in Images using a Semantic Ontology
Anthony Hoogs, Roderic Collins |
AAAI | 1 |
| 2006 | Joint Recognition of Complex Events and Track MatchingabstractWe present a novel method for jointly performing recognition of complex events and linking fragmented tracks into coherent, long-duration tracks. Many event recognition methods require highly accurate tracking, and may fail when tracks corresponding to event actors are fragmented or partially missing. However, these conditions occur frequently from occlusions, traffic and tracking errors. Recently, methods have been proposed for linking track fragments from multiple objects under these difficult conditions. Here, we develop a method for solving these two problems jointly. A hypothesized event model, represented as a Dynamic Bayes Net, supplies data-driven constraints on the likelihood of proposed track fragment matches. These event-guided constraints are combined with appearance and kinematic constraints used in the previous track linking formulation. The result is the most likely track linking solution given the event model, and the highest event score given all of the track fragments. The event model with the highest score is determined to have occurred, if the score exceeds a threshold. Results demonstrated on a busy scene of airplane servicing activities, where many non-event movers and long fragmented tracks are present, show the promise of the approach to solving the joint problem. Michael T. Chan, Anthony Hoogs, Rahul Bhotika, A. G. Amitha Perera, John Schmiederer, Gianfranco Doretto |
CVPR (2) | 2 |
| 2006 | Multi-Object Tracking Through Simultaneous Long Occlusions and Split-Merge ConditionsabstractA fundamental requirement for effective automated analysis of object behavior and interactions in video is that each object must be consistently identified over time. This is difficult when the objects are often occluded for long periods: nearly all tracking algorithms will terminate a track with loss of identity on a long gap. The problem is further confounded by objects in close proximity, tracking failures due to shadows, etc. Recently, some work has been done to address these issues using higher level reasoning, by linking tracks from multiple objects over long gaps. However, these efforts have assumed a one-to-one correspondence between tracks on either side of the gap. This is often not true in real scenarios of interest, where the objects are closely spaced and dynamically occlude each other, causing trackers to merge objects into single tracks. In this paper, we show how to efficiently handle splitting and merging during track linking. Moreover, we show that we can maintain the identities of objects that merge together and subsequently split. This enables the identity of objects to be maintained throughout long sequences with difficult conditions. We demonstrate our approach on a highly challenging, oblique-view video sequence of dense traffic of a highway interchange. We successfully track the large majority of the hundreds of moving vehicles in the scene, many in close proximity, through long occlusions and shadows. A. G. Amitha Perera, Chukka Srinivas, Anthony Hoogs, Glen Brooksby, Wensheng Hu |
CVPR (1) | 3 |
| 2006 | Image Comparison by Compound Disjoint InformationabstractIn this paper, we study disjoint information as a metric for image comparison and its applications in image matching, alignment, and video tracking. Disjoint information is the joint entropy of random variables excluding the mutual information. This measure of statistical dependence and information redundancy satisfies more rigorous metric conditions than mutual information. For image comparison, compound disjoint information is derived from the marginal densities of the image distributions. By using marginal densities other than color histograms, it can overcome the difficulties (such as a lack of spatial information) inherent in histogram-based mutual information methods and enrich the vocabulary of image description. Disjoint information is not sensitive to illumination and appearance changes, and it is particularly suited for multimodal applications. Zhaohui Sun, Anthony Hoogs |
CVPR (1) | 2 |
| 2005 | A Unified Framework for Tracking through Occlusions and across Sensor GapsabstractA common difficulty encountered in tracking applications is how to track an object that becomes totally occluded, possibly for a significant period of time. Another problem is how to associate objects, or tracklets, across non-overlapping cameras, or between observations of a moving sensor that switches fields of regard. A third problem is how to update appearance models for tracked objects over time. As opposed to using a comprehensive multi-object tracker that must simultaneously deal with these tracking challenges, we present a novel, modular framework that handles each of these problems in a unified manner by the initialization, tracking, and linking of high-confidence tracklets. In this track/suspend/match paradigm, we first analyze the scene to identify areas where tracked objects are likely to become occluded. Tracking is then suspended on occluded objects and re-initiated when they emerge from behind the occlusion. We then associate, or match, suspended tracklets with the new tracklets using full kinematic models for object motion and Gibbsian distributions for object appearance in order to complete the track through the occlusion. Sensor gaps are handled in a similar manner, where tracking is suspended when the sensor looks away and then re-initiated when the sensor returns. Changes in object appearance and orientation during tracking are also seamlessly handled in this framework. Tracklets with low lock scores are terminated. Tracking then resumes on untracked movers with corresponding updated appearance models. These new tracklets are then linked back to the terminated ones as appropriate. Fully automatic tracking results from a moving sensor are presented. Robert Kaucic, A. G. Amitha Perera, Glen Brooksby, John P. Kaufhold, Anthony Hoogs |
CVPR (1) | 5 |
| 2004 | Learning to Segment Images Using Region-Based Perceptual Features
John P. Kaufhold, Anthony Hoogs |
CVPR (2) | 2 |
| 2003 | Video Content Annotation Using Visual Analysis and a Large Semantic KnowledgebaseabstractWe present a novel approach to automatically annotating broadcast video. To manage the enormous variety of objects, events and scenes in video problem domains such as news video, we couple generic image analysis with a semantic database, WordNet, containing huge amounts of real-world information. Object and event recognition are performed by searching WordNet for concepts jointly supported by image evidence and topic context derived from the video transcript. No object- specific or event-specific training is required, and only a few object models and detection algorithms are required to label much of the significant content of news video. The hierarchical structure of WordNet yields hierarchical recognition, dynamically tailored to the level of supporting image evidence. The potential of the approach is demonstrated by analyzing a wide variety of scenes in news video. Anthony Hoogs, Jens Rittscher, Gees C. Stein, John Schmiederer |
CVPR (2) | 1 |
| 2003 | Enabling video annotation using a semantic database extended with visual knowledgeabstractA semantic database has been extended with visual information to enable video annotation. This paper describes a lexical database, WordNet. We show its limitations with respect to describing visual characteristics, and describe an extension to WordNet that contains specific visual information. Having such a semantic database makes video annotation possible for broadcast news: a domain that can cover any topic and involve a wide variety of events, objects and scenes. Combining basic visual analysis techniques and a semantic database containing visual descriptions avoids the problem developing large numbers of specific object and event detectors. Such a semantic database can be of great value for the analysis of multi-modal information. As far as we know, such a database has not been developed before. Gees C. Stein, Jens Rittscher, Anthony Hoogs |
ICME | 3 |
| 2003 | A Common Set of Perceptual Observables for Grouping, Figure-Ground Discrimination, and Texture ClassificationabstractWe present a complete set of perceptual observables that provides a unified image description for grouping, figure-ground separation, and texture analysis. Although much progress has been made recently in treating contours and texture simultaneously for image segmentation and grouping, current approaches rely on different models for contours, regions, and texture such as one-dimensional intensity discontinuities for contours and filter bank responses for texture. This results in expensive computation that arbitrates between these disparate representations at each pixel. In our approach, salient image content such as contours, regions, and texture are represented in a common, low-level framework of image observables. We model the image as a partition of surfaces bounded by intensity discontinuities and derive perceptual measures as relations between neighboring surfaces. This enables us to extend the traditional Gestalt measures based on local edge geometry and contrast to region-based measures that jointly exploit large scale image topology, photometry, and geometry. These measures provide a natural basis for grouping on multidimensional similarity criteria and texture is directly derived as relational properties on local region neighborhoods. The viability of our model is demonstrated by applying the common observables to texture recognition, figure-ground separation, and generic image segmentation. Anthony Hoogs, Roderic Collins, Robert Kaucic, Joseph L. Mundy |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2000 | An Integrated Boundary and Region Approach to Perceptual GroupingabstractThe primary focus of work on perceptual grouping has been geometric constraints derived from projected object boundaries. Although boundaries can be robustly extracted under some conditions, much intensity information in the image is ignored. In this work, we incorporate image topology and intensity into grouping for object detection and recognition. Using boundary-based region segmentation, region intensity, geometry and topology are exploited in forming groups of regions that satisfy high-level constraints at arbitrary scales. It is demonstrated that regions provide a natural and computationally effective framework for discovering localized image structure, such as man-made objects in natural scenes. Anthony Hoogs, Joseph L. Mundy |
ICPR | 1 |
| 1997 | Analysis of Learning Using Segmentation Models
Anthony Hoogs |
CAIP | 1 |
| 1996 | Model-based learning of segmentationsabstractA method for integrating image segmentation information into geometric models is presented. The resulting object representation has advantages of both model-based and view-based representations, in that model geometry plus learned appearance information is used to improve the prediction of object appearance over purely geometric methods. The combined models are constructed over a training set of imagery using prior geometric models. Segmentation features are matched to the geometric models, and an evidential framework is used to characterize the segmentations of model features. To test the validity of the models, a pose adjustment system was modified to incorporate the prior segmentation information. Results indicate that the inclusion of the segmentation information significantly improves pose adjustment accuracy over using purely geometric information for model appearance. Anthony Hoogs, Ruzena Bajcsy |
ICPR | 1 |
| 1995 | Segmentation Modeling
Anthony Hoogs, Ruzena Bajcsy |
CAIP | 1 |
| 1992 | On using CAD models to compute the pose of curved 3D objects
Jean Ponce, Anthony Hoogs, David J. Kriegman |
CVGIP Image Underst. | 2 |