VLDB 2026 Research / reviewers in the wild / expert
Jason J. Corso
dblp:68/2447
· DBLP profile ↗
114ranked-venue papers
11as first author
18since 2021 · last 2026
0000-0001-6454-9594ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 77 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 73 · 7 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 4 first-authorSystems, architecture and hardware · 9 · 1 first-authorDatabases, data management, data science and information retrieval · 5Human-computer interaction and ubiquitous computing · 4 · 2 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Zero-Shot Coreset Selection via Iterative Subspace SamplingabstractDeep learning increasingly relies on massive data with substantial storage, annotation, and training costs. To reduce costs, coreset selection finds a representative subset of data to train models while ideally performing on par with the full data training. To maximize performance, current state-of-the-art coreset methods select data using dataset-specific ground truth labels and training. However, these methodological requirements prevent selection at scale on real-world, unlabeled data. To that end, this paper addresses the selection of coresets that achieve state-of-the-art performance but without using any labels or training on candidate data. Instead, our solution, Zero-Shot Coreset Selection via Iterative Subspace Sampling (ZCore), uses previously-trained foundation models to generate zero-shot, high-dimensional embedding spaces to interpret unlabeled data. ZCore then iteratively quantifies the relative value of all candidate data based on coverage and redundancy in numerous subspace distributions. Finally, ZCore selects a coreset sized for any data budget to train downstream models. We evaluate ZCore on four datasets and outperform several state-of-the-art label-based methods, especially at low data rates that provide the most substantial cost reduction. On ImageNet, ZCore selections for 10% training data achieve a downstream validation accuracy of 53.99%, which outperforms prior label-based methods and removes annotation and training costs for 1.15 million images.‡ Brent A. Griffin, Jacob Marks, Jason J. Corso |
WACV | 3 |
| 2025 | Transparent and Coherent Procedural Mistake DetectionabstractProcedural mistake detection (PMD) is a challenging problem of classifying whether a human user (observed through egocentric video) has successfully executed a task (specified by a procedural text).Despite significant recent efforts, machine performance in the wild remains nonviable, and the reasoning processes underlying this performance are opaque.As such, we extend PMD to require generating visual self-dialog rationales to inform decisions.Given the impressive, mature image understanding capabilities observed in recent visionand-language models (VLMs), we curate a suitable benchmark dataset for PMD based on individual frames.As our reformulation enables unprecedented transparency, we leverage a natural language inference (NLI) model to formulate two automated metrics for the coherence of generated rationales.We establish baselines for this reframed task, showing that VLMs struggle off-the-shelf, but with some trade-offs, their accuracy, coherence, and efficiency can be improved by incorporating these metrics into common inference and finetuning methods.Lastly, our multi-faceted metrics visualize common outcomes, highlighting areas for further improvement. Shane Storks, Itamar Bar-Yossef, Yayuan Li, Jason J. Corso, Joyce Y. Chai |
EMNLP | 5 |
| 2025 | VITRO: Vocabulary Inversion for Time-series Representation OptimizationabstractAlthough LLMs have demonstrated remarkable capabilities in processing and generating textual data, their pretrained vocabularies are ill-suited for capturing the nuanced temporal dynamics and patterns inherent in time series. The discrete, symbolic nature of natural language tokens, which these vocabularies are designed to represent, does not align well with the continuous, numerical nature of time series data. To address this fundamental limitation, we propose VITRO. Our method adapts textual inversion optimization from the vision-language domain in order to learn a new time series per-dataset vocabulary that bridges the gap between the discrete, semantic nature of natural language and the continuous, numerical nature of time series data. We show that learnable time series-specific pseudo-word embeddings represent time series data better than existing general language model vocabularies, with VITRO-enhanced methods achieving state-of-the-art performance in long-term forecasting across most datasets. Filippos Bellos, Nam H. Nguyen, Jason J. Corso |
ICASSP | 3 |
| 2024 | Measuring Physical Plausibility of 3D Human Poses Using Physics Simulation
Nathan Louis, Mahzad Khoshlessan, Jason J. Corso |
BMVC | 3 |
| 2024 | Slimming Neural Networks Using Adaptive Connectivity ScoresabstractIn general, deep neural network (DNN) pruning methods fall into two categories: 1) weight-based deterministic constraints and 2) probabilistic frameworks. While each approach has its merits and limitations, there are a set of common practical issues such as trial-and-error to analyze sensitivity and hyper-parameters to prune DNNs, which plague them both. In this work, we propose a new single-shot, fully automated pruning algorithm called slimming neural networks using adaptive connectivity scores (SNACS). Our proposed approach combines a probabilistic pruning framework with constraints on the underlying weight matrices, via a novel connectivity measure, at multiple levels to capitalize on the strengths of both approaches while solving their deficiencies. In SNACS, we propose a fast hash-based estimator of adaptive conditional mutual information (ACMI), that uses a weight-based scaling criterion, to evaluate the connectivity between filters and prune unimportant ones. To automatically determine the limit up to which a layer can be pruned, we propose a set of operating constraints that jointly define the upper pruning percentage limits across all the layers in a deep network. Finally, we define a novel sensitivity criterion for filters that measures the strength of their contributions to the succeeding layer and highlights critical filters that need to be completely protected from pruning. Through our experimental validation, we show that SNACS is faster by over 17× the nearest comparable method and is the state-of-the-art single-shot pruning method across four standard Dataset-DNN pruning benchmarks: CIFAR10-VGG16, CIFAR10-ResNet56, CIFAR10-MobileNetv2, and ILSVRC2012-ResNet50. Madan Ravi Ganesh, Dawsin Blanchard, Jason J. Corso, Salimeh Yasaei Sekeh |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Evaluating and Improving Interactions with Hazy OraclesabstractMany AI systems integrate sensor inputs, world knowledge, and human-provided information to perform inference. While such systems often treat the human input as flawless, humans are better thought of as hazy oracles whose input may be ambiguous or outside of the AI system's understanding. In such situations it makes sense for the AI system to defer its inference while it disambiguates the human-provided information by, for example, asking the human to rephrase the query. Though this approach has been considered in the past, current work is typically limited to application-specific methods and non-standardized human experiments. We instead introduce and formalize a general notion of deferred inference. Using this formulation, we then propose a novel evaluation centered around the Deferred Error Volume (DEV) metric, which explicitly considers the tradeoff between error reduction and the additional human effort required to achieve it. We demonstrate this new formalization and an innovative deferred inference method on the disparate tasks of Single-Target Video Object Tracking and Referring Expression Comprehension, ultimately reducing error by up to 48% without any change to the underlying model or its parameters. Stephan J. Lemmer, Jason J. Corso |
AAAI | 2 |
| 2023 | Can Deep Networks be Highly Performant, Efficient and Robust simultaneously?
Madan Ravi Ganesh, Salimeh Yasaei Sekeh, Jason J. Corso |
BMVC | 3 |
| 2023 | Iterative Vision-and-Language NavigationabstractWe present Iterative Vision-and-Language Navigation (IVLN), a paradigm for evaluating language-guided agents navigating in a persistent environment over time. Existing Vision-and-Language Navigation (VLN) benchmarks erase the agent's memory at the beginning of every episode, testing the ability to perform cold-start navigation with no prior information. However, deployed robots occupy the same environment for long periods of time. The IVLN paradigm addresses this disparity by training and evaluating VLN agents that maintain memory across tours of scenes that consist of up to 100 ordered instruction-following Room-to-Room (R2R) episodes, each defined by an individual language instruction and a target path. We present discrete and continuous Iterative Room-to-Room (IR2R) benchmarks comprising about 400 tours each in 80 indoor scenes. We find that extending the implicit memory of high-performing transformer VLN agents is not sufficient for IVLN, but agents that build maps can benefit from environment persistence, motivating a renewed focus on map-building agents in VLN. Jacob Krantz, Shurjo Banerjee, Wang Zhu 0001, Jason J. Corso, Stefan Lee, Jesse Thomason |
CVPR | 4 |
| 2023 | Human-Centered Deferred Inference: Measuring User Interactions and Setting Deferral Criteria for Human-AI TeamsabstractAlthough deep learning holds the promise of novel and impactful interfaces, realizing such promise in practice remains a challenge: since dataset-driven deep-learned models assume a one-time human input, there is no recourse when they do not understand the input provided by the user. Works that address this via deferred inference—soliciting additional human input when uncertain—show meaningful improvement, but ignore key aspects of how users and models interact. In this work, we focus on the role of users in deferred inference and argue that the deferral criteria should be a function of the user and model as a team, not simply the model itself. In support of this, we introduce a novel mathematical formulation, validate it via an experiment analyzing the interactions of 25 individuals with a deep learning-based visiolinguistic model, and identify user-specific dependencies that are under-explored in prior work. We conclude by demonstrating two human-centered procedures for setting deferral criteria that are simple to implement, applicable to a wide variety of tasks, and perform equal to or better than equivalent procedures that use much larger datasets. Stephan J. Lemmer, Anhong Guo, Jason J. Corso |
IUI | 3 |
| 2022 | The DEVIL is in the Details: A Diagnostic Evaluation Benchmark for Video InpaintingabstractQuantitative evaluation has increased dramatically among recent video inpainting work, but the video and mask content used to gauge performance has received relatively little attention. Although attributes such as camera and background scene motion inherently change the difficulty of the task and affect methods differently, existing evaluation schemes fail to control for them, thereby providing minimal insight into inpainting failure modes. To address this gap, we propose the Diagnostic Evaluation of Video Inpainting on Landscapes (DEVIL) benchmark, which consists of two contributions: (i) a novel dataset of videos and masks labeled according to several key inpainting failure modes, and (ii) an evaluation scheme that samples slices of the dataset characterized by a fixed content attribute, and scores performance on each slice according to reconstruction, realism, and temporal consistency quality. By revealing systematic changes in performance induced by particular characteristics of the input content, our challenging benchmark enables more insightful analysis into video inpainting methods and serves as an invaluable diagnostic tool for the field. Our code and data are available at github.com/MichiganCOG/devil. Ryan Szeto, Jason J. Corso |
CVPR | 2 |
| 2022 | Learning to Estimate External Forces of Human Motion in VideoabstractAnalyzing sports performance or preventing injuries requires capturing ground reaction forces (GRFs) exerted by the human body during certain movements. Standard practice uses physical markers paired with force plates in a controlled environment, but this is marred by high costs, lengthy implementation time, and variance in repeat experiments; hence, we propose GRF inference from video. While recent work has used LSTMs to estimate GRFs from 2D viewpoints, these can be limited in their modeling and representation capacity. First, we propose using a transformer architecture to tackle the GRF from video task, being the first to do so. Then we introduce a new loss to minimize high impact peaks in regressed curves. We also show that pre-training and multi-task learning on 2D-to-3D human pose estimation improves generalization to unseen motions. And pre-training on this different task provides good initial weights when finetuning on smaller (rarer) GRF datasets. We evaluate on LAAS Parkour and a newly collected ForcePose dataset; we show up to 19% decrease in error compared to prior approaches. Nathan Louis, Jason J. Corso, Tylan N. Templin, Travis D. Eliason, Daniel P. Nicolella |
ACM Multimedia | 2 |
| 2022 | Guest Editorial Introduction to the Special Section on Video and LanguageabstractComputer Vision (CV) and Natural Language Processing (NLP) are two most fundamental disciplines under a broad area of artificial intelligence (AI). CV is regarded as a field of research that explores the techniques to teach computers to see and understand digital content such as images and videos. NLP is a branch of linguistics that enables computers to process, interpret, and even generate human language. With the rise and development of deep learning over the past decade, there has been a steady momentum of innovation and breakthroughs that convincingly push the limits and improve the state-of-the-art of both vision and language modeling. An interesting observation is that the research in the two areas starts to interact, with a significant growth in both the volume of publications and extensive applications. Meanwhile, many previous experiences have shown that this can naturally build up the circle of human intelligence. Tao Mei 0001, Jason J. Corso, Gunhee Kim, Jiebo Luo 0001, Chunhua Shen, Hanwang Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Crowdsourcing More Effective Initializations for Single-Target Trackers Through Automatic Re-queryingabstractIn single-target video object tracking, an initial bounding box is drawn around a target object and propagated through a video. When this bounding box is provided by a careful human expert, it is expected to yield strong overall tracking performance that can be mimicked at scale by novice crowd workers with the help of advanced quality control methods. However, we show through an investigation of 900 crowdsourced initializations that such quality control strategies are inadequate for this task in two major ways: first, the high level of redundancy in these methods (e.g., averaging multiple responses to reduce error) is unnecessary, as 23% of crowdsourced initializations perform just as well as the gold-standard initialization. Second, even nearly perfect initializations can lead to degraded long-term performance due to the complexity of object tracking. Considering these findings, we evaluate novel approaches for automatically selecting bounding boxes to re-query, and introduce Smart Replacement, an efficient method that decides whether to use the crowdsourced replacement initialization. Stephan J. Lemmer, Jean Y. Song, Jason J. Corso |
CHI | 3 |
| 2021 | Depth From Camera Motion and Object DetectionabstractThis paper addresses the problem of learning to estimate the depth of detected objects given some measurement of camera motion (e.g., from robot kinematics or vehicle odometry). We achieve this by 1) designing a recurrent neural network (DBox) that estimates the depth of objects using a generalized representation of bounding boxes and uncalibrated camera movement and 2) introducing the Object Depth via Motion and Detection Dataset (ODMD). ODMD training data are extensible and configurable, and the ODMD benchmark includes 21,600 examples across four validation and test sets. These sets include mobile robot experiments using an end-effector camera to locate objects from the YCB dataset and examples with perturbations added to camera motion or bounding box data. In addition to the ODMD benchmark, we evaluate DBox in other monocular application domains, achieving state-of-the-art results on existing driving and robotics benchmarks and estimating the depth of objects using a camera phone. Brent A. Griffin, Jason J. Corso |
CVPR | 2 |
| 2021 | Ground-truth or DAER: Selective Re-query of Secondary InformationabstractMany vision tasks use secondary information at inference time—a seed—to assist a computer vision model in solving a problem. For example, an initial bounding box is needed to initialize visual object tracking. To date, all such work makes the assumption that the seed is a good one. However, in practice, from crowdsourcing to noisy automated seeds, this is often not the case. We hence propose the problem of seed rejection—determining whether to reject a seed based on the expected performance degradation when it is provided in place of a gold-standard seed. We provide a formal definition to this problem, and focus on two meaningful subgoals: understanding causes of error and understanding the model’s response to noisy seeds conditioned on the primary input. With these goals in mind, we propose a novel training method and evaluation metrics for the seed rejection problem. We then use seeded versions of the viewpoint estimation and fine-grained classification tasks to evaluate these contributions. In these experiments, we show our method can reduce the number of seeds that need to be reviewed for a target performance by over 23% compared to strong baselines. Stephan J. Lemmer, Jason J. Corso |
ICCV | 2 |
| 2021 | Cross-View Exocentric to Egocentric Video SynthesisabstractCross-view video synthesis task seeks to generate video sequences of one view from another dramatically different view. In this paper, we investigate the exocentric (third-person) view to egocentric (first-person) view video generation task. This is challenging because egocentric view sometimes is remarkably different from the exocentric view. Thus, transforming the appearances across the two different views is a non-trivial task. Particularly, we propose a novel Bi-directional Spatial Temporal Attention Fusion Generative Adversarial Network (STA-GAN) to learn both spatial and temporal information to generate egocentric video sequences from the exocentric view. The proposed STA-GAN consists of three parts: temporal branch, spatial branch, and attention fusion. First, the temporal and spatial branches generate a sequence of fake frames and their corresponding features. The fake frames are generated in both downstream and upstream directions for both temporal and spatial branches. Next, the generated four different fake frames and their corresponding features (spatial and temporal branches in two directions) are fed into a novel multi-generation attention fusion module to produce the final video sequence. Meanwhile, we also propose a novel temporal and spatial dual-discriminator for more robust network optimization. Extensive experiments on the Side2Ego and Top2Ego datasets show that the proposed STA-GAN significantly outperforms the existing methods. Gaowen Liu, Hao Tang 0005, Hugo Latapie, Jason J. Corso, Yan Yan 0002 |
ACM Multimedia | 4 |
| 2021 | Integrating Human Gaze into Attention for Egocentric Activity RecognitionabstractIt is well known that human gaze carries significant information about visual attention. However, there are three main difficulties in incorporating the gaze data in an attention mechanism of deep neural networks: (i) the gaze fixation points are likely to have measurement errors due to blinking and rapid eye movements; (ii) it is unclear when and how much the gaze data is correlated with visual attention; and (iii) gaze data is not always available in many real-world situations. In this work, we introduce an effective probabilistic approach to integrate human gaze into spatiotemporal attention for egocentric activity recognition. Specifically, we represent the locations of gaze fixation points as structured discrete latent variables to model their uncertainties. In addition, we model the distribution of gaze fixations using a variational method. The gaze distribution is learned during the training process so that the ground-truth annotations of gaze locations are no longer needed in testing situations since they are predicted from the learned gaze distribution. The predicted gaze locations are used to provide informative attentional cues to improve the recognition performance. Our method outperforms all the previous state-of-the-art approaches on EGTEA, which is a large-scale dataset for egocentric activity recognition provided with gaze measurements. We also perform an ablation study and qualitative analysis to demonstrate that our attention mechanism is effective. Kyle Min 0001, Jason J. Corso |
WACV | 2 |
| 2021 | HyperCon: Image-To-Video Model Transfer for Video-To-Video Translation TasksabstractVideo-to-video translation is more difficult than image-to-image translation due to the temporal consistency problem that, if unaddressed, leads to distracting flickering effects. Although video models designed from scratch produce temporally consistent results, training them to match the vast visual knowledge captured by image models requires an intractable number of videos. To combine the benefits of image and video models, we propose an image-to-video model transfer method called Hyperconsistency (HyperCon) that transforms any well-trained image model into a temporally consistent video model without fine-tuning. HyperCon works by translating a temporally interpolated video frame-wise and then aggregating over temporally localized windows on the interpolated video. It handles both masked and unmasked inputs, enabling support for even more video-to-video translation tasks than prior image-to-video model transfer techniques. We demonstrate HyperCon on video style transfer and inpainting, where it performs favorably compared to prior state-of-the-art methods without training on a single stylized or incomplete video. Our project website is available at ryanszeto.com/projects/hypercon. Ryan Szeto, Mostafa El-Khamy, Jason J. Corso |
WACV | 4 |
| 2020 | Unified Vision-Language Pre-Training for Image Captioning and VQAabstractThis paper presents a unified Vision-Language Pre-training (VLP) model. The model is unified in that (1) it can be fine-tuned for either vision-language generation (e.g., image captioning) or understanding (e.g., visual question answering) tasks, and (2) it uses a shared multi-layer transformer network for both encoding and decoding, which differs from many existing methods where the encoder and decoder are implemented using separate models. The unified VLP model is pre-trained on a large amount of image-text pairs using the unsupervised learning objectives of two tasks: bidirectional and sequence-to-sequence (seq2seq) masked vision-language prediction. The two tasks differ solely in what context the prediction conditions on. This is controlled by utilizing specific self-attention masks for the shared transformer network. To the best of our knowledge, VLP is the first reported model that achieves state-of-the-art results on both vision-language generation and understanding tasks, as disparate as image captioning and visual question answering, across three challenging benchmark datasets: COCO Captions, Flickr30k Captions, and VQA 2.0. The code and the pre-trained models are available at https://github.com/LuoweiZhou/VLP. Luowei Zhou, Hamid Palangi, Lei Zhang 0001, Houdong Hu, Jason J. Corso, Jianfeng Gao 0001 |
AAAI | 5 |
| 2020 | Rethinking Curriculum Learning with Incremental Labels and Adaptive Compensation
Madan Ravi Ganesh, Jason J. Corso |
BMVC | 2 |
| 2020 | Novel Object Viewpoint Estimation Through Reconstruction Alignment
Mohamed El Banani, Jason J. Corso, David F. Fouhey |
CVPR | 2 |
| 2020 | Learning Object Depth from Camera Motion and Video Object Segmentation
Brent A. Griffin, Jason J. Corso |
ECCV (7) | 2 |
| 2020 | Adversarial Background-Aware Loss for Weakly-Supervised Temporal Activity Localization
Kyle Min 0001, Jason J. Corso |
ECCV (14) | 2 |
| 2020 | MINT: Deep Network Compression via Mutual Information-based Neuron TrimmingabstractMost approaches to deep neural network compression via pruning either directly evaluate a filter's importance using its weights or optimize an alternative objective function with sparsity constraints. While these methods offer a useful way to approximate contributions from similar filters, they often either ignore the dependency between layers or solve a more difficult optimization objective than standard cross-entropy. Our method, Mutual Information-based Neuron Trimming (MINT), approaches deep compression via pruning by enforcing sparsity based on the strength of the dependency between filters of adjacent layers, across every pair of layers in the network. The dependency is calculated using conditional geometric mutual information which evaluates the amount of similar information exchanged between filters using a graph-based criterion. When pruning a network, we ensure that retained filters contribute the majority of the information towards succeeding layers which ensures high performance. Our novel approach is highly competitive with existing state-of-the-art compression-via-pruning methods on standard benchmarks for this task: MNIST, CIFAR-10, and ILSVRC2012, across a variety of network architectures despite using only a single retraining pass. Also, we discuss our observations of a common denominator between our pruning methodology's response to adversarial attacks and calibration statistics when compared to the original network. Madan Ravi Ganesh, Jason J. Corso, Salimeh Yasaei Sekeh |
ICPR | 2 |
| 2020 | Robot-Supervised Learning for Object SegmentationabstractTo be effective in unstructured and changing environments, robots must learn to recognize new objects. Deep learning has enabled rapid progress for object detection and segmentation in computer vision; however, this progress comes at the price of human annotators labeling many training examples. This paper addresses the problem of extending learning-based segmentation methods to robotics applications where annotated training data is not available. Our method enables pixelwise segmentation of grasped objects. We factor the problem of segmenting the object from the background into two sub-problems: (1) segmenting the robot manipulator and object from the background and (2) segmenting the object from the manipulator. We propose a kinematics-based foreground segmentation technique to solve (1). To solve (2), we train a self-recognition network that segments the robot manipulator. We train this network without human supervision, leveraging our foreground segmentation technique from (1) to label a training set of images containing the robot manipulator without a grasped object. We demonstrate experimentally that our method outperforms state-of-the-art adaptable in-hand object segmentation. We also show that a training set composed of automatically labelled images of grasped objects improves segmentation performance on a test set of images of the same objects in the environment. Victoria Florence, Jason J. Corso, Brent A. Griffin |
ICRA | 2 |
| 2020 | Video Object Segmentation-based Visual Servo Control and Object Depth Estimation on a Mobile RobotabstractTo be useful in everyday environments, robots must be able to identify and locate real-world objects. In recent years, video object segmentation has made significant progress on densely separating such objects from background in real and challenging videos. Building off of this progress, this paper addresses the problem of identifying generic objects and locating them in 3D using a mobile robot with an RGB camera. We achieve this by, first, introducing a video object segmentation-based approach to visual servo control and active perception and, second, developing a new Hadamard-Broyden update formulation. Our segmentation-based methods are simple but effective, and our update formulation lets a robot quickly learn the relationship between actuators and visual features without any camera calibration. We validate our approach in experiments by learning a variety of actuator-camera configurations on a mobile HSR robot, which subsequently identifies, locates, and grasps objects from the YCB dataset and tracks people and other dynamic articulated objects in real-time. Brent A. Griffin, Victoria Florence, Jason J. Corso |
WACV | 3 |
| 2020 | A Weakly Supervised Multi-task Ranking Framework for Actor-Action Semantic Segmentation
Yan Yan 0002, Chenliang Xu, Dawen Cai, Jason J. Corso |
Int. J. Comput. Vis. | 4 |
| 2020 | A Temporally-Aware Interpolation Network for Video Frame InpaintingabstractIn this work, we explore video frame inpainting, a task that lies at the intersection of general video inpainting, frame interpolation, and video prediction. Although our problem can be addressed by applying methods from other video interpolation or extrapolation tasks, doing so fails to leverage the additional context information that our problem provides. To this end, we devise a method specifically designed for video frame inpainting that is composed of two modules: a bidirectional video prediction module and a temporally-aware frame interpolation module. The prediction module makes two intermediate predictions of the missing frames, each conditioned on the preceding and following frames respectively, using a shared convolutional LSTM-based encoder-decoder. The interpolation module blends the intermediate predictions by using time information and hidden activations from the video prediction module to resolve disagreements between the predictions. Our experiments demonstrate that our approach produces smoother and more accurate results than state-of-the-art methods for general video inpainting, frame interpolation, and video prediction. Ryan Szeto, Ximeng Sun, Kunyi Lu, Jason J. Corso |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | Dynamic Graph Modules for Modeling Object-Object Interactions in Activity Recognition
Hao Huang 0003, Luowei Zhou, Jason J. Corso, Chenliang Xu |
BMVC | 4 |
| 2019 | BubbleNets: Learning to Select the Guidance Frame in Video Object Segmentation by Deep Sorting FramesabstractSemi-supervised video object segmentation has made significant progress on real and challenging videos in recent years. The current paradigm for segmentation methods and benchmark datasets is to segment objects in video provided a single annotation in the first frame. However, we find that segmentation performance across the entire video varies dramatically when selecting an alternative frame for annotation. This paper addresses the problem of learning to suggest the single best frame across the video for user annotation-this is, in fact, never the first frame of video. We achieve this by introducing BubbleNets, a novel deep sorting network that learns to select frames using a performance-based loss function that enables the conversion of expansive amounts of training examples from already existing datasets. Using BubbleNets, we are able to achieve an 11% relative improvement in segmentation performance on the DAVIS benchmark without any changes to the underlying method of segmentation. Brent A. Griffin, Jason J. Corso |
CVPR | 2 |
| 2019 | Multi-Channel Attention Selection GAN With Cascaded Semantic Guidance for Cross-View Image TranslationabstractCross-view image translation is challenging because it involves images with drastically different views and severe deformation. In this paper, we propose a novel approach named Multi-Channel Attention SelectionGAN (SelectionGAN) that makes it possible to generate images of natural scenes in arbitrary viewpoints, based on an image of the scene and a novel semantic map. The proposed SelectionGAN explicitly utilizes the semantic information and consists of two stages. In the first stage, the condition image and the target semantic map are fed into a cycled semantic-guided generation network to produce initial coarse results. In the second stage, we refine the initial results by using a multi-channel attention selection mechanism. Moreover, uncertainty maps automatically learned from attentions are used to guide the pixel loss for better network optimization. Extensive experiments on Dayton, CVUSA and Ego2Top datasets show that our model is able to generate significantly better results than the state-of-the-art methods. The source code, data and trained models are available at https://github.com/Ha0Tang/SelectionGAN. Hao Tang 0005, Dan Xu 0002, Nicu Sebe, Yanzhi Wang 0001, Jason J. Corso, Yan Yan 0002 |
CVPR | 5 |
| 2019 | Grounded Video DescriptionabstractVideo description is one of the most challenging problems in vision and language understanding due to the large variability both on the video and language side. Models, hence, typically shortcut the difficulty in recognition and generate plausible sentences that are based on priors but are not necessarily grounded in the video. In this work, we explicitly link the sentence to the evidence in the video by annotating each noun phrase in a sentence with the corresponding bounding box in one of the frames of a video. Our dataset, ActivityNet-Entities, augments the challenging ActivityNet Captions dataset with 158k bounding box annotations, each grounding a noun phrase. This allows training video description models with this data, and importantly, evaluate how grounded or "true" such model are to the video they describe. To generate grounded captions, we propose a novel video description model which is able to exploit these bounding box annotations. We demonstrate the effectiveness of our model on our dataset, but also show how it can be applied to image description on the Flickr30k Entities dataset. We achieve state-of-the-art performance on video description, video paragraph description, and image description and demonstrate our generated sentences are better grounded in the video. Luowei Zhou, Yannis Kalantidis, Xinlei Chen, Jason J. Corso, Marcus Rohrbach |
CVPR | 4 |
| 2019 | Attribute-Guided Sketch GenerationabstractFacial attributes are important since they provide a detailed description and determine the visual appearance of human faces. In this paper, we aim at converting a face image to a sketch while simultaneously generating facial attributes. To this end, we propose a novel Attribute-Guided Sketch Generative Adversarial Network (ASGAN) which is an end-to-end framework and contains two pairs of generators and discriminators, one of which is used to generate faces with attributes while the other one is employed for image-to-sketch translation. The two generators form a W-shaped network (W-net) and they are trained jointly with a weight-sharing constraint. Additionally, we also propose two novel discriminators, the residual one focusing on attribute generation and the triplex one helping to generate realistic looking sketches. To validate our model, we have created a new large dataset with 8,804 images, named the Attribute Face Photo & Sketch (AFPS) dataset which is the first dataset containing attributes associated to face sketch images. The experimental results demonstrate that the proposed network (i) generates more photo-realistic faces with sharper facial attributes than baselines and (ii) has good generalization capability on different generative tasks. Hao Tang 0005, Xinya Chen, Wei Wang 0108, Dan Xu 0002, Jason J. Corso, Nicu Sebe, Yan Yan 0002 |
FG | 5 |
| 2019 | TASED-Net: Temporally-Aggregating Spatial Encoder-Decoder Network for Video Saliency DetectionabstractTASED-Net is a 3D fully-convolutional network architecture for video saliency detection. It consists of two building blocks: first, the encoder network extracts low-resolution spatiotemporal features from an input clip of several consecutive frames, and then the following prediction network decodes the encoded features spatially while aggregating all the temporal information. As a result, a single prediction map is produced from an input clip of multiple frames. Frame-wise saliency maps can be predicted by applying TASED-Net in a sliding-window fashion to a video. The proposed approach assumes that the saliency map of any frame can be predicted by considering a limited number of past frames. The results of our extensive experiments on video saliency detection validate this assumption and demonstrate that our fully-convolutional model with temporal aggregation method is effective. TASED-Net significantly outperforms previous state-of-the-art approaches on all three major large-scale datasets of video saliency detection: DHF1K, Hollywood2, and UCFSports. After analyzing the results qualitatively, we observe that our model is especially better at attending to salient moving objects. Kyle Min 0001, Jason J. Corso |
ICCV | 2 |
| 2019 | Popup: reconstructing 3D video using particle filtering to aggregate crowd responsesabstractCollecting a sufficient amount of 3D training data for autonomous vehicles to handle rare, but critical, traffic events (e.g., collisions) may take decades of deployment. Abundant video data of such events from municipal traffic cameras and video sharing sites (e.g., YouTube) could provide a potential alternative, but generating realistic training data in the form of 3D video reconstructions is a challenging task beyond the current capabilities of computer vision. Crowdsourcing the annotation of necessary information could bridge this gap, but the level of accuracy required to obtain usable reconstructions makes this task nearly impossible for non-experts. In this paper, we propose a novel hybrid intelligence method that combines annotations from workers viewing different instances (video frames) of the same target (3D object), and uses particle filtering to aggregate responses. Our approach can leveraging temporal dependencies between video frames, enabling higher quality through more aggressive filtering. The proposed method results in a 33% reduction in the relative error of position estimation compared to a state-of-the-art baseline. Moreover, our method enables skipping (self-filtering) challenging annotations, reducing the total annotation time for hard-to-annotate frames by 16%. Our approach provides a generalizable means of aggregating more accurate crowd responses in settings where annotation is especially challenging or error-prone. Jean Y. Song, Stephan J. Lemmer, Michael Xieyang Liu, Shiyan Yan, Juho Kim 0001, Jason J. Corso, Walter S. Lasecki |
IUI | 6 |
| 2019 | Tukey-Inspired Video Object SegmentationabstractWe investigate the problem of strictly unsupervised video object segmentation, i.e., the separation of a primary object from background in video without a user-provided object mask or any training on an annotated dataset. We find foreground objects in low-level vision data using a John Tukey-inspired measure of "outlierness." This Tukey-inspired measure also estimates the reliability of each data source as video characteristics change (e.g., a camera starts moving). The proposed method achieves state-of-the-art results for strictly unsupervised video object segmentation on the challenging DAVIS dataset. Finally, we use a variant of the Tukey-inspired measure to combine the output of multiple segmentation methods, including those using supervision during training, runtime, or both. This collectively more robust method of segmentation improves the Jaccard measure of its constituent methods by as much as 28%. Brent A. Griffin, Jason J. Corso |
WACV | 2 |
| 2018 | Towards Automatic Learning of Procedures From Web Instructional VideosabstractThe potential for agents, whether embodied or software, to learn by observing other agents performing procedures involving objects and actions is rich. Current research on automatic procedure learning heavily relies on action labels or video subtitles, even during the evaluation phase, which makes them infeasible in real-world scenarios. This leads to our question: can the human-consensus structure of a procedure be learned from a large set of long, unconstrained videos (e.g., instructional videos from YouTube) with only visual evidence? To answer this question, we introduce the problem of procedure segmentation---to segment a video procedure into category-independent procedure segments. Given that no large-scale dataset is available for this problem, we collect a large-scale procedure segmentation dataset with procedure segments temporally localized and described; we use cooking videos and name the dataset YouCook2. We propose a segment-level recurrent network for generating procedure segments by modeling the dependencies across segments. The generated segments can be used as pre-processing for other tasks, such as dense video captioning and event parsing. We show in our experiments that the proposed model outperforms competitive baselines in procedure segmentation. Luowei Zhou, Chenliang Xu, Jason J. Corso |
AAAI | 3 |
| 2018 | A Temporally-Aware Interpolation Network for Video Frame Inpainting
Ximeng Sun, Ryan Szeto, Jason J. Corso |
ACCV (3) | 3 |
| 2018 | Weakly-Supervised Video Object Grounding from Text by Loss Weighting and Object Interaction
Luowei Zhou, Nathan Louis, Jason J. Corso |
BMVC | 3 |
| 2018 | End-to-End Dense Video Captioning With Masked TransformerabstractDense video captioning aims to generate text descriptions for all events in an untrimmed video. This involves both detecting and describing events. Therefore, all previous methods on dense video captioning tackle this problem by building two models, i.e. an event proposal and a captioning model, for these two sub-problems. The models are either trained separately or in alternation. This prevents direct influence of the language description to the event proposal, which is important for generating accurate descriptions. To address this problem, we propose an end-to-end transformer model for dense video captioning. The encoder encodes the video into appropriate representations. The proposal decoder decodes from the encoding with different anchors to form video event proposals. The captioning decoder employs a masking network to restrict its attention to the proposal event over the encoding feature. This masking network converts the event proposal to a differentiable mask, which ensures the consistency between the proposal and captioning during training. In addition, our model employs a self-attention mechanism, which enables the use of efficient non-recurrent structure during encoding and leads to performance improvements. We demonstrate the effectiveness of this end-to-end model on ActivityNet Captions and YouCookII datasets, where we achieved 10.12 and 6.58 METEOR score, respectively. Luowei Zhou, Yingbo Zhou 0002, Jason J. Corso, Richard Socher, Caiming Xiong |
CVPR | 3 |
| 2018 | The Wrong Tool for Inference - A Critical View of Gaussian Graphical Models
Kevin R. Keane, Jason J. Corso |
ICPRAM | 2 |
| 2018 | Machine learning for big visual analysis
Jun Yu 0002, Xue Mei, Fatih Porikli, Jason J. Corso |
Mach. Vis. Appl. | 4 |
| 2018 | Learning Compositional Sparse Bimodal ModelsabstractVarious perceptual domains have underlying compositional semantics that are rarely captured in current models. We suspect this is because directly learning the compositional structure has evaded these models. Yet, the compositional structure of a given domain can be grounded in a separate domain thereby simplifying its learning. To that end, we propose a new approach to modeling bimodal perceptual domains that explicitly relates distinct projections across each modality and then jointly learns a bimodal sparse representation. The resulting model enables compositionality across these distinct projections and hence can generalize to unobserved percepts spanned by this compositional basis. For example, our model can be trained on red triangles and blue squares; yet, implicitly will also have learned red squares and blue triangles. The structure of the projections and hence the compositional basis is learned automatically; no assumption is made on the ordering of the compositional elements in either modality. Although our modeling paradigm is general, we explicitly focus on a tabletop building-blocks setting. To test our model, we have acquired a new bimodal dataset comprising images and spoken utterances of colored shapes (blocks) in the tabletop setting. Our experiments demonstrate the benefits of explicitly leveraging compositionality in both quantitative and human evaluation studies. Suren Kumar, Vikas Dhiman, Parker A. Koch, Jason J. Corso |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | Weakly Supervised Actor-Action Segmentation via Robust Multi-task RankingabstractFine-grained activity understanding in videos has attracted considerable recent attention with a shift from action classification to detailed actor and action understanding that provides compelling results for perceptual needs of cutting-edge autonomous systems. However, current methods for detailed understanding of actor and action have significant limitations: they require large amounts of finely labeled data, and they fail to capture any internal relationship among actors and actions. To address these issues, in this paper, we propose a novel, robust multi-task ranking model for weakly supervised actor-action segmentation where only video-level tags are given for training samples. Our model is able to share useful information among different actors and actions while learning a ranking matrix to select representative supervoxels for actors and actions respectively. Final segmentation results are generated by a conditional random field that considers various ranking scores for different video parts. Extensive experimental results on the Actor-Action Dataset (A2D) demonstrate that the proposed approach outperforms the state-of-the-art weakly supervised methods and performs as well as the top-performing fully supervised method. Yan Yan 0002, Chenliang Xu, Dawen Cai, Jason J. Corso |
CVPR | 4 |
| 2017 | Click Here: Human-Localized Keypoints as Guidance for Viewpoint EstimationabstractWe motivate and address a human-in-the-loop variant of the monocular viewpoint estimation task in which the location and class of one semantic object keypoint is available at test time. In order to leverage the keypoint information, we devise a Convolutional Neural Network called Click-Here CNN (CH-CNN) that integrates the keypoint information with activations from the layers that process the image. It transforms the keypoint information into a 2D map that can be used to weigh features from certain parts of the image more heavily. The weighted sum of these spatial features is combined with global image features to provide relevant information to the prediction layers. To train our network, we collect a novel dataset of 3D keypoint annotations on thousands of CAD models, and synthetically render millions of images with 2D keypoint information. On test instances from PASCAL 3D+, our model achieves a mean class accuracy of 90.7%, whereas the state-of-the-art baseline only obtains 85.7% mean class accuracy, justifying our argument for human-in-the-loop inference. Ryan Szeto, Jason J. Corso |
ICCV | 2 |
| 2017 | Active Clustering with Model-Based Uncertainty ReductionabstractSemi-supervised clustering seeks to augment traditional clustering methods by incorporating side information provided via human expertise in order to increase the semantic meaningfulness of the resulting clusters. However, most current methods are passive in the sense that the side information is provided beforehand and selected randomly. This may require a large number of constraints, some of which could be redundant, unnecessary, or even detrimental to the clustering results. Thus in order to scale such semi-supervised algorithms to larger problems it is desirable to pursue an active clustering method-i.e., an algorithm that maximizes the effectiveness of the available human labor by only requesting human input where it will have the greatest impact. Here, we propose a novel online framework for active semi-supervised spectral clustering that selects pairwise constraints as clustering proceeds, based on the principle of uncertainty reduction. Using a first-order Taylor expansion, we decompose the expected uncertainty reduction problem into a gradient and a step-scale, computed via an application of matrix perturbation theory and cluster-assignment entropy, respectively. The resulting model is used to estimate the uncertainty reduction potential of each sample in the dataset. We then present the human user with pairwise queries with respect to only the best candidate sample. We evaluate our method using three different image datasets (faces, leaves and dogs), a set of common UCI machine learning datasets and a gene dataset. The results validate our decomposition formulation and show that our method is consistently superior to existing state-of-the-art techniques, as well as being robust to noise and to unknown numbers of clusters. Caiming Xiong, David M. Johnson 0001, Jason J. Corso |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | Joint occlusion boundary detection and figure/ground assignment by extracting common-fate fragments in a back-projection scheme
Jason J. Corso |
Pattern Recognit. | 2 |
| 2017 | Editorial for special section of video analytics with deep learning
Tao Mei 0001, Jason J. Corso, Jiebo Luo 0001 |
Pattern Recognit. | 2 |
| 2017 | Detection and Localization of Robotic Tools in Robot-Assisted Surgery Videos Using Deep Neural Networks for Region Proposal and DetectionabstractVideo understanding of robot-assisted surgery (RAS) videos is an active research area. Modeling the gestures and skill level of surgeons presents an interesting problem. The insights drawn may be applied in effective skill acquisition, objective skill assessment, real-time feedback, and human-robot collaborative surgeries. We propose a solution to the tool detection and localization open problem in RAS video understanding, using a strictly computer vision approach and the recent advances of deep learning. We propose an architecture using multimodal convolutional neural networks for fast detection and localization of tools in RAS videos. To the best of our knowledge, this approach will be the first to incorporate deep neural networks for tool detection and localization in RAS videos. Our architecture applies a region proposal network (RPN) and a multimodal two stream convolutional network for object detection to jointly predict objectness and localization on a fusion of image and temporal motion cues. Our results with an average precision of 91% and a mean computation time of 0.1 s per test frame detection indicate that our study is superior to conventionally used methods for medical imaging while also emphasizing the benefits of using RPN for precision and efficiency. We also introduce a new data set, ATLAS Dione, for RAS video understanding. Our data set provides video data of ten surgeons from Roswell Park Cancer Institute, Buffalo, NY, USA, performing six different surgical tasks on the daVinci Surgical System (dVSS) with annotations of robotic tools per frame. Duygu Sarikaya, Jason J. Corso, Khurshid A. Guru |
IEEE Trans. Medical Imaging | 2 |
| 2017 | Dancelets Mining for Video Recommendation Based on Dance StylesabstractDance is a unique and meaningful type of human expression, composed of abundant and various action elements. However, existing methods based on associated texts and spatial visual features have difficulty capturing the highly articulated motion patterns. To overcome this limitation, we propose to take advantage of the intrinsic motion information in dance videos to solve the video recommendation problem. We present a novel system that recommends dance videos based on a mid-level action representation, termed Dancelets. The Dancelets are used to bridge the semantic gap between video content and high-level concept, dance style, which plays a significant role in characterizing different types of dances. The proposed method executes automatic mining of dancelets with a concatenation of normalized cut clustering and linear discriminant analysis. This ensures that the discovered dancelets are both representative and discriminative. Additionally, to exploit the motion cues in videos, we employ motion boundaries as saliency priors to generate volumes of interest and extract C3D features to capture spatiotemporal information from the mid-level patches. Extensive experiments validated on our proposed large dance dataset, HIT Dances dataset, demonstrate the effectiveness of the proposed methods for dance style-based video recommendation. Tingting Han 0003, Hongxun Yao, Chenliang Xu, Xiaoshuai Sun, Yanhao Zhang 0001, Jason J. Corso |
IEEE Trans. Multim. | 6 |
| 2016 | A Continuous Occlusion Model for Road Scene UnderstandingabstractWe present a physically interpretable, continuous threedimensional (3D) model for handling occlusions with applications to road scene understanding. We probabilistically assign each point in space to an object with a theoretical modeling of the reflection and transmission probabilities for the corresponding camera ray. Our modeling is unified in handling occlusions across a variety of scenarios, such as associating structure from motion (SFM) point tracks with potentially occluding objects or modeling object detection scores in applications such as 3D localization. For point track association, our model uniformly handles static and dynamic objects, which is an advantage over motion segmentation approaches traditionally used in multibody SFM. Detailed experiments on the KITTI raw dataset show the superiority of the proposed method over both state-of-the-art motion segmentation and a baseline that heuristically uses detection bounding boxes for resolving occlusions. We also demonstrate how our continuous occlusion model may be applied to the task of 3D localization in road scenes. Vikas Dhiman, Quoc-Huy Tran, Jason J. Corso, Manmohan Krishna Chandraker |
CVPR | 3 |
| 2016 | Actor-Action Semantic Segmentation with Grouping Process ModelsabstractActor-action semantic segmentation made an important step toward advanced video understanding: what action is happening, who is performing the action, and where is the action happening in space-time. Current methods based on layered CRFs for this problem are local and unable to capture the long-ranging interactions of video parts. We propose a new model that combines the labeling CRF with a supervoxel hierarchy, where supervoxels at various scales provide cues for possible groupings of nodes in the CRF to encourage adaptive and long-ranging interactions. The new model defines a dynamic and continuous process of information exchange: the CRF influences what supervoxels in the hierarchy are active, and these active supervoxels, in turn, affect the connectivities in the CRF, we hence call it a grouping process model. By further incorporating the video-level recognition, the proposed method achieves a large margin of 60% relative improvement over the state of the art on the recent A2D large-scale video labeling dataset, which demonstrates the effectiveness of our modeling. Chenliang Xu, Jason J. Corso |
CVPR | 2 |
| 2016 | LIBSVX: A Supervoxel Library and Benchmark for Early Video Processing
Chenliang Xu, Jason J. Corso |
Int. J. Comput. Vis. | 2 |
| 2016 | Compositional models and Structured learning for visual recognition
Liang Lin 0004, Jason J. Corso, Wangmeng Zuo, David Zhang 0001, Benjamin Z. Yao |
Pattern Recognit. | 2 |
| 2016 | Semi-Supervised Nonlinear Distance Metric Learning via Forests of Max-Margin Cluster HierarchiesabstractMetric learning is a key problem for many data mining and machine learning applications, and has long been dominated by Mahalanobis methods. Recent advances in nonlinear metric learning have demonstrated the potential power of non-Mahalanobis distance functions, particularly tree-based functions. We propose a novel nonlinear metric learning method that uses an iterative, hierarchical variant of semi-supervised max-margin clustering to construct a forest of cluster hierarchies, where each individual hierarchy can be interpreted as a weak metric over the data. By introducing randomness during hierarchy training and combining the output of many of the resulting semi-random weak hierarchy metrics, we can obtain a powerful and robust nonlinear metric model. This method has two primary contributions: first, it is semi-supervised, incorporating information from both constrained and unconstrained points. Second, we take a relaxed approach to constraint satisfaction, allowing the method to satisfy different subsets of the constraints at different levels of the hierarchy rather than attempting to simultaneously satisfy all of them. This leads to a more robust learning algorithm. We compare our method to a number of state-of-the-art benchmarks on $k$-nearest neighbor classification, large-scale image retrieval and semi-supervised clustering problems, and find that our algorithm yields results comparable or superior to the state-of-the-art. David M. Johnson 0001, Caiming Xiong, Jason J. Corso |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2015 | Jointly Modeling Deep Video and Compositional Text to Bridge Vision and Language in a Unified FrameworkabstractRecently, joint video-language modeling has been attracting more and more attention. However, most existing approaches focus on exploring the language model upon on a fixed visual model. In this paper, we propose a unified framework that jointly models video and the corresponding text sentences. The framework consists of three parts: a compositional semantics language model, a deep video model and a joint embedding model. In our language model, we propose a dependency-tree structure model that embeds sentence into a continuous vector space, which preserves visually grounded meanings and word order. In the visual model, we leverage deep neural networks to capture essential semantic information from videos. In the joint embedding model, we minimize the distance of the outputs of the deep video model and compositional language model in the joint space, and update these two models jointly. Based on these three parts, our system is able to accomplish three tasks: 1) natural language generation, and 2) video retrieval and 3) language retrieval. In the experiments, the results show our approach outperforms SVM, CRF and CCA baselines in predicting Subject-Verb-Object triplet and natural sentence generation, and is better than CCA in video retrieval and language retrieval tasks. Ran Xu 0001, Caiming Xiong, Wei Chen 0134, Jason J. Corso |
AAAI | 4 |
| 2015 | Human action segmentation with hierarchical supervoxel consistencyabstractDetailed analysis of human action, such as action classification, detection and localization has received increasing attention from the community; datasets like JHMDB have made it plausible to conduct studies analyzing the impact that such deeper information has on the greater action understanding problem. However, detailed automatic segmentation of human action has comparatively been unexplored. In this paper, we take a step in that direction and propose a hierarchical MRF model to bridge low-level video fragments with high-level human motion and appearance; novel higher-order potentials connect different levels of the supervoxel hierarchy to enforce the consistency of the human segmentation by pulling from different segment-scales. Our single layer model significantly outperforms the current state-of-the-art on actionness, and our full model improves upon the single layer baselines in action segmentation. Jiasen Lu, Ran Xu 0001, Jason J. Corso |
CVPR | 3 |
| 2015 | Can humans fly? Action understanding with multiple classes of actorsabstractCan humans fly? Emphatically no. Can cars eat? Again, absolutely not. Yet, these absurd inferences result from the current disregard for particular types of actors in action understanding. There is no work we know of on simultaneously inferring actors and actions in the video, not to mention a dataset to experiment with. Our paper hence marks the first effort in the computer vision community to jointly consider various types of actors undergoing various actions. To start with the problem, we collect a dataset of 3782 videos from YouTube and label both pixel-level actors and actions in each video. We formulate the general actor-action understanding problem and instantiate it at various granularities: both video-level single- and multiple-label actor-action recognition and pixel-level actor-action semantic segmentation. Our experiments demonstrate that inference jointly over actors and actions outperforms inference independently over them, and hence concludes our argument of the value of explicit consideration of various actors in comprehensive action understanding. Chenliang Xu, Shao-Hang Hsieh, Caiming Xiong, Jason J. Corso |
CVPR | 4 |
| 2015 | Action Detection by Implicit Intentional Motion ClusteringabstractExplicitly using human detection and pose estimation has found limited success in action recognition problems. This may be due to the complexity in the articulated motion human exhibit. Yet, we know that action requires an actor and intention. This paper hence seeks to understand the spatiotemporal properties of intentional movement and how to capture such intentional movement without relying on challenging human detection and tracking. We conduct a quantitative analysis of intentional movement, and our findings motivate a new approach for implicit intentional movement extraction that is based on spatiotemporal trajectory clustering by leveraging the properties of intentional movement. The intentional movement clusters are then used as action proposals for detection. Our results on three action detection benchmarks indicate the relevance of focusing on intentional movement for action detection, our method significantly outperforms the state of the art on the challenging MSR-II multi-action video benchmark. Wei Chen 0134, Jason J. Corso |
ICCV | 2 |
| 2015 | The Multimodal Brain Tumor Image Segmentation Benchmark (BRATS)abstractIn this paper we report the set-up and results of the Multimodal Brain Tumor Image Segmentation Benchmark (BRATS) organized in conjunction with the MICCAI 2012 and 2013 conferences. Twenty state-of-the-art tumor segmentation algorithms were applied to a set of 65 multi-contrast MR scans of low- and high-grade glioma patients-manually annotated by up to four raters-and to 65 comparable scans generated using tumor image simulation software. Quantitative evaluations revealed considerable disagreement between the human raters in segmenting various tumor sub-regions (Dice scores in the range 74%-85%), illustrating the difficulty of this task. We found that different algorithms worked best for different sub-regions (reaching performance comparable to human inter-rater variability), but that no single algorithm ranked in the top for all sub-regions simultaneously. Fusing several good algorithms using a hierarchical majority vote yielded segmentations that consistently ranked above all individual algorithms, indicating remaining opportunities for further methodological improvements. The BRATS image data and manual annotations continue to be publicly available through an online evaluation system as an ongoing benchmarking resource. Bjoern Menze, András Jakab, Stefan Bauer, Jayashree Kalpathy-Cramer, Keyvan Farahani, Justin S. Kirby, Yuliya Burren, Nicole Porz, Johannes Slotboom, Roland Wiest, Levente Lanczi, Elizabeth R. Gerstner, Marc-André Weber, Tal Arbel, Brian B. Avants, Nicholas Ayache, Patricia Buendia, D. Louis Collins, Nicolas Cordier, Jason J. Corso, Antonio Criminisi, Tilak Das, Hervé Delingette, Çagatay Demiralp, Christopher R. Durst, Michel Dojat, Senan Doyle, Joana Festa, Florence Forbes, Ezequiel Geremia, Ben Glocker, Polina Golland, Xiaotao Guo, Andac Hamamci, Khan M. Iftekharuddin, Raj Jena, Nigel M. John, Ender Konukoglu, Danial Lashkari, José Antonio Mariz, Raphael Meier, Sérgio Pereira, Doina Precup, Stephen J. Price, Tammy Riklin-Raviv, Syed M. S. Reza, Michael T. Ryan, Duygu Sarikaya, Lawrence H. Schwartz, Hoo-Chang Shin, Jamie Shotton, Carlos A. Silva 0002, Nuno J. Sousa, Nagesh K. Subbanna, Gábor Székely, Thomas J. Taylor, Owen M. Thomas, Nicholas J. Tustison, Gozde Unal, Flor Vasseur, Max Wintermark, Dong Hye Ye, Liang Zhao 0018, Binsheng Zhao, Darko Zikic, Marcel Prastawa, Mauricio Reyes 0001, Koenraad Van Leemput |
IEEE Trans. Medical Imaging | 20 |
| 2014 | Learning Compositional Sparse Models of Bimodal PerceptsabstractVarious perceptual domains have underlying compositional semantics that are rarely captured in current models. We suspect this is because directly learning the compositional structure has evaded these models. Yet, the compositional structure of a given domain can be grounded in a separate domain thereby simplifying its learning. To that end, we propose a new approach to modeling bimodal percepts that explicitly relates distinct projections across each modality and then jointly learns a bimodal sparse representation. The resulting model enables compositionality across these distinct projections and hence can generalize to unobserved percepts spanned by this compositional basis. For example, our model can be trained on 'red triangles' and 'blue squares'; yet, implicitly will also have learned 'red squares' and 'blue triangles'. The structure of the projections and hence the compositional basis is learned automatically for a given language model. To test our model, we have acquired a new bimodal dataset comprising images and spoken utterances of colored shapes in a tabletop setup. Our experiments demonstrate the benefits of explicitly leveraging compositionality in both quantitative and human evaluation studies. Suren Kumar, Vikas Dhiman, Jason J. Corso |
AAAI | 3 |
| 2014 | Latent Domains Modeling for Visual Domain AdaptationabstractTo improve robustness to significant mismatches between source domain and target domain - arising from changes such as illumination, pose and image quality - domain adaptation is increasingly popular in computer vision. But most of methods assume that the source data is from single domain, or that multi-domain datasets provide the domain label for training instances. In practice, most datasets are mixtures of multiple latent domains, and difficult to manually provide the domain label of each data point. In this paper, we propose a model that automatically discovers latent domains in visual datasets. We first assume the visual images are sampled from multiple manifolds, each of which represents different domain, and which are represented by different subspaces. Using the neighborhood structure estimated from images belonging to the same category, we approximate the local linear invariant subspace for each image based on its local structure, eliminating the category-specific elements of the feature. Based on the effectiveness of this representation, we then propose a squared-loss mutual information based clustering model with category distribution prior in each domain to infer the domain assignment for images. In experiment, we test our approach on two common image datasets, the results show that our method outperforms the existing state-of-the-art methods, and also show the superiority of multiple latent domain discovery. Caiming Xiong, Scott McCloskey, Shao-Hang Hsieh, Jason J. Corso |
AAAI | 4 |
| 2014 | Actionness Ranking with Lattice Conditional Ordinal Random FieldsabstractAction analysis in image and video has been attracting more and more attention in computer vision. Recognizing specific actions in video clips has been the main focus. We move in a new, more general direction in this paper and ask the critical fundamental question: what is action, how is action different from motion, and in a given image or video where is the action? We study the philosophical and visual characteristics of action, which lead us to define actionness: intentional bodily movement of biological agents (people, animals). To solve the general problem, we propose the lattice conditional ordinal random field model that incorporates local evidence as well as neighboring order agreement. We implement the new model in the continuous domain and apply it to scoring actionness in both image and video datasets. Our experiments demonstrate not only that our new model can outperform the popular ranking SVM but also that indeed action is distinct from motion. Wei Chen 0134, Caiming Xiong, Ran Xu 0001, Jason J. Corso |
CVPR | 4 |
| 2014 | Seeing is Worse than Believing: Reading People's Minds Better than Computer-Vision Methods Recognize Actions
Andrei Barbu, Daniel Paul Barrett, Wei Chen 0134, N. Siddharth 0001, Caiming Xiong, Jason J. Corso, Christiane Fellbaum, Catherine Hanson, Stephen Jose Hanson, Sébastien Hélie, Evguenia Malaia, Barak A. Pearlmutter, Jeffrey Mark Siskind, Thomas M. Talavage, Ronnie B. Wilbur |
ECCV (5) | 6 |
| 2014 | Systemic test and evaluation of a hard+soft information fusion framework: Challenges and current approaches
Geoff A. Gross, Ketan Date, Daniel R. Schlegel, Jason J. Corso, James Llinas, Rakesh Nagi, Stuart C. Shapiro |
FUSION | 4 |
| 2014 | Modern MAP inference methods for accurate and fast occupancy grid mapping on higher order factor graphsabstractUsing the inverse sensor model has been popular in occupancy grid mapping. However, it is widely known that applying the inverse sensor model to mapping requires certain assumptions that are not necessarily true. Even the works that use forward sensor models have relied on methods like expectation maximization or Gibbs sampling which have been succeeded by more effective methods of maximum a posteriori (MAP) inference over graphical models. In this paper, we propose the use of modern MAP inference methods along with the forward sensor model. Our implementation and experimental results demonstrate that these modern inference methods deliver more accurate maps more efficiently than previously used methods. Vikas Dhiman, Abhijit Kundu, Frank Dellaert, Jason J. Corso |
ICRA | 4 |
| 2014 | Surgical tool attributes from monocular videoabstractHD Video from the (monocular or binocular) endoscopic camera provides a rich real-time sensing channel from surgical site to the surgeon console in various Minimally Invasive Surgery (MIS) procedures. However, a real-time framework for video understanding would be critical for tapping into the rich information-content provided by the non-invasive and well-established digital endoscopic video-streaming modality. While contemporary research focuses on enhancing aspects such as tool-tracking within the challenging visual scenes, we consider the associated problem of using that rich (but often compromised) streaming visual data to discover the underlying semantic attributes of the tools. Directly analyzing the surgical videos to extract more realistic attributes online can aid in the decision-making and feedback aspects. We propose a novel probabilistic attribute labelling framework with Bayesian filtering to identify associated semantics (open/closed, stained with blood etc.) to ultimately give semantic feedback to the surgeon. Our robust video-understanding framework overcomes many of the challenges (tissue deformations, image specularities, clutter, tool-occlusion due to blood and/or organs) under realistic in-vivo surgical conditions. Specifically, this manuscript performs rigorous experimental analysis of the resulting method with varying parameters and different visual features on a data-corpus consisting of real surgical procedures performed on patients with da Vinci Surgical System [9]. Suren Kumar, Madusudanan Sathia Narayanan, Pankaj Singhal, Jason J. Corso, Venkat N. Krovi |
ICRA | 4 |
| 2014 | Adaptive Quantization for Hashing: An Information-Based Approach to Learning Binary CodesabstractLarge-scale data mining and retrieval applications have increasingly turned to compact binary data representations as a way to achieve both fast queries and efficient data storage; many algorithms have been proposed for learning effective binary encodings. Most of these algorithms focus on learning a set of projection hyperplanes for the data and simply binarizing the result from each hyperplane, but this neglects the fact that informativeness may not be uniformly distributed across the projections. In this paper, we address this issue by proposing a novel adaptive quantization (AQ) strategy that adaptively assigns varying numbers of bits to different hyperplanes based on their information content. Our method provides an information-based schema that preserves the neighborhood structure of data points, and we jointly find the globally optimal bit-allocation for all hyperplanes. In our experiments, we compare with state-of-the-art methods on four large-scale datasets and find that our adaptive quantization approach significantly improves on traditional hashing methods. Caiming Xiong, Wei Chen 0134, Gang Chen 0032, David M. Johnson 0001, Jason J. Corso |
SDM | 5 |
| 2014 | Multimedia event detection with multimodal feature fusion and temporal concept localization
Sangmin Oh, Scott McCloskey, Ilseo Kim, Arash Vahdat, Kevin J. Cannons, Hossein Hajimirsadeghi, Greg Mori, A. G. Amitha Perera, Megha Pandey, Jason J. Corso |
Mach. Vis. Appl. | 10 |
| 2013 | A Thousand Frames in Just a Few Words: Lingual Description of Videos through Latent Topics and Sparse Object StitchingabstractThe problem of describing images through natural language has gained importance in the computer vision community. Solutions to image description have either focused on a top-down approach of generating language through combinations of object detections and language models or bottom-up propagation of keyword tags from training images to test images through probabilistic or nearest neighbor techniques. In contrast, describing videos with natural language is a less studied problem. In this paper, we combine ideas from the bottom-up and top-down approaches to image description and propose a method for video description that captures the most relevant contents of a video in a natural language description. We propose a hybrid system consisting of a low level multimodal latent topic model for initial keyword annotation, a middle level of concept detectors and a high level module to produce final lingual descriptions. We compare the results of our system to human descriptions in both short and long forms on two datasets, and demonstrate that final system output has greater agreement with the human descriptions than any single level. Pradipto Das, Chenliang Xu, Richard F. Doell, Jason J. Corso |
CVPR | 4 |
| 2013 | Flattening Supervoxel Hierarchies by the Uniform Entropy SliceabstractSupervoxel hierarchies provide a rich multiscale decomposition of a given video suitable for subsequent processing in video analysis. The hierarchies are typically computed by an unsupervised process that is susceptible to under-segmentation at coarse levels and over-segmentation at fine levels, which make it a challenge to adopt the hierarchies for later use. In this paper, we propose the first method to overcome this limitation and flatten the hierarchy into a single segmentation. Our method, called the uniform entropy slice, seeks a selection of supervoxels that balances the relative level of information in the selected supervoxels based on some post hoc feature criterion such as object-ness. For example, with this criterion, in regions nearby objects, our method prefers finer supervoxels to capture the local details, but in regions away from any objects we prefer coarser supervoxels. We formulate the uniform entropy slice as a binary quadratic program and implement four different feature criteria, both unsupervised and supervised, to drive the flattening. Although we apply it only to supervoxel hierarchies in this paper, our method is generally applicable to segmentation tree hierarchies. Our experiments demonstrate both strong qualitative performance and superior quantitative performance to state of the art baselines on benchmark internet videos. Chenliang Xu, Spencer Whitt, Jason J. Corso |
ICCV | 3 |
| 2013 | Ascending stairway modeling from dense depth imagery for traversability analysisabstractLocalization and modeling of stairways by mobile robots can enable multi-floor exploration for those platforms capable of stair traversal. Existing approaches focus on either stairway detection or traversal, but do not address these problems in the context of path planning for the autonomous exploration of multi-floor buildings. We propose a system for detecting and modeling ascending stairways while performing simultaneous localization and mapping, such that the traversability of each stairway can be assessed by estimating its physical properties. The long-term objective of our approach is to enable exploration of multiple floors of a building by allowing stairways to be considered during path planning as traversable portals to new frontiers. We design a generative model of a stairway as a single object. We localize these models with respect to the map, and estimate the dimensions of the stairway as a whole, as well as its steps. With these estimates, a robot can determine if the stairway is traversable based on its climbing capabilities. Our system consists of two parts: a computationally efficient detector that leverages geometric cues from dense depth imagery to detect sets of ascending stairs, and a stairway modeler that uses multiple detections to infer the location and parameters of a stairway that is discovered during exploration. We demonstrate the performance of this system when deployed on several mobile platforms using a Microsoft Kinect sensor. Jeffrey A. Delmerico, David Baran, Philip David, Julian Ryde, Jason J. Corso |
ICRA | 5 |
| 2013 | Mutual localization: Two camera relative 6-DOF pose estimation from reciprocal fiducial observationabstractConcurrently estimating the 6-DOF pose of multiple cameras or robots - cooperative localization - is a core problem in contemporary robotics. Current works focus on a set of mutually observable world landmarks and often require inbuilt egomotion estimates; situations in which both assumptions are violated often arise, for example, robots with erroneous low quality odometry and IMU exploring an unknown environment. In contrast to these existing works in cooperative localization, we propose a cooperative localization method, which we call mutual localization, that uses reciprocal observations of camera-fiducials to obviate the need for egomotion estimates and mutually observable world landmarks. We formulate and solve an algebraic formulation for the pose of the two camera mutual localization setup under these assumptions. Our experiments demonstrate the capabilities of our proposal egomotion-free cooperative localization method: for example, the method achieves 2cm range and 0.7 degree accuracy at 2m sensing for 6-DOF pose. To demonstrate the applicability of the proposed work, we deploy our method on Turtlebots and we compare our results with ARToolKit [1] and Bundler [2], over which our method achieves a tenfold improvement in translation estimation accuracy. Vikas Dhiman, Julian Ryde, Jason J. Corso |
IROS | 3 |
| 2013 | Semi-automatic Brain Tumor Segmentation by Constrained MRFs Using Structural Trajectories
Liang Zhao 0018, Jason J. Corso |
MICCAI (3) | 3 |
| 2013 | Translating related words to videos and back through latent topicsabstractDocuments containing video and text are becoming more and more widespread and yet content analysis of those documents depends primarily on the text. Although automated discovery of semantically related words from text improves free text query understanding, translating videos into text summaries facilitates better video search particularly in the absence of accompanying text. In this paper, we propose a multimedia topic modeling framework suitable for providing a basis for automatically discovering and translating semantically related words obtained from textual metadata of multimedia documents to semantically related videos or frames from videos. Pradipto Das, Rohini K. Srihari, Jason J. Corso |
WSDM | 3 |
| 2013 | Hamiltonian streamline-guided features
Yingjie Miao, Jason J. Corso |
Neurocomputing | 2 |
| 2013 | Building facade detection, segmentation, and parameter estimation for mobile robot stereo vision
Jeffrey A. Delmerico, Philip David, Jason J. Corso |
Image Vis. Comput. | 3 |
| 2013 | Toward parts-based scene understanding with pixel-support parts-sparse pictorial structures
Jason J. Corso |
Pattern Recognit. Lett. | 1 |
| 2012 | Action bank: A high-level representation of activity in videoabstractActivity recognition in video is dominated by low- and mid-level features, and while demonstrably capable, by nature, these features carry little semantic meaning. Inspired by the recent object bank approach to image representation, we present Action Bank, a new high-level representation of video. Action bank is comprised of many individual action detectors sampled broadly in semantic space as well as viewpoint space. Our representation is constructed to be semantically rich and even when paired with simple linear SVM classifiers is capable of highly discriminative performance. We have tested action bank on four major activity recognition benchmarks. In all cases, our performance is better than the state of the art, namely 98.2% on KTH (better by 3.3%), 95.0% on UCF Sports (better by 3.7%), 57.9% on UCF50 (baseline is 47.9%), and 26.9% on HMDB51 (baseline is 23.2%). Furthermore, when we analyze the classifiers, we find strong transfer of semantics from the constituent action detectors to the bank classifier. Sreemanananth Sadanand, Jason J. Corso |
CVPR | 2 |
| 2012 | Evaluation of super-voxel methods for early video processingabstractSupervoxel segmentation has strong potential to be incorporated into early video analysis as superpixel segmentation has in image analysis. However, there are many plausible supervoxel methods and little understanding as to when and where each is most appropriate. Indeed, we are not aware of a single comparative study on supervoxel segmentation. To that end, we study five supervoxel algorithms in the context of what we consider to be a good supervoxel: namely, spatiotemporal uniformity, object/region boundary detection, region compression and parsimony. For the evaluation we propose a comprehensive suite of 3D volumetric quality metrics to measure these desirable supervoxel characteristics. We use three benchmark video data sets with a variety of content-types and varying amounts of human annotations. Our findings have led us to conclusive evidence that the hierarchical graph-based and segmentation by weighted aggregation methods perform best and almost equally-well on nearly all the metrics and are the methods of choice given our proposed assumptions. Chenliang Xu, Jason J. Corso |
CVPR | 2 |
| 2012 | Streaming Hierarchical Video Segmentation
Chenliang Xu, Caiming Xiong, Jason J. Corso |
ECCV (6) | 3 |
| 2012 | Dictionary transfer for image denoising via domain adaptationabstractThe idea of using overcomplete dictionaries with prototype signal atoms for sparse representation has found many applications, among which image denoising is considered as an active research topic. However, the standard process to train a new dictionary for image denoising requires the whole image (or most parts) as input, which is costly; training the dictionary on just a few patches would result in overfitting. We instead propose a dictionary learning approach for image denoising via transfer learning. We transfer the source domain dictionary to a target domain for image denoising via a dictionary-regularization term in the energy function. Thus, we have a new dictionary that is trained from only a few patches of the target noisy image. We measure the performance on various corrupted images, and show that our method is fast and comparable to the state of the art. We also demonstrate cross-domain transfer (photo to medical image). Gang Chen 0032, Caiming Xiong, Jason J. Corso |
ICIP | 3 |
| 2012 | Maintaining Prior Distributions across Evolving Eigenspaces: An Application to Portfolio ConstructionabstractTemporal evolution in the generative distribution of nonstationary sequential data is challenging to model. This paper presents a method for retaining the information in prior distributions of matrix variate dynamic linear models (MVDLMs) as the eigenspace of sequential data evolves. The method starts by constructing sliding windows â" matrices composed of a fixed number of columns containing the most recent point-in-time multivariate observation vectors. Characteristic time series, the right singular vectors, are extracted from a window using singular value decomposition (SVD). Then, a sequence of matrices capturing the rotation and scaling of the eigenspace is specified as a function of adjacent windowsâ characteristic time series. The method is tested on observations derived from daily US stock prices spanning 25 years. The results indicate that models constructed using sliding window SVD and MVDLMs, as extended in this paper, are resistant to over-fitting and perform well when used in portfolio construction applications. Kevin R. Keane, Jason J. Corso |
ICMLA (2) | 2 |
| 2012 | Dynamically Mixing Dynamic Linear Models with Applications in Finance
Kevin R. Keane, Jason J. Corso |
ICPRAM (2) | 2 |
| 2012 | Ascending stairway modeling: A first step toward autonomous multi-floor explorationabstractMany robotics platforms are capable of ascending stairways, but all existing approaches for autonomous stair climbing use stairway detection as a trigger for immediate traversal. In the broader context of autonomous exploration, the ability to travel between floors of a building should be compatible with path planning, such that the robot can traverse a stairway at a time that is appropriate to its navigation goals. No system yet presented is capable of both localizing stairways on a map and estimating their properties, functions that in combination would enable stairways to be considered as traversable terrain in a path planning algorithm. We propose a method for modeling stairways as objects and localizing them on a map, such that they can be subsequently traversed if they are of dimensions that the robotic platform is capable of climbing. Our system consists of two parts: a computationally efficient detector that leverages geometric cues from depth imagery to detect sets of ascending stairs, and a stairway modeler that uses multiple detections to infer the location and parameters of a stairway that is discovered during exploration. This video demonstrates the performance of the system in a number of real-world situations, modeling and localizing a variety of stairway types in both indoor and outdoor environments. Jeffrey A. Delmerico, Jason J. Corso, David Baran, Philip David, Julian Ryde |
IROS | 2 |
| 2012 | Fast voxel maps with counting bloom filtersabstractIn order to achieve good and timely volumetric mapping for mobile robots, we improve the speed and accuracy of multi-resolution voxel map building from 3D data. Mobile robot capabilities, such as SLAM and path planning, often involve algorithms that query a map many times and this lookup is often the bottleneck limiting the execution speed. As such, fast spatial proximity queries has been the topic of much active research. Various data structures have been researched including octrees, k-d trees, approximate nearest neighbours and even dense 3D arrays. We tackle this problem by extending previous work that stores the map as a hash table containing occupied voxels at multiple resolutions. We apply Bloom filters to the problem of spatial querying and voxel maps for the example application of SLAM. Their efficacy is demonstrated building 3D maps with both simulated and real 3D point cloud data. Looking up whether a voxel is occupied is three times faster than the hash table and within 10% of the speed of querying a dense 3D array, potentially the upper limit to query speed. Map generation was done with scan to map alignment on simulated depth images, for which the true pose is available. The calculated poses exhibited sub-voxel error of 0.02m and 0.3 degrees for a typical indoor scene with a map resolution of 0.04m. Julian Ryde, Jason J. Corso |
IROS | 2 |
| 2012 | Random forests for metric learning with implicit pairwise position dependenceabstractMetric learning makes it plausible to learn semantically meaningful distances for complex distributions of data using label or pairwise constraint information. However, to date, most metric learning methods are based on a single Mahalanobis metric, which cannot handle heterogeneous data well. Those that learn multiple metrics throughout the feature space have demonstrated superior accuracy, but at a severe cost to computational efficiency. Here, we adopt a new angle on the metric learning problem and learn a single metric that is able to implicitly adapt its distance function throughout the feature space. This metric adaptation is accomplished by using a random forest-based classifier to underpin the distance function and incorporate both absolute pairwise position and standard relative position into the representation. We have implemented and tested our method against state of the art global and multi-metric methods on a variety of data sets. Overall, the proposed method outperforms both types of method in terms of accuracy (consistently ranked first) and is an order of magnitude faster than state of the art multi-metric methods (16x faster in the worst case). Caiming Xiong, David M. Johnson 0001, Ran Xu 0001, Jason J. Corso |
KDD | 4 |
| 2011 | Building facade detection, segmentation, and parameter estimation for mobile robot localization and guidanceabstractBuilding facade detection is an important problem in computer vision, with applications in mobile robotics and semantic scene understanding. In particular, mobile platform localization and guidance in urban environments can be enabled with an accurate segmentation of the various building facades in a scene. Toward that end, we present a system for segmenting and labeling an input image that for each pixel, seeks to answer the question ¿Is this pixel part of a building facade, and if so, which one?¿ The proposed method determines a set of candidate planes by sampling and clustering points from the image with Random Sample Consensus (RANSAC), using local normal estimates derived from Principal Component Analysis (PCA) to inform the planar model. The corresponding disparity map and a discriminative classification provide prior information for a two-layer Markov Random Field model. This MRF problem is solved via Graph Cuts to obtain a labeling of building facade pixels at the mid-level, and a segmentation of those pixels into particular planes at the high-level. The results indicate a strong improvement in the accuracy of the binary building detection problem over the discriminative classifier alone, and the planar surface estimates provide a good approximation to the ground truth planes. Jeffrey A. Delmerico, Philip David, Jason J. Corso |
IROS | 3 |
| 2011 | Temporally consistent multi-class video-object segmentation with the Video Graph-Shifts algorithmabstractWe present the Video Graph-Shifts (VGS) approach for efficiently incorporating temporal consistency into MRF energy minimization for multi-class video object segmentation. In contrast to previous methods, our dynamic temporal links avoid the computational overhead of using a fully connected spatiotemporal MRF, while still being able to deal with the uncertainties of the exact inter-frame pixel correspondence issues. The dynamic temporal links are initialized flexibly for balancing between speed and accuracy, and are automatically revised whenever a label change (shift) occurs during the energy minimization process. We show in the benchmark CamVid database and our own wintry driving dataset that VGS improves the issue of temporally inconsistent segmentation effectively - enhancements of up to 5% to 10% for those semantic classes with high intra-class variance. Furthermore, VGS processes each frame at pixel resolution in about one second, which provides a practical way of modeling complex probabilistic relationships in videos and solving it in near real-time. Albert Y. C. Chen 0002, Jason J. Corso |
WACV | 2 |
| 2011 | AirTouch: Interacting with computer systems at a distanceabstractWe present AirTouch, a new vision-based interaction system. AirTouch uses computer vision techniques to extend commonly used interaction metaphors, such as multitouch screens, yet removes any need to physically touch the display. The user interacts with a virtual plane that rests in between the user and the display. On this plane, hands and fingers are tracked and gestures are recognized in a manner similar to a multitouch surface. Many of the other vision and gesture-based human-computer interaction systems presented in the literature have been limited by requirements that users do not leave the frame or do not perform gestures accidentally, as well as by cost or specialized equipment. AirTouch does not suffer from these drawbacks. Instead, it is robust, easy to use, builds on a familiar interaction paradigm, and can be implemented using a single camera with off-the-shelf equipment such as a webcam-enabled laptop. In order to maintain usability and accessibility while minimizing cost, we present a set of basic AirTouch guidelines. We have developed two interfaces using these guidelines-one for general computer interaction, and one for searching an image database. We present the workings of these systems along with observational results regarding their usability. Daniel R. Schlegel, Albert Y. C. Chen 0002, Caiming Xiong, Jeffrey A. Delmerico, Jason J. Corso |
WACV | 5 |
| 2011 | Labeling of Lumbar Discs Using Both Pixel- and Object-Level Features With a Two-Level Probabilistic ModelabstractBackbone anatomical structure detection and labeling is a necessary step for various analysis tasks of the vertebral column. Appearance, shape and geometry measurements are necessary for abnormality detection locally at each disc and vertebrae (such as herniation) as well as globally for the whole spine (such as spinal scoliosis). We propose a two-level probabilistic model for the localization of discs from clinical magnetic resonance imaging (MRI) data that captures both pixel- and object-level features. Using a Gibbs distribution, we model appearance and spatial information at the pixel level, and at the object level, we model the spatial distribution of the discs and the relative distances between them. We use generalized expectation-maximization for optimization, which achieves efficient convergence of disc labels. Our two-level model allows the assumption of conditional independence at the pixel-level to enhance efficiency while maintaining robustness. We use a dataset that contains 105 MRI clinical normal and abnormal cases for the lumbar area. We thoroughly test our model and achieve encouraging results on normal and abnormal cases. Raja' S. Alomari, Jason J. Corso, Vipin Chaudhary |
IEEE Trans. Medical Imaging | 2 |
| 2010 | A Framework for Hand Gesture Recognition and Spotting Using Sub-gesture ModelingabstractHand gesture interpretation is an open research problem in Human Computer Interaction (HCI), which involves locating gesture boundaries (Gesture Spotting) in a continuous video sequence and recognizing the gesture. Existing techniques model each gesture as a temporal sequence of visual features extracted from individual frames which is not efficient due to the large variability of frames at different timestamps. In this paper, we propose a new sub-gesture modeling approach which represents each gesture as a sequence of fixed sub-gestures (a group of consecutive frames with locally coherent context) and provides a robust modeling of the visual features. We further extend this approach to the task of gesture spotting where the gesture boundaries are identified using a filler model and gesture completion model. Experimental results show that the proposed method outperforms state-of-the-art Hidden Conditional Random Fields (HCRF) based methods and baseline gesture spotting techniques. Manavender R. Malgireddy, Jason J. Corso, Srirangaraj Setlur, Venu Govindaraju, Dinesh Mandalapu |
ICPR | 2 |
| 2009 | Geometric tomography: a limited-view approach for computed tomographyabstractNo abstract available. Peter B. Noël, Jinhui Xu 0001, Kenneth R. Hoffmann, Jason J. Corso |
SCG | 4 |
| 2009 | Robust unsupervised segmentation of degraded document images with topic modelsabstractSegmentation of document images remains a challenging vision problem. Although document images have a structured layout, capturing enough of it for segmentation can be difficult. Most current methods combine text extraction and heuristics for segmentation, but text extraction is prone to failure and measuring accuracy remains a difficult challenge. Furthermore, when presented with significant degradation many common heuristic methods fall apart. In this paper, we propose a Bayesian generative model for document images which seeks to overcome some of these drawbacks. Our model automatically discovers different regions present in a document image in a completely unsupervised fashion. We attempt no text extraction, but rather use discrete patch-based codebook learning to make our probabilistic representation feasible. Each latent region topic is a distribution over these patch indices. We capture rough document layout with an MRF Potts model. We take an analysis by synthesis approach to examine the model, and provide quantitative segmentation results on a manually labeled document image data set. We illustrate our model's robustness by providing results on a highly degraded version of our test set. Timothy J. Burns, Jason J. Corso |
CVPR | 2 |
| 2009 | Multi-level Ground Glass Nodule Detection and Segmentation in CT Lung Images
Yimo Tao, Le Lu 0001, Maneesh Dewan, Albert Y. Chen, Jason J. Corso, Jianhua Xuan, Marcos Salganicoff, Arun Krishnan |
MICCAI (1) | 5 |
| 2009 | Image description with features that summarize
Jason J. Corso, Gregory D. Hager |
Comput. Vis. Image Underst. | 1 |
| 2008 | Discriminative modeling by Boosting on Multilevel AggregatesabstractThis paper presents a new approach to discriminative modeling for classification and labeling. Our method, called boosting on multilevel aggregates (BMA), adds a new class of hierarchical, adaptive features into boosting-based discriminative models. Each pixel is linked with a set of aggregate regions in a multilevel coarsening of the image. The coarsening is adaptive, rapid and stable. The multilevel aggregates present additional information rich features on which to boost, such as shape properties, neighborhood context, hierarchical characteristics, and photometric statistics. We implement and test our approach on three two-class problems: classifying documents in office scenes, buildings and horses in natural images. In all three cases, the majority, about 75%, of features selected during boosting are our proposed BMA features rather than patch-based features. This large percentage demonstrates the discriminative power of the multilevel aggregate features over conventional patch-based features. Our quantitative performance measures show the proposed approach gives superior results to the state-of-the-art in all three applications. Jason J. Corso |
CVPR | 1 |
| 2008 | Graph-shifts: Natural image labeling by dynamic hierarchical computingabstractIn this paper, we present a new approach for image labeling based on the recently introduced graph-shifts algorithm. Graph-shifts is an energy minimization algorithm that does labeling by dynamically manipulating, or shifting, the parent-child relationships in a hierarchical decomposition of the image. Each shift optimally reduces the energy by indirectly causing a change to the labeling; graph-shifts is able to rapidly compute and select this optimal shift at every iteration. There are no constraints on the terms of the (pairwise) energy function. The algorithm was originally presented in the context of medical image labeling using conditional random field models. In this paper, we consider the algorithm in the context of both low- and high-level natural image labeling. We show that for examples in both classes of problems, graph-shifts does labeling both accurately and rapidly. For low-level vision, we explore image restoration, and for high-level vision, we make use of a hybrid discriminative-generative model to segment and label images into semantically meaningful regions (e.g., trees, buildings, etc.). For both problems, we obtain comparable or superior results to the state-of-the-art computed in just a few seconds per image. Jason J. Corso, Alan L. Yuille, Zhuowen Tu |
CVPR | 1 |
| 2008 | (BP)2: Beyond pairwise Belief Propagation labeling by approximating Kikuchi free energiesabstractBelief propagation (BP) can be very useful and efficient for performing approximate inference on graphs. But when the graph is very highly connected with strong conflicting interactions, BP tends to fail to converge. Generalized Belief Propagation (GBP) provides more accurate solutions on such graphs, by approximating Kikuchi free energies, but the clusters required for the Kikuchi approximations are hard to generate. We propose a new algorithmic way of generating such clusters from a graph without exponentially increasing the size of the graph during triangulation. In order to perform the statistical region labeling, we introduce the use of superpixels for the nodes of the graph, as it is a more natural representation of an image than the pixel grid. This results in a smaller but much more highly interconnected graph where BP consistently fails. We demonstrate how our version of the GBP algorithm outperforms BP on synthetic and natural images and in both cases, GBP converges after only a few iterations. Ifeoma Nwogu, Jason J. Corso |
CVPR | 2 |
| 2008 | HOPS: Efficient region labeling using Higher Order Proxy NeighborhoodsabstractWe present the Higher Order Proxy Neighborhoods (HOPS) approach to modeling higher order neighborhoods in Markov Random Fields (MRFs). HOPS incorporates more context information into the energy function in a recursive and cached manner. It induces little or no additional computational cost in the overall minimization process, and can better represent the underlying energy leading to fewer total computations. Indeed, when integrated with the Graph-Shifts energy minimization algorithm we observe a 30% average decrease to the convergence time. We apply HOPS to high-level labeling of natural and geospatial images; our results show that HOPS leads to smoother labelings that better follow object boundaries. HOPS can label an image with an average 75% accuracy in a couple of seconds. Albert Y. C. Chen 0002, Jason J. Corso |
ICPR | 2 |
| 2008 | Integrating minutiae based fingerprint matching with local mutual informationabstractMinutiae based fingerprint matching algorithms are wildly used in fingerprint identification and verification applications. However, they may suffer from spurious matches because they do not use the rich local image information. In this paper, we extend minutiae based methods to incorporate such local image information. Our method uses local mutual information, a proven similarity measure in various applications, to improve the matching rate. The overall minutiae distribution pattern between two fingerprints is represented by the initial minutiae matching result, while the mutual information measures the similarity between neighborhoods of matched minutiae, thus enhancing the final matching decision. FVC2002 DB1 and DB3 databases are used to test the proposed approach. Experimental result shows the improvement when combining minutiae matching scores with mutual information scores. Sergey Tulyakov, Faisal Farooq, Jason J. Corso, Venu Govindaraju |
ICPR | 4 |
| 2008 | MRF Labeling with a Graph-Shifts Algorithm
Jason J. Corso, Zhuowen Tu, Alan L. Yuille |
IWCIA | 1 |
| 2008 | Labeling Irregular Graphs with Belief Propagation
Ifeoma Nwogu, Jason J. Corso |
IWCIA | 2 |
| 2008 | Lumbar Disc Localization and Labeling with a Probabilistic Model on Both Pixel and Object Features
Jason J. Corso, Raja' S. Alomari, Vipin Chaudhary |
MICCAI (1) | 1 |
| 2008 | Exploratory Identification of Image-Based Biomarkers for Solid Mass Pulmonary Tumors
Ifeoma Nwogu, Jason J. Corso |
MICCAI (1) | 2 |
| 2008 | Efficient Multilevel Brain Tumor Segmentation With Integrated Bayesian Model ClassificationabstractWe present a new method for automatic segmentation of heterogeneous image data that takes a step toward bridging the gap between bottom-up affinity-based segmentation methods and top-down generative model based approaches. The main contribution of the paper is a Bayesian formulation for incorporating soft model assignments into the calculation of affinities, which are conventionally model free. We integrate the resulting model-aware affinities into the multilevel segmentation by weighted aggregation algorithm, and apply the technique to the task of detecting and segmenting brain tumor and edema in multichannel magnetic resonance (MR) volumes. The computationally efficient method runs orders of magnitude faster than current state-of-the-art techniques giving comparable or improved results. Our quantitative results indicate the benefit of incorporating model-aware affinities into the segmentation process for the difficult case of glioblastoma multiforme brain tumor. Jason J. Corso, Eitan Sharon, Shishir Dube, Suzie El-Saden, Usha S. Sinha, Alan L. Yuille |
IEEE Trans. Medical Imaging | 1 |
| 2007 | Detection and Segmentation of Pathological Structures by the Extended Graph-Shifts Algorithm
Jason J. Corso, Alan L. Yuille, Nancy L. Sicotte, Arthur W. Toga |
MICCAI (1) | 1 |
| 2006 | Multilevel Segmentation and Integrated Bayesian Model Classification with an Application to Brain Tumor Segmentation
Jason J. Corso, Eitan Sharon, Alan L. Yuille |
MICCAI (2) | 1 |
| 2005 | Coherent Regions for Concise and Stable Image DescriptionabstractWe present a new method for summarizing images for the purposes of matching and registration. We take the point of view that large, coherent regions in the image provide a concise and stable basis for image description. We develop a new algorithm for image segmentation that operates on several projections (feature spaces) of the image, using kernel-based optimization techniques to locate local extrema of a continuous scale-space of image regions. Descriptors of these image regions and their relative geometry then form the basis of an image description. We present experimental results of these methods applied to the problem of image retrieval. On a moderate sized database, we find that our method performs comparably to two published techniques: Blobworld and SIFT features. However, compared to these techniques two significant advantages of our method are its 1) stability under large changes in the images and 2) its representational efficiency. As a result we argue our proposed method will scale well with larger image sets. Jason J. Corso, Gregory D. Hager |
CVPR (2) | 1 |
| 2004 | Stereo-Based Endoscopic Tracking of Cardiac Surface Deformation
William W. Lau, Nicholas A. Ramey, Jason J. Corso, Nitish V. Thakor, Gregory D. Hager |
MICCAI (2) | 3 |
| 2004 | VICs: A modular HCI framework using spatiotemporal dynamics
Guangqi Ye, Jason J. Corso, Darius Burschka, Gregory D. Hager |
Mach. Vis. Appl. | 2 |
| 2003 | Direct plane tracking in stereo images for mobile navigationabstractWe present a novel plane tracking algorithm based on the direct update of surface parameters from two stereo images. The plane tracking algorithm is posed as an optimization problem, and maintains an iteratively re-weighted least squares approximation of the plane's orientation using direct pixel measurements. To facilitate autonomous operation, we include an algorithm for robust detection of significant planes in the environment. The algorithms have been implemented in a robot navigation system. Jason J. Corso, Darius Burschka, Gregory D. Hager |
ICRA | 1 |
| 2003 | VICs: A Modular Vision-Based HCI Framework
Guangqi Ye, Jason J. Corso, Darius Burschka, Gregory D. Hager |
ICVS | 2 |
| 2003 | VisHap: augmented reality combining haptics and visionabstractHaptic devices have been successfully incorporated into the human-computer interaction model. However, a drawback common to almost all haptic systems is that the user must be attached to the haptic devices at all times even though force feedback is not always being rendered. This constant contact hinders perception of the virtual environment, primarily because it prevents the user from feeling new tactile sensations upon contact with virtual objects. We present the design and implementation of an augmented reality system called VisHap that uses visual tracking to seamlessly integrate force feedback with tactile feedback to generate a "complete" haptic experience. The VisHap framework allows the user to interact with combinations of virtual and real objects naturally, thereby combining active and passive haptics. An example application of this framework is also presented. The flexibility and extensibility of our framework is promising in that it supports many interaction modes and allows further integration with other augmented reality system. Guangqi Ye, Jason J. Corso, Gregory D. Hager, Allison M. Okamura |
SMC | 2 |