VLDB 2026 Research / reviewers in the wild / expert
Vlad I. Morariu
dblp:27/6671
· DBLP profile ↗
60ranked-venue papers
5as first author
20since 2021 · last 2026
0000-0001-7937-7748ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 48 · 4 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 39 · 3 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Agentic Design Review System
Sayan Nag, K. J. Joseph, Koustava Goswami, Vlad I. Morariu, Balaji Vasan Srinivasan |
AAAI | 4 |
| 2025 | DELOC: Document Element LocalizerabstractEditing documents and PDFs using natural language instructions is desirable for many reasons -ease of use, increasing accessibility to non-technical users, and for creativity.To do this automatically, a system needs to first understand the user's intent and convert this to an executable plan or command, and then the system needs to identify or localize the elements that the user desires to edit.While there exist methods that can accomplish these tasks, a major bottleneck in these systems is the inability to ground the spatial edit location effectively.We address this gap through our proposed system, DELOC (Document Element LOCalizer).DELOC adapts the grounding capabilities of existing Multimodal Large Language Model (MLLM) from natural images to PDFs.This adaptation involves two novel contributions: 1) synthetically generating PDF-grounding instruction tuning data from partially annotated datasets; and 2) synthetic data cleaning via Code-NLI, an NLI-inspired process to clean data using generated Python code.The effectiveness of DELOC is apparent in the >2x zeroshot improvement it achieves over the next best MLLM, GPT-4o. Hammad A. Ayyubi, Puneet Mathur, Md. Mehrab Tanjim, Vlad I. Morariu |
EMNLP | 4 |
| 2025 | FormA11y - Research and Development of a Tool for Remediating PDF Forms for AccessibilityabstractPDF documents are usually not born-accessible, and so document authors need to put in additional work (remediation) to make them accessible for people with disabilities. Unfortunately, this step is often overlooked and hard to execute, resulting in a large number of inaccessible PDF documents on the internet. Previously, there have been research efforts to investigate potential solutions for remediating PDF documents for accessibility. However, most of the existing research focuses on accessibility of long or scientific PDF documents meant for passive reading. PDF documents come in different types, and this research project focuses on a distinct type of PDF document-forms-where the user is required to interact with the PDF document and enter data. Through our research work we identified that the PDF form remediation process is non-intuitive, repetitive, and overwhelming due to the high-information density of PDF forms, and existing research and tools do not yet address the challenges. Our research work culminated in the creation of a tool - FormA11y - that addresses these challenges by making the repetitive and painstaking process of form remediation easier. To evaluate the effectiveness and efficiency of FormA11y against the industry standard tool - Adobe Acrobat - for PDF form remediation, we performed a within-subject user study with 20 participants. With FormA11y, users remediated forms 2.8 times faster while creating more accurately accessible PDF forms. Sparsh Paliwal, Joshua Hoeflich, J. Bern Jordan, Rajiv Jain, Vlad I. Morariu, Alexa F. Siu, Jonathan Lazar |
ACM Trans. Comput. Hum. Interact. | 5 |
| 2024 | TutoAI: a cross-domain framework for AI-assisted mixed-media tutorial creation on physical tasksabstractMixed-media tutorials, which integrate videos, images, text, and diagrams to teach procedural skills, offer more browsable alternatives than timeline-based videos. However, manually creating such tutorials is tedious, and existing automated solutions are often restricted to a particular domain. While AI models hold promise, it is unclear how to effectively harness their powers, given the multi-modal data involved and the vast landscape of models. We present TutoAI, a cross-domain framework for AI-assisted mixed-media tutorial creation on physical tasks. First, we distill common tutorial components by surveying existing work; then, we present an approach to identify, assemble, and evaluate AI models for component extraction; finally, we propose guidelines for designing user interfaces (UI) that support tutorial creation based on AI-generated components. We show that TutoAI has achieved higher or similar quality compared to a baseline model in preliminary user studies. Yuexi Chen, Vlad I. Morariu, Anh Truong, Zhicheng Liu 0001 |
CHI | 2 |
| 2024 | DocScript: Document-level Script Event PredictionabstractWe present a novel task of document-level script event prediction, which aims to predict the next event given a candidate list of narrative events in long-form documents. To enable this, we introduce DocSEP, a challenging dataset in two new domains - contractual documents and Wikipedia articles, where timeline events may be paragraphs apart and may require multi-hop temporal and causal reasoning. We benchmark existing baselines and present a novel architecture called DocScript to learn sequential ordering between events at the document scale. Our experimental results on the DocSEP dataset demonstrate that learning longer-range dependencies between events is a key challenge and show that contemporary LLMs such as ChatGPT and FlanT5 struggle to solve this task, indicating their lack of reasoning abilities for understanding causal relationships and temporal sequences within long texts. Puneet Mathur, Vlad I. Morariu, Aparna Garimella, Franck Dernoncourt, Jiuxiang Gu, Ramit Sawhney, Preslav Nakov, Dinesh Manocha, Rajiv Jain |
LREC/COLING | 2 |
| 2024 | DocEdit-v2: Document Structure Editing Via Multimodal LLM GroundingabstractManan Suri, Puneet Mathur, Franck Dernoncourt, Rajiv Jain, Vlad I Morariu, Ramit Sawhney, Preslav Nakov, Dinesh Manocha. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Manan Suri, Puneet Mathur, Franck Dernoncourt, Rajiv Jain, Vlad I. Morariu, Ramit Sawhney, Preslav Nakov, Dinesh Manocha |
EMNLP | 5 |
| 2024 | Localizing and Editing Knowledge In Text-to-Image Generative ModelsabstractText-to-Image Diffusion Models such as Stable-Diffusion and Imagen have achieved unprecedented quality of photorealism with state-of-the-art FID scores on MS-COCO and other generation benchmarks. Given a caption, image generation requires fine-grained knowledge about attributes such as object structure, style, and viewpoint amongst others. Where does this information reside in text-to-image generative models? In our paper, we tackle this question and understand how knowledge corresponding to distinct visual attributes is stored in large-scale text-to-image diffusion models. We adapt Causal Mediation Analysis for text-to-image models and trace knowledge about distinct visual attributes to various (causal) components in the (i) UNet and (ii) text-encoder of the diffusion model.
In particular, we show that unlike large-language models, knowledge about different attributes is not localized in isolated components, but is instead distributed amongst a set of components in the conditional UNet. These sets of components are often distinct for different visual attributes (e.g., style} / objects). Remarkably, we find that the text-encoder in public text-to-image models such as Stable-Diffusion contains {\it only} one causal state across different visual attributes, and this is the first self-attention layer corresponding to the last subject token of the attribute in the caption. This is in stark contrast to the causal states in other language models which are often the mid-MLP layers. Based on this observation of only one causal state in the text-encoder, we introduce a fast, data-free model editing method DiffQuickFix which can effectively edit concepts (remove or update knowledge) in text-to-image models. DiffQuickFix can edit (ablate) concepts in under a second with a closed-form update, providing a significant 1000x speedup and comparable editing performance to existing fine-tuning based editing methods. Samyadeep Basu, Nanxuan Zhao, Vlad I. Morariu, Soheil Feizi, Varun Manjunatha |
ICLR | 3 |
| 2024 | On Mechanistic Knowledge Localization in Text-to-Image Generative ModelsabstractIdentifying layers within text-to-image models which control visual attributes can facilitate efficient model editing through closed-form updates. Recent work, leveraging causal tracing show that early Stable-Diffusion variants confine knowledge primarily to the first layer of the CLIP text-encoder, while it diffuses throughout the UNet. Extending this framework, we observe that for recent models (e.g., SD-XL, DeepFloyd), causal tracing fails in pinpointing localized knowledge, highlighting challenges in model editing. To address this issue, we introduce the concept of mechanistic localization in text-to-image models, where knowledge about various visual attributes (e.g., "style", "objects", "facts") can be mechanistically localized to a small fraction of layers in the UNet, thus facilitating efficient model editing. We localize knowledge using our method LocoGen which measures the direct effect of intermediate layers to output generation by performing interventions in the cross-attention layers of the UNet. We then employ LocoEdit, a fast closed-form editing method across popular open-source text-to-image models (including the latest SD-XL) and explore the possibilities of neuron-level model editing. Using mechanistic localization, our work offers a better view of successes and failures in localization-based text-to-image model editing. Samyadeep Basu, Keivan Rezaei, Priyatham Kattakinda, Vlad I. Morariu, Nanxuan Zhao, Ryan Rossi, Varun Manjunatha, Soheil Feizi |
ICML | 4 |
| 2024 | FlexDoc: Flexible Document Adaptation through Optimizing both Content and LayoutabstractDesigning adaptive documents that are visually appealing across various devices and for diverse viewers is a challenging task. This is due to the wide variety of devices and different viewer requirements and preferences. Alterations to a document’s content, style, or layout often necessitate numerous adjustments, potentially leading to a complete layout redesign. We introduce FlexDoc, a framework for creating and consuming documents that seamlessly adapt to different devices, author, and viewer preferences and interactions. It eliminates the need to manually create multiple document layouts, as FlexDoc enables authors to define desired document properties using templates and employs both discrete and continuous optimization in a novel comprehensive optimization process, which leverages automatic text summarization and image carving techniques to adapt both layout and content during consumption dynamically. Further, we demonstrate FlexDoc in real-world scenarios. Yue Jiang 0002, Christof Lutteroth, Rajiv Jain, Chris Tensmeyer, Varun Manjunatha, Wolfgang Stuerzlinger, Vlad I. Morariu |
VL/HCC | 7 |
| 2023 | DocEdit: Language-Guided Document EditingabstractProfessional document editing tools require a certain level of expertise to perform complex edit operations. To make editing tools accessible to increasingly novice users, we investigate intelligent document assistant systems that can make or suggest edits based on a user's natural language request. Such a system should be able to understand the user's ambiguous requests and contextualize them to the visual cues and textual content found in a document image to edit localized unstructured text and structured layouts. To this end, we propose a new task of language-guided localized document editing, where the user provides a document and an open vocabulary editing request, and the intelligent system produces a command that can be used to automate edits in real-world document editing software. In support of this task, we curate the DocEdit dataset, a collection of approximately 28K instances of user edit requests over PDF and design templates along with their corresponding ground truth software executable commands. To our knowledge, this is the first dataset that provides a diverse mix of edit operations with direct and indirect references to the embedded text and visual objects such as paragraphs, lists, tables, etc. We also propose DocEditor, a Transformer-based localization-aware multimodal (textual, spatial, and visual) model that performs the new task. The model attends to both document objects and related text contents which may be referred to in a user edit request, generating a multimodal embedding that is used to predict an edit command and associated bounding box localizing it. Our proposed model empirically outperforms other baseline deep learning approaches by 15-18%, providing a strong starting point for future work. Puneet Mathur, Rajiv Jain, Jiuxiang Gu, Franck Dernoncourt, Dinesh Manocha, Vlad I. Morariu |
AAAI | 6 |
| 2023 | DocDancer: Authoring Ultra-Responsive Documents with Layout GenerationabstractResponsive design enhances user experience by adapting layout and content to different display factors. However, existing tools for authoring responsive design primarily adapt to screen width only, and they either rely on predefined templates or require significant manual effort. To expedite responsive design creation, we introduce an authoring tool called DocDancer. DocDancer supports creating ultra-responsive documents where both layout and content adapt to multiple factors (screen width, font properties, customer segments, etc), and provides layout suggestions based on user-provided content and popular responsive patterns. A comparative user study with 16 participants shows that authoring responsive documents in DocDancer takes significantly less time and effort than a commercial tool. Yuexi Chen, Zhicheng Liu 0001, Chris Tensmeyer, Niklas Elmqvist, Vlad I. Morariu |
VL/HCC | 5 |
| 2023 | LayerDoc: Layer-wise Extraction of Spatial Hierarchical Structure in Visually-Rich DocumentsabstractDigital documents often contain images and scanned text. Parsing such visually-rich documents is a core task for work-flow automation, but it remains challenging since most documents do not encode explicit layout information, e.g., how characters and words are grouped into boxes and ordered into larger semantic entities. Current state-of-the-art layout extraction methods are challenged by such documents as they rely on word sequences to have correct reading order and do not exploit their hierarchical structure. We propose LayerDoc, an approach that uses visual features, textual semantics, and spatial coordinates along with constraint inference to extract the hierarchical layout structure of documents in a bottom-up layer-wise fashion. LayerDoc recursively groups smaller regions into larger semantic elements in 2D to infer complex nested hierarchies. Experiments show that our approach outperforms competitive baselines by 10-15% on three diverse datasets of forms and mobile app screen layouts for the tasks of spatial region classification, higher-order group identification, layout hierarchy extraction, reading order detection, and word grouping. Puneet Mathur, Rajiv Jain, Ashutosh Mehra 0002, Jiuxiang Gu, Franck Dernoncourt, Anandhavelu Natarajan, Quan Hung Tran, Verena Kaynig, Ani Nenkova, Dinesh Manocha, Vlad I. Morariu |
WACV | 11 |
| 2022 | MGDoc: Pre-training with Multi-granular Hierarchy for Document Image UnderstandingabstractZilong Wang, Jiuxiang Gu, Chris Tensmeyer, Nikolaos Barmpalios, Ani Nenkova, Tong Sun, Jingbo Shang, Vlad Morariu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Zilong Wang 0002, Jiuxiang Gu, Chris Tensmeyer, Nikolaos Barmpalios, Ani Nenkova, Tong Sun 0005, Jingbo Shang, Vlad I. Morariu |
EMNLP | 8 |
| 2022 | DocLayoutTTS: Dataset and Baselines for Layout-informed Document-level Neural Speech Synthesis
Puneet Mathur, Franck Dernoncourt, Quan Hung Tran, Jiuxiang Gu, Ani Nenkova, Vlad I. Morariu, Rajiv Jain, Dinesh Manocha |
INTERSPEECH | 6 |
| 2022 | DocTime: A Document-level Temporal Dependency Graph ParserabstractPuneet Mathur, Vlad Morariu, Verena Kaynig-Fittkau, Jiuxiang Gu, Franck Dernoncourt, Quan Tran, Ani Nenkova, Dinesh Manocha, Rajiv Jain. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Puneet Mathur, Vlad I. Morariu, Verena Kaynig, Jiuxiang Gu, Franck Dernoncourt, Quan Hung Tran, Ani Nenkova, Dinesh Manocha, Rajiv Jain |
NAACL-HLT | 2 |
| 2021 | Syntopical Graphs for Computational Argumentation TasksabstractJoe Barrow, Rajiv Jain, Nedim Lipka, Franck Dernoncourt, Vlad Morariu, Varun Manjunatha, Douglas Oard, Philip Resnik, Henning Wachsmuth. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Joe Barrow, Rajiv Jain, Nedim Lipka, Franck Dernoncourt, Vlad I. Morariu, Varun Manjunatha, Douglas W. Oard, Philip Resnik, Henning Wachsmuth |
ACL/IJCNLP (1) | 5 |
| 2021 | SelfDoc: Self-Supervised Document Representation LearningabstractWe propose SelfDoc, a task-agnostic pre-training framework for document image understanding. Because documents are multimodal and are intended for sequential reading, our framework exploits the positional, textual, and visual information of every semantically meaningful component in a document, and it models the contextualization between each block of content. Unlike existing document pre-training models, our model is coarse-grained instead of treating individual words as input, therefore avoiding an overly fine-grained with excessive contextualization. Beyond that, we introduce cross-modal learning in the model pre-training phase to fully leverage multimodal information from unlabeled documents. For downstream usage, we propose a novel modality-adaptive attention mechanism for multimodal feature fusion by adaptively emphasizing language and vision signals. Our framework benefits from self-supervised pre-training on documents without requiring annotations by a feature masking training strategy. It achieves superior performance on multiple downstream tasks with significantly fewer document images used in the pre-training stage compared to previous works. Peizhao Li, Jiuxiang Gu, Jason Kuen, Vlad I. Morariu, Handong Zhao, Rajiv Jain, Varun Manjunatha, Hongfu Liu 0001 |
CVPR | 4 |
| 2021 | Black-Box Explanation of Object Detectors via Saliency MapsabstractWe propose D-RISE, a method for generating visual explanations for the predictions of object detectors. Utilizing the proposed similarity metric that accounts for both localization and categorization aspects of object detection allows our method to produce saliency maps that show image areas that most affect the prediction. D-RISE can be considered "black-box" in the software testing sense, as it only needs access to the inputs and outputs of an object detector. Compared to gradient-based methods, D-RISE is more general and agnostic to the particular type of object detector being tested, and does not need knowledge of the inner workings of the model. We show that D-RISE can be easily applied to different object detectors including one-stage detectors such as YOLOv3 and two-stage detectors such as Faster-RCNN. We present a detailed analysis of the generated visual explanations to highlight the utilization of context and possible biases learned by object detectors. Vitali Petsiuk, Rajiv Jain, Varun Manjunatha, Vlad I. Morariu, Ashutosh Mehra 0002, Vicente Ordonez, Kate Saenko |
CVPR | 4 |
| 2021 | IGA: An Intent-Guided Authoring AssistantabstractSimeng Sun, Wenlong Zhao, Varun Manjunatha, Rajiv Jain, Vlad Morariu, Franck Dernoncourt, Balaji Vasan Srinivasan, Mohit Iyyer. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Simeng Sun, Wenlong Zhao 0001, Varun Manjunatha, Rajiv Jain, Vlad I. Morariu, Franck Dernoncourt, Balaji Vasan Srinivasan, Mohit Iyyer |
EMNLP (1) | 5 |
| 2021 | UniDoc: Unified Pretraining Framework for Document UnderstandingabstractDocument intelligence automates the extraction of information from documents and supports many business applications. Recent self-supervised learning methods on large-scale unlabeled document datasets have opened up promising directions towards reducing annotation efforts by training models with self-supervised objectives. However, most of the existing document pretraining methods are still language-dominated. We present UDoc, a new unified pretraining framework for document understanding. UDoc is designed to support most document understanding tasks, extending the Transformer to take multimodal embeddings as input. Each input element is composed of words and visual features from a semantic region of the input document image. An important feature of UDoc is that it learns a generic representation by making use of three self-supervised losses, encouraging the representation to model sentences, learn similarities, and align modalities. Extensive empirical analysis demonstrates that the pretraining procedure learns better joint representations and leads to improvements in downstream tasks. Jiuxiang Gu, Jason Kuen, Vlad I. Morariu, Handong Zhao, Rajiv Jain, Nikolaos Barmpalios, Ani Nenkova, Tong Sun 0005 |
NeurIPS | 3 |
| 2020 | A Joint Model for Document Segmentation and Segment LabelingabstractText segmentation aims to uncover latent structure by dividing text from a document into coherent sections.Where previous work on text segmentation considers the tasks of document segmentation and segment labeling separately, we show that the tasks contain complementary information and are best addressed jointly.We introduce the Segment Pooling LSTM (S-LSTM) model, which is capable of jointly segmenting a document and labeling segments.In support of joint training, we develop a method for teaching the model to recover from errors by aligning the predicted and ground truth segments.We show that S-LSTM reduces segmentation error by 30% on average, while also improving segment labeling. Joe Barrow, Rajiv Jain, Vlad I. Morariu, Varun Manjunatha, Douglas W. Oard, Philip Resnik |
ACL | 3 |
| 2020 | Cross-Domain Document Object Detection: Benchmark Suite and MethodabstractDecomposing images of document pages into high-level semantic regions (e.g., figures, tables, paragraphs), document object detection (DOD) is fundamental for downstream tasks like intelligent document editing and understanding. DOD remains a challenging problem as document objects vary significantly in layout, size, aspect ratio, texture, etc. An additional challenge arises in practice because large labeled training datasets are only available for domains that differ from the target domain. We investigate cross-domain DOD, where the goal is to learn a detector for the target domain using labeled data from the source domain and only unlabeled data from the target domain. Documents from the two domains may vary significantly in layout, language, and genre. We establish a benchmark suite consisting of different types of PDF document datasets that can be utilized for cross-domain DOD model training and evaluation. For each dataset, we provide the page images, bounding box annotations, PDF files, and the rendering layers extracted from the PDF files. Moreover, we propose a novel cross-domain DOD model which builds upon the standard detection model and addresses domain shifts by incorporating three novel alignment modules: Feature Pyramid Alignment (FPA) module, Region Alignment (RA) module and Rendering Layer alignment (RLA) module. Extensive experiments on the benchmark suite substantiate the efficacy of the three proposed modules and the proposed method significantly outperforms the baseline methods. The project page is at \url{https://github.com/kailigo/cddod}. Kai Li 0012, Curtis Wigington, Chris Tensmeyer, Handong Zhao, Nikolaos Barmpalios, Vlad I. Morariu, Varun Manjunatha, Tong Sun 0005, Yun Fu 0001 |
CVPR | 6 |
| 2020 | Generative-Discriminative Feature Representations for Open-Set RecognitionabstractWe address the problem of open-set recognition, where the goal is to determine if a given sample belongs to one of the classes used for training a model (known classes). The main challenge in open-set recognition is to disentangle open-set samples that produce high class activations from known-set samples. We propose two techniques to force class activations of open-set samples to be low. First, we train a generative model for all known classes and then augment the input with the representation obtained from the generative model to learn a classifier. This network learns to associate high classification probabilities both when image content is from the correct class as well as when the input and the reconstructed image are consistent with each other. Second, we use self-supervision to force the network to learn more informative featues when assigning class scores to improve separation of classes from each other and from open-set samples. We evaluate the performance of the proposed method with recent open-set recognition works across three datasets, where we obtain state-of-the-art results. Pramuditha Perera, Vlad I. Morariu, Rajiv Jain, Varun Manjunatha, Curtis Wigington, Vicente Ordonez, Vishal M. Patel |
CVPR | 2 |
| 2020 | Self-Supervised Relationship ProbingabstractStructured representations of images that model visual relationships are beneficial for many vision and vision-language applications. However, current human-annotated visual relationship datasets suffer from the long-tailed predicate distribution problem which limits the potential of visual relationship models. In this work, we introduce a self-supervised method that implicitly learns the visual relationships without relying on any ground-truth visual relationship annotations. Our method relies on 1) intra- and inter-modality encodings to respectively model relationships within each modality separately and jointly, and 2) relationship probing, which seeks to discover the graph structure within each modality. By leveraging masked language modeling, contrastive learning, and dependency tree distances for self-supervision, our method learns better object features as well as implicit visual relationships. We verify the effectiveness of our proposed method on various vision-language tasks that benefit from improved visual relationship understanding. Jiuxiang Gu, Jason Kuen, Shafiq R. Joty, Jianfei Cai 0001, Vlad I. Morariu, Handong Zhao, Tong Sun 0005 |
NeurIPS | 5 |
| 2019 | Layout-Induced Video Representation for Recognizing Agent-in-Place ActionsabstractWe address scene layout modeling for recognizing agent-in-place actions, which are actions associated with agents who perform them and the places where they occur, in the context of outdoor home surveillance. We introduce a novel representation to model the geometry and topology of scene layouts so that a network can generalize from the layouts observed in the training scenes to unseen scenes in the test set. This Layout-Induced Video Representation (LIVR) abstracts away low-level appearance variance and encodes geometric and topological relationships of places to explicitly model scene layout. LIVR partitions the semantic features of a scene into different places to force the network to learn generic place-based feature descriptions which are independent of specific scene layouts; then, LIVR dynamically aggregates features based on connectivities of places in each specific scene to model its layout. We introduce a new Agent-in-Place Action (APA) dataset to show that our method allows neural network models to generalize significantly better to unseen scenes. Ruichi Yu, Ang Li 0001, Jingxiao Zheng, Vlad I. Morariu, Larry Davis 0001 |
ICCV | 5 |
| 2019 | Deep Splitting and Merging for Table Structure DecompositionabstractGiven the large variety and complexity of tables, table structure extraction is a challenging task in automated document analysis systems. We present a pair of novel deep learning models (Split and Merge models) that given an input image, 1) predicts the basic table grid pattern and 2) predicts which grid elements should be merged to recover cells that span multiple rows or columns. We propose projection pooling as a novel component of the Split model and grid pooling as a novel part of the Merge model. While most Fully Convolutional Networks rely on local evidence, these unique pooling regions allow our models to take advantage of the global table structure. We achieve state-of-the-art performance on the public ICDAR 2013 Table Competition dataset of PDF documents. On a much larger private dataset which we used to train the models, we significantly outperform both a state-ofthe-art deep model and a major commercial software system. Chris Tensmeyer, Vlad I. Morariu, Brian L. Price, Scott Cohen, Tony R. Martinez |
ICDAR | 2 |
| 2018 | Dynamic Zoom-In Network for Fast Object Detection in Large ImagesabstractWe introduce a generic framework that reduces the computational cost of object detection while retaining accuracy for scenarios where objects with varied sizes appear in high resolution images. Detection progresses in a coarse-to-fine manner, first on a down-sampled version of the image and then on a sequence of higher resolution regions identified as likely to improve the detection accuracy. Built upon reinforcement learning, our approach consists of a model (R-net) that uses coarse detection results to predict the potential accuracy gain for analyzing a region at a higher resolution and another model (Q-net) that sequentially selects regions to zoom in. Experiments on the Caltech Pedestrians dataset show that our approach reduces the number of processed pixels by over 50% without a drop in detection accuracy. The merits of our approach become more significant on a high resolution test set collected from YFCC100M dataset, where our approach maintains high detection performance while reducing the number of processed pixels by about 70% and the detection time by over 50%. Mingfei Gao, Ruichi Yu, Ang Li 0001, Vlad I. Morariu, Larry Davis 0001 |
CVPR | 4 |
| 2018 | Learning a Discriminative Filter Bank Within a CNN for Fine-Grained RecognitionabstractCompared to earlier multistage frameworks using CNN features, recent end-to-end deep approaches for fine-grained recognition essentially enhance the mid-level learning capability of CNNs. Previous approaches achieve this by introducing an auxiliary network to infuse localization information into the main classification network, or a sophisticated feature encoding method to capture higher order feature statistics. We show that mid-level representation learning can be enhanced within the CNN framework, by learning a bank of convolutional filters that capture class-specific discriminative patches without extra part or bounding box annotations. Such a filter bank is well structured, properly initialized and discriminatively learned through a novel asymmetric multi-stream architecture with convolutional filter supervision and a non-random layer initialization. Experimental results show that our approach achieves state-of-the-art on three publicly available fine-grained recognition datasets (CUB-200-2011, Stanford Cars and FGVC-Aircraft). Ablation studies and visualizations are provided to understand our approach. Yaming Wang, Vlad I. Morariu, Larry Davis 0001 |
CVPR | 2 |
| 2018 | NISP: Pruning Networks Using Neuron Importance Score PropagationabstractTo reduce the significant redundancy in deep Convolutional Neural Networks (CNNs), most existing methods prune neurons by only considering the statistics of an individual layer or two consecutive layers (e.g., prune one layer to minimize the reconstruction error of the next layer), ignoring the effect of error propagation in deep networks. In contrast, we argue that for a pruned network to retain its predictive power, it is essential to prune neurons in the entire neuron network jointly based on a unified goal: minimizing the reconstruction error of important responses in the "final response layer" (FRL), which is the second-to-last layer before classification. Specifically, we apply feature ranking techniques to measure the importance of each neuron in the FRL, formulate network pruning as a binary integer optimization problem, and derive a closed-form solution to it for pruning neurons in earlier layers. Based on our theoretical analysis, we propose the Neuron Importance Score Propagation (NISP) algorithm to propagate the importance scores of final responses to every neuron in the network. The CNN is pruned by removing neurons with least importance, and it is then fine-tuned to recover its predictive power. NISP is evaluated on several datasets with multiple CNN models and demonstrated to achieve significant acceleration and compression with negligible accuracy loss. Ruichi Yu, Ang Li 0001, Chun-Fu Chen 0001, Jui-Hsin Lai, Vlad I. Morariu, Xintong Han, Mingfei Gao, Ching-Yung Lin, Larry Davis 0001 |
CVPR | 5 |
| 2018 | Learning Rich Features for Image Manipulation DetectionabstractImage manipulation detection is different from traditional semantic object detection because it pays more attention to tampering artifacts than to image content, which suggests that richer features need to be learned. We propose a two-stream Faster R-CNN network and train it end-to-end to detect the tampered regions given a manipulated image. One of the two streams is an RGB stream whose purpose is to extract features from the RGB image input to find tampering artifacts like strong contrast difference, unnatural tampered boundaries, and so on. The other is a noise stream that leverages the noise features extracted from a steganalysis rich model filter layer to discover the noise inconsistency between authentic and tampered regions. We then fuse features from the two streams through a bilinear pooling layer to further incorporate spatial co-occurrence of these two modalities. Experiments on four standard image manipulation datasets demonstrate that our two-stream framework outperforms each individual stream, and also achieves state-of-the-art performance compared to alternative methods with robustness to resizing and compression. Peng Zhou 0009, Xintong Han, Vlad I. Morariu, Larry Davis 0001 |
CVPR | 3 |
| 2018 | C-WSL: Count-Guided Weakly Supervised Localization
Mingfei Gao, Ang Li 0001, Ruichi Yu, Vlad I. Morariu, Larry Davis 0001 |
ECCV (1) | 4 |
| 2017 | Generating Holistic 3D Scene Abstractions for Text-Based Image RetrievalabstractSpatial relationships between objects provide important information for text-based image retrieval. As users are more likely to describe a scene from a real world perspective, using 3D spatial relationships rather than 2D relationships that assume a particular viewing direction, one of the main challenges is to infer the 3D structure that bridges images with users text descriptions. However, direct inference of 3D structure from images requires learning from large scale annotated data. Since interactions between objects can be reduced to a limited set of atomic spatial relations in 3D, we study the possibility of inferring 3D structure from a text description rather than an image, applying physical relation models to synthesize holistic 3D abstract object layouts satisfying the spatial constraints present in a textual description. We present a generic framework for retrieving images from a textual description of a scene by matching images with these generated abstract object layouts. Images are ranked by matching object detection outputs (bounding boxes) to 2D layout candidates (also represented by bounding boxes) which are obtained by projecting the 3D scenes with sampled camera directions. We validate our approach using public indoor scene datasets and show that our method outperforms baselines built upon object occurrence histograms and learned 2D pairwise relations. Ang Li 0001, Jin Sun 0011, Joe Yue-Hei Ng, Ruichi Yu, Vlad I. Morariu, Larry Davis 0001 |
CVPR | 5 |
| 2017 | Generalized Deep Image to Image RegressionabstractWe present a Deep Convolutional Neural Network architecture which serves as a generic image-to-image regressor that can be trained end-to-end without any further machinery. Our proposed architecture, the Recursively Branched Deconvolutional Network (RBDN), develops a cheap multi-context image representation very early on using an efficient recursive branching scheme with extensive parameter sharing and learnable upsampling. This multi-context representation is subjected to a highly non-linear locality preserving transformation by the remainder of our network comprising of a series of convolutions/deconvolutions without any spatial downsampling. The RBDN architecture is fully convolutional and can handle variable sized images during inference. We provide qualitative/quantitative results on 3 diverse tasks: relighting, denoising and colorization and show that our proposed RBDN architecture obtains comparable results to the state-of-the-art on each of these tasks when used off-the-shelf without any post processing or task-specific architectural modifications. Venkataraman Santhanam, Vlad I. Morariu, Larry Davis 0001 |
CVPR | 2 |
| 2017 | Visual Relationship Detection with Internal and External Linguistic Knowledge DistillationabstractUnderstanding the visual relationship between two objects involves identifying the subject, the object, and a predicate relating them. We leverage the strong correlations between the predicate and the hsubj; obji pair (both semantically and spatially) to predict predicates conditioned on the subjects and the objects. Modeling the three entities jointly more accurately reflects their relationships compared to modeling them independently, but it complicates learning since the semantic space of visual relationships is huge and training data is limited, especially for longtail relationships that have few instances. To overcome this, we use knowledge of linguistic statistics to regularize visual model learning. We obtain linguistic knowledge by mining from both training annotations (internal knowledge) and publicly available text, e.g., Wikipedia (external knowledge), computing the conditional probability distribution of a predicate given a (subj, obj) pair. As we train the visual model, we distill this knowledge into the deep model to achieve better generalization. Our experimental results on the Visual Relationship Detection (VRD) and Visual Genome datasets suggest that with this linguistic knowledge distillation, our model outperforms the stateof- the-art methods significantly, especially when predicting unseen relationships (e.g., recall improved from 8.45% to 19.17% on VRD zero-shot testing set). Ruichi Yu, Ang Li 0001, Vlad I. Morariu, Larry Davis 0001 |
ICCV | 3 |
| 2017 | VRFP: On-the-Fly Video Retrieval Using Web Images and Fast Fisher Vector ProductsabstractOn-the-fly video retrieval using Web images and fast Fisher Vector products (VRFP) is a real-time video retrieval framework based on short text input queries, which obtains weakly labeled training images from the Web after the query is known. The retrieved Web images representing the query and each database video are treated as unordered collections of images, and each collection is represented using a single Fisher Vector built on CNN features. Our experiments show that a Fisher Vector is robust to noise present in Web images and compares favorably in terms of accuracy to other standard representations. While a Fisher Vector can be constructed efficiently for a new query, matching against the test set is slow due to its high dimensionality. To perform matching in real time, we present a lossless algorithm that accelerates the inner product computation between high-dimensional Fisher Vectors. We prove that the expected number of multiplications required decreases quadratically with the sparsity of Fisher Vectors. We are not only able to construct and apply query models in real time, but with the help of a simple reranking scheme, we also outperform state-of-the-art automatic retrieval methods by a significant margin on TRECVID MED13 (3.5%), MED14 (1.3%), and CCV datasets (5.2%). We also provide a direct comparison on standard datasets between two different paradigms for automatic video retrieval: zero-shot learning and on-the-fly retrieval. Xintong Han, Vlad I. Morariu, Larry Davis 0001 |
IEEE Trans. Multim. | 3 |
| 2016 | The Role of Context Selection in Object Detection
Ruichi Yu, Xi Chen 0016, Vlad I. Morariu, Larry Davis 0001 |
BMVC | 3 |
| 2016 | Mining Discriminative Triplets of Patches for Fine-Grained ClassificationabstractFine-grained classification involves distinguishing between similar sub-categories based on subtle differences in highly localized regions, therefore, accurate localization of discriminative regions remains a major challenge. We describe a patch-based framework to address this problem. We introduce triplets of patches with geometric constraints to improve the accuracy of patch localization, and automatically mine discriminative geometrically-constrained triplets for classification. The resulting approach only requires object bounding boxes. Its effectiveness is demonstrated using four publicly available fine-grained datasets, on which it outperforms or achieves comparable performance to the state-of-the-art in classification. Yaming Wang, Vlad I. Morariu, Larry Davis 0001 |
CVPR | 3 |
| 2016 | Modeling Context Between Objects for Referring Expression Understanding
Varun K. Nagaraja, Vlad I. Morariu, Larry Davis 0001 |
ECCV (4) | 2 |
| 2016 | A non-parametric approach to extending generic binary classifiers for multi-classification
Venkataraman Santhanam, Vlad I. Morariu, David Harwood, Larry Davis 0001 |
Pattern Recognit. | 2 |
| 2015 | Searching for Objects using Structure in Indoor ScenesabstractTo identify the location of objects of a particular class, a passive computer vision system generally processes all the regions in an image to finally output few regions. However, we can use structure in the scene to search for objects without processing the entire image. We propose a search technique that sequentially processes image regions such that the regions that are more likely to correspond to the query class object are explored earlier. We frame the problem as a Markov decision process and use an imitation learning algorithm to learn a search strategy. Since structure in the scene is essential for search, we work with indoor scene images as they contain both unary scene context information and object-object context in the scene. We perform experiments on the NYU-depth v2 dataset and show that the unary scene context features alone can achieve a significantly high average precision while processing only 20-25\% of the regions for classes like bed and sofa. By considering object-object context along with the scene context features, the performance is further improved for classes like counter, lamp, pillow and sofa. Varun K. Nagaraja, Vlad I. Morariu, Larry Davis 0001 |
BMVC | 2 |
| 2015 | Selective Encoding for Recognizing Unreliably Localized FacesabstractMost existing face verification systems rely on precise face detection and registration. However, these two components are fallible under unconstrained scenarios (e.g., mobile face authentication) due to partial occlusions, pose variations, lighting conditions and limited view-angle coverage of mobile cameras. We address the unconstrained face verification problem by encoding face images directly without any explicit models of detection or registration. We propose a selective encoding framework which injects relevance information (e.g., foreground/background probabilities) into each cluster of a descriptor codebook. An additional selector component also discards distractive image patches and improves spatial robustness. We evaluate our framework using Gaussian mixture models and Fisher vectors on challenging face verification datasets. We apply selective encoding to Fisher vector features, which in our experiments degrade quickly with inaccurate face localization, our framework improves robustness with no extra test time computation. We also apply our approach to mobile based active face authentication task, demonstrating its utility in real scenarios. Ang Li 0001, Vlad I. Morariu, Larry Davis 0001 |
ICCV | 2 |
| 2015 | Selecting Relevant Web Trained Concepts for Automated Event RetrievalabstractComplex event retrieval is a challenging research problem, especially when no training videos are available. An alternative to collecting training videos is to train a large semantic concept bank a priori. Given a text description of an event, event retrieval is performed by selecting concepts linguistically related to the event description and fusing the concept responses on unseen videos. However, defining an exhaustive concept lexicon and pre-training it requires vast computational resources. Therefore, recent approaches automate concept discovery and training by leveraging large amounts of weakly annotated web data. Compact visually salient concepts are automatically obtained by the use of concept pairs or, more generally, n-grams. However, not all visually salient n-grams are necessarily useful for an event query -- some combinations of concepts may be visually compact but irrelevant -- and this drastically affects performance. We propose an event retrieval algorithm that constructs pairs of automatically discovered concepts and then prunes those concepts that are unlikely to be helpful for retrieval. Pruning depends both on the query and on the specific video instance being evaluated. Our approach also addresses calibration and domain adaptation issues that arise when applying concept detectors to unseen videos. We demonstrate large improvements over other vision based systems on the TRECVID MED 13 dataset. Xintong Han, Zhe Wu 0001, Vlad I. Morariu, Larry Davis 0001 |
ICCV | 4 |
| 2015 | Clauselets: Leveraging Temporally Related Actions for Video Event AnalysisabstractWe propose clause lets, sets of concurrent actions and their temporal relationships, and explore their application to video event analysis. We train clause lets in two stages. We initially train first level clause let detectors that find a limited set of actions in particular qualitative temporal configurations based on Allen's interval relations. In the second stage, we apply the first level detectors to training videos, and discriminatively learn temporal patterns between activations that involve more actions over longer durations and lead to improved second level clause let models. We demonstrate the utility of clause lets by applying them to the task of "in-the-wild" video event recognition on the TRECVID MED 11 dataset. Not only do clause lets achieve state-of-the-art results on this task, but qualitative results suggest that they may also lead to semantically meaningful descriptions of videos in terms of detected actions and their temporal relationships. Hyungtae Lee, Vlad I. Morariu, Larry Davis 0001 |
WACV | 2 |
| 2015 | Unsupervised Feature Extraction Inspired by Latent Low-Rank RepresentationabstractLatent Low-Rank Representation (Lat LRR) has the empirical capability of identifying "salient" features. However, the reason behind this feature extraction effect is still not understood. Its optimization leads to non-unique solutions and has high computational complexity, limiting its potential in practice. We show that Lat LRR learns a transformation matrix which suppresses the most significant principal components corresponding to the largest singular values while preserving the details captured by the components with relatively smaller singular values. Based on this, we propose a novel feature extraction method which directly designs the transformation matrix and has similar behavior to Lat LRR. Our method has a simple analytical solution and can achieve better performance with little computational cost. The effectiveness and efficiency of our method are validated on two face recognition datasets. Yaming Wang, Vlad I. Morariu, Larry Davis 0001 |
WACV | 2 |
| 2014 | Planar Structure Matching under Projective Uncertainty for Geolocation
Ang Li 0001, Vlad I. Morariu, Larry Davis 0001 |
ECCV (7) | 2 |
| 2014 | Jointly Optimizing 3D Model Fitting and Fine-Grained Classification
Yen-Liang Lin, Vlad I. Morariu, Winston H. Hsu, Larry Davis 0001 |
ECCV (4) | 2 |
| 2014 | Interactive video segmentation using occlusion boundaries and temporally coherent superpixelsabstractWe propose an interactive video segmentation system built on the basis of occlusion and long term spatio-temporal structure cues. User supervision is incorporated in a superpixel graph clustering framework that differs crucially from prior art in that it modifies the graph according to the output of an occlusion boundary detector. Working with long temporal intervals (up to 100 frames) enables our system to significantly reduce annotation effort with respect to state of the art systems. Even though the segmentation results are less than perfect, they are obtained efficiently and can be used in weakly supervised learning from video or for video content description. We do not rely on a discriminative object appearance model and allow extracting multiple foreground objects together, saving user time if more than one object is present. Additional experiments with unsupervised clustering based on occlusion boundaries demonstrate the importance of this cue for video segmentation and thus validate our system design. Radu Dondera, Vlad I. Morariu, Yulu Wang, Larry Davis 0001 |
WACV | 2 |
| 2014 | Composite Discriminant Factor analysisabstractWe propose a linear dimensionality reduction method, Composite Discriminant Factor (CDF) analysis, which searches for a discriminative but compact feature subspace that can be used as input to classifiers that suffer from problems such as multi-collinearity or the curse of dimensionality. The subspace selected by CDF maximizes the performance of the entire classification pipeline, and is chosen from a set of candidate subspaces that are each discriminative. Our method is based on Partial Least Squares (PLS) analysis, and can be viewed as a generalization of the PLS1 algorithm, designed to increase discrimination in classification tasks. We demonstrate our approach on the UCF50 action recognition dataset, two object detection datasets (INRIA pedestrians and vehicles from aerial imagery), and machine learning datasets from the UCI Machine Learning repository. Experimental results show that the proposed approach improves significantly in terms of accuracy over linear SVM, and also over PLS in terms of compactness and efficiency, while maintaining or improving accuracy. Vlad I. Morariu, Ejaz Ahmed 0002, Venkataraman Santhanam, David Harwood, Larry Davis 0001 |
WACV | 1 |
| 2013 | Sampling for unsupervised domain adaptive object detectionabstractWe explore the problem of extreme class imbalance present when performing fully unsupervised domain adaptation for object detection. The main challenge arises from the fact that images in unconstrained settings are mostly occupied by the background (negative class). Therefore, random sampling will not typically result in a sufficient number of positive samples from the target domain, which is required by domain adaptation methods. Motivated by traditional semi-supervised learning algorithms that aim for better classification using both labeled and unlabeled data, we propose a variation of co-learning technique that automatically constructs a more balanced set of samples from the target domain. We evaluate the effectiveness of our approach using a vehicle detection task in an urban surveillance dataset. Furthermore, we compare the performance of our technique with two other approaches-one based on unbiased learning on multiple training data sets and the other on self-learning. Fatemeh Mirrashed, Vlad I. Morariu, Larry Davis 0001 |
ICIP | 2 |
| 2013 | Domain adaptive object detectionabstractWe study the use of domain adaptation and transfer learning techniques as part of a framework for adaptive object detection. Unlike recent applications of domain adaptation work in computer vision, which generally focus on image classification, we explore the problem of extreme class imbalance present when performing domain adaptation for object detection. The main difficulty caused by this imbalance is that test images contain millions or billions of negative image subwindows but just a few image subwindows containing positive instances, which makes it difficult to adapt to changes in the positive classes present new domains by simple techniques such as random sampling. We propose an initial approach to addressing this problem and apply our technique to vehicle detection in a challenging urban surveillance dataset, demonstrating the performance of our approach with various amounts of supervision, including the fully unsupervised case. Fatemeh Mirrashed, Vlad I. Morariu, Behjat Siddiquie, Rogério Feris, Larry Davis 0001 |
WACV | 2 |
| 2013 | Tracking People's Hands and Feet Using Mixed Network AND/OR SearchabstractWe describe a framework that leverages mixed probabilistic and deterministic networks and their AND/OR search space to efficiently find and track the hands and feet of multiple interacting humans in 2D from a single camera view. Our framework detects and tracks multiple people's heads, hands, and feet through partial or full occlusion; requires few constraints (does not require multiple views, high image resolution, knowledge of performed activities, or large training sets); and makes use of constraints and AND/OR Branch-and-Bound with lazy evaluation and carefully computed bounds to efficiently solve the complex network that results from the consideration of interperson occlusion. Our main contributions are: 1) a multiperson part-based formulation that emphasizes extremities and allows for the globally optimal solution to be obtained in each frame, and 2) an efficient and exact optimization scheme that relies on AND/OR Branch-and-Bound, lazy factor evaluation, and factor cost sensitive bound computation. We demonstrate our approach on three datasets: the public single person HumanEva dataset, outdoor sequences where multiple people interact in a group meeting scenario, and outdoor one-on-one basketball videos. The first dataset demonstrates that our framework achieves state-of-the-art performance in the single person setting, while the last two demonstrate robustness in the presence of partial and full occlusion and fast nontrivial motion. Vlad I. Morariu, David Harwood, Larry Davis 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2012 | Qualitative Pose Estimation by Discriminative Deformable Part Models
Hyungtae Lee, Vlad I. Morariu, Larry Davis 0001 |
ACCV (2) | 2 |
| 2012 | A flow model for joint action recognition and identity maintenanceabstractWe propose a framework that performs action recognition and identity maintenance of multiple targets simultaneously. Instead of first establishing tracks using an appearance model and then performing action recognition, we construct a network flow-based model that links detected bounding boxes across video frames while inferring activities, thus integrating identity maintenance and action recognition. Inference in our model reduces to a constrained minimum cost flow problem, which we solve exactly and efficiently. By leveraging both appearance similarity and action transition likelihoods, our model improves on state-of-the-art results on action recognition for two datasets. Sameh Khamis, Vlad I. Morariu, Larry Davis 0001 |
CVPR | 2 |
| 2012 | Combining Per-frame and Per-track Cues for Multi-person Action Recognition
Sameh Khamis, Vlad I. Morariu, Larry Davis 0001 |
ECCV (1) | 2 |
| 2011 | Multi-agent event recognition in structured scenariosabstractWe present a framework for the automatic recognition of complex multi-agent events in settings where structure is imposed by rules that agents must follow while performing activities. Given semantic spatio-temporal descriptions of what generally happens (i.e., rules, event descriptions, physical constraints), and based on video analysis, we determine the events that occurred. Knowledge about spatio-temporal structure is encoded using first-order logic using an approach based on Allen's Interval Logic, and robustness to low-level observation uncertainty is provided by Markov Logic Networks (MLN). Our main contribution is that we integrate interval-based temporal reasoning with probabilistic logical inference, relying on an efficient bottom-up grounding scheme to avoid combinatorial explosion. Applied to one-on-one basketball, our framework detects and tracks players, their hands and feet, and the ball, generates event observations from the resulting trajectories, and performs probabilistic logical inference to determine the most consistent sequence of events. We demonstrate our approach on 1hr (100,000 frames) of outdoor videos. Vlad I. Morariu, Larry Davis 0001 |
CVPR | 1 |
| 2011 | Birdlets: Subordinate categorization using volumetric primitives and pose-normalized appearanceabstractSubordinate-level categorization typically rests on establishing salient distinctions between part-level characteristics of objects, in contrast to basic-level categorization, where the presence or absence of parts is determinative. We develop an approach for subordinate categorization in vision, focusing on an avian domain due to the fine-grained structure of the category taxonomy for this domain. We explore a pose-normalized appearance model based on a volumetric poselet scheme. The variation in shape and appearance properties of these parts across a taxonomy provides the cues needed for subordinate categorization. Training pose detectors requires a relatively large amount of training data per category when done from scratch; using a subordinate-level approach, we exploit a pose classifier trained at the basic-level, and extract part appearance and shape information to build subordinate-level models. Our model associates the underlying image pattern parameters used for detection with corresponding volumetric part location, scale and orientation parameters. These parameters implicitly define a mapping from the image pixels into a pose-normalized appearance space, removing view and pose dependencies, facilitating fine-grained categorization from relatively few training examples. Ryan Farrell, Om Oza, Ning Zhang 0014, Vlad I. Morariu, Trevor Darrell, Larry Davis 0001 |
ICCV | 4 |
| 2008 | Automatic online tuning for fast Gaussian summationabstractMany machine learning algorithms require the summation of Gaussian kernel functions, an expensive operation if implemented straightforwardly. Several methods have been proposed to reduce the computational complexity of evaluating such sums, including tree and analysis based methods. These achieve varying speedups depending on the bandwidth, dimension, and prescribed error, making the choice between methods difficult for machine learning tasks. We provide an algorithm that combines tree methods with the Improved Fast Gauss Transform (IFGT). As originally proposed the IFGT suffers from two problems: (1) the Taylor series expansion does not perform well for very low bandwidths, and (2) parameter selection is not trivial and can drastically affect performance and ease of use. We address the first problem by employing a tree data structure, resulting in four evaluation methods whose performance varies based on the distribution of sources and targets and input parameters such as desired accuracy and bandwidth. To solve the second problem, we present an online tuning approach that results in a black box method that automatically chooses the evaluation method and its parameters to yield the best performance for the input data, desired accuracy, and bandwidth. In addition, the new IFGT parameter selection approach allows for tighter error bounds. Our approach chooses the fastest method at negligible additional cost, and has superior performance in comparisons with previous approaches. Vlad I. Morariu, Balaji Vasan Srinivasan, Vikas C. Raykar, Ramani Duraiswami, Larry Davis 0001 |
NIPS | 1 |
| 2007 | Using vision, acoustics, and natural language for disambiguationabstractCreating a human-robot interface is a daunting experience. Capabilities and functionalities of the interface are dependent on the robustness of many different sensor and input modalities. For example, object recognition poses problems for state-of-the-art vision systems. Speech recognition in noisy environments remains problematic for acoustic systems. Natural language understanding and dialog are often limited to specific domains and baffled by ambiguous or novel utterances. Plans based on domain-specific tasks limit the applicability of dialog managers. The types of sensors used limit spatial knowledge and understanding, and constrain cognitive issues, such as perspective-taking.In this research, we are integrating several modalities, such as vision, audition, and natural language understanding to leverage the existing strengths of each modality and overcome individual weaknesses. We are using visual, acoustic, and linguistic inputs in various combinations to solve such problems as the disambiguation of referents (objects in the environment), localization of human speakers, and determination of the source of utterances and appropriateness of responses when humans and robots interact. For this research, we limit our consideration to the interaction of two humans and one robot in a retrieval scenario. This paper will describe the system and integration of the various modules prior to future testing. Benjamin R. Fransen, Vlad I. Morariu, Eric Martinson, Samuel Blisard, Matthew Marge, Scott Thomas, Alan C. Schultz, Dennis Perzanowski |
HRI | 2 |
| 2006 | Dynamic Appearance Modeling for Human TrackingabstractDynamic appearance is one of the most important cues for tracking and identifying moving people. However, direct modeling spatio-temporal variations of such appearance is often a difficult problem due to their high dimensionality and nonlinearities. In this paper we present a human tracking system that uses a dynamic appearance and motion modeling framework based on the use of robust system dynamics identification and nonlinear dimensionality reduction techniques. The proposed system learns dynamic appearance and motion models from a small set of initial frames and does not require prior knowledge such as gender or type of activity. The advantages of the proposed tracking system are illustrated with several examples where the learned dynamics accurately predict the location and appearance of the targets in future frames, preventing tracking failures due to model drifting, target occlusion and scene clutter. Hwasup Lim, Octavia I. Camps, Mario Sznaier, Vlad I. Morariu |
CVPR (1) | 4 |
| 2006 | Modeling Correspondences for Multi-Camera Tracking Using Nonlinear Manifold Learning and Target DynamicsabstractMulti-camera tracking systems often must maintain consistent identity labels of the targets across views to recover 3D trajectories and fully take advantage of the additional information available from the multiple sensors. Previous approaches to the "correspondence across views" problem include matching features, using camera calibration information, and computing homographies between views under the assumption that the world is planar. However, it can be difficult to match features across significantly different views. Furthermore, calibration information is not always available and planar world hypothesis can be too restrictive. In this paper, a new approach is presented for matching correspondences based on the use of nonlinear manifold learning and system dynamics identification. The proposed approach does not require similar views, calibration nor geometric assumptions of the 3D environment, and is robust to noise and occlusion. Experimental results demonstrate the use of this approach to generate and predict views in cases where identity labels become ambiguous. Vlad I. Morariu, Octavia I. Camps |
CVPR (1) | 1 |