VLDB 2026 Research / reviewers in the wild / expert
Yezhou Yang
dblp:78/7455
· DBLP profile ↗
88ranked-venue papers
8as first author
41since 2021 · last 2026
0000-0003-0126-8976ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 69 · 7 first-author · 34 since 2021Graphics, computer vision, multimedia, augmented reality and games · 50 · 5 first-author · 25 since 2021Systems, architecture and hardware · 14 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VOCAL: Visual Odometry via ContrAstive LearningabstractBreakthroughs in visual odometry (VO) have fundamentally reshaped the landscape of robotics, enabling ultra-precise camera state estimation that is crucial for modern autonomous systems. Despite these advances, many learning-based VO techniques rely on rigid geometric assumptions, which often fall short in interpretability and lack a solid theoretical basis within fully data-driven frameworks. To overcome these limitations, we introduce VOCAL (Visual Odometry via ContrAstive Learning), a novel framework that reimagines VO as a label ranking challenge. By integrating Bayesian inference with a representation learning framework, VOCAL organizes visual features to mirror camera states. The ranking mechanism compels similar camera states to converge into consistent and spatially coherent representations within the latent space. This strategic alignment not only bolsters the interpretability of the learned features but also ensures compatibility with multimodal data sources. Extensive evaluations on the KITTI dataset highlight VOCAL's enhanced interpretability and flexibility, pushing VO toward more general and explainable spatial intelligence. Chi-Yao Huang, Zeel Bhatt, Yezhou Yang |
WACV | 3 |
| 2026 | Event-based Graph Representation with Spatial and Motion Vectors for Asynchronous Object DetectionabstractEvent-based sensors offer high temporal resolution and low latency by generating sparse, asynchronous data. However, converting this irregular data into dense tensors for use in standard neural networks diminishes these inherent advantages, motivating research into graph representations. While such methods preserve sparsity and support asynchronous inference, their performance on downstream tasks remains limited due to suboptimal modeling of spatiotemporal dynamics. In this work, we propose a novel spatiotemporal multigraph representation to better capture spatial structure and temporal changes. Our approach constructs two decoupled graphs: a spatial graph leveraging B-spline basis functions to model global structure, and a temporal graph utilizing motion vector-based attention for local dynamic changes. This design enables the use of efficient 2D kernels in place of computationally expensive 3D kernels. We evaluate our method on the Gen1 automotive and eTraM datasets for event-based object detection, achieving over a 6% improvement in detection accuracy compared to previous graph-based works, with a 5× speedup, reduced parameter count, and no increase in computational cost. These results highlight the effectiveness of structured graph modeling for asynchronous vision. Project page: eventbasedvision.github.io/eGSMV. Aayush Atul Verma, Arpitsinh Vaghela, Bharatesh Chakravarthi, Kaustav Chanda, Yezhou Yang |
WACV | 5 |
| 2025 | AcT2I: Evaluating and Improving Action Depiction in Text-to-Image ModelsabstractText-to-Image (T2I) models have recently achieved remarkable success in generating images from textual descriptions.However, challenges still persist in accurately rendering complex scenes where actions and interactions form the primary semantic focus.Our key observation in this work is that T2I models frequently struggle to capture nuanced and often implicit attributes inherent in action depiction, leading to generating images that lack key contextual details.To enable systematic evaluation, we introduce AcT2I, a benchmark designed to evaluate the performance of T2I models in generating images from action-centric prompts.We experimentally validate that leading T2I models do not fare well on AcT2I.We further hypothesize that this shortcoming arises from the incomplete representation of the inherent attributes and contextual dependencies in the training corpora of existing T2I models.We build upon this by developing a trainingfree, knowledge distillation technique utilizing Large Language Models to address this limitation.Specifically, we enhance prompts by incorporating dense information across three dimensions, observing that injecting prompts with temporal details significantly improves image generation accuracy, with our best model achieving an increase of 72%.Our findings highlight the limitations of current T2I methods in generating images that require complex reasoning and demonstrate that integrating linguistic knowledge in a systematic way can notably advance the generation of nuanced and contextually accurate images. Vatsal Malaviya, Agneet Chatterjee, Maitreya Patel, Yezhou Yang, Chitta Baral |
EMNLP | 4 |
| 2025 | FlowChef: Steering of Rectified Flow Models for Controlled Generations
Maitreya Patel, Song Wen 0001, Dimitris N. Metaxas, Yezhou Yang |
ICCV | 4 |
| 2025 | RefEdit: A Benchmark and Method for Improving Instruction-Based Image Editing Model on Referring ExpressionsabstractDespite recent advances in inversion and instruction-based image editing, existing approaches primarily excel at editing single, prominent objects but significantly struggle when applied to complex scenes containing multiple entities. To quantify this gap, we first introduce RefEdit-Bench, a rigorous real-world benchmark rooted in RefCOCO, where even baselines trained on millions of samples perform poorly. To overcome this limitation, we introduce RefEdit -- an instruction-based editing model trained on our scalable synthetic data generation pipeline. Our RefEdit, trained on only 20,000 editing triplets, outperforms the Flux/SD3 model-based baselines trained on millions of data. Extensive evaluations across various benchmarks demonstrate that our model not only excels in referring expression tasks but also enhances performance on traditional benchmarks, achieving state-of-the-art results comparable to closed-source methods. We release data \& checkpoint for reproducibility. Bimsara Pathiraja, Maitreya Patel, Shivam Singh, Yezhou Yang, Chitta Baral |
ICCV | 4 |
| 2025 | VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical ReasoningabstractMultimodal Large Language Models (MLLMs) have become a powerful tool for integrating visual and textual information. Despite their exceptional performance on visual understanding benchmarks, measuring their ability to reason abstractly across multiple images remains a significant challenge. To address this, we introduce VOILA, a large-scale, open-ended, dynamic benchmark designed to evaluate MLLMs' perceptual understanding and abstract relational reasoning. VOILA employs an analogical mapping approach in the visual domain, requiring models to generate an image that completes an analogy between two given image pairs, reference and application, without relying on predefined choices. Our experiments demonstrate that the analogical reasoning tasks in VOILA present a challenge to MLLMs. Through multi-step analysis, we reveal that current MLLMs struggle to comprehend inter-image relationships and exhibit limited capabilities in high-level relational reasoning. Notably, we observe that performance improves when following a multi-step strategy of least-to-most prompting. Comprehensive evaluations on open-source models and GPT-4o show that on text-based answers, the best accuracy for challenging scenarios is 13% (LLaMa 3.2) and even for simpler tasks is only 29% (GPT-4o), while human performance is significantly higher at 70% across both difficulty levels. Nilay Yilmaz, Maitreya Patel, Yiran Luo 0001, Tejas Gokhale, Chitta Baral, Suren Jayasuriya, Yezhou Yang |
ICLR | 7 |
| 2025 | DeepShade: Enable Shade Simulation by Text-conditioned Image GenerationabstractHeatwaves pose a significant threat to public health, especially as global warming intensifies. However, current routing systems (e.g., online maps) fail to incorporate shade information due to the difficulty of estimating shades directly from noisy satellite imagery and the limited availability of training data for generative models. In this paper, we address these challenges through two main contributions. First, we build an extensive dataset covering diverse longitude-latitude regions, varying levels of building density, and different urban layouts. Leveraging Blender-based 3D simulations alongside building outlines, we capture building shadows under various solar zenith angles throughout the year and at different times of day. These simulated shadows are aligned with satellite images, providing a rich resource for learning shade patterns. Second, we propose the DeepShade, a diffusion-based model designed to learn and synthesize shade variations over time. It emphasizes the nuance of edge features by jointly considering RGB with the Canny edge layer, and incorporates contrastive learning to capture the temporal change rules of shade. Then, by conditioning on textual descriptions of known conditions (e.g., time of day, solar angles), our framework provides improved performance in generating shade images. We demonstrate the utility of our approach by using our shade predictions to calculate shade ratios for real-world route planning in Tempe, Arizona. We believe this work will benefit society by providing a reference for urban planning in extreme heat weather and its potential practical applications in the environment. Longchao Da, Xiangrui Liu, Mithun Shivakoti, Thirulogasankar Pranav Kutralingam, Yezhou Yang, Hua Wei 0001 |
IJCAI | 5 |
| 2025 | Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video GenerationabstractRecent advances in video generation have enabled high-fidelity video synthesis from user provided prompts. However, existing models and benchmarks fail to capture the complexity and requirements of professional video generation. Towards that goal, we introduce Stable Cinemetrics, a structured evaluation framework that formalizes filmmaking controls into four disentangled, hierarchical taxonomies: Setup, Event, Lighting, and Camera. Together, these taxonomies define 76 fine-grained control nodes grounded in industry practices. Using these taxonomies, we construct a benchmark of prompts aligned with professional use cases and develop an automated pipeline for prompt categorization and question generation, enabling independent evaluation of each control dimension. We conduct a large-scale human study spanning 10+ models and 20K videos, annotated by a pool of 80+ film professionals. Our analysis, both coarse and fine-grained reveal that even the strongest current models exhibit significant gaps, particularly in Events and Camera-related controls. To enable scalable evaluation, we train an automatic evaluator, a vision-language model aligned with expert annotations that outperforms existing zero-shot baselines. SCINE is the first approach to situate professional video generation within the landscape of video generative models, introducing taxonomies centered around cinematic controls and supporting them with structured evaluation pipelines and detailed analyses to guide future research. Agneet Chatterjee, Rahim Entezari, Maksym Zhuravinskyi, Maksim Lapin, Reshinth Adithyan, Amit Raj, Chitta Baral, Yezhou Yang, Varun Jampani |
NeurIPS | 8 |
| 2025 | EraseFlow: Learning Concept Erasure Policies via GFlowNet-Driven AlignmentabstractErasing harmful or proprietary concepts from powerful text‑to‑image generators is an emerging safety requirement, yet current ``concept erasure'' techniques either collapse image quality, rely on brittle adversarial losses, or demand prohibitive retraining cycles. We trace these limitations to a myopic view of the denoising trajectories that govern diffusion‑based generation. We introduce EraseFlow, the first framework that casts concept unlearning as exploration in the space of denoising paths and optimizes it with a GFlowNets equipped with the trajectory‑balance objective. By sampling entire trajectories rather than single end states, EraseFlow learns a stochastic policy that steers generation away from target concepts while preserving the model’s prior. EraseFlow eliminates the need for carefully crafted reward models and by doing this, it generalizes effectively to unseen concepts and avoids hackable rewards while improving the performance. Extensive empirical results demonstrate that EraseFlow outperforms existing baselines and achieves an optimal trade-off between performance and prior preservation. Abhiram Kusumba, Maitreya Patel, Kyle Min 0001, Changhoon Kim, Chitta Baral, Yezhou Yang |
NeurIPS | 6 |
| 2025 | Deep Geometric Moments Promote Shape Consistency in Text-to-3D GenerationabstractTo address the data scarcity associated with 3D assets, 2D-lifting techniques such as Score Distillation Sampling (SDS) have become a widely adopted practice in text-to-3D generation pipelines. However, the diffusion models used in these techniques are prone to viewpoint bias and thus lead to geometric inconsistencies such as the Janus problem. To counter this, we introduce MT3D, a text-to-3D generative model that leverages a high-fidelity 3D object to overcome viewpoint bias and explicitly infuse geometric understanding into the generation pipeline. Firstly, we employ depth maps derived from a high-quality 3D model as control signals to guarantee that the generated 2D images preserve the funda-mental shape and structure, thereby reducing the inherent viewpoint bias. Next, we utilize deep geometric moments to ensure geometric consistency in the 3D representation explicitly. By incorporating geometric details from a 3D asset, MT3D enables the creation of diverse and geometri-cally consistent objects, thereby improving the quality and usability of our 3D representations. Project page and code: https://moment-3d.github.io/ Utkarsh Nath, Rajeev Goel, Eun Som Jeon, Changhoon Kim, Kyle Min 0001, Yezhou Yang, Yingzhen Yang, Pavan Turaga |
WACV | 6 |
| 2024 | ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion ModelsabstractThe ability to understand visual concepts and replicate and compose these concepts from images is a central goal for computer vision. Recent advances in text-to-image (T2I) models have lead to high definition and realistic image quality generation by learning from large databases of images and their descriptions. However, the evaluation of T2I models has focused on photorealism and limited qualitative measures of visual understanding. To quantify the ability of T2I models in learning and synthesizing novel visual concepts (a.k.a. personalized T2I), we introduce ConceptBed, a large-scale dataset that consists of 284 unique visual concepts, and 33K composite text prompts. Along with the dataset, we propose an evaluation metric, Concept Confidence Deviation (CCD), that uses the confidence of oracle concept classifiers to measure the alignment between concepts generated by T2I generators and concepts contained in target images. We evaluate visual concepts that are either objects, attributes, or styles, and also evaluate four dimensions of compositionality: counting, attributes, relations, and actions. Our human study shows that CCD is highly correlated with human understanding of concepts. Our results point to a trade-off between learning the concepts and preserving the compositionality which existing approaches struggle to overcome. The data, code, and interactive demo is available at: https://conceptbed.github.io/ Maitreya Patel, Tejas Gokhale, Chitta Baral, Yezhou Yang |
AAAI | 4 |
| 2024 | On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth EstimationabstractRecent advances in monocular depth estimation have been made by incorporating natural language as additional guidance. Although yielding impressive results, the impact of the language prior, particularly in terms of generalization and robustness, remains unexplored. In this paper, we address this gap by quantifying the impact of this prior and introduce methods to benchmark its effectiveness across various settings. We generate “low-level” sentences that convey object-centric, three-dimensional spatial relationships, incorporate them as additional language priors and evaluate their downstream impact on depth estimation. Our key finding is that current language-guided depth estimators perform optimally only with scene-level descriptions and counter-intuitively fare worse with low level descriptions. Despite leveraging additional data, these methods are not robust to directed adversarial attacks and decline in performance with an increase in distribution shift. Finally, to provide a foundation for future research, we identify points of failures and offer insights to better understand these shortcomings. With an increasing number of methods using language for depth estimation, our findings highlight the opportunities and pitfalls that require careful consideration for effective deployment in real-world settings.11Code/Data: https://github.com/agneet42/lang_depth Agneet Chatterjee, Tejas Gokhale, Chitta Baral, Yezhou Yang |
CVPR | 4 |
| 2024 | WOUAF: Weight Modulation for User Attribution and Fingerprinting in Text-to-Image Diffusion ModelsabstractThe rapid advancement of generative models, facilitating the creation of hyper-realistic images from textual de-scriptions, has concurrently escalated critical societal con-cerns such as misinformation. Although providing some mitigation, traditional fingerprinting mechanisms fall short in attributing responsibility for the malicious use of syn-thetic images. This paper introduces a novel approach to model fingerprinting that assigns responsibility for the gen-erated images, thereby serving as a potential countermea-sure to model misuse. Our method modifies generative mod-els based on each user's unique digital fingerprint, imprinting a unique identifier onto the resultant content that can be traced back to the user. This approach, incorporating fine-tuning into Text-to-Image (T2I) tasks using the Stable Diffusion Model, demonstrates near-perfect attribution ac-curacy with a minimal impact on output quality. Through extensive evaluation, we show that our method outperforms baseline methods with an average improvement of 11 % in handling image post-processes. Our method presents a promising and novel avenue for accountable model distribution and responsible use. Our code is available in https://github.com/kylemin/WQUAF. Changhoon Kim, Kyle Min 0001, Maitreya Patel, Yezhou Yang |
CVPR | 5 |
| 2024 | ECLIPSE: A Resource-Efficient Text-to-Image Prior for Image GenerationsabstractText-to-image (T2I) diffusion models, notably the unCLIP models (e.g., DALL-E-2), achieve state-of-the-art (SOTA) performance on various compositional T2I benchmarks, at the cost of significant computational resources. The unCLIP stack comprises T2I prior and diffusion image decoder. The T2I prior model alone adds a billion parameters compared to the Latent Diffusion Models, which increases the computational and high-quality data requirements. We introduce ECLIPSE11Our strategy, ECLIPSE, draws an analogy from the way a smaller prior model, akin to a celestial entity, offers a glimpse of the grandeur within the larger pre-trained vision-language model, mirroring how an eclipse reveals the vastness of the cosmos., a novel contrastive learning method that is both parameter and dataefficient. ECLIPSE leverages pre-trained vision-language models (e.g., CLIP) to distill the knowledge into the prior model. We demonstrate that the ECLIPSE trained prior, with only 3.3% of the parameters and trained on a mere 2.8% of the data, surpasses the baseline T2I priors with an average of 71.6% preference score under resource-limited setting. It also attains performance on par with SOTA big models, achieving an average of 63.36% preference score in terms of the ability to follow the text compositions. Extensive experiments on two unCLIP diffusion image decoders, Karlo and Kandinsky, affirm that ECLIPSE priors consistently deliver high performance while significantly reducing resource dependency. Project page: https://eclipse-t2i.vercel.app/ Maitreya Patel, Changhoon Kim, Chitta Baral, Yezhou Yang |
CVPR | 5 |
| 2024 | eTraM: Event-Based Traffic Monitoring DatasetabstractEvent cameras, with their high temporal and dynamic range and minimal memory usage, have found applications in various fields. However, their potential in static traffic monitoring remains largely unexplored. To facil-itate this exploration, we present eTraM - a first-of-its- kind, fully event-based traffic monitoring dataset. eTraM offers 10 hr of data from different traffic scenarios in various lighting and weather conditions, providing a compre-hensive overview of real-world situations. Providing 2M bounding box annotations, it covers eight distinct classes of traffic participants, ranging from vehicles to pedestri-ans and micro-mobility. eTraM's utility has been assessed using state-of-the-art methods for traffic participant detection, including RVT, RED, and YOLOv8. We quantitatively evaluate the ability of event-based models to gener-alize on nighttime and unseen scenes. Our findings sub-stantiate the compelling potential of leveraging event cam-eras for traffic monitoring, opening new avenues for research and application. eTraM is available at https://eventbasedvision.github.io/eTraM. Aayush Atul Verma, Bharatesh Chakravarthi, Arpitsinh Vaghela, Hua Wei 0001, Yezhou Yang |
CVPR | 5 |
| 2024 | REVISION: Rendering Tools Enable Spatial Fidelity in Vision-Language Models
Agneet Chatterjee, Yiran Luo 0001, Tejas Gokhale, Yezhou Yang, Chitta Baral |
ECCV (30) | 4 |
| 2024 | Getting it Right: Improving Spatial Consistency in Text-to-Image Models
Agneet Chatterjee, Gabriela Ben Melech Stan, Estelle Aflalo, Sayak Paul, Dhruba Ghosh, Tejas Gokhale, Ludwig Schmidt, Hannaneh Hajishirzi, Vasudev Lal, Chitta Baral, Yezhou Yang |
ECCV (22) | 11 |
| 2024 | R.A.C.E. : Robust Adversarial Concept Erasure for Secure Text-to-Image Diffusion Model
Changhoon Kim, Kyle Min 0001, Yezhou Yang |
ECCV (83) | 3 |
| 2024 | TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language NegativesabstractContrastive Language-Image Pretraining (CLIP) models maximize the mutual information between text and visual modalities to learn representations. This makes the nature of the training data a significant factor in the efficacy of CLIP for downstream tasks. However, the lack of compositional diversity in contemporary image-text datasets limits the compositional reasoning ability of CLIP. We show that generating ``hard'' negative captions via in-context learning and synthesizing corresponding negative images with text-to-image generators offers a solution. We introduce a novel contrastive pre-training strategy that leverages these hard negative captions and images in an alternating fashion to train CLIP. We demonstrate that our method, named TripletCLIP, when applied to existing datasets such as CC3M and CC12M, enhances the compositional capabilities of CLIP, resulting in an absolute improvement of over 9% on the SugarCrepe benchmark on an equal computational budget, as well as improvements in zero-shot image classification and image retrieval. Our code, models, and data are available at: tripletclip.github.io. Maitreya Patel, Abhiram Kusumba, Changhoon Kim, Tejas Gokhale, Chitta Baral, Yezhou Yang |
NeurIPS | 7 |
| 2024 | Towards Addressing the Misalignment of Object Proposal Evaluation for Vision-Language Tasks via Semantic GroundingabstractObject proposal generation serves as a standard preprocessing step in Vision-Language (VL) tasks (image captioning, visual question answering, etc.). The performance of object proposals generated for VL tasks is currently evaluated across all available annotations, a protocol that we show is "misaligned" - higher scores do not necessarily correspond to improved performance on downstream VL tasks. Our work serves as a study of this phenomenon and explores the effectiveness of semantic grounding to mitigate its effects. To this end, we propose evaluating object proposals against only a subset of available annotations, selected by thresholding an annotation importance score. Importance of object annotations to VL tasks is quantified by extracting relevant semantic information from text describing the image. We show that our method is consistent and demonstrates greatly improved alignment with annotations selected by image captioning metrics and human annotation when compared against existing techniques. Lastly, we compare current detectors used in the Scene Graph Generation (SGG) benchmark as a use case, which serves as an example of when traditional object proposal evaluation techniques are misaligned1. Joshua Feinglass, Yezhou Yang |
WACV | 2 |
| 2023 | End-to-end Knowledge Retrieval with Multi-modal QueriesabstractWe investigate knowledge retrieval with multimodal queries, i.e. queries containing information split across image and text inputs, a challenging task that differs from previous work on cross-modal retrieval.We curate a new dataset called ReMuQ 1 for benchmarking progress on this task.ReMuQ requires a system to retrieve knowledge from a large corpus by integrating contents from both text and image queries.We introduce a retriever model "ReViz" that can directly process input text and images to retrieve relevant knowledge in an end-to-end fashion without being dependent on intermediate modules such as object detectors or caption generators.We introduce a new pretraining task that is effective for learning knowledge retrieval with multimodal queries and also improves performance on downstream tasks.We demonstrate superior performance in retrieval on two datasets (ReMuQ and OK-VQA) under zeroshot settings as well as further improvements when finetuned on these datasets. Man Luo 0003, Zhiyuan Fang, Tejas Gokhale, Yezhou Yang, Chitta Baral |
ACL (1) | 4 |
| 2023 | Adversarial Bayesian Augmentation for Single-Source Domain GeneralizationabstractGeneralizing to unseen image domains is a challenging problem primarily due to the lack of diverse training data, inaccessible target data, and the large domain shift that may exist in many real-world settings. As such data augmentation is a critical component of domain generalization methods that seek to address this problem. We present Adversarial Bayesian Augmentation (ABA), a novel algorithm that learns to generate image augmentations in the challenging single-source domain generalization setting. ABA draws on the strengths of adversarial learning and Bayesian neural networks to guide the generation of diverse data augmentations –these synthesized image domains aid the classifier in generalizing to unseen domains. We demonstrate the strength of ABA on several types of domain shift including style shift, subpopulation shift, and shift in the medical imaging setting. ABA outperforms all previous state-of-the-art methods, including pre-specified augmentations, pixel-based and convolutional-based augmentations. Code: https://github.com/shengcheng/ABA. Tejas Gokhale, Yezhou Yang |
ICCV | 3 |
| 2023 | Attributing Image Generative Models using Latent FingerprintsabstractGenerative models have enabled the creation of contents that are indistinguishable from those taken from nature. Open-source development of such models raised concerns about the risks of their misuse for malicious purposes. One potential risk mitigation strategy is to attribute generative models via fingerprinting. Current fingerprinting methods exhibit a significant tradeoff between robust attribution accuracy and generation quality while lacking design principles to improve this tradeoff. This paper investigates the use of latent semantic dimensions as fingerprints, from where we can analyze the effects of design variables, including the choice of fingerprinting dimensions, strength, and capacity, on the accuracy-quality tradeoff. Compared with previous SOTA, our method requires minimum computation and is more applicable to large-scale models. We use StyleGAN2 and the latent diffusion model to demonstrate the efficacy of our method. Guangyu Nie, Changhoon Kim, Yezhou Yang |
ICML | 3 |
| 2023 | CAROM Air - Vehicle Localization and Traffic Scene Reconstruction from Aerial VideosabstractRoad traffic scene reconstruction from videos has been desirable by road safety regulators, city planners, researchers, and autonomous driving technology developers. However, it is expensive and unnecessary to cover every mile of the road with cameras mounted on the road infrastructure. This paper presents a method that can process aerial videos to vehicle trajectory data so that a traffic scene can be automatically reconstructed and accurately re-simulated using computers. On average, the vehicle localization error is about 0.1 m to 0.3 m using a consumer-grade drone flying at 120 meters. This project also compiles a dataset of 50 reconstructed road traffic scenes from about 100 hours of aerial videos to enable various downstream traffic analysis applications and facilitate further road traffic related research. The dataset is available at https://github.com/duolu/CAROM. Duo Lu, Eric Eaton, Matt Weg, Steven Como, Jeffrey Wishart, Yezhou Yang |
ICRA | 8 |
| 2023 | Improving Diversity with Adversarially Learned Transformations for Domain GeneralizationabstractTo be successful in single source domain generalization (SSDG), maximizing diversity of synthesized domains has emerged as one of the most effective strategies. Recent success in SSDG comes from methods that pre-specify diversity inducing image augmentations during training, so that it may lead to better generalization on new domains. However, naïve pre-specified augmentations are not always effective, either because they cannot model large domain shift, or be-cause the specific choice of transforms may not cover the types of shift commonly occurring in domain generalization. To address this issue, we present a novel framework called ALT: adversarially learned transformations, that uses an adversary neural network to model plausible, yet hard image transformations that fool the classifier. ALT learns image transformations by randomly initializing the adversary net-work for each batch and optimizing it for a fixed number of steps to maximize classification error. The classifier is trained by enforcing a consistency between its predictions on the clean and transformed images. With extensive empirical analysis, we find that this new form of adversarial transformations achieves both objectives of diversity and hardness simultaneously, outperforming all existing techniques on competitive benchmarks for SSDG. We also show that ALT can seamlessly work with existing diversity modules to produce highly distinct, and large transformations of the source domain leading to state-of-the-art performance. Code: https://github.com/tejas-gokhale/ALT Tejas Gokhale, Rushil Anirudh, Jayaraman J. Thiagarajan, Bhavya Kailkhura, Chitta Baral, Yezhou Yang |
WACV | 6 |
| 2022 | Injecting Semantic Concepts into End-to-End Image CaptioningabstractTremendous progresses have been made in recent years in developing better image captioning models, yet most of them rely on a separate object detector to extract regional features. Recent vision-language studies are shifting towards the detector-free trend by leveraging grid representations for more flexible model training and faster inference speed. However, such development is primarily focused on image understanding tasks, and remains less investigated for the caption generation task. In this paper, we are concerned with a better-performing detector-free image captioning model, and propose a pure vision transformer-based image captioning model, dubbed as ViTCAP, in which grid representations are used without extracting the regional features. For improved performance, we introduce a novel Concept Token Network (CTN) to predict the semantic concepts and then incorporate them into the end-to-end captioning. In particular, the CTN is built on the basis of a vision transformer, and is designed to predict the concept tokens through a classification task, from which the rich semantic information contained greatly benefits the captioning task. Compared with the previous detector-based models, ViTCAP drastically simplifies the architectures and at the same time achieves competitive performance on various challenging image captioning datasets. In particular, ViTCAP reaches 138.1 CIDEr scores on COCO-caption Karpathy-split, 93.8 and 108.6 CIDEr scores on nocaps and Google-CC captioning datasets, respectively. Zhiyuan Fang, Xiaowei Hu 0006, Zhe Gan, Yezhou Yang, Zicheng Liu 0001 |
CVPR | 7 |
| 2022 | CRIPP-VQA: Counterfactual Reasoning about Implicit Physical Properties via Video Question AnsweringabstractVideos often capture objects, their visible properties, their motion, and the interactions between different objects.Objects also have physical properties such as mass, which the imaging pipeline is unable to directly capture.However, these properties can be estimated by utilizing cues from relative object motion and the dynamics introduced by collisions.In this paper, we introduce CRIPP-VQA 1 , a new video question answering dataset for reasoning about the implicit physical properties of objects in a scene.CRIPP-VQA contains videos of objects in motion, annotated with questions that involve counterfactual reasoning about the effect of actions, questions about planning in order to reach a goal, and descriptive questions about visible properties of objects.The CRIPP-VQA test set enables evaluation under several outof-distribution settings -videos with objects with masses, coefficients of friction, and initial velocities that are not observed in the training distribution.Our experiments reveal a surprising and significant performance gap in terms of answering questions about implicit properties (the focus of this paper) and explicit properties of objects (the focus of prior work). Maitreya Patel, Tejas Gokhale, Chitta Baral, Yezhou Yang |
EMNLP | 4 |
| 2022 | Attributable Watermarking of Speech Generative ModelsabstractGenerative models are now capable of synthesizing images, speeches, and videos that are hardly distinguishable from authentic contents. Such capabilities cause concerns such as malicious impersonation and IP theft. This paper investigates a solution for model attribution, i.e., the classification of synthetic contents by their source models via watermarks embedded in the contents. Building on past success of model attribution in the image domain, we discuss algorithmic improvements for generating user-end speech models that empirically achieve high attribution accuracy, while maintaining high generation quality. We show the tradeoff between attributability and generation quality under a variety of attacks on generated speech signals attempting to remove the watermarks, and the feasibility of learning robust watermarks against these attacks. Watermarked speech samples are available at https://attdemo.github.io/attdemofull.github.io. Yongbaek Cho, Changhoon Kim, Yezhou Yang |
ICASSP | 3 |
| 2022 | CAVAN: Commonsense Knowledge Anchored Video CaptioningabstractIt is not merely an aggregation of static entities that a video clip carries, but also a variety of interactions and relations among these entities. Challenges still remain for a video captioning system to generate descriptions focusing on the prominent interest and aligning with the latent aspects beyond observations. In this work, we present a Commonsense knowledge Anchored Video cAptioNing(dubbed as CAVAN) approach. CAVAN exploits inferential commonsense knowledge to assist the training of video captioning model with a novel paradigm for sentence-level semantic alignment. Specifically, we acquire commonsense knowledge complementing per training caption by querying a generic knowledge atlas (ATOMIC [1]), and form the commonsense-caption entailment corpus. A BERT [2] based language entailment model trained from this corpus then serves as a commonsense discriminator for the training of video captioning model, and penalizes the model from generating semantically misaligned captions. Experimental results with ablations on MSRVTT [3], V2C [4] and VATEX [5] datasets validate the effectiveness of CAVAN and reveal that the use of commonsense knowledge benefits video caption generation. Huiliang Shao, Zhiyuan Fang, Yezhou Yang |
ICPR | 3 |
| 2022 | Targeted Attack on Deep RL-based Autonomous Driving with Learned Visual PatternsabstractRecent studies demonstrated the vulnerability of control policies learned through deep reinforcement learning against adversarial attacks, raising concerns about the application of such models to risk-sensitive tasks such as autonomous driving. Threat models for these demonstrations are limited to (1) targeted attacks through real-time manipulation of the agent's observation, and (2) untargeted attacks through manipulation of the physical environment. The former assumes full access to the agent's states/observations at all times, while the latter has no control over attack outcomes. This paper investigates the feasibility of targeted attacks through visually learned patterns placed on physical objects in the environment, a threat model that combines the practicality and effectiveness of the existing ones. Through analysis, we demonstrate that a pre-trained policy can be hijacked within a time window, e.g., performing an unintended self-parking, when an adversarial object is present. To enable the attack, we adopt an assumption that the dynamics of both the environment and the agent can be learned by the attacker. Lastly, we empirically show the effectiveness of the proposed attack on different driving scenarios, perform a location robustness test, and study the tradeoff between the attack strength and its effectiveness Code is available at https://github.com/ASU-APG/ Targeted-Physical-Adversarial-Attacks-on-AD Prasanth Buddareddygari, Travis Zhang, Yezhou Yang |
ICRA | 3 |
| 2022 | A Coulomb Force Inspired Loss Function for High-Performance Pedestrian DetectionabstractPedestrian detection has received considerable research interest due to its wide application and has made significant progress along with the development of deep neural networks. However, crowd occlusion still remains a significant challenge to current state-of-the-art pedestrian detectors due to the complication in formulating interactions between occluded instances. Inspired by the Coulomb force, we in this work set each proposal as a single electric charge and define the attractive and repulsive forces to model the interaction between ground truths and assigned proposals. This design is driven by two motivations: the attractive force pulls bounding boxes toward their assigned targets, aggregating them compactly around the ground truths. The repulsive force pushes bounding boxes away from other instances, preventing them from shifting to surrounding pedestrians. With this insight, we propose a novel bounding box regression loss and achieve more robust localization performance in crowded scenes without introducing any computational overhead. Extensive experimental evaluations on the CityPersons and CrowdHuman benchmarks demonstrate consistent state-of-the-art performance. Zhe Wang 0013, Jun Wang 0041, Yezhou Yang, Junliang Xing |
IEEE Signal Process. Lett. | 3 |
| 2021 | Attribute-Guided Adversarial Training for Robustness to Natural PerturbationsabstractWhile existing work in robust deep learning has focused on small pixel-level norm-based perturbations, this may not account for perturbations encountered in several real world settings. In many such cases although test data might not be available, broad specifications about the types of perturbations (such as an unknown degree of rotation) may be known. We consider a setup where robustness is expected over an unseen test domain that is not i.i.d. but deviates from the training domain. While this deviation may not be exactly known, its broad characterization is specified a priori, in terms of attributes. We propose an adversarial training approach which learns to generate new samples so as to maximize exposure of the classifier to the attributes-space, without having access to the data from the test domain. Our adversarial training solves a min-max optimization problem, with the inner maximization generating adversarial perturbations, and the outer minimization finding model parameters by optimizing the loss on adversarial perturbations generated from the inner maximization. We demonstrate the applicability of our approach on three types of naturally occurring perturbations --- object-related shifts, geometric transformations, and common image corruptions. Our approach enables deep neural networks to be robust against a wide range of naturally occurring perturbations. We demonstrate the usefulness of the proposed approach by showing the robustness gains of deep neural networks trained using our adversarial training on MNIST, CIFAR-10, and a new variant of the CLEVR dataset. Tejas Gokhale, Rushil Anirudh, Bhavya Kailkhura, Jayaraman J. Thiagarajan, Chitta Baral, Yezhou Yang |
AAAI | 6 |
| 2021 | SMURF: SeMantic and linguistic UndeRstanding Fusion for Caption Evaluation via Typicality AnalysisabstractJoshua Feinglass, Yezhou Yang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Joshua Feinglass, Yezhou Yang |
ACL/IJCNLP (1) | 2 |
| 2021 | First Workshop on Knowledge Injection in Neural Networks (KINN)abstractDeep learning (DL) has made rapid progress in the last decade, with neural network-based language and vision models achieving state-of-the-art performance in various tasks. Yet purely data-driven neural network models exhibit several issues impacting real-world deployment of such models adversely. These include reliance on large quantities of training data, poor robustness, lack of generalization, poor explainability, and glaring gaps in implicit and commonsense knowledge. The availability of rich structured (or semi-structured) knowledge sources has spurred the research community into exploring Knowledge Injection in Neural Networks (KINN) as a means of mitigating the above-mentioned challenges. This has led to the development of hybrid AI systems that combine the purely data-driven learning of the neural network models with an infusion of knowledge from external sources. Such KINN systems include the development of retrieval augmented neural models, neuro symbolic systems and a plethora of combinations of NNs and knowledge graphs and structured knowledge bases. Vasudev Lal, Somak Aditya, Yezhou Yang, Pasquale Minervini, Sandya Mannarswamy |
CIKM | 3 |
| 2021 | Hierarchical and Partially Observable Goal-Driven Policy Learning With Goals Relational GraphabstractWe present a novel two-layer hierarchical reinforcement learning approach equipped with a Goals Relational Graph (GRG) for tackling the partially observable goal-driven task, such as goal-driven visual navigation. Our GRG captures the underlying relations of all goals in the goal space through a Dirichlet-categorical process that facilitates: 1) the high-level network raising a sub-goal towards achieving a designated final goal; 2) the low-level network towards an optimal policy; and 3) the overall system generalizing unseen environments and goals. We evaluate our approach with two settings of partially observable goal-driven tasks — a grid-world domain and a robotic object search task. Our experimental results show that our approach exhibits superior generalization performance on both unseen environments and new goals1. Yezhou Yang |
CVPR | 2 |
| 2021 | Weakly Supervised Relative Spatial Reasoning for Visual Question AnsweringabstractVision-and-language (V&L) reasoning necessitates perception of visual concepts such as objects and actions, understanding semantics and language grounding, and reasoning about the interplay between the two modalities. One crucial aspect of visual reasoning is spatial understanding, which involves understanding relative locations of objects, i.e. implicitly learning the geometry of the scene. In this work, we evaluate the faithfulness of V&L models to such geometric understanding, by formulating the prediction of pair-wise relative locations of objects as a classification as well as a regression task. Our findings suggest that state-of-the-art transformer-based V&L models lack sufficient abilities to excel at this task. Motivated by this, we design two objectives as proxies for 3D spatial reasoning (SR) – object centroid estimation, and relative position estimation, and train V&L with weak supervision from off-the-shelf depth estimators. This leads to considerable improvements in accuracy for the "GQA" visual question answering challenge (in fully supervised, few-shot, and O.O.D settings) as well as improvements in relative spatial reasoning. Code and data will be released here. Pratyay Banerjee, Tejas Gokhale, Yezhou Yang, Chitta Baral |
ICCV | 3 |
| 2021 | Compressing Visual-linguistic Model via Knowledge DistillationabstractDespite exciting progress in pre-training for visual-linguistic (VL) representations, very few aspire to a small VL model. In this paper, we study knowledge distillation (KD) to effectively compress a transformer based large VL model into a small VL model. The major challenge arises from the inconsistent regional visual tokens extracted from different detectors of Teacher and Student, resulting in the misalignment of hidden representations and attention distributions. To address the problem, we retrain and adapt the Teacher by using the same region proposals from Student’s detector while the features are from Teacher’s own object detector. With aligned network inputs, the adapted Teacher is capable of transferring the knowledge through the intermediate representations. Specifically, we use the mean square error loss to mimic the attention distribution inside the transformer block, and present a token-wise noise contrastive loss to align the hidden state by contrasting with negative representations stored in a sample queue. To this end, we show that our proposed distillation significantly improves the performance of small VL models on image captioning and visual question answering tasks. It reaches 120.8 in CIDEr score on COCO captioning, an improvement of 5.1 over its non-distilled counterpart; and an accuracy of 69.8 on VQA 2.0, a 0.8 gain from the baseline. Our extensive experiments and ablations confirm the effectiveness of VL distillation in both pre-training and fine-tuning stages. Zhiyuan Fang, Xiaowei Hu 0006, Yezhou Yang, Zicheng Liu 0001 |
ICCV | 5 |
| 2021 | SEED: Self-supervised Distillation For Visual Representation
Zhiyuan Fang, Lei Zhang 0001, Yezhou Yang, Zicheng Liu 0001 |
ICLR | 5 |
| 2021 | Decentralized Attribution of Generative Models
Changhoon Kim, Yezhou Yang |
ICLR | 3 |
| 2021 | CAROM - Vehicle Localization and Traffic Scene Reconstruction from Monocular Cameras on Road InfrastructuresabstractTraffic monitoring cameras are powerful tools for traffic management and essential components of intelligent road infrastructure systems. In this paper, we present a vehicle localization and traffic scene reconstruction framework using these cameras, dubbed as CAROM, i.e., "CARs On the Map". CAROM processes traffic monitoring videos and converts them to anonymous data structures of vehicle type, 3D shape, position, and velocity for traffic scene reconstruction and replay. Through collaborating with a local department of transportation in the United States, we constructed a benchmarking dataset containing GPS data, roadside camera videos, and drone videos to validate the vehicle tracking results. On average, the localization error is approximately 0.8 m and 1.7 m within the range of 50 m and 120 m from the cameras, respectively. Duo Lu, Varun Chandra Jammula, Steven Como, Jeffrey Wishart, Yan Chen 0013, Yezhou Yang |
ICRA | 6 |
| 2021 | CLEVR_HYP: A Challenge Dataset and Baselines for Visual Question Answering with Hypothetical Actions over ImagesabstractShailaja Keyur Sampat, Akshay Kumar, Yezhou Yang, Chitta Baral. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Shailaja Sampat, Yezhou Yang, Chitta Baral |
NAACL-HLT | 3 |
| 2020 | VQA-LOL: Visual Question Answering Under the Lens of Logic
Tejas Gokhale, Pratyay Banerjee, Chitta Baral, Yezhou Yang |
ECCV (21) | 4 |
| 2020 | ViTAA: Visual-Textual Attributes Alignment in Person Search by Natural Language
Zhe Wang 0013, Zhiyuan Fang, Jun Wang 0041, Yezhou Yang |
ECCV (12) | 4 |
| 2020 | Video2Commonsense: Generating Commonsense Descriptions to Enrich Video CaptioningabstractCaptioning is a crucial and challenging task for video understanding.In videos that involve active agents such as humans, the agent's actions can bring about myriad changes in the scene.Observable changes such as movements, manipulations, and transformations of the objects in the scene, are reflected in conventional video captioning.Unlike images, actions in videos are also inherently linked to social aspects such as intentions (why the action is taking place), effects (what changes due to the action), and attributes that describe the agent.Thus for video understanding, such as when captioning videos or when answering questions about videos, one must have an understanding of these commonsense aspects.We present the first work on generating commonsense captions directly from videos, to describe latent aspects such as intentions, effects, and attributes.We present a new dataset "Video-to-Commonsense (V2C)" that contains ∼ 9k videos of human agents performing various actions, annotated with 3 types of commonsense descriptions.Additionally we explore the use of open-ended video-based commonsense question answering (V2C-QA) as a way to enrich our captions.Both the generation task and the QA task can be used to enrich video captions.. frame frame frame CNN Zhiyuan Fang, Tejas Gokhale, Pratyay Banerjee, Chitta Baral, Yezhou Yang |
EMNLP (1) | 5 |
| 2020 | MUTANT: A Training Paradigm for Out-of-Distribution Generalization in Visual Question AnsweringabstractWhile progress has been made on the visual question answering leaderboards, models often utilize spurious correlations and priors in datasets under the i.i.d.setting.As such, evaluation on out-of-distribution (OOD) test samples has emerged as a proxy for generalization.In this paper, we present MUTANT, a training paradigm that exposes the model to perceptually similar, yet semantically distinct mutations of the input, to improve OOD generalization, such as the VQA-CP challenge.Under this paradigm, models utilize a consistency-constrained training objective to understand the effect of semantic changes in input (question-image pair) on the output (answer).Unlike existing methods on VQA-CP, MUTANT does not rely on the knowledge about the nature of train and test answer distributions.MUTANT establishes a new state-ofthe-art accuracy on VQA-CP with a 10.57% improvement.Our work opens up avenues for the use of semantic input mutations for OOD generalization in question answering. Tejas Gokhale, Pratyay Banerjee, Chitta Baral, Yezhou Yang |
EMNLP (1) | 4 |
| 2020 | Learning hierarchical behavior and motion planning for autonomous drivingabstractLearning-based driving solution, a new branch for autonomous driving, is expected to simplify the modeling of driving by learning the underlying mechanisms from data. To improve the tactical decision-making for learning-based driving solution, we introduce hierarchical behavior and motion planning (HBMP) to explicitly model the behavior in learning-based solution. Due to the coupled action space of behavior and motion, it is challenging to solve HBMP problem using reinforcement learning (RL) for long-horizon driving tasks. We transform HBMP problem by integrating a classical sampling-based motion planner, of which the optimal cost is regarded as the rewards for high-level behavior learning. As a result, this formulation reduces action space and diversifies the rewards without losing the optimality of HBMP. In addition, we propose a sharable representation for input sensory data across simulation platforms and real-world environment, so that models trained in a fast event-based simulator, SUMO, can be used to initialize and accelerate the RL training in a dynamics based simulator, CARLA. Experimental results demonstrate the effectiveness of the method. Besides, the model is successfully transferred to the real-world, validating the generalization capability. Jingke Wang, Yue Wang 0020, Dongkun Zhang, Yezhou Yang, Rong Xiong |
IROS | 4 |
| 2020 | TKD: Temporal Knowledge Distillation for Active PerceptionabstractDeep neural network-based methods have been proved to achieve outstanding performance on object detection and classification tasks. Despite the significant performance improvement using the deep structures, they still require prohibitive runtime to process images and maintain the highest possible performance for real-time applications. Observing the phenomenon that human visual system (HVS) relies heavily on the temporal dependencies among frames from the visual input to conduct recognition efficiently, we propose a novel framework dubbed as TKD: temporal knowledge distillation. This framework distills the temporal knowledge from a heavy neural network-based model over selected video frames (the perception of the moments) to a light-weight model. To enable the distillation, we put forward two novel procedures: 1) a Long-short Term Memory (LSTM)-based keyframe selection method; and 2) a novel teacher-bounded loss design. To validate our approach, we conduct comprehensive empirical evaluations using different object detection methods over multiple datasets including Youtube-Objects and Hollywood scene dataset. Our results show consistent improvement in accuracy-speed trad- offs for object detection over the frames of the dynamic scene, compared to other modern object recognition methods. It can maintain the desired accuracy with the throughput of around 220 images per second. Implementation: https://github.com/mfarhadi/TKD-Cloud. Mohammad Farhadi, Yezhou Yang |
WACV | 2 |
| 2020 | Fine-grained visual understanding and reasoning
Jun Yu 0002, Yezhou Yang, Fionn Murtagh, Xinbo Gao 0001 |
Neurocomputing | 2 |
| 2020 | Neural Style Transfer: A ReviewabstractThe seminal work of Gatys et al. demonstrated the power of Convolutional Neural Networks (CNNs) in creating artistic imagery by separating and recombining image content and style. This process of using CNNs to render a content image in different styles is referred to as Neural Style Transfer (NST). Since then, NST has become a trending topic both in academic literature and industrial applications. It is receiving increasing attention and a variety of approaches are proposed to either improve or extend the original NST algorithm. In this paper, we aim to provide a comprehensive overview of the current progress towards NST. We first propose a taxonomy of current algorithms in the field of NST. Then, we present several evaluation methods and compare different NST algorithms both qualitatively and quantitatively. The review concludes with a discussion of various applications of NST and open problems for future research. A list of papers discussed in this review, corresponding codes, pre-trained models and more comparison results are publicly available at: https://osf.io/f8tu4/. Yongcheng Jing, Yezhou Yang, Zunlei Feng, Jingwen Ye, Yizhou Yu, Mingli Song |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2019 | Modularized Textual Grounding for Counterfactual ResilienceabstractComputer Vision applications often require a textual grounding module with precision, interpretability, and resilience to counterfactual inputs/queries. To achieve high grounding precision, current textual grounding methods heavily rely on large-scale training data with manual annotations at the pixel level. Such annotations are expensive to obtain and thus severely narrow the model's scope of real-world applications. Moreover, most of these methods sacrifice interpretability, generalizability, and they neglect the importance of being resilient to counterfactual inputs. To address these issues, we propose a visual grounding system which is 1) end-to-end trainable in a weakly supervised fashion with only image-level annotations, and 2) counterfactually resilient owing to the modular design. Specifically, we decompose textual descriptions into three levels: entity, semantic attribute, color information, and perform compositional grounding progressively. We validate our model through a series of experiments and demonstrate its improvement over the state-of-the-art methods. In particular, our model's performance not only surpasses other weakly/un-supervised methods and even approaches the strongly supervised ones, but also is interpretable for decision making and performs much better in face of counterfactual classes than all the others. Zhiyuan Fang, Shu Kong, Charless C. Fowlkes, Yezhou Yang |
CVPR | 4 |
| 2019 | Image Decomposition and Classification Through a Generative ModelabstractWe demonstrate in this paper that a generative model can be designed to perform classification tasks under challenging settings, including adversarial attacks and input distribution shifts. Specifically, we propose a conditional variational autoencoder that learns both the decomposition of inputs and the distributions of the resulting components. During test, we jointly optimize the latent variables of the generator and the relaxed component labels to find the best match between the given input and the output of the generator. The model demonstrates promising performance at recognizing overlapping components from the multiMNIST dataset, and novel component combinations from a traffic sign dataset. Experiments also show that the proposed model achieves high robustness on MNIST and NORB datasets, in particular for high-strength gradient attacks and non-gradient attacks. Houpu Yao, Malcolm Regan, Yezhou Yang |
ICIP | 3 |
| 2019 | How Shall I Drive? Interaction Modeling and Motion Planning towards Empathetic and Socially-Graceful DrivingabstractWhile intelligence of autonomous vehicles (AVs) has significantly advanced in recent years, accidents involving AVs suggest that these autonomous systems lack gracefulness in driving when interacting with human drivers. In the setting of a two-player game, we propose model predictive control based on social gracefulness, which is measured by the discrepancy between the actions taken by the AV and those that could have been taken in favor of the human driver. We define social awareness as the ability of an agent to infer such favorable actions based on knowledge about the other agent's intent, and further show that empathy, i.e., the ability to understand others' intent by simultaneously inferring others' understanding of the agent's self intent, is critical to successful intent inference. Lastly, through an intersection case, we show that the proposed gracefulness objective allows an AV to learn more sophisticated behavior, such as passive-aggressive motions that gently force the other agent to yield. Steven Elliott, Yiwei Wang 0002, Yezhou Yang |
ICRA | 4 |
| 2019 | Integrating Knowledge and Reasoning in Image UnderstandingabstractDeep learning based data-driven approaches have been successfully applied in various image understanding applications ranging from object recognition, semantic segmentation to visual question answering. However, the lack of knowledge integration as well as higher-level reasoning capabilities with the methods still pose a hindrance. In this work, we present a brief survey of a few representative reasoning mechanisms, knowledge integration methods and their corresponding image understanding applications developed by various groups of researchers, approaching the problem from a variety of angles. Furthermore, we discuss upon key efforts on integrating external knowledge with neural networks. Taking cues from these efforts, we conclude by discussing potential pathways to improve reasoning capabilities. Somak Aditya, Yezhou Yang, Chitta Baral |
IJCAI | 2 |
| 2019 | Spatial Knowledge Distillation to Aid Visual ReasoningabstractFor tasks involving language and vision, the current state-of-the-art methods tend not to leverage any additional information that might be present to gather relevant (commonsense) knowledge. A representative task is Visual Question Answering where large diagnostic datasets have been proposed to test a system's capability of answering questions about images. The training data is often accompanied by annotations of individual object properties and spatial locations. In this work, we take a step towards integrating this additional privileged information in the form of spatial knowledge to aid in visual reasoning. We propose a framework that combines recent advances in knowledge distillation (teacher-student framework), relational reasoning and probabilistic logical languages to incorporate such knowledge in existing neural networks for the task of Visual Question Answering. Specifically, for a question posed against an image, we use a probabilistic logical language to encode the spatial knowledge and the spatial understanding about the question in the form of a mask that is directly provided to the teacher network. The student network learns from the ground-truth information as well as the teachers prediction via distillation. We also demonstrate the impact of predicting such a mask inside the teachers network using attention. Empirically, we show that both the methods improve the test accuracy over a state-of-the-art approach on a publicly available dataset. Somak Aditya, Rudra Saha, Yezhou Yang, Chitta Baral |
WACV | 3 |
| 2019 | Interpretable Partitioned Embedding for Intelligent Multi-item Fashion Outfit CompositionabstractIntelligent fashion outfit composition has become more popular in recent years. Some deep-learning-based approaches reveal competitive composition. However, the uninterpretable characteristic makes such a deep-learning-based approach fail to meet the businesses’, designers’, and consumers’ urges to comprehend the importance of different attributes in an outfit composition. To realize interpretable and intelligent multi-item fashion outfit compositions, we propose a partitioned embedding network to learn interpretable embeddings from clothing items. The network contains two vital components: attribute partition module and partition adversarial module. In the attribute partition module, multiple attribute labels are adopted to ensure that different parts of the overall embedding correspond to different attributes. In the partition adversarial module, adversarial operations are adopted to achieve the independence of different parts. With the interpretable and partitioned embedding, we then construct an outfit-composition graph and an attribute matching map. Extensive experiments demonstrate that (1) the partitioned embedding have unmingled parts that correspond to different attributes and (2) outfits recommended by our model are more desirable in comparison with the existing methods. Zunlei Feng, Zhenyun Yu, Yongcheng Jing, Sai Wu, Mingli Song, Yezhou Yang, Junxiao Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2018 | Explicit Reasoning over End-to-End Neural Architectures for Visual Question AnsweringabstractMany vision and language tasks require commonsense reasoning beyond data-driven image and natural language processing. Here we adopt Visual Question Answering (VQA) as an example task, where a system is expected to answer a question in natural language about an image. Current state-of-the-art systems attempted to solve the task using deep neural architectures and achieved promising performance. However, the resulting systems are generally opaque and they struggle in understanding questions for which extra knowledge is required. In this paper, we present an explicit reasoning layer on top of a set of penultimate neural network based systems. The reasoning layer enables reasoning and answering questions where additional knowledge is required, and at the same time provides an interpretable interface to the end users. Specifically, the reasoning layer adopts a Probabilistic Soft Logic (PSL) based engine to reason over a basket of inputs: visual relations, the semantic parse of the question, and background ontological knowledge from word2vec and ConceptNet. Experimental analysis of the answers and the key evidential predicates generated on the VQA dataset validate our approach. Somak Aditya, Yezhou Yang, Chitta Baral |
AAAI | 2 |
| 2018 | Transductive Unbiased Embedding for Zero-Shot LearningabstractMost existing Zero-Shot Learning (ZSL) methods have the strong bias problem, in which instances of unseen (target) classes tend to be categorized as one of the seen (source) classes. So they yield poor performance after being deployed in the generalized ZSL settings. In this paper, we propose a straightforward yet effective method named Quasi-Fully Supervised Learning (QFSL) to alleviate the bias problem. Our method follows the way of transductive learning, which assumes that both the labeled source images and unlabeled target images are available for training. In the semantic embedding space, the labeled source images are mapped to several fixed points specified by the source categories, and the unlabeled target images are forced to be mapped to other points specified by the target categories. Experiments conducted on AwA2, CUB and SUN datasets demonstrate that our method outperforms existing state-of-the-art approaches by a huge margin of 9.3 ~ 24.5% following generalized ZSL settings, and by a large margin of 0.2 ~ 16.2% following conventional ZSL settings. Jie Song 0011, Chengchao Shen, Yezhou Yang, Yang Liu 0212, Mingli Song |
CVPR | 3 |
| 2018 | Stroke Controllable Fast Style Transfer with Adaptive Receptive Fields
Yongcheng Jing, Yang Liu 0212, Yezhou Yang, Zunlei Feng, Yizhou Yu, Dacheng Tao, Mingli Song |
ECCV (13) | 3 |
| 2018 | DeepSSH: Deep Semantic Structured Hashing for Explainable Person Re-IdentificationabstractFor large collections of gallery images captured by sparsely distributed cameras, we often employ hashing based approaches to enhance the efficiency of person re-identification (re-id). However, these hashing based approaches fail to provide semantically explainable encoding in solving the re-id problem, which makes it infeasible to identify the correct matches in the collection by just using a semantic query. To overcome this limitation, we propose a new deep hashing network called Deep Semantic Structured Hashing (DeepSSH) to obtain the semantic structured representation of human. In the proposed DeepSSH framework, both the mid-level human attributes and the high-level ID labels are used to learn a deep hashing network. Then, based on the obtained semantic structured hash code and the attribute labels, we learn a decoder to find the partial hash code corresponding to the specified attributes. Finally, a new grain scalable re-id framework is constructed to support semantic query of a person by providing partial or full semantic description of a person instead of the whole photo. Experimental results show that DeepSSH is comparable with state-of-the-art hashing-based person re-id approaches, and the experiment in semantic analysis shows that our hash code owns semantic meaning indeed. Sihui Luo 0001, Yezhou Yang, Mingli Song |
ICIP | 3 |
| 2018 | DeepSIC: Deep Semantic Image Compression
Sihui Luo 0001, Yezhou Yang, Yanling Yin, Chengchao Shen, Mingli Song |
ICONIP (1) | 2 |
| 2018 | Extrinsic Dexterity Through Active Slip Control Using Deep Predictive ModelsabstractWe present a machine learning methodology for actively controlling slip, in order to increase robot dexterity. Leveraging recent insights in deep learning, we propose a Deep Predictive Model that uses tactile sensor information to reason about slip and its future influence on the manipulated object. The obtained information is then used to precisely manipulate objects within a robot end-effector using external perturbations imposed by gravity or acceleration. We show in a set of experiments that this approach can be used to increase a robot's repertoire of motor skills. Simon Stepputtis, Yezhou Yang, Heni Ben Amor |
ICRA | 2 |
| 2018 | Active Object Perceiver: Recognition-Guided Policy Learning for Object Searching on Mobile RobotsabstractWe study the problem of learning a navigation policy for a robot to actively search for an object of interest in an indoor environment solely from its visual inputs. While scene-driven visual navigation has been widely studied, prior efforts on learning navigation policies for robots to find objects are limited. The problem is often more challenging than target scene finding as the target objects can be very small in the view and can be in an arbitrary pose. We approach the problem from an active perceiver perspective, and propose a novel framework that integrates a deep neural network based object recognition module and a deep reinforcement learning based action prediction mechanism. To validate our method, we conduct experiments on both a simulation dataset (AI2-THOR)and a real-world environment with a physical robot. We further propose a new decaying reward function to learn the control policy specific to the object searching task. Experimental results validate the efficacy of our method, which outperforms competing methods in both average trajectory length and success rate. Xin Ye 0024, Zhe Lin 0001, Shibin Zheng, Yezhou Yang |
IROS | 5 |
| 2018 | Weakly-Supervised Learning-Based Feature Localization for Confocal Laser Endomicroscopy Glioma Images
Mohammadhassan Izadyyazdanabadi, Evgenii Belykh, Claudio Cavallo, Xiaochun Zhao, Sirin Gandhi, Leandro Borba Moreira, Jennifer Eschbacher, Peter Nakaji, Mark C. Preul, Yezhou Yang |
MICCAI (2) | 10 |
| 2018 | Interpretable Partitioned Embedding for Customized Multi-item Fashion Outfit CompositionabstractIntelligent fashion outfit composition becomes more and more popular in these years. Some deep learning based approaches reveal competitive composition recently. However, the uninterpretable characteristic makes such deep learning based approach cannot meet the designers, businesses and consumers' urge to comprehend the importance of different attributes in an outfit composition. To realize interpretable and customized multi-item fashion outfit compositions, we propose a partitioned embedding network to learn interpretable embeddings from clothing items. The network consists of two vital components: attribute partition module and partition adversarial module. In the attribute partition module, multiple attribute labels are adopted to ensure that different parts of the overall embedding correspond to different attributes. In the partition adversarial module, adversarial operations are adopted to achieve the independence of different parts. With the interpretable and partitioned embedding, we then construct an outfit composition graph and an attribute matching map. Extensive experiments demonstrate that 1) the partitioned embedding have unmingled parts which corresponding to different attributes and 2) outfits recommended by our model are more desirable in comparison with the existing methods. Zunlei Feng, Zhenyun Yu, Yezhou Yang, Yongcheng Jing, Junxiao Jiang, Mingli Song |
ICMR | 3 |
| 2018 | Combining Knowledge and Reasoning through Probabilistic Soft Logic for Image Puzzle Solving
Somak Aditya, Yezhou Yang, Chitta Baral, Yiannis Aloimonos |
UAI | 2 |
| 2018 | Image Understanding using vision and reasoning through Scene Description Graph
Somak Aditya, Yezhou Yang, Chitta Baral, Yiannis Aloimonos, Cornelia Fermüller |
Comput. Vis. Image Underst. | 2 |
| 2018 | Prediction of Manipulation Actions
Cornelia Fermüller, Yezhou Yang, Konstantinos Zampogiannis, Francisco Barranco, Michael Pfeiffer 0001 |
Int. J. Comput. Vis. | 3 |
| 2018 | Convolutional neural networks: Ensemble modeling, fine-tuning and unsupervised semantic localization for neurosurgical CLE images
Mohammadhassan Izadyyazdanabadi, Evgenii Belykh, Michael Mooney, Nikolay Martirosyan, Jennifer Eschbacher, Peter Nakaji, Mark C. Preul, Yezhou Yang |
J. Vis. Commun. Image Represent. | 8 |
| 2017 | Fast task-specific target detection via graph based constraints representation and checkingabstractWe present a framework for fast target detection in real-world robotics applications. Considering that an intelligent agent attends to a task-specific object target during execution, our goal is to detect the object efficiently. We propose the concept of early recognition, which influences the candidate proposal process to achieve fast and reliable detection performance. To check the target constraints efficiently, we put forward a novel policy which generates a sub-optimal checking order, and we prove that it has bounded time cost compared to the optimal checking sequence, which is not achievable in polynomial time. Experiments on two different scenarios: 1) rigid object and 2) non-rigid body part detection validate our pipeline. To show that our method is widely applicable, we further present a human-robot interaction system based on our non-rigid body part detection. Wentao Luan, Yezhou Yang, Cornelia Fermüller, John S. Baras |
ICRA | 2 |
| 2017 | What can i do around here? Deep functional scene understanding for cognitive robotsabstractFor robots that have the capability to interact with the physical environment through their end effectors, understanding the surrounding scenes is not merely a task of image classification or object recognition. To perform actual tasks, it is critical for the robot to have a functional understanding of the visual scene. Here, we address the problem of localization and recognition of functional areas in an arbitrary indoor scene, formulated as a two-stage deep learning based detection pipeline. A new scene functionality testing-bed, which is compiled from two publicly available indoor scene datasets, is used for evaluation. Our method is evaluated quantitatively on the new dataset, demonstrating the ability to perform efficient recognition of functional areas from arbitrary indoor scenes. We also demonstrate that our detection model can be generalized to novel indoor scenes by cross validating it with images from two different datasets. Chengxi Ye, Yezhou Yang, Ren Mao, Cornelia Fermüller, Yiannis Aloimonos |
ICRA | 2 |
| 2016 | Reliable Attribute-Based Object Recognition Using High Predictive Value Classifiers
Wentao Luan, Yezhou Yang, Cornelia Fermüller, John S. Baras |
ECCV (3) | 2 |
| 2016 | LightNet: A Versatile, Standalone Matlab-based Environment for Deep LearningabstractLightNet is a lightweight, versatile, purely Matlab-based deep learning framework. The idea underlying its design is to provide an easy-to-understand, easy-to-use and efficient computational platform for deep learning research. The implemented framework supports major deep learning architectures such as Multilayer Perceptron Networks (MLP), Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN). The framework also supports both CPU and GPU computation, and the switch between them is straightforward. Different applications in computer vision, natural language processing and robotics are demonstrated as experiments. Chengxi Ye, Chen Zhao 0009, Yezhou Yang, Cornelia Fermüller, Yiannis Aloimonos |
ACM Multimedia | 3 |
| 2015 | Robot Learning Manipulation Action Plans by "Watching" Unconstrained Videos from the World Wide WebabstractIn order to advance action generation and creation in robots beyond simple learned schemas we need computational tools that allow us to automatically interpret and represent human actions. This paper presents a system that learns manipulation action plans by processing unconstrained videos from the World Wide Web. Its goal is to robustly generate the sequence of atomic actions of seen longer actions in video in order to acquire knowledge for robots. The lower level of the system consists of two convolutional neural network (CNN) based recognition modules, one for classifying the hand grasp type and the other for object recognition. The higher level is a probabilistic manipulation action grammar based parsing module that aims at generating visual sentences for robot manipulation. Experiments conducted on a publicly available unconstrained video dataset show that the system is able to learn manipulation actions by ``watching'' unconstrained videos with high accuracy. Yezhou Yang, Cornelia Fermüller, Yiannis Aloimonos |
AAAI | 1 |
| 2015 | Learning the Semantics of Manipulation ActionabstractIn this paper we present a formal computational framework for modeling manipulation actions. The introduced formalism leads to semantics of manipulation action and has applications to both observing and understanding human manipulation actions as well as executing them with a robotic mechanism (e.g. a humanoid robot). It is based on a Combinatory Categorial Grammar. The goal of the introduced framework is to: (1) represent manipulation actions with both syntax and semantic parts, where the semantic part employs $\lambda$-calculus; (2) enable a probabilistic semantic parsing schema to learn the $\lambda$-calculus representation of manipulation action from an annotated action corpus of videos; (3) use (1) and (2) to develop a system that visually observes manipulation actions and understands their meaning while it can reason beyond observations using propositional logic and axiom schemata. The experiments conducted on a public available large manipulation action dataset validate the theoretical framework and our implementation. Yezhou Yang, Yiannis Aloimonos, Cornelia Fermüller, Eren Erdal Aksoy |
ACL (1) | 1 |
| 2015 | Grasp type revisited: A modern perspective on a classical feature for visionabstractThe grasp type provides crucial information about human action. However, recognizing the grasp type from unconstrained scenes is challenging because of the large variations in appearance, occlusions and geometric distortions. In this paper, first we present a convolutional neural network to classify functional hand grasp types. Experiments on a public static scene hand data set validate good performance of the presented method. Then we present two applications utilizing grasp type classification: (a) inference of human action intention and (b) fine level manipulation action segmentation. Experiments on both tasks demonstrate the usefulness of grasp type as a cognitive feature for computer vision. This study shows that the grasp type is a powerful symbolic representation for action understanding, and thus opens new avenues for future research. Yezhou Yang, Cornelia Fermüller, Yiannis Aloimonos |
CVPR | 1 |
| 2015 | Learning the spatial semantics of manipulation actions through preposition groundingabstractIn this paper, we introduce an abstract representation for manipulation actions that is based on the evolution of the spatial relations between involved objects. Object tracking in RGBD streams enables straightforward and intuitive ways to model spatial relations in 3D space. Reasoning in 3D overcomes many of the limitations of similar previous approaches, while providing significant flexibility in the desired level of abstraction. At each frame of a manipulation video, we evaluate a number of spatial predicates for all object pairs and treat the resulting set of sequences (Predicate Vector Sequences, PVS) as an action descriptor. As part of our representation, we introduce a symmetric, time-normalized pairwise distance measure that relies on finding an optimal object correspondence between two actions. We experimentally evaluate the method on the classification of various manipulation actions in video, performed at different speeds and timings and involving different objects. The results demonstrate that the proposed representation is remarkably descriptive of the high-level manipulation semantics. Konstantinos Zampogiannis, Yezhou Yang, Cornelia Fermüller, Yiannis Aloimonos |
ICRA | 2 |
| 2014 | Low-level and high-level prior learning for visual saliency estimation
Mingli Song, Chun Chen 0001, Senlin Wang, Yezhou Yang |
Inf. Sci. | 4 |
| 2013 | Action Attribute Detection from Sports Videos with Contextual ConstraintsabstractIn this paper, we are interested in detecting action attributes from sports videos for event understanding and video analysis. Action attribute is a middle layer between low level motion features and high level action classes, which includes various motion patterns of human limbs and bodies and the interaction between human and objects. Successfully detecting action attributes provides a richer video description that facilitates many other important tasks, such action classification, video understanding, automatic video transcript, etc. A naive approach to deal with this challenging problem is to train a classifier for each attribute and then use them to detect attributes in novel videos independently. However, this independence assumption is often too strong, and as we show in our experiments, produces a large number of false positives in practice. We propose a novel approach that incorporates the contextual constraints for activity attribute detection. The temporal contexts within an attribute and the co-occurrence contexts between different attributes are modelled by a factorial conditional random field, which encourages agreement between different time points and attributes. The effectiveness of our methods are clearly illustrated by the experimental evaluations. Xiaodong Yu 0002, Ching Lik Teo, Yezhou Yang, Cornelia Fermüller, Yiannis Aloimonos |
BMVC | 3 |
| 2013 | Detection of Manipulation Action Consequences (MAC)abstractThe problem of action recognition and human activity has been an active research area in Computer Vision and Robotics. While full-body motions can be characterized by movement and change of posture, no characterization, that holds invariance, has yet been proposed for the description of manipulation actions. We propose that a fundamental concept in understanding such actions, are the consequences of actions. There is a small set of fundamental primitive action consequences that provides a systematic high-level classification of manipulation actions. In this paper a technique is developed to recognize these action consequences. At the heart of the technique lies a novel active tracking and segmentation method that monitors the changes in appearance and topological structure of the manipulated object. These are then used in a visual semantic graph (VSG) based procedure applied to the time sequence of the monitored object to recognize the action consequence. We provide a new dataset, called Manipulation Action Consequences (MAC 1.0), which can serve as test bed for other studies on this topic. Several experiments on this dataset demonstrates that our method can robustly track objects and detect their deformations and division during the manipulation. Quantitative tests prove the effectiveness and efficiency of the method. Yezhou Yang, Cornelia Fermüller, Yiannis Aloimonos |
CVPR | 1 |
| 2013 | Robots with language: Multi-label visual recognition using NLPabstractThere has been a recent interest in utilizing contextual knowledge to improve multi-label visual recognition for intelligent agents like robots. Natural Language Processing (NLP) can give us labels, the correlation of labels, and the ontological knowledge about them, so we can automate the acquisition of contextual knowledge. In this paper we show how to use tools from NLP in conjunction with Vision to improve visual recognition. There are two major approaches: First, different language databases organize words according to various semantic concepts. Using these, we can build special purpose databases that can predict the labels involved given a certain context. Here we build a knowledge base for the purpose of describing common daily activities. Second, statistical language tools can provide the correlations of different labels. We show a way to learn a language model from large corpus data that exploits these correlations and propose a general optimization scheme to integrate the language model into the system. Experiments conducted on three multi-label everyday recognition tasks support the effectiveness and efficiency of our approach, with significant gains in recognition accuracies when correlation information is used. Yezhou Yang, Ching Lik Teo, Cornelia Fermüller, Yiannis Aloimonos |
ICRA | 1 |
| 2013 | Minimalist plans for interpreting manipulation actionsabstractHumans attribute meaning to actions, and can recognize, imitate, predict, compose from parts, and analyse complex actions performed by other humans. We have built a model of action representation and understanding which takes as input perceptual data of humans performing manipulatory actions and finds a semantic interpretation of it. It achieves this by representing actions as minimal plans based on a few primitives. The motivation for our approach is to have a description, that abstracts away the variations in the way humans perform actions. The model can be used to represent complex activities on the basis of simple actions. The primitives of these minimal plans are embodied in the physicality of the system doing the analysis. The model understands an action under observation by recognising which plan is occurring. Using primitives thus rooted in its own physical structure, the model has a semanticist and causal understanding of what it observes. Using plans, the model considers actions as well as complex activities in terms of causality, compositions, and goal achievement, enabling it to perform complex tasks like prediction of primitives, separation of interleaved actions and filtering of perceptual input. We use our model over an action dataset involving humans using hand tools on objects in a constrained universe to understand an activity it has not seen before in terms of actions whose plans it knows of. The model thus illustrates a novel approach of understanding human actions by a robot. Anupam Guha, Yezhou Yang, Cornelia Fermüller, Yiannis Aloimonos |
IROS | 2 |
| 2013 | Color-to-gray based on chance of happening preservation
Mingli Song, Dapeng Tao, Chun Chen 0001, Jiajun Bu, Yezhou Yang |
Neurocomputing | 5 |
| 2012 | Towards a Watson that sees: Language-guided action recognition for robotsabstractFor robots of the future to interact seamlessly with humans, they must be able to reason about their surroundings and take actions that are appropriate to the situation. Such reasoning is only possible when the robot has knowledge of how the World functions, which must either be learned or hard-coded. In this paper, we propose an approach that exploits language as an important resource of high-level knowledge that a robot can use, akin to IBM's Watson in Jeopardy!. In particular, we show how language can be leveraged to reduce the ambiguity that arises from recognizing actions involving hand-tools from video data. Starting from the premise that tools and actions are intrinsically linked, with one explaining the existence of the other, we trained a language model over a large corpus of English newswire text so that we can extract this relationship directly. This model is then used as a prior to select the best tool and action that explains the video. We formalize the approach in the context of 1) an unsupervised recognition and 2) a supervised classification scenario by an EM formulation for the former and integrating language features for the latter. Results are validated over a new hand-tool action dataset, and comparisons with state of the art STIP features showed significantly improved results when language is used. In addition, we discuss the implications of these results and how it provides a framework for integrating language into vision on other robotic applications. Ching Lik Teo, Yezhou Yang, Hal Daumé III, Cornelia Fermüller, Yiannis Aloimonos |
ICRA | 2 |
| 2012 | Using a minimal action grammar for activity understanding in the real worldabstractThere is good reason to believe that humans use some kind of recursive grammatical structure when we recognize and perform complex manipulation activities. We have built a system to automatically build a tree structure from observations of an actor performing such activities. The activity trees that result form a framework for search and understanding, tying action to language. We explore and evaluate the system by performing experiments over a novel complex activity dataset taken using synchronized Kinect and SR4000 Time of Flight cameras. Processing of the combined 3D and 2D image data provides the necessary terminals and events to build the tree from the bottom-up. Experimental results highlight the contribution of the action grammar in: 1) providing a robust structure for complex activity recognition over real data and 2) disambiguating interleaved activities from within the same sequence. Douglas Summers-Stay, Ching Lik Teo, Yezhou Yang, Cornelia Fermüller, Yiannis Aloimonos |
IROS | 3 |
| 2011 | Corpus-Guided Sentence Generation of Natural Images
Yezhou Yang, Ching Lik Teo, Hal Daumé III, Yiannis Aloimonos |
EMNLP | 1 |
| 2011 | Active scene recognition with vision and languageabstractThis paper presents a novel approach to utilizing high level knowledge for the problem of scene recognition in an active vision framework, which we call active scene recognition. In traditional approaches, high level knowledge is used in the post-processing to combine the outputs of the object detectors to achieve better classification performance. In contrast, the proposed approach employs high level knowledge actively by implementing an interaction between a reasoning module and a sensory module (Figure 1). Following this paradigm, we implemented an active scene recognizer and evaluated it with a dataset of 20 scenes and 100+ objects. We also extended it to the analysis of dynamic scenes for activity recognition with attributes. Experiments demonstrate the effectiveness of the active paradigm in introducing attention and additional constraints into the sensing process. Xiaodong Yu 0002, Cornelia Fermüller, Ching Lik Teo, Yezhou Yang, Yiannis Aloimonos |
ICCV | 4 |
| 2010 | What Is the Chance of Happening: A New Way to Predict Where People Look
Yezhou Yang, Mingli Song, Jiajun Bu, Chun Chen 0001 |
ECCV (5) | 1 |
| 2009 | Visual attention analysis by pseudo gravitational fieldabstractAs a crucial step of the visual cognition and perception, visual attention analysis shows great importance in many research or application areas. In this paper, we treat this problem from a new angle, inspired by the classic gravitational field theory. By defining "mass" of each pixel, we compute "force" between them to obtain a so-called pseudo gravitational field over an image. Then, we propose an iteration algorithm to simulate the movement of the fixation points affected by this field. Finally, stable visual attention points or areas are obtained when those fixation points finally aggregate around some special pixels or areas. The main contributions are threefold: (1) by introducing classic gravitational field theory into visual attention analysis, a new point of view is proposed; (2) a competition scheme is constructed and the meaning of attraction can be applied into visual attention analysis intuitively; (3) by using pixels into computation through down sampling, a faster analysis method is achieved. The experimental result shows that our method is effective and consistent with the generally accepted definition of visual attention. Yezhou Yang, Mingli Song, Jiajun Bu, Chun Chen 0001 |
ACM Multimedia | 1 |