EDBT 2026 Demo / reviewers in the wild / expert
Alan L. Yuille
dblp:y/AlanLYuille · also Alan Loddon Yuille
· DBLP profile ↗
472ranked-venue papers
48as first author
151since 2021 · last 2026
0000-0001-5207-9249ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 394 · 46 first-author · 125 since 2021Graphics, computer vision, multimedia, augmented reality and games · 300 · 15 first-author · 101 since 2021Applied, interdisciplinary, general and emerging computing · 40 · 18 since 2021Systems, architecture and hardware · 4Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | X-LRM: X-Ray Large Reconstruction Model for Extremely Sparse-View Computed Tomography Recovery in One SecondabstractSparse-view 3D CT reconstruction aims to recover volumetric structures from a limited number of 2D X-ray projections. Existing feedforward methods are constrained by the scarcity of large-scale training datasets and the absence of direct and consistent 3D representations. In this paper, we propose an X-ray Large Reconstruction Model (X-LRM) for extremely sparse-view ($X$-LRM consists of two key components:$X$-former and$X$triplane.$\quad X$-former can handle an arbitrary number of input views using an MLP-based image tokenizer and a Transformer-based encoder. The output tokens are then upsampled into our X-triplane representation, which models the 3D radiodensity as an implicit neural field. To support the training of X-LRM, we introduce Torso-16K, a largescale dataset comprising over 16 K volume-projection pairs of various torso organs. Extensive experiments demonstrate that X-LRM outperforms the state-of-the-art method by 1.5$d B$and achieves$27 \times$faster speed with better flexibility. Furthermore, the evaluation of lung segmentation tasks also suggests the practical value of our approach. Our code and dataset will be released at https://github.com/Richard-Guofeng-Zhang/X-LRM. Guofeng Zhang 0025, Ruyi Zha, Hao He 0011, Yixun Liang, Alan L. Yuille, Hongdong Li, Yuanhao Cai |
3DV | 5 |
| 2026 | 4D-Animal: Freely Reconstructing Animatable 3D Animals from VideosabstractReconstructing animatable 3D animals from videos traditionally depends on sparse semantic keypoints to fit parametric models. Acquiring these keypoints is labor-intensive, and detectors trained on limited animal datasets are often unreliable. We propose 4D-Animal, a keypoint-free framework that reconstructs animatable 3D animals directly from videos. Our method employs a dense feature network to map 2D image representations to SMAL parameters, improving both efficiency and stability. Additionally, we introduce a hierarchical alignment strategy that leverages silhouette, part-level, pixel-level, and temporal cues from pretrained 2D models, ensuring accurate and temporally coherent reconstructions. Extensive experiments demonstrate that 4D-Animal outperforms both model-based and model-free baselines on dog dataset. Moreover, the high-quality 3D assets generated by our method can benefit other 3D tasks, underscoring its potential for large-scale applications. The code is released at https://github.com/zhongshsh/4D-Animal. Shanshan Zhong, Zehan Zheng, Zhongzhan Huang, Wufei Ma, Guofeng Zhang 0025, Qihao Liu, Alan L. Yuille, Jieneng Chen |
WACV | 8 |
| 2026 | A comprehensive survey of AI agents in healthcareabstractOBJECTIVE: This survey aims to systematically map the rapidly evolving landscape of AI agents in healthcare. It addresses the critical need to adapt general-purpose agentic frameworks characterized by autonomy, planning, and tool use to the high-stakes, safety-critical constraints of medical decision-making and patient care. METHODS: We conducted a comprehensive review of over 200 recent studies, synthesizing literature from major academic databases. We developed a holistic taxonomy that traces the full lifecycle of healthcare agents, analyzing perception modalities, core technical architectures, and evaluation protocols specific to autonomous systems. RESULTS: The review presents a quantitative landscape analysis showing exponential growth in the field. We structure the domain into three pillars: (1) Perception of multi-modal clinical data (e.g., EHR, imaging, genomics); (2) Agent Capabilities, including tool use, reasoning, memory, and multi-agent collaboration; and (3) an Application Ecosystem organized by stakeholder roles (clinicians, patients, researchers, and administrators). Additionally, we categorize evaluation frameworks, and discuss the deployment readiness of current systems across technical, evidentiary, and governance dimensions. Finally, we identify challenges for advancing healthcare agents from controlled evaluation toward real-world clinical integration. A continuously updated repository of related papers is available at https://github.com/AgenticHealthAI/Awesome-AI-Agents-for-Healthcare. CONCLUSION: AI agents offer significant potential to enhance healthcare through autonomous reasoning and workflow integration. However, current research remains largely concentrated in benchmark and controlled evaluation settings, and the translation into clinical practice will require advances in reliability, privacy protection, governance, and operational integration. Gelei Xu, Yixiong Chen, Yuying Duan, Shuqing Wu, Haoxinran Yu, Ching-Hao Chiu, Juntong Ni, Ningzhi Tang, Toby Jia-Jun Li, Alan L. Yuille, Wei Jin 0009, Yiyu Shi 0001 |
J. Biomed. Informatics | 11 |
| 2026 | Text-Driven Tumor SynthesisabstractTumor synthesis can generate challenging cases that AI often misses or over-detects. Training on these cases improves AI performance. However, most existing synthesis methods are either unconditional- generating images from random variables-or conditioned only on tumor shape. As a result, they lack control over clinically important tumor characteristics, such as texture, heterogeneity, boundary, and pathology. The generated tumors are therefore overly similar or duplicates of existing training cases, failing to effectively address AI's weaknesses. We propose a new text-driven tumor synthesis approach, termed TextoMorph, that provides textual control over tumor characteristics in conjunction with mask control. This approach is particularly beneficial for examples that confuse the AI the most, such as early tumor detection (improving Sensitivity by + 6.5%), tumor segmentation for precise radiotherapy (improving NSD by + 3.1%), and classification between benign and malignant tumors (improving Sensitivity by + 8.2%). By incorporating text mined from radiology reports into the synthesis process, we increase the variability and controllability of the synthetic tumors to target AI's failure cases more precisely. Moreover, TextoMorph uses contrastive learning across different texts and CT scans, significantly reducing dependence on scarce image-report pairs (only 141 pairs used in this study) by leveraging a large corpus of 34,035 radiology reports. Finally, we have developed rigorous tests to evaluate synthetic tumors, showing that our synthetic tumors is realistic and diverse in texture, heterogeneity, boundary, and pathology. Code and models are available at https://github.com/MrGiovanni/TextoMorph. Yi Shuai, Qi Chen 0014, Dong Yang 0005, Can Zhao 0001, Pedro R. A. S. Bassi, Daguang Xu, Kang Wang 0016, Yang Yang 0009, Alan L. Yuille, Zongwei Zhou |
IEEE Trans. Medical Imaging | 13 |
| 2025 | Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality EvolutionabstractDiffusion models, and their generalization, flow matching, have had a remarkable impact on the field of media generation. Here, the conventional approach is to learn the complex mapping from a simple source distribution of Gaussian noise to the target media distribution. For cross-modal tasks such as text-to-image generation, this same mapping from noise to image is learnt whilst including a conditioning mechanism in the model. One key and thus far relatively unexplored feature of flow matching is that, unlike Diffusion models, they are not constrained for the source distribution to be noise. Hence, in this paper, we propose a paradigm shift, and ask the question of whether we can instead train flow matching models to learn a direct mapping from the distribution of one modality to the distribution of another, thus obviating the need for both the noise distribution and conditioning mechanism. We present a general and simple framework, CrossFlow, for cross-modal flow matching. We show the importance of applying Variational Encoders to the input data, and introduce a method to enable Classifier-free guidance. Surprisingly, for text-to-image, CrossFlow with a vanilla transformer without cross attention slightly outperforms standard flow matching, and we show that it scales better with training steps and model size, while also allowing for interesting latent arithmetic which results in semantically meaningful edits in the output space. To demonstrate the generalizability of our approach, we also show that CrossFlow is on par with or outperforms the state-of-the-art for various cross-modal / intra-modal mapping tasks, viz. image captioning, depth estimation, and image super-resolution. We hope this paper contributes to accelerating progress in cross-modal media generation. Qihao Liu, Xi Yin 0001, Alan L. Yuille, Mannat Singh |
CVPR | 3 |
| 2025 | SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal ModelsabstractHumans naturally understand 3D spatial relationships, enabling complex reasoning like predicting collisions of vehicles from different directions. Current large multimodal models (LMMs), however, lack of this capability of 3D spatial reasoning. This limitation stems from the scarcity of 3D training data and the bias in current model designs toward 2D data. In this paper, we systematically study the impact of 3D-informed data, architecture, and training setups, introducing SpatialLLM, a large multi-modal model with advanced 3D spatial reasoning abilities. To address data limitations, we develop two types of 3D-informed training datasets: (1) 3D-informed probing data focused on object’s 3D location and orientation, and (2) 3D-informed conversation data for complex spatial relationships. Notably, we are the first to curate VQA data that incorporate 3D orientation relationships on real images. Furthermore, we systematically integrate these two types of training data with the architectural and training designs of LMMs, providing a roadmap for optimal design aimed at achieving superior 3D reasoning capabilities. Our SpatialLLM advances machines toward highly capable 3D-informed reasoning, surpassing GPT-4o performance by 8.7%. Our systematic empirical design and the resulting findings offer valuable insights for future research in this direction. Our project page is available at this link. Wufei Ma, Luoxin Ye, Celso de Melo, Alan L. Yuille, Jieneng Chen |
CVPR | 4 |
| 2025 | Mamba-Reg: Vision Mamba Also Needs RegistersabstractSimilar to Vision Transformers, this paper identifies artifacts also present within the feature maps of Vision Mamba. These artifacts, corresponding to high-norm tokens emerging in low-information background areas of images, appear much more severe in Vision Mamba—they exist prevalently even with the tiny-sized model and activate extensively across background regions. To mitigate this issue, we follow the prior solution of introducing register tokens into Vision Mamba. To better cope with Mamba blocks’ uni-directional inference paradigm, two key modifications are introduced: 1) evenly inserting registers throughout the input token sequence, and 2) recycling registers for final decision predictions. We term this new architecture Mamba®. Qualitative observations suggest, compared to vanilla Vision Mamba, Mamba®’s feature maps appear cleaner and more focused on semantically meaningful regions. Quantitatively, Mamba®attains stronger performance and scales better. For example, on the ImageNet benchmark, our Mamba®-B attains 83.0% accuracy, significantly outperforming Vim-B’s 81.8%; furthermore, we provide the first successful scaling to the large model size with 341M parameters, attaining competitive accuracies of 83.6% and 84.5% for 224×224 and 384×384 inputs, respectively. Additional validation on the downstream semantic segmentation task also supports Mamba®’s efficacy. Code is available at https://github.com/wangf3014/Mamba-Reg. Feng Wang 0047, Jiahao Wang 0001, Sucheng Ren, Guoyizhe Wei, Jieru Mei, Wei Shao 0008, Yuyin Zhou, Alan L. Yuille, Cihang Xie |
CVPR | 8 |
| 2025 | Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Mutimodal ModelsabstractAlthough large multimodal models (LMMs) have demonstrated remarkable capabilities in visual scene interpretation and reasoning, their capacity for complex and precise 3-dimensional spatial reasoning remains uncertain. Existing benchmarks focus predominantly on 2D spatial understanding and lack a framework to comprehensively evaluate 6D spatial reasoning across varying complexities. To address this limitation, we present Spatial457, a scalable and unbiased synthetic dataset designed with 4 key capability for spatial reasoning: multi-object recognition, 2D location, 3D location, and 3D orientation. We develop a cascading evaluation structure, constructing 7 question types across 5 difficulty levels that range from basic single object recognition to our new proposed complex 6D spatial reasoning tasks. We evaluated various large multimodal models (LMMs) on Spatial457, observing a general decline in performance as task complexity increases, particularly in 3D reasoning and 6D spatial tasks. To quantify these challenges, we introduce the Relative Performance Dropping Rate (RPDR), highlighting key weaknesses in 3D reasoning capabilities. Leveraging the unbiased attribute design of our dataset, we also uncover prediction biases across different attributes, with similar patterns observed in real-world image settings.1The code is released in https://github.com/XingruiWang/Spatial457. Xingrui Wang, Wufei Ma, Tiezheng Zhang, Celso de Melo, Jieneng Chen, Alan L. Yuille |
CVPR | 6 |
| 2025 | Adventurer: Optimizing Vision Mamba Architecture Designs for Efficiencyabstractseries models where we treat images as sequences of patch tokens and employ uni-directional language models to learn visual representations. This modeling paradigm allows us to process images in a recurrent formulation with linear complexity relative to the sequence length, which can effectively address the memory and computation explosion issues posed by high-resolution and fine-grained images. In detail, we introduce two simple designs that seamlessly integrate image inputs into the causal inference framework: a global pooling token placed at the beginning of the sequence and a flipping operation between every two layers. Extensive empirical studies highlight that compared with the existing plain architectures such as DeiT [46] and Vim [57], Adventurer offers an optimal efficiency-accuracy trade-off. For example, our Adventurer-Base attains a competitive test accuracy of 84.3% on the standard ImageNet-1k benchmark with 216 images/s training throughput, which is 3.8× and 6.2× faster than Vim and DeiT to achieve the same result. As Adventurer offers great computation and memory efficiency and allows scaling with linear complexity, we hope this architecture can benefit future explorations in modeling long sequences for high-resolution or fine-grained images. Code is available at https://github.com/wangf3014/Adventurer. Feng Wang 0047, Timing Yang, Yaodong Yu, Sucheng Ren, Guoyizhe Wei, Angtian Wang, Wei Shao 0008, Yuyin Zhou, Alan L. Yuille, Cihang Xie |
CVPR | 9 |
| 2025 | RadGPT: Constructing 3D Image-Text Tumor DatasetsabstractWith over 85 million CT scans performed annually in the United States, creating tumor-related reports is a challenging and time-consuming task for radiologists. To address this need, we present RadGPT, an Anatomy-Aware Vision-Language AI Agent for generating detailed reports from CT scans. RadGPT first segments tumors, including benign cysts and malignant tumors, and their surrounding anatomical structures, then transforms this information into both structured reports and narrative reports. These reports provide tumor size, shape, location, attenuation, volume, and interactions with surrounding blood vessels and organs. Extensive evaluation on unseen hospitals shows that RadGPT can produce accurate reports, with high sensitivity/specificity for small tumor (<2 cm) detection: 80/73% for liver tumors, 92/78% for kidney tumors, and 77/77% for pancreatic tumors. For large tumors, sensitivity ranges from 89% to 97%. The results significantly surpass the state-of-the-art in abdominal CT report generation. RadGPT generated reports for 17 public datasets. Through radiologist review and refinement, we have ensured the reports' accuracy, and created the first publicly available image-text 3D medical dataset, comprising over 1.8 million text tokens and 2.7 million images from 9,262 CT scans, including 2,947 tumor scans/reports of 8,562 tumor instances. Our reports can: (1) localize tumors in eight liver sub-segments and three pancreatic sub-segments annotated per-voxel; (2) determine pancreatic tumor stage (T1-T4) in 260 reports; and (3) present individual analyses of multiple tumors--rare in human-made reports. Importantly, 948 of the reports are for early-stage tumors. Pedro R. A. S. Bassi, Mehmet Can Yavuz, Ibrahim Ethem Hamamci, Sezgin Er, Xiaoxi Chen, Bjoern Menze, Sergio Decherchi, Andrea Cavalli, Kang Wang 0016, Yang Yang 0009, Alan L. Yuille, Zongwei Zhou |
ICCV | 12 |
| 2025 | Baking Gaussian Splatting Into Diffusion Denoiser for Fast and Scalable Single-Stage Image-to-3D Generation and Reconstruction
Yuanhao Cai, He Zhang 0004, Kai Zhang 0045, Yixun Liang, Mengwei Ren, Fujun Luan, Qing Liu 0017, Soo Ye Kim, Jianming Zhang 0001, Yuqian Zhou, Yulun Zhang 0001, Xiaokang Yang 0001, Zhe Lin 0001, Alan L. Yuille |
ICCV | 15 |
| 2025 | Scaling Tumor Segmentation: Best Lessons from Real and Synthetic Data
Qi Chen 0014, Xinze Zhou, Hao Chen 0011, Zekun Jiang, Ziyan Huang, Dexin Yu, Junjun He, Yefeng Zheng 0001, Ling Shao 0001, Alan L. Yuille, Zongwei Zhou |
ICCV | 13 |
| 2025 | 3DSRBENCH: A Comprehensive 3D Spatial Reasoning Benchmarkabstract3D spatial reasoning is the ability to analyze and interpret the positions, orientations, and spatial relationships of objects within the 3D space. This allows models to develop a comprehensive understanding of the 3D scene, enabling their applicability to a broader range of areas, such as autonomous navigation, robotics, and AR/VR. While large multi-modal models (LMMs) have achieved remarkable progress in a wide range of image and video understanding tasks, their capabilities to perform 3D spatial reasoning on diverse natural images are less studied. In this work we present the first comprehensive 3D spatial reasoning benchmark, 3DSRBench, with 2,772 manually annotated visual question-answer pairs across 12 question types. We conduct robust and thorough evaluation of 3D spatial reasoning abilities by balancing data distribution and adopting a novel FlipEval strategy. To further study the robustness of 3D spatial reasoning w.r.t. camera 3D viewpoints, our 3DSRBench includes two subsets with 3D spatial reasoning questions on paired images with common and uncommon viewpoints. We benchmark a wide range of open-sourced and proprietary LMMs, uncovering their limitations in various aspects of 3D awareness, such as height, orientation, location, and multi-object reasoning, as well as their degraded performance on images from uncommon 6D viewpoints. Our 3DSRBench provide valuable findings and insights about future development of LMMs with strong spatial reasoning abilities. Our project page is available at https://3dsrbench.github.io/. Wufei Ma, Guofeng Zhang 0025, Yu-Cheng Chou, Jieneng Chen, Celso de Melo, Alan L. Yuille |
ICCV | 7 |
| 2025 | Beyond Next-Token: Next-X Prediction for Autoregressive Visual GenerationabstractAutoregressive (AR) modeling, known for its next-token prediction paradigm, underpins state-of-the-art language and visual generative models. Traditionally, a ``token'' is treated as the smallest prediction unit, often a discrete symbol in language or a quantized patch in vision. However, the optimal token definition for 2D image structures remains an open question. Moreover, AR models suffer from exposure bias, where teacher forcing during training leads to error accumulation at inference. In this paper, we propose xAR, a generalized AR framework that extends the notion of a token to an entity X, which can represent an individual patch token, a cell (a $k\times k$ grouping of neighboring patches), a subsample (a non-local grouping of distant patches), a scale (coarse-to-fine resolution), or even a whole image. Additionally, we reformulate discrete token classification as continuous entity regression, leveraging flow-matching methods at each AR step. This approach conditions training on noisy entities instead of ground truth tokens, leading to Noisy Context Learning, which effectively alleviates exposure bias. As a result, xAR offers two key advantages: (1) it enables flexible prediction units that capture different contextual granularity and spatial structures, and (2) it mitigates exposure bias by avoiding reliance on teacher forcing. On ImageNet-256 generation benchmark, our base model, xAR-B (172M), outperforms DiT-XL/SiT-XL (675M) while achieving 20$\times$ faster inference. Meanwhile, xAR-H sets a new state-of-the-art with an FID of 1.24, running 2.2$\times$ faster than the previous best-performing model without relying on vision foundation modules (e.g., DINOv2) or advanced guidance interval sampling. Sucheng Ren, Qihang Yu, Ju He, Xiaohui Shen, Alan L. Yuille, Liang-Chieh Chen |
ICCV | 5 |
| 2025 | Videoauteur: Towards Long Narrative Video GenerationabstractRecent video generation models have shown promising results in producing high-quality video clips lasting several seconds. However, these models face challenges in generating long sequences that convey clear and informative events, limiting their ability to support coherent narrations. In this paper, we present a large-scale cooking video dataset designed to advance long-form narrative generation in the cooking domain. We validate the quality of our proposed dataset in terms of visual fidelity and textual caption accuracy using state-of-the-art Vision-Language Models (VLMs) and video generation models, respectively. We further introduce a Long Narrative Video Director to enhance both visual and semantic coherence in generated videos and emphasize the role of aligning visual embeddings to achieve improved overall video quality. Our method demonstrates substantial improvements in generating visually detailed and semantically aligned keyframes, supported by finetuning techniques that integrate text and image embeddings within the video generation process. Project page: https://videoauteur.github.io/ Junfei Xiao, Liangke Gui, Yang Zhao 0003, Shanchuan Lin, Jiepeng Cen, Zhibei Ma, Alan L. Yuille |
ICCV | 9 |
| 2025 | Medical World Model
Zhao-Yang Wang, Qiuping Liu, Shuwen Sun, Kang Wang 0016, Rama Chellappa, Zongwei Zhou, Alan L. Yuille, Lei Zhu 0003, Jieneng Chen |
ICCV | 8 |
| 2025 | Scaling 3D Compositional Models for Robust Classification and Pose EstimationabstractDeep learning algorithms for object classification and 3D object pose estimation lack robustness to out-of-distribution factors such as synthetic stimuli, changes in weather conditions, and partial occlusion. Recently, a class of Neural Mesh Models have been developed where objects are represented in terms of 3D meshes with learned features at the vertices. These models have shown robustness in small-scale settings, involving 10 objects, but it is unclear that they can be scaled up to 100s of object classes. The main problem is that their training involves contrastive learning among the vertices of all object classes, which scales quadratically with the number of classes. We present a strategy which exploits the compositionality of the objects, i.e. the independence of the feature vectors of the vertices, which greatly reduces the training time while also improving the performance of the algorithms. We first restructure the per-vertex contrastive learning into contrasting within class and between classes. Then we propose a process that dynamically decouples the contrast between classes which are rarely confused, and enhances the contrast between the vertices of classes that are most confused. Our large-scale 3D compositional model not only achieves state-of-the-art performance on the task of predicting classification and pose estimation simultaneously, surpassing Neural Mesh Models and standard DNNs, but is also more robust to out-of-distribution testing including occlusion, weather conditions, synthetic data, and generalization to unknown classes. Xiaoding Yuan, Guofeng Zhang 0020, Prakhar Kaushik, Artur Jesslen, Adam Kortylewski, Alan L. Yuille |
ICCV | 6 |
| 2025 | GenEx: Generating an Explorable WorldabstractUnderstanding, navigating, and exploring the 3D physical real world has long been a central challenge in the development of artificial intelligence. In this work, we take a step toward this goal by introducing *GenEx*, a system capable of planning complex embodied world exploration, guided by its generative imagination that forms expectations about the surrounding environments. *GenEx* generates high-quality, continuous 360-degree virtual environments, achieving robust loop consistency and active 3D mapping over extended trajectories. Leveraging generative imagination, GPT-assisted agents can undertake complex embodied tasks, including goal-agnostic exploration and goal-driven navigation. Agents utilize imagined observations to update their beliefs, simulate potential outcomes, and enhance their decision-making. Training on the synthetic urban dataset *GenEx-DB* and evaluation on *GenEx-EQA* demonstrate that our approach significantly improves agents' planning capabilities, providing a transformative platform toward intelligent, imaginative embodied exploration. Taiming Lu, Tianmin Shu, Alan L. Yuille, Daniel Khashabi, Jieneng Chen |
ICLR | 3 |
| 2025 | Autoregressive Pretraining with Mamba in VisionabstractThe vision community has started to build with the recently developed state space model, Mamba, as the new backbone for a range of tasks. This paper shows that Mamba's visual capability can be significantly enhanced through autoregressive pretraining, a direction not previously explored. Efficiency-wise, the autoregressive nature can well capitalize on the Mamba's unidirectional recurrent structure, enabling faster overall training speed compared to other training strategies like mask modeling. Performance-wise, autoregressive pretraining equips the Mamba architecture with markedly higher accuracy over its supervised-trained counterparts and, more importantly, successfully unlocks its scaling potential to large and even huge model sizes. For example, with autoregressive pretraining, a base-size Mamba attains 83.2\% ImageNet accuracy, outperforming its supervised counterpart by 2.0\%; our huge-size Mamba, the largest Vision Mamba to date, attains 85.0\% ImageNet accuracy (85.5\% when finetuned with $384\times384$ inputs), notably surpassing all other Mamba variants in vision. The code is available at \url{https://github.com/OliverRensu/ARM}. Sucheng Ren, Xianhang Li, Haoqin Tu, Fangxun Shu, Jieru Mei, Alan L. Yuille, Cihang Xie |
ICLR | 11 |
| 2025 | Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question AnsweringabstractFor vision-language models (VLMs), understanding the dynamic properties of objects and their interactions in 3D scenes from videos is crucial for effective reasoning about high-level temporal and action semantics. Although humans are adept at understanding these properties by constructing 3D and temporal (4D) representations of the world, current video understanding models struggle to extract these dynamic semantics, arguably because these models use cross-frame reasoning without underlying knowledge of the 3D/4D scenes.
In this work, we introduce **DynSuperCLEVR**, the first video question answering dataset that focuses on language understanding of the dynamic properties of 3D objects. We concentrate on three physical concepts—*velocity*, *acceleration*, and *collisions*—within 4D scenes. We further generate three types of questions, including factual queries, future predictions, and counterfactual reasoning that involve different aspects of reasoning on these 4D dynamic properties.
To further demonstrate the importance of explicit scene representations in answering these 4D dynamics questions, we propose **NS-4DPhysics**, a **N**eural-**S**ymbolic VideoQA model integrating **Physics** prior for **4D** dynamic properties with explicit scene representation of videos.
Instead of answering the questions directly from the video text input, our method first estimates the 4D world states with a 3D generative model powered by a physical prior, and then uses neural symbolic reasoning to answer the questions based on the 4D world states.
Our evaluation on all three types of questions in DynSuperCLEVR shows that previous video question answering models and large multimodal models struggle with questions about 4D dynamics, while our NS-4DPhysics significantly outperforms previous state-of-the-art models. Xingrui Wang, Wufei Ma, Angtian Wang, Adam Kortylewski, Alan L. Yuille |
ICLR | 6 |
| 2025 | Scaling Laws in Patchification: An Image Is Worth 50, 176 Tokens And MoreabstractSince the introduction of Vision Transformer (ViT), patchification has long been regarded as a common image pre-processing approach for plain visual architectures. By compressing the spatial size of images, this approach can effectively shorten the token sequence and reduce the computational cost of ViT-like plain architectures. In this work, we aim to thoroughly examine the information loss caused by this patchification-based compressive encoding paradigm and how it affects visual understanding. We conduct extensive patch size scaling experiments and excitedly observe an intriguing scaling law in patchification: the models can consistently benefit from decreased patch sizes and attain improved predictive performance, until it reaches the minimum patch size of 1*1, i.e., pixel tokenization. This conclusion is broadly applicable across different vision tasks, various input scales, and diverse architectures such as ViT and the recent Mamba models. Moreover, as a by-product, we discover that with smaller patches, task-specific decoder heads become less critical for dense prediction. In the experiments, we successfully scale up the visual sequence to an exceptional length of 50,176 tokens, achieving a competitive test accuracy of 84.6% with a base-sized model on the ImageNet-1k benchmark. We hope this study can provide insights and theoretical foundations for future works of building non-compressive vision models. Feng Wang 0047, Yaodong Yu, Wei Shao 0008, Yuyin Zhou, Alan L. Yuille, Cihang Xie |
ICML | 5 |
| 2025 | FlowAR: Scale-wise Autoregressive Image Generation Meets Flow MatchingabstractAutoregressive (AR) modeling has achieved remarkable success in natural language processing by enabling models to generate text with coherence and contextual understanding through next token prediction. Recently, in image generation, VAR proposes scale-wise autoregressive modeling, which extends the next token prediction to the next scale prediction, preserving the 2D structure of images. However, VAR encounters two primary challenges: (1) its complex and rigid scale design limits generalization in next scale prediction, and (2) the generator’s dependence on a discrete tokenizer with the same complex scale structure restricts modularity and flexibility in updating the tokenizer. To address these limitations, we introduce FlowAR, a general next scale prediction method featuring a streamlined scale design, where each subsequent scale is simply double the previous one. This eliminates the need for VAR’s intricate multi-scale residual tokenizer and enables the use of any off-the-shelf Variational AutoEncoder (VAE). Our simplified design enhances generalization in next scale prediction and facilitates the integration of Flow Matching for high-quality image synthesis. We validate the effectiveness of FlowAR on the challenging ImageNet-256 benchmark, demonstrating superior generation performance compared to previous methods. Codes is available at \href{https://github.com/OliverRensu/FlowAR}{https://github.com/OliverRensu/FlowAR}. Sucheng Ren, Qihang Yu, Ju He, Xiaohui Shen, Alan L. Yuille, Liang-Chieh Chen |
ICML | 5 |
| 2025 | Learning Segmentation from Radiology Reports
Pedro R. A. S. Bassi, Jieneng Chen, Zheren Zhu, Sergio Decherchi, Andrea Cavalli, Kang Wang 0016, Yang Yang 0009, Alan L. Yuille, Zongwei Zhou |
MICCAI (5) | 10 |
| 2025 | CoCa-CXR: Contrastive Captioners Learn Strong Temporal Structures for Chest X-Ray Vision-Language Understanding
Yixiong Chen, Shawn Xu, Andrew Sellergren, Yossi Matias, Avinatan Hassidim, Shravya Shetty, Daniel Golden, Alan L. Yuille |
MICCAI (6) | 8 |
| 2025 | OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control ConditionsabstractExisting feedforward subject-driven video customization methods mainly study single-subject scenarios due to the difficulty of constructing multi-subject training data pairs. Another challenging problem that how to use the signals such as depth, mask, camera, and text prompts to control and edit the subject in the customized video is still less explored. In this paper, we first propose a data construction pipeline, VideoCus-Factory, to produce training data pairs for multi-subject customization from raw videos without labels and control signals such as depth-to-video and mask-to-video pairs. Based on our constructed data, we develop an Image-Video Transfer Mixed (IVTM) training with image editing data to enable instructive editing for the subject in the customized video. Then we propose a diffusion Transformer framework, OmniVCus, with two embedding mechanisms, Lottery Embedding (LE) and Temporally Aligned Embedding (TAE). LE enables inference with more subjects by using the training subjects to activate more frame embeddings. TAE encourages the generation process to extract guidance from temporally aligned control signals by assigning the same frame embeddings to the control and noise tokens. Experiments demonstrate that our method significantly surpasses state-of-the-art methods in both quantitative and qualitative evaluations. Project page is at https://caiyuanhao1998.github.io/project/OmniVCus/ Yuanhao Cai, He Zhang 0004, Jinbo Xing, Yuqian Zhou, Soo Ye Kim, Yulun Zhang 0001, Xiaokang Yang 0001, Alan L. Yuille |
NeurIPS | 14 |
| 2025 | PanTS: The Pancreatic Tumor Segmentation DatasetabstractPanTS is a large-scale, multi-institutional dataset curated to advance research in pancreatic CT analysis. It contains 36,390 CT scans from 145 medical centers, with expert-validated, voxel-wise annotations of over 993,000 anatomical structures, covering pancreatic tumors, pancreas head, body, and tail, and 24 surrounding anatomical structures such as vascular/skeletal structures and abdominal/thoracic organs. Each scan includes metadata such as patient age, sex, diagnosis, contrast phase, in-plane spacing, slice thickness, etc. AI models trained on PanTS achieve significantly better performance in pancreatic tumor detection, localization, and segmentation than those trained on existing public datasets. Our analysis indicates that these gains are directly attributable to the 16× larger-scale tumor annotations and indirectly supported by the 24 additional surrounding anatomical structures. As the largest and most comprehensive resource of its kind, PanTS offers a new benchmark for developing and evaluating AI models in pancreatic CT analysis. Xinze Zhou, Qi Chen 0014, Pedro R. A. S. Bassi, Xiaoxi Chen, Zheren Zhu, Kang Wang 0016, Yang Yang 0009, Yucheng Tang, Daguang Xu, Alan L. Yuille, Zongwei Zhou |
NeurIPS | 15 |
| 2025 | Are Pixel-Wise Metrics Reliable for Computerized Tomography Reconstruction?abstractWidely adopted evaluation metrics for sparse-view CT reconstruction, such as Structural Similarity Index Measure and Peak Signal-to-Noise Ratio, prioritize pixel-wise fidelity but often fail to capture the completeness of critical anatomical structures, particularly small or thin regions that are easily missed. To address this limitation, we propose a suite of novel anatomy-aware evaluation metrics designed to assess structural completeness across anatomical structures, including large organs, small organs, intestines, and vessels. Building on these metrics, we introduce CARE, a Completeness-Aware Reconstruction Enhancement framework that incorporates structural penalties during training to encourage anatomical preservation of significant structures. CARE is model-agnostic and can be seamlessly integrated into analytical, implicit, and generative methods.
When applied to these methods, CARE substantially improves structural completeness in CT reconstructions, achieving up to **32%** improvement for large organs, **22%** for small organs, **40%** for intestines, and **36%** for vessels. Chuntung Zhuang, Qi Chen 0014, Yuanhao Cai, Alan L. Yuille, Zongwei Zhou |
NeurIPS | 7 |
| 2025 | SpatialReasoner: Towards Explicit and Generalizable 3D Spatial ReasoningabstractDespite recent advances on multi-modal models, 3D spatial reasoning remains a challenging task for state-of-the-art open-source and proprietary models. Recent studies explore data-driven approaches and achieve enhanced spatial reasoning performance by fine-tuning models on 3D-related visual question-answering data. However, these methods typically perform spatial reasoning in an implicit manner and often fail on questions that are trivial to humans, even with long chain-of-thought reasoning. In this work, we introduce SpatialReasoner, a novel large vision-language model (LVLM) that addresses 3D spatial reasoning with explicit 3D representations shared between multiple stages--3D perception, computation, and reasoning. Explicit 3D representations provide a coherent interface that supports advanced 3D spatial reasoning and improves the generalization ability to novel question types. Furthermore, by analyzing the explicit 3D representations in multi-step reasoning traces of SpatialReasoner, we study the factual errors and identify key shortcomings of current LVLMs. Results show that our SpatialReasoner achieves improved performance on a variety of spatial reasoning benchmarks, outperforming Gemini 2.0 by 9.2% on 3DSRBench, and generalizes better when evaluating on novel 3D spatial reasoning questions. Our study bridges the 3D parsing capabilities of prior visual foundation models with the powerful reasoning abilities of large language models, opening new directions for 3D spatial reasoning. Wufei Ma, Yu-Cheng Chou, Qihao Liu, Xingrui Wang, Celso de Melo, Jianwen Xie, Alan L. Yuille |
NeurIPS | 7 |
| 2025 | Vision‑Language‑Vision Auto‑Encoder: Scalable Knowledge Distillation from Diffusion ModelsabstractBuilding state-of-the-art Vision-Language Models (VLMs) with strong captioning capabilities typically necessitates training on billions of high-quality image-text pairs, requiring millions of GPU hours. This paper introduces the Vision-Language-Vision **(VLV)** auto-encoder framework, which strategically leverages key pretrained components: a vision encoder, the decoder of a Text-to-Image (T2I) diffusion model, and subsequently, a Large Language Model (LLM). Specifically, we establish an information bottleneck by regularizing the language representation space, achieved through freezing the pretrained T2I diffusion decoder. Our VLV pipeline effectively distills knowledge from the text-conditioned diffusion model using continuous embeddings, demonstrating comprehensive semantic understanding via high-quality reconstructions. Furthermore, by fine-tuning a pretrained LLM to decode the intermediate language representations into detailed descriptions, we construct a state-of-the-art (SoTA) captioner comparable to leading models like GPT-4o and Gemini 2.0 Flash. Our method demonstrates exceptional cost-efficiency and significantly reduces data requirements; by primarily utilizing single-modal images for training and maximizing the utility of existing pretrained models (image encoder, T2I diffusion model, and LLM), it circumvents the need for massive paired image-text datasets, keeping the total training expenditure under $1,000 USD. Tiezheng Zhang, Yu-Cheng Chou, Jieneng Chen, Alan L. Yuille, Junfei Xiao |
NeurIPS | 5 |
| 2025 | EasyRet3D: Uncalibrated Multi-View Multi-Human 3D Reconstruction and TrackingabstractCurrent methods performing 3D human pose estimation from multi-view still bear several key limitations. First, most methods require manual intrinsic and extrinsic camera calibration, which is laborious and difficult in many settings. Second, more accurate models rely on further training on the same datasets they evaluate, severely limiting their generalizability in real-world settings. We address these limitations with EasyRet3D (Easy REconstruction and Tracking in 3D), which simultaneously reconstructs and tracks 3D humans in a global coordinate frame across all views with uncalibrated cameras and videos in the wild. EasyRet3D is a compositional framework that composes our proposed modules (Automatic Calibration module, Adaptive Stitching Module, and Optimization Module) and off-the-shelf, large pre-trained models at intermediate steps to avoid manual intrinsic and extrinsic calibration and task-specific training. EasyRet3D outperforms all existing multi-view 3D tracking or pose estimation methods in Panoptic, EgoHumans, Shelf, and Human3.6M datasets. Code and demos will be released on the project website. Junjie Oscar Yin, Jiahao Wang 0001, Yi Zhang 0099, Alan L. Yuille |
WACV | 5 |
| 2024 | HISR: Hybrid Implicit Surface Representation for Photorealistic 3D Human ReconstructionabstractNeural reconstruction and rendering strategies have demonstrated state-of-the-art performances due, in part, to their ability to preserve high level shape details. Existing approaches, however, either represent objects as implicit surface functions or neural volumes and still struggle to recover shapes with heterogeneous materials, in particular human skin, hair or clothes. To this aim, we present a new hybrid implicit surface representation to model human shapes. This representation is composed of two surface layers that represent opaque and translucent regions on the clothed human body. We segment different regions automatically using visual cues and learn to reconstruct two signed distance functions (SDFs). We perform surface-based rendering on opaque regions (e.g., body, face, clothes) to preserve high-fidelity surface normals and volume rendering on translucent regions (e.g., hair). Experiments demonstrate that our approach obtains state-of-the-art results on 3D human reconstructions, and also shows competitive performances on other objects. Angtian Wang, Yuanlu Xu, Nikolaos Sarafianos, Robert Maier 0001, Edmond Boyer, Alan L. Yuille, Tony Tung |
AAAI | 6 |
| 2024 | De-Diffusion Makes Text a Strong Cross-Modal InterfaceabstractWe demonstrate text as a strong cross-modal interface. Rather than relying on deep embeddings to connect image and language as the interface representation, our approach represents an image as text, from which we enjoy the interpretability and flexibility inherent to natural language. We employ an autoencoder that uses a pre-trained text-to-image diffusion model for decoding. The encoder is trained to transform an input image into text, which is then fed into the fixed text-to-image diffusion decoder to reconstruct the original input ─a process we term De-Diffusion. Experiments validate both the precision and comprehensiveness of De-Diffusion text representing images, such that it can be readily ingested by off-the-shelf text-to-image tools and LLMs for diverse multi-modal tasks. For example, a single De-Diffusion model can generalize to provide transferable prompts for different text-to-image tools, and also achieves a new state of the art on open-ended vision-language tasks by simply prompting large language models with few-shot examples. Project page: dediffusion.github.io. Chen Wei 0005, Chenxi Liu 0001, Siyuan Qiao, Zhishuai Zhang, Alan L. Yuille |
CVPR | 5 |
| 2024 | Sequential Modeling Enables Scalable Learning for Large Vision ModelsabstractWe introduce a novel sequential modeling approach which enables learning a Large Vision Model (LVM) without making use of any linguistic data. To do this, we define a common format, “visual sentences ”, in which we can represent raw images and videos as well as annotated data sources such as semantic segmentations and depth reconstructions with-out needing any meta-knowledge beyond the pixels. Once this wide variety of visual data (comprising 420 billion to-kens) is represented as sequences, the model can be trained to minimize a cross-entropy loss for next token prediction. By training across various scales of model architecture and data diversity, we provide empirical evidence that our models scale effectively. Many different vision tasks can be solved by designing suitable visual prompts at test time. Yutong Bai, Xinyang Geng, Karttikeya Mangalam, Amir Bar, Alan L. Yuille, Trevor Darrell, Jitendra Malik, Alexei A. Efros |
CVPR | 5 |
| 2024 | Structure-Aware Sparse-View X-Ray 3D ReconstructionabstractX-ray, known for its ability to reveal internal structures of objects, is expected to provide richer information for 3D reconstruction than visible light. Yet, existing NeRF algorithms overlook this nature of X-ray, leading to their limitations in capturing structural contents of imaged objects. In this paper, we propose a framework, Structure-Aware X-ray Neural Radiodensity Fields (SAX-NeRF), for sparse-view X-ray 3D reconstruction. Firstly, we design a Line Segment-based Transformer (Lineformer) as the backbone of SAX-NeRF. Linefomer captures internal structures of objects in 3D space by modeling the dependencies within each line segment of an X-ray. Secondly, we present a Masked Local-Global (MLG) ray sampling strategy to extract contextual and geometric information in 2D projection. Plus, we collect a larger-scale dataset X3D covering wider X-ray applications. Experiments on X3D show that SAX-NeRF surpasses previous NeRF-based methods by 12.56 and 2.49 dB on novel view synthesis and CT reconstruction. https://github.com/caiyuanhao1998/SAX-NeRF Yuanhao Cai, Jiahao Wang 0001, Alan L. Yuille, Zongwei Zhou, Angtian Wang |
CVPR | 3 |
| 2024 | Towards Generalizable Tumor SynthesisabstractTumor synthesis enables the creation of artificial tumors in medical images, facilitating the training of AI models for tumor detection and segmentation. However, success in tumor synthesis hinges on creating visually realistic tumors that are generalizable across multiple organs and, furthermore, the resulting AI models being capable of detecting real tumors in images sourced from different domains (e.g., hospitals). This paper made a progressive stride toward generalizable tumor synthesis by leveraging a critical observation: early-stage tumors (< 2cm) tend to have similar imaging characteristics in computed tomography (CT), whether they originate in the liver, pancreas, or kidneys. We have ascertained that generative AI models, e.g., Diffusion Models, can create realistic tumors generalized to a range of organs even when trained on a limited number of tumor examples from only one organ. Moreover, we have shown that AI models trained on these synthetic tumors can be generalized to detect and segment real tumors from CT volumes, encompassing a broad spectrum of patient demographics, imaging protocols, and healthcare facilities. Qi Chen 0014, Xiaoxi Chen, Haorui Song, Zhiwei Xiong, Alan L. Yuille, Chen Wei 0002, Zongwei Zhou |
CVPR | 5 |
| 2024 | ViTamin: Designing Scalable Vision Models in the Vision-Language EraabstractRecent breakthroughs in vision-language models (VLMs) start a new page in the vision community. The VLMs provide stronger and more generalizable feature embeddings compared to those from ImageNet-pretrained models, thanks to the training on the large-scale Internet image-text pairs. However, despite the amazing achievement from the VLMs, vanilla Vision Transformers (ViTs) remain the default choice for the image encoder. Although pure transformer proves its effectiveness in the text encoding area, it remains questionable whether it is also the case for image encoding, especially considering that various types of networks are proposed on the ImageNet benchmark, which, unfortunately, are rarely studied in VLMs. Due to small data/model scale, the original conclusions of model design on ImageNet can be limited and biased. In this paper, we aim at building an evaluation protocol of vision models in the vision-language era under the contrastive language-image pretraining (CLIP) framework. We provide a comprehensive way to benchmark different vision models, covering their zero-shot performance and scalability in both model and training data sizes. To this end, we introduce ViTamin, a new vision models tailored for VLMs. ViTamin-L significantly outperforms ViT-L by 2.0% ImageNet zero-shot accuracy, when using the same publicly available DataComp-1B dataset and the same OpenCLIP training scheme. ViTamin-L presents promising results on 60 diverse benchmarks, including classification, retrieval, open-vocabulary detection and segmentation, and large multi-modal models. When further scaling up the model size, our ViTamin-XL with only 436M parameters attains 82.9% ImageNet zero-shot accuracy, surpassing 82.0% achieved by EVA-E that has ten times more parameters (4.4B). Jieneng Chen, Qihang Yu, Xiaohui Shen, Alan L. Yuille, Liang-Chieh Chen |
CVPR | 4 |
| 2024 | A Bayesian Approach to OOD Robustness in Image ClassificationabstractAn important and unsolved problem in computer vision is to ensure that the algorithms are robust to changes in image domains. We address this problem in the scenario where we only have access to images from the target domains. Motivated by the challenges of the OOD-CV [45] benchmark where we encounter real world Out-of-Domain (OOD) nuisances and occlusion, we introduce a novel Bayesian approach to OOD robustness for object classification. Our work extends Compositional Neural Networks (CompNets), which have been shown to be robust to occlusion but degrade badly when tested on OOD data. We exploit the fact that CompNets contain a generative head defined over feature vectors represented by von Mises-Fisher (vMF) kernels, which correspond roughly to object parts, and can be learned without supervision. We obverse that some vMF kernels are similar between different domains, while others are not. This enables us to learn a transitional dictionary of vMF kernels that are intermediate between the source and target domains and train the generative model on this dictionary using the annotations on the source domain, followed by iterative refinement. This approach, termed Unsupervised Generative Transition (UGT), performs very well in OOD scenarios even when occlusion is present. UGT is evaluated on different OOD benchmarks including the OOD-CV dataset, several popular datasets (e.g., ImageNet-C [9]), artificial image corruptions (including adding occluders), and synthetic-to-real domain transfer, and does well in all scenarios. Prakhar Kaushik, Adam Kortylewski, Alan L. Yuille |
CVPR | 3 |
| 2024 | DIRECT-3D: Learning Direct Text-to-3D Generation on Massive Noisy 3D DataabstractWe present DIRECT-3D, a diffusion-based 3D generative model for creating high-quality 3D assets (represented by Neural Radiance Fields) from text prompts. Unlike recent 3D generative models that rely on clean and well-aligned 3D data, limiting them to single or few-class generation, our model is directly trained on extensive noisy and unaligned ‘in-the-wild’ 3D assets, mitigating the key challenge (i.e., data scarcity) in large-scale 3D generation. In particular, DIRECT-3D is a tri-plane diffusion model that integrates two innovations: 1) A novel learning framework where noisy data are filtered and aligned automatically during the training process. Specifically, after an initial warm-up phase using a small set of clean data, an iterative optimization is introduced in the diffusion process to explicitly estimate the 3D pose of objects and select beneficial data based on conditional density. 2) An efficient 3D representation that is achieved by disentangling object geometry and color features with two separate conditional diffusion models that are optimized hierarchically. Given a prompt input, our model generates high-quality, high-resolution, realistic, and complex 3D objects with accurate geometric details in seconds. We achieve state-of-the-art performance in both single-class generation and text-to-3D generation. We also demonstrate that DIRECT-3D can serve as a useful 3D geometric prior of objects, for example to alleviate the well-known Janus problem in 2D-lifting methods such as DreamFusion. The code and models are available for research purposes at: https://github.com/qihao067/direct3d. Qihao Liu, Yi Zhang 0099, Song Bai 0001, Adam Kortylewski, Alan L. Yuille |
CVPR | 5 |
| 2024 | Causal-CoG: A Causal-Effect Look at Context Generation for Boosting Multi-Modal Language ModelsabstractWhile Multi-modal Language Models (MLMs) demonstrate impressive multimodal ability, they still struggle on providing factual and precise responses for tasks like visual question answering (VQA). In this paper, we address this challenge from the perspective of contextual information. We propose Causal Context Generation, Causal-CoG, which is a prompting strategy that engages contextual information to enhance precise VQA during inference. Specifically, we prompt MLMs to generate contexts, i.e, text description of an image, and engage the generated contexts for question answering. Moreover, we investigate the ad-vantage of contexts on VQA from a causality perspective, introducing causality filtering to select samples for which contextual information is helpful. To show the effectiveness of Causal-CoG, we run extensive experiments on 10 multimodal benchmarks and show consistent improvements, e.g., +6.30% on POPE, +13.69% on Vizwiz and +6.43% on VQAv2 compared to direct decoding, surpassing existing methods. We hope Casual-CoG inspires explorations of context knowledge in multimodal models, and serves as a plug-and-play strategy for MLM decoding.11Code is released zhaoshitian/Causal-CoG Shitian Zhao, Zhuowan Li, Yadong Lu, Alan L. Yuille, Yan Wang 0033 |
CVPR | 4 |
| 2024 | Localization vs. Semantics: Visual Representations in Unimodal and Multimodal ModelsabstractDespite the impressive advancements achieved through vision-and-language pretraining, it remains unclear whether this joint learning paradigm can help understand each individual modality.In this work, we conduct a comparative analysis of the visual representations in existing vision-and-language models and visiononly models by probing a broad range of tasks, aiming to assess the quality of the learned representations in a nuanced manner.Interestingly, our empirical observations suggest that visionand-language models are better at label prediction tasks like object and attribute prediction, while vision-only models are stronger at dense prediction tasks that require more localized information.We hope our study sheds light on the role of language in visual learning, and serves as an empirical guide for various pretrained models.Code will be released at https: //github.com/Lizw14/visual_probing. Zhuowan Li, Cihang Xie, Benjamin Van Durme, Alan L. Yuille |
EACL (1) | 4 |
| 2024 | Radiative Gaussian Splatting for Efficient X-Ray Novel View Synthesis
Yuanhao Cai, Yixun Liang, Jiahao Wang 0001, Angtian Wang, Yulun Zhang 0001, Xiaokang Yang 0001, Zongwei Zhou, Alan L. Yuille |
ECCV (1) | 8 |
| 2024 | iNeMo: Incremental Neural Mesh Models for Robust Class-Incremental Learning
Tom Fischer, Yaoyao Liu 0001, Artur Jesslen, Prakhar Kaushik, Angtian Wang, Alan L. Yuille, Adam Kortylewski, Eddy Ilg |
ECCV (77) | 7 |
| 2024 | NOVUM: Neural Object Volumes for Robust Object Classification
Artur Jesslen, Guofeng Zhang 0025, Angtian Wang, Wufei Ma, Alan L. Yuille, Adam Kortylewski |
ECCV (4) | 5 |
| 2024 | Rethinking Video-Text Understanding: Retrieval from Counterfactually Augmented Data
Wufei Ma, Kai Li 0016, Zhongshi Jiang, Moustafa Meshry, Qihao Liu, Christian Häne, Alan L. Yuille |
ECCV (13) | 8 |
| 2024 | SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference
Feng Wang 0047, Jieru Mei, Alan L. Yuille |
ECCV (21) | 3 |
| 2024 | A Semantic Space is Worth 256 Language Descriptions: Make Stronger Segmentation Models with Descriptive Properties
Junfei Xiao, Shiyi Lan, Jieru Mei, Zhiding Yu, Bingchen Zhao, Alan L. Yuille, Yuyin Zhou, Cihang Xie |
ECCV (38) | 8 |
| 2024 | From Pixels to Objects: A Hierarchical Approach for Part and Object Segmentation Using Local and Global Aggregation
Yunfei Xie, Cihang Xie, Alan L. Yuille, Jieru Mei |
ECCV (66) | 3 |
| 2024 | IG Captioner: Information Gain Captioners Are Strong Zero-Shot Classifiers
Siyuan Qiao, Yuan Cao 0007, Yu Zhang 0033, Tao Zhu 0005, Alan L. Yuille |
ECCV (64) | 6 |
| 2024 | Source-Free and Image-Only Unsupervised Domain Adaptation for Category Level Object Pose EstimationabstractWe consider the problem of source-free unsupervised category-level 3D pose estimation from only RGB images to an non-annotated and unlabelled target domain without any access to source domain data or annotations during adaptation. Collecting and annotating real world 3D data and corresponding images is laborious, expensive yet unavoidable process since even 3D pose domain adaptation methods require 3D data in the target domain. We introduce a method which is capable of adapting to a nuisance ridden target domain without any 3D data or annotations. We represent object categories as simple cuboid meshes, and harness a generative model of neural feature activations modeled as a von Mises Fisher distribution at each mesh vertex learnt using differential rendering. We focus on individual mesh vertex features and iteratively update them based on their proximity to corresponding features in the target domain. Our key insight stems from the observation that specific object subparts remain stable across out-of-domain (OOD) scenarios, enabling strategic utilization of these invariant subcomponents for effective model updates. Our model is then trained in an EM fashion alternating between updating the vertex features and feature extractor. We show that our method simulates fine-tuning on a global-pseudo labelled dataset under mild assumptions which converges to the target domain asymptotically. Through extensive empirical validation, we demonstrate the potency of our simple approach in addressing the domain shift challenge and significantly enhancing pose estimation accuracy. By accentuating robust and less changed object subcomponents, our framework contributes to the evolution of UDA techniques in the context of 3D pose estimation using only images from the target domain. Prakhar Kaushik, Aayush Mishra, Adam Kortylewski, Alan L. Yuille |
ICLR | 4 |
| 2024 | How Well Do Supervised 3D Models Transfer to Medical Imaging Tasks?abstractThe pre-training and fine-tuning paradigm has become prominent in transfer learning. For example, if the model is pre-trained on ImageNet and then fine-tuned to PASCAL, it can significantly outperform that trained on PASCAL from scratch. While ImageNet pre-training has shown enormous success, it is formed in 2D, and the learned features are for classification tasks; when transferring to more diverse tasks, like 3D image segmentation, its performance is inevitably compromised due to the deviation from the original ImageNet context. A significant challenge lies in the lack of large, annotated 3D datasets rivaling the scale of ImageNet for model pre-training. To overcome this challenge, we make two contributions. Firstly, we construct AbdomenAtlas 1.1 that comprises **9,262** three-dimensional computed tomography (CT) volumes with high-quality, per-voxel annotations of 25 anatomical structures and pseudo annotations of seven tumor types. Secondly, we develop a suite of models that are pre-trained on our AbdomenAtlas 1.1 for transfer learning. Our preliminary analyses indicate that the model trained only with 21 CT volumes, 672 masks, and 40 GPU hours has a transfer learning ability similar to the model trained with 5,050 (unlabeled) CT volumes and 1,152 GPU hours. More importantly, the transfer learning ability of supervised models can further scale up with larger annotated datasets, achieving significantly better performance than preexisting pre-trained models, irrespective of their pre-training methodologies or data sources. We hope this study can facilitate collective efforts in constructing larger 3D medical datasets and more releases of supervised pre-trained models. Alan L. Yuille, Zongwei Zhou |
ICLR | 2 |
| 2024 | Discovering Failure Modes of Text-guided Diffusion Models via Adversarial SearchabstractText-guided diffusion models (TDMs) are widely applied but can fail unexpectedly. Common failures include: _(i)_ natural-looking text prompts generating images with the wrong content, or _(ii)_ different random samples of the latent variables that generate vastly different, and even unrelated, outputs despite being conditioned on the same text prompt. In this work, we aim to study and understand the failure modes of TDMs in more detail. To achieve this, we propose SAGE, the first adversarial search method on TDMs that systematically explores the discrete prompt space and the high-dimensional latent space, to automatically discover undesirable behaviors and failure cases in image generation. We use image classifiers as surrogate loss functions during searching, and employ human inspections to validate the identified failures. For the first time, our method enables efficient exploration of both the discrete and intricate human language space and the challenging latent space, overcoming the gradient vanishing problem. Then, we demonstrate the effectiveness of SAGE on five widely used generative models and reveal four typical failure modes that have not been systematically studied before: (1) We find a variety of natural text prompts that generate images failing to capture the semantics of input texts. We further discuss the underlying causes and potential solutions based on the results. (2) We find regions in the latent space that lead to distorted images independent of the text prompt, suggesting that parts of the latent space are not well-structured. (3) We also find latent samples that result in natural-looking images unrelated to the text prompt, implying a possible misalignment between the latent and prompt spaces. (4) By appending a single adversarial token embedding to any input prompts, we can generate a variety of specified target objects, with minimal impact on CLIP scores, demonstrating the fragility of language representations. Qihao Liu, Adam Kortylewski, Yutong Bai, Song Bai 0001, Alan L. Yuille |
ICLR | 5 |
| 2024 | Generating Images with 3D Annotations Using Diffusion ModelsabstractDiffusion models have emerged as a powerful generative method, capable of producing stunning photo-realistic images from natural language descriptions. However, these models lack explicit control over the 3D structure in the generated images. Consequently, this hinders our ability to obtain detailed 3D annotations for the generated images or to craft instances with specific poses and distances. In this paper, we propose 3D Diffusion Style Transfer (3D-DST), which incorporates 3D geometry control into diffusion models. Our method exploits ControlNet, which extends diffusion models by using visual prompts in addition to text prompts. We generate images of the 3D objects taken from 3D shape repositories~(e.g., ShapeNet and Objaverse), render them from a variety of poses and viewing directions, compute the edge maps of the rendered images, and use these edge maps as visual prompts to generate realistic images. With explicit 3D geometry control, we can easily change the 3D structures of the objects in the generated images and obtain ground-truth 3D annotations automatically. This allows us to improve a wide range of vision tasks, e.g., classification and 3D pose estimation, in both in-distribution (ID) and out-of-distribution (OOD) settings. We demonstrate the effectiveness of our method through extensive experiments on ImageNet-100/200, ImageNet-R, PASCAL3D+, ObjectNet3D, and OOD-CV. The results show that our method significantly outperforms existing methods, e.g., 3.8 percentage points on ImageNet-100 using DeiT-B. Our code is available at <https://ccvl.jhu.edu/3D-DST/> Wufei Ma, Qihao Liu, Jiahao Wang 0001, Angtian Wang, Xiaoding Yuan, Yi Zhang 0099, Zihao Xiao 0001, Guofeng Zhang 0020, Beijia Lu, Ruxiao Duan, Yongrui Qi, Adam Kortylewski, Yaoyao Liu 0001, Alan L. Yuille |
ICLR | 14 |
| 2024 | Rejuvenating image-GPT as Strong Visual Representation LearnersabstractThis paper enhances image-GPT (iGPT), one of the pioneering works that introduce autoregressive pretraining to predict the next pixels for visual representation learning. Two simple yet essential changes are made. First, we shift the prediction target from raw pixels to semantic tokens, enabling a higher-level understanding of visual content. Second, we supplement the autoregressive modeling by instructing the model to predict not only the next tokens but also the visible tokens. This pipeline is particularly effective when semantic tokens are encoded by discriminatively trained models, such as CLIP. We introduce this novel approach as D-iGPT. Extensive experiments showcase that D-iGPT excels as a strong learner of visual representations: A notable achievement is its compelling performance on the ImageNet-1K dataset — by training on publicly available datasets, D-iGPT unprecedentedly achieves 90.0% top-1 accuracy with a vanilla ViT-H. Additionally, D-iGPT shows strong generalization on the downstream task. Code is available at https://github.com/OliverRensu/D-iGPT. Sucheng Ren, Zeyu Wang 0008, Hongru Zhu, Junfei Xiao, Alan L. Yuille, Cihang Xie |
ICML | 5 |
| 2024 | Embracing Massive Medical Data
Yu-Cheng Chou, Zongwei Zhou, Alan L. Yuille |
MICCAI (1) | 3 |
| 2024 | From Pixel to Cancer: Cellular Automata in Computed Tomography
Yuxiang Lai, Xiaoxi Chen, Angtian Wang, Alan L. Yuille, Zongwei Zhou |
MICCAI (1) | 4 |
| 2024 | Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?abstractHow can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified metrics, unfair comparisons, and short-term outcome pressure. As a consequence, good performance on standard benchmarks does not guarantee success in real-world scenarios. To address these problems, we present Touchstone, a large-scale collaborative segmentation benchmark of 9 types of abdominal organs. This benchmark is based on 5,195 training CT scans from 76 hospitals around the world and 5,903 testing CT scans from 11 additional hospitals. This diverse test set enhances the statistical significance of benchmark results and rigorously evaluates AI algorithms across various out-of-distribution scenarios. We invited 14 inventors of 19 AI algorithms to train their algorithms, while our team, as a third party, independently evaluated these algorithms on three test sets. In addition, we also evaluated pre-existing AI frameworks---which, differing from algorithms, are more flexible and can support different algorithms—including MONAI from NVIDIA, nnU-Net from DKFZ, and numerous other open-source frameworks. We are committed to expanding this benchmark to encourage more innovation of AI algorithms for the medical domain. Pedro R. A. S. Bassi, Yucheng Tang, Fabian Isensee, Zifu Wang, Jieneng Chen, Yu-Cheng Chou, Yannick Kirchhoff, Maximilian Rokuss, Ziyan Huang, Jin Ye 0002, Junjun He, Tassilo Wald, Constantin Ulrich, Michael Baumgartner 0001, Saikat Roy, Klaus H. Maier-Hein, Paul F. Jaeger, Yiwen Ye, Yutong Xie 0001, Ziyang Chen 0003, Yong Xia 0001, Zhaohu Xing, Lei Zhu 0003, Yousef Sadegheih, Afshin Bozorgpour, Pratibha Kumari 0001, Reza Azad, Dorit Merhof, Yuxin Du 0001, Fan Bai 0008, Tiejun Huang 0001, Bo Zhao 0015, Xiaomeng Li 0001, Hanxue Gu, Haoyu Dong 0003, Maciej A. Mazurowski, Saumya Gupta, Linshan Wu, Jiaxin Zhuang, Hao Chen 0011, Holger Roth, Daguang Xu, Matthew B. Blaschko, Sergio Decherchi, Andrea Cavalli, Alan L. Yuille, Zongwei Zhou |
NeurIPS | 52 |
| 2024 | HDR-GS: Efficient High Dynamic Range Novel View Synthesis at 1000x Speed via Gaussian SplattingabstractHigh dynamic range (HDR) novel view synthesis (NVS) aims to create photorealistic images from novel viewpoints using HDR imaging techniques. The rendered HDR images capture a wider range of brightness levels containing more details of the scene than normal low dynamic range (LDR) images. Existing HDR NVS methods are mainly based on NeRF. They suffer from long training time and slow inference speed. In this paper, we propose a new framework, High Dynamic Range Gaussian Splatting (HDR-GS), which can efficiently render novel HDR views and reconstruct LDR images with a user input exposure time. Specifically, we design a Dual Dynamic Range (DDR) Gaussian point cloud model that uses spherical harmonics to fit HDR color and employs an MLP-based tone-mapper to render LDR color. The HDR and LDR colors are then fed into two Parallel Differentiable Rasterization (PDR) processes to reconstruct HDR and LDR views. To establish the data foundation for the research of 3D Gaussian splatting-based methods in HDR NVS, we recalibrate the camera parameters and compute the initial positions for Gaussian point clouds. Comprehensive experiments show that HDR-GS surpasses the state-of-the-art NeRF-based method by 3.84 and 1.91 dB on LDR and HDR NVS while enjoying 1000$\times$ inference speed and only costing 6.3\% training time. Code and data are released at https://github.com/caiyuanhao1998/HDR-GS Yuanhao Cai, Zihao Xiao 0001, Yixun Liang, Minghan Qin, Yulun Zhang 0001, Xiaokang Yang 0001, Yaoyao Liu 0001, Alan L. Yuille |
NeurIPS | 8 |
| 2024 | Efficient Large Multi-modal Models via Visual Context CompressionabstractWhile significant advancements have been made in compressed representations for text embeddings in large language models (LLMs), the compression of visual tokens in multi-modal LLMs (MLLMs) has remained a largely overlooked area. In this work, we present the study on the analysis of redundancy concerning visual tokens and efficient training within these models. Our initial experiments
show that eliminating up to 70% of visual tokens at the testing stage by simply average pooling only leads to a minimal 3% reduction in visual question answering accuracy on the GQA benchmark, indicating significant redundancy in visual context. Addressing this, we introduce Visual Context Compressor, which reduces the number of visual tokens to enhance training and inference efficiency without sacrificing performance. To minimize information loss caused by the compression on visual tokens while maintaining training efficiency, we develop LLaVolta as a light and staged training scheme that incorporates stage-wise visual context compression to progressively compress the visual tokens from heavily to lightly compression during training, yielding no loss of information when testing. Extensive experiments demonstrate that our approach enhances the performance of MLLMs in both image-language and video-language understanding, while also significantly cutting training costs and improving inference efficiency. Jieneng Chen, Luoxin Ye, Ju He, Daniel Khashabi, Alan L. Yuille |
NeurIPS | 6 |
| 2024 | ImageNet3D: Towards General-Purpose Object-Level 3D UnderstandingabstractA vision model with general-purpose object-level 3D understanding should be capable of inferring both 2D (e.g., class name and bounding box) and 3D information (e.g., 3D location and 3D viewpoint) for arbitrary rigid objects in natural images. This is a challenging task, as it involves inferring 3D information from 2D signals and most importantly, generalizing to rigid objects from unseen categories. However, existing datasets with object-level 3D annotations are often limited by the number of categories or the quality of annotations. Models developed on these datasets become specialists for certain categories or domains, and fail to generalize. In this work, we present ImageNet3D, a large dataset for general-purpose object-level 3D understanding. ImageNet3D augments 200 categories from the ImageNet dataset with 2D bounding box, 3D pose, 3D location annotations, and image captions interleaved with 3D information. With the new annotations available in ImageNet3D, we could (i) analyze the object-level 3D awareness of visual foundation models, and (ii) study and develop general-purpose models that infer both 2D and 3D information for arbitrary rigid objects in natural images, and (iii) integrate unified 3D models with large language models for 3D-related reasoning. We consider two new tasks, probing of object-level 3D awareness and open vocabulary pose estimation, besides standard classification and pose estimation. Experimental results on ImageNet3D demonstrate the potential of our dataset in building vision models with stronger general-purpose object-level 3D understanding. Our dataset and project page are available here: https://imagenet3d.github.io. Wufei Ma, Guofeng Zhang 0020, Qihao Liu, Guanning Zeng, Adam Kortylewski, Yaoyao Liu 0001, Alan L. Yuille |
NeurIPS | 7 |
| 2024 | Neural Textured Deformable Meshes for Robust Analysis-by-SynthesisabstractHuman vision demonstrates higher robustness than current AI algorithms under out-of-distribution scenarios. It has been conjectured such robustness benefits from performing analysis-by-synthesis. Our paper formulates triple vision tasks in a consistent manner using approximate analysis-by-synthesis by render-and-compare algorithms on neural features. In this work, we introduce Neural Textured Deformable Meshes (NTDM), which involve the object model with deformable geometry that allows optimization on both camera parameters and object geometries. The deformable mesh is parameterized as a neural field, and covered by whole-surface neural texture maps, which are trained to have spatial discriminability. During inference, we extract the feature map of the test image and subsequently optimize the 3D pose and shape parameters of our model using differentiable rendering to best reconstruct the target feature map. We show that our analysis-by-synthesis is much more robust than conventional neural networks when evaluated on real-world images and even in challenging out-of-distribution scenarios, such as occlusion and domain shift. Our algorithms are competitive with standard algorithms when tested on conventional performance measures. Angtian Wang, Wufei Ma, Alan L. Yuille, Adam Kortylewski |
WACV | 3 |
| 2024 | Robust Category-Level 3D Pose Estimation from Diffusion-Enhanced Synthetic DataabstractObtaining accurate 3D object poses is vital for numerous computer vision applications, such as 3D reconstruction and scene understanding. However, annotating real-world objects is time-consuming and challenging. While synthetically generated training data is a viable alternative, the domain shift between real and synthetic data is a significant challenge. In this work, we aim to narrow the performance gap between models trained on synthetic data and fully supervised models trained on a large amount of real data. We achieve this by approaching the problem from two perspectives: 1) We introduce P3D-Diffusion, a new synthetic dataset with accurate 3D annotations generated with a graphics-guided diffusion model. 2) We propose Cross-domain 3D Consistency, CC3D, for unsupervised domain adaptation of neural mesh models. In particular, we exploit the spatial relationships between features on the mesh surface and a contrastive learning scheme to guide the domain adaptation process. Combined, these two approaches enable our models to perform competitively with state-of-the-art models using only 10% of the respective real training images, while outperforming the SOTA model by a wide margin using only 50% of the real training data. By encouraging the diversity of synthetic data and generating the images with an OOD-aware manner, our model further demonstrates robust generalization to out-of-distribution scenarios despite being trained with minimal real data. The code is available at https://github.com/YangYY06/synthetic_3d. Wufei Ma, Angtian Wang, Xiaoding Yuan, Alan L. Yuille, Adam Kortylewski |
WACV | 5 |
| 2024 | TransUNet: Rethinking the U-Net architecture design for medical image segmentation through the lens of transformersabstractMedical image segmentation is crucial for healthcare, yet convolution-based methods like U-Net face limitations in modeling long-range dependencies. To address this, Transformers designed for sequence-to-sequence predictions have been integrated into medical image segmentation. However, a comprehensive understanding of Transformers' self-attention in U-Net components is lacking. TransUNet, first introduced in 2021, is widely recognized as one of the first models to integrate Transformer into medical image analysis. In this study, we present the versatile framework of TransUNet that encapsulates Transformers' self-attention into two key modules: (1) a Transformer encoder tokenizing image patches from a convolution neural network (CNN) feature map, facilitating global context extraction, and (2) a Transformer decoder refining candidate regions through cross-attention between proposals and U-Net features. These modules can be flexibly inserted into the U-Net backbone, resulting in three configurations: Encoder-only, Decoder-only, and Encoder+Decoder. TransUNet provides a library encompassing both 2D and 3D implementations, enabling users to easily tailor the chosen architecture. Our findings highlight the encoder's efficacy in modeling interactions among multiple abdominal organs and the decoder's strength in handling small targets like tumors. It excels in diverse medical applications, such as multi-organ segmentation, pancreatic tumor segmentation, and hepatic vessel segmentation. Notably, our TransUNet achieves a significant average Dice improvement of 1.06% and 4.30% for multi-organ segmentation and pancreatic tumor segmentation, respectively, when compared to the highly competitive nn-UNet, and surpasses the top-1 solution in the BrasTS2021 challenge. 2D/3D Code and models are available at https://github.com/Beckschen/TransUNet and https://github.com/Beckschen/TransUNet-3D, respectively. Jieneng Chen, Jieru Mei, Xianhang Li, Yongyi Lu, Qihang Yu, Qingyue Wei, Xiangde Luo, Yutong Xie 0001, Ehsan Adeli-Mosabbeb, Yan Wang 0033, Matthew P. Lungren, Shaoting Zhang 0001, Lei Xing 0001, Le Lu 0001, Alan L. Yuille, Yuyin Zhou |
Medical Image Anal. | 15 |
| 2024 | AbdomenAtlas: A large-scale, detailed-annotated, & multi-center dataset for efficient transfer learning and open algorithmic benchmarking
Chongyu Qu, Xiaoxi Chen, Pedro R. A. S. Bassi, Yijia Shi, Yuxiang Lai, Qian Yu 0012, Huimin Xue, Yixiong Chen, Xiaorui Lin, Yutong Tang, Yining Cao, Haoqi Han, Tiezheng Zhang, Yujiu Ma, Alan L. Yuille, Zongwei Zhou |
Medical Image Anal. | 20 |
| 2024 | Universal and extensible language-vision models for organ segmentation and tumor detection from abdominal computed tomography
Jie Liu 0044, Yixiao Zhang 0001, Kang Wang 0016, Mehmet Can Yavuz, Xiaoxi Chen, Yixuan Yuan, Haoliang Li, Yang Yang 0009, Alan L. Yuille, Yucheng Tang, Zongwei Zhou |
Medical Image Anal. | 9 |
| 2024 | Exploiting Structural Consistency of Chest Anatomy for Unsupervised Anomaly Detection in Radiography ImagesabstractRadiography imaging protocols focus on particular body regions, therefore producing images of great similarity and yielding recurrent anatomical structures across patients. Exploiting this structured information could potentially ease the detection of anomalies from radiography images. To this end, we propose a Simple Space-Aware Memory Matrix for In-painting and Detecting anomalies from radiography images (abbreviated as SimSID). We formulate anomaly detection as an image reconstruction task, consisting of a space-aware memory matrix and an in-painting block in the feature space. During the training, SimSID can taxonomize the ingrained anatomical structures into recurrent visual patterns, and in the inference, it can identify anomalies (unseen/modified visual patterns) from the test image. Our SimSID surpasses the state of the arts in unsupervised anomaly detection by +8.0%, +5.0%, and +9.9% AUC scores on ZhangLab, COVIDx, and CheXpert benchmark datasets, respectively. Tiange Xiang, Yixiao Zhang 0001, Yongyi Lu, Alan L. Yuille, Chaoyi Zhang, Tom Weidong Cai, Zongwei Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | OOD-CV-v2 : An Extended Benchmark for Robustness to Out-of-Distribution Shifts of Individual Nuisances in Natural ImagesabstractEnhancing the robustness of vision algorithms in real-world scenarios is challenging. One reason is that existing robustness benchmarks are limited, as they either rely on synthetic data or ignore the effects of individual nuisance factors. We introduce OOD-CV-v2, a benchmark dataset that includes out-of-distribution examples of 10 object categories in terms of pose, shape, texture, context and the weather conditions, and enables benchmarking of models for image classification, object detection, and 3D pose estimation. In addition to this novel dataset, we contribute extensive experiments using popular baseline methods, which reveal that: 1) Some nuisance factors have a much stronger negative effect on the performance compared to others, also depending on the vision task. 2) Current approaches to enhance robustness have only marginal effects, and can even reduce robustness. 3) We do not observe significant differences between convolutional and transformer architectures. We believe our dataset provides a rich test bed to study robustness and will help push forward research in this area. Bingchen Zhao, Jiahao Wang 0001, Wufei Ma, Artur Jesslen, Siwei Yang, Shaozuo Yu, Oliver Zendel 0001, Christian Theobalt, Alan L. Yuille, Adam Kortylewski |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2024 | Guest Editorial Introduction to the Issue on Pre-Trained Models for Multi-Modality UnderstandingabstractIn the ever-evolving domain of multimedia, the significance of multi-modality understanding cannot be overstated. As multimedia content becomes increasingly sophisticated and ubiquitous, the ability to effectively combine and analyze the diverse information from different types of data, such as text, audio, image, video and point clouds, will be paramount in pushing the boundaries of what technology can achieve in understanding and interacting with the world around us. Accordingly, multi-modality understanding has attracted a tremendous amount of research, establishing itself as an emerging topic. Pre-trained models, in particular, have revolutionized this field, providing a way to leverage vast amounts of data without task-specific annotation to facilitate various downstream tasks. Wengang Zhou 0001, Jiajun Deng, Nicu Sebe, Qi Tian 0001, Alan L. Yuille, Concetto Spampinato, Zakia Hammal |
IEEE Trans. Multim. | 5 |
| 2023 | Masked Autoencoders Enable Efficient Knowledge DistillersabstractThis paper studies the potential of distilling knowledge from pre-trained models, especially Masked Autoencoders. Our approach is simple: in addition to optimizing the pixel reconstruction loss on masked inputs, we minimize the distance between the intermediate feature map of the teacher model and that of the student model. This design leads to a computationally efficient knowledge distillation framework, given 1) only a small visible subset of patches is used, and 2) the (cumbersome) teacher model only needs to be partially executed, i.e., forward propagate inputs through the first few layers, for obtaining intermediate feature maps. Compared to directly distilling fine-tuned models, distilling pre-trained models substantially improves downstream performance. For example, by distilling the knowledge from an MAE pre-trained ViT-L into a ViT-B, our method achieves 84.0% ImageNet top-1 accuracy, outperforming the baseline of directly distilling afinetuned ViT-L by 1.2%. More intriguingly, our method can robustly distill knowledge from teacher models even with extremely high masking ratios: e.g., with 95% masking ratio where merely TEN patches are visible during distillation, our ViT-B competitively attains a top-1 ImageNet accuracy of 83.6%; surprisingly, it can still secure 82.4% top-1 ImageNet accuracy by aggressively training with just FOUR visible patches (98% masking ratio). The code and models are publicly available at https://github.com/UCSC-VLAA/DMAE. Yutong Bai, Zeyu Wang 0008, Junfei Xiao, Chen Wei 0005, Alan L. Yuille, Yuyin Zhou, Cihang Xie |
CVPR | 6 |
| 2023 | Compositor: Bottom-Up Clustering and Compositing for Robust Part and Object SegmentationabstractIn this work, we present a robust approach for joint part and object segmentation. Specifically, we reformulate object and part segmentation as an optimization problem and build a hierarchical feature representation including pixel, part, and object-level embeddings to solve it in a bottom-up clustering manner. Pixels are grouped into several clusters where the part-level embeddings serve as cluster centers. Afterwards, object masks are obtained by compositing the part proposals. This bottom-up interaction is shown to be effective in integrating information from lower semantic levels to higher semantic levels. Based on that, our novel approach Compositor produces part and object segmentation masks simultaneously while improving the mask quality. Compositor achieves state-of-the-art performance on PartImageNet and Pascal-Part by outperforming previous methods by around 0.9% and 1.3% on PartImageNet, 0.4% and 1.7% on Pascal-Part in terms of part and object mIoU and demonstrates better robustness against occlusion by around 4.4% and 7.1% on part and object respectively. Ju He, Jieneng Chen, Ming-Xian Lin, Qihang Yu, Alan L. Yuille |
CVPR | 5 |
| 2023 | Label-Free Liver Tumor SegmentationabstractWe demonstrate that AI models can accurately segment liver tumors without the need for manual annotation by using synthetic tumors in CT scans. Our synthetic tumors have two intriguing advantages: (I) realistic in shape and texture, which even medical professionals can confuse with real tumors; (II) effective for training AI models, which can perform liver tumor segmentation similarly to the model trained on real tumors—this result is exciting because no existing work, using synthetic tumors only, has thus far reached a similar or even close performance to real tumors. This result also implies that manual efforts for annotating tumors voxel by voxel (which took years to create) can be significantly reduced in the future. Moreover, our synthetic tumors can automatically generate many examples of small (or even tiny) synthetic tumors and have the potential to im-prove the success rate of detecting small liver tumors, which is critical for detecting the early stages of cancer. In addition to enriching the training data, our synthesizing strategy also enables us to rigorously assess the AI robustness. Qixin Hu, Yixiong Chen, Junfei Xiao, Shuwen Sun, Jieneng Chen, Alan L. Yuille, Zongwei Zhou |
CVPR | 6 |
| 2023 | Multispectral Video Semantic Segmentation: A Benchmark Dataset and BaselineabstractRobust and reliable semantic segmentation in complex scenes is crucial for many real-life applications such as autonomous safe driving and nighttime rescue. In most approaches, it is typical to make use of RGB images as input. They however work well only in preferred weather conditions; when facing adverse conditions such as rainy, overexposure, or low-light, they often fail to deliver satisfactory results. This has led to the recent investigation into multispectral semantic segmentation, where RGB and thermal infrared (RGBT) images are both utilized as input. This gives rise to significantly more robust segmentation of image objects in complex scenes and under adverse conditions. Nevertheless, the present focus in single RGBT image input restricts existing methods from well addressing dynamic real-world scenes. Motivated by the above observations, in this paper, we set out to address a relatively new task of semantic segmentation of multispectral video input, which we refer to as Multispectral Video Semantic Segmentation, or MVSS in short. An in-house MVSeg dataset is thus curated, consisting of 738 calibrated RGB and thermal videos, accompanied by 3,545 fine-grained pixel-level semantic annotations of 26 categories. Our dataset contains a wide range of challenging urban scenes in both daytime and nighttime. Moreover, we propose an effective MVSS baseline, dubbed MVNet, which is to our knowledge the first model to jointly learn semantic representations from multispectral and temporal contexts. Comprehensive experiments are conducted using various semantic segmentation models on the MVSeg dataset. Empirically, the engagement of multispectral video input is shown to lead to significant improvement in semantic segmentation; the effectiveness of our MVNet baseline has also been verified. Wei Ji 0011, Cheng Bian, Zongwei Zhou, Jiaying Zhao, Alan L. Yuille, Li Cheng 0001 |
CVPR | 6 |
| 2023 | Super-CLEVR: A Virtual Benchmark to Diagnose Domain Robustness in Visual ReasoningabstractVisual Question Answering (VQA) models often perform poorly on out-of-distribution data and struggle on domain generalization. Due to the multi-modal nature of this task, multiple factors of variation are intertwined, making generalization difficult to analyze. This motivates us to introduce a virtual benchmark, Super-CLEVR, where different factors in VQA domain shifts can be isolated in order that their effects can be studied independently. Four factors are considered: visual complexity, question redundancy, concept distribution and concept compositionality. With controllably generated data, Super-CLEVR enables us to test VQA methods in situations where the test data differs from the training data along each of these axes. We study four existing methods, including two neural symbolic methods NSCL [45] and NSVQA [59], and two non-symbolic methods FiLM [50] and mDETR [29]; and our proposed method, probabilistic NSVQA (P-NSVQA), which extends NSVQA with uncertainty reasoning. P-NSVQA outperforms other methods on three of the four domain shift factors. Our results suggest that disentangling reasoning and perception, combined with probabilistic uncertainty, form a strong VQA model that is more robust to domain shifts. The dataset and code are released at https://github.com/Lizw14/Super-CLEVR. Zhuowan Li, Xingrui Wang, Elias Stengel-Eskin, Adam Kortylewski, Wufei Ma, Benjamin Van Durme, Alan L. Yuille |
CVPR | 7 |
| 2023 | PoseExaminer: Automated Testing of Out-of-Distribution Robustness in Human Pose and Shape EstimationabstractHuman pose and shape (HPS) estimation methods achieve remarkable results. However, current HPS bench-marks are mostly designed to test models in scenarios that are similar to the training data. This can lead to critical situations in real-world applications when the observed data differs significantly from the training data and hence is out-of-distribution (OOD). It is therefore important to test and improve the OOD robustness of HPS methods. To address this fundamental problem, we develop a simulator that can be controlled in a fine-grained manner using interpretable parameters to explore the manifold of images of human pose, e.g. by varying poses, shapes, and clothes. We introduce a learning-based testing method, termed PoseExaminer, that automatically diagnoses HPS algorithms by searching over the parameter space of human pose images to find the failure modes. Our strategy for exploring this high-dimensional parameter space is a multiagent reinforcement learning system, in which the agents collaborate to explore different parts of the parameter space. We show that our PoseExaminer discovers a variety of limitations in current state-of-the-art models that are relevant in real-world scenarios but are missed by current benchmarks. For example, it finds large regions of realistic human poses that are not predicted correctly, as well as reduced performance for humans with skinny and corpulent body shapes. In addition, we show that fine-tuning HPS methods by exploiting the failure modes found by PoseExaminer improve their robustness and even their performance on standard benchmarks by a significant margin. The code are available for research purposes at https://github.com/qihao0671PoseExaminer. Qihao Liu, Adam Kortylewski, Alan L. Yuille |
CVPR | 3 |
| 2023 | InstMove: Instance Motion for Object-centric Video SegmentationabstractDespite significant efforts, cutting-edge video segmentation methods still remain sensitive to occlusion and rapid movement, due to their reliance on the appearance of objects in the form of object embeddings, which are vulnerable to these disturbances. A common solution is to use optical flow to provide motion information, but essentially it only considers pixel-level motion, which still relies on appearance similarity and hence is often inaccurate under occlusion and fast movement. In this work, we study the instance-level motion and present InstMove, which stands for Instance Motion for Object-centric Video Segmentation. In comparison to pixel-wise motion, Inst-Move mainly relies on instance-level motion information that is free from image feature embeddings, and features physical interpretations, making it more accurate and robust toward occlusion and fast-moving objects. To better fit in with the video segmentation tasks, InstMove uses instance masks to model the physical presence of an object and learns the dynamic model through a memory network to predict its position and shape in the next frame. With only a few lines of code, InstMove can be integrated into current SOTA methods for three different video segmentation tasks and boost their performance. Specifically, we improve the previous arts by 1.5 AP on OVIS dataset, which features heavy occlusions, and 4.9 AP on YouTube VIS-Long dataset, which mainly contains fast moving objects. These results suggest that instance-level motion is robust and accurate, and hence serving as a powerful solution in complex scenarios for object-centric video segmentation. Qihao Liu, Junfeng Wu 0003, Yi Jiang 0009, Xiang Bai, Alan L. Yuille, Song Bai 0001 |
CVPR | 5 |
| 2023 | SQUID: Deep Feature In-Painting for Unsupervised Anomaly DetectionabstractRadiography imaging protocols focus on particular body regions, therefore producing images of great similarity and yielding recurrent anatomical structures across patients. To exploit this structured information, we propose the use of Space-aware Memory Queues for In-painting and Detecting anomalies from radiography images (abbreviated as SQUID). We show that SQUID can taxonomize the ingrained anatomical structures into recurrent patterns; and in the inference, it can identify anomalies (unseen/modified patterns) in the image. SQUID surpasses 13 state-of-the-art methods in unsupervised anomaly detection by at least 5 points on two chest X-ray benchmark datasets measured by the Area Under the Curve (AUC). Additionally, we have created a new dataset (DigitAnatomy), which synthesizes the spatial correlation and consistent shape in chest anatomy. We hope DigitAnatomy can prompt the development, evaluation, and interpretability of anomaly detection methods. Tiange Xiang, Yixiao Zhang 0001, Yongyi Lu, Alan L. Yuille, Chaoyi Zhang, Tom Weidong Cai, Zongwei Zhou |
CVPR | 4 |
| 2023 | Diffusion Models as Masked AutoencodersabstractThere has been a longstanding belief that generation can facilitate a true understanding of visual data. In line with this, we revisit generatively pre-training visual representations in light of recent interest in denoising diffusion models. While directly pre-training with diffusion models does not produce strong representations, we condition diffusion models on masked input and formulate diffusion models as masked autoencoders (DiffMAE). Our approach is capable of (i) serving as a strong initialization for downstream recognition tasks, (ii) conducting high-quality image inpainting, and (iii) being effortlessly extended to video where it produces state-of-the-art classification accuracy. We further perform a comprehensive study on the pros and cons of design choices and build connections between diffusion models and masked autoencoders. Project page. Chen Wei 0005, Karttikeya Mangalam, Po-Yao Huang 0001, Yanghao Li, Haoqi Fan 0001, Hu Xu 0001, Cihang Xie, Alan L. Yuille, Christoph Feichtenhofer |
ICCV | 9 |
| 2023 | CancerUniT: Towards a Single Unified Model for Effective Detection, Segmentation, and Diagnosis of Eight Major Cancers Using a Large Collection of CT ScansabstractHuman readers or radiologists routinely perform full-body multi-organ multi-disease detection and diagnosis in clinical practice, while most medical AI systems are built to focus on single organs with a narrow list of a few diseases. This might severely limit AI’s clinical adoption. A certain number of AI models need to be assembled nontrivially to match the diagnostic process of a human reading a CT scan. In this paper, we construct a Unified Tumor Transformer (CancerUniT) model to jointly detect tumor existence & location and diagnose tumor characteristics for eight major cancers in CT scans. CancerUniT is a query-based Mask Transformer model with the output of multi-tumor prediction. We decouple the object queries into organ queries, tumor detection queries and tumor diagnosis queries, and further establish hierarchical relationships among the three groups. This clinically-inspired architecture effectively assists inter- and intra-organ representation learning of tumors and facilitates the resolution of these complex, anatomically related multi-organ cancer image reading tasks. CancerUniT is trained end-to-end using a curated large-scale CT images of 10,042 patients including eight major types of cancers and occurring non-cancer tumors (all are pathology-confirmed with 3D tumor masks annotated by radiologists). On the test set of 631 patients, CancerUniT has demonstrated strong performance under a set of clinically relevant evaluation metrics, substantially outperforming both multi-disease methods and an assembly of eight single-organ expert models in tumor detection, segmentation, and diagnosis. This moves one step closer towards a universal high performance cancer screening tool. Jieneng Chen, Yingda Xia, Jiawen Yao, Ke Yan 0006, Le Lu 0001, Fakai Wang, Bo Zhou 0009, Mingyan Qiu, Qihang Yu, Mingze Yuan, Wei Fang 0005, Yuxing Tang, Minfeng Xu, Xianghua Ye, Xiaoli Yin, Xin Chen 0058, Jingren Zhou 0001, Alan L. Yuille, Zaiyi Liu, Ling Zhang 0002 |
ICCV | 23 |
| 2023 | SMAUG: Sparse Masked Autoencoder for Efficient Video-Language Pre-trainingabstractVideo-language pre-training is crucial for learning powerful multi-modal representation. However, it typically requires a massive amount of computation. In this paper, we develop SMAUG, an efficient pre-training framework for video-language models. The foundation component in SMAUG is masked autoencoders. Different from prior works which only mask textual inputs, our masking strategy considers both visual and textual modalities, providing a better cross-modal alignment and saving more pre-training costs. On top of that, we introduce a space-time token sparsification module, which leverages context information to further select only "important" spatial regions and temporal frames for pre-training. Coupling all these designs allows our method to enjoy both competitive performances on text-to-video retrieval and video question answering tasks, and much less pre-training costs by 1.9× or more. For example, our SMAUG only needs ~50 NVIDIA A6000 GPU hours for pre-training to attain competitive performances on these two video-language tasks across six popular benchmarks. Yuanze Lin, Chen Wei 0005, Alan L. Yuille, Cihang Xie |
ICCV | 4 |
| 2023 | CLIP-Driven Universal Model for Organ Segmentation and Tumor DetectionabstractAn increasing number of public datasets have shown a marked impact on automated organ segmentation and tumor detection. However, due to the small size and partially labeled problem of each dataset, as well as a limited investigation of diverse types of tumors, the resulting models are often limited to segmenting specific organs/tumors and ignore the semantics of anatomical structures, nor can they be extended to novel domains. To address these issues, we propose the CLIP-Driven Universal Model, which incorporates text embedding learned from Contrastive Language-Image Pre-training (CLIP) to segmentation models. This CLIP-based label encoding captures anatomical relationships, enabling the model to learn a structured feature embedding and segment 25 organs and 6 types of tumors. The proposed model is developed from an assembly of 14 datasets, using a total of 3,410 CT scans for training and then evaluated on 6,162 external CT scans from 3 additional datasets. We rank first on the Medical Segmentation Decathlon (MSD) public leaderboard and achieve state-of-the-art results on Beyond The Cranial Vault (BTCV). Additionally, the Universal Model is computationally more efficient (6× faster) compared with dataset-specific models, generalized better to CT scans from varying sites, and shows stronger transfer learning performance on novel tasks. Jie Liu 0044, Yixiao Zhang 0001, Jieneng Chen, Junfei Xiao, Yongyi Lu, Bennett A. Landman, Yixuan Yuan, Alan L. Yuille, Yucheng Tang, Zongwei Zhou |
ICCV | 8 |
| 2023 | Animal3D: A Comprehensive Dataset of 3D Animal Pose and ShapeabstractAccurately estimating the 3D pose and shape is an essential step towards understanding animal behavior, and can potentially benefit many downstream applications, such as wildlife conservation. However, research in this area is held back by the lack of a comprehensive and diverse dataset with high-quality 3D pose and shape annotations. In this paper, we propose Animal3D, the first comprehensive dataset for mammal animal 3D pose and shape estimation. Animal3D consists of 3379 images collected from 40 mammal species, high-quality annotations of 26 key-points, and importantly the pose and shape parameters of the SMAL [50] model. All annotations were labeled and checked manually in a multi-stage process to ensure highest quality results. Based on the Animal3D dataset, we benchmark representative shape and pose estimation models at: (1) supervised learning from only the Animal3D data, (2) synthetic to real transfer from synthetically generated images, and (3) fine-tuning human pose and shape estimation models. Our experimental results demonstrate that predicting the 3D shape and pose of animals across species remains a very challenging task, despite significant advances in human pose estimation. Our results further demonstrate that synthetic pre-training is a viable strategy to boost the model performance. Overall, Animal3D opens new directions for facilitating future research in animal 3D pose and shape estimation, and is publicly available. Jiacong Xu, Yi Zhang 0099, Wufei Ma, Artur Jesslen, Pengliang Ji, Qixin Hu, Qihao Liu, Jiahao Wang 0001, Wei Ji 0011, Chen Wang 0049, Xiaoding Yuan, Prakhar Kaushik, Guofeng Zhang 0020, Jie Liu 0044, Yushan Xie, Yawen Cui, Alan L. Yuille, Adam Kortylewski |
ICCV | 19 |
| 2023 | 3D-Aware Neural Body Fitting for Occlusion Robust 3D Human Pose EstimationabstractRegression-based methods for 3D human pose estimation directly predict the 3D pose parameters from a 2D image using deep networks. While achieving state-of-the-art performance on standard benchmarks, their performance degrades under occlusion. In contrast, optimization-based methods fit a parametric body model to 2D features in an iterative manner. The localized reconstruction loss can potentially make them robust to occlusion, but they suffer from the 2D-3D ambiguity. Motivated by the recent success of generative models in rigid object pose estimation, we propose 3D-aware Neural Body Fitting (3DNBF) - an approximate analysis-by-synthesis approach to 3D human pose estimation with SOTA performance and occlusion robustness. In particular, we propose a generative model of deep features based on a volumetric human representation with Gaussian ellipsoidal kernels emitting 3D pose-dependent feature vectors. The neural features are trained with contrastive learning to become 3D-aware and hence to overcome the 2D-3D ambiguity. Experiments show that 3DNBF outperforms other approaches on both occluded and standard benchmarks. Code is available at https://github.com/edz-o/3DNBF Yi Zhang 0099, Pengliang Ji, Angtian Wang, Jieru Mei, Adam Kortylewski, Alan L. Yuille |
ICCV | 6 |
| 2023 | Which Layer is Learning Faster? A Systematic Exploration of Layer-wise Convergence Rate for Deep Neural Networks
Yixiong Chen, Alan L. Yuille, Zongwei Zhou |
ICLR | 2 |
| 2023 | VoGE: A Differentiable Volume Renderer using Gaussian Ellipsoids for Analysis-by-Synthesis
Angtian Wang, Peng Wang 0001, Adam Kortylewski, Alan L. Yuille |
ICLR | 5 |
| 2023 | MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models
Siyuan Qiao, Qihang Yu, Xiaoding Yuan, Yukun Zhu, Alan L. Yuille, Hartwig Adam, Liang-Chieh Chen |
ICLR | 6 |
| 2023 | SwinMM: Masked Multi-view with Swin Transformers for 3D Medical Image Segmentation
Jieru Mei, Zihao Wei, Li Liu 0046, Chen Wang 0049, Shengtian Sang, Alan L. Yuille, Cihang Xie, Yuyin Zhou |
MICCAI (3) | 8 |
| 2023 | Continual Learning for Abdominal Multi-organ and Tumor Segmentation
Yixiao Zhang 0001, Huimiao Chen, Alan L. Yuille, Yaoyao Liu 0001, Zongwei Zhou |
MICCAI (2) | 4 |
| 2023 | AbdomenAtlas-8K: Annotating 8, 000 CT Volumes for Multi-Organ Segmentation in Three WeeksabstractAnnotating medical images, particularly for organ segmentation, is laborious and time-consuming. For example, annotating an abdominal organ requires an estimated rate of 30-60 minutes per CT volume based on the expertise of an annotator and the size, visibility, and complexity of the organ. Therefore, publicly available datasets for multi-organ segmentation are often limited in data size and organ diversity. This paper proposes an active learning procedure to expedite the annotation process for organ segmentation and creates the largest multi-organ dataset (by far) with the spleen, liver, kidneys, stomach, gallbladder, pancreas, aorta, and IVC annotated in 8,448 CT volumes, equating to 3.2 million slices. The conventional annotation methods would take an experienced annotator up to 1,600 weeks (or roughly 30.8 years) to complete this task. In contrast, our annotation procedure has accomplished this task in three weeks (based on an 8-hour workday, five days a week) while maintaining a similar or even better annotation quality. This achievement is attributed to three unique properties of our method: (1) label bias reduction using multiple pre-trained segmentation models, (2) effective error detection in the model predictions, and (3) attention guidance for annotators to make corrections on the most salient errors. Furthermore, we summarize the taxonomy of common errors made by AI algorithms and annotators. This allows for continuous improvement of AI and annotations, significantly reducing the annotation costs required to create large-scale datasets for a wider variety of medical imaging tasks. Code and dataset are available at https://github.com/MrGiovanni/AbdomenAtlas Chongyu Qu, Tiezheng Zhang, Hualin Qiao, Jie Liu 0044, Yucheng Tang, Alan L. Yuille, Zongwei Zhou |
NeurIPS | 6 |
| 2023 | 3D-Aware Visual Question Answering about Parts, Poses and OcclusionsabstractDespite rapid progress in Visual question answering (\textit{VQA}), existing datasets and models mainly focus on testing reasoning in 2D. However, it is important that VQA models also understand the 3D structure of visual scenes, for example to support tasks like navigation or manipulation. This includes an understanding of the 3D object pose, their parts and occlusions. In this work, we introduce the task of 3D-aware VQA, which focuses on challenging questions that require a compositional reasoning over the 3D structure of visual scenes. We address 3D-aware VQA from both the dataset and the model perspective. First, we introduce Super-CLEVR-3D, a compositional reasoning dataset that contains questions about object parts, their 3D poses, and occlusions. Second, we propose PO3D-VQA, a 3D-aware VQA model that marries two powerful ideas: probabilistic neural symbolic program execution for reasoning and deep neural networks with 3D generative representations of objects for robust visual recognition. Our experimental results show our model PO3D-VQA outperforms existing methods significantly, but we still observe a significant performance gap compared to 2D VQA benchmarks, indicating that 3D-aware VQA remains an important open research area. Xingrui Wang, Wufei Ma, Zhuowan Li, Adam Kortylewski, Alan L. Yuille |
NeurIPS | 5 |
| 2023 | CoKe: Contrastive Learning for Robust Keypoint DetectionabstractIn this paper, we introduce a contrastive learning framework for keypoint detection (CoKe). Keypoint detection differs from other visual tasks where contrastive learning has been applied because the input is a set of images in which multiple keypoints are annotated. This requires the contrastive learning to be extended such that the keypoints are represented and detected independently, which enables the contrastive loss to make the keypoint features different from each other and from the background. Our approach has two benefits: It enables us to exploit contrastive learning for keypoint detection, and by detecting each key-point independently the detection becomes more robust to occlusion compared to holistic methods, such as stacked hourglass networks, which attempt to detect all keypoints jointly. Our CoKe framework introduces several technical innovations. In particular, we introduce: (i) A clutter bank to represent non-keypoint features; (ii) a keypoint bank that stores prototypical representations of keypoints to approximate the contrastive loss between keypoints; and (iii) a cumulative moving average update to learn the key-point prototypes while training the feature extractor. Our experiments on a range of diverse datasets (PASCAL3D+, MPII, ObjectNet3D) show that our approach works as well, or better than, alternative methods for keypoint detection, even for human keypoints, for which the literature is vast. Moreover, we observe that CoKe is exceptionally robust to partial occlusion and previously unseen object poses. Yutong Bai, Angtian Wang, Adam Kortylewski, Alan L. Yuille |
WACV | 4 |
| 2023 | CORL: Compositional Representation Learning for Few-Shot ClassificationabstractFew-shot image classification consists of two consecutive learning processes: 1) In the meta-learning stage, the model acquires a knowledge base from a set of training classes. 2) During meta-testing, the acquired knowledge is used to recognize unseen classes from very few examples. Inspired by the compositional representation of objects in humans, we train a neural network architecture that explicitly represents objects as a dictionary of shared components and their spatial composition. In particular, during meta-learning, we train a knowledge base that consists of a dictionary of component representations and a dictionary of component activation maps that encode common spatial activation patterns of components. The elements of both dictionaries are shared among the training classes. During meta-testing, the representation of unseen classes is learned using the component representations and the component activation maps from the knowledge base. Finally, an attention mechanism is used to strengthen those components that are most important for each category. We demonstrate the value of our interpretable compositional learning framework for a few-shot classification using miniImageNet, tieredImageNet, CIFAR-FS, and FC100, where we achieve comparable performance. Ju He, Adam Kortylewski, Alan L. Yuille |
WACV | 3 |
| 2023 | Delving into Masked Autoencoders for Multi-Label Thorax Disease ClassificationabstractVision Transformer (ViT) has become one of the most popular neural architectures due to its great scalability, computational efficiency, and compelling performance in many vision tasks. However, ViT has shown inferior performance to Convolutional Neural Network (CNN) on medical tasks due to its data-hungry nature and the lack of an-notated medical data. In this paper, we pre-train ViTs on 266,340 chest X-rays using Masked Autoencoders (MAE) which reconstruct missing pixels from a small part of each image. For comparison, CNNs are also pre-trained on the same 266,340 X-rays using advanced self-supervised methods (e.g. MoCo v2). The results show that our pre-trained ViT performs comparably (sometimes better) to the state-of-the-art CNN (DenseNet-121) for multi-label thorax dis-ease classification. This performance is attributed to the strong recipes extracted from our empirical studies for pre-training and fine-tuning ViT. The pre-training recipe signifies that medical reconstruction requires a much smaller proportion of an image (10% vs. 25%) and a more moderate random resized crop range (0.5∼1.0 vs. 0.2∼1.0) compared with natural imaging. Furthermore, we remark that in-domain transfer learning is preferred whenever possible. The fine-tuning recipe discloses that layer-wise LR decay, RandAug magnitude, and DropPath rate are significant factors to consider. We hope that this study can direct future research on the application of Transformers to a larger variety of medical imaging tasks. Junfei Xiao, Yutong Bai, Alan L. Yuille, Zongwei Zhou |
WACV | 3 |
| 2023 | BNET: Batch Normalization With Enhanced Linear TransformationabstractBatch normalization (BN) is a fundamental unit in modern deep neural networks. However, BN and its variants focus on normalization statistics but neglect the recovery step that uses linear transformation to improve the capacity of fitting complex data distributions. In this paper, we demonstrate that the recovery step can be improved by aggregating the neighborhood of each neuron rather than just considering a single neuron. Specifically, we propose a simple yet effective method named batch normalization with enhanced linear transformation (BNET) to embed spatial contextual information and improve representation ability. BNET can be easily implemented using the depth-wise convolution and seamlessly transplanted into existing architectures with BN. To our best knowledge, BNET is the first attempt to enhance the recovery step for BN. Furthermore, BN is interpreted as a special case of BNET from both spatial and spectral views. Experimental results demonstrate that BNET achieves consistent performance gains based on various backbones in a wide range of visual tasks. Moreover, BNET can accelerate the convergence of network training and enhance spatial information by assigning important neurons with large weights accordingly. Yuhui Xu 0002, Lingxi Xie, Cihang Xie, Wenrui Dai, Jieru Mei, Siyuan Qiao, Wei Shen 0002, Hongkai Xiong, Alan L. Yuille |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2022 | Masked Feature Prediction for Self-Supervised Visual Pre-TrainingabstractWe present Masked Feature Prediction (MaskFeat) for self-supervised pre-training of video models. Our approach first randomly masks out a portion of the input sequence and then predicts the feature of the masked regions. We study five different types of features and find Histograms of Oriented Gradients (HOG), a hand-crafted feature descriptor, works particularly well in terms of both performance and efficiency. We observe that the local contrast normalization in HOG is essential for good results, which is in line with earlier work using HOG for visual recognition. Our approach can learn abundant visual knowledge and drive large-scale Transformer based models. Without using extra model weights or supervision, MaskFeat pretrained on unlabeled videos achieves unprecedented results of 86.7% with MViTv2-L on Kinetics-400, 88.3% on Kinetics 600, 80.4% on Kinetics-700, 38.8 mAP on AVA, and 75.0% on SSv2. MaskFeat further generalizes to image input, which can be interpreted as a video with a single frame and obtains competitive results on ImageN et. Chen Wei 0005, Haoqi Fan 0001, Saining Xie, Chao-Yuan Wu, Alan L. Yuille, Christoph Feichtenhofer |
CVPR | 5 |
| 2022 | Learning Part Segmentation through Unsupervised Domain Adaptation from Synthetic VehiclesabstractPart segmentations provide a rich and detailed part-level description of objects. However, their annotation requires an enormous amount of work, which makes it difficult to apply standard deep learning methods. In this paper, we propose the idea of learning part segmentation through unsupervised domain adaptation (UDA) from synthetic data. We first introduce UDA-Part, a comprehensive part segmentation dataset for vehicles that can serve as an adequate benchmark for UDA11https://qliu24.github.io/udapart/. In UDA-Part, we label parts on 3D CAD models which enables us to generate a large set of annotated synthetic images. We also annotate parts on a number of real images to provide a real test set. Secondly, to advance the adaptation of part models trained from the synthetic data to the real images, we introduce a new UDA algorithm that leverages the object's spatial structure to guide the adaptation process. Our experimental results on two real test datasets confirm the superiority of our approach over existing works, and demonstrate the promise of learning part segmentation for general objects from synthetic data. We believe our dataset provides a rich testbed to study UDA for part segmentation and will help to significantly push forward research in this area. Qing Liu 0017, Adam Kortylewski, Zhishuai Zhang, Zizhang Li, Mengqi Guo, Qihao Liu, Xiaoding Yuan, Jiteng Mu, Weichao Qiu, Alan L. Yuille |
CVPR | 10 |
| 2022 | Point-Level Region Contrast for Object Detection Pre-TrainingabstractIn this work we present point-level region contrast, a self-supervised pre-training approach for the task of object detection. This approach is motivated by the two key factors in detection: localization and recognition. While accurate localization favors models that operate at the pixel- or point-level, correct recognition typically relies on a more holistic, region-level view of objects. Incorporating this perspective in pre-training, our approach performs contrastive learning by directly sampling individual point pairs from different regions. Compared to an aggregated representation per region, our approach is more robust to the change in input region quality, and further enables us to implicitly improve initial region assignments via online knowledge distillation during training. Both advantages are important when dealing with imperfect regions encountered in the unsupervised setting. Experiments show point-level region contrast improves on state-of-the-art pre-training methods for object detection and segmentation across multiple tasks and datasets, and we provide extensive ablation studies and visualizations to aid understanding. Code will be made available. Yutong Bai, Xinlei Chen, Alexander Kirillov, Alan L. Yuille, Alexander C. Berg |
CVPR | 4 |
| 2022 | TransMix: Attend to Mix for Vision TransformersabstractMixup-based augmentation has been found to be effective for generalizing models during training, especially for Vision Transformers (ViTs) since they can easily overfit. However, previous mixup-based methods have an underlying prior knowledge that the linearly interpolated ratio of targets should be kept the same as the ratio proposed in input interpolation. This may lead to a strange phenomenon that sometimes there is no valid object in the mixed image due to the random process in augmentation but there is still response in the label space. To bridge such gap between the input and label spaces, we propose TransMix, which mixes labels based on the attention maps of Vision Transformers. The confidence of the label will be larger if the corresponding input image is weighted higher by the attention map. TransMix is embarrassingly simple and can be implemented in just a few lines of code without introducing any extra parameters and FLOPs to ViT-based models. Experimental results show that our method can consistently improve various ViT-based models at scales on ImageNet classification. After pre-trained with TransMix on ImageNet, the ViT-based models also demonstrate better transferability to semantic segmentation, object detection and instance segmentation. TransMix also exhibits to be more robust when evaluating on 4 different benchmarks. Code is publicly available at https://github.com/Beckschen/TransMix. Jieneng Chen, Shuyang Sun, Ju He, Philip Torr 0001, Alan L. Yuille, Song Bai 0001 |
CVPR | 5 |
| 2022 | SwapMix: Diagnosing and Regularizing the Over-Reliance on Visual Context in Visual Question AnsweringabstractWhile Visual Question Answering (VQA) has progressed rapidly, previous works raise concerns about robustness of current VQA models. In this work, we study the robustness of VQA models from a novel perspective: visual context. We suggest that the models over-rely on the visual context, i.e., irrelevant objects in the image, to make predictions. To diagnose the models' reliance on visual context and measure their robustness, we propose a simple yet effective perturbation technique, SwapMix. SwapMix perturbs the visual context by swapping features of irrelevant context objects with features from other objects in the dataset. Using SwapMix we are able to change answers to more than 45% of the questions for a representative VQA model. Additionally, we train the models with perfect sight and find that the context over-reliance highly depends on the quality of visual representations. In addition to diagnosing, SwapMix can also be applied as a data augmentation strategy during training in order to regularize the context over-reliance. By swapping the context object features, the model reliance on context can be suppressed effectively. Two representative VQA models are studied using SwapMix: a co-attention model MCAN and a large-scale pretrained model LXMERT. Our experiments on the popular GQA dataset show the effectiveness of SwapMix for both diagnosing model robustness, and regularizing the over-reliance on visual context. The code for our method is available at https://github.com/vipulgupta1011/swapmix Zhuowan Li, Adam Kortylewski, Yingwei Li 0002, Alan L. Yuille |
CVPR | 6 |
| 2022 | DeepFusion: Lidar-Camera Deep Fusion for Multi-Modal 3D Object DetectionabstractLidars and cameras are critical sensors that provide complementary information for 3D detection in autonomous driving. While prevalent multi-modal methods [34], [36] simply decorate raw lidar point clouds with camera features and feed them directly to existing 3D detection models, our study shows that fusing camera features with deep lidar features instead of raw points, can lead to better performance. However, as those features are often augmented and aggregated, a key challenge in fusion is how to effectively align the transformed features from two modalities. In this paper, we propose two novel techniques: InverseAug that inverses geometric-related augmentations, e.g., rotation, to enable accurate geometric alignment between lidar points and image pixels, and LearnableAlign that leverages cross-attention to dynamically capture the correlations between image and lidar features during fusion. Based on InverseAug and LearnableAlign, we develop a family of generic multi-modal 3D detection models named DeepFusion, which is more accurate than previous methods. For example, DeepFusion improves Point-Pillars, CenterPoint, and 3D-MAN baselines on Pedestrian detection for 6.7,8.9, and 6.2 LEVEL_2 APH, respectively. Notably, our models achieve state-of-the-art performance on Waymo Open Dataset, and show strong model robustness against input corruptions and out-of-distribution data. Code will be publicly available at https://github.com/tensorflow/lingvo. Yingwei Li 0002, Adams Wei Yu, Tianjian Meng, Benjamin Caine, Jiquan Ngiam, Daiyi Peng, Junyang Shen, Yifeng Lu, Denny Zhou, Quoc V. Le, Alan L. Yuille, Mingxing Tan |
CVPR | 11 |
| 2022 | A Simple Data Mixing Prior for Improving Self-Supervised LearningabstractData mixing (e.g., Mixup, Cutmix, ResizeMix) is an essential component for advancing recognition models. In this paper, we focus on studying its effectiveness in the self-supervised setting. By noticing the mixed images that share the same source images are intrinsically related to each other, we hereby propose SDMP, short for Simple Data Mixing Prior, to capture this straightforward yet essential prior, and position such mixed images as additional positive pairs to facilitate self-supervised representation learning. Our experiments verify that the proposed SDMP enables data mixing to help a set of self-supervised learning frameworks (e.g., MoCo) achieve better accuracy and out-of-distribution robustness. More notably, our SDMP is the first method that successfully leverages data mixing to improve (rather than hurt) the performance of Vision Transformers in the self-supervised setting. Code is publicly available at https://github.com/OliverRensu/SDMP. Sucheng Ren, Zhengqi Gao, Shengfeng He, Alan L. Yuille, Yuyin Zhou, Cihang Xie |
CVPR | 5 |
| 2022 | Simulated Adversarial Testing of Face Recognition ModelsabstractMost machine learning models are validated and tested on fixed datasets. This can give an incomplete picture of the capabilities and weaknesses of the model. Such weaknesses can be revealed at test time in the real world. The risks involved in such failures can be loss of profits, loss of time or even loss of life in certain critical applications. In order to alleviate this issue, simulators can be controlled in a finegrained manner using interpretable parameters to explore the semantic image manifold. In this work, we propose a framework for learning how to test machine learning algorithms using simulators in an adversarial manner in order to find weaknesses in the model before deploying it in critical scenarios. We apply this method in a face recognition setup. We show that certain weaknesses of models trained on real data can be discovered using simulated samples. Using our proposed method, we can find adversarial synthetic faces that fool contemporary face recognition models. This demonstrates the fact that these models have weaknesses that are not measured by commonly used validation datasets. We hypothesize that this type of adversarial examples are not isolated, but usually lie in connected spaces in the latent space of the simulator. We present a method to find these adversarial regions as opposed to the typical adversarial points found in the adversarial example literature. Nataniel Ruiz, Adam Kortylewski, Weichao Qiu, Cihang Xie, Sarah Adel Bargal, Alan L. Yuille, Stan Sclaroff |
CVPR | 6 |
| 2022 | Amodal Segmentation through Out-of-Task and Out-of-Distribution Generalization with a Bayesian ModelabstractAmodal completion is a visual task that humans perform easily but which is difficult for computer vision algorithms. The aim is to segment those object boundaries which are occluded and hence invisible. This task is particularly challenging for deep neural networks because data is difficult to obtain and annotate. Therefore, we formulate amodal segmentation as an out-of-task and out-of-distribution generalization problem. Specifically, we replace the fully connected classifier in neural networks with a Bayesian generative model of the neural network features. The model is trained from non-occluded images using bounding box annotations and class labels only, but is applied to generalize out-of-task to object segmentation and to generalize out-of-distribution to segment occluded objects. We demonstrate how such Bayesian models can naturally generalize beyond the training task labels when they learn a prior that models the object's background context and shape. Moreover, by leveraging an outlier process, Bayesian models can further generalize out-of-distribution to segment partially occluded objects and to predict their amodal object boundaries. Our algorithm outperforms alternative methods that use the same supervision by a large margin, and even outperforms methods where annotated amodal segmentations are used during training, when the amount of occlusion is large. Code is publicly available at https://github.com/YihongSun/Bayesian-Amodal. Yihong Sun, Adam Kortylewski, Alan L. Yuille |
CVPR | 3 |
| 2022 | Learning from Temporal Gradient for Semi-supervised Action RecognitionabstractSemi-supervised video action recognition tends to enable deep neural networks to achieve remarkable performance even with very limited labeled data. However, existing methods are mainly transferred from current image-based methods (e.g., FixMatch). Without specifically utilizing the temporal dynamics and inherent multimodal attributes, their results could be suboptimal. To better leverage the encoded temporal information in videos, we introduce temporal gradient as an additional modality for more attentive feature extraction in this paper. To be specific, our method explicitly distills the fine-grained motion representations from temporal gradient (TG) and imposes consistency across different modalities (i.e., RGB and TG). The performance of semi-supervised action recognition is significantly improved without additional computation or parameters during inference. Our method achieves the state-of-the-art performance on three video action recognition benchmarks (i.e., Kinetics-400, UCF-101, and HMDB-51) under several typical semi-supervised settings (i.e., different ratios of labeled data). Code is made available at https://github.com/lambert-x/video-semisup. Junfei Xiao, Longlong Jing, Lin Zhang 0040, Ju He, Qi She, Zongwei Zhou, Alan L. Yuille, Yingwei Li 0002 |
CVPR | 7 |
| 2022 | Lite Vision Transformer with Enhanced Self-AttentionabstractDespite the impressive representation capacity of vision transformer models, current light-weight vision transformer models still suffer from inconsistent and incorrect dense predictions at local regions. We suspect that the power of their self-attention mechanism is limited in shallower and thinner networks. We propose Lite Vision Transformer (LVT), a novel light-weight transformer network with two enhanced self-attention mechanisms to improve the model performances for mobile deployment. For the low-level features, we introduce Convolutional Self-Attention (CSA). Unlike previous approaches of merging convolution and self-attention, CSA introduces local self-attention into the convolution within a kernel of size$3\times 3$to enrich low-level features in the first stage of LVT. For the high-level features, we propose Recursive Atrous Self-Attention (RASA), which utilizes the multi-scale context when calculating the similarity map and a recursive mechanism to increase the representation capability with marginal extra parameter cost. The superiority of LVT is demonstrated on ImageNet recognition, ADE20K semantic segmentation, and COCO panoptic segmentation. The code is made publicly available11https://github.com/Chenglin-Yang/LVT. Yilin Wang 0002, Jianming Zhang 0001, He Zhang 0004, Zijun Wei, Zhe Lin 0001, Alan L. Yuille |
CVPR | 7 |
| 2022 | CMT-DeepLab: Clustering Mask Transformers for Panoptic SegmentationabstractWe propose Clustering Mask Transformer (CMT-DeepLab), a transformer-based framework for panoptic segmentation designed around clustering. It rethinks the existing transformer architectures used in segmentation and detection; CMT-DeepLab considers the object queries as cluster centers, which fill the role of grouping the pixels when applied to segmentation. The clustering is computed with an alternating procedure, by first assigning pixels to the clusters by their feature affinity, and then updating the cluster centers and pixel features. Together, these operations comprise the Clustering Mask Transformer (CMT) layer, which produces cross-attention that is denser and more consistent with the final segmentation task. CMT-DeepLab improves the performance over prior art significantly by 4.4% PQ, achieving a new state-of-the-art of 55.7% PQ on the COCO test-dev set. Qihang Yu, Dahun Kim, Siyuan Qiao, Maxwell D. Collins, Yukun Zhu, Hartwig Adam, Alan L. Yuille, Liang-Chieh Chen |
CVPR | 8 |
| 2022 | Context-Enhanced Stereo Transformer
Weiyu Guo, Zhaoshuo Li, Yongkui Yang, Zheng Wang 0027, Russell H. Taylor, Mathias Unberath, Alan L. Yuille, Yingwei Li 0002 |
ECCV (32) | 7 |
| 2022 | PartImageNet: A Large, High-Quality Dataset of Parts
Ju He, Shaokang Yang, Adam Kortylewski, Xiaoding Yuan, Jieneng Chen, Qihang Yu, Alan L. Yuille |
ECCV (8) | 10 |
| 2022 | In Defense of Image Pre-Training for Spatiotemporal Recognition
Xianhang Li, Chen Wei 0005, Jieru Mei, Alan L. Yuille, Yuyin Zhou, Cihang Xie |
ECCV (25) | 5 |
| 2022 | Explicit Occlusion Reasoning for Multi-person 3D Human Pose Estimation
Qihao Liu, Yi Zhang 0099, Song Bai 0001, Alan L. Yuille |
ECCV (5) | 4 |
| 2022 | Robust Category-Level 6D Pose Estimation with Coarse-to-Fine Rendering of Neural Features
Wufei Ma, Angtian Wang, Alan L. Yuille, Adam Kortylewski |
ECCV (9) | 3 |
| 2022 | CP2: Copy-Paste Contrastive Pretraining for Semantic Segmentation
Feng Wang 0047, Chen Wei 0005, Alan L. Yuille, Wei Shen 0002 |
ECCV (30) | 4 |
| 2022 | In Defense of Online Models for Video Instance Segmentation
Junfeng Wu 0003, Qihao Liu, Yi Jiang 0009, Song Bai 0001, Alan L. Yuille, Xiang Bai |
ECCV (28) | 5 |
| 2022 | Coarse-To-Fine Incremental Few-Shot Learning
Xiang Xiang 0001, Yuwen Tan, Alan L. Yuille, Gregory D. Hager |
ECCV (31) | 5 |
| 2022 | k-means Mask Transformer
Qihang Yu, Siyuan Qiao, Maxwell D. Collins, Yukun Zhu, Hartwig Adam, Alan L. Yuille, Liang-Chieh Chen |
ECCV (29) | 7 |
| 2022 | OOD-CV: A Benchmark for Robustness to Out-of-Distribution Shifts of Individual Nuisances in Natural Images
Bingchen Zhao, Shaozuo Yu, Wufei Ma, Mingxin Yu, Shenxiao Mei, Angtian Wang, Ju He, Alan L. Yuille, Adam Kortylewski |
ECCV (8) | 8 |
| 2022 | Fast AdvProp
Jieru Mei, Yucheng Han, Yutong Bai, Yixiao Zhang 0001, Yingwei Li 0002, Xianhang Li, Alan L. Yuille, Cihang Xie |
ICLR | 7 |
| 2022 | Image BERT Pre-training with Online Tokenizer
Jinghao Zhou, Chen Wei 0005, Wei Shen 0002, Cihang Xie, Alan L. Yuille, Tao Kong |
ICLR | 6 |
| 2022 | Occluded Video Instance Segmentation: A BenchmarkabstractAbstract Can our video understanding systems perceive objects when a heavy occlusion exists in a scene? To answer this question, we collect a large-scale dataset called OVIS for occluded video instance segmentation, that is, to simultaneously detect, segment, and track instances in occluded scenes. OVIS consists of 296k high-quality instance masks from 25 semantic categories, where object occlusions usually occur. While our human vision systems can understand those occluded instances by contextual reasoning and association, our experiments suggest that current video understanding systems cannot. On the OVIS dataset, the highest AP achieved by state-of-the-art algorithms is only 16.3, which reveals that we are still at a nascent stage for understanding objects, instances, and videos in a real-world scenario. We also present a simple plug-and-play module that performs temporal feature calibration to complement missing object cues caused by occlusion. Built upon MaskTrack R-CNN and SipMask, we obtain a remarkable AP improvement on the OVIS dataset. The OVIS dataset and project code are available at http://songbai.site/ovis . Jiyang Qi, Yan Gao 0017, Yao Hu 0002, Xinggang Wang, Xiang Bai, Serge J. Belongie, Alan L. Yuille, Philip Torr 0001, Song Bai 0001 |
Int. J. Comput. Vis. | 8 |
| 2022 | Imbalanced regression for intensity series of pain expression from videos by regularizing spatio-temporal face nets
Xiang Xiang 0001, Feng Wang 0015, Yuwen Tan, Alan L. Yuille |
Pattern Recognit. Lett. | 4 |
| 2022 | External Attention Assisted Multi-Phase Splenic Vascular Injury Segmentation With Limited DataabstractThe spleen is one of the most commonly injured solid organs in blunt abdominal trauma. The development of automatic segmentation systems from multi-phase CT for splenic vascular injury can augment severity grading for improving clinical decision support and outcome prediction. However, accurate segmentation of splenic vascular injury is challenging for the following reasons: 1) Splenic vascular injury can be highly variant in shape, texture, size, and overall appearance; and 2) Data acquisition is a complex and expensive procedure that requires intensive efforts from both data scientists and radiologists, which makes large-scale well-annotated datasets hard to acquire in general. In light of these challenges, we hereby design a novel framework for multi-phase splenic vascular injury segmentation, especially with limited data. On the one hand, we propose to leverage external data to mine pseudo splenic masks as the spatial attention, dubbed external attention, for guiding the segmentation of splenic vascular injury. On the other hand, we develop a synthetic phase augmentation module, which builds upon generative adversarial networks, for populating the internal data by fully leveraging the relation between different phases. By jointly enforcing external attention and populating internal data representation during training, our proposed method outperforms other competing methods and substantially improves the popular DeepLab-v3+ baseline by more than 7% in terms of average DSC, which confirms its effectiveness. Yuyin Zhou, David Dreizin, Yan Wang 0033, Fengze Liu, Wei Shen 0002, Alan L. Yuille |
IEEE Trans. Medical Imaging | 6 |
| 2021 | DASZL: Dynamic Action Signatures for Zero-shot LearningabstractThere are many realistic applications of activity recognition where the set of potential activity descriptions is combinatorially large. This makes end-to-end supervised training of a recognition system impractical as no training set is practically able to encompass the entire label set. In this paper, we present an approach to fine-grained recognition that models activities as compositions of dynamic action signatures. This compositional approach allows us to reframe fine-grained recognition as zero-shot activity recognition, where a detector is composed "on the fly" from simple first-principles state machines supported by deep-learned components. We evaluate our method on the Olympic Sports and UCF101 datasets, where our model establishes a new state of the art under multiple experimental paradigms. We also extend this method to form a unique framework for zero-shot joint segmentation and classification of activities in video and demonstrate the first results in zero- shot decoding of complex action sequences on a widely-used surgical dataset. Lastly, we show that we can use off-the-shelf object detectors to recognize activities in completely de-novo settings with no additional training. Tae Soo Kim 0001, Jonathan D. Jones, Michael Peven, Zihao Xiao 0001, Jin Bai 0001, Yi Zhang 0099, Weichao Qiu, Alan L. Yuille, Gregory D. Hager |
AAAI | 8 |
| 2021 | CAKES: Channel-wise Automatic KErnel Shrinking for Efficient 3D Networksabstract3D Convolution Neural Networks (CNNs) have been widely applied to 3D scene understanding, such as video analysis and volumetric image recognition. However, 3D networks can easily lead to over-parameterization which incurs expensive computation cost. In this paper, we propose Channel-wise Automatic KErnel Shrinking (CAKES), to enable efficient 3D learning by shrinking standard 3D convolutions into a set of economic operations (e.g., 1D, 2D convolutions). Unlike previous methods, CAKES performs channel-wise kernel shrinkage, which enjoys the following benefits: 1) enabling operations deployed in every layer to be heterogeneous, so that they can extract diverse and complementary information to benefit the learning process; and 2) allowing for an efficient and flexible replacement design, which can be generalized to both spatial-temporal and volumetric data. Further, we propose a new search space based on CAKES, so that the configuration can be determined automatically for simplifying 3D networks. CAKES shows superior performance to other methods with similar model size, and it also achieves comparable performance to state-of-the-art methods with much fewer parameters and computational costs on tasks including 3D medical imaging segmentation and video action recognition. Codes and models are available at https://github.com/yucornetto/CAKES Qihang Yu, Yingwei Li 0002, Jieru Mei, Yuyin Zhou, Alan L. Yuille |
AAAI | 5 |
| 2021 | Searching for TrioNet: Combining Convolution with Local and Global Self-Attention
Huaijin Pi, Yingwei Li 0002, Zizhang Li, Alan L. Yuille |
BMVC | 5 |
| 2021 | Visual Analogy: Deep Learning Versus Compositional Models
Nicholas Ichien, Qing Liu 0017, Shuhao Fu, Keith J. Holyoak, Alan L. Yuille, Hongjing Lu |
CogSci | 5 |
| 2021 | Three-dimensional pose discrimination in natural images of humans
Hongru Zhu, Alan L. Yuille, Daniel J. Kersten |
CogSci | 2 |
| 2021 | Weakly Supervised Instance Segmentation for Videos With Temporal Mask ConsistencyabstractWeakly supervised instance segmentation reduces the cost of annotations required to train models. However, existing approaches which rely only on image-level class labels predominantly suffer from errors due to (a) partial segmentation of objects and (b) missing object predictions. We show that these issues can be better addressed by training with weakly labeled videos instead of images. In videos, motion and temporal consistency of predictions across frames provide complementary signals which can help segmentation. We are the first to explore the use of these video signals to tackle weakly supervised instance segmentation. We propose two ways to leverage this information in our model. First, we adapt inter-pixel relation network (IRN) [1] to effectively incorporate motion information during training. Second, we introduce a new MaskConsist module, which addresses the problem of missing object instances by transferring stable predictions between neighboring frames during training. We demonstrate that both approaches together improve the instance segmentation metric AP50on video frames of two datasets: Youtube-VIS and Cityscapes by 5% and 3% respectively. Qing Liu 0017, Vignesh Ramanathan, Dhruv Mahajan 0001, Alan L. Yuille, Zhenheng Yang |
CVPR | 4 |
| 2021 | Deeply Shape-Guided Cascade for Instance SegmentationabstractThe key to a successful cascade architecture for precise instance segmentation is to fully leverage the relationship between bounding box detection and mask segmentation across multiple stages. Although modern instance segmentation cascades achieve leading performance, they mainly make use of a unidirectional relationship, i.e., mask segmentation can benefit from iteratively refined bounding box detection. In this paper, we investigate an alternative direction, i.e., how to take the advantage of precise mask segmentation for bounding box detection in a cascade architecture. We propose a Deeply Shape-guided Cascade (DSC) for instance segmentation, which iteratively imposes the shape guidances extracted from mask prediction at previous stage on bounding box detection at current stage. It forms a bi-directional relationship between the two tasks by introducing three key components: (1) Initial shape guidance: A mask-supervised Region Proposal Network (mPRN) with the ability to generate class-agnostic masks; (2) Explicit shape guidance: A mask-guided regionof-interest (RoI) feature extractor, which employs mask segmentation at previous stage to focus feature extraction at current stage within a region aligned well with the shape of the instance-of-interest rather than a rectangular RoI; (3) Implicit shape guidance: A feature fusion operation which feeds intermediate mask features at previous stage to the bounding box head at current stage. Experimental results show that DSC outperforms the state-of-the-art instance segmentation cascade, Hybrid Task Cascade (HTC), by a large margin and achieves 51.8 box AP and 45.5 mask AP on COCO test-dev. The code is released at: https://github.com/hding2455/DSC. Hao Ding 0021, Siyuan Qiao, Alan L. Yuille, Wei Shen 0002 |
CVPR | 3 |
| 2021 | Progressive Stage-Wise Learning for Unsupervised Feature Representation EnhancementabstractUnsupervised learning methods have recently shown their competitiveness against supervised training. Typically, these methods use a single objective to train the en-tire network. But one distinct advantage of unsupervised over supervised learning is that the former possesses more variety and freedom in designing the objective. In this work, we explore new dimensions of unsupervised learning by proposing the Progressive Stage-wise Learning (PSL) framework. For a given unsupervised task, we design multi-level tasks and define different learning stages for the deep network. Early learning stages are forced to focus on low-level tasks while late stages are guided to extract deeper information through harder tasks. We discover that by progressive stage-wise learning, unsupervised feature representation can be effectively enhanced. Our extensive experiments show that PSL consistently improves results for the leading unsupervised learning methods. Zefan Li, Chenxi Liu 0001, Alan L. Yuille, Bingbing Ni, Wenjun Zhang 0001, Wen Gao 0001 |
CVPR | 3 |
| 2021 | Self-Supervised Pillar Motion Learning for Autonomous DrivingabstractAutonomous driving can benefit from motion behavior comprehension when interacting with diverse traffic participants in highly dynamic environments. Recently, there has been a growing interest in estimating class-agnostic motion directly from point clouds. Current motion estimation methods usually require vast amount of annotated training data from self-driving scenes. However, manually labeling point clouds is notoriously difficult, error-prone and time-consuming. In this paper, we seek to answer the research question of whether the abundant unlabeled data collections can be utilized for accurate and efficient motion learning. To this end, we propose a learning framework that leverages free supervisory signals from point clouds and paired camera images to estimate motion purely via self-supervision. Our model involves a point cloud based structural consistency augmented with probabilistic motion masking as well as a cross-sensor motion regularization to realize the desired self-supervision. Experiments reveal that our approach performs competitively to supervised methods, and achieves the state-of-the-art result when combining our self-supervised model with supervised fine-tuning. Chenxu Luo, Alan L. Yuille |
CVPR | 3 |
| 2021 | DetectoRS: Detecting Objects With Recursive Feature Pyramid and Switchable Atrous ConvolutionabstractMany modern object detectors demonstrate outstanding performances by using the mechanism of looking and thinking twice. In this paper, we explore this mechanism in the backbone design for object detection. At the macro level, we propose Recursive Feature Pyramid, which incorporates extra feedback connections from Feature Pyramid Networks into the bottom-up backbone layers. At the micro level, we propose Switchable Atrous Convolution, which convolves the features with different atrous rates and gathers the results using switch functions. Combining them results in DetectoRS, which significantly improves the performances of object detection. On COCO test-dev, DetectoRS achieves state-of-the-art 55.7% box AP for object detection, 48.5% mask AP for instance segmentation, and 50.0% PQ for panoptic segmentation. The code is made publicly available1. Siyuan Qiao, Liang-Chieh Chen, Alan L. Yuille |
CVPR | 3 |
| 2021 | VIP-DeepLab: Learning Visual Perception With Depth-Aware Video Panoptic SegmentationabstractIn this paper, we present ViP-DeepLab, a unified model attempting to tackle the long-standing and challenging inverse projection problem in vision, which we model as restoring the point clouds from perspective image sequences while providing each point with instance-level semantic interpretations. Solving this problem requires the vision models to predict the spatial location, semantic class, and temporally consistent instance label for each 3D point. ViP-DeepLab approaches it by jointly performing monocular depth estimation and video panoptic segmentation. We name this joint task as Depth-aware Video Panoptic Segmentation, and propose a new evaluation metric along with two derived datasets for it, which will be made available to the public. On the individual sub-tasks, ViP-DeepLab also achieves state-of-the-art results, outperforming previous methods by 5.1% VPQ on Cityscapes-VPS, ranking 1st on the KITTI monocular depth estimation benchmark, and 1st on KITTI MOTS pedestrian. The datasets and the evaluation codes are made publicly available1. Siyuan Qiao, Yukun Zhu, Hartwig Adam, Alan L. Yuille, Liang-Chieh Chen |
CVPR | 4 |
| 2021 | MaX-DeepLab: End-to-End Panoptic Segmentation With Mask TransformersabstractWe present MaX-DeepLab, the first end-to-end model for panoptic segmentation. Our approach simplifies the current pipeline that depends heavily on surrogate sub-tasks and hand-designed components, such as box detection, non-maximum suppression, thing-stuff merging, etc. Although these sub-tasks are tackled by area experts, they fail to comprehensively solve the target task. By contrast, our MaX-DeepLab directly predicts class-labeled masks with a mask transformer, and is trained with a panoptic quality inspired loss via bipartite matching. Our mask transformer employs a dual-path architecture that introduces a global memory path in addition to a CNN path, allowing direct communication with any CNN layers. As a result, MaX-DeepLab shows a significant 7.1% PQ gain in the box-free regime on the challenging COCO dataset, closing the gap between box-based and box-free methods for the first time. A small variant of MaX-DeepLab improves 3.0% PQ over DETR with similar parameters and M-Adds. Furthermore, MaX-DeepLab, without test time augmentation, achieves new state-of-the-art 51.3% PQ on COCO test-dev set. Yukun Zhu, Hartwig Adam, Alan L. Yuille, Liang-Chieh Chen |
CVPR | 4 |
| 2021 | CReST: A Class-Rebalancing Self-Training Framework for Imbalanced Semi-Supervised LearningabstractSemi-supervised learning on class-imbalanced data, although a realistic problem, has been under studied. While existing semi-supervised learning (SSL) methods are known to perform poorly on minority classes, we find that they still generate high precision pseudo-labels on minority classes. By exploiting this property, in this work, we propose Class-Rebalancing Self-Training (CReST), a simple yet effective framework to improve existing SSL methods on class-imbalanced data. CReST iteratively retrains a baseline SSL model with a labeled set expanded by adding pseudolabeled samples from an unlabeled set, where pseudolabeled samples from minority classes are selected more frequently according to an estimated class distribution. We also propose a progressive distribution alignment to adaptively adjust the rebalancing strength dubbed CReST+. We show that CReST and CReST+ improve state-of-the-art SSL algorithms on various class-imbalanced datasets and consistently outperform other popular rebalancing methods. Code has been made available at https://github.com/google-research/crest. Chen Wei 0005, Kihyuk Sohn, Clayton Mellina, Alan L. Yuille |
CVPR | 4 |
| 2021 | Mask Guided Matting via Progressive Refinement NetworkabstractWe propose Mask Guided (MG) Matting, a robust matting framework that takes a general coarse mask as guidance. MG Matting leverages a network (PRN) design which encourages the matting model to provide self-guidance to progressively refine the uncertain regions through the decoding process. A series of guidance mask perturbation operations are also introduced in the training to further enhance its robustness to external guidance. We show that PRN can generalize to unseen types of guidance masks such as trimap and low-quality alpha matte, making it suitable for various application pipelines. In addition, we revisit the foreground color prediction problem for matting and propose a surprisingly simple improvement to address the dataset issue. Evaluation on real and synthetic benchmarks shows that MG Matting achieves state-of-the-art performance using various types of guidance inputs. Code and models are available at https://github.com/yucornetto/MGMatting. Qihang Yu, Jianming Zhang 0001, He Zhang 0004, Yilin Wang 0002, Zhe Lin 0001, Ning Xu 0007, Yutong Bai, Alan L. Yuille |
CVPR | 8 |
| 2021 | Robust Instance Segmentation Through Reasoning About Multi-Object OcclusionabstractAnalyzing complex scenes with Deep Neural Networks is a challenging task, particularly when images contain multiple objects that partially occlude each other. Existing approaches to image analysis mostly process objects independently and do not take into account the relative occlusion of nearby objects. In this paper, we propose a deep network for multi-object instance segmentation that is robust to occlusion and can be trained from bounding box supervision only. Our work builds on Compositional Networks, which learn a generative model of neural feature activations to locate occluders and to classify objects based on their non-occluded parts. We extend their generative model to include multiple objects and introduce a framework for efficient inference in challenging occlusion scenarios. In particular, we obtain feed-forward predictions of the object classes and their instance and occluder segmentations. We introduce an Occlusion Reasoning Module (ORM) that locates erroneous segmentations and estimates the occlusion order to correct them. The improved segmentation masks are, in turn, integrated into the network in a top-down manner to improve the image classification. Our experiments on the KITTI INStance dataset (KINS) and a synthetic occlusion dataset demonstrate the effectiveness and robustness of our model at multi-object instance segmentation under occlusion. Code is publically available at https://github.com/XD7479/Multi-Object-Occlusion. Xiaoding Yuan, Adam Kortylewski, Yihong Sun, Alan L. Yuille |
CVPR | 4 |
| 2021 | Calibrating Concepts and Operations: Towards Symbolic Reasoning on Real ImagesabstractWhile neural symbolic methods demonstrate impressive performance in visual question answering on synthetic images, their performance suffers on real images. We identify that the long-tail distribution of visual concepts and unequal importance of reasoning steps in real data are the two key obstacles that limit the models’ real-world potentials. To address these challenges, we propose a new paradigm, Calibrating Concepts and Operations (CCO), which enables neural symbolic models to capture underlying data characteristics and to reason with hierarchical importance. Specifically, we introduce an executor with learnable concept embedding magnitudes for handling distribution imbalance, and an operation calibrator for highlighting important operations and suppressing redundant ones.Our experiments show CCO substantially boosts the performance of neural symbolic methods on real images. By evaluating models on the real world dataset GQA, CCO helps the neural symbolic method NSCL outperforms its vanilla counterpart by 9.1% (from 47.0% to 56.1%); this result also largely reduces the performance gap between symbolic and non-symbolic methods. Additionally, we create a perturbed test set for better understanding and analyzing model performance on real images. Code is available at https://lizw14.github.io/project/ccosr. Zhuowan Li, Elias Stengel-Eskin, Yixiao Zhang 0001, Cihang Xie, Quan Tran, Benjamin Van Durme, Alan L. Yuille |
ICCV | 7 |
| 2021 | Exploring Simple 3D Multi-Object Tracking for Autonomous Drivingabstract3D multi-object tracking in LiDAR point clouds is a key ingredient for self-driving vehicles. Existing methods are predominantly based on the tracking-by-detection pipeline and inevitably require a heuristic matching step for the detection association. In this paper, we present SimTrack to simplify the hand-crafted tracking paradigm by proposing an end-to-end trainable model for joint detection and tracking from raw point clouds. Our key design is to predict the first-appear location of each object in a given snippet to get the tracking identity and then update the location based on motion estimation. In the inference, the heuristic matching step can be completely waived by a simple read-off operation. SimTrack integrates the tracked object association, newborn object detection, and dead track killing in a single unified model. We conduct extensive evaluations on two large-scale datasets: nuScenes and Waymo Open Dataset. Experimental results reveal that our simple approach compares favorably with the state-of-the-art methods while ruling out the heuristic matching rules. Chenxu Luo, Alan L. Yuille |
ICCV | 3 |
| 2021 | A-SDF: Learning Disentangled Signed Distance Functions for Articulated Shape RepresentationabstractRecent work has made significant progress on using implicit functions, as a continuous representation for 3D rigid object shape reconstruction. However, much less effort has been devoted to modeling general articulated objects. Compared to rigid objects, articulated objects have higher degrees of freedom, which makes it hard to generalize to unseen shapes. To deal with the large shape variance, we introduce Articulated Signed Distance Functions (A-SDF) to represent articulated shapes with a disentangled latent space, where we have separate codes for encoding shape and articulation. With this disentangled continuous representation, we demonstrate that we can control the articulation input and animate unseen instances with unseen joint angles. Furthermore, we propose a Test-Time Adaptation inference algorithm to adjust our model during inference. We demonstrate our model generalize well to out-of-distribution and unseen data, e.g., partial point clouds and real-world depth images. Project page with code: https://jitengmu.github.io/A-SDF/. Jiteng Mu, Weichao Qiu, Adam Kortylewski, Alan L. Yuille, Nuno Vasconcelos, Xiaolong Wang 0004 |
ICCV | 4 |
| 2021 | CO2: Consistent Contrast for Unsupervised Visual Representation Learning
Chen Wei 0005, Wei Shen 0002, Alan L. Yuille |
ICLR | 4 |
| 2021 | Shape-Texture Debiased Neural Network Training
Yingwei Li 0002, Qihang Yu, Mingxing Tan, Jieru Mei, Peng Tang 0005, Wei Shen 0002, Alan L. Yuille, Cihang Xie |
ICLR | 7 |
| 2021 | NeMo: Neural Mesh Models of Contrastive Features for Robust 3D Pose Estimation
Angtian Wang, Adam Kortylewski, Alan L. Yuille |
ICLR | 3 |
| 2021 | Sequential Learning on Liver Tumor Boundary Semantics and Prognostic Biomarker Mining
Jieneng Chen, Ke Yan 0006, Youbao Tang, Shuwen Sun, Qiuping Liu, Lingyun Huang, Jing Xiao 0006, Alan L. Yuille, Ya Zhang 0002, Le Lu 0001 |
MICCAI (7) | 10 |
| 2021 | SAME: Deformable Image Registration Based on Self-supervised Anatomical Embeddings
Fengze Liu, Ke Yan 0006, Adam P. Harrison, Dazhou Guo, Le Lu 0001, Alan L. Yuille, Lingyun Huang, Guo Tong Xie, Jing Xiao 0006, Xianghua Ye, Dakai Jin |
MICCAI (4) | 6 |
| 2021 | Effective Pancreatic Cancer Screening on Non-contrast CT Scans via Anatomy-Aware Transformers
Yingda Xia, Jiawen Yao, Le Lu 0001, Lingyun Huang, Guo Tong Xie, Jing Xiao 0006, Alan L. Yuille, Ling Zhang 0002 |
MICCAI (5) | 7 |
| 2021 | ADVM'21: 1st International Workshop on Adversarial Learning for MultimediaabstractDeep learning has achieved significant success in multimedia fields involving computer vision, natural language processing, and acoustics. However research in adversarial learning also shows that they are highly vulnerable to adversarial examples. Extensive works have demonstrated that adversarial examples could easily fool deep neural networks to wrong predictions threatening practical deep learning applications in both digital and physical world. Though challenging, discovering and harnessing adversarial attacks is beneficial for diagnosing model blind-spots and further understanding as well as improving multimedia systems in practice. In this workshop, we aim to bring together researchers from the fields of adversarial machine learning, model robustness, and explainable AI to discuss recent research and future directions for adversarial robustness of deep learning models, with a particular focus on multimedia applications, including computer vision, acoustics, etc. As far as we know, we are the first workshop to focus on adversarial learning of multimedia deep learning systems, which is of great significance and we hope will be held annually in conjunction with ACM MM. Aishan Liu, Yingwei Li 0002, Chaowei Xiao, Xun Yang 0001, Xianglong Liu 0001, Dawn Song, Dacheng Tao, Alan L. Yuille, Anima Anandkumar |
ACM Multimedia | 9 |
| 2021 | Are Transformers more robust than CNNs?abstractTransformer emerges as a powerful tool for visual recognition. In addition to demonstrating competitive performance on a broad range of visual benchmarks, recent works also argue that Transformers are much more robust than Convolutions Neural Networks (CNNs). Nonetheless, surprisingly, we find these conclusions are drawn from unfair experimental settings, where Transformers and CNNs are compared at different scales and are applied with distinct training frameworks. In this paper, we aim to provide the first fair & in-depth comparisons between Transformers and CNNs, focusing on robustness evaluations. With our unified training setup, we first challenge the previous belief that Transformers outshine CNNs when measuring adversarial robustness. More surprisingly, we find CNNs can easily be as robust as Transformers on defending against adversarial attacks, if they properly adopt Transformers' training recipes. While regarding generalization on out-of-distribution samples, we show pre-training on (external) large-scale datasets is not a fundamental request for enabling Transformers to achieve better performance than CNNs. Moreover, our ablations suggest such stronger generalization is largely benefited by the Transformer's self-attention-like architectures per se, rather than by other training setups. We hope this work can help the community better understand and benchmark the robustness of Transformers and CNNs. The code and models are publicly available at: https://github.com/ytongbai/ViTs-vs-CNNs. Yutong Bai, Jieru Mei, Alan L. Yuille, Cihang Xie |
NeurIPS | 3 |
| 2021 | Neural View Synthesis and Matching for Semi-Supervised Few-Shot Learning of 3D PoseabstractWe study the problem of learning to estimate the 3D object pose from a few labelled examples and a collection of unlabelled data. Our main contribution is a learning framework, neural view synthesis and matching, that can transfer the 3D pose annotation from the labelled to unlabelled images reliably, despite unseen 3D views and nuisance variations such as the object shape, texture, illumination or scene context. In our approach, objects are represented as 3D cuboid meshes composed of feature vectors at each mesh vertex. The model is initialized from a few labelled images and is subsequently used to synthesize feature representations of unseen 3D views. The synthesized views are matched with the feature representations of unlabelled images to generate pseudo-labels of the 3D pose. The pseudo-labelled data is, in turn, used to train the feature extractor such that the features at each mesh vertex are more invariant across varying 3D views of the object. Our model is trained in an EM-type manner alternating between increasing the 3D pose invariance of the feature extractor and annotating unlabelled data through neural view synthesis and matching. We demonstrate the effectiveness of the proposed semi-supervised learning framework for 3D pose estimation on the PASCAL3D+ and KITTI datasets. We find that our approach outperforms all baselines by a wide margin, particularly in an extreme few-shot setting where only 7 annotated images are given. Remarkably, we observe that our model also achieves an exceptional robustness in out-of-distribution scenarios that involve partial occlusion. Angtian Wang, Shenxiao Mei, Alan L. Yuille, Adam Kortylewski |
NeurIPS | 3 |
| 2021 | Glance-and-Gaze Vision TransformerabstractRecently, there emerges a series of vision Transformers, which show superior performance with a more compact model size than conventional convolutional neural networks, thanks to the strong ability of Transformers to model long-range dependencies. However, the advantages of vision Transformers also come with a price: Self-attention, the core part of Transformer, has a quadratic complexity to the input sequence length. This leads to a dramatic increase of computation and memory cost with the increase of sequence length, thus introducing difficulties when applying Transformers to the vision tasks that require dense predictions based on high-resolution feature maps.In this paper, we propose a new vision Transformer, named Glance-and-Gaze Transformer (GG-Transformer), to address the aforementioned issues. It is motivated by the Glance and Gaze behavior of human beings when recognizing objects in natural scenes, with the ability to efficiently model both long-range dependencies and local context. In GG-Transformer, the Glance and Gaze behavior is realized by two parallel branches: The Glance branch is achieved by performing self-attention on the adaptively-dilated partitions of the input, which leads to a linear complexity while still enjoying a global receptive field; The Gaze branch is implemented by a simple depth-wise convolutional layer, which compensates local image context to the features obtained by the Glance mechanism. We empirically demonstrate our method achieves consistently superior performance over previous state-of-the-art Transformers on various vision tasks and benchmarks. Qihang Yu, Yingda Xia, Yutong Bai, Yongyi Lu, Alan L. Yuille, Wei Shen 0002 |
NeurIPS | 5 |
| 2021 | Compositional Convolutional Neural Networks: A Robust and Interpretable Model for Object Recognition Under Occlusion
Adam Kortylewski, Qing Liu 0017, Angtian Wang, Yihong Sun, Alan L. Yuille |
Int. J. Comput. Vis. | 5 |
| 2021 | Deep Nets: What have They Ever Done for Vision?
Alan L. Yuille, Chenxi Liu 0001 |
Int. J. Comput. Vis. | 1 |
| 2021 | Deep Differentiable Random Forests for Age EstimationabstractAge estimation from facial images is typically cast as a label distribution learning or regression problem, since aging is a gradual progress. Its main challenge is the facial feature space w.r.t. ages is inhomogeneous, due to the large variation in facial appearance across different persons of the same age and the non-stationary property of aging. In this paper, we propose two Deep Differentiable Random Forests methods, Deep Label Distribution Learning Forest (DLDLF) and Deep Regression Forest (DRF), for age estimation. Both of them connect split nodes to the top layer of convolutional neural networks (CNNs) and deal with inhomogeneous data by jointly learning input-dependent data partitions at the split nodes and age distributions at the leaf nodes. This joint learning follows an alternating strategy: (1) Fixing the leaf nodes and optimizing the split nodes and the CNN parameters by Back-propagation; (2) Fixing the split nodes and optimizing the leaf nodes by Variational Bounding. Two Deterministic Annealing processes are introduced into the learning of the split and leaf nodes, respectively, to avoid poor local optima and obtain better estimates of tree parameters free of initial values. Experimental results show that DLDLF and DRF achieve state-of-the-art performance on three age estimation datasets. Wei Shen 0002, Yilu Guo, Yan Wang 0033, Kai Zhao 0012, Bo Wang 0044, Alan L. Yuille |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2021 | Learning Inductive Attention Guidance for Partially Supervised Pancreatic Ductal Adenocarcinoma PredictionabstractPancreatic ductal adenocarcinoma (PDAC) is the third most common cause of cancer death in the United States. Predicting tumors like PDACs (including both classification and segmentation) from medical images by deep learning is becoming a growing trend, but usually a large number of annotated data are required for training, which is very labor-intensive and time-consuming. In this paper, we consider a partially supervised setting, where cheap image-level annotations are provided for all the training data, and the costly per-voxel annotations are only available for a subset of them. We propose an Inductive Attention Guidance Network (IAG-Net) to jointly learn a global image-level classifier for normal/PDAC classification and a local voxel-level classifier for semi-supervised PDAC segmentation. We instantiate both the global and the local classifiers by multiple instance learning (MIL), where the attention guidance, indicating roughly where the PDAC regions are, is the key to bridging them: For global MIL based normal/PDAC classification, attention serves as a weight for each instance (voxel) during MIL pooling, which eliminates the distraction from the background; For local MIL based semi-supervised PDAC segmentation, the attention guidance is inductive, which not only provides bag-level pseudo-labels to training data without per-voxel annotations for MIL training, but also acts as a proxy of an instance-level classifier. Experimental results show that our IAG-Net boosts PDAC segmentation accuracy by more than 5% compared with the state-of-the-arts. Yan Wang 0033, Peng Tang 0005, Yuyin Zhou, Wei Shen 0002, Elliot K. Fishman, Alan L. Yuille |
IEEE Trans. Medical Imaging | 6 |
| 2020 | Learning Transferable Adversarial Examples via Ghost NetworksabstractRecent development of adversarial attacks has proven that ensemble-based methods outperform traditional, non-ensemble ones in black-box attack. However, as it is computationally prohibitive to acquire a family of diverse models, these methods achieve inferior performance constrained by the limited number of models to be ensembled.In this paper, we propose Ghost Networks to improve the transferability of adversarial examples. The critical principle of ghost networks is to apply feature-level perturbations to an existing model to potentially create a huge set of diverse models. After that, models are subsequently fused by longitudinal ensemble. Extensive experimental results suggest that the number of networks is essential for improving the transferability of adversarial examples, but it is less necessary to independently train different networks and ensemble them in an intensive aggregation way. Instead, our work can be used as a computationally cheap and easily applied plug-in to improve adversarial approaches both in single-model and multi-model attack, compatible with residual and non-residual networks. By reproducing the NeurIPS 2017 adversarial competition, our method outperforms the No.1 attack submission by a large margin, demonstrating its effectiveness and efficiency. Code is available at https://github.com/LiYingwei/ghost-network. Yingwei Li 0002, Song Bai 0001, Yuyin Zhou, Cihang Xie, Zhishuai Zhang, Alan L. Yuille |
AAAI | 6 |
| 2020 | Identifying Model Weakness with Adversarial ExaminerabstractMachine learning models are usually evaluated according to the average case performance on the test set. However, this is not always ideal, because in some sensitive domains (e.g. autonomous driving), it is the worst case performance that matters more. In this paper, we are interested in systematic exploration of the input data space to identify the weakness of the model to be evaluated. We propose to use an adversarial examiner in the testing stage. Different from the existing strategy to always give the same (distribution of) test data, the adversarial examiner will dynamically select the next test data to hand out based on the testing history so far, with the goal being to undermine the model's performance. This sequence of test data not only helps us understand the current model, but also serves as constructive feedback to help improve the model in the next iteration. We conduct experiments on ShapeNet object classification. We show that our adversarial examiner can successfully put more emphasis on the weakness of the model, preventing performance estimates from being overly optimistic. Michelle Shu, Chenxi Liu 0001, Weichao Qiu, Alan L. Yuille |
AAAI | 4 |
| 2020 | When Radiology Report Generation Meets Knowledge GraphabstractAutomatic radiology report generation has been an attracting research problem towards computer-aided diagnosis to alleviate the workload of doctors in recent years. Deep learning techniques for natural image captioning are successfully adapted to generating radiology reports. However, radiology image reporting is different from the natural image captioning task in two aspects: 1) the accuracy of positive disease keyword mentions is critical in radiology image reporting in comparison to the equivalent importance of every single word in a natural image caption; 2) the evaluation of reporting quality should focus more on matching the disease keywords and their associated attributes instead of counting the occurrence of N-gram. Based on these concerns, we propose to utilize a pre-constructed graph embedding module (modeled with a graph convolutional neural network) on multiple disease findings to assist the generation of reports in this work. The incorporation of knowledge graph allows for dedicated feature learning for each disease finding and the relationship modeling between them. In addition, we proposed a new evaluation metric for radiology image reporting with the assistance of the same composed graph. Experimental results demonstrate the superior performance of the methods integrated with the proposed graph embedding module on a publicly accessible dataset (IU-RR) of chest radiographs compared with previous approaches using both the conventional evaluation metrics commonly adopted for image captioning and our proposed ones. Yixiao Zhang 0001, Xiaosong Wang 0001, Ziyue Xu 0001, Qihang Yu, Alan L. Yuille, Daguang Xu |
AAAI | 5 |
| 2020 | ASAP-Net: Attention and Structure Aware Point Cloud Sequence Segmentation
Yongyi Lu, Bo Pang 0003, Cewu Lu, Alan L. Yuille, Gongshen Liu |
BMVC | 5 |
| 2020 | Universal Physical Camouflage Attacks on Object DetectorsabstractIn this paper, we study physical adversarial attacks on object detectors in the wild. Previous works mostly craft instance-dependent perturbations only for rigid or planar objects. To this end, we propose to learn an adversarial pattern to effectively attack all instances belonging to the same object category, referred to as Universal Physical Camouflage Attack (UPC). Concretely, UPC crafts camouflage by jointly fooling the region proposal network, as well as misleading the classifier and the regressor to output errors. In order to make UPC effective for non-rigid or non-planar objects, we introduce a set of transformations for mimicking deformable properties. We additionally impose optimization constraint to make generated patterns look natural to human observers. To fairly evaluate the effectiveness of different physical-world attacks, we present the first standardized virtual database, AttackScenes, which simulates the real 3D world in a controllable and reproducible environment. Extensive experiments suggest the superiority of our proposed UPC compared with existing physical adversarial attackers not only in virtual environments (AttackScenes), but also in real-world physical environments. Lifeng Huang, Chengying Gao, Yuyin Zhou, Cihang Xie, Alan L. Yuille, Changqing Zou |
CVPR | 5 |
| 2020 | Compositional Convolutional Neural Networks: A Deep Architecture With Innate Robustness to Partial OcclusionabstractRecent work has shown that deep convolutional neural networks (DCNNs) do not generalize well under partial occlusion. Inspired by the success of compositional models at classifying partially occluded objects, we propose to integrate compositional models and DCNNs into a unified deep model with innate robustness to partial occlusion. We term this architecture Compositional Convolutional Neural Network. In particular, we propose to replace the fully connected classification head of a DCNN with a differentiable compositional model. The generative nature of the compositional model enables it to localize occluders and subsequently focus on the non-occluded parts of the object. We conduct classification experiments on artificially occluded images as well as real images of partially occluded objects from the MS-COCO dataset. The results show that DCNNs do not classify occluded objects robustly, even when trained with data that is strongly augmented with partial occlusions. Our proposed model outperforms standard DCNNs by a large margin at classifying partially occluded objects, even when it has not been exposed to occluded objects during training. Additional experiments demonstrate that CompositionalNets can also localize the occluders accurately, despite being trained with class labels only. The code and data used in this work are publicly available. Adam Kortylewski, Ju He, Qing Liu 0017, Alan L. Yuille |
CVPR | 4 |
| 2020 | Neural Architecture Search for Lightweight Non-Local NetworksabstractNon-Local (NL) blocks have been widely studied in various vision tasks. However, it has been rarely explored to embed the NL blocks in mobile neural networks, mainly due to the following challenges: 1) NL blocks generally have heavy computation cost which makes it difficult to be applied in applications where computational resources are limited, and 2) it is an open problem to discover an optimal configuration to embed NL blocks into mobile neural networks. We propose AutoNL to overcome the above two obstacles. Firstly, we propose a Lightweight Non-Local (LightNL) block by squeezing the transformation operations and incorporating compact features. With the novel design choices, the proposed LightNL block is 400 times computationally cheaper} than its conventional counterpart without sacrificing the performance. Secondly, by relaxing the structure of the LightNL block to be differentiable during training, we propose an efficient neural architecture search algorithm to learn an optimal configuration of LightNL blocks in an end-to-end manner. Notably, using only 32 GPU hours, the searched AutoNL model achieves 77.7% top-1 accuracy on ImageNet under a typical mobile setting (350M FLOPs), significantly outperforming previous mobile models including MobileNetV2 (+5.7%), FBNet (+2.8%) and MnasNet (+2.1%). Code and models are available at https://github.com/LiYingwei/AutoNL. Yingwei Li 0002, Xiaojie Jin 0004, Jieru Mei, Xiaochen Lian, Cihang Xie, Qihang Yu, Yuyin Zhou, Song Bai 0001, Alan L. Yuille |
CVPR | 10 |
| 2020 | Context-Aware Group Captioning via Self-Attention and Contrastive FeaturesabstractWhile image captioning has progressed rapidly, existing works focus mainly on describing single images. In this paper, we introduce a new task, context-aware group captioning, which aims to describe a group of target images in the context of another group of related reference images. Context-aware group captioning requires not only summarizing information from both the target and reference image group but also contrasting between them. To solve this problem, we propose a framework combining self-attention mechanism with contrastive feature construction to effectively summarize common information from each image group while capturing discriminative information between them. To build the dataset for this task, we propose to group the images and generate the group captions based on single image captions using scene graphs matching. Our datasets are constructed on top of the public Conceptual Captions dataset and our new Stock Captions dataset. Experiments on the two datasets show the effectiveness of our method on this new task. Zhuowan Li, Quan Tran, Long Mai, Zhe Lin 0001, Alan L. Yuille |
CVPR | 5 |
| 2020 | Learning From Synthetic AnimalsabstractDespite great success in human parsing, progress for parsing other deformable articulated objects, like animals, is still limited by the lack of labeled data. In this paper, we use synthetic images and ground truth generated from CAD animal models to address this challenge. To bridge the domain gap between real and synthetic images, we propose a novel consistency-constrained semi-supervised learning method (CC-SSL). Our method leverages both spatial and temporal consistencies, to bootstrap weak models trained on synthetic data with unlabeled real images. We demonstrate the effectiveness of our method on highly deformable animals, such as horses and tigers. Without using any real image label, our method allows for accurate keypoint prediction on real images. Moreover, we quantitatively show that models using synthetic data achieve better generalization performance than models trained on real images across different domains in the Visual Domain Adaptation Challenge dataset. Our synthetic dataset contains 10+ animals with diverse poses and rich ground truth, which enables us to use the multi-task learning strategy to further boost models' performance. Jiteng Mu, Weichao Qiu, Gregory D. Hager, Alan L. Yuille |
CVPR | 4 |
| 2020 | Robust Object Detection Under Occlusion With Context-Aware CompositionalNetsabstractDetecting partially occluded objects is a difficult task.Our experimental results show that deep learning approaches, such as Faster R-CNN, are not robust at object detection under occlusion.Compositional convolutional neural networks (CompositionalNets) have been shown to be robust at classifying occluded objects by explicitly representing the object as a composition of parts.In this work, we propose to overcome two limitations of Compositional-Nets which will enable them to detect partially occluded objects: 1) CompositionalNets, as well as other DCNN architectures, do not explicitly separate the representation of the context from the object itself.Under strong object occlusion, the influence of the context is amplified which can have severe negative effects for detection at test time.In order to overcome this, we propose to segment the context during training via bounding box annotations.We then use the segmentation to learn a context-aware CompositionalNet that disentangles the representation of the context and the object.2) We extend the part-based voting scheme in Compo-sitionalNets to vote for the corners of the object's bounding box, which enables the model to reliably estimate bounding boxes for partially occluded objects.Our extensive experiments show that our proposed model can detect objects robustly, increasing the detection performance of strongly occluded vehicles from PASCAL3D+ and MS-COCO by 41% and 35% respectively in absolute performance relative to Faster R-CNN. Angtian Wang, Yihong Sun, Adam Kortylewski, Alan L. Yuille |
CVPR | 4 |
| 2020 | Deep Distance Transform for Tubular Structure Segmentation in CT ScansabstractTubular structure segmentation in medical images, e.g., segmenting vessels in CT scans, serves as a vital step in the use of computers to aid in screening early stages of related diseases. But automatic tubular structure segmentation in CT scans is a challenging problem, due to issues such as poor contrast, noise and complicated background. A tubular structure usually has a cylinder-like shape which can be well represented by its skeleton and cross-sectional radii (scales). Inspired by this, we propose a geometry-aware tubular structure segmentation method, Deep Distance Transform (DDT), which combines intuitions from the classical distance transform for skeletonization and modern deep segmentation networks. DDT first learns a multi-task network to predict a segmentation mask for a tubular structure and a distance map. Each value in the map represents the distance from each tubular structure voxel to the tubular structure surface. Then the segmentation mask is refined by leveraging the shape prior reconstructed from the distance map. We apply our DDT on six medical image datasets. Results show that (1) DDT can boost tubular structure segmentation performance significantly (e.g., over 13% DSC improvement for pancreatic duct segmentation), and (2) DDT additionally provides a geometrical measurement for a tubular structure, which is important for clinical diagnosis (e.g., the cross-sectional scale of a pancreatic duct can be an indicator for pancreatic cancer). Yan Wang 0033, Fengze Liu, Jieneng Chen, Yuyin Zhou, Wei Shen 0002, Elliot K. Fishman, Alan L. Yuille |
CVPR | 8 |
| 2020 | Adversarial Examples Improve Image RecognitionabstractAdversarial examples are commonly viewed as a threat to ConvNets. Here we present an opposite perspective: adversarial examples can be used to improve image recognition models if harnessed in the right manner. We propose AdvProp, an enhanced adversarial training scheme which treats adversarial examples as additional examples, to prevent overfitting. Key to our method is the usage of a separate auxiliary batch norm for adversarial examples, as they have different underlying distributions to normal examples. We show that AdvProp improves a wide range of models on various image recognition tasks and performs better when the models are bigger. For instance, by applying AdvProp to the latest EfficientNet-B7 [28] on ImageNet, we achieve significant improvements on ImageNet (+0.7%), ImageNet-C (+6.5%), ImageNet-A (+7.0%), Stylized-ImageNet (+4.8%). With an enhanced EfficientNet-B8, our method achieves the state-of-the-art 85.5% ImageNet top-1 accuracy without extra data. This result even surpasses the best model in [20] which is trained with 3.5B Instagram images (~3000X more than ImageNet) and ~9.4X more parameters. Code and models will be made publicly available. Cihang Xie, Mingxing Tan, Boqing Gong, Jiang Wang 0001, Alan L. Yuille, Quoc V. Le |
CVPR | 5 |
| 2020 | C2FNAS: Coarse-to-Fine Neural Architecture Search for 3D Medical Image Segmentationabstract3D convolution neural networks (CNN) have been proved very successful in parsing organs or tumours in 3D medical images, but it remains sophisticated and time-consuming to choose or design proper 3D networks given different task contexts. Recently, Neural Architecture Search (NAS) is proposed to solve this problem by searching for the best network architecture automatically. However, the inconsistency between search stage and deployment stage often exists in NAS algorithms due to memory constraints and large search space, which could become more serious when applying NAS to some memory and time-consuming tasks, such as 3D medical image segmentation. In this paper, we propose a coarse-to-fine neural architecture search (C2FNAS) to automatically search a 3D segmentation network from scratch without inconsistency on network size or input size. Specifically, we divide the search procedure into two stages: 1) the coarse stage, where we search the macro-level topology of the network, i.e. how each convolution module is connected to other modules; 2) the fine stage, where we search at micro-level for operations in each cell based on previous searched macro-level topology. The coarse-to-fine manner divides the search procedure into two consecutive stages and meanwhile resolves the inconsistency. We evaluate our method on 10 public datasets from Medical Segmentation Decalthon (MSD) challenge, and achieve state-of-the-art performance with the network searched using one dataset, which demonstrates the effectiveness and generalization of our searched models. Qihang Yu, Dong Yang 0005, Holger Roth, Yutong Bai, Yixiao Zhang 0001, Alan L. Yuille, Daguang Xu |
CVPR | 6 |
| 2020 | Object as Hotspots: An Anchor-Free 3D Object Detection Approach via Firing of Hotspots
Qi Chen 0014, Lin Sun 0004, Zhixin Wang, Kui Jia, Alan L. Yuille |
ECCV (21) | 5 |
| 2020 | Regional Homogeneity: Towards Learning Transferable Universal Adversarial Perturbations Against Defenses
Yingwei Li 0002, Song Bai 0001, Cihang Xie, Zhenyu Liao 0003, Xiaohui Shen, Alan L. Yuille |
ECCV (11) | 6 |
| 2020 | JSSR: A Joint Synthesis, Segmentation, and Registration System for 3D Multi-modal Image Alignment of Large-Scale Pathological CT Scans
Fengze Liu, Jinzheng Cai, Yuankai Huo, Chi-Tung Cheng, Ashwin Raju, Dakai Jin, Jing Xiao 0006, Alan L. Yuille, Le Lu 0001, Chien-Hung Liao, Adam P. Harrison |
ECCV (13) | 8 |
| 2020 | Are Labels Necessary for Neural Architecture Search?
Chenxi Liu 0001, Piotr Dollár, Kaiming He, Ross B. Girshick, Alan L. Yuille, Saining Xie |
ECCV (4) | 5 |
| 2020 | Axial-DeepLab: Stand-Alone Axial-Attention for Panoptic Segmentation
Yukun Zhu, Bradley Green, Hartwig Adam, Alan L. Yuille, Liang-Chieh Chen |
ECCV (4) | 5 |
| 2020 | Synthesize Then Compare: Detecting Failures and Anomalies for Semantic Segmentation
Yingda Xia, Yi Zhang 0099, Fengze Liu, Wei Shen 0002, Alan L. Yuille |
ECCV (1) | 5 |
| 2020 | PatchAttack: A Black-Box Texture-Based Attack with Reinforcement Learning
Adam Kortylewski, Cihang Xie, Yinzhi Cao, Alan L. Yuille |
ECCV (26) | 5 |
| 2020 | AtomNAS: Fine-Grained End-to-End Neural Architecture Search
Jieru Mei, Yingwei Li 0002, Xiaochen Lian, Xiaojie Jin 0004, Alan L. Yuille, Jianchao Yang |
ICLR | 6 |
| 2020 | Intriguing Properties of Adversarial Training at Scale
Cihang Xie, Alan L. Yuille |
ICLR | 2 |
| 2020 | Probabilistic Multi-modal Trajectory Prediction with Lane Attention for Autonomous VehiclesabstractTrajectory prediction is crucial for autonomous vehicles. The planning system not only needs to know the current state of the surrounding objects but also their possible states in the future. As for vehicles, their trajectories are significantly influenced by the lane geometry and how to effectively use the lane information is of active interest. Most of the existing works use rasterized maps to explore road information, which does not distinguish different lanes. In this paper, we propose a novel instance-aware representation for lane representation. By integrating the lane features and trajectory features, a goal-oriented lane attention module is proposed to predict the future locations of the vehicle. We show that the proposed lane representation together with the lane attention module can be integrated into the widely used encoder-decoder framework to generate diverse predictions. Most importantly, each generated trajectory is associated with a probability to handle the uncertainty. Our method does not suffer from collapsing to one behavior modal and can cover diverse possibilities. Extensive experiments and ablation studies on the benchmark datasets corroborate the effectiveness of our proposed method. Notably, our proposed method ranks third place in the Argoverse motion forecasting competition at NeurIPS 20191. Chenxu Luo, Lin Sun 0004, Dariush Dabiri, Alan L. Yuille |
IROS | 4 |
| 2020 | Lymph Node Gross Tumor Volume Detection in Oncology Imaging via Relationship Learning Using Graph Neural Network
Chun-Hung Chao, Zhuotun Zhu, Dazhou Guo, Ke Yan 0006, Tsung-Ying Ho, Jinzheng Cai, Adam P. Harrison, Xianghua Ye, Jing Xiao 0006, Alan L. Yuille, Min Sun 0001, Le Lu 0001, Dakai Jin |
MICCAI (7) | 10 |
| 2020 | Domain Adaptive Relational Reasoning for 3D Multi-organ Segmentation
Shuhao Fu, Yongyi Lu, Yan Wang 0033, Yuyin Zhou, Wei Shen 0002, Elliot K. Fishman, Alan L. Yuille |
MICCAI (1) | 7 |
| 2020 | Detecting Pancreatic Ductal Adenocarcinoma in Multi-phase CT Scans via Alignment Ensemble
Yingda Xia, Qihang Yu, Wei Shen 0002, Yuyin Zhou, Elliot K. Fishman, Alan L. Yuille |
MICCAI (3) | 6 |
| 2020 | Lymph Node Gross Tumor Volume Detection and Segmentation via Distance-Based Gating Using 3D CT/PET Imaging in Radiotherapy
Zhuotun Zhu, Dakai Jin, Ke Yan 0006, Tsung-Ying Ho, Xianghua Ye, Dazhou Guo, Chun-Hung Chao, Jing Xiao 0006, Alan L. Yuille, Le Lu 0001 |
MICCAI (7) | 9 |
| 2020 | Every View Counts: Cross-View Consistency in 3D Object Detection with Hybrid-Cylindrical-Spherical VoxelizationabstractRecent voxel-based 3D object detectors for autonomous vehicles learn point cloud representations either from bird eye view (BEV) or range view (RV, a.k.a. the perspective view). However, each view has its own strengths and weaknesses. In this paper, we present a novel framework to unify and leverage the benefits from both BEV and RV. The widely-used cuboid-shaped voxels in Cartesian coordinate system only benefit learning BEV feature map. Therefore, to enable learning both BEV and RV feature maps, we introduce Hybrid-Cylindrical-Spherical voxelization. Our findings show that simply adding detection on another view as auxiliary supervision will lead to poor performance. We proposed a pair of cross-view transformers to transform the feature maps into the other view and introduce cross-view consistency loss on them. Comprehensive experiments on the challenging NuScenes Dataset validate the effectiveness of our proposed method by virtue of joint optimization and complementary information on both views. Remarkably, our approach achieved mAP of 55.8%, outperforming all published approaches by at least 3% in overall performance and up to 16.5% in safety-crucial categories like cyclist. Qi Chen 0014, Lin Sun 0004, Ernest Cheung, Alan L. Yuille |
NeurIPS | 4 |
| 2020 | Adversarial Examples for Edge Detection: They Exist, and They TransferabstractConvolutional neural networks have recently advanced the state of the art in many tasks including edge and object boundary detection. However, in this paper, we demonstrate that these edge detectors inherit a troubling property of neural networks: they can be fooled by adversarial examples. We show that adding small perturbations to an image causes HED [42], a CNN-based edge detection model, to fail to locate edges, to detect nonexistent edges, and even to hallucinate arbitrary configurations of edges. More importantly, we find that these adversarial examples blindly transfer to other CNN-based vision models. In particular, attacks on edge detection result in significant drops in accuracy in models trained to perform unrelated, high-level tasks like image classification and semantic segmentation. Christian Cosgrove, Alan L. Yuille |
WACV | 2 |
| 2020 | Combining Compositional Models and Deep Networks For Robust Object Classification under OcclusionabstractDeep convolutional neural networks (DCNNs) are powerful models that yield impressive results at object classification. However, recent work has shown that they do not generalize well to partially occluded objects and to mask attacks. In contrast to DCNNs, compositional models are robust to partial occlusion, however, they are not as discriminative as deep models. In this work, we combine DC-NNs and compositional object models to retain the best of both approaches: a discriminative model that is robust to partial occlusion and mask attacks. Our model is learned in two steps. First, a standard DCNN is trained for image classification. Subsequently, we cluster the DCNN features into dictionaries. We show that the dictionary components resemble object part detectors and learn the spatial distribution of parts for each object class. We propose mixtures of compositional models to account for large changes in the spatial activation patterns (e.g. due to changes in the 3D pose of an object). At runtime, an image is first classified by the DCNN in a feedforward manner. The prediction uncertainty is used to detect partially occluded objects, which in turn are classified by the compositional model. Our experimental results demonstrate that combining compositional models and DCNNs resolves a fundamental problem of current deep learning approaches to computer vision: The combined model recognizes occluded objects, even when it has not been exposed to occluded objects during training, while at the same time maintaining high discriminative performance for non-occluded objects. Adam Kortylewski, Qing Liu 0017, Zhishuai Zhang, Alan L. Yuille |
WACV | 5 |
| 2020 | 3D Semi-Supervised Learning with Uncertainty-Aware Multi-View Co-TrainingabstractWhile making a tremendous impact in various fields, deep neural networks usually require large amounts of labeled data for training which are expensive to collect in many applications, especially in the medical domain. Un-labeled data, on the other hand, is much more abundant. Semi-supervised learning techniques, such as co-training, could provide a powerful tool to leverage unlabeled data. In this paper, we propose a novel framework, uncertainty-aware multi-view co-training (UMCT), to address semi-supervised learning on 3D data, such as volumetric data from medical imaging. In our work, co-training is achieved by exploiting multi-viewpoint consistency of 3D data. We generate different views by rotating or permuting the 3D data and utilize asymmetrical 3D kernels to encourage diversified features in different sub-networks. In addition, we propose an uncertainty-weighted label fusion mechanism to estimate the reliability of each view's prediction with Bayesian deep learning. As one view requires the supervision from other views in co-training, our self-adaptive approach computes a confidence score for the prediction of each unlabeled sample in order to assign a reliable pseudo label. Thus, our approach can take advantage of unlabeled data during training. We show the effectiveness of our proposed semi-supervised method on several public datasets from medical image segmentation tasks (NIH pancreas & LiTS liver tumor dataset). Meanwhile, a fully-supervised method based on our approach achieved state-of-the-art performances on both the LiTS liver tumor segmentation and the Medical Segmentation Decathlon (MSD) challenge, demonstrating the robustness and value of our framework, even when fully supervised training is feasible. Yingda Xia, Fengze Liu, Dong Yang 0005, Jinzheng Cai, Lequan Yu, Zhuotun Zhu, Daguang Xu, Alan L. Yuille, Holger Roth |
WACV | 8 |
| 2020 | Robust Face Detection via Learning Small Faces on Hard ImagesabstractRecent anchor-based deep face detectors have achieved promising performance, but they are still struggling to detect hard faces, such as small, blurred and partially occluded faces. One reason is that they treat all images and faces equally, and ignore the imbalance between easy images and hard images; however large amounts of training images only contain easy faces, which are less helpful to learn robust detectors for hard faces. In this paper, we propose that the robustness of a face detector against hard faces can be improved by learning small faces on hard images. Our intuitions are (1) hard images are the images which contain at least one hard face, thus they facilitate training robust face detectors; (2) most hard faces are small faces and other types of hard faces can be easily shrunk to small faces. To this end, we build an anchor-based deep face detector, which only outputs a single high-resolution feature map with small anchors, to specifically learn small faces and train it by a novel hard image mining strategy which automatically adjusts training weights on images according to their difficulties. Extensive experiments have been conducted on WIDER FACE, FDDB, Pascal Faces, and AFW datasets and our method achieves APs of 95.7, 94.9 and 89.7 on easy, medium and hard WIDER FACE val dataset respectively, which verify the effectiveness of our methods, especially on detecting hard faces. Our detector is also lightweight and enjoys a fast inference speed. Code and model are available at https://github.com/bairdzhang/smallhardface. Zhishuai Zhang, Wei Shen 0002, Siyuan Qiao, Yan Wang 0033, Bo Wang 0044, Alan L. Yuille |
WACV | 6 |
| 2020 | Resisting Large Data Variations via Introspective Transformation NetworkabstractTraining deep networks that generalize to a wide range of variations in test data is essential to building accurate and robust image classifiers. Data variations in this paper include but not limited to unseen affine transformations and warping in the training data. One standard strategy to overcome this problem is to apply data augmentation to synthetically enlarge the training set. However, data augmentation is essentially a brute-force method which generates uniform samples from some pre-defined set of transformations. In this paper, we propose a principled approach named introspective transformation network (ITN) that significantly improves network resistance to large variations between training and testing data. This is achieved by embedding a learnable transformation module into the introspective network, which is a convolutional neural network (CNN) classifier empowered with generative capabilities. Our approach alternates between synthesizing pseudo-negative samples and transformed positive examples based on the current model, and optimizing model predictions on these synthesized samples. Experimental results verify that our approach significantly improves the ability of deep networks to resist large variations between training and testing data and achieves classification accuracy improvements on several benchmark datasets, including MNIST, affNIST, SVHN, CIFAR-10 and miniImageNet. Yunhan Zhao, Charless C. Fowlkes, Wei Shen 0002, Alan L. Yuille |
WACV | 5 |
| 2020 | Uncertainty-aware multi-view co-training for semi-supervised medical image segmentation and domain adaptation
Yingda Xia, Dong Yang 0005, Zhiding Yu, Fengze Liu, Jinzheng Cai, Lequan Yu, Zhuotun Zhu, Daguang Xu, Alan L. Yuille, Holger Roth |
Medical Image Anal. | 9 |
| 2020 | Every Pixel Counts ++: Joint Learning of Geometry and Motion with 3D Holistic UnderstandingabstractLearning to estimate 3D geometry in a single frame and optical flow from consecutive frames by watching unlabeled videos via deep convolutional network has made significant progress recently. Current state-of-the-art (SoTA) methods treat the two tasks independently. One typical assumption of the existing depth estimation methods is that the scenes contain no independent moving objects. while object moving could be easily modeled using optical flow. In this paper, we propose to address the two tasks as a whole, i.e., to jointly understand per-pixel 3D geometry and motion. This eliminates the need of static scene assumption and enforces the inherent geometrical consistency during the learning process, yielding significantly improved results for both tasks. We call our method as “Every Pixel Counts++” or “EPC++”. Specifically, during training, given two consecutive frames from a video, we adopt three parallel networks to predict the camera motion (MotionNet), dense depth map (DepthNet), and per-pixel optical flow between two frames (OptFlowNet) respectively. The three types of information, are fed into a holistic 3D motion parser (HMP), and per-pixel 3D motion of both rigid background and moving objects are disentangled and recovered. Various loss terms are formulated to jointly supervise the three networks. An effective adaptive training strategy is proposed to achieve better performance and more efficient convergence. Comprehensive experiments were conducted on datasets with different scenes, including driving scenario (KITTI 2012 and KITTI 2015 datasets), mixed outdoor/indoor scenes (Make3D) and synthetic animation (MPI Sintel dataset). Performance on the five tasks of depth estimation, optical flow estimation, odometry, moving object segmentation and scene flow estimation shows that our approach outperforms other SoTA methods, demonstrating the effectiveness of each module of our proposed method. Code will be available at: https://github.com/chenxuluo/EPC. Chenxu Luo, Zhenheng Yang, Peng Wang 0001, Yang Wang 0046, Wei Xu 0017, Ramakant Nevatia, Alan L. Yuille |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2020 | PCL: Proposal Cluster Learning for Weakly Supervised Object DetectionabstractWeakly Supervised Object Detection (WSOD), using only image-level annotations to train object detectors, is of growing importance in object recognition. In this paper, we propose a novel deep network for WSOD. Unlike previous networks that transfer the object detection problem to an image classification problem using Multiple Instance Learning (MIL), our strategy generates proposal clusters to learn refined instance classifiers by an iterative process. The proposals in the same cluster are spatially adjacent and associated with the same object. This prevents the network from concentrating too much on parts of objects instead of whole objects. We first show that instances can be assigned object or background labels directly based on proposal clusters for instance classifier refinement, and then show that treating each cluster as a small new bag yields fewer ambiguities than the directly assigning label method. The iterative instance classifier refinement is implemented online using multiple streams in convolutional neural networks, where the first is an MIL network and the others are for instance classifier refinement supervised by the preceding one. Experiments are conducted on the PASCAL VOC, ImageNet detection, and MS-COCO benchmarks for WSOD. Results show that our method outperforms the previous state of the art significantly. Peng Tang 0005, Xinggang Wang, Song Bai 0001, Wei Shen 0002, Xiang Bai, Wenyu Liu 0001, Alan L. Yuille |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2020 | Unsupervised learning of optical flow with patch consistency and occlusion estimation
Zhe Ren, Junchi Yan, Xiaokang Yang 0001, Alan L. Yuille, Hongyuan Zha |
Pattern Recognit. | 4 |
| 2020 | STFlow: Self-Taught Optical Flow Estimation Using Pseudo LabelsabstractThe Deep learning of optical flow has been an active area for its empirical success. For the difficulty of obtaining accurate dense correspondence labels, unsupervised learning of optical flow has drawn more and more attention, while the accuracy is still far from satisfaction. By holding the philosophy that better estimation models can be trained with betterapproximated labels, which in turn can be obtained from better estimation models, we propose a self-taught learning framework to continually improve the accuracy using self-generated pseudo labels. The estimated optical flow is first filtered by bidirectional flow consistency validation and occlusion-aware dense labels are then generated by edge-aware interpolation from selected sparse matches. Moreover, by combining reconstruction loss with regression loss on the generated pseudo labels, the performance is further improved. The experimental results demonstrate that our models achieve state-of-the-art results among unsupervised methods on the public KITTI, MPI-Sintel and Flying Chairs datasets. Zhe Ren, Wenhan Luo, Junchi Yan, Wenlong Liao, Xiaokang Yang 0001, Alan L. Yuille, Hongyuan Zha |
IEEE Trans. Image Process. | 6 |
| 2020 | Recurrent Saliency Transformation Network for Tiny Target Segmentation in Abdominal CT ScansabstractWe aim at segmenting a wide variety of organs, including tiny targets (e.g., adrenal gland), and neoplasms (e.g., pancreatic cyst), from abdominal CT scans. This is a challenging task in two aspects. First, some organs (e.g., the pancreas), are highly variable in both anatomy and geometry, and thus very difficult to depict. Second, the neoplasms often vary a lot in its size, shape, as well as its location within the organ. Third, the targets (organs and neoplasms) can be considerably small compared to the human body, and so standard deep networks for segmentation are often less sensitive to these targets and thus predict less accurately especially around their boundaries. In this paper, we present an end-to-end framework named recurrent saliency transformation network (RSTN) for segmenting tiny and/or variable targets. The RSTN is a coarse-to-fine approach that uses prediction from the first (coarse) stage to shrink the input region for the second (fine) stage. A saliency transformation module is inserted between these two stages so that 1) the coarse-scaled segmentation mask can be transferred as spatial weights and applied to the fine stage and 2) the gradients can be back-propagated from the loss layer to the entire network so that the two stages are optimized in a joint manner. In the testing stage, we perform segmentation iteratively to improve accuracy. In this extended journal paper, we allow a gradual optimization to improve the stability of the RSTN, and introduce a hierarchical version named H-RSTN to segment tiny and variable neoplasms such as pancreatic cysts. Experiments are performed on several CT datasets including a public pancreas segmentation dataset, our own multi-organ dataset, and a cystic pancreas dataset. In all these cases, the RSTN outperforms the baseline (a stage-wise coarse-to-fine approach) significantly. Confirmed by the radiologists in our team, these promising segmentation results can help early diagnosis of pancreatic cancer. The code and pre-trained models of our project were made available at https://github.com/198808xc/OrganSegRSTN. Lingxi Xie, Qihang Yu, Yuyin Zhou, Yan Wang 0033, Elliot K. Fishman, Alan L. Yuille |
IEEE Trans. Medical Imaging | 6 |
| 2019 | Learning to Refine 3D Human Pose SequencesabstractWe present a basis approach to refine noisy 3D human pose sequences by jointly projecting them onto a non-linear pose manifold, which is represented by a number of basis dictionaries with each covering a small manifold region. We learn the dictionaries by jointly minimizing the distance between the original poses and their projections on the dictionaries, along with the temporal jittering of the projected poses. During testing, given a sequence of noisy poses which are probably off the manifold, we project them to the manifold using the same strategy as in training for refinement. We apply our approach to the monocular 3D pose estimation and the long term motion prediction tasks. The experimental results on the benchmark dataset shows the estimated 3D poses are notably improved in both tasks. In particular, the smoothness constraint helps generate more robust refinement results even when some poses in the original sequence have large errors. Jieru Mei, Xingyu Chen 0001, Chunyu Wang 0001, Alan L. Yuille, Xuguang Lan, Wenjun Zeng 0001 |
3DV | 4 |
| 2019 | V-NAS: Neural Architecture Search for Volumetric Medical Image SegmentationabstractDeep learning algorithms, in particular 2D and 3D fully convolutional neural networks (FCNs), have rapidly become the mainstream methodology for volumetric medical image segmentation. However, 2D convolutions cannot fully leverage the rich spatial information along the third axis, while 3D convolutions suffer from the demanding computation and high GPU memory consumption. In this paper, we propose to automatically search the network architecture tailoring to volumetric medical image segmentation problem. Concretely, we formulate the structure learning as differentiable neural architecture search, and let the network itself choose between 2D, 3D or Pseudo-3D (P3D) convolutions at each layer. We evaluate our method on 3 public datasets, i.e., the NIH Pancreas dataset, the Lung and Pancreas dataset from the Medical Segmentation Decathlon (MSD) Challenge. Our method, named V-NAS, consistently outperforms other state-of-the-arts on the segmentation tasks of both normal organ (NIH Pancreas) and abnormal organs (MSD Lung tumors and MSD Pancreas tumors), which shows the power of chosen architecture. Moreover, the searched architecture on one dataset can be well generalized to other datasets, which demonstrates the robustness and practical use of our proposed method. Zhuotun Zhu, Chenxi Liu 0001, Dong Yang 0005, Alan L. Yuille, Daguang Xu |
3DV | 4 |
| 2019 | Learning Basis Representation to Refine 3D Human Pose EstimationsabstractEstimating 3D human poses from 2D joint positions is an illposed problem, and is further complicated by the fact that the estimated 2D joints usually have errors to which most of the 3D pose estimators are sensitive. In this work, we present an approach to refine inaccurate 3D pose estimations. The core idea of the approach is to learn a number of bases to obtain tight approximations of the low-dimensional pose manifold where a 3D pose is represented by a convex combination of the bases. The representation requires that globally the refined poses are close to the pose manifold thus avoiding generating illegitimate poses. Second, the designed bases also have the property to guarantee that the distances among the body joints of a pose are within reasonable ranges. Experiments on benchmark datasets show that our approach obtains more legitimate poses over the baselines. In particular, the limb lengths are closer to the ground truth. Chunyu Wang 0001, Haibo Qiu, Alan L. Yuille, Wenjun Zeng 0001 |
AAAI | 3 |
| 2019 | Training Deep Neural Networks in Generations: A More Tolerant Teacher Educates Better StudentsabstractWe focus on the problem of training a deep neural network in generations. The flowchart is that, in order to optimize the target network (student), another network (teacher) with the same architecture is first trained, and used to provide part of supervision signals in the next stage. While this strategy leads to a higher accuracy, many aspects (e.g., why teacher-student optimization helps) still need further explorations.This paper studies this problem from a perspective of controlling the strictness in training the teacher network. Existing approaches mostly used a hard distribution (e.g., one-hot vectors) in training, leading to a strict teacher which itself has a high accuracy, but we argue that the teacher needs to be more tolerant, although this often implies a lower accuracy. The implementation is very easy, with merely an extra loss term added to the teacher network, facilitating a few secondary classes to emerge and complement to the primary class. Consequently, the teacher provides a milder supervision signal (a less peaked distribution), and makes it possible for the student to learn from inter-class similarity and potentially lower the risk of over-fitting. Experiments are performed on standard image classification tasks (CIFAR100 and ILSVRC2012). Although the teacher network behaves less powerful, the students show a persistent ability growth and eventually achieve higher classification accuracies than other competitors. Model ensemble and transfer feature extraction also verify the effectiveness of our approach. Lingxi Xie, Siyuan Qiao, Alan L. Yuille |
AAAI | 4 |
| 2019 | Seeing the Meaning: Vision Meets Semantics in Solving Pictorial Analogy Problems
Hongjing Lu, Qing Liu 0017, Nicholas Ichien, Alan L. Yuille, Keith J. Holyoak |
CogSci | 4 |
| 2019 | Robustness of Object Recognition under Extreme Occlusion in Humans and Computational Models
Hongru Zhu, Peng Tang 0005, Alan L. Yuille, Soojin Park |
CogSci | 3 |
| 2019 | NDDR-CNN: Layerwise Feature Fusing in Multi-Task CNNs by Neural Discriminative Dimensionality ReductionabstractIn this paper, we propose a novel Convolutional Neural Network (CNN) structure for general-purpose multi-task learning (MTL), which enables automatic feature fusing at every layer from different tasks. This is in contrast with the most widely used MTL CNN structures which empirically or heuristically share features on some specific layers (e.g., share all the features except the last convolutional layer). The proposed layerwise feature fusing scheme is formulated by combining existing CNN components in a novel way, with clear mathematical interpretability as discriminative dimensionality reduction, which is referred to as Neural Discriminative Dimensionality Reduction (NDDR). Specifically, we first concatenate features with the same spatial resolution from different tasks according to their channel dimension. Then, we show that the discriminative dimensionality reduction can be fulfilled by 1×1 Convolution, Batch Normalization, and Weight Decay in one CNN. The use of existing CNN components ensures the end-to-end training and the extensibility of the proposed NDDR layer to various state-of-the-art CNN architectures in a "plug-and-play" manner. The detailed ablation analysis shows that the proposed NDDR layer is easy to train and also robust to different hyperparameters. Experiments on different task sets with various base network architectures demonstrate the promising performance and desirable generalizability of our proposed method. The code of our paper is available at https://github.com/ethanygao/NDDR-CNN. Yuan Gao 0015, Jiayi Ma 0001, Ming-Bo Zhao, Wei Liu 0005, Alan L. Yuille |
CVPR | 5 |
| 2019 | Auto-DeepLab: Hierarchical Neural Architecture Search for Semantic Image SegmentationabstractRecently, Neural Architecture Search (NAS) has successfully identified neural network architectures that exceed human designed ones on large-scale image classification. In this paper, we study NAS for semantic image segmentation. Existing works often focus on searching the repeatable cell structure, while hand-designing the outer network structure that controls the spatial resolution changes. This choice simplifies the search space, but becomes increasingly problematic for dense image prediction which exhibits a lot more network level architectural variations. Therefore, we propose to search the network level structure in addition to the cell level structure, which forms a hierarchical architecture search space. We present a network level search space that includes many popular designs, and develop a formulation that allows efficient gradient-based architecture search (3 P100 GPU days on Cityscapes images). We demonstrate the effectiveness of the proposed method on the challenging Cityscapes, PASCAL VOC 2012, and ADE20K datasets. Auto-DeepLab, our architecture searched specifically for semantic image segmentation, attains state-of-the-art performance without any ImageNet pretraining. Chenxi Liu 0001, Liang-Chieh Chen, Florian Schroff, Hartwig Adam, Alan L. Yuille, Li Fei-Fei 0001 |
CVPR | 6 |
| 2019 | CLEVR-Ref+: Diagnosing Visual Reasoning With Referring ExpressionsabstractReferring object detection and referring image segmentation are important tasks that require joint understanding of visual information and natural language. Yet there has been evidence that current benchmark datasets suffer from bias, and current state-of-the-art models cannot be easily evaluated on their intermediate reasoning process. To address these issues and complement similar efforts in visual question answering, we build CLEVR-Ref+, a synthetic diagnostic dataset for referring expression comprehension. The precise locations and attributes of the objects are readily available, and the referring expressions are automatically associated with functional programs. The synthetic nature allows control over dataset bias (through sampling strategy), and the modular programs enable intermediate reasoning ground truth without human annotators. In addition to evaluating several state-of-the-art models on CLEVR-Ref+, we also propose IEP-Ref, a module network approach that significantly outperforms other models on our dataset. In particular, we present two interesting and important findings using IEP-Ref: (1) the module trained to transform feature maps into segmentation masks can be attached to any intermediate module to reveal the entire reasoning process step-by-step; (2) even if all training data has at least one object referred, IEP-Ref can correctly predict no-foreground when presented with false-premise referring expressions. To the best of our knowledge, this is the first direct and quantitative proof that neural modules behave in the way they are intended. We will release data and code for CLEVR-Ref+. Runtao Liu, Chenxi Liu 0001, Yutong Bai, Alan L. Yuille |
CVPR | 4 |
| 2019 | Elastic Boundary Projection for 3D Medical Image SegmentationabstractWe focus on an important yet challenging problem: using a 2D deep network to deal with 3D segmentation for medical image analysis. Existing approaches either applied multi-view planar (2D) networks or directly used volumetric (3D) networks for this purpose, but both of them are not ideal: 2D networks cannot capture 3D contexts effectively, and 3D networks are both memory-consuming and less stable arguably due to the lack of pre-trained models. In this paper, we bridge the gap between 2D and 3D using a novel approach named Elastic Boundary Projection (EBP). The key observation is that, although the object is a 3D volume, what we really need in segmentation is to find its boundary which is a 2D surface. Therefore, we place a number of pivot points in the 3D space, and for each pivot, we determine its distance to the object boundary along a dense set of directions. This creates an elastic shell around each pivot which is initialized as a perfect sphere. We train a 2D deep network to determine whether each ending point falls within the object, and gradually adjust the shell so that it gradually converges to the actual shape of the boundary and thus achieves the goal of segmentation. EBP allows boundary-based segmentation without cutting a 3D volume into slices or patches, which stands out from conventional 2D and 3D approaches. EBP achieves promising accuracy in abdominal organ segmentation. Our code will be released on https://github.com/twni2016/Elastic-Boundary-Projection . Tianwei Ni, Lingxi Xie, Huangjie Zheng, Elliot K. Fishman, Alan L. Yuille |
CVPR | 5 |
| 2019 | Neural Rejuvenation: Improving Deep Network Training by Enhancing Computational Resource UtilizationabstractIn this paper, we study the problem of improving computational resource utilization of neural networks. Deep neural networks are usually over-parameterized for their tasks in order to achieve good performances, thus are likely to have underutilized computational resources. This observation motivates a lot of research topics, e.g. network pruning, architecture search, etc. As models with higher computational costs (e.g. more parameters or more computations) usually have better performances, we study the problem of improving the resource utilization of neural networks so that their potentials can be further realized. To this end, we propose a novel optimization method named Neural Rejuvenation. As its name suggests, our method detects dead neurons and computes resource utilization in real time, rejuvenates dead neurons by resource reallocation and reinitialization, and trains them with new training schemes. By simply replacing standard optimizers with Neural Rejuvenation, we are able to improve the performances of neural networks by a very large margin while using similar training efforts and maintaining their original resource usages. The code is available here: https://github.com/joe-siyuan-qiao/NeuralRejuvenation-CVPR19 Siyuan Qiao, Zhe Lin 0001, Jianming Zhang 0001, Alan L. Yuille |
CVPR | 4 |
| 2019 | ELASTIC: Improving CNNs With Dynamic Scaling PoliciesabstractScale variation has been a challenge from traditional to modern approaches in computer vision. Most solutions to scale issues have a similar theme: a set of intuitive and manually designed policies that are generic and fixed (e.g. SIFT or feature pyramid). We argue that the scaling policy should be learned from data. In this paper, we introduce Elastic, a simple, efficient and yet very effective approach to learn a dynamic scale policy from data. We formulate the scaling policy as a non-linear function inside the network's structure that (a) is learned from data, (b) is instance specific, (c) does not add extra computation, and (d) can be applied on any network architecture. We applied Elastic to several state-of-the-art network architectures and showed consistent improvement without extra (sometimes even lower) computation on ImageNet classification, MSCOCO multi-label classification, and PASCAL VOC semantic segmentation. Our results show major improvement for images with scale challenges. Our code is available here: https://github.com/allenai/elastic. Aniruddha Kembhavi, Ali Farhadi, Alan L. Yuille, Mohammad Rastegari |
CVPR | 4 |
| 2019 | Iterative Reorganization With Weak Spatial Constraints: Solving Arbitrary Jigsaw Puzzles for Unsupervised Representation LearningabstractLearning visual features from unlabeled image data is an important yet challenging task, which is often achieved by training a model on some annotation-free information. We consider spatial contexts, for which we solve so-called jigsaw puzzles, i.e., each image is cut into grids and then disordered, and the goal is to recover the correct configuration. Existing approaches formulated it as a classification task by defining a fixed mapping from a small subset of configurations to a class set, but these approaches ignore the underlying relationship between different configurations and also limit their applications to more complex scenarios. This paper presents a novel approach which applies to jigsaw puzzles with an arbitrary grid size and dimensionality. We provide a fundamental and generalized principle, that weaker cues are easier to be learned in an unsupervised manner and also transfer better. In the context of puzzle recognition, we use an iterative manner which, instead of solving the puzzle all at once, adjusts the order of the patches in each step until convergence. In each step, we combine both unary and binary features of each patch into a cost function judging the correctness of the current configuration. Our approach, by taking similarity between puzzles into consideration, enjoys a more efficient way of learning visual knowledge. We verify the effectiveness of our approach from two aspects. First, it solves arbitrarily complex puzzles, including high-dimensional puzzles, that prior methods are difficult to handle. Second, it serves as a reliable way of network initialization, which leads to better transfer performance in visual recognition tasks including classification, detection and segmentation. Chen Wei 0005, Lingxi Xie, Xutong Ren, Yingda Xia, Chi Su, Jiaying Liu 0001, Qi Tian 0001, Alan L. Yuille |
CVPR | 8 |
| 2019 | Feature Denoising for Improving Adversarial RobustnessabstractAdversarial attacks to image classification systems present challenges to convolutional networks and opportunities for understanding them. This study suggests that adversarial perturbations on images lead to noise in the features constructed by these networks. Motivated by this observation, we develop new network architectures that increase adversarial robustness by performing feature denoising. Specifically, our networks contain blocks that denoise the features using non-local means or other filters; the entire networks are trained end-to-end. When combined with adversarial training, our feature denoising networks substantially improve the state-of-the-art in adversarial robustness in both white-box and black-box attack settings. On ImageNet, under 10-iteration PGD white-box attacks where prior art has 27.9% accuracy, our method achieves 55.7%; even under extreme 2000-iteration PGD white-box attacks, our method secures 42.6% accuracy. Our method was ranked first in Competition on Adversarial Attacks and Defenses (CAAD) 2018 --- it achieved 50.6% classification accuracy on a secret, ImageNet-like test dataset against 48 unknown attackers, surpassing the runner-up approach by ~10%. Code is available at https://github.com/facebookresearch/ImageNet-Adversarial-Training. Cihang Xie, Yuxin Wu 0004, Laurens van der Maaten, Alan L. Yuille, Kaiming He |
CVPR | 4 |
| 2019 | Improving Transferability of Adversarial Examples With Input DiversityabstractThough CNNs have achieved the state-of-the-art performance on various vision tasks, they are vulnerable to adversarial examples --- crafted by adding human-imperceptible perturbations to clean images. However, most of the existing adversarial attacks only achieve relatively low success rates under the challenging black-box setting, where the attackers have no knowledge of the model structure and parameters. To this end, we propose to improve the transferability of adversarial examples by creating diverse input patterns. Instead of only using the original images to generate adversarial examples, our method applies random transformations to the input images at each iteration. Extensive experiments on ImageNet show that the proposed attack method can generate adversarial examples that transfer much better to different networks than existing baselines. By evaluating our method against top defense solutions and official baselines from NIPS 2017 adversarial competition, the enhanced attack reaches an average success rate of 73.0%, which outperforms the top-1 attack submission in the NIPS competition by a large margin of 6.6%. We hope that our proposed attack strategy can serve as a strong benchmark baseline for evaluating the robustness of networks to adversaries and the effectiveness of different defense methods in the future. Code is available at https://github.com/cihangxie/DI-2-FGSM. Cihang Xie, Zhishuai Zhang, Yuyin Zhou, Song Bai 0001, Jianyu Wang 0001, Zhou Ren, Alan L. Yuille |
CVPR | 7 |
| 2019 | Snapshot Distillation: Teacher-Student Optimization in One GenerationabstractOptimizing a deep neural network is a fundamental task in computer vision, yet direct training methods often suffer from over-fitting. Teacher-student optimization aims at providing complementary cues from a model trained previously, but these approaches are often considerably slow due to the pipeline of training a few generations in sequence, i.e., time complexity is increased by several times. This paper presents snapshot distillation (SD), the first framework which enables teacher-student optimization in one generation. The idea of SD is very simple: instead of borrowing supervision signals from previous generations, we extract such information from earlier epochs in the same generation, meanwhile make sure that the difference between teacher and student is sufficiently large so as to prevent under-fitting. To achieve this goal, we implement SD in a cyclic learning rate policy, in which the last snapshot of each cycle is used as the teacher for all iterations in the next cycle, and the teacher signal is smoothed to provide richer information. In standard image classification benchmarks such as CIFAR100 and ILSVRC2012, SD achieves consistent accuracy gain without heavy computational overheads. We also verify that models pre-trained with SD transfers well to object detection and semantic segmentation in the PascalVOC dataset. Lingxi Xie, Chi Su, Alan L. Yuille |
CVPR | 4 |
| 2019 | Adversarial Attacks Beyond the Image SpaceabstractGenerating adversarial examples is an intriguing problem and an important way of understanding the working mechanism of deep neural networks. Most existing approaches generated perturbations in the image space, i.e., each pixel can be modified independently. However, in this paper we pay special attention to the subset of adversarial examples that correspond to meaningful changes in 3D physical properties (like rotation and translation, illumination condition, etc.). These adversaries arguably pose a more serious concern, as they demonstrate the possibility of causing neural network failure by easy perturbations of real-world 3D objects and scenes. In the contexts of object classification and visual question answering, we augment state-of-the-art deep neural networks that receive 2D input images with a rendering module (either differentiable or not) in front, so that a 3D scene (in the physical space) is rendered into a 2D image (in the image space), and then mapped to a prediction (in the output space). The adversarial perturbations can now go beyond the image space, and have clear meanings in the 3D physical world. Though image-space adversaries can be interpreted as per-pixel albedo change, we verify that they cannot be well explained along these physically meaningful dimensions, which often have a non-local effect. But it is still possible to successfully attack beyond the image space on the physical space, though this is more difficult than image-space attacks, reflected in lower success rates and heavier perturbations required. Xiaohui Zeng, Chenxi Liu 0001, Yu-Siang Wang, Weichao Qiu, Lingxi Xie, Yu-Wing Tai, Chi-Keung Tang, Alan L. Yuille |
CVPR | 8 |
| 2019 | CRAVES: Controlling Robotic Arm With a Vision-Based Economic SystemabstractTraining a robotic arm to accomplish real-world tasks has been attracting increasing attention in both academia and industry. This work discusses the role of computer vision algorithms in this field. We focus on low-cost arms on which no sensors are equipped and thus all decisions are made upon visual recognition, e.g., real-time 3D pose estimation. This requires annotating a lot of training data, which is not only time-consuming but also laborious. In this paper, we present an alternative solution, which uses a 3D model to create a large number of synthetic data, trains a vision model in this virtual domain, and applies it to real-world images after domain adaptation. To this end, we design a semi-supervised approach, which fully leverages the geometric constraints among keypoints. We apply an iterative algorithm for optimization. Without any annotations on real images, our algorithm generalizes well and produces satisfying results on 3D pose estimation, which is evaluated on two real-world datasets. We also construct a vision-based control system for task accomplishment, for which we train a reinforcement learning agent in a virtual environment and apply it to the real-world. Moreover, our approach, with merely a 3D model being required, has the potential to generalize to other types of multi-rigid-body dynamic systems. Yiming Zuo 0001, Weichao Qiu, Lingxi Xie, Fangwei Zhong, Yizhou Wang 0001, Alan L. Yuille |
CVPR | 6 |
| 2019 | Semantic-Aware Knowledge Preservation for Zero-Shot Sketch-Based Image RetrievalabstractSketch-based image retrieval (SBIR) is widely recognized as an important vision problem which implies a wide range of real-world applications. Recently, research interests arise in solving this problem under the more realistic and challenging setting of zero-shot learning. In this paper, we investigate this problem from the viewpoint of domain adaptation which we show is critical in improving feature embedding in the zero-shot scenario. Based on a framework which starts with a pre-trained model on ImageNet and fine-tunes it on the training set of SBIR benchmark, we advocate the importance of preserving previously acquired knowledge, e.g., the rich discriminative features learned from ImageNet, to improve the model's transfer ability. For this purpose, we design an approach named Semantic-Aware Knowledge prEservation (SAKE), which fine-tunes the pre-trained model in an economical way and leverages semantic information, e.g., inter-class relationship, to achieve the goal of knowledge preservation. Zero-shot experiments on two extended SBIR datasets, TU-Berlin and Sketchy, verify the superior performance of our approach. Extensive diagnostic experiments validate that knowledge preserved benefits SBIR in zero-shot settings, as a large fraction of the performance gain is from the more properly structured feature embedding for photo images. Qing Liu 0017, Lingxi Xie, Alan L. Yuille |
ICCV | 4 |
| 2019 | Semantic Part Detection via Matching: Learning to Generalize to Novel Viewpoints From Limited Training DataabstractDetecting semantic parts of an object is a challenging task, particularly because it is hard to annotate semantic parts and construct large datasets. In this paper, we present an approach which can learn from a small annotated dataset containing a limited range of viewpoints and generalize to detect semantic parts for a much larger range of viewpoints. The approach is based on our matching algorithm, which is used for finding accurate spatial correspondence between two images and transplanting semantic parts annotated on one image to the other. Images in the training set are matched to synthetic images rendered from a 3D CAD model, following which a clustering algorithm is used to automatically annotate semantic parts of the CAD model. During the testing period, this CAD model can synthesize annotated images under every viewpoint. These synthesized images are matched to images in the testing set to detect semantic parts in novel viewpoints. Our algorithm is simple, intuitive, and contains very few parameters. Experiments show our method outperforms standard deep learning approaches and, in particular, performs much better on novel viewpoints. For facilitating the future research, code is available: https://github.com/ytongbai/SemanticPartDetection. Yutong Bai, Qing Liu 0017, Lingxi Xie, Weichao Qiu, Alan L. Yuille |
ICCV | 6 |
| 2019 | An Alarm System for Segmentation Algorithm Based on Shape ModelabstractIt is usually hard for a learning system to predict correctly on rare events that never occur in the training data, and there is no exception for segmentation algorithms. Meanwhile, manual inspection of each case to locate the failures becomes infeasible due to the trend of large data scale and limited human resource. Therefore, we build an alarm system that will set off alerts when the segmentation result is possibly unsatisfactory, assuming no corresponding ground truth mask is provided. One plausible solution is to project the segmentation results into a low dimensional feature space; then learn classifiers/regressors to predict their qualities. Motivated by this, in this paper, we learn a feature space using the shape information which is a strong prior shared among different datasets and robust to the appearance variation of input data. The shape feature is captured using a Variational Auto-Encoder (VAE) network that trained with only the ground truth masks. During testing, the segmentation results with bad shapes shall not fit the shape prior well, resulting in large loss values. Thus, the VAE is able to evaluate the quality of segmentation result on unseen data, without using ground truth. Finally, we learn a regressor in the one-dimensional feature space to predict the qualities of segmentation results. Our alarm system is evaluated on several recent state-of-art segmentation algorithms for 3D medical segmentation tasks. Compared with other standard quality assessment methods, our system consistently provides more reliable prediction on the qualities of segmentation results. Fengze Liu, Yingda Xia, Dong Yang 0005, Alan L. Yuille, Daguang Xu |
ICCV | 4 |
| 2019 | Grouped Spatial-Temporal Aggregation for Efficient Action RecognitionabstractTemporal reasoning is an important aspect of video analysis. 3D CNN shows good performance by exploring spatial-temporal features jointly in an unconstrained way, but it also increases the computational cost a lot. Previous works try to reduce the complexity by decoupling the spatial and temporal filters. In this paper, we propose a novel decomposition method that decomposes the feature channels into spatial and temporal groups in parallel. This decomposition can make two groups focus on static and dynamic cues separately. We call this grouped spatial-temporal aggregation (GST). This decomposition is more parameter-efficient and enables us to quantitatively analyze the contributions of spatial and temporal features in different layers. We verify our model on several action recognition tasks that require temporal reasoning and show its effectiveness. Chenxu Luo, Alan L. Yuille |
ICCV | 2 |
| 2019 | Prior-Aware Neural Network for Partially-Supervised Multi-Organ SegmentationabstractAccurate multi-organ abdominal CT segmentation is essential to many clinical applications such as computer-aided intervention. As data annotation requires massive human labor from experienced radiologists, it is common that training data is usually partially-labeled. However, these background labels can be misleading in multi-organ segmentation since the ``background'' usually contains some other organs of interest. To address the background ambiguity in these partially-labeled datasets, we propose Prior-aware Neural Network (PaNN) via explicitly incorporating anatomical priors on abdominal organ sizes, guiding the training process with domain-specific knowledge. More specifically, PaNN assumes that the average organ size distributions in the abdomen should approximate their empirical distributions, a prior statistics obtained from the fully-labeled dataset. As our objective is difficult to be directly optimized using stochastic gradient descent, it is reformulated as a min-max form and optimized via the stochastic primal-dual gradient algorithm. PaNN achieves state-of-the-art performance on the MICCAI2015 challenge ``Multi-Atlas Labeling Beyond the Cranial Vault'', a competition on organ segmentation in the abdomen. We report an average Dice score of 84.97%, surpassing the prior art by a large margin of 3.27%. Code and models will be made publicly available. Yuyin Zhou, Song Bai 0001, Xinlei Chen, Elliot K. Fishman, Alan L. Yuille |
ICCV | 8 |
| 2019 | Hyper-Pairing Network for Multi-phase Pancreatic Ductal Adenocarcinoma Segmentation
Yuyin Zhou, Yingwei Li 0002, Zhishuai Zhang, Yan Wang 0033, Angtian Wang, Elliot K. Fishman, Alan L. Yuille, Seyoun Park |
MICCAI (2) | 7 |
| 2019 | Multi-scale Coarse-to-Fine Segmentation for Screening Pancreatic Ductal Adenocarcinoma
Zhuotun Zhu, Yingda Xia, Lingxi Xie, Elliot K. Fishman, Alan L. Yuille |
MICCAI (6) | 5 |
| 2019 | Semi-Supervised 3D Abdominal Multi-Organ Segmentation Via Deep Multi-Planar Co-TrainingabstractIn multi-organ segmentation of abdominal CT scans, most existing fully supervised deep learning algorithms require lots of voxel-wise annotations, which are usually difficult, expensive, and slow to obtain. In comparison, massive unlabeled 3D CT volumes are usually easily accessible. Current mainstream works to address semi-supervised biomedical image segmentation problem are mostly graph-based. By contrast, deep network based semi-supervised learning methods have not drawn much attention in this field. In this work, we propose Deep Multi-Planar Co-Training (DMPCT), whose contributions can be divided into two folds: 1) The deep model is learned in a co-training style which can mine consensus information from multiple planes like the sagittal, coronal, and axial planes; 2) Multi-planar fusion is applied to generate more reliable pseudo-labels, which alleviates the errors occurring in the pseudo-labels and thus can help to train better segmentation networks. Experiments are done on our newly collected large dataset with 100 unlabeled cases as well as 210 labeled cases where 16 anatomical structures are manually annotated by four radiologists and confirmed by a senior expert. The results suggest that DMPCT significantly outperforms the fully supervised method by more than 4% especially when only a small set of annotations is used. Yuyin Zhou, Yan Wang 0033, Peng Tang 0005, Song Bai 0001, Wei Shen 0002, Elliot K. Fishman, Alan L. Yuille |
WACV | 7 |
| 2019 | Estimation of 3D Category-Specific Object Structure: Symmetry, Manhattan and/or Multiple Images
Yuan Gao 0015, Alan L. Yuille |
Int. J. Comput. Vis. | 2 |
| 2019 | Abdominal multi-organ segmentation with organ-attention networks and statistical fusion
Yan Wang 0033, Yuyin Zhou, Wei Shen 0002, Seyoun Park, Elliot K. Fishman, Alan L. Yuille |
Medical Image Anal. | 6 |
| 2019 | Robust 3D Human Pose Estimation from Single Images or Video SequencesabstractWe propose a method for estimating 3D human poses from single images or video sequences. The task is challenging because: (a) many 3D poses can have similar 2D pose projections which makes the lifting ambiguous, and (b) current 2D joint detectors are not accurate which can cause big errors in 3D estimates. We represent 3D poses by a sparse combination of bases which encode structural pose priors to reduce the lifting ambiguity. This prior is strengthened by adding limb length constraints. We estimate the 3D pose by minimizing an$L_1$norm measurement error between the 2D pose and the 3D pose because it is less sensitive to inaccurate 2D poses. We modify our algorithm to output$K$3D pose candidates for an image, and for videos, we impose a temporal smoothness constraint to select the best sequence of 3D poses from the candidates. We demonstrate good results on 3D pose estimation from static images and improved performance by selecting the best 3D pose from the$K$proposals. Our results on video sequences also show improvements (over static images) of roughly 15%. Chunyu Wang 0001, Yizhou Wang 0001, Zhouchen Lin, Alan L. Yuille |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2018 | UnrealStereo: Controlling Hazardous Factors to Analyze Stereo VisionabstractA reliable stereo algorithm is critical for many robotics applications. But textureless and specular regions can easily cause failure by making feature matching difficult. Understanding whether an algorithm is robust to these hazardous regions is important. Although many stereo benchmarks have been developed to evaluate performance, it is hard to quantify the effect of hazardous regions in real images because the location and severity of these regions are unknown. In this paper, we develop a synthetic image generation tool enabling to control hazardous factors, such as making objects more specular or transparent, to produce hazardous regions at different degrees. The densely controlled sampling strategy in virtual worlds enables to effectively stress test stereo algorithms by varying the types and degrees of the hazard. We generate a large synthetic image dataset with automatically computed hazardous regions and analyze algorithms on these regions. The observations from synthetic images are further validated by annotating hazardous regions in real-world datasets Middlebury and KITTI (which gives a sparse sampling of the hazards). Our synthetic image generation tool is based on a game engine Unreal Engine 4 and will be open-source along with the virtual scenes in our experiments. Many publicly available realistic game contents can be used by our tool to provide an enormous resource for development and evaluation of algorithms. Yi Zhang 0099, Weichao Qiu, Qi Chen 0014, Xiaolin Hu 0001, Alan L. Yuille |
3DV | 5 |
| 2018 | A 3D Coarse-to-Fine Framework for Volumetric Medical Image SegmentationabstractIn this paper, we adopt 3D Convolutional Neural Networks to segment volumetric medical images. Although deep neural networks have been proven to be very effective on many 2D vision tasks, it is still challenging to apply them to 3D tasks due to the limited amount of annotated 3D data and limited computational resources. We propose a novel 3D-based coarse-to-fine framework to effectively and efficiently tackle these challenges. The proposed 3D-based framework outperforms the 2D counterpart to a large margin since it can leverage the rich spatial information along all three axes. We conduct experiments on two datasets which include healthy and pathological pancreases respectively, and achieve the current state-of-the-art in terms of Dice-Sørensen Coefficient (DSC). On the NIH pancreas segmentation dataset, we outperform the previous best by an average of over 2%, and the worst case is improved by 7% to reach almost 70%, which indicates the reliability of our framework in clinical applications. Zhuotun Zhu, Yingda Xia, Wei Shen 0002, Elliot K. Fishman, Alan L. Yuille |
3DV | 5 |
| 2018 | SampleAhead: Online Classifier-Sampler Communication for Learning from Synthesized Data
Qi Chen 0014, Weichao Qiu, Yi Zhang 0099, Lingxi Xie, Alan L. Yuille |
BMVC | 5 |
| 2018 | OriNet: A Fully Convolutional Network for 3D Human Pose Estimation
Chenxu Luo, Xiao Chu, Alan L. Yuille |
BMVC | 3 |
| 2018 | Deep Regression Forests for Age EstimationabstractAge estimation from facial images is typically cast as a nonlinear regression problem. The main challenge of this problem is the facial feature space w.r.t. ages is inhomogeneous, due to the large variation in facial appearance across different persons of the same age and the non-stationary property of aging patterns. In this paper, we propose Deep Regression Forests (DRFs), an end-to-end model, for age estimation. DRFs connect the split nodes to a fully connected layer of a convolutional neural network (CNN) and deal with inhomogeneous data by jointly learning input-dependant data partitions at the split nodes and data abstractions at the leaf nodes. This joint learning follows an alternating strategy: First, by fixing the leaf nodes, the split nodes as well as the CNN parameters are optimized by Back-propagation; Then, by fixing the split nodes, the leaf nodes are optimized by iterating a step-size free update rule derived from Variational Bounding. We verify the proposed DRFs on three standard age estimation benchmarks and achieve state-of-the-art results on all of them. Wei Shen 0002, Yilu Guo, Yan Wang 0033, Kai Zhao 0012, Bo Wang 0044, Alan L. Yuille |
CVPR | 6 |
| 2018 | Few-Shot Image Recognition by Predicting Parameters From ActivationsabstractIn this paper, we are interested in the few-shot learning problem. In particular, we focus on a challenging scenario where the number of categories is large and the number of examples per novel category is very limited, e.g. 1, 2, or 3. Motivated by the close relationship between the parameters and the activations in a neural network associated with the same category, we propose a novel method that can adapt a pre-trained neural network to novel categories by directly predicting the parameters from the activations. Zero training is required in adaptation to novel categories, and fast inference is realized by a single forward pass. We evaluate our method by doing few-shot image recognition on the ImageNet dataset, which achieves the state-of-the-art classification accuracy on novel categories by a significant margin while keeping comparable performance on the large-scale categories. We also test our method on the MiniImageNet dataset and it strongly outperforms the previous state-of-the-art methods. Siyuan Qiao, Chenxi Liu 0001, Wei Shen 0002, Alan L. Yuille |
CVPR | 4 |
| 2018 | Recurrent Saliency Transformation Network: Incorporating Multi-Stage Visual Cues for Small Organ SegmentationabstractWe aim at segmenting small organs (e.g., the pancreas) from abdominal CT scans. As the target often occupies a relatively small region in the input image, deep neural networks can be easily confused by the complex and variable background. To alleviate this, researchers proposed a coarse-to-fine approach [46], which used prediction from the first (coarse) stage to indicate a smaller input region for the second (fine) stage. Despite its effectiveness, this algorithm dealt with two stages individually, which lacked optimizing a global energy function, and limited its ability to incorporate multi-stage visual cues. Missing contextual information led to unsatisfying convergence in iterations, and that the fine stage sometimes produced even lower segmentation accuracy than the coarse stage. This paper presents a Recurrent Saliency Transformation Network. The key innovation is a saliency transformation module, which repeatedly converts the segmentation probability map from the previous iteration as spatial weights and applies these weights to the current iteration. This brings us two-fold benefits. In training, it allows joint optimization over the deep networks dealing with different input scales. In testing, it propagates multi-stage visual information throughout iterations to improve segmentation accuracy. Experiments in the NIH pancreas segmentation dataset demonstrate the state-of-the-art accuracy, which outperforms the previous best by an average of over 2%. Much higher accuracies are also reported on several small organs in a larger dataset collected by ourselves. In addition, our approach enjoys better convergence properties, making it more efficient and reliable in practice. Qihang Yu, Lingxi Xie, Yan Wang 0033, Yuyin Zhou, Elliot K. Fishman, Alan L. Yuille |
CVPR | 6 |
| 2018 | Single-Shot Object Detection With Enriched SemanticsabstractWe propose a novel single shot object detection network named Detection with Enriched Semantics (DES). Our motivation is to enrich the semantics of object detection features within a typical deep detector, by a semantic segmentation branch and a global activation module. The segmentation branch is supervised by weak segmentation ground-truth, i.e., no extra annotation is required. In conjunction with that, we employ a global activation module which learns relationship between channels and object classes in a self-supervised manner. Comprehensive experimental results on both PASCAL VOC and MS COCO detection datasets demonstrate the effectiveness of the proposed method. In particular, with a VGG16 based DES, we achieve an mAP of 81.7 on VOC2007 test and an mAP of 32.8 on COCO test-dev with an inference speed of 31.5 milliseconds per image on a Titan Xp GPU. With a lower resolution version, we achieve an mAP of 79.7 on VOC2007 with an inference speed of 13.0 milliseconds per image. Zhishuai Zhang, Siyuan Qiao, Cihang Xie, Wei Shen 0002, Bo Wang 0044, Alan L. Yuille |
CVPR | 6 |
| 2018 | DeepVoting: A Robust and Explainable Deep Network for Semantic Part Detection Under Partial OcclusionabstractIn this paper, we study the task of detecting semantic parts of an object, e.g., a wheel of a car, under partial occlusion. We propose that all models should be trained without seeing occlusions while being able to transfer the learned knowledge to deal with occlusions. This setting alleviates the difficulty in collecting an exponentially large dataset to cover occlusion patterns and is more essential. In this scenario, the proposal-based deep networks, like RCNN-series, often produce unsatisfactory results, because both the proposal extraction and classification stages may be confused by the irrelevant occluders. To address this, [25] proposed a voting mechanism that combines multiple local visual cues to detect semantic parts. The semantic parts can still be detected even though some visual cues are missing due to occlusions. However, this method is manually-designed, thus is hard to be optimized in an end-to-end manner. In this paper, we present DeepVoting, which incorporates the robustness shown by [25] into a deep network, so that the whole pipeline can be jointly optimized. Specifically, it adds two layers after the intermediate features of a deep network, e.g., the pool-4 layer of VGGNet. The first layer extracts the evidence of local visual cues, and the second layer performs a voting mechanism by utilizing the spatial relationship between visual cues and semantic parts. We also propose an improved version DeepVoting+ by learning visual cues from context outside objects. In experiments, DeepVoting achieves significantly better performance than several baseline methods, including Faster-RCNN, for semantic part detection under occlusion. In addition, DeepVoting enjoys explainability as the detection results can be diagnosed via looking up the voting cues. Zhishuai Zhang, Cihang Xie, Jianyu Wang 0001, Lingxi Xie, Alan L. Yuille |
CVPR | 5 |
| 2018 | Progressive Neural Architecture Search
Chenxi Liu 0001, Barret Zoph, Maxim Neumann, Jonathon Shlens, Li-Jia Li 0001, Li Fei-Fei 0001, Alan L. Yuille, Jonathan Huang, Kevin Murphy 0002 |
ECCV (1) | 8 |
| 2018 | Deep Co-Training for Semi-Supervised Image Recognition
Siyuan Qiao, Wei Shen 0002, Zhishuai Zhang, Bo Wang 0044, Alan L. Yuille |
ECCV (15) | 5 |
| 2018 | Weakly Supervised Region Proposal Network and Object Detection
Peng Tang 0005, Xinggang Wang, Angtian Wang, Yongluan Yan, Wenyu Liu 0001, Junzhou Huang, Alan L. Yuille |
ECCV (11) | 7 |
| 2018 | Multi-scale Spatially-Asymmetric Recalibration for Image Classification
Yan Wang 0033, Lingxi Xie, Siyuan Qiao, Ya Zhang 0002, Wenjun Zhang 0001, Alan L. Yuille |
ECCV (13) | 6 |
| 2018 | PreCo: A Large-scale Dataset in Preschool Vocabulary for Coreference ResolutionabstractWe introduce PreCo, a large-scale English dataset for coreference resolution.The dataset is designed to embody the core challenges in coreference, such as entity representation, by alleviating the challenge of low overlap between training and test sets and enabling separated analysis of mention detection and mention clustering.To strengthen the training-test overlap, we collect a large corpus of 38K documents and 12.5M words which are mostly from the vocabulary of Englishspeaking preschoolers.Experiments show that with higher training-test overlap, error analysis on PreCo is more efficient than the one on OntoNotes, a popular existing dataset.Furthermore, we annotate singleton mentions making it possible for the first time to quantify the influence that a mention detector makes on coreference resolution performance. Zhenhua Fan, Alan L. Yuille, Shu Rong |
EMNLP | 4 |
| 2018 | Mitigating Adversarial Effects Through Randomization
Cihang Xie, Jianyu Wang 0001, Zhishuai Zhang, Zhou Ren, Alan L. Yuille |
ICLR (Poster) | 5 |
| 2018 | Gradually Updated Neural Networks for Large-Scale Image RecognitionabstractDepth is one of the keys that make neural networks succeed in the task of large-scale image recognition. The state-of-the-art network architectures usually increase the depths by cascading convolutional layers or building blocks. In this paper, we present an alternative method to increase the depth. Our method is by introducing computation orderings to the channels within convolutional layers or blocks, based on which we gradually compute the outputs in a channel-wise manner. The added orderings not only increase the depths and the learning capacities of the networks without any additional computation costs, but also eliminate the overlap singularities so that the networks are able to converge faster and perform better. Experiments show that the networks based on our method achieve the state-of-the-art performances on CIFAR and ImageNet datasets. Siyuan Qiao, Zhishuai Zhang, Wei Shen 0002, Bo Wang 0044, Alan L. Yuille |
ICML | 5 |
| 2018 | Stochastic Functional Verification of DNN Design through Progressive Virtual Dataset GenerationabstractDeep Neural Networks have emerged as state-of-the-art solutions for complex intelligence problems. DNNs derive their predictive power by learning from millions of training examples in either a supervised or semi-supervised fashion. As such, a critical aspect of the DNN system design procedure is the collection of large annotated training datasets that exhibit high coverage of the problem space. While data synthesis and annotation techniques have been proposed to mitigate the burden of acquiring large datasets, these methods do not quantify the usefulness of each generated dataset and its subsequent impact on training effort. In this work we establish parallels between the autonomous design of DNNs for machine vision applications and the task of functionally verifying a hardware design. Similar to automatic test vector generation, we propose a technique that progressively generates training datasets using virtual synthetic models. Furthermore, we propose an automated DNN design framework that jointly tries to stochastically maximize training coverage while minimizing the number of training and validation cycles utilizing insights from functional verification. Jinhang Choi, Kevin M. Irick, Justin Hardin, Weichao Qiu, Alan L. Yuille, Jack Sampson, Narayanan Vijaykrishnan |
ISCAS | 5 |
| 2018 | Training Multi-organ Segmentation Networks with Sample Selection by Relaxed Upper Confident Bound
Yan Wang 0033, Yuyin Zhou, Peng Tang 0005, Wei Shen 0002, Elliot K. Fishman, Alan L. Yuille |
MICCAI (4) | 6 |
| 2018 | Bridging the Gap Between 2D and 3D Organ Segmentation with Volumetric Fusion Net
Yingda Xia, Lingxi Xie, Fengze Liu, Zhuotun Zhu, Elliot K. Fishman, Alan L. Yuille |
MICCAI (4) | 6 |
| 2018 | Scene Graph Parsing as Dependency ParsingabstractYu-Siang Wang, Chenxi Liu, Xiaohui Zeng, Alan Yuille. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Yu-Siang Wang, Chenxi Liu 0001, Xiaohui Zeng, Alan L. Yuille |
NAACL-HLT | 4 |
| 2018 | FIPIP: A novel fine-grained parallel partition based intra-frame prediction on heterogeneous many-core systems
Wenbin Jiang 0001, Laurence T. Yang, Xiaobai Liu, Hai Jin 0001, Alan L. Yuille, Ye Chi |
Future Gener. Comput. Syst. | 6 |
| 2018 | DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFsabstractIn this work we address the task of semantic image segmentation with Deep Learning and make three main contributions that are experimentally shown to have substantial practical merit. First, we highlight convolution with upsampled filters, or 'atrous convolution', as a powerful tool in dense prediction tasks. Atrous convolution allows us to explicitly control the resolution at which feature responses are computed within Deep Convolutional Neural Networks. It also allows us to effectively enlarge the field of view of filters to incorporate larger context without increasing the number of parameters or the amount of computation. Second, we propose atrous spatial pyramid pooling (ASPP) to robustly segment objects at multiple scales. ASPP probes an incoming convolutional feature layer with filters at multiple sampling rates and effective fields-of-views, thus capturing objects as well as image context at multiple scales. Third, we improve the localization of object boundaries by combining methods from DCNNs and probabilistic graphical models. The commonly deployed combination of max-pooling and downsampling in DCNNs achieves invariance but has a toll on localization accuracy. We overcome this by combining the responses at the final DCNN layer with a fully connected Conditional Random Field (CRF), which is shown both qualitatively and quantitatively to improve localization performance. Our proposed "DeepLab" system sets the new state-of-art at the PASCAL VOC-2012 semantic image segmentation task, reaching 79.7 percent mIOU in the test set, and advances the results on three other datasets: PASCAL-Context, PASCAL-Person-Part, and Cityscapes. All of our code is made publicly available online. Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy 0002, Alan L. Yuille |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2018 | A Novel Linelet-Based Representation for Line Segment DetectionabstractThis paper proposes a method for line segment detection in digital images. We propose a novel linelet-based representation to model intrinsic properties of line segments in rasterized image space. Based on this, line segment detection, validation, and aggregation frameworks are constructed. For a numerical evaluation on real images, we propose a new benchmark dataset of real images with annotated lines called YorkUrban-LineSegment. The results show that the proposed method outperforms state-of-the-art methods numerically and visually. To our best knowledge, this is the first report of numerical evaluation of line segment detection on real images. Nam-Gyu Cho, Alan L. Yuille, Seong-Whan Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Ground-Truth Data Set and Baseline Evaluations for Base-Detail Separation Algorithms at the Part LevelabstractBase-detail separation is a fundamental image processing problem, which models the image by a smooth base layer for the coarse structure and a detail layer for the texturelike structures. Base-detail separation is hierarchical and can be performed from the fine level to the coarse level. The separation at coarse level, in particular at the part level, is important for many applications, but currently lacks ground-truth data sets that are needed for comparing algorithms quantitatively. Thus, we propose a procedure to construct such data sets and provide two examples: Pascal Part UCLA and Fashionista, containing 1000 and 250 images, respectively. Our assumption is that the base is piecewise smooth, and we label the appearance of each piece by a polynomial model. The pieces are objects and parts of objects obtained from human annotations. Finally, we propose a way to evaluate different separation methods with our data sets and compared the performances of seven state-of-the-art algorithms. Xuan Dong 0001, Boyan Bonev 0001, Weixin Li 0001, Weichao Qiu, Xianjie Chen, Alan L. Yuille |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2018 | Guest Editorial Introduction to the Special Issue on Large Scale and Nonlinear Similarity Learning for Intelligent Video AnalysisabstractLearning similarity and distance measures has become increasingly important for the analysis, matching, retrieval, recognition, and categorization of video and multimedia data. With the ubiquitous use of digital imaging devices, mobile terminals and social networks, there are massive volumes of heterogeneous and homogeneous video and multimedia data from multiple sources, views, and domains, e.g., news media websites, microblog, mobile phone, social networking, etc. Similarity and distance-based constraints can also be extended and incorporated to boost classification and relationship learning. Moreover, the spatio-temporal coherence among video data can also be utilized for self-supervised learning of similarity and distance metrics. This trend has brought several challenging issues for developing similarity and metric learning methods for large scale and weakly annotated data, where outliers and incorrectly annotated data are inevitable. Recently, scalability has been investigated to cope with lightweight and large scale metric learning, while nonlinear similarity models have shown their great potentials in learning invariant representation and nonlinear measures of video and multimedia data. Wangmeng Zuo, Liang Lin 0004, Alan L. Yuille, Horst Bischof, Lei Zhang 0006, Fatih Porikli |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | Attention Correctness in Neural Image CaptioningabstractAttention mechanisms have recently been introduced in deep learning for various tasks in natural language processing and computer vision. But despite their popularity, the ``correctness'' of the implicitly-learned attention maps has only been assessed qualitatively by visualization of several examples. In this paper we focus on evaluating and improving the correctness of attention in neural image captioning models. Specifically, we propose a quantitative evaluation metric for the consistency between the generated attention maps and human annotations, using recently released datasets with alignment between regions in images and entities in captions. We then propose novel models with different levels of explicit supervision for learning attention maps during training. The supervision can be strong when alignment between regions and caption entities are available, or weak when only object segments and categories are provided. We show on the popular Flickr30k and COCO datasets that introducing supervision of attention maps during training solidly improves both attention correctness and caption quality, showing the promise of making machine perception more human-like. Chenxi Liu 0001, Junhua Mao, Fei Sha, Alan L. Yuille |
AAAI | 4 |
| 2017 | Multiple Instance Visual-Semantic Embedding
Zhou Ren, Hailin Jin, Zhe Lin 0001, Alan L. Yuille |
BMVC | 5 |
| 2017 | Detecting Semantic Parts on Partially Occluded Objects
Jianyu Wang 0001, Zhishuai Zhang, Cihang Xie, Jun Zhu 0001, Lingxi Xie, Alan L. Yuille |
BMVC | 6 |
| 2017 | Exploiting Protrusion Cues for Fast and Effective Shape Modeling via Ellipses
Alex Wong 0001, Alan L. Yuille, Brian Taylor 0001 |
BMVC | 2 |
| 2017 | Multi-context Attention for Human Pose EstimationabstractIn this paper, we propose to incorporate convolutional neural networks with a multi-context attention mechanism into an end-to-end framework for human pose estimation. We adopt stacked hourglass networks to generate attention maps from features at multiple resolutions with various semantics. The Conditional Random Field (CRF) is utilized to model the correlations among neighboring regions in the attention map. We further combine the holistic attention model, which focuses on the global consistency of the full human body, and the body part attention model, which focuses on detailed descriptions for different body parts. Hence our model has the ability to focus on different granularity from local salient regions to global semantic consistent spaces. Additionally, we design novel Hourglass Residual Units (HRUs) to increase the receptive field of the network. These units are extensions of residual units with a side branch incorporating filters with larger receptive field, hence features with various scales are learned and combined within the HRUs. The effectiveness of the proposed multi-context attention mechanism and the hourglass residual units is evaluated on two widely used human pose estimation benchmarks. Our approach outperforms all existing methods on both benchmarks over all the body parts. Code has been made publicly available. Xiao Chu, Wei Yang 0019, Wanli Ouyang, Alan L. Yuille, Xiaogang Wang 0001 |
CVPR | 5 |
| 2017 | Exploiting Symmetry and/or Manhattan Properties for 3D Object Structure Estimation from Single and Multiple ImagesabstractMany man-made objects have intrinsic symmetries and Manhattan structure. By assuming an orthographic projection model, this paper addresses the estimation of 3D structures and camera projection using symmetry and/or Manhattan structure cues, which occur when the input is single-or multiple-image from the same category, e.g., multiple different cars. Specifically, analysis on the single image case implies that Manhattan alone is sufficient to recover the camera projection, and then the 3D structure can be reconstructed uniquely exploiting symmetry. However, Manhattan structure can be difficult to observe from a single image due to occlusion. To this end, we extend to the multiple-image case which can also exploit symmetry but does not require Manhattan axes. We propose a novel rigid structure from motion method, exploiting symmetry and using multiple images from the same category as input. Experimental results on the Pascal3D+ dataset show that our method significantly outperforms baseline methods. Yuan Gao 0015, Alan L. Yuille |
CVPR | 2 |
| 2017 | Joint Multi-person Pose Estimation and Semantic Part SegmentationabstractHuman pose estimation and semantic part segmentation are two complementary tasks in computer vision. In this paper, we propose to solve the two tasks jointly for natural multi-person images, in which the estimated pose provides object-level shape prior to regularize part segments while the part-level segments constrain the variation of pose locations. Specifically, we first train two fully convolutional neural networks (FCNs), namely Pose FCN and Part FCN, to provide initial estimation of pose joint potential and semantic part potential. Then, to refine pose joint location, the two types of potentials are fused with a fully-connected conditional random field (FCRF), where a novel segment-joint smoothness term is used to encourage semantic and spatial consistency between parts and joints. To refine part segments, the refined pose and the original part potential are integrated through a Part FCN, where the skeleton feature from pose serves as additional regularization cues for part segments. Finally, to reduce the complexity of the FCRF, we induce human detection boxes and infer the graph inside each box, making the inference forty times faster. Since theres no dataset that contains both part segments and pose labels, we extend the PASCAL VOC part dataset [6] with human pose joints and perform extensive experiments to compare our method against several most recent strategies. We show that our algorithm surpasses competing methods by 10.6% in pose estimation with much faster speed and by 1.5% in semantic part segmentation. Fangting Xia, Peng Wang 0001, Xianjie Chen, Alan L. Yuille |
CVPR | 4 |
| 2017 | Recurrent Multimodal Interaction for Referring Image SegmentationabstractIn this paper we are interested in the problem of image segmentation given natural language descriptions, i.e. referring expressions. Existing works tackle this problem by first modeling images and sentences independently and then segment images by combining these two types of representations. We argue that learning word-to-image interaction is more native in the sense of jointly modeling two modalities for the image segmentation task, and we propose convolutional multimodal LSTM to encode the sequential interactions between individual words, visual information, and spatial information. We show that our proposed model outperforms the baseline model on benchmark datasets. In addition, we analyze the intermediate output of the proposed multimodal LSTM approach and empirically explain how this approach enforces a more effective word-to-image interaction. Chenxi Liu 0001, Zhe Lin 0001, Xiaohui Shen, Jimei Yang, Xin Lu 0006, Alan L. Yuille |
ICCV | 6 |
| 2017 | ScaleNet: Guiding Object Proposal Generation in Supermarkets and BeyondabstractMotivated by product detection in supermarkets, this paper studies the problem of object proposal generation in supermarket images and other natural images. We argue that estimation of object scales in images is helpful for generating object proposals, especially for supermarket images where object scales are usually within a small range. Therefore, we propose to estimate object scales of images before generating object proposals. The proposed method for predicting object scales is called ScaleNet. To validate the effectiveness of ScaleNet, we build three supermarket datasets, two of which are real-world datasets used for testing and the other one is a synthetic dataset used for training. In short, we extend the previous state-of-the-art object proposal methods by adding a scale prediction phase. The resulted method outperforms the previous state-of-the-art on the supermarket datasets by a large margin. We also show that the approach works for object proposal on other natural images and it outperforms the previous state-of-the-art object proposal methods on the MS COCO dataset. The supermarket datasets, the virtual supermarkets, and the tools for creating more synthetic datasets will be made public. Siyuan Qiao, Wei Shen 0002, Weichao Qiu, Chenxi Liu 0001, Alan L. Yuille |
ICCV | 5 |
| 2017 | Multi-stage Multi-recursive-input Fully Convolutional Networks for Neuronal Boundary DetectionabstractIn the field of connectomics, neuroscientists seek to identify cortical connectivity comprehensively. Neuronal boundary detection from the Electron Microscopy (EM) images is often done to assist the automatic reconstruction of neuronal circuit. But the segmentation of EM images is a challenging problem, as it requires the detector to be able to detect both filament-like thin and blob-like thick membrane, while suppressing the ambiguous intracellular structure. In this paper, we propose multi-stage multi-recursiveinput fully convolutional networks to address this problem. The multiple recursive inputs for one stage, i.e., the multiple side outputs with different receptive field sizes learned from the lower stage, provide multi-scale contextual boundary information for the consecutive learning. This design is biologically-plausible, as it likes a human visual system to compare different possible segmentation solutions to address the ambiguous boundary issue. Our multi-stage networks are trained end-to-end. It achieves promising results on two public available EM segmentation datasets, the mouse piriform cortex dataset and the ISBI 2012 EM dataset. Wei Shen 0002, Bin Wang 0027, Yuan Jiang 0002, Yan Wang 0033, Alan L. Yuille |
ICCV | 5 |
| 2017 | SORT: Second-Order Response Transform for Visual RecognitionabstractIn this paper, we reveal the importance and benefits of introducing second-order operations into deep neural networks. We propose a novel approach named Second-Order Response Transform (SORT), which appends element-wise product transform to the linear sum of a two-branch network module. A direct advantage of SORT is to facilitate cross-branch response propagation, so that each branch can update its weights based on the current status of the other branch. Moreover, SORT augments the family of transform operations and increases the nonlinearity of the network, making it possible to learn flexible functions to fit the complicated distribution of feature space. SORT can be applied to a wide range of network architectures, including a branched variant of a chain-styled network and a residual network, with very light-weighted modifications. We observe consistent accuracy gain on both small (CIFAR10, CIFAR100 and SVHN) and big (ILSVRC2012) datasets. In addition, SORT is very efficient, as the extra computation overhead is less than 5%. Yan Wang 0033, Lingxi Xie, Chenxi Liu 0001, Siyuan Qiao, Ya Zhang 0002, Wenjun Zhang 0001, Qi Tian 0001, Alan L. Yuille |
ICCV | 8 |
| 2017 | Adversarial Examples for Semantic Segmentation and Object DetectionabstractIt has been well demonstrated that adversarial examples, i.e., natural images with visually imperceptible perturbations added, cause deep networks to fail on image classification. In this paper, we extend adversarial examples to semantic segmentation and object detection which are much more difficult. Our observation is that both segmentation and detection are based on classifying multiple targets on an image (e.g., the target is a pixel or a receptive field in segmentation, and an object proposal in detection). This inspires us to optimize a loss function over a set of targets for generating adversarial perturbations. Based on this, we propose a novel algorithm named Dense Adversary Generation (DAG), which applies to the state-of-the-art networks for segmentation and detection. We find that the adversarial perturbations can be transferred across networks with different training data, based on different architectures, and even for different recognition tasks. In particular, the transfer ability across networks with the same architecture is more significant than in other cases. Besides, we show that summing up heterogeneous perturbations often leads to better transfer performance, which provides an effective method of black-box adversarial attack. Cihang Xie, Jianyu Wang 0001, Zhishuai Zhang, Yuyin Zhou, Lingxi Xie, Alan L. Yuille |
ICCV | 6 |
| 2017 | Genetic CNNabstractThe deep convolutional neural network (CNN) is the state-of-the-art solution for large-scale visual recognition. Following some basic principles such as increasing network depth and constructing highway connections, researchers have manually designed a lot of fixed network architectures and verified their effectiveness. In this paper, we discuss the possibility of learning deep network structures automatically. Note that the number of possible network structures increases exponentially with the number of layers in the network, which motivates us to adopt the genetic algorithm to efficiently explore this large search space. The core idea is to propose an encoding method to represent each network structure in a fixed-length binary string. The genetic algorithm is initialized by generating a set of randomized individuals. In each generation, we define standard genetic operations, e.g., selection, mutation and crossover, to generate competitive individuals and eliminate weak ones. The competitiveness of each individual is defined as its recognition accuracy, which is obtained via a standalone training process on a reference dataset. We run the genetic process on CIFAR10, a small-scale dataset, demonstrating its ability to find high-quality structures which are little studied before. The learned powerful structures are also transferrable to the ILSVRC2012 dataset for large-scale visual recognition. Lingxi Xie, Alan L. Yuille |
ICCV | 2 |
| 2017 | Regularizing face verification nets for pain intensity regressionabstractLimited labeled data are available for the research of estimating facial expression intensities. For instance, the ability to train deep networks for automated pain assessment is limited by small datasets with labels of patient-reported pain intensities. Fortunately, fine-tuning from a data-extensive pre-trained domain, such as face verification, can alleviate this problem. In this paper, we propose a network that fine-tunes a state-of-the-art face verification network using a regularized regression loss and additional data with expression labels. In this way, the expression intensity regression task can benefit from the rich feature representations trained on a huge amount of data for face verification. The proposed regularized deep regressor is applied to estimate the pain expression intensity and verified on the widely-used UNBC-McMaster Shoulder-Pain dataset, achieving the state-of-the-art performance. A weighted evaluation metric is also proposed to address the imbalance issue of different pain intensities. Feng Wang 0015, Xiang Xiang 0001, Trac D. Tran, Austin Reiter, Gregory D. Hager, Harry Quon, Jian Cheng 0003, Alan L. Yuille |
ICIP | 9 |
| 2017 | Transfer of View-manifold Learning to Similarity Perception of Novel Objects
Alan L. Yuille, Tai Sing Lee |
ICLR (Poster) | 5 |
| 2017 | MAT: A Multimodal Attentive Translator for Image CaptioningabstractIn this work we formulate the problem of image captioning as a multimodal translation task. Analogous to machine translation, we present a sequence-to-sequence recurrent neural networks (RNN) model for image caption generation. Different from most existing work where the whole image is represented by convolutional neural network (CNN) feature, we propose to represent the input image as a sequence of detected objects which feeds as the source sequence of the RNN model. In this way, the sequential representation of an image can be naturally translated to a sequence of words, as the target sequence of the RNN model. To represent the image in a sequential way, we extract the objects features in the image and arrange them in a order using convolutional neural networks. To further leverage the visual information from the encoded objects, a sequential attention layer is introduced to selectively attend to the objects that are related to generate corresponding words in the sentences. Extensive experiments are conducted to validate the proposed approach on popular benchmark dataset, i.e., MS COCO, and the proposed model surpasses the state-of-the-art methods in all metrics following the dataset splits of previous work. The proposed approach is also evaluated by the evaluation server of MS COCO captioning challenge, and achieves very competitive results, e.g., a CIDEr of 1.029 (c5) and 1.064 (c40). Fuchun Sun 0001, Changhu Wang, Feng Wang 0015, Alan L. Yuille |
IJCAI | 5 |
| 2017 | Object Recognition with and without ObjectsabstractWhile recent deep neural networks have achieved a promising performance on object recognition, they rely implicitly on the visual contents of the whole image. In this paper, we train deep neural networks on the foreground (object) and background (context) regions of images respectively. Considering human recognition in the same situations, networks trained on the pure background without objects achieves highly reasonable recognition performance that beats humans by a large margin if only given context. However, humans still outperform networks with pure object available, which indicates networks and human beings have different mechanisms in understanding an image. Furthermore, we straightforwardly combine multiple trained networks to explore different visual cues learned by different networks. Experiments show that useful visual hints can be explicitly learned separately and then combined to achieve higher performance, which verifies the advantages of the proposed framework. Zhuotun Zhu, Lingxi Xie, Alan L. Yuille |
IJCAI | 3 |
| 2017 | Deep Supervision for Pancreatic Cyst Segmentation in Abdominal CT Scans
Yuyin Zhou, Lingxi Xie, Elliot K. Fishman, Alan L. Yuille |
MICCAI (3) | 4 |
| 2017 | A Fixed-Point Model for Pancreas Segmentation in Abdominal CT Scans
Yuyin Zhou, Lingxi Xie, Wei Shen 0002, Yan Wang 0033, Elliot K. Fishman, Alan L. Yuille |
MICCAI (1) | 6 |
| 2017 | NormFace: L2 Hypersphere Embedding for Face VerificationabstractThanks to the recent developments of Convolutional Neural Networks, the performance of face verification methods has increased rapidly. In a typical face verification method, feature normalization is a critical step for boosting performance. This motivates us to introduce and study the effect of normalization during training. But we find this is non-trivial, despite normalization being differentiable. We identify and study four issues related to normalization through mathematical analysis, which yields understanding and helps with parameter settings. Based on this analysis we propose two strategies for training using normalized features. The first is a modification of softmax loss, which optimizes cosine similarity instead of inner-product. The second is a reformulation of metric learning by introducing an agent vector for each class. We show that both strategies, and small variants, consistently improve performance by between 0.2% to 0.4% on the LFW dataset based on two models. This is significant because the performance of the two models on LFW dataset is close to saturation at over 98%. Feng Wang 0015, Xiang Xiang 0001, Jian Cheng 0003, Alan L. Yuille |
ACM Multimedia | 4 |
| 2017 | Label Distribution Learning ForestsabstractLabel distribution learning (LDL) is a general learning framework, which assigns to an instance a distribution over a set of labels rather than a single label or multiple labels. Current LDL methods have either restricted assumptions on the expression form of the label distribution or limitations in representation learning, e.g., to learn deep features in an end-to-end manner. This paper presents label distribution learning forests (LDLFs) - a novel label distribution learning algorithm based on differentiable decision trees, which have several advantages: 1) Decision trees have the potential to model any general form of label distributions by a mixture of leaf node predictions. 2) The learning of differentiable decision trees can be combined with representation learning. We define a distribution-based loss function for a forest, enabling all the trees to be learned jointly, and show that an update function for leaf node predictions, which guarantees a strict decrease of the loss function, can be derived by variational bounding. The effectiveness of the proposed LDLFs is verified on several LDL tasks and a computer vision application, showing significant improvements to the state-of-the-art LDL methods. Wei Shen 0002, Kai Zhao 0012, Yilu Guo, Alan L. Yuille |
NIPS | 4 |
| 2017 | PASCAL Boundaries: A Semantic Boundary Dataset with a Deep Semantic Boundary DetectorabstractIn this paper, we address the task of instance-level semantic boundary detection. To this end, we generate a large database consisting of more than 10k images (which is 20 bigger than existing edge detection databases) along with ground truth boundaries between 459 semantic classes including instances from both foreground objects and different types of background, and call it the PASCAL Boundaries dataset. This is a timely introduction of the dataset since results on the existing standard dataset for measuring edge detection performance, i.e. BSDS500, have started to saturate. In addition to generating a new dataset, we propose a deep network-based multi-scale semantic boundary detector and name it Multi-scale Deep Semantic Boundary Detector (M-DSBD). We evaluate M-DSBD on PASCAL boundaries and compare it to baselines. We test the transfer capabilities of our model by evaluating on MS-COCO and BSDS500 and show that our model transfers well. Vittal Premachandran, Boyan Bonev 0001, Xiaochen Lian, Alan L. Yuille |
WACV | 4 |
| 2017 | Editorial- Deep Learning for Computer Vision
Ross B. Girshick, Iasonas Kokkinos, Ivan Laptev, Jitendra Malik, George Papandreou, Andrea Vedaldi, Xiaogang Wang 0001, Shuicheng Yan, Alan L. Yuille |
Comput. Vis. Image Underst. | 9 |
| 2017 | Empirical Minimum Bayes Risk PredictionabstractWhen building vision systems that predict structured objects such as image segmentations or human poses, a crucial concern is performance under task-specific evaluation measures (e.g., Jaccard Index or Average Precision). An ongoing research challenge is to optimize predictions so as to maximize performance on such complex measures. In this work, we present a simple meta-algorithm that is surprisingly effective - Empirical Min Bayes Risk. EMBR takes as input a pre-trained model that would normally be the final product and learns three additional parameters so as to optimize performance on the complex instance-level high-order task-specific measure. We demonstrate EMBR in several domains, taking existing state-of-the-art algorithms and improving performance up to 8 percent, simply by learning three extra parameters. Our code is publicly available and the results presented in this paper can be replicated from our code-release. Vittal Premachandran, Daniel Tarlow, Alan L. Yuille, Dhruv Batra |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | Semi-Supervised Sparse Representation Based Classification for Face Recognition With Insufficient Labeled SamplesabstractThis paper addresses the problem of face recognition when there is only few, or even only a single, labeled examples of the face that we wish to recognize. Moreover, these examples are typically corrupted by nuisance variables, both linear (i.e., additive nuisance variables, such as bad lighting and wearing of glasses) and non-linear (i.e., non-additive pixel-wise nuisance variables, such as expression changes). The small number of labeled examples means that it is hard to remove these nuisance variables between the training and testing faces to obtain good recognition performance. To address the problem, we propose a method called semi-supervised sparse representation-based classification. This is based on recent work on sparsity, where faces are represented in terms of two dictionaries: a gallery dictionary consisting of one or more examples of each person, and a variation dictionary representing linear nuisance variables (e.g., different lighting conditions and different glasses). The main idea is that: 1) we use the variation dictionary to characterize the linear nuisance variables via the sparsity framework and 2) prototype face images are estimated as a gallery dictionary via a Gaussian mixture model, with mixed labeled and unlabeled samples in a semi-supervised manner, to deal with the non-linear nuisance variations between labeled and unlabeled samples. We have done experiments with insufficient labeled samples, even when there is only a single labeled sample per person. Our results on the AR, Multi-PIE, CAS-PEAL, and LFW databases demonstrate that the proposed method is able to deliver significantly improved performance over existing methods. Yuan Gao 0015, Jiayi Ma 0001, Alan L. Yuille |
IEEE Trans. Image Process. | 3 |
| 2017 | DeepSkeleton: Learning Multi-Task Scale-Associated Deep Side Outputs for Object Skeleton Extraction in Natural ImagesabstractObject skeletons are useful for object representation and object detection. They are complementary to the object contour, and provide extra information, such as how object scale (thickness) varies among object parts. But object skeleton extraction from natural images is very challenging, because it requires the extractor to be able to capture both local and non-local image context in order to determine the scale of each skeleton pixel. In this paper, we present a novel fully convolutional network with multiple scale-associated side outputs to address this problem. By observing the relationship between the receptive field sizes of the different layers in the network and the skeleton scales they can capture, we introduce two scale-associated side outputs to each stage of the network. The network is trained by multi-task learning, where one task is skeleton localization to classify whether a pixel is a skeleton pixel or not, and the other is skeleton scale prediction to regress the scale of each skeleton pixel. Supervision is imposed at different stages by guiding the scale-associated side outputs toward the ground-truth skeletons at the appropriate scales. The responses of the multiple scale-associated side outputs are then fused in a scale-specific way to detect skeleton pixels using multiple scales effectively. Our method achieves promising results on two skeleton extraction datasets, and significantly outperforms other competitors. In addition, the usefulness of the obtained skeletons and scales (thickness) are verified on two object detection applications: foreground object segmentation and object proposal detection. Wei Shen 0002, Kai Zhao 0012, Yuan Jiang 0002, Yan Wang 0033, Xiang Bai, Alan L. Yuille |
IEEE Trans. Image Process. | 6 |
| 2016 | Recognizing Actions in 3D Using Action-Snippets and Activated SimplicesabstractPose-based action recognition in 3D is the task of recognizing an action (e.g., walking or running) from a sequence of 3D skeletal poses. This is challenging because of variations due to different ways of performing the same action and inaccuracies in the estimation of the skeletal poses. The training data is usually small and hence complex classifiers risk over-fitting the data. We address this task by action-snippets which are short sequences of consecutive skeletal poses capturing the temporal relationships between poses in an action. We propose a novel representation for action-snippets, called activated simplices. Each activity is represented by a manifold which is approximated by an arrangement of activated simplices. A sequence (of action-snippets) is classified by selecting the closest manifold and outputting the corresponding activity. This is a simple classifier which helps avoid over-fitting the data but which significantly outperforms state-of-the-art methods on standard benchmarks. Chunyu Wang 0001, John Flynn, Yizhou Wang 0001, Alan L. Yuille |
AAAI | 4 |
| 2016 | Pose-Guided Human Parsing by an AND/OR Graph Using Pose-Context FeaturesabstractParsing human into semantic parts is crucial to human-centric analysis. In this paper, we propose a human parsing pipeline that uses pose cues, e.g., estimates of human joint locations, to provide pose-guided segment proposals for semantic parts. These segment proposals are ranked using standard appearance cues, deep-learned semantic feature, and a novel pose feature called pose-context. Then these proposals are selected and assembled using an And-Or graph to output a parse of the person. The And-Or graph is able to deal with large human appearance variability due to pose, choice of clothing, etc. We evaluate our approach on the popular Penn-Fudan pedestrian parsing dataset, showing that it significantly outperforms the state of the art, and perform diagnostics to demonstrate the effectiveness of different stages of our pipeline. Fangting Xia, Jun Zhu 0001, Peng Wang 0001, Alan L. Yuille |
AAAI | 4 |
| 2016 | Semantic Image Segmentation with Task-Specific Edge Detection Using CNNs and a Discriminatively Trained Domain TransformabstractDeep convolutional neural networks (CNNs) are the backbone of state-of-art semantic image segmentation systems. Recent work has shown that complementing CNNs with fully-connected conditional random fields (CRFs) can significantly enhance their object localization accuracy, yet dense CRF inference is computationally expensive. We propose replacing the fully-connected CRF with domain transform (DT), a modern edge-preserving filtering method in which the amount of smoothing is controlled by a reference edge map. Domain transform filtering is several times faster than dense CRF inference and we show that it yields comparable semantic segmentation results, accurately capturing object boundaries. Importantly, our formulation allows learning the reference edge map from intermediate CNN features instead of using the image gradient magnitude as in standard DT filtering. This produces task-specific edges in an end-to-end trainable system optimizing the target semantic segmentation quality. Liang-Chieh Chen, Jonathan T. Barron, George Papandreou, Kevin Murphy 0002, Alan L. Yuille |
CVPR | 5 |
| 2016 | Attention to Scale: Scale-Aware Semantic Image SegmentationabstractIncorporating multi-scale features in fully convolutional neural networks (FCNs) has been a key element to achieving state-of-the-art performance on semantic image segmentation. One common way to extract multi-scale features is to feed multiple resized input images to a shared deep network and then merge the resulting features for pixelwise classification. In this work, we propose an attention mechanism that learns to softly weight the multi-scale features at each pixel location. We adapt a state-of-the-art semantic image segmentation model, which we jointly train with multi-scale input images and the attention model. The proposed attention model not only outperforms averageand max-pooling, but allows us to diagnostically visualize the importance of features at different positions and scales. Moreover, we show that adding extra supervision to the output at each scale is essential to achieving excellent performance when merging multi-scale features. We demonstrate the effectiveness of our model with extensive experiments on three challenging datasets, including PASCAL-Person-Part, PASCAL VOC 2012 and a subset of MS-COCO 2014. Liang-Chieh Chen, Yi Yang 0007, Jiang Wang 0001, Wei Xu 0017, Alan L. Yuille |
CVPR | 5 |
| 2016 | Generation and Comprehension of Unambiguous Object DescriptionsabstractWe propose a method that can generate an unambiguous description (known as a referring expression) of a specific object or region in an image, and which can also comprehend or interpret such an expression to infer which object is being described. We show that our method outperforms previous methods that generate descriptions of objects without taking into account other potentially ambiguous objects in the scene. Our model is inspired by recent successes of deep learning methods for image captioning, but while image captioning is difficult to evaluate, our task allows for easy objective evaluation. We also present a new large-scale dataset for referring expressions, based on MSCOCO. We have released the dataset and a toolbox for visualization and evaluation, see https://github.com/ mjhucla/Google_Refexp_toolbox. Junhua Mao, Jonathan Huang, Alexander Toshev, Oana-Maria Camburu, Alan L. Yuille, Kevin Murphy 0002 |
CVPR | 5 |
| 2016 | Mining 3D Key-Pose-Motifs for Action RecognitionabstractRecognizing an action from a sequence of 3D skeletal poses is a challenging task. First, different actors may perform the same action in various styles. Second, the estimated poses are sometimes inaccurate. These challenges can cause large variations between instances of the same class. Third, the datasets are usually small, with only a few actors performing few repetitions of each action. Hence training complex classifiers risks over-fitting the data. We address this task by mining a set of key-pose-motifs for each action class. A key-pose-motif contains a set of ordered poses, which are required to be close but not necessarily adjacent in the action sequences. The representation is robust to style variations. The key-pose-motifs are represented in terms of a dictionary using soft-quantization to deal with inaccuracies caused by quantization. We propose an efficient algorithm to mine key-pose-motifs taking into account of these probabilities. We classify a sequence by matching it to the motifs of each class and selecting the class that maximizes the matching score. This simple classifier obtains state-of the-art performance on two benchmark datasets. Chunyu Wang 0001, Yizhou Wang 0001, Alan L. Yuille |
CVPR | 3 |
| 2016 | InterActive: Inter-Layer Activeness PropagationabstractAn increasing number of computer vision tasks can be tackled with deep features, which are the intermediate outputs of a pre-trained Convolutional Neural Network. Despite the astonishing performance, deep features extracted from low-level neurons are still below satisfaction, arguably because they cannot access the spatial context contained in the higher layers. In this paper, we present InterActive, a novel algorithm which computes the activeness of neurons and network connections. Activeness is propagated through a neural network in a top-down manner, carrying highlevel context and improving the descriptive power of lowlevel and mid-level neurons. Visualization indicates that neuron activeness can be interpreted as spatial-weighted neuron responses. We achieve state-of-the-art classification performance on a wide range of image datasets. Lingxi Xie, Liang Zheng 0001, Jingdong Wang 0001, Alan L. Yuille, Qi Tian 0001 |
CVPR | 4 |
| 2016 | Symmetric Non-rigid Structure from Motion for Category-Specific Object Structure Estimation
Yuan Gao 0015, Alan L. Yuille |
ECCV (2) | 2 |
| 2016 | DOC: Deep OCclusion Estimation from a Single Image
Peng Wang 0001, Alan L. Yuille |
ECCV (1) | 2 |
| 2016 | Zoom Better to See Clearer: Human and Object Parsing with Hierarchical Auto-Zoom Net
Fangting Xia, Peng Wang 0001, Liang-Chieh Chen, Alan L. Yuille |
ECCV (5) | 4 |
| 2016 | Geometric Neural Phrase Pooling: Modeling the Spatial Co-occurrence of Neurons
Lingxi Xie, Qi Tian 0001, John Flynn, Jingdong Wang 0001, Alan L. Yuille |
ECCV (1) | 5 |
| 2016 | Joint Image-Text Representation by Gaussian Visual-Semantic EmbeddingabstractHow to jointly represent images and texts is important for tasks involving both modalities. Visual-semantic embedding models have been recently proposed and shown to be effective. The key idea is that by learning a mapping from images into a semantic text space, the algorithm is able to learn a compact and effective joint representation. However, existing approaches simply map each text concept to a single point in the semantic space. Mapping instead to a density distribution provides many interesting advantages, including better capturing uncertainty about each text concept, and enabling better geometric interpretation of concepts such as inclusion, intersection, etc. In this work, we present a novel Gaussian Visual-Semantic Embedding (GVSE) model, which leverages the visual information to model text concepts as Gaussian distributions in semantic space. Experiments in two tasks, image classification and text-based image retrieval on the large scale MIT Places205 dataset, have demonstrated the superiority of our method over existing approaches, with higher accuracy and better robustness. Zhou Ren, Hailin Jin, Zhe Lin 0001, Alan L. Yuille |
ACM Multimedia | 5 |
| 2016 | SURGE: Surface Regularized Geometry Estimation from a Single ImageabstractThis paper introduces an approach to regularize 2.5D surface normal and depth predictions at each pixel given a single input image. The approach infers and reasons about the underlying 3D planar surfaces depicted in the image to snap predicted normals and depths to inferred planar surfaces, all while maintaining fine detail within objects. Our approach comprises two components: (i) a fourstream convolutional neural network (CNN) where depths, surface normals, and likelihoods of planar region and planar boundary are predicted at each pixel, followed by (ii) a dense conditional random field (DCRF) that integrates the four predictions such that the normals and depths are compatible with each other and regularized by the planar region and planar boundary information. The DCRF is formulated such that gradients can be passed to the surface normal and depth CNNs via backpropagation. In addition, we propose new planar wise metrics to evaluate geometry consistency within planar surfaces, which are more tightly related to dependent 3D editing applications. We show that our regularization yields a 30% relative improvement in planar consistency on the NYU v2 dataset. Peng Wang 0001, Xiaohui Shen, Bryan C. Russell, Scott Cohen, Brian L. Price, Alan L. Yuille |
NIPS | 6 |
| 2016 | Training and Evaluating Multimodal Word Embeddings with Large-scale Web Annotated ImagesabstractIn this paper, we focus on training and evaluating effective word embeddings with both text and visual information. More specifically, we introduce a large-scale dataset with 300 million sentences describing over 40 million images crawled and downloaded from publicly available Pins (i.e. an image with sentence descriptions uploaded by users) on Pinterest. This dataset is more than 200 times larger than MS COCO, the standard large-scale image dataset with sentence descriptions. In addition, we construct an evaluation dataset to directly assess the effectiveness of word embeddings in terms of finding semantically similar or related words and phrases. The word/phrase pairs in this evaluation dataset are collected from the click data with millions of users in an image search system, thus contain rich semantic relationships. Based on these datasets, we propose and compare several Recurrent Neural Networks (RNNs) based multimodal (text and image) models. Experiments show that our model benefits from incorporating the visual information into the word embeddings, and a weight sharing strategy is crucial for learning such multimodal embeddings. The project page is: http://www.stat.ucla.edu/~junhua.mao/multimodal_embedding.html (The datasets introduced in this work will be gradually released on the project page.). Junhua Mao, Yushi Jing, Alan L. Yuille |
NIPS | 4 |
| 2016 | Complexity of Representation and Inference in Compositional Models with Part SharingabstractThis paper performs a complexity analysis of a class of serial and parallel compositional models of multiple objects and shows that they enable efficient representation and rapid inference. Compositional models are generative and represent objects in a hierarchically distributed manner in terms of parts and subparts, which are constructed recursively by part-subpart compositions. Parts are represented more coarsely at higher level of the hierarchy, so that the upper levels give coarse summary descriptions (e.g., there is a horse in the image) while the lower levels represents the details (e.g., the positions of the legs of the horse). This hierarchically distributed representation obeys the executive summary principle, meaning that a high level executive only requires a coarse summary description and can, if necessary, get more details by consulting lower level executives. The parts and subparts are organized in terms of hierarchical dictionaries which enables part sharing between different objects allowing efficient representation of many objects. The first main contribution of this paper is to show that compositional models can be mapped onto a parallel visual architecture similar to that used by bio- inspired visual models such as deep convolutional networks but more explicit in terms of representation, hence enabling part detection as well as object detection, and suitable for complexity analysis. Inference algorithms can be run on this architecture to exploit the gains caused by part sharing and executive summary. Effectively, this compositional architecture enables us to perform exact inference simultaneously over a large class of generative models of objects. The second contribution is an analysis of the complexity of compositional models in terms of computation time (for serial computers) and numbers of nodes (e.g., "neurons") for parallel computers. In particular, we compute the complexity gains by part sharing and executive summary and their dependence on how the dictionary scales with the level of the hierarchy. We explore three regimes of scaling behavior where the dictionary size (i) increases exponentially with the level of the hierarchy, (ii) is determined by an unsupervised compositional learning algorithm applied to real data, (iii) decreases exponentially with scale. This analysis shows that in some regimes the use of shared parts enables algorithms which can perform inference in time linear in the number of levels for an exponential number of objects. In other regimes part sharing has little advantage for serial computers but can enable linear processing on parallel computers. Alan L. Yuille, Roozbeh Mottaghi |
J. Mach. Learn. Res. | 1 |
| 2016 | Human-Machine CRFs for Identifying Bottlenecks in Scene UnderstandingabstractRecent trends in image understanding have pushed for scene understanding models that jointly reason about various tasks such as object detection, scene recognition, shape analysis, contextual reasoning, and local appearance based classifiers. In this work, we are interested in understanding the roles of these different tasks in improved scene understanding, in particular semantic segmentation, object detection and scene recognition. Towards this goal, we "plug-in" human subjects for each of the various components in a conditional random field model. Comparisons among various hybrid human-machine CRFs give us indications of how much "head room" there is to improve scene understanding by focusing research efforts on various individual tasks. Roozbeh Mottaghi, Sanja Fidler, Alan L. Yuille, Raquel Urtasun, Devi Parikh |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2016 | Non-Rigid Point Set Registration by Preserving Global and Local StructuresabstractIn previous work on point registration, the input point sets are often represented using Gaussian mixture models and the registration is then addressed through a probabilistic approach, which aims to exploit global relationships on the point sets. For non-rigid shapes, however, the local structures among neighboring points are also strong and stable and thus helpful in recovering the point correspondence. In this paper, we formulate point registration as the estimation of a mixture of densities, where local features, such as shape context, are used to assign the membership probabilities of the mixture model. This enables us to preserve both global and local structures during matching. The transformation between the two point sets is specified in a reproducing kernel Hilbert space and a sparse approximation is adopted to achieve a fast implementation. Extensive experiments on both synthesized and real data show the robustness of our approach under various types of distortions, such as deformation, noise, outliers, rotation, and occlusion. It greatly outperforms the state-of-the-art methods, especially when the data is badly degraded. Jiayi Ma 0001, Ji Zhao 0001, Alan L. Yuille |
IEEE Trans. Image Process. | 3 |
| 2015 | A Computational Model of Jazz Improvisation Inspired by Language
Cody Komers, Alan L. Yuille |
CogSci | 2 |
| 2015 | Parsing occluded people by flexible compositionsabstractThis paper presents an approach to parsing humans when there is significant occlusion. We model humans using a graphical model which has a tree structure building on recent work [32, 6] and exploit the connectivity prior that, even in presence of occlusion, the visible nodes form a connected subtree of the graphical model. We call each connected subtree a flexible composition of object parts. This involves a novel method for learning occlusion cues. During inference we need to search over a mixture of different flexible models. By exploiting part sharing, we show that this inference can be done extremely efficiently requiring only twice as many computations as searching for the entire object (i.e., not modeling occlusion). We evaluate our model on the standard benchmarked “We Are Family” Stickmen dataset and obtain significant performance improvements over the best alternative algorithms. Xianjie Chen, Alan L. Yuille |
CVPR | 2 |
| 2015 | Region-based temporally consistent video post-processingabstractWe study the problem of temporally consistent video post-processing. Previous post-processing algorithms usually either fail to keep high fidelity or fail to keep temporal consistency of output videos. In this paper, we observe experimentally that many image/video enhancement algorithms enforce a spatially consistent prior on the enhancement. More precisely, within a local region, the enhancement is consistent, i.e., pixels with the same RGB values will get the same enhancement values. Using this prior, we segment each frame into several regions and temporally-spatially adjust the enhancement of regions of different frames, taking into account fidelity, temporal consistency and spatial consistency. User study, objective measurement and visual quality comparisons are conducted. The experimental results demonstrate that our output videos can keep high fidelity and temporal consistency at the same time. Xuan Dong 0001, Boyan Bonev 0001, Yu Zhu 0004, Alan L. Yuille |
CVPR | 4 |
| 2015 | Towards unified depth and semantic prediction from a single imageabstractDepth estimation and semantic segmentation are two fundamental problems in image understanding. While the two tasks are strongly correlated and mutually beneficial, they are usually solved separately or sequentially. Motivated by the complementary properties of the two tasks, we propose a unified framework for joint depth and semantic prediction. Given an image, we first use a trained Convolutional Neural Network (CNN) to jointly predict a global layout composed of pixel-wise depth values and semantic labels. By allowing for interactions between the depth and semantic information, the joint network provides more accurate depth prediction than a state-of-the-art CNN trained solely for depth prediction [6]. To further obtain fine-level details, the image is decomposed into local segments for region-level depth and semantic prediction under the guidance of global layout. Utilizing the pixel-wise global prediction and region-wise local prediction, we formulate the inference problem in a two-layer Hierarchical Conditional Random Field (HCRF) to produce the final depth and semantic map. As demonstrated in the experiments, our approach effectively leverages the advantages of both tasks and provides the state-of-the-art results. Peng Wang 0001, Xiaohui Shen, Zhe Lin 0001, Scott Cohen, Brian L. Price, Alan L. Yuille |
CVPR | 6 |
| 2015 | Semantic part segmentation using compositional model combining shape and appearanceabstractIn this paper, we study the problem of semantic part segmentation for animals. This is more challenging than standard object detection, object segmentation and pose estimation tasks because semantic parts of animals often have similar appearance and highly varying shapes. To tackle these challenges, we build a mixture of compositional models to represent the object boundary and the boundaries of semantic parts. And we incorporate edge, appearance, and semantic part cues into the compositional model. Given part-level segmentation annotation, we develop a novel algorithm to learn a mixture of compositional models under various poses and viewpoints for certain animal classes. Furthermore, a linear complexity algorithm is offered for efficient inference of the compositional model using dynamic programming. We evaluate our method for horse and cow using a newly annotated dataset on Pascal VOC 2010 which has pixelwise part labels. Experimental results demonstrate the effectiveness of our method. Jianyu Wang 0001, Alan L. Yuille |
CVPR | 2 |
| 2015 | Modeling deformable gradient compositions for single-image super-resolutionabstractWe propose a single-image super-resolution method based on the gradient reconstruction. To predict the gradient field, we collect a dictionary of gradient patterns from an external set of images. We observe that there are patches representing singular primitive structures (e.g. a single edge), and non-singular ones (e.g. a triplet of edges). Based on the fact that singular primitive patches are more invariant to the scale change (i.e. have less ambiguity across different scales), we represent the non-singular primitives as compositions of singular ones, each of which is allowed some deformation. Both the input patches and dictionary elements are decomposed to contain only singular primitives. The compositional aspect of the model makes the gradient field more reliable. The deformable aspect makes the dictionary more expressive. As shown in our experimental results, the proposed method outperforms the state-of-the-art methods. Yu Zhu 0004, Yanning Zhang 0001, Boyan Bonev 0001, Alan L. Yuille |
CVPR | 4 |
| 2015 | Learning Like a Child: Fast Novel Visual Concept Learning from Sentence Descriptions of ImagesabstractIn this paper, we address the task of learning novel visual concepts, and their interactions with other concepts, from a few images with sentence descriptions. Using linguistic context and visual features, our method is able to efficiently hypothesize the semantic meaning of new words and add them to its word dictionary so that they can be used to describe images which contain these novel concepts. Our method has an image captioning module based on [38] with several improvements. In particular, we propose a transposed weight sharing scheme, which not only improves performance on image captioning, but also makes the model more suitable for the novel concept learning task. We propose methods to prevent overfitting the new concepts. In addition, three novel concept datasets are constructed for this new task, and are publicly available on the project page. In the experiments, we show that our method effectively learns novel visual concepts from a few examples without disturbing the previously learned concepts. The project page is: www.stat.ucla.edu/junhua. mao/projects/child_learning.html. Junhua Mao, Yi Yang 0007, Jiang Wang 0001, Zhiheng Huang, Alan L. Yuille |
ICCV | 6 |
| 2015 | Weakly-and Semi-Supervised Learning of a Deep Convolutional Network for Semantic Image SegmentationabstractDeep convolutional neural networks (DCNNs) trained on a large number of images with strong pixel-level annotations have recently significantly pushed the state-of-art in semantic image segmentation. We study the more challenging problem of learning DCNNs for semantic image segmentation from either (1) weakly annotated training data such as bounding boxes or image-level labels or (2) a combination of few strongly labeled and many weakly labeled images, sourced from one or multiple datasets. We develop Expectation-Maximization (EM) methods for semantic image segmentation model training under these weakly supervised and semi-supervised settings. Extensive experimental evaluation shows that the proposed techniques can learn models delivering competitive results on the challenging PASCAL VOC 2012 image segmentation benchmark, while requiring significantly less annotation effort. We share source code implementing the proposed system at https://bitbucket.org/deeplab/deeplab-public. George Papandreou, Liang-Chieh Chen, Kevin Murphy 0002, Alan L. Yuille |
ICCV | 4 |
| 2015 | Scene-Domain Active Part Models for Object RepresentationabstractIn this paper, we are interested in enhancing the expressivity and robustness of part-based models for object representation, in the common scenario where the training data are based on 2D images. To this end, we propose scene-domain active part models (SDAPM), which reconstruct and characterize the 3D geometric statistics between object's parts in 3D scene-domain by using 2D training data in the image-domain alone. And on top of this, we explicitly model and handle occlusions in SDAPM. Together with the developed learning and inference algorithms, such a model provides rich object descriptions, including 2D object and parts localization, 3D landmark shape and camera viewpoint, which offers an effective representation to various image understanding tasks, such as object and parts detection, 3D landmark shape and viewpoint estimation from images. Experiments on the above tasks show that SDAPM outperforms previous part-based models, and thus demonstrates the potential of the proposed technique. Zhou Ren, Chaohui Wang, Alan L. Yuille |
ICCV | 3 |
| 2015 | Joint Object and Part Segmentation Using Deep Learned PotentialsabstractSegmenting semantic objects from images and parsing them into their respective semantic parts are fundamental steps towards detailed object understanding in computer vision. In this paper, we propose a joint solution that tackles semantic object and part segmentation simultaneously, in which higher object-level context is provided to guide part segmentation, and more detailed part-level localization is utilized to refine object segmentation. Specifically, we first introduce the concept of semantic compositional parts (SCP) in which similar semantic parts are grouped and shared among different objects. A two-stream fully convolutional network (FCN) is then trained to provide the SCP and object potentials at each pixel. At the same time, a compact set of segments can also be obtained from the SCP predictions of the network. Given the potentials and the generated segments, in order to explore long-range context, we finally construct an efficient fully connected conditional random field (FCRF) to jointly predict the final object and part labels. Extensive evaluation on three different datasets shows that our approach can mutually enhance the performance of object and part segmentation, and outperforms the current state-of-the-art on both tasks. Peng Wang 0001, Xiaohui Shen, Zhe Lin 0001, Scott Cohen, Brian L. Price, Alan L. Yuille |
ICCV | 6 |
| 2015 | One Shot Learning via Compositions of Meaningful PatchesabstractThe task of discriminating one object from another is almost trivial for a human being. However, this task is computationally taxing for most modern machine learning methods, whereas, we perform this task at ease given very few examples for learning. It has been proposed that the quick grasp of concept may come from the shared knowledge between the new example and examples previously learned. We believe that the key to one-shot learning is the sharing of common parts as each part holds immense amounts of information on how a visual concept is constructed. We propose an unsupervised method for learning a compact dictionary of image patches representing meaningful components of an objects. Using those patches as features, we build a compositional model that outperforms a number of popular algorithms on a one-shot learning task. We demonstrate the effectiveness of this approach on hand-written digits and show that this model generalizes to multiple datasets. Alex Wong 0001, Alan L. Yuille |
ICCV | 2 |
| 2015 | Temporally consistent region-based video exposure correctionabstractWe analyze the problem of temporally consistent video exposure correction. Existing methods usually either fail to evaluate optimal exposure for every region or cannot get temporally consistent correction results. In addition, the contrast is often lost when the detail is not preserved properly during correction. In this paper, we use the block-based energy minimization to evaluate the temporally consistent exposure, which considers 1) the maximization of the visibility of all contents, 2) keeping the relative difference between neighboring regions, and 3) temporally consistent exposure of corresponding contents in different frames. Then, based on Weber contrast definition, we propose a contrast preserving exposure correction method. Experimental results show that our method enables better temporally consistent exposure evaluation and produces contrast preserving outputs. Xuan Dong 0001, Lu Yuan 0001, Weixin Li 0001, Alan L. Yuille |
ICME | 4 |
| 2015 | Learning Deep Structured ModelsabstractMany problems in real-world applications involve predicting several random variables that are statistically related. Markov random fields (MRFs) are a great mathematical tool to encode such dependencies. The goal of this paper is to combine MRFs with deep learning to estimate complex representations while taking into account the dependencies between the output random variables. Towards this goal, we propose a training algorithm that is able to learn structured models jointly with deep features that form the MRF potentials. Our approach is efficient as it blends learning and inference and makes use of GPU acceleration. We demonstrate the effectiveness of our algorithm in the tasks of predicting words from noisy images, as well as tagging of Flickr photographs. We show that joint learning of the deep features and the MRF parameters results in significant performance gains. Liang-Chieh Chen, Alexander G. Schwing, Alan L. Yuille, Raquel Urtasun |
ICML | 3 |
| 2015 | Error Factor Analysis for Wild Scene Image-LabellingabstractPASCAL VOC Segmentation Challenge [10] is currently considered as one of the datasets that reflect the image segmentation difficulties for real world scenarios [29]. However, current evaluation is simply based on a single Inter-section Over Union (IOU) score. In this paper, we try to discover the error factors under the IOU, which makes the results more informative to understand rather than a black box. Specifically, we decompose the error into three error types in terms of object characteristics, i.e. general, appearance and shape. Each error type is composed of respective factors, e.g. size and aspect ratio for general, appearance distinctiveness for appearance, etc. Finally, for each factor and error type, we perform analysis over its impact on and correlation with the final IOU through robust regression. Our experiments show that these error factors have significant relationship with the given IOU accuracy, and the analysis provides practical guidance on further improvement of the given algorithm. Peng Wang 0001, Alan L. Yuille |
WACV | 2 |
| 2014 | Parsing Semantic Parts of Cars Using Graphical Models and Segment Appearance Consistency
Wenhao Lu, Xiaochen Lian, Alan L. Yuille |
BMVC | 3 |
| 2014 | Detect What You Can: Detecting and Representing Objects Using Holistic Models and Body PartsabstractDetecting objects becomes difficult when we need to deal with large shape deformation, occlusion and low resolution. We propose a novel approach to i) handle large deformations and partial occlusions in animals (as examples of highly deformable objects), ii) describe them in terms of body parts, and iii) detect them when their body parts are hard to detect (e.g., animals depicted at low resolution). We represent the holistic object and body parts separately and use a fully connected model to arrange templates for the holistic object and body parts. Our model automatically decouples the holistic object or body parts from the model when they are hard to detect. This enables us to represent a large number of holistic object and body part combinations to better deal with different "detectability" patterns caused by deformations, occlusion and/or low resolution. We apply our method to the six animal categories in the PASCAL VOC dataset and show that our method significantly improves state-of-the-art (by 4.1% AP) and provides a richer representation for objects. During training we use annotations for body parts (e.g., head, torso, etc.), making use of a new dataset of fully annotated object parts for PASCAL VOC 2010, which provides a mask for each part. Xianjie Chen, Roozbeh Mottaghi, Xiaobai Liu, Sanja Fidler, Raquel Urtasun, Alan L. Yuille |
CVPR | 6 |
| 2014 | The Secrets of Salient Object SegmentationabstractIn this paper we provide an extensive evaluation of fixation prediction and salient object segmentation algorithms as well as statistics of major datasets. Our analysis identifies serious design flaws of existing salient object benchmarks, called the dataset design bias, by over emphasising the stereotypical concepts of saliency. The dataset design bias does not only create the discomforting disconnection between fixations and salient object segmentation, but also misleads the algorithm designing. Based on our analysis, we propose a new high quality dataset that offers both fixation and salient object segmentation ground-truth. With fixations and salient object being presented simultaneously, we are able to bridge the gap between fixations and salient objects, and propose a novel method for salient object segmentation. Finally, we report significant benchmark progress on 3 existing datasets of segmenting salient objects. Yin Li 0003, Christof Koch, James M. Rehg, Alan L. Yuille |
CVPR | 5 |
| 2014 | The Role of Context for Object Detection and Semantic Segmentation in the WildabstractIn this paper we study the role of context in existing state-of-the-art detection and segmentation approaches. Towards this goal, we label every pixel of PASCAL VOC 2010 detection challenge with a semantic category. We believe this data will provide plenty of challenges to the community, as it contains 520 additional classes for semantic segmentation and object detection. Our analysis shows that nearest neighbor based approaches perform poorly on semantic segmentation of contextual classes, showing the variability of PASCAL imagery. Furthermore, improvements of existing contextual models for detection is rather modest. In order to push forward the performance in this difficult scenario, we propose a novel deformable part-based model, which exploits both local context around each candidate detection as well as global context at the level of the scene. We show that this contextual reasoning significantly helps in detecting objects at all scales. Roozbeh Mottaghi, Xianjie Chen, Xiaobai Liu, Nam-Gyu Cho, Seong-Whan Lee, Sanja Fidler, Raquel Urtasun, Alan L. Yuille |
CVPR | 8 |
| 2014 | Modeling Image Patches with a Generic Dictionary of Mini-epitomesabstractThe goal of this paper is to question the necessity of features like SIFT in categorical visual recognition tasks. As an alternative, we develop a generative model for the raw intensity of image patches and show that it can support image classification performance on par with optimized SIFT-based techniques in a bag-of-visual-words setting. Key ingredient of the proposed model is a compact dictionary of mini-epitomes, learned in an unsupervised fashion on a large collection of images. The use of epitomes allows us to explicitly account for photometric and position variability in image appearance. We show that this flexibility considerably increases the capacity of the dictionary to accurately approximate the appearance of image patches and support recognition tasks. For image classification, we develop histogram-based image encoding methods tailored to the epitomic representation, as well as an "epitomic footprint" encoding which is easy to visualize and highlights the generative nature of our model. We discuss in detail computational aspects and develop efficient algorithms to make the model scalable to large tasks. The proposed techniques are evaluated with experiments on the challenging PASCAL VOC 2007 image classification benchmark. George Papandreou, Liang-Chieh Chen, Alan L. Yuille |
CVPR | 3 |
| 2014 | Robust Estimation of 3D Human Poses from a Single ImageabstractHuman pose estimation is a key step to action recognition. We propose a method of estimating 3D human poses from a single image, which works in conjunction with an existing 2D pose/joint detector. 3D pose estimation is challenging because multiple 3D poses may correspond to the same 2D pose after projection due to the lack of depth information. Moreover, current 2D pose estimators are usually inaccurate which may cause errors in the 3D estimation. We address the challenges in three ways: (i) We represent a 3D pose as a linear combination of a sparse set of bases learned from 3D human skeletons. (ii) We enforce limb length constraints to eliminate anthropomorphically implausible skeletons. (iii) We estimate a 3D pose by minimizing the 1-norm error between the projection of the 3D pose and the corresponding 2D detection. The 1-norm loss term is robust to inaccurate 2D joint estimations. We use the alternating direction method (ADM) to solve the optimization problem efficiently. Our approach outperforms the state-of-the-arts on three benchmark datasets. Chunyu Wang 0001, Yizhou Wang 0001, Zhouchen Lin, Alan L. Yuille, Wen Gao 0001 |
CVPR | 4 |
| 2014 | Single Image Super-resolution Using Deformable PatchesabstractWe proposed a deformable patches based method for single image super-resolution. By the concept of deformation, a patch is not regarded as a fixed vector but a flexible deformation flow. Via deformable patches, the dictionary can cover more patterns that do not appear, thus becoming more expressive. We present the energy function with slow, smooth and flexible prior for deformation model. During example-based super-resolution, we develop the deformation similarity based on the minimized energy function for basic patch matching. For robustness, we utilize multiple deformed patches combination for the final reconstruction. Experiments evaluate the deformation effectiveness and super-resolution performance, showing that the deformable patches help improve the representation accuracy and perform better than the state-of-art methods. Yu Zhu 0004, Yanning Zhang 0001, Alan L. Yuille |
CVPR | 3 |
| 2014 | A Fast and Simple Algorithm for Producing Candidate Regions
Boyan Bonev 0001, Alan L. Yuille |
ECCV (3) | 2 |
| 2014 | Towards Unified Object Detection and Semantic Segmentation
Jian Dong 0011, Qiang Chen 0007, Shuicheng Yan, Alan L. Yuille |
ECCV (5) | 4 |
| 2014 | An Active Patch Model for Real World Texture and Appearance Classification
Junhua Mao, Jun Zhu 0001, Alan L. Yuille |
ECCV (3) | 3 |
| 2014 | Articulated Pose Estimation by a Graphical Model with Image Dependent Pairwise Relations
Xianjie Chen, Alan L. Yuille |
NIPS | 2 |
| 2014 | Learning From Weakly Supervised Data by The Expectation Loss SVM (e-SVM) algorithm
Jun Zhu 0001, Junhua Mao, Alan L. Yuille |
NIPS | 3 |
| 2014 | Scale-Space SIFT flowabstractThe state-of-the-art SIFT flow has been widely adopted for the general image matching task, especially in dealing with image pairs from similar scenes but with different object configurations. However, the way in which the dense SIFT features are computed at a fixed scale in the SIFT flow method limits its capability of dealing with scenes of large scale changes. In this paper, we propose a simple, intuitive, and very effective approach, Scale-Space SIFT flow, to deal with the large scale differences in different image locations. We introduce a scale field to the SIFT flow function to automatically explore the scale deformations. Our approach achieves similar performance as the SIFT flow method on general natural scenes but obtains significant improvement on the images with large scale differences. Compared with a recent method that addresses the similar problem, our approach shows its clear advantage being more effective, and significantly less demanding in memory and time requirement. Weichao Qiu, Xinggang Wang, Xiang Bai, Alan L. Yuille, Zhuowen Tu |
WACV | 4 |
| 2014 | A Shape Reconstructability Measure of Object Part Importance with Applications to Object Detection and Localization
Ge Guo 0002, Yizhou Wang 0001, Tingting Jiang 0001, Alan L. Yuille, Fang Fang 0003, Wen Gao 0001 |
Int. J. Comput. Vis. | 4 |
| 2014 | Guest Editorial: Geometry, Lighting, Motion, and Learning
Alan L. Yuille, Jiebo Luo 0001 |
Int. J. Comput. Vis. | 1 |
| 2014 | Robust Point Matching via Vector Field ConsensusabstractIn this paper, we propose an efficient algorithm, called vector field consensus, for establishing robust point correspondences between two sets of points. Our algorithm starts by creating a set of putative correspondences which can contain a very large number of false correspondences, or outliers, in addition to a limited number of true correspondences (inliers). Next, we solve for correspondence by interpolating a vector field between the two point sets, which involves estimating a consensus of inlier points whose matching follows a nonparametric geometrical constraint. We formulate this a maximum a posteriori (MAP) estimation of a Bayesian model with hidden/latent variables indicating whether matches in the putative set are outliers or inliers. We impose nonparametric geometrical constraints on the correspondence, as a prior distribution, using Tikhonov regularizers in a reproducing kernel Hilbert space. MAP estimation is performed by the EM algorithm which by also estimating the variance of the prior model (initialized to a large value) is able to obtain good estimates very quickly (e.g., avoiding many of the local minima inherent in this formulation). We illustrate this method on data sets in 2D and 3D and demonstrate that it is robust to a very large number of outliers (even up to 90%). We also show that in the special case where there is an underlying parametric geometrical model (e.g., the epipolar line constraint) that we obtain better results than standard alternatives like RANSAC if a large number of outliers are present. This suggests a two-stage strategy, where we use our nonparametric model to reduce the size of the putative set and then apply a parametric variant of our approach to estimate the geometric parameters. Our algorithm is computationally efficient and we provide code for others to use it. In addition, our approach is general and can be applied to other problems, such as learning with a badly corrupted training data set. Jiayi Ma 0001, Ji Zhao 0001, Jinwen Tian, Alan L. Yuille, Zhuowen Tu |
IEEE Trans. Image Process. | 4 |
| 2013 | Bottom-Up Segmentation for Top-Down DetectionabstractIn this paper we are interested in how semantic segmentation can help object detection. Towards this goal, we propose a novel deformable part-based model which exploits region-based segmentation algorithms that compute candidate object regions by bottom-up clustering followed by ranking of those regions. Our approach allows every detection hypothesis to select a segment (including void), and scores each box in the image using both the traditional HOG filters as well as a set of novel segmentation features. Thus our model ``blends'' between the detector and segmentation models. Since our features can be computed very efficiently given the segments, we maintain the same complexity as the original DPM. We demonstrate the effectiveness of our approach in PASCAL VOC 2010, and show that when employing only a root filter our approach outperforms Dalal & Triggs detector on all classes, achieving 13% higher average AP. When employing the parts, we outperform the original DPM in $19$ out of $20$ classes, achieving an improvement of 8% AP. Furthermore, we outperform the previous state-of-the-art on VOC 2010 test by 4%. Sanja Fidler, Roozbeh Mottaghi, Alan L. Yuille, Raquel Urtasun |
CVPR | 3 |
| 2013 | Boundary Detection Benchmarking: Beyond F-MeasuresabstractFor an ill-posed problem like boundary detection, human labeled datasets play a critical role. Compared with the active research on finding a better boundary detector to refresh the performance record, there is surprisingly little discussion on the boundary detection benchmark itself. The goal of this paper is to identify the potential pitfalls of today's most popular boundary benchmark, BSDS 300. In the paper, we first introduce a psychophysical experiment to show that many of the "weak" boundary labels are unreliable and may contaminate the benchmark. Then we analyze the computation of f-measure and point out that the current benchmarking protocol encourages an algorithm to bias towards those problematic "weak" boundary labels. With this evidence, we focus on a new problem of detecting strong boundaries as one alternative. Finally, we assess the performances of 9 major algorithms on different ways of utilizing the dataset, suggesting new directions for improvements. Alan L. Yuille, Christof Koch |
CVPR | 2 |
| 2013 | Robust Region Grouping via Internal Patch StatisticsabstractIn this work, we present an efficient multi-scale low-rank representation for image segmentation. Our method begins with partitioning the input images into a set of super pixels, followed by seeking the optimal super pixel-pair affinity matrix, both of which are performed at multiple scales of the input images. Since low-level super pixel features are usually corrupted by image noises, we propose to infer the low-rank refined affinity matrix. The inference is guided by two observations on natural images. First, looking into a single image, local small-size image patterns tend to recur frequently within the same semantic region, but may not appear in semantically different regions. We call this internal image statistics as replication prior, and quantitatively justify it on real image databases. Second, the affinity matrices at different scales should be consistently solved, which leads to the cross-scale consistency constraint. We formulate these two purposes with one unified formulation and develop an efficient optimization procedure. Our experiments demonstrate the presented method can substantially improve segmentation accuracy. Xiaobai Liu, Liang Lin 0004, Alan L. Yuille |
CVPR | 3 |
| 2013 | Robust Estimation of Nonrigid Transformation for Point Set RegistrationabstractWe present a new point matching algorithm for robust nonrigid registration. The method iteratively recovers the point correspondence and estimates the transformation between two point sets. In the first step of the iteration, feature descriptors such as shape context are used to establish rough correspondence. In the second step, we estimate the transformation using a robust estimator called L_2E. This is the main novelty of our approach and it enables us to deal with the noise and outliers which arise in the correspondence step. The transformation is specified in a functional space, more specifically a reproducing kernel Hilbert space. We apply our method to nonrigid sparse image feature correspondence on 2D images and 3D surfaces. Our results quantitatively show that our approach outperforms state-of-the-art methods, particularly when there are a large number of outliers. Moreover, our method of robustly estimating transformations from correspondences is general and has many other applications. Jiayi Ma 0001, Ji Zhao 0001, Jinwen Tian, Zhuowen Tu, Alan L. Yuille |
CVPR | 5 |
| 2013 | An Approach to Pose-Based Action RecognitionabstractWe address action recognition in videos by modeling the spatial-temporal structures of human poses. We start by improving a state of the art method for estimating human joint locations from videos. More precisely, we obtain the K-best estimations output by the existing method and incorporate additional segmentation cues and temporal constraints to select the ``best'' one. Then we group the estimated joints into five body parts (e.g. the left arm) and apply data mining techniques to obtain a representation for the spatial-temporal structures of human actions. This representation captures the spatial configurations of body parts in one frame (by spatial-part-sets) as well as the body part movements(by temporal-part-sets) which are characteristic of human actions. It is interpretable, compact, and also robust to errors on joint estimations. Experimental results first show that our approach is able to localize body joints more accurately than existing methods. Next we show that it outperforms state of the art action recognizers on the UCF sport, the Keck Gesture and the MSR-Action3D datasets. Chunyu Wang 0001, Yizhou Wang 0001, Alan L. Yuille |
CVPR | 3 |
| 2013 | Learning a Dictionary of Shape Epitomes with Applications to Image LabelingabstractThe first main contribution of this paper is a novel method for representing images based on a dictionary of shape epitomes. These shape epitomes represent the local edge structure of the image and include hidden variables to encode shift and rotations. They are learnt in an unsupervised manner from groundtruth edges. This dictionary is compact but is also able to capture the typical shapes of edges in natural images. In this paper, we illustrate the shape epitomes by applying them to the image labeling task. In other work, described in the supplementary material, we apply them to edge detection and image modeling. We apply shape epitomes to image labeling by using Conditional Random Field (CRF) Models. They are alternatives to the superpixel or pixel representations used in most CRFs. In our approach, the shape of an image patch is encoded by a shape epitome from the dictionary. Unlike the superpixel representation, our method avoids making early decisions which cannot be reversed. Our resulting hierarchical CRFs efficiently capture both local and global class co-occurrence properties. We demonstrate its quantitative and qualitative properties of our approach with image labeling experiments on two standard datasets: MSRC-21 and Stanford Background. Liang-Chieh Chen, George Papandreou, Alan L. Yuille |
ICCV | 3 |
| 2013 | Adaptive occlusion state estimation for human pose tracking under self-occlusions
Nam-Gyu Cho, Alan L. Yuille, Seong-Whan Lee |
Pattern Recognit. | 2 |
| 2012 | Self-occlusion robust 3D human pose tracking from monocular image sequenceabstractPose tracking technique has great potential for many applications such as marker-free human motion capture system, Human Computer Interactions (HCI), and video surveillance. Though many methods are introduced during last decades, self-occlusion - one body part is occluded by another one - is still considered one of the most difficult problems for 3D human pose tracking. In this paper, we propose a self-occlusion state estimation method. A MRF (Markov Random Field) is used to model the occlusion state which represents the pairwise depth order between two human body parts. A novel estimation method is proposed to infer a body pose and an occlusion state separately. HumanEva dataset is used for testing the proposed method. In order to evaluate and quantify how often the occlusion state changes, we label the ground truth of occlusion state. Nam-Gyu Cho, Alan L. Yuille, Seong-Whan Lee |
SMC | 2 |
| 2012 | Computer vision needs a core and foundationsabstractI argue that computer vision needs a core of techniques and foundational research to enable it to build on its current successes and achieve its enormous potential. “How do I know what papers to read in computer vision? There are so many. And they are so different.” Graduate Student. Xi'An. China. November, 2011. Alan L. Yuille |
Image Vis. Comput. | 1 |
| 2012 | Recursive Segmentation and Recognition Templates for Image ParsingabstractIn this paper, we propose a Hierarchical Image Model (HIM) which parses images to perform segmentation and object recognition. The HIM represents the image recursively by segmentation and recognition templates at multiple levels of the hierarchy. This has advantages for representation, inference, and learning. First, the HIM has a coarse-to-fine representation which is capable of capturing long-range dependency and exploiting different levels of contextual information (similar to how natural language models represent sentence structure in terms of hierarchical representations such as verb and noun phrases). Second, the structure of the HIM allows us to design a rapid inference algorithm, based on dynamic programming, which yields the first polynomial time algorithm for image labeling. Third, we learn the HIM efficiently using machine learning methods from a labeled data set. We demonstrate that the HIM is comparable with the state-of-the-art methods by evaluation on the challenging public MSRC and PASCAL VOC 2007 image data sets. Long Zhu, Yuanhao Chen, Alan L. Yuille |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2011 | Perturb-and-MAP random fields: Using discrete optimization to learn and sample from energy modelsabstractWe propose a novel way to induce a random field from an energy function on discrete labels. It amounts to locally injecting noise to the energy potentials, followed by finding the global minimum of the perturbed energy function. The resulting Perturb-and-MAP random fields harness the power of modern discrete energy minimization algorithms, effectively transforming them into efficient random sampling algorithms, thus extending their scope beyond the usual deterministic setting. In this fashion we can enjoy the benefits of a sound probabilistic framework, such as the ability to represent the solution uncertainty or learn model parameters from training data, while completely bypassing costly Markov-chain Monte-Carlo procedures typically associated with discrete label Gibbs Markov random fields (MRFs). We study some interesting theoretical properties of the proposed model in juxtaposition to those of Gibbs MRFs and address the issue of principled design of the perturbation process. We present experimental results in image segmentation and scene labeling that illustrate the new qualitative aspects and the potential of the proposed model for practical computer vision applications. George Papandreou, Alan L. Yuille |
ICCV | 2 |
| 2011 | AdaBoost for Text Detection in Natural SceneabstractDetecting text regions in natural scenes is an important part of computer vision. We propose a novel text detection algorithm that extracts six different classes features of text, and uses Modest AdaBoost with multi-scale sequential search. Experiments show that our algorithm can detect text regions with a f= 0.70, from the ICDAR 2003 datasets which include images with text of various fonts, sizes, colors, alphabets and scripts. Jung-Jin Lee, Pyoung-Hean Lee, Seong-Whan Lee, Alan L. Yuille, Christof Koch |
ICDAR | 4 |
| 2011 | Inference and Learning with Hierarchical Shape ModelsabstractIn this work we introduce a hierarchical representation for object detection. We represent an object in terms of parts composed of contours corresponding to object boundaries and symmetry axes; these are in turn related to edge and ridge features that are extracted from the image. We propose a coarse-to-fine algorithm for efficient detection which exploits the hierarchical nature of the model. This provides a tractable framework to combine bottom-up and top-down computation. We learn our models from training images where only the bounding box of the object is provided. We automate the decomposition of an object category into parts and contours, and discriminatively learn the cost function that drives the matching of the object to the image using Multiple Instance Learning. Using shape-based information, we obtain state-of-the-art localization results on the UIUC and ETHZ datasets. Iasonas Kokkinos, Alan L. Yuille |
Int. J. Comput. Vis. | 2 |
| 2011 | Max Margin Learning of Hierarchical Configural Deformable Templates (HCDTs) for Efficient Object Parsing and Pose EstimationabstractIn this paper we formulate a hierarchical configurable deformable template (HCDT) to model articulated visual objects—such as horses and baseball players—for tasks such as parsing, segmentation, and pose estimation. HCDTs represent an object by an AND/OR graph where the OR nodes act as switches which enables the graph topology to vary adaptively. This hierarchical representation is compositional and the node variables represent positions and properties of subparts of the object. The graph and the node variables are required to obey the summarization principle which enables an efficient compositional inference algorithm to rapidly estimate the state of the HCDT. We specify the structure of the AND/OR graph of the HCDT by hand and learn the model parameters discriminatively by extending Max-Margin learning to AND/OR graphs. We illustrate the three main aspects of HCDTs—representation, inference, and learning—on the tasks of segmenting, parsing, and pose (configuration) estimation for horses and humans. We demonstrate that the inference algorithm is fast and that max-margin learning is effective. We show that HCDTs gives state of the art results for segmentation and pose estimation when compared to other methods on benchmarked datasets. Long Zhu, Yuanhao Chen, Alan L. Yuille |
Int. J. Comput. Vis. | 4 |
| 2010 | Part and appearance sharing: Recursive Compositional Models for multi-viewabstractWe propose Recursive Compositional Models (RCMs) for simultaneous multi-view multi-object detection and parsing (e.g. view estimation and determining the positions of the object subparts). We represent the set of objects by a family of RCMs where each RCM is a probability distribution defined over a hierarchical graph which corresponds to a specific object and viewpoint. An RCM is constructed from a hierarchy of subparts/subgraphs which are learnt from training data. Part-sharing is used so that different RCMs are encouraged to share subparts/subgraphs which yields a compact representation for the set of objects and which enables efficient inference and learning from a limited number of training samples. In addition, we use appearance-sharing so that RCMs for the same object, but different viewpoints, share similar appearance cues which also helps efficient learning. RCMs lead to a multi-view multi-object detection system. We illustrate RCMs on four public datasets and achieve state-of-the-art performance. Long Zhu, Yuanhao Chen, Antonio Torralba 0001, William T. Freeman, Alan L. Yuille |
CVPR | 5 |
| 2010 | Latent hierarchical structural learning for object detectionabstractWe present a latent hierarchical structural learning method for object detection. An object is represented by a mixture of hierarchical tree models where the nodes represent object parts. The nodes can move spatially to allow both local and global shape deformations. The models can be trained discriminatively using latent structural SVM learning, where the latent variables are the node positions and the mixture component. But current learning methods are slow, due to the large number of parameters and latent variables, and have been restricted to hierarchies with two layers. In this paper we describe an incremental concave-convex procedure (iCCCP) which allows us to learn both two and three layer models efficiently. We show that iCCCP leads to a simple training algorithm which avoids complex multi-stage layer-wise training, careful part selection, and achieves good performance without requiring elaborate initialization. We perform object detection using our learnt models and obtain performance comparable with state-of-the-art methods when evaluated on challenging public PASCAL datasets. We demonstrate the advantages of three layer hierarchies - outperforming Felzenszwalb et al.'s two layer models on all 20 classes. Long Zhu, Yuanhao Chen, Alan L. Yuille, William T. Freeman |
CVPR | 3 |
| 2010 | Active Mask Hierarchies for Object Detection
Yuanhao Chen, Long Zhu, Alan L. Yuille |
ECCV (5) | 3 |
| 2010 | Occlusion Boundary Detection Using Pseudo-depth
Xuming He 0001, Alan L. Yuille |
ECCV (4) | 2 |
| 2010 | Functional form of motion priors in human motion perceptionabstractIt has been speculated that the human motion system combines noisy measurements with prior expectations in an optimal, or rational, manner. The basic goal of our work is to discover experimentally which prior distribution is used. More specifically, we seek to infer the functional form of the motion prior from the performance of human subjects on motion estimation tasks. We restricted ourselves to priors which combine three terms for motion slowness, first-order smoothness, and second-order smoothness. We focused on two functional forms for prior distributions: L2-norm and L1-norm regularization corresponding to the Gaussian and Laplace distributions respectively. In our first experimental session we estimate the weights of the three terms for each functional form to maximize the fit to human performance. We then measured human performance for motion tasks and found that we obtained better fit for the L1-norm (Laplace) than for the L2-norm (Gaussian). We note that the L1-norm is also a better fit to the statistics of motion in natural environments. In addition, we found large weights for the second-order smoothness term, indicating the importance of high-order smoothness compared to slowness and lower-order smoothness. To validate our results further, we used the best fit models using the L1-norm to predict human performance in a second session with different experimental setups. Our results showed excellent agreement between human performance and model prediction -- ranging from 3\% to 8\% for five human subjects over ten experimental conditions -- and give further support that the human visual system uses an L1-norm (Laplace) prior. Hongjing Lu, Tungyou Lin, Alan L. F. Lee, Luminita A. Vese, Alan L. Yuille |
NIPS | 5 |
| 2010 | Gaussian sampling by local perturbationsabstractWe present a technique for exact simulation of Gaussian Markov random fields (GMRFs), which can be interpreted as locally injecting noise to each Gaussian factor independently, followed by computing the mean/mode of the perturbed GMRF. Coupled with standard iterative techniques for the solution of symmetric positive definite systems, this yields a very efficient sampling algorithm with essentially linear complexity in terms of speed and memory requirements, well suited to extremely large scale probabilistic models. Apart from synthesizing data under a Gaussian model, the proposed technique directly leads to an efficient unbiased estimator of marginal variances. Beyond Gaussian models, the proposed algorithm is also very useful for handling highly non-Gaussian continuously-valued MRFs such as those arising in statistical image modeling or in the first layer of deep belief networks describing real-valued data, where the non-quadratic potentials coupling different sites can be represented as finite or infinite mixtures of Gaussians with the help of local or distributed latent mixture assignment variables. The Bayesian treatment of such models most naturally involves a block Gibbs sampler which alternately draws samples of the conditionally independent latent mixture assignments and the conditionally multivariate Gaussian continuous vector and we show that it can directly benefit from the proposed methods. George Papandreou, Alan L. Yuille |
NIPS | 2 |
| 2010 | A unified model of short-range and long-range motion perceptionabstractThe human vision system is able to effortlessly perceive both short-range and long-range motion patterns in complex dynamic scenes. Previous work has assumed that two different mechanisms are involved in processing these two types of motion. In this paper, we propose a hierarchical model as a unified framework for modeling both short-range and long-range motion perception. Our model consists of two key components: a data likelihood that proposes multiple motion hypotheses using nonlinear matching, and a hierarchical prior that imposes slowness and spatial smoothness constraints on the motion field at multiple scales. We tested our model on two types of stimuli, random dot kinematograms and multiple-aperture stimuli, both commonly used in human vision research. We demonstrate that the hierarchical model adequately accounts for human performance in psychophysical experiments. Xuming He 0001, Hongjing Lu, Alan L. Yuille |
NIPS | 4 |
| 2010 | Detecting object boundaries using low-, mid-, and high-level information
Songfeng Zheng, Alan L. Yuille, Zhuowen Tu |
Comput. Vis. Image Underst. | 2 |
| 2010 | Learning a Hierarchical Deformable Template for Rapid Deformable Object ParsingabstractIn this paper, we address the tasks of detecting, segmenting, parsing, and matching deformable objects. We use a novel probabilistic object model that we call a hierarchical deformable template (HDT). The HDT represents the object by state variables defined over a hierarchy (with typically five levels). The hierarchy is built recursively by composing elementary structures to form more complex structures. A probability distribution--a parameterized exponential model--is defined over the hierarchy to quantify the variability in shape and appearance of the object at multiple scales. To perform inference--to estimate the most probable states of the hierarchy for an input image--we use a bottom-up algorithm called compositional inference. This algorithm is an approximate version of dynamic programming where approximations are made (e.g., pruning) to ensure that the algorithm is fast while maintaining high performance. We adapt the structure-perceptron algorithm to estimate the parameters of the HDT in a discriminative manner (simultaneously estimating the appearance and shape parameters). More precisely, we specify an exponential distribution for the HDT using a dictionary of potentials, which capture the appearance and shape cues. This dictionary can be large and so does not require handcrafting the potentials. Instead, structure-perceptron assigns weights to the potentials so that less important potentials receive small weights (this is like a "soft" form of feature selection). Finally, we provide experimental evaluation of HDTs on different visual tasks, including detection, segmentation, matching (alignment), and parsing. We show that HDTs achieve state-of-the-art performance for these different tasks when evaluated on data sets with groundtruth (and when compared to alternative algorithms, which are typically specialized to each task). Long Zhu, Yuanhao Chen, Alan L. Yuille |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2009 | HOP: Hierarchical object parsingabstractIn this paper we consider the problem of object parsing, namely detecting an object and its components by composing them from image observations. Apart from object localization, this involves the question of combining top-down (model-based) with bottom-up (image-based) information. We use an hierarchical object model, that recursively decomposes an object into simple structures. Our first contribution is the formulation of composition rules to build the object structures, while addressing problems such as contour fragmentation and missing parts. Our second contribution is an efficient inference method for object parsing that addresses the combinatorial complexity of the problem. For this we exploit our hierarchical object representation to efficiently compute a coarse solution to the problem, which we then use to guide search at a finer level. This rules out a large portion of futile compositions and allows us to parse complex objects in heavily cluttered scenes. Iasonas Kokkinos, Alan L. Yuille |
CVPR | 2 |
| 2009 | Compositional noisy-logical learningabstractWe describe a new method for learning the conditional probability distribution of a binary-valued variable from labelled training examples. Our proposed Compositional Noisy-Logical Learning (CNLL) approach learns a noisy-logical distribution in a compositional manner. CNLL is an alternative to the well-known AdaBoost algorithm which performs coordinate descent on an alternative error measure. We describe two CNLL algorithms and test their performance compared to AdaBoost on two types of problem: (i) noisy-logical data (such as noisy exclusive-or), and (ii) four standard datasets from the UCI repository. Our results show that we outperform AdaBoost while using significantly fewer weak classifiers, thereby giving a more transparent classifier suitable for knowledge extraction. Alan L. Yuille, Songfeng Zheng |
ICML | 1 |
| 2009 | Modeling the spacing effect in sequential category learningabstractWe develop a Bayesian sequential model for category learning. The sequential model updates two category parameters, the mean and the variance, over time. We define conjugate temporal priors to enable closed form solutions to be obtained. This model can be easily extended to supervised and unsupervised learning involving multiple categories. To model the spacing effect, we introduce a generic prior in the temporal updating stage to capture a learning preference, namely, less change for repetition and more change for variation. Finally, we show how this approach can be generalized to efficiently performmodel selection to decide whether observations are from one or multiple categories. Hongjing Lu, Matthew Weiden, Alan L. Yuille |
NIPS | 3 |
| 2009 | Unsupervised Learning of Probabilistic Object Models (POMs) for Object Classification, Segmentation, and Recognition Using Knowledge PropagationabstractWe present a method to learn probabilistic object models (POMs) with minimal supervision, which exploit different visual cues and perform tasks such as classification, segmentation, and recognition. We formulate this as a structure induction and learning task and our strategy is to learn and combine elementary POMs that make use of complementary image cues. We describe a novel structure induction procedure, which uses knowledge propagation to enable POMs to provide information to other POMs and "teach them" (which greatly reduces the amount of supervision required for training and speeds up the inference). In particular, we learn a POM-IP defined on Interest Points using weak supervision [1], [2] and use this to train a POM-mask, defined on regional features, which yields a combined POM that performs segmentation/localization. This combined model can be used to train POM-edgelets, defined on edgelets, which gives a full POM with improved performance on classification. We give detailed experimental analysis on large data sets for classification and segmentation with comparison to other methods. Inference takes five seconds while learning takes approximately four hours. In addition, we show that the full POM is invariant to scale and rotation of the object (for learning and inference) and can learn hybrid objects classes (i.e., when there are several objects and the identity of the object in each image is unknown). Finally, we show that POMs can be used to match between different objects of the same category, and hence, enable objects recognition. Yuanhao Chen, Long Zhu, Alan L. Yuille, HongJiang Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2009 | Unsupervised Learning of Probabilistic Grammar-Markov Models for Object CategoriesabstractWe introduce a Probabilistic Grammar-Markov Model (PGMM) which couples probabilistic context free grammars and Markov Random Fields. These PGMMs are generative models defined over attributed features and are used to detect and classify objects in natural images. PGMMs are designed so that they can perform rapid inference, parameter learning, and the more difficult task of structure induction. PGMMs can deal with unknown 2D pose (position, orientation, and scale) in both inference and learning, different appearances, or aspects, of the model. The PGMMs can be learnt in an unsupervised manner where the image can contain one of an unknown number of objects of different categories or even be pure background. We first study the weakly supervised case, where each image contains an example of the (single) object of interest, and then generalize to less supervised cases. The goal of this paper is theoretical but, to provide proof of concept, we demonstrate results from this approach on a subset of the Caltech dataset (learning on a training set and evaluating on a testing set). Our results are generally comparable with the current state of the art, and our inference is performed in less than five seconds. Long Zhu, Yuanhao Chen, Alan L. Yuille |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2008 | Unsupervised learning of probabilistic object models (POMs) for object classification, segmentation and recognitionabstractWe present a new unsupervised method to learn unified probabilistic object models (POMs) which can be applied to classification, segmentation, and recognition. We formulate this as a structure learning task and our strategy is to learn and combine basic POM's that make use of complementary image cues. Each POM has algorithms for inference and parameter learning, but: (i) the structure of each POM is unknown, and (ii) the inference and parameter learning algorithm for a POM may be impractical without additional information. We address these problems by a novel structure induction procedure which uses knowledge propagation to enable POM's to provide information to other POM's and "teach them" (which greatly reduced the amount of supervision required for training). In particular, we learn a POM-IP defined on interest points using weak supervision [1, 2] and use this to train a POM- mask, defined on regional features, which yields a combined POM which performs segmentation/localization. This combined model can be used to train POM-edgelets, defined on edgelets, which gives a full POM with improved performance on classification. We give detailed experimental analysis on large datasets which show that the full POM is invariant to scale and rotation of the object (for learning and inference) and performs inference rapidly. In addition, we show that we can apply POM's to learn objects classes (i.e. when there are several objects and the identity of the object in each image is unknown). We emphasize that these models can match between different objects from the same category and hence enable object recognition. Yuanhao Chen, Long Zhu, Alan L. Yuille, HongJiang Zhang |
CVPR | 3 |
| 2008 | Graph-shifts: Natural image labeling by dynamic hierarchical computingabstractIn this paper, we present a new approach for image labeling based on the recently introduced graph-shifts algorithm. Graph-shifts is an energy minimization algorithm that does labeling by dynamically manipulating, or shifting, the parent-child relationships in a hierarchical decomposition of the image. Each shift optimally reduces the energy by indirectly causing a change to the labeling; graph-shifts is able to rapidly compute and select this optimal shift at every iteration. There are no constraints on the terms of the (pairwise) energy function. The algorithm was originally presented in the context of medical image labeling using conditional random field models. In this paper, we consider the algorithm in the context of both low- and high-level natural image labeling. We show that for examples in both classes of problems, graph-shifts does labeling both accurately and rapidly. For low-level vision, we explore image restoration, and for high-level vision, we make use of a hybrid discriminative-generative model to segment and label images into semantically meaningful regions (e.g., trees, buildings, etc.). For both problems, we obtain comparable or superior results to the state-of-the-art computed in just a few seconds per image. Jason J. Corso, Alan L. Yuille, Zhuowen Tu |
CVPR | 2 |
| 2008 | Scale invariance without scale selectionabstractIn this work we construct scale invariant descriptors (SIDs) without requiring the estimation of image scale; we thereby avoid scale selection which is often unreliable. Our starting point is a combination of log-polar sampling and spatially-varying smoothing that converts image scalings and rotations into translations. Scale invariance can then be guaranteed by estimating the Fourier transform modulus (FTM) of the formed signal as the FTM is translation invariant. We build our descriptors using phase, orientation and amplitude features that compactly capture the local image structure. Our results show that the constructed SIDs outperform state-of-the-art descriptors on standard datasets. A main advantage of SIDs is that they are applicable to a broader range of image structures, such as edges, for which scale selection is unreliable. We demonstrate this by combining SIDs with contour segments and show that the performance of a boundary-based model is systematically improved on an object detection task. Iasonas Kokkinos, Alan L. Yuille |
CVPR | 2 |
| 2008 | Max Margin AND/OR Graph learning for parsing the human bodyabstractWe present a novel structure learning method, Max Margin AND/OR graph (MM-AOG), for parsing the human body into parts and recovering their poses. Our method represents the human body and its parts by an AND/OR graph, which is a multi-level mixture of Markov random fields (MRFs). Max-margin learning, which is a generalization of the training algorithm for support vector machines (SVMs), is used to learn the parameters of the AND/OR graph model discriminatively. There are four advantages from this combination of AND/OR graphs and max-margin learning. Firstly, the AND/OR graph allows us to handle enormous articulated poses with a compact graphical model. Secondly, max-margin learning has more discriminative power than the traditional maximum likelihood approach. Thirdly, the parameters of the AND/OR graph model are optimized globally. In particular, the weights of the appearance model for individual nodes and the relative importance of spatial relationships between nodes are learnt simultaneously. Finally, the kernel trick can be used to handle high dimensional features and to enable complex similarity measure of shapes. We perform comparison experiments on the base ball datasets, showing significant improvements over state of the art methods. Long Zhu, Yuanhao Chen, Alan L. Yuille |
CVPR | 5 |
| 2008 | Structure-perceptron learning of a hierarchical log-linear modelabstractIn this paper, we address the problems of deformable object matching (alignment) and segmentation with cluttered background. We propose a novel hierarchical log-linear model (HLLM) which represents both shape and appearance features at multiple levels of a hierarchy. This model enables us to combine appearance cues at multiple scales directly into the hierarchy and to model shape deformations at short-range, medium range, and long-range. We introduce the structure-perceptron algorithm to estimate the parameters of the HLLM in a discriminative way. The learning is able to estimate the appearance and shape parameters simultaneously in a global manner. Moreover, the structure-perceptron learning has a feature selection aspect (similar to AdaBoost) which enables us to specify a class of appearance/shape features and allow the algorithm to select which features to use and weight their importance. This method was applied to the tasks of deformable object localization, segmentation, matching (alignment), and parsing. We demonstrate that the algorithm achieves the state of the art performance by evaluation on public dataset (horse and multi-view face). Long Zhu, Yuanhao Chen, Xingyao Ye, Alan L. Yuille |
CVPR | 4 |
| 2008 | Unsupervised Structure Learning: Hierarchical Recursive Composition, Suspicious Coincidence and Competitive Exclusion
Long Zhu, Haoda Huang, Yuanhao Chen, Alan L. Yuille |
ECCV (2) | 5 |
| 2008 | MRF Labeling with a Graph-Shifts Algorithm
Jason J. Corso, Zhuowen Tu, Alan L. Yuille |
IWCIA | 3 |
| 2008 | Model selection and velocity estimation using novel priors for motion patternsabstractPsychophysical experiments show that humans are better at perceiving rotation and expansion than translation. These findings are inconsistent with standard models of motion integration which predict best performance for translation [6]. To explain this discrepancy, our theory formulates motion perception at two levels of inference: we first perform model selection between the competing models (e.g. translation, rotation, and expansion) and then estimate the velocity using the selected model. We define novel prior models for smooth rotation and expansion using techniques similar to those in the slow-and-smooth model [17] (e.g. Green functions of differential operators). The theory gives good agreement with the trends observed in human experiments. Hongjing Lu, Alan L. Yuille |
NIPS | 3 |
| 2008 | Recursive Segmentation and Recognition Templates for 2D ParsingabstractLanguage and image understanding are two major goals of artificial intelligence which can both be conceptually formulated in terms of parsing the input signal into a hierarchical representation. Natural language researchers have made great progress by exploiting the 1D structure of language to design efficient polynomial- time parsing algorithms. By contrast, the two-dimensional nature of images makes it much harder to design efficient image parsers and the form of the hierarchical representations is also unclear. Attempts to adapt representations and algorithms from natural language have only been partially successful. In this paper, we propose a Hierarchical Image Model (HIM) for 2D image pars- ing which outputs image segmentation and object recognition. This HIM is rep- resented by recursive segmentation and recognition templates in multiple layers and has advantages for representation, inference, and learning. Firstly, the HIM has a coarse-to-fine representation which is capable of capturing long-range de- pendency and exploiting different levels of contextual information. Secondly, the structure of the HIM allows us to design a rapid inference algorithm, based on dy- namic programming, which enables us to parse the image rapidly in polynomial time. Thirdly, we can learn the HIM efficiently in a discriminative manner from a labeled dataset. We demonstrate that HIM outperforms other state-of-the-art methods by evaluation on the challenging public MSRC image dataset. Finally, we sketch how the HIM architecture can be extended to model more complex image phenomena. Long Zhu, Yuanhao Chen, Alan L. Yuille |
NIPS | 5 |
| 2008 | Shape matching and registration by data-driven EM
Zhuowen Tu, Songfeng Zheng, Alan L. Yuille |
Comput. Vis. Image Underst. | 3 |
| 2008 | Efficient Multilevel Brain Tumor Segmentation With Integrated Bayesian Model ClassificationabstractWe present a new method for automatic segmentation of heterogeneous image data that takes a step toward bridging the gap between bottom-up affinity-based segmentation methods and top-down generative model based approaches. The main contribution of the paper is a Bayesian formulation for incorporating soft model assignments into the calculation of affinities, which are conventionally model free. We integrate the resulting model-aware affinities into the multilevel segmentation by weighted aggregation algorithm, and apply the technique to the task of detecting and segmenting brain tumor and edema in multichannel magnetic resonance (MR) volumes. The computationally efficient method runs orders of magnitude faster than current state-of-the-art techniques giving comparable or improved results. Our quantitative results indicate the benefit of incorporating model-aware affinities into the segmentation process for the difficult case of glioblastoma multiforme brain tumor. Jason J. Corso, Eitan Sharon, Shishir Dube, Suzie El-Saden, Usha S. Sinha, Alan L. Yuille |
IEEE Trans. Medical Imaging | 6 |
| 2007 | Detecting Object Boundaries Using Low-, Mid-, and High-level InformationabstractObject boundary detection and segmentation is a central problem in computer vision. The importance of combining low-level, mid-level, and high-level cues has been realized in recent literature. However, it is unclear how to efficiently and effectively engage and fuse different levels of information. In this paper, we emphasize a learning based approach to explore different levels of information, both implicitly and explicitly. First, we learn low-level cues for object boundaries and interior regions using a probabilistic boosting tree (PBT). Second, we learn short and long range context information based on the results from the first stage. Both stages implicitly contain object-specific information such as texture and local geometry, and it is shown that this implicit knowledge is extremely powerful. Third, we use high-level shape information explicitly to further refine the object segmentation and to parse the object into components. The algorithm is trained and tested on a challenging dataset of horses [2], and the results obtained are very encouraging compared with other approaches. In detailed experiments we show significantly better performance (e.g. F-values of 0.75 compared to 0.66) than the best comparable reported performance on this dataset. Furthermore, the system only needs 1.5 minutes for a typical image. Although our system is illustrated on horse images, the approach can be directly applied to detecting/segmenting other types of objects. Songfeng Zheng, Zhuowen Tu, Alan L. Yuille |
CVPR | 3 |
| 2007 | Unsupervised Learning of Object Deformation ModelsabstractThe aim of this work is to learn generative models of object deformations in an unsupervised manner. Initially, we introduce an Expectation Maximization approach to estimate a linear basis for deformations by maximizing the likelihood of the training set under an Active Appearance Model (AAM). This approach is shown to successfully capture the global shape variations of objects like faces, cars and hands. However the AAM representation cannot deal with articulated objects, like cows and horses. We therefore extend our approach to a representation that allows for multiple parts with the relationships between them modeled by a Markov Random Field (MRF). Finally, we propose an algorithm for efficiently performing inference on part-based MRF object models by speeding up the estimation of observation potentials. We use manually collected landmarks to compare the alternative models and quantify learning performance. Iasonas Kokkinos, Alan L. Yuille |
ICCV | 2 |
| 2007 | Detection and Segmentation of Pathological Structures by the Extended Graph-Shifts Algorithm
Jason J. Corso, Alan L. Yuille, Nancy L. Sicotte, Arthur W. Toga |
MICCAI (1) | 2 |
| 2007 | Rapid Inference on a Novel AND/OR graph for Object Detection, Segmentation and ParsingabstractIn this paper we formulate a novel AND/OR graph representation capable of describing the different configurations of deformable articulated objects such as horses. The representation makes use of the summarization principle so that lower level nodes in the graph only pass on summary statistics to the higher level nodes. The probability distributions are invariant to position, orientation, and scale. We develop a novel inference algorithm that combined a bottom-up process for proposing configurations for horses together with a top-down process for refining and validating these proposals. The strategy of surround suppression is applied to ensure that the inference time is polynomial in the size of input data. The algorithm was applied to the tasks of detecting, segmenting and parsing horses. We demonstrate that the algorithm is fast and comparable with the state of the art approaches. Yuanhao Chen, Long Zhu, Alan L. Yuille, HongJiang Zhang |
NIPS | 4 |
| 2007 | The Noisy-Logical Distribution and its Application to Causal InferenceabstractWe describe a novel noisy-logical distribution for representing the distribution of a binary output variable conditioned on multiple binary input variables. The distribution is represented in terms of noisy-or's and noisy-and-not's of causal features which are conjunctions of the binary inputs. The standard noisy-or and noisy-and-not models, used in causal reasoning and artificial intelligence, are special cases of the noisy-logical distribution. We prove that the noisy-logical distribution is complete in the sense that it can represent all conditional distributions provided a sufficient number of causal factors are used. We illustrate the noisy-logical distribution by showing that it can account for new experimental findings on how humans perform causal reasoning in more complex contexts. Finally, we speculate on the use of the noisy-logical distribution for causal reasoning and artificial intelligence. Alan L. Yuille, Hongjing Lu |
NIPS | 1 |
| 2007 | Automated Extraction of the Cortical Sulci Based on a Supervised Learning ApproachabstractIt is important to detect and extract the major cortical sulci from brain images, but manually annotating these sulci is a time-consuming task and requires the labeler to follow complex protocols. This paper proposes a learning-based algorithm for automated extraction of the major cortical sulci from magnetic resonance imaging (MRI) volumes and cortical surfaces. Unlike alternative methods for detecting the major cortical sulci, which use a small number of predefined rules based on properties of the cortical surface such as the mean curvature, our approach learns a discriminative model using the probabilistic boosting tree algorithm (PBT). PBT is a supervised learning approach which selects and combines hundreds of features at different scales, such as curvatures, gradients and shape index. Our method can be applied to either MRI volumes or cortical surfaces. It first outputs a probability map which indicates how likely each voxel lies on a major sulcal curve. Next, it applies dynamic programming to extract the best curve based on the probability map and a shape prior. The algorithm has almost no parameters to tune for extracting different major sulci. It is very fast (it runs in under 1 min per sulcus including the time to compute the discriminative models) due to efficient implementation of the features (e.g., using the integral volume to rapidly compute the responses of 3-D Haar filters). Because the algorithm can be applied to MRI volumes directly, there is no need to perform preprocessing such as tissue segmentation or mapping to a canonical space. The learning aspect of our approach makes the system very flexible and general. For illustration, we use volumes of the right hemisphere with several major cortical sulci manually labeled. The algorithm is tested on two groups of data, including some brains from patients with Williams Syndrome, and the results are very encouraging. Zhuowen Tu, Songfeng Zheng, Alan L. Yuille, Allan L. Reiss, Rebecca A. Dutton, Agatha D. Lee, Albert M. Galaburda, Ivo D. Dinov, Paul M. Thompson, Arthur W. Toga |
IEEE Trans. Medical Imaging | 3 |
| 2006 | Bottom-Up & Top-down Object Detection using Primal Sketch Features and Graphical ModelsabstractA combination of techniques that is becoming increasingly popular is the construction of part-based object representations using the outputs of interest-point detectors. Our contributions in this paper are twofold: first, we propose a primal-sketch-based set of image tokens that are used for object representation and detection. Second, top-down information is introduced based on an efficient method for the evaluation of the likelihood of hypothesized part locations. This allows us to use graphical model techniques to complement bottom-up detection, by proposing and finding the parts of the object that were missed by the front-end feature detection stage. Detection results for four object categories validate the merits of this joint top-down and bottom-up approach. Iasonas Kokkinos, Petros Maragos, Alan L. Yuille |
CVPR (2) | 3 |
| 2006 | Multilevel Segmentation and Integrated Bayesian Model Classification with an Application to Brain Tumor Segmentation
Jason J. Corso, Eitan Sharon, Alan L. Yuille |
MICCAI (2) | 3 |
| 2006 | A Learning Based Algorithm for Automatic Extraction of the Cortical Sulci
Songfeng Zheng, Zhuowen Tu, Alan L. Yuille, Allan L. Reiss, Rebecca A. Dutton, Agatha D. Lee, Albert M. Galaburda, Paul M. Thompson, Ivo D. Dinov, Arthur W. Toga |
MICCAI (1) | 3 |
| 2006 | Unsupervised Learning of a Probabilistic Grammar for Object Detection and ParsingabstractWe describe an unsupervised method for learning a probabilistic grammar of an object from a set of training examples. Our approach is invariant to the scale and rotation of the objects. We illustrate our approach using thirteen objects from the Caltech 101 database. In addition, we learn the model of a hybrid object class where we do not know the specific object or its position, scale or pose. This is illustrated by learning a hybrid class consisting of faces, motorbikes, and airplanes. The individual objects can be recovered as different aspects of the grammar for the object class. In all cases, we validate our results by learning the probability grammars from training datasets and evaluating them on the test datasets. We compare our method to alternative approaches. The advantages of our approach is the speed of inference (under one second), the parsing of the object, and increased accuracy of performance. Moreover, our approach is very general and can be applied to a large range of objects and structures. Long Zhu, Yuanhao Chen, Alan L. Yuille |
NIPS | 3 |
| 2005 | Ideal Observers for Detecting Motion: Correspondence NoiseabstractWe derive a Bayesian Ideal Observer (BIO) for detecting motion and solving the correspondence problem. We obtain Barlow and Tripathy’s classic model as an approximation. Our psychophysical experiments show that the trends of human performance are similar to the Bayesian Ideal, but overall human performance is far worse. We investigate ways to degrade the Bayesian Ideal but show that even extreme degradations do not approach human performance. Instead we propose that humans perform motion tasks using generic, general purpose, models of motion. We perform more psychophysical experiments which are consistent with humans using a Slow-and-Smooth model and which rule out an alterna- tive model using Slowness. Hongjing Lu, Alan L. Yuille |
NIPS | 2 |
| 2005 | Augmented Rescorla-Wagner and Maximum Likelihood EstimationabstractWe show that linear generalizations of Rescorla-Wagner can perform Maximum Likelihood estimation of the parameters of all generative models for causal reasoning. Our approach involves augmenting variables to deal with conjunctions of causes, similar to the agumented model of Rescorla. Our results involve genericity assumptions on the distributions of causes. If these assumptions are violated, for example for the Cheng causal power theory, then we show that a linear Rescorla-Wagner can estimate the parameters of the model up to a nonlinear transformtion. Moreover, a nonlinear Rescorla-Wagner is able to estimate the parameters directly to within arbitrary accuracy. Previous results can be used to determine convergence and to estimate convergence rates. Alan L. Yuille |
NIPS | 1 |
| 2005 | A Hierarchical Compositional System for Rapid Object DetectionabstractWe describe a hierarchical compositional system for detecting deformable objects in images. Objects are represented by graphical models. The algorithm uses a hierarchical tree where the root of the tree corresponds to the full object and lower-level elements of the tree correspond to simpler features. The algorithm proceeds by passing simple messages up and down the tree. The method works rapidly, in under a second, on 320 240 images. We demonstrate the approach on detecting cats, horses, and hands. The method works in the presence of background clutter and occlusions. Our approach is contrasted with more traditional methods such as dynamic programming and belief propagation. Long Zhu, Alan L. Yuille |
NIPS | 2 |
| 2005 | The DLR Hierarchy of Approximate Inference
Michal Rosen-Zvi, Michael I. Jordan, Alan L. Yuille |
UAI | 3 |
| 2005 | Image Parsing: Unifying Segmentation, Detection, and Recognition
Zhuowen Tu, Alan L. Yuille, Song-Chun Zhu |
Int. J. Comput. Vis. | 3 |
| 2004 | Motion Estimation by Swendsen-Wang Cuts
Adrian Barbu, Alan L. Yuille |
CVPR (1) | 2 |
| 2004 | Detecting and Reading Text in Natural Scenes
Alan L. Yuille |
CVPR (2) | 2 |
| 2004 | Shape Matching and Recognition - Using Generative Models and Informative Features
Zhuowen Tu, Alan L. Yuille |
ECCV (3) | 2 |
| 2004 | The Rescorla-Wagner Algorithm and Maximum Likelihood Estimation of Causal ParametersabstractThis paper analyzes generalization of the classic Rescorla-Wagner (R- W) learning algorithm and studies their relationship to Maximum Like- lihood estimation of causal parameters. We prove that the parameters of two popular causal models, P and P C, can be learnt by the same generalized linear Rescorla-Wagner (GLRW) algorithm provided gener- icity conditions apply. We characterize the fixed points of these GLRW algorithms and calculate the fluctuations about them, assuming that the input is a set of i.i.d. samples from a fixed (unknown) distribution. We describe how to determine convergence conditions and calculate conver- gence rates for the GLRW algorithms under these conditions. Alan L. Yuille |
NIPS | 1 |
| 2004 | The Convergence of Contrastive DivergencesabstractThis paper analyses the Contrastive Divergence algorithm for learning statistical parameters. We relate the algorithm to the stochastic approxi- mation literature. This enables us to specify conditions under which the algorithm is guaranteed to converge to the optimal solution (with proba- bility 1). This includes necessary and sufficient conditions for the solu- tion to be unbiased. Alan L. Yuille |
NIPS | 1 |
| 2003 | A Bayesian Network Framework for Relational Shape MatchingabstractA Bayesian network formulation for relational shape matching is presented. The main advantage of the relational shape matching approach is the obviation of the nonrigid spatial mappings used by recent nonrigid matching approaches. The basic variables that need to be estimated in the relational shape matching objective function are the global rotation and scale and the local displacements and correspondences. The new Bethe free energy approach is used to estimate the pairwise correspondences between links of the template graphs and the data. The resulting framework is useful in both registration and recognition contexts. Results are shown on hand-drawn templates and on 2D transverse T1-weighted MR images. Anand Rangarajan 0001, James M. Coughlan, Alan L. Yuille |
ICCV | 3 |
| 2003 | Image Parsing: Unifying Segmentation, Detection, and Recognition
Zhuowen Tu, Alan L. Yuille, Song-Chun Zhu |
ICCV | 3 |
| 2003 | Human and Ideal Observers for Detecting Image CurvesabstractThis paper compares the ability of human observers to detect target im- age curves with that of an ideal observer. The target curves are sam- pled from a generative model which specifies (probabilistically) the ge- ometry and local intensity properties of the curve. The ideal observer performs Bayesian inference on the generative model using MAP esti- mation. Varying the probability model for the curve geometry enables us investigate whether human performance is best for target curves that obey specific shape statistics, in particular those observed on natural shapes. Experiments are performed with data on both rectangular and hexagonal lattices. Our results show that human observers’ performance approaches that of the ideal observer and are, in general, closest to the ideal for con- ditions where the target curve tends to be straight or similar to natural statistics on curves. This suggests a bias of human observers towards straight curves and natural statistics. Alan L. Yuille, Fang Fang 0003, Paul Schrater, Daniel J. Kersten |
NIPS | 1 |
| 2003 | Algorithms from statistical physics for generative models of images
James M. Coughlan, Alan L. Yuille |
Image Vis. Comput. | 2 |
| 2003 | A statistical approach to multi-scale edge detection
Scott Konishi, Alan L. Yuille, James M. Coughlan |
Image Vis. Comput. | 2 |
| 2003 | Manhattan World: Orientation and Outlier Detection by Bayesian InferenceabstractThis letter argues that many visual scenes are based on a "Manhattan" three-dimensional grid that imposes regularities on the image statistics. We construct a Bayesian model that implements this assumption and estimates the viewer orientation relative to the Manhattan grid. For many images, these estimates are good approximations to the viewer orientation (as estimated manually by the authors). These estimates also make it easy to detect outlier structures that are unaligned to the grid. To determine the applicability of the Manhattan world model, we implement a null hypothesis model that assumes that the image statistics are independent of any three-dimensional scene structure. We then use the log-likelihood ratio test to determine whether an image satisfies the Manhattan world assumption. Our results show that if an image is estimated to be Manhattan, then the Bayesian model's estimates of viewer direction are almost always accurate (according to our manual estimates), and vice versa. James M. Coughlan, Alan L. Yuille |
Neural Comput. | 2 |
| 2003 | The Concave-Convex ProcedureabstractThe concave-convex procedure (CCCP) is a way to construct discrete-time iterative dynamical systems that are guaranteed to decrease global optimization and energy functions monotonically. This procedure can be applied to almost any optimization problem, and many existing algorithms can be interpreted in terms of it. In particular, we prove that all expectation-maximization algorithms and classes of Legendre minimization and variational bounding algorithms can be reexpressed in terms of CCCP. We show that many existing neural network and mean-field theory algorithms are also examples of CCCP. The generalized iterative scaling algorithm and Sinkhorn's algorithm can also be expressed as CCCP by changing variables. CCCP can be used both as a new way to understand, and prove the convergence of, existing optimization algorithms and as a procedure for generating new algorithms. Alan L. Yuille, Anand Rangarajan 0001 |
Neural Comput. | 1 |
| 2003 | Statistical Edge Detection: Learning and Evaluating Edge CuesabstractWe formulate edge detection as statistical inference. This statistical edge detection is data driven, unlike standard methods for edge detection which are model based. For any set of edge detection filters (implementing local edge cues), we use presegmented images to learn the probability distributions of filter responses conditioned on whether they are evaluated on or off an edge. Edge detection is formulated as a discrimination task specified by a likelihood ratio test on the filter responses. This approach emphasizes the necessity of modeling the image background (the off-edges). We represent the conditional probability distributions nonparametrically and illustrate them on two different data sets of 100 (Sowerby) and 50 (South Florida) images. Multiple edges cues, including chrominance and multiple-scale, are combined by using their joint distributions. Hence, this cue combination is optimal in the statistical sense. We evaluate the effectiveness of different visual cues using the Chernoff information and Receiver Operator Characteristic (ROC) curves. This shows that our approach gives quantitatively better results than the Canny edge detector when the image background contains significant clutter. In addition, it enables us to determine the effectiveness of different edge cues and gives quantitative measures for the advantages of multilevel processing, for the use of chrominance, and for the relative effectiveness of different detectors. Furthermore, we show that we can learn these conditional distributions on one data set and adapt them to the other with only slight degradation of performance without knowing the ground truth on the second data set. This shows that our results are not purely domain specific. We apply the same approach to the spatial grouping of edge cues and obtain analogies to nonmaximal suppression and hysteresis. Scott Konishi, Alan L. Yuille, James M. Coughlan, Song-Chun Zhu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2003 | The Generic Viewpoint Assumption and Planar BiasabstractWe show that generic viewpoint and lighting assumptions resolve standard visual ambiguities by biasing toward planar surfaces. Our model uses orthographic projection with a two-dimensional affine warp and Lambertian reflectance functions, including cast and attached shadows. We use uniform priors on nuisance variables such as viewpoint direction and the light source. Limitations of using uniform priors on nuisance variables are discussed. Alan L. Yuille, James M. Coughlan, Scott Konishi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2002 | Bayesian A* Tree Search with Expected O(N) Node Expansions: Applications to Road TrackingabstractMany perception, reasoning, and learning problems can be expressed as Bayesian inference. We point out that formulating a problem as Bayesian inference implies specifying a probability distribution on the ensemble of problem instances. This ensemble can be used for analyzing the expected complexity of algorithms and also the algorithm-independent limits of inference. We illustrate this problem by analyzing the complexity of tree search. In particular, we study the problem of road detection, as formulated by Geman and Jedynak (1996). We prove that the expected convergence is linear in the size of the road (the depth of the tree) even though the worst-case performance is exponential. We also put a bound on the constant of the convergence and place a bound on the error rates. James M. Coughlan, Alan L. Yuille |
Neural Comput. | 2 |
| 2002 | CCCP Algorithms to Minimize the Bethe and Kikuchi Free Energies: Convergent Alternatives to Belief PropagationabstractThis article introduces a class of discrete iterative algorithms that are provably convergent alternatives to belief propagation (BP) and generalized belief propagation (GBP). Our work builds on recent results by Yedidia, Freeman, and Weiss (2000), who showed that the fixed points of BP and GBP algorithms correspond to extrema of the Bethe and Kikuchi free energies, respectively. We obtain two algorithms by applying CCCP to the Bethe and Kikuchi free energies, respectively (CCCP is a procedure, introduced here, for obtaining discrete iterative algorithms by decomposing a cost function into a concave and a convex part). We implement our CCCP algorithms on two- and three-dimensional spin glasses and compare their results to BP and GBP. Our simulations show that the CCCP algorithms are stable and converge very quickly (the speed of CCCP is similar to that of BP and GBP). Unlike CCCP, BP will often not converge for these problems (GBP usually, but not always, converges). The results found by CCCP applied to the Bethe or Kikuchi free energies are equivalent, or slightly better than, those found by BP or GBP, respectively (when BP and GBP converge). Note that for these, and other problems, BP and GBP give very accurate results (see Yedidia et al., 2000), and failure to converge is their major error mode. Finally, we point out that our algorithms have a large range of inference and learning applications. Alan L. Yuille |
Neural Comput. | 1 |
| 2001 | The KGBR Viewpoint-Lighting Ambiguity and its Resolution by Generic ConstraintsabstractWe describe a novel viewpoint-lighting ambiguity which we call the KGBR. This ambiguity assumes orthographic projecting or an affine camera, and uses Lambertian reflectance functions including case/attached shadows and multiple light sources. A KGBR transform alters the geometry (by a three-dimensional affine transformation) and albedo properties of objects. If two objects are related by a KGBR transform then for any viewpoint and lighting of the first object there exists a corresponding viewpoint and lighting of the second object so that the images are identical up to an affine transformation. The Generalized Bas Relief (GBR) ambiguity is obtained as a special case of the KGBR. We describe generic viewpoint and lighting assumptions and show that either, or both, resolve this ambiguity by biasing towards objects with planar geometry. Alan L. Yuille, James M. Coughlan, Scott Konishi |
ICCV | 1 |
| 2001 | The g Factor: Relating Distributions on Features to Distributions on ImagesabstractWe describe the g-factor, which relates probability distributions on image features to distributions on the images themselves. The g-factor depends only on our choice of features and lattice quanti(cid:173) zation and is independent of the training image data. We illustrate the importance of the g-factor by analyzing how the parameters of Markov Random Field (i.e. Gibbs or log-linear) probability models of images are learned from data by maximum likelihood estimation. In particular, we study homogeneous MRF models which learn im(cid:173) age distributions in terms of clique potentials corresponding to fea(cid:173) ture histogram statistics (d. Minimax Entropy Learning (MEL) by Zhu, Wu and Mumford 1997 [11]) . We first use our analysis of the g-factor to determine when the clique potentials decouple for different features . Second, we show that clique potentials can be computed analytically by approximating the g-factor. Third, we demonstrate a connection between this approximation and the Generalized Iterative Scaling algorithm (GIS), due to Darroch and Ratcliff 1972 [2], for calculating potentials. This connection en(cid:173) ables us to use GIS to improve our multinomial approximation, using Bethe-Kikuchi[8] approximations to simplify the GIS proce(cid:173) dure. We support our analysis by computer simulations. James M. Coughlan, Alan L. Yuille |
NIPS | 2 |
| 2001 | MIME: Mutual Information Minimization and Entropy Maximization for Bayesian Belief PropagationabstractBayesian belief propagation in graphical models has been recently shown to have very close ties to inference methods based in statis- tical physics. After Yedidia et al. demonstrated that belief prop- agation (cid:12)xed points correspond to extrema of the so-called Bethe free energy, Yuille derived a double loop algorithm that is guar- anteed to converge to a local minimum of the Bethe free energy. Yuille’s algorithm is based on a certain decomposition of the Bethe free energy and he mentions that other decompositions are possi- ble and may even be fruitful. In the present work, we begin with the Bethe free energy and show that it has a principled interpre- tation as pairwise mutual information minimization and marginal entropy maximization (MIME). Next, we construct a family of free energy functions from a spectrum of decompositions of the original Bethe free energy. For each free energy in this family, we develop a new algorithm that is guaranteed to converge to a local min- imum. Preliminary computer simulations are in agreement with this theoretical development. Anand Rangarajan 0001, Alan L. Yuille |
NIPS | 2 |
| 2001 | The Concave-Convex Procedure (CCCP)abstractWe introduce the Concave-Convex procedure (CCCP) which con(cid:173) structs discrete time iterative dynamical systems which are guar(cid:173) anteed to monotonically decrease global optimization/energy func(cid:173) tions. It can be applied to (almost) any optimization problem and many existing algorithms can be interpreted in terms of CCCP. In particular, we prove relationships to some applications of Legendre transform techniques. We then illustrate CCCP by applications to Potts models, linear assignment, EM algorithms, and Generalized Iterative Scaling (GIS). CCCP can be used both as a new way to understand existing optimization algorithms and as a procedure for generating new algorithms. Alan L. Yuille, Anand Rangarajan 0001 |
NIPS | 1 |
| 2001 | Order Parameters for Detecting Target Curves in Images: When Does High Level Knowledge Help?
Alan L. Yuille, James M. Coughlan, Ying Nian Wu, Song-Chun Zhu |
Int. J. Comput. Vis. | 1 |
| 2001 | Introduction by Guest Editors
Alan L. Yuille, Song-Chun Zhu, David Mumford |
Int. J. Comput. Vis. | 1 |
| 2000 | Statistical cues for Domain Specific Image Segmentation with Performance AnalysisabstractThis paper investigates the use of colour and texture cues for segmentation of images within two specified domains. The first is the Sowerby dataset, which contains one hundred colour photographs of country roads in England that have been interactively segmented and classified into six classes-edge, vegetation, air, road, building, and other. The second domain is a set of thirty five-images, taken in San Francisco, which have been interactively segmented into similar classes. In each domain we learn the joint probability distributions of filter responses, based on colour and texture, for each class. These distributions are then used for classification. We restrict ourselves to a limited number of filters in order to ensure that the learnt filter responses do not overfit the training data (our region classes are chosen so as to ensure that there is enough data to avoid over fitting). We do performance analysis on the two datasets by evaluating the false positive and false negative error rates for the classification. This shows that the learnt models achieve high accuracy in classifying individual pixels into those classes for which the filter responses are approximately spatially homogeneous (i.e. road, vegetation, and air but not edge and building). A more sensitive performance measure, the Chernoff information, is calculated in order to quantify how well the cues for edge and building are doing. This demonstrates that statistical knowledge of the domain is a powerful tool for segmentation. Scott Konishi, Alan L. Yuille |
CVPR | 2 |
| 2000 | Order Parameters for Minimax Entropy Distributions: When Does High Level Knowledge Help?abstractMany problems in vision can be formulated as Bayesian inference. It is important to determine the accuracy of these inferences and how they depend on the problem domain. In recent work, Coughlan and Yuille showed that, for a restricted class of problems, the performance of Bayesian inference could be summarized by an order parameter K which depends on the probability distributions which characterize the problem domain. In this paper we generalize the theory of order parameters so that it applies to domains for which the probability models can be obtained by Minimax Entropy learning theory. By analyzing order parameters it is possible to determine whether a target can be detected using a general purpose "generic" model or whether a more specific "high-level" model is needed. At critical values of the order parameters the problem becomes unsolvable without the addition of extra prior knowledge. Alan L. Yuille, James M. Coughlan, Song-Chun Zhu, Ying Nian Wu |
CVPR | 1 |
| 2000 | The Manhattan World Assumption: Regularities in Scene Statistics which Enable Bayesian InferenceabstractPreliminary work by the authors made use of the so-called "Man(cid:173) hattan world" assumption about the scene statistics of city and indoor scenes. This assumption stated that such scenes were built on a cartesian grid which led to regularities in the image edge gra(cid:173) dient statistics. In this paper we explore the general applicability of this assumption and show that, surprisingly, it holds in a large variety of less structured environments including rural scenes. This enables us, from a single image, to determine the orientation of the viewer relative to the scene structure and also to detect target ob(cid:173) jects which are not aligned with the grid. These inferences are performed using a Bayesian model with probability distributions (e.g. on the image gradient statistics) learnt from real data. James M. Coughlan, Alan L. Yuille |
NIPS | 2 |
| 2000 | Efficient Deformable Template Detection and Localization without User Initialization
James M. Coughlan, Alan L. Yuille, Camper English, Daniel Snow |
Comput. Vis. Image Underst. | 2 |
| 2000 | Guest Editorial: Statistical and Computational Theories of Vision: Modeling, Learning, Sampling and Computing, Part I
Song-Chun Zhu, Alan L. Yuille, David Mumford |
Int. J. Comput. Vis. | 2 |
| 2000 | Probabilistic Motion Estimation Based on Temporal CoherenceabstractWe develop a theory for the temporal integration of visual motion motivated by psychophysical experiments. The theory proposes that input data are temporally grouped and used to predict and estimate the motion flows in the image sequence. This temporal grouping can be considered a generalization of the data association techniques that engineers use to study motion sequences. Our temporal grouping theory is expressed in terms of the Bayesian generalization of standard Kalman filtering. To implement the theory, we derive a parallel network that shares some properties of cortical networks. Computer simulations of this network demonstrate that our theory qualitatively accounts for psychophysical experiments on motion occlusion and motion outliers. In deriving our theory, we assumed spatial factorizability of the probability distributions and made the approximation of updating the marginal distributions of velocity at each point. This allowed us to perform local computations and simplified our implementation. We argue that these approximations are suitable for the stimuli we are considering (for which spatial coherence effects are negligible). Pierre-Yves Burgi, Alan L. Yuille, Norberto M. Grzywacz |
Neural Comput. | 2 |
| 2000 | Fundamental Limits of Bayesian Inference: Order Parameters and Phase Transitions for Road TrackingabstractThere is a growing interest in formulating vision problems in terms of Bayesian inference and, in particular, the maximum a posteriori (MAP) estimator. In this paper, we consider the special case of detecting roads from aerial images and demonstrate that analysis of this ensemble enables us to determine fundamental bounds on the performance of the MAP estimate. We demonstrate that there is a phase transition at a critical value of the order parameter; below this phase transition, it is impossible to detect the road by any algorithm. We derive closely related order parameters which determine the time and memory complexity of search and the accuracy of the solution using the n* search strategy. Our approach can be applied to other vision problems, and we briefly summarize the results when the model uses the "wrong prior". We comment on how our work relates to studies of the complexity of visual search and the critical behaviour in the computational cost of solving NP-complete problems. Alan L. Yuille, James M. Coughlan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |