Yizhou Yu

dblp:90/6896 · DBLP profile ↗
← Back
243ranked-venue papers
15as first author
104since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 182 · 13 first-author · 57 since 2021Artificial intelligence and machine learning · 98 · 3 first-author · 60 since 2021Applied, interdisciplinary, general and emerging computing · 49 · 32 since 2021Computer networks · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2026 Effective registration-free dual-phase segmentation for pancreas and pancreatic mass via symmetrical selective feature integration
Fuze Cong, Wenyi Deng, Xiuli Li, Zaiyi Liu, Longjiang Zhang, Zhengyu Jin, Yizhou Yu, Huadan Xue
Medical Image Anal.8
2025 SparX: A Sparse Cross-Layer Connection Mechanism for Hierarchical Vision Mamba and Transformer Networks
abstract
Due to the capability of dynamic state space models (SSMs) in capturing long-range dependencies with linear-time computational complexity, Mamba has shown notable performance in NLP tasks. This has inspired the rapid development of Mamba-based vision models, resulting in promising results in visual recognition tasks. However, such models are not capable of distilling features across layers through feature aggregation, interaction, and selection. Moreover, existing cross-layer feature aggregation methods designed for CNNs or ViTs are not practical in Mamba-based models due to high computational costs. Therefore, this paper aims to introduce an efficient cross-layer feature aggregation mechanism for vision backbone networks. Inspired by the Retinal Ganglion Cells (RGCs) in the human visual system, we propose a new sparse cross-layer connection mechanism termed SparX to effectively improve cross-layer feature interaction and reuse. Specifically, we build two different types of network layers: ganglion layers and normal layers. The former has higher connectivity and complexity, enabling multi-layer feature aggregation and interaction in an input-dependent manner. In contrast, the latter has lower connectivity and complexity. By interleaving these two types of layers, we design a new family of vision backbone networks with sparsely cross-connected layers, achieving an excellent trade-off among model size, computational cost, memory cost, and accuracy in comparison to its counterparts. For instance, with fewer parameters, SparX-Mamba-T improves the top-1 accuracy of VMamba-T from 82.5% to 83.5%, while SparX-Swin-T achieves a 1.3% increase in top-1 accuracy compared to Swin-T. Extensive experimental results demonstrate that our new connection mechanism possesses both superior performance and generalization capabilities on various vision tasks.
Meng Lou, Yunxiang Fu, Yizhou Yu
AAAI3
2025 Autoregressive Sequence Modeling for 3D Medical Image Representation
abstract
Three-dimensional (3D) medical images, such as Computed Tomography (CT) and Magnetic Resonance Imaging (MRI), are essential for clinical applications. However, the need for diverse and comprehensive representations is particularly pronounced when considering the variability across different organs, diagnostic tasks, and imaging modalities. How to effectively interpret the intricate contextual information and extract meaningful insights from these images remains an open challenge to the community. While current self-supervised learning methods have shown potential, they often consider an image as a whole thereby overlooking the extensive, complex relationships among local regions from one or multiple images. In this work, we introduce a pioneering method for learning 3D medical image representations through an autoregressive pre-training framework. Our approach sequences various 3D medical images based on spatial, contrast, and semantic correlations, treating them as interconnected visual tokens within a token sequence. By employing an autoregressive sequence modeling task, we predict the next visual token in the sequence, which allows our model to deeply understand and integrate the contextual information inherent in 3D medical images. Additionally, we implement a random startup strategy to avoid overestimating token relationships and to enhance the robustness of learning. The effectiveness of our approach is demonstrated by the superior performance over others on nine downstream tasks in public datasets.
Chu-ran Wang, Lixian Su, Fandong Zhang, Yizhou Wang 0001, Yizhou Yu
AAAI7
2025 Accurate Coronary Microvascular Segmentation with Parallel Local-Global Chains
abstract
3D vessel segmentation models aid physicians with the analysis, diagnosis, and intervention of coronary microvascular disease. Existing methods for the general medical image segmentation often produce inaccurate and discontinuous results on the complex vessel patterns, particularly for tiny vessel structures. To overcome this, recent studies have combined the convolutional neural network (CNN) and transformers in a serial architecture to enhance continuity and accuracy by leveraging long-range dependencies, i.e., the topological relationships of vessel fragments. However, the sequential architectures, where the CNN and transformers bottleneck each other, limits the ability to maintain both local and global features, which are vital in our task. In this paper, we collected the largest and highest-quality CT coronary artery dataset to date, ASACA500. Based on that, we introduce twinSeg, which enables the parallel learning of the local and global chains. To facilitate effective and efficient interaction between these chains, we propose the bidirectional attention fusion module, which enables the fusion of global and local features. Extensive evaluations conducted on our in-house ASACA and a public dataset demonstrate that twinSeg achieves state-of-the-art performance across all evaluation metrics and exhibits an exceptional average symmetric surface distance (ASSD) due to its ability to model sparse and anisotropic vessel structures. Our code and pretrained model will be released after the anonymity period.
Yaling Tao, Yanwu Xu 0001, Yizhou Yu, Jinpeng Li 0002
BIBM3
2025 SegMAN: Omni-scale Context Modeling with State Space Models and Local Attention for Semantic Segmentation
abstract
High-quality semantic segmentation relies on three key capabilities: global context modeling, local detail encoding, and multi-scale feature extraction. However, recent methods struggle to possess all these capabilities simultaneously. Hence, we aim to empower segmentation networks to simultaneously carry out efficient global context modeling, high-quality local detail encoding, and rich multi-scale feature representation for varying input resolutions. In this paper, we introduce SegMAN, a novel linear-time model comprising a hybrid feature encoder dubbed SegMAN Encoder, and a decoder based on state space models. Specifically, the SegMAN Encoder synergistically integrates sliding local attention with dynamic state space models, enabling highly efficient global context modeling while preserving fine-grained local details. Meanwhile, the MMSCopE module in our decoder enhances multi-scale context feature extraction and adaptively scales with the input resolution. Our SegMAN-B Encoder achieves 85.1% ImageNet-1k accuracy (+1.5% over VMamba-S with fewer parameters). When paired with our decoder, the full SegMAN-B model achieves 52.6% mIoU on ADE20K (+1.6% over SegNeXt-L with 15% fewer GFLOPs), 83.8% mIoU on Cityscapes (+2.1% over SegFormer-B3 with half the GFLOPs), and 1.6% higher mIoU than VWFormer-B3 on COCO-Stuff with lower GFLOPs. Our code is available at https://github.com/yunxiangfu2001/SegMAN.
Yunxiang Fu, Meng Lou, Yizhou Yu
CVPR3
2025 OverLoCK: An Overview-first-Look-Closely-next ConvNet with Context-Mixing Dynamic Kernels
abstract
Top-down attention plays a crucial role in the human vision system, wherein the brain initially obtains a rough overview of a scene to discover salient cues (i.e., overview first), followed by a more careful finer-grained examination (i.e., look closely next). However, modern ConvNets remain confined to a pyramid structure that successively downsamples the feature map for receptive field expansion, neglecting this crucial biomimetic principle. We present OverLoCK, the first pure ConvNet backbone architecture that explicitly incorporates a top-down attention mechanism. Unlike pyramid backbone networks, our design features a branched architecture with three synergistic sub-networks: 1) a Base-Net that encodes low/mid-level features; 2) a lightweight Overview-Net that generates dynamic top-down attention through coarse global context modeling (i.e., overview first); and 3) a robust Focus-Net that performs finer-grained perception guided by top-down attention (i.e., look closely next). To fully unleash the power of top-down attention, we further propose a novel context-mixing dynamic convolution (ContMix) that effectively models long-range dependencies while preserving inherent local inductive biases even when the input resolution increases, addressing critical limitations in existing convolutions. Our OverLoCK exhibits a notable performance improvement over existing methods. For instance, OverLoCK-T achieves a Top-1 accuracy of 84.2%, significantly surpassing ConvNeXt-B while using only around one-third of the FLOPs/parameters. On object detection, our OverLoCK-S clearly surpasses MogaNet-B by 1% in APb. On semantic segmentation, our OverLoCK-T remarkably improves UniRepLKNet-T by 1.7% in mIoU. Code is publicly available at https://rb.gy/wit4jh.
Meng Lou, Yizhou Yu
CVPR2
2025 Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis
abstract
Recent advances in large language models (LLMs) have spurred interests in encoding images as discrete tokens and leveraging autoregressive (AR) frameworks for visual generation. However, the quantization process in AR-based visual generation models inherently introduces information loss that degrades image fidelity. To mitigate this limitation, recent studies have explored to autoregressively predict continuous tokens. Unlike discrete tokens that reside in a structured and bounded space, continuous representations exist in an unbounded, high-dimensional space, making density estimation more challenging and increasing the risk of generating out-of-distribution artifacts. Based on the above findings, this work introduces DisCon (Discrete-Conditioned Continuous Autoregressive Model), a novel framework that reinterprets discrete tokens as conditional signals rather than generation targets. By modeling the conditional probability of continuous representations conditioned on discrete tokens, DisCon circumvents the optimization challenges of continuous token modeling while avoiding the information loss caused by quantization. DisCon achieves a gFID score of 1.38 on ImageNet 256$\times$256 generation, outperforming state-of-the-art autoregressive approaches by a clear margin. Project page: https://pengzheng0707.github.io/DisCon.
Yizhou Yu, Rui Ma 0011, Zuxuan Wu
ICCV4
2025 Vision Function Layer in Multimodal LLMs
abstract
This study identifies that visual-related functional decoding is distributed across different decoder layers in Multimodal Large Language Models (MLLMs). Typically, each function, such as counting, grounding, or OCR recognition, narrows down to two or three layers, which we define as Vision Function Layers (VFL). Additionally, the depth and its order of different VFLs exhibits a consistent pattern across different MLLMs, which is well-aligned with human behaviors (e.g., recognition occurs first, followed by counting, and then grounding). These findings are derived from Visual Token Swapping, our novel analytical framework that modifies targeted KV cache entries to precisely elucidate layer-specific functions during decoding. Furthermore, these insights offer substantial utility in tailoring MLLMs for real-world downstream applications. For instance, when LoRA training is selectively applied to VFLs whose functions align with the training data, VFL-LoRA not only outperform full-LoRA but also prevent out-of-domain function forgetting. Moreover, by analyzing the performance differential on training data when particular VFLs are ablated, VFL-select automatically classifies data by function, enabling highly efficient data selection to directly bolster corresponding capabilities. Consequently, VFL-select surpasses human experts in data selection, and achieves 98% of full-data performance with only 20% of the original dataset. This study delivers deeper comprehension of MLLM visual processing, fostering the creation of more efficient, interpretable, and robust models.
Yizhou Yu, Sibei Yang
NeurIPS2
2025 SDR-Former: A Siamese Dual-Resolution Transformer for liver lesion classification using 3D multi-phase imaging
Meng Lou, Hanning Ying, Yizhou Yu
Neural Networks6
2025 Swin-UMamba†: Adapting Mamba-Based Vision Foundation Models for Medical Image Segmentation
abstract
Vision foundation models have shown great potential in improving generalizability and data efficiency, especially for medical image segmentation since medical image datasets are relatively small due to high annotation costs and privacy concerns. However, current research on foundation models predominantly relies on transformers. The high quadratic complexity and large parameter counts make these models computationally expensive, limiting their potential for clinical applications. In this work, we introduce Swin-UMamba†, a novel Mamba-based model for medical image segmentation that seamlessly leverages the power of the vision foundation model, which is also computationally efficient with the linear complexity of Mamba. Moreover, we investigated and verified the impact of the vision foundation model on medical image segmentation, in which a self-supervised model adaptation scheme was designed to bridge the gap between natural and medical data. Notably, Swin-UMamba† outperforms 7 state-of-the-art methods, including CNN-based, transformer-based, and Mamba-based approaches across AbdomenMRI, Encoscopy, and Microscopy datasets. The code and models are publicly available at: https://github.com/JiarunLiu/Swin-UMamba.
Jiarun Liu, Hao Yang 0026, Lequan Yu, Yong Liang 0001, Yizhou Yu, Shaoting Zhang 0001, Hairong Zheng, Shanshan Wang 0002
IEEE Trans. Medical Imaging6
2025 TransXNet: Learning Both Global and Local Dynamics With a Dual Dynamic Token Mixer for Visual Recognition
abstract
Recent studies have integrated convolutions into transformers to introduce inductive bias and improve generalization performance. However, the static nature of conventional convolution prevents it from dynamically adapting to input variations, resulting in a representation discrepancy between convolution and self-attention as self-attention calculates attention matrices dynamically. Furthermore, when stacking token mixers that consist of convolution and self-attention to form a deep network, the static nature of convolution hinders the fusion of features previously generated by self-attention into convolution kernels. These two limitations result in a suboptimal representation capacity of the constructed networks. To find a solution, we propose a lightweight dual dynamic token mixer (D-Mixer) to simultaneously learn global and local dynamics, that is, mechanisms that compute weights for aggregating global contexts and local details in an input-dependent manner. D-Mixer works by applying an efficient global attention module and an input-dependent depthwise convolution separately on evenly split feature segments, endowing the network with strong inductive bias and an enlarged effective receptive field. We use D-Mixer as the basic building block to design TransXNet, a novel hybrid CNN-transformer vision backbone network that delivers compelling performance. In the ImageNet-1K image classification task, TransXNet-T surpasses Swin-T by 0.3% in top-1 accuracy while requiring less than half of the computational cost. Furthermore, TransXNet-S and TransXNet-B exhibit excellent model scalability, achieving top-1 accuracy of 83.8% and 84.6%, respectively, with reasonable computational costs. In addition, our proposed network architecture demonstrates strong generalization capabilities in various dense prediction tasks, outperforming other state-of-the-art networks while having lower computational costs. Code is publicly available at https://github.com/LMMMEng/TransXNet.
Meng Lou, Shu Zhang 0001, Sibei Yang, Chuan Wu 0001, Yizhou Yu
IEEE Trans. Neural Networks Learn. Syst.6
2024 FedDiv: Collaborative Noise Filtering for Federated Learning with Noisy Labels
abstract
Federated Learning with Noisy Labels (F-LNL) aims at seeking an optimal server model via collaborative distributed learning by aggregating multiple client models trained with local noisy or clean samples. On the basis of a federated learning framework, recent advances primarily adopt label noise filtering to separate clean samples from noisy ones on each client, thereby mitigating the negative impact of label noise. However, these prior methods do not learn noise filters by exploiting knowledge across all clients, leading to sub-optimal and inferior noise filtering performance and thus damaging training stability. In this paper, we present FedDiv to tackle the challenges of F-LNL. Specifically, we propose a global noise filter called Federated Noise Filter for effectively identifying samples with noisy labels on every client, thereby raising stability during local training sessions. Without sacrificing data privacy, this is achieved by modeling the global distribution of label noise across all clients. Then, in an effort to make the global model achieve higher performance, we introduce a Predictive Consistency based Sampler to identify more credible local data for local model training, thus preventing noise memorization and further boosting the training stability. Extensive experiments on CIFAR-10, CIFAR-100, and Clothing1M demonstrate that FedDiv achieves superior performance over state-of-the-art F-LNL methods under different label noise settings for both IID and non-IID data partitions. Source code is publicly available at https://github.com/lijichang/FLNL-FedDiv.
Jichang Li, Guanbin Li, Zicheng Liao, Yizhou Yu
AAAI5
2024 RegionGPT: Towards Region Understanding Vision Language Model
abstract
Vision language models (VLMs) have experienced rapid advancements through the integration of large language models (LLMs) with image-text pairs, yet they struggle with detailed regional visual understanding due to limited spatial awareness of the vision encoder, and the use of coarse-grained training data that lacks detailed, region-specific captions. To address this, we introduce RegionGPT (short as RGPT), a novel framework designed for complex region-level captioning and understanding. RGPT enhances the spatial awareness of regional representation with simple yet effective modifications to existing visual encoders in VLMs. We further improve performance on tasks requiring a specific output scope by integrating task-guided instruction prompts during both training and inference phases, while maintaining the model's versatility for general-purpose tasks. Additionally, we develop an automated region caption data generation pipeline, enriching the training set with detailed region-level captions. We demonstrate that a universal RGPT model can be effectively applied and significantly enhancing performance across a range of region-level tasks, including but not limited to complex region descriptions, reasoning, object classification, and referring expressions comprehension. Code will be released at the project page.
Qiushan Guo, Shalini De Mello, Hongxu Yin, Wonmin Byeon, Ka Chun Cheung, Yizhou Yu, Ping Luo 0002, Sifei Liu
CVPR6
2024 OVER-NAV: Elevating Iterative Vision-and-Language Navigation with Open-Vocabulary Detection and StructurEd Representation
abstract
Recent advances in Iterative Vision-and-Language Navigation (IVLN) introduce a more meaningful and practical paradigm of VLN by maintaining the agent's memory across tours of scenes. Although the long-term memory aligns better with the persistent nature of the VLN task, it poses more challenges on how to utilize the highly unstructured navigation memory with extremely sparse supervision. Towards this end, we propose OVER-NAV, which aims to go over and beyond the current arts of IVLN techniques. In particular, we propose to incorporate LLMs and open-vocabulary detectors to distill key information and establish correspondence between multi-modal signals. Such a mechanism introduces reliable cross-modal supervision and enables on-the-fly generalization to unseen scenes without the need of extra annotation and re-training. To fully exploit the interpreted navigation data, we further introduce a structured representation, coded Omnigraph, to effectively integrate multi-modal information along the tour. Accompanied with a novel omnigraph fusion mechanism, OVER-NAV is able to extract the most relevant knowledge from omnigraph for a more accurate navigating action. In addition, OVER-NAV seamlessly supports both discrete and continuous environments under a unified framework. We demonstrate the superiority of OVER-NAV in extensive experiments.
Ganlong Zhao, Guanbin Li, Weikai Chen 0001, Yizhou Yu
CVPR4
2024 Cross-dimensional Medical Self-supervised Representation Learning Based on a Pseudo-3D Transformation
Fandong Zhang, Yizhou Wang 0001, Chu-ran Wang, Yizhou Yu
MICCAI (11)8
2024 Swin-UMamba: Mamba-Based UNet with ImageNet-Based Pretraining
Jiarun Liu, Hao Yang 0026, Yan Xi, Lequan Yu, Cheng Li 0008, Yong Liang 0001, Guangming Shi, Yizhou Yu, Shaoting Zhang 0001, Hairong Zheng, Shanshan Wang 0002
MICCAI (9)9
2024 Exploration and Exploitation of Unlabeled Data for Open-Set Semi-supervised Learning
Ganlong Zhao, Guanbin Li, Yipeng Qin, Zhenhua Chai, Xiaolin Wei, Liang Lin 0004, Yizhou Yu
Int. J. Comput. Vis.8
2024 A Survey on Graph Neural Networks and Graph Transformers in Computer Vision: A Task-Oriented Perspective
abstract
Graph Neural Networks (GNNs) have gained momentum in graph representation learning and boosted the state of the art in a variety of areas, such as data mining (e.g., social network analysis and recommender systems), computer vision (e.g., object detection and point cloud learning), and natural language processing (e.g., relation extraction and sequence learning), to name a few. With the emergence of Transformers in natural language processing and computer vision, graph Transformers embed a graph structure into the Transformer architecture to overcome the limitations of local neighborhood aggregation while avoiding strict structural inductive biases. In this paper, we present a comprehensive review of GNNs and graph Transformers in computer vision from a task-oriented perspective. Specifically, we divide their applications in computer vision into five categories according to the modality of input data, i.e., 2D natural images, videos, 3D data, vision + language, and medical images. In each category, we further divide the applications according to a set of vision tasks. Such a task-oriented taxonomy allows us to examine how each task is tackled by different GNN-based approaches and how well these approaches perform. Based on the necessary preliminaries, we provide the definitions and challenges of the tasks, in-depth coverage of the representative approaches, as well as discussions regarding insights, limitations, and future directions.
Chaoqi Chen, Yushuang Wu, Qiyuan Dai 0001, Mutian Xu, Sibei Yang, Xiaoguang Han 0001, Yizhou Yu
IEEE Trans. Pattern Anal. Mach. Intell.8
2024 I2F: A Unified Image-to-Feature Approach for Domain Adaptive Semantic Segmentation
abstract
Unsupervised domain adaptation (UDA) for semantic segmentation is a promising task freeing people from heavy annotation work. However, domain discrepancies in low-level image statistics and high-level contexts compromise the segmentation performance over the target domain. A key idea to tackle this problem is to perform both image-level and feature-level adaptation jointly. Unfortunately, there is a lack of such unified approaches for UDA tasks in the existing literature. This paper proposes a novel UDA pipeline for semantic segmentation that unifies image-level and feature-level adaptation. Concretely, for image-level domain shifts, we propose a global photometric alignment module and a global texture alignment module that align images in the source and target domains in terms of image-level properties. For feature-level domain shifts, we perform global manifold alignment by projecting pixel features from both domains onto the feature manifold of the source domain; and we further regularize category centers in the source domain through a category-oriented triplet loss, and perform target domain consistency regularization over augmented target domain images. Experimental results demonstrate that our pipeline significantly outperforms previous methods. In the commonly tested GTA5 →Cityscapes task, our proposed method using Deeplab V3+ as the backbone surpasses previous SOTA by 8%, achieving 58.2% in mIoU.
Xiangru Lin, Yizhou Yu
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Inter-domain mixup for semi-supervised domain adaptation
Jichang Li, Guanbin Li, Yizhou Yu
Pattern Recognit.3
2024 Correction to "A Structure-Aware Relation Network for Thoracic Diseases Detection and Segmentation"
abstract
In the above article[1], there are errors on pages 2045 and 2046. Section METHOD.D and Section METHOD.E should be the subsections of Section METHOD.C, i.e., METHOD.C: 1) Relation Graph Construction; 2) Message Passing via Relation Graph; and 3) Mapping Disease Relation to Regions.
Jingyu Liu 0004, Shu Zhang 0001, Dingwen Zhang, Yizhou Yu
IEEE Trans. Medical Imaging7
2024 MVCNet: Multiview Contrastive Network for Unsupervised Representation Learning for 3-D CT Lesions
abstract
With the renaissance of deep learning, automatic diagnostic algorithms for computed tomography (CT) have achieved many successful applications. However, they heavily rely on lesion-level annotations, which are often scarce due to the high cost of collecting pathological labels. On the other hand, the annotated CT data, especially the 3-D spatial information, may be underutilized by approaches that model a 3-D lesion with its 2-D slices, although such approaches have been proven effective and computationally efficient. This study presents a multiview contrastive network (MVCNet), which enhances the representations of 2-D views contrastively against other views of different spatial orientations. Specifically, MVCNet views each 3-D lesion from different orientations to collect multiple 2-D views; it learns to minimize a contrastive loss so that the 2-D views of the same 3-D lesion are aggregated, whereas those of different lesions are separated. To alleviate the issue of false negative examples, the uninformative negative samples are filtered out, which results in more discriminative features for downstream tasks. By linear evaluation, MVCNet achieves state-of-the-art accuracies on the lung image database consortium and image database resource initiative (LIDC-IDRI) (88.62%), lung nodule database (LNDb) (76.69%), and TianChi (84.33%) datasets for unsupervised representation learning. When fine-tuned on 10% of the labeled data, the accuracies are comparable to the supervised learning models (89.46% versus 85.03%, 73.85% versus 73.44%, 83.56% versus 83.34% on the three datasets, respectively), indicating the superiority of MVCNet in learning representations with limited annotations. Our findings suggest that contrasting multiple 2-D views is an effective approach to capturing the original 3-D information, which notably improves the utilization of the scarce and valuable annotated CT data.
Penghua Zhai, Huaiwei Cong, Enwei Zhu, Gangming Zhao, Yizhou Yu, Jinpeng Li 0002
IEEE Trans. Neural Networks Learn. Syst.5
2024 SketchMetaFace: A Learning-Based Sketching Interface for High-Fidelity 3D Character Face Modeling
abstract
Modeling 3D avatars benefits various application scenarios such as AR/VR, gaming, and filming. Character faces contribute significant diversity and vividity as a vital component of avatars. However, building 3D character face models usually requires a heavy workload with commercial tools, even for experienced artists. Various existing sketch-based tools fail to support amateurs in modeling diverse facial shapes and rich geometric details. In this article, we present SketchMetaFace - a sketching system targeting amateur users to model high-fidelity 3D faces in minutes. We carefully design both the user interface and the underlying algorithm. First, curvature-aware strokes are adopted to better support the controllability of carving facial details. Second, considering the key problem of mapping a 2D sketch map to a 3D model, we develop a novel learning-based method termed "Implicit and Depth Guided Mesh Modeling" (IDGMM). It fuses the advantages of mesh, implicit, and depth representations to achieve high-quality results with high efficiency. In addition, to further support usability, we present a coarse-to-fine 2D sketching interface design and a data-driven stroke suggestion tool. User studies demonstrate the superiority of our system over existing modeling tools in terms of the ease to use and visual quality of results. Experimental analyses also show that IDGMM reaches a better trade-off between accuracy and efficiency.
Zhongjin Luo, Dong Du 0002, Heming Zhu, Yizhou Yu, Hongbo Fu 0001, Xiaoguang Han 0001
IEEE Trans. Vis. Comput. Graph.4
2023 RankDNN: Learning to Rank for Few-Shot Learning
abstract
This paper introduces a new few-shot learning pipeline that casts relevance ranking for image retrieval as binary ranking relation classification. In comparison to image classification, ranking relation classification is sample efficient and domain agnostic. Besides, it provides a new perspective on few-shot learning and is complementary to state-of-the-art methods. The core component of our deep neural network is a simple MLP, which takes as input an image triplet encoded as the difference between two vector-Kronecker products, and outputs a binary relevance ranking order. The proposed RankMLP can be built on top of any state-of-the-art feature extractors, and our entire deep neural network is called the ranking deep neural network, or RankDNN. Meanwhile, RankDNN can be flexibly fused with other post-processing methods. During the meta test, RankDNN ranks support images according to their similarity with the query samples, and each query sample is assigned the class label of its nearest neighbor. Experiments demonstrate that RankDNN can effectively improve the performance of its baselines based on a variety of backbones and it outperforms previous state-of-the-art algorithms on multiple few-shot learning benchmarks, including miniImageNet, tieredImageNet, Caltech-UCSD Birds, and CIFAR-FS. Furthermore, experiments on the cross-domain challenge demonstrate the superior transferability of RankDNN.The code is available at: https://github.com/guoqianyu-alberta/RankDNN.
Haotong Gong, Xujun Wei, Yanwei Fu 0001, Yizhou Yu, Weifeng Ge
AAAI5
2023 Geometry-Aware Network for Domain Adaptive Semantic Segmentation
abstract
Measuring and alleviating the discrepancies between the synthetic (source) and real scene (target) data is the core issue for domain adaptive semantic segmentation. Though recent works have introduced depth information in the source domain to reinforce the geometric and semantic knowledge transfer, they cannot extract the intrinsic 3D information of objects, including positions and shapes, merely based on 2D estimated depth. In this work, we propose a novel Geometry-Aware Network for Domain Adaptation (GANDA), leveraging more compact 3D geometric point cloud representations to shrink the domain gaps. In particular, we first utilize the auxiliary depth supervision from the source domain to obtain the depth prediction in the target domain to accomplish structure-texture disentanglement. Beyond depth estimation, we explicitly exploit 3D topology on the point clouds generated from RGB-D images for further coordinate-color disentanglement and pseudo-labels refinement in the target domain. Moreover, to improve the 2D classifier in the target domain, we perform domain-invariant geometric adaptation from source to target and unify the 2D semantic and 3D geometric segmentation results in two domains. Note that our GANDA is plug-and-play in any existing UDA framework. Qualitative and quantitative results demonstrate that our model outperforms state-of-the-arts on GTA5->Cityscapes and SYNTHIA->Cityscapes.
Yinghong Liao, Wending Zhou, Xu Yan 0005, Zhen Li 0026, Yizhou Yu, Shuguang Cui
AAAI5
2023 START: Automatic Sleep Staging with Attention-based Cross-modal Learning Transformer
abstract
Automatic sleep staging is vital to scale up sleep assessment and diagnosis to serve millions experiencing sleep deprivation and disorders and enable longitudinal sleep monitoring in home environments. However, how to learn from multi-channel raw physiological signal inputs (e.g., EEG and EOG) to capture the sleep stage and physiological signal relations remains a big challenge. In this paper, we propose a sleep staging model, named Sleep Staging Cross-modal Transformer (START), which is a transformer-only method for sleep stage classification. Our model is capable of learning a joint representation from both EEG and EOG signals by using a cross-modal fusion strategy. Experimental results show that our model outperforms the state-of-the-art methods on two public datasets. Furthermore, our model provides considerable reductions in parameters and training time compared to previous methods.
Jingpeng Sun, Rongxiao Wang, Gangming Zhao, Chen Chen 0036, Yixiao Qu, Xiyuan Hu, Yizhou Yu
BIBM8
2023 Leveraging Frequency Domain Learning in 3D Vessel Segmentation
abstract
Coronary microvascular disease constitutes a substantial risk to human health. Employing computer-aided analysis and diagnostic systems, medical professionals can intervene early in disease progression, with 3D vessel segmentation serving as a crucial component. Nevertheless, conventional U-Net architectures tend to yield incoherent and imprecise segmentation outcomes, particularly for small vessel structures. While models with attention mechanisms, such as Transformers and large convolutional kernels, demonstrate superior performance, their extensive computational demands during training and inference lead to increased time complexity. In this study, we leverage Fourier domain learning as a substitute for multi-scale convo-lutional kernels in 3D hierarchical segmentation models, which can reduce computational expenses while preserving global receptive fields within the network. Furthermore, a zero-parameter frequency domain fusion method is designed to improve the skip connections in U-Net architecture. Experimental results on a public dataset and an in-house dataset indicate that our novel Fourier transformation-based network achieves remarkable dice performance (84.37% on ASACA500 and 80.32% on ImageCAS) in tubular vessel segmentation tasks and substantially reduces computational requirements without compromising global receptive fields.
Xinyuan Wang 0009, Chengwei Pan, Hongming Dai, Gangming Zhao, Yizhou Yu
BIBM7
2023 MISC210K: A Large-Scale Dataset for Multi-Instance Semantic Correspondence
abstract
Semantic correspondence have built up a new way for object recognition. However current single-object matching schema can be hard for discovering commonalities for a category and far from the real-world recognition tasks. To fill this gap, we design the multi-instance semantic correspondence task which aims at constructing the correspondence between multiple objects in an image pair. To support this task, we build a multi-instance semantic correspondence (MISC) dataset from COCO Detection 2017 task called MISC210K. We construct our dataset as three steps: (1) category selection and data cleaning; (2) keypoint design based on 3D models and object description rules; (3) human-machine collaborative annotation. Following these steps, we select 34 classes of objects with 4,812 challenging images annotated via a well designed semi-automatic workflow, and finally acquire 218,179 image pairs with instance masks and instance-level keypoint pairs annotated. We design a dual-path collaborative learning pipeline to train instance-level co-segmentation task and fine-grained level correspondence task together. Benchmark evaluation and further ablation results with detailed analysis are provided with three future directions proposed. Our project is available on https://github.com/YXSUNMADMAX/MISC210K.
Yixuan Sun, Haijing Guo, Yuzhou Zhao, Runmin Wu, Yizhou Yu, Weifeng Ge
CVPR6
2023 Improved Distribution Matching for Dataset Condensation
abstract
Dataset Condensation aims to condense a large dataset into a smaller one while maintaining its ability to train a well-performing model, thus reducing the storage cost and training effort in deep learning applications. However, conventional dataset condensation methods are optimization-oriented and condense the dataset by performing gradient or parameter matching during model optimization, which is computationally intensive even on small datasets and models. In this paper, we propose a novel dataset condensation method based on distribution matching, which is more efficient and promising. Specifically, we identify two important shortcomings of naive distribution matching (i.e., imbalanced feature numbers and unvalidated embeddings for distance computation) and address them with three novel techniques (i.e., partitioning and expansion augmentation, efficient and enriched model sampling, and class-aware distribution regularization). Our simple yet effective method outperforms most previous optimization-oriented methods with much fewer computational resources, thereby scaling data condensation to larger datasets and models. Extensive experiments demonstrate the effectiveness of our method. Codes are available at https://github.com/uitrbn/IDM
Ganlong Zhao, Guanbin Li, Yipeng Qin, Yizhou Yu
CVPR4
2023 Activate and Reject: Towards Safe Domain Generalization under Category Shift
abstract
Albeit the notable performance on in-domain test points, it is non-trivial for deep neural networks to attain satisfactory accuracy when deploying in the open world, where novel domains and object classes often occur. In this paper, we study a practical problem of Domain Generalization under Category Shift (DGCS), which aims to simultaneously detect unknown-class samples and classify known-class samples in the target domains. Compared to prior DG works, we face two new challenges: 1) how to learn the concept of "unknown " during training with only source known- class samples, and 2) how to adapt the source-trained model to unseen environments for safe model deployment. To this end, we propose a novel Activate and Reject (ART) framework to reshape the model’s decision boundary to accommodate unknown classes and conduct post hoc modification to further discriminate known and unknown classes using unlabeled test data. Specifically, during training, we promote the response to the unknown by optimizing the unknown probability and then smoothing the overall output to mitigate the overconfidence issue. At test time, we introduce a step-wise online adaptation method that predicts the label by virtue of the cross-domain nearest neighbor and class prototype information without updating the network’s parameters or using threshold-based mechanisms. Experiments reveal that ART consistently improves the generalization capability of deep networks on different vision tasks. For image classification, ART improves the H-score by 6.1% on average compared to the previous best method. For object detection and semantic segmentation, we establish new benchmarks and achieve competitive performance.
Chaoqi Chen, Luyao Tang, Leitian Tao, Yue Huang 0001, Xiaoguang Han 0001, Yizhou Yu
ICCV7
2023 EGC: Image Generation and Classification via a Diffusion Energy-Based Model
abstract
Learning image classification and image generation using the same set of network parameters presents a formidable challenge. Recent advanced approaches perform well in one task often exhibit poor performance in the other. This work introduces an energy-based classifier and generator, namely EGC, which can achieve superior performance in both tasks using a single neural network. Unlike conventional classifiers that produce a label given an image (i.e., a conditional distribution p(y|x)), the forward pass in EGC is a classification model that yields a joint distribution p(x,y), enabling a diffusion model in its backward pass by marginalizing out the label y to estimate the score function. Furthermore, EGC can be adapted for unsupervised learning by considering the label as latent variables. EGC achieves competitive generation results compared with state-of-the-art approaches on ImageNet-1k, CelebA-HQ and LSUN Church, while achieving superior classification accuracy and robustness against adversarial attacks on CIFAR-10. This work marks the inaugural success in mastering both domains using a unified network parameter set. We believe that EGC bridges the gap between discriminative and generative learning. Code will be released at https://github.com/GuoQiushan/EGC.
Qiushan Guo, Chuofan Ma, Yi Jiang 0009, Zehuan Yuan, Yizhou Yu, Ping Luo 0002
ICCV5
2023 Learning Domain-Agnostic Representation for Disease Diagnosis
Chu-ran Wang, Jing Li 0091, Xinwei Sun 0001, Fandong Zhang, Yizhou Yu, Yizhou Wang 0001
ICLR5
2023 Protein Representation Learning via Knowledge Enhanced Primary Structure Reasoning
Yunxiang Fu, Zhicheng Zhang 0005, Cheng Bian, Yizhou Yu
ICLR5
2023 Advancing Radiograph Representation Learning with Masked Record Modeling
Chenyu Lian, Yizhou Yu
ICLR4
2023 Dynamic Triple Reweighting Network for Automatic Femoral Head Necrosis Diagnosis from Computed Tomography
abstract
Avascular necrosis of the femoral head (AVNFH) is a common orthopedic disease that seriously affects the life quality of middle-aged and elderly people. Early AVNFH is difficult to diagnose due to its complex symptoms. In recent years, some works have applied deep learning algorithms to find traces of early AVNFH in X-rays or magnetic resonance imaging (MRI). However, X-rays are difficult to reflect hidden features due to the tissue overlap; MRI is sensitive but requires more time for imaging and is expensive. This study aims to develop a computer-aided diagnosis system for early AVNFH based on computed tomography (CT), which provides layer-wise features and is less costly. To achieve this, a large-scale dataset for AVNFH was collected and annotated by experienced doctors. We propose the Dynamic Triple Reweighting Network (DTRNet) that integrates the AVNFH classification and weakly-supervised localization. DTRNet incorporates nested multi-instance learning as the first and second reweighting, and structure regularization as the third reweighting to identify diseases and localize the lesion region. Since nested multi-instance learning is inapplicable in situations with few positive samples in the patch set, we propose a dynamic pseudo-package module to compensate for this limitation. Experimental results show that DTRNet is superior to the baselines in AVNFH classification. In addition, it can locate lesions to provide more information for assisting clinical decisions. The desensitized data and codes has been made available at: https://github.com/tomas-lilingfeng/DTRNet.
Gangming Zhao, Yizhou Yu, Jinpeng Li 0002
ACM Multimedia3
2023 CODA: Generalizing to Open and Unseen Domains with Compaction and Disambiguation
abstract
The generalization capability of machine learning systems degenerates notably when the test distribution drifts from the training distribution. Recently, Domain Generalization (DG) has been gaining momentum in enabling machine learning models to generalize to unseen domains. However, most DG methods assume that training and test data share an identical label space, ignoring the potential unseen categories in many real-world applications. In this paper, we delve into a more general but difficult problem termed Open Test-Time DG (OTDG), where both domain shift and open class may occur on the unseen test data. We propose Compaction and Disambiguation (CODA), a novel two-stage framework for learning compact representations and adapting to open classes in the wild. To meaningfully regularize the model's decision boundary, CODA introduces virtual unknown classes and optimizes a new training objective to insert unknowns into the latent space by compacting the embedding space of source known classes. To adapt target samples to the source model, we then disambiguate the decision boundaries between known and unknown classes with a test-time training objective, mitigating the adaptivity gap and catastrophic forgetting challenges. Experiments reveal that CODA can significantly outperform the previous best method on standard DG datasets and harmonize the classification accuracy between known and unknown classes.
Chaoqi Chen, Luyao Tang, Yue Huang 0001, Xiaoguang Han 0001, Yizhou Yu
NeurIPS5
2023 Advancing 3D medical image analysis with variable dimension transform based supervised 3D pre-training
Shu Zhang 0001, Jiechao Ma, Yizhou Yu
Neurocomputing5
2023 Transformer guided progressive fusion network for 3D pancreas and pancreatic mass segmentation
Taiping Qu, Xiuli Li, Xiheng Wang, Wenyi Deng, Zaiyi Liu, Longjiang Zhang, Zhengyu Jin, Huadan Xue, Yizhou Yu
Medical Image Anal.13
2023 Relation Matters: Foreground-Aware Graph-Based Relational Reasoning for Domain Adaptive Object Detection
abstract
Domain Adaptive Object Detection (DAOD) focuses on improving the generalization ability of object detectors via knowledge transfer. Recent advances in DAOD strive to change the emphasis of the adaptation process from global to local in virtue of fine-grained feature alignment methods. However, both the global and local alignment approaches fail to capture the topological relations among different foreground objects as the explicit dependencies and interactions between and within domains are neglected. In this case, only seeking one-vs-one alignment does not necessarily ensure the precise knowledge transfer. Moreover, conventional alignment-based approaches may be vulnerable to catastrophic overfitting regarding those less transferable regions (e.g., backgrounds) due to the accumulation of inaccurate localization results in the target domain. To remedy these issues, we first formulate DAOD as an open-set domain adaptation problem, in which the foregrounds and backgrounds are seen as the "known classes" and "unknown class" respectively. Accordingly, we propose a new and general framework for DAOD, named Foreground-aware Graph-based Relational Reasoning (FGRR), which incorporates graph structures into the detection pipeline to explicitly model the intra- and inter-domain foreground object relations on both pixel and semantic spaces, thereby endowing the DAOD model with the capability of relational reasoning beyond the popular alignment-based paradigm. FGRR first identifies the foreground pixels and regions by searching reliable correspondence and cross-domain similarity regularization respectively. The inter-domain visual and semantic correlations are hierarchically modeled via bipartite graph structures, and the intra-domain relations are encoded via graph attention mechanisms. Through message-passing, each node aggregates semantic and contextual information from the same and opposite domain to substantially enhance its expressive power. Empirical results demonstrate that the proposed FGRR exceeds the state-of-the-art performance on four DAOD benchmarks.
Chaoqi Chen, Jiongcheng Li, Xiaoguang Han 0001, Yue Huang 0001, Xinghao Ding, Yizhou Yu
IEEE Trans. Pattern Anal. Mach. Intell.7
2023 A Unified Visual Information Preservation Framework for Self-supervised Pre-Training in Medical Image Analysis
abstract
Recent advances in self-supervised learning (SSL) in computer vision are primarily comparative, whose goal is to preserve invariant and discriminative semantics in latent representations by comparing siamese image views. However, the preserved high-level semantics do not contain enough local information, which is vital in medical image analysis (e.g., image-based diagnosis and tumor segmentation). To mitigate the locality problem of comparative SSL, we propose to incorporate the task of pixel restoration for explicitly encoding more pixel-level information into high-level semantics. We also address the preservation of scale information, a powerful tool in aiding image understanding but has not drawn much attention in SSL. The resulting framework can be formulated as a multi-task optimization problem on the feature pyramid. Specifically, we conduct multi-scale pixel restoration and siamese feature comparison in the pyramid. In addition, we propose non-skip U-Net to build the feature pyramid and develop sub-crop to replace multi-crop in 3D medical imaging. The proposed unified SSL framework (PCRLv2) surpasses its self-supervised counterparts on various tasks, including brain tumor segmentation (BraTS 2018), chest pathology identification (ChestX-ray, CheXpert), pulmonary nodule detection (LUNA), and abdominal organ segmentation (LiTS), sometimes outperforming them by large margins with limited annotations. Codes and models are available at https://github.com/RL4M/PCRLv2.
Chixiang Lu, Chaoqi Chen, Sibei Yang, Yizhou Yu
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Adaptive Betweenness Clustering for Semi-Supervised Domain Adaptation
abstract
Compared to unsupervised domain adaptation, semi-supervised domain adaptation (SSDA) aims to significantly improve the classification performance and generalization capability of the model by leveraging the presence of a small amount of labeled data from the target domain. Several SSDA approaches have been developed to enable semantic-aligned feature confusion between labeled (or pseudo labeled) samples across domains; nevertheless, owing to the scarcity of semantic label information of the target domain, they were arduous to fully realize their potential. In this study, we propose a novel SSDA approach named Graph-based Adaptive Betweenness Clustering (G-ABC) for achieving categorical domain alignment, which enables cross-domain semantic alignment by mandating semantic transfer from labeled data of both the source and target domains to unlabeled target samples. In particular, a heterogeneous graph is initially constructed to reflect the pairwise relationships between labeled samples from both domains and unlabeled ones of the target domain. Then, to degrade the noisy connectivity in the graph, connectivity refinement is conducted by introducing two strategies, namely Confidence Uncertainty based Node Removal and Prediction Dissimilarity based Edge Pruning. Once the graph has been refined, Adaptive Betweenness Clustering is introduced to facilitate semantic transfer by using across-domain betweenness clustering and within-domain betweenness clustering, thereby propagating semantic label information from labeled samples across domains to unlabeled target data. Extensive experiments on three standard benchmark datasets, namely DomainNet, Office-Home, and Office-31, indicated that our method outperforms previous state-of-the-art SSDA approaches, demonstrating the superiority of the proposed G-ABC algorithm.
Jichang Li, Guanbin Li, Yizhou Yu
IEEE Trans. Image Process.3
2023 nnFormer: Volumetric Medical Image Segmentation via a 3D Transformer
abstract
Transformer, the model of choice for natural language processing, has drawn scant attention from the medical imaging community. Given the ability to exploit long-term dependencies, transformers are promising to help atypical convolutional neural networks to learn more contextualized visual representations. However, most of recently proposed transformer-based segmentation approaches simply treated transformers as assisted modules to help encode global context into convolutional representations. To address this issue, we introduce nnFormer (i.e., not-another transFormer), a 3D transformer for volumetric medical image segmentation. nnFormer not only exploits the combination of interleaved convolution and self-attention operations, but also introduces local and global volume-based self-attention mechanism to learn volume representations. Moreover, nnFormer proposes to use skip attention to replace the traditional concatenation/summation operations in skip connections in U-Net like architecture. Experiments show that nnFormer significantly outperforms previous transformer-based counterparts by large margins on three public datasets. Compared to nnUNet, the most widely recognized convnet-based 3D medical segmentation model, nnFormer produces significantly lower HD95 and is much more computationally efficient. Furthermore, we show that nnFormer and nnUNet are highly complementary to each other in model ensembling. Codes and models of nnFormer are available at https://git.io/JSf3i.
Jiansen Guo, Xiaoguang Han 0001, Lequan Yu, Liansheng Wang 0002, Yizhou Yu
IEEE Trans. Image Process.7
2023 A Knowledge-Guided Framework for Fine-Grained Classification of Liver Lesions Based on Multi-Phase CT Images
abstract
Automatic and accurate differentiation of liver lesions from multi-phase computed tomography imaging is critical for the early detection of liver cancer. Multi-phase data can provide more diagnostic information than single-phase data, and the effective use of multi-phase data can significantly improve diagnostic accuracy. Current fusion methods usually fuse multi-phase information at the image level or feature level, ignoring the specificity of each modality, therefore, the information integration capacity is always limited. In this paper, we propose a Knowledge-guided framework, named MCCNet, which adaptively integrates multi-phase liver lesion information from three different stages to fully utilize and fuse multi-phase liver information. Specifically, 1) a multi-phase self-attention module was designed to adaptively combine and integrate complementary information from three phases using multi-level phase features; 2) a cross-feature interaction module was proposed to further integrate multi-phase fine-grained features from a global perspective; 3) a cross-lesion correlation module was proposed for the first time to imitate the clinical diagnosis process by exploiting inter-lesion correlation in the same patient. By integrating the above three modules into a 3D backbone, we constructed a lesion classification network. The proposed lesion classification network was validated on an in-house dataset containing 3,683 lesions from 2,333 patients in 9 hospitals. Extensive experimental results and evaluations on real-world clinical applications demonstrate the effectiveness of the proposed modules in exploiting and fusing multi-phase information.
Xingxin Xu, Qikui Zhu, Hanning Ying, Jiongcheng Li, Xiujun Cai, Shuo Li 0001, Yizhou Yu
IEEE J. Biomed. Health Informatics8
2023 Graph Convolution Based Cross-Network Multiscale Feature Fusion for Deep Vessel Segmentation
abstract
Vessel segmentation is widely used to help with vascular disease diagnosis. Vessels reconstructed using existing methods are often not sufficiently accurate to meet clinical use standards. This is because 3D vessel structures are highly complicated and exhibit unique characteristics, including sparsity and anisotropy. In this paper, we propose a novel hybrid deep neural network for vessel segmentation. Our network consists of two cascaded subnetworks performing initial and refined segmentation respectively. The second subnetwork further has two tightly coupled components, a traditional CNN-based U-Net and a graph U-Net. Cross-network multi-scale feature fusion is performed between these two U-shaped networks to effectively support high-quality vessel segmentation. The entire cascaded network can be trained from end to end. The graph in the second subnetwork is constructed according to a vessel probability map as well as appearance and semantic similarities in the original CT volume. To tackle the challenges caused by the sparsity and anisotropy of vessels, a higher percentage of graph nodes are distributed in areas that potentially contain vessels while a higher percentage of edges follow the orientation of potential nearby vessels. Extensive experiments demonstrate our deep network achieves state-of-the-art 3D vessel segmentation performance on multiple public and in-house datasets.
Gangming Zhao, Kongming Liang, Chengwei Pan, Fandong Zhang, Xianpeng Wu, Xinyang Hu, Yizhou Yu
IEEE Trans. Medical Imaging7
2023 EMS: 3D Eyebrow Modeling from Single-View Images
abstract
Eyebrows play a critical role in facial expression and appearance. Although the 3D digitization of faces is well explored, less attention has been drawn to 3D eyebrow modeling. In this work, we propose EMS, the first learning-based framework for single-view 3D eyebrow reconstruction. Following the methods of scalp hair reconstruction, we also represent the eyebrow as a set of fiber curves and convert the reconstruction to fibers growing problem. Three modules are then carefully designed: RootFinder firstly localizes the fiber root positions which indicate where to grow; OriPredictor predicts an orientation field in the 3D space to guide the growing of fibers; FiberEnder is designed to determine when to stop the growth of each fiber. Our OriPredictor directly borrows the method used in hair reconstruction. Considering the differences between hair and eyebrows, both RootFinder and FiberEnder are newly proposed. Specifically, to cope with the challenge that the root location is severely occluded, we formulate root localization as a density map estimation task. Given the predicted density map, a density-based clustering method is further used for finding the roots. For each fiber, the growth starts from the root point and moves step by step until the ending, where each step is defined as an oriented line segment with a constant length according to the predicted orientation field. To determine when to end, a pixel-aligned RNN architecture is designed to form a binary classifier, which outputs stop or not for each growing step. To support the training of all proposed networks, we build the first 3D synthetic eyebrow dataset that contains 400 high-quality eyebrow models manually created by artists. Extensive experiments have demonstrated the effectiveness of the proposed EMS pipeline on a variety of different eyebrow styles and lengths, ranging from short and sparse to long bushy eyebrows.
Chenghong Li, Leyang Jin 0001, Yujian Zheng, Yizhou Yu, Xiaoguang Han 0001
ACM Trans. Graph.4
2022 A Causal Inference Look at Unsupervised Video Anomaly Detection
abstract
Unsupervised video anomaly detection, a task that requires no labeled normal/abnormal training data in any form, is challenging yet of great importance to both industrial applications and academic research. Existing methods typically follow an iterative pseudo label generation process. However, they lack a principled analysis of the impact of such pseudo label generation on training. Furthermore, the long-range temporal dependencies also has been overlooked, which is unreasonable since the definition of an abnormal event depends on the long-range temporal context. To this end, first, we propose a causal graph to analyze the confounding effect of the pseudo label generation process. Then, we introduce a simple yet effective causal inference based framework to disentangle the noisy pseudo label's impact. Finally, we perform counterfactual based model ensemble that blends long-range temporal context with local image context in inference to make final anomaly detection. Extensive experiments on six standard benchmark datasets show that our proposed method significantly outperforms previous state-of-the-art methods, demonstrating our framework's effectiveness.
Xiangru Lin, Guanbin Li, Yizhou Yu
AAAI4
2022 A Causal Debiasing Framework for Unsupervised Salient Object Detection
abstract
Unsupervised Salient Object Detection (USOD) is a promising yet challenging task that aims to learn a salient object detection model without any ground-truth labels. Self-supervised learning based methods have achieved remarkable success recently and have become the dominant approach in USOD. However, we observed that two distribution biases of salient objects limit further performance improvement of the USOD methods, namely, contrast distribution bias and spatial distribution bias. Concretely, contrast distribution bias is essentially a confounder that makes images with similar high-level semantic contrast and/or low-level visual appearance contrast spuriously dependent, thus forming data-rich contrast clusters and leading the training process biased towards the data-rich contrast clusters in the data. Spatial distribution bias means that the position distribution of all salient objects in a dataset is concentrated on the center of the image plane, which could be harmful to off-center objects prediction. This paper proposes a causal based debiasing framework to disentangle the model from the impact of such biases. Specifically, we use causal intervention to perform de-confounded model training to minimize the contrast distribution bias and propose an image-level weighting strategy that softly weights each image's importance according to the spatial distribution bias map. Extensive experiments on 6 benchmark datasets show that our method significantly outperforms previous unsupervised state-of-the-art methods and even surpasses some of the supervised methods, demonstrating our debiasing framework's effectiveness.
Xiangru Lin, Guanqi Chen, Guanbin Li, Yizhou Yu
AAAI5
2022 BOAT: Bilateral Local Attention Vision Transformer
Gangming Zhao, Ping Li 0001, Yizhou Yu
BMVC4
2022 Compound Domain Generalization via Meta-Knowledge Encoding
abstract
Domain generalization (DG) aims to improve the generalization performance for an unseen target domain by using the knowledge of multiple seen source domains. Mainstream DG methods typically assume that the domain label of each source sample is known a priori, which is challenged to be satisfied in many real-world applications. In this paper, we study a practical problem of compound DG, which relaxes the discrete domain assumption to the mixed source domains setting. On the other hand, current DG algorithms prioritize the focus on semantic invariance across domains (one-vs-one), while paying less attention to the holistic semantic structure (many-vs-many). Such holistic semantic structure, referred to as meta-knowledge here, is crucial for learning generalizable representations. To this end, we present COmpound domain generalization via Meta-knowledge ENcoding (COMEN), a general approach to automatically discover and model latent domains in two steps. Firstly, we introduce Style-induced Domain-specific Normalization (SDNorm) to re-normalize the multi-modal underlying distributions, thereby dividing the mixture of source domains into latent clusters. Secondly, we harness the prototype representations, the centroids of classes, to perform relational modeling in the embedding space with two parallel and complementary modules, which explicitly encode the semantic structure for the out-of-distribution generalization. Experiments on four standard DG benchmarks reveal that COMEN exceeds the state-of-the-art performance without the need of domain supervision.
Chaoqi Chen, Jiongcheng Li, Xiaoguang Han 0001, Yizhou Yu
CVPR5
2022 Scale-Equivalent Distillation for Semi-Supervised Object Detection
abstract
Recent Semi-Supervised Object Detection (SS-OD) methods are mainly based on self-training, i.e., generating hard pseudo-labels by a teacher model on unlabeled data as supervisory signals. Although they achieved certain success, the limited labeled data in semi-supervised learning scales up the challenges of object detection. We analyze the challenges these methods meet with the empirical experiment results. We find that the massive False Negative samples and inferior localization precision lack consideration. Besides, the large variance of object sizes and class imbalance (i.e., the extreme ratio between back-ground and object) hinder the performance of prior arts. Further, we overcome these challenges by introducing a novel approach, Scale-Equivalent Distillation (SED), which is a simple yet effective end-to-end knowledge distillation framework robust to large object size variance and class imbalance. SED has several appealing benefits compared to the previous works. (1) SED imposes a consistency regularization to handle the large scale variance problem. (2) SED alleviates the noise problem from the False Negative samples and inferior localization precision. (3) A re-weighting strategy can implicitly screen the potential foreground regions of the unlabeled data to reduce the effect of class imbalance. Extensive experiments show that SED consistently outperforms the recent state-of-the-art methods on different datasets with significant margins. For example, it surpasses the supervised counterpart by more than 10 mAP when using 5% and 10% labeled data on MS-COCO.
Qiushan Guo, Yao Mu 0001, Jianyu Chen 0002, Yizhou Yu, Ping Luo 0002
CVPR5
2022 Attribute Surrogates Learning and Spectral Tokens Pooling in Transformers for Few-shot Learning
abstract
This paper presents new hierarchically cascaded transformers that can improve data efficiency through attribute surrogates learning and spectral tokens pooling. Vision transformers have recently been thought of as a promising alternative to convolutional neural networks for visual recognition. But when there is no sufficient data, it gets stuck in overfitting and shows inferior performance. To improve data efficiency, we propose hierarchically cascaded transformers that exploit intrinsic image structures through spectral tokens pooling and optimize the learnable parameters through latent attribute surrogates. The intrinsic image structure is utilized to reduce the ambiguity between foreground content and background noise by spectral tokens pooling. And the attribute surrogate learning scheme is designed to benefit from the rich visual information in image-label pairs instead of simple visual concepts assigned by their labels. Our Hierarchically Cascaded Transformers, called HCTransformers, is built upon a self-supervised learning framework DINO and is tested on several popular few-shot learning benchmarks. In the inductive setting, HCTransformers surpass the DINO baseline by a large margin of 9.7% 5-way 1-shot accuracy and 9.17% 5-way 5-shot accuracy on miniImageNet, which demonstrates HCTransformers are efficient to extract discriminative features. Also, HCTransformers show clear advantages over SOTA few-shot classification methods in both 5-way 1-shot and 5-way 5-shot settings on four popular benchmark datasets, including miniImageNet, tieredImageNet, FC100, and CIFAR-FS. The trained weights and codes are available at https://github.com/StomachCold/HCTransformers.
Yangji He, Weihan Liang, Dongyang Zhao, Weifeng Ge, Yizhou Yu
CVPR6
2022 Neighborhood Collective Estimation for Noisy Label Identification and Correction
Jichang Li, Guanbin Li, Feng Liu 0036, Yizhou Yu
ECCV (24)4
2022 One-Shot Medical Landmark Localization by Edge-Guided Transform and Noisy Landmark Refinement
Ping Gong 0002, Chunyu Wang 0001, Yizhou Yu, Yizhou Wang 0001
ECCV (21)4
2022 Centrality and Consistency: Two-Stage Clean Samples Identification for Learning with Instance-Dependent Noisy Labels
Ganlong Zhao, Guanbin Li, Yipeng Qin, Feng Liu 0036, Yizhou Yu
ECCV (25)5
2022 Disentangling Disease-related Representation from Obscure for Disease Prediction
abstract
Disease-related representations play a crucial role in image-based disease prediction such as cancer diagnosis, due to its considerable generalization capacity. However, it is still a challenge to identify lesion characteristics in obscured images, as many lesions are obscured by other tissues. In this paper, to learn the representations for identifying obscured lesions, we propose a disentanglement learning strategy under the guidance of alpha blending generation in an encoder-decoder framework (DAB-Net). Specifically, we take mammogram mass benign/malignant classification as an example. In our framework, composite obscured mass images are generated by alpha blending and then explicitly disentangled into disease-related mass features and interference glands features. To achieve disentanglement learning, features of these two parts are decoded to reconstruct the mass and the glands with corresponding reconstruction losses, and only disease-related mass features are fed into the classifier for disease prediction. Experimental results on one public dataset DDSM and three in-house datasets demonstrate that the proposed strategy can achieve state-of-the-art performance. DAB-Net achieves substantial improvements of 3.9%~4.4% AUC in obscured cases. Besides, the visualization analysis shows the model can better disentangle the mass and glands in the obscured image, suggesting the effectiveness of our solution in exploring the hidden characteristics in this challenging problem.
Chu-ran Wang, Fandong Zhang, Fangwei Zhong, Yizhou Yu, Yizhou Wang 0001
ICML5
2022 Computer-Aided Tuberculosis Diagnosis with Attribute Reasoning Assistance
Chengwei Pan, Gangming Zhao, Junjie Fang, Baolian Qi, Chaowei Fang, Dingwen Zhang, Jinpeng Li 0002, Yizhou Yu
MICCAI (1)9
2022 Mix and Reason: Reasoning over Semantic Topology with Data Mixing for Domain Generalization
abstract
Domain generalization (DG) enables generalizing a learning machine from multiple seen source domains to an unseen target one. The general objective of DG methods is to learn semantic representations that are independent of domain labels, which is theoretically sound but empirically challenged due to the complex mixture of common and domain-specific factors. Although disentangling the representations into two disjoint parts has been gaining momentum in DG, the strong presumption over the data limits its efficacy in many real-world scenarios. In this paper, we propose Mix and Reason (MiRe), a new DG framework that learns semantic representations via enforcing the structural invariance of semantic topology. MiRe consists of two key components, namely, Category-aware Data Mixing (CDM) and Adaptive Semantic Topology Refinement (ASTR). CDM mixes two images from different domains in virtue of activation maps generated by two complementary classification losses, making the classifier focus on the representations of semantic objects. ASTR introduces relation graphs to represent semantic topology, which is progressively refined via the interactions between local feature aggregation and global cross-domain relational reasoning. Experiments on multiple DG benchmarks validate the effectiveness and robustness of the proposed MiRe.
Chaoqi Chen, Luyao Tang, Feng Liu 0036, Gangming Zhao, Yue Huang 0001, Yizhou Yu
NeurIPS6
2022 MASS: Modality-collaborative semi-supervised segmentation by exploiting cross-modal consistency from unpaired CT and MRI images
Feng Liu 0036, Jiansen Guo, Liansheng Wang 0002, Yizhou Yu
Medical Image Anal.6
2022 M3Net: A multi-scale multi-view framework for multi-phase pancreas segmentation based on cross-phase non-local attention
Taiping Qu, Xiheng Wang, Chaowei Fang, Jinrong Qu, Xiuli Li, Huadan Xue, Yizhou Yu, Zhengyu Jin
Medical Image Anal.10
2022 Hierarchical deep network with uncertainty-aware semi-supervised learning for vessel segmentation
Chenxin Li, Wenao Ma, Liyan Sun, Xinghao Ding, Yue Huang 0001, Guisheng Wang, Yizhou Yu
Neural Comput. Appl.7
2022 A teacher-student framework for liver and tumor segmentation under mixed supervision from abdominal CT scans
Liyan Sun, Jianxiong Wu, Xinghao Ding, Yue Huang 0001, Zhong Chen 0005, Guisheng Wang, Yizhou Yu
Neural Comput. Appl.7
2022 Act Like a Radiologist: Towards Reliable Multi-View Correspondence Reasoning for Mammogram Mass Detection
abstract
Mammogram mass detection is crucial for diagnosing and preventing the breast cancers in clinical practice. The complementary effect of multi-view mammogram images provides valuable information about the breast anatomical prior structure and is of great significance in digital mammography interpretation. However, unlike radiologists who can utilize the natural reasoning ability to identify masses based on multiple mammographic views, how to endow the existing object detection models with the capability of multi-view reasoning is vital for decision-making in clinical diagnosis but remains the boundary to explore. In this paper, we propose an anatomy-aware graph convolutional network (AGN), which is tailored for mammogram mass detection and endows existing detection methods with multi-view reasoning ability. The proposed AGN consists of three steps. First, we introduce a bipartite graph convolutional network (BGN) to model the intrinsic geometric and semantic relations of ipsilateral views. Second, considering that the visual asymmetry of bilateral views is widely adopted in clinical practice to assist the diagnosis of breast lesions, we propose an inception graph convolutional network (IGN) to model the structural similarities of bilateral views. Finally, based on the constructed graphs, the multi-view information is propagated through nodes methodically, which equips the features learned from the examined view with multi-view reasoning ability. Experiments on two standard benchmarks reveal that AGN significantly exceeds the state-of-the-art performance. Visualization results show that AGN provides interpretable visual cues for clinical diagnosis.
Fandong Zhang, Chaoqi Chen, Yizhou Wang 0001, Yizhou Yu
IEEE Trans. Pattern Anal. Mach. Intell.6
2022 Diagnose Like a Radiologist: Hybrid Neuro-Probabilistic Reasoning for Attribute-Based Medical Image Diagnosis
abstract
During clinical practice, radiologists often use attributes, e.g., morphological and appearance characteristics of a lesion, to aid disease diagnosis. Effectively modeling attributes as well as all relationships involving attributes could boost the generalization ability and verifiability of medical image diagnosis algorithms. In this paper, we introduce a hybrid neuro-probabilistic reasoning algorithm for verifiable attribute-based medical image diagnosis. There are two parallel branches in our hybrid algorithm, a Bayesian network branch performing probabilistic causal relationship reasoning and a graph convolutional network branch performing more generic relational modeling and reasoning using a feature representation. Tight coupling between these two branches is achieved via a cross-network attention mechanism and the fusion of their classification results. We have successfully applied our hybrid reasoning algorithm to two challenging medical image diagnosis tasks. On the LIDC-IDRI benchmark dataset for benign-malignant classification of pulmonary nodules in CT images, our method achieves a new state-of-the-art accuracy of 95.36% and an AUC of 96.54%. Our method also achieves a 3.24% accuracy improvement on an in-house chest X-ray image dataset for tuberculosis diagnosis. Our ablation study indicates that our hybrid algorithm achieves a much better generalization performance than a pure neural network architecture under very limited training data.
Gangming Zhao, Quanlong Feng, Chaoqi Chen, Yizhou Yu
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 3D Graph-Connectivity Constrained Network for Hepatic Vessel Segmentation
abstract
Segmentation of hepatic vessels from 3D CT images is necessary for accurate diagnosis and preoperative planning for liver cancer. However, due to the low contrast and high noises of CT images, automatic hepatic vessel segmentation is a challenging task. Hepatic vessels are connected branches containing thick and thin blood vessels, showing an important structural characteristic or a prior: the connectivity of blood vessels. However, this is rarely applied in existing methods. In this paper, we segment hepatic vessels from 3D CT images by utilizing the connectivity prior. To this end, a graph neural network (GNN) used to describe the connectivity prior of hepatic vessels is integrated into a general convolutional neural network (CNN). Specifically, a graph attention network (GAT) is first used to model the graphical connectivity information of hepatic vessels, which can be trained with the vascular connectivity graph constructed directly from the ground truths. Second, the GAT is integrated with a lightweight 3D U-Net by an efficient mechanism called the plug-in mode, in which the GAT is incorporated into the U-Net as a multi-task branch and is only used to supervise the training procedure of the U-Net with the connectivity prior. The GAT will not be used in the inference stage, and thus will not increase the hardware and time costs of the inference stage compared with the U-Net. Therefore, hepatic vessel segmentation can be well improved in an efficient mode. Extensive experiments on two public datasets show that the proposed method is superior to related works in accuracy and connectivity of hepatic vessel segmentation.
Ruikun Li 0004, Yi-Jie Huang, Huai Chen, Yizhou Yu, Dahong Qian, Lisheng Wang
IEEE J. Biomed. Health Informatics5
2022 GREN: Graph-Regularized Embedding Network for Weakly-Supervised Disease Localization in X-Ray Images
abstract
Locating diseases in chest X-ray images with few careful annotations saves large human effort. Recent works approached this task with innovative weakly-supervised algorithms such as multi-instance learning (MIL) and class activation maps (CAM), however, these methods often yield inaccurate or incomplete regions. One of the reasons is the neglection of the pathological implications hidden in the relationship across anatomical regions within each image and the relationship across images. In this paper, we argue that the cross-region and cross-image relationship, as contextual and compensating information, is vital to obtain more consistent and integral regions. To model the relationship, we propose the Graph Regularized Embedding Network (GREN), which leverages the intra-image and inter-image information to locate diseases on chest X-ray images. GREN uses a pre-trained U-Net to segment the lung lobes, and then models the intra-image relationship between the lung lobes using an intra-image graph to compare different regions. Meanwhile, the relationship between in-batch images is modeled by an inter-image graph to compare multiple images. This process mimics the training and decision-making process of a radiologist: comparing multiple regions and images for diagnosis. In order for the deep embedding layers of the neural network to retain structural information (important in the localization task), we use the Hash coding and Hamming distance to compute the graphs, which are used as regularizers to facilitate training. By means of this, our approach achieves the state-of-the-art result on NIH chest X-ray dataset for weakly-supervised disease localization. Our codes are accessible online.
Baolian Qi, Gangming Zhao, Changde Du, Chengwei Pan, Yizhou Yu, Jinpeng Li 0002
IEEE J. Biomed. Health Informatics6
2022 Harmonizing Pathological and Normal Pixels for Pseudo-Healthy Synthesis
abstract
Synthesizing a subject-specific pathology-free image from a pathological image is valuable for algorithm development and clinical practice. In recent years, several approaches based on the Generative Adversarial Network (GAN) have achieved promising results in pseudo-healthy synthesis. However, the discriminator (i.e., a classifier) in the GAN cannot accurately identify lesions and further hampers from generating admirable pseudo-healthy images. To address this problem, we present a new type of discriminator, the segmentor, to accurately locate the lesions and improve the visual quality of pseudo-healthy images. Then, we apply the generated images into medical image enhancement and utilize the enhanced results to cope with the low contrast problem existing in medical image segmentation. Furthermore, a reliable metric is proposed by utilizing two attributes of label noise to measure the health of synthetic images. Comprehensive experiments on the T2 modality of BraTS demonstrate that the proposed method substantially outperforms the state-of-the-art methods. The method achieves better performance than the existing methods with only 30% of the training data. The effectiveness of the proposed method is also demonstrated on the LiTS and the T1 modality of BraTS. The code and the pre-trained model of this study are publicly available at https://github.com/Au3C2/Generator-Versus-Segmentor.
Yihong Zhuang, Liyan Sun, Yue Huang 0001, Xinghao Ding, Guisheng Wang, Lin Yang 0002, Yizhou Yu
IEEE Trans. Medical Imaging9
2022 GraVIS: Grouping Augmented Views From Independent Sources for Dermatology Analysis
Chixiang Lu, Liansheng Wang 0002, Yizhou Yu
IEEE Trans. Medical Imaging4
2022 Structure-aware Meta-fusion for Image Super-resolution
abstract
There are two main categories of image super-resolution algorithms: distortion oriented and perception oriented. Recent evidence shows that reconstruction accuracy and perceptual quality are typically in disagreement with each other. In this article, we present a new image super-resolution framework that is capable of striking a balance between distortion and perception. The core of our framework is a deep fusion network capable of generating a final high-resolution image by fusing a pair of deterministic and stochastic images using spatially varying weights. To make a single fusion model produce images with varying degrees of stochasticity, we further incorporate meta-learning into our fusion network. Once equipped with the kernel produced by a kernel prediction module, our meta fusion network is able to produce final images at any desired level of stochasticity. Experimental results indicate that our meta fusion network outperforms existing state-of-the-art SISR algorithms on widely used datasets, including PIRM-val, DIV2K-val, Set5, Set14, Urban100, Manga109, and B100. In addition, it is capable of producing high-resolution images that achieve low distortion and high perceptual quality simultaneously.
Bingchen Gong, Yizhou Yu
ACM Trans. Multim. Comput. Commun. Appl.3
2022 SAniHead: Sketching Animal-Like 3D Character Heads Using a View-Surface Collaborative Mesh Generative Network
abstract
In the game and film industries, modeling 3D heads plays a very important role in designing characters. Although human head modeling has been researched for a long time, few works have focused on animal-like heads, which are of more diverse shapes and richer geometric details. In this article, we present SAniHead, an interactive system for creating animal-like heads with a mesh representation from dual-view sketches. Our core technical contribution is a view-surface collaborative mesh generative network. Initially, a graph convolutional neural network (GCNN) is trained to learn the deformation of a template mesh to fit the shape of sketches, giving rise to a coarse model. It is then projected into vertex maps where image-to-image translation networks are performed for detail inference. After back-projecting the inferred details onto the meshed surface, a new GCNN is trained for further detail refinement. The modules of view-based detail inference and surface-based detail refinement are conducted in an alternating cascaded fashion, collaboratively improving the model. A refinement sketching interface is also implemented to support direct mesh manipulation. Experimental results show the superiority of our approach and the usability of our interactive system. Our work also contributes a 3D animal head dataset with corresponding line drawings.
Dong Du 0002, Xiaoguang Han 0001, Hongbo Fu 0001, Feiyang Wu, Yizhou Yu, Shuguang Cui, Ligang Liu 0001
IEEE Trans. Vis. Comput. Graph.5
2021 I3Net: Implicit Instance-Invariant Network for Adapting One-Stage Object Detectors
abstract
Recent works on two-stage cross-domain detection have widely explored the local feature patterns to achieve more accurate adaptation results. These methods heavily rely on the region proposal mechanisms and ROI-based instance-level features to design fine-grained feature alignment modules with respect to the foreground objects. However, for one-stage detectors, it is hard or even impossible to obtain explicit instance-level features in the detection pipelines. Motivated by this, we propose an Implicit Instance-Invariant Network (I3Net), which is tailored for adapting one-stage detectors and implicitly learns instance-invariant features via exploiting the natural characteristics of deep features in different layers. Specifically, we facilitate the adaptation from three aspects: (1) Dynamic and Class-Balanced Reweighting (DCBR) strategy, which considers the coexistence of intra-domain and intra-class variations to assign larger weights to those sample-scarce categories and easy-to-adapt samples; (2) Category-aware Object Pattern Matching (COPM) module, which boosts the cross-domain foreground objects matching guided by the categorical information and suppresses the uninformative background features; (3) Regularized Joint Category Alignment (RJCA) module, which jointly enforces the category alignment at different domain-specific layers with a consistency regularization. Experiments reveal that I3Net exceeds the state-of-the-art performance on benchmark datasets.
Chaoqi Chen, Zebiao Zheng, Yue Huang 0001, Xinghao Ding, Yizhou Yu
CVPR5
2021 Cross-Domain Adaptive Clustering for Semi-Supervised Domain Adaptation
abstract
In semi-supervised domain adaptation, a few labeled samples per class in the target domain guide features of the remaining target samples to aggregate around them. However, the trained model cannot produce a highly discriminative feature representation for the target domain because the training data is dominated by labeled samples from the source domain. This could lead to disconnection between the labeled and unlabeled target samples as well as misalignment between unlabeled target samples and the source domain. In this paper, we propose a novel approach called Cross-domain Adaptive Clustering to address this problem. To achieve both inter-domain and intra-domain adaptation, we first introduce an adversarial adaptive clustering loss to group features of unlabeled target data into clusters and perform cluster-wise feature alignment across the source and target domains. We further apply pseudo labeling to unlabeled samples in the target domain and retain pseudo-labels with high confidence. Pseudo labeling expands the number of "labeled" samples in each class in the target domain, and thus produces a more robust and powerful cluster core for each class to facilitate adversarial learning. Extensive experiments on benchmark datasets, including DomainNet, Office-Home and Office, demonstrate that our proposed approach achieves the state-of-the-art performance in semi-supervised domain adaptation.
Jichang Li, Guanbin Li, Yemin Shi 0001, Yizhou Yu
CVPR4
2021 Scene-Intuitive Agent for Remote Embodied Visual Grounding
abstract
Humans learn from life events to form intuitions towards the understanding of visual environments and languages. Envision that you are instructed by a high-level instruction, "Go to the bathroom in the master bedroom and replace the blue towel on the left wall", what would you possibly do to carry out the task? Intuitively, we comprehend the semantics of the instruction to form an overview of where a bathroom is and what a blue towel is in mind; then, we navigate to the target location by consistently matching the bathroom appearance in mind with the current scene. In this paper, we present an agent that mimics such human behaviors. Specifically, we focus on the Remote Embodied Visual Referring Expression in Real Indoor Environments task, called REVERIE, where an agent is asked to correctly localize a remote target object specified by a concise high-level natural language instruction, and propose a two-stage training pipeline. In the first stage, we pretrain the agent with two cross-modal alignment sub-tasks, namely the Scene Grounding task and the Object Grounding task. The agent learns where to stop in the Scene Grounding task and what to attend to in the Object Grounding task respectively. Then, to generate action sequences, we propose a memory-augmented attentive action decoder to smoothly fuse the pre-trained vision and language representations with the agent’s past memory experiences. Without bells and whistles, experimental results show that our method outperforms previous state-of-the-art(SOTA) significantly, demonstrating the effectiveness of our method.
Xiangru Lin, Guanbin Li, Yizhou Yu
CVPR3
2021 Refer-It-in-RGBD: A Bottom-Up Approach for 3D Visual Grounding in RGBD Images
abstract
Grounding referring expressions in RGBD image has been an emerging field. We present a novel task of 3D visual grounding in single-view RGBD image where the referred objects are often only partially scanned due to occlusion. In contrast to previous works that directly generate object proposals for grounding in the 3D scenes, we propose a bottom-up approach to gradually aggregate content-aware information, effectively addressing the challenge posed by the partial geometry. Our approach first fuses the language and the visual features at the bottom level to generate a heatmap that coarsely localizes the relevant regions in the RGBD image. Then our approach conducts an adaptive feature learning based on the heatmap and performs the object-level matching with another visio-linguistic fusion to finally ground the referred object. We evaluate the proposed method by comparing to the state-of-the-art methods on both the RGBD images extracted from the ScanRefer dataset and our newly collected SUNRefer dataset. Experiments show that our method outperforms the previous methods by a large margin (by 11.2% and 15.6% [email protected]) on both datasets.
Haolin Liu 0004, Anran Lin, Xiaoguang Han 0001, Yizhou Yu, Shuguang Cui
CVPR5
2021 Coarse-To-Fine Domain Adaptive Semantic Segmentation With Photometric Alignment and Category-Center Regularization
abstract
Unsupervised domain adaptation (UDA) in semantic segmentation is a fundamental yet promising task relieving the need for laborious annotation works. However, the domain shifts/discrepancies problem in this task compromise the final segmentation performance. Based on our observation, the main causes of the domain shifts are differences in imaging conditions, called image-level domain shifts, and differences in object category configurations called category-level domain shifts. In this paper, we propose a novel UDA pipeline that unifies image-level alignment and category-level feature distribution regularization in a coarse-to-fine manner. Specifically, on the coarse side, we propose a photometric alignment module that aligns an image in the source domain with a reference image from the target domain using a set of image-level operators; on the fine side, we propose a category-oriented triplet loss that imposes a soft constraint to regularize category centers in the source domain and a self-supervised consistency regularization method in the target domain. Experimental results show that our proposed pipeline improves the generalization capability of the final segmentation model and significantly outperforms all previous state-of-the-arts.
Xiangru Lin, Zifeng Wu, Yizhou Yu
CVPR4
2021 Bottom-Up Shift and Reasoning for Referring Image Segmentation
abstract
Referring image segmentation aims to segment the referent that is the corresponding object or stuff referred by a natural language expression in an image. Its main challenge lies in how to effectively and efficiently differentiate between the referent and other objects of the same category as the referent. In this paper, we tackle the challenge by jointly performing compositional visual reasoning and accurate segmentation in a single stage via the proposed novel Bottom-Up Shift (BUS) and Bidirectional Attentive Refinement (BIAR) modules. Specifically, BUS progressively locates the referent along hierarchical reasoning steps implied by the expression. At each step, it locates the corresponding visual region by disambiguating between similar regions, where the disambiguation bases on the relationships between regions. By the explainable visual reasoning, BUS explicitly aligns linguistic components with visual regions so that it can identify all the mentioned entities in the expression. BIAR fuses multi-level features via a two-way attentive message passing, which captures the visual details relevant to the referent to refine segmentation results. Experimental results demonstrate that the proposed method consisting of BUS and BIAR modules, can not only consistently surpass all existing state-of-the-art algorithms across common benchmark datasets but also visualize interpretable reasoning steps for stepwise segmentation. Code is available at https://github.com/incredibleXM/BUSNet.
Sibei Yang, Guanbin Li, Yizhou Yu
CVPR5
2021 Dual Bipartite Graph Learning: A General Approach for Domain Adaptive Object Detection
abstract
Domain Adaptive Object Detection (DAOD) relieves the reliance on large-scale annotated data by transferring the knowledge learned from a labeled source domain to a new unlabeled target domain. Recent DAOD approaches resort to local feature alignment in virtue of domain adversarial training in conjunction with the ad-hoc detection pipelines to achieve feature adaptation. However, these methods are limited to adapt the specific types of object detectors and do not explore the cross-domain topological relations. In this paper, we first formulate DAOD as an open-set domain adaptation problem in which foregrounds (pixel or region) can be seen as the “known class”, while backgrounds (pixel or region) are referred to as the “unknown class”. To this end, we present a new and general perspective for DAOD named Dual Bipartite Graph Learning (DBGL), which captures the cross-domain interactions on both pixel-level and semantic-level via increasing the distinction between foregrounds and backgrounds and modeling the cross-domain dependencies among different semantic categories. Experiments reveal that the proposed DBGL in conjunction with one-stage and two-stage detectors exceeds the state-of-the-art performance on standard DAOD benchmarks.
Chaoqi Chen, Jiongcheng Li, Zebiao Zheng, Yue Huang 0001, Xinghao Ding, Yizhou Yu
ICCV6
2021 ME-PCN: Point Completion Conditioned on Mask Emptiness
abstract
Point completion refers to completing the missing geometries of an object from incomplete observations. Mainstream methods predict the missing shapes by decoding a global feature learned from the input point cloud, which often leads to deficient results in preserving topology consistency and surface details. In this work, we present MEPCN, a point completion network that leverages emptiness in 3D shape space. Given a single depth scan, previous methods often encode the occupied partial shapes while ignoring the empty regions (e.g. holes) in depth maps. In contrast, we argue that these ‘emptiness’ clues indicate shape boundaries that can be used to improve topology representation and detail granularity on surfaces. Specifically, our ME-PCN encodes both the occupied point cloud and the neighboring ‘empty points’. It estimates coarse-grained but complete and reasonable surface points in the first stage, followed by a refinement stage to produce fine-grained surface details. Comprehensive experiments verify that our ME-PCN presents better qualitative and quantitative performance against the state-of-the-art. Besides, we further prove that our ‘emptiness’ design is lightweight and easy to embed in existing methods, which shows consistent effectiveness in improving the CD and EMD scores.
Bingchen Gong, Yinyu Nie, Yiqun Lin, Xiaoguang Han 0001, Yizhou Yu
ICCV5
2021 GraphFPN: Graph Feature Pyramid Network for Object Detection
abstract
Feature pyramids have been proven powerful in image understanding tasks that require multi-scale features. State-of-the-art methods for multi-scale feature learning focus on performing feature interactions across space and scales using neural networks with a fixed topology. In this paper, we propose graph feature pyramid networks that are capable of adapting their topological structures to varying intrinsic image structures, and supporting simultaneous feature interactions across all scales. We first define an image specific superpixel hierarchy for each input image to represent its intrinsic image structures. The graph feature pyramid network inherits its structure from this superpixel hierarchy. Contextual and hierarchical layers are designed to achieve feature interactions within the same scale and across different scales. To make these layers more powerful, we introduce two types of local channel attention for graph neural networks by generalizing global channel attention for convolutional neural networks. The proposed graph feature pyramid network can enhance the multiscale features from a convolutional feature pyramid network.We evaluate our graph feature pyramid network in the object detection task by integrating it into the Faster R-CNN algorithm. The modified algorithm outperforms not only previous state-of-the-art feature pyramid based methods with a clear margin but also other popular detection methods on both MS-COCO 2017 validation and test datasets.
Gangming Zhao, Weifeng Ge, Yizhou Yu
ICCV3
2021 Multi-scale Matching Networks for Semantic Correspondence
abstract
Deep features have been proven powerful in building accurate dense semantic correspondences in various previous works. However, the multi-scale and pyramidal hierarchy of convolutional neural networks has not been well studied to learn discriminative pixel-level features for semantic correspondence. In this paper, we propose a multi-scale matching network that is sensitive to tiny semantic differences between neighboring pixels. We follow the coarse-to-fine matching strategy and build a top-down feature and matching enhancement scheme that is coupled with the multi-scale hierarchy of deep convolutional neural networks. During feature enhancement, intra-scale enhancement fuses same-resolution feature maps from multiple layers together via local self-attention and cross-scale enhancement hallucinates higher-resolution feature maps along the top-down pathway. Besides, we learn complementary matching details at different scales thus the overall matching score is refined by features of different semantic levels gradually. Our multi-scale matching network can be trained end-to-end easily with few additional learnable parameters. Experimental results demonstrate that the proposed method achieves state-of-the-art performance on three popular benchmarks with high computational efficiency. The code has been released at https://github.com/wintersun661/MMNet.
Dongyang Zhao, Zhenghao Ji, Gangming Zhao, Weifeng Ge, Yizhou Yu
ICCV6
2021 Preservational Learning Improves Self-supervised Medical Image Models by Reconstructing Diverse Contexts
abstract
Preserving maximal information is one of principles of designing self-supervised learning methodologies. To reach this goal, contrastive learning adopts an implicit way which is contrasting image pairs. However, we believe it is not fully optimal to simply use the contrastive estimation for preservation. Moreover, it is necessary and complemental to introduce an explicit solution to preserve more information. From this perspective, we introduce Preservational Learning to reconstruct diverse image contexts in order to preserve more information in learned representations. Together with the contrastive loss, we present Preservational Contrastive Representation Learning (PCRL) for learning self-supervised medical representations. PCRL provides very competitive results under the pretraining-finetuning protocol, outperforming both self-supervised and supervised counterparts in 5 classification/segmentation tasks substantially. Codes are available at https://github.com/Luchixiang/PCRL.
Chixiang Lu, Sibei Yang, Xiaoguang Han 0001, Yizhou Yu
ICCV5
2021 CCF-Net: Composite Context Fusion Network with Inter-Slice Correlative Fusion for Multi-Disease Lesion Detection
abstract
Detecting lesions from computed tomography (CT) scans relies on two aspects of the input: intra-slice texture information from the key slice and inter-slice structural context information from the adjacent slices. However, most existing methods ignore the correlation and complementarity between texture and structural information resulting in unexpected loss of performance. In this paper, a novel Composite Context Fusion Network (CCF-Net) is proposed to jointly model intra-slice and inter-slice features so as to prove the effectiveness of the two-steam framework. To extract both texture and structural information, two streams of 2D and 3D convolutional modules are employed in each stage. Moreover, a Composite Fusion architecture equipped with Inter-slice Correlative Fusion (ICF) modules is proposed to achieve stage-by-stage feature fusion in order to excavate and exchange information between texture-aware and context-aware features. Extensive experiments show that the proposed CCF-Net is able to achieve state-of-the-art detection performance on the multi-disease CT lesion detection task and significantly surpass the baseline methods.1
Jiechao Ma, Shu Zhang 0001, Yemin Shi 0001, Junge Zhang, Kaiqi Huang, Yizhou Yu
ICIP7
2021 Noise2Grad: Extract Image Noise to Denoise
abstract
In many image denoising tasks, the difficulty of collecting noisy/clean image pairs limits the application of supervised CNNs. We consider such a case in which paired data and noise statistics are not accessible, but unpaired noisy and clean images are easy to collect. To form the necessary supervision, our strategy is to extract the noise from the noisy image to synthesize new data. To ease the interference of the image background, we use a noise removal module to aid noise extraction. The noise removal module first roughly removes noise from the noisy image, which is equivalent to excluding much background information. A noise approximation module can therefore easily extract a new noise map from the removed noise to match the gradient of the noisy input. This noise map is added to a random clean image to synthesize a new data pair, which is then fed back to the noise removal module to correct the noise removal process. These two modules cooperate to extract noise finely. After convergence, the noise removal module can remove noise without damaging other background details, so we use it as our final denoising network. Experiments show that the denoising performance of the proposed method is competitive with other supervised CNNs.
Huangxing Lin, Yihong Zhuang, Yue Huang 0001, Xinghao Ding, Yizhou Yu
IJCAI6
2021 Symmetry-Enhanced Attention Network for Acute Ischemic Infarct Segmentation with Non-contrast CT Images
Kongming Liang, Kai Han 0010, Xiuli Li, Xiaoqing Cheng, Yizhou Wang 0001, Yizhou Yu
MICCAI (7)7
2021 Improved Brain Lesion Segmentation with Anatomical Priors from Healthy Subjects
Xiangzhu Zeng, Kongming Liang, Yizhou Yu, Chuyang Ye
MICCAI (1)4
2021 CA-Net: Leveraging Contextual Features for Lung Cancer Prediction
Mingzhou Liu 0001, Fandong Zhang, Xinwei Sun 0001, Yizhou Yu, Yizhou Wang 0001
MICCAI (5)4
2021 Fast Magnetic Resonance Imaging on Regions of Interest: From Sensing to Reconstruction
Liyan Sun, Xinghao Ding, Yue Huang 0001, Yizhou Yu
MICCAI (6)6
2021 DAE-GCN: Identifying Disease-Related Features for Disease Prediction
Chu-ran Wang, Xinwei Sun 0001, Fandong Zhang, Yizhou Yu, Yizhou Wang 0001
MICCAI (5)4
2021 Generator Versus Segmentor: Pseudo-healthy Synthesis
Chenxin Li, Liyan Sun, Yihong Zhuang, Yue Huang 0001, Xinghao Ding, Yizhou Yu
MICCAI (6)9
2021 CarveMix: A Simple Data Augmentation Method for Brain Lesion Segmentation
Xinru Zhang 0001, Ni Ou, Xiangzhu Zeng, Xiaoliang Xiong, Yizhou Yu, Chuyang Ye
MICCAI (1)6
2021 Self-supervised Correction Learning for Semi-supervised Biomedical Image Segmentation
Ruifei Zhang, Sishuo Liu, Yizhou Yu, Guanbin Li
MICCAI (2)3
2021 Cross-Modal Self-Attention with Multi-Task Pre-Training for Medical Visual Question Answering
abstract
Due to the severe lack of labeled data, existing methods of medical visual question answering usually rely on transfer learning to obtain effective image feature representation and use cross-modal fusion of visual and linguistic features to achieve question-related answer prediction. These two phases are performed independently and without considering the compatibility and applicability of the pre-trained features for cross-modal fusion. Thus, we reformulate image feature pre-training as a multi-task learning paradigm and witness its extraordinary superiority, forcing it to take into account the applicability of features for the specific image comprehension task. Furthermore, we introduce a cross-modal self-attention~(CMSA) module to selectively capture the long-range contextual relevance for more effective fusion of visual and linguistic features. Experimental results demonstrate that the proposed method outperforms existing state-of-the-art methods. Our code and models are available at https://github.com/haifangong/CMSA-MTPT-4-MedicalVQA.
Haifan Gong, Guanqi Chen, Sishuo Liu, Yizhou Yu, Guanbin Li
ICMR4
2021 Learning Part Generation and Assembly for Sketching Man-Made Objects
abstract
Abstract Modeling 3D objects on existing software usually requires a heavy amount of interactions, especially for users who lack basic knowledge of 3D geometry. Sketch‐based modeling is a solution to ease the modelling procedure and thus has been researched for decades. However, modelling a man‐made shape with complex structures remains challenging. Existing methods adopt advanced deep learning techniques to map holistic sketches to 3D shapes. They are still bottlenecked to deal with complicated topologies. In this paper, we decouple the task of sketch2shape into a part generation module and a part assembling module, where deep learning methods are leveraged for the implementation of both modules. By changing the focus from holistic shapes to individual parts, it eases the learning process of the shape generator and guarantees high‐quality outputs. With the learned automated part assembler, users only need a little manual tuning to obtain a desired layout. Extensive experiments and user studies demonstrate the usefulness of our proposed system.
Dong Du 0002, Heming Zhu, Yinyu Nie, Xiaoguang Han 0001, Shuguang Cui, Yizhou Yu, Ligang Liu 0001
Comput. Graph. Forum6
2021 Instance-level salient object segmentation
Guanbin Li, Pengxiang Yan, Yuan Xie 0004, Guisheng Wang, Liang Lin 0004, Yizhou Yu
Comput. Vis. Image Underst.6
2021 A Nonintrusive Elderly Home Monitoring System
abstract
Home anomaly monitoring is crucial for the elderly who live alone. A number of IoT-based home monitoring systems have been available, but most rely on privacy-intrusive cameras. With more and more concerns on privacy and security of human data, anomaly detection based on nonintrusive IoT devices becomes more desirable. Considering the elderly consumers, a low-cost system with good detection accuracy is further critical for the system's acceptability by elderly users. We propose a smart home monitoring system for living-alone senior citizens, relying on carefully designed, low-cost infrared sensor devices, as well as a cloud-based data processing and anomaly detection platform. Our PIR sensor device is effective in continuous monitoring of motion data in a user's apartment, and an open-hardware software platform is devised to support sensors manufactured by various vendors in the IoT system, all for cost reduction purpose. For privacy preservation, we encrypt collected data and store data indices in a blockchain system, to achieve efficient data access control and auditing. For motion anomaly detection, we propose a simple but effective environment adaptation method to work with the one-class support vector machine (OCSVM) method. Experiments driven by real-world traces show good reliability, accuracy, and efficiency of our system.
Le Fang 0003, Yu Wu 0010, Chuan Wu 0001, Yizhou Yu
IEEE Internet Things J.4
2021 Identification of pediatric respiratory diseases using a fine-grained diagnosis system
Zhongzhi Yu, Yemin Shi 0001, Yingshuo Wang, Zheming Li, Yonggen Zhao, Fenglei Sun, Yizhou Yu, Qiang Shu
J. Biomed. Informatics9
2021 Compare and contrast: Detecting mammographic soft-tissue lesions with C2-Net
Changsheng Zhou, Fandong Zhang, Qianyi Zhang, Fugeng Sheng, Wanhua Liu, Yizhou Wang 0001, Yizhou Yu, Guangming Lu 0001
Medical Image Anal.11
2021 SSMD: Semi-Supervised medical image detection with adaptive consistency and heterogeneous perturbation
Chengdi Wang, Haofeng Li, Shu Zhang 0001, Weimin Li 0003, Yizhou Yu
Medical Image Anal.7
2021 Relationship-Embedded Representation Learning for Grounding Referring Expressions
abstract
Grounding referring expressions in images aims to locate the object instance in an image described by a referring expression. It involves a joint understanding of natural language and image content, and is essential for a range of visual tasks related to human-computer interaction. As a language-to-vision matching task, the core of this problem is to not only extract all the necessary information (i.e., objects and the relationships among them) in both the image and referring expression, but also make full use of context information to align cross-modal semantic concepts in the extracted information. Unfortunately, existing work on grounding referring expressions fails to accurately extract multi-order relationships from the referring expression and associate them with the objects and their related contexts in the image. In this paper, we propose a cross-modal relationship extractor (CMRE) to adaptively highlight objects and relationships (spatial and semantic relations) related to the given expression with a cross-modal attention mechanism, and represent the extracted information as a language-guided visual relation graph. In addition, we propose a Gated Graph Convolutional Network (GGCN) to compute multimodal semantic contexts by fusing information from different modes and propagating multimodal information in the structured relation graph. Experimental results on three common benchmark datasets show that our Cross-Modal Relationship Inference Network, which consists of CMRE and GGCN, significantly surpasses all existing state-of-the-art methods.
Sibei Yang, Guanbin Li, Yizhou Yu
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 Cross-modality deep feature learning for brain tumor segmentation
Dingwen Zhang, Guohai Huang, Qiang Zhang 0020, Jungong Han, Junwei Han 0001, Yizhou Yu
Pattern Recognit.6
2021 Depthwise Nonlocal Module for Fast Salient Object Detection Using a Single Thread
abstract
Recently, deep convolutional neural networks have achieved significant success in salient object detection. However, existing state-of-the-art methods require high-end GPUs to achieve real-time performance, which makes it hard to adapt to low cost or portable devices. Although generic network architectures have been proposed to speed up inference on mobile devices, they are tailored to the task of image classification or semantic segmentation, and struggle to capture intrachannel and interchannel correlations that are essential for contrast modeling in salient object detection. Motivated by the above observations, we design a new deep-learning algorithm for fast salient object detection. The proposed algorithm for the first time achieves competitive accuracy and high inference efficiency simultaneously with a single CPU thread. Specifically, we propose a novel depthwise nonlocal module (DNL), which implicitly models contrast via harvesting intrachannel and interchannel correlations in a self-attention manner. In addition, we introduce a depthwise nonlocal network architecture that incorporates both DNLs module and inverted residual blocks. The experimental results show that our proposed network attains very competitive accuracy on a wide range of salient object detection datasets while achieving state-of-the-art efficiency among all existing deep-learning-based algorithms.
Haofeng Li, Guanbin Li, Guanqi Chen, Liang Lin 0004, Yizhou Yu
IEEE Trans. Cybern.6
2021 Bilateral Asymmetry Guided Counterfactual Generating Network for Mammogram Classification
abstract
Mammogram benign or malignant classification with only image-level labels is challenging due to the absence of lesion annotations. Motivated by the symmetric prior that the lesions on one side of breasts rarely appear in the corresponding areas on the other side, we explore to answer a counterfactual question to identify the lesion areas. This counterfactual question means: given an image with lesions, how would the features have behaved if there were no lesions in the image? To answer this question, we derive a new theoretical result based on the symmetric prior. Specifically, by building a causal model that entails such a prior for bilateral images, we identify to optimize the distances in distribution between i) the counterfactual features and the target side's features in lesion-free areas; and ii) the counterfactual features and the reference side's features in lesion areas. To realize these optimizations for better benign/malignant classification, we propose a counterfactual generative network, which is mainly composed of Generator Adversarial Network and a prediction feedback mechanism, they are optimized jointly and prompt each other. Specifically, the former can further improve the classi?cation performance by generating counterfactual features to calculate lesion areas. On the other hand, the latter helps counterfactual generation by the supervision of classification loss. The utility of our method and the effectiveness of each module in our model can be verified by state-of-the-art performance on INBreast and an in-house dataset and ablation studies.
Chu-ran Wang, Jing Li 0091, Fandong Zhang, Xinwei Sun 0001, Hao Dong 0003, Yizhou Yu, Yizhou Wang 0001
IEEE Trans. Image Process.6
2021 Curriculum Feature Alignment Domain Adaptation for Epithelium-Stroma Classification in Histopathological Images
abstract
In recent years, deep learning methods have received more attention in epithelial-stroma (ES) classification tasks. Traditional deep learning methods assume that the training and test data have the same distribution, an assumption that is seldom satisfied in complex imaging procedures. Unsupervised domain adaptation (UDA) transfers knowledge from a labelled source domain to a completely unlabeled target domain, and is more suitable for ES classification tasks to avoid tedious annotation. However, existing UDA methods for this task ignore the semantic alignment across domains. In this paper, we propose a Curriculum Feature Alignment Network (CFAN) to gradually align discriminative features across domains through selecting effective samples from the target domain and minimizing intra-class differences. Specifically, we developed the Curriculum Transfer Strategy (CTS) and Adaptive Centroid Alignment (ACA) steps to train our model iteratively. We validated the method using three independent public ES datasets, and experimental results demonstrate that our method achieves better performance in ES classification compared with commonly used deep learning methods and existing deep domain adaptation methods.
Qi Qi 0005, Chaoqi Chen, Weiping Xie, Yue Huang 0001, Xinghao Ding, Yizhou Yu
IEEE J. Biomed. Health Informatics8
2021 A Structure-Aware Relation Network for Thoracic Diseases Detection and Segmentation
abstract
Instance level detection and segmentation of thoracic diseases or abnormalities are crucial for automatic diagnosis in chest X-ray images. Leveraging on constant structure and disease relations extracted from domain knowledge, we propose a structure-aware relation network (SAR-Net) extending Mask R-CNN. The SAR-Net consists of three relation modules: 1. the anatomical structure relation module encoding spatial relations between diseases and anatomical parts. 2. the contextual relation module aggregating clues based on query-key pair of disease RoI and lung fields. 3. the disease relation module propagating co-occurrence and causal relations into disease proposals. Towards making a practical system, we also provide ChestX-Det, a chest X-Ray dataset with instance-level annotations (boxes and masks). ChestX-Det is a subset of the public dataset NIH ChestX-ray14. It contains ~3500 images of 13 common disease categories labeled by three board-certified radiologists. We evaluate our SAR-Net on it and another dataset DR-Private. Experimental results show that it can enhance the strong baseline of Mask R-CNN with significant improvements. The ChestX-Det is released at https://github.com/Deepwise-AILab/ChestX-Det-Dataset.
Jingyu Liu 0004, Shu Zhang 0001, Dingwen Zhang, Yizhou Yu
IEEE Trans. Medical Imaging7
2021 Contralaterally Enhanced Networks for Thoracic Disease Detection
abstract
Identifying and locating diseases in chest X-rays are very challenging, due to the low visual contrast between normal and abnormal regions, and distortions caused by other overlapping tissues. An interesting phenomenon is that there exist many similar structures in the left and right parts of the chest, such as ribs, lung fields and bronchial tubes. This kind of similarities can be used to identify diseases in chest X-rays, according to the experience of broad-certificated radiologists. Aimed at improving the performance of existing detection methods, we propose a deep end-to-end module to exploit the contralateral context information for enhancing feature representations of disease proposals. First of all, under the guidance of the spine line, the spatial transformer network is employed to extract local contralateral patches, which can provide valuable context information for disease proposals. Then, we build up a specific module, based on both additive and subtractive operations, to fuse the features of the disease proposal and the contralateral patch. Our method can be integrated into both fully and weakly supervised disease detection frameworks. It achieves 33.17 AP50 on a carefully annotated private chest X-ray dataset which contains 31,000 images. Experiments on the NIH chest X-ray dataset indicate that our method achieves state-of-the-art performance in weakly-supervised disease localization.
Gangming Zhao, Chaowei Fang, Guanbin Li, Licheng Jiao, Yizhou Yu
IEEE Trans. Medical Imaging5
2020 Cross-View Correspondence Reasoning Based on Bipartite Graph Convolutional Network for Mammogram Mass Detection
abstract
Mammogram mass detection is of great clinical significance due to its high proportion in breast cancers. The information from cross views (i.e., mediolateral oblique and cranio-caudal) is highly related and complementary, and is helpful to make comprehensive decisions. However, unlike radiologists who are able to recognize masses with reasoning ability in cross-view images, most existing methods lack the ability to reason under the guidance of domain knowledge, thus it limits the performance. In this paper, we introduce bipartite graph convolutional network to endow existing methods with cross-view reasoning ability of radiologists in mammogram mass detection. The bipartite node sets are constructed by cross-view images respectively to represent relatively consistent regions in breasts, while the bipartite edge learns to model both inherent cross-view geometric constraints and appearance similarities between correspondences. Based on the bipartite graph, the information propagates methodically through correspondences and enables spatial visual features equipped with customized cross-view reasoning ability. Experimental results on DDSM dataset demonstrate that the proposed algorithm achieves state-of-the-art performance. Besides, visual analysis shows the model has a clear physical meaning, which is helpful for radiologists in clinical interpretation.
Fandong Zhang, Qianyi Zhang, Yizhou Wang 0001, Yizhou Yu
CVPR6
2020 Graph-Structured Referring Expression Reasoning in the Wild
abstract
Grounding referring expressions aims to locate in an image an object referred to by a natural language expression. The linguistic structure of a referring expression provides a layout of reasoning over the visual contents, and it is often crucial to align and jointly understand the image and the referring expression. In this paper, we propose a scene graph guided modular network (SGMN), which performs reasoning over a semantic graph and a scene graph with neural modules under the guidance of the linguistic structure of the expression. In particular, we model the image as a structured semantic graph, and parse the expression into a language scene graph. The language scene graph not only decodes the linguistic structure of the expression, but also has a consistent representation with the image semantic graph. In addition to exploring structured solutions to grounding referring expressions, we also propose Ref-Reasoning, a large-scale real-world dataset for structured referring expression reasoning. We automatically generate referring expressions over the scene graphs of images using diverse expression templates and functional programs. This dataset is equipped with real-world visual contents as well as semantically rich expressions with different reasoning layouts. Experimental results show that our SGMN not only significantly outperforms existing state-of-the-art algorithms on the new Ref-Reasoning dataset, but also surpasses state-of-the-art structured methods on commonly used benchmark datasets. It can also provide interpretable visual evidences of reasoning.
Sibei Yang, Guanbin Li, Yizhou Yu
CVPR3
2020 Propagating Over Phrase Relations for One-Stage Visual Grounding
Sibei Yang, Guanbin Li, Yizhou Yu
ECCV (19)3
2020 An End-To-End Network For Detecting Multi-Domain Fractures On X-Ray Images
abstract
Automated fracture detection on medical images is a crucial prerequisite for orthopedic diagnosis. However, due to the considerable variation of bone structures, it is challenging to detect fractures on images filmed from various body parts utilizing a single model. In this paper, we treat each body part as a domain and propose a novel Multi-domain Fracture Detection Network (MFDN), which is composed of two sub-networks, namely, a domain classification network for predicting the domain type of an image and a fracture detection network for detecting fractures on X-ray images of different domains. By constructing Feature Enhancement Modules and Multi-Feature-Enhanced R-CNN, the proposed MFDN extracts better feature representations for each domain. Experimental results on real-world datasets show the effectiveness of our model which has been used in clinical diagnosis with the best performance on all the domains.
Shukai Wu, Lifeng Yan, Yizhou Yu, Sanyuan Zhang
ICIP4
2020 Towards Robust Bone Age Assessment: Rethinking Label Noise and Ambiguity
Ping Gong 0002, Yizhou Wang 0001, Yizhou Yu
MICCAI (6)4
2020 Learning Hybrid Representations for Automatic 3D Vessel Centerline Extraction
Jiafa He, Chengwei Pan, Can Yang 0002, Ming Zhang 0004, Yang Wang 0020, Xiaowei Zhou 0001, Yizhou Yu
MICCAI (6)7
2020 Multi-stream Progressive Up-Sampling Network for Dense CT Image Reconstruction
Qiuyue Liu, Feng Liu 0036, Xiangming Fang, Yizhou Yu, Yizhou Wang 0001
MICCAI (6)5
2020 Context-Aware Refinement Network Incorporating Structural Connectivity Prior for Brain Midline Delineation
Kongming Liang, Yizhou Yu, Yizhou Wang 0001
MICCAI (7)4
2020 BR-GAN: Bilateral Residual Generating Adversarial Network for Mammogram Classification
Chu-ran Wang, Fandong Zhang, Yizhou Yu, Yizhou Wang 0001
MICCAI (2)3
2020 Adaptive Context Selection for Polyp Segmentation
Ruifei Zhang, Guanbin Li, Zhen Li 0026, Shuguang Cui, Dahong Qian, Yizhou Yu
MICCAI (6)6
2020 Revisiting 3D Context Modeling with Supervised Pre-training for Universal Lesion Detection in CT Slices
Shu Zhang 0001, Jincheng Xu, Yu-Chun Chen, Jiechao Ma, Yizhou Wang 0001, Yizhou Yu
MICCAI (4)7
2020 EdgeStereo: An Effective Multi-task Learning Network for Stereo Matching and Edge Detection
Xiao Song 0002, Xu Zhao 0001, Liangji Fang, Hanwen Hu, Yizhou Yu
Int. J. Comput. Vis.5
2020 ROSA: Robust Salient Object Detection Against Adversarial Attacks
abstract
Recently, salient object detection has witnessed remarkable improvement owing to the deep convolutional neural networks which can harvest powerful features for images. In particular, the state-of-the-art salient object detection methods enjoy high accuracy and efficiency from fully convolutional network (FCN)-based frameworks which are trained from end to end and predict pixel-wise labels. However, such framework suffers from adversarial attacks which confuse neural networks via adding quasi-imperceptible noises to input images without changing the ground truth annotated by human subjects. To our knowledge, this paper is the first one that mounts successful adversarial attacks on salient object detection models and verifies that adversarial samples are effective on a wide range of existing methods. Furthermore, this paper proposes a novel end-to-end trainable framework to enhance the robustness for arbitrary FCN-based salient object detection models against adversarial attacks. The proposed framework adopts a novel idea that first introduces some new generic noise to destroy adversarial perturbations, and then learns to predict saliency maps for input images with the introduced noise. Specifically, our proposed method consists of a segment-wise shielding component, which preserves boundaries and destroys delicate adversarial noise patterns and a context-aware restoration component, which refines saliency maps through global contrast modeling. The experimental results suggest that our proposed framework improves the performance significantly for state-of-the-art models on a series of datasets.
Haofeng Li, Guanbin Li, Yizhou Yu
IEEE Trans. Cybern.3
2020 Self-Enhanced Convolutional Network for Facial Video Hallucination
abstract
As a domain-specific super-resolution problem, facial image hallucination has enjoyed a series of breakthroughs thanks to the advances of deep convolutional neural networks. However, the direct migration of existing methods to video is still difficult to achieve good performance due to its lack of alignment and consistency modelling in temporal domain. Taking advantage of high inter-frame dependency in videos, we propose a selfenhanced convolutional network for facial video hallucination. It is implemented by making full usage of preceding super-resolved frames and a temporal window of adjacent low-resolution frames. Specifically, the algorithm first obtains the initial high-resolution inference of each frame by taking into consideration a sequence of consecutive low-resolution inputs through temporal consistency modelling. It further recurrently exploits the reconstructed results and intermediate features of a sequence of preceding frames to improve the initial super-resolution of the current frame by modelling the coherence of structural facial features across frames. Quantitative and qualitative evaluations demonstrate the superiority of the proposed algorithm against state-of-theart methods. Moreover, our algorithm also achieves excellent performance in the task of general video super-resolution in a single-shot setting.
Chaowei Fang, Guanbin Li, Xiaoguang Han 0001, Yizhou Yu
IEEE Trans. Image Process.4
2020 Residual Learning for Salient Object Detection
abstract
Recent deep learning based salient object detection methods improve the performance by introducing multi-scale strategies into fully convolutional neural networks (FCNs). The final result is obtained by integrating all the predictions at each scale. However, the existing multi-scale based methods suffer from several problems: 1) it is difficult to directly learn discriminative features and filters to regress high-resolution saliency masks for each scale; 2) rescaling the multi-scale features could pull in many redundant and inaccurate values, and this weakens the representational ability of the network. In this paper, we propose a residual learning strategy and introduce to gradually refine the coarse prediction scale-by-scale. Concretely, instead of directly predicting the finest-resolution result at each scale, we learn to predict residuals to remedy the errors between coarse saliency map and scale-matching ground truth masks. We employ a Dilated Convolutional Pyramid Pooling (DCPP) module to generate the coarse prediction and guide the the residual learning process through several novel Attentional Residual Modules (ARMs). We name our network as Residual Refinement Network (R2Net). We demonstrate the effectiveness of the proposed method against other state-of-the-art algorithms on five released benchmark datasets. Our R2Net is a fully convolutional network which does not need any post-processing and achieves a real-time speed of 33 FPS when it is run on one GPU.
Mengyang Feng, Huchuan Lu, Yizhou Yu
IEEE Trans. Image Process.3
2020 Online Alternate Generator Against Adversarial Attacks
abstract
The field of computer vision has witnessed phenomenal progress in recent years partially due to the development of deep convolutional neural networks. However, deep learning models are notoriously sensitive to adversarial examples which are synthesized by adding quasi-perceptible noises on real images. Some existing defense methods require to re-train attacked target networks and augment the train set via known adversarial attacks, which is inefficient and might be unpromising with unknown attack types. To overcome the above issues, we propose a portable defense method, online alternate generator, which does not need to access or modify the parameters of the target networks. The proposed method works by online synthesizing another image from scratch for an input image, instead of removing or destroying adversarial noises. To avoid pretrained parameters exploited by attackers, we alternately update the generator and the synthesized image at the inference stage. Experimental results demonstrate that the proposed defensive scheme and method outperforms a series of state-of-the-art defending models against gray-box adversarial attacks.
Haofeng Li, Yirui Zeng, Guanbin Li, Liang Lin 0004, Yizhou Yu
IEEE Trans. Image Process.5
2020 Exploring Task Structure for Brain Tumor Segmentation From Multi-Modality MR Images
abstract
Brain tumor segmentation, which aims at segmenting the whole tumor area, enhancing tumor core area, and tumor core area from each input multi-modality bioimaging data, has received considerable attention from both academia and industry. However, the existing approaches usually treat this problem as a common semantic segmentation task without taking into account the underlying rules in clinical practice. In reality, physicians tend to discover different tumor areas by weighing different modality volume data. Also, they initially segment the most distinct tumor area, and then gradually search around to find the other two. We refer to the first property as the task-modality structure while the second property as the task-task structure, based on which we propose a novel task-structured brain tumor segmentation network (TSBTS net). Specifically, to explore the task-modality structure, we design a modality-aware feature embedding mechanism to infer the important weights of the modality data during network learning. To explore the tasktask structure, we formulate the prediction of the different tumor areas as conditional dependency sub-tasks and encode such dependency in the network stream. Experiments on BraTS benchmarks show that the proposed method achieves superior performance in segmenting the desired brain tumor areas while requiring relatively lower computational costs, compared to other state-of-the-art methods and baseline models.
Dingwen Zhang, Guohai Huang, Qiang Zhang 0020, Jungong Han, Junwei Han 0001, Yizhou Wang 0001, Yizhou Yu
IEEE Trans. Image Process.7
2020 CaricatureShop: Personalized and Photorealistic Caricature Sketching
abstract
In this paper, we propose the first sketching system for interactively personalized and photorealistic face caricaturing. Input an image of a human face, the users can create caricature photos by manipulating its facial feature curves. Our system first performs exaggeration on the recovered 3D face model, which is conducted by assigning the laplacian of each vertex a scaling factor according to the edited sketches. The mapping between 2D sketches and the vertex-wise scaling field is constructed by a novel deep learning architecture. Our approach allows outputting different exaggerations when applying the same sketching on different input figures in term of their different geometric characteristics, which makes the generated results "personalized". With the obtained 3D caricature model, two images are generated, one obtained by applying 2D warping guided by the underlying 3D mesh deformation and the other obtained by re-rendering the deformed 3D textured model. These two images are then seamlessly integrated to produce our final output. Due to the severe stretching of meshes, the rendered texture is of blurry appearances. A deep learning approach is exploited to infer the missing details for enhancing these blurry regions. Moreover, a relighting operation is invented to further improve the photorealism of the result. These further make our results "photorealistic". The qualitative experiment results validated the efficiency of our sketching system.
Xiaoguang Han 0001, Kangcheng Hou, Dong Du 0002, Yuda Qiu, Shuguang Cui, Kun Zhou 0001, Yizhou Yu
IEEE Trans. Vis. Comput. Graph.7
2020 Neural Style Transfer: A Review
abstract
The seminal work of Gatys et al. demonstrated the power of Convolutional Neural Networks (CNNs) in creating artistic imagery by separating and recombining image content and style. This process of using CNNs to render a content image in different styles is referred to as Neural Style Transfer (NST). Since then, NST has become a trending topic both in academic literature and industrial applications. It is receiving increasing attention and a variety of approaches are proposed to either improve or extend the original NST algorithm. In this paper, we aim to provide a comprehensive overview of the current progress towards NST. We first propose a taxonomy of current algorithms in the field of NST. Then, we present several evaluation methods and compare different NST algorithms both qualitatively and quantitatively. The review concludes with a discussion of various applications of NST and open problems for future research. A list of papers discussed in this review, corresponding codes, pre-trained models and more comparison results are publicly available at: https://osf.io/f8tu4/.
Yongcheng Jing, Yezhou Yang, Zunlei Feng, Jingwen Ye, Yizhou Yu, Mingli Song
IEEE Trans. Vis. Comput. Graph.5
2019 Non-Local Context Encoder: Robust Biomedical Image Segmentation against Adversarial Attacks
abstract
Recent progress in biomedical image segmentation based on deep convolutional neural networks (CNNs) has drawn much attention. However, its vulnerability towards adversarial samples cannot be overlooked. This paper is the first one that discovers that all the CNN-based state-of-the-art biomedical image segmentation models are sensitive to adversarial perturbations. This limits the deployment of these methods in safety-critical biomedical fields. In this paper, we discover that global spatial dependencies and global contextual information in a biomedical image can be exploited to defend against adversarial attacks. To this end, non-local context encoder (NLCE) is proposed to model short- and long-range spatial dependencies and encode global contexts for strengthening feature activations by channel-wise attention. The NLCE modules enhance the robustness and accuracy of the non-local context encoding network (NLCEN), which learns robust enhanced pyramid feature representations with NLCE modules, and then integrates the information across different levels. Experiments on both lung and skin lesion segmentation datasets have demonstrated that NLCEN outperforms any other state-of-the-art biomedical image segmentation methods against adversarial attacks. In addition, NLCE modules can be applied to improve the robustness of other CNN-based biomedical image segmentation methods.
Sibei Yang, Guanbin Li, Haofeng Li, HuiYou Chang, Yizhou Yu
AAAI6
2019 Weakly Supervised Complementary Parts Models for Fine-Grained Image Classification From the Bottom Up
abstract
Given a training dataset composed of images and corresponding category labels, deep convolutional neural networks show a strong ability in mining discriminative parts for image classification. However, deep convolutional neural networks trained with image level labels only tend to focus on the most discriminative parts while missing other object parts, which could provide complementary information. In this paper, we approach this problem from a different perspective. We build complementary parts models in a weakly supervised manner to retrieve information suppressed by dominant object parts detected by convolutional neural networks. Given image level labels only, we first extract rough object instances by performing weakly supervised object detection and instance segmentation using Mask R-CNN and CRF-based segmentation. Then we estimate and search for the best parts model for each object instance under the principle of preserving as much diversity as possible. In the last stage, we build a bi-directional long short-term memory (LSTM) network to fuze and encode the partial information of these complementary parts into a comprehensive feature for image classification. Experimental results indicate that the proposed method not only achieves significant improvement over our baseline models, but also outperforms state-of-the-art algorithms by a large margin (6.7%, 2.8%, 5.2% respectively) on Stanford Dogs 120, Caltech-UCSD Birds 2011-200 and Caltech 256.
Weifeng Ge, Xiangru Lin, Yizhou Yu
CVPR3
2019 Cross-Modal Relationship Inference for Grounding Referring Expressions
abstract
Grounding referring expressions is a fundamental yet challenging task facilitating human-machine communication in the physical world. It locates the target object in an image on the basis of the comprehension of the relationships between referring natural language expressions and the image. A feasible solution for grounding referring expressions not only needs to extract all the necessary information (i.e. objects and the relationships among them) in both the image and referring expressions, but also compute and represent multimodal contexts from the extracted information. Unfortunately, existing work on grounding referring expressions cannot extract multi-order relationships from the referring expressions accurately and the contexts they obtain have discrepancies with the contexts described by referring expressions. In this paper, we propose a Cross-Modal Relationship Extractor (CMRE) to adaptively highlight objects and relationships, that have connections with a given expression, with a cross-modal attention mechanism, and represent the extracted information as a language-guided visual relation graph. In addition, we propose a Gated Graph Convolutional Network (GGCN) to compute multimodal semantic contexts by fusing information from different modes and propagating multimodal information in the structured relation graph. Experiments on various common benchmark datasets show that our Cross-Modal Relationship Inference Network, which consists of CMRE and GGCN, outperforms all existing state-of-the-art methods.
Sibei Yang, Guanbin Li, Yizhou Yu
CVPR3
2019 Multi-Source Weak Supervision for Saliency Detection
abstract
The high cost of pixel-level annotations makes it appealing to train saliency detection models with weak supervision. However, a single weak supervision source usually does not contain enough information to train a well-performing model. To this end, we propose a unified framework to train saliency detection models with diverse weak supervision sources. In this paper, we use category labels, captions, and unlabelled data for training, yet other supervision sources can also be plugged into this flexible framework. We design a classification network (CNet) and a caption generation network (PNet), which learn to predict object categories and generate captions, respectively, meanwhile highlight the most important regions for corresponding tasks. An attention transfer loss is designed to transmit supervision signal between networks, such that the network designed to be trained with one supervision source can benefit from another. An attention coherence loss is defined on unlabelled data to encourage the networks to detect generally salient regions instead of task-specific regions. We use CNet and PNet to generate pixel-level pseudo labels to train a saliency prediction network (SNet). During the testing phases, we only need SNet to predict saliency maps. Experiments demonstrate the performance of our method compares favourably against unsupervised and weakly supervised methods and even some supervised methods.
Yu Zeng 0001, Yunzhi Zhuge, Huchuan Lu, Lihe Zhang, Mingyang Qian, Yizhou Yu
CVPR6
2019 Cascaded Generative and Discriminative Learning for Microcalcification Detection in Breast Mammograms
abstract
Accurate microcalcification (μC) detection is of great importance due to its high proportion in early breast cancers. Most of the previous μC detection methods belong to discriminative models, where classifiers are exploited to distinguish μCs from other backgrounds. However, it is still challenging for these methods to tell the μCs from amounts of normal tissues because they are too tiny (at most 14 pixels). Generative methods can precisely model the normal tissues and regard the abnormal ones as outliers, while they fail to further distinguish the μCs from other anomalies, i.e. vessel calcifications. In this paper, we propose a hybrid approach by taking advantages of both generative and discriminative models. Firstly, a generative model named Anomaly Separation Network (ASN) is used to generate candidate μCs. ASN contains two major components. A deep convolutional encoder-decoder network is built to learn the image reconstruction mapping and a t-test loss function is designed to separate the distributions of the reconstruction residuals of μCs from normal tissues. Secondly, a discriminative model is cascaded to tell the μCs from the false positives. Finally, to verify the effectiveness of our method, we conduct experiments on both public and in-house datasets, which demonstrates that our approach outperforms previous state-of-the-art methods.
Fandong Zhang, Xinwei Sun 0001, Xiuli Li, Yizhou Yu, Yizhou Wang 0001
CVPR6
2019 Motion Guided Attention for Video Salient Object Detection
abstract
Video salient object detection aims at discovering the most visually distinctive objects in a video. How to effectively take object motion into consideration during video salient object detection is a critical issue. Existing state-of-the-art methods either do not explicitly model and harvest motion cues or ignore spatial contexts within optical flow images. In this paper, we develop a multi-task motion guided video salient object detection network, which learns to accomplish two sub-tasks using two sub-networks, one sub-network for salient object detection in still images and the other for motion saliency detection in optical flow images. We further introduce a series of novel motion guided attention modules, which utilize the motion saliency sub-network to attend and enhance the sub-network for still images. These two sub-networks learn to adapt to each other by end-to-end training. Experimental results demonstrate that the proposed method significantly outperforms existing state-of-the-art algorithms on a wide range of benchmarks. We hope our simple and effective approach will serve as a solid baseline and help ease future research in video salient object detection. Code and models will be made available.
Haofeng Li, Guanqi Chen, Guanbin Li, Yizhou Yu
ICCV4
2019 Align, Attend and Locate: Chest X-Ray Diagnosis via Contrast Induced Attention Network With Limited Supervision
abstract
Obstacles facing accurate identification and localization of diseases in chest X-ray images lie in the lack of high-quality images and annotations. In this paper, we propose a Contrast Induced Attention Network (CIA-Net), which exploits the highly structured property of chest X-ray images and localizes diseases via contrastive learning on the aligned positive and negative samples. To force the attention module to focus only on sites of abnormalities, we also introduce a learnable alignment module to adjust all the input images, which eliminates variations of scales, angles, and displacements of X-ray images generated under bad scan conditions. We show that the use of contrastive attention and alignment module allows the model to learn rich identification and localization information using only a small amount of location annotations, resulting in state-of-the-art performance in NIH chest X-ray dataset.
Jingyu Liu 0004, Gangming Zhao, Ming Zhang 0004, Yizhou Wang 0001, Yizhou Yu
ICCV6
2019 Dynamic Graph Attention for Referring Expression Comprehension
abstract
Referring expression comprehension aims to locate the object instance described by a natural language referring expression in an image. This task is compositional and inherently requires visual reasoning on top of the relationships among the objects in the image. Meanwhile, the visual reasoning process is guided by the linguistic structure of the referring expression. However, existing approaches treat the objects in isolation or only explore the first-order relationships between objects without being aligned with the potential complexity of the expression. Thus it is hard for them to adapt to the grounding of complex referring expressions. In this paper, we explore the problem of referring expression comprehension from the perspective of language-driven visual reasoning, and propose a dynamic graph attention network to perform multi-step reasoning by modeling both the relationships among the objects in the image and the linguistic structure of the expression. In particular, we construct a graph for the image with the nodes and edges corresponding to the objects and their relationships respectively, propose a differential analyzer to predict a language-guided visual reasoning process, and perform stepwise reasoning on top of the graph to update the compound object representation at every node. Experimental results demonstrate that the proposed method can not only significantly surpass all existing state-of-the-art algorithms across three common benchmark datasets, but also generate interpretable visual evidences for stepwise locating the objects referred to in complex language descriptions.
Sibei Yang, Guanbin Li, Yizhou Yu
ICCV3
2019 Harnessing 2D Networks and 3D Features for Automated Pancreas Segmentation from Volumetric CT Images
Huai Chen, Xiuying Wang 0001, Xiyi Wu, Yizhou Yu, Lisheng Wang
MICCAI (6)5
2019 Globally Guided Progressive Fusion Network for 3D Pancreas Segmentation
Chaowei Fang, Guanbin Li, Chengwei Pan, Yizhou Yu
MICCAI (2)5
2019 MVP-Net: Multi-view FPN with Position-Aware Attention for Deep Universal Lesion Detection
Shu Zhang 0001, Junge Zhang, Kaiqi Huang, Yizhou Wang 0001, Yizhou Yu
MICCAI (6)6
2019 From Unilateral to Bilateral Learning: Detecting Mammogram Masses with Contrasted Bilateral Network
Shu Zhang 0001, Qianyi Zhang, Fandong Zhang, Xiuli Li, Yizhou Wang 0001, Yizhou Yu
MICCAI (6)9
2019 Simultaneous Lung Field Detection and Segmentation for Pediatric Chest Radiographs
Guanbin Li, Fuyu Wang 0001, Longjiang E, Yizhou Yu, Liang Lin 0004, Huiying Liang
MICCAI (6)5
2019 Transductive Zero-Shot Learning with Visual Structure Constraint
abstract
To recognize objects of the unseen classes, most existing Zero-Shot Learning (ZSL) methods first learn a compatible projection function between the common semantic space and the visual space based on the data of source seen classes, then directly apply it to the target unseen classes. However, in real scenarios, the data distribution between the source and target domain might not match well, thus causing the well-known domain shift problem. Based on the observation that visual features of test instances can be separated into different clusters, we propose a new visual structure constraint on class centers for transductive ZSL, to improve the generality of the projection function (\ie alleviate the above domain shift problem). Specifically, three different strategies (symmetric Chamfer-distance,Bipartite matching distance, and Wasserstein distance) are adopted to align the projected unseen semantic centers and visual cluster centers of test instances. We also propose a new training strategy to handle the real cases where many unrelated images exist in the test dataset, which is not considered in previous methods. Experiments on many widely used datasets demonstrate that the proposed visual structure constraint can bring substantial performance gain consistently and achieve state-of-the-art results.
Ziyu Wan, Dongdong Chen 0001, Yan Li 0043, Xingguang Yan, Junge Zhang, Yizhou Yu, Jing Liao 0001
NeurIPS6
2019 Piecewise Flat Embedding for Image Segmentation
abstract
We introduce a new multi-dimensional nonlinear embedding-Piecewise Flat Embedding (PFE)-for image segmentation. Based on the theory of sparse signal recovery, piecewise flat embedding with diverse channels attempts to recover a piecewise constant image representation with sparse region boundaries and sparse cluster value scattering. The resultant piecewise flat embedding exhibits interesting properties such as suppressing slowly varying signals, and offers an image representation with higher region identifiability which is desirable for image segmentation or high-level semantic analysis tasks. We formulate our embedding as a variant of the Laplacian Eigen-map embedding with an L1,p(01,1-regularized piecewise flat embeddings. We further generalize this algorithm through iterative reweighting to solve the general L1,p-regularized problem. To demonstrate its efficacy, we integrate PFE into two existing image segmentation frameworks, segmentation based on clustering and hierarchical segmentation based on contour detection. Experiments on four major benchmark datasets, BSDS500, MSRC, Stanford Background Dataset, and PASCAL Context, show that segmentation algorithms incorporating our embedding achieve significantly improved results.
Chaowei Fang, Zicheng Liao, Yizhou Yu
IEEE Trans. Pattern Anal. Mach. Intell.3
2019 Context-Aware Semantic Inpainting
abstract
In recent times, image inpainting has witnessed rapid progress due to the generative adversarial networks (GANs) that are able to synthesize realistic contents. However, most existing GAN-based methods for semantic inpainting apply an auto-encoder architecture with a fully connected layer, which cannot accurately maintain spatial information. In addition, the discriminator in existing GANs struggles to comprehend high-level semantics within the image context and yields semantically consistent content. Existing evaluation criteria are biased toward blurry results and cannot well characterize edge preservation and visual authenticity in the inpainting results. In this paper, we propose an improved GAN to overcome the aforementioned limitations. Our proposed GAN-based framework consists of a fully convolutional design for the generator which helps to better preserve spatial structures and a joint loss function with a revised perceptual loss to capture high-level semantics in the context. Furthermore, we also introduce two novel measures to better assess the quality of image inpainting results. The experimental results demonstrate that our method outperforms the state-of-the-art under a wide range of criteria.
Haofeng Li, Guanbin Li, Liang Lin 0004, Hongchuan Yu, Yizhou Yu
IEEE Trans. Cybern.5
2019 A Benchmark for Edge-Preserving Image Smoothing
abstract
Edge-preserving image smoothing is an important step for many low-level vision problems. Though many algorithms have been proposed, there are several difficulties hindering its further development. First, most existing algorithms cannot perform well on a wide range of image contents using a single parameter setting. Second, the performance evaluation of edge-preserving image smoothing remains subjective, and there lacks a widely accepted datasets to objectively compare the different algorithms. To address these issues and further advance the state of the art, in this work we propose a benchmark for edge-preserving image smoothing. This benchmark includes an image dataset with groundtruth image smoothing results as well as baseline algorithms that can generate competitive edge-preserving smoothing results for a wide range of image contents. The established dataset contains 500 training and testing images with a number of representative visual object categories, while the baseline methods in our benchmark are built upon representative deep convolutional network architectures, on top of which we design novel loss functions well suited for edge-preserving image smoothing. The trained deep networks run faster than most state-of-the-art smoothing algorithms with leading smoothing results both qualitatively and quantitatively. The benchmark will be made publicly accessible.
Feida Zhu 0002, Zhetong Liang, Xixi Jia, Lei Zhang 0006, Yizhou Yu
IEEE Trans. Image Process.5
2019 Traffic Sign Detection Using a Multi-Scale Recurrent Attention Network
abstract
Traffic sign detection plays an important role in intelligent transportation systems. But traffic signs are still not well-detected by deep convolution neural network-based methods because the sizes of their feature maps are constrained, and the environmental context information has not been fully exploited by other researchers. What we need is a way to incorporate relevant context detail from the neighboring layers into the detection architecture. We have developed a novel traffic sign detection approach based on recurrent attention for multi-scale analysis and use of local context in the image. Experiments on the German traffic sign detection benchmark and the Tsinghua-Tencent 100K data set demonstrated that our approach obtained an accuracy comparable to the state-of-the-art approaches in traffic sign detection.
Judith Gelernter, Xun Wang 0007, Jianyuan Li, Yizhou Yu
IEEE Trans. Intell. Transp. Syst.5
2019 Facial Landmark Machines: A Backbone-Branches Architecture With Progressive Representation Learning
abstract
Facial landmark localization plays a critical role in face recognition and analysis. In this paper, we propose a novel cascaded backbone-branches fully convolutional neural network (BB-FCN) for rapidly and accurately localizing facial landmarks in unconstrained and cluttered settings. Our proposed BB-FCN generates facial landmark response maps directly from raw images without any preprocessing. BB-FCN follows a coarse-to-fine cascaded pipeline, which consists of a backbone network to roughly detect the locations of all facial landmarks and one branch network for each type of detected landmark to further refine its location. Furthermore, to facilitate the facial landmark localization under unconstrained settings, we propose a large-scale benchmark named SYSU16K, which contains 16 000 faces with large variations in pose, expression, illumination, and resolution. Extensive experimental evaluations demonstrate that our proposed BB-FCN can significantly outperform the state of the art under both constrained (i.e., within detected facial regions only) and unconstrained settings. We further confirm that high-quality facial landmarks localized with our proposed network can also improve the precision and recall of face detection.
Lingbo Liu, Guanbin Li, Yuan Xie 0004, Yizhou Yu, Qing Wang 0018, Liang Lin 0004
IEEE Trans. Multim.4
2019 Harvesting Visual Objects from Internet Images via Deep-Learning-Based Objectness Assessment
abstract
The collection of internet images has been growing in an astonishing speed. It is undoubted that these images contain rich visual information that can be useful in many applications, such as visual media creation and data-driven image synthesis. In this article, we focus on the methodologies for building a visual object database from a collection of internet images. Such database is built to contain a large number of high-quality visual objects that can help with various data-driven image applications. Our method is based on dense proposal generation and objectness-based re-ranking. A novel deep convolutional neural network is designed for the inference of proposal objectness , the probability of a proposal containing optimally located foreground object. In our work, the objectness is quantitatively measured in regard of completeness and fullness , reflecting two complementary features of an optimal proposal: a complete foreground and relatively small background. Our experiments indicate that object proposals re-ranked according to the output of our network generally achieve higher performance than those produced by other state-of-the-art methods. As a concrete example, a database of over 1.2 million visual objects has been built using the proposed method, and has been successfully used in various data-driven image applications.
Kan Wu 0005, Guanbin Li, Haofeng Li, Jian J. Zhang 0001, Yizhou Yu
ACM Trans. Multim. Comput. Commun. Appl.5
2018 ReCoNet: Real-Time Coherent Video Style Transfer Network
Chang Gao 0001, Derun Gu, Fangjun Zhang, Yizhou Yu
ACCV (6)4
2018 Multi-Evidence Filtering and Fusion for Multi-Label Classification, Object Detection and Semantic Segmentation Based on Weakly Supervised Learning
abstract
Supervised object detection and semantic segmentation require object or even pixel level annotations. When there exist image level labels only, it is challenging for weakly supervised algorithms to achieve accurate predictions. The accuracy achieved by top weakly supervised algorithms is still significantly lower than their fully supervised counterparts. In this paper, we propose a novel weakly supervised curriculum learning pipeline for multi-label object recognition, detection and semantic segmentation. In this pipeline, we first obtain intermediate object localization and pixel labeling results for the training images, and then use such results to train task-specific deep networks in a fully supervised manner. The entire process consists of four stages, including object localization in the training images, filtering and fusing object instances, pixel labeling for the training images, and task-specific network training. To obtain clean object instances in the training images, we propose a novel algorithm for filtering, fusing and classifying object instances collected from multiple solution mechanisms. In this algorithm, we incorporate both metric learning and density-based clustering to filter detected object instances. Experiments show that our weakly supervised pipeline achieves state-of-the-art results in multi-label image classification as well as weakly supervised object detection and very competitive results in weakly supervised semantic segmentation on MS-COCO, PASCAL VOC 2007 and PASCAL VOC 2012.
Weifeng Ge, Sibei Yang, Yizhou Yu
CVPR3
2018 Stroke Controllable Fast Style Transfer with Adaptive Receptive Fields
Yongcheng Jing, Yang Liu 0212, Yezhou Yang, Zunlei Feng, Yizhou Yu, Dacheng Tao, Mingli Song
ECCV (13)5
2018 Automatic object extraction from images using deep neural networks and the level-set method
abstract
The authors propose an automatic method for extracting objects with fine quality from photographs. The authors’ method starts with finding bounding boxes that enclose potential objects, which is achievable by state‐of‐the‐art object proposal methods. To further segment objects within obtained bounding boxes, the authors propose a new multi‐pass level‐set method based on saliency detection and foreground pixel classification. The level‐set function is initially constructed with respect to the automatically detected salient parts within the bounding box, which eliminates potential user interaction and predicts an initial set of pixels on the object. The input features for foreground pixel classifiers are constructed as a combination of classical texture features from the Gabor filter banks and convolutional features from a pre‐trained deep neural network. Through multi‐pass evolution of the level‐set function and re‐training of the foreground pixel classifier, the authors’ method is able to overcome possible inaccuracies in the initial level‐set function and converge to the real object boundary.
Kan Wu 0005, Yizhou Yu
IET Image Process.2
2018 Moiré Photo Restoration Using Multiresolution Convolutional Neural Networks
abstract
Digital cameras and mobile phones enable us to conveniently record precious moments. While digital image quality is constantly being improved, taking high-quality photos of digital screens still remains challenging because the photos are often contaminated with moiré patterns, a result of the interference between the pixel grids of the camera sensor and the device screen. Moiré patterns can severely damage the visual quality of photos. However, few studies have aimed to solve this problem. In this paper, we introduce a novel multiresolution fully convolutional network for automatically removing moiré patterns from photos. Since a moiré pattern spans over a wide range of frequencies, our proposed network performs a nonlinear multiresolution analysis of the input image before computing how to cancel moiré artefacts within every frequency band. We also create a large-scale benchmark dataset with 100,000+ image pairs for investigating and evaluating moiré pattern removal algorithms. Our network achieves state-of-the-art performance on this dataset in comparison to existing learning architectures for image restoration problems.
Yujing Sun 0001, Yizhou Yu, Wenping Wang 0001
IEEE Trans. Image Process.2
2018 Contrast-Oriented Deep Neural Networks for Salient Object Detection
abstract
Deep convolutional neural networks (CNNs) have become a key element in the recent breakthrough of salient object detection. However, existing CNN-based methods are based on either patchwise (regionwise) training and inference or fully convolutional networks. Methods in the former category are generally time-consuming due to severe storage and computational redundancies among overlapping patches. To overcome this deficiency, methods in the second category attempt to directly map a raw input image to a predicted dense saliency map in a single network forward pass. Though being very efficient, it is arduous for these methods to detect salient objects of different scales or salient regions with weak semantic information. In this paper, we develop hybrid contrast-oriented deep neural networks to overcome the aforementioned limitations. Each of our deep networks is composed of two complementary components, including a fully convolutional stream for dense prediction and a segment-level spatial pooling stream for sparse saliency inference. We further propose an attentional module that learns weight maps for fusing the two saliency predictions from these two streams. A tailored alternate scheme is designed to train these deep networks by fine-tuning pretrained baseline models. Finally, a customized fully connected conditional random field model incorporating a salient contour feature embedding can be optionally applied as a postprocessing step to improve spatial coherence and contour positioning in the fused result from these two streams. Extensive experiments on six benchmark data sets demonstrate that our proposed model can significantly outperform the state of the art in terms of all popular evaluation metrics.
Guanbin Li, Yizhou Yu
IEEE Trans. Neural Networks Learn. Syst.2
2018 Image super-resolution via deterministic-stochastic synthesis and local statistical rectification
abstract
Single image superresolution has been a popular research topic in the last two decades and has recently received a new wave of interest due to deep neural networks. In this paper, we approach this problem from a different perspective. With respect to a downsampled low resolution image, we model a high resolution image as a combination of two components, a deterministic component and a stochastic component. The deterministic component can be recovered from the low-frequency signals in the downsampled image. The stochastic component, on the other hand, contains the signals that have little correlation with the low resolution image. We adopt two complementary methods for generating these two components. While generative adversarial networks are used for the stochastic component, deterministic component reconstruction is formulated as a regression problem solved using deep neural networks. Since the deterministic component exhibits clearer local orientations, we design novel loss functions tailored for such properties for training the deep regression network. These two methods are first applied to the entire input image to produce two distinct high-resolution images. Afterwards, these two images are fused together using another deep neural network that also performs local statistical rectification, which tries to make the local statistics of the fused image match the same local statistics of the groundtruth image. Quantitative results and a user study indicate that the proposed method outperforms existing state-of-the-art algorithms with a clear margin.
Weifeng Ge, Bingchen Gong, Yizhou Yu
ACM Trans. Graph.3
2017 Borrowing Treasures from the Wealthy: Deep Transfer Learning through Selective Joint Fine-Tuning
abstract
Deep neural networks require a large amount of labeled training data during supervised learning. However, collecting and labeling so much data might be infeasible in many cases. In this paper, we introduce a deep transfer learning scheme, called selective joint fine-tuning, for improving the performance of deep learning tasks with insufficient training data. In this scheme, a target learning task with insufficient training data is carried out simultaneously with another source learning task with abundant training data. However, the source learning task does not use all existing training data. Our core idea is to identify and use a subset of training images from the original source learning task whose low-level characteristics are similar to those from the target learning task, and jointly fine-tune shared convolutional layers for both tasks. Specifically, we compute descriptors from linear or nonlinear filter bank responses on training images from both tasks, and use such descriptors to search for a desired subset of training samples for the source learning task. Experiments demonstrate that our deep transfer learning scheme achieves state-of-the-art performance on multiple visual classification tasks with insufficient training data for deep learning. Such tasks include Caltech 256, MIT Indoor 67, and fine-grained classification problems (Oxford Flowers 102 and Stanford Dogs 120). In comparison to fine-tuning without a source domain, the proposed method can improve the classification accuracy by 2% - 10% using a single model. Codes and models are available at https://github.com/ZYYSzj/Selective-Joint-Fine-tuning.
Weifeng Ge, Yizhou Yu
CVPR2
2017 Instance-Level Salient Object Segmentation
abstract
Image saliency detection has recently witnessed rapid progress due to deep convolutional neural networks. However, none of the existing methods is able to identify object instances in the detected salient regions. In this paper, we present a salient instance segmentation method that produces a saliency mask with distinct object instance labels for an input image. Our method consists of three steps, estimating saliency map, detecting salient object contours and identifying salient object instances. For the first two steps, we propose a multiscale saliency refinement network, which generates high-quality salient region masks and salient object contours. Once integrated with multiscale combinatorial grouping and a MAP-based subset optimization framework, our method can generate very promising salient object instance segmentation results. To promote further research and evaluation of salient instance segmentation, we also construct a new database of 1000 images and their pixelwise salient instance annotations. Experimental results demonstrate that our proposed method is capable of achieving state-of-the-art performance on all public benchmarks for salient region detection as well as on our new dataset for salient instance segmentation.
Guanbin Li, Yuan Xie 0004, Liang Lin 0004, Yizhou Yu
CVPR4
2017 High-Resolution Shape Completion Using Deep Neural Networks for Global Structure and Local Geometry Inference
abstract
We propose a data-driven method for recovering missing parts of 3D shapes. Our method is based on a new deep learning architecture consisting of two sub-networks: a global structure inference network and a local geometry refinement network. The global structure inference network incorporates a long short-term memorized context fusion module (LSTM-CF) that infers the global structure of the shape based on multi-view depth information provided as part of the input. It also includes a 3D fully convolutional (3DFCN) module that further enriches the global structure representation according to volumetric information in the input. Under the guidance of the global structure network, the local geometry refinement network takes as input local 3D patches around missing regions, and progressively produces a high-resolution, complete surface through a volumetric encoder-decoder architecture. Our method jointly trains the global structure inference and local geometry refinement networks in an end-to-end manner. We perform qualitative and quantitative evaluations on six object categories, demonstrating that our method outperforms existing state-of-the-art work on shape completion.
Xiaoguang Han 0001, Zhen Li 0026, Evangelos Kalogerakis, Yizhou Yu
ICCV5
2017 Folding Membrane Proteins by Deep Transfer Learning
Zhen Li 0026, Sheng Wang 0001, Yizhou Yu, Jinbo Xu
RECOMB3
2017 A fast propagation scheme for approximate geodesic paths
Xiaoguang Han 0001, Hongchuan Yu, Yizhou Yu, Jian J. Zhang 0001
Graph. Model.3
2017 Editorial Special issue on the fifth Computational Visual Media conference (CVM 2017)
Niloy J. Mitra, Yizhou Yu, Ming-Ming Cheng
Graph. Model.2
2017 Preface
Shi-Min Hu 0001, Niloy J. Mitra, Yizhou Yu
J. Comput. Sci. Technol.3
2017 Exemplar-Based Image and Video Stylization Using Fully Convolutional Semantic Features
abstract
Color and tone stylization in images and videos strives to enhance unique themes with artistic color and tone adjustments. It has a broad range of applications from professional image postprocessing to photo sharing over social networks. Mainstream photo enhancement softwares, such as Adobe Lightroom and Instagram, provide users with predefined styles, which are often hand-crafted through a trial-and-error process. Such photo adjustment tools lack a semantic understanding of image contents and the resulting global color transform limits the range of artistic styles it can represent. On the other hand, stylistic enhancement needs to apply distinct adjustments to various semantic regions. Such an ability enables a broader range of visual styles. In this paper, we first propose a novel deep learning architecture for exemplar-based image stylization, which learns local enhancement styles from image pairs. Our deep learning architecture consists of fully convolutional networks for automatic semantics-aware feature extraction and fully connected neural layers for adjustment prediction. Image stylization can be efficiently accomplished with a single forward pass through our deep network. To extend our deep network from image stylization to video stylization, we exploit temporal superpixels to facilitate the transfer of artistic styles from image exemplars to videos. Experiments on a number of data sets for image stylization as well as a diverse set of video clips demonstrate the effectiveness of our deep learning architecture.
Feida Zhu 0002, Zhicheng Yan 0001, Jiajun Bu, Yizhou Yu
IEEE Trans. Image Process.4
2017 DeepSketch2Face: a deep learning based sketching system for 3D face and caricature modeling
abstract
Face modeling has been paid much attention in the field of visual computing. There exist many scenarios, including cartoon characters, avatars for social media, 3D face caricatures as well as face-related art and design, where low-cost interactive face modeling is a popular approach especially among amateur users. In this paper, we propose a deep learning based sketching system for 3D face and caricature modeling. This system has a labor-efficient sketching interface, that allows the user to draw freehand imprecise yet expressive 2D lines representing the contours of facial features. A novel CNN based deep regression network is designed for inferring 3D face models from 2D sketches. Our network fuses both CNN and shape based features of the input sketch, and has two independent branches of fully connected layers generating independent subsets of coefficients for a bilinear face representation. Our system also supports gesture based interactions for users to further manipulate initial face models. Both user studies and numerical results indicate that our sketching system can help users create face models quickly and effectively. A significantly expanded face database with diverse identities, expressions and levels of exaggeration is constructed to promote further research and evaluation of face modeling techniques.
Xiaoguang Han 0001, Chang Gao 0001, Yizhou Yu
ACM Trans. Graph.3
2017 Stereoscopic Thumbnail Creation via Efficient Stereo Saliency Detection
abstract
In this paper, we propose a framework for automatically producing thumbnails from stereo image pairs. It has two components focusing respectively on stereo saliency detection and stereo thumbnail generation. The first component analyzes stereo saliency through various saliency stimuli, stereoscopic perception and the relevance between two stereo views. The second component uses stereo saliency to guide stereo thumbnail generation. We develop two types of thumbnail generation methods, both changing image size automatically. The first method is called content-persistent cropping (CPC), which aims at cropping stereo images for display devices with different aspect ratios while preserving as much content as possible. The second method is an object-aware cropping method (OAC) for generating the smallest possible thumbnail pair that retains the most important content only and facilitates quick visual exploration of a stereo image database. Quantitative and qualitative experimental evaluations demonstrate promising performance of our thumbnail generation methods in comparison to state-of-the-art algorithms.
Wenguan Wang, Jianbing Shen, Yizhou Yu, Kwan-Liu Ma
IEEE Trans. Vis. Comput. Graph.3
2016 Deep Contrast Learning for Salient Object Detection
abstract
Salient object detection has recently witnessed substantial progress due to powerful features extracted using deep convolutional neural networks (CNNs). However, existing CNN-based methods operate at the patch level instead of the pixel level. Resulting saliency maps are typically blurry, especially near the boundary of salient objects. Furthermore, image patches are treated as independent samples even when they are overlapping, giving rise to significant redundancy in computation and storage. In this paper, we propose an end-to-end deep contrast network to overcome the aforementioned limitations. Our deep network consists of two complementary components, a pixel-level fully convolutional stream and a segment-wise spatial pooling stream. The first stream directly produces a saliency map with pixel-level accuracy from an input image. The second stream extracts segment-wise features very efficiently, and better models saliency discontinuities along object boundaries. Finally, a fully connected CRF model can be optionally incorporated to improve spatial coherence and contour localization in the fused result from these two streams. Experimental results demonstrate that our deep model significantly improves the state of the art.
Guanbin Li, Yizhou Yu
CVPR2
2016 LSTM-CF: Unifying Context Modeling and Fusion with LSTMs for RGB-D Scene Labeling
Zhen Li 0026, Yukang Gan, Xiaodan Liang, Yizhou Yu, Liang Lin 0004
ECCV (2)4
2016 Protein Secondary Structure Prediction Using Cascaded Convolutional and Recurrent Neural Networks
Zhen Li 0026, Yizhou Yu
IJCAI2
2016 Visual Saliency Detection Based on Multiscale Deep CNN Features
abstract
Visual saliency is a fundamental problem in both cognitive and computational sciences, including computer vision. In this paper, we discover that a high-quality visual saliency model can be learned from multiscale features extracted using deep convolutional neural networks (CNNs), which have had many successes in visual recognition tasks. For learning such saliency models, we introduce a neural network architecture, which has fully connected layers on top of CNNs responsible for feature extraction at three different scales. The penultimate layer of our neural network has been confirmed to be a discriminative high-level feature vector for saliency detection, which we call deep contrast feature. To generate a more robust feature, we integrate handcrafted low-level features with our deep contrast feature. To promote further research and evaluation of visual saliency models, we also construct a new large database of 4447 challenging images and their pixelwise saliency annotations. Experimental results demonstrate that our proposed method is capable of achieving the state-of-the-art performance on all public benchmarks, improving the F-measure by 6.12% and 10%, respectively, on the DUT-OMRON data set and our new data set (HKU-IS), and lowering the mean absolute error by 9% and 35.3%, respectively, on these two data sets.
Guanbin Li, Yizhou Yu
IEEE Trans. Image Process.2
2016 Fast and exact discrete geodesic computation based on triangle-oriented wavefront propagation
abstract
Computing discrete geodesic distance over triangle meshes is one of the fundamental problems in computational geometry and computer graphics. In this problem, an effective window pruning strategy can significantly affect the actual running time. Due to its importance, we conduct an in-depth study of window pruning operations in this paper, and produce an exhaustive list of scenarios where one window can make another window partially or completely redundant. To identify a maximal number of redundant windows using such pairwise cross checking, we propose a set of procedures to synchronize local window propagation within the same triangle by simultaneously propagating a collection of windows from one triangle edge to its two opposite edges. On the basis of such synchronized window propagation, we design a new geodesic computation algorithm based on a triangle-oriented region growing scheme. Our geodesic algorithm can remove most of the redundant windows at the earliest possible stage, thus significantly reducing computational cost and memory usage at later stages. In addition, by adopting triangles instead of windows as the primitive in propagation management, our algorithm significantly cuts down the data management overhead. As a result, it runs 4--15 times faster than MMP and ICH algorithms, 2-4 times faster than FWP-MMP and FWP-CH algorithms, and also incurs the least memory usage.
Yipeng Qin, Xiaoguang Han 0001, Hongchuan Yu, Yizhou Yu, Jian J. Zhang 0001
ACM Trans. Graph.4
2016 Automatic Photo Adjustment Using Deep Neural Networks
abstract
Photo retouching enables photographers to invoke dramatic visual impressions by artistically enhancing their photos through stylistic color and tone adjustments. However, it is also a time-consuming and challenging task that requires advanced skills beyond the abilities of casual photographers. Using an automated algorithm is an appealing alternative to manual work, but such an algorithm faces many hurdles. Many photographic styles rely on subtle adjustments that depend on the image content and even its semantics. Further, these adjustments are often spatially varying. Existing automatic algorithms are still limited and cover only a subset of these challenges. Recently, deep learning has shown unique abilities to address hard problems. This motivated us to explore the use of deep neural networks (DNNs) in the context of photo editing. In this article, we formulate automatic photo adjustment in a manner suitable for this approach. We also introduce an image descriptor accounting for the local semantics of an image. Our experiments demonstrate that training DNNs using these descriptors successfully capture sophisticated photographic styles. In particular and unlike previous techniques, it can model local adjustments that depend on image semantics. We show that this yields results that are qualitatively and quantitatively better than previous work.
Zhicheng Yan 0001, Hao Zhang 0025, Baoyuan Wang, Sylvain Paris, Yizhou Yu
ACM Trans. Graph.5
2016 Surface Mosaic Synthesis with Irregular Tiles
abstract
Mosaics are widely used for surface decoration to produce appealing visual effects. We present a method for synthesizing digital surface mosaics with irregularly shaped tiles, which are a type of tiles often used for mosaics design. Our method employs both continuous optimization and combinatorial optimization to improve tile arrangement. In the continuous optimization step, we iteratively partition the base surface into approximate Voronoi regions of the tiles and optimize the positions and orientations of the tiles to achieve a tight fit. Combination optimization performs tile permutation and replacement to further increase surface coverage and diversify tile selection. The alternative applications of these two optimization steps lead to rich combination of tiles and high surface coverage. We demonstrate the effectiveness of our solution with extensive experiments and comparisons.
Wenchao Hu, Zhonggui Chen, Hao Pan 0001, Yizhou Yu, Eitan Grinspun, Wenping Wang 0001
IEEE Trans. Vis. Comput. Graph.4
2016 Medial Meshes - A Compact and Accurate Representation of Medial Axis Transform
abstract
The medial axis transform has long been known as an intrinsic shape representation supporting a variety of shape analysis and synthesis tasks. However, for a given shape, it is hard to obtain its faithful, concise and stable medial axis, which hinders the application of the medial axis. In this paper, we introduce the medial mesh, a new discrete representation of the medial axis. A medial mesh is a 2D simplicial complex coupled with a radius function that provides a piecewise linear approximation to the medial axis. We further present an effective algorithm for computing a concise and stable medial mesh for a given shape. Our algorithm is quantitatively driven by a shape approximation error metric, and progressively simplifies an initial medial mesh by iteratively contracting edges until the approximation error reaches a predefined threshold. We further demonstrate the superior efficiency and accuracy of our method over existing methods for medial axis simplification.
Feng Sun 0006, Yi-King Choi, Yizhou Yu, Wenping Wang 0001
IEEE Trans. Vis. Comput. Graph.3
2016 A fast modal space transform for robust nonrigid shape retrieval
Jianbo Ye, Yizhou Yu
Vis. Comput.2
2015 Visual saliency based on multiscale deep features
abstract
Visual saliency is a fundamental problem in both cognitive and computational sciences, including computer vision. In this paper, we discover that a high-quality visual saliency model can be learned from multiscale features extracted using deep convolutional neural networks (CNNs), which have had many successes in visual recognition tasks. For learning such saliency models, we introduce a neural network architecture, which has fully connected layers on top of CNNs responsible for feature extraction at three different scales. We then propose a refinement method to enhance the spatial coherence of our saliency results. Finally, aggregating multiple saliency maps computed for different levels of image segmentation can further boost the performance, yielding saliency maps better than those generated from a single segmentation. To promote further research and evaluation of visual saliency models, we also construct a new large database of 4447 challenging images and their pixelwise saliency annotations. Experimental results demonstrate that our proposed method is capable of achieving state-of-the-art performance on all public benchmarks, improving the F-Measure by 5.0% and 13.2% respectively on the MSRA-B dataset and our new dataset (HKU-IS), and lowering the mean absolute error by 5.7% and 35.1% respectively on these two datasets.
Guanbin Li, Yizhou Yu
CVPR2
2015 Harvesting Discriminative Meta Objects with Deep CNN Features for Scene Classification
abstract
Recent work on scene classification still makes use of generic CNN features in a rudimentary manner. In this paper, we present a novel pipeline built upon deep CNN features to harvest discriminative visual objects and parts for scene classification. We first use a region proposal technique to generate a set of high-quality patches potentially containing objects, and apply a pre-trained CNN to extract generic deep features from these patches. Then we perform both unsupervised and weakly supervised learning to screen these patches and discover discriminative ones representing category-specific objects and parts. We further apply discriminative clustering enhanced with local CNN fine-tuning to aggregate similar objects and parts into groups, called meta objects. A scene image representation is constructed by pooling the feature response maps of all the learned meta objects at multiple spatial scales. We have confirmed that the scene image representation obtained using this new pipeline is capable of delivering state-of-the-art performance on two popular scene benchmark datasets, MIT Indoor 67 [22] and Sun397 [31].
Ruobing Wu, Baoyuan Wang, Yizhou Yu
ICCV4
2015 HD-CNN: Hierarchical Deep Convolutional Neural Networks for Large Scale Visual Recognition
abstract
In image classification, visual separability between different object categories is highly uneven, and some categories are more difficult to distinguish than others. Such difficult categories demand more dedicated classifiers. However, existing deep convolutional neural networks (CNN) are trained as flat N-way classifiers, and few efforts have been made to leverage the hierarchical structure of categories. In this paper, we introduce hierarchical deep CNNs (HD-CNNs) by embedding deep CNNs into a two-level category hierarchy. An HD-CNN separates easy classes using a coarse category classifier while distinguishing difficult classes using fine category classifiers. During HDCNN training, component-wise pretraining is followed by global fine-tuning with a multinomial logistic loss regularized by a coarse category consistency term. In addition, conditional executions of fine category classifiers and layer parameter compression make HD-CNNs scalable for largescale visual recognition. We achieve state-of-the-art results on both CIFAR100 and large-scale ImageNet 1000-class benchmark datasets. In our experiments, we build up three different two-level HD-CNNs, and they lower the top-1 error of the standard CNNs by 2:65%, 3:1%, and 1:1%.
Zhicheng Yan 0001, Hao Zhang 0025, Robinson Piramuthu, Vignesh Jagadeesh, Dennis DeCoste, Wei Di, Yizhou Yu
ICCV7
2015 Piecewise Flat Embedding for Image Segmentation
abstract
Image segmentation is a critical step in many computer vision tasks, including high-level visual recognition and scene understanding as well as low-level photo and video processing. In this paper, we propose a new nonlinear embedding, called piecewise flat embedding, for image segmentation. Based on the theory of sparse signal recovery, piecewise flat embedding attempts to identify segment boundaries while significantly suppressing variations within segments. We adopt an L1-regularized energy term in the formulation to promote sparse solutions. We further devise an effective two-stage numerical algorithm based on Bregman iterations to solve the proposed embedding. Piecewise flat embedding can be easily integrated into existing image segmentation frameworks, including segmentation based on spectral clustering and hierarchical segmentation based on contour detection. Experiments on BSDS500 indicate that segmentation algorithms incorporating this embedding can achieve significantly improved results in both frameworks.
Yizhou Yu, Chaowei Fang, Zicheng Liao
ICCV1
2015 Content-Aware Video2Comics With Manga-Style Layout
abstract
We introduce in this paper a new approach that conveniently converts conversational videos into comics with manga-style layout. With our approach, the manga-style layout of a comic page is achieved in a content-driven manner, and the main components, including panels and word balloons, that constitute a visually pleasing comic page are intelligently organized . Our approach extracts key frames on speakers by using a speaker detection technique such that word balloons can be placed near the corresponding speakers. We qualitatively measure the information contained in a comic page. With the initial layout automatically determined, the final comic page is obtained by maximizing such a measure and optimizing the parameters relating to the optimal display of comics. An efficient Markov chain Monte Carlo sampling algorithm is designed for the optimization. Our user study demonstrates that users much prefer our manga-style comics to purely Western style comics. Extensive experiments and comparisons against previous work also verify the effectiveness of our approach.
Guangmei Jing, Yongtao Hu 0001, Yanwen Guo 0001, Yizhou Yu, Wenping Wang 0001
IEEE Trans. Multim.4
2015 An L1 image transform for edge-preserving smoothing and scene-level intrinsic decomposition
abstract
Identifying sparse salient structures from dense pixels is a longstanding problem in visual computing. Solutions to this problem can benefit both image manipulation and understanding. In this paper, we introduce an image transform based on theL1norm for piecewise image flattening. This transform can effectively preserve and sharpen salient edges and contours while eliminating insignificant details, producing a nearly piecewise constant image with sparse structures. A variant of this image transform can perform edge-preserving smoothing more effectively than existing state-of-the-art algorithms. We further present a new method for complex scene-level intrinsic image decomposition. Our method relies on the above image transform to suppress surface shading variations, and perform probabilistic reflectance clustering on the flattened image instead of the original input image to achieve higher accuracy. Extensive testing on theIntrinsic-Images-in-the-Wilddatabase indicates our method can perform significantly better than existing techniques both visually and numerically. The obtained intrinsic images have been successfully used in two applications, surface retexturing and 3D object compositing in photographs.
Sai Bi, Xiaoguang Han 0001, Yizhou Yu
ACM Trans. Graph.3
2015 Magic decorator: automatic material suggestion for indoor digital scenes
abstract
Assigning textures and materials within 3D scenes is a tedious and labor-intensive task. In this paper, we present Magic Decorator , a system that automatically generates material suggestions for 3D indoor scenes. To achieve this goal, we introduce local material rules , which describe typical material patterns for a small group of objects or parts, and global aesthetic rules , which account for the harmony among the entire set of colors in a specific scene. Both rules are obtained from collections of indoor scene images. We cast the problem of material suggestion as a combinatorial optimization considering both local material and global aesthetic rules. We have tested our system on various complex indoor scenes. A user study indicates that our system can automatically and efficiently produce a series of visually plausible material suggestions which are comparable to those produced by artists.
Kun Xu 0003, Yizhou Yu, Tian-Yi Wang, Shi-Min Hu 0001
ACM Trans. Graph.3
2015 audeosynth: music-driven video montage
abstract
We introduce music-driven video montage, a media format that offers a pleasant way to browse or summarize video clips collected from various occasions, including gatherings and adventures. In music-driven video montage, the music drives the composition of the video content. According to musical movement and beats, video clips are organized to form a montage that visually reflects the experiential properties of the music. Nonetheless, it takes enormous manual work and artistic expertise to create it. In this paper, we develop a framework for automatically generating music-driven video montages. The input is a set of video clips and a piece of background music. By analyzing the music and video content, our system extracts carefully designed temporal features from the input, and casts the synthesis problem as an optimization and solves the parameters through Markov Chain Monte Carlo sampling. The output is a video montage whose visual activities are cut and synchronized with the rhythm of the music, rendering a symphony of audio-visual resonance.
Zicheng Liao, Yizhou Yu, Bingchen Gong, Lechao Cheng
ACM Trans. Graph.2
2014 Action-Gons: Action Recognition with a Discriminative Dictionary of Structured Elements with Varying Granularity
Yuwang Wang, Baoyuan Wang, Yizhou Yu, Qionghai Dai, Zhuowen Tu
ACCV (5)3
2014 Speaker-Following Video Subtitles
abstract
We propose a new method for improving the presentation of subtitles in video (e.g., TV and movies). With conventional subtitles, the viewer has to constantly look away from the main viewing area to read the subtitles at the bottom of the screen, which disrupts the viewing experience and causes unnecessary eyestrain. Our method places on-screen subtitles next to the respective speakers to allow the viewer to follow the visual content while simultaneously reading the subtitles. We use novel identification algorithms to detect the speakers based on audio and visual information. Then the placement of the subtitles is determined using global optimization. A comprehensive usability study indicated that our subtitle placement method outperformed both conventional fixed-position subtitling and another previous dynamic subtitling method in terms of enhancing the overall viewing experience and reducing eyestrain.
Yongtao Hu 0001, Jan Kautz, Yizhou Yu, Wenping Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2014 Optimized Synthesis of Art Patterns and Layered Textures
abstract
Line drawings and digital arts appear everywhere, from simple icons and logos to cartoons, maps, and illustrations. We define art patterns as the subset of line drawings and digital arts that are comprised of repeated elements. There exist textures that share characteristics with art patterns. Examples of such textures include piled discrete elements with curved contours. Inspired by recent success of exemplar-based texture synthesis, in this paper, we focus on synthesizing art patterns and textures with curvilinear features from exemplars, which we cast as a global optimization problem. Our energy function for this problem measures both the appearance similarity of color patterns and shape similarity of curvilinear features between an input exemplar and a synthesized image. We develop an overall expectation-maximization-style algorithm for minimizing this energy function. The shape similarity part of the energy is minimized through an innovative application of the level set method. We further generalize our energy function and optimization algorithm to multilayer pattern and texture synthesis. Our generalized optimization can effectively handle multiple layers and synthesize valid instances of interaction.
Ruobing Wu, Yizhou Yu
IEEE Trans. Vis. Comput. Graph.3
2013 SCaLE: Supervised and Cascaded Laplacian Eigenmaps for Visual Object Recognition Based on Nearest Neighbors
abstract
Recognizing the category of a visual object remains a challenging computer vision problem. In this paper we develop a novel deep learning method that facilitates example-based visual object category recognition. Our deep learning architecture consists of multiple stacked layers and computes an intermediate representation that can be fed to a nearest-neighbor classifier. This intermediate representation is discriminative and structure-preserving. It is also capable of extracting essential characteristics shared by objects in the same category while filtering out nonessential differences among them. Each layer in our model is a nonlinear mapping, whose parameters are learned through two sequential steps that are designed to achieve the aforementioned properties. The first step computes a discrete mapping called supervised Laplacian Eigenmap. The second step computes a continuous mapping from the discrete version through nonlinear regression. We have extensively tested our method and it achieves state-of-the-art recognition rates on a number of benchmark datasets.
Ruobing Wu, Yizhou Yu, Wenping Wang 0001
CVPR2
2013 Sparse similarity matrix learning for visual object retrieval
abstract
Tf-idf weighting scheme is adopted by state-of-the-art object retrieval systems to reflect the difference in discriminability between visual words. However, we argue it is only suboptimal by noting that tf-idf weighting scheme does not take quantization error into account and exploit word correlation. We view tf-idf weights as an example of diagonal Mahalanobis-type similarity matrix and generalize it into a sparse one by selectively activating off-diagonal elements. Our goal is to separate similarity of relevant images from that of irrelevant ones by a safe margin. We satisfy such similarity constraints by learning an optimal similarity metric from labeled data. An effective scheme is developed to collect training data with an emphasis on cases where the tf-idf weights violates the relative relevance constraints. Experimental results on benchmark datasets indicate the learnt similarity metric consistently and significantly outperforms the tf-idf weighting scheme.
Zhicheng Yan 0001, Yizhou Yu
IJCNN2
2013 Fast nonrigid 3D retrieval using modal space transform
abstract
Nonrigid or deformable 3D objects are common in many application domains. Retrieval of such objects in large databases based on shape similarity is still a challenging problem. In this paper, we first analyze the advantages of functional operators, and further propose a framework to design novel shape signatures for encoding nonrigid object structures. Our approach constructs a context-aware integral kernel operator on a manifold, then applies modal analysis to map this operator into a low-frequency functional representation, called fast functional transform, and finally computes its spectrum as the shape signature. Our method is fast, isometry-invariant, discriminative, and numerically stable with respect to multiple types of perturbations.
Jianbo Ye, Zhicheng Yan 0001, Yizhou Yu
ICMR3
2013 A robust high-resolution details preserving denoising algorithm for meshes
Hanqi Fan, Qunsheng Peng 0001, Yizhou Yu
Sci. China Inf. Sci.3
2013 A compact random-access representation for urban modeling and rendering
abstract
We propose a highly memory-efficient representation for modeling and rendering urban buildings composed predominantly of rectangular block structures, which can be used to completely or partially represent most modern buildings. With the proposed representation, the data size required for modeling most buildings is more than two orders of magnitude less than using the conventional mesh representation. In addition, it substantially reduces the dependency on conventional texture maps, which are not space-efficient for defining visual details of building facades. The proposed representation can be stored and transmitted as images and can be rendered directly without any mesh reconstruction. A ray-casting based shader has been developed to render buildings thus represented on the GPU with a high frame rate to support interactive fly-by as well as street-level walk-through. Comparisons with standard geometric representations and recent urban modeling techniques indicate the proposed representation performs well when viewed from a short and long distance.
Zhengzheng Kuang, Bin Chan, Yizhou Yu
ACM Trans. Graph.3
2012 A Robust Algorithm for Denoising Meshes with High-Resolution Details
Hanqi Fan, Qunsheng Peng 0001, Yizhou Yu
CVM3
2012 Learning image-specific parameters for interactive segmentation
abstract
In this paper, we present a novel interactive image segmentation technique that automatically learns segmentation parameters tailored for each and every image. Unlike existing work, our method does not require any offline parameter tuning or training stage, and is capable of determining image-specific parameters according to some simple user interactions with the target image. We formulate the segmentation problem as an inference of a conditional random field (CRF) over a segmentation mask and the target image, and parametrize this CRF by different weights (e.g., color, texture and smoothing). The weight parameters are learned via an energy margin maximization, which is solved using a constraint approximation scheme and the cutting plane method. Experimental results show that our method, by learning image-specific parameters automatically, outperforms other state-of-the-art interactive image segmentation techniques.
Zhanghui Kuang, Dirk Schnieders, Hao Zhou 0010, Kwan-Yee Kenneth Wong, Yizhou Yu
CVPR5
2012 Subspace segmentation with a Minimal Squared Frobenius Norm Representation
Siming Wei, Yizhou Yu
ICPR2
2012 Single-view hair modeling for portrait manipulation
abstract
Human hair is known to be very difficult to model or reconstruct. In this paper, we focus on applications related to portrait manipulation and take an application-driven approach to hair modeling. To enable an average user to achieve interesting portrait manipulation results, we develop a single-view hair modeling technique with modest user interaction to meet the unique requirements set by portrait manipulation. Our method relies on heuristics to generate a plausible high-resolution strand-based 3D hair model. This is made possible by an effective high-precision 2D strand tracing algorithm, which explicitly models uncertainty and local layering during tracing. The depth of the traced strands is solved through an optimization, which simultaneously considers depth constraints, layering constraints as well as regularization terms. Our single-view hair modeling enables a number of interesting applications that were previously challenging, including transferring the hairstyle of one subject to another in a potentially different pose, rendering the original portrait in a novel view and image-space hair editing.
Menglei Chai, Lvdi Wang, Yanlin Weng, Yizhou Yu, Baining Guo, Kun Zhou 0001
ACM Trans. Graph.4
2012 Object-space multiphase implicit functions
abstract
Implicit functions have a wide range of applications in entertainment, engineering and medical imaging. A standard two-phase implicit function only represents the interior and exterior of a single object. To facilitate solid modeling of heterogeneous objects with multiple internal regions, object-space multiphase implicit functions are much desired. Multiphase implicit functions have much potential in modeling natural organisms, heterogeneous mechanical parts and anatomical atlases. In this paper, we introduce a novel class of object-space multiphase implicit functions that are capable of accurately and compactly representing objects with multiple internal regions. Our proposed multiphase implicit functions facilitate true object-space geometric modeling of heterogeneous objects with non-manifold features. We present multiple methods to create object-space multiphase implicit functions from existing data, including meshes and segmented medical images. Our algorithms are inspired by machine learning algorithms for training multicategory max-margin classifiers. Comparisons demonstrate that our method achieves an error rate one order of magnitude smaller than alternative techniques.
Zhan Yuan, Yizhou Yu
ACM Trans. Graph.2
2012 Detail-Preserving Controllable Deformation from Sparse Examples
abstract
Recent advances in laser scanning technology have made it possible to faithfully scan a real object with tiny geometric details, such as pores and wrinkles. However, a faithful digital model should not only capture static details of the real counterpart but also be able to reproduce the deformed versions of such details. In this paper, we develop a data-driven model that has two components; the first accommodates smooth large-scale deformations and the second captures high-resolution details. Large-scale deformations are based on a nonlinear mapping between sparse control points and bone transformations. A global mapping, however, would fail to synthesize realistic geometries from sparse examples, for highly deformable models with a large range of motion. The key is to train a collection of mappings defined over regions locally in both the geometry and the pose space. Deformable fine-scale details are generated from a second nonlinear mapping between the control points and per-vertex displacements. We apply our modeling scheme to scanned human hand models, scanned face models, face models reconstructed from multiview video sequences, and manually constructed dinosaur models. Experiments show that our deformation models, learned from extremely sparse training data, are effective and robust in synthesizing highly deformable models with rich fine features, for keyframe animation as well as performance-driven animation. We also compare our results with those obtained by alternative techniques.
Hao-Da Huang, KangKang Yin, Ling Zhao 0006, Yizhou Yu, Xin Tong 0001
IEEE Trans. Vis. Comput. Graph.5
2012 A Subdivision-Based Representation for Vector Image Editing
abstract
Vector graphics has been employed in a wide variety of applications due to its scalability and editability. Editability is a high priority for artists and designers who wish to produce vector-based graphical content with user interaction. In this paper, we introduce a new vector image representation based on piecewise smooth subdivision surfaces, which is a simple, unified and flexible framework that supports a variety of operations, including shape editing, color editing, image stylization, and vector image processing. These operations effectively create novel vector graphics by reusing and altering existing image vectorization results. Because image vectorization yields an abstraction of the original raster image, controlling the level of detail of this abstraction is highly desirable. To this end, we design a feature-oriented vector image pyramid that offers multiple levels of abstraction simultaneously. Our new vector image representation can be rasterized efficiently using GPU-accelerated subdivision. Experiments indicate that our vector image representation achieves high visual quality and better supports editing operations than existing representations.
Zicheng Liao, Hugues Hoppe, David A. Forsyth, Yizhou Yu
IEEE Trans. Vis. Comput. Graph.4
2012 Interactive Image Segmentation Based on Level Sets of Probabilities
abstract
In this paper, we present a robust and accurate algorithm for interactive image segmentation. The level set method is clearly advantageous for image objects with a complex topology and fragmented appearance. Our method integrates discriminative classification models and distance transforms with the level set method to avoid local minima and better snap to true object boundaries. The level set function approximates a transformed version of pixelwise posterior probabilities of being part of a target object. The evolution of its zero level set is driven by three force terms, region force, edge field force, and curvature force. These forces are based on a probabilistic classifier and an unsigned distance transform of salient edges. We further propose a technique that improves the performance of both the probabilistic classifier and the level set method over multiple passes. It makes the final object segmentation less sensitive to user interactions. Experiments and comparisons demonstrate the effectiveness of our method.
Yugang Liu, Yizhou Yu
IEEE Trans. Vis. Comput. Graph.2
2011 Interactive 2D and volume image segmentation using level sets of probabilities
abstract
In this technical sketch, we adopt the level set method for image segmentation that integrates region statistics and edge responses. It is well-known that a serious limitation of existing level set algorithms for image segmentation is that the final result is sensitive to the location of the initialization. This is because level set evolution is typically driven by forces computed from local image data. We overcome this problem by adopting a novel level set function based on foreground probabilities, and further integrating the level set method with a probabilistic pixel classifier [Liu and Yu 2012]. Since an accurate classifier does not exist at the beginning, the segmentation framework is based on the expectation-maximization (EM) algorithm. In summary, the motivations for our method based on level sets of probabilities are manifold.
Yugang Liu, Yizhou Yu
SIGGRAPH Asia Sketches2
2011 New Technique: Sketch-based rotation editing
abstract
We present a sketch-based rotation editing system for enriching rotational motion in keyframe animations. Given a set of keyframe orientations of a rigid object, the user first edits its angular velocity trajectory by sketching curves, and then the system computes the altered rotational motion by solving a variational curve fitting problem. The solved rotational motion not only satisfies the orientation constraints at the keyframes, but also fits well the user-specified angular velocity trajectory. Our system is simple and easy to use. We demonstrate its usefulness by adding interesting and realistic rotational details to several keyframe animations.
Weiwei Xu 0003, Yizhou Yu, Yanlin Weng
J. Zhejiang Univ. Sci. C3
2011 Diffusion Kurtosis Imaging Based on Adaptive Spherical Integral
abstract
Diffusion kurtosis imaging (DKI) is a recent approach in medical engineering that has potential value for both neurological diseases and basic neuroscience research. In this letter, we develop a robust method based on adaptive spherical integral that can compute kurtosis based quantities more precisely and efficiently. Our method integrates spherical trigonometry with a recursive computational scheme to make numerical estimations in kurtosis imaging convergent. Our algorithm improves the efficiency of computing integral invariants based on reconstructed diffusion kurtosis tensors and makes DKI better prepared for further clinical applications.
Yugang Liu, Leiting Chen, Yizhou Yu
IEEE Signal Process. Lett.3
2011 Example-based image color and tone style enhancement
abstract
Color and tone adjustments are among the most frequent image enhancement operations. We define a color and tone style as a set of explicit or implicit rules governing color and tone adjustments. Our goal in this paper is to learn implicit color and tone adjustment rules from examples. That is, given a set of examples, each of which is a pair of corresponding images before and after adjustments, we would like to discover the underlying mathematical relationships optimally connecting the color and tone of corresponding pixels in all image pairs. We formally define tone and color adjustment rules as mappings, and propose to approximate complicated spatially varying nonlinear mappings in a piecewise manner. The reason behind this is that a very complicated mapping can still be locally approximated with a low-order polynomial model. Parameters within such low-order models are trained using data extracted from example image pairs. We successfully apply our framework in two scenarios, low-quality photo enhancement by transferring the style of a high-end camera, and photo enhancement using styles learned from photographers and designers.
Baoyuan Wang, Yizhou Yu, Ying-Qing Xu
ACM Trans. Graph.2
2011 Multiscale vector volumes
abstract
We introduce multiscale vector volumes , a compact vector representation for volumetric objects with complex internal structures spanning a wide range of scales. With our representation, an object is decomposed into components and each component is modeled as an SDF tree , a novel data structure that uses multiple signed distance functions (SDFs) to further decompose the volumetric component into regions. Multiple signed distance functions collectively can represent non-manifold surfaces and deliver a powerful vector representation for complex volumetric features. We use multiscale embedding to combine object components at different scales into one complex volumetric object. As a result, regions with dramatically different scales and complexities can co-exist in an object. To facilitate volumetric object authoring and editing, we have also developed a scripting language and a GUI prototype. With the help of a recursively defined spatial indexing structure, our vector representation supports fast random access, and arbitrary cross sections of complex volumetric objects can be visualized in real time.
Lvdi Wang, Yizhou Yu, Kun Zhou 0001, Baining Guo
ACM Trans. Graph.2
2010 Reconstructing diffusion kurtosis tensors from sparse noisy measurements
abstract
Diffusion kurtosis imaging (DKI) is a recent MRI based method that can quantify deviation from Gaussian behavior using a kurtosis tensor. DKI has potential value for the assessment of neurologic diseases. Existing techniques for diffusion kurtosis imaging typically need to capture hundreds of MRI images, which is not clinically feasible on human subjects. In this paper, we develop robust denoising and model fitting methods that make it possible to accurately reconstruct a kurtosis tensor from 75 or less noisy measurements. Our denoising method is based on subspace learning for multi-dimensional signals and our model fitting technique uses iterative reweighting to effectively discount the influences of outliers. The total data acquisition time thus drops significantly, making diffusion kurtosis imaging feasible for many clinical applications involving human subjects.
Yugang Liu, Siming Wei, Quan Jiang, Yizhou Yu
ICIP4
2010 Bayesian regularization of diffusion tensor images using hierarchical MCMC and loopy belief propagation
abstract
Based on the theory of Markov Random Fields, a Bayesian regularization model for diffusion tensor images (DTI) is proposed in this paper. The low-degree parameterization of diffusion tensors in our model makes it less computationally intensive to obtain a maximum a posteriori (MAP) estimation. An approximate solution to the problem is achieved efficiently using hierarchical Markov Chain Monte Carlo (HMCMC), and a loopy belief propagation algorithm is applied to a coarse grid to obtain a good initial solution for hierarchical MCMC. Experiments on synthetic and real data demonstrate the effectiveness of our methods.
Siming Wei, Jing Hua 0001, Jiajun Bu, Chun Chen 0001, Yizhou Yu
ICIP5
2010 Automatic detection of malignant prostatic gland units in cross-sectional microscopic images
abstract
Prostate cancer is the second most frequent cause of cancer deaths among men in the US. In the most reliable screening method, histological images from a biopsy are examined under a microscope by pathologists. In an early stage of prostate cancer, only relatively few gland units in a large region become malignant. Discovering such sparse malignant gland units using a microscope is a labor-intensive and error-prone task for pathologists. In this paper, we develop effective image segmentation and classification methods for automatic detection of malignant gland units in microscopic images. Both segmentation and classification methods are based on carefully designed feature descriptors, including color histograms and texton co-occurrence tables.
Yizhou Yu, Jing Hua 0001
ICIP2
2010 Feature-preserving triangular geometry images for level-of-detail representation of static and skinned meshes
abstract
Geometry images resample meshes to represent them as texture for efficient GPU processing by forcing a regular parameterization that often incurs a large amount of distortion. Previous approaches broke the geometry image into multiple rectangular or irregular charts to reduce distortion, but complicated the automatic level of detail one gets from MIP-maps of the geometry image. We introduce triangular-chart geometry images and show this new approach better supports the GPU-side representation and display of skinned dynamic meshes, with support for feature preservation, bounding volumes, and view-dependent level of detail. Triangular charts pack efficiently, simplify the elimination of T-junctions, arise naturally from an edge-collapse simplification base mesh, and layout more flexibly to allow their edges to follow curvilinear mesh features. To support the construction and application of triangular-chart geometry images, this article introduces a new spectral clustering method for feature detection, and new methods for incorporating skinning weights and skinned bounding boxes into the representation. This results in a tenfold improvement in fidelity when compared to quad-chart geometry images.
Wei-Wen Feng, Byung-Uck Kim, Yizhou Yu, John Hart
ACM Trans. Graph.3
2010 A deformation transformer for real-time cloth animation
abstract
Achieving interactive performance in cloth animation has significant implications in computer games and other interactive graphics applications. Although much progress has been made, it is still much desired to have real-time high-quality results that well preserve dynamic folds and wrinkles. In this paper, we introduce a hybrid method for real-time cloth animation. It relies on data-driven models to capture the relationship between cloth deformations at two resolutions. Such data-driven models are responsible for transforming low-quality simulated deformations at the low resolution into high-resolution cloth deformations with dynamically introduced fine details. Our data-driven transformation is trained using rotation invariant quantities extracted from the cloth models, and is independent of the simulation technique chosen for the lower resolution model. We have also developed a fast collision detection and handling scheme based on dynamically transformed bounding volumes. All the components in our algorithm can be efficiently implemented on programmable graphics hardware to achieve an overall real-time performance on high-resolution cloth models.
Wei-Wen Feng, Yizhou Yu, Byung-Uck Kim
ACM Trans. Graph.2
2010 Data-driven image color theme enhancement
abstract
It is often important for designers and photographers to convey or enhance desired color themes in their work. A color theme is typically defined as a template of colors and an associated verbal description. This paper presents a data-driven method for enhancing a desired color theme in an image. We formulate our goal as a unified optimization that simultaneously considers a desired color theme, texture-color relationships as well as automatic or user-specified color constraints. Quantifying the difference between an image and a color theme is made possible by color mood spaces and a generalization of an additivity relationship for two-color combinations. We incorporate prior knowledge, such as texture-color relationships, extracted from a database of photographs to maintain a natural look of the edited images. Experiments and a user study have confirmed the effectiveness of our method.
Baoyuan Wang, Yizhou Yu, Tien-Tsin Wong, Chun Chen 0001, Ying-Qing Xu
ACM Trans. Graph.2
2010 Vector solid textures
abstract
In this paper, we introduce a compact random-access vector representation for solid textures made of intermixed regions with relatively smooth internal color variations. It is feature-preserving and resolution-independent. In this representation, a texture volume is divided into multiple regions. Region boundaries are implicitly defined using a signed distance function. Color variations within the regions are represented using compactly supported radial basis functions (RBFs). With a spatial indexing structure, such RBFs enable efficient color evaluation during real-time solid texture mapping. Effective techniques have been developed for generating such a vector representation from bitmap solid textures. Data structures and techniques have also been developed to compactly store region labels and distance values for efficient random access during boundary and color evaluation.
Lvdi Wang, Kun Zhou 0001, Yizhou Yu, Baining Guo
ACM Trans. Graph.3
2010 Robust Feature-Preserving Mesh Denoising Based on Consistent Subneighborhoods
abstract
In this paper, we introduce a feature-preserving denoising algorithm. It is built on the premise that the underlying surface of a noisy mesh is piecewise smooth, and a sharp feature lies on the intersection of multiple smooth surface regions. A vertex close to a sharp feature is likely to have a neighborhood that includes distinct smooth segments. By defining the consistent subneighborhood as the segment whose geometry and normal orientation most consistent with those of the vertex, we can completely remove the influence from neighbors lying on other segments during denoising. Our method identifies piecewise smooth subneighborhoods using a robust density-based clustering algorithm based on shared nearest neighbors. In our method, we obtain an initial estimate of vertex normals and curvature tensors by robustly fitting a local quadric model. An anisotropic filter based on optimal estimation theory is further applied to smooth the normal field and the curvature tensor field. This is followed by second-order bilateral filtering, which better preserves curvature details and alleviates volume shrinkage during denoising. The support of these filters is defined by the consistent subneighborhood of a vertex. We have applied this algorithm to both generic and CAD models, and sharp features, such as edges and corners, are very well preserved.
Hanqi Fan, Yizhou Yu, Qunsheng Peng 0001
IEEE Trans. Vis. Comput. Graph.2
2010 Real-time data driven deformation with affine bones
Byung-Uck Kim, Wei-Wei Feng, Yizhou Yu
Vis. Comput.3
2010 Erratum to: Real-time data driven deformation with affine bones
Byung-Uck Kim, Wei-Wen Feng, Yizhou Yu
Vis. Comput.3
2010 Lazy texture selection based on active learning
Qing Wu 0006, Chun Chen 0001, Yizhou Yu
Vis. Comput.4
2009 Hierarchical and wavelet-based multilinear models for multi-dimensional visual data approximation
abstract
With advances in imaging technologies such as CCD, laser, magnetic resonance, and diffusion tensor; visual data of multiple dimensions have been produced at an unprecedented rate and scale. These new technologies bring new challenges to existing multidimensional image compression techniques. In ["Hierarchical tensor approximation of multidimensional images" and "Hierarchical tensor approximation of multidimensional visual data" by Q. Wu et. al.] we exploit the aforementioned characteristics of visual data and develop a compact representation technique based on a hierarchical tensor based transformation. In this technique, an original multidimensional dataset is transformed into a hierarchy of signals to expose its multiscale structures. In ["Wavelet based hybrid multilinear models for multidimensional image approximation" by Q. Wu et. al.] we propose hybrid multilinear models in the wavelet domain to harness the power of both wavelet (packet) transforms and tensor approximation.
Yizhou Yu
CAD/Graphics1
2009 Example-based hair geometry synthesis
abstract
We present an example-based approach to hair modeling because creating hairstyles either manually or through image-based acquisition is a costly and time-consuming process. We introduce a hierarchical hair synthesis framework that views a hairstyle both as a 3D vector field and a 2D arrangement of hair strands on the scalp. Since hair forms wisps, a hierarchical hair clustering algorithm has been developed for detecting wisps in example hairstyles. The coarsest level of the output hairstyle is synthesized using traditional 2D texture synthesis techniques. Synthesizing finer levels of the hierarchy is based on cluster oriented detail transfer. Finally, we compute a discrete tangent vector field from the synthesized hair at every level of the hierarchy to remove undesired inconsistencies among hair trajectories. Improved hair trajectories can be extracted from the vector field. Based on our automatic hair synthesis method, we have also developed simple user-controlled synthesis and editing techniques including feature-preserving combing as well as detail transfer between different hairstyles.
Lvdi Wang, Yizhou Yu, Kun Zhou 0001, Baining Guo
ACM Trans. Graph.2
2009 Patch-based image vectorization with automatic curvilinear feature alignment
abstract
Raster image vectorization is increasingly important since vector-based graphical contents have been adopted in personal computers and on the Internet. In this paper, we introduce an effective vector-based representation and its associated vectorization algorithm for full-color raster images. There are two important characteristics of our representation. First, the image plane is decomposed into nonoverlapping parametric triangular patches with curved boundaries. Such a simplicial layout supports a flexible topology and facilitates adaptive patch distribution. Second, a subset of the curved patch boundaries are dedicated to faithfully representing curvilinear features. They are automatically aligned with the features. Because of this, patches are expected to have moderate internal variations that can be well approximated using smooth functions. We have developed effective techniques for patch boundary optimization and patch color fitting to accurately and compactly approximate raster images with both smooth variations and curvilinear features. A real-time GPU-accelerated parallel algorithm based on recursive patch subdivision has also been developed for rasterizing a vectorized image. Experiments and comparisons indicate our image vectorization algorithm achieves a more accurate and compact vector-based representation than existing ones do.
Binbin Liao, Yizhou Yu
ACM Trans. Graph.3
2008 Wavelet-based hybrid multilinear models for multidimensional image approximation
abstract
The wavelet transform hierarchically decomposes images with prescribed bases, while multilinear models search for optimal bases to adapt visual data. In this paper, we integrate these two approaches to compactly represent 2D images and 3D volume data. Once a wavelet (packet) decomposition has been performed, the coefficients are subdivided into small blocks most of which have small energy and are pruned. Surviving blocks usually exhibit strong redundancy among different channels and subbands. To exploit this property, we organize the surviving blocks into small tensors, group the tensors into clusters using an EM algorithm, and compactly approximate each cluster using tensor ensemble approximation. Experimental results on images and medical volume data indicate that our approach achieves better approximation quality than wavelet (packet) transforms.
Qing Wu 0006, Chun Chen 0001, Yizhou Yu
ICIP3
2008 Shape-constrained flock animation
abstract
Abstract We propose a novel shape‐constrained flock animation system for interactively controlling flock navigation in virtual environments. This system is capable of making the spatial distribution of a flock meet static or deforming shape constraints while performing flock simulation. Such a capability can find many applications in the entertainment industry. Given a 3D constraining shape, our system first draws a set of uniform sample points through a 3D surface mosaicing process or a stratified point sampling strategy. Once correspondences between flock members and sample points have been established, points on the target shape are used as homing destinations to guide flock migration. Under a global path control scheme, an effective fuzzy control logic, which dynamically adjusts steering forces and control forces, has been developed to create visually pleasing shape‐constrained flock animations. Copyright © 2008 John Wiley & Sons, Ltd.
Jiayi Xu 0002, Xiaogang Jin 0001, Yizhou Yu, Tian Shen, Mingdong Zhou
Comput. Animat. Virtual Worlds3
2008 Real-time data driven deformation using kernel canonical correlation analysis
abstract
Achieving intuitive control of animated surface deformation while observing a specific style is an important but challenging task in computer graphics. Solutions to this task can find many applications in data-driven skin animation, computer puppetry, and computer games. In this paper, we present an intuitive and powerful animation interface to simultaneously control the deformation of a large number of local regions on a deformable surface with a minimal number of control points. Our method learns suitable deformation subspaces from training examples, and generate new deformations on the fly according to the movements of the control points. Our contributions include a novel deformation regression method based on kernel Canonical Correlation Analysis (CCA) and a Poisson-based translation solving technique for easy and fast deformation control based on examples. Our run-time algorithm can be implemented on GPUs and can achieve a few hundred frames per second even for large datasets with hundreds of training examples.
Wei-Wen Feng, Byung-Uck Kim, Yizhou Yu
ACM Trans. Graph.3
2008 Hierarchical Tensor Approximation of Multi-Dimensional Visual Data
abstract
Visual data comprise of multi-scale and inhomogeneous signals. In this paper, we exploit these characteristics and develop a compact data representation technique based on a hierarchical tensor-based transformation. In this technique, an original multi-dimensional dataset is transformed into a hierarchy of signals to expose its multi-scale structures. The signal at each level of the hierarchy is further divided into a number of smaller tensors to expose its spatially inhomogeneous structures. These smaller tensors are further transformed and pruned using a tensor approximation technique. Our hierarchical tensor approximation supports progressive transmission and partial decompression. Experimental results indicate that our technique can achieve higher compression ratios and quality than previous methods, including wavelet transforms, wavelet packet transforms, and single-level tensor approximation. We have successfully applied our technique to multiple tasks involving multi-dimensional visual data, including medical and scientific data visualization, data-driven rendering and texture synthesis.
Qing Wu 0006, Chun Chen 0001, Hsueh-Yi Sean Lin, Yizhou Yu
IEEE Trans. Vis. Comput. Graph.6
2007 Hierarchical Tensor Approximation of Multidimensional Images
abstract
Visual data comprises of multi-scale and inhomogeneous signals. In this paper, we exploit these characteristics and develop an adaptive data approximation technique based on a hierarchical tensor-based transformation. In this technique, an original multi-dimensional image is transformed into a hierarchy of signals to expose its multi-scale structures. The signal at each level of the hierarchy is further divided into a number of smaller tensors to expose its spatially inhomogeneous structures. These smaller tensors are further transformed and pruned using a collective tensor approximation technique. Experimental results indicate that our technique can achieve higher compression ratios than existing functional approximation methods, including wavelet transforms, wavelet packet transforms and single-level tensor approximation.
Qing Wu 0006, Yizhou Yu
ICIP (4)3
2007 Transplanting and Editing Animations on Skinned Meshes
abstract
Skinned Mesh Animation (SMA) well approximates a mesh animation with extracted bones and their transformations. However, unlike skeleton, bones in SMA are not organized in hierarchies, thus they need mesh dependent translation vectors which prevent other sources of motion (i.e. skeletal animations, MoCAP, SMAs etc) from being applied to the skinned mesh. In this paper, we propose a new and fast method to transplant motion to skinned meshes. By efficiently solving a linear least-squares system, we can compute new translation vectors which enable the motion to work on the skinned mesh. Based on the same idea, we have also devised a SMA editing tool which allows users to edit frames of the SMA interactively. Furthermore, the editing can be propagated to all subsequent frames.
Yuntao Jia, Wei-Wen Feng, Yizhou Yu
PG3
2007 Laplacian Guided Editing, Synthesis, and Simulation
abstract
Summary form only given. The Laplacian has been playing a central role in numerous scientific and engineering problems. It has also become popular in computer graphics. This talk presents a series of our work that exploits the Laplacian in mesh editing, texture synthesis and flow simulation. First, a review is given on mesh editing using differential coordinates and the Poisson equation, which involves the Laplacian. The distinctive feature of this approach is that it modifies the original geometry implicitly through gradient field manipulation. This approach can produce desired results for both global and local editing operations, such as deformation, object merging, and denoising. This technique is computationally involved since it requires solving a large sparse linear system. To overcome this difficulty, an efficient multigrid algorithm specifically tailored for geometry processing has been developed. This multigrid algorithm is capable of interactively processing meshes with hundreds of thousands of vertices. In our latest work, Laplacian-based editing has been generalized to deforming mesh sequences, and efficient user interaction techniques have also been designed. Second, this talk presents a Laplacian-based method for surface texture synthesis and mixing from multiple sources. Eliminating seams among texture patches is important during texture synthesis. In our technique, it is solved by performing Laplacian texture reconstruction, which retains the high frequency details but computes new consistent low frequency components. Third, a method for inviscid flow simulation over manifold surfaces is presented. This method enforces incompressibility on closed surfaces by solving a discrete Poisson equation. Different from previous work, it performs simulations directly on triangle meshes and thus eliminates parametrization distortions.
Yizhou Yu
PG1
2007 Large-Scale Data Management for PRT-Based Real-Time Rendering of Dynamically Skinned Models
Wei-Wen Feng, Yuntao Jia, Yizhou Yu
Rendering Techniques4
2007 Gradient domain editing of deforming mesh sequences
abstract
Many graphics applications, including computer games and 3D animated films, make heavy use of deforming mesh sequences. In this paper, we generalize gradient domain editing to deforming mesh sequences. Our framework is keyframe based. Given sparse and irregularly distributed constraints at unevenly spaced keyframes, our solution first adjusts the meshes at the keyframes to satisfy these constraints, and then smoothly propagate the constraints and deformations at keyframes to the whole sequence to generate new deforming mesh sequence. To achieve convenient keyframe editing, we have developed an efficient alternating least-squares method. It harnesses the power of subspace deformation and two-pass linear methods to achieve high-quality deformations. We have also developed an effective algorithm to define boundary conditions for all frames using handle trajectory editing. Our deforming mesh editing framework has been successfully applied to a number of editing scenarios with increasing complexity, including footprint editing, path editing, temporal filtering, handle-based deformation mixing, and spacetime morphing.
Weiwei Xu 0003, Kun Zhou 0001, Yizhou Yu, Qifeng Tan, Qunsheng Peng 0001, Baining Guo
ACM Trans. Graph.3
2006 A fast multigrid algorithm for mesh deformation
abstract
In this paper, we present a multigrid technique for efficiently deforming large surface and volume meshes. We show that a previous least-squares formulation for distortion minimization reduces to a Laplacian system on a general graph structure for which we derive an analytic expression. We then describe an efficient multigrid algorithm for solving the relevant equations. Here we develop novel prolongation and restriction operators used in the multigrid cycles. Combined with a simple but effective graph coarsening strategy, our algorithm can outperform other multigrid solvers and the factorization stage of direct solvers in both time and memory costs for large meshes. It is demonstrated that our solver can trade off accuracy for speed to achieve greater interactivity, which is attractive for manipulating large meshes. Our multigrid solver is particularly well suited for a mesh editing environment which does not permit extensive precomputation. Experimental evidence of these advantages is provided on a number of meshes with a wide range of size. With our mesh deformation solver, we also successfully demonstrate that visually appealing mesh animations can be generated from both motion capture data and a single base mesh even when they are inconsistent.
Yizhou Yu, Nathan Bell, Wei-Wen Feng
ACM Trans. Graph.2
2005 Shadow Graphs and 3D Texture Reconstruction
Yizhou Yu, Johnny T. Chang
Int. J. Comput. Vis.1
2005 Controllable smoke animation with guiding objects
abstract
This article addresses the problem of controlling the density and dynamics of smoke (a gas phenomenon) so that the synthetic appearance of the smoke (gas) resembles a still or moving object. Both the smoke region and the target object are represented as implicit functions. As a part of the target implicit function, a shape transformation is generated between an initial smoke region and the target object. In order to match the smoke surface with the target surface, we impose carefully designed velocity constraints on the smoke boundary during a dynamic fluid simulation. The velocity constraints are derived from an iterative functional minimization procedure for shape matching. The dynamics of the smoke is formulated using a novel compressible fluid model which can effectively absorb the discontinuities in the velocity field caused by imposed velocity constraints while reproducing realistic smoke appearances. As a result, a smoke region can evolve into a regular object and follow the motion of the object, while maintaining its smoke appearance.
Yizhou Yu
ACM Trans. Graph.2
2005 Out-of-core tensor approximation of multi-dimensional matrices of visual data
abstract
Tensor approximation is necessary to obtain compact multilinear models for multi-dimensional visual datasets. Traditionally, each multi-dimensional data item is represented as a vector. Such a scheme flattens the data and partially destroys the internal structures established throughout the multiple dimensions. In this paper, we retain the original dimensionality of the data items to more effectively exploit existing spatial redundancy and allow more efficient computation. Since the size of visual datasets can easily exceed the memory capacity of a single machine, we also present an out-of-core algorithm for higher-order tensor approximation. The basic idea is to partition a tensor into smaller blocks and perform tensor-related operations blockwise. We have successfully applied our techniques to three graphics-related data-driven models, including 6D bidirectional texture functions, 7D dynamic BTFs and 4D volume simulation sequences. Experimental results indicate that our techniques can not only process out-of-core data, but also achieve higher compression ratios and quality than previous methods.
Qing Wu 0006, Yizhou Yu, Narendra Ahuja
ACM Trans. Graph.4
2005 Controllable motion synthesis in a gaseous medium
Yizhou Yu, Christopher Wojtan, Stephen Chenney
Vis. Comput.2
2005 Photogrammetric reconstruction of free-form objects with curvilinear structures
Yizhou Yu
Vis. Comput.2
2004 Reconstruction of 3-D Symmetric Curves from Perspective Images without Discrete Features
Wei Hong 0003, Yi Ma 0001, Yizhou Yu
ECCV (3)3
2004 Inviscid and incompressible fluid simulation on triangle meshes
abstract
Abstract Simulating fluid motion on manifold surfaces is an interesting but rarely explored area because of the difficulty of establishing plausible physical models. In this paper, we introduce a novel method for inviscid fluid simulation over meshes. It can enforce incompressibility on closed surfaces by utilizing a discrete vector field decomposition algorithm. It also includes effective implementations of semi‐Lagrangian tracing and velocity interpolation schemes. Different from previous work, our method performs simulations directly on triangle meshes and thus eliminates parametrization distortions. Our implementation can produce convincing fluid motion on surfaces and has interactive performance for meshes with tens of thousands of faces. Copyright © 2004 John Wiley & Sons, Ltd.
Yizhou Yu
Comput. Animat. Virtual Worlds2
2004 Video metamorphosis using dense flow fields
abstract
Abstract When we perform video metamorphosis, it would be desirable to make smooth morphing transitions simultaneously along with the original motion in the image sequences. In this paper, we present a novel semi‐automatic video morphing technique that exhibits this behavior. Our technique effectively exploits temporal coherence and automatic image matching. One‐to‐one dense mappings between pairs of corresponding frames are obtained by applying a compositing procedure and a hierarchical image matching technique. These dense mappings can be initialized with sparse frame‐to‐frame feature correspondences obtained semi‐automatically by integrating a friendly user interface with a robust feature tracking algorithm. Experimental results show that our approach to video metamorphosis can produce superior results. Copyright © 2004 John Wiley & Sons, Ltd.
Yizhou Yu, Qing Wu 0006
Comput. Animat. Virtual Worlds1
2004 Feature matching and deformation for texture synthesis
abstract
One significant problem in patch-based texture synthesis is the presence of broken features at the boundary of adjacent patches. The reason is that optimization schemes for patch merging may fail when neighborhood search cannot find satisfactory candidates in the sample texture because of an inaccurate similarity measure. In this paper, we consider both curvilinear features and their deformation. We develop a novel algorithm to perform feature matching and alignment by measuring structural similarity. Our technique extracts a feature map from the sample texture, and produces both a new feature map and texture map. Texture synthesis guided by feature maps can significantly reduce the number of feature discontinuities and related artifacts, and gives rise to satisfactory results.
Qing Wu 0006, Yizhou Yu
ACM Trans. Graph.2
2004 Mesh editing with poisson-based gradient field manipulation
abstract
In this paper, we introduce a novel approach to mesh editing with the Poisson equation as the theoretical foundation. The most distinctive feature of this approach is that it modifies the original mesh geometry implicitly through gradient field manipulation. Our approach can produce desirable and pleasing results for both global and local editing operations, such as deformation, object merging, and smoothing. With the help from a few novel interactive tools, these operations can be performed conveniently with a small amount of user interaction. Our technique has three key components, a basic mesh solver based on the Poisson equation, a gradient field manipulation scheme using local transforms, and a generalized boundary condition representation based on local frames. Experimental results indicate that our framework can outperform previous related mesh editing techniques.
Yizhou Yu, Kun Zhou 0001, Dong Xu 0001, Hujun Bao, Baining Guo, Harry Shum
ACM Trans. Graph.1
2002 Shadow Graphs and Surface Reconstruction
Yizhou Yu, Johnny T. Chang
ECCV (2)1
2002 Pattern-Based Texture Metamorphosis
abstract
In this paper we study texture metamorphosis, or how to generate texture samples that smoothly transform from a source texture image to a target. We propose a pattern-based approach to specify the feature correspondence between two textures, based on the observation that man), texture images have stochastically distributed patterns which are similar to each other First, the user selects a pattern in the source and target textures, and establishes the "local feature correspondence" between these two patterns by specifying landmarks. Then, repeated patterns are automatically detected and localized in the source and target textures. The "pattern correspondence" between two textures is formulated as an integer programming problem and solved using the Hungarian algorithm. Finally, we obtain a warp function between two textures by combining "local feature correspondence" and "pattern correspondence". Experiments demonstrate that our technique produces visually appealing morphing sequences, with moderate amount of user interaction.
Ce Liu 0001, Harry Shum, Yizhou Yu
PG4
2001 Modeling Realistic Virtual Hairstyles
abstract
The author presents an effective method for modeling realistic curly hairstyles, taking into account both artificial hairstyling processes and natural curliness. The result is a detailed geometric model of hairs that can be rendered and animated via existing methods. Our technique exploits the analogy between hairs and a vector field; it interactively and efficiently models global and local hair flows by superimposing procedurally defined vector field primitives that have local influence. Usually only a very small number of vector field primitives are needed to model a complicated hairstyle. An initial model of hair strands is extracted from the superimposed vector fields by tracing their field lines. Random natural or artificial curliness can be added to the initial model through a parametric hair offset function with a randomized distribution of parameters over the scalp. Techniques for shearing and clustering are also designed to improve the overall appearance of the hair model. Our technique has been successfully applied to generate a variety of realistic hairstyles with different curliness and length distributions.
Yizhou Yu
PG1
2001 Synthesizing bidirectional texture functions for real-world surfaces
abstract
In this paper, we present a novel approach to synthetically generating bidirectional texture functions (BTFs) of real-world surfaces. Unlike a conventional two-dimensional texture, a BTF is a six-dimensional function that describes the appearance of texture as a function of illumination and viewing directions. The BTF captures the appearance change caused by visible small-scale geometric details on surfaces. From a sparse set of images under different viewing/lighting settings, our approach generates BTFs in three steps. First, it recovers approximate 3D geometry of surface details using a shape-from-shading method. Then, it generates a novel version of the geometric details that has the same statistical properties as the sample surface with a non-parametric sampling method. Finally, it employs an appearance preserving procedure to synthesize novel images for the recovered or generated geometric details under various viewing/lighting settings, which then define a BTF. Our experimental results demonstrate the effectiveness of our approach.
Xinguo Liu, Yizhou Yu, Harry Shum
SIGGRAPH2
2001 Extracting Objects from Range and Radiance Images
abstract
In this paper, we present a pipeline and several key techniques necessary for editing a real scene captured with both cameras and laser range scanners. We develop automatic algorithms to segment the geometry from range images into distinct surfaces, register texture from radiance images with the geometry, and synthesize compact high-quality texture maps. The result is an object-level representation of the scene which can be rendered with modifications to structure via traditional rendering methods. The segmentation algorithm for geometry operates directly on the point cloud from multiple registered 3D range images instead of a reconstructed mesh. It is a top-down algorithm which recursively partitions a point set into two subsets using a pairwise similarity measure. The result is a binary tree with individual surfaces as leaves. Our image registration technique performs a very efficient search to automatically find the camera poses for arbitrary position and orientation relative to the geometry. Thus, we can take photographs from any location without precalibration between the scanner and the camera. The algorithms have been applied to large-scale real data. We demonstrate our ability to edit a captured scene by moving, inserting, and deleting objects.
Yizhou Yu, Andras Ferencz, Jitendra Malik
IEEE Trans. Vis. Comput. Graph.1
1999 Inverse Global Illumination: Recovering Reflectance Models of Real Scenes from Photographs
abstract
In this paper we present a method for recovering the reflectance properties of all surfaces in a real scene from a sparse set of photographs, taking into account both direct and indirect illumination.The result is a lighting-independent model of the scene's geometry and reflectance properties, which can be rendered with arbitrary modifications to structure and lighting via traditional rendering methods.Our technique models reflectance with a lowparameter reflectance model, and allows diffuse albedo to vary arbitrarily over surfaces while assuming that non-diffuse characteristics remain constant across particular regions.The method's input is a geometric model of the scene and a set of calibrated high dynamic range photographs taken with known direct illumination.The algorithm hierarchically partitions the scene into a polygonal mesh, and uses image-based rendering to construct estimates of both the radiance and irradiance of each patch from the photographic data.The algorithm computes the expected location of specular highlights, and then analyzes the highlight areas in the images by running a novel iterative optimization procedure to recover the diffuse and specular reflectance parameters for each region.Lastly, these parameters are used in constructing high-resolution diffuse albedo maps for each surface.The algorithm has been applied to both real and synthetic data, including a synthetic cubical room and a real meeting room.Rerenderings are produced using a global illumination system under both original and novel lighting, and with the addition of synthetic objects.Side-by-side comparisons show success at predicting the appearance of the scene under novel lighting conditions.
Yizhou Yu, Paul E. Debevec, Jitendra Malik, Tim Hawkins
SIGGRAPH1
1999 Efficient visibility processing for projective texture mapping
Yizhou Yu
Comput. Graph.1
1998 Recovering Photometric Properties of Architectural Scenes from Photographs
abstract
In this paper, we present a new approach to producing photorealistic computer renderings of real architectural scenes under novel lighting conditions, such as at different times of day, starting from a small set of photographs of the real scene. Traditional texture mapping approaches to image-based modeling and rendering are unable to do this because texture maps are the product of the interaction between lighting and surface reflectance and one cannot deal with novel lighting without dissecting their respective contributions. To obtain this decomposition into lighting and reflectance, our basic approach is to solve a series of optimization problems to find the parameters of appropriate lighting and reflectance models that best explain the measured values in the various photographs of the scene. The lighting models include the radiance distributions from the sun and the sky, as well as the landscape to consider the effect of secondary illumination from the environment. The reflectance models are for the surfaces of the architecture. Photographs are taken for the sun, the sky, the landscape, as well as the architecture at a few different times of day to collect enough data for recovering the various lighting and reflectance models. We can predict novel illumination conditions with the recovered lighting models and use these together with the recovered reflectance values to produce renderings of the scene. Our results show that our goal of generating photorealistic renderings of real architectural scenes under novel lighting conditions has been achieved.
Yizhou Yu, Jitendra Malik
SIGGRAPH1
1997 A Rendering Equation for Specular Transfers and Its Integration into Global Illumination
abstract
In this paper, we present a rigorous theoretical formulation of the fundamental problem—indirect illumination from area sources via curved ideal specular surfaces. Intensity and area factors are introduced to clarify this problem and to rectify the radiance from these specular surfaces. They take surface geometry, such as Gaussian curvature, into account. Based on this formulation, an algorithm for integrating ideal specular transfers into global illumination is also presented. This algorithm can deal with curved specular reflectors and transmitters. An implementation is described based on wavefront tracing and progressive radiosity. Sample images generated by this method are presented.
Yizhou Yu
Comput. Graph. Forum1
1997 Parallel Progressive Radiosity with Adaptive Meshing
Yizhou Yu, Oscar H. Ibarra, Tao Yang 0009
J. Parallel Distributed Comput.1
1995 Multiresolution B-spline Radiosity
abstract
Abstract This paper introduces a kind of new wavelet radiosity method called multiresolution B‐spline radiosity, which uses B‐splines of different scales to represent radiosity distribution functions. A set of techniques and algorithms, such as function extrapolation, adaptive quadrature, scale adjustment and octree, are proposed to implement it. This method sets up hierarchical structures on surfaces, keeps radiosity distribution continuous at element boundaries, does not need postprocessing, and does not prevent the use of any surface whose parameter domain is rectilinear.
Yizhou Yu, Qunsheng Peng 0001
Comput. Graph. Forum1