Yuanhao Chen

dblp:93/479 · DBLP profile ↗
← Back
22ranked-venue papers
5as first author
3since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 4 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
16 papers
Image recognition and object detection · 35% Segmentation and scene understanding · 34% Probabilistic and Bayesian machine learning · 13%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 29 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection
object detection
0.452010
Active Mask Hierarchies for Object Detection · ECCV (5) 2010
Latent hierarchical structural learning for object detection · CVPR 2010
Unsupervised Learning of Probabilistic Grammar-Markov Models for Object Categories · IEEE Trans. Pattern Anal. Mach. Intell. 2009
Computer vision › Image recognition and object detection
object recognition
0.442012
Recursive Segmentation and Recognition Templates for Image Parsing · IEEE Trans. Pattern Anal. Mach. Intell. 2012
Unsupervised Learning of Probabilistic Object Models (POMs) for Object Classification, Segmentation, and Recognition Using Knowledge Propagation · IEEE Trans. Pattern Anal. Mach. Intell. 2009
Recursive Segmentation and Recognition Templates for 2D Parsing · NIPS 2008
Computer vision › Segmentation and scene understanding › object segmentation
object parsing
0.442011
Max Margin Learning of Hierarchical Configural Deformable Templates (HCDTs) for Efficient Object Parsing and Pose Estimation · Int. J. Comput. Vis. 2011
Part and appearance sharing: Recursive Compositional Models for multi-view · CVPR 2010
Rapid Inference on a Novel AND/OR graph for Object Detection, Segmentation and Parsing · NIPS 2007
Computer vision › Segmentation and scene understanding
object segmentation
0.332011
Max Margin Learning of Hierarchical Configural Deformable Templates (HCDTs) for Efficient Object Parsing and Pose Estimation · Int. J. Comput. Vis. 2011
Unsupervised Learning of Probabilistic Object Models (POMs) for Object Classification, Segmentation, and Recognition Using Knowledge Propagation · IEEE Trans. Pattern Anal. Mach. Intell. 2009
Unsupervised learning of probabilistic object models (POMs) for object classification, segmentation and recognition · CVPR 2008
Computer vision › Segmentation and scene understanding
scene parsing
0.222012
Recursive Segmentation and Recognition Templates for Image Parsing · IEEE Trans. Pattern Anal. Mach. Intell. 2012
Recursive Segmentation and Recognition Templates for 2D Parsing · NIPS 2008
Computer vision › Face, body and person analysis
human pose estimation
0.222011
Max Margin Learning of Hierarchical Configural Deformable Templates (HCDTs) for Efficient Object Parsing and Pose Estimation · Int. J. Comput. Vis. 2011
Max Margin AND/OR Graph learning for parsing the human body · CVPR 2008
Computer vision › Image recognition and object detection › image classification
object classification
0.232009
Unsupervised Learning of Probabilistic Object Models (POMs) for Object Classification, Segmentation, and Recognition Using Knowledge Propagation · IEEE Trans. Pattern Anal. Mach. Intell. 2009
Unsupervised learning of probabilistic object models (POMs) for object classification, segmentation and recognition · CVPR 2008
Unsupervised Learning of Probabilistic Grammar-Markov Models for Object Categories · IEEE Trans. Pattern Anal. Mach. Intell. 2009
Computer vision › Image recognition and object detection › object detection
deformable object detection
0.222010
Learning a Hierarchical Deformable Template for Rapid Deformable Object Parsing · IEEE Trans. Pattern Anal. Mach. Intell. 2010
Rapid Inference on a Novel AND/OR graph for Object Detection, Segmentation and Parsing · NIPS 2007
Computer vision › Segmentation and scene understanding
instance segmentation
0.112010
Active Mask Hierarchies for Object Detection · ECCV (5) 2010
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
latent structure discovery
0.112010
Latent hierarchical structural learning for object detection · CVPR 2010
Computer vision › Image recognition and object detection › object detection
multi-view object detection
0.112010
Part and appearance sharing: Recursive Compositional Models for multi-view · CVPR 2010
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian nonparametric model
0.112009
Nonparametric Bayesian Texture Learning and Synthesis · NIPS 2009
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian nonparametric model
hierarchical dirichlet process
0.112009
Nonparametric Bayesian Texture Learning and Synthesis · NIPS 2009
Computer vision › Segmentation and scene understanding
image segmentation
0.112009
Nonparametric Bayesian Texture Learning and Synthesis · NIPS 2009
Visual content generation and editing
texture synthesis
0.112009
Nonparametric Bayesian Texture Learning and Synthesis · NIPS 2009
Computer vision › Segmentation and scene understanding › image segmentation › model-based segmentation
deformable model segmentation
0.112008
Structure-perceptron learning of a hierarchical log-linear model · CVPR 2008
Computer vision › 3D vision › shape matching
deformable object matching
0.112008
Structure-perceptron learning of a hierarchical log-linear model · CVPR 2008
Machine learning › Probabilistic and Bayesian machine learning › hierarchical modeling › hierarchical model
hierarchical image model
0.112008
Recursive Segmentation and Recognition Templates for 2D Parsing · NIPS 2008
Computer vision › Image recognition and object detection › part-based model
hierarchical shape model
0.112008
Structure-perceptron learning of a hierarchical log-linear model · CVPR 2008
Computer vision › Segmentation and scene understanding
human parsing
0.112008
Max Margin AND/OR Graph learning for parsing the human body · CVPR 2008
Machine learning › Generative modeling › generative model
probabilistic object model
0.112008
Unsupervised learning of probabilistic object models (POMs) for object classification, segmentation and recognition · CVPR 2008
Machine learning › Representation and self-supervised learning
structure induction
0.112008
Unsupervised learning of probabilistic object models (POMs) for object classification, segmentation and recognition · CVPR 2008
Machine learning › Learning paradigms › unsupervised learning
unsupervised structure learning
0.112008
Unsupervised Structure Learning: Hierarchical Recursive Composition, Suspicious Coincidence and Competitive Exclusion · ECCV (2) 2008
Natural language and speech › Language models and text generation › grammar formalisms
probabilistic grammar
0.112006
Unsupervised Learning of a Probabilistic Grammar for Object Detection and Parsing · NIPS 2006
Machine learning › Probabilistic and Bayesian machine learning › hierarchical modeling
hierarchical model
0.012012
Recursive Segmentation and Recognition Templates for Image Parsing · IEEE Trans. Pattern Anal. Mach. Intell. 2012
Computer vision › 3D vision › 3d shape modeling
articulated object modeling
0.012011
Max Margin Learning of Hierarchical Configural Deformable Templates (HCDTs) for Efficient Object Parsing and Pose Estimation · Int. J. Comput. Vis. 2011
Computer vision › 3D vision › object modeling
shape and appearance modeling
0.012010
Learning a Hierarchical Deformable Template for Rapid Deformable Object Parsing · IEEE Trans. Pattern Anal. Mach. Intell. 2010
Machine learning › Generative modeling › variational autoencoder › hierarchical latent variable model
hierarchical compositional model
0.012008
Unsupervised Structure Learning: Hierarchical Recursive Composition, Suspicious Coincidence and Competitive Exclusion · ECCV (2) 2008
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models
0.012007
Rapid Inference on a Novel AND/OR graph for Object Detection, Segmentation and Parsing · NIPS 2007

Methods — techniques the papers use, named apart from their topics

and-or graph · 0.3dynamic programming · 0.2max-margin learning · 0.2structure-perceptron · 0.2structure induction · 0.2markov random field · 0.2knowledge propagation · 0.2recursive segmentation · 0.1deformable template · 0.1appearance sharing · 0.12d hidden markov model · 0.1
YearPublicationVenuePosition
2026 Tri-cross modal fusion network for multimodal sentiment analysis in short videos
Heyong Wang, Yuanhao Chen
Signal Process. Image Commun.2
2025 Frequency-Domain Convolutional Network With Historical Data Fusion Module for Regional Streamflow Prediction
abstract
Accurate runoff prediction is essential for effective water resource management, particularly in addressing flood control and monitoring drought conditions. However, the diverse nature of land types and varying climate conditions often complicate this task, requiring frequent adaptations to prediction models for local applications. Existing methods primarily focus on modeling for individual regions, while regional runoff prediction models cannot often learn long-term patterns, limiting their regional adaptability. To overcome this challenge, we present the temporal fusion runoff network (TFRN), a new framework designed to enhance long short-term memory (LSTM) models by enabling them to incorporate distant historical information. This innovation offers a promising framework for regional runoff prediction by enhancing model performance and minimizing computational demands. In this study, the proposed TFRN utilizes convolutional networks to extract and integrate both long-term and short-term trends from input sequences, and by merging the strengths of LSTM and Transformer architectures, TFRN achieves a thorough integration of historical data. Specifically, our method employs convolutional networks across both time and frequency domains to capture multi-scale features. Within the Transformer component, we introduce an adaptive fusion module to improve the integration of historical information. We validated the effectiveness of our model using two extensive hydrological datasets for a 7-day runoff prediction task. The results underscore the superiority of our approach, demonstrating its advantages over several leading methods. The source code is available at https://github.com/redtea-code/TFRN.
Yuanhao Chen, Haoqi Yu, Jingrong Dai, Nannan Li 0001, Changmiao Wang, Ahmed El-Azab
IEEE Trans. Geosci. Remote. Sens.2
2022 Pose estimation for workpieces in complex stacking industrial scene based on RGB images
Jianjun Yi, Yuanhao Chen, Zhiyong Dai, Shuqing Cao
Appl. Intell.3
2016 Highly efficient epidemic spreading model based LPA threshold community detection method
Yuanhao Chen
Neurocomputing3
2014 Epdemic spreading model based overlapping community detection
abstract
Community detection in inhomogeneous structured network is an attractive research problem that searches for methods to discover groups in which individuals are more densely interconnected with each other with higher probability of internal information propagation. While most of the previous approaches attempt to divide networks into communities according to the algorithm results of network or edge measurement, Label Propagation Algorithm (LPA) adopts semi-supervised machine learning and implements community detection in an intelligent way with the automatic convergent process of network entity label iteration. In this work, we study the early community detection approaches, explore the low efficacy and stagnant converging rate of LPA in its response to network with overlapped communities, and propose a new approach for community detection using epidemic spreading virus to discover groups with super positioned members. Extensive experiments in synthetic signed network and real-life large networks derived from Internet social media are conducted to explore the optimal mechanism of the most suitable community-detecting virus infection.
Yuanhao Chen
ASONAM2
2012 Recursive Segmentation and Recognition Templates for Image Parsing
abstract
In this paper, we propose a Hierarchical Image Model (HIM) which parses images to perform segmentation and object recognition. The HIM represents the image recursively by segmentation and recognition templates at multiple levels of the hierarchy. This has advantages for representation, inference, and learning. First, the HIM has a coarse-to-fine representation which is capable of capturing long-range dependency and exploiting different levels of contextual information (similar to how natural language models represent sentence structure in terms of hierarchical representations such as verb and noun phrases). Second, the structure of the HIM allows us to design a rapid inference algorithm, based on dynamic programming, which yields the first polynomial time algorithm for image labeling. Third, we learn the HIM efficiently using machine learning methods from a labeled data set. We demonstrate that the HIM is comparable with the state-of-the-art methods by evaluation on the challenging public MSRC and PASCAL VOC 2007 image data sets.
Long Zhu, Yuanhao Chen, Alan L. Yuille
IEEE Trans. Pattern Anal. Mach. Intell.2
2011 Max Margin Learning of Hierarchical Configural Deformable Templates (HCDTs) for Efficient Object Parsing and Pose Estimation
abstract
In this paper we formulate a hierarchical configurable deformable template (HCDT) to model articulated visual objects—such as horses and baseball players—for tasks such as parsing, segmentation, and pose estimation. HCDTs represent an object by an AND/OR graph where the OR nodes act as switches which enables the graph topology to vary adaptively. This hierarchical representation is compositional and the node variables represent positions and properties of subparts of the object. The graph and the node variables are required to obey the summarization principle which enables an efficient compositional inference algorithm to rapidly estimate the state of the HCDT. We specify the structure of the AND/OR graph of the HCDT by hand and learn the model parameters discriminatively by extending Max-Margin learning to AND/OR graphs. We illustrate the three main aspects of HCDTs—representation, inference, and learning—on the tasks of segmenting, parsing, and pose (configuration) estimation for horses and humans. We demonstrate that the inference algorithm is fast and that max-margin learning is effective. We show that HCDTs gives state of the art results for segmentation and pose estimation when compared to other methods on benchmarked datasets.
Long Zhu, Yuanhao Chen, Alan L. Yuille
Int. J. Comput. Vis.2
2010 Part and appearance sharing: Recursive Compositional Models for multi-view
abstract
We propose Recursive Compositional Models (RCMs) for simultaneous multi-view multi-object detection and parsing (e.g. view estimation and determining the positions of the object subparts). We represent the set of objects by a family of RCMs where each RCM is a probability distribution defined over a hierarchical graph which corresponds to a specific object and viewpoint. An RCM is constructed from a hierarchy of subparts/subgraphs which are learnt from training data. Part-sharing is used so that different RCMs are encouraged to share subparts/subgraphs which yields a compact representation for the set of objects and which enables efficient inference and learning from a limited number of training samples. In addition, we use appearance-sharing so that RCMs for the same object, but different viewpoints, share similar appearance cues which also helps efficient learning. RCMs lead to a multi-view multi-object detection system. We illustrate RCMs on four public datasets and achieve state-of-the-art performance.
Long Zhu, Yuanhao Chen, Antonio Torralba 0001, William T. Freeman, Alan L. Yuille
CVPR2
2010 Latent hierarchical structural learning for object detection
abstract
We present a latent hierarchical structural learning method for object detection. An object is represented by a mixture of hierarchical tree models where the nodes represent object parts. The nodes can move spatially to allow both local and global shape deformations. The models can be trained discriminatively using latent structural SVM learning, where the latent variables are the node positions and the mixture component. But current learning methods are slow, due to the large number of parameters and latent variables, and have been restricted to hierarchies with two layers. In this paper we describe an incremental concave-convex procedure (iCCCP) which allows us to learn both two and three layer models efficiently. We show that iCCCP leads to a simple training algorithm which avoids complex multi-stage layer-wise training, careful part selection, and achieves good performance without requiring elaborate initialization. We perform object detection using our learnt models and obtain performance comparable with state-of-the-art methods when evaluated on challenging public PASCAL datasets. We demonstrate the advantages of three layer hierarchies - outperforming Felzenszwalb et al.'s two layer models on all 20 classes.
Long Zhu, Yuanhao Chen, Alan L. Yuille, William T. Freeman
CVPR2
2010 Active Mask Hierarchies for Object Detection
Yuanhao Chen, Long Zhu, Alan L. Yuille
ECCV (5)1
2010 Learning a Hierarchical Deformable Template for Rapid Deformable Object Parsing
abstract
In this paper, we address the tasks of detecting, segmenting, parsing, and matching deformable objects. We use a novel probabilistic object model that we call a hierarchical deformable template (HDT). The HDT represents the object by state variables defined over a hierarchy (with typically five levels). The hierarchy is built recursively by composing elementary structures to form more complex structures. A probability distribution--a parameterized exponential model--is defined over the hierarchy to quantify the variability in shape and appearance of the object at multiple scales. To perform inference--to estimate the most probable states of the hierarchy for an input image--we use a bottom-up algorithm called compositional inference. This algorithm is an approximate version of dynamic programming where approximations are made (e.g., pruning) to ensure that the algorithm is fast while maintaining high performance. We adapt the structure-perceptron algorithm to estimate the parameters of the HDT in a discriminative manner (simultaneously estimating the appearance and shape parameters). More precisely, we specify an exponential distribution for the HDT using a dictionary of potentials, which capture the appearance and shape cues. This dictionary can be large and so does not require handcrafting the potentials. Instead, structure-perceptron assigns weights to the potentials so that less important potentials receive small weights (this is like a "soft" form of feature selection). Finally, we provide experimental evaluation of HDTs on different visual tasks, including detection, segmentation, matching (alignment), and parsing. We show that HDTs achieve state-of-the-art performance for these different tasks when evaluated on data sets with groundtruth (and when compared to alternative algorithms, which are typically specialized to each task).
Long Zhu, Yuanhao Chen, Alan L. Yuille
IEEE Trans. Pattern Anal. Mach. Intell.2
2009 Nonparametric Bayesian Texture Learning and Synthesis
abstract
We present a nonparametric Bayesian method for texture learning and synthesis. A texture image is represented by a 2D-Hidden Markov Model (2D-HMM) where the hidden states correspond to the cluster labeling of textons and the transition matrix encodes their spatial layout (the compatibility between adjacent textons). 2D-HMM is coupled with the Hierarchical Dirichlet process (HDP) which allows the number of textons and the complexity of transition matrix grow as the input texture becomes irregular. The HDP makes use of Dirichlet process prior which favors regular textures by penalizing the model complexity. This framework (HDP-2D-HMM) learns the texton vocabulary and their spatial layout jointly and automatically. The HDP-2D-HMM results in a compact representation of textures which allows fast texture synthesis with comparable rendering quality over the state-of-the-art image-based rendering methods. We also show that HDP-2D-HMM can be applied to perform image segmentation and synthesis.
Long Zhu, Yuanhao Chen, William T. Freeman, Antonio Torralba 0001
NIPS2
2009 Unsupervised Learning of Probabilistic Object Models (POMs) for Object Classification, Segmentation, and Recognition Using Knowledge Propagation
abstract
We present a method to learn probabilistic object models (POMs) with minimal supervision, which exploit different visual cues and perform tasks such as classification, segmentation, and recognition. We formulate this as a structure induction and learning task and our strategy is to learn and combine elementary POMs that make use of complementary image cues. We describe a novel structure induction procedure, which uses knowledge propagation to enable POMs to provide information to other POMs and "teach them" (which greatly reduces the amount of supervision required for training and speeds up the inference). In particular, we learn a POM-IP defined on Interest Points using weak supervision [1], [2] and use this to train a POM-mask, defined on regional features, which yields a combined POM that performs segmentation/localization. This combined model can be used to train POM-edgelets, defined on edgelets, which gives a full POM with improved performance on classification. We give detailed experimental analysis on large data sets for classification and segmentation with comparison to other methods. Inference takes five seconds while learning takes approximately four hours. In addition, we show that the full POM is invariant to scale and rotation of the object (for learning and inference) and can learn hybrid objects classes (i.e., when there are several objects and the identity of the object in each image is unknown). Finally, we show that POMs can be used to match between different objects of the same category, and hence, enable objects recognition.
Yuanhao Chen, Long Zhu, Alan L. Yuille, HongJiang Zhang
IEEE Trans. Pattern Anal. Mach. Intell.1
2009 Unsupervised Learning of Probabilistic Grammar-Markov Models for Object Categories
abstract
We introduce a Probabilistic Grammar-Markov Model (PGMM) which couples probabilistic context free grammars and Markov Random Fields. These PGMMs are generative models defined over attributed features and are used to detect and classify objects in natural images. PGMMs are designed so that they can perform rapid inference, parameter learning, and the more difficult task of structure induction. PGMMs can deal with unknown 2D pose (position, orientation, and scale) in both inference and learning, different appearances, or aspects, of the model. The PGMMs can be learnt in an unsupervised manner where the image can contain one of an unknown number of objects of different categories or even be pure background. We first study the weakly supervised case, where each image contains an example of the (single) object of interest, and then generalize to less supervised cases. The goal of this paper is theoretical but, to provide proof of concept, we demonstrate results from this approach on a subset of the Caltech dataset (learning on a training set and evaluating on a testing set). Our results are generally comparable with the current state of the art, and our inference is performed in less than five seconds.
Long Zhu, Yuanhao Chen, Alan L. Yuille
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 Unsupervised learning of probabilistic object models (POMs) for object classification, segmentation and recognition
abstract
We present a new unsupervised method to learn unified probabilistic object models (POMs) which can be applied to classification, segmentation, and recognition. We formulate this as a structure learning task and our strategy is to learn and combine basic POM's that make use of complementary image cues. Each POM has algorithms for inference and parameter learning, but: (i) the structure of each POM is unknown, and (ii) the inference and parameter learning algorithm for a POM may be impractical without additional information. We address these problems by a novel structure induction procedure which uses knowledge propagation to enable POM's to provide information to other POM's and "teach them" (which greatly reduced the amount of supervision required for training). In particular, we learn a POM-IP defined on interest points using weak supervision [1, 2] and use this to train a POM- mask, defined on regional features, which yields a combined POM which performs segmentation/localization. This combined model can be used to train POM-edgelets, defined on edgelets, which gives a full POM with improved performance on classification. We give detailed experimental analysis on large datasets which show that the full POM is invariant to scale and rotation of the object (for learning and inference) and performs inference rapidly. In addition, we show that we can apply POM's to learn objects classes (i.e. when there are several objects and the identity of the object in each image is unknown). We emphasize that these models can match between different objects from the same category and hence enable object recognition.
Yuanhao Chen, Long Zhu, Alan L. Yuille, HongJiang Zhang
CVPR1
2008 Max Margin AND/OR Graph learning for parsing the human body
abstract
We present a novel structure learning method, Max Margin AND/OR graph (MM-AOG), for parsing the human body into parts and recovering their poses. Our method represents the human body and its parts by an AND/OR graph, which is a multi-level mixture of Markov random fields (MRFs). Max-margin learning, which is a generalization of the training algorithm for support vector machines (SVMs), is used to learn the parameters of the AND/OR graph model discriminatively. There are four advantages from this combination of AND/OR graphs and max-margin learning. Firstly, the AND/OR graph allows us to handle enormous articulated poses with a compact graphical model. Secondly, max-margin learning has more discriminative power than the traditional maximum likelihood approach. Thirdly, the parameters of the AND/OR graph model are optimized globally. In particular, the weights of the appearance model for individual nodes and the relative importance of spatial relationships between nodes are learnt simultaneously. Finally, the kernel trick can be used to handle high dimensional features and to enable complex similarity measure of shapes. We perform comparison experiments on the base ball datasets, showing significant improvements over state of the art methods.
Long Zhu, Yuanhao Chen, Alan L. Yuille
CVPR2
2008 Structure-perceptron learning of a hierarchical log-linear model
abstract
In this paper, we address the problems of deformable object matching (alignment) and segmentation with cluttered background. We propose a novel hierarchical log-linear model (HLLM) which represents both shape and appearance features at multiple levels of a hierarchy. This model enables us to combine appearance cues at multiple scales directly into the hierarchy and to model shape deformations at short-range, medium range, and long-range. We introduce the structure-perceptron algorithm to estimate the parameters of the HLLM in a discriminative way. The learning is able to estimate the appearance and shape parameters simultaneously in a global manner. Moreover, the structure-perceptron learning has a feature selection aspect (similar to AdaBoost) which enables us to specify a class of appearance/shape features and allow the algorithm to select which features to use and weight their importance. This method was applied to the tasks of deformable object localization, segmentation, matching (alignment), and parsing. We demonstrate that the algorithm achieves the state of the art performance by evaluation on public dataset (horse and multi-view face).
Long Zhu, Yuanhao Chen, Xingyao Ye, Alan L. Yuille
CVPR2
2008 Unsupervised Structure Learning: Hierarchical Recursive Composition, Suspicious Coincidence and Competitive Exclusion
Long Zhu, Haoda Huang, Yuanhao Chen, Alan L. Yuille
ECCV (2)4
2008 Recursive Segmentation and Recognition Templates for 2D Parsing
abstract
Language and image understanding are two major goals of artificial intelligence which can both be conceptually formulated in terms of parsing the input signal into a hierarchical representation. Natural language researchers have made great progress by exploiting the 1D structure of language to design efficient polynomial- time parsing algorithms. By contrast, the two-dimensional nature of images makes it much harder to design efficient image parsers and the form of the hierarchical representations is also unclear. Attempts to adapt representations and algorithms from natural language have only been partially successful. In this paper, we propose a Hierarchical Image Model (HIM) for 2D image pars- ing which outputs image segmentation and object recognition. This HIM is rep- resented by recursive segmentation and recognition templates in multiple layers and has advantages for representation, inference, and learning. Firstly, the HIM has a coarse-to-fine representation which is capable of capturing long-range de- pendency and exploiting different levels of contextual information. Secondly, the structure of the HIM allows us to design a rapid inference algorithm, based on dy- namic programming, which enables us to parse the image rapidly in polynomial time. Thirdly, we can learn the HIM efficiently in a discriminative manner from a labeled dataset. We demonstrate that HIM outperforms other state-of-the-art methods by evaluation on the challenging public MSRC image dataset. Finally, we sketch how the HIM architecture can be extended to model more complex image phenomena.
Long Zhu, Yuanhao Chen, Alan L. Yuille
NIPS2
2007 Rapid Inference on a Novel AND/OR graph for Object Detection, Segmentation and Parsing
abstract
In this paper we formulate a novel AND/OR graph representation capable of describing the different configurations of deformable articulated objects such as horses. The representation makes use of the summarization principle so that lower level nodes in the graph only pass on summary statistics to the higher level nodes. The probability distributions are invariant to position, orientation, and scale. We develop a novel inference algorithm that combined a bottom-up process for proposing configurations for horses together with a top-down process for refining and validating these proposals. The strategy of surround suppression is applied to ensure that the inference time is polynomial in the size of input data. The algorithm was applied to the tasks of detecting, segmenting and parsing horses. We demonstrate that the algorithm is fast and comparable with the state of the art approaches.
Yuanhao Chen, Long Zhu, Alan L. Yuille, HongJiang Zhang
NIPS1
2006 Automatic Classification of Photographs and Graphics
abstract
In general, digital images can be classified into photographs and computer graphics. This taxonomy is very useful in many applications, such as Web image search. However, there are no effective methods to perform this classification automatically. In this paper, we manage to solve this problem from two aspects. At first, we propose some novel low-level features that can reveal perceptional differences between photographs and graphics. Then, we adopt an effective algorithm to perform the classification. The experiments conducted on a large-scale image database indicate the effectiveness of our algorithm
Yuanhao Chen, Zhiwei Li 0006, Mingjing Li, Wei-Ying Ma
ICME1
2006 Unsupervised Learning of a Probabilistic Grammar for Object Detection and Parsing
abstract
We describe an unsupervised method for learning a probabilistic grammar of an object from a set of training examples. Our approach is invariant to the scale and rotation of the objects. We illustrate our approach using thirteen objects from the Caltech 101 database. In addition, we learn the model of a hybrid object class where we do not know the specific object or its position, scale or pose. This is illustrated by learning a hybrid class consisting of faces, motorbikes, and airplanes. The individual objects can be recovered as different aspects of the grammar for the object class. In all cases, we validate our results by learning the probability grammars from training datasets and evaluating them on the test datasets. We compare our method to alternative approaches. The advantages of our approach is the speed of inference (under one second), the parsing of the object, and increased accuracy of performance. Moreover, our approach is very general and can be applied to a large range of objects and structures.
Long Zhu, Yuanhao Chen, Alan L. Yuille
NIPS2