EDBT 2026 Demo / reviewers in the wild / expert
Long Zhu
dblp:z/LongZhu · also Leo Zhu
· DBLP profile ↗
24ranked-venue papers
14as first author
5since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 13 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
19 papers |
Image recognition and object detection · 36% Segmentation and scene understanding · 31% Probabilistic and Bayesian machine learning · 12% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 100% |
Topics — the 30 heaviest of 36, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection
object detection |
0.5 | 6 | 2010 | Active Mask Hierarchies for Object Detection · ECCV (5) 2010 Latent hierarchical structural learning for object detection · CVPR 2010 Unsupervised Learning of Probabilistic Grammar-Markov Models for Object Categories · IEEE Trans. Pattern Anal. Mach. Intell. 2009 |
Computer vision › Image recognition and object detection
object recognition |
0.4 | 4 | 2012 | Recursive Segmentation and Recognition Templates for Image Parsing · IEEE Trans. Pattern Anal. Mach. Intell. 2012 Unsupervised Learning of Probabilistic Object Models (POMs) for Object Classification, Segmentation, and Recognition Using Knowledge Propagation · IEEE Trans. Pattern Anal. Mach. Intell. 2009 Recursive Segmentation and Recognition Templates for 2D Parsing · NIPS 2008 |
Computer vision › Segmentation and scene understanding › object segmentation
object parsing |
0.4 | 4 | 2011 | Max Margin Learning of Hierarchical Configural Deformable Templates (HCDTs) for Efficient Object Parsing and Pose Estimation · Int. J. Comput. Vis. 2011 Part and appearance sharing: Recursive Compositional Models for multi-view · CVPR 2010 Rapid Inference on a Novel AND/OR graph for Object Detection, Segmentation and Parsing · NIPS 2007 |
Computer vision › Segmentation and scene understanding
object segmentation |
0.3 | 3 | 2011 | Max Margin Learning of Hierarchical Configural Deformable Templates (HCDTs) for Efficient Object Parsing and Pose Estimation · Int. J. Comput. Vis. 2011 Unsupervised Learning of Probabilistic Object Models (POMs) for Object Classification, Segmentation, and Recognition Using Knowledge Propagation · IEEE Trans. Pattern Anal. Mach. Intell. 2009 Unsupervised learning of probabilistic object models (POMs) for object classification, segmentation and recognition · CVPR 2008 |
Computer vision › Image recognition and object detection › object detection
deformable object detection |
0.2 | 3 | 2010 | Learning a Hierarchical Deformable Template for Rapid Deformable Object Parsing · IEEE Trans. Pattern Anal. Mach. Intell. 2010 Rapid Inference on a Novel AND/OR graph for Object Detection, Segmentation and Parsing · NIPS 2007 A Hierarchical Compositional System for Rapid Object Detection · NIPS 2005 |
Computer vision › Segmentation and scene understanding
scene parsing |
0.2 | 2 | 2012 | Recursive Segmentation and Recognition Templates for Image Parsing · IEEE Trans. Pattern Anal. Mach. Intell. 2012 Recursive Segmentation and Recognition Templates for 2D Parsing · NIPS 2008 |
Computer vision › Face, body and person analysis
human pose estimation |
0.2 | 2 | 2011 | Max Margin Learning of Hierarchical Configural Deformable Templates (HCDTs) for Efficient Object Parsing and Pose Estimation · Int. J. Comput. Vis. 2011 Max Margin AND/OR Graph learning for parsing the human body · CVPR 2008 |
Computer vision › Image recognition and object detection › image classification
object classification |
0.2 | 3 | 2009 | Unsupervised Learning of Probabilistic Object Models (POMs) for Object Classification, Segmentation, and Recognition Using Knowledge Propagation · IEEE Trans. Pattern Anal. Mach. Intell. 2009 Unsupervised learning of probabilistic object models (POMs) for object classification, segmentation and recognition · CVPR 2008 Unsupervised Learning of Probabilistic Grammar-Markov Models for Object Categories · IEEE Trans. Pattern Anal. Mach. Intell. 2009 |
Computer vision › Segmentation and scene understanding
instance segmentation |
0.1 | 1 | 2010 | Active Mask Hierarchies for Object Detection · ECCV (5) 2010 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
latent structure discovery |
0.1 | 1 | 2010 | Latent hierarchical structural learning for object detection · CVPR 2010 |
Computer vision › Image recognition and object detection › object detection
multi-view object detection |
0.1 | 1 | 2010 | Part and appearance sharing: Recursive Compositional Models for multi-view · CVPR 2010 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian nonparametric model |
0.1 | 1 | 2009 | Nonparametric Bayesian Texture Learning and Synthesis · NIPS 2009 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian nonparametric model
hierarchical dirichlet process |
0.1 | 1 | 2009 | Nonparametric Bayesian Texture Learning and Synthesis · NIPS 2009 |
Computer vision › Segmentation and scene understanding
image segmentation |
0.1 | 1 | 2009 | Nonparametric Bayesian Texture Learning and Synthesis · NIPS 2009 |
Visual content generation and editing
texture synthesis |
0.1 | 1 | 2009 | Nonparametric Bayesian Texture Learning and Synthesis · NIPS 2009 |
Computer vision › Segmentation and scene understanding › image segmentation › model-based segmentation
deformable model segmentation |
0.1 | 1 | 2008 | Structure-perceptron learning of a hierarchical log-linear model · CVPR 2008 |
Computer vision › 3D vision › shape matching
deformable object matching |
0.1 | 1 | 2008 | Structure-perceptron learning of a hierarchical log-linear model · CVPR 2008 |
Machine learning › Probabilistic and Bayesian machine learning › hierarchical modeling › hierarchical model
hierarchical image model |
0.1 | 1 | 2008 | Recursive Segmentation and Recognition Templates for 2D Parsing · NIPS 2008 |
Computer vision › Image recognition and object detection › part-based model
hierarchical shape model |
0.1 | 1 | 2008 | Structure-perceptron learning of a hierarchical log-linear model · CVPR 2008 |
Computer vision › Segmentation and scene understanding
human parsing |
0.1 | 1 | 2008 | Max Margin AND/OR Graph learning for parsing the human body · CVPR 2008 |
Machine learning › Generative modeling › generative model
probabilistic object model |
0.1 | 1 | 2008 | Unsupervised learning of probabilistic object models (POMs) for object classification, segmentation and recognition · CVPR 2008 |
Machine learning › Representation and self-supervised learning
structure induction |
0.1 | 1 | 2008 | Unsupervised learning of probabilistic object models (POMs) for object classification, segmentation and recognition · CVPR 2008 |
Machine learning › Learning paradigms › unsupervised learning
unsupervised structure learning |
0.1 | 1 | 2008 | Unsupervised Structure Learning: Hierarchical Recursive Composition, Suspicious Coincidence and Competitive Exclusion · ECCV (2) 2008 |
Machine learning › Generative modeling › variational autoencoder › hierarchical latent variable model
hierarchical compositional model |
0.1 | 2 | 2008 | A Hierarchical Compositional System for Rapid Object Detection · NIPS 2005 Unsupervised Structure Learning: Hierarchical Recursive Composition, Suspicious Coincidence and Competitive Exclusion · ECCV (2) 2008 |
Computer vision › Face, body and person analysis
face detection |
0.1 | 2 | 2003 | Boosting Chain Learning for Object Detection · ICCV 2003 Statistical Learning of Multi-view Face Detection · ECCV (4) 2002 |
Natural language and speech › Language models and text generation › grammar formalisms
probabilistic grammar |
0.1 | 1 | 2006 | Unsupervised Learning of a Probabilistic Grammar for Object Detection and Parsing · NIPS 2006 |
Machine learning › Probabilistic and Bayesian machine learning › hierarchical modeling
hierarchical model |
0.0 | 1 | 2012 | Recursive Segmentation and Recognition Templates for Image Parsing · IEEE Trans. Pattern Anal. Mach. Intell. 2012 |
Computer vision › Image recognition and object detection › object detection
boosting-based detection |
0.0 | 1 | 2003 | Boosting Chain Learning for Object Detection · ICCV 2003 |
Machine learning › Kernel, tree and ensemble methods › classifier combination
boosting cascade |
0.0 | 1 | 2003 | Boosting Chain Learning for Object Detection · ICCV 2003 |
Computer vision › 3D vision › 3d shape modeling
articulated object modeling |
0.0 | 1 | 2011 | Max Margin Learning of Hierarchical Configural Deformable Templates (HCDTs) for Efficient Object Parsing and Pose Estimation · Int. J. Comput. Vis. 2011 |
Methods — techniques the papers use, named apart from their topics
dynamic programming · 0.3and-or graph · 0.3max-margin learning · 0.2structure-perceptron · 0.2structure induction · 0.2markov random field · 0.2knowledge propagation · 0.2recursive segmentation · 0.1deformable template · 0.1appearance sharing · 0.12d hidden markov model · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Reinforcement Learning for Option Hedging Using Quantile Regression and Curriculum Learning with Historical Data FusionabstractIn the financial field, options hedging has gained significant attention due to its crucial role in risk management and corporate operations. Traditional option hedging methods, such as delta hedging based on the Black-Scholes model, face limitations in practical application due to assumptions of constant volatility and the neglect of transaction costs. In recent years, reinforcement learning (RL) has become a hot topic in option hedging research because it can interact with the environment and adjust hedging strategies based on environmental feedback, better reflecting real market conditions. However, it still faces several challenges, such as the neglect of historical market information, insufficient evaluation of hedging cost distributions, and high randomness in model training. This paper proposes a reinforcement learning method for option hedging using quantile regression and curriculum learning with historical data fusion, aiming to address the current shortcomings in historical information utilization, risk assessment of the hedging strategy, and robustness of the model. Firstly, historical market information is incorporated into state variables, and the time-trend multi-head self-attention mechanism (TiTrMHSA) and Gated Recurrent Units (GRU) are introduced to capture the dynamic changing trends of historical information and integrate current features, significantly improving the hedging strategy’s sensitivity to market fluctuations. Secondly, quantile regression is used in the value network, combining the Quantile Huber Loss function to fit the hedging cost distribution. This enables a comprehensive evaluation of strategy performance and enhances risk control capabilities. Finally, by incorporating the concept of curriculum learning, a two-stage training method for the agent is designed to progressively optimize the agent’s learning process from simple to complex, addressing the issue of high randomness in early-stage training. The Experimental results show that this method significantly outperforms traditional BS-Delta hedging and other reinforcement learning models regarding hedging costs and the balance between returns and risks, demonstrating superior robustness and adaptability. Qiao Pan, Long Zhu, Zhaoju Wang |
IJCNN | 2 |
| 2024 | Robust Design Optimization of PMSLM Based on I-MEABP Neural NetworkabstractUncertain factors such as manufacturing error, assembly error and operational wear will cause the dimension deviation of the optimal structural parameters for permanent magnet synchronous linear motor (PMSLM), which will result in the deterioration of the thrust performance for PMSLM. A robust optimization method based on interval-mind evolutionary algorithm back propagation (I-MEABP) neural network is proposed in this paper. First, the analysis model of thrust performance for PMSLM is established, and the key structural parameters that have greater impacts on the thrust performance of PMSLM are obtained. Second, the I-MEABP neural network is formed to establish the robust numerical model of thrust performance for PMSLM based on the key structural parameters. It integrates the interval analysis theory into the Mind Evolutionary Algorithm Back Propagation (MEABP) neural network and a robust modeling method I-MEABP is formed. It can take the interval structural parameters that include structural dimension deviations as input variables and takes the interval thrust performance values as output values based on the interval sample library. Third, the adaptive whale optimization algorithm (AWOA) is applied to optimize the robust model, and the robust optimal interval values of structural parameters are obtained. Compared with the deterministic optimization method, the robustness of the results in this paper is higher; compared with other robust methods, this robust method obtains the optimal interval of structural parameters for PMSLM, which is different from the optimal point. Finally, Finite element analysis and motor prototype experiments prove the robustness and superiority of the proposed method. Weiguo Zhou, Jing Zhao 0023, Wangliang Qian, Long Zhu, Chen Chen 0106 |
IEEE Trans. Ind. Informatics | 5 |
| 2023 | A Single-sample Pruning and Clustering Method for Neural Network VerificationabstractThe verification techniques based on formal methods can provide deterministic guarantees for the robustness of Deep Neural Networks(DNNS). However, the enormous scale of DNNS makes the application of such methods in this field a huge challenge. To address this problem, this study proposes a single-sample sub-network pruning method, which can identify redundant nodes by combining neuron coverage and the symbolic interval propagation method to reduce the network verification scale. In addition, to solve the problem of too many sub-networks to be pruned, according to the similarity of neuron coverage between samples, we propose a corresponding clustering algorithm to establish sub-networks for different categories of samples to improve the verification efficiency. We combine the MIPverify verification tool to validate the above method. Experiments show that the sub-networks can give the same robust validation results and similar robustness bounds as the original network, while greatly reducing the validation time and network size. Huanzhang Xiong, Gang Hou, Long Zhu, Jie Wang 0004, Weiqiang Kong |
APSEC | 3 |
| 2023 | Accumulated Error Reduction of Linear Motor Mover Position Measurement Based on SKLMabstractCumulative measurement error is the most critical factor affecting the accuracy of long-stroke displacement measurement. Based on the 1-D image gradient method and smooth Kalman filter algorithm, this article proposes a long-stroke linear motor displacement cumulative error reduction and high-precision measurement methods. First, according to the motion characteristics of the linear motor and the principle of image measurement, a 1-D speckle target image is generated, and the 1-D Barron gradient algorithm is used to calculate the displacement of adjacent frames quickly. Second, according to the reasons for the accumulated error of long-stroke displacement measurement, the smooth Kalman algorithm for optimized autoregressive data processing is introduced to optimally estimate the measured displacement of adjacent frames to realize the reduction of accumulated error. To improve the robustness of the measurement system, a wavelet soft-threshold image filter is introduced to perform noise reduction and restoration processing on the collected signal and further realize the high-precision displacement measurement of the long-stroke linear motor. Simulations and experiments show that the method presented in this article not only improves the measurement accuracy of adjacent frames but also reduces the cumulative error of long-stroke displacement measurement. And under different working conditions, compared with other methods, it has higher accuracy and anti-interference. Jing Zhao 0023, Long Zhu, Chen Chen 0106 |
IEEE Trans. Ind. Informatics | 2 |
| 2021 | Integration of a physiologically-based pharmacokinetic model with a whole-body, organ-resolved genome-scale model for characterization of ethanol and acetaldehyde metabolismabstractEthanol is one of the most widely used recreational substances in the world and due to its ubiquitous use, ethanol abuse has been the cause of over 3.3 million deaths each year. In addition to its effects, ethanol's primary metabolite, acetaldehyde, is a carcinogen that can cause symptoms of facial flushing, headaches, and nausea. How strongly ethanol or acetaldehyde affects an individual depends highly on the genetic polymorphisms of certain genes. In particular, the genetic polymorphisms of mitochondrial aldehyde dehydrogenase, ALDH2, play a large role in the metabolism of acetaldehyde. Thus, it is important to characterize how genetic variations can lead to different exposures and responses to ethanol and acetaldehyde. While the pharmacokinetics of ethanol metabolism through alcohol dehydrogenase have been thoroughly explored in previous studies, in this paper, we combined a base physiologically-based pharmacokinetic (PBPK) model with a whole-body genome-scale model (WBM) to gain further insight into the effect of other less explored processes and genetic variations on ethanol metabolism. This combined model was fit to clinical data and used to show the effect of alcohol concentrations, organ damage, ALDH2 enzyme polymorphisms, and ALDH2-inhibiting drug disulfiram on ethanol and acetaldehyde exposure. Through estimating the reaction rates of auxiliary processes with dynamic Flux Balance Analysis, The PBPK-WBM was able to navigate around a lack of kinetic constants traditionally associated with PK modelling and demonstrate the compensatory effects of the body in response to decreased liver enzyme expression. Additionally, the model demonstrated that acetaldehyde exposure increased with higher dosages of disulfiram and decreased ALDH2 efficiency, and that moderate consumption rates of ethanol could lead to unexpected accumulations in acetaldehyde. This modelling framework combines the comprehensive steady-state analyses from genome-scale models with the dynamics of traditional PK models to create a highly personalized form of PBPK modelling that can push the boundaries of precision medicine. Long Zhu, William Pei, Ines Thiele, Radhakrishnan Mahadevan |
PLoS Comput. Biol. | 1 |
| 2012 | Recursive Segmentation and Recognition Templates for Image ParsingabstractIn this paper, we propose a Hierarchical Image Model (HIM) which parses images to perform segmentation and object recognition. The HIM represents the image recursively by segmentation and recognition templates at multiple levels of the hierarchy. This has advantages for representation, inference, and learning. First, the HIM has a coarse-to-fine representation which is capable of capturing long-range dependency and exploiting different levels of contextual information (similar to how natural language models represent sentence structure in terms of hierarchical representations such as verb and noun phrases). Second, the structure of the HIM allows us to design a rapid inference algorithm, based on dynamic programming, which yields the first polynomial time algorithm for image labeling. Third, we learn the HIM efficiently using machine learning methods from a labeled data set. We demonstrate that the HIM is comparable with the state-of-the-art methods by evaluation on the challenging public MSRC and PASCAL VOC 2007 image data sets. Long Zhu, Yuanhao Chen, Alan L. Yuille |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | Max Margin Learning of Hierarchical Configural Deformable Templates (HCDTs) for Efficient Object Parsing and Pose EstimationabstractIn this paper we formulate a hierarchical configurable deformable template (HCDT) to model articulated visual objects—such as horses and baseball players—for tasks such as parsing, segmentation, and pose estimation. HCDTs represent an object by an AND/OR graph where the OR nodes act as switches which enables the graph topology to vary adaptively. This hierarchical representation is compositional and the node variables represent positions and properties of subparts of the object. The graph and the node variables are required to obey the summarization principle which enables an efficient compositional inference algorithm to rapidly estimate the state of the HCDT. We specify the structure of the AND/OR graph of the HCDT by hand and learn the model parameters discriminatively by extending Max-Margin learning to AND/OR graphs. We illustrate the three main aspects of HCDTs—representation, inference, and learning—on the tasks of segmenting, parsing, and pose (configuration) estimation for horses and humans. We demonstrate that the inference algorithm is fast and that max-margin learning is effective. We show that HCDTs gives state of the art results for segmentation and pose estimation when compared to other methods on benchmarked datasets. Long Zhu, Yuanhao Chen, Alan L. Yuille |
Int. J. Comput. Vis. | 1 |
| 2010 | Part and appearance sharing: Recursive Compositional Models for multi-viewabstractWe propose Recursive Compositional Models (RCMs) for simultaneous multi-view multi-object detection and parsing (e.g. view estimation and determining the positions of the object subparts). We represent the set of objects by a family of RCMs where each RCM is a probability distribution defined over a hierarchical graph which corresponds to a specific object and viewpoint. An RCM is constructed from a hierarchy of subparts/subgraphs which are learnt from training data. Part-sharing is used so that different RCMs are encouraged to share subparts/subgraphs which yields a compact representation for the set of objects and which enables efficient inference and learning from a limited number of training samples. In addition, we use appearance-sharing so that RCMs for the same object, but different viewpoints, share similar appearance cues which also helps efficient learning. RCMs lead to a multi-view multi-object detection system. We illustrate RCMs on four public datasets and achieve state-of-the-art performance. Long Zhu, Yuanhao Chen, Antonio Torralba 0001, William T. Freeman, Alan L. Yuille |
CVPR | 1 |
| 2010 | Latent hierarchical structural learning for object detectionabstractWe present a latent hierarchical structural learning method for object detection. An object is represented by a mixture of hierarchical tree models where the nodes represent object parts. The nodes can move spatially to allow both local and global shape deformations. The models can be trained discriminatively using latent structural SVM learning, where the latent variables are the node positions and the mixture component. But current learning methods are slow, due to the large number of parameters and latent variables, and have been restricted to hierarchies with two layers. In this paper we describe an incremental concave-convex procedure (iCCCP) which allows us to learn both two and three layer models efficiently. We show that iCCCP leads to a simple training algorithm which avoids complex multi-stage layer-wise training, careful part selection, and achieves good performance without requiring elaborate initialization. We perform object detection using our learnt models and obtain performance comparable with state-of-the-art methods when evaluated on challenging public PASCAL datasets. We demonstrate the advantages of three layer hierarchies - outperforming Felzenszwalb et al.'s two layer models on all 20 classes. Long Zhu, Yuanhao Chen, Alan L. Yuille, William T. Freeman |
CVPR | 1 |
| 2010 | Active Mask Hierarchies for Object Detection
Yuanhao Chen, Long Zhu, Alan L. Yuille |
ECCV (5) | 2 |
| 2010 | Learning a Hierarchical Deformable Template for Rapid Deformable Object ParsingabstractIn this paper, we address the tasks of detecting, segmenting, parsing, and matching deformable objects. We use a novel probabilistic object model that we call a hierarchical deformable template (HDT). The HDT represents the object by state variables defined over a hierarchy (with typically five levels). The hierarchy is built recursively by composing elementary structures to form more complex structures. A probability distribution--a parameterized exponential model--is defined over the hierarchy to quantify the variability in shape and appearance of the object at multiple scales. To perform inference--to estimate the most probable states of the hierarchy for an input image--we use a bottom-up algorithm called compositional inference. This algorithm is an approximate version of dynamic programming where approximations are made (e.g., pruning) to ensure that the algorithm is fast while maintaining high performance. We adapt the structure-perceptron algorithm to estimate the parameters of the HDT in a discriminative manner (simultaneously estimating the appearance and shape parameters). More precisely, we specify an exponential distribution for the HDT using a dictionary of potentials, which capture the appearance and shape cues. This dictionary can be large and so does not require handcrafting the potentials. Instead, structure-perceptron assigns weights to the potentials so that less important potentials receive small weights (this is like a "soft" form of feature selection). Finally, we provide experimental evaluation of HDTs on different visual tasks, including detection, segmentation, matching (alignment), and parsing. We show that HDTs achieve state-of-the-art performance for these different tasks when evaluated on data sets with groundtruth (and when compared to alternative algorithms, which are typically specialized to each task). Long Zhu, Yuanhao Chen, Alan L. Yuille |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2009 | Nonparametric Bayesian Texture Learning and SynthesisabstractWe present a nonparametric Bayesian method for texture learning and synthesis. A texture image is represented by a 2D-Hidden Markov Model (2D-HMM) where the hidden states correspond to the cluster labeling of textons and the transition matrix encodes their spatial layout (the compatibility between adjacent textons). 2D-HMM is coupled with the Hierarchical Dirichlet process (HDP) which allows the number of textons and the complexity of transition matrix grow as the input texture becomes irregular. The HDP makes use of Dirichlet process prior which favors regular textures by penalizing the model complexity. This framework (HDP-2D-HMM) learns the texton vocabulary and their spatial layout jointly and automatically. The HDP-2D-HMM results in a compact representation of textures which allows fast texture synthesis with comparable rendering quality over the state-of-the-art image-based rendering methods. We also show that HDP-2D-HMM can be applied to perform image segmentation and synthesis. Long Zhu, Yuanhao Chen, William T. Freeman, Antonio Torralba 0001 |
NIPS | 1 |
| 2009 | Unsupervised Learning of Probabilistic Object Models (POMs) for Object Classification, Segmentation, and Recognition Using Knowledge PropagationabstractWe present a method to learn probabilistic object models (POMs) with minimal supervision, which exploit different visual cues and perform tasks such as classification, segmentation, and recognition. We formulate this as a structure induction and learning task and our strategy is to learn and combine elementary POMs that make use of complementary image cues. We describe a novel structure induction procedure, which uses knowledge propagation to enable POMs to provide information to other POMs and "teach them" (which greatly reduces the amount of supervision required for training and speeds up the inference). In particular, we learn a POM-IP defined on Interest Points using weak supervision [1], [2] and use this to train a POM-mask, defined on regional features, which yields a combined POM that performs segmentation/localization. This combined model can be used to train POM-edgelets, defined on edgelets, which gives a full POM with improved performance on classification. We give detailed experimental analysis on large data sets for classification and segmentation with comparison to other methods. Inference takes five seconds while learning takes approximately four hours. In addition, we show that the full POM is invariant to scale and rotation of the object (for learning and inference) and can learn hybrid objects classes (i.e., when there are several objects and the identity of the object in each image is unknown). Finally, we show that POMs can be used to match between different objects of the same category, and hence, enable objects recognition. Yuanhao Chen, Long Zhu, Alan L. Yuille, HongJiang Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2009 | Unsupervised Learning of Probabilistic Grammar-Markov Models for Object CategoriesabstractWe introduce a Probabilistic Grammar-Markov Model (PGMM) which couples probabilistic context free grammars and Markov Random Fields. These PGMMs are generative models defined over attributed features and are used to detect and classify objects in natural images. PGMMs are designed so that they can perform rapid inference, parameter learning, and the more difficult task of structure induction. PGMMs can deal with unknown 2D pose (position, orientation, and scale) in both inference and learning, different appearances, or aspects, of the model. The PGMMs can be learnt in an unsupervised manner where the image can contain one of an unknown number of objects of different categories or even be pure background. We first study the weakly supervised case, where each image contains an example of the (single) object of interest, and then generalize to less supervised cases. The goal of this paper is theoretical but, to provide proof of concept, we demonstrate results from this approach on a subset of the Caltech dataset (learning on a training set and evaluating on a testing set). Our results are generally comparable with the current state of the art, and our inference is performed in less than five seconds. Long Zhu, Yuanhao Chen, Alan L. Yuille |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2008 | Unsupervised learning of probabilistic object models (POMs) for object classification, segmentation and recognitionabstractWe present a new unsupervised method to learn unified probabilistic object models (POMs) which can be applied to classification, segmentation, and recognition. We formulate this as a structure learning task and our strategy is to learn and combine basic POM's that make use of complementary image cues. Each POM has algorithms for inference and parameter learning, but: (i) the structure of each POM is unknown, and (ii) the inference and parameter learning algorithm for a POM may be impractical without additional information. We address these problems by a novel structure induction procedure which uses knowledge propagation to enable POM's to provide information to other POM's and "teach them" (which greatly reduced the amount of supervision required for training). In particular, we learn a POM-IP defined on interest points using weak supervision [1, 2] and use this to train a POM- mask, defined on regional features, which yields a combined POM which performs segmentation/localization. This combined model can be used to train POM-edgelets, defined on edgelets, which gives a full POM with improved performance on classification. We give detailed experimental analysis on large datasets which show that the full POM is invariant to scale and rotation of the object (for learning and inference) and performs inference rapidly. In addition, we show that we can apply POM's to learn objects classes (i.e. when there are several objects and the identity of the object in each image is unknown). We emphasize that these models can match between different objects from the same category and hence enable object recognition. Yuanhao Chen, Long Zhu, Alan L. Yuille, HongJiang Zhang |
CVPR | 2 |
| 2008 | Max Margin AND/OR Graph learning for parsing the human bodyabstractWe present a novel structure learning method, Max Margin AND/OR graph (MM-AOG), for parsing the human body into parts and recovering their poses. Our method represents the human body and its parts by an AND/OR graph, which is a multi-level mixture of Markov random fields (MRFs). Max-margin learning, which is a generalization of the training algorithm for support vector machines (SVMs), is used to learn the parameters of the AND/OR graph model discriminatively. There are four advantages from this combination of AND/OR graphs and max-margin learning. Firstly, the AND/OR graph allows us to handle enormous articulated poses with a compact graphical model. Secondly, max-margin learning has more discriminative power than the traditional maximum likelihood approach. Thirdly, the parameters of the AND/OR graph model are optimized globally. In particular, the weights of the appearance model for individual nodes and the relative importance of spatial relationships between nodes are learnt simultaneously. Finally, the kernel trick can be used to handle high dimensional features and to enable complex similarity measure of shapes. We perform comparison experiments on the base ball datasets, showing significant improvements over state of the art methods. Long Zhu, Yuanhao Chen, Alan L. Yuille |
CVPR | 1 |
| 2008 | Structure-perceptron learning of a hierarchical log-linear modelabstractIn this paper, we address the problems of deformable object matching (alignment) and segmentation with cluttered background. We propose a novel hierarchical log-linear model (HLLM) which represents both shape and appearance features at multiple levels of a hierarchy. This model enables us to combine appearance cues at multiple scales directly into the hierarchy and to model shape deformations at short-range, medium range, and long-range. We introduce the structure-perceptron algorithm to estimate the parameters of the HLLM in a discriminative way. The learning is able to estimate the appearance and shape parameters simultaneously in a global manner. Moreover, the structure-perceptron learning has a feature selection aspect (similar to AdaBoost) which enables us to specify a class of appearance/shape features and allow the algorithm to select which features to use and weight their importance. This method was applied to the tasks of deformable object localization, segmentation, matching (alignment), and parsing. We demonstrate that the algorithm achieves the state of the art performance by evaluation on public dataset (horse and multi-view face). Long Zhu, Yuanhao Chen, Xingyao Ye, Alan L. Yuille |
CVPR | 1 |
| 2008 | Unsupervised Structure Learning: Hierarchical Recursive Composition, Suspicious Coincidence and Competitive Exclusion
Long Zhu, Haoda Huang, Yuanhao Chen, Alan L. Yuille |
ECCV (2) | 1 |
| 2008 | Recursive Segmentation and Recognition Templates for 2D ParsingabstractLanguage and image understanding are two major goals of artificial intelligence which can both be conceptually formulated in terms of parsing the input signal into a hierarchical representation. Natural language researchers have made great progress by exploiting the 1D structure of language to design efficient polynomial- time parsing algorithms. By contrast, the two-dimensional nature of images makes it much harder to design efficient image parsers and the form of the hierarchical representations is also unclear. Attempts to adapt representations and algorithms from natural language have only been partially successful. In this paper, we propose a Hierarchical Image Model (HIM) for 2D image pars- ing which outputs image segmentation and object recognition. This HIM is rep- resented by recursive segmentation and recognition templates in multiple layers and has advantages for representation, inference, and learning. Firstly, the HIM has a coarse-to-fine representation which is capable of capturing long-range de- pendency and exploiting different levels of contextual information. Secondly, the structure of the HIM allows us to design a rapid inference algorithm, based on dy- namic programming, which enables us to parse the image rapidly in polynomial time. Thirdly, we can learn the HIM efficiently in a discriminative manner from a labeled dataset. We demonstrate that HIM outperforms other state-of-the-art methods by evaluation on the challenging public MSRC image dataset. Finally, we sketch how the HIM architecture can be extended to model more complex image phenomena. Long Zhu, Yuanhao Chen, Alan L. Yuille |
NIPS | 1 |
| 2007 | Rapid Inference on a Novel AND/OR graph for Object Detection, Segmentation and ParsingabstractIn this paper we formulate a novel AND/OR graph representation capable of describing the different configurations of deformable articulated objects such as horses. The representation makes use of the summarization principle so that lower level nodes in the graph only pass on summary statistics to the higher level nodes. The probability distributions are invariant to position, orientation, and scale. We develop a novel inference algorithm that combined a bottom-up process for proposing configurations for horses together with a top-down process for refining and validating these proposals. The strategy of surround suppression is applied to ensure that the inference time is polynomial in the size of input data. The algorithm was applied to the tasks of detecting, segmenting and parsing horses. We demonstrate that the algorithm is fast and comparable with the state of the art approaches. Yuanhao Chen, Long Zhu, Alan L. Yuille, HongJiang Zhang |
NIPS | 2 |
| 2006 | Unsupervised Learning of a Probabilistic Grammar for Object Detection and ParsingabstractWe describe an unsupervised method for learning a probabilistic grammar of an object from a set of training examples. Our approach is invariant to the scale and rotation of the objects. We illustrate our approach using thirteen objects from the Caltech 101 database. In addition, we learn the model of a hybrid object class where we do not know the specific object or its position, scale or pose. This is illustrated by learning a hybrid class consisting of faces, motorbikes, and airplanes. The individual objects can be recovered as different aspects of the grammar for the object class. In all cases, we validate our results by learning the probability grammars from training datasets and evaluating them on the test datasets. We compare our method to alternative approaches. The advantages of our approach is the speed of inference (under one second), the parsing of the object, and increased accuracy of performance. Moreover, our approach is very general and can be applied to a large range of objects and structures. Long Zhu, Yuanhao Chen, Alan L. Yuille |
NIPS | 1 |
| 2005 | A Hierarchical Compositional System for Rapid Object DetectionabstractWe describe a hierarchical compositional system for detecting deformable objects in images. Objects are represented by graphical models. The algorithm uses a hierarchical tree where the root of the tree corresponds to the full object and lower-level elements of the tree correspond to simpler features. The algorithm proceeds by passing simple messages up and down the tree. The method works rapidly, in under a second, on 320 240 images. We demonstrate the approach on detecting cats, horses, and hands. The method works in the presence of background clutter and occlusions. Our approach is contrasted with more traditional methods such as dynamic programming and belief propagation. Long Zhu, Alan L. Yuille |
NIPS | 1 |
| 2003 | Boosting Chain Learning for Object DetectionabstractA general classification framework, called boosting chain, is proposed for learning boosting cascade. In this framework, a "chain" structure is introduced to integrate historical knowledge into successive boosting learning. Moreover, a linear optimization scheme is proposed to address the problems of redundancy in boosting learning and threshold adjusting in cascade coupling. By this means, the resulting classifier consists of fewer weak classifiers yet achieves lower error rates than boosting cascade in both training and test. Experimental comparisons of boosting chain and boosting cascade are provided through a face detection problem. The promising results clearly demonstrate the effectiveness made by boosting chain. Long Zhu, HongJiang Zhang |
ICCV | 2 |
| 2002 | Statistical Learning of Multi-view Face Detection
Stan Z. Li, Long Zhu, ZhenQiu Zhang, Andrew Blake 0001, HongJiang Zhang, Harry Shum |
ECCV (4) | 2 |