Chaohui Wang

dblp:51/711 · DBLP profile ↗
← Back
32ranked-venue papers
9as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 5 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 4 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
19 papers
3D vision · 31% Segmentation and scene understanding · 20% Generative modeling · 11%
Computer graphics and multimedia
4 papers
Image and video processing · 43% Rendering · 37% Visual content generation and editing · 14%
Databases, data mining, and information retrieval
2 papers
Graph data management · 59% Query processing and optimization · 22% Indexing and storage engines · 19%

Topics — the 30 heaviest of 44, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding › boundary detection
occlusion boundary detection
0.822022
Occlusion Boundary: A Formal Definition & Its Detection via Deep Exploration of Context · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Occlusion Boundary Detection via Deep Exploration of Context · CVPR 2016
Graph data management › graph query
subgraph query
0.822020
FERRARI: an efficient framework for visual exploratory subgraph search in graph databases · VLDB J. 2020
An Indexing Framework for Efficient Visual Exploratory Subgraph Search in Graph Databases · ICDE 2019
Machine learning › Generative modeling
generative adversarial network
0.832019
Geometry-Consistent Generative Adversarial Networks for One-Sided Unsupervised Domain Mapping · CVPR 2019
Perceptual Adversarial Networks for Image-to-Image Transformation · IEEE Trans. Image Process. 2018
Tag Disentangled Generative Adversarial Network for Object Image Re-rendering · IJCAI 2017
Computer vision › Segmentation and scene understanding
scene understanding
0.722022
Occlusion Boundary: A Formal Definition & Its Detection via Deep Exploration of Context · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Simultaneous Cast Shadows, Illumination and Geometry Inference Using Hypergraphs · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Rendering › novel view synthesis
real-time view synthesis
0.612022
Digging into Radiance Grid for Real-Time View Synthesis with Detail Preservation · ECCV (15) 2022
Computer vision › 3D vision › geometric estimation › registration
non-rigid registration
0.532016
Higher-Order Graph Principles towards Non-Rigid Surface Registration · IEEE Trans. Pattern Anal. Mach. Intell. 2016
A Generic Deformation Model for Dense Non-rigid Surface Registration: A Higher-Order MRF-Based Approach · ICCV 2013
Dense non-rigid surface registration using high-order graph matching · CVPR 2010
Machine learning › Deep learning architectures and training
attention mechanism
0.412020
Compressed Self-Attention for Deep Metric Learning · AAAI 2020
Computer vision › 3D vision › local feature descriptor
local descriptor learning
0.412020
Compressed Self-Attention for Deep Metric Learning · AAAI 2020
Computer vision › Face, body and person analysis
person re-identification
0.412020
Compressed Self-Attention for Deep Metric Learning · AAAI 2020
Machine learning › Deep learning architectures and training › attention mechanism
self-attention
0.412020
Compressed Self-Attention for Deep Metric Learning · AAAI 2020
Computer vision › 3D vision › geometric estimation › 3d registration
surface registration
0.422016
Higher-Order Graph Principles towards Non-Rigid Surface Registration · IEEE Trans. Pattern Anal. Mach. Intell. 2016
A Generic Deformation Model for Dense Non-rigid Surface Registration: A Higher-Order MRF-Based Approach · ICCV 2013
Computer vision › Video understanding and tracking
object tracking
0.422015
MUlti-Store Tracker (MUSTer): A cognitive psychology inspired approach to object tracking · CVPR 2015
Tracking Using Multilevel Quantizations · ECCV (6) 2014
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain mapping
0.412019
Geometry-Consistent Generative Adversarial Networks for One-Sided Unsupervised Domain Mapping · CVPR 2019
Indexing and storage engines
feature-based indexing
0.412019
An Indexing Framework for Efficient Visual Exploratory Subgraph Search in Graph Databases · ICDE 2019
Graph data management
graph indexing
0.412019
An Indexing Framework for Efficient Visual Exploratory Subgraph Search in Graph Databases · ICDE 2019
Machine learning › Generative modeling › generative adversarial network
image-to-image translation
0.312018
Perceptual Adversarial Networks for Image-to-Image Transformation · IEEE Trans. Image Process. 2018
Computer vision › 3D vision › depth estimation
monocular depth estimation
0.312018
Deep Ordinal Regression Network for Monocular Depth Estimation · CVPR 2018
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression
ordinal regression
0.312018
Deep Ordinal Regression Network for Monocular Depth Estimation · CVPR 2018
Image and video processing › image restoration
image deraining
0.312018
Perceptual Adversarial Networks for Image-to-Image Transformation · IEEE Trans. Image Process. 2018
Image and video processing
image restoration
0.312018
Perceptual Adversarial Networks for Image-to-Image Transformation · IEEE Trans. Image Process. 2018
Visual content generation and editing
image generation
0.312017
Tag Disentangled Generative Adversarial Network for Object Image Re-rendering · IJCAI 2017
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
markov random field
0.212016
Higher-Order Graph Principles towards Non-Rigid Surface Registration · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Graph algorithms and graph theory
graph matching
0.212016
Higher-Order Graph Principles towards Non-Rigid Surface Registration · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Computer vision › 3D vision
shape matching
0.222011
Intrinsic dense 3D surface tracking · CVPR 2011
Dense non-rigid surface registration using high-order graph matching · CVPR 2010
Computer vision › Video understanding and tracking › object tracking › appearance modeling
appearance model adaptation
0.212015
MUlti-Store Tracker (MUSTer): A cognitive psychology inspired approach to object tracking · CVPR 2015
Computer vision › Image recognition and object detection
object detection
0.212015
Scene-Domain Active Part Models for Object Representation · ICCV 2015
Computer vision › 3D vision
object representation
0.212015
Scene-Domain Active Part Models for Object Representation · ICCV 2015
Computer vision › Image recognition and object detection
part-based model
0.212015
Scene-Domain Active Part Models for Object Representation · ICCV 2015
Computer vision › Image recognition and object detection › object detection
part-based object detection
0.212015
Scene-Domain Active Part Models for Object Representation · ICCV 2015
Rendering
neural radiance fields
0.212022
Digging into Radiance Grid for Real-Time View Synthesis with Detail Preservation · ECCV (15) 2022

Methods — techniques the papers use, named apart from their topics

generative adversarial network · 1.2convolutional neural network · 0.9conditional random field · 0.8MAP inference · 0.6radiance grid · 0.6detail preservation · 0.6deep model · 0.6contextual information · 0.6graph indexing · 0.4compressed self-attention · 0.4channel grouping · 0.4markov random field · 0.4cycle consistency · 0.4VACCINE · 0.4ADVISE · 0.4perceptual adversarial loss · 0.3disentangled representation learning · 0.3dual decomposition · 0.2
YearPublicationVenuePosition
2026 Occlusion Boundary and Depth: Mutual Enhancement via Multi-Task Learning
abstract
Occlusion Boundary Estimation (OBE) identifies boundaries arising from both inter-object occlusions and self-occlusion within individual objects. This task is closely related to Monocular Depth Estimation (MDE), which infers depth from a single image, as Occlusion Boundaries (OBs) provide critical geometric cues for resolving depth ambiguities, while depth can conversely refine occlusion reasoning. In this paper, we aim to systematically model and exploit this mutually beneficial relationship. To this end, we propose MoDOT, a novel framework for joint estimation of depth and OBs, which incorporates a new Cross-Attention Strip Module (CASM) to leverage mid-level OB features for depth prediction, and a novel OB-Depth Constraint Loss (OBDCL) to enforce geometric consistency. To facilitate this study, we contribute OB-Hypersim, a large-scale photorealistic dataset with precise depth and self-occlusion-handled OB annotations. Extensive experiments on two synthetic datasets and NYUD-v2 demonstrate that MoDOT achieves significantly better performance than single-task baselines and multi-task competitors. Furthermore, models trained solely on our synthetic data demonstrate strong generalization to real-world scenes without fine-tuning, producing depth maps with sharper boundaries and improved geometric fidelity. Collectively, these results underscore the significant benefits of jointly modeling OBs and depth. Code and resources are available at HERE.
Lintao Xu, Yinghao Wang, Chaohui Wang
WACV3
2025 GMTPM: A General Multitask Pretrained Model for Electricity Data in Various Scenarios
abstract
The smart grid has become increasingly complex due to the integration of diverse energy sources and loads. Recognizing the limitations of traditional data sharing for future applications, the need to share model between different systems becomes unavoidable. An universal pretrained model for electricity data representation meets two primary challenges: The diversity of electricity data types and the range of issues in electricity analysis. To address these issues, this study presents a comprehensive univariate time-series representation learning framework called GMTPM for multiple tasks, for the first time in electricity analysis. This method extracts comprehensive feature extraction from univariate electricity data and shows improved performance across various tasks and data scenarios. Our approach incorporates a hybrid contextual attention mechanism to learn both local and global contextual relationships within the data. To adapt to various tasks, we optimize the model through three objectives: Reconstruction-based regression, context-consistent comparative learning, and time-consistent comparative learning. At last, we conducted extensive experiments on multiple univariate time-series datasets and achieved competitive performance in time-series analysis tasks such as imputation, forecasting, classification, anomaly detection, zero-shot forecasting, and few-shot classification.
Siying Zhou, Chaohui Wang, Fei Wang 0054
IEEE Trans. Ind. Informatics2
2022 Digging into Radiance Grid for Real-Time View Synthesis with Detail Preservation
Jinchi Huang, Bowen Cai 0001, Huan Fu, Mingming Gong, Chaohui Wang, Hongchen Luo, Rongfei Jia, Binqiang Zhao
ECCV (15)6
2022 Occlusion Boundary: A Formal Definition & Its Detection via Deep Exploration of Context
abstract
Occlusion boundaries contain rich perceptual information about the underlying scene structure and provide important cues in many visual perception-related tasks such as object recognition, segmentation, motion estimation, scene understanding, and autonomous navigation. However, there is no formal definition of occlusion boundaries in the literature, and state-of-the-art occlusion boundary detection is still suboptimal. With this in mind, in this paper we propose a formal definition of occlusion boundaries for related studies. Further, based on a novel idea, we develop two concrete approaches with different characteristics to detect occlusion boundaries in video sequences via enhanced exploration of contextual information (e.g, local structural boundary patterns, observations from surrounding regions, and temporal context) with deep models and conditional random fields. Experimental evaluations of our methods on two challenging occlusion boundary benchmarks (CMU and VSB100) demonstrate that our detectors significantly outperform the current state-of-the-art. Finally, we empirically assess the roles of several important components of the proposed detectors to validate the rationale behind these approaches.
Chaohui Wang, Huan Fu, Dacheng Tao, Michael J. Black
IEEE Trans. Pattern Anal. Mach. Intell.1
2020 Compressed Self-Attention for Deep Metric Learning
abstract
In this paper, we aim to enhance self-attention (SA) mechanism for deep metric learning in visual perception, by capturing richer contextual dependencies in visual data. To this end, we propose a novel module, named compressed self-attention (CSA), which significantly reduces the computation and memory cost with a neglectable decrease in accuracy with respect to the original SA mechanism, thanks to the following two characteristics: i) it only needs to compute a small number of base attention maps for a small number of base feature vectors; and ii) the output at each spatial location can be simply obtained by an adaptive weighted average of the outputs calculated from the base attention maps. The high computational efficiency of CSA enables the application to high-resolution shallow layers in convolutional neural networks with little additional cost. In addition, CSA makes it practical to further partition the feature maps into groups along the channel dimension and compute attention maps for features in each group separately, thus increasing the diversity of long-range dependencies and accordingly boosting the accuracy. We evaluate the performance of CSA via extensive experiments on two metric learning tasks: person re-identification and local descriptor learning. Qualitative and quantitative comparisons with latest methods demonstrate the significance of CSA in this topic.
Ziye Chen, Mingming Gong, Yanwu Xu 0003, Chaohui Wang, Kun Zhang 0001, Bo Du 0001
AAAI4
2020 Pixel-Pair Occlusion Relationship Map (P2ORM): Formulation, Inference and Application
Xuchong Qiu, Yang Xiao 0009, Chaohui Wang, Renaud Marlet
ECCV (4)3
2020 FERRARI: an efficient framework for visual exploratory subgraph search in graph databases
Chaohui Wang, Miao Xie, Sourav S. Bhowmick, Byron Choi, Xiaokui Xiao, Shuigeng Zhou
VLDB J.1
2019 Geometry-Consistent Generative Adversarial Networks for One-Sided Unsupervised Domain Mapping
abstract
Unsupervised domain mapping aims to learn a function GXY to translate domain X to Y in the absence of paired examples. Finding the optimal GXY without paired data is an ill-posed problem, so appropriate constraints are required to obtain reasonable solutions. While some prominent constraints such as cycle consistency and distance preservation successfully constrain the solution space, they overlook the special properties of images that simple geometric transformations do not change the image's semantic structure. Based on this special property, we develop a geometry-consistent generative adversarial network (Gc-GAN), which enables one-sided unsupervised domain mapping. GcGAN takes the original image and its counterpart image transformed by a predefined geometric transformation as inputs and generates two images in the new domain coupled with the corresponding geometry-consistency constraint. The geometry-consistency constraint reduces the space of possible solutions while keep the correct solutions in the search space. Quantitative and qualitative comparisons with the baseline (GAN alone) and the state-of-the-art methods including CycleGAN [66] and DistanceGAN [5] demonstrate the effectiveness of our method.
Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Batmanghelich, Kun Zhang 0001, Dacheng Tao
CVPR3
2019 An Indexing Framework for Efficient Visual Exploratory Subgraph Search in Graph Databases
abstract
Although exploratory search has received significant attention recently in the context of structured data, scant attention has been paid for graph-structured data. In this paper, we present two novel index structures called VACCINE and ADVISE to efficiently support exploratory subgraph search in a visual environment (VESS). VACCINE is an offline, feature-based index that stores rich information related to frequent and infrequent subgraphs in the underlying graph database and how they can be transformed from one subgraph to another. ADVISE, on the other hand, is an adaptive, compact, on-the-fly index instantiated during iterative visual formulation/reformulation of a subgraph query for exploratory search and records relevant information to efficiently support its repeated evaluation. These indexes engender more efficient and scalable visual exploratory subgraph search framework compared to a state-of-the-art technique.
Chaohui Wang, Miao Xie, Sourav S. Bhowmick, Byron Choi, Xiaokui Xiao, Shuigeng Zhou
ICDE1
2018 Robust Angular Local Descriptor Learning
Yanwu Xu 0001, Mingming Gong, Tongliang Liu, Kayhan Batmanghelich, Chaohui Wang
ACCV (5)5
2018 Deep Ordinal Regression Network for Monocular Depth Estimation
abstract
Monocular depth estimation, which plays a crucial role in understanding 3D scene geometry, is an ill-posed problem. Recent methods have gained significant improvement by exploring image-level information and hierarchical features from deep convolutional neural networks (DCNNs). These methods model depth estimation as a regression problem and train the regression networks by minimizing mean squared error, which suffers from slow convergence and unsatisfactory local solutions. Besides, existing depth estimation networks employ repeated spatial pooling operations, resulting in undesirable low-resolution feature maps. To obtain high-resolution depth maps, skip-connections or multilayer deconvolution networks are required, which complicates network training and consumes much more computations. To eliminate or at least largely reduce these problems, we introduce a spacing-increasing discretization (SID) strategy to discretize depth and recast depth network learning as an ordinal regression problem. By training the network using an ordinary regression loss, our method achieves much higher accuracy and faster convergence in synch. Furthermore, we adopt a multi-scale network structure which avoids unnecessary spatial pooling and captures multi-scale information in parallel. The proposed deep ordinal regression network (DORN) achieves state-of-the-art results on three challenging benchmarks, i.e., KITTI [16], Make3D [49], and NYU Depth v2 [41], and outperforms existing methods by a large margin.
Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Batmanghelich, Dacheng Tao
CVPR3
2018 MoE-SPNet: A mixture-of-experts scene parsing network
Huan Fu, Mingming Gong, Chaohui Wang, Dacheng Tao
Pattern Recognit.3
2018 Perceptual Adversarial Networks for Image-to-Image Transformation
abstract
In this paper, we propose Perceptual Adversarial Networks (PAN) for image-to-image transformations. Different from existing application driven algorithms, PAN provides a generic framework of learning to map from input images to desired images (Fig. 1), such as a rainy image to its de-rained counterpart, object edges to photos, semantic labels to a scenes image, etc. The proposed PAN consists of two feed-forward convolutional neural networks (CNNs): the image transformation network T and the discriminative network D. Besides the generative adversarial loss widely used in GANs, we propose the perceptual adversarial loss, which undergoes an adversarial training process between the image transformation network T and the hidden layers of the discriminative network D. The hidden layers and the output of the discriminative network D are upgraded to constantly and automatically discover the discrepancy between the transformed image and the corresponding ground-truth, while the image transformation network T is trained to minimize the discrepancy explored by the discriminative network D. Through integrating the generative adversarial loss and the perceptual adversarial loss, D and T can be trained alternately to solve image-to-image transformation tasks. Experiments evaluated on several image-to-image transformation tasks (e.g., image de-raining, image inpainting, etc) demonstrate the effectiveness of the proposed PAN and its advantages over many existing works.
Chang Xu 0002, Chaohui Wang, Dacheng Tao
IEEE Trans. Image Process.3
2017 Tag Disentangled Generative Adversarial Network for Object Image Re-rendering
abstract
In this paper, we propose a principled Tag Disentangled Generative Adversarial Networks (TD-GAN) for re-rendering new images for the object of interest from a single image of it by specifying multiple scene properties (such as viewpoint, illumination, expression, etc.). The whole framework consists of a disentangling network, a generative network, a tag mapping net, and a discriminative network, which are trained jointly based on a given set of images that are completely/partially tagged (i.e., supervised/semi-supervised setting). Given an input image, the disentangling network extracts disentangled and interpretable representations, which are then used to generate images by the generative network. In order to boost the quality of disentangled representations, the tag mapping net is integrated to explore the consistency between the image and its tags. Furthermore, the discriminative network is introduced to implement the adversarial training strategy for generating more realistic images. Experiments on two challenging datasets demonstrate the state-of-the-art performance of the proposed framework in the problem of interest.
Chaohui Wang, Chang Xu 0002, Dacheng Tao
IJCAI2
2017 Consistency-Constrained Nonnegative Coding for Tracking
abstract
A novel visual object tracking method based on consistency-constrained nonnegative coding (CNC) is proposed in this paper. For the purpose of computational efficiency, superpixels are first extracted from each observed video frame. And then CNC is performed based on those obtained superpixels, where the locality on manifold is preserved by enforcing the temporal and spatial smoothness. The coding result is achieved via an iterative update scheme, which is proved to converge. The proposed method enhances the coding stability and makes the tracker more robust for object tracking. The tracking performance has been evaluated based on ten challenging benchmark sequences involving drastic motion, partial or severe occlusions, large variation in pose, and illumination variation. The experimental results demonstrate the superior performance of our method in comparison with ten state-of-art trackers.
Xiaolin Tian 0002, Licheng Jiao, Zhipeng Gan, Chaohui Wang, Xiaoli Zheng
IEEE Trans. Circuits Syst. Video Technol.4
2016 Occlusion Boundary Detection via Deep Exploration of Context
abstract
Occlusion boundaries contain rich perceptual information about the underlying scene structure. They also provide important cues in many visual perception tasks such as scene understanding, object recognition, and segmentation. In this paper, we improve occlusion boundary detection via enhanced exploration of contextual information (e.g., local structural boundary patterns, observations from surrounding regions, and temporal context), and in doing so develop a novel approach based on convolutional neural networks (CNNs) and conditional random fields (CRFs). Experimental results demonstrate that our detector significantly outperforms the state-of-the-art (e.g., improving the F-measure from 0.62 to 0.71 on the commonly used CMU benchmark). Last but not least, we empirically assess the roles of several important components of the proposed detector, so as to validate the rationale behind this approach.
Huan Fu, Chaohui Wang, Dacheng Tao, Michael J. Black
CVPR2
2016 Inference and Learning of Graphical Models: Theory and Applications in Computer Vision and Image Analysis
Chaohui Wang, Nikos Komodakis, Hiroshi Ishikawa 0002, Olga Veksler, Endre Boros
Comput. Vis. Image Underst.1
2016 Higher-Order Graph Principles towards Non-Rigid Surface Registration
abstract
This paper casts surface registration as the problem of finding a set of discrete correspondences through the minimization of an energy function, which is composed of geometric and appearance matching costs, as well as higher-order deformation priors. Two higher-order graph-based formulations are proposed under different deformation assumptions. The first formulation encodes isometric deformations using conformal geometry in a higher-order graph matching problem, which is solved through dual-decomposition and is able to handle partial matching. Despite the isometry assumption, this approach is able to robustly match sparse feature point sets on surfaces undergoing highly anisometric deformations. Nevertheless, its performance degrades significantly when addressing anisometric registration for a set of densely sampled points. This issue is rigorously addressed subsequently through a novel deformation model that is able to handle arbitrary diffeomorphisms between two surfaces. Such a deformation model is introduced into a higher-order Markov Random Field for dense surface registration, and is inferred using a new parallel and memory efficient algorithm. To deal with the prohibitive search space, we also design an efficient way to select a number of matching candidates for each point of the source surface based on the matching results of a sparse set of points. A series of experiments demonstrate the accuracy and the efficiency of the proposed framework, notably in challenging cases of large and/or anisometric deformations, or surfaces that are partially occluded.
Chaohui Wang, Xianfeng Gu, Dimitris Samaras, Nikos Paragios
IEEE Trans. Pattern Anal. Mach. Intell.2
2015 MUlti-Store Tracker (MUSTer): A cognitive psychology inspired approach to object tracking
abstract
Variations in the appearance of a tracked object, such as changes in geometry/photometry, camera viewpoint, illumination, or partial occlusion, pose a major challenge to object tracking. Here, we adopt cognitive psychology principles to design a flexible representation that can adapt to changes in object appearance during tracking. Inspired by the well-known Atkinson-Shiffrin Memory Model, we propose MUlti-Store Tracker (MUSTer), a dual-component approach consisting of short- and long-term memory stores to process target appearance memories. A powerful and efficient Integrated Correlation Filter (ICF) is employed in the short-term store for short-term tracking. The integrated long-term component, which is based on keypoint matching-tracking and RANSAC estimation, can interact with the long-term memory and provide additional information for output control. MUSTer was extensively evaluated on the CVPR2013 Online Object Tracking Benchmark (OOTB) and ALOV++ datasets. The experimental results demonstrated the superior performance of MUSTer in comparison with other state-of-art trackers.
Zhibin Hong, Zhe Chen 0013, Chaohui Wang, Xue Mei, Danil V. Prokhorov, Dacheng Tao
CVPR3
2015 Scene-Domain Active Part Models for Object Representation
abstract
In this paper, we are interested in enhancing the expressivity and robustness of part-based models for object representation, in the common scenario where the training data are based on 2D images. To this end, we propose scene-domain active part models (SDAPM), which reconstruct and characterize the 3D geometric statistics between object's parts in 3D scene-domain by using 2D training data in the image-domain alone. And on top of this, we explicitly model and handle occlusions in SDAPM. Together with the developed learning and inference algorithms, such a model provides rich object descriptions, including 2D object and parts localization, 3D landmark shape and camera viewpoint, which offers an effective representation to various image understanding tasks, such as object and parts detection, 3D landmark shape and viewpoint estimation from images. Experiments on the above tasks show that SDAPM outperforms previous part-based models, and thus demonstrates the potential of the proposed technique.
Zhou Ren, Chaohui Wang, Alan L. Yuille
ICCV2
2014 Tracking Using Multilevel Quantizations
Zhibin Hong, Chaohui Wang, Xue Mei, Danil V. Prokhorov, Dacheng Tao
ECCV (6)2
2013 Nonlinearly Constrained MRFs: Exploring the Intrinsic Dimensions of Higher-Order Cliques
abstract
This paper introduces an efficient approach to integrating non-local statistics into the higher-order Markov Random Fields (MRFs) framework. Motivated by the observation that many non-local statistics (e.g., shape priors, color distributions) can usually be represented by a small number of parameters, we reformulate the higher-order MRF model by introducing additional latent variables to represent the intrinsic dimensions of the higher-order cliques. The resulting new model, called NC-MRF, not only provides the flexibility in representing the configurations of higher-order cliques, but also automatically decomposes the energy function into less coupled terms, allowing us to design an efficient algorithmic framework for maximum a posteriori (MAP) inference. Based on this novel modeling/ inference framework, we achieve state-of-the-art solutions to the challenging problems of class-specific image segmentation and template-based 3D facial expression tracking, which demonstrate the potential of our approach.
Chaohui Wang, Stefano Soatto, Shing-Tung Yau
CVPR2
2013 A Generic Deformation Model for Dense Non-rigid Surface Registration: A Higher-Order MRF-Based Approach
abstract
We propose a novel approach for dense non-rigid 3D surface registration, which brings together Riemannian geometry and graphical models. To this end, we first introduce a generic deformation model, called Canonical Distortion Coefficients (CDCs), by characterizing the deformation of every point on a surface using the distortions along its two principle directions. This model subsumes the deformation groups commonly used in surface registration such as isometry and conformality, and is able to handle more complex deformations. We also derive its discrete counterpart which can be computed very efficiently in a closed form. Based on these, we introduce a higher-order Markov Random Field (MRF) model which seamlessly integrates our deformation model and a geometry/texture similarity metric. Then we jointly establish the optimal correspondences for all the points via maximum a posteriori (MAP) inference. Moreover, we develop a parallel optimization algorithm to efficiently perform the inference for the proposed higher-order MRF model. The resulting registration algorithm outperforms state-of-the-art methods in both dense non-rigid 3D surface registration and tracking.
Chaohui Wang, Xianfeng Gu, Dimitris Samaras, Nikos Paragios
ICCV2
2013 Markov Random Field modeling, inference & learning in computer vision & image understanding: A survey
Chaohui Wang, Nikos Komodakis, Nikos Paragios
Comput. Vis. Image Underst.1
2013 Simultaneous Cast Shadows, Illumination and Geometry Inference Using Hypergraphs
abstract
The cast shadows in an image provide important information about illumination and geometry. In this paper, we utilize this information in a novel framework in order to jointly recover the illumination environment, a set of geometry parameters, and an estimate of the cast shadows in the scene given a single image and coarse initial 3D geometry. We model the interaction of illumination and geometry in the scene and associate it with image evidence for cast shadows using a higher order Markov Random Field (MRF) illumination model, while we also introduce a method to obtain approximate image evidence for cast shadows. Capturing the interaction between light sources and geometry in the proposed graphical model necessitates higher order cliques and continuous-valued variables, which make inference challenging. Taking advantage of domain knowledge, we provide a two-stage minimization technique for the MRF energy of our model. We evaluate our method in different datasets, both synthetic and real. Our model is robust to rough knowledge of geometry and inaccurate initial shadow estimates, allowing a generic coarse 3D model to represent a whole class of objects for the task of illumination estimation, or the estimation of geometry parameters to refine our initial knowledge of scene geometry, simultaneously with illumination estimation.
Alexandros Panagopoulos, Chaohui Wang, Dimitris Samaras, Nikos Paragios
IEEE Trans. Pattern Anal. Mach. Intell.2
2011 Illumination estimation and cast shadow detection through a higher-order graphical model
abstract
In this paper, we propose a novel framework to jointly recover the illumination environment and an estimate of the cast shadows in a scene from a single image, given coarse 3D geometry. We describe a higher-order Markov Random Field (MRF) illumination model, which combines low-level shadow evidence with high-level prior knowledge for the joint estimation of cast shadows and the illumination environment. First, a rough illumination estimate and the structure of the graphical model in the illumination space is determined through a voting procedure. Then, a higher order approach is considered where illumination sources are coupled with the observed image and the latent variables corresponding to the shadow detection. We examine two inference methods in order to effectively minimize the MRF energy of our model. Experimental evaluation shows that our approach is robust to rough knowledge of geometry and reflectance and inaccurate initial shadow estimates. We demonstrate the power of our MRF illumination model on various datasets and show that we can estimate the illumination in images of objects belonging to the same class using the same coarse 3D model to represent all instances of the class.
Alexandros Panagopoulos, Chaohui Wang, Dimitris Samaras, Nikos Paragios
CVPR2
2011 Intrinsic dense 3D surface tracking
abstract
This paper presents a novel intrinsic 3D surface distance and its use in a complete probabilistic tracking framework for dynamic 3D data. Registering two frames of a deforming 3D shape relies on accurate correspondences between all points across the two frames. In the general case such correspondence search is computationally intractable. Common prior assumptions on the nature of the deformation such as near-rigidity, isometry or learning from a training set, reduce the search space but often at the price of loss of accuracy when it comes to deformations not in the prior assumptions. If we consider the set of all possible 3D surface matchings defined by specifying triplets of correspondences in the uniformization domain, then we introduce a new matching cost between two 3D surfaces. The lowest feature differences across this set of matchings that cause two points to correspond, become the matching cost of that particular correspondence. We show that for surface tracking applications, the matching cost can be efficiently computed in the uniformization domain. This matching cost is then combined with regularization terms that enforce spatial and temporal motion consistencies, into a maximum a posteriori (MAP) problem which we approximate using a Markov Random Field (MRF). Compared to previous 3D surface tracking approaches that either assume isometric deformations or consistent features, our method achieves dense, accurate tracking results, which we demonstrate through a series of dense, anisometric 3D surface tracking experiments.
Chaohui Wang, Yang Wang 0001, Xianfeng Gu, Dimitris Samaras, Nikos Paragios
CVPR2
2011 Viewpoint invariant 3D landmark model inference from monocular 2D images using higher-order priors
abstract
In this paper, we propose a novel one-shot optimization approach to simultaneously determine both the optimal 3D landmark model and the corresponding 2D projections without explicit estimation of the camera viewpoint, which is also able to deal with misdetections as well as partial occlusions. To this end, a 3D shape manifold is built upon fourth-order interactions of landmarks from a training set where pose-invariant statistics are obtained in this space. The 3D-2D consistency is also encoded in such high-order interactions, which eliminate the necessity of viewpoint estimation. Furthermore, the modeling of visibility improves further the performance of the method by handling missing correspondences and occlusions. The inference is addressed through a MAP formulation which is naturally transformed into a higher-order MRF optimization problem and is solved using a dual-decomposition-based method. Promising results on standard face benchmarks demonstrate the potential of our approach.
Chaohui Wang, Loïc Simon, Ioannis A. Kakadiaris, Dimitris Samaras, Nikos Paragios
ICCV1
2011 Pose-Invariant 3D Proximal Femur Estimation through Bi-planar Image Segmentation with Hierarchical Higher-Order Graph-Based Priors
Chaohui Wang, Haithem Boussaid, Loïc Simon, Jean-Yves Lazennec, Nikos Paragios
MICCAI (3)1
2010 Dense non-rigid surface registration using high-order graph matching
abstract
In this paper, we propose a high-order graph matching formulation to address non-rigid surface matching. The singleton terms capture the geometric and appearance similarities (e.g., curvature and texture) while the high-order terms model the intrinsic embedding energy. The novelty of this paper includes: 1. casting 3D surface registration into a graph matching problem that combines both geometric and appearance similarities and intrinsic embedding information, 2. the first implementation of high-order graph matching algorithm that solves a non-convex optimization problem, and 3. an efficient two-stage optimization approach to constrain the search space for dense surface registration. Our method is validated through a series of experiments demonstrating its accuracy and efficiency, notably in challenging cases of large and/or non-isometric deformations, or meshes that are partially occluded.
Chaohui Wang, Yang Wang 0001, Xianfeng Gu, Dimitris Samaras, Nikos Paragios
CVPR2
2010 3D Knowledge-Based Segmentation Using Pose-Invariant Higher-Order Graphs
Chaohui Wang, Olivier Teboul, Fabrice Michel, Salma Essafi, Nikos Paragios
MICCAI (3)1
2009 Segmentation, ordering and multi-object tracking using graphical models
abstract
In this paper, we propose a unified graphical-model framework to interpret a scene composed of multiple objects in monocular video sequences. Using a single pairwise Markov random field (MRF), all the observed and hidden variables of interest such as image intensities, pixels' states (associated object's index and relative depth), objects' states (model motion parameters and relative depth) are jointly considered. Particular attention is given to occlusion handling by introducing a rigorous visibility modeling within the MRF formulation. Through minimizing the MRF's energy, we simultaneously segment, track and sort by depth the objects. Promising experimental results demonstrate the potential of this framework and its robustness to image noise, cluttered background, moving camera and background, and even complete occlusions.
Chaohui Wang, Martin de La Gorce, Nikos Paragios
ICCV1