EDBT 2026 Demo / reviewers in the wild / expert
Wei Yang 0019
dblp:03/1094-19
· DBLP profile ↗
28ranked-venue papers
9as first author
10since 2021 · last 2024
0000-0003-3975-2472ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 8 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 5 since 2021Systems, architecture and hardware · 7 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Towards Detailed Text-to-Motion Synthesis via Basic-to-Advanced Hierarchical Diffusion ModelabstractText-guided motion synthesis aims to generate 3D human motion that not only precisely reflects the textual description but reveals the motion details as much as possible. Pioneering methods explore the diffusion model for text-to-motion synthesis and obtain significant superiority. However, these methods conduct diffusion processes either on the raw data distribution or the low-dimensional latent space, which typically suffer from the problem of modality inconsistency or detail-scarce. To tackle this problem, we propose a novel Basic-to-Advanced Hierarchical Diffusion Model, named B2A-HDM, to collaboratively exploit low-dimensional and high-dimensional diffusion models for high quality detailed motion synthesis. Specifically, the basic diffusion model in low-dimensional latent space provides the intermediate denoising result that to be consistent with the textual description, while the advanced diffusion model in high-dimensional latent space focuses on the following detail-enhancing denoising process. Besides, we introduce a multi-denoiser framework for the advanced diffusion model to ease the learning of high-dimensional model and fully explore the generative potential of the diffusion model. Quantitative and qualitative experiment results on two text-to-motion benchmarks (HumanML3D and KIT-ML) demonstrate that B2A-HDM can outperform existing state-of-the-art methods in terms of fidelity, modality consistency, and diversity. Zhenyu Xie, Yang Wu 0001, Xuehao Gao, Zhongqian Sun, Wei Yang 0019, Xiaodan Liang |
AAAI | 5 |
| 2024 | FoundationPose: Unified 6D Pose Estimation and Tracking of Novel ObjectsabstractWe present FoundationPose, a unified foundation model for 6D object pose estimation and tracking, supporting both model-based and model-free setups. Our approach can be instantly applied at test-time to a novel object without finetuning, as long as its CAD model is given, or a small number of reference images are captured. Thanks to the unified framework, the downstream pose estimation modules are the same in both setups, with a neural implicit representation used for efficient novel view synthesis when no CAD model is available. Strong generalizability is achieved via large-scale synthetic training, aided by a large language model (LLM), a novel transformer-based architecture, and contrastive learning formulation. Extensive evaluation on multiple public datasets involving challenging scenarios and objects indicate our unified approach outperforms existing methods specialized for each task by a large margin. In addition, it even achieves comparable results to instance-level methods despite the reduced assumptions. Project page: https://nvlabs.github.io/FoundationPose/ Bowen Wen, Wei Yang 0019, Jan Kautz, Stanley T. Birchfield |
CVPR | 2 |
| 2024 | Learning Pseudo 3D Guidance for View-Consistent Texturing with 2D Diffusion
Kehan Li 0002, Yanbo Fan, Yang Wu 0001, Zhongqian Sun, Wei Yang 0019, Xiangyang Ji, Li Yuan 0007, Jie Chen 0001 |
ECCV (86) | 5 |
| 2024 | SynH2R: Synthesizing Hand-Object Motions for Learning Human-to-Robot HandoversabstractVision-based human-to-robot handover is an important and challenging task in human-robot interaction. Recent work has attempted to train robot policies by interacting with dynamic virtual humans in simulated environments, where the policies can later be transferred to the real world. However, a major bottleneck is the reliance on human motion capture data, which is expensive to acquire and difficult to scale to arbitrary objects and human grasping motions. In this paper, we introduce a framework that can generate plausible human grasping motions suitable for training the robot. To achieve this, we propose a hand-object synthesis method that is designed to generate handover-friendly motions similar to humans. This allows us to generate synthetic training and testing data with 100x more objects than previous work. In our experiments, we show that our method trained purely with synthetic data is competitive with state-of-the-art methods that rely on real human motion data both in simulation and on a real system. In addition, we can perform evaluations on a larger scale compared to prior work. With our newly introduced test set, we show that our model can better scale to a large variety of unseen objects and human motions compared to the baselines. Sammy Joe Christen, Lan Feng, Wei Yang 0019, Yu-Wei Chao, Otmar Hilliges, Jie Song 0006 |
ICRA | 3 |
| 2023 | Learning Human-to-Robot Handovers from Point CloudsabstractWe propose the first framework to learn control policies for vision-based human-to-robot handovers, a critical task for human-robot interaction. While research in Embodied AI has made significant progress in training robot agents in simulated environments, interacting with humans remains challenging due to the difficulties of simulating humans. Fortunately, recent research has developed realistic simulated environments for human-to-robot handovers. Leveraging this result, we introduce a method that is trained with a human-in-the-loop via a two-stage teacher-student framework that uses motion and grasp planning, reinforcement learning, and self-supervision. We show significant performance gains over baselines on a simulation benchmark, sim-to-sim transfer and sim-to-real transfer. Video and code are available at https://handover-sim2real.github.io. Sammy Joe Christen, Wei Yang 0019, Claudia Pérez-D'Arpino, Otmar Hilliges, Dieter Fox, Yu-Wei Chao |
CVPR | 2 |
| 2023 | Act As You Wish: Fine-Grained Control of Motion Diffusion Model with Hierarchical Semantic GraphsabstractMost text-driven human motion generation methods employ sequential modeling approaches, e.g., transformer, to extract sentence-level text representations automatically and implicitly for human motion synthesis. However, these compact text representations may overemphasize the action names at the expense of other important properties and lack fine-grained details to guide the synthesis of subtly distinct motion. In this paper, we propose hierarchical semantic graphs for fine-grained control over motion generation. Specifically, we disentangle motion descriptions into hierarchical semantic graphs including three levels of motions, actions, and specifics. Such global-to-local structures facilitate a comprehensive understanding of motion description and fine-grained control of motion generation. Correspondingly, to leverage the coarse-to-fine topology of hierarchical semantic graphs, we decompose the text-to-motion diffusion process into three semantic levels, which correspond to capturing the overall motion, local actions, and action specifics. Extensive experiments on two benchmark human motion datasets, including HumanML3D and KIT, with superior performances, justify the efficacy of our method. More encouragingly, by modifying the edge weights of hierarchical semantic graphs, our method can continuously refine the generated motion, which may have a far-reaching impact on the community. Code and pre-trained weights are available at https://github.com/jpthu17/GraphMotion. Peng Jin 0001, Yang Wu 0001, Yanbo Fan, Zhongqian Sun, Wei Yang 0019, Li Yuan 0007 |
NeurIPS | 5 |
| 2022 | HandoverSim: A Simulation Framework and Benchmark for Human-to-Robot Object HandoversabstractWe introduce a new simulation benchmark “Han-doverSim” for human-to-robot object handovers. To simulate the giver's motion, we leverage a recent motion capture dataset of hand grasping of objects. We create training and evaluation environments for the receiver with standardized protocols and metrics. We analyze the performance of a set of baselines and show a correlation with a real-world evaluation.11Code is open sourced at https://handover-sim.github.io. Yu-Wei Chao, Chris Paxton 0001, Yu Xiang 0001, Wei Yang 0019, Balakumar Sundaralingam, Tao Chen 0046, Adithyavairavan Murali, Maya Cakmak, Dieter Fox |
ICRA | 4 |
| 2022 | Model Predictive Control for Fluid Human-to-Robot HandoversabstractHuman-robot handover is a fundamental yet challenging task in human-robot interaction and collaboration. Recently, remarkable progressions have been made in human-to-robot handovers of unknown objects by using learning-based grasp generators. However, how to responsively generate smooth motions to take an object from a human is still an open question. Specifically, planning motions that take human comfort into account is not a part of the human-robot handover process in most prior works. In this paper, we propose to generate smooth motions via an efficient model-predictive control (MPC) framework that integrates perception and complex domain-specific constraints into the optimization problem. We introduce a learning-based grasp reachability model to select candidate grasps which maximize the robot's manipulability, giving it more freedom to satisfy these constraints. Finally, we integrate a neural net force/torque classifier that detects contact events from noisy data. We conducted human-to-robot handover experiments on a diverse set of objects with several users ($N=4$) and performed a systematic evaluation of each module. The study shows that the users preferred our MPC approach over the baseline system by a large margin. Wei Yang 0019, Balakumar Sundaralingam, Chris Paxton 0001, Iretiayo Akinola, Yu-Wei Chao, Maya Cakmak, Dieter Fox |
ICRA | 1 |
| 2021 | DexYCB: A Benchmark for Capturing Hand Grasping of ObjectsabstractWe introduce DexYCB, a new dataset for capturing hand grasping of objects. We first compare DexYCB with a related one through cross-dataset evaluation. We then present a thorough benchmark of state-of-the-art approaches on three relevant tasks: 2D object and keypoint detection, 6D object pose estimation, and 3D hand pose estimation. Finally, we evaluate a new robotics-relevant task: generating safe robot grasps in human-to-robot object handover.1 Yu-Wei Chao, Wei Yang 0019, Yu Xiang 0001, Pavlo Molchanov 0001, Ankur Handa, Jonathan Tremblay, Yashraj Narang, Karl Van Wyk, Umar Iqbal 0001, Stanley T. Birchfield, Jan Kautz, Dieter Fox |
CVPR | 2 |
| 2021 | Reactive Human-to-Robot Handovers of Arbitrary ObjectsabstractHuman-robot object handovers have been an actively studied area of robotics over the past decade; however, very few techniques and systems have addressed the challenge of handing over diverse objects with arbitrary appearance, size, shape, and deformability. In this paper, we present a vision-based system that enables reactive human-to-robot handovers of unknown objects. Our approach combines closed-loop motion planning with real-time, temporally consistent grasp generation to ensure reactivity and motion smoothness. Our system is robust to different object positions and orientations, and can grasp both rigid and non-rigid objects. We demonstrate the generalizability, usability, and robustness of our approach on a novel benchmark set of 26 diverse household objects, a user study with six participants handing over a subset of 15 objects, and a systematic evaluation examining different ways of handing objects. Wei Yang 0019, Chris Paxton 0001, Arsalan Mousavian, Yu-Wei Chao, Maya Cakmak, Dieter Fox |
ICRA | 1 |
| 2020 | DexPilot: Vision-Based Teleoperation of Dexterous Robotic Hand-Arm SystemabstractTeleoperation offers the possibility of imparting robotic systems with sophisticated reasoning skills, intuition, and creativity to perform tasks. However, teleoperation solutions for high degree-of-actuation (DoA), multi-fingered robots are generally cost-prohibitive, while low-cost offerings usually offer reduced degrees of control. Herein, a low-cost, depth-based teleoperation system, DexPilot, was developed that allows for complete control over the full 23 DoA robotic system by merely observing the bare human hand. DexPilot enabled operators to solve a variety of complex manipulation tasks that go beyond simple pick-and-place operations and performance was measured through speed and reliability metrics. DexPilot cost-effectively enables the production of high dimensional, multi-modality, state-action data that can be leveraged in the future to learn sensorimotor policies for challenging manipulation tasks. The videos of the experiments can be found at https://sites.google.com/view/dex-pilot. Ankur Handa, Karl Van Wyk, Wei Yang 0019, Jacky Liang, Yu-Wei Chao, Stanley T. Birchfield, Nathan D. Ratliff, Dieter Fox |
ICRA | 3 |
| 2020 | Collaborative Interaction Models for Optimized Human-Robot TeamworkabstractEffective human-robot collaboration requires informed anticipation. The robot must anticipate the human's actions, but also react quickly and intuitively when its predictions are wrong. The robot must plan its actions to account for the human's own plan, with the knowledge that the human's behavior will change based on what the robot actually does. This cyclical game of predicting a human's future actions and generating a corresponding motion plan is extremely difficult to model using standard techniques. In this work, we describe a novel Model Predictive Control (MPC)-based framework for finding optimal trajectories in a collaborative, multi-agent setting, in which we simultaneously plan for the robot while predicting the actions of its external collaborators. We use human-robot handovers to demonstrate that with a strong model of the collaborator, our framework produces fluid, reactive human-robot interactions in novel, cluttered environments. Our method efficiently generates coordinated trajectories, and achieves a high success rate in handover, even in the presence of significant sensor noise. Adam Fishman, Chris Paxton 0001, Wei Yang 0019, Dieter Fox, Byron Boots, Nathan D. Ratliff |
IROS | 3 |
| 2020 | Human Grasp Classification for Reactive Human-to-Robot HandoversabstractTransfer of objects between humans and robots is a critical capability for collaborative robots. Although there has been a recent surge of interest in human-robot handovers, most prior research focus on robot-to-human handovers. Further, work on the equally critical human-to-robot handovers often assumes humans can place the object in the robot's gripper. In this paper, we propose an approach for human-to-robot handovers in which the robot meets the human halfway, by classifying the human's grasp of the object and quickly planning a trajectory accordingly to take the object from the human's hand according to their intent. To do this, we collect a human grasp dataset which covers typical ways of holding objects with various hand shapes and poses, and learn a deep model on this dataset to classify the hand grasps into one of these categories. We present a planning and execution approach that takes the object from the human hand according to the detected grasp and hand position, and replans as necessary when the handover is interrupted. Through a systematic evaluation, we demonstrate that our system results in more fluent handovers versus two baselines. We also present findings from a user study (N = 9) demonstrating the effectiveness and usability of our approach with naive users in different scenarios. More information can be found at http://wyang.me/handovers. Wei Yang 0019, Chris Paxton 0001, Maya Cakmak, Dieter Fox |
IROS | 1 |
| 2019 | Visual Semantic Navigation using Scene Priors
Wei Yang 0019, Xiaolong Wang 0004, Ali Farhadi, Abhinav Gupta 0001, Roozbeh Mottaghi |
ICLR (Poster) | 1 |
| 2019 | Progressively diffused networks for semantic visual parsing
Ruimao Zhang, Wei Yang 0019, Zhanglin Peng, Pengxu Wei, Xiaogang Wang 0001, Liang Lin 0004 |
Pattern Recognit. | 2 |
| 2018 | 3D Human Pose Estimation in the Wild by Adversarial LearningabstractRecently, remarkable advances have been achieved in 3D human pose estimation from monocular images because of the powerful Deep Convolutional Neural Networks (DCNNs). Despite their success on large-scale datasets collected in the constrained lab environment, it is difficult to obtain the 3D pose annotations for in-the-wild images. Therefore, 3D human pose estimation in the wild is still a challenge. In this paper, we propose an adversarial learning framework, which distills the 3D human pose structures learned from the fully annotated dataset to in-the-wild images with only 2D pose annotations. Instead of defining hard-coded rules to constrain the pose estimation results, we design a novel multi-source discriminator to distinguish the predicted 3D poses from the ground-truth, which helps to enforce the pose estimator to generate anthropometrically valid poses even with images in the wild. We also observe that a carefully designed information source for the discriminator is essential to boost the performance. Thus, we design a geometric descriptor, which computes the pairwise relative locations and distances between body joints, as a new information source for the discriminator. The efficacy of our adversarial learning framework with the new geometric descriptor has been demonstrated through extensive experiments on widely used public benchmarks. Our approach significantly improves the performance compared with previous state-of-the-art approaches. Wei Yang 0019, Wanli Ouyang, Xiaolong Wang 0004, Jimmy S. J. Ren, Hongsheng Li 0001, Xiaogang Wang 0001 |
CVPR | 1 |
| 2017 | Multi-context Attention for Human Pose EstimationabstractIn this paper, we propose to incorporate convolutional neural networks with a multi-context attention mechanism into an end-to-end framework for human pose estimation. We adopt stacked hourglass networks to generate attention maps from features at multiple resolutions with various semantics. The Conditional Random Field (CRF) is utilized to model the correlations among neighboring regions in the attention map. We further combine the holistic attention model, which focuses on the global consistency of the full human body, and the body part attention model, which focuses on detailed descriptions for different body parts. Hence our model has the ability to focus on different granularity from local salient regions to global semantic consistent spaces. Additionally, we design novel Hourglass Residual Units (HRUs) to increase the receptive field of the network. These units are extensions of residual units with a side branch incorporating filters with larger receptive field, hence features with various scales are learned and combined within the HRUs. The effectiveness of the proposed multi-context attention mechanism and the hourglass residual units is evaluated on two widely used human pose estimation benchmarks. Our approach outperforms all existing methods on both benchmarks over all the body parts. Code has been made publicly available. Xiao Chu, Wei Yang 0019, Wanli Ouyang, Alan L. Yuille, Xiaogang Wang 0001 |
CVPR | 2 |
| 2017 | Identity-Aware Textual-Visual Matching with Latent Co-attentionabstractTextual-visual matching aims at measuring similarities between sentence descriptions and images. Most existing methods tackle this problem without effectively utilizing identity-level annotations. In this paper, we propose an identity-aware two-stage framework for the textual-visual matching problem. Our stage-1 CNN-LSTM network learns to embed cross-modal features with a novel Cross-Modal Cross-Entropy (CMCE) loss. The stage-1 network is able to efficiently screen easy incorrect matchings and also provide initial training point for the stage-2 training. The stage-2 CNN-LSTM network refines the matching results with a latent co-attention mechanism. The spatial attention relates each word with corresponding image regions while the latent semantic attention aligns different sentence structures to make the matching results more robust to sentence structure variations. Extensive experiments on three datasets with identity-level annotations show that our framework outperforms state-of-the-art approaches by large margins. Shuang Li 0013, Tong Xiao 0003, Hongsheng Li 0001, Wei Yang 0019, Xiaogang Wang 0001 |
ICCV | 4 |
| 2017 | Learning Feature Pyramids for Human Pose EstimationabstractArticulated human pose estimation is a fundamental yet challenging task in computer vision. The difficulty is particularly pronounced in scale variations of human body parts when camera view changes or severe foreshortening happens. Although pyramid methods are widely used to handle scale changes at inference time, learning feature pyramids in deep convolutional neural networks (DCNNs) is still not well explored. In this work, we design a Pyramid Residual Module (PRMs) to enhance the invariance in scales of DCNNs. Given input features, the PRMs learn convolutional filters on various scales of input features, which are obtained with different subsampling ratios in a multibranch network. Moreover, we observe that it is inappropriate to adopt existing methods to initialize the weights of multi-branch networks, which achieve superior performance than plain networks in many tasks recently. Therefore, we provide theoretic derivation to extend the current weight initialization scheme to multi-branch network structures. We investigate our method on two standard benchmarks for human pose estimation. Our approach obtains state-of-the-art results on both benchmarks. Code is available at https://github.com/bearpaw/PyraNet. Wei Yang 0019, Shuang Li 0013, Wanli Ouyang, Hongsheng Li 0001, Xiaogang Wang 0001 |
ICCV | 1 |
| 2016 | End-to-End Learning of Deformable Mixture of Parts and Deep Convolutional Neural Networks for Human Pose EstimationabstractRecently, Deep Convolutional Neural Networks (DCNNs) have been applied to the task of human pose estimation, and have shown its potential of learning better feature representations and capturing contextual relationships. However, it is difficult to incorporate domain prior knowledge such as geometric relationships among body parts into DCNNs. In addition, training DCNN-based body part detectors without consideration of global body joint consistency introduces ambiguities, which increases the complexity of training. In this paper, we propose a novel end-to-end framework for human pose estimation that combines DCNNs with the expressive deformable mixture of parts. We explicitly incorporate domain prior knowledge into the framework, which greatly regularizes the learning process and enables the flexibility of our framework for loopy models or tree-structured models. The effectiveness of jointly learning a DCNN with a deformable mixture of parts model is evaluated through intensive experiments on several widely used benchmarks. The proposed approach significantly improves the performance compared with state-of-the-art approaches, especially on benchmarks with challenging articulations. Wei Yang 0019, Wanli Ouyang, Hongsheng Li 0001, Xiaogang Wang 0001 |
CVPR | 1 |
| 2016 | Inference With Collaborative Model for Interactive Tumor Segmentation in Medical Image SequencesabstractSegmenting organisms or tumors from medical data (e.g., computed tomography volumetric images, ultrasound, or magnetic resonance imaging images/image sequences) is one of the fundamental tasks in medical image analysis and diagnosis, and has received long-term attentions. This paper studies a novel computational framework of interactive segmentation for extracting liver tumors from image sequences, and it is suitable for different types of medical data. The main contributions are twofold. First, we propose a collaborative model to jointly formulate the tumor segmentation from two aspects: 1) region partition and 2) boundary presence. The two terms are complementary but simultaneously competing: the former extracts the tumor based on its appearance/texture information, while the latter searches for the palpable tumor boundary. Moreover, in order to adapt the data variations, we allow the model to be discriminatively trained based on both the seed pixels traced by the Lucas-Kanade algorithm and the scribbles placed by the user. Second, we present an effective inference algorithm that iterates to: 1) solve tumor segmentation using the augmented Lagrangian method and 2) propagate the segmentation across the image sequence by searching for distinctive matches between images. We keep the collaborative model updated during the inference in order to well capture the tumor variations over time. We have verified our system for segmenting liver tumors from a number of clinical data, and have achieved very promising results. The software developed with this paper can be found at http://vision.sysu.edu.cn/projects/med-interactive-seg/. Liang Lin 0004, Wei Yang 0019, Chenglong Li 0002, Jin Tang 0001, Xiaochun Cao |
IEEE Trans. Cybern. | 2 |
| 2016 | Clothes Co-Parsing Via Joint Image Segmentation and Labeling With Application to Clothing RetrievalabstractThis paper aims at developing an integrated system for clothing co-parsing (CCP), in order to jointly parse a set of clothing images (unsegmented but annotated with tags) into semantic configurations. A novel data-driven system consisting of two phases of inference is proposed. The first phase, referred as “image cosegmentation,” iterates to extract consistent regions on images and jointly refines the regions over all images by employing the exemplar-SVM technique [1]. In the second phase (i.e., “region colabeling”), we construct a multiimage graphical model by taking the segmented regions as vertices, and incorporating several contexts of clothing configuration (e.g., item locations and mutual interactions). The joint label assignment can be solved using the efficient Graph Cuts algorithm. In addition to evaluate our framework on the Fashionista dataset [2], we construct a dataset called the SYSU-Clothes dataset consisting of 2098 high-resolution street fashion photos to demonstrate the performance of our system. We achieve 90.29%/88.23% segmentation accuracy and 65.52%/63.89% recognition rate on the Fashionista and the SYSU-Clothes datasets, respectively, which are superior compared with the previous methods. Furthermore, we apply our method on a challenging task, i.e., cross-domain clothing retrieval: given user photo depicting a clothing image, retrieving the same clothing items from online shopping stores based on the fine-grained parsing results. Xiaodan Liang, Liang Lin 0004, Wei Yang 0019, Ping Luo 0002, Junshi Huang, Shuicheng Yan |
IEEE Trans. Multim. | 3 |
| 2015 | Multi-task Recurrent Neural Network for Immediacy PredictionabstractIn this paper, we propose to predict immediacy for interacting persons from still images. A complete immediacy set includes interactions, relative distance, body leaning direction and standing orientation. These measures are found to be related to the attitude, social relationship, social interaction, action, nationality, and religion of the communicators. A large-scale dataset with 10,000 images is constructed, in which all the immediacy measures and the human poses are annotated. We propose a rich set of immediacy representations that help to predict immediacy from imperfect 1-person and 2-person pose estimation results. A multi-task deep recurrent neural network is constructed to take the proposed rich immediacy representation as input and learn the complex relationship among immediacy predictions multiple steps of refinement. The effectiveness of the proposed approach is proved through extensive experiments on the large scale dataset. Xiao Chu, Wanli Ouyang, Wei Yang 0019, Xiaogang Wang 0001 |
ICCV | 3 |
| 2015 | Discriminatively Trained And-Or Graph Models for Object Shape DetectionabstractIn this paper, we investigate a novel reconfigurable part-based model, namely And-Or graph model, to recognize object shapes in images. Our proposed model consists of four layers: leaf-nodes at the bottom are local classifiers for detecting contour fragments; or-nodes above the leaf-nodes function as the switches to activate their child leaf-nodes, making the model reconfigurable during inference; and-nodes in a higher layer capture holistic shape deformations; one root-node on the top, which is also an or-node, activates one of its child and-nodes to deal with large global variations (e.g. different poses and views). We propose a novel structural optimization algorithm to discriminatively train the And-Or model from weakly annotated data. This algorithm iteratively determines the model structures (e.g. the nodes and their layouts) along with the parameter learning. On several challenging datasets, our model demonstrates the effectiveness to perform robust shape-based object detection against background clutter and outperforms the other state-of-the-art approaches. We also release a new shape database with annotations, which includes more than 1500 challenging shape instances, for recognition and detection. Liang Lin 0004, Xiaolong Wang 0004, Wei Yang 0019, Jian-Huang Lai |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | Clothing Co-parsing by Joint Image Segmentation and LabelingabstractThis paper aims at developing an integrated system of clothing co-parsing, in order to jointly parse a set of clothing images (unsegmented but annotated with tags) into semantic configurations. We propose a data-driven framework consisting of two phases of inference. The first phase, referred as "image co-segmentation", iterates to extract consistent regions on images and jointly refines the regions over all images by employing the exemplar-SVM (ESVM) technique [23]. In the second phase (i.e. "region colabeling"), we construct a multi-image graphical model by taking the segmented regions as vertices, and incorporate several contexts of clothing configuration (e.g., item location and mutual interactions). The joint label assignment can be solved using the efficient Graph Cuts algorithm. In addition to evaluate our framework on the Fashionista dataset [30], we construct a dataset called CCP consisting of 2098 high-resolution street fashion photos to demonstrate the performance of our system. We achieve 90.29% / 88.23% segmentation accuracy and 65.52% / 63.89% recognition rate on the Fashionista and the CCP datasets, respectively, which are superior compared with state-of-the-art methods. Wei Yang 0019, Ping Luo 0002, Liang Lin 0004 |
CVPR | 1 |
| 2014 | Data-driven scene understanding by adaptive exemplar retrievalabstractThis article studies a data-driven approach for semantically scene understanding, without pixelwise annotation and classifier pre-training. Our framework parses a target image with two steps: (i) retrieving its exemplars (i.e. references) from an image database, where all images are unsegmented but annotated with tags; (ii) recovering its pixel labels by propagating semantics from the references. We present a novel framework making the two steps mutually conditional and bootstrapped under the probabilistic Expectation-maximization (EM) formulation. In the first step, the references are selected by jointly matching their appearances with the target as well as the semantics. We process the second step via a combinatorial graphical representation, in which the vertices are superpixels extracted from the target and its selected references. Then we derive the potentials of assigning labels to one vertex of the target, which depends upon the graph edges that connect the vertex to its spatial neighbors of the target and to its similar vertices of the references. Two steps can be both solved analytically, and the inference is conducted in a self-driven fashion. In the experiments, we validate our approach on two public databases, and demonstrate superior performances over the state-of-the-art methods. Xionghao Liu, Wei Yang 0019, Qing Wang 0018, Liang Lin 0004, Jian-Huang Lai |
ICME | 2 |
| 2012 | Learning contour-fragment-based shape model with And-Or tree representationabstractThis paper proposes a simple yet effective method to learn the hierarchical object shape model consisting of local contour fragments, which represents a category of shapes in the form of an And-Or tree. This model extends the traditional hierarchical tree structures by introducing the “switch” variables (i.e. the or-nodes) that explicitly specify production rules to capture shape variations. We thus define the model with three layers: the leaf-nodes for detecting local contour fragments, the or-nodes specifying selection of leaf-nodes, and the root-node encoding the holistic distortion. In the training stage, for optimization of the And-Or tree learning, we extend the concave-convex procedure (CCCP) by embedding the structural clustering during the iterative learning steps. The inference of shape detection is consistent with the model optimization, which integrates the local testings via the leaf-nodes and or-nodes with the global verification via the root-node. The advantages of our approach are validated on the challenging shape databases (i.e., ETHZ and INRIA Horse) and summarized as follows. (1) The proposed method is able to accurately localize shape contours against unreliable edge detection and edge tracing. (2) The And-Or tree model enables us to well capture the intraclass variance. Liang Lin 0004, Xiaolong Wang 0004, Wei Yang 0019, Jian-Huang Lai |
CVPR | 3 |
| 2011 | Interactive CT image segmentation with online discriminative learningabstractAlthough interactive image segmentation has been widely exploited, current approaches present unsatisfactory results in medical image processing. This paper proposes a fast method for interactive CT image segmentation in which the tumor regions should be partitioned as foreground against the healthy tissues. In contrast to natural images, we have the following observation on CT images: (1) CT images often include discontinuous silhouette or cluttered spots caused by input de- vices or patient corporeity; (2) Disease areas often have varying appearance and shape. We thus train a discriminative fore- ground/background model based on user-placed scribbles. In our method, we extract positive and negative samples according to the foreground and background scribbles respectively, and use dense SIFT descriptors plus gray-level histogram as candidate features. With online learning, segmentation can be fast solved by the Bregman iteration. We test our method on CT liver images and demonstrate the advantage by comparing to state-of-the-art approaches. Wei Yang 0019, Xiaolong Wang 0004, Liang Lin 0004, Chengying Gao |
ICIP | 1 |