EDBT 2026 Demo / reviewers in the wild / expert
Yangang Wang 0001
dblp:83/9429
· DBLP profile ↗
53ranked-venue papers
7as first author
35since 2021 · last 2026
0000-0002-1325-9252ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 47 · 6 first-author · 31 since 2021Artificial intelligence and machine learning · 22 · 1 first-author · 17 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Two-stage diffusion for hands and articulated objects interaction synthesis
Wenqian Sun, Binghui Zuo, Zimeng Zhao, Yangang Wang 0001 |
Pattern Recognit. | 5 |
| 2026 | FOSUP: Dynamic Garments Diffusion With Fourier Spherical Unwrapping From Monocular VideoabstractRecent advances in dynamic garment reconstruction boost virtual try-on with monocular video streams as inputs. However, existing literature has been intensively conducted on its sub-tasks, including static garment reconstruction and dynamic clothed human reconstruction, which are difficult to extend to dynamic garment reconstruction. The former bottleneck is mainly the lack of cross-frame correspondences and independent clothing topology on implicit garment fields, which results in the inability to obtain accurate motion information during dynamic clothing reconstruction and the absence of stable topology critical for downstream tasks, such as animation with physics engines. The latter usually binds the garment motion with body or skeleton motions, leading to rigid artifacts for loose-fitting garments. Our key idea is to build a diffusion based T-pose garments generator with a strong prior on garments structure. The garment generator is trained to generate 2D clothing representation, termed FOSUP, conditioned by a monocular video. FOSUP, defined as FOurier Spherical Unwrapping, enables a bidirectional mapping between FOSUP and the mesh through FFT and inverse FFT, which maintain spatial order and adjacency. Subsequently, this FOSUP is mapped back to 3D meshes through an inverse FFT process and transformed into pose space through a point transformation network to guide the three-dimensional reconstruction of the entire sequence. To sufficiently train our framework and address the lack of domain-specific data, we have constructed a large-scale garment MoCap dataset. This dataset captures the motion of various loose garments and includes multi-view raw images, frame-by-frame human motion annotations, raw scanned point clouds, topology-independent garment templates, and garment meshes with cross-frame correspondences. Comprehensive experiments have demonstrated that our unwrapping-based representation and diffusion-based framework significantly improve the performance and robustness of dynamic garment reconstruction. Zimeng Zhao, Yangang Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | Reconstructing Close Human Interaction with Appearance and Proxemics ReasoningabstractDue to visual ambiguities and inter-person occlusions, existing human pose estimation methods cannot recover plausible close interactions from in-the-wild videos. Even state-of-the-art large foundation models (e.g., SAM) cannot accurately distinguish human semantics in such challenging scenarios. In this work, we find that human appearance can provide a straightforward cue to address these obstacles. Based on this observation, we propose a dual-branch optimization framework to reconstruct accurate interactive motions with plausible body contacts constrained by human appearances, social proxemics, and physical laws. Specifically, we first train a diffusion model to learn the human proxemic behavior and pose prior knowledge. The trained network and two optimizable tensors are then incorporated into a dual-branch optimization framework to reconstruct human motions and appearances. Several constraints based on 3D Gaussians, 2D keypoints, and mesh penetrations are also designed to assist the optimization. With the proxemics prior and diverse constraints, our method is capable of estimating accurate interactions from in-the-wild videos captured in complex environments. We further build a dataset with pseudo ground-truth interaction annotations, which may promote future research on pose estimation and human behavior understanding. Experimental results on several benchmarks demonstrate that our method outperforms existing approaches. The code and data are available at https://www.buzhenhuang.com/works/CloseApp.html. Buzhen Huang, Chen Li 0038, Chongyang Xu, Dongyue Lu, Jinnan Chen, Yangang Wang 0001, Gim Hee Lee |
CVPR | 6 |
| 2025 | Metalwork: A Synthetic Dataset and Baseline for Stereo Matching of Metal WorkpiecesabstractStereo matching has received significant progress in recent years. However, traditional methods designed for conventional scenarios are unsuitable for metal workpieces due to challenges such as metal reflections, lack of distinct textures, and small holes. To address the challenge of stereo matching of metal workpieces, we propose (1) a large-scale synthetic dataset named MetalWork, and (2) a novel stereo matching framework for metal workpieces. Our MetalWork is the first synthetic dataset specifically designed for metal workpieces, offering a more realistic simulation of industrial challenges to support the training of stereo matching frameworks better. Meanwhile, we introduce a transformer-based context network that leverages multi-scale, multi-head attention mechanisms to effectively exchange features across both spatial and channel dimensions, enabling more accurate capture of intricate structural details. Extensive experiments show that, with the aid of MetalWork, our proposed method outperforms existing baselines in stereo matching of metal workpieces and achieves superior accuracy and robustness in disparity estimation for real metal workpieces. Binghui Zuo, Yangang Wang 0001 |
ICIP | 3 |
| 2025 | Cabin-HMR: Single-view Multi-person Human Mesh Estimation in Cabin SpacesabstractSevere occlusion remains a fundamental challenge in single-view human pose estimation, particularly in cabin environments where strong distortion further exacerbates the ill-posed nature of the problem. Leveraging the abundance of 2D pose data, we can impose structured modeling of the human body to infer plausible estimates for occluded regions, thereby guiding the reconstruction of a complete human body mesh. To address this issue, we propose Cabin-HMR, a novel method for multi-person 3D pose reconstruction from a single view. Our approach effectively incorporates human structural priors derived from 2D poses and sitting postures to infer the most plausible full-body pose under occlusion. Furthermore, by integrating depth information as a corrective signal for local image patches, our method significantly mitigates the impact of camera distortion in cabin environments. To enhance generalization to diverse and complex seated postures, we construct a large-scale dataset comprising paired 2D and 3D sitting pose annotations collected from synchronized multi-view camera systems in vehicle interiors. Experimental results demonstrate that Cabin-HMR achieves robust performance across various scenarios, particularly excelling in cabin environments where occlusion and distortion are prevalent. Zhengheng Rui, Buzhen Huang, Ziyazhuo Wang, Yangang Wang 0001 |
SMC | 5 |
| 2025 | NP-Hand: Novel Perspective Hand Image Synthesis Guided by NormalsabstractSynthesizing multi-view images that are geometrically consistent with a given single-view image is one of the hot issues in AIGC in recent years. Existing methods have achieved impressive performance on objects with symmetry or rigidity, but they are inappropriate for the human hand. Because an image-captured human hand has more diverse poses and less attractive textures. In this paper, we propose NP-Hand, a framework that elegantly combines the diffusion model and generative adversarial network: The multi-step diffusion is trained to synthesize low-resolution novel perspective, while the single-step generator is exploited to further enhance synthesis quality. To maintain the consistency between inputs and synthesis, we creatively introduce normal maps into NP-Hand to guide the whole synthesizing process. Comprehensive evaluations have demonstrated that the proposed framework is superior to existing state-of-the-art models and more suitable for synthesizing hand images with faithful structures and realistic appearance details. The code will be released on our website. Binghui Zuo, Wenqian Sun, Zimeng Zhao, Yangang Wang 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | T2C: Text-guided 4D Cloth GenerationabstractIn the age of AIGC, the creation process is increasingly automated. Generating vivid characters with clothing and motions according to scripts or novels is no exception. Unfortunately, the diversity of fabric topologies, the complexity of fabric layering, and the flexibility of fabric motion make most approaches only applicable to motion generation for characters in undressing or tight-fitting clothing. This article introduces a novel approach named T2C , which employs a multi-layered clothing representation and a physics-based clothing animation paradigm to generate text-controlled Clothed 4D Humans, expanding the boundaries of the aforementioned issues. The hierarchical representation of clothing utilizes Fourier spherical mapping to define the geometric information of garments within a standard pose space, mapping it onto several 2D frequency domain subspaces. The motion of clothing in tandem with the human body is realized through a hybrid forward dynamic solution, where the internal virtual mechanic’s parameters driving the clothing are learned from text features. A series of qualitative and quantitative experiments reveal that T2C can generate dynamic clothing with a sense of layering, realistic details, and rich textures. The code will be publicly available at https://zhipengyu28.github.io/t2c/ . Zimeng Zhao, Yanxi Du, Yuzhou Zheng, Binghui Zuo, Yangang Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2025 | GraspDiff: Grasping Generation for Hand-Object Interaction With Multimodal Guided DiffusionabstractGrasping generation holds significant importance in both robotics and AI-generated content. While pure network paradigms based on VAEs or GANs ensure diversity in outcomes, they often fall short of achieving plausibility. Additionally, although those two-step paradigms that first predict contact and then optimize distance yield plausible results, they are always known to be time-consuming. This paper introduces a novel paradigm powered by DDPM, accommodating diverse modalities with varying interaction granularities as its generating conditions, including 3D object, contact affordance, and image content. Our key idea is that the iterative steps inherent to diffusion models can supplant the iterative optimization routines in existing optimization methods, thereby endowing the generated results from our method with both diversity and plausibility. Using the same training data, our paradigm achieves superior generation performance and competitive generation speed compared to optimization-based paradigms. Extensive experiments on both in-domain and out-of-domain objects demonstrate that our method receives significant improvement over the SOTA method. We will release the code for research purposes. Binghui Zuo, Zimeng Zhao, Wenqian Sun, Yangang Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | Synthesizing Physically Plausible Human Motions in 3D ScenesabstractWe present a physics-based character control framework for synthesizing human-scene interactions. Recent advances adopt physics simulation to mitigate artifacts produced by data-driven kinematic approaches. However, existing physics-based methods mainly focus on single-object environments, resulting in limited applicability in realistic 3D scenes with multi-objects. To address such challenges, we propose a framework that enables physically simulated characters to perform long-term interaction tasks in diverse, cluttered, and unseen 3D scenes. The key idea is to decouple human-scene interactions into two fundamental processes, Interacting and Navigating, which motivates us to construct two reusable Controllers, namely InterCon and NavCon. Specifically, InterCon uses two complementary policies to enable characters to enter or leave the interacting state with a particular object (e.g., sitting on a chair or getting up). To realize navigation in cluttered environments, we introduce NavCon, where a trajectory following policy enables characters to track pre-planned collision-free paths. Benefiting from the divide and conquer strategy, we can train all policies in simple environments and directly apply them in complex multi-object scenes through coordination from a rule-based scheduler. Video and code are available at https://liangpan99.github.io/InterScene/. Liang Pan, Buzhen Huang, Yangang Wang 0001 |
3DV | 7 |
| 2024 | TexDC: Text-Driven Disease-Aware 4D Cardiac Cine MRI Images Generation
Yangang Wang 0001 |
ACCV (2) | 4 |
| 2024 | Closely Interactive Human Reconstruction with Proxemics and Physics-Guided AdaptionabstractExisting multi-person human reconstruction approaches mainly focus on recovering accurate poses or avoiding penetration, but overlook the modeling of close interactions. In this work, we tackle the task of reconstructing closely interactive humans from a monocular video. The main challenge of this task comes from insufficient visual information caused by depth ambiguity and severe inter-person occlusion. In view of this, we propose to leverage knowledge from proxemic behavior and physics to compensate the lack of visual information. This is based on the observation that human interaction has specific patterns following the social proxemics. Specifically, we first design a latent representation based on Vector Quantised-Variational AutoEncoder (VQ-VAE) to model human interaction. A proxemics and physics guided diffusion model is then introduced to denoise the initial distribution. We design the diffusion model as dual branch with each branch representing one individual such that the interaction can be modeled via cross attention. With the learned priors of VQ-VAE and physical constraint as the additional information, our proposed approach is capable of estimating accurate poses that are also proxemics and physics plausible. Experimental results on Hi4D, 3DPW, and CHI3D demonstrate that our method outperforms existing approaches. The code is available at https://github.com/boycehbz/HumanInteraction. Buzhen Huang, Chen Li 0038, Chongyang Xu, Liang Pan, Yangang Wang 0001, Gim Hee Lee |
CVPR | 5 |
| 2024 | Music Conditioned Generation for Human-Centric VideoabstractMusic and human-centric video are two fundamental signals across languages. Correlation analysis between the two is currently used in choreography and film accompaniment. This letter explores this correlation in a new task: human-centric video generation from a start-end image pair and transitional music. Existing human-centric generation methods are not competent for this task because they require frame-wise pose as input or have difficulty handling long-duration videos. Our key idea is to build a temporal generation framework dominated by DDPM and assisted by VAE and GAN. It reduces the computational cost of music-image diffusion by utilizing the latent space compactness of VAE and the image translation efficiency of GAN. To produce videos with both long duration and high quality, our framework first generates small-scale keyframes and then generates high-resolution videos. To strengthen the frame-wise consistency of the human body, a frame-aligned correspondence map is adopted as an intermediate supervision. Extensive experiments compared with the SOTA method have demonstrated the rationality and effectiveness of this signal generation framework. Zimeng Zhao, Binghui Zuo, Yangang Wang 0001 |
IEEE Signal Process. Lett. | 3 |
| 2024 | Simultaneously Recovering Multi-Person Meshes and Multi-View Cameras With Human SemanticsabstractDynamic multi-person mesh recovery has broad applications in sports broadcasting, virtual reality, and video games. However, current multi-view frameworks rely on a time-consuming camera calibration procedure. In this work, we focus on multi-person motion capture with uncalibrated cameras, which mainly faces two challenges: one is that inter-person interactions and occlusions introduce inherent ambiguities for both camera calibration and motion capture; the other is that a lack of dense correspondences can be used to constrain sparse camera geometries in a dynamic multi-person scene. Our key idea is to incorporate motion prior knowledge to simultaneously estimate camera parameters and human meshes from noisy human semantics. We first utilize human information from 2D images to initialize intrinsic and extrinsic parameters. Thus, the approach does not rely on any other calibration tools or background features. Then, a pose-geometry consistency is introduced to associate the detected humans from different views. Finally, a latent motion prior is proposed to refine the camera parameters and human motions. Experimental results show that accurate camera parameters and human motions can be obtained through a one-step reconstruction. The code are publicly available at https://github.com/boycehbz/DMMR. Buzhen Huang, Jingyi Ju, Yuan Shu, Yangang Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Skeleton Extraction for Articulated Objects With the Spherical Unwrapping ProfilesabstractEmbedding unified skeletons into unregistered scans is fundamental to finding correspondences, depicting motions, and capturing underlying structures among the articulated objects in the same category. Some existing approaches rely on laborious registration to adapt a predefined LBS model to each input, while others require the input to be set to a canonical pose, e.g., T-pose or A-pose. However, their effectiveness is always influenced by the water-tightness, face topology, and vertex density of the input mesh. At the core of our approach lies a novel unwrapping method, named SUPPLE (Spherical UnwraPping ProfiLEs), which maps a surface into image planes independent of mesh topologies. Based on this lower-dimensional representation, a learning-based framework is further designed to localize and connect skeletal joints with fully convolutional architectures. Experiments demonstrate that our framework yields reliable skeleton extractions across a broad range of articulated categories, from raw scans to online CADs. Zimeng Zhao, Wei Xie 0012, Binghui Zuo, Yangang Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | Semi-Supervised Hand Appearance Recovery via Structure Disentanglement and Dual Adversarial DiscriminationabstractEnormous hand images with reliable annotations are collected through marker-based MoCap. Unfortunately, degradations caused by markers limit their application in hand appearance reconstruction. A clear appearance recovery insight is an image-to-image translation trained with unpaired data. However, most frameworks fail because there exists structure inconsistency from a degraded hand to a bare one. The core of our approach is to first disentangle the bare hand structure from those degraded images and then wrap the appearance to this structure with a dual adversarial discrimination (DAD) scheme. Both modules take full advantage of the semi-supervised learning paradigm: The structure disentanglement benefits from the modeling ability of ViT, and the translator is enhanced by the dual discrimination on both translation processes and translation results. Comprehensive evaluations have been conducted to prove that our framework can robustly recover photo-realistic hand appearance from diverse marker-contained and even object-occluded datasets. It provides a novel avenue to acquire bare hand appearance data for other down-stream learning problems. Zimeng Zhao, Binghui Zuo, Zhiyu Long, Yangang Wang 0001 |
CVPR | 4 |
| 2023 | Reconstructing Groups of People with Hypergraph Relational ReasoningabstractDue to the mutual occlusion, severe scale variation, and complex spatial distribution, the current multi-person mesh recovery methods cannot produce accurate absolute body poses and shapes in large-scale crowded scenes. To address the obstacles, we fully exploit crowd features for reconstructing groups of people from a monocular image. A novel hypergraph relational reasoning network is proposed to formulate the complex and high-order relation correlations among individuals and groups in the crowd. We first extract compact human features and location information from the original high-resolution image. By conducting the relational reasoning on the extracted individual features, the underlying crowd collectiveness and interaction relationship can provide additional group information for the reconstruction. Finally, the updated individual features and the localization information are used to regress human meshes in camera coordinates. To facilitate the network training, we further build pseudo ground-truth on two crowd datasets, which may also promote future research on pose estimation and human behavior understanding in crowded scenes. The experimental results show that our approach outperforms other baseline methods both in crowded and common scenarios. The code and datasets are publicly available at https://github.com/boycehbz/GroupRec. Buzhen Huang, Jingyi Ju, Zhihao Li 0002, Yangang Wang 0001 |
ICCV | 4 |
| 2023 | Nonrigid Object Contact Estimation With Regional Unwrapping TransformerabstractAcquiring contact patterns between hands and nonrigid objects is a common concern in the vision and robotics community. However, existing learning-based methods focus more on contact with rigid ones from monocular images. When adopting them for nonrigid contact, a major problem is that the existing contact representation is restricted by the geometry of the object. Consequently, contact neighborhoods are stored in an unordered manner and contact features are difficult to align with image cues. At the core of our approach lies a novel hand-object contact representation called RUPs (Region Unwrapping Profiles), which unwrap the roughly estimated hand-object surfaces as multiple high-resolution 2D regional profiles. The region grouping strategy is consistent with the hand kinematic bone division because they are the primitive initiators for a composite contact pattern. Based on this representation, our Regional Unwrapping Transformer (RUFormer) learns the correlation priors across regions from monocular inputs and predicts corresponding contact and deformed transformations. Our experiments demonstrate that the proposed framework can robustly estimate the deformed degrees and deformed transformations, which makes it suitable for both nonrigid and rigid contact. Wei Xie 0012, Zimeng Zhao, Binghui Zuo, Yangang Wang 0001 |
ICCV | 5 |
| 2023 | 4D Myocardium Reconstruction with Decoupled Motion and Shape ModelabstractEstimating the shape and motion state of the myocardium is essential in diagnosing cardiovascular diseases. However, cine magnetic resonance (CMR) imaging is dominated by 2D slices, whose large slice spacing challenges inter-slice shape reconstruction and motion acquisition. To address this problem, we propose a 4D reconstruction method that decouples motion and shape, which can predict the inter-/intra- shape and motion estimation from a given sparse point cloud sequence obtained from limited slices. Our framework comprises a neural motion model and an end-diastolic (ED) shape model. The implicit ED shape model can learn a continuous boundary and encourage the motion model to predict without the supervision of ground truth deformation, and the motion model enables canonical input of the shape model by deforming any point from any phase to the ED phase. Additionally, the constructed ED-space enables pre-training of the shape model, thereby guiding the motion model and addressing the issue of data scarcity. We propose the first 4D myocardial dataset as we know and verify our method on the proposed, public, and cross-modal datasets, showing superior reconstruction performance and enabling various clinical applications. Yangang Wang 0001 |
ICCV | 3 |
| 2023 | Reconstructing Interacting Hands with Interaction Prior from Monocular ImagesabstractReconstructing interacting hands from monocular images is indispensable in AR/VR applications. Most existing solutions rely on the accurate localization of each skeleton joint. However, these methods tend to be unreliable due to the severe occlusion and confusing similarity among adjacent hand parts. This also defies human perception because humans can quickly imitate an interaction pattern without localizing all joints. Our key idea is to first construct a two-hand interaction prior and recast the interaction reconstruction task as the conditional sampling from the prior. To expand more interaction states, a large-scale multimodal dataset with physical plausibility is proposed. Then a VAE is trained to further condense these interaction patterns as latent codes in a prior distribution. When looking for image cues that contribute to interaction prior sampling, we propose the interaction adjacency heatmap (IAH). Compared with a joint-wise heatmap for localization, IAH assigns denser visible features to those invisible joints. Compared with an all-in-one visible heatmap, it provides more fine-grained local interaction information in each interaction region. Finally, the correlations between the extracted features and corresponding interaction codes are linked by the ViT module. Comprehensive evaluations on benchmark datasets have verified the effectiveness of this framework. The code and dataset are publicly available at https://github.com/binghui-z/InterPrior_pytorch. Binghui Zuo, Zimeng Zhao, Wenqian Sun, Wei Xie 0012, Zhou Xue, Yangang Wang 0001 |
ICCV | 6 |
| 2023 | Implicit Representation for Interacting Hands Reconstruction from Monocular Color Images
Binghui Zuo, Zimeng Zhao, Wei Xie 0012, Yangang Wang 0001 |
ICIG (1) | 4 |
| 2023 | Physics-Guided Human Motion Capture with Pose Probability ModelingabstractIncorporating physics in human motion capture to avoid artifacts like floating, foot sliding, and ground penetration is a promising direction. Existing solutions always adopt kinematic results as reference motions, and the physics is treated as a post-processing module. However, due to the depth ambiguity, monocular motion capture inevitably suffers from noises, and the noisy reference often leads to failure for physics-based tracking. To address the obstacles, our key-idea is to employ physics as denoising guidance in the reverse diffusion process to reconstruct physically plausible human motion from a modeled pose probability distribution. Specifically, we first train a latent gaussian model that encodes the uncertainty of 2D-to-3D lifting to facilitate reverse diffusion. Then, a physics module is constructed to track the motion sampled from the distribution. The discrepancies between the tracked motion and image observation are used to provide explicit guidance for the reverse diffusion model to refine the motion. With several iterations, the physics-based tracking and kinematic denoising promote each other to generate a physically plausible human motion. Experimental results show that our method outperforms previous physics-based methods in both joint accuracy and success rate. More information can be found at https://github.com/Me-Ditto/Physics-Guided-Mocap. Jingyi Ju, Buzhen Huang, Zhihao Li 0002, Yangang Wang 0001 |
IJCAI | 5 |
| 2023 | HMDO : Markerless multi-view hand manipulation capture with deformable objectsabstractWe construct the first markerless deformable interaction dataset recording interactive motions of the hands and deformable objects, called HMDO (Hand Manipulation with Deformable Objects). With our built multi-view capture system, it captures the deformable interactions with multiple perspectives, various object shapes, and diverse interactive forms. Our motivation is the current lack of hand and deformable object interaction datasets, as 3D hand and deformable object reconstruction is challenging. Mainly due to mutual occlusion, the interaction area is difficult to observe, the visual features between the hand and the object are entangled, and the reconstruction of the interaction area deformation is difficult. To tackle this challenge, we propose a method to annotate our captured data. Our key idea is to collaborate with estimated hand features to guide the object global pose estimation, and then optimize the deformation process of the object by analyzing the relationship between the hand and the object. Through comprehensive evaluation, the proposed method can reconstruct interactive motions of hands and deformable objects with high quality. HMDO currently consists of 21600 frames over 12 sequences. In the future, this dataset could boost the research of learning-based reconstruction of deformable interaction scenes. Wei Xie 0012, Zimeng Zhao, Binghui Zuo, Yangang Wang 0001 |
Graph. Model. | 5 |
| 2023 | A causal convolutional neural network for multi-subject motion modeling and generationabstractInspired by the success of WaveNet in multi-subject speech synthesis, we propose a novel neural network based on causal convolutions for multi-subject motion modeling and generation. The network can capture the intrinsic characteristics of the motion of different subjects, such as the influence of skeleton scale variation on motion style. Moreover, after fine-tuning the network using a small motion dataset for a novel skeleton that is not included in the training dataset, it is able to synthesize high-quality motions with a personalized style for the novel skeleton. The experimental results demonstrate that our network can model the intrinsic characteristics of motions well and can be applied to various motion modeling and synthesis tasks. Shuaiying Hou, Congyi Wang, Wenlin Zhuang, Yangang Wang 0001, Hujun Bao, Jinxiang Chai, Weiwei Xu 0003 |
Comput. Vis. Media | 5 |
| 2023 | Object-Occluded Human Shape and Pose Estimation With Probabilistic Latent ConsistencyabstractOcclusions between human and objects, especially for the activities of human-object interactions, are very common in practical applications. However, most of the existing approaches for 3D human shape and pose estimation require that human bodies are well captured without occlusions or with minor self-occlusions. In this paper, we focus on the problem of directly estimating the object-occluded human shape and pose from single color images. Our key idea is to utilize a partial UV map to represent an object-occluded human body, and the full 3D human shape estimation is ultimately converted as an image inpainting problem. We propose a novel two-branch network architecture to train an end-to-end regressor via a latent distribution consistency, which also includes a novel visible feature sub-net to extract the human information from object-occluded color images. To supervise the network training, we further build a novel dataset named as 3DOH50K. Several experiments are conducted to reveal the effectiveness of the proposed method. Experimental results demonstrate that the proposed method achieves state-of-the-art compared with previous methods. The dataset and codes are publicly available at https://www.yangangwang.com/papers/ZHANG-OOH-2020-03.html. Buzhen Huang, Yangang Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | DeepCloth: Neural Garment Representation for Shape and Style EditingabstractGarment representation, editing and animation are challenging topics in the area of computer vision and graphics. It remains difficult for existing garment representations to achieve smooth and plausible transitions between different shapes and topologies. In this work, we introduce, DeepCloth, a unified framework for garment representation, reconstruction, animation and editing. Our unified framework contains 3 components: First, we represent the garment geometry with a "topology-aware UV-position map", which allows for the unified description of various garments with different shapes and topologies by introducing an additional topology-aware UV-mask for the UV-position map. Second, to further enable garment reconstruction and editing, we contribute a method to embed the UV-based representations into a continuous feature space, which enables garment shape reconstruction and editing by optimization and control in the latent space, respectively. Finally, we propose a garment animation method by unifying our neural garment representation with body shape and pose, which achieves plausible garment animation results leveraging the dynamic information encoded by our shape and style representation, even under drastic garment editing operations. To conclude, with DeepCloth, we move a step forward in establishing a more flexible and general 3D garment digitization framework. Experiments demonstrate that our method can achieve state-of-the-art garment representation performance compared with previous methods. Zhaoqi Su, Tao Yu 0007, Yangang Wang 0001, Yebin Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Slice-Mask Based 3D Cardiac Shape Reconstruction from CT Volume
Fu Feng, Yinsu Zhu, Yangang Wang 0001 |
ACCV (6) | 5 |
| 2022 | Neural MoCon: Neural Motion Control for Physically Plausible Human Motion CaptureabstractDue to the visual ambiguity, purely kinematic formulations on monocular human motion capture are often physically incorrect, biomechanically implausible, and can not reconstruct accurate interactions. In this work, we focus on exploiting the high-precision and non-differentiable physics simulator to incorporate dynamical constraints in motion capture. Our key-idea is to use real physical supervisions to train a target pose distribution prior for sampling-based motion control to capture physically plausible human motion. To obtain accurate reference motion with terrain interactions for the sampling, we first introduce an interaction constraint based on SDF (Signed Distance Field) to enforce appropriate ground contact modeling. We then design a novel two-branch decoder to avoid stochastic error from pseudo ground-truth and train a distribution prior with the non-differentiable physics simulator. Finally, we regress the sampling distribution from the current state of the physical character with the trained prior and sample satisfied target poses to track the estimated reference motion. Qualitative and quantitative results show that we can obtain physically plausible human motion with complex terrain interactions, human shape variations, and diverse behaviors. More information can be found ar https://www.yangangwang.com/papers/HBZ-NM-2022-03.html Buzhen Huang, Liang Pan, Jingyi Ju, Yangang Wang 0001 |
CVPR | 5 |
| 2022 | Stability-driven Contact Reconstruction From Monocular Color ImagesabstractPhysical contact provides additional constraints for hand-object state reconstruction as well as a basis for further understanding of interaction affordances. Estimating these severely occluded regions from monocular images presents a considerable challenge. Existing methods optimize the hand-object contact driven by distance threshold or prior from contact-labeled datasets. However, due to the number of subjects and objects involved in these indoor datasets being limited, the learned contact patterns could not be generalized easily. Our key idea is to reconstruct the contact pattern directly from monocular images, and then utilize the physical stability criterion in the simulation to optimize it. This criterion is defined by the resultant forces and contact distribution computed by the physics engine. Compared to existing solutions, our framework can be adapted to more personalized hands and diverse object shapes. Furthermore, an interaction dataset with extra physical attributes is created to verify the sim-to-real consistency of our methods. Through comprehensive evaluations, hand-object contact can be reconstructed with both accuracy and stability by the proposed framework. Zimeng Zhao, Binghui Zuo, Wei Xie 0012, Yangang Wang 0001 |
CVPR | 4 |
| 2022 | Pose2UV: Single-Shot Multiperson Mesh Recovery With Deep UV PriorabstractIn this work, we focus on the task of multi-person mesh recovery from a single color image, where the key issue is to tackle the pixel-level ambiguities caused by inter-person occlusions. Overall, there are two main technical challenges when addressing the ambiguities: how to extract valid target features under occlusions and how to reconstruct reasonable human meshes with only a handful of body cues? To deal with these problems, our key idea is to utilize the predicted 2D poses to locate and separate the target person, and reconstruct them with a novel learning-based UV prior. Specifically, we propose a visible pose-mask module to help extract valid target features, then train a dense body mesh prior to promote reconstructing natural mesh represented by the UV position map. To evaluate the performance of our proposed method under occlusions, we further build an in-the-wild 3D multi-person benchmark named as 3DMPB. Experimental results demonstrate that our method achieves state-of-the-art compared with previous methods. The dataset, codes are publicly available on our website. Buzhen Huang, Yangang Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | Music2Dance: DanceNet for Music-Driven Dance GenerationabstractSynthesize human motions from music (i.e., music to dance) is appealing and has attracted lots of research interests in recent years. It is challenging because of the requirement for realistic and complex human motions for dance, but more importantly, the synthesized motions should be consistent with the style, rhythm, and melody of the music. In this article, we propose a novel autoregressive generative model, DanceNet, to take the style, rhythm, and melody of music as the control signals to generate 3D dance motions with high realism and diversity. Due to the high long-term spatio-temporal complexity of dance, we propose the dilated convolution to improve the receptive field, and adopt the gated activation unit as well as separable convolution to enhance the fusion of motion features and control signals. To boost the performance of our proposed model, we capture several synchronized music-dance pairs by professional dancers and build a high-quality music-dance pair dataset. Experiments have demonstrated that the proposed method can achieve state-of-the-art results. Wenlin Zhuang, Congyi Wang, Jinxiang Chai, Yangang Wang 0001, Ming Shao, Si-Yu Xia |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2021 | Dynamic Multi-Person Mesh Recovery From Uncalibrated Multi-View CamerasabstractDynamic multi-person mesh recovery has been a hot topic in 3D vision recently. However, few works focus on the multi-person motion capture from uncalibrated cameras, which mainly faces two challenges: the one is that inter-person interactions and occlusions introduce inherent ambiguities for both camera calibration and motion capture; The other is that a lack of dense correspondences can be used to constrain sparse camera geometries in a dynamic multi-person scene. Our key idea is incorporating motion prior knowledge into simultaneous optimization of extrinsic camera parameters and human meshes from noisy human semantics. First, we introduce a physics-geometry consistency to reduce the low and high frequency noises of the detected human semantics. Then a novel latent motion prior is proposed to simultaneously optimize extrinsic camera parameters and coherent human motions from slightly noisy inputs. Experimental results show that accurate camera parameters and human motions can be obtained through one-stage optimization. The codes will be publicly available at https://www.yangangwang.com. Buzhen Huang, Yuan Shu, Yangang Wang 0001 |
3DV | 4 |
| 2021 | SUPPLE: Extracting Hand Skeleton with Spherical Unwrapping ProfilesabstractEmbedding a unified skeleton into diverse hand meshes is a prominent task both for animation and pose estimation. Most existing methods extracted skeletons from humanoid characters under simple poses, e.g. T- pose or A- pose. Applying them directly to hand meshes may yield inaccurate or implausible results because hands have higher dexterity and similar endpoints. Furthermore, these methods did not attempt to extract skeleton directly from a scan model which may be not watertight and has much more vertices. Our key idea is to unwrap meshes with different topologies in the same image-based representation, named SUPPLE (Spherical UnwraPping ProfiLEs), and then train a convolutional encoder-decoder to extract skeleton under this representation. Experiments demonstrate that our framework produces reliable and accurate skeleton estimation results across a broad range of datasets, from raw scans to artist-designed models. Zimeng Zhao, Ruting Rao, Yangang Wang 0001 |
3DV | 3 |
| 2021 | Interacting Two-Hand 3D Pose and Shape Reconstruction from Single Color ImageabstractIn this paper, we propose a novel deep learning framework to reconstruct 3D hand poses and shapes of two interacting hands from a single color image. Previous methods designed for single hand cannot be easily applied for the two hand scenario because of the heavy inter-hand occlusion and larger solution space. In order to address the occlusion and similar appearance between hands that may confuse the network, we design a hand pose-aware attention module to extract features associated to each individual hand respectively. We then leverage the two hand context presented in interaction to propose a context-aware cascaded refinement that improves the hand pose and shape accuracy of each hand conditioned on the context between interacting hands. Extensive experiments on the main benchmark datasets demonstrate that our method predicts accurate 3D hand pose and shape from single color image, and achieves the state-of-the-art performance. Code is available in project webpage https://baowenz.github.io/Intershape/. Baowen Zhang, Yangang Wang 0001, Xiaoming Deng 0001, Yinda Zhang 0001, Ping Tan 0002, CuiXia Ma, Hongan Wang |
ICCV | 2 |
| 2021 | TravelNet: Self-supervised Physically Plausible Hand Motion Learning from Monocular Color ImagesabstractThis paper aims to reconstruct physically plausible hand motion from monocular color images. Existing frame-by-frame estimating approaches can not guarantee the physical plausibility (e.g. penetration, jittering) directly. In this paper, we embed physical constraints on the per-frame estimated motions in both spatial and temporal space. Our key idea is to adopt a self-supervised learning strategy to train a novel encoder-decoder, named TravelNet, whose training motion data is prepared by the physics engine using discrete pose states. TravelNet captures key pose states from hand motion sequences as compact motion descriptors, inspired by the concept of keyframes in animation. Finally, it manages to extract those key states out of perturbations without manual annotations, and reconstruct the motions preserving details and physical plausibility. In the experiments, we show that the outputs of the TravelNet contain both finger synergism and time consistency. Through the proposed framework, hand motions can be accurately reconstructed and flexibly re-edited, which is superior to the state-of-the-art methods. Zimeng Zhao, Yangang Wang 0001 |
ICCV | 3 |
| 2021 | Hair Salon: A Geometric Example-Based Method to Generate 3D Hair Data
Qiaomu Ren, Haikun Wei, Yangang Wang 0001 |
ICIG (3) | 3 |
| 2020 | Object-Occluded Human Shape and Pose Estimation From a Single Color ImageabstractOcclusions between human and objects, especially for the activities of human-object interactions, are very common in practical applications. However, most of the existing approaches for 3D human shape and pose estimation require human bodies are well captured without occlusions or with minor self-occlusions. In this paper, we focus on the problem of directly estimating the object-occluded human shape and pose from single color images. Our key idea is to utilize a partial UV map to represent an object-occluded human body, and the full 3D human shape estimation is ultimately converted as an image inpainting problem. We propose a novel two-branch network architecture to train an end-to-end regressor via the latent feature supervision, which also includes a novel saliency map sub-net to extract the human information from object-occluded color images. To supervise the network training, we further build a novel dataset named as 3DOH50K. Several experiments are conducted to reveal the effectiveness of the proposed method. Experimental results demonstrate that the proposed method achieves the state-of-the-art comparing with previous methods. The dataset, codes are publicly available at https://www.yangangwang.com. Buzhen Huang, Yangang Wang 0001 |
CVPR | 3 |
| 2020 | Hand-3d-Studio: A New Multi-View System for 3d Hand ReconstructionabstractThis paper proposes a new system named as Hand-3D-Studio to capture the 3D hand pose and shape information. Our system includes 15 synchronized DSLR cameras, which can acquire high quality multi-view 4K resolution color images in a circular manner. We then introduce a 2D hand keypoints guided iterative pixel growth matching strategy for 3D reconstruction, where the 2D keypoints are obtained via convolution neural network. We find that the pre-detected 2D hand keypoints can greatly remove the matching noise, and thus improve the performance of reconstruction. After that, a non-rigid iterative closest points algorithm is performed to drive a template hand to fit the point clouds and register all the hand meshes. As a consequence, we captured more than 20K high quality hand color images, annotated 2D hand key-points, 3D point cloud as well as the registered hand meshes (>200). All the data are public on the website http://www.yangangwang.com for future research. Tianyao Wang, Si-Yu Xia, Yangang Wang 0001 |
ICASSP | 4 |
| 2020 | Personalized Hand Modeling from Multiple Postures with Multi-View Color ImagesabstractAbstract Personalized hand models can be utilized to synthesize high quality hand datasets, provide more possible training data for deep learning and improve the accuracy of hand pose estimation. In recent years, parameterized hand models, e.g., MANO, are widely used for obtaining personalized hand models. However, due to the low resolution of existing parameterized hand models, it is still hard to obtain high‐fidelity personalized hand models. In this paper, we propose a new method to estimate personalized hand models from multiple hand postures with multi‐view color images. The personalized hand model is represented by a personalized neutral hand, and multiple hand postures. We propose a novel optimization strategy to estimate the neutral hand from multiple hand postures. To demonstrate the performance of our method, we have built a multi‐view system and captured more than 35 people, and each of them has 30 hand postures. We hope the estimated hand models can boost the research of high‐fidelity parameterized hand modeling in the future. All the hand models are publicly available on www.yangangwang.com . Yangang Wang 0001, Ruting Rao, Changqing Zou |
Comput. Graph. Forum | 1 |
| 2020 | A Three-Switch-Based Single-Input Dual-Output Converter With Simultaneous Boost & Buck Voltage ConversionabstractWith the increasing demand of applications that have two different output voltages, the single-input dual-output (SIDO) converter with fewer components is becoming the cost-effective option instead of employing two single-input single-output converters. However, the cross-regulation of different outputs is still a challenge in SIDO converters design. To save cost and obtain improved cross-regulation performance, a novel SIDO converter consisting of three active switches is proposed in this article. Owing to the use of a voltage multiplier circuit, a high step-up voltage conversion ratio is achieved with relatively low voltage stress of switches. Thanks to the independent power flow and two control variables, the cross-regulation performance is improved, and simultaneous buck as well as boost output voltages is realized. Additionally, all switches can achieve zero-voltage switching operation, which contributes to a significant switching loss reduction. The operation characteristics, design considerations, and control strategy of the proposed converter are analyzed. To verify the theoretical analysis and measure the power efficiency, a prototype circuit with 48 V input and 24 V/2 A, 200 V/1 A outputs is built. Sen Song, Guipeng Chen, Yihua Hu 0004, Kai Ni 0003, Yangang Wang 0001 |
IEEE Trans. Ind. Informatics | 6 |
| 2020 | SRHandNet: Real-Time 2D Hand Pose Estimation With Simultaneous Region LocalizationabstractThis paper introduces a novel method for real-time 2D hand pose estimation from monocular color images, which is named as SRHandNet. Existing methods can not time efficiently obtain appropriate results for small hand. Our key idea is to simultaneously regress the hand region of interests (RoIs) and hand keypoints for a given color image, and iteratively take the hand RoIs as feedback information for boosting the performance of hand keypoints estimation with a single encoder-decoder network architecture. Different from previous region proposal network (RPN), a new lightweight bounding box representation, which is called region map, is proposed. The proposed bounding box representation map together with hand keypoints heatmaps are combined into the unified multi-channel feature maps, which can be easily acquired with only one forward network inference and thus improve the runtime efficiency of the network. Our proposed SRHandNet can run at 40fps for hand bounding box detection and up to 30fps accurate hand keypoints estimation under the desktop environment without implementation optimization. Experiments demonstrate the effectiveness of the proposed method. State-of-the-art results are also achieved out competing all recent methods. Yangang Wang 0001, Baowen Zhang, Cong Peng 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Neural Hand Reconstruction Using A Single RGB ImageabstractWe present a neural hand reconstruction method for monocular 3D hand pose and shape estimation in this paper. Instead of directly representing hand with 3D data, a novel UV position map is introduced to represent hand pose and shape with 2D data, which maps 3D hand surface points to 2D image space. Furthermore, an encoder-decoder neural network is proposed to infer such UV position map from only single image. To train such network with the lack of ground truth training pairs, we propose a novel MANOReg module which employs MANO model as shape prior to constrain high-dimensional space of UV position map. Both quantitative and qualitative experiments demonstrate the effectiveness of our UV position map representation and MANOReg module. Mengcheng Li, Liang An 0001, Tao Yu 0007, Yangang Wang 0001, Feng Chen 0007, Yebin Liu |
Virtual Real. Intell. Hardw. | 4 |
| 2019 | Mask-Pose Cascaded CNN for 2D Hand Pose Estimation From Single Color ImageabstractWe present a cascaded convolutional neural network for 2D hand pose estimation from single in-the-wild RGB images. Inspired by the commonly used silhouette information in the generative pose estimation approaches, we build the cascaded network with two stages, including mask prediction stage as well as pose estimation stage. We find that the two stages network architecture for end-to-end training could benefit from each other for detecting the hand mask and 2D pose. To further improve the hand pose detection accuracy, we contribute a new RGB hand dataset named OneHand10K, which contains 10K RGB images. Each image contains one single hand. We manually obtain the segmented mask and labeled keypoints for guided learning. We hope that this dataset will be a benchmark and encourage more people to conduct research on this challenging topic. Experiments on the validation dataset have demonstrated the superior performance of the proposed cascaded convolutional neural network. Yangang Wang 0001, Cong Peng 0001, Yebin Liu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Multi-scale Adaptive Structure Network for Human Pose Estimation from Color Images
Wenlin Zhuang, Cong Peng 0001, Si-Yu Xia, Yangang Wang 0001 |
ACCV (1) | 4 |
| 2018 | Shape and Pose Estimation for Closely Interacting Persons Using Multi-view ImagesabstractAbstract Multi‐person pose and shape estimation is very challenging, especially when the persons have close interactions. Existing methods only work well when people are well spaced out in the captured images. However, close interaction among people is very common in real life, which is more challenge due to complex articulation, frequent occlusion and inherent ambiguities. We present a fully‐automatic markerless motion capture method to simultaneously estimate 3D poses and shapes of closely interacting people from multi‐view sequences. We first predict the 2D joints for each person in an image, and then design a spatio‐temporal tracker for multi‐person pose tracking based on multi‐view videos. Finally, we estimate 3D poses and shapes of all the persons with multi‐view constraints using a skinned multi‐person linear model (SMPL). Experimental results demonstrate that our method achieves fast but accurate pose and shape estimation results for multi‐person close interaction cases. Compared with existing methods, our method does not need pre‐segmentation for each person and manual intervention, which greatly reduces the complexity of the system including time complexity and system processing complexity. Kun Li 0001, Nianhong Jiao, Yebin Liu, Yangang Wang 0001, Jing-Yu Yang 0002 |
Comput. Graph. Forum | 4 |
| 2018 | Robust Non-Rigid Motion Tracking and Surface Reconstruction Using L0 RegularizationabstractWe present a new motion tracking technique to robustly reconstruct non-rigid geometries and motions from a single view depth input recorded by a consumer depth sensor. The idea is based on the observation that most non-rigid motions (especially human-related motions) are intrinsically involved in articulate motion subspace. To take this advantage, we propose a novel based motion regularizer with an iterative solver that implicitly constrains local deformations with articulate structures, leading to reduced solution space and physical plausible deformations. The strategy is integrated into the available non-rigid motion tracking pipeline, and gradually extracts articulate joints information online with the tracking, which corrects the tracking errors in the results. The information of the articulate joints is used in the following tracking procedure to further improve the tracking accuracy and prevent tracking failures. Extensive experiments over complex human body motions with occlusions, facial and hand motions demonstrate that our approach substantially improves the robustness and accuracy in motion tracking. Feng Xu 0005, Yangang Wang 0001, Yebin Liu, Qionghai Dai |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2018 | Errata to "Robust Non-Rigid Motion Tracking and Surface Reconstruction Using L0 Regularization"abstractPresents corrections to grant number information from the paper, “Robust non-rigid motion tracking and surface reconstruction using L0 regularization,” (Guo, K., et al), IEEE Trans. Vis. Comput. Graph., vol. 24, no. 5, pp. 1770–1783, May 2018. Feng Xu 0005, Yangang Wang 0001, Yebin Liu, Qionghai Dai |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2018 | Outdoor Markerless Motion Capture with Sparse Handheld Video CamerasabstractWe present a method for outdoor markerless motion capture with sparse handheld video cameras. In the simplest setting, it only involves two mobile phone cameras following the character. This setup can maximize the flexibilities of data capture and broaden the applications of motion capture. To solve the character pose under such challenge settings, we exploit the generative motion capture methods and propose a novel model-view consistency that considers both foreground and background in the tracking stage. The background is modeled as a deformable 2D grid, which allows us to compute the background-view consistency for sparse moving cameras. The 3D character pose is tracked with a global-local optimization through minimizing our consistency cost. A novel motion regularizer is also proposed in the optimization to constrain the solution pose space. The whole process of the proposed method is simple as frame by frame video segmentation is not required. Our method outperforms several alternative methods on various examples demonstrated in the paper. Yangang Wang 0001, Yebin Liu, Xin Tong 0001, Qionghai Dai, Ping Tan 0002 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2017 | Depth Estimation by Parameter Transfer With a Lightweight Model for Single Still ImagesabstractIn this paper, we propose a novel method for automatic depth estimation from color images using parameter transfer. By modeling the correlation between color images and their depth maps with a set of parameters, we get a database of parameter sets. Given an input image, we extract the high-level features to find the best matched image sets from the database. Then the set of parameters corresponding to the best match are used to estimate the depth of the input image. Compared with the past learning-based methods, our trained model consists only of trained features and parameter sets, which occupy little space. We evaluate our depth estimation method on several benchmark RGB-D (RGB + depth) data sets. The experimental results are comparable to the state-of-the-art results, while the model size is very small and very suitable for mobile devices, demonstrating the promising performance of our proposed method. Hongwei Qin, Xiu Li 0001, Yangang Wang 0001, Yongbing Zhang 0002, Qionghai Dai |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | Normalized filter pool for prior modeling of nature images
Yangang Wang 0001, Jin-Li Suo, Qionghai Dai |
Mach. Vis. Appl. | 1 |
| 2015 | Robust Non-rigid Motion Tracking and Surface Reconstruction Using L0 RegularizationabstractWe present a new motion tracking method to robustly reconstruct non-rigid geometries and motions from single view depth inputs captured by a consumer depth sensor. The idea comes from the observation of the existence of intrinsic articulated subspace in most of non-rigid motions. To take advantage of this characteristic, we propose a novel L0based motion regularizer with an iterative optimization solver that can implicitly constrain local deformation only on joints with articulated motions, leading to reduced solution space and physical plausible deformations. The L0strategy is integrated into the available non-rigid motion tracking pipeline, forming the proposed L0-L2non-rigid motion tracking method that can adaptively stop the tracking error propagation. Extensive experiments over complex human body motions with occlusions, face and hand motions demonstrate that our approach substantially improves tracking robustness and surface reconstruction accuracy. Feng Xu 0005, Yangang Wang 0001, Yebin Liu, Qionghai Dai |
ICCV | 3 |
| 2014 | DEPT: Depth Estimation by Parameter Transfer for Single Still Images
Xiu Li 0001, Hongwei Qin, Yangang Wang 0001, Yongbing Zhang 0002, Qionghai Dai |
ACCV (2) | 3 |
| 2014 | A Parametric Model for Describing the Correlation Between Single Color Images and Depth MapsabstractThis letter introduces a new approach for modeling the correlation between a single color image and its depth map with a set of parameters. The proposed model treats the color image as a set of patches and describes the correlation with a kernel function in a non-linear mapping space. We also present how to estimate the model parameters from sampled color image patches as well as the corresponding depth values. The proposed approach is tested on different color images and experimental results are comparable to the state-of-the-art, which demonstrates the power of the proposed method. Furthermore, we validate the efficiency of the proposed parametric model by evaluating each of its component, including the filters optimization, the choice of the patches and the kernel function. Yangang Wang 0001, Ruiping Wang 0001, Qionghai Dai |
IEEE Signal Process. Lett. | 1 |
| 2013 | Video-based hand manipulation capture through composite motion controlabstractThis paper describes a new method for acquiring physically realistic hand manipulation data from multiple video streams. The key idea of our approach is to introduce a composite motion control to simultaneously model hand articulation, object movement, and subtle interaction between the hand and object. We formulate video-based hand manipulation capture in an optimization framework by maximizing the consistency between the simulated motion and the observed image data. We search an optimal motion control that drives the simulation to best match the observed image data. We demonstrate the effectiveness of our approach by capturing a wide range of high-fidelity dexterous manipulation data. We show the power of our recovered motion controllers by adapting the captured motion data to new objects with different properties. The system achieves superior performance against alternative methods such as marker-based motion capture and kinematic hand motion tracking. Yangang Wang 0001, Jianyuan Min, Jianjie Zhang, Yebin Liu, Feng Xu 0005, Qionghai Dai, Jinxiang Chai |
ACM Trans. Graph. | 1 |