EDBT 2026 Demo / reviewers in the wild / expert
Zimeng Zhao
dblp:310/5441
· DBLP profile ↗
16ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0001-6570-0620ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 14 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Two-stage diffusion for hands and articulated objects interaction synthesis
Wenqian Sun, Binghui Zuo, Zimeng Zhao, Yangang Wang 0001 |
Pattern Recognit. | 3 |
| 2026 | FOSUP: Dynamic Garments Diffusion With Fourier Spherical Unwrapping From Monocular VideoabstractRecent advances in dynamic garment reconstruction boost virtual try-on with monocular video streams as inputs. However, existing literature has been intensively conducted on its sub-tasks, including static garment reconstruction and dynamic clothed human reconstruction, which are difficult to extend to dynamic garment reconstruction. The former bottleneck is mainly the lack of cross-frame correspondences and independent clothing topology on implicit garment fields, which results in the inability to obtain accurate motion information during dynamic clothing reconstruction and the absence of stable topology critical for downstream tasks, such as animation with physics engines. The latter usually binds the garment motion with body or skeleton motions, leading to rigid artifacts for loose-fitting garments. Our key idea is to build a diffusion based T-pose garments generator with a strong prior on garments structure. The garment generator is trained to generate 2D clothing representation, termed FOSUP, conditioned by a monocular video. FOSUP, defined as FOurier Spherical Unwrapping, enables a bidirectional mapping between FOSUP and the mesh through FFT and inverse FFT, which maintain spatial order and adjacency. Subsequently, this FOSUP is mapped back to 3D meshes through an inverse FFT process and transformed into pose space through a point transformation network to guide the three-dimensional reconstruction of the entire sequence. To sufficiently train our framework and address the lack of domain-specific data, we have constructed a large-scale garment MoCap dataset. This dataset captures the motion of various loose garments and includes multi-view raw images, frame-by-frame human motion annotations, raw scanned point clouds, topology-independent garment templates, and garment meshes with cross-frame correspondences. Comprehensive experiments have demonstrated that our unwrapping-based representation and diffusion-based framework significantly improve the performance and robustness of dynamic garment reconstruction. Zimeng Zhao, Yangang Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | Fine-Grained Change Point Detection for Topic Modeling with Pitman-Yor ProcessabstractIdentifying change points in dynamic text data is crucial for understanding the evolving nature of topics across various sources, such as news articles, scientific papers, and social media posts. While topic modeling has become a widely used technique for this purpose, capturing fine-grained shifts in individual topics over time remains a significant challenge. Traditional approaches typically use a two-stage process, separating topic modeling and change point detection. However, this separation can lead to information loss and inconsistency in capturing subtle changes in topic evolution. To address this issue, we propose TOPIC-PYP, a change point detection model specifically designed for fine-grained topic-level analysis, i.e., detecting change points for each individual topic. By leveraging the Pitman-Yor process, TOPIC-PYP effectively captures the dynamic evolution of topic meanings over time. Unlike traditional methods, TOPIC-PYP integrates topic modeling and change point detection into a unified framework, facilitating a more comprehensive understanding of the relationship between topic evolution and change points. Experimental evaluations on both synthetic and real-world datasets demonstrate the effectiveness of TOPIC-PYP in accurately detecting change points and generating high-quality topics. Zimeng Zhao, Ruimin Ye, Xiaoge Gu, Xiaoling Lu |
J. Mach. Learn. Res. | 2 |
| 2025 | NP-Hand: Novel Perspective Hand Image Synthesis Guided by NormalsabstractSynthesizing multi-view images that are geometrically consistent with a given single-view image is one of the hot issues in AIGC in recent years. Existing methods have achieved impressive performance on objects with symmetry or rigidity, but they are inappropriate for the human hand. Because an image-captured human hand has more diverse poses and less attractive textures. In this paper, we propose NP-Hand, a framework that elegantly combines the diffusion model and generative adversarial network: The multi-step diffusion is trained to synthesize low-resolution novel perspective, while the single-step generator is exploited to further enhance synthesis quality. To maintain the consistency between inputs and synthesis, we creatively introduce normal maps into NP-Hand to guide the whole synthesizing process. Comprehensive evaluations have demonstrated that the proposed framework is superior to existing state-of-the-art models and more suitable for synthesizing hand images with faithful structures and realistic appearance details. The code will be released on our website. Binghui Zuo, Wenqian Sun, Zimeng Zhao, Yangang Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | T2C: Text-guided 4D Cloth GenerationabstractIn the age of AIGC, the creation process is increasingly automated. Generating vivid characters with clothing and motions according to scripts or novels is no exception. Unfortunately, the diversity of fabric topologies, the complexity of fabric layering, and the flexibility of fabric motion make most approaches only applicable to motion generation for characters in undressing or tight-fitting clothing. This article introduces a novel approach named T2C , which employs a multi-layered clothing representation and a physics-based clothing animation paradigm to generate text-controlled Clothed 4D Humans, expanding the boundaries of the aforementioned issues. The hierarchical representation of clothing utilizes Fourier spherical mapping to define the geometric information of garments within a standard pose space, mapping it onto several 2D frequency domain subspaces. The motion of clothing in tandem with the human body is realized through a hybrid forward dynamic solution, where the internal virtual mechanic’s parameters driving the clothing are learned from text features. A series of qualitative and quantitative experiments reveal that T2C can generate dynamic clothing with a sense of layering, realistic details, and rich textures. The code will be publicly available at https://zhipengyu28.github.io/t2c/ . Zimeng Zhao, Yanxi Du, Yuzhou Zheng, Binghui Zuo, Yangang Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | GraspDiff: Grasping Generation for Hand-Object Interaction With Multimodal Guided DiffusionabstractGrasping generation holds significant importance in both robotics and AI-generated content. While pure network paradigms based on VAEs or GANs ensure diversity in outcomes, they often fall short of achieving plausibility. Additionally, although those two-step paradigms that first predict contact and then optimize distance yield plausible results, they are always known to be time-consuming. This paper introduces a novel paradigm powered by DDPM, accommodating diverse modalities with varying interaction granularities as its generating conditions, including 3D object, contact affordance, and image content. Our key idea is that the iterative steps inherent to diffusion models can supplant the iterative optimization routines in existing optimization methods, thereby endowing the generated results from our method with both diversity and plausibility. Using the same training data, our paradigm achieves superior generation performance and competitive generation speed compared to optimization-based paradigms. Extensive experiments on both in-domain and out-of-domain objects demonstrate that our method receives significant improvement over the SOTA method. We will release the code for research purposes. Binghui Zuo, Zimeng Zhao, Wenqian Sun, Yangang Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Music Conditioned Generation for Human-Centric VideoabstractMusic and human-centric video are two fundamental signals across languages. Correlation analysis between the two is currently used in choreography and film accompaniment. This letter explores this correlation in a new task: human-centric video generation from a start-end image pair and transitional music. Existing human-centric generation methods are not competent for this task because they require frame-wise pose as input or have difficulty handling long-duration videos. Our key idea is to build a temporal generation framework dominated by DDPM and assisted by VAE and GAN. It reduces the computational cost of music-image diffusion by utilizing the latent space compactness of VAE and the image translation efficiency of GAN. To produce videos with both long duration and high quality, our framework first generates small-scale keyframes and then generates high-resolution videos. To strengthen the frame-wise consistency of the human body, a frame-aligned correspondence map is adopted as an intermediate supervision. Extensive experiments compared with the SOTA method have demonstrated the rationality and effectiveness of this signal generation framework. Zimeng Zhao, Binghui Zuo, Yangang Wang 0001 |
IEEE Signal Process. Lett. | 1 |
| 2024 | Skeleton Extraction for Articulated Objects With the Spherical Unwrapping ProfilesabstractEmbedding unified skeletons into unregistered scans is fundamental to finding correspondences, depicting motions, and capturing underlying structures among the articulated objects in the same category. Some existing approaches rely on laborious registration to adapt a predefined LBS model to each input, while others require the input to be set to a canonical pose, e.g., T-pose or A-pose. However, their effectiveness is always influenced by the water-tightness, face topology, and vertex density of the input mesh. At the core of our approach lies a novel unwrapping method, named SUPPLE (Spherical UnwraPping ProfiLEs), which maps a surface into image planes independent of mesh topologies. Based on this lower-dimensional representation, a learning-based framework is further designed to localize and connect skeletal joints with fully convolutional architectures. Experiments demonstrate that our framework yields reliable skeleton extractions across a broad range of articulated categories, from raw scans to online CADs. Zimeng Zhao, Wei Xie 0012, Binghui Zuo, Yangang Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | Semi-Supervised Hand Appearance Recovery via Structure Disentanglement and Dual Adversarial DiscriminationabstractEnormous hand images with reliable annotations are collected through marker-based MoCap. Unfortunately, degradations caused by markers limit their application in hand appearance reconstruction. A clear appearance recovery insight is an image-to-image translation trained with unpaired data. However, most frameworks fail because there exists structure inconsistency from a degraded hand to a bare one. The core of our approach is to first disentangle the bare hand structure from those degraded images and then wrap the appearance to this structure with a dual adversarial discrimination (DAD) scheme. Both modules take full advantage of the semi-supervised learning paradigm: The structure disentanglement benefits from the modeling ability of ViT, and the translator is enhanced by the dual discrimination on both translation processes and translation results. Comprehensive evaluations have been conducted to prove that our framework can robustly recover photo-realistic hand appearance from diverse marker-contained and even object-occluded datasets. It provides a novel avenue to acquire bare hand appearance data for other down-stream learning problems. Zimeng Zhao, Binghui Zuo, Zhiyu Long, Yangang Wang 0001 |
CVPR | 1 |
| 2023 | Nonrigid Object Contact Estimation With Regional Unwrapping TransformerabstractAcquiring contact patterns between hands and nonrigid objects is a common concern in the vision and robotics community. However, existing learning-based methods focus more on contact with rigid ones from monocular images. When adopting them for nonrigid contact, a major problem is that the existing contact representation is restricted by the geometry of the object. Consequently, contact neighborhoods are stored in an unordered manner and contact features are difficult to align with image cues. At the core of our approach lies a novel hand-object contact representation called RUPs (Region Unwrapping Profiles), which unwrap the roughly estimated hand-object surfaces as multiple high-resolution 2D regional profiles. The region grouping strategy is consistent with the hand kinematic bone division because they are the primitive initiators for a composite contact pattern. Based on this representation, our Regional Unwrapping Transformer (RUFormer) learns the correlation priors across regions from monocular inputs and predicts corresponding contact and deformed transformations. Our experiments demonstrate that the proposed framework can robustly estimate the deformed degrees and deformed transformations, which makes it suitable for both nonrigid and rigid contact. Wei Xie 0012, Zimeng Zhao, Binghui Zuo, Yangang Wang 0001 |
ICCV | 2 |
| 2023 | Reconstructing Interacting Hands with Interaction Prior from Monocular ImagesabstractReconstructing interacting hands from monocular images is indispensable in AR/VR applications. Most existing solutions rely on the accurate localization of each skeleton joint. However, these methods tend to be unreliable due to the severe occlusion and confusing similarity among adjacent hand parts. This also defies human perception because humans can quickly imitate an interaction pattern without localizing all joints. Our key idea is to first construct a two-hand interaction prior and recast the interaction reconstruction task as the conditional sampling from the prior. To expand more interaction states, a large-scale multimodal dataset with physical plausibility is proposed. Then a VAE is trained to further condense these interaction patterns as latent codes in a prior distribution. When looking for image cues that contribute to interaction prior sampling, we propose the interaction adjacency heatmap (IAH). Compared with a joint-wise heatmap for localization, IAH assigns denser visible features to those invisible joints. Compared with an all-in-one visible heatmap, it provides more fine-grained local interaction information in each interaction region. Finally, the correlations between the extracted features and corresponding interaction codes are linked by the ViT module. Comprehensive evaluations on benchmark datasets have verified the effectiveness of this framework. The code and dataset are publicly available at https://github.com/binghui-z/InterPrior_pytorch. Binghui Zuo, Zimeng Zhao, Wenqian Sun, Wei Xie 0012, Zhou Xue, Yangang Wang 0001 |
ICCV | 2 |
| 2023 | Implicit Representation for Interacting Hands Reconstruction from Monocular Color Images
Binghui Zuo, Zimeng Zhao, Wei Xie 0012, Yangang Wang 0001 |
ICIG (1) | 2 |
| 2023 | HMDO : Markerless multi-view hand manipulation capture with deformable objectsabstractWe construct the first markerless deformable interaction dataset recording interactive motions of the hands and deformable objects, called HMDO (Hand Manipulation with Deformable Objects). With our built multi-view capture system, it captures the deformable interactions with multiple perspectives, various object shapes, and diverse interactive forms. Our motivation is the current lack of hand and deformable object interaction datasets, as 3D hand and deformable object reconstruction is challenging. Mainly due to mutual occlusion, the interaction area is difficult to observe, the visual features between the hand and the object are entangled, and the reconstruction of the interaction area deformation is difficult. To tackle this challenge, we propose a method to annotate our captured data. Our key idea is to collaborate with estimated hand features to guide the object global pose estimation, and then optimize the deformation process of the object by analyzing the relationship between the hand and the object. Through comprehensive evaluation, the proposed method can reconstruct interactive motions of hands and deformable objects with high quality. HMDO currently consists of 21600 frames over 12 sequences. In the future, this dataset could boost the research of learning-based reconstruction of deformable interaction scenes. Wei Xie 0012, Zimeng Zhao, Binghui Zuo, Yangang Wang 0001 |
Graph. Model. | 3 |
| 2022 | Stability-driven Contact Reconstruction From Monocular Color ImagesabstractPhysical contact provides additional constraints for hand-object state reconstruction as well as a basis for further understanding of interaction affordances. Estimating these severely occluded regions from monocular images presents a considerable challenge. Existing methods optimize the hand-object contact driven by distance threshold or prior from contact-labeled datasets. However, due to the number of subjects and objects involved in these indoor datasets being limited, the learned contact patterns could not be generalized easily. Our key idea is to reconstruct the contact pattern directly from monocular images, and then utilize the physical stability criterion in the simulation to optimize it. This criterion is defined by the resultant forces and contact distribution computed by the physics engine. Compared to existing solutions, our framework can be adapted to more personalized hands and diverse object shapes. Furthermore, an interaction dataset with extra physical attributes is created to verify the sim-to-real consistency of our methods. Through comprehensive evaluations, hand-object contact can be reconstructed with both accuracy and stability by the proposed framework. Zimeng Zhao, Binghui Zuo, Wei Xie 0012, Yangang Wang 0001 |
CVPR | 1 |
| 2021 | SUPPLE: Extracting Hand Skeleton with Spherical Unwrapping ProfilesabstractEmbedding a unified skeleton into diverse hand meshes is a prominent task both for animation and pose estimation. Most existing methods extracted skeletons from humanoid characters under simple poses, e.g. T- pose or A- pose. Applying them directly to hand meshes may yield inaccurate or implausible results because hands have higher dexterity and similar endpoints. Furthermore, these methods did not attempt to extract skeleton directly from a scan model which may be not watertight and has much more vertices. Our key idea is to unwrap meshes with different topologies in the same image-based representation, named SUPPLE (Spherical UnwraPping ProfiLEs), and then train a convolutional encoder-decoder to extract skeleton under this representation. Experiments demonstrate that our framework produces reliable and accurate skeleton estimation results across a broad range of datasets, from raw scans to artist-designed models. Zimeng Zhao, Ruting Rao, Yangang Wang 0001 |
3DV | 1 |
| 2021 | TravelNet: Self-supervised Physically Plausible Hand Motion Learning from Monocular Color ImagesabstractThis paper aims to reconstruct physically plausible hand motion from monocular color images. Existing frame-by-frame estimating approaches can not guarantee the physical plausibility (e.g. penetration, jittering) directly. In this paper, we embed physical constraints on the per-frame estimated motions in both spatial and temporal space. Our key idea is to adopt a self-supervised learning strategy to train a novel encoder-decoder, named TravelNet, whose training motion data is prepared by the physics engine using discrete pose states. TravelNet captures key pose states from hand motion sequences as compact motion descriptors, inspired by the concept of keyframes in animation. Finally, it manages to extract those key states out of perturbations without manual annotations, and reconstruct the motions preserving details and physical plausibility. In the experiments, we show that the outputs of the TravelNet contain both finger synergism and time consistency. Through the proposed framework, hand motions can be accurately reconstructed and flexibly re-edited, which is superior to the state-of-the-art methods. Zimeng Zhao, Yangang Wang 0001 |
ICCV | 1 |