VLDB 2026 Research / reviewers in the wild / expert
Zhenbo Yu
dblp:177/5426
· DBLP profile ↗
19ranked-venue papers
7as first author
15since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 12 since 2021Artificial intelligence and machine learning · 11 · 4 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LaTexBlend: Scaling Multi-concept Customized Generation with Latent Textual BlendingabstractCustomized text-to-image generation renders user-specified concepts into novel contexts based on textual prompts. Scaling the number of concepts in customized generation meets a broader demand for user creation, whereas existing methods face challenges with generation quality and computational efficiency. In this paper, we propose LaTexBlend, a novel framework for effectively and efficiently scaling multi-concept customized generation. The core idea of LaTexBlend is to represent single concepts and blend multiple concepts within a Latent Textual space, which is positioned after the text encoder and a linear projection. LaTexBlend customizes each concept individually, storing them in a concept bank with a compact representation of latent textual features that captures sufficient concept information to ensure high fidelity. At inference, concepts from the bank can be freely and seamlessly combined in the latent textual space, offering two key merits for multi-concept generation: 1) excellent scalability, and 2) significant reduction of denoising deviation, preserving coherent layouts. Extensive experiments demonstrate that LaTexBlend can flexibly integrate multiple customized concepts with harmonious structures and high subject fidelity, substantially outperforming baselines in both generation quality and computational efficiency. Project page: https://jinjianrick.github.io/latexblend/ Zhenbo Yu, Yang Shen 0006, Zhenyong Fu, Jian Yang 0003 |
CVPR | 2 |
| 2025 | SSAIM: Not All Self-Attentions Contain Effective Spatial Structure in Diffusion Models for Text-to-Image EditingabstractWith the rapid progress of diffusion-based Text-to-Image Generation (TIG), Text-to-Image Editing (TIE) has become increasingly important for enabling controllable visual content creation. A core challenge in TIE is generating text-guided edits while preserving the spatial structure of the original image. Recent methods attempt to address this by leveraging self-attention maps from diffusion models, as these encode rich spatial information. However, we identify two key limitations: (1) not all self-attention maps contribute meaningfully to spatial structure, and (2) over-reliance on them can suppress desired editing effects. To address this, we propose the Spatial Information Score (SIS), a novel metric that quantifies the spatial structure encoded in each self-attention map. Leveraging SIS, we develop Selective Self-Attention-based Image Manipulation (SSAIM), which selectively utilizes self-attention maps with effective spatial structure (high SIS) to preserve the structural of the original image and reduce excessive reliance on self-attention maps with ineffective spatial structure (low SIS) to enhance editing performance in TIE tasks. Extensive experiments across diverse TIE tasks demonstrate that SSAIM significantly improves both structural fidelity and editing quality. Zhenbo Yu, Jimin Dai, Yingzhen Zhang, Jian Yang 0003, Lei Luo 0001 |
ACM Multimedia | 1 |
| 2025 | Exploring multi-semantic disentangled controls in GANs using conjugate gradient optimization
Zhenbo Yu, Zhenyong Fu, Jian Yang 0003 |
Pattern Recognit. Lett. | 1 |
| 2025 | Mesh2Animation: Unsupervised Animating for Quadruped 3D ObjectsabstractAnimating quadruped 3D objects, such as chairs and tables, typically involves three steps in the traditional computer graphics pipeline: Rigging, Skinning, and Retargeting. Commonly, prevailing methods for each specific step are conceived in isolation. For rigging and skinning steps, optimization-based methods are typically used, but these approaches tend to be slow and susceptible to variations in 3D mesh surfaces. For the retargeting step, the obtained results often fall short of expectations, especially when dealing with dissimilar source and target skeletons, leading to issues like joint twisting. The devised procedure is also time-intensive, resulting in a complex final pipeline. To this end, we present a unified framework, termed Mesh2Animation, providing an end-to-end solution to these challenges. In Mesh2Animation, a learning-based method is proposed for quadruped 3D skeleton estimation. We introduce both skeleton-level and mesh-level loss, allowing the rigging, skinning, and retargeting steps to be optimized simultaneously. Specifically, a general predicted estimation from the rigging step initializes the skeleton, making the skinning step faster and more accurate, which in turn leads to better results in the retargeting step. Finally, the rigging, skinning and retargeting processes are optimized simultaneously under static and temporal constraints. Additionally, we can construct a novel animating dataset termed ShapeNet2Animation (SN2Animation) based on the proposed method, which shows potential application for pose transfer. Qualitative and quantitative results on SN2Animation, ShapeNet, Object3D and ModelNet10 datasets for animation demonstrate that our method achieves competitive performance and shows promising generalization ability on quadruped 3D objects. Our project is available athttps://sites.google.com/view/mesh2animation. Zhenbo Yu, Jinxian Liu, Zefan Li, Bingbing Ni, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Variational Adversarial Defense: A Bayes Perspective for Adversarial TrainingabstractVarious methods have been proposed to defend against adversarial attacks. However, there is a lack of enough theoretical guarantee of the performance, thus leading to two problems: First, deficiency of necessary adversarial training samples might attenuate the normal gradient's back-propagation, which leads to overfitting and gradient masking potentially. Second, point-wise adversarial sampling offers an insufficient support region for adversarial data and thus cannot form a robust decision-boundary. To solve these issues, we provide a theoretical analysis to reveal the relationship between robust accuracy and the complexity of the training set in adversarial training. As a result, we propose a novel training scheme called Variational Adversarial Defense. Based on the distribution of adversarial samples, this novel construction upgrades the defend scheme from local point-wise to distribution-wise, yielding an enlarged support region for safeguarding robust training, thus possessing a higher promising to defense attacks. The proposed method features the following advantages: 1) Instead of seeking adversarial examples point-by-point (in a sequential way), we draw diverse adversarial examples from the inferred distribution; and 2) Augmenting the training set by a larger support region consolidates the smoothness of the decision boundary. Finally, the proposed method is analyzed via the Taylor expansion technique, which casts our solution with natural interpretability. Chenglong Zhao, Shibin Mei, Bingbing Ni, Shengchao Yuan, Zhenbo Yu, Jun Wang 0159 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Fast Fluid Simulation via Dynamic Multi-Scale GriddingabstractRecent works on learning-based frameworks for Lagrangian (i.e., particle-based) fluid simulation, though bypassing iterative pressure projection via efficient convolution operators, are still time-consuming due to excessive amount of particles. To address this challenge, we propose a dynamic multi-scale gridding method to reduce the magnitude of elements that have to be processed, by observing repeated particle motion patterns within certain consistent regions. Specifically, we hierarchically generate multi-scale micelles in Euclidean space by grouping particles that share similar motion patterns/characteristics based on super-light motion and scale estimation modules. With little internal motion variation, each micelle is modeled as a single rigid body with convolution only applied to a single representative particle. In addition, a distance-based interpolation is conducted to propagate relative motion message among micelles. With our efficient design, the network produces high visual fidelity fluid simulations with the inference time to be only 4.24 ms/frame (with 6K fluid particles), hence enables real-time human-computer interaction and animation. Experimental results on multiple datasets show that our work achieves great simulation acceleration with negligible prediction error increase. Jinxian Liu, Ye Chen 0006, Bingbing Ni, Zhenbo Yu |
AAAI | 5 |
| 2023 | Learning by Restoring Broken 3D GeometryabstractThe key point for an experienced craftsman to repair broken objects effectively is that he must know about them deeply. Similarly, we believe that a model can capture rich geometry information from a shape/scene and generate discriminative representations if it is able to find distorted parts of shapes/scenes and restore them. Inspired by this observation, we propose a novel self-supervised 3D learning paradigm named learning by restoring broken shapes/scenes (collectively called 3D geometry). We first develop a destroy-method cluster, from which we sample methods to break some local parts of an object. Then the destroyed object and the normal object are both sent into a point cloud network to get representations, which are employed to segment points that belong to distorted parts and further reconstruct/restore them to normal. To perform better in these two associated pretext tasks, the model is constrained to capture useful object features, such as rich geometric and contextual information. The object representations learned by this self-supervised paradigm transfer well to different datasets and perform well on downstream classification, segmentation and detection tasks. Experimental results on shape datasets and scene datasets demonstrate that our method achieves state-of-the-art performance among unsupervised methods. We also show experimentally that pre-training with our framework significantly boosts the performance of supervised models. Jinxian Liu, Bingbing Ni, Ye Chen 0006, Zhenbo Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Joint Global and Dynamic Pseudo Labeling for Semi-Supervised Point Cloud Sequence SegmentationabstractSupervised learning is a mainstay for large discriminative models in 3D computer vision, while large amounts of human-annotated data are the key to achieve state-of-the-art performance. This limitation is particularly notable for large-scale point cloud sequence segmentation tasks, because point-level annotations are very time-consuming and especially expensive. To overcome this challenge, we develop a novel semi-supervised framework for point cloud sequences segmentation. Specifically, we develop two kinds of pseudo labeling methods with extracting global semantic information from labeled frames and dynamic information from each sequence respectively. Then the two kinds of generated labels are combined as more robust pseudo labels (GD-Pseudo labels) for unlabeled frames. We finally apply an efficient iterative learning scheme to train a model with a small quantity of human-annotated data and large-scale pseudo-labeled data. Equipped with our framework, the model achieves significant performance improvement (+12—25 mIoU) on SemanticKITTI and Synthia when compared with frameworks that do not utilize large amounts of unlabeled data. Moreover, our method achieves comparable performance with only 20% annotated frames on SemanticKITTI to state-of-the-art models trained with 100% human-annotated frames. Jinxian Liu, Ye Chen 0006, Bingbing Ni, Zhenbo Yu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Cross-Species 3D Face Morphing via Alignment-Aware ControllerabstractWe address cross-species 3D face morphing (i.e., 3D face morphing from human to animal), a novel problem with promising applications in social media and movie industry. It remains challenging how to preserve target structural information and source fine-grained facial details simultaneously. To this end, we propose an Alignment-aware 3D Face Morphing (AFM) framework, which builds semantic-adaptive correspondence between source and target faces across species, via an alignment-aware controller mesh (Explicit Controller, EC) with explicit source/target mesh binding. Based on EC, we introduce Controller-Based Mapping (CBM), which builds semantic consistency between source and target faces according to the semantic importance of different face regions. Additionally, an inference-stage coarse-to-fine strategy is exploited to produce fine-grained meshes with rich facial details from rough meshes. Extensive experimental results in multiple people and animals demonstrate that our method produces high-quality deformation results. Xirui Yan, Zhenbo Yu, Bingbing Ni |
AAAI | 2 |
| 2022 | Object Wake-Up: 3D Object Rigging from a Single Image
Xinxin Zuo, Sen Wang 0003, Zhenbo Yu, Bingbing Ni, Minglun Gong, Li Cheng 0001 |
ECCV (2) | 4 |
| 2022 | Skeleton2Humanoid: Animating Simulated Characters for Physically-plausible Motion In-betweeningabstractHuman motion synthesis is a long-standing problem with various applications in digital twins and the Metaverse. However, modern deep learning based motion synthesis approaches barely consider the physical plausibility of synthesized motions and consequently they usually produce unrealistic human motions. In order to solve this problem, we propose a system "Skeleton2Humanoid" which performs physics-oriented motion correction at test time by regularizing synthesized skeleton motions in a physics simulator. Concretely, our system consists of three sequential stages: (I) test time motion synthesis network adaptation, (II) skeleton to humanoid matching and (III) motion imitation based on reinforcement learning (RL). Stage I introduces a test time adaptation strategy, which improves the physical plausibility of synthesized human skeleton motions by optimizing skeleton joint locations. Stage II performs an analytical inverse kinematics strategy, which converts the optimized human skeleton motions to humanoid robot motions in a physics simulator, then the converted humanoid robot motions can be served as reference motions for the RL policy to imitate. Stage III introduces a curriculum residual force control policy, which drives the humanoid robot to mimic complex converted reference motions in accordance with the physical law. We verify our system on a typical human motion synthesis task, motion-in-betweening. Experiments on the challenging LaFAN1 dataset show our system can outperform prior methods significantly in terms of both physical plausibility and accuracy. Code will be released for research purposes at: https://github.com/michaelliyunhao/Skeleton2Humanoid. Zhenbo Yu, Yucheng Zhu, Bingbing Ni, Guangtao Zhai, Wei Shen 0002 |
ACM Multimedia | 2 |
| 2022 | OCR-Pose: Occlusion-aware Contrastive Representation for Unsupervised 3D Human Pose EstimationabstractOcclusion is a significant problem in 3D human pose estimation from the 2D counterpart. On one hand, without explicit annotation, the 3D skeleton is hard to be accurately estimated from the occluded 2D pose. On the other hand, one occluded 2D pose might correspond to multiple 3D skeletons with low confidence parts. To address these issues, we decouple the 3D representation feature into view-invariant part termed occlusion-aware feature and view-dependent part termed rotation feature to facilitate subsequent optimization of the former. Then we propose an occlusion-aware contrastive representation based scheme (OCR-Pose) consisting of Topology Invariant Contrastive Learning module (TiCLR) and View Equivariant Contrastive Learning module (VeCLR). Specifically, TiCLR drives invariance to topology transformation, i.e., bridging the gap between an occluded 2D pose and the unoccluded one. While VeCLR encourages equivariance to view transformation, i.e., capturing the geometric similarity of the 3D skeleton in two views. Both modules optimize occlusion-aware constrastive representation with pose filling and lifting networks via an iterative training strategy in an end-to-end manner. OCR-Pose not only achieves superior performance against state-of-the-art unsupervised methods on unoccluded benchmarks, but also obtains significant improvements when occlusion is involved. Our project is available at https://sites.google.com/view/ocr-pose. Zhenbo Yu, Zhengyan Tong, Jinxian Liu, Wenjun Zhang 0001 |
ACM Multimedia | 2 |
| 2022 | Limb Pose Aware Networks for Monocular 3D Pose EstimationabstractIn the task of monocular 3D pose estimation, the estimation errors of limb joints (i.e., wrist, ankle, etc) with a higher degree of freedom(DOF) are larger than that of others (i.e., hip, thorax, etc). Specifically, errors may accumulate along the physiological structure of human body parts, and trajectories of joints with higher DOF bring in higher complexity. To address this problem, we propose a limb pose aware framework, involving a kinematic constraint aware network as well as a trajectory aware temporal module, to improve the 3D prediction accuracy of limb joint positions. Two kinematic constraints named relative bone angles and absolute bone angles are introduced in this paper, the former being used for building the angular relation between adjacent bones and the latter for building the angular relation between bones and the camera plane. As a joint result of two constraints, our work suppresses errors accumulated along limbs. Furthermore, we propose a trajectory-aware network, named as Hierarchical Transformer, which takes temporal trajectories of joints as input and generates fused trajectory estimation as a result. The Hierarchical Transformer consists of Transformer Encoder blocks and aims at improving the performance of fusing temporal features. Under the effect of kinematic constraints and trajectory network, we alleviate the problem of errors accumulated along limbs and achieve promising results. Most of the off-the-shelf 2D pose estimators can be easily integrated into our framework. We perform extensive experiments on public datasets and validate the effectiveness of the framework. The ablation studies show the strength of each individual sub-module. Lele Wu, Zhenbo Yu, Yijiang Liu, Qingshan Liu 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Towards Alleviating the Modeling Ambiguity of Unsupervised Monocular 3D Human Pose EstimationabstractIn this work, we study the ambiguity problem in the task of unsupervised 3D human pose estimation from 2D counterpart. On one hand, without explicit annotation, the scale of 3D pose is difficult to be accurately captured (scale ambiguity). On the other hand, one 2D pose might correspond to multiple 3D gestures, where the lifting procedure is inherently ambiguous (pose ambiguity). Previous methods generally use temporal constraints (e.g., constant bone length and motion smoothness) to alleviate the above issues. However, these methods commonly enforce the outputs to fulfill multiple training objectives simultaneously, which often lead to sub-optimal results. In contrast to the majority of previous works, we propose to split the whole problem into two sub-tasks, i.e., optimizing 2D input poses via a scale estimation module and then mapping optimized 2D pose to 3D counterpart via a pose lifting module. Furthermore, two temporal constraints are proposed to alleviate the scale and pose ambiguity respectively. These two modules are optimized via a iterative training scheme with corresponding temporal constraints, which effectively reduce the learning difficulty and lead to better performance. Results on the Human3.6M dataset demonstrate that our approach improves upon the prior art by 23.1% and also outperforms several weakly supervised approaches that rely on 3D annotations. Our project is available at https://sites.google.com/view/ambiguity-aware-hpe. Zhenbo Yu, Bingbing Ni, Jingwei Xu 0005, Chenglong Zhao, Wenjun Zhang 0001 |
ICCV | 1 |
| 2021 | Skeleton2Mesh: Kinematics Prior Injected Unsupervised Human Mesh RecoveryabstractIn this paper, we decouple unsupervised human mesh recovery into the well-studied problems of unsupervised 3D pose estimation, and human mesh recovery from estimated 3D skeletons, focusing on the latter task. The challenges of the latter task are two folds: (1) pose failure (i.e., pose mismatching – different skeleton definitions in dataset and SMPL , and pose ambiguity – endpoints have arbitrary joint angle configurations for the same 3D joint coordinates). (2) shape ambiguity (i.e., the lack of shape constraints on body configuration). To address these issues, we propose Skeleton2Mesh, a novel lightweight framework that recovers human mesh from a single image. Our Skeleton2Mesh contains three modules, i.e., Differentiable Inverse Kinematics (DIK), Pose Refinement (PR) and Shape Refinement (SR) modules. DIK is designed to transfer 3D rotation from estimated 3D skeletons, which relies on a minimal set of kinematics prior knowledge. Then PR and SR modules are utilized to tackle the pose ambiguity and shape ambiguity respectively. All three modules can be incorporated into Skeleton2Mesh seamlessly via an end-to-end manner. Furthermore, we utilize an adaptive joint regressor to alleviate the effects of skeletal topology from different datasets. Results on the Human3.6M dataset for human mesh recovery demonstrate that our method improves upon the previous unsupervised methods by 32.6% under the same setting. Qualitative results on in-the-wild datasets exhibit that the recovered 3D meshes are natural, realistic. Our project is available at https://sites.google.com/view/skeleton2mesh. Zhenbo Yu, Jingwei Xu 0005, Bingbing Ni, Chenglong Zhao, Minsi Wang, Wenjun Zhang 0001 |
ICCV | 1 |
| 2020 | Deep Kinematics Analysis for Monocular 3D Human Pose EstimationabstractFor monocular 3D pose estimation conditioned on 2D detection, noisy/unreliable input is a key obstacle in this task. Simple structure constraints attempting to tackle this problem, e.g., symmetry loss and joint angle limit, could only provide marginal improvements and are commonly treated as auxiliary losses in previous researches. Thus it still remains challenging about how to effectively utilize the power of human prior knowledge for this task. In this paper, we propose to address above issue in a systematic view. Firstly, we show that optimizing the kinematics structure of noisy 2D inputs is critical to obtain accurate 3D estimations. Secondly, based on corrected 2D joints, we further explicitly decompose articulated motion with human topology, which leads to more compact 3D static structure easier for estimation. Finally, temporal refinement emphasizing the validity of 3D dynamic structure is naturally developed to pursue more accurate result. Above three steps are seamlessly integrated into deep neural models, which form a deep kinematics analysis pipeline concurrently considering the static/dynamic structure of 2D inputs and 3D outputs. Extensive experiments show that proposed framework achieves state-of-the-art performance on two widely used 3D human action datasets. Meanwhile, targeted ablation study shows that each former step is critical for the latter one to obtain promising results. Jingwei Xu 0005, Zhenbo Yu, Bingbing Ni, Jiancheng Yang, Xiaokang Yang 0001, Wenjun Zhang 0001 |
CVPR | 2 |
| 2020 | Deep convolutional BiLSTM fusion network for facial expression recognition
Dandan Liang, Huagang Liang, Zhenbo Yu, Yipu Zhang 0001 |
Vis. Comput. | 3 |
| 2018 | Spatio-temporal convolutional features with nested LSTM for facial expression recognition
Zhenbo Yu, Guangcan Liu, Qingshan Liu 0001, Jiankang Deng |
Neurocomputing | 1 |
| 2018 | Deeper cascaded peak-piloted network for weak expression recognition
Zhenbo Yu, Qinshan Liu, Guangcan Liu |
Vis. Comput. | 1 |