VLDB 2026 Research / reviewers in the wild / expert
Zhenyu Zhang 0005
dblp:01/1844-5
· DBLP profile ↗
44ranked-venue papers
10as first author
33since 2021 · last 2026
0000-0001-5727-9450ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 9 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 8 first-author · 25 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OmniDrag: Enabling Motion Control for Omnidirectional Image-to-Video Generation
Shijie Zhao 0001, Chong Mou, Xuhan Sheng, Zhenyu Zhang 0005, Li Zhang 0006, Jian Zhang 0018 |
Int. J. Comput. Vis. | 5 |
| 2026 | AddSR: Accelerating diffusion-based blind super-resolution with adversarial diffusion distillation
Ying Tai, Rui Xie 0005, Chen Zhao 0002, Kai Zhang 0008, Zhenyu Zhang 0005, Jian Yang 0003 |
Pattern Recognit. | 5 |
| 2026 | DreamBarbie: Text to Barbie-Style 3D AvatarsabstractTo integrate digital humans into everyday life, there is a strong demand for generating high-quality, fine-grained disentangled 3D avatars that support expressive animation and simulation capabilities, ideally from low-cost textual inputs. Although text-driven 3D avatar generation has made significant progress by leveraging 2D generative priors, existing methods still struggle to fulfill all these requirements simultaneously. To address this challenge, we propose DreamBarbie, a novel text-driven framework for generating animatable 3D avatars with separable shoes, accessories, and simulation-ready garments, truly capturing the iconic "Barbie doll" aesthetic. The core of our framework lies in an expressive 3D representation combined with appropriate modeling constraints. Unlike prior methods, we use G-Shell to uniformly model watertight components (e.g., bodies, shoes) and non-watertight garments. By reformulating boundaries as euclidean field intersections instead of manifold geodesics, we propose an SDF-based initialization and a hole regularization loss that together achieve a $100\times$100× speedup and stable open topology without image input. These disentangled 3D representations are then optimized by specialized expert diffusion models tailored to each domain, ensuring high-fidelity outputs. To mitigate geometric artifacts and texture conflicts when combining different expert models, we further propose several effective geometric losses and strategies. Extensive experiments demonstrate that DreamBarbie outperforms existing methods in both dressed human and outfit generation. Our framework further enables diverse applications, including apparel combination, editing, expressive animation, and physical simulation. Xiaokun Sun, Zhenyu Zhang 0005, Ying Tai, Hao Tang 0005, Zili Yi, Jian Yang 0003 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | Exploiting Multimodal Spatial-temporal Patterns for Video Object TrackingabstractMultimodal tracking has garnered widespread attention as a result of its ability to effectively address the inherent limitations of traditional RGB tracking. However, existing multimodal trackers mainly focus on the fusion and enhancement of spatial features or merely leverage the sparse temporal relationships between video frames. These approaches do not fully exploit the temporal correlations in multimodal videos, making it difficult to capture the dynamic changes and motion information of targets in complex scenarios. To alleviate this problem, we propose a unified multimodal spatial-temporal tracking approach named STTrack. In contrast to previous paradigms that solely relied on updating reference information, we introduced a temporal state generator (TSG) that continuously generates a sequence of tokens containing multimodal temporal information. These temporal information tokens are used to guide the localization of the target in the next time state, establish long-range contextual relationships between video frames, and capture the temporal trajectory of the target. Furthermore, at the spatial level, we introduced the mamba fusion and background suppression interactive (BSI) modules. These modules establish a dual-stage mechanism for coordinating information interaction and fusion between modalities. Extensive comparisons on five benchmark datasets illustrate that STTrack achieves state-of-the-art performance across various multimodal tracking scenarios. Xiantao Hu, Ying Tai, Xu Zhao 0001, Chen Zhao 0002, Zhenyu Zhang 0005, Jun Li 0027, Bineng Zhong 0001, Jian Yang 0003 |
AAAI | 5 |
| 2025 | Anywhere: A Multi-Agent Framework for User-Guided, Reliable, and Diverse Foreground-Conditioned Image GenerationabstractRecent advancements in image-conditioned image generation have demonstrated substantial progress. However, foreground-conditioned image generation remains underexplored, encountering challenges such as compromised object integrity, foreground-background inconsistencies, limited diversity, and reduced control flexibility. These challenges arise from current end-to-end inpainting models, which suffer from inaccurate training masks, limited foreground semantic understanding, data distribution biases, and inherent interference between visual and textual prompts. To overcome these limitations, we present Anywhere, a multi-agent framework that departs from the traditional end-to-end approach. In this framework, each agent is specialized in a distinct aspect, such as foreground understanding, diversity enhancement, object integrity protection, and textual prompt consistency. Our framework is further enhanced with the ability to incorporate optional user textual inputs, perform automated quality assessments, and initiate re-generation as needed. Comprehensive experiments demonstrate that this modular design effectively overcomes the limitations of existing end-to-end models, resulting in higher fidelity, quality, diversity and controllability in foreground-conditioned image generation. Additionally, the Anywhere framework is extensible, allowing it to benefit from future advancements in each individual agent. Tianyidan Xie, Rui Ma 0011, Xiaoqian Ye, Feixuan Liu, Ying Tai, Zhenyu Zhang 0005, Lanjun Wang, Zili Yi |
AAAI | 7 |
| 2025 | Describe, Don't Dictate: Semantic Image Editing with Natural Language IntentabstractDespite the progress in text-to-image generation, semantic image editing remains a challenge. Inversion-based algorithms unavoidably introduce reconstruction errors, while instruction-based models mainly suffer from limited dataset quality and scale. To address these problems, we propose a descriptive-prompt-based editing framework, named DescriptiveEdit. The core idea is to re-frame `instruction-based image editing' as `reference-image-based text-to-image generation', which preserves the generative power of well-trained Text-to-Image models without architectural modifications or inversion. Specifically, taking the reference image and a prompt as input, we introduce a Cross-Attentive UNet, which newly adds attention bridges to inject reference image features into the prompt-to-edit-image generation process. Owing to its text-to-image nature, DescriptiveEdit overcomes limitations in instruction dataset quality, integrates seamlessly with ControlNet, IP-Adapter, and other extensions, and is more scalable. Experiments on the Emu Edit benchmark show it improves editing accuracy and consistency. En Ci, Shanyan Guan, Yanhao Ge, Zhenyu Zhang 0005, Jian Yang 0003, Ying Tai |
ICCV | 6 |
| 2025 | StrandHead: Text to Hair-Disentangled 3D Head Avatars Using Human-Centric Priors
Xiaokun Sun, Ying Tai, Jian Yang 0003, Zhenyu Zhang 0005 |
ICCV | 5 |
| 2025 | Star: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-ResolutionabstractImage diffusion models have been adapted for real-world video super-resolution to tackle over-smoothing issues in GAN-based methods. However, these models struggle to maintain temporal consistency, as they are trained on static images, limiting their ability to capture temporal dynamics effectively. Integrating text-to-video (T2V) models into video super-resolution for improved temporal modeling is straightforward. However, two key challenges remain: artifacts introduced by complex degradations in real-world scenarios, and compromised fidelity due to the strong generative capacity of powerful T2V models (\textit{e.g.}, CogVideoX-5B). To enhance the spatio-temporal quality of restored videos, we introduce\textbf{~\name} (\textbf{S}patial-\textbf{T}emporal \textbf{A}ugmentation with T2V models for \textbf{R}eal-world video super-resolution), a novel approach that leverages T2V models for real-world video super-resolution, achieving realistic spatial details and robust temporal consistency. Specifically, we introduce a Local Information Enhancement Module (LIEM) before the global attention block to enrich local details and mitigate degradation artifacts. Moreover, we propose a Dynamic Frequency (DF) Loss to reinforce fidelity, guiding the model to focus on different frequency components across diffusion steps. Extensive experiments demonstrate\textbf{~\name}~outperforms state-of-the-art methods on both synthetic and real-world datasets. Rui Xie 0005, Yinhong Liu, Penghao Zhou, Chen Zhao 0002, Kai Zhang 0008, Zhenyu Zhang 0005, Jian Yang 0003, Zhenheng Yang, Ying Tai |
ICCV | 7 |
| 2025 | EasySplat: View-Adaptive Learning makes 3D Gaussian Splatting Easyabstract3D Gaussian Splatting (3DGS) techniques have achieved satisfactory 3D scene representation. Despite their impressive performance, they confront challenges due to the limitation of structure-from-motion (SfM) methods on acquiring accurate scene initialization, or the inefficiency of densification strategy. In this paper, we introduce a novel framework EasySplat to achieve high-quality 3DGS modeling. Instead of using SfM for scene initialization, we employ a novel method to release the power of large-scale pointmap approaches. Specifically, we propose an efficient grouping strategy based on view similarity, and use robust pointmap priors to obtain high-quality point clouds and camera poses for 3D scene initialization. After obtaining a reliable scene structure, we propose a novel densification approach that adaptively splits Gaussian primitives based on the average shape of neighboring Gaussian ellipsoids, utilizing KNN scheme. In this way, the proposed method tackles the limitation on initialization and optimization, leading to an efficient and accurate 3DGS modeling. Extensive experiments demonstrate that EasySplat outperforms the current state-of-the-art (SOTA) in handling novel view synthesis. Ao Gao, Luosong Guo, Ying Tai, Jian Yang 0003, Zhenyu Zhang 0005 |
ICME | 7 |
| 2025 | Toward Zero-Shot Learning for Visual Dehazing of Urological Surgical RobotsabstractRobot-assisted surgery has profoundly influenced current forms of minimally invasive surgery. However, in transurethral urological surgical robots, they need to work in a liquid environment. This causes vaporization of the liquid when shearing and heating is performed, resulting in bubble atomization that affects the visual perception of the robot. This can lead to the need for uninterrupted pauses in the surgical procedure, which makes the surgery take longer. To address the atomization characteristics of liquids under urological surgical robotic vision, we propose an unsupervised zero-shot dehaze method (RSF-Dehaze). Specifically, the proposed Region Similarity Filling Module (RSFM) of RSF-Dehaze significantly improves the recovery of blurred region tissues. In addition, we organize and propose a dehaze dataset for robotic vision in urological surgery (USRobot-Dehaze dataset). In particular, this dataset contains the three most common urological surgical robot operation scenarios. To the best of our knowledge, we are the first to organize and propose a publicly available dehaze dataset for urological surgical robot vision. The proposed RSF-Dehaze proves the effectiveness of our method in three urological surgical robot operation scenarios with extensive comparative experiments with 20 most classical and advanced dehazing and image recovery algorithms. The proposed source code and dataset are available at https://github.com/wurenkai/RSF-Dehaze. Renkai Wu, Xianjin Wang, Pengchen Liang, Zhenyu Zhang 0005, Qing Chang 0004, Hao Tang 0005 |
ICRA | 4 |
| 2025 | Semantic-Guided Diffusion Model for Single-Step Image Super-ResolutionabstractDiffusion-based image super-resolution (SR) methods have demonstrated remarkable performance. Recent advancements have introduced deterministic sampling processes that reduce inference from 15 iterative steps to a single step, thereby significantly improving the inference speed of existing diffusion models. However, their efficiency remains limited when handling complex semantic regions due to the single-step inference. To address this limitation, we propose SAMSR, a semantic-guided diffusion framework that incorporates semantic segmentation masks into the sampling process. Specifically, we introduce the SAM-Noise Module, which refines Gaussian noise using segmentation masks to preserve spatial and semantic features. Furthermore, we develop a pixel-wise sampling strategy that dynamically adjusts the residual transfer rate and noise strength based on pixel-level semantic weights, prioritizing semantically rich regions during the diffusion process. To enhance model training, we also propose a semantic consistency loss, which aligns pixel-wise semantic weights between predictions and ground truth. Extensive experiments on both real-world and synthetic datasets demonstrate that SAMSR significantly improves perceptual quality and detail recovery, particularly in semantically complex images. Zhenyu Zhang 0005, Hao Tang 0005 |
IJCAI | 2 |
| 2025 | RigNet++: Semantic Assisted Repetitive Image Guided Network for Depth Completion
Zhiqiang Yan 0001, Xiang Li 0041, Le Hui, Zhenyu Zhang 0005, Jun Li 0027, Jian Yang 0003 |
Int. J. Comput. Vis. | 4 |
| 2025 | Invertible Diffusion Models for Compressed SensingabstractWhile deep neural networks (NNs) significantly advance image compressed sensing (CS) by improving reconstruction quality, the necessity of training current CS NNs from scratch constrains their effectiveness and hampers rapid deployment. Although recent methods utilize pre-trained diffusion models for image reconstruction, they struggle with slow inference and restricted adaptability to CS. To tackle these challenges, this paper proposes Invertible Diffusion Models (IDM), a novel efficient, end-to-end diffusion-based CS method. IDM repurposes a large-scale diffusion sampling process as a reconstruction model, and fine-tunes it end-to-end to recover original images directly from CS measurements, moving beyond the traditional paradigm of one-step noise estimation learning. To enable such memory-intensive end-to-end fine-tuning, we propose a novel two-level invertible design to transform both 1) multi-step sampling process and 2) noise estimation U-Net in each step into invertible networks. As a result, most intermediate features are cleared during training to reduce up to 93.8% GPU memory. In addition, we develop a set of lightweight modules to inject measurements into noise estimator to further facilitate reconstruction. Experiments demonstrate that IDM outperforms existing state-of-the-art CS networks by up to 2.64 dB in PSNR. Compared to the recent diffusion-based approach DDNM, our IDM achieves up to 10.09 dB PSNR gain and 14.54 times faster inference. Bin Chen 0006, Zhenyu Zhang 0005, Chen Zhao 0002, Jiwen Yu, Shijie Zhao 0001, Jie Chen 0001, Jian Zhang 0018 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Spatial-Temporal Graph Mamba for Music-Guided Dance Video SynthesisabstractWe propose a novel spatial-temporal graph Mamba (STG-Mamba) for the music-guided dance video synthesis task, i.e., to translate the input music to a dance video. STG-Mamba consists of two translation mappings: music-to-skeleton translation and skeleton-to-video translation. In the music-to-skeleton translation, we introduce a novel spatial-temporal graph Mamba (STGM) block to effectively construct skeleton sequences from the input music, capturing dependencies between joints in both the spatial and temporal dimensions. For the skeleton-to-video translation, we propose a novel self-supervised regularization network to translate the generated skeletons, along with a conditional image, into a dance video. Lastly, we collect a new skeleton-to-video translation dataset from the Internet, containing 54,944 video clips. Extensive experiments demonstrate that STG-Mamba achieves significantly better results than existing methods. Hao Tang 0005, Ling Shao 0001, Zhenyu Zhang 0005, Luc Van Gool, Nicu Sebe |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | AltNeRF: Learning Robust Neural Radiance Field via Alternating Depth-Pose OptimizationabstractNeural Radiance Fields (NeRF) have shown promise in generating realistic novel views from sparse scene images. However, existing NeRF approaches often encounter challenges due to the lack of explicit 3D supervision and imprecise camera poses, resulting in suboptimal outcomes. To tackle these issues, we propose AltNeRF---a novel framework designed to create resilient NeRF representations using self-supervised monocular depth estimation (SMDE) from monocular videos, without relying on known camera poses. SMDE in AltNeRF masterfully learns depth and pose priors to regulate NeRF training. The depth prior enriches NeRF's capacity for precise scene geometry depiction, while the pose prior provides a robust starting point for subsequent pose refinement. Moreover, we introduce an alternating algorithm that harmoniously melds NeRF outputs into SMDE through a consistence-driven mechanism, thus enhancing the integrity of depth priors. This alternation empowers AltNeRF to progressively refine NeRF representations, yielding the synthesis of realistic novel views. Extensive experiments showcase the compelling capabilities of AltNeRF in generating high-fidelity and robust novel views that closely resemble reality. Kun Wang 0042, Zhiqiang Yan 0001, Huang Tian, Zhenyu Zhang 0005, Xiang Li 0041, Jun Li 0027, Jian Yang 0003 |
AAAI | 4 |
| 2024 | Tri-Perspective view Decomposition for Geometry-Aware Depth CompletionabstractDepth completion is a vital taskfor autonomous driving, as it involves reconstructing the precise 3D geometry of a scene from sparse and noisy depth measurements. How-ever, most existing methods either rely only on 2D depth representations or directly incorporate raw 3D point clouds for compensation, which are still insufficient to capture the fine-grained 3D geometry of the scene. To address this chal-lenge, we introduce Tri-Perspective View Decomposition (TPVD), a novel framework that can explicitly model 3D geometry. In particular, (1) TPVD ingeniously decomposes the original point cloud into three 2D views, one of which corresponds to the sparse depth input. (2) We design TPV Fusion to update the 2D TPV features through recurrent 2D-3D-2D aggregation, where a Distance-Aware Spherical Convolution (DASC) is applied. (3) By adaptively choosing TPVaffinitive neighbors, the newly proposed Geometric Spatial Propagation Network (GSPN) further improves the geometric consistency. As a result, our TPVD outperforms existing methods on KITTI, NYUv2, and SUN RGBD. Fur-thermore, we build a novel depth completion dataset named TOFDC, which is acquired by the time-of-flight (TOF) sen-sor and the color camera on smart phones. Project page. Zhiqiang Yan 0001, Yuankai Lin, Kun Wang 0042, Yupeng Zheng, Zhenyu Zhang 0005, Jun Li 0027, Jian Yang 0003 |
CVPR | 6 |
| 2024 | ConsistentAvatar: Learning to Diffuse Fully Consistent Talking Head Avatar with Temporal GuidanceabstractDiffusion models have shown impressive potential on talking head generation. While plausible appearance and talking effect are achieved, these methods still suffer from temporal, 3D or expression inconsistency due to the error accumulation and inherent limitation of single-image generation ability. In this paper, we propose ConsistentAvatar, a novel framework for fully consistent and high-fidelity talking avatar generation. Instead of directly employing multi-modal conditions to the diffusion process, our method learns to first model the temporal representation for stability between adjacent frames. Specifically, we propose a Temporally-Sensitive Detail (TSD) map containing high-frequency feature and contours that vary significantly along the time axis. Using a temporal consistent diffusion module, we learn to align TSD of the initial result to that of the video frame ground truth. The final avatar is generated by a fully consistent diffusion module, conditioned on the aligned TSD, rough head normal, and emotion prompt embedding. We find that the aligned TSD, which represents the temporal patterns, constrains the diffusion process to generate temporally stable talking head. Further, its reliable guidance complements the inaccuracy of other conditions, suppressing the accumulated error while improving the consistency on various aspects. Extensive experiments demonstrate that ConsistentAvatar outperforms the state-of-the-art methods on the generated appearance, 3D, expression and temporal consistency. Haijie Yang, Zhenyu Zhang 0005, Hao Tang 0005, Jianjun Qian, Jian Yang 0003 |
ACM Multimedia | 2 |
| 2024 | Learning to Decouple the Lights for 3D Face Texture ModelingabstractExisting research has made impressive strides in reconstructing human facial shapes and textures from images with well-illuminated faces and minimal external occlusions.
Nevertheless, it remains challenging to recover accurate facial textures from scenarios with complicated illumination affected by external occlusions, \eg a face that is partially obscured by items such as a hat.
Existing works based on the assumption of single and uniform illumination cannot correctly process these data.
In this work, we introduce a novel approach to model 3D facial textures under such unnatural illumination. Instead of assuming single illumination, our framework learns to imitate the unnatural illumination as a composition of multiple separate light conditions combined with learned neural representations, named Light Decoupling.
According to experiments on both single images and video sequences, we demonstrate the effectiveness of our approach in modeling facial textures under challenging illumination affected by occlusions. Tianxin Huang, Zhenyu Zhang 0005, Ying Tai, Gim Hee Lee |
NeurIPS | 2 |
| 2024 | Learning Complementary Correlations for Depth Super-Resolution With Incomplete Data in Real WorldabstractDepth information is a significant ingredient to visually perceive the physical world. However, mainstream depth sensors, e.g., time-of-flight (ToF) cameras, often measure incomplete and low-resolution depth data, resulting in low-quality visual perception. In this article, we try to address a potentially valuable task, i.e., depth super-resolution (DSR) with incomplete data, which recovers dense and high-resolution depth map from incomplete and low-resolution one. To tackle this task, we introduce a novel incomplete DSR (IDSR) framework, including a primary branch for DSR to recover high-frequency details, and an auxiliary branch for depth completion (DC) to fill missing pixels. More importantly, we propose two modules, joint correlation learning (JCL) and iterative-cross (IC), to enhance the learning of complementary information flows between the two branches. The former module aims to learn the correlative relationships of the two branches, whilst the latter module adequately fuses higher level representations for more precise predictions. Extensive experiments show that our framework is effective and achieves the state-of-the-art performance on the real-world RGB-D-D and the synthetic NYUv2 datasets. Zhiqiang Yan 0001, Kun Wang 0042, Xiang Li 0041, Zhenyu Zhang 0005, Jun Li 0027, Jian Yang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | DesNet: Decomposed Scale-Consistent Network for Unsupervised Depth CompletionabstractUnsupervised depth completion aims to recover dense depth from the sparse one without using the ground-truth annotation. Although depth measurement obtained from LiDAR is usually sparse, it contains valid and real distance information, i.e., scale-consistent absolute depth values. Meanwhile, scale-agnostic counterparts seek to estimate relative depth and have achieved impressive performance. To leverage both the inherent characteristics, we thus suggest to model scale-consistent depth upon unsupervised scale-agnostic frameworks. Specifically, we propose the decomposed scale-consistent learning (DSCL) strategy, which disintegrates the absolute depth into relative depth prediction and global scale estimation, contributing to individual learning benefits. But unfortunately, most existing unsupervised scale-agnostic frameworks heavily suffer from depth holes due to the extremely sparse depth input and weak supervisory signal. To tackle this issue, we introduce the global depth guidance (GDG) module, which attentively propagates dense depth reference into the sparse target via novel dense-to-sparse attention. Extensive experiments show the superiority of our method on outdoor KITTI, ranking 1st and outperforming the best KBNet more than 12% in RMSE. Additionally, our approach achieves state-of-the-art performance on indoor NYUv2 benchmark as well. Zhiqiang Yan 0001, Kun Wang 0042, Xiang Li 0041, Zhenyu Zhang 0005, Jun Li 0027, Jian Yang 0003 |
AAAI | 4 |
| 2023 | Graph Transformer GANs for Graph-Constrained House GenerationabstractWe present a novel graph Transformer generative adversarial network (GTGAN) to learn effective graph node relations in an end-to-end fashion for the challenging graph-constrained house generation task. The proposed graph-Transformer-based generator includes a novel graph Transformer encoder that combines graph convolutions and self-attentions in a Transformer to model both local and global interactions across connected and non-connected graph nodes. Specifically, the proposed connected node attention (CNA) and non-connected node attention (NNA) aim to capture the global relations across connected nodes and non-connected nodes in the input graph, respectively. The proposed graph modeling block (GMB) aims to exploit local vertex interactions based on a house layout topology. More-over, we propose a new node classification-based discriminator to preserve the high-level semantic and discriminative node features for different house components. Finally, we propose a novel graph-based cycle-consistency loss that aims at maintaining the relative spatial relationships between ground truth and predicted graphs. Experiments on two challenging graph-constrained house generation tasks (i.e., house layout and roof generation) with two public datasets demonstrate the effectiveness of GTGAN in terms of objective quantitative scores and subjective visual realism. New state-of-the-art results are established by large margins on both tasks. Hao Tang 0005, Zhenyu Zhang 0005, Humphrey Shi, Ling Shao 0001, Nicu Sebe, Radu Timofte, Luc Van Gool |
CVPR | 2 |
| 2023 | Learning Neural Proto-Face Field for Disentangled 3D Face Modeling in the WildabstractGenerative models show good potential for recovering 3D faces beyond limited shape assumptions. While plausible details and resolutions are achieved, these models easily fail under extreme conditions of pose, shadow or appearance, due to the entangled fitting or lack of multi-view priors. To address this problem, this paper presents a novel Neural Proto-face Field (NPF) for unsupervised robust 3D face modeling. Instead of using constrained images as Neural Radiance Field (NeRF), NPF disentangles the common/specific facial cues, i.e., ID, expression and scene-specific details from in-the-wild photo collections. Specifically, NPF learns a face prototype to aggregate 3D-consistent identity via uncertainty modeling, extracting multi-image priors from a photo collection. NPF then learns to deform the prototype with the appropriate facial expressions, constrained by a loss of expression consistency and personal idiosyncrasies. Finally, NPF is optimized to fit a target image in the collection, recovering specific details of appearance and geometry. In this way, the generative model benefits from multi-image priors and meaningful facial structures. Extensive experiments on benchmarks show that NPF recovers superior or competitive facial shapes and textures, compared to state-of-the-art methods. Zhenyu Zhang 0005, Renwang Chen, Weijian Cao, Ying Tai, Chengjie Wang 0001 |
CVPR | 1 |
| 2023 | Learning to Measure the Point Cloud Reconstruction Loss in a Representation SpaceabstractFor point cloud reconstruction-related tasks, the reconstruction losses to evaluate the shape differences between reconstructed results and the ground truths are typically used to train the task networks. Most existing works measure the training loss with point-to-point distance, which may introduce extra defects as predefined matching rules may deviate from the real shape differences. Although some learning-based works have been proposed to overcome the weaknesses of manually-defined rules, they still measure the shape differences in 3D Euclidean space, which may limit their ability to capture defects in reconstructed shapes. In this work, we propose a learning-based Contrastive Adver-sarial Loss (CALoss) to measure the point cloud reconstruction loss dynamically in a non-linear representation space by combining the contrastive constraint with the adversarial strategy. Specifically, we use the contrastive constraint to help CALoss learn a representation space with shape similarity, while we introduce the adversarial strategy to help CALoss mine differences between reconstructed results and ground truths. According to experiments on reconstruction-related tasks, CALoss can help task networks improve re-construction performances and learn more representative representations. Tianxin Huang, Zhonggan Ding, Jiangning Zhang, Ying Tai, Zhenyu Zhang 0005, Mingang Chen, Chengjie Wang 0001, Yong Liu 0007 |
CVPR | 5 |
| 2023 | Learning Versatile 3D Shape Generation with Improved Auto-regressive ModelsabstractAuto-Regressive (AR) models have achieved impressive results in 2D image generation by modeling joint distributions in the grid space. While this approach has been extended to the 3D domain for powerful shape generation, it still has two limitations: expensive computations on volumetric grids and ambiguous auto-regressive order along grid dimensions. To overcome these limitations, we propose the Improved Auto-regressive Model (ImAM) for 3D shape generation, which applies discrete representation learning based on a latent vector instead of volumetric grids. Our approach not only reduces computational costs but also preserves essential geometric details by learning the joint distribution in a more tractable order. Moreover, thanks to the simplicity of our model architecture, we can naturally extend it from unconditional to conditional generation by concatenating various conditioning inputs, such as point clouds, categories, images, and texts. Extensive experiments demonstrate that ImAM can synthesize diverse and faithful shapes of multiple categories, achieving state-of-the-art performance. Simian Luo, Xuelin Qian, Yanwei Fu 0001, Yinda Zhang 0001, Ying Tai, Zhenyu Zhang 0005, Chengjie Wang 0001, Xiangyang Xue 0001 |
ICCV | 6 |
| 2022 | Physically-guided Disentangled Implicit Rendering for 3D Face ModelingabstractThis paper presents a novel Physically-guided Disentangled Implicit Rendering (PhyDIR) framework for highfidelity 3D face modeling. The motivation comes from two observations: Widely-used graphics renderers yield excessive approximations against photo-realistic imaging, while neural rendering methods produce superior appearances but are highly entangled to perceive 3D-aware operations. Hence, we learn to disentangle the implicit rendering via explicit physical guidance, while guaranteeing the properties of: (1) 3D-aware comprehension and (2) high-reality image formation. For the former one, PhyDIR explicitly adopts 3D shading and rasterizing modules to control the renderer, which disentangles the light, facial shape, and viewpoint from neural reasoning. Specifically, PhyDIR proposes a novel multi-image shading strategy to compensate for the monocular limitation, so that the lighting variations are accessible to the neural renderer. For the latter, PhyDIR learns the face-collection implicit texture to avoid ill-posed intrinsic factorization, then leverages a series of consistency losses to constrain the rendering robustness. With the disentangled method, we make 3D face modeling benefit from both kinds of rendering strategies. Extensive experiments on benchmarks show that PhyDIR obtains superior performance than state-of-the-art explicit/implicit methods on geometry/texture modeling. Zhenyu Zhang 0005, Yanhao Ge, Ying Tai, Weijian Cao, Renwang Chen, Kunlin Liu, Hao Tang 0005, Chengjie Wang 0001, Dongjin Huang |
CVPR | 1 |
| 2022 | Learning to Restore 3D Face from In-the-Wild Degraded ImagesabstractIn-the-wild 3D face modelling is a challenging problem as the predicted facial geometry and texture suffer from a lack of reliable clues or priors, when the input images are degraded. To address such a problem, in this paper we propose a novel Learning to Restore (L2R) 3D face framework for unsupervised high-quality face reconstruction from low-resolution images. Rather than directly refining 2D image appearance, L2R learns to recover fine-grained 3D details on the proxy against degradation via extracting generative facial priors. Concretely, L2R proposes a novel albedo restoration network to model high-quality 3D facial texture, in which the diverse guidance from the pre-trained Generative Adversarial Networks (GANs) is leveraged to complement the lack of input facial clues. With the finer details of the restored 3D texture, L2R then learns displacement maps from scratch to enhance the significant facial structure and geometry. Both of the procedures are mutually optimized with a novel 3D-aware adversarial loss, which further improves the modelling performance and suppresses the potential uncertainty. Extensive experiments on benchmarks show that L2R outperforms state-of-the-art methods under the condition of low-quality inputs, and obtains superior performances than 2D pre-processed modelling approaches with limited 3D proxy. Zhenyu Zhang 0005, Yanhao Ge, Ying Tai, Chengjie Wang 0001, Hao Tang 0005, Dongjin Huang |
CVPR | 1 |
| 2022 | Multi-modal Masked Pre-training for Monocular Panoramic Depth Completion
Zhiqiang Yan 0001, Xiang Li 0041, Kun Wang 0042, Zhenyu Zhang 0005, Jun Li 0027, Jian Yang 0003 |
ECCV (1) | 4 |
| 2022 | RigNet: Repetitive Image Guided Network for Depth Completion
Zhiqiang Yan 0001, Kun Wang 0042, Xiang Li 0041, Zhenyu Zhang 0005, Jun Li 0027, Jian Yang 0003 |
ECCV (27) | 4 |
| 2022 | 3QNet: 3D Point Cloud Geometry Quantization Compression NetworkabstractSince the development of 3D applications, the point cloud, as a spatial description easily acquired by sensors, has been widely used in multiple areas such as SLAM and 3D reconstruction. Point Cloud Compression (PCC) has also attracted more attention as a primary step before point cloud transferring and saving, where the geometry compression is an important component of PCC to compress the points geometrical structures. However, existing non-learning-based geometry compression methods are often limited by manually pre-defined compression rules. Though learning-based compression methods can significantly improve the algorithm performances by learning compression rules from data, they still have some defects. Voxel-based compression networks introduce precision errors due to the voxelized operations, while point-based methods may have relatively weak robustness and are mainly designed for sparse point clouds. In this work, we propose a novel learning-based point cloud compression framework named 3D Point Cloud Geometry Quantiation Compression Network (3QNet), which overcomes the robustness limitation of existing point-based methods and can handle dense points. By learning a codebook including common structural features from simple and sparse shapes, 3QNet can efficiently deal with multiple kinds of point clouds. According to experiments on object models, indoor scenes, and outdoor scans, 3QNet can achieve better compression performances than many representative methods. Tianxin Huang, Jiangning Zhang, Jun Chen 0023, Zhonggan Ding, Ying Tai, Zhenyu Zhang 0005, Chengjie Wang 0001, Yong Liu 0007 |
ACM Trans. Graph. | 6 |
| 2021 | Learning To Aggregate and Personalize 3D Face From In-the-Wild Photo CollectionabstractNon-parametric face modeling aims to reconstruct 3D face only from images without shape assumptions. While plausible facial details are predicted, the models tend to over-depend on local color appearance and suffer from ambiguous noise. To address such problem, this paper presents a novel Learning to Aggregate and Personalize (LAP) framework for unsupervised robust 3D face modeling. Instead of using controlled environment, the proposed method implicitly disentangles ID-consistent and scene-specific face from unconstrained photo set. Specifically, to learn ID-consistent face, LAP adaptively aggregates intrinsic face factors of an identity based on a novel curriculum learning approach with relaxed consistency loss. To adapt the face for a personalized scene, we propose a novel attribute-refining network to modify ID-consistent face with target attribute and details. Based on the proposed method, we make unsupervised 3D face modeling benefit from meaningful image facial structure and possibly higher resolutions. Extensive experiments on benchmarks show LAP recovers superior or competitive face shape and texture, compared with state-of-the-art (SOTA) methods with or without prior and supervision. Zhenyu Zhang 0005, Yanhao Ge, Renwang Chen, Ying Tai, Yan Yan 0002, Jian Yang 0003, Chengjie Wang 0001, Feiyue Huang |
CVPR | 1 |
| 2021 | Regularizing Nighttime Weirdness: Efficient Self-supervised Monocular Depth Estimation in the DarkabstractMonocular depth estimation aims at predicting depth from a single image or video. Recently, self-supervised methods draw much attention since they are free of depth annotations and achieve impressive performance on several daytime benchmarks. However, they produce weird outputs in more challenging nighttime scenarios because of low visibility and varying illuminations, which bring weak textures and break brightness-consistency assumption, respectively. To address these problems, in this paper we propose a novel framework with several improvements: (1) we introduce Priors-Based Regularization to learn distribution knowledge from unpaired depth maps and prevent model from being incorrectly trained; (2) we leverage Mapping-Consistent Image Enhancement module to enhance image visibility and contrast while maintaining brightness consistency; and (3) we present Statistics-Based Mask strategy to tune the number of removed pixels within textureless regions, using dynamic statistics. Experimental results demonstrate the effectiveness of each component. Mean-while, our framework achieves remarkable improvements and state-of-the-art results on two nighttime datasets. Code is available at https://github.com/w2kun/RNW. Kun Wang 0042, Zhenyu Zhang 0005, Zhiqiang Yan 0001, Xiang Li 0041, Baobei Xu, Jun Li 0027, Jian Yang 0003 |
ICCV | 2 |
| 2021 | ASFD: Automatic and Scalable Face DetectorabstractAlong with current multi-scale based detectors, Feature Aggregation and Enhancement (FAE) modules have shown superior performance gains for cutting-edge object detection. However, these hand-crafted FAE modules show inconsistent improvements on face detection, which is mainly due to the significant distribution difference between its training and applying corpus, i.e. COCO vs. WIDER Face. To tackle this problem, we essentially analyse the effect of data distribution, and consequently propose to search an effective FAE architecture, termed AutoFAE by a differentiable architecture search, which outperforms all existing FAE modules in face detection with a considerable margin. Upon the found AutoFAE and existing backbones, a supernet is further built and trained, which automatically obtains a family of detectors under the different complexity constraints. Extensive experiments conducted on popular benchmarks, i.e. WIDER Face and FDDB, demonstrate the state-of-the-art performance-efficiency trade-off for the proposed automatic and scalable face detector (ASFD) family. In particular, our strong ASFD-D6 outperforms the best competitor with AP 96.7/96.2/92.1 on WIDER Face test, and the lightweight ASFD-D0 costs about 3.1 ms, i.e. more than 320 FPS, on the V100 GPU with VGA-resolution images. Jian Li 0062, Yabiao Wang, Ying Tai, Zhenyu Zhang 0005, Chengjie Wang 0001, Yili Xia |
ACM Multimedia | 5 |
| 2021 | Discriminative Cross-Modality Attention Network for Temporal Inconsistent Audio-Visual Event LocalizationabstractIt is theoretically insufficient to construct a complete set of semantics in the real world using single-modality data. As a typical application of multi-modality perception, the audio-visual event localization task aims to match audio and visual components to identify the simultaneous events of interest. Although some recent methods have been proposed to deal with this task, they cannot handle the practical situation of temporal inconsistency that is widespread in the audio-visual scene. Inspired by the human system which automatically filters out event-unrelated information when performing multi-modality perception, we propose a discriminative cross-modality attention network to simulate such a process. Similar to human mechanism, our network can adaptively select "where" to attend, "when" to attend and "which" to attend for audio-visual event localization. In addition, to prevent our network from getting trivial solutions, a novel eigenvalue-based objective function is proposed to train the whole network to better fuse audio and visual signals, which can obtain discriminative and nonlinear multi-modality representation. In this way, even with large temporal inconsistency between audio and visual sequence, our network is able to adaptively select event-valuable information for audio-visual event localization. Furthermore, we systemically investigate three subtasks of audio-visual event localization, i.e., temporal localization, weakly-supervised spatial localization and cross-modality localization. The visualization results also help us better understand how our network works. Hanyu Xuan, Lei Luo 0001, Zhenyu Zhang 0005, Jian Yang 0003, Yan Yan 0002 |
IEEE Trans. Image Process. | 3 |
| 2020 | Cross-Modal Attention Network for Temporal Inconsistent Audio-Visual Event LocalizationabstractIn human multi-modality perception systems, the benefits of integrating auditory and visual information are extensive as they provide plenty supplementary cues for understanding the events. Despite some recent methods proposed for such application, they cannot deal with practical conditions with temporal inconsistency. Inspired by human system which puts different focuses at specific locations, time segments and media while performing multi-modality perception, we provide an attention-based method to simulate such process. Similar to human mechanism, our network can adaptively select “where” to attend, “when” to attend and “which” to attend for audio-visual event localization. In this way, even with large temporal inconsistent between vision and audio, our network is able to adaptively trade information between different modalities and successfully achieve event localization. Our method achieves state-of-the-art performance on AVE (Audio-Visual Event) dataset collected in the real life. In addition, we also systemically investigate audio-visual event localization tasks. The visualization results also help us better understand how our model works. Hanyu Xuan, Zhenyu Zhang 0005, Shuo Chen 0003, Jian Yang 0003, Yan Yan 0002 |
AAAI | 2 |
| 2020 | Online Depth Learning Against Forgetting in Monocular VideosabstractOnline depth learning is the problem of consistently adapting a depth estimation model to handle a continuously changing environment. This problem is challenging due to the network easily overfits on the current environment and forgets its past experiences. To address such problem, this paper presents a novel Learning to Prevent Forgetting (LPF) method for online mono-depth adaptation to new target domains in unsupervised manner. Instead of updating the universal parameters, LPF learns adapter modules to efficiently adjust the feature representation and distribution without losing the pre-learned knowledge in online condition. Specifically, to adapt temporal-continuous depth patterns in videos, we introduce a novel meta-learning approach to learn adapter modules by combining online adaptation process into the learning objective. To further avoid overfitting, we propose a novel temporal-consistent regularization to harmonize the gradient descent procedure at each online learning step. Extensive evaluations on real-world datasets demonstrate that the proposed method, with very limited parameters, significantly improves the estimation quality. Zhenyu Zhang 0005, Stéphane Lathuilière, Elisa Ricci 0001, Nicu Sebe, Yan Yan 0002, Jian Yang 0003 |
CVPR | 1 |
| 2020 | Pattern-Structure Diffusion for Multi-Task LearningabstractInspired by the observation that pattern structures high-frequently recur within intra-task also across tasks, we propose a pattern-structure diffusion (PSD) framework to mine and propagate task-specific and task-across pattern structures in the task-level space for joint depth estimation, segmentation and surface normal prediction. To represent local pattern structures, we model them as small-scale graphlets, and propagate them in two different ways, i.e., intra-task and inter-task PSD. For the former, to overcome the limit of the locality of pattern structures, we use the high-order recursive aggregation on neighbors to multiplicatively increase the spread scope, so that long-distance patterns are propagated in the intra-task space. In the inter-task PSD, we mutually transfer the counterpart structures corresponding to the same spatial position into the task itself based on the matching degree of paired pattern structures therein. Finally, the intra-task and inter-task pattern structures are jointly diffused among the task-level patterns, and encapsulated into an end-to-end PSD network to boost the performance of multi-task learning. Extensive experiments on two widely-used benchmarks demonstrate that our proposed PSD is more effective and also achieves the state-of-the-art or competitive results. Zhen Cui 0001, Chunyan Xu, Zhenyu Zhang 0005, Chaoqun Wang 0012, Tong Zhang 0021, Jian Yang 0003 |
CVPR | 4 |
| 2020 | Local-Global Feature for Video-Based One-Shot Person Re-IdentificationabstractOne-shot video-based re-identification, which uses only one labeled tracklet for each identity, is challenging since the framework usually suffers misalignment and inefficient utilizing of unlabeled data. In this paper we propose a novel local-global progressive learning framework to overcome the limitations. To obtain robust features in a tracklet, we first design sub-networks to learn four discriminative part-based feature maps and one global feature map which is insensitive to misalignment. Then a novel adaptive loss is proposed to balance the part-based and global feature properly. To utilize unlabeled data, our framework gradually select most reliable pseudo-labeled tracklets to the training set for iterative training. Extensive experiments are conducted on two video-based Re-ID datasets, MARS and DukeMTMC-VideoReID. The mAP of our model outperforms the state-of-the-art methods by 20.8% on the DukeMTMC-VideoReID dataset. Zhenyu Zhang 0005, Jian Yan 0009, Yan Yan 0002 |
ICASSP | 2 |
| 2020 | Decoupled Self-attention Module for Person Re-identificationabstractPerson re-identification aims to identify the same person from different cameras, which needs to integrate whole-body information and capture global correlation. However, convolutional neural network is able to only capture short-distance information because of the size of filters. Self-attention is introduced to capture long-distance correlation, but inner-product similarity calculation in self-attention mingles semantic response and semantic difference together. Semantic difference is more important for person re-identification, because it is robust to illumination without the effect of semantic response. However, we find the scale of norms measuring semantic response is much larger than angle measuring semantic difference by decoupling inner-product similarity into norms and angle. To balance the importance of semantic response and semantic difference in self-attention, we propose the decoupled self-attention module for person re-identification to make the most of self-attention. Extensive experiments show that the decoupled self-attention module obtains significant performance with easier convergence and stronger robustness. Zhenyu Zhang 0005, Jian Yan 0009, Yan Yan 0002 |
ICPR | 2 |
| 2020 | Joint Task-Recursive Learning for RGB-D Scene UnderstandingabstractRGB-D scene understanding under monocular camera is an emerging and challenging topic with many potential applications. In this paper, we propose a novel Task-Recursive Learning (TRL) framework to jointly and recurrently conduct three representative tasks therein containing depth estimation, surface normal prediction and semantic segmentation. TRL recursively refines the prediction results through a series of task-level interactions, where one-time cross-task interaction is abstracted as one network block of one time stage. In each stage, we serialize multiple tasks into a sequence and then recursively perform their interactions. To adaptively enhance counterpart patterns, we encapsulate interactions into a specific Task-Attentional Module (TAM) to mutually-boost the tasks from each other. Across stages, the historical experiences of previous states of tasks are selectively propagated into the next stages by using Feature-Selection unit (FS-Unit), which takes advantage of complementary information across tasks. The sequence of task-level interactions is also evolved along a coarse-to-fine scale space such that the required details may be refined progressively. Finally the task-abstracted sequence problem of multi-task prediction is framed into a recursive network. Extensive experiments on NYU-Depth v2 and SUN RGB-D datasets demonstrate that our method can recursively refines the results of the triple tasks and achieves state-of-the-art performance. Zhenyu Zhang 0005, Zhen Cui 0001, Chunyan Xu, Zequn Jie, Xiang Li 0041, Jian Yang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | Pattern-Affinitive Propagation Across Depth, Surface Normal and Semantic SegmentationabstractIn this paper, we propose a novel Pattern-Affinitive Propagation (PAP) framework to jointly predict depth, surface normal and semantic segmentation. The motivation behind it comes from the statistic observation that pattern-affinitive pairs recur much frequently across different tasks as well as within a task. Thus, we can conduct two types of propagations, cross-task propagation and task-specific propagation, to adaptively diffuse those similar patterns. The former integrates cross-task affinity patterns to adapt to each task therein through the calculation on non-local relationships. Next the latter performs an iterative diffusion in the feature space so that the cross-task affinity patterns can be widely-spread within the task. Accordingly, the learning of each task can be regularized and boosted by the complementary task-level affinities. Extensive experiments demonstrate the effectiveness and the superiority of our method on the joint three tasks. Meanwhile, we achieve the state-of-the-art or competitive results on the three related datasets, NYUD-v2, SUN-RGBD and KITTI. Zhenyu Zhang 0005, Zhen Cui 0001, Chunyan Xu, Yan Yan 0002, Nicu Sebe, Jian Yang 0003 |
CVPR | 1 |
| 2019 | LRR for Subspace Segmentation via Tractable Schatten- $p$ Norm Minimization and FactorizationabstractRecently, nuclear norm-based low rank representation (LRR) methods have been popular in several applications, such as subspace segmentation. However, there exist two limitations: one is that nuclear norm as the relaxation of rank function will lead to the suboptimal solution since nuclear norm-based minimization subproblem tends to the over-relaxations of singular value elements and treats each of them equally; the other is that solving LRR problems may cause more time consumption due to involving singular value decomposition of the large scale matrix at each iteration. To overcome both disadvantages, this paper mainly considers two tractable variants of LRR: one is Schatten-p norm minimization-based LRR (i.e., SpNM_LRR) and the other is Schatten-p norm factorization-based LRR (i.e., SpNFLRR) for p=1, 2/3 and 1/2. By introducing two or more auxiliary variables in the constraints, the alternating direction method of multiplier (ADMM) with multiple updating variables can be devised to solve these variants of LRR. Furthermore, both computational complexity and convergence property are given to evaluate nonconvex multiblocks ADMM algorithms. Several experiments finally validate the efficacy and efficiency of our methods on both synthetic data and real world data. Hengmin Zhang, Jian Yang 0003, Fanhua Shang, Chen Gong 0002, Zhenyu Zhang 0005 |
IEEE Trans. Cybern. | 5 |
| 2018 | Joint Task-Recursive Learning for Semantic Segmentation and Depth Estimation
Zhenyu Zhang 0005, Zhen Cui 0001, Chunyan Xu, Zequn Jie, Xiang Li 0041, Jian Yang 0003 |
ECCV (10) | 1 |
| 2018 | Deep hierarchical guidance and regularization learning for end-to-end depth estimation
Zhenyu Zhang 0005, Chunyan Xu, Jian Yang 0003, Ying Tai, Liang Chen 0003 |
Pattern Recognit. | 1 |
| 2018 | Progressive Hard-Mining Network for Monocular Depth EstimationabstractDepth estimation from the monocular RGB image is a challenging task for computer vision due to no reliable cues as the prior knowledge. Most existing monocular depth estimation works including various geometric or network learning methods lack of an effective mechanism to preserve the cross-border details of depth maps, which yet is very important for the performance promotion. In this paper, we propose a novel end-to-end progressive hard-mining network (PHN) framework to address this problem. Specifically, we construct the hard-mining objective function, the intra-scale and inter-scale refinement subnetworks to accurately localize and refine those hard-mining regions. The intra-scale refining block recursively recovers details of depth maps from different semantic features in the same receptive field while the inter-scale block favors a complementary interaction among multi-scale depth cues of different receptive fields. For further reducing the uncertainty of the network, we design a difficulty-ware refinement loss function to guide the depth learning process, which can adaptively focus on mining these hard-regions where accumulated errors easily occur. All three modules collaborate together to progressively reduce the error propagation in the depth learning process, and then, boost the performance of monocular depth estimation to some extent. We conduct comprehensive evaluations on several public benchmark data sets (including NYU Depth V2, KITTI, and Make3D). The experiment results well demonstrate the superiority of our proposed PHN framework over other state of the arts for monocular depth estimation task. Zhenyu Zhang 0005, Chunyan Xu, Jian Yang 0003, Junbin Gao, Zhen Cui 0001 |
IEEE Trans. Image Process. | 1 |