VLDB 2026 Research / reviewers in the wild / expert
Yongwei Nie
dblp:31/8005
· DBLP profile ↗
79ranked-venue papers
15as first author
49since 2021 · last 2026
0000-0002-8922-3205ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 66 · 13 first-author · 38 since 2021Artificial intelligence and machine learning · 23 · 4 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Large model-assisted video summarization via global entity unification and robust importance scoring
Donglei Chen, Shaoyu Huang, Xuemiao Xu, Yongwei Nie, Ping Li 0016, C. L. Philip Chen |
Vis. Comput. | 4 |
| 2025 | Occlusion-Insensitive Talking Head Video Generation via Facelet CompensationabstractTalking head video generation involves animating a still face image using facial motion cues derived from a driving video to replicate target poses and expressions. Traditional methods often rely on the assumption that the relative positions of facial keypoints remain unchanged. However, this assumption fails when keypoints are occluded or when the head is in a profile pose, leading to inconsistencies in identity and blurring in certain facial regions. In this paper, we introduce Occlusion-Insensitive Talking Head Video Generation, a novel approach that eliminates the reliance on spatial correlation of keypoints and instead leverages semantic correlation. Our method transforms facial features into a facelet semantic bank, where each facelet token represents a specific facial semantic. This bank is devoid of spatial information, allowing it to compensate for any invisible or occluded face regions during motion warping. The facelet compensation module then populates the facelet tokens within the initially warped features by learning a correlation matrix between facial semantics and the facelet bank. This approach enables precise compensation for occlusions and pose changes, enhancing the fidelity of the generated videos. Extensive experiments demonstrate that our method achieves state-of-the-art results, preserving source identity, maintaining fine-grained facial details, and capturing nuanced facial expressions with remarkable accuracy. Yuhui Deng 0005, Yuqin Lu, Yangyang Xu 0003, Yongwei Nie, Shengfeng He |
AAAI | 4 |
| 2025 | When Shadow Removal Meets Intrinsic Image Decomposition: A Joint Learning Framework Using Unpaired DataabstractWe present a framework that achieves shadow removal by learning intrinsic image decomposition (IID) from unpaired shadow and shadow-free images. Although it is well-known that intrinsic images, \ie, illumination and reflectance, are highly beneficial to shadow removal, IID is rarely adopted by previous work due to its inherent ambiguity and the scarcity of training data. However, we find that by properly coupling shadow removal and IID into a joint learning framework, they can reinforce each other and enable promising results on both tasks, even with unpaired training data. Our framework is comprised of an IID network for separating the shadow input image into illumination and reflectance, and an illumination recovery network for predicting shadow-free illumination with which we are able to produce the shadow removal output by recombining with the estimated reflectance. We perform extensive experiments on various benchmark datasets to demonstrate the effectiveness of our method in shadow removal, and also showcase our advantage over previous IID methods in handling images with complex shadows. Rongjia Zheng, Qing Zhang 0006, Yongwei Nie, Wei-Shi Zheng 0001 |
AAAI | 3 |
| 2025 | Training-Free Language-Guided Video Summarization via Multi-Grained Saliency Scoring
Yongwei Nie, Fei Ma 0006, Keke Tang, F. Richard Yu, Hongmin Cai, Ping Li 0016 |
CVM (3) | 2 |
| 2025 | Simplification Is All You Need against Out-of-Distribution OverconfidenceabstractDeep neural networks (DNNs) often exhibit out-of-distribution (OOD) overconfidence, producing overly confident predictions on OOD samples. We attribute this issue to the inherent over-complexity of DNNs and investigate two key aspects: capacity and nonlinearity. First, we demonstrate that reducing model capacity through knowledge distillation can effectively mitigate OOD overconfidence. Second, we show that selectively reducing nonlinearity by removing ReLU operations further alleviates the issue. Building on these findings, we present a practical guide to model simplification, combining both strategies to significantly reduce OOD overconfidence. Extensive experiments validate the effectiveness of this approach in mitigating OOD overconfidence and demonstrate its superiority over state-of-the-art methods. Additionally, our simplification strategies can be combined with existing OOD detection techniques to further enhance OOD detection performance. Keke Tang, Weilong Peng, Zhize Wu, Yongwei Nie, Wenping Wang 0001, Zhihong Tian 0001 |
CVPR | 6 |
| 2025 | EntityErasure: Erasing Entity Cleanly via Amodal Entity Segmentation and CompletionabstractThis paper presents EntityErasure, a novel diffusion-based inpainting method that can effectively erase entities without inducing unwanted sundries. To this end, we propose to address this problem by dividing it into amodal entity segmentation and completion, such that the region to inpaint takes only entities in the non-inpainting area as reference, avoiding the possibility to generate unpredictable sundries. Moreover, we develop two entity segmentation based metrics for quantitatively assessing the performance of object erasure, which are shown be more effective than existing metrics. Experimental results demonstrate that our approach outperforms other state-of-the-art object erasure methods. Our code and data are available at https://zyxunh.github.io/EntityErasure-ProjectPage/. Yixing Zhu, Qing Zhang 0006, Yongwei Nie, Wei-Shi Zheng 0001 |
CVPR | 4 |
| 2025 | Attribute-Based Out-of-Distribution Detection Using LLaVA
Daojie Zhao, Yongwei Nie, Peican Zhu, Keke Tang |
ICIC (19) | 3 |
| 2025 | RecDreamer: Consistent Text-to-3D Generation via Uniform Score DistillationabstractCurrent text-to-3D generation methods based on score distillation often suffer from geometric inconsistencies, leading to repeated patterns across different poses of 3D assets. This issue, known as the Multi-Face Janus problem, arises because existing methods struggle to maintain consistency across varying poses and are biased toward a canonical pose. While recent work has improved pose control and approximation, these efforts are still limited by this inherent bias, which skews the guidance during generation.
To address this, we propose a solution called RecDreamer, which reshapes the underlying data distribution to achieve more consistent pose representation. The core idea behind our method is to rectify the prior distribution, ensuring that pose variation is uniformly distributed rather than biased toward a canonical form. By modifying the prescribed distribution through an auxiliary function, we can reconstruct the density of the distribution to ensure compliance with specific marginal constraints. In particular, we ensure that the marginal distribution of poses follows a uniform distribution, thereby eliminating the biases introduced by the prior knowledge.
We incorporate this rectified data distribution into existing score distillation algorithms, a process we refer to as uniform score distillation. To efficiently compute the posterior distribution required for the auxiliary function, RecDreamer introduces a training-free classifier that estimates pose categories in a plug-and-play manner. Additionally, we utilize various approximation techniques for noisy states, significantly improving system performance.
Our experimental results demonstrate that RecDreamer effectively mitigates the Multi-Face Janus problem, leading to more consistent 3D asset generation across different poses. Chenxi Zheng, Yihong Lin, Bangzhen Liu, Xuemiao Xu, Yongwei Nie, Shengfeng He |
ICLR | 5 |
| 2025 | EOOD: Entropy-based Out-of-distribution DetectionabstractDeep neural networks (DNNs) often exhibit overconfidence when encountering out-of-distribution (OOD) samples, posing significant challenges for deployment. Since DNNs are trained on in-distribution (ID) datasets, the information flow of ID samples through DNNs inevitably differs from that of OOD samples. In this paper, we propose an Entropy-based Out-Of-distribution Detection (EOOD) framework. EOOD first identifies specific block where the information flow differences between ID and OOD samples are more pronounced, using both ID and pseudo-OOD samples. It then calculates the conditional entropy on the selected block as the OOD confidence score. Comprehensive experiments conducted across various ID and OOD settings demonstrate the effectiveness of EOOD in OOD detection and its superiority over state-of-the-art methods. Guide Yang, Weilong Peng, Yongwei Nie, Peican Zhu, Keke Tang |
IJCNN | 5 |
| 2025 | Registration is a Powerful Rotation-Invariance Learner for 3D Anomaly Detectionabstract3D anomaly detection in point-cloud data is critical for industrial quality control, aiming to identify structural defects with high reliability. However, current memory bank-based methods often suffer from inconsistent feature transformations and limited discriminative capacity, particularly in capturing local geometric details and achieving rotation invariance. These limitations become more pronounced when registration fails, leading to unreliable detection results. We argue that point-cloud registration plays an essential role not only in aligning geometric structures but also in guiding feature extraction toward rotation-invariant and locally discriminative representations. To this end, we propose a registration-induced, rotation-invariant feature extraction framework that integrates the objectives of point-cloud registration and memory-based anomaly detection. Our key insight is that both tasks rely on modeling local geometric structures and leveraging feature similarity across samples. By embedding feature extraction into the registration learning process, our framework jointly optimizes alignment and representation learning. This integration enables the network to acquire features that are both robust to rotations and highly effective for anomaly detection. Extensive experiments on the Anomaly-ShapeNet and Real3D-AD datasets demonstrate that our method consistently outperforms existing approaches in effectiveness and generalizability. Yuyang Yu, Zhengwei Chen, Xuemiao Xu, Lei Zhang 0006, Haoxin Yang, Yongwei Nie, Shengfeng He |
NeurIPS | 6 |
| 2025 | Implicit-based collision-aware clothed human reconstruction from a single image
Guiqing Li, Yongwei Nie, Feiran Yu, Ping Li 0016, Tonglai Liu, Zhao Zhang 0001 |
Comput. Graph. | 3 |
| 2025 | Occlusion-Preserved Surveillance Video Synopsis with Flexible Object Graph
Yongwei Nie, Siming Zeng, Qing Zhang 0006, Guiqing Li, Ping Li 0016, Hongmin Cai |
Int. J. Comput. Vis. | 1 |
| 2025 | Gaussian Prompter: Linking 2D Prompts for 3D Gaussian SegmentationabstractInteractive 3D segmentation in radiance fields is crucial for advanced 3D scene understanding and manipulation. However, existing methods often struggle to achieve both volumetric completeness and segmentation accuracy, primarily because they fail to consider the critical links between 2D prompt-based segmentations across multiple views. Motivated by this gap, we introduce Gaussian Prompter, a novel approach specifically designed for 3D Gaussian Splatting. The core idea behind Gaussian Prompter is to seamlessly integrate a Gaussian-centric segmentation paradigm by effectively linking various 2D prompts from multi-view segmentations to ensure consistent 3D segmentation. To realize this, we employ two tailored approaches: GaussBlend and PinPrompt. GaussBlend aggregates multi-view 2D segmentation masks into a cohesive 3D segmentation, ensuring both accuracy and completeness. PinPrompt leverages high-confidence prompts from adjacent views to enhance segmentation precision further. Additionally, to address the lack of complex datasets in 3D segmentation, we introduce the SegMip-360 dataset, which includes over 350 precisely annotated masks across seven scenes. Extensive experiments demonstrate that the Gaussian Prompter significantly outperforms state-of-the-art methods in both segmentation accuracy and completeness. Our code and video demonstrations can be found at our repository and project page. Honghan Pan, Bangzhen Liu, Xuemiao Xu, Chenxi Zheng, Yongwei Nie, Shengfeng He |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | SGG-Nets: Generic Rotation-Invariant Plugin Networks for Point Cloud AnalysisabstractRotation invariance is a crucial requirement for the analysis of 3D point clouds. However, current methods often achieve rotation invariance by employing specific network designs. These networks, though perform well on rotation-aware tasks, is inferior in general tasks such as classification and segmentation. On the other hand, many powerful point processing networks, such as PointNet++, DGCNN, etc., have general point processing abilities, but do not own the property of rotation invariance. In this paper, we propose a standalone rotation-invariant convolution operator called SGGConv (Spherical Geometric Graph-based Convolution) and two ways integrating it with common point-based networks. The networks equipped with SGGConvs are called SGG-Nets which promote the rotation-invariance ability of regular point networks without modifying their network architectures much. Our contributions are three-fold. First, we propose a rotation-invariant feature descriptor, namely Spherical Geometry Descriptor (SGD), which captures point-pair features in a Local Spherical Coordinate System (LSCS). Second, we propose the SGGConv based on SGD and LSCS with an efficient Graph-based Spherical Feature Passing (GSFP) mechanism. Thirdly, we define two modules S-SGGConvMdl and M-SGGConvMdl, which are used to integrate SGGConv into baseline point nets. We test SGG-Nets, such as SGG-PointNet++, SGG-DGCNN, SGG-RIConv++, on representative point cloud datasets. These models, equipped with our SGGConvs, not only enhance the rotation-invariance of the baseline network but also improve its performance on point cloud analysis tasks such as classification and part segmentation, without incurring too much computational overhead. Jian Zhu 0001, Jianrong Yan, Jiebin Huang, Yongwei Nie, Bin Sheng 0001, Tong-Yee Lee |
IEEE Trans. Multim. | 4 |
| 2025 | Modality-Aware Discriminative Fusion Network for Integrated Analysis of Brain Imaging GenomicsabstractMild cognitive impairment (MCI) represents an early stage of Alzheimer's disease (AD), characterized by subtle clinical symptoms that pose challenges for accurate diagnosis. The quest for the identification of MCI individuals has highlighted the importance of comprehending the underlying mechanisms of disease causation. Integrated analysis of brain imaging and genomics offers a promising avenue for predicting MCI risk before clinical symptom onset. However, most existing methods face challenges in: 1) mining the brain network-specific topological structure and addressing the single nucleotide polymorphisms (SNPs)-related noise contamination and 2) extracting the discriminative properties of brain imaging genomics, resulting in limited accuracy for MCI diagnosis. To this end, a modality-aware discriminative fusion network (MA-DFN) is proposed to integrate the complementary information from brain imaging genomics to diagnose MCI. Specifically, we first design two modality-specific feature extraction modules: the graph convolutional network with edge-augmented self-attention module (GCN-EASA) and the deep adversarial denoising autoencoder module (DAD-AE), to capture the topological structure of brain networks and the intrinsic distribution of SNPs. Subsequently, a discriminative-enhanced fusion network with correlation regularization module (DFN-CorrReg) is employed to enhance inter-modal consistency and between-class discrimination in brain imaging and genomics. Compared to other state-of-the-art approaches, MA-DFN not only exhibits superior performance in stratifying cognitive normal (CN) and MCI individuals but also identifies disease-related brain regions and risk SNPs locus, which hold potential as putative biomarkers for MCI diagnosis. Xiaoqi Sheng, Hongmin Cai, Yongwei Nie, Shengfeng He, Yiu-Ming Cheung, Jiazhou Chen 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Self-supervised Texture FilteringabstractDecomposing an image I into the combination of structure S and texture T components is an important problem in computational photography and image analysis. Traditional solutions are basically non-learning based, because it is difficult to construct datasets containing ground-truth decompositions or find effective structure/texture supervisions. In this article, we present a self-supervised framework for smoothing out textures while maintaining the image structures. At the core of our method is a texture-inversion observation — if structure S and texture T are well disentangled, then S-T will produce a texture-inverted image that is symmetric to the input image I=S+T and the two will be visually highly similar, while for other conditions that structure and texture are not effectively separated, the generated texture-inverted images will be less similar to the input. Based on the observation, we propose to learn texture filtering from unlabeled data by encouraging the texture inverted image generated from the filtering output to be visually more similar to the input via contrastive learning. Experiments show that our method can robustly produce high-quality texture smoothing results, and also enables various applications. Hao Jiang 0057, Rongjia Zheng, Yongwei Nie, Chunxia Xiao, Wei-Shi Zheng 0001, Qing Zhang 0006 |
ACM Trans. Graph. | 3 |
| 2025 | Single-Image SVBRDF Estimation Using Auxiliary Renderings as Intermediate TargetsabstractRecently, single-image SVBRDF capture is formulated as a regression problem, which uses a network to infer four SVBRDF maps from a flash-lit image. However, the accuracy is still not satisfactory since previous approaches usually adopt end-to-end inference strategies. To mitigate the challenge, we propose "auxiliary renderings" as the intermediate regression targets, through which we divide the original end-to-end regression task into several easier sub-tasks, thus achieving better inference accuracy. Our contributions are threefold. First, we design three (or two pairs of) auxiliary renderings and summarize the motivations behind the designs. By our design, the auxiliary images are bumpiness-flattened or highlight-removed, containing disentangled visual cues about the final SVBRDF maps and can be easily transformed to the final maps. Second, to help estimate the auxiliary targets from the input image, we propose two mask images including a bumpiness mask and a highlight mask. Our method thus first infers mask images, then with the help of the mask images infers auxiliary renderings, and finally transforms the auxiliary images to SVBRDF maps. Third, we propose backbone UNets to infer mask images, and gated deformable UNets for estimating auxiliary targets. Thanks to the well-designed networks and intermediate images, our method outputs better SVBRDF maps than previous approaches, validated by the extensive comparisonal and ablation experiments. Yongwei Nie, Chengjiang Long, Qing Zhang 0006, Guiqing Li, Hongmin Cai |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2025 | Towards Photorealistic Portrait Style Transfer in Unconstrained ConditionsabstractWe present a photorealistic portrait style transfer approach that allows for producing high-quality results in previously challenging unconstrained conditions, e.g., large facial perspective difference between portraits, faces with complex illumination (e.g., shadow and highlight) and occlusion, and can test without portrait parsing masks. We achieve this by developing a framework to learn robust dense correspondence across portraits for semantically aligned style transfer, where a regional style contrastive learning strategy is devised to boost the effectiveness of semantic-aware style transfer while enhancing the robustness to complex illumination. Extensive experiments demonstrate the superiority of our method. Xinbo Wang, Qing Zhang 0006, Yongwei Nie, Wei-Shi Zheng 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | FunduSAM: A Specialized Deep Learning Model for Enhanced Optic Disc and Cup Segmentation in Fundus ImagesabstractThe Segment Anything Model (SAM) has gained popularity as a versatile image segmentation method, thanks to its strong generalization capabilities across various domains. However, when applied to optic disc (OD) and optic cup (OC) segmentation tasks, SAM encounters challenges due to the complex structures, low contrast, and blurred boundaries typical of fundus images, leading to suboptimal performance. To over-come these challenges, we introduce a novel model, FunduSAM, which incorporates several Adapters into SAM to create a deep network specifically designed for OD and OC segmentation. The FunduSAM utilizes Adapter into each transformer block after encoder for parameter fine-tuning (PEFT). It enhances SAM’s feature extraction capabilities by designing a Convolutional Block Attention Module (CBAM), addressing issues related to blurred boundaries and low contrast. Given the unique requirements of OD and OC segmentation, polar transformation is used to convert the original fundus OD images into a format better suited for training and evaluating FunduSAM. A joint loss is used to achieve structure preservation between the OD and OC, while accurate segmentation. Extensive experiments on the REFUGE dataset, comprising 1,200 fundus images, demonstrate the superior performance of FunduSAM compared to five mainstream approaches. Jinchen Yu, Yongwei Nie, Fei Qi 0007, Wenxiong Liao, Hongmin Cai |
BIBM | 2 |
| 2024 | Face Expression Recognition via Product-Cross Dual Attention and Neutral-Aware Anchor Loss
Yongwei Nie, Qing Zhang 0006, Xuemiao Xu, Guiqing Li, Hongmin Cai |
CVM (2) | 1 |
| 2024 | Interleaving One-Class and Weakly-Supervised Models with Adaptive Thresholding for Unsupervised Video Anomaly Detection
Yongwei Nie, Chengjiang Long, Qing Zhang 0006, Pradipta Maji, Hongmin Cai |
ECCV (30) | 1 |
| 2024 | Multi-RoI Human Mesh Recovery with Camera Consistency and Contrastive Losses
Yongwei Nie, Changzhen Liu, Chengjiang Long, Qing Zhang 0006, Guiqing Li, Hongmin Cai |
ECCV (47) | 1 |
| 2024 | Incorporating Test-Time Optimization into Training with Dual Networks for Human Mesh RecoveryabstractHuman Mesh Recovery (HMR) is the task of estimating a parameterized 3D human mesh from an image. There is a kind of methods first training a regression model for this problem, then further optimizing the pretrained regression model for any specific sample individually at test time. However, the pretrained model may not provide an ideal optimization starting point for the test-time optimization. Inspired by meta-learning, we incorporate the test-time optimization into training, performing a step of test-time optimization for each sample in the training batch before really conducting the training optimization over all the training samples. In this way, we obtain a meta-model, the meta-parameter of which is friendly to the test-time optimization. At test time, after several test-time optimization steps starting from the meta-parameter, we obtain much higher HMR accuracy than the test-time optimization starting from the simply pretrained regression model. Furthermore, we find test-time HMR objectives are different from training-time objectives, which reduces the effectiveness of the learning of the meta-model. To solve this problem, we propose a dual-network architecture that unifies the training-time and test-time objectives. Our method, armed with meta-learning and the dual networks, outperforms state-of-the-art regression-based and optimization-based HMR approaches, as validated by the extensive experiments. The codes are available at https://github.com/fmx789/Meta-HMR. Yongwei Nie, Mingxian Fan, Chengjiang Long, Qing Zhang 0006, Jian Zhu 0001, Xuemiao Xu |
NeurIPS | 1 |
| 2024 | Make static person walk again via separating pose action from shapeabstractThis paper addresses the problem of animating a person in static images, the core task of which is to infer future poses for the person. Existing approaches predict future poses in the 2D space, suffering from entanglement of pose action and shape. We propose a method that generates actions in the 3D space and then transfers them to the 2D person. We first lift the 2D pose of the person to a 3D skeleton, then propose a 3D action synthesis network predicting future skeletons, and finally devise a self-supervised action transfer network that transfers the actions of 3D skeletons to the 2D person. Actions generated in the 3D space look plausible and vivid. More importantly, self-supervised action transfer allows our method to be trained only on a 3D MoCap dataset while being able to process images in different domains. Experiments on three image datasets validate the effectiveness of our method. Yongwei Nie, Meihua Zhao, Qing Zhang 0006, Ping Li 0016, Jian Zhu 0001, Hongmin Cai |
Graph. Model. | 1 |
| 2024 | Spatial attention for human-centric visual understanding: An Information Bottleneck method
Qiuxia Lai, Yongwei Nie, Yu Li 0007, Hanqiu Sun, Qiang Xu 0001 |
Comput. Vis. Image Underst. | 2 |
| 2024 | PatchMixing Masked Autoencoders for 3D Point Cloud Self-Supervised LearningabstractRecently, Point-MAE has extended Masked Autoencoders (MAE) to point clouds for 3D self-supervised learning, which however faces two problems: (1) the shape similarity between the masked point cloud and original point cloud is high, and (2) the pretext task of reconstructing the original point cloud is straightforward which fails to compel the network to learn deep representative features. In this paper, we tackle these problems by proposing a PatchMixing strategy and a teacher-student training framework. First, with PatchMixing, we mix selected point patches of multiple point clouds and attempt to infer the object information from the resulting mixed point cloud. Due to the interference of other objects, the task is challenging but facilitates representation learning. Second, rather than directly restoring the original point cloud, we propose a novel pretext task that involves a two-branch teacher model and a student model. These models process the multiple input point clouds in different ways (no mixing, mixing + unmixing, mixing + masking), but are expected to output similar features, thereby compelling the network to extract essential features from the input. Extensive experiments show that our well-designed PatchMixing strategy and effective teacher-student learning architecture yield impressive results. Specifically, our model achieves a remarkable 92.9% classification accuracy in the Linear SVM task on the ModelNet40 dataset. Through pre-training and fine-tuning on downstream tasks, our method achieves an 89.8% classification accuracy on the most challenging split of ScanObjectNN and an outstanding 94.0% on ModelNet40. Chengxing Lin 0001, Wenju Xu, Jian Zhu 0001, Yongwei Nie, Ruichu Cai, Xuemiao Xu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Disentangled Representation Learning for Controllable Person Image GenerationabstractIn this paper, we propose a novel framework named DRL-CPG to learn disentangled latent representation for controllable person image generation, which can produce realistic person images with desired poses and human attributes (e.g. pose, head, upper clothes, and pants) provided by various source persons. Unlike the existing works leveraging the semantic masks to obtain the representation of each component, we propose to generate disentangled latent code via a novel attribute encoder with transformers trained in a manner of curriculum learning from a relatively easy step to a gradually hard one. A random component mask-agnostic strategy is introduced to randomly remove component masks from the person segmentation masks, which aims at increasing the difficulty of training and promoting the transformer encoder to recognize the underlying boundaries between each component. This enables the model to transfer both the shape and texture of the components. Furthermore, we propose a novel attribute decoder network to integrate multi-level attributes (e.g. the structure feature and the attribute representation) with well-designed Dual Adaptive Denormalization (DAD) residual blocks. Extensive experiments strongly demonstrate that the proposed approach is able to transfer both the texture and shape of different human parts and yield realistic results. To our knowledge, we are the first to learn disentangled latent representations with transformers for person image generation. Wenju Xu, Chengjiang Long, Yongwei Nie, Guanghui Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Building Coarse to Fine Convex Hulls With Auxiliary Vertices for Palette-Based Image RecoloringabstractConstructing a convex hull for the pixel colors of an image by viewing them as 3D points can extract a set of palette colors for the image, then image recoloring can be achieved by modifying the palette colors. For better recoloring effect, the convex hull should contain more pixels (inclusive) and be more compact. Otherwise, reconstruction error would occur or the extracted palette color would be less representative, yielding wrong recoloring results or less effective edit. We observe that convex hulls constructed by prior methods can contain all the image pixels, but are far from compact. Efforts have been made to optimize the vertices of convex hull to increase the compactness but are still not perfect. In this paper, we propose a novel coarse to fine convex hull construction scheme with auxiliary vertices. We start by constructing a coarse convex hull whose vertices are directly image pixels which is thus the most compact but cannot contain all pixels. We then make a remedy by adding auxiliary vertices into the coarse convex hull to obtain a fine convex hull. More auxiliary vertices are added, more image pixels will be contained into the fine convex hull. The auxiliary vertices are image pixels too so that the compactness can still be maintained. During editing, the auxiliary vertices are not allowed to be edited for edit convenience, but deformed as-rigid-as-possible with the adjusting of other vertices. Our convex hull is both inclusive and compact. Extensive experiments validate the effectiveness of the proposed method. Qiwei Sun, Yongwei Nie, Qing Zhang 0006, Guiqing Li |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | METRO-X: Combining Vertex and Parameter Regressions for Recovering 3D Human Meshes with Full Motions
Guiqing Li, Chenhao Yao, Huiqian Zhang, Juncheng Zeng, Yongwei Nie, Chuhua Xian |
CGI | 5 |
| 2023 | MANet: Multi-level Attention Network for 3D Human Shape and Pose Estimation
Chenhao Yao, Guiqing Li, Juncheng Zeng, Yongwei Nie, Chuhua Xian |
CGI (1) | 4 |
| 2023 | Fine-Grained Face Swapping Via Regional GAN InversionabstractWe present a novel paradigm for high-fidelity face swapping that faithfully preserves the desired subtle geometry and texture details. We rethink face swapping from the perspective of fine-grained face editing, i.e., “editing for swapping” (E4S), and propose a framework that is based on the explicit disentanglement of the shape and texture of facial components. Following the E4S principle, our framework enables both global and local swapping of facial features, as well as controlling the amount of partial swapping specified by the user. Furthermore, the E4S paradigm is in-herently capable of handling facial occlusions by means of facial masks. At the core of our system lies a novel Regional GAN Inversion (RGI) method, which allows the explicit disentanglement of shape and texture. It also allows face swapping to be performed in the latent space of Style-GAN. Specifically, we design a multi-scale mask-guided encoder to project the texture of each facial component into regional style codes. We also design a mask-guided injection module to manipulate the feature maps with the style codes. Based on the disentanglement, face swapping is re-formulated as a simplified problem of style and mask swapping. Extensive experiments and comparisons with current state-of-the-art methods demonstrate the superiority of our approach in preserving texture and shape details, as well as working with high resolution images. The project page is https://e4s2022.github.io Zhian Liu, Maomao Li, Yong Zhang 0034, Cairong Wang, Qi Zhang 0029, Jue Wang 0001, Yongwei Nie |
CVPR | 7 |
| 2023 | Learning Dynamic Style Kernels for Artistic Style TransferabstractArbitrary style transfer has been demonstrated to be efficient in artistic image generation. Previous methods either globally modulate the content feature ignoring local details, or overly focus on the local structure details leading to style leakage. In contrast to the literature, we propose a new scheme “style kernel” that learns spatially adaptive kernels for per-pixel stylization, where the convolutional kernels are dynamically generated from the global style-content aligned feature and then the learned kernels are applied to modulate the content feature at each spatial position. This new scheme allows flexible both global and local interactions between the content and style features such that the wanted styles can be easily transferred to the content image while at the same time the content structure can be easily preserved. To further enhance the flexibility of our style transfer method, we propose a Style Alignment Encoding (SAE) module complemented with a Content-based Gating Modulation (CGM) module for learning the dynamic style kernels in focusing regions. Extensive experiments strongly demonstrate that our proposed method outperforms state-of-the-art methods and exhibits superior performance in terms of visual quality and efficiency. Wenju Xu, Chengjiang Long, Yongwei Nie |
CVPR | 3 |
| 2023 | PSPDNet: Part-aware shape and pose disentanglement neural network for 3D human animating meshes
Guiqing Li, Juncheng Zeng, Fanzhong Zeng, Chenhao Yao, Bixia Kuang, Yongwei Nie |
Comput. Aided Geom. Des. | 6 |
| 2023 | Learning to Remove Shadows from a Single Image
Hao Jiang 0057, Qing Zhang 0006, Yongwei Nie, Lei Zhu 0003, Wei-Shi Zheng 0001 |
Int. J. Comput. Vis. | 3 |
| 2023 | Pyramid Texture FilteringabstractWe present a simple but effective technique to smooth out textures while preserving the prominent structures. Our method is built upon a key observation---the coarsest level in a Gaussian pyramid often naturally eliminates textures and summarizes the main image structures. This inspires our central idea for texture filtering, which is to progressively upsample the very low-resolution coarsest Gaussian pyramid level to a full-resolution texture smoothing result with well-preserved structures, under the guidance of each fine-scale Gaussian pyramid level and its associated Laplacian pyramid level. We show that our approach is effective to separate structure from texture of different scales, local contrasts, and forms, without degrading structures or introducing visual artifacts. We also demonstrate the applicability of our method on various applications including detail enhancement, image abstraction, HDR tone mapping, inverse halftoning, and LDR image enhancement. Code is available at https://rewindl.github.io/pyramid_texture_filtering/. Qing Zhang 0006, Hao Jiang 0057, Yongwei Nie, Wei-Shi Zheng 0001 |
ACM Trans. Graph. | 3 |
| 2023 | High-fidelity facial expression transfer using part-based local-global conditional gans
Muhammad Mamunur Rashid, Yongwei Nie, Guiqing Li |
Vis. Comput. | 3 |
| 2022 | Progressively Generating Better Initial Guesses Towards Next Stages for High-Quality Human Motion PredictionabstractThis paper presents a high-quality human motion pre-diction method that accurately predicts future human poses given observed ones. Our method is based on the observation that a good “initial guess” of the future poses is very helpful in improving the forecasting accuracy. This mo-tivates us to propose a novel two-stage prediction frame-work, including an init-prediction network that just computes the good guess and then a formal-prediction network that predicts the target future poses based on the guess. More importantly, we extend this idea further and design a multi-stage prediction framework where each stage pre-dicts initial guess for the next stage, which brings more performance gain. To fulfill the prediction task at each stage, we propose a network comprising Spatial Dense Graph Convolutional Networks (S-DGCN) and Temporal Dense Graph Convolutional Networks (T-DGCN). Alternatively executing the two networks helps extract spatiotem-poral features over the global receptive field of the whole pose sequence. All the above design choices cooperating together make our method outperform previous approaches by large margins: 6%-7% on Human3.6M, 5%-10% on CMU-MoCap, and 13%-16% on 3DPW. Code is available at https://github.com/705062791/PGBIG. Tiezheng Ma, Yongwei Nie, Chengjiang Long, Qing Zhang 0006, Guiqing Li |
CVPR | 2 |
| 2022 | Diverse Human Motion Prediction via Gumbel-Softmax Sampling from an Auxiliary SpaceabstractDiverse human motion prediction aims at predicting multiple possible future pose sequences from a sequence of observed poses. Previous approaches usually employ deep generative networks to model the conditional distribution of data, and then randomly sample outcomes from the distribution. While different results can be obtained, they are usually the most likely ones which are not diverse enough. Recent work explicitly learns multiple modes of the conditional distribution via a deterministic network, which however can only cover a fixed number of modes within a limited range. In this paper, we propose a novel sampling strategy for sampling very diverse results from an imbalanced multimodal distribution learned by a deep generative model. Our method works by generating an auxiliary space and smartly making randomly sampling from the auxiliary space equivalent to the diverse sampling from the target distribution. We propose a simple yet effective network architecture that implements this novel sampling strategy, which incorporates a Gumbel-Softmax coefficient matrix sampling method and an aggressive diversity promoting hinge loss function. Extensive experiments demonstrate that our method significantly improves both the diversity and accuracy of the samplings compared with previous state-of-the-art sampling approaches. Code and pre-trained models are available at https://github.com/Droliven/diverse_sampling. Lingwei Dang, Yongwei Nie, Chengjiang Long, Qing Zhang 0006, Guiqing Li |
ACM Multimedia | 2 |
| 2022 | GPU-Driven Real-Time Mesh Contour Vectorization
Wangziwei Jiang, Guiqing Li, Yongwei Nie, Chuhua Xian |
EGSR (ST) | 3 |
| 2022 | Learning Multi-Scale Deep Image Prior for High-Quality Unsupervised Image DenoisingabstractAbstract Recent methods on image denoising have achieved remarkable progress, benefiting mostly from supervised learning on massive noisy/clean image pairs and unsupervised learning on external noisy images. However, due to the domain gap between the training and testing images, these methods typically have limited applicability on unseen images. Although several attempts have been made to avoid the domain gap issue by learning denoising from singe noisy image itself, they are less effective in handling real‐world noise because of assuming the noise corruptions are independent and zero mean. In this paper, we go step further beyond prior work by presenting a novel unsupervised image denoising framework trained from single noisy image without making any explicit assumptions on the noise statistics. Our approach is built upon the deep image prior (DIP), which enables diverse image restoration tasks. However, as is, the denoising performance of DIP will significantly deteriorate on nonzero‐mean noise and is sensitive to the number of iterations. To overcome this problem, we propose to utilize multi‐scale deep image prior by imposing DIP across different image scales under the constraint of a scale consistency. Experiments on synthetic and real datasets demonstrate that our method performs favorably against the state‐of‐the‐art methods for image denoising. Hao Jiang 0057, Qing Zhang 0006, Yongwei Nie, Lei Zhu 0003, Wei-Shi Zheng 0001 |
Comput. Graph. Forum | 3 |
| 2022 | A Blind Color Separation Model for Faithful Palette-Based Image RecoloringabstractPalette-based image recoloring provides a simple yet effective way for color adjustment, which allows users to interactively manipulate the color of an image by editing a compact color palette. While remarkable progress has been made by previous methods, they have the common limitations that may produce unfaithful image recoloring results i.e., the obtained result does not respond faithfully to the palette adjustment, and tend to induce visual artifacts such as color bleeding and distortion. To address these limitations, we in this paper present a novel color separation model for palette-based recoloring. Akin to previous methods, our color separation model is built upon the assumption that color of each pixel in an image can be formulated as a linear combination of a small set of same basis colors. However, different from previous palette-based recoloring methods which typically rely on heuristic rules to build the color separation model, we experimentally reveal the underlying relationship between the color separation and the palette-based recoloring, and summarize three specialized color separation priors that allow more faithful palette-based recoloring. Based on these priors, we devise a blind color separation model that not only does not require known palette as input as done in previous methods, but also enables more effective palette-based recoloring with much less visual artifacts. Experiments on two datasets demonstrate that our method outperforms the state-of-the-art palette-based recoloring methods. In addition, we show some applications enabled by the proposed color separation model, including automatic pattern coloring generation, green screen keying and region-controllable color transfer. Qing Zhang 0006, Yongwei Nie, Lei Zhu 0003, Chunxia Xiao, Wei-Shi Zheng 0001 |
IEEE Trans. Multim. | 2 |
| 2021 | MSR-GCN: Multi-Scale Residual Graph Convolution Networks for Human Motion PredictionabstractHuman motion prediction is a challenging task due to the stochasticity and aperiodicity of future poses. Recently, graph convolutional network has been proven to be very effective to learn dynamic relations among pose joints, which is helpful for pose prediction. On the other hand, one can abstract a human pose recursively to obtain a set of poses at multiple scales. With the increase of the abstraction level, the motion of the pose becomes more stable, which benefits pose prediction too. In this paper, we propose a novel Multi-Scale Residual Graph Convolution Network (MSR-GCN) for human pose prediction task in the manner of end-to-end. The GCNs are used to extract features from fine to coarse scale and then from coarse to fine scale. The extracted features at each scale are then combined and decoded to obtain the residuals between the input and target poses. Intermediate supervisions are imposed on all the predicted poses, which enforces the network to learn more representative features. Our proposed approach is evaluated on two standard benchmark datasets, i.e., the Human3.6M dataset and the CMU Mocap dataset. Experimental results demonstrate that our method outperforms the state-of-the-art approaches. Code and pre-trained models are available at https://github.com/Droliven/MSRGCN. Lingwei Dang, Yongwei Nie, Chengjiang Long, Qing Zhang 0006, Guiqing Li |
ICCV | 2 |
| 2021 | A Hybrid Video Anomaly Detection Framework via Memory-Augmented Flow Reconstruction and Flow-Guided Frame PredictionabstractIn this paper, we propose HF2-VAD, a Hybrid framework that integrates Flow reconstruction and Frame prediction seamlessly to handle Video Anomaly Detection. Firstly, we design the network of ML-MemAE-SC (Multi-Level Memory modules in an Autoencoder with Skip Connections) to memorize normal patterns for optical flow reconstruction so that abnormal events can be sensitively identified with larger flow reconstruction errors. More importantly, conditioned on the reconstructed flows, we then employ a Conditional Variational Autoencoder (CVAE), which captures the high correlation between video frame and optical flow, to predict the next frame given several previous frames. By CVAE, the quality of flow reconstruction essentially influences that of frame prediction. Therefore, poorly reconstructed optical flows of abnormal events further deteriorate the quality of the final predicted future frame, making the anomalies more detectable. Experimental results demonstrate the effectiveness of the proposed method. Code is available at https://github.com/LiUzHiAn/hf2vad. Zhian Liu, Yongwei Nie, Chengjiang Long, Qing Zhang 0006, Guiqing Li |
ICCV | 2 |
| 2021 | BPA-GAN: Human motion transfer using body-part-aware generative adversarial networks
Jinfeng Jiang, Guiqing Li, Huiqian Zhang, Yongwei Nie |
Graph. Model. | 5 |
| 2021 | Understanding More About Human and Machine Attention in Deep Neural NetworksabstractHuman visual system can selectively attend to parts of a scene for quick perception, a biological mechanism known asHuman attention. Inspired by this, recent deep learning models encode attention mechanisms to focus on the most task-relevant parts of the input signal for further processing, which is calledMachine/Neural/Artificial attention. Understanding the relation between human and machine attention is important for interpreting and designing neural networks. Many works claim that the attention mechanism offers an extra dimension of interpretability by explaining where the neural networks look. However, recent studies demonstrate that artificial attention maps do not always coincide with common intuition. In view of these conflicting evidence, here we make a systematic study on using artificial attention and human attention in neural network design. With three example computer vision tasks (i.e., salient object segmentation, video action recognition, and fine-grained image classification), diverse representative backbones (i.e., AlexNet, VGGNet, ResNet) and famous architectures (i.e., Two-stream, FCN), corresponding real human gaze data, and systematically conducted large-scale quantitative studies, we quantify the consistency between artificial attention and human visual attention and offer novel insights into existing artificial attention mechanisms by giving preliminary answers to several key questions related to human and artificial attention mechanisms. Overall results demonstrate that human attention can benchmark the meaningful ‘ground-truth’ in attention-driven tasks, where the more the artificial attention is close to human attention, the better the performance; for higher-level vision tasks, it is case-by-case. It would be advisable for attention-driven tasks to explicitly force a better alignment between artificial and human attention to boost the performance; such alignment would also improve the network explainability for higher-level computer vision tasks. Qiuxia Lai, Salman Khan 0001, Yongwei Nie, Hanqiu Sun, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Multim. | 3 |
| 2021 | Enhancing Underexposed Photos Using Perceptually Bidirectional SimilarityabstractAlthough remarkable progress has been made, existing methods for enhancing underexposed photos tend to produce visually unpleasing results due to the existence of visual artifacts (e.g., color distortion, loss of details and uneven exposure). We observed that this is because they fail to ensure the perceptual consistency of visual information between the source underexposed image and its enhanced output. To obtain high-quality results free of these artifacts, we present a novel underexposed photo enhancement approach that is able to maintain the perceptual consistency. We achieve this by proposing an effective criterion, referred to as perceptually bidirectional similarity, which explicitly describes how to ensure the perceptual consistency. Particularly, we adopt the Retinex theory and cast the enhancement problem as a constrained illumination estimation optimization, where we formulate perceptually bidirectional similarity as constraints on illumination and solve for the illumination which can recover the desired artifact-free enhancement results. In addition, we describe a video enhancement framework that adopts the presented illumination estimation for handling underexposed videos. To this end, a probabilistic approach is introduced to propagate illuminations of sampled keyframes to the entire video by tackling a Bayesian Maximum A Posteriori problem. Extensive experiments demonstrate the superiority of our method over the state-of-the-art methods. Qing Zhang 0006, Yongwei Nie, Lei Zhu 0003, Chunxia Xiao, Wei-Shi Zheng 0001 |
IEEE Trans. Multim. | 2 |
| 2021 | PanoMan: Sparse Localized Components-based Model for Full Human MotionsabstractParameterizing Variations of human shapes and motions is a long-standing problem in computer graphics and vision. Most of the existing methods only deal with a specific kind of motion, such as body poses, facial expressions, or hand gestures. We propose PanoMan (sParse locAlized compoNents based mOdel for full huMAn motioNs) to handle shape variation and full-motion across body, face, and hand in a unified framework. Like previous approaches, we factor shape variation into principal components to obtain a human shape space that approximates the shape of arbitrary identity. We then analyze sparse localized components in terms of relative edge length and dihedral angle to capture full motions of body poses, facial expressions, and hand gestures. The final piece of our model is a multilayer perceptron (MLP) that fits the residual between the ground truth and the aforementioned two-level approximation. As an application, we employ the discrete-shell deformation to drive the model to fit sparse constraints such as joint positions and surface feature points. We thoroughly evaluate PanoMan on body, face, and hand motion benchmarks as well as scanned data. The existing skinning-based techniques suffer from joint collapsing when encountering twisting motion of joints. Experiments show that PanoMan can capture all kinds of full human motions with high quality and is easier than the state-of-the-art models in recovering poses with wide joint twisting and complex hand gestures. Yupan Wang, Guiqing Li, Huiqian Zhang, Xinyi Zou, Yongwei Nie |
ACM Trans. Graph. | 6 |
| 2021 | Monte Carlo denoising via auxiliary feature guided self-attentionabstractWhile self-attention has been successfully applied in a variety of natural language processing and computer vision tasks, its application in Monte Carlo (MC) image denoising has not yet been well explored. This paper presents a self-attention based MC denoising deep learning network based on the fact that self-attention is essentially non-local means filtering in the embedding space which makes it inherently very suitable for the denoising task. Particularly, we modify the standard self-attention mechanism to an auxiliary feature guided self-attention that considers the by-products (e.g., auxiliary feature buffers) of the MC rendering process. As a critical prerequisite to fully exploit the performance of self-attention, we design a multi-scale feature extraction stage, which provides a rich set of raw features for the later self-attention module. As self-attention poses a high computational complexity, we describe several ways that accelerate it. Ablation experiments validate the necessity and effectiveness of the above design choices. Comparison experiments show that the proposed self-attention based MC denoising method outperforms the current state-of-the-art methods. Yongwei Nie, Chengjiang Long, Wenjun Xu 0002, Qing Zhang 0006, Guiqing Li |
ACM Trans. Graph. | 2 |
| 2021 | 3D hand reconstruction from a single image based on biomechanical constraints
Guiqing Li, Zihui Wu, Huiqian Zhang, Yongwei Nie, Aihua Mao |
Vis. Comput. | 5 |
| 2020 | Deep Camouflage ImagesabstractThis paper addresses the problem of creating camouflage images. Such images typically contain one or more hidden objects embedded into a background image, so that viewers are required to consciously focus to discover them. Previous methods basically rely on hand-crafted features and texture synthesis to create camouflage images. However, due to lack of reliable understanding of what essentially makes an object recognizable, they typically result in either complete standout or complete invisible hidden objects. Moreover, they may fail to produce seamless and natural images because of the sensitivity to appearance differences. To overcome these limitations, we present a novel neural style transfer approach that adopts the visual perception mechanism to create camouflage images, which allows us to hide objects more effectively while producing natural-looking results. In particular, we design an attention-aware camouflage loss to adaptively mask out information that make the hidden objects visually standout, and also leave subtle yet enough feature clues for viewers to perceive the hidden objects. To remove the appearance discontinuities between the hidden objects and the background, we formulate a naturalness regularization to constrain the hidden objects to maintain the manifold structure of the covered background. Extensive experiments show the advantages of our approach over existing camouflage methods and state-of-the-art neural style transfer algorithms. Qing Zhang 0006, Gelin Yin, Yongwei Nie, Wei-Shi Zheng 0001 |
AAAI | 3 |
| 2020 | Video super-resolution via pre-frame constrained and deep-feature enhanced sparse reconstruction
Qiuxia Lai, Yongwei Nie, Hanqiu Sun, Qiang Xu 0001, Zhensong Zhang, Mingyu Xiao 0001 |
Pattern Recognit. | 2 |
| 2020 | Interactive Contour Extraction via Sketch-Alike Dense-Validation OptimizationabstractWe propose an interactive contour extraction method inspired by a skill often adopted in sketching: an artist usually sketches an object by first drawing lots of short, directional, and redundant strokes, then following these small strokes to draw the final outline of the object. Our method simulates this process. To extract a contour, our method relies on user interaction, which provides us with a narrow band containing the target contour. Then, we densely sample sub-bands from the whole band, with each sub-band containing a local segment of the target contour. We design a curve-centered coordinate system in which a dynamic programming algorithm is proposed to extract the local segment in each sub-band. The local segment is guaranteed to be as evident and smooth as possible, to mimic the strokes sketched by the artist. Finally, we integrate all local segments of all sub-bands together to obtain the whole target contour based on the weighted principal component analysis. Our method can extract high-quality object contours due to the dense validations among local segments. That is, even if one segment deviates from the right location, several other segments in its local neighborhood can correct it in the integration stage. Both quantitative experiments and a user study demonstrate the effectiveness of the proposed method. Yongwei Nie, Ping Li 0016, Qing Zhang 0006, Zhensong Zhang, Guiqing Li, Hanqiu Sun |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Collision-Free Video Synopsis Incorporating Object Speed and Size ChangesabstractThis paper presents a new surveillance video synopsis method which performs much better than previous approaches in terms of both compression ratio and artifact. Previously, a surveillance video was usually compressed by shifting the moving objects of that video forward along the time axis, which inevitably yielded serious collision and chronological disorder artifacts between the shifted objects. The main observation of this paper is that these artifacts can be alleviated by changing the speed or size of the objects, since with varied speed and size the objects can move more flexibly to avoid collision points or to keep chronological relationships. Based on this observation, we propose a video synopsis method that performs object shifting, speed changing, and size scaling simultaneously. We show how to integrate the three heterogeneous operations into a single optimization framework and achieve high-quality synopsis results. Unlike previous approaches that usually use alternative optimization strategies to solve synopsis optimizations, we develop a Metropolis sampling algorithm to find the solution for our three-variable optimization problem. A variety of experiments demonstrate the effectiveness of our method. Yongwei Nie, Zhenkai Li, Zhensong Zhang, Qing Zhang 0006, Tiezheng Ma, Hanqiu Sun |
IEEE Trans. Image Process. | 1 |
| 2020 | Multi-View Video Synopsis via Simultaneous Object-Shifting and View-Switching OptimizationabstractWe present a method for synopsizing multiple videos captured by a set of surveillance cameras with some overlapped field-of-views. Currently, object-based approaches that directly shift objects along the time axis are already able to compute compact synopsis results for multiple surveillance videos. The challenge is how to present the multiple synopsis results in a more compact and understandable way. Previous approaches show them side by side on the screen, which however is difficult for user to comprehend. In this paper, we solve the problem by joint object-shifting and camera view-switching. Firstly, we synchronize the input videos, and group the same object in different videos together. Then we shift the groups of objects along the time axis to obtain multiple synopsis videos. Instead of showing them simultaneously, we just show one of them at each time, and allow to switch among the views of different synopsis videos. In this view switching way, we obtain just a single synopsis results consisting of content from all the input videos, which is much easier for user to follow and understand. To obtain the best synopsis result, we construct a simultaneous object-shifting and view-switching optimization framework instead of solving them separately. We also present an alternative optimization strategy composed of graph cuts and dynamic programming to solve the unified optimization. Experiments demonstrate that our single synopsis video generated from multiple input videos is compact, complete, and easy to understand. Zhensong Zhang, Yongwei Nie, Hanqiu Sun, Qing Zhang 0006, Qiuxia Lai, Guiqing Li, Mingyu Xiao 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Effective Video Stabilization via Joint Trajectory Smoothing and Frame WarpingabstractVideo stabilization is usually composed of three stages: feature trajectory extraction, trajectory smoothing, and frame warping. Most previous approaches view them as three separate stages. This paper proposes a method combining the last two stages, namely the trajectory smoothing and frame warping stages, into a single optimization framework. The novelty exists in the way of how we combine them: the trajectory smoothing part plays a major role while the frame warping part plays an auxiliary role. With this kind of design, we can conveniently increase the strength of the trajectory smoothing part by a robust first-order derivative term, which makes it possible to produce very aggressive stabilization effects. On the other hand, we adopt adaptive weighting mechanisms in the frame warping part, to follow the smoothed trajectories as much as possible while regularizing other places as similar as possible. Our method is robust to utilize both foreground and background features, and very short trajectories. The utilization of all these information in turn increases the accuracy of the proposed method. We also provide a simplified implementation of our method, which is less accurate but more efficient. Experiments on various kinds of videos demonstrate the effectiveness of our method. Tiezheng Ma, Yongwei Nie, Qing Zhang 0006, Zhensong Zhang, Hanqiu Sun, Guiqing Li |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2019 | Data-driven 3D human head reconstruction
Huayun He, Guiqing Li, Zehao Ye, Aihua Mao, Chuhua Xian, Yongwei Nie |
Comput. Graph. | 6 |
| 2019 | Discrete shell deformation driven by adaptive sparse localized components
Guiqing Li, Yupan Wang, Yongwei Nie, Aihua Mao |
Comput. Graph. | 4 |
| 2019 | Dual Illumination Estimation for Robust Exposure CorrectionabstractAbstract Exposure correction is one of the fundamental tasks in image processing and computational photography. While various methods have been proposed, they either fail to produce visually pleasing results, or only work well for limited types of image (e.g., underexposed images). In this paper, we present a novel automatic exposure correction method, which is able to robustly produce high‐quality results for images of various exposure conditions (e.g., underexposed, overexposed, and partially under‐ and over‐exposed). At the core of our approach is the proposed dual illumination estimation, where we separately cast the under‐and over‐exposure correction as trivial illumination estimation of the input image and the inverted input image. By performing dual illumination estimation, we obtain two intermediate exposure correction results for the input image, with one fixes the underexposed regions and the other one restores the overexposed regions. A multi‐exposure image fusion technique is then employed to adaptively blend the visually best exposed parts in the two intermediate exposure correction images and the input image into a globally well‐exposed image. Experiments on a number of challenging images demonstrate the effectiveness of the proposed approach and its superiority over the state‐of‐the‐art methods and popular automatic exposure correction tools. Qing Zhang 0006, Yongwei Nie, Wei-Shi Zheng 0001 |
Comput. Graph. Forum | 2 |
| 2019 | SRNPD: Spatial rendering network for pencil drawing stylizationabstractAbstract Pencil drawing is a simple yet effective way to depict what people see by clearly presenting details of the scene. Existing methods usually extract strokes of the input image and adjust the result image tone to make it look like a pencil drawing. However, they do not consider the quality of the stroke image and the geometry information of lines in the stroke image, which unavoidably results in the violation of original essential structures and in a flatten pencil drawing with unrealistic appearance. We put forward a spatial rendering network for pencil drawing stylization. Spatial stroke images are extracted from the image pyramid by a single‐shot bottom‐up neural network to improve the quality of these stroke images. Unlike the former tone adjustment–based methods, we analyze perceptual cues of strokes at different stroke image levels and use the obtained geometry information to constrain the stroke shading procedure. The final pencil drawing result is achieved by the stroke shading fusion of different levels' shading results. The effectiveness of our spatial rendering network for pencil drawing stylization is demonstrated by an ablation study, comparison to the state of the art, and a user study. Yuxi Jin, Ping Li 0016, Bin Sheng 0001, Yongwei Nie, Jinman Kim, Enhua Wu |
Comput. Animat. Virtual Worlds | 4 |
| 2019 | Multiview-coherent disocclusion synthesis using connected regions optimizationabstractAbstract Handling of missing areas is a key step for depth‐based rendering to synthesize virtual views. Existing methods usually consider finding candidate pixels from only one reference view to fill missing areas. However, the information provided by one reference view is restricted by the position of the view. By utilizing two reference views located on both the left and right sides of the virtual view, we propose to synthesize the missing areas at hole level with connected regions optimization. To avoid the appearance of ghost boundary, we apply morphological operations to generate a boundary band map for the depth map, which restricts the warping of the virtual view. We use a binary map to mark unknown pixels in a hole, label the connected unknown regions, and count the area of each connected region, which decide the order in our enhanced inpainting synthesis. Besides, we separate the foreground and background regions of the depth map to constrain the searching of candidate pixels. Multiple experiments on virtual view synthesis have shown the effectiveness and high quality of our multiview‐coherent disocclusion synthesis. Ping Li 0016, Yuxi Jin, Bin Sheng 0001, Di Lin 0002, Yongwei Nie, Enhua Wu |
Comput. Animat. Virtual Worlds | 5 |
| 2019 | Structure-preserving image completion with multi-level dynamic patches
Bowen Liu 0015, Ping Li 0016, Bin Sheng 0001, Yongwei Nie, Enhua Wu |
Vis. Comput. | 4 |
| 2018 | Temporal Coherent Video Super-resolution via Pre-frame-constrained Sparse ReconstructionabstractIn this paper, we extend the sparse representation based image super-resolution method to process videos, mainly aiming at obtaining temporally consistent consecutive high-resolution (HR) video frames. In our formulation, the previous estimated HR frame is used to guide the sparse reconstruction of current low-resolution (LR) frame, which is able to obtain more consistent representations. We show that such guidance is robust and effective by incorporating with a non-rigid dense correspondence based motion compensation schema. We also propose a dictionary updating strategy which regularly updates the dictionaries that are critical for the sparse representation procedure using the newly reconstructed HR frames. To further preserve sharp edges and remove reconstruction errors, once a HR image is recovered, we refine it with a L0-norm based optimization that constrains the final HR output with relatively sparse gradients. Experimental results on natural videos demonstrated the effectiveness of our proposed method. Qiuxia Lai, Yongwei Nie, Zhensong Zhang, Hanqiu Sun |
CGI | 2 |
| 2018 | Rolling normal filtering for point clouds
Yinglong Zheng, Guiqing Li, Xuemiao Xu, Yongwei Nie |
Comput. Aided Geom. Des. | 5 |
| 2018 | Dynamic Video Stitching via Shakiness RemovingabstractStitching videos captured by hand-held mobile cameras can essentially enhance entertainment experience of ordinary users. However, such videos usually contain heavy shakiness and large parallax, which are challenging to stitch. In this paper, we propose a novel approach of video stitching and stabilization for videos captured by mobile devices. The main component of our method is a unified video stitching and stabilization optimization that computes stitching and stabilization simultaneously rather than does each one individually. In this way, we can obtain the best stitching and stabilization results relative to each other without any bias to one of them. To make the optimization robust, we propose a method to identify background of input videos, and also common background of them. This allows us to apply our optimization on background regions only, which is the key to handle large parallax problem. Since stitching relies on feature matches between input videos, and there inevitably exist false matches, we thus propose a method to distinguish between right and false matches, and encapsulate the false match elimination scheme and our optimization into a loop, to prevent the optimization from being affected by bad feature matches. We test the proposed approach on videos that are causally captured by smartphones when walking along busy streets, and use stitching and stability scores to evaluate the produced panoramic videos quantitatively. Experiments on a diverse of examples show that our results are much better than (challenging cases) or at least on par with (simple cases) the results of previous approaches.Stitching videos captured by hand-held mobile cameras can essentially enhance entertainment experience of ordinary users. However, such videos usually contain heavy shakiness and large parallax, which are challenging to stitch. In this paper, we propose a novel approach of video stitching and stabilization for videos captured by mobile devices. The main component of our method is a unified video stitching and stabilization optimization that computes stitching and stabilization simultaneously rather than does each one individually. In this way, we can obtain the best stitching and stabilization results relative to each other without any bias to one of them. To make the optimization robust, we propose a method to identify background of input videos, and also common background of them. This allows us to apply our optimization on background regions only, which is the key to handle large parallax problem. Since stitching relies on feature matches between input videos, and there inevitably exist false matches, we thus propose a method to distinguish between right and false matches, and encapsulate the false match elimination scheme and our optimization into a loop, to prevent the optimization from being affected by bad feature matches. We test the proposed approach on videos that are causally captured by smartphones when walking along busy streets, and use stitching and stability scores to evaluate the produced panoramic videos quantitatively. Experiments on a diverse of examples show that our results are much better than (challenging cases) or at least on par with (simple cases) the results of previous approaches. Yongwei Nie, Tan Su, Zhensong Zhang, Hanqiu Sun, Guiqing Li |
IEEE Trans. Image Process. | 1 |
| 2018 | Corrections to "Dynamic Video Stitching via Shakiness Removing"abstractIn[1], the biographies of Hanqiu Sun and Guiqing Li included incorrect information. The correct biographies are as follows. Yongwei Nie, Tan Su, Zhensong Zhang, Hanqiu Sun, Guiqing Li |
IEEE Trans. Image Process. | 1 |
| 2017 | How Does a Camera Look at One 3D CAD Object?abstractCamera pose and the camera’s rotation angles and translation vector (RT), are one-to-one relation with a 2D real image when the intrinsic parameter is fixed. In this paper, we propose a novel convolutional neural network (CNN) based framework to intelligently estimate the 6-DOF RTs from images taken on one 3D CAD object directly and indirectly, as well as visually verifying the correctness of the predicted RTs. Such a framework enables us to accurately interpret how a camera looks at the object. The direct way is simple and obtains lower average errors for the predicted RTs experimentally, while the indirect way utilizes the POSIT algorithm via landmarks and is able to avoid the non-Euclidean issue in rotation angles. To our best knowledge, we are the first one to estimate camera’s RTs and effectively interprets how a camera looks at one 3D CAD object from the images taken on it. The experiments on four models quantitatively and qualitatively demonstrate the efficacy of our proposed approach. Chuang Xing, Chengjiang Long, Hao Guo 0005, Yongwei Nie, Dehai Zhu, Qin Ma 0001, Mengxiao Tian |
ICTAI | 4 |
| 2017 | Homography Propagation and Optimization for Wide-Baseline Street Image InterpolationabstractWide-baseline street image interpolation is useful but very challenging. Existing approaches either rely on heavyweight 3D reconstruction or computationally intensive deep networks. We present a lightweight and efficient method which uses simple homography computing and refining operators to estimate piecewise smooth homographies between input views. To achieve the goal, we show how to combine homography fitting and homography propagation together based on reliable and unreliable superpixel discrimination. Such a combination, other than using homography fitting only, dramatically increases the accuracy and robustness of the estimated homographies. Then, we integrate the concepts of homography and mesh warping, and propose a novel homography-constrained warping formulation which enforces smoothness between neighboring homographies by utilizing the first-order continuity of the warped mesh. This further eliminates small artifacts of overlapping, stretching, etc. The proposed method is lightweight and flexible, allows wide-baseline interpolation. It improves the state of the art and demonstrates that homography computation suffices for interpolation. Experiments on city and rural datasets validate the efficiency and effectiveness of our method. Yongwei Nie, Zhensong Zhang, Hanqiu Sun, Tan Su, Guiqing Li |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2016 | Stylistic indoor colour design via Bayesian network
Guangming Chen, Guiqing Li, Yongwei Nie, Chuhua Xian, Aihua Mao |
Comput. Graph. | 3 |
| 2016 | Underexposed Video Enhancement via Perception-Driven Progressive FusionabstractUnderexposed video enhancement aims at revealing hidden details that are barely noticeable in LDR video frames with noise. Previous work typically relies on a single heuristic tone mapping curve to expand the dynamic range, which inevitably leads to uneven exposure and visual artifacts. In this paper, we present a novel approach for underexposed video enhancement using an efficient perception-driven progressive fusion. For an input underexposed video, we first remap each video frame using a series of tentative tone mapping curves to generate an multi-exposure image sequence that contains different exposed versions of the original video frame. Guided by some visual perception quality measures encoding the desirable exposed appearance, we locate all the best exposed regions from multi-exposure image sequences and then integrate them into a well-exposed video in a temporally consistent manner. Finally, we further perform an effective texture-preserving spatio-temporal filtering on this well-exposed video to obtain a high-quality noise-free result. Experimental results have shown that the enhanced video exhibits uniform exposure, brings out noticeable details, preserves temporal coherence, and avoids visual artifacts. Besides, we demonstrate applications of our approach to a set of problems including video dehazing, video denoising and HDR video reconstruction. Qing Zhang 0006, Yongwei Nie, Ling Zhang 0017, Chunxia Xiao |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2015 | Color correspondence of image warping using plane constraintsabstractImage-based rendering (IBR) can render novel views from several existing photographs of the same scene. Current IBR methods can be roughly classified into reconstruction-based methods and warping-based methods. Warping-based methods [Liu et al. 2009; Chaurasia et al. 2011] directly warped the triangle mesh overlaid on the source image to make the source image match with the target one. Then the novel views between the two given images can be computed by interpolating and retexturing the intermediate mesh. Zhensong Zhang, Yongwei Nie, Hanqiu Sun |
VRST | 2 |
| 2015 | Content-aware model resizing with symmetry-preservation
Chunxia Xiao, Liqiang Jin, Yongwei Nie, Renfang Wang, Hanqiu Sun, Kwan-Liu Ma |
Vis. Comput. | 3 |
| 2014 | Object Movements Synopsis viaPart Assembling and StitchingabstractVideo synopsis aims at removing video's less important information, while preserving its key content for fast browsing, retrieving, or efficient storing. Previous video synopsis methods, including frame-based and object-based approaches that remove valueless whole frames or combine objects from time shots, cannot handle videos with redundancies existing in the movements of video object. In this paper, we present a novel part-based object movements synopsis method, which can effectively compress the redundant information of a moving video object and represent the synopsized object seamlessly. Our method works by part-based assembling and stitching. The object movement sequence is first divided into several part movement sequences. Then, we optimally assemble moving parts from different part sequences together to produce an initial synopsis result. The optimal assembling is formulated as a part movement assignment problem on a Markov Random Field (MRF), which guarantees the most important moving parts are selected while preserving both the spatial compatibility between assembled parts and the chronological order of parts. Finally, we present a non-linear spatiotemporal optimization formulation to stitch the assembled parts seamlessly, and achieve the final compact video object synopsis. The experiments on a variety of input video objects have demonstrated the effectiveness of the presented synopsis method. Yongwei Nie, Hanqiu Sun, Ping Li 0016, Chunxia Xiao, Kwan-Liu Ma |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2013 | Compact Video Synopsis via Global Spatiotemporal OptimizationabstractVideo synopsis aims at providing condensed representations of video data sets that can be easily captured from digital cameras nowadays, especially for daily surveillance videos. Previous work in video synopsis usually moves active objects along the time axis, which inevitably causes collisions among the moving objects if compressed much. In this paper, we propose a novel approach for compact video synopsis using a unified spatiotemporal optimization. Our approach globally shifts moving objects in both spatial and temporal domains, which shifting objects temporally to reduce the length of the video and shifting colliding objects spatially to avoid visible collision artifacts. Furthermore, using a multilevel patch relocation (MPR) method, the moving space of the original video is expanded into a compact background based on environmental content to fit with the shifted objects. The shifted objects are finally composited with the expanded moving space to obtain the high-quality video synopsis, which is more condensed while remaining free of collision artifacts. Our experimental results have shown that the compact video synopsis we produced can be browsed quickly, preserves relative spatiotemporal relationships, and avoids motion collisions. Yongwei Nie, Chunxia Xiao, Hanqiu Sun, Ping Li 0016 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2013 | Video retargeting combining warping and summarizing optimization
Yongwei Nie, Qing Zhang 0006, Renfang Wang, Chunxia Xiao |
Vis. Comput. | 1 |
| 2012 | Interactive image/video retexturing using GPU parallelism
Ping Li 0016, Hanqiu Sun, Jianbing Shen, Yongwei Nie |
Comput. Graph. | 5 |
| 2011 | Fast Exact Nearest Patch Matching for Patch-Based Image Editing and ProcessingabstractThis paper presents an efficient exact nearest patch matching algorithm which can accurately find the most similar patch-pairs between source and target image. Traditional match matching algorithms treat each pixel/patch as an independent sample and build a hierarchical data structure, such as kd-tree, to accelerate nearest patch finding. However, most of these approaches can only find approximate nearest patch and do not explore the sequential overlap between patches. Hence, they are neither accurate in quality nor optimal in speed. By eliminating redundant similarity computation of sequential overlap between patches, our method finds the exact nearest patch in brute-force style but reduces its running time complexity to be linear on the patch size. Furthermore, relying on recent multicore graphics hardware, our method can be further accelerated by at least an order of magnitude (≥10×). This greatly improves performance and ensures that our method can be efficiently applied in an interactive editing framework for moderate-sized image even video. To our knowledge, this approach is the fastest exact nearest patch matching method for high-dimensional patch and also its extra memory requirement is minimal. Comparisons with the popular nearest patch matching methods in the experimental results demonstrate the merits of our algorithm. Chunxia Xiao, Yongwei Nie, Zhao Dong 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2011 | Efficient Edit Propagation Using Hierarchical Data StructureabstractThis paper presents a novel unified hierarchical structure for scalable edit propagation. Our method is based on the key observation that in edit propagation, appearance varies very smoothly in those regions where the appearance is different from the user-specified pixels. Uniformly sampling in these regions leads to redundant computation. We propose to use a quadtree-based adaptive subdivision method such that more samples are selected in similar regions and less in those that are different from the user-specified regions. As a result, both the computation and the memory requirement are significantly reduced. In edit propagation, an edge-preserving propagation function is first built, and the full solution for all the pixels can be computed by interpolating from the solution obtained from the adaptively subdivided domain. Furthermore, our approach can be easily extended to accelerate video edit propagation using an adaptive octree structure. In order to improve user interaction, we introduce several new Gaussian Mixture Model (GMM) brushes to find pixels that are similar to the user-specified regions. Compared with previous methods, our approach requires significantly less time and memory, while achieving visually same results. Experimental results demonstrate the efficiency and effectiveness of our approach on high-resolution photographs and videos. Chunxia Xiao, Yongwei Nie |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2010 | Fast multi-scale joint bilateral texture upsampling
Chunxia Xiao, Yongwei Nie, Wenting Zheng |
Vis. Comput. | 2 |
| 2009 | Multie-scale joint bilateral image and video texture upsamplingabstractWe present a novel approach for upsampling the synthesized image and video texture using a multi-scale joint bilateral filter. Our method is based on the motivation: if the available exemplar texture is used as a prior to upsample the synthesized texture, a high resolution result that better preventing image blurring can be obtained. Our joint bilateral upsampling applies a spatial filter on the synthesized texture, and jointly applies a similar range filter on exemplar texture which guides the interpolation from low to high resolution. To further enhance the detail of the upsampled texture, we propose multi-scale joint bilateral upsampling method which progressively enhances the detail of the upsampled texture. In addition, we present a detail-aware texture optimization approach which combines texture optimization and histogram matching of the image detail to improve the quality of the synthesized results. Finally, we present an accelerated joint bilateral filter, which enables our upsampling process to interactively generate a large texture. We show results for upsampling image and video texture and compare them to traditional upsampling methods, which illustrate that our methods require low computational and memory costs while receive better results. Chunxia Xiao, Yongwei Nie, Guangpu Feng |
CAD/Graphics | 2 |