VLDB 2026 Research / reviewers in the wild / expert
Kun Li 0001
dblp:75/1458-1
· DBLP profile ↗
86ranked-venue papers
18as first author
37since 2021 · last 2026
0000-0003-2326-0166ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 69 · 14 first-author · 28 since 2021Artificial intelligence and machine learning · 21 · 2 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | InterCoser: Interactive 3D Character Creation with Disentangled Fine-Grained FeaturesabstractThis paper aims to interactively generate and edit disentangled 3D characters based on precise user instructions. Existing methods generate and edit 3D characters via rough and simple editing guidance and entangled representations, making it difficult to achieve precise and comprehensive control over fine-grained local editing and free clothing transfer for characters. To enable accurate and intuitive control over the generation and editing of high-quality 3D characters with freely interchangeable clothing, we propose a novel user-interactive approach for disentangled 3D character creation. Specifically, to achieve precise control over 3D character generation and editing, we introduce two user-friendly interaction approaches: a sketch-based layered character generation/editing method, which supports clothing transfer; and a 3D-proxy-based part-level editing method, enabling fine-grained disentangled editing. To enhance 3D character quality, we propose a 3D Gaussian reconstruction strategy guided by geometric priors, ensuring that 3D characters exhibit detailed local geometry and smooth global surfaces. Extensive experiments on both public datasets and in-the-wild data demonstrate that our approach not only generates high-quality disentangled 3D characters but also supports precise and fine-grained editing through user interaction. Zhuo Su 0006, Guidong Wang, Jing-Yu Yang 0002, Yukun Lai, Kun Li 0001 |
AAAI | 7 |
| 2026 | LoGAvatar: Local Gaussian Splatting for human avatar modeling from monocular video
Xiongzheng Li, Hailong Jia, Zhuo Su 0006, Guidong Wang, Kun Li 0001 |
Comput. Aided Des. | 7 |
| 2026 | SMixNet: Style Mixture Network for Exemplar-Based Image TranslationabstractExemplar-based image translation, which aims to transfer the style of an exemplar image to an input semantic image, is challenging and important in many applications. Most current methods build coarse correspondences and overlook extracting faithful style information from the exemplar image, leading to unsatisfactory results with style inconsistent with the exemplar image. In this paper, we propose a novel and efficient style mixture block to extract faithful style information and build reliable correspondences progressively. Specifically, instead of modeling explicit correspondences, we extract faithful style descriptors by considering global information about the exemplar features. Then, we generate coefficients for these style descriptors by modeling the interaction between the exemplar image and the input image, and efficiently compose these descriptors using the coefficients. The efficiency of the style mixture block allows a multi-scale architecture to extract and transform style descriptors at different resolutions, deforming the features of the exemplar image and refining the correspondences progressively. Experimental results on several datasets show that our SMixNet outperforms the current state-of-the-art, and is faster. Code is available for research purposes at https://github.com/Zhangjinso/SMixNet. Yukun Lai, Hongjiang Xiao, Kun Li 0001 |
Comput. Vis. Media | 4 |
| 2025 | RESCUE: Crowd Evacuation Simulation via Controlling SDM-United Characters
Joey Tianyi Zhou, Hongbo Kang, Wenguo Weng, Yukun Lai, Kun Li 0001 |
ICCV | 9 |
| 2025 | FRNeRF: Fusion and Regularization Fields for Dynamic View SynthesisabstractNovel space-time view synthesis for monocular video is a highly challenging task: both static and dynamic objects usually appear in the video, but only a single view of the current scene is available, resulting in inaccurate synthesis results. To address this challenge, we propose FRNeRF, a novel space-time view synthesis method with a fusion regularization field. Specifically, we design a 2D-3D fusion regularization field for the original dynamic neural field, which helps reduce blurring of dynamic objects in the scene. In addition, we add image prior features to the hierarchical sampling to solve the problem that the traditional hierarchical sampling strategy cannot obtain sufficient sampling points during training. We evaluate our method extensively on multiple datasets and show the results of dynamic space-time view synthesis. Our method achieves state-of-the-art performance both qualitatively and quantitatively. Code is available for research purposes at https://cic.tju.edu.cn/faculty/likun/projects/FRNerf. Xinyi Jing, Tao Yu 0007, Renyuan He, Yukun Lai, Kun Li 0001 |
Comput. Vis. Media | 5 |
| 2025 | JASRNet: Learning Joint Adaptive Sampling and Reconstruction for Depth SensingabstractRecent attempts to exploit irregular sampling strategies for depth sensing have shown prominent merits over the uniform rectangular sampling in terms of depth reconstruction quality, particularly at low sampling rates. However, the separate treatment of depth sampling and reconstruction did not enjoy potential merits of joint optimization. In this article, we propose a joint adaptive depth sampling and reconstruction network, named JASRNet , for the RGB-D sensing configuration, to simultaneously optimize both the sampling and reconstruction of the depth information in an end-to-end manner. The sampling sub-network infers the locations to sample according to the significance distribution generated from the associated RGB image without any prior information of the underlying depth maps. The depth reconstruction sub-network learns and then fuses global and local depth features with attention guidance, which helps to obtain more accurate depth reconstruction results at boundaries. A hybrid loss function is further proposed to promote sharp discontinuities of the reconstructed depth maps. The qualitative and quantitative results show that our method achieves better depth sensing quality than several state-of-the-art methods for various indoor and outdoor scenes. Chunyang Bi, Mingnuo Teng, Tianhao Xie, Kun Li 0001, Jing-Yu Yang 0002 |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2025 | High-Quality Reconstruction of Depth Maps From Graph-Based Non-Uniform SamplingabstractDepth sensing is essential for intelligent computer vision applications, but it often suffers from low range precision and spatial resolution. To address this problem, we propose a novel framework that combines non-uniform sampling and reconstruction based on graph theory. Our framework consists of two main components: (1) a graph Laplacian induced non-uniform sampling (GLINUS) scheme that samples depth signals more densely around edges and contours than in smooth regions, and (2) an ensemble of priors (EoP) model that reconstructs the high-quality depth map using adaptive dual-tree discrete wavelet packets (ADDWP) transform, graph total variation regularizer, and graph Laplacian regularizer with color guidance. We solve the reconstruction problem using the alternating direction method of multipliers (ADMM). Our experiments demonstrate that our framework can capture fine structures and global information in depth signals and produce superior depth reconstruction results. Jing-Yu Yang 0002, Yusen Hou, Xinchen Ye, Pascal Frossard, Kun Li 0001 |
IEEE Trans. Multim. | 6 |
| 2025 | SpeechAct: Towards Generating Whole-Body Motion From SpeechabstractWhole-body motion generation from speech audio is crucial for computer graphics and immersive VR/AR. Prior methods struggle to produce natural and diverse whole-body motions from speech. In this paper, we introduce a novel method, named SpeechAct, based on a hybrid point representation and contrastive motion learning to boost realism and diversity in motion generation. Our hybrid point representation leverages the advantages of keypoint representation and surface points of 3D body model, which is easy to learn and helps to achieve smooth and natural motion generation from speech audio. We design a VQ-VAE to learn a motion codebook using our hybrid presentation, and then regress the motion from the input audio using a translation model. To boost diversity in motion generation, we propose a contrastive motion learning method according to the intuitive idea that the generated motion should be different from the motions of other audios and other speakers. We collect negative samples from other audio inputs and other speakers using our translation model. With these negative samples, we pull the current motion away from them using a contrastive loss to produce more distinctive representations. In addition, we compose a face generator to generate deterministic face motion due to the strong connection between the face movements and the speech audio. Experimental results validate the superior performance of our model. The code is available at http://cic.tju.edu.cn/faculty/likun/projects/SpeechAct/index.html. Minjie Zhu, Yuxiang Zhang 0006, Zerong Zheng, Yebin Liu, Kun Li 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | KeDuSR: Real-World Dual-Lens Super-Resolution via Kernel-Free MatchingabstractDual-lens super-resolution (SR) is a practical scenario for reference (Ref) based SR by utilizing the telephoto image (Ref) to assist the super-resolution of the low-resolution wide-angle image (LR input). Different from general RefSR, the Ref in dual-lens SR only covers the overlapped field of view (FoV) area. However, current dual-lens SR methods rarely utilize these specific characteristics and directly perform dense matching between the LR input and Ref. Due to the resolution gap between LR and Ref, the matching may miss the best-matched candidate and destroy the consistent structures in the overlapped FoV area. Different from them, we propose to first align the Ref with the center region (namely the overlapped FoV area) of the LR input by combining global warping and local warping to make the aligned Ref be sharp and consistent. Then, we formulate the aligned Ref and LR center as value-key pairs, and the corner region of the LR is formulated as queries. In this way, we propose a kernel-free matching strategy by matching between the LR-corner (query) and LR-center (key) regions, and the corresponding aligned Ref (value) can be warped to the corner region of the target. Our kernel-free matching strategy avoids the resolution gap between LR and Ref, which makes our network have better generalization ability. In addition, we construct a DuSR-Real dataset with (LR, Ref, HR) triples, where the LR and HR are well aligned. Experiments on three datasets demonstrate that our method outperforms the second-best method by a large margin. Our code and dataset are available at https://github.com/ZifanCui/KeDuSR. Huanjing Yue, Zifan Cui, Kun Li 0001, Jing-Yu Yang 0002 |
AAAI | 3 |
| 2024 | LPSNet: End-to-End Human Pose and Shape Estimation with Lensless ImagingabstractHuman pose and shape (HPS) estimation with lensless imaging is not only beneficial to privacy protection but also can be used in covert surveillance scenarios due to the small size and simple structure of this device. However, this task presents significant challenges due to the inherent ambiguity of the captured measurements and lacks effective methods for directly estimating human pose and shape from lensless data. In this paper, we propose the first end-to-end framework to recover 3D human poses and shapes from lensless measurements to our knowledge. We specifically design a multi-scale lensless feature decoder to decode the lensless measurements through the optically encoded mask for efficient feature extraction. We also propose a double-head auxiliary supervision mechanism to improve the estimation accuracy of human limb ends. Besides, we establish a lensless imaging system and verify the effectiveness of our method on various datasets acquired by our lensless imaging system. The code and dataset are available at https://cic.tju.edu.cn/faculty/likun/projects/LPSNet. Haoyang Ge, Qiao Feng 0001, Hailong Jia, Xiongzheng Li, Xiangjun Yin, Jing-Yu Yang 0002, Kun Li 0001 |
CVPR | 8 |
| 2024 | Joint2Human: High-quality 3D Human Generation via Compact Spherical Embedding of 3D Jointsabstract3D human generation is increasingly significant in var-ious applications. However, the direct use of 2D genera-tive methods in 3D generation often results in losing lo-cal details, while methods that reconstruct geometry from generated images struggle with global view consistency. In this work, we introduce joint2Human, a novel method that leverages 2D diffusion models to generate detailed 3D human geometry directly, ensuring both global structure and local details. To achieve this, we employ the Fourier occupancy field (FOF) representation, enabling the direct generation of 3D shapes as preliminary results with 2D generative models. With the proposed high-frequency enhancer and the multi-view recarving strategy, our method can seamlessly integrate the details from different views into a uniform global shape. To better utilize the 3D human prior and enhance control over the generated geometry, we introduce a compact spherical embedding of 3D joints. This allows for an effective guidance of pose during the gener-ation process. Additionally, our method can generate 3D humans guided by textual inputs. Our experimental results demonstrate the capability of our method to ensure global structure, local details, high resolution, and low computational cost simultaneously. More results and the code can be found on our project page at http://cic.tju.edu.cn/faculty/likun/projects/Joint2Human. Muxin Zhang, Qiao Feng 0001, Zhuo Su 0006, Zhou Xue, Kun Li 0001 |
CVPR | 6 |
| 2024 | HumanCoser: Layered 3D Human Generation via Semantic-Aware Diffusion ModelabstractThis paper aims to generate physically-layered 3D humans from text prompts. Existing methods either generate 3D clothed humans as a whole or support only tight and simple clothing generation, which limits their applications to virtual try-on and partlevel editing. To achieve physically-layered 3D human generation with reusable and complex clothing, we propose a novel layer-wise dressed human representation based on a physically-decoupled diffusion model. Specifically, to achieve layer-wise clothing generation, we propose a dual-representation decoupling framework for generating clothing decoupled from the human body, in conjunction with an innovative multi-layer fusion volume rendering method. To match the clothing with different body shapes, we propose an SMPL-driven implicit field deformation network that enables the free transfer and reuse of clothing. Extensive experiments demonstrate that our approach not only achieves state-of-the-art layered 3D human generation with complex clothing but also supports virtual try-on and layered human animation. More results and the code can be found on our project page at https: //cic.tju.edu.cn/faculty/likun/projects/HumanCoser Ruizhi Shao, Qiao Feng 0001, Yukun Lai, Kun Li 0001 |
ISMAR | 6 |
| 2024 | R2Human: Real-Time 3D Human Appearance Rendering from a Single ImageabstractRendering 3D human appearance from a single image in real-time is crucial for achieving holographic communication and immersive VR/AR. Existing methods either rely on multi-camera setups or are constrained to offline operations. In this paper, we propose R2Human, the first approach for real-time inference and rendering of photorealistic 3D human appearance from a single image. The core of our approach is to combine the strengths of implicit texture fields and explicit neural rendering with our novel representation, namely Z-map. Based on this, we present an end-to-end network that performs high-fidelity color reconstruction of visible areas and provides reliable color inference for occluded regions. To further enhance the 3D perception ability of our network, we leverage the Fourier occupancy field as a prior for generating the texture field and providing a sampling surface in the rendering stage. We also propose a consistency loss and a spatial fusion strategy to ensure the multi-view coherence. Experimental results show that our method outperforms the state-of-the-art methods on both synthetic data and challenging real-world images, in real-time. The project page can be found at http://cic.tju. edu.cn/faculty/likun/projects/R2Human. Yuanwang Yang, Qiao Feng 0001, Yukun Lai, Kun Li 0001 |
ISMAR | 4 |
| 2024 | High-Quality Animatable Dynamic Garment Reconstruction From Monocular VideosabstractMuch progress has been made in reconstructing garments from an image or a video. However, none of existing works meet the expectations of digitizing high-quality animatable dynamic garments that can be adjusted to various unseen poses. In this paper, we propose the first method to recover high-quality animatable dynamic garments from monocular videos without depending on scanned data. To generate reasonable deformations for various unseen poses, we propose a learnable garment deformation network that formulates the garment reconstruction task as a pose-driven deformation problem. To alleviate the ambiguity estimating 3D garments from monocular videos, we design a multi-hypothesis deformation module that learns spatial representations of multiple plausible deformations. Experimental results on several public datasets demonstrate that our method can reconstruct high-quality dynamic garments with coherent surface details, which can be easily animated under unseen poses. The code will be provided for research purposes. Xiongzheng Li, Yukun Lai, Jing-Yu Yang 0002, Kun Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Toward Grouping in Large Scenes With Occlusion-Aware Spatio-Temporal TransformersabstractGroup detection, especially for large-scale scenes, has many potential applications for public safety and smart cities. Existing methods fail to cope with frequent occlusions in large-scale scenes with multiple people, and are difficult to effectively utilize spatio-temporal information. In this paper, we propose an end-to-end framework,GroupTransformer, for group detection in large-scale scenes. To deal with the frequent occlusions caused by multiple people, we design an occlusion encoder to detect and suppress severely occluded person crops. To explore the potential spatio-temporal relationship, we propose spatio-temporal transformers to simultaneously extract trajectory information and fuse inter-person features in a hierarchical manner. Experimental results on both large-scale and small-scale scenes demonstrate that our method achieves better performance compared with state-of-the-art methods. On large-scale scenes, our method significantly boosts the performance in terms of precision and F1 score by more than 10%. On small-scale scenes, our method still improves the performance of F1 score by more than 5%.We will release the code for research purposes. Lingfeng Gu, Yukun Lai, Kun Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | HDhuman: High-Quality Human Novel-View Rendering From Sparse ViewsabstractIn this paper, we aim to address the challenge of novel view rendering of human performers that wear clothes with complex texture patterns using a sparse set of camera views. Although some recent works have achieved remarkable rendering quality on humans with relatively uniform textures using sparse views, the rendering quality remains limited when dealing with complex texture patterns as they are unable to recover the high-frequency geometry details that are observed in the input views. To this end, we propose HDhuman, which uses a human reconstruction network with a pixel-aligned spatial transformer and a rendering network with geometry-guided pixel-wise feature integration to achieve high-quality human reconstruction and rendering. The designed pixel-aligned spatial transformer calculates the correlations between the input views and generates human reconstruction results with high-frequency details. Based on the surface reconstruction results, the geometry-guided pixel-wise visibility reasoning provides guidance for multi-view feature integration, enabling the rendering network to render high-quality images at 2k resolution on novel views. Unlike previous neural rendering works that always need to train or fine-tune an independent network for a different scene, our method is a general framework that is able to generalize to novel subjects. Experiments show that our approach outperforms all the prior generic or specific methods on both synthetic data and real-world data. Source code and test data will be made publicly available for research purposes at http://cic.tju.edu.cn/faculty/likun/projects/HDhuman/index.html. Tiansong Zhou, Tao Yu 0007, Ruizhi Shao, Kun Li 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | Learning Semantic-Aware Disentangled Representation for Flexible 3D Human Body Editingabstract3D human body representation learning has received increasing attention in recent years. However, existing works cannot flexibly, controllably and accurately represent human bodies, limited by coarse semantics and unsatisfactory representation capability, particularly in the absence of supervised data. In this paper, we propose a human body representation with fine-grained semantics and high reconstruction-accuracy in an unsupervised setting. Specifically, we establish a correspondence between latent vectors and geometric measures of body parts by designing a part-aware skeleton-separated decoupling strategy, which facilitates controllable editing of human bodies by modifying the corresponding latent codes. With the help of a bone-guided auto-encoder and an orientation-adaptive weighting strategy, our representation can be trained in an unsupervised manner. With the geometrically meaningful latent space, it can be applied to a wide range of applications, from human body editing to latent code interpolation and shape style transfer. Experimental results on public datasets demonstrate the accurate reconstruction and flexible editing abilities of the proposed method. The code will be available at http://cic.tju.edu.cn/faculty/likun/projects/SemanticHuman. Xiaokun Sun, Qiao Feng 0001, Xiongzheng Li, Yukun Lai, Jing-Yu Yang 0002, Kun Li 0001 |
CVPR | 7 |
| 2023 | Crowd3D: Towards Hundreds of People Reconstruction from a Single ImageabstractImage-based multi-person reconstruction in wide-field large scenes is critical for crowd analysis and security alert. However, existing methods cannot deal with large scenes containing hundreds of people, which encounter the challenges of large number of people, large variations in human scale, and complex spatial distribution. In this paper, we propose Crowd3D, the first framework to reconstruct the 3D poses, shapes and locations of hundreds of people with global consistency from a single large-scene image. The core of our approach is to convert the problem of complex crowd localization into pixel localization with the help of our newly defined concept, Human-scene Virtual Interaction Point (HVIP). To reconstruct the crowd with global consistency, we propose a progressive reconstruction network based on HVIP by pre-estimating a scene-level camera and a ground plane. To deal with a large number of persons and various human sizes, we also design an adaptive human-centric cropping scheme. Besides, we contribute a benchmark dataset, LargeCrowd, for crowd reconstruction in a large scene. Experimental results demonstrate the effectiveness of the proposed method. The code and the dataset are available at http://cic.tju.edu.cn/faculty/likun/projects/Crowd3D. Huili Cui, Haozhe Lin, Yukun Lai, Lu Fang 0001, Kun Li 0001 |
CVPR | 7 |
| 2023 | Narrator: Towards Natural Control of Human-Scene Interaction Generation via Relationship ReasoningabstractNaturally controllable human-scene interaction (HSI) generation has an important role in various fields, such as VR/AR content creation and human-centered AI. However, existing methods are unnatural and unintuitive in their controllability, which heavily limits their application in practice. Therefore, we focus on a challenging task of naturally and controllably generating realistic and diverse HSIs from textual descriptions. From human cognition, the ideal generative model should correctly reason about spatial relationships and interactive actions. To that end, we propose Narrator, a novel relationship reasoning-based generative approach using a conditional variation autoencoder for naturally controllable generation given a 3D scene and a textual description. Also, we model global and local spatial relationships in a 3D scene and a textual description respectively based on the scene graph, and introduce a part-level action mechanism to represent interactions as atomic body part states. In particular, benefiting from our relationship reasoning, we further propose a simple yet effective multi-human generation strategy, which is the first exploration for controllable multi-human scene interaction generation. Our extensive experiments and perceptual studies show that Narrator can controllably generate diverse interactions and significantly outperform existing works. Haibiao Xuan, Xiongzheng Li, Hongwen Zhang 0001, Yebin Liu, Kun Li 0001 |
ICCV | 6 |
| 2023 | TSNeRF: Text-driven stylized neural radiance fields via semantic contrastive learning
Jing-Song Cheng, Qiao Feng 0001, Wen-Yuan Tao, Yukun Lai, Kun Li 0001 |
Comput. Graph. | 6 |
| 2023 | STATE: Learning structure and texture representations for novel view synthesisabstractNovel viewpoint image synthesis is very challenging, especially from sparse views, due to large changes in viewpoint and occlusion. Existing image-based methods fail to generate reasonable results for invisible regions, while geometry-based methods have difficulties in synthesizing detailed textures. In this paper, we propose STATE, an end-to-end deep neural network, for sparse view synthesis by learning structure and texture representations. Structure is encoded as a hybrid feature field to predict reasonable structures for invisible regions while maintaining original structures for visible regions, and texture is encoded as a deformed feature map to preserve detailed textures. We propose a hierarchical fusion scheme with intra-branch and inter-branch aggregation, in which spatio-view attention allows multi-view fusion at the feature level to adaptively select important information by regressing pixel-wise or voxel-wise confidence maps. By decoding the aggregated features, STATE is able to generate realistic images with reasonable structures and detailed textures. Experimental results demonstrate that our method achieves qualitatively and quantitatively better results than state-of-the-art methods. Our method also enables texture and structure editing applications benefiting from implicit disentanglement of structure and texture. Our code is available at http://cic.tju.edu.cn/faculty/likun/projects/STATE . Xinyi Jing, Qiao Feng 0001, Yukun Lai, Yuanqiang Yu, Kun Li 0001 |
Comput. Vis. Media | 6 |
| 2023 | Learning to Infer Inner-Body Under Clothing From Monocular VideoabstractAccurately estimating the human inner-body under clothing is very important for body measurement, virtual try-on and VR/AR applications. In this article, we propose the first method to allow everyone to easily reconstruct their own 3D inner-body under daily clothing from a self-captured video with the mean reconstruction error of 0.73cm within 15s. This avoids privacy concerns arising from nudity or minimal clothing. Specifically, we propose a novel two-stage framework with a Semantic-guided Undressing Network (SUNet) and an Intra-Inter Transformer Network (IITNet). SUNet learns semantically related body features to alleviate the complexity and uncertainty of directly estimating 3D inner-bodies under clothing. IITNet reconstructs the 3D inner-body model by making full use of intra-frame and inter-frame information, which addresses the misalignment of inconsistent poses in different frames. Experimental results on both public datasets and our collected dataset demonstrate the effectiveness of the proposed method. The code and dataset is available for research purposes at http://cic.tju.edu.cn/faculty/likun/projects/Inner-Body. Xiongzheng Li, Xiaokun Sun, Haibiao Xuan, Yukun Lai, Yingdi Xie, Jing-Yu Yang 0002, Kun Li 0001 |
IEEE Trans. Vis. Comput. Graph. | 9 |
| 2022 | High-Fidelity Human Avatars from a Single RGB CameraabstractIn this paper, we propose a coarse-to-fine framework to reconstruct a personalized high-fidelity human avatar from a monocular video. To deal with the misalignment problem caused by the changed poses and shapes in different frames, we design a dynamic surface network to recover pose-dependent surface deformations, which help to decouple the shape and texture of the person. To cope with the complexity of textures and generate photo-realistic results, we propose a reference-based neural rendering network and exploit a bottom-up sharpening-guided fine-tuning strategy to obtain detailed textures. Our frame-work also enables photo-realistic novel view/pose syn-thesis and shape editing applications. Experimental re-sults on both the public dataset and our collected dataset demonstrate that our method outperforms the state-of-the-art methods. The code and dataset will be available at http://cic.tju.edu.cn/faculty/likun/projects/HF-Avatar. Yukun Lai, Zerong Zheng, Yingdi Xie, Yebin Liu, Kun Li 0001 |
CVPR | 7 |
| 2022 | FOF: Learning Fourier Occupancy Field for Monocular Real-time Human ReconstructionabstractThe advent of deep learning has led to significant progress in monocular human reconstruction. However, existing representations, such as parametric models, voxel grids, meshes and implicit neural representations, have difficulties achieving high-quality results and real-time speed at the same time. In this paper, we propose Fourier Occupancy Field (FOF), a novel, powerful, efficient and flexible 3D geometry representation, for monocular real-time and accurate human reconstruction. A FOF represents a 3D object with a 2D field orthogonal to the view direction where at each 2D position the occupancy field of the object along the view direction is compactly represented with the first few terms of Fourier series, which retains the topology and neighborhood relation in the 2D domain. A FOF can be stored as a multi-channel image, which is compatible with 2D convolutional neural networks and can bridge the gap between 3D geometries and 2D images. A FOF is very flexible and extensible, \eg, parametric models can be easily integrated into a FOF as a prior to generate more robust results. Meshes and our FOF can be easily inter-converted. Based on FOF, we design the first 30+FPS high-fidelity real-time monocular human reconstruction framework. We demonstrate the potential of FOF on both public datasets and real captured data. The code is available for research purposes at http://cic.tju.edu.cn/faculty/likun/projects/FOF. Qiao Feng 0001, Yebin Liu, Yukun Lai, Jing-Yu Yang 0002, Kun Li 0001 |
NeurIPS | 5 |
| 2022 | Cloud Detection From Remote Sensing Imagery Based on Domain Translation NetworkabstractCloud detection in optical imagery has drawn remarkable attention in the era of big Earth observation data analytic. While multiple supervised learning models have been developed for such purpose, large volumes of paired training samples annotated at the pixel level are essential to ensure the model’s generalization capacity. However, constructing a comprehensive cloud detection training database is a tedious and time-consuming process. To tackle this dilemma, we simply regard cloud-contaminated remote sensing (RS) imagery as the combination of cloud and background domains and propose a cloud detection framework based on image-to-image domain translation network (DTNet) to separate cloud-contaminated RS imagery into two target domains of cloud and background object images without using any paired and pixel-level annotation training data. The framework was evaluated with multispectral images from two types of sensors, Landsat-8 Operational Land Imager (OLI) (30 m) and GaoFen-1 (16 m), and demonstrated superior or comparable performance compared with several state-of-the-art cloud detection models. Jianhua Guo 0002, Jing-Yu Yang 0002, Huanjing Yue, Yang Chen 0015, Chunping Hou, Kun Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | Unsupervised Domain Adaptation for Cloud Detection Based on Grouped Features Alignment and Entropy MinimizationabstractMost convolutional neural network (CNN)-based cloud detection methods are built upon the supervised learning framework that requires a large number of pixel-level labels. However, it is expensive and time-consuming to manually annotate pixelwise labels for massive remote sensing images. To reduce the labeling cost, we propose an unsupervised domain adaptation (UDA) approach to generalize the model trained on labeled images of source satellite to unlabeled images of the target satellite. To effectively address the domain shift problem on cross-satellite images, we develop a novel UDA method based on grouped features alignment (GFA) and entropy minimization (EM) to extract domain-invariant representations to improve the cloud detection accuracy of cross-satellite images. The proposed UDA method is evaluated on “Landsat-$8~\rightarrow $ZY-3” and “GF-$1\rightarrow $ZY-3” domain adaptation tasks. Experimental results demonstrate the effectiveness of our method against existing state-of-the-art UDA approaches. The code of this paper has been made available online (https://github.com/nkszjx/grouped-features-alignment). Jianhua Guo 0002, Jing-Yu Yang 0002, Huanjing Yue, Kun Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Unsupervised Domain-Invariant Feature Learning for Cloud Detection of Remote Sensing ImagesabstractThe detection of clouds in remote sensing (RS) images is an important task, and convolutional neural networks (CNNs) have been used to perform it. However, supervised cloud detection CNNs rely heavily on a large number of samples annotated at pixel level to tune their parameter. Annotating RS images is a labor-intensive procedure and requires expert-level human knowledge. To reduce the labeling cost, we propose an unsupervised domain adaptation (UDA) approach to enable the model trained on labeled source satellite images to generalize to unlabeled target satellite images. Specifically, we propose a fine-grained feature alignment (FGFA) domain adaptation strategy that encourages a cloud detection network to extract domain-invariant representations, which improves the accuracy of cloud detection in unlabeled target satellite images. The proposed FGFA strategy consists of two steps: 1) fine-grained class-relevant feature selection based on an attention-guided mechanism and 2) class-relevant feature alignment (FA) based on a proposed grouped FA approach. Experimental results on the “Landsat-$8~\rightarrow $ZY-3” and “GF-$1\rightarrow $ZY-3” domain adaptation tasks demonstrate the effectiveness of our method and its superiority to existing state-of-the-art UDA approaches. Jianhua Guo 0002, Jing-Yu Yang 0002, Huanjing Yue, Xin Liu 0012, Kun Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Geometry-Guided Dense Perspective Network for Speech-Driven Facial AnimationabstractRealistic speech-driven 3D facial animation is a challenging problem due to the complex relationship between speech and face. In this paper, we propose a deep architecture, called Geometry-guided Dense Perspective Network (GDPnet), to achieve speaker-independent realistic 3D facial animation. The encoder is designed with dense connections to strengthen feature propagation and encourage the re-use of audio features, and the decoder is integrated with an attention mechanism to adaptively recalibrate point-wise feature responses by explicitly modeling interdependencies between different neuron units. We also introduce a non-linear face reconstruction representation as a guidance of latent space to obtain more accurate deformation, which helps solve the geometry-related deformation and is good for generalization across subjects. Huber and HSIC (Hilbert-Schmidt Independence Criterion) constraints are adopted to promote the robustness of our model and to better exploit the non-linear and high-order correlations. Experimental results on the public dataset and real scanned dataset validate the superiority of our proposed GDPnet compared with state-of-the-art model. The code is available for research purposes at http://cic.tju.edu.cn/faculty/likun/projects/GDPnet. Jingying Liu, Binyuan Hui, Kun Li 0001, Yunke Liu, Yukun Lai, Yuxiang Zhang 0006, Yebin Liu, Jing-Yu Yang 0002 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2021 | PISE: Person Image Synthesis and Editing With Decoupled GANabstractPerson image synthesis, e.g., pose transfer, is a challenging problem due to large variation and occlusion. Existing methods have difficulties predicting reasonable invisible regions and fail to decouple the shape and style of clothing, which limits their applications on person image editing. In this paper, we propose PISE, a novel two-stage generative model for Person Image Synthesis and Editing, which is able to generate realistic person images with desired poses, textures, or semantic layouts. For human pose transfer, we first synthesize a human parsing map aligned with the target pose to represent the shape of clothing by a parsing generator, and then generate the final image by an image generator. To decouple the shape and style of clothing, we propose joint global and local per-region encoding and normalization to predict the reasonable style of clothing for invisible regions. We also propose spatial-aware normalization to retain the spatial context relationship in the source image. The results of qualitative and quantitative experiments demonstrate the superiority of our model on human pose transfer. Besides, the results of texture transfer and region editing show that our model can be applied to person image editing. The code is available for research purposes at https://github.com/Zhangjinso/PISE. Kun Li 0001, Yukun Lai, Jing-Yu Yang 0002 |
CVPR | 2 |
| 2021 | Cross-MPI: Cross-Scale Stereo for Image Super-Resolution Using Multiplane ImagesabstractVarious combinations of cameras enrich computational photography, among which reference-based super-resolution (RefSR) plays a critical role in multiscale imaging systems. However, existing RefSR approaches fail to accomplish high-fidelity super-resolution under a large resolution gap, e.g., 8× upscaling, due to the lower consideration of the underlying scene structure. In this paper, we aim to solve the RefSR problem in actual multiscale camera systems inspired by multiplane image (MPI) representation. Specifically, we propose Cross-MPI, an end-to-end RefSR network composed of a novel plane-aware attention-based MPI mechanism, a multiscale guided upsampling module as well as a super-resolution (SR) synthesis and fusion module. Instead of using a direct and exhaustive matching between the cross-scale stereo, the proposed plane-aware attention mechanism fully utilizes the concealed scene structure for efficient attention-based correspondence searching. Further combined with a gentle coarse-to-fine guided upsampling strategy, the proposed Cross-MPI can achieve a robust and accurate detail transmission. Experimental results on both digitally synthesized and optical zoom cross-scale data show that the Cross-MPI framework can achieve superior performance against the existing RefSR methods and is a real fit for actual multiscale camera systems even with large-scale differences. Yuemei Zhou, Gaochang Wu, Ying Fu 0001, Kun Li 0001, Yebin Liu |
CVPR | 4 |
| 2021 | Implicit Transformer Network for Screen Content Image Continuous Super-ResolutionabstractNowadays, there is an explosive growth of screen contents due to the wide application of screen sharing, remote cooperation, and online education. To match the limited terminal bandwidth, high-resolution (HR) screen contents may be downsampled and compressed. At the receiver side, the super-resolution (SR)of low-resolution (LR) screen content images (SCIs) is highly demanded by the HR display or by the users to zoom in for detail observation. However, image SR methods mostly designed for natural images do not generalize well for SCIs due to the very different image characteristics as well as the requirement of SCI browsing at arbitrary scales. To this end, we propose a novel Implicit Transformer Super-Resolution Network (ITSRN) for SCISR. For high-quality continuous SR at arbitrary ratios, pixel values at query coordinates are inferred from image features at key coordinates by the proposed implicit transformer and an implicit position encoding scheme is proposed to aggregate similar neighboring pixel values to the query one. We construct benchmark SCI1K and SCI1K-compression datasets withLR and HR SCI pairs. Extensive experiments show that the proposed ITSRN significantly outperforms several competitive continuous and discrete SR methods for both compressed and uncompressed SCIs. Jing-Yu Yang 0002, Sheng Shen 0010, Huanjing Yue, Kun Li 0001 |
NeurIPS | 4 |
| 2021 | Low-light image enhancement based on Retinex decomposition and adaptive gamma correctionabstractAbstract Low‐light images suffer from poor visibility and noise. In this paper, a low‐light image enhancement method based on Retinex decomposition is proposed. A pyramid network is first utilized to extract multi‐scale features to improve the quality of Retinex decomposition. Then the decomposed illumination is refined via an adaptive Gamma correction network to handle non‐uniform illumination, while the decomposed reflectance is refined with a lightweight network. Finally, the enhanced image is obtained by element‐wise multiplication between the refined illumination and reflectance components. Quantitative and qualitative experiments demonstrate the superiority of our method over state‐of‐the‐art image enhancement methods. Jing-Yu Yang 0002, Huanjing Yue, Zhongyu Jiang, Kun Li 0001 |
IET Image Process. | 5 |
| 2021 | Sparse intrinsic decomposition and applications
Kun Li 0001, Xinchen Ye, Chenggang Yan 0001, Jing-Yu Yang 0002 |
Signal Process. Image Commun. | 1 |
| 2021 | Landsat-8 OLI Multispectral Image Dehazing Based on Optimized Atmospheric Scattering ModelabstractOptical satellite images are often affected by haze atmospheric conditions, which degrades the quality of remote sensing (RS) data and reduces the accuracy of interpretation and classification. Hence, haze removal becomes a necessary preprocessing step for most of the applications of RS image. In this article, we propose a novel haze removal method for Landsat-8 OLI multispectral image based on an optimized atmospheric scattering model. We focus on adaptively estimating the haze transmission map of each band by taking into account the effect of both wavelength and haze atmospheric conditions (haze particle size and haze particle concentration) thus improving dehazing performance. The experimental results on Landsat-8 OLI multispectral images show that the proposed dehazing model is able to remove haze successfully and significantly improve the image visibility as well as correct the spectral bias to some degree. Moreover, this method is simple and feasible, and has good practical value. Jianhua Guo 0002, Jing-Yu Yang 0002, Huanjing Yue, Chunping Hou, Kun Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | CDnetV2: CNN-Based Cloud Detection for Remote Sensing Imagery With Cloud-Snow CoexistenceabstractCloud detection is a crucial preprocessing step for optical satellite remote sensing (RS) images. This article focuses on the cloud detection for RS imagery with cloud-snow coexistence and the utilization of the satellite thumbnails that lose considerable amount of high resolution and spectrum information of original RS images to extract cloud mask efficiently. To tackle this problem, we propose a novel cloud detection neural network with an encoder-decoder structure, named CDnetV2, as a series work on cloud detection. Compared with our previous CDnetV1, CDnetV2 contains two novel modules, that is, adaptive feature fusing model (AFFM) and high-level semantic information guidance flows (HSIGFs). AFFM is used to fuse multilevel feature maps by three submodules: channel attention fusion model (CAFM), spatial attention fusion model (SAFM), and channel attention refinement model (CARM). HSIGFs are designed to make feature layers at decoder of CDnetV2 be aware of the locations of the cloud objects. The high-level semantic information of HSIGFs is extracted by a proposed high-level feature fusing model (HFFM). By being equipped with these two proposed key modules, AFFM and HSIGFs, CDnetV2 is able to fully utilize features extracted from encoder layers and yield accurate cloud detection results. Experimental results on the ZY-3 satellite thumbnail data set demonstrate that the proposed CDnetV2 achieves accurate detection accuracy and outperforms several state-of-the-art methods. Jianhua Guo 0002, Jing-Yu Yang 0002, Huanjing Yue, Chunping Hou, Kun Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2021 | RSDehazeNet: Dehazing Network With Channel Refinement for Multispectral Remote Sensing ImagesabstractMultispectral remote sensing (RS) images are often contaminated by the haze that degrades the quality of RS data and reduces the accuracy of interpretation and classification. Recently, the emerging deep convolutional neural networks (CNNs) provide us new approaches for RS image dehazing. Unfortunately, the power of CNNs is limited by the lack of sufficient hazy-clean pairs of RS imagery, which makes supervised learning impractical. To meet the data hunger of supervised CNNs, we propose a novel haze synthesis method to generate realistic hazy multispectral images by modeling the wavelength-dependent and spatial-varying characteristics of haze in RS images. The proposed haze synthesis method not only alleviates the lack of realistic training pairs in multispectral RS image dehazing but also provides a benchmark data set for quantitative evaluation. Furthermore, we propose an end-to-end RSDehazeNet for haze removal. We utilize both local and global residual learning strategies in RSDehazeNet for fast convergence with superior performance. Channel attention modules are incorporated to exploit strong channel correlation in multispectral RS images. Experimental results show that the proposed network outperforms the state-of-the-art methods for synthetic data and real Landsat-8 OLI multispectral RS images. Jianhua Guo 0002, Jing-Yu Yang 0002, Huanjing Yue, Chunping Hou, Kun Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2021 | Image-Guided Human Reconstruction via Multi-Scale Graph Transformation Networksabstract3D human reconstruction from a single image is a challenging problem. Existing methods have difficulties to infer 3D clothed human models with consistent topologies for various poses. In this paper, we propose an efficient and effective method using a hierarchical graph transformation network. To deal with large deformations and avoid distorted geometries, rather than using Euclidean coordinates directly, 3D human shapes are represented by a vertex-based deformation representation that effectively encodes the deformation and copes well with large deformations. To infer a 3D human mesh consistent with the input real image, we also use a perspective projection layer to incorporate perceptual image features into the deformation representation. Our model is easy to train and fast to converge with short test time. Besides, we present the$D^{2}Human$(Dynamic Detailed Human) dataset, including variously posed 3D human meshes with consistent topologies and rich geometry details, together with the captured color images and SMPL models, which is useful for training and evaluation of deep frameworks, particularly for graph neural networks. Experimental results demonstrate that our method achieves more plausible and complete 3D human reconstruction from a single image, compared with several state-of-the-art methods. The code and dataset are available for research purposes athttp://cic.tju.edu.cn/faculty/likun/projects/MGTnet. Kun Li 0001, Qiao Feng 0001, Yuxiang Zhang 0006, Xiongzheng Li, Cunkuan Yuan, Yukun Lai, Yebin Liu |
IEEE Trans. Image Process. | 1 |
| 2020 | 4D Association Graph for Realtime Multi-Person Motion Capture Using Multiple Video Camerasabstracthis paper contributes a novel realtime multi-person motion capture algorithm using multiview video inputs. Due to the heavy occlusions and closely interacting motions in each view, joint optimization on the multiview images and multiple temporal frames is indispensable, which brings up the essential challenge of realtime efficiency. To this end, for the first time, we unify per-view parsing, cross-view matching, and temporal tracking into a single optimization framework, i.e., a 4D association graph that each dimension (image space, viewpoint and time) can be treated equally and simultaneously. To solve the 4D association graph efficiently, we further contribute the idea of 4D limb bundle parsing based on heuristic searching, followed with limb bundle assembling by proposing a bundle Kruskal's algorithm. Our method enables a realtime motion capture system running at 30fps using 5 cameras on a 5-person scene. Benefiting from the unified parsing, matching and tracking constraints, our method is robust to noisy detection due to severe occlusions and close interacting motions, and achieves high-quality online pose reconstruction quality. The proposed method outperforms state-of-the-art methods quantitatively without using high-level appearance information. Yuxiang Zhang 0006, Liang An 0001, Tao Yu 0007, Xiu Li 0001, Kun Li 0001, Yebin Liu |
CVPR | 5 |
| 2020 | 3D Motion Recovery via Low Rank Matrix Restoration with Hankel-Like AugmentationabstractThis paper proposes a 3D skeleton recovery model equipped with a joint augmented low-rank and sparse prior and an articulation-graph-based isometric constraint to exploit temporal and spatial correlation, respectively. The corrupted 3D skeleton sequence is represented as a matrix and constrained by several priors in the proposed model. A Hankel-like augmentation is adopted to strengthen the low-rankness and we integrate a decoupling technique to reduce the internal interferences of the data. We solve our model via an alternating direction method under the augmented Lagrangian multiplier framework with a Gauss-Newton solver for the subproblem of isometric optimization. Experimental results on two skeleton datasets demonstrate the effectiveness and superiority of the proposed model in motion reconstruction and skeleton recovery, compared with state-of-the-art methods. Jing-Yu Yang 0002, Jiabin Shi, Yuyuan Zhu, Kun Li 0001, Chunping Hou |
ICME | 4 |
| 2020 | GPS-Net: Graph-based Photometric Stereo NetworkabstractLearning-based photometric stereo methods predict the surface normal either in a per-pixel or an all-pixel manner. Per-pixel methods explore the inter-image intensity variation of each pixel but ignore features from the intra-image spatial domain. All-pixel methods explore the intra-image intensity variation of each input image but pay less attention to the inter-image lighting variation. In this paper, we present a Graph-based Photometric Stereo Network, which unifies per-pixel and all-pixel processings to explore both inter-image and intra-image information. For per-pixel operation, we propose the Unstructured Feature Extraction Layer to connect an arbitrary number of input image-light pairs into graph structures, and introduce Structure-aware Graph Convolution filters to balance the input data by appropriately weighting shadows and specular highlights. For all-pixel operation, we propose the Normal Regression Network to make efficient use of the intra-image spatial information for predicting a surface normal map with rich details. Experimental results on the real-world benchmark show that our method achieves excellent performance under both sparse and dense lighting distributions. Zhuokun Yao, Kun Li 0001, Ying Fu 0001, Haofeng Hu, Boxin Shi |
NeurIPS | 2 |
| 2020 | SHREC'20: Shape correspondence with non-isometric deformations abstractEstimating correspondence between two shapes continues to be a challenging problem in geometry processing. Most current methods assume deformation to be near-isometric, however this is often not the case. For this paper, a collection of shapes of different animals has been curated, where parts of the animals (e.g., mouths, tails & ears) correspond yet are naturally non-isometric. Ground-truth correspondences were established by asking three specialists to independently label corresponding points on each of the models with respect to a previously labelled reference model. We employ an algorithmic strategy to select a single point for each correspondence that is representative of the proposed labels. A novel technique that characterises the sparsity and distribution of correspondences is employed to measure the performance of ten shape correspondence methods. Roberto M. Dyke, Yukun Lai, Paul L. Rosin, Stefano Zappalà, Seana Dykes, Daoliang Guo, Kun Li 0001, Riccardo Marin, Simone Melzi, Jing-Yu Yang 0002 |
Comput. Graph. | 7 |
| 2020 | Human Pose Transfer by Adaptive Hierarchical DeformationabstractAbstract Human pose transfer, as a misaligned image generation task, is very challenging. Existing methods cannot effectively utilize the input information, which often fail to preserve the style and shape of hair and clothes. In this paper, we propose an adaptive human pose transfer network with two hierarchical deformation levels. The first level generates human semantic parsing aligned with the target pose, and the second level generates the final textured person image in the target pose with the semantic guidance. To avoid the drawback of vanilla convolution that treats all the pixels as valid information, we use gated convolution in both two levels to dynamically select the important features and adaptively deform the image layer by layer. Our model has very few parameters and is fast to converge. Experimental results demonstrate that our model achieves better performance with more consistent hair, face and clothes with fewer parameters than state‐of‐the‐art methods. Furthermore, our method can be applied to clothing texture transfer. The code is available for research purposes at https://github.com/Zhangjinso/PINet_PG . Xingzi Liu, Kun Li 0001 |
Comput. Graph. Forum | 3 |
| 2020 | Full-body motion capture for multiple closely interacting persons
Kun Li 0001, Yali Mao, Yunke Liu, Ruizhi Shao, Yebin Liu |
Graph. Model. | 1 |
| 2020 | Spatiotemporally scalable matrix recovery for background modeling and moving object detection
Jing-Yu Yang 0002, Huanjing Yue, Kun Li 0001, Chunping Hou |
Signal Process. | 4 |
| 2020 | Spatio-Temporal Reconstruction for 3D Motion RecoveryabstractThis paper addresses the challenge of 3D motion recovery by exploiting the spatio-temporal correlations of corrupted 3D skeleton sequences. We propose a new 3D motion recovery method using spatio-temporal reconstruction, which uses joint low-rank and sparse priors to exploit temporal correlation and an isometric constraint for spatial correlation. The proposed model is formulated as a constrained optimization problem, which is efficiently solved by the augmented Lagrangian method with a Gauss-Newton solver for the subproblem of isometric optimization. The experimental results on the CMU motion capture dataset, Edinburgh dataset, and two Kinect datasets demonstrate that the proposed approach achieves better motion recovery than the state-of-the-art methods. The proposed method is applicable to Kinect-like skeleton tracking devices and pose estimation methods that cannot provide accurate estimation of complex motions, especially in the presence of occlusion. Jing-Yu Yang 0002, Kun Li 0001, Meiyuan Wang, Yukun Lai, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Discern Depth Under Foul Weather: Estimate PM2.5 for Depth InferenceabstractNowadays, haze is a common and serious problem and PM$_{2.5}$is a main measurement for air quality. Current methods estimate the level of primary pollutant with professional instruments, which is expensive and inconvenient. Moreover, with haze, the captured images will be unclear and are difficult to estimate the depth of the scene using passive methods. This article proposes a cheap, fast, and convenient PM$_{2.5}$estimation method that only need a captured image using daily-life devices, and further, discerns the depth of the scene using the estimated PM$_{2.5}$. We learn haze-relevant classified mapping via the hybrid convolutional neural network and combine the high-level features extracted from the convolutional layer with ground-truth PM$_{2.5}$to train support vector regression. The transmission map is computed using nonlocal sparse priors, and the depth map is inferred using the estimated PM$_{2.5}$value through the atmospheric scattering model. Experimental results demonstrate that our method achieves accurate PM$_{2.5}$estimation and depth inference. This could be very useful in many applications, for both clean and foul weather. Kun Li 0001, Yahong Han, Xibin Yue, Jing-Yu Yang 0002 |
IEEE Trans. Ind. Informatics | 1 |
| 2020 | PoNA: Pose-Guided Non-Local Attention for Human Pose TransferabstractHuman pose transfer, which aims at transferring the appearance of a given person to a target pose, is very challenging and important in many applications. Previous work ignores the guidance of pose features or only uses local attention mechanism, leading to implausible and blurry results. We propose a new human pose transfer method using a generative adversarial network (GAN) with simplified cascaded blocks. In each block, we propose a pose-guided non-local attention (PoNA) mechanism with a long-range dependency scheme to select more important regions of image features to transfer. We also design pre-posed image-guided pose feature update and post-posed pose-guided image feature update to better utilize the pose and image features. Our network is simple, stable, and easy to train. Quantitative and qualitative results on Market-1501 and DeepFashion datasets show the efficacy and efficiency of our model. Compared with state-of-the-art methods, our model generates sharper and more realistic images with rich details, while having fewer parameters and faster speed. Furthermore, our generated images can help to alleviate data insufficiency for person re-identification. Kun Li 0001, Yebin Liu, Yukun Lai, Qionghai Dai |
IEEE Trans. Image Process. | 1 |
| 2020 | Learning to Reconstruct and Understand Indoor Scenes From Sparse ViewsabstractThis paper proposes a new method for simultaneous 3D reconstruction and semantic segmentation for indoor scenes. Unlike existing methods that require recording a video using a color camera and/or a depth camera, our method only needs a small number of (e.g., 3~5) color images from uncalibrated sparse views, which significantly simplifies data acquisition and broadens applicable scenarios. To achieve promising 3D reconstruction from sparse views with limited overlap, our method first recovers the depth map and semantic information for each view, and then fuses the depth maps into a 3D scene. To this end, we design an iterative deep architecture, named IterNet, to estimate the depth map and semantic segmentation alternately. To obtain accurate alignment between views with limited overlap, we further propose a joint global and local registration method to reconstruct a 3D scene with semantic information. We also make available a new indoor synthetic dataset, containing photorealistic high-resolution RGB images, accurate depth maps and pixel-level semantic labels for thousands of complex layouts. Experimental results on public datasets and our dataset demonstrate that our method achieves more accurate depth estimation, smaller semantic segmentation errors, and better 3D reconstruction results over state-of-the-art methods. Jing-Yu Yang 0002, Kun Li 0001, Yukun Lai, Huanjing Yue, Jianzhi Lu, Hao Wu 0042, Yebin Liu |
IEEE Trans. Image Process. | 3 |
| 2020 | View Synthesis from multi-view RGB data using multilayered representation and volumetric estimationabstractAiming at free-view exploration of complicated scenes, this paper presents a method for interpolating views among multi RGB cameras. In this study, we combine the idea of cost volume, which represent 3D information, and 2D semantic segmentation of the scene, to accomplish view synthesis of complicated scenes. We use the idea of cost volume to estimate the depth and confidence map of the scene, and use a multi-layer representation and resolution of the data to optimize the view synthesis of the main object. /Conclusions By applying different treatment methods on different layers of the volume, we can handle complicated scenes containing multiple persons and plentiful occlusions. We also propose the view-interpolation→multi-view reconstruction→view interpolation pipeline to iteratively optimize the result. We test our method on varying data of multi-view scenes and generate decent results. Zhaoqi Su, Tiansong Zhou, Kun Li 0001, David J. Brady, Yebin Liu |
Virtual Real. Intell. Hardw. | 3 |
| 2019 | Graph Based Non-Uniform Sampling and Reconstruction of Depth MapsabstractHigh-quality depth sensing is highly demanded in intelligent computer vision, 3DTV, and many other related fields. However, prevalent time-of-fly (ToF) depth sensors are of low resolution as the number of pixel-level demodulators is limited. Moreover, the rectangular sampling does not consider the signal characteristics of depth maps. Being a departure of previous resolution enhancement on rectangular sampling, this paper investigates the non-uniform sampling of depth maps, and the high-resolution depth reconstruction from limited non-uniformly distributed samples. The proposed depth sampling and reconstruction schemes are developed based on graph signal processing. We first propose a graph-based non-uniform sampling (GNS) scheme, where depth signals are sampled based on the response of a high-pass graph filter, which results in denser sampling around discontinuities such as edges and contours than in smooth regions. We then propose a graph-based depth reconstruction (GDR) framework,where a graph Laplacian regularizer is designed to fully exploit structural correlation between the depth and photometric images. To solve the reconstruction problem, we derive an efficient algorithm based on the alternating direction method of multipliers (ADMM). Experimental results show that the GNS-GDR non-uniform sampling and reconstruction method achieves high-quality depth sensing, outperforming several state-of-the-art schemes. Jing-Yu Yang 0002, Xinchen Ye, Pascal Frossard, Kun Li 0001 |
ICIP | 5 |
| 2019 | Global as-Conformal-as-Possible Non-Rigid Registration of Multi-view ScansabstractIn this paper, we present a novel framework for global non-rigid registration of multi-view scans captured using consumer-level depth cameras. In our method, all scans from different viewpoints are allowed to undergo large non-rigid deformations and finally fused into a complete high quality model. To avoid the well-known loop closure problem, we simultaneously optimize a global alignment problem instead of pairwise non-rigid registration in succession. We employ a joint point-to-point and point-to-plane positional constraint to reduce the influence of wrong correspondences, and incorporate an as-conformal-as-possible constraint to avoid mesh distortions during deformation. We also design a reweighting scheme on position and transformation to reduce registration errors. Experimental results on both public datasets and real scanned datasets demonstrate that our approach outperforms state-of-the-art methods through extensive quantitative and qualitative evaluations. Zhenchao Wu, Kun Li 0001, Yukun Lai, Jing-Yu Yang 0002 |
ICME | 2 |
| 2019 | 3D Face Reprentation and Reconstruction with Multi-scale Graph Convolutional AutoencodersabstractEffective representation and reconstruction for human faces are very important in many applications. Existing linear representation methods cannot reconstruct high quality 3D faces with details, while the newest non-linear representation method is less suitable for real shapes since spectral decompositions are unstable across different graphs. To address these problems, we propose a multi-scale graph convolutional autoencoder for face representation and reconstruction. Our autoencoder uses graph convolution, which is easily trained for the data with graph structures and can be used for other deformable models. Our model can also be used for variational training to generate high quality face shapes. Experimental results demonstrate that our model can generate more plausible, complex, and stable 3D shapes, and achieves higher quality face reconstruction compared with state-of-the-art methods. Cunkuan Yuan, Kun Li 0001, Yukun Lai, Yebin Liu, Jing-Yu Yang 0002 |
ICME | 2 |
| 2019 | Generating 3D Faces using Multi-column Graph Convolutional NetworksabstractAbstract In this work, we introduce multi‐column graph convolutional networks (MGCNs), a deep generative model for 3D mesh surfaces that effectively learns a non‐linear facial representation. We perform spectral decomposition of meshes and apply convolutions directly in the frequency domain. Our network architecture involves multiple columns of graph convolutional networks (GCNs), namely large GCN (L‐GCN), medium GCN (M‐GCN) and small GCN (S‐GCN), with different filter sizes to extract features at different scales. L‐GCN is more useful to extract large‐scale features, whereas S‐GCN is effective for extracting subtle and fine‐grained features, and M‐GCN captures information in between. Therefore, to obtain a high‐quality representation, we propose a selective fusion method that adaptively integrates these three kinds of information. Spatially non‐local relationships are also exploited through a self‐attention mechanism to further improve the representation ability in the latent vector space. Through extensive experiments, we demonstrate the superiority of our end‐to‐end framework in improving the accuracy of 3D face reconstruction. Moreover, with the help of variational inference, our model has excellent generating ability. Kun Li 0001, Jingying Liu, Yukun Lai, Jing-Yu Yang 0002 |
Comput. Graph. Forum | 1 |
| 2019 | CDnet: CNN-Based Cloud Detection for Remote Sensing ImageryabstractCloud detection is one of the important tasks for remote sensing image (RSI) preprocessing. In this paper, we utilize the thumbnail (i.e., preview image) of RSI, which contains the information of original multispectral or panchromatic imagery, to extract cloud mask efficiently. Compared with detection cloud mask from original RSI, it is more challenging to detect cloud mask using thumbnails due to the loss of resolution and spectrum information. To tackle this problem, we propose a cloud detection neural network (CDnet) with an encoder-decoder structure, a feature pyramid module (FPM), and a boundary refinement (BR) block. The FPM extracts the multiscale contextual information without the loss of resolution and coverage; the BR block refines object boundaries; and the encoder-decoder structure gradually recovers segmentation results with the same size as input image. Experimental results on the ZY-3 satellite thumbnails cloud cover validation data set and two other validation data sets (GF-1 WFV Cloud and Cloud Shadow Cover Validation Data and Landsat-8 Cloud Cover Assessment Validation Data) demonstrate that the proposed method achieves accurate detection accuracy and outperforms several state-of-the-art methods. Jing-Yu Yang 0002, Jianhua Guo 0002, Huanjing Yue, Haofeng Hu, Kun Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2019 | Global 3D Non-Rigid Registration of Deformable Objects Using a Single RGB-D CameraabstractWe present a novel global non-rigid registration method for dynamic 3D objects. Our method allows objects to undergo large non-rigid deformations and achieves high-quality results even with substantial pose change or camera motion between views. In addition, our method does not require a template prior and uses less raw data than tracking-based methods since only a sparse set of scans is needed. We simultaneously compute the deformations of all the scans by optimizing a global alignment problem to avoid the well-known loop closure problem and use an as-rigid-as-possible constraint to eliminate the shrinkage problem of the deformed shapes, especially near open boundaries of scans. To cope with large-scale problems, we design a coarse-to-fine multi-resolution scheme, which also avoids the optimization being trapped into local minima. The proposed method is evaluated on public datasets and real datasets captured by an RGB-D sensor. The experimental results demonstrate that the proposed method obtains better results than several state-of-the-art methods. Jing-Yu Yang 0002, Daoliang Guo, Kun Li 0001, Zhenchao Wu, Yukun Lai |
IEEE Trans. Image Process. | 3 |
| 2019 | Robust Non-Rigid Registration with Reweighted Position and Transformation SparsityabstractNon-rigid registration is challenging because it is ill-posed with high degrees of freedom and is thus sensitive to noise and outliers. We propose a robust non-rigid registration method using reweighted sparsities on position and transformation to estimate the deformations between 3-D shapes. We formulate the energy function with position and transformation sparsity on both the data term and the smoothness term, and define the smoothness constraint using local rigidity. The double sparsity based non-rigid registration model is enhanced with a reweighting scheme, and solved by transferring the model into four alternately-optimized subproblems which have exact solutions and guaranteed convergence. Experimental results on both public datasets and real scanned datasets show that our method outperforms the state-of-the-art methods and is more robust to noise and outliers than conventional non-rigid registration methods. Kun Li 0001, Jing-Yu Yang 0002, Yukun Lai, Daoliang Guo |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2018 | Image Alignment via Multi-Model Geometric Fitting and Hierarchical Homography EstimationabstractIt is challenging to achieve accurate alignment for building images containing multiple planes. We propose a multi-model geometric fitting and hierarchical homography estimation method to improve the alignment performance for building images. We first extract scale-invariant feature transform (SIFT) features of the images, and then adopt the multi-homography fitting algorithm to classify the feature points into different deformation models. According to the deduced deformation models, we partition the source image into base and transition regions. For the base regions, we adopt the moving direct linear transformation (Moving DLT) to estimate homographies. For the transition regions, we propose a hierarchical homography estimation method to select appropriate homographies. Experimental results show that our method achieves more accurate alignment results compared with state-of-the-art alignment methods for building images. Jing-Yu Yang 0002, Huanjing Yue, Kun Li 0001, Chunping Hou |
ICASSP | 4 |
| 2018 | Image-Based PM2.5 Estimation and its Application on Depth EstimationabstractAir pollution is still a big threat to human health particularly for developing countries. It is highly demanding to measure air quality with daily-used devices such as smartphones. On the other hand, it is difficult to estimate the scene depth under the foul weather using traditional vision-based methods. This paper proposes an image-based method for PM2.5 estimation by capturing a single image. We extract high-level features based on convolutional neural network (CNN) and learn the mapping between the features and PM2.5 by support vector regression (SVR). Given a captured image, we can estimate the PM2.5 value in real time. With the estimated PM2.5, we can estimate the depth of scene using sparse prior and nonlocal bilateral kernel. Experimental results demonstrate that the proposed method achieves the same accuracy of PM2.5 estimation as commodity measurement devices, and estimates the accurate depth information that is even better than the “ground-truth” captured by a laser in the no-haze condition. Kun Li 0001, Yahong Han, Pufeng Du, Jing-Yu Yang 0002 |
ICASSP | 2 |
| 2018 | Image-based Air Pollution Estimation Using Hybrid Convolutional Neural NetworkabstractAir pollution has a serious impact on our daily life, and how to quickly and easily measure the air pollution level without any expensive equipment is a quite challenging task. This paper proposes an air pollution estimation method using deep hybrid convolutional neural network from a single image, e.g., captured by a smartphone. The captured image is input to the main network, a very deep network, which solves the side effects of increased depth (degradation issues) by skip connection. This can improve network performance by simply increasing the depth of the network. Dark channel map is computed and fed into a secondary network to enrich the features with implicit representation. We have collected 1575 images of different scenes with different values of PM2.5to train the network in the end-to-end fusion mode. Experimental results on synthetic dataset and real captured dataset demonstrate that our method achieves excellent performance on classification of air pollution levels from a single captured image. Kun Li 0001, Yahong Han, Jing-Yu Yang 0002 |
ICPR | 2 |
| 2018 | Shape and Pose Estimation for Closely Interacting Persons Using Multi-view ImagesabstractAbstract Multi‐person pose and shape estimation is very challenging, especially when the persons have close interactions. Existing methods only work well when people are well spaced out in the captured images. However, close interaction among people is very common in real life, which is more challenge due to complex articulation, frequent occlusion and inherent ambiguities. We present a fully‐automatic markerless motion capture method to simultaneously estimate 3D poses and shapes of closely interacting people from multi‐view sequences. We first predict the 2D joints for each person in an image, and then design a spatio‐temporal tracker for multi‐person pose tracking based on multi‐view videos. Finally, we estimate 3D poses and shapes of all the persons with multi‐view constraints using a skinned multi‐person linear model (SMPL). Experimental results demonstrate that our method achieves fast but accurate pose and shape estimation results for multi‐person close interaction cases. Compared with existing methods, our method does not need pre‐segmentation for each person and manual intervention, which greatly reduces the complexity of the system including time complexity and system processing complexity. Kun Li 0001, Nianhong Jiao, Yebin Liu, Yangang Wang 0001, Jing-Yu Yang 0002 |
Comput. Graph. Forum | 1 |
| 2017 | KD-Tree and HEALPix-Based Distributed Cone Search Indexing System for Multi-Band Astronomical Catalogs
Ce Yu, Jian Xiao 0001, Xiaoteng Hu, Hao Fu 0021, Kun Li 0001, Yanyan Huang |
ICA3PP | 6 |
| 2017 | Estimation of signal-dependent noise level function using multi-column convolutional neural networkabstractTo estimate the levels of signal-dependent noise (SDN) from a single image is challenging. This paper proposes a novel method to estimate the noise level function (NLF) from a single image using a Multi-column Convolutional Neural Network (MC-Net) with an end-to-end architecture. The MC-Net is trained on a synthesized dataset containing noisy images with known NLFs, and it allows to learn rich hierarchical features using three sub-networks. Moreover, this method performs end-to-end training to retain more details for pixel-wise noise level estimation. Experimental results indicate that our method is accurate and robust to estimate NLFs of SDN for various types of images. Jing-Yu Yang 0002, Xin Liu 0012, Kun Li 0001 |
ICIP | 4 |
| 2017 | Low-rank matrix completion against missing rows and columns with separable 2-D sparsity priorsabstractMost existing matrix completion approaches assume that entries of matrices are missing at random, which could be violated in practical applications. This paper proposes a novel matrix completion method equipped with Joint Priors of LOw-rank and Separable 2-D Sparsity (JPLOSS) to complete missing rows and columns besides random missing. The underlying matrix is regularized by a low-rank prior, and its rows and columns are regularized by a row and a column dictionary, respectively. An reweighting scheme is incorporated into both the low-rank and sparsity terms to promote the low-rankness and sparseness simultaneously. The proposed model is effectively solved by an alternating direction method under the augmented Lagrangian multiplier framework. Experiments on both synthetic data and real images demonstrate the effectiveness and superiority of the proposed model in completing matrices with missing rows and columns compared with state-of-the-art matrix completion approaches. Jiaoru Yang, Kun Li 0001, Jing-Yu Yang 0002 |
ICIP | 3 |
| 2017 | Global alignment of deformable objects captured by a single RGB-D cameraabstractWe present a novel global registration method for deformable objects captured using a single RGB-D camera. Our algorithm allows objects to undergo large non-rigid deformations, and achieves high quality results without constraining the actor's pose or camera motion. We compute the deformations of all the scans simultaneously by optimizing a global alignment problem to avoid the well-known loop closure problem, and use an as-rigid-as-possible constraint to eliminate the shrinkage problem of the deformed model. To attack large scale problems, we design a coarse-to-fine multi-resolution scheme, which also avoids the optimization being trapped into local minima. The proposed method is evaluated on public datasets and real datasets captured by an RGB-D sensor. Experimental results demonstrate that the proposed method obtains better results than the state-of-the-art methods. Daoliang Guo, Kun Li 0001, Yukun Lai, Jing-Yu Yang 0002 |
ICME | 2 |
| 2017 | 3-D motion recovery via low rank matrix restoration on articulation graphsabstractThis paper addresses the challenge of 3-D skeleton recovery by exploiting the spatio-temporal correlations of corrupted 3D skeleton sequences. A skeleton sequence is represented as a matrix. We propose a novel low-rank solution that effectively integrates both a low-rank model for robust skeleton recovery based on temporal coherence, and an articulation-graph-based isometric constraint for spatial coherence, namely consistency of bone lengths. The proposed model is formulated as a constrained optimization problem, which is efficiently solved by the Augmented Lagrangian Method with a Gauss-Newton solver for the subproblem of isometric optimization. Experimental results on the CMU motion capture dataset and a Kinect dataset show that the proposed approach achieves better recovery accuracy over a state-of-the-art method. The proposed method has wide applicability for skeleton tracking devices, such as the Kinect, because these devices cannot provide accurate reconstructions of complex motions, especially in the presence of occlusion. Kun Li 0001, Meiyuan Wang, Yukun Lai, Jing-Yu Yang 0002, Feng Wu 0001 |
ICME | 1 |
| 2017 | Intrinsic decomposition from a single RGB-D image with sparse and non-local priorsabstractThis paper proposes a new intrinsic image decomposition method that decomposes a single RGB-D image into reflectance and shading components. We observe and verify that, a shading image mainly contains smooth regions separated by curves, and its gradient distribution is sparse. We therefore use ℓ1-norm to model the direct irradiance component - the main sub-component extracted from shading component. Moreover, a non-local prior weighted by a bilateral kernel on a larger neighborhood is designed to fully exploit structural correlation in the reflectance component to improve the decomposition performance. The model is solved by the alternating direction method under the augmented Lagrangian multiplier (ADM-ALM) framework. Experimental results on both synthetic and real datasets demonstrate that the proposed method yields better results and enjoys lower complexity compared with two state-of-the-art methods. Kun Li 0001, Jing-Yu Yang 0002, Xinchen Ye |
ICME | 2 |
| 2017 | Depth super-resolution via fully edge-augmented guidanceabstractRecently, convolutional neural networks (CNNs) have been widely used for image processing problems. In this work, we present an end-to-end depth map super-resolution method based on CNN. Standing on a residual learning architecture, the proposed network learns joint features to get a high-resolution (HR) depth map from a low-resolution (LR) one with the multi-layers guidance of a HR color image. Furthermore, in order to focus on the boundaries of depth map, we generate an edge-attention map from the associated HR color images as a guidance. Experimental results show that the proposed network outperforms the state-of-the-art depth map super-resolution methods. Jing-Yu Yang 0002, Kun Li 0001 |
VCIP | 4 |
| 2017 | Demoiréing for screen-shot images with multi-channel layer decompositionabstractMoiré patterns on screen-shot images are mainly due to the aliasing of the grid of the display and the camera sensor, which heavily degenerated the image quality. This paper proposes an demoiréing method for screen-shot images via layer decomposition on polyphase components (LDPC). The layer decomposition model separates the image into a background layer and a moiré layer, which are both regularized by a patch-based Gaussian Mixture Model (GMM) prior. To enhance the distinguishability between the image patches and moiré patches, the input image is first subsampled into four polyphase components, each of which is decomposed with the GMM-based layer decomposition model. The proposed model is applied on luminance (Y) channel to weaken the intensity of moiré patterns, and on red, green, blue channels respectively to further remove moiré patterns. Experimental results demonstrate that the proposed method is able to efficiently remove moiré artifacts for screen-shot images and outperform several other methods. Jing-Yu Yang 0002, Changrui Cai, Kun Li 0001 |
VCIP | 4 |
| 2017 | SPA: Sparse Photorealistic Animation Using a Single RGB-D CameraabstractPhotorealistic animation is a desirable technique for computer games and movie production. We propose a new method to synthesize plausible videos of human actors with new motions using a single cheap RGB-D camera. A small database is captured in a usual office environment, which happens only once for synthesizing different motions. We propose a marker-less performance capture method using sparse deformation to obtain the geometry and pose of the actor for each time instance in the database. Then, we synthesize an animation video of the actor performing the new motion that is defined by the user. An adaptive model-guided texture synthesis method based on weighted low-rank matrix completion is proposed to be less sensitive to noise and outliers, which enables us to easily create photorealistic animation videos with new motions that are different from the motions in the database. Experimental results on the public data set and our captured data set have verified the effectiveness of the proposed method. Kun Li 0001, Jing-Yu Yang 0002, Leijie Liu, Ronan Boulic, Yukun Lai, Yebin Liu, Eray Molla |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2016 | 3-D motion recovery via low rank matrix analysisabstractSkeleton tracking is a useful and popular application of Kinect. However, it cannot provide accurate reconstructions for complex motions, especially in the presence of occlusion. This paper proposes a new 3-D motion recovery method based on low-rank matrix analysis to correct invalid or corrupted motions. We address this problem by representing a motion sequence as a matrix, and introducing a convex low-rank matrix recovery model, which fixes erroneous entries and finds the correct low-rank matrix by minimizing nuclear norm and norm of constituent clean motion and error matrices. Experimental results show that our method recovers the corrupted skeleton joints, achieving accurate and smooth reconstructions even for complicated motions. Meiyuan Wang, Kun Li 0001, Feng Wu 0001, Yukun Lai, Jing-Yu Yang 0002 |
VCIP | 2 |
| 2016 | Video super-resolution using an adaptive superpixel-guided auto-regressive model
Kun Li 0001, Yanming Zhu 0001, Jing-Yu Yang 0002, Jianmin Jiang |
Pattern Recognit. | 1 |
| 2015 | Sparse Non-rigid Registration of 3D ShapesabstractAbstract Non‐rigid registration of 3D shapes is an essential task of increasing importance as commodity depth sensors become more widely available for scanning dynamic scenes. Non‐rigid registration is much more challenging than rigid registration as it estimates a set of local transformations instead of a single global transformation, and hence is prone to the overfitting issue due to underdetermination. The common wisdom in previous methods is to impose an ℓ2‐norm regularization on the local transformation differences. However, the ℓ2‐norm regularization tends to bias the solution towards outliers and noise with heavy‐tailed distribution, which is verified by the poor goodness‐of‐fit of the Gaussian distribution over transformation differences. On the contrary, Laplacian distribution fits well with the transformation differences, suggesting the use of a sparsity prior. We propose a sparse non‐rigid registration (SNR) method with an ℓ1‐norm regularized model for transformation estimation, which is effectively solved by an alternate direction method (ADM) under the augmented Lagrangian framework. We also devise a multi‐resolution scheme for robust and progressive registration. Results on both public datasets and our scanned datasets show the superiority of our method, particularly in handling large‐scale deformations as well as outliers and noise. Jing-Yu Yang 0002, Kun Li 0001, Yukun Lai |
Comput. Graph. Forum | 3 |
| 2015 | Foreground-Background Separation From Video Clips via Motion-Assisted Matrix RestorationabstractSeparation of video clips into foreground and background components is a useful and important technique, making recognition, classification, and scene analysis more efficient. In this paper, we propose a motion-assisted matrix restoration (MAMR) model for foreground-background separation in video clips. In the proposed MAMR model, the backgrounds across frames are modeled by a low-rank matrix, while the foreground objects are modeled by a sparse matrix. To facilitate efficient foreground-background separation, a dense motion field is estimated for each frame, and mapped into a weighting matrix which indicates the likelihood that each pixel belongs to the background. Anchor frames are selected in the dense motion estimation to overcome the difficulty of detecting slowly moving objects and camouflages. In addition, we extend our model to a robust MAMR model against noise for practical applications. Evaluations on challenging datasets demonstrate that our method outperforms many other state-of-the-art methods, and is versatile for a wide range of surveillance videos. Xinchen Ye, Jing-Yu Yang 0002, Kun Li 0001, Chunping Hou, Yao Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2015 | Nonrigid Structure From Motion via Sparse RepresentationabstractThis paper proposes a new approach for nonrigid structure from motion with occlusion, based on sparse representation. We address the occlusion problem based on the latest developments on sparse representation: matrix completion, which can recover the observation matrix that has high percentages of missing data and can also reduce the noises and outliers in the known elements. We introduce sparse transform to the joint estimation of 3-D shapes and motions. 3-D shape trajectory space is fit by wavelet basis to achieve better modeling of complex motion. Experimental results on datasets without and with occlusion show that our method can better estimate the 3-D shapes and motions, compared with state-of-the-art algorithms. Kun Li 0001, Jing-Yu Yang 0002, Jianmin Jiang |
IEEE Trans. Cybern. | 1 |
| 2015 | Graph-Based Segmentation for RGB-D Data Using 3-D Geometry Enhanced SuperpixelsabstractWith the advances of depth sensing technologies, color image plus depth information (referred to as RGB-D data hereafter) is more and more popular for comprehensive description of 3-D scenes. This paper proposes a two-stage segmentation method for RGB-D data: 1) oversegmentation by 3-D geometry enhanced superpixels and 2) graph-based merging with label cost from superpixels. In the oversegmentation stage, 3-D geometrical information is reconstructed from the depth map. Then, a K-means-like clustering method is applied to the RGB-D data for oversegmentation using an 8-D distance metric constructed from both color and 3-D geometrical information. In the merging stage, treating each superpixel as a node, a graph-based model is set up to relabel the superpixels into semantically-coherent segments. In the graph-based model, RGB-D proximity, texture similarity, and boundary continuity are incorporated into the smoothness term to exploit the correlations of neighboring superpixels. To obtain a compact labeling, the label term is designed to penalize labels linking to similar superpixels that likely belong to the same object. Both the proposed 3-D geometry enhanced superpixel clustering method and the graph-based merging method from superpixels are evaluated by qualitative and quantitative results. By the fusion of color and depth information, the proposed method achieves superior segmentation performance over several state-of-the-art algorithms. Jing-Yu Yang 0002, Ziqiao Gan, Kun Li 0001, Chunping Hou |
IEEE Trans. Cybern. | 3 |
| 2014 | Background extraction from video sequences via motion-assisted matrix completionabstractBackground extraction from video sequences is a useful and important technique in video surveillance. This paper proposes a motion-assisted matrix completion model for background extraction from video sequences. A binary motion map is first calculated for each frame by optical flow. By excluding areas associated with moving objects with the binary motion maps, the background extraction is formulated into a motion-assisted matrix completion (MAMC) problem. Experimental results show that our method not only extracts promising backgrounds but also outperforms many state-of-the-art methods in distinguishing moving objects on challenging datasets. Jing-Yu Yang 0002, Xinchen Ye, Kun Li 0001 |
ICIP | 4 |
| 2014 | Non-rigid structure from motion via sparse representationabstractThis paper proposes a new approach for non-rigid structure from motion with occlusion, based on sparse representation. We introduce sparse transform to the joint estimation of 3D shapes and motions. 3D shape trajectory space is fit by wavelet basis to achieve better modeling of complex motion. We address the occlusion problem based on the latest developments on sparse representation: matrix completion, which can recover the observation matrix that has high percentages of missing data and can also reduce the noises and outliers in the known elements. Experimental results on datasets without and with occlusion show that our method can better estimate the 3D shapes and motions, compared with state-of-the-art algorithms. Kun Li 0001, Jing-Yu Yang 0002, Jianmin Jiang |
ICME | 1 |
| 2014 | Video super-resolution based on automatic key-frame selection and feature-guided variational optical flow
Yanming Zhu 0001, Kun Li 0001, Jianmin Jiang |
Signal Process. Image Commun. | 2 |
| 2014 | Color-Guided Depth Recovery From RGB-D Data Using an Adaptive Autoregressive ModelabstractThis paper proposes an adaptive color-guided autoregressive (AR) model for high quality depth recovery from low quality measurements captured by depth cameras. We observe and verify that the AR model tightly fits depth maps of generic scenes. The depth recovery task is formulated into a minimization of AR prediction errors subject to measurement consistency. The AR predictor for each pixel is constructed according to both the local correlation in the initial depth map and the nonlocal similarity in the accompanied high quality color image. We analyze the stability of our method from a linear system point of view, and design a parameter adaptation scheme to achieve stable and accurate depth recovery. Quantitative and qualitative evaluation compared with ten state-of-the-art schemes show the effectiveness and superiority of our method. Being able to handle various types of depth degradations, the proposed method is versatile for mainstream depth sensors, time-of-flight camera, and Kinect, as demonstrated by experiments on real systems. Jing-Yu Yang 0002, Xinchen Ye, Kun Li 0001, Chunping Hou, Yao Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2012 | Depth Recovery Using an Adaptive Color-Guided Auto-Regressive Model
Jing-Yu Yang 0002, Xinchen Ye, Kun Li 0001, Chunping Hou |
ECCV (5) | 3 |
| 2012 | Optimized image super-resolution based on sparse representation
Yanming Zhu 0001, Jianmin Jiang, Kun Li 0001 |
ICPR | 3 |
| 2012 | Three-Dimensional Motion Estimation via Matrix CompletionabstractThree-dimensional motion estimation from multiview video sequences is of vital importance to achieve high-quality dynamic scene reconstruction. In this paper, we propose a new 3-D motion estimation method based on matrix completion. Taking a reconstructed 3-D mesh as the underlying scene representation, this method automatically estimates motions of 3-D objects. A "separating + merging" framework is introduced to multiview 3-D motion estimation. In the separating step, initial motions are first estimated for each view with a neighboring view. Then, in the merging step, the motions obtained by each view are merged together and optimized by low-rank matrix completion method. The most accurate motion estimation for each vertex in the recovered matrix is further selected by three spatiotemporal criteria. Experimental results on data sets with synthetic motions and real motions show that our method can reliably estimate 3-D motions. Kun Li 0001, Qionghai Dai, Wenli Xu, Jing-Yu Yang 0002, Jianmin Jiang |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2011 | Collaborative color calibration for multi-camera systems
Kun Li 0001, Qionghai Dai, Wenli Xu |
Signal Process. Image Commun. | 1 |
| 2011 | Markerless Shape and Motion Capture From Multiview Video SequencesabstractWe propose a new markerless shape and motion capture approach from multiview video sequences. The shape recovery method consists of two steps: separating and merging. In the separating step, the depth map represented with a point cloud for each view is generated by solving a proposed variational model, which is regularized by four constraints to ensure the accuracy and completeness of the reconstruction. Then, in the merging step, the point clouds of all the views are merged together and reconstructed into a 3-D mesh using a marching cubes method with silhouette constraints. Experiments show that the geometric details are faithfully preserved in each estimated depth map. The 3-D meshes reconstructed from the estimated depth maps are watertight and present rich geometric details, even for non-convex objects. Taking the reconstructed 3-D mesh as the underlying scene representation, a volumetric deformation method with a new positional-constraint computation scheme is proposed to automatically capture motions of the 3-D object. Our method can capture non-rigid motions even for loosely dressed humans without the aid of markers. Kun Li 0001, Qionghai Dai, Wenli Xu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2010 | High quality color calibration for multi-camera systems with an omnidirectional color checkerabstractThis paper proposes a new color calibration method for multi-camera systems with a novel omnidirectional color checker. The designed cylindrical color checker contains a periodic array of color patches, and is visible for all the cameras without manual adjustment. For color calibration, accurate global correspondences are first generated by local descriptors and area-based correlation methods. Then, the multi-camera color calibration problem is formulated as an overdetermined linear system, in which the dynamic range shaping is incorporated to ensure the high contrasts of captured images. The cameras are calibrated with the parameters obtained by solving the linear system. According to experimental results on a real multi-camera system, the proposed method shows high performance in achieving inter-camera color consistency and high dynamic range. Kun Li 0001, Qionghai Dai, Wenli Xu |
ICASSP | 1 |
| 2007 | A comparative study of image compression based on directional waveletsabstractDiscrete wavelet transform is an effective tool to generate scalable stream, but it cannot efficiently represent edges which are not aligned in horizontal or vertical directions, while natural images often contain rich edges and textures of this kind. Hence, recently, intensive research has been focused particularly on the directional wavelets which can effectively represent directional attributes of images. Specifically, there are two categories of directional wavelets: redundant wavelets (RW) and adaptive directional wavelets (ADW). One representative redundant wavelet is the dual-tree discrete wavelet transform (DDWT), while adaptive directional wavelets can be further categorized into two types: with or without side information. In this paper, we briefly introduce directional wavelets and compare their directional bases and image compression performances. Kun Li 0001, Wenli Xu, Qionghai Dai, Yao Wang 0001 |
VCIP | 1 |