Ke Xian

dblp:178/1416 · DBLP profile ↗
← Back
48ranked-venue papers
4as first author
30since 2021 · last 2025
0000-0002-0884-5126ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 40 · 3 first-author · 24 since 2021Artificial intelligence and machine learning · 26 · 3 first-author · 17 since 2021
YearPublicationVenuePosition
2025 PoseMamba: Monocular 3D Human Pose Estimation with Bidirectional Global-Local Spatio-Temporal State Space Model
abstract
Transformers have significantly advanced the field of 3D human pose estimation (HPE). However, existing transformer-based methods primarily use self-attention mechanisms for spatio-temporal modeling, leading to a quadratic complexity, unidirectional modeling of spatio-temporal relationships, and insufficient learning of spatial-temporal correlations. Recently, the Mamba architecture, utilizing the state space model (SSM), has exhibited superior long-range modeling capabilities in a variety of vision tasks with linear complexity. In this paper, we propose PoseMamba, a novel purely SSM-based approach with linear complexity for 3D human pose estimation in monocular video. Specifically, we propose a bidirectional global-local spatio-temporal SSM block that comprehensively models human joint relations within individual frames as well as temporal correlations across frames. Within this bidirectional global-local spatio-temporal SSM block, we introduce a reordering strategy to enhance the local modeling capability of the SSM. This strategy provides a more logical geometric scanning order and integrates it with the global SSM, resulting in a combined global-local spatial scan. We have quantitatively and qualitatively evaluated our approach using two benchmark datasets: Human3.6M and MPI-INF-3DHP. Extensive experiments demonstrate that PoseMamba achieves state-of-the-art performance on both datasets while maintaining a smaller model size and reducing computational costs.
Yunlong Huang, Junshuo Liu, Ke Xian, Robert C. Qiu
AAAI3
2025 BokehMe++: Harmonious Fusion of Classical and Neural Rendering for Versatile Bokeh Creation
abstract
Despite significant advancements in simulating the bokeh effect of Digital Single Lens Reflex Camera (DSLR) from an all-in-focus image, challenges remain in processing highlight points, preserving boundary details for in-focus objects and processing high-resolution images efficiently. To tackle these issues, we first develop a ray-tracing-based bokeh simulator. An innovative pipeline with weight redistribution is introduced to handle highlight rendering. By considering the front length of lens barrel, we can simulate realistic cat-eye effect. This bokeh simulator serves as the foundation for creating our training dataset. Building on this dataset, we introduce a hybrid framework BokehMe++, combining a classical renderer and a neural renderer. The classical renderer is implemented by a hierarchical scattering-based method, which suffers from boundary inaccuracies. These erroneous areas will be identified by an error map generator and be corrected by a two-stage neural renderer. Adaptive resizing and iterative upsampling are introduced in the neural renderer to process arbitrary blur size efficiently. Extensive experiments demonstrate that BokehMe++ outperforms existing methods and provides highly customizable rendering features, such as adjustable blur amount, focal plane, highlight mode and cat-eye effect. Furthermore, BokehMe++ can maintain the sharpness of hair details in portraits through an auxiliary alpha map input.
Juewen Peng, Zhiguo Cao 0001, Xianrui Luo, Ke Xian, Wenfeng Tang, Jianming Zhang 0001, Guosheng Lin
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 NVDS$^{\mathbf{+}}$+: Towards Efficient and Versatile Neural Stabilizer for Video Depth Estimation
abstract
Video depth estimation aims to infer temporally consistent depth. One approach is to finetune a single-image model on each video with geometry constraints, which proves inefficient and lacks robustness. An alternative is learning to enforce consistency from data, which requires well-designed models and sufficient video depth data. To address both challenges, we introduce NVDS that stabilizes inconsistent depth estimated by various single-image models in a plug-and-play manner. We also elaborate a large-scale Video Depth in the Wild (VDW) dataset, which contains 14,203 videos with over two million frames, making it the largest natural-scene video depth dataset. Additionally, a bidirectional inference strategy is designed to improve consistency by adaptively fusing forward and backward predictions. We instantiate a model family ranging from small to large scales for different applications. The method is evaluated on VDW dataset and three public benchmarks. To further prove the versatility, we extend NVDS to video semantic segmentation and several downstream applications like bokeh rendering, novel view synthesis, and 3D reconstruction. Experimental results show that our method achieves significant improvements in consistency, accuracy, and efficiency. Our work serves as a solid baseline and data foundation for learning-based video depth estimation.
Yiran Wang 0005, Min Shi 0004, Jiaqi Li 0007, Chaoyi Hong, Zihao Huang 0001, Juewen Peng, Zhiguo Cao 0001, Jianming Zhang 0001, Ke Xian, Guosheng Lin
IEEE Trans. Pattern Anal. Mach. Intell.9
2025 Dynamic View Synthesis From Small Camera Motion Videos
abstract
Novel view synthesis for dynamic 3D scenes poses a significant challenge. Many notable efforts use NeRF-based approaches to address this task and yield impressive results. However, these methods rely heavily on sufficient motion parallax in the input images or videos. When the camera motion range becomes limited or even stationary (i.e., small camera motion), existing methods encounter two primary challenges: incorrect representation of scene geometry and inaccurate estimation of camera parameters. These challenges make prior methods struggle to produce satisfactory results or even become ineffective. To address the first challenge, we propose a novel Distribution-based Depth Regularization (DDR) that ensures the rendering weight distribution to align with the true distribution. Specifically, unlike previous methods that use depth loss to calculate the error of the expectation, we calculate the expectation of the error by using Gumbel-softmax to differentiably sample points from discrete rendering weight distribution. Additionally, we introduce constraints that enforce the volume density of spatial points before the object boundary along the ray to be near zero, ensuring that our model learns the correct geometry of the scene. To demystify the DDR, we further propose a visualization tool that enables observing the scene geometry representation at the rendering weight level. For the second challenge, we incorporate camera parameter learning during training to enhance the robustness of our model to camera parameters. We conduct extensive experiments to demonstrate the effectiveness of our approach in representing scenes with small camera motion input, and our results compare favorably to state-of-the-art methods.
Huiqiang Sun, Xingyi Li 0005, Juewen Peng, Liao Shen, Zhiguo Cao 0001, Ke Xian, Guosheng Lin
IEEE Trans. Vis. Comput. Graph.6
2024 Semi-supervised Class-Agnostic Motion Prediction with Pseudo Label Regeneration and BEVMix
abstract
Class-agnostic motion prediction methods aim to comprehend motion within open-world scenarios, holding significance for autonomous driving systems. However, training a high-performance model in a fully-supervised manner always requires substantial amounts of manually annotated data, which can be both expensive and time-consuming to obtain. To address this challenge, our study explores the potential of semi-supervised learning (SSL) for class-agnostic motion prediction. Our SSL framework adopts a consistency-based self-training paradigm, enabling the model to learn from unlabeled data by generating pseudo labels through test-time inference. To improve the quality of pseudo labels, we propose a novel motion selection and re-generation module. This module effectively selects reliable pseudo labels and re-generates unreliable ones. Furthermore, we propose two data augmentation strategies: temporal sampling and BEVMix. These strategies facilitate consistency regularization in SSL. Experiments conducted on nuScenes demonstrate that our SSL method can surpass the self-supervised approach by a large margin by utilizing only a tiny fraction of labeled data. Furthermore, our method exhibits comparable performance to weakly and some fully supervised methods. These results highlight the ability of our method to strike a favorable balance between annotation costs and performance. Code will be available at https://github.com/kwwcv/SSMP.
Kewei Wang 0001, Yizheng Wu, Xingyi Li 0005, Ke Xian, Zhe Wang 0006, Zhiguo Cao 0001, Guosheng Lin
AAAI5
2024 S-DyRF: Reference-Based Stylized Radiance Fields for Dynamic Scenes
abstract
Current 3D stylization methods often assume static scenes, which violates the dynamic nature of our real world. To address this limitation, we present S-DyRF, a reference-based spatio-temporal stylization method for dynamic neu-ral radiance fields. However, stylizing dynamic 3D scenes is inherently challenging due to the limited availability of stylized reference images along the temporal axis. Our key insight lies in introducing additional temporal cues besides the provided reference. To this end, we generate temporal pseudo-references from the given stylized reference. These pseudo-references facilitate the propagation of style infor-mation from the reference to the entire dynamic 3D scene. For coarse style transfer, we enforce novel views and times to mimic the style details present in pseudo-references at the feature level. To preserve high-frequency details, we create a collection of stylized temporal pseudo-rays from temporal pseudo-references. These pseudo-rays serve as detailed and explicit stylization guidance for achieving fine style trans-fer. Experiments on both synthetic and real-world datasets demonstrate that our method yields plausible stylized re-sults of space-time view synthesis on dynamic 3D scenes.
Xingyi Li 0005, Zhiguo Cao 0001, Yizheng Wu, Kewei Wang 0001, Ke Xian, Zhe Wang 0006, Guosheng Lin
CVPR5
2024 Dr.Bokeh: DiffeRentiable Occlusion-Aware Bokeh Rendering
abstract
Bokeh is widely used in photography to draw attention to the subject while effectively isolating distractions in the background. Computational methods can simulate bokeh effects without relying on a physical camera lens, but the inaccurate lens modeling in existing filtering-based meth-ods leads to artifacts that need post-processing or learning-based methods to fix. We propose Dr.Bokeh, a novel ren-dering method that addresses the issue by directly correcting the defect that violates physics in the current filtering-based bokeh rendering equation. Dr.Bokeh first preprocesses the input RGBD to obtain a layered scene representation. Dr.Bokeh then takes the layered representation and user-defined lens parameters to render photo-realistic lens blur based on the novel occlusion-aware bokeh rendering method. Experiments show that the non-learning based renderer Dr.Bokeh outperforms state-of-the-art bokeh ren-dering algorithms in terms of photo-realism. In addition, extensive quantitative and qualitative evaluations show that the more accurate lens model pushes the limit of depth-from-defocus.
Yichen Sheng, Zixun Yu, Lu Ling, Zhiwen Cao, Xuaner Cecilia Zhang, Xin Lu 0006, Ke Xian, Haiting Lin, Bedrich Benes
CVPR7
2024 DyBluRF: Dynamic Neural Radiance Fields from Blurry Monocular Video
abstract
Recent advancements in dynamic neural radiance field methods have yielded remarkable outcomes. However, these approaches rely on the assumption of sharp input images. When faced with motion blur, existing dynamic NeRF methods often struggle to generate high-quality novel views. In this paper, we propose DyBluRF, a dynamic radiance field approach that synthesizes sharp novel views from a monocular video affected by motion blur. To account for motion blur in input images, we simultaneously capture the camera trajectory and object Discrete Cosine Transform (DCT) trajectories within the scene. Additionally, we employ a global cross-time rendering approach to ensure consistent temporal coherence across the entire scene. We curate a dataset comprising diverse dynamic scenes that are specifically tailored for our task. Experimental results on our dataset demonstrate that our method outperforms existing approaches in generating sharp novel views from motion-blurred inputs while maintaining spatial-temporal consistency of the scene.
Huiqiang Sun, Xingyi Li 0005, Liao Shen, Ke Xian, Zhiguo Cao 0001
CVPR5
2024 iControl3D: An Interactive System for Controllable 3D Scene Generation
abstract
3D content creation has long been a complex and time-consuming process, often requiring specialized skills and resources. While re- cent advancements have allowed for text-guided 3D object and scene generation, they still fall short of providing sufficient control over the generation process, leading to a gap between the user’s creative vision and the generated results. In this paper, we present iControl3D, a novel interactive system that empowers users to gen- erate and render customizable 3D scenes with precise control. To this end, a 3D creator interface has been developed to provide users with fine-grained control over the creation process. Technically, we leverage 3D meshes as an intermediary proxy to iteratively merge individual 2D diffusion-generated images into a cohesive and uni- fied 3D scene representation. To ensure seamless integration of 3D meshes, we propose to perform boundary-aware depth alignment before fusing the newly generated mesh with the existing one in 3D space. Additionally, to effectively manage depth discrepancies between remote content and foreground, we propose to model re- mote content separately with an environment map instead of 3D meshes. Finally, our neural rendering interface enables users to build a radiance field of their scene online and navigate the entire scene. Extensive experiments have been conducted to demonstrate the effectiveness of our system. The code will be made available at https://github.com/xingyi- li/iControl3D.
Xingyi Li 0005, Yizheng Wu, Jun Cen, Juewen Peng, Kewei Wang 0001, Ke Xian, Zhe Wang 0006, Zhiguo Cao 0001, Guosheng Lin
ACM Multimedia6
2024 Self-Distilled Depth Refinement with Noisy Poisson Fusion
abstract
Depth refinement aims to infer high-resolution depth with fine-grained edges and details, refining low-resolution results of depth estimation models. The prevailing methods adopt tile-based manners by merging numerous patches, which lacks efficiency and produces inconsistency. Besides, prior arts suffer from fuzzy depth boundaries and limited generalizability. Analyzing the fundamental reasons for these limitations, we model depth refinement as a noisy Poisson fusion problem with local inconsistency and edge deformation noises. We propose the Self-distilled Depth Refinement (SDDR) framework to enforce robustness against the noises, which mainly consists of depth edge representation and edge-based guidance. With noisy depth predictions as input, SDDR generates low-noise depth edge representations as pseudo-labels by coarse-to-fine self-distillation. Edge-based guidance with edge-guided gradient loss and edge-based fusion loss serves as the optimization objective equivalent to Poisson fusion. When depth maps are better refined, the labels also become more noise-free. Our model can acquire strong robustness to the noises, achieving significant improvements in accuracy, edge quality, efficiency, and generalizability on five different benchmarks. Moreover, directly training another model with edge labels produced by SDDR brings improvements, suggesting that our method could help with training robust refinement models in future works.
Jiaqi Li 0007, Yiran Wang 0005, Jinghong Zheng 0002, Zihao Huang 0001, Ke Xian, Zhiguo Cao 0001, Jianming Zhang 0001
NeurIPS5
2024 Towards Robust Monocular Depth Estimation: A New Baseline and Benchmark
Ke Xian, Zhiguo Cao 0001, Chunhua Shen, Guosheng Lin
Int. J. Comput. Vis.1
2024 Hierarchical Feature Warping and Blending for Talking Head Animation
abstract
Talking head animation transforms a source anime image to a target pose, where the transformation includes the change of facial expression and head movement. In contrast to existing approaches that operate on the low-resolution image (256 × 256), we study this task at a higher resolution,e.g., 512 × 512. High-resolution talking head animation, however, raises two major challenges: i) how to achieve smooth global transformation while maintaining rich details of anime characters under large-displacement pose variations; ii) how to address the shortage of data, because no related dataset is publicly available. In this paper, we present a Hierarchical Feature Warping and Blending (HFWB) model, which tackles talking head animation hierarchically. Specifically, we use low-level features to control global transformation and high-level features to determine the details of anime characters, under the guidance of feature flow fields. These features are then blended by selective fusion units, outputting transformed anime images. In addition, we construct an anime pose dataset–AniTalk-2K, aiming to alleviate the shortage of data. It contains around 2000 anime characters with thousands of different face/head poses at a resolution of 512 × 512. Extensive experiments on AniTalk-2K demonstrate the superiority of our approach in generating high-quality anime talking heads over state-of-the-art methods.
Ke Xian, Zhiguo Cao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 ViTA: Video Transformer Adaptor for Robust Video Depth Estimation
abstract
Depth information plays a pivotal role in numerous computer vision applications, including autonomous driving, 3D reconstruction, and 3D content generation. When deploying depth estimation models in practical applications, it is essential to ensure that the models have strong generalization capabilities. However, existing depth estimation methods primarily concentrate on robust single-image depth estimation, leading to the occurrence of flickering artifacts when applied to video inputs. On the other hand, video depth estimation methods either consume excessive computational resources or lack robustness. To address the above issues, we propose ViTA, a video transformer adaptor, to estimate temporally consistent video depth in the wild. In particular, we leverage a pre-trained image transformer (i.e., DPT) and introduce additional temporal embeddings in the transformer blocks. Such designs enable our ViTA to output reliable results given an unconstrained video. Besides, we present a spatio-temporal consistency loss for supervision. The spatial loss computes the per-pixel discrepancy between the prediction and the ground truth in space, while the temporal loss regularizes the inconsistent outputs of the same point in consecutive frames. To find the correspondences between consecutive frames, we design a bi-directional warping strategy based on the forward and backward optical flow. During inference, our ViTA no longer requires optical flow estimation, which enables it to estimate spatially accurate and temporally consistent video depth maps with fine-grained details in real time. We conduct a detailed ablation study to verify the effectiveness of the proposed components. Extensive experiments on the zero-shot cross-dataset evaluation demonstrate that the proposed method is superior to previous methods.
Ke Xian, Juewen Peng, Zhiguo Cao 0001, Jianming Zhang 0001, Guosheng Lin
IEEE Trans. Multim.1
2023 3D Cinemagraphy from a Single Image
abstract
We present 3D Cinemagraphy, a new technique that mar-ries 2D image animation with 3D photography. Given a single still image as input, our goal is to generate a video that contains both visual content animation and camera motion. We empirically find that naively combining existing 2D image animation and 3D photography methods leads to obvious artifacts or inconsistent animation. Our key insight is that representing and animating the scene in 3D space offers a natural solution to this task. To this end, we first convert the input image into feature-based layered depth images using predicted depth values, followed by unprojecting them to a feature point cloud. To animate the scene, we perform motion estimation and lift the 2D motion into the 3D scene flow. Finally, to resolve the problem of hole emer-gence as points move forward, we propose to bidirectionally displace the point cloud as per the scene flow and synthe-size novel views by separately projecting them into target image planes and blending the results. Extensive experiments demonstrate the effectiveness of our method. A user study is also conducted to validate the compelling rendering results of our method.
Xingyi Li 0005, Zhiguo Cao 0001, Huiqiang Sun, Jianming Zhang 0001, Ke Xian, Guosheng Lin
CVPR5
2023 Neural Video Depth Stabilizer
abstract
Video depth estimation aims to infer temporally consistent depth. Some methods achieve temporal consistency by finetuning a single-image depth model during test time using geometry and re-projection constraints, which is inefficient and not robust. An alternative approach is to learn how to enforce temporal consistency from data, but this requires well-designed models and sufficient video depth data. To address these challenges, we propose a plug-and-play framework called Neural Video Depth Stabilizer (NVDS) that stabilizes inconsistent depth estimations and can be applied to different single-image depth models without extra effort. We also introduce a large-scale dataset, Video Depth in the Wild (VDW), which consists of 14,203 videos with over two million frames, making it the largest natural-scene video depth dataset to our knowledge. We evaluate our method on the VDW dataset as well as two public benchmarks and demonstrate significant improvements in consistency, accuracy, and efficiency compared to previous approaches. Our work serves as a solid baseline and provides a data foundation for learning-based video depth models. We will release our dataset and code for future research.
Yiran Wang 0005, Min Shi 0004, Jiaqi Li 0007, Zihao Huang 0001, Zhiguo Cao 0001, Jianming Zhang 0001, Ke Xian, Guosheng Lin
ICCV7
2023 SimHMR: A Simple Query-based Framework for Parameterized Human Mesh Reconstruction
abstract
Human Mesh Reconstruction (HMR) aims to recover 3D human poses and shapes from a single image. Existing parameterized HMR approaches follow the "representation-to-reasoning'' paradigm to predict human body and pose parameters. This paradigm typically involves intermediate representation and complex pipeline, where potential side effects may occur that could hinder performance. In contrast, query-based non-parameterized methods directly output 3D joints and mesh vertices, but they rely on excessive queries for prediction, leading to low efficiency and robustness. In this work, we propose a simple query-based framework, dubbed SimHMR, for parameterized human mesh reconstruction. This framework streamlines the prediction process by using a few parameterized queries, which effectively removes the need for hand-crafted intermediate representation and reasoning pipeline. Different from query-based non-parameterized HMR that uses excessive coordinate queries, SimHMR only requires a few semantic queries, which physically correspond to pose, shape, and camera. The use of semantic queries significantly improves the efficiency and robustness in extreme scenarios, e.g., occlusions. Without bells and whistles, øurs achieves state-of-the-art performance on 3DPW and Human3.6M benchmarks, and surpasses existing methods on challenging 3DPW-OCC. Code available at https://github.com/inso-13/SimHMR github.com/inso-13/SimHMR
Zihao Huang 0001, Min Shi 0004, Ke Xian, Zhiguo Cao 0001
ACM Multimedia4
2023 Diffusion-Augmented Depth Prediction with Sparse Annotations
abstract
Depth estimation aims to predict dense depth maps. In autonomous driving scenes, sparsity of annotations makes the task challenging. Supervised models produce concave objects due to insufficient structural information. They overfit to valid pixels and fail to restore spatial structures. Self-supervised methods are proposed for the problem. Their robustness is limited by pose estimation, leading to erroneous results in natural scenes. In this paper, we propose a supervised framework termed Diffusion-Augmented Depth Prediction (DADP). We leverage the structural characteristics of diffusion model to enforce depth structures of depth models in a plug-and-play manner. An object-guided integrality loss is also proposed to further enhance regional structure integrality by fetching objective information. We evaluate DADP on three driving benchmarks and achieve significant improvements in depth structures and robustness. Our work provides a new perspective on depth estimation with sparse annotations in autonomous driving scenes.
Jiaqi Li 0007, Yiran Wang 0005, Zihao Huang 0001, Jinghong Zheng 0002, Ke Xian, Zhiguo Cao 0001, Jianming Zhang 0001
ACM Multimedia5
2023 Make-It-4D: Synthesizing a Consistent Long-Term Dynamic Scene Video from a Single Image
abstract
We study the problem of synthesizing a long-term dynamic video from only a single image. This is challenging since it requires consistent visual content movements given large camera motions. Existing methods either hallucinate inconsistent perpetual views or struggle with long camera trajectories. To address these issues, it is essential to estimate the underlying 4D (including 3D geometry and scene motion) and fill in the occluded regions. To this end, we present Make-It-4D, a novel method that can generate a consistent long-term dynamic video from a single image. On the one hand, we utilize layered depth images (LDIs) to represent a scene, and they are then unprojected to form a feature point cloud. To animate the visual content, the feature point cloud is displaced based on the scene flow derived from motion estimation and the corresponding camera pose. Such 4D representation enables our method to maintain the global consistency of the generated dynamic video. On the other hand, we fill in the occluded regions by using a pre-trained diffusion model to inpaint and outpaint the input image. This enables our method to work under large camera motions. Benefiting from our design, our method can be training-free which saves a significant amount of training time. Experimental results demonstrate the effectiveness of our approach, which showcases compelling rendering results.
Liao Shen, Xingyi Li 0005, Huiqiang Sun, Juewen Peng, Ke Xian, Zhiguo Cao 0001, Guosheng Lin
ACM Multimedia5
2023 Large motion anime head animation using a cascade pose transform network
Ke Xian, Zhiguo Cao 0001
Pattern Recognit.3
2023 Point-and-Shoot All-in-Focus Photo Synthesis From Smartphone Camera Pair
abstract
All-in-Focus (AIF) photography is expected to be a commercial selling point for modern smartphones. Standard AIF synthesis requires manual, time-consuming operations such as focal stack compositing, which is unfriendly to ordinary people. To achieve point-and-shoot AIF photography with a smartphone, we expect that an AIF photo can be generated from one shot of the scene, instead of from multiple photos captured by the same camera. Benefiting from the multi-camera module in modern smartphones, we introduce a new task of AIF synthesis from main (wide) and ultra-wide cameras. The goal is to recover sharp details from defocused regions in the main-camera photo with the help of the ultra-wide-camera one. The camera setting poses new challenges such as parallax-induced occlusions and inconsistent color between cameras. To overcome the challenges, we introduce a predict-and-refine network to mitigate occlusions and propose dynamic frequency-domain alignment for color correction. To enable effective training and evaluation, we also build an AIF dataset with 2686 unique scenes. Each scene includes two photos captured by the main camera, one photo captured by the ultra-wide camera, and a synthesized AIF photo. Results show that our solution, termed EasyAIF, can produce high-quality AIF photos and outperforms strong baselines quantitatively and qualitatively. For the first time, we demonstrate point-and-shoot AIF photo synthesis successfully from main and ultra-wide cameras.
Xianrui Luo, Juewen Peng, Weiyue Zhao, Ke Xian, Hao Lu 0003, Zhiguo Cao 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 SymmNeRF: Learning to Explore Symmetry Prior for Single-View View Synthesis
Xingyi Li 0005, Chaoyi Hong, Yiran Wang 0005, Zhiguo Cao 0001, Ke Xian, Guosheng Lin
ACCV (1)5
2022 BokehMe: When Neural Rendering Meets Classical Rendering
abstract
We propose BokehMe, a hybrid bokeh rendering framework that marries a neural renderer with a classical physically motivated renderer. Given a single image and a potentially imperfect disparity map, BokehMe generates high-resolution photo-realistic bokeh effects with adjustable blur size, focal plane, and aperture shape. To this end, we analyze the errors from the classical scattering-based method and derive a formulation to calculate an error map. Based on this formulation, we implement the classical renderer by a scattering-based method and propose a two-stage neural renderer to fix the erroneous areas from the classical renderer. The neural renderer employs a dynamic multi-scale scheme to efficiently handle arbitrary blur sizes, and it is trained to handle imperfect disparity input. Experiments show that our method compares favorably against previous methods on both synthetic image data and real image data with predicted disparity. A user study is further conducted to validate the advantage of our method.
Juewen Peng, Zhiguo Cao 0001, Xianrui Luo, Hao Lu 0003, Ke Xian, Jianming Zhang 0001
CVPR5
2022 MPIB: An MPI-Based Bokeh Rendering Framework for Realistic Partial Occlusion Effects
Juewen Peng, Jianming Zhang 0001, Xianrui Luo, Hao Lu 0003, Ke Xian, Zhiguo Cao 0001
ECCV (6)5
2022 Discriminate Clearer To Rank Better: Image Cropping By Amplifying View-Wise Differences
abstract
Image cropping aims to enhance the aesthetic quality of a given image by searching for the good cropping views. One common routine is to score and rank the candidate views by the neural network. The network is expected to discriminate the subtle view-wise differences. However, the image-wise differences and the ambiguity in the annotations render difficulties in discriminating the view-wise differences. To focus on the view-wise differences, we propose a feature spliter to build image-wise and view-wise feature and evaluate the candidate views only based on the view-wise feature. Then, we propose the ranking gain loss that alleviates the ambiguity in annotations to amplify the view-wise differences. The remarkable improvement compared with prior arts on public benchmarks illustrates that the view-wise differences matter in cropping view recommendation.
Zhiguo Cao 0001, Ke Xian, Hao Lu 0003, Weicai Zhong
ICIP3
2022 Less is More: Consistent Video Depth Estimation with Masked Frames Modeling
abstract
Temporal consistency is the key challenge of video depth estimation. Previous works are based on additional optical flow or camera poses, which is time-consuming. By contrast, we derive consistency with less information. Since videos inherently exist with heavy temporal redundancy, a missing frame could be recovered from neighboring ones. Inspired by this, we propose the frame masking network (FMNet), a spatial-temporal transformer network predicting the depth of masked frames based on their neighboring frames. By reconstructing masked temporal features, the FMNet can learn intrinsic inter-frame correlations, which leads to consistency. Compared with prior arts, experimental results demonstrate that our approach achieves comparable spatial accuracy and higher temporal consistency without any additional information. Our work provides a new perspective on consistent video depth estimation.
Yiran Wang 0005, Xingyi Li 0005, Zhiguo Cao 0001, Ke Xian, Jianming Zhang 0001
ACM Multimedia5
2021 Composing Photos Like a Photographer
abstract
We show that explicit modeling of composition rules benefits image cropping. Image cropping is considered a promising way to automate aesthetic composition in professional photography. Existing efforts, however, only model such professional knowledge implicitly, e.g., by ranking from comparative candidates. Inspired by the observation that natural composition traits always follow a specific rule, we propose to learn such rules in a discriminative manner, and more importantly, to incorporate learned composition clues explicitly in the model. To this end, we introduce the concept of the key composition map (KCM) to encode the composition rules. The KCM can reveal the common laws hidden behind different composition rules and can inform the cropping model of what is important in composition. With the KCM, we present a novel cropping-by-composition paradigm and instantiate a network to implement composition-aware image cropping. Extensive experiments on two benchmarks justify that our approach enables effective, interpretable, and fast image cropping.
Chaoyi Hong, Shuaiyuan Du, Ke Xian, Hao Lu 0003, Zhiguo Cao 0001, Weicai Zhong
CVPR3
2021 Robust Image Cropping by Filtering Composition Irrelevant Factors
Ke Xian, Hao Lu 0003, Zhiguo Cao 0001
ICIG (3)2
2021 Context-Aware Candidates for Image Cropping
abstract
Image cropping aims to enhance the aesthetic quality of a given image by removing unwanted areas. Existing image cropping methods can be divided into two groups: candidate-based and candidate-free methods. For candidate-based methods, dense predefined candidate boxes can indeed cover good boxes, but most candidates with low aesthetic quality may disturb the following judgment and lead to an undesirable result. For candidate-free methods, the cropping box is directly acquired according to certain prior knowledge. However, the effect of only one box is not stable enough due to the subjectivity of image cropping. In order to combine the advantages of the above methods and overcome these shortcomings, we need fewer but more representative candidate boxes. To this end, we propose FCRNet, a fully convolutional regression network, which predicts several context-aware cropping boxes in an ensemble manner as candidates. A multi-task loss is employed to supervise the generation of candidates. Unlike previous candidate-based works, FCRNet outputs a small number of context-aware candidates without any predefined box and the final result is selected from these candidates by an aesthetic evaluation network or even manual selection. Extensive experiments show the superiority of our context-aware candidates based method over the state-of-the-art approaches.
Tianpei Lian, Zhiguo Cao 0001, Ke Xian, Weicai Zhong
ICIP3
2021 Interactive Portrait Bokeh Rendering System
abstract
Portrait bokeh rendering has become a hot topic in computer vision and graphics in recent years. Existing methods usually suffer from noticeable artifacts around foreground boundaries and unrealistic rendering effects. To tackle these problems, we design a brand new bokeh system in this paper. The system is comprised of three modules: depth estimation, portrait matting, and bokeh rendering. The introduction of the portrait matting module makes it possible to preserve the details of portraits in final rendering results. In bokeh rendering modules, we propose two pixelwise rendering methods which are based on light gathering and light scattering to render realistic bokeh effect. For flexibility and interactivity. We provide two parameter interfaces, i.e., aperture size and bokeh salience to adjust rendering details according to the preferences of different users. Finally, experimental results on our synthetic dataset and real images demonstrate the effectiveness of our proposed method.
Juewen Peng, Xianrui Luo, Ke Xian, Zhiguo Cao 0001
ICIP3
2021 A Performance Evaluation of Correspondence Grouping Methods for 3D Rigid Data Matching
abstract
Seeking consistent point-to-point correspondences between 3D rigid data (point clouds, meshes, or depth maps) is a fundamental problem in 3D computer vision. While a number of correspondence selection methods have been proposed in recent years, their advantages and shortcomings remain unclear regarding different applications and perturbations. To fill this gap, this paper gives a comprehensive evaluation of nine state-of-the-art 3D correspondence grouping methods. A good correspondence grouping algorithm is expected to retrieve as many as inliers from initial feature matches, giving a rise in both precision and recall as well as facilitating accurate transformation estimation. Toward this rule, we deploy experiments on three benchmarks with different application contexts, including shape retrieval, 3D object recognition, and point cloud registration. We also investigate various perturbations such as noise, point density variation, clutter, occlusion, partial overlap, different scales of initial correspondences, and different combinations of keypoint detectors and descriptors. The rich variety of application scenarios and nuisances result in different spatial distributions and inlier ratios of initial feature correspondences, thus enabling a thorough evaluation. Based on the outcomes, we give a summary of the traits, merits, and demerits of evaluated approaches and indicate some potential future research directions.
Jiaqi Yang 0002, Ke Xian, Peng Wang 0015, Yanning Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 CPTNet: Cascade Pose Transform Network for Single Image Talking Head Animation
Ke Xian, Yinpeng Chen, Zhiguo Cao 0001, Weicai Zhong
ACCV (4)2
2020 Structure-Guided Ranking Loss for Single Image Depth Prediction
abstract
Single image depth prediction is a challenging task due to its ill-posed nature and challenges with capturing ground truth for supervision. Large-scale disparity data generated from stereo photos and 3D videos is a promising source of supervision, however, such disparity data can only approximate the inverse ground truth depth up to an affine transformation. To more effectively learn from such pseudo-depth data, we propose to use a simple pair-wise ranking loss with a novel sampling strategy. Instead of randomly sampling point pairs, we guide the sampling to better characterize structure of important regions based on the low-level edge maps and high-level object instance masks. We show that the pair-wise ranking loss, combined with our structure-guided sampling strategies, can significantly improve the quality of depth map prediction. In addition, we introduce a new relative depth dataset of about 21K diverse high-resolution web stereo photos to enhance the generalization ability of our model. In experiments, we conduct cross-dataset evaluation on six benchmark datasets and show that our method consistently improves over the baselines, leading to superior quantitative and qualitative results.
Ke Xian, Jianming Zhang 0001, Oliver Wang, Long Mai, Zhe Lin 0001, Zhiguo Cao 0001
CVPR1
2020 Sparse-to-Dense Depth Completion Revisited: Sampling Strategy and Graph Construction
Haipeng Xiong, Ke Xian, Chen Zhao 0025, Zhiguo Cao 0001, Xin Li 0005
ECCV (21)3
2020 Multi - Direction Convolution for Semantic Segmentation
Dehui Li, Zhiguo Cao 0001, Ke Xian, Xinyuan Qi, Hao Lu 0003
ICPR3
2020 Monocular Depth Estimation With Augmented Ordinal Depth Relationships
abstract
Most existing algorithms for depth estimation from single monocular images need large quantities of metric ground-truth depths for supervised learning. We show that relative depth can be an informative cue for metric depth estimation and can be easily obtained from vast stereo videos. Acquiring metric depths from stereo videos are sometimes impracticable due to the absence of camera parameters. In this paper, we propose to improve the performance of metric depth estimation with relative depths collected from stereo movie videos using existing stereo matching algorithm. We introduce a new “relative depth in stereo” (RDIS) dataset densely labeled with relative depths. We first pretrain a ResNet model on our RDIS dataset. Then, we finetune the model on RGB-D datasets with metric ground-truth depths. During our finetuning, we formulate depth estimation as a classification task. This re-formulation scheme enables us to obtain the confidence of a depth prediction in the form of probability distribution. With this confidence, we propose an information gain loss to make use of the predictions that are close to ground-truth. We evaluate our approach on both indoor and outdoor benchmark RGB-D datasets and achieve the state-of-the-art performance.
Yuanzhouhan Cao, Ke Xian, Chunhua Shen, Zhiguo Cao 0001, Shugong Xu
IEEE Trans. Circuits Syst. Video Technol.3
2020 Counting Objects by Blockwise Classification
abstract
In this paper, we introduce the idea of blockwise classification to count objects. The current mainstream method for counting objects is to regress the density map or to regress the redundant count map via a deep convolutional neural network (CNN). However, these methods suffer from two critical issues: inaccurately generated regression targets and serious sample imbalances. First, the ground truth density map is generated by convolving the dot map using a Gaussian kernel. Because an inappropriate kernel can cover the background or uncover objects, this approach introduces a form of noise, and therefore results in ambiguities when training the networks. Second, inhomogeneously distributed objects often exist in images, which gives rise to a data collection bias. This leads to a long-tailed distribution of region counts, which is a typical characteristic that occurs with imbalanced samples; therefore, underestimations in high-density regions and overestimations in low-density regions are common. In this paper, we address these two issues within one framework-blockwise count level classification. The intuition behind this idea is that while it may not be possible to provide an exact count of pixels or patches, it is possible to provide a count of a region that falls within a certain interval with high confidence. Our method classifies the count levels of each block produced by nonlinearly quantizing the continuous counts, thus transforming the imbalance of sample patch counts into a class imbalance of count levels. Consequently, an information-entropy-inspired loss can be applied to alleviate this issue. Through ablative studies, we analyze the impact of imbalanced data, Gaussian kernel sizes, quantization errors, and the effectiveness of each module in our method. Without bells and whistles, our method outperforms or performs competitively with other state-of-the-art approaches on seven object-counting benchmarks, including four crowd-counting datasets from ShanghaiTech, WorldExpo'10, UCF-QNRF and UCF_CC_50, one vehicle-counting dataset (TRANCOS), one maize-tassel-counting dataset (MTC), and one challenging sonar fish-counting dataset that we constructed. The results suggest that our framework provides a strong and improved baseline for object counting.
Liang Liu 0001, Hao Lu 0003, Haipeng Xiong, Ke Xian, Zhiguo Cao 0001, Chunhua Shen
IEEE Trans. Circuits Syst. Video Technol.4
2020 Image Feature Correspondence Selection: A Comparative Study and a New Contribution
abstract
Image feature correspondence selection is pivotal to many computer vision tasks from object recognition to 3D reconstruction. Although many correspondence selection algorithms have been developed in the past decade, there still lacks an in-depth evaluation and comparison in the open literature, which makes it difficult to choose the appropriate algorithm for a specific application. This paper attempts to fill this gap by evaluating eight competing correspondence selection algorithms including both classical methods and current state-of-the-art ones. In addition to preselected correspondences, we have compared different combinations of detector and descriptor on four standard datasets. The diversity of those datasets cover a wide range of uncertainty factors including zoom, rotation, blur, viewpoint change, JPEG compression, light change, different rendering styles and multiple structures. We have measured the quality of competing correspondence selection algorithms in terms of four performance metrics -i.e., precision, recall, F-measure and efficiency. Moreover, we propose to combine the strengths of eight competing methods by combining their correspondence selection results. Extensive experimental results are reported to demonstrate the superiority of several fusion strategies to individual methods, which suggests the possibility of adaptively combining those methods for even better performance.
Chen Zhao 0025, Zhiguo Cao 0001, Jiaqi Yang 0002, Ke Xian, Xin Li 0005
IEEE Trans. Image Process.4
2019 Binoboost: Boosting Self-Supervised Monocular Depth Prediction with Binocular Guidance
abstract
In this paper, we study the problem of self-supervised monocular depth prediction. Owing to the fact that a vast quantity of expensive ground-truth depth data is required in supervised deep learning methods, self-supervised deep learning methods are what we apply to and are more approachable as well. Specifically, we base our model on a monocular disparity network that generates disparity images by training with an image reconstruction loss. Because binocular images implicitly provide epipolar geometry constraints, we find that binocular depth estimators always perform better than monocular ones. Therefore, we propose a novel module (i.e., BinoBoost) to boost our monocular disparity network with binocular guidance during training. In particular, the binocular disparity network is trained in a similar way to the base model to generate proxy ground truth disparity. We drive the outputs of the base model to be the same as relatively better outputs produced by the binocular disparity network. By training our base model and BinoBoost in an end-to-end fashion, we improve the performance on our base model and achieve the state-of-the-art results on the KITTI dataset.
Zhiguo Cao 0001, Ke Xian, Hongwei Zou
ICIP4
2019 Salient Object Detection via Deep Hierarchical Context Aggregation and Multi-Layer Supervision
abstract
The aggregation of hierarchical information is vital for saliency detection. To achieve this, most existing saliency detectors apply various network structures to fuse features. But most of them utilize shallow skip connections and only concentrate on the final results, which can not guarantee the model to learn the rich and accurate contextual information. To address these problems, we propose a network with deep layer aggregation and multi-layer intermediate supervision. We utilize deep layer aggregation to fuse features iteratively and hierarchically across layers to obtain richer information. Then we add multi-layer intermediate supervision on each side-output layer to capture more accurate contextual information. We evaluate our method on six benchmark datasets under various metrics and it achieves the new state-of-the-art.
Zhiguo Cao 0001, Ke Xian, Xinyuan Qi
ICIP4
2019 Mean-Variance Loss for Monocular Depth Estimation
abstract
Monocular depth estimation is a widely studied computer vision problem with a vast variety of applications. In this paper, we formulate it as a pixel-wise classification task and use a mean-variance loss for robust depth estimation via distribution learning. More precisely, the mean-variance loss is composed of a mean loss that penalizes the difference between the mean of predicted depth distribution and the ground-truth depth, and a variance loss that penalizes the variance of predicted depth distribution to obtain a more focused distribution. The mean-variance loss is jointly trained with the soft-max loss to supervise a Deep Convolutional Neural Networks (DCNN) for depth estimation. Experimental results on the NYUDv2 dataset show that the proposed method outperforms previous state-of-the-art approaches.
Hongwei Zou, Ke Xian, Jiaqi Yang 0002, Zhiguo Cao 0001
ICIP2
2019 Limited Receptive Field Network for Real-Time Driving Scene Semantic Segmentation
Dehui Li, Zhiguo Cao 0001, Ke Xian, Jiaqi Yang 0002, Xinyuan Qi, Wei Li 0132
PRICAI (3)3
2018 Deep Attention-Based Classification Network for Robust Depth Prediction
Ruibo Li, Ke Xian, Chunhua Shen, Zhiguo Cao 0001, Hao Lu 0003, Lingxiao Hang
ACCV (4)2
2018 Monocular Relative Depth Perception With Web Stereo Data Supervision
abstract
In this paper we study the problem of monocular relative depth perception in the wild. We introduce a simple yet effective method to automatically generate dense relative depth annotations from web stereo images, and propose a new dataset that consists of diverse images as well as corresponding dense relative depth maps. Further, an improved ranking loss is introduced to deal with imbalanced ordinal relations, enforcing the network to focus on a set of hard pairs. Experimental results demonstrate that our proposed approach not only achieves state-of-the-art accuracy of relative depth perception in the wild, but also benefits other dense per-pixel prediction tasks, e.g., metric depth estimation and semantic segmentation.
Ke Xian, Chunhua Shen, Zhiguo Cao 0001, Hao Lu 0003, Yang Xiao 0007, Ruibo Li, Zhenbo Luo
CVPR1
2017 Performance Evaluation of 3D Correspondence Grouping Algorithms
abstract
This paper presents a thorough evaluation of several widely-used 3D correspondence grouping algorithms, motived by their significance in vision tasks relying on correct feature correspondences. A good correspondence grouping algorithm is desired to retrieve as many as inliers from initial feature matches, giving a rise in both precision and recall. Towards this rule, we deploy the experiments on three benchmarks respectively addressing shape retrieval, 3D object recognition and point cloud registration scenarios. The variety in application context brings a rich category of nuisances including noise, varying point densities, clutter, occlusion and partial overlaps. It also results to different ratios of inliers and correspondence distributions for comprehensive evaluation. Based on the quantitative outcomes, we give a summarization of the merits/demerits of the evaluated algorithms from both performance and efficiency perspectives.
Jiaqi Yang 0002, Ke Xian, Yang Xiao 0007, Zhiguo Cao 0001
3DV2
2017 When Unsupervised Domain Adaptation Meets Tensor Representations
abstract
Domain adaption (DA) allows machine learning methods trained on data sampled from one distribution to be applied to data sampled from another. It is thus of great practical importance to the application of such methods. Despite the fact that tensor representations are widely used in Computer Vision to capture multi-linear relationships that affect the data, most existing DA methods are applicable to vectors only. This renders them incapable of reflecting and preserving important structure in many problems. We thus propose here a learning-based method to adapt the source and target tensor representations directly, without vectorization. In particular, a set of alignment matrices is introduced to align the tensor representations from both domains into the invariant tensor subspace. These alignment matrices and the tensor subspace are modeled as a joint optimization problem and can be learned adaptively from the data using the proposed alternative minimization scheme. Extensive experiments show that our approach is capable of preserving the discriminative power of the source domain, of resisting the effects of label noise, and works effectively for small sample sizes, and even one-shot DA. We show that our method outperforms the state-of-the-art on the task of cross-domain visual recognition in both efficacy and efficiency, and particularly that it outperforms all comparators when applied to DA of the convolutional activations of deep convolutional networks.
Hao Lu 0003, Lei Zhang 0054, Zhiguo Cao 0001, Wei Wei 0008, Ke Xian, Chunhua Shen, Anton van den Hengel
ICCV5
2017 Rotational contour signatures for both real-valued and binary feature representations of 3D local shape
Jiaqi Yang 0002, Qian Zhang 0046, Ke Xian, Yang Xiao 0007, Zhiguo Cao 0001
Comput. Vis. Image Underst.3
2016 Rotational contour signatures for robust local surface description
abstract
This paper presents a novel local surface descriptor called rotational contour signatures (RCS) for 3D rigid objects. RCS comprises several signatures that characterize the 2D contour information derived from 3D-to-2D projection of the local surface. The inspiration of our encoding technique comes from that, viewing towards an object, its contour is an effective and robust cue for representing its shape. In order to achieve a comprehensive geometry encoding, the local surface is continually rotated in a predefined local reference frame (LRF) so that multi-view information is obtained. Experiments on two publicly available datasets demonstrate the effectiveness and robustness of the proposed descriptor. Further, comparisons with five state-of-the-art descriptors show the superiority of our RCS descriptor.
Jiaqi Yang 0002, Qian Zhang 0046, Ke Xian, Yang Xiao 0007, Zhiguo Cao 0001
ICIP3
2016 Exploiting Attribute Dependency for Attribute Assignment in Crowded Scenes
abstract
Attributes now play a vital role for characterizing a crowded scene. Compared to low-level visual features, processing informed by attributes can capture rich semantic information. However, to effectively assign attributes to a crowded scene still remains a challenging task. In this letter, inspired by a recently proposed zero-shot learning framework, a novel attribute assignment method that maps low-level features to predefined attributes is proposed. In particular, we propose to exploit the attribute dependency during the phase of attribute assignment, which can be regarded as our main contribution. In addition, to further enhance the performance, an effective low-level feature extraction mechanism is also proposed. More precisely, appearance and motion features are first simultaneously extracted from several sampled video frames and corresponding optical flow fields via deep convolutional neural network and then, respectively, aggregated by using Fisher vector encoding to form the low-level representation of crowded scenes. Experimental results on the challenging WWW dataset demonstrate that both the proposed attribute assignment method and the low-level feature extraction mechanism outperform the state of the art.
Chunhua Deng, Zhiguo Cao 0001, Yang Xiao 0007, Hao Lu 0003, Ke Xian
IEEE Signal Process. Lett.5