VLDB 2026 Research / reviewers in the wild / expert
Xun Cao
dblp:78/7658
· DBLP profile ↗
102ranked-venue papers
6as first author
64since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 81 · 4 first-author · 52 since 2021Artificial intelligence and machine learning · 58 · 2 first-author · 42 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Split-Layer: Enhancing Implicit Neural Representation by Maximizing the Dimensionality of Feature SpaceabstractImplicit neural representation (INR) models signals as continuous functions using neural networks, offering efficient and differentiable optimization for inverse problems across diverse disciplines. However, the representational capacity of INR—defined by the range of functions the neural network can characterize—is inherently limited by the low-dimensional feature space in conventional multilayer perceptron (MLP) architectures. While widening the MLP can linearly increase feature space dimensionality, it also leads to a quadratic growth in computational and memory costs. To address this limitation, we propose the split-layer, a novel reformulation of MLP construction. The split-layer divides each layer into multiple parallel branches and integrates their outputs via Hadamard product, effectively constructing a high-degree polynomial space. This approach significantly enhances INR’s representational capacity by expanding the feature space dimensionality without incurring prohibitive computational overhead. Extensive experiments demonstrate that the split-layer substantially improves INR performance, surpassing existing methods across multiple tasks, including 2D image fitting, 2D CT reconstruction, 3D shape representation, and 5D novel view synthesis. Zhicheng Cai, Hao Zhu 0004, Linsen Chen, Qiu Shen, Xun Cao |
AAAI | 5 |
| 2026 | Spike Imaging Velocimetry: Dense Motion Estimation of Fluids Using Spike StreamsabstractParticle Image Velocimetry (PIV) is a widely adopted non-invasive imaging technique that tracks the motion of tracer particles across image sequences to capture the velocity distribution of fluid flows. It is commonly employed to analyze complex flow structures and validate numerical simulations. This study explores the untapped potential of spike cameras—ultra-high-speed, high-dynamic-range vision sensors—in high-speed fluid velocimetry. We propose a deep learning framework, Spike Imaging Velocimetry (SIV), tailored for high-resolution fluid motion estimation. To enhance the network’s performance, we design three novel modules specifically adapted to the characteristics of fluid dynamics and spike streams: the Detail-Preserving Hierarchical Transform (DPHT), the Graph Encoder (GE), and the Multi-scale Velocity Refinement (MSVR). Furthermore, we introduce a spike-based PIV dataset, Particle Scenes with Spike and Displacement (PSSD), which contains labeled samples from three representative fluid-dynamics scenarios: steady turbulence, high-speed flow, and high-dynamic-range conditions. Our proposed method outperforms existing baselines across all these scenarios, demonstrating its effectiveness. Yunzhong Zhang, Changqing Su, Zhen Cheng 0005, Zhaofei Yu, Tiejun Huang 0001, Xun Cao |
AAAI | 8 |
| 2026 | RHINO: regularizing the hash-based implicit neural representation
Hao Zhu 0004, Qi Zhang 0029, Zhan Ma 0001, Xun Cao |
Sci. China Inf. Sci. | 5 |
| 2026 | Toward the Spectral Bias Alleviation by Normalizations in Coordinate NetworksabstractRepresenting signals using coordinate networks dominates the area of inverse problems recently, and is widely applied in various scientific computing tasks. Still, there exists an issue of spectral bias in coordinate networks, limiting the capacity to learn high-frequency components. This problem is caused by the pathological distribution of the neural tangent kernel's (NTK's) eigenvalues of coordinate networks. We find that, this pathological distribution could be improved using classical normalization techniques (batch normalization and layer normalization), which are commonly used in convolutional neural networks but rarely used in coordinate networks. We prove that normalization techniques greatly reduces the maximum and variance of NTK's eigenvalues while slightly modifies the mean value, considering the max eigenvalue is much larger than the most, this variance change results in a shift of eigenvalues' distribution from a lower one to a higher one, therefore the spectral bias could be alleviated (see Fig. 1). Furthermore, we propose two new normalization techniques by combining these two techniques in different ways. The efficacy of these normalization techniques is substantiated by the significant improvements and new state-of-the-arts achieved by applying normalization-based coordinate networks to various tasks, including the image compression, computed tomography reconstruction, shape representation, magnetic resonance imaging, novel view synthesis and multi-view stereo reconstruction. Zhicheng Cai, Hao Zhu 0005, Qiu Shen, Xun Cao |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Hierarchical Bayesian Guided Spatial-, Angular- and Temporal-Consistent View SynthesisabstractNeural Radiance Fields (NeRF) have gained significant attention due to their precise reconstruction and rapid inference capabilities, making them highly promising for applications in virtual reality and gaming. However, extending NeRF's capabilities to dynamic scenes remains underexplored, particularly in ensuring consistent and coherent reconstructions across space, time, and viewing angles. To address this challenge, we propose Scale-NeRF, a novel approach that organizes the training of dynamic NeRFs as a progressive, scale-based refinement process, grounded in hierarchical Bayesian theory. Scale-NeRF begins by reconstructing the radiance fields using coarse, large-scale frames and iteratively refines them with progressively smaller-scale frames. This hierarchical strategy, combined with a corresponding sampling approach and a newly introduced structural loss, ensures consistency and integrity throughout the reconstruction process. Experiments on public datasets validate the superiority of Scale-NeRF over traditional methods, especially in terms of the proposed metrics evaluating spatial, angular, and temporal consistency. Furthermore, Scale-NeRF demonstrates excellent dynamic reconstruction capabilities with real-time rendering, offering a significant advancement for applications demanding both high fidelity and real-time performance. Junyu Zhu, Hao Zhu 0005, Zhan Ma 0001, Xun Cao |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid PriorabstractAudio-driven talking head generation has drawn much attention in recent years, and many efforts have been made in lip-sync, facial motion, head pose generation, and video quality. However, no model has yet led or tied on all these metrics due to the one-to-many mapping between audio and motion. In this paper, we propose VividTalk, a two-stage generic framework that supports generating high-visual quality talking head videos with all the above properties. Specifically, in the first stage, we map the audio to mesh by learning two motions, including non-rigid facial motion and rigid head motion. For facial motion, both blendshape and vertex are adopted as the intermediate representation to maximize the representation ability of the model. For head motion, a novel learnable head pose codebook with a two-phase training mechanism is proposed. In the second stage, we proposed a dual branch motion-vae and a generator to transform the meshes into dense motion and synthesize high-quality video frame-by-frame. Extensive experiments show that the proposed VividTalk can generate high-visual quality talking head videos with lip-sync and realistic enhanced by a large margin, and outperforms previous state-of-the-art works in objective and subjective comparisons. The code will be publicly released upon publication. Xusen Sun, Longhao Zhang, Hao Zhu 0004, Peng Zhang 0080, Bang Zhang, Xinya Ji, Kangneng Zhou, Daiheng Gao, Liefeng Bo, Xun Cao |
3DV | 10 |
| 2025 | Matrix3D: Large Photogrammetry Model All-in-OneabstractWe present Matrix3D, a unified model that performs several photogrammetry subtasks, including pose estimation, depth prediction, and novel view synthesis using just the same model. Matrix3D utilizes a multi-modal diffusion transformer (DiT) to integrate transformations across several modalities, such as images, camera parameters, and depth maps. The key to Matrix3D’s large-scale multi-modal training lies in the incorporation of a mask learning strategy. This enables full-modality model training even with partially complete data, such as bi-modality data of image-pose and image-depth pairs, thus significantly increases the pool of available training data. Matrix3D demonstrates state-of-the-art performance in pose estimation and novel view synthesis tasks. Additionally, it offers fine-grained control through multi-round interactions, making it an innovative tool for 3D content creation. Project page: https://nju-3dv.github.io/projects/matrix3d. Yuanxun Lu, Jingyang Zhang, Tian Fang, Jean-Daniel Nahmias, Yanghai Tsin, Long Quan, Xun Cao, Yao Yao 0008, Shiwei Li 0001 |
CVPR | 7 |
| 2025 | MotionPRO: Exploring the Role of Pressure in Human MoCap and BeyondabstractExisting human Motion Capture (MoCap) methods mostly focus on the visual similarity while neglecting the physical plausibility. As a result, downstream tasks such as driving virtual human in 3D scene or humanoid robots in real world suffer from issues such as timing drift and jitter, spatial problems like sliding and penetration, and poor global trajectory accuracy. In this paper, we revisit human MoCap from the perspective of interaction between human body and physical world by exploring the role of pressure. Firstly, we construct a large-scale human Motion capture dataset with Pressure, RGB and Optical sensors (named MotionPRO), which comprises 70 volunteers performing 400 types of motion, encompassing a total of 12.4M pose frames. Secondly, we examine both the necessity and effectiveness of the pressure signal through two challenging tasks: (1) pose and trajectory estimation based solely on pressure: We propose a network that incorporates a small kernel decoder and a long-short-term attention module, and proof that pressure could provide accurate global trajectory and plausible lower body pose. (2) pose and trajectory estimation by fusing pressure and RGB: We impose constraints on orthographic similarity along the camera axis and whole-body contact along the vertical axis to enhance the cross-attention strategy to fuse pressure and RGB feature maps. Experiments demonstrate that fusing pressure with RGB features not only significantly improves performance in terms of objective metrics, but also plausibly drives virtual humans (SMPL) in 3D scene. Furthermore, we demonstrate that incorporating physical perception enables humanoid robots to perform more precise and stable actions, which is highly beneficial for the development of embodied artificial intelligence. Project page is available at: https://nju-cite-mocaphumanoid.github.io/MotionPRO/ Shenghao Ren, Qiu Shen, Xun Cao |
CVPR | 8 |
| 2025 | FATE: Full-head Gaussian Avatar with Textural Editing from Monocular VideoabstractReconstructing high-fidelity, animatable 3D head avatars from effortlessly captured monocular videos is a pivotal yet formidable challenge. Although significant progress has been made in rendering performance and manipulation capabilities, notable challenges remain, including incomplete reconstruction and inefficient Gaussian representation. To address these challenges, we introduce FATE — a novel method for reconstructing an editable full-head avatar from a single monocular video. FATE integrates a sampling-based densification strategy to ensure optimal positional distribution of points, improving rendering efficiency. A neural baking technique is introduced to convert discrete Gaussian representations into continuous attribute maps, facilitating intuitive appearance editing. Furthermore, we propose a universal completion framework to recover non-frontal appearance, culminating in a 360° -renderable 3D head avatar. FATE outperforms previous approaches in both qualitative and quantitative evaluations, achieving state-of-the-art performance. To the best of our knowledge, FATE is the first animatable and 360° full-head monocular reconstruction method for a 3D head avatar. Project page and code are available at this link. Zhiyang Liang 0002, Dongfang Hu, Yao Yao 0008, Xun Cao, Hao Zhu 0004 |
CVPR | 7 |
| 2025 | Mitigating Ambiguities in 3D Classification with Gaussian Splattingabstract3D classification with point cloud input is a fundamental problem in 3D vision. However, due to the discrete nature and the insufficient material description of point cloud representations, there are ambiguities in distinguishing wire-like and flat surfaces, as well as transparent or reflective objects. To address these issues, we propose Gaussian Splatting (GS) point cloud-based 3D classification. We find that the scale and rotation coefficients in the GS point cloud help characterize surface types. Specifically, wire-like surfaces consist of multiple slender Gaussian ellipsoids, while flat surfaces are composed of a few flat Gaussian ellipsoids. Additionally, the opacity in the GS point cloud represents the transparency characteristics of objects. As a result, ambiguities in point cloud-based 3D classification can be mitigated utilizing GS point cloud as input. To verify the effectiveness of GS point cloud input, we construct the first real-world GS point cloud dataset in the community, which includes 20 categories with 200 objects in each category. Experiments not only validate the superiority of GS point cloud input, especially in distinguishing ambiguous objects, but also demonstrate the generalization ability across different classification methods. Our project page: https://ruiqi-nju.github.io/MACGS. Hao Zhu 0004, Qi Zhang 0029, Xun Cao, Zhan Ma 0001 |
CVPR | 5 |
| 2025 | IDOL: Instant Photorealistic 3D Human Creation from a Single ImageabstractCreating a high-fidelity, animatable 3D full-body avatar from a single image is a challenging task due to the diverse appearance and poses of humans and the limited availability of high-quality training data. To achieve fast and high-quality human reconstruction, this work rethinks the task from the perspectives of dataset, model, and representation. First, we introduce a large-scale HUman-centric GEnerated dataset, HuGe100K, consisting of 100K diverse, photorealistic sets of human images. Each set contains 24-view frames in specific human poses, generated using a pose-controllable image-to-multi-view model. Next, leveraging the diversity in views, poses, and appearances within HuGe100K, we develop a scalable feed-forward transformer model to predict a 3D human Gaussian representation in a uniform space from a given human image. This model is trained to disentangle human pose, body shape, clothing geometry, and texture. The estimated Gaussians can be animated without post-processing. We conduct comprehensive experiments to validate the effectiveness of the proposed dataset and method. Our model demonstrates the ability to efficiently reconstruct photorealistic humans at 1K resolution from a single input image using a single GPU instantly. Additionally, it seamlessly supports various applications, as well as shape and texture editing tasks. Yiyu Zhuang, Jiaxi Lv, Hao Wen 0005, Qing Shuai, Ailing Zeng, Hao Zhu 0004, Shifeng Chen, Yujiu Yang 0001, Xun Cao, Wei Liu 0005 |
CVPR | 9 |
| 2025 | Tera: Rethinking Text-Guided Realistic 3D Avatar GenerationabstractIn this paper, we rethink text-to-avatar generative models by proposing TeRA, a more efficient and effective framework than the previous SDS-based models and general large 3D generative models. Our approach employs a two-stage training strategy for learning a native 3D avatar generative model. Initially, we distill a decoder to derive a structured latent space from a large human reconstruction model. Subsequently, a text-controlled latent diffusion model is trained to generate photorealistic 3D human avatars within this latent space. TeRA enhances the model performance by eliminating slow iterative optimization and enables text-based partial customization through a structured 3D human representation. Experiments have proven our approach's superiority over previous text-to-avatar generative models in subjective and objective evaluation. Yiyu Zhuang, Yifei Zeng, Xun Cao, Xinxin Zuo, Hao Zhu 0004 |
ICCV | 6 |
| 2025 | Epona: Autoregressive Diffusion World Model for Autonomous DrivingabstractDiffusion models have demonstrated exceptional visual quality in video generation, making them promising for autonomous driving world modeling. However, existing video diffusion-based world models struggle with flexible-length, long-horizon predictions and integrating trajectory planning. This is because conventional video diffusion models rely on global joint distribution modeling of fixed-length frame sequences rather than sequentially constructing localized distributions at each timestep. In this work, we propose Epona, an autoregressive diffusion world model that enables localized spatiotemporal distribution modeling through two key innovations: 1) Decoupled spatiotemporal factorization that separates temporal dynamics modeling from fine-grained future world generation, and 2) Modular trajectory and video prediction that seamlessly integrate motion planning with visual modeling in an end-to-end framework. Our architecture enables high-resolution, long-duration generation while introducing a novel chain-of-forward training strategy to address error accumulation in autoregressive loops. Experimental results demonstrate state-of-the-art performance with 7.4\% FVD improvement and minutes longer prediction duration compared to prior works. The learned world model further serves as a real-time motion planner, outperforming strong end-to-end planners on NAVSIM benchmarks. Code will be publicly available at \href{https://github.com/Kevin-thu/Epona/}{https://github.com/Kevin-thu/Epona/}. Kaiwen Zhang 0015, Zhenyu Tang 0004, Xiaotao Hu, Xingang Pan, Yuan Liu 0025, Li Yuan 0007, Qian Zhang 0001, Xiao-Xiao Long, Xun Cao, Wei Yin 0006 |
ICCV | 11 |
| 2025 | WIPES: Wavelet-based Visual PrimitivesabstractPursuing a continuous visual representation that offers flexible frequency modulation and fast rendering speed has recently garnered increasing attention in the fields of 3D vision and graphics. However, existing representations often rely on frequency guidance or complex neural network decoding, leading to spectrum loss or slow rendering. To address these limitations, we propose WIPES, a universal Wavelet-based vIsual PrimitivES for representing multi-dimensional visual signals. Building on the spatial-frequency localization advantages of wavelets, WIPES effectively captures both the low-frequency "forest" and the high-frequency "trees." Additionally, we develop a wavelet-based differentiable rasterizer to achieve fast visual rendering. Experimental results on various visual tasks, including 2D image representation, 5D static and 6D dynamic novel view synthesis, demonstrate that WIPES, as a visual primitive, offers higher rendering quality and faster inference than INR-based methods, and outperforms Gaussian-based representations in rendering quality. Hao Zhu 0004, Delong Wu, Linchao Bao, Xun Cao, Zhan Ma 0001 |
ICCV | 6 |
| 2025 | M-SpecGene: Generalized Foundation Model for RGBT Multispectral Vision
Kailai Zhou, Fuqiang Yang, Shixian Wang, Bihan Wen, Chongde Zi, Linsen Chen, Qiu Shen, Xun Cao |
ICCV | 8 |
| 2025 | Flow Distillation Sampling: Regularizing 3D Gaussians with Pre-trained Matching Priorsabstract3D Gaussian Splatting (3DGS) has achieved excellent rendering quality with fast training and rendering speed. However, its optimization process lacks explicit geometric constraints, leading to suboptimal geometric reconstruction in regions with sparse or no observational input views. In this work, we try to mitigate the issue by incorporating a pre-trained matching prior to the 3DGS optimization process. We introduce Flow Distillation Sampling (FDS), a technique that leverages pre-trained geometric knowledge to bolster the accuracy of the Gaussian radiance field. Our method employs a strategic sampling technique to target unobserved views adjacent to the input views, utilizing the optical flow calculated from the matching model (Prior Flow) to guide the flow analytically calculated from the 3DGS geometry (Radiance Flow). Comprehensive experiments in depth rendering, mesh reconstruction, and novel view synthesis showcase the significant advantages of FDS over state-of-the-art methods. Additionally, our interpretive experiments and analysis aim to shed light on the effects of FDS on geometric accuracy and rendering quality, potentially providing readers with insights into its performance. Lin-Zhuo Chen, Kangjie Liu, Youtian Lin, Zhihao Li 0002, Siyu Zhu 0001, Xun Cao, Yao Yao 0008 |
ICLR | 6 |
| 2025 | RealityAvatar: Comprehensive Head Avatar Generation with 360° RenderingabstractWe present RealityAvatar, a novel model for 360°-renderable head avatar generation. RealityAvatar supports both 3D Gaussians and textured polygen 3D mesh as representation. A multi-stage framework is proposed, beginning with a pretrained 3D GAN generator that produces multi-view images from random noise, text, or image inputs, which are then used as training data. 3D Gaussians are bound to initialized FLAME faces, with their scales and rotations determined by the corresponding faces. FLAME parameters and vertex positions are optimized for a detailed mesh consistent with the FLAME topology. Then, the mesh and images are refined via a differentiable renderer to produce texture maps. Finally, 3D Gaussians undergo unconstrained optimization for enhanced rendering quality. To the best of our knowledge, RealityAvatar is the first generative model for 360°-renderable and animatable human heads. The experimental results show high rendering quality and robust drivability. Houteng Yu, Xun Cao |
ICME | 3 |
| 2025 | Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse AttentionabstractGenerating high-resolution 3D shapes using volumetric representations such as Signed Distance Functions (SDFs) presents substantial computational and memory challenges. We introduce Direct3D-S2, a scalable 3D generation framework based on sparse volumes that achieves superior output quality with dramatically reduced training costs.
Our key innovation is the Spatial Sparse Attention (SSA) mechanism, which greatly enhances the efficiency of Diffusion Transformer (DiT) computations on sparse volumetric data. SSA allows the model to effectively process large token sets within sparse volumes, significantly reducing computational overhead and achieving a 3.9$\times$ speedup in the forward pass and a 9.6$\times$ speedup in the backward pass.
Our framework also includes a variational autoencoder (VAE) that maintains a consistent sparse volumetric format across input, latent, and output stages. Compared to previous methods with heterogeneous representations in 3D VAE, this unified design significantly improves training efficiency and stability.
Our model is trained on public datasets, and experiments demonstrate that Direct3D-S2 not only surpasses state-of-the-art methods in generation quality and efficiency, but also enables training at 1024³ resolution using only 8 GPUs—a task typically requiring at least 32 GPUs for volumetric representations at $256^3$ resolution, thus making gigascale 3D generation both practical and accessible. Project page: https://www.neural4d.com/research-page/direct3d-s2. Youtian Lin, Feihu Zhang, Yifei Zeng, Yajie Bao, Jiachen Qian, Siyu Zhu 0001, Xun Cao, Philip Torr 0001, Yao Yao 0008 |
NeurIPS | 9 |
| 2025 | Sketch2PoseNet: Efficient and Generalized Sketch to 3D Human Pose Predictionabstract3D human pose estimation from sketches has broad applications in computer animation and film production. Unlike traditional human pose estimation, this task presents unique challenges due to the abstract and disproportionate nature of sketches. Previous sketch-to-pose methods, constrained by the lack of large-scale sketch-3D pose annotations, primarily relied on optimization with heuristic rules—an approach that is both time-consuming and limited in generalizability. To address these challenges, we propose a novel approach leveraging a "learn from synthesis" strategy. Firstly, a diffusion model is learned to synthesize sketch images from 2D poses projected from 3D human poses, mimicking disproportionate human structures in sketches. This process enables the creation of a synthetic dataset, SKEP-120K, consisting of 120k accurate sketch-3D pose annotation pairs across various sketch styles. Building on this synthetic dataset, we introduce an end-to-end data-driven framework for estimating human poses and shapes from diverse sketch styles. Our framework combines existing 2D pose detectors and generative diffusion priors for sketch feature extraction with a feed-forward neural network for efficient 2D pose estimation. Multiple heuristic loss functions have been incorporated to guarantee geometric coherence between the derived 3D poses and the detected 2D poses while preserving accurate self-contacts. Qualitative, quantitative, and subjective evaluations collectively affirm that our proposed model substantially surpasses previous ones in both estimation accuracy and speed for sketch-to-pose tasks. Yiyu Zhuang, Xun Cao, Chuan Guo 0002, Xinxin Zuo, Hao Zhu 0004 |
SIGGRAPH Asia | 4 |
| 2025 | Towards Edge Deployment: An Ultra Lightweight Encoder for Learned Image CompressionabstractLearned Image Compression has shown superior performance over traditional codecs. However, the high computational complexity of existing LIC methods hinders their deployment on resource-constrained edge devices. To address this challenge, we propose an ultra lightweight encoder with an asymmetric architecture: while the encoder is designed to be extremely simple, the decoder employs a more complex synthesis network to maintain high reconstruction quality. The proposed method achieves better rate-distortion performance than BPG with an encoder even lighter than JPEG2000. Peijie Diao, Qiu Shen, Xun Cao |
VCIP | 4 |
| 2025 | Prior Image Guided Snapshot High-Resolution Spectral Imaging in Near InfraredabstractNear-Infrared (NIR) hyperspectral imaging opens up numerous possibilities for wide applications. Despite Compressive Spectral Imaging (CSI) being a promising technique, which enables the acquisition of three-dimensional (3D) spatio-spectral information from dynamic scenes, applying it to the NIR spectrum remains challenging. The bottleneck lies in the high cost and limited resolution of InGaAs Focal Plane Arrays (FPAs), which further degrade the high-frequency information of the compressed measurements. Here we demonstrate a novel Effective Prior Image-guided Spectral imager, termed EpiSpec, towards high-resolution spectral imaging in the NIR. Our key observation is that the tail response of low-cost silicon-based sensors tends to capture similar image, offering high spatial resolution guidance for retrieving details. Hence, the degraded measurement of hyperspectral scene, guided by the prior image, is capable of obtaining high-quality reconstructions. Since the prior image integrates only a partial spectrum of the target scene, introducing content-aware chromatic errors, we propose the Prior Image Guided Deep Unfolding Framework (PIUF) for high-fidelity spectral reconstruction. This framework implicitly models the underlying non-linear relationship between the degraded measurements and the Non-Panchromatic (NPA) prior image. We also introduce a new NIR Spectral Images Dataset (NISID), which features a broad selection of real-world NIR spectral interesting scenes. Based on the dataset in hand, we evaluate the sparse structure of such spectra, which can serve as a guide for efficient CSI sensing matrices design. Extensive evaluations on representative CSI systems demonstrate the effectiveness of the proposed EpiSpec framework. Subsequently, lab prototypes are built for real-world imaging validation, further supporting the viability of high-resolution spectral imaging in the NIR. Lijing Cai, Linsen Chen, Qiu Shen, Xun Cao |
IEEE Trans. Image Process. | 7 |
| 2025 | Joint Spatial and Frequency Domain Learning for Lightweight Spectral Image DemosaicingabstractConventional spectral image demosaicing algorithms rely on pixels' spatial or spectral correlations for reconstruction. Due to the missing data in the multispectral filter array (MSFA), the estimation of spatial or spectral correlations is inaccurate, leading to poor reconstruction results, and these algorithms are time-consuming. Deep learning-based spectral image demosaicing methods directly learn the nonlinear mapping relationship between 2D spectral mosaic images and 3D multispectral images. However, these learning-based methods focused only on learning the mapping relationship in the spatial domain, but neglected valuable image information in the frequency domain, resulting in limited reconstruction quality. To address the above issues, this paper proposes a novel lightweight spectral image demosaicing method based on joint spatial and frequency domain information learning. First, a novel parameter-free spectral image initialization strategy based on the Fourier transform is proposed, which leads to better initialized spectral images and eases the difficulty of subsequent spectral image reconstruction. Furthermore, an efficient spatial-frequency transformer network is proposed, which jointly learns the spatial correlations and the frequency domain characteristics. Compared to existing learning-based spectral image demosaicing methods, the proposed method significantly reduces the number of model parameters and computational complexity. Extensive experiments on simulated and real-world data show that the proposed method notably outperforms existing spectral image demosaicing methods. Xun Cao, Weisheng Dong, Guangming Shi |
IEEE Trans. Image Process. | 4 |
| 2025 | RefConv: Reparameterized Refocusing Convolution for Powerful ConvNetsabstractWe propose reparameterized refocusing convolution (RefConv) as a replacement for regular convolutional layers, which is a plug-and-play module to improve the performance without any inference costs. Specifically, given a pretrained model, RefConv applies a trainable Refocusing Transformation to the basis kernels inherited from the pretrained model to establish connections among the parameters. For example, a depthwise RefConv can relate the parameters of a specific channel of convolution kernel to the parameters of the other kernel, i.e., make them refocus on the other parts of the model they have never attended to, rather than focus on the input features only. From another perspective, RefConv augments the priors of existing model structures by utilizing the representations encoded in the pretrained parameters as the priors and refocusing on them to learn novel representations, thus further enhancing the representational capacity of the pretrained model. The experimental results validated that RefConv can improve multiple convolutional neural network (CNN)-based models by a clear margin on image classification (up to 1.47% higher top-1 accuracy on ImageNet), object detection, semantic segmentation, and adversarial attacks without introducing any extra inference costs or altering the original model structure. Further studies demonstrated that RefConv can strengthen the spatial skeletons of kernels, reduce the redundancy of channels, and smooth the loss landscape, which explains its effectiveness. Zhicheng Cai, Xiaohan Ding, Qiu Shen, Xun Cao |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | A Pre-convolved Representation for Plug-and-Play Neural Illumination FieldsabstractRecent advances in implicit neural representation have demonstrated the ability to recover detailed geometry and material from multi-view images. However, the use of simplified lighting models such as environment maps to represent non-distant illumination, or using a network to fit indirect light modeling without a solid basis, can lead to an undesirable decomposition between lighting and material. To address this, we propose a fully differentiable framework named Neural Illumination Fields (NeIF) that uses radiance fields as a lighting model to handle complex lighting in a physically based way. Together with integral lobe encoding for roughness-adaptive specular lobe and leveraging the pre-convolved background for accurate decomposition, the proposed method represents a significant step towards integrating physically based rendering into the NeRF representation. The experiments demonstrate the superior performance of novel-view rendering compared to previous works, and the capability to re-render objects under arbitrary NeRF-style environments opens up exciting possibilities for bridging the gap between virtual and real-world scenes. Yiyu Zhuang, Qi Zhang 0029, Xuan Wang 0009, Hao Zhu 0004, Xiaoyu Li 0002, Ying Shan, Xun Cao |
AAAI | 8 |
| 2024 | Batch Normalization Alleviates the Spectral Bias in Coordinate NetworksabstractRepresenting signals using coordinate networks domi-nates the area of inverse problems recently, and is widely applied in various scientific computing tasks. Still, there exists an issue of spectral bias in coordinate networks, lim-iting the capacity to learn high-frequency components. This problem is caused by the pathological distribution of the neural tangent kernel's (NTK's) eigenvalues of coordinate networks. We find that, this pathological distribution could be improved using the classical batch normalization (BN), which is a common deep learning technique but rarely used in coordinate networks. BN greatly reduces the maximum and variance of NTK's eigenvalues while slightly modifies the mean value, considering the max eigenvalue is much larger than the most, this variance change results in a shift of eigenvalues' distribution from a lower one to a higher one, therefore the spectral bias could be alleviated (see Fig. 1). This observation is substantiated by the significant improvements of applying BN-based coordinate networks to various tasks, including the image compression, computed tomography reconstruction, shape representation, magnetic resonance imaging and novel view synthesis. Zhicheng Cai, Hao Zhu 0004, Qiu Shen, Xun Cao |
CVPR | 5 |
| 2024 | FINER: Flexible Spectral-Bias Tuning in Implicit NEural Representation by Variableperiodic Activation FunctionsabstractImplicit Neural Representation (INR), which utilizes a neural network to map coordinate inputs to corresponding attributes, is causing a revolution in the field of signal processing. However, current INR techniques suffer from a re-stricted capability to tune their supported frequency set, re-sulting in imperfect performance when representing complex signals with multiple frequencies. We have identified that this frequency-related problem can be greatly alleviated by introducing variableperiodic activation functions, for which we propose FINER. By initializing the bias of the neural network within different ranges, sub-functions with various frequencies in the variableperiodic function are selected for activation. Consequently, the supported frequency set of FINER can be flexibly tuned, leading to improved performance in signal representation. We demon-strate the capabilities of FINER in the contexts of2D image fitting, 3D signed distance field representation, and 5D neural radiance fields optimization, and we show that it outper-forms existing INRs. Zhen Liu 0031, Hao Zhu 0005, Qi Zhang 0029, Jingde Fu, Weibing Deng, Zhan Ma 0001, Yanwen Guo 0001, Xun Cao |
CVPR | 8 |
| 2024 | Direct2.5: Diverse Text-to-3D Generation via Multi-view 2.5D DiffusionabstractRecent advances in generative AI have unveiled significant potential for the creation of 3D content. However, current methods either apply a pre-trained 2D diffusion model with the time-consuming score distillation sampling (SDS), or a direct 3D diffusion model trained on limited 3D data losing generation diversity. In this work, we approach the problem by employing a multi-view 2.5D diffusion fine-tuned from a pre-trained 2D diffusion model. The multi-view 2.5D diffusion directly models the structural distribution of 3D data, while still maintaining the strong generalization ability of the original 2D diffusion model, filling the gap between 2D diffusion-based and direct 3D diffusion-based methods for 3D content generation. During inference, multi-view normal maps are generated using the 2.5D diffusion, and a novel differentiable rasterization scheme is introduced to fuse the almost consistent multi-view normal maps into a consistent 3D model. We further design a normal-conditioned multi-view image generation module for fast appearance generation given the 3D geometry. Our method is a one-pass diffusion process and does not require any SDS optimization as post-processing. We demonstrate through extensive experiments that, our direct 2.5D generation with the specially-designed fusion scheme can achieve diverse, mode-seeking-free, and high-fidelity 3D content generation in only 10 seconds. Project page: https://nju-3dv.github.io/projects/direct25. Yuanxun Lu, Jingyang Zhang, Shiwei Li 0001, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan, Xun Cao, Yao Yao 0008 |
CVPR | 8 |
| 2024 | MMVP: A Multimodal MoCap Dataset with Vision and Pressure SensorsabstractFoot contact is an important cue for human motion capture, understanding, and generation. Existing datasets tend to annotate dense foot contact using visual matching with thresholding or incorporating pressure signals. However, these approaches either suffer from low accuracy or are only designed for small-range and slow motion. There is still a lack of a vision-pressure multimodal dataset with large-range and fast human motion, as well as accurate and dense foot-contact annotation. To fill this gap, we propose a Multimodal MoCap Dataset with Vision and Pressure sensors, named MMVP. MMVP provides accurate and dense plantar pressure signals synchronized with RGBD observations, which is especially useful for both plausible shape estimation, robust pose fitting without foot drifting, and accurate global translation tracking. To validate the dataset, we propose an RGBD-P SMPL fitting method and also a monocular-video-based baseline framework, VP-MoCap, for human motion capture. Experiments demonstrate that our RGBD-P SMPL Fitting results significantly outperform pure visual motion capture. Moreover, VP-MoCap outperforms SOTA methods in foot-contact and global translation estimation accuracy. We believe the configuration of the dataset and the baseline frameworks will stimulate the research in this direction and also provide a good reference for MoCap applications in various domains. Project page: https://metaverse-ai-lab-thu.github.io/MMVP-Dataset/ He Zhang 0015, Shenghao Ren, Haolei Yuan, Jianhui Zhao 0002, Fan Li 0023, Shuangpeng Sun, Zhenghao Liang, Tao Yu 0007, Qiu Shen, Xun Cao |
CVPR | 10 |
| 2024 | Relightable 3D Gaussians: Realistic Point Cloud Relighting with BRDF Decomposition and Ray Tracing
Jian Gao 0009, Chun Gu, Youtian Lin, Zhihao Li 0002, Hao Zhu 0004, Xun Cao, Li Zhang 0040, Yao Yao 0008 |
ECCV (45) | 6 |
| 2024 | EmoTalk3D: High-Fidelity Free-View Synthesis of Emotional 3D Talking Head
Qianyun He, Xinya Ji, Yuanxun Lu, Zhengyu Diao, Linjia Huang, Yao Yao 0008, Siyu Zhu 0001, Zhan Ma 0001, Songcen Xu, Zixiao Zhang, Xun Cao, Hao Zhu 0004 |
ECCV (57) | 13 |
| 2024 | Head360: Learning a Parametric 3D Full-Head for Free-View Synthesis in 360$^\circ $
Yuxiao He, Yiyu Zhuang, Yao Yao 0008, Siyu Zhu 0001, Xiaoyu Li 0002, Qi Zhang 0029, Xun Cao, Hao Zhu 0004 |
ECCV (56) | 8 |
| 2024 | Efficient Snapshot Spectral Imaging: Calibration-Free Parallel Structure with Aperture Diffraction Fusion
Lihao Hu, Shiqiao Li, Xun Cao |
ECCV (51) | 5 |
| 2024 | Neural Poisson Solver: A Universal and Continuous Framework for Natural Signal Blending
Delong Wu, Hao Zhu 0005, Qi Zhang 0029, Zhan Ma 0001, Xun Cao |
ECCV (80) | 6 |
| 2024 | STAG4D: Spatial-Temporal Anchored Generative 4D Gaussians
Yifei Zeng, Yanqin Jiang, Siyu Zhu 0001, Yuanxun Lu, Youtian Lin, Hao Zhu 0004, Weiming Hu 0004, Xun Cao, Yao Yao 0008 |
ECCV (36) | 8 |
| 2024 | Joint RGB-Spectral Decomposition Model Guided Image Enhancement in Mobile Photography
Kailai Zhou, Lijing Cai, Yibo Wang 0004, Bihan Wen, Qiu Shen, Xun Cao |
ECCV (13) | 7 |
| 2024 | Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance
Shenhao Zhu, Junming Leo Chen, Zuozhuo Dai, Zilong Dong, Xun Cao, Yao Yao 0008, Hao Zhu 0004, Siyu Zhu 0001 |
ECCV (55) | 6 |
| 2024 | Token-Based Spatiotemporal Representation of the EventsabstractThe event camera’s low power consumption and ability to capture microsecond brightness changes make it attractive for various computer vision tasks. Existing event representation methods typically convert events into frames, voxel grids, or spikes for deep neural networks (DNNs). However, these approaches often sacrifice temporal granularity or require specialized devices for processing. This work introduces a novel token-based event representation, where each event is considered a fundamental processing unit termed an event-token. This approach preserves the sequence’s intricate spatiotemporal attributes at the event level. Moreover, we propose a Three-way Attention mechanism in the Event Transformer Block (ETB) to collaboratively construct temporal and spatial correlations between events. We compare our proposed token-based event representation extensively with other prevalent methods for object classification and optical flow estimation. The experimental results showcase its competitive performance while demanding minimal computational resources on standard devices. Bin Jiang 0018, Muhammad Salman Asif, Xun Cao, Zhan Ma 0001 |
ICASSP | 4 |
| 2024 | NeRI: Implicit Neural Representation of LiDAR Point Cloud Using Range Image SequenceabstractThis paper proposes the NeRI, an implicit neural representation (INR) based LiDAR point cloud compressor. In NeRI, we first transform a sequence of 3D LiDAR frames into a 2D range image sequence through range image projection over time. Then, we employ a neural network conditioned on the temporal frame index and associated LiDAR sensor pose to fit input range images as closely as possible. The optimized network parameters, which implicitly represent the input LiDAR data, are later lossily compressed. NeRI decoder is then initialized using decoded parameters to generate range images for reconstructing the 3D LiDAR sequence accordingly. Extensive experimental results demonstrate the significant superiority of NeRI regarding the compression efficiency and decoding speed compared to state-of-the-art 2D and 3D compressors for LiDAR point cloud. Ruixiang Xue, Tong Chen 0004, Dandan Ding, Xun Cao, Zhan Ma 0001 |
ICASSP | 5 |
| 2024 | Leveraging RGB-Pressure for Whole-body Human-to-Humanoid Motion Imitation
Shenghao Ren, Qiu Shen, Xun Cao |
ACM Multimedia | 4 |
| 2024 | HINER: Neural Representation for Hyperspectral ImageabstractThis paper introduces HINER, a novel neural representation for compressing HSI and ensuring high-quality downstream tasks on compressed HSI. HINER fully exploits inter-spectral correlations by explicitly encoding of spectral wavelengths and achieves a compact representation of the input HSI sample through joint optimization with a learnable decoder. By additionally incorporating the Content Angle Mapper with the L1 loss, we can supervise the global and local information within each spectral band, thereby enhancing the overall reconstruction quality. For downstream classification on compressed HSI, we theoretically demonstrate the task accuracy is not only related to the classification loss but also to the reconstruction fidelity through a first-order expansion of the accuracy degradation, and accordingly adapt the reconstruction by introducing Adaptive Spectral Weighting. Owing to the monotonic mapping of HINER between wavelengths and spectral bands, we propose Implicit Spectral Interpolation for data augmentation by adding random variables to input wavelengths during classification model training. Experimental results on various HSI datasets demonstrate the superior compression performance of our HINER compared to the existing learned methods and also the traditional codecs. Our model is lightweight and computationally efficient, which maintains high accuracy for downstream classification task even on decoded HSIs at high compression ratios. Our materials will be released at https://github.com/Eric-qi/HINER. Junqi Shi, Mingyi Jiang, Ming Lu 0003, Tong Chen 0004, Xun Cao, Zhan Ma 0001 |
ACM Multimedia | 5 |
| 2024 | Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion TransformerabstractGenerating high-quality 3D assets from text and images has long been challenging, primarily due to the absence of scalable 3D representations capable of capturing intricate geometry distributions. In this work, we introduce Direct3D, a native 3D generative model scalable to in-the-wild input images, without requiring a multi-view diffusion model or SDS optimization. Our approach comprises two primary components: a Direct 3D Variational Auto-Encoder (D3D-VAE) and a Direct 3D Diffusion Transformer (D3D-DiT). D3D-VAE efficiently encodes high-resolution 3D shapes into a compact and continuous latent triplane space. Notably, our method directly supervises the decoded geometry using a semi-continuous surface sampling strategy, diverging from previous methods relying on rendered images as supervision signals. D3D-DiT models the distribution of encoded 3D latents and is specifically designed to fuse positional information from the three feature maps of the triplane latent, enabling a native 3D generative model scalable to large-scale 3D datasets. Additionally, we introduce an innovative image-to-3D generation pipeline incorporating semantic and pixel-level image conditions, allowing the model to produce 3D shapes consistent with the provided conditional image input. Extensive experiments demonstrate the superiority of our large-scale pre-trained Direct3D over previous image-to-3D approaches, achieving significantly better generation quality and generalization ability, thus establishing a new state-of-the-art for 3D content creation. Project page: https://www.neural4d.com/research/direct3d. Youtian Lin, Yifei Zeng, Feihu Zhang, Jingxi Xu 0001, Philip Torr 0001, Xun Cao, Yao Yao 0008 |
NeurIPS | 7 |
| 2024 | Gaseous Object DetectionabstractObject detection, a fundamental and challenging problem in computer vision, has experienced rapid development due to the effectiveness of deep learning. The current objects to be detected are mostly rigid solid substances with apparent and distinct visual characteristics. In this paper, we endeavor on a scarcely explored task named Gaseous Object Detection (GOD), which is undertaken to explore whether the object detection techniques can be extended from solid substances to gaseous substances. Nevertheless, the gas exhibits significantly different visual characteristics: 1) saliency deficiency, 2) arbitrary and ever-changing shapes, 3) lack of distinct boundaries. To facilitate the study on this challenging task, we construct a GOD-Video dataset comprising 600 videos (141,017 frames) that cover various attributes with multiple types of gases. A comprehensive benchmark is established based on this dataset, allowing for a rigorous evaluation of frame-level and video-level detectors. Deduced from the Gaussian dispersion model, the physics-inspired Voxel Shift Field (VSF) is designed to model geometric irregularities and ever-changing shapes in potential 3D space. By integrating VSF into Faster RCNN, the VSF RCNN serves as a simple but strong baseline for gaseous object detection. Our work aims to attract further research into this valuable albeit challenging area. Kailai Zhou, Yibo Wang 0004, Qiu Shen, Xun Cao |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Disorder-Invariant Implicit Neural RepresentationabstractImplicit neural representation (INR) characterizes the attributes of a signal as a function of corresponding coordinates which emerges as a sharp weapon for solving inverse problems. However, the expressive power of INR is limited by the spectral bias in the network training. In this paper, we find that such a frequency-related problem could be greatly solved by re-arranging the coordinates of the input signal, for which we propose the disorder-invariant implicit neural representation (DINER) by augmenting a hash-table to a traditional INR backbone. Given discrete signals sharing the same histogram of attributes and different arrangement orders, the hash-table could project the coordinates into the same distribution for which the mapped signal can be better modeled using the subsequent INR network, leading to significantly alleviated spectral bias. Furthermore, the expressive power of the DINER is determined by the width of the hash-table. Different width corresponds to different geometrical elements in the attribute space, e.g., 1D curve, 2D curved-plane and 3D curved-volume when the width is set as 1, 2 and 3, respectively. More covered areas of the geometrical elements result in stronger expressive power. Experiments not only reveal the generalization of the DINER for different INR backbones (MLP versus SIREN) and various tasks (image/video representation, phase retrieval, refractive index recovery, and neural radiance field optimization) but also show the superiority over the state-of-the-art algorithms both in quality and speed. Hao Zhu 0005, Shaowen Xie, Zhen Liu 0031, Qi Zhang 0029, Zhan Ma 0001, Xun Cao |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2024 | Exploring Video Denoising in Thermal Infrared Imaging: Physics-Inspired Noise Generator, Dataset, and ModelabstractWe endeavor on a rarely explored task named thermal infrared video denoising. Perception in the thermal infrared significantly enhances the capabilities of machine vision. Nonetheless, noise in imaging systems is one of the factors that hampers the large-scale application of equipment. Existing thermal infrared denoising methods, primarily focusing on the image level, inadequately utilize time-domain information and insufficiently conduct investigation of system-level mixed noise, presenting the inferior ability in the video-recorded era; while video denoising methods, commonly applied to RGB cameras, exhibit uncertain effectiveness owing to substantial dissimilarities in the noise models and modalities between RGB and thermal infrared images. In sight of this, we initially revisit the imaging mechanism, while concurrently introducing a physics-inspired noise generator based on the sources and characteristics of system noise. Subsequently, a thermal infrared video denoising dataset consisting of 518 real-world videos is constructed. Lastly, we propose a denoising model called multi-domain infrared video denoising network, capable of concentrating features from the time, space, and frequency domains to restore high-fidelity videos. Extensive experiments demonstrate that the proposed method achieves state-of-the-art denoising quality and can be successfully applied to commercial cameras and downstream vision tasks, providing a new avenue for clear videography in the thermal infrared world. The dataset and code will be available. Lijing Cai, Kailai Zhou, Xun Cao |
IEEE Trans. Image Process. | 4 |
| 2024 | Continuous 3D Myocardial Motion Tracking via EchocardiographyabstractMyocardial motion tracking stands as an essential clinical tool in the prevention and detection of cardiovascular diseases (CVDs), the foremost cause of death globally. However, current techniques suffer from incomplete and inaccurate motion estimation of the myocardium in both spatial and temporal dimensions, hindering the early identification of myocardial dysfunction. To address these challenges, this paper introduces the Neural Cardiac Motion Field (NeuralCMF). NeuralCMF leverages implicit neural representation (INR) to model the 3D structure and the comprehensive 6D forward/backward motion of the heart. This method surpasses pixel-wise limitations by offering the capability to continuously query the precise shape and motion of the myocardium at any specific point throughout the cardiac cycle, enhancing the detailed analysis of cardiac dynamics beyond traditional speckle tracking. Notably, NeuralCMF operates without the need for paired datasets, and its optimization is self-supervised through the physics knowledge priors in both space and time dimensions, ensuring compatibility with both 2D and 3D echocardiogram video inputs. Experimental validations across three representative datasets support the robustness and innovative nature of the NeuralCMF, marking significant advantages over existing state-of-the-art methods in cardiac imaging and motion tracking. Code is available at: https://njuvision.github.io/NeuralCMF. Chengkang Shen, Hao Zhu 0005, Si Yi, Weipeng Zhao, David J. Brady, Xun Cao, Zhan Ma 0001 |
IEEE Trans. Medical Imaging | 9 |
| 2023 | RAFaRe: Learning Robust and Accurate Non-parametric 3D Face Reconstruction from Pseudo 2D&3D PairsabstractWe propose a robust and accurate non-parametric method for single-view 3D face reconstruction (SVFR). While tremendous efforts have been devoted to parametric SVFR, a visible gap still lies between the result 3D shape and the ground truth. We believe there are two major obstacles: 1) the representation of the parametric model is limited to a certain face database; 2) 2D images and 3D shapes in the fitted datasets are distinctly misaligned. To resolve these issues, a large-scale pseudo 2D&3D dataset is created by first rendering the detailed 3D faces, then swapping the face in the wild images with the rendered face. These pseudo 2D&3D pairs are created from publicly available datasets which eliminate the gaps between 2D and 3D data while covering diverse appearances, poses, scenes, and illumination. We further propose a non-parametric scheme to learn a well-generalized SVFR model from the created dataset, and the proposed hierarchical signed distance function turns out to be effective in predicting middle-scale and small-scale 3D facial geometry. Our model outperforms previous methods on FaceScape-wild/lab and MICC benchmarks and is well generalized to various appearances, poses, expressions, and in-the-wild environments. The code is released at https://github.com/zhuhao-nju/rafare. Longwei Guo, Hao Zhu 0004, Yuanxun Lu, Menghua Wu, Xun Cao |
AAAI | 5 |
| 2023 | High-fidelity 3D Face Generation from Natural Language DescriptionsabstractSynthesizing high-quality 3D face models from natural language descriptions is very valuable for many applications, including avatar creation, virtual reality, and telepresence. However, little research ever tapped into this task. We argue the major obstacle lies in 1) the lack of high-quality 3D face data with descriptive text annotation, and 2) the complex mapping relationship between descriptive language space and shape/appearance space. To solve these problems, we build Describe3D dataset, the first large-scale dataset with fine-grained text descriptions for text-to-3D face generation task. Then we propose a two-stage framework to first generate a 3D face that matches the concrete descriptions, then optimize the parameters in the 3D shape and texture space with abstract description to refine the 3D face model. Extensive experimental results show that our method can produce a faithful 3D face that conforms to the input descriptions with higher accuracy and quality than previous methods. The code and Describe3D dataset are released at https://github.com/zhuhao-nju/describe3D. Menghua Wu, Hao Zhu 0004, Linjia Huang, Yiyu Zhuang, Yuanxun Lu, Xun Cao |
CVPR | 6 |
| 2023 | DINER: Disorder-Invariant Implicit Neural RepresentationabstractImplicit neural representation (INR) characterizes the attributes of a signal as a function of corresponding coordinates which emerges as a sharp weapon for solving inverse problems. However, the capacity of INR is limited by the spectral bias in the network training. In this paper, we find that such a frequency-related problem could be largely solved by re-arranging the coordinates of the input signal, for which we propose the disorder-invariant implicit neural representation (DINER) by augmenting a hash-table to a traditional INR backbone. Given discrete signals sharing the same histogram of attributes and different arrangement orders, the hash-table could project the coordinates into the same distribution for which the mapped signal can be better modeled using the subsequent INR network, leading to significantly alleviated spectral bias. Experiments not only reveal the generalization of the DINER for different INR backbones (MLP vs. SIREN) and various tasks (image/video representation, phase retrieval, and refractive index recovery) but also show the superiority over the state-of-the-art algorithms both in quality and speed. Project page: https://ezio77.github.io/DINER-website/ Shaowen Xie, Hao Zhu 0004, Zhen Liu 0031, Qi Zhang 0029, Xun Cao, Zhan Ma 0001 |
CVPR | 6 |
| 2023 | Compact Self-adaptive Coding for Spectral Compressive SensingabstractSpectral snapshot compressive imaging (SCI) has been extensively studied and applied to various fields. Although the typical coded aperture snapshot spectral imaging (CASSI) presents an effective paradigm, its fixed code designs do not sufficiently exploit flexible and optimal modulation in consideration of scene sparsity. In this paper, we present a novel Compact Self-adaptive optical Coding framework for Spectral Compressive Sensing, termed 3CS, to optimize the coded pattern adaptively for better hyper-spectral videos perception. Our framework enables extracting context high-frequency components from the compressed domain without requiring hybrid guiding camera. The specifically designed mask distribution enables higher light efficiency, and is robust against temporal correlation reduction when processing dynamic spectral videos. Extensive experiments and model discussions validate the superiority of the proposed framework over traditional end-to-end (E2E) methods in various aspects for the spectral reconstruction. Yibo Wang 0004, Xun Cao |
ICCP | 5 |
| 2023 | Aperture Diffraction for Compact Snapshot Spectral ImagingabstractWe demonstrate a compact, cost-effective snapshot spectral imaging system named Aperture Diffraction Imaging Spectrometer (ADIS), which consists only of an imaging lens with an ultra-thin orthogonal aperture mask and a mosaic filter sensor, requiring no additional physical footprint compared to common RGB cameras. Then we introduce a new optical design that each point in the object space is multiplexed to discrete encoding locations on the mosaic filter sensor by diffraction-based spatial-spectral projection engineering generated from the orthogonal mask. The orthogonal projection is uniformly accepted to obtain a weakly calibration-dependent data form to enhance modulation robustness. Meanwhile, the Cascade Shift-Shuffle Spectral Transformer (CSST) with strong perception of the diffraction degeneration is designed to solve a sparsity-constrained inverse problem, realizing the volume reconstruction from 2D measurements with Large amount of aliasing. Our system is evaluated by elaborating the imaging optical theory and reconstruction algorithm with demonstrating the experimental imaging under a single exposure. Ultimately, we achieve the sub-super-pixel spatial resolution and high spectral resolution imaging. The code will be available at: https://github.com/Krito-ex/CSST. Yibo Wang 0004, Shuming Wang, Xun Cao |
ICCV | 7 |
| 2023 | Anableps: Adapting Bitrate for Real-Time Communication Using VBR-encoded VideoabstractContent providers increasingly replace traditional constant bitrate with variable bitrate (VBR) encoding in real-time video communication systems for better video quality. However, VBR encoding often leads to large and frequent bitrate fluctuation, inevitably deteriorating the efficiency of existing adaptive bitrate (ABR) methods. To tackle it, we propose the Anableps to consider the network dynamics and VBR-encoding-induced video bitrate fluctuations jointly for deploying the best ABR policy. With this aim, Anableps uses sender-side information from the past to predict the video bitrate range of upcoming frames. Such bitrate range is then combined with the receiver-side observations to set the proper bitrate target for video encoding using a reinforcement-learning-based ABR model. As revealed by extensive experiments on a real-world trace-driven testbed, our Anableps outperforms the de facto GCC with significant improvement of quality of experience, e.g., 1.88× video quality, 57% less bitrate consumption, 85% less stalling, and 74% shorter interaction delay. Hao Chen 0036, Xun Cao, Zhan Ma 0001 |
ICME | 3 |
| 2023 | WormTrack: Dataset and Benchmark for Multi-Object Tracking in Worm CrowdsabstractCurrently, multimedia systems and computer vision algorithms are increasingly playing a crucial role in biological research. However, due to the significant difference between macro and micro scenarios, it is impractical to directly transfer existing computer vision methods to the images captured by microscopes. Taking social behavior analysis of worm for example, it heavily depends on accurate and efficient Multi-object tracking (MOT) methods. Meanwhile, it faces great challenges due to the unique physical characteristics of worm, such as small size, highly uniform appearance, rapid deformation and overlapping movement. This paper studies on the challenges and existing solutions for MOT in worm crowds by building a well-designed dataset ("WormTrack") and a tracking-by-detection benchmark. We observed that the state-of-the-art MOT methods suffers from considerable performance drop on the new dataset. Therefore, we propose a customized MOT method for worm crowds by deeply understanding the physical characteristics of worms and scenes. The method is composed by an instance segmentation based detector, a multiple model fused Kalman filter based tracker and a multi-constraint based trajectory repairer. The experimental results demonstrate that our method can accurately track over 100 worms with almost identical appearance for a long period, which is exceptional compared to existing methods. We hope our work will attract further researches to explore more in this new field, and promote the crossing field researches with biology and medicine. Our code and data is available at https://github.com/Jeerrzy/wormstudio. Zhiyu Jin, Hanyang Yu, Chen Haul, Linxiang Wang, Zuobin Zhu, Qiu Shen, Xun Cao |
ACM Multimedia | 7 |
| 2023 | Anti-Aliased Neural Implicit Surfaces with Encoding Level of DetailabstractWe present LoD-NeuS, an efficient neural representation for high-frequency geometry detail recovery and anti-aliased novel view rendering. Drawing inspiration from voxel-based representations with the level of detail (LoD), we introduce a multi-scale tri-plane-based scene representation that is capable of capturing the LoD of the signed distance function (SDF) and the space radiance. Our representation aggregates space features from a multi-convolved featurization within a conical frustum along a ray and optimizes the LoD feature volume through differentiable rendering. Additionally, we propose an error-guided sampling strategy to guide the growth of the SDF during the optimization. Both qualitative and quantitative evaluations demonstrate that our method achieves superior surface reconstruction and photorealistic view synthesis compared to state-of-the-art approaches. Yiyu Zhuang, Qi Zhang 0029, Hao Zhu 0004, Yao Yao 0008, Xiaoyu Li 0002, Yan-Pei Cao 0001, Ying Shan, Xun Cao |
SIGGRAPH Asia | 9 |
| 2023 | Pyramid NeRF: Frequency Guided Fast Radiance Field Optimization
Junyu Zhu, Hao Zhu 0004, Qi Zhang 0029, Zhan Ma 0001, Xun Cao |
Int. J. Comput. Vis. | 6 |
| 2023 | FaceScape: 3D Facial Dataset and Benchmark for Single-View 3D Face ReconstructionabstractIn this article, we present a large-scale detailed 3D face dataset, FaceScape, and the corresponding benchmark to evaluate single-view facial 3D reconstruction. By training on FaceScape data, a novel algorithm is proposed to predict elaborate riggable 3D face models from a single image input. FaceScape dataset releases 16,940 textured 3D faces, captured from 847 subjects and each with 20 specific expressions. The 3D models contain the pore-level facial geometry that is also processed to be topologically uniform. These fine 3D facial models can be represented as a 3D morphable model for coarse shapes and displacement maps for detailed geometry. Taking advantage of the large-scale and high-accuracy dataset, a novel algorithm is further proposed to learn the expression-specific dynamic details using a deep neural network. The learned relationship serves as the foundation of our 3D face prediction system from a single image input. Different from most previous methods, our predicted 3D models are riggable with highly detailed geometry under different expressions. We also use FaceScape data to generate the in-the-wild and in-the-lab benchmark to evaluate recent methods of single-view face reconstruction. The accuracy is reported and analyzed on the dimensions of camera pose and focal length, which provides a faithful and comprehensive evaluation and reveals new challenges. The unprecedented dataset, benchmark, and code have been released to the public for research purpose. Hao Zhu 0004, Longwei Guo, Mingkai Huang, Menghua Wu, Qiu Shen, Ruigang Yang, Xun Cao |
IEEE Trans. Pattern Anal. Mach. Intell. | 10 |
| 2022 | Detailed Facial Geometry Recovery from Multi-View Images by Learning an Implicit FunctionabstractRecovering detailed facial geometry from a set of calibrated multi-view images is valuable for its wide range of applications. Traditional multi-view stereo (MVS) methods adopt an optimization-based scheme to regularize the matching cost. Recently, learning-based methods integrate all these into an end-to-end neural network and show superiority of efficiency. In this paper, we propose a novel architecture to recover extremely detailed 3D faces within dozens of seconds. Unlike previous learning-based methods that regularize the cost volume via 3D CNN, we propose to learn an implicit function for regressing the matching cost. By fitting a 3D morphable model from multi-view images, the features of multiple images are extracted and aggregated in the mesh-attached UV space, which makes the implicit function more effective in recovering detailed facial shape. Our method outperforms SOTA learning-based MVS in accuracy by a large margin on the FaceScape dataset. The code and data are released in https://github.com/zhuhao-nju/mvfr. Yunze Xiao, Hao Zhu 0004, Zhengyu Diao, Xiangju Lu, Xun Cao |
AAAI | 6 |
| 2022 | Explore Spatio-temporal Aggregation for Insubstantial Object Detection: Benchmark Dataset and BaselineabstractWe endeavor on a rarely explored task named Insubstantial Object Detection (IOD), which aims to localize the object with following characteristics: (1) amorphous shape with indistinct boundary; (2) similarity to surroundings; (3) absence in color. Accordingly, it is far more challenging to distinguish insubstantial objects in a single static frame and the collaborative representation of spatial and temporal information is crucial. Thus, we construct an IOD-Video dataset comprised of 600 videos (141,017 frames) covering various distances, sizes, visibility, and scenes captured by different spectral ranges. In addition, we develop a spatio-temporal aggregation framework for IOD, in which different backbones are deployed and a spatio-temporal aggregation loss (STAloss) is elaborately designed to leverage the consistency along the time axis. Experiments conducted on IOD-Video dataset demonstrate that spatio-temporal aggregation can significantly improve the performance of IOD. We hope our work will attract further researches into this valuable yet challenging task. The code will be available at: https://github.com/CalayZhou/IOD-Video. Kailai Zhou, Yibo Wang 0004, Yunqian Li, Linsen Chen, Qiu Shen, Xun Cao |
CVPR | 7 |
| 2022 | MoFaNeRF: Morphable Facial Neural Radiance Field
Yiyu Zhuang, Hao Zhu 0004, Xusen Sun, Xun Cao |
ECCV (3) | 4 |
| 2022 | Detailed Avatar Recovery From Single ImageabstractThis paper presents a novel framework to recover detailed avatar from a single image. It is a challenging task due to factors such as variations in human shapes, body poses, texture, and viewpoints. Prior methods typically attempt to recover the human body shape using a parametric-based template that lacks the surface details. As such resulting body shape appears to be without clothing. In this paper, we propose a novel learning-based framework that combines the robustness of the parametric model with the flexibility of free-form 3D deformation. We use the deep neural networks to refine the 3D shape in a Hierarchical Mesh Deformation (HMD) framework, utilizing the constraints from body joints, silhouettes, and per-pixel shading information. Our method can restore detailed human body shapes with complete textures beyond skinned models. Experiments demonstrate that our method has outperformed previous state-of-the-art approaches, achieving better accuracy in terms of both 2D IoU number and 3D metric distance. Hao Zhu 0004, Xinxin Zuo, Sen Wang 0003, Xun Cao, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | End-to-End Neural Video Coding Using a Compound Spatiotemporal RepresentationabstractRecent years have witnessed rapid advances in learnt video coding. Most algorithms have solely relied on the vector-based motion representation and resampling (e.g., optical flow based bilinear sampling) for exploiting the inter frame redundancy. In spite of the great success of adaptive kernel-based resampling (e.g., adaptive convolutions and deformable convolutions) in video prediction for uncompressed videos, integrating such approaches with rate-distortion optimization for inter frame coding has been less successful. Recognizing that each resampling solution offers unique advantages in regions with different motion and texture characteristics, we propose a hybrid motion compensation (HMC) method that adaptively combines the predictions generated by these two approaches. Specifically, we generate a compound spatiotemporal representation (CSTR) through a recurrent information aggregation (RIA) module using information from the current and multiple past frames. We further design a one-to-many decoder pipeline to generate multiple predictions from the CSTR, including vector-based resampling, adaptive kernel-based resampling, compensation mode selection maps and texture enhancements, and combines them adaptively to achieve more accurate inter prediction. Experiments show that our proposed inter coding system can provide better motion-compensated prediction and is more robust to occlusions and complex motions. Together with jointly trained intra coder and residual coder, the overall learnt hybrid coder yields the state-of-the-art coding efficiency in low-delay scenario, compared to the traditional H.264/AVC and H.265/HEVC, as well as recently published learning-based methods, in terms of both PSNR and MS-SSIM metrics. Ming Lu 0003, Zhiqi Chen 0001, Xun Cao, Zhan Ma 0001, Yao Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Audio-Driven Emotional Video PortraitsabstractDespite previous success in generating audio-driven talking heads, most of the previous studies focus on the correlation between speech content and the mouth shape. Facial emotion, which is one of the most important features on natural human faces, is always neglected in their methods. In this work, we present Emotional Video Portraits (EVP), a system for synthesizing high-quality video portraits with vivid emotional dynamics driven by audios. Specifically, we propose the Cross-Reconstructed Emotion Disentanglement technique to decompose speech into two decoupled spaces, i.e., a duration-independent emotion space and a duration- dependent content space. With the disentangled features, dynamic 2D emotional facial landmarks can be deduced. Then we propose the Target-Adaptive Face Synthesis technique to generate the final high-quality video portraits, by bridging the gap between the deduced landmarks and the natural head poses of target videos. Extensive experiments demonstrate the effectiveness of our method both qualitatively and quantitatively.1 Xinya Ji, Hang Zhou 0009, Kaisiyuan Wang, Wayne Wu, Chen Change Loy, Xun Cao, Feng Xu 0005 |
CVPR | 6 |
| 2021 | Neural Video Coding Using Multiscale Motion Compensation and Spatiotemporal Context ModelabstractOver the past two decades, traditional block-based video coding has made remarkable progress and spawned a series of well-known standards such as MPEG-4, H.264/AVC and H.265/HEVC. On the other hand, deep neural networks (DNNs) have shown their powerful capacity for visual content understanding, feature extraction and compact representation. Some previous works have explored the learnt video coding algorithms in an end-to-end manner, which show the great potential compared with traditional methods. In this paper, we propose an end-to-end deep neural video coding framework (NVC), which uses variational autoencoders (VAEs) with joint spatial and temporal prior aggregation (PA) to exploit the correlations in intra-frame pixels, inter-frame motions and inter-frame compensation residuals, respectively. Novel features of NVC include: 1) To estimate and compensate motion over a large range of magnitudes, we propose an unsupervised multiscale motion compensation network (MS-MCN) together with a pyramid decoder in the VAE for coding motion features that generates multiscale flow fields, 2) we design a novel adaptive spatiotemporal context model for efficient entropy coding for motion information, 3) we adopt nonlocal attention modules (NLAM) at the bottlenecks of the VAEs for implicit adaptive feature extraction and activation, leveraging its high transformation capacity and unequal weighting with joint global and local information, and 4) we introduce multi-module optimization and a multi-frame training strategy to minimize the temporal error propagation among P-frames. NVC is evaluated for the low-delay causal settings and compared with H.265/HEVC, H.264/AVC and the other learnt video compression methods following the common test conditions, demonstrating consistent gains across all popular test sequences for both PSNR and MS-SSIM distortion metrics. Ming Lu 0003, Zhan Ma 0001, Zhihuang Xie, Xun Cao, Yao Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | End-to-End Learnt Image Compression via Non-Local Attention Optimization and Improved Context ModelingabstractThis article proposes an end-to-end learnt lossy image compression approach, which is built on top of the deep nerual network (DNN)-based variational auto-encoder (VAE) structure with Non-Local Attention optimization and Improved Context modeling (NLAIC). Our NLAIC 1) embeds non-local network operations as non-linear transforms in both main and hyper coders for deriving respective latent features and hyperpriors by exploiting both local and global correlations, 2) applies attention mechanism to generate implicit masks that are used to weigh the features for adaptive bit allocation, and 3) implements the improved conditional entropy modeling of latent features using joint 3D convolutional neural network (CNN)-based autoregressive contexts and hyperpriors. Towards the practical application, additional enhancements are also introduced to speed up the computational processing (e.g., parallel 3D CNN-based context prediction), decrease the memory consumption (e.g., sparse non-local processing) and reduce the implementation complexity (e.g., a unified model for variable rates without re-training). The proposed model outperforms existing learnt and conventional (e.g., BPG, JPEG2000, JPEG) image compression methods, on both Kodak and Tecnick datasets with the state-of-the-art compression efficiency, for both PSNR and MS-SSIM quality measurements. We have made all materials publicly accessible at https://njuvision.github.io/NIC for reproducible research. Tong Chen 0004, Zhan Ma 0001, Qiu Shen, Xun Cao, Yao Wang 0001 |
IEEE Trans. Image Process. | 5 |
| 2021 | Live speech portraits: real-time photorealistic talking-head animationabstractTo the best of our knowledge, we first present a live system that generates personalized photorealistic talking-head animation only driven by audio signals at over 30 fps. Our system contains three stages. The first stage is a deep neural network that extracts deep audio features along with a manifold projection to project the features to the target person's speech space. In the second stage, we learn facial dynamics and motions from the projected audio features. The predicted motions include head poses and upper body motions, where the former is generated by an autoregressive probabilistic model which models the head pose distribution of the target person. Upper body motions are deduced from head poses. In the final stage, we generate conditional feature maps from previous predictions and send them with a candidate image set to an image-to-image translation network to synthesize photorealistic renderings. Our method generalizes well to wild audio and successfully synthesizes high-fidelity personalized facial details, e.g., wrinkles, teeth. Our method also allows explicit control of head poses. Extensive qualitative and quantitative evaluations, along with user studies, demonstrate the superiority of our method over state-of-the-art techniques. Yuanxun Lu, Jinxiang Chai, Xun Cao |
ACM Trans. Graph. | 3 |
| 2020 | FaceScape: A Large-Scale High Quality 3D Face Dataset and Detailed Riggable 3D Face PredictionabstractIn this paper, we present a large-scale detailed 3D face dataset, FaceScape, and propose a novel algorithm that is able to predict elaborate riggable 3D face models from a single image input. FaceScape dataset provides 18,760 textured 3D faces, captured from 938 subjects and each with 20 specific expressions. The 3D models contain the pore-level facial geometry that is also processed to be topologically uniformed. These fine 3D facial models can be represented as a 3D morphable model for rough shapes and displacement maps for detailed geometry. Taking advantage of the large-scale and high-accuracy dataset, a novel algorithm is further proposed to learn the expression-specific dynamic details using a deep neural network. The learned relationship serves as the foundation of our 3D face prediction system from a single image input. Different than the previous methods, our predicted 3D models are riggable with highly detailed geometry under different expressions. The unprecedented dataset and code will be released to public for research purpose. Hao Zhu 0004, Mingkai Huang, Qiu Shen, Ruigang Yang, Xun Cao |
CVPR | 7 |
| 2020 | Improving Multispectral Pedestrian Detection by Addressing Modality Imbalance Problems
Kailai Zhou, Linsen Chen, Xun Cao |
ECCV (18) | 3 |
| 2020 | Interactive free-viewpoint video generationabstractFree-viewpoint video (FVV) is processed video content in which viewers can freely select the viewing position and angle. FVV delivers an improved visual experience and can also help synthesize special effects and virtual reality content. In this paper, a complete FVV system is proposed to interactively control the viewpoints of video relay programs through multimedia terminals such as computers and tablets. The hardware of the FVV generation system is a set of synchronously controlled cameras, and the software generates videos in novel viewpoints from the captured video using view interpolation. The interactive interface is designed to visualize the generated video in novel viewpoints and enable the viewpoint to be changed interactively. Experiments show that our system can synthesize plausible videos in intermediate viewpoints with a view range of up to 180°. Hao Zhu 0004, Wei Li 0111, Xun Cao, Ruigang Yang |
Virtual Real. Intell. Hardw. | 5 |
| 2019 | Detailed Human Shape Estimation From a Single Image by Hierarchical Mesh DeformationabstractThis paper presents a novel framework to recover detailed human body shapes from a single image. It is a challenging task due to factors such as variations in human shapes, body poses, and viewpoints. Prior methods typically attempt to recover the human body shape using a parametric based template that lacks the surface details. As such the resulting body shape appears to be without clothing. In this paper, we propose a novel learning-based framework that combines the robustness of parametric model with the flexibility of free-form 3D deformation. We use the deep neural networks to refine the 3D shape in a Hierarchical Mesh Deformation (HMD) framework, utilizing the constraints from body joints, silhouettes, and per-pixel shading information. We are able to restore detailed human body shapes beyond skinned models. Experiments demonstrate that our method has outperformed previous state-of-the-art approaches, achieving better accuracy in terms of both 2D IoU number and 3D metric distance. The code is available in https://github.com/zhuhao-nju/hmd.git. Hao Zhu 0004, Xinxin Zuo, Sen Wang 0003, Xun Cao, Ruigang Yang |
CVPR | 4 |
| 2019 | Hyperspectral Imaging With Random Printed MaskabstractHyperspectral images can provide rich clues for various computer vision tasks. However, the requirements of professional and expensive hardware for capturing hyperspectral images impede its wide applications. In this paper, based on a simple but not widely noticed phenomenon that the color printer can print color masks with a large number of independent spectral transmission responses, we propose a simple and low-budget scheme to capture the hyperspectral images with a random mask printed by the consumer-level color printer. Specifically, we notice that the printed dots with different colors are stacked together, forming multiplicative, instead of additive, spectral transmission responses. Therefore, new spectral transmission response uncorrelated with that of the original printer dyes are generated. With the random printed color mask, hyperspectral images could be captured in a snapshot way. A convolutional neural network (CNN) based method is developed to reconstruct the hyperspectral images from the captured image. The effectiveness and accuracy of the proposed system are verified on both synthetic and real captured images. Zhan Ma 0001, Xun Cao, Tao Yue 0003 |
CVPR | 4 |
| 2019 | Spectral Reconstruction From Dispersive Blur: A Novel Light Efficient Spectral ImagerabstractDeveloping high light efficiency imaging techniques to retrieve high dimensional optical signal is a long-term goal in computational photography. Multispectral imaging, which captures images of different wavelengths and boosting the abilities for revealing scene properties, has developed rapidly in the last few decades. From scanning method to snapshot imaging, the limit of light collection efficiency is kept being pushed which enables wider applications especially under the light-starved scenes. In this work, we propose a novel multispectral imaging technique, that could capture the multispectral images with a high light efficiency. Through investigating the dispersive blur caused by spectral dispersers and introducing the difference of blur (DoB) constraints, we propose a basic theory for capturing multispectral information from a single dispersive-blurred image and an additional spectrum of an arbitrary point in the scene. Based on the theory, we design a prototype system and develop an optimization algorithm to realize snapshot multispectral imaging. The effectiveness of the proposed method is verified on both the synthetic data and real captured images. Zhan Ma 0001, Tao Yue 0003, Xun Cao |
CVPR | 6 |
| 2019 | Self-Adaptive Clustering and Load-Bandwidth Management for Uplink Enhancement in Heterogeneous Vehicular NetworksabstractDue to the diversity of traffic scenes and high mobility of vehicles, the propagation environment between moving vehicles and road side access points can be highly dynamic, which causes unstable uplink connectivity and time-varying uplink data rates in vehicular networks. In this paper, we study the uplink performance of heterogenous vehicular networks, which integrate dedicated short-range communications (DSRCs) and Long Term Evolution vehicle-to-everything (LTE V2X) into a single vehicular network to provide high reliability, low latency, and wide area coverage. Here, we adopt a cluster-based approach, in which DSRC and LTE are utilized to provide vehicle-to-vehicle communications between vehicles within a cluster and vehicle-to-infrastructure communications between vehicles and access points, respectively. Specifically, a self-adaptive clustering method is proposed based on the iterative self-organizing data analysis technique algorithm, in which the number of clusters can automatically adjust to the optimal value according to the mobility information. Also, a joint load-bandwidth management scheme is proposed to distribute traffic load and bandwidth resources between DSRC and LTE. Simulation results show that the proposed algorithm outperforms the traditional section-based and ${K}$ -means clustering methods, and a tradeoff between average uplink data rate and signaling overhead can be achieved. Tianyu Wang 0001, Xun Cao, Shaowei Wang 0001 |
IEEE Internet Things J. | 2 |
| 2018 | View Extrapolation of Human Body From a Single ImageabstractWe study how to synthesize novel views of human body from a single image. Though recent deep learning based methods work well for rigid objects, they often fail on objects with large articulation, like human bodies. The core step of existing methods is to fit a map from the observable views to novel views by CNNs; however, the rich articulation modes of human body make it rather challenging for CNNs to memorize and interpolate the data well. To address the problem, we propose a novel deep learning based pipeline that explicitly estimates and leverages the geometry of the underlying human body. Our new pipeline is a composition of a shape estimation network and an image generation network, and at the interface a perspective transformation is applied to generate a forward flow for pixel value transportation. Our design is able to factor out the space of data variation and makes learning at each step much easier. Empirically, we show that the performance for pose-varying objects can be improved dramatically. Our method can also be applied on real data captured by 3D sensors, and the flow generated by our methods can be used for generating high quality results in higher resolution. Hao Zhu 0004, Peng Wang 0001, Xun Cao, Ruigang Yang |
CVPR | 4 |
| 2018 | Multispectral Image Intrinsic Decomposition via Subspace ConstraintabstractMultispectral images contain many clues of surface characteristics of the objects, thus can be used in many computer vision tasks, e.g., recolorization and segmentation. However, due to the complex geometry structure of natural scenes, the spectra curves of the same surface can look very different under different illuminations and from different angles. In this paper, a new Multispectral Image Intrinsic Decomposition model (MIID) is presented to decompose the shading and reflectance from a single multispectral image. We extend the Retinex model, which is proposed for RGB image intrinsic decomposition, for multispectral domain. Based on this, a subspace constraint is introduced to both the shading and reflectance spectral space to reduce the ill-posedness of the problem and make the problem solvable. A dataset of 22 scenes is given with the ground truth of shadings and reflectance to facilitate objective evaluations. The experiments demonstrate the effectiveness of the proposed method. Weixin Zhu, Linsen Chen, Yao Wang 0001, Tao Yue 0003, Xun Cao |
CVPR | 7 |
| 2018 | Modeling the Perceptual Quality of Immersive Images Rendered on Head Mounted Displays: Resolution and CompressionabstractWe develop a model that expresses the joint impact of spatial resolution s and JPEG compression quality factor qf on immersive image quality. The model is expressed as the product of optimized exponential functions of these factors. The model is tested on a subjective database of immersive image contents rendered on a head mounted display (HMD). High Pearson correlation and Spearman correlation (> 0.95) and small relative root mean squared error (< 5.6%) are achieved between the model predictions and the subjective quality judgements. The immersive ground-truth images along with the rest of the database are made available for future research and comparisons. Mingkai Huang, Qiu Shen, Zhan Ma 0001, Alan C. Bovik, Praful Gupta, Rongbing Zhou, Xun Cao |
IEEE Trans. Image Process. | 7 |
| 2017 | Why Did They Do That?: Exploring Attribution Mismatches Between Native and Non-Native Speakers Using VideoconferencingabstractThe meaning we attribute to another's actions significantly impact our subsequent behaviors and interactions towards that person. Distributed teams often combine native speakers (NS) and non-native speakers (NNS) and are particularly prone to making attribution errors. Language difficulties place NNS under a higher cognitive load, potentially leading NS to make inaccurate attributions of NNS. We conducted an exploratory laboratory study to investigate the attributions NS and NNS form about each other in multiparty videoconferencing. Our findings revealed significant mismatches in NS' attributions of NNS behavior, but no significant mismatch in NNS' attributions of NS behavior. Due to cognitive overload stemming from language challenges, NNS were only able to engage in "compromised" impression management during the task. Yet, NS were relatively unaware of how profoundly language difficulties impacted NNS' behaviors. Our findings identify opportunities for technology support for NS-NNS interactions, particularly with regards to impression construction and impression management. Helen Ai He, Naomi Yamashita, Ari Hautasaari, Xun Cao, Elaine M. Huang |
CSCW | 4 |
| 2017 | Multispectral focal stack acquisition using a chromatic aberration enlarged cameraabstractCapturing more information, e.g. geometry and material, using optical cameras can greatly help the perception and understanding of complex scenes. This paper proposes a novel method to capture the spectral and light field information simultaneously. By using a delicately designed chromatic aberration enlarged camera, the spectral-varying slices at different depths of the scene can be easily captured. Afterwards, the multispectral focal stack, which is composed of a stack of multispectral slice images focusing on different depths, can be recovered from the spectral-varying slices by using a Local Linear Transformation (LLT) based algorithm. The experiments verify the effectiveness of the proposed method. Yunqian Li, Linsen Chen, Xiaoming Zhong, Jin-Li Suo, Zhan Ma 0001, Tao Yue 0003, Xun Cao |
ICIP | 8 |
| 2017 | DeepCoder: A deep neural network based video compressionabstractInspired by recent advances in deep learning, we present the DeepCoder - a Convolutional Neural Network (CNN) based video compression framework. We apply separate CNN nets for predictive and residual signals respectively. Scalar quantization and Huffman coding are employed to encode the quantized feature maps (fMaps) into binary stream. We use the fixed 32 × 32 block in this work to demonstrate our ideas, and performance comparison is conducted with the well-known H.264/AVC video coding standard with comparable rate-distortion performance. Here distortion is measured using Structural Similarity (SSIM) because it is more close to perceptual response. Tong Chen 0004, Qiu Shen, Tao Yue 0003, Xun Cao, Zhan Ma 0001 |
VCIP | 5 |
| 2017 | Modeling peripheral vision impact on perceptual quality of immersive imagesabstractConventional images/videos are often rendered within the central area of human visual system (HVS) with uniform quality. Recent virtual reality (VR) device with head mounted display (HMD) extends the field of view (FoV) significantly to include both central and peripheral areas. It exhibits the unequal acuity of the quality sensation because of the non-uniform distribution of photoreceptors in our retina. Hence, we propose to study the impact of image qualities (with respect to the quantization stepsize q or spatial resolution s) in peripheral vision and conclude self-adaptive analytical models that have shown quite impressive accuracy through independent cross validations. These models can further be applied to assign different quality weights at different regions, so as to significantly reduce the transmission data size but without subjective quality loss. Peiyao Guo, Qiu Shen, Mingkai Huang, Rongbing Zhou, Xun Cao, Zhan Ma 0001 |
VCIP | 5 |
| 2017 | A Practical System Towards the Secure, Robust and Pervasive Mobile WorkstyleabstractWe develop an innovative PC2PC (personal computer to pervasive computing) system to enable the secure, robust and pervasive mobile workstyle. PC2PC server compresses the desktop screens of any virtualized system, and delivers the stream through any popular networks to PC2PC client remotely for stream decoding, rendering and end-user interaction (such as keyboard/mouse commands). We have implemented the overall system from the scratch, where the emerging screen content coding (SCC) extension of the High-Efficiency Video Coding (HEVC) is implemented to compress and stream the desktop screens in real-time, and three core asset channels (i.e., system, display, inputs, etc) are defined to enable systematic end-to-end communication. Compared with the commercial Red Hat SPICE virtual desktop infrastructure (VDI) scheme, our PC2PC could save the network bandwidth by a factor of 2, 7 and 4 respectively for typical video streaming, web browsing and stationary office applications at same visual quality. Meanwhile, we have also measured the delays in the system and presented the preliminary study on the user experience impact. A simple network estimation is applied to optimize the quality-bandwidth adaptation for both single user and multiuser scenarios to combat the network dynamics. Zhan Ma 0001, Tao Yue 0003, Xun Cao, Yiling Xu, Xin Li 0106, Yongjin Wang |
VTC Spring | 3 |
| 2017 | The role of prior in image based 3D modeling: a survey
Hao Zhu 0004, Yongming Nie, Tao Yue 0003, Xun Cao |
Frontiers Comput. Sci. | 4 |
| 2017 | Robust multi-view stereo synthesized by various parameters model
Yongming Nie, Tao Yue 0003, Hao Zhu 0004, Sidan Du, Xun Cao |
J. Vis. Commun. Image Represent. | 5 |
| 2017 | High-resolution spectral video acquisitionabstractCompared with conventional cameras, spectral imagers provide many more features in the spectral domain. They have been used in various fields such as material identification, remote sensing, precision agriculture, and surveillance. Traditional imaging spectrometers use generally scanning systems. They cannot meet the demands of dynamic scenarios. This limits the practical applications for spectral imaging. Recently, with the rapid development in computational photography theory and semiconductor techniques, spectral video acquisition has become feasible. This paper aims to offer a review of the state-of-the-art spectral imaging technologies, especially those capable of capturing spectral videos. Finally, we evaluate the performances of the existing spectral acquisition systems and discuss the trends for future work. Linsen Chen, Tao Yue 0003, Xun Cao, Zhan Ma 0001, David J. Brady |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2017 | Efficient Method for High-Quality Removal of Nonuniform Blur in the Wavelet DomainabstractThis paper presents a novel nonuniform deblurring approach, which defines the blur model and calculates regularized nonuniform deconvolution in the wavelet domain to achieve high efficiency and high accuracy simultaneously. Targeting high computation efficiency, we derive a wavelet-domain hierarchical blur model, which can be calculated efficiently by exploiting the sparsity property of natural images in the wavelet domain. Correspondingly, the blur model is incorporated into a multilayer framework and at each layer spatially varying step sizes are introduced to further accelerate the convergence of the algorithm. In addition to the efficiency advantages, the proposed approach deals with intensely nonuniform blur with high accuracy due to the intrinsic tight supportness of wavelet basis. We conduct a series of experiments and comparisons to validate the efficiency and effectiveness of our algorithm. Tao Yue 0003, Jin-Li Suo, Xun Cao, Qionghai Dai |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | Video-Based Outdoor Human ReconstructionabstractA human body scanning system of great practical convenience, which can be used in an outdoor environment, is proposed. The system uses only a single conventional video camera without the aid of special sensors or controlled illuminations. We leverage the structure from motion calibration results directly and improve the available video-based dense 3D reconstruction by integrating the surface smoothness constraints. The point cloud reinforcement is proposed to detect and adjust the conflict point data for the slender and shaky body parts. Combined with the silhouette adaptation, the proposed point cloud reinforcement achieves reasonable and plausible mesh reconstruction on these challenging parts. We further introduce the close-shot frames to refine the prereconstructed mesh model, leading to a colored watertight model. The overall system is approximate to automatic since only one or two times of painting brush interaction are required for robust and high-quality multiview image segmentation. The experiment results on various test sequences demonstrate the effectiveness and the robustness of the proposed method, even under very challenging scenarios when shaking body, varying illumination, and textureless regions occur. Hao Zhu 0004, Yebin Liu, Jingtao Fan, Qionghai Dai, Xun Cao |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2017 | Interactive Screen Video Streaming-Based Pervasive Mobile WorkstyleabstractIn this paper, we develop an interactive screen video streaming-based system to enable the ubiquitous mobile workstyle, which is referred to as personal computer to pervasive computing (PC2PC). The desktop screens of virtualized systems are compressed in the PC2PC servers and delivered to remote end users for stream decoding, rendering, and interactions. We have implemented a system from the scratch, where the emerging screen content coding extension of high-efficiency video coding is implemented to compress and stream the desktop screens of the virtualized system in real time. Three core asset channels, system, display, and inputs, are defined to enable systematic end-to-end communication. Compared with Red Hat SPICE virtual desktop infrastructure scheme, the proposed PC2PC could save network bandwidth consumption by a factor of 2, 7, and 4, respectively, in terms of typical video streaming, web browsing, and stationary office applications at the same visual quality. Meanwhile, we have also measured the delays of the system and presented preliminary results on the user experience aspect. A simple network estimation is applied to optimize the quality bandwidth adaptation for both single user and multiuser scenarios to consider the network dynamics. Zhan Ma 0001, Tao Yue 0003, Xun Cao, Yiling Xu, Xin Li 0106, Yongjin Wang |
IEEE Trans. Multim. | 3 |
| 2016 | Investigating the impact of automated transcripts on non-native speakers' listening comprehensionabstractReal-time transcripts generated by automatic speech recognition (ASR) technologies hold potential to facilitate non-native speakers’ (NNSs) listening comprehension. While introducing another modality (i.e., ASR transcripts) to NNSs provides supplemental information to understand speech, it also runs the risk of overwhelming them with excessive information. The aim of this paper is to understand the advantages and disadvantages of presenting ASR transcripts to NNSs and to study how such transcripts affect listening experiences. To explore these issues, we conducted a laboratory experiment with 20 NNSs who engaged in two listening tasks in different conditions: audio only and audio+ASR transcripts. In each condition, the participants described the comprehension problems they encountered while listening. From the analysis, we found that ASR transcripts helped NNSs solve certain problems (e.g., “do not recognize words they know”), but imperfect ASR transcripts (e.g., errors and no punctuation) sometimes confused them and even generated new problems. Furthermore, post-task interviews and gaze analysis of the participants revealed that NNSs did not have enough time to fully exploit the transcripts. For example, NNSs had difficulty shifting between multimodal contents. Based on our findings, we discuss the implications for designing better multimodal interfaces for NNSs. Xun Cao, Naomi Yamashita, Toru Ishida 0001 |
ICMI | 1 |
| 2016 | Hyperspectral Image Super-Resolution via Non-Negative Structured Sparse RepresentationabstractHyperspectral imaging has many applications from agriculture and astronomy to surveillance and mineralogy. However, it is often challenging to obtain high-resolution (HR) hyperspectral images using existing hyperspectral imaging techniques due to various hardware limitations. In this paper, we propose a new hyperspectral image super-resolution method from a low-resolution (LR) image and a HR reference image of the same scene. The estimation of the HR hyperspectral image is formulated as a joint estimation of the hyperspectral dictionary and the sparse codes based on the prior knowledge of the spatial-spectral sparsity of the hyperspectral image. The hyperspectral dictionary representing prototype reflectance spectra vectors of the scene is first learned from the input LR image. Specifically, an efficient non-negative dictionary learning algorithm using the block-coordinate descent optimization technique is proposed. Then, the sparse codes of the desired HR hyperspectral image with respect to learned hyperspectral basis are estimated from the pair of LR and HR reference images. To improve the accuracy of non-negative sparse coding, a clustering-based structured sparse coding method is proposed to exploit the spatial correlation among the learned sparse codes. The experimental results on both public datasets and real LR hypspectral images suggest that the proposed method substantially outperforms several existing HR hyperspectral image recovery techniques in the literature in terms of both objective quality metrics and computational efficiency. Weisheng Dong, Fazuo Fu, Guangming Shi, Xun Cao, Jinjian Wu, Xin Li 0005 |
IEEE Trans. Image Process. | 4 |
| 2015 | Blind optical aberration correction by exploring geometric and visual priorsabstractOptical aberration widely exists in optical imaging systems, especially in consumer-level cameras. In contrast to previous solutions using hardware compensation or pre-calibration, we propose a computational approach for blind aberration removal from a single image, by exploring various geometric and visual priors. The global rotational symmetry allows us to transform the non-uniform degeneration into several uniform ones by the proposed radial splitting and warping technique. Locally, two types of symmetry constraints, i.e. central symmetry and reflection symmetry are defined as geometric priors in central and surrounding regions, respectively. Furthermore, by investigating the visual artifacts of aberration degenerated images captured by consumer-level cameras, the non-uniform distribution of sharpness across color channels and the image lattice is exploited as visual priors, resulting in a novel strategy to utilize the guidance from the sharpest channel and local image regions to improve the overall performance and robustness. Extensive evaluation on both real and synthetic data suggests that the proposed method outperforms the state-of-the-art techniques. Tao Yue 0003, Jin-Li Suo, Jue Wang 0001, Xun Cao, Qionghai Dai |
CVPR | 4 |
| 2015 | Semantic Matching in APP SearchabstractPast years, with the growth of smart-phones and applications, APP market has become an important mobile internet portal. As an important function in application market, APP search gains lots of attentions.However, mismatch between queries and APP is the most critical problem in APP search because of less text within term matching search engine. In this talk, we describe a semantic matching architecture in APP search--which mining topics and tags in big data. It enriches query and APP representations with topics and tags to achieve semantic matching in search. Some challenge must be considered: 1) How to extract tag-APP relationship from large web text. 2) How to use machine learning technologies to process de-noising and computing confidence. 3) How to hybrid ranking apps retrieved by different matching method. These will be introduced in some of our related works and as examples to describe how semantic matching is used in Tencent MyApp, an application market which serving hundreds of millions of users. Juchao Zhuo, Zeqian Huang, Zhanhui Kang, Xun Cao, Mingzhi Li |
WSDM | 5 |
| 2015 | Toward Naturalistic 2D-to-3D ConversionabstractNatural scene statistics (NSSs) models have been developed that make it possible to impose useful perceptually relevant priors on the luminance, colors, and depth maps of natural scenes. We show that these models can be used to develop 3D content creation algorithms that can convert monocular 2D videos into statistically natural 3D-viewable videos. First, accurate depth information on key frames is obtained via human annotation. Then, both forward and backward motion vectors are estimated and compared to decide the initial depth values, and a compensation process is applied to further improve the depth initialization. Then, the luminance/chrominance and initial depth map are decomposed by a Gabor filter bank. Each subband of depth is modeled to produce a NSS prior term. The statistical color-depth priors are combined with the spatial smoothness constraint in the depth propagation target function as a prior regularizing term. The final depth map associated with each frame of the input 2D video is optimized by minimizing the target function over all subbands. In the end, stereoscopic frames are rendered from the color frames and their associated depth maps. We evaluated the quality of the generated 3D videos using both subjective and objective quality assessment methods. The experimental results obtained on various sequences show that the presented method outperforms several state-of-the-art 2D-to-3D conversion methods. Xun Cao, Ke Lu 0002, Qionghai Dai, Alan C. Bovik |
IEEE Trans. Image Process. | 2 |
| 2014 | Acquisition of High Spatial and Spectral Resolution Video with a Hybrid Camera System
Chenguang Ma, Xun Cao, Xin Tong 0001, Qionghai Dai, Stephen Lin 0001 |
Int. J. Comput. Vis. | 2 |
| 2012 | Iterative Feedback Estimation of Depth and Radiance from Defocused Images
Jin-Li Suo, Xun Cao, Qionghai Dai |
ACCV (4) | 3 |
| 2012 | 3D spatial reconstruction and communication from vision fieldabstractVision field describes the real world visual information by summarizing the seven-dimensional plenoptic function into three domains: view, light, time, from which a better understanding of previous 3D capture and reconstruction systems can be provided. In this paper, we first show how to reconstruct 3D spatial information from all the three attributes of the vision field, namely full-space vision field reconstruction. Then, based on Laplacian iterative geometry prediction, a 3D mesh coding algorithm with cascaded quantization is presented to facilitate the communication of the reconstructed 3D models from vision field. At last, experimental results of both the 3D spatial reconstruction and the 3D mesh coding are demonstrated. Xun Cao, Qifei Wang, Xiangyang Ji, Qionghai Dai |
ICASSP | 1 |
| 2012 | A novel method for 2D-to-3D video conversion using bi-directional motion estimationabstractIn this paper, we proposed a novel semi-automatic 2D-to-3D video conversion method. Our method requires just a few user-scribbles to generate depth maps for key frames and propagates these depth maps to non-key frames automatically. For key frames, foreground objects and the corresponding depth maps can be obtained by an interactive method. Then, both forward and backward motion vectors are estimated and compared to decide the depth propagation strategy. For pixels that failed the motion vectors comparison, a compensation process is adopted to refine their depth propagation results. Finally, stereoscopic pairs are generated by the warping method based on the original frames and associated depth maps. Our method is validated by both subjective and objective quality assessments. The experimental results show that our method outperforms several state-of-the-art 2D-to-3D video conversion methods. Zhenyao Li, Xun Cao, Qionghai Dai |
ICASSP | 2 |
| 2011 | High resolution multispectral video capture with a hybrid camera systemabstractWe present a new approach to capture video at high spatial and spectral resolutions using a hybrid camera system. Composed of an RGB video camera, a grayscale video camera and several optical elements, the hybrid camera system simultaneously records two video streams: an RGB video with high spatial resolution, and a multispectral video with low spatial resolution. After registration of the two video streams, our system propagates the multispectral information into the RGB video to produce a video with both high spectral and spatial resolution. This propagation between videos is guided by color similarity of pixels in the spectral domain, proximity in the spatial domain, and the consistent color of each scene point in the temporal domain. The propagation algorithm is designed for rapid computation to allow real-time video generation at the original frame rate, and can thus facilitate real-time video analysis tasks such as tracking and surveillance. Hardware implementation details and design tradeoffs are discussed. We evaluate the proposed system using both simulations with ground truth data and on real-world scenes. The utility of this high resolution multispectral video data is demonstrated in dynamic white balance adjustment and tracking. Xun Cao, Xin Tong 0001, Qionghai Dai, Stephen Lin 0001 |
CVPR | 1 |
| 2011 | Vision field capture for advanced 3DTV applicationsabstractThe seven-dimensional plenoptic function provides a full description of the visual information for the real world. In this paper, we present a novel concept called vision field, which simplifies the seven-dimensional plenoptic function into its three subspaces, namely, view, light, time. Based on this concept, we found that most previous 3D capture systems can be related to the vision field capture. This paper first gives a brief survey on the previous 3D capture systems, categorizes them from the vision field perspective. Then, we introduce a system which is able to capture the vision field. A Multi-View-Multi-Lighting (MVML) capture system is built to obtain the multiview images of the 3D scenes or objects under different steerable light conditions. Finally, we show how the vision field capture can be used for advanced 3DTV applications. Xun Cao, Yebin Liu, Xiangyang Ji, Qionghai Dai |
VCIP | 1 |
| 2011 | A Prism-Mask System for Multispectral Video AcquisitionabstractThis paper presents a prism-mask system for capturing multispectral videos. The system is composed of a triangular prism, a monochrome camera, and an occlusion mask. Incoming light beams from the scene are sampled by the occlusion mask, dispersed into their constituent spectra by the triangular prism, and then captured by the monochrome camera. Our system is capable of capturing frames with high spectral resolution at video rates. It also allows for different trade-offs between spectral and spatial resolution by adjusting the focal length of the camera. We demonstrate multispectral video acquisition with various spectral resolutions and spatial resolutions, as well as different frame rates. The effectiveness of our system is further evaluated with several applications, including human skin detection, physical material recognition, video segmentation, RGB video generation, and illumination identification. Xun Cao, Hao Du 0004, Xin Tong 0001, Qionghai Dai, Stephen Lin 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2010 | Multiview video depth estimation with spatial-temporal consistency
Mingjin Yang, Xun Cao, Qionghai Dai |
BMVC | 2 |
| 2010 | Vision field capturing and its applications in 3DTVabstract3D video capturing acquires the visual information in 3D manner, which possesses the first step of the entire 3DTV system chain before 3D coding, transmission and visualization. The 3D capturing plays an important role because precise 3D visual capturing will benefit the whole 3DTV system. During the past decades, various kinds of capturing system have been built for different applications such as FTV[1], 3DTV, 3D movie, etc. As the cost of sensors reduces in recent years, a lot of systems utilize multiple cameras to acquire visual information, which is called multiview capturing. 3D information can be further extracted through multiview geometry. We will first give a brief review of these multiview systems and analyze their relationship from the perspective of plenoptic function [2]. Along with the multiple cameras, a lot of systems also make use of multiple lights to control the illumination condition. A new concept of vision field is presented in this talk according to the view-light-time subspace, which can be derived from the plenoptic function. The features and applications for each capturing system will be emphasized as well as the important issues in capturing like synchronization and calibration. Besides the multiple camera systems, some new techniques using TOF (time-off-light) camera [3] and 3D scanner will also be included in this talk. Qionghai Dai, Xiangyang Ji, Xun Cao |
PCS | 3 |
| 2009 | Continuous depth estimation for multi-view stereoabstractDepth-map merging approaches have become more and more popular in multi-view stereo (MVS) because of their flexibility and superior performance. The quality of depth map used for merging is vital for accurate 3D reconstruction. While traditional depth map estimation has been performed in a discrete manner, we suggest the use of a continuous counterpart. In this paper, we first integrate silhouette information and epipolar constraint into the variational method for continuous depth map estimation. Then, several depth candidates are generated based on a multiple starting scales (MSS) framework. From these candidates, refined depth maps for each view are synthesized according to path-based NCC (normalized cross correlation) metric. Finally, the multiview depth maps are merged to produce 3D models. Our algorithm excels at detail capture and produces one of the most accurate results among the current algorithms for sparse MVS datasets according to the Middlebury benchmark. Additionally, our approach shows its outstanding robustness and accuracy in free-viewpoint video scenario. Yebin Liu, Xun Cao, Qionghai Dai, Wenli Xu |
CVPR | 2 |
| 2009 | A prism-based system for multispectral video acquisitionabstractIn this paper, we propose a prism-based system for capturing multispectral videos. The system consists of a triangular prism, a monochrome camera, and an occlusion mask. Incoming light beams from the scene are sampled by the occlusion mask, dispersed into their constituent spectra by the triangular prism, and then captured by the monochrome camera. Our system is capable of capturing videos of high spectral resolution. It also allows for different tradeoffs between spectral and spatial resolution by adjusting the focal length of the camera. We demonstrate the effectiveness of our system with several applications, including human skin detection, physical material recognition, and RGB video generation. Hao Du 0004, Xin Tong 0001, Xun Cao, Stephen Lin 0001 |
ICCV | 3 |
| 2007 | All-Clear Image Based Synthesis using Clarity DegreeabstractImage based rendering (IBR) usually produces severe artifacts, typically blur and ghost, for objects not located on the focal plane. To obtain an all clear result, previous works either endeavor to recover scene depth or rest on iterations. These methods are difficult and time-consuming, not suitable for real-time IBR applications. In this paper, we propose a novel all-clear synthesis scheme bypassing tedious depth estimation. Our algorithm directly measures the clarity degree of local regions in pre-rendered images and then optimal blocks are combined. This measurement is motivated by the observation that mean change energies (MCE) are well consistent with the human vision system's feeling of "clarity". Furthermore, an efficient pseudo-depth filtering is proposed to alleviate the block effect during combination. Our algorithm gives outstanding performances on both synthetic and real world data. It is also proved to be fast enough for real-time implementation. Xun Cao, Su Xue, Qionghai Dai |
ICASSP (1) | 1 |