VLDB 2026 Research / reviewers in the wild / expert
Qiu Shen
dblp:85/8116
· DBLP profile ↗
39ranked-venue papers
2as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 1 first-author · 22 since 2021Artificial intelligence and machine learning · 15 · 1 first-author · 13 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Split-Layer: Enhancing Implicit Neural Representation by Maximizing the Dimensionality of Feature SpaceabstractImplicit neural representation (INR) models signals as continuous functions using neural networks, offering efficient and differentiable optimization for inverse problems across diverse disciplines. However, the representational capacity of INR—defined by the range of functions the neural network can characterize—is inherently limited by the low-dimensional feature space in conventional multilayer perceptron (MLP) architectures. While widening the MLP can linearly increase feature space dimensionality, it also leads to a quadratic growth in computational and memory costs. To address this limitation, we propose the split-layer, a novel reformulation of MLP construction. The split-layer divides each layer into multiple parallel branches and integrates their outputs via Hadamard product, effectively constructing a high-degree polynomial space. This approach significantly enhances INR’s representational capacity by expanding the feature space dimensionality without incurring prohibitive computational overhead. Extensive experiments demonstrate that the split-layer substantially improves INR performance, surpassing existing methods across multiple tasks, including 2D image fitting, 2D CT reconstruction, 3D shape representation, and 5D novel view synthesis. Zhicheng Cai, Hao Zhu 0004, Linsen Chen, Qiu Shen, Xun Cao |
AAAI | 4 |
| 2026 | Stability optimization in action imitation for humanoid robot
Shenghao Ren, Zhiyu Jin, Qiu Shen |
J. Vis. Commun. Image Represent. | 4 |
| 2026 | Toward the Spectral Bias Alleviation by Normalizations in Coordinate NetworksabstractRepresenting signals using coordinate networks dominates the area of inverse problems recently, and is widely applied in various scientific computing tasks. Still, there exists an issue of spectral bias in coordinate networks, limiting the capacity to learn high-frequency components. This problem is caused by the pathological distribution of the neural tangent kernel's (NTK's) eigenvalues of coordinate networks. We find that, this pathological distribution could be improved using classical normalization techniques (batch normalization and layer normalization), which are commonly used in convolutional neural networks but rarely used in coordinate networks. We prove that normalization techniques greatly reduces the maximum and variance of NTK's eigenvalues while slightly modifies the mean value, considering the max eigenvalue is much larger than the most, this variance change results in a shift of eigenvalues' distribution from a lower one to a higher one, therefore the spectral bias could be alleviated (see Fig. 1). Furthermore, we propose two new normalization techniques by combining these two techniques in different ways. The efficacy of these normalization techniques is substantiated by the significant improvements and new state-of-the-arts achieved by applying normalization-based coordinate networks to various tasks, including the image compression, computed tomography reconstruction, shape representation, magnetic resonance imaging, novel view synthesis and multi-view stereo reconstruction. Zhicheng Cai, Hao Zhu 0005, Qiu Shen, Xun Cao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | MotionPRO: Exploring the Role of Pressure in Human MoCap and BeyondabstractExisting human Motion Capture (MoCap) methods mostly focus on the visual similarity while neglecting the physical plausibility. As a result, downstream tasks such as driving virtual human in 3D scene or humanoid robots in real world suffer from issues such as timing drift and jitter, spatial problems like sliding and penetration, and poor global trajectory accuracy. In this paper, we revisit human MoCap from the perspective of interaction between human body and physical world by exploring the role of pressure. Firstly, we construct a large-scale human Motion capture dataset with Pressure, RGB and Optical sensors (named MotionPRO), which comprises 70 volunteers performing 400 types of motion, encompassing a total of 12.4M pose frames. Secondly, we examine both the necessity and effectiveness of the pressure signal through two challenging tasks: (1) pose and trajectory estimation based solely on pressure: We propose a network that incorporates a small kernel decoder and a long-short-term attention module, and proof that pressure could provide accurate global trajectory and plausible lower body pose. (2) pose and trajectory estimation by fusing pressure and RGB: We impose constraints on orthographic similarity along the camera axis and whole-body contact along the vertical axis to enhance the cross-attention strategy to fuse pressure and RGB feature maps. Experiments demonstrate that fusing pressure with RGB features not only significantly improves performance in terms of objective metrics, but also plausibly drives virtual humans (SMPL) in 3D scene. Furthermore, we demonstrate that incorporating physical perception enables humanoid robots to perform more precise and stable actions, which is highly beneficial for the development of embodied artificial intelligence. Project page is available at: https://nju-cite-mocaphumanoid.github.io/MotionPRO/ Shenghao Ren, Qiu Shen, Xun Cao |
CVPR | 7 |
| 2025 | M-SpecGene: Generalized Foundation Model for RGBT Multispectral Vision
Kailai Zhou, Fuqiang Yang, Shixian Wang, Bihan Wen, Chongde Zi, Linsen Chen, Qiu Shen, Xun Cao |
ICCV | 7 |
| 2025 | Multi-worm Tracking with Hierarchical Cues in Complex Microenvironments
Zhiyu Jin, Qiu Shen |
PRCV (17) | 3 |
| 2025 | Towards Edge Deployment: An Ultra Lightweight Encoder for Learned Image CompressionabstractLearned Image Compression has shown superior performance over traditional codecs. However, the high computational complexity of existing LIC methods hinders their deployment on resource-constrained edge devices. To address this challenge, we propose an ultra lightweight encoder with an asymmetric architecture: while the encoder is designed to be extremely simple, the decoder employs a more complex synthesis network to maintain high reconstruction quality. The proposed method achieves better rate-distortion performance than BPG with an encoder even lighter than JPEG2000. Peijie Diao, Qiu Shen, Xun Cao |
VCIP | 3 |
| 2025 | Prior Image Guided Snapshot High-Resolution Spectral Imaging in Near InfraredabstractNear-Infrared (NIR) hyperspectral imaging opens up numerous possibilities for wide applications. Despite Compressive Spectral Imaging (CSI) being a promising technique, which enables the acquisition of three-dimensional (3D) spatio-spectral information from dynamic scenes, applying it to the NIR spectrum remains challenging. The bottleneck lies in the high cost and limited resolution of InGaAs Focal Plane Arrays (FPAs), which further degrade the high-frequency information of the compressed measurements. Here we demonstrate a novel Effective Prior Image-guided Spectral imager, termed EpiSpec, towards high-resolution spectral imaging in the NIR. Our key observation is that the tail response of low-cost silicon-based sensors tends to capture similar image, offering high spatial resolution guidance for retrieving details. Hence, the degraded measurement of hyperspectral scene, guided by the prior image, is capable of obtaining high-quality reconstructions. Since the prior image integrates only a partial spectrum of the target scene, introducing content-aware chromatic errors, we propose the Prior Image Guided Deep Unfolding Framework (PIUF) for high-fidelity spectral reconstruction. This framework implicitly models the underlying non-linear relationship between the degraded measurements and the Non-Panchromatic (NPA) prior image. We also introduce a new NIR Spectral Images Dataset (NISID), which features a broad selection of real-world NIR spectral interesting scenes. Based on the dataset in hand, we evaluate the sparse structure of such spectra, which can serve as a guide for efficient CSI sensing matrices design. Extensive evaluations on representative CSI systems demonstrate the effectiveness of the proposed EpiSpec framework. Subsequently, lab prototypes are built for real-world imaging validation, further supporting the viability of high-resolution spectral imaging in the NIR. Lijing Cai, Linsen Chen, Qiu Shen, Xun Cao |
IEEE Trans. Image Process. | 6 |
| 2025 | RefConv: Reparameterized Refocusing Convolution for Powerful ConvNetsabstractWe propose reparameterized refocusing convolution (RefConv) as a replacement for regular convolutional layers, which is a plug-and-play module to improve the performance without any inference costs. Specifically, given a pretrained model, RefConv applies a trainable Refocusing Transformation to the basis kernels inherited from the pretrained model to establish connections among the parameters. For example, a depthwise RefConv can relate the parameters of a specific channel of convolution kernel to the parameters of the other kernel, i.e., make them refocus on the other parts of the model they have never attended to, rather than focus on the input features only. From another perspective, RefConv augments the priors of existing model structures by utilizing the representations encoded in the pretrained parameters as the priors and refocusing on them to learn novel representations, thus further enhancing the representational capacity of the pretrained model. The experimental results validated that RefConv can improve multiple convolutional neural network (CNN)-based models by a clear margin on image classification (up to 1.47% higher top-1 accuracy on ImageNet), object detection, semantic segmentation, and adversarial attacks without introducing any extra inference costs or altering the original model structure. Further studies demonstrated that RefConv can strengthen the spatial skeletons of kernels, reduce the redundancy of channels, and smooth the loss landscape, which explains its effectiveness. Zhicheng Cai, Xiaohan Ding, Qiu Shen, Xun Cao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Batch Normalization Alleviates the Spectral Bias in Coordinate NetworksabstractRepresenting signals using coordinate networks domi-nates the area of inverse problems recently, and is widely applied in various scientific computing tasks. Still, there exists an issue of spectral bias in coordinate networks, lim-iting the capacity to learn high-frequency components. This problem is caused by the pathological distribution of the neural tangent kernel's (NTK's) eigenvalues of coordinate networks. We find that, this pathological distribution could be improved using the classical batch normalization (BN), which is a common deep learning technique but rarely used in coordinate networks. BN greatly reduces the maximum and variance of NTK's eigenvalues while slightly modifies the mean value, considering the max eigenvalue is much larger than the most, this variance change results in a shift of eigenvalues' distribution from a lower one to a higher one, therefore the spectral bias could be alleviated (see Fig. 1). This observation is substantiated by the significant improvements of applying BN-based coordinate networks to various tasks, including the image compression, computed tomography reconstruction, shape representation, magnetic resonance imaging and novel view synthesis. Zhicheng Cai, Hao Zhu 0004, Qiu Shen, Xun Cao |
CVPR | 3 |
| 2024 | MMVP: A Multimodal MoCap Dataset with Vision and Pressure SensorsabstractFoot contact is an important cue for human motion capture, understanding, and generation. Existing datasets tend to annotate dense foot contact using visual matching with thresholding or incorporating pressure signals. However, these approaches either suffer from low accuracy or are only designed for small-range and slow motion. There is still a lack of a vision-pressure multimodal dataset with large-range and fast human motion, as well as accurate and dense foot-contact annotation. To fill this gap, we propose a Multimodal MoCap Dataset with Vision and Pressure sensors, named MMVP. MMVP provides accurate and dense plantar pressure signals synchronized with RGBD observations, which is especially useful for both plausible shape estimation, robust pose fitting without foot drifting, and accurate global translation tracking. To validate the dataset, we propose an RGBD-P SMPL fitting method and also a monocular-video-based baseline framework, VP-MoCap, for human motion capture. Experiments demonstrate that our RGBD-P SMPL Fitting results significantly outperform pure visual motion capture. Moreover, VP-MoCap outperforms SOTA methods in foot-contact and global translation estimation accuracy. We believe the configuration of the dataset and the baseline frameworks will stimulate the research in this direction and also provide a good reference for MoCap applications in various domains. Project page: https://metaverse-ai-lab-thu.github.io/MMVP-Dataset/ He Zhang 0015, Shenghao Ren, Haolei Yuan, Jianhui Zhao 0002, Fan Li 0023, Shuangpeng Sun, Zhenghao Liang, Tao Yu 0007, Qiu Shen, Xun Cao |
CVPR | 9 |
| 2024 | Joint RGB-Spectral Decomposition Model Guided Image Enhancement in Mobile Photography
Kailai Zhou, Lijing Cai, Yibo Wang 0004, Bihan Wen, Qiu Shen, Xun Cao |
ECCV (13) | 6 |
| 2024 | Encoding Semantic Priors into the Weights of Implicit Neural RepresentationabstractImplicit neural representation (INR) has recently emerged as a promising paradigm for signal representations, which takes coordinates as inputs and generates corresponding signal values. Since these coordinates contain no semantic features, INR fails to take any semantic information into consideration. However, semantic information has been proven critical in many vision tasks, especially for visual signal representation. This paper proposes a reparameterization method termed as SPW, which encodes the semantic priors to the weights of INR, thus making INR contain semantic information implicitly and enhancing its representational capacity. Specifically, SPW uses the Semantic Neural Network (SNN) to extract both low- and high-level semantic information of the target visual signal and generates the semantic vector, which is input into the Weight Generation Network (WGN) to generate the weights of INR model. Finally, INR uses the generated weights with semantic priors to map the coordinates to the signal values. After training, we only retain the generated weights while abandoning both SNN and WGN, thus SPW introduces no extra costs in inference. Experimental results show that SPW can improve the performance of various INR models significantly on various tasks, including image fitting, CT reconstruction, MRI reconstruction, and novel view synthesis. Further experiments illustrate that model with SPW has lower weight redundancy and learns more novel representations, validating the effectiveness of SPW. Zhicheng Cai, Qiu Shen |
ICME | 2 |
| 2024 | Leveraging RGB-Pressure for Whole-body Human-to-Humanoid Motion Imitation
Shenghao Ren, Qiu Shen, Xun Cao |
ACM Multimedia | 3 |
| 2024 | Gaseous Object DetectionabstractObject detection, a fundamental and challenging problem in computer vision, has experienced rapid development due to the effectiveness of deep learning. The current objects to be detected are mostly rigid solid substances with apparent and distinct visual characteristics. In this paper, we endeavor on a scarcely explored task named Gaseous Object Detection (GOD), which is undertaken to explore whether the object detection techniques can be extended from solid substances to gaseous substances. Nevertheless, the gas exhibits significantly different visual characteristics: 1) saliency deficiency, 2) arbitrary and ever-changing shapes, 3) lack of distinct boundaries. To facilitate the study on this challenging task, we construct a GOD-Video dataset comprising 600 videos (141,017 frames) that cover various attributes with multiple types of gases. A comprehensive benchmark is established based on this dataset, allowing for a rigorous evaluation of frame-level and video-level detectors. Deduced from the Gaussian dispersion model, the physics-inspired Voxel Shift Field (VSF) is designed to model geometric irregularities and ever-changing shapes in potential 3D space. By integrating VSF into Faster RCNN, the VSF RCNN serves as a simple but strong baseline for gaseous object detection. Our work aims to attract further research into this valuable albeit challenging area. Kailai Zhou, Yibo Wang 0004, Qiu Shen, Xun Cao |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Learn to Enhance the Negative Information in Convolutional Neural Network
Zhicheng Cai, Chenglei Peng, Qiu Shen |
ICIG (1) | 3 |
| 2023 | FalconNet: Factorization for the Light-Weight ConvNets
Zhicheng Cai, Qiu Shen |
ICONIP (1) | 2 |
| 2023 | WormTrack: Dataset and Benchmark for Multi-Object Tracking in Worm CrowdsabstractCurrently, multimedia systems and computer vision algorithms are increasingly playing a crucial role in biological research. However, due to the significant difference between macro and micro scenarios, it is impractical to directly transfer existing computer vision methods to the images captured by microscopes. Taking social behavior analysis of worm for example, it heavily depends on accurate and efficient Multi-object tracking (MOT) methods. Meanwhile, it faces great challenges due to the unique physical characteristics of worm, such as small size, highly uniform appearance, rapid deformation and overlapping movement. This paper studies on the challenges and existing solutions for MOT in worm crowds by building a well-designed dataset ("WormTrack") and a tracking-by-detection benchmark. We observed that the state-of-the-art MOT methods suffers from considerable performance drop on the new dataset. Therefore, we propose a customized MOT method for worm crowds by deeply understanding the physical characteristics of worms and scenes. The method is composed by an instance segmentation based detector, a multiple model fused Kalman filter based tracker and a multi-constraint based trajectory repairer. The experimental results demonstrate that our method can accurately track over 100 worms with almost identical appearance for a long period, which is exceptional compared to existing methods. We hope our work will attract further researches to explore more in this new field, and promote the crossing field researches with biology and medicine. Our code and data is available at https://github.com/Jeerrzy/wormstudio. Zhiyu Jin, Hanyang Yu, Chen Haul, Linxiang Wang, Zuobin Zhu, Qiu Shen, Xun Cao |
ACM Multimedia | 6 |
| 2023 | AU-Oriented Expression Decomposition Learning for Facial Expression Recognition
Zehao Lin, Jiahui She, Qiu Shen |
PRCV (5) | 3 |
| 2023 | Real emotion seeker: recalibrating annotation for facial expression recognition
Zehao Lin, Jiahui She, Qiu Shen |
Multim. Syst. | 3 |
| 2023 | 2C-Net: integrate image compression and classification via deep neural network
Tong Chen 0004, Shiliang Pu, Qiu Shen |
Multim. Syst. | 6 |
| 2023 | FaceScape: 3D Facial Dataset and Benchmark for Single-View 3D Face ReconstructionabstractIn this article, we present a large-scale detailed 3D face dataset, FaceScape, and the corresponding benchmark to evaluate single-view facial 3D reconstruction. By training on FaceScape data, a novel algorithm is proposed to predict elaborate riggable 3D face models from a single image input. FaceScape dataset releases 16,940 textured 3D faces, captured from 847 subjects and each with 20 specific expressions. The 3D models contain the pore-level facial geometry that is also processed to be topologically uniform. These fine 3D facial models can be represented as a 3D morphable model for coarse shapes and displacement maps for detailed geometry. Taking advantage of the large-scale and high-accuracy dataset, a novel algorithm is further proposed to learn the expression-specific dynamic details using a deep neural network. The learned relationship serves as the foundation of our 3D face prediction system from a single image input. Different from most previous methods, our predicted 3D models are riggable with highly detailed geometry under different expressions. We also use FaceScape data to generate the in-the-wild and in-the-lab benchmark to evaluate recent methods of single-view face reconstruction. The accuracy is reported and analyzed on the dimensions of camera pose and focal length, which provides a faithful and comprehensive evaluation and reveals new challenges. The unprecedented dataset, benchmark, and code have been released to the public for research purpose. Hao Zhu 0004, Longwei Guo, Mingkai Huang, Menghua Wu, Qiu Shen, Ruigang Yang, Xun Cao |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2022 | Explore Spatio-temporal Aggregation for Insubstantial Object Detection: Benchmark Dataset and BaselineabstractWe endeavor on a rarely explored task named Insubstantial Object Detection (IOD), which aims to localize the object with following characteristics: (1) amorphous shape with indistinct boundary; (2) similarity to surroundings; (3) absence in color. Accordingly, it is far more challenging to distinguish insubstantial objects in a single static frame and the collaborative representation of spatial and temporal information is crucial. Thus, we construct an IOD-Video dataset comprised of 600 videos (141,017 frames) covering various distances, sizes, visibility, and scenes captured by different spectral ranges. In addition, we develop a spatio-temporal aggregation framework for IOD, in which different backbones are deployed and a spatio-temporal aggregation loss (STAloss) is elaborately designed to leverage the consistency along the time axis. Experiments conducted on IOD-Video dataset demonstrate that spatio-temporal aggregation can significantly improve the performance of IOD. We hope our work will attract further researches into this valuable yet challenging task. The code will be available at: https://github.com/CalayZhou/IOD-Video. Kailai Zhou, Yibo Wang 0004, Yunqian Li, Linsen Chen, Qiu Shen, Xun Cao |
CVPR | 6 |
| 2022 | A Learning-based Approach for Martian Image CompressionabstractFor the scientific exploration and research on Mars, it is an indispensable step to transmit high-quality Martian images from distant Mars to Earth. Image compression is the key technique given the extremely limited Mars-Earth bandwidth. Recently, deep learning has demonstrated remarkable performance in natural image compression, which provides a possibility for efficient Martian image compression. However, deep learning usually requires large training data. In this paper, we establish the first large-scale high-resolution Martian image compression (MIC) dataset. Through analyzing this dataset, we observe an important non-local self-similarity prior for Marian images. Benefiting from this prior, we propose a deep Martian image compression network with the non-local block to explore both local and non-local dependencies among Martian image patches. Experimental results verify the effectiveness of the proposed network in Martian image compression, which outperforms both the deep learning based compression methods and HEVC codec. Mai Xu, Shengxi Li, Xin Deng 0002, Qiu Shen |
VCIP | 5 |
| 2021 | Dive Into Ambiguity: Latent Distribution Mining and Pairwise Uncertainty Estimation for Facial Expression RecognitionabstractDue to the subjective annotation and the inherent interclass similarity of facial expressions, one of key challenges in Facial Expression Recognition (FER) is the annotation ambiguity. In this paper, we proposes a solution, named DMUE, to address the problem of annotation ambiguity from two perspectives: the latent Distribution Mining and the pairwise Uncertainty Estimation. For the former, an auxiliary multi-branch learning framework is introduced to better mine and describe the latent distribution in the label space. For the latter, the pairwise relationship of semantic feature between instances are fully exploited to estimate the ambiguity extent in the instance space. The proposed method is independent to the backbone architectures, and brings no extra burden for inference. The experiments are conducted on the popular real-world benchmarks and the synthetic noisy datasets. Either way, the proposed DMUE stably achieves leading performance. Jiahui She, Yibo Hu 0003, Hailin Shi, Jun Wang 0127, Qiu Shen, Tao Mei 0001 |
CVPR | 5 |
| 2021 | End-to-End Learnt Image Compression via Non-Local Attention Optimization and Improved Context ModelingabstractThis article proposes an end-to-end learnt lossy image compression approach, which is built on top of the deep nerual network (DNN)-based variational auto-encoder (VAE) structure with Non-Local Attention optimization and Improved Context modeling (NLAIC). Our NLAIC 1) embeds non-local network operations as non-linear transforms in both main and hyper coders for deriving respective latent features and hyperpriors by exploiting both local and global correlations, 2) applies attention mechanism to generate implicit masks that are used to weigh the features for adaptive bit allocation, and 3) implements the improved conditional entropy modeling of latent features using joint 3D convolutional neural network (CNN)-based autoregressive contexts and hyperpriors. Towards the practical application, additional enhancements are also introduced to speed up the computational processing (e.g., parallel 3D CNN-based context prediction), decrease the memory consumption (e.g., sparse non-local processing) and reduce the implementation complexity (e.g., a unified model for variable rates without re-training). The proposed model outperforms existing learnt and conventional (e.g., BPG, JPEG2000, JPEG) image compression methods, on both Kodak and Tecnick datasets with the state-of-the-art compression efficiency, for both PSNR and MS-SSIM quality measurements. We have made all materials publicly accessible at https://njuvision.github.io/NIC for reproducible research. Tong Chen 0004, Zhan Ma 0001, Qiu Shen, Xun Cao, Yao Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | Learned Resolution Scaling Powered Gaming-as-a-Service at ScaleabstractBuilt on the explosive advancement of cloud and telecommunication technologies, Gaming-as-a-Service (GaaS) or cloud gaming system is expected to revolutionize the traditional multi-billion video game market in the near future. This wave is analogous to the rise of live-video-streaming-based-Netflix to replace conventional DVD rental business for movies and TVs. In practice, a successful GaaS platform need to operate in a transparent mode without requiring substantial efforts from both content providers and end users, and offer the pristine quality of experience (QoE) at an affordable cost. Our analysis suggests that GaaS provisioning cost can be reduced significantly by enforcing the game video rendering and streaming at a lower resolution (so as to increase the user concurrency in the cloud and reduce the streaming bandwidth over the network). However, streaming video at a lower resolution may deteriorate the QoE. To maintain the client QoE at the level using the default-native resolution for streaming or even enhance it, we introduce the learned resolution scaling (LRS), which leverages the computational capabilities at clients/edges to restore/improve the reconstructed image/video quality via stacked deep neural networks (DNN). We integrate this LRS into a commercialized GaaS platform - AnyGame, to study its efficiency and complexity quantitatively. Extensive real-life experiments have shown that LRS-powered AnyGame offers the state-of-the-art performance, and the lower operational cost, paving the road for a potential success of GaaS over the Internet. Additionally, we dive into proposed LRS via ablation studies to further demonstrate its consistent performance, including the discussions on trade-off between efficiency and complexity, alternative training sets, etc. Hao Chen 0036, Ming Lu 0003, Zhan Ma 0001, Xu Zhang 0006, Yiling Xu, Qiu Shen, Wenjun Zhang 0001 |
IEEE Trans. Multim. | 6 |
| 2020 | FaceScape: A Large-Scale High Quality 3D Face Dataset and Detailed Riggable 3D Face PredictionabstractIn this paper, we present a large-scale detailed 3D face dataset, FaceScape, and propose a novel algorithm that is able to predict elaborate riggable 3D face models from a single image input. FaceScape dataset provides 18,760 textured 3D faces, captured from 938 subjects and each with 20 specific expressions. The 3D models contain the pore-level facial geometry that is also processed to be topologically uniformed. These fine 3D facial models can be represented as a 3D morphable model for rough shapes and displacement maps for detailed geometry. Taking advantage of the large-scale and high-accuracy dataset, a novel algorithm is further proposed to learn the expression-specific dynamic details using a deep neural network. The learned relationship serves as the foundation of our 3D face prediction system from a single image input. Different than the previous methods, our predicted 3D models are riggable with highly detailed geometry under different expressions. The unprecedented dataset and code will be released to public for research purpose. Hao Zhu 0004, Mingkai Huang, Qiu Shen, Ruigang Yang, Xun Cao |
CVPR | 5 |
| 2020 | Modeling the Perceptual Quality of Viewport Adaptive Omnidirectional Video StreamingabstractInstead of streaming the entire OmniDirectional Videos (ODVs) that are often sampled at ultra high definition and high frame rate, a viewport adaptive streaming is preferred in practice. We usually stream the High-Quality (HQ) content within current viewport, while Low-Quality (LQ) elsewhere to save the network bandwidth consumption. Such scheme would lead to a quality refinement after user adapts his/her focus to a new viewport. In this paper, we thus model the perceptual impact of the quality variations (through adjusting the Quantization Stepsize (QS or q) and Spatial Resolution (SR or s)) with respect to the Refinement Duration (RD or τ) when performing the refinement from an arbitrary LQ scale to an arbitrary HQ one. A number of quality variations are studied to cover sufficient use cases in practice, resulting in a unified analytical model, as a product of separable exponential functions that measure the QS and SR induced perceptual impacts in terms of the RD, and a perceptual index measuring the subjective quality of corresponding viewport video after refinement. This model is first validated in a managed lab environment via independent subjective assessments by constraining user's navigation to avoid unexpected noise, where both Pearson Correlation Coefficient (PCC) and Spearman's Rank Correlation Coefficient (SRCC) are around 0.97. We then extend the validations in a real-life viewport-dependent streaming system, still yielding PCC and SRCC about 0.96 when comparing collected subjective scores with model predictions. Shaowei Xie, Yiling Xu, Qiu Shen, Zhan Ma 0001, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Learned Quality Enhancement via Multi-Frame Priors for HEVC Compliant Low-Delay ApplicationsabstractNetworked video applications, e.g., video conferencing, often suffer from poor visual quality due to unexpected network fluctuation and limited bandwidth. In this paper, we have developed a Quality Enhancement Network (QENet) to reduce the video compression artifacts, leveraging the spatial and temporal priors generated by respective multi-scale convolutions spatially and warped temporal predictions in a recurrent fashion temporally. We have integrated this QENet as a stand-alone post-processing subsystem to the High-Efficiency Video Coding (HEVC) compliant decoder. Experimental results show that our QENet demonstrates the state-of-the-art performance against default in-loop filters in HEVC and other deep learning based methods with noticeable objective gains in Peak Signal-to-Noise Ratio (PSNR) and subjective gains visually. Ming Lu 0003, Yiling Xu, Shiliang Pu, Qiu Shen, Zhan Ma 0001 |
ICIP | 5 |
| 2019 | Looking-Ahead: Neural Future Video Frame PredictionabstractWe have developed a Looking-Ahead system to facilitate the future video frame prediction via deep learning, which is of practical value in the domain like autonomous driving etc. The overall problem is decomposed into cascaded optical flow prediction and subsequent predictive frame post-processing for quality refinement. A pyramid flow calculation across existing frames is used to efficiently infer the motion of target frame; while a universal inpainting network is applied to restore those motion-induced occluded pixels. Compared with those published methods, our Looking-Ahead offers the state-of-the-art performance measured objectively with better Peak-Signal-to-Noise Ratio (PSNR) and Structural Similarity (SSIM), and more appealing reconstructions. Changxu Zhang, Tong Chen 0004, Qiu Shen, Zhan Ma 0001 |
ICIP | 4 |
| 2019 | Extreme Image Coding via Multiscale Autoencoders with Generative Adversarial OptimizationabstractWe propose a MultiScale AutoEncoder (MSAE) based extreme image coding/compression framework to offer visually pleasing reconstruction at a very low bitrate. Our method leverages the "priors" at different resolution scale to improve the compression efficiency, and also employs the generative adversarial network (GAN) with multiscale discriminators to perform the end-to-end trainable rate-distortion optimization. We compare the perceptual quality of our reconstructions with traditional compression algorithms using High-Efficiency Video Coding (HEVC) based Intra Profile and JPEG2000 on the public Cityscapes, ADE20K and Kodak datasets, demonstrating the significant subjective quality improvement. However, objective measurements, such as PSNR, SSIM, etc, are often deteriorated by applying the generative adversarial optimization. Tong Chen 0004, Qiu Shen, Zhan Ma 0001 |
VCIP | 4 |
| 2019 | Codedretrieval: Joint Image Compression and Retrieval with Neural NetworksabstractWith the explosive increase of image data, the efficiency of both image compression and retrieval becomes unprecedentedly significant. However, these two tasks are usually isolated executed, which waste great computational resources in large-scale image applications. In this work, we propose a joint framework called CodedRetrieval, which can find a general feature expression for both compression and retrieval based on neural network. Additionally, a two stage training strategy is designed to achieve better balance between the two distinct tasks. Experimental results show that our method can achieve competitive perfomance on both compression and retrieval comparing to classic methods, while saving great amount of computation time. Tong Chen 0004, Qiu Shen, Zhan Ma 0001 |
VCIP | 4 |
| 2018 | Modeling the Perceptual Impact of Viewport Adaptation for Immersive VideoabstractImmersive video offers the freedom to navigate inside the virtualized environment. Instead of streaming the entire bulky content, a viewport or field of view (FoV) adaptive streaming is preferred. We often stream the high-quality content within current viewport, but degraded-quality representation elsewhere, so as to reduce the network bandwidth consumption. We then could refine the quality when focusing to a new FoV. Therefore, in this work, we have attempted to model the perceptual response of the quality variations (through adapting the quantization and spatial resolution) with respect to the refinement duration, and reach at a product of two closed-form exponential functions that well explain the joint quantization and resolution induced quality impact. Analytical model is also cross-validated using another set of data with both Pearson and Spearman's rank rank correlations over 0.98. Our work would be devised to guide the bandwidth-quality optimized immersive video streaming. Shaowei Xie, Yiling Xu, Qiaojian Qian, Qiu Shen, Zhan Ma 0001, Wenjun Zhang 0001 |
ISCAS | 4 |
| 2018 | Modeling the Perceptual Quality of Immersive Images Rendered on Head Mounted Displays: Resolution and CompressionabstractWe develop a model that expresses the joint impact of spatial resolution s and JPEG compression quality factor qf on immersive image quality. The model is expressed as the product of optimized exponential functions of these factors. The model is tested on a subjective database of immersive image contents rendered on a head mounted display (HMD). High Pearson correlation and Spearman correlation (> 0.95) and small relative root mean squared error (< 5.6%) are achieved between the model predictions and the subjective quality judgements. The immersive ground-truth images along with the rest of the database are made available for future research and comparisons. Mingkai Huang, Qiu Shen, Zhan Ma 0001, Alan C. Bovik, Praful Gupta, Rongbing Zhou, Xun Cao |
IEEE Trans. Image Process. | 2 |
| 2017 | DeepCoder: A deep neural network based video compressionabstractInspired by recent advances in deep learning, we present the DeepCoder - a Convolutional Neural Network (CNN) based video compression framework. We apply separate CNN nets for predictive and residual signals respectively. Scalar quantization and Huffman coding are employed to encode the quantized feature maps (fMaps) into binary stream. We use the fixed 32 × 32 block in this work to demonstrate our ideas, and performance comparison is conducted with the well-known H.264/AVC video coding standard with comparable rate-distortion performance. Here distortion is measured using Structural Similarity (SSIM) because it is more close to perceptual response. Tong Chen 0004, Qiu Shen, Tao Yue 0003, Xun Cao, Zhan Ma 0001 |
VCIP | 3 |
| 2017 | Modeling peripheral vision impact on perceptual quality of immersive imagesabstractConventional images/videos are often rendered within the central area of human visual system (HVS) with uniform quality. Recent virtual reality (VR) device with head mounted display (HMD) extends the field of view (FoV) significantly to include both central and peripheral areas. It exhibits the unequal acuity of the quality sensation because of the non-uniform distribution of photoreceptors in our retina. Hence, we propose to study the impact of image qualities (with respect to the quantization stepsize q or spatial resolution s) in peripheral vision and conclude self-adaptive analytical models that have shown quite impressive accuracy through independent cross validations. These models can further be applied to assign different quality weights at different regions, so as to significantly reduce the transmission data size but without subjective quality loss. Peiyao Guo, Qiu Shen, Mingkai Huang, Rongbing Zhou, Xun Cao, Zhan Ma 0001 |
VCIP | 2 |
| 2017 | 3D motion estimation via optimized feature point selection
Qiu Shen, Yuxi Dai, Fanqiang Kong |
Neurocomputing | 1 |
| 2009 | Content-based hierarchical motion description for multiple video adaptationabstractVideo adaptation has been considered as a promising technique to tackle challenging problems in pervasive multimedia applications. However, the styles of video representation and description in existing framework are not flexible enough to adapt diversified application environment. In this paper, we propose a novel solution based on intermediate description, which can support fast multiple video adaptation operations in signal level (e.g., temporal resolution reduction, bit-rate adaptation), structural level (e.g., random access of any shots, fast preview of any key-frames or thumbnails), as well as joint level. Experimental results show that the presented solution can support the operations in real-time environment, while maintaining coding efficiency, which demonstrates its feasibility and effectiveness. Qiu Shen, Houqiang Li, Feng Wu 0001 |
ICME | 1 |