Xin Jin 0002

dblp:68/3340-2 · DBLP profile ↗
← Back
76ranked-venue papers
18as first author
19since 2021 · last 2025
0000-0001-6655-3888ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 68 · 17 first-author · 16 since 2021Systems, architecture and hardware · 6 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 ScatterSplatting: Enhanced View Synthesis in Scattering Scenarios via Joint NeRF and Gaussian Splatting
abstract
View synthesis in the scattering medium faces challenges due to image degradation caused by medium scattering, and existing methods struggle to effectively address this issue. Although Neural Radiance Fields (NeRFs) excel in novel view synthesis and high-quality 3D reconstruction, they suffer from slow training speeds and poor handling of scenes with scattering. However, 3D Gaussian Splatting (3DGS) naturally exhibits sparsity, allowing for efficient representation of complex opaque geometries and surface details, but it is limited in modeling medium scattering. To address these limitations, we propose a novel method that combines the volumetric rendering strengths of NeRF with the sparse representation capabilities of 3DGS to effectively handle scattering media. By using 3DGS to represent objects and leveraging volumetric rendering for medium modeling, our method can mitigate the effects of medium scattering, rendering clear scenes in scattering media and synthesizing the appearance and depth of distant objects from novel viewpoints. The experimental results demonstrate a significant improvement in the performance of view-synthesis tasks, with an average increase of 3 to 4 dB. The rendering efficiency improved, with the training time reduced to 8 minutes, while preserving fine details.
Renrong Hu, Qianyue He, Dongyu Du, Xin Jin 0002
ISCAS4
2025 Atmospheric Scattered Light Field Sampling for Improving Reconstruction Efficiency
abstract
Light field (LF) atmospheric descattering methods using multi-view images from camera arrays offer significant advantages for solving strong scattering due to their ability in exploiting high-dimensional light information. However, the relationship between performance and scattered LF sampling rate (i.e., the density of samples per unit area) is an unknown coupling, affecting acquisition and processing complexity. In this paper, we define the minimum atmospheric scattered LF sampling rate under optimal descattering quality, based on attenuated spectral support in scattering scenarios derived from the proposed atmospheric point spread function (APSF). The proposed APSF integrates the camera model, radiative transfer equation, and modified generalized Gaussian distribution (GGD) to describe multiple scattering. For any scattering parameters, the proposed APSF can be directly derived without infinite series, ensuring full adaptability to all acquisition systems through the integration of system model. Combining APSF with scene and acquisition system information, the scattered LF spectrum is determined, and consequently the minimum atmospheric scattered LF sampling rate is derived for the first time. Experimental results demonstrate the accuracy, effectiveness, and robustness of the proposed atmospheric scattered LF sampling theory through comparisons of atmospheric descattering performance across different LF sampling rates, object types, scene depths, and scattering intensities. The proposed method achieves a reduction in the number of acquisition cameras by an average of 78.4% while maintaining processing quality, which significantly enhances the applicability of LF atmospheric descattering methods.
Yihui Fan, Dongyu Du, Hongkun Cao, Jiayu Xie, Xin Jin 0002
IEEE Trans. Circuits Syst. Video Technol.5
2024 Multi-Dimensional Geometric Feature-Based Calibration Method for LiDAR and Camera Fusion
abstract
Extrinsic calibration between LiDAR and camera has become an indispensable task across diverse domains, including autonomous vehicles, robotics, and surveillance systems. However, existing methods suffer from limited precision due to the inaccurate and insufficient detected features caused by the sparsity of point-clouds and the inherent ranging errors of LiDAR. In this paper, we propose a multi-dimensional geometric feature-based calibration method between LiDAR and camera. First, a 3D structured calibration target is proposed with multi-normal surfaces, edges in different directions, and distinctive corner features. Secondly, a point-plane and angle error-based point-cloud feature detection method is designed to establish 3D-2D feature point pairs with image features. Finally, a Perspective-n-Point (PnP) problem is solved to estimate the extrinsic parameters. The experimental results show that the proposed method reduces Mean Reprojection Error (MRE) by 0.05 pixels and achieves a 70% reduction in Normalized Reprojection Error (NRE) compared with state-of-the-art (SOTA) methods under the conditions of smaller training size, larger test size, and more repeat times.
Yuhan Hao, Xin Jin 0002, Dongyu Du
ICASSP2
2024 Optimizing Limited-Stop Transit Operations Through Simulation: A Case Study of Shenzhen, China
abstract
Limited-stop service represents a paradigm for enhancing efficiency in transportation systems, with applications spanning public transit, logistics, and supply chain management. This service model operates along a designated route featuring a reduced number of stops in comparison to standard services, thereby facilitating expedited transportation. In this research, we utilize accumulated data from Intelligent Transportation Systems for the development and deployment of a simulation-driven framework tailored towards augmenting the operational efficiency of limited-stop services. To illustrate the practical application of this framework, we focus on the paradigm of limited-stop bus services, elucidating the intricate technical facets of the proposed system. The efficacy and viability of this framework have been validated through a series of empirical investigations conducted within the urban bus transit network of Shenzhen, a prominent metropolis situated in China.
Xiancai Tian, Xin Jin 0002
MobiCom4
2024 AdWeatherNet: Adverse Weather Denoising with Point Cloud Spatiotemporal Attention
abstract
Adverse weather introduces disruptive noise into LiDAR data within autonomous driving systems, compromising the accuracy and range of 3D perception. Mitigating this challenge for high-precision noise removal becomes intricate due to the varying noise distributions at different distances. A novel spatiotemporal denoising network, AdWeatherNet, is proposed to address this problem. The Spatial Encoder module dynamically encodes spatial features using a designed density evaluation model. Additionally, the Temporal Differential Attention module effectively leverages temporal variation in adjacent point clouds to identify and accurately remove noise. To drive the research, we also introduce an adverse weather dataset, named the AdScenes dataset, which features point-wise annotations and a wide variety of weather conditions, making it one of the largest comprehensive datasets in this domain. The experimental results demonstrate the effectiveness of our method, with a remarkable improvement of +9.5% of IoU in rainy scenes, +5.9% of IoU in snowy scenes, and +1.9% of IoU in foggy scenes. Compared to the SOTA, AdWeatherNet enhances the mAP of object detection by an average of +1.8% across all weather conditions. Our method contributes to the development of reliable LiDAR perception systems, fostering the development of autonomous vehicles.
Haozheng Han, Dongyu Du, Xin Jin 0002
VCIP4
2024 Light Fields Stitching for Windowed-6DoF VR Content
abstract
Windowed six degrees of Freedom (Windowed-6DoF) virtual reality (VR) content that provides users an immersive feeling of walking through a 3D 360 VR space with constrained rotational movements around X and Y axes and constrained translational movements along Z axis is important for the development of VR. To facilitate this windowed-6DoF immersive feeling, light fields (LFs) from multiple perspectives within the windowed 6-DoF space are captured. In contrast to employing a large-scale camera array for LF capture, utilizing hand-held plenoptic cameras offers a more portable and versatile solution, thereby promoting practical applications. However, how to stitch the LFs at different rotational angles containing motion parallax is challenging. In this paper, a novel LF stitching method is proposed to generate windowed-6DoF LFs. First, multi-concentric spherical modeling is proposed to parameterize the recorded LFs to eliminate projection biases in the registration process. Then, a global-local adaptive LF registration is proposed by developing incremental multi-layer global-local adaptive homographies based on the 4D light field feature (LiFF), incremental strategy and depth layer maps (DLMs) to eliminate parallax errors. Testing on the LFs captured in both indoor and outdoor scenes with different focal lengths, quantities of LFs and scales of translation and rotation, the proposed method outperforms the existing approaches in terms of subjective quality, objective quality, light field consistency and content production robustness, which can produce VR content of superior quality more reliably.
Yihui Fan, Xin Jin 0002, Siyao Zhou 0003, Shun Zou
IEEE Trans. Circuits Syst. Video Technol.2
2024 Learned Focused Plenoptic Image Compression With Microimage Preprocessing and Global Attention
abstract
Focused plenoptic cameras can record spatial and angular information of the light field (LF) simultaneously with higher spatial resolution relative to traditional plenoptic cameras, which facilitate various applications in computer vision. However, the existing plenoptic image compression methods present ineffectiveness to the captured images due to the complex micro-textures generated by the microlens relay imaging and long-distance correlations among the microimages. In this article, a lossy end-to-end learning architecture is proposed to compress the focused plenoptic images efficiently. First, a data preprocessing scheme is designed according to the imaging principle to remove the sub-aperture image ineffective pixels in the recorded light field and align the microimages to the rectangular grid. Then, the global attention module with large receptive field is proposed to capture the global correlation among the feature maps using pixel-wise vector attention computed in the resampling process. Also, a new image dataset consisting of 1910 focused plenoptic images with content and depth diversity is built to benefit training and testing. Extensive experimental evaluations demonstrate the effectiveness of the proposed approach. It outperforms intra coding of HEVC and VVC by an average of 62.57% and 51.67% bitrate reduction on the 20 preprocessed focused plenoptic images, respectively. Also, it achieves 18.73% bitrate saving and generates perceptually pleasant reconstructions compared to the state-of-the-art end-to-end image compression methods, which benefits the applications of focused plenoptic cameras greatly. The dataset and code are publicly available athttps://github.com/VincentChandelier/GACN.
Kedeng Tong, Xin Jin 0002, Jinshi Kang, Fan Jiang 0012
IEEE Trans. Multim.2
2024 DARTS: Diffusion Approximated Residual Time Sampling for Time-of-flight Rendering in Homogeneous Scattering Media
abstract
Time-of-flight (ToF) devices have greatly propelled the advancement of various multi-modal perception applications. However, achieving accurate rendering of time-resolved information remains a challenge, particularly in scenes involving complex geometries, diverse materials and participating media. Existing ToF rendering works have demonstrated notable results, yet they struggle with scenes involving scattering media and camera-warped settings. Other steady-state volumetric rendering methods exhibit significant bias or variance when directly applied to ToF rendering tasks. To address these challenges, we integrate transient diffusion theory into path construction and propose novel sampling methods for free-path distance and scattering direction, via resampled importance sampling and offline tabulation. An elliptical sampling method is further adapted to provide controllable vertex connection satisfying any required photon traversal time. In contrast to the existing temporal uniform sampling strategy, our method is the first to consider the contribution of transient radiance to importance-sample the full path, and thus enables improved temporal path construction under multiple scattering settings. The proposed method can be integrated into both path tracing and photon-based frameworks, delivering significant improvements in quality and efficiency with at least a 5x MSE reduction versus SOTA methods in equal rendering time.
Qianyue He, Dongyu Du, Haitian Jiang, Xin Jin 0002
ACM Trans. Graph.4
2023 Discrete Point-Wise Attack is Not Enough: Generalized Manifold Adversarial Attack for Face Recognition
abstract
Classical adversarial attacks for Face Recognition (FR) models typically generate discrete examples for target identity with a single state image. However, such paradigm of point-wise attack exhibits poor generalization against numerous unknown states of identity and can be easily defended. In this paper, by rethinking the inherent relationship between the face of target identity and its variants, we introduce a new pipeline of Generalized Manifold Adversarial Attack (GMAA)11https://github.com/tokaka22/GMAA to achieve a better attack performance by expanding the attack range. Specifically, this expansion lies on two aspects - GMAA not only expands the target to be attacked from one to many to encourage a good generalization ability for the generated adversarial examples, but it also expands the latter from discrete points to manifold by leveraging the domain knowledge that face expression change can be continuous, which enhances the attack effect as a data augmentation mechanism did. Moreover, we further design a dual supervision with local and global constraints as a minor contribution to improve the visual quality of the generated adversarial examples. We demonstrate the effectiveness of our method based on extensive experiments, and reveal that GMAA promises a semantic continuous adversarial space with a higher generalization ability and visual quality.
Yuxiao Hu 0003, Dongxiao Zhang, Xin Jin 0002, Yuntian Chen
CVPR5
2023 Denoising Point Clouds with Intensity and Spatial Features in Rainy Weather
abstract
LiDAR is important for 3D vision in autonomous vehicles, but rain causes inaccurate LiDAR point clouds due to reflection and scattering. Rain noise removal without loss of environmental features becomes an inevitable challenge. This paper presents a novel point cloud denoising method with intensity and spatial features to solve the problem. It utilizes a weighted edge-preserving filter to recover distorted contours and intensities of point clouds due to the reflection of the surface attached by raindrops. A low-intensity filtering method is also proposed to remove low-intensity noise due to the reflection of rainfall. In addition, a semi-synthetic rainy point cloud dataset with point-wise annotations is created, which benefits the research on improving LiDAR perception in adverse weather. Our method outperforms existing methods in terms of precision when it achieves a high recall of 99.28%. Using denoised data by our method can improve target detection accuracy by 5.37%. It is also faster than the state-of-the-art methods and shows the potential for use in snowy weather, making it suitable for all-weather LiDAR applications.
Haozheng Han, Xin Jin 0002, Zhiheng Li 0001
ICIP2
2023 QVRF: A Quantization-Error-Aware Variable Rate Framework for Learned Image Compression
abstract
Learned image compression has exhibited promising compression performance, but variable bitrates over a wide range remain a challenge. State-of-the-art variable rate methods compromise the loss of model performance and require numerous additional parameters. In this paper, we present a Quantization-error-aware Variable Rate Framework (QVRF) that utilizes a univariate quantization regulator a to achieve wide-range variable rates within a single model. Specifically, QVRF defines a quantization regulator vector coupled with predefined Lagrange multipliers to control quantization error of all latent representation for discrete variable rates. Additionally, a reparameterization method makes QVRF compatible with round quantizer and integer entropy coding. Exhaustive experiments demonstrate that existing fixed-rate VAE-based methods equipped with QVRF can achieve wide-range continuous variable rates within a single model without significant performance degradation. Furthermore, QVRF outperforms contemporary variable-rate methods in rate-distortion performance with minimal additional parameters. The code is available at https://github.com/bytedance/QRAF.
Kedeng Tong, Yaojun Wu 0001, Yue Li 0015, Kai Zhang 0007, Li Zhang 0006, Xin Jin 0002
ICIP6
2023 Thermal Infrared Guided Color Image Dehazing
abstract
Owing to the superior wavelength properties in penetrating haze, the thermal infrared images maintain high contrast and sharp edge even in the presence of dense haze, holding significant potential for improving color image dehazing. Thus, we propose a thermal infrared guided color image dehazing algorithm. Rather than direct fusion, an optimization framework is established, in which the regional contrast information of the thermal infrared guides contrast enhancement of color image, and the edge information is used to transmission map refinement and edge preservation. In conjunction with a color fidelity constraint, the optimization framework is solved with gradient descent. Additionally, we propose the Thermal infrared/Visible Images in Haze dataset (TVIH), which consists of registered high-resolution image pairs in outdoor dense haze and mist scenarios. Our method outperforms single image dehazing and image fusion methods on our dataset regarding subjective quality and objective metrics.
Jiayu Xie, Xin Jin 0002
ICIP2
2023 Microimage-based Two-step Search For Plenoptic 2.0 Video Coding
abstract
The plenoptic 2.0 video can record a time-varying dense light field, which benefits many immersive visual applications such as AR/VR. However, traditional inter motion estimation methods perform inefficiently in such kinds of video sequences due to the distinctive temporal characteristics caused by the imaging principle. In this paper, a microimage-based two- step search (MTSS) is proposed to achieve a better trade-off between coding performance and coding complexity. Based on microimage focus variation analysis in imaging dynamic scenes, a microlens-diameter and matching-distance spatial search with local refinement is proposed to exploit the image correlations among the microimage and to compensate the defocused inaccuracy. Implementing the proposed motion estimation in H.266 platform VTM-11.0 and comparing with the state-of-the-art methods, obvious compression efficiency improvements are achieved with limited complexity increment, which benefits the standardization of plenoptic video coding.
Xin Jin 0002, Kedeng Tong, Haitian Huang
ICME2
2023 Gravity-Shift-VIO: Adaptive Acceleration Shift and Multi-Modal Fusion with Transformer in Visual-Inertial Odometry
abstract
Visual-inertial odometry (VIO) estimates the 6-degree-of-freedom (6-DoF) ego-motion of an agent based on sequential data from cameras and inertial measurement units (IMUs). The acceleration measured through IMUs is affected by gravity, which is typically addressed by initialization methods in traditional VIO approaches. However, this problem has not received much attention in recent end-to-end deep learning methods. For raw accelerations, gravity causes overlapping be-tween different motion patterns, degenerating the representation embedding, which limits the performance of pose estimation. In this paper, we propose Gravity-Shift-VIO, an attention-based approach that addresses this issue by adaptively shifting the acceleration vector before the representation embedding. Further, a cross-frame multimodal transformer is introduced to fuse multimodal information. Experimentation on the KITTI dataset shows that Gravity-Shift-VIO exhibits strong performance and shows promising results in terms of ego-motion estimation. Further ablation study indicates that the Gravity-Shift- Viois highly effective in reducing the overlap of acceleration representation caused by gravity. And the cross-frame transformer effectively improves the multi-sensor fusion and time-series feature extraction.
Zhiheng Li 0001, Xin Jin 0002
IJCNN4
2022 SADN: Learned Light Field Image Compression with Spatial-Angular Decorrelation
abstract
Light field image becomes one of the most promising media types for immersive video applications. In this paper, we propose a novel end-to-end spatial-angular-decorrelated network (SADN) for high-efficiency light field image compression. Different from the existing methods that exploit either spatial or angular consistency in the light field image, SADN decouples the angular and spatial information by dilation convolution and stride convolution in spatial-angular interaction, and performs feature fusion to compress spatial and angular information jointly. To train a stable and robust algorithm, a large-scale dataset consisting of 7549 light field images is proposed and built. The proposed method provides 2.137 times and 2.849 times higher compression efficiency relative to H.266/VVC and H.265/HEVC inter coding, respectively. It also outperforms the end-to-end image compression networks by an average of 79.6% bitrate saving with much higher subjective quality and light field consistency.
Kedeng Tong, Xin Jin 0002, Fan Jiang 0012
ICASSP2
2022 Piecewise Linear Model Based Local Illumination Compensation Inter Prediction for Video Coding
abstract
In video coding, the illumination information of video scenes is usually hard to be compressed due to the complex and unpredictable illumination variations. To simplify the problem, many prediction algorithms assume a linear correlation of illumination variations existing between frames by constructing corresponding linear model (LM), such as Weighted Prediction (WP) method. The assumption is suitable for uniform illumination variations with lager area, but ineffective in sharp illumination variations in small area. In this paper, we propose a piecewise linear model based local illumination compensation (PLMLIC) approach to further compensate sharp illumination variations in small area. When PLMLIC is applied to a coding unit (CU), we first use multiple reference lines, i.e. allow not only the nearest reference line but also long-distance reference lines to be the neighbouring samples of the current CU. Then the neighbouring samples and their corresponding reference samples are classified into 2 groups, based on which PLMLIC parameters are derived for each group. Finally, the current CU is predicted by the corresponding PLMLIC parameters. Experimental results show that 0.23% BD-rate savings on average can be achieved for lowdelay configuration based on Enhanced Compression Model (ECM) beyond VVC.
Xin Jin 0002, Huanbang Chen, Haitao Yang 0001, Kedeng Tong
PCS2
2021 Pixel Gradient Based Zooming Method for Plenoptic Intra Prediction
abstract
Plenoptic 2.0 videos that record time-varying light fields by focused plenoptic cameras are prospective to immersive visual applications due to capturing dense sampled light fields with high spatial resolution in the rendered sub-apertures. In this paper, an intra prediction method is proposed for compressing multi-focus plenoptic 2.0 videos efficiently. Based on the estimation of zooming factor, novel gradient-feature-based zooming, adaptive-bilinear-interpolation-based tailoring and inverse-gradient-based boundary filtering are proposed and executed sequentially to generate accurate prediction candidates for weighted prediction working with adaptive skipping strategy. Experimental results demonstrate the superior performance of the proposed method relative to HEVC and state-of-the-art methods.
Fan Jiang 0012, Xin Jin 0002, Kedeng Tong
VCIP2
2021 SMRD: A Local Feature Descriptor for Multi-modal Image Registration
abstract
Image registration among multimodality has received increasing attention in the scope of computer vision and computational photography nowadays. However, the non-linear intensity variations prohibit the accurate feature points matching between modal-different image pairs. Thus, a robust image descriptor for multi-modal image registration is proposed, named shearlet-based modality robust descriptor(SMRD). The anisotropic feature of edge and texture information in multi-scale is encoded to describe the region around a point of interest based on discrete shearlet transform. We conducted the experiments to verify the proposed SMRD compared with several state-of-the-art multi-modal/multispectral descriptors on four different multi-modal datasets. The experimental results showed that our SMRD achieves superior performance than other methods in terms of precision, recall and F1-score.
Jiayu Xie, Xin Jin 0002, Hongkun Cao
VCIP2
2021 Deep Learning-Based Chroma Prediction for Intra Versatile Video Coding
abstract
Color images always exhibit a high correlation between luma and chroma components. Cross component linear model (CCLM) has been introduced to exploit such correlation for removing redundancy in the on-going video coding standard, i.e., versatile video coding (VVC). To further improve the coding performance, this paper presents a deep learning based intra chroma prediction method, termed as convolutional neural network based chroma prediction (CNNCP). More specifically, the process of chroma prediction is formulated to produce the colorful version from available information input. CNNCP includes two sub-networks for luma down-sampling and chroma prediction, which are jointly optimized to fully exploit spatial and cross component information. In addition, the outputs of CCLM are adopted as chroma initialization for performance enhancement, and the coding distortion level characterized by quantization parameter is fed into the network to release the negative affect from compression artifacts. To further improve the coding performance, the competition is performed between the conventional chroma prediction and CNNCP in terms of rate-distortion cost with a binary flag signalled. The learned CNNCP is incorporated into both video encoder and decoder. Extensive experimental results demonstrate that the proposed scheme can achieve 4.283%, 3.343%, and 4.634% bit rate savings for luma and two chroma components, compared with the VVC test model version 4.0 (VTM 4.0).
Linwei Zhu, Yun Zhang 0002, Shiqi Wang 0001, Sam Kwong, Xin Jin 0002, Yu Qiao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2020 Chromatic Aberration Correction Using Cross-Channel Prior in Shearlet Domain
Kunyi Li, Xin Jin 0002
ACCV (2)2
2020 EPI-Neighborhood Distribution Based Light Field Depth Estimation
abstract
In this paper, a novel depth estimation algorithm tackling foreground occlusion is proposed based on the neighborhood distribution in the sheared epipolar images (EPIs). First, the EPI is sheared to perform refocusing. Next a series of sheared EPI's neighboring pixels in a local window are selected and the corresponding histogram distributions are analyzed by the proposed novel tensor, Kullback-Leibler Divergence (KLD). Then, depths calculated from vertical and horizontal EPIs' tensors are fused according to the tensors' variation scale for a high quality depth map. Finally, confident depth points are propagated to the whole image by global optimization. Experimental results show that the proposed algorithm achieves better performance relative to state-of-the-art algorithms.
Junke Li, Xin Jin 0002
ICASSP2
2020 Light Field Stitching Based On Concentric Spherical Modeling
abstract
VR image in form of the spherical panoramic image is already widely available while enhancing its immersive experience with six degrees of freedom (6-DoF) is fundamentally required. Spherical panoramic light field (LF) becomes a potential solution because of recording the spatial and angular information of the light rays in the 360° spherical space. In this paper, a novel method is proposed to generate spherical panoramic LF by stitching LFs captured at different rotational angles. First, concentric spherical modeling is proposed to parameterize the recorded rays to eliminate the projection biases in registration. Then, the concentric spherical model-based LF registration which is insensitive to the ordering is introduced to transform each 4D LFs mesh accurately. Finally, the stitching result is projected to Two-parallel-plane (TPP) coordinates for viewing. Experimental results show that the proposed method outperforms the existing methods in terms of subjective quality and continuity in the stitched LF.
Siyao Zhou 0003, Xin Jin 0002
ICIP2
2020 Imaging-Correlated Intra Prediction for Plenoptic 2.0 Video Coding
abstract
The plenoptic 2.0 video records time-varying light fields varying with much higher spatial resolution for each subaperture image. However, the huge data volume becomes a significant obstacle of the plenoptic 2.0 video application. Therefore, this paper proposes an efficient compression method designed for plenoptic 2.0 video coding. First, the matching blocks relative to the current block are determined by imaging principles. The matching blocks and current block are imaged by the same object under neighboring microlens, which are highly correlated with each other. Then, weighted prediction by using the matching blocks as reference blocks is adopted for the intra prediction of current block. Compared with HEVC, experimental results show that the proposed method can save 9.63% and 4.73% bitrate under the All Intra and Random Access configuration, respectively.
Lingjun Li, Xin Jin 0002, Tingting Zhong
ICME2
2020 3D-CNN Autoencoder for Plenoptic Image Compression
abstract
Recently, plenoptic image has attracted great attentions because of its applications in various scenarios. However, high resolution and special pixel distribution structure bring huge challenges to its storage and transmission. In order to adapt compression to the structural characteristic of plenoptic image, in this paper, we propose a Data Structure Adaptive 3D-convolutional(DSA-3D) autoencoder. The DSA-3D autoencoder enables up-sampling and down-samping the sub-aperture sequence along the angular resolution or spatial resolution, thereby avoiding the artifacts caused by directly compressing plenoptic image and achieving better compression efficiency. In addition, we propose a special and efficient Square rearrangement to generate sub-aperture sequence. We compare Square with Zigzag sub-aperture sequence rearrangements, and analyzed the compression efficiency of block image compression and whole image compression. Compared with traditional hybrid encoders HEVC, JPEG2000 and JPEG PLENO(WaSP), the proposed DSA-3D(Square) autoencoder achieves a superior performance in terms of PSNR metrics.
Tingting Zhong, Xin Jin 0002, Kedeng Tong
VCIP2
2020 Parallax Tolerant Light Field Stitching for Hand-Held Plenoptic Cameras
abstract
Light field (LF) stitching is a potential solution to improve the field of view (FOV) for hand-held plenoptic cameras. Existing LF stitching methods cannot provide accurate registration for scenes with large depth variation. In this paper, a novel LF stitching method is proposed to handle parallax in the LFs more flexibly and accurately. First, a depth layer map (DLM) is proposed to guarantee adequate feature points on each depth layer. For the regions of nondeterministic depth, superpixel layer map (SLM) is proposed based on LF spatial correlation analysis to refine the depth layer assignments. Then, DLM-SLM-based LF registration is proposed to derive the location dependent homography transforms accurately and to warp LFs to its corresponding position without parallax interference. 4D graph-cut is further applied to fuse the registration results for higher LF spatial continuity and angular continuity. Horizontal, vertical and multi-LF stitching are tested for different scenes, which demonstrates the superior performance provided by the proposed method in terms of subjective quality of the stitched LFs, epipolar plane image consistency in the stitched LF, and perspective-averaged correlation between the stitched LF and the input LFs.
Xin Jin 0002, Qionghai Dai
IEEE Trans. Image Process.1
2019 Light Field Image Compression Using Depth-based CNN in Intra Prediction
abstract
Recently, light field images have received extensive attention due to their potential applications. Since they take up a huge memory because of its super-high resolution, efficient compression methods are fundamentally required. In this paper, we propose a novel intra prediction mode by using depth-adaptive convolutional neuro network (DCNN). Light field projection finds the imaging response distribution for each object point using the depth estimated from each macropixel in the light field image. The highly correlated imaging responses are used to select the neural network structure. The network structure also adapts to the to-be-encoded block size. Adding the proposed DCNN-based prediction mode into the rate-distortion optimization loop with other 35 intra prediction modes of HEVC, the proposed encoding scheme achieves a significant bit-rate saving compared to representative compression approaches with limited computational complexity increment. Statistical data are also provided and analyzed to demonstrate the efficiency of the proposed method.
Tingting Zhong, Xin Jin 0002, Lingjun Li, Qionghai Dai
ICASSP2
2019 Light Field Stitching Targeting Focal Length Inconsistency
abstract
Focal length inconsistency is a common problem in capturing multiple light fields (LFs), which reduces the accuracy in feature matching and LF registration, and limits the quality of LF stitching. In this paper, a novel method is proposed to stitch the LFs captured at different focal length accurately. First, refocus matching is proposed to find the best-matched-focused regions in the focal stack to improve the quality of feature matching. Then, depth distances and refocusing features are exploited to register central sub-aperture images (CSI) to handle focal length variation. Finally, CSI-based 4D warping is applied to preserve the LF angular consistency. Experimental results show that the proposed method outperforms the existing approaches obviously in terms of the stitched LFs and the coherence in epipolar plane image (EPI), especially for the LFs with the focal length inconsistency.
Xin Jin 0002, Qionghai Dai
ICIP2
2019 F-Number Adaptation for Maximizing the Sensor Usage of Light Field Cameras
abstract
Since the effective imaging area in micro lens array (MLA) of the light field camera cannot completely cover the MLA plane, the imaging response of MLA cannot fully cover the sensor plane, which decreases the sensor usage. Hence, an f-number adaptation model is proposed to maximize the sensor usage without changing the structure of light field cameras. By deriving the relationship between the main lens f-number, the macropixel diameter and the sensor usage, a constrained concave optimization model is proposed. A pixel extraction method is also proposed to make a full use of the increased effective pixels. The experimental results demonstrate the effectiveness of the proposed model in terms of the effective pixel number, the subaperture number and improving 3D reconstruction quality. It is also robust to different light field camera architectures and different arrangements of MLA.
Chuanpu Li, Xin Jin 0002, Junke Li, Qionghai Dai
ICME2
2019 Blind Calibration for Focused Plenoptic Cameras
abstract
Because of the subtle structure of focused plenoptic camera, the exact geometry parameters cannot be retrieved, which leads to inaccurate f-number matching or refocusing errors in light field processing. In this paper, a novel blind calibration method is proposed to calculate the geometry parameters for the focused plenoptic cameras. The blind calibration model is derived based on the geometry projection analysis to establish the relationship between the patch-size of each micro-image used in subaperture image rendering and geometry parameters. A triple-level calibration board is designed to realize calibration via single shot. A gradient-SSIM-based fractional-pixel matching is proposed to retrieve the precise rendering patch-size for the calibration model. Experimental results demonstrate that the proposed method is robust to different focused plenoptic cameras and can get the geometry parameters with high accuracy.
Xufu Sun, Xin Jin 0002, Yanqin Chen, Qionghai Dai
ICME2
2019 Live Demonstration: 4-DoF Parallax Tolerant Light Field Stitching
abstract
This demonstration shows a 4-Degree-of-Freedom (4-DoF) parallax tolerant light field (LF) stitching system to generate large field of view (FoV) LF with interactive LF experiences. It can capture the LFs in real-time, generate the large FoV LF with high visual quality and visualize the advantage of large FoV LF by interactive 3D applications like 4-DoF stereo panorama, digital refocusing and 3D point cloud reconstruction.
Xin Jin 0002, Qionghai Dai
ISCAS2
2019 Macropixel-constrained Collocated Position Search for Plenoptic Video Coding
abstract
The plenoptic video recording light fields varying with time has the large-range motion and complex macropixel structure, which brings a great challenge to efficient compression. Based on the analysis of the relationship between temporal motion and macropixel arrangement, an efficient motion estimation algorithm is proposed for plenoptic video compression. It finetunes motion vector predictor (MVP) by the macropixel-constrained collocated position of the current prediction unit in the nearest macropixel, and uses all the macropixel-constrained collocated positions, derived from the finetuned MVP, as motion search candidates for better complexity and compression-efficiency trade-off. The experimental results demonstrate that the proposed algorithm outperforms multi-view based compression and the pseudovideo based compression by an average of 56.40% and 78.72% bitrate reduction, respectively. Compared with HEVC, the proposed method can also save 7.65% bitrate.
Lingjun Li, Xin Jin 0002
VCIP2
2019 Frequency Descriptor based Light Field Depth Estimation
abstract
Depth estimation plays an important role in light field data processing. However, conventional focus measurement based approaches fail at the angular patches containing occlusion boundaries. In this paper, a novel depth estimation algorithm is proposed based on frequency descriptors. On the basis of the imaging process analysis, we propose to first perform the occlusion discrimination and edge orientation extraction in the frequency domain for the spatial patch from the central sub-aperture image. Then, according to the occlusion orientation, a variable-block-size angular patch is selected in the normal direction to construct the frequency descriptors for focus measurement in the focal stack. Experimental results demonstrate superior performance of the proposed method in robustness and depth accuracy.
Junke Li, Xin Jin 0002
VCIP2
2018 High-Speed Light Field Image Formation Analysis Using Wavefield Modeling with Flexible Sampling
abstract
Understanding the image formation inside plenoptic cameras is significant for the investigations of improving the low spatial resolution. Most researches explore the image formation from the perspective of geometric optics. However, as the hardware components in combination with low-aperture optical systems become smaller and smaller, geometric analysis will no longer be valid due to diffraction effects. In this paper, a wave-optic-based model is proposed that uses the Fresnel diffraction equation to propagate the whole object field into the plenoptic systems. The proposed model employs averaging of intensities on the sensor from uncorrelated coherent wave to avoid interference during propagations among the optical component planes. Besides, by utilizing the method of multiple partial propagations, the proposed model is much flexible at sampling on propagation planes. In order to verify the effectiveness of the proposed model, numerical simulations are conducted by comparing with existing wave optic model under different optical configurations of plenoptic cameras. Results demonstrate that the proposed model can describe the light field image formation properly. In addition, the time for image formation has been reduced by a factor of 19.22 using the proposed model.
Yanqin Chen, Xin Jin 0002, Qionghai Dai
ICASSP2
2018 Light Field Stitching for Parallax Tolerance
abstract
In this paper, a novel light field (LF) stitching method is proposed to handle parallax more flexibly and accurately. The depth guided feature point filtering and depth-based motion model are proposed to warp the 4D meshes in the LFs adapting to the relative depth layer distance between the mesh center and the feature points. Also, 4D graph-cut is applied in LF fusion to reduce ghosting effects and produce better refocusing effects. Experimental results show that the proposed method works well without producing visual artifacts like ghosting or misalignments. It outperforms the existing approaches obviously in terms of the subjective quality of the stitched LFs, RMSE and the coherence in epipolar plane image (EPI).
Xin Jin 0002, Chuanpu Li, Yanqin Chen, Qionghai Dai
ICIP2
2018 Plenoptic Image Coding Using Macropixel-Based Intra Prediction
abstract
The plenoptic image in a super high resolution is composed of a number of macropixels recording both spatial and angular light radiance. Based on the analysis of spatial correlations of macropixel structure, this paper proposes a macropixel-based intra prediction method for plenoptic image coding. After applying an invertible image reshaping method to the plenoptic image, the macropixel structures are aligned with the coding unit grids of a block-based video coding standard. The reshaped and regularized image is compressed by the video encoder comprising the proposed macropixel-based intra prediction, which includes three modes: multi-block weighted prediction mode (MWP), co-located single-block prediction mode (CSP), and boundary matching based prediction mode (BMP). In the MWP mode and BMP mode, the predictions are generated by minimizing spatial Euclidean distance and boundary error among the reference samples, respectively, which can fully exploit spatial correlations among the pixels beneath the neighboring microlens. The proposed approach outperforms HEVC by an average of 47.0% bitrate reduction. Compared with other state-of-the-art methods, like pseudo-video based on tiling and arrangement method (PVTA), intra block copy (IBC) mode, and locally linear embedding (LLE) based prediction, it can also achieve 45.0%, 27.7% and 22.7% bitrate savings on average, respectively.
Xin Jin 0002, Haixu Han, Qionghai Dai
IEEE Trans. Image Process.1
2018 Depth Assisted Adaptive Workload Balancing for Parallel View Synthesis
abstract
Depth image-based rendering has been adopted by MPEG as the recommended view synthesis technique for free viewpoint TV applications. In this paper, a workload balancing algorithm is proposed for parallel view synthesis on multicore platforms. First, view synthesis workload is defined as the function of the number of hole-pixels in the warped images. Then, a novel depth assisted prediction method is proposed to predict the number of hole-pixels in the current frame by exploiting the depth differences between the neighboring frames, which reflects the movement of objects in video content. Feeding the predicted workload to the proposed cost function, each input frame is partitioned adaptively to balance the synthesis workload among the cores. The proposed workload prediction method outperforms the existing approaches both in terms of frame average prediction error and standard deviation in prediction error. Applying the proposed workload balancing method, the parallel view synthesis system provides higher acceleration ratio and better synchronization performance among the cores compared with other parallel processing systems without sacrificing the subjective and objective quality. It is also robust to different platforms, which shows high potential in being applied to mobile oriented applications.
Xin Jin 0002, Zhanqi Liu, Qionghai Dai
IEEE Trans. Multim.1
2017 Dynamic cloud Offloading for View Synthesis
abstract
In this paper, a dynamic offloading model is proposed to minimize the energy consumption of mobile devices by exploiting cloud computational resources for view synthesis. The computational complexity of view synthesis, the processing capability of the cloud, the processing capability and the power consumption of the mobile are considered jointly into the model to provide an optimized solution. Several simulations based on parameters of real mobile devices demonstrate that the proposed method can save an average of 42.59%(4-partition case) and 46.40%(8-partition case) of total energy on different mobile devices and an average of 67.58%(4-partition case) and 69.76%(8-partition case) of total energy under different transmitting rates than the existing algorithms for view synthesis, respectively.
Xin Jin 0002, Zhanqi Liu, Qionghai Dai
ICASSP2
2017 Enhanced depth estimation for hand-held light field cameras
abstract
Conventional depth estimation methods are confined by the occluder and homogenous regions in the scene. In this paper, we propose a new depth estimation and enhancement method. The raw depth is calculated from analyzing the Consistency Metric Range (CMR) in the angular patch. Confident depth map is obtained by analyzing the variation of CMR within a neighborhood around the lowest CMR curves. Confident depth points are propagated to the whole image by global optimization with weighted neighborhood smoothness, gradient and second derivative constraints. Finally, depth is enhanced by using weighted median filter. The experimental results demonstrated the effectiveness of the proposed approach in providing much clearer transitions of texture regions and much smoother homogenous regions after depth propagation.
Yanwen Qin, Xin Jin 0002, Yanqin Chen, Qionghai Dai
ICASSP2
2017 Lenslet image compression using adaptive macropixel prediction
abstract
In this paper, an efficient compression method is proposed for lenslet images captured by plenoptic cameras for recording the spatial and angular light information at a super-high-resolution. After applying a reversible image reshaping method to the lenslet image, a reshaped and regularized image will be generated and compressed by the video codec comprising the proposed adaptive macropixel prediction mode. Based on the analysis of spatial correlations among adjacent macropixels, two spatial prediction modes are proposed as: multi-block weighted prediction mode and co-located single-block prediction mode, to predict the coding unit by minimizing the coding cost. The multi-block weighted prediction is formulated by minimizing the Euclidean distance between the coding unit and co-located blocks in the macropixel structure. Performance evaluations have shown that the proposed method achieves 50.9% of bit-savings on average compared to HEVC. It also outperforms state-of-the-art coding methods drastically.
Haixu Han, Xin Jin 0002, Qionghai Dai
ICIP2
2017 Lenslet image compression based on image reshaping and macro-pixel Intra prediction
abstract
Lenslet images that record both spatial and angular light radiance in a super high definition with distinct macropixel structures desire efficient compression methods for promoting the applications of handheld plenoptic cameras urgently. In this paper, a lenslet image compression method is proposed. First, a reversible image reshaping and adaptive interpolation is proposed to align the macropixel structures with coding unit grids in the block based video coding standards. Then, based on the reshaped and regularized lenslet images, a macro-pixel Intra prediction mode, in which the coding unit is predicted by minimizing spatial boundary error among the adjacent macropixels, is proposed to fully exploit spatial correlations among the pixels beneath the neighboring microlens. The proposed approach outperforms HEVC by an average of 35.56% bitrate reduction. Compared with the existing coding approaches, like Intra block coding (IBC) and locally linear embedding-based (LLE) prediction, it achieves an average of 17.72%/23.06% bitrate reduction, which demonstrates its efficiency explicitly.
Haixu Han, Xin Jin 0002, Qionghai Dai
ICME2
2017 Non-invasive imaging based on speckle pattern estimation and deconvolution
abstract
Non-invasive imaging through scattering media is still a challenging task, especially when the imaging target is complex. This paper presents a novel imaging method through scattering media based on speckle pattern estimation and deconvolution. Different from previous frameworks based on PSF-adjusting or PSF-capturing that need to invasively put a point source at the imaging area, we utilize the non-invasive imaging system based on speckle scanning and memory effect. Phase retrieval results of simple targets behind the scattering media are used as the input of the proposed speckle pattern estimation model, in which speckle modeling and constrained least square optimization are applied to estimate the distribution of speckle pattern. The estimated speckle pattern is exploited in deconvoluting the integrated intensity matrices (IIMs) of the scattered images to recover the complex targets. Experimental results show that the proposed method can recover the imaging targets behind the scattering media much more accurate and stable than the existing phase retrieval methods.
Zhouping Wang, Xin Jin 0002, Yifu Hu, Qionghai Dai
VCIP2
2017 VideoSet: A large-scale compressed video quality dataset based on JND measurement
abstract
• A large-scale JND-based coded video quality dataset is presented. • The VideoSet contains 220 5-s sequences in four resolutions coded by H.264/AVC. • The subjective test procedure, JND data cleaning and properties are described. • The significance and implications of the VideoSet are discussed. • This work points out a clear path to data-driven perceptual coding. A new methodology to measure coded image/video quality using the just-noticeable-difference (JND) idea was proposed in Lin et al. (2015). Several small JND-based image/video quality datasets were released by the Media Communications Lab at the University of Southern California in Jin et al. (2016) and Wang et al. (2016) [3]. In this work, we present an effort to build a large-scale JND-based coded video quality dataset. The dataset consists of 220 5-s sequences in four resolutions (i.e., 1920 × 1080 , 1280 × 720 , 960 × 540 and 640 × 360 ). For each of the 880 video clips, we encode it using the H.264/AVC codec with QP = 1 , … , 51 and measure the first three JND points with 30 + subjects. The dataset is called the “VideoSet”, which is an acronym for “Video Subject Evaluation Test (SET)”. This work describes the subjective test procedure, detection and removal of outlying measured data, and the properties of collected JND data. Finally, the significance and implications of the VideoSet to future video coding research and standardization efforts are pointed out. All source/coded video clips as well as measured JND data included in the VideoSet are available to the public in the IEEE DataPort (Wang et al., 2016 [4]).
Haiqiang Wang, Ioannis Katsavounidis, Jiantong Zhou, Jeong-Hoon Park, Shawmin Lei, Xin Zhou 0001, Man-On Pun, Xin Jin 0002, Ronggang Wang, Xu Wang 0006, Yun Zhang 0002, Jiwu Huang, Sam Kwong, C.-C. Jay Kuo
J. Vis. Commun. Image Represent.8
2017 Depth map super-resolution via low-resolution depth guided joint trilateral up-sampling
Xin Jin 0002, Yangguang Li 0001, Chun Yuan 0003
J. Vis. Commun. Image Represent.2
2016 Depth fused from intensity range and blur estimation for light-field cameras
abstract
Light-field cameras attract great attention because of its refocusing and perspective-shifting functions after capturing. The special 4D-structured data contains depth information. In this paper, a novel depth estimation algorithm is proposed for light-field cameras by fully exploiting the characteristics of 4D light-field data. A novel tensor, intensity range of pixels within a microlens, is proposed, which presents strong correlation with the transition on focus, especially for texture-complex regions. Meanwhile, the other tensor, defocus blur amount is utilized to estimate the focus level, which generates more accurate depth estimation especially for homogeneous regions. Then, the depths calculated from the two tensors are fused according to the variation scale of intensity range and the minimal defocus blur amount under spatial smoothness constraints. Compared with the representative approaches, the depth generated by the proposed approach presents richer details for texture regions and higher consistency for unified regions.
Yatong Xu, Xin Jin 0002, Qionghai Dai
ICASSP2
2016 Imaging through scattering media with intensity modulated incoherent sources
abstract
Imaging through scattering media is a significant challenge in computational imaging. Recently, a breakthrough technique based on speckle scanning was proposed with outstanding imaging performance. However, the dense angular scanning of the incident laser beam leads to a lengthy scanning process. In this paper, we propose a method based on compressive sensing (CS) to accelerate the data acquisition process. A new imaging system is proposed which adopts mutually incoherent laser beams as light sources and the intensity of each beam is modulated to perform CS measurement. High-resolution fluctuation of the total fluorescence can be well reconstructed from CS measurements with the learned dictionary as basis function. Experimental results show that the proposed method can reduce nearly 70 percent of the data acquisition complexity without impairing the imaging quality.
Yifu Hu, Xin Jin 0002, Kaiyun Wei, Qionghai Dai
ICIP2
2016 A SVR based quality metric for depth quality assessment
abstract
A depth map generally can be divided into the region with sharp edges and the region consisting of nearly constant or slowly varying samples. In this paper, in order to investigate how the distortion in the two regions affect the perceived 3D quality of synthesized stereopairs, a dataset is first built based on distinctive coding for two regions, and then a subjective test is conducted. Based on the subjective evaluation results, a support vector regression (SVR) based model is built to estimate the perceived 3D quality of the synthesized stereopairs with the features extracted from the depth maps. As a metric for depth quality assessment, the proposed model outperforms the conventional 2D QA metrics applied to the depth maps as well as that applied to the stereopairs.
Xin Jin 0002, Qionghai Dai
ISCAS2
2016 Efficient imaging through scattering media by random sampling
abstract
Imaging through scattering media is a tough task in computational imaging. A recent breakthrough technique based on speckle scanning was proposed with outstanding imaging performance. However, to achieve high imaging quality, dense sampling of the integrated intensity matrix is needed, which leads to a time-consuming scanning process. In this paper, we propose a method that exploits spatial redundancy of the integrated intensity matrix and reconstructs the complete matrix from few random samples. A reconstruction model that jointly penalizes total variation and weighted sum of nuclear norm of local patches is built with improved reconstruction quality. Experiments are performed to verify the effectiveness of the proposed method and results demonstrate that the proposed method can achieve a same imaging quality with 80% reduction of the data acquisition complexity.
Yifu Hu, Xin Jin 0002, Qionghai Dai
MMSP2
2016 A 3D subjective quality prediction model based on depth distortion
abstract
Depth map quality plays an important role in 3D subjective quality. In this paper, a 3D subjective quality prediction model is proposed to estimate the 3D quality of synthesized stereopairs based on depth map distortion and neural mechanism, instead of performing view synthesis directly, which benefits 3D processing. In order to build the model, a dataset is first built to include distinctive distortion features for depth map coding, and then a subjective test is conducted. Based on the subjective evaluation results, a prediction model is built to estimate the perceived 3D quality of the synthesized stereopairs with the features extracted from the texture characteristics of decoded depth maps and neural population coding model. In terms of the correlation coefficient with actual 3D subjective quality, the proposed model outperforms the conventional 2D QA metrics applied to the depth maps as well as that applied to the stereopairs.
Xin Jin 0002, Qionghai Dai
VCIP2
2016 Depth dithering based on texture edge-assisted classification
Xin Jin 0002, Yatong Xu, Qionghai Dai
Signal Process. Image Commun.1
2016 Clustering-Based Content Adaptive Tiles Under On-chip Memory Constraints
abstract
Tiles have been introduced to the next generation video coding standard, high-efficiency video coding (HEVC) standard, as a fundamental tool to reduce on-chip memory requirement during encoding and decoding high-definition video. In this paper, a content adaptive tile partitioning approach is proposed to improve the compression efficiency for HEVC under the on-chip memory constraint. Local competition optimization-based rectangular clustering is proposed to partition the frames into a required number of tiles adapting to content variations. Under the same memory constraint, the adaptive scheme improves compression efficiency by up to 1.8% bitrate saving relative to uniformly spaced tiles with negligible complexity increment. It especially benefits the videos with regional high spatial correlations and no penalty in compression efficiency is observed for other types of videos.
Xin Jin 0002, Qionghai Dai
IEEE Trans. Multim.1
2015 A workload balanced parallel view synthesis for FTV
abstract
In this paper, a parallel system together with an adaptive workload balancing algorithm is proposed for view synthesis on multi-core platforms. Based on system level data parallelism, an adaptive workload balancing method is proposed for depth image based rendering by evaluating the number of non-hole pixels after warping. Experimental results demonstrated that with the proposed workload balancing algorithm, the workload difference among the cores is reduced by 90.65% on average for 2-core systems and by 79.57% on average for 4-core systems, respectively. Compared with the parallel system without the proposed balancing algorithm, synthesis speed is further improved by 7.5% for 2-core systems and 8.9% for 4-core systems at maximum, respectively, without degradation in the subjective and objective quality.
Zhanqi Liu, Xin Jin 0002, Qionghai Dai
ICASSP2
2015 Depth estimation by analyzing intensity distribution for light-field cameras
abstract
In this paper, a novel depth estimation algorithm is proposed for light-field cameras by fully exploiting the characteristics of 4D light-field data. Based on a standard light-field acquisition system, a novel tensor, intensity range of pixels within a microlens, is proposed, which presents much stronger correlation with the transition on focus. Then, the tendency of intensity range is combined with the texture gradient to further improve the accuracy and consistency of the estimated depth and preserve edges through global optimization. Compared with the representative approaches, the depth generated by the proposed approach presents clearer object boundaries with much richer details both for indoor and outdoor contents. Moreover, the execution time is on average 12 times lower than the existing approaches.
Yatong Xu, Xin Jin 0002, Qionghai Dai
ICIP2
2015 Non-invasive imaging based on sparse representation
abstract
In this paper, we propose a method based on sparse representation to accelerate the scanning technique to perform non-invasive imaging through scattering layers. The scanning time is proportional to the size of the collected integrated intensity pattern and it usually takes tens of hours to finish the collecting process. To speed up the scanning technique, only a much smaller integrated intensity pattern is collected. A training set of integrated intensity pattern pairs with sizes corresponding to the necessary integrated intensity pattern that makes the scanning technique work well and the much smaller one is constructed. And a pair of dictionaries is trained from the constructed set exploiting the K-SVD algorithm. Based on sparse representation, the necessary integrated intensity pattern can be recovered from the much smaller one using the trained dictionaries, thus realizing non-invasive imaging successfully. Experimental results show that our method can successfully make the scanning time reduced by 8/9 without deteriorating the imaging quality of the scanning technique.
Kaiyun Wei, Xin Jin 0002, Yifu Hu, Qionghai Dai
MMSP2
2015 Region adaptive workload prediction for parallel view synthesis
abstract
In this paper, a parallel system together with a real-time workload balancing algorithm is proposed for view synthesis on multi-core platforms. First, a numerical relationship between the number of holes after warping and the workload of view synthesis is derived based on correlation analysis for the texture regions. Then, according to the location of the holes (whether they lie around an edge or at a homogeneous region), different models are proposed to predict the synthesis workload accurately. Experimental results show that the workload difference among cores is reduced largely and higher speedup ratio is achieved with negligible quality degradation by the proposed workload balancing system.
Zhanqi Liu, Xin Jin 0002, Qionghai Dai
VCIP2
2015 A fast encoder of frame-compatible format based on content similarity for 3D distribution
Xin Jin 0002, Zhuoying Zeng, Satoshi Goto, Qionghai Dai
Signal Process. Image Commun.1
2014 A novel distortion model for depth coding in 3D-HEVC
abstract
In 3D-HEVC, the latest 3D video coding project of MPEG, the coding mode of depth maps is determined by view synthesis optimization (VSO) which integrates view synthesis distortion into rate distortion optimization (RDO) to improve the compression efficiency and synthesis quality simultaneously. However, it introduces a big burden in coding complexity of depth maps. In this paper, a novel distortion model for depth coding in 3D-HEVC is proposed which estimates the distortion of synthesized views using an adaptive model for different pixel intervals. It outperforms existing depth distortion estimation algorithms by 16.6% BD-BR saving in max and 10.2% BD-BR saving on average. Compared with VSO, it reduces encoding complexity of depth maps by 46.2% with the highest efficiency-complexity performance among existing depth distortion estimation algorithms.
Xin Jin 0002, Qionghai Dai
ICIP2
2014 A fast coding algorithm based on inter-view correlations for 3D-HEVC
abstract
The newly published 3D-HEVC has received a remarkable response due to its high compression efficiency which is based on High Efficiency Video Coding (HEVC). However, the complexity of its encoding process is also large as a result of introducing the coding units (CU) size decision process together with the rate distortion optimization (RDO) process. In this paper, a fast coding algorithm making good use of the interview correlations is proposed. With the inter-view correlation statistical analysis, the CU depth candidates of the dependent views can be predicted from the independent view instead of the brute force RDO process in determining CU depth. The experimental results show that the proposed method saves 51% time in texture coding and the loss is negligible.
Guangsheng Chi, Xin Jin 0002, Qionghai Dai
VCIP2
2014 A fast mode decision algorithm applied to Coarse-Grain quality Scalable Video Coding
Chun Yuan 0003, Yu Sun 0003, Jian Zhang 0002, Xin Jin 0002
J. Vis. Commun. Image Represent.5
2013 Retinex based visual identicalness detection for videos corrupted by imaging noise
Xin Jin 0002, Satoshi Goto, Qionghai Dai
Signal Process. Image Commun.1
2012 Motion robust rain detection and removal from videos
abstract
Weather such as rain and snow cause difficulties in processing the videos captured. Since the appearance of rain drops can affect the performance of human tracking and reduce the efficiency of video compression, detection and removal of rain is a challenging problem in outdoor surveillance systems. In this paper, we propose a new algorithm for rain detection, which is based on joint spatial and wavelet domain features. This approach is robust to the videos with moving objects in the rain. Experimental results demonstrated its better performance in comparison with the existing approaches in the subjective quality.
Xinwei Xue, Xin Jin 0002, Satoshi Goto
MMSP2
2012 A 18.42 times faster video encoder based on Retinex theory
abstract
Because of the limitation in the imaging system, videos captured by mobile devices are usually noisy. In this paper, a noise robust fast video encoding scheme is proposed based on Retinex theory. The reflectance image is estimated in a co-analysis module by Retinex theory, within which the content similarity is detected to conduct the encoding process. Experimental results demonstrate that the proposed encoding scheme provides an average of 18.42 times higher encoding speed relative to the existing approaches without sacrificing the subjective quality.
Xin Jin 0002, Satoshi Goto
PCS1
2012 Hilbert Transform-Based Workload Prediction and Dynamic Frequency Scaling for Power-Efficient Video Encoding
abstract
With the popularity of mobile devices with embedded video cameras, real-time video encoding on hand-held devices becomes increasingly popular. Reducing the power consumption during real-time video encoding to suspend the battery life with the same encoding performance is very important to improve the quality of service. Although some workload estimation techniques have been developed for video decoding to reduce power consumption for video playback applications, they present inefficiency in being transferred to video encoding directly because the compressed information cannot be retrieved before encoding and the future input video content is often nondeterministic. In this paper, a workload estimation scheme targeting video encoding applications is proposed. Based on the definition of video encoding workload and the analysis of the features, a Hilbert transform-based workload estimation model is proposed to predict the overall variation trend in the encoding workload to overcome the workload fluctuations and the nondeterministic content variations, e.g., burst motion. The effectiveness of the proposed algorithm is demonstrated on two H.264/AVC encoders on PC and an embedded platform by encoding different video contents at different bit-rates. The proposed algorithm provides a negligible deadline missing ratio around 4.8%, which is much lower than the previous solutions, together with platform and content robustness. Compared with the previous solutions, the proposed algorithm provides up to 61.69% power reduction under the same performance constraint.
Xin Jin 0002, Satoshi Goto
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2011 A 98 GMACs/W 32-core vector processor in 65nm CMOS
Dajiang Zhou, Xin Jin 0002, Satoshi Goto
ISLPED3
2011 A Novel Depth-Image Based View Synthesis Scheme for Multiview and 3DTV
Xin Jin 0002, Satoshi Goto
MMM (1)2
2011 Adaptive raster scan for slice/frame coding
abstract
In this paper, new raster scan orders other than the conventional horizontal scan from the top-left corner of the image/slice are proposed to fully exploit content correlation from the regions right to and below the current coding unit. By being applied to each slice/frame adaptively with the conventional one, features influencing the scan order selection are investigated. Different adaptive scan schemes are tested and compared. Up to 3.4% bit-rate reduction can be achieved for Intra coding with limited increase in complexity.
Xin Jin 0002, Satoshi Goto
VCIP1
2011 Low power parallel surveillance video encoding system based on joint power-speed scheduling
abstract
In this paper, a low power parallel surveillance video encoding system based on joint power-speed scheduling is proposed. The relative relationships among the CPU statuses, total power consumption and encoding speed are analyzed and modeled for multi-core processors. Based on the power directional graph and the relative encoding speed model, the working statuses of the cores are controlled jointly adapting to video encoding workload to minimize the total power consumption. It provides more than 20% power reduction compared with the latest existing system without a penalty on the encoding speed.
Xin Jin 0002, Satoshi Goto
VCIP1
2011 A fast encoder of frame-compatible format based on content similarity for 3D distribution
abstract
Through packing the two neighboring views after down-sampling into one frame, frame-compatible format coding allows stereo video to be encoded on the conventional video applications. Using the content similarity which leads to similar prediction between the two packed planes, a fast algorithm is discussed in this paper. The target of the proposed algorithm is to reduce the computational complexity and maintain the compression performance. By operating the candidate predictions solely on the half of a frame (plane_1) according to the encoded half of the frame (plane_0), it introduces an average of 75.7% and 90.4% reduction in planet's computational complexity for Intra and Inter encoding, respectively. The statistic analysis illustrates the prediction correlation due to the content similarity, and contributes to the prediction candidate sets for Intra and Inter coding. In addition, a shift obtaining method integrated in the fast algorithm is also introduced.
Zhuoying Zeng, Xin Jin 0002, Satoshi Goto
VCIP2
2011 Encoder adaptable difference detection for low power video compression in surveillance system
Xin Jin 0002, Satoshi Goto
Signal Process. Image Commun.1
2011 Composite Model-Based DC Dithering for Suppressing Contour Artifacts in Decompressed Video
abstract
Because of the outstanding contribution in improving compression efficiency, block-based quantization has been widely accepted in state-of-the-art image/video coding standards. However, false contour artifacts are introduced, which result in reducing the fidelity of the decoded image/video especially in terms of subjective quality. In this paper, a block-based decontouring method is proposed to reduce the false contour artifacts in the decoded image/video by automatically dithering its direct current (DC) value according to a composite model established between gradient smoothness and block-edge smoothness. Feature points on the model with the corresponding criteria in suppressing contour artifacts are compared to show a good consistency between the model and the actual processing effects. Discrete cosine transform (DCT)-based block level contour artifacts detection mechanism ensures the blocks within the texture region are not affected by the DC dithering. Both the implementation method and the algorithm complexity are analyzed to present the feasibility in integrating the proposed method into an existing video decoder on an embedded platform or system-on-chip (SoC). Experimental results demonstrate the effectiveness of the proposed method both in terms of subjective quality and processing complexity in comparison with the previous methods.
Xin Jin 0002, Satoshi Goto, King Ngi Ngan
IEEE Trans. Image Process.1
2010 Hilbert transform based workload estimation for low power surveillance video compression
abstract
In this paper, a workload estimation scheme is proposed for surveillance video encoding by using difference detection and Hilbert transform-based workload estimation model. Difference detection distributes the input video data according to their content similarity features and retrieves the encoding workload for the coded frames. Workload estimation model predicts the encoding workload for the following time slot using Hilbert transform and error control. Experimental results indicate that the proposed workload estimation model can provide accurate estimation results without performance hit to maintain the video encoding performance under a given performance constraint during power reduction.
Xin Jin 0002, Satoshi Goto
ICIP1
2010 Low power surveillance video coding system
abstract
In this demo, a low power H.264/AVC video encoding system is demonstrated for surveillance video compression, which is based on our proposed algorithms as difference detection and Hilbert transform based workload estimation. Two surveillance video coding schemes (with and without our proposed methods, respectively) are implemented on multi-core embedded platform and compared side by side. The power consumption of each encoding scheme is measured during video encoding and displayed explicitly to show a more than 60% power reduction introduced by the proposed methods. The reconstructed video is also synchronized and displayed as the video is encoded to present a maximum of 50 times higher encoding speed.
Xin Jin 0002, Kun Ba, Satoshi Goto
ICME1
2010 An adaptive bandwidth reduction scheme for video coding
abstract
In high definition video decoders for the standard like H.264, bandwidth requirement is a critical design issue due to the overwhelming amount of memory data access. This paper proposes a new adaptive reference frame compression scheme to reduce external memory bandwidth consumption. An efficient variable length coding is proposed to compress each processing unit efficiently. Moreover, an adaptive compression mode decision unit is proposed to adaptively choose the best compression mode according to the image characteristic and bandwidth requirement. As a result, the proposed scheme achieves efficient bandwidth reduction with little image quality loss. The experimental results show that, at the same bandwidth reduction ratio the proposed algorithm achieves up-to 0.7 dB gain compared to the existing approaches.
Liu Song, Dajiang Zhou, Xin Jin 0002, Satoshi Goto
ISCAS3
2010 Difference detection based early mode termination for depth map coding in MVC
abstract
Depth map coding is a new topic in multiview video coding (MVC) following the development of depth-image-based rendering (DIBR). Since depth map is monochromatic and has less texture than color map, fast algorithm is necessary and possible to reduce the computation burden of the encoder. This paper proposed difference detection based early mode termination strategy. The difference detection (DD) algorithms are categorized to reconstructed frame based (RDD) and original frame based (ODD). A simplified ODD (sODD) strategy is also proposed. Early mode termination based on these three DD algorithms are implemented and evaluated in the reference software of Joint Multiview Video Coding (JMVC) version 8.0 respectively. Simulation results indicate that RDD based one has no performance lost and reduce 25% runtime on average. ODD and sODD based ones can save 54.3% and 43.6% runtime respectively and have an acceptable R-D performance lost.
Xin Jin 0002, Satoshi Goto
PCS2
2009 Composite modeling of optical flow for artifacts reduction
abstract
Because of the outstanding contribution in removing pixel correlation, block-based transform and quantization has been widely accepted in the state-of-art image/video coding standards. However, artifacts are introduced, which result in reducing the subjective quality of the decoded image/video. In this paper, a composite model of the optical flow velocity is mathematically derived and proved to further reduce the blocking artifacts by compensating the DC surface of the decoded image according to the model. The functional relationship among the optical flow velocity of the DC surface of the decoded image, the value discontinuity across the block boundary in the decoded image and the quantity of DC surface adjustment is analyzed and mathematically modeled as a composite function with second-degree polynomial. Four feature points on the model with the corresponding criteria in artifacts reduction are compared to show a good consistency between the model and the actual filtering effects. Finding a point on the model with a better trade-off between the optical flow smoothness and the edge smoothness is promising in producing a more pleasing and natural looking image/video.
Xin Jin 0002, Satoshi Goto, King Ngi Ngan
ICME1
2009 Optical flow based DC surface compensation for artifacts reduction
abstract
The block-based prediction, transform and quantization reduce the redundancy in video data efficiently. However, the blocking artifacts are introduced because of quantization error and prediction scheme. Even if the de-blocking filter is applied like that designed in H.264 and AVS, the blocking artifacts are still visible especially in videos with smooth area coded at low bit-rate. In this paper, an optical flow based technique is proposed to further reduce the blocking artifacts by compensating the DC surfaces of the decoded images using joint optimization of optical flow and edge difference. The proposed algorithm automatically classifies the macroblocks in the decoded images according to their texture feature and endures the smooth area with smooth brightness variation by minimizing the joint function of optical flow magnitude and edge smoothness. It produces a more pleasing image without over-filtering, which can be further applied as an in loop filter for video coding.
Xin Jin 0002, Satoshi Goto, King Ngi Ngan
PCS1
2009 Platform-independent MB-based AVS video standard implementation
Xin Jin 0002, Songnan Li, King Ngi Ngan
Signal Process. Image Commun.1