Jianjun Lei 0001

dblp:09/267-1 · DBLP profile ↗
← Back
109ranked-venue papers
30as first author
56since 2021 · last 2026
0000-0003-3171-7680ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 85 · 25 first-author · 44 since 2021Artificial intelligence and machine learning · 15 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 4 since 2021Computer networks · 7 · 1 first-author · 6 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Lightweight stereo image super-resolution via adaptive pruning and bridge distillation
Zhe Zhang 0041, Bingzheng Liu, Lei Chen 0091, Pengzhi Li, Yidan Zhang 0002, Jianjun Lei 0001
Knowl. Based Syst.6
2026 Spinal Lesion Detection in X-Ray Images via Uncertainty-Guided Classification and Localization
Lisha Guo, Bo Peng 0007, Jianjun Lei 0001, Xu Zhang 0045, Qingming Huang
IEEE Signal Process. Lett.3
2026 Mining Temporal Redundancy Using Long Short-Term Motion Aggregation and Global-Local Decorrelation for Learned Video Compression
abstract
The conditional coding paradigm is widely used in learned video compression, which shows superior performance in capturing redundancies within a large context space. However, existing Conditional coding-based Learned Video Compression (C-LVC) methods ignore that the predicted motion vectors usually contain large uncertainty due to complex motions, occlusions, etc., which consequently decrease the accuracy of the generated temporal contexts. In addition, existing C-LVC methods have a weak ability to mine diverse dependencies within the context space, which are closely related to the coding efficiency. To address these issues, an efficient temporal redundancy mining method is proposed to improve the coding efficiency of C-LVC in this paper. To generate accurate temporal contexts, a Long Short-Term Motion Aggregation (LSTMA) model is proposed, in which an LSTMA-based motion estimation module is developed to capture both current and aggregated long short-term motion information to reduce the uncertainty of predicted motion vectors. Based on the dual motion information, an LSTMA-based temporal context mining module is developed to exploit the aggregated long short-term motion information and increase the accuracy of the generated temporal contexts. In order to fully eliminate spatial-temporal redundancies in a video, a Global-Local Information Decorrelation Module (GLIDM)-based context codec is proposed, in which the GLIDM is designed based on the visual state space block (namely vmamba), the residual block, and the squeeze-and-excitation block to effectively capture long-range, short-range spatial-temporal dependencies and channel-wise dependencies. Experimental results demonstrate that our proposed method can effectively improve the coding performance of C-LVC, and outperforms other state-of-the-art LVC methods.
Zhaoqing Pan, Jianjun Lei 0001, Bo Peng 0007, Haoran Xie 0001, Fu Lee Wang, Sam Kwong
IEEE Trans. Circuits Syst. Video Technol.3
2026 Uncertainty-Aware Multi-View Graph Clustering
abstract
Multi-view clustering has achieved advanced progress over the years, which typically integrates multi-view information to learn discriminative common representations or a unified clustering distribution for clustering. However, existing methods either simply regard each view as equally important or assign fixed weights to each view, which are insufficient to dynamically assess the sample quality variations caused by noise in multi-view data. To address this issue, by effectively modeling the uncertainty of different samples across different views, this paper proposes a novel uncertainty-aware multi-view graph clustering network, termed UMGC-Net, which achieves trusted multi-view clustering in an unsupervised manner. Specifically, by measuring the clustering distribution entropy, an uncertainty-guided common feature learning mechanism is proposed to estimate the uncertainty for each sample of each view, thus learning multi-view features friendly to clustering. Besides, a cross-view trusted distribution fusion module is designed to obtain robust clustering distribution by exploring the trusted consistency among multi-view clustering distributions based on uncertainty. Finally, experimental results on four popular multi-view datasets validate the superior performance of the proposed UMGC-Net.
Bo Peng 0007, Shaobo Bai, Jianjun Lei 0001, Changqing Zhang 0002, Nam Ling
IEEE Trans. Multim.3
2026 Depth-Aware Transformer for Aerial Localization
abstract
Recently, deep learning-based visual localization has gained significant attention and made remarkable advancements. Although previous visual localization methods have obtained promising performance on indoor or outdoor street scenes, there have been few attempts at visual localization on aerial scenes. In this article, a depth-aware aerial localization transformer (DALTR) is proposed to learn camera poses in real-world aerial scenes assisted by the depth map. To improve the ability of network to perceive on aerial scenes, a multi-level depth embedding transformer module is presented by adaptively incorporating depth information into multiple levels of transformer. In addition, to encourage the piece-wise smooth geometric characteristic of the scene coordinates, a depth-guided smoothness constraint is developed to provide additional supervision for scene coordinate regression. Extensive experimental results on aerial localization benchmark datasets demonstrate that the proposed DALTR achieves superior aerial localization performance.
Jianjun Lei 0001, Duohui Tu, Bo Peng 0007, Zhe Zhang 0041, Chong Wu 0004, Qingming Huang
ACM Trans. Multim. Comput. Commun. Appl.1
2026 Meta-Learned Zero-Shot Sketch-Based Point Cloud Retrieval via Perspective-Predicted Feature Learning
abstract
In recent times, sketch-based 3D shape retrieval has emerged as a pivotal theme and garnered considerable attention within the area of cross-modal retrieval. As a prevalent 3D shape modality, the exponential growth in the quantity of 3D point clouds has boosted a substantial increase in the demand for 3D point cloud retrieval. Simultaneously, due to the absence of prior knowledge about unseen classes, transferring models learned from seen classes to tackle the data from unseen classes effectively remains a significant hurdle in cross-modal retrieval. In light of this, a novel meta-learned zero-shot sketch-based point cloud retrieval (MetaZS-SBPR) network is proposed in this article for exploring cross-modal consistent feature representation from 2D sketches and 3D point clouds, while effectively transferring the knowledge from seen classes to unseen classes. Specifically, a perspective-predicted point cloud feature learning module is presented to capture discriminative features of point clouds from predicted perspectives, thereby mitigating the modal differences across point clouds and sketches. Additionally, a meta zero-shot retrieval strategy is introduced to investigate the knowledge transfer from seen classes to unseen classes harnessing meta-learning, thereby enabling the efficient retrieval of the point clouds from unseen classes. Experimental evaluations conducted on the ZS-SBPR benchmark dataset affirm the effectiveness of the proposed MetaZS-SBPR.
Bo Peng 0007, Menglei Zhao, Qingming Huang, Jianjun Lei 0001
ACM Trans. Multim. Comput. Commun. Appl.6
2025 Hypergraph Contrastive Learning for Large-Scale Hyperspectral Image Clustering
abstract
Large-scale hyperspectral image (HSI) clustering has become an important research task owing to its promising applications in various fields. Recently, beneficial from the correlation modeling capability of graphs, graph contrastive learning methods have received increasing attention in the clustering task. However, these methods usually have limited ability to explore the high-order correlation as well as beneficial clustering information of large-scale HSI, thus limiting the clustering performance on large-scale HSI. To this end, a novel hypergraph contrastive learning network (HCL-Net) for large-scale HSI clustering is proposed in this paper. Specifically, a diffusion hypergraph-based contrastive clustering mechanism is presented, in which a diffusion hypergraph is constructed to model the high-order correlation in large-scale HSI, thus guiding contrastive learning for obtaining more discriminative representations. Besides, by mining the confident clustering information, a confidence-guided positive-negative updating strategy is designed to dynamically update positives and negatives for contrastive learning, thereby obtaining a more compact clustering structure. The proposed method is evaluated on three public large-scale HSI datasets. The experimental results have demonstrated the superior performance of the proposed HCL-Net over state-of-the-art methods.
Bo Peng 0007, Tianyi Qin, Yanfeng Gu, Nam Ling, Jianjun Lei 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 Adversarially Robust Object Detection via Deviation Calibration and Content Preservation
abstract
Object detection has achieved a promising development in recent years and played an important role in various applications. However, the performance of object detection networks generally drops significantly when subjected to adversarial attacks. As an effective technique for defending against adversarial attacks, adversarially robust object detection has attracted increasing interest. In this paper, a novel deviation-calibrated and content-preserved network (DCCP-Net) is proposed for adversarially robust object detection by effectively exploring and mitigating the essential negative impact of noise disturbance in the feature space. Specifically, a deviation-calibrated robust feature enhancement module is designed to enhance the feature robustness of adversarial images by removing noise disturbance and supplementing rectified information. Besides, by enabling adversarial image features to imitate corresponding clean image features, a content-preserved consistency information imitation mechanism is proposed to obtain more accurate content information of adversarial images. Extensive experiment results have verified the superiority of the proposed DCCP-Net.
Xu Zhang 0045, Bo Peng 0007, Jianjun Lei 0001, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.3
2025 Cross-Modal Aligned Identity-Discriminative Feature Learning Network for Face Sketch Recognition
abstract
Face sketch recognition focuses on retrieving face photos that have the same identity as query face sketches, and plays a vital role in the field of information forensics and security. Owing to the large cross-modal differences between face sketches and photos, extracting and aligning cross-modal features is still considered a challenging task in the face sketch recognition community. This paper presents a novel cross-modal aligned identity-discriminative feature learning network (CAIFL-Net) for face sketch recognition. Specifically, in this paper, an identity-discriminative feature preservation module is designed to capture the identity-discriminative features of face sketches and photos by eliminating features that are weakly related to recognition. In addition, a sketch-photo cross-reconstructed feature alignment module is proposed to obtain cross-modal aligned features for effective recognition by reconstructing and embedding global features of one modality into another. Extensive experiments on the Uom-SGFS and CUFSF datasets demonstrate the effectiveness of the proposed CAIFL-Net.
Jianjun Lei 0001, Menglei Zhao, Bo Peng 0007, Qingming Huang
IEEE Trans. Inf. Forensics Secur.1
2025 Advancing Real-World Stereoscopic Image Super-Resolution via Vision-Language Model
abstract
Recent years have witnessed the remarkable success of the vision-language model in various computer vision tasks. However, how to exploit the semantic language knowledge of the vision-language model to advance real-world stereoscopic image super-resolution remains a challenging problem. This paper proposes a vision-language model-based stereoscopic image super-resolution (VLM-SSR) method, in which the semantic language knowledge in CLIP is exploited to facilitate stereoscopic image SR in a training-free manner. Specifically, by designing visual prompts for CLIP to infer the region similarity, a prompt-guided information aggregation mechanism is presented to capture inter-view information among relevant regions between the left and right views. Besides, driven by the prior knowledge of CLIP, a cognition prior-driven iterative enhancing mechanism is presented to optimize fuzzy regions adaptively. Experimental results on four datasets verify the effectiveness of the proposed method.
Zhe Zhang 0041, Jianjun Lei 0001, Bo Peng 0007, Liying Xu, Qingming Huang
IEEE Trans. Image Process.2
2025 Advancing Generalizable Occlusion Modeling for Neural Human Radiance Field
abstract
Generalizable human neural rendering aims to render the target views of the human body by leveraging source views and the skinned multi-person linear (SMPL) model. Despite exhibiting promising performance, the target views rendered by previous methods usually contain corrupted parts of the human body. Two primary challenges hinder high-quality human neural rendering. These challenges involve non-correspondences between 2D pixels and 3D SMPL vertices induced by self-occlusion of the human body and erroneous appearance predictions caused by occlusion between the source and target views. To solve these two challenges, we propose an advancing generalizable occlusion modeling method for the neural human radiance field, in which the hurdles from the self-occlusion of the human body and the occlusion between source and target views are explored and solved. Specifically, to alleviate the non-correspondence problem induced by self-occlusion, a geometry perception module is designed to obtain 3D geometric representations of SMPL vertices, enabling the prediction of accurate density values. Furthermore, a visibility aggregation module is designed to estimate the visibility maps with respect to different source views by utilizing the predicted density. Then, the complementary information among multiple source views is integrated with the support of the visibility maps in the visibility aggregation module, thus effectively addressing the occlusion between views. Experiments on the ZJU-MoCap and THUman datasets show that the proposed method achieves promising performance compared with the existing state-of-the-art methods.
Bingzheng Liu, Jianjun Lei 0001, Bo Peng 0007, Zhe Zhang 0041, Qingming Huang
IEEE Trans. Multim.2
2025 Efficient Chroma Intra Prediction via Exemplar Colorization Network for Versatile Video Coding
abstract
Chroma intra prediction aims to reduce chroma redundancies within a frame, which plays an important role in improving the coding efficiency of intra coding. Existing chroma intra prediction methods typically utilize the spatial relationship between the current luma block and its neighboring reference luma blocks to predict its chroma samples. However, the spatial properties of luma components differ from those of chroma components, which limits the accuracy of chroma intra prediction. To tackle this issue, an efficient Exemplar Colorization Network (ECNet)-based chroma intra prediction method is proposed in this paper, in which the colorization relationship between reference luma and chroma components is exploited to predict the chroma components for the current luma component. Inspired by the principle that semantic information in an image exhibits short-range continuity, a Spatial-consistency-based Colorization Transfer Network (SCTNet) is proposed, which builds and transfers colorization representations of neighboring reference blocks for chroma prediction. To improve the chroma prediction capability of SCTNet, a colorization learning module is developed to learn the robust mapping relationship from the luma component to the chroma component in a region-to-pixel manner, and a weight-adaptive reconstruction module is designed to adaptively utilize reference information from neighboring blocks to generate an initial prediction result. In addition, to further improve the accuracy of chroma intra prediction, a multi-reference-based chroma refinement network is proposed, which simultaneously uses the spatial information of neighboring reference chroma blocks and the current luma block to eliminate blocking and color-bleeding artifacts in the initial prediction result. Experimental results demonstrate that our proposed ECNet outperforms the state-of-the-art chroma intra prediction methods in terms of coding performance.
Zhaoqing Pan, Jixing Chen, Bo Peng 0007, Jianjun Lei 0001, Fu Lee Wang, Nam Ling, Sam Kwong
IEEE Trans. Multim.4
2025 Modeling Intra- and Inter-Modal Correlations for Incomplete Multi-Modal 3D Shape Clustering
abstract
The investigation for incomplete multi-modal 3D shape clustering is evolving as a promising task for the field of recognizing massive unlabeled 3D shapes. As two widely adopted 3D shape modalities, point clouds and multiple views not only exhibit rich intra-modal correlations but also encompass complementary structures and appearances of 3D shapes. By effectively modeling the intra-modal and inter-modal correlations, this paper proposes a novel incomplete multi-modal 3D shape clustering method to reveal the underlying clustering associations from incomplete multi-modal 3D shapes. In detail, a similarity-transferred feature prediction module is presented to recover the features of missing instances within one modality with the assistance of similarity exploring from another modality. Then, an intra-to-inter progressive feature fusion module is designed to mine the correlations within the modality as well as between different modalities, thereby obtaining comprehensive 3D shape features for clustering. Extensive experiments on two public 3D shape datasets have demonstrated that the proposed method has achieved promising clustering results under different missing rates.
Tianyi Qin, Bo Peng 0007, Jianjun Lei 0001, Qingming Huang
IEEE Trans. Multim.3
2025 Adaptive Multi-Exposure Image Correction via Joint Lightness and Structure Awareness
abstract
In order to alleviate the impact of ambient light on the quality of captured images, correcting multi-exposure images has become a popular topic. Most existing multi-exposure image correction methods mainly focus on the adjustment of lightness levels, but ignore the significant issue of structural information loss in incorrectly exposed images. Taking into consideration both lightness adjustment and structural reconstruction, this article proposes an adaptive multi-exposure image correction network by jointly exploring the lightness and structure information, named LSANet. Specifically, the proposed LSANet first extracts lightness and structure representations of the input image in the frequency domain, and then performs exposure level adjustment and structure detail reconstruction based on the lightness and structure representations. In the proposed network, the lightness- and structure-aware adaptive module is designed to achieve adaptive correction by predicting dynamic kernels under the guidance of the lightness and structure representations. Experimental results on the widely used ME and SICE datasets demonstrate that the proposed LSANet achieves excellent performance and generates images with well-exposed levels and rich structural details.
Bo Peng 0007, Jia Zhang 0025, Zhe Zhang 0041, Liying Xu, Qingming Huang, Tao Wang 0119, Jianjun Lei 0001
ACM Trans. Multim. Comput. Commun. Appl.7
2024 Saliency Map-Guided End-to-End Image Coding for Machines
abstract
Existing end-to-end image coding for machines (ICM) methods generally use joint training strategies to promote the compression efficiency for machine vision without considering the influence of different regions in the image. To encourage the image compression network to focus on the regions that are critical to the subsequent visual task, this paper proposes a saliency map-guided image compression network (SMIC-Net) for ICM. Specifically, a saliency map-guided transform module (SMTM) is proposed to improve the representation ability of image features for object detection task by exploring the semantic and structural information of the detected object. Besides, a saliency map-guided mean square error (SM-MSE) loss is designed to place more emphasis on the detected object regions. Experimental results demonstrate that the proposed SMIC-Net effectively promotes the compression efficiency for machine vision.
Bo Peng 0007, Tianxiang Lin, Dengchao Jin, Zhaoqing Pan, Jianjun Lei 0001
IEEE Signal Process. Lett.5
2024 Unsupervised Single-View Synthesis Network via Style Guidance and Prior Distillation
abstract
View synthesis aims to learn a view transformation and synthesize the target views from a single or multiple source views. Although previous view synthesis methods have obtained promising performance, they heavily rely on the supervision of the target view. In this paper, we propose an unsupervised single-view synthesis network (USVS-Net) to learn the view transformation without the supervision of the target view. Specifically, with the usage of only a single source view, a style-guidance view synthesis model is proposed to learn an intrinsic representation, which intends to describe the object from a reference pose. With the intrinsic representation, the view transformation is learned to boost the learning of the unsupervised single-view synthesis. Then, taking the style-guidance view synthesis model as the teacher, a prior-distillation view synthesis model is further presented as the student to learn a more direct view transformation. By utilizing the proposed method, high-quality target views are synthesized in a time-efficient manner. Experiments on both synthetic and real-scene datasets show that despite the lack of supervision of the target view, the proposed method achieves promising results compared with the existing view synthesis methods.
Bingzheng Liu, Bo Peng 0007, Zhe Zhang 0041, Qingming Huang, Nam Ling, Jianjun Lei 0001
IEEE Trans. Circuits Syst. Video Technol.6
2024 PIPC-3Ddet: Harnessing Perspective Information and Proposal Correlation for 3D Point Cloud Object Detection
abstract
As a fundamental technology in autonomous driving and robotic sensing system, 3D point cloud object detection has received increasing attention. In this paper, a novel 3D detection method that harnesses perspective information and proposal correlation (PIPC-3Ddet) is proposed for detecting 3D objects from point clouds. Specifically, a perspective information embedding module is designed to enhance the voxel features by capturing and embedding the perspective information of range images, so as to effectively distinguish the objects and backgrounds. Besides, by revealing the correlation among 3D proposals, a proposal correlation reasoning module is presented to learn high-quality proposal features for better 3D proposal refinement. With the designed perspective information embedding and proposal correlation reasoning modules, the proposed PIPC-3Ddet is able to better perceive the objects in the 3D scene, thus boosting the 3D object detection performance. Extensive experiments on the KITTI and Waymo benchmarks have demonstrated the superiority of the proposed PIPC-3Ddet.
Chuanbo Yu, Bo Peng 0007, Qingming Huang, Jianjun Lei 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 SWGNet: Step-Wise Reference Frame Generation Network for Multiview Video Coding
abstract
In multiview video coding, the coding performance highly depends on the quality of the reference frames. In view of this, a step-wise reference frame generation network (SWGNet) is designed to improve the quality of the reference frame for efficient multiview video coding. In particular, a frame-level to block-level learning paradigm is proposed to step-wisely generate a high-quality reference frame. In the frame-level stage, by exploiting parallax correlations between temporal and inter-view references on the basis of image alignment, a parallax-guided frame-level synthesis module is proposed to generate an elementary reference frame. Then, in the block-level stage, a transformer-based block-level aggregation module is designed to further refine the texture details of the reference frame by modeling long-range dependencies among pixels. The proposed SWGNet is integrated into 3D-HEVC, and extensive experiments demonstrate that the proposed method achieves significant bitrate saving compared with 3D-HEVC.
Jing Zhang 0017, Yonghong Hou, Zhaoqing Pan, Bo Peng 0007, Nam Ling, Jianjun Lei 0001
IEEE Trans. Circuits Syst. Video Technol.6
2024 Self-Constructing Stereo Correspondences for Unsupervised Multi-View Stereo
abstract
Existing unsupervised Multi-View Stereo (MVS) methods generally construct supervision on the basis of the photometric consistency loss, which suffers from unreliable supervision and limited scalability. In this paper, a novel unsupervised MVS framework with Self-constructed Stereo Correspondences, termed SSC-MVS, is proposed to provide reliable supervision for the network and improve scalability of unsupervised MVS. Specifically, a pseudo depth-based learning strategy is first presented to supervise the MVS network with a pseudo depth, which is used to characterize the accurate stereo correspondences. Additionally, a consistency-based training mechanism is designed, where the depth consistency between two differently-augmented inputs is constrained to further improve the robustness of the network in real MVS scenes. Experimental results on widely-used MVS datasets demonstrate that the proposed SSC-MVS obtains the state-of-the-art performance among the unsupervised methods and has the potential to outperform the fully-supervised methods. The code is available athttps://github.com/jzhu98/ssc-mvs.
Bo Peng 0007, Bingzheng Liu, Qingming Huang, Jianjun Lei 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 λ-Domain Rate Control via Wavelet-Based Residual Neural Network for VVC HDR Intra Coding
abstract
High dynamic range (HDR) video offers a more realistic visual experience than standard dynamic range (SDR) video, while introducing new challenges to both compression and transmission. Rate control is an effective technology to overcome these challenges, and ensure optimal HDR video delivery. However, the rate control algorithm in the latest video coding standard, versatile video coding (VVC), is tailored to SDR videos, and does not produce well coding results when encoding HDR videos. To address this problem, a data-driven λ -domain rate control algorithm is proposed for VVC HDR intra frames in this paper. First, the coding characteristics of HDR intra coding are analyzed, and a piecewise R- λ model is proposed to accurately determine the correlation between the rate (R) and the Lagrange parameter λ for HDR intra frames. Then, to optimize bit allocation at the coding tree unit (CTU)-level, a wavelet-based residual neural network (WRNN) is developed to accurately predict the parameters of the piecewise R- λ model for each CTU. Third, a large-scale HDR dataset is established for training WRNN, which facilitates the applications of deep learning in HDR intra coding. Extensive experimental results show that our proposed HDR intra frame rate control algorithm achieves superior coding results than the state-of-the-art algorithms. The source code of this work will be released at https://github.com/TJU-Videocoding/WRNN.git.
Jianjun Lei 0001, Zhaoqing Pan, Bo Peng 0007, Haoran Xie 0001
IEEE Trans. Image Process.2
2024 Contrastive Multi-View Learning for 3D Shape Clustering
abstract
Unsupervised 3D shape clustering is emerging as a promising research topic in multimedia and computer vision field. Considering the flexibility of acquiring multiple views for 3D shapes, this paper proposes a contrastive multi-view learning network (CMVL-Net) to cluster unlabeled 3D shapes from multiple views. To the best of our knowledge, this is the first multi-view-oriented 3D shape deep clustering method. The key to this method lies in how to capture highly discriminative 3D shape features suitable for clustering. By exploring consistency and complementarity among multiple views, a cross-view contrastive clustering mechanism is proposed to learn clustering-specified discriminative 3D shape features. To obtain a more compact 3D shape clustering structure, a consensus graph-guided contrastive constraint is designed to encourage cluster-wise consistency learning under the guidance of potential category associations among shapes. Experimental results on two widely used benchmark datasets demonstrate the effectiveness of the proposed method.
Bo Peng 0007, Guoting Lin, Jianjun Lei 0001, Tianyi Qin, Xiaochun Cao, Nam Ling
IEEE Trans. Multim.3
2024 Multi-Projection Fusion and Refinement Network for Salient Object Detection in 360° Omnidirectional Image
abstract
Salient object detection (SOD) aims to determine the most visually attractive objects in an image. With the development of virtual reality (VR) technology, 360° omnidirectional image has been widely used, but the SOD task in 360° omnidirectional image is seldom studied due to its severe distortions and complex scenes. In this article, we propose a multi-projection fusion and refinement network (MPFR-Net) to detect the salient objects in 360° omnidirectional image. Different from the existing methods, the equirectangular projection (EP) image and four corresponding cube-unfolding (CU) images are embedded into the network simultaneously as inputs, where the CU images not only provide supplementary information for EP image but also ensure the object integrity of cube-map projection. In order to make full use of these two projection modes, a dynamic weighting fusion (DWF) module is designed to adaptively integrate the features of different projections in a complementary and dynamic manner from the perspective of inter and intrafeatures. Furthermore, in order to fully explore the way of interaction between encoder and decoder features, a filtration and refinement (FR) module is designed to suppress the redundant information of the feature itself and between the features. Experimental results on two omnidirectional datasets demonstrate that the proposed approach outperforms the state-of-the-art methods both qualitatively and quantitatively. The code and results can be found from the link of https://rmcong.github.io/proj_MPFRNet.html.
Runmin Cong, Jianjun Lei 0001, Yao Zhao 0001, Qingming Huang, Sam Kwong
IEEE Trans. Neural Networks Learn. Syst.3
2024 Self-Supervised Monocular Depth Estimation via Binocular Geometric Correlation Learning
abstract
Monocular depth estimation aims to infer a depth map from a single image. Although supervised learning-based methods have achieved remarkable performance, they generally rely on a large amount of labor-intensively annotated data. Self-supervised methods, on the other hand, do not require any annotation of ground-truth depth and have recently attracted increasing attention. In this work, we propose a self-supervised monocular depth estimation network via binocular geometric correlation learning. Specifically, considering the inter-view geometric correlation, a binocular cue prediction module is presented to generate the auxiliary vision cue for the self-supervised learning of monocular depth estimation. Then, to deal with the occlusion in depth estimation, an occlusion interference attenuated constraint is developed to guide the supervision of the network by inferring the occlusion region and producing paired occlusion masks. Experimental results on two popular benchmark datasets have demonstrated that the proposed network obtains competitive results compared to state-of-the-art self-supervised methods and achieves comparable results to some popular supervised methods.
Bo Peng 0007, Jianjun Lei 0001, Bingzheng Liu, Haifeng Shen, Wanqing Li 0001, Qingming Huang
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Deep In-Loop Filtering via Multi-Domain Correlation Learning and Partition Constraint for Multiview Video Coding
abstract
The deep learning-based in-loop filtering methods have greatly improved the coding efficiency for High Efficiency Video Coding (HEVC). However, directly applying these HEVC-orientated in-loop filtering methods to multiview video coding may not obtain satisfactory performance due to the characteristics of multiview video. In this paper, a deep in-loop filtering method based on multi-domain correlation learning and partition constraint network (MDP-Net) is proposed to boost the multiview video coding performance. To the best of our knowledge, this work is the first attempt at deep in-loop filtering for multiview video coding. Specifically, a multi-domain correlation learning module is presented to restore the high-frequency details of the distorted frame by exploring the multi-domain correlations. Besides, based on the block partition information generated in video coding, a partition-constrained reconstruction module is proposed to better attenuate the compression artifacts by designing a partition loss. Finally, the proposed MDP-Net is integrated into 3D-HEVC reference software, and the experimental results demonstrate that the proposed method achieves considerable performance improvement compared with 3D-HEVC.
Bo Peng 0007, Renjie Chang, Zhaoqing Pan, Ge Li 0006, Nam Ling, Jianjun Lei 0001
IEEE Trans. Circuits Syst. Video Technol.6
2023 RGB-D Human Matting: A Real-World Benchmark Dataset and a Baseline Method
abstract
The last decade has witnessed an increasing exploration and development of human matting. However, existing matting works primarily focus on predicting better alpha mattes from RGB images. So far few efforts have been devoted to tackling human matting in real-world activity scenarios with RGB-D information. To this end, this paper concentrates on the RGB-D human matting task, and provides the first public RGB-D human matting benchmark dataset as well as a baseline method for deep learning-based RGB-D human matting. To support the research on RGB-D human matting, a new RGB-D human-matting dataset (HDM-2K) is collected and released, which contains 2,270 high-resolution human images in various real-world scenarios and the corresponding depth maps. Additionally, a baseline method for RGB-D human matting is further proposed, which automatically generates the alpha matte by jointly exploiting the spatial structure information in the depth map and detailed texture information in the RGB image. Finally, extensive experiments conducted on the HDM-2K dataset demonstrate that the depth maps are effective for the matting task and the proposed baseline method achieves promising performance on human matting.
Bo Peng 0007, Jianjun Lei 0001, Huazhu Fu, Haifeng Shen, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.3
2023 Recurrent Interaction Network for Stereoscopic Image Super-Resolution
abstract
Recently, deep learning-based stereoscopic image super-resolution has attracted extensive attention and made great progress. However, existing methods have not adequately explored the inter-view dependency among two-view multi-level features. In this paper, a recurrent interaction network for stereoscopic image super-resolution (RISSRnet) is proposed to learn the inter-view dependency. To efficiently utilize the relationship between the two views, a recurrent interaction module is designed to achieve recurrent interaction among two-view multi-level features from the regrouped sequences, which are generated by a coupled queue-regroup mechanism. In addition, to recursively enhance features in the recurrent interaction module, an iterative propagation strategy is developed for sufficient interaction. Extensive experimental results demonstrate the effectiveness and superiority of the proposed RISSRnet.
Zhe Zhang 0041, Bo Peng 0007, Jianjun Lei 0001, Haifeng Shen, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.3
2023 Reducing Background Induced Domain Shift for Adaptive Person Re-Identification
abstract
Cross-domain person re-identification (Re-ID) is a challenging and important task in monitoring safety and procedure compliance of industrial work places. In this article, a novel method is proposed to reduce background induced domain shift for adaptive person Re-ID. Specifically, a foreground-background joint clustering module is proposed to extract discriminative foreground and background features and an attention-based feature disentanglement module is designed to reduce the interference of background with the extraction of discriminative foreground features. Experimental results on three widely used person Re-ID benchmarking datasets (Market-1501, DukeMTMC-reID, and MSMT17) have demonstrated that the proposed method achieves promising performance compared with the state-of-the-art methods.
Jianjun Lei 0001, Tianyi Qin, Bo Peng 0007, Wanqing Li 0001, Zhaoqing Pan, Haifeng Shen, Sam Kwong
IEEE Trans. Ind. Informatics1
2023 ZS-SBPRnet: A Zero-Shot Sketch-Based Point Cloud Retrieval Network Based on Feature Projection and Cross-Reconstruction
abstract
With the widespread deployment of 3D sensors, point cloud analysis has become an important topic in the field of industrial information. This article proposes a novel zero-shot sketch-based point cloud retrieval network based on feature projection and cross reconstruction, termed as ZS-SBPRnet. As far as we know, the proposed ZS-SBPRnet is the first attempt at retrieving point clouds based on sketches under the zero-shot scenario. To tackle the problem of the cross-modal differences, a structure-preserving learnable feature projection module is designed to obtain view feature representations from point cloud features containing spatial structure information through feature projection. Besides, to achieve efficient cross-modal feature alignment under the zero-shot scenario, a sketch-point cloud cross-reconstruction mechanism is presented to promote cross-modal feature alignment between sketches and point clouds in visual space. Experimental results on the benchmark datasets validate the superiority of the proposed ZS-SBPRnet.
Bo Peng 0007, Haifeng Shen, Qingming Huang, Jianjun Lei 0001
IEEE Trans. Ind. Informatics6
2023 Learned Video Compression With Efficient Temporal Context Learning
abstract
In contrast to image compression, the key of video compression is to efficiently exploit the temporal context for reducing the inter-frame redundancy. Existing learned video compression methods generally rely on utilizing short-term temporal correlations or image-oriented codecs, which prevents further improvement of the coding performance. This paper proposed a novel temporal context-based video compression network (TCVC-Net) for improving the performance of learned video compression. Specifically, a global temporal reference aggregation (GTRA) module is proposed to obtain an accurate temporal reference for motion-compensated prediction by aggregating long-term temporal context. Furthermore, in order to efficiently compress the motion vector and residue, a temporal conditional codec (TCC) is proposed to preserve structural and detailed information by exploiting the multi-frequency components in temporal context. Experimental results show that the proposed TCVC-Net outperforms public state-of-the-art methods in terms of both PSNR and MS-SSIM metrics.
Dengchao Jin, Jianjun Lei 0001, Bo Peng 0007, Zhaoqing Pan, Li Li 0040, Nam Ling
IEEE Trans. Image Process.2
2023 Novel View Synthesis from a Single Unposed Image via Unsupervised Learning
abstract
Novel view synthesis aims to generate novel views from one or more given source views. Although existing methods have achieved promising performance, they usually require paired views with different poses to learn a pixel transformation. This article proposes an unsupervised network to learn such a pixel transformation from a single source image. In particular, the network consists of a token transformation module that facilities the transformation of the features extracted from a source image into an intrinsic representation with respect to a pre-defined reference pose and a view generation module that synthesizes an arbitrary view from the representation. The learned transformation allows us to synthesize a novel view from any single source image of an unknown pose. Experiments on the widely used view synthesis datasets have demonstrated that the proposed network is able to produce comparable results to the state-of-the-art methods despite the fact that learning is unsupervised and only a single source image is required for generating a novel view. The code will be available upon the acceptance of the article.
Bingzheng Liu, Jianjun Lei 0001, Bo Peng 0007, Chuanbo Yu, Wanqing Li 0001, Nam Ling
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Modeling Long-range Dependencies and Epipolar Geometry for Multi-view Stereo
abstract
This article proposes a network, referred to as Multi-View Stereo TRansformer (MVSTR) for depth estimation from multi-view images. By modeling long-range dependencies and epipolar geometry, the proposed MVSTR is capable of extracting dense features with global context and 3D consistency, which are crucial for reliable matching in multi-view stereo (MVS). Specifically, to tackle the problem of the limited receptive field of existing CNN-based MVS methods, a global-context Transformer module is designed to establish intra-view long-range dependencies so that global contextual features of each view are obtained. In addition, to further enable features of each view to be 3D consistent, a 3D-consistency Transformer module with an epipolar feature sampler is built, where epipolar geometry is modeled to effectively facilitate cross-view interaction. Experimental results show that the proposed MVSTR achieves the best overall performance on the DTU dataset and demonstrates strong generalization on the Tanks & Temples benchmark dataset.
Bo Peng 0007, Wanqing Li 0001, Haifeng Shen, Qingming Huang, Jianjun Lei 0001
ACM Trans. Multim. Comput. Commun. Appl.6
2022 Deep Stereo Image Compression via Bi-directional Coding
abstract
Existing learning-based stereo compression methods usually adopt a unidirectional approach to encoding one image independently and the other image conditioned upon the first. This paper proposes a novel bidirectional coding-based end-to-end stereo image compression network (BCSIC-Net). BCSIC-Net consists of a novel bidirectional contextual transform module which performs nonlinear transform conditioned upon the inter-view context in a latent space to reduce inter-view redundancy, and a bidirectional conditional entropy model that employs interview correspondence as a conditional prior to improve coding efficiency. Experimental results on the InStereo2K and KITTI datasets demonstrate that the proposed BCSIC-Net can effectively reduce the inter-view redundancy and out-performs state-of-the-art methods.
Jianjun Lei 0001, Xiangrui Liu, Bo Peng 0007, Dengchao Jin, Wanqing Li 0001, Jingxiao Gu
CVPR1
2022 Texture-Guided End-to-End Depth Map Compression
abstract
End-to-end compression methods designed for the texture image have achieved excellent coding performances. Due to the characteristic differences between the depth map and the texture image, the texture-oriented methods have limitations in depth map compression. To address this problem, this paper proposes a texture-guided end-to-end depth map compression network (TDMC-Net). Specifically, the proposed TDMC-Net is mainly composed of the texture-guided transform module (TTM) which performs the nonlinear transform with providing the textual context to reduce the redundancy in depth feature, and a texture-guided conditional entropy model (TCEM) which is designed to improve the entropy model by introducing the texture conditional prior. Experimental results show that the proposed TDMC-Net boosts the depth coding efficiency by utilizing the texture information and achieves superior performance.
Bo Peng 0007, Yuying Jing, Dengchao Jin, Xiangrui Liu, Zhaoqing Pan, Jianjun Lei 0001
ICIP6
2022 Advanced Dropout: A Model-Free Methodology for Bayesian Dropout Optimization
abstract
Due to lack of data, overfitting ubiquitously exists in real-world applications of deep neural networks (DNNs). We propose advanced dropout, a model-free methodology, to mitigate overfitting and improve the performance of DNNs. The advanced dropout technique applies a model-free and easily implemented distribution with parametric prior, and adaptively adjusts dropout rate. Specifically, the distribution parameters are optimized by stochastic gradient variational Bayes in order to carry out an end-to-end training. We evaluate the effectiveness of the advanced dropout against nine dropout techniques on seven computer vision datasets (five small-scale datasets and two large-scale datasets) with various base models. The advanced dropout outperforms all the referred techniques on all the datasets. We further compare the effectiveness ratios and find that advanced dropout achieves the highest one on most cases. Next, we conduct a set of analysis of dropout rate characteristics, including convergence of the adaptive dropout rate, the learned distributions of dropout masks, and a comparison with dropout rate generation without an explicit distribution. In addition, the ability of overfitting prevention is evaluated and confirmed. Finally, we extend the application of the advanced dropout to uncertainty inference, network pruning, text classification, and regression. The proposed advanced dropout is also superior to the corresponding referred methods. Codes are available at https://github.com/PRIS-CV/AdvancedDropout.
Jiyang Xie 0001, Zhanyu Ma, Jianjun Lei 0001, Guoqiang Zhang 0003, Jing-Hao Xue, Zheng-Hua Tan, Jun Guo 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Deep Affine Motion Compensation Network for Inter Prediction in VVC
abstract
In video coding, it is a challenge to deal with scenes with complex motions, such as rotation and zooming. Although affine motion compensation (AMC) is employed in Versatile Video Coding (VVC), it is still difficult to handle non-translational motions due to the adopted hand-craft block-based motion compensation. In this paper, we propose a deep affine motion compensation network (DAMC-Net) for inter prediction in video coding to effectively improve the prediction accuracy. To the best of our knowledge, our work is the first attempt to deal with the deformable motion compensation based on CNN in VVC. Specifically, a deformable motion-compensated prediction (DMCP) module is proposed to compensate the current encoding block through a learnable way to estimate accurate motion fields. Meanwhile, the spatial neighboring information and the temporal reference block as well as the initial motion field are fully exploited. By effectively fusing the multi-channel feature maps from DMCP, an attention-based fusion and reconstruction (AFR) module is designed to reconstruct the output block. The proposed DAMC-Net is integrated into VVC and the experimental results demonstrate that the proposed method considerably enhances the coding performance.
Dengchao Jin, Jianjun Lei 0001, Bo Peng 0007, Wanqing Li 0001, Nam Ling, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.2
2022 Multiple Resolution Prediction With Deep Up-Sampling for Depth Video Coding
abstract
The depth video contains large smooth contents with sharp edges. Since the deep learning-based color video orientated intra prediction methods pay no attention to the characteristics of depth video, they are unsuitable for optimizing the coding efficiency of depth video. In this paper, a multiple resolution prediction method with deep up-sampling is proposed to promote the coding efficiency of depth video. To efficiently encode the depth blocks of different complexity, the depth block is selectively encoded at different resolutions, including$\times 1$,$\times 1$/2, and$\times 1$/4 resolutions. If the block is encoded with a low-resolution (LR), the resolution of reconstructed LR depth block is recovered by an up-sampling network. To constrain the quality of both reconstructed high-resolution depth block and its synthesized view, a view synthesis distortion guidance mechanism is proposed for the up-sampling network. In addition, a distillation-based lightweight up-sampling network is proposed to reduce the computational complexity. Experimental results demonstrate that the proposed multiple resolution prediction method obtains an average of 10.84% BD-rate saving in comparison with 3D-HEVC.
Ge Li 0006, Jianjun Lei 0001, Zhaoqing Pan, Bo Peng 0007, Nam Ling
IEEE Trans. Circuits Syst. Video Technol.2
2022 TSAN: Synthesized View Quality Enhancement via Two-Stream Attention Network for 3D-HEVC
abstract
In three-dimensional video system, the texture and depth videos are jointly encoded, and then the Depth Image Based Rendering (DIBR) is utilized to realize view synthesis. However, the compression distortion of texture and depth videos, as well as the disocclusion problem in DIBR degrade the visual quality of the synthesized view. To address this problem, a Two-stream Attention Network (TSAN)-based synthesized view quality enhancement method is proposed for 3D-High Efficiency Video Coding (3D-HEVC) in this article. First, the shortcomings of the view synthesis technique and traditional convolutional neural networks are analyzed. Then, based on these analyses, a TSAN with two information extraction streams is proposed for enhancing the quality of the synthesized view, in which the global information extraction stream learns the contextual information, and the local information extraction stream extracts the texture information from the rendered image. Third, a Multi-Scale Residual Attention Block (MSRAB) is proposed, which can efficiently detect features in different scales, and adaptively refine features by considering interdependencies among spatial dimensions. Extensive experimental results show that the proposed synthesized view quality enhancement method achieves significantly better performance than the state-of-the-art methods.
Zhaoqing Pan, Jianjun Lei 0001, Nam Ling, Sam Kwong
IEEE Trans. Circuits Syst. Video Technol.3
2022 RDEN: Residual Distillation Enhanced Network-Guided Lightweight Synthesized View Quality Enhancement for 3D-HEVC
abstract
In the three-dimensional video system, the depth image-based rendering is a key technique for generating synthesized views, which provides audiences with depth perception and interactivity. However, the inaccuracy of depth information leads to geometrical rendering position errors, and the compression distortion of texture and depth videos degrades the quality of the synthesized views. Although existing quality enhancement methods can eliminate the distortions in the synthesized views, their huge computational complexity hinders their applications in real-time multimedia systems. To this end, a residual distillation enhanced network (RDEN)-guided lightweight synthesized view quality enhancement (SVQE) method is proposed to minimize holes and compression distortions in the synthesized views while reducing the model complexity. First, a rethinking on the deep-learning-based SVQE methods is performed. Then, a feature distillation attention block is proposed to effectively reduce the distortions in the synthesized views and make the model fulfill more real-time tasks, which is a lightweight and flexible feature extraction block using an information distillation mechanism and a lightweight multi-scale spatial attention mechanism. Third, a residual feature fusion block is proposed to improve the enhancement performance by using the feature fusion mechanism, which efficiently improves the feature extraction capability without introducing any additional parameters. Experimental results prove that the proposed RDEN efficiently improves the SVQE performance while consuming few computational complexities compared with the state-of-the-art SVQE methods.
Zhaoqing Pan, Jianjun Lei 0001, Nam Ling, Sam Kwong
IEEE Trans. Circuits Syst. Video Technol.4
2022 DACNN: Blind Image Quality Assessment via a Distortion-Aware Convolutional Neural Network
abstract
Deep neural networks have achieved great performance on blind Image Quality Assessment (IQA), but it is still challenging for using one network to accurately predict the quality of images with different distortions. In this paper, a Distortion-Aware Convolutional Neural Network (DACNN) is proposed for blind IQA, which works effectively for not only synthetically distorted images but also authentically distorted images. The proposed DACNN consists of a distortion aware module, a distortion fusion module, and a quality prediction module. In the distortion aware module, a Siamese network-based pretraining strategy is proposed to design a synthetic distortion-aware network for full learning the synthetic distortions, and an authentic distortion-aware network is used for extracting the authentic distortions. To efficiently fuse the learned distortion features, and make the network pay more attention to the essential features, a weight-adaptive fusion network is proposed to adaptively adjust the weight of each distortion. Finally, the quality prediction module is adopted to map the fused features to a quality score. Extensive experiments on four authentic IQA databases and four synthetic IQA databases have proved the effectiveness of the proposed DACNN.
Zhaoqing Pan, Jianjun Lei 0001, Yuming Fang 0001, Xiao Shao, Nam Ling, Sam Kwong
IEEE Trans. Circuits Syst. Video Technol.3
2022 LVE-S2D: Low-Light Video Enhancement From Static to Dynamic
abstract
Recently, deep-learning-based low-light video enhancement methods have drawn wide attention and achieved remarkable performance. However, limited by the difficulty in collecting dynamic low-light and well-lighted video pairs in real scenes, how to construct video sequences for supervised learning and design a low-light enhancement network for real dynamic video remains a challenge. In this paper, we propose a simple yet effective low-light video enhancement method (LVE-S2D), which generates dynamic video training pairs from static videos, and enhances the low-light video by mining dynamic temporal information. To obtain low-light and well-lighted video pairs, a sliding window-based dynamic video generation mechanism is designed to produce pseudo videos with rich dynamic temporal information. Then, a siamese dynamic low-light video enhancement network is presented, which effectively utilizes temporal correlation between adjacent frames to enhance the video frames. Extensive experimental results demonstrate that the proposed method not only achieves superior performance on static low-light videos, but also outperforms the state-of-the-art methods on real dynamic low-light videos.
Bo Peng 0007, Jianjun Lei 0001, Zhe Zhang 0041, Nam Ling, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.3
2022 SIEV-Net: A Structure-Information Enhanced Voxel Network for 3D Object Detection From LiDAR Point Clouds
abstract
As one of the fundamental tasks in scene understanding, 3D object detection from LiDAR point clouds has drawn extensive attention in the past few years. Although the existing voxel-based methods have achieved remarkable performance, how to effectively exploit geometric structure information of the point clouds to boost the detection performance remains to be explored. In this paper, we propose a novel structure-information enhanced voxel network (SIEV-Net) for 3D object detection from LiDAR point clouds. The proposed SIEV-Net learns feature representations of 3D objects by jointly considering uneven spatial distribution and height information of the point clouds. Specifically, considering the uneven spatial distribution characteristics of point clouds, a hierarchical-voxel feature encoding module is proposed to effectively extract features of voxels in both sparse and dense regions. Besides, by utilizing the Bird’s Eye View (BEV) map of point clouds, a height information complement module is designed to minimize the height information lost in the process of point feature aggregation in a voxel network. Experimental results on the widely used KITTI benchmark dataset have demonstrated the efficacy of the proposed SIEV-Net.
Chuanbo Yu, Jianjun Lei 0001, Bo Peng 0007, Haifeng Shen, Qingming Huang
IEEE Trans. Geosci. Remote. Sens.2
2022 C2FNet: A Coarse-to-Fine Network for Multi-View 3D Point Cloud Generation
abstract
Generation of a 3D model of an object from multiple views has a wide range of applications. Different parts of an object would be accurately captured by a particular view or a subset of views in the case of multiple views. In this paper, a novel coarse-to-fine network (C2FNet) is proposed for 3D point cloud generation from multiple views. C2FNet generates subsets of 3D points that are best captured by individual views with the support of other views in a coarse-to-fine way, and then fuses these subsets of 3D points to a whole point cloud. It consists of a coarse generation module where coarse point clouds are constructed from multiple views by exploring the cross-view spatial relations, and a fine generation module where the coarse point cloud features are refined under the guidance of global consistency in appearance and context. Extensive experiments on the benchmark datasets have demonstrated that the proposed method outperforms the state-of-the-art methods.
Jianjun Lei 0001, Bo Peng 0007, Wanqing Li 0001, Zhaoqing Pan, Qingming Huang
IEEE Trans. Image Process.1
2022 Disparity-Aware Reference Frame Generation Network for Multiview Video Coding
abstract
Multiview video coding (MVC) aims to compress the multiview video through the elimination of video redundancies, where the quality of the reference frame directly affects the compression efficiency. In this paper, we propose a deep virtual reference frame generation method based on a disparity-aware reference frame generation network (DAG-Net) to transform the disparity relationship between different viewpoints and generate a more reliable reference frame. The proposed DAG-Net consists of a multi-level receptive field module, a disparity-aware alignment module, and a fusion reconstruction module. First, a multi-level receptive field module is designed to enlarge the receptive field, and extract the multi-scale deep features of the temporal and inter-view reference frames. Then, a disparity-aware alignment module is proposed to learn the disparity relationship, and perform disparity shift on the inter-view reference frame to align it with the temporal reference frame. Finally, a fusion reconstruction module is utilized to fuse the complementary information and generate a more reliable virtual reference frame. Experiments demonstrate that the proposed reference frame generation method achieves superior performance for multiview video coding.
Jianjun Lei 0001, Zongqian Zhang, Zhaoqing Pan, Dong Liu 0002, Xiangrui Liu, Ying Chen 0011, Nam Ling
IEEE Trans. Image Process.1
2022 VCRNet: Visual Compensation Restoration Network for No-Reference Image Quality Assessment
abstract
Guided by the free-energy principle, generative adversarial networks (GAN)-based no-reference image quality assessment (NR-IQA) methods have improved the image quality prediction accuracy. However, the GAN cannot well handle the restoration task for the free-energy principle-guided NR-IQA methods, especially for the severely destroyed images, which results in that the quality reconstruction relationship between the distorted image and its restored image cannot be accurately built. To address this problem, a visual compensation restoration network (VCRNet)-based NR-IQA method is proposed, which uses a non-adversarial model to efficiently handle the distorted image restoration task. The proposed VCRNet consists of a visual restoration network and a quality estimation network. To accurately build the quality reconstruction relationship between the distorted image and its restored image, a visual compensation module, an optimized asymmetric residual block, and an error map-based mixed loss function, are proposed for increasing the restoration capability of the visual restoration network. For further addressing the NR-IQA problem of severely destroyed images, the multi-level restoration features which are obtained from the visual restoration network are used for the image quality estimation. To prove the effectiveness of the proposed VCRNet, seven representative IQA databases are used, and experimental results show that the proposed VCRNet achieves the state-of-the-art image quality prediction accuracy. The implementation of the proposed VCRNet has been released at https://github.com/NUIST-Videocoding/VCRNet.
Zhaoqing Pan, Jianjun Lei 0001, Yuming Fang 0001, Xiao Shao, Sam Kwong
IEEE Trans. Image Process.3
2022 Multi-Modality MR Image Synthesis via Confidence-Guided Aggregation and Cross-Modality Refinement
abstract
Magnetic resonance imaging (MRI) can provide multi-modality MR images by setting task-specific scan parameters, and has been widely used in various disease diagnosis and planned treatments. However, in practical clinical applications, it is often difficult to obtain multi-modality MR images simultaneously due to patient discomfort, and scanning costs, etc. Therefore, how to effectively utilize the existing modality images to synthesize missing modality image has become a hot research topic. In this paper, we propose a novel confidence-guided aggregation and cross-modality refinement network (CACR-Net) for multi-modality MR image synthesis, which effectively utilizes complementary and correlative information of multiple modalities to synthesize high-quality target-modality images. Specifically, to effectively utilize the complementary modality-specific characteristics, a confidence-guided aggregation module is proposed to adaptively aggregate the multiple target-modality images generated from multiple source-modality images by using the corresponding confidence maps. Based on the aggregated target-modality image, a cross-modality refinement module is presented to further refine the target-modality image by mining correlative information among the multiple source-modality images and aggregated target-modality image. By training the proposed CACR-Net in an end-to-end manner, high-quality and sharp target-modality MR images are effectively synthesized. Experimental results on the widely used benchmark demonstrate that the proposed method outperforms state-of-the-art methods.
Bo Peng 0007, Bingzheng Liu, Yi Bin, Lili Shen, Jianjun Lei 0001
IEEE J. Biomed. Health Informatics5
2022 MIEGAN: Mobile Image Enhancement via a Multi-Module Cascade Neural Network
abstract
Visual quality of images captured by mobile devices is often inferior to that of images captured by a Digital Single Lens Reflex (DSLR) camera. This paper presents a novel generative adversarial network-based mobile image enhancement method, referred to as MIEGAN. It consists of a novel multi-module cascade generative network and a novel adaptive multi-scale discriminative network. The multi-module cascade generative network is built upon a two-stream encoder, a feature transformer, and a decoder. In the two-stream encoder, a luminance-regularizing stream is proposed to help the network focus on low-light areas. In the feature transformation module, two networks effectively capture both global and local information of an image. To further assist the generative network to generate the high visual quality images, a multi-scale discriminator is used instead of a regular single discriminator to distinguish whether an image is fake or real globally and locally. To balance the global and local discriminators, an adaptive weight allocation is proposed. In addition, a contrast loss is proposed, and a new mixed loss function is developed to improve the visual quality of the enhanced images. Extensive experiments on the popular DSLR photo enhancement dataset and MIT-FiveK dataset have verified the effectiveness of the proposed MIEGAN.
Zhaoqing Pan, Jianjun Lei 0001, Wanqing Li 0001, Nam Ling, Sam Kwong
IEEE Trans. Multim.3
2021 Depth-Assisted Joint Detection Network For Monocular 3d Object Detection
abstract
In the past few years, monocular 3D object detection has attracted increasing attention due to the merit of low cost and wide range of applications. In this paper, a depth-assisted joint detection network (MonoDAJD) is proposed for monocular 3D object detection. Specifically, a consistency-aware joint detection mechanism is proposed to jointly detect objects in the image and depth map, and exploit the localization information from the depth detection stream to optimize the detection results. To obtain more accurate 3D bounding boxes, an orientation-embedded NMS is designed by introducing the orientation confidence prediction and embedding the orientation confidence into the traditional NMS. Experimental results on the widely used KITTI benchmark demonstrate that the proposed method achieves promising performance compared with the state-of-the-art monocular 3D object detection methods.
Jianjun Lei 0001, Tingyi Guo, Bo Peng 0007, Chuanbo Yu
ICIP1
2021 Unsupervised stereoscopic image retargeting via view synthesis and stereo cycle consistency losses
Xiaoting Fan, Jianjun Lei 0001, Jie Liang 0001, Yuming Fang 0001, Xiaochun Cao, Nam Ling
Neurocomputing2
2021 Deep video action clustering via spatio-temporal feature learning
Bo Peng 0007, Jianjun Lei 0001, Huazhu Fu, Yalong Jia, Zongqian Zhang
Neurocomputing2
2021 No-reference stereoscopic image quality assessment based on global and local content characteristics
Lili Shen, Xiongfei Chen, Zhaoqing Pan, Kefeng Fan, Jianjun Lei 0001
Neurocomputing6
2021 RGB-D salient object detection via cross-modal joint feature extraction and low-bound fusion loss
Huazhu Fu, Xiaoting Fan, Yanan Shi, Jianjun Lei 0001
Neurocomputing6
2021 A CNN-Based Fast Inter Coding Method for VVC
abstract
The Versatile Video Coding (VVC) achieves superior coding efficiency as compared with the High Efficiency Video Coding (HEVC), while its excellent coding performance is at the cost of several high computational complexity coding tools, such as Quad-Tree plus Multi-type Tree (QTMT)-based Coding Units (CUs) and multiple inter prediction modes. To reduce the computational complexity of VVC, a CNN-based fast inter coding method is proposed in this paper. First, a multi-information fusion CNN (MF-CNN) model is proposed to early terminate the QTMT-based CU partition process by jointly using the multi-domain information. Then, a content complexity-based early Merge mode decision is proposed to skip the time-consuming inter prediction modes by considering the CU prediction residuals and the confidence of MF-CNN. Experimental results show that the proposed method reduces an average of 30.63% VVC encoding time, and the Bjøontegaard Delta Bit Rate (BDBR) increases about 3%.
Zhaoqing Pan, Peihan Zhang, Bo Peng 0007, Nam Ling, Jianjun Lei 0001
IEEE Signal Process. Lett.5
2021 Stereoscopic Image Retargeting Based on Deep Convolutional Neural Network
abstract
Stereoscopic image retargeting aims at converting stereoscopic images to the target resolution adaptively. Different from 2D image retargeting, stereoscopic image retargeting needs to preserve both the shape structure of salient objects and depth consistency of 3D scenes. In this paper, we present a stereoscopic image retargeting method based on deep convolutional neural network to obtain high-quality retargeted images with both object shape preservation and scene depth preservation. First, a cross-attention extraction mechanism is constructed to generate attention map, which contains the valuable attention features of the left and right images and the common attention features between them. Second, since the disparity map can provide accurate depth information of objects in 3D scenes, a disparity-assisted 3D significance map generation module is utilized to further preserve the valuable depth information of stereoscopic images. Finally, in order to predict the retargeted stereoscopic images accurately, an image consistency loss is developed to preserve the geometric structure of salient objects, and a disparity consistency loss is introduced to eliminate depth distortions. Experimental results demonstrate that the proposed deep convolutional neural network can provide favorable stereoscopic image retargeting results.
Xiaoting Fan, Jianjun Lei 0001, Jie Liang 0001, Yuming Fang 0001, Nam Ling, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.2
2021 Perceptual Quality Assessment for Asymmetrically Distorted Stereoscopic Video by Temporal Binocular Rivalry
abstract
In this paper, we propose a two-stage weighting based perceptual quality assessment framework for asymmetrically distorted stereoscopic video (SV) sequences by temporal binocular rivalry. Firstly, a traditional 2D image quality assessment (IQA) method is employed to measure spatial distortion, and the temporal distortion is evaluated by the magnitude differences between motion vectors of distorted and reference video frames. Secondly, the structural strength (SS) computed by gradient map and the motion energy (ME) computed by frame difference map are used to estimate the intensity of visual stimulus in spatial and temporal domain respectively. Then, SS and ME are considered as the importance indexes to combine the quality scores of spatial and temporal distortion to estimate perceived distortion of single-view video sequences, which is denoted as the first-stage weighting. Finally, considering that the difference of intensity of visual stimulus between two eyes results in binocular rivalry, a novel temporal binocular rivalry inspired weighting method is designed to integrate the quality scores of left- and right-views for the final visual quality prediction of SV sequences, which is denoted as the second-stage weighting. Experimental results on Waterloo-IVC SV quality databases show that several specific examples of 2D-IQA methods within the proposed framework can obtain highly competitive performance over other existing ones.
Yuming Fang 0001, Xiangjie Sui, Jiheng Wang, Jiebin Yan, Jianjun Lei 0001, Patrick Le Callet
IEEE Trans. Circuits Syst. Video Technol.5
2021 Deep Spatial-Spectral Subspace Clustering for Hyperspectral Image
abstract
Hyperspectral image (HSI) clustering is a challenging task due to the complex characteristics in HSI data, such as spatial-spectral structure, high-dimension, and large spectral variability. In this paper, we propose a novel deep spatial-spectral subspace clustering network (DS3C-Net), which explores spatial-spectral information via the multi-scale auto-encoder and collaborative constraint. Considering the structure correlations of HSI, the multi-scale auto-encoder is first designed to extract spatial-spectral features with different-scale pixel blocks which are selected as the inputs. Then, the collaborative constrained self-expressive layers are introduced between the encoder and decoder, to capture the self-expressive subspace structures. By designing a self-expressiveness similarity constraint, the proposed network is trained collaboratively, and the affinity matrices of the feature representation are learned in an end-to-end manner. Based on the affinity matrices, the spectral clustering algorithm is utilized to obtain the final HSI clustering result. Experimental results on three widely used hyperspectral image datasets demonstrate that the proposed method outperforms state-of-the-art methods.
Jianjun Lei 0001, Bo Peng 0007, Leyuan Fang, Nam Ling, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.1
2021 Deep Stereoscopic Image Super-Resolution via Interaction Module
abstract
Deep learning-based methods have achieved remarkable performance in single image super-resolution. However, these methods cannot be effectively applied in stereoscopic image super-resolution without considering the characteristics of stereoscopic images. In this article, an interaction module-based stereoscopic image super-resolution network (IMSSRnet) is proposed to effectively utilize the correlation information in stereoscopic images. The key insight of the network lies with how to explore the complementary information of one view to help the reconstruction of another view. Thus, an interaction module is designed to acquire the enhanced features by utilizing complementary information between different views. Specifically, the interaction module is composed of a series of interaction units with a residual structure. In addition, the single image features of left and right views are obtained by a spatial feature extraction module, which can be realized by any existing single image super-resolution models. In order to obtain high-quality stereoscopic images, a gradient loss is introduced to preserve the texture details in a view, and a disparity loss is developed to constrain the disparity relationship between different views. Experimental results demonstrate that the proposed method achieves a promising performance and outperforms the state-of-the-art methods.
Jianjun Lei 0001, Zhe Zhang 0041, Xiaoting Fan, Bolan Yang, Ying Chen 0011, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.1
2020 Deep Virtual Reference Frame Generation For Multiview Video Coding
abstract
Multiview video has a large amount of data which brings great challenges to both the storage and transmission. Thus, it is essential to increase the compression efficiency of multiview video coding. In this paper, a deep virtual reference frame generation method is proposed to improve the performance of multiview video coding. Specifically, a parallax-guided generation network (PGG-Net) is designed to transform the parallax relation between different viewpoints and generate a high-quality virtual reference frame. In the network, a multilevel receptive field module is designed to enlarge the receptive field and extract the multi-scale deep features. After that, a parallax attention fusion module is used to transform the parallax and merge the features. The proposed method is integrated into the platform of 3D-HEVC and the generated virtual reference frame is inserted into the reference picture list as an additional reference. Experimental results show that the proposed method achieves 5.31% average BD-rate reduction compared to the 3D-HEVC.
Jianjun Lei 0001, Zongqian Zhang, Dong Liu 0002, Ying Chen 0011, Nam Ling
ICIP1
2020 Attention-Guided Fusion Network of Point Cloud and Multiple Views for 3D Shape Recognition
abstract
With the dramatic growth of 3D shape data, 3D shape recognition has become a hot research topic in the field of computer vision. How to effectively utilize the multimodal characteristics of 3D shape has been one of the key problems to boost the performance of 3D shape recognition. In this paper, we propose a novel attention-guided fusion network of point cloud and multiple views for 3D shape recognition. Specifically, in order to obtain more discriminative descriptor for 3D shape data, the inter-modality attention enhancement module and view-context attention fusion module are proposed to gradually refine and fuse the features of the point cloud and multiple views. In the inter-modality attention enhancement module, the inter-modality attention mask based on the joint feature representation is computed, so that the features of each modality are enhanced by fusing the correlative information between two modalities. After that, the view-context attention fusion module is proposed to explore the context information of multiple views, and fuse the enhanced features to obtain more discriminative descriptor for 3D shape data. Experimental results on the ModelNet40 dataset demonstrate that the proposed method achieves promising performance compared with state-of-the-art methods.
Bo Peng 0007, Zengrui Yu, Jianjun Lei 0001
VCIP3
2020 Joint spatial-spectral hyperspectral image classification based on convolutional neural network
Mengxin Han, Runmin Cong, Huazhu Fu, Jianjun Lei 0001
Pattern Recognit. Lett.5
2020 Semi-Heterogeneous Three-Way Joint Embedding Network for Sketch-Based Image Retrieval
abstract
Sketch-based image retrieval (SBIR) is a challenging task due to the large cross-domain gap between sketches and natural images. How to align abstract sketches and natural images into a common high-level semantic space remains a key problem in SBIR. In this paper, we propose a novel semi-heterogeneous three-way joint embedding network (Semi3-Net), which integrates three branches (a sketch branch, a natural image branch, and an edgemap branch) to learn more discriminative cross-domain feature representations for the SBIR task. The key insight lies with how we cultivate the mutual and subtle relationships amongst the sketches, natural images, and edgemaps. A semi-heterogeneous feature mapping is designed to extract bottom features from each domain, where the sketch and edgemap branches are shared while the natural image branch is heterogeneous to the other branches. In addition, a joint semantic embedding is introduced to embed the features from different domains into a common high-level semantic space, where all of the three branches are shared. To further capture informative features common to both natural images and the corresponding edgemaps, a co-attention model is introduced to conduct common channel-wise feature recalibration between different domains. A hybrid-loss mechanism is designed to align the three branches, where an alignment loss and a sketch-edgemap contrastive loss are presented to encourage the network to learn invariant cross-domain representations. Experimental results on two widely used category-level datasets (Sketchy and TU-Berlin Extension) demonstrate that the proposed method outperforms state-of-the-art methods.
Jianjun Lei 0001, Bo Peng 0007, Zhanyu Ma, Ling Shao 0001, Yi-Zhe Song
IEEE Trans. Circuits Syst. Video Technol.1
2020 Unsupervised Video Action Clustering via Motion-Scene Interaction Constraint
abstract
In the past few years, scene contextual information has been increasingly used for action understanding with promising results. However, unsupervised video action clustering using context has been less explored, and existing clustering methods cannot achieve satisfactory performances. In this paper, we propose a novel unsupervised video action clustering method by using the motion-scene interaction constraint (MSIC). The proposed method takes the unique static scene and dynamic motion characteristics of video action into account, and develops a contextual interaction constraint model under a self-representation subspace clustering framework. First, the complementarity of multi-view subspace representation in each context is explored by single-view and multi-view constraints. Afterward, the context-constrained affinity matrix is calculated and the MSIC is introduced to mutually regularize the disagreement of subspace representation in scene and motion. Finally, by jointly constraining the complementarity of multi-views and the consistency of multi-contexts, an overall objective function is constructed to guarantee the video action clustering result. The experiments on four video benchmark datasets (Weizmann, KTH, UCFsports, and Olympic) demonstrate that the proposed method outperforms the state-of-the-art methods.
Bo Peng 0007, Jianjun Lei 0001, Huazhu Fu, Changqing Zhang 0002, Tat-Seng Chua, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2020 Going From RGB to RGBD Saliency: A Depth-Guided Transformation Model
abstract
Depth information has been demonstrated to be useful for saliency detection. However, the existing methods for RGBD saliency detection mainly focus on designing straightforward and comprehensive models, while ignoring the transferable ability of the existing RGB saliency detection models. In this article, we propose a novel depth-guided transformation model (DTM) going from RGB saliency to RGBD saliency. The proposed model includes three components, that is: 1) multilevel RGBD saliency initialization; 2) depth-guided saliency refinement; and 3) saliency optimization with depth constraints. The explicit depth feature is first utilized in the multilevel RGBD saliency model to initialize the RGBD saliency by combining the global compactness saliency cue and local geodesic saliency cue. The depth-guided saliency refinement is used to further highlight the salient objects and suppress the background regions by introducing the prior depth domain knowledge and prior refined depth shape. Benefiting from the consistency of the entire object in the depth map, we formulate an optimization model to attain more consistent and accurate saliency results via an energy function, which integrates the unary data term, color smooth term, and depth consistency term. Experiments on three public RGBD saliency detection benchmarks demonstrate the effectiveness and performance improvement of the proposed DTM from RGB to RGBD saliency.
Runmin Cong, Jianjun Lei 0001, Huazhu Fu, Junhui Hou, Qingming Huang, Sam Kwong
IEEE Trans. Cybern.2
2020 Region-Enhanced Convolutional Neural Network for Object Detection in Remote Sensing Images
abstract
The convolutional neural networks (CNNs) have recently demonstrated to be a powerful tool for object detection. However, with the complex scenes in remote sensing images, feature extraction of the object in the CNN will be seriously affected by background information. To address this issue, in this article, a region-enhanced CNN (RECNN) is proposed for the object detection of remote sensing images. The RECNN introduces the saliency constraint and multilayer fusion strategy into the CNN model, which can effectively enhance the object regions for better detection. Specifically, the saliency map is extracted and utilized to guide the training of the proposed model to strengthen saliency regions in feature maps. In addition, since different layers can reflect the object regions in varied resolutions, a multilayer fusion strategy is introduced to connect different convolutional layers and explore the context, where the feature maps of object regions are further enhanced. Experimental results on a publicly available ten-class object detection data set demonstrate the superiority of the RECNN over several competitive object detection methods.
Jianjun Lei 0001, Leyuan Fang, Yanfeng Gu
IEEE Trans. Geosci. Remote. Sens.1
2020 A Recursive Constrained Framework for Unsupervised Video Action Clustering
abstract
Video action understanding is an active field of intelligent video analytics, and contextual information in the videos has gained lots of attention for better action understanding. However, most existing works focus on using contextual information for supervised or semi-supervised analysis, and how to effectively use contextual information to boost the unsupervised action clustering performance is still a challenging problem. In this article, we propose a recursive constrained framework for unsupervised video action clustering by utilizing the contextual information of the action and scene. Considering the unique contextual characteristics of video action, action context clustering solution and scene context clustering solution are obtained simultaneously. Based on these two solutions, a recursive priori propagation is proposed to exploit information gain of the priori clustering solutions, and then the information gain is fed back into the procedures of both subspace representation and spectral clustering. Specifically, to explore the unknown relationships in the priori clustering solutions, the constraint-guided subspace representation is introduced by fusing the recursive priori constraint into the self-representation model. Taking priori information and multiview features into consideration, the priori-inherited multiview spectral clustering is proposed to obtain more discriminative spectral embeddings for action clustering. Experiments on three video benchmark datasets demonstrate that the proposed method outperforms state-of-the-art methods.
Bo Peng 0007, Jianjun Lei 0001, Huazhu Fu, Ling Shao 0001, Qingming Huang
IEEE Trans. Ind. Informatics2
2020 Stereoscopic Image Stitching via Disparity-Constrained Warping and Blending
abstract
As a significant branch of virtual reality, stereoscopic image stitching aims to generating wide perspectives and natural-looking scenes. Existing 2D image stitching methods cannot be successfully applied to the stereoscopic images without considering the disparity consistency of stereoscopic images. To address this issue, this paper presents a stereoscopic image stitching method based on disparity-constrained warping and blending, which could avoid visual distortion and preserve disparity consistency. First, a point-line-driven homography based disparity minimization method is designed to pre-align the left and right images and reduce vertical disparity. Afterwards, a multi-constraint warping is proposed to further align the left and right images, where the initial disparity map is introduced to control the consistency of disparities. Finally, a disparity consistency seam-cutting and blending method is presented to determine the optimal seam and conduct stereoscopic image stitching. Experimental results demonstrate that the proposed method achieves competitive performance compared with other state-of-the-art methods.
Xiaoting Fan, Jianjun Lei 0001, Yuming Fang 0001, Qingming Huang, Nam Ling, Chunping Hou
IEEE Trans. Multim.2
2019 Channel-wise Temporal Attention Network for Video Action Recognition
abstract
Recently, video action recognition receives lots of attention, and deep learning based methods have achieved promising performance. Most existing methods focus on spatiotemporal information encoding to learn video representation, which ignore the relevance among channels. In this paper, we propose a novel Channel-wise Temporal Attention Network (CTAN) to explore the fine-grained key information for action recognition. First, the channel-wise attention generation module is proposed to emphasize the fine-grained informative features in each frame. Then, the temporal information aggregation module is introduced before attention generation to exploit the interaction of different frames. Finally, a discriminative video-level representation for action recognition is generated by end-to-end training. Experimental results on two benchmarks, UCF101 and HMDB51, demonstrate the effectiveness of the proposed CTAN.
Jianjun Lei 0001, Yalong Jia, Bo Peng 0007, Qingming Huang
ICME1
2019 Convolutional Neural Network Based Up-Sampling for Depth Video Intra Coding
abstract
Depth video contains depth and disparity information of a scene, which is critical for 3D video systems. In this paper, a convolutional neural network (CNN) based block upsampling method is proposed to improve the efficiency of depth video intra coding. For each largest coding tree in a depth map, it is down-sampled before sent into encoder and recovered into the original size in an intelligent way after low-resolution coding. A novel texture-assisted CNN (TACNN) is presented to handle the depth block up-sampling. The network is made up of several residual coding units and the features of texture block are extracted to assist the reconstruction of the corresponding depth block. Experimental results show that the proposed method achieves competitive rate-distortion performance compared with the state-of-the-art approaches.
Jianjun Lei 0001, Xiao-huan Liu, Kaiming Zhang, Ge Li 0006, Nam Ling
VCIP1
2019 Transferred deep learning based waveform recognition for cognitive passive radar
Qing Wang 0015, Panfei Du, Jing-Yu Yang 0002, Guohua Wang 0002, Jianjun Lei 0001, Chunping Hou
Signal Process.5
2019 Review of Visual Saliency Detection With Comprehensive Information
abstract
The visual saliency detection model simulates the human visual system to perceive the scene and has been widely used in many vision tasks. With the development of acquisition technology, more comprehensive information, such as depth cue, inter-image correspondence, or temporal relationship, is available to extend image saliency detection to RGBD saliency detection, co-saliency detection, or video saliency detection. The RGBD saliency detection model focuses on extracting the salient regions from RGBD images by combining the depth information. The co-saliency detection model introduces the inter-image correspondence constraint to discover the common salient object in an image group. The goal of the video saliency detection model is to locate the motion-related salient object in video sequences, which considers the motion cue and spatiotemporal constraint jointly. In this paper, we review different types of saliency detection algorithms, summarize the important issues of the existing methods, and discuss the existent problems and future works. Moreover, the evaluation datasets and quantitative measurements are briefly introduced, and the experimental analysis and discussion are conducted to provide a holistic overview of different saliency detection methods.
Runmin Cong, Jianjun Lei 0001, Huazhu Fu, Ming-Ming Cheng, Weisi Lin, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.2
2019 Person Re-Identification by Semantic Region Representation and Topology Constraint
abstract
Person re-identification is a popular research topic which aims at matching the specific person in a multi-camera network automatically. Feature representation and metric learning are two important issues for person re-identification. In this paper, we propose a novel person re-identification method, which consists of a reliable representation called semantic region representation (SRR), and an effective metric learning with mapping space topology constraint (MSTC). The SRR integrates semantic representations to achieve effective similarity comparison between the corresponding regions via parsing the body into multiple parts, which focuses on the foreground context against the background interference. To learn a discriminant metric, the MSTC is proposed to consider the topological relationship among all samples in the feature space. It considers two-fold constraints: the distribution of positive pairs should be more compact than the average distribution of negative pairs with regard to the same probe, while the average distance between different classes should be larger than that between same classes. These two aspects cooperate to maintain the compactness of the intra-class as well as the sparsity of the inter-class. Extensive experiments conducted on five challenging person re-identification datasets, VIPeR, SYSU-sReID, QUML GRID, CUHK03, and Market-1501, show that the proposed method achieves competitive performance with the state-of-the-art approaches.
Jianjun Lei 0001, Lijie Niu, Huazhu Fu, Bo Peng 0007, Qingming Huang, Chunping Hou
IEEE Trans. Circuits Syst. Video Technol.1
2019 An Iterative Co-Saliency Framework for RGBD Images
abstract
As a newly emerging and significant topic in computer vision community, co-saliency detection aims at discovering the common salient objects in multiple related images. The existing methods often generate the co-saliency map through a direct forward pipeline which is based on the designed cues or initialization, but lack the refinement-cycle scheme. Moreover, they mainly focus on RGB image and ignore the depth information for RGBD images. In this paper, we propose an iterative RGBD co-saliency framework, which utilizes the existing single saliency maps as the initialization, and generates the final RGBD co-saliency map by using a refinement-cycle model. Three schemes are employed in the proposed RGBD co-saliency framework, which include the addition scheme, deletion scheme, and iteration scheme. The addition scheme is used to highlight the salient regions based on intra-image depth propagation and saliency propagation, while the deletion scheme filters the saliency regions and removes the non-common salient regions based on interimage constraint. The iteration scheme is proposed to obtain more homogeneous and consistent co-saliency map. Furthermore, a novel descriptor, named depth shape prior, is proposed in the addition scheme to introduce the depth information to enhance identification of co-salient objects. The proposed method can effectively exploit any existing 2-D saliency model to work well in RGBD co-saliency scenarios. The experiments on two RGBD co-saliency datasets demonstrate the effectiveness of our proposed framework.
Runmin Cong, Jianjun Lei 0001, Huazhu Fu, Weisi Lin, Qingming Huang, Xiaochun Cao, Chunping Hou
IEEE Trans. Cybern.2
2019 Video Saliency Detection via Sparsity-Based Reconstruction and Propagation
abstract
Video saliency detection aims to continuously discover the motion-related salient objects from the video sequences. Since it needs to consider the spatial and temporal constraints jointly, video saliency detection is more challenging than image saliency detection. In this paper, we propose a new method to detect the salient objects in video based on sparse reconstruction and propagation. With the assistance of novel static and motion priors, a single-frame saliency model is first designed to represent the spatial saliency in each individual frame via the sparsity-based reconstruction. Then, through a progressive sparsity-based propagation, the sequential correspondence in the temporal space is captured to produce the inter-frame saliency map. Finally, these two maps are incorporated into a global optimization model to achieve spatio-temporal smoothness and global consistency of the salient object in the whole video. The experiments on three large-scale video saliency datasets demonstrate that the proposed method outperforms the state-of-the-art algorithms both qualitatively and quantitatively.
Runmin Cong, Jianjun Lei 0001, Huazhu Fu, Fatih Porikli, Qingming Huang, Chunping Hou
IEEE Trans. Image Process.2
2019 Visual Attention Prediction for Stereoscopic Video by Multi-Module Fully Convolutional Network
abstract
Visual attention is an important mechanism in the human visual system (HVS) and there have been numerous saliency detection algorithms designed for 2D images/video recently. However, the research for fixation detection of stereoscopic video is still limited and challenging due to the complicated depth and motion information. In this paper, we design a novel multi-module fully convolutional network (MM-FCN) for fixation detection of stereoscopic video. Specifically, we design a fully convolutional network for spatial saliency prediction (S-FCN), where the initial spatial saliency map of stereoscopic video is learned by image database of object detection. Furthermore, the fully convolutional network for temporal saliency prediction (T-FCN) is constructed by combining saliency results from S-FCN and motion information from video frames. Finally, the fully convolutional network for depth fixation prediction (D-FCN) is designed to compute the final fixation map of stereoscopic video by learning depth features with spatiotemporal features from T-FCN. The experimental results show that the proposed MM-FCN can predict fixation results for stereoscopic video more effectively and efficiently than other related fixation prediction methods.
Yuming Fang 0001, Chi Zhang 0027, Hanqin Huang, Jianjun Lei 0001
IEEE Trans. Image Process.4
2019 Salient Object Detection via Fuzzy Theory and Object-Level Enhancement
abstract
This paper proposes a bottom-up saliency detection method via effective integration of regional saliency measure and object-level information using fuzzy theory. First, we generate an initial saliency map by fusing multiple prior maps. Second, to emphasize the object-level concept of saliency, we further generate many object proposals of the input image. A fuzzy set theory is then applied to measure the objectness score of the object proposals and integrate them into an objectness map. Third, an optimization framework is proposed to effectively fuse various prior saliency cues and object-level information to produce a clean and uniform saliency map as well as to maintain the salient object completeness. Experimental studies in several benchmark datasets confirmed the superiority of the proposed method over state-of-the-art saliency detection methods.
Yuan Zhou 0006, Ailing Mao, Shuwei Huo, Jianjun Lei 0001, Sun-Yuan Kung
IEEE Trans. Multim.4
2019 HSCS: Hierarchical Sparsity Based Co-saliency Detection for RGBD Images
abstract
Co-saliency detection aims to discover common and salient objects in an image group containing more than two relevant images. Moreover, depth information has been demonstrated to be effective for many computer vision tasks. In this paper, we propose a novel co-saliency detection method for RGBD images based on hierarchical sparsity reconstruction and energy function refinement. With the assistance of the intrasaliency map, the inter-image correspondence is formulated as a hierarchical sparsity reconstruction framework. The global sparsity reconstruction model with a ranking scheme focuses on capturing the global characteristics among the whole image group through a common foreground dictionary. The pairwise sparsity reconstruction model aims to explore the corresponding relationship between pairwise images through a set of pairwise dictionaries. In order to improve the intra-image smoothness and inter-image consistency, an energy function refinement model is proposed, which includes the unary data term, spatial smooth term, and holistic consistency term. Experiments on two RGBD co-saliency detection benchmarks demonstrate that the proposed method outperforms the state-of-the-art algorithms both qualitatively and quantitatively.
Runmin Cong, Jianjun Lei 0001, Huazhu Fu, Qingming Huang, Xiaochun Cao, Nam Ling
IEEE Trans. Multim.2
2018 Fast Mode Decision Based on Grayscale Similarity and Inter-View Correlation for Depth Map Coding in 3D-HEVC
abstract
The 3D extension of High Efficiency Video Coding significantly improves the coding efficiency of 3D video at the expense of computational complexity. This paper presents a novel fast mode decision algorithm for depth map coding based on the grayscale similarity and inter-view correlation. First, depth map grayscale similarity is adopted to judge whether the reference frame could assist the coding of the current frame. When the difference in the average grayscale between the co-located coding unit (CU) and the current CU is smaller than the similarity threshold, the depth level of the current CU will be restricted by that of the coded reference CU. Second, the grayscale similarity and inter-view correlation are jointly used for dependent views to achieve early decision on the best prediction unit (PU) mode. The mode decision procedure will be determined early when the co-located CU, which has a grayscale similarity with the current CU, selects Merge or Inter 2N ×2N as the best prediction mode. Moreover, when the corresponding CU in the independent view selects Merge or Inter 2N × 2N as the best prediction mode, the current CU will skip other PU modes checking based on the strong inter-view correlation. Finally, different strategies are proposed for the P-frames and B-frames of dependent views in view of the characteristics of different prediction structures. For B frames, the PU mode information of the coded independent view is utilized as reference to skip the unnecessary mode decision processes. For P frames, the spatial-temporal correlation is considered in the process of early mode decision to determine whether to choose the Merge mode or Inter 2N × 2N as the best mode. Experimental results show that our proposed scheme achieves considerable time saving with negligible degradation of coding performance.
Jianjun Lei 0001, Jinhui Duan, Feng Wu 0001, Nam Ling, Chunping Hou
IEEE Trans. Circuits Syst. Video Technol.1
2018 Region Adaptive R-λ Model-Based Rate Control for Depth Maps Coding
abstract
In this paper, a novel rate-control algorithm based on the region adaptive R-λ model is proposed for depth maps coding. First, in order to obtain an accurate rate control for depth maps coding, a modified frame level bit allocation method based on coding bits statistical distribution of depth maps is proposed. Second, considering that different areas in a depth map have an imparity effect on virtual view rendering, the blocks of the depth map are divided into two types, namely, interested blocks for virtual view rending (IBV) and noninterested blocks for virtual view rending (NIBV). Then, two different R-λ models are derived for IBV and NIBV, respectively. The optimal bitrates for IBV and NIBV are determined by solving an optimization problem. After that, based on the regional R-λ models, the optimal Lagrange multipliers are calculated for both IBV and NIBV. Finally, the largest coding unit (LCU) level rate control is performed by adaptively adjusting the Lagrange multiplier to avoid blocking artifacts and smooth the quality of coding. Experimental results demonstrate that the proposed method can achieve considerable BD-PSNR gains compared with the unified rate-quantization model and conventional R-λ modelbased algorithms in terms of rendered virtual views quality.
Jianjun Lei 0001, Xiaoxu He, Hui Yuan 0001, Feng Wu 0001, Nam Ling, Chunping Hou
IEEE Trans. Circuits Syst. Video Technol.1
2018 Shape-Preserving Object Depth Control for Stereoscopic Images
abstract
In the field of 3-D technology, it is interesting as well as meaningful issue to control object depth in 3-D space. Recently, some depth control methods for stereoscopic images have been proposed, which usually employ depth map or directly process color images to implement depth control. There are two main disadvantages for these methods. First, the results of these methods usually suffer from object deformation and holes. Second, these methods are prone to cause undesired object size changing in 3-D space. To address these issues, we propose a shape-preserving object depth control method for stereoscopic images. First, a novel depth mapping model is presented for calculating the ideal coordinates of the key points in depth control, so that the shape of the object can be well preserved. Afterward, the image content-based constraints are used to further preserve the structure of the object and its background. Finally, the warping technology is introduced to deal with images optimally as well as to avoid holes. Experimental results show that the proposed method can control object depth and preserve the shape of the object effectively without sensible background distortion.
Jianjun Lei 0001, Bo Peng 0007, Changqing Zhang 0002, Xuguang Mei, Xiaochun Cao, Xiaoting Fan, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.1
2018 Optimal Region Selection for Stereoscopic Video Subtitle Insertion
abstract
Stereoscopic subtitle insertion is a fundamental and essential element in stereoscopic film and TV industry. However, little work has been dedicated to the optimal region selection for stereoscopic subtitle insertion. In addition, there is no public database reported for the performance evaluation of it. In this paper, we build the first large-scale video database (TJU3D) for stereoscopic video subtitle insertion, which includes 50 video sequences with rich screen scenes. Compared with 2D subtitle region selection, there are several problems we have to consider in stereoscopic subtitle region selection: 1) the subtitle should avoid depth cue collision and occlusion from objects in stereoscopic video sequences; 2) the disparity value of the subtitle must be minimized to reduce visual discomfort; and 3) the temporal coherence constraint must be considered during region selection for subtitles in video sequences. By considering these constraints, we propose an optimal region selection algorithm for stereoscopic subtitle insertion. First, we compute the disparity map of each video frame in video sequences. For each frame, the optimal position and disparity value of the subtitle are determined by a subtitle region selection algorithm, which contains two parts (i.e., the coarse selection and fine selection). After that, by considering the temporal consistency between adjacent frames, the position and disparity value of each frame are further classified and processed in order to avoid the subtitle jitter. We evaluate the proposed method on TJU3D video database through two visual discomfort prediction metrics and one subjective experiment. To further verify the effectiveness of the proposed method, we also validate the performance of the proposed method on video comfort assessment database, i.e., IEEE-SA Stereo Database. Experimental results demonstrate that the visual discomfort is greatly reduced when using the proposed method compared with the basic method.
Guanghui Yue 0001, Chunping Hou, Jianjun Lei 0001, Yuming Fang 0001, Weisi Lin
IEEE Trans. Circuits Syst. Video Technol.3
2018 Co-Saliency Detection for RGBD Images Based on Multi-Constraint Feature Matching and Cross Label Propagation
abstract
Co-saliency detection aims at extracting the common salient regions from an image group containing two or more relevant images. It is a newly emerging topic in computer vision community. Different from the most existing co-saliency methods focusing on RGB images, this paper proposes a novel co-saliency detection model for RGBD images, which utilizes the depth information to enhance identification of co-saliency. First, the intra saliency map for each image is generated by the single image saliency model, while the inter saliency map is calculated based on the multi-constraint feature matching, which represents the constraint relationship among multiple images. Then, the optimization scheme, namely cross label propagation, is used to refine the intra and inter saliency maps in a cross way. Finally, all the original and optimized saliency maps are integrated to generate the final co-saliency result. The proposed method introduces the depth information and multi-constraint feature matching to improve the performance of co-saliency detection. Moreover, the proposed method can effectively exploit any existing single image saliency model to work well in co-saliency scenarios. Experiments on two RGBD co-saliency datasets demonstrate the effectiveness of our proposed model.
Runmin Cong, Jianjun Lei 0001, Huazhu Fu, Qingming Huang, Xiaochun Cao, Chunping Hou
IEEE Trans. Image Process.2
2018 Iterative Feedback Control-Based Salient Object Segmentation
abstract
In this paper, we establish a mathematical model that relates the control states and the saliency values in salient object detection. We show that a linear feedback control system (LFCS) is amenable to saliency detection tasks owing to its functional properties. This inspired us to employ an LFCS to detect salient objects in static images. Based on the novel iteration method, the system gradually converges to an optimized stable state, which is associated with an accurate saliency map. In addition, to initialize the system, we propose a so-called boundary homogeneity based on a priori knowledge of the boundary in order to estimate the background likelihood and indirectly obtain a foreground (saliency) map. The experimental results indicate that such a feedback control model can offer significant improvement in salient object detection performance.
Shuwei Huo, Yuan Zhou 0006, Jianjun Lei 0001, Nam Ling, Chunping Hou
IEEE Trans. Multim.3
2018 Adaptive Fractional-Pixel Motion Estimation Skipped Algorithm for Efficient HEVC Motion Estimation
abstract
High-Efficiency Video Coding (HEVC) efficiently addresses the storage and transmit problems of high-definition videos, especially for 4K videos. The variable-size Prediction Units (PUs)--based Motion Estimation (ME) contributes a significant compression rate to the HEVC encoder and also generates a huge computation load. Meanwhile, high-level encoding complexity prevents widespread adoption of the HEVC encoder in multimedia systems. In this article, an adaptive fractional-pixel ME skipped scheme is proposed for low-complexity HEVC ME. First, based on the property of the variable-size PUs--based ME process and the video content partition relationship among variable-size PUs, all inter-PU modes during a coding unit encoding process are classified into root-type PU mode and children-type PU modes. Then, according to the ME result of the root-type PU mode, the fractional-pixel ME of its children-type PU modes is adaptively skipped. Simulation results show that, compared to the original ME in HEVC reference software, the proposed algorithm reduces ME encoding time by an average of 63.22% while encoding efficiency performance is maintained.
Zhaoqing Pan, Jianjun Lei 0001, Fu Lee Wang
ACM Trans. Multim. Comput. Commun. Appl.2
2017 Sketch based image retrieval via image-aided cross domain learning
abstract
Existing methods on sketch based image retrieval (SBIR) are usually based on the hand-crafted features whose ability of representation is limited. In this paper, we propose a sketch based image retrieval method via image-aided cross domain learning. First, the deep learning model is introduced to learn the discriminative features. However, it needs a large number of images to train the deep model, which is not suitable for the sketch images. Thus, we propose to extend the sketch training images via introducing the real images. Specifically, we initialize the deep models with extra image data, and then extract the generalized boundary from real images as the sketch approximation. The using of generalized boundary is under the assumption that their domain is similar with sketch domain. Finally, the neural network is fine-tuned with the sketch approximation data. Experimental results on Flicker15 show that the proposed method has a strong ability to link the associated image-sketch pairs and the results outperform state-of-the-arts methods.
Jianjun Lei 0001, Kaifu Zheng, Hua Zhang 0008, Xiaochun Cao, Nam Ling, Yonghong Hou
ICIP1
2017 Simplified search algorithm for explicit wedgelet signalization mode in 3D-HEVC
abstract
As the latest 3D video coding standard, 3D High Efficiency Video Coding (3D-HEVC) achieves efficient coding. Depth coding plays an important role in 3D-HEVC. For better prediction of edges in depth maps, new depth intra modes were proposed, such as Depth Modeling Modes (DMMs). DMM Mode 1, namely explicit wedgelet signalization mode, can improve the performance of synthesized views, but leads to unaffordable encoding computation complexity. In this paper, a simplified search algorithm for explicit wedgelet signalization mode is proposed to predigest the complex search process. First, the proposed algorithm only searches the partition patterns in a limited set based on edge detection in the coarse search stage, since the separation line in the best matching pattern should be similar to the edge in an actual depth prediction unit (PU). Then, a wide range of refinement is implemented to guarantee the coding performance. In fact, up to 24 refinements are tested to provide accurate prediction in the refinement process. Experimental results show that considerable encoding time saving is achieved with negligible performance loss. For all intra test cases, the proposed algorithm achieves an average encoding time saving of 75% for DMM Mode 1 with negligible bitrate increase on synthesized views, compared with the default coarse-refinement algorithm in HTM.
Jianjun Lei 0001, Zhenyan Sun, Zhouye Gu, Nam Ling, Feng Wu 0001
ICME1
2017 Learning visual saliency from human fixations for stereoscopic images
Yuming Fang 0001, Jianjun Lei 0001, Jia Li 0003, Long Xu 0001, Weisi Lin, Patrick Le Callet
Neurocomputing2
2017 Rate control for HEVC based on spatio-temporal context and motion complexity
Yonghong Hou, Jianjun Lei 0001, Wei Xiang 0001, Yao Guo 0006
Multim. Tools Appl.3
2017 Region-based bit allocation and rate control for depth video in HEVC
Jianjun Lei 0001, Xiaoxu He, Chunping Hou
Multim. Tools Appl.1
2017 A divide-and-conquer hole-filling method for handling disocclusion in single-view rendering
abstract
Large holes are unavoidably generated in depth image based rendering (DIBR) using a single color image and its associated depth map. Such holes are mainly caused by disocclusion, which occurs around the sharp depth discontinuities in the depth map. We propose a divide-and-conquer hole-filling method which refines the background depth pixels around the sharp depth discontinuities to address the disocclusion problem. Firstly, the disocclusion region is detected according to the degree of depth discontinuity, and the target area is marked as a binary mask. Then, the depth pixels located in the target area are modified by a linear interpolation process, whose pixel values decrease from the foreground depth value to the background depth value. Finally, in order to remove the isolated depth pixels, median filtering is adopted to refine the depth map. In these ways, disocclusion regions in the synthesized view are divided into several small holes after DIBR, and are easily filled by image inpainting. Experimental results demonstrate that the proposed method can effectively improve the quality of the synthesized view subjectively and objectively.
Jianjun Lei 0001, Cuicui Zhang, Kefeng Fan, Chunping Hou
Multim. Tools Appl.1
2017 Near-Optimal Cross-Layer Forward Error Correction Using Raptor and RCPC Codes for Prioritized Video Transmission Over Wireless Channels
abstract
Cross-layer forward error correction (FEC) aims at utilizing available bandwidth more efficiently, which has been applied to error-prone video transmission over imperfect wireless channels. In this paper we propose a new near-optimal cross-layer FEC scheme in which systematic Raptor codes are used at the application layer and rate compatible punctured convolutional (RCPC) codes are used at the physical layer for H.264/AVC encoded video streaming with channel bandwidth constraints. In the proposed scheme, in order to fully exploit the unequal importance of compressed video data, we assign each source packet a different priority according to its contribution to the reconstructed video quality. We first obtain the transmission parameters, which satisfies the conditions for optimal video transmission, for the optimal cross-layer Raptor-RCPC FEC in the ideal situation through a theoretical analysis, and then we propose a heuristic algorithm searching from the optimal solution point to obtain the transmission parameters, which are near optimal in the practical situation. Computer simulation results show that the proposed scheme can achieve significant performance improvements in both the additive white Gaussian noise and Rayleigh channels compared with the previous work.
Yonghong Hou, Wei Xiang 0001, Maode Ma, Jianjun Lei 0001
IEEE Trans. Circuits Syst. Video Technol.5
2017 Stereoscopic Image Stitching Based on a Hybrid Warping Model
abstract
Traditional image editing techniques cannot be directly used to process stereoscopic media, as extra constraints are required to ensure consistent changes between the left and right images. In this paper, we propose a hybrid warping model for stereoscopic image stitching by combining projective and content-preserving warping. First, a uniform homography algorithm is proposed to prewarp the left and right images, and thus ensure consistent changes. Second, a content-preserving warping is introduced to locally refine alignment and reduce vertical disparities. Finally, a seam-cutting-based algorithm is used to find a blending seam, and the multiband blending algorithm is used to produce the final stitched image. Experimental results show that the proposed method can effectively stitch stereoscopic images, which not only avoids local distortions, but also reduces vertical disparities reasonably.
Weiqing Yan, Chunping Hou, Jianjun Lei 0001, Yuming Fang 0001, Zhouye Gu, Nam Ling
IEEE Trans. Circuits Syst. Video Technol.3
2017 Visual Attention Modeling for Stereoscopic Video: A Benchmark and Computational Model
abstract
In this paper, we investigate the visual attention modeling for stereoscopic video from the following two aspects. First, we build one large-scale eye tracking database as the benchmark of visual attention modeling for stereoscopic video. The database includes 47 video sequences and their corresponding eye fixation data. Second, we propose a novel computational model of visual attention for stereoscopic video based on Gestalt theory. In the proposed model, we extract the low-level features, including luminance, color, texture, and depth, from discrete cosine transform coefficients, which are used to calculate feature contrast for the spatial saliency computation. The temporal saliency is calculated by the motion contrast from the planar and depth motion features in the stereoscopic video sequences. The final saliency is estimated by fusing the spatial and temporal saliency with uncertainty weighting, which is estimated by the laws of proximity, continuity, and common fate in Gestalt theory. Experimental results show that the proposed method outperforms the state-of-the-art stereoscopic video saliency detection models on our built large-scale eye tracking database and one other database (DML-ITRACK-3D).
Yuming Fang 0001, Chi Zhang 0027, Jing Li 0026, Jianjun Lei 0001, Matthieu Perreira Da Silva, Patrick Le Callet
IEEE Trans. Image Process.4
2017 Depth Map Super-Resolution Considering View Synthesis Quality
abstract
Accurate and high-quality depth maps are required in lots of 3D applications, such as multi-view rendering, 3D reconstruction and 3DTV. However, the resolution of captured depth image is much lower than that of its corresponding color image, which affects its application performance. In this paper, we propose a novel depth map super-resolution (SR) method by taking view synthesis quality into account. The proposed approach mainly includes two technical contributions. First, since the captured low-resolution (LR) depth map may be corrupted by noise and occlusion, we propose a credibility based multi-view depth maps fusion strategy, which considers the view synthesis quality and interview correlation, to refine the LR depth map. Second, we propose a view synthesis quality based trilateral depth-map up-sampling method, which considers depth smoothness, texture similarity and view synthesis quality in the up-sampling filter. Experimental results demonstrate that the proposed method outperforms state-of-the-art depth SR methods for both super-resolved depth maps and synthesized views. Furthermore, the proposed method is robust to noise and achieves promising results under noise-corruption conditions.
Jianjun Lei 0001, Lele Li, Huanjing Yue, Feng Wu 0001, Nam Ling, Chunping Hou
IEEE Trans. Image Process.1
2017 Depth-Preserving Stereo Image Retargeting Based on Pixel Fusion
abstract
In this paper, we propose a pixel fusion-based stereo image retargeting method, which could adaptively retarget stereo images with flexible aspect ratios, simultaneously preserving the depth. Retargeting each image independently by the pixel fusion method ignores the disparity relationship between pixels in the image pair and hence will introduce the distortion of disparity. To address this issue, we advocate to extend the single pixel fusion-based way to be applicable for stereo image pair. First, seams are selected based on the energy function, which simultaneously considers the seam selecting and seam matching. Second, a seam-matching-based matching map is proposed to preserve the disparity relationship between image pair. Then, the scaling factors for the left image are assigned considering both the important object and depth preservation. Subsequently, the scaling factors for the right image are obtained according to the proposed matching map. Based on these scaling factors, the stereo image pair is retargeted with pixel fusion. In contrast to removing pixels to resize image, the way of pixel fusion can obtain more smooth results with less depth distortion. Experimental results demonstrate that our method achieves more preferable qualities in both depth and shape preservation for stereo image retargeting.
Jianjun Lei 0001, Changqing Zhang 0002, Feng Wu 0001, Nam Ling, Chunping Hou
IEEE Trans. Multim.1
2016 Early DIRECT mode decision based on all-zero block and rate distortion cost for multiview video coding
abstract
The exhaustive variable‐block‐size mode decision can efficiently remove the redundancies among the multiview videos, while it also leads to significant increase of computational complexity in the multiview video coding (MVC) encoder, and the high encoding complexity becomes a bottleneck for the MVC encoder to achieve real‐time multimedia applications. To address this bottleneck, many fast mode decision methods have been proposed. However, most of them are only suitable for optimising the encoding complexity of the odd views of the MVC encoder. In this study, based on the property of the all‐zero block and rate distortion (RD) cost of the DIRECT mode as well as the correlations between the current macroblock (MB) and its spatial–temporal nearby MBs, an early DIRECT mode decision method is proposed for reducing the encoding complexity of the MVC. Experimental results show that the proposed method achieves 48.25 and 55.64% on average encoding time saving for the even and odd views, respectively, whereas the RD performance degradation is quite acceptable. In summary, the proposed method efficiently reduces the encoding complexity for the MVC encoder.
Zhaoqing Pan, Yun Zhang 0002, Jianjun Lei 0001, Long Xu 0001, Xingming Sun
IET Image Process.3
2016 Saliency-based stereoscopic image retargeting
Yuming Fang 0001, Junle Wang, Yuan Yuan 0029, Jianjun Lei 0001, Weisi Lin, Patrick Le Callet
Inf. Sci.4
2016 Fast reference frame selection based on content similarity for low complexity HEVC encoder
Zhaoqing Pan, Jianjun Lei 0001, Yun Zhang 0002, Xingming Sun, Sam Kwong
J. Vis. Commun. Image Represent.3
2016 Saliency Detection for Stereoscopic Images Based on Depth Confidence Analysis and Multiple Cues Fusion
abstract
Stereoscopic perception is an important part of human visual system that allows the brain to perceive depth. However, depth information has not been well explored in existing saliency detection models. In this letter, a novel saliency detection method for stereoscopic images is proposed. First, we propose a measure to evaluate the reliability of depth map, and use it to reduce the influence of poor depth map on saliency detection. Then, the input image is represented as a graph, and the depth information is introduced into graph construction. After that, a new definition of compactness using color and depth cues is put forward to compute the compactness saliency map. In order to compensate the detection errors of compactness saliency when the salient regions have similar appearances with background, foreground saliency map is calculated based on depth-refined foreground seeds' selection (DRSS) mechanism and multiple cues contrast. Finally, these two saliency maps are integrated into a final saliency map through weighted-sum method according to their importance. Experiments on two publicly available stereo data sets demonstrate that the proposed method performs better than other ten state-of-the-art approaches.
Runmin Cong, Jianjun Lei 0001, Changqing Zhang 0002, Qingming Huang, Xiaochun Cao, Chunping Hou
IEEE Signal Process. Lett.2
2016 A Universal Framework for Salient Object Detection
abstract
In this paper, we propose a novel universal framework for salient object detection, which aims to enhance the performance of any existing saliency detection method. First, rough salient regions are extracted from any existing saliency detection model with distance weighting, adaptive binarization, and morphological closing. With the superpixel segmentation, a Bayesian decision model is adopted to refine the rough saliency map to obtain a more accurate saliency map. An iterative optimization method is designed to obtain better saliency results by exploiting the characteristics of the output saliency map each time. Through the iterative optimization process, the rough saliency map is updated step by step with better and better performance until an optimal saliency map is obtained. Experimental results on the public salient object detection datasets with ground truth demonstrate the promising performance of the proposed universal framework subjectively and objectively.
Jianjun Lei 0001, Bingren Wang, Yuming Fang 0001, Weisi Lin, Patrick Le Callet, Nam Ling, Chunping Hou
IEEE Trans. Multim.1
2015 Fast Transform Unit Depth Decision Based on Quantized Coefficients for HEVC
abstract
The quad tree structure based Transform Unit (TU) helps high efficiency video coding to improve the coding efficiency. However, the achieved coding efficiency comes at the cost of the increased computational complexity. In this paper, based on the quantizated coefficients of the TU, we propose an early termination for the quad tree structure based TU encoding process. If the quantized coefficients of the luminance components are all zeros, the TU encoding process will be terminated. Experimental results show that the proposed method achieves about 55.13% on average TU encoding time saving, while the rate distortion performance degradation is negligible.
Zhaoqing Pan, Jianjun Lei 0001, Yun Zhang 0002, Sam Kwong
SMC2
2015 View generation with DIBR for 3D display system
Laihua Wang, Chunping Hou, Jianjun Lei 0001, Weiqing Yan
Multim. Tools Appl.3
2015 Depth Coding Based on Depth-Texture Motion and Structure Similarities
abstract
This paper addresses high performance depth coding in 3D video by making good use of its coded texture video counterpart. The relationship between the depth and its associated texture video in terms of coding mode and motion vector is carefully examined. Our statistical study suggests that the skip-coding mode and its associated motion vectors in the coded texture can be shared for depth coding by saving bit rate at the cost of little increase of distortion, which subsequently results in a nonsequential coding of the depth map. In this sense, coding/prediction of a block can be performed using the skip-coded blocks below and right, which are not available in the conventional sequential coding, thus producing the so-called omnidirectional blocks predicted in the intra-coding by making the best use of (at most) four neighboring blocks. Moreover, in view of the depth-texture structure similarity, a depth-texture cooperative clustering-based prediction method is proposed for cluster-based depth prediction in the intra-coding, which exploits the structure similarity for the current coding block and its neighboring pixels around the block. On the other hand, some large prediction errors may be present for the depth-texture misaligned pixels, which may greatly compromise the coding performance. To deal with these large residuals induced by the depth-texture misalignment, a simple yet effective detection and rectification approach is incorporated in the proposed depth coding scheme. Experimental results show that our proposed depth coding scheme achieves superior rate-distortion performance compared with other relevant coding methods.
Jianjun Lei 0001, Shuai Li 0005, Ce Zhu, Ming-Ting Sun, Chunping Hou
IEEE Trans. Circuits Syst. Video Technol.1
2015 Fast Mode Decision Using Inter-View and Inter-Component Correlations for Multiview Depth Video Coding
abstract
With the development of three-dimensional (3-D) display technologies, 3-D video has attracted more and more interest. Multiview video plus depth (MVD) is one of the most popular representation formats of 3-D video. In MVD coding system, multiview depth video needs to be coded and transmitted in addition to the texture video. This paper presents a novel fast mode decision (FMD) method for odd views in multiview depth video coding. First, the inter-view and inter-component coding correlations are analyzed to provide efficient reference information. Then, with a view to the characteristics of different types of frames, different early termination strategies are proposed. For the nonanchor frame, the early termination criterion is based on the rate-distortion cost information of the even views and the coded block pattern information. For the anchor frame, the criterion is set stricter to maintain the coding accuracy. Experimental results show that the proposed method can reduce 78.07% coding time on average, without significant loss of video quality.
Jianjun Lei 0001, Jing Sun 0010, Zhaoqing Pan, Sam Kwong, Jinhui Duan, Chunping Hou
IEEE Trans. Ind. Informatics1
2015 Depth Sensation Enhancement for Multiple Virtual View Rendering
abstract
Depth information is an indispensable element in depth image-based rendering (DIBR) for three-dimensional (3-D) display. In this paper, we propose a novel depth sensation enhancement method to address the problems in multiple virtual view rendering. First, as the depth sensation is decreased when rendering intermediate multiple virtual views, the basic principle of depth sensation enhancement is derived according to the number of rendering views. Second, with the increase of the scene complexity, it is difficult to ensure the depth sensation of all neighboring objects. The saliency analysis is adopted to give preferred guarantee to the depth sensation between the salient object and its neighbors. Then, the depth sensation enhancement for multiple virtual view rendering is performed based on a defined energy function built by the number of rendering views and the saliency analysis. Finally, considering the temporal consistency between adjacent frames, the depth sensation enhancement is extended to video applications with a newly designed energy function with energy term of temporal consistency preservation. Experimental results on a public database demonstrate that the proposed method can obtain promising performance in depth sensation.
Jianjun Lei 0001, Cuicui Zhang, Yuming Fang 0001, Zhouye Gu, Nam Ling, Chunping Hou
IEEE Trans. Multim.1
2014 Fast Coding Tree Unit depth decision for high efficiency video coding
abstract
High Efficiency Video Coding (HEVC) is the latest video coding standard, which adapts quadtree structure based Coding Tree Unit (CTU) to improve the coding efficiency. In HEVC encoding process, the CTU is recursively partitioned into coding units according to the quadtree depth. This technique increases the coding efficiency of HEVC, however, the achieved coding efficiency comes at the cost of high computational complexity. In this paper, we propose a fast C-TU quadtree depth decision algorithm to reduce the computational complexity of HEVC. Firstly, based on the best C-TU depth correlation among spatial and temporal neighboring CTUs, an early quadtree depth 0 decision algorithm is proposed. Then, according to the correlation between the prediction unit mode and the best CTU depth selection, a quadtree depth 3 skipped decision algorithm is proposed. Experimental results show that the proposed algorithm can achieve 40% on average encoding time saving, while maintaining a comparable rate-distortion performance.
Zhaoqing Pan, Sam Kwong, Yun Zhang 0002, Jianjun Lei 0001, Hui Yuan 0001
ICIP4
2014 Rate control of hierarchical B prediction structure for multi-view video coding
Jianjun Lei 0001, Meimin Wu, Shuai Li 0005, Chunping Hou
Multim. Tools Appl.1
2014 Pixel-Based Inter Prediction in Coded Texture Assisted Depth Coding
abstract
This letter presents a pixel-based motion estimation scheme assisted with the coded texture video for depth inter-prediction, in view of motion similarity between depth and texture video. The proposed scheme can achieve higher inter-prediction gain without transmitting any motion vector in the pixel-based motion estimation. Coupled with depth-texture structure similarity, the inter prediction method is further extended to an integrated prediction approach by making use of both intra and inter information. Experimental results show that our proposed method achieves superior rate-distortion performance.
Shuai Li 0005, Jianjun Lei 0001, Ce Zhu, Lu Yu 0003, Chunping Hou
IEEE Signal Process. Lett.2
2013 Evaluation and modeling of depth feature incorporated visual attention for salient object segmentation
Jianjun Lei 0001, Hailong Zhang 0011, Chunping Hou, Laihua Wang
Neurocomputing1
2013 The objective quality assessment of stereo image
Nan Yun, Zhiyong Feng 0002, Jianjun Lei 0001
Neurocomputing4
2006 Multi-scale Support Vector Machine for Regression Estimation
Zhen Yang 0004, Jun Guo 0002, Weiran Xu, Xiangfei Nie, Jianjun Lei 0001
ISNN (1)6