Bo Peng 0007

dblp:03/5954-7 · DBLP profile ↗
← Back
52ranked-venue papers
16as first author
45since 2021 · last 2026
0000-0002-6616-453XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 44 · 12 first-author · 38 since 2021Computer networks · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Spinal Lesion Detection in X-Ray Images via Uncertainty-Guided Classification and Localization
Lisha Guo, Bo Peng 0007, Jianjun Lei 0001, Xu Zhang 0045, Qingming Huang
IEEE Signal Process. Lett.2
2026 Mining Temporal Redundancy Using Long Short-Term Motion Aggregation and Global-Local Decorrelation for Learned Video Compression
abstract
The conditional coding paradigm is widely used in learned video compression, which shows superior performance in capturing redundancies within a large context space. However, existing Conditional coding-based Learned Video Compression (C-LVC) methods ignore that the predicted motion vectors usually contain large uncertainty due to complex motions, occlusions, etc., which consequently decrease the accuracy of the generated temporal contexts. In addition, existing C-LVC methods have a weak ability to mine diverse dependencies within the context space, which are closely related to the coding efficiency. To address these issues, an efficient temporal redundancy mining method is proposed to improve the coding efficiency of C-LVC in this paper. To generate accurate temporal contexts, a Long Short-Term Motion Aggregation (LSTMA) model is proposed, in which an LSTMA-based motion estimation module is developed to capture both current and aggregated long short-term motion information to reduce the uncertainty of predicted motion vectors. Based on the dual motion information, an LSTMA-based temporal context mining module is developed to exploit the aggregated long short-term motion information and increase the accuracy of the generated temporal contexts. In order to fully eliminate spatial-temporal redundancies in a video, a Global-Local Information Decorrelation Module (GLIDM)-based context codec is proposed, in which the GLIDM is designed based on the visual state space block (namely vmamba), the residual block, and the squeeze-and-excitation block to effectively capture long-range, short-range spatial-temporal dependencies and channel-wise dependencies. Experimental results demonstrate that our proposed method can effectively improve the coding performance of C-LVC, and outperforms other state-of-the-art LVC methods.
Zhaoqing Pan, Jianjun Lei 0001, Bo Peng 0007, Haoran Xie 0001, Fu Lee Wang, Sam Kwong
IEEE Trans. Circuits Syst. Video Technol.4
2026 Uncertainty-Aware Multi-View Graph Clustering
abstract
Multi-view clustering has achieved advanced progress over the years, which typically integrates multi-view information to learn discriminative common representations or a unified clustering distribution for clustering. However, existing methods either simply regard each view as equally important or assign fixed weights to each view, which are insufficient to dynamically assess the sample quality variations caused by noise in multi-view data. To address this issue, by effectively modeling the uncertainty of different samples across different views, this paper proposes a novel uncertainty-aware multi-view graph clustering network, termed UMGC-Net, which achieves trusted multi-view clustering in an unsupervised manner. Specifically, by measuring the clustering distribution entropy, an uncertainty-guided common feature learning mechanism is proposed to estimate the uncertainty for each sample of each view, thus learning multi-view features friendly to clustering. Besides, a cross-view trusted distribution fusion module is designed to obtain robust clustering distribution by exploring the trusted consistency among multi-view clustering distributions based on uncertainty. Finally, experimental results on four popular multi-view datasets validate the superior performance of the proposed UMGC-Net.
Bo Peng 0007, Shaobo Bai, Jianjun Lei 0001, Changqing Zhang 0002, Nam Ling
IEEE Trans. Multim.1
2026 Depth-Aware Transformer for Aerial Localization
abstract
Recently, deep learning-based visual localization has gained significant attention and made remarkable advancements. Although previous visual localization methods have obtained promising performance on indoor or outdoor street scenes, there have been few attempts at visual localization on aerial scenes. In this article, a depth-aware aerial localization transformer (DALTR) is proposed to learn camera poses in real-world aerial scenes assisted by the depth map. To improve the ability of network to perceive on aerial scenes, a multi-level depth embedding transformer module is presented by adaptively incorporating depth information into multiple levels of transformer. In addition, to encourage the piece-wise smooth geometric characteristic of the scene coordinates, a depth-guided smoothness constraint is developed to provide additional supervision for scene coordinate regression. Extensive experimental results on aerial localization benchmark datasets demonstrate that the proposed DALTR achieves superior aerial localization performance.
Jianjun Lei 0001, Duohui Tu, Bo Peng 0007, Zhe Zhang 0041, Chong Wu 0004, Qingming Huang
ACM Trans. Multim. Comput. Commun. Appl.3
2026 Meta-Learned Zero-Shot Sketch-Based Point Cloud Retrieval via Perspective-Predicted Feature Learning
abstract
In recent times, sketch-based 3D shape retrieval has emerged as a pivotal theme and garnered considerable attention within the area of cross-modal retrieval. As a prevalent 3D shape modality, the exponential growth in the quantity of 3D point clouds has boosted a substantial increase in the demand for 3D point cloud retrieval. Simultaneously, due to the absence of prior knowledge about unseen classes, transferring models learned from seen classes to tackle the data from unseen classes effectively remains a significant hurdle in cross-modal retrieval. In light of this, a novel meta-learned zero-shot sketch-based point cloud retrieval (MetaZS-SBPR) network is proposed in this article for exploring cross-modal consistent feature representation from 2D sketches and 3D point clouds, while effectively transferring the knowledge from seen classes to unseen classes. Specifically, a perspective-predicted point cloud feature learning module is presented to capture discriminative features of point clouds from predicted perspectives, thereby mitigating the modal differences across point clouds and sketches. Additionally, a meta zero-shot retrieval strategy is introduced to investigate the knowledge transfer from seen classes to unseen classes harnessing meta-learning, thereby enabling the efficient retrieval of the point clouds from unseen classes. Experimental evaluations conducted on the ZS-SBPR benchmark dataset affirm the effectiveness of the proposed MetaZS-SBPR.
Bo Peng 0007, Menglei Zhao, Qingming Huang, Jianjun Lei 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2025 Hypergraph Contrastive Learning for Large-Scale Hyperspectral Image Clustering
abstract
Large-scale hyperspectral image (HSI) clustering has become an important research task owing to its promising applications in various fields. Recently, beneficial from the correlation modeling capability of graphs, graph contrastive learning methods have received increasing attention in the clustering task. However, these methods usually have limited ability to explore the high-order correlation as well as beneficial clustering information of large-scale HSI, thus limiting the clustering performance on large-scale HSI. To this end, a novel hypergraph contrastive learning network (HCL-Net) for large-scale HSI clustering is proposed in this paper. Specifically, a diffusion hypergraph-based contrastive clustering mechanism is presented, in which a diffusion hypergraph is constructed to model the high-order correlation in large-scale HSI, thus guiding contrastive learning for obtaining more discriminative representations. Besides, by mining the confident clustering information, a confidence-guided positive-negative updating strategy is designed to dynamically update positives and negatives for contrastive learning, thereby obtaining a more compact clustering structure. The proposed method is evaluated on three public large-scale HSI datasets. The experimental results have demonstrated the superior performance of the proposed HCL-Net over state-of-the-art methods.
Bo Peng 0007, Tianyi Qin, Yanfeng Gu, Nam Ling, Jianjun Lei 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Adversarially Robust Object Detection via Deviation Calibration and Content Preservation
abstract
Object detection has achieved a promising development in recent years and played an important role in various applications. However, the performance of object detection networks generally drops significantly when subjected to adversarial attacks. As an effective technique for defending against adversarial attacks, adversarially robust object detection has attracted increasing interest. In this paper, a novel deviation-calibrated and content-preserved network (DCCP-Net) is proposed for adversarially robust object detection by effectively exploring and mitigating the essential negative impact of noise disturbance in the feature space. Specifically, a deviation-calibrated robust feature enhancement module is designed to enhance the feature robustness of adversarial images by removing noise disturbance and supplementing rectified information. Besides, by enabling adversarial image features to imitate corresponding clean image features, a content-preserved consistency information imitation mechanism is proposed to obtain more accurate content information of adversarial images. Extensive experiment results have verified the superiority of the proposed DCCP-Net.
Xu Zhang 0045, Bo Peng 0007, Jianjun Lei 0001, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.2
2025 Cross-Modal Aligned Identity-Discriminative Feature Learning Network for Face Sketch Recognition
abstract
Face sketch recognition focuses on retrieving face photos that have the same identity as query face sketches, and plays a vital role in the field of information forensics and security. Owing to the large cross-modal differences between face sketches and photos, extracting and aligning cross-modal features is still considered a challenging task in the face sketch recognition community. This paper presents a novel cross-modal aligned identity-discriminative feature learning network (CAIFL-Net) for face sketch recognition. Specifically, in this paper, an identity-discriminative feature preservation module is designed to capture the identity-discriminative features of face sketches and photos by eliminating features that are weakly related to recognition. In addition, a sketch-photo cross-reconstructed feature alignment module is proposed to obtain cross-modal aligned features for effective recognition by reconstructing and embedding global features of one modality into another. Extensive experiments on the Uom-SGFS and CUFSF datasets demonstrate the effectiveness of the proposed CAIFL-Net.
Jianjun Lei 0001, Menglei Zhao, Bo Peng 0007, Qingming Huang
IEEE Trans. Inf. Forensics Secur.3
2025 Advancing Real-World Stereoscopic Image Super-Resolution via Vision-Language Model
abstract
Recent years have witnessed the remarkable success of the vision-language model in various computer vision tasks. However, how to exploit the semantic language knowledge of the vision-language model to advance real-world stereoscopic image super-resolution remains a challenging problem. This paper proposes a vision-language model-based stereoscopic image super-resolution (VLM-SSR) method, in which the semantic language knowledge in CLIP is exploited to facilitate stereoscopic image SR in a training-free manner. Specifically, by designing visual prompts for CLIP to infer the region similarity, a prompt-guided information aggregation mechanism is presented to capture inter-view information among relevant regions between the left and right views. Besides, driven by the prior knowledge of CLIP, a cognition prior-driven iterative enhancing mechanism is presented to optimize fuzzy regions adaptively. Experimental results on four datasets verify the effectiveness of the proposed method.
Zhe Zhang 0041, Jianjun Lei 0001, Bo Peng 0007, Liying Xu, Qingming Huang
IEEE Trans. Image Process.3
2025 Advancing Generalizable Occlusion Modeling for Neural Human Radiance Field
abstract
Generalizable human neural rendering aims to render the target views of the human body by leveraging source views and the skinned multi-person linear (SMPL) model. Despite exhibiting promising performance, the target views rendered by previous methods usually contain corrupted parts of the human body. Two primary challenges hinder high-quality human neural rendering. These challenges involve non-correspondences between 2D pixels and 3D SMPL vertices induced by self-occlusion of the human body and erroneous appearance predictions caused by occlusion between the source and target views. To solve these two challenges, we propose an advancing generalizable occlusion modeling method for the neural human radiance field, in which the hurdles from the self-occlusion of the human body and the occlusion between source and target views are explored and solved. Specifically, to alleviate the non-correspondence problem induced by self-occlusion, a geometry perception module is designed to obtain 3D geometric representations of SMPL vertices, enabling the prediction of accurate density values. Furthermore, a visibility aggregation module is designed to estimate the visibility maps with respect to different source views by utilizing the predicted density. Then, the complementary information among multiple source views is integrated with the support of the visibility maps in the visibility aggregation module, thus effectively addressing the occlusion between views. Experiments on the ZJU-MoCap and THUman datasets show that the proposed method achieves promising performance compared with the existing state-of-the-art methods.
Bingzheng Liu, Jianjun Lei 0001, Bo Peng 0007, Zhe Zhang 0041, Qingming Huang
IEEE Trans. Multim.3
2025 Efficient Chroma Intra Prediction via Exemplar Colorization Network for Versatile Video Coding
abstract
Chroma intra prediction aims to reduce chroma redundancies within a frame, which plays an important role in improving the coding efficiency of intra coding. Existing chroma intra prediction methods typically utilize the spatial relationship between the current luma block and its neighboring reference luma blocks to predict its chroma samples. However, the spatial properties of luma components differ from those of chroma components, which limits the accuracy of chroma intra prediction. To tackle this issue, an efficient Exemplar Colorization Network (ECNet)-based chroma intra prediction method is proposed in this paper, in which the colorization relationship between reference luma and chroma components is exploited to predict the chroma components for the current luma component. Inspired by the principle that semantic information in an image exhibits short-range continuity, a Spatial-consistency-based Colorization Transfer Network (SCTNet) is proposed, which builds and transfers colorization representations of neighboring reference blocks for chroma prediction. To improve the chroma prediction capability of SCTNet, a colorization learning module is developed to learn the robust mapping relationship from the luma component to the chroma component in a region-to-pixel manner, and a weight-adaptive reconstruction module is designed to adaptively utilize reference information from neighboring blocks to generate an initial prediction result. In addition, to further improve the accuracy of chroma intra prediction, a multi-reference-based chroma refinement network is proposed, which simultaneously uses the spatial information of neighboring reference chroma blocks and the current luma block to eliminate blocking and color-bleeding artifacts in the initial prediction result. Experimental results demonstrate that our proposed ECNet outperforms the state-of-the-art chroma intra prediction methods in terms of coding performance.
Zhaoqing Pan, Jixing Chen, Bo Peng 0007, Jianjun Lei 0001, Fu Lee Wang, Nam Ling, Sam Kwong
IEEE Trans. Multim.3
2025 Modeling Intra- and Inter-Modal Correlations for Incomplete Multi-Modal 3D Shape Clustering
abstract
The investigation for incomplete multi-modal 3D shape clustering is evolving as a promising task for the field of recognizing massive unlabeled 3D shapes. As two widely adopted 3D shape modalities, point clouds and multiple views not only exhibit rich intra-modal correlations but also encompass complementary structures and appearances of 3D shapes. By effectively modeling the intra-modal and inter-modal correlations, this paper proposes a novel incomplete multi-modal 3D shape clustering method to reveal the underlying clustering associations from incomplete multi-modal 3D shapes. In detail, a similarity-transferred feature prediction module is presented to recover the features of missing instances within one modality with the assistance of similarity exploring from another modality. Then, an intra-to-inter progressive feature fusion module is designed to mine the correlations within the modality as well as between different modalities, thereby obtaining comprehensive 3D shape features for clustering. Extensive experiments on two public 3D shape datasets have demonstrated that the proposed method has achieved promising clustering results under different missing rates.
Tianyi Qin, Bo Peng 0007, Jianjun Lei 0001, Qingming Huang
IEEE Trans. Multim.2
2025 Adaptive Multi-Exposure Image Correction via Joint Lightness and Structure Awareness
abstract
In order to alleviate the impact of ambient light on the quality of captured images, correcting multi-exposure images has become a popular topic. Most existing multi-exposure image correction methods mainly focus on the adjustment of lightness levels, but ignore the significant issue of structural information loss in incorrectly exposed images. Taking into consideration both lightness adjustment and structural reconstruction, this article proposes an adaptive multi-exposure image correction network by jointly exploring the lightness and structure information, named LSANet. Specifically, the proposed LSANet first extracts lightness and structure representations of the input image in the frequency domain, and then performs exposure level adjustment and structure detail reconstruction based on the lightness and structure representations. In the proposed network, the lightness- and structure-aware adaptive module is designed to achieve adaptive correction by predicting dynamic kernels under the guidance of the lightness and structure representations. Experimental results on the widely used ME and SICE datasets demonstrate that the proposed LSANet achieves excellent performance and generates images with well-exposed levels and rich structural details.
Bo Peng 0007, Jia Zhang 0025, Zhe Zhang 0041, Liying Xu, Qingming Huang, Tao Wang 0119, Jianjun Lei 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2024 Few-Shot Object Detection with Instance Feature Generation and Hybrid Contrastive Learning
abstract
Few-shot object detection aims to effectively detect novel classes with limited annotated samples. Due to the low-quality instance features obtained by deep learning in data-scarce scenarios, few-shot object detection remains a significant challenge. In this paper, a novel method is proposed to enhance the few-shot object detection performance by focusing on the generalizability and discriminability of instance features. In detail, by synthesizing auxiliary instances that embed diverse attributes, a mask-guided instance feature generation module is presented to alleviate the overfitting on sample-specific characteristics, thereby facilitating the acquisition of generalizable object-relevant knowledge. Then, to learn discriminative object-relevant attributes, a hybrid cross-layer and intra-layer contrastive learning mechanism is designed to enhance the discriminability of instance features by building contrastive constraints between instances within and across layers. Experimental results on two widely used benchmarks demonstrate the effectiveness of the proposed method.
Bo Peng 0007, Tianyi Qin, Xu Zhang 0045
ECAI2
2024 Saliency Map-Guided End-to-End Image Coding for Machines
abstract
Existing end-to-end image coding for machines (ICM) methods generally use joint training strategies to promote the compression efficiency for machine vision without considering the influence of different regions in the image. To encourage the image compression network to focus on the regions that are critical to the subsequent visual task, this paper proposes a saliency map-guided image compression network (SMIC-Net) for ICM. Specifically, a saliency map-guided transform module (SMTM) is proposed to improve the representation ability of image features for object detection task by exploring the semantic and structural information of the detected object. Besides, a saliency map-guided mean square error (SM-MSE) loss is designed to place more emphasis on the detected object regions. Experimental results demonstrate that the proposed SMIC-Net effectively promotes the compression efficiency for machine vision.
Bo Peng 0007, Tianxiang Lin, Dengchao Jin, Zhaoqing Pan, Jianjun Lei 0001
IEEE Signal Process. Lett.1
2024 Unsupervised Single-View Synthesis Network via Style Guidance and Prior Distillation
abstract
View synthesis aims to learn a view transformation and synthesize the target views from a single or multiple source views. Although previous view synthesis methods have obtained promising performance, they heavily rely on the supervision of the target view. In this paper, we propose an unsupervised single-view synthesis network (USVS-Net) to learn the view transformation without the supervision of the target view. Specifically, with the usage of only a single source view, a style-guidance view synthesis model is proposed to learn an intrinsic representation, which intends to describe the object from a reference pose. With the intrinsic representation, the view transformation is learned to boost the learning of the unsupervised single-view synthesis. Then, taking the style-guidance view synthesis model as the teacher, a prior-distillation view synthesis model is further presented as the student to learn a more direct view transformation. By utilizing the proposed method, high-quality target views are synthesized in a time-efficient manner. Experiments on both synthetic and real-scene datasets show that despite the lack of supervision of the target view, the proposed method achieves promising results compared with the existing view synthesis methods.
Bingzheng Liu, Bo Peng 0007, Zhe Zhang 0041, Qingming Huang, Nam Ling, Jianjun Lei 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 PIPC-3Ddet: Harnessing Perspective Information and Proposal Correlation for 3D Point Cloud Object Detection
abstract
As a fundamental technology in autonomous driving and robotic sensing system, 3D point cloud object detection has received increasing attention. In this paper, a novel 3D detection method that harnesses perspective information and proposal correlation (PIPC-3Ddet) is proposed for detecting 3D objects from point clouds. Specifically, a perspective information embedding module is designed to enhance the voxel features by capturing and embedding the perspective information of range images, so as to effectively distinguish the objects and backgrounds. Besides, by revealing the correlation among 3D proposals, a proposal correlation reasoning module is presented to learn high-quality proposal features for better 3D proposal refinement. With the designed perspective information embedding and proposal correlation reasoning modules, the proposed PIPC-3Ddet is able to better perceive the objects in the 3D scene, thus boosting the 3D object detection performance. Extensive experiments on the KITTI and Waymo benchmarks have demonstrated the superiority of the proposed PIPC-3Ddet.
Chuanbo Yu, Bo Peng 0007, Qingming Huang, Jianjun Lei 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 SWGNet: Step-Wise Reference Frame Generation Network for Multiview Video Coding
abstract
In multiview video coding, the coding performance highly depends on the quality of the reference frames. In view of this, a step-wise reference frame generation network (SWGNet) is designed to improve the quality of the reference frame for efficient multiview video coding. In particular, a frame-level to block-level learning paradigm is proposed to step-wisely generate a high-quality reference frame. In the frame-level stage, by exploiting parallax correlations between temporal and inter-view references on the basis of image alignment, a parallax-guided frame-level synthesis module is proposed to generate an elementary reference frame. Then, in the block-level stage, a transformer-based block-level aggregation module is designed to further refine the texture details of the reference frame by modeling long-range dependencies among pixels. The proposed SWGNet is integrated into 3D-HEVC, and extensive experiments demonstrate that the proposed method achieves significant bitrate saving compared with 3D-HEVC.
Jing Zhang 0017, Yonghong Hou, Zhaoqing Pan, Bo Peng 0007, Nam Ling, Jianjun Lei 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 Self-Constructing Stereo Correspondences for Unsupervised Multi-View Stereo
abstract
Existing unsupervised Multi-View Stereo (MVS) methods generally construct supervision on the basis of the photometric consistency loss, which suffers from unreliable supervision and limited scalability. In this paper, a novel unsupervised MVS framework with Self-constructed Stereo Correspondences, termed SSC-MVS, is proposed to provide reliable supervision for the network and improve scalability of unsupervised MVS. Specifically, a pseudo depth-based learning strategy is first presented to supervise the MVS network with a pseudo depth, which is used to characterize the accurate stereo correspondences. Additionally, a consistency-based training mechanism is designed, where the depth consistency between two differently-augmented inputs is constrained to further improve the robustness of the network in real MVS scenes. Experimental results on widely-used MVS datasets demonstrate that the proposed SSC-MVS obtains the state-of-the-art performance among the unsupervised methods and has the potential to outperform the fully-supervised methods. The code is available athttps://github.com/jzhu98/ssc-mvs.
Bo Peng 0007, Bingzheng Liu, Qingming Huang, Jianjun Lei 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 λ-Domain Rate Control via Wavelet-Based Residual Neural Network for VVC HDR Intra Coding
abstract
High dynamic range (HDR) video offers a more realistic visual experience than standard dynamic range (SDR) video, while introducing new challenges to both compression and transmission. Rate control is an effective technology to overcome these challenges, and ensure optimal HDR video delivery. However, the rate control algorithm in the latest video coding standard, versatile video coding (VVC), is tailored to SDR videos, and does not produce well coding results when encoding HDR videos. To address this problem, a data-driven λ -domain rate control algorithm is proposed for VVC HDR intra frames in this paper. First, the coding characteristics of HDR intra coding are analyzed, and a piecewise R- λ model is proposed to accurately determine the correlation between the rate (R) and the Lagrange parameter λ for HDR intra frames. Then, to optimize bit allocation at the coding tree unit (CTU)-level, a wavelet-based residual neural network (WRNN) is developed to accurately predict the parameters of the piecewise R- λ model for each CTU. Third, a large-scale HDR dataset is established for training WRNN, which facilitates the applications of deep learning in HDR intra coding. Extensive experimental results show that our proposed HDR intra frame rate control algorithm achieves superior coding results than the state-of-the-art algorithms. The source code of this work will be released at https://github.com/TJU-Videocoding/WRNN.git.
Jianjun Lei 0001, Zhaoqing Pan, Bo Peng 0007, Haoran Xie 0001
IEEE Trans. Image Process.4
2024 Contrastive Multi-View Learning for 3D Shape Clustering
abstract
Unsupervised 3D shape clustering is emerging as a promising research topic in multimedia and computer vision field. Considering the flexibility of acquiring multiple views for 3D shapes, this paper proposes a contrastive multi-view learning network (CMVL-Net) to cluster unlabeled 3D shapes from multiple views. To the best of our knowledge, this is the first multi-view-oriented 3D shape deep clustering method. The key to this method lies in how to capture highly discriminative 3D shape features suitable for clustering. By exploring consistency and complementarity among multiple views, a cross-view contrastive clustering mechanism is proposed to learn clustering-specified discriminative 3D shape features. To obtain a more compact 3D shape clustering structure, a consensus graph-guided contrastive constraint is designed to encourage cluster-wise consistency learning under the guidance of potential category associations among shapes. Experimental results on two widely used benchmark datasets demonstrate the effectiveness of the proposed method.
Bo Peng 0007, Guoting Lin, Jianjun Lei 0001, Tianyi Qin, Xiaochun Cao, Nam Ling
IEEE Trans. Multim.1
2024 Self-Supervised Monocular Depth Estimation via Binocular Geometric Correlation Learning
abstract
Monocular depth estimation aims to infer a depth map from a single image. Although supervised learning-based methods have achieved remarkable performance, they generally rely on a large amount of labor-intensively annotated data. Self-supervised methods, on the other hand, do not require any annotation of ground-truth depth and have recently attracted increasing attention. In this work, we propose a self-supervised monocular depth estimation network via binocular geometric correlation learning. Specifically, considering the inter-view geometric correlation, a binocular cue prediction module is presented to generate the auxiliary vision cue for the self-supervised learning of monocular depth estimation. Then, to deal with the occlusion in depth estimation, an occlusion interference attenuated constraint is developed to guide the supervision of the network by inferring the occlusion region and producing paired occlusion masks. Experimental results on two popular benchmark datasets have demonstrated that the proposed network obtains competitive results compared to state-of-the-art self-supervised methods and achieves comparable results to some popular supervised methods.
Bo Peng 0007, Jianjun Lei 0001, Bingzheng Liu, Haifeng Shen, Wanqing Li 0001, Qingming Huang
ACM Trans. Multim. Comput. Commun. Appl.1
2023 Local to non-local: Multi-scale progressive attention network for image restoration
Lili Shen, Qunxia Li, Chuhe Zhang, Xichun Sun, Bo Peng 0007
Comput. Vis. Image Underst.6
2023 Deep In-Loop Filtering via Multi-Domain Correlation Learning and Partition Constraint for Multiview Video Coding
abstract
The deep learning-based in-loop filtering methods have greatly improved the coding efficiency for High Efficiency Video Coding (HEVC). However, directly applying these HEVC-orientated in-loop filtering methods to multiview video coding may not obtain satisfactory performance due to the characteristics of multiview video. In this paper, a deep in-loop filtering method based on multi-domain correlation learning and partition constraint network (MDP-Net) is proposed to boost the multiview video coding performance. To the best of our knowledge, this work is the first attempt at deep in-loop filtering for multiview video coding. Specifically, a multi-domain correlation learning module is presented to restore the high-frequency details of the distorted frame by exploring the multi-domain correlations. Besides, based on the block partition information generated in video coding, a partition-constrained reconstruction module is proposed to better attenuate the compression artifacts by designing a partition loss. Finally, the proposed MDP-Net is integrated into 3D-HEVC reference software, and the experimental results demonstrate that the proposed method achieves considerable performance improvement compared with 3D-HEVC.
Bo Peng 0007, Renjie Chang, Zhaoqing Pan, Ge Li 0006, Nam Ling, Jianjun Lei 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 RGB-D Human Matting: A Real-World Benchmark Dataset and a Baseline Method
abstract
The last decade has witnessed an increasing exploration and development of human matting. However, existing matting works primarily focus on predicting better alpha mattes from RGB images. So far few efforts have been devoted to tackling human matting in real-world activity scenarios with RGB-D information. To this end, this paper concentrates on the RGB-D human matting task, and provides the first public RGB-D human matting benchmark dataset as well as a baseline method for deep learning-based RGB-D human matting. To support the research on RGB-D human matting, a new RGB-D human-matting dataset (HDM-2K) is collected and released, which contains 2,270 high-resolution human images in various real-world scenarios and the corresponding depth maps. Additionally, a baseline method for RGB-D human matting is further proposed, which automatically generates the alpha matte by jointly exploiting the spatial structure information in the depth map and detailed texture information in the RGB image. Finally, extensive experiments conducted on the HDM-2K dataset demonstrate that the depth maps are effective for the matting task and the proposed baseline method achieves promising performance on human matting.
Bo Peng 0007, Jianjun Lei 0001, Huazhu Fu, Haifeng Shen, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.1
2023 Recurrent Interaction Network for Stereoscopic Image Super-Resolution
abstract
Recently, deep learning-based stereoscopic image super-resolution has attracted extensive attention and made great progress. However, existing methods have not adequately explored the inter-view dependency among two-view multi-level features. In this paper, a recurrent interaction network for stereoscopic image super-resolution (RISSRnet) is proposed to learn the inter-view dependency. To efficiently utilize the relationship between the two views, a recurrent interaction module is designed to achieve recurrent interaction among two-view multi-level features from the regrouped sequences, which are generated by a coupled queue-regroup mechanism. In addition, to recursively enhance features in the recurrent interaction module, an iterative propagation strategy is developed for sufficient interaction. Extensive experimental results demonstrate the effectiveness and superiority of the proposed RISSRnet.
Zhe Zhang 0041, Bo Peng 0007, Jianjun Lei 0001, Haifeng Shen, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.2
2023 Reducing Background Induced Domain Shift for Adaptive Person Re-Identification
abstract
Cross-domain person re-identification (Re-ID) is a challenging and important task in monitoring safety and procedure compliance of industrial work places. In this article, a novel method is proposed to reduce background induced domain shift for adaptive person Re-ID. Specifically, a foreground-background joint clustering module is proposed to extract discriminative foreground and background features and an attention-based feature disentanglement module is designed to reduce the interference of background with the extraction of discriminative foreground features. Experimental results on three widely used person Re-ID benchmarking datasets (Market-1501, DukeMTMC-reID, and MSMT17) have demonstrated that the proposed method achieves promising performance compared with the state-of-the-art methods.
Jianjun Lei 0001, Tianyi Qin, Bo Peng 0007, Wanqing Li 0001, Zhaoqing Pan, Haifeng Shen, Sam Kwong
IEEE Trans. Ind. Informatics3
2023 ZS-SBPRnet: A Zero-Shot Sketch-Based Point Cloud Retrieval Network Based on Feature Projection and Cross-Reconstruction
abstract
With the widespread deployment of 3D sensors, point cloud analysis has become an important topic in the field of industrial information. This article proposes a novel zero-shot sketch-based point cloud retrieval network based on feature projection and cross reconstruction, termed as ZS-SBPRnet. As far as we know, the proposed ZS-SBPRnet is the first attempt at retrieving point clouds based on sketches under the zero-shot scenario. To tackle the problem of the cross-modal differences, a structure-preserving learnable feature projection module is designed to obtain view feature representations from point cloud features containing spatial structure information through feature projection. Besides, to achieve efficient cross-modal feature alignment under the zero-shot scenario, a sketch-point cloud cross-reconstruction mechanism is presented to promote cross-modal feature alignment between sketches and point clouds in visual space. Experimental results on the benchmark datasets validate the superiority of the proposed ZS-SBPRnet.
Bo Peng 0007, Haifeng Shen, Qingming Huang, Jianjun Lei 0001
IEEE Trans. Ind. Informatics1
2023 Learned Video Compression With Efficient Temporal Context Learning
abstract
In contrast to image compression, the key of video compression is to efficiently exploit the temporal context for reducing the inter-frame redundancy. Existing learned video compression methods generally rely on utilizing short-term temporal correlations or image-oriented codecs, which prevents further improvement of the coding performance. This paper proposed a novel temporal context-based video compression network (TCVC-Net) for improving the performance of learned video compression. Specifically, a global temporal reference aggregation (GTRA) module is proposed to obtain an accurate temporal reference for motion-compensated prediction by aggregating long-term temporal context. Furthermore, in order to efficiently compress the motion vector and residue, a temporal conditional codec (TCC) is proposed to preserve structural and detailed information by exploiting the multi-frequency components in temporal context. Experimental results show that the proposed TCVC-Net outperforms public state-of-the-art methods in terms of both PSNR and MS-SSIM metrics.
Dengchao Jin, Jianjun Lei 0001, Bo Peng 0007, Zhaoqing Pan, Li Li 0040, Nam Ling
IEEE Trans. Image Process.3
2023 Novel View Synthesis from a Single Unposed Image via Unsupervised Learning
abstract
Novel view synthesis aims to generate novel views from one or more given source views. Although existing methods have achieved promising performance, they usually require paired views with different poses to learn a pixel transformation. This article proposes an unsupervised network to learn such a pixel transformation from a single source image. In particular, the network consists of a token transformation module that facilities the transformation of the features extracted from a source image into an intrinsic representation with respect to a pre-defined reference pose and a view generation module that synthesizes an arbitrary view from the representation. The learned transformation allows us to synthesize a novel view from any single source image of an unknown pose. Experiments on the widely used view synthesis datasets have demonstrated that the proposed network is able to produce comparable results to the state-of-the-art methods despite the fact that learning is unsupervised and only a single source image is required for generating a novel view. The code will be available upon the acceptance of the article.
Bingzheng Liu, Jianjun Lei 0001, Bo Peng 0007, Chuanbo Yu, Wanqing Li 0001, Nam Ling
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Modeling Long-range Dependencies and Epipolar Geometry for Multi-view Stereo
abstract
This article proposes a network, referred to as Multi-View Stereo TRansformer (MVSTR) for depth estimation from multi-view images. By modeling long-range dependencies and epipolar geometry, the proposed MVSTR is capable of extracting dense features with global context and 3D consistency, which are crucial for reliable matching in multi-view stereo (MVS). Specifically, to tackle the problem of the limited receptive field of existing CNN-based MVS methods, a global-context Transformer module is designed to establish intra-view long-range dependencies so that global contextual features of each view are obtained. In addition, to further enable features of each view to be 3D consistent, a 3D-consistency Transformer module with an epipolar feature sampler is built, where epipolar geometry is modeled to effectively facilitate cross-view interaction. Experimental results show that the proposed MVSTR achieves the best overall performance on the DTU dataset and demonstrates strong generalization on the Tanks & Temples benchmark dataset.
Bo Peng 0007, Wanqing Li 0001, Haifeng Shen, Qingming Huang, Jianjun Lei 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2022 Deep Stereo Image Compression via Bi-directional Coding
abstract
Existing learning-based stereo compression methods usually adopt a unidirectional approach to encoding one image independently and the other image conditioned upon the first. This paper proposes a novel bidirectional coding-based end-to-end stereo image compression network (BCSIC-Net). BCSIC-Net consists of a novel bidirectional contextual transform module which performs nonlinear transform conditioned upon the inter-view context in a latent space to reduce inter-view redundancy, and a bidirectional conditional entropy model that employs interview correspondence as a conditional prior to improve coding efficiency. Experimental results on the InStereo2K and KITTI datasets demonstrate that the proposed BCSIC-Net can effectively reduce the inter-view redundancy and out-performs state-of-the-art methods.
Jianjun Lei 0001, Xiangrui Liu, Bo Peng 0007, Dengchao Jin, Wanqing Li 0001, Jingxiao Gu
CVPR3
2022 Texture-Guided End-to-End Depth Map Compression
abstract
End-to-end compression methods designed for the texture image have achieved excellent coding performances. Due to the characteristic differences between the depth map and the texture image, the texture-oriented methods have limitations in depth map compression. To address this problem, this paper proposes a texture-guided end-to-end depth map compression network (TDMC-Net). Specifically, the proposed TDMC-Net is mainly composed of the texture-guided transform module (TTM) which performs the nonlinear transform with providing the textual context to reduce the redundancy in depth feature, and a texture-guided conditional entropy model (TCEM) which is designed to improve the entropy model by introducing the texture conditional prior. Experimental results show that the proposed TDMC-Net boosts the depth coding efficiency by utilizing the texture information and achieves superior performance.
Bo Peng 0007, Yuying Jing, Dengchao Jin, Xiangrui Liu, Zhaoqing Pan, Jianjun Lei 0001
ICIP1
2022 Entity Slot Filling for Visual Captioning
abstract
To explore the specific visual aspects and the language consistency at the same time, this paper introduces a new image captioning task, dubbedentity slot filling captioning (ESFCap). It is similar to the masked entity completion tasks in NLP, which are widely used to study language context and has been successfully employed to improve language understanding. Specifically, given a sentence with blank for describing an image, the ESFCap task aims to fill the blank with proper text content according to the visual information. The filled text should be grounded to correct visual entities and also in concordance with the sentence structure. To support the ESFCap research, we collect and release an entity slot filling captioning dataset,Flickr30k-EnFi, based on Flickr30k-Entities. The Flickr30k-EnFi dataset consists of 31,783 images and 565,750 masked sentences, as well as the text snippets for the masked slot. For tackling the ESFCap task, we propose a multi-modal fusion model equipped with a novel adaptive dynamic attention module, termed AdaMFN. The AdaMFN model effectively leverages both global and local information from vision and language. It is also able to adaptively focus on the key linguistic knowledge and visual regions to generate correct filling results. The experimental results and analysis demonstrate the effectiveness of our proposed model.
Yi Bin, Yujuan Ding, Bo Peng 0007, Yang Yang 0002, Tat-Seng Chua
IEEE Trans. Circuits Syst. Video Technol.3
2022 Deep Affine Motion Compensation Network for Inter Prediction in VVC
abstract
In video coding, it is a challenge to deal with scenes with complex motions, such as rotation and zooming. Although affine motion compensation (AMC) is employed in Versatile Video Coding (VVC), it is still difficult to handle non-translational motions due to the adopted hand-craft block-based motion compensation. In this paper, we propose a deep affine motion compensation network (DAMC-Net) for inter prediction in video coding to effectively improve the prediction accuracy. To the best of our knowledge, our work is the first attempt to deal with the deformable motion compensation based on CNN in VVC. Specifically, a deformable motion-compensated prediction (DMCP) module is proposed to compensate the current encoding block through a learnable way to estimate accurate motion fields. Meanwhile, the spatial neighboring information and the temporal reference block as well as the initial motion field are fully exploited. By effectively fusing the multi-channel feature maps from DMCP, an attention-based fusion and reconstruction (AFR) module is designed to reconstruct the output block. The proposed DAMC-Net is integrated into VVC and the experimental results demonstrate that the proposed method considerably enhances the coding performance.
Dengchao Jin, Jianjun Lei 0001, Bo Peng 0007, Wanqing Li 0001, Nam Ling, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.3
2022 Multiple Resolution Prediction With Deep Up-Sampling for Depth Video Coding
abstract
The depth video contains large smooth contents with sharp edges. Since the deep learning-based color video orientated intra prediction methods pay no attention to the characteristics of depth video, they are unsuitable for optimizing the coding efficiency of depth video. In this paper, a multiple resolution prediction method with deep up-sampling is proposed to promote the coding efficiency of depth video. To efficiently encode the depth blocks of different complexity, the depth block is selectively encoded at different resolutions, including$\times 1$,$\times 1$/2, and$\times 1$/4 resolutions. If the block is encoded with a low-resolution (LR), the resolution of reconstructed LR depth block is recovered by an up-sampling network. To constrain the quality of both reconstructed high-resolution depth block and its synthesized view, a view synthesis distortion guidance mechanism is proposed for the up-sampling network. In addition, a distillation-based lightweight up-sampling network is proposed to reduce the computational complexity. Experimental results demonstrate that the proposed multiple resolution prediction method obtains an average of 10.84% BD-rate saving in comparison with 3D-HEVC.
Ge Li 0006, Jianjun Lei 0001, Zhaoqing Pan, Bo Peng 0007, Nam Ling
IEEE Trans. Circuits Syst. Video Technol.4
2022 LVE-S2D: Low-Light Video Enhancement From Static to Dynamic
abstract
Recently, deep-learning-based low-light video enhancement methods have drawn wide attention and achieved remarkable performance. However, limited by the difficulty in collecting dynamic low-light and well-lighted video pairs in real scenes, how to construct video sequences for supervised learning and design a low-light enhancement network for real dynamic video remains a challenge. In this paper, we propose a simple yet effective low-light video enhancement method (LVE-S2D), which generates dynamic video training pairs from static videos, and enhances the low-light video by mining dynamic temporal information. To obtain low-light and well-lighted video pairs, a sliding window-based dynamic video generation mechanism is designed to produce pseudo videos with rich dynamic temporal information. Then, a siamese dynamic low-light video enhancement network is presented, which effectively utilizes temporal correlation between adjacent frames to enhance the video frames. Extensive experimental results demonstrate that the proposed method not only achieves superior performance on static low-light videos, but also outperforms the state-of-the-art methods on real dynamic low-light videos.
Bo Peng 0007, Jianjun Lei 0001, Zhe Zhang 0041, Nam Ling, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.1
2022 SIEV-Net: A Structure-Information Enhanced Voxel Network for 3D Object Detection From LiDAR Point Clouds
abstract
As one of the fundamental tasks in scene understanding, 3D object detection from LiDAR point clouds has drawn extensive attention in the past few years. Although the existing voxel-based methods have achieved remarkable performance, how to effectively exploit geometric structure information of the point clouds to boost the detection performance remains to be explored. In this paper, we propose a novel structure-information enhanced voxel network (SIEV-Net) for 3D object detection from LiDAR point clouds. The proposed SIEV-Net learns feature representations of 3D objects by jointly considering uneven spatial distribution and height information of the point clouds. Specifically, considering the uneven spatial distribution characteristics of point clouds, a hierarchical-voxel feature encoding module is proposed to effectively extract features of voxels in both sparse and dense regions. Besides, by utilizing the Bird’s Eye View (BEV) map of point clouds, a height information complement module is designed to minimize the height information lost in the process of point feature aggregation in a voxel network. Experimental results on the widely used KITTI benchmark dataset have demonstrated the efficacy of the proposed SIEV-Net.
Chuanbo Yu, Jianjun Lei 0001, Bo Peng 0007, Haifeng Shen, Qingming Huang
IEEE Trans. Geosci. Remote. Sens.3
2022 C2FNet: A Coarse-to-Fine Network for Multi-View 3D Point Cloud Generation
abstract
Generation of a 3D model of an object from multiple views has a wide range of applications. Different parts of an object would be accurately captured by a particular view or a subset of views in the case of multiple views. In this paper, a novel coarse-to-fine network (C2FNet) is proposed for 3D point cloud generation from multiple views. C2FNet generates subsets of 3D points that are best captured by individual views with the support of other views in a coarse-to-fine way, and then fuses these subsets of 3D points to a whole point cloud. It consists of a coarse generation module where coarse point clouds are constructed from multiple views by exploring the cross-view spatial relations, and a fine generation module where the coarse point cloud features are refined under the guidance of global consistency in appearance and context. Extensive experiments on the benchmark datasets have demonstrated that the proposed method outperforms the state-of-the-art methods.
Jianjun Lei 0001, Bo Peng 0007, Wanqing Li 0001, Zhaoqing Pan, Qingming Huang
IEEE Trans. Image Process.3
2022 Multi-Modality MR Image Synthesis via Confidence-Guided Aggregation and Cross-Modality Refinement
abstract
Magnetic resonance imaging (MRI) can provide multi-modality MR images by setting task-specific scan parameters, and has been widely used in various disease diagnosis and planned treatments. However, in practical clinical applications, it is often difficult to obtain multi-modality MR images simultaneously due to patient discomfort, and scanning costs, etc. Therefore, how to effectively utilize the existing modality images to synthesize missing modality image has become a hot research topic. In this paper, we propose a novel confidence-guided aggregation and cross-modality refinement network (CACR-Net) for multi-modality MR image synthesis, which effectively utilizes complementary and correlative information of multiple modalities to synthesize high-quality target-modality images. Specifically, to effectively utilize the complementary modality-specific characteristics, a confidence-guided aggregation module is proposed to adaptively aggregate the multiple target-modality images generated from multiple source-modality images by using the corresponding confidence maps. Based on the aggregated target-modality image, a cross-modality refinement module is presented to further refine the target-modality image by mining correlative information among the multiple source-modality images and aggregated target-modality image. By training the proposed CACR-Net in an end-to-end manner, high-quality and sharp target-modality MR images are effectively synthesized. Experimental results on the widely used benchmark demonstrate that the proposed method outperforms state-of-the-art methods.
Bo Peng 0007, Bingzheng Liu, Yi Bin, Lili Shen, Jianjun Lei 0001
IEEE J. Biomed. Health Informatics1
2021 Depth-Assisted Joint Detection Network For Monocular 3d Object Detection
abstract
In the past few years, monocular 3D object detection has attracted increasing attention due to the merit of low cost and wide range of applications. In this paper, a depth-assisted joint detection network (MonoDAJD) is proposed for monocular 3D object detection. Specifically, a consistency-aware joint detection mechanism is proposed to jointly detect objects in the image and depth map, and exploit the localization information from the depth detection stream to optimize the detection results. To obtain more accurate 3D bounding boxes, an orientation-embedded NMS is designed by introducing the orientation confidence prediction and embedding the orientation confidence into the traditional NMS. Experimental results on the widely used KITTI benchmark demonstrate that the proposed method achieves promising performance compared with the state-of-the-art monocular 3D object detection methods.
Jianjun Lei 0001, Tingyi Guo, Bo Peng 0007, Chuanbo Yu
ICIP3
2021 Multi-Perspective Video Captioning
abstract
This work targets at the problems of comprehensive video captioning and the generation of multiple descriptions from different perspectives, termed asMulti-Perspective Video Captioning. We build and release a dataset named VidOR-MPVC, the first dataset for multi-perspective video captioning, where each video is annotated with multiple descriptions from different perspectives. We also propose a novel model, dubbedperspective-aware captioner (PAC), which is capable of mining the various perspectives in a video and generating a description from each perspective. More specifically, a perspective generator is designed to perceive video content with perspective preferences, and followed by a language generator equipped with perspective-aware attention mechanism. As our new task expects to produce multiple descriptions for a video, existing evaluation metrics are fail to handle this situation. To address this problem, we devise the maximum matching scores based on existing metrics for an overall evaluation which aims to cover the aspects of semantic similarity, completeness and compactness. The experimental results demonstrate that our model is able to describe videos with multiple descriptions from different perspectives.
Yi Bin, Xindi Shang, Bo Peng 0007, Yujuan Ding, Tat-Seng Chua
ACM Multimedia3
2021 Deep video action clustering via spatio-temporal feature learning
Bo Peng 0007, Jianjun Lei 0001, Huazhu Fu, Yalong Jia, Zongqian Zhang
Neurocomputing1
2021 A CNN-Based Fast Inter Coding Method for VVC
abstract
The Versatile Video Coding (VVC) achieves superior coding efficiency as compared with the High Efficiency Video Coding (HEVC), while its excellent coding performance is at the cost of several high computational complexity coding tools, such as Quad-Tree plus Multi-type Tree (QTMT)-based Coding Units (CUs) and multiple inter prediction modes. To reduce the computational complexity of VVC, a CNN-based fast inter coding method is proposed in this paper. First, a multi-information fusion CNN (MF-CNN) model is proposed to early terminate the QTMT-based CU partition process by jointly using the multi-domain information. Then, a content complexity-based early Merge mode decision is proposed to skip the time-consuming inter prediction modes by considering the CU prediction residuals and the confidence of MF-CNN. Experimental results show that the proposed method reduces an average of 30.63% VVC encoding time, and the Bjøontegaard Delta Bit Rate (BDBR) increases about 3%.
Zhaoqing Pan, Peihan Zhang, Bo Peng 0007, Nam Ling, Jianjun Lei 0001
IEEE Signal Process. Lett.3
2021 Deep Spatial-Spectral Subspace Clustering for Hyperspectral Image
abstract
Hyperspectral image (HSI) clustering is a challenging task due to the complex characteristics in HSI data, such as spatial-spectral structure, high-dimension, and large spectral variability. In this paper, we propose a novel deep spatial-spectral subspace clustering network (DS3C-Net), which explores spatial-spectral information via the multi-scale auto-encoder and collaborative constraint. Considering the structure correlations of HSI, the multi-scale auto-encoder is first designed to extract spatial-spectral features with different-scale pixel blocks which are selected as the inputs. Then, the collaborative constrained self-expressive layers are introduced between the encoder and decoder, to capture the self-expressive subspace structures. By designing a self-expressiveness similarity constraint, the proposed network is trained collaboratively, and the affinity matrices of the feature representation are learned in an end-to-end manner. Based on the affinity matrices, the spectral clustering algorithm is utilized to obtain the final HSI clustering result. Experimental results on three widely used hyperspectral image datasets demonstrate that the proposed method outperforms state-of-the-art methods.
Jianjun Lei 0001, Bo Peng 0007, Leyuan Fang, Nam Ling, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.3
2020 Attention-Guided Fusion Network of Point Cloud and Multiple Views for 3D Shape Recognition
abstract
With the dramatic growth of 3D shape data, 3D shape recognition has become a hot research topic in the field of computer vision. How to effectively utilize the multimodal characteristics of 3D shape has been one of the key problems to boost the performance of 3D shape recognition. In this paper, we propose a novel attention-guided fusion network of point cloud and multiple views for 3D shape recognition. Specifically, in order to obtain more discriminative descriptor for 3D shape data, the inter-modality attention enhancement module and view-context attention fusion module are proposed to gradually refine and fuse the features of the point cloud and multiple views. In the inter-modality attention enhancement module, the inter-modality attention mask based on the joint feature representation is computed, so that the features of each modality are enhanced by fusing the correlative information between two modalities. After that, the view-context attention fusion module is proposed to explore the context information of multiple views, and fuse the enhanced features to obtain more discriminative descriptor for 3D shape data. Experimental results on the ModelNet40 dataset demonstrate that the proposed method achieves promising performance compared with state-of-the-art methods.
Bo Peng 0007, Zengrui Yu, Jianjun Lei 0001
VCIP1
2020 Semi-Heterogeneous Three-Way Joint Embedding Network for Sketch-Based Image Retrieval
abstract
Sketch-based image retrieval (SBIR) is a challenging task due to the large cross-domain gap between sketches and natural images. How to align abstract sketches and natural images into a common high-level semantic space remains a key problem in SBIR. In this paper, we propose a novel semi-heterogeneous three-way joint embedding network (Semi3-Net), which integrates three branches (a sketch branch, a natural image branch, and an edgemap branch) to learn more discriminative cross-domain feature representations for the SBIR task. The key insight lies with how we cultivate the mutual and subtle relationships amongst the sketches, natural images, and edgemaps. A semi-heterogeneous feature mapping is designed to extract bottom features from each domain, where the sketch and edgemap branches are shared while the natural image branch is heterogeneous to the other branches. In addition, a joint semantic embedding is introduced to embed the features from different domains into a common high-level semantic space, where all of the three branches are shared. To further capture informative features common to both natural images and the corresponding edgemaps, a co-attention model is introduced to conduct common channel-wise feature recalibration between different domains. A hybrid-loss mechanism is designed to align the three branches, where an alignment loss and a sketch-edgemap contrastive loss are presented to encourage the network to learn invariant cross-domain representations. Experimental results on two widely used category-level datasets (Sketchy and TU-Berlin Extension) demonstrate that the proposed method outperforms state-of-the-art methods.
Jianjun Lei 0001, Bo Peng 0007, Zhanyu Ma, Ling Shao 0001, Yi-Zhe Song
IEEE Trans. Circuits Syst. Video Technol.3
2020 Unsupervised Video Action Clustering via Motion-Scene Interaction Constraint
abstract
In the past few years, scene contextual information has been increasingly used for action understanding with promising results. However, unsupervised video action clustering using context has been less explored, and existing clustering methods cannot achieve satisfactory performances. In this paper, we propose a novel unsupervised video action clustering method by using the motion-scene interaction constraint (MSIC). The proposed method takes the unique static scene and dynamic motion characteristics of video action into account, and develops a contextual interaction constraint model under a self-representation subspace clustering framework. First, the complementarity of multi-view subspace representation in each context is explored by single-view and multi-view constraints. Afterward, the context-constrained affinity matrix is calculated and the MSIC is introduced to mutually regularize the disagreement of subspace representation in scene and motion. Finally, by jointly constraining the complementarity of multi-views and the consistency of multi-contexts, an overall objective function is constructed to guarantee the video action clustering result. The experiments on four video benchmark datasets (Weizmann, KTH, UCFsports, and Olympic) demonstrate that the proposed method outperforms the state-of-the-art methods.
Bo Peng 0007, Jianjun Lei 0001, Huazhu Fu, Changqing Zhang 0002, Tat-Seng Chua, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.1
2020 A Recursive Constrained Framework for Unsupervised Video Action Clustering
abstract
Video action understanding is an active field of intelligent video analytics, and contextual information in the videos has gained lots of attention for better action understanding. However, most existing works focus on using contextual information for supervised or semi-supervised analysis, and how to effectively use contextual information to boost the unsupervised action clustering performance is still a challenging problem. In this article, we propose a recursive constrained framework for unsupervised video action clustering by utilizing the contextual information of the action and scene. Considering the unique contextual characteristics of video action, action context clustering solution and scene context clustering solution are obtained simultaneously. Based on these two solutions, a recursive priori propagation is proposed to exploit information gain of the priori clustering solutions, and then the information gain is fed back into the procedures of both subspace representation and spectral clustering. Specifically, to explore the unknown relationships in the priori clustering solutions, the constraint-guided subspace representation is introduced by fusing the recursive priori constraint into the self-representation model. Taking priori information and multiview features into consideration, the priori-inherited multiview spectral clustering is proposed to obtain more discriminative spectral embeddings for action clustering. Experiments on three video benchmark datasets demonstrate that the proposed method outperforms state-of-the-art methods.
Bo Peng 0007, Jianjun Lei 0001, Huazhu Fu, Ling Shao 0001, Qingming Huang
IEEE Trans. Ind. Informatics1
2019 Channel-wise Temporal Attention Network for Video Action Recognition
abstract
Recently, video action recognition receives lots of attention, and deep learning based methods have achieved promising performance. Most existing methods focus on spatiotemporal information encoding to learn video representation, which ignore the relevance among channels. In this paper, we propose a novel Channel-wise Temporal Attention Network (CTAN) to explore the fine-grained key information for action recognition. First, the channel-wise attention generation module is proposed to emphasize the fine-grained informative features in each frame. Then, the temporal information aggregation module is introduced before attention generation to exploit the interaction of different frames. Finally, a discriminative video-level representation for action recognition is generated by end-to-end training. Experimental results on two benchmarks, UCF101 and HMDB51, demonstrate the effectiveness of the proposed CTAN.
Jianjun Lei 0001, Yalong Jia, Bo Peng 0007, Qingming Huang
ICME3
2019 Person Re-Identification by Semantic Region Representation and Topology Constraint
abstract
Person re-identification is a popular research topic which aims at matching the specific person in a multi-camera network automatically. Feature representation and metric learning are two important issues for person re-identification. In this paper, we propose a novel person re-identification method, which consists of a reliable representation called semantic region representation (SRR), and an effective metric learning with mapping space topology constraint (MSTC). The SRR integrates semantic representations to achieve effective similarity comparison between the corresponding regions via parsing the body into multiple parts, which focuses on the foreground context against the background interference. To learn a discriminant metric, the MSTC is proposed to consider the topological relationship among all samples in the feature space. It considers two-fold constraints: the distribution of positive pairs should be more compact than the average distribution of negative pairs with regard to the same probe, while the average distance between different classes should be larger than that between same classes. These two aspects cooperate to maintain the compactness of the intra-class as well as the sparsity of the inter-class. Extensive experiments conducted on five challenging person re-identification datasets, VIPeR, SYSU-sReID, QUML GRID, CUHK03, and Market-1501, show that the proposed method achieves competitive performance with the state-of-the-art approaches.
Jianjun Lei 0001, Lijie Niu, Huazhu Fu, Bo Peng 0007, Qingming Huang, Chunping Hou
IEEE Trans. Circuits Syst. Video Technol.4
2018 Shape-Preserving Object Depth Control for Stereoscopic Images
abstract
In the field of 3-D technology, it is interesting as well as meaningful issue to control object depth in 3-D space. Recently, some depth control methods for stereoscopic images have been proposed, which usually employ depth map or directly process color images to implement depth control. There are two main disadvantages for these methods. First, the results of these methods usually suffer from object deformation and holes. Second, these methods are prone to cause undesired object size changing in 3-D space. To address these issues, we propose a shape-preserving object depth control method for stereoscopic images. First, a novel depth mapping model is presented for calculating the ideal coordinates of the key points in depth control, so that the shape of the object can be well preserved. Afterward, the image content-based constraints are used to further preserve the structure of the object and its background. Finally, the warping technology is introduced to deal with images optimally as well as to avoid holes. Experimental results show that the proposed method can control object depth and preserve the shape of the object effectively without sensible background distortion.
Jianjun Lei 0001, Bo Peng 0007, Changqing Zhang 0002, Xuguang Mei, Xiaochun Cao, Xiaoting Fan, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.2