Jialun Pei

dblp:209/1971 · DBLP profile ↗
← Back
27ranked-venue papers
10as first author
26since 2021 · last 2026
0000-0002-2630-2838ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Anatomy-Aware Text-Visual Fusion with Dual-Perspective Prompts for Fine-Grained Lumbar Spine Segmentation
Sheng Lian, Jianlong Cai, Dengfeng Pan, Guang-Yong Chen, Fan Zhang 0045, Jialun Pei, Shuo Li 0001
Int. J. Comput. Vis.8
2026 Depth-induced prompt learning for laparoscopic liver landmark detection
abstract
• A new liver landmark detection dataset, L3D-2K, comprising 2,000 keyframes sourced from surgical videos with professional annotations. • A novel deep learning framework D2GPLand+ that utilizes RGB-D information for laparoscopic liver landmark detections. • Proposing the DPE module, which incorporates learnable prompts with contrastive learning to discriminate the geometric features of different landmark categories from depth clues. • Introducing the CUMamba block that concurrently conducts cross-modal interactions on spatial dimension and feature reparameterization on channel dimension for effective RGB-D fusion. • Introducing the AFA scheme to highlight anatomical structures by implicit and explicit edge emphasis and controlling detail levels. Laparoscopic liver surgery presents a highly intricate intraoperative environment with significant liver deformation, posing challenges for surgeons in locating critical liver structures. Anatomical liver landmarks can greatly assist surgeons in spatial perception in laparoscopic scenarios and facilitate preoperative-to-intraoperative registration. To advance research in liver landmark detection, we develop a new dataset called L3D-2K , comprising 2,000 keyframes with expert landmark annotations from surgical videos of 47 patients. Accordingly, we propose a baseline, D 2 GPLand+, which effectively leverages depth modality to boost landmark detection performance. Concretely, we introduce a Depth-aware Prompt Embedding (DPE) scheme, which dynamically extracts class-related global geometric cues with the guidance of self-supervised prompts from the SAM encoder. Further, a Cross-dimension Unified Mamba (CUMamba) block is designed to comprehensively incorporate RGB and depth features with the concurrent spatial and channel scanning mechanism. Besides, we bring out an Anatomical Feature Augmentation (AFA) module that captures anatomical cues and emphasizes key structures by optimizing feature granularity. For benchmarking purposes, we evaluate our method and 17 mainstream detection models on L3D, L3D-2K, and P2ILF datasets. Experimental results demonstrate that D 2 GPLand+ obtains superior performance on all three datasets. Our approach provides surgeons with guiding clues that facilitate surgical operations and decision-making in complex laparoscopic surgery. Our code and dataset are available at https://github.com/cuiruize/D2GPLand-Plus .
Ruize Cui, Weixin Si, Zhixi Li, Kai Wang 0092, Jialun Pei, Pheng-Ann Heng, Harry Qin
Medical Image Anal.5
2026 LungRes80: Towards tangled surgical workflow recognition in video-assisted thoracoscopic surgery
abstract
Video-Assisted Thoracoscopic Surgery (VATS) is a minimally invasive procedure developed to remove specific lung segments for the treatment of early-stage lung diseases. The surgical procedure involves intricate vascular and bronchial anatomy to preserve as much lung tissue as possible, minimizing impact on the pulmonary function. To assist in monitoring and early warning of this high-risk surgical workflow, we build a new dataset, LungRes80, including 269,806 video frames with phase annotations sampled from 80 VATS cases. LungRes80 presents unique challenges for hierarchical temporal modeling due to diverse short-term transitions between segmentectomy phases and latent long-term causal relations. To this end, we introduce an online baseline model termed LungReco. This framework employs Masked Causal Reasoning (MCR) to perform causal reasoning with semantic modeling from continuously updated memories along with pre-trained Large Language Models (LLMs), and combines it with Concurrent Spatial-Temporal encoding (CoST) for holistic bi-modal co-spatial-temporal aggregation across short- and long-term memories. Furthermore, a new metric, called the Attentional Distraction Coefficient (ADC), is proposed to quantify the costs of intraoperative distraction and postoperative corrections by wrong predictions. We establish a comprehensive benchmark for surgical workflow recognition by evaluating representative models on LungRes80, AutoLaparo, and Cholec80, where our method consistently achieves state-of-the-art performance. Code and data are available at LungRes80.
Diandian Guo, Jialun Pei, Jiaao Li, Yanhui Wan, Hao Chen 0011, Pheng-Ann Heng
Medical Image Anal.3
2026 Spatio-Temporal Representation Decoupling and Enhancement for Federated Instrument Segmentation in Surgical Videos
abstract
Surgical instrument segmentation under Federated Learning (FL) is a promising direction, which enables multiple surgical sites to collaboratively train the model without centralizing datasets. However, there exist very limited FL works in surgical data science, and FL methods for other modalities do not consider inherent characteristics in surgical domain: i) different scenarios show diverse anatomical backgrounds while highly similar instrument representation; ii) there exist surgical simulators which promote large-scale synthetic data generation with minimal efforts. In this paper, we propose a novel Personalized FL scheme, Spatio-Temporal Representation Decoupling and Enhancement (FedST), which wisely leverages surgical domain knowledge during both local-site and global-server training to boost segmentation. Concretely, our model embraces a Representation Separation and Cooperation (RSC) mechanism in local-site training, which decouples the query embedding layer to be trained privately, to encode respective backgrounds. Meanwhile, other parameters are optimized globally to capture the consistent representations of instruments, including the temporal layer to capture similar motion patterns. A textual-guided channel selection is further designed to highlight site-specific features, facilitating model adaptation to each site. Moreover, in global-server training, we propose Synthesis-based Explicit Representation Quantification (SERQ), which defines an explicit representation target based on synthetic data to synchronize the model convergence during fusion for improving model generalization. We construct a new PFL benchmark comprising five surgical sites from public datasets covering four types, with one out-of-federation site. FedST outperforms other state-of-the-art methods on federated sites (1.84% on IoU) and achieves a remarkable improvement on the out-of-federation site (45.29% on IoU). Our source code can be made available at: https://github.com/Meaw0415/FedST.
Xiaoming Qi, Chun-Mei Feng 0001, Jialun Pei, Weixin Si, Yueming Jin
IEEE Trans. Medical Imaging4
2025 Surgical Workflow Recognition and Blocking Effectiveness Detection in Laparoscopic Liver Resection with Pringle Maneuver
abstract
Pringle maneuver (PM) in laparoscopic liver resection aims to reduce blood loss and provide a clear surgical view by intermittently blocking blood inflow of the liver, whereas prolonged PM may cause ischemic injury. To comprehensively monitor this surgical procedure and provide timely warnings of ineffective and prolonged blocking, we suggest two complementary AI-assisted surgical monitoring tasks: workflow recognition and blocking effectiveness detection in liver resections. The former presents challenges in real-time capturing of short-term PM, while the latter involves the intraoperative discrimination of long-term liver ischemia states. To address these challenges, we meticulously collect a novel dataset, called PmLR50, consisting of 25,037 video frames covering various surgical phases from 50 laparoscopic liver resection procedures. Additionally, we develop an online baseline for PmLR50, termed PmNet. This model embraces Masked Temporal Encoding (MTE) and Compressed Sequence Modeling (CSM) for efficient short-term and long-term temporal information modeling, and embeds Contrastive Prototype Separation (CPS) to enhance action discrimination between similar intraoperative operations. Experimental results demonstrate that PmNet outperforms existing state-of-the-art surgical workflow recognition methods on the PmLR50 benchmark. Our research offers potential clinical applications for the laparoscopic liver surgery community.
Diandian Guo, Weixin Si, Zhixi Li, Jialun Pei, Pheng-Ann Heng
AAAI4
2025 Rethinking Detecting Salient and Camouflaged Objects in Unconstrained Scenes
Zhangjun Zhou, Chunlin Zhong, Jianuo Huang, Jialun Pei, He Tang 0002
ICCV5
2025 Topology-Constrained Learning for Efficient Laparoscopic Liver Landmark Detection
Ruize Cui, Jiaan Zhang, Jialun Pei, Kai Wang 0092, Pheng-Ann Heng, Harry Qin
MICCAI (10)3
2025 Boosting Few-Shot Semantic Segmentation of 3D Medical Images via Collaborative Slice Alignment
abstract
Few-shot semantic segmentation (FSS) of 3D medical images requires finding a 2D slice from the labeled volume as support to 'query' slices of the unlabeled one. Accurately determining support slices is crucial for learning representative prototypical features, thereby enhancing segmentation accuracy. The existing methods typically resort to the true position of the query target to align the query with support slices or simply exploit one key support slice to segment all query slices, which inevitably results in poor practicality and mis-segmentation. In this regard, we seek a practical and efficient solution by proposing a novel Collaborative Slice Alignment (CSA) module, which densely assigns each query slice its own fittest support without knowing the target prior. Concretely, our CSA first estimates the confidence scores of slices from the sorting task to implicitly reflect their physical location in the human body. The estimated scores are considered as spatial references for aligning support slices and query slices so that each matching pair shares the most similar image contents. Moreover, the self-learnable ranking objective allows CSA to transfer internal knowledge into both support and query features to further boost the FSS performance. Additionally, we introduce an Information Reconciliation (InRe) module to mitigate the inconsistent feature distribution caused by the individual differences between support and query images. Experimental results demonstrate that the combination of CSA and InRe achieves an average Dice score improvement of at least 8.61% across three datasets, consistently outperforming other state-of-the-art methods.
Jialun Pei, Zhiwei Wang 0002, Qiang Li 0018, Pheng-Ann Heng
IEEE J. Biomed. Health Informatics2
2025 Toward Reliable AR-Guided Surgical Navigation: Interactive Deformation Modeling With Data-Driven Biomechanics and Prompts
abstract
In augmented reality (AR)-guided surgical navigation, preoperative organ models are superimposed onto the patient's intraoperative anatomy to visualize critical structures such as vessels and tumors. Accurate deformation modeling is essential to maintain the reliability of AR overlays by ensuring alignment between preoperative models and the dynamically changing anatomy. Although the finite element method (FEM) offers physically plausible modeling, its high computational cost limits intraoperative applicability. Moreover, existing algorithms often fail to handle large anatomical changes, such as those induced by pneumoperitoneum or ligament dissection, leading to inaccurate anatomical correspondences and compromised AR guidance. To address these challenges, we propose a data-driven biomechanics algorithm that preserves FEM-level accuracy while improving computational efficiency. In addition, we introduce a novel human-in-the-loop mechanism into the deformation modeling process. This enables surgeons to interactively provide prompts to correct anatomical misalignments, thereby incorporating clinical expertise and allowing the model to adapt dynamically to complex surgical scenarios. Experiments on a publicly available dataset demonstrate that our algorithm achieves a mean target registration error of 3.42 mm. Incorporating surgeon prompts through the interactive framework further reduces the error to 2.78 mm, surpassing state-of-the-art methods in volumetric accuracy. These results highlight the ability of our framework to deliver efficient and accurate deformation modeling while enhancing surgeon-algorithm collaboration, paving the way for safer and more reliable computer-assisted surgeries.
Jun Zhou 0029, Jialun Pei, Harry Qin, Yingfang Fan, Qi Dou 0001
IEEE Trans. Medical Imaging3
2025 S²Former-OR: Single-Stage Bi-Modal Transformer for Scene Graph Generation in OR
abstract
Scene graph generation (SGG) of surgical procedures is crucial in enhancing holistically cognitive intelligence in the operating room (OR). However, previous works have primarily relied on multi-stage learning, where the generated semantic scene graphs depend on intermediate processes with pose estimation and object detection. This pipeline may potentially compromise the flexibility of learning multimodal representations, consequently constraining the overall effectiveness. In this study, we introduce a novel single-stage bi-modal transformer framework for SGG in the OR, termed S2Former-OR, aimed to complementally leverage multi-view 2D scenes and 3D point clouds for SGG in an end-to-end manner. Concretely, our model embraces a View-Sync Transfusion scheme to encourage multi-view visual information interaction. Concurrently, a Geometry-Visual Cohesion operation is designed to integrate the synergic 2D semantic features into 3D point cloud features. Moreover, based on the augmented feature, we propose a novel relation-sensitive transformer decoder that embeds dynamic entity-pair queries and relational trait priors, which enables the direct prediction of entity-pair relations for graph generation without intermediate steps. Extensive experiments have validated the superior SGG performance and lower computational cost of S2Former-OR on 4D-OR benchmark, compared with current OR-SGG methods, e.g., 3 percentage points Precision increase and 24.2M reduction in model parameters. We further compared our method with generic single-stage SGG methods with broader metrics for a comprehensive evaluation, with consistently better performance achieved. Our source code can be made available at: https://github.com/PJLallen/S2Former-OR.
Jialun Pei, Diandian Guo, Jingyang Zhang, Manxi Lin, Yueming Jin, Pheng-Ann Heng
IEEE Trans. Medical Imaging1
2025 Instrument-Tissue-Guided Surgical Action Triplet Detection via Textual-Temporal Trail Exploration
abstract
Surgical action triplet detection offers intuitive intraoperative scene analysis for dynamically perceiving laparoscopic surgical workflows and analyzing the interaction between instruments and tissues. The current challenge of this task lies in simultaneously localizing surgical instruments while performing more accurate surgical triplet recognition to enhance a comprehensive understanding of intraoperative surgical scenes. To fully leverage the spatial localization of surgical instruments for associating with triplet detection, we propose an Instrument-Tissue-Guided Triplet detector, termed ITG-Trip, which navigates the confluence of surgical action cues through instrument and tissue pseudo-localization labeling to optimize action triplet detection. For exploiting textual and temporal trails, our framework embraces a Visual-Linguistic Association (VLA) module that exploits a pre-trained text encoder to distill textual prior knowledge, enhancing semantic information in global visual features and compensating rare interaction class perception. Besides, we introduce a Mamba-enhanced Spatial-temporal Perception (MSP) decoder, which weaves Mamba and Transformer blocks to explore subject- and object-aware spatial and temporal information to improve the accuracy of action triplet detection in long-time sequence surgical videos. Experimental results on the CholecT50 benchmark indicate that our method significantly outperforms existing state-of-the-art methods in both instrument localization and action triplet detection. The code is available at: github.com/PJLallen/ITG-Trip.
Jialun Pei, Jiaan Zhang, Guanyi Qin, Kai Wang 0092, Yueming Jin, Pheng-Ann Heng
IEEE Trans. Medical Imaging1
2025 DC²T: Disentanglement-Guided Consolidation and Consistency Training for Semi-Supervised Cross-Site Continual Segmentation
abstract
Continual Learning (CL) is recognized to be a storage-efficient and privacy-protecting approach for learning from sequentially-arriving medical sites. However, most existing CL methods assume that each site is fully labeled, which is impractical due to budget and expertise constraint. This paper studies the Semi-Supervised Continual Learning (SSCL) that adopts partially-labeled sites arriving over time, with each site delivering only limited labeled data while the majority remains unlabeled. In this regard, it is challenging to effectively utilize unlabeled data under dynamic cross-site domain gaps, leading to intractable model forgetting on such unlabeled data. To address this problem, we introduce a novel Disentanglement-guided Consolidation and Consistency Training (DC2T) framework, which roots in an Online Semi-Supervised representation Disentanglement (OSSD) perspective to excavate content representations of partially labeled data from sites arriving over time. Moreover, these content representations are required to be consolidated for site-invariance and calibrated for style-robustness, in order to alleviate forgetting even in the absence of ground truth. Specifically, for the invariance on previous sites, we retain historical content representations when learning on a new site, via a Content-inspired Parameter Consolidation (CPC) method that prevents altering the model parameters crucial for content preservation. For the robustness against style variation, we develop a Style-induced Consistency Training (SCT) scheme that enforces segmentation consistency over style-related perturbations to recalibrate content encoding. We extensively evaluate our method on fundus and cardiac image segmentation, indicating the advantage over existing SSCL methods for alleviating forgetting on unlabeled data.
Jingyang Zhang, Jialun Pei, Dunyuan Xu, Yueming Jin, Pheng-Ann Heng
IEEE Trans. Medical Imaging2
2025 Landmark-Free Preoperative-to-Intraoperative Registration in Laparoscopic Liver Resection
abstract
Liver registration by overlaying preoperative 3D models onto intraoperative 2D frames can assist surgeons in perceiving the spatial anatomy of the liver clearly for a higher surgical success rate. Existing registration methods rely heavily on anatomical landmark-based workflows, which encounter two major limitations: 1) ambiguous landmark definitions fail to provide efficient markers for registration; 2) insufficient integration of intraoperative liver visual information in shape deformation modeling. To address these challenges, in this paper, we propose a landmark-free preoperative-to-intraoperative registration framework utilizing effective self-supervised learning, termed Self-P2IR. This framework transforms the conventional 3D-2D workflow into a 3D-3D registration pipeline, which is then decoupled into rigid and non-rigid registration subtasks. Self-P2IR first introduces a feature-disentangled transformer to learn robust correspondences for recovering rigid transformations. Further, a structure-regularized deformation network is designed to adjust the preoperative model to align with the intraoperative liver surface. This network captures structural correlations through geometry similarity modeling in a low-rank transformer network. To facilitate the validation of the registration performance, we also construct an in-vivo registration dataset containing liver resection videos of 21 patients, called P2I-LReg, which contains 346 keyframes that provide a global view of the liver together with liver mask annotations and calibrated camera intrinsic parameters. Extensive experiments and user studies on both synthetic and in-vivo datasets demonstrate the superiority and potential clinical applicability of our method. The code and dataset are available at https://github.com/junzastar/Self-P2IR.
Jun Zhou 0029, Bingchen Gao, Kai Wang 0092, Jialun Pei, Pheng-Ann Heng, Harry Qin
IEEE Trans. Medical Imaging4
2025 Delving Into Quaternion Wavelet Transformer for Facial Expression Recognition in the Wild
abstract
The Facial Expression Recognition (FER) technique has increasingly matured over time. However, recognizing facial expressions in wild environments poses great challenges in achieving promising performance. The main obstacles arise from various factors, such as illumination changes, head pose variations, and occlusions. To overcome interferences from external environments and improve recognition accuracy, we propose a novel Quaternion Wavelet TRansformer (QWTR) model for FER in the wild. Specifically, we present a Quaternion Value Transformer (QVT) network that combines quaternion multi-head attention with quaternion CNN to capture emotional cues from global and local perception. To preserve the color structure while enhancing image contrast and brightness, we introduce a Quaternion Histogram Equalization (QHE) representation to transform color images into quaternion matrices representation. After that, to alleviate the impact of head pose and occlusion together with feature redundancy, a Quaternion Wavelet Feature Selection (QWFS) scheme is designed to decompose quaternion features and select the most correlated signals. Extensive experiments have been conducted on four in-the-wild FER datasets and several specific FER benchmarks under various conditions. The qualitative and quantitative results demonstrate thatQWTRoutperforms other state-of-the-art methods in FER benchmarks, e.g., 68.37% vs. 66.31% accuracy on the AffectNet dataset.
Yu Zhou 0049, Jialun Pei, Weixin Si, Harry Qin, Pheng-Ann Heng
IEEE Trans. Multim.2
2024 Tri-Modal Confluence with Temporal Dynamics for Scene Graph Generation in Operating Rooms
Diandian Guo, Manxi Lin, Jialun Pei, He Tang 0002, Yueming Jin, Pheng-Ann Heng
MICCAI (6)3
2024 Epicardium Prompt-Guided Real-Time Cardiac Ultrasound Frame-to-Volume Registration
Long Lei, Jun Zhou 0007, Jialun Pei, Baoliang Zhao, Yueming Jin, Jeremy Yuen-Chun Teoh, Harry Qin, Pheng-Ann Heng
MICCAI (2)3
2024 Depth-Driven Geometric Prompt Learning for Laparoscopic Liver Landmark Detection
Jialun Pei, Ruize Cui, Yaoqian Li, Weixin Si, Harry Qin, Pheng-Ann Heng
MICCAI (6)1
2024 CalibNet: Dual-Branch Cross-Modal Calibration for RGB-D Salient Instance Segmentation
abstract
In this study, we propose a novel approach for RGB-D salient instance segmentation using a dual-branch cross-modal feature calibration architecture called CalibNet. Our method simultaneously calibrates depth and RGB features in the kernel and mask branches to generate instance-aware kernels and mask features. CalibNet consists of three simple modules, a dynamic interactive kernel (DIK) and a weight-sharing fusion (WSF), which work together to generate effective instance-aware kernels and integrate cross-modal features. To improve the quality of depth features, we incorporate a depth similarity assessment (DSA) module prior to DIK and WSF. In addition, we further contribute a new DSIS dataset, which contains 1,940 images with elaborate instance-level annotations. Extensive experiments on three challenging benchmarks show that CalibNet yields a promising result, i.e., 58.0% AP with 320×480 input size on the COME15K-E test set, which significantly surpasses the alternative frameworks. Our code and dataset will be publicly available at: https://github.com/PJLallen/CalibNet.
Jialun Pei, Tao Jiang 0002, He Tang 0002, Nian Liu 0002, Yueming Jin, Deng-Ping Fan, Pheng-Ann Heng
IEEE Trans. Image Process.1
2023 A Unified Query-based Paradigm for Camouflaged Instance Segmentation
abstract
Due to the high similarity between camouflaged instances and the background, the recently proposed camouflaged instance segmentation (CIS) faces challenges in accurate localization and instance segmentation. To this end, inspired by query-based transformers, we propose a unified query-based multi-task learning framework for camouflaged instance segmentation, termed UQFormer, which builds a set of mask queries and a set of boundary queries to learn a shared composed query representation and efficiently integrates global camouflaged object region and boundary cues, for simultaneous instance segmentation and instance boundary detection in camouflaged scenarios. Specifically, we design a composed query learning paradigm that learns a shared representation to capture object region and boundary features by the cross-attention interaction of mask queries and boundary queries in the designed multi-scale unified learning transformer decoder. Then, we present a transformer-based multi-task learning framework for simultaneous camouflaged instance segmentation and camouflaged instance boundary detection based on the learned composed query representation, which also forces the model to learn a strong instance-level query representation. Notably, our model views the instance segmentation as a query-based direct set prediction problem, without other post-processing such as non-maximal suppression. Compared with 14 state-of-the-art approaches, our UQFormer significantly improves the performance of camouflaged instance segmentation. Our code will be available at: https://github.com/dongbo811/UQFormer.
Jialun Pei, Rongrong Gao, Tian-Zhu Xiang, Shuo Wang 0010, Huan Xiong
ACM Multimedia2
2023 Unite-Divide-Unite: Joint Boosting Trunk and Structure for High-accuracy Dichotomous Image Segmentation
abstract
High-accuracy Dichotomous Image Segmentation (DIS) aims to pinpoint category-agnostic foreground objects from natural scenes. The main challenge for DIS involves identifying the highly accurate dominant area while rendering detailed object structure. However, directly using a general encoder-decoder architecture may result in an oversupply of high-level features and neglect the shallow spatial information necessary for partitioning meticulous structures. To fill this gap, we introduce a novel Unite-Divide-Unite Network (UDUN) that restructures and bipartitely arranges complementary features to simultaneously boost the effectiveness of trunk and structure identification. The proposed UDUN proceeds from several strengths. First, a dual-size input feeds into the shared backbone to produce more holistic and detailed features while keeping the model lightweight. Second, a simple Divide-and-Conquer Module (DCM) is proposed to decouple multiscale low- and high-level features into our structure decoder and trunk decoder to obtain structure and trunk information respectively. Moreover, we design a Trunk-Structure Aggregation module (TSA) in our union decoder that performs cascade integration for uniform high-accuracy segmentation. As a result, UDUN performs favorably against state-of-the-art competitors in all six evaluation metrics on overall DIS-TE, i.e., achieving 0.772 weighted F-measure and 977 HCE. Using 1024X1024 input, our model enables real-time inference at 65.3 fps with ResNet-18. The source code is available at https://github.com/PJLallen/UDUN.
Jialun Pei, Zhangjun Zhou, Yueming Jin, He Tang 0002, Pheng-Ann Heng
ACM Multimedia1
2023 Partitioned Saliency Ranking with Dense Pyramid Transformers
abstract
In recent years, saliency ranking has emerged as a challenging task focusing on assessing the degree of saliency at instance-level. Being subjective, even humans struggle to identify the precise order of all salient instances. Previous approaches undertake the saliency ranking by directly sorting the rank scores of salient instances, which have not explicitly resolved the inherent ambiguities. To overcome this limitation, we propose the ranking by partition paradigm, which segments unordered salient instances into partitions and then ranks them based on the correlations among these partitions. The ranking by partition paradigm alleviates ranking ambiguities in a general sense, as it consistently improves the performance of other saliency ranking models. Additionally, we introduce the Dense Pyramid Transformer (DPT) to enable global cross-scale interactions, which significantly enhances feature interactions with reduced computational burden. Extensive experiments demonstrate that our approach outperforms all existing methods. The code for our method is available at https://github.com/ssecv/PSR.
Chengxiao Sun, Jialun Pei, Haopeng Fang, He Tang 0002
ACM Multimedia3
2023 FGO-Net: Feature and Gaussian Optimization Network for visual saliency prediction
Jialun Pei, He Tang 0002, Chao Liu 0063, Chuanbo Chen
Appl. Intell.1
2023 Depth-Induced Gap-Reducing Network for RGB-D Salient Object Detection: An Interaction, Guidance and Refinement Approach
abstract
Depth provides complementary information for salient object detection (SOD). However, the performance of RGB-D SOD methods is usually hindered by low quality depth map, semantic gap cross-modality and intrinsic gap between multi-level features. Although recent RGB-D SOD methods have been embedded into depth quality assessment, these methods do not consider the inconsistency of the depth format across datasets. In this paper, we propose an interpretable and effective mechanism called interference degree (ID) to assess depth quality and reweight the contribution of single-modality features without extra annotation. Then, a cross-modality interaction block (CMIB) is designed to reduce the semantic gap between RGB and depth features with the help of ID mechanism, and a mutually guided cross-level fusion (MGCF) module is designed to reduce the intrinsic gap among multi-level features. Finally, a refinement branch is proposed to enhance the salient regions and suppress the non-salient regions of fused features. Extensive experiments on six benchmark datasets show that the proposed depth-induced gap-reducing network (DIGR-Net) outperforms 20 recent state-of-the-art methods.
Jialun Pei, He Tang 0002, Zehua Lyu, Chuanbo Chen
IEEE Trans. Multim.3
2023 Transformer-Based Efficient Salient Instance Segmentation Networks With Orientative Query
abstract
Salient instance segmentation (SIS) can be considered as the next generation task for the saliency detection community. Most of the existing state-of-the-art methods used for this novel challenging task are built on the mainstream Mask R-CNN architecture. However, this mechanism relies heavily on hand-designed anchors and NMS post-processing. In this paper, we provide a one stage SIS framework with transformers, termed Orientative Query Transformer (OQTR). To leverage the long-range dependencies of transformers, a cross fusion module is designed to efficiently fuse the global features in the encoder and salient query features for salient mask prediction. Furthermore, derived from the center prior in traditional saliency models, we propose an orientative query that is considered as the initial salient object query to accelerate convergence. In addition, to mitigate the issue of the lack of a large-scale dataset with salient instance labels, we collect a new SIS dataset (SIS10 K) containing over 10 K images elaborately annotated with both object- and instance-level labels to promote the community. Without any post-processing, our end-to-end OQTR framework significantly surpasses the top-1 RDPNet by an average of 13.1% AP scores across all three challenging datasets, demonstrating the strong performance of the proposed OQTR. The code and the dataset proposed in this work are available at:https://github.com/ssecv/OQTR.
Jialun Pei, Tianyang Cheng, He Tang 0002, Chuanbo Chen
IEEE Trans. Multim.1
2022 OSFormer: One-Stage Camouflaged Instance Segmentation with Transformers
Jialun Pei, Tianyang Cheng, Deng-Ping Fan, He Tang 0002, Chuanbo Chen, Luc Van Gool
ECCV (18)1
2022 Salient instance segmentation with region and box-level annotations
Jialun Pei, He Tang 0002, Tianyang Cheng, Chuanbo Chen
Neurocomputing1
2020 Salient instance segmentation via subitizing and clustering
Jialun Pei, He Tang 0002, Chao Liu 0063, Chuanbo Chen
Neurocomputing1