Xibin Jia

dblp:141/0898 · also XiBin Jia · DBLP profile ↗
← Back
29ranked-venue papers
14as first author
20since 2021 · last 2026
0000-0001-8799-8042ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 8 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Semi-supervised boundary-aware medical image segmentation via symmetric boundary-foreground collaboration
Xibin Jia, Luo Wang, Chuanxu Yang, Xunjie Yin, Zhenghan Yang, Min Hong
Expert Syst. Appl.1
2026 Domain generalization for lesion classification via frequency swapping and soft-mask disentanglement
Xibin Jia, Shaowu Xu, Chao Fan 0001, Zhenghan Yang
Expert Syst. Appl.1
2026 VLAlignSeg: Vision-language alignment for few-shot medical image segmentation
Haipeng Qiao, Xibin Jia, Chao Fan 0001
Expert Syst. Appl.3
2026 Guided Hierarchical Interaction Network for RGB-D semantic segmentation
Binhui Wang, Yiheng Cai, Xibin Jia
Image Vis. Comput.4
2025 SPENet: Self-guided Prototype Enhancement Network for Few-Shot Medical Image Segmentation
Chao Fan 0001, Xibin Jia, Anqi Xiao, Hongyuan Yu, Zhenghan Yang, Yan Huang 0008, Liang Wang 0001
MICCAI (5)2
2025 Cross-Modal Dual-Causal Learning for Long-Term Action Recognition
abstract
Long-term action recognition (LTAR) is challenging due to extended temporal spans with complex atomic action correlations and visual confounders. Although vision-language models (VLMs) have shown promise, they often rely on statistical correlations instead of causal mechanisms. Moreover, existing causality-based methods address modal-specific biases but lack cross-modal causal modeling, limiting their utility in VLM-based LTAR. This paper proposes Cross-Modal Dual-Causal Learning (CMDCL), which introduces a structural causal model to uncover causal relationships between videos and label texts. CMDCL addresses cross-modal biases in text embeddings via textual causal intervention and removes confounders inherent in the visual modality through visual causal intervention guided by the debiased text. These dual-causal interventions enable robust action representations to address LTAR challenges. Experimental results on three benchmarks including Charades, Breakfast and COIN, demonstrate the effectiveness of the proposed model. Our code is available at https://github.com/xushaowu/CMDCL.
Shaowu Xu, Xibin Jia, Junyu Gao 0002, Qianmei Sun, Jing Chang 0006, Chao Fan 0001
ACM Multimedia2
2025 WaveConvX: Multi-Level Wavelet Enhancement for Histopathology Image Classification
abstract
Artificial histopathological image classification is emerging as a promising approach for assisting early cancer diagnosis and informing treatment planning. Although convolutional neural network (CNN) based methods have achieved significant success, they still struggle to capture the rich multi-scale and high-frequency details inherent in pathological tissues. To address the above issue, we propose WaveConvX, a novel histopathological image classification model that integrates multi-level wavelet decomposition with the ConvNeXt backbone. Our approach applies a two-stage discrete wavelet transform (DWT) to intermediate feature maps, enabling the decomposition of features into multiple frequency sub-bands. High-frequency components are enhanced using Adaptive Power Gabor Convolution (APGConv), while mid-frequency ranges are refined through tailored attention mechanisms. These processed sub-bands are then fused via inverse wavelet transforms, producing feature maps enriched with both global context and local morphological cues essential for accurate diagnosis. Comprehensive experiments are conducted on three public datasets (BreakHis for breast cancer, KBSMC for gastric cancer, and LC25000 for lung and colon cancer). The results demonstrate that WaveConvX consistently outperforms ten state-of-the-art benchmarks, achieving superior accuracy, F1 scores, and robustness across multiple types and magnifications of cancer. Our work demonstrates the significant potential of wavelet-enhanced CNNs in histopathological image analysis.
Feiran Liu, Xibin Jia
SMC5
2025 Temporal fidelity enhancement for video action recognition
abstract
Temporal attention mechanisms are essential for video action recognition, enabling models to focus on semantically informative moments. However, these models frequently exhibit temporal infidelity—misaligned attention weights caused by limited training diversity and the absence of fine-grained temporal supervision. While video-level labels provide coarse-grained action guidance, the lack of detailed constraints allows attention noise to persist, especially in complex scenarios with distracting spatial elements. To address this issue, we propose temporal fidelity enhancement (TFE), a competitive learning paradigm based on the disentangled information bottleneck (DisenIB) theory. TFE mitigates temporal infidelity by decoupling action-relevant semantics from spurious correlations through adversarial feature disentanglement. Using pre-trained representations for initialization, TFE establishes an adversarial process in which segments with elevated temporal attention compete against contexts with diminished action relevance. This mechanism ensures temporal consistency and enhances the fidelity of attention patterns without requiring explicit fine-grained supervision. Extensive studies on UCF101, HMDB-51, and Charades benchmarks validate the effectiveness of our method, with significant improvements in action recognition accuracy.
Shaowu Xu, Xibin Jia, Qianmei Sun, Jing Chang 0006
Frontiers Inf. Technol. Electron. Eng.2
2025 SliceMamba With Neural Architecture Search for Medical Image Segmentation
abstract
Despite the progress made in Mamba-based medical image segmentation models, existing methods utilizing unidirectional or multi-directional feature scanning mechanisms struggle to effectively capture dependencies between neighboring positions, limiting the discriminant representation learning of local features. These local features are crucial for medical image segmentation as they provide critical structural information about lesions and organs. To address this limitation, we propose SliceMamba, a simple yet effective locally sensitive Mamba-based medical image segmentation model. SliceMamba features an efficient Bidirectional Slicing and Scanning (BSS) module, which performs bidirectional feature slicing and employs varied scanning mechanisms for sliced features with distinct shapes. This design keeps spatially adjacent features close in the scan sequence, preserving the local structure of the image and enhancing segmentation performance. Additionally, to fit the varying sizes and shapes of lesions and organs, we introduce an Adaptive Slicing Search method that automatically identifies the optimal feature slicing method based on the characteristics of the target data. Extensive experiments on two skin lesion datasets (ISIC2017 and ISIC2018), two polyp segmentation datasets (Kvasir and ClinicDB), one ultra-wide field retinal hemorrhage segmentation dataset (UWF-RHS), and one multi-organ segmentation dataset (Synapse) demonstrate the effectiveness of our method.
Chao Fan 0001, Hongyuan Yu, Yan Huang 0008, Liang Wang 0001, Zhenghan Yang, Xibin Jia
IEEE J. Biomed. Health Informatics6
2023 A Few-Shot Medical Image Segmentation Network with Boundary Category Correction
Xibin Jia, Xiong Guo, Luo Wang
PRCV (10)2
2023 An improved unified domain adversarial category-wise alignment network for unsupervised cross-domain sentiment classification
Xibin Jia, Meng Zeng, Luo Wang, Qing Mi
Eng. Appl. Artif. Intell.1
2023 Graph Contrastive Learning with Constrained Graph Data Augmentation
Shaowu Xu, Luo Wang, Xibin Jia
Neural Process. Lett.3
2023 A novel dual-channel graph convolutional neural network for facial action unit recognition
Xibin Jia, Shaowu Xu, Luo Wang, Weiting Li
Pattern Recognit. Lett.1
2023 What makes a readable code? A causal analysis method
abstract
Abstract Context Code readability is one of the most important quality attributes for software source code. To investigate which features affect code readability, most existing studies rely on correlation‐based methods. However, spurious correlations (a mathematical relationship wherein two variables appear to be causal but are not) involved in correlation‐based methods may affect research conclusions. Objective In order to remove spurious correlations and obtain conclusions from the perspective of causation as to what makes a readable code, we propose a causal theory‐based approach to analyze the relationship between code features and code readability scores. Method First, we adopt the PC algorithm and additive noise models to construct the causal graph on the basis of the selected code features. Then, we use the linear regression algorithm based on the back‐door criterion to obtain the causal effect of different features on code readability. Result We conduct a set of experiments on readability data labeled by human annotators. The experimental results show that the average number of comments positively impacts code readability, with each additional unit increasing the code readability score by 0.799 points. Whereas the average number of assignments, identifiers, and periods have a negative impact, with each additional unit decreasing the code readability score by 0.528, 0.281, and 0.170 points respectively. Conclusion We believe that our findings will provide developers with a better understanding of the patterns behind code readability, and guide developers to optimize their code as the ultimate goal.
Qing Mi, Zhi Cai, Xibin Jia
Softw. Pract. Exp.4
2022 Robust Liver Segmentation Using Boundary Preserving Dual Attention Network
Xibin Jia, Luo Wang
PRCV (2)2
2022 Data-aware relation learning-based graph convolution neural network for facial action unit recognition
Xibin Jia, Weiting Li
Pattern Recognit. Lett.1
2022 A Multimodality-Contribution-Aware TripNet for Histologic Grading of Hepatocellular Carcinoma
abstract
Hepatocellular carcinoma (HCC) is a type of primary liver malignant tumor with a high recurrence rate and poor prognosis even undergoing resection or transplantation. Accurate discrimination of the histologic grades of HCC plays a critical role in the management and therapy of HCC patients. In this paper, we discuss a deep learning-based diagnostic model for HCC histologic grading with multimodal Magnetic Resonance Imaging (MRI) images to overcome the problem of limited well-annotated data and extract the discriminated fusion feature referring to the clinical diagnosis experience of radiologists. Accordingly, we propose a novel Multimodality-Contribution-Aware TripNet (MCAT) based on the metric learning and the attention-aware weighted multimodal fusion. The novelty of the method lies in the multimodality small-shot learning architecture designation and the multimodality adaptive weighted computing scheme. The comprehensive experiments are done on the clinic dataset with the well-annotation of lesion location by the professional radiologist. The experimental results show that our proposed MCAT is not only able to achieve acceptable quantitative measuring of HCC histologic grading based on the MRI sequences with small cases but also outperforms previous models in HCC histologic grading, reaching an accuracy of 84 percent, a sensitivity of 87 percent and precision of 89 percent.
Xibin Jia, Qing Mi, Zhenghan Yang
IEEE ACM Trans. Comput. Biol. Bioinform.1
2021 MAGAN: Multi-attention Generative Adversarial Networks for Text-to-Image Generation
Xibin Jia, Qing Mi
PRCV (4)1
2021 An unsupervised person re-identification approach based on cross-view distribution alignment
abstract
Abstract Unsupervised clustering is a kind of popular solution for unsupervised person re‐identification (re‐ID). However, due to the influence of cross‐view differences, the results of clustering labels are not accurate. To solve this problem, an unsupervised re ID method based on cross‐view distributed alignment (CV‐DA) to reduce the influence of unsupervised cross‐view is proposed. Specifically, based on a popular unsupervised clustering method, density clustering DBSCAN is used to obtain pseudo labels. By calculating the similarity scores of images in the target domain and the source domain, the similarity distribution of different camera views is obtained and is aligned with the distribution with the consistency constraint of pseudo labels. The cross‐view distribution alignment constraint is used to guide the clustering process to obtain a more reliable pseudo label. The comprehensive comparative experiments are done in two public datasets, i.e. Market‐1501 and DukeMTMC‐reID. The comparative results show that the proposed method outperforms several state‐of‐the‐art approaches with mAP reaching 52.6% and rank1 71.1%. In order to prove the effectiveness of the proposed CV‐DA, the proposed constraint is added into two advanced re‐ID methods. The experimental results demonstrate that the mAP and rank increase by 0.5–2% after using the cross‐view distribution alignment constraint comparing with that of the associated original methods without using CV‐DA.
Xibin Jia, Qing Mi
IET Image Process.1
2021 The effectiveness of data augmentation in code readability classification
Qing Mi, Yan Xiao 0002, Zhi Cai, Xibin Jia
Inf. Softw. Technol.4
2019 Style Consistency Constrained Fusion Feature Learning for Liver Tumor Segmentation
Xibin Jia, Zhenghan Yang
PRCV (3)2
2019 Domain-invariant representation learning using an unsupervised domain adversarial adaptation deep neural network
Xibin Jia, Ya Jin, Xing Su 0001, Yongli Hu
Neurocomputing1
2018 Multi-vehicles dynamic navigating method for large-scale event crowd evacuations
Zhi Cai, Fujie Ren, Yuanying Chi, Xibin Jia, Lijuan Duan, Zhiming Ding
GeoInformatica4
2018 Words alignment based on association rules for cross-domain sentiment classification
abstract
Automatic classification of sentiment data (e.g., reviews, blogs) has many applications in enterprise user management systems, and can help us understand people’s attitudes about products or services. However, it is difficult to train an accurate sentiment classifier for different domains. One of the major reasons is that people often use different words to express the same sentiment in different domains, and we cannot easily find a direct mapping relationship between them to reduce the differences between domains. So, the accuracy of the sentiment classifier will decline sharply when we apply a classifier trained in one domain to other domains. In this paper, we propose a novel approach called words alignment based on association rules (WAAR) for cross-domain sentiment classification, which can establish an indirect mapping relationship between domain-specific words in different domains by learning the strong association rules between domain-shared words and domain-specific words in the same domain. In this way, the differences between the source domain and target domain can be reduced to some extent, and a more accurate cross-domain classifier can be trained. Experimental results on Amazon® datasets show the effectiveness of our approach on improving the performance of cross-domain sentiment classification.
Xibin Jia, Ya Jin, Xing Su 0001, Barry Cardiff, Bir Bhanu
Frontiers Inf. Technol. Electron. Eng.1
2017 A Multiway Semi-supervised Online Sequential Extreme Learning Machine for Facial Expression Recognition with Kinect RGB-D Images
Xibin Jia
ICIC (2)1
2016 Local Invariance Representation Learning Algorithm with Multi-layer Extreme Learning Machine
Xibin Jia, Hua Du, Bir Bhanu
ICONIP (4)1
2016 A semi-supervised online sequential extreme learning machine method
Xibin Jia, Runyuan Wang, Junfa Liu, David M. W. Powers
Neurocomputing1
2016 A novel edge detection approach using a fusion model
Xibin Jia, Haiyong Huang, Jianming Yuan, David M. W. Powers
Multim. Tools Appl.1
2004 Audio to visual signal mappings with HMM
abstract
There has been a large amount of research on speech driven face animation. Particularly, recently research efforts have been demonstrated that the hidden Markov model techniques could achieve a high level of success in the field of audio/visual mapping without language information. In this paper, firstly a linear model based facial representation method was applied, which extracts face features as a global feature. Secondly, a HMM based method was presented, which includes a two-level frame to promote the audio-visual mapping result.
Xibin Jia, Dehui Kong
ICASSP (5)3