EDBT 2026 Demo / reviewers in the wild / expert
Jinhai Xiang
dblp:150/7211
· DBLP profile ↗
21ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0002-8923-5302ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Understanding In-Context Learning of Transformers Under Non-I.I.D. ScenariosabstractUnderstanding the generalization behavior of in-context learning (ICL) in Transformers remains a fundamental challenge, as most existing theoretical analyses are based on the assumption that data are independently and identically distributed (i.i.d.), an assumption that often does not hold in practice. Motivated by the theoretical insight that ICL operates similarly to gradient-based optimization, we leverage the concept of gradient stability to establish generalization error bounds for ICL under a general non-i.i.d. setting. Our analysis shows that two factors play a central role in ICL generalization: the number of demonstrations in the prompt and their distributional alignment with the query. In particular, increasing the number of demonstrations and improving their alignment with the query distribution lead to better generalization, even without any parameter tuning. Under mild conditions, we further prove that the generalization error can achieve the optimal convergence rate of O(N^(-1/2)), where N is the number of demonstrations. Our empirical evaluations validate the effectiveness of our theoretical findings. Qilu Shen, Jinhai Xiang |
AAAI | 3 |
| 2025 | Tetrahedron-Net for Medical Image RegistrationabstractMedical image registration plays a vital role in medical image processing. Extracting expressive representations for medical images is crucial for improving the registration quality. One common practice for this end is constructing a convolutional backbone to enable interactions with skip connections among feature extraction layers. The de facto structure, U-Net-like networks, has attempted to design skip connections such as nested or full-scale ones to connect one single encoder and one single decoder to improve its representation capacity. Despite being effective, it still does not fully explore interactions with a single encoder and decoder architectures. In this paper, we embrace this observation and introduce a simple yet effective alternative strategy to enhance the representations for registrations by appending one additional decoder. The new decoder is designed to interact with both the original encoder and decoder. In this way, it not only reuses feature presentation from corresponding layers in the encoder but also interacts with the original decoder to corporately give more accurate registration results. The new architecture is concise yet generalized, with only one encoder and two decoders forming a “Tetrahedron” structure, thereby dubbed Tetrahedron-Net. Three instantiations of Tetrahedron-Net are further constructed regarding the different structures of the appended decoder. Our extensive experiments prove that superior performance can be obtained on several representative benchmarks of medical image registration. Finally, such a “Tetrahedron” design can also be easily integrated into popular U-Net-like architectures including VoxelMorph, ViT-V-Net, and TransMorph, leading to consistent performance gains. Jinhai Xiang, Dantong Shi, Xinwei He 0001 |
BIBM | 1 |
| 2025 | TeDA: Boosting Vision-Lanuage Models for Zero-Shot 3D Object Retrieval via Testing-time Distribution AlignmentabstractLearning discriminative 3D representations that generalize well to unknown testing categories is an emerging requirement for many real-world 3D applications. Existing well-established methods often struggle to attain this goal due to insufficient 3D training data from broader concepts. Meanwhile, pre-trained large vision-language models (e.g., CLIP) have shown remarkable zero-shot generalization capabilities. Yet, they are limited in extracting suitable 3D representations due to substantial gaps between their 2D training and 3D testing distributions. To address these challenges, we propose Testing-time Distribution Alignment (TeDA), a novel framework that adapts a pretrained 2D vision-language model CLIP for unknown 3D object retrieval at test time. To our knowledge, it is the first work that studies the test-time adaptation of a vision-language model for 3D feature learning. TeDA projects 3D objects into multi-view images, extracts features using CLIP, and refines 3D query embeddings with an iterative optimization strategy by confident query-target sample pairs in a self-boosting manner. Additionally, TeDA integrates textual descriptions generated by a multimodal language model (InternVL) to enhance 3D object understanding, leveraging CLIP's aligned feature space to fuse visual and textual cues. Extensive experiments on four open-set 3D object retrieval benchmarks demonstrate that TeDA greatly outperforms state-of-the-art methods, even those requiring extensive training. We also experimented with depth maps on Objaverse-LVIS, further validating its effectiveness. Code is available at https://github.com/wangzhichuan123/TeDA. Yang Zhou 0007, Jinhai Xiang, Yulong Wang 0002, Xinwei He 0001 |
ICMR | 3 |
| 2025 | Graph Contrastive Learning via Hierarchical Multiview Enhancement for RecommendationabstractIn the field of recommender systems, self-supervised learning has become an effective framework. In response to the noisy interaction behaviors in realworld scenarios, as well as the skewed distribution influenced by data sparsity and popularity bias, graph contrastive learning has been introduced as a powerful self-supervised method in collaborative filtering (CF) to learn enhanced user and item representations. Despite their success, neither heuristic manual enhancement methods nor the use of final node representations to construct contrastive pairs are sufficient to provide effective and rich self-supervised signals to regulate the training process. Therefore, the learned representations of users and items are either fragile or lack heuristic guidance. In light of this, we propose the Hierarchical multiview graph contrastive learning framework HMCF, which leverages the message passing mechanism at the layer level to introduce different granularity levels of view augmentation using supervised signals, thus better enhancing the CF paradigm. HMCF leverages rich, high-quality self-supervised signals from different granularity views for accurate contrastive optimization, helping to alleviate data sparsity and noise issues. It also explains the hierarchical topology and relative distances between nodes in the original graph. Comprehensive experiments on three public datasets shows that our model significantly outperforms the state-of-the-art baselines. Zhi Liu 0011, Hengjing Xiang, Ruxia Liang, Jinhai Xiang, Chaodong Wen, Sannyuya Liu |
IEEE Trans. Ind. Informatics | 4 |
| 2024 | BcMatch: Semi-supervised Medical Image Segmentation with Bias CorrectionabstractSemi-supervised semantic segmentation (SSS), which allows for learning a better model with a small fraction of labeled samples and a large amount of unlabeled ones, is valuable yet challenging in medical image analysis. Recent works (e.g., UniMatch) have found that weak-to-strong consistency via augmentation is especially conducive to SSS training. However, they inadvertently introduce cognitive biases for unlabeled images, making it difficult to segment accurately near the target’s edge regions. In this paper we present a lightweight bias-correct module to self-correct these mistakes between the strong perturbations. Based on it, we design a new framework named BcMatch by plugging it into UniMatch to reduce such cognitive biases stemming from incorrect pseudo-labels for unlabeled images. Moreover, we also introduce a bias correction loss, which works in tandem with the consistency loss to guide the model learning, focusing more on the edge regions of the targets. Experiments on the representative semi-supervised segmentation dataset, ACDC, demonstrate our BcMatch surpasses UniMatch by a large margin, attaining new state-of-the-art performance. The code is at https://github.com/zhangyan498/BcMatch. Jinhai Xiang, Jiakun Yu, Xinwei He 0001 |
BIBM | 1 |
| 2024 | Dual decoder UNet with Contrastive learning for Brain Image RegistrationabstractThe core of the image registration task is to accurately extract and compare the spatial feature information between moving and fixed images. The model must not only be able to capture features inside an image, but also be able to distinguish similarities and differences between moving image and fixed image. Although UNet and its variant networks are widely used in the field of registration, they still present challenges in terms of fine-grained and multi-level feature extraction. In order to address these issues, we propose a dual-decoder UNet network based on contrastive learning. For enhancing the discriminant ability of the network at the feature level, the proposed model applies contrastive learning to the encoder stage of UNet, which strengthens the differential learning in the process of feature extraction. In addition, a multi-level feature fusion method is introduced in the encoder to consider the feature differences at each level more comprehensively. Experimental results show that the proposed method has achieved a significant improvement in the registration accuracy. On the LPBA40 dataset, compared with Affine, VoxelMorph and CLMorph, the Dice scores were increased by 15.2%, 2.5% and 1.9%, respectively, which verified the effectiveness of the proposed method. Our code implementation is publicly available at https://github.com/LaneAzur/DD-CLUNet. Dantong Shi, Jinhai Xiang |
BIBM | 4 |
| 2024 | SimpleFusion: A Simple Fusion Framework for Infrared and Visible Images
Yuxuan Cheng, Xinwei He 0001, Yan Aze, Jinhai Xiang |
PRCV (8) | 6 |
| 2023 | AiA-UNet: Attention in Attention for Medical Image SegmentationabstractIn medical image segmentation, U-Net has consistently played a vital role. Recently, the U-Net networks based on the Vision Transformer (ViT) architecture have become more and more popular. ViT exhibits superior capabilities in handling long-range dependencies and capturing global contextual information. However, it requires significant computational cost, and does not explore the optimal matching and the potential dependencies between different patches. To address the aforementioned issues, we propose a novel network framework, called AiA-UNet, for medical image segmentation. The AiA-UNet makes two main contributions. A convolutional self-attention mechanism is proposed to replace the self-attention module in ViT, effectively reducing computational complexity. Moreover, an Attention in Attention module (AiA) is applied within the ViT block. Experimental results on the Synapse multi-organ segmentation dataset demonstrate that AiA-UNet outperforms Trans-UNet by 5.40% and Swin-UNet by 3.75%. Code and models are available at https://github.com/xiaoqin1998/AiA-UNet. Jianfeng Qin, Xinwei He 0001, Jiakun Yu, Jinhai Xiang, Lulu Wu |
BIBM | 5 |
| 2023 | Trans-UNeter: A new Decoder of TransUNet for Medical Image SegmentationabstractRecently, how to integrate convolutional neural networks and transformers into a U-Net-like encoder-decoder structure has drawn growing interest in medical image segmentation, as transformer is more efficient in capturing longrange relations. Following this line of research, TransUNet is one representative work. However, it still insufficiently explores the rich relations of features from the encoder layer and the decoder layer with just a simple concatenation, which weakens their effectiveness to some extent. To address this issue, we propose two important design improvements to strengthen TransUNet: 1) a novel skip connection module, which upsamples the high-level semantic features and fuse it with low-level features, producing comprehensive semantic-aware features for the decoder. 2) an improved decoder network cascades reverse attention and spatial attention to adaptively combines features from the corresponding encoder layer and the previously decoded outputs.The results of the abdominal multi-organ segmentation experiment on the Synapse multi-organ segmentation dataset indicated that Trans-UNeter improved the mean similarity coefficient(DSC) by 3.71% compared to TransUNet. Code and models are available at https://github.com/iaoqin/Trans-UNeter. Jiakun Yu, Xinwei He 0001, Jianfeng Qin, Jinhai Xiang, Weiming Zhao |
BIBM | 5 |
| 2023 | CCAC: Contrastive Learning with Channel Attention and Contour Loss for Brain Image RegistrationabstractIn medical image analysis, image registration plays a key role in establishing the deformation field between image pairs. A good deformation model is essential for high-quality estimates. However, most existing approaches still face great challenges in extracting better image features and generating detailed information on deformation fields. In this paper, we present a brain image registration model based on contrastive learning with channel attention and contour loss (CCAC). To integrate lower-level features more effectively, we propose a channel attention-based skip connection to obtain weight information of features at different encoder layers. To enhance the feature consistency of semantics in different stages, we then introduce multi-stage contrastive learning to capture semantics at different stages. Moreover, in order to reduce the influence caused by the disparities of inter-image gray value in registered images, contour feature loss is employed. Both qualitative and quantitative experiments on the popular LPBA40 dataset demonstrate that our method can achieve much better performance than the baseline methods. Dantong Shi, Jinhai Xiang |
BIBM | 4 |
| 2022 | GL-GAN: Adaptive global and local bilevel optimization for generative adversarial network
Liu Ying, Heng Fan 0001, Xiaohui Yuan 0001, Jinhai Xiang |
Pattern Recognit. | 4 |
| 2021 | ClsGAN: Selective Attribute Editing Model based on Classification Adversarial Network
Liu Ying, Heng Fan 0001, Fuchuan Ni, Jinhai Xiang |
Neural Networks | 4 |
| 2019 | Complementary Siamese Networks for Robust Visual TrackingabstractIn this paper, we propose the novel complementary Siamese networks (CoSNet) for visual tracking by exploiting complementary global and local representations to learn a matching function. In specific, the proposed CoSNet is two-fold: a global Siamese network (GSNet) and a local Siamese network (LSNet). The GSNet aims to match the target with candidates using holistic representation. By contrast, the LSNet explores partial object representation for matching. Instead of simply decomposing the object into regular patches in LSNet, we propose a novel attentional local part network, which automatically generates salient object parts for local representation and adaptively weights each part according to its importance in matching. In CoSNet, the GSNet and LSNet are jointly trained in an end-to-end manner. By coupling two complementary Siamese networks, our CoSNet learns a robust matching function which can effectively handle various appearance changes in visual tracking. Extensive experiments on a large-scale dataset with 100 sequences show that CoSNet outperforms other state-of-the-art trackers. Heng Fan 0001, Jinhai Xiang |
ICASSP | 3 |
| 2019 | Hierarchical Multi-Task Network For Race, Gender and Facial Attractiveness RecognitionabstractDeep learning has powered many face related tasks and shown state-of-the-art performance. However, existing deep models are often trained separately for different problems, which results in heavy computational burden. To address this problem, we propose a novel multi-task network with fully convolutional architecture-Hierarchical Multi-task Network (HMT-Net), that simultaneously recognizes a person's gender, race and facial attractiveness from a given portrait image. Aiming to improve the robustness to outliers in facial beauty prediction task, a novel loss is introduced into HMTNet. Compared to existing deep approaches, the proposed HMTNet achieves state-of-the-art performance on several datasets, and it can learn more discriminative feature representation through joint training and feature aggregation. Extensive experiments evidence the effectiveness of HMTNet. Heng Fan 0001, Jinhai Xiang |
ICIP | 3 |
| 2017 | Robust Visual Tracking via Local-Global Correlation FilterabstractCorrelation filter has drawn increasing interest in visual tracking due to its high efficiency, however, it is sensitive to partial occlusion, which may result in tracking failure. To address this problem, we propose a novel local-global correlation filter (LGCF) for object tracking. Our LGCF model utilizes both local-based and global-based strategies, and effectively combines these two strategies by exploiting the relationship of circular shifts among local object parts and global target for their motion models to preserve the structure of object. In specific, our proposed model has two advantages: (1) Owing to the benefits of local-based mechanism, our method is robust to partial occlusion by leveraging visible parts. (2) Taking into account the relationship of motion models among local parts and global target, our LGCF model is able to capture the inner structure of object, which further improves its robustness to occlusion. In addition, to alleviate the issue of drift away from object, we incorporate temporal consistencies of both local parts and global target in our LGCF model. Besides, we adopt an adaptive method to accurately estimate the scale of object. Extensive experiments on OTB15 with 100 videos demonstrate that our tracking algorithm performs favorably against state-of-the-art methods. Heng Fan 0001, Jinhai Xiang |
AAAI | 2 |
| 2017 | Robust visual tracking via deep discriminative modelabstractIn this paper, we exploit deep convolutional features for object appearance modeling and propose a simple while effective deep discriminative model (DDM) for visual tracking. The proposed DDM takes as input the deep features and outputs an object-background confidence map. Considering that both spatial information from lower convolutional layers and semantic information from higher layers benefit object tracking, we construct multiple deep discriminative models (DDMs) for each layer and combine these confidence maps from each layer to obtain the final object-background confidence map. To reduce the risk of model drift, we propose to adopt a saliency method to generate object candidates. Object tracking is then achieved by finding the candidate with the largest confidence value. Experiments on a large-scale tracking benchmark demonstrate that the propose method performs favorably against state-of-the-art trackers. Heng Fan 0001, Jinhai Xiang, Guoliang Li 0002, Fuchuan Ni |
ICASSP | 2 |
| 2017 | Robust Visual Tracking With Multitask Joint Dictionary LearningabstractDictionary learning for sparse representation has been increasingly applied to object tracking, however, the existing methods only utilize one modality of the object to learn a single dictionary. In this paper, we propose a robust tracking method based on multitask joint dictionary learning. Through extracting different features of the target, multiple linear sparse representations are obtained. Each sparse representation can be learned by a corresponding dictionary. Instead of separately learning the multiple dictionaries, we adopt a multitask learning approach to learn the multiple linear sparse representations, which provide additional useful information to the classification problem. Because different tasks may favor different sparse representation coefficients, yet the joint sparsity may enforce the robustness in coefficient estimation. During tracking, a classifier is constructed based on a joint linear representation, and the candidate with the smallest joint decision error is selected to be the tracked object. In addition, reliable tracking results and augmented training samples are accumulated into two sets to update the dictionaries for classification, which helps our tracker adapt to the fast time-varying object appearance. Both qualitative and quantitative evaluations on CVPR2013 visual tracking benchmark demonstrate that our method performs favorably against state-of-the-art trackers. Heng Fan 0001, Jinhai Xiang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Robust visual tracking via bag of superpixels
Heng Fan 0001, Jinhai Xiang |
Multim. Tools Appl. | 2 |
| 2015 | Patch-Based Visual Tracking with Two-Stage Multiple Kernel Learning
Heng Fan 0001, Jinhai Xiang |
ICIG (3) | 2 |
| 2015 | Robust tracking based on local structural cell graph
Heng Fan 0001, Jinhai Xiang, Hong-Hong Liao, Xiaoping Du |
J. Vis. Commun. Image Represent. | 2 |
| 2011 | An Abnormal Event Recognition in Crowd SceneabstractReal-world actions occur often in crowded, dynamic environments. This poses a difficult challenge for current approaches to video event detection because crowd scenes are always extremely cluttered. In this paper, we design a video content analysis method for fighting event recognition in crowd scene. Our method begins with four MPEG-7 descriptors: crowd kinetic energy, motion directions histogram, spatial distribution parameter and spatial localization parameter of two adjacent frames. Then the support vector machines (SVMs) method is introduced to train and test these descriptors for fighting event recognition. Extensive experimental results have demonstrated that our method is effective in fighting events recognition with low error rates and can be easily adopted in fixed camera environment with real time application. Hong-Hong Liao, Jinhai Xiang, Wei-Ping Sun, Jianghua Dai |
ICIG | 2 |