VLDB 2026 Research / reviewers in the wild / expert
Xi Jia
dblp:134/4847
· DBLP profile ↗
23ranked-venue papers
5as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AEPL: Adaptive empirical prototype learning with dynamic margins for deep face recognition
Weijia Fan, Zhixiang Cai, Chunsong Chen, Yanxi Liu 0004, Jiajun Wen 0001, Xi Jia, LinLin Shen, Jiancan Zhou, Qiufu Li |
Pattern Anal. Appl. | 7 |
| 2025 | Decoder-Only Image RegistrationabstractIn unsupervised medical image registration, encoder-decoder architectures are widely used to predict dense, full-resolution displacement fields from paired images. Despite their popularity, we question the necessity of making both the encoder and decoder learnable. To address this, we propose LessNet, a simplified network architecture with only a learnable decoder, while completely omitting a learnable encoder. Instead, LessNet replaces the encoder with simple, handcrafted features, eliminating the need to optimize encoder parameters. This results in a compact, efficient, and decoder-only architecture for 3D medical image registration. We evaluate our decoder-only LessNet on five registration tasks: 1) inter-subject brain registration using the OASIS-1 dataset, 2) atlas-based brain registration using the IXI dataset, 3) cardiac ES-ED registration using the ACDC dataset, 4) inter-subject abdominal MR registration using the CHAOS dataset, and 5) multi-study, multi-site brain registration using images from 13 public datasets. Our results demonstrate that LessNet can effectively and efficiently learn both dense displacement and diffeomorphic deformation fields. Furthermore, our decoder-only LessNet can achieve comparable registration performance to benchmarking methods such as VoxelMorph and TransMorph, while requiring significantly fewer computational resources. Our code and pre-trained models are available at https://github.com/xi-jia/LessNet. Xi Jia, Wenqi Lu 0001, Xinxing Cheng, Jinming Duan 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2024 | MERG: Multi-Dimensional Edge Representation Generation Layer for Graph Neural NetworksabstractEdges are essential in describing relationships among nodes. While existing graphs frequently use a single-value edge to describe association between each pair of node vectors, crucial relationships may be disregarded if they are not linearly correlated, which may limit graph analysis performance. Although some recent Graph Neural Networks (GNNs) can process graphs containing multi-dimensional edge features, they cannot convert single-value edge graphs to multi-dimensional edge graphs during propagation. This paper proposes a generic Multi-dimensional Edge Representation Generation (MERG) layer that can be inserted into any GNNs for heterogeneous graph analysis. It assigns multi-dimensional edge features for the input single-value edge graph, describing multiple task-specific and global context-aware relationship cues between each connected node pair. Results on eight graph benchmark datasets demonstrate that inserting the MERG layer into widely-used GNNs (e.g., GatedGCN and GAT) leads to major performance improvements, resulting in state-of-the-art (SOTA) results on seven out of eight evaluated datasets. Our code is publicly available at1. YuXin Song 0001, Aaron S. Jackson, Xi Jia, Weicheng Xie 0001, LinLin Shen, Hatice Gunes, Siyang Song |
ICASSP | 4 |
| 2024 | WiNet: Wavelet-Based Incremental Learning for Efficient Medical Image Registration
Xinxing Cheng, Xi Jia, Wenqi Lu 0001, Qiufu Li, LinLin Shen, Alexander Krull, Jinming Duan 0001 |
MICCAI (2) | 2 |
| 2024 | Structure and Intensity Unbiased Translation for 2D Medical Image SegmentationabstractData distribution gaps often pose significant challenges to the use of deep segmentation models. However, retraining models for each distribution is expensive and time-consuming. In clinical contexts, device-embedded algorithms and networks, typically unretrainable and unaccessable post-manufacture, exacerbate this issue. Generative translation methods offer a solution to mitigate the gap by transferring data across domains. However, existing methods mainly focus on intensity distributions while ignoring the gaps due to structure disparities. In this paper, we formulate a new image-to-image translation task to reduce structural gaps. We propose a simple, yet powerful Structure-Unbiased Adversarial (SUA) network which accounts for both intensity and structural differences between the training and test sets for segmentation. It consists of a spatial transformation block followed by an intensity distribution rendering module. The spatial transformation block is proposed to reduce the structural gaps between the two images. The intensity distribution rendering module then renders the deformed structure to an image with the target intensity distribution. Experimental results show that the proposed SUA method has the capability to transfer both intensity distribution and structural content between multiple pairs of datasets and is superior to prior arts in closing the gaps for improving segmentation. Tianyang Miller, Shaoming Zheng, Jun Cheng 0003, Xi Jia, Joseph Bartlett, Xinxing Cheng, Zhaowen Qiu, Huazhu Fu, Jiang Liu 0001, Ales Leonardis, Jinming Duan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Fourier-Net: Fast Image Registration with Band-Limited DeformationabstractUnsupervised image registration commonly adopts U-Net style networks to predict dense displacement fields in the full-resolution spatial domain. For high-resolution volumetric image data, this process is however resource-intensive and time-consuming. To tackle this problem, we propose the Fourier-Net, replacing the expansive path in a U-Net style network with a parameter-free model-driven decoder. Specifically, instead of our Fourier-Net learning to output a full-resolution displacement field in the spatial domain, we learn its low-dimensional representation in a band-limited Fourier domain. This representation is then decoded by our devised model-driven decoder (consisting of a zero padding layer and an inverse discrete Fourier transform layer) to the dense, full-resolution displacement field in the spatial domain. These changes allow our unsupervised Fourier-Net to contain fewer parameters and computational operations, resulting in faster inference speeds. Fourier-Net is then evaluated on two public 3D brain datasets against various state-of-the-art approaches. For example, when compared to a recent transformer-based method, named TransMorph, our Fourier-Net, which only uses 2.2% of its parameters and 6.66% of the multiply-add operations, achieves a 0.5% higher Dice score and an 11.48 times faster inference speed. Code is available at https://github.com/xi-jia/Fourier-Net. Xi Jia, Joseph Bartlett, Wei Chen 0092, Siyang Song, Tianyang Miller, Xinxing Cheng, Wenqi Lu 0001, Zhaowen Qiu, Jinming Duan 0001 |
AAAI | 1 |
| 2023 | UniFace: Unified Cross-Entropy Loss for Deep Face RecognitionabstractAs a widely used loss function in deep face recognition, the softmax loss cannot guarantee that the minimum positive sample-to-class similarity is larger than the maximum negative sample-to-class similarity. As a result, no unified threshold is available to separate positive sample-to-class pairs from negative sample-to-class pairs. To bridge this gap, we design a UCE (Unified Cross-Entropy) loss for face recognition model training, which is built on the vital constraint that all the positive sample-to-class similarities shall be larger than the negative ones. Our UCE loss can be integrated with margins for a further performance boost. The face recognition model trained with the proposed UCE loss, UniFace, was intensively evaluated using a number of popular public datasets like MFR, IJB-C, LFW, CFP-FP, AgeDB, and MegaFace. Experimental results show that our approach outperforms SOTA methods like SphereFace, CosFace, ArcFace, Partial FC, etc. Especially, till the submission of this work (Mar. 8, 2023), the proposed UniFace achieves the highest TAR@MR-All on the academic track of the MFR-ongoing challenge. $\color{Blue}{\mathbf{Code}}$ is publicly available. Jiancan Zhou, Xi Jia, Qiufu Li, LinLin Shen, Jinming Duan 0001 |
ICCV | 2 |
| 2023 | UniTSFace: Unified Threshold Integrated Sample-to-Sample Loss for Face RecognitionabstractSample-to-class-based face recognition models can not fully explore the cross-sample relationship among large amounts of facial images, while sample-to-sample-based models require sophisticated pairing processes for training. Furthermore, neither method satisfies the requirements of real-world face verification applications, which expect a unified threshold separating positive from negative facial pairs. In this paper, we propose a unified threshold integrated sample-to-sample based loss (USS loss), which features an explicit unified threshold for distinguishing positive from negative pairs. Inspired by our USS loss, we also derive the sample-to-sample based softmax and BCE losses, and discuss their relationship. Extensive evaluation on multiple benchmark datasets, including MFR, IJB-C, LFW, CFP-FP, AgeDB, and MegaFace, demonstrates that the proposed USS loss is highly efficient and can work seamlessly with sample-to-class-based losses. The embedded loss (USS and sample-to-class Softmax loss) overcomes the pitfalls of previous approaches and the trained facial model UniTSFace exhibits exceptional performance, outperforming state-of-the-art methods, such as CosFace, ArcFace, VPL, AnchorFace, and UNPG. Our code is available at https://github.com/CVI-SZU/UniTSFace. Qiufu Li, Xi Jia, Jiancan Zhou, LinLin Shen, Jinming Duan 0001 |
NeurIPS | 2 |
| 2023 | Arbitrary Order Total Variation for Deformable Image RegistrationabstractIn this work, we investigate image registration in a variational framework and focus on regularization generality and solver efficiency. We first propose a variational model combining the state-of-the-art sum of absolute differences (SAD) and a new arbitrary order total variation regularization term. The main advantage is that this variational model preserves discontinuities in the resultant deformation while being robust to outlier noise. It is however non-trivial to optimize the model due to its non-convexity, non-differentiabilities, and generality in the derivative order. To tackle these, we propose to first apply linearization to the problem to formulate a convex objective function and then break down the resultant convex optimization into several point-wise, closed-form subproblems using a fast, over-relaxed alternating direction method of multipliers (ADMM). With our proposed algorithm, we show that solving higher-order variational formulations is similar to solving their lower-order counterparts. Extensive experiments show that our ADMM is significantly more efficient than both the subgradient and primal-dual algorithms particularly when higher-order derivatives are used, and that our new models outperform state-of-the-art methods based on deep learning and free-form deformation. Our code implemented in both Matlab and Pytorch is publicly available at https://github.com/j-duan/AOTV. Jinming Duan 0001, Xi Jia, Joseph Bartlett, Wenqi Lu 0001, Zhaowen Qiu |
Pattern Recognit. | 2 |
| 2022 | Learning a Model-Driven Variational Network for Deformable Image RegistrationabstractData-driven deep learning approaches to image registration can be less accurate than conventional iterative approaches, especially when training data is limited. To address this issue and meanwhile retain the fast inference speed of deep learning, we propose VR-Net, a novel cascaded variational network for unsupervised deformable image registration. Using a variable splitting optimization scheme, we first convert the image registration problem, established in a generic variational framework, into two sub-problems, one with a point-wise, closed-form solution and the other one being a denoising problem. We then propose two neural layers (i.e. warping layer and intensity consistency layer) to model the analytical solution and a residual U-Net (termed generalized denoising layer) to formulate the denoising problem. Finally, we cascade the three neural layers multiple times to form our VR-Net. Extensive experiments on three (two 2D and one 3D) cardiac magnetic resonance imaging datasets show that VR-Net outperforms state-of-the-art deep learning methods on registration accuracy, whilst maintaining the fast inference speed of deep learning and the data-efficiency of variational models. Xi Jia, Alexander Thorley, Wei Chen 0092, Huaqi Qiu, LinLin Shen, Iain B. Styles, Hyung Jin Chang, Ales Leonardis, Antonio M. Simoes Monteiro de Marvao, Declan P. O'Regan, Daniel Rueckert, Jinming Duan 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2021 | FS-Net: Fast Shape-Based Network for Category-Level 6D Object Pose Estimation With Decoupled Rotation MechanismabstractIn this paper, we focus on category-level 6D pose and size estimation from a monocular RGB-D image. Previous methods suffer from inefficient category-level pose feature extraction, which leads to low accuracy and inference speed. To tackle this problem, we propose a fast shape-based network (FS-Net) with efficient category-level feature extraction for 6D pose estimation. First, we design an orientation aware autoencoder with 3D graph convolution for latent feature extraction. Thanks to the shift and scale-invariance properties of 3D graph convolution, the learned latent feature is insensitive to point shift and object size. Then, to efficiently decode category-level rotation information from the latent feature, we propose a novel decoupled rotation mechanism that employs two decoders to complementarily access the rotation information. For translation and size, we estimate them by two residuals: the difference between the mean of object points and ground truth translation, and the difference between the mean size of the category and ground truth size, respectively. Finally, to increase the generalization ability of the FS-Net, we propose an on-line box-cage based 3D deformation mechanism to augment the training data. Extensive experiments on two benchmark datasets show that the proposed method achieves state-of-the-art performance in both category- and instance-level 6D object pose estimation. Especially in category-level pose estimation, without extra synthetic data, our method outperforms existing methods by 6.3% on the NOCS-REAL dataset1. Wei Chen 0092, Xi Jia, Hyung Jin Chang, Jinming Duan 0001, LinLin Shen, Ales Leonardis |
CVPR | 2 |
| 2021 | Nesterov Accelerated ADMM for Fast Diffeomorphic Image Registration
Alexander Thorley, Xi Jia, Hyung Jin Chang, Karina Bunting, Victoria Stoll, Antonio M. Simoes Monteiro de Marvao, Declan P. O'Regan, Georgios V. Gkoutos, Dipak Kotecha, Jinming Duan 0001 |
MICCAI (4) | 2 |
| 2020 | G2L-Net: Global to Local Network for Real-Time 6D Pose Estimation With Embedding Vector FeaturesabstractIn this paper, we propose a novel real-time 6D object pose estimation framework, named G2L-Net. Our network operates on point clouds from RGB-D detection in a divide-and-conquer fashion. Specifically, our network consists of three steps. First, we extract the coarse object point cloud from the RGB-D image by 2D detection. Second, we feed the coarse object point cloud to a translation localization network to perform 3D segmentation and object translation prediction. Third, via the predicted segmentation and translation, we transfer the fine object point cloud into a local canonical coordinate, in which we train a rotation localization network to estimate initial object rotation. In the third step, we define point-wise embedding vector features to capture viewpoint-aware information. To calculate more accurate rotation, we adopt a rotation residual estimator to estimate the residual between initial rotation and ground truth, which can boost initial pose estimation performance. Our proposed G2L-Net is real-time despite the fact multiple steps are stacked via the proposed coarse-to-fine framework. Extensive experiments on two benchmark datasets show that G2L-Net achieves state-of-the-art performance in terms of both accuracy and speed. Wei Chen 0092, Xi Jia, Hyung Jin Chang, Jinming Duan 0001, Ales Leonardis |
CVPR | 2 |
| 2020 | Geometry Constrained Weakly Supervised Object Localization
Weizeng Lu, Xi Jia, Weicheng Xie 0001, LinLin Shen, Yicong Zhou, Jinming Duan 0001 |
ECCV (26) | 2 |
| 2020 | One-stage Multi-task Detector for 3D Cardiac MR ImagingabstractFast and accurate landmark location and bounding box detection are important steps in 3D medical imaging. In this paper, we propose a novel multi-task learning framework, for real-time, simultaneous landmark location and bounding box detection in 3D space. Our method extends the famous single-shot multibox detector (SSD) from single-task learning to multitask learning and from 2D to 3D. Furthermore, we propose a post-processing approach to refine the network landmark output, by averaging the candidate landmarks. Owing to these settings, the proposed framework is fast and accurate. For 3D cardiac magnetic resonance (MR) images with size 224×224×64, our framework runs ~128 volumes per second (VPS) on GPU and achieves 6.75mm average point-to-point distance error for landmark location, which outperforms both state-of-the-art and baseline methods. We also show that segmenting the 3D image cropped with the bounding box results in both improved performance and efficiency. Weizeng Lu, Xi Jia, Wei Chen 0092, Nicolò Savioli, Antonio M. Simoes Monteiro de Marvao, LinLin Shen, Declan P. O'Regan, Jinming Duan 0001 |
ICPR | 2 |
| 2020 | Car-Following Safe Headway Strategy with Battery-Health Conscious: A Reinforcement Learning ApproachabstractThis paper proposes an optimal car-following strategy for pure electric vehicles (EVs) with the aim of keeping an expected headway of the leader and reducing vehicle battery loss. In particular, a car-following system model is established. The primary task of the automatic vehicle is to follow the trajectory of the preceding car and maintain an expected headway. Then, the paper analyzes the powertrain of the electric vehicle. The loss of battery life over a period of time is proportional to the acceleration, so it takes the battery life into consideration. The Q-learning algorithm is conducted for the optimal car-following strategy using system data instead of system dynamics information. It utilizes reward function and greedy strategy to select actions to train the following vehicle to achieve car-following safety. When there is no collision in these two cars, acceleration is considered into reward function to reduce battery loss. Finally, it is verified by simulation that the proposed car-following strategy can keep good tracking, maintain the expected headway from the preceding vehicle, and reduce battery loss. Xi Jia, Jun Peng 0001, Yongjie Liu, Mengfei Wen, Zhiwu Huang |
SMC | 1 |
| 2019 | Local Normalization Based BN Layer Pruning
Xi Jia, LinLin Shen, Zhong Ming 0001, Jinming Duan 0001 |
ICANN (2) | 2 |
| 2019 | Adversarial Feature Distillation for Facial Expression Recognition
Mengchao Bai, Xi Jia, Weicheng Xie 0001, LinLin Shen |
PRICAI (3) | 2 |
| 2019 | Texture Deformation Based Generative Adversarial Networks for Multi-domain Face Editing
Wenting Chen, Xinpeng Xie, Xi Jia, LinLin Shen |
PRICAI (1) | 3 |
| 2019 | Sparse deep feature learning for facial expression recognition
Weicheng Xie 0001, Xi Jia, LinLin Shen, Meng Yang 0001 |
Pattern Recognit. | 2 |
| 2018 | Hand-Crafted Feature Guided Deep Learning for Facial Expression RecognitionabstractA number of facial expression recognition algorithms based on hand-crafted features and deep neutral networks have been developed. Motivated by the similarity between the hand-crafted features and features learned by deep network, a new feature loss is proposed to embed the information of hand-crafted features into the training process of network, which tries to reduce the difference between the two features. Based on the feature loss, a general framework for embedding the traditional feature information was developed and tested using CK+, JAFFE and FER2013 datasets. Experimental results show that the proposed network achieves much better accuracy than the original hand-crafted feature and the network without using our feature loss. When compared with other algorithms in literature, our network also achieved the best performance on CK+ dataset, i.e. 97.35% accuracy has been achieved. Guohang Zeng, Jiancan Zhou, Xi Jia, Weicheng Xie 0001, LinLin Shen |
FG | 3 |
| 2018 | Deep cross residual network for HEp-2 cell staining pattern classification
LinLin Shen, Xi Jia, Yuexiang Li |
Pattern Recognit. | 2 |
| 2016 | Deep convolutional neural network based HEp-2 cell classificationabstractAs different staining patterns of HEp-2 cells indicate different diseases, the classification of Indirect Immune Fluorescence (IIF) images on Human Epithelial-2 (HEp-2) cell is important for clinical applications. Different from traditional pattern recognition techniques, we use CNN to extract more high-level features for cell images classification. Compared to the existing CNN based HEp-2 classification methods, we proposed a network with deeper architecture. A class-balanced approach is also proposed to augment the HEp-2 cell dataset for network training. The proposed framework achieves an average class accuracy of 79.29% on ICPR 2012 HEp-2 dataset and a mean class accuracy of 98.26% on ICPR 2016 HEp-2 training set. Xi Jia, LinLin Shen, Xiande Zhou, Shiqi Yu 0001 |
ICPR | 1 |