Jiaxin Xie

dblp:168/9631 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2026 MSTN-ISNet: A probability-guided spatio-temporal decoding for alleviating imbalance in sleep stage classification
Dongrui Gao, Shibing Li, Haokai Zhang, Zongyao Peng, Shihong Liu, Shaofei Ying, Jiaxin Xie
Appl. Intell.9
2025 VideoVAE+: Large Motion Video Autoencoding with Cross-Modal Video VAE
Yazhou Xing, Yingqing He, Jingye Chen, Jiaxin Xie, Xiaowei Chi, Qifeng Chen 0001
ICCV5
2025 Pan-LUT: Efficient Pan-sharpening via Learnable Look-Up Tables
abstract
Recently, deep learning-based pan-sharpening algorithms have achieved notable advancements over traditional methods. However, deep learning-based methods incur substantial computational overhead during inference, especially with large images. This excessive computational demand limits the applicability of these methods in real-world scenarios, particularly in the absence of dedicated computing devices such as GPUs and TPUs. To address these challenges, we propose Pan-LUT, a novel learnable look-up table (LUT) framework for pan-sharpening that strikes a balance between performance and computational efficiency for large remote sensing images. Our method makes it possible to process 15K$\times$15K remote sensing images on a 24GB GPU. To finely control the spectral transformation, we devise the PAN-guided look-up table (PGLUT) for channel-wise spectral mapping. To effectively capture fine-grained spatial details, we introduce the spatial details look-up table (SDLUT). Furthermore, to adaptively aggregate channel information for generating high-resolution multispectral images, we design an adaptive output look-up table (AOLUT). Our model contains fewer than 700K parameters and processes a 9K$\times$9K image in under 1 ms using one RTX 2080 Ti GPU, demonstrating significantly faster performance compared to other methods. Experiments reveal that Pan-LUT efficiently processes large remote sensing images in a lightweight manner, bridging the gap to real-world applications. Furthermore, our model surpasses SOTA methods in full-resolution scenes under real-world conditions, highlighting its effectiveness and efficiency. We also extend our method to general image fusion tasks.
Zhongnan Cai, Yingying Wang 0005, Hui Zheng 0003, Panwang Pan, Zixu Lin, Ge Meng, Chenxin Li, Chunming He, Jiaxin Xie, Yunlong Lin, Junbin Lu, Yue Huang 0001, Xinghao Ding
NeurIPS9
2025 Spatial-frequency dual-domain Kolmogorov-Arnold networks for multimodal medical image fusion
Lewu Lin, Jiaxin Xie, Yingying Wang 0005, Jialing Huang, Rongjin Zhuang, Xiaotong Tu, Xinghao Ding, Na Shen
Neurocomputing2
2024 C3T: Contrastive Consistency Cross-Network Learning for Semi-Supervised Semantic Segmentation
abstract
Semi-supervised image semantic segmentation, a vital but challenging task in multimedia applications, aims to accurately classify pixels with limited labeled data. Traditional approaches in this domain often grapple with the confirmation bias problem, where models, influenced by their own predictions, become prone to replicating errors. To address this critical issue, our research introduces a cross-network-crossview consistency learning framework. This novel paradigm significantly reduce the confirmation bias through diversifying the learning perspectives. Integral to our approach are two components: a pseudo-label validation and filtering mechanism, and a cross-contrastive learning module within the feature domain. These elements work in synergy to not only amplify the accuracy of the model but also its robustness against varied data scenarios. Extensive evaluations, conducted across multiple datasets, clearly demonstrate the effectiveness of our method. In comparison to existing state-of-the-art models, our approach exhibits marked improvements, especially in the challenging contexts of semisupervised image semantic segmentation. The code is available at https://github.com/Sstar2orchid/C3T.
Yucheng Shu, Jiaxin Xie, Lihong Qiao, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001
ICME2
2024 Robust Principal Component Analysis Based on Fuzzy Local Information Reservation
abstract
Principal Component Analysis (PCA) aims to acquire the principal component space containing the essential structure of data, instead of being used for mining and extracting the essential structure of data. In other words, the principal component space contains not only information related to the essential structure of data but also some unrelated information. This frequently occurs when the intrinsic dimensionality of data is unknown or when it has complex distribution characteristics such as multi-modalities, manifolds, etc. Therefore, it is unreasonable to identify noise and useful information based solely on reconstruction error. For this reason, PCA is unsuitable as a preprocessing technique for most applications, especially in noisy environment. To solve this problem, this paper proposes robust PCA based on fuzzy local information reservation (FLIPCA). By analyzing the impact of reconstruction error on sample discriminability, FLIPCA provides a theoretical basis for noise identification and processing. This not only greatly improves its robustness but also extends its applicability and effectiveness as a data preprocessing technique. Meanwhile, FLIPCA maintains consistent mathematical descriptions with traditional PCA while having few adjustable hyperparameters and low algorithmic complexity. Finally, we conducted comprehensive experiments on synthetic and real-world datasets, which substantiated the superiority of our proposed algorithm.
Yunlong Gao 0001, Xinjing Wang, Jiaxin Xie, Peng Yan 0006, Feiping Nie 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 High-fidelity 3D GAN Inversion by Pseudo-multi-view Optimization
abstract
We present a high-fidelity 3D generative adversarial network (GAN) inversion framework that can synthesize photorealistic novel views while preserving specific details of the input image. High-fidelity 3D GAN inversion is inherently challenging due to the geometry-texture trade-off, where overfitting to a single view input image often damages the estimated geometry during the latent optimization. To solve this challenge, we propose a novel pipeline that builds on the pseudo-multi-view estimation with visibility analysis. We keep the original textures for the visible parts and utilize generative priors for the occluded parts. Extensive experiments show that our approach achieves advantageous reconstruction and novel view synthesis quality over prior work, even for images with out-of-distribution textures. The proposed pipeline also enables image attribute editing with the inverted latent code and 3D-aware texture modification. Our approach enables high-fidelity 3D rendering from a single image, which is promising for various applications of AI-generated 3D content. The source code is at https://github.com/jiaxinxie97/HFGI3D/.
Jiaxin Xie, Hao Ouyang, Jingtan Piao, Chenyang Lei, Qifeng Chen 0001
CVPR1
2022 Shape from Polarization for Complex Scenes in the Wild
abstract
We present a new data-driven approach with physics based priors to scene-level normal estimation from a single polarization image. Existing shape from polarization (SfP) works mainly focus on estimating the normal of a single object rather than complex scenes in the wild. A key barrier to high-quality scene-level SfP is the lack of real-world SfP data in complex scenes. Hence, we contribute the first real world scene-level SfP dataset with paired input polarization images and ground-truth normal maps. Then we propose a learning-based framework with a multi-head self-attention module and viewing encoding, which is designed to handle increasing polarization ambiguities caused by complex materials and non-orthographic projection in scene-level SfP. Our trained model can be generalized to far-field outdoor scenes as the relationship between polarized light and surface normals is not affected by distance. Experimental results demonstrate that our approach significantly outperforms existing SfP models on two datasets. Our dataset and source code will be publicly available at https://github.com/ChenyangLEI/sfp-wild.
Chenyang Lei, Jiaxin Xie, Na Fan 0002, Vladlen Koltun, Qifeng Chen 0001
CVPR3
2022 A new robust fuzzy c-means clustering method based on adaptive elastic distance
Yunlong Gao 0001, Jiaxin Xie
Knowl. Based Syst.3
2021 Dual-Camera Super-Resolution with Aligned Attention Modules
abstract
We present a novel approach to reference-based super-resolution (RefSR) with the focus on dual-camera super-resolution (DCSR), which utilizes reference images for high-quality and high-fidelity results. Our proposed method generalizes the standard patch-based feature matching with spatial alignment operations. We further explore the dual-camera super-resolution that is one promising application of RefSR, and build a dataset that consists of 146 image pairs from the main and telephoto cameras in a smartphone. To bridge the domain gaps between real-world images and the training images, we propose a self-supervised domain adaptation strategy for real-world images. Extensive experiments on our dataset and a public benchmark demonstrate clear improvement achieved by our method over state of the art in both quantitative evaluation and visual comparisons. Our code and data are available at https://tengfei-wang.github.io/Dual-Camera-SR/index.html.
Tengfei Wang 0002, Jiaxin Xie, Wenxiu Sun, Qiong Yan, Qifeng Chen 0001
ICCV2
2020 Depth Sensing Beyond LiDAR Range
abstract
Depth sensing is a critical component of autonomous driving technologies, but today's LiDAR- or stereo camera- based solutions have limited range. We seek to increase the maximum range of self-driving vehicles' depth perception modules for the sake of better safety. To that end, we propose a novel three-camera system that utilizes small field of view cameras. Our system, along with our novel algorithm for computing metric depth, does not require full pre-calibration and can output dense depth maps with practically acceptable accuracy for scenes and objects at long distances not well covered by most commercial LiDARs.
Kai Zhang 0045, Jiaxin Xie, Noah Snavely, Qifeng Chen 0001
CVPR2
2020 Video Depth Estimation by Fusing Flow-to-Depth Proposals
abstract
Depth from a monocular video can enable billions of devices and robots with a single camera to see the world in 3D. In this paper, we present a model for video depth estimation, which consists of a flow-to-depth layer, a camera pose refinement module, and a depth fusion network. Given optical flow and camera poses, our flow-to-depth layer generates depth proposals and their corresponding confidence maps by explicitly solving an epipolar geometry optimization problem. Our flow-to-depth layer is differentiable, and thus we can refine camera poses by maximizing the aggregated confidence in the camera pose refinement module. Our depth fusion network can utilize the target frame, depth proposals, and confidence maps inferred from different neighboring frames to produce the final depth map. Furthermore, the depth fusion network can additionally take the depth proposals generated by other methods to further improve the results. The experiments on three public datasets show that our approach outperforms state-of-the-art depth estimation methods, and has reasonable crossdataset generalization ability: our model trained on KITTI still performs well on the unseen Waymo dataset.
Jiaxin Xie, Chenyang Lei, Zhuwen Li, Li Erran Li, Qifeng Chen 0001
IROS1
2020 A symmetric alternating minimization algorithm for total variation minimization
Jiaxin Xie
Signal Process.2
2018 A new accelerated alternating minimization method for analysis sparse recovery
Jiaxin Xie, An-ping Liao
Signal Process.1