VLDB 2026 Research / reviewers in the wild / expert
Jian-Xun Mi
dblp:03/6206
· DBLP profile ↗
55ranked-venue papers
37as first author
31since 2021 · last 2026
0000-0002-7531-4341ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 18 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 13 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Interpretable object detection via integrated heatmap, concept attribution, and sobol sensitivity analysis
Muhammad Imran Khalid, Jian-Xun Mi, Ghulam Ali, Tariq Ali, Mohammad Hijji, Muhammad Ayaz, Zia-ur-Rehman |
Inf. Sci. | 2 |
| 2026 | EEL-RIS: Explain everything on localization via randomized input sampling
Jian-Xun Mi, Shenglei Shi, Lin Luo 0008, Weisheng Li 0001 |
Inf. Sci. | 1 |
| 2026 | Decomposing decisions: causal concept explanations for deep learning modelsabstractThe deployment of deep learning models in high-stakes applications such as autonomous driving is critically important, yet their black-box nature remains a fundamental barrier to trust and accountability. Existing explainability methods typically produce ambiguous, pixel-based heatmaps that capture correlation rather than establishing a causal link between high-level, human-interpretable concepts and model outputs. This paper introduces the Causal Concept Decomposer (CCD), a three-stage framework for concept-driven causal explanation of object detectors. CCD first employs semantic segmentation to isolate the target object, then applies Non-negative Matrix Factorization to discover constituent semantic parts, and finally uses Sobol sensitivity analysis to quantify the causal influence of each part on the detector’s decision. Evaluated on the MS COCO dataset, CCD produces explanations that are both visually coherent and quantitatively more faithful than existing approaches, achieving a Deletion AUC of 0.11 and an Insertion AUC of 0.91. By moving beyond correlational attribution towards principled causal analysis, this work represents an important step towards more trustworthy and reliable AI systems. Muhammad Imran Khalid, Jian-Xun Mi |
J. Exp. Theor. Artif. Intell. | 2 |
| 2026 | IGCDet: Independence guided co-training for sparsely annotated object detection
Jian-Xun Mi, Ranzhi Zhao |
Knowl. Based Syst. | 1 |
| 2026 | Multi-view subspace clustering via tensor nuclear norm factorization
Tinghe Yan, Qiang Guo 0003, Jian-Xun Mi, Weisheng Li 0001 |
Pattern Recognit. | 3 |
| 2026 | Adaptive ensemble attack: breaking Large Multimodal Models via dynamic caption selection and weighted gradientsabstractAbstract Large Multimodal Models (LMMs) have achieved remarkable performance across vision-language tasks, yet their robustness against adversarial attacks remains critically underexplored. While LMMs are vulnerable to visual encoder attacks, they exhibit surprising resilience due to encoder diversity—attacks optimized for CLIP fail to transfer to EVA-CLIP, especially when textual context is provided. We introduce the Adaptive Ensemble PGD (AE-PGD) attack, which simultaneously targets both encoders through three key innovations: (1) dynamic adversarial caption selection , combining gradient magnitude with global semantic displacement to identify the most attack-effective caption per model; (2) an adaptive weight controller , dynamically balancing each encoder’s contribution using real-time loss, gradient norm, and confidence metrics; and (3) an Expectation over Transforms (EoT) gradient update ensuring robustness against input-transformation defenses. Evaluated on COCO 2014 images, AE-PGD reduces accuracy from a 75.42% baseline to 0.0% across all three evaluation metrics—visual encoding, image-to-text recall, and LLM answer recall—achieving complete model collapse. Manifold analysis confirms that adversarial perturbations push image embeddings to antipodal regions of the joint embedding space, activating semantically opposite concept clusters and producing structured hallucinations. WordNet WUP similarity analysis reveals a 33.5 percentage point semantic drop across the test set. AE-PGD causes state-of-the-art LMMs (LLaVA, Qwen-VL, GPT-4V) to catastrophically misidentify a bullet train as a “helicopter crash,” with strong black-box transfer yielding a 65 percentage point recall collapse on unseen encoders. This work exposes critical vulnerabilities in current LMM architectures and underscores the urgent need for ensemble-aware defense mechanisms. Sudhir Kumar Pandey, Jian-Xun Mi, Muhammad Salman Pathan |
Vis. Comput. | 2 |
| 2025 | Learning Discriminative Features with VAE-GAN for Zero-Shot Learning
Jian-Xun Mi, Shenglei Shi |
ICIC (9) | 1 |
| 2025 | StealthMask: Highly stealthy adversarial attack on face recognition system
Jian-Xun Mi, Mingxuan Chen |
Appl. Intell. | 1 |
| 2025 | Rethinking the CNN and transformer for deformable image registration
Weisheng Li 0001, Yucheng Shu, Jian-Xun Mi, Guofen Wang, Bin Xiao 0002 |
Expert Syst. Appl. | 4 |
| 2025 | Improving the sparse coding model via hybrid Gaussian priors
Jian-Xun Mi, Weisheng Li 0001, Guofen Wang, Bin Xiao 0002 |
Pattern Recognit. | 2 |
| 2025 | Directing model attention to the discriminative foreground features
Jian-Xun Mi, Shenglei Shi |
J. Supercomput. | 1 |
| 2024 | Window-Based Convolutional Sparse Coding: Towards A Unified FrameworkabstractSparse Coding (SC) and Convolution Sparse Coding (CSC) are two widely studied sparse methods in computer vision and signal processing. SC encodes the image patches independently, however fails to utilize the correlation among them. CSC adopts a convolution operator to connect the overlapping patches but in an inflexible manner. In this paper, a novel integrated framework for the two sparse models is proposed, wherein the local correlations among patches are controllable by manipulating a window function. Moreover, the inherent border effect of a convolution model is mitigated with a carefully designed weight function. It can be demonstrated that both SC and CSC are two distinct implementations of this framework. Consequently, our unified framework provides a balanced solution by addressing the strengths and limitations of both SC and CSC. Extensive experimental results are presented to demonstrate the superiority and effectiveness of the proposed method for image inpainting tasks. Jian-Xun Mi, Guofen Wang, Weisheng Li 0001 |
ICASSP | 2 |
| 2024 | Non-targeted Adversarial Attacks on Object Detection Models
Jian-Xun Mi, Xiangjin Zhao, Yongtao Chen, Xiaohong Lv, Jiayong Zhong |
ICIC (9) | 1 |
| 2024 | ShiftMorph: A Fast and Robust Convolutional Neural Network for 3D Deformable Medical Image Registration
Weisheng Li 0001, Yucheng Shu, Jian-Xun Mi, Bin Xiao 0002 |
ACM Multimedia | 4 |
| 2024 | Toward explainable artificial intelligence: A survey and overview on their intrinsic properties
Jian-Xun Mi, Xilai Jiang, Lin Luo 0008 |
Neurocomputing | 1 |
| 2024 | An elastic competitive and discriminative collaborative representation method for image classification
Jian-Xun Mi, Shijie Yin, Weisheng Li 0001 |
Neural Networks | 1 |
| 2024 | Deep Cross-View Reconstruction GAN Based on Correlated Subspace for Multi-View TransformationabstractIn scenarios where identifying face information in the visible spectrum (VIS) is challenging due to poor lighting conditions, the use of near-infrared (NIR) and thermal (TH) cameras can provide viable alternatives. However, the unique data distribution of images captured by these cameras compared to VIS images presents challenges in matching face identities. To address these challenges, we propose a novel image transformation framework. The framework includes feature extraction from the input image, followed by a transformation network that generates target domain images with perceptual fidelity. Additionally, a reconstruction network preserves original information by reconstructing the original domain image from the extracted features. By considering the correlation between features from both domains, our framework utilizes paired data obtained from the same individual. We apply this framework to two well-established image-to-image transformation models, pix2pix and CycleGAN, known as CRC-pix2pix and CRC-CycleGAN respectively. The versatility of our approach allows extension to other models based on pix2pix or CycleGAN architectures. Our models generate high-quality images while preserving the identity information of the original face. Performance evaluation on TFW and BUAA NIR-VIS datasets demonstrates the superiority of our models in terms of generated image face matching and evaluation metrics such as SSIM, MSE, PSNR, and LPIPS. Moreover, we introduce the CQUPT-VIS-TH dataset, which enriches the paired dataset with thermal-visual face data capturing various angles and expressions. Jian-Xun Mi, Junchang He, Weisheng Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | A Multi-granularity Decision Fusion Method Based on Category Hierarchy
Jian-Xun Mi, Ke-Yang Huang |
ICIC (2) | 1 |
| 2023 | Attribute self-representation steered by exclusive lasso for zero-shot learning
Jian-Xun Mi, Debao Tai, Li-Fang Zhou |
Appl. Intell. | 1 |
| 2023 | Adversarial examples based on object detection tasks: A survey
Jian-Xun Mi, Xu-Dong Wang, Lifang Zhou |
Neurocomputing | 1 |
| 2023 | Hierarchical neural network with efficient selection inference
Jian-Xun Mi, Ke-Yang Huang, Weisheng Li 0001, Lifang Zhou |
Neural Networks | 1 |
| 2023 | Accurate and Robust Eye Center Localization by Deep VotingabstractEye Center Localization (ECL) is one of the most crucial technologies for various computer vision applications, such as eye gazing estimation and eye-tracking. Current conventional implementations consist of two phases, including locating the approximate eye regions and finding the eye center position by extracting the semantic features around the corresponding eye region. However, the combination pipeline results in the ECL accuracy being influenced by not only the environmental factors, such as the variability of photographing angles, illuminations, and the probable occlusions by eyelids or glasses, but also the quality of preceding procedures. Inspired by the ensemble mechanism in machine learning, we formulate ECL problem as a process of end-to-end voting, and the core is to select a set of local descriptors which can capture efficient independent information to vote for eye centers. With the help of deep convolutional neural networks, we are able to determine semantic descriptors around the eye regions. Each descriptor proposes a vote pointing to the corresponding eye center, and all the votes indicate the eye centers finally. The experimental results on the public databases, BioID and GI4E, show that our method achieves 80.3% and 95.2% accuracy, respectively, which outperforms the existing state-of-the-art methods, and the results based on our customized challenging database verify the robustness of our method. Jian-Xun Mi, Shiyao Yuan, Weisheng Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Inverse Sparse Object Tracking via Adaptive Representation
Jian-Xun Mi |
ICIC (3) | 1 |
| 2022 | Designing efficient convolutional neural network structure: A survey
Jian-Xun Mi, Jie Feng 0011, Ke-Yang Huang |
Neurocomputing | 1 |
| 2022 | Deep cross-view autoencoder network for multi-view learning
Jian-Xun Mi, Chang-Qing Fu, Tingting Gou |
Multim. Tools Appl. | 1 |
| 2022 | Local spatial continuity steered sparse representation for occluded face recognition
Jian-Xun Mi, Lifang Zhou |
Multim. Tools Appl. | 1 |
| 2021 | Exploring Multi-scale Temporal and Spectral CSP Feature for Multi-class Motion Imagination Task Classification
Jian-Xun Mi, Rongfeng Li 0006 |
ICIC (3) | 1 |
| 2021 | A Robust and Automatic Recognition Method of Pointer Instruments in Power System
Jian-Xun Mi, Xu-Dong Wang, Qing-Yun Yang, Xin Deng 0003 |
ICIC (1) | 1 |
| 2021 | Symmetrical feature extraction via novel Mirror PCA
Jian-Xun Mi, Lifang Zhou, Yueru Sun, Heng Kong |
Neurocomputing | 1 |
| 2021 | IoU-guided Siamese region proposal network for real-time visual tracking
Lifang Zhou, Weisheng Li 0001, Jian-Xun Mi, Bang Jun Lei |
Neurocomputing | 4 |
| 2021 | A Lightweight SE-YOLOv3 Network for Multi-Scale Object Detection in Remote Sensing ImageryabstractCurrent state-of-the-art detectors achieved impressive performance in detection accuracy with the use of deep learning. However, most of such detectors cannot detect objects in real time due to heavy computational cost, which limits their wide application. Although some one-stage detectors are designed to accelerate the detection speed, it is still not satisfied for task in high-resolution remote sensing images. To address this problem, a lightweight one-stage approach based on YOLOv3 is proposed in this paper, which is named Squeeze-and-Excitation YOLOv3 (SE-YOLOv3). The proposed algorithm maintains high efficiency and effectiveness simultaneously. With an aim to reduce the number of parameters and increase the ability of feature description, two customized modules, lightweight feature extraction and attention-aware feature augmentation, are embedded by utilizing global information and suppressing redundancy features, respectively. To meet the scale invariance, a spatial pyramid pooling method is used to aggregate local features. The evaluation experiments on two remote sensing image data sets, DOTA and NWPU VHR-10, reveal that the proposed approach achieves more competitive detection effect with less computational consumption. Lifang Zhou, Guang Deng, Weisheng Li 0001, Jian-Xun Mi, Bang Jun Lei |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2020 | SharedNet: A Novel Efficient Convolutional Architecture Based on Group Sharing Convolution
Jian-Xun Mi, Jie Feng 0011 |
ICIC (1) | 1 |
| 2019 | Mirror PCA: Exploiting Facial Symmetry for Feature Extraction
Jian-Xun Mi, Yueru Sun |
ICIC (1) | 1 |
| 2019 | Principal component analysis based on block-norm minimization
Jian-Xun Mi, Quanwei Zhu |
Appl. Intell. | 1 |
| 2019 | Bilateral structure based matrix regression classification for face recognition
Jian-Xun Mi, Zhiheng Luo, Lifang Zhou, Fujin Zhong |
Neurocomputing | 1 |
| 2019 | Principal Component Analysis based on Nuclear norm Minimization
Jian-Xun Mi, Zhihui Lai 0001, Weisheng Li 0001, Lifang Zhou, Fujin Zhong |
Neural Networks | 1 |
| 2019 | Matrix regression-based classification with block-norm
Jian-Xun Mi, Quanwei Zhu, Zhiheng Luo |
Pattern Recognit. Lett. | 1 |
| 2018 | Robust Face Recognition Based on Supervised Sparse Representation
Jian-Xun Mi, Yueru Sun |
ICIC (3) | 1 |
| 2018 | Sparse factorial code representation using independent component analysis for face recognition
Jian-Xun Mi |
Multim. Tools Appl. | 2 |
| 2018 | Pose-robust face recognition with Huffman-LBP enhanced by Divide-and-Rule strategy
Lifang Zhou, Yue-Wei Du, Weisheng Li 0001, Jian-Xun Mi, Xiao Luan |
Pattern Recognit. | 4 |
| 2017 | Adaptive Class Preserving Representation for Image ClassificationabstractIn linear representation-based image classification, an unlabeled sample is represented by the entire training set. To obtain a stable and discriminative solution, regularization on the vector of representation coefficients is necessary. For example, the representation in sparse representation-based classification (SRC) uses L1 norm penalty as regularization, which is equal to lasso. However, lasso overemphasizes the role of sparseness while ignoring the inherent structure among samples belonging to a same class. Many recent developed representation classifications have adopted lasso-type regressions to improve the performance. In this paper, we propose the adaptive class preserving representation for classification (ACPRC). Our method is related to group lasso based classification but different in two key points: When training samples in a class are uncorrelated, ACPRC turns into SRC, when samples in a class are highly correlated, it obtains similar result as group lasso. The superiority of ACPRC over other state-of-the-art regularization techniques including lasso, group lasso, sparse group lasso, etc. are evaluated by extensive experiments. Jian-Xun Mi, Qiankun Fu, Weisheng Li 0001 |
CVPR | 1 |
| 2016 | Extraction of Independent Components from Sparse Mixture
Jian-Xun Mi |
ICIC (2) | 1 |
| 2016 | Multi-step linear representation-based classification for face recognitionabstractError detection is an important approach to improve the robustness of face recognition method. However, it is hard to directly detect the invalid pixels in a facial image. The authors decompose the hard problem into many simpler sub‐problems in this study. That is, the error detection process of pixels is divided into multiple phases and a portion of invalid pixels are detected in each phase. The goal is to decrease the ratio of invalid pixels to the whole pixels in a testing image, which progressively improves the final recognition accuracy. The performance that their method deals with occlusion and corruption problems is evaluated on different databases. In addition, the comparison with other state‐of‐the‐art studies shows that the proposed method achieves the best results in face occlusion and disguise issues. Jian-Xun Mi |
IET Comput. Vis. | 1 |
| 2016 | Robust face recognition via sparse boosting representation
Jian-Xun Mi |
Neurocomputing | 2 |
| 2014 | A comparative study and improvement of two ICA using reference signal methods
Jian-Xun Mi, Yong Xu 0001 |
Neurocomputing | 1 |
| 2013 | Robust PCA based method for discovering differentially expressed genesabstractHow to identify a set of genes that are relevant to a key biological process is an important issue in current molecular biology. In this paper, we propose a novel method to discover differentially expressed genes based on robust principal component analysis (RPCA). In our method, we treat the differentially and non-differentially expressed genes as perturbation signals S and low-rank matrix A, respectively. Perturbation signals S can be recovered from the gene expression data by using RPCA. To discover the differentially expressed genes associated with special biological progresses or functions, the scheme is given as follows. Firstly, the matrix D of expression data is decomposed into two adding matrices A and S by using RPCA. Secondly, the differentially expressed genes are identified based on matrix S. Finally, the differentially expressed genes are evaluated by the tools based on Gene Ontology. A larger number of experiments on hypothetical and real gene expression data are also provided and the experimental results show that our method is efficient and effective. Jin-Xing Liu 0001, Chun-Hou Zheng 0001, Wen Sha, Jian-Xun Mi, Yong Xu 0001 |
BMC Bioinform. | 5 |
| 2013 | Noisy component extraction with reference
Jian-Xun Mi |
Frontiers Comput. Sci. | 3 |
| 2013 | The nearest-farthest subspace classification for face recognition
Jian-Xun Mi, De-Shuang Huang, Bing Wang 0004, Xingjie Zhu |
Neurocomputing | 1 |
| 2013 | Using the idea of the sparse representation to perform coarse-to-fine face recognition
Yong Xu 0001, Qi Zhu 0001, Zizhu Fan, David Zhang 0001, Jian-Xun Mi, Zhihui Lai 0001 |
Inf. Sci. | 5 |
| 2012 | A Comparative Study of Two Independent Component Analysis Using Reference Signal Methods
Jian-Xun Mi, Yanxin Yang |
ICIC (3) | 1 |
| 2012 | Identifying Characteristic Genes Based on Robust Principal Component Analysis
Chun-Hou Zheng 0001, Jin-Xing Liu 0001, Jian-Xun Mi, Yong Xu 0001 |
ICIC (3) | 3 |
| 2011 | A New Subspace Approach for Face Recognition
Jian-Xun Mi |
ICIC (3) | 1 |
| 2010 | A Method for ICA with Reference Signals
Jian-Xun Mi, Jie Gui |
ICIC (2) | 1 |
| 2007 | A New Constrained Independent Component Analysis MethodabstractConstrained independent component analysis (cICA) is a general framework to incorporate a priori information from problem into the negentropy contrast function as constrained terms to form an augmented Lagrangian function. In this letter, a new improved algorithm for cICA is presented through the investigation of the inequality constraints, in which different closeness measurements are compared. The utility of our proposed algorithm is demonstrated by the experiments with synthetic data and electroencephalogram (EEG) data. De-Shuang Huang, Jian-Xun Mi |
IEEE Trans. Neural Networks | 2 |
| 2004 | Image compression using principal component neural networkabstractThis paper presents a comparison of three kinds of principal component neural networks, which are used for image compression. Principal component analysis (PCA), which is a statistical processing technique, is used in many engineering and scientific fields. Application of computing principal components includes data compression, pattern recognition and signal processing, etc. In this paper we use it for image compression. Jian-Xun Mi, De-Shuang Huang |
ICARCV | 1 |