Mingxuan Cai

dblp:209/8183 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0003-4011-8292ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PV-MLLM: A Generalized Intelligent Framework for Zero-Shot Photovoltaic Fault Diagnosis
abstract
Existing zero-shot fault diagnosis methods are typically system-specific and numerically sensitive, which lack adaptive deployment capabilities across heterogeneous photovoltaic (PV) system scales and topologies. Multimodal large language models (MLLMs) emerge as a powerful solution in cross-system generalization, but their adoption in PV fault diagnosis has been limited by the lack of PV knowledge integration and challenges in processing diverse operating conditions. To bridge this gap, an MLLMs-empowered framework for zero-shot PV fault diagnosis is proposed for the first time, which jointly integrates data-driven and knowledge-driven schemes. First, a chain-of-thought-based data augmentation pipeline is constructed to achieve data-knowledge alignment and interpretable results. Second, a two-stage adaptation strategy is specifically designed for PV data to overcome system scales, diverse topologies, and numerical differences. It consists of a Kolmogorov–Arnold networks-based condition adaptive layer embedded in vision transformer and a low-rank adaptation-based PV domain fine-tuning. Third, we design a microservices-based architecture for PV-MLLM deployment that enables flexible component decoupling and adaptive inference, significantly reducing hardware requirements and resource consumption. The proposed method achieves 99.66% and 97.25% diagnostic accuracy on simulated and real-world datasets.
Qi Liu 0014, Bo Yang 0006, Mengqi Han, Mingxuan Cai, Kai Ma 0001, Xin-Ping Guan
IEEE Trans. Ind. Informatics4
2025 Event2Audio: Event-Based Optical Vibration Sensing
abstract
Small vibrations observed in video can unveil information beyond what is visual, such as sound and material properties. It is possible to passively record these vibrations when they are visually perceptible, or actively amplify their visual contribution with a laser beam when they are not perceptible. In this paper, we improve upon the active sensing approach by leveraging event-based cameras, which are designed to efficiently capture fast motion. We demonstrate our method experimentally by recovering audio from vibrations, even for multiple simultaneous sources, and in the presence of environmental distortions. Our approach matches the state-of-the-art reconstruction quality at much faster speeds, approaching real-time processing.
Mingxuan Cai, Dekel Galor, Amit P. S. Kohli, Jacob L. Yates, Laura Waller
ICCP1
2025 A Collaborative Framework Based on MLLM for Generalization Photovoltaic Fault Diagnosis
abstract
Data-driven fault diagnosis methods for photovoltaic modules often encounter issues of sample scarcity and insufficient generalization, whereas knowledge-driven multimodal large language models (MLLMs) can comprehend human-summarized prior knowledge for logical reasoning, significantly enhancing the model’s general applicability. This paper proposes a high-generalization diagnostic framework based on knowledge-driven approaches, featuring a state correction vision transformer on MLLMs to solve the problem of sample scarcity, thereby enabling fault diagnosis grounded in photovoltaic knowledge. To mitigate the high computational cost of the large model inference, the framework incorporates an edge-based small model using Support Vector Machines to filter faulty samples. Additionally, a carefully designed photovoltaic knowledge datasets, along with an adaptive fine-tuning method, facilitates efficient domain-specific refinement of the pre-trained model. To address the potential issue of frequent model updates, each framework component is encapsulated as a microservice for easy orchestration. The framework has been deployed and tested on real cloud-edge-end cluster, achieving a diagnostic accuracy of 97.25% and reducing the invocation of large models by 90%.
Mengqi Han, Bo Yang 0006, Qi Liu 0014, Mingxuan Cai
IECON4
2025 Physics-Data Fusion for Long-Term Voltage Prediction in Vanadium Redox Flow Batteries
abstract
Voltage prediction is critical for ensuring both safety and operational efficiency of vanadium redox flow batteries (VRFBs) in long-duration energy storage. In this paper, we propose a physics-data fusion model framework for long-term voltage prediction of VRFBs. First, a parameterized Nernst equation is introduced to construct a physically meaningful latent space. Second, the temporal dynamics of hidden state variables are modeled based on electrochemical principles to precisely capture their time-dependent behavior. Subsequently, auxiliary variables are constructed using a physics-guided approach based on the polarization. Finally, a deep neural network is employed as the output layer of the model to establish the mapping between multidimensional feature variables and voltage. Experimental results demonstrate that this approach effectively captures the long-term voltage aging trends under diverse operational conditions, improving both accuracy and generalization performance.
Bo Yang 0006, Mingxuan Cai, Qi Liu 0014, Peng Wang 0029
INDIN3
2025 Funmap: integrating high-dimensional functional annotations to improve fine-mapping
abstract
MOTIVATION: Fine-mapping aims to prioritize causal variants underlying complex traits by accounting for the linkage disequilibrium of genome-wide association study risk locus. The expanding resources of functional annotations serve as auxiliary evidence to improve the power of fine-mapping. However, existing fine-mapping methods tend to generate many false positive results when integrating a large number of annotations. RESULTS: In this study, we propose a unified method to integrate high-dimensional functional annotations with fine-mapping (Funmap). Funmap can effectively improve the power of fine-mapping by borrowing information from hundreds of functional annotations. Meanwhile, it relates the annotation to the causal probability with a random effects model that avoids the over-fitting issue, thereby producing a well-controlled false positive rate. Paired with a fast algorithm, Funmap enables scalable integration of a large number of annotations to facilitate prioritizing multiple causal single nucleotide polymorphisms. Our comprehensive simulations across a wide range of annotation relevance settings demonstrate that Funmap is the only method that produces well-calibrated false discovery rate under the setting of high-dimensional annotations while achieving better or comparable power gains as compared to existing methods. By integrating genome-wide association studies of 4 lipid traits with 187 functional annotations, Funmap consistently identified more variants that can be replicated in an independent cohort, achieving 15.5%-26.2% improvement over the runner-up in terms of replication rate. AVAILABILITY AND IMPLEMENTATION: The Funmap software and all analysis code are available at https://github.com/LeeHITsz/Funmap.
Yuekai Li, Jiashun Xiao, Jingsi Ming, Yicheng Zeng, Mingxuan Cai
Bioinform.5
2025 Self-Correcting-Guided Generalized Contrastive Learning Framework for Small-Sample PV Fault Diagnosis With Cloud-Edge Collaboration
abstract
Intelligent fault diagnosis of photovoltaic (PV) arrays in small-sample scenarios remains challenging due to poor model accuracy and generalization. Existing methods fail to simultaneously address issues of varied operation conditions and insufficient samples, leading to the limited applicability of models built by few-shot learning. In addition, factors, such as data transmission and computation costs, also need to be considered. Therefore, this article proposes a cloud-edge collaborative self-correcting-guided generalized contrastive learning framework for small-sample PV fault diagnosis. First, an end-to-end self-correcting model is proposed to eliminate the influence of variable environments. Then, a self-correcting scheme is integrated with contrastive learning to achieve model generalization, and a type screening method is designed to improve model accuracy. Furthermore, a fast fault filtering mechanism is proposed to enhance the algorithm efficiency with cloud-edge collaboration. Both simulation and real data are utilized to validate the proposed method.
Qi Liu 0014, Bo Yang 0006, Mingxuan Cai, Kai Ma 0001, Xin-Ping Guan
IEEE Trans. Ind. Informatics3
2024 RL-SeqISP: Reinforcement Learning-Based Sequential Optimization for Image Signal Processing
abstract
Hardware image signal processing (ISP), aiming at converting RAW inputs to RGB images, consists of a series of processing blocks, each with multiple parameters. Traditionally, ISP parameters are manually tuned in isolation by imaging experts according to application-specific quality and performance metrics, which is time-consuming and biased towards human perception due to complex interaction with the output image. Since the relationship between any single parameter’s variation and the output performance metric is a complex, non-linear function, optimizing such a large number of ISP parameters is challenging. To address this challenge, we propose a novel Sequential ISP parameter optimization model, called the RL-SeqISP model, which utilizes deep reinforcement learning to jointly optimize all ISP parameters for a variety of imaging applications. Concretely, inspired by the sequential tuning process of human experts, the proposed model can progressively enhance image quality by seamlessly integrating information from both the image feature space and the parameter space. Furthermore, a dynamic parameter optimization module is introduced to avoid ISP parameters getting stuck into local optima, which is able to more effectively guarantee the optimal parameters resulting from the sequential learning strategy. These merits of the RL-SeqISP model as well as its high efficiency are substantiated by comprehensive experiments on a wide range of downstream tasks, including two visual analysis tasks (instance segmentation and object detection), and image quality assessment (IQA), as compared with representative methods both quantitatively and qualitatively. In particular, even using only 10% of the training data, our model outperforms other SOTA methods by an average of 7% mAP on two visual analysis tasks.
Zhikun Zhao, Congyan Lang, Mingxuan Cai, Longfei Han, Juan Wang 0012, Bing Li 0001
AAAI5
2023 PALM: a powerful and adaptive latent model for prioritizing risk variants with functional annotations
abstract
MOTIVATION: The findings from genome-wide association studies (GWASs) have greatly helped us to understand the genetic basis of human complex traits and diseases. Despite the tremendous progress, much effects are still needed to address several major challenges arising in GWAS. First, most GWAS hits are located in the non-coding region of human genome, and thus their biological functions largely remain unknown. Second, due to the polygenicity of human complex traits and diseases, many genetic risk variants with weak or moderate effects have not been identified yet. RESULTS: To address the above challenges, we propose a powerful and adaptive latent model (PALM) to integrate cell-type/tissue-specific functional annotations with GWAS summary statistics. Unlike existing methods, which are mainly based on linear models, PALM leverages a tree ensemble to adaptively characterize non-linear relationship between functional annotations and the association status of genetic variants. To make PALM scalable to millions of variants and hundreds of functional annotations, we develop a functional gradient-based expectation-maximization algorithm, to fit the tree-based non-linear model in a stable manner. Through comprehensive simulation studies, we show that PALM not only controls false discovery rate well, but also improves statistical power of identifying risk variants. We also apply PALM to integrate summary statistics of 30 GWASs with 127 cell type/tissue-specific functional annotations. The results indicate that PALM can identify more risk variants as well as rank the importance of functional annotations, yielding better interpretation of GWAS results. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/YangLabHKUST/PALM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jiashun Xiao, Mingxuan Cai, Yuling Jiao, Jin Liu 0011, Can Yang 0002
Bioinform.3
2022 XPXP: improving polygenic prediction by cross-population and cross-phenotype analysis
abstract
MOTIVATION: As increasing sample sizes from genome-wide association studies (GWASs), polygenic risk scores (PRSs) have shown great potential in personalized medicine with disease risk prediction, prevention and treatment. However, the PRS constructed using European samples becomes less accurate when it is applied to individuals from non-European populations. It is an urgent task to improve the accuracy of PRSs in under-represented populations, such as African populations and East Asian populations. RESULTS: In this article, we propose a cross-population and cross-phenotype (XPXP) method for construction of PRSs in under-represented populations. XPXP can construct accurate PRSs by leveraging biobank-scale datasets in European populations and multiple GWASs of genetically correlated phenotypes. XPXP also allows to incorporate population-specific and phenotype-specific effects, and thus further improves the accuracy of PRS. Through comprehensive simulation studies and real data analysis, we demonstrated that our XPXP outperformed existing PRS approaches. We showed that the height PRSs constructed by XPXP achieved 9% and 18% improvement over the runner-up method in terms of predicted R2 in East Asian and African populations, respectively. We also showed that XPXP substantially improved the stratification ability in identifying individuals at high genetic risk of type 2 diabetes. AVAILABILITY AND IMPLEMENTATION: The XPXP software and all analysis code are available at github.com/YangLabHKUST/XPXP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jiashun Xiao, Mingxuan Cai, Xianghong Hu 0002, Can Yang 0002
Bioinform.2
2018 LSMM: a statistical approach to integrating functional annotations with genome-wide association studies
abstract
Motivation: Thousands of risk variants underlying complex phenotypes (quantitative traits and diseases) have been identified in genome-wide association studies (GWAS). However, there are still two major challenges towards deepening our understanding of the genetic architectures of complex phenotypes. First, the majority of GWAS hits are in non-coding region and their biological interpretation is still unclear. Second, accumulating evidence from GWAS suggests the polygenicity of complex traits, i.e. a complex trait is often affected by many variants with small or moderate effects, whereas a large proportion of risk variants with small effects remain unknown. Results: The availability of functional annotation data enables us to address the above challenges. In this study, we propose a latent sparse mixed model (LSMM) to integrate functional annotations with GWAS data. Not only does it increase the statistical power of identifying risk variants, but also offers more biological insights by detecting relevant functional annotations. To allow LSMM scalable to millions of variants and hundreds of functional annotations, we developed an efficient variational expectation-maximization algorithm for model parameter estimation and statistical inference. We first conducted comprehensive simulation studies to evaluate the performance of LSMM. Then we applied it to analyze 30 GWAS of complex phenotypes integrated with nine genic category annotations and 127 cell-type specific functional annotations from the Roadmap project. The results demonstrate that our method possesses more statistical power than conventional methods, and can help researchers achieve deeper understanding of genetic architecture of these complex phenotypes. Availability and implementation: The LSMM software is available at https://github.com/mingjingsi/LSMM. Supplementary information: Supplementary data are available at Bioinformatics online.
Jingsi Ming, Mingwei Dai, Mingxuan Cai, Jin Liu 0011, Can Yang 0002
Bioinform.3
2017 IGESS: a statistical approach to integrating individual-level genotype data and summary statistics in genome-wide association studies
abstract
MOTIVATION: Results from genome-wide association studies (GWAS) suggest that a complex phenotype is often affected by many variants with small effects, known as 'polygenicity'. Tens of thousands of samples are often required to ensure statistical power of identifying these variants with small effects. However, it is often the case that a research group can only get approval for the access to individual-level genotype data with a limited sample size (e.g. a few hundreds or thousands). Meanwhile, summary statistics generated using single-variant-based analysis are becoming publicly available. The sample sizes associated with the summary statistics datasets are usually quite large. How to make the most efficient use of existing abundant data resources largely remains an open question. RESULTS: In this study, we propose a statistical approach, IGESS, to increasing statistical power of identifying risk variants and improving accuracy of risk prediction by i ntegrating individual level ge notype data and s ummary s tatistics. An efficient algorithm based on variational inference is developed to handle the genome-wide analysis. Through comprehensive simulation studies, we demonstrated the advantages of IGESS over the methods which take either individual-level data or summary statistics data as input. We applied IGESS to perform integrative analysis of Crohns Disease from WTCCC and summary statistics from other studies. IGESS was able to significantly increase the statistical power of identifying risk variants and improve the risk prediction accuracy from 63.2% ( ±0.4% ) to 69.4% ( ±0.1% ) using about 240 000 variants. AVAILABILITY AND IMPLEMENTATION: The IGESS software is available at https://github.com/daviddaigithub/IGESS . CONTACT: [email protected] or [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mingwei Dai, Jingsi Ming, Mingxuan Cai, Jin Liu 0011, Can Yang 0002, Zongben Xu
Bioinform.3