EDBT 2026 Demo / reviewers in the wild / expert
Yuhao Xia
dblp:307/2081
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | M-DeepAssembly: enhanced DeepAssembly based on multi-objective multi-domain protein conformation samplingabstractBACKGROUND: Association and cooperation among structural domains play an important role in protein function and drug design. Despite remarkable advancements in highly accurate single-domain protein structure prediction through the collaborative efforts of the community using deep learning, challenges still exist in predicting multi-domain protein structures when the evolutionary signal for a given domain pair is weak or the protein structure is large. RESULTS: To alleviate the above challenges, we proposed M-DeepAssembly, a protocol based on multi-objective protein conformation sampling algorithm for multi-domain protein structure prediction. Firstly, the inter-domain interactions and full-length sequence distance features are extracted through DeepAssembly and AlphaFold2, respectively. Secondly, subject to these features, we constructed a multi-objective energy model and designed a sampling algorithm for exploring and exploiting conformational space to generate ensembles. Finally, the output protein structure was selected from the ensembles using our in-house developed model quality assessment algorithm. On the test set of 164 multi-domain proteins, the results show that the average TM-score of M-DeepAssembly is 15.4% and 2.0% higher than AlphaFold2 and DeepAssembly, respectively. It is worth noting that there are models with higher accuracy in ensembles, achieving an improvement of 20.3% and 6.4% relative to the two baseline methods, although these models were not selected. Furthermore, when compared to the prediction results of AlphaFold2 for CASP15 multi-domain targets, M-DeepAssembly demonstrates certain performance advantages. CONCLUSIONS: M-DeepAssembly provides a distinctive multi-domain protein assembly algorithm, which can alleviate the current challenges of weak evolutionary signals and large structures to some extent by forming diverse ensembles using multi-objective protein conformation sampling algorithm. The proposed method contributes to exploring the functions of multi-domain proteins, especially providing new insights into targets with multiple conformational states. Yuhao Xia, Minghua Hou, Xuanfeng Zhao, Suhui Wang, Guijun Zhang |
BMC Bioinform. | 2 |
| 2025 | A Novel EAGLe Framework for Robust UAV-View Geo-LocalizationabstractThis paper addresses the UAV-view geo-localization task, which focuses on bi-directional retrieval between UAV-view and satellite-view images. Generally, existing methods aim to learn image representations that can distinguish between different locations while effectively mitigating the cross-view domain gap. However, these methods often struggle in noisy UAV flight environments, as they fail to account for environmental domain shifts caused by varying weather and lighting conditions. To this end, we propose a novel Environment-Agnostic Geo-Localization (EAGLe) framework, which integrates a dual-objective discriminator and a style mixture module into diverse UAV-view geo-localization networks to enhance their robustness in dynamic environments. Specifically, the dual-objective discriminator not only distinguishes between UAV and satellite views but also identifies various environmental styles in UAV-view images. Through adversarial learning, the dual-objective discriminator encourages the feature encoder to produce features that remain invariant to both viewpoint and environmental variations. Furthermore, the style mixture module is integrated into the feature encoder to extend diversity at the feature level, allowing EAGLe to learn a broader range of environmental styles beyond the training data. Extensive experiments on the University-1652 and SUES-200 datasets demonstrate that the proposed EAGLe significantly improves the reliability of UAV-view geo-localization networks under dynamic and unpredictable environmental conditions, while maintaining inference efficiency. Cuiwei Liu, Shiting Peng, Shishen Li, Huaijun Qiu, Yuhao Xia, Zhaokui Li, Liang Zhao 0004 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Effect of Overlapping Degree and Distribution of High-Accuracy Images on Combined RFM-Based Geometric Positioning of Multi-Resolution Satellite ImageryabstractThis study investigated the effect of the distribution of high-accuracy images and the overlapping degree on combined geometric positioning by further experiments using Geoeye-1 and ZY-3 satellite images. The experimental results revealed that: (a) For combined positioning, reference stereo imagery with different degrees of overlap have different geometric positioning capabilities, especially in the planar direction. Moreover, the overall disparities gradually decrease with the increase of the overlapping degree for the reference stereo images, and (b) the geometric positioning accuracy can be improved by adding more high-accuracy images as references, and the diagonally distributed reference images are more helpful than the single-cornered reference image in improving the combined positioning accuracy. In future work, more and wider range of images would be used for the further validation. Wenping Song, Shijie Liu 0001, Xiaohua Tong, Yuhao Xia |
IGARSS | 5 |
| 2024 | Mixed Pixel Spatial Unmixing with Hyperspectral LiDAR Echo Waveform AnalysisabstractHyperspectral LiDAR (HSL) has demonstrated significant promise in achieving super range resolution, yet the effects of measurement variables like multi-target spectral relationships and signal-to-noise ratio (SNR) on its limits are unclear. This research establishes a mathematical model for HSL’s multi-layer target detection, validated by a 94% match with measurements. A novel method for spatial unmixing is introduced, informed by prior-knowledge acquisition and waveform decomposition, demonstrating that with SNR over 10 dB, HSL can resolve targets 10 cm apart with distinct spectral signatures, using a 4 ns pulse width. These findings advance the understanding of HSL’s capabilities for detailed environmental awareness and object discrimination. Yuhao Xia, Shilong Xu, Shengjie Ma, Wenxin Tian, Yihua Hu 0001 |
IGARSS | 1 |
| 2023 | Inter-domain distance prediction based on deep learning for domain assemblyabstractAlphaFold2 achieved a breakthrough in protein structure prediction through the end-to-end deep learning method, which can predict nearly all single-domain proteins at experimental resolution. However, the prediction accuracy of full-chain proteins is generally lower than that of single-domain proteins because of the incorrect interactions between domains. In this work, we develop an inter-domain distance prediction method, named DeepIDDP. In DeepIDDP, we design a neural network with attention mechanisms, where two new inter-domain features are used to enhance the ability to capture the interactions between domains. Furthermore, we propose a data enhancement strategy termed DPMSA, which is employed to deal with the absence of co-evolutionary information on targets. We integrate DeepIDDP into our previously developed domain assembly method SADA, termed SADA-DeepIDDP. Tested on a given multi-domain benchmark dataset, the accuracy of SADA-DeepIDDP inter-domain distance prediction is 11.3% and 21.6% higher than trRosettaX and trRosetta, respectively. The accuracy of the domain assembly model is 2.5% higher than that of SADA. Meanwhile, we reassemble 68 human multi-domain protein models with TM-score ≤ 0.80 from the AlphaFold protein structure database, where the average TM-score is improved by 11.8% after the reassembly by our method. The online server is at http://zhanglab-bioinf.com/DeepIDDP/. Fengqi Ge, Chunxiang Peng, Yuhao Xia, Guijun Zhang |
Briefings Bioinform. | 4 |
| 2023 | Pathfinder: Protein folding pathway prediction based on conformational samplingabstractThe study of protein folding mechanism is a challenge in molecular biology, which is of great significance for revealing the movement rules of biological macromolecules, understanding the pathogenic mechanism of folding diseases, and designing protein engineering materials. Based on the hypothesis that the conformational sampling trajectory contain the information of folding pathway, we propose a protein folding pathway prediction algorithm named Pathfinder. Firstly, Pathfinder performs large-scale sampling of the conformational space and clusters the decoys obtained in the sampling. The heterogeneous conformations obtained by clustering are named seed states. Then, a resampling algorithm that is not constrained by the local energy basin is designed to obtain the transition probabilities of seed states. Finally, protein folding pathways are inferred from the maximum transition probabilities of seed states. The proposed Pathfinder is tested on our developed test set (34 proteins). For 11 widely studied proteins, we correctly predicted their folding pathways and specifically analyzed 5 of them. For 13 proteins, we predicted their folding pathways to be further verified by biological experiments. For 6 proteins, we analyzed the reasons for the low prediction accuracy. For the other 4 proteins without biological experiment results, potential folding pathways were predicted to provide new insights into protein folding mechanism. The results reveal that structural analogs may have different folding pathways to express different biological functions, homologous proteins may contain common folding pathways, and α-helices may be more prone to early protein folding than β-strands. Zhaohong Huang, Yuhao Xia, Kailong Zhao, Guijun Zhang |
PLoS Comput. Biol. | 3 |
| 2023 | Range Resolution Enhanced Method With Spectral Properties for Hyperspectral LiDARabstractWaveform decomposition is needed as a first step in the extraction of various types of geometric and spectral information from hyperspectral full-waveform LiDAR echoes. We present a new approach to deal with the ”Pseudo-monopulse” waveform formed by the overlapped waveforms from multi-targets when they are very close. We use one single skew-normal distribution (SND) model to fit waveforms of all spectral channels first and count the geometric center position distribution of the echoes to decide whether it contains multi-targets. The geometric center position distribution of the ”Pseudo-monopulse” presents aggregation and asymmetry with the change of wavelength, while such an asymmetric phenomenon cannot be found from the echoes of the single target. Both theoretical and experimental data verify the point. Based on such observation, we further propose a hyperspectral waveform decomposition method utilizing the SND mixture model with: 1) initializing new waveform component parameters and their ranges based on the distinction of the three characteristics (geometric center position, pulse width, and skew-coefficient) between the echo and fitted SND waveform and 2) conducting single-channel waveform decomposition for all channels and 3) setting thresholds to find outlier channels based on statistical parameters of all single-channel decomposition results (the standard deviation and the means of geometric center position) and 4) re-conducting single-channel waveform decomposition for these outlier channels. The proposed method significantly improves the range resolution from 60cm to 5cm at most for a 4ns width laser pulse and represents the state-of-the-art in ”Pseudo-monopulse” waveform decomposition. Yuhao Xia, Shilong Xu, Ahui Hou, Jiajie Fang, Youlong Chen, Jiaqi Wen, Fashuai Li, Yuwei Chen 0005, Yihua Hu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Structural analogue-based protein structure domain assembly assisted by deep learningabstractMOTIVATION: With the breakthrough of AlphaFold2, the protein structure prediction problem has made remarkable progress through deep learning end-to-end techniques, in which correct folds could be built for nearly all single-domain proteins. However, the full-chain modelling appears to be lower on average accuracy than that for the constituent domains and requires higher demand on computing hardware, indicating the performance of full-chain modelling still needs to be improved. In this study, we investigate whether the predicted accuracy of the full-chain model can be further improved by domain assembly assisted by deep learning. RESULTS: In this article, we developed a structural analogue-based protein structure domain assembly method assisted by deep learning, named SADA. In SADA, a multi-domain protein structure database was constructed for the full-chain analogue detection using individual domain models. Starting from the initial model constructed from the analogue, the domain assembly simulation was performed to generate the full-chain model through a two-stage differential evolution algorithm guided by the energy function with an inter-residue distance potential predicted by deep learning. SADA was compared with the state-of-the-art domain assembly methods on 356 benchmark proteins, and the average TM-score of SADA models is 8.1% and 27.0% higher than that of DEMO and AIDA, respectively. We also assembled 293 human multi-domain proteins, where the average TM-score of the full-chain model after the assembly by SADA is 1.1% higher than that of the model by AlphaFold2. To conclude, we find that the domains often interact in the similar way in the quaternary orientations if the domains have similar tertiary structures. Furthermore, homologous templates and structural analogues are complementary for multi-domain protein full-chain modelling. AVAILABILITY AND IMPLEMENTATION: http://zhanglab-bioinf.com/SADA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chunxiang Peng, Yuhao Xia, Jun Liu 0078, Minghua Hou, Guijun Zhang |
Bioinform. | 3 |
| 2021 | Distance-guided protein folding based on generalized descent directionabstractAdvances in the prediction of the inter-residue distance for a protein sequence have increased the accuracy to predict the correct folds of proteins with distance information. Here, we propose a distance-guided protein folding algorithm based on generalized descent direction, named GDDfold, which achieves effective structural perturbation and potential minimization in two stages. In the global stage, random-based direction is designed using evolutionary knowledge, which guides conformation population to cross potential barriers and explore conformational space rapidly in a large range. In the local stage, locally rugged potential landscape can be explored with the aid of conjugate-based direction integrated into a specific search strategy, which can improve the exploitation ability. GDDfold is tested on 347 proteins of a benchmark set, 24 template-free modeling (FM) approaches targets of CASP13 and 20 FM targets of CASP14. Results show that GDDfold correctly folds [template modeling (TM) score ≥ = 0.5] 316 out of 347 proteins, where 65 proteins have TM scores that are greater than 0.8, and significantly outperforms Rosetta-dist (distance-assisted fragment assembly method) and L-BFGSfold (distance geometry optimization method). On CASP FM targets, GDDfold is comparable with five state-of-the-art full-version methods, namely, Quark, RaptorX, Rosetta, MULTICOM and trRosetta in the CASP 13 and 14 server groups. Liujing Wang, Jun Liu 0078, Yuhao Xia, Guijun Zhang |
Briefings Bioinform. | 3 |
| 2021 | A sequential niche multimodal conformational sampling algorithm for protein structure predictionabstractMOTIVATION: Massive local minima on the protein energy landscape often cause traditional conformational sampling algorithms to be easily trapped in local basin regions, because they find it difficult to overcome high-energy barriers. Also, the lowest energy conformation may not correspond to the native structure due to the inaccuracy of energy models. This study investigates whether these two problems can be alleviated by a sequential niche technique without loss of accuracy. RESULTS: A sequential niche multimodal conformational sampling algorithm for protein structure prediction (SNfold) is proposed in this study. In SNfold, a derating function is designed based on the knowledge learned from the previous sampling and used to construct a series of sampling-guided energy functions. These functions then help the sampling algorithm overcome high-energy barriers and avoid the re-sampling of the explored regions. In inaccurate protein energy models, the high-energy conformation that may correspond to the native structure can be sampled with successively updated sampling-guided energy functions. The proposed SNfold is tested on 300 benchmark proteins, 24 CASP13 and 19 CASP14 FM targets. Results show that SNfold correctly folds (TM-score ≥ 0.5) 231 out of 300 proteins. In particular, compared with Rosetta restrained by distance (Rosetta-dist), SNfold achieves higher average TM-score and improves the sampling efficiency by more than 100 times. On several CASP FM targets, SNfold also shows good performance compared with four state-of-the-art servers in CASP. As a plug-in conformational sampling algorithm, SNfold can be extended to other protein structure prediction methods. AVAILABILITY AND IMPLEMENTATION: The source code and executable versions are freely available at https://github.com/iobio-zjut/SNfold. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yuhao Xia, Chunxiang Peng, Guijun Zhang |
Bioinform. | 1 |