Yadong Liu 0001

dblp:75/5732-1 · DBLP profile ↗
← Back
42ranked-venue papers
6as first author
29since 2021 · last 2026
0000-0002-6322-3944ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 24 · 4 first-author · 20 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
YearPublicationVenuePosition
2026 cuteSV-OL: a real-time structural variation detection framework for nanopore sequencing devices
abstract
SUMMARY: Nanopore sequencing technology enables real-time sequencing and is widely used in rapid detection applications. However, in clinical scenarios, existing structural variant (SV) detection tools typically separate sequencing from computation, limiting their timeliness for clinical applications. To address this, we introduce cuteSV-OL, a novel framework designed for real-time SV discovery, which can be embedded within nanopore sequencing instruments to analyze data concurrently with its generation. Additionally, cuteSV-OL features a real-time SV detection rate evaluation module, allowing users to terminate sequencing early when appropriate, thereby reducing time and cost. Experimental results show that on a standard desktop computer, cuteSV-OL can perform real-time analysis during sequencing and complete SV calling within min after sequencing ends, achieving performance comparable to offline methods. This approach has the potential to enhance rapid clinical diagnostics. AVAILABILITY AND IMPLEMENTATION: cuteSV-OL is released under the MIT license and is available at https://github.com/gwmHIT/cuteSV-OL. It can also be installed via Bioconda or accessed through https://doi.org/10.5281/zenodo.17777436.
Weimin Guo, Yadong Liu 0001, Yadong Wang 0001, Tao Jiang 0021
Bioinform.2
2026 Self-Supervised Disentangled Representation Learning via Compositional Invariance
abstract
Images serve as a crucial information source for machine intelligence to understand the world, while how to represent images significantly impacting the generalizability and interpretability of intelligent systems. Disentangled representation learning offers a promising approach to improve both aspects. However, most of existing methods predominantly rely on statistical independence assumptions. This poses two key limitations: first, it fails to capture the reality that many concepts are both disentangled yet interrelated; second, it conflicts with human cognitive patterns where concepts naturally exhibit complex dependencies. These limitations further hinder collaboration between machine and human beings. To overcome these limitations, we propose Compositional Invariant Disentanglement (CID), a novel self-supervised learning method that enables models to learn composable representations aligned with human cognitive habits. Inspired by humans’ ability to flexibly recombine concepts, we reframe the definition of disentanglement through the lens of compositional invariance rather than statistical independence. This paradigm shift allows effective disentanglement even with correlated factors, achieving state-of-the-art disentanglement performance across multiple standard benchmarks (improved by 4.0% on Shapes3D, 4.5% on Dsprites, and 28.6% on MPI3D). Furthermore, by building upon and extending the successful self-supervised learning framework BYOL, CID demonstrates potential for large-scale disentanglement pre-training on unlabeled data. This work contributes to extracting more robust and interpretable representations from images for machine intelligence.
Haoqiang Chen, Jianxiang Sun, Yadong Liu 0001, Dewen Hu
IEEE Trans. Circuits Syst. Video Technol.3
2026 Decoding Driving Intentions via a Novel Brain-Computer Interface Paradigm With Low Cognitive Load and High Robustness
Jianxiang Sun, Zongtan Zhou, Yadong Liu 0001, Daxue Liu, Haoqiang Chen, Yingxin Liu, Dewen Hu
IEEE Trans. Syst. Man Cybern. Syst.3
2025 RVC: A Real-Time Variant Calling Framework for Short-Read Sequencing Data
abstract
Accurate detection of single-nucleotide variants (SNVs) and small insertions/deletions (indels) from second-generation sequencing (NGS) data is essential for clinical applications such as cancer diagnostics, infectious disease monitoring, and rapid genetic screening. However, conventional variant calling pipelines, such as GATK, decouple analysis from sequencing, deferring detection until sequencing is fully completed. We introduce RVC, a real-time variant calling framework tailored for cycle-based NGS workflows. RVC incrementally processes partially sequenced reads and continuously updates variant evidence using a scanline-based alignment algorithm and a lightweight binomial scoring model. This design enables progressive, low-latency SNVs and indels detection during sequencing, without disrupting the sequencing pipeline. In benchmark experiments using the HG002 dataset, RVC completed variant calling within tens of minutes after sequencing, significantly outperforming GATK in runtime. By tightly integrating analysis with sequencing output, RVC bridges the gap between sequencing speed and clinical responsiveness, offering a scalable and practical solution for real-time genomic diagnostics.
Miao Cui 0005, Tao Jiang 0021, Yadong Wang 0001, Bo Liu 0023, Guohua Wang 0001, Yadong Liu 0001
BIBM7
2025 MethSV: A Long-Read-Based Framework for Profiling DNA Methylation in Structural Variation
abstract
DNA methylation plays a crucial role in regulating diverse molecular processes in living organisms. However, methylation patterns associated with structural variations (SVs) remain poorly understood due to technical limitations. with the advent of long-read sequencing technologies, it is now feasible to simultaneously detect SVs and methylation at single-molecule resolution. Here, we propose methSV, a novel framework to profile DNA methylation within SV regions, revealing potential interactions between genetic and epigenetic regulation. Using methSV, we achieved a$\sim 3.7$-fold increase in the detection of SV-associated population-level differentially methylated regions (pDMRs) and \~{}4.2-fold increase in high-confidence haplotype-resolved differentially methylated regions (hDMRs). Notably, distinct methylation signatures in high-type deletions (DELs) and low-type insertions (INSs) suggested regulatory potential and were enriched in disease-associated genes such as GP1BA, CHRNE, and HCN2. This framework extends the analytical capabilities of long-read epigenomic studies and offers new insights into the functional consequences of SV.
Weize Kong, Yadong Liu 0001, Yadong Wang 0001, Tao Jiang 0021
BIBM3
2025 Bameth: A Bilstm-Cross Attention Network for Methylation-Based Progression-Level Classification of Colon Adenocarcinoma
abstract
Colon adenocarcinoma (COAD) is a prevalent malignancy with high morbidity and mortality, largely due to its asymptomatic onset and late diagnosis. DNA methylation has emerged as a promising biomarker for cancer progression and molecular classification. To capture the stage-based methylation patterns, we propose BAMeth, a novel deep learning framework combining bidirectional Long ShortTerm Memory (BiLSTM) and cross-attention mechanisms. In this study, COAD samples were categorized into four types based on clinical stages and pathological characteristics, representing early (Type I), intermediate (Type II), late (Type III), and normal (Type IV) conditions. BAMeth performs multi-class classification across these stage-based categories using genome-wide methylation profiles. The architecture integrates dimensionality reduction, sequential modeling, and inter-feature dependency learning, enabling extraction of both sequential and contextual dependencies from highdimensional methylation data. Experimental results demonstrate that BAMeth achieves superior performance compared with some existing methods, particularly in terms of accuracy (ACC) by leveraging differentially methylated positions. Functional enrichment analysis of the genes annotated to the key CpG sites further reveals biological relevance in Wnt signaling and ubiquitin-mediated proteolysis pathways, highlighting its interpretability. Overall, BAMeth provides a robust and biologically interpretable framework for stage-based molecular classification of colon adenocarcinoma, offering potential for early detection and progression monitoring in epigenomics-driven oncology research.
Yue Liu 0034, Zhongyu Liu, Tao Jiang 0014, Yadong Wang 0001, Yadong Liu 0001
BIBM5
2025 SimPG: A Pangenome-Guided Population-Specific Human Genome Simulation Tool
abstract
The rapid advancement of high-throughput genome sequencing has enabled large-scale reconstruction of genome sequences at both individual and population levels. However, accurately evaluating genome assemblies and associated variants remains a significant challenge due to sequencing errors, assembly artifacts, and limited experimental validation, which hinder the development of novel algorithms, benchmarking of genomic pipelines, and validation of biological hypotheses. Existing linear genome simulation tools, typically based on a single reference genome, offer limited biological realism and fail to capture the genomic diversity and structural complexity needed for comprehensive evaluation. To address these limitations, we introduce SimPG, a novel simulation framework that generates individual genomes with populationlevel characteristics by leveraging the rich variant and structural information embedded in pangenomes. SimPG produces realistic, high-quality simulated genomes that support diverse applications such as structural variant detection and population genetics research. By providing a reproducible, controllable, and standardized environment, SimPG facilitates the development, testing, and benchmarking of genome analysis tools, enabling robust performance assessment and error analysis, and ultimately advancing computational genomics and genome biology.
Yadong Liu 0001, Yadong Wang 0001, Tao Jiang 0021
BIBM3
2025 DNAMatch: An Ultra-Fast and Memory-Efficient Deep Learning Framework for Aligning Ultra-Long DNA Fragments
abstract
Efficient and accurate alignment of DNA fragments is fundamental to genomics research. With the advent of advanced sequencing technologies, the length of sequencing reads and assembled contigs has increased significantly, posing substantial challenges for existing alignment algorithms. These methods often struggle with megabase-scale DNA fragments due to the computational burden of global searches and exhaustive chromosomal queries. To address this, we propose DNAMatch, a novel alignment framework that integrates the DNABERT2 pre-trained model for feature extraction with a deep residual network for chromosome identification. DNAMatch introduces a chromosome pre-localization strategy, which effectively narrows the search space and significantly reduces the memory footprint required by the downstream aligner, minimap2. This design enables fast and precise alignment of ultra-long DNA fragments. Benchmarking on simulated ultra-long reads from multiple model organisms-including human, Drosophila melanogaster, and Arabidopsis thaliana-demonstrates that DNAMatch achieves 9 8-9 9% accuracy in chromosome identification, accelerates the alignment process by 52.7%, and reduces memory usage by 73%. Importantly, when applied to downstream structural variation detection, DNAMatch maintains high accuracy, with only a marginal 0.05% decrease in F1-score compared to the standard minimap2 pipeline. These results highlight DNAMatch as a powerful and efficient tool for aligning ultra-long reads and analyzing complex genomes.
Yuansong Zhu, Chuanmin Wu, Yadong Wang 0001, Tao Jiang 0021, Yadong Liu 0001
BIBM5
2025 CKG-TPI: integrating collaborative knowledge graph with sequence interactions for TCR-peptide binding specificity
abstract
Accurately identifying interactions between T-cell receptors (TCRs) and peptides is a fundamental challenge in immunology, with significant implications for vaccine design and immunotherapy. While computational methods offer efficient alternatives to labor-intensive experimental screening, achieving robust and accurate TCR-peptide binding prediction remains a challenging task. To address this, we propose collaborative knowledge graph (CKG-TPI), a novel prediction framework based on graph neural networks that integrates both interaction patterns between TCR and peptide sequences and their higher-order biological context through a constructed collaborative knowledge graph. Experimental results on multiple publicly available independent datasets demonstrate that CKG-TPI consistently outperforms state-of-the-art models. Specifically, it achieves a 9.89% improvement in area under the ROC curve compared to the strongest baseline model UnifyImmun, and a 23.93% increase in area under the precision-recall curve over the leading baseline method. Moreover, attention weight visualization and peptide-specific TCR screening validate the model's effectiveness, underscoring its potential as a powerful tool for immunological research and therapeutic discovery.
Yue Liu 0034, Haoyan Wang, Guohua Wang 0001, Yadong Liu 0001, Tao Jiang 0021, Yadong Wang 0001
Briefings Bioinform.4
2025 Pessimistic policy iteration with bounded uncertainty
abstract
Offline Reinforcement Learning (RL) aims to learn policies by using static datasets. The extrapolation error in out-of-distribution (OOD) samples can cause off-policy RL algorithms to perform poorly on offline datasets. Hence, it is critical to avoid visiting OOD states and taking OOD actions in offline RL. Several recent methods have used uncertainty estimation to distinguish OOD samples. However, errors in the uncertainty estimation make the purely uncertainty-based method unstable and require additional components to ensure sufficient pessimism . In this study, we propose a Bounded Uncertainty based Pessimistic policy iteration algorithm (BUP). The BUP pessimistically estimates the value function via bounded uncertainty, and the uncertainty bound is achieved by constraining the actor from taking highly uncertain actions. The suboptimality bound of BUP is theoretically guaranteed in linear Markov Decision Processes (MDPs), and experiments on D4RL datasets show that BUP matches the state-of-the-art performance. Moreover, BUP is simple to implement with low computational cost and does not require any additional components.
Zhiyong Peng 0002, Changlin Han, Yadong Liu 0001, Jingsheng Tang, Zongtan Zhou
Expert Syst. Appl.3
2025 An Early Warning Approach for Pilots' Cognitive Tipping Points Based Multi-Modal Signals
abstract
When executing complex missions in emergency scenarios, pilots’ cognitive state may deteriorate, posing significant challenges to flight safety and mission execution. This paper proposes a cross-subject and cross-session early warning approach based on small-sample and multi-modal signals for predicting cognitive collapse state. We designs an experimental paradigm that has been demonstrated to induce cognitive collapse in 87$\%$of trials by analyzing questionnaire scores, physiological signals, and task performance. The extracted multi-modal features are fused and selected to construct the optimal feature set. Further, a two-step early warning method is introduced to identify critical slowing down and to predict tipping points by detecting the high cognitive workload state and classifying the corresponding signals into early warning signal (EWS) and normal state. The early warning of pilot’s cognitive tipping point obtains the result that recall is 77.76$\%$, false alarm rate is 28.43$\%$, and the AUC of the warning method can reach 0.83. Our method shows better early warning performance compared with other classification models and has strong generalizability on other datasets.Note to Practitioners—The pilots’ cognitive state is crucial for mission completion and flight safety. Previous studies mainly focused on assessing current cognitive state, and there is little research on early warning of cognitive collapse that occurs in emergency scenarios (e.g., faults of system and the sudden increase in flight missions). In particular, identification is not effective for small-sample and instable data. Therefore, this paper presents an early warning method for cognitive tipping points, which can provide timely warning of pilots’ cognitive collapse and reduce human-caused risks in emergency scenarios.
Yadong Liu 0001, Dewen Hu
IEEE Trans Autom. Sci. Eng.2
2025 Causal Confusion in Pedestrian Crossing Intention Prediction for Autonomous Vehicles: The Role of Ego-Vehicle Speed
abstract
Pedestrian crossing intention prediction is crucial for autonomous vehicles due to their inherent inertia, yet challenging. The prevailing practice is to leverage multi-modal data that correlate with pedestrian crossing intention as input to infer it, with ego-vehicle speed being a commonly used modality. However, from a causal perspective, we identify two critical issues overlooked by existing methods: 1) The causal relationship between ego-vehicle speed and pedestrian crossing intention is inconsistent across the training and testing phases, leading to a distribution shift in the data; 2) There exists a bidirectional causality between ego-vehicle speed and pedestrian crossing intention, comprising forward causality and anti-causality. The imbalanced distribution of these two causal directions in natural datasets results in causal confusion, further exacerbating the distribution shift. These issues lead to a counter-intuitive hypothesis: removing ego-vehicle speed as input can actually benefit prediction performance. Therefore, we propose LIM (Less Is More), a uni-modal model that utilizes only skeleton sequences as input. LIM features specially designed modules for efficient skeleton sequence processing, eliminating the reliance on unstable ego-vehicle speed. Moreover, LIM employs adversarial training to identify and remove any latent ego-vehicle speed information embedded within the skeleton sequences. LIM achieves competitive prediction performance compared to state-of-the-art multi-modal models while offering superior real-time capabilities and lower computational costs. By revealing critical issues overlooked by most existing work and providing a more robust crossing intention prediction model, our work contributes to the development of safer autonomous vehicles.
Haoqiang Chen, Jianxiang Sun, Yadong Liu 0001, Zhiyong Peng 0002, Dewen Hu
IEEE Trans. Intell. Transp. Syst.3
2024 Comprehensive Benchmarking of Genotype Imputation Tools Using a Large-Scale Chinese Reference Panel
abstract
Genome-wide association studies (GWAS) remain as one of the most essential and potent strategies for identifying genetic markers associated with common human diseases. Gene arrays, which are widely used in GWAS, target predesigned known genomic mutations to enhance clinical diagnostics. The approach of genotype imputation which leverages reference panels has shown great promise in extrapolating additional meaningful variants from existing data. Therefore, accurately broadening the spectrum of mutations beyond those included in gene arrays can unlock significant potential for cost-effective and advanced GWAS analysis. In this study, we focus on the four most popular imputation tools currently to assess their performances based on a large-scale Chinese reference panel. Benchmarking results demonstrate that Minimac4 is the best imputation method, which achieved higher accuracy with more well-imputed loci and maintained lower computational consumption. Moreover, we further revealed the minimum sample size required in the reference panel and the mutations involved in arrays that ensure the acquisition of acceptable imputed call sets. We believe this study can provide practical significance guidelines for the research community to select the suitable imputation tool via an existing reference panel, or construct their own reference panel and genotype array in the most effective way.
Yadong Liu 0001, Zhongbo Yang, Yadong Wang 0001, Tao Jiang 0021
BIBM2
2024 TDLM: A Diffusion Language Model for TCR Sequence Exploration and Generation
abstract
The adaptive immune response relies on the ability of T-cell receptors (TCRs) to recognize specific antigens. The vast diversity of TCRs allows T-cells to recognize a broad spectrum of antigens, but this complexity also poses challenges for understanding and predicting TCR-antigen binding specificity. Despite the development of various machine learning and deep learning methods for prediction and clustering, there remains a need for a versatile and effective TCR language framework that can be flexibly applied to various downstream tasks, including sequence generation. Here we present TDLM, a T-cell Receptor (TCR) diffusion language model, designed to decode complex patterns within TCR sequences and apply them across various downstream tasks. Firstly, TDLM can be trained on unlabeled TCR sequence data, enabling it to utilize vast datasets to generate comprehensive embeddings. When compared to other embedding methods, TDLM embeddings enhance TCR-antigen binding prediction accuracy and enable effective TCR sequence clustering and similarity analysis, helping identify TCRs with shared antigen specificity. Furthermore, as a diffusion-based generative model, TDLM can generate highly diverse and specific TCR sequences. This ability is invaluable for the rapid screening and optimization of TCRs with target antigen specificities, offering significant potential in disease diagnosis, personalized immunotherapy, and vaccine research. The code is available at: https://github.com/skybluewhy/TDLM
Haoyan Wang, Tianyi Zang, Yadong Liu 0001
BIBM5
2024 Comprehensive evaluation of haplotype phasing tools with different strategies across diverse sequence technologies
abstract
Haplotype phasing is a computational technique used to determine the allelic arrangement in a chromosome from genotype data. This process is crucial for understanding genetic diversity and inheritance patterns, which are essential for fields like genetic research, personalized medicine, and population genetic studies. In this study, we conducted a comprehensive evaluation of five state-of-the-art haplotype phasing tools under different strategies including statistical-based Beagle5, Eagle2 and Shapeit5, and read-based WhatsHap, and HapCut2, across various sequencing platforms including Illumina, PacBio, and ONT. Our findings indicate that statistical-based tools generally provide better phasing continuity and are less dependent on the sequencing technology used, whereas read-based tools offer superior performance with long-read datasets but struggle with short reads due to insufficient sequence length. In addition, statistical-based tools were found to be less resource-intensive compared to their read-based counterparts. The insights gained from this benchmarking provide a valuable guide for researchers in selecting the most appropriate phasing tools, potentially enhancing the accuracy and efficiency of genetic analysis and supporting advancements in genetic research methodologies.
Zhongbo Yang, Zhenhao Lu, Tao Jiang 0021, Yadong Wang 0001, Yadong Liu 0001
BIBM6
2024 miniSNV: accurate and fast single nucleotide variant calling from nanopore sequencing data
abstract
Nanopore sequence technology has demonstrated a longer read length and enabled to potentially address the limitations of short-read sequencing including long-range haplotype phasing and accurate variant calling. However, there is still room for improvement in terms of the performance of single nucleotide variant (SNV) identification and computing resource usage for the state-of-the-art approaches. In this work, we introduce miniSNV, a lightweight SNV calling algorithm that simultaneously achieves high performance and yield. miniSNV utilizes known common variants in populations as variation backgrounds and leverages read pileup, read-based phasing, and consensus generation to identify and genotype SNVs for Oxford Nanopore Technologies (ONT) long reads. Benchmarks on real and simulated ONT data under various error profiles demonstrate that miniSNV has superior sensitivity and comparable accuracy on SNV detection and runs faster with outstanding scalability and lower memory than most state-of-the-art variant callers. miniSNV is available from https://github.com/CuiMiao-HIT/miniSNV.
Miao Cui 0005, Yadong Liu 0001, Hongzhe Guo, Tao Jiang 0021, Yadong Wang 0001, Bo Liu 0023
Briefings Bioinform.2
2024 Kled: an ultra-fast and sensitive structural variant detection tool for long-read sequencing data
abstract
Structural Variants (SVs) are a crucial type of genetic variant that can significantly impact phenotypes. Therefore, the identification of SVs is an essential part of modern genomic analysis. In this article, we present kled, an ultra-fast and sensitive SV caller for long-read sequencing data given the specially designed approach with a novel signature-merging algorithm, custom refinement strategies and a high-performance program structure. The evaluation results demonstrate that kled can achieve optimal SV calling compared to several state-of-the-art methods on simulated and real long-read data for different platforms and sequencing depths. Furthermore, kled excels at rapid SV calling and can efficiently utilize multiple Central Processing Unit (CPU) cores while maintaining low memory usage. The source code for kled can be obtained from https://github.com/CoREse/kled.
Tao Jiang 0021, Shuqi Cao, Yadong Liu 0001, Bo Liu 0023, Yadong Wang 0001
Briefings Bioinform.5
2024 MEHunter: transformer-based mobile element variant detection from long reads
abstract
SUMMARY: Mobile genetic elements (MEs) are heritable mutagens that significantly contribute to genetic diseases. The advent of long-read sequencing technologies, capable of resolving large DNA fragments, offers promising prospects for the comprehensive detection of ME variants (MEVs). However, achieving high precision while maintaining recall performance remains challenging mainly brought by the variable length and similar content of MEV signatures, which are often obscured by the noise in long reads. Here, we propose MEHunter, a high-performance MEV detection approach utilizing a fine-tuned transformer model adept at identifying potential MEVs with fragmented features. Benchmark experiments on both simulated and real datasets demonstrate that MEHunter consistently achieves higher accuracy and sensitivity than the state-of-the-art tools. Furthermore, it is capable of detecting novel potentially individual-specific MEVs that have been overlooked in published population projects. AVAILABILITY AND IMPLEMENTATION: MEHunter is available from https://github.com/120L021101/MEHunter.
Tao Jiang 0021, Zuji Zhou, Shuqi Cao, Yadong Wang 0001, Yadong Liu 0001
Bioinform.6
2024 Deadly triad matters for offline reinforcement learning
Zhiyong Peng 0002, Yadong Liu 0001, Zongtan Zhou
Knowl. Based Syst.2
2023 Weighted Policy Constraints for Offline Reinforcement Learning
abstract
Offline reinforcement learning (RL) aims to learn policy from the passively collected offline dataset. Applying existing RL methods on the static dataset straightforwardly will raise distribution shift, causing these unconstrained RL methods to fail. To cope with the distribution shift problem, a common practice in offline RL is to constrain the policy explicitly or implicitly close to behavioral policy. However, the available dataset usually contains sub-optimal or inferior actions, constraining the policy near all these actions will make the policy inevitably learn inferior behaviors, limiting the performance of the algorithm. Based on this observation, we propose a weighted policy constraints (wPC) method that only constrains the learned policy to desirable behaviors, making room for policy improvement on other parts. Our algorithm outperforms existing state-of-the-art offline RL algorithms on the D4RL offline gym datasets. Moreover, the proposed algorithm is simple to implement with few hyper-parameters, making the proposed wPC algorithm a robust offline RL method with low computational complexity.
Zhiyong Peng 0002, Changlin Han, Yadong Liu 0001, Zongtan Zhou
AAAI3
2023 Comprehensive evaluation of RNA-seq alignment methods based on long-read sequencing data
abstract
Long-read RNA sequencing (RNA-seq) has revolutionized our ability to comprehensively study transcriptomes, enabling the detection of full-length transcripts. The accurate alignment of long-read RNA-seq is a critical step in downstream analysis. However, with the development of alignment tools and each tool declaring the ability to handle long-read RNA-seq alignment on their own, researchers face the challenge of distinguishing and selecting the most suitable tool for their tasks. But to the best of our knowledge, there is still a lack of comprehensive evaluations on the aligners designed for long-read sequencing data. Here, we conducted a benchmark on five state-of-the-art tools on simulated long-read RNA-seq data from three levels to comprehensively assess the performance of each aligner, and provide valuable insights into the strengths and limitations of each aligner, aiding researchers in selecting the most appropriate tool for their long-read RNA-seq studies. We expected our study to advance the field of transcriptomics and promote the accurate analysis of long-read RNA-seq data in diverse biological contexts.
Yadong Liu 0001, Hongzhe Guo, Zhenhao Lu, Yadong Wang 0001, Zhongyu Liu, Tao Jiang 0021
BIBM1
2023 Overfitting-avoiding goal-guided exploration for hard-exploration multi-goal reinforcement learning
Changlin Han, Zhiyong Peng 0002, Yadong Liu 0001, Jingsheng Tang, Yang Yu 0014, Zongtan Zhou
Neurocomputing3
2023 Conservative network for offline reinforcement learning
Zhiyong Peng 0002, Yadong Liu 0001, Haoqiang Chen, Zongtan Zhou
Knowl. Based Syst.2
2022 Comparison of the Nanopore and PacBio sequencing technologies for DNA 5-methylcytosine detection
abstract
DNA methylation provides a pivotal layer of epigenetic regulation in eukaryotes that has significant involvement for numerous biological processes in health and disease. Recent long-read sequencing technology including Oxford Nanopore sequencing and PacBio HiFi sequencing greatly expands the capacity of long-range, single-molecule, and direct DNA modification detection from reads without extra laboratory techniques. A growing number of analytical pipelines including base-calling and 5mC methylation detection have been developed, but there is still a lack of comprehensive evaluations of the two sequencing technologies. Here, we assess the performance of different methylation-calling pipelines based on Nanopore and HiFi sequencing datasets to provide a systematic evaluation to guide researchers on how to select the long-read sequencing technologies in performing human epigenome-wide studies.
Yadong Liu 0001, Zhongyu Liu, Tao Jiang 0021, Tianyi Zang, Yadong Wang 0001
BIBM1
2022 StackCirRNAPred: computational classification of long circRNA from other lncRNA based on stacking strategy
abstract
BACKGROUND: CircRNAs are essential for the regulation of post-transcriptional gene expression, including as miRNA sponges, and play an important role in disease development. Some computational tools have been proposed recently to predict circRNA, since only one classifier is used, there is still much that can be done to improve the performance. RESULTS: StackCirRNAPred was proposed, the computational classification of long circRNA from other lncRNA based on stacking strategy. In order to cope with the potential problem that a single feature might not be able to distinguish circRNA well from other lncRNA, we first extracted features from different sources, including nucleic acid composition, sequence spatial features and physicochemical properties, Alu and tandem repeats. We innovatively apply the stacking strategy to integrate the more advantageous classifiers of RF, LightGBM, XGBoost. This allows the model to incorporate these features more flexibly. StackCirRNAPred was found to be significantly better than other tools, with precision, accuracy, F1, recall and MCC of 0.843, 0.833, 0.831, 0.819 and 0.666 respectively. We tested it directly on the mouse dataset. StackCirRNAPred was still significantly better than other methods, with precision, accuracy, F1, recall and MCC of 0.837, 0.839, 0.839, 0.841, 0.677. CONCLUSIONS: We proposed StackCirRNAPred based on stacking strategy to distinguish long circRNAs from other lncRNAs. With the test results demonstrating the validity and robustness of StackCirRNAPred, we hope StackCirRNAPred will complement existing circRNA prediction methods and is helpful in down-stream research.
Xin Wang 0124, Yadong Liu 0001, Jie Li 0055, Guohua Wang 0001
BMC Bioinform.2
2022 Giant magneto-impedance sensor with working point selfadaptation for unshielded human bio-magnetic detection
abstract
Compared with traditional biomagnetic field detection devices, such as superconducting quantum interference devices (SQUIDs) and atomic magnetometers, only giant magnetoimpedance (GMI) sensors can be applied for unshielded human brain biomagnetic detection, and they have the potential for application in next-generation wearable equipment for brain-computer interfaces (BCIs). Achieving a better GMI sensor without magnetic shielding requires the stimulation of the GMI effect to be maximized and environmental noise interference to be minimized. Moreover, the GMI effect stimulated in an amorphous filament is closely related to its working point, which is sensitive to both the external magnetic field and the drive current of the filament. In this paper, we propose a new noisereducing GMI gradiometer with a dual-loop self-adapting structure. Noise reduction is realized by a direction-flexible differential probe, and the dual-loop structure optimizes and stabilizes the working point by automatically controlling the external magnetic field and drive current. This dual-loop structure is fully program controlled by a micro control unit (MCU), which not only simplifies the traditional constantparameter sensor circuit, saving the time required to adjust the circuit component parameters, but also improves the sensor performance and environmental adaptation. In the performance test, within 2 min of self-adaptation, our sensor showed a better sensitivity and signal-to-noise ratio (SNR) than those of the traditional designs and achieved a background noise of 12 pT/√Hz at 10 Hz and 7pT/√Hz at 200 Hz. To the best of our knowledge, our sensor is the first to realize self-adaptation of both the external magnetic field and the drive current.
Changlin Han, Ming Xu 0022, Jingsheng Tang, Yadong Liu 0001, Zongtan Zhou
Virtual Real. Intell. Hardw.4
2021 Brain-computer interface for human-multirobot strategic consensus with a differential world model
Wei Dai 0014, Huimin Lu 0002, Yadong Liu 0001, Zongtan Zhou
Appl. Intell.4
2021 SKSV: ultrafast structural variation detection from circular consensus sequencing reads
abstract
SUMMARY: Circular consensus sequencing reads are promising for the comprehensive detection of structural variants (SVs). However, alignment-based SV calling pipelines are computationally intensive due to the generation of complete read-alignments and its post-processing. Herein, we propose a SKeleton-based analysis toolkit for Structural Variation detection (SKSV). Benchmarks on real and simulated datasets demonstrate that SKSV has an order of magnitude of faster speed than state-of-the-art SV calling approaches; moreover, it achieves higher F1 scores for various types of SVs. AVAILABILITY AND IMPLEMENTATION: SKSV is available from https://github.com/ydLiu-HIT/SKSV. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yadong Liu 0001, Tao Jiang 0021, Junhao Su, Bo Liu 0023, Tianyi Zang, Yadong Wang 0001
Bioinform.1
2021 Long-read sequencing settings for efficient structural variation detection based on comprehensive evaluation
abstract
BACKGROUND: With the rapid development of long-read sequencing technologies, it is possible to reveal the full spectrum of genetic structural variation (SV). However, the expensive cost, finite read length and high sequencing error for long-read data greatly limit the widespread adoption of SV calling. Therefore, it is urgent to establish guidance concerning sequencing coverage, read length, and error rate to maintain high SV yields and to achieve the lowest cost simultaneously. RESULTS: In this study, we generated a full range of simulated error-prone long-read datasets containing various sequencing settings and comprehensively evaluated the performance of SV calling with state-of-the-art long-read SV detection methods. The benchmark results demonstrate that almost all SV callers perform better when the long-read data reach 20× coverage, 20 kbp average read length, and approximately 10-7.5% or below 1% error rates. Furthermore, high sequencing coverage is the most influential factor in promoting SV calling, while it also directly determines the expensive costs. CONCLUSIONS: Based on the comprehensive evaluation results, we provide important guidelines for selecting long-read sequencing settings for efficient SV calling. We believe these recommended settings of long-read sequencing will have extraordinary guiding significance in cutting-edge genomic studies and clinical practices.
Tao Jiang 0021, Shuqi Cao, Yadong Liu 0001, Yadong Wang 0001, Hongzhe Guo
BMC Bioinform.4
2020 GAMA: Graph Attention Multi-agent reinforcement learning algorithm for cooperation
Haoqiang Chen, Yadong Liu 0001, Zongtan Zhou, Dewen Hu, Ming Zhang 0027
Appl. Intell.2
2020 A Dynamic User Interface Based BCI Environmental Control System
abstract
In this study, a dynamic user interface (UI) is proposed in visual P300 Brain-Computer Interface (BCI) based environmental control system. A head-mounted Augmented Reality (AR) glass is used as the interactive media, which is used to assists the BCI system to build the dynamic UI with the scene in subject’s field of view. In the dynamic UI, based on the objects detected by the AR glass, options are dynamically generated. The subject can assign tasks by selecting different options in the dynamic UI. Five subjects successfully completed the task of controlling household appliances and navigating wheelchairs to designated destinations. Compared to static UI, the proposed dynamic UI has a 17.4% improvement in time delay. On average, only 1.9% of the commands resulted in incorrect operations. The dynamic UI makes progress in reducing time delay and incorrect operations. The proposed system provides a brand-new interactive method in BCI based applications.
Saisai Zhong, Yadong Liu 0001, Yang Yu 0014, Jingsheng Tang, Zongtan Zhou, Dewen Hu
Int. J. Hum. Comput. Interact.2
2019 Dense-CAM: Visualize the Gender of Brains with MRI Images
abstract
Studying the gender differences of brains is important to understand the brain cognitive mechanism. With the increasing amounts of brain imaging data, researchers start to use deep learning for brain image processing. However, the interpretability of deep models is the big gap between them. At present, deep model visualization is the main method to solve the interpretability problem of deep networks. Class Activation Map (CAM) is a commonly-used deep model visualization method. But the resolution of CAMs are restricted since it only uses the last layer for visualization. In this paper, we proposed a novel convolutional network named Dense-CAM by combining DenseNet and CAM and realized the visualization of the whole network so that it can generate more accurate and more robust deep model visualization. The network was tested with the gender classification problem using more than 6000 samples and achieved an accuracy of 92.93%. Brain regions with significant differences between men and women are found with the proposed method, which can be used for future brain imaging studies.
Kai Gao 0011, Hui Shen 0004, Yadong Liu 0001, Dewen Hu
IJCNN3
2019 An Tactile ERP-Based Brain-Computer Interface for Communication
abstract
A classical visual event relative potential (ERP) brain–computer interface (BCI) system relies on visual stimuli to choose commands. Users obtain most information about their surroundings visually as well. This large amount of information can aggravate visual burden and fatigue. In our study, we proposed a novel approach to evoke ERP with a tactile stimulus. To achieve this approach, we first designed a wireless stimulus module with vibrators to provide a tactile stimulus for the system. The vibrators were located on the subject’s arm to imitate the joint motion of a robotic arm. Then, the ERP feature and the parameters of classifiers were obtained through offline experimental data analysis. Based on the analysis, the suitable electrode channels, stimulus onset asynchrony (SOA), and filter upper limit were different for different subjects. According to those outcomes, a unique classifier was designed for each subject. Finally, 10 healthy BCI-naive subjects participated in online experiments to evaluate the performance of our tactile BCI system; they achieved an accuracy range from 78.67% to 100% with an average of 89.1% and an instantaneous transmission rate (ITR) range from 7.77 to 28.70 bits/min with an average of 14.77 bits/min. The accuracy of different subjects and SOAs remained relatively stable, the ITR fluctuated mainly due to the different SOAs, and we achieved balance between ITR and accuracy.
Yadong Liu 0001, Jingjun Wang, Erwei Yin, Yang Yu 0014, Zongtan Zhou, Dewen Hu
Int. J. Hum. Comput. Interact.1
2019 Toward Brain-Actuated Mobile Platform
abstract
This study presents a brain–computer interface (BCI) system aimed at providing disabled patients with mobile solutions for practical use. The proposed system employs an omnidirectional chassis and a bionic robot arm to construct a multi-functional mobile platform. In addition, the system is equipped with a Kinect and 12 ultrasonic sensors to capture environment information. Based on artificial intelligence technology, the mobile system can understand the environment and smartly completes certain tasks. A hybrid BCI combined with movement imagery paradigm and asynchronous P300 paradigm is designed to translate human intent to computer commands. The users interact with the system in a flexible way: on the one hand, the user issues commands to drive the system directly; on the other hand, the system searches for predefined operable targets and reports the results to the user. Once the user confirms the target, the system will automatically complete the associated operation. To evaluate the system’s performance, a testing environment with a small room, aisle, and an elevator was built to simulate the mobile tasks in the daily scene. Participants were instructed to operate the mobile system in the room, aisle, and using the elevator to go outdoors. In this study, four subjects participated in the test, and all of them completed the task.
Jingsheng Tang, Yadong Liu 0001, Jun Jiang 0001, Yang Yu 0014, Dewen Hu, Zongtan Zhou
Int. J. Hum. Comput. Interact.2
2019 Towards a Hybrid BCI Gaming Paradigm Based on Motor Imagery and SSVEP
abstract
Brain-computer interfaces (BCIs) not only can allow individuals to voluntarily control external devices, helping to restore lost motor functions of the disabled, but can also be used by healthy users for entertainment and gaming applications. In this study, we proposed a hybrid BCI paradigm to explore a feasible and natural way to play games by using electroencephalogram (EEG) signals in a practical environment. In this paradigm, we combined motor imagery (MI) and steady-state visually evoked potentials (SSVEPs) to generate multiple commands. A classic game, Tetris, was chosen as the control object. The novelty of this study includes the effective usage of a “dwell time” approach and fusion rules to design BCI games. To demonstrate the feasibility of the proposed hybrid paradigm, ten subjects were chosen to participate in online control experiments. The experimental results showed that all subjects successfully completed the predefined tasks with high accuracy. This proposed hybrid BCI paradigm could potentially provide those who suffer disability or paralysis with additional entertainment options, such as brain-actuated games, that could improve their happiness and quality of life.Abbreviations: BCI: brain-computer interface; EEG: electroencephalogram; MI: motor imagery; SSVEP: steady-state visually evoked potential; ERP: event-related potential; SMR: sensorimotor rhythm; VEP: visual evoked potential; TCP/IP: transmission control protocol/internet protocol; GUI: graphical user interface; ERD/ERS: event-related desynchronization/synchronization; CIC: control intention classifier; LRC: left/right classifier; CSP: common spatial pattern; LDA: linear discriminant analysis; ROC: receiver operating characteristic; TPR: true positive rate; FPR: false positive rate; CCA: canonical correlation analysis.
Zhihua Wang 0002, Yang Yu 0014, Ming Xu 0022, Yadong Liu 0001, Erwei Yin, Zongtan Zhou
Int. J. Hum. Comput. Interact.4
2017 Detect visual field using eye tracking and steady-state visual evoked potential
abstract
This paper makes the subjects' sight locked in a certain area using an eye tracker, getting Steady-state visual evoked potential (SSVEP) from flickering stimuli with a fixed frequency but at random positions, in order to observe the impact of stimulus at different positions and their distances on the electroencephalogram (EEG). The result suggests that if human have to select the positions of stimuli of SSVEP-BCI, it is an agreeable strategy to separate them at least 4 ° for avoiding the possible mistakes. We hope that it could help in setting distances between stimuli or updating pattern selection algorithms in the future BCI system and other paradigms.
Yadong Liu 0001, Zongtan Zhou, Dewen Hu, Erwei Yin
SMC2
2017 Toward a Hybrid BCI: Self-Paced Operation of a P300-based Speller by Merging a Motor Imagery-Based "Brain Switch" into a P300 Spelling Approach
abstract
This study presents the self-paced operation of a brain–computer interface (BCI) speller, which can be voluntarily turned on/off by merging a motor imagery (MI)-based brain switch into a P300-based BCI speller. From an off state (idle state), the users can generate a “control signal” by consciously changing the cognitive state differential from the idle state to turn on a P300-based spelling system when he or she wants to spell words. With the system turned on, the user can spell words, and then, the spelling system can be voluntarily turned off and switched to the initial state using a command. In this paradigm, the participants tried to perform the two different cognitive tasks sequentially, rather than simultaneously, and multiple EEG components were processed sequentially. The practicability and effectiveness of the proposed approach were validated by eleven participants, and all of them achieved a satisfactory performance. For the P300 speller, they achieved an average PITR of 42.61 bits/min. The preliminary results indicated that the proposed hybrid BCI system with different mental strategies operating sequentially is feasible and has potential applications for practical self-paced control.
Yang Yu 0014, Zongtan Zhou, Jun Jiang 0001, Erwei Yin, Kunjia Liu, Jingjun Wang, Yadong Liu 0001, Dewen Hu
Int. J. Hum. Comput. Interact.7
2016 A P300-Based Brain-Computer Interface for Chinese Character Input
abstract
The majority of previously developed assistive communication brain–computer interface systems have primarily focused on languages that are written in alphabetic scripts. However, languages that are written in logographic scripts, such as those in Chinese hanzi (or sinograms), pose a challenge for the implementation of visual spelling systems because it is impossible to simultaneously display thousands of items in a stimulus matrix of a reasonable size. In this study, a P300 visual spelling system that uses a novel method to input Chinese sinograms developed with a Hanyu Pinyin-based method is presented. This method transcribes a Chinese Pinyin into initial consonant and vowel components according to its Mandarin pronunciation. In this paradigm, each sinogram is input by selecting the initial consonant and then the vowel components and subsequently selecting the sinogram itself. Ten healthy subjects participated in the study and achieved an average offline accuracy of 92.6% with a mean information transfer rate of 39.2 bits/min and an average online input speed of one sinogram per 43.9 s. The preliminary results presented here indicated that the online input of Chinese text using a Pinyin-based visual speller is feasible.
Yang Yu 0014, Zongtan Zhou, Erwei Yin, Jun Jiang 0001, Yadong Liu 0001, Dewen Hu
Int. J. Hum. Comput. Interact.5
2015 Including Signal Intensity Increases the Performance of Blind Source Separation on Brain Imaging Data
abstract
When analyzing brain imaging data, blind source separation (BSS) techniques critically depend on the level of dimensional reduction. If the reduction level is too slight, the BSS model would be overfitted and become unavailable. Thus, the reduction level must be set relatively heavy. This approach risks discarding useful information and crucially limits the performance of BSS techniques. In this study, a new BSS method that can work well even at a slight reduction level is presented. We proposed the concept of "signal intensity" which measures the significance of the source. Only picking the sources with significant intensity, the new method can avoid the overfitted solutions which are nonexistent artifacts. This approach enables the reduction level to be set slight and retains more useful dimensions in the preliminary reduction. Comparisons between the new and conventional algorithms were performed on both simulated and real data.
Ming Li 0028, Yadong Liu 0001, Fanglin Chen 0001, Dewen Hu
IEEE Trans. Medical Imaging2
2011 Cerebral Artery-Vein Separation Using 0.1-Hz Oscillation in Dual-Wavelength Optical Imaging
abstract
We present a novel artery-vein separation method using 0.1-Hz oscillation at two wavelengths with optical imaging of intrinsic signals (OIS). The 0.1-Hz oscillation at a green light wavelength of 546 nm exhibits greater amplitude in arteries than in veins and is primarily caused by vasomotion, whereas the 0.1-Hz oscillation at a red light wavelength of 630 nm exhibits greater amplitude in veins than in arteries and is primarily caused by changes of deoxyhemoglobin concentration. This spectral feature enables cortical arteries and veins to be segmented independently. The arteries can be segmented on the 0.1-Hz amplitude image at 546 nm using matched filters of a modified dual Gaussian model combining with a single Gaussian model. The veins are a combination of vessels segmented on both amplitude images at the two wavelengths using multiscale matched filters of single Gaussian model. Our method can separate most of the thin arteries and veins from each other, especially the thin arteries with low contrast in raw gray images. In vivo OIS experiments demonstrate the separation ability of the 0.1-Hz based segmentation method in cerebral cortex of eight rats. Two validation studies were undertaken to evaluate the performance of the method by quantifying the arterial and venous length based on a reference standard. The results indicate that our 0.1-Hz method is very effective in separating both large and thin arteries and veins regardless of vessel crossover or overlapping to great extent in comparison with previous methods.
Dewen Hu, Yadong Liu 0001, Ming Li 0028
IEEE Trans. Medical Imaging3
2009 Local region structured noise reduction for cortical optical imaging
Yadong Liu 0001, Dewen Hu, Zongtan Zhou, Fayi Liu
Neurocomputing1
2005 A novel method for spatio-temporal pattern analysis of brain fMRI data
abstract
A novel data processing procedure for fMRI was suggested in this paper, by which spatial and temporal characteristics of stimuli-induced signal dynamic responses can be investigated simultaneously. First the multitaper spectral estimation was utilized to estimate the spectrum of each voxel; the significance of the line frequency components at the interested frequency was tested to detect the task-related cortex areas; the temporal independent component analysis (tICA) was then applied to the activated voxels to obtain stimuli-induced signal dynamic responses. The advantages of this procedure are: few assumptions are needed for the cerebral hemodynamics and spatial distribution of task-related areas, problems which often appear in tICA analysis of fMRI data, such as the lack of stability, reliability and robustness, are overcome by the suggested method.
Yadong Liu 0001, Zongtan Zhou, Dewen Hu, Lirong Yan, Changlian Tan, Daxing Wu, Shuqiao Yao
Sci. China Ser. F Inf. Sci.1