Yaochen Xu

dblp:192/4800 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0002-5039-2781ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Theory of computation · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Spatial histology and gene-expression representation and generative learning via online self-distillation contrastive learning
abstract
Spatial transcriptomics quantifies spatial molecular profiles alongside histology, enabling computational prediction of spatial gene expression distribution directly from whole slide images. Inspired by image-to-text alignment and generation, we introduce Magic, a self-training contrastive learning model designed for histology-to-gene expression prediction. Magic (i) employs contrastive learning to derive shared embeddings for histology and gene expression while utilizing a momentum-based module to generate pseudo-targets to reduce the impact of noise; and (ii) leverages a transformer-based decoder to predict the expression of 300 genes based on histological features. Trained on 75 760 spots from 56 breast cancer slices and validated on 11 026 spots from five independent slices, Magic outperforms existing methods in aligning and generating histology-gene expression data, achieving a 10% improvement over the second-best approach. Furthermore, Magic demonstrates robust generalization, effectively predicting gene expression in colorectal cancer samples and The Cancer Genome Atlas (TCGA) datasets through zero-shot learning. Notably, Magic's predicted gene expression captures interpatient differences, highlighting its strong potential for clinical applications.
Qianyi Yan, Jiangnan Cui, Jianming Rong, Jingsong Zhang, Pingting Gao, Yaochen Xu, Fufang Qiu, Chunman Zuo
Briefings Bioinform.7
2025 Action-Driven Semantic Representation and Aggregation for Video Captioning
abstract
Video captioning, a challenging task that entails generating natural language descriptions of visual content, often fails to effectively grasp the essence of action semantics. To harness the power of action detection to facilitate a deeper understanding of the video content, we propose an action-driven method, named Hierarchical Semantic Representation and Aggregation (HSRA) network. This method explicitly exploits action clues with a hierarchical semantic representation module, which models visual semantics in a three-level structure: “object-action-event”. By employing learnable action queries, our approach injects extensive action semantics into the model, thereby enabling more accurate and context-rich captions. To further enhance semantic alignment and understanding, we introduce a semantic aggregation composed of a semantic interaction module and a semantic refinement module. This component facilitates the alignment of semantics across different levels and emphasizes key information, ultimately leading to significant improvements in semantic consistency between the video and generated captions. We performed extensive evaluations on two well-established public datasets, MSVD and MSR-VTT, and the findings consistently demonstrate that our proposed HSRA network outperforms contemporary state-of-the-art methods.
Tingting Han 0003, Yaochen Xu, Jun Yu 0002, Zhou Yu 0001, Sicheng Zhao
IEEE Trans. Circuits Syst. Video Technol.2
2023 TRAmHap: accurate prediction of transcriptional activity from DNA methylation haplotypes in bisulfite-sequencing data
abstract
Deoxyribonucleic acid (DNA) methylation (DNAm) is an important epigenetic mechanism that plays a role in chromatin structure and transcriptional regulation. Elucidating the relationship between DNAm and gene expression is of great importance for understanding its role in transcriptional regulation. The conventional approach is to construct machine-learning-based methods to predict gene expression based on mean methylation signals in promoter regions. However, this type of strategy only explains about 25% of gene expression variation, and hence is inadequate in elucidating the relationship between DNAm and transcriptional activity. In addition, using mean methylation as input features neglects the heterogeneity of cell populations that can be reflected by DNAm haplotypes. We here developed TRAmaHap, a novel deep-learning framework that predicts gene expression by utilizing the characteristics of DNAm haplotypes in proximal promoters and distal enhancers. Using benchmark data of human and mouse normal tissues, TRAmHap shows much higher accuracy than existing machine-learning based methods, by explaining 60~80% of gene expression variation across tissue types and disease conditions. Our model demonstrated that gene expression can be accurately predicted by DNAm patterns in promoters and long-range enhancers as far as 25 kb away from transcription start site, especially in the presence of intra-gene chromatin interactions.
Hanwen Zhu, Kangwen Cai, Leiqin Liu, Yaochen Xu, Xiaoqi Zheng
Briefings Bioinform.7
2022 A Mechanical Method for Isolating Locally Optimal Points of Certain Radical Functions
Zhenbing Zeng, Yaochen Xu, Zhengfeng Yang
CASC2
2021 The DNA methylation haplotype (mHap) format and mHapTools
abstract
SUMMARY: Bisulfite sequencing (BS-seq) is currently the gold standard for measuring genome-wide DNA methylation profiles at single-nucleotide resolution. Most analyses focus on mean CpG methylation and ignore methylation states on the same DNA fragments [DNA methylation haplotypes (mHaps)]. Here, we propose mHap, a simple DNA mHap format for storing DNA BS-seq data. This format reduces the size of a BAM file by 40- to 140-fold while retaining complete read-level CpG methylation information. It is also compatible with the Tabix tool for fast and random access. We implemented a command-line tool, mHapTools, for converting BAM/SAM files from existing platforms to mHap files as well as post-processing DNA methylation data in mHap format. With this tool, we processed all publicly available human reduced representation bisulfite sequencing data and provided these data as a comprehensive mHap database. AVAILABILITY AND IMPLEMENTATION: https://jiantaoshi.github.io/mHap/index.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yuhao Dan, Yaochen Xu, Xiaoqi Zheng
Bioinform.3
2017 Searching approximate global optimal Heilbronn configurations of nine points in the unit square via GPGPU computing
Liangyu Chen 0001, Yaochen Xu, Zhenbing Zeng
J. Glob. Optim.2
2016 CMIP: a software package capable of reconstructing genome-wide regulatory networks using gene expression data
abstract
BACKGROUND: A gene regulatory network (GRN) represents interactions of genes inside a cell or tissue, in which vertexes and edges stand for genes and their regulatory interactions respectively. Reconstruction of gene regulatory networks, in particular, genome-scale networks, is essential for comparative exploration of different species and mechanistic investigation of biological processes. Currently, most of network inference methods are computationally intensive, which are usually effective for small-scale tasks (e.g., networks with a few hundred genes), but are difficult to construct GRNs at genome-scale. RESULTS: Here, we present a software package for gene regulatory network reconstruction at a genomic level, in which gene interaction is measured by the conditional mutual information measurement using a parallel computing framework (so the package is named CMIP). The package is a greatly improved implementation of our previous PCA-CMI algorithm. In CMIP, we provide not only an automatic threshold determination method but also an effective parallel computing framework for network inference. Performance tests on benchmark datasets show that the accuracy of CMIP is comparable to most current network inference methods. Moreover, running tests on synthetic datasets demonstrate that CMIP can handle large datasets especially genome-wide datasets within an acceptable time period. In addition, successful application on a real genomic dataset confirms its practical applicability of the package. CONCLUSIONS: This new software package provides a powerful tool for genomic network reconstruction to biological community. The software can be accessed at http://www.picb.ac.cn/CMIP/ .
Guangyong Zheng, Yaochen Xu, Zhi-Ping Liu, Luonan Chen, Xin-Guang Zhu
BMC Bioinform.2