Yichao Ma

dblp:17/1334 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Generative modeling · 61% Efficient and distributed learning · 30% Deep learning architectures and training · 9%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
0.912025
Efficient Diffusion Models: A Comprehensive Survey From Principles to Practices · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Machine learning › Generative modeling › diffusion model
efficient diffusion model
0.912025
Efficient Diffusion Models: A Comprehensive Survey From Principles to Practices · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
Efficient Diffusion Models: A Comprehensive Survey From Principles to Practices · IEEE Trans. Pattern Anal. Mach. Intell. 2025

Methods — techniques the papers use, named apart from their topics

model deployment · 0.9fast inference · 0.9
YearPublicationVenuePosition
2025 EventLens: Enhancing Visual Commonsense Reasoning by Leveraging Event-Aware Pretraining and Cross-modal Linking
abstract
Visual Commonsense Reasoning (VCR) is a cognitive task, challenging models to answer visual questions, and to explain the rationale behind their answers. While Large Language Models (LLMs) offer potential for this task, VCR’s complex scenes require specialized approaches to activate their commonsense reasoning abilities, as existing Multimodal LLMs struggle with VCR’s visual events and unique reference tags. To address these challenges, we propose EventLens, which enhances VCR through Event-Aware Pretraining and Cross-modal Linking in Supervised Fine-tuning. First, we introduce a new pretraining stage that emulates human cognitive processes to improve LLM comprehension of complex scenarios. Second, during supervised fine-tuning, we leverage reference tags to explicitly bridge RoI features with text, maintaining semantic integrity across modalities. Additionally, instruct prompts and task-specific adapters help integrate LLMs’ knowledge with new commonsense reasoning. Experimental results demonstrate competitive performance with state-of-the-art methods and ablation studies verify the effectiveness of proposed EventLens.
Zhihuan Yu, Yichao Ma, Guohui Li 0001, Zhong Yang 0004
ICASSP3
2025 Efficient Diffusion Models: A Comprehensive Survey From Principles to Practices
abstract
As one of the most popular and sought-after generative models in recent years, diffusion models have sparked the interests of many researchers and steadily shown excellent advantage in various generative tasks such as image synthesis, video generation, bioinformatics engineering, 3D scene rendering and multimodal generation, relying on their dense theoretical principles and reliable application practices. The remarkable success of these recent efforts on diffusion models comes largely from progressive design principles and efficient architecture, training, inference, and deployment methodologies. However, there has not been a comprehensive and in-depth review to summarize these principles and practices to help the rapid understanding and application of diffusion models. In this survey, we provide a new efficiency-oriented perspective on these existing efforts, which mainly focuses on the profound principles and efficient practices in architecture designs, model training, fast inference and reliable deployment, to guide further theoretical research, algorithm migration and model application for new scenarios in a reader-friendly way.
Zhiyuan Ma 0005, Yuzhu Zhang, Guoli Jia, Yichao Ma, Gaofeng Liu, Ning Ding 0002, Jianjun Li 0010, Bowen Zhou 0002
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Novel m7G-related lncRNA signature for predicting overall survival in patients with gastric cancer
abstract
Presenting with a poor prognosis, gastric cancer (GC) remains one of the leading causes of disease and death worldwide. Long non-coding RNAs (lncRNAs) regulate tumor formation and have been long used to predict tumor prognosis. N7-methylguanosine (m7G) is the most prevalent RNA modification. m7G-lncRNAs regulate GC onset and progression, but their precise mechanism in GC is unclear. The objective of this research was the development of a new m7G-related lncRNA signature as a biomarker for predicting GC survival rate and guiding treatment. The Cancer Genome Atlas database helped extract gene expression data and clinical information for GC. Pearson correlation analysis helped point out m7G-related lncRNAs. Univariate Cox analysis helped in identifying m7G-related lncRNA with predictive capability. The Lasso-Cox method helped point out seven lncRNAs for the purpose of establishing an m7G-related lncRNA prognostic signature (m7G-LPS), followed by the construction of a nomogram. Kaplan-Meier analysis, univariate and multivariate Cox regression analysis, calibration plot of the nomogram model, receiver operating characteristic curve and principal component analysis were utilized for the verification of the risk model's reliability. Furthermore, q-PCR helped verify the lncRNAs expression of m7G-LPS in-vitro. The study subjects were classified into high and low-risk groups based on the median value of the risk score. Gene enrichment analysis confirmed the constructed m7G-LPS' correlation with RNA transcription and translation and multiple immune-related pathways. Analysis of the clinicopathological features revealed more progressive features in the high-risk group. CIBERSORT analysis showed the involvement of m7G-LPS in immune cell infiltration. The risk score was correlated with immune checkpoint gene expression, immune cell and immune function score, immune cell infiltration, and chemotherapy drug sensitivity. Therefore, our study shows that m7G-LPS constructed using seven m7G-related lncRNAs can predict the survival time of GC patients and guide chemotherapy and immunotherapy regimens as biomarker.
Yiqun Liao, Yuji Chen, Yichao Ma, Daorong Wang
BMC Bioinform.6
2007 Usage-Oriented Performance Evaluation for Text Localization Algorithms
abstract
The localization of texts in image/video is the first step in a text processing system. Its effect will do great impact on the following processing steps. Although many studies have been done on text localization algorithms, there is not a universally accepted performance evaluation method. In this paper we propose two sets of metrics to evaluate the performance of text localization algorithms in different usage conditions. The metrics also consider the text distribution characteristics, and the difficulties of the underlying task. Some experiments on the proposed metrics are also given.
Yichao Ma, Chunheng Wang, Baihua Xiao, Ruwei Dai
ICDAR1