Gongshun Yang

dblp:388/0546 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 67% Computational science and engineering · 33%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
genomics
1.012026
Robust prioritization of genomic features with stability selection · Bioinform. 2026
Bioinformatics and computational biology
statistical genetics
1.012026
Robust prioritization of genomic features with stability selection · Bioinform. 2026
Computational science and engineering › model selection
variable selection
1.012026
Robust prioritization of genomic features with stability selection · Bioinform. 2026

Methods — techniques the papers use, named apart from their topics

stability selection · 1.0robust regression · 1.0LAD LASSO · 1.0
YearPublicationVenuePosition
2026 Robust prioritization of genomic features with stability selection
abstract
MOTIVATION: The heterogeneity of complex diseases including cancer leads to heavy-tailed distributions in the disease traits. In such settings, non-robust variable selection methods are inherently susceptible to data contamination and can yield unstable or misleading results. This vulnerability becomes more severe for recently proposed approaches that introduce pseudo-features as negative controls, as these methods further amplify the curse of dimensionality by expanding the genotype matrix in the presence of outliers and high-dimensional genomic features. RESULTS: We develop a robust variable selection framework with stability selection to prioritize genomic features in the presence of contamination. In contrast to existing approaches that rely on pseudo-features for error control, the proposed method achieves double robustness. First, it adopts least absolute deviation (LAD) LASSO to ensure robustness against outliers and heavy-tailed errors in disease traits. Second, it avoids augmenting the genotype matrix with pseudo-features, thereby mitigating the curse of dimensionality that is particularly problematic in high-dimensional genomic data. The proposed method has been extensively evaluated in simulation studies to demonstrate its effectiveness over multiple competing methods for variable selection. In addition, we have applied the proposed method and competing approaches to two real-data case studies: the The Cancer Genome Atlas (TCGA) Skin Cutaneous Melanoma (SKCM) dataset and an eQTL dataset. The results demonstrate that the proposed method achieves superior performance by identifying genomic features with higher reproducibility. AVAILABILITY AND IMPLEMENTATION: The source code for implementing the proposed methods is publicly available at https://github.com/cenwu/RSS with an archival DOI https://doi.org/10.6084/m9.figshare.32306883.
Gongshun Yang, Cen Wu
Bioinform.1