VLDB 2026 Research / reviewers in the wild / expert
Mengyun Wu
dblp:244/3409
· DBLP profile ↗
8ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0002-0970-4712ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bilevel Network Learning via Hierarchically Structured SparsityabstractAccurate network estimation serves as the cornerstone for understanding complex systems across scientific domains, from decoding gene regulatory networks in systems biology to identifying social relationship patterns in computational sociology. Modern applications demand methods that simultaneously address two critical challenges: capturing nonlinear dependencies between variables and reconstructing inherent hierarchical structures where higher-level entities coordinate lower-level components (e.g., functional pathways organizing gene clusters). Traditional Gaussian graphical models fundamentally fail in these aspects due to their restrictive linear assumptions and flat network representations. We propose NNBLNet, a neural network-based learning framework for bi-level network inference. The core innovation lies in hierarchical selection layers that enforce structural consistency between high-level coordinator groups and their constituent low-level connections via adaptive sparsity constraints. This architecture is integrated with a compositional neural network architecture that learn cross-level association patterns through constrained nonlinear transformations, explicitly preserving hierarchical dependencies while overcoming the representational limitations of linear methods. Crucially, we establish formal theoretical guarantees for the consistent recovery of both high-level connections and their internal low-level structures under general statistical regimes. Extensive validation demonstrates NNBLNet's effectiveness across synthetic and real-world scenarios, achieving superior F1 scores compared to competitive methods and particularly beneficial for complex systems analysis through its interpretable bi-level structure discovery. Jingyuan Yang 0021, Shuangge Ma, Mengyun Wu |
NeurIPS | 4 |
| 2022 | Single Image Dehazing Based on Generative Adversarial Networks
Mengyun Wu |
ICIC (2) | 1 |
| 2022 | Analysis of cancer omics data: a selective review of statistical techniquesabstractCancer is an omics disease. The development in high-throughput profiling has fundamentally changed cancer research and clinical practice. Compared with clinical, demographic and environmental data, the analysis of omics data-which has higher dimensionality, weaker signals and more complex distributional properties-is much more challenging. Developments in the literature are often 'scattered', with individual studies focused on one or a few closely related methods. The goal of this review is to assist cancer researchers with limited statistical expertise in establishing the 'overall framework' of cancer omics data analysis. To facilitate understanding, we mainly focus on intuition, concepts and key steps, and refer readers to the original publications for mathematical details. This review broadly covers unsupervised and supervised analysis, as well as individual-gene-based, gene-set-based and gene-network-based analysis. We also briefly discuss 'special topics' including interaction analysis, multi-datasets analysis and multi-omics analysis. Chenjin Ma, Mengyun Wu, Shuangge Ma |
Briefings Bioinform. | 2 |
| 2022 | Network-based cancer heterogeneity analysis incorporating multi-view of prior informationabstractMOTIVATION: Cancer genetic heterogeneity analysis has critical implications for tumour classification, response to therapy and choice of biomarkers to guide personalized cancer medicine. However, existing heterogeneity analysis based solely on molecular profiling data usually suffers from a lack of information and has limited effectiveness. Many biomedical and life sciences databases have accumulated a substantial volume of meaningful biological information. They can provide additional information beyond molecular profiling data, yet pose challenges arising from potential noise and uncertainty. RESULTS: In this study, we aim to develop a more effective heterogeneity analysis method with the help of prior information. A network-based penalization technique is proposed to innovatively incorporate a multi-view of prior information from multiple databases, which accommodates heterogeneity attributed to both differential genes and gene relationships. To account for the fact that the prior information might not be fully credible, we propose a weighted strategy, where the weight is determined dependent on the data and can ensure that the present model is not excessively disturbed by incorrect information. Simulation and analysis of The Cancer Genome Atlas glioblastoma multiforme data demonstrate the practical applicability of the proposed method. AVAILABILITY AND IMPLEMENTATION: R code implementing the proposed method is available at https://github.com/mengyunwu2020/PECM. The data that support the findings in this paper are openly available in TCGA (The Cancer Genome Atlas) at https://portal.gdc.cancer.gov/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yang Li 0070, Shaodong Xu, Shuangge Ma, Mengyun Wu |
Bioinform. | 4 |
| 2021 | Vertical integration methods for gene expression data analysisabstractGene expression data have played an essential role in many biomedical studies. When the number of genes is large and sample size is limited, there is a 'lack of information' problem, leading to low-quality findings. To tackle this problem, both horizontal and vertical data integrations have been developed, where vertical integration methods collectively analyze data on gene expressions as well as their regulators (such as mutations, DNA methylation and miRNAs). In this article, we conduct a selective review of vertical data integration methods for gene expression data. The reviewed methods cover both marginal and joint analysis and supervised and unsupervised analysis. The main goal is to provide a sketch of the vertical data integration paradigm without digging into too many technical details. We also briefly discuss potential pitfalls, directions for future developments and application notes. Mengyun Wu, Huangdi Yi, Shuangge Ma |
Briefings Bioinform. | 1 |
| 2021 | GEInter: an R package for robust gene-environment interaction analysisabstractSUMMARY: For understanding complex diseases, gene-environment (G-E) interactions have important implications beyond main G and E effects. Most of the existing analysis approaches and software packages cannot accommodate data contamination/long-tailed distribution. We develop GEInter, a comprehensive R package tailored to robust G-E interaction analysis. For both marginal and joint analysis, for data without and with missingness, for continuous and censored survival responses, it comprehensively conducts identification, estimation, visualization and prediction. It can fill an important gap in the existing literature and enjoy broad applicability. AVAILABILITY AND IMPLEMENTATION: TCGA data is analyzed as demonstrating examples. It is well known that such data is publicly available https://cran.r-project.org/web/packages/GEInter/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mengyun Wu, Xing Qin, Shuangge Ma |
Bioinform. | 1 |
| 2020 | NCutYX: a package for clustering analysis of multilayer omics dataabstractSUMMARY: Multilayer omics profiling has become a major venue for understanding complex diseases. We develop NCutYX, an R package for clustering analysis of multilayer omics data. The package and methods jointly analyze multiple layers of omics measurements and effectively accommodate their regulations. They systematically conduct a series of analysis based on the normalized cut technique, including the clusterings of subjects and omics measurements and biclustering. The package can be valuable for its timely context, novel methods, and comprehensiveness. AVAILABILITY: https://cran.r-project.org/web/packages/NCutYX/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sebastian J. Teran Hidalgo, Mengyun Wu, Shuangge Ma |
Bioinform. | 2 |
| 2019 | Robust genetic interaction analysisabstractFor the risk, progression, and response to treatment of many complex diseases, it has been increasingly recognized that genetic interactions (including gene-gene and gene-environment interactions) play important roles beyond the main genetic and environmental effects. In practical genetic interaction analyses, model mis-specification and outliers/contaminations in response variables and covariates are not uncommon, and demand robust analysis methods. Compared with their nonrobust counterparts, robust genetic interaction analysis methods are significantly less popular but are gaining attention fast. In this article, we provide a comprehensive review of robust genetic interaction analysis methods, on their methodologies and applications, for both marginal and joint analysis, and for addressing model mis-specification as well as outliers/contaminations in response variables and covariates. Mengyun Wu, Shuangge Ma |
Briefings Bioinform. | 1 |