EDBT 2026 Demo / reviewers in the wild / expert
Wen Zhu
dblp:51/3657
· DBLP profile ↗
27ranked-venue papers
2as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Decomposing the Fairness Gap: Disentangling Data Scarcity from Validity Bias in AI-Driven Formative Assessment
Wen Zhu |
AIED (6) | 2 |
| 2026 | Artificial intelligence-guided design of manufactured sand concrete with targeted performance
Qing-feng Liu, Qing Xiang Xiong, Kanghao Tan, Wen Zhu, Weibo Shen |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | Dual-Prompt Learning with Cross-Modal Decoders for Few-Shot Whole Slide Image ClassificationabstractFew-shot learning offers a promising solution for computational pathology by alleviating the reliance on large an-notated datasets, but faces challenges from the high redundancy in whole slide images and underutilized cross-modal knowledge. Existing methods typically use foundation models only for pre-liminary feature extraction while employing fixed or single-level prompts that lack multi-scale pathological representation. To address these limitations, we propose a Hierarchical Vision-Text Prompt (H- VTP) framework that enables multi-level cross-modal interaction through GPT-4 generated Local Instance Prompts for patch-level morphological details and Global Semantic Prompts for slide-level diagnostic context. A dual-branch decoding mech-anism with Text-Guided-Patch Decoder and Patch-Augmented-Text Decoder facilitates closed-loop vision-text fusion, while a parameter-efficient adaptation strategy trains only lightweight prompts and adapters. Extensive experiments on three cancer subtype datasets demonstrate the superiority of H-VTP in few-shot WSI classification, confirming its effectiveness for clinical applications. Bingbing Zhang 0001, Wen Zhu, Bin Liu 0040, Jianxin Zhang 0001, Qiang Zhang 0008 |
BIBM | 3 |
| 2025 | Two efficient beamforming methods for hybrid IRS-aided AF relay wireless networks
Qingbo Li, Wen Zhu, Feng Shu 0002, Mengxing Huang, Fuhui Zhou, Riqing Chen, Cunhua Pan, Yongpeng Wu 0001, Jiangzhou Wang |
Sci. China Inf. Sci. | 3 |
| 2024 | Second-Order Self-Supervised Learning for Breast Cancer ClassificationabstractMultiple instance learning (MIL), which adopts instance-level feature encoders in a weakly supervised manner, efficiently classifies cancers on whole slide images (WSis). Nevertheless, the lack of fine-grained annotation samples still limits the performance of existing MIL models. Therefore, we propose a second-order self-supervised learning method integrated with the MIL structure, called SoS2MIL, for breast cancer classification. Specifically, SoS2MIL explores intrinsic structural information within WSI instances using self-supervised learning that does not require label information to identify discriminative instances. Then, it improves the instance-level classifier concentration by introducing second-order pooling in a visual perception module. Besides, a grouping strategy is proposed to reduce the computational cost of second-order calculations and effectively approximate feature distributions by combining first-order features. Experimental results on public and private datasets verify its effectiveness and competitiveness with the state-of-the-art MIL works. Haitao Yao, Wen Zhu |
ICME | 4 |
| 2024 | Identifying Associations Between Small Nucleolar RNAs and Diseases via Graph Convolutional Network and Attention MechanismabstractResearch has shown that small nucleolar RNAs (snoRNAs) play crucial roles in various biological processes, and understanding disease pathogenesis by studying their relationship with diseases is beneficial. Currently, known associations are insufficient, and conventional biological experiments are costly and time-consuming. Therefore, developing efficient computational methods is crucial for identifying potential snoRNA-disease associations. In this paper, a method to identify snoRNA-disease associations based on graph convolutional network and multi-view graph attention mechanism (GCASDA) is proposed. Firstly, the similarity matrices of snoRNAs and diseases are calculated based on biological entity-related information, and the weights of the edges between the snoRNA nodes and the disease nodes are supplemented by random forest. Then two homogeneous graphs and one heterogeneous graph are constructed. Subsequently, different types of embedded features are extracted from the graphs using specific graph convolutional network structure and integrated through a multi-view graph attention mechanism to obtain node embedded feature representations. Finally, for each pair of nodes, in addition to their global features, node interaction features are passed together to a multilayer perceptron neural network (MLP) to identify snoRNA-disease associations. Experimental results show that GCASDA achieves 0.9356 and 0.9294 in AUC and AUPR, respectively, and significantly outperformed other state-of-the-art methods on the basis of different evaluation metrics. Furthermore, the case study could further demonstrate the realistic feasibility of GCASDA. Shuchen Liu, Wen Zhu, Shaoyou Yu, Fang-Xiang Wu |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | SAMMS: Multi-modality Deep Learning with the Foundation Model for the Prediction of Cancer Patient SurvivalabstractCancer survival prediction is pivotal in tailoring individualized treatment strategies and guiding clinician decision-making. Yet, existing methodologies grapple with efficiently harnessing the intricate distribution of medical data spanning various modalities. In response, we present SAMMS, an advanced multi-omics multimodal deep learning framework tailored for survival prediction. SAMMS leverages the robust image segmentation model, "Segment Anything" to adeptly characterize pathological images. This prowess is further enhanced by integrating multi-omics data and clinical insights, facilitating holistic modeling across a diverse modal spectrum. The framework weaves a modality-specific subnetwork with a cross-modality common subnetwork, meticulously capturing intra-modality nuances and inter-modality correlations. SAMMS eclipsed its contemporaries by delivering remarkable performance on TCGA’s LGG and KIRC tumor datasets. A battery of analyses underscored SAMMS’s unparalleled capability to distill multifaceted insights from multimodal datasets, yielding richer and more integrative multimodal representations. Such strides promise significant advancements in cancer survival analytics, bolstering the precision and efficacy of patient-centric treatments, disease oversight, and clinical decision processes. Wen Zhu, Shanling Nie, Hai Yang 0002 |
BIBM | 1 |
| 2023 | A Robust Oversampling Approach for Class Imbalance Problem With Small DisjunctsabstractClass imbalance is one of the important challenges for machine learning because of it's learning to bias toward the majority classes. The oversampling method is a fundamental imbalance-learning technique with many real-world applications. However, when the small disjuncts problem occurs, how to effectively avoiding the negative oversampling results rather than using clusters previously, remains a challenging task. Thus, this study introduces a disjuncts-robust oversampling (DROS) method. The novel method shows that the data filling of new synthetic samples to the minority class areas in data space can be thought of as the searchlight illuminating with light cones to the restricted areas in real life. In the first step, DROS computes a series of light-cone structures that is first started from the inner minority class area, then passes through the boundary minority class area, last is stopped by the majority class area. In the second step, DROS generates new synthetic samples in those light-cone structures. Experiments considering both real-world and 2D emulational datasets demonstrate that our method outperforms the current state-of-the-art oversampling methods and suggest that our method is able to deal with the small disjuncts. Bo Liao 0001, Wen Zhu, Junlin Xu |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2022 | Attention multiple instance learning with Transformer aggregation for breast cancer whole slide image classificationabstractRecently, attention-based multiple instance learning (MIL) methods have received more concentration in histopathology whole slide image (WSI) applications. However, existing attention-based MIL methods rarely consider the cross-channel information interaction of pathology images when identifying discriminant patches. Additionally, they also have limitations on capturing the correlation between different discriminant instances for the bag-level classification. To address these challenges, we present a novel attention-based MIL model (AMIL-Trans) for breast cancer WSI classification. AMIL-Trans first embeds the efficient channel attention to realize the cross-channel interaction of pathology images, thus computing more robust features for instance selection without introducing too much computation cost. Then, it leverages vision Transformer encoder to directly aggregate selected instance features for better bag-level prediction, which effectively considers the correlation between different discriminant instances. Experiment results illustrate that AMIL-Trans respectively achieves its optimal AUC of 94.27% and 84.22% on the Camelyon-16 dataset and MSK external validation dataset, demonstrating the competitive performance compared with state-of-the-art MIL methods on breast cancer WSI classification task. The code will be available at https://github.con CunqiaoHou/AMIL-Trans. Jianxin Zhang 0001, Cunqiao Hou, Wen Zhu, Ying Zou 0015, Qiang Zhang 0008 |
BIBM | 3 |
| 2022 | Minority Sub-Region Estimation-Based Oversampling for Imbalance LearningabstractClass imbalance problem that characterized with the skew distribution towards the majority arises as one challenge in recent years. Many oversampling techniques have been proposed to cope with this problem and some of them combine the oversampling procedure with the clustering algorithm which guaranteeing new synthetic samples being generated in clusters. However far-away samples but with the same minority sub-region are generally clustered into different groups owing to the characteristic of clustering algorithm itself. Therefore, the following oversampling procedure is mostly carried in incomplete minority sub-regions that synthetic samples not well cover the integral minority region. And to our best knowledge, none of existing algorithm is designed to directly estimate minority sub-regions for class imbalance problem. Thus, one new grouping algorithm, named Direction Distribution-based Minority Sub-region Estimation (DDMSE), is first proposed. The new algorithm exploits the intuitive observation, that the minority with the same sub-region almost distribute within the same direction when compared to other majority, to estimate minority sub-regions that tactfully ignoring negative impacts brought by the distance factor like in clustering algorithms. Finally, new synthetic samples are generated in those minority sub-regions. And experimental results on real-world datasets show the comparable performance with other state-of-the-art oversampling methods. Bo Liao 0001, Wen Zhu |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2021 | CMF-Impute: an accurate imputation tool for single cell RNA-seq data
Junlin Xu, Wen Zhu, Jialiang Yang |
Bioinform. | 4 |
| 2020 | CMF-Impute: an accurate imputation tool for single-cell RNA-seq dataabstractMOTIVATION: Single-cell RNA-sequencing (scRNA-seq) technology provides a powerful tool for investigating cell heterogeneity and cell subpopulations by allowing the quantification of gene expression at single-cell level. However, scRNA-seq data analysis remains challenging because of various technical noises such as dropout events (i.e. excessive zero counts in the expression matrix). RESULTS: By taking consideration of the association among cells and genes, we propose a novel collaborative matrix factorization-based method called CMF-Impute to impute the dropout entries of a given scRNA-seq expression matrix. We test CMF-Impute and compare it with the other five state-of-the-art methods on six popular real scRNA-seq datasets of various sizes and three simulated datasets. For simulated datasets, CMF-Impute outperforms other methods in imputing the closest dropouts to the original expression values as evaluated by both the sum of squared error and Pearson correlation coefficient. For real datasets, CMF-Impute achieves the most accurate cell classification results in spite of the choice of different clustering methods like SC3 or T-SNE followed by K-means as evaluated by both adjusted rand index and normalized mutual information. Finally, we demonstrate that CMF-Impute is powerful in reconstructing cell-to-cell and gene-to-gene correlation, and in inferring cell lineage trajectories. AVAILABILITY AND IMPLEMENTATION: CMF-Impute is written as a Matlab package which is available at https://github.com/xujunlin123/CMFImpute.git. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Junlin Xu, Wen Zhu, Jialiang Yang |
Bioinform. | 4 |
| 2019 | A new resource allocation strategy based on the relationship between subproblems for MOEA/D
Peng Wang 0035, Wen Zhu, Haihua Liu, Bo Liao 0002, Xiaohui Wei 0001, Siqi Ren, Jialiang Yang |
Inf. Sci. | 2 |
| 2019 | Selection-based resampling ensemble algorithm for nonstationary imbalanced stream data learning
Siqi Ren, Wen Zhu, Bo Liao 0002, Peng Wang 0035, Keqin Li 0001, Min Chen 0028 |
Knowl. Based Syst. | 2 |
| 2019 | Scalable One-Pass Self-Representation Learning for Hyperspectral Band SelectionabstractFor applications based on hyperspectral imagery (HSI), selecting informative and representative bands without the degradation of performance is a challenging task in the context of big data. In this paper, an unsupervised band selection method, scalable one-pass self-representation learning (SOP-SRL), is proposed to address this problem by processing data in a streaming fashion without storing the entire data. SOP-SRL embeds band selection into a scalable self-representation learning, which is formulated as an adaptive linear combination of regression-based loss functions, with the row-sparsity constraint. To further enhance the representativeness of bands, the local similarity between samples constructed by the selected bands is dynamically measured by means of graph-based regularization term in the embedded space. Moreover, a cache with memory function that reflects the quality of bands in the historical data is designed to keep the consistency between data coming at different times and guide subsequent band selection. An efficient algorithm is developed to optimize the SOP-SRL model. The HSI classification is conducted on three public data sets, and the experimental results validate the superiority of SOP-SRL in terms of performance and time when compared with other state-of-the-art band selection methods. Xiaohui Wei 0001, Wen Zhu, Bo Liao 0002 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | The Gradual Resampling Ensemble for mining imbalanced data streams with concept drift
Siqi Ren, Bo Liao 0002, Wen Zhu, Wei Liu 0245, Keqin Li 0001 |
Neurocomputing | 3 |
| 2018 | Novel architecture for long short-term memory used in question classification
Wen Zhu, Bo Liao 0002, Min Chen 0028, Lei Huang 0018 |
Neurocomputing | 2 |
| 2018 | Knowledge-maximized ensemble algorithm for different types of concept drift
Siqi Ren, Bo Liao 0002, Wen Zhu, Keqin Li 0001 |
Inf. Sci. | 3 |
| 2018 | Matrix-Based Margin-Maximization Band Selection With Data-Driven Diversity for Hyperspectral Image ClassificationabstractFor hyperspectral image classification, high-dimensional spectral features not only increase the computational and storage burden but also degrade the classification accuracy due to the Hughes phenomenon. Band selection is an important technique to solve these issues without destroying the interpretation of data. This paper presents a matrix-based margin-maximization method of band selection with data-driven diversity. In particular, the matrices composed of adjacent pixels in space are fed to the hinge loss function with a row-sparse constraint. This constraint is used to select the expected bands and preserve the spatial structure information simultaneously while maximizing the margin between the classes. In consideration of the continuity of bands in the spectral dimension, a novel regularization term that is continually updated according to the current context is added to promote the differential expression of dissimilar bands. Finally, the one-against-all parallel mechanism is used to learn a coefficient matrix for each class, and class-related bands are then carefully selected by a partitioning strategy (e.g., k-means clustering) based on the learned coefficient matrix. Experiments are conducted on three hyperspectral data sets and four widely used classifiers. The experimental results have shown that our proposed method is superior to several state-of-the-art methods, especially when the number of selected bands is relatively small. Xiaohui Wei 0001, Wen Zhu, Bo Liao 0002 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2017 | Seeksv: an accurate tool for somatic structural variation and virus integration detectionabstractMOTIVATION: Many forms of variations exist in the human genome including single nucleotide polymorphism, small insert/deletion (DEL) (indel) and structural variation (SV). Somatically acquired SV may regulate the expression of tumor-related genes and result in cell proliferation and uncontrolled growth, eventually inducing tumor formation. Virus integration with host genome sequence is a type of SV that causes the related gene instability and normal cells to transform into tumor cells. Cancer SVs and viral integration sites must be discovered in a genome-wide scale for clarifying the mechanism of tumor occurrence and development. RESULTS: In this paper, we propose a new tool called seeksv to detect somatic SVs and viral integration events. Seeksv simultaneously uses split read signal, discordant paired-end read signal, read depth signal and the fragment with two ends unmapped. Seeksv can detect DEL, insertion, inversion and inter-chromosome transfer at single-nucleotide resolution. Different types of sequencing data, such as single-end sequencing data or paired-end sequencing data can accommodate to detect SV. Seeksv develops a rescue model for SV with breakpoints located in sequence homology regions. Results on simulated and real data from the 1000 Genomes Project and esophageal squamous cell carcinoma samples show that seeksv has higher efficiency and precision compared with other similar software in detecting SVs. For the discovery of hepatitis B virus integration sites from probe capture data, the verified experiments show that more than 90% viral integration sequences detected by seeksv are true. AVAILABILITY AND IMPLEMENTATION: seeksv is implemented in C ++ and can be downloaded from https://github.com/qkl871118/seeksv CONTACT: : [email protected] information: Supplementary data are available at Bioinformatics online. Kunlong Qiu, Bo Liao 0002, Wen Zhu, Xuanlin Huang, Xiangtao Chen, Keqin Li 0001 |
Bioinform. | 4 |
| 2017 | Predicting Drug-Target Interactions With Multi-Information FusionabstractIdentifying potential associations between drugs and targets is a critical prerequisite for modern drug discovery and repurposing. However, predicting these associations is difficult because of the limitations of existing computational methods. Most models only consider chemical structures and protein sequences, and other models are oversimplified. Moreover, datasets used for analysis contain only true-positive interactions, and experimentally validated negative samples are unavailable. To overcome these limitations, we developed a semi-supervised based learning framework called NormMulInf through collaborative filtering theory by using labeled and unlabeled interaction information. The proposed method initially determines similarity measures, such as similarities among samples and local correlations among the labels of the samples, by integrating biological information. The similarity information is then integrated into a robust principal component analysis model, which is solved using augmented Lagrange multipliers. Experimental results on four classes of drug-target interaction networks suggest that the proposed approach can accurately classify and predict drug-target interactions. Part of the predicted interactions are reported in public databases. The proposed method can also predict possible targets for new drugs and can be used to determine whether atropine may interact with alpha1B- and beta1- adrenergic receptors. Furthermore, the developed technique identifies potential drugs for new targets and can be used to assess whether olanzapine and propiomazine may target 5HT2B. Finally, the proposed method can potentially address limitations on studies of multitarget drugs and multidrug targets. Lihong Peng, Bo Liao 0002, Wen Zhu, Keqin Li 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2014 | Gene Selection Using Locality Sensitive Laplacian ScoreabstractGene selection based on microarray data, is highly important for classifying tumors accurately. Existing gene selection schemes are mainly based on ranking statistics. From manifold learning standpoint, local geometrical structure is more essential to characterize features compared with global information. In this study, we propose a supervised gene selection method called locality sensitive Laplacian score (LSLS), which incorporates discriminative information into local geometrical structure, by minimizing local within-class information and maximizing local between-class information simultaneously. In addition, variance information is considered in our algorithm framework. Eventually, to find more superior gene subsets, which is significant for biomarker discovery, a two-stage feature selection method that combines the LSLS and wrapper method (sequential forward selection or sequential backward selection) is presented. Experimental results of six publicly available gene expression profile data sets demonstrate the effectiveness of the proposed approach compared with a number of state-of-the-art gene selection methods. Bo Liao 0002, Wei Liang 0005, Wen Zhu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2013 | Picocell-density based energy-saving for QoS provisioning in heterogeneous networksabstractIn order to reduce the energy consumption, we propose a novel Macro Base Station (MBS) sleep scheme based on the density of Pico Base Station (PBS) including three sleep approaches for current heterogeneous networks (het-net). Our scheme consists of two parts. In the first part, by setting thresholds for PBS density according to the total coverage of PBSs in a macro-coverage, we divide the network into three scenarios where PBSs are deployed densely, sparsely, and commonly correspondingly. In the second part, the MBS chooses one sleep approach of our scheme by comparing its PBS density with the thresholds. In each approach, minimizing the system energy consumption and meeting the QoS requirements are both considered as our goals. Particularly, the offset value is designed to guarantee blocking probability, especially when the neighboring MBSs are in charge of taking over the migrated users from the sleep MBS. Finally, simulation results demonstrate that the proposed scheme is much more efficient than the other existing schemes in terms of the power consumption and the blocking probability. Li Wang 0039, Xi Zhang 0005, Wen Zhu |
WCNC | 3 |
| 2013 | Informative SNPs Selection Based on Two-Locus and Multilocus Linkage Disequilibrium: Criteria of Max-Correlation and Min-RedundancyabstractCurrently, there are lots of methods to select informative SNPs for haplotype reconstruction. However, there are still some challenges that render them ineffective for large data sets. First, some traditional methods belong to wrappers which are of high computational complexity. Second, some methods ignore linkage disequilibrium that it is hard to interpret selection results. In this study, we innovatively derive optimization criteria by combining two-locus and multilocus LD measure to obtain the criteria of Max-Correlation and Min-Redundancy (MCMR). Then, we use a greedy algorithm to select the candidate set of informative SNPs constrained by the criteria. Finally, we use backward scheme to refine the candidate subset. We separately use small and middle (>1,000 SNPs) data sets to evaluate MCMR in terms of the reconstuction accuracy, the time complexity, and the compactness. Additionally, to demonstrate that MCMR is practical for large data sets, we design a parameter w to adapt to various platforms and introduce another replacement scheme for larger data sets, which sharply narrow down the computational complexity of evaluating the reconstruct ratio. Then, we first apply our method based on haplotype reconstruction for large size (>5,000 SNPs) data sets. The results confirm that MCMR leads to promising improvement in informative SNPs selection and prediction accuracy. Wen Zhu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2012 | Multiple ant colony algorithm method for selecting tag SNPs
Xiong Li 0002, Wen Zhu, Renfa Li, Shulin Wang |
J. Biomed. Informatics | 3 |
| 2012 | A Novel Method to Select Informative SNPs and Their Application in Genetic Association StudiesabstractThe association studies between complex diseases and single nucleotide polymorphisms (SNPs) or haplotypes have recently received great attention. However, these studies are limited by the cost of genotyping all SNPs. Therefore, it is essential to find a small subset of tag SNPs representing the rest of the SNPs. The presence of linkage disequilibrium between tag SNPs and the disease variant (genotyped or not), may allow fine mapping study. In this paper, we combine a nearest-means classifier (NMC) and ant colony algorithm to select tags. Results show that our method (ACO/NMC) can get a similar prediction accuracy with method BPSO/SVM and is better than BPSO/STAMPA for small data sets. For large data sets, although the prediction accuracy of our method is lower than BPSO/SVM, ACO/NMC can reach a high accuracy (>99 percent) in a relatively short time. when the number of tags increases, the time complexity of NMC is nearly linear growth. To find out that the ability of tags to locate disease locus, we simulate a case-control study and use two-locus haplotype analysis to quantitatively assess the power. The result showed that 20 percent of all SNPs selected by NMC have about 10 percent higher power than random tags, on average. Wen Zhu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2009 | Refactoring J2EE Application for JBI-Based ESB: A Case StudyabstractEnterprise service bus (ESB) plays an important role in enterprise SOA, and Java business integration (JBI) is the standard for Java based ESBs. As enterprises adopt ESB as the integration hub, how to preserve and leverage existing J2EE assets becomes an important question. While most ESBs allow J2EE applications to be deployed as-is, refactoring these applications allows them to take advantage of many service provided by the ESB platform and prompts service reuse. We will discuss our experience with a major US federal agencypsilas SOA efforts where a legacy J2EE application was refactored to be integrated in a JBI environment. We will also show how infrastructure concerns such quality of services (QoS) were moved from the J2EE application to the JBI container. Wen Zhu, Walcélio L. Melo |
EDOC | 1 |