EDBT 2026 Demo / reviewers in the wild / expert
Bowei Yan
dblp:44/10575
· DBLP profile ↗
22ranked-venue papers
8as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 9 · 6 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Interpretable integration of unpaired multi-omics for Alzheimer's diagnosis via cross-modal transformer reconstructionabstractAlzheimer's disease (AD) is a progressive neurodegenerative disorder with limited diagnostic tools and poorly understood molecular underpinnings. Although multi-omics technologies hold promise for early detection, integrating unpaired transcriptomic and epigenetic data remains a major challenge due to modality heterogeneity and small sample sizes. We present AE-Trans, an interpretable dual-channel Transformer framework that aligns RNA and DNA methylation data through cross-modal reconstruction and multi-head attention. AE-Trans achieves superior performance on prefrontal cortex datasets (accuracy = 0.9736, AUC = 0.9910) and demonstrates strong generalizability to external regions temporal cortex cohorts across brain regions (accuracy = 0.7389, AUC = 0.8432). To validate the performance within the same brain region, we tested AE-Trans on an external unpaired multi-omics dataset from the prefrontal cortex. Additionally, we validated the model on a paired multi-omics dataset to assess whether it could achieve good results in real-world scenarios. In the unpaired dataset from the external same brain region, AE-Trans achieved an accuracy of (accuracy = 0.87) and AUC of (AUC = 0.94), while in the real-world paired multi-omics dataset, the accuracy was (accuracy = 0.88) and AUC was (AUC = 0.93). These results demonstrate that AE-Trans not only validates well on external unpaired datasets, but also generalizes effectively to real-world multi-omics paired datasets, highlighting its robustness in practical applications. Through counterfactual integrated gradients, we identified key features associated with immune regulation, hormonal signaling, and neuronal metabolism. These were validated via pathway enrichment and logistic regression (AUC = 0.9749), confirming the biological relevance of model-derived markers. Furthermore, AE-Trans generalized well to two independent RNA datasets, where latent representations not only improved classification (AUCs = 0.92 and 0.89) but also stratified patients into subgroups with significantly different prognoses. These results highlight AE-Trans as a robust and explainable tool for multi-omics integration, supporting early diagnosis, biomarker discovery, and individualized risk prediction in Alzheimer's disease. Danfeng Du, Xiaodan Fan, Changshui Chen, Bowei Yan |
PLoS Comput. Biol. | 8 |
| 2026 | Semantic contrastive learning via VLM for few-shot remote sensing object detection
Bowei Yan, Chunbo Lang, Gong Cheng 0003 |
Pattern Recognit. | 1 |
| 2025 | Multi-task multi-view and iterative error-correcting random forest for acute toxicity prediction
Lianlian Wu, Guangyi Lin, Jiayu Zou, Bowei Yan, Kunhong Liu 0001, Xiaochen Bo |
Expert Syst. Appl. | 5 |
| 2025 | Constrained multi-agent evasion using deep reinforcement learning
Bowei Yan, Runle Du, Xiaojun Ban |
Neurocomputing | 1 |
| 2025 | Global-Integrated and Drift-Rectified Imprinting for Few-Shot Remote Sensing Object DetectionabstractFew-shot object detection (FSOD) in remote sensing images is a marginally explored but highly challenging task that focuses on identifying unseen classes of objects with a limited number of annotations. Current FSOD approaches often fail to accurately localize the foreground and misalign targets with various orientations, resulting in poor detection performance. For this purpose, we develop a fresh and powerful meta-learning framework based on the idea of imprinting, which leverages tailored support information to model the regional correlation between query and support objects in different stages. Specifically, a global-integrated scheme is first proposed to guide the generation of high-quality proposals by increasing the activation of foreground features and integrating global support information. Considering the orientation discrepancy of objects in query and support sets, we introduce a drift-rectified technique to achieve adaptive alignment by implicitly capturing the positional correspondence between the instances in two sets. In stark contrast to conventional FSOD approaches, our method can extract key clues and establish directional relationships between objects from different training sets, leading to better generalization capability. Extensive experiments on two standard benchmarks (DIOR and NWPU VHR-10.V2) manifest the effectiveness, and our proposed method exhibits superior performance to other competitors with similar motivation. The source code is available athttps://github.com/Ybowei/GIDR Bowei Yan, Gong Cheng 0003, Chunbo Lang, Zhongling Huang, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Understanding Negative Proposals in Generic Few-Shot Object DetectionabstractRecently, Few-Shot Object Detection (FSOD) has received considerable research attention as a strategy for reducing reliance on extensively labeled bounding boxes. However, current approaches encounter significant challenges due to the intrinsic issue of incomplete annotation while building the instance-level training benchmark. In such cases, the instances with missing annotations are regarded as background, resulting in erroneous training gradients back-propagated through the detector, thereby compromising the detection performance. To mitigate this challenge, we introduce a simple and highly efficient method that can be plugged into both meta-learning-based and transfer-learning-based methods. Our method incorporates two innovative components: Confusing Proposals Separation (CPS) and Affinity-Driven Gradient Relaxation (ADGR). Specifically, CPS effectively isolates confusing negatives while ensuring the contribution of hard negatives during model fine-tuning; ADGR then adjusts their gradients based on the affinity to different category prototypes. As a result, false-negative samples are assigned lower weights than other negatives, alleviating their harmful impacts on the few-shot detector without the requirement of additional learnable parameters. Extensive experiments conducted on the PASCAL VOC and MS-COCO datasets consistently demonstrate that our method significantly outperforms both the baseline and recent FSOD methods. Furthermore, its versatility and efficiency suggest the potential to become a stronger new baseline in the field of FSOD. Code is available at https://github.com/Ybowei/UNP. Bowei Yan, Chunbo Lang, Gong Cheng 0003, Junwei Han 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | A Multi-View Learning-Based Rule Extraction Algorithm For Accurate Hepatotoxicity PredictionabstractHepatotoxicity prediction is key to diseases with the high mortality rate. However, most of the algorithms used by now are black box in nature and lack of clear interpretability. This paper proposes a genetic algorithm-based interpretable algorithm based on rules extracted from a random forest. To take advantages from different types of omics data and molecular representations gathered from various datasets, our algorithm utilizes multiple distinct features to form a multi-view learning strategy. In detail, the genetic algorithm is designed to select optimal rules from each view, which are then used to form the ensemble of multi-view rule sets. The experiments are carried out to verify the performance of our algorithm on the hepatotoxicity data. The results confirm that our algorithm can gain high accuracy in most cases with more compact and shorter rules, compared with the original random forest or other rule-based algorithms. Our python source code and the related Supplementary Materials are available at: github.com/MLDMXM2017/MVR-GA. Yuting Zhong, Bowei Yan, Kunhong Liu 0001, Yong Xu 0009, Xiaochen Bo |
BIBM | 2 |
| 2022 | An enhanced cascade-based deep forest model for drug combination predictionabstractCombination therapy has shown an obvious curative effect on complex diseases, whereas the search space of drug combinations is too large to be validated experimentally even with high-throughput screens. With the increase of the number of drugs, artificial intelligence techniques, especially machine learning methods, have become applicable for the discovery of synergistic drug combinations to significantly reduce the experimental workload. In this study, in order to predict novel synergistic drug combinations in various cancer cell lines, the cell line-specific drug-induced gene expression profile (GP) is added as a new feature type to capture the cellular response of drugs and reveal the biological mechanism of synergistic effect. Then, an enhanced cascade-based deep forest regressor (EC-DFR) is innovatively presented to apply the new small-scale drug combination dataset involving chemical, physical and biological (GP) properties of drugs and cells. Verified by the dataset, EC-DFR outperforms two state-of-the-art deep neural network-based methods and several advanced classical machine learning algorithms. Biological experimental validation performed subsequently on a set of previously untested drug combinations further confirms the performance of EC-DFR. What is more prominent is that EC-DFR can distinguish the most important features, making it more interpretable. By evaluating the contribution of each feature type, GP feature contributes 82.40%, showing the cellular responses of drugs may play crucial roles in synergism prediction. The analysis based on the top contributing genes in GP further demonstrates some potential relationships between the transcriptomic levels of key genes under drug regulation and the synergism of drug combinations. Weiping Lin, Lianlian Wu, Yuqi Wen, Bowei Yan, Chong Dai, Kunhong Liu 0001, Xiaochen Bo |
Briefings Bioinform. | 5 |
| 2022 | Computational methods, databases and tools for synthetic lethality predictionabstractSynthetic lethality (SL) occurs between two genes when the inactivation of either gene alone has no effect on cell survival but the inactivation of both genes results in cell death. SL-based therapy has become one of the most promising targeted cancer therapies in the last decade as PARP inhibitors achieve great success in the clinic. The key point to exploiting SL-based cancer therapy is the identification of robust SL pairs. Although many wet-lab-based methods have been developed to screen SL pairs, known SL pairs are less than 0.1% of all potential pairs due to large number of human gene combinations. Computational prediction methods complement wet-lab-based methods to effectively reduce the search space of SL pairs. In this paper, we review the recent applications of computational methods and commonly used databases for SL prediction. First, we introduce the concept of SL and its screening methods. Second, various SL-related data resources are summarized. Then, computational methods including statistical-based methods, network-based methods, classical machine learning methods and deep learning methods for SL prediction are summarized. In particular, we elaborate on the negative sampling methods applied in these models. Next, representative tools for SL prediction are introduced. Finally, the challenges and future work for SL prediction are discussed. Junshan Han, Yanpeng Zhao, Caiyun Zhao, Bowei Yan, Chong Dai, Lianlian Wu, Yuqi Wen, Dongjin Leng, Zhongming Wang, Xiaoxi Yang, Xiaochen Bo |
Briefings Bioinform. | 6 |
| 2022 | Machine learning methods, databases and tools for drug combination predictionabstractCombination therapy has shown an obvious efficacy on complex diseases and can greatly reduce the development of drug resistance. However, even with high-throughput screens, experimental methods are insufficient to explore novel drug combinations. In order to reduce the search space of drug combinations, there is an urgent need to develop more efficient computational methods to predict novel drug combinations. In recent decades, more and more machine learning (ML) algorithms have been applied to improve the predictive performance. The object of this study is to introduce and discuss the recent applications of ML methods and the widely used databases in drug combination prediction. In this study, we first describe the concept and controversy of synergism between drug combinations. Then, we investigate various publicly available data resources and tools for prediction tasks. Next, ML methods including classic ML and deep learning methods applied in drug combination prediction are introduced. Finally, we summarize the challenges to ML methods in prediction tasks and provide a discussion on future work. Lianlian Wu, Yuqi Wen, Dongjin Leng, Chong Dai, Zhongming Wang, Bowei Yan, Xiaochen Bo |
Briefings Bioinform. | 8 |
| 2022 | Prototype-CNN for Few-Shot Object Detection in Remote Sensing ImagesabstractRecently, due to the excellent representation ability of convolutional neural networks (CNNs), object detection in remote sensing images has undergone remarkable development. However, when trained with a small number of samples, the performance of the object detectors drops sharply. In this article, we focus on the following three main challenges of few-shot object detection in remote sensing images: 1) since the sample number of novel classes is far less than base classes, object detectors would fail to quickly adapt to the features of novel classes, which would result in overfitting; 2) the scarcity of samples in novel classes leads to a sparse orientation space, while the objects in remote sensing images usually have arbitrary orientations; and 3) the distribution of object instances in remote sensing images is scattered and, therefore, it is hard to identify foreground objects from the complex background. To tackle these problems, we propose a simple yet effective method named prototype-CNN (P-CNN), which mainly consists of three parts: a prototype learning network (PLN) converting support images to class-aware prototypes, a prototype-guided region proposal network (P-G RPN) for better generation of region proposals, and a detector head extending the head of Faster region-based CNN (R-CNN) to further boost the performance. Comprehensive evaluations on the large-scale DIOR dataset demonstrate the effectiveness of our P-CNN. The source code is available athttps://github.com/Ybowei/P-CNN. Gong Cheng 0003, Bowei Yan, Peizhen Shi, Ke Li 0005, Xiwen Yao, Lei Guo 0002, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Multi-dimensional data integration algorithm based on random walk with restartabstractBACKGROUND: The accumulation of various multi-omics data and computational approaches for data integration can accelerate the development of precision medicine. However, the algorithm development for multi-omics data integration remains a pressing challenge. RESULTS: Here, we propose a multi-omics data integration algorithm based on random walk with restart (RWR) on multiplex network. We call the resulting methodology Random Walk with Restart for multi-dimensional data Fusion (RWRF). RWRF uses similarity network of samples as the basis for integration. It constructs the similarity network for each data type and then connects corresponding samples of multiple similarity networks to create a multiplex sample network. By applying RWR on the multiplex network, RWRF uses stationary probability distribution to fuse similarity networks. We applied RWRF to The Cancer Genome Atlas (TCGA) data to identify subtypes in different cancer data sets. Three types of data (mRNA expression, DNA methylation, and microRNA expression data) are integrated and network clustering is conducted. Experiment results show that RWRF performs better than single data type analysis and previous integrative methods. CONCLUSIONS: RWRF provides powerful support to users to decipher the cancer molecular subtypes, thus may benefit precision treatment of specific patients in clinical practice. Yuqi Wen, Xinyu Song 0002, Bowei Yan, Xiaoxi Yang, Lianlian Wu, Dongjin Leng, Xiaochen Bo |
BMC Bioinform. | 3 |
| 2021 | Synthetic Lethal Interactions Prediction Based on Multiple Similarity Measures Fusion
Lianlian Wu, Yuqi Wen, Xiaoxi Yang, Bowei Yan, Xiaochen Bo |
J. Comput. Sci. Technol. | 4 |
| 2021 | COMSUC: A web server for the identification of consensus molecular subtypes of cancer based on multiple methods and multi-omics dataabstractExtensive amounts of multi-omics data and multiple cancer subtyping methods have been developed rapidly, and generate discrepant clustering results, which poses challenges for cancer molecular subtype research. Thus, the development of methods for the identification of cancer consensus molecular subtypes is essential. The lack of intuitive and easy-to-use analytical tools has posed a barrier. Here, we report on the development of the COnsensus Molecular SUbtype of Cancer (COMSUC) web server. With COMSUC, users can explore consensus molecular subtypes of more than 30 cancers based on eight clustering methods, five types of omics data from public reference datasets or users' private data, and three consensus clustering methods. The web server provides interactive and modifiable visualization, and publishable output of analysis results. Researchers can also exchange consensus subtype results with collaborators via project IDs. COMSUC is now publicly and freely available with no login requirement at http://comsuc.bioinforai.tech/ (IP address: http://59.110.25.27/). For a video summary of this web server, see S1 Video and S1 File. Xinyu Song 0002, Xiaoxi Yang, Jijun Yu, Yuqi Wen, Lianlian Wu, Bowei Yan, Jiannan Feng, Xiaochen Bo |
PLoS Comput. Biol. | 7 |
| 2018 | Provable Estimation of the Number of Blocks in Block ModelsabstractCommunity detection is a fundamental unsupervised learning problem for unlabeled networks which has a broad range of applications. Many community detection algorithms assume that the number of clusters r is known apriori. In this paper, we propose an approach based on semi-definite relaxations, which does not require prior knowledge of model parameters like many existing convex relaxation methods and recovers the number of clusters and the clustering matrix exactly under a broad parameter regime, with probability tending to one. On a variety of simulated and real data experiments, we show that the proposed method often outperforms state-of-the-art techniques for estimating the number of clusters. Bowei Yan, Purnamrita Sarkar, Xiuyuan Cheng |
AISTATS | 1 |
| 2018 | Binary Classification with Karmic, Threshold-Quasi-Concave MetricsabstractComplex performance measures, beyond the popular measure of accuracy, are increasingly being used in the context of binary classification. These complex performance measures are typically not even decomposable, that is, the loss evaluated on a batch of samples cannot typically be expressed as a sum or average of losses evaluated at individual samples, which in turn requires new theoretical and methodological developments beyond standard treatments of supervised learning. In this paper, we advance this understanding of binary classification for complex performance measures by identifying two key properties: a so-called Karmic property, and a more technical threshold-quasi-concavity property, which we show is milder than existing structural assumptions imposed on performance measures. Under these properties, we show that the Bayes optimal classifier is a threshold function of the conditional probability of positive class. We then leverage this result to come up with a computationally practical plug-in classifier, via a novel threshold estimator, and further, provide a novel statistical analysis of classification error with respect to complex performance measures. Bowei Yan, Oluwasanmi Koyejo, Pradeep Ravikumar |
ICML | 1 |
| 2018 | Mean Field for the Stochastic Blockmodel: Optimization Landscape and Convergence IssuesabstractVariational approximation has been widely used in large-scale Bayesian inference recently, the simplest kind of which involves imposing a mean field assumption to approximate complicated latent structures. Despite the computational scalability of mean field, theoretical studies of its loss function surface and the convergence behavior of iterative updates for optimizing the loss are far from complete. In this paper, we focus on the problem of community detection for a simple two-class Stochastic Blockmodel (SBM). Using batch co-ordinate ascent (BCAVI) for updates, we give a complete characterization of all the critical points and show different convergence behaviors with respect to initializations. When the parameters are known, we show a significant proportion of random initializations will converge to ground truth. On the other hand, when the parameters themselves need to be estimated, a random initialization will converge to an uninformative local optimum. Soumendu Sundar Mukherjee, Purnamrita Sarkar, Y. X. Rachel Wang, Bowei Yan |
NeurIPS | 4 |
| 2017 | Fast Classification with Binary PrototypesabstractIn this work, we propose a new technique for \emphfast k-nearest neighbor (k-NN) classification in which the original database is represented via a small set of learned binary prototypes. The training phase simultaneously learns a hash function which maps the data points to binary codes, and a set of representative binary prototypes. In the prediction phase, we first hash the query into a binary code and then do the k-NN classification using the binary prototypes as the database. Our approach speeds up k-NN classification in two aspects. First, we compress the database into a smaller set of prototypes such that k-NN search only goes through a smaller set rather than the whole dataset. Second, we reduce the original space to a compact binary embedding, where the Hamming distance between two binary codes is very efficient to compute. We propose a formulation to learn the hash function and prototypes such that the classification error is minimized. We also provide a novel theoretical analysis of the proposed technique in terms of Bayes error consistency. Empirically, our method is much faster than the state-of-the-art k-NN compression methods with comparable accuracy. Sanjiv Kumar, Bowei Yan, David Simcha, Inderjit S. Dhillon |
AISTATS | 4 |
| 2017 | Convergence of Gradient EM on Multi-component Mixture of GaussiansabstractIn this paper, we study convergence properties of the gradient variant of Expectation-Maximization algorithm~\cite{lange1995gradient} for Gaussian Mixture Models for arbitrary number of clusters and mixing coefficients. We derive the convergence rate depending on the mixing coefficients, minimum and maximum pairwise distances between the true centers, dimensionality and number of components; and obtain a near-optimal local contraction radius. While there have been some recent notable works that derive local convergence rates for EM in the two symmetric mixture of Gaussians, in the more general case, the derivations need structurally different and non-trivial arguments. We use recent tools from learning theory and empirical processes to achieve our theoretical results. Bowei Yan, Mingzhang Yin, Purnamrita Sarkar |
NIPS | 1 |
| 2016 | On Robustness of Kernel ClusteringabstractClustering is an important unsupervised learning problem in machine learning and statistics. Among many existing algorithms, kernel \km has drawn much research attention due to its ability to find non-linear cluster boundaries and its inherent simplicity. There are two main approaches for kernel k-means: SVD of the kernel matrix and convex relaxations. Despite the attention kernel clustering has received both from theoretical and applied quarters, not much is known about robustness of the methods. In this paper we first introduce a semidefinite programming relaxation for the kernel clustering problem, then prove that under a suitable model specification, both K-SVD and SDP approaches are consistent in the limit, albeit SDP is strongly consistent, i.e. achieves exact recovery, whereas K-SVD is weakly consistent, i.e. the fraction of misclassified nodes vanish. Also the error bounds suggest that SDP is more resilient towards outliers, which we also demonstrate with experiments. Bowei Yan, Purnamrita Sarkar |
NIPS | 1 |
| 2012 | HodgeRank on Random Graphs for Subjective Video Quality AssessmentabstractThis paper introduces a novel framework, HodgeRank on Random Graphs, based on paired comparison, for subjective video quality assessment. Two types of random graph models are studied, i.e., Erdös-Rényi random graphs and random regular graphs. Hodge decomposition of paired comparison data may derive, from incomplete and imbalanced data, quality scores of videos and inconsistency of participants' judgments. We demonstrate the effectiveness of the proposed framework on LIVE video database. Both of the two random designs are promising sampling methods without jeopardizing the accuracy of the results. In particular, due to balanced sampling, random regular graphs may achieve better performances when sampling rates are small. However, when the number of videos is large or when sampling rates are large, their performances are so close that Erdös-Rényi random graphs, as the simplest independent and identically distributed sampling scheme, could provide good approximations to random regular graphs, as a dependent sampling scheme. In contrast to the traditional deterministic incomplete block designs, our random design is not only suitable for traditional laboratory studies, but also for crowdsourcing experiments on Internet where the raters are distributive and it is hard to control with fixed designs. Qianqian Xu 0001, Qingming Huang, Tingting Jiang 0001, Bowei Yan, Weisi Lin, Yuan Yao 0011 |
IEEE Trans. Multim. | 4 |
| 2011 | Random partial paired comparison for subjective video quality assessment via hodgerankabstractSubjective visual quality evaluation provides the groundtruth and source of inspiration in building objective visual quality metrics. Paired comparison is expected to yield more reliable results; however, this is an expensive and timeconsuming process. In this paper, we propose a novel framework of HodgeRank on Random Graphs (HRRG) to achieve efficient and reliable subjective Video Quality Assessment (VQA). To address the challenge of a potentially large number of combinations of videos to be assessed, the proposed methodology does not require the participants to perform the complete comparison of all the paired videos. Instead, participants only need to perform a random sample of all possible paired comparisons, which saves a great amount of time and labor. In contrast to the traditional deterministic incomplete block designs, our random design is not only suitable for traditional laboratory and focus-group studies, but also fit for crowdsourcing experiments on Internet where the raters are distributive over Internet and it is hard to control with precise experimental designs. Qianqian Xu 0001, Tingting Jiang 0001, Yuan Yao 0011, Qingming Huang, Bowei Yan, Weisi Lin |
ACM Multimedia | 5 |