VLDB 2026 Research / reviewers in the wild / expert
Haifeng Hu 0004
dblp:28/5938-4
· DBLP profile ↗
15ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0002-6585-3106ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A 3-D Correlated Blockage Model for Satellite Networks
Bin Liu 0056, Haifeng Hu 0004, Rongfang Song |
IEEE Trans. Commun. | 2 |
| 2024 | OLB-AC: toward optimizing ligand bioactivities through deep graph learning and activity cliffsabstractMOTIVATION: Deep graph learning (DGL) has been widely employed in the realm of ligand-based virtual screening. Within this field, a key hurdle is the existence of activity cliffs (ACs), where minor chemical alterations can lead to significant changes in bioactivity. In response, several DGL models have been developed to enhance ligand bioactivity prediction in the presence of ACs. Yet, there remains a largely unexplored opportunity within ACs for optimizing ligand bioactivity, making it an area ripe for further investigation. RESULTS: We present a novel approach to simultaneously predict and optimize ligand bioactivities through DGL and ACs (OLB-AC). OLB-AC possesses the capability to optimize ligand molecules located near ACs, providing a direct reference for optimizing ligand bioactivities with the matching of original ligands. To accomplish this, a novel attentive graph reconstruction neural network and ligand optimization scheme are proposed. Attentive graph reconstruction neural network reconstructs original ligands and optimizes them through adversarial representations derived from their bioactivity prediction process. Experimental results on nine drug targets reveal that out of the 667 molecules generated through OLB-AC optimization on datasets comprising 974 low-activity, noninhibitor, or highly toxic ligands, 49 are recognized as known highly active, inhibitor, or nontoxic ligands beyond the datasets' scope. The 27 out of 49 matched molecular pairs generated by OLB-AC reveal novel transformations not present in their training sets. The adversarial representations employed for ligand optimization originate from the gradients of bioactivity predictions. Therefore, we also assess OLB-AC's prediction accuracy across 33 different bioactivity datasets. Results show that OLB-AC achieves the best Pearson correlation coefficient (r2) on 27/33 datasets, with an average improvement of 7.2%-22.9% against the state-of-the-art bioactivity prediction methods. AVAILABILITY AND IMPLEMENTATION: The code and dataset developed in this work are available at github.com/Yueming-Yin/OLB-AC. Yueming Yin, Haifeng Hu 0004, Jitao Yang, Chun Ye, Wilson Wen Bin Goh, Adams Wai-Kin Kong |
Bioinform. | 2 |
| 2024 | Macro-Diversity in Cellular Networks With Correlated Blockage ModelabstractBlockage effects can dramatically deteriorate high frequency cellular network performance. One practical solution is to apply macro-diversity technique, where the user associates with more than one Base Station (BS), to improve the chance of Light-Of-Sight (LOS) link availability. However, how the correlations are brought by both blockages and cooperative BSs is still unclear due to the analytical complexity. In addition, these two kinds of correlations are highly coupled and further impact macro-diversity gain. On basis of the cascade blockage model, this paper proposes a new functional framework to evaluate macro-diversity gains from the perspective of a typical user subjected to three metrics: joint LOS probability, conditional LOS probability and coverage probability. We first derive a set of tractable functional equations to describe three performance metrics. Then corresponding iterative algorithms that can be solved efficiently are designed respectively for three metrics. This analytical framework gains insight into contradictory requirements for macro-diversity, which not only guide the usage of Coordinated MultiPoint (CoMP) in realistic environments, but also determine the beamforming usage in mmWave band. We demonstrate the feasibility and flexibility of this functional framework by a simplified joint transmission example. Bin Liu 0056, Haifeng Hu 0004, Hongkui Shi, Hong Wang 0011, Rongfang Song |
IEEE Trans. Commun. | 2 |
| 2023 | Effectiveness Analysis of Multiple Initial States Simulated Annealing Algorithm, a Case Study on the Molecular Docking Tool AutoDock VinaabstractSimulated Annealing (SA) algorithm is not effective with large optimization problems for its slow convergence. Hence, several parallel Simulated Annealing (pSA) methods have been proposed, where the increase of searching threads can boost the speed of convergence. Although satisfactory solutions can be obtained by these methods, there is no rigorous mathematical analyses on their effectiveness. Thus, this article introduces a probabilistic model, on which a theorem about the effectiveness of multiple initial states parallel SA (MISPSA) has been proven. The theorem also demonstrates that the increasing parallelism in pSA algorithm with the reducing of search depth in each thread could obtain almost the same probability of finding the global optimal solution. We validated our theorem on AutoDock Vina, a widely used molecular docking tool with high accuracy and docking speed. AutoDock Vina uses a pSA strategy to find optimal molecular conformations. Under the premise that the total searching workload (i.e., thread number * iteration depth of each thread) remains unchanged, the docking accuracy from an aggressively parallelized SA searching method is almost the same or even better than those from the default exhaustiveness (parallelism degree) configuration of AutoDock Vina. Taking complex '1hnn' as an example,with the increase (125x) in the number of initial states (from 8 to 1000) and the decrease in the search depth for each thread (from 15540 to 124, or 1/125 of the original search depth), the mean energy is -7.80 and -7.94, while the mean RMSD is 3.4 and 3.14, respectively. The result also implies that a considerable speedup (in this case 125x in theory) can be obtained by a highly parallelized SA algorithm implementation. Xingxing Zhou, Qingde Lin, Shidi Tang, Haifeng Hu 0004 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2022 | AFSE: towards improving model generalization of deep graph learning of ligand bioactivities targeting GPCR proteinsabstractLigand molecules naturally constitute a graph structure. Recently, many excellent deep graph learning (DGL) methods have been proposed and used to model ligand bioactivities, which is critical for the virtual screening of drug hits from compound databases in interest. However, pharmacists can find that these well-trained DGL models usually are hard to achieve satisfying performance in real scenarios for virtual screening of drug candidates. The main challenges involve that the datasets for training models were small-sized and biased, and the inner active cliff cases would worsen model performance. These challenges would cause predictors to overfit the training data and have poor generalization in real virtual screening scenarios. Thus, we proposed a novel algorithm named adversarial feature subspace enhancement (AFSE). AFSE dynamically generates abundant representations in new feature subspace via bi-directional adversarial learning, and then minimizes the maximum loss of molecular divergence and bioactivity to ensure local smoothness of model outputs and significantly enhance the generalization of DGL models in predicting ligand bioactivities. Benchmark tests were implemented on seven state-of-the-art open-source DGL models with the potential of modeling ligand bioactivities, and precisely evaluated by multiple criteria. The results indicate that, on almost all 33 GPCRs datasets and seven DGL models, AFSE greatly improved their enhancement factor (top-10%, 20% and 30%), which is the most important evaluation in virtual screening of hits from compound databases, while ensuring the superior performance on RMSE and $r^2$. The web server of AFSE is freely available at http://noveldelta.com/AFSE for academic purposes. Yueming Yin, Haifeng Hu 0004, Zhen Yang 0001, Feihu Jiang, Yihe Huang |
Briefings Bioinform. | 2 |
| 2022 | Metric learning for domain adversarial network
Haifeng Hu 0004, Yueming Yin |
Frontiers Comput. Sci. | 1 |
| 2022 | Disclosing incoherent sparse and low-rank patterns inside homologous GPCR tasks for better modelling of ligand bioactivities
Chuangchuang Lan, Xuelin Ye, Jiale Deng, Wanqing Huang, Xueni Yang, Yanxiang Zhu, Haifeng Hu 0004 |
Frontiers Comput. Sci. | 8 |
| 2022 | Universal multi-Source domain adaptation for image classification
Yueming Yin, Zhen Yang 0001, Haifeng Hu 0004, Xiaofu Wu |
Pattern Recognit. | 3 |
| 2021 | Metric-learning-assisted domain adaptation
Yueming Yin, Zhen Yang 0001, Haifeng Hu 0004, Xiaofu Wu |
Neurocomputing | 3 |
| 2021 | Pseudo-margin-based universal domain adaptation
Yueming Yin, Zhen Yang 0001, Xiaofu Wu, Haifeng Hu 0004 |
Knowl. Based Syst. | 4 |
| 2019 | Precise modelling and interpretation of bioactivities of ligands targeting G protein-coupled receptorsabstractMOTIVATION: Accurate prediction and interpretation of ligand bioactivities are essential for virtual screening and drug discovery. Unfortunately, many important drug targets lack experimental data about the ligand bioactivities; this is particularly true for G protein-coupled receptors (GPCRs), which account for the targets of about a third of drugs currently on the market. Computational approaches with the potential of precise assessment of ligand bioactivities and determination of key substructural features which determine ligand bioactivities are needed to address this issue. RESULTS: A new method, SED, was proposed to predict ligand bioactivities and to recognize key substructures associated with GPCRs through the coupling of screening for Lasso of long extended-connectivity fingerprints (ECFPs) with deep neural network training. The SED pipeline contains three successive steps: (i) representation of long ECFPs for ligand molecules, (ii) feature selection by screening for Lasso of ECFPs and (iii) bioactivity prediction through a deep neural network regression model. The method was examined on a set of 16 representative GPCRs that cover most subfamilies of human GPCRs, where each has 300-5000 ligand associations. The results show that SED achieves excellent performance in modelling ligand bioactivities, especially for those in the GPCR datasets without sufficient ligand associations, where SED improved the baseline predictors by 12% in correlation coefficient (r2) and 19% in root mean square error. Detail data analyses suggest that the major advantage of SED lies on its ability to detect substructures from long ECFPs which significantly improves the predictive performance. AVAILABILITY AND IMPLEMENTATION: The source code and datasets of SED are freely available at https://zhanglab.ccmb.med.umich.edu/SED/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Wallace K. B. Chan, Weijian Wu, Haifeng Hu 0004, Shancheng Yan, Xiaoyan Ke, Yang Zhang 0040 |
Bioinform. | 6 |
| 2019 | Semi-Supervised Metric Learning-Based Anchor Graph Hashing for Large-Scale Image RetrievalabstractHashing-based image retrieval methods have become a cutting-edge topic in the information retrieval domain due to their high efficiency and low cost. In order to perform efficient hash learning by simultaneously preserving the semantic similarity and data structures in the feature space, this paper presents the semi-supervised metric learning-based anchor graph hashing method. Our proposed approach can be divided into three parts. First, we exploit a transformation matrix to construct the anchor-based similarity graph of the training set. Second, we propose the objective function based on the triplet relationship, in which the optimal transformation matrix can be learned by using the smoothness of labels and the margin hinge loss incurred by the triplet constraint. Moreover, the stochastic gradient descent (SGD) method leverages the gradient on each triplet to update the transformation matrix. Finally, a penalty factor is designed to accelerate the execution speed of SGD. Through comparison with the retrieval results of several state-of-the-art methods on several image benchmarks, the experiments validate the feasibility and advantages of our proposed methods. Haifeng Hu 0004, Kun Wang 0005, Chenggang Lv, Zhen Yang 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | WDL-RF: predicting bioactivities of ligand molecules acting with G protein-coupled receptors by combining weighted deep learning and random forestabstractMotivation: Precise assessment of ligand bioactivities (including IC50, EC50, Ki, Kd, etc.) is essential for virtual screening and lead compound identification. However, not all ligands have experimentally determined activities. In particular, many G protein-coupled receptors (GPCRs), which are the largest integral membrane protein family and represent targets of nearly 40% drugs on the market, lack published experimental data about ligand interactions. Computational methods with the ability to accurately predict the bioactivity of ligands can help efficiently address this problem. Results: We proposed a new method, WDL-RF, using weighted deep learning and random forest, to model the bioactivity of GPCR-associated ligand molecules. The pipeline of our algorithm consists of two consecutive stages: (i) molecular fingerprint generation through a new weighted deep learning method, and (ii) bioactivity calculations with a random forest model; where one uniqueness of the approach is that the model allows end-to-end learning of prediction pipelines with input ligands being of arbitrary size. The method was tested on a set of twenty-six non-redundant GPCRs that have a high number of active ligands, each with 200-4000 ligand associations. The results from our benchmark show that WDL-RF can generate bioactivity predictions with an average root-mean square error 1.33 and correlation coefficient (r2) 0.80 compared to the experimental measurements, which are significantly more accurate than the control predictors with different molecular fingerprints and descriptors. In particular, data-driven molecular fingerprint features, as extracted from the weighted deep learning models, can help solve deficiencies stemming from the use of traditional hand-crafted features and significantly increase the efficiency of short molecular fingerprints in virtual screening. Availability and implementation: The WDL-RF web server, as well as source codes and datasets of WDL-RF, is freely available at https://zhanglab.ccmb.med.umich.edu/WDL-RF/ for academic purposes. Supplementary information: Supplementary data are available at Bioinformatics online. Weijian Wu, Haifeng Hu 0004, Wallace K. B. Chan, Xiaoyan Ke, Yang Zhang 0040 |
Bioinform. | 5 |
| 2015 | Fast Multi-label Learning via HashingabstractMulti-label learning (MLL) copes with the classification problems where each in-stance can be tagged with multiple labels simultaneously. During the last several years, many MLL algorithms were proposed and they achieved excellent performance in multiple applications. However, these approaches are usually time-consuming and cannot handle large-scale data. In this paper, we propose a fast multi-label learning algorithm HashMLL based on hashing schemes. The approach HashMLL takes advantage of a Locality Sensitive Hashing (LSH) to identify its neighboring instances for each unseen instance, and exploits label correlation by estimating the similarity of labels through a minwise independent permutations locality sensitive hashing (MinHash). After that, relied on statistical information attained from all related labels of the neighboring instances, maxi-mum a posteriori (MAP) principle is used to determine the label set for each unseen instance. Experiments show that the performance of HashMLL is highly competitive to state-of-the-art techniques, whereas its time cost is much less. Particularly, on the dataset NUS-WIDE with 269,648 instances and the dataset Flickr with 565,444 instances where none of existing methods can return results in 24 hours, HashMLL takes only 90 secs and 23266 secs respectively. Haifeng Hu 0004 |
KSEM | 1 |
| 2013 | Distributed multi-sensor tracking in wireless networks using nonparametric variant of sum-product algorithmabstractGraphical models have been widely applied in solving distributed inference problems in wireless sensor networks (WSNs). In this paper, we formulate the distributed multi-sensor tracking problem in a WSN as an inference problem on a factor graph. Using particle filtering methods, we propose a nonparametric variant of sum-product algorithm (SPA), called sequential particle-based SPA (SPSPA), for factor graphs to infer the multi-sensor target states over time. In the proposed algorithm, importance sampling methods are used to sample from message products, and the computational complexity of SPSPA is thus linear in the number of particles. We apply the SPSPA to a distributed multi-sensor tracking problem, and evaluate its performance in terms of the measurement noise and the number of particles. Wei Li 0064, Zhen Yang 0001, Haifeng Hu 0004 |
APCC | 3 |