VLDB 2026 Research / reviewers in the wild / expert
Fang Bai
dblp:47/7235
· DBLP profile ↗
15ranked-venue papers
9as first author
8since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 6 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 5 since 2021Systems, architecture and hardware · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | DiffPROTACs is a deep learning-based generator for proteolysis targeting chimerasabstractPROteolysis TArgeting Chimeras (PROTACs) has recently emerged as a promising technology. However, the design of rational PROTACs, especially the linker component, remains challenging due to the absence of structure-activity relationships and experimental data. Leveraging the structural characteristics of PROTACs, fragment-based drug design (FBDD) provides a feasible approach for PROTAC research. Concurrently, artificial intelligence-generated content has attracted considerable attention, with diffusion models and Transformers emerging as indispensable tools in this field. In response, we present a new diffusion model, DiffPROTACs, harnessing the power of Transformers to learn and generate new PROTAC linkers based on given ligands. To introduce the essential inductive biases required for molecular generation, we propose the O(3) equivariant graph Transformer module, which augments Transformers with graph neural networks (GNNs), using Transformers to update nodes and GNNs to update the coordinates of PROTAC atoms. DiffPROTACs effectively competes with existing models and achieves comparable performance on two traditional FBDD datasets, ZINC and GEOM. To differentiate the molecular characteristics between PROTACs and traditional small molecules, we fine-tuned the model on our self-built PROTACs dataset, achieving a 93.86% validity rate for generated PROTACs. Additionally, we provide a generated PROTAC database for further research, which can be accessed at https://bailab.siais.shanghaitech.edu.cn/service/DiffPROTACs-generated.tgz. The corresponding code is available at https://github.com/Fenglei104/DiffPROTACs and the server is at https://bailab.siais.shanghaitech.edu.cn/services/diffprotacs. Fenglei Li, Qiaoyu Hu, Yongqi Zhou, Hao Yang 0060, Fang Bai |
Briefings Bioinform. | 5 |
| 2024 | TEFDTA: a transformer encoder and fingerprint representation combined prediction method for bonded and non-bonded drug-target affinitiesabstractMOTIVATION: The prediction of binding affinity between drug and target is crucial in drug discovery. However, the accuracy of current methods still needs to be improved. On the other hand, most deep learning methods focus only on the prediction of non-covalent (non-bonded) binding molecular systems, but neglect the cases of covalent binding, which has gained increasing attention in the field of drug development. RESULTS: In this work, a new attention-based model, A Transformer Encoder and Fingerprint combined Prediction method for Drug-Target Affinity (TEFDTA) is proposed to predict the binding affinity for bonded and non-bonded drug-target interactions. To deal with such complicated problems, we used different representations for protein and drug molecules, respectively. In detail, an initial framework was built by training our model using the datasets of non-bonded protein-ligand interactions. For the widely used dataset Davis, an additional contribution of this study is that we provide a manually corrected Davis database. The model was subsequently fine-tuned on a smaller dataset of covalent interactions from the CovalentInDB database to optimize performance. The results demonstrate a significant improvement over existing approaches, with an average improvement of 7.6% in predicting non-covalent binding affinity and a remarkable average improvement of 62.9% in predicting covalent binding affinity compared to using BindingDB data alone. At the end, the potential ability of our model to identify activity cliffs was investigated through a case study. The prediction results indicate that our model is sensitive to discriminate the difference of binding affinities arising from small variances in the structures of compounds. AVAILABILITY AND IMPLEMENTATION: The codes and datasets of TEFDTA are available at https://github.com/lizongquan01/TEFDTA. Zongquan Li, Pengxuan Ren, Hao Yang 0060, Jie Zheng 0002, Fang Bai |
Bioinform. | 5 |
| 2024 | GENNDTI: Drug-Target Interaction Prediction Using Graph Neural Network Enhanced by Router NodesabstractIdentifying drug-target interactions (DTI) is crucial in drug discovery and repurposing, and in silico techniques for DTI predictions are becoming increasingly important for reducing time and cost. Most interaction-based DTI models rely on the guilt-by-association principle that "similar drugs can interact with similar targets". However, such methods utilize precomputed similarity matrices and cannot dynamically discover intricate correlations. Meanwhile, some methods enrich DTI networks by incorporating additional networks like DDI and PPI networks, enriching biological signals to enhance DTI prediction. While these approaches have achieved promising performance in DTI prediction, such coarse-grained association data do not explain the specific biological mechanisms underlying DTIs. In this work, we propose GENNDTI, which constructs biologically meaningful routers to represent and integrate the salient properties of drugs and targets. Similar drugs or targets connect to more same router nodes, capturing property sharing. In addition, heterogeneous encoders are designed to distinguish different types of interactions, modeling both real and constructed interactions. This strategy enriches graph topology and enhances prediction efficiency as well. We evaluate the proposed method on benchmark datasets, demonstrating comparative performance over existing methods. We specifically analyze router nodes to validate their efficacy in improving predictions and providing biological explanations. Beiyuan Yang, Yule Liu, Fang Bai, Mingyue Zheng, Jie Zheng 0002 |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | The Proxy Step-Size Technique for Regularized Optimization on the Sphere ManifoldabstractWe give an effective solution to the regularized optimization problem$g (\boldsymbol{x}) + h (\boldsymbol{x})$, where$\boldsymbol{x}$is constrained on the unit sphere$\Vert \boldsymbol{x} \Vert _{2} = 1$. Here$g (\cdot )$is a smooth cost with Lipschitz continuous gradient within the unit ball$\lbrace \boldsymbol{x} : \Vert \boldsymbol{x} \Vert _{2} \le 1 \rbrace$whereas$h (\cdot )$is typically non-smooth but convex and absolutely homogeneous,e.g.,norm regularizers and their combinations. Our solution is based on the Riemannian proximal gradient, using an idea we callproxy step-size– a scalar variable which we prove is monotone with respect to the actual step-size within an interval. The proxy step-size exists ubiquitously for convex and absolutely homogeneous$h(\cdot )$, and decides the actual step-size and the tangent update in closed-form, thus the complete proximal gradient iteration. Based on these insights, we design a Riemannian proximal gradient method using the proxy step-size. We prove that our method converges to a critical point, guided by a line-search technique based on the$g(\cdot )$cost only. The proposed method can be implemented in a couple of lines of code. We show its usefulness by applying nuclear norm,$\ell _{1}$norm, and nuclear-spectral norm regularization to three classical computer vision problems. The improvements are consistent and backed by numerical experiments. Fang Bai, Adrien Bartoli |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Scanline Homographies for Rolling-Shutter Plane Absolute PoseabstractCameras on portable devices are manufactured with a rolling-shutter (RS) mechanism, where the image rows (aka. scanlines) are read out sequentially. The unknown camera motions during the imaging process cause the so-called RS effects which are solved by motion assumptions in the literature. In this work, we give a solution to the absolute pose problem free of motion assumptions. We categorically demonstrate that the only requirement is motion smoothness instead of stronger constraints on the camera motion. To this end, we propose a novel mathematical abstraction for RS cameras observing a planar scene, called the scanline-homography, a 3 × 2 matrix with 5 DOFs. We establish the relationship between a scanline-homography and the corresponding plane-homography, a 3 × 3 matrix with 6 DOFs assuming the camera is calibrated. We estimate the scanline-homographies of an RS frame using a smooth image warp powered by B-Splines, and recover the plane-homographies afterwards to obtain the scanline-poses based on motion smoothness. We back our claims with various experiments. Code and new datasets: https://bitbucket.org/clermontferrand/planarscanlinehomography/src/master/. Fang Bai, Agniva Sengupta, Adrien Bartoli |
CVPR | 1 |
| 2022 | PiLSL: pairwise interaction learning-based graph neural network for synthetic lethality prediction in human cancersabstractMOTIVATION: Synthetic lethality (SL) is a type of genetic interaction in which the simultaneous inactivation of two genes leads to cell death, while the inactivation of a single gene does not affect the cell viability. It can effectively expand the range of anti-cancer therapeutic targets. SL interactions are identified mainly by experimental screening and computational prediction. Recent machine-learning methods mostly learn the representation of each gene individually, ignoring the representation of the pairwise interaction between two genes. In addition, the mechanisms of SL, the key to translating SL into cancer therapeutics, are often unclear. RESULTS: To fill the gaps, we propose a pairwise interaction learning-based graph neural network (GNN) named PiLSL to learn the representation of pairwise interaction between two genes for SL prediction. First, we construct an enclosing graph for each pair of genes from a knowledge graph. Secondly, we design an attentive embedding propagation layer in a GNN to discriminate the importance among the edges in the enclosing graph and to learn the latent features of the pairwise interaction from the weighted enclosing graph. Finally, we further fuse the latent features with explicit features extracted from multi-omics data to obtain powerful gene representations for SL prediction. Extensive experimental results demonstrate that PiLSL outperforms the best baseline by a large margin and generalizes well under three realistic scenarios. Besides, PiLSL provides an explanation of SL mechanisms via the weighted paths in the enclosing graphs by attention mechanism. AVAILABILITY AND IMPLEMENTATION: Our source code is available at https://github.com/JieZheng-ShanghaiTech/PiLSL. Xin Liu 0027, Jiale Yu, Beiyuan Yang, Shike Wang, Fang Bai, Jie Zheng 0002 |
Bioinform. | 7 |
| 2022 | Procrustes Analysis with Deformations: A Closed-Form Solution by Eigenvalue Decomposition
Fang Bai, Adrien Bartoli |
Int. J. Comput. Vis. | 1 |
| 2021 | Sparse Pose Graph Optimization in Cycle SpaceabstractThe state-of-the-art modern pose-graph optimization (PGO) systems are vertex based. In this context, the number of variables might be high, albeit the number of cycles in the graph (loop closures) is relatively low. For sparse problems particularly, the cycle space has a significantly smaller dimension than the number of vertices. By exploiting this observation, in this article, we propose an alternative solution to PGO that directly exploits the cycle space. We characterize the topology of the graph as a cycle matrix, and reparameterize the problem using relative poses, which are further constrained by a cycle basis of the graph. We show that by using a minimum cycle basis, the cycle-based approach has superior convergence properties against its vertex-based counterpart, in terms of convergence speed and convergence to the global minimum. For sparse graphs, our cycle-based approach is also more time efficient than the vertex-based. As an additional contribution of this work, we present an effective algorithm to compute the minimum cycle basis. Albeit known in computer science, we believe that this algorithm is not familiar to the robotics community. All the claims are validated by experiments on both standard benchmarks and simulated datasets. To foster the reproduction of the results, we provide a complete open-source C++ implementation1of our approach. Fang Bai, Teresa Vidal-Calleja, Giorgio Grisetti |
IEEE Trans. Robotics | 1 |
| 2020 | Change of Optimal Values: A Pre-calculated Metric
Fang Bai |
ICRA | 1 |
| 2020 | Efficient two step optimization for large embedded deformation graph based SLAMabstractEmbedded deformation graph is a widely used technique in deformable geometry and graphical problems. Although the technique has been transmitted to stereo (or RGB-D) camera based SLAM applications, it remains challenging to compromise the computational cost as the model grows. In practice, the processing time grows rapidly in accordance with the expansion of maps. In this paper, we propose an approach to decouple the nodes of deformation graph in large scale dense deformable SLAM and keep the estimation time to be constant. We observe that only partial deformable nodes in the graph are connected to visible points. Based on this fact, the sparsity of the original Hessian matrix is utilized to split the parameter estimation into two independent steps. With this new technique, we achieve faster parameter estimation with amortized computation complexity reduced from O(n2) to almost O(1). As a result, the computational cost barely increases as the map keeps growing. Based on our strategy, the computational bottleneck in large scale embedded deformation graph based applications will be greatly mitigated. The effectiveness is validated by experiments, featuring large scale deformation scenarios. Jingwei Song, Fang Bai, Liang Zhao 0003, Shoudong Huang, Rong Xiong |
ICRA | 2 |
| 2019 | Research on Game Model of Wireless Sensor Network Intrusion Detection
Fang Bai, Da Peng Lang |
EWSN | 1 |
| 2018 | Predicting Objective Function Change in Pose-Graph OptimizationabstractRobust online incremental SLAM applications require metrics to evaluate the impact of current measurements. Despite its prevalence in graph pruning, information-theoretic metrics solely are insufficient to detect outliers. The optimal value of the objective function is a better choice to detect outliers but cannot be computed unless the problem is solved. In this paper, we show how the objective function change can be predicted in an incremental pose-graph optimization scheme, without actually solving the problem. The predicted objective function change can be used to guide online decisions or detect outliers. Experiments validate the accuracy of the predicted objective function, and an application to outlier detection is also provided, showing its advantages over M-estimators. Fang Bai, Teresa Vidal-Calleja, Shoudong Huang, Rong Xiong |
IROS | 1 |
| 2016 | Incremental SQP method for constrained optimization formulation in SLAMabstract© 2016 IEEE. The simultaneous localization and mapping (SLAM) problem has been a research focus for many years and have reached a mature state. However, more robust solutions to the SLAM problem are still required, especially in large noise level scenarios. Because of the strong non-linearity of the SLAM problem, it is vital to start from a good initial value to avoid being trapped in local minima. In this paper, we propose a new SLAM formulation transforming the unconstrained Least Squares formulation into a constrained optimization problem. Algorithms based on this new formulation can naturally start from good initial value. Different from other constrained optimization problem, this new formulation can be efficiently solved with Sequential Quadratic Programming (SQP) methods. Based on SQP, we propose an incremental SQP algorithm to solve SLAM, which shows great advantage over Gauss Newton (g2o implementation) when working in large noise level scenarios. Experimental results show the validity of the proposed approach. Fang Bai, Shoudong Huang, Teresa Vidal-Calleja, Qingling Zhang 0001 |
ICARCV | 1 |
| 2010 | Bioactive Conformational Generation of Small Molecules: A Comparative Analysis between Force-Field and Multiple Empirical Criteria Based MethodsabstractBACKGROUND: Conformational sampling for small molecules plays an essential role in drug discovery research pipeline. Based on multi-objective evolution algorithm (MOEA), we have developed a conformational generation method called Cyndi in the previous study. In this work, in addition to Tripos force field in the previous version, Cyndi was updated by incorporation of MMFF94 force field to assess the conformational energy more rationally. With two force fields against a larger dataset of 742 bioactive conformations of small ligands extracted from PDB, a comparative analysis was performed between pure force field based method (FFBM) and multiple empirical criteria based method (MECBM) hybrided with different force fields. RESULTS: Our analysis reveals that incorporating multiple empirical rules can significantly improve the accuracy of conformational generation. MECBM, which takes both empirical and force field criteria as the objective functions, can reproduce about 54% (within 1Å RMSD) of the bioactive conformations in the 742-molecule testset, much higher than that of pure force field method (FFBM, about 37%). On the other hand, MECBM achieved a more complete and efficient sampling of the conformational space because the average size of unique conformations ensemble per molecule is about 6 times larger than that of FFBM, while the time scale for conformational generation is nearly the same as FFBM. Furthermore, as a complementary comparison study between the methods with and without empirical biases, we also tested the performance of the three conformational generation methods in MacroModel in combination with different force fields. Compared with the methods in MacroModel, MECBM is more competitive in retrieving the bioactive conformations in light of accuracy but has much lower computational cost. CONCLUSIONS: By incorporating different energy terms with several empirical criteria, the MECBM method can produce more reasonable conformational ensemble with high accuracy but approximately the same computational cost in comparison with FFBM method. Our analysis also reveals that the performance of conformational generation is irrelevant to the types of force field adopted in characterization of conformational accessibility. Moreover, post energy minimization is not necessary and may even undermine the diversity of conformational ensemble. All the results guide us to explore more empirical criteria like geometric restraints during the conformational process, which may improve the performance of conformational generation in combination with energetic accessibility, regardless of force field types adopted. Fang Bai, Xiaofeng Liu 0005, Haoyun Zhang, Hualiang Jiang, Xicheng Wang, Honglin Li 0003 |
BMC Bioinform. | 1 |
| 2009 | Cyndi: a multi-objective evolution algorithm based method for bioactive molecular conformational generationabstractBACKGROUND: Conformation generation is a ubiquitous problem in molecule modelling. Many applications require sampling the broad molecular conformational space or perceiving the bioactive conformers to ensure success. Numerous in silico methods have been proposed in an attempt to resolve the problem, ranging from deterministic to non-deterministic and systemic to stochastic ones. In this work, we described an efficient conformation sampling method named Cyndi, which is based on multi-objective evolution algorithm. RESULTS: The conformational perturbation is subjected to evolutionary operation on the genome encoded with dihedral torsions. Various objectives are designated to render the generated Pareto optimal conformers to be energy-favoured as well as evenly scattered across the conformational space. An optional objective concerning the degree of molecular extension is added to achieve geometrically extended or compact conformations which have been observed to impact the molecular bioactivity (J Comput -Aided Mol Des 2002, 16: 105-112). Testing the performance of Cyndi against a test set consisting of 329 small molecules reveals an average minimum RMSD of 0.864 A to corresponding bioactive conformations, indicating Cyndi is highly competitive against other conformation generation methods. Meanwhile, the high-speed performance (0.49 +/- 0.18 seconds per molecule) renders Cyndi to be a practical toolkit for conformational database preparation and facilitates subsequent pharmacophore mapping or rigid docking. The copy of precompiled executable of Cyndi and the test set molecules in mol2 format are accessible in Additional file 1. CONCLUSION: On the basis of MOEA algorithm, we present a new, highly efficient conformation generation method, Cyndi, and report the results of validation and performance studies comparing with other four methods. The results reveal that Cyndi is capable of generating geometrically diverse conformers and outperforms other four multiple conformer generators in the case of reproducing the bioactive conformations against 329 structures. The speed advantage indicates Cyndi is a powerful alternative method for extensive conformational sampling and large-scale conformer database preparation. Xiaofeng Liu 0005, Fang Bai, Sisheng Ouyang, Xicheng Wang, Honglin Li 0003, Hualiang Jiang |
BMC Bioinform. | 2 |