Bingding Huang

dblp:65/6254 · DBLP profile ↗
← Back
19ranked-venue papers
0as first author
17since 2021 · last 2026
0000-0002-4748-2882ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 11 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 UMC-FM: Uncertainty-Guided Mamba-Consistent Feature Mixing for Semi-supervised Pulmonary Vessel Segmentation
Zhiqing Liang, Haomin Liang, Inayatul Haq, Liyilei Su, Bingding Huang
ICIC (17)7
2026 PIPCDiff: Pocket-Informed Pharmacophore Conditioning for Diffusion-Based Molecular Generation
Ningxin Qin, Chenran Jiang, jiean Chen, Bingding Huang
ICIC (28)7
2026 CoCo-GAN: CodeBERT-driven collaborative generative adversarial learning for software defect prediction
Xiaoxing Yang, Liwei Xiao, Jianmin Su, Bingding Huang
Softw. Qual. J.4
2025 Diffusion-Based Pre-Training for Label-Efficient Abdominal Multi-Organ Segmentation
abstract
Accurate multi-organ segmentation in Computed Tomography (CT) images is critical for computer-aided diagnosis systems. However, existing supervised methods heavily rely on costly, high-quality labeled data. To address this, we propose a label-efficient segmentation method for abdominal organs in CT images, leveraging knowledge transfer from a pre-trained diffusion model. Specifically, we pre-train a denoising diffusion model on 207,029 unlabeled 2D CT slices to capture anatomical patterns, which is then fine-tuned on limited labeled data for abdominal organ segmentation. During fine-tuning, two strate-gies-linear probing and decoder fine-tuning-are employed to adapt the model for segmentation while preserving learned representations. Quantitative results demonstrate that the pre-trained diffusion model can generate diverse and realistic$256 \times 256$CT images (FID: 11.32, sFID: 46.93, F1-score: 73.1%). Moreover, our method achieves competitive performance on the FLARE 2022 dataset for organ segmentation, particularly excelling in limited labeled data scenarios. With only 10% and 1% labeled data, our method achieves DSCs of 78.51% and 71.56% on 13 abdominal organs, respectively. Remarkably, with only four labeled 2D slices, our method still achieves a DSC of 51.81%, highlighting the efficacy of our method in alleviating the reliance of supervised learning on large-scale labeled data.
Yongzhi Huang 0002, Jinxin Zhu, Haseeb Hassan, Liyilei Su, Bingding Huang
BIBM6
2025 BioLedger: A Token Incentive-Based Framework for the Secure Sharing of Genetic Data in Consortium Blockchain
abstract
Current health data sharing frameworks encounter significant challenges in genetic data management, including issues related to privacy protection, constraints on access control, and a lack of effective incentive mechanisms. These deficiencies expose patients and researchers to risks of information leakage and result in inefficiencies in collaboration throughout the entire lifecycle of genetic data—from creation and access to sharing and revocation. Such problems have impeded the advancement of precision medicine and escalated the costs of data acquisition in medical research. Notably, the security and efficiency of data circulation become particularly acute when genetic information is shared across institutional boundaries to support the development of personalized treatment regimens or pharmaceutical research. To tackle these challenges, we propose a genetic data sharing framework based on consortium blockchain, BioLedger. In contrast to traditional centralized storage models, this framework leverages the FISCO BCOS consortium blockchain to construct a decentralized architecture. By integrating smart contracts with HCE (Hybrid CP-ABE with ECIES Encryption), it achieves fine-grained access control while ensuring a high standard of security and privacy for genetic data. Each user retains a high degree of control over their genetic data and can earn token rewards through data contribution, thereby incentivizing user participation in data sharing. Our solution aims to establish an ecosystem that not only safeguards data security and privacy but also enables efficient data exchange, with applications spanning genetic research, pharmaceutical development, and personalized medicine.
Xin Wang 0168, Ruixuan Lin, Ruisheng Huang, Huankai Chen, Lixin Liang, Bingding Huang
BIBM7
2025 NanoLAS 2.0: A Comprehensive Update on a Nanobody-Focused Platform with Advanced Visualization and Docking Simulation Features
abstract
Nanobodies are a unique class of antibodies with significant therapeutic potential. However, existing nanobody databases often lack standardized data, advanced search capabilities, and integrated analysis tools. We previously developed NanoLAS to address data integration. Here, we present NanoLAS 2.0, a comprehensive update that introduces a multi-condition associative search, an enhanced 3D viewer with molecular docking simulation, an improved sequence viewer, and integrated binding site prediction via NanoBERTa-ASP. The platform features an updated database backend and a redesigned user interface. NanoLAS 2.0 (available at https://www.nanolas2.online) provides a powerful all-in-one platform to accelerate nanobody research.
Xin Wang 0168, Zebiao Zheng, Kangrui Yu, Yangqi Hong, Yongqi Tang, Tiantai Wang, Lixin Liang, Bingding Huang
BIBM9
2025 SegSAM-3D: Integrating Semantic Point Prompts and Multi-Layer Feature Sampling in SAM for Medical Imaging Segmentation
abstract
Medical imaging segmentation, particularly in 3D, poses significant challenges related to accuracy and computational cost. Existing 2D prompt segmentation models, such as Segment Anything Model (SAM), are limited by their dependence on positional information and struggle to capture complex 3D spatial dependencies. In this study, we present SegSAM-3D, an advanced framework that integrates semantic point prompts and multi-layer feature sampling within the SAM architecture to boost the performance of 3D medical image segmentation. We have expanded the SAM model to three dimensions (3D) by adapting its 2 D components into their 3D counterparts. Semantic point prompts were incorporated to encode positional and semantic information through visual sampling. Multi-layer feature sampling integrates features from shallow and deep network layers, improving the delineation of intricate anatomical structures. We evaluate SegSAM-3D on five datasets: AbdomenCT-1K, BTCV-Abdomen, BTCVCervix, FLARE22, and KiPA22. SegSAM-3D demonstrated a significant enhancement in performance when compared to the original SAM and other baseline models (MedSAM, SAMMed3D). The proposed method achieves a Dice score of 85.08%, representing a 59.27% improvement over SAM, while maintaining computational efficiency. Notably, SegSAM-3D excels in segmenting regions with ambiguous boundaries, such as abdominal organs. These results highlight the potential of integrating semantic guidance and hierarchical features within the SAM framework, advancing the state of 3D medical image segmentation for clinical diagnosis and treatment planning.
Rashid Khan, Liyilei Su, Bingding Huang
BIBM6
2025 Enhancing Early Detection of Tractional Retinal Lesions in OCT via Self-Supervised Learning
abstract
Optical Coherence Tomography (OCT) plays a vital role in the early detection and monitoring of tractional retinal lesions (TRL), providing high-resolution visualization of retinal structures. However, automated TRL diagnosis remains challenging due to complex lesion morphology, large low-entropy background regions, and the scarcity of high-quality labeled data. Existing Self-Supervised Learning (SSL) approaches often treat all image patches equally, making them sensitive to background noise and limiting their ability to capture fine-grained lesion features. To address these issues, we propose Clustering Hetero-geneous Masked Image Modeling (CH-MIM), a novel SSL frame-work tailored for OCT-based TRL analysis. Our method lever-ages a large-scale clinical dataset containing 11,861 OCT scans collected over five years, including 3,950 expert-annotated images across six TRL severity levels (TO- T5). CH - MIM introduces a Weighted Feature Space Clustering (WFSC) module to selectively mask high-entropy regions, effectively filtering out irrelevant background information. A heterogeneous progressive masking strategy combines binary, Gaussian, and Poisson noise masks to provide diverse, informative reconstruction tasks. Furthermore, a Consistency Regularization Module (CRM) enforces stable predictions across masking branches, improving representation robustness and transferability to downstream classification. Ex-tensive experiments demonstrate that CH - MIM achieves a top-l accuracy of 97.7% and top-5 accuracy of 99.8%, surpassing state-of-the-art supervised and self-supervised baselines. These results highlight the potential of CH - MIM as an effective pretraining strategy for automated TRL screening and its applicability to broader OCT-based retinal disease diagnosis.
Yu Lu 0001, Qianying Liu, Bingding Huang, Zhaoshun Zhang, Liyilei Su
BIBM6
2025 BiGAMR-Net: Bidirectional Gated Attention and Multi-scale Residual Network for Polyp Segmentation
Liuyi Yang, Shao-Chi Pao, Xiaoxing Yang, Lixin Liang, Bingding Huang
ICIC (25)6
2025 Attention-Enhanced Few-Shot Diagnosis of Pathological Myopia and MTM
Yu Lu 0001, Bingding Huang, Caifen Wang
ICIC (16)5
2025 JanusDNA: A Powerful Bi-directional Hybrid DNA Foundation Model
abstract
Large language models (LLMs) have revolutionized natural language processing and are increasingly applied to other sequential data types, including genetic sequences. However, adapting LLMs to genetics presents significant challenges. Capturing complex genomic interactions requires modeling long-range global dependencies within DNA sequences, where interactions often span over 10,000 base pairs, even within a single gene. This poses substantial computational demands under conventional model architectures and training paradigms. Additionally, traditional LLM training approaches are suboptimal for DNA sequences: autoregressive training, while efficient for training, only supports unidirectional sequence understanding. However, DNA is inherently bidirectional. For instance, bidirectional promoters regulate gene expression in both directions and govern approximately 11% of human gene expression. Masked language models (MLMs) enable bidirectional understanding. However, they are inefficient since only masked tokens contribute to loss calculations at each training step. To address these limitations, we introduce JanusDNA, the first bidirectional DNA foundation model built upon a novel pretraining paradigm, integrating the optimization efficiency of autoregressive modeling with the bidirectional comprehension capability of masked modeling. JanusDNA's architecture leverages a Mamba-Attention Mixture-of-Experts (MoE) design, combining the global, high-resolution context awareness of attention mechanisms with the efficient sequential representation learning capabilities of Mamba. The MoE layers further enhance the model's capacity through sparse parameter scaling, while maintaining manageable computational costs. Notably, JanusDNA can process up to 1 million base pairs at single-nucleotide resolution on a single 80GB GPU using its hybrid architecture. Extensive experiments and ablation studies demonstrate that JanusDNA achieves new state-of-the-art performance on three genomic representation benchmarks. Remarkably, JanusDNA surpasses models with 250x more activated parameters, underscoring its efficiency and effectiveness. Code available at https://anonymous.4open.science/r/JanusDNA/.
Qihao Duan, Bingding Huang, Zhenqiao Song, Irina Lehmann, Roland Eils, Benjamin Wild
NeurIPS2
2025 A review of breast cancer histopathology image analysis with deep learning: Challenges, innovations, and clinical integration
abstract
Breast cancer (BC) is the most frequently diagnosed cancer among women and a leading cause of cancer-related mortality globally. Accurate and timely diagnosis is essential for improving patient outcomes. However, traditional histopathological assessments are labor-intensive and subjective, leading to inter-observer variability and diagnostic inconsistencies, especially in resource-limited settings. Furthermore, variability in tissue staining, limited availability of standardized annotated datasets, and subtle morphological patterns complicate the consistent characterization of tumors. Deep learning (DL) has recently emerged as a transformative technology in breast cancer pathology, providing automated and objective solutions for cancer detection, classification, and segmentation from histopathological images. This review systematically evaluates advanced deep learning (DL) architectures, including convolutional neural networks (CNNs), generative adversarial networks (GANs), autoencoders, deep belief networks (DBNs), extreme learning machines (ELMs), and transformer-based models such as Vision Transformers (ViTs) as well as transfer learning, attention-based explainable AI techniques, and multimodal integration to address these diagnostic challenges. Analyzing 199 references, including 182 peer-reviewed studies published between 2014 and 2025 and 17 reputable online sources (websites, databases, etc.), we identify key innovations, limitations, and opportunities for future research. Furthermore, we explore the critical roles of synthetic data augmentation, explainable AI (XAI), and multimodal integration to enhance clinical trust, model interpretability, and diagnostic precision, ultimately facilitating personalized and efficient patient care.
Inayatul Haq, Haomin Liang, Rashid Khan, Roland Eils, Bingding Huang
Image Vis. Comput.9
2024 NanoBERTa-ASP: predicting nanobody paratope based on a pretrained RoBERTa model
abstract
BACKGROUND: Nanobodies, also known as VHH or single-domain antibodies, are unique antibody fragments derived solely from heavy chains. They offer advantages of small molecules and conventional antibodies, making them promising therapeutics. The paratope is the specific region on an antibody that binds to an antigen. Paratope prediction involves the identification and characterization of the antigen-binding site on an antibody. This process is crucial for understanding the specificity and affinity of antibody-antigen interactions. Various computational methods and experimental approaches have been developed to predict and analyze paratopes, contributing to advancements in antibody engineering, drug development, and immunotherapy. However, existing predictive models trained on traditional antibodies may not be suitable for nanobodies. Additionally, the limited availability of nanobody datasets poses challenges in constructing accurate models. METHODS: To address these challenges, we have developed a novel nanobody prediction model, named NanoBERTa-ASP (Antibody Specificity Prediction), which is specifically designed for predicting nanobody-antigen binding sites. The model adopts a training strategy more suitable for nanobodies, based on an advanced natural language processing (NLP) model called BERT (Bidirectional Encoder Representations from Transformers). To be more specific, the model utilizes a masked language modeling approach named RoBERTa (Robustly Optimized BERT Pretraining Approach) to learn the contextual information of the nanobody sequence and predict its binding site. RESULTS: NanoBERTa-ASP achieved exceptional performance in predicting nanobody binding sites, outperforming existing methods, indicating its proficiency in capturing sequence information specific to nanobodies and accurately identifying their binding sites. Furthermore, NanoBERTa-ASP provides insights into the interaction mechanisms between nanobodies and antigens, contributing to a better understanding of nanobodies and facilitating the design and development of nanobodies with therapeutic potential. CONCLUSION: NanoBERTa-ASP represents a significant advancement in nanobody paratope prediction. Its superior performance highlights the potential of deep learning approaches in nanobody research. By leveraging the increasing volume of nanobody data, NanoBERTa-ASP can further refine its predictions, enhance its performance, and contribute to the development of novel nanobody-based therapeutics. Github repository: https://github.com/WangLabforComputationalBiology/NanoBERTa-ASP.
Shangru Li, Xiangpeng Meng, Bingding Huang, Xin Wang 0168
BMC Bioinform.4
2024 Correction: NanoBERTa-ASP: predicting nanobody paratope based on a pretrained RoBERTa model
Shangru Li, Xiangpeng Meng, Bingding Huang, Xin Wang 0168
BMC Bioinform.4
2024 Representation Learning and Reinforcement Learning for Dynamic Complex Motion Planning System
abstract
Indoor motion planning challenges researchers because of the high density and unpredictability of moving obstacles. Classical algorithms work well in the case of static obstacles but suffer from collisions in the case of dense and dynamic obstacles. Recent reinforcement learning (RL) algorithms provide safe solutions for multiagent robotic motion planning systems. However, these algorithms face challenges in convergence: slow convergence speed and suboptimal converged result. Inspired by RL and representation learning, we introduced the ALN-DSAC: a hybrid motion planning algorithm where attention-based long short-term memory (LSTM) and novel data replay combine with discrete soft actor-critic (SAC). First, we implemented a discrete SAC algorithm, which is the SAC in the setting of discrete action space. Second, we optimized existing distance-based LSTM encoding by attention-based encoding to improve the data quality. Third, we introduced a novel data replay method by combining the online learning and offline learning to improve the efficacy of data replay. The convergence of our ALN-DSAC outperforms that of the trainable state of the arts. Evaluations demonstrate that our algorithm achieves nearly 100% success with less time to reach the goal in motion planning tasks when compared to the state of the arts. The test code is available at https://github.com/CHUENGMINCHOU/ALN-DSAC.
Chengmin Zhou, Bingding Huang, Pasi Fränti
IEEE Trans. Neural Networks Learn. Syst.2
2023 A Framework based on Deep Neural Network for Ranking-oriented Software Defect Prediction
abstract
Software systems are getting larger and more complex than ever before. In order to improve software reliability, software defect prediction is applied to assist developers in bug discovery. The ranking-oriented software defect prediction aims to rank software modules according to the predicted defect counts. However, existing ranking-oriented defect prediction models are constructed based on traditional hand-crafted features, which might overlook the rich syntactic information buried inside the source codes. In this paper, we propose a universal deep learning-based framework called United Deep Network for ranking-oriented software defect prediction. This framework utilizes deep neural networks to automatically generate features from source code with the syntactic and structural information preserved, and it can combine extracted features with traditional hand-crafted features in order to take advantage of both kinds of features to construct prediction models. Experimental results over 29 sets of data show the good performance of the proposed framework for building ranking-oriented defect prediction models.
Jiapeng Dai, Xiaoxing Yang, Bingding Huang, Xiaofen Lu
QRS3
2022 Effects of haze and dehazing on deep learning-based vision models
Haseeb Hassan, Pranshu Mishra, Muhammad Ahmad 0002, Ali Kashif Bashir, Bingding Huang, Bin Luo 0001
Appl. Intell.5
2019 BreakID: genomics breakpoints identification to detect gene fusion events using discordant pairs and split reads
abstract
SUMMARY: Here we developed a tool called Breakpoint Identification (BreakID) to identity fusion events from targeted sequencing data. Taking discordant read pairs and split reads as supporting evidences, BreakID can identify gene fusion breakpoints at single nucleotide resolution. After validation with confirmed fusion events in cancer cell lines, we have proved that BreakID can achieve high sensitivity of 90.63% along with PPV of 100% at sequencing depth of 500× and perform better than other available fusion detection tools. We anticipate that BreakID will have an extensive popularity in the detection and analysis of fusions involved in clinical and research sequencing scenarios. AVAILABILITY AND IMPLEMENTATION: Source code is freely available at https://github.com/SinOncology/BreakID. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Linfang Jin, Jinhuo Lai, Yang Zhang 0056, Ying Fu 0002, Shuhang Wang, Heng Dai, Bingding Huang
Bioinform.7
2011 Identification of cavities on protein surface using multiple computational approaches for drug binding site prediction
abstract
MOTIVATION: Protein-ligand binding sites are the active sites on protein surface that perform protein functions. Thus, the identification of those binding sites is often the first step to study protein functions and structure-based drug design. There are many computational algorithms and tools developed in recent decades, such as LIGSITE(cs/c), PASS, Q-SiteFinder, SURFNET, and so on. In our previous work, MetaPocket, we have proved that it is possible to combine the results of many methods together to improve the prediction result. RESULTS: Here, we continue our previous work by adding four more methods Fpocket, GHECOM, ConCavity and POCASA to further improve the prediction success rate. The new method MetaPocket 2.0 and the individual approaches are all tested on two datasets of 48 unbound/bound and 210 bound structures as used before. The results show that the average success rate has been raised 5% at the top 1 prediction compared with previous work. Moreover, we construct a non-redundant dataset of drug-target complexes with known structure from DrugBank, DrugPort and PDB database and apply MetaPocket 2.0 to this dataset to predict drug binding sites. As a result, >74% drug binding sites on protein target are correctly identified at the top 3 prediction, and it is 12% better than the best individual approach. AVAILABILITY: The web service of MetaPocket 2.0 and all the test datasets are freely available at http://projects.biotec.tu-dresden.de/metapocket/ and http://sysbio.zju.edu.cn/metapocket.
Zengming Zhang, Biaoyang Lin, Michael Schroeder 0001, Bingding Huang
Bioinform.5