EDBT 2026 Demo / reviewers in the wild / expert
Bin Liu 0023
dblp:35/837-23
· DBLP profile ↗
20ranked-venue papers
7as first author
14since 2021 · last 2026
0000-0001-9388-0198ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FedCED: Consensus enhancement in decentralized federated learning via distillation
Bin Liu 0023, Changle Li, Qi Chu 0013, Banghao Zhai, Zeyu Ji, Keqin Li 0001 |
Knowl. Based Syst. | 1 |
| 2026 | A Thread-Level Stream Scheduling Method for Accelerating LVMs' Inference on a Resource-Constrained PlatformabstractAs a new generation of edge devices, the integrated CPU/GPU architecture has opened up new opportunities for deploying different scale vision models. In order to reduce models’ inference time on the integrated devices, this article first compresses deep learning models using model quantization. The quantization process greatly reduces the computation requirements of a model, which enables its deployment on embedded development boards. However, quantization also leads to lower GPU resource utilization during inference on integrated devices. This insufficient utilization results in slower inference speed. To address this problem, this article first depicts the data flow of model inference within an integrated device. Secondly, this article implements a unified memory management between the CPU and GPU based on managed memory strategy. Finally, this article designs a thread-level stream scheduling method to improve GPU utilization and throughput during model inference in a pipeline way. Experimental results show that the proposed method achieves a 2x-10x improvement in throughput compared with the TensorRT’s default scheduling method, which is crucial for realizing real-time inference tasks on edge devices. Bin Liu 0023, Xinzhe Zhang, Rongyu Dou, Keqin Li 0001 |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2026 | LIBPipe: Efficient Load Imbalance Pipeline Model Parallelism for Large Models TrainingabstractWith the increasing size of datasets and the expansion of Deep Neural Networks (DNNs), the training process has become exceedingly time-consuming. Distributed training, specifically the Pipeline Model Parallelism (PMP) method, commonly mitigates this problem but suffers from bubble time delays. This paper proposes LIBPipe, a pipeline training framework that explicitly incorporates a load-imbalance method to reduce bubble time in PMP. Within LIBPipe, a model-unequal-partitioning method is designed from the perspective of load imbalance to reshape the pipeline execution pattern and significantly shorten idle periods during training. On top of this method, a performance-guided unequal-partitioning search algorithm is developed to efficiently identify near-optimal partitioning strategies under memory constraints. The paper theoretically proves the time efficiency of adopting load imbalance in pipeline models. Comprehensive experiments are conducted on an 8-GPU server to evaluate the efficiency of LIBPipe, using the IMDB and mini-ImageNet datasets as well as well-known models such as BERT and ResNet. The BERT-series models achieve a maximum throughput improvement of 60.3%, while the ResNet-series models achieve a maximum improvement of 74.1%. Bin Liu 0023, Hengzhao Li, Zeyu Ji, Hongming Zhang 0002, Keqin Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2025 | PRT: An Efficient Pipeline Reuse Technology for Large Models TrainingabstractThe rapid evolution of large models and the widespread application of extensive datasets have made the cost of training increasingly prohibitive. While pipeline model parallelism makes it possible to train large models, existing pipeline techniques find it difficult to reduce bubble time due to their strong dependence on the number of GPUs for pipeline depth. This paper introduces a novel pipeline reuse technology, PRT, which breaks the limitation of pipeline depth being dependent on the number of GPUs, allowing for deeper pipelines even when the number of GPUs is limited. This paper also theoretically demonstrates the feasibility of PRT. Furthermore, the high orthogonality of PRT allows it to be implemented in both unidirectional and bidirectional pipelines, further enhancing pipeline efficiency. It is evaluated on a server equipped with 8 GPUs, using the BERT series models and ResNet series models with datasets including the IMDB dataset and the mini-ImageNet dataset. Experimental results show that for the BERT series models, unidirectional and bidirectional pipelines with PRT achieve throughput improvements of up to 54.78% and 30.38%, respectively. For the ResNet series models, the improvements reached up to 76.59% and 26.45%, respectively. Additionally, PRT achieves more balanced memory usage, validating its efficiency. Zeyu Ji, Banghao Zhai, Qi Chu 0013, Bin Liu 0023 |
CLUSTER | 5 |
| 2025 | GroPipe: A Grouped Pipeline Hybrid Parallel Method for Accelerating DCNNs TrainingabstractTraining large Deep Convolutional Neural Networks (DCNNs) with increasingly large datasets to improve model accuracy has become extremely time-consuming. Distributed training methods, such as data parallelism (DP) and pipeline model parallelism (PMP), offer potential solutions but face challenges like load imbalance and significant communication overhead. This paper introduces GroPipe, a novel architecture that synergistically integrates PMP and DP, markedly improving training speeds. GroPipe employs an automatic model partitioning algorithm based on a performance projection technique, ensuring load balance and facilitating quantitative performance evaluation in PMP. Additionally, it adopts a group-based delayed asynchronous communication strategy to efficiently reduce communication overhead in DP. Using the ResNet and VGG models with the ImageNet dataset, extensive experiments are performed on an 8-GPU server and demonstrate GroPipe’s effectiveness. GroPipe achieves substantial improvements in time to accuracy, showing an average improvement of 42.2% and 14.0% on the ResNet series, and 79.2% and 43.9% on the VGG series, without compromising Top-1 accuracy. Bin Liu 0023, Yongyao Ma, Zeyu Ji, Zhenli He, Keqin Li 0001 |
IEEE Trans. Computers | 1 |
| 2024 | Incremental RPN: Hierarchical Region Proposal Network for Apple Leaf Disease Detection in Natural EnvironmentsabstractApple leaf diseases can seriously affect apple production and quality, and accurately detecting them can improve the efficiency of disease monitoring. Owing to the complex natural growth environment, apple leaf lesions may be easily confused with background noise, leading to poor performance. In this study, a cascaded Incremental Region Proposal Network (Inc-RPN) is proposed to accurately detect apple leaf diseases in natural environments. The proposed Inc-RPN has a two-layer RPN architecture, where the precursor RPN is leveraged to generate diseased leaf proposals, and the successor RPN focuses on extracting target disease spots based on diseased leaf proposals. In the successor RPN, a low-level feature aggregation module is designed to fully utilize the bridged features and preserve the semantic information of the target disease spots. An incremental module is also leveraged to extract aggregated diseased leaf features and target disease spot features. Finally, a novel position anchor generator is designed to generate anchors based on diseased leaf proposals. The experimental results show that the proposed Inc-RPN performs very well on the FALD_CED and Apple Leaf Disease datasets, showing that it can accurately perform apple leaf disease detection tasks. Haixi Zhang, Chenyan Lv, Haibin Han, Bin Liu 0023 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2024 | LBB: load-balanced batching for efficient distributed learning on heterogeneous GPU cluster
Feixiang Yao, Zeyu Ji, Bin Liu 0023, Haoyuan Gao |
J. Supercomput. | 4 |
| 2023 | DAC-PPYOLOE+: A Lightweight Real-time Detection Model for Early Apple Leaf Pests and Diseases under Complex Background
Bin Liu 0023, Xinyue Su, Chenxi Song, Zhuohan Yao, Haixi Zhang |
COMPSAC | 1 |
| 2023 | A new local search algorithm with greedy crossover restart for the dominating tree problem
Dangdang Niu, Bin Liu 0023, Minghao Yin, Yupeng Zhou |
Expert Syst. Appl. | 2 |
| 2023 | VMF-SSD: A Novel V-Space Based Multi-Scale Feature Fusion SSD for Apple Leaf Disease DetectionabstractApple leaf diseases seriously affect the quality of apples and may lead to yield losses, detecting apple leaf diseases accurately can prevent diseases from spreading and promote the healthy growth of the industry. However, recent studies cannot achieve accurate detection of leaf diseases with high accuracy because the lesions are of different sizes. So, this paper proposed a novel apple leaf disease detection method called VMF-SSD (V-space-based Multi-scale Feature-fusion SSD), which is designed to extract more reliable multi-scale feature representations for varied sizes of diseased spots and improve the final detection performance. The multi-scale feature extraction is established with multi-scale feature representation to further improve the disease detection performance, especially for small spots. After that, a V-space-based location branch is presented to enhance the texture feature information and help further identify disease spot location. Finally, attention mechanisms are utilized to automatically learn the importance of feature channels at different scales for distinguishing diseased spots of different sizes. Experimental results showed that the VMF-SSD method achieves 83.19% mAP and obtains the detection speed of 27.53 FPS on the test set, which indicates that the proposed VMF-SSD method can achieve competitive performance on apple leaf diseases detection task and satisfy the requirements of agricultural production applications. Liangliang Tian, Haixi Zhang, Bin Liu 0023, Nannan Duan, Aihong Yuan, Yingqiu Huo |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2023 | LAD-Net: A Novel Light Weight Model for Early Apple Leaf Pests and Diseases ClassificationabstractAphids, brown spots, mosaics, rusts, powdery mildew and Alternaria blotches are common types of early apple leaf pests and diseases that severely affect the yield and quality of apples. Recently, deep learning has been regarded as the best classification model for apple leaf pests and diseases. However, these models with large parameters have difficulty providing an accurate and fast diagnosis of apple leaf pests and diseases on mobile terminals. This paper proposes a novel and real-time early apple leaf disease recognition model. AD Convolution is firstly utilized to replace standard convolution to make smaller number of parameters and calculations. Meanwhile, a LAD-Inception is built to enhance the ability of extracting multiscale features of different sizes of disease spots. Finally, the LAD-Net model is built by the LR-CBAM and the LAD-Inception modules, replacing a full connection with global average pooling to further reduce parameters. The results show that the LAD-Net, with a size of only 1.25MB, can achieve a recognition performance of 98.58%. Additionally, it is only delayed by 15.2ms on HUAWEI P40 and by 100.1ms on Jetson Nano, illustrating that the LAD-Net can accurately recognize early apple leaf pests and diseases on mobile devices in real-time, providing portable technical support. Xianyu Zhu, Runchang Jia, Bin Liu 0023, Zhuohan Yao, Aihong Yuan, Yingqiu Huo, Haixi Zhang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2022 | Apple-YOLO: A Novel Mobile Terminal Detector Based on YOLOv5 for Early Apple Leaf DiseasesabstractEarly detection of apple leaf diseases is the basis for timely precautions, which can inhibit the spread of the diseases and minimize the severe economic loss. Nowadays, CNN-based models are used for apple leaf diseases detection. However, due to the large model size and inference delay, the model is challenging to be transplanted to mobile terminals with good detection performance. This paper proposes a lightweight detection model Apple-YOLO on mobile terminals for real-time apple leaf diseases detection. First, a dataset named AppleSet8 is constructed using digital image processing and Mosaic data augmentation to improve the robustness and generalization ability of the model. Then the double-branch Apple-CSP module is presented to reduce the model parameters and guarantee feature extraction capability. Fur-thermore, the improved FDSA (Focus layer with depthwise separable convolution and attention mechanism) module effectively decreases the model's FLOPs and enhances the network's attention to the disease spots. Finally, the Skip-Spp (Skip-connection and Spatial pyramid pooling) module is built to strengthen the detection performance for multi-scale disease spots. The experiment results show that mobile-based Apple-YOLO has achieved 96.04% mAP, the inference speed of 34 FPS, and the size is only 5.33 ME, indicating that Apple-YOLO is suitable for the real-time detection of early apple leaf diseases in the real scenario. Xianyu Zhu, Runchang Jia, Bin Liu 0023, Cong Yu 0016 |
COMPSAC | 4 |
| 2022 | Improving local search for the weighted sum coloring problem using the branch-and-bound algorithm
Dangdang Niu, Bin Liu 0023, Hongming Zhang 0002, Minghao Yin |
Knowl. Based Syst. | 2 |
| 2021 | CGAN-IRB: A Novel Data Augmentation Method for Apple Leaf DiseasesabstractAt present, the identification of apple leaf diseases plays an important role in controlling apple leaf diseases and improving apple yield. CNNs(Convolutional Neural Networks) have been widely used in apple leaf diseases identification, but the training of the CNNs requires a large number of images. The lack of images would make the CNNs hard to generalize. Thus the CNNs are unable to recognize new disease images. Focusing on this problem, this paper proposes a new model named CGAN-IRB(Conditional Generative Adversarial Network with the Improved Residual Block) for data augmentation. Firstly, various improvements have been made based on CGAN to generate high-quality, robust, and specific-category images of apple leaf diseases. Among which the embedding of the residual block has been found to significantly improve the model performance. Then the interpolation algorithm is used instead of deconvolution to increase the image size. Finally, the TTUR(Two-Timescale Update Rule) training strategy is employed and all the convolutional layers of the network are spectrally normalized to stabilize the training of the network. The performance of CGAN-IRB was tested both on image generation and classification tasks. Experiment results show that the images generated by the network possess high quality and robust features, pro-viding a novel solution for the data augmentation of apple leaf diseases. The new GAN-based data augmentation method leads to significant improvements in the classification accuracy of CNNs. In the case of all tested CNNs, the classification accuracy improvements are 11.75% and 2.17% on average over non-augmented and traditional-augmented, respectively. Among them, the classification accuracy of GoogLeNet V2 and ShuffleNet V2 is 99.34% and 99.67%, respectively. The data augmentation approach proposed in this paper can be used more widely in the field of disease identification, solving the problem of insufficient data sets, and can be extended to related fields where data sets are difficult to obtain. Xinbin Yuan, Cong Yu 0016, Bin Liu 0023, Henan Sun, Xianyu Zhu |
COMPSAC | 3 |
| 2020 | Kiwifruit Leaf Disease Identification Using Improved Deep Convolutional Neural NetworksabstractBrown spot, Mosaic and Anthracnose are three common kiwifruit leaf diseases, which causes serious economic losses in the kiwifruit industry. The timely and precise identification approach of kiwifruit leaf diseases is significant for controlling the spread of disease and ensuring the healthy growth of the kiwifruit industry. In this paper, a novel identification approach based on improved convolutional neural networks is proposed for kiwifruit leaf diseases. A dataset consisting of 11322 kiwifruit leaf images is firstly generated using image augmentation. And then, a novel CNNs-based model named Kiwi-ConvNet is built with Kiwi-Inception structures and dense connectivity strategy, which can enhance the capability of multi-scale feature extraction and ensure multi-dimensional feature fusion. Under the hold-out test set, the experimental results show that the proposed model realizes an accuracy of 98.54%, gaining a better accuracy of 2.29% and 9.51% than GoogLeNet and ResNet-20 respectively. This research indicates that the proposed model achieves accurate diagnosis of kiwifruit leaf diseases automatically, and provides a viable solution in the field of crop leaf disease identification with high recognition accuracy. Bin Liu 0023, Zefeng Ding, Dongjian He, Jinrong He |
COMPSAC | 1 |
| 2019 | A Parallel Retinex Image Enhancement Algorithm Based on OpenMP
Shixiong Cheng, Bin Liu 0023, Dongjian He, Jinrong He, Yanning Du |
NPC | 2 |
| 2015 | Optimization of thread partitioning parameters in speculative multithreading based on artificial immune algorithmabstractThread partition plays an important role in speculative multithreading (SpMT) for automatic parallelization of irregular programs. Using unified values of partition parameters to partition different applications leads to the fact that every application cannot own its optimal partition scheme. In this paper, five parameters affecting thread partition are extracted from heuristic rules. They are the dependence threshold (DT), lower limit of thread size (TSL), upper limit of thread size (TSU), lower limit of spawning distance (SDL), and upper limit of spawning distance (SDU). Their ranges are determined in accordance with heuristic rules, and their step-sizes are set empirically. Under the condition of setting speedup as an objective function, all combinations of five threshold values form the solution space, and our aim is to search for the best combination to obtain the best thread granularity, thread dependence, and spawning distance, so that every application has its best partition scheme. The issue can be attributed to a single objective optimization problem. We use the artificial immune algorithm (AIA) to search for the optimal solution. On Prophet, which is a generic SpMT processor to evaluate the performance of multithreaded programs, Olden benchmarks are used to implement the process. Experiments show that we can obtain the optimal parameter values for every benchmark, and Olden benchmarks partitioned with the optimized parameter values deliver a performance improvement of 3.00% on a 4-core platform compared with a machine learning based approach, and 8.92% compared with a heuristics-based approach. Yinliang Zhao, Bin Liu 0023 |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2014 | Similar Samples Cleaning in Speculative Multithreading
Yinliang Zhao, Bin Liu 0023 |
ICA3PP (2) | 3 |
| 2014 | A thread partitioning approach for speculative multithreading
Bin Liu 0023, Yingliang Zhao, Yanjun Sun, Boqin Feng |
J. Supercomput. | 1 |
| 2012 | A Virtual Sample Generation Approach for Speculative Multithreading Using Feature Sets and Abstract Syntax TreesabstractSpeculative multithreading (SpMT) is a thread level automatic parallelization technique to accelerate sequential programs. Since approaches based on heuristic rules only get the local optimal speculative thread solution and have reached their speedup performance limit, machine learning approaches have been introduced into speculative multithreading to avoid the shortcomings of the heuristic rules relied on experience. However, few irregular programs can meet the need for training model of machine learning. To solve this problem, we first build feature sets based on Olden benchmarks and then disturb them into new sets. With the new sets, virtual samples are generated by abstract syntax trees (ASTs). By this means, we effectively resolve the shortage of samples for speculative multithreading based on machine learning. On Prophet, which is a generic SpMT processor to evaluate the performance of multithread programs, the validity of virtual samples is verified and reaches an average speedup of 1.47. Experiments show that the virtual samples can simulate a variety of procedure structures of Olden benchmarks and this sample generation technique can provide sufficient samples for training model. Bin Liu 0023, Yingliang Zhao, Meirong Li, Yanzhao Liu, Boqin Feng |
PDCAT | 1 |