EDBT 2026 Demo / reviewers in the wild / expert
Jianguo Chen 0001
dblp:18/733-1
· DBLP profile ↗
55ranked-venue papers
13as first author
49since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 15 · 1 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 11 since 2021Databases, data management, data science and information retrieval · 5 · 4 first-author · 3 since 2021Computer networks · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Edge large language models: a comprehensive survey
Shan Jiang 0005, Xuecheng Zhou, Mingjin Zhang, Changfu Xu, Guocheng Liao, Jianguo Chen 0001, Jiannong Cao 0001 |
CCF Trans. Pervasive Comput. Interact. | 6 |
| 2026 | iJTyper: An effective type inference framework for incomplete java codes by integrating constraint- and statistics-based methods
Zhixiang Chen 0016, Anji Li 0001, Neng Zhang 0001, Jianguo Chen 0001, Yuan Huang 0002, Zibin Zheng |
Expert Syst. Appl. | 4 |
| 2026 | A multilevel alignment and cross-fusion knowledge distillation framework for vision transformer-based medical image segmentation
Pengchen Liang, Jianguo Chen 0001, Renkai Wu, Zhuangzhuang Chen, Bin Pu, Qing Chang 0004, Guo Ran |
Future Gener. Comput. Syst. | 2 |
| 2026 | Adaptive-oriented mutation snake optimizer for scheduling budget-constrained workflows in heterogeneous cloud environments
Yanfen Zhang, Longxin Zhang, Buqing Cao, Jing Liu 0032, Jianguo Chen 0001, Keqin Li 0001 |
Future Gener. Comput. Syst. | 6 |
| 2026 | Multi-agent reinforcement learning for resource allocation in NOMA-enhanced aerial edge computing networks
Longxin Zhang, Xiaotong Lu, Jing Liu 0032, Yanfen Zhang, Jianguo Chen 0001, Buqing Cao, Keqin Li 0001 |
J. Syst. Archit. | 5 |
| 2026 | Task-specific knowledge distillation from the vision foundation model for enhanced medical image segmentation
Pengchen Liang, Haishan Huang, Bin Pu, Quanhong Zeng, Jianguo Chen 0001 |
Knowl. Based Syst. | 5 |
| 2026 | Collaborative Coarse-to-Fine Disease Learning With Discharge Summary Awareness for EHR Event PredictionabstractDeep learning-based models have been widely used to predict electronic health record (EHR) events by exploiting diagnostic characteristics. Despite significant progress, three limitations remain: 1) effectively modeling dynamic relationships among diseases, 2) fully leveraging diagnosis code ontologies from multiple perspectives, and 3) incorporating unstructured discharge summaries. To address these challenges, we propose a coarse-to-fine disease learning framework with patient notes for EHR event prediction, tailored to capture both dynamic and static disease characteristics. First, we construct a fine-grained dynamic disease graph by removing disease weakly correlated disease pairs based on co-occurrence distributions. Second, disease embeddings are refined by integrating coarse and fine-grained information within the hierarchical structure of ICD-9-CM codes. In addition, discharge summaries are combined with auxiliary patient notes for collaborative disease learning. Finally, gated recurrent units, location-based attention, and soft attention mechanisms are utilized to further enhance embedding representations. Experiments on two real-world EHR datasets, MIMIC-III and MIMIC-IV, demonstrate that our model consistently outperforms nine baseline methods in EHR prediction. The source code can be found at https://github.com/YNU-L/CCDLD. Yan Kang 0003, Zhuolun Li, Bin Pu, Xingbo Dong, Jiewen Yang, Lei Zhao 0013, Benteng Ma, Ningshu Li, Jianguo Chen 0001, Philip S. Yu |
IEEE Trans. Cybern. | 9 |
| 2026 | Not All Data are What You Need: A Data-Efficient Training Method Using Heterogeneous Hardware
Zulong Diao, Mingyu Qiao, Xin Wang 0001, Guangxing Zhang, Wei Liang 0005, Jianguo Chen 0001, Changhua Pei, Yanbiao Li 0001, Zhenyu Li 0001, Gaogang Xie |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2026 | A Federated Adaptive Large Language Model Fine-Tuning Framework for Software DevelopmentabstractLarge Language Models (LLMs) have achieved remarkable progress in code intelligence tasks, significantly en hancing the efficiency of software development. However, several challenges remain. First, fine-tuning LLMs for specific tasks requires a large amount of task-specific labeled data, which is often costly and time-consuming to acquire. Second, due to the sensitivity of code data, high-quality internal datasets from different organizations cannot be directly shared or combined for fine tuning. Moreover, variations in programming languages across organizations can introduce interference during the fine-tuning process. To address these challenges, we propose F-CodeLLM, a federated adaptive large language model fine-tuning framework designed for real-world software development scenarios. To the best of our knowledge, this is the first approach to apply federated learning to the fine-tuning of code LLMs, enabling collaborative model optimization while preserving the privacy of each organization's code data. We design an efficient LLM fine tuning method to mitigate the computational and communication overhead associated with collaborative fine-tuning. Experimental results demonstrate that F-CodeLLM effectively allows LLMs to learn from each organization's dataset, achieving performance comparable to centralized fine-tuning. Furthermore, F-CodeLLM is well-suited for multilingual data environments, as it can lever age shared knowledge across programming languages to enhance performance within individual language domains. Our code is publicly available at: https://github.com/AAnony/F-CodeLLM. Jianguo Chen 0001, Zeju Cai, Wenqing Chen, Zibin Zheng, Philip S. Yu |
IEEE Trans. Serv. Comput. | 1 |
| 2026 | Cooperative and Competitive Pricing in Collaborative Edge ComputingabstractA user with limited computation resources can address his delay-sensitive and computation-intensive tasks through task offloading to nearby edge servers, by purchasing both network and computation resources from profitseeking providers. We identify a substitutability property of computation and network resources for realizing the delay requirement. That is, to reduce task delay, the user can purchase more network resources to reduce transmission delay or more computation resources to reduce computation delay. This property significantly affects the user's purchase behavior and leads to strategic interactions between the computation service provider (CSP) and the network service provider (NSP), which have not been systematically studied yet. To this end, we formulate a two-stage Stackelberg game. In Stage I, one CSP and one NSP set their prices. In Stage II, each user decides offloading ratio and the amount of resources to purchase. By deriving the closed-form solutions in Stage II, we analytically conclude that the substitutability affects the user's decision through the network price to computation price ratio. We then incorporate the solution in Stage II into Stage I and analyze the service providers' pricing under two market structures. In the cooperative setting, where two service providers are integrated and jointly maximize their total profit, they would flexibly adjust the price ratio based on computation and network costs. In the competitive setting, where they are separate firms and aim to maximize their own profit, we formulate a pricing game and characterize a counter-intuitive equilibrium: the service providers would set high prices instead of low prices. Experimental results show that users benefit from service providers' competitive interactions. Guocheng Liao, Peng Sun 0003, Qian Ma 0002, Jianguo Chen 0001, Xu Chen 0004 |
IEEE Trans. Serv. Comput. | 4 |
| 2025 | Improved Convergence-relaxed Mechanism for Handling Imbalance Between Convergence and Diversity in the Decision Space in Multimodal Multi-objective optimizationabstractBalancing convergence and diversity in the decision space is essential in solving multimodal multi-objective optimization problems (MMOPs), which have multiple equivalent Pareto optimal sets (PSs) with the same Pareto optimal front (PF). For MMOPs with an imbalance between convergence and diversity in the decision space (MMOP-ICD), numerous efficient multimodal multiobjective evolutionary algorithms (MMEAs) avoid premature convergence and search for the imbalanced PS by relaxing the traditional convergence-first selection mechanism. Unfortunately, existing MMEAs suffer from convergence degradation due to excessive relaxation of the convergence-first selection mechanism. Therefore, this paper proposes an improved convergence-relaxed mechanism that includes an enhanced local convergence indicator and a two-stage mating selection. The enhanced local convergence indicator introduces the global convergence indicator into the local convergence indicator. The local convergence indicator can locate more equivalent PSs and prevent premature convergence caused by the global convergence indicator. The global convergence indicator can improve the convergence quality of the solution selected by the local convergence indicator. Then, the two-stage mating selection is used to enhance the diversity in the decision space and balance the improved convergence. Experimental results and statistical analysis show that the proposed algorithm is significantly superior to other state-of-the-art MMEAs. Zhipan Li, Wenkai Mao, Huigui Rong, Jianguo Chen 0001, Shengxu Huo, Zilu Zhao |
GECCO | 4 |
| 2025 | A Deep Reinforcement Learning Algorithm with Ordered Action Space for Budget-Aware Workflow Scheduling in Heterogeneous Clouds
Yanfen Zhang, Longxin Zhang, Lili Du, Zhihua Wen, Buqing Cao, Jianguo Chen 0001 |
ICA3PP (3) | 7 |
| 2025 | A Privacy-Preserving Edge Inference Framework for Low-Altitude UAV Swarm Intelligence
Jianguo Chen 0001, Guoqing Xiao 0001, Longxin Zhang, Guocheng Liao, Bodong Wang, Weijian You |
NPC (1) | 1 |
| 2025 | Lightweight train image fault detection model based on location information enhancement
Longxin Zhang, Runti Tan, Wenliang Zeng, Jianguo Chen 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | Enhancing long-term memory in federated class continual learning with lightweight adapters
Ji Wang 0002, Zhengyi Zhong, Weidong Bao 0001, Yaohong Zhang, Jianguo Chen 0001 |
Neurocomputing | 6 |
| 2025 | Blockchain-Empowered Federated Learning: Benefits, Challenges, and SolutionsabstractFederated learning (FL) is a distributed machine learning approach that protects user data privacy by training models locally on clients and aggregating them on a parameter server. While effective at preserving privacy, FL systems face limitations such as single points of failure, lack of incentives, and inadequate security. To address these challenges, blockchain technology is integrated into FL systems to provide stronger security, fairness, and scalability. However, blockchain-empowered FL (BC-FL) systems introduce additional demands on network, computing, and storage resources. This survey provides a comprehensive review of recent research on BC-FL systems, analyzing the benefits and challenges associated with blockchain integration. We explore why blockchain is applicable to FL, how it can be implemented, and the challenges and existing solutions for its integration. Additionally, we offer insights on future research directions for the BC-FL system. Zeju Cai, Jianguo Chen 0001, Yuting Fan, Zibin Zheng, Keqin Li 0001 |
IEEE Trans. Big Data | 2 |
| 2025 | TS-RePSO: A Three-Stage Feature Selection Method Combing ReliefF and PSO in BioinformaticsabstractThe inherent characteristics of high-dimensional feature redundancy of biomedical data lead to the "curse of dimensionality" in bioinformatics, which brings new challenges to feature selection problems. Recently, the two-stage approach combining the filter and wrapper methods has become popular for feature selection tasks. However, these two-stage or previous one-stage algorithms suffer from blindness in the setting of thresholds, and the search methods tend to fall into local optimum solutions. To this end, we propose a three-stage feature selection method that combines ReliefF and Particle swarm optimization as a specific case, called TS-RePSO, including the filter stage, grouping stage, and wrapper stage. Specifically, in the filter stage, ReliefF is utilized to compute the weights of the features and sort them in descending order. In the grouping stage, the ranked features are grouped based on the density equalization strategy so that the weight of groups in all groups is equal. In the wrapper stage, the proposed grouping PSO is employed to search for the grouped features and select them according to the in-group and out-group evaluation strategies. Extensive experiments are conducted on 5 benchmark datasets and 6 real-world datasets, and experiment results show that the proposed method achieves the best performance. Bin Pu, Haining Wang 0006, Zhaozhao Xu, Fangyuan Yang, Xiangqiong Wu, Jianguo Chen 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 6 |
| 2025 | MambaSAM: A Visual Mamba-Adapted SAM Framework for Medical Image SegmentationabstractThe Segment Anything Model (SAM) has shown exceptional versatility in segmentation tasks across various natural image scenarios. However, its application to medical image segmentation poses significant challenges due to the intricate anatomical details and domain-specific characteristics inherent in medical images. To address these challenges, we propose a novel VMamba adapter framework that integrates a lightweight, trainable Visual Mamba (VMamba) branch with the pre-trained SAM ViT encoder. The VMamba adapter accurately captures multi-scale contextual correlations, integrates global and local information, and reduces ambiguities arising from local features only. Specifically, we propose a novel cross-branch attention (CBA) mechanism to facilitate effective interaction between the SAM and VMamba branches. This mechanism enables the model to learn and adapt more efficiently to the nuances of medical images, extracting rich, complementary features that enhance its representational capacity. Beyond architectural enhancements, we streamline the segmentation workflow by eliminating the need for prompt-driven input mechanisms. This results in an autonomous prediction model that reduces manual input requirements and improves operational efficiency. In addition, our method introduces only minimal additional trainable parameters, offering an efficient solution for medical image segmentation. Extensive evaluations of four medical image datasets demonstrate that our VMamba adapter framework achieves state-of-the-art performance. Specifically, on the ACDC dataset with limited training data, our method achieves an average Dice coefficient improvement of 0.18 and reduces the Hausdorff distance by 20.38 mm compared to the AutoSAM. Pengchen Liang, Leijun Shi, Bin Pu, Renkai Wu, Jianguo Chen 0001, Lite Xu, Zhuangzhuang Chen, Qing Chang 0004 |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | SacFL: Self-Adaptive Federated Continual Learning for Resource-Constrained End DevicesabstractThe proliferation of end devices has led to a distributed computing paradigm, wherein on-device machine learning models continuously process diverse data generated by these devices. The dynamic nature of this data, characterized by continuous changes or data drift, poses significant challenges for on-device models. To address this issue, continual learning (CL) is proposed, enabling machine learning models to incrementally update their knowledge and mitigate catastrophic forgetting. However, the traditional centralized approach to CL is unsuitable for end devices due to privacy and data volume concerns. In this context, federated CL (FCL) emerges as a promising solution, preserving user data locally while enhancing models through collaborative updates. Aiming at the challenges of limited storage resources for CL, poor autonomy in task shift detection, and difficulty in coping with new adversarial tasks in the FCL scenario, we propose a novel FCL framework named self-adaptive federated CL (SacFL). $\rm {SacFL}$ employs an encoder-decoder architecture to separate task-robust and task-sensitive components, significantly reducing storage demands by retaining lightweight task-sensitive components for resource-constrained end devices. Moreover, $\rm {SacFL}$ leverages contrastive learning to introduce an autonomous data shift detection mechanism, enabling it to discern whether a new task has emerged and whether it is a benign task. This capability ultimately allows the device to autonomously trigger CL or attack defense strategy without additional information, which is more practical for end devices. Comprehensive experiments conducted on multiple text and image datasets, such as Cifar100 and THUCNews, have validated the effectiveness of $\rm {SacFL}$ in both class-incremental and domain-incremental scenarios. Furthermore, a demo system has been developed to verify its practicality. Zhengyi Zhong, Weidong Bao 0001, Ji Wang 0002, Jianguo Chen 0001, Lingjuan Lyu, Wei Yang Bryan Lim |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Joint Optimization of Scheduling Length and Cost Based on White Shark Optimization in Heterogeneous CloudsabstractIn the era of the Internet of Things, the significant increase in data volume, time, and space complexity presents great challenges to workflow scheduling in resource-constrained clouds. This study proposes an efficient hybrid algorithm, denoted white shark optimization (WSO) algorithm with budget constraints (BC-WSO), designed to adhere to budget constraints. The primary objective of BC-WSO is to optimize the scheduling length and cost. This objective is achieved by employing a heuristic algorithm that utilizes the predicted makespan matrix (PMMS) alongside the WSO algorithm as its foundation. The PMMS can minimize the scheduling length of a workflow application and satisfy the task prioritization dependencies. BC-WSO incorporates PMMS into the population initialization phase to improve the accuracy of WSO and accelerate the convergence process. Extensive experiments in two real-world scientific workflow applications show that BC-WSO outperforms current state-of-the-art meta-heuristic algorithms in simultaneously optimizing scheduling length and cost. Longxin Zhang, Minghui Ai, Yanfen Zhang, Buqing Cao, Jianguo Chen 0001, Lihua Ai |
HPCC | 5 |
| 2024 | Budget-aware Scheduling Algorithm Using Negative Offset Mechanism for Snake Optimization in Heterogeneous CloudabstractCloud computing, as a cutting-edge computing paradigm, offers substantial data processing and storage capabilities. In a heterogeneous cloud environment, the diversity among cloud platforms results in varying task execution times, posing challenges in minimizing workflow makespan under budget constraints. On this basis, a novel meta-heuristic optimization algorithm, named snake optimizer (SO), is proposed for workflow scheduling in the cloud. Then, a negative offset mechanism is designed to dynamically guide the offset of individual positions during population update to prevent falling into local optimums, thereby optimizing the search for feasible solutions and improving the success rate. Finally, using the negative offset mechanism, a snake optimization budget-aware scheduling algorithm (NO-SO) is developed to schedule budget-constrained workflows in heterogeneous cloud computing environments and minimize the makespan. A series of comparative experiments conducted on real-world scientific workflows demonstrates that the NO-SO algorithm enhances the success rate in finding a feasible solution by 38.89% and 34.45% compared with the advanced MG-PRO algorithm and the original SO algorithm, respectively. Moreover, it achieves an average reduction in makespan of 30.30% and 32.19%. Longxin Zhang, Yanfen Zhang, Xiaotong Lu, Runti Tan, Xianming Huang, Jianguo Chen 0001 |
ISPA | 6 |
| 2024 | LDD-Net: Lightweight printed circuit board defect detection network fusing multi-scale features
Longxin Zhang, Jingsheng Chen, Jianguo Chen 0001, Zhicheng Wen, Xusheng Zhou |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | Integration of preferences in multimodal multi-objective optimization
Zhipan Li, Huigui Rong, Jianguo Chen 0001, Zilu Zhao, Yupeng Huang |
Expert Syst. Appl. | 3 |
| 2024 | RSKD: Enhanced medical image segmentation via multi-layer, rank-sensitive knowledge distillation in Vision Transformer models
Pengchen Liang, Jianguo Chen 0001, Qing Chang 0004 |
Knowl. Based Syst. | 2 |
| 2024 | ParaCPI: A Parallel Graph Convolutional Network for Compound-Protein Interaction PredictionabstractIdentifying compound-protein interactions (CPIs) is critical in drug discovery, as accurate prediction of CPIs can remarkably reduce the time and cost of new drug development. The rapid growth of existing biological knowledge has opened up possibilities for leveraging known biological knowledge to predict unknown CPIs. However, existing CPI prediction models still fall short of meeting the needs of practical drug discovery applications. A novel parallel graph convolutional network model for CPI prediction (ParaCPI) is proposed in this study. This model constructs feature representation of compounds using a unique approach to predict unknown CPIs from known CPI data more effectively. Experiments are conducted on five public datasets, and the results are compared with current state-of-the-art (SOTA) models under three different experimental settings to evaluate the model's performance. In the three cold-start settings, ParaCPI achieves an average performance gain of 26.75%, 23.84%, and 14.68% in terms of area under the curve compared with the other SOTA models. In addition, the results of the experiments in the case study show ParaCPI's superior ability to predict unknown CPIs based on known data, with higher accuracy and stronger generalization compared with the SOTA models. Researchers can leverage ParaCPI to accelerate the drug discovery process. Longxin Zhang, Wenliang Zeng, Jingsheng Chen, Jianguo Chen 0001, Keqin Li 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2024 | The End-to-End Fetal Head Circumference Detection and Estimation in Ultrasound ImagesabstractIn prenatal examinations, the fetal head circumference (HC) measurement is essential for assessing fetal weight and health conditions. The sonographers obtain the fetal HC manually by fitting peripheral skull ellipse in clinical practice, which is highly subjective, time-consuming, and experience-dependent. Recently, many fetal HC automatic measurement algorithms have been proposed to improve workflow efficiency in prenatal examination. But most automatic measurement algorithms focus on using fetal head segmentation as an intermediate processing step, and HC estimation relies heavily on segmentation results, which causes the accumulation of errors in the above two stages. Independent of the segmentation method, we design a regression network to generate the oriented bounding box to detect the head contour, and directly obtain the fetal head parameters with a pixel-based ellipse regression (PER) loss. Moreover, an effective 3D attention mechanism is integrated into the network to estimate HC more precisely without adding parameters in complex ultrasound images. The extensive experimental results on the public HC18 and our clinical dataset show that the proposed network provides a feasible scheme for end-to-end estimating fetal HC, and avoids the mistake brought by the intermediary processes. Lei Zhao 0013, Ningshu Li, Guanghua Tan, Jianguo Chen 0001, Shengli Li 0001, Mingxing Duan |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2024 | MVSTT: A Multiview Spatial-Temporal Transformer Network for Traffic-Flow ForecastingabstractAccurate traffic-flow prediction remains a critical challenge due to complicated spatial dependencies, temporal factors, and unpredictable events. Most existing approaches focus on single- or dual-view learning and thus face limitations in systematically learning complex spatial-temporal features. In this work, we propose a novel multiview spatial-temporal transformer (MVSTT) network that can effectively learn complex spatial-temporal domain correlations and potential patterns from multiple views. First, we examine a temporal view and design a short-range gated convolution component from a short-term subview, and a long-range gated convolution component from a long-term subview. These two components effectively aggregate knowledge of the temporal domain at multiple granularities and mine patterns of node evolution across time steps. Meanwhile, in the spatial view, we design a dual-graph spatial learning module that captures fixed and dynamic spatial dependencies of nodes, as well as the evolution patterns of edges, from the static and dynamic graph subviews, respectively. In addition, we further design a spatial-temporal transformer to mine different levels of spatial-temporal features through multiview knowledge fusion. Extensive experiments on four real-world traffic datasets show that our method consistently outperforms the state-of-the-art baseline. The code of MVSTT is available at https://github.com/JianSoL/MVSTT. Bin Pu, Jiansong Liu, Yan Kang 0003, Jianguo Chen 0001, Philip S. Yu |
IEEE Trans. Cybern. | 4 |
| 2024 | HFSCCD: A Hybrid Neural Network for Fetal Standard Cardiac Cycle Detection in Ultrasound VideosabstractIn the fetal cardiac ultrasound examination, standard cardiac cycle (SCC) recognition is the essential foundation for diagnosing congenital heart disease. Previous studies have mostly focused on the detection of adult CCs, which may not be applicable to the fetus. In clinical practice, localization of SCCs needs to recognize end-systole (ES) and end-diastole (ED) frames accurately, ensuring that every frame in the cycle is a standard view. Most existing methods are not based on the detection of key anatomical structures, which may not recognize irrelevant views and background frames, results containing non-standard frames, or even it does not work in clinical practice. We propose an end-to-end hybrid neural network based on an object detector to detect SCCs from fetal ultrasound videos efficiently, which consists of 3 modules, namely Anatomical Structure Detection (ASD), Cardiac Cycle Localization (CCL), and Standard Plane Recognition (SPR). Specifically, ASD uses an object detector to identify 9 key anatomical structures, 3 cardiac motion phases, and the corresponding confidence scores from fetal ultrasound videos. On this basis, we propose a joint probability method in the CCL to learn the cardiac motion cycle based on the 3 cardiac motion phases. In SPR, to reduce the impact of structure detection errors on the accuracy of the standard plane recognition, we use XGBoost algorithm to learn the relation knowledge of the detected anatomical structures. We evaluate our method on the test fetal ultrasound video datasets and clinical examination cases and achieve remarkable results. This study may pave the way for clinical practices. Bin Pu, Kenli Li 0001, Jianguo Chen 0001, Yuhuan Lu 0002, Qing Zeng 0005, Jiewen Yang, Shengli Li 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | DMSTG: Dynamic Multiview Spatio-Temporal Networks for Traffic ForecastingabstractTraffic sensor networks are widely applied in smart cities to monitor traffic in real-time and record huge volumes of traffic data. Exploiting such data to forecast future traffic conditions have the potential to enhance the decision-making capabilities of intelligent transportation systems, which attracts widespread attention from both industries and academia. Among them, network-wide prediction based on graph convolutional neural networks(GCN) has become mainstream. It models the spatial dependencies of sensors in a graph with a pre-defined Laplacian matrix based on the distances among sensors. However, understanding spatio-temporal traffic patterns is quite challenging as there is a huge difference in terms of traffic patterns during different periods or in different regions. In addition, the actual data collected can be polluted due to unavoidable data loss from severe communication conditions or sensor failures. Considering these issues, we propose a novel dynamic multiview spatial-temporal prediction framework which takes into consideration various factors, including local/global, short/long term spatio-temporal dependencies and their dynamic changes. To comprehensively track the dynamic spatio-temporal dependencies among traffic data, we creatively design two different modules to perceive the changes in traffic patterns. We first propose a dynamic Laplacian matrix learning module based on our theoretical derivation to estimate the Laplacian matrix of the graph for GCN timely. We creatively incorporate tensor decomposition into this module, where real-time traffic data are decomposed into a global component that is stable and depends on long-term temporal-spatial traffic relationships and a local component that captures the traffic fluctuations. We also design a self-attention based module to dynamically assign a weight to each part in traffic data. The spatio-temporal features from multiple views are deeply fused by a feature fusion module. The forecasting performance is evaluated with 5 real-time traffic datasets. Experiment results demonstrate that our framework can consistently outperform the state-of-the-art baselines and be more robust under noisy environments. Zulong Diao, Xin Wang 0001, Da-Fang Zhang 0001, Gaogang Xie, Jianguo Chen 0001, Changhua Pei, Xuying Meng, Kun Xie 0001, Guangxing Zhang |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | TSBG: A Two-Stage Stackelberg Game Algorithm for QoE-Awareness Video Streaming TransmissionabstractDynamic Adaptive Streaming over HTTP (DASH) stands as a leading streaming technology embraced by major video platforms and smart TV manufacturers worldwide. Despite its widespread use, the inherent diversity in both the video content and the client devices poses challenges, hindering DASH from consistently delivering top-notch playback quality for all users. This oversight often leads to network congestion, compromising the playback quality for users. To tackle these issues, we propose a Two-stage Stackelberg Game (TSBG) algorithm for personalized video streaming transmission in Edge Computing (EC) environments. The TSBG algorithm aims to optimize the Quality of Experience (QoE) of users by tailoring video streaming services between EC servers and clients. Initially, we establish the system model and define the video stream transmission problem as a multi-objective optimization problem, balancing server downlink resource scheduling and client adaptive bit rate. Subsequently, we design the TSBG algorithm, where an edge server allocation mechanism is adopted in the first stage to maximize overall user QoE, while users adjust their video bit rates based on the edge server's distribution plan to enhance their individual QoE in the second stage. We prove the existence and uniqueness of the equilibrium solution of the two-stage Starkelberg game and design an optimal pricing algorithm to maximize the benefits of edge servers. Extensive simulation experiments validate the effectiveness of the TSBG algorithm, showcasing its superiority in achieving enhanced QoE, network efficiency, and fairness compared to alternative approaches. Shuzhen Xiang, Huigui Rong, Jianguo Chen 0001, Daibo Liu, Hongbo Jiang 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | A Deep Graph Network with Multiple Similarity for User Clustering in Human-Computer InteractionabstractUser counterparts, such as user attributes in social networks or user interests, are the keys to more natural Human–Computer Interaction (HCI) . In addition, users’ attributes and social structures help us understand the complex interactions in HCI. Most previous studies have been based on supervised learning to improve the performance of HCI. However, in the real world, owing to signal malfunctions in user devices, large amounts of abnormal information, unlabeled data, and unsupervised approaches (e.g., the clustering method) based on mining user attributes are particularly crucial. This paper focuses on improving the clustering performance of users’ attributes in HCI and proposes a deep graph embedding network with feature and structure similarity (called DGENFS ) to cluster users’ attributes in HCI applications based on feature and structure similarity. The DGENFS model consists of a Feature Graph Autoencoder (FGA) module, a Structure Graph Attention Network (SGAT) module, and a Dual Self-supervision (DSS) module. First, we design an attributed graph clustering method to divide users into clusters by making full use of their attributes. To take full advantage of the information of human feature space, a k-neighbor graph is generated as a feature graph based on the similarity between human features. Then, the FGA and SGAT modules are utilized to extract the representations of human features and topological space, respectively. Next, an attention mechanism is further developed to learn the importance weights of different representations to effectively integrate human features and social structures. Finally, to learn cluster-friendly features, the DSS module unifies and integrates the features learned from the FGA and SGAT modules. DSS explores the high-confidence cluster assignment as a soft label to guide the optimization of the entire network. Extensive experiments are conducted on five real-world data sets on user attribute clustering. The experimental results demonstrate that the proposed DGENFS model achieves the most advanced performance compared with nine competitive baselines. Yan Kang 0003, Bin Pu, Yongqi Kou, Yun Yang 0003, Jianguo Chen 0001, Khan Muhammad 0001, Po Yang 0001, Mohammad Hijji |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Reliability Enhancement Strategies for Workflow Scheduling Under Energy Consumption Constraints in CloudsabstractAs the demand for Big Data analysis and artificial intelligence technology continues to surge, a significant amount of research has been conducted on cloud computing services. An effective workflow scheduling strategy stands as the pivotal factor in ensuring the quality of cloud services. Dynamic voltage and frequency scaling (DVFS) is an effective energy-saving technology that is extensively used in the development of workflow scheduling algorithms. However, DVFS reduces the processor's running frequency, which increases the possibility of soft errors in workflow execution, thereby lowering the workflow execution reliability. This study proposes an energy-aware reliability enhancement scheduling (EARES) method with a checkpoint mechanism to improve system reliability while meeting the workflow deadline and the energy consumption constraints. The proposed EARES algorithm consists of three phases, namely, workflow application initialization, deadline partitioning, and energy partitioning and virtual machine selection. Numerous experiments are conducted to assess the performance of the EARES algorithm using three real-world scientific workflows. Experimental results demonstrate that the EARES algorithm remarkably improves reliability in comparison with other state-of-the-art algorithms while meeting the deadline and satisfying the energy consumption requirement. Longxin Zhang, Minghui Ai, Jianguo Chen 0001, Kenli Li 0001 |
IEEE Trans. Sustain. Comput. | 4 |
| 2023 | HN-PPISP: a hybrid network based on MLP-Mixer for protein-protein interaction site predictionabstractMOTIVATION: Biological experimental approaches to protein-protein interaction (PPI) site prediction are critical for understanding the mechanisms of biochemical processes but are time-consuming and laborious. With the development of Deep Learning (DL) techniques, the most popular Convolutional Neural Networks (CNN)-based methods have been proposed to address these problems. Although significant progress has been made, these methods still have limitations in encoding the characteristics of each amino acid in protein sequences. Current methods cannot efficiently explore the nature of Position Specific Scoring Matrix (PSSM), secondary structure and raw protein sequences by processing them all together. For PPI site prediction, how to effectively model the PPI context with attention to prediction remains an open problem. In addition, the long-distance dependencies of PPI features are important, which is very challenging for many CNN-based methods because the innate ability of CNN is difficult to outperform auto-regressive models like Transformers. RESULTS: To effectively mine the properties of PPI features, a novel hybrid neural network named HN-PPISP is proposed, which integrates a Multi-layer Perceptron Mixer (MLP-Mixer) module for local feature extraction and a two-stage multi-branch module for global feature capture. The model merits Transformer, TextCNN and Bi-LSTM as a powerful alternative for PPI site prediction. On the one hand, this is the first application of an advanced Transformer (i.e. MLP-Mixer) with a hybrid network for sequence-based PPI prediction. On the other hand, unlike existing methods that treat global features altogether, the proposed two-stage multi-branch hybrid module firstly assigns different attention scores to the input features and then encodes the feature through different branch modules. In the first stage, different improved attention modules are hybridized to extract features from the raw protein sequences, secondary structure and PSSM, respectively. In the second stage, a multi-branch network is designed to aggregate information from both branches in parallel. The two branches encode the features and extract dependencies through several operations such as TextCNN, Bi-LSTM and different activation functions. Experimental results on real-world public datasets show that our model consistently achieves state-of-the-art performance over seven remarkable baselines. AVAILABILITY: The source code of HN-PPISP model is available at https://github.com/ylxu05/HN-PPISP. Yan Kang 0003, Xinchao Wang, Bin Pu, Xuekun Yang, Yulong Rao, Jianguo Chen 0001 |
Briefings Bioinform. | 7 |
| 2023 | Hamiltonian paths and Hamiltonian cycles passing through prescribed linear forests in star graph with fault-tolerant edges
Shudan Xue, Qingying Deng, Pingshan Li, Jianguo Chen 0001 |
Discret. Appl. Math. | 4 |
| 2023 | Interval-enhanced Graph Transformer solution for session-based recommendation
Huanwen Wang, Yawen Zeng, Jianguo Chen 0001, Ning Han 0005, Hao Chen 0051 |
Expert Syst. Appl. | 3 |
| 2023 | Non-cooperative game algorithms for computation offloading in mobile edge computing environments
Jianguo Chen 0001, Qingying Deng, Xulei Yang |
J. Parallel Distributed Comput. | 1 |
| 2023 | A Hybrid Two-Stage Teaching-Learning-Based Optimization Algorithm for Feature Selection in BioinformaticsabstractThe "curse of dimensionality" brings new challenges to the feature selection (FS) problem, especially in bioinformatics filed. In this paper, we propose a hybrid Two-Stage Teaching-Learning-Based Optimization (TS-TLBO) algorithm to improve the performance of bioinformatics data classification. In the selection reduction stage, potentially informative features, as well as noisy features, are selected to effectively reduce the search space. In the following comparative self-learning stage, the teacher and the worst student with self-learning evolve together based on the duality of the FS problems to enhance the exploitation capabilities. In addition, an opposition-based learning strategy is utilized to generate initial solutions to rapidly improve the quality of the solutions. We further develop a self-adaptive mutation mechanism to improve the search performance by dynamically adjusting the mutation rate according to the teacher's convergence ability. Moreover, we integrate a differential evolutionary method with TLBO to boost the exploration ability of our algorithm. We conduct comparative experiments on 31 public data sets with different data dimensions, including 7 bioinformatics datasets, and evaluate our TS-TLBO algorithm compared with 11 related methods. The experimental results show that the TS-TLBO algorithm obtains a good feature subset with better classification performance, and indicates its generality to the FS problems. Yan Kang 0003, Haining Wang 0006, Bin Pu, Liu Tao, Jianguo Chen 0001, Philip S. Yu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2022 | MANET: Mitral Annulus Point Tracking Network in Cardiac Magnetic ResonanceabstractCardiac magnetic resonance (CMR) imaging is frequently recommended for patients at intermediate risk of cardiovascular disease to triage them for medication or invasive aggressive treatment. Mitral annulus (MA) motion and velocities represent the cardiac contraction and relaxation, and hold potential to improve the detection of subtle cardiac dysfunction. However, conventional interpretation of CMR images requires expert manipulation and is often operator-dependent. In this paper, we propose an end-to-end MA Point Tracking Network (MANet) to automatically detect and track MA motion during cardiac cycle. The MANet model consists of MA point detection module and motion tracking module. In MA point detection, we design the convolutional-based feature extraction and elastic regression to detect MA points frame by frame of each CMR video. Then, in MA tracking, we adopt the Deep SORT model to capture spatio-temporal continuity between frames and fine-tune the coordinate position of MA points. 171 CMR videos with 4275 frames are used in comparison experiments, and the results demonstrate that our MANet model achieves promising performance in reference to clinical ground truth (r=0.71, P<0.001). This work provides an important preamble for cardiac motion tracking and cardiac function evaluation. Jianguo Chen 0001, Xulei Yang, Shuang Leng, Ru-San Tan, Zeng Zeng, Liang Zhong 0001 |
ICIP | 1 |
| 2022 | A Spatiotemporal Graph Neural Network for session-based recommendation
Huanwen Wang, Yawen Zeng, Jianguo Chen 0001, Zhouting Zhao, Hao Chen 0051 |
Expert Syst. Appl. | 3 |
| 2022 | An ultrasound standard plane detection model of fetal head based on multi-task learning and hybrid knowledge graph
Lei Zhao 0013, Kenli Li 0001, Bin Pu, Jianguo Chen 0001, Shengli Li 0001, Xiangke Liao |
Future Gener. Comput. Syst. | 4 |
| 2022 | HEA-PAS: A hybrid energy allocation strategy for parallel applications scheduling on heterogeneous computing systems
Jiwu Peng, Kenli Li 0001, Jianguo Chen 0001, Keqin Li 0001 |
J. Syst. Archit. | 3 |
| 2022 | A configurable deep learning framework for medical image analysis
Jianguo Chen 0001, Mimi Zhou, Zhaolei Zhang, Xulei Yang |
Neural Comput. Appl. | 1 |
| 2022 | MobileUNet-FPN: A Semantic Segmentation Model for Fetal Ultrasound Four-Chamber Segmentation in Edge Computing EnvironmentsabstractThe apical four-chamber (A4C) view in fetal echocardiography is a prenatal examination widely used for the early diagnosis of congenital heart disease (CHD). Accurate segmentation of A4C key anatomical structures is the basis for automatic measurement of growth parameters and necessary disease diagnosis. However, due to the ultrasound imaging arising from artefacts and scattered noise, the variability of anatomical structures in different gestational weeks, and the discontinuity of anatomical structure boundaries, accurately segmenting the fetal heart organ in the A4C view is a very challenging task. To this end, we propose to combine an explicit Feature Pyramid Network (FPN), MobileNet and UNet, i.e., MobileUNet-FPN, for the segmentation of 13 key heart structures. To our knowledge, this is the first AI-based method that can segment so many anatomical structures in fetal A4C view. We split the MobileNet backbone network into four stages and use the features of these four phases as the encoder and the upsampling operation as the decoder. We build an explicit FPN network to enhance multi-scale semantic information and ultimately generate segmentation masks of key anatomical structures. In addition, we design a multi-level edge computing system and deploy the distributed edge nodes in different hospitals and city servers, respectively. Then, we train the MobileUNet-FPN model in parallel at each edge node to effectively reduce the network communication overhead. Extensive experiments are conducted and the results show the superior performance of the proposed model on the fetal A4C and femoral-length images. Bin Pu, Yuhuan Lu 0002, Jianguo Chen 0001, Shengli Li 0001, Ningbo Zhu, Wei Wei 0006, Kenli Li 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | Privacy-Preserving Deep Learning Model for Decentralized VANETs Using Fully Homomorphic Encryption and BlockchainabstractIn Vehicular Ad-hoc Networks (VANETs), privacy protection and data security during network transmission and data analysis have attracted attention. In this paper, we apply deep learning, blockchain, and fully homomorphic encryption (FHE) technologies in VANETs and propose a Decentralized Privacy-preserving Deep Learning (DPDL) model. We propose a Decentralized VANETs (DVANETs) architecture, where computing tasks are decomposed from centralized cloud services to edge computing (EC) nodes, thereby effectively reducing network communication overhead and congestion delay. We use blockchain to establish a secure and trusted data communication mechanism among vehicles, roadside units, and EC nodes. In addition, we propose a DPDL model to provide privacy-preserving data analysis for DVANET, where the FHE algorithm is used to encrypt the transportation data on each EC node and input it into the local DPDL models, thereby effectively protecting the privacy and credibility of vehicles. Moreover, we further use blockchain to provide a decentralized and trusted DPDL model update mechanism, where the parameters of each local DPDL model are stored in the blockchain for sharing with other distributed models. In this way, all distributed models can update their models in a credible and asynchronous manner, avoiding possible threats and attacks. Extensive simulations are conducted to evaluate the effectiveness, practicality, and robustness of the proposed DVANET system and DPDL models. Jianguo Chen 0001, Kenli Li 0001, Philip S. Yu |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Reliability/Performance-Aware Scheduling for Parallel Applications With Energy Constraints on Heterogeneous Computing SystemsabstractHeterogeneous Computing Systems (HCSs) have developed rapidly due to their high performance and low cost, and have been adopted by more and more applications. Energy consumption, reliability, and schedule length are the core issues of HCSs. Due to the negative correlation between frequency and reliability, DVFS-supported HCSs requires high energy consumption and a long schedule length to obtain high reliability, which resulting in performance degradation. In this paper, we focus on the reliability and performance-aware scheduling for energy-constrained parallel applications on HCSs. First, we design an energy pre-allocation mechanism based on Energy Demand Rate (EDR) to pre-allocate energy reasonably. Second, we propose an EDR-aware Maximizing Reliability of Energy-Constrained parallel applications (EMREC) scheduling algorithm. Third, considering that maximize reliability will cause the schedule length to be too long and unacceptable, we further highlight the concept of Reliability Performance Ratio (RPR). Finally, we propose a Maximizing RPR with Energy-Constrained parallel applications (MRPEC) scheduling algorithm, which enables parallel applications have a smaller schedule length while with high reliability. Extensive experimental results in real-world and randomly generated applications show the effectiveness of the proposed algorithms under different conditions. Jiwu Peng, Kenli Li 0001, Jianguo Chen 0001, Keqin Li 0001 |
IEEE Trans. Sustain. Comput. | 3 |
| 2021 | Coalition formation for deadline-constrained resource procurement in cloud computing
Junyan Hu, Kenli Li 0001, Chubo Liu, Jianguo Chen 0001, Keqin Li 0001 |
J. Parallel Distributed Comput. | 4 |
| 2021 | Dynamic Bicycle Dispatching of Dockless Public Bicycle-sharing Systems Using Multi-objective Reinforcement LearningabstractAs a new generation of Public Bicycle-sharing Systems (PBS), the Dockless PBS (DL-PBS) is an important application of cyber-physical systems and intelligent transportation. How to use artificial intelligence to provide efficient bicycle dispatching solutions based on dynamic bicycle rental demand is an essential issue for DL-PBS. In this article, we propose MORL-BD, a dynamic bicycle dispatching algorithm based on multi-objective reinforcement learning to provide the optimal bicycle dispatching solution for DL-PBS. We model the DL-PBS system from the perspective of cyber-physical systems and use deep learning to predict the layout of bicycle parking spots and the dynamic demand of bicycle dispatching. We define the multi-route bicycle dispatching problem as a multi-objective optimization problem by considering the optimization objectives of dispatching costs, dispatch truck's initial load, workload balance among the trucks, and the dynamic balance of bicycle supply and demand. On this basis, the collaborative multi-route bicycle dispatching problem among multiple dispatch trucks is modeled as a multi-agent and multi-objective reinforcement learning model. All dispatch paths between parking spots are defined as state spaces, and the reciprocal of dispatching costs is defined as a reward. Each dispatch truck is equipped with an agent to learn the optimal dispatch path in the dynamic DL-PBS network. We create an elite list to store the Pareto optimal solutions of bicycle dispatch paths found in each action, and finally get the Pareto frontier. Experimental results on the actual DL-PBS show that compared with existing methods, MORL-BD can find a higher quality Pareto frontier with less execution time. Jianguo Chen 0001, Kenli Li 0001, Keqin Li 0001, Philip S. Yu, Zeng Zeng |
ACM Trans. Cyber Phys. Syst. | 1 |
| 2021 | Dynamic Planning of Bicycle Stations in Dockless Public Bicycle-sharing System Using Gated Graph Neural NetworkabstractBenefiting from convenient cycling and flexible parking locations, the Dockless Public Bicycle-sharing (DL-PBS) network becomes increasingly popular in many countries. However, redundant and low-utility stations waste public urban space and maintenance costs of DL-PBS vendors. In this article, we propose a Bicycle Station Dynamic Planning (BSDP) system to dynamically provide the optimal bicycle station layout for the DL-PBS network. The BSDP system contains four modules: bicycle drop-off location clustering, bicycle-station graph modeling, bicycle-station location prediction, and bicycle-station layout recommendation. In the bicycle drop-off location clustering module, candidate bicycle stations are clustered from each spatio-temporal subset of the large-scale cycling trajectory records. In the bicycle-station graph modeling module, a weighted digraph model is built based on the clustering results and inferior stations with low station revenue and utility are filtered. Then, graph models across time periods are combined to create a graph sequence model. In the bicycle-station location prediction module, the GGNN model is used to train the graph sequence data and dynamically predict bicycle stations in the next period. In the bicycle-station layout recommendation module, the predicted bicycle stations are fine-tuned according to the government urban management plan, which ensures that the recommended station layout is conducive to city management, vendor revenue, and user convenience. Experiments on actual DL-PBS networks verify the effectiveness, accuracy, and feasibility of the proposed BSDP system. Jianguo Chen 0001, Kenli Li 0001, Keqin Li 0001, Philip S. Yu, Zeng Zeng |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2021 | A Domain Adaptive Density Clustering Algorithm for Data With Varying Density DistributionabstractAs one type of efficient unsupervised learning methods, clustering algorithms have been widely used in data mining and knowledge discovery with noticeable advantages. However, clustering algorithms based on density peak have limited clustering effect on data with varying density distribution (VDD), equilibrium distribution (ED), and multiple domain-density maximums (MDDM), leading to the problems of sparse cluster loss and cluster fragmentation. To address these problems, we propose a Domain-Adaptive Density Clustering (DADC) algorithm, which consists of three steps: domain-adaptive density measurement, cluster center self-identification, and cluster self-ensemble. For data with VDD features, clusters in sparse regions are often neglected by using uniform density peak thresholds, which results in the loss of sparse clusters. We define a domain-adaptive density measurement method based on K K-Nearest Neighbors (KNN) to adaptively detect the density peaks of different density regions. We treat each data point and its KNN neighborhood as a subgroup to better reflect its density distribution in a domain view. In addition, for data with ED or MDDM features, a large number of density peaks with similar values can be identified, which results in cluster fragmentation. We propose a cluster center self-identification and cluster self-ensemble method to automatically extract the initial cluster centers and merge the fragmented clusters. Experimental results demonstrate that compared with other comparative algorithms, the proposed DADC algorithm can obtain more reasonable clustering results on data with VDD, ED and MDDM features. Benefitting from a few parameter requirement and non-iterative nature, DADC achieves low computational complexity and is suitable for large-scale data clustering. Jianguo Chen 0001, Philip S. Yu |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2020 | A Decision Support System to Provide Criminal Pattern Based Suggestions to Travelers
Khin Nandar Win, Jianguo Chen 0001, Mingxing Duan, Guoqing Xiao 0001, Kenli Li 0001, Philippe Fournier-Viger, Keqin Li 0001 |
IEA/AIE | 2 |
| 2020 | Fingerprint classification and identification algorithms for criminal investigation: A survey
Khin Nandar Win, Kenli Li 0001, Jianguo Chen 0001, Philippe Fournier-Viger, Keqin Li 0001 |
Future Gener. Comput. Syst. | 3 |
| 2019 | A periodicity-based parallel time series prediction algorithm in cloud computing environments
Jianguo Chen 0001, Kenli Li 0001, Huigui Rong, Kashif Bilal, Keqin Li 0001, Philip S. Yu |
Inf. Sci. | 1 |
| 2019 | A Bi-layered Parallel Training Architecture for Large-Scale Convolutional Neural NetworksabstractBenefitting from large-scale training datasets and the complex training network, Convolutional Neural Networks (CNNs) are widely applied in various fields with high accuracy. However, the training process of CNNs is very time-consuming, where large amounts of training samples and iterative operations are required to obtain high-quality weight parameters. In this paper, we focus on the time-consuming training process of large-scale CNNs and propose a Bi-layered Parallel Training (BPT-CNN) architecture in distributed computing environments. BPT-CNN consists of two main components: (a) an outer-layer parallel training for multiple CNN subnetworks on separate data subsets, and (b) an inner-layer parallel training for each subnetwork. In the outer-layer parallelism, we address critical issues of distributed and parallel computing, including data communication, synchronization, and workload balance. A heterogeneous-aware Incremental Data Partitioning and Allocation (IDPA) strategy is proposed, where large-scale training datasets are partitioned and allocated to the computing nodes in batches according to their computing power. To minimize the synchronization waiting during the global weight update process, an Asynchronous Global Weight Update (AGWU) strategy is proposed. In the inner-layer parallelism, we further accelerate the training process for each CNN subnetwork on each computer, where computation steps of convolutional layer and the local weight training are parallelized based on task-parallelism. We introduce task decomposition and scheduling strategies with the objectives of thread-level load balancing and minimum waiting time for critical paths. Extensive experimental results indicate that the proposed BPT-CNN effectively improves the training performance of CNNs while maintaining the accuracy. Jianguo Chen 0001, Kenli Li 0001, Kashif Bilal, Xu Zhou 0001, Keqin Li 0001, Philip S. Yu |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2018 | A disease diagnosis and treatment recommendation system based on big data mining and cloud computing
Jianguo Chen 0001, Kenli Li 0001, Huigui Rong, Kashif Bilal, Keqin Li 0001 |
Inf. Sci. | 1 |
| 2017 | A Parallel Random Forest Algorithm for Big Data in a Spark Cloud Computing EnvironmentabstractWith the emergence of the big data age, the issue of how to obtain valuable knowledge from a dataset efficiently and accurately has attracted increasingly attention from both academia and industry. This paper presents a Parallel Random Forest (PRF) algorithm for big data on the Apache Spark platform. The PRF algorithm is optimized based on a hybrid approach combining dataparallel and task-parallel optimization. From the perspective of data-parallel optimization, a vertical data-partitioning method is performed to reduce the data communication cost effectively, and a data-multiplexing method is performed is performed to allow the training dataset to be reused and diminish the volume of data. From the perspective of task-parallel optimization, a dual parallel approach is carried out in the training process of RF, and a task Directed Acyclic Graph (DAG) is created according to the parallel training process of PRF and the dependence of the Resilient Distributed Datasets (RDD) objects. Then, different task schedulers are invoked for the tasks in the DAG. Moreover, to improve the algorithm's accuracy for large, high-dimensional, and noisy data, we perform a dimension-reduction approach in the training process and a weighted voting approach in the prediction process prior to parallelization. Extensive experimental results indicate the superiority and notable advantages of the PRF algorithm over the relevant algorithms implemented by Spark MLlib and other studies in terms of the classification accuracy, performance, and scalability. With the expansion of the scale of the random forest model and the Spark cluster, the advantage of the PRF algorithm is more obvious. Jianguo Chen 0001, Kenli Li 0001, Zhuo Tang, Kashif Bilal, Shui Yu 0001, Chuliang Weng, Keqin Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |