Gangyong Jia

dblp:57/8298 · DBLP profile ↗
← Back
62ranked-venue papers
18as first author
32since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 8 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 3 first-author · 11 since 2021Computer networks · 9 · 5 first-author · 1 since 2021Artificial intelligence and machine learning · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 LungNoduleAgent: A Collaborative Multi-Agent System for Precision Diagnosis of Lung Nodules
abstract
Diagnosing lung cancer typically involves physicians identifying lung nodules in Computed tomography (CT) scans and generating diagnostic reports based on their morphological features and medical expertise. Although advancements have been made in using multimodal large language models for analyzing lung CT scans, challenges remain in accurately describing nodule morphology and incorporating medical expertise. These limitations affect the reliability and effectiveness of these models in clinical settings. Collaborative multi-agent systems offer a promising strategy for achieving a balance between generality and precision in medical applications, yet their potential in pathology has not been thoroughly explored. To bridge these gaps, we introduce LungNoduleAgent, an innovative collaborative multi-agent system specifically designed for analyzing lung CT scans. LungNoduleAgent streamlines the diagnostic process into sequential components, improving precision in describing nodules and grading malignancy through three primary modules. The first module, the Nodule Spotter, coordinates clinical detection models to accurately identify nodules. The second module, the Radiologist, integrates localized image description techniques to produce comprehensive CT reports. Finally, the Doctor Agent System performs malignancy reasoning by using images and CT reports, supported by a pathology knowledge base and a multi-agent system framework. Extensive testing on two private datasets and the public LIDC-IDRI dataset indicates that LungNoduleAgent surpasses mainstream vision-language models, agent systems, and advanced expert models such as GPT-4o, Claude 3.7 Sonnet, LLaMA-3.2 Vision, Qwen2.5-VL, Med-R1, MedGemma, MedAgent-Pro, MedAgents, MDAgent and LLaVA-Med. These results highlight the importance of region-level semantic alignment and multi-agent collaboration in diagnosing nodules. LungNoduleAgent stands out as a promising foundational tool for supporting clinical analyses of lung nodules.
Yaoqun Liu, Fenglei Fan, Dajiang Lei, Gangyong Jia, Changmiao Wang, Ruiquan Ge
AAAI8
2026 A plug-and-play intra-class variance suppression framework for industrial anomaly detection
abstract
Anomaly detection methods leveraging unsupervised learning are expected to find broad application across diverse sectors, especially in inspecting defects of industrial products. This potential is largely due to their resilience against the unpredictability of anomaly types and the imbalances of learning data across classes. Central to these methods is the premise that feature extractors or image reconstructors, when trained solely on normal data, are incapable of fully replicating the features or inputs of anomalous data. As a result, anomalies could be detected by thresholding the deviations in the extracted features or the reconstructed outputs. However, finding an optimal threshold that effectively separates anomalous from normal data remains a substantial challenge in real-world scenarios. The inherent variability within normal data itself is a significant factor contributing to this challenge. In this study, we introduce a simple yet powerful intra-class variance suppression framework that enables anomaly detection models to suppress intra-class variability by learning compact representations of normal data. We evaluate the proposed framework on three established unsupervised anomaly detection paradigms, namely generative adversarial learning, knowledge distillation, and reverse distillation. Experiments are conducted on multiple benchmark datasets, including handwritten digit images, natural object images, industrial anomaly detection benchmarks, and two additional real-world industrial datasets. The results demonstrate that the proposed framework consistently improves anomaly detection and localization performance, particularly in practical industrial quality inspection scenarios.
Yixuan Ju, Prawit Buayai, Gangyong Jia, Xiaoyang Mao
Eng. Appl. Artif. Intell.4
2025 3D-Telepathy: Reconstructing 3D Objects from EEG Signals
Yuxiang Ge, Jionghao Cheng, Ruiquan Ge, Zhaojie Fang, Gangyong Jia, Nannan Li 0001, Ahmed El-Azab, Changmiao Wang
IJCNN5
2025 CRISP-SAM2: SAM2 with Cross-Modal Interaction and Semantic Prompting for Multi-Organ Segmentation
abstract
Multi-organ medical segmentation is a crucial component of medical image processing, essential for doctors to make accurate diagnoses and develop effective treatment plans. Despite significant progress in this field, current multi-organ segmentation models often suffer from inaccurate details, dependence on geometric prompts and loss of spatial information. Addressing these challenges, we introduce a novel model named CRISP-SAM2 with CR oss-modal Interaction and Semantic Prompting based on SAM2. This model represents a promising approach to multi-organ medical segmentation guided by textual descriptions of organs. Our method begins by converting visual and textual inputs into cross-modal contextualized semantics using a progressive cross-attention interaction mechanism. These semantics are then injected into the image encoder to enhance the detailed understanding of visual information. To eliminate reliance on geometric prompts, we use a semantic prompting strategy, replacing the original prompt encoder to sharpen the perception of challenging targets. In addition, a similarity-sorting self-updating strategy for memory and a mask-refining process is applied to further adapt to medical imaging and enhance localized details. Comparative experiments conducted on seven public datasets indicate that CRISP-SAM2 outperforms existing models. Extensive analysis also demonstrates the effectiveness of our method, thereby confirming its superior performance, especially in addressing the limitations mentioned earlier. Our code is available at: https://github.com/YU-deep/CRISP_SAM2.git.
Changmiao Wang, Ahmed El-Azab, Gangyong Jia, Changqing Zou, Ruiquan Ge
ACM Multimedia5
2025 LPUWF-LDM: Enhanced latent diffusion model for precise late-phase UWF-FA generation on limited dataset
Zhaojie Fang, Guanyu Zhou, Ke Zhuang, Yifei Chen 0019, Ruiquan Ge, Changmiao Wang, Gangyong Jia, Qing Wu 0008, Juan Ye, Maimaiti Nuliqiman, Peifang Xu, Ahmed El-Azab
Expert Syst. Appl.8
2025 ICH-PRNet: a cross-modal intracerebral haemorrhage prognostic prediction method using joint-attention interaction mechanism
Ahmed El-Azab, Ruiquan Ge, Jichao Zhu, Gangyong Jia, Qing Wu 0008, Changmiao Wang
Neural Networks6
2025 Spatial Temporal Attention-Based Target Vehicle Trajectory Prediction for Internet of Vehicles
abstract
Forecasting vehicle behavior within complex traffic environments is pivotal within Intelligent Transportation Systems (ITS). Though this technology plays a significant role in alleviating the prevalent operational difficulties in logistics and transportation systems, the precise prediction of vehicle trajectories still poses a substantial challenge. To address this, our study introduces the Spatio Temporal Attention-based methodology for Target Vehicle Trajectory Prediction (STATVTPred). This approach integrates Global Positioning System(GPS) localization technology to track target movement and dynamically predict the vehicle’s future path using comprehensive spatio-temporal trajectory data. We map the vehicle trajectory onto a directed graph, after which spatial attributes are extracted via a Graph Attention Networks(GATs). The Transformer technology is employed to yield temporal features from the sequence. These elements are then amalgamated with local road network structure maps to filter and deliver a smooth trajectory sequence, resulting in precise vehicle trajectory prediction.This study validates our proposed STATVTPred method on T-Drive and Chengdu taxi-trajectory datasets. STATVTPred achieves an AMR of 73.07% on the Beijing dataset, surpassing the Transformer by 6.38% and the LSTM Encoder-Decoder by 37.45%, while also reducing Distance Error (DE) by 26.93% and 20.95% in Beijing and Chengdu, respectively, also much lower than the baseline results. This is expected to establish STATVTPred as a new approach for handling trajectory prediction of targets in logistics and transportation scenarios, thereby enhancing prediction accuracy.
Ouhan Huang, Huanle Rao, Tianyun Wang, Aolong Sun, Sizhe Xing, Gangyong Jia
IEEE Trans Autom. Sci. Eng.8
2025 Unlabeled data augmentation with diffusion model for semi-supervised object detection
Zhanyun Lu, Renshu Gu, Huimin Cheng, Peifang Xu, Yuichiro Kinoshita, Juan Ye, Gangyong Jia, Qing Wu 0008
Vis. Comput.8
2024 ICH-SCNet: Intracerebral Hemorrhage Segmentation and Prognosis Classification Network Using CLIP-guided SAM mechanism
abstract
Intracerebral hemorrhage (ICH) is the most fatal subtype of stroke and is characterized by a high incidence of disability. Accurate segmentation of the ICH region and prognosis prediction are critically important for developing and refining treatment plans for post-ICH patients. However, existing approaches address these two tasks independently and predominantly focus on imaging data alone, thereby neglecting the intrinsic correlation between the tasks and modalities. This paper introduces a multi-task network, ICH-SCNet, designed for both ICH segmentation and prognosis classification. Specifically, we integrate a SAM-CLIP cross-modal interaction mechanism that combines medical text and segmentation auxiliary information with neuroimaging data to enhance cross-modal feature recognition. Additionally, we develop an effective feature fusion module and a multi-task loss function to improve performance further. Extensive experiments on an ICH dataset reveal that our approach surpasses other state-of-the-art methods. It excels in the overall performance of classification tasks and outperforms competing models in all segmentation task metrics.
Ahmed El-Azab, Ruiquan Ge, Xinchen Jiang, Gangyong Jia, Qing Wu 0008, Qinglei Shi, Changmiao Wang
BIBM6
2024 Diffusers Generated Unlabeled Images Improves Semi-supervised Object Detection
abstract
In the field of object detection, particularly in medical imaging, the scarcity of data often poses a significant challenge to model performance. To address this issue, this study proposes a semi-supervised learning approach based on a generative model. We begin by fine-tuning a pre-trained generative model using our dataset to better adapt the generative model to our specific data distribution. The fine-tuned generative model is then used to generate additional unlabeled data. These generated unlabeled data, combined with the original dataset, are employed in a semi-supervised training process. Experimental results demonstrate that our method significantly enhances the performance of the object detection model, especially in scenarios with limited labeled data, such as medical imaging. By incorporating the generated unlabeled training data into the semi-supervised framework, we observed a notable improvement in model accuracy. Specifically, our experiments showed an increase of up to $6.92 \%$ after adding the generated iamges. Moreover, it is foreseeable that incorporating a higher proportion of generated unlabeled data could lead to even more significant improvements in performance.
Zhanyun Lu, Renshu Gu, Huimin Cheng, Peifang Xu, Yuichiro Kinoshita, Juan Ye, Gangyong Jia, Qing Wu 0008
CW8
2024 Heterogeneous Graph Modeling for Resource-Aware Prediction of DRL Training Time
Gangyong Jia, Yuxia Cheng, Qing Wu 0008
ICA3PP (5)4
2024 PGKD-Net: Prior-guided and Knowledge Diffusive Network for Choroid Segmentation
abstract
The thickness of the choroid is considered to be an important indicator of clinical diagnosis. Therefore, accurate choroid segmentation in retinal OCT images is crucial for monitoring various ophthalmic diseases. However, this is still challenging due to the blurry boundaries and interference from other lesions. To address these issues, we propose a novel prior-guided and knowledge diffusive network (PGKD-Net) to fully utilize retinal structural information to highlight choroidal region features and boost segmentation performance. Specifically, it is composed of two parts: a Prior-mask Guided Network (PG-Net) for coarse segmentation and a Knowledge Diffusive Network (KD-Net) for fine segmentation. In addition, we design two novel feature enhancement modules, Multi-Scale Context Aggregation (MSCA) and Multi-Level Feature Fusion (MLFF). The MSCA module captures the long-distance dependencies between features from different receptive fields and improves the model's ability to learn global context. The MLFF module integrates the cascaded context knowledge learned from PG-Net to benefit fine-level segmentation. Comprehensive experiments are conducted to evaluate the performance of the proposed PGKD-Net. Experimental results show that our proposed method achieves superior segmentation accuracy over other state-of-the-art methods. Our code is made up publicly available at: https://github.com/yzh-hdu/choroid-segmentation.
Yaqi Wang 0002, Zehua Yang, Xindi Liu, Dechao Chen, Gangyong Jia, Juan Ye, Xingru Huang
Artif. Intell. Medicine9
2024 FlexibleCP: A data augmentation strategy for traffic sign detection
abstract
Abstract In the field of traffic sign detection, effective data augmentation can improve the model's detection capacity, enabling the model to distinguish and locate traffic signs more precisely and enhancing driving safety. However, due to the small size and low representation of traffic signs in the dataset, standard common data augmentation techniques are not suitable for traffic sign detection. To address this issue, a novel data augmentation strategy called flexible cut and paste (FlexibleCP) is proposed. The overall enhancement approach is shifted from multi‐image fusion to target cropping and pasting. By introducing parameters to control the target pasting ratio and scaling ratio, the diversity of small target data and their size variations are enriched. Additionally, target size and type filters are added to enable targeted enhancement for different sizes and types of targets. This study, evaluates the proposed strategy using two representative traffic sign detection datasets, namely CTSD and GTSDB. The experimental results demonstrate a significant improvement in both detection and recognition performance of the model: on the CTSD dataset, the models trained with FlexibleCP data enhancement achieve 88.9% and 64.5% mAP0.5 and mAP0.5:0.95, respectively, which are 3.5% and 2.5% better than those trained with mosaic data enhancement; on the GTSDB dataset mAP0.5 and mAP0.5:0.95 reached 89.2% and 56.0%, respectively, an improvement of 4.0% and 3.9% over mosaic.
Huanle Rao, Qinyang Jing, Ziqiang Wen, Gangyong Jia
IET Image Process.5
2024 Mmy-net: a multimodal network exploiting image and patient metadata for simultaneous segmentation and diagnosis
Renshu Gu, Yueyu Zhang, Lisha Wang, Dechao Chen, Yaqi Wang 0002, Ruiquan Ge, Zicheng Jiao, Juan Ye, Gangyong Jia, Linyan Wang
Multim. Syst.9
2024 HybAVPnet: A Novel Hybrid Network Architecture for Antiviral Peptides Prediction
abstract
Viruses pose a great threat to human production and life, thus the research and development of antiviral drugs is urgently needed. Antiviral peptides play an important role in drug design and development. Compared with the time-consuming and laborious wet chemical experiment methods, it is critical to use computational methods to predict antiviral peptides accurately and rapidly. However, due to limited data, accurate prediction of antiviral peptides is still challenging and extracting effective feature representations from sequences is crucial for creating accurate models. This study introduces a novel two-step approach, named HybAVPnet, to predict antiviral peptides with a hybrid network architecture based on neural networks and traditional machine learning methods. We adopted a stacking-like structure to capture both the long-term dependencies and local evolution information to achieve a comprehensive and diverse prediction using the predicted labels and probabilities. Using an ensemble technique with the different kinds of features can reduce the variance without increasing the bias. The experimental result shows HybAVPnet can achieve better and more robust performance compared with the state-of-the-art methods, which makes it useful for the research and development of antiviral drugs. Meanwhile, it can also be extended to other peptide recognition problems because of its generalization ability.
Ruiquan Ge, Yixiao Xia, Minchao Jiang, Gangyong Jia, Xiaoyang Jing, Ye Li 0002, Yunpeng Cai
IEEE ACM Trans. Comput. Biol. Bioinform.4
2024 A Self-Supervised Learning Based Framework for Eyelid Malignant Melanoma Diagnosis in Whole Slide Images
abstract
Eyelid malignant melanoma (MM) is a rare disease with high mortality. Accurate diagnosis of such disease is important but challenging. In clinical practice, the diagnosis of MM is currently performed manually by pathologists, which is subjective and biased. Since the heavy manual annotation workload, most pathological whole slide image (WSI) datasets are only partially labeled (without region annotations), which cannot be directly used in supervised deep learning. For these reasons, it is of great practical significance to design a laborsaving and high data utilization diagnosis method. In this paper, a self-supervised learning (SSL) based framework for automatically detecting eyelid MM is proposed. The framework consists of a self-supervised model for detecting MM areas at the patch-level and a second model for classifying lesion types at the slide level. A squeeze-excitation (SE) attention structure and a feature-projection (FP) structure are integrated to boost learning on details of pathological images and improve model performance. In addition, this framework also provides visual heatmaps with high quality and reliability to highlight the likely areas of the lesion to assist the evaluation and diagnosis of the eyelid MM. Extensive experimental results on different datasets show that our proposed method outperforms other state-of-the-art SSL and fully supervised methods at both patch and slide levels when only a subset of WSIs are annotated. It should be noted that our method is even comparable to supervised methods when all WSIs are fully annotated. To the best of our knowledge, our work is the first SSL method for automatic diagnosis of MM at the eyelid and has a great potential impact on reducing the workload of human annotations in clinical practice.
Zijing Jiang, Linyan Wang, Yaqi Wang 0002, Gangyong Jia, Guodong Zeng, Jun Wang 0072, Dechao Chen, Guiping Qian, Qun Jin
IEEE Trans. Comput. Biol. Bioinform.4
2024 Heter-Train: A Distributed Training Framework Based on Semi-Asynchronous Parallel Mechanism for Heterogeneous Intelligent Transportation Systems
abstract
Transportation big data (TBD) are increasingly combined with artificial intelligence to mine novel patterns and information due to the powerful representational capabilities of deep neural networks (DNNs), especially for anti-COVID19 applications. The distributed cloud-edge-vehicle training architecture has been applied to accelerate DNNs training while ensuring low latency and high privacy for TBD processing. However, multiple intelligent devices (e.g., intelligent vehicles, edge computing chips at base stations) and different networks in intelligent transportation systems lead to computing power and communication heterogeneity among distributed nodes. Existing parallel training mechanisms perform poorly on heterogeneous cloud-edge-vehicle clusters. The synchronous parallel mechanism may force fast workers to wait for the slowest worker for synchronization, thus wasting their computing power. The asynchronous mechanism has communication bottlenecks and can exacerbate the straggler problem, causing increased training iterations and even incorrect convergence. In this paper, we introduce a distributed training framework, Heter-Train. First, a communication-efficient semi-asynchronous parallel mechanism (SAP-SGD) is proposed, which can take full advantage of acceleration effect of asynchronous strategy on heterogeneous training and constrain the straggler problem by using global interval synchronization. Second, Considering the difference in node bandwidth, we design a solution for heterogeneous communication. Moreover, a novel weighted aggregation strategy is proposed to aggregate the model parameters with different versions. Finally, experimental results show that our proposed strategy can achieve up to$6.74 \times $speedups on training time, with almost no accuracy decrease.
Jiawei Geng, Haipeng Jia, Zongwei Zhu, Hai Fang, Chengxi Gao, Cheng Ji 0002, Gangyong Jia, Guangjie Han, Xuehai Zhou
IEEE Trans. Intell. Transp. Syst.8
2024 An adaptive service deployment algorithm for cloud-edge collaborative system based on speedup weights
Zhichao Hu, Huanle Rao, Chenjie Hong, Ouhan Huang, Gangyong Jia
J. Supercomput.7
2024 Aphto: a task offloading strategy for autonomous driving under mobile edge
JiaCheng Lin, Huanle Rao, SongSong Liang, Yumiao Zhao, Qing Ren, Gangyong Jia
J. Supercomput.6
2024 A container optimal matching deployment algorithm based on CN-Graph for mobile edge computing
Huanle Rao, Gangyong Jia
J. Supercomput.6
2024 RCFS: rate and cost fair CPU scheduling strategy in edge nodes
Yumiao Zhao, Huanle Rao, Kelei Le, Youqing Xu, Gangyong Jia
J. Supercomput.6
2024 Boosting Research for Carbon Neutral on Edge UWB Nodes Integration Communication Localization Technology of IoV
Ouhan Huang, Huanle Rao, Renshu Gu, Hong Xu 0014, Gangyong Jia
IEEE Trans. Sustain. Comput.6
2023 GCS-ICHNet: Assessment of Intracerebral Hemorrhage Prognosis using Self-Attention with Domain Knowledge Integration
abstract
Intracerebral Hemorrhage (ICH) is a severe condition resulting from damaged brain blood vessel ruptures, often leading to complications and fatalities. Timely and accurate prognosis and management are essential due to its high mortality rate. However, conventional methods heavily rely on subjective clinician expertise, which can lead to inaccurate diagnoses and delays in treatment. Artificial intelligence (AI) models have been explored to assist clinicians, but many prior studies focused on model modification without considering domain knowledge. This paper introduces a novel deep learning algorithm, GCS-ICHNet, which integrates multimodal brain CT image data and the Glasgow Coma Scale (GCS) score to improve ICH prognosis. The algorithm utilizes a transformer-based fusion module for assessment. GCS-ICHNet demonstrates high sensitivity 81.03% and specificity 91.59%, outperforming average clinicians and other state-of-the-art methods. The code is available at https://github.com/Windbelll/Prognosis-analysis-of-cerebral-hemorrhage.
Xuhao Shan, Ruiquan Ge, Shibin Wu, Ahmed El-Azab, Jichao Zhu, Gangyong Jia, Qingying Xiao, Changmiao Wang
BIBM8
2023 UWAT-GAN: Fundus Fluorescein Angiography Synthesis via Ultra-Wide-Angle Transformation Multi-scale GAN
Zhaojie Fang, Zhanghao Chen, Pengxue Wei, Wangting Li, Shaochong Zhang, Ahmed El-Azab, Gangyong Jia, Ruiquan Ge, Changmiao Wang
MICCAI (7)7
2023 VTP: volumetric transformer for multi-view multi-person 3D pose estimation
Renshu Gu, Ouhan Huang, Gangyong Jia
Appl. Intell.4
2023 GOMPS: Global Attention-based Ophthalmic Image Measurement and Postoperative Appearance Prediction System
abstract
Accurate measurements of ophthalmic parameters and postoperative appearance prediction are essential for the diagnosis and treatment of many ophthalmic diseases. Nevertheless, it remains challenging due to (1) inconsistent ophthalmic image sampling standards, including ocular-camera distance, facial angle, and patient number, (2) complicated ocular morphology, such as subconjunctival hemorrhage, ocular movements, lighting effects, and morphological aging. It is difficult for a model to measure parameters and make predictions in variable sampling methods and morphology conditions. Therefore, the Global attention-based Ophthalmic Image Measurement and Postoperative Appearance Prediction System (GOMPS) is proposed, which quantifies ophthalmic image parameters to diagnose disease and simultaneously predict postoperative appearance of blepharoptosis. By perceiving the global structure of the ophthalmic image, GOMPS makes logical inference predictions of the sclera and cornea morphology, to overcome the above difficulties. Concretely, a global attention unit (GAU) and a novel global attention structure-aware network (GASA-Net) are designed to enhance GOMPS’s global structure awareness ability to perform logical reasoning. Extensive experimental results on our collected ophthalmic dataset for diagnosis & prediction (OD2P) demonstrate that GOMPS surpasses the state-of-the-art methods in segmentation accuracy and achieves the current optimal performance in measurement and postoperative prediction under many clinical scenes.
Xingru Huang, Lixia Lou, Ruilong Dan, Lingxiao Chen, Guodong Zeng, Gangyong Jia, Qun Jin, Juan Ye, Yaqi Wang 0002
Expert Syst. Appl.7
2023 CDNet: Contrastive Disentangled Network for Fine-Grained Image Categorization of Ocular B-Scan Ultrasound
abstract
Precise and rapid categorization of images in the B-scan ultrasound modality is vital for diagnosing ocular diseases. Nevertheless, distinguishing various diseases in ultrasound still challenges experienced ophthalmologists. Thus a novel contrastive disentangled network (CDNet) is developed in this work, aiming to tackle the fine-grained image categorization (FGIC) challenges of ocular abnormalities in ultrasound images, including intraocular tumor (IOT), retinal detachment (RD), posterior scleral staphyloma (PSS), and vitreous hemorrhage (VH). Three essential components of CDNet are the weakly-supervised lesion localization module (WSLL), contrastive multi-zoom (CMZ) strategy, and hyperspherical contrastive disentangled loss (HCD-Loss), respectively. These components facilitate feature disentanglement for fine-grained recognition in both the input and output aspects. The proposed CDNet is validated on our ZJU Ocular Ultrasound Dataset (ZJUOUSD), consisting of 5213 samples. Furthermore, the generalization ability of CDNet is validated on two public and widely-used chest X-ray FGIC benchmarks. Quantitative and qualitative results demonstrate the efficacy of our proposed CDNet, which achieves state-of-the-art performance in the FGIC task.
Ruilong Dan, Gangyong Jia, Shuai Wang 0003, Ruiquan Ge, Guiping Qian, Qun Jin, Juan Ye, Yaqi Wang 0002
IEEE J. Biomed. Health Informatics5
2022 Juggler-ResNet: A Flexible and High-Speed ResNet Optimization Method for Intrusion Detection System in Software-Defined Industrial Networks
abstract
ResNetsare widely used in the intrusion detection system (IDS) of software-defined industrial network to construct accurate intelligence detection of network attacks. However, the IDS based on ResNets has a long detecting interval because of the fine-grained operator and intermediate outcomes of the multi-branch architecture of ResNets. To address this problem, in this article, we propose Juggler-ResNet with a fusible residual structure that preserves the feature extraction ability of the residual structure and enables equivalent transformation to linear topology to support low latency inference service in the industrial application (e.g., malicious network behavior detection, fault diagnosis, etc.). First, we propose a fusible multibranch residual structure to avoid gradient vanishing problems in the training phase. Second, we convert it to linear-topology by using a set of equivalent fusion operators. Finally, the linear-topology model is deployed to accelerate inference speed. Our experimental results on CIFAR-10 and CIFAR-100 show that fusible residual structure can achieve 2.08-4.3x acceleration with state-of-the-art level accuracy performance.
Zongwei Zhu, Wenjie Zhai, Huanghe Liu, Jiawei Geng, Mingliang Zhou 0001, Cheng Ji 0002, Gangyong Jia
IEEE Trans. Ind. Informatics7
2022 A Secure Dynamic Mix Zone Pseudonym Changing Scheme Based on Traffic Context Prediction
abstract
Traffic context plays an important role in supporting automated driving and intelligent transportation systems. Smart vehicles explore surrounding environments by analyzing sensor data and periodically communicating with neighbors and road infrastructures. The context can be well learned in this way to support driving, but the vehicle trajectory can be also easily exposed under eavesdropping attacks. The pseudonym is proposed to hide the real identity of the vehicles. However, the effectiveness of anonymity, the safety of driving, the convenience of implementation and the utilization of resources in previous approaches have not been well-balanced. Therefore, focusing on efficiently replacing pseudonyms with the premise of ensuring driving safety, we propose a secure dynamic silent mix zone pseudonym changing scheme (TLAS) based on the real-time traffic context prediction for urban regions. It naturally takes the area in front of the red traffic light as a silent mix zone, which avoids the driving security issue caused by signal silence. Besides, the area length is dynamically configured according to the traffic context predicted in the last green light cycle, so the anonymous effect can be improved. In addition, considering the resource utilization and accuracy requirement, the adaptive prediction algorithm is applied. We conduct simulation experiments with real-world traffic history using SUMO and OMNET++, the results show that TLAS strategy can indeed achieve a better anonymous effect (reducing standardized traceability rate by 8.2%) with lower driving speed for safety concern.
Youhuizi Li, Yuyu Yin, Xu Chen 0048, Jian Wan 0001, Gangyong Jia, Kewei Sha
IEEE Trans. Intell. Transp. Syst.5
2021 Touch Point Prediction for Interactive Public Displays Based on Camera Images
abstract
Feedback latency during the use of interactive displays is an issue currently being considered in the HCI field. Several studies have focused on reducing latency using various approaches. This paper proposes a framework that uses a convolutional neural network to predict user touch points for interactive public displays. The framework predicts user touch events before the finger reaches the display surface to reduce the latency in feedback. As a training dataset, 1,651 tapping actions were collected from 18 participants in front of a display. The training of the convolutional neural network architecture was performed using the collected tapping actions. Validation test results showed that reasonable accuracy could be achieved at 390 ms before touching the display.
Yuichiro Kinoshita, Kentaro Go, Gangyong Jia
CW4
2021 SuccSPred: Succinylation Sites Prediction Using Fused Feature Representation and Ranking Method
Ruiquan Ge, Yizhang Luo, Guanwen Feng, Gangyong Jia, Gang Xu 0001
ISBRA4
2021 Edge Network Routing Protocol Base on Target Tracking Scenario
abstract
Abstract Edge computing perfectly integrates cloud computing centers and edge-end devices together, but there are not many related researches on how the edge-end node devices work to form an edge network and what the protocols used to implement the communication among nodes in the edge network. Aiming at the problem of coordinated communication among edge nodes in the current edge computing network architecture, this paper proposes an edge network routing and forwarding protocol based on target tracking scenarios. This protocol can meet the dynamic changes of node locations, and the elastic expansion of node scale. Individual node failures will not affect the overall network, and the network ensures efficient real-time with less communication overhead. The experimental results display that the protocol can effectively reduce the communications volume of the edge network, improve the overall efficiency of the network, and set the optimal sampling period, so as to ensure that the network delay is minimized.
Weihua Zhao, Ouhan Huang, Gangyong Jia, Youhuizi Li, Songzhu Mei, Duan Zhao
Mob. Networks Appl.4
2020 Modified DenseNet for Automatic Fabric Defect Detection With Edge Computing for Minimizing Latency
abstract
As an essential step in quality control, fabric defect detection plays an important role in the textile manufacturing industry. The traditional manual detection method is inaccurate and incurs a high cost; as a result, it is gradually being replaced by deep learning algorithms based on cloud computing. However, a high data transmission latency between end devices and the cloud has a significant impact on textile production efficiency. In contrast, edge computing, which provides services near end devices by deploying network, computing and storage facilities at the edge of the Internet, can effectively solve the above-mentioned problem. In this article, we propose a deep-learning-based fabric defect detection method for edge computing scenarios. First, this article modifies the structure of DenseNet to better suit a resource-constrained edge computing scenario. To better assess the proposed model, an optimized cross-entropy loss function is also formulated. Afterward, six feasible expansion schemes are utilized to enhance the data set according to the characteristics of various defects in fabric samples. To balance the distribution of samples, proportions of various defect types are used to determine the number of enhancements. Finally, a fabric defect detection system is established to test the performance of the optimized model used on edge devices in a real-world textile industry scenario. Experimental results demonstrate that compared with the conventional convolutional neural network (CNN), the proposed optimized model attains an average improvement of 18% in the area under the curve (AUC) metric for 11 defects. Data transmission is reduced by approximately 50% and latency is reduced by 32% in the Cambricon 1H8 platform compared with a cloud platform.
Zongwei Zhu, Guangjie Han, Gangyong Jia, Lei Shu 0001
IEEE Internet Things J.3
2020 DPAM: A Demand-Based Page-Level Address Mappings Algorithm in Flash Memory for Smart Industrial Edge Devices
abstract
Edge computing brings data storage closer to the location where it is needed. Therefore, the edge devices, especially smart industrial edge devices, require higher storage systems. NAND flash memory has the advantages of small size, high speed, and strong shock resistance, which is widely used in various storage systems, providing a good choice for edge devices. NAND flash has unique physical characteristics, such as “out-of-place updates” and “prewrite erasure,” therefore, the traditional address mapping methods require improvement. This article presents a novel demand-based page-level address mapping algorithm called DPAM. The goal of DPAM is to provide efficient address translation by using a smaller address mapping table. Due to the high service cost of block-level address mapping and hybrid address mapping, a page-level address mapping scheme is proposed. The algorithm is implemented and tested on the flash simulation platform FlashSim. The results indicate that our algorithm provides improvements of 7.11% for the hit ratio and 7% for the number of block erasures compared with other approaches.
Gangyong Jia, Guangjie Han, Jinfang Jiang, Li Liu 0022, Lei Shu 0001
IEEE Trans. Ind. Informatics1
2019 A Collaborative Anomaly Detection Approach of Marine Vessel Trajectory (Short Paper)
Jian Wan 0001, Jie Huang 0014, Gangyong Jia, Wei Zhang 0138
CollaborateCom4
2019 An NB-IoT-based smart trash can system for improved health in smart cities
abstract
The intelligent treatment of urban garbage is an important component of creating a smart city and also solves several problems associated with urban garbage. Many traditional garbage cans are widely distributed, resulting in a waste of human and material resources, untimely government. Therefore, in this paper, we propose an intelligent system based on edge computing and the narrow-band Internet of things (NB-IoT) for monitoring smart trash cans (STCs). The deployed intelligent garbage cans are distributed throughout the city and are equipped with a variety of sensors, including a compression sensor, a location sensor, an infrared sensor, and an alarm sensor. The data sent from the smart bins are preprocessed through edge nodes for data classification and priority transmission, which reduces the required network transmission bandwidth and the computational tasks at the centralized data center. The NB-IoT is a narrow-band communication technology with low power consumption, wide coverage, low cost, and large capacity. The experimental results show that the proposed STC system shows good system performance, and allows for intelligent management of garbage in smart cities.
Gangyong Jia, Guangjie Han, Zeren Zhou, Mohsen Guizani
IWCMC2
2019 A Maximum Cache Value Policy in Hybrid Memory-Based Edge Computing for Mobile Devices
abstract
Edge computing is proposed to bridge mobile devices with cloud computing data centers in the era of mobile big data, as an intermediate level of computing power. One important issue in edge computing is how to improve performance for mobile devices. Current systems utilize cache in multicore systems to reduce memory access cost with an acceptable hardware cost. However, existing cache management policies are unable to maximize cache value in the newly developed hybrid memory platform that combines phase-change memory and dynamic random-access memory. In this paper, we propose maximizes cache value (MCV), an efficient cache management policy, which MCV to minimize memory access cost in a hybrid main memory platform for edge computing. Extensive simulation studies indicate that this strategy can improve performance in hybrid main memory-based edge computing for mobile devices.
Gangyong Jia, Guangjie Han, Sammy Chan
IEEE Internet Things J.1
2019 Hybrid-LRU Caching for Optimizing Data Storage and Retrieval in Edge Computing-Based Wearable Sensors
abstract
In the era of the Internet of Things, edge computing-based wearable sensors are rapidly emerging for smart health. The collection, storage, and retrieval of data are the key components of wearable sensors. Therefore, it is important to optimize data storage and retrieval. Phase change memory (PRAM) is a kind of phase change memory that is widely used as a new storage medium. It has the characteristics of nonvolatility, high-density storage. However, it has the disadvantages of asymmetry in reading and writing and limited life. In recent years, PRAM and DRAM were combined into PDRAM as a hybrid memory architecture, to solve the problems caused by PRAM. This paper proposes a new cache policy named hybrid-LRU to adapt PDRAM. Hybrid-LRU uses two different LRU cache policies to distinguish PRAM and DRAM as two different storage mediums. The experimental results show that the hybrid-LRU cache policy improves the performance by 4.2%, and reduces the utilization rate of PRAM in PDRAM by 11.8%. In addition, the energy consumption of writing and reading can be reduced to 87.8%.
Gangyong Jia, Guangjie Han, Hongtianchen Xie
IEEE Internet Things J.1
2019 Coordinate Memory Deduplication and Partition for Improving Performance in Cloud Computing
abstract
Both limited main memory size and memory interference are considered as the major bottlenecks in virtualization environments. Memory deduplication, detecting pages with same content and being shared into one single copy, reduces memory requirements; memory partition, allocating unique colors for each virtual machine according to page color, reduces memory interference among virtual machines to improve performance. In this paper, we propose a coordinate memory deduplication and partition approach named CMDP to reduce memory requirement and interference simultaneously for improving performance in virtualization. Moreover, CMDP adopts a lightweight page behavior-based memory deduplication approach named BMD to reduce futile page comparison overhead meanwhile to detect page sharing opportunities efficiently. And a virtual machine based memory partition called VMMP is added into CMDP to reduce interference among virtual machines. According to page color, VMMP allocates unique page colors to applications, virtual machines and hypervisor. The experimental results show that CMDP can efficiently improve performance (by about 15.8 percent) meanwhile accommodate more virtual machines concurrently.
Gangyong Jia, Guangjie Han, Joel J. P. C. Rodrigues, Jaime Lloret Mauri, Wei Li 0064
IEEE Trans. Cloud Comput.1
2019 Special Section on Emerging Trends Issues and Challenges in Edge Artificial Intelligence
abstract
The papers in this special section focus on edge computing and the challenges that exist for artificial intelligence applications. Edge computing has the advantages of real-time response and less network demand for computing closer to the edge of the network, while bridging the physical and digital worlds. The core of the edge computing is to provide the edge intelligent service. Therefore, edge artificial intelligence is becoming a popular trend for the future, such as intelligent sound box, and so on. Edge artificial intelligence combines edge computing with artificial intelligence, while taking both advantages. However, there are some problems, which need to be solved for the edge artificial intelligence. Addresses these issues and examines future areas of development in this area.
Guangjie Han, Mohsen Guizani, Gangyong Jia, Jaime Lloret Mauri
IEEE Trans. Ind. Informatics3
2019 Effective shortest travel-time path caching and estimating for location-based services
Detian Zhang, An Liu 0002, Zhixu Li, Gangyong Jia, Fei Chen 0010, Qing Li 0001
World Wide Web4
2018 Edge Computing-Based Intelligent Manhole Cover Management System for Smart Cities
abstract
An intelligent manhole cover management system (IMCS) is one of the most important basic platforms in a smart city to prevent frequent manhole cover accidents. Manhole cover displacement, loss, and damage pose threats to personal safety, which is contrary to the aim of smart cities. This paper proposes an edge computing-based IMCS for smart cities. A unique radio frequency identification tag with tilt and vibration sensors is used for each manhole cover, and a Narrowband Internet of Things is adopted for communication. Meanwhile, edge computing servers interact with corresponding management personnel through mobile devices based on the collected information. A demonstration application of the proposed IMCS in the Xiasha District of Hangzhou, China, showed its high efficiency. It efficiently reduced the average repair time, which could improve the security for both people and manhole covers.
Gangyong Jia, Guangjie Han, Huanle Rao, Lei Shu 0001
IEEE Internet Things J.1
2018 Resource-utilization-aware energy efficient server consolidation algorithm for green computing in IIOT
Guangjie Han, Wenhui Que, Gangyong Jia, Wenbo Zhang 0001
J. Netw. Comput. Appl.3
2018 Dynamic cloud resource management for efficient media applications in mobile computing environments
Gangyong Jia, Guangjie Han, Jinfang Jiang, Sammy Chan
Pers. Ubiquitous Comput.1
2018 SSL: Smart Street Lamp Based on Fog Computing for Smarter Cities
abstract
Both safety and energy conservation are very important advantages of smart cities. Namely, the city street lamp is correlated with both safety and energy conservation. Therefore, a street lamp is an indispensable part of the smart cities. However, current street lamps have lack of smart characteristics, which increases both danger and energy consumption. In order to address these problems, a smart street lamp (SSL) based on the fog computing for smarter cities is proposed in this paper. The advantages of the proposed SSL are as follows: 1) fine management, because every street lamp can be operated independently; 2) dynamic brightness adjustment, all street lamps can be adjusted dynamically; and 3) autonomous alarm on abnormal states, each street lamp can report the abnormal status independently, such as broken, stolen, and so on. The experimental results showed that the proposed SSL can improve the energy efficiency and reduce danger.
Gangyong Jia, Guangjie Han, Aohan Li
IEEE Trans. Ind. Informatics1
2017 Effective Caching of Shortest Travel-Time Paths for Web Mapping Mashup Systems
Detian Zhang, An Liu 0002, Gangyong Jia, Fei Chen 0010, Qing Li 0001
WISE (1)3
2017 Dynamic Adaptive Replacement Policy in Shared Last-Level Cache of DRAM/PCM Hybrid Memory for Big Data Storage
abstract
The increasing demand on the main memory capacity is one of the main big data challenges. Dynamic random access memory (DRAM) does not represent the best choice for a main memory, due to high power consumption and low density. However, the nonvolatile memory, such as the phase-change memory (PCM), represents an additional choice because of the low power consumption and high-density characteristic. Nevertheless, the high access latency and limited write endurance have disabled the PCM to replace the DRAM currently. Therefore, a hybrid memory, which combines both the DRAM and the PCM, has become a good alternative to the traditional DRAM memory. Both DRAM and PCM disadvantages are challenges for the hybrid memory. In this paper, a dynamic adaptive replacement policy (DARP) in the shared last-level cache for the DRAM/PCM hybrid main memory is proposed. The DARP distinguishes the cache data into the PCM data and the DRAM data, then, the algorithm adopts different replacement policies for each data type. Specifically, for the PCM data, the least recently used (LRU) replacement policy is adopted, and for the DRAM data, the DARP is employed according to the process behavior. Experimental results have shown that the DARP improved the memory access efficiency by 25.4%.
Gangyong Jia, Guangjie Han, Jinfang Jiang, Li Liu 0022
IEEE Trans. Ind. Informatics1
2016 Virtual Page Behavior Based Page Management Policy for Hybrid Main Memory in Cloud Computing
abstract
A new generation memory, Non-Volatile Memory (NVM), such as Phase-Change Memory (PCM), has been adopted together with DRAM in the main memory to form the hybrid main memory for low energy consumption and high capacity. The biggest challenge of hybrid memory is how to decrease the average memory access cost for the higher cost of NVM's read/write operation. Currently, most researches are based on migration. However, the page migration itself is a high cost operation. And the migration based policy produces many migration operations, which induces high cost in memory access. Therefore, in order to decrease the cost, we present a virtual page behavior based page management policy (VBPM) in this paper. According to the virtual pages' behavior, we allocate virtual pages into DRAM or PCM physical pages correspondingly. The whole process is migration independent. The experimental results show our VBPM decreases the average memory access time by 24%, moreover, VBPM improves real-time performance in critical path.
Jie Huang 0014, Guangjie Han, Gangyong Jia, Huizi Liyou, Jian Wan 0001
MSN4
2015 PARS: A scheduling of periodically active rank to optimize power efficiency for main memory
Gangyong Jia, Guangjie Han, Jinfang Jiang, Joel J. P. C. Rodrigues
J. Netw. Comput. Appl.1
2015 Dynamic Time-slice Scaling for Addressing OS Problems Incurred by Main Memory DVFS in Intelligent System
Gangyong Jia, Guangjie Han, Jinfang Jiang, Aohan Li
Mob. Networks Appl.1
2014 Temperature-Aware Scheduling Based on Dynamic Time-Slice Scaling
Gangyong Jia, Youwei Yuan, Jian Wan 0001, Congfeng Jiang, Xi Li 0003, Dong Dai 0001
ICA3PP (1)1
2014 Combine thread with memory scheduling for maximizing performance in multi-core systems
abstract
The growing gap between microprocessor speed and DRAM speed is a major problem that computer designers are facing. In order to narrow the gap, it is necessary to improve DRAM's speed and throughput. Moreover, on multi-core platforms, DRAM memory shared by all cores usually suffers from the memory contention and interference problem, which can cause serious performance degradation and unfairness among parallel running threads. To address these problems, this paper proposes techniques to take both advantages of partitioning cores, threads and memory banks into groups to reduce interference among different groups and grouping the memory accesses of the same row together to reduce cache miss rate. A memory optimization framework combined thread scheduling with memory scheduling (CTMS) is proposed in this paper, which simultaneously minimizes memory access schedule length, memory access time and reduce interference to maximize performance for multi-core systems. Experimental results show CTMS is 12.6% shorter in memory access time, while improving 11.8% throughput on average. Moreover, CTMS also saves 5.8% of the energy consumption.
Gangyong Jia, Guangjie Han, Liang Shi 0001, Jian Wan 0001, Dong Dai 0001
ICPADS1
2014 PUMA: Pseudo unified memory architecture for single-ISA heterogeneous multi-core systems
abstract
Single-ISA heterogeneous multi-core processors have advantages over cost-equivalent homogeneous ones, which integrate cores having the same instruction set architecture (ISA) but offer different performance and power characteristics. When these cores share the off-chip main memory, requests from different cores will interfere with each other, leading to low system performance and unfairness even starvation. Unfortunately, state-of-the-art memory scheduling and thread scheduling algorithms are ineffective at solving these problems. This paper proposes a fundamentally new memory architecture of pseudo unified memory (PUMA), which partitions the memory into regions according cores' different performance, each core mostly requests only one memory region seldom exceeding, reducing interfere among cores while retaining bank level parallelism for improving performance and fairness. We evaluate the design trade-offs involved in our PUMA and compare it against three state-of-the-art memory management methods. Our experimental results show that PUMA improves both system performance and fairness among cores while reducing memory power.
Gangyong Jia, Liang Shi 0001, Jian Wan 0001, Youwei Yuan, Xi Li 0003, Dong Dai 0001
RTCSA1
2013 Coordinate Task and Memory Management for Improving Power Efficiency
Gangyong Jia, Xi Li 0003, Jian Wan 0001, Chao Wang 0003, Dong Dai 0001, Congfeng Jiang
ICA3PP (1)1
2013 Power-aware buddy system and task group scheduler
abstract
Memory is responsible for a large and increasing fraction of the energy consumed by computers. To address this challenge, memory manufacturers have developed memory devices with different power states. In order to more effectively manage the power states in the operating system, in this paper, we propose a rank-sensitive buddy system (RS-Buddy) which clusters pages together to prolong the idle time of memory ranks without breaking defragmentation characteristics. For the purpose of decreasing unnecessary frequent mode transitions, we introduce a power-aware task group scheduler (PATGS) that groups the threads which access the same rank together to schedule while sustaining system fairness. Finally, we integrate state-of-the-art mode control policies with our RS-Buddy and PATGS, with experimental results demonstrating that our algorithms can improve the power efficiency from 25.31% to 27.35% compared with state-of-the-art studies.
Xi Li 0003, Zongwei Zhu, Gangyong Jia, Xuehai Zhou
ISCAS3
2012 Cache Promotion Policy Using Re-reference Interval Prediction
abstract
The last-level cache (LLC) mitigates the long latencies of memory access in today's chip multi-core processor (CMP). The promotion policy in the LLC largely affects cache efficiency, while an inappropriate promotion policy may lead useless blocks to remain in the cache longer than necessary, in turn result into inefficiency. Currently state-of-the-art promotion policies are unaware of the re-reference interval of cache accesses. Applications that exhibit a long re-reference interval perform poorly with these promotion policies. In this paper, we propose a promotion policy that uses re-reference interval prediction (RRIP) information. Such technique requires minor hardware modification over the least-recently-used (LRU) replacement policy. Our evaluation shows that RRIP improves IPCsumby 2.58%, Weighted Speedup by 3.54% and IPCnorm_hmeanby 6.2% on average over single-step promotion policy.
Gangyong Jia, Xi Li 0003, Chao Wang 0003, Xuehai Zhou, Zongwei Zhu
CLUSTER1
2012 Memory Affinity: Balancing Performance, Power, Thermal and Fairness for Multi-core Systems
abstract
Main memory is expected to grow significantly in both speed and capacity for it is a major shared resource among cores in a multi-core system, which will lead to increasing power consumption. Therefore, it is critical to address the power issue without seriously decreasing performance in the memory subsystem. In this paper, we firstly propose memory affinity which retains the active and low power memory ranks as long as possible to avoid frequently switching between active and low power status, and then present a memory affinity aware scheduling (MAS) to balance performance, power, thermal and fairness for multi-core systems. Experimental results demonstrate our memory affinity aware scheduling algorithms well adapt to system loading to maximize power saving and avoid memory hotspot at the same time while sustaining the system bandwidth demand and preserving fairness among threads.
Gangyong Jia, Xi Li 0003, Chao Wang 0003, Xuehai Zhou, Zongwei Zhu
CLUSTER1
2012 Phase Detection for Loop-Based Programs on Multicore Architectures
abstract
Phase detection and behavior analysis have been major concerned to improve the performance as well as the system throughputs. However, for the distributed acceleration engines, the execution among different phases is much more difficult to be analyzed, especially for the loop based programs. With respect to the tasks in different iterations, how to efficiently detect the phases belonging to the same loop iteration or even across iterations is posing significant challenge. In this paper we propose a phase detection method for loop-based programs on multiprocessor system-on-chip (MPSoC). A cross compiling tool based on state-of-the-art ARM RVDS is employed to locate the hot spot function of the program. Based on the hot spots, we target the function optimization on a hadoop cluster for performance evaluation. The preliminary experimental results demonstrate that our proposed techniques can extract the hot block function with high accuracy and modest overheads. The method can be applied to guide the optimization and adaptive mapping scheme on MPSoC architectures.
Chao Wang 0003, Xi Li 0003, Dong Dai 0001, Gangyong Jia, Xuehai Zhou
CLUSTER4
2012 Share memory aware scheduler: balancing performance and fairness
abstract
Optimizing system performance through scheduling has received a lot of attention. However, none of the existing approaches can balance the system performance improvement and the fair share of CPU time among threads. We present in this paper a share memory aware scheduler (SMAS). The key idea is to adopt thread group scheduling which partitions threads based on memory address space to reduce switching overhead and to give each thread a fair chance to occupy CPU time. There are three main contributions: 1) SMAS does well in balancing system performance and fairness among all threads; 2) to our knowledge, this is the first attempt to use share memory aware scheduler for system performance improvement; 3) we implement SMAS both in testbed and simulator for evaluation. The testbed results on a 2-core processor show that our proposed scheduler can improve performance of different performance parameters with neglected overhead in fairness, which reduced 0.128% in cache miss rate, 2.62% in run time, 13.15% in DTBL misses, 31.68% in ITLB misses and 46.15% in ITLB flushes maximum. Furthermore, our extensive simulation results for 4 and 8 cores demonstrate that SMAS is highly scalable.
Xi Li 0003, Gangyong Jia, Zongwei Zhu, Xuehai Zhou
ACM Great Lakes Symposium on VLSI2
2012 Behavior Aware Data Locality for Caches
abstract
Optimizing cache performance through improving data locality has been receiving a lot of attention. However, none of the existing approaches can combine each task's behavior to optimize data locality for caches. We present a behavior aware data locality (BADL) to optimize cache performance in this paper. The key idea is to add each task's behavior when allocating memory, which can take advantage of each task's different locality to optimize cache performance. There are five main contributions: 1. to our best knowledge, this is the first attempt to improve cache performance through combining task behavior, 2. BADL detailed analyzes low performance derived from internal of the cache line, which is more fine-grained than the current state-of-the-art fine-grained in hardware angle, 3. BADL optimizes the cache performance through improving internal of cache line efficiency, 4. we implement BADL both in single-threaded application and multi-threaded applications scenarios, 5. BADL can be combined to most of the cache optimizing researches. The experiment results show our proposed BADL can improve 18.6% performance on average in single-threaded application situation and improve 20.8% performance on average in multi-threaded application situation.
Gangyong Jia, Xi Li 0003, Chao Wang 0003, Xuehai Zhou, Zongwei Zhu
ICPADS1
2012 Frequency Affinity: Analyzing and Maximizing Power Efficiency in Multi-core Systems
abstract
Performance optimization and energy efficiency are the major challenges in multi-core system design. Of the state-of-the-art approaches, cache affinity aware scheduling and techniques based on dynamic voltage frequency scaling (DVFS) are widely applied to improve performance and save energy consumptions respectively. In modern operating systems, schedulers exploit high cache affinity by allocating a process on a recently used processor whenever possible. When a process runs on a high-affinity processor it will find most of its states already in the cache and will thus achieve more efficiency. However, most state-of-the-art DVFS techniques do not concentrate on the cost analysis for DVFS mechanism. In this paper, we firstly propose frequency affinity which retains the voltage frequency as long as possible to avoid frequently switching, and then present a frequency affinity aware scheduling (FAS) to maximize power efficiency for multi-core systems. Experimental results demonstrate our frequency affinity aware scheduling algorithms are much more power efficient than single-ISA heterogeneous multi-core processors.
Gangyong Jia, Xi Li 0003, Chao Wang 0003, Xuehai Zhou, Zongwei Zhu
MASCOTS1
2012 Analyzing Parallelization and Program Performance in Heterogeneous MPSoCs
abstract
In this paper we extend and analyze Amdahl's law to general heterogeneous MPSoC era, to find out how the speedup is affected by the parameters, including amount and speedup for microprocessors and accelerators, as well as the task partition characteristics. We also analyze the theoretical results about how the extended Amdahl's Law is applied to leverage load balancing of a heterogeneous MPSoC without the abstract limitation of base core equivalents (BCEs). A prototype on FPGA is constructed with Microblaze processors and JPEG hardware accelerators. The experimental results demonstrate that our extended model reinforces state-of-the-art performance evaluation methods for hybrid MPSoC architectures and also provide creditable new insights on the heterogeneous research communities, in particular for scalable FPGA based reconfigurable MPSoCs.
Chao Wang 0003, Xi Li 0003, Junneng Zhang, Gangyong Jia, Peng Chen 0004, Xuehai Zhou
MASCOTS4