Zhenjiang Dong

dblp:140/9501 · DBLP profile ↗
← Back
53ranked-venue papers
2as first author
37since 2021 · last 2027
0009-0001-1213-8686ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 15 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 7 since 2021Artificial intelligence and machine learning · 7 · 3 since 2021Systems, architecture and hardware · 7 · 1 first-author · 6 since 2021Security and privacy · 3 · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2027 A federated contrastive bifocal distillation approach for heterogeneous IIoT devices
Xiaoxuan Hu, Songhao Hu, Zhenjiang Dong, Jialin Hua, Yanfei Sun
Future Gener. Comput. Syst.4
2026 TrustTable: A Neuro-Symbolic Auditing Framework for Faithful Table QA
abstract
Large Language Models based Table Question Answering (LLMs-based TableQA) models excel in NLP field, however, they occasionally exhibit unfaithful behavior where correct answers are derived through erroneous reasoning paths.In this condition, we propose Trust-Table, a neuro-symbolic framework designed to ensure reasoning faithfulness by auditing the reasoning processes of LLMs.Unlike monolithic LLM-based auditors, TrustTable decouples the auditing operation into two orthogonal dimensions.It enforces factual grounding by executing neurally generated Pandas code against the table, and ensures logical soundness by verifying reasoning chains through a LLM-synthesized formal solver.By integrating these symbolic checks, TrustTable enables a "Label-Free Audit Loop" that systematically identifies and rectifies reasoning flaws without human supervision.In addition, we present the TrustTable-Bench, a diagnostic dataset containing diverse error categories that range from calculation discrepancies to schema misalignments.This benchmark allows for a rigorous quantification of reasoning limitations.Experiments demonstrate that our symbolic audit detects reasoning flaws more accurately than advanced baselines.More broadly, TrustTable outperforms LLM judges in both majority voting with logical weighting and rejection sampling with process supervision.
Guangzhen Zhao, Dechang Kong, Zhenjiang Dong
ACL (1)4
2026 A granular approach for enhancing node representation in heterogeneous graph learning
abstract
Heterogeneous graph learning aims to generate meaningful node representations for graph-structured data with diverse node types and complex relations, facilitating downstream tasks such as node classification and clustering. However, existing methods often emphasize either coarse-grained relational structures or fine-grained node attributes, paying limited attention to the other, which constrains their ability to fully capture the intricate interplay between nodes and relations. To address this limitation, we propose a novel Granular Interaction Heterogeneous Graph Auto-Encoder (GIHGAE), which effectively balances granular fusion and interactions in heterogeneous graph learning. Specifically, GIHGAE employs a relation-level encoder as the primary structure extractor to capture coarse-grained relational dependencies across the graph. Complementarily, we design a node-level encoder that integrates fine-grained contextual details from diverse node attributes, refining representations. These multi-granular features are fused into holistic node embeddings. Additionally, to ensure seamless integration of fine-grained and coarse-grained information, we introduce a global-level decoder to model interactions between nodes and relations explicitly. Finally, to further enhance GIHGAE, we incorporate a dual-loss mechanism, combining reconstruction loss for feature preservation and prediction loss to enhance downstream task performance. Extensive experimental evaluations in heterogeneous graph learning tasks highlight the strong performance of GIHGAE, which consistently outperforms current state-of-the-art methods in classification accuracy, clustering quality, and link prediction performance.
Ying Sun 0023, Hongjiang Ye, Feiyi Xu, Zhenjiang Dong, Yanfei Sun
Future Gener. Comput. Syst.4
2026 PFLSE: A Personalized Federated Learning Framework Based on Shannon Entropy Metric for Intrusion Detection in IIoT
abstract
Intrusion detection is a crucial method for addressing the security risks of the Industrial Internet of Things (IIoT). However, acquiring substantial and high-quality training data can be challenging for centralized schemes. While federated learning has shown great application prospects as a secure distributed solution, it also encounters the problems of heterogeneous and imbalanced data in real-world production environments. In this article, we propose a personalized federated learning scheme based on Shannon entropy metric (PFLSE), aimed at providing a high-accuracy customized detection model for local organizations. This scheme introduces Shannon entropy into the aggregation mechanism, allowing the edge agent model, which contains richer global information, to carry greater weight in the aggregation process. In the local training process, a two-stage training strategy based on the concept of personalized layer is firstly applied to strengthen the global features and local personalized representations. Secondly, considering the differential balance degree between various edge agent data, a Shannon entropy based dynamic loss function (SDL) is proposed, which combines focal loss and cross-entropy loss, to improve training stability and alleviate the difficulty of training on imbalanced data. Finally, a comprehensive experiment simulating a real-world environment shows that PFLSE exhibits reliable intrusion detection performance across metrics such as accuracy, precision, andF1-Score. Furthermore, it outperforms other methods in the scenarios involving non-independent and identically distributed (non-IID) data.
Xingjian Zhu, Jialin Hua, Tian Li 0008, Zhenjiang Dong, Yanfei Sun
IEEE Internet Things J.6
2026 Multi-Component Fault Tolerance and Path Construction in Interconnection Networks
abstract
In the realm of interconnection networks, reliability analysis is of utmost importance, especially considering the increasing vulnerability of components as the network scales. Fault tolerance is a key aspect in this regard, and extra connectivity and component connectivity are two crucial metrics for its assessment. In this paper, we establish a theoretical framework for multi-component fault tolerance in the augmentedk-aryn-cubeAQn,k, a hypercube-derived interconnection network commonly used in distributed-memory architectures. We derive a general result for ther-component connectivity ofAQn,kas$4n(n - 1) - \lfloor{\frac{{5{{(r - 1)}^2}}}{2}}\rfloor$(n≥ 4,k≥ 4, and 2 ≤r≤n). Furthermore, we extend the result to explore theh-extrar-component connectivity ofAQn,kas (8n− 10)(r−1) −2(r−2) (n≥ 4,k≥ 4,h= 1 and 2≤r≤n). Based on these theoretical results, we propose a novel fault-tolerant path algorithm forAQn,kthat handlesh-extrar-component faults. The algorithm first preprocesses and classifies fault-free components, which efficiently determines whether two fault-free nodes belong to the same component, thereby avoiding ineffective path searches. When two fault-free nodes are in the same component, we employ a hybrid greedy-BFS algorithm to construct fault-free paths between them. To validate the algorithm, we conduct comprehensive simulations onAQn,kwith varying parameters. The experimental results demonstrate that the proposed algorithm achieves constant-time path existence queries after preprocessing, significantly reduces path discovery time in multi-query scenarios compared to conventional methods, and maintains near-optimal path lengths while exhibiting superior scalability as network dimensions increase. Furthermore, the algorithm demonstrates robust and highly efficient performance even under fault conditions significantly exceeding theoretical connectivity limits. Additionally, the greedy strategy effectively resolves the vast majority of pathfinding scenarios, confirming its effectiveness underh-extrar-component fault conditions.
Xueli Sun, Shuangxiang Kan, Jianxi Fan, Weibei Fan, Zhenjiang Dong
IEEE Trans. Computers6
2026 HI-CKKS: Is High-Throughput Neglected? Reimagining CKKS Efficiency With Parallelism
abstract
The rapid advancement of the Industrial Internet of Things (IIoT) has positioned data privacy protection as a critical challenge in smart manufacturing. The Cheon–Kim–Kim–Song (CKKS) homomorphic encryption scheme, with its floating-point approximation and single instruction multiple data support capabilities, has emerged as a vital technology for IIoT privacy preservation. However, the high computational complexity of homomorphic multiplication significantly constrains its application performance in real-world industrial scenarios. While existing optimizations primarily focus on latency reduction, research on high-throughput parallel optimization remains notably insufficient. To address this, we propose HI-CKKS, a GPU-based high-throughput CKKS homomorphic multiplication optimization scheme for IIoT, achieving performance breakthroughs through multilevel technical innovations. By designing a batch-processing asynchronous execution architecture, we resolve host-server interaction delays in edge-cloud collaboration. A hierarchical hybrid optimization strategy combining instruction set acceleration, kernel fusion, and dynamic memory management significantly enhances (I)NTT operation efficiency. In addition, we construct a multidimensional parallel homomorphic multiplication model tailored for IIoT high-concurrency task characteristics. Experimental results show that HI-CKKS improves the throughput of NTT, INTT, and homomorphic multiplication by$175.08\times$,$191.27\times$, and$679.57\times$over CPU baselines. When compared to state-of-the-art GPU implementations on similar platforms, HI-CKKS attains$1.54\times$–$14.70\times$speedup for (I)NTT and$55.8\times$–$266.87\times$speedup for homomorphic multiplication. Our work provides an efficient solution for secure computation in IIoT and offers a new approach to performance optimization of homomorphic encryption in edge-cloud collaborative scenarios.
Fuyuan Chen, Jiankuo Dong, Zhenjiang Dong, Wangchen Dai
IEEE Trans. Ind. Informatics4
2026 RIGHT: GPU-Optimized Parallel PQC HAETAE for High-Throughput Cryptographic Acceleration
abstract
The rapid development of quantum computing poses a significant threat to the security of the Industrial Internet of Things (IIoT), rendering traditional cryptographic systems inadequate for ensuring long-term security. As a fundamental technology for establishing trust in IIoT networks, digital signatures are essential for secure device authentication, data integrity verification, and the protection of communication channels. However, implementing efficient digital signature schemes in resource-constrained embedded devices presents significant challenges, particularly when faced with the demands of postquantum cryptography (PQC). We propose a scheme for resource-intensive GPU optimization for HAETAE throughput tuning (RIGHT). Specifically, we employ coarse-grained parallelism and kernel fusion to maximize the parallel processing capabilities of GPUs. Furthermore, we present a hierarchical data locality optimization method for NTT/INTT and FFT, and optimize the SHAKE256 using CUDA PTX instructions, further enhancing throughput to meet the high concurrency and computational performance requirements of IIoT applications. On the Jetson Xavier, the signing performance of HAETAE is 1.25 x, 1.09 x, and 1.31 x that of the Intel i7-10700 K CPU with AVX2, while verification performance is about 4 times faster, demonstrating low power consumption and high efficiency, making it suitable for edge computing. Specifically, on the NVIDIA RTX 4090, the signature throughput reaches 155 324 ops/s, and the verification throughput reaches 4 276 075 ops/s, showcasing significant parallel computing capabilities, which is ideal for large-scale digital signature verification tasks in IIoT.
Jiankuo Dong, Yuze Hou, Mengke Liu, Lunjie Li, Zhenjiang Dong
IEEE Trans. Ind. Informatics6
2026 A Scalable and High-Performance Architecture for Data Center Networks
Xuanli Liu, Weibei Fan, Zhenjiang Dong, Fu Xiao 0001, Mengjie Lv, Xueli Sun, Sun-Yuan Hsieh
IEEE Trans. Netw.3
2026 A Graph Neural Network Approach for Hybrid Node-Edge Fault Diagnosis in Interconnection Networks Under the HPMC* Model
abstract
Fault diagnosis is crucial for ensuring the reliability of interconnection networks. Traditional diagnostic models usually assume that edges connected to faulty nodes are fault-free, which is unrealistic in practice where both node and edge failures can occur simultaneously. The recently proposed HPMC* diagnostic model provides a more realistic framework by considering both node and edge failures simultaneously, but existing diagnostic approaches under this model have significant limitations in handling complex fault scenarios. This paper proposes HYBRID-GNN, the first graph neural network-based approach for hybrid fault diagnosis under the HPMC* model. HYBRID-GNN employs an edge-enhanced GraphSAGE with comprehensive feature engineering that extracts diagnostic characteristics from HPMC* syndrome data and enables joint training for node and edge fault prediction. HYBRID-GNN learns complex fault patterns from syndrome data, overcoming traditional diagnosability constraints. Experiments on multiple interconnection network topologies show that HYBRID-GNN matches the traditional algorithm in node fault diagnosis (achieving over 99% accuracy within the hybrid diagnosability bound), while delivering substantially higher performance in link fault diagnosis (with accuracy above 97%). Even beyond the diagnosability bound, HYBRID-GNN remains robust, maintaining over 98% node accuracy and over 83% link precision under high fault rates. Furthermore, results on real-world networks further validate its practical effectiveness, achieving over 99% node accuracy and over 95% link accuracy.
Xueli Sun, Shuangxiang Kan, Weibei Fan, Zhenjiang Dong, Jianxi Fan
IEEE Trans. Netw.5
2025 SAKGR: Structure-Augmented Knowledge Graph Representations for Reasoning Over Large Language Models in Medical Question Answering
abstract
Medical question answering (QA) systems must deliver accurate, interpretable, and trustworthy responses to support clinical decision-making and patient education. However, existing knowledge-enhanced QA methods often overlook the structural richness of biomedical knowledge graphs (KGs) or delegate all retrieval and reasoning to large language models (LLMs), resulting in hallucinations and significant computational overhead. To address these limitations, we propose SAKGR (Structure-Augmented Knowledge Graph Representations), a novel framework that integrates structured knowledge into LLMbased reasoning. SAKGR constructs structure-aware representations for medical entities by combining semantic embeddings with graph structural features through contrastive learning. These enriched representations enable accurate entity linking and facilitate efficient evidence subgraph retrieval via a hybrid approach that combines path-based search with neighborhood pruning. The extracted subgraphs are then converted into natural language prompts to guide the LLM's reasoning process, effectively incorporating both contextual and structural cues. Experimental results demonstrate that SAKGR significantly enhances retrieval quality and medical QA accuracy while reducing reliance on computationally expensive LLM-driven graph operations, thus offering a scalable and interpretable solution for medical question answering. Our code and data are available on https://github.com/deeomnjobs/sakgr.
Kang Xu 0001, Peihan Cai, Liqun Lu, Zhenjiang Dong
BIBM4
2025 PIESCCA: A Chinese Dataset for Pharmacovigilance Information Extraction Based on Standardised Case Causality Assessment of National Medical Products Administration
abstract
Adverse drug reactions (ADRs) can reduce therapeutic efficacy and pose serious risks. Pharmacovigilance monitors and prevents ADRs, but current evaluations rely on manual text assessments, causing inefficiency and potential errors. To address this, we introduce PIESCCA, a Chinese dataset for pharmacovigilance information extraction aligned with the Standardised Case Causality Assessment (SCCA) guidelines from the NMPA. PIESCCA frames causality assessment as four NLP tasks: named entity recognition (NER), relation extraction (RE), event extraction (EE), and question answering (QA). It is curated from over 50,000 real-world reports collected in the past five years, with$2,100+$representative cases selected and annotated under strict quality standards. We benchmark PIESCCA using four LLMs and seven PLM-based models, establishing a foundation for automated pharmacovigilance research. Our code and data are available at https://github.com/deeomnjobs/PIESCCA.
Kang Xu 0001, Peihan Cai, Liqun Lu, Zitian Wang, Lingli Jiao, Danhua Ma, Zhenjiang Dong
BIBM9
2025 Power allocation optimization for hybrid IRS-assisted 6G V2V communication
Xiaoxuan Hu, Xianyu Wei, Liang Shan 0020, Zhenjiang Dong, Yanfei Sun
Comput. Networks5
2025 FDSS: Flight data sharing scheme based on blockchain with dynamic, secure and efficient consensus algorithm
Feiyi Xu, Shihao Hu, Ying Sun 0023, Xiaoxuan Hu, Yanfei Sun, Zhenjiang Dong
Comput. Networks7
2025 Spatio-Temporal Dynamic Interlaced Network for 3D human pose estimation in video
Feiyi Xu, Jifan Wang, Ying Sun 0023, Zhenjiang Dong, Yanfei Sun
Comput. Vis. Image Underst.5
2025 Cross-Domain Open-Set Fault Diagnosis for Rotating Machinery Based on Frequency-Aware Model With Neighborhood Invariance
abstract
Domain adaptation (DA) is a frequently used technique in intelligent fault diagnosis. However, existing DA methods presume that the source and target domains have the same label space. Due to the complexity of industrial operation conditions, new fault types will inevitably occur. Thus, the above assumption is only sometimes satisfied. To overcome this issue, we propose a novel Frequency-Aware Model with Neighborhood Invariance (FAN) for cross-domain open-set fault diagnosis. Firstly, we comprehensively consider the domain shift phenomenon in time and frequency features and construct an encoder based on the Fourier Neural Operator (FNO) to extract potential invariant information efficiently. Secondly, we expect known class samples to be mapped to an invariant neighborhood to separate unknown classes. Based on this, we adopt neighborhood invariance learning to reduce the intra-domain variations in the target domain and form robust discriminative boundaries. Extensive experiments on public and real-world datasets demonstrate that FAN outperforms the comparison methods and has flexibility.
Yu Gao 0015, Ying Sun 0023, Xingjian Zhu, Genxin Chen, Zhenjiang Dong, Yanfei Sun
IEEE Internet Things J.6
2025 Optimized Cross-Chain Transactions With Aggregated Zero-Knowledge Proofs: Enhancing Efficiency and Security
abstract
With the rapid development of the blockchain industry and the widespread adoption of IoT devices, which are often deployed on different blockchains, the need for cross-chain value and data exchange has become increasingly important. However, existing cross-chain transactions face challenges such as low efficiency, high costs, and insufficient security. To address these issues, this paper proposes a cross-chain transaction scheme based on aggregated zero-knowledge proofs. This scheme optimizes the allocation of computing resources in a distributed environment and employs a multi-branch balanced Merkle tree to construct aggregated zero-knowledge proofs, significantly reducing the verification costs for batch cross-chain transactions.To further enhance data privacy and integrity, this paper introduces the Secure Aggregated Block Verification (SABV) algorithm and improves system consistency and reliability through the Local Merkle Tree Rebalance (LMTR) algorithm. In addition, this paper analyzes the basic security of the proposed scheme when implemented in adversarial environments and provides countermeasures for common threats in distributed systems. Finally, simulations and actual deployment on the Ethereum test network were conducted. The results indicate that our method reduces CPU usage, memory consumption, and time expenditure by 50.10%, 99.03%, and 99.47%, respectively, during the generation of zero-knowledge proofs for batch cross-chain transactions. At the same time, building upon the performance improvements of the existing zero-knowledge proofs, our approach also demonstrates significant enhancements in contract deployment and cross-chain transaction efficiency.
Xiaoxuan Hu, Xiangting Chen, Zhenjiang Dong, Yanfei Sun, Bingyi Fang
IEEE Internet Things J.3
2025 REMODT: Reputation-Driven Efficient Many-to-One Data Trading Based on Blockchain
Xiaoxuan Hu, Yinchuan Hai, Tian Li 0008, Zhenjiang Dong, Yanfei Sun
IEEE Internet Things J.4
2025 Advancing rule learning in knowledge graphs with structure-aware graph transformer
Kang Xu 0001, Miqi Chen, Zhenjiang Dong
Inf. Process. Manag.4
2025 MITU: Locating relevant tutorial fragments of APIs with multi-source API knowledge
Di Wu 0014, Hongyu Zhang 0002, Yang Feng 0003, Zhenjiang Dong
J. Syst. Softw.4
2025 AMFMER: A multimodal full transformer for unifying aesthetic assessment tasks
Can Su, Xiaoxuan Hu, Mengwei Chen, Yanfei Sun, Zhenjiang Dong, Tianliang Liu, Jiebo Luo 0001
Signal Process. Image Commun.6
2025 A Highly Reliable Multiplexing Scheme in Hypercube-Structured Hierarchical Networks
abstract
The design and optimization of network topologies play a critical role in ensuring the performance and efficiency of high-performance computing (HPC) systems. Traditional topology designs often fall short in satisfying the stringent requirements of HPC environments, particularly with respect to fault tolerance, latency, and bandwidth. To address these limitations, we propose a novel class of hierarchical networks, termed Hypercube-Structured Hierarchical Networks (HHNs). This architecture generalizes and extends existing architectures such as half hypercube networks and complete cubic networks, while also introducing previously unexplored hierarchical designs. HHNs exhibit several advantages, particularly in high-performance computing. Most notably, their high connectivity enables efficient parallel data processing, and their hierarchical structure supports scalability to accommodate growing computational demands. Furthermore, we present a unicast routing strategy and a broadcast algorithm for HHNs. A fault-tolerant algorithm is also designed based on the construction of disjoint paths. Experimental evaluations demonstrate that HHNs consistently outperform mainstream architectures in critical performance metrics, including scalability, latency, and robustness to failures.
Xuanli Liu, Zhenjiang Dong, Weibei Fan, Mengjie Lv, Xueli Sun, Sun-Yuan Hsieh
IEEE Trans. Computers2
2025 CSCR: A Cross-View Intelligent Scheduling Method Implemented via Cloud Computing Workflow Reduction
abstract
The surge in the development of artificial intelligence has led to increases in the complexity of computational tasks and the resource demands within cloud computing scenarios. Therefore, intelligent scheduling methods have formed a crucial research area. Solving complex scheduling problems requires many problem feature and long-sequence decision-making observations as possible. To address the workflow scheduling problem under the limited capabilities of models, workflow reduction and cross-view workflow scheduling problems are first proposed in this paper, with the optimization objectives and constraints of each problem described. Second, a cross-view intelligent scheduling method implemented via cloud computing workflow reduction (CSCR), including a workflow reduction sorting algorithm (Task-priority ranker), an intelligent reduction algorithm (Workflow view-transformer), and a cross-view intelligent scheduling algorithm (Joint-scheduler), is proposed. We also propose an intelligent scheduling architecture under the workflow reduction paradigm. By reducing the workflow, we provide multiple views that support the decision-making processes of deep reinforcement learning-based scheduling models and coordinate workflow views before and after the reduction step to achieve cross-view joint scheduling. Experimental results show that CSCR achieves minimum advantages of 42.1%, 43.2%, and 33.3% in terms of three workflow reduction indicators over four other algorithms, significantly optimizing the effect of the employed scheduling model.
Genxin Chen, Xingjian Zhu, Jialin Hua, Zhenjiang Dong, Yanfei Sun
IEEE Trans. Cloud Comput.5
2025 GOLF: Unleashing GPU-Driven Acceleration for FALCON Post-Quantum Cryptography
abstract
Quantum computers leverage qubits to solve certain computational problems significantly faster than classical computers. This capability poses a severe threat to traditional cryptographic algorithms, leading to the rise of post-quantum cryptography (PQC) designed to withstand quantum attacks. FALCON, a lattice-based signature algorithm, has been selected by the National Institute of Standards and Technology (NIST) as part of its post-quantum cryptography standardization process. However, due to the computational complexity of PQC, especially in cloud-based environments, throughput limitations during peak demand periods have become a bottleneck, particularly for FALCON. In this paper, we introduce GOLF (GPU-accelerated Optimization for Lattice-based FALCON), a novel GPU-based parallel acceleration framework for FALCON. GOLF includes algorithm porting to the GPU, compatibility modifications, multi-threaded parallelism with distinct data, single-thread optimization for single tasks, and specific enhancements to the Fast Fourier Transform (FFT) module within FALCON. Our approach achieves unprecedented performance in FALCON acceleration on GPUs, setting the highest throughput record in the history of FALCON digital signature generation and verification. On the NVIDIA RTX 4090, GOLF reaches a signature generation throughput of 420.25 kops/s and a signature verification throughput of 10,311.04 kops/s. These results represent a 58.05× / 73.14× improvement over the reference FALCON implementation and a 7.17× / 3.79× improvement compared to the fastest known GPU implementation to date. Additionally, since we have not modified the content of the algorithm, but only optimized its engineering implementation, the security of the algorithm has not changed, and the security of the original algorithm has been maintained. GOLF demonstrates that GPU acceleration is not only feasible for post-quantum cryptography but also crucial for addressing throughput bottlenecks in real-world applications.
Ruihao Dai, Jiankuo Dong, Mingrui Qiu, Zhenjiang Dong, Fu Xiao 0001, Jingqiang Lin 0001
IEEE Trans. Inf. Forensics Secur.4
2025 Symphony of Speeds: Harmonizing Classic McEliece Cryptography With GPU Innovation
abstract
The Classic McEliece key encapsulation mechanism (KEM), a candidate in the fourth-round post-quantum cryptography (PQC) standardization process by the National Institute of Standards and Technology (NIST), stands out for its conservative design and robust security guarantees. Its deployment is impeded by exceptionally large public and secret keys. Modern GPUs offer abundant parallelism and global memory, making them well suited to such key sizes and to high-throughput cryptographic workloads. However, there has not been a systematic implementation of Classic McEliece on GPU platforms. This paper presents the first high-performance implementation of Classic McEliece on NVIDIA GPUs. Firstly, we present the first GPU-based implementation of Classic McEliece, utilizing a “CPU-GPU” heterogeneous approach and a kernel fusion strategy. We significantly reduce global memory accesses, optimizing memory access patterns. This results in encapsulation and decapsulation performance of 28,628,195 ops/s and 3,051,701 ops/s, respectively, for McEliece348864. Secondly, core operations like Additive Fast Fourier Transforms (AFFT), and Transpose AFFT (TAFFT) are optimized. We introduce the concept of the (T)AFFT stepping chain and propose two universal schemes: Memory Access Stepping Strategy (MASS) and Layer-Fused Memory Access Stepping Strategy (LFMASS), which achieve a speedup of 30.56% and 38.37%, respectively, compared to the native GPU-based McEliece6960119 implementation. Thirdly, extensive experiments on the NVIDIA RTX4090 show significant performance gains, achieving up to 344× higher encapsulation and 125× higher decapsulation compared to the official CPU-based AVX implementation, decisively outperforming existing ARM Cortex-M4 and FPGA implementations.
Jiankuo Dong, Zhenjiang Dong, Dung Hoang Duong, Fu Xiao 0001, Jingqiang Lin 0001
IEEE Trans. Inf. Forensics Secur.4
2025 DDCC: Synergizing Denoising Diffusion Probabilistic Models and Curriculum-Based Complexity Control for Insider Threat Detection
abstract
Insider threat detection aims to identify malicious activities by employees that may compromise the confidentiality, integrity, or availability of organizational data. Detecting insider threats poses unique challenges compared to external threats, as internal actors often possess authorized system access and familiarity with organizational systems, enabling them to execute attacks discreetly. This article presents a novel approach to insider threat detection, synergizing denoising diffusion probabilistic models (DDPM) with a curriculum-based complexity control strategy (DDCC). While DDPM has shown promise in anomaly detection, its application to insider threat detection remains relatively unexplored. The proposed approach leverages DDPM for context-aware behavioral modeling and anomaly scoring. The methodology encompasses three primary components: First, employee behavioral logs undergo fusion to aggregate context information across various time scales. Second, a curriculum learning module controls the training process, gradually exposing the model to increasingly complex samples. The training data organization progresses from simple to complex sequences, facilitating more effective feature learning. Third, the denoising diffusion probabilistic model reconstructs the fused employee behavioral sequences for anomaly detection. Experimental validation on the CMU CERT r5.2 and r6.2 datasets demonstrates robust performance in insider threat detection across multiple time granularities of aggregated employee logs. The results underscore the effectiveness of this approach in addressing the nuanced challenges associated with detecting internal threats, highlighting its potential for real-world deployment in organizational security frameworks.
Jiankuo Dong, Zhenjiang Dong, Fuyuan Chen
IEEE Trans. Ind. Informatics4
2025 Heterogeneous Federated Learning Driven by Multi-Knowledge Distillation
abstract
In a fully heterogeneous federated learning environment, the client has significant differences in model structure and local data distribution (Non-IID), and the joint learning of the client model is blocked due to the limited communication content available for interaction in a fully heterogeneous scenario. In this context, the global knowledge constructed by the server through the simple aggregation of the client logits is essentially a fuzzy representation containing a lot of noise and information loss, which is difficult to effectively guide the client model update. To solve these problems, this paper proposes a heterogeneous federated learning framework (FedMkd) based on multi-knowledge distillation fusion to cope with multiple challenges in heterogeneous environments. The FedMkd framework uses a class-grained logits interaction architecture (CLIA) and introduces an efficient knowledge sharing mechanism. It innovatively integrates two knowledge distillation methods: 1) Temperature-Adaptive Knowledge Distillation (TAKD), which provides differentiated temperatures for teacher and student models by adaptively adjusting the distillation temperature, maximizing knowledge transfer between them; 2) Class-related Knowledge Distillation (CRKD), which introduces batch-level sample correlation loss to reduce over-reliance on specific samples or classes and improve the model's understanding of overall data features. We conducted a large number of experiments on four public data sets. The results show that in a variety of data and model heterogeneous scenarios, FedMkd still performs better than the comparison method when the communication overhead is reduced by more than one order of magnitude.
Bin Xu 0014, Longgang Cheng, Qing Wen, Zhensheng Zou, Xiaoxuan Hu, Zhenjiang Dong
IEEE Trans. Mob. Comput.6
2025 AFAS: Arbitrary-Freedom Adaptive Scheduling for Multiworkflow Cloud Computing via Deep Reinforcement Learning
abstract
The in-depth development of artificial intelligence models has supported the high-quality allocation of cloud computing resources. The optimization of workflow scheduling issues in cloud computing has become increasingly critical due to the complexity of computing tasks, constraints on computing resources, and the growing demand for high-quality service. To address the increasingly complex workflow scheduling problems in cloud computing, this paper presents an arbitrary-freedom adaptive scheduling method for cloud computing with multiple workflows based on deep reinforcement learning (termed AFAS), with the workflow makespan and response time as the optimization objectives. First, we define the concept of degrees of freedom in the scheduling context to establish the feature space and foundational decision patterns relevant to multiworkflow scheduling. Second, an adaptive real-time scheduling strategy generation (ARS) algorithm is proposed for multiworkflow scheduling tasks. Third, a composite reward mechanism with an advanced-time-window real-time-reward (ATR) algorithm is designed for intelligent model optimization. Finally, the generation algorithm and intelligent model are fused to perform arbitrary-freedom multiworkflow adaptive scheduling. The experiments show that ATR can significantly increase the frequency of reward generation, AFAS can achieve at least 6.6% better performance than existing methods can achieve, and the incorporation of intelligent models improves the performance of ARS by 2.7%.
Genxin Chen, Jialin Hua, Ying Sun 0023, Zhenjiang Dong, Yanfei Sun
IEEE Trans. Netw. Serv. Manag.5
2025 Research on storage optimization for efficient training of large language models
Biyun Shang, Mo Xu, Junning Xu, Xinyuan Sun, Zhenjiang Dong
World Wide Web (WWW)6
2024 Reliability of Half Hypercube Networks under Cluster Faults
abstract
Malicious attackers frequently aim to partition the network into disjointed segments to facilitate specific attacks. Consequently, enhancing network reliability stands as an effective preventive measure. Connectivity serves as a crucial metric for gauging network reliability, yet classical connectivity inadequately captures a network's fault tolerance in the face of such attacks. To address this, cluster connectivity has been proposed, considering the faults within clusters to improve fault tolerance assessment. In this paper, we establish the cluster connectivity of the half hypercube network HHn. In detail, we show that the K1,1-cluster connectivity of HHnis $\left\lfloor {n/2} \right\rfloor + 1$, where n ≥ 3, and the K1,r- cluster connectivity of HHnis $\left\lceil {\frac{{\left\lceil {n/2} \right\rceil }}{2}} \right\rceil + 1$, where n ≥ 5 and 2 ≤ r ≤ 4, which is almost r times the classical connectivity. This indicates that the network possesses an enhanced capacity to accommodate a greater number of faulty nodes, potentially enabling more effective orchestration of attacks.
Xuanli Liu, Mengjie Lv, Weibei Fan, Xueli Sun, Zhenjiang Dong, Fu Xiao 0001
CSCWD5
2024 The future of API analytics
Di Wu 0014, Hongyu Zhang 0002, Yang Feng 0003, Zhenjiang Dong, Ying Sun 0023
Autom. Softw. Eng.4
2024 DGTRL: Deep graph transfer reinforcement learning method based on fusion of knowledge and data
Genxin Chen, Yu Gao 0015, Xingjian Zhu, Zhenjiang Dong, Yanfei Sun
Inf. Sci.5
2024 Industrial process fault diagnosis based on feature enhanced meta-learning toward domain generalization scenarios
Yu Gao 0015, Ying Sun 0023, Xiaoxuan Hu, Zhenjiang Dong, Yanfei Sun
Knowl. Based Syst.5
2024 ECO-BIKE: Bridging the Gap Between PQC BIKE and GPU Acceleration
abstract
Advancements in quantum computing pose a threat to public-key cryptosystems, leading to the development of post-quantum cryptography. NIST is standardizing candidate algorithms, with BIKE, a code-based key encapsulation mechanism, among those under consideration. Performance is crucial in NIST PQC standardization process, and researchers have introduced a range of optimization techniques for BIKE across various platforms. To the best of our knowledge, our Efficient CryptOgraphy BIKE (ECO-BIKE) represents the first attempt at optimizing the implementation of BIKE on GPU architecture. In this paper, we introduce a comprehensive construction of a 3-threading parallel architecture tailored for the BIKE cryptosystem. This architecture covers a range of computational tasks, addressing operations from low-level to high-level computations. These include a parallel dense polynomial multiplication scheme with a better memory access pattern and a better XOR calculation, which forms the basis for a comprehensive parallel execution framework for the entire BIKE algorithm. Targeted optimizations are implemented for specific modules (KEYGEN, ENCAPS, DECAPS), which collectively enhance the overall efficiency of the algorithm. Our ECO-BIKE exhibits exceptional throughput performance on the NVIDIA GeForce RTX 4090. In the 3-thread mode, the throughput of the KEYGEN, ENCAPS, and DECAPS modules reaches 24.033 kops/s, 277.789 kops/s, and 5.817 kops/s, respectively. Our proposed optimal parallel multiplication scheme achieves a significantly higher overall throughput of 481.302 kops/s. These results highlight the substantial computational advantages our approach provides for cryptographic workloads.
Jiankuo Dong, Yusheng Fu, Xusheng Qin, Zhenjiang Dong, Fu Xiao 0001, Jingqiang Lin 0001
IEEE Trans. Inf. Forensics Secur.4
2024 AUTH: An Adversarial Autoencoder Based Unsupervised Insider Threat Detection Scheme for Multisource Logs
abstract
Deep learning has shown broad research prospects in addressing insider threats, a serious problem currently facing industrial information systems. Although deep learning is able to capture effective feature representations from complex multidimensional data, there are still issues such as strong stealth of insider threat behavior and the imbalance data that need to be solved. Therefore, we propose an adversarial Autoencoder based Unsupervised insider Threat detection scHeme (AUTH). Compared to other methods, AUTH fully considers the role of time feature and event feature in threat detection. In addition, in order to improve the performance of autoencoder models to detect covert threat behaviors, AUTH drives a temporal convolutional network and long short-term memory network-based Adversarial Autoencoder (TL-AAE). Generative Adversarial Theory is introduced to solve the problem of uncertainty in the latent feature of the encoder. Finally, with the sufficient experiments on public datasets, we demonstrate that the usefulness of adding time features and the proposed TL-AAE model to improve threat detection performance. Compared with the baseline, AUTH obtains the area under curve value of 0.932, which is 4.95% higher than the highest result obtained by the baseline. In addition, AUTH obtains the EER value of 0.146, which is 12.57% lower than the lowest result of the baseline.
Xingjian Zhu, Jiankuo Dong, Zhen-Guo Zhou, Zhenjiang Dong, Yanfei Sun, Moyu Wang
IEEE Trans. Ind. Informatics5
2024 A Blockchain Cross-Chain Transaction Method Based on Decentralized Dynamic Reputation Value Assessment
abstract
With the vigorous development of the blockchain industry, cross-chain transactions can effectively solve the problem of “islands of value” caused by the inability to interact between different chains. However, security risks in reputation management caused by cross-chain transactions implemented through notary solutions have always existed. Consequently, this paper proposes a blockchain cross-chain transaction method based on decentralized dynamic reputation value assessment. The notary election phase addresses the issue of the continually changing behaviour of notaries in actual transactions by designing a dynamic evaluation window mechanism based on an RNN. Moreover, a reputation-rating decay mechanism is introduced to avoid the problem of reputation value recovery caused by malicious notaries being inactive for a long time. Relative to alternative reputation assessment models, the proposed method offers a thorough evaluation of user behavior and effectively identifies malicious activities in real-time. Finally, the method was tested by deploying it on the Ethereum blockchain. Our approach offers more dynamic settings for window parameters, adapting to changes in notary behavior and reducing the number of detections within the same timeframe by approximately 59.14%. The weight factor settings are also optimized, allowing for adjustments based on specific situations to achieve accurate reputation values. Overall, this method not only enhances the security of cross-chain transactions but also reduces operational costs by 53.3% compared to traditional technologies.
Xiaoxuan Hu, Yaochen Ling, Jialin Hua, Zhenjiang Dong, Yanfei Sun
IEEE Trans. Netw. Serv. Manag.4
2023 A collaborative scheduling method for cloud computing heterogeneous workflows based on deep reinforcement learning
Genxin Chen, Ying Sun 0023, Xiaoxuan Hu, Zhenjiang Dong, Yanfei Sun
Future Gener. Comput. Syst.5
2022 Research on an Intelligent Computing Offloading Model for the Internet of Vehicles Based on Blockchain
abstract
Aiming at the problems of computing power, reliability and cost when intelligent vehicles deal with computationally intensive and delay-sensitive emerging applications in multiple business scenarios in the Internet of Vehicles, an intelligent computing offloading model is proposed. This can minimize the total system cost under the constraint of time delay and energy consumption. Considering the cost of blockchain and the cost of intelligent vehicles, the DDPG algorithm is used to solve the proposed model. Simulation results show that the method proposed in this paper can effectively reduce the total cost of computing offload and further improve the success rate of computational offloading under the premise of computational offloading safety.
Yaochen Ling, Bin Xu 0014, Zhenjiang Dong, Yanfei Sun
IEEE Trans. Netw. Serv. Manag.5
2019 Seeing Isn't Believing: QoE Evaluation for Privacy-Aware Users
abstract
More and more network media users concern about their privacy issues since they know that their network behaviors are being observed, and thus the observable users' data are not reliable and sufficient in this case. How to evaluate the true quality of experience (QoE) of the privacy-aware users has become a significant technical challenge because of the most majority of existing data-driven QoE evaluation schemes based on the premise of the true and adequate users' observations. To get over this dilemma, this paper proposes a systematic and robust QoE evaluation scheme with unreliable and insufficient observation data. Specifically, we first translate the subjective privacy-aware QoE evaluation problem into an objective rational user analysis procedure. Then, a semantics-based similarity measurement for multidimensional correlation analysis is constructed to classify the observable data. Subsequently, the highlight of this paper lies in proposing a class-level joint user classification and data cleaning strategy by frequently updating the training processes. Through elaborately designing an iterative framework, it can effectively resolve the data sparsity and inconsistency problems due to the user privacy-aware preferences. Importantly, we also introduce an efficient QoE model construction method for online implementation, and numerical results validate its efficiency for different kinds of privacy-aware users.
Liang Zhou 0002, Dan Wu 0001, Xin Wei 0001, Zhenjiang Dong
IEEE J. Sel. Areas Commun.4
2018 Mining IPTV User Behaviors with an Enhanced LDA Model
abstract
With the increasing popularity of IPTV industry, QoE has been regarded as one of the most promising evaluation indicators for IPTV service. However, due to the increasing amount of TV programs and users' mixed preferences, how to recommend interesting programs for users is still a challenging and urgent problem. Existing related researches ignore the personalized recommendation and the prediction of prospective interests for different users from large amounts of TV programs. To solve this problem, this work proposes an enhanced latent Dirichlet allocation (LDA) model to analyze user behaviors and recommend personalized programs of the users' mixed interests. Specifically, we put forward a new attribute called viewing ratio to calculate the proportion of program's time viewed by the user, which could measure users' subjective viewing experience from objective indicators. Based on the proposed model, we improve the accuracy of user behaviors modeling and prediction of prospective interests. Experimental results show that our model has better performances of programs recommendation and TV viewing experience than other models.
Xin Wei 0001, Liang Zhou 0002, Zhenjiang Dong
GLOBECOM5
2018 When Computation Hugs Intelligence: Content-Aware Data Processing for Industrial IoT
abstract
Data service has been considered as one the most prominent characteristics for Industrial Internet of Things (IIoT). This paper studies how to design an optimal computing manner for a general IIoT system. On the theory end, we analyze the relationship between the data processing and the energy consumption through investigating the content correlation of the captured data. Importantly, we derive an exact expression for the performance of IIoT by combining computation with intelligence. On the application end, we design an efficient way to obtain a threshold by approximating the performances of different computing manners, and show how to apply it to practical IIoT applications. We believe that the proposed computation rules hold great significance for the IIoT designer, that is, it is better to use distributed computing manner when the content correlation is high, otherwise, centralized computing manner is better.
Liang Zhou 0002, Dan Wu 0001, Zhenjiang Dong
IEEE Internet Things J.4
2018 Image-to-Video Person Re-Identification With Temporally Memorized Similarity Learning
abstract
With the development of video surveillance in public safety field, there is an increasing research on person re-identification (re-id). In this paper, we address the image-to-video person re-id, in which the probe is an image and the gallery is consists of videos captured by nonoverlapping cameras. Compared with image, video sequence contains more temporal information that can be explored to improve the performance of re-identification system. However, it is challenging to model temporal information in the matching process of image-to-video person re-id. In this paper, we proposed a novel temporally memorized similarity learning neural network for this problem. In specific, the proposed network mainly consisted of two parts, including feature representation sub-network and similarity sub-network. In the first part, we adopted a convolutional neural network (CNN) to extract features from the input image. Given a video sequence of a person, features were first extracted from each its frame by using CNN and further forward to a long shot term memory (LSTM) network to encode the temporal information of video sequence. The outputs of LSTM were concatenated together as the feature vector of video sequences. Finally, the feature vectors of probe image and the video sequence were further forward to the similarity sub-network for distance metric learning. In the proposed framework, the feature representation and the similarity metric learning can be learned and optimized simultaneously. We evaluated the proposed framework on three public person re-id data sets, and the experimental results showed that the proposed approach is effective for the image-to-video person re-id.
Dongyu Zhang 0002, Wenxi Wu, Hui Cheng 0002, Ruimao Zhang, Zhenjiang Dong, Zhaoquan Cai 0001
IEEE Trans. Circuits Syst. Video Technol.5
2018 Greening the Smart Cities: Energy-Efficient Massive Content Delivery via D2D Communications
abstract
Massive multimedia services have been considered as one the most prominent characteristics for smart cities. In this paper, we propose an energy-efficient content delivery system via the device-to-device communications, which realizes the large-scale content delivery among mobile devices with constrained energy, unpredictable demand, limited storage, random mobility, and opportunistic transmission. The highlights of this paper lie in two parts. On the theoretical end, through exploring the relationship among the coding, storage, and transmission, a systematic energy-saving content delivery fashion is investigated. On the technical end, a totally distributed content delivery system is designed in a simple and efficient manner, in which each device only utilizes local information to make decisions and implements its own scheme individually. Importantly, the proposed scheme is realized in a practical smart city system, and numerical results demonstrate that it is flexible to various users' needs and communication environments.
Liang Zhou 0002, Dan Wu 0001, Zhenjiang Dong
IEEE Trans. Ind. Informatics4
2017 Cascaded LSTMs Based Deep Reinforcement Learning for Goal-Driven Dialogue
Xiaojie Wang 0006, Zhenjiang Dong
NLPCC3
2017 Jointly Modeling Intent Identification and Slot Filling with Contextual and Hierarchical Information
Liyun Wen, Xiaojie Wang 0006, Zhenjiang Dong
NLPCC3
2017 An Integrated Quality Assessment for IPTV Operation and Maintenance
abstract
This paper proposes a novel quality assessment scheme for IPTV operation and maintenance. It is an integrated system to make the IPTV network fault location diagnosis more efficiently and accurately. Specifically, the potential user complaint and potential warning facility are integrated for constructing the IPTV service quality assessment system. When handling the determination of the potential warning facility, objective Quality of Experience (QoE) indicators reflecting users' viewing behaviors are considered and integrated with traditional Quality of Service (QoS) indicators. Based on this, a novel feature selection algorithm is proposed for replacing existing ones in decision tree generation and pruning, efficiently and feasibly realizing faulted equipment prediction. Experimental results show that the prediction accuracy for faulted equipments can be further enhanced when compared with existing algorithms.
Xin Wei 0001, Zhifeng Wu, Liang Zhou 0002, Zhenjiang Dong
VTC Spring4
2017 A Dictionary-Based Approach for Identifying Biomedical Concepts
abstract
In this research, we provided a dictionary-based approach for identifying biomedical concepts from the literature. The approach first crawled experimental corpus by E-utilities and built a concept dictionary. Then, we developed an algorithm called Variable-step Window Identification Algorithm (VWIA) for matching biomedical concepts based on preprocessing, POS tagging and the formation of phrase block. The approach could identify embedded biomedical concepts and new concepts, which could identify concepts more completely. The proposed approach obtain 95.0% F-measure overall for the test dataset. Thus, it is promising for the method of biomedical text mining.
Lejun Gong, Ronggen Yang, Zhenjiang Dong
Int. J. Pattern Recognit. Artif. Intell.4
2016 A Demonstration of QA System Based on Knowledge Base
Zhenjiang Dong, Jingqiang Chen, Huakang Li, Tao Li 0001
APWeb (2)1
2015 What Causes Different Emotion Distributions of a Hot Event? A Deep Event-Emotion Analysis System on Microblogs
abstract
Current online public opinion analysis systems can explore lots of hot events and present the public emotion distribution for each event, which are useful for the governments and companies. However, the public emotion distributions are just the shallow analysis of the hot events, more and more people want to know the hidden causation behind the emotion distributions. Thus, this paper presents a deep Event-Emotion analysis system on Microblogs to reveal what causes different emotions of a hot event. We here use several related sub-events to describe a hot event in different perspectives, accordingly these sub-events combined with their different emotion distributions can be used to explain the total emotion distribution of a hot event. Experiments on 15 hot events show that the above idea is reasonable to exploit the emotion causation and can help people better understand the evolution of the hot event. Furthermore, this deep Event-Emotion analysis system also tracks the amount treads and emotion treads of the hot event, and presents the deep analysis based on the user profile.
Bing Qin 0001, Zhenjiang Dong, Ting Liu 0001
NLPCC3
2015 Learning Visual-Spatial Saliency for Multiple-Shot Person Re-Identification
abstract
Recognizing persons across non-overlapping camera views, known as person re-identification, has received increasing attentions for its importance in many surveillance applications. However, most of existing methods rely on pre-training steps to ensure their performance and ignore the body prior knowledge of pedestrians. In this letter, we propose a novel non-training method for person re-identification which learns visual-spatial saliency from voter images and the given query image. First we segment pedestrian images into small regions and use two hypergraphs to represent the visual and spatial relationship among regions. Then we formulate the visual-spatial saliency learning as a joint hypergraph ranking problem by simultaneously considering the human body prior and the appearance similarity among pedestrians. Finally, the visual-spatial saliency is incorporated in region-based matching to improve the performance of person re-identification. Experimental evaluation on three publicly available datasets demonstrates the effectiveness of our approach.
Xiaojin Gong, Zhenjiang Dong
IEEE Signal Process. Lett.4
2014 Manifold: A parallel simulation framework for multicore systems
abstract
This paper presents Manifold, an open-source parallel simulation framework for multicore architectures. It consists of a parallel simulation kernel, a set of microarchitecture components, and an integrated library of power, thermal, reliability, and energy models. Using the components as building blocks, users can assemble multicore architecture simulation models and perform serial or parallel simulations to study the architectural and/or the physical characteristics of the models. Users can also create new components for Manifold or port existing models. Importantly, Manifold's component-based design provides the user with the ability to easily replace a component with another for efficient explorations of the design space. It also allows components to evolve independently and making it easy for simulators to incorporate new components as they become available. The distinguishing features of Manifold include i) transparent parallel execution, ii) integration of power, thermal, reliability, and energy models, iii) full system simulation, e.g., operating system and system binaries, and iv) component-based design. In this paper we provide a description of the software architecture of Manifold, and its main elements - a parallel multicore emulator front-end and a parallel component-based back-end timing model. We describe a few simulators that are built with Manifold components to illustrate its flexibility, and present test results of the scalability obtained on full-system simulation of coherent shared-memory multicore models with 16, 32, and 64 cores executing PARSEC and SPLASH-2 benchmarks.
Jun Wang 0077, Jesse G. Beu, Rishiraj A. Bheda, Thomas M. Conte, Zhenjiang Dong, Chad D. Kersey, Mitchelle Rasquinha, George F. Riley, William J. Song, Sudhakar Yalamanchili
ISPASS5
2013 A Study of the Effect of Partitioning on Parallel Simulation of Multicore Systems
abstract
There has been little research that studies the effect of partitioning on parallel simulation of multicore systems. This paper presents our study of this important problem in the context of Null-message-based synchronization algorithm for parallel multicore simulation. This paper focuses on coarse grain parallel simulation where each core and its cache slices are modeled within a single logical process (LP) and different partitioning schemes are only applied to the interconnection network. In this paper we show that encapsulating the entire on-chip interconnection network into a single logical process is an impediment to scalable simulation. This baseline partitioning and two other schemes are investigated. Experiments are conducted on a subset of the PARSEC benchmarks with 16-, 32-, 64- and 128-core models. Results show that the partitioning scheme has a significant impact on simulation performance and parallel efficiency. Beyond a certain system scale, one scheme consistently outperforms the other two schemes, and the performance as well as efficiency gaps increases as the size of the model increases - with up to 4.1 times faster speed and 277% better efficiency for 128-core models. We explain the reasons for this behavior, which can be traced to the features of the Null-message-based synchronization algorithm. Because of this, we believe that, if a component has increasing number of inter-LP interactions with increasing system size, such components should be partitioned into several sub-components to achieve better performance.
Zhenjiang Dong, Jun Wang 0077, George F. Riley, Sudhakar Yalamanchili
MASCOTS1
2013 Optimizing parallel simulation of multicore systems using domain-specific knowledge
abstract
This paper presents two optimization techniques for the basic Null-message algorithm in the context of parallel simulation of multicore computer architectures. Unlike the general, application-independent optimization methods, these are application-specific optimizations that make use of system properties of the simulation application. We demonstrate in two aspects that the domain-specific knowledge offers great potential for optimization. First, it allows us to send Null-messages much less eagerly, thus greatly reducing the amount of Null-messages. Second, the internal state of the simulation application allows us to make conservative forecast of future outgoing events. This leads to the creation of an enhanced synchronization algorithm called Forecast Null-message algorithm, which, by combining the forecast from both sides of a link, can greatly improve the simulation look-ahead. Compared with the basic Null-message algorithm, our optimizations greatly reduce the number of Null-messages and increase simulation performance significantly as a result. On a subset of the PARSEC benchmarks, a maximum speedup of about 6 is achieved with 17 LPs.
Jun Wang 0077, Zhenjiang Dong, Sudhakar Yalamanchili, George F. Riley
SIGSIM-PADS2
2012 Spam Short Messages Detection via Mining Social Networks
Jianyun Liu, Zhaoxiang Zhang 0001, Yunhong Wang 0001, Xue-Mei Yuan, Zhenjiang Dong
J. Comput. Sci. Technol.7