Junsheng Chang

dblp:95/5495 · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 2 first-author · 7 since 2021Software engineering, systems software and programming languages · 6 · 3 since 2021Computer networks · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Row-wise Inter-Phase Pipelining for Hardware-Efficient GCN Acceleration
abstract
Graph Convolutional Networks (GCNs) are widely deployed for learning on graph-structured data. Yet, their canonical two-phase execution, consisting of feature transformation followed by neighborhood aggregation, suffers from severe inefficiencies on existing accelerators. The prevailing phase-decoupled model explicitly stores dense intermediate results to off-chip memory after transformation and reloads them for aggregation, incurring substantial inter-phase data movement. Additionally, irregular access patterns cause intra-phase redundancy through repeated fetches of sparse inputs. To overcome both forms of redundancy, we propose IPRS-GCN, a streaming accelerator built on Inter-Phase Row-Streaming (IPRS). IPRS fuses transformation and aggregation at row granularity into a single pipeline, forwarding each node’s transformed features directly to its neighbors via the Row Queue as soon as they are computed. This eliminates off-chip storage of intermediates and enforces single-pass streaming access to inputs while keeping weights resident on-chip. The architecture realizes this dataflow through lightweight mechanisms including conflict-aware non-zero packing and degree-adaptive buffering, achieving near-ideal compute utilization with minimal on-chip footprint. Evaluated across five real-world datasets, IPRS-GCN achieves a geometric mean speedup of 107.74 × over an NVIDIA V100 GPU, outperforms HyGCN and AWB-GCN by 7.74 × and 2.12 ×, respectively, and improves energy efficiency by 1.37 ×.
Junsheng Chang, Yang Guo 0003, Li Shen 0007
CF2
2025 Accelerating DFS-based Subgraph Matching on GPU via Reusing Intersection
abstract
Subgraph matching is a well-known NP-hard problem widely applied in fields such as bioinformatics, cheminformatics, and social network analysis. It aims to enumerate all embeddings in a data graph that are isomorphic to a query graph. Subgraph matching algorithms can be roughly classified into BFS-based and DFS-based algorithms. The intersection operation is the core operation in both types of algorithms and consumes a significant amount of time. There are numerous repeated intersection operations in subgraph matching, and their results can be reused. Recent studies have focused on implementing the DFS-based algorithm on GPUs with an explicit stack. We categorize the reuse in the DFS-based algorithm into two types: within-stack reuse and across-stack reuse. Previous works only considered within-stack reuse, which has a narrow scope of application and lacks generality for some query graphs. In this paper, we are the first to propose a method for across-stack reuse in DFS-based algorithms on GPUs. We introduce a tree-structured copy stack to reuse duplicate intersections and design a secondary checking mechanism to ensure the correctness of the reuse. We use the non-blocking lock mechanism to avoid read-write races and conduct a rigorous theoretical analysis to prove that the impact of the non-blocking lock on reuse is negligible. Moreover, we propose a GPU-specific optimization method that uses key nodes to reduce memory transactions during intersection operations. Compared with the state-of-the-art DFS-based matching algorithms, our work achieves $1.12 \times$ to $1.95 \times$ speedup, and the reuse rate reaches 78.68%.ACM Reference Format:Chen Chen, Shanzhi Gu, Junsheng Chang, and Li Shen. 2025. Accelerating DFS-based Subgraph Matching on GPU via Reusing Intersection. In. ACM, New York, NY, USA, 12 pages. https://doi.org/10.1145/nnnnnnn.nnnnnnnn
Chen Chen 0016, Shanzhi Gu, Junsheng Chang, Li Shen 0007
PACT3
2025 Combination of Storage and Accumulation for Synchronous SpMV Acceleration on FPGAs with HBM
abstract
Sparse matrix-vector multiplication (SpMV) is a crucial computational operation in various fields, such as graph computation, machine learning, and molecular dynamics.However, due to the irregular data distribution and low density of non-zero elements, the performance of SpMV is typically inferior to that of dense matrix computations.To tackle this issue, numerous optimization efforts have been made on FPGAs equipped with high-bandwidth memory (HBM), addressing problems like excessive transmission latency and load imbalance between channels.Nevertheless, several key challenges remain that hinder the performance of FPGA-based SpMV accelerators, including (1) the tendency for cache blocks to frequently miss and be replaced due to the irregular distribution of non-zero elements in sparse matrices; (2) control divergence issues within the single instruction multiple data (SIMD) SpMV accelerator architecture.To overcome these challenges, this study introduces CoSpMV, an accelerator design for FPGAs with HBM.It incorporates (1) a matrix block synchronous processing technique under ping-pong buffering, (2) a more efficient data compression format known as row-column compressed coordinate format (R3Coo), and (3) a storage accumulation module.R3Coo enhances data transmission efficiency by compressing bit width and improving the efficiency of flag bits; the matrix block synchronous processing technique under ping-pong buffering conceals replacement overhead and prevents irregular cache access by dividing and synchronously processing matrix data into rows and columns; the storage accumulation module eliminates inconsistent control flow by decoupling the addition
DongHuan Xie, Qingjie Lang, Dunbo Zhang, Junsheng Chang, Li Shen 0007
CF5
2025 MAAU-UIE: Multiple Attention Aggregation U-Net for Underwater Image Enhancement
Junsheng Chang, Yijun Zhang 0003, Zongtang Hu, Xunlun Ye
CVM (1)1
2025 X-SA: An Efficient Configurable Systolic Array Computing Architecture for GPGPU
abstract
GPGPUs are pivotal for edge AI, but resource constraints demand efficient low-precision computation. Conventional GPGPUs face challenges in resource utilization, particularly with irregular matrices common in AI, and memory bandwidth limitations on edge devices. Traditional fixed-size systolic arrays often suffer from underutilization under varying workloads. This paper introduces X-SA, a configurable systolic array architecture tailored for INT8 matrix multiplication on GPGPUs in resource-constrained edge environments. X-SA distinctively employs a parameterized$2 \times N$processing element design enabling dynamic computational scaling, unlike fixed systolic arrays. It integrates an interleaved matrix buffer to alleviate memory bottlenecks and optimize dataflow. Experimental results demonstrate X-SA achieves a$2.83 \times$performance speedup over the Vortex baseline with minimal Look-Up Table overhead of 2.8% and Flip-Flops overhead of 1.4%.. It offers comparable performance to a standard$4 \times 4$systolic array but with significantly reduced area by 46.26% and power by 39.42%, and superior processing element utilization for irregular matrices. X-SA provides a approach to help improve the performance of some AI applications running on edge GPGPUs relatively in resource-constrained environments.
Yingsong Wang, Zhenzhen Jia, Ling Yang 0008, Hongbing Tan, Junsheng Chang, Junbo Tie, Libo Huang 0002
HPCC5
2025 PolyPE: An Efficient Multi-Precision Multi-Mode Floating-Point Processing Element for HPC and AI
abstract
In this paper, an efficient multi-precision multimode floating-point Processing Element is designed for HPCenabled AI workloads, called PolyPE, in which Poly means multiprecision multi-mode. It supports both conventional and mixedprecision FMA operations, including single-FMA, dual-FMA, and quad-FMA modes, as well as quad-FMA-add for enhanced throughput. The supported precisions include double precision, single precision, half precision, TF32, and BF16. At each clock cycle, the processing element can perform one double-precision, two single-precision, or four half-precision operations. Compared to existing designs, it offers broader precision support, including TF32 and BF16, with higher throughput and lower hardware overhead, achieving up to 5× improvement over standard FMA. We integrated the design into an open-source GPGPU and extended its instruction set. Experimental results show up to 2.17× performance gain, with 27.2% and 41.2% reductions in LUT and FF usage, respectively, while preserving functional equivalence.
Zhenzhen Jia, Hongbing Tan, Ling Yang 0008, Hui Guo 0004, Junsheng Chang, Yongwen Wang, Libo Huang 0002
ICCD6
2024 A survey of compute nodes with 100 TFLOPS and beyond for supercomputers
Junsheng Chang, Kai Lu 0001, Yang Guo 0003, Yongwen Wang, Libo Huang 0002, Yao Wang 0002, Biwei Zhang
CCF Trans. High Perform. Comput.1
2022 Full-credit Flow Control: A Novel Technique to Implement Deadlock-free Adaptive Routing
abstract
Deadlock-free adaptive routing is extensively adopted in interconnection networks to improve communication bandwidth and reduce latency. However, existing deadlock-free flow control schemes either underutilize memory resources due to inefficient buffer management for simple hardware implementations, or rely on complicated coordination and synchronization mechanisms with high hardware complexity. In this work, we solve the deadlock problem from a different perspective by considering the deadlock as a lack of credit. With minor modifications of the credit accumulation procedure, our proposed full-credit flow control (FFC) ensures atomic buffer usage only based on local credit status while making full use of the buffer space. FFC can be easily integrated in the industrial router to achieve deadlock freedom with less area and power consumption, but 112% higher throughput, compared to the critical bubble scheme (CBS). We further propose a credit reservation strategy to eliminate the escape virtual channel (VC) cost for fully adaptive routing implementation. The synthesizing results demonstrate that FFC along with credit reservation (FFC-CR) can reduce the area by 29% and power consumption by 26% compared with CBS.
Kai Lu 0001, Sheng Ma, Junsheng Chang
DATE4
2022 Trusted-Committee- Based Secure and Scalable BFT Consensus for Consortium Blockchain
abstract
Compared with public blockchain, consortium blockchain is more secure and controllable deployed in an enterprise scenario. Byzantine fault tolerance (BFT) consensus is widely applied in consortium blockchain. Although PBFT is the most classic practical BFT consensus with message complexity O(n2), it still faces some security threats and has low consensus efficiency. To address these issues, we propose a secure and trusted BFT (S2BFT) consensus based on trusted committees. S2BFT generates a trusted anonymous number using trust execution environment (TEE) for each server node and selects committees by pseudo-random algorithm. S2BFT can efficiently reach consensus by the committees with an O(m*n) message complexity. In addition, correctness analysis proves that S2BFT can resist more attacks than traditional BFT consensus and tolerate 1/2 byzantine server nodes. Results further demonstrate the efficiency of the simulated S2BFT implementation.
Liaoliao Feng, Yusong Tan, Xiang Fu 0002, Keming Wang, Junsheng Chang
MSN6
2022 A Distributed Graph Inference Computation Framework Based on Graph Neural Network Model
abstract
A graph is a structure that can effectively represent objects and the relationships between them.Graph Neural Networks (GNNs) enable deep learning to be applied in the graph domain.However, most GNN models are trained offline and cannot be directly used in real-time monitoring scenarios.In addition, due to the very large data scale of the graph, a single machine cannot meet the demand, and there is a performance bottleneck.Therefore, we propose a distributed graph neural network inference computing framework, which can be applied to GNN models in the form of Encoder-Decoder.We propose the idea of "single-point inference, message passing, distributed computing", which enables the system to use offline-trained GNNs for real-time inference computations on graph data.To maintain the model effect, we add the second-degree subgraph and mailbox mechanism to the continuous iterative calculation.Finally, our results on public datasets show that this method greatly improves the upper limit of inference computation and has better timeliness.And it maintains a good model effect on three types of classical tasks.The source code is published in a Github repository.
Zeting Pan, Yue Yu 0001, Junsheng Chang
SEKE3
2021 PFT: A Congestion Avoidance Method based on Proactive Flow Throttling at Endpoints
Xingyun Qi, Dezun Dong, Junsheng Chang, Jijun Cao
IM5
2021 JointCloud Cross-chain Verification Model of Decentralized Identifiers
abstract
When multiple entities communicate or collaborate in JointCloud, identities are the very prior basis to build trust with each other. Decentralized identifier (DID) can provide a trusted identity with blockchain technology and a complete method of identity verification based on verifiable credentials, which solves problems of conventional centralized identity. However, current DIDs can only conduct verification within a single blockchain, which limits the interoperability of DIDs on different blockchains. Network isolation hinders the verification of DIDs on different blockchains and thus there is a need to break the barrier between blockchains. In this paper, we propose a model to conduct cross-chain verification of DIDs. We build a system of credit evaluation to describe the credibility of DIDs in a unified way and deploy smart contracts to implement cross-chain verification of DIDs. Experimental results verifies the feasibility of the model, which realizes cross-chain verification of DIDs in the networks of blockchain.
Peichang Shi, Junsheng Chang
IPCCC3
2021 A Technical Capability Evaluation Model Based Concept and Prerequisite Relation in Computer Education(SEKEEO) (S)
abstract
Effectively assessing the results of users' online learning and enhancing social recognition has become a major development direction for online education platforms.For computer education, this article constructs a technical capability assessment model.This model integrates professional concepts in the field of computer science and extracts knowledge concepts from educational resources.The model first extracts candidate concepts, then uses a graph propagation algorithm to quantify candidate concepts and obtains concepts from them, and finally uses prerequisite relationships to further quantify the concepts mastered by students.The model combines the prerequisite relationship among concepts to quantify the skills that students have mastered.It can not only effectively evaluate the user's skill mastery but also lays a foundation for subsequent course recommendations and career recommendations for users.The model is tested in the real learning environment of 250 students.This model has been proved to own certain practicability and reliability by Kendall rank correlation coefficient, which is used as an evaluation index.
Jiwen Luo, Tao Wang 0006, Junsheng Chang
SEKE3
2020 Using Configuration Semantic Features and Machine Learning Algorithms to Predict Build Result in Cloud-Based Container Environment
abstract
Container technologies are being widely used in large scale production cloud environments, of which Docker has become the de-facto industry standard. In practice, Docker builds often break, and a large amount of efforts are put into troubleshooting broken builds. Prior studies have evaluated the rate at which builds in large organizations fail. However, there is still a lack of early warning methods for predicting the Docker build result before the build starts. This paper provides a first attempt to propose an automatic method named PDBR. It aims to use the configuration semantic features extracted by AST and the machine learning algorithms to predict build result in the cloud-based container environment. The evaluation experiments based on more than 36,000 collected Docker builds show that PDBR achieves 73.45%-91.92% in F1 and 29.72%-72.16% in AUC. We also demonstrate that different ML classifiers have significant and large effects on the PDBR AUC performance.
Yiwen Wu 0001, Yang Zhang 0026, Junsheng Chang, Bo Ding 0001, Tao Wang 0006, Huaimin Wang 0001
ICPADS3
2019 Multi-reviewing pull-requests: An exploratory study on GitHub OSS projects
Dongyang Hu, Yang Zhang 0026, Junsheng Chang, Gang Yin, Yue Yu 0001, Tao Wang 0006
Inf. Softw. Technol.3
2018 Recommending Similar Bug Reports: A Novel Approach Using Document Embedding Model
abstract
In the software development, it is not uncommon to find that several bug reports are related to many common code files, i.e., similar bugs. Similar bug recommendation is a meaningful task which can assist developers in bug triaging and fixing. As the state of the art, Yang et al.'s work presented an approach that combines TF-IDF method with word embedding model and achieved a good result. To further improve the performance of their approach, in this paper, we propose a novel approach using Document Embedding model. In our preliminary evaluation, we conduct the experiment on 13,090 bug reports from the Eclipse platform and the results show that our approach outperforms Yang et al.'s, with 7.89-8.96% of improvement.
Dongyang Hu, Tao Wang 0006, Junsheng Chang, Gang Yin, Yue Yu 0001, Yang Zhang 0026
APSEC4
2018 Multi-Discussing across Issues in GitHub: A Preliminary Study
abstract
Social coding sites like GitHub has enabled developers to easily contribute their comments on multiple issues and switch their discussion between issues, i.e., multi-discussing. Discussing multiple issues simultaneously may enhance the work efficiency of developers. However, multi-discussing also relies on developers' rationally allocating their time and focus, which may bring different influence to the resolution of issues. Therefore, investigating how multi-discussing affects the issue resolution is a meaningful research question which can help developers understand the benefits and limitations when they switch their discussion between issues. In this paper, we present a preliminary study of the impact of multi-discussing on issue resolution in GitHub projects, by using quantitative methods. First, we collect and analyzed data from 631 GitHub projects to explore how multi-discussing affects the average resolution latency of project issues. Further, we develop method for measuring the rate and breadth of a developers' discussionswitching behavior, and we use regression modeling to study how discussion-switching affects the single issue resolution latency. We find that multi-discussing is a common behavior of developers in GitHub projects. Also, multi-discussing is associated with shorter average issue resolution latency of project. However, during a single issue resolution, more participants' discussion-switching tend to bring longer issue resolution latency. Our study motivates the need for further research on the multi-discussing.
Dongyang Hu, Tao Wang 0006, Junsheng Chang, Gang Yin, Yang Zhang 0026
APSEC3
2014 A Low Overhead Last-Write-Touch Prediction Scheme
abstract
Last-write-touch prediction can reduce cache-to-cache transfer latency by converting 3-hop misses into 2-hop misses in directory-based shared-memory multiprocessors. By predicting a last-write-touch and self-downgrading a cache block in advance, a processor can get the data from the memory directly and the coherence overhead is significantly reduced. In this paper, we propose a new low overhead last-write-touch prediction scheme that exploits the inherent write burst characteristics of programs. The scheme uses write burst numbers to compute history traces and generate signatures. Compared with the existing instruction-based prediction technique, much storage overhead can be reduced. The experimental results show that our last-write-touch prediction scheme can achieve almost the same prediction accuracy as the instruction-based prediction scheme with the storage overheads of the history table reduced by 69% and the storage overheads of the signature table reduced by 36%.
Zhengbin Pang, Junsheng Chang
DASC5
2014 An incentive compatible reputation mechanism for P2P systems
Junsheng Chang, Zhengbin Pang, Huaimin Wang 0001, Gang Yin
J. Supercomput.1
2007 A New Reputation Mechanism Against Dishonest Recommendations in P2P Systems
Junsheng Chang, Huaimin Wang 0001, Gang Yin, Yang-Bin Tang
WISE1