EDBT 2026 Demo / reviewers in the wild / expert
Chunye Gong
dblp:79/7926
· DBLP profile ↗
39ranked-venue papers
3as first author
29since 2021 · last 2026
0000-0003-0349-1100ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 2 first-author · 15 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Security and privacy · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive density clustering for data-driven password mangling rule generation
Yongtao Luo, Chunye Gong, Jie Liu 0002, Tun Li 0002 |
Comput. Secur. | 2 |
| 2026 | Flux-conserved physics-informed neural networks for electromagnetic scattering computation
Chenyu Peng, Tiaojie Xiao, Sifan Wang, Xinhai Chen 0001, Chunye Gong |
Eng. Appl. Artif. Intell. | 8 |
| 2026 | GNNRL-smoothing: A prior-free reinforcement learning model for mesh optimization
Xinhai Chen 0001, Chunye Gong, Bo Yang 0023, Liang Deng, Yufei Pang, Xiang Zhang 0008, Jie Liu 0002 |
Neural Networks | 3 |
| 2026 | ST-FlowNet: A lightweight framework for long-term spatio-temporal flow field prediction
Qisong Xiao, Xinhai Chen 0001, Haijian Yang, Chunye Gong, Jie Liu 0002 |
Neural Networks | 4 |
| 2025 | Physics-Guided Multimodal Neural Networks for Big Data - Driven Magnetic Component Design
Jin Zhang 0018, Cong Yao, Wengen Li, Qiyou Xie, Qiuzhen Wan, Chunye Gong |
IEEE Big Data | 6 |
| 2025 | A Parallel Implementation of ChaCha20 on MT-3000 Heterogeneous Multi-zone Processor
Yongtao Luo, Jie Liu 0002, Tun Li 0002, Chunye Gong |
ICA3PP (1) | 4 |
| 2025 | AccuGraph: Memory-Efficient Full-Graph GNN Training on a Single GPU via Subgraph Accumulation
Liyang Wu, Menghan Jia, Yahui Wu, Gongqingjian Jiang, Jiezhong He, Chunye Gong, Yinghui Gao |
ICA3PP (3) | 7 |
| 2025 | YH-Light: Yielding Hierarchy-aware Partitioner for Large-scale Graph ProcessingabstractLarge-scale graph tasks often have to be conducted in parallel on partitioned graphs.However, current partitioning methods, designed for smaller-scale clusters with a limited number of computing nodes, struggle to scale effectively due to their inability to handle messages transferred through traditional partitioning grids.We present YH-Light, an hierarchy-aware partitioning engine on Tianhe supercomputers to minimize communication.The key idea of YH-Light is to take advantage of the hierarchical communication topology to perform a graph partition based on communication hierarchies, where scattered messages are (i) clustered with hub vertices according to the organization of the computing nodes and (ii) grouped and then exchanged messages according to hierarchical communication domains.We demonstrate YH-Light's effectiveness with synthetic benchmarks and real-world graphs.In particular, the YH-Light-based Graph 500 tests on the Tianhe supercomputer outperform the leading systems in the latest Graph 500 list.Furthermore, YH-Light significantly advances graph processing, surpassing the current state-of-the-art graph partitioning engines and graph systems by orders of magnitude. Xinbiao Gan, Chunye Gong, Jie Liu 0002, Kai Lu 0001 |
ICS | 3 |
| 2025 | Informative Discrimination Network for Efficient Single Image Super-ResolutionabstractDeploying convolutional neural networks on low-resource mobile devices for single image super-resolution (SISR) faces the issue of how to balance the parameter amount and performance. The default solution is simultaneously condensing both hierarchical representation and attention features into their respective light proxies. The insight underlying this solution lies in the fact that features are redundant since the super-resolution needs plenty of similar pixels. This work takes it to the next step from the viewpoint of informativeness and discrimination. In detail, we propose an informative disrcimination network (IDNet) for SISR. For informativeness, a multi-scale residual block (MRB) is explored to capture informative spatial details via the scale-in-scale structure. It mines rich intra-layer spatial details based on inter-layer ones of the default hierarchical representation. However, it also incurs feature redundancy. Though attention serves to reduce this redundancy, feature discrimination and pixel-wise structural preservation cannot be guaranteed. Here spatial discrimination attention behaves like the biased discriminant classifier to induce spatial discrimination, while the nuclear-norm regularization recovers the image low-rank structure to reduce artifacts or noises. Importantly, no extra network weights are introduced for model efficiency. Experiments show that IDNet delivers sound performance with fewer parameters, as compared to its cousins. Yuzheng Tu, Xinhai Chen 0001, Chunye Gong, Jie Liu 0002, Bo Yang 0023, Xiang Gao 0020, Xiang Zhang 0008 |
IJCNN | 3 |
| 2025 | GraphWorld: Ultra-fast Graph Engine for World-Wide Web SearchingabstractGraph has recently enabled substantial advances to the Web. Processing worldwide graphs with millions to billions, even trillions of edges in large-scale high-performance systems is pressing, but current graph processing engines are designed for small-scale graph processing beyond a few tens of computing nodes and are unable to scale well to large parallel systems because they are oblivious to imbalanced communication across the communication grid. Therefore, we present GraphWorld, a better approach to optimizing graph search in large parallel systems for world-wide web crawling and indexing.GraphWorld (i) features a new graph partitioning method to achieve better load balancing and minimize communication overhead across the row and column directions; (ii) designs an efficient hardware prefetching and caching mechanism that can gather, traverse, and scatter pipeline vertices to accelerate graph processing; and (iii) proposes υBFS: vectorization-based BFS for leveraging vectorization units equipped in modern high performance processors to further improve graph search.In addition, we used real-world graphs and benchmarks to demonstrate the effectiveness of GraphWorld. In particular, the GraphWorld-based Graph 500 tests on the Tianhe supercomputer are superior to the fastest systems in the latest Graph 500 lists. We finally apply GraphWorld to real-life graphs for the worldwide search of the Web, which outperforms the state-of-the-art graph partitioning and graph system by orders of magnitude. Xinbiao Gan, Qiang Zhang 0053, Chunye Gong, Kai Lu 0001 |
ACM Multimedia | 4 |
| 2025 | TianheEngine: Hierarchy-aware Adaptive Partitioning System for Trillion-scale Graph ProcessingabstractGraph partitioning is essential for effectively managing multi-trillion-edge graphs in distributed computing systems, particularly those spanning hierarchical architectures with thousands of computing nodes. Traditional partitioning strategies neglect hierarchical communication variances across modern high-performance computing (HPC) systems, leading to prohibitive cross overhead when processing trillion-scale graphs. We propose TianheEngine, a hierarchy-aware adaptive partitioning system that takes advantage of the communication hierarchy of the underlying distributed computing system and the sparsity characteristics of the input graphs to improve communication efficiency. We evaluated TianheEngine on fundamental graph operations using both synthetic and real-world datasets. Our extensive experiments use up to 79,024 computing nodes and over 1.2 million processor cores. Experimental results show that TianheEngine is superior to state-of-the-art graph partitioning methods and parallel graph systems and outperforms top-ranked systems on the latest Graph 500 list. Xinbiao Gan, Yiqi Wang 0001, Qiang Zhang 0053, Yongming Yi, Chunye Gong, Jie Liu 0002, Kai Lu 0001 |
SC | 6 |
| 2025 | GraphCSR: A Space and Time-Efficient Sparse Matrix Representation for Web-scale Graph ProcessingabstractGraph data processing is essential for web-scale applications, including social networks, recommendation systems, and web of things (WoT) systems, where large, sparsely connected graphs dominate. Traditional sparse matrix storage formats like compressed sparse row (CSR) face significant memory and performance bottlenecks in distributed, federated, and edge-based computing environments, which are increasingly central to the web. To address this challenge, we propose GraphCSR, a novel storage format that clusters vertices with identical edge degrees and stores only the starting index of each group. This approach minimizes memory overhead and facilitates batch memory access while enhancing overall performance, making it particularly suitable for federated systems and resource-constrained edge nodes. Our experiments across various graph operations and large datasets show that GraphCSR achieves considerable memory savings and performance gains of large-scale, distributed graph processing. When deployed GraphCSR on two production-grade supercomputers, demonstrating its potential for scaling web and WoT graph processing in large-scale distributed computing systems. Xinbiao Gan, Qiang Zhang 0053, Bo Yang 0023, Chunye Gong, Jie Liu 0002, Kai Lu 0001 |
WWW | 6 |
| 2025 | Fine-grained vectorized merge sorting on RISC-V: from register to cacheabstractAbstract Merge sort as a divide-sort-merge paradigm has been widely applied in computer science fields. As modern reduced instruction set computing architectures like the fifth generation (RISC-V) regard multiple registers as a vector register group for wide instruction parallelism, optimizing merge sort with this vectorized property is becoming increasingly common. In this paper, we overhaul the divide-sort-merge paradigm, from its register-level sort to the cache-aware merge, to develop a fine-grained RISC-V vectorized merge sort (RVMS). From the register-level view, the inline vectorized transpose instruction is missed in RISC-V, so implementing it efficiently is non-trivial. Besides, the vectorized comparisons do not always work well in the merging networks. Both issues primarily stem from the expensive data shuffle instruction. To bypass it, RVMS strides to take register data as the proxy of data shuffle to accelerate the transpose operation, and meanwhile replaces vectorized comparisons with scalar cousin for more light real value swap. On the other hand, as cache-aware merge makes larger data merge in the cache, most merge schemes have two drawbacks: the in-cache merge usually has low cache utilization, while the out-of-cache merging network remains an ineffectively symmetric structure. To this end, we propose the half-merge scheme to employ the auxiliary space of in-place merge to halve the footprint of naïve merge sort, and meanwhile copy one sequence to this space to avoid the former data exchange. Furthermore, an asymmetric merging network is developed to adapt to two different input sizes. Experiments on the RISC-V processor SG2042 show that four fine-grained optimization schemes including register strided transpose, hybrid merging network, half-merge strategy, and asymmetric merging network, improve performance by 4.05%, 19.88%, 12.23%, and 11.04% respectively. Importantly, the overall performance is 1.34x faster than the parallel sorting in the Boost C++ library, and 1.85x faster than std::sort. Jin Zhang 0018, Jincheng Zhou, Xiang Zhang 0008, Chunye Gong |
CCF Trans. High Perform. Comput. | 5 |
| 2025 | FMCC-RT: a scalable and fine-grained all-reduce algorithm for large-scale SMP clusters
Jintao Peng, Jie Liu 0002, Jianbin Fang, Zhiquan Lai, Bo Yang 0023, Chunye Gong, Xinjun Mao, Guo Mao, Jie Ren 0007 |
Sci. China Inf. Sci. | 8 |
| 2025 | Efficient adaptive Cartesian mesh generation for complex boundary representation models
Xiang Gao 0020, Qingyang Zhang 0009, Chunye Gong, Chao Li 0002, Jie Liu 0002 |
Graph. Model. | 3 |
| 2025 | GraphCSR: A Degree-Equalized CSR Format for Large-scale Graph ProcessingabstractGraph processing underpins a vast array of data-centric applications, serving as a crucial component in fields such as social network analysis, recommendation systems, bio-informatics, and search engines. As graph data grows in scale and complexity, high-performance graph processing is increasingly essential. Many graph processing tasks depend on efficient data structures to manage the sparsity typical of real-world graphs, where most vertices have limited connectivity. This sparsity poses challenges for memory and computational efficiency in large-scale graph processing, and conventional sparse formats like Compressed Sparse Row (CSR) often struggle with memory and computation inefficiencies when handling massive graphs. To address these challenges, we introduce GraphCSR, a degree-equalized CSR format specifically tailored to enhance the spatio-temporal efficiency of distributed graph processing across various tasks. GraphCSR aggregates low-degree vertices into synthetic high-degree ones and applies group-wise compression to reduce storage overhead by recording only the starting index for each aggregated group. This reduces memory usage and supports batch-memory access to improve performance. Our extensive evaluations in various graph processing algorithms and datasets demonstrate that GraphCSR not only reduces the memory footprint required for large-scale graphs, but also improves performance across multiple types of graph processing tasks, outperforming popular sparse storage formats. Furthermore, when deployed on a production-scale supercomputer with 79,024 nodes, GraphCSR achieved a graph processing throughput that exceeded the top-ranked system on the Graph500 benchmark. Xinbiao Gan, Chunye Gong, Dezun Dong, Jie Liu 0002, Kai Lu 0001 |
Proc. VLDB Endow. | 3 |
| 2025 | TransCL: An Automatic CUDA-to-OpenCL Programs Transformation FrameworkabstractWith the rising demand for computational power and the increasing variety of computational scenarios, considerable interest has emerged in transforming existing CUDA programs into more general-purpose OpenCL programs, enabling them to run across diverse hardware platforms. However, manual methods, typically designed for specific applications, lack flexibility. Current automated conversion techniques also face considerable challenges, particularly in handling diverse programming interfaces, memory management, and so on, and are insufficient for converting large-scale, complex CUDA projects. In this article, we propose a novel source-to-source program transformation framework, TransCL, which automates the conversion of CUDA programs in four key aspects: source code, execution model, programming model, and memory model. To achieve this, we abstract a set of conversion rules aligned with the latest CUDA standards, develop a transcoder, implement an OpenCL-compatible programming interface library, and establish a memory mapping mechanism between CUDA and OpenCL. Experiments demonstrate that TransCL provides a high level of automation in converting CUDA-based applications and is effective in handling large, complex projects such as TensorFlow. Moreover, the converted AI framework successfully conducted model training for the first time. The experiment also validates that the converted program can execute correctly across multiple platforms and demonstrate good performance. Changqing Shi, Chunye Gong, Yicheng Sui, Yutong Jin |
ACM Trans. Archit. Code Optim. | 6 |
| 2025 | Deep Learning Operators Performance Tuning for Changeable Sized Input Data on Tensor Accelerate HardwareabstractThe operator library is the fundamental infrastructure of deep learning acceleration hardware. Automatically generating the library and tuning its performance is promising because the manual development by well-trained and skillful programmers is costly in terms of both time and money. Tensor hardware has the best computing efficiency for deep learning applications, but the operator library programs are hard to tune because the tensor hardware primitives have many limitations. Otherwise, the performance is difficult to be fully explored. The recent advancement in LLM exacerbates this problem because the size of input data is not fixed. Therefore, mapping the computing tasks of operators to tensor hardware units is a significant challenge when the shape of the input tensor is unknown before the runtime. We propose DSAT, a deep learning operator performance autotuning technique for changeable-sized input data on tensor hardware. To match the input tensor's undetermined shape, we choose a group of abstract computing units as the basic building blocks of operators for changeable-sized input tensor shapes. We design a group of programming tuning rules to construct a large exploration space of the variant implementation of the operator programs. Based on these rules, we construct an intermediate representation of computing and memory access to describe the computing process and use it to map the abstract computing units to tensor primitives. To speed up the tuning process, we narrow down the optimization space by predicting the actual hardware resource requirement and providing an optimized cost model for performance prediction. DSAT achieves performance comparable to the vendor's manually tuned operator libraries. Compared to state-of-the-art deep learning compilers, it improves the performance of inference by 13% on average and decreases the tuning time by an order of magnitude. Pengyu Mu, Yi Liu 0013, Rui Wang 0014, Hangcheng An, Qianhe Zhao, Hailong Yang 0002, Chenhao Xie 0001, Zhongzhi Luan, Chunye Gong, Depei Qian 0001 |
IEEE Trans. Computers | 10 |
| 2025 | An efficient heterogeneous parallel password recovery system on MT-3000
Yongtao Luo, Jie Liu 0002, Chunye Gong, Tun Li 0002 |
J. Supercomput. | 3 |
| 2024 | A Hybrid Vectorized Merge Sort on ARM NEON
Jincheng Zhou, Jin Zhang 0018, Xiang Zhang 0008, Tiaojie Xiao, Chunye Gong |
ICA3PP (6) | 6 |
| 2024 | GraphCube: Interconnection Hierarchy-aware Graph ProcessingabstractProcessing large-scale graphs with billions to trillions of edges requires efficiently utilizing parallel systems. However, current graph processing engines do not scale well beyond a few tens of computing nodes because they are oblivious to the communication cost variations across the interconnection hierarchy. We introduce GraphCube, a better approach to optimizing graph processing on large-scale parallel systems with complex interconnections. GraphCube features a new graph partitioning approach to achieve better load balancing and minimize communication overhead across multiple levels of the interconnection hierarchy. We evaluate GraphCube by applying it to fundamental graph operations performed on synthetic and real-world graph datasets. Our evaluation used up to 79,024 computing nodes and 1.2+ million processor cores. Our large-scale experiments show that GraphCube outperforms state-of-the-art parallel graph processing methods in throughput and scalability. Furthermore, GraphCube outperformed the top-ranked systems on the Graph 500 list. Xinbiao Gan, Shenghao Qiu, Jiaqi Si, Jianbin Fang, Dezun Dong, Chunye Gong, Zheng Wang 0001 |
PPoPP | 8 |
| 2024 | Transplantation and optimization of molecular dynamics simulation on MT-3000
Jianjiang Li, Hongyaoxing Gu, Lin Qiao, Chunye Gong |
Future Gener. Comput. Syst. | 5 |
| 2024 | MST: Topology-Aware Message Aggregation for Exascale Graph Processing of Traversal-Centric AlgorithmsabstractThis article presents MST, a communication-efficient message library for fast graph traversal on exascale clusters. The key idea is to follow the multi-level network topology to perform topology-aware message aggregation, where small messages are gathered and scattered at each level of domain. To facilitate message aggregation, we equip MST with flexible buffer management including active buffer switching and dynamic buffer expansion. We implement MST on the newest-generation Tianhe supercomputer and evaluated its performance using various traversal-centric algorithms on both synthetic trillion-scale graphs and real-world big graphs. The results show that MST-based graph traversal is orders of magnitude faster than that based on Active Messages Library (AML). For the Graph500-BFS benchmark, MST-based Tianhe (with 77.2 K nodes) outperforms the Fugaku supercomputer (with 148.5 K nodes) by 18.53%, while Fugaku is ranked No. 1 in the latest Graph500-BFS ranking (June 2023). MST also greatly improves graph processing performance on other commercial large-scale computing systems at the National Supercomputing Center in Changsha (NSCC) and WuzhenLight. Xinbiao Gan, Bo Yang 0023, Xinhai Chen 0001, Chunye Gong, Shijie Li 0002, Kai Lu 0001, Qiao Li 0001, Yiming Zhang 0003 |
ACM Trans. Archit. Code Optim. | 6 |
| 2023 | FT-topo: Architecture-Driven Folded-Triangle Partitioning for Communication-efficient Graph ProcessingabstractAs graph size (numbers of vertices and edges) is increasing from billions to trillions, efficient graph processing requires exascale computing clusters, which consist of hundreds of thousands of nodes connected via hierarchical networks with multiple levels of communication domains, e.g., multilevel triangle communication domains. While the computation of traversal-centric graph algorithms is relatively simple (e.g., status check), communication is the bottleneck due to the transfer of numerous small messages among hierarchical triangle communication domains. Xinbiao Gan, Ruigeng Zeng, Jiaqi Si, Ji Liu 0003, Daxiang Dong, Chunye Gong, Cong Liu 0047 |
ICS | 7 |
| 2023 | MT-office: parallel password recovery program for office on domestic heterogeneous multi-core processor
Yongtao Luo, Bo Yang 0023, Jie Liu 0002, Ruibo Wang, Jinmin Wen, Tiaojie Xiao, Xuguang Chen, Chunye Gong |
CCF Trans. High Perform. Comput. | 8 |
| 2022 | A novel neural network approach for airfoil mesh quality evaluation
Xinhai Chen 0001, Chunye Gong, Jie Liu 0002, Yufei Pang, Liang Deng, Lihua Chi, Kenli Li 0001 |
J. Parallel Distributed Comput. | 2 |
| 2022 | A Scalable Parallel Algorithm for 3-D Magnetotelluric Finite Element Modeling in Anisotropic Mediaabstract3-D magnetotelluric (MT) forward modeling has always been faced with the problems of high memory requirements and long computing time. In this article, we design a scalable parallel algorithm for 3-D MT finite element modeling in anisotropic media. The parallel algorithm is based on the distributed mesh storage, including multiple parallel granularities, and is implemented through multiple tools. Message-passing interface (MPI) is used to exploit process parallelisms for subdomains, frequencies, and solving equations. Thread parallelisms for merge sorting, element analysis, matrix assembly, and imposing Dirichlet boundary conditions are developed by Open Multi-Processing (OpenMP). We validate the algorithm through several model simulations and study the effects of topography and conductivity anisotropy on apparent resistivities and phase responses. Scalability tests are performed on the Tianhe-2 supercomputer to analyze the parallel performance of different parallel granularities. Three parallel direct solvers Supernodal LU (SUPERLU), MUltifrontal Massively Parallel sparse direct Solver (MUMPS), and Parallel Sparse matriX package (PASTIX) are compared in solving sparse systems of equations. As a result, reasonable parallel parameters are suggested for practical applications. The developed parallel algorithm is proven to be efficient and scalable. Xiaoxiong Zhu, Jie Liu 0002, Yi-an Cui, Chunye Gong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | An efficient image to column algorithm for convolutional neural networksabstractConvolutional Neural Networks (CNNs) are a class of deep neural networks. The image to column (im2col) procedure is an important step for CNN and consumes about 28.8% of the whole inference time. In this paper, we present an efficient im2col algorithm, name im2cole (word “e” means efficient). The condition with different stride and pad in im2cole is well handled and the judgements in the innermost loop are removed. The procedure with pad = 1 is split into three conditions. This will reduce the pause of CPU instruction pipeline. The performances of the presented im2cole algorithm are reported with different inputs. Some discussion and performance issues are also reported. The experimental results show that the overall performance speedup of im2cole ranges from 2.12 to 4.33 compared with the original algorithm. The real application with Darknet shows that im2cole can get 20.75% whole performance improvement. Chunye Gong, Xinhai Chen 0001, Shuling Lv, Jie Liu 0002, Bo Yang 0023, Weimin Bao, Yufei Pang |
IJCNN | 1 |
| 2021 | MVE-Net: An Automatic 3-D Structured Mesh Validity Evaluation Framework Using Deep Neural Networks
Xinhai Chen 0001, Jie Liu 0002, Chunye Gong, Shengguo Li, Yufei Pang |
Comput. Aided Des. | 3 |
| 2020 | OHTMA: an optimized heuristic topology-aware mapping algorithm on the Tianhe-3 exascale supercomputer prototypeabstractWith the rapid increase of the size of applications and the complexity of the supercomputer architecture, topology-aware process mapping becomes increasingly important. High communication cost has become a dominant constraint of the performance of applications running on the supercomputer. To avoid a bad mapping strategy which can lead to terrible communication performance, we propose an optimized heuristic topology-aware mapping algorithm (OHTMA). The algorithm attempts to minimize the hop-byte metric that we use to measure the mapping results. OHTMA incorporates a new greedy heuristic method and pair-exchange-based optimization. It reduces the number of long-distance communications and effectively enhances the locality of the communication. Experimental results on the Tianhe-3 exascale supercomputer prototype indicate that OHTMA can significantly reduce the communication costs. Yishui Li, Xinhai Chen 0001, Jie Liu 0002, Bo Yang 0023, Chunye Gong, Xinbiao Gan, Shengguo Li, Han Xu 0008 |
Frontiers Inf. Technol. Electron. Eng. | 5 |
| 2020 | VBSF: a new storage format for SIMD sparse matrix-vector multiplication on modern processors
Yishui Li, Peizhen Xie, Xinhai Chen 0001, Jie Liu 0002, Bo Yang 0023, Shengguo Li, Chunye Gong, Xinbiao Gan, Han Xu 0008 |
J. Supercomput. | 7 |
| 2019 | Parallel convolution algorithm using implicit matrix multiplication on multi-core CPUsabstractConvolution neural networks (CNNs) have been extensively used in machine learning applications. The most time-consuming part of CNNs are convolution operations. A common approach to implementing convolution operations is to recast them as general matrix multiplication, known as the im2col+GEMM approach. There are two main drawbacks of this approach. One is that large additional memory space is required. The other is the packing on the input elements of convolution operations are not memory-efficient enough. In this paper, we present a new parallel convolution algorithm using implicit matrix multiplication on multi-core CPUs. In comparison with Im2col+GEMM, our new algorithm can reduce the memory footprints and improve the packing efficiency. The experiment results on two ARV8-based multi-core CPUs demonstrate that our new algorithm gives much better performance and scalability than the im2col+GEMM method in most cases. Songzhu Mei, Jie Liu 0002, Chunye Gong |
IJCNN | 4 |
| 2019 | Succinct Representations in Collaborative Filtering: A Case Study using Wavelet Tree on 1, 000 CoresabstractUser-Item (U-I) matrix has been used as the dominant data infrastructure of Collaborative Filtering (CF). To reduce space consumption in runtime and storage, caused by data sparsity and growing need to accommodate side information in CF design, one needs to go beyond the U-I Matrix. In this paper, we took a case study of Succinct Representations in Collaborative Filtering, rather than using a U-I Matrix. Our key insight is to introduce Succinct Data Structures as a new infrastructure of CF. Towards this, we implemented a User-based K-Nearest-Neighbor CF prototype via Wavelet Tree, by first designing a Accessible Compressed Documents (ACD) to compress U-I data in Wavelet Tree, which is efficient in both storage and runtime. Then, we showed that ACD can be applied to develop an efficient intersection algorithm without decompression, by taking advantage of ACD's characteristics. We evaluated our design on 1,000 cores of Tianhe-II supercomputer, with one of the largest public data set ml-20m. The results showed that our prototype could achieve 3.7 minutes on average to deliver the results. Xiangjun Peng, Qingfeng Wang 0002, Xu Sun 0002, Chunye Gong |
PDCAT | 4 |
| 2018 | Customizing the HPL for China accelerator
Xinbiao Gan, Yikun Hu 0001, Jie Liu 0002, Lihua Chi, Han Xu 0008, Chunye Gong, Shengguo Li, Yihui Yan |
Sci. China Inf. Sci. | 6 |
| 2018 | An efficient SIMD compression format for sparse matrix-vector multiplicationabstractSummary Sparse matrix‐vector multiplication (SpMV) is an essential kernel in sparse linear algebra and has been studied extensively on all modern processor and accelerator architectures. Compressed Sparse Row (CSR) is a frequently used format for sparse matrices storage. However, CSR‐based SpMV has poor performance on processors with vector units. In order to take full advantage of SIMD acceleration technology in SpMV, we proposed a new matrix storage format called CSR‐SIMD. The new storage format compresses the non‐zero elements into many variable‐length data fragments with consecutive memory access addresses. Thus, the data locality of sparse matrix A and dense vector x expands and the floating‐point operations for each fragment can be completely calculated by vectorized implementation on wide SIMD units. Our experimental results indicate that CSR‐SIMD has better storage efficiency and low‐overhead for format conversion. Besides, the new format achieves high scalability on wide SIMD units. In comparison with the CSR‐based and BCSR‐based SpMV, CSR‐SIMD obtains better performance on FT1500A, Intel Xeon, and Intel Xeon Phi. Xinhai Chen 0001, Peizhen Xie, Lihua Chi, Jie Liu 0002, Chunye Gong |
Concurr. Comput. Pract. Exp. | 5 |
| 2014 | An efficient parallel solution for Caputo fractional reaction-diffusion equation
Chunye Gong, Weimin Bao, Guojian Tang, Bo Yang 0023, Jie Liu 0002 |
J. Supercomput. | 1 |
| 2012 | High-Performance Matrix Multiply on a Massively Multithreaded Fiteng1000 Processor
Jie Liu 0002, Lihua Chi, Chunye Gong, Han Xu 0008, Yihui Yan, Qingfeng Hu |
ICA3PP (2) | 3 |
| 2010 | Optimizing Sweep3D for Graphic Processor Unit
Chunye Gong, Jie Liu 0002, Zhenghu Gong |
ICA3PP (1) | 1 |
| 2009 | Differential Fault Analysis on SHACAL-1abstractSHACAL-1, known as one of the finalists of the NESSIE project, originates from the compression component of the widely used hash function SHA-1. The requirements of confusion and diffusion are implemented through mixing operations and rotations other than substitution and permutation, thus there exists little literature on its immunity against fault attacks. In this paper, we apply differential fault analysis on SHACAL-1 in a synthetic approach. We introduce the random word fault model, present some theoretical arguments, and give an efficient fault attack based on the characteristic of the cipher. Both theoretical predications and experimental results demonstrate that, 72 random faults are needed to obtain 512 bits key with successful probability more than 60%, while 120 random faults are enough to obtain 512 bits key with successful probability more than 99%. Ruilin Li 0002, Chao Li 0002, Chunye Gong |
FDTC | 3 |