Ke-Kun Hu

dblp:197/9852 · also Kekun Hu · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0003-3433-9566ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Vertex Partitioning Algorithm for Large-Scale Uncertain Graphs
abstract
ABSTRACT With the exponential growth of graph‐structured data, single‐machine efficient analysis has become increasingly impractical, making high‐performance distributed graph computing systems indispensable. The efficacy of these systems hinges critically on high‐quality graph partitioning. The edges of many graphs stemmed from real applications are uncertain, but many existing graph partitioning algorithms are only for deterministic graphs without considering uncertainty. This paper presents a novel partitioning algorithm, PAUG (Partitioning Algorithm for Uncertain Graphs), tailored for uncertain graphs. First, it formalizes the partitioning problem as an optimization task to minimize the cut‐edge ratio while balancing load. Second, it introduces probabilistic similarity to quantify vertex relationships under uncertainty. Finally, it details the PAUG algorithm which consists of initial partition phase and score‐function‐guided refinement strategy. Experimental results shows that PAUG achieves an average 23.2% reduction in cut‐edge ratio and a 26.2% improvement in load balance over state‐of‐the‐art algorithms.
Huanqing Cui, Anfu Chang, Jinbin Zhu, Ruixia Liu, Ke-Kun Hu
Concurr. Comput. Pract. Exp.5
2026 Memory access optimization for the dynamics EVP model of the sea ice model on the SW39000 on-chip heterogeneous many-core processor
abstract
To improve the performance of the sea ice model in the Community Earth System Model (CESM) under a heterogeneous computing environment, this work conducts an in-depth study on the memory access optimization of the Elastic-Viscous-Plastic (EVP) dynamics model in Community Ice Code(CICE) on the SW39000 heterogeneous many-core processor, which is deployed in the new-generation Sunway supercomputer. The processor’s complex on-chip heterogeneous architecture, multi-level memory hierarchy, and unique inter-core communication mechanism present significant challenges for the parallel optimization of the sea ice dynamics simulation. To address the inefficiencies caused by diverse data access patterns, a differentiated processing strategy based on data read/write characteristics is proposed to reduce unnecessary data transfers. In addition, to alleviate load imbalance arising from the sparsity of sea ice boundary update data, a local dynamic compression method incorporating the probability density of data sparsity is designed. This method dynamically compresses data according to the probability density of Direct Memory Access (DMA) data transfers, thereby reducing communication volume and balancing the workload across slave cores. Finally, to enhance the computational intensity of the slave cores and reduce data dependencies between master and slave cores, an operator fusion algorithm based on Remote Memory Access (RMA) communication is introduced to achieve efficient data caching and transmission between operators. Experimental results demonstrate that, under the standard gx3 grid configuration, the optimized EVP model achieves a 27.54×speedup over the serial version running on a single master core when executed with a single core group. Multi-core group parallel tests validate the excellent scalability of the proposed optimization strategies, achieving up to a 123.93×speedup with a 10-core group, while also exhibiting effective load balancing in terms of both clock cycles and instruction counts across the slave-core array.
Jianzhi Yu, Jianguo Liang, You Fu, Ke-Kun Hu
Future Gener. Comput. Syst.5
2026 SDGraph: A scalable training system for GNNs with GPU sampling and parallel feature access
Jianzhi Yu, You Fu, Ke-Kun Hu, Jianguo Liang
Future Gener. Comput. Syst.4
2025 FrigateBird: Decoupling Metadata/Data Services for Continuously Fast Object Storage
abstract
Currently, ride-hailing service has tens of millions of registered drivers and hundreds of millions of registered passengers, serving tens of millions of rides per day. Different from social media applications like Facebook and LinkedIn, ride hailing needs to read/write/query a large number of small objects (like photos and audio/video pieces)always fast, so that it can support critical online computations such as face comparison and sentiment analysis on audio/video records. This is of particular importance for ride-hailing service to recognize and avoid potential dangers. Existing object stores (like Haystack and Tectonic) usually store object data in files and place object metadata in a separate key-value store (like RocksDB), which is unsuitable for ride-hailing service mainly because the objects' file-related information is placed together with the object data. This severely affects the I/O performance of object storage: first, for crash consistency, the writes of object data and object metadata must be conducted inseparatephases of one transaction, which significantly increases I/O latency; second, the space of deleted objects needs to be reclaimed viacompaction, which could sharply lower I/O performance when the system is busy in serving normal read/write requests. This paper describes FrigateBird, a continuously fast object store for ride-hailing service. FrigateBird differs from existing object stores in three aspects. First, we present a metadata/data decoupled service architecture for object storage, where therichmetadata service realizes efficient queries and updates of object metadata, and therawdata service purely performs disk I/O to read/write object data from/to raw disks. Second, we propose a rich metadata structure (calledRichmeta) taking the write operation logs as part of object metadata, which allows FrigateBird tosimultaneouslywrite the object data to raw disks (without filesystem overhead) and write the metadata to a key-value store, guaranteeing crash consistency by checking whether the transaction is completed and rolling back if not. Third, we design a compaction-free deletion mechanism which can efficiently delete an object by only updating the metadata without involving the data service, so that FrigateBird can efficiently support ride-hailing service's frequent delete operations while avoiding data-migration-caused performance hiccup. Evaluation shows that FrigateBird outperforms the state-of-the-art object stores by up to$3.36\times$and$27.9\times$in the mean I/O latency for normal and in-compaction scenarios, respectively.
Yiming Zhang 0003, Ke-Kun Hu, Gang Dong, RenGang Li
IEEE Trans. Serv. Comput.2
2023 Hierarchically stacked graph convolution for emotion recognition in conversation
abstract
Accurate emotion recognition can drive the robot to understand human affection intentions precisely and deliver the emotional response when communicating with a person. Recently, graph structure has been applied to explicitly capture the self and inter-dependencies of speakers in the conversation. However, the performance of the method is limited by inadequate discriminative information extraction based on naive graph convolution. In this paper, we propose a novel Hierarchically Stacked Graph Convolution Framework (HSGCF), which leverages hierarchical structure to extract emotional discriminative features. The proposed HSGCF uses five graph convolution layers connected hierarchically to establish a more discriminative emotional feature extractor. More importantly, to mitigate the over-smooth problem caused by deeper networks, Transformer structures with residual connection are introduced into HSGCF. Experimental results on the IEMOCAP benchmark dataset indicate the proposed framework achieves a 4.12% improvement in accuracy and a 4.80% improvement in F1 score compared with the baseline method.
Binqiang Wang, Gang Dong, Yaqian Zhao, RenGang Li, Qichun Cao, Ke-Kun Hu, Dongdong Jiang
Knowl. Based Syst.6
2019 Placing big graph into cloud for parallel processing with a two-phase community-aware approach
Ke-Kun Hu, Guosun Zeng
Future Gener. Comput. Syst.1
2018 Partitioning big graph with respect to arbitrary proportions in a streaming manner
Ke-Kun Hu, Guosun Zeng, Huo-wen Jiang, Wei Wang 0033
Future Gener. Comput. Syst.1