Ke-shi Ge

dblp:204/0971 · also Keshi Ge · DBLP profile ↗
← Back
22ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0002-0669-6892ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 2 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 AdaCheck: An Adaptive Checkpointing System for Efficient LLM Training with Redundancy Utilization
Zhiquan Lai, Ke-shi Ge, Qiaoling Chen, Peng Sun 0006, Dongsheng Li 0001, Kai Lu 0001
FAST4
2026 Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-Scale MoE Models
abstract
The size of deep learning models has been increasing to enhance model quality. The linear increase in training computation budgets with model size means that training an extremely large-scale model is exceedingly time-consuming. Recently, the Mixture of Experts (MoE) has drawn significant attention as it can scale models to extra-large sizes with a near-stable computation budget. However, inefficient distributed training of large-scale MoE models hinders their broader application. Specifically, a considerable dynamic load imbalance occurs among devices during training, significantly reducing throughput. Several load-balancing works have been proposed to address the challenge. System-level solutions draw more attention for their hardware affinity and non-disruption of model convergence compared to algorithm-level ones. However, they are troubled by high communication costs and poor communication-computation overlap. To address these challenges, we propose a systematic load-balancing method, Pro-Prophet, which consists of a planner and a scheduler for efficient parallel training of large-scale MoE models. To adapt to the dynamic load imbalance, we have profiled training statistics and utilized them to design Pro-Prophet. For lower communication volume, the Pro-Prophet planner determines a series of lightweight load-balancing strategies and efficiently searches for a communication-efficient one for training based on the statistics. For sufficient overlapping of communication and computation, the Pro-Prophet scheduler schedules the data-dependent operations based on the statistics and operation characteristics, further improving the training throughput. We conduct extensive experiments in various clusters and MoE models. The results indicate that Pro-Prophet achieves up to 2.66x speedup on MoE-GPT models compared to two popular MoE frameworks, namely Deepspeed-MoE and FasterMoE. Furthermore, Pro-Prophet has demonstrated a load-balancing improvement of up to 11.01x and speedups on modern MoE models up to 1.22x compared to a representative load-balancing work, FasterMoE.
Zhiquan Lai, Dongsheng Li 0001, Ke-shi Ge, Huayou Su
IEEE Trans. Parallel Distributed Syst.6
2025 Efficient deep neural network training via decreasing precision with layer capacity
abstract
Abstract Low-precision training has emerged as a practical approach, saving the cost of time, memory, and energy during deep neural networks (DNNs) training. Typically, the use of lower precision introduces quantization errors that need to be minimized to maintain model performance, often neglecting to consider the potential benefits of reducing training precision. This paper rethinks low-precision training, highlighting the potential benefits of lowering precision: (1) low precision can serve as a form of regularization in DNN training by constraining excessive variance in the model; (2) layer-wise low precision can be seen as an alternative dimension of sparsity, orthogonal to pruning, contributing to improved generalization in DNNs. Based on these analyses, we propose a simple yet powerful technique–DPC (Decreasing Precision with layer Capacity), which directly assigns different bit-widths to model layers, without the need for an exhaustive analysis of the training process or any delicate low-precision criteria. Thorough extensive experiments on five datasets and fourteen models across various applications consistently demonstrate the effectiveness of the proposed DPC technique in saving computational cost (−16.21%–−44.37%) while achieving comparable or even superior accuracy (up to +0.68%, +0.21% on average). Furthermore, we offer feature embedding visualizations and conduct further analysis with experiments to investigate the underlying mechanisms behind DPC’s effectiveness, enhancing our understanding of low-precision training. Our source code will be released upon paper acceptance.
Zhiquan Lai, Tao Sun 0005, Ke-shi Ge, Dongsheng Li 0001
Frontiers Comput. Sci.5
2025 AutoPipe-H: A Heterogeneity-Aware Data-Paralleled Pipeline Approach on Commodity GPU Servers
abstract
Recently, the data-parallel pipeline approach has been widely used in training DNN models on commodity GPU servers. However, there are still three challenges for hybrid parallelism on commodity GPU servers: i) a balanced model partition is crucial for efficiency, whereas prior works lack a sound solution to generate a balanced partition automatically; ii) an orchestrated device mapping is essential to reduce communication contention, however, prior works ignore server heterogeneity, exacerbating communication contention; iii) the startup overhead is inevitable and especially significant for deep pipelines, which is an essential source of pipeline bubbles and severely affects pipeline scalability. We proposeAutoPipe-Hto solve these three problems, which contains i) apipeline partitionercomponent for automatically and quickly generating a balanced sub-block partition scheme; ii) adevice mappingcomponent that assigns pipeline stages to devices, considering server heterogeneity, to reduce communication contention; and iii) adistributed training runtimecomponent that reduces pipeline startup overhead by splitting the micro-batch evenly. The experimental results show that AutoPipe-H can accelerate training by up to 1.26x over the hybrid parallelism framework DAPPLE and Piper, with a 2.73x-12.7x improvement in the partition balance and an order-of-magnitude time reduction in partition scheme searching.
Kai Lu 0001, Zhiquan Lai, Ke-shi Ge, Dongsheng Li 0001, Xicheng Lu
IEEE Trans. Computers5
2025 Oases: Efficient Large-Scale Model Training on Commodity Servers via Overlapped and Automated Tensor Model Parallelism
abstract
Deep learning is experiencing a rise in large-scale models. Training large-scale models is costly, prompting researchers to train large-scale models on commodity servers that more researchers can access. The massive number of parameters necessitates the use of model parallelism training methods. Existing studies focus on training with pipeline model parallelism. However, the tensor model parallelism (TMP) is inevitable when the model size keeps increasing, where frequent data-dependent communication and computation operations significantly reduce the training efficiency. In this paper, we present Oases, an automated TMP method with overlapped communication to accelerate large-scale model training on commodity servers. Oases proposes a fine-grained training operation schedule to maximize overlapping communication and computation that have data dependence. Additionally, we design the Oases planner that searches for the best model parameter partition strategy of TMP to achieve further accelerations. Unlike existing methods, Oases planner is tailored to model the cost of overlapped communication-computation operations. We evaluate Oases on various model settings and two commodity clusters, and compare Oases to four state-of-the-art implementations. Experimental results show that Oases achieves speedups of 1.01–1.48 × over the fastest baseline, and speedups of up to 1.95 × over Megatron.
Zhiquan Lai, Dongsheng Li 0001, Yanqi Hao, Ke-shi Ge, Xiaoge Deng, Kai Lu 0001
IEEE Trans. Parallel Distributed Syst.6
2024 HSDP: Accelerating Large-scale Model Training via Efficient Sharded Data Parallelism
abstract
Large deep neural network (DNN) models have demonstrated exceptional performance across diverse downstream tasks. Sharded data parallelism (SDP) has been widely used to reduce the memory footprint of model states. In a DNN training cluster, a device usually has multiple inter-device links that connect to other devices, like NVLink and InfiniBand. However, existing SDP approaches employ a single link at any given time, encountering challenges in efficient training due to significant communication overheads. We observe that the inter-device links can work independently without affecting each other. To reduce the fatal communication overhead of distributed training of large DNNs, this paper introduces HSDP, an efficient SDP training approach that enables the simultaneous utilization of multiple inter-device links. HSDP partitions models in a novel fine-grained manner and orchestrates the communication processes of partitioned parameters while considering inter-device links. This design enables concurrent communication execution and reduces communication overhead. To further optimize the training performance of HSDP, we propose a HSDP planner. The HSDP planner first abstracts the model partition and execution of HSDP into a communication parallel strategy, and builds a cost model to estimate the performance of each strategy. We then formulate the strategy searching as an optimization problem and solve it with an off-the-shelf solver. Evaluations on representative DNN workloads demonstrate that HSDP achieves up to 1.30× speedup compared to the state-of-the-art SDP training approaches.
Yanqi Hao, Zhiquan Lai, Ke-shi Ge, Dongsheng Li 0001
ISPA6
2024 Advances of Pipeline Model Parallelism for Deep Learning Training: An Overview
Jiye Liang, Ke-shi Ge, Xicheng Lu
J. Comput. Sci. Technol.5
2024 A Multidimensional Communication Scheduling Method for Hybrid Parallel DNN Training
abstract
The transformer-based deep neural network (DNN) models have shown considerable success across diverse tasks, prompting widespread adoption of distributed training methods such as data parallelism and pipeline parallelism. With the increasing parameter number, hybrid parallel training becomes imperative to scale training. The primary bottleneck in scaling remains the communication overhead. The communication scheduling technique, emphasizing the overlap of communication with computation, has demonstrated its benefits in scaling. However, most existing works focus on data parallelism, overlooking the nuances of hybrid parallel training. In this paper, we proposeTriRace, an efficient communication scheduling framework for accelerating communications in hybrid parallel training of asynchronous pipeline parallelism and data parallelism. To achieve effective computation-communication overlap,TriRaceintroduces3D communication scheduling, which adeptly leverages data dependencies between communication and computations, efficiently scheduling AllReduce communication, sparse communication, and peer-to-peer communication in hybrid parallel training. To avoid possible communication contentions,TriRacealso incorporates atopology-aware runtimewhich optimizes the execution of communication operations by considering ongoing communication operations and real-time network status. We have implemented a prototype ofTriRacebased on PyTorch and Pipedream-2BW, and conducted comprehensive evaluations with three representative baselines. Experimental results show thatTriRaceachieves up to 1.07–1.45× speedup compared to the state-of-the-art pipeline parallelism training baseline Pipedream-2BW, and 1.24–1.81× speedup compared to the Megatron.
Kai Lu 0001, Zhiquan Lai, Ke-shi Ge, Dongsheng Li 0001
IEEE Trans. Parallel Distributed Syst.5
2023 Prophet: Fine-grained Load Balancing for Parallel Training of Large-scale MoE Models
abstract
Mixture of Expert (MoE) has received increasing attention for scaling DNN models to extra-large size with negligible increases in computation. The MoE model has achieved the highest accuracy in several domains. However, a significant load imbalance occurs in the device during the training of a MoE model, resulting in significantly reduced throughput. Previous works on load balancing either harm model convergence or suffer from high execution overhead. To address these issues, we present Prophet: a fine-grained load balancing method for parallel training of large-scale MoE models, which consists of a planner and a scheduler. Prophet planner first employs a fine-grained resource allocation method to determine the possible scenarios for the expert placement in a fine-grained manner, and then efficiently searches for a well-balanced expert placement to balance the load without introducing additional overhead. Prophet scheduler exploits the locality of the token distribution to schedule the resource allocation operations using a layer-wise fine-grained schedule strategy to hide their overhead. We conduct extensive experiments in four clusters and five representative models. The results indicate that Prophet gains up to 2.3x speedup compared to the state-of-the-art MoE frameworks including Deepspeed-MoE and FasterMoE. Additionally, Prophet achieves a load balancing enhancement of up to 12.06x when compared to FasterMoE.
Zhiquan Lai, Ke-shi Ge, Dongsheng Li 0001
CLUSTER5
2023 Auto-Divide GNN: Accelerating GNN Training with Subgraph Division
Zhejiang Ran, Ke-shi Ge, Zhiquan Lai, Jingfei Jiang, Dongsheng Li 0001
Euro-Par3
2023 Compressed Collective Sparse-Sketch for Distributed Data-Parallel Training of Deep Learning Models
abstract
Distributed data-parallel training (DDP) is prevalent in large-scale deep learning. To increase the training throughput and scalability, high-performance collective communication methods such as AllReduce have recently proliferated for DDP use. However, these approaches require long communication periods with increasing model sizes. Collective communication transmits many sparse gradient values that can be efficiently compressed to reduce the required training time. State-of-the-art compression approaches do not provide mergeable compression for AllReduce and lack convergence bounds. We present a sparse sketch reducer (S2Reducer), a sparsity-preserving sketch-based collective communication method. S2Reducer preserves gradient sparsity and reduces communication costs via a bitmap informed count sketch structure and adapts to efficient AllReduce operators. We tune the count sketch organization to minimize the hash conflicts in a fixed-size budget. We prove that our method has the same convergence rate as vanilla data-parallel training and a much smaller communication overhead than those of state-of-the-art methods. We implement a GPU-accelerated S2Reducer for the Ring AllReduce-based DDP system. We perform extensive evaluations against four state-of-the-art methods across seven deep learning models. Our results show that S2Reducer converges to the same accuracy as that of state-of-the-art approaches while reducing the sparse communication overhead by up to 86% and achieving a speedup of up to$3.5\times $in distributed training.
Ke-shi Ge, Kai Lu 0001, Yongquan Fu, Xiaoge Deng, Zhiquan Lai, Dongsheng Li 0001
IEEE J. Sel. Areas Commun.1
2023 Merak: An Efficient Distributed DNN Training Framework With Automated 3D Parallelism for Giant Foundation Models
abstract
Foundation models are in the process of becoming the dominant deep learning technology. Pretraining a foundation model is always time-consuming due to the large scale of both the model parameter and training dataset. Besides being computing-intensive, the pretraining process is extremely memory- and communication-intensive. These challenges make it necessary to apply 3D parallelism, which integrates data parallelism, pipeline model parallelism, and tensor model parallelism, to achieve high training efficiency. However, current 3D parallelism frameworks still encounter two issues: i) they are not transparent to model developers, requiring manual model modification to parallelize training, and ii) their utilization of computation resources, GPU memory, and network bandwidth is insufficient. We proposeMerak, an automated 3D parallelism deep learning training framework with high resource utilization. Merak automatically deploys 3D parallelism with an automatic model partitioner, which includes a graph-sharding algorithm and proxy node-based model graph. Merak also offers a non-intrusive API to scale out foundation model training with minimal code modification. In addition, we design a high-performance 3D parallel runtime engine that employs several techniques to exploit available training resources, including a shifted critical path pipeline schedule that increases computation utilization, stage-aware recomputation that makes use of idle worker memory, and sub-pipelined tensor model parallelism that overlaps communication and computation. Experiments on 64 GPUs demonstrate Merak's capability to speed up training performance over state-of-the-art 3D parallelism frameworks of models with 1.5, 2.5, 8.3, and 20 billion parameters by up to 1.42, 1.39, 1.43, and 1.61×, respectively.
Zhiquan Lai, Xudong Tang, Ke-shi Ge, Yabo Duan, Linbo Qiao, Dongsheng Li 0001
IEEE Trans. Parallel Distributed Syst.4
2022 HPH: Hybrid Parallelism on Heterogeneous Clusters for Accelerating Large-scale DNNs Training
abstract
As the deep learning model grows larger, training model with a single computational resource becomes impractical. To solve this, hybrid parallelism, which combines data and pipeline parallelism emerges to train large models with multiple GPUs. In practice, using heterogeneous GPU clusters to train large models is a common need due to the upgrade of a part of hardware. However, existing hybrid parallelism approaches in the heterogeneous environment do not work well in communication efficacy, workload balance among GPUs and utilizing the memory constrained GPU. To address these problems, we present a parallel DNN training approach, Hybrid Parallelism on Heterogeneous clusters (HPH). In HPH, we propose a topology designer that minimizes the communication time cost. Furthermore, HPH uses a partition algorithm that automatically partitions DNN layers among workers to maximize throughput. Besides, HPH adopts recomputation-aware scheduling to reduce memory consumption and further reschedule the pipeline to eliminate the extra time overhead of recomputation. Our experimental results on a 32-GPU heterogeneous cluster show that HPH achieves up to 1.42x training speed-ups compared with the state-of-the-art approach.
Yabo Duan, Zhiquan Lai, Ke-shi Ge, Peng Liang 0017, Dongsheng Li 0001
CLUSTER5
2022 AutoPipe: A Fast Pipeline Parallelism Approach with Balanced Partitioning and Micro-batch Slicing
abstract
Recently, pipeline parallelism has been widely used in training large DNN models. However, there are still two main challenges for efficient pipeline parallelism: i) a balanced model partition is crucial for pipeline efficiency, whereas prior works lack a sound solution to generate a balanced partition automatically. ii) the startup overhead is inevitable and especially significant for deep pipelines, which is an essential source of pipeline bubbles and severely affects pipeline scalability. We propose AutoPipe to solve these two problems, which contains i) a planner for automatically and quickly generating a balanced pipeline partition scheme with a fine-grained partitioner. This partitioner groups DNN in the sub-layer granularity and finds the balanced scheme with a heuristic search algorithm; and ii) a micro-batch slicer that reduces pipeline startup overhead according to the planner results by splitting the micro-batch evenly. This slicer automatically solves an appropriate number of micro-batches to split. The experimental results show that AutoPipe can accelerate training by up to 1.30x over the state-of-the-art distributed training framework Megatron-LM, with a 50% reduction in startup overhead and an order-of-magnitude reduction in pipeline planning time. Furthermore, AutoPipe Planner improves the partition balance by 2.73x-12.7x compared to DAPPLE Planner and Piper.
Zhiquan Lai, Yabo Duan, Ke-shi Ge, Dongsheng Li 0001
CLUSTER5
2022 S2 Reducer: High-Performance Sparse Communication to Accelerate Distributed Deep Learning
abstract
Distributed stochastic gradient descent (SGD) approach has been widely used in large-scale deep learning, and the gradient collective method is vital to ensure the training scalability of the distributed deep learning system. Collective communication such as AllReduce has been widely adopted for the distributed SGD process to reduce the communication time. However, AllReduce incurs large bandwidth resources while most gradients are sparse in many cases since many gradient values are zeros and should be efficiently compressed for bandwidth saving. To reduce the sparse gradient communication overhead, we propose Sparse-Sketch Reducer (S2 Reducer), a novel sketch-based sparse gradient aggregation method with convergence guarantees. S2 Reducer reduces the communication cost by only compressing the non-zero gradients with count-sketch and bitmap, and enables the efficient AllReduce operators for parallel SGD training. We perform extensive evaluation against four state-of-the-art methods over five training models. Our results show that S2 Reducer converges to the same accuracy, reduces 81% sparse communication overhead, and achieves 1.8× distributed training speedup compared to state-of-the-art approaches.
Ke-shi Ge, Yongquan Fu, Yiming Zhang 0003, Zhiquan Lai, Xiaoge Deng, Dongsheng Li 0001
ICASSP1
2022 BRGraph: An efficient graph neural network training system by reusing batch data on GPU
abstract
Summary With the increasing adoption of graph neural networks (GNNs) in the community, various GPU‐based graph programming systems have been developed to improve the productivity of GNNs. However, sampling‐based GNN training is still inefficient, and we observe that the main bottleneck comes from the data transferring, where vertex features are transferred from host memory to GPU through limited bandwidth. In this article, we propose BRGraph, a sampling‐based GNN training system that supports efficient data transferring. BRGraph leverages the duplicate vertices between mini‐batches by batch reusing (BR) strategy to avoid duplicate data transmission. Furthermore, to reduce the overhead of detecting duplicate vertices, we design an efficient parallel batch reusing algorithm based on GPU. BRGraph also exploits the data reusing potential of the non‐duplicate vertex features by the two‐level batch reusing (two‐level BR) strategy. Comprehensive evaluations on three representative GNN models show that BRGraph reduces data transferring time by up to 60% and delivers up to 1.79 GNN training speedup over the state‐of‐the‐art baselines. Besides, it can save GPU memory by up to 40% while reaching the same training time compared with the static cache strategy. When applying the two‐level BR, BRGraph further reduces 20% of the data transferring time compared with the BR.
Ke-shi Ge, Zhejiang Ran, Zhiquan Lai, Dongsheng Li 0001
Concurr. Comput. Pract. Exp.1
2021 CASQ: Accelerate Distributed Deep Learning with Sketch-Based Gradient Quantization
abstract
Gradient quantization has been widely used in distributed training of deep neural network (DNN) models to reduce communication costs. However, existing quantization methods overlook that gradients have a nonuniform distribution changing over time, which can lead to significant gradient variance that requires a higher number of quantization bits (and consequently higher communication cost) to keep the validation accuracy as high as stochastic gradient descent (SGD). In this paper, we propose Cluster-Aware Sketch Quantization (CASQ), a novel sketch-based gradient quantization method for SGD. CASQ models the nonuniform distribution of gradients via clustering, and adaptively allocates appropriate numbers of hash buckets based on the statistics of different clusters to compress gradients. The extensive evaluation shows that compared to existing quantization methods CASQ-based SGD (i) achieves the same validation accuracy when decreasing quantization level from 3 bits to 2 bits, and (ii) reduces the training time to convergence by up to 43% for the same training loss.
Ke-shi Ge, Yiming Zhang 0003, Yongquan Fu, Zhiquan Lai, Xiaoge Deng, Dongsheng Li 0001
CLUSTER1
2020 Tag Pollution Detection in Web Videos via Cross-Modal Relevance Estimation
abstract
In the era of big data, web videos are known for their astronomical volume and great difficulties to be understood by computers. Therefore it is challenging to detect and curb tag pollution on social networking platforms. Intuitively, the pollution in the tags of a video can be identified by exploiting the tags of visually similar videos. From this intuition, we develop a semi-supervised approach to estimate the relevance between Internet videos and user-provided labels accurately and detect polluted labels accordingly. To further enhance the accuracy of relevance estimation and pollution detection, we introduce three multi-view multi-label models, which employ the coherence and differences between various similarity relations of videos. Compared with two quintessential multi-view fusion models, the proposed models consistently outperform or achieve comparable performance.
Yixin Chen 0004, Xinye Lin, Ke-shi Ge, Wenbo He 0003, Dongsheng Li 0001
IWQoS3
2020 An efficient parallel and distributed solution to nonconvex penalized linear SVMs
abstract
Support vector machines (SVMs) have been recognized as a powerful tool to perform linear classification. When combined with the sparsity-inducing nonconvex penalty, SVMs can perform classification and variable selection simultaneously. However, the nonconvex penalized SVMs in general cannot be solved globally and efficiently due to their nondifferentiability, nonconvexity, and nonsmoothness. Existing solutions to the nonconvex penalized SVMs typically solve this problem in a serial fashion, which are unable to fully use the parallel computing power of modern multi-core machines. On the other hand, the fact that many real-world data are stored in a distributed manner urgently calls for a parallel and distributed solution to the nonconvex penalized SVMs. To circumvent this challenge, we propose an efficient alternating direction method of multipliers (ADMM) based algorithm that solves the nonconvex penalized SVMs in a parallel and distributed way. We design many useful techniques to decrease the computation and synchronization cost of the proposed parallel algorithm. The time complexity analysis demonstrates the low time complexity of the proposed parallel algorithm. Moreover, the convergence of the parallel algorithm is guaranteed. Experimental evaluations on four LIBSVM benchmark datasets demonstrate the efficiency of the proposed parallel algorithm.
Lei Guan 0001, Tao Sun 0005, Linbo Qiao, Zhi-hui Yang, Dongsheng Li 0001, Ke-shi Ge, Xicheng Lu
Frontiers Inf. Technol. Electron. Eng.6
2019 HPDL: Towards a General Framework for High-performance Distributed Deep Learning
abstract
With growing scale of the data volume and neural network size, we have come into the era of distributed deep learning. High-performance training and inference on distributed computing systems has been attracting increasing research attention in both academia and industry. Meanwhile, diversity of existing machine learning frameworks (e.g. TensorFlow, Pytorch and MXNet) and the explosion of deep learning hardwares (e.g. CPUs, GPUs, FPGAs and ASICs) bring more challenges for users to leverage new deep learning technologies and accelerating capability of hardware devices. We firstly search around the state-of-the-art work in the area which open our mind to take a vision upon the future deep learning framework. Then, we propose HPDL, a general framework for high-performance distributed deep learning which is compatible with existing frameworks and adaptive to various hardware architectures. At last, we discuss and foresee the key technologies fulfilling high-performance and large-scale deep learning, including optimization algorithm, hybrid communication mechanism, model parallelization, resource scheduling and single-node execution optimization.
Dongsheng Li 0001, Zhiquan Lai, Ke-shi Ge, Yiming Zhang 0003, Zhaoning Zhang 0001, Huaimin Wang 0001
ICDCS3
2018 Deep Discriminative Clustering Network
abstract
Deep clustering aims to cluster unlabeled data by embedding them into a subspace based on deep model. The key challenge of deep clustering is to learn discriminative representations for input data with high dimensions. In this paper, we present a deep discriminative clustering network for clustering the real-world images. We use a convolutional auto-encoder stacked with a softmax layer to predict clustering assignments. To learn a discriminative representations, the proposed approach adds discriminative loss as embedded regularization with relative entropy minimization. With the discriminative loss, the network can not only produce clustering assignments, but also learn discriminative features by reducing intra-cluster distance and increasing inter-cluster distance. We evaluate the proposed method on three datasets: MNIST-full, YTF and FRGC-v2.0. We outperform state-of-the-art results on MNIST-full and FRGC-v2.0 and achieve competitive result on YTF. The source code has been made publicly available at https://github.com/shaoxuying/DeepDiscriminativeClusteringNetwork.
Xuying Shaol, Ke-shi Ge, Huayou Su, Lei Luo 0002, Baoyun Peng, Dongsheng Li 0001
IJCNN2
2017 Efficient parallel implementation of a density peaks clustering algorithm on graphics processing unit
abstract
The density peak (DP) algorithm has been widely used in scientific research due to its novel and effective peak density-based clustering approach. However, the DP algorithm uses each pair of data points several times when determining cluster centers, yielding high computational complexity. In this paper, we focus on accelerating the time-consuming density peaks algorithm with a graphics processing unit (GPU). We analyze the principle of the algorithm to locate its computational bottlenecks, and evaluate its potential for parallelism. In light of our analysis, we propose an efficient parallel DP algorithm targeting on a GPU architecture and implement this parallel method with compute unified device architecture (CUDA), called the ‘CUDA-DP platform’. Specifically, we use shared memory to improve data locality, which reduces the amount of global memory access. To exploit the coalescing accessing mechanism of GPU, we convert the data structure of the CUDA-DP program from array of structures to structure of arrays. In addition, we introduce a binary search-and-sampling method to avoid sorting a large array. The results of the experiment show that CUDA-DP can achieve a 45-fold acceleration when compared to the central processing unit based density peaks implementation.
Ke-shi Ge, Huayou Su, Dongsheng Li 0001, Xicheng Lu
Frontiers Inf. Technol. Electron. Eng.1