Jinkun Geng

dblp:171/1013 · DBLP profile ↗
← Back
29ranked-venue papers
8as first author
13since 2021 · last 2025
0000-0002-6574-8349ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 14 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 3 since 2021Systems, architecture and hardware · 4 · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2025 FlexPie: Accelerate Distributed Inference on Edge Devices with Flexible Combinatorial Optimization
Runhua Zhang 0002, Jinkun Geng, Chenhui Zhu
DASFAA (1)3
2025 Beyond Lamport, Towards Probabilistic Fair Ordering
abstract
A growing class of applications demands fair ordering of events, which ensures that events generated earlier are processed before later events. However, achieving such sequencing is challenging due to the inherent errors in clock synchronization: two events at two clients generated close together may have timestamps that cannot be compared confidently. We advocate for an approach that embraces, rather than eliminates, clock synchronization errors. Instead of attempting to remove the error from a timestamp, Tommy, our proposed system, leverages a statistical model to compare two noisy timestamps probabilistically by learning per-clock synchronization error distributions. Our preliminary statistical model computes the probability that one event precedes another by only relying on local clocks of clients. This serves as a foundation for a new relation: likely-happened-before denoted by →p where p represents the probability that an event happened before another. The →p relation provides a basis for ordering multiple events which are otherwise considered concurrent by Lamport's happened-before (→) relation. We highlight various related challenges including the intransitivity of the →p relation as opposed to the transitive → relation. We outline several research directions: online fair sequencing, stochastically fair total ordering, and handling byzantine clients.
Jinkun Geng, Radhika Mittal, Aurojit Panda, Srinivas Narayana, Anirudh Sivaraman
HotNets2
2025 Network Support For Scalable And High Performance Cloud Exchanges
abstract
Financial exchanges are migrating to the public cloud, but the best-effort nature of the cloud fabric is at odds with the stringent networking requirements of the exchanges. We present Onyx, a system for meeting such requirements which uses many well-studied techniques in a new context as well as introduces new techniques that enable a scalable cloud financial exchange. An overlay multicast tree is used to disseminate data to 1000 participants with ≤ 1 μs difference in data reception time between any two participants, crucial for maintaining fair competition. Several techniques for mitigating latency variance are introduced. Onyx also presents a scheduling policy for trade orders that enhances an exchange's performance and gracefully services bursty traffic. Onyx achieves ≈50% lower latency than the AWS multicast service [1]. Onyx outperforms an existing system, CloudEx [2] in terms of supported number of participants, exchange's throughput and multicast latency. Onyx's techniques can be applied to other existing systems (e.g., DBO) to enhance their performance.
Jinkun Geng, Daniel Duclos-Cavalcanti, Xiyu Hao, Ulysses Butler, Radhika Mittal, Srinivas Narayana, Anirudh Sivaraman
SIGCOMM2
2025 Tiga: Accelerating Geo-Distributed Transactions with Synchronized Clocks
abstract
This paper presents Tiga, a new design for geo-replicated and scalable transactional databases such as Google Spanner. Tiga aims to commit transactions within 1 wide-area roundtrip time, or 1 WRTT, for a wide range of scenarios, while maintaining high throughput with minimal computational overhead. Tiga consolidates concurrency control and consensus, completing both strictly serializable execution and consistent replication in a single round. It uses synchronized clocks to proactively order transactions by assigning each a future timestamp at submission. In most cases, transactions arrive at servers before their future timestamps and are serialized according to the designated timestamp, requiring 1 WRTT to commit. In rare cases, transactions are delayed and proactive ordering fails, in which case Tiga falls back to a slow path, committing in 1.5–2 WRTTs. Compared to state-of-the-art solutions, Tiga can commit more transactions at 1-WRTT latency, and incurs much less throughput overhead. Evaluation results show that Tiga outperforms all baselines, achieving 1.3–7.2× higher throughput and 1.4–4.6× lower latency. Tiga is open-sourced at https://github.com/New-Consensus-Concurrency-Control/Tiga.
Jinkun Geng, Shuai Mu 0001, Anirudh Sivaraman, Balaji Prabhakar
SOSP1
2024 A high-performance dataflow-centric optimization framework for deep learning inference on the edge
Runhua Zhang 0002, Jinkun Geng
J. Syst. Archit.3
2023 Xenos : Dataflow-Centric Optimization to Accelerate Model Inference on Edge Devices
Runhua Zhang 0002, Jinkun Geng, Chenhui Zhu
DASFAA (1)4
2023 Diffusion Policies as Multi-Agent Reinforcement Learning Strategies
Jinkun Geng, Xiubo Liang, Hongzhi Wang 0009
ICANN (3)1
2023 Light: A Compatible, high-performance and scalable user-level network stack
Dan Li 0001, Huiyou Jiang, Du Lin, Jinkun Geng, K. K. Ramakrishnan, Kai Zheng 0003
Comput. Networks5
2022 Nezha: Deployable and High-Performance Consensus Using Synchronized Clocks
abstract
This paper presents a high-performance consensus protocol, Nezha, which can be deployed by cloud tenants without support from cloud providers. Nezha bridges the gap between protocols such as Multi-Paxos and Raft, which can be readily deployed, and protocols such as NOPaxos and Speculative Paxos, that provide better performance, but require access to technologies such as programmable switches and in-network prioritization, which cloud tenants do not have. Nezha uses a new multicast primitive called deadline-ordered multicast (DOM). DOM uses high-accuracy software clock synchronization to synchronize sender and receiver clocks. Senders tag messages with deadlines in synchronized time; receivers process messages in deadline order, on or after their deadline. We compare Nezha with Multi-Paxos, Fast Paxos, Raft, (optimized) NOPaxos, and 2 recent protocols, Domino and TOQ-EPaxos, that use synchronized clocks. In throughput, Nezha outperforms all baselines by a median of 5.4X (range: 1.9--20.9X). In latency, Nezha outperforms five baselines by a median of 2.3X (range: 1.3--4.0X), with one exception: it sacrifices 33% of latency compared with our optimized NOPaxos in one test. We also prototype two applications, a key-value store and a fair-access stock exchange, on top of Nezha to show that Nezha only modestly reduces their performance relative to an unreplicated system.
Jinkun Geng, Anirudh Sivaraman, Balaji Prabhakar, Mendel Rosenblum
Proc. VLDB Endow.1
2022 Impact of Synchronization Topology on DML Performance: Both Logical Topology and Physical Topology
abstract
To tackle the increasingly larger training data and models, researchers and engineers resort to multiple servers in a data center for distributed machine learning (DML). On one hand, DML enables us to leverage the computation power of multiple servers, which can effectively accelerate those computation-intensive tasks. On the other hand, DML also incurs significant communication cost due to parameter synchronization among these servers. In this paper, we want to explore the impact of synchronization topology, including both logical topology and physical topology, on the DML performance. First, we revisit the existing logical topologies, e.g., parameter server and ring allreduce, for parameter synchronization, and we find that theseflatsynchronization topologies is inefficient when running a large-scale DML training. Therefore, we propose a hierarchical parameter synchronization topology, called HiPS, which can achieve efficient parameter synchronization even on a large scale. Then, we compare two representative physical network topologies, namely, Fat-Tree and BCube. Based on our analyses, BCube has many advantages over Fat-Tree, e.g., higher bandwidth, better load balance, and lower hardware cost. The simulation results also show that BCube is more friendly to RDMA. Relying on the advantages of HiPS and BCube, the GST of “HiPS+BCube” is 12% ~ 70% lower than other combinations. Moreover, when the cluster size increases from 16 to 1024, the performance of “HiPS+BCube” only drops by 6.5%, while the performance of “Ring+BCube” drops by 44.6%. Hence, we believe “HiPS+BCube” is the optimal solution to benefit DML in large scale.
Shuai Wang 0028, Jinkun Geng, Dan Li 0001
IEEE/ACM Trans. Netw.2
2021 CloudEx: a fair-access financial exchange in the cloud
abstract
Financial exchanges have begun a move from on-premise and custom-engineered datacenters to the public cloud, accelerated by a rush of new investors, the rise of remote work, cost savings from the cloud, and the desire for more resilient infrastructure. While the promise of the cloud is enticing, the cloud's varying network latencies can lead to market unfairness: orders can be processed out of sequence, and market data can be disseminated to market participants at incorrect times due to varying latencies between participants and the exchange. We present CloudEx, a fair-access cloud exchange, which leverages high-precision software clock synchronization to compensate for noisy network conditions in the public cloud. We also discuss refinements to the CloudEx design that were informed by lessons learned from deploying CloudEx in two academic courses and conclude by outlining future research directions.
Ahmad Ghalayini, Jinkun Geng, Vighnesh Sachidananda, Vinay Sriram, Yilong Geng, Balaji Prabhakar, Mendel Rosenblum, Anirudh Sivaraman
HotOS2
2021 Sphinx: A transport protocol for high-speed and lossy mobile networks
Dan Li 0001, Wenfei Wu, K. K. Ramakrishnan, Jinkun Geng, Fanzhao Wang, Kai Zheng 0003
Comput. Networks5
2021 Accelerating End-to-End Deep Learning Workflow With Codesign of Data Preprocessing and Scheduling
abstract
In this article, we investigate the performance bottleneck of existing deep learning (DL) systems and propose DLBooster to improve the running efficiency of deploying DL applications on GPU clusters. At its core, DLBooster leverages two-level optimizations to boost the end-to-end DL workflow. On the one hand, DLBooster selectively offloads some key decoding workloads to FPGAs to provide high-performance online data preprocessing services to the computing engine. On the other hand, DLBooster reorganizes the computational workloads of training neural networks with the backpropagation algorithm and schedules them according to their dependencies to improve the utilization of GPUs at runtime. Based on our experiments, we demonstrate that compared with baselines, DLBooster can improve the image processing throughput by 1.4× - 2.5× and reduce the processing latency by 1/3 in several real-world DL applications and datasets. Moreover, DLBooster consumes less than 1 CPU core to manage FPGA devices at runtime, which is at least 90 percent less than the baselines in some cases. DLBooster shows its potential to accelerate DL workflows in the cloud.
Dan Li 0001, Binyao Jiang, Jinkun Geng, Wei Bai 0001, Yongqiang Xiong
IEEE Trans. Parallel Distributed Syst.5
2020 Fela: Incorporating Flexible Parallelism and Elastic Tuning to Accelerate Large-Scale DML
abstract
Distributed machine learning (DML) has become the common practice in industry, because of the explosive volume of training data and the growing complexity of training model. Traditional DML follows data parallelism but causes significant communication cost, due to the huge amount of parameter transmission. The recently emerging model-parallel solutions can reduce the communication workload, but leads to load imbalance and serious straggler problems. More importantly, the existing solutions, either data-parallel or model-parallel, ignore the nature of flexible parallelism for most DML tasks, thus failing to fully exploit the GPU computation power. Targeting at these existing drawbacks, we propose Fela, which incorporates both flexible parallelism and elastic tuning mechanism to accelerate DML. In order to fully leverage GPU power and reduce communication cost, Fela adopts hybrid parallelism and uses flexible parallel degrees to train different parts of the model. Meanwhile, Fela designs token-based scheduling policy to elastically tune the workload among different workers, thus mitigating the straggler effect and achieve better load balance. Our comparative experiments show that Fela can significantly improve the training throughput and outperforms the three main baselines (i.e. dataparallel, model-parallel, and hybrid-parallel) by up to 3.23×, 12.22×, and 1.85× respectively.
Jinkun Geng, Dan Li 0001, Shuai Wang 0028
ICDE1
2020 Geryon: Accelerating Distributed CNN Training by Network-Level Flow Scheduling
abstract
Increasingly rich data sets and complicated models make distributed machine learning more and more important. However, the cost of extensive and frequent parameter synchronizations can easily diminish the benefits of distributed training across multiple machines. In this paper, we present Geryon, a network-level flow scheduling scheme to accelerate distributed Convolutional Neural Network (CNN) training. Geryon leverages multiple flows with different priorities to transfer parameters of different urgency levels, which can naturally coordinate multiple parameter servers and prioritize the urgent parameter transfers in the entire network fabric. Geryon requires no modification in CNN models and does not affect the training accuracy. Based on the experimental results of four representative CNN models on a testbed of 8 GPU servers, Geryon achieves up to 95.7% scaling efficiency even with 10GbE bandwidth. In contrast, for most models, the scaling efficiency of vanilla TensorFlow is no more than 37% and that of TensorFlow with parameter partition and slicing is around 80%. In terms of training throughput, Geryon enhanced with parameter partition and slicing achieves up to 4.37x speedup, where the flow scheduling algorithm itself achieves up to 1.2x speedup over parameter partition and slicing.
Shuai Wang 0028, Dan Li 0001, Jinkun Geng
INFOCOM3
2020 Incorporating Intra-flow Dependencies and Inter-flow Correlations for Traffic Matrix Prediction
abstract
Traffic matrix (TM) prediction is essential for effective traffic engineering and network management. Based on our analysis of real traffic traces from Wide Area Network, the traffic flows in TM are both time-varying (i.e. with intra-flow dependencies) and correlated with each other (i.e. with inter-flow correlations). However, most existing works in TM prediction ignore inter-flow correlations. In this paper, we propose a novel Attention-based Convolutional Recurrent Neural Network (ACRNN) model to capture both intra-flow dependencies and inter-flow correlations. ACRNN mainly contains two components: 1) Correlational Modeling employs attention-based convolutional structures to capture the correlation of any two flows in TMs; 2) Temporal Modeling uses attention-based recurrent structures to model the long-term temporal dependencies of each flow, and then predicts TMs according inter-flow correlations and intra-flow dependencies. Experiments on two real-world datasets show that, when predicting the next TM, ACRNN model reduces the Mean Squared Error by up to 44.8% and reduces the Mean Absolute Error by up to 30.6%, compared to state-of-the-art method; and the gap is even larger when predicting the next multiple TMs. Besides, simulation results demonstrate that ACRNN's accurate prediction can help traffic engineering to mitigate traffic congestion.
Kaihui Gao, Dan Li 0001, Li Chen 0008, Jinkun Geng, Fei Gui
IWQoS4
2020 Web Service Recommendation based on Knowledge Graph Convolutional Network and Doc2Vec
abstract
With the rapid development of Internet, the number of Web services is increasing sharply, which makes it more difficult for Mashup developers to find suitable Web services. Nowadays, there are numerous methods to improve Web service recommendation, but it is still a challenging problem to recommend Web services with both good accuracy and satisfying diversity. Collaborative filtering is a common algorithm in recommendation system, but it often faces serious cold start and sparsity problems. To alleviate the above problems, this paper proposes a Web service recommendation method based on knowledge graph convolutional network and Doc2Vec. First of all, it constructs the knowledge graph of Web services based on the additional information such as the categories, developers, scope of application of Web services, and adopts knowledge graph convolutional networks to mine the higher-order relationship between Web service and the preference information of Mashups. Secondly, it employs Doc2Vec to mine the semantics of Web service description documents, and integrates the Mashup preference information and the Mashup semantic information in the training process, so as to predict Web services needed for Mashup development. Finally, the experiment is conducted on the latest Programmable Web dataset and the experimental results show that the recommended performance of the proposed method is better than that of FM, NCF, CKE, RippleNet, KGCN.
Jinkun Geng, Buqing Cao, Hongfan Ye, Mi Peng, Jianxun Liu 0001
SERVICES1
2020 A Scalable, High-Performance, and Fault-Tolerant Network Architecture for Distributed Machine Learning
abstract
In large-scale distributed machine learning (DML), the network performance between machines significantly impacts the speed of iterative training. In this paper we propose BML, a scalable, high-performance and fault-tolerant DML network architecture on top of Ethernet and commodity devices. BML builds on BCube topology, and runs a fully-distributed gradient synchronization algorithm. Compared to a Fat-Tree network with the same size, a BML network is expected to take much less time for gradient synchronization, for both low theoretical synchronization time and its benefit to RDMA transport. With server/link failures, the performance of BML degrades in a graceful way. Experiments of MNIST and VGG-19 benchmarks on a testbed with 9 dual-GPU servers show that, BML reduces the job completion time of DML training by up to 56.4% compared with Fat-Tree running state-of-the-art gradient synchronization algorithm.
Dan Li 0001, Jinkun Geng, Yanshu Wang, Shuai Wang 0028, Shutao Xia
IEEE/ACM Trans. Netw.4
2019 Accelerating Distributed Machine Learning by Smart Parameter Server
abstract
Parameter Server (PS)-based architecture is widely applied in distributed machine learning (DML), but it is still an open issue how to improve the DML performance in this frame-work. Existing works mainly focus on the view of workers. In this paper, we tackle this problem from another perspective, by leveraging the central control on the PS. Specifically, we propose SmartPS, which transforms the passive role of PS in traditional DML and fully exploits the intelligence of PS. Firstly, the PS holds the global view of parameter dependency, facilitating it to update workers' parameters selectively and proactively. Secondly, the PS records the workers' speeds, and prioritizes parameter transmission to narrow the gap between stragglers and fast workers. Thirdly, the PS considers the parameter dependency in consecutive training iterations, and opportunistically blocks unnecessary pushes from workers. We conduct comparative experiments with two typical benchmarks, Matrix Factorization (MF) and PageRank (PR). The experimental results prove that, compared with all the baseline algorithms (i.e. standard BSP, ASP and SSP), SmartPS can reduce the overall training time by 65.7%~84.9%, with the same training accuracy.
Jinkun Geng, Dan Li 0001, Shuai Wang 0028
APNet1
2019 Rima: An RDMA-Accelerated Model-Parallelized Solution to Large-Scale Matrix Factorization
abstract
Matrix factorization (MF) is a fundamental technique in machine learning and data mining, which gains wide application in many fields. When the matrix becomes large, MF cannot be processed on a single machine. Considering this, many distributed SGD algorithms (e.g. DSGD) have been developed to solve large-scale MF on multiple machines in a model-parallel way. Existing distributed algorithms are primarily implemented under Map/Reduce or PS (parameter server)-based architectures, which incur significant communication overheads. Besides, existing solutions cannot well embrace the benefit of RDMA/RoCE transport and suffer from scalability problems. Targeting at these drawbacks, we propose Rima, which uses ring-based model parallelism to solve large-scale MF with higher communication efficiency. Compared with PS-based SGD algorithms, Rima also consumes less queue pairs (QPs) and can thus better leverage the power of RDMA/RoCE to accelerate the training speed. Our experiment shows that, compared with PS-based DSGD when solving 1M × 1M MF, Rima achieves comparable convergence performance after equal number of iterations, but reduces the training time by 68.7% and 85.4% via TCP and RDMA respectively.
Jinkun Geng, Dan Li 0001, Shuai Wang 0028
ICDE1
2019 DLBooster: Boosting End-to-End Deep Learning Workflows with Offloading Data Preprocessing Pipelines
abstract
In recent years, deep learning (DL) has prospered again due to improvements in both computing and learning theory. Emerging studies mostly focus on the acceleration of refining DL models but ignore data preprocessing issues. However, data preprocessing can significantly affect the overall performance of end-to-end DL workflows. Our studies on several image DL workloads show that existing preprocessing backends are quite inefficient: they either perform poorly in throughput (30% degradation) or burn too many (>10) CPU cores. Based on these observations, we propose DLBooster, a high-performance data preprocessing pipeline that selectively offloads key workloads to FPGAs, to fit the stringent demands on data preprocessing for cutting-edge DL applications. Our testbed experiments show that, compared with the existing baselines, DLBooster can achieve 1.35×~2.4× image processing throughput in several DL workloads, but consumes only 1/10 CPU cores. Besides, it also reduces the latency by 1/3 in online image inference.
Dan Li 0001, Binyao Jiang, Xi Fan, Jinkun Geng, Wei Bai 0001, Ran Shu 0001, Peng Cheng 0005, Yongqiang Xiong
ICPP7
2019 Impact of Network Topology on the Performance of DML: Theoretical Analysis and Practical Factors
abstract
To deal with the increasingly larger input data and model sizes, it has become necessary to scale the training of machine learning models to multiple nodes, even a server cluster, which we call distributed machine learning, or DML. However, DML utilizes more computation power at the cost of high communication overhead, which may limit the overall performance in turn. In this paper, we study the impact of network topology on the DML performance both in theory and in practice. We compare two representative network topologies, namely, Fat-Tree which is widely-used in modern data centers, and BCube, which is a low-cost and server-centric network topology, both running on top of RDMA. The results show that Fat-Tree not only has theoretically higher global synchronization time (GST) than BCube, but its practical GST (by NS-3 based simulation) is also considerably larger than the theoretical one. By analyzing the large-scale simulation traces, we find that the root cause for the gap in Fat-Tree comes from the load imbalance among the multiple parallel paths as well as the inevitable PFC frames, both of which do not appear in BCube. For a cluster of around 250 servers, BCube achieves 53%\sim 70% lower GST than Fat-Tree from the simulation. As a result, we suggest using server-centric network topology such as BCube, instead of the common Fat-Tree network, to build a special-purpose DML cluster, due to its parallel synchronization, RDMA friendliness, natural load balance, as well as low economical cost.
Shuai Wang 0028, Dan Li 0001, Jinkun Geng
INFOCOM3
2019 Sphinx: A Transport Protocol for High-Speed and Lossy Mobile Networks
abstract
Modern mobile wireless networks have been demonstrated to be high-speed but lossy, while mobile applications have more strict requirements including reliability, goodput guarantee, bandwidth efficiency, and computation efficiency. Such a complicated combination of requirements and conditions in networks pushes the pressure to transport layer protocol design. We analyze and argue that few of existing network transport layer solutions are able to handle all these requirements. We design and implement Sphinx to satisfy the four requirements in high-speed and lossy networks. Sphinx has (1) a proactive coding-based method named semi-random LT codes for loss recovery, which estimates packet loss rate and adjusts the redundancy level accordingly, (2) a reactive retransmission method named Instantaneous Compensation Mechanism (ICM) for loss retransmission, which compensates the lost packets once actual loss exceeds the estimation, and (3) a parallel coding architecture, which leverages multi-core, shared memory and kernel-bypass DPDK. Prototype and evaluation show that Sphinx outperforms TCP and other coding solutions significantly in microbenchmarks across all four requirements, and improves the performance of applications such as video streaming and block data transfer.
Dan Li 0001, Wenfei Wu, K. K. Ramakrishnan, Jinkun Geng, Fei Gui, Fanzhao Wang, Kai Zheng 0003
IPCCC5
2019 HiPower: A High-Performance RDMA Acceleration Solution for Distributed Transaction Processing
Runhua Zhang 0002, Jinkun Geng, Shuai Wang 0028, Kaihui Gao, Guowei Shen
NPC3
2018 Dante: Enabling FOV-Aware Adaptive FEC Coding for 360-Degree Video Streaming
abstract
As 360-degree videos grow dramatically in popularity, more applications demand the ability to stream 360-degree videos to wirelessly connected devices, such as smartphone headsets. However, the limited capacity and the unstable network conditions make wireless networks ill-suited to the requirements of 360-degree videos--high resolution and low delay. One common approach is to take advantage of the fact that the viewer only watches a small portion of the video around the field of view (FOV). This allows for better allocation of network bandwidth by prioritizing content the viewer actually watches. Previous efforts on 360-degree videos have largely focused on adapting the encoded bitrate to optimize video quality in the time-varying FOV. This paper follows the general FOV-aware approach but uses a different technique. Rather than adapting bitrate, we explore the opportunities of a custom underlying transport protocol for 360-degree videos. In particular, we make a case for using Forward Error Correction (FEC) coding over UDP to reduce video streaming delay (a key limitation of all TCP-based approaches). We present Dante, an FOV-aware UDP-based video streaming protocol that adapts to changing network conditions by dynamically choosing FEC redundancy levels based on how close the video content is to the FOV region. Experimental results show that Dante improves video quality (PSNR) by 20% to 30% over traditional UDP-based video streaming protocols and 40% over FOV-aware DASH.
Zhetao Li, Fei Gui, Jinkun Geng, Dan Li 0001, Zhibo Wang 0001, Usama Zafar
APNet3
2018 BML: A High-performance, Low-cost Gradient Synchronization Algorithm for DML Training
abstract
In distributed machine learning (DML), the network performance between machines significantly impacts the speed of iterative training. In this paper we propose BML, a new gradient synchronization algorithm with higher network performance and lower network cost than the current practice. BML runs on BCube network, instead of using the traditional Fat-Tree topology. BML algorithm is designed in such a way that, compared to the parameter server (PS) algorithm on a Fat-Tree network connecting the same number of server machines, BML achieves theoretically 1/k of the gradient synchronization time, with k/5 of switches (the typical number of k is 2∼4). Experiments of LeNet-5 and VGG-19 benchmarks on a testbed with 9 dual-GPU servers show that, BML reduces the job completion time of DML training by up to 56.4%.
Dan Li 0001, Jinkun Geng, Yanshu Wang, Shuai Wang 0028, Shutao Xia
NeurIPS4
2018 Towards full virtualization of SDN infrastructure
Dan Li 0001, Yirong Yu, Jing Zhu 0007, Jinkun Geng
Comput. Networks6
2017 LOS: A High Performance and Compatible User-level Network Operating System
abstract
With the ever growing speed of Ethernet NIC and more and more CPU cores on commodity X86 servers, the processing capability of the network stack in Linux kernel has become the bottleneck. Recently there is a trend on moving the network stack up to user level and bypassing the kernel. However, most of these stacks require changing the APIs or modifying the source code of applications, and hence are difficult to support legacy applications. In this work, we design and develop LOS, a user-level network operating system that not only gains high throughput and low latency by kernel-bypass technologies but also achieves compatibility with legacy applications. We successfully run Nginx and NetPIPE on top of LOS without touching the source code, and the experimental results show that LOS achieves significant throughput and latency gains compared with Linux kernel.
Jinkun Geng, Du Lin, Ruilin Ling, Dan Li 0001
APNet2
2015 A Novel Clustering Algorithm for Database Anomaly Detection
Jinkun Geng, Daren Ye
SecureComm1