Jiangming Jin

dblp:56/9726 · DBLP profile ↗
← Back
16ranked-venue papers
1as first author
8since 2021 · last 2024
0000-0001-7552-6937ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 4 since 2021Computer networks · 5 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2024 Hashing-Based Multi-Modal Semantic Communication
abstract
The advanced sixth-generation (6G) wireless network is considered as an indispensable part of the Metaverse, where a substantial volume of communication content is transmitted through multiple modalities, placing significant transmission loads on communication channels. In this paper, we propose a framework for multi-modal semantic communication using hashing-based semantic extraction approach to produce optimal binary signatures (hash codes). Instead of directly using coarse-grained feature fusion methods, we capture deep semantics in self-attention manner, achieving fine-grained multi-modal feature fusion thereby strengthening the representation ability of hash codes. To enhance adaptability in practical situations, we then design a modality-completion module to address missing modalities in data, accommodating scenarios with both single-modal and cross-modal data. We evaluate the proposed semantic extraction framework on two popular multi-modal datasets, comparing it with the latest hashing methods and then demonstrate the effectiveness in various channel conditions.
Hongyu Gu, Jiangtian Nie, Jianhang Tang, Jiangming Jin, Yang Zhang 0025
WCNC5
2024 Diffusion-Model-Based Incentive Mechanism With Prospect Theory for Edge AIGC Services in 6G IoT
abstract
The fusion of the Internet of Things (IoT) with sixth-generation (6G) technology has significant potential to revolutionize the IoT landscape. With the ultrareliable and low-latency communication capabilities of 6G, 6G-IoT networks can transmit high-quality and diverse data to enhance edge learning. Artificial intelligence-generated content (AIGC) harnesses advanced artificial intelligence (AI) algorithms to automatically generate various types of content. The emergence of edge AIGC integrates with edge networks, facilitating real-time provision of customized AIGC services by deploying AIGC models on edge devices. However, the current practice of edge devices as AIGC service providers (ASPs) lacks incentives, hindering the sustainable provision of high-quality edge AIGC services amidst information asymmetry. In this article, we develop a user-centric incentive mechanism framework for edge AIGC services in 6G-IoT networks. Specifically, we first propose a contract theory model for incentivizing ASPs to provide AIGC services to clients. Recognizing the irrationality of clients toward personalized AIGC services, we utilize prospect theory (PT) to capture their subjective utility better. Furthermore, we adopt the diffusion-based soft actor-critic algorithm to generate the optimal contract design under PT, outperforming traditional deep reinforcement learning algorithms. Our numerical results demonstrate the effectiveness of the proposed scheme.
Jinbo Wen, Jiangtian Nie, Changyan Yi, Xiaohuan Li 0001, Jiangming Jin, Yang Zhang 0025, Dusit Niyato
IEEE Internet Things J.6
2023 Privacy-Aware Double Auction With Time-Dependent Valuation for Blockchain-Based Dynamic Spectrum Sharing in IoT Systems
abstract
For future Internet of Things (IoT) systems, data-driven and dynamic spectrum-sharing schemes can significantly improve the spectrum utilization and efficiency. However, conventional centralized architecture of such dynamic IoT spectrum-sharing systems is often considered to be nontransparent, costly, and vulnerable to potential attacks and single-point failures. To address the aforementioned issues, a blockchain-based dynamic spectrum-sharing scheme has been proposed and investigated in this work, which aims at enhancing the system by providing desirable features, such as decentralization, transparency, immutability, and auditability. By considering the privacy and transaction dynamics issues when blockchain is integrated into spectrum-sharing systems, a privacy-preserving double auction mechanism based on differential privacy is developed for incentivizing spectrum sharing, where the time-varying valuations of the spectrum resources are also taken into consideration. In the proposed auction, a winner determination problem (WDP) is formulated to decide the winning bidders and spectrum allocation. A deep reinforcement learning (DRL)-based method is then proposed for efficiently solving the WDP. The proposed auction mechanism can be integrated with smart contracts on blockchain platforms. Furthermore, the computation of the DRL-based method for solving the WDP is designed as part of the consensus mechanism in the blockchain. Theoretical analysis show that the proposed privacy-aware double auction mechanism satisfies the properties of differential privacy, individual rationality, and truthfulness. Finally, simulation results are provided to validate the performance of the spectrum-sharing approach.
Kun Zhu 0001, Lu Huang 0001, Jiangtian Nie, Yang Zhang 0025, Zehui Xiong, Hongning Dai, Jiangming Jin
IEEE Internet Things J.7
2022 TCUDA: A QoS-based GPU Sharing Framework for Autonomous Navigation Systems
abstract
Autonomous navigation systems (ANS) consist of several software modules, such as sensing, perception, and planning to achieve traffic perception and fast decision making. These modules are required to process large amounts of data, such as images, in real-time. GPUs are commonly exploited on ANS to speed up data processing. As GPUs provides tremendous computation resources, it is common that one software module uses only partial GPU resources, leading to low GPU utilization and energy inefficiency. Traditionally, GPU sharing is a method to address this problem. However, GPU sharing is not supported on current embedded GPU platforms, which are widely used by ANS. Furthermore, GPU sharing challenges task Quality of Service (QoS) that requires task execution in a fixed latency. To achieve QoS requirement, we propose a progress bar scheduling policy to provide GPU tasks with QoS guarantee in GPU sharing environments. Based on this policy, a GPU sharing framework named TCUDA, is proposed to endow existing GPU tasks with QoS guarantee and improve GPU utilization. Finally, results show TCUDA reduces GPU task latency by 16.8% with QoS guarantee on an embedded GPU platform Xavier.
Pangbo Sun, Hao Wu 0012, Jiangming Jin, Ziyue Jiang 0002, Yifan Gong 0003
SBAC-PAD3
2022 Zoro: A robotic middleware combining high performance and high reliability
Wei Liu 0143, Jiangming Jin, Hao Wu 0012, Yifan Gong 0003, Ziyue Jiang 0002, Jidong Zhai
J. Parallel Distributed Comput.2
2022 Decentralized Edge Intelligence: A Dynamic Resource Allocation Framework for Hierarchical Federated Learning
abstract
To enable the large scale and efficient deployment of Artificial Intelligence (AI), the confluence of AI and Edge Computing has given rise to Edge Intelligence, which leverages on the computation and communication capabilities of end devices and edge servers to process data closer to where it is produced. One of the enabling technologies of Edge Intelligence is the privacy preserving machine learning paradigm known as Federated Learning (FL), which enables data owners to conduct model training without having to transmit their raw data to third-party servers. However, the FL network is envisioned to involve thousands of heterogeneous distributed devices. As a result, communication inefficiency remains a key bottleneck. To reduce node failures and device dropouts, the Hierarchical Federated Learning (HFL) framework has been proposed whereby cluster heads are designated to support the data owners through intermediate model aggregation. This decentralized learning approach reduces the reliance on a central controller, e.g., the model owner. However, the issues of resource allocation and incentive design are not well-studied in the HFL framework. In this article, we consider a two-level resource allocation and incentive mechanism design problem. In the lower level, the cluster heads offer rewards in exchange for the data owners' participation, and the data owners are free to choose which cluster to join. Specifically, we apply the evolutionary game theory to model the dynamics of the cluster selection process. In the upper level, each cluster head can choose to serve a model owner, whereas the model owners have to compete amongst each other for the services of the cluster heads. As such, we propose a deep learning based auction mechanism to derive the valuation of each cluster head's services. The performance evaluation shows the uniqueness and stability of our proposed evolutionary game, as well as the revenue maximizing properties of the deep learning based auction.
Wei Yang Bryan Lim, Jer Shyuan Ng, Zehui Xiong, Jiangming Jin, Yang Zhang 0025, Dusit Niyato, Cyril Leung, Chunyan Miao
IEEE Trans. Parallel Distributed Syst.4
2022 Reputation-Aware Hedonic Coalition Formation for Efficient Serverless Hierarchical Federated Learning
abstract
Amid growing concerns on data privacy, Federated Learning (FL) has emerged as a promising privacy preserving distributed machine learning paradigm. Given that the FL network is expected to be implemented at scale, several studies have proposed system architectures towards improving the network scalability and efficiency. Specifically, the Hierarchical FL (HFL) network utilizes cluster heads, e.g., base stations, for the intermediate aggregation and relay of model parameters. Serverless FL is also proposed recently, in which the data owners, i.e., workers, exchange the local model parameters among a neighborhood of workers. This decentralized approach reduces the risk of a single point of failure but inevitably incurs significant communication overheads. To achieve the best of both worlds, we propose the Serverless Hierarchical Federated Learning (SHFL) framework in this paper. The SHFL framework adopts a two-layer system architecture. In the lower layer, the FL workers are grouped into clusters under cluster heads. In the upper layer, the cluster heads exchange the intermediate parameters with their one-hop neighbors without the aid of a central server. To improve the sustainable efficiency of the FL system while taking into account the incentive design for workers marginal contributions in the system, we propose the reputation-aware hedonic coalition formation game in this paper. Specifically, the workers are rewarded for their marginal contribution to the cluster, whereas the reputation opinions of each cluster head is updated in a decentralized manner, thereby deterring malicious behaviors by the cluster head. This improves the performance of the network since cluster heads with higher reputation scores are more reliable in relaying the intermediate model parameters. The simulation results show that our proposed hedonic coalition formation algorithm converges to a Nash-stable partition and improves the network efficiency.
Jer Shyuan Ng, Wei Yang Bryan Lim, Zehui Xiong, Xianbin Cao 0001, Jiangming Jin, Dusit Niyato, Cyril Leung, Chunyan Miao
IEEE Trans. Parallel Distributed Syst.5
2021 Accelerating GPU Message Communication for Autonomous Navigation Systems
abstract
Autonomous navigation systems consist of multiple software modules, such as sensing, object detection, and planning, to achieve traffic perception and fast decision making. Such a system generates a large amount of data and requires data processing and communication in real-time. Although accelerators, such as GPUs, have been exploited to speed up data processing, communicating GPU messages between modules is still lacking support, leading to high communication latency and resource contention. For such a latency-sensitive and resource-limited autonomous navigation system, high performance and lightweight message communication are crucial and demanding. To obtain both high performance and low resource usages, we first propose a novel pub-centric memory pool and an on-the-fly offset conversion algorithm to avoid unnecessary data movement. Secondly, we combine these two techniques and propose an efficient message communication on a single GPU. Finally, we extend this approach to multi-GPU and design a framework that natively supports GPU message communication for Inter-Process Communication. With comprehensive evaluation, results show our approach is able to reduce communication latency by 53.7% for PointCloud and Image messages compared to the state-of-the-art approach. Moreover, in the real autonomous navigation scenario, our approach reduces the end-to-end latency by 29.2% and decreases resource usage up to 58.9%.
Hao Wu 0012, Jiangming Jin, Jidong Zhai, Yifan Gong 0003, Wei Liu 0143
CLUSTER2
2020 Safe Process Quitting for GPU Multi-Process Service (MPS)
abstract
GPUs have been widely adopted to speedup various throughput-originated applications running on HPC platforms, where typically there are a number of tasks sharing GPUs to maximize GPU utilization. To facilitate GPU sharing, GPU vendors provide tools, allowing multiple processes concurrently to use GPUs. For example, Nvidia provides MPS (Multi-Process Service) managing all GPU processes to achieve high throughput by fully exploiting hardware resources. However, such tool leads to undesired single point of failure for all GPU processes, namely, one process’s exception makes other processes abnormal. In this work, we investigate the seriousness of this GPU process interferences caused by MPS, and propose an approach to address one of these interferences, which takes place during process quitting. By using signal handling and thread synchronization techniques in this approach, GPU processes are able to quit safely without interfering other GPU processes.
Hao Wu 0012, Wei Liu 0143, Yifan Gong 0003, Jiangming Jin
ICDCS4
2020 Memory-Centric Communication Mechanism for Real-time Autonomous Navigation Applications
abstract
There has been a remarkable increase in the speed of AI development over the past few years. Artificial intelligence and deep learning techniques are blooming and expanding in all forms to every sector possible. With the emerging intelligent autonomous navigation systems, both memory allocation and data movement are becoming the main bottlenecks in inter-process communication procedures, especially in supporting various types of messages between multiple programming languages. To reduce significant memory allocation and data movement cost, we propose a novel memory-centric mechanism, which includes a virtual layer based architecture and a pre-record memory allocation algorithm. Furthermore, we implement a memory-centric communication framework named Z-framework based on the proposed mechanism to achieve high efficient IPC procedures in autonomous navigation systems. Experimental results show that Z-framework is able to gain up to 41% and 35% performance improvement compared with the approach used in ROS2, which is an industry standard and the state-of-the-art approach used in CyberRT, respectively.
Wei Liu 0143, Yifan Gong 0003, Hao Wu 0012, Jidong Zhai, Jiangming Jin
ICPP5
2020 Poster: A Light Weight Service Discovery Mechanism in Robot Systems
abstract
Benefiting from the major breakthrough of AI technology and increasing application of AIoT (AI + IoT), the intelligent robot system achieves high developing speed in recent years. As a core component of the intelligent system, service discovery plays a crucial role in system reliability. Because of the high CPU usages in conventional service discovery, a light weight service discovery mechanism is required in such a resourcelimited robot system. To improve the efficiency of CPU usages, we propose a weak centralized mechanism and a socket based notification mechanism, to reduce the amount of event in service discovery. The evaluation results show that our proposed light weight service discovery mechanism can reduce 95% CPU usages in average, compared with the conventional service discovery used in ROS2, which is an industry standard in robot systems.
Yifan Gong 0003, Wei Liu 0143, Jiangming Jin
SEC4
2020 A Robotic Communication Middleware Combining High Performance and High Reliability
abstract
With the significant advances of AI technology, intelligent robotic systems have achieved remarkable development and profound effects. To enable massive data transmissionin an efficient and reliable way, both high performance andhigh reliability should be taken into account in system design. However, the conventional communication middleware used in the majority of autonomous robotic systems, is based on socked-based methods, which always lead to high latency. Moreover, some sophisticated communication middleware utilizes shared memory upon ring buffers for high performance without consideration of the reliability. To obtain both high performance and high reliability, we employ shared memory for performance improvement and propose a novel socket-based communication control algorithm to improve reliability during data transmission. Furthermore, based on the proposed algorithm, we implement a novel robotic communication middleware, named Robust-Z, combining both high performance and high reliability. Experimental results show that (1) Robust-Z is able to gain up to 41% and 5% performance improvement compared to ROS2 and Apollo CyberRT, respectively; (2) Robust-Z is able to provide crash safety and reduce 5.2% data missing rate compared with CyberRT.
Wei Liu 0143, Hao Wu 0012, Ziyue Jiang 0002, Yifan Gong 0003, Jiangming Jin
SBAC-PAD5
2018 BitFlow: Exploiting Vector Parallelism for Binary Neural Networks on CPU
abstract
Deep learning has revolutionized computer vision and other fields since its big bang in 2012. However, it is challenging to deploy Deep Neural Networks (DNNs) into real-world applications due to their high computational complexity. Binary Neural Networks (BNNs) dramatically reduce computational complexity by replacing most arithmetic operations with bitwise operations. Existing implementations of BNNs have been focusing on GPU or FPGA, and using the conventional image-to-column method that doesn't perform well for binary convolution due to low arithmetic intensity and unfriendly pattern for bitwise operations. We propose BitFlow, a gemm-operator-network three-level optimization framework for fully exploiting the computing power of BNNs on CPU. BitFlow features a new class of algorithm named PressedConv for efficient binary convolution using locality-aware layout and vector parallelism. We evaluate BitFlow with the VGG network. On a single core of Intel Xeon Phi, BitFlow obtains 1.8x speedup over unoptimized BNN implementations, and 11.5x speedup over counterpart full-precision DNNs. Over 64 cores, BitFlow enables BNNs to run 1.1x faster than counterpart full-precision DNNs on GPU (GTX 1080).
Jidong Zhai, Dinghua Li, Yifan Gong 0003, Yuhao Zhu 0001, Wei Liu 0143, Jiangming Jin
IPDPS8
2018 Joint optimization of information trading in Internet of Things (IoT) market with externalities
abstract
Internet of Things (IoT) technology enables various physical devices to collect, process and exchange information. Market oriented models become important for IoT systems to efficiently utilize information, as IoT network nodes operate in a highly distributed and autonomous manner. In this work, we propose a three-player game theoretic market model for IoT information trading, considering direct and indirect externalities among market participants. In the model, an IoT service provider collects and processes IoT information, and then delivers the processed information as IoT services to IoT users. Then, an IoT content vendor senses and generates raw information for the IoT service provider to collect, and receives rewards from the provider. Finally, an IoT user pays a fixed service fee to the IoT service provider to access the IoT services. To jointly derive the optimal market decisions of the three participants in the model, we employ a Stackelberg game approach. The equilibria are obtained as the closed form solutions of the game, with which the existence and uniqueness properties are proved. The analytical results show that the IoT service provider operates as an intermediary agent between the IoT content vendor and users, reducing the information trading complexity of both user and vendor sides.
Yang Zhang 0025, Zehui Xiong, Dusit Niyato, Ping Wang 0001, Jiangming Jin
WCNC5
2017 A Game-Theoretic Analysis of Complementarity, Substitutability and Externalities in Cloud Services
abstract
In cloud computing, cloud services can be allocated to users upon requests in an on-demand basis. Heterogeneous cloud service providers may join the cloud systems to serve various types of users. Cloud services can be complementary or substitutable. For the complementary services, users may request for a bundle of the services, e.g., CPU and storage, to gain higher benefit from requesting them alone. The substitutable services have similar functionalities to serve users, e.g., different cloud database services, obtaining one of them can replace another one. Furthermore, the users of the cloud systems also influence each other because of externalities, particularly, network effect and congestion effect. From the perspective of each user, the existence of other users may introduce positive or negative impacts on the user utility, in the case of network and congestion effects, respectively. In this work, the participants in the cloud systems are treated as social enabled rational individuals. We model the complementarity, substitutability and externalities in cloud services by employing a multiple-leader multiple- follower Stackelberg game approach, including a two-stage service transaction process where service providers and users make their transaction decisions in a distributed manner. The analytical expressions of equilibria, service pricing strategies, and service allocations are derived with numerical results. We also find in the numerical results that both collusive and competitive service pricing schemes may lead to the optimized provider and user performances simultaneously.
Yang Zhang 0025, Zehui Xiong, Dusit Niyato, Ping Wang 0001, Jiangming Jin
GLOBECOM5
2013 Simulation of Information Propagation over Complex Networks: Performance Studies on Multi-GPU
abstract
General Purpose Graphics Processing Units (GPGPU) have been used in high performance computing platforms to accelerate the performance of scientific applications such as simulations. With the increased computing resources required for large-scale network simulation, one GPU device may not have enough memory and computation capacities. It is therefore necessary to enhance the system scalability by introducing multiple GPU devices. It is also attractive to investigate the performance scalability of Multi-GPU simulations. This paper describes the simulation of information propagation on multiple GPU devices, including the optimized network simulation algorithms, the network partitioning and replication strategy, and the data synchronization scheme. The experimental results for scalable random networks show that the number of simulation steps, computation time, synchronization time, and data transfer time all affect the overall simulation performance. In order to compare with random networks, we also conduct simulations of scale-free networks. We can observe that the node replication ratio in scale-free networks is smaller than that in random networks and therefore the cost of data transfer and synchronization is significantly reduced. This indicates that the network structure is also an important factor that influences the simulation performance in a Multi-GPU system.
Jiangming Jin, Stephen John Turner, Bu-Sung Lee, Jianlong Zhong, Bingsheng He
DS-RT1