Guangyao Ding

dblp:241/7328 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
7since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 Hybrid Evaluation for Occlusion-based Explanations on CNN Inference Queries
abstract
Deep CNNs are increasingly prevalent in various application domains such as image processing. To explain a CNN prediction, it is popular to employ occlusion-based explanations (OBE). OBE helps users understand which parts of an image are important to a CNN prediction. Existing systems have explored incremental evaluation to accelerate CNN inference in OBE. However, they are oblivious that incremental evaluation does not always outperform full evaluation for certain layers. To address this issue, we propose a hybrid evaluation to efficiently interleave full and incremental evaluations during the CNN inference. Ad-ditionally, it employs a cost model to compare the overhead costs of two types of evaluations and a heuristic method to determine the efficient plan combination for common CNNs. More impor-tantly, hybrid evaluation adopts a dynamic programming-based method for attention-based CNNs. In particular, the dynamic programming-based method significantly reduces the overhead of searching for the efficient plan combination on the complex DAG structure. To demonstrate the efficiency of our techniques, we implement HyInJ, a hybrid CNN inf erence system based on PyTorch. Our experiments show that HyInf reduces execution time by up to 22% on GPU and 55% on CPU in comparison to the state-of-the-art incremental evaluation.
Guangyao Ding, Chen Xu 0001, Weining Qian
ICDE1
2024 CTS: Sim-to-Real Unsupervised Domain Adaptation on 3D Detection
abstract
Simulation data can be accurately labeled and have been expected to improve the performance of data-driven algorithms, including object detection. However, due to the various domain inconsistencies from simulation to reality (sim-to-real), cross-domain object detection algorithms usually suffer from dramatic performance drops. While numerous unsupervised domain adaptation (UDA) methods have been developed to address cross-domain tasks between real-world datasets, progress in sim-to-real remains limited. This paper presents a novel Complex-to-Simple (CTS) framework to transfer models from labeled simulation (source) to unlabeled reality (target) domains. Based on a two-stage detector, the novelty of this work is threefold: 1) developing fixed-size anchor heads and RoI augmentation to address size bias and feature diversity between two domains, thereby improving the quality of pseudo-label; 2) developing a novel corner-format representation of aleatoric uncertainty (AU) for the bounding box, to uniformly quantify pseudo-label quality; 3) developing a noise-aware mean teacher domain adaptation method based on AU, as well as object-level and frame-level sampling strategies, to migrate the impact of noisy labels. Experimental results demonstrate that our proposed approach significantly enhances the sim-to-real domain adaptation capability of 3D object detection models, outperforming state-of-the-art cross-domain algorithms, which are usually developed for real-to-real UDA tasks.
Meiying Zhang, Weiyuan Peng, Guangyao Ding, Chenyang Lei, Chunlin Ji, Qi Hao 0003
IROS3
2024 Joint Beamforming Design and Blocklength Optimization for Low-Latency Multiuser MISO URLLC Systems
abstract
To satisfy the requirements of many industrial applications, realizing ultrareliable low-latency communication (URLLC) has become one of the major challenges for future wireless networks. This article considers a downlink multiuser multiple-input-single-output (MISO) system in the Internet of Things (IoT) networks, in which a multiantenna base station (BS) serves multiple delay-sensitive IoT users, each equipped with a single antenna. To minimize the overall end-to-end delay, we jointly optimize the beamforming vectors and the packet blocklength to balance the queuing delay and the transmission delay. The problem is formulated as a Markov decision process (MDP), whose optimal solution can be theoretically found. However, the complexity on finding the optimal resource allocation and blocklength selection strategy is prohibitively high for real-system deployments due to the large state and action space. To overcome this issue, we simplify the original problem and develop an iterative algorithm to solve the simplified problem based on the uplink-downlink duality theory. Since solving the simplified problem would result in suboptimal solutions and may degrade the latency performance, we further develop a deep-reinforcement-learning (DRL)-based beamforming and blocklength selection framework to efficiently learn the optimal strategy of the original MDP. Simulation results demonstrate that the proposed algorithms can effectively improve the latency performance compared with the benchmark algorithm.
Guangyao Ding, Guanding Yu, Jiantao Yuan, Shengli Liu 0002
IEEE Internet Things J.1
2024 Joint URLLC Traffic Scheduling and Resource Allocation for Semantic Communication Systems
abstract
Recently, deep learning (DL) based semantic communication systems have shown great potential to improve transmission efficiency in various tasks. However, the coexisting mechanism between semantic communications and other services remains unexplored, which limits the application of semantic communications in practical communication systems. In this paper, we propose a dynamic multiplexing and co-scheduling scheme for the semantic and ultra-reliable low-latency communication (URLLC) traffic coexisting systems. In particular, a joint resource allocation and model training problem is formulated, which aims at maximizing the utility of semantic service while satisfying the latency requirement of URLLC traffic. To reduce the computational complexity, the original problem is simplified and decoupled into a joint resource allocation and model selection problem and a robust model training problem. In the resource allocation and model selection phase, the original problem is decomposed into three subproblems and an alternating optimization algorithm is then proposed to obtain the optimal resource allocation result. In the model training phase, a two-stage semantic communication network is designed, which can efficiently mitigate the impact of feature erasure brought by the random arrival of URLLC traffic. Simulation results show that the proposed method can effectively improve the quality of semantic service while satisfying the latency requirement of URLLC traffic.
Guangyao Ding, Shengli Liu 0002, Jiantao Yuan, Guanding Yu
IEEE Trans. Wirel. Commun.1
2022 JST: Joint Self-training for Unsupervised Domain Adaptation on 2D&3D Object Detection
abstract
2D&3D object detection always suffers from a dramatic performance drop when transferring the model trained in the source domain to the target domain due to various domain shifts. In this paper, we propose a Joint Self-Training (JST) framework to improve 2D image and 3D point cloud detectors with aligned outputs simultaneously during the transferring. The proposed framework contains three novelties to overcome object biases and unstable self-training processes: 1) an anchor scaling scheme is developed to efficiently eliminate the object size biases without any modification on point clouds; 2) a 2D&3D bounding box alignment method is proposed to generate high-quality pseudo labels for the self-training process; 3) a model smoothing based training strategy is developed to reduce the training oscillation properly. Experiment results show that the proposed approach improves the performance of 2D and 3D detectors in the target domain simultaneously; especially the superior accuracy of 3D detection can be achieved on benchmark datasets over the state-of-the-art methods.
Guangyao Ding, Meiying Zhang, E. Li, Qi Hao 0003
ICRA1
2022 Two-Timescale Resource Management for Ultrareliable and Low-Latency Vehicular Communications
abstract
Ultra-reliable low-latency communication (URLLC) is essential for future vehicle-to-vehicle (V2V) networks to improve traffic safety and enhance driving experience. Due to the fast-varying channel caused by high mobility, guaranteeing latency and reliability performance of the V2V links is a tremendous challenge. In this paper, we propose a novel resource allocation framework to support ultra-reliable low-latency V2V communications. The proposed framework includes both large-scale and small-scale resource optimizations. The large-scale resource allocation is performed at the central base station based on large-scale channel information periodically collected from vehicles. On the other hand, the small-scale resource allocation is performed at the vehicles according to instantaneous channel and queuing information. We develop optimal solutions for both resource allocation problems. With the proposed optimal solutions, the latency performance at the occurrence of extreme events is enhanced by enabling spectrum sharing among the vehicles. Simulation results demonstrate that the proposed algorithm can effectively improve the URLLC performance compared against the benchmark algorithm.
Guangyao Ding, Jiantao Yuan, Guanding Yu, Yuan Jiang 0008
IEEE Trans. Commun.1
2021 Accelerating DNN Training in Wireless Federated Edge Learning Systems
abstract
Training task in classical machine learning models, such as deep neural networks, is generally implemented at a remote cloud center for centralized learning, which is typically time-consuming and resource-hungry. It also incurs serious privacy issue and long communication latency since a large amount of data are transmitted to the centralized node. To overcome these shortcomings, we consider a newly-emerged framework, namely federated edge learning, to aggregate local learning updates at the network edge in lieu of users' raw data. Aiming at accelerating the training process, we first define a novel performance evaluation criterion, called learning efficiency. We then formulate a training acceleration optimization problem in the CPU scenario, where each user device is equipped with CPU. The closed-form expressions for joint batchsize selection and communication resource allocation are developed and some insightful results are highlighted. Further, we extend our learning framework to the GPU scenario. The optimal solution in this scenario is manifested to have the similar structure as that of the CPU scenario, recommending that our proposed algorithm is applicable in more general systems. Finally, extensive experiments validate the theoretical analysis and demonstrate that the proposed algorithm can reduce the training time and improve the learning accuracy simultaneously.
Jinke Ren, Guanding Yu, Guangyao Ding
IEEE J. Sel. Areas Commun.3
2019 IDFT-VFDM for LTE FDD-NR SUL Co-existence
abstract
In the paper, an inverse discrete Fourier transform-based Vandermonde-subspace frequency division multiplexing (IDFT-VFDM) waveform is proposed for the new radio (NR) supplementary uplink (SUL) to share the same time and frequency resources with the frequency division duplex (FDD) based long-term evolution (LTE) network. To avoid the co-channel interference to the LTE user equipment (UE) uplink transmission, the interference channel state information (CSI) is necessary for the NR UE to design the interference-free precoder. Since the operating band used for NR SUL corresponds to LTE FDD mode, the channel reciprocity condition in the time division duplex (TDD) mode is no longer held. To deal with it, the reciprocity on the channel related parameters for each path, i.e. amplitude, initial phase, propagation distance, angle of arrival, angle of departure, is exploited to estimate the uplink CSI from the NR UE to the LTE base station (BS) via the downlink CSI. Accordingly, the uplink waveform is designed for the NR UE to guarantee the absence of interference towards the LTE BS with the knowledge of uplink CSI. Numerical results are presented to validate the accuracy of the CSI estimation and the merit of the IDFT-VFDM as a potential waveform to achieve the LTE FDD-NR SUL co-existence.
Jiyong Pang, Yinghui He, Qiyu Hu, Guangyao Ding, Rui Yin 0001, Guanding Yu
PIMRC5