EDBT 2026 Demo / reviewers in the wild / expert
Yongdong Zhu
dblp:295/2921
· DBLP profile ↗
18ranked-venue papers
0as first author
18since 2021 · last 2025
0000-0002-5420-7926ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Computer networks · 7 · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | 6G autonomous radio access network empowered by artificial intelligence and network digital twinabstractAbstract The sixth-generation (6G) mobile network implements the social vision of digital twins and ubiquitous intelligence. Contrary to the fifth-generation (5G) mobile network that focuses only on communications, 6G mobile networks must natively support new capabilities such as sensing, computing, artificial intelligence (AI), big data, and security while facilitating Everything as a Service. Although 5G mobile network deployment has demonstrated that network automation and intelligence can simplify network operation and maintenance (O&M), the addition of external functionalities has resulted in low service efficiency and high operational costs. In this study, a technology framework for a 6G autonomous radio access network (RAN) is proposed to achieve a high-level network autonomy that embraces the design of native cloud, native AI, and network digital twin (NDT). First, a service-based architecture is proposed to re-architect the protocol stack of RAN, which flexibly orchestrates the services and functions on demand as well as customizes them into cloud-native services. Second, a native AI framework is structured to provide AI support for the diverse use cases of network O&M by orchestrating communications, AI models, data, and computing power demanded by AI use cases. Third, a digital twin network is developed as a virtual environment for the training, pre-validation, and tuning of AI algorithms and neural networks, avoiding possible unexpected losses of the network O&M caused by AI applications. The combination of native AI and NDT can facilitate network autonomy by building closed-loop management and optimization for RAN. Guangyi Liu 0001, Juan Deng, Yanhong Zhu, Boxiao Han, Shoufeng Wang, Hua Rui, Jingyu Wang 0001, Jianhua Zhang 0001, Ying Cui 0001, Yingping Cui, Yang Yang 0001, Jiangzhou Wang, Ye Ouyang, Xiaozhou Ye, Tao Chen 0011, Rongpeng Li, Yongdong Zhu, Sen Bian, Wanfei Sun, Qingbi Zheng, Zhou Tong, Zecai Shao, Jiajun Wu 0021, Mancong Kang |
Frontiers Inf. Technol. Electron. Eng. | 19 |
| 2025 | OWRT-DETR: A Novel Real-Time Transformer Network for Small-Object Detection in Open-Water Search and Rescue From UAV Aerial ImageryabstractUAV object detection is crucial in open water search and rescue missions. Due to varying perspectives and altitudes of UAV images, the apparent size of objects varies significantly. Challenges such as insufficient feature representation and background confusion make open water object detection particularly difficult. Currently, deep learning-based detection methods rely on convolution to extract features at a fixed spatial scale. This limited receptive field leads to insufficient feature representation, causing false detections and missed detections, which severely impact detection accuracy. This paper proposes an efficient, feature-enhanced real-time detection network based on transformer architecture, called OWRT-DETR, to address the challenges of diverse UAV image detection in open water. To the best of our knowledge, a Transformer-based detection network has not yet been explored for open water UAV images. OWRT-DETR incorporates a cross-scale feature pyramid module (CFPIM), multi scale sensing fusion (MSSF), and small object enhancement module (SOEM). These modules enhance cross-scale interaction, cross-channel spatial global association, and local perception of the network, while avoiding increased complexity, improving weak feature representation of small targets, and suppressing easily confused backgrounds. Three public datasets are used to validate the effectiveness of OWRT-DETR. OWRT-DETR achieves an averaged precision (AP) of 51.5%, 45.9%, 50.6% on the SeaDronesSee, Aerial Dataset of Floating Objects, and Aerialbus Ship datasets, exceeding the performance of several state-of-the-art models. To ensure efficiency and reduce computational resources, OWRT-DETR is optimized by reconstructing the backbone network using PConv and Rep, resulting in Light-OWRT-DETR. Compared with OWRT-DETR, Light-OWRT-DETR is faster, uses fewer parameters, requires less computational power, and achieves higher accuracy. The code will be available at https://github.com/mshauima/OWRT-DETR. Yihong Zhang 0002, Baolong Ding, Yongdong Zhu |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Real-Time Driving Style Integration in Deep Reinforcement Learning for Traffic Signal ControlabstractMost existing reinforcement learning-based traffic signal control approaches overlook vehicle-specific information in the state representation. This study addresses this gap by integrating real-time driving style information into the deep reinforcement learning (DRL) framework. We introduce a model-based framework that captures real-time driving styles and converts them into Intelligent Driver Model (IDM) parameters. Our proposed method demonstrates superior performance across various reinforcement algorithms and traffic flow scenarios, with statistical tests confirming a significant reduction in average queue length. The contributions of this paper can be summarized as follows: 1) proposing a model-based method for real-time driving style recognition, significantly reducing the requirements for trajectory data duration and computational resources, and 2) proposing a new state variable called transformed occupancy (o*) that allows the DRL-based traffic signal controller to be trained with driving style information, thereby enhancing the performance of the traffic signal control system. The proposed framework is so flexible that other car-following models, machine learning algorithms, and various downstream tasks can be incorporated. Tu Xu, Yuqi Pang, Yongdong Zhu, Rui Jiang 0008 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | HDA-IDS: A Hybrid DoS Attacks Intrusion Detection System for IoT by using semi-supervised CL-GAN
Yue Cao 0002, Shuohan Liu, Yuping Lai, Yongdong Zhu, Naveed Ahmad 0003 |
Expert Syst. Appl. | 5 |
| 2024 | Outage Analysis of IRS-Assisted UAV NOMA Downlink Wireless NetworksabstractThis article studies an intelligent reflecting surface (IRS)-assisted unmanned aerial vehicle (UAV) network, where the ground users (GUs) desire to receive information from a UAV. Downlink nonorthogonal multiple access (NOMA) is considered typically with two GUs being selected according to whether a Line-of-Sight (LoS) link between GUs and UAV exists. As the accurate channel information of LoS or Non-LoS (NLoS) links for multiple GUs is difficult to acquire, an approximate LoS region-based method is designed to select GUs as an alternative. In order to enhance the communication quality of the far GU, an IRS is deployed to assist the NLoS transmission. For such a system, we evaluate its outage performance in Nakagami-m fading. First, the central limit theorem (CLT) and Laplace transform (LT) are employed to derive the channel statistics of the UAV- IRS-user link. Then, asymptotic closed-form expressions of the outage probabilities are derived for the selected GUs based on Gaussian–Chebyshev quadrature approximation. Monte Carlo simulations validate the validness of our derived outage probabilities. It shows that the approximate LoS region-based scheme provides similar outage performance laws as the accurate LoS region-based one. Moreover, the outage probabilities of selected GUs in terms of NOMA-based protocol and orthogonal multiple access (OMA)-based protocol are analyzed. Simulation results confirm that the proposed NOMA-based protocol is capable of achieving superior performance compared with the OMA-based protocol by setting power allocation factor and targeted acrlong SINR thresholds of near GU and far GU properly. Specifically, when the rate threshold of near GU is relatively large or the rate threshold of far GU is relatively small, the outage performance derived by NOMA-based protocol performs better than OMA-based protocol in most of cases. Yuan Liu 0030, Ke Xiong 0001, Yongdong Zhu, Hong-Chuan Yang, Pingyi Fan, Khaled Ben Letaief |
IEEE Internet Things J. | 3 |
| 2024 | Maximizing Data Collection and Rental Requests in Drone-Based IIoT NetworksabstractMany industries now rely on drones to monitor infrastructures. In this respect, this article considers maximizing the revenue of an Industrial Internet of Things operator that provides two services: 1) data trading; and 2) drones rental. In service 1), the operator sells data of locations/points it acquired via drones. For service 2), it rents idle drones to users. The problem at hand is to determine the allocation of drones to services 1) and 2) that maximizes the operator's revenue over a given planning horizon. We outline a novel integer linear program (ILP) to solve the said problem, which can be used to determine the optimal number of drones assigned to both services. The ILP, however, requires an exhaustive collection of drone trajectories. We therefore present two heuristics called weighted-based algorithm (WBA) and genetic algorithm (GA) to generate trajectories for data collection. The results show that WBA earns 95.6% of the optimal revenue. GA is able to achieve 99% of the revenue of WBA at best. Chuyu Li, Kwan-Wu Chin, Yongdong Zhu |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | Efficient Revocable Anonymous Authentication Mechanism for Edge Intelligent ControllersabstractIn the field of industrial control, edge computing puts forward higher requirements for the local processing ability of data. The traditional industrial programmable logic controller (PLC) cannot complete this task. The edge intelligent controller (EIC) is developed according to the demand for edge computing. As the core component of edge computing, the security and reliable operation of EIC has great significance and influence on the promotion and development of edge computing. For the access authentication of EIC, an efficient revocable anonymous authentication scheme for EICs by utilizing a group signature technique is proposed in this article, which can trace the identity information of the EIC and protect the privacy of the EIC, and apply an efficient revocation mechanism to revoke any illegal or abnormal EICs. Through security analysis, we prove that the authentication scheme is anonymity, integrity, traceability, forward secrecy, resistance to the replay attack, and efficient revocability. Furthermore, the performance analysis and comparisons show that the authentication scheme is efficient and feasible for EICs. Zhong Cao 0002, Yongdong Zhu |
IEEE Internet Things J. | 4 |
| 2023 | Multiagent Meta-Reinforcement Learning for Optimized Task Scheduling in Heterogeneous Edge Computing SystemsabstractMobile-edge computing (MEC) brings the potential to address the ever increasing computation demands from the mobile users (MUs). In addition to local processing, the resource-constrained MUs in an MEC system can also offload computation to the nearby servers for remote execution. With the explosive growth of mobile devices, computation offloading faces the challenge of spectrum congestion, which, in turn, deteriorates the overall quality of computation experience. This article, hence, investigates computation task scheduling in a heterogeneous cellular and WiFi MEC system. Such a system provides both licensed and unlicensed spectrum opportunities. Due to the sharing of communication and computation resources as well as the uncertainties, we formulate the problem of computation task scheduling among the competing MUs in a stationary heterogeneous edge computing system as a noncooperative stochastic game. We propose an approximation-based multiagent Markov decision process without the global system state observations, under which a multiagent proximal policy optimization (PPO) algorithm is derived to solve the corresponding Nash equilibrium. When expanding to a nonstationary heterogeneous edge computing system, the obtained algorithm suffers from the slow convergence due to constrained adaptability. Accordingly, we explore meta-learning and propose a multiagent meta-PPO algorithm, which rapidly adapts the control policy learning to the nonstationarity. Numerical experiments demonstrate performance gains from our proposed algorithms. Liwen Niu, Xianfu Chen, Ning Zhang 0007, Yongdong Zhu, Rui Yin 0001, Celimuge Wu, Yangjie Cao |
IEEE Internet Things J. | 4 |
| 2023 | Autonomous valet parking optimization with two-step reservation and pricing strategy
Ziyi Hu, Yue Cao 0002, Yongdong Zhu, Naveed Ahmad 0003 |
J. Netw. Comput. Appl. | 4 |
| 2023 | Predicting Urban Region Heat via Learning Arrive-Stay-Leave Behaviors of Private CarsabstractUrban region heat refers to the extent of which people congregate in various regions when they travel to and stay in a specified place. Predicting urban region heat facilitates broad applications ranging from location-based services to intelligent transportation management. The region heat is essentially characterized by the ‘arrive-stay-leave (ASL)’ behaviors, while it is a challenging task to well capture the spatial-temporal evolution of region heat since the following issues remain: i) ASL behaviors of private cars is usually heterogeneous resulting in a hierarchical distribution of region heat. ii) Urban region heat contains complex spatial-temporal correlations hidden in ASL behaviors and how to collaboratively integrate them is challenging. To address these challenges, we propose a Hierarchical Spatial-Temporal Network (HierSTNet) to forecast urban region heat, which contains two representations, namely, grid region from micro perspective and node region from macro perspective. For the grids, three-dimension spatial and temporal convolutional network (3D-STCNN) is proposed to model multi-scale properties in temporal dimension of ASL behaviors. For the nodes, multi-head graph attention networks are utilized to model the periodicity and spatial heterogeneity among macro region. Hierarchical structures are designed for multi-view modeling spatial-temporal distribution of ASL behaviors, by which they capture small-scale features in micro regions and embeds the global representation into graph propagation. Finally, we design an interaction decoder layer to integrate the external factors and aggregate spatial-temporal information across hierarchical structures. Extensive experiments based on real-world private car trajectory dataset demonstrate the superiority and effectiveness of proposed framework. Zhu Xiao, Hongbo Jiang 0001, Mamoun Alazab, Yongdong Zhu, Schahram Dustdar |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | Trans-RL: A Prediction-Control Approach for QoE-Aware Point Cloud Video StreamingabstractIn point cloud video streaming systems, the field of view (FoV) prediction is critical for selecting the tiles, the objective of which is to optimize the expected long-term quality-of-experience (QoE) from the perspective of a user. On one hand, a satisfactory QoE accounts for not only the playback quality but also the playback smoothness. On the other hand, the large data volume of a selected tile requires the transmission to be adaptive to the system uncertainties. This paper applies a Markov decision process to formulate the problem of tile selection across the infinite discrete time horizon. In particular, a system state includes the FoV information, which is predicted from the Transformer. To alleviate the dependence on system uncertainty statistics, a deep reinforcement learning approach is derived for solving the optimal control policy. Under different settings, we conduct experiments based on the real throughput and head-mounted display data. The results show that compared to the existing baselines, our proposed prediction-control approach achieves a higher FoV prediction accuracy, better playback quality as well as smoothness, and hence a better average QoE for the user. Cunhui Zhang, Yangjie Cao, Zhi Liu 0002, Rui Yin 0001, Yongdong Zhu, Xianfu Chen |
GLOBECOM | 5 |
| 2022 | Towards Event-driven Misbehavior Detection Mechanism in Social Internet of VehiclesabstractDue to inadequate management of Vehicular Ad hoc Networks (VANETs), malicious nodes could participate in communications along with misbehavior, e.g., dropping packets and spreading fake information. Therefore, it is essential to detect misbehavior of internal attackers that will cause network performance degradation (e.g., taking longer time to receive messages or reaching destinations with detours). Apart from the capture of dynamic network topology of VANETs, the social relationship among nodes can also be applied as a relatively stable metric to qualify nodes. This paper proposes a misbehavior detection mechanism based on social relationships, from which nodes determine trust for the receiver or transmitter. Based on the proposed mechanism, road traffic control applications can avoid the interference from malicious nodes. The construction of social relationships depends on the geographic information reflected by the movement of nodes, including contact frequency and trajectory similarity, since the geographic information can accurately indicate the relevance among nodes. In addition to the social relationship, the proposed mechanism also evaluates the data trust from time and spatial factors to reduce the interference of fake data. Finally, the proposed mechanism integrates data trust and social relationships to enable misbehavior detection decisions. Extensive results of simulations show that the proposed mechanism has outstanding malicious nodes detection rates under various proportions of malicious nodes and movement patterns. Chenchen Lv, Yue Cao 0002, Lexi Xu, Shitao Zou, Yongdong Zhu, Zhili Sun |
MSN | 5 |
| 2022 | Few-Shot Hyperspectral Image Classification Based on Adaptive Subspaces and Feature TransformationabstractIn the field of hyperspectral image (HSI) classification, deep learning has helped achieve great successes. However, most of these achievements are made with very large amounts of labeled training data. Manual annotation of HSIs is labor intensive and time consuming. In practical HSI classification, there may only be a few labeled samples available. To perform HSI classification with a small number of labeled samples, a new few-shot classification model based on adaptive subspaces and featurewise transformation is proposed in this article. First, we design a 3-D local channel attention residual network to obtain the spatial–spectral features of HSIs. Then, a featurewise transformation strategy is introduced to enhance feature diversity to avoid model overfitting problems and to mitigate the impact of cross-domain problems. Finally, a subspace classifier is implemented to construct different subspace categories based on the embedded features of the limited labeled samples. Classification of an HSI sample is performed using spatial projection and a distance metric. The proposed model is trained using the metalearning mechanism to perform few-shot classification of HSIs. Four public datasets are utilized to construct a sufficient few-shot classification task named episodes for training. The other three public datasets are used to test the proposed model. Experiments show that our proposed method can outperform mainstream small sample HSI classification methods. Jing Bai 0003, Shaojie Huang, Zhu Xiao, Xianmin Li, Yongdong Zhu, Amelia Regan, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Hyperspectral Image Classification Based on Superpixel Feature Subdivision and Adaptive Graph StructureabstractThe graph-based hyperspectral image classification (HSIC) method has attracted wide attention because it can extract information with a non-Euclidean structure. Many graph-based HSIC works have achieved good results, but unresolved technical issues remain. For example, many graph nodes lead to high computational costs, and the mining of non-Euclidean structures is not sufficient. To solve these problems, we propose a graph attention network with an adaptive graph structure mining (GAT-AGSM) approach. Specifically, we first propose an HSIC framework with a superpixel feature subdivision (SFS) mechanism. In this framework, the number of nodes in the graph structure is reduced by using superpixel segmentation algorithms, and the SFS mechanism is designed to generate finer classification results. Second, we design the spatial–spectral attention layer with an adaptive graph structure mining (AGSM) mechanism for the graph attention network. The spatial–spectral attention layer can filter information in both spatial and spectral dimensions. The AGSM mechanism requires less manual intervention to dynamically generate non-Euclidean graph structures that better aggregate information. We conduct excessive experiments to compare the proposed GAT-AGSM with seven nongraph methods and three graph-based methods on widely used datasets. On the Indian Pines, Pavia University, and Salinas datasets, compared to the comparison method, the overall accuracy of GAT-AGSM is improved by at least 4.26%, 2.59%, and 1.41%, respectively. Experimental results show that GAT-AGSM has the best performance compared to the baselines in terms of various metrics. Jing Bai 0003, Zhu Xiao, Amelia Regan, Talal Ahmed Ali Ali, Yongdong Zhu, Rui Zhang 0066, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Hyperspectral Image Classification Based on Multibranch Attention Transformer NetworksabstractDeep learning has become a mainstream method of hyperspectral image (HSI) classification. Many DL-based methods exploit spatial-spectral features to achieve better classification results. However, due to the complex backgrounds in HSIs, existing methods usually show unsatisfactory performance for the class pixels located on the land-cover category boundary area. In large part, this is because the network is susceptible to interference by the irrelevant information around the target pixel in the training stage, resulting in inaccurate feature extraction. In this paper, a new multibranch transformer architecture (SST-M) that assembles spatial attention and extracts spectral features is proposed to address this problem. The transformer model has a global receptive field and thus can integrate global spatial position information in the HSI cube. Meanwhile, we design a spatial sequence attention model to enhance the useful spatial location features and weaken invalid information. Considering that HSIs contain considerable spectral information, a spectral feature extraction model is designed to extract discriminative spectral features, replacing the widely used PCA method and obtaining better classification results than it. Finally, inspired by semantic segmentation, a mask prediction model is designed to classify all of the pixels in the HSI cube; this guides the neural network to learn precise pixel characteristics and spatial distributions. To verify the effectiveness of our algorithm (SST-M), quantitative experiments were conducted in three well-known datasets, namely, IP, PU, and KSC. The experimental results demonstrate that the proposed model achieves better performance than the other state-of-the-art methods. Jing Bai 0003, Zhu Xiao, Fawang Ye, Yongdong Zhu, Mamoun Alazab, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | Binarizing Super-Resolution Networks by Pixel-Correlation Knowledge DistillationabstractConvolutional neural networks (CNNs) have been widely used in single image super-resolution (SR) and obtained remarkable performance. However, most CNN-based SR models require heavy computation, which limits their real-world applications. In this paper, we address the computation problem of SR by network binarization, which converts the full-precision network into the binary network, thus intensively reducing computation. We propose the pixel-correlation distillation for SR network binarization, which distills the knowledge of pixel relationship from the original full-precision network to the binary network. In addition, we further reduce the quantization errors of the binary network by introducing trainable scaling factors to replace the fixed scaling factors in most existing binarization methods. We carry out extensive experiments on SRResNet [1] and VDSR [2], which are two commonly used SR networks. It is shown that the proposed method generates more visually pleasing SR images, and consistently outperforms other state-of-the-art methods in PSNR and SSIM. Qiu Huang, Haoji Hu, Yongdong Zhu, Zhifeng Zhao |
ICIP | 4 |
| 2021 | Design aspects on physical layer structure for 5G V2X and related issuesabstractThe 3rd Generation Partnership Project (3GPP) is currently studying the support of advanced Vehicle-to-Everything (V2X) services based on 5G New Radio (NR). The standardization work in Release 16 has been completed in June 2020. In order to achieve the high reliability and low latency requirements of the advanced V2X services, substantial changes are required to extend the NR framework to sidelink. This paper elaborates the state-of-art design aspects on physical layer structure for 5G V2X in Release 16. Some technical issues and future enhancements are also presented. Zhenting Li, Yongdong Zhu |
VTC Spring | 4 |
| 2021 | A geometry-based stochastic channel model and its application for intelligent reflecting surface assisted wireless communicationabstractAbstract Intelligent reflecting surface (IRS) is a new concept originating from metamaterials, which can achieve beamforming through controllable passive reflecting. This device makes it possible to engineer the wireless communication environment, and has drawn increasing attention. However, the associated channel models in current literature are mainly borrowed from conventional wireless channel models directly, omitting the unique features of IRS. In this paper, a geometry‐based stochastic channel model for IRS‐assisted wireless communication system is employed. The model has certain accuracy and low computational complexity. In particular, it captures the correlations of subchannels associated with different IRS elements, which is typically not considered in current works. Based on this channel model and the derived channel spatial correlation functions (CFs), an iterative reflection coefficients configuration method is proposed exploiting statistical channel state information to maximise the ergodic channel capacity. The impacts of the IRS spatial positions as well as the number of the IRS elements on the ergodic channel capacity is investigated through simulations. It is found that to obtain a larger ergodic channel capacity, the IRS should be placed in the vicinity of either the transmitter side or the receiver side, which is a useful guideline for practical deployment. Jian Dang, Shicheng Gao, Yongdong Zhu, Rongbin Guo, Hao Jiang 0006, Zaichen Zhang, Liang Wu 0001, Bingcheng Zhu, Lei Wang 0182 |
IET Commun. | 3 |