Jing Yang 0017

dblp:62/5839-17 · DBLP profile ↗
← Back
23ranked-venue papers
4as first author
23since 2021 · last 2026
0000-0003-1915-9487ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Systems, architecture and hardware · 5 · 5 since 2021Computer networks · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-Design
abstract
Efficiently training large-scale models (LMs) in GPU clusters involves two separate avenues: inter-job dynamic scheduling and intra-job adaptive parallelism (AP). However, existing dynamic schedulers struggle with large-model scheduling due to the mismatch between static parallelism (SP)-aware scheduling and AP-based execution, leading to cluster inefficiencies such as degraded throughput and prolonged job queuing. This paper presents Arena, a large-model training system that co-designs dynamic scheduling and adaptive parallelism to achieve high cluster efficiency. To reduce scheduling costs while improving decision quality, Arena designs low-cost, disaggregated profiling and AP-tailored, load-aware performance estimation, while unifying them by sharding the joint scheduling-parallelism optimization space via a grid abstraction. Building on this, Arena dynamically schedules profiled jobs in elasticity and heterogeneity dimensions, and executes them using efficient AP with pruned search space. Evaluated on heterogeneous testbeds and production workloads, Arena reduces job completion time by up to 49.3% and improves cluster throughput by up to 1.6X.
Chunyu Xue, Weihao Cui, Quan Chen 0002, Chen Chen 0067, Han Zhao 0005, Shulai Zhang, Linmei Wang, Limin Xiao 0001, Weifeng Zhang 0003, Jing Yang 0017, Bingsheng He, Minyi Guo
EuroSys11
2026 AR-FRSG: Fine-grained agricultural image recognition via feature space reconstruction and semantically guided fusion
Banghong Wu, Jing Yang 0017, Jiheng Wu
Expert Syst. Appl.3
2026 MMT-SNN: Markovian decision and multi-threshold spike delivery integrated adaptive spiking neural network for tactile object recognition
Jing Yang 0017, Zukun Yu, Changfu Zhang, Shaobo Li 0001, Zhidong Su, Yixiong Feng
Expert Syst. Appl.1
2026 Graph neural networks with flow conservation constraints for real-time origin-destination matrix completion
Jing Yang 0017
Pattern Recognit.2
2026 LHRS-LTW: A Load-Aware Hybrid-Cloud Resource Scheduling Framework for Large-Scale Training Workloads
Ao Wei, Jing Yang 0017, Pu Pang, Shixuan Sun, Xiaoli Ruan, Yuling Chen 0002, Minyi Guo
IEEE Trans. Parallel Distributed Syst.2
2025 DACO: Unlocking Latent Dataflow Opportunities in Edge-Side SIMT Accelerators
Han Zhao 0005, Yiying Xiang, Xiaochun Ye, Deze Zeng, Jing Yang 0017, Weihao Cui, Quan Chen 0002, Jingwen Leng, Minyi Guo
APPT6
2025 EC-TRL: Evolutionary-Weighted Clustering and Transformer-Augmented Reinforcement Learning for Dynamic Resource Scheduling in Edge Cloud Environments
abstract
With the rapid development of edge computing, devices now offer powerful computing capabilities and diverse applications. However, the surge in smart devices accessing the Internet overwhelms edge servers, which have limited and unevenly distributed resources. This results in challenges like energy management, load balancing (LB), real-time performance, and system complexity. Existing research fails to comprehensively consider these challenges’ combined impact, making it difficult to maximize performance when facing real complex scenarios. To address the above issues, this article proposes an edge cloud resource scheduling scheme based on evolutionary-weighted clustering and transformer-augmented reinforcement learning (EC-TRL). First, server nodes are deployed at the center of user clusters, based on user device locations, to optimize communication delay and evenly distribute resources. Second, the multiobjective scheduling optimization problem under delay constraints is converted into a Markov decision problem, and a deep reinforcement learning method based on soft actor-critic (SAC) is proposed. Finally, actor transformer (AT) and critic transformer (CT) are proposed to improve the network structure of SAC, capture long-term dependencies and complex patterns in long task scheduling sequences, and improve the model’s adaptability and generalization performance in complex dynamic environments. Through comparison experiments with round robin, random, proximal policy optimization, dueling double deep Q-learning network, SAC-L, and SAC-M, the results show that the proposed method improves the optimization performance of energy consumption, LB, and rejection rate of edge cloud resource scheduling by at least 9.57%, 10.90%, and 5.05%.
Jing Yang 0017, Shaobo Li 0001, Zhidong Su, Jialin Lu
IEEE Internet Things J.2
2025 MSM-TFL: A Multiservice, Multitask Transformer Framework for Edge Load Prediction
abstract
Accurate load prediction ensures edge cluster stability, which is crucial for resource allocation and task scheduling. However, varied service levels in multi-service clusters cause uneven resource distribution, overloading some devices. To address this issue, this paper introduces a multiservice, multitask transformer framework for edge load prediction, referred to as MSM-TFL. This framework applies dual encoders to capture variations in service features over time and leverages the Continuous Time Screening (CTS) method to calculate adjacent time points or identify the most similar time points as inputs to the dual encoders using the Similarity Search Using Service Features (SSUSF) method. In tests conducted on 30 server clusters with five loads, MSM-TFL-Continue outperforms the SOTA approach in terms of prediction errors for 1-, 5-, and 15-minute loads, total CPU time, and user CPU time. Compared with iTransformer, MSM-TFL-Continue reduces the prediction error in terms of the MAE, RSE, RMSE, and MAPE by 77.29%, 70.06%, 73.90%, and 77.47%, respectively, when a 5-minute load is used as an example. The prediction errors in terms of the MAE, RSE, RMSE, and MAPE are decreased by 74.46%, 64.75%, 69.31%, and 72.45%, respectively, by applying MSM-TFL-Similarity.
Shuanglin Zou, Jing Yang 0017, Xiaoli Ruan, Yongbin Qin, Chengjiang Li, Wei Zhang 0363
IEEE Internet Things J.2
2025 Effective Generative Replay with Strong Memory for Continual Learning
Jing Yang 0017, Qinglang Li, Zhidong Su, Xiaoli Ruan, Changfu Zhang
Knowl. Based Syst.1
2025 Indoor scene multi-object tracking based on region search and memory buffer pool
Yang Li 0046, Guanci Yang, Zhidong Su, Shaobo Li 0001, Jing Yang 0017, Ling He 0002
Pattern Recognit.5
2025 Taming Flexible Job Packing in Deep Learning Training Clusters
abstract
Job packing is an effective technique to harvest the idle resources allocated to the deep learning (DL) training jobs but not fully utilized, especially when clusters may experience low utilization, and users may overestimate their resource needs. However, existing job packing techniques tend to be conservative due to the mismatch in scope and granularity between job packing and cluster scheduling. In particular, tapping the potential of job packing in the training cluster requires a local and fine-grained coordination mechanism. To this end, we propose a novel job-packing middleware named Gimbal , which operates between the cluster scheduler and the hardware resources. As middleware, Gimbal must not only facilitate coordination among the packed jobs but also support various scheduling objectives of different schedulers. Gimbal achieves dual functionality by introducing a set of worker calibration primitives designed to calibrate workers’ execution status in a fine-grained manner. The primitives obscure the complexity of the underlying job and resource management mechanisms, thus offering the generality and extensibility for crafting coordination policies tailored to various scheduling objectives. We implement Gimbal on a real-world GPU cluster and evaluate it with a set of representative DL training jobs. The results show that Gimbal improves different scheduling objectives up to 1.32× compared with the state-of-the-art job packing techniques.
Pengyu Yang, Weihao Cui, Chunyu Xue, Han Zhao 0005, Chen Chen 0067, Quan Chen 0002, Jing Yang 0017, Minyi Guo
ACM Trans. Archit. Code Optim.7
2025 LRDTN: Spectral-Spatial Convolutional Fusion Long-Range Dependence Transformer Network for Hyperspectral Image Classification
abstract
Recently, deep learning has achieved remarkable breakthroughs in hyperspectral image (HSI) classification tasks, particularly with methods based on convolutional neural networks (CNNs) and transformers. However, these methods have several limitations: 1) the limited receptive field inherent in the convolutional layer greatly hampers capturing feature contextual information on a large scale and 2) transformers cannot establish strong local relationships, making it challenging to characterize complex dependencies between distant pixels and different bands in HSIs. Moreover, as the network complexity increases, so does the number of network parameters. We propose a novel network called the spectral-spatial convolutional fusion long-range dependence transformer network (LRDTN) for HSI classification to address these challenges. LRDTN comprises three key components: dynamic-dependent convolutional (DDC) module, the multiscale enhanced fusion (MsEF) module, and the local–global perception transformer (LGPT). Specifically, the DDC dynamically models local features, while the MsEF integrates information from different scales to capture contextual relationships in HSI features effectively. Additionally, the ability to mine and utilize HSI local-global features and complex long-range dependencies is enhanced by the proposed transformer variant, LGPT. Ultimately, through the ingeniously designed structure of the LRDTN, the model effectively maintains its performance while reducing the number of network parameters. Extensive experiments conducted on four typical HSI datasets, including urban areas, agricultural areas, and swamps, demonstrate the superiority of LRDTN over other state-of-the-art networks. The code is available athttps://github.com/ybyangjing/LRDTN.
Shujie Ding, Xiaoli Ruan, Jing Yang 0017, Chengjiang Li, Jie Sun 0033, Xianghong Tang, Zhidong Su
IEEE Trans. Geosci. Remote. Sens.3
2025 EMC-LSP: A Novel Lightweight Architecture for Edge Multi-Node Long Sequence Prediction
abstract
Edge device traffic prediction is crucial for autonomous network control and management. However, the rapid proliferation of smart 5G networks results in increasingly heterogeneous, dynamic, and complex traffic loads on edge nodes, rendering traditional short-term prediction methods insufficient for medium- and long-term network resource scheduling. To address this, we propose a novel multi-node lightweight long-sequence deep learning-based prediction architecture (EMC-LSP) to effectively capture complex long- and short-term correlations in edge environments. Specifically, EMC-LSP employs frequency-domain hard-attention decomposition to separately model non-stationary high-frequency and low-frequency traffic, utilizes two-layer null frequency-domain convolution for long-term low-frequency similarity, and designs a high-frequency interpolation-based prediction method. In extensive tests on 18 datasets, EMC-LSP demonstrated superior performance, reducing the average prediction error of MSE and MAE by 15.20% while decreasing model parameters by 50 times.
Chuanyue Xiong, Jing Yang 0017, Jiahao Zhong, Zirui He, Pu Pang, Minyi Guo
IEEE Trans. Sustain. Comput.2
2024 Deep reinforcement learning-based resource scheduling for energy optimization and load balancing in SDN-driven edge computing
Jing Yang 0017, Shaobo Li 0001, Zhidong Su
Comput. Commun.2
2024 A2C-DRL: Dynamic Scheduling for Stochastic Edge-Cloud Environments Using A2C and Deep Reinforcement Learning
abstract
Resource management challenges frequently manifest in systems and networks as tough online decision tasks, for which the proper solution is dependent on an understanding of the workload and environment and facilitates smooth use of mobile edge and cloud resources. Due to the geographical dispersion of resources, constrained resource capacity, unpredictable nature of tasks, and network hierarchy present in such contexts, it is difficult to efficiently schedule jobs in edge environments. Unfortunately, existing heuristic-based methods lack generality and fast adaptability and thus cannot optimally solve such problems. The advantage actor–critic (A2C) method, on the one hand, can quickly adapt to dynamic circumstances based on relatively few data, and deep reinforcement learning (DRL) agents can on the other hand rapidly learn from their experience of environmental interactions to make better judgments. Therefore, we present an A2C-DRL real-time task scheduling technique for stochastic edge–cloud environments that enables decentralized learning and simultaneous work scheduling across multiple servers. With the aim of producing efficient scheduling decisions, we develop reward values for various resources and model the update policy, server resource scheduling method, and policy learning method. The model is adaptive and includes various hyperparameters that can be adjusted in accordance with the application requirements. We evaluate the load balancing capability of the model by introducing a load balancing factor. Experiments on real datasets show that the proposed A2C-DRL method outperforms seven state-of-the-art algorithms in terms of the reward value, task rejection, and the load balancing factor.
Jialin Lu, Jing Yang 0017, Shaobo Li 0001, Wu Jiang, Jiangtian Dai, Jianjun Hu
IEEE Internet Things J.2
2024 GGT-SNN: Graph learning and Gaussian prior integrated spiking graph neural network for event-driven tactile object recognition
Jing Yang 0017, Zukun Yu, Shaobo Li 0001, Jianjun Hu, Ji Xu 0001
Inf. Sci.1
2024 SPIRF-CTA: Selection of parameter importance levels for reasonable forgetting in continuous task adaptation
Qinglang Li, Jing Yang 0017, Xiaoli Ruan, Shaobo Li 0001, Jianjun Hu, Bingqi Hu
Knowl. Based Syst.2
2024 A class-incremental learning approach for learning feature-compatible embeddings
Hongchao An, Jing Yang 0017, Xiuhua Zhang, Xiaoli Ruan, Shaobo Li 0001, Jianjun Hu
Neural Networks2
2024 LCTCS: Low-Cost and Two-Channel Sparse Network for Hyperspectral Image Classification
abstract
Abstract Using convolutional neural networks (CNNs) in classifying hyperspectral images (HSIs) has achieved quite good results in recent years. It is widely used in agricultural remote sensing, geological exploration, environmental monitoring, and marine remote sensing. Unfortunately, the complexity of network structures used for hyperspectral image classification challenges the efficient delivery of HSI data extremely, and existing methods suffer from a large amount of redundancy in the network weight parameters during training, as they either require huge computational resources or make inefficient use of storage space when designing the network structure, and many of the parameters that waste computational resources contribute less to the rich spectral and spatial information transfer in HSI. So we introduce LCTCS, a better low-memory and less-parametric network approach. LCTCS aims to improve the efficiency of computational resource utilization with advanced classification performance and lower levels of computational resources. Unlike the conventional 2D and 3D convolution used previously, we use simple and efficient 3D grouped convolution as a vehicle to convey the semantic features of HSIs. More specifically, we design a novel two-channel sparse network to classify HSIs since grouped 3D convolution conveys the properties of hyperspectral data well in the time and space domains.We have compared LCTCS with eight widely used network methods on four publicly available hyperspectral datasets for learning HSI information. A series of experiments shows that the model architecture designed has $$65.89 \%$$ 65.89 % less storage space than the DBDA method, consumes $$67.36 \%$$ 67.36 % fewer computational resources than the SSRN method on the IP dataset, and accomplishes a highly accurate classification task with the number of parameters accounting for only $$1.99 \%$$ 1.99 % that of the DBMA method.
Jie Sun 0033, Jing Yang 0017, Shujie Ding, Shaobo Li 0001, Jianjun Hu
Neural Process. Lett.2
2024 LVAR-CZSL: Learning Visual Attributes Representation for Compositional Zero-Shot Learning
abstract
Compositional Zero-Shot Learning (CZSL) has been applied to various scenarios, including scene understanding, visual-language representation, and domain adaptation. Despite numerous endeavours and significant advancements, the crucial issues of fuzzy conceptualization of visual attributes and insufficient inter-class connectivity, have remained insufficiently addressed. To address these issues, we propose Learning Visual Attributes Representation for Compositional Zero-Shot Learning (LVAR-CZSL), which has the ability to learn visual attributes and inter-class dependencies. LVAR-CZSL is mainly composed of two key components: the Visual Attribute Representation Module (VARM) and the Connected Learning Module (CLM). Specifically, VARM extracts detailed attributes and object visual features from global visual features, resolving the issue of fuzzy visual attribute concepts. Moreover, CLM endows LVAR-CZSL with the capability to perceive connectivity between different attributes and objects, effectively enhancing inter-class connectivity. To establish a close connection between VARM and CLM and minimize the gap between image and text features, we introduce the composition-attribute-object Joint Scoring Function (JSF). Additionally, we propose Joint Loss Function (JLF) to optimize the learning process of VARM and CLM. The experiment results on four datasets show that LVAR-CZSL achieves state-of-the-art performance. The code is available athttps://github.com/mxjmxj1/LVAR-CZSL.
Xingjiang Ma, Jing Yang 0017, Jiacheng Lin, Zhenzhe Zheng 0001, Shaobo Li 0001, Bingqi Hu, Xianghong Tang
IEEE Trans. Circuits Syst. Video Technol.2
2024 Long-Term Airport Network Performance Forecasting With Linear Diffusion Graph Networks
abstract
Precise forecasting of airport performances, such as landing rates and delays, is essential for the smooth operation of air traffic management systems and for improving the passenger experience. While current efforts predominantly address short-term predictions, the imperative for long-term forecasting is undeniable, particularly for strategic operational planning and resource management. Equally important is the explainability of these forecasts, which is critical for effective decision-making. To meet these needs, our study introduces an innovative approach to airport performance forecasting with the Linear-Diffusion Graph Network (LDGN), an explainable and probabilistic model. The LDGN is intricately structured, comprising stacked temporal linear layers and graph diffusion layers that harness the clarity of linear time series models. This configuration adeptly captures the nuanced interactions between graph-based diffusion processes and the dynamic spread of conditions across airport performances. Departing from conventional point forecasts, the LDGN produces a probabilistic output, prioritizing predictability and a strong capacity for generalization. The model’s pre-training is enhanced with stochastic mask reconstruction, a technique that significantly improves its ability to generalize. Through rigorous testing on real-world datasets, we have validated the LDGN’s superior performance in both long-term and very long-term forecasting. Our results demonstrate not only high accuracy and explainability but also a robust capacity for uncertainty quantification.
Jing Yang 0017, Yi Lin 0006, Hongyu Yang 0002
IEEE Trans. Intell. Transp. Syst.2
2023 A TL_FLAT Model for Chinese Text Datasets of UAV Power Systems: Optimization and Performance
abstract
The manufacturing processes of unmanned aerial vehicle (UAV) power systems generate large amounts of data and knowledge. The extraction of useful information or patterns from redundant data and knowledge texts has become a challenge in intelligent manufacturing. Unfortunately, graphics processing unit (GPU)‐based parallel computing is limited, and the inference speeds of the available named entity recognition (NER) models for Chinese text datasets are low because they are mainly based on the long short‐term memory (LSTM) algorithm. Herein, first, the flat‐lattice transformer (FLAT) model was optimized by using a stochastic gradient descent with momentum (SGDM) optimizer and adjusting the model hyperparameters. Compared with the existing NER methods, the proposed optimization algorithm achieved better performance on the available dataset. Then, an NER method named the TL_FLAT model based on transfer learning and the abovementioned optimization model was introduced. Finally, a Chinese text dataset from a UAV power system created by the authors was used to validate the proposed method. The F1 score was 76.26%, the precision value was 76.98%, and the recall value was 75.56%, indicating that the TL_FLAT model was suitable for Chinese text entity recognition for UAV power systems.
Mingming Shen, Shaobo Li 0001, Jing Yang 0017, Ansi Zhang, Qiuchen He, Ruiqiang Pu
Int. J. Intell. Syst.3
2023 A New Multinetwork Mean Distillation Loss Function for Open-World Domain Incremental Object Detection
abstract
The development of object detection networks has reached a high point, and there have been significant improvements in accuracy and detection speed. Object detection is widely used in intelligent robots, self‐driving cars, and other edge‐intelligent terminals. Unfortunately, when a detector is allowed to learn new objects in an unfamiliar environment, it can catastrophically forget the objects it has already learned. In particular, reliable and stable knowledge cannot be extracted from old models. Based on this, a new multinetwork mean distillation loss function for open‐world domain incremental object detection is presented. To better extract reliable and stable knowledge from old models, we enhanced the distillation output of the detector with a ResNet50 backbone and an output RoI head. The distillation output of the intermediate RPN is softened by adaptive distillation. To obtain more stable results, the ResNet50 backbone and RPN on the channel are zero‐averaged. Various incremental steps and stability experiments are performed on two benchmark datasets, PASCAL VOC and MS COCO. The experimental results show the excellent performance of our method in different experimental scenarios, and it is superior to the most advanced methods. For example, in the setting of the batch task, incremental object detection on the PASCAL VOC and MS COCO datasets is improved by 3.4% and 2.1%, respectively.
Jing Yang 0017, Suhao Chen, Qinglang Li, Shaobo Li 0001, Xiuhua Zhang
Int. J. Intell. Syst.1