VLDB 2026 Research / reviewers in the wild / expert
Jinghui Zhang 0001
dblp:44/6686-1
· DBLP profile ↗
49ranked-venue papers
16as first author
35since 2021 · last 2026
0000-0002-9067-7896ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 14 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Systems, architecture and hardware · 8 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 8 since 2021Computer networks · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Otter: Mitigating Background Distractions of Wide-Angle Few-Shot Action Recognition with Enhanced RWKV
Wenbo Huang 0001, Jinghui Zhang 0001, Guang Li 0008, Lei Zhang 0130, Fang Dong 0001, Takahiro Ogawa 0001, Miki Haseyama |
AAAI | 2 |
| 2026 | FedTPA: Tackling Data Heterogeneity with Adaptive Parameter Allocation in Federated Instruction TuningabstractFederated instruction tuning of large language models (LLMs) has recently emerged as a promising research direction for preserving data privacy while enabling collaborative model adaptation. However, due to the heterogeneity of local instruction data across clients in federated settings, assigning the same trainable parameter size to all clients may compromise local learning effectiveness and limit the overall performance of the global model. To address this challenge, we propose FedTPA, a dynamic pruning-based strategy that allocates and adjusts the adapter dimensions of local models based on the distribution of local instruction data and trends in training loss. This allows the trainable parameter size on each client to better align with the complexity and characteristics of its local data. We evaluate FedTPA across multilingual, multi-task, and varying degrees of data heterogeneity scenarios. Experimental results demonstrate that FedTPA outperforms existing federated instruction tuning methods, achieving up to a 3% improvement in Rouge-L scores. Jinghui Zhang 0001, Ding Ding 0002 |
DATE | 2 |
| 2026 | Gbp-llm: gaze behavior prediction in 6DoF VR via large language models
Ding Ding 0002, Chang Qi, Zheyu Cao, Jinghui Zhang 0001 |
Multim. Syst. | 6 |
| 2025 | Manta: Enhancing Mamba for Few-Shot Action Recognition of Long Sub-SequenceabstractIn few-shot action recognition (FSAR), long sub-sequences of video naturally express entire actions more effectively. However, the high computational complexity of mainstream Transformer-based methods limits their application. Recent Mamba demonstrates efficiency in modeling long sequences, but directly applying Mamba to FSAR overlooks the importance of local feature modeling and alignment. Moreover, long sub-sequences within the same class accumulate intra-class variance, which adversely impacts FSAR performance. To solve these challenges, we propose a Matryoshka MAmba and CoNtrasTive LeArning framework (Manta). Firstly, the Matryoshka Mamba introduces multiple Inner Modules to enhance local feature representation, rather than directly modeling global features. An Outer Module captures dependencies of timeline between these local features for implicit temporal alignment. Secondly, a hybrid contrastive learning paradigm, combining both supervised and unsupervised methods, is designed to mitigate the negative effects of intra-class variance accumulation. The Matryoshka Mamba and the hybrid contrastive learning paradigm operate in two parallel branches within Manta, enhancing Mamba for FSAR of long sub-sequence. Manta achieves new state-of-the-art performance on prominent benchmarks, including SSv2, Kinetics, UCF101, and HMDB51. Extensive empirical studies prove that Manta significantly improves FSAR of long sub-sequence from multiple perspectives. Wenbo Huang 0001, Jinghui Zhang 0001, Guang Li 0008, Lei Zhang 0130, Shuoyuan Wang, Fang Dong 0001, Jiahui Jin 0001, Takahiro Ogawa 0001, Miki Haseyama |
AAAI | 2 |
| 2025 | Bilateral Virtual Companions: The Impact of Virtual Humans' Movement and Voice Realism on User Perception and Experience in Multi-user VR CinemasabstractDespite the increasing prevalence of online social interaction, challenges such as insufficient immersion and lack of interactivity still persist. To overcome these limitations, this study developed a multi-user virtual reality (VR) cinema system with motion capture (Mocap) and multi-user VR technology. The system demonstrated strengths in overcoming physical space restrictions and saving travel costs, providing a more enriched interactive experience and fulfilling social needs under special circumstances, and facilitating metaverse applications. Furthermore, to give insight into the further design of multi-user VR cinemas, this study investigated the impact of virtual humans' (VHs') characteristics on user perception and experience of bilateral virtual companions, by evaluating the impact of movement and voice realism on immersion, social presence and intimacy. The results show that while neither movement nor voice realism significantly influences immersion, voice realism rather than movement realism significantly affects social presence and intimacy. Jingfeng Hu, Ding Ding 0002, Xiangyu Xu 0001, Jinghui Zhang 0001, Jiahui Jin 0001, Fang Dong 0001 |
CSCWD | 4 |
| 2025 | Urban Region Pre-training and Prompting: A Graph-based ApproachabstractUrban region representation is crucial for various urban downstream tasks. However, despite the proliferation of methods and their success, acquiring general urban region knowledge and adapting to different tasks remains challenging. Existing work pays limited attention to the fine-grained functional layout semantics in urban regions, limiting their ability to capture transferable knowledge across regions. Further, inadequate handling of the unique features and relationships required for different downstream tasks may also hinder effective task adaptation. In this paper, we propose a Graph-based Urban Region Pre-training and Prompting framework (GURPP) for region representation learning. Specifically, we first construct an urban region graph and develop a subgraph-centric urban region pre-training model to capture the heterogeneous and transferable patterns of entity interactions. This model pre-trains knowledge-rich region embeddings using contrastive learning and multi-view learning methods. To further refine these representations, we design two graph-based prompting methods: a manually-defined prompt to incorporate explicit task knowledge and a task-learnable prompt to discover hidden knowledge, which enhances the adaptability of these embeddings to different tasks. Extensive experiments on various urban region prediction tasks and different cities demonstrate the superior performance of our framework. Jiahui Jin 0001, Yifan Song 0003, Dong Kan, Haojia Zhu, Xiangguo Sun, Xigang Sun, Jinghui Zhang 0001 |
KDD (2) | 8 |
| 2025 | Spatiotemporal Generalization Graph Neural Network-Based Prediction Models by Considering Morphological Diversity in Traffic NetworksabstractThe morphological diversity, referring to the variations in traffic network topologies defined in this paper, often emerges and brings difficulties in successfully transferring a pre-trained prediction model from one traffic network to another. Moreover, most existing research primarily assumes that traffic data in source and target networks follow independent and identically distributed (i.i.d.) patterns, which is usually not consistent with real-world situations, particularly when considering morphological diversity. For this inconsistency, many efforts have been made, but they mainly concentrate on temporal aspects, which significantly differ from traffic prediction due to spatial and temporal correlations among road segments, influenced by variations in road topology and traffic behavior. This paper introduces a causality-based spatiotemporal out-of-distribution (OOD) generalization method, which is adaptable to most GNNs for diverse, large-scale, dynamic traffic systems with zero-shot. Furthermore, to enhance the generalization and adaptability of the proposed method, we introduce graph matching and equal-sized graph partitioning to alleviate spatial shift between the source and target traffic networks, reduce and align the scale of the networks. Experiments carried out on traffic flow datasets demonstrate that our method significantly improves the performance of various GNN-based traffic predictors in the situation of morphological diversity, achieving a maximum reduction in MAE of 33.08%. Compared to other OOD-driven baselines, our approach also shows a notable improvement, with up to a 40.58% decrease in MAE. Limei Liu, Peibo Duan, Zhuo Chen 0019, Jinghui Zhang 0001, Siyuan Feng 0006, Wenwei Yue, Jia Rong |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | An Event-Centric Framework for Predicting Crime Hotspots With Flexible Time Intervals
Jiahui Jin 0001, Yi Hong 0003, Guandong Xu, Jinghui Zhang 0001, Hancheng Wang |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2025 | Multi-Dimensional Training Optimization for Efficient Federated Synergy LearningabstractEdge learning (EL) is an end-to-edge collaborative learning paradigm enabling devices to participate in model training and data analysis, opening countless opportunities for edge intelligence. As a promising EL framework, federated synergy learning (FSyL) mitigates the computation and communication overhead on resource-constrained devices by offloading partial model layers to the edge server for synergistic training. Nevertheless, due to the system and statistical heterogeneity, naively using existing FSyL methods is significantly time-consuming and causes accuracy degradation. Motivated by this issue, this paper introduces a novel FSyL framework that integrates multi-dimensional training optimization and formulates the edge learning cost minimization (ELCM) problem. To tackle the ELCM efficiently, we designOL-MG, anOnLineModel Splitting and Resource ProvisioningGame. Specifically, we first reformulate and decompose the original ELCM based on data quality evaluation. Then, given a model splitting decision, we determine the optimal resource provisioning in Sub-problem1, based on which optimal model splitting in Sub-problem2 is modeled as a potential game. Subsequently, we introduce a decentralized algorithm to find a Nash equilibrium (NE) solution. Furthermore, we further extendOL-MGto support a budget-aware multi-edge scenario. Extensive experiments demonstrate that the proposed mechanism significantly outperforms state-of-the-art methods in cost-saving and accuracy improvement. Shucun Fu, Fang Dong 0001, Runze Chen 0001, Dian Shen, Jinghui Zhang 0001, Qiang He 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | Resource-Efficient DNN Inference With Early Exiting in Serverless Edge ComputingabstractServerless Edge Computing (SEC) has gained widespread adoption in improving resource utilization due to its triggered event-driven model. However, deploying deep neural network (DNN) inference services directly in SEC leads to resource inefficiencies, which stem from two key factors. First, existing methods adopt model-wise function encapsulation, which requires the entire DNN model to occupy memory throughout its execution lifecycle. This increases both memory footprint and occupancy time. Second, uniform DNN inference for diversity input leads to redundant computations and additional inference time. To this end, we propose REDI, a novel framework that leverages fine-grained block-wise function encapsulation and progressive inference to provide resource-efficient DNN inference while ensuring latency requirements. REDI enables the release of memory from already inferred shallow networks and allows each request to exit early based on input data complexity, eliminating redundant computations. To fully unleash the potential, REDI jointly considers resource heterogeneity, data diversity, and environment dynamics to investigate the block-wise function placement problem. We introduce an uncertainty-aware online learning-driven algorithm with bounded regret. Finally, we conduct extensive trace-driven experiments to evaluate our methods, demonstrating that REDI achieves a significant speedup of up to$6.52\times$in terms of resource usage cost compared to state-of-the-art methods. Xiaolin Guo, Fang Dong 0001, Dian Shen, Zhaowu Huang, Jinghui Zhang 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | DiG-In-GNN: Discriminative Feature Guided GNN-Based Fraud Detector against Inconsistencies in Multi-Relation Fraud GraphabstractFraud detection on multi-relation graphs aims to identify fraudsters in graphs. Graph Neural Network (GNN) models leverage graph structures to pass messages from neighbors to the target nodes, thereby enriching the representations of those target nodes. However, feature and structural inconsistency in the graph, owing to fraudsters' camouflage behaviors, diminish the suspiciousness of fraud nodes which hinders the effectiveness of GNN-based models. In this work, we propose DiG-In-GNN, Discriminative Feature Guided GNN against Inconsistency, to dig into graphs for fraudsters. Specifically, we use multi-scale contrastive learning from the perspective of the neighborhood subgraph where the target node is located to generate guidance nodes to cope with the feature inconsistency. Then, guided by the guidance nodes, we conduct fine-grained neighbor selection through reinforcement learning for each neighbor node to precisely filter nodes that can enhance the message passing and therefore alleviate structural inconsistency. Finally, the two modules are integrated together to obtain discriminable representations of the nodes. Experiments on three fraud detection datasets demonstrate the superiority of the proposed method DiG-In-GNN, which obtains up to 20.73% improvement over previous state-of-the-art methods. Our code can be found at https://github.com/GraphBerry/DiG-In-GNN. Jinghui Zhang 0001, Zhengjia Xu, Dingyang Lyu 0001, Dian Shen, Jiahui Jin 0001, Fang Dong 0001 |
AAAI | 1 |
| 2024 | A Deep Prediction Framework for Multi-Source Information via Heterogeneous GNNabstractPredicting information diffusion is a fundamental task in online social networks (OSNs).Recent studies mainly focus on the popularity prediction of specific content but ignore the correlation between multiple pieces of information.The topic is often used to correlate such information and can correspond to multi-source information.The popularity of a topic relies not only on information diffusion time but also on users' followership.Current solutions concentrate on hard time partition, lacking versatility.Meanwhile, the hop-based sampling adopted in state-of-the-art (SOTA) methods encounters redundant user followership.Moreover, many SOTA methods are not designed with good modularity and lack evaluation for each functional module and enlightening discussion.This paper presents a novel extensible framework, coined as HIF, for effective popularity prediction in OSNs with four original contributions.First, HIF adopts a soft partition of users and time intervals to better learn users' behavioral preferences over time.Second, HIF utilizes weighted sampling to optimize the construction of heterogeneous graphs and reduce redundancy.Furthermore, HIF supports multi-task collaborative optimization to improve its learning capability.Finally, as an extensible framework, HIF provides generic module slots to combine different submodules (e.g., RNNs, Zhen Wu 0001, Jingya Zhou, Jinghui Zhang 0001, Ling Liu 0001, Chizhou Huang |
KDD | 3 |
| 2024 | SOAP: Enhancing Spatio-Temporal Relation and Motion Information Capturing for Few-Shot Action RecognitionabstractHigh frame-rate~(HFR) videos of action recognition improve fine-grained expression while reducing the spatio-temporal relation and motion information density. Thus, large amounts of video samples are continuously required for traditional data-driven training. However, samples are not always sufficient in real-world scenarios, promoting few-shot action recognition~(FSAR) research. We observe that most recent FSAR works build spatio-temporal relation of video samples via temporal alignment after spatial feature extraction, cutting apart spatial and temporal features within samples. They also capture motion information via narrow perspectives between adjacent frames without considering density, leading to insufficient motion information capturing. Therefore, we propose a novel plug-and-play architecture for FSAR called Spatio-tempOral frAme tuPle enhancer (SOAP) in this paper. The model we designed with such architecture refers to SOAP-Net. Temporal connections between different feature channels and spatio-temporal relation of features are considered instead of simple feature extraction. Comprehensive motion information is also captured, using frame tuples with multiple frames containing more motion information than adjacent frames. Combining frame tuples of diverse frame counts further provides a broader perspective. SOAP-Net achieves new state-of-the-art performance across well-known benchmarks such as SthSthV2, Kinetics, UCF101, and HMDB51. Extensive empirical evaluations underscore the competitiveness, pluggability, generalization, and robustness of SOAP. The code is released at https://github.com/wenbohuang1002/SOAP. Wenbo Huang 0001, Jinghui Zhang 0001, Xuwei Qian, Zhen Wu 0001, Meng Wang 0009, Lei Zhang 0130 |
ACM Multimedia | 2 |
| 2024 | Learning context-aware region similarity with effective spatial normalization over Point-of-Interest data
Jiahui Jin 0001, Yifan Song 0003, Dong Kan, Binjie Zhang, Jinghui Zhang 0001, Hongru Lu |
Inf. Process. Manag. | 6 |
| 2024 | Non-parameter clustering algorithm based on chain propagation and natural neighbor
Tianshuo Li, Juntao Yang, Rui Pu, Jinghui Zhang 0001, Dongming Tang, Tao Liu 0027 |
Inf. Sci. | 5 |
| 2024 | NMNN: Newtonian Mechanics-based Natural Neighbor algorithm
Wentong Wang, Juntao Yang, Jinghui Zhang 0001, Dongming Tang, Tao Liu 0027 |
Inf. Sci. | 4 |
| 2024 | Federated variational generative learning for heterogeneous data in distributed environments
Wei Xie 0011, Runqun Xiong, Jinghui Zhang 0001, Jiahui Jin 0001, Junzhou Luo |
J. Parallel Distributed Comput. | 3 |
| 2024 | Natural local density-based adaptive oversampling algorithm for imbalanced classification
Wentong Wang, Jinghui Zhang 0001, Juntao Yang, Dongming Tang, Tao Liu 0027 |
Knowl. Based Syst. | 3 |
| 2024 | GNaN: A natural neighbor search algorithm based on universal gravitation
Juntao Yang, Jinghui Zhang 0001, Qiwen Liang, Wentong Wang, Dongming Tang, Tao Liu 0027 |
Pattern Recognit. | 3 |
| 2024 | Addressing Heterogeneity in Federated Learning with Client Selection via Submodular OptimizationabstractFederated learning (FL) has been proposed as a privacy-preserving distributed learning paradigm, which differs from traditional distributed learning in two main aspects: the systems heterogeneity, meaning that clients participating in training have significant differences in systems performance including CPU frequency, dataset size, and transmission power, and the statistical heterogeneity, indicating that the data distribution among clients exhibits Non-Independent Identical Distribution. Therefore, the random selection of clients will significantly reduce the training efficiency of FL. In this article, we propose a client selection mechanism considering both systems and statistical heterogeneity, which aims to improve the time-to-accuracy performance by trading off the impact of systems performance differences and data distribution differences among the clients on training efficiency. First, client selection is formulated as a combinatorial optimization problem that jointly optimizes systems and statistical performance. Then, we generalize it to a submodular maximization problem with knapsack constraint, and propose the Iterative Greedy with Partial Enumeration (IGPE) algorithm to greedily select the suitable clients. Then, the approximation ratio of IGPE is analyzed theoretically. Extensive experiments verify that the time-to-accuracy performance of the IGPE algorithm outperforms other compared algorithms in a variety of heterogeneous environments. Jinghui Zhang 0001, Fa Xin, Fang Dong 0001, Junzhou Luo |
ACM Trans. Sens. Networks | 1 |
| 2024 | Joint Optimization of Device Selection and Resource Allocation for Multiple Federations in Federated Edge LearningabstractFederated edge learning (FEEL) is a promising collaborative paradigm, which employs edge devices (EDs) to train machine learning models for a federation. It opens countless opportunities to enable edge intelligence. The increasingly diversified demands for intelligent services are driving the deployment of various federations at the edge. Existing works on FEEL focus on a single federation and ignore inter-federation device competition and intra-device resource allocation, which hinders the applications of FEEL. To address this issue, this article first investigates the bottlenecks of executing multiple federations and builds a joint optimization model as a two-stage Stackelberg game involving device selection and resource allocation. To tackle the problem efficiently, we present a game-theoretical approach namedDeviceSelection andResourceAllocation forMultipleFederationsGame (DSRAMF-G). First, following the arbitrary device selection of leaders (i.e., federations), the time cost minimization of followers (i.e., EDs) is modeled as a convex problem to obtain the optimal resource allocation. Then, based on followers’ optimal responses, device selection is modeled as a congestion game. We prove the existence of the Nash equilibrium and propose a decentralized mechanism. Finally, extensive experiments show that DSRAMF-G significantly outperforms the state-of-the-art methods, achieving up to 5.9x training speedup and 2.8x resource-savings. Shucun Fu, Fang Dong 0001, Dian Shen, Jinghui Zhang 0001, Zhaowu Huang, Qiang He 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2023 | FreezePipe: An Efficient Dynamic Pipeline Parallel Approach Based on Freezing Mechanism for Distributed DNN TrainingabstractDeep Neural Network (DNN) training on a large scale is extremely time-consuming and computationally intensive, which is accelerated by distributed training. In recent years, pipeline parallelism has been developed, which enables partitioning the model across several devices, e.g. GPU, and training efficiency is improved by dividing data batches into micro-batches, with each of them processed by a different stage of the model. Currently, parallel training assumes pipeline placement and partitioning are static, with parameters updating each iteration, without accounting for freezing. This results in computational resources not being fully utilized. In this paper, we propose FreezePipe, a novel method for optimizing deep learning training that combines the freezing mechanism with pipeline parallel training. In FreezePipe, a lightweight method for determining the freezing strategy based on gradient changes is employed. Considering that resources need to be released based on the frozen layer, a lightweight model partitioning algorithm was designed to determine the optimal strategy for pipeline partitioning. Experimental results show that FreezePipe can reduce the training time by 64.5% compared to Torchgpipe on CIFAR-10 dataset without compromising any model performance. Caishan Weng, Zhiyang Shu, Zhengjia Xu, Jinghui Zhang 0001, Junzhou Luo, Fang Dong 0001, Zhengang Wang |
CSCWD | 4 |
| 2023 | Label Information Enhanced Fraud Detection against Low Homophily in GraphsabstractNode classification is a substantial problem in graph-based fraud detection. Many existing works adopt Graph Neural Networks (GNNs) to enhance fraud detectors. While promising, currently most GNN-based fraud detectors fail to generalize to the low homophily setting. Besides, label utilization has been proved to be significant factor for node classification problem. But we find they are less effective in fraud detection tasks due to the low homophily in graphs. In this work, we propose GAGA, a novel Group AGgregation enhanced TrAnsformer, to tackle the above challenges. Specifically, the group aggregation provides a portable method to cope with the low homophily issue. Such an aggregation explicitly integrates the label information to generate distinguishable neighborhood information. Along with group aggregation, an attempt towards end-to-end trainable group encoding is proposed which augments the original feature space with the class labels. Meanwhile, we devise two additional learnable encodings to recognize the structural and relational context. Then, we combine the group aggregation and the learnable encodings into a Transformer encoder to capture the semantic information. Experimental results clearly show that GAGA outperforms other competitive graph-based fraud detectors by up to 24.39% on two trending public datasets and a real-world industrial dataset from Baidu. Even more, the group aggregation is demonstrated to outperform other label utilization methods (e.g., C&S, BoT/UniMP) in the low homophily setting. Jinghui Zhang 0001, Zhengjie Huang, Weibin Li 0004, Shikun Feng, Ziheng Ma, Yu Sun 0029, Dianhai Yu, Fang Dong 0001, Jiahui Jin 0001, Beilun Wang, Junzhou Luo |
WWW | 2 |
| 2023 | Sensor-based implicit authentication through learning user physiological and behavioral characteristics
Jinghui Zhang 0001, Hancheng Zhang, Zhen Ling 0001, Ming Yang 0001 |
Comput. Commun. | 1 |
| 2023 | Wi-Fi device identification based on multi-domain physical layer fingerprint
Jinghui Zhang 0001, Zhengjia Xu, Junhe Li, Qiangsheng Dai, Zhen Ling 0001, Ming Yang 0001 |
Comput. Commun. | 1 |
| 2023 | PipePar: Enabling fast DNN pipeline parallel training in heterogeneous GPU clusters
Jinghui Zhang 0001, Geng Niu, Qiangsheng Dai, Fang Dong 0001, Zhiang Wu 0001 |
Neurocomputing | 1 |
| 2023 | Mobile applications identification using autoencoder based electromagnetic side channel analysis
Jinghui Zhang 0001, Boxi Liang, Hancheng Zhang, Zhen Ling 0001, Ming Yang 0001 |
J. Inf. Secur. Appl. | 1 |
| 2023 | PADP-FedMeta: A personalized and adaptive differentially private federated meta learning mechanism for AIoT
Fang Dong 0001, Xinghua Ge, Qinya Li, Jinghui Zhang 0001, Dian Shen, Xiao Liu 0004, Gang Li 0009, Fan Wu 0006, Junzhou Luo |
J. Syst. Archit. | 4 |
| 2023 | Noise-aware Local Model Training Mechanism for Federated LearningabstractAs a new paradigm in training intelligent models, federated learning is widely used to train a global model without requiring local data to be uploaded from end devices. However, there are often mislabeled samples (i.e., noisy samples) in the dataset, which will cause the model update to deviate from the correct direction during the training process, thus reducing the convergence accuracy of the global model. Existing works employ noisy label correction techniques to reduce the impact of noisy samples on model updates by correcting labels; however, such methods necessitate the use of prior knowledge and additional communication costs, which cannot be directly applied to federated learning due to data privacy concerns and limited communication resources. Therefore, this paper proposes a noise-aware local model training method that corrects the noisy labels directly at the end device under the constraints of federated learning. By constructing a label correction model, a joint optimization problem is formally defined for optimizing both the label correction model and the client-side local training model (e.g., classification model). As a solution to this optimization problem, we propose a robustness training algorithm using label correction, along with a cross-validation data sampling algorithm that updates both models simultaneously. It is verified through experiments that the mechanism can effectively improve the model convergence accuracy on noisy datasets in federated learning scenarios. Jinghui Zhang 0001, Dingyang Lyu 0001, Qiangsheng Dai, Fa Xin, Fang Dong 0001 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2023 | Multi-Exit DNN Inference Acceleration Based on Multi-Dimensional Optimization for Edge IntelligenceabstractEdge intelligence, as a prospective paradigm for accelerating DNN inference, is mostly implemented by model partitioning which inevitably incurs the large transmission overhead of DNN's intermediate data. A popular solution introduces multi-exit DNNs to reduce latency by enabling early exits. However, existing work ignores the correlation between exit settings and synergistic inference, causing incoordination of device-to-edge. To address this issue, this paper first investigates the bottlenecks of executing multi-exit DNNs in edge computing and builds a novel model for inference acceleration with exit selection, model partition, and resource allocation. To tackle the intractable coupling subproblems, we propose a Multi-exit DNN inference Acceleration framework based on Multi-dimensional Optimization (MAMO). In MAMO, the exit selection subproblem is first extracted from the original problem. Then, bidirectional dynamic programming is employed to determine the optimal exit setting for an arbitrary multi-exit DNN. Finally, based on the optimal exit setting, a DRL-based policy is developed to learn joint decisions of model partition and resource allocation. We deploy MAMO on a real-world testbed and evaluate its performance in various scenarios. Extensive experiments show that it can adapt to heterogeneous tasks and dynamic networks, and accelerate DNN inference by up to 13.7x compared with the state-of-the-art. Fang Dong 0001, Huitian Wang, Dian Shen, Zhaowu Huang, Qiang He 0001, Jinghui Zhang 0001, Liangsheng Wen |
IEEE Trans. Mob. Comput. | 6 |
| 2022 | Federated Learning Client Selection Mechanism Under System and Data HeterogeneityabstractFederated learning (FL) has been proposed to train a global model by distributed architecture, while keeping the training data local. Owing to the large scale of clients in FL, all clients to participate in training is not feasible. The heterogeneity of clients, including system and data heterogeneity, also poses huge challenge to the client selection problem. Traditional client selection mechanisms can’t handle these heterogeneities effectively, which lead to poor training efficiency. Hence, this paper comprehensively considers system and data heterogeneity to select clients and dynamically adjust the number of selected clients. For system heterogeneity, we build the latency model to predict the training time for selecting clients with best performance including CPU frequency, the size of dataset and transmission power. Besides, for data heterogeneity, the cluster model is established to cluster clients for alleviating the accuracy jitter owing to the non independent and identically distributed (Non-IID) dataset. We formulate the client selection problem aiming to minimize the overall training time on the premise of accuracy, and design the Federated Client Cluster and latency-Prediction Selection (FCCPS) algorithm to solve this problem. With extensive simulations, we show that the FCCPS algorithm can reduce the training time by up to 21% on Cifar-10 dataset and 13% on FashionMNIST dataset, as compared to FedAvg. Fan Xin, Jinghui Zhang 0001, Junzhou Luo, Fang Dong 0001 |
CSCWD | 2 |
| 2022 | FedAda: Fast-convergent adaptive federated learning in heterogeneous mobile edge computing environment
Jinghui Zhang 0001, Jiahui Jin 0001, Aibo Song, Wei Zhao 0023, Liangsheng Wen |
World Wide Web | 1 |
| 2021 | PipePar: A Pipelined Hybrid Parallel Approach for Accelerating Distributed DNN TrainingabstractLarge scale DNN training tasks are exceedingly compute-intensive and time-consuming, which are usually executed on highly-parallel platforms. Data and model parallelization is a common way to speed up the training progress across devices. However, they tend to achieve sub-optimal performance due to the communication overheads and unbalanced load among servers. Recent emerging pipelining solutions mitigate the above issues, incorporating the advantages of data and model parallelism. In this paper, we make a step further towards optimizing the execution of pipelining. We introduce PipePar, a pipeline-parallel DNN training method that provides optimized execution strategies of layer-stacked DNNs. PipePar considers the entire tensor partition space of pipelining and explores potential hybrid parallel configurations of each stage in the pipeline. Additionally, we notice the network heterogeneity between different GPU servers and it is inevitable to transfer tensors with different bandwidths and latency. So, taking into account both computation and communication capacity of different GPU servers, PipePar is intended to find a elastic load distribution strategy at different levels. We evaluate PipePar with a set of real-world DNNs on 4 GPU servers. Our experimental results show that PipePar is able to find an efficient strategy that are up to 2.16× faster than state-of-the-art hybrid parallelization approaches. Jiange Li, Jinghui Zhang 0001, Jiahui Jin 0001, Fang Dong 0001 |
CSCWD | 3 |
| 2021 | Region-aware POI Recommendation with Semantic Spatial GraphabstractThe development of Location-Based Social Networks (LBSNs) offers an opportunity for understanding user preferences and promoting Point-of-Interest (POI) recommendation. The user preferences usually change when the environment changes, which will affect the performance of POI recommendation. Many existing methods extract the environmental features from static regions, but they cannot capture user preferences in real time with the fine-grained changes in user locations. Meanwhile, the similarity of POI categories, which is significant to capture user preferences, is usually ignored. To address these issues, we propose RegDM, a region-aware POI recommendation model that employs a semantic spatial graph to model the relations among POIs. With the semantic spatial graph, RegDM uses a Graph Neural Network (GNN) to extract fine-grained region features and user preferences for personalized recommendation. We evaluate RegDM on two datasets, and the experiment results demonstrate the effectiveness of our model. Jiakai Tang, Jiahui Jin 0001, Zijia Miao, Binjie Zhang, Jinghui Zhang 0001 |
CSCWD | 6 |
| 2021 | Towards Tunable RDMA Parameter Selection at Runtime for Datacenter ApplicationsabstractBecause of the low-latency and high-throughput benefits of RDMA, an increasing number of collaborative applications in datacenters are re-designed with RDMA to boost the performance. Among various low-level hardware primitives provided by RDMA, exposed as parameters of APIs, the application designers select and hardcode them to exploit all the performance benefits of RDMA. However, with the dynamic nature of datacenter application, the hardcoded and fixed parameter selection fails to take full advantages of RDMA capabilities, which can cause up to 35% throughput performance loss. To address this issue, we present a tunable RDMA parameter selection framework, which allows parameter tuning at runtime, adaptive to the dynamic application and server status. To attain the native RDMA performance, we use a lightweight decision tree to reduce the overhead of RDMA parameter selection. Finally, we implement the tunable RDMA parameter selection framework with native RDMA API to provide a more abstract API. To demonstrate the effectiveness of our method, we implement a key-value service based on the abstract API. Experiment results show that our implementation has only a very small overhead compared with the native RDMA, while the optimized key-value service achieves 112% more throughput than Pilaf and 66% more throughput than FaRM. Fang Dong 0001, Dian Shen, Chengtian Zhang, Jinghui Zhang 0001, Junzhou Luo |
CSCWD | 5 |
| 2020 | Optimizing execution for pipelined-based distributed deep learning in a heterogeneously networked GPU clusterabstractSummary Exorbitant resources (computing and memory) are required to train a deep neural network (DNN). Often researchers deploy an approach that uses distributed parallel training to acquire larger models faster on GPUs. This approach has its detriments, though; on one hand, a GPU's expanded capacity to compute also produces bigger bottlenecks in inter‐GPU's communications during model training, and multi‐GPU systems lead to complex connectivity. Workload schedulers then end up having to consider hardware topology and requirements for workload communication, in hopes of allocating GPU resources to optimize execution time and improve usage in a heterogeneous environment. On the other hand, the high memory requirements to train a DNN model make running the training processes on GPUs onerous. To contend with this, we introduce two execution optimization methods based on pipeline‐hybrid parallelism (using both data and model parallelism) in a GPU cluster with heterogeneous networking. First, we propose a model partition algorithm that accelerates pipeline‐hybrid parallelism training between heterogeneously network‐connected GPUs. Second, we introduce a cost‐balanced recomputing algorithm to reduce memory usage in the pipeline mode. Experiments show that our solution (Pipe‐Torch) averages a speedup of 1.4× compared with data parallelism, and reduces the memory footprint while maintaining pipelined load‐balanced training. Jinghui Zhang 0001, Jiange Li, Jiahui Jin 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2019 | MBECN: Enabling ECN with Micro-burst Traffic in Multi-queue Data CenterabstractModern multi-queue data centers often use the standard Explicit Congestion Notification (ECN) scheme to achieve high network performance. However, one substantial drawback of this approach is that micro-burst traffic can cause the instantaneous queue length to exceed the ECN's threshold, resulting in numerous mismarkings. After enduring too many mismarkings, senders may overreact, leading to severe throughput loss. As a solution to this dilemma, we propose our own adaptationthe Micro-burst ECN (MBECN) scheme-to mitigate mismarking. MBECN finds a more appropriate threshold baseline for each queue to absorb micro-bursts, based on steady-state analysis and an ideal generalized processor sharing (GPS) model. By adopting a queue-occupation-based dynamically adjusting algorithm, MBECN effectively handles packet backlog without hurting latency. Through testbed experiments, we find that MBECN improves throughput by ~20% and reduces flow completion time (FCT) by ~40%. Using large scale simulations, we find that throughput can be improved by 1.5~2.4× with DCTCP and 1.26~1.35× with ECN*. We also measure network delay and find that latency only increases by 7.36%. Kexi Kang, Jinghui Zhang 0001, Jiahui Jin 0001, Dian Shen, Junzhou Luo, Zhiang Wu 0001 |
CLUSTER | 2 |
| 2019 | QAECN: Dynamically Tuning ECN Threshold with Micro-burst in Multi-queue Data CentersabstractPacket loss is a common problem in data center networks. The factors causing packet loss are various. Among them, micro-burst is the most important reason. Some previous works have studied the causes and influence o f micro-burst in single queue data center. However, through simulations and experiments, we find that micro-burst could bring m ore serious performance degradation in multi-queue data centers. The micro-burst traffic could cause E CN marking ratio rising from 4% to 22%, and cause throughput loss by up to 40%. Through observing queue length, we find that the standard E CN, which adopts immutable threshold, is not suitable for micro-burst traffic because micro-burst could trigger spurious congestion signals frequently, especially in DCTCP. In this paper, we not only show how much influence the micro-burst brings, but also propose Queue-length Aware ECN (QA-ECN) scheme to mitigate micro-burst. Finally, the simulations and experiments show that QAECN could reduce ECN marking ratio to 2.5%. In addition, the throughput and flow completion time could be improved by up to 22.9% and 34.1%, respectively. Kexi Kang, Jinghui Zhang 0001, Jiahui Jin 0001, Dian Shen, Runqun Xiong, Junzhou Luo |
CSCWD | 2 |
| 2019 | Graph partition-based data and task co-scheduling of scientific workflow in geo-distributed datacentersabstractSummary Most large‐scale scientific workflows take place in multiple collaborative datacenters for access to community‐wide resources, while adhering to each datacenter's non‐uniform resource limits. However, moving both initial input datasets with predetermined locations and intermediate datasets needing placement decisions across geo‐distributed datacenters hinders efficient execution of large‐scale data‐intensive scientific workflows. Thus, scientific workflow's data and task co‐scheduling deal with situations such as pre‐placed initial input datasets, placement of intermediate datasets and each datacenter's non‐uniform computation and storage constraint, while minimizing the cross‐datacenter data transfer. Since this scheduling problem is known to be NP‐hard, here, we propose a novel approach, based on the multilevel graph coarsening and uncoarsening framework, together with a specialized hybrid genetic algorithm having distinctive graph partition driven features of repair and local improvement, for scheduling data‐intensive scientific workflows in geo‐distributed datacenters and optimizing the cross‐datacenter data transfer volume. Extensive simulations, based on four real‐world workflow traces, show that our algorithm significantly reduces the overall geo‐distributed data transfer and demonstrate its effectiveness. Jinghui Zhang 0001, Jiahui Jin 0001, Aibo Song |
Concurr. Comput. Pract. Exp. | 1 |
| 2019 | A data-locality-aware task scheduler for distributed social graph queries
Jiahui Jin 0001, Junzhou Luo, Mingyang Du, Yongcheng Dang, Jinghui Zhang 0001, Aibo Song |
Future Gener. Comput. Syst. | 6 |
| 2016 | Resource provisioning optimization for service hosting on cloud platformabstractWith the popularity of cloud computing technology, service hosting is used as a typical model to deploy different kinds of services on cloud platform. In recent years, how to effectively provide resources for service hosting has attracted more and more attention. However, most of the existing works only focused on how to effectively provide virtual machines for service hosting. They ignored how to efficiently place these virtual machines into physical servers, when considering multidimensional resource requirements. This may result in unreasonable virtual machine placement in servers, thereby causing the underutilization of resource. To address this problem, we propose a novel resource provisioning method including virtual machine provisioning for hosting service and virtual machine placement in servers. The proposed method decides how many virtual machines should be provided for each service by utilizing queuing theory. Then based on the virtual machines to be provided, the proposed method models the virtual machine placement problem as a variant of cutting stock problem, and decides how many servers should be provided by solving this problem. The proposed method is evaluated by simulations. Experimental results show the proposed method achieves a better performance than these baseline methods. Jiyuan Shi, Fang Dong 0001, Jinghui Zhang 0001, Jiahui Jin 0001, Junzhou Luo |
CSCWD | 3 |
| 2015 | Two-Phase Online Virtual Machine Placement in Heterogeneous Cloud Data CenterabstractWith the rapid development and popularity of cloud computing technology, more and more Collaborative Virtual Environment (CVE) systems are migrated to cloud computing environment to improve the effectiveness of resource usage. Virtual Machine (VM) placement in cloud data center is a key issue of providing high-efficient cloud platform for CVE system. However, most existing VM placement algorithms ignore the following characteristics of actual cloud environment: (1) VMs deployment requests arrive and leave dynamically, (2) Cloud data center usually consists of many heterogeneous Physical Machines (PMs). Ignoring these two characteristics result in an inefficient and unbalanced use of multiple resources of PMs. Thus using these algorithms directly will lead to a poor resource utilization. In this article, we propose a two-phase online VM placement algorithm, which helps the cloud data center to minimize different resource usages and aims at a more efficient use of multiple resources. Our algorithm selects the most suitable PM type for VM based on Cosine Similarity, and adaptively maps VMs to PMs by using an approximation algorithm. The proposed algorithm is evaluated by simulations. Experimental results show our proposed algorithm ensures a more efficient use of multiple resources over the existing approaches. Jiyuan Shi, Fang Dong 0001, Jinghui Zhang 0001, Junzhou Luo, Ding Ding 0002 |
SMC | 3 |
| 2015 | Towards optimized scheduling for data-intensive scientific workflow in multiple datacenter environmentabstractSummary In the big data era, scientific workflow exhibits the characteristics of data intensity and becomes increasingly popular in scientific domains. Efficient scheduling of data‐intensive scientific workflow in a multiple datacenter (DC) environment has been a long‐standing challenge. Most of previous work on data‐intensive scientific workflow scheduling primarily focused on the optimization of reducing the volumes of data transfer between workflow tasks. In this paper, novel scheduling strategies for the execution of data‐intensive scientific workflow in multi‐DC environment are proposed aiming at the optimization of the overall data transfer time. A novel DC selection approach is proposed to minimize the number of DCs having enough storage capacity for the execution of scientific workflow as well as optimized inter‐DC network bandwidth for efficient data transfer between workflow tasks. A k‐means clustering‐based data placement strategy is adopted to intelligently place the initial data of scientific workflow thereby reducing the volume of initial data transfer between different DCs. A multilevel task replication scheduling strategy is invented to reduce the volumes of intermediate data transfer between DCs during the runtime of the scientific workflow. Simulations spanning a broad range of scientific workflow and multi‐DC settings are performed in order to verify the proposed approaches. The numerical results show that our combined scheduling strategy significantly reduces the overall data transfer time and data transfer volume when scientific workflow is scheduled in multi‐DC environment. Copyright © 2015 John Wiley & Sons, Ltd. Jinghui Zhang 0001, Junzhou Luo, Fang Dong 0001, Junxue Zhang 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2014 | A budget and deadline aware scientific workflow resource provisioning and scheduling mechanism for cloudabstractCurrently in large-scale scientific experiments, scientists often submit scientific workflow jobs at different time. From the view of system, the entire workload is a stream of jobs submitted at an unpredictable time and different job has different priority and deadline. Moreover the cost of performing these jobs cannot exceed a certain budget constraint. Therefore how to perform scientific workflow applications efficiently in cloud has become the urgent problem. However most of existing work didn't consider unpredictable submission time of jobs, as well as budget and deadline constrains. In this paper, we design an elastic resource provisioning and task scheduling mechanism to perform scientific workflows in cloud. Our goal is to complete as many high-priority workflows as possible under budget and deadline constrains. This mechanism consists of three phases: workflow preprocessing, elastic resource provisioning and task scheduling. We perform evaluation with real AMS experiment scientific computing data under different budget constraints. We also consider inaccurate task execution time, VM provisioning delays and task failures in evaluation. The results show that our mechanism achieves a better performance than these reference mechanisms. In addition, the inaccurate task execution time, VM provisioning delays, and task failures do not bring significant impact to mechanism's performance. Jiyuan Shi, Junzhou Luo, Fang Dong 0001, Jinghui Zhang 0001 |
CSCWD | 4 |
| 2013 | Scheduling Parallel Task Graphs on non-dedicated heterogeneous multicluster platform with Moldable Task DuplicationabstractWorkflow applications structured as Parallel Task Graphs (PTG) exhibit both data and task parallelism and arise in scientific and industrial domains. Most of previous works regarding PTG scheduling only target dedicated multicluster platform. In this paper we develop a scheduling algorithm, MTD (Moldable Task Duplication with forward migration of duplicated predecessors), which applies to non-dedicated heterogeneous multicluster platforms. Our novel contribution is that in MTD, dynamic critical task determination accounts for the heterogeneity and fluctuations of multicluster platform within the hypothetical deadline, and the strategy of moldable task duplication with forward migrations of duplicated predecessors is invented to fully exploit the flexibility of data-parallel tasks. Simulations show that our approach can achieve better average PTG makespan than its competitors. Jinghui Zhang 0001, Junzhou Luo, Fang Dong 0001 |
CSCWD | 1 |
| 2013 | Scheduling of scientific workflow in non-dedicated heterogeneous multicluster platform
Jinghui Zhang 0001, Junzhou Luo, Fang Dong 0001 |
J. Syst. Softw. | 1 |
| 2011 | Scheduling mixed-parallel application onto multicluster grid with background workloadsabstractIn a shared multicluster grid where mixed-parallel application workload and background workload co-exist, available processors for executing mixed-parallel application workload are those time-varying residual processors after reservations for background workloads. The grid scheduler aims to minimize the makespan of the mixed-parallel application and allocate processors to tasks belong to the mixed-parallel application with a coordinated manner, whereas advance reservation capability from Local Resource Manager(LRM) are fully exploited. We develop a heuristic algorithm MHEFT-RSV (ReSerVation) adapted from MHEFT (Mixed-Heterogeneous Earliest Finish Time) to multicluster grid with background workloads. Based on MHEFT-RSV, we further propose an exact branch-and-cut scheduling algorithm, which exploits the intertask precedence and resource constraints as much as possible, to accelerate the process of obtaining the schedule with minimized makespan. From the detailed simulation experiment we find that on average the exact branch-and-cut algorithm obtains shorter or equal makespan as MHEFT-RSV while MHEFT-RSV achieves better tradeoff between makespan and computation time. Jinghui Zhang 0001, Junzhou Luo |
CSCWD | 1 |
| 2008 | Agent based automated negotiation for gridabstractIn a dynamic service oriented environment like grid, automated negotiation between service providers and consumers becomes a crucial issue that aims to minimize manual intervening for grid environments. Currently the foundation of automatic negotiation framework is missing in grid, as well as a widely adopted grid negotiation model. In this paper, a multi-agent based automated negotiation framework is introduced into the grid environment to facilitate automated negotiation, with the advantage of its independence from concrete grid negotiation models. A refined negotiation model for grid environment is proposed as a substantial negotiation implementation based on the proposed framework and the result suggests that agent based grid automated negotiation bring remarkable flexibility. Jinghui Zhang 0001, Junzhou Luo |
CSCWD | 1 |
| 2007 | An Adaptive QoS Group Guided Grid Scheduling Algorithm with Task ReplicasabstractTo balance resource loads and minimize makespan are two vital goals in grid scheduling. However, it comes to be difficult for the dynamicity of grid resources, especially with meeting the QoS requirements of tasks considered. In this paper we propose a scheduling algorithm called the QoS group guided grid scheduling algorithm with task replicas (QGTR). The proposed algorithm makes scheduling decision that bases on recent QoS status feedback of resources, dividing tasks into groups according to the feedback QoS status of resources and then scheduling tasks in different groups to resources accordingly. In addition task replica is adopted to improve the resource utilization and gain a better schedule result. The simulation results show QGTR can effectively reduce the makespan and enhance the resource utilization in addition to meeting the QoS requirement of tasks with its best efforts. Jinghui Zhang 0001, Junzhou Luo |
CSCWD | 1 |