Xiao Zhang 0015

dblp:49/4478-15 · DBLP profile ↗
← Back
62ranked-venue papers
12as first author
48since 2021 · last 2026
0000-0003-0824-9284ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 1 first-author · 15 since 2021Computer networks · 16 · 1 first-author · 14 since 2021Databases, data management, data science and information retrieval · 14 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 5 since 2021Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Theory of computation · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 DGTF: Cross-Domain Decentralized Graph Learning with Topology-Aware Knowledge Fusion
abstract
Cross-Domain Decentralized Graph Learning (CD-DGL) is a promising paradigm that enables efficient, privacy-preserving collaboration among multiple parties to unlock the value of cross-domain graph data. However, it faces two fundamental challenges. First, inconsistent label spaces across domains drive local models to learn domain-specific biases, which means domain-invariant topological knowledge extraction beyond label constraints is difficult. Second, existing domain topology shift and heterogeneous model architectures make direct model aggregation infeasible. To address these issues, we first use Extended Persistent Homology (EPH) to reveal and quantify the problem of domain topology shift induced by the cross-domain setting. Building on this insight, we present Decentralized Graph Learning with Topology-Aware Knowledge Fusion (DGTF), a novel framework designed to facilitate positive topological knowledge transfer in CD-DGL. Our framework achieves this by integrating two core strategies: first, a contrastive learning-based approach to extract task-agnostic topological knowledge, and second, a topology-aware, model-independent knowledge fusion method to effectively integrate this topological information. Extensive experiments conducted under various cross-domain and model-heterogeneous settings validate the superiority and effectiveness of our proposed framework.
Ruisheng Zheng, Xiao Zhang 0015, Hongjian Shi, Yanjie Fu, Yuan Yuan 0040, Dongxiao Yu
AAAI3
2026 Resource-Aware Decentralized Learning with Rate-Adaptive Quantization
Jing Qiao, Yu Liu 0085, Yuan Yuan 0040, Yifei Zou, Xiao Zhang 0015, Dongxiao Yu
INFOCOM5
2026 FedPDM: Representation enhanced federated learning with privacy preserving diffusion models
Fuzhen Zhuang, Yiqi Tong, Xiao Zhang 0015, Zhaojun Hu, Jiejie Zhao, Jin Dong 0004
Knowl. Based Syst.4
2026 A provably robust framework for distributed minimax learning in adversarial environments
Xiao Zhang 0015, Yingfan Deng, Yifei Zou, Zhipeng Cai 0001, Dongxiao Yu
Theor. Comput. Sci.3
2026 Unity is Power: Semi-Asynchronous Collaborative Training of Large-Scale Models With Structured Pruning in Resource-Limited Clients
abstract
In this work, we study to release the potential of massive heterogeneous weak computing power to collaboratively train large-scale models on dispersed datasets. In order to improve both efficiency and accuracy in resource-adaptive collaborative learning, we take the first step to consider the unstructured pruning, varying submodel architectures, knowledge loss, and straggler challenges simultaneously. We propose a novel semiasynchronous collaborative training framework, namely Co-S2P, with data distribution-aware structured pruning and cross-block knowledge transfer mechanism to address the above concerns. Furthermore, we provide theoretical proof that Co-S2P can achieve asymptotic optimal convergence rate of O(1/√ N∗EQ). Finally, we conduct extensive experiments on two types of tasks with a real-world hardware testbed including diverse IoT devices. The experimental results demonstrate that Co-S2P improves accuracy by up to 8.8% and resource utilization by up to 1.2× compared to state-of-the-art methods, while reducing memory consumption by approximately 22% and training time by about 24% on all resource-limited devices.
Xiao Zhang 0015, Feng Chen 0005, Yuan Yuan 0040, Yifei Zou, Mengying Zhao, Jianbo Lu 0001, Dongxiao Yu
IEEE Trans. Mob. Comput.2
2026 Federated Bilevel Learning Against Model Poisoning Attacks
Yuan Yuan 0040, Yingfan Deng, Xiao Zhang 0015, Yifei Zou, Yangguang Shi, Dongxiao Yu
IEEE Trans. Netw.3
2025 A Robust Distributed Minimax Learning Method Against Model Poisoning Attacks
Yuan Yuan 0040, Xiao Zhang 0015, Yifei Zou, Zhipeng Cai 0001, Dongxiao Yu
COCOON (2)3
2025 H2Tune: Federated Foundation Model Fine-Tuning with Hybrid Heterogeneity
abstract
Different from existing federated fine-tuning (FFT) methods for foundation models, hybrid heterogeneous federated fine-tuning (HHFFT) is an under-explored scenario where clients exhibit double heterogeneity in model architectures and downstream tasks. This hybrid heterogeneity introduces two significant challenges: 1) heterogeneous matrix aggregation, where clients adopt different large-scale foundation models based on their task requirements and resource limitations, leading to dimensional mismatches during LoRA parameter aggregation; and 2) multi-task knowledge interference, where local shared parameters, trained with both task-shared and task-specific knowledge, cannot ensure only task-shared knowledge is transferred between clients. To address these challenges, we propose H2Tune, a federated foundation model fine-tuning with hybrid heterogeneity. Our framework H2Tune consists of three key components: (i) sparsified triple matrix decomposition to align hidden dimensions across clients through constructing rank-consistent middle matrices, with adaptive sparsification based on client resources; (ii) relation-guided matrix layer alignment to handle heterogeneous layer structures and representation capabilities; and (iii) alternating task-knowledge disentanglement mechanism to decouple shared and specific knowledge of local model parameters through alternating optimization. Theoretical analysis proves a convergence rate of O(1/√T). Extensive experiments show our method achieves up to 15.4% accuracy improvement compared to state-of-the-art baselines.
Yiqi Tong, Zhaojun Hu, Fuzhen Zhuang, Xiao Zhang 0015, Jin Dong 0004
ECAI6
2025 Adaptive Strategy Weighting with Fault Tolerant Localization for Object Navigation
abstract
End-to-end navigation models commonly incorporate multiple sub-modules, each designed for distinct purposes such as searching, obstacle avoidance, and target localization. However, agents equipped with these modules may still struggle to apply the appropriate strategies at the right locations and stages. For instance, agent might incorrectly rely on the search or localization module for obstacle avoidance, reducing adaptability in dynamic environments. Additionally, existing methods assume the recognition for target object is always correct, neglecting the unavoidable misclassification caused by visually similar objects. To apply appropriate strategy for a given situation, we introduce Adaptive Strategy Feature Fusion (ASFF). It heuristically assigns appropriate weights to different sub-modules based on current observation and memory state, enabling flexible integration with arbitrary sub-module combinations. To improve localization in the presence of misclassification, we propose Fault Tolerant Target Memory Aggregator (FTTMA), a module that uses clustering-based sparse self-attention and target cross-attention to minimize interference from misclassified object, providing accurate target orientation to the agent. Experiments on the AI2THOR and RoboTHOR datasets, including both typical and zero-shot navigation tasks, demonstrate that our model outperforms the state-of-the-art (SOTA) methods in both success rate and navigation efficiency.
Yanwei Zheng, Shaopu Feng, Changrui Li, Xiao Zhang 0015, Dongxiao Yu
ICME5
2025 How Distributed Collaboration Influences the Diffusion Model Training? A Theoretical Perspective
abstract
This paper examines the theoretical performance of distributed diffusion models in environments where computational resources and data availability vary significantly among workers. Traditional models centered on single-worker scenarios fall short in such distributed settings, particularly when some workers are resource-constrained. This discrepancy in resources and data diversity challenges the assumption of accurate score function estimation foundational to single-worker models. We establish the inaugural generation error bound for distributed diffusion models in resource-limited settings, establishing a linear relationship with the data dimension $d$ and consistency with established single-worker results. Our analysis highlights the critical role of hyperparameter selection in influencing the training dynamics, which are key to the performance of model generation. This study provides a streamlined theoretical approach to optimizing distributed diffusion models, paving the way for future research in this area.
Jing Qiao, Yu Liu 0085, Yuan Yuan 0040, Xiao Zhang 0015, Zhipeng Cai 0001, Dongxiao Yu
ICML4
2025 PDUDT: Provable Decentralized Unlearning under Dynamic Topologies
abstract
This paper investigates decentralized unlearning, aiming to eliminate the impact of a specific client on the whole decentralized system. However, decentralized communication characterizations pose new challenges for effective unlearning: the indirect connections make it difficult to trace the specific client's impact, while the dynamic topology limits the scalability of retraining-based unlearning methods. In this paper, we propose the first **P**rovable **D**ecentralized **U**nlearning algorithm under **D**ynamic **T**opologies called PDUDT. It allows clients to eliminate the influence of a specific client without additional communication or retraining. We provide rigorous theoretical guarantees for PDUDT, showing it is statistically indistinguishable from perturbed retraining. Additionally, it achieves an efficient convergence rate of $\mathcal{O}(\frac{1}{T})$ in subsequent learning, where $T$ is the total communication rounds. This rate matches state-of-the-art results. Experimental results show that compared with the Retrain method, PDUDT saves more than 99\% of unlearning time while achieving comparable unlearning performance.
Jing Qiao, Yu Liu 0085, Zengzhe Chen, Yuan Yuan 0014, Xiao Zhang 0015, Dongxiao Yu
ICML6
2025 FEZE: Alignment-Flexible Zero-Shot Vertical Federated Learning
abstract
Different from existing vertical federated learning (VFL), zero-shot VFL (ZVFL) is an under-explored scenario where test classes are absent from partial parties' training sets. In extreme cases, some test classes even have no training samples for all parties. Traditionally, existing zero-shot methods require abundant seen-class samples for effective knowledge transfer to recognize unseen classes. However, both the limited aligned samples and different seen classes pose several unique challenges to ZVFL. The primary challenge lies in the seen-to-unseen transfer insufficiency, as the scarcity of aligned samples and diverse seen-class distributions across parties severely limits the model's capability to learn discriminative features that can generalize to unseen classes. Moreover, the multi-party bias inconsistency arises as different parties tend to be biased towards their own seen classes during prediction, leading to skewed classification results at the active party. To address these challenges, we propose FEZE, an alignment-flexible zero-shot vertical federated learning framework. Specifically, we introduce a relation learning network to capture class-feature relationships between class labels and feature representations across heterogeneous feature spaces, enabling unseen class recognition through relationship inference. Additionally, we design a meta-relation learning mechanism that leverages diverse class-feature patterns to tackle the insufficient feature generalization from limited seen-class samples. Finally, we propose an alignment-flexible adaptive aggregation strategy that achieves adaptively aggregation based on inconsistent prediction spaces with arbitrary number of aligned samples. Theoretical analysis proves that FEZE can achieve a convergence rate of O(1/T). In the most challenging zero-shot scenario without aligned samples, FEZE surpasses state-of-the-art baselines by an average of 7.47% across three datasets.
Yiqi Tong, Yiyang Duan, Fuzhen Zhuang, Xiao Zhang 0015, Zhaojun Hu, Jin Dong 0004
KDD (2)5
2025 Efficient Strategy Learning by Decoupling Searching and Pathfinding for Object Navigation
abstract
Inspired by human-like behaviors for navigation: first searching to explore unknown areas before discovering the target, and then the pathfinding of moving towards the discovered target, recent studies design parallel submodules to achieve different functions in the searching and pathfinding stages, while ignoring the differences in reward signals between the two stages. As a result, these models often cannot be fully trained or are overfitting on training scenes. Another bottleneck that restricts agents from learning two-stage strategies is spatial perception ability, since the studies used generic visual encoders without considering the depth information of navigation scenes. To release the potential of the model on strategy learning, we propose the Two-Stage Reward Mechanism (TSRM) for object navigation that decouples the searching and pathfinding behaviours in an episode, enabling the agent to explore larger area in searching stage and seek the optimal path in pathfinding stage. Also, we propose a pretraining method Depth Enhanced Masked Autoencoders (DE-MAE) that enables agent to determine explored and unexplored areas during the searching stage, locate target object and plan paths during the pathfinding stage more accurately. In addition, we propose a new metric of Searching Success weighted by Searching Path Length (SSSPL) that assesses agent’s searching ability and exploring efficiency. Finally, we evaluated our method on AI2-Thor and RoboTHOR extensively and demonstrated it can outperform the state-of-the-art (SOTA) methods in both the success rate and the navigation efficiency.
Yanwei Zheng, Shaopu Feng, Chuanlin Lan, Xiao Zhang 0015, Dongxiao Yu
SMC5
2025 Pruning-Based Adaptive Federated Learning at the Edge
abstract
Federated Learning (FL) is a new learning framework in which$s$clients collaboratively train a model under the guidance of a central server. Meanwhile, with the advent of the era of large models, the parameters of models are facing explosive growth. Therefore, it is important to design federated learning algorithms for edge environment. However, the edge environment is severely limited in computing, storage, and network bandwidth resources. Concurrently, adaptive gradient methods show better performance than constant learning rate in non-distributed settings. In this paper, we propose a pruning-based distributed Adam (PD-Adam) algorithm, which combines model pruning and adaptive learning steps to achieve asymptotically optimal convergence rate of$O(1/\sqrt[4]{K})$. At the same time, the algorithm can achieve convergence consistent with the centralized model. Finally, extensive experiments have confirmed the convergence of our algorithm, demonstrating its reliability and effectiveness across various scenarios. Specially, our proposed algorithm is$2$% and$18$% more accurate than the current state-of-the-art FedAvg algorithm on the ResNet and CIFAR datasets.
Dongxiao Yu, Yuan Yuan 0014, Yifei Zou, Xiao Zhang 0015, Yu Liu 0085, Li-Zhen Cui 0001, Xiuzhen Cheng
IEEE Trans. Computers4
2025 BaDFL: Mitigating Model Poisoning in Decentralized Federated Learning
abstract
Decentralized federated learning (DFL) has gained significant attention due to its ability to facilitate collaborative model training without relying on a central server. However, it is highly vulnerable to backdoor attacks, where malicious participants can manipulate model updates to embed hidden functionalities. In this paper, we propose BaDFL, a novel Backdoor Attack defense mechanism for Decentralized Federated Learning. BaDFL enhances robustness by applying strategic model clipping at the local update level. To the best of our knowledge, BaDFL is the first decentralized federated learning algorithm with theoretical guarantees against model poisoning attacks. Specifically, BaDFL achieves an asymptotically optimal convergence rate of$O(\frac{1}{\sqrt{nT}})$, wherenis the number of nodes andTis the global maximum iteration number. Furthermore, we provide a comprehensive analysis under two different attack scenarios, showing that BaDFL maintains robustness within a specific defense radius. Extensive experimental results show that, on average, BaDFL can effectively defend against model poisoning within 6 mitigation rounds, with less than a 1% drop in accuracy.
Yuan Yuan 0014, Anhao Zhou, Xiao Zhang 0015, Yifei Zou, Yangguang Shi, Dongxiao Yu
IEEE Trans. Computers3
2025 Temporal-Spatial Object Relations Modeling for Vision-and-Language Navigation
abstract
Vision-and-Language Navigation (VLN) is a challenging task where an agent is required to navigate to a natural language described location via vision observations. The navigation abilities of the agent can be enhanced by the relations between objects, which are usually learned using internal objects or external datasets. The relationships between internal objects are modeled employing graph convolutional network (GCN) in traditional studies. However, GCN tends to be shallow, limiting its modeling ability. To address this issue, we utilize a cross attention mechanism to learn the connections between objects over a trajectory, which takes temporal continuity into account, termed as Temporal Object Relations (TOR). The external datasets have a gap with the navigation environment, leading to inaccurate modeling of relations. To avoid this problem, we construct object connections based on observations from all viewpoints in the navigational environment, which ensures complete spatial coverage and eliminates the gap, called Spatial Object Relations (SOR). Additionally, we observe that agents may repeatedly visit the same location during navigation, significantly hindering their performance. For resolving this matter, we introduce the Turning Back Penalty (TBP) loss function, which penalizes the agent’s repetitive visiting behavior, substantially reducing the navigational distance. Experimental results on the REVERIE, SOON, Touchdown and R2R datasets demonstrate the effectiveness of the proposed method.
Yanwei Zheng, Dongchen Sui, Chuanlin Lan, Xinpeng Zhao 0001, Xiao Zhang 0015, Jingke Meng, Mengbai Xiao, Yifei Zou, Dongxiao Yu
IEEE Trans. Intell. Transp. Syst.6
2025 Convergence-Guaranteed Federated Learning through Gradient Trajectory Smoothing with Triple-Objective Decomposition
abstract
Federated Learning (FL) has been widely adopted as a distributed machine learning paradigm aiming to derive a global model without transferring local data to the server. In the context of heterogeneous environments typical of many FL deployments, our research has identified the performance oscillation problem in existing FL methods, resulting in slow convergence and severe performance drop. In this article, we first investigate the global optimizing objective in FL and demonstrate that, due to data heterogeneity and partial client participation, the global updates in a single training epoch may diverge from the intended objectives of conventional FL methods. To address this problem, we introduce a triple-objective decomposition mechanism to decompose the overarching global objective into three distinct local objectives aimed at aligning client gradients. Subsequently, we propose a gradient trajectory smoothing technique known as FedGTS, which refines local updates by estimating a pseudo-gradient leveraging historical global update trajectories. This approach is designed to mitigate performance oscillations and enhance the stability of the learning process. We theoretically demonstrate that our approach reduces variance of local updates and achieves a guaranteed convergence rate. We experimentally show that the proposed method outperforms the baselines with faster convergence and higher accuracy. Extensive experiments validate the effectiveness of the proposed approach across various heterogeneity settings. Our codes are publicly available at GitHub ( https://github.com/ZongHR/FedGTS ).
Haoran Zong, Xiao Zhang 0015, Jianhui Duan, Derun Zou
ACM Trans. Knowl. Discov. Data2
2025 Action-Aware Visual-Textual Alignment for Long-Instruction Vision-and-Language Navigation
abstract
Traditional Vision-and-Language Navigation (VLN) requires an agent to navigate to a target location solely based on visual observations, guided by natural language instructions. Compared to this task, long-instruction VLN involves longer instructions, extended trajectories, and the need to consider more contextual information for global path planning. As a result, it is more challenging and requires accurately aligning the instructions with the agent’s current visual observations, which is accompanied by two significant issues. Firstly, there is a misalignment between actions. The visual observations of the agent at each step lack explicit action-related details, while the instructions contain action-oriented words. Secondly, there is a misalignment between global instructions and local visual observations. The instructions describe the entire navigation trajectory, whereas the agent’s visual observations only provide localized information about a specific position along the trajectory. To address these issues, this article introduces the Action-Perception Alignment Framework (APAF). In this framework, we first design the Action-Contextual Encoding Module (ACEM), which enriches the agent’s visual perception by encoding potential actions with relative heading and elevation angles. We then propose the Dynamic Instruction Weighting Module (DIWM), which adjusts the importance of instruction words based on the agent’s current visual observations, emphasizing those words most relevant to the agent’s visual observations. Our approach significantly outperforms existing methods, achieving state-of-the-art results with improvements of 8.5% and 4.0% in Success Rate (SR) on the long-instruction R4R and RxR datasets, respectively.
Yanwei Zheng, Chuanlin Lan, Dongchen Sui, Xinpeng Zhao 0001, Xiao Zhang 0015, Mengbai Xiao, Dongxiao Yu
ACM Trans. Multim. Comput. Commun. Appl.6
2025 TAMO:Fine-Grained Root Cause Analysis via Tool-Assisted LLM Agent With Multi-Modality Observation Data in Cloud-Native Systems
abstract
Implementing large language models (LLMs)-driven root cause analysis (RCA) in cloud-native systems has become a key topic of modern software operations and maintenance. However, existing LLM-based approaches face three key challenges: multi-modality input constraint, context window limitation, and dynamic dependence graph. To address these issues, we propose a tool-assisted LLM agent with multi-modality observation data for fine-grained RCA, namely TAMO, including multi-modality alignment tool, root cause localization tool, and fault types classification tool. In detail, TAMO unifies multi-modal observation data into time-aligned representations for cross-modal feature consistency. Based on the unified representations, TAMO then invokes its specialized root cause localization tool and fault types classification tool for further identifying root cause and fault type underlying system context. This approach overcomes the limitations of LLMs in processing real-time raw observational data and dynamic service dependencies, guiding the model to generate repair strategies that align with system context through structured prompt design. Experiments on two benchmark datasets demonstrate that TAMO outperforms state-of-the-art (SOTA) approaches with comparable performance.
Xiao Zhang 0015, Yuan Yuan 0040, Mengbai Xiao, Fuzhen Zhuang, Dongxiao Yu
IEEE Trans. Serv. Comput.1
2024 Self-Paced Unified Representation Learning for Hierarchical Multi-Label Classification
abstract
Hierarchical Multi-Label Classification (HMLC) is a well-established problem that aims at assigning data instances to multiple classes stored in a hierarchical structure. Despite its importance, existing approaches often face two key limitations: (i) They employ dense networks to solely explore the class hierarchy as hard criterion for maintaining taxonomic consistency among predicted classes, yet without leveraging rich semantic relationships between instances and classes; (ii) They struggle to generalize in settings with deep class levels, since the mini-batches uniformly sampled from different levels ignore the varying complexities of data and result in a non-smooth model adaptation to sparse data. To mitigate these issues, we present a Self-Paced Unified Representation (SPUR) learning framework, which focuses on the interplay between instance and classes to flexibly organize the training process of HMLC algorithms. Our framework consists of two lightweight encoders designed to capture the semantics of input features and the topological information of the class hierarchy. These encoders generate unified embeddings of instances and class hierarchy, which enable SPUR to exploit semantic dependencies between them and produce predictions in line with taxonomic constraints. Furthermore, we introduce a dynamic hardness measurement strategy that considers both class hierarchy and instance features to estimate the learning difficulty of each instance. This strategy is achieved by incorporating the propagation loss obtained at each hierarchical level, allowing for a more comprehensive assessment of learning complexity. Extensive experiments on several empirical benchmarks demonstrate the effectiveness and efficiency of SPUR compared to state-of-the-art methods, especially in scenarios with missing features.
Zixuan Yuan, Hao Liu 0026, Haoyi Zhou, Xiao Zhang 0015, Hao Wang 0073, Hui Xiong 0001
AAAI5
2024 BR-DeFedRL: Byzantine-Robust Decentralized Federated Reinforcement Learning with Fast Convergence and Communication Efficiency
abstract
In this paper, we propose Byzantine-Robust Decentralized Federated Reinforcement Learning (BR-DeFedRL), an innovative framework that effectively combats the harmful influence of Byzantine agents by adaptively adjusting communication weights, thereby significantly enhancing the robustness of the learning system. By leveraging decentralized learning, our approach eliminates the dependence on a central server. Striking a harmonious balance between communication round count and sample complexity, BR-DeFedRL achieves efficient convergence with a rate of $\mathcal{O}\left( {\frac{1}{{TN}}} \right)$, where T denotes the communication rounds and N represents the local steps related to variance reduction. Notably, each agent attains an ϵ-approximation with a state-of-the-art sample complexity of $\mathcal{O}\left( {\frac{1}{{\varepsilon N}} + \frac{1}{\varepsilon }} \right)$. Extensive experimental validations further affirm the efficacy of BR-DeFedRL, making it a promising and practical solution for Byzantine-robust decentralized federated reinforcement learning.
Jing Qiao, Zuyuan Zhang, Sheng Yue 0001, Yuan Yuan 0014, Zhipeng Cai 0001, Xiao Zhang 0015, Ju Ren 0001, Dongxiao Yu
INFOCOM6
2024 Federating from History in Streaming Federated Learning
abstract
To address the online learning problem in distributed systems, Streaming Federated learning (SFL) enables immediate model training by clients upon collecting new data, finding wide applications in AI-enabled Internet-of-Things and sensor networks. Given the variability in data distribution across different historical periods, the ability to recall and rapidly apply previously encountered data distributions significantly enhances the efficiency and accuracy of model training. In this paper, a demo based on the real-world temperature datasets is presented to demonstrate the importance of history knowledge in local training and the federating process of SFL, which also shows that vanilla federated learning without considering the history knowledge may even be harmful to model training. Observing this, we propose Fed-HIST, a Federated learning framework that enables the clients to learn from the HISTory knowledge of the whole distributed learning system. Unlike direct raw data storage, Fed-HIST employs model architectures to capture the data distributions, offering a more space-efficient and privacy-preserving method of knowledge storage on a server pool. Additionally, a model similarity comparison scheme is designed to retrieve beneficial knowledge from the pool uploaded by the clients in the past. Such a history-aware federation can enhance the efficiency of training each client, only requiring the recurrence of similar data distributions among SFL participants. We validate our framework through extensive simulations on MNIST, Fashion-MINST, CIFAR10, and CIFAR100 datasets, benchmarking against 9 baselines and highlighting the importance of federating from history in SFL problem through necessary ablation studies.
Ruirui Zhang 0003, Yifei Zou, Zhenzhen Xie 0002, Xiao Zhang 0015, Peng Li 0017, Zhipeng Cai 0001, Xiuzhen Cheng, Dongxiao Yu
MobiHoc4
2024 Resource-Aware Federated Self-Supervised Learning with Global Class Representations
abstract
Due to the heterogeneous architectures and class skew, the global representation models training in resource-adaptive federated self-supervised learning face with tricky challenges: $\textit{deviated representation abilities}$ and $\textit{inconsistent representation spaces}$. In this work, we are the first to propose a multi-teacher knowledge distillation framework, namely $\textit{FedMKD}$, to learn global representations with whole class knowledge from heterogeneous clients even under extreme class skew. Firstly, the adaptive knowledge integration mechanism is designed to learn better representations from all heterogeneous models with deviated representation abilities. Then the weighted combination of the self-supervised loss and the distillation loss can support the global model to encode all classes from clients into a unified space. Besides, the global knowledge anchored alignment module can make the local representation spaces close to the global spaces, which further improves the representation abilities of local ones. Finally, extensive experiments conducted on two datasets demonstrate the effectiveness of $\textit{FedMKD}$ which outperforms state-of-the-art baselines 4.78\% under linear evaluation on average.
Xiao Zhang 0015, Tengfei Liu 0007, Weiqiang Wang 0002, Fuzhen Zhuang, Hui Xiong 0001, Dongxiao Yu
NeurIPS2
2024 Federated Dynamic Graph Fusion Framework for Remaining Useful Life Prediction
Xiao Zhang 0015, Dongxiao Yu
WASA (2)4
2024 A comprehensive survey of federated transfer learning: challenges, methods and applications
abstract
Abstract Federated learning (FL) is a novel distributed machine learning paradigm that enables participants to collaboratively train a centralized model with privacy preservation by eliminating the requirement of data sharing. In practice, FL often involves multiple participants and requires the third party to aggregate global information to guide the update of the target participant. Therefore, many FL methods do not work well due to the training and test data of each participant may not be sampled from the same feature space and the same underlying distribution. Meanwhile, the differences in their local devices (system heterogeneity), the continuous influx of online data (incremental data), and labeled data scarcity may further influence the performance of these methods. To solve this problem, federated transfer learning (FTL), which integrates transfer learning (TL) into FL, has attracted the attention of numerous researchers. However, since FL enables a continuous share of knowledge among participants with each communication round while not allowing local data to be accessed by other participants, FTL faces many unique challenges that are not present in TL. In this survey, we focus on categorizing and reviewing the current progress on federated transfer learning, and outlining corresponding solutions and applications. Furthermore, the common setting of FTL scenarios, available datasets, and significant related research are summarized in this survey.
Fuzhen Zhuang, Xiao Zhang 0015, Yiqi Tong, Jin Dong 0004
Frontiers Comput. Sci.3
2024 Adaptive Clustering Based Personalized Federated Learning Framework for Next POI Recommendation With Location Noise
abstract
Next point-of-interest (POI) recommendation has been a hot research topic, which enables new paradigms for kinds of location-based services in real-world scenarios. Due to the privacy concerns and rigorous data regulations, federated learning provides a distributed learning framework to collaboratively train the recommendation model without sharing the highly sensitive POI data with others. However, there exist two main challenges, namelylocation noise, andbalance between personalization and knowledge sharing, seriously restrict the development of the federated next POI recommendation. To this end, in this work, we propose an adaptive clustering based personalized federated learning framework for next POI recommendation with location noise, namedCPF-POI, to address the above challenges. In detail, within the local client, a location recovery module can efficiently remove noises under the given assumption from the noisy POI data in which the recovery error bound can be theoretically proved. Then, within the parameter server, an adaptive clustering scheme is proposed to capture the internal relatedness among all clients to augment positive knowledge sharing. In order to make a balance between personalization and knowledge sharing under personalized federated learning framework, we design an alternative optimization process between clustering similar clients and minimizing local personalized loss functions. Finally, extensive experiments are conducted on two diverse real-world datasets to show the advantages ofCPF-POIover state-of-the-art methods. improvement across all metrics on average.
Ziming Ye, Xiao Zhang 0015, Xu Chen 0004, Hui Xiong 0001, Dongxiao Yu
IEEE Trans. Knowl. Data Eng.2
2024 Rethinking Robust Multivariate Time Series Anomaly Detection: A Hierarchical Spatio-Temporal Variational Perspective
abstract
The robust multivariate time series anomaly detection can facilitate intelligent decisions and timely maintenance in various kinds of monitor systems. However, the robustness is highly restricted by the stochasticity in multivariate time series, which is summarized astemporal stochasticityandspatial stochasticityspecifically. In this paper, we explicitly model the temporal stochasticity variables and the latent graph relationship variables into a unified graphical framework, which can achieve better robustness to dynamicity from both the spatial and temporal perspective. First, within the spatial encoder, every connection exists or not is modeled as a binary stochastic variable, and the graph structure can be learnt automatically. Then, the temporal encoder would embed the highly structured time series into latent stochastic variables to capture both complex temporal dependencies and neighbors information. Moreover, we design a history-future combined anomaly score mechanism with both reconstruction decoder and forecasting decoder to improve the anomaly detection performance. By weighting the historical anomaly factor, the future anomaly factor, and the prediction error of current timestamp, the anomaly detection at current timestamp could be more sensitive to anomaly detection. Finally, extensive experiments on three publicly available anomaly detection datasets demonstrate our proposed method can achieve the best performance in terms of recall and F1 compared with state-of-the-arts baselines.
Xiao Zhang 0015, Shuqing Xu, Huashan Chen, Zekai Chen 0005, Fuzhen Zhuang, Hui Xiong 0001, Dongxiao Yu
IEEE Trans. Knowl. Data Eng.1
2024 De-RPOTA: Decentralized Learning With Resource Adaptation and Privacy Preservation Through Over-the-Air Computation
abstract
In this paper, we propose De-RPOTA, a novel algorithm designed for decentralized learning, equipped with mechanisms for resource adaptation and privacy protection through over-the-air computation. We theoretically analyze the combined effects of limited resources and lossy communication on decentralized learning, showing it converges towards a contraction region defined by a scaled errors version. Remarkably, De-RPOTA achieves a convergence rate of$\mathcal {O}\left ({{\frac {1}{\sqrt {nT}}}}\right)$in scenarios devoid of errors, matching the state-of-the-arts. Additionally, we tackle a power control challenge, breaking it down into transmitter and receiver sub-problems to hasten the De-RPOTA algorithm’s convergence. We also offer a quantifiable privacy assurance for our over-the-air computation methodology. Intriguingly, our findings suggest that network noise can actually strengthen the privacy of aggregated information, with over-the-air computation providing extra security for individual updates. Comprehensive experimental validation confirms De-RPOTA’s efficacy in communication resources limited environments. Specifically, the results on the CIFAR-10 dataset reveal nearly 30% reduction in communication costs compared to the state-of-the-arts, all while maintaining similar levels of learning accuracy, even under resource restrictions.
Jing Qiao, Shikun Shen, Shuzhen Chen 0001, Xiao Zhang 0015, Tian Lan 0001, Xiuzhen Cheng, Dongxiao Yu
IEEE/ACM Trans. Netw.4
2023 Data Quality Aware Hierarchical Federated Reinforcement Learning Framework for Dynamic Treatment Regimes
abstract
Due to the privacy concerns and rigorous data regulations, dynamic treatment regimes across hospitals have become increasingly difficult. Fortunately, federated learning provides a distributed learning framework to collaboratively train the model without sharing the highly sensitive electronic health record (EHR) data with others. However, there exist two main challenges, namely data quality discrepancy, and heterogeneous data distribution, which seriously restrict the development of federated dynamic treatment regimes. To this end, we develop a global data quality aware dynamic treatment regime based on hierarchical federated reinforcement learning across different hospitals. In detail, we first quantify data quality in EHR using immediate health status changes, which are then utilized as rewards to encourage the high-quality treatment actions in the offline actor-critic reinforcement learning model. Within the parameter server, an online reinforcement learning based clustering scheme is proposed to capture the internal similarities to augment the positive knowledge transfer of high-quality hospitals while neglecting the heterogeneity. Extensive experiments are conducted on two diverse real-world datasets to show the advantages of DFR-DTR over state-of-the-art baselines.
Xiao Zhang 0015, Haochao Ying, Xu Han 0025, Dongxiao Yu
ICDM2
2023 Robust Image Ordinal Regression with Controllable Image Generation
abstract
Image ordinal regression has been mainly studied along the line of exploiting the order of categories. However, the issues of class imbalance and category overlap that are very common in ordinal regression were largely overlooked. As a result, the performance on minority categories is often unsatisfactory. In this paper, we propose a novel framework called CIG based on controllable image generation to directly tackle these two issues. Our main idea is to generate extra training samples with specific labels near category boundaries, and the sample generation is biased toward the less-represented categories. To achieve controllable image generation, we seek to separate structural and categorical information of images based on structural similarity, categorical similarity, and reconstruction constraints. We evaluate the effectiveness of our new CIG approach in three different image ordinal regression scenarios. The results demonstrate that CIG can be flexibly integrated with off-the-shelf image encoders or ordinal regression models to achieve improvement, and further, the improvement is more significant for minority categories.
Haochao Ying, Renjun Hu, Xiao Zhang 0015, Danny Ziyi Chen, Jian Wu 0001
IJCAI6
2023 Theoretical Convergence Guaranteed Resource-Adaptive Federated Learning with Mixed Heterogeneity
abstract
In this paper, we propose an adaptive learning paradigm for resource-constrained cross-device federated learning, in which heterogeneous local submodels with varying resources can be jointly trained to produce a global model. Different from existing studies, the submodel structures of different clients are formed by arbitrarily assigned neurons according to their local resources. Along this line, we first design a general resource-adaptive federated learning algorithm, namely RA-Fed, and rigorously prove its convergence with asymptotically optimal rate O(1/√Γ*TQ) under loose assumptions. Furthermore, to address both submodels heterogeneity and data heterogeneity challenges under non-uniform training, we come up with a new server aggregation mechanism RAM-Fed with the same theoretically proved convergence rate. Moreover, we shed light on several key factors impacting convergence, such as minimum coverage rate, data heterogeneity level, submodel induced noises. Finally, we conduct extensive experiments on two types of tasks with three widely used datasets under different experimental settings. Compared with the state-of-the-arts, our methods improve the accuracy up to 10% on average. Particularly, when submodels jointly train with 50% parameters, RAM-Fed achieves comparable accuracy to FedAvg trained with the full model.
Xiao Zhang 0015, Tian Lan 0001, Huashan Chen, Hui Xiong 0001, Xiuzhen Cheng, Dongxiao Yu
KDD2
2023 Communication Resources Limited Decentralized Learning with Privacy Guarantee through Over-the-Air Computation
abstract
In this paper, we propose a novel decentralized learning algorithm, namely DLLR-OA, for resource-constrained over-the-air computation with formal privacy guarantee. Theoretically, we characterize how the limited resources induced model-components selection error and compound communication errors jointly impact decentralized learning, making the iterates of DLLR-OA converge to a contraction region centered around a scaled version of the errors. In particular, the convergence rate of the DLLR-OA algorithm in the error-free case [EQUATION] achieves the state-of-the-arts. Besides, we formulate a power control problem and decouple it into two sub-problems of transmitter and receiver to accelerate the convergence of the DLLR-OA algorithm. Furthermore, we provide quantitative privacy guarantee for the proposed over-the-air computation approach. Interestingly, we show that network noise can indeed enhance privacy of aggregated updates while over-the-air computation can further protect individual updates. Finally, the extensive experiments demonstrate that DLLR-OA performs well in the communication resources constrained setting. In particular, numerical results on CIFAR-10 dataset shows nearly 30% communication cost reduction over state-of-the-art baselines with comparable learning accuracy even in resource constrained settings.
Jing Qiao, Shikun Shen, Shuzhen Chen 0001, Xiao Zhang 0015, Tian Lan 0001, Xiuzhen Cheng, Dongxiao Yu
MobiHoc4
2023 Fine-Grained Preference-Aware Personalized Federated POI Recommendation with Data Sparsity
abstract
With the raised privacy concerns and rigorous data regulations, federated learning has become a hot collaborative learning paradigm for the recommendation model without sharing the highly sensitive POI data. However, the time-sensitive, heterogeneous, and limited POI records seriously restrict the development of federated POI recommendation. To this end, in this paper, we design the fine-grained preference-aware personalized federated POI recommendation framework, namely PrefFedPOI, under extremely sparse historical trajectories to address the above challenges. In details, PrefFedPOI extracts the fine-grained preference of current time slot by combining historical recent preferences and periodic preferences within each local client. Due to the extreme lack of POI data in some time slots, a data amount aware selective strategy is designed for model parameters uploading. Moreover, a performance enhanced clustering mechanism with reinforcement learning is proposed to capture the preference relatedness among all clients to encourage the positive knowledge sharing. Furthermore, a clustering teacher network is designed for improving efficiency by clustering guidance. Extensive experiments are conducted on two diverse real-world datasets to demonstrate the effectiveness of proposed PrefFedPOI comparing with state-of-the-arts. In particular, personalized PrefFedPOI can achieve 7% accuracy improvement on average among data-sparsity clients.
Xiao Zhang 0015, Ziming Ye, Jianfeng Lu 0002, Fuzhen Zhuang, Yanwei Zheng, Dongxiao Yu
SIGIR1
2023 Robust decentralized stochastic gradient descent over unstable networks
Yanwei Zheng, Liangxu Zhang, Shuzhen Chen 0001, Xiao Zhang 0015, Zhipeng Cai 0001, Xiuzhen Cheng
Comput. Commun.4
2023 Trustworthy decentralized collaborative learning for edge intelligence: A survey
abstract
Edge intelligence is an emerging technology that enables artificial intelligence on connected systems and devices in close proximity to the data sources. Decentralized Collaborative Learning (DCL) is a novel edge intelligence technique that allows distributed clients to cooperatively train a global learning model without revealing their data. DCL has a wide range of applications in various domains, such as smart city and autonomous driving. However, DCL faces significant challenges in ensuring its trustworthiness, as data isolation and privacy issues make DCL systems vulnerable to adversarial attacks that aim to breach system confidentiality, undermine learning reliability or violate data privacy. Therefore, it is crucial to design DCL in a trustworthy manner, with a focus on security, robustness, and privacy. In this survey, we present a comprehensive review of existing efforts for designing trustworthy DCL systems from the three key aformentioned aspects: security, robustness, and privacy. We analyze the threats that affect the trustworthiness of DCL across different scenarios and assess specific technical solutions for achieving each aspect of Trustworthy DCL (TDCL). Finally, we highlight open challenges and future directions for advancing TDCL research and practice.
Dongxiao Yu, Zhenzhen Xie 0002, Yuan Yuan 0014, Shuzhen Chen 0001, Jing Qiao, Yong Yu 0002, Yifei Zou, Xiao Zhang 0015
High Confid. Comput.9
2023 Robust communication-efficient decentralized learning with heterogeneity
Xiao Zhang 0015, Shuzhen Chen 0001, Dongxiao Yu, Xiuzhen Cheng
J. Syst. Archit.1
2023 Federated Representation Learning With Data Heterogeneity for Human Mobility Prediction
abstract
The advancement of smart wearable devices and location-based smart services has enabled a new paradigm for smart human mobility prediction (HMP), which has a broad range of applications in smart healthcare and smart cities. Due to the privacy concerns and rigorous data regulations, federated learning provides a distributed learning framework to collaboratively train the HMP model without sharing the highly sensitive location data with others. However, in real-world scenarios, federated human mobility prediction suffers from data heterogeneity challenge, which includes two main aspects: heterogeneity mobility patterns, and data scarcity. In this paper, we propose an end-to-end federated representation learning framework for human mobility prediction, named FR-HMP, to overcome all the above obstacles. Specially, in order to enhance the representation abilities of data-scarcity clients, a two-phase learning process is proposed. The clustering module could cluster similar clients together on the parameter server to address the heterogeneous mobility patterns, and the representation learning module learns the enhanced representations of each client through the graph learning layer and graph convolution layer on the third-part server. Finally, extensive experiments are conducted using two diverse real-world HMP datasets to show the advantages of FR-HMP over state-of-the-art methods.
Xiao Zhang 0015, Ziming Ye, Haochao Ying, Dongxiao Yu
IEEE Trans. Intell. Transp. Syst.1
2023 Time-Aware Context-Gated Graph Attention Network for Clinical Risk Prediction
abstract
Clinical risk prediction based on Electronic Health Records (EHR) can assist doctors in better judgment and can make sense of early diagnosis. However, the prediction performance heavily relies on effective representations from multi-dimensional time-series EHR data. Existing solutions usually focus on temporal features or inherent relations between clinical event variables or extract both information in two separate phases. This usually leads to insufficient patient feature information and results in poor prediction performance. Moreover, existing methods based on Heterogeneous Graph Neural Network usually require manual selection of proper Meta-Paths. To solve these problems, we propose the Time-aware Context-Gated Graph Attention Network (T-ContextGGAN). Specifically, we design a GNN based module with Time-aware Meta-Paths and self-attention mechanism to extract both temporal semantic information and inherent relations of EHR data simultaneously and perform automatic Meta-Path selection. To evaluate the proposed model, we extract the first 48 hour EHR data in the first Intensive Care Unit (ICU) admission of three different tasks from two open-source datasets and model various clinical variables on the proposed EHRGraph. Extensive experimental results show the proposed model can effectively extract informative features, and outperform existing state-of-art models in terms of various prediction measures. Our code is available in https://github.com/OwlCitizen/TContext-GGAN.
Yuyang Xu, Haochao Ying, Siyi Qian, Fuzhen Zhuang, Xiao Zhang 0015, Deqing Wang 0001, Jian Wu 0001, Hui Xiong 0001
IEEE Trans. Knowl. Data Eng.5
2023 FedHAR: Semi-Supervised Online Learning for Personalized Federated Human Activity Recognition
abstract
The advancement of smartphone sensors and wearable devices has enabled a new paradigm for smart human activity recognition (HAR), which has a broad range of applications in healthcare and smart cities. However, there are four challenges,privacy preservation,label scarcity,real-timing, andheterogeneity patterns, to be addressed before HAR can be more applicable in real-world scenarios. To this end, in this paper, we propose a personalized federated HAR framework, namedFedHAR, to overcome all the above obstacles. Specially, as federated learning,FedHARperforms distributed learning, which allows training data to be kept local to protect users’ privacy. Also, for each client without activity labels, inFedHAR, we design an algorithm to compute unsupervised gradients under theconsistency trainingproposition and an unsupervised gradient aggregation strategy is developed for overcoming the concept drift and convergence instability issues in online federated learning process. Finally, extensive experiments are conducted using two diverse real-world HAR datasets to show the advantages ofFedHARover state-of-the-art methods. In addition, when fine-tuning each unlabeled client, personalizedFedHARcan achieve additional 10% improvement across all metrics on average.
Hongzheng Yu, Zekai Chen 0005, Xiao Zhang 0015, Xu Chen 0004, Fuzhen Zhuang, Hui Xiong 0001, Xiuzhen Cheng
IEEE Trans. Mob. Comput.3
2022 ASM2TV: An Adaptive Semi-supervised Multi-Task Multi-View Learning Framework for Human Activity Recognition
abstract
Many real-world scenarios, such as human activity recognition (HAR) in IoT, can be formalized as a multi-task multi-view learning problem. Each specific task consists of multiple shared feature views collected from multiple sources, either homogeneous or heterogeneous. Common among recent approaches is to employ a typical hard/soft sharing strategy at the initial phase separately for each view across tasks to uncover common knowledge, underlying the assumption that all views are conditionally independent. On the one hand, multiple views across tasks possibly relate to each other under practical situations. On the other hand, supervised methods might be insufficient when labeled data is scarce. To tackle these challenges, we introduce a novel framework ASM2TV for semi-supervised multi-task multi-view learning. We present a new perspective named gating control policy, a learnable task-view-interacted sharing policy that adaptively selects the most desirable candidate shared block for any view across any task, which uncovers more fine-grained task-view-interacted relatedness and improves inference efficiency. Significantly, our proposed gathering consistency adaption procedure takes full advantage of large amounts of unlabeled fragmented time-series, making it a general framework that accommodates a wide range of applications. Experiments on two diverse real-world HAR benchmark datasets collected from various subjects and sources demonstrate our framework's superiority over other state-of-the-arts. Anonymous codes are available at https://github.com/zachstarkk/ASM2TV.
Zekai Chen 0005, Xiao Zhang 0015, Xiuzhen Cheng
AAAI2
2022 EdgeViT: Efficient Visual Modeling for Edge Computing
Zekai Chen 0005, Fangtian Zhong, Xiao Zhang 0015, Yanwei Zheng
WASA (3)4
2022 Forecasting fine-grained city-scale cellular traffic with sparse crowdsourced measurements
Jian-Hui Duan, Xiao Zhang 0015, Sanglu Lu
Comput. Networks3
2022 Learning Graph Structures With Transformer for Multivariate Time-Series Anomaly Detection in IoT
abstract
Many real-world IoT systems, which include a variety of internet-connected sensory devices, produce substantial amounts of multivariate time series data. Meanwhile, vital IoT infrastructures like smart power grids and water distribution networks are frequently targeted by cyber-attacks, making anomaly detection an important study topic. Modeling such relatedness is, nevertheless, unavoidable for any efficient and effective anomaly detection system, given the intricate topological and nonlinear connections that are originally unknown among sensors. Furthermore, detecting anomalies in multivariate time series is difficult due to their temporal dependency and stochasticity. This paper presented GTA, a new framework for multivariate time series anomaly detection that involves automatically learning a graph structure, graph convolution, and modeling temporal dependency using a Transformer-based architecture. The connection learning policy, which is based on the Gumbel-softmax sampling approach to learn bi-directed links among sensors directly, is at the heart of learning graph structure. To describe the anomaly information flow between network nodes, we introduced a new graph convolution called Influence Propagation convolution. In addition, to tackle the quadratic complexity barrier, we suggested a multi-branch attention mechanism to replace the original multi-head self-attention method. Extensive experiments on four publicly available anomaly detection benchmarks further demonstrate the superiority of our approach over alternative state-of-the-arts. Codes are available at https://github.com/ZEKAICHEN/GTA.
Zekai Chen 0005, Dingshuo Chen, Xiao Zhang 0015, Zixuan Yuan, Xiuzhen Cheng
IEEE Internet Things J.3
2022 HarMI: Human Activity Recognition Via Multi-Modality Incremental Learning
abstract
Nowadays, with the development of various kinds of sensors in smartphones or wearable devices, human activity recognition (HAR) has been widely researched and has numerous applications in healthcare, smart city, etc. Many techniques based on hand-crafted feature engineering or deep neural network have been proposed for sensor based HAR. However, these existing methods usually recognize activities offline, which means the whole data should be collected before training, occupying large-capacity storage space. Moreover, once the offline model training finished, the trained model can't recognize new activities unless retraining from the start, thus with a high cost of time and space. In this paper, we propose a multi-modality incremental learning model, called HarMI, with continuous learning ability. The proposed HarMI model can start training quickly with little storage space and easily learn new activities without storing previous training data. In detail, we first adopt attention mechanism to align heterogeneous sensor data with different frequencies. In addition, to overcome catastrophic forgetting in incremental learning, HarMI utilizes the elastic weight consolidation and canonical correlation analysis from a multi-modality perspective. Extensive experiments based on two public datasets demonstrate that HarMI can achieve a superior performance compared with several state-of-the-arts.
Xiao Zhang 0015, Hongzheng Yu, Yang Yang 0129, Jingjing Gu, Fuzhen Zhuang, Dongxiao Yu, Zhaochun Ren
IEEE J. Biomed. Health Informatics1
2022 Question Tagging via Graph-guided Ranking
abstract
With the increasing prevalence of portable devices and the popularity of community Question Answering (cQA) sites, users can seamlessly post and answer many questions. To effectively organize the information for precise recommendation and easy searching, these platforms require users to select topics for their raised questions. However, due to the limited experience, certain users fail to select appropriate topics for their questions. Thereby, automatic question tagging becomes an urgent and vital problem for the cQA sites, yet it is non-trivial due to the following challenges. On the one hand, vast and meaningful topics are available yet not utilized in the cQA sites; how to model and tag them to relevant questions is a highly challenging problem. On the other hand, related topics in the cQA sites may be organized into a directed acyclic graph. In light of this, how to exploit relations among topics to enhance their representations is critical. To settle these challenges, we devise a graph-guided topic ranking model to tag questions in the cQA sites appropriately. In particular, we first design a topic information fusion module to learn the topic representation by jointly considering the name and description of the topic. Afterwards, regarding the special structure of topics, we propose an information propagation module to enhance the topic representation. As the comprehension of questions plays a vital role in question tagging, we design a multi-level context-modeling-based question encoder to obtain the enhanced question representation. Moreover, we introduce an interaction module to extract topic-aware question information and capture the interactive information between questions and topics. Finally, we utilize the interactive information to estimate the ranking scores for topics. Extensive experiments on three Chinese cQA datasets have demonstrated that our proposed model outperforms several state-of-the-art competitors.
Xiao Zhang 0015, Meng Liu 0006, Jianhua Yin 0001, Zhaochun Ren, Liqiang Nie
ACM Trans. Inf. Syst.1
2021 Modeling Heterogeneous Relations across Multiple Modes for Potential Crowd Flow Prediction
abstract
Potential crowd flow prediction for new planned transportation sites is a fundamental task for urban planners and administrators. Intuitively, the potential crowd flow of the new coming site can be implied by exploring the nearby sites. However, the transportation modes of nearby sites (e.g. bus stations, bicycle stations) might be different from the target site (e.g. subway station), which results in severe data scarcity issues. To this end, we propose a data-driven approach, named MOHER, to predict the potential crowd flow in a certain mode for a new planned site. Specifically, we first identify the neighbor regions of the target site by examining the geographical proximity as well as the urban function similarity. Then, to aggregate these heterogeneous relations, we devise a cross-mode relational GCN, a novel relation-specific transformation model, which can learn not only the correlation but also the differences between different transportation modes. Afterward, we design an aggregator for inductive potential flow representation. Finally, an LTSM module is used for sequential flow prediction. Extensive experiments on real-world data sets demonstrate the superiority of the MOHER framework compared with the state-of-the-art algorithms.
Qiang Zhou 0007, Jingjing Gu, Xinjiang Lu, Fuzhen Zhuang, Yanchao Zhao, Xiao Zhang 0015
AAAI7
2021 DCAP: Deep Cross Attentional Product Network for User Response Prediction
abstract
User response prediction, which aims to predict the probability that a user will provide a predefined positive response in a given context such as clicking on an ad or purchasing an item, is crucial to many industrial applications such as online advertising, recommender systems, and search ranking. For these tasks and many other machine learning tasks, an indispensable part of success is feature engineering, where cross features are a significant type of feature transformations. However, due to the high dimensionality and super sparsity of the data collected in these tasks, handcrafting cross features is inevitably time expensive. Prior studies in predicting user response leveraged the feature interactions by enhancing feature vectors with products of features to model second-order or high-order cross features, either explicitly or implicitly. However, these existing methods can be hindered by not learning sufficient cross features due to model architecture limitations or modeling all high-order feature interactions with equal weights. Different features should contribute differently to the prediction, and not all cross features are with the same prediction power.
Zekai Chen 0005, Fangtian Zhong, Zhumin Chen, Xiao Zhang 0015, Robert Pless, Xiuzhen Cheng
CIKM4
2021 Cascaded SE-ResUnet for segmentation of thoracic organs at risk
Zheng Cao 0005, Bohan Yu, Biwen Lei, Haochao Ying, Xiao Zhang 0015, Danny Ziyi Chen, Jian Wu 0001
Neurocomputing5
2020 PersonalitySensing: A Multi-View Multi-Task Learning Approach for Personality Detection based on Smartphone Usage
abstract
Assessing individual's personality traits has important implications in psychology, sociology, and economics. Conventional personality measurement methods were questionnaire-based, which are time-consuming and manpower-expensive. With the pervasive deployment of mobile communication applications, smartphone usage data was found to relate to people's social behavioral and psychological aspects. In this paper, we propose a deep learning approach to infer people's Big Five personality traits based on smartphone data. Specifically, we collect smartphone usage snapshots with an Android App, and extract features from the collected data. We propose a multi-view multi-task learning approach with a deep neural network model to fuse the extracted features and learn the Big Five personality traits jointly. Extensive experiments based on the real-world smartphone data collected from university volunteers show that the proposed approach significantly outperforms the state-of-the-art algorithms in personality prediction.
Songcheng Gao, Lynda Jiwen Song, Xiao Zhang 0015, Mingkai Lin, Sanglu Lu
ACM Multimedia4
2020 Predicting and Recommending the next Smartphone Apps based on Recurrent Neural Network
Shijian Xu, Xiao Zhang 0015, Songcheng Gao, Tong Zhan, Sanglu Lu
CCF Trans. Pervasive Comput. Interact.3
2020 Emotion Detection in Online Social Networks: A Multilabel Learning Approach
abstract
Emotion detection in online social networks (OSNs) can benefit kinds of applications, such as personalized advertisement services, recommendation systems, etc. Conventionally, emotion analysis mainly focuses on the sentence level polarity prediction or single emotion label classification, however, ignoring the fact that emotions might coexist from users' perspective. To this end, in this work, we address the multiple emotions detection in OSNs from user-level view, and formulate this problem as a multilabel learning problem. First, we discover emotion labels correlations, social correlations, and temporal correlations from an annotated Twitter data set. Second, based on the above observations, we adopt a factor graph-based emotion recognition model to incorporate emotion labels correlations, social correlations, and temporal correlations into a general framework, and detect the multiple emotions based on the multilabel learning approach. Performance evaluation demonstrates that the factor graph-based emotion detection model can outperform the existing baselines.
Xiao Zhang 0015, Haochao Ying, Feng Li 0002, Siyi Tang, Sanglu Lu
IEEE Internet Things J.1
2019 AttnSense: Multi-level Attention Mechanism For Multimodal Human Activity Recognition
abstract
Sensor-based human activity recognition is a fundamental research problem in ubiquitous computing, which uses the rich sensing data from multimodal embedded sensors such as accelerometer and gyroscope to infer human activities. The existing activity recognition approaches either rely on domain knowledge or fail to address the spatial-temporal dependencies of the sensing signals. In this paper, we propose a novel attention-based multimodal neural network model called AttnSense for multimodal human activity recognition. AttnSense introduce the framework of combining attention mechanism with a convolutional neural network (CNN) and a Gated Recurrent Units (GRU) network to capture the dependencies of sensing signals in both spatial and temporal domains, which shows advantages in prioritized sensor selection and improves the comprehensibility. Extensive experiments based on three public datasets show that AttnSense achieves a competitive performance in activity recognition compared with several state-of-the-art methods.
Haojie Ma, Xiao Zhang 0015, Songcheng Gao, Sanglu Lu
IJCAI3
2019 Inferring Mood Instability via Smartphone Sensing: A Multi-View Learning Approach
abstract
A high correlation between mood instability (MI), the rapid and constant fluctuation in mood, and mental health has been demonstrated. However, conventional approaches to measure MI are limited owing to the high manpower and time cost required. In this paper, we propose a smartphone-based MI detection that can automatically and passively detect MI with minimal human involvement. The proposed method trains a multi-view learning classification model using features extracted from the smartphone sensing data of volunteers and their self-reported moods. The trained classifier is then used to detect the MI of unseen users efficiently, thereby reducing the human involvement and time cost significantly. Based on extensive experiments conducted with the dataset collected from 68 volunteers, we demonstrate that the proposed multi-view learning model outperforms the baseline classifiers.
Xiao Zhang 0015, Fuzhen Zhuang, Haochao Ying, Hui Xiong 0001, Sanglu Lu
ACM Multimedia1
2019 Supervised representation learning for multi-label classification
Fuzhen Zhuang, Xiao Zhang 0015, Xiang Ao 0001, Zhengyu Niu, Min-Ling Zhang, Qing He 0003
Mach. Learn.3
2019 Time-aware metric embedding with asymmetric projection for successive POI recommendation
Haochao Ying, Jian Wu 0001, Guandong Xu, Yanchi Liu, Tingting Liang, Xiao Zhang 0015, Hui Xiong 0001
World Wide Web6
2018 An Integrated Model for Crime Prediction Using Temporal and Spatial Factors
abstract
Given its importance, crime prediction has attracted a lot of attention in the literature, and several methods have been proposed to discover different aspects of characteristics for crime prediction. In this paper, we propose a Clustered Continuous Conditional Random Field (Clustered-CCRF) model which is able to effectively exploit both spatial and temporal factors for crime prediction in an integrated way. In particular, we observe that the crime number at one specific area is not only conditioned on its own historical records but also has high correlation to crime records from similar areas. Therefore, we propose two factors: an auto-regressed temporal correlation and a feature-based inter-area spatial correlation, to measure such patterns for crime prediction. Further, we present a tree-structured clustering algorithm to discover high similar areas based on spatial characteristics to improve the performance of our proposed model. Experiments on real-world crime dataset demonstrate the superiority of our proposed model over the state-of-the-art methods.
Fei Yi, Zhiwen Yu 0001, Fuzhen Zhuang, Xiao Zhang 0015, Hui Xiong 0001
ICDM4
2018 Label-Sensitive Task Grouping by Bayesian Nonparametric Approach for Multi-Task Multi-Label Learning
abstract
Multi-label learning is widely applied in many real-world applications, such as image and gene annotation. While most of the existing multi-label learning models focus on the single-task learning problem, there are always some tasks that share some commonalities, which can help each other to improve the learning performances if the knowledge in the similar tasks can be smartly shared. In this paper, we propose a LABel-sensitive TAsk Grouping framework, named LABTAG, based on Bayesian nonparametric approach for multi-task multi-label classification. The proposed framework explores the label correlations to capture feature-label patterns, and clusters similar tasks into groups with shared knowledge, which are learned jointly to produce a strengthened multi-task multi-label model. We evaluate the model performance on three public multi-task multi-label data sets, and the results show that LABTAG outperforms the compared baselines with a significant margin.
Xiao Zhang 0015, Vu Nguyen 0001, Fuzhen Zhuang, Hui Xiong 0001, Sanglu Lu
IJCAI1
2018 Predicting Smartphone App Usage with Recurrent Neural Networks
Shijian Xu, Xiao Zhang 0015, Songcheng Gao, Tong Zhan, Yongzhu Zhao, Wei-wei Zhu, Tianzi Sun
WASA3
2017 Emotion Detection in Online Social Network Based on Multi-label Learning
Xiao Zhang 0015, Sanglu Lu
DASFAA (1)1
2017 Ambula: Build Communication Lifeline of Corporations During Emergency
abstract
Many corporations rely on Internet service provider (ISP) network to provide reliable communication services. However, the current communication networks are vulnerable to disruptive events, such as natural disaster or power outage. Such disastrous events may destroy multiple network facilities in a specific region and result in a long term recovery of ISP networks. The disconnected communication will lead to enormous economic loss even if corporation's infrastructure is not directly destroyed during the disaster. Therefore, corporations need a self-rescue mechanism to actively respond to the emergency instead of simply relying on the ISP. This paper proposes Ambula, an easy-to-deploy platform to realize fast congestion-aware recovery for corporation's communication lifeline. Our platform leverages current widely-deployed public cloud services to build a scalable peer-to-peer overlay routing system. By so doing, the corporation is capable of controlling the packets forwarding path to bypass the affected region and congested routes. To this end, Ambula first carefully selects a small set of virtual machines (VMs) from geographically distributed public clouds, and then apply the self-developed congestion-aware routing protocol to achieve automatic and fast routing recovery. Simulations on both random generated and real network topologies show that the high recovery ratio of 80% can be achieved. The congestion avoidance algorithm can significantly reduce the impact of congestion. Our prototype on Emulab shows it can recover within hundreds of milliseconds. To the best of our knowledge, no effective disaster recovery mechanism currently exists for corporations during emergency. Ambula will facilitate the business continuity management of corporations in present of hazard events.
An Xie, Xiao Zhang 0015, Xiaoliang Wang 0001, Zhuzhong Qian, Sanglu Lu
ICPADS2
2017 Predicting Happiness State Based on Emotion Representative Mining in Online Social Networks
Xiao Zhang 0015, Hong Huang 0001, Cam-Tu Nguyen, Xu Chen 0004, Xiaoliang Wang 0001, Sanglu Lu
PAKDD (1)1
2016 Academic Paper Recommendation Based on Community Detection in Citation-Collaboration Networks
Xiao Zhang 0015, Sanglu Lu
APWeb (2)3