Tao Chang

dblp:02/6285 · DBLP profile ↗
← Back
23ranked-venue papers
3as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 12 since 2021Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Do as You See, Not Just as Told: Multimodal Fusion for Proactive Decision-Making in Dynamic Environments
Tao Chang, Xinxin Dong, Yongxue Shan
ICIC (15)3
2025 FedVLA: Federated Vision-Language-Action Learning with Dual Gating Mixture-of-Experts for Robotic Manipulation
abstract
Vision-language-action (VLA) models have significantly advanced robotic manipulation by enabling robots to interpret language instructions for task execution. However, training these models often relies on large-scale user-specific data, raising concerns about privacy and security, which in turn limits their broader adoption. To address this, we propose FedVLA, the first federated VLA learning framework, enabling distributed model training that preserves data privacy without compromising performance. Our framework integrates task-aware representation learning, adaptive expert selection, and expert-driven federated aggregation, enabling efficient and privacy-preserving training of VLA models. Specifically, we introduce an Instruction Oriented Scene-Parsing mechanism, which decomposes and enhances object-level features based on task instructions, improving contextual understanding. To effectively learn diverse task patterns, we design a Dual Gating Mixture-of-Experts (DGMoE) mechanism, where not only input tokens but also self-aware experts adaptively decide their activation. Finally, we propose an Expert-Driven Aggregation strategy at the federated server, where model aggregation is guided by activated experts, ensuring effective cross-client knowledge transfer.Extensive simulations and real-world robotic experiments demonstrate the effectiveness of our proposals. Notably, DGMoE significantly improves computational efficiency compared to its vanilla counterpart, while FedVLA achieves task success rates comparable to centralized training, effectively preserving data privacy.
Cui Miao, Tao Chang, Meihan Wu, Ming Li 0073, Xiaodong Wang 0002
ICCV2
2025 EFTViT: Efficient Federated Training of Vision Transformers with Masked Images on Resource-Constrained Clients
Meihan Wu, Tao Chang, Cui Miao, Jie Zhou 0001, Xiangyu Xu 0002, Ming Li 0073, Xiaodong Wang 0002
ICCV2
2025 $U^2$ Frame: A Unified and Unsupervised Learning Framework for LiDAR-Based Loop Closing
abstract
Loop closing is critically important in Simultaneous Localization and Mapping (SLAM) due to its ability to correct accumulated localization errors. However, existing methods are hindered by the difficulty of acquiring pose labels and the unreliability of ground truth data. In this paper, we propose$U^{2}$Frame, a unified LiDAR-based loop closing framework that handles both loop closure detection and relative pose estimation without any ground truth training data. Specifically, the natural temporal-spatial correlation in point cloud sequences is first leveraged to supervise the network training, where near scans are treated as positives and vice versa as negatives. A new neural architecture is then constructed to jointly learn highly discriminative local and global features for loop closure detection. Additionally, an effective candidate verification module that exploits high-order geometric information is presented to further filter out false loop closures and estimate precise poses. We extensively evaluate$U^{2}$Frame on multiple datasets according to two tasks derived from loop closing: loop closure detection and loop pose estimation. Comparative experiments demonstrate that our method outperforms existing state-of-the-art supervised techniques and has a strong generalization ability across unseen scenarios. Our code is released at https://github.com/yxin-zhang/U2Frame.
Sheng Ao, Ye Zhang 0037, Qingyong Hu, Tao Chang, Yulan Guo
ICRA6
2025 3D Whole-Body Pose Estimation Using Graph High-Resolution Network for Humanoid Robot Teleoperation
abstract
In the realm of robotics, teleoperation plays a pivotal role in performing high-risk or intricate tasks, and obtaining precise 3D whole-body pose is crucial for this purpose. Traditional two-stage methods have limitations in estimating different body parts, leading to complex systems and higher estimation errors. In order to address these issues,the paper introduces a novel framework called Graph High-Resolution Network (GraphHRNet) for accurate 3D whole-body pose estimation, which is essential for the teleoperation of humanoid robots. GraphHRNet effectively captures global structural information and local details by integrating a High-Resolution Module and a Multi-branch Regression Module. The High-Resolution Module utilizes an enhanced graph convolution kernel to fuse multi-scale features, capturing global information, while the Multi-branch Regression Module focuses on refining and predicting accurate 3D coordinates for intricate body parts such as hands and face. Experimental results on the H3WB dataset demonstrate that GraphHRNet surpasses state-of-the-art (SOTA) methods in 3D whole-body pose estimation, significantly improving performance. Furthermore, the paper explores the potential application of this approach in a tele-operation system for humanoid robots, providing an intuitive and high-fidelity solution for remotely executing complex tasks. The code have been publicly available at https://github.com/Z-mingyu/GraphHRNet.git
Qing Gao 0002, Yuanchuan Lai, Ye Zhang 0037, Tao Chang, Yulan Guo
ICRA5
2025 FedETE: Privacy-Preserving Federated Event Trigger Extraction
Fei Hu 0005, Tao Chang, Meihan Wu, Shenpo Dong, Jie Zhou 0032, Jiaqian Yin, Xiaodong Wang 0002
NLPCC (4)2
2024 FedEKT: Ensemble Knowledge Transfer for Model-Heterogeneous Federated Learning
abstract
Federated Learning (FL) enables multiple clients to collaboratively train a shared server model while preserving data privacy. Most existing FL systems rely on the assumption that the server model and client models have homogeneous architecture. However, intensive resource requirements during the training process prevent low-end devices from contributing to the server model with their own data. On the other hand, the resource constraints on participating clients can significantly limit the size of the server model in the model-homogeneous setting, thereby restricting the application scope of FL. In this work, we propose FedEKT, a novel model-heterogeneous FL system designed to obtain a high-performance large server model while benefiting heterogeneous small client models. Specifically, a new aggregation approach is designed to enable the integration of knowledge from heterogeneous client models to a large server model while mitigating the adverse effects of biases stemming from data heterogeneity. Subsequently, to enhance the performance of client models by benefiting from the high-performance server model, FedEKT distills this large server model into multiple heterogeneous client models, facilitating the transfer of integrated knowledge back to the client models. In addition, we design specialized modules within the model and communication strategy to accomplish aggregation and transfer of knowledge in a data-free manner. The evaluation results demonstrate that FedEKT enhances the accuracy of the server model and client models by up to 53.96% and 12.35%, respectively, compared with the state-of-the-art FL approach on CIFAR-100.
Meihan Wu, Li Li 0064, Tao Chang, Peng Qiao, Cui Miao, Jie Zhou 0032, Jingnan Wang, Xiaodong Wang 0002
IWQoS3
2024 PFed-DBA: Distribution Bias Aware Personalized Federated Learning for Data Heterogeneity
abstract
Personalized Federated Learning (PFL) aims to learn a custom model for each distributed client while benefiting from collaborative training in order to overcome the detrimental impact of data heterogeneity. Despite the promising benefits, the existing approaches often compromise the generalization performance of personalized models, as they solely focus on enhancing the personalization capability of models or merely aim to strike a balance between personalization and generalization. Indeed, increasing the personalization capability while preserving the strong generalization performance enabled by collaborative training remains a challenge for PFL, as the two objectives seem to compete with each other. To tackle this challenge, we investigate the relationship between model generalization and personalization under different degrees of heterogeneity. We find that besides the client-specific data distribution, the distribution bias between the unique data distribution of each client and that of the whole population is another critical factor that prominently impacts these two performances. Motivated by the above finding, we propose PFed-DBA, a novel PFL framework that effectively perceives this distribution bias to guide the training process. Concretely, we design the PFL models as a skip-connection network between a shared module for learning the shared representations delivering the common distribution of data across all clients and a personalized module for learning the personalized representations of the heterogeneous distribution bias. Then, we devise corresponding loss functions, aggregation strategy, and updating strategy in order to make the two modules intelligently complement each other. Moreover, we conduct extensive experiments to evaluate the effectiveness of PFed-DBA. The results show that PFed-DBA improves model accuracy to 12.34% at best compared with the state-of-the-art.
Meihan Wu, Li Li 0064, Tao Chang, Jie Zhou 0032, Eric Rigall, Cui Miao, Xiaodong Wang 0002, Cheng-Zhong Xu 0001
IWQoS3
2024 Vision-language navigation: a survey and taxonomy
Wansen Wu, Tao Chang, Xinmeng Li, Quanjun Yin, Yue Hu 0016
Neural Comput. Appl.2
2024 Dynamic attention augmented graph network for video accident anticipation
Wenfeng Song, Shuai Li 0001, Tao Chang, Ke Xie 0005, Aimin Hao, Hong Qin 0001
Pattern Recognit.3
2023 UATR: An Uncertainty Aware Two-Stage Refinement Model for Targeted Sentiment Analysis
Qingsong Yin, Qingbing Ji, Tao Chang
ICONIP (5)6
2023 FedEAE: Federated Learning Based Privacy-Preserving Event Argument Extraction
Fei Hu 0005, Shenpo Dong, Tao Chang, Jie Zhou 0032, Haili Li, Jingnan Wang, Haijiao Liu, Xiaodong Wang 0002
NLPCC (2)3
2023 FedHybrid: Hierarchical Hybrid Training for High-Performance Federated Learning
abstract
Federated Learning coordinates multiple devices to train a shared model while preserving data privacy. Despite its potential benefit, the increasing number of participating devices poses new challenges to the deployment in real-world cases. The highly limited amount of data located on each device coupled with significantly unbalanced data across different devices severely impede the performance of the shared model and the overall training progress at the same time.In this paper, we propose FedHybrid, a hierarchical hybrid training framework for high-performance Federated Learning on a wide scale. Unlike the existing work that mainly focuses on the statistical challenge, FedHybrid establishes a hierarchical hybrid training framework that effectively utilizes the fragmented and unbalanced data located on the participating devices on a wide scale. Specifically, FedHybrid consists of the following two core components, a global coordinator deployed on the central server and a local coordinator deployed on each participating device. The global coordinator organizes the participating devices into different groups through jointly considering the system heterogeneity and unbalanced training data in order to accelerate the overall training progress while guaranteeing the model performance. Within each group, a novel device-to-device (D2D) sequential training procedure is coordinated by the local coordinator to effectively utilize the fragmented and unbalanced training data in order to intelligently update the local models. At the same time, we provide the theoretical analysis of FedHybrid and conduct extensive experiments to evaluate its effectiveness. The results show that FedHybrid effectively improves model accuracy up to 27% and accelerates the whole training process by 20% on average.
Tao Chang, Li Li 0064, Meihan Wu, Wei Yu 0029, Xiaodong Wang 0002
SECON1
2023 PAGroup: Privacy-aware grouping framework for high-performance federated learning
Tao Chang, Li Li 0064, Meihan Wu, Wei Yu 0029, Xiaodong Wang 0002, Cheng-Zhong Xu 0001
J. Parallel Distributed Comput.1
2023 GraphCS: Graph-based client selection for heterogeneity in federated learning
Tao Chang, Li Li 0064, Meihan Wu, Wei Yu 0029, Xiaodong Wang 0002, Cheng-Zhong Xu 0001
J. Parallel Distributed Comput.1
2022 ArgumentPrompt: Activating Multi-category of Information for Event Argument Extraction with Automatically Generated Prompts
Shenpo Dong, Wei Yu 0029, Hongkui Tu, Xiaodong Wang 0002, Yunyan Zhou, Haili Li, Jie Zhou 0032, Tao Chang
NLPCC (1)8
2022 FedGosp: A Novel Framework of Gossip Federated Learning for Data Heterogeneity
abstract
Federated learning (FL) provides the possibility to solve the problem of data privacy, but it suffers much from the data heterogeneity among different participants. Currently, some promising FL algorithms improve the effectiveness of learning under the non independent-and-identically-distributed (Non-IID) data settings. However, they require a large number of communication rounds between the server and clients for an acceptable accuracy. Inspired by the training paradigm of gossip learning, this paper proposes a new FL framework, named FedGosp. It first classifies the clients into different categories based on the model weights trained by the locally stored data. Then FedGosp utilizes the communication not only between clients and the server, but also between different classes of clients themselves. This training process enables instilling knowledge about various data distributions in the passed models. We evaluate the performance of FedGosp in multiple Non-IID settings on CIFAR10 and MNIST datasets, and compare it with the recently popular algorithms such as SCAFFOLD, FedAvg and FedProx. The experimental results show that FedGosp can improve the model accuracy by 6.53% and save 5.6 × communication costs at best compared to the second-ranked baseline.
Yue Hu 0016, Miao Zhang 0037, Li Li 0064, Tao Chang, Quanjun Yin
SMC5
2021 Similar Questions Correspond to Similar SQL Queries: A Case-Based Reasoning Approach for Text-to-SQL Translation
Wei Yu 0029, Tao Chang, Mengzhu Wang, Xiaodong Wang 0002
ICCBR4
2021 An interaction-modeling mechanism for context-dependent Text-to-SQL translation based on heterogeneous graph aggregation
Wei Yu 0029, Tao Chang, Mengzhu Wang, Xiaodong Wang 0002
Neural Networks2
2020 Cross-View Contextual Relation Transferred Network for Unsupervised Vehicle Tracking in Drone Videos
abstract
Recently CNN-centric object tracking methods have been gaining tremendous success in ground-view videos, however, it remains hard to cope with vehicle tracking in unmanned aerial vehicle (UAV) videos. The key difficulties mainly stem from lacking large-scale well-labeled training datasets and view-invariant appearance model for fast-moving drone-view vehicles. We enhance the vehicle's cross-view feature by exploring relations between the pivotal context and the target to facilitate unsupervised vehicle tracking. The relation is modeled as the relevance of the target and its contextual regions in the tracking task. Specifically, we propose a contextual relation actor-critic (CRAC) framework integrates an actor-critic agent with a dual GAN learning mechanism, which aims to dynamically search the related contextual regions and transfer the relations from ground-view to drone-view videos while retaining the discriminative features. We demonstrate that CRAC could be applied to several state-of-the-art trackers by extensive experiments and ablation studies on four public benchmarks. All the experiments confirm that, our CRAC can improve the performance of state-of-the-art methods in terms of accuracy, robustness, and versatility.
Wenfeng Song, Shuai Li 0001, Tao Chang, Aimin Hao, Qinping Zhao, Hong Qin 0001
WACV3
2020 Context-Interactive CNN for Person Re-Identification
abstract
Despite growing progresses in recent years, cross-scenario person re-identification remains challenging, mainly due to the pedestrians commonly surrounded by highly-complex environment contexts. In reality, the human perception mechanism could adaptively find proper contextualized spatial-temporal clues towards pedestrian recognition. However, conventional methods fall short in adaptively leveraging the long-term spatial-temporal information due to ever-increasing computational cost. Moreover, CNN-based deep learning methods are hard to conduct optimization due to the non-differentiable property of the built-in context search operation. To ameliorate, this paper proposes a novel Context-Interactive CNN (CI-CNN) to dynamically find both spatial and temporal contexts by embedding multi-task Reinforcement Learning (MTRL). The CI-CNN streamlines the multi-task reinforcement learning by using an actor-critic agent to capture the temporal-spatial context simultaneously, which comprises a context-policy network and a context-critic network. The former network learns policies to determine the optimal spatial context region and temporal sequence range. Based on the inferred temporal-spatial cues, the latter one focuses on the identification task and provides feedback for the policy network. Thus, CI-CNN can simultaneously zoom in/out the perception field in spatial and temporal domain for the context interaction with the environment. By fostering the collaborative interaction between the person and context, our method could achieve outstanding performance on various public benchmarks, which confirms the rationality of our hypothesis, and verifies the effectiveness of our CI-CNN framework.
Wenfeng Song, Shuai Li 0001, Tao Chang, Aimin Hao, Qinping Zhao, Hong Qin 0001
IEEE Trans. Image Process.3
2019 Research on Fine-Grained Sentiment Classification
Zhihui Wang 0016, Xiaodong Wang 0002, Tao Chang, Shaohe Lv
NLPCC (2)3
2003 A Comparative Framework for EB Systems Development Methodologies
abstract
EB systems development methodology is one of active researches in the field of software engineering, information systems and EB. Firstly, we introduce the existing EB systems development methodologies briefly. Then, we put forward the comparative framework to be as criteria to compare methodologies. Finally, we compare the development methodologies by the framework. The conclusion of the comparison can be used to design a new development methodology and improve existing EB systems development methodologies.
Jinghua Huang, Tao Chang, Chunjun Zhao
Web Intelligence3