Han Hu 0003

dblp:59/3754-3 · DBLP profile ↗
← Back
133ranked-venue papers
17as first author
96since 2021 · last 2026
0000-0001-7532-0496ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 50 · 9 first-author · 37 since 2021Graphics, computer vision, multimedia, augmented reality and games · 46 · 7 first-author · 27 since 2021Artificial intelligence and machine learning · 31 · 29 since 2021Systems, architecture and hardware · 3 · 1 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Learning-Based Resource Allocation for Integrated Sensing, Communication, and Computation Networks: A Delay-Aware Approach
abstract
Integrated sensing, communication, and computation (ISCC) network has been recognized as a key enabler to realize the vision of Internet-of-Things. In this paper, we explore the resource allocation problem in ISCC networks, where the task execution workflow consists of multiple dependent processes, i.e., wireless sensing, signal processing, data delivery, and data processing. To this end, a tandem-parallel queuing model is first proposed to characterize the end-to-end (E2E) task execution process. Given the model, the E2E delay upper bound is derived according to the stochastic network calculus theory. Based on the analytical results, the joint allocation problem of the sensing, communication, and computation (SCC) resources is formulated to minimize the E2E delay while satisfying the constraints of network resources, tolerable delay, and sensing mutual information, etc. Further, this non-convex optimization problem is parameterized to enable a learning-based optimization approach. Next, we design the unsupervised learning (UL) framework based on multilevel decomposition architecture (MDA) and residual network (RN) to accelerate training speed and ensure effective primal-dual learning. Numerical results demonstrate that the proposed UL-MDA-RN framework is superior to existing baselines with excellent convergence efficiency and lower achieved E2E delay. In addition, our results analyze the impacts of the network parameters on the E2E delay performance to guide the design of appropriate SCC resource provisioning patterns.
Mengxin Yang, Yixiao Gu, Han Hu 0003, Dan Zeng 0001
IEEE Internet Things J.3
2026 Adaptive Batch Size Time Evolving Stochastic Gradient Descent for Federated Learning
abstract
Variance reduction has been shown to improve the performance of Stochastic Gradient Descent (SGD) in centralized machine learning. However, when it is extended to federated learning systems, many issues may arise, including (i) mega-batch size settings; (ii) additional noise introduced by the gradient difference between the current iteration and the snapshot point; and (iii) gradient (statistical) heterogeneity. In this paper, we propose a lightweight algorithm termed federated adaptive batch size time evolving variance reduction (FedATEVR) to tackle these issues, consisting of an adaptive batch size setting scheme and a time-evolving variance reduction gradient estimator. In particular, we use the historical gradient information to set an appropriate mega-batch size for each client, which can steadily accelerate the local SGD process and reduce the computation cost. The historical information involves both global and local gradient, which mitigates unstable varying in mega-batch size introduced by gradient heterogeneity among the clients. For each client, the gradient difference between the current iteration and the snapshot point is used to tune the time-evolving weight of the variance reduction term in the gradient estimator. This can avoid meaningless variance reduction caused by the out-of-date snapshot point gradient. We theoretically prove that our algorithm can achieve a linear speedup of of $\mathcal {O}(\frac{1}{\sqrt{SKT}})$O(1SKT) for non-convex objective functions under partial client participation. Extensive experiments demonstrate that our proposed method can achieve higher test accuracy than the baselines and decrease communication rounds greatly.
Xuming An 0001, Li Shen 0008, Yong Luo 0002, Han Hu 0003, Dacheng Tao
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Hypergraph Foundation Model
abstract
Hypergraph neural networks (HGNNs) effectively model complex high-order relationships in domains like protein interactions and social networks by connecting multiple vertices through hyperedges, enhancing modeling capabilities, and reducing information loss. Developing foundation models for hypergraphs is challenging due to their distinct data, which includes both vertex features and intricate structural information. We present Hyper-FM, a Hypergraph Foundation Model for multi-domain knowledge extraction, featuring Hierarchical High-Order Neighbor Guided Vertex Knowledge Embedding for vertex feature representation and Hierarchical Multi-Hypergraph Guided Structural Knowledge Extraction for structural information. Additionally, we curate 11 text-attributed hypergraph datasets to advance research between HGNNs and LLMs. Experiments on these datasets show that Hyper-FM outperforms baseline methods by approximately 13.4%, validating our approach. Furthermore, we propose the first scaling law for hypergraph foundation models, demonstrating that increasing domain diversity significantly enhances performance, unlike merely augmenting vertex and hyperedge counts. This underscores the critical role of domain diversity in scaling hypergraph models.
Yue Gao 0002, Yifan Feng 0001, Shiquan Liu, Xiangmin Han, Shaoyi Du, Zongze Wu 0001, Han Hu 0003
IEEE Trans. Pattern Anal. Mach. Intell.7
2026 HGNNv2: Stable Hypergraph Neural Networks
abstract
Hypergraph neural networks (HGNNs) are widely used models for analyzing higher-order relational data. HGNNs suffer from the rapid performance degradation with increasing layers. Hypergraph dynamic system (HDS) is a potential way to deal with this challenge. However, hypergraph dynamic system is confined to a time-continuous isotropic model, lacking positional information in the structural space of the hypergraph. In contrast, anisotropic diffusion can capture structural space differences among vertices, providing a more precise representation of the information propagation process in hypergraph structures than isotropic diffusion. In this paper, we introduce HGNNv2, a stable hypergraph neural network, which is built as a hypergraph dynamic system with partial differential equation (PDE). This model incorporates a position-aware anisotropic diffusion term and an external control term. We further present the vertex-rooted subtree method to determine anisotropic diffusion intensity. HGNNv2 has properties that vertices occupying equivalent positions in the structural space share equivalent structural labels and positional features. Experiments on 6 hypergraph datasets and 3 graph datasets reveal that HGNNv2 outperforms all 12 compared methods. HGNNv2 is capable of achieving stable final representations and task accuracy even under noisy conditions. HGNNv2 achieves stable performance with fewer layers than hypergraph dynamic systems employing isotropic diffusion. We provide feature visualizations to illustrate the evolution of representations.
Yue Gao 0002, Jielong Yan, Yifan Feng 0001, Xiangmin Han, Shihui Ying, Zongze Wu 0001, Han Hu 0003
IEEE Trans. Pattern Anal. Mach. Intell.7
2026 Toward Effective Knowledge Distillation: Navigating Beyond Small-Data Pitfall
abstract
The spectacular success of training large models on extensive datasets highlights the potential of scaling up for exceptional performance. To deploy these models on edge devices, knowledge distillation (KD) is commonly used to create a compact model from a larger, pretrained teacher model. However, as models and datasets rapidly scale up in practical applications, it is crucial to consider the applicability of existing KD approaches originally designed for limited-capacity architectures and small-scale datasets. In this paper, we revisit current KD methods and identify the presence of a small-data pitfall, where most modifications to vanilla KD prove ineffective on large-scale datasets. To guide the design of consistently effective KD methods across different data scales, we conduct a meticulous evaluation of the knowledge transfer process. Our findings reveal that incorporating more useful information is crucial for achieving consistently effective KD methods, while modifications in loss functions show relatively less significance. In light of this, we present a paradigmatic example that combines vanilla KD with deep supervision, incorporating additional information into the student during distillation. This approach surpasses almost all recent KD methods. We believe our study will offer valuable insights to guide the community in navigating beyond the small-data pitfall and toward consistently effective KD.
Zhiwei Hao 0001, Jianyuan Guo, Kai Han 0002, Han Hu 0003, Chang Xu 0002, Yunhe Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Zero-Shot Sparse Mixture of Low-Rank Experts Construction From Pre-Trained Foundation Models
abstract
Deep model training on extensive datasets is increasingly becoming cost-prohibitive, prompting the widespread adoption of deep model fusion techniques to leverage knowledge from pre-existing models. From simple weight averaging to more sophisticated methods like AdaMerging, model fusion effectively improves model performance and accelerates the development of new models. However, potential interference between parameters of individual models and the lack of interpretability in the fusion progress remain significant challenges. Existing methods often try to resolve the parameter interference issue by evaluating attributes of parameters, such as their magnitude or sign, or by parameter pruning. In this study, we begin by examining the fine-tuning of linear layers through the lens of subspace analysis and explicitly define parameter interference as an optimization problem to shed light on this subject. Subsequently, we introduce an innovative approach to model fusion called zero-shot Sparse MIxture of Low-rank Experts (SMILE) construction, which allows for the upscaling of source models into an MoE model without extra data or further training. Our approach relies on the observation that fine-tuning mostly keeps the important parts from the pre-training, but it uses less significant or unused areas to adapt to new tasks. Additionally, the issue of parameter interference, which is intrinsically challenging in the original parameter space, can be managed by expanding the dimensions. We conduct extensive experiments across diverse scenarios, such as image classification and text generation tasks, using full fine-tuning and LoRA fine-tuning, and we apply our method to large language models (CLIP models, Flan-T5 models, and Mistral-7B models), highlighting the adaptability and scalability of SMILE. For full fine-tuned models, about 50% additional parameters can achieve around 98-99% of the performance of eight individual fine-tuned ViT models, while for LoRA fine-tuned Flan-T5 models, maintaining 99% performance with only 2% extra parameters. Code is available athttps://github.com/tanganke/fusion_bench.
Anke Tang, Li Shen 0008, Yong Luo 0002, Shuai Xie, Han Hu 0003, Lefei Zhang, Bo Du 0001, Dacheng Tao
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Joint Sensing, Communication, and Computation for Vertical Federated Edge Learning in Edge Perception Networks
abstract
Combining wireless sensing and edge intelligence, edge perception networks enable intelligent data collection and processing at the network edge. However, traditional sample partition based horizontal federated edge learning (HFEEL) struggles to effectively fuse complementary multi-view information from distributed devices. To address this limitation, we propose a vertical federated edge learning (VFEEL) framework tailored for feature-partitioned sensing data. In this paper, we consider an integrated sensing, communication, and computation (ISCC)-enabled edge perception network, where multiple edge devices utilize wireless signals to sense environmental information for updating their local models, and the edge server aggregates feature embeddings via over-the-air computation (AirComp) for global model training. First, we analyze the convergence behavior of the ISCC-enabled VFEEL in terms of the loss function degradation in the presence of wireless sensing noise and aggregation distortions during AirComp. Then, to accelerate convergence, we aim to optimize the batch size, sensing power, and transmission power control at edge devices as well as the denoising factors at the edge server under limited network constraints on overall energy consumption and per-round latency. Due to the tight coupling of variables, the problem is non-convex. To address this problem, we design an alternating optimization-based algorithm to efficiently obtain a high-quality solution. Numerical results are conducted based on a human motion recognition task to verify that the proposed ISCC-enabled VFEEL algorithm achieves higher accuracy compared with other benchmarking schemes including ISCC-enabled HFEEL approach.
Xiaowen Cao 0001, Dingzhu Wen, Suzhi Bi, Yuanhao Cui, Guangxu Zhu, Han Hu 0003, Yonina C. Eldar
IEEE Trans. Mob. Comput.6
2026 Joint UAV Placement and Dependent Task Offloading in Multi-UAV MEC Networks: A Graph Attention Enhanced DRL Approach
abstract
Unmanned aerial vehicles (UAVs) have emerged as effective platforms for mobile edge computing (MEC), offering flexible and efficient computational support to ground users (GUs). Many practical applications, such as deep neural network inference tasks, generate subtasks with complex dependencies, significantly complicating scheduling and offloading decisions. In this paper, we study the joint optimization of UAV deployment, UAV-GU associations, and dependent task offloading decisions within a multi-UAV-enabled MECsystem, aiming to minimize the end time of the overall tasks. The tasks generated by GUs are modeled using directed acyclic graphs (DAGs), explicitly capturing subtask dependencies and execution orders. To address the resulting complex optimization problem, we first propose a Joint Successive convex approximation and Penalty dual decomposition-based Optimization (JSPO) algorithm to determine the initial UAV deployment and UAV-GU associations. Next, we formulate the dependent task offloading decision process as a Markov decision process (MDP), which is solved by employing deep reinforcement learning (DRL). To effectively exploit the structural information within DAG tasks, we integrate a graph attention network (GAT) to provide enhanced state representations for DRL. JSPO and the DRL framework were executed in turns to gradually improve the performance. Extensive simulation results verify that our proposed framework significantly reduces the end time compared to existing methods, demonstrating its superiority in multi-UAV MEC systems.
Cheng Zhan, Kaifeng Song, Rongfei Fan, Jun Liu 0006, Han Hu 0003
IEEE Trans. Mob. Comput.6
2026 EEformer: Early Exiting for Transformer With Global-Local Exits and Progressive Fine-Tuning
abstract
Recently, the efficient deployment and acceleration of transformer-based pre-trained models (TPMs) on resource-constrained edge devices for multimedia services have gained significant interest. Although early exiting is a feasible solution, it may lead to extra computational cost and substantial performance degradation compared to the original models. To tackle these issues, we propose a framework termed EEformer, which incorporates global-local heads (GLHs) into intermediate layers to construct the early exiting dynamic neural network (EDNN). The GLH can efficiently extract global and local information from hidden states produced by the backbone layer, thereby achieving a better performance-efficiency trade-off for the EDNN. Moreover, we propose a novel progressive fine-tuning strategy to steadily improve the efficiency of the EDNN while maintaining its performance comparable to the original mode through three fine-tuning stages. We conduct extensive experiments on image classification and natural language processing tasks, demonstrating the superiority of the proposed framework. In particular, the proposed framework achieves 1.87× speed-up while maintaining 99.0% performance on the CIFAR-100 dataset, and 3.05× speed-up while maintaining 98.5% performance on the SST-2 dataset.
Guanyu Xu, Yong Luo 0002, Li Shen 0008, Han Hu 0003, Dan Zeng 0001
IEEE Trans. Multim.5
2026 Deep Model Fusion: A Survey
abstract
Deep model fusion/merging is an emerging technique that integrates parameters or predictions from multiple deep learning (DL) models into a unified framework. It combines the abilities of different models to compensate for the biases and errors of an individual model, improving overall performance. However, deep model fusion, especially on large-scale DL models such as large language models (LLMs) and foundation models, faces several challenges, including high computational cost and interference between different heterogeneous models. In order to understand it better, we present a comprehensive survey to summarize the recent progress. We categorize existing model fusion methods as fourfold: 1) weight average (WA) averages the parameters of multiple models to obtain results closer to the optimal solution; 2) considering that direct averaging of models often yields suboptimal results, "mode connectivity" connects networks via paths of nonincreasing loss in weight spaces before the fusion. Along these paths, initial models are transformed into forms with consistent functions and better fusion effects; 3) similarly, for models with poor direct fusion results, "alignment" matches the corresponding units and merges these models, thus fully exploiting the corresponding relationships between the models; and 4) in addition to the above-mentioned methods of parameter fusion, "ensemble learning" fuses the outputs of multiple models in the inference stage to improve the accuracy and robustness of networks. In addition, we analyze the challenges of deep model fusion and illuminate the possible research directions in the future.
Yong Peng 0006, Miao Zhang 0037, Liang Ding 0006, Han Hu 0003, Li Shen 0008
IEEE Trans. Neural Networks Learn. Syst.5
2026 Cooperative ISAC Systems With Extended Targets: Performance Analysis and Beamforming Design
abstract
This paper investigates a cooperative integrated sensing and communication (ISAC) system, where multiple base stations (BSs) employ coordinated transmit beamforming to communicate with their respective users and jointly sense a set of extended targets (ETs). Different from prior cooperative ISAC works considering point targets (PTs), the visible scatterers on the same extended target (ET) and the radar cross section (RCS) of the same scatterers observed by multiple BSs are considered to be different. Given the model, we first derive the Cramér-Rao bound (CRB) for the BSs to cooperatively estimate the ET’s parameters, thus quantifying the cooperative sensing gains. Next, based on the derived CRB, we formulate a joint node selection and coordinated transmit beamforming design problem with the object of minimizing the average trace of sensing CRB, while satisfying the minimum communication rate constraint, the maximum transmit power constraint, and the node selection constraints. To solve this non-convex optimization problem, we first utilize the block coordinate descent (BCD) method to decompose it into node selection sub-problem and beamforming design sub-problem. Next, the continuous relaxation and linear programming (LP) approach are employed to handle the node selection sub-problem, and a decentralized augmented Lagrangian manifold optimization algorithm is developed to solve the beamforming sub-problem with reduced computation complexity. Numerical simulations demonstrate that the proposed design outperforms benchmark designs with larger CRB-rate region. Moreover, our results show the impacts of the ET’s state and the number of network nodes on network performance to enable valuable ISAC beamforming design insights.
Yixiao Gu, Han Hu 0003, Jie Xu 0002, Dan Zeng 0001
IEEE Trans. Wirel. Commun.3
2026 SigGen: Signal Generation for Wireless Sensing Based on Disentangled Representation
abstract
With the thriving artificial intelligence-generated content (AIGC), it is becoming increasingly appealing to exploit generative AI to generate wireless signals for facilitating wireless sensing. However, this is a challenging task, as wireless signals are highly random in general and contain rich physical information. To tackle these challenges, we propose a novel signal disentanglement and generation framework termed SigGen, which is inspired by the Fourier Transform (FT) that converts signals to the frequency domain and accordingly separates objectives by distinct frequency bands. In our proposed framework, we first disentangle the features of objects embedded in the signal and subsequently modify these features to generate the desired signals. Specifically, we devise a neural network based on the vision transformer (ViT) to extract effective features for signal generation. In this neural network, we incorporate both local and global frequency attention modules to adaptively leverage frequency features, and introduce a hybrid patch embedding module to enhance information interaction for the ViT architecture. Furthermore, we propose a novel sequential training method to improve the disentanglement and generation capability of the neural network. Finally, extensive experiments on two benchmark public wireless sensing datasets demonstrate that our framework can effectively decouple wireless signals and generate diverse signals closely resembling real ones, surpassing state-of-the-art methods by 30.83%. A practical case study further demonstrates that our framework can be used as a data augmentation method to improve gesture recognition accuracy by 12.74%.
Hanxiang He, Xintao Huan, Yong Luo 0002, Rongfei Fan, Jie Xu 0002, Han Hu 0003
IEEE Trans. Wirel. Commun.6
2026 UAV-Enabled Aerial Monitoring Aided by STAR-RIS: A Stochastic Optimization Framework
abstract
This paper studies the unmanned aerial vehicle (UAV)-enabled aerial monitoring assisted by simultaneous transmitting and reflecting reconfigurable intelligent surfaces (STAR-RISs), in which one UAV aims to monitor a number of moving targets, and one STAR-RIS is installed on a building for assisting the UAV to broadcast the monitored information to both indoor and outdoor users. Due to the randomness of target movements over time, the UAV needs to adaptively adjust its flight trajectory to track them. This thus results in highly dynamic channel conditions and uncertain UAV energy consumption, which accordingly make the efficient aerial monitoring a challenging task. To address these challenges, we propose a STAR-RIS-aided UAV-enabled aerial monitoring framework, which aims to maximize the long-term average throughput for all users, through joint optimization of transmit beamforming, UAV trajectory, and STAR-RIS configuration, while ensuring the monitoring requirements under strict energy constraints. The formulated problem is a multi-stage stochastic optimization problem, due to the randomness of various system parameters. To handle this problem, we apply the Lyapunov optimization technique and introduce a virtual energy queue to transform it into a series of single-slot optimization subproblems that are solvable online. For each subproblem, we develop efficient algorithms to obtain a near-optimal solution, in which a penalty dual decomposition (PDD) approach is used for the transmit beamforming and STAR-RIS configuration optimization, and a sequential parametric convex approximation (SPCA) method is used for UAV trajectory optimization. Extensive simulations demonstrate that the proposed framework significantly outperforms benchmark schemes, effectively maximizing the throughput and energy efficiency under dynamic operational conditions.
Cheng Zhan, Kaifeng Song, Rongfei Fan, Han Hu 0003, Jie Xu 0002
IEEE Trans. Wirel. Commun.5
2025 Cross-Silo Feature Space Alignment for Federated Learning on Clients with Imbalanced Data
abstract
Data imbalance across clients in federated learning often leads to different local feature space partitions, harming the global model's generalization ability. Existing methods either employ knowledge distillation to guide consistent local training or performs procedures to calibrate local models before aggregation. However, they overlook the ill-posed model aggregation caused by imbalanced representation learning. To address this issue, this paper presents a cross-silo feature space alignment method (FedFSA), which learns a unified feature space for clients to bridge inconsistency. Specifically, FedFSA consists of two modules, where the in-silo prototypical space learning (ISPSL) module uses predefined text embeddings to regularize representation learning, which can improve the distinguishability of representations on imbalanced data. Subsequently, it introduces a variance transfer approach to construct the prototypical space, which aids in calibrating minority classes feature distribution and provides necessary information for the cross-silo feature space alignment (CSFSA) module. Moreover, the CSFSA module utilizes augmented features learned from the ISPSL module to learn a generalized mapping and align these features from different sources into a common space, which mitigates the negative impact caused by imbalanced factors. Experimental results from three datasets verified that FedFSA improves the consistency between diverse spaces on imbalanced data, which results in superior performance compared to existing methods.
Zhuang Qi, Lei Meng 0001, Zhaochuan Li, Han Hu 0003, Xiangxu Meng
AAAI4
2025 Targeted Low-rank Refinement: Enhancing Sparse Language Models with Precision
abstract
Pruning is a widely used technique for compressing large neural networks that eliminates weights that have minimal impact on the model's performance. Current pruning methods, exemplified by magnitude pruning, assign an importance score to each weight based on its magnitude and remove weights with scores below a certain threshold. Nonetheless, these methods often create a gap between the original dense and the pruned sparse model, potentially impairing performance. Especially when the sparsity ratio is high, the gap becomes more pronounced. To mitigate this issue, we introduce a method to bridge the gap left by pruning by utilizing a low-rank approximation of the difference between the dense and sparse matrices. Our method entails the iterative refinement of the sparse weight matrix augmented by a low-rank adjustment. This technique captures and retains the essential information often lost during pruning, thereby improving the performance of the pruned model. Furthermore, we offer a comprehensive theoretical analysis of our approach, emphasizing its convergence properties and establishing a solid basis for its efficacy. Experimental results on LLaMa models validate its effectiveness on large language models across various pruning techniques and sparsity levels. Our method shows significant improvements: at 50\% sparsity, it reduces perplexity by 53.9\% compared to conventional magnitude pruning on LLaMa-7B. Furthermore, to achieve a specific performance target, our approach enables an 8.6\% reduction in model parameters while maintaining a sparsity ratio of about 50\%.
Li Shen 0008, Anke Tang, Yong Luo 0002, Tao Sun 0005, Han Hu 0003, Xiaochun Cao
ICML5
2025 Federated Deconfounding and Debiasing Learning for Out-of-Distribution Generalization
abstract
Attribute bias in federated learning (FL) typically leads local models to optimize inconsistently due to the learning of non-causal associations, resulting degraded performance. Existing methods either use data augmentation for increasing sample diversity or knowledge distillation for learning invariant representations to address this problem. However, they lack a comprehensive analysis of the inference paths, and the interference from confounding factors limits their performance. To address these limitations, we propose the Federated Deconfounding and Debiasing Learning (FedDDL) method. It constructs a structured causal graph to analyze the model inference process, and performs backdoor adjustment to eliminate confounding paths. Specifically, we design an intra-client deconfounding learning module for computer vision tasks to decouple background and objects, generating counterfactual samples that establish a connection between the background and any label, which stops the model from using the background to infer the label. Moreover, we design an inter-client debiasing learning module to construct causal prototypes to reduce the proportion of the background in prototype components. Notably, it bridges the gap between heterogeneous representations via causal prototypical regularization. Extensive experiments on 2 benchmarking datasets demonstrate that FedDDL significantly enhances the model capability to focus on main objects in unseen data, leading to 4.5% higher Top-1 Accuracy on average over 9 state-of-the-art existing methods.
Zhuang Qi, Sijin Zhou, Lei Meng 0001, Han Hu 0003, Han Yu 0001, Xiangxu Meng
IJCAI4
2025 MixPrompt: Efficient Mixed Prompting for Multimodal Semantic Segmentation
abstract
Recent advances in multimodal semantic segmentation show that incorporating auxiliary inputs—such as depth or thermal images—can significantly improve performance over single-modality (RGB-only) approaches. However, most existing solutions rely on parallel backbone networks and complex fusion modules, greatly increasing model size and computational demands. Inspired by prompt tuning in large language models, we introduce \textbf{MixPrompt}: a prompting-based framework that integrates auxiliary modalities into a pretrained RGB segmentation model without modifying its architecture. MixPrompt uses a lightweight prompting module to extract and fuse information from auxiliary inputs into the main RGB backbone. This module is initialized using the early layers of a pretrained RGB feature extractor, ensuring a strong starting point. At each backbone layer, MixPrompt aligns RGB and auxiliary features in multiple low-rank subspaces, maximizing information use with minimal parameter overhead. An information mixing scheme enables cross-subspace interaction for further performance gains. During training, only the prompting module and segmentation head are updated, keeping the RGB backbone frozen for parameter efficiency. Experiments across NYU Depth V2, SUN-RGBD, MFNet, and DELIVER datasets show that MixPrompt achieves improvements of 4.3, 1.1, 0.4, and 1.1 mIoU, respectively, over two-branch baselines, while using nearly half the parameters. MixPrompt also outperforms recent prompting-based methods under similar compute budgets.
Zhiwei Hao 0001, Zhongyu Xiao, Jianyuan Guo, Li Shen 0008, Yong Luo 0002, Han Hu 0003, Dan Zeng 0001
NeurIPS6
2025 Merging on the Fly Without Retraining: A Sequential Approach to Scalable Continual Model Merging
abstract
Deep model merging represents an emerging research direction that combines multiple fine-tuned models to harness their specialized capabilities across different tasks and domains. Current model merging techniques focus on merging all available models simultaneously, with weight interpolation-based methods being the predominant approach. However, these conventional approaches are not well-suited for scenarios where models become available sequentially, and they often suffer from high memory requirements and potential interference between tasks. In this study, we propose a training-free projection-based continual merging method that processes models sequentially through orthogonal projections of weight matrices and adaptive scaling mechanisms. Our method operates by projecting new parameter updates onto subspaces orthogonal to existing merged parameter updates while using an adaptive scaling mechanism to maintain stable parameter distances, enabling efficient sequential integration of task-specific knowledge. Our approach maintains constant memory complexity to the number of models, minimizes interference between tasks through orthogonal projections, and retains the performance of previously merged models through adaptive task vector scaling. Extensive experiments on CLIP-ViT models demonstrate that our method achieves a 5-8% average accuracy improvement while maintaining robust performance in different task orderings. Code is publicly available at https://github.com/tanganke/opcm .
Anke Tang, Enneng Yang, Li Shen 0008, Yong Luo 0002, Han Hu 0003, Lefei Zhang, Bo Du 0001, Dacheng Tao
NeurIPS5
2025 Self-Evolving Pseudo-Rehearsal for Catastrophic Forgetting with Task Similarity in LLMs
abstract
Continual learning for large language models (LLMs) demands a precise balance between $\textbf{plasticity}$ - the ability to absorb new tasks - and $\textbf{stability}$ - the preservation of previously learned knowledge. Conventional rehearsal methods, which replay stored examples, are limited by long-term data inaccessibility; earlier pseudo-rehearsal methods require additional generation modules, while self-synthesis approaches often generate samples that poorly align with real tasks, suffer from unstable outputs, and ignore task relationships. We present $\textbf{\textit{Self-Evolving Pseudo-Rehearsal for Catastrophic Forgetting with Task Similarity}}(\textbf{SERS})$, a lightweight framework that 1) decouples pseudo-input synthesis from label creation, using semantic masking and template guidance to produce diverse, task-relevant prompts without extra modules; 2) applies label self-evolution, blending base-model priors with fine-tuned outputs to prevent over-specialization; and 3) introduces a dynamic regularizer driven by the Wasserstein distance between task distributions, automatically relaxing or strengthening constraints in proportion to task similarity. Experiments across diverse tasks on different LLMs show that our SERS reduces forgetting by over 2\% points against strong pseudo-rehearsal baselines, by ensuring efficient data utilization and wisely transferring knowledge. The code will be released at https://github.com/JerryWangJun/LLM_CL_SERS/.
Liang Ding 0006, Shuai Wang 0011, Hongyu Li 0004, Yong Luo 0002, Huangxuan Zhao, Han Hu 0003, Bo Du 0001
NeurIPS7
2025 Efficient Federated Learning against Byzantine Attacks and Data Heterogeneity via Aggregating Normalized Gradients
abstract
Federated Learning (FL) enables multiple clients to collaboratively train models without sharing raw data, but is vulnerable to Byzantine attacks and data heterogeneity, which can severely degrade performance. Existing Byzantine-robust approaches tackle data heterogeneity, but incur high computational overhead during gradient aggregation, thereby slowing down the training process. To address this issue, we propose a simple yet effective Federated Normalized Gradients Algorithm (Fed-NGA), which performs aggregation by merely computing the weighted mean of the normalized gradients from each client. This approach yields a favorable time complexity of $\mathcal{O}(pM)$, where $p$ is the model dimension and $M$ is the number of clients. We rigorously prove that Fed-NGA is robust to both Byzantine faults and data heterogeneity. For non-convex loss functions, Fed-NGA achieves convergence to a neighborhood of stationary points under general assumptions, and further attains zero optimality gap under some mild conditions, which is an outcome rarely achieved in existing literature. In both cases, the convergence rate is $\mathcal{O}(1/T^{\frac{1}{2} - \delta})$, where $T$ denotes the number of iterations and $\delta \in (0, 1/2)$. Experimental results on benchmark datasets confirm the superior time efficiency and convergence performance of Fed-NGA over existing methods.
Shiyuan Zuo, Xingrun Yan, Rongfei Fan, Li Shen 0008, Puning Zhao, Jie Xu 0002, Han Hu 0003
NeurIPS7
2025 DM-PCL: Text-Driven Dual-Modal Prototype Consistency Learning for Weakly-Supervised Few-Shot Part Segmentation
Mengya Han, Yong Luo 0002, Han Hu 0003, Zengmao Wang, Lefei Zhang, Bo Du 0001, Ling-Yu Duan, Dacheng Tao
Int. J. Comput. Vis.3
2025 ADEM-VL: Adaptive and Embedded Fusion for Efficient Vision-Language Tuning
Zhiwei Hao 0001, Jianyuan Guo, Li Shen 0008, Yong Luo 0002, Han Hu 0003, Yonggang Wen 0001
Int. J. Comput. Vis.5
2025 Data-Adaptive Weight-Ensembling for Multi-task Model Fusion
Anke Tang, Li Shen 0008, Yong Luo 0002, Shiwei Liu 0003, Han Hu 0003, Bo Du 0001, Dacheng Tao
Int. J. Comput. Vis.5
2025 Sense+: A Plug-and-Play Signal Preprocessing Approach for Enhancing Human-Centered Wireless Sensing
abstract
Human-centered wireless sensing has been significantly advanced by artificial intelligence (AI) technologies. To enhance AI model performance, signal preprocessing, as a fundamental procedure, is widely employed for improving signal quality. However, existing methods are time-consuming, labor-intensive, and exhibit limited generalization. To address this issue, we first investigate the frequency spectrum of various signals. The results demonstrate that, in human-centered wireless applications, human activities significantly affect the low-frequency components in the signal spectrum. Motivated by this observation, we propose Sense+, a concise and versatile signal preprocessing module that can be seamlessly integrated into existing models to enhance sensing performance. Specifically, we transform the raw signals into a unified frequency domain, apply a learnable filter to process their frequency spectra, and then convert them back to the original signal domain. To accurately extract low-frequency features, we further propose a low-pass weight initialization method for the filter. Extensive experiments are conducted across various sensing tasks and signal types, including IR-UWB signals for person identification, mmWave radar signals for gesture recognition, and Wi-Fi signals for action recognition. The results highlight the effectiveness of Sense+ in enabling preprocessing across diverse wireless signals. Specifically, when equipped with Sense+, the average accuracy improves by 21.84% compared to conventional preprocessing methods. Additionally, Sense+ accelerates convergence and exhibits consistent generalization across different neural network models.
Hanxiang He, Xintao Huan, Heng Liu 0001, Han Hu 0003, Jianping An
IEEE Internet Things J.6
2025 FusionBench: A Unified Library and Comprehensive Benchmark for Deep Model Fusion
abstract
Deep model fusion is an emerging technique that unifies the predictions or parameters of several deep neural networks into a single better-performing model in a cost-effective and data-efficient manner. Although a variety of deep model fusion techniques have been introduced, their evaluations tend to be inconsistent and often inadequate to validate their effectiveness and robustness. We present FusionBench, the first benchmark and a unified library designed specifically for deep model fusion. Our benchmark consists of multiple tasks, each with different settings of models and datasets. This variety allows us to compare fusion methods across different scenarios and model scales. Additionally, FusionBench serves as a unified library for easy implementation and testing of new fusion techniques. FusionBench is open source and actively maintained, with community contributions encouraged.
Anke Tang, Li Shen 0008, Yong Luo 0002, Enneng Yang, Han Hu 0003, Lefei Zhang, Bo Du 0001, Dacheng Tao
J. Mach. Learn. Res.5
2025 ViF-SD2E: a robust weakly-supervised framework for neural decoding
Jingyi Feng, Yong Luo 0002, Han Hu 0003
Neural Comput. Appl.4
2025 Aligning Text-to-Image Diffusion Models With Constrained Reinforcement Learning
abstract
Reward finetuning has emerged as a powerful technique for aligning diffusion models with specific downstream objectives or user preferences. However, current approaches suffer from a persistent challenge of reward overoptimization, where models exploit imperfect reward feedback at the expense of overall performance. In this work, we identify three key contributors to overoptimization: (1) a granularity mismatch between the multi-step diffusion process and sparse rewards; (2) a loss of plasticity that limits the model's ability to adapt and generalize; and (3) an overly narrow focus on a single reward objective that neglects complementary performance criteria. Accordingly, we introduce Constrained Diffusion Policy Optimization (CDPO), a novel reinforcement learning framework that addresses reward overoptimization from multiple angles. Firstly, CDPO tackles the granularity mismatch through a temporal policy optimization strategy that delivers step-specific rewards throughout the entire diffusion trajectory, thereby reducing the risk of overfitting to sparse final-step rewards. Then we incorporate a neuron reset strategy that selectively resets overactive neurons in the model, preventing overoptimization induced by plasticity loss. Finally, to avoid overfitting to a narrow reward objective, we integrate constrained reinforcement learning with auxiliary reward objectives serving as explicit constraints, ensuring a balanced optimization across diverse performance metrics.
Ziyi Zhang 0001, Sen Zhang 0006, Li Shen 0008, Yibing Zhan, Yong Luo 0002, Han Hu 0003, Bo Du 0001, Yonggang Wen 0001, Dacheng Tao
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 PartSeg: Few-shot part segmentation via part-aware prompt learning
Mengya Han, Heliang Zheng, Yong Luo 0002, Han Hu 0003, Jing Zhang 0037, Bo Du 0001
Pattern Recognit.5
2025 CoFormer: Collaborating With Heterogeneous Edge Devices for Scalable Transformer Inference
abstract
The impressive performance of transformer models has sparked the deployment of intelligent applications on resource-constrained edge devices. However, ensuring high-quality service for real-time edge systems is a significant challenge due to the considerable computational demands and resource requirements of these models. Existing strategies typically either offload transformer computations to other devices or directly deploy compressed models on individual edge devices. These strategies, however, result in either considerable communication overhead or suboptimal trade-offs between accuracy and efficiency. To tackle these challenges, we propose a collaborative inference system for general transformer models, termed CoFormer. The central idea behind CoFormer is to exploit the divisibility and integrability of transformer. An off-the-shelf large transformer can be decomposed into multiple smaller models for distributed inference, and their intermediate results are aggregated to generate the final output. We formulate an optimization problem to minimize both inference latency and accuracy degradation under heterogeneous hardware constraints. DeBo algorithm is proposed to first solve the optimization problem to derive the decomposition policy, and then progressively calibrate decomposed models to restore performance. We demonstrate the capability to support a wide range of transformer models on heterogeneous edge devices, achieving up to 3.1× inference speedup with large transformer models. Notably, CoFormer enables the efficient inference of GPT2-XL with 1.6 billion parameters on edge devices, reducing memory requirements by 76.3%. CoFormer can also reduce energy consumption by approximately 40% while maintaining satisfactory inference performance.
Guanyu Xu, Zhiwei Hao 0001, Li Shen 0008, Yong Luo 0002, Fuhui Sun, Han Hu 0003, Yonggang Wen 0001
IEEE Trans. Computers7
2025 Energy-Efficient Image Semantic Communication: Architecture Design and Optimal Joint Allocation of Communication and Computation Resources
abstract
Semantic communication is an emerging paradigm with significant potential for image transmission. However, resource-efficient architecture design and resource allocation in this field have not received adequate research attention. This paper proposes a resource-efficient multi-branch semantic communication architecture based on saliency detection, aimed at optimizing computational efficiency in image transmission. The architecture leverages models with varying capacities to process regions of images with different complexities. We further address the problem of multi-user uplink semantic communication and resource allocation, focusing on minimizing the total energy consumption for communication and computation. The optimization problem, subject to user demand, computation, delay, and transmission power constraints, is non-convex due to the coupling of variables, making it challenging to solve. To tackle this, we introduce a two-level decomposition approach. The lower-level problem, given a fixed compression rate, is solved using Karush-Kuhn-Tucker (KKT) conditions to derive closed-form solutions for transmission power and computation frequency. The upper-level problem, which optimizes the compression rate, is reformulated as a monotone optimization problem for efficient solution finding. Numerical results demonstrate that the proposed architecture significantly reduces computational resource usage while maintaining image quality, and the resource allocation strategy effectively minimizes energy consumption, outperforming baseline schemes in terms of energy efficiency.
Han Hu 0003, Kaifeng Song, Rongfei Fan, Cheng Zhan, Jie Xu 0002, Jian Yang 0014
IEEE Trans. Circuits Syst. Video Technol.1
2025 Viewport Prediction for Volumetric Video Streaming by Exploring Video Saliency and User Trajectory Information
abstract
Volumetric video, also referred to as hologram video, is an emerging medium that represents 3D content in extended reality. As a next-generation video technology, it is poised to become a key application in 5G and future wireless communication networks. Because each user generally views only a specific portion of the volumetric video, known as the viewport, accurate prediction of the viewport is crucial for ensuring an optimal streaming performance. Despite its significance, research in this area is still in the early stages. To this end, this paper introduces a novel approach called Saliency and Trajectory-based Viewport Prediction (STVP), which enhances the accuracy of viewport prediction in volumetric video streaming by effectively leveraging both video saliency and viewport trajectory information. In particular, we first introduce a novel sampling method, Uniform Random Sampling (URS), which efficiently preserves video features while minimizing computational complexity. Next, we propose a saliency detection technique that integrates both spatial and temporal information to identify visually static and dynamic geometric and luminance-salient regions. Finally, we fuse saliency and trajectory information to achieve more accurate viewport prediction. Extensive experimental results validate the superiority of our method over existing state-of-the-art schemes. To the best of our knowledge, this is the first comprehensive study of viewport prediction in volumetric video streaming. We also make the source code of this work publicly available.
Jie Li 0015, Zhi Liu 0002, Peng Yuan Zhou, Richang Hong, Qiyue Li 0001, Han Hu 0003
IEEE Trans. Circuits Syst. Video Technol.7
2025 SNICK: Secure Node Identification Based on Covert Clock Feature Extraction for Cross-Environment Wireless IoT
abstract
Node identification is the first line of defense for the security of wireless Internet-of-Things (IoT), which prevents illegal devices from accessing the network and launching attacks. Hardware features originating from innate hardware manufacturing imperfections are considered promising fingerprints for identification; among which, the hardware clock feature has been put under the spotlight due to its practicality and ease of extraction. However, current extractions of hardware clock features over wireless networks rely on the transmissions of time information, which, per se, enable significant vulnerabilities such as spoofing and replay attacks. In this paper, we propose a covert method to extract the hardware clock features, which does not rely on the insecure time information transmissions that are adopted in most existing schemes. We also analyze the security of the proposed covert extraction. We further propose SNICK, a secure node identification scheme based on our tailored implementation of covert clock feature extraction and machine learning. We implement and evaluate the proposed approach on a real IoT testbed consisting of a Long Range (LoRa) gateway and heterogeneous end nodes. We conduct experiments to prove the security of the proposed scheme and evaluate the proposed scheme under three scenarios: short-term, long-term, and cross-environment. Experimental results of three scenarios demonstrate average identification accuracies of 98.53%, 85.9%, and 88.3%. We further reveal the identification performance under parameter and environmental variations.
Xintao Huan, Yixuan Zou, Shengkang Zhang, Han Hu 0003, Alan Marshall 0001
IEEE Trans. Inf. Forensics Secur.5
2025 KGPT: Wireless Key Generation Based on Power Tuning for Static Internet-of-Things
abstract
Wireless key generation is an emerging key sharing solution for Internet-of-Things (IoT) devices, which heavily relies on wireless signal fluctuations. However, most IoT devices have remained static ever since being deployed, where their stable wireless signals seriously deteriorate the effectiveness of key generation. In this article, we proposea new wireless Key Generation approach based on Power Tuning (KGPT) for statice IoT. Unlike the existing methods, KGPT does not require any additional helpers or any hardware modifications, thus, suiting practical static IoT deployments. We analyze the possibility of eavesdropping on the existing wireless key generation. We propose three power tuning strategies to generate wireless signal fluctuations for key generation in static IoT while defending against eavesdropping. In a static IoT scenario, we conduct experimental evaluations to first verify the eavesdropping on multi-antenna wireless key generation and then demonstrate the effectiveness of the proposed KGPT in producing sufficient signal fluctuations for key generation. Results show that KGPT can achieve a correlation value up to 0.97 for legitimate links while resisting the correlation of eavesdropping links down below 0.4.
Xintao Huan, Kaitao Miao, Changfan Wu, Rui Pei, Han Hu 0003
IEEE Trans. Ind. Informatics7
2025 ScaleNet: Scaling up Pretrained Neural Networks With Incremental Parameters
abstract
Recent advancements in vision transformers (ViTs) have demonstrated that larger models often achieve superior performance. However, training these models remains computationally intensive and costly. To address this challenge, we introduce ScaleNet, an efficient approach for scaling ViT models. Unlike conventional training from scratch, ScaleNet facilitates rapid model expansion with negligible increases in parameters, building on existing pretrained models. This offers a cost-effective solution for scaling up ViTs. Specifically, ScaleNet achieves model expansion by inserting additional layers into pretrained ViTs, utilizing layer-wise weight sharing to maintain parameters efficiency. Each added layer shares its parameter tensor with a corresponding layer from the pretrained model. To mitigate potential performance degradation due to shared weights, ScaleNet introduces a small set of adjustment parameters for each layer. These adjustment parameters are implemented through parallel adapter modules, ensuring that each instance of the shared parameter tensor remains distinct and optimized for its specific function. Experiments on the ImageNet-1K dataset demonstrate that ScaleNet enables efficient expansion of ViT models. With a $2\times $ depth-scaled DeiT-Base model, ScaleNet achieves a 7.42% accuracy improvement over training from scratch while requiring only one-third of the training epochs, highlighting its efficiency in scaling ViTs. Beyond image classification, our method shows significant potential for application in downstream vision areas, as evidenced by the validation in object detection task.
Zhiwei Hao 0001, Jianyuan Guo, Li Shen 0008, Kai Han 0002, Yehui Tang 0001, Han Hu 0003, Yunhe Wang 0001
IEEE Trans. Image Process.6
2025 MHR: A Multi-Modal Hyperbolic Representation Framework for Fake News Detection
abstract
The rapid growth of the internet has led to an alarming increase in the dissemination of fake news, which has had many negative effects on society. Various methods have been proposed for detecting fake news. However, these approaches suffer from several limitations. First, most existing works only consider news as separate entities and do not consider the correlations between fake news and real news. Moreover, these works are usually conducted in the Euclidean space, which is unable to capture complex relationships between news, in particular the hierarchical relationships. To tackle these issues, we introduce a novelMulti-modalHyperbolicRepresentation framework (MHR) for fake news detection. Specifically, we capture the correlations between news for graph construction to arrange and analyze different news. To fully utilize the multi-modal characteristics, we first extract the textual and visual information, and then design a Lorentzian multi-modal fusion module to fuse them as the node information in the graph. By utilizing the fully hyperbolic graph neural networks, we learn the graph’s representation in hyperbolic space, followed by a detector for detecting fake news. The experimental results on three real-world datasets demonstrate that our proposed MHR model achieves state-of-the-art performance, indicating the benefits of hyperbolic representation.
Shanshan Feng 0001, Guoxin Yu, Han Hu 0003, Yong Luo 0002, Yew-Soon Ong
IEEE Trans. Knowl. Data Eng.4
2025 P3ID: A Privacy-Preserving Person Identification Framework Towards Multi-Environments Based on Transfer Learning
abstract
Concerns surrounding privacy leakages caused by prevalent vision-based person identifications are countless. A promising privacy-preserving solution is to identify the wireless signals reflecting persons, which, however, faces a major challenge of losing efficacy in multi-environments. In this paper, we work on person identification based on wireless signals using transfer learning, toward tackling the performance deterioration across environments. We investigate the feature variations induced by environmental shifts based on data measurements. Lay our foundation on the feature alignment concept, we propose a novel wireless-based person identification framework using transfer learning. In the framework, we integrate a series of signal processing methods including signal selection, pre-processing, and augmentation, where the first includes a reference environment to assist the feature extraction while the latter two respectively reduce the data noise and improve the data diversity. We also propose a model generalization method where a neural network is employed to align features from different environments, which facilitates the extraction of environment-independent features while incorporating both person and environment information. On a real wireless testbed consisting of an Impulse Radio Ultra-WideBand (IR-UWB) radar, we build and publicly release a dataset with 22,264 samples of ten individuals from three environments, varying in testing distance and obstruction condition. Extensive experimental evaluations demonstrate that the proposed framework can improve the identification accuracy across environments, and surpasses state-of-the-art methods by up to 18.06%.
Hanxiang He, Xintao Huan, Jing Wang 0055, Yong Luo 0002, Han Hu 0003, Jianping An
IEEE Trans. Mob. Comput.5
2025 An Efficient Two-Stage Networking Topology Design for Mega-Constellation of Low Earth Orbit Satellites
abstract
Low Earth Orbit (LEO) satellites play a crucial role in providing high-speed internet to remote areas and ensuring network resilience during outages. The design of efficient satellite constellations requires optimizing network topology, which is a complex task due to the large solution space and the need for fault tolerance. This paper presents the AlphaSat algorithm, a two-phase approach to improve latency and network robustness in LEO constellations. In the initialization phase, Monte Carlo Tree Search (MCTS) is used to generate an initial topology by selecting links from a vast search space. In the refinement phase, an edge-switching method is applied to enhance network resilience and performance. AlphaSat is evaluated on OneWeb, Starlink, and Telesat mega-constellations, demonstrating superior performance over existing algorithms. The results show significant reductions in latency ranging from 4.7% to 44.5% and improvements in network robustness, increasing by 3.3% to 28.3%. Furthermore, AlphaSat effectively balances network load and optimizes power consumption, offering a promising solution for efficient and resilient LEO satellite network design.
Han Hu 0003, Yifeng Lyu, Kaifeng Song, Rongfei Fan, Cheng Zhan, Jian Yang 0014
IEEE Trans. Mob. Comput.1
2025 Joint Service Caching and Resource Allocation Over Different Timescales in Satellite Edge Computing Networks
abstract
The integration of edge computing into satellite networks offers a promising solution for extending computational services to remote and underserved areas. To effectively provide a variety of computing services, it is essential to cache the corresponding services on satellites. However, challenges exist such as dynamic computing requests that vary over time and space, energy constraints due to restricted power supply, as well as limited storage capacity on satellites and the impracticality of frequently adjusting service deployments. To tackle such challenges, this paper proposes a two-timescale joint optimization framework to minimize energy consumption in satellite edge computing networks while ensuring the delay requirements, by jointly optimizing service placement and task offloading, as well as computation resource and power allocation. On a larger timescale, we optimize service caching placement by strategically deploying services on satellites and ground devices (GDs) based on long-term service request statistics, aiming to minimize the total average delay over each time frame. We develop an efficient iterative algorithm by employing penalty-based methods and Lagrange duality techniques to achieve suboptimal service deployment. On a smaller timescale, we optimize task offloading and resource allocation in shorter time slots, adapting to dynamic traffic fluctuations to minimize energy consumption while meeting delay constraints. We utilize alternating optimization and quadratic transform methods to efficiently allocate resources and schedule tasks. Extensive simulations demonstrate the effectiveness and superiority of our framework over benchmark schemes, revealing significant reductions in delay and energy consumption. The results also highlight the trade-offs between task delay and energy consumption, as well as between transmit power and energy consumption.
Han Hu 0003, Kaifeng Song, Cheng Zhan, Rongfei Fan, Jian Yang 0014
IEEE Trans. Mob. Comput.1
2025 A Near-Optimal Category Information Sampling in RFID Systems
abstract
In many RFID-enabled applications, objects are classified into different categories, and the information associated with each object's category (called category information) is written into the attached tag, allowing the reader to access it later. The category information sampling in such RFID systems, which is to randomly choose (sample) a few tags from each category and collect their category information, is fundamental for providing real-time monitoring and analysis in RFID. However, to the best of our knowledge, two technical challenges, i.e., how to guarantee a minimized execution time and reduce collection failure caused by missing tags, remain unsolved for this problem. In this paper, we address these two limitations by considering how to use the shortest possible time to sample a different number of random tags from each category and collect their category information sequentially in small batches. In particular, we first obtain a lower bound on the execution time of any protocol that can solve this problem. Subsequently, we present a near-OPTimalCategory information sampling protocol (OPT-C) that solves the problem with an execution time close to the lower bound. Finally, extensive simulation results demonstrate the superiority of OPT-C over existing protocols, while real-world experiments further validate its practicality.
Xiujun Wang, Zhi Liu 0002, Xiaokang Zhou, Yong Liao 0003, Han Hu 0003, Jie Li 0002
IEEE Trans. Mob. Comput.5
2025 Sequential Federated Learning in Hierarchical Architecture on Non-IID Datasets
abstract
In a real federated learning (FL) system, communication overhead for passing model parameters between the clients and the parameter server (PS) is often a bottleneck. Hierarchical federated learning (HFL) that poses multiple edge servers (ESs) between clients and the PS can partially alleviate communication pressure but still needs the aggregation of model parameters from multiple ESs at the PS. To further reduce communication overhead, we remove the central PS, so that each iteration only completes model training by transmitting the global model between two adjacent ES. We call this serial learning method Sequential FL (SFL). For the first time, we introduced SFL into HFL and proposed a novel algorithm adapted to this combined framework, called Fed-CHS. Convergence results are derived for strongly convex and non-convex loss functions under various data heterogeneity setups, which show comparable convergence performance with the algorithms for HFL or SFL solely. Experimental results provide evidence of the superiority of our proposed Fed-CHS on both communication overhead saving and test accuracy over baseline methods.
Xingrun Yan, Shiyuan Zuo, Rongfei Fan, Han Hu 0003, Li Shen 0008, Puning Zhao, Yong Luo 0002
IEEE Trans. Mob. Comput.4
2025 Online Energy and Interference Management for Dynamic Target Tracking With Cellular-Connected UAV
abstract
Cellular-connected Unmanned Aerial Vehicles (UAVs) have significant potential for target tracking in future cellular networks due to their broad coverage and operational flexibility. In this paper, we consider a multi-cell cellular network with a cellular-connected UAV for target tracking, which encounters challenges such as unpredictable flight energy consumption from the stochastic movements of the tracking target and severe uplink interference from ground devices (GDs). To tackle these challenges, we propose a multi-stage stochastic optimization framework focused on energy-efficient target tracking with interference coordination. Our objective is to optimize the long-term average uplink throughput of both aerial users and GDs by jointly optimizing the UAV's trajectory, power allocation, and cell association across multiple orthogonal communication resource blocks (RBs). The formulated stochastic non-convex problem is first transformed into a deterministic problem for each time slot by using the Lyapunov optimization framework. An online optimization strategy is proposed, utilizing the optimal structure, alternative optimization, and successive convex approximation (SCA) techniques. Simulation results show that the proposed approach significantly enhances network throughput and UAV energy queue stability compared to existing baseline schemes.
Cheng Zhan, Rongfei Fan, Han Hu 0003, Shubin Xu, Jian Yang 0014
IEEE Trans. Mob. Comput.4
2025 Federated Learning Resilient to Byzantine Attacks and Data Heterogeneity
abstract
This paper addresses federated learning (FL) in the context of malicious Byzantine attacks and data heterogeneity. We introduce a novel Robust Average Gradient Algorithm (RAGA), which uses the geometric median for aggregation and allows flexible round number for local updates. Unlike most existing resilient approaches, which base their convergence analysis on strongly-convex loss functions or homogeneously distributed datasets, this work conducts convergence analysis for both strongly-convex and non-convex loss functions over heterogeneous datasets. The theoretical analysis indicates that as long as the fraction of the data from malicious users is less than half, RAGA can achieve convergence at a rate of$\mathcal {O}({1}/{T^{2/3- \delta }})$for non-convex loss functions, where$T$is the iteration number and$\delta \in (0, 2/3)$. For strongly-convex loss functions, the convergence rate is linear. Furthermore, the stationary point or global optimal solution is shown to be attainable as data heterogeneity diminishes. Experimental results validate the robustness of RAGA against Byzantine attacks and demonstrate its superior convergence performance compared to baselines under varying intensities of Byzantine attacks on heterogeneous datasets.
Shiyuan Zuo, Xingrun Yan, Rongfei Fan, Han Hu 0003, Hangguan Shan, Tony Q. S. Quek, Puning Zhao
IEEE Trans. Mob. Comput.4
2025 Federated Learning With Only Positive Labels by Exploring Label Correlations
abstract
Federated learning (FL) aims to collaboratively learn a model by using the data from multiple users under privacy constraints. In this article, we study the multilabel classification (MLC) problem under the FL setting, where trivial solution and extremely poor performance may be obtained, especially when only positive data with respect to a single class label is provided for each client. This issue can be addressed by adding a specially designed regularizer on the server side. Although effective sometimes, the label correlations are simply ignored and thus suboptimal performance may be obtained. Besides, it is expensive and unsafe to exchange user's private embeddings between server and clients frequently, especially when training model in the contrastive way. To remedy these drawbacks, we propose a novel and generic method termed federated averaging (FedAvg) by exploring label correlations (FedALCs). Specifically, FedALC estimates the label correlations in the class embedding learning for different label pairs and utilizes it to improve the model training. To further improve the safety and also reduce the communication overhead, we propose a variant to learn fixed class embedding for each client, so that the server and clients only need to exchange class embeddings once. Extensive experiments on multiple popular datasets demonstrate that our FedALC can significantly outperform the existing counterparts.
Xuming An 0001, Dui Wang, Li Shen 0008, Yong Luo 0002, Han Hu 0003, Bo Du 0001, Yonggang Wen 0001, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.5
2025 Integrated Sensing, Communication, and Powering Over Multi-Antenna OFDM Systems
abstract
This paper considers a multi-functional orthogonal frequency division multiplexing (OFDM) system with integrated sensing, communication, and powering (ISCAP), in which a multi-antenna base station (BS) transmits OFDM signals to simultaneously deliver information to multiple information receivers (IRs), provide energy supply to multiple energy receivers (ERs), and sense potential targets based on the echo signals. To facilitate ISCAP, the BS employs the joint transmit beamforming design by sending dedicated sensing/energy beams jointly with information beams. Furthermore, we consider the beam scanning for sensing, in which the joint beams scan in different directions over time to sense potential targets. In order to ensure the sensing beam scanning performance and meet the communication and powering requirements, it is essential to properly schedule IRs and ERs and design the resource allocation over time, frequency, and space. More specifically, we optimize the joint transmit beamforming over multiple OFDM symbols and subcarriers, with the objective of minimizing the average beampattern matching error of beam scanning for sensing, subject to the constraints on the average communication rates at IRs and the average harvested power at ERs. We find converged high-quality solutions to the formulated problem by proposing efficient iterative algorithms based on advanced optimization techniques. We also develop various heuristic designs based on the principles of zero-forcing (ZF) beamforming, round-robin user scheduling, and time switching, respectively. Numerical results show that our proposed algorithms adaptively generate information and sensing/energy beams at each time-frequency slot to match the scheduled IRs/ERs with the desired scanning beam, significantly outperforming the heuristic designs.
Yilong Chen 0003, Zixiang Ren, Han Hu 0003, Jie Xu 0002, Lexi Xu, Shuguang Cui
IEEE Trans. Wirel. Commun.4
2025 Dynamic Access Control in Multi-Layer Satellite Remote Sensing System Using Multi-Agent Deep Reinforcement Learning
abstract
The Multi-Layer Satellite Remote Sensing (SRS) integrates data collection by Low Earth Orbit (LEO) satellites and data processing assistance from Medium Earth Orbit (MEO) satellites, thereby playing a crucial role in scientific exploration. However, effectively controlling access to LEO satellites for processing data, especially considering the frequent handovers caused by speed differences, presents a significant challenge to achieving high energy efficiency services. To address this challenge, we explore cooperative dynamic access control based on efficient communication mechanisms, with the aim of prioritizing processed data volume and meeting energy consumption requirements for satellites. Specifically, we formulate the access control issue as an optimization problem and integrate it into the framework of partially observable Markov decision process (POMDP), considering MEO satellites’ limited observation ability. By employing Multi-agent Deep Reinforcement Learning (MADRL), we propose a novel dynamic access control algorithm named DAC to solve our featured problem. Specifically, for improving performance, communication-efficient cooperation among MEOs is enhanced through modeling decision-relevant information of fellow MEO satellites and maximizing mutual information with their actual data to extract precise awareness and enable the generation of concise message. Finally, we conduct comprehensive experiments and an ablation study spanning the Starlink, OneWeb, and Telesat mega-constellations. The results demonstrate that DAC increases the average system data processing volume by at least 13.5%, while meeting energy consumption constraints and outperforming baseline algorithms.
Han Hu 0003, Yifeng Lyu, Rongfei Fan, Xiufeng Sui, Cheng Zhan, Dusit Niyato
IEEE Trans. Wirel. Commun.1
2024 POCE: Primal Policy Optimization with Conservative Estimation for Multi-constraint Offline Reinforcement Learning
abstract
Multi-constraint offline reinforcement learning (RL) promises to learn policies that satisfy both cumulative and state- wise costs from offline datasets. This arrangement provides an effective approach for the widespread appli-cation of RL in high-risk scenarios where both cumulative and state-wise costs need to be considered simulta-neously. However, previously constrained offline RL algorithms are primarily designed to handle single-constraint problems related to cumulative cost, which faces challenges when addressing multi-constraint tasks that involve both cumulative and state-wise costs. In this work, we pro-pose a novel Primal policy Optimization with Conservative Estimation algorithm (POCE) to address the problem of multi-constraint offline RL. Concretely, we reframe the ob-jective of multi-constraint offline RL by introducing the con-cept of Maximum Markov Decision Processes (MMDP). Subsequently, we present a primal policy optimization al-gorithm to confront the multi-constraint problems, which improves the stability and convergence speed of model training. Furthermore, we propose a conditional Bell-man operator to estimate cumulative and state-wise Q-values, reducing the extrapolation error caused by out-of-distribution (OOD) actions. Finally, extensive experiments demonstrate that the POCE algorithm achieves competitive performance across multiple experimental tasks, particu-larly outperforming baseline algorithms in terms of safety. Our code is available at github. POCE.
Jiayi Guan, Li Shen 0008, Ao Zhou 0005, Lusong Li, Han Hu 0003, Xiaodong He 0001, Guang Chen 0001, Changjun Jiang 0002
CVPR5
2024 Parameter-Efficient Multi-Task Model Fusion with Partial Linearization
abstract
Large pre-trained models have enabled significant advances in machine learning and served as foundation components. Model fusion methods, such as task arithmetic, have been proven to be powerful and scalable to incorporate fine-tuned weights from different tasks into a multi-task model. However, efficiently fine-tuning large pre-trained models on multiple downstream tasks remains challenging, leading to inefficient multi-task model fusion. In this work, we propose a novel method to improve multi-task fusion for parameter-efficient fine-tuning techniques like LoRA fine-tuning. Specifically, our approach partially linearizes only the adapter modules and applies task arithmetic over the linearized adapters. This allows us to leverage the the advantages of model fusion over linearized fine-tuning, while still performing fine-tuning and inference efficiently. We demonstrate that our partial linearization technique enables a more effective fusion of multiple tasks into a single model, outperforming standard adapter tuning and task arithmetic alone. Experimental results demonstrate the capabilities of our proposed partial linearization technique to effectively construct unified multi-task models via the fusion of fine-tuned task vectors. We evaluate performance over an increasing number of tasks and find that our approach outperforms standard parameter-efficient fine-tuning techniques. The results highlight the benefits of partial linearization for scalable and efficient multi-task model fusion.
Anke Tang, Li Shen 0008, Yong Luo 0002, Yibing Zhan, Han Hu 0003, Bo Du 0001, Yixin Chen 0001, Dacheng Tao
ICLR5
2024 Collaborative Edge Caching in LEO Satellites Networks: A MAPPO Based Approach
abstract
Low Earth Orbit satellite networks, as a crucial component of global low-latency internet access, are expected to carry significant user traffic in the future. Caching frequently requested content, e.g., popular videos on short-video platforms, in satellite networks can significantly alleviate traffic congestion. However, the satellite’s brief overhead passing time, which is less than ten minutes, makes it difficult for satellites to capture the content popularity distribution. And the changing relative position between satellites poses challenges for cooperation. To address the challenges, we propose a method called SEC_MAPPO for deploying cooperative edge caching in satellite networks. First, we model this novel scenario and transform it into a Partially Observable Markov Decision Process (POMDP). Then, we design a multi-agent reinforcement learning algorithm specifically tailored for this scenario. Trace-driven simulation using a real-world LEO satellite constellation and video request dataset demonstrated that our proposed algorithm could achieve a reduction in average video request latency ranging from 4.53% to 9.31% compared to the baseline solutions.
Mingzhou Wu, Shiqi Dai, Han Hu 0003, Zhi Wang 0001
ICME3
2024 Joint Input and Output Coordination for Class-Incremental Learning
Shuai Wang 0011, Yibing Zhan, Yong Luo 0002, Han Hu 0003, Wei Yu 0004, Yonggang Wen 0001, Dacheng Tao
IJCAI4
2024 PrimKD: Primary Modality Guided Multimodal Fusion for RGB-D Semantic Segmentation
abstract
The recent advancements in cross-modal transformers have demonstrated their superior performance in RGB-D segmentation tasks by effectively integrating information from both RGB and depth modalities. However, existing methods often overlook the varying levels of informative content present in each modality, treating them equally and using models of the same architecture. This oversight can potentially hinder segmentation performance, especially considering that RGB images typically contain significantly more information than depth images. To address this issue, we propose PrimKD, a knowledge distillation based approach that focuses on guided multimodal fusion, with an emphasis on leveraging the primary RGB modality. In our approach, we utilize a model trained exclusively on the RGB modality as the teacher, guiding the learning process of a student model that fuses both RGB and depth modalities. To prioritize information from the primary RGB modality while leveraging the depth modality, we incorporate primary focused feature reconstruction and a selective alignment scheme. This integration enhances the overall freature fusion, resulting in improved segmentation results. We evaluate our proposed method on the NYU Depth V2 and SUN-RGBD datasets, and the experimental results demonstrate the effectiveness of PrimKD. Specifically, our approach achieves mIoU scores of 57.8 and 52.5 on these two datasets, respectively, surpassing existing counterparts by 1.5 and 0.4 mIoU. The code is available at https://github.com/xiaoshideta/PrimKD.
Zhiwei Hao 0001, Zhongyu Xiao, Yong Luo 0002, Jianyuan Guo, Jing Wang 0055, Li Shen 0008, Han Hu 0003
ACM Multimedia7
2024 WisdoM: Improving Multimodal Sentiment Analysis by Fusing Contextual World Knowledge
Liang Ding 0006, Li Shen 0008, Yong Luo 0002, Han Hu 0003, Dacheng Tao
ACM Multimedia5
2024 Multi-Granularity Hand Action Detection
abstract
Detecting hand actions in videos is crucial for understanding video content and has diverse real-world applications. Existing approaches often focus on whole-body actions or coarse-grained action categories, lacking fine-grained hand-action localization information. To fill this gap, we introduce the FHA-Kitchens (Fine-Grained Hand Actions in Kitchen Scenes) dataset, providing both coarse- and fine-grained hand action categories along with localization annotations. This dataset comprises 2,377 video clips and 30,047 frames, annotated with approximately 200k bounding boxes and 880 action categories. Evaluation of existing action detection methods on FHA-Kitchens reveals varying generalization capabilities across different granularities. To handle multi-granularity in hand actions, we propose MG-HAD, an End-to-End Multi-Granularity Hand Action Detection method. It incorporates two new designs: Multi-dimensional Action Queries and Coarse-Fine Contrastive Denoising. Extensive experiments demonstrate MG-HAD's effectiveness for multi-granularity hand action detection, highlighting the significance of FHA-Kitchens for future research and real-world applications. The dataset and source code are available at MG-HAD.
Ting Zhe, Jing Zhang 0037, Yongqian Li, Yong Luo 0002, Han Hu 0003, Dacheng Tao
ACM Multimedia5
2024 Joint Power Control and Data Size Selection for Over-the-Air Computation-Aided Federated Learning
abstract
Federated learning (FL) has emerged as an appealing machine learning approach to deal with massive raw data generated at multiple mobile devices, which needs to aggregate the training parameter of every mobile device at one base station (BS) iteratively. For parameter aggregating in FL, over-the-air computation is a spectrum-efficient solution, which allows all mobile devices to transmit their parameter-mapped signals concurrently to a BS. Due to heterogeneous channel fading and noise, there exists difference between the BS’s received signal and its desired signal, measured as the mean-squared error (MSE). To minimize the MSE, we propose to jointly optimize the signal amplification factors at the BS and the mobile devices as well as the data size (the number of data samples involved in local training) at every mobile device. The formulated problem is difficult to address due to its nonconvexity. To find the optimal solution, we perform cost function simplification and variable transformation, and solve the transformed problem in a two-level structure. Optimal solution of the lower level problem is found by analyzing every candidate solution from the Karush–Kuhn–Tucker (KKT) condition. Optimal solution of the upper level problem is found by exploring its piecewise convexity. Numerical results show that our proposed method can greatly reduce the MSE and can help to enhance the training performance of FL compared with benchmark methods.
Xuming An 0001, Rongfei Fan, Shiyuan Zuo, Han Hu 0003, Hai Jiang 0001, Ning Zhang 0007
IEEE Internet Things J.4
2024 Carrier Frequency Offset in Internet of Things Radio Frequency Fingerprint Identification: An Experimental Review
abstract
Radio frequency fingerprint (RFF) identification has become a promising security solution for resource-constrained Internet-of-Things (IoT) devices, which relies on hardware impairments-induced radio frequency features for identification; among which, a hotspot feature is the carrier frequency offset (CFO). Existing research, however, advocates contradictory perspectives on the usage of CFO: For identification and for compensation; the former employs CFO in the feature space while the latter eliminates the CFO from the feature space, both for improving the RFF identification accuracy. In this review, we first discuss the RFF identification procedures and investigate the origination of the CFO and further its relationship with the clock skew of the crystal oscillator. We then provide a review of the state-of-the-art RFF identification schemes, in two categories respectively employing CFO for identification and compensation. Finally, on a real testbed, we experimentally investigate the impact of the usage of CFO on RFF identification accuracy. Experimental results reveal that, the stabilities of CFOs are quite different on hardware platforms from different manufacturers; CFOs can be used for identification when they are relatively distinguishable; compensating CFO alone is inadequate for long-term identification.
Xintao Huan, Kaitao Miao, Hanxiang He, Han Hu 0003
IEEE Internet Things J.5
2024 Kerra: An Internet of Things Wireless Key Generation Resistant to Replay Attacks
abstract
Wireless key generation is a promising security solution for Internet-of-Things (IoT) networks to share identical secret keys between communication pairs, whose foundation is wireless channel randomness and reciprocity. Its security, however, affects not only the generated keys but more importantly, the security of the IoT networks. So far, a number of major attacks threatening the wireless key generation have been studied in the literature, but not the replay attack. In this paper, we reveal the replay attack can penetrate conventional defense measures and invade the wireless key generation through both analysis and experiments, which can deteriorate the channel measurement correlation and result in a high key disagreement rate. We propose a wireless key generation approach named Kerra where we integrate a synchronized time measurement to defend against the replay attack on it. On a real IoT testbed composed of Long Range (LoRa) nodes, we implement the proposed Kerra and evaluate it in terms of both key generation performance and replay attack defense. Experimental results demonstrate first the impact of the replay attack on both channel measurement correlation and key disagreement rate, then the effects of quantization and pre-processing on key disagreement rate under replay attacks, and finally, the effectiveness of the proposed Kerra whose key disagreement rates under replay attacks are maintained to a similar level as without attacks.
Xintao Huan, Kaitao Miao, Pengyi Jia, Han Hu 0003
IEEE Internet Things J.5
2024 Dynamic Routing for Integrated Satellite-Terrestrial Networks: A Constrained Multi-Agent Reinforcement Learning Approach
abstract
The integrated satellite-terrestrial network (ISTN) system has experienced significant growth, offering seamless communication services in remote areas with limited terrestrial infrastructure. However, designing a routing scheme for ISTN is exceedingly difficult, primarily due to the heightened complexity resulting from the inclusion of additional ground stations, along with the requirement to satisfy various constraints related to satellite service quality. To address these challenges, we study packet routing with ground stations and satellites working jointly to transmit packets, while prioritizing fast communication and meeting energy efficiency and packet loss requirements. Specifically, we formulate the problem of packet routing with constraints as a max-min problem using the Lagrange method. Then we propose a novel constrained Multi-Agent reinforcement learning (MARL) dynamic routing algorithm named CMADR, which efficiently balances objective improvement and constraint satisfaction during the updating of policy and Lagrange multipliers. Finally, we conduct extensive experiments and an ablation study using the OneWeb and Telesat mega-constellations. Results demonstrate that CMADR reduces the packet delay by a minimum of 21% and 15%, while meeting stringent energy consumption and packet loss rate constraints, outperforming several baseline algorithms.
Yifeng Lyu, Han Hu 0003, Rongfei Fan, Zhi Liu 0002, Jianping An, Shiwen Mao
IEEE J. Sel. Areas Commun.2
2024 SGDA: Towards 3-D Universal Pulmonary Nodule Detection via Slice Grouped Domain Attention
abstract
Lung cancer is the leading cause of cancer death worldwide. The best solution for lung cancer is to diagnose the pulmonary nodules in the early stage, which is usually accomplished with the aid of thoracic computed tomography (CT). As deep learning thrives, convolutional neural networks (CNNs) have been introduced into pulmonary nodule detection to help doctors in this labor-intensive task and demonstrated to be very effective. However, the current pulmonary nodule detection methods are usually domain-specific, and cannot satisfy the requirement of working in diverse real-world scenarios. To address this issue, we propose a slice grouped domain attention (SGDA) module to enhance the generalization capability of the pulmonary nodule detection networks. This attention module works in the axial, coronal, and sagittal directions. In each direction, we divide the input feature into groups, and for each group, we utilize a universal adapter bank to capture the feature subspaces of the domains spanned by all pulmonary nodule datasets. Then the bank outputs are combined from the perspective of domain to modulate the input group. Extensive experiments demonstrate that SGDA enables substantially better multi-domain pulmonary nodule detection performance compared with the state-of-the-art multi-domain learning methods.
Rui Xu 0031, Zhi Liu 0002, Yong Luo 0002, Han Hu 0003, Li Shen 0008, Bo Du 0001, Kaiming Kuang, Jiancheng Yang
IEEE Trans. Comput. Biol. Bioinform.4
2024 Age of Information Minimization for Opportunistic Channel Access
abstract
This paper investigates how to suppress the Age of Information (AoI) in an opportunistic channel access system, which allows multiple mobile devices to access a base station without central coordination while being aware of instant channel quality. An optimization problem is formulated to minimize the average AoI by optimizing each mobile device’s probability of contending for channel access opportunity and the threshold of offload rate. We derive the exact expression of the average AoI and generate a reformulated optimization problem. Although being non-convex, the reformulated problem is tackled by the following operations. First, we leverage the Dinkelbach method and the block coordinate descent method to convert the reformulated problem into an iterative solving procedure of two non-convex sub-problems, which optimize the contending probability and the threshold of offload rate respectively. Second, for each non-convex sub-problem, we explore the piecewise differential monotonicity for the cost function, and achieve the associated optimal solution by transforming them into standard monotonic optimization problems. Numerical results can verify the effectiveness of the proposed method through the comparison with benchmark methods.
Rongfei Fan, Han Hu 0003, Gongpu Wang, Julian Cheng 0001
IEEE Trans. Commun.3
2024 DeViT: Decomposing Vision Transformers for Collaborative Inference in Edge Devices
abstract
Recent years have witnessed the great success of vision transformer (ViT), which has achieved state-of-the-art performance on multiple computer vision benchmarks. However, ViT models suffer from vast amounts of parameters and high computation cost, leading to difficult deployment on resource-constrained edge devices. Existing solutions mostly compress ViT models to a compact model but still cannot achieve real-time inference. To tackle this issue, we propose to explore the divisibility of transformer structure, and decompose the large ViT into multiple small models for collaborative inference at edge devices. Our objective is to achieve fast and energy-efficient collaborative inference while maintaining comparable accuracy compared with large ViTs. To this end, we first propose a collaborative inference framework termedDeViTto facilitate edge deployment by decomposing large ViTs. Subsequently, we design a decomposition-and-ensemble algorithm based on knowledge distillation, termed DEKD, to fuse multiple small decomposed models while dramatically reducing communication overheads, and handle heterogeneous models by developing a feature matching module to promote the imitations of decomposed models from the large ViT. Extensive experiments for three representative ViT backbones on four widely-used datasets demonstrate our method achieves efficient collaborative inference for ViTs and outperforms existing lightweight ViTs, striking a good trade-off between efficiency and accuracy. For example, our DeViTs improves end-to-end latency by 2.89× with only 1.65% accuracy sacrifice using CIFAR-100 compared to the large ViT, ViT-L/16, on the GPU server. DeDeiTs surpasses the recent efficient ViT, MobileViT-S, by 3.54% in accuracy on ImageNet-1 K, while running 1.72× faster and requiring 55.28% lower energy consumption on the edge device.
Guanyu Xu, Zhiwei Hao 0001, Yong Luo 0002, Han Hu 0003, Jianping An, Shiwen Mao
IEEE Trans. Mob. Comput.4
2024 Interference-Aware Online Optimization for Cellular-Connected Multiple UAV Networks With Energy Constraints
abstract
The incorporation of Unmanned Aerial Vehicles (UAVs) into cellular networks opens up new possibilities to enhance their ubiquitous operations and establish superior performance owing to the high probability of line-of-sight (LoS) for air-to-ground channels. However, this also results in the UAV inducing more significant uplink interference to non-associated Base Stations (BSs). This paper explores the online design policy in cellular-connected multiple UAV communications in the absence of channel conditions, focusing on wireless resource allocation and dynamic three-dimensional (3-D) path planning. Our objective is to maximize the minimum uplink throughput for all UAVs while considering the energy constraints of the UAVs. First, we implement an online design utilizing the achievable rate based on the estimated instantaneous channel state information (CSI) for the current time slot, and the expected data rate for future time slots based on channel distribution information (CDI). Our solution employs the exact penalty method along with alternating optimization and successive convex optimization methods. Second, we formulate an online design by merely using the achievable rate based on the estimated instantaneous CSI for the current time slot. We introduce an energy-triggered penalty term to regulate the energy consumption of the UAVs, resulting in a low-complexity solution even if the CDI is unavailable before the flight. Lastly, we conduct extensive simulations to corroborate our findings and provide comprehensive comparisons with other baseline schemes to underline the effectiveness of the proposed designs.
Cheng Zhan, Han Hu 0003, Zhi Liu 0002, Jing Wang 0055, Rongfei Fan
IEEE Trans. Mob. Comput.2
2024 Tradeoff Between Age of Information and Operation Time for UAV Sensing Over Multi-Cell Cellular Networks
abstract
Unmanned aerial vehicles (UAVs) have a significant potential for sensing applications in further cellular networks due to their extensive coverage and flexible deployment. In this paper, we consider a multi-cell cellular network with a cellular-connected UAV, which senses data with onboard sensors and uploads sensory data to the ground base stations (BSs). To evaluate the freshness of sensory data, we employ the concept of age of information (AoI), which is defined as the time elapsed since the latest successful transmission of sensory data. A lower AoI implies fresher sensory data, which may lead to the increase of UAV operation time. To balance such tradeoff, we aim to minimize the weighted sum of operation time and total AoI for the UAV by jointly optimizing transmission scheduling, BS association, as well as UAV trajectory. The problem is formulated as a mixed-integer nonlinear programming (MINLP) problem, which is difficult to solve due to the time-varying propagation channels. To this end, we first characterize the average communication performance with statistic channel information, and then develop a search algorithm to obtain the optimal solution via employing the optimal structure as well as convex optimization techniques, while a low-complexity Double Graph based Algorithm (DGA) is developed to obtain a suboptimal solution. Then, by taking into account the site-specific performance and making fast decisions online, we propose a Deep reinforcement Learning Algorithm (DLA). Compared to DGA, DLA can adapt to the specific local environment and obtain a solution more rapidly once the training process is completed. Simulation results show that the proposed algorithms outperform the benchmarks about 30%, and achieve flexible tradeoff between operation time and AoI of UAV sensing, which is not available by considering just one objective.
Cheng Zhan, Han Hu 0003, Jing Wang 0055, Zhi Liu 0002, Shiwen Mao
IEEE Trans. Mob. Comput.2
2024 Textual Enhanced Adaptive Meta-Fusion for Few-Shot Visual Recognition
abstract
Few-shot learning (FSL) is a challenging task that aims to train a classifier to recognize novel categories, where only a few annotated examples are available in each category. Recently, many FSL approaches have been proposed based on the meta-learning paradigm, which attempts to learn transferable knowledge from similar tasks by designing a meta-learner. However, most of these approaches only exploit the information from visual modality and do not utilize ones from additional modalities (e.g., textual description). Since the labeled examples in FSL are limited, increasing the information on the examples is a probable solution to improve the classification performance. This motivates us to propose a novel meta-learning method, termed textual enhanced adaptive meta-fusion FSL (TAMF-FSL), which leverages both the visual information from the visual image and semantic information from language supervision. Specifically, TAMF-FSL exploits the semantic information of textual description to improve the visual-based models. We first employ a text encoder to learn the semantic features of each visual category, and then design a modality alignment module and meta-fusion module to align and fuse the visual and semantic features for final prediction. Extensive experiments show that the proposed method outperforms many recent or competitive FSL counterparts on two popular datasets.
Mengya Han, Yibing Zhan, Yong Luo 0002, Han Hu 0003, Kehua Su, Bo Du 0001
IEEE Trans. Multim.4
2024 Not All Instances Contribute Equally: Instance-Adaptive Class Representation Learning for Few-Shot Visual Recognition
abstract
Few-shot visual recognition refers to recognize novel visual concepts from a few labeled instances. Many few-shot visual recognition methods adopt the metric-based meta-learning paradigm by comparing the query representation with class representations to predict the category of query instance. However, the current metric-based methods generally treat all instances equally and consequently often obtain biased class representation, considering not all instances are equally significant when summarizing the instance-level representations for the class-level representation. For example, some instances may contain unrepresentative information, such as too much background and information of unrelated concepts, which skew the results. To address the above issues, we propose a novel metric-based meta-learning framework termed instance-adaptive class representation learning network (ICRL-Net) for few-shot visual recognition. Specifically, we develop an adaptive instance revaluing network (AIRN) with the capability to address the biased representation issue when generating the class representation, by learning and assigning adaptive weights for different instances according to their relative significance in the support set of corresponding class. In addition, we design an improved bilinear instance representation and incorporate two novel structural losses, i.e., intraclass instance clustering loss and interclass representation distinguishing loss, to further regulate the instance revaluation process and refine the class representation. We conduct extensive experiments on four commonly adopted few-shot benchmarks: miniImageNet, tieredImageNet, CIFAR-FS, and FC100 datasets. The experimental results compared with the state-of-the-art approaches demonstrate the superiority of our ICRL-Net.
Mengya Han, Yibing Zhan, Yong Luo 0002, Bo Du 0001, Han Hu 0003, Yonggang Wen 0001, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.5
2024 Aerial Video Streaming Over 3D Cellular Networks: An Environment and Channel Knowledge Map Approach
abstract
Aerial video streaming is a promising application of unmanned aerial vehicles (UAVs), which extends video service from ground to three-dimensional (3D) airspaces. However, high data rates and smooth transmission are required along with ubiquitous and environment-aware communications. To this end, we study the quality of experience (QoE) maximization problem in this paper for aerial video streaming over 3D cellular networks in urban environments with building avoidance. Different from the typical channel model based optimization in prior works, we tackle the joint design of 3D UAV trajectory and transmission scheduling as well as playback rate adaption with an environment and channel knowledge map (ECKM) approach, which provides rich information about the location-specific channel for enabling environment-aware communications. Specifically, we first consider the scenario with perfect ECKM, and propose efficient algorithms to obtain suboptimal solutions by utilizing two graph models and the iterative parameter-enabled block coordinate descent method. For the scenario without such map information, we propose a dueling Deep Q-learning (DQL) solution with map construction such that the learning process can be facilitated for path planning. Simulation results are provided to demonstrate the improvement in QoE by the proposed solutions over baseline schemes, as well as a tradeoff between video quality and rate variation.
Cheng Zhan, Han Hu 0003, Zhi Liu 0002, Jing Wang 0055, Nan Cheng 0001, Shiwen Mao
IEEE Trans. Wirel. Commun.2
2023 FedABC: Targeting Fair Competition in Personalized Federated Learning
abstract
Federated learning aims to collaboratively train models without accessing their client's local private data. The data may be Non-IID for different clients and thus resulting in poor performance. Recently, personalized federated learning (PFL) has achieved great success in handling Non-IID data by enforcing regularization in local optimization or improving the model aggregation scheme on the server. However, most of the PFL approaches do not take into account the unfair competition issue caused by the imbalanced data distribution and lack of positive samples for some classes in each client. To address this issue, we propose a novel and generic PFL framework termed Federated Averaging via Binary Classification, dubbed FedABC. In particular, we adopt the ``one-vs-all'' training strategy in each client to alleviate the unfair competition between classes by constructing a personalized binary classification problem for each class. This may aggravate the class imbalance challenge and thus a novel personalized binary classification loss that incorporates both the under-sampling and hard sample mining strategies is designed. Extensive experiments are conducted on two popular datasets under different settings, and the results demonstrate that our FedABC can significantly outperform the existing counterparts.
Dui Wang, Li Shen 0008, Yong Luo 0002, Han Hu 0003, Kehua Su, Yonggang Wen 0001, Dacheng Tao
AAAI4
2023 RotDiff: A Hyperbolic Rotation Representation Model for Information Diffusion Prediction
abstract
The massive amounts of online user behavior data on social networks allow for the investigation of information diffusion prediction, which is essential to comprehend how information propagates among users. The main difficulty in diffusion prediction problem is to effectively model the complex social factors in social networks and diffusion cascades. However, existing methods are mainly based on Euclidean space, which cannot well preserve the underlying hierarchical structures that could better reflect the strength of user influence. Meanwhile, existing methods cannot accurately model the obvious asymmetric features of the diffusion process. To alleviate these limitations, we utilize rotation transformation in the hyperbolic to model complex diffusion patterns. The modulus of representations in the hyperbolic space could effectively describe the strength of the user's influence. Rotation transformations could represent a variety of complex asymmetric features. Further, rotation transformation could model various social factors without changing the strength of influence. In this paper, we propose a novel hyperbolic rotation representation model RotDiff for the diffusion prediction problem. Specifically, we first map each social user to a Lorentzian vector and use two groups of transformations to encode global social factors in the social graph and the diffusion graph. Then, we combine attention mechanism in the hyperbolic space with extra rotation transformations to capture local diffusion dependencies within a given cascade. Experimental results on five real-world datasets demonstrate that the proposed model RotDiff outperforms various state-of-the-art diffusion prediction models.
Hongliang Qiao, Shanshan Feng 0001, Xutao Li 0003, Huiwei Lin, Han Hu 0003, Wei Wei 0002, Yunming Ye
CIKM5
2023 AI Generated Signal for Wireless Sensing
abstract
Deep learning has significantly advanced wireless sensing technology by leveraging substantial amounts of high-quality training data. However, collecting wireless sensing data encounters diverse challenges, including unavoidable data noise, limited data scale due to significant collection overhead, and the necessity to reacquire data in new environments. Taking inspiration from the achievements of AI-generated content, this paper introduces a signal generation method that achieves data denoising, augmentation, and synthesis by disentangling distinct attributes within the signal, such as individual and environment. The approach encompasses two pivotal modules: structured signal selection and signal disentanglement generation. Structured signal selection establishes a minimal signal set with the target attributes for subsequent attribute disentanglement. Signal disentanglement generation disentangles the target attributes and reassembles them to generate novel signals. Extensive experimental results demonstrate that the proposed method can generate data that closely resembles real-world data on two wireless sensing datasets, exhibiting state-of-the-art performance. Our approach presents a robust framework for comprehending and manipulating attribute-specific information in wireless sensing.
Hanxiang He, Han Hu 0003, Xintao Huan, Heng Liu 0001, Jianping An, Shiwen Mao
GLOBECOM2
2023 QoE Maximization for Aerial Video Streaming with Multiple Cellular Connected UAVs
abstract
In this paper, we consider an aerial video streaming scenario where multiple cellular connected UAVs are employed to capture videos from different Point of Interest (PoI) areas. The videos are transmitted to the base stations (BSs) such that ground users can share the visions of the UAVs. We aim to maximize the minimum quality of experience (QoE) of all users by optimizing transmission scheduling jointly with video playback rate and UAV trajectory design, where the uplink interference as well as the trade off between video quality and video smoothness are taken into account. The formulation problem is a mixed integer nonconvex optimization problem that is difficult to solve. We address it through an inexact block coordinate descent method with overlapped blocks of variables to improve optimization flexibly. To relax binary constraints, we adopt the exact penalty method with equilibrium constraints, where the exactness of the penalty function is guaranteed. In addition, successive convex approximation method is adopted to tackle the non-convexity of the optimization problem. Simulation results indicate that the proposed scheme achieves significant performance improvement compared with the baseline schemes, and reveal the tradeoff between video quality and playback smoothness.
Cheng Zhan, Han Hu 0003, Liyue Zhu, Shubin Xu
ICME3
2023 Improving Heterogeneous Model Reuse by Density Estimation
abstract
This paper studies multiparty learning, aiming to learn a model using the private data of different participants. Model reuse is a promising solution for multiparty learning, assuming that a local model has been trained for each party. Considering the potential sample selection bias among different parties, some heterogeneous model reuse approaches have been developed. However, although pre-trained local classifiers are utilized in these approaches, the characteristics of the local data are not well exploited. This motivates us to estimate the density of local data and design an auxiliary model together with the local classifiers for reuse. To address the scenarios where some local models are not well pre-trained, we further design a multiparty cross-entropy loss for calibration. Upon existing works, we address a challenging problem of heterogeneous model reuse from a decision theory perspective and take advantage of recent advances in density estimation. Experimental results on both synthetic and benchmark data demonstrate the superiority of the proposed method.
Anke Tang, Yong Luo 0002, Han Hu 0003, Fengxiang He, Kehua Su, Bo Du 0001, Yixin Chen 0001, Dacheng Tao
IJCAI3
2023 Cross-Silo Prototypical Calibration for Federated Learning with Non-IID Data
abstract
Federated Learning aims to learn a global model on the server side that generalizes to all clients in a privacy-preserving manner, by leveraging the local models from different clients. Existing solutions focus on either regularizing the objective functions among clients or improving the aggregation mechanism for the improved model generalization capability. However, their performance is typically limited by the dataset biases, such as the heterogeneous data distributions and the missing classes. To address this issue, this paper presents a cross-silo prototypical calibration method (FedCSPC), which takes additional prototype information from the clients to learn a unified feature space on the server side. Specifically, FedCSPC first employs the Data Prototypical Modeling (DPM) module to learn data patterns via clustering to aid calibration. Subsequently, the cross-silo prototypical calibration (CSPC) module develops an augmented contrastive learning method to improve the robustness of the calibration, which can effectively project cross-source features into a consistent space while maintaining clear decision boundaries. Moreover, the CSPC module's ease of implementation and plug-and-play characteristics make it even more remarkable. Experiments were conducted on four datasets in terms of performance comparison, ablation study, in-depth analysis and case study, and the results verified that FedCSPC is capable of learning the consistent features across different data sources of the same class under the guidance of calibrated model, which leads to better performance than the state-of-the-art methods. The source codes have been released at https://github.com/qizhuang-qz/FedCSPC.
Zhuang Qi, Lei Meng 0001, Zitan Chen, Han Hu 0003, Xiangxu Meng
ACM Multimedia4
2023 Rethinking the Localization in Weakly Supervised Object Localization
abstract
Weakly supervised object localization (WSOL) is one of the most popular and challenging tasks in computer vision. This task is to localize the objects in the images given only the image-level supervision. Recently, dividing WSOL into two parts (class-agnostic object localization and object classification) has become the state-of-the-art pipeline for this task. However, existing solutions under this pipeline usually suffer from the following drawbacks: 1) they are not flexible since they can only localize one object for each image due to the adopted single-class regression (SCR) for localization; 2) the generated pseudo bounding boxes may be noisy, but the negative impact of such noise is not well addressed. To remedy these drawbacks, we first propose to replace SCR with a binary-class detector (BCD) for localizing multiple objects, where the detector is trained by discriminating the foreground and background. Then we design a weighted entropy (WE) loss using the unlabeled data to reduce the negative impact of noisy bounding boxes. Extensive experiments on the popular CUB-200-2011 and ImageNet-1K datasets demonstrate the effectiveness of our method.
Rui Xu 0031, Yong Luo 0002, Han Hu 0003, Bo Du 0001, Jialie Shen 0001, Yonggang Wen 0001
ACM Multimedia3
2023 LGViT: Dynamic Early Exiting for Accelerating Vision Transformer
abstract
Recently, the efficient deployment and acceleration of powerful vision transformers (ViTs) on resource-limited edge devices for providing multimedia services have become attractive tasks. Although early exiting is a feasible solution for accelerating inference, most works focus on convolutional neural networks (CNNs) and transformer models in natural language processing (NLP). Moreover, the direct application of early exiting methods to ViTs may result in substantial performance degradation. To tackle this challenge, we systematically investigate the efficacy of early exiting in ViTs and point out that the insufficient feature representations in shallow internal classifiers and the limited ability to capture target semantic information in deep internal classifiers restrict the performance of these methods. We then propose an early exiting framework for general ViTs termed LGViT, which incorporates heterogeneous exiting heads, namely, local perception head and global aggregation head, to achieve an efficiency-accuracy trade-off. In particular, we develop a novel two-stage training scheme, including end-to-end training and self-distillation with the backbone frozen to generate early exiting ViTs, which facilitates the fusion of global and local information extracted by the two types of heads. We conduct extensive experiments using three popular ViT backbones on three vision datasets. Results demonstrate that our LGViT can achieve competitive performance with approximately 1.8 × speed-up.
Guanyu Xu, Li Shen 0008, Han Hu 0003, Yong Luo 0002, Jialie Shen 0001
ACM Multimedia4
2023 Federated Learning with Manifold Regularization and Normalized Update Reaggregation
abstract
Federated Learning (FL) is an emerging collaborative machine learning framework where multiple clients train the global model without sharing their own datasets. In FL, the model inconsistency caused by the local data heterogeneity across clients results in the near-orthogonality of client updates, which leads to the global update norm reduction and slows down the convergence. Most previous works focus on eliminating the difference of parameters (or gradients) between the local and global models, which may fail to reflect the model inconsistency due to the complex structure of the machine learning model and the Euclidean space's limitation in meaningful geometric representations. In this paper, we propose FedMRUR by adopting the manifold model fusion scheme and a new global optimizer to alleviate the negative impacts. Concretely, FedMRUR adopts a hyperbolic graph manifold regularizer enforcing the representations of the data in the local and global models are close to each other in a low-dimensional subspace. Because the machine learning model has the graph structure, the distance in hyperbolic space can reflect the model bias better than the Euclidean distance. In this way, FedMRUR exploits the manifold structures of the representations to significantly reduce the model inconsistency. FedMRUR also aggregates the client updates norms as the global update norm, which can appropriately enlarge each client's contribution to the global update, thereby mitigating the norm reduction introduced by the near-orthogonality of client updates. Furthermore, we theoretically prove that our algorithm can achieve a linear speedup property $\mathcal{O}(\frac{1}{\sqrt{SKT}})$ for non-convex setting under partial client participation, where $S$ is the participated clients number, $K$ is the local interval and $T$ is the total number of communication rounds. Experiments demonstrate that FedMRUR can achieve a new state-of-the-art (SOTA) accuracy with less communication.
Xuming An 0001, Li Shen 0008, Han Hu 0003, Yong Luo 0002
NeurIPS3
2023 Robust Task Offloading and Resource Allocation in Mobile Edge Computing With Uncertain Distribution of Computation Burden
abstract
In mobile edge computing (MEC) supporting multiple mobile users (MUs), it is essential to optimize the offloading policy and communication and computation resource allocation. A main challenge is that the computation burden of a computation task may be random and even with uncertain probabilistic distribution. To address this challenge, we investigate a multiple-MU MEC system with random computation burden. For the random computation burden of an MU, only the mean and variance are known, but its distribution is unknown. Robustness is provided such that computation outage probabilities (due to uncertain distribution of computation burden) are bounded by a predefined threshold. We minimize the weighted sum of the MUs’ energy consumption. The formulated optimization problem is non-deterministic and non-convex, and thus, is hard to solve. To deal with the challenge, we transform the formulated problem into a deterministic and convex problem by applying the Chebyshev-Cantelli inequality and some mathematical manipulations. We further decompose the convex problem to a lower-level and an upper-level problem. Low-complexity algorithms are developed for the lower-level and upper-level problems. The overall complexity of our proposed method is linear with the number of MUs.
Rongfei Fan, Bizheng Liang, Shiyuan Zuo, Han Hu 0003, Hai Jiang 0001, Ning Zhang 0007
IEEE Trans. Commun.4
2023 A One-Way Time Synchronization Scheme for Practical Energy-Efficient LoRa Network Based on Reverse Asymmetric Framework
abstract
Long Range (LoRa) network has been thriving in the IoT era due to its long-range coverage and energy efficiency. Prevalent LoRa applications rely on time synchronization to achieve accurate data ordering and coordination among the LoRa network. The energy-efficient nature of the LoRa network where end nodes in sleep scheduling always initiate communications, however, contradicts most conventional time synchronization methods based on message exchanges. In this paper, we propose a one-way time synchronization scheme tailored for the energy-efficient LoRa network based on the reverse asymmetric framework. We first discuss the delay minimization and compensation for the reverse one-way time synchronization in the LoRa network. We then propose a time synchronization scheme consisting of two time translation methods with different computational complexities and error bounds, respectively for resource-abundant and -constrained LoRa gateway and end nodes. Experiment results on a real LoRa testbed consisting of LoRa gateway and end nodes demonstrate that the proposed scheme could achieve microsecond-level synchronization accuracy in both scenarios between the end node and gateway and between the end node and end node; the latter scenario is advocated in the recent multi-hop LoRa network research.
Xintao Huan, Han Hu 0003, Yuanqing Zheng
IEEE Trans. Commun.4
2023 Multi-Agent Collaborative Inference via DNN Decoupling: Intermediate Feature Compression and Edge Learning
abstract
Recently, deploying deep neural network (DNN) models via collaborative inference, which splits a pre-trained model into two parts and executes them on user equipment (UE) and edge server respectively, becomes attractive. However, the large intermediate feature of DNN impedes flexible decoupling, and existing approaches either focus on the single UE scenario or simply define tasks considering the required CPU cycles, but ignore the indivisibility of a single DNN layer. In this article, we study the multi-agent collaborative inference scenario, where a single edge server coordinates the inference of multiple UEs. Our goal is to achieve fast and energy-efficient inference for all UEs. To achieve this goal, we design a lightweight autoencoder-based method to compress the large intermediate feature at first. Then we define tasks according to the inference overhead of DNNs and formulate the problem as a Markov decision process (MDP). Finally, we propose a multi-agent hybrid proximal policy optimization (MAHPPO) algorithm to solve the optimization problem with a hybrid action space. We conduct extensive experiments with different types of networks, and the results show that our method can reduce up to 56% of inference latency and save up to 72% of energy consumption.
Zhiwei Hao 0001, Guanyu Xu, Yong Luo 0002, Han Hu 0003, Jianping An, Shiwen Mao
IEEE Trans. Mob. Comput.4
2023 Optimal Volumetric Video Streaming With Hybrid Saliency Based Tiling
abstract
Volumetric video enables a six-degree-of-freedom (6DoF) immersive viewing experience and has a wide range of applications in entertainment and education, among others. Most existing approaches to volumetric video streaming are extensions of VR video streaming solutions that do not take into account user behavior and the properties of the video during the tiling process, and the complexity of decoding is high. To this end, we study volumetric video streaming in this paper and address the research questions mentioned above. In particular, we first propose a hybrid visual saliency and hierarchical clustering empowered 3D tiling scheme that better matches the user’s field of view (FoV). Then, we build a quality of experience (QoE) model considering the volumetric video features as the optimization objective. In addition to the usual encoded version, we introduce the reconstructed version (i.e., decoded version, which allows the user to skip the decoding process and thus reduces the decoding overhead) and propose a joint computational and communication resource allocation scheme to achieve a trade-off between communication and computational resources to maximize the QoE. We perform exhaustive simulations and build a prototype system to verify the performance of the proposed tiling and transmission scheme. The results show that the proposed tiling and transmission scheme performs significantly better than the comparison schemes.
Jie Li 0015, Cong Zhang 0002, Zhi Liu 0002, Richang Hong, Han Hu 0003
IEEE Trans. Multim.5
2023 Two-Stream Prototype Learning Network for Few-Shot Face Recognition Under Occlusions
abstract
Few-shot face recognition under occlusion (FSFRO) aims to recognize novel subjects given only a few, probably occluded face images, and it is challenging and common in real-world scenarios. Unknown occlusions may deteriorate the class prototypes, while an occluded image in the support set may be critical for recognition if the query image is occluded. This motivates us to propose a novel Two-stream Prototype Learning Network (TSPLN) for FSFR under occlusions by simultaneously considering the quality of support images and their relevance to the query i mage. Specifically, we design a two-stream architecture, which mainly consists of a support-centered stream and query-centered stream, to learn the optimal class prototypes. The former stream is to reduce the negative impact of occluded images on the prototype. This is achieved by exploring the similarities between different images in the support set. In the query-centered stream, we exploit the relevance between the query and support set based on feature alignment (FA). We conduct extensive experiments on two popular datasets: CASIA-WebFace and RMFRD. The experimental results show that our proposed method achieves the state-of-the-art performance for occluded face recognition in the few-shot setting.
Mengya Han, Yong Luo 0002, Han Hu 0003, Yonggang Wen 0001
IEEE Trans. Multim.4
2023 Optimizing Energy Efficiency for Data Center via Parameterized Deep Reinforcement Learning
abstract
The rapid advancements in cloud computing, Big Data and their related applications have led to a skyrocketing increase in data center energy consumption year by year. The prior approaches for improving data center energy efficiency mostly suffer from high system dynamics or the complexity of data centers. In this paper, we propose an optimization framework based on deep reinforcement learning, named DeepEE, to jointly optimize energy consumption from the perspectives of task scheduling and cooling control. In DeepEE, a PArameterized action space based Deep Q-Network (PADQN) algorithm is proposed to tackle the hybrid action space problem. Then, a dynamic time factor mechanism for adjusting cooling control interval is introduced into PADQN (PADQN-D) to achieve more accurate and efficient coordination of IT and cooling subsystems. Finally, in order to train and evaluate the proposed algorithms safely and quickly, a simulation platform is built to model the dynamics of IT and cooling subsystems. Extensive real-trace based experiments illustrate that: 1) the proposed PADQN algorithm can save up to 15% and 10% energy consumption compared with the baseline siloed and joint optimization approaches respectively; 2) the proposed PADQN-D algorithm with dynamic cooling control interval can better adapt to the change of IT workload; 3) our proposed algorithms achieve more stable performance gain in terms of power consumption by adopting the parameterized action space.
Yongyi Ran, Han Hu 0003, Yonggang Wen 0001, Xin Zhou 0003
IEEE Trans. Serv. Comput.2
2023 Optimizing Data Center Energy Efficiency via Event-Driven Deep Reinforcement Learning
abstract
To reduce the skyrocketing energy consumption of data centers, the prevailing approaches adopt the time-driven manner to control IT and cooling subsystems. These methods suffer from highly dynamic system states, complex action spaces and the risk of instability caused by frequent and unnecessary control operations. To tackle these problems, we propose a novel event-driven control paradigm and an optimization algorithm, under the deep reinforcement learning (DRL) framework. The principle is to make decisions based on certain critical events (e.g., overheating), rather than fixed periodic control. Specifically, we design an event-driven optimization framework to trigger control operations. Then, we present several models to describe IT and cooling subsystems, and mathematically define events to capture four types of prior factors that impact system performance. Furthermore, we develop an event-driven DRL (E-DRL) optimization algorithm to dispatch jobs and regulate cooling facilities for energy efficiency. Using two different types of real workload traces, we conduct extensive experiments to demonstrate that: 1) E-DRL reduces the number of regulating decisions by 70%$\sim$95% while achieving a comparable or even better energy efficiency in comparison with the state-of-the-art algorithm; and 2) E-DRL can adapt the control frequency to the changing operational conditions and diverse workloads.
Yongyi Ran, Xin Zhou 0003, Han Hu 0003, Yonggang Wen 0001
IEEE Trans. Serv. Comput.3
2022 Optimal Task Offloading for Deep Neural Network Driven Application in Space-Air-Ground Integrated Network
abstract
Running intelligent applications on a satellite is in urgent need, which can help to extract useful information from massive surveillance or remote sensing data and return it to ground in time. However, the limited computing ability on a satellite prohibits it from completing the whole application by itself quickly. Within the circumstance of space-air-ground integrated network (SAGIN), we propose to offload part of the computation task from the satellite to the ground station with strong computing ability, through the introduction of airship, which can assist the satellite not only by relaying but also in computing. To save the energy consumption of the satellite and airship, task offloading policy and resource allocation, are investigated for a special task model supporting deep neural network (DNN), which is popular in intelligent application. An optimization problem is formulated, which is difficult to solve. We achieve the global optimal solution through the following operations: 1) Transform the formulated problem into two levels, with every level dealing with discrete or continuous variables exclusively; 2) Explore implicit monotonicity and convexity of concerned functions so as to solve the non-convex lower level problem optimally only with several rounds of bisection or Golden search methods; 3) Solve the upper level problem optimally by enumeration but with polynomial complexity. Numerical results verify the effectiveness of our proposed method.
Rongfei Fan, Xiang Li 0024, Zhi Liu 0002, Cheng Zhan, Han Hu 0003
HPSR5
2022 Energy-Efficient Trajectory Optimization for Aerial Video Surveillance under QoS Constraints
abstract
Surveillance drones are unmanned aerial vehicles (UAVs) that are utilized to collect video recordings of targets. In this paper, we propose a novel design framework for aerial video surveillance in urban areas, where a cellular-connected UAV captures and transmits videos to the cellular network that services users. Fundamental challenges arise due to the limited onboard energy and quality of service (QoS) requirements over environment-dependent air-to-ground cellular links, where UAVs are usually served by the sidelobes of base stations (BSs). We aim to minimize the energy consumption of the UAV by jointly optimizing the mission completion time and UAV trajectory as well as transmission scheduling and association, subject to QoS constraints. The problem is formulated as a mixed-integer nonlinear programming (MINLP) problem by taking into account building blockage and BS antenna patterns. We first consider the average performance for uncertain local environments, and obtain an efficient sub-optimal solution by employing graph theory and convex optimization techniques. Next, we investigate the site-specific performance for specific urban local environments. By reformulating the problem as a Markov decision process (MDP), a deep reinforcement learning (DRL) algorithm is proposed by employing a dueling deep Q-network (DQN) neural network model with only local observations of sampled rate measurements. Simulation results show that the proposed solutions achieve significant performance gains over baseline schemes.
Cheng Zhan, Han Hu 0003, Shiwen Mao, Jing Wang 0055
INFOCOM2
2022 Leveraging GAN Priors for Few-Shot Part Segmentation
abstract
Few-shot part segmentation aims to separate different parts of an object given only a few annotated samples. Due to the challenge of limited data, existing works mainly focus on learning classifiers over pre-trained features, failing to learn task-specific features for part segmentation. In this paper, we propose to learn task-specific features in a "pre-training"-"fine-tuning" paradigm. We conduct prompt designing to reduce the gap between the pre-train task (i.e., image generation) and the downstream task (i.e., part segmentation), so that the GAN priors for generation can be leveraged for segmentation. This is achieved by projecting part segmentation maps into the RGB space and conducting interpolation between RGB segmentation maps and original images. Specifically, we design a fine-tuning strategy to progressively tune an image generator into a segmentation generator, where the supervision of the generator varying from images to segmentation maps by interpolation. Moreover, we propose a two-stream architecture, i.e., a segmentation stream to generate task-specific features, and an image stream to provide spatial constraints. The image stream can be regarded as a self-supervised auto-encoder, and this enables our model to benefit from large-scale support images. Overall, this work is an attempt to explore the internal relevance between generation tasks and perception tasks by prompt designing. Extensive experiments show that our model can achieve state-of-the-art performance on several part segmentation datasets.
Mengya Han, Heliang Zheng, Yong Luo 0002, Han Hu 0003, Bo Du 0001
ACM Multimedia5
2022 Robust Metric Boosts Transfer
abstract
Transfer metric learning (TML) aims to improve the metric learning in target domains by transferring knowledge from related tasks, where the distance metrics are strong and reliable. Existing TML approaches only focus on how to transfer the source metric knowledge, which is often prone to be over-fitting to the source domain. In this paper, we study how to train a source metric that is appropriate for transfer and then design a general deep TML method for effective metric transfer. In particular, we propose to learn the source metric parameterized by a deep neural network in an adversarial way and then transfer the metric to the target domain by embedding imitation, which allows the inputs of source and target domains to be heterogeneous. Besides, we restrict the size of the target metric network to be small so that the inference is efficient in the target domain. Results in the popular face verification application demonstrate the effectiveness of our method.
Qiancheng Yang, Yong Luo 0002, Han Hu 0003, Xin Zhou 0003, Bo Du 0001, Dacheng Tao
MMSP3
2022 Optimizing Data Centre Energy Efficiency via Event Driven Deep Reinforcement Learning
abstract
[J1C2 Presentation Abstract at IEEE SERVICES 2022 for IEEE Transactions on Services Computing DOI 10.1109/TSC.2022.3157145]
Yongyi Ran, Xin Zhou 0003, Han Hu 0003, Yonggang Wen 0001
SERVICES3
2022 Joint Task Offloading and Resource Allocation for IoT Edge Computing With Sequential Task Dependency
abstract
Incorporating mobile-edge computing (MEC) in the Internet of Things (IoT) enables resource-limited IoT devices to offload their computation tasks to a nearby edge server. In this article, we investigate an IoT system assisted by the MEC technique with its computation task subjected to sequential task dependency, which is critical for video stream processing and other intelligent applications. To minimize energy consumption per IoT device while limiting task processing delay, task offloading strategy, communication resource, and computation resource are optimized jointly under both slow and fast-fading channels. In slow fading channels, an optimization problem is formulated, which is nonconvex and involves one integer variable. To solve this challenging problem, we decompose it as a 1-D search of task offloading decision problem and a nonconvex optimization problem with task offloading decision given. Through mathematical manipulations, the nonconvex problem is transformed to be a convex one, which is shown to be solvable only with the simple Golden search method. In fast-fading channels, optimal online policies depending on the instant channel state are derived even though they are entangled. In addition, it is proved that the derived policy will converge to the offline policy when the channel coherence time is low, which can help save extra computation complexity. Numerical results verify the correctness of our analysis and the effectiveness of our proposed strategies over the existing methods.
Xuming An 0001, Rongfei Fan, Han Hu 0003, Ning Zhang 0007, Saman Atapattu, Theodoros A. Tsiftsis
IEEE Internet Things J.3
2022 Joint Task Offloading and Resource Allocation for Cooperative Mobile-Edge Computing Under Sequential Task Dependency
abstract
The emergence of mobile-edge computing (MEC) makes it possible to run intelligent applications on Internet of Things (IoT) devices. However, due to blockage or deep fading, one IoT device may not have direct link with the edge server. In this case, many surrounding wireless devices can serve as a cooperative node. In this article, we study a cooperative MEC system running sequential task, which is composed of a series of subtasks and can support many intelligent applications. To minimize the energy consumption of the IoT device and cooperative node, a task offloading policy together with the allocation of communication and computation resources is designed jointly. The cases when the cooperative node has no/has private task to complete are investigated, which are denoted as cases I and II, respectively. Although both cases involve the optimization of integer variables, their optimal solutions are achieved. For the first case, the associated problem is simplified equivalently and then decomposed into two levels, with the upper level dealing with integer variables and the lower level handling continuous variables. Bisection search is employed to reach optimality in the lower level and the searching space is compressed in the upper level. For the second case, the associated problem is subdivided into three subproblems. To solve every subproblem optimally, a similar operation like case I is followed, with a semiclosed form solution derived in the lower level. Numerical results verify the effectiveness of our proposed methods compared with benchmark methods and our effort on reducing computation complexity.
Xiang Li 0024, Rongfei Fan, Han Hu 0003, Ning Zhang 0007
IEEE Internet Things J.3
2022 Energy-Efficient Resource Allocation for Mobile Edge Computing With Multiple Relays
Xiang Li 0024, Rongfei Fan, Han Hu 0003, Ning Zhang 0007, Xianfu Chen, Anqi Meng
IEEE Internet Things J.3
2022 Editorial: Heterogeneous Cloud-Based Intelligent Computing for Next-Generation 5G Applications
Qiang Liu 0004, Ryan Shea, Zhi Liu 0002, Zehua Wang 0001, Han Hu 0003
Mob. Networks Appl.5
2022 A Timestamp-Free Time Synchronization Scheme Based on Reverse Asymmetric Framework for Practical Resource-Constrained Wireless Sensor Networks
abstract
Energy-efficient time synchronizations for wireless sensor networks (WSNs) have been put under the spotlight for years. A promising technique among which is the timestamp-free approach where no timestamps are required to establish the synchronization, thereby sparing the transmissions of the timing messages for conserving significant transmission energy. In this paper, we first investigate the feasibility of adopting timestamp-free time synchronization in practical resource-constrained WSNs; we then identify the issue of inaccuracy in maintaining the pre-defined response interval which affects the foundations of most existing timestamp-free schemes. Based on the investigation and our previously proposed reverse asymmetric time synchronization framework, we further propose an asymmetric timestamp-free time synchronization scheme with two estimation methods tailored for resource-constrained WSNs. We as well introduce the centralized and distributed multi-hop extension methods for the proposed scheme to cover diverse multi-hop scenarios. Experimental results on a real WSN testbed consisting of TelosB motes running TinyOS demonstrate that the proposed scheme achieves high energy efficiency while maintaining microsecond-level time synchronization accuracy compared to three other conventional schemes.
Xintao Huan, Hanxiang He, Qigang Wu, Han Hu 0003
IEEE Trans. Commun.5
2022 CDFKD-MFS: Collaborative Data-Free Knowledge Distillation via Multi-Level Feature Sharing
abstract
Recently, the compression and deployment of powerful deep neural networks (DNNs) on resource-limited edge devices to provide intelligent services have become attractive tasks. Although knowledge distillation (KD) is a feasible solution for compression, its requirement on the original dataset raises privacy concerns. In addition, it is common to integrate multiple pretrained models to achieve satisfactory performance. How to compress multiple models into a tiny model is challenging, especially when the original data are unavailable. To tackle this challenge, we propose a framework termed collaborative data-free knowledge distillation via multi-level feature sharing (CDFKD-MFS), which consists of a multi-header student module, an asymmetric adversarial data-free KD module, and an attention-based aggregation module. In this framework, the student model equipped with a multi-level feature-sharing structure learns from multiple teacher models and is trained together with a generator in an asymmetric adversarial manner. When some real samples are available, the attention module adaptively aggregates predictions of the student headers, which can further improve performance. We conduct extensive experiments on three popular computer visual datasets. In particular, compared with the most competitive alternative, the accuracy of the proposed framework is 1.18% higher on the CIFAR-100 dataset, 1.67% higher on the Caltech-101 dataset, and 2.99% higher on the mini-ImageNet dataset.
Zhiwei Hao 0001, Yong Luo 0002, Zhi Wang 0001, Han Hu 0003, Jianping An
IEEE Trans. Multim.4
2021 Model Compression via Collaborative Data-Free Knowledge Distillation for Edge Intelligence
abstract
Model compression without the original data for fine-tuning is challenging for deploying large-size models on resource constrained edge devices. To this end, we propose a novel data-free model compression framework based on knowledge distillation (KD), where multiple teachers are utilized in a collaborative manner to enable reliable distillation. It mainly consists of three components: adversarial data generation, multi-teacher KD, and adaptive outputs aggregation. In particular, some synthesized data are generated in an adversarial manner to mimic the original data for model compression. Then a multi-header module is developed to simultaneously leverage diverse knowledge from multiple teachers. The distillation outputs are adaptively aggregated for final prediction. The experimental results demonstrate that our framework outperforms the data-free counterpart significantly (4.48% on MNIST and 2.96% on CIFAR-10). Effectiveness of different components of our method is also verified via carefully designed ablation study.
Zhiwei Hao 0001, Yong Luo 0002, Zhi Wang 0001, Han Hu 0003, Jianping An
ICME4
2021 Joint Cache Size Scaling and Replacement Adaptation for Small Content Providers
abstract
Elastic Content Delivery Networks (Elastic CDNs) have been introduced to support explosive Internet traffic growth by providing small Content Providers (CPs) with just-in-time services. Due to the diverse requirements of small CPs, they need customized adaptive caching modules to help them adjust the cached contents to maximize their long-term utility. The traditional adaptive caching module is usually a built-in service in a cloud CDN. They adaptively change cache contents using size-scaling-only methods or strategy-adaptation-only methods. A natural question is: can we jointly optimize size and strategy to achieve tradeoff and better performance for small CPs when renting services from elastic CDNs? The problem is challenging because the two decision variables could involve both discrete and categorical variables, where discrete variables have an intrinsic order while categorical variables do not. In this paper, we propose a distribution-guided reinforcement learning framework JEANA to learn the joint cache size scaling and strategy adaptation policy. We design a distribution-guided regularizer to keep the intrinsic order of discrete variables. More importantly, we prove that our algorithm has a theoretical guarantee of performance improvement. Trace-driven experimental results demonstrate our method can improve the hit ratio while reducing the rental cost.
Jiahui Ye, Zichun Li, Zhi Wang 0001, Zhuobin Zheng, Han Hu 0003, Wenwu Zhu 0001
INFOCOM5
2021 Data-Free Ensemble Knowledge Distillation for Privacy-conscious Multimedia Model Compression
abstract
Recent advances in deep learning bring impressive performance for multimedia applications. Hence, compressing and deploying these applications on resource-limited edge devices via model compression becomes attractive. Knowledge distillation (KD) is one of the most popular model compression techniques. However, most well-behaved KD approaches require the original dataset, which is usually unavailable due to privacy issues, while existing data-free KD methods perform much worse than data-required counterparts. In this paper, we analyze previous data-free KD methods from the data perspective and point out that using a single pre-trained model limits the performance of these approaches. We then propose a Data-Free Ensemble knowledge Distillation (DFED) framework, which contains a student network, a generator network, and multiple pre-trained teacher networks. During training, the student mimics behaviors of the ensemble of teachers using samples synthesized by a generator, which aims to enlarge the prediction discrepancy between the student and teachers. A moment matching loss term assists the generator training by minimizing the distance between activations of synthesized samples and real samples. We evaluate DFED on three popular image classification datasets. Results demonstrate that our method achieves significant performance improvements compared with previous works. We also design an ablation study to verify the effectiveness of each component of the proposed framework.
Zhiwei Hao 0001, Yong Luo 0002, Han Hu 0003, Jianping An, Yonggang Wen 0001
ACM Multimedia3
2021 Multi-UAV-Enabled Mobile-Edge Computing for Time-Constrained IoT Applications
abstract
Unmanned-aerial-vehicle (UAV)-enabled mobile-edge computing (MEC) has emerged as a promising paradigm to extend the coverage of computation service for Internet of Things (IoT) applications, which are usually time sensitive and computation intensive. In this article, a novel design framework is proposed for a multi-UAV-enabled MEC system, where edge servers are equipped on multiple UAVs to provide flexible computation assistance to IoT devices with hard deadlines. The aim is to maximize the number of served IoT devices through jointly optimizing UAV trajectory and service indicator as well as resource allocation and computation offloading, where the chosen IoT devices will complete their computation tasks on time under given energy budgets and co-channel interference is taken into account. We formulate the optimization problem as a mixed integer nonlinear programming (MINLP), which is challenging to solve directly. The problem is first reformulated to a more mathematically tractable form by adding a penalty term to the objective function. We then decouple the problem into two subproblems and develop an iterative algorithm by solving the two subproblems with alternating optimization and successive convex approximation techniques, where the proposed algorithm converges to a Karush–Kuhn–Tucker (KKT) solution. In addition, an efficient initialization scheme is proposed based on multiple traveling salesman problem with time windows (m-TSPTWs) method. Finally, simulation results are provided to demonstrate that the proposed joint design achieves significant performance gains over baseline schemes.
Cheng Zhan, Han Hu 0003, Zhi Liu 0002, Zhi Wang 0001, Shiwen Mao
IEEE Internet Things J.2
2021 Joint Resource Allocation and 3D Aerial Trajectory Design for Video Streaming in UAV Communication Systems
abstract
Unmanned aerial vehicles (UAVs) can be flexibly deployed to offload cellular traffic or to provide video services for emergency scenarios without infrastructure. However, the inherent resource allocation and three-dimensional (3D) aerial trajectory design have not been formally studied. In this paper, we study the joint resource allocation and 3D aerial trajectory design for dynamic adaptive streaming over HTTP (DASH)-enabled services in a UAV communication system, where a UAV is employed as a base station for multiuser video streaming. Various factors are taken into account, including video data rate, quality variation, communication outage, play interruption, etc. By adopting a video streaming utility model, two fundamental problems are formulated with different practical aims: the first problem maximizes the minimum utility for all users within a given time horizon such that max-min fairness can be provided, and the second problem minimizes the UAV operation time subject to the individual utility requirement for all users to prolong UAV endurance. To tackle the first non-convex problem, we decouple it into three sub-problems, and a three-stage iterative algorithm is proposed to obtain a suboptimal solution by solving the three sub-problems with successive convex approximation and alternating optimization techniques. An exponential search based algorithm is proposed for the second problem by utilizing the structure of the considered problem and a similar three-stage iterative algorithm. Extensive simulations are carried out to evaluate the performance, and the results show that our proposed designs significantly outperform baseline schemes. Furthermore, our results reveal new insights of UAV movement for video streaming and unveil the tradeoff between utility and quality variance.
Cheng Zhan, Han Hu 0003, Xiufeng Sui, Zhi Liu 0002, Honggang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2020 Deep Heterogeneous Multi-Task Metric Learning for Visual Recognition and Retrieval
abstract
How to estimate the distance between data instances is a fundamental problem in many artificial intelligence algorithms, and critical in diverse multimedia applications. A major challenge in the estimation is how to find an appropriate distance function when labeled data are insufficient for a certain task. Multi-task metric learning (MTML) is able to alleviate such data deficiency issue by learning distance metrics for multiple tasks together and sharing information between the different tasks. Recently, heterogeneous MTML (HMTML) has attracted much attention since it can handle multiple tasks with varied data representations. A major drawback of the current HMTML approaches is that only linear transformations are learned to connect different domains. This is suboptimal since the correlations between different domains may be very complex and highly nonlinear. To overcome this drawback, we propose a deep heterogeneous MTML (DHMTML) method, in which a nonlinear mapping is learned for each task by using a deep neural network. The correlations of different domains are exploited by sharing some parameters at the top layers of different networks. More importantly, the auto-encoder scheme and the adversarial learning mechanism are integrated and incorporated to help exploit the feature correlations in and between different tasks and the specific properties are preserved by learning additional task-specific layers together with the common layers. Experiments demonstrated that the proposed method outperforms single-task deep metric learning algorithms and other HMTML approaches consistently on several benchmark datasets.
Shikang Gan, Yong Luo 0002, Yonggang Wen 0001, Tongliang Liu, Han Hu 0003
ACM Multimedia5
2020 Semi-supervised Online Multi-Task Metric Learning for Visual Recognition and Retrieval
abstract
Distance metric learning (DML) is critial in many multimedia application tasks. However, it is hard to learn a satisfactory distance metric given only a few labeled samples for each task. In this paper, we proposed a novel semi-supervised online multi-task DML method termed SOMTML, which enables the models describing different tasks to help each other during the metric learning procedure and thus improving their respective performance. Besides, unlabeled data are leveraged to further help alleviate the data deficiency issue in different tasks by designing a novel regularization term, which also allows some prior information to be incorporated. More importantly, a quite efficient algorithm is developed to update the metrics of all tasks adaptively. The proposed SOMTML is experimentally validated in two popular visual analytic-based applications: handwriting digits recognition and face retrieval. We compared the proposed method with competitive single-task and multi-task metric learning approaches. Extensive experimental results demonstrate the effectiveness and efficiency of the proposed SOMTML.
Yangxi Li, Han Hu 0003, Jin Li 0014, Yong Luo 0002, Yonggang Wen 0001
ACM Multimedia2
2020 Look, Read and Feel: Benchmarking Ads Understanding with Multimodal Multitask Learning
abstract
Given the massive market of advertising and the sharply increasing online multimedia content (such as videos), it is now fashionable to promote advertisements (ads) together with the multimedia content. However, manually finding relevant ads to match the provided content is labor-intensive, and hence some automatic advertising techniques are developed. Since ads are usually hard to understand only according to its visual appearance due to the contained visual metaphor, some other modalities, such as the contained texts, should be exploited for understanding. To further improve user experience, it is necessary to understand both the ads' topic and sentiment. This motivates us to develop a novel deep multimodal multitask framework that integrates multiple modalities to achieve effective topic and sentiment prediction simultaneously for ads understanding. In particular, in our framework termed Deep$M^2$Ad, we first extract multimodal information from ads and learn high-level and comparable representations. The visual metaphor of the ad is decoded in an unsupervised manner. The obtained representations are then fed into the proposed hierarchical multimodal attention modules to learn task-specific representations for final prediction. A multitask loss function is also designed to jointly train both the topic and sentiment prediction models in an end-to-end manner, where bottom-layer parameters are shared to alleviate over-fitting. We conduct extensive experiments on a large-scale advertisement dataset and achieve state-of-the-art performance for both prediction tasks. The obtained results could be utilized as a benchmark for ads understanding.
Huaizheng Zhang, Yong Luo 0002, Qiming Ai, Yonggang Wen 0001, Han Hu 0003
ACM Multimedia5
2020 Transforming Device Fingerprinting for Wireless Security via Online Multitask Metric Learning
abstract
Device fingerprinting is a crucial part in the Internet of Things applications. Existing device-fingerprinting solutions either ignore the influence of the type of network traffic or separately learn a fingerprinting model for each traffic type. This often leads to suboptimal solutions, especially when training data are limited. Considering that the data distributions of different traffic types may be different but related, we propose a novel multitask learning method to learn the fingerprinting models for several traffic types simultaneously. Specifically, we first design a system for device fingerprinting using the popular k-nearest neighbor (KNN) approach. Then, a novel distance metric learning (DML) algorithm termed online multitask metric learning (OMTML) is developed to improve the distance estimation in our system. OMTML enables the models describing different traffic types to help each other during the metric learning procedure, and thus improving their respective accuracies. OMTML can also be updated adaptively, and the updating process is efficient. The experimental results show that the proposed KNN-based system outperforms the artificial neural network (ANN)-based counterpart significantly. Besides, the comparisons of our OMTML and other representative online and multitask DML approaches demonstrate both effectiveness and efficiency of the proposed metric learning method.
Yong Luo 0002, Han Hu 0003, Yonggang Wen 0001, Dacheng Tao
IEEE Internet Things J.2
2020 Completion Time and Energy Optimization in the UAV-Enabled Mobile-Edge Computing System
abstract
Completion time and energy consumption of the unmanned aerial vehicle (UAV) are two important design aspects in UAV-enabled applications. In this article, we consider a UAV-enabled mobile-edge computing (MEC) system for Internet-of-Things (IoT) computation offloading with limited or no common cloud/edge infrastructure. We study the joint design of computation offloading and resource allocation, as well as UAV trajectory for minimization of energy consumption and completion time of the UAV, subject to the IoT devices' task and energy budget constraints. We first consider the UAV energy minimization problem without predetermined completion time, a discretized nonconvex equivalent problem is obtained by using the path discretization technique. An efficient alternating optimization algorithm for the discretized problem is proposed by decoupling it into two subproblems and addressing the two subproblems with successive convex approximation (SCA)-based algorithms iteratively. Subsequently, we focus on the completion time minimization problem, which is nonconvex and challenging to solve. By using the same path discretization approximation model to reformulate problem, a similar alternating optimization algorithm is proposed. Furthermore, we study the Pareto-optimal solution that balances the tradeoff between the UAV energy and completion time. The simulation results are provided to corroborate this article and show that the proposed designs outperform the baseline schemes. Our results unveil the tradeoff between completion time and energy consumption of the UAV for the MEC system, and the proposed solution can provide the performance close to the lower bound.
Cheng Zhan, Han Hu 0003, Xiufeng Sui, Zhi Liu 0002, Dusit Niyato
IEEE Internet Things J.2
2020 Unmanned Aircraft System Aided Adaptive Video Streaming: A Joint Optimization Approach
abstract
Due to the coverage constraint of a wireless base station, mobile users suffer from the unstable network connection and poor service quality, especially for the prevalent video services. As an alternative solution, an unmanned aerial vehicle (UAV) is able to reach the cell edge and serve ground users (GUs). In this paper, we extend the UAV applications to the more challenging adaptive streaming service over fading channel. First, we decompose the system into different modules, and present mathematical models for each of them, including a trajectory model of the UAV, fading channels between the UAV and GUs, and video streaming utility. Second, we formulate the problem as a non-convex optimization problem by optimizing the UAV trajectory and transmit power allocation, jointly with transmission schedule and rate allocation for multiple users. The objective is to maximize the overall utility while guaranteeing the fairness among multiple users under the UAV energy budget and rate-outage probability constraints. Third, to tackle this problem, we first analyze the relationship between transmission rate and rate-outage probability over the fading channel, and then divide the original problem into three subproblems, which can be solved by leveraging the successive convex approximation technique. Furthermore, an overall iterative algorithm over the three subproblems is proposed to obtain a locally optimal solution by applying the block coordinate descent technique. Finally, through extensive experiments, we demonstrate that the proposed design can achieve almost 30% performance gain in terms of max-min streaming utility for all users, compared with other benchmark schemes.
Cheng Zhan, Han Hu 0003, Zhi Wang 0001, Rongfei Fan, Dusit Niyato
IEEE Trans. Multim.2
2020 DeepQoE: A Multimodal Learning Framework for Video Quality of Experience (QoE) Prediction
abstract
Recently, many models have been developed to predict video Quality of Experience (QoE), yet the applicability of these models still faces significant challenges. Firstly, many models rely on features that are unique to a specific dataset and thus lack the capability to generalize. Due to the intricate interactions among these features, a unified representation that is independent of datasets with different modalities is needed. Secondly, existing models often lack the configurability to perform both classification and regression tasks. Thirdly, the sample size of the available datasets to develop these models is often very small, and the impact of limited data on the performance of QoE models has not been adequately addressed. To address these issues, in this work we develop a novel and end-to-end framework termed as DeepQoE. The proposed framework first uses a combination of deep learning techniques, such as word embedding and 3D convolutional neural network (C3D), to extract generalized features. Next, these features are combined and fed into a neural network for representation learning. A learned representation will then serve as input for classification or regression tasks. We evaluate the performance of DeepQoE with three datasets. The results show that for small datasets (e.g., WHU-MVQoE2016 and Live-Netflix Video Database), the performance of state-of-the-art machine learning algorithms is greatly improved by using the QoE representation from DeepQoE (e.g., 35.71% to 44.82%); while for the large dataset (e.g., VideoSet), our DeepQoE framework achieves significant performance improvement in comparison to the best baseline method (90.94% vs. 82.84%). In addition to the much improved performance, DeepQoE has the flexibility to fit different datasets, to learn QoE representation, and to perform both classification and regression problems. We also develop a DeepQoE based adaptive bitrate streaming (ABR) system to verify that our framework can be easily applied to multimedia communication service. The software package of the DeepQoE framework has been released to facilitate the current research on QoE.
Huaizheng Zhang, Linsen Dong, Guanyu Gao, Han Hu 0003, Yonggang Wen 0001, Kyle Guan
IEEE Trans. Multim.4
2020 Coordinating Workload Scheduling of Geo-Distributed Data Centers and Electricity Generation of Smart Grid
abstract
With the rapidly increasing computing demand, data centers become more and more power-hungry, which incurs substantial electricity cost. Meanwhile, due to the time-dependent demand preference, power grid is suffering high load variations, which results in a large profit loss. In this paper, we consider a cost-efficient workload scheduling with a coordination between a cloud service provider operating multiple geo-distributed data centers and smart grids. The aim is to explore the flexibility of data center power demands to reduce the cost of the cloud service provider and smooth the load variations of smart grids simultaneously. We first present the penalty model of the computation workload scheduling at each data center, and introduce the cost model of smart grids, including power generation cost and the cost due to the power load variations. To jointly minimize the cost of smart grids and penalty of the cloud service provider resulted from workload scheduling, we formulate the objective function as a weighted sum of the cost and the penalty to study the tradeoffs, and obtain the optimal offline solution by the dual decomposition technique. In order to make the coordination implemented in an online fashion, we propose a Receding Horizon Control (RHC) based online algorithm to obtain the suboptimal workload management based on the predicted information, including the future amounts of interactive workload, batch workload, and power load, in the prediction horizon. The simulation results show that with the coordination between the cloud service provider and smart grids, the cost of smart grids can be significantly reduced, by up to 20 percent, and the load variations of smart grids can be well smoothed simultaneously.
Han Hu 0003, Yonggang Wen 0001, Ling Qiu 0003, Dusit Niyato
IEEE Trans. Serv. Comput.1
2019 Optimization for HTTP Adaptive Video Streaming in UAV-Enabled Relaying System
abstract
To guarantee quality of experience (QoE) of video streaming for ground users with large obstacles that deteriorate the quality of links, unmanned aerial vehicle (UAV) is introduced in this paper as a relay to serve these users for HTTP adaptive streaming. By introducing the QoE utility model for users, we study the average QoE maximization problem in UAV-enabled relaying system by optimizing the UAV position along with the bandwidth and transmit power allocation for ground users, subject to the information causality constraint at the UAV relay. The optimization problem is formulated with a non-convex programming which is difficult to solve in general. By applying the successive convex approximation and block coordinate descent techniques, an efficient iterative algorithm is proposed to simultaneously update the UAV's position, transmit power and bandwidth allocation at each iteration, where the convergence is analyzed. Extensive simulations are conducted to show that the proposed solution can achieve significant gains for QoE in terms of the average streaming utility.
Han Hu 0003, Cheng Zhan, Jianping An, Yonggang Wen 0001
ICC1
2019 DeepEE: Joint Optimization of Job Scheduling and Cooling Control for Data Center Energy Efficiency Using Deep Reinforcement Learning
abstract
The past decade witnessed the tremendous growth of power consumption in data centers due to the rapid development of cloud computing, big data analytics, and machine learning, etc. The prior approaches that optimize the power consumption of the information technology (IT) system and/or the cooling system always fail to capture the system dynamics or suffer from the complexity of system states and action spaces. In this paper, we propose a Deep Reinforcement Learning (DRL) based optimization framework, named DeepEE, to improve the energy efficiency for data centers by considering the IT and cooling systems concurrently. In DeepEE, we first propose a PArameterized action space based Deep Q-Network (PADQN) algorithm to solve the hybrid action space problem and jointly optimize the job scheduling for the IT system and the airflow rate adjustment for the cooling system. Then, a two-time-scale control mechanism is applied in PADQN to coordinate the IT and cooling systems more accurately and efficiently. In addition, to train and evaluate the proposed PADQN in a safe and quick way, we build a simulation platform to model the dynamics of IT workload and cooling systems simultaneously. Through extensive real-trace based simulations, we demonstrate that: 1) our algorithm can save up to 15% and 10% energy consumption in comparison with the baseline siloed and joint optimization approaches respectively; 2) our algorithm achieves more stable performance gain in terms of power consumption by adopting the parameterized action space; and 3) our algorithm leads to a better tradeoff between energy saving and service quality.
Yongyi Ran, Han Hu 0003, Xin Zhou 0003, Yonggang Wen 0001
ICDCS2
2019 Emotion Recognition from Physiological Signals using Multi-Hypergraph Neural Networks
abstract
Emotion recognition from physiological signals is an effective way to discern the inner state of users. Existing works are lack in the exploration of latent correlation among multiple physiological signals and relationship among different subjects. To tackle this issue, we propose to recognize emotion from physiological signals using multi-hypergraph neural networks (MHGNN). In this method, the correlation among different subjects is formulated in the multi-hypergraph structure, where each type of physiological signal is used to generate one hypergraph. In each hypergraph, the hyperedges are used to represent the connections among the vertices (subject, stimuli). Thus, the emotion recognition task is modeled as classifying each vertex in the multi-hypergraph. Experimental results and comparisons with the state-of-the-art methods in the DEAP dataset demonstrate the superior performance of our method. The comparative experiments based on available biological knowledge verify that MHGNN can depict the real biological response process in a much more precise way.
Xibin Zhao, Han Hu 0003, Yue Gao 0002
ICME3
2019 Orchestrating Caching, Transcoding and Request Routing for Adaptive Video Streaming Over ICN
abstract
Information-centric networking (ICN) has been touted as a revolutionary solution for the future of the Internet, which will be dominated by video traffic. This work investigates the challenge of distributing video content of adaptive bitrate (ABR) over ICN. In particular, we use the in-network caching capability of ICN routers to serve users; in addition, with the help of named function, we enable ICN routers to transcode videos to lower-bitrate versions to improve the cache hit ratio. Mathematically, we formulate this design challenge into a constrained optimization problem, which aims to maximize the cache hit ratio for service providers and minimize the service delay for endusers. We design a two-step iterative algorithm to find the optimum. First, given a content management scheme, we minimize the service delay via optimally configuring the routing scheme. Second, we maximize the cache hits for a given routing policy. Finally, we rigorously prove its convergence. Through extensive simulations, we verify the convergence and the performance gains over other algorithms. We also find that more resources should be allocated to ICN routers with a heavier request rate, and the routing scheme favors the shortest path to schedule more traffic.
Han Hu 0003, Yichao Jin 0002, Yonggang Wen 0001, Cédric Westphal
ACM Trans. Multim. Comput. Commun. Appl.1
2018 Deepqoe: A Unified Framework for Learning to Predict Video QoE
abstract
Motivated by the prowess of deep learning (DL) based techniques in prediction, generalization, and representation learning, we develop a novel framework called DeepQoE to predict video quality of experience (QoE). The end-to-end framework first uses a combination of DL techniques (e.g., word embeddings) to extract generalized features. Next, these features are combined and fed into a neural network for representation learning. Such representations serve as inputs for classification or regression tasks. Evaluating the performance of DeepQoE with two datasets, we show that for the small dataset, the accuracy of all shallow learning algorithms is improved by using the representation derived from DeepQoE. For the large dataset, our DeepQoE framework achieves significant performance improvement in comparison to the best baseline method (90.94% vs. 82.84%). Moreover, DeepQoE, also released as an open source tool, provides video QoE research much-needed flexibility in fitting different datasets, extracting generalized features, and learning representations.
Huaizheng Zhang, Han Hu 0003, Guanyu Gao, Yonggang Wen 0001, Kyle Guan
ICME2
2018 Optimizing Personalized Interaction Experience in Crowd-Interactive Livecast: A Cloud-Edge Approach
abstract
Enabling users to interact with broadcasters and audience, the crowd-interactive livecast greatly improves viewer's quality of experience (QoE) and attracts millions of daily active users recently. In addition to striking the balance between resource utilization and viewers' QoE met in the traditional video streaming service, this novel service needs to take supererogatory efforts to improve the interaction QoE, which reflects the viewer interaction experience. To tackle this issue, we conduct measurement studies over a large-scale dataset crawled from a representative livecast service provider. We observe that the individual's interaction pattern is quite heterogeneous: only 10% viewers proactively participate in the interaction, and the rest viewers usually watch passively. Incorporating the insight into the emerging cloud-edge architecture, we propose a framework PIECE, which optimizes the Personalized Interaction Experience with Cloud-Edge architecture (PIECE) for intelligent user access control and livecast distribution. In particular, we first devise a novel deep neural network based algorithm to predict users' interaction intensity using the historical viewer pattern. We then design an algorithm to maximize the individual's QoE, by strategically matching viewer sessions and transcoding-delivery paths over cloud-edge infrastructure. Finally, we use trace-driven experiments to verify the effectiveness of PIECE. Our results show that our prediction algorithm outperforms the state-of-the-art algorithms with a much smaller mean absolute error (40% reduction). Furthermore, in comparison with the cloud-based video delivery strategy, the proposed framework can simultaneously improve the average viewers QoE (26% improvement) and interaction QoE (21% improvement), while maintaining a high streaming bitrate.
Haitian Pang, Cong Zhang 0002, Fangxin Wang 0001, Han Hu 0003, Zhi Wang 0001, Jiangchuan Liu, Lifeng Sun
ACM Multimedia4
2018 Budget-Efficient Viral Video Distribution Over Online Social Networks: Mining Topic-Aware Influential Users
abstract
Marketing over online social networks (OSNs) has become an essential tool for spreading product information in a “word of mouth” way. In particular, campaigns normally adopt a pragmatic approach of seeding videos with a selected list of influential users, hoping to create a viral distribution to reach as many users as possible. In this paper, we propose a multitopic-aware influence maximization framework to identify a fixed number of influential users and assign video clips of specific topics to them, with an ultimate objective to maximize the number of message deliveries, defined as expected posting number (EPN). We first prove the submodularity of the EPN function, resulting in a general greedy algorithm with a performance bound of 1-1/e. We further develop two faster algorithms to accelerate the computing speed for large-scale social networks. The first algorithm leverages two estimation methods to compute the upper bound for marginal EPN without a loss of accuracy. The second algorithm generates an approximation solution based on the upper bound and lower bound estimation, with a performance bound of ε(1-1/e). We have implemented a prototype system based on a private data center at the Nanyang Technological University campus in Singapore to enable video clip extraction and sharing among social users. Furthermore, we conduct experiments on four real large-scale social networks (with different scales and structures) and the results show that the proposed methods are much faster than previous algorithms but with high accuracy.
Han Hu 0003, Yonggang Wen 0001, Shanshan Feng 0001
IEEE Trans. Circuits Syst. Video Technol.1
2018 Can We Speculate Running Application With Server Power Consumption Trace?
abstract
In this paper, we propose to detect the running applications in a server by classifying the observed power consumption series for the purpose of data center energy consumption monitoring and analysis. Time series classification problem has been extensively studied with various distance measurements developed; also recently the deep learning-based sequence models have been proved to be promising. In this paper, we propose a novel distance measurement and build a time series classification algorithm hybridizing nearest neighbor and long short term memory (LSTM) neural network. More specifically, first we propose a new distance measurement termed as local time warping (LTW), which utilizes a user-specified index set for local warping, and is designed to be noncommutative and nondynamic programming. Second, we hybridize the 1-nearest neighbor (1NN)-LTW and LSTM together. In particular, we combine the prediction probability vector of 1NN-LTW and LSTM to determine the label of the test cases. Finally, using the power consumption data from a real data center, we show that the proposed LTW can improve the classification accuracy of dynamic time warping (DTW) from about 84% to 90%. Our experimental results prove that the proposed LTW is competitive on our data set compared with existed DTW variants and its noncommutative feature is indeed beneficial. We also test a linear version of LTW and find out that it can perform similar to state-of-the-art DTW-based method while it runs as fast as the linear runtime lower bound methods like LB_Keogh for our problem. With the hybrid algorithm, for the power series classification task we achieve an accuracy up to about 93%. Our research can inspire more studies on time series distance measurement and the hybrid of the deep learning models with other traditional models.
Han Hu 0003, Yonggang Wen 0001, Jun Zhang 0003
IEEE Trans. Cybern.2
2018 Optimizing Quality of Experience for Adaptive Bitrate Streaming via Viewer Interest Inference
abstract
Rate adaptation is widely adopted in video streaming to improve the quality of experience (QoE). However, most of the existing rate adaptation approaches neglect the underlying video semantic information. In fact, influenced by video semantics and viewer preferences, the viewer may have different degrees of interest on different parts of a video. The interesting parts of a video can draw more visual attention from the viewer and have higher visual importance. As such, delivering the parts of a video that are interesting to the viewer in a higher quality can improve the perceptual video quality, compared with the semantics-agnostic approaches that treat each part of a video equally. Thus, it is natural to wonder: how to allocate bitrate budgets temporally over a video session under time-varying bandwidth while considering viewer interest? As an exploratory study, we propose an interest-aware rate adaptation approach for improving QoE by inferring viewer interest based on video semantics. We adopt the deep learning method to recognize the scenes of video frames and leverage the term frequency-inverse document frequency method to analyze the degrees of an individual viewer's interest on different types of scenes. The bandwidth, buffer occupancy, and viewer interest are jointly considered under the model predictive control framework for selecting appropriate bitrates for maximizing QoE. The objective and subjective evaluations measured in a real environment show that our method can achieve a higher QoE compared with the semantics-agnostic approaches.
Guanyu Gao, Huaizheng Zhang, Han Hu 0003, Yonggang Wen 0001, Jianfei Cai 0001, Chong Luo 0001, Wenjun Zeng 0001
IEEE Trans. Multim.3
2018 Toward Rendering-Latency Reduction for Composable Web Services via Priority-Based Object Caching
abstract
Web services serve as the cornerstone of the Internet for rendering webpages. The initial rendering latency of webpages, which depends on a subset of critical objects required by the webpage, is a key metric for web services. In this work, we propose to identify this set of critical objects systematically with the goal of caching them at a higher priority to reduce the initial rendering time. We first conduct a measurement study on a mainstream content delivery network provider, the results of which suggest that not all currently cached objects are critical and that only a small portion of the critical objects are cached. Thus, we model the critical-object aware caching scheme as a constrained optimization problem. Using the stochastic optimization framework, we decompose the problem into a set of one-shot optimization problems, which are proved to be NP-hard. We then develop two greedy algorithms with different computational complexity but the same performance bound. Finally, we integrate the resulting approximation algorithms into an online algorithm. Through trace-based simulations, we verify that our proposed algorithm can reduce service latency and network traffic by ensuring a higher cache hit ratio.
Han Hu 0003, Yonggang Wen 0001
IEEE Trans. Multim.1
2017 GECKO: Gamer Experience-Centric Bitrate Control Algorithm for Cloud Gaming
Yi-Hao Ke, Guoqiao Ye, Di Wu 0001, Yipeng Zhou, Edith C. H. Ngai, Han Hu 0003
ICIG (2)6
2017 QDLCoding: QoS-differentiated low-cost video encoding scheme for online video service
abstract
Adaptive bitrate (ABR) streaming is the de facto solution in online video services to cope with heterogeneous devices and varying network connections. However, this solution is computation intensive, demanding a large number of servers for encoding videos. Moreover, due to the time-varying nature of video generation, intelligent strategies are required in order to determine the right amount of resources for encoding. The situation is further complicated by the fact that, the two types of co-existing video content, live content and Video-on-Demand (VoD) content, have different QoS requirements for encoding. These observations posit daunting challenges for meeting the heterogeneous QoS requirements with a minimum computing capacity. This paper proposes the QoS-differentiated low-cost video encoding (QDLCoding) scheme to address these challenges. We develop a framework for scheduling the encoding workloads of the two types of videos with statistical QoS guarantees. Each type of videos is specified with a QoS criterion and a QoS loss bound. The objective is to provision the minimum amount of resources while keeping the QoS loss probabilities within the prescribed bounds. We design an online algorithm that can determine the minimum required capacity by learning content arrival distributions. The experiment results demonstrate that our method can greatly reduce the required capacity for encoding online videos while controlling the likelihood of QoS loss precisely.
Guanyu Gao, Yonggang Wen 0001, Han Hu 0003
INFOCOM3
2017 Toward Joint Compression-Transmission Optimization for Green Wearable Devices: An Energy-Delay Tradeoff
abstract
Small-size and light-weight, as the modern design concept for the emerging wearable devices, has become a trend. However, such trend puts physical limitations to the battery, and the resulting short battery lifetime becomes the bottleneck for most wearable devices today. In this paper, we aim to optimize the energy usage through data compression and transmission rate control. We propose a novel joint compression-transmission approach, which not only minimizes the energy consumption of both compression and transmission, but also maintains the corresponding data distortion and transmission delay within a certain tolerant level. By adopting the Lyapunov framework, we develop an online algorithm to minimize the one-slot drift-plus-penalty function. We conduct numerical analysis and experimental study for our proposed approach. The results show that the size of queuing buffer has the significant impact on the energy cost. Next, we verify a fundamental tradeoff between the energy expenditure and transmission delay, and derive the theoretical performance bounds. After that, we show that the energy cost is also determined by the wireless channel gain and the data compression ratio. Finally, compared to a strategy without compression, our approach can save up to 92% of energy.
Weizheng Hu, Wei Zhang 0082, Han Hu 0003, Yonggang Wen 0001, King-Jet Tseng
IEEE Internet Things J.3
2017 Public Cloud Storage-Assisted Mobile Social Video Sharing: A Supermodular Game Approach
abstract
Mobile social video sharing enables mobile users to create ultra-short video clips and instantly share them with social friends, which poses significant pressure to the content distribution infrastructure. In this paper, we propose a public cloud-assisted architecture to tackle this problem. In particular, by motivating mobile users to upload videos to the local public cloud to serve requests, and, therefore, having a permission to access friends' videos stored in the cloud, our method can alleviate the traffic burden to the social service providers, while reducing the service latency of mobile users. First, we present a general framework to model the information diffusion and utility function of each user on the proposed architecture, and formulate the problem as a decentralized social utility maximization game. Second, we show that this problem is a supermodular game and there exists at least one socially aware Nash equilibrium (SNE). We then develop two decentralized algorithms to solve this problem. The first algorithm can find an SNE with less computation complexity, and the second algorithm can find the Pareto-optimal SNE with better performance. Finally, through extensive experiments, we demonstrate that the overall system performance can be significantly improved by exploiting the selflessness among social friends.
Han Hu 0003, Yonggang Wen 0001, Dusit Niyato
IEEE J. Sel. Areas Commun.1
2017 Spectrum Allocation and Bitrate Adjustment for Mobile Social Video Sharing: Potential Game With Online QoS Learning Approach
abstract
With the recent progress on mobile networking and devices, mobile social video sharing (MSVS) has emerged as one of the most important social media services. It enables mobile users to create ultra-short video clips and instantly share them with social friends. Due to the huge volume of videos and limited available bandwidth of wireless infrastructure, it is challenging to distribute these massive videos to mobile users with satisfactory quality of service (QoS). In this paper, we present a general framework to model the video diffusion among mobile users and user QoS of the MSVS service over the wireless infrastructure. Then, we utilize the hierarchical structure to decompose this problem into two subproblems, including a bitrate adjustment and spectrum allocation problems. For the bitrate adjustment problem, we propose a QoS estimation model based on the large deviation principle. By introducing a sliding window method to derive the online estimation, we develop an online bitrate adjustment strategy without relying on any prior knowledge of neither network environment nor video traffic. For the spectrum allocation problem, we prove that such a problem is a potential game. We devise a decentralized algorithm to find the Nash equilibrium, and analyze the convergence rate and the performance gap with the centralized optimization solution. Through extensive real trace driven simulations, we demonstrate that our proposed algorithm can guarantee smooth video playback with a higher PSNR.
Han Hu 0003, Yonggang Wen 0001, Dusit Niyato
IEEE J. Sel. Areas Commun.1
2017 Cost-Optimized Microblog Distribution over Geo-Distributed Data Centers: Insights from Cross-Media Analysis
abstract
The unprecedent growth of microblog services poses significant challenges on network traffic and service latency to the underlay infrastructure (i.e., geo-distributed data centers). Furthermore, the dynamic evolution in microblog status generates a huge workload on data consistence maintenance. In this article, motivated by insights of cross-media analysis-based propagation patterns, we propose a novel cache strategy for microblog service systems to reduce the inter-data center traffic and consistence maintenance cost, while achieving low service latency. Specifically, we first present a microblog classification method, which utilizes the external knowledge from correlated domains, to categorize microblogs. Then we conduct a large-scale measurement on a representative online social network system to study the category-based propagation diversity on region and time scales. These insights illustrate social common habits on creating and consuming microblogs and further motivate our architecture design. Finally, we formulate the content cache problem as a constrained optimization problem. By jointly using the Lyapunov optimization framework and simplex gradient method, we find the optimal online control strategy. Extensive trace-driven experiments further demonstrate that our algorithm reduces the system cost by 24.5% against traditional approaches with the same service latency.
Han Hu 0003, Yonggang Wen 0001, Tat-Seng Chua, Xuelong Li 0001
ACM Trans. Intell. Syst. Technol.1
2017 Resource Provisioning and Profit Maximization for Transcoding in Clouds: A Two-Timescale Approach
abstract
Transcoding is widely adopted for content adaptation; however, it may incur excessive resource consumption and processing delays. Taking advantage of cloud infrastructure, cloud-based transcoding can elastically allocate resources under time-varying workloads and perform multiple transcodings in parallel to reduce delays. To provide transcoding as a cloud service, cloud transcoding systems require some intelligent mechanisms to provision resources and schedule tasks to satisfy user requirements while maximizing financial profit. To this end, we propose a two-timescale stochastic optimization framework for maximizing service profit while achieving performance requirements by jointly provisioning resources and scheduling tasks under a hierarchical control architecture. Our method analytically integrates service revenue, processing delay, and resource consumption in one optimization framework. We derive the offline exact solution and design some approximate online solutions for task scheduling and resource provisioning. We implement an open source cloud transcoding system, called Morph, and evaluate the performance of our method in a real environment. Empirical studies verify that our method can reduce resource consumption and achieve a higher profit compared with baseline schemes.
Guanyu Gao, Han Hu 0003, Yonggang Wen 0001, Cédric Westphal
IEEE Trans. Multim.2
2016 Towards cost-efficient workload scheduling for a Tango between geo-distributed data center and power grid
abstract
Nowadays, data centers consume substantial power, which takes up a considerable portion of local power supply (e.g., smart grid). In this paper, we leverage data center workload scheduling for the coordination between data centers and the smart grid, aiming to reduce the electricity cost of data centers and smooth the load variation of the smart grid simultaneously. We first build cost models of workload scheduling at data centers and the power generation and variation at the smart grid. We formulate the objective function as a weighted sum of the cost of the smart grid and the penalty caused by workload scheduling. Using the dual decomposition method, we then derive the optimal offline solution. To facilitate online implementation, we finally propose a Receding Horizon Control (RHC) based algorithm to obtain the suboptimal solution using limited predicted information. Extensive simulation results show that our proposed scheme can significantly reduce the cost of the smart grid, by up to 20%, while smoothing the load variation simultaneously.
Han Hu 0003, Yonggang Wen 0001, Ling Qiu 0003
ICC1
2016 Joint Content Replication and Request Routing for Social Video Distribution Over Cloud CDN: A Community Clustering Method
abstract
The increasing popularity of online social networks (OSNs) has been transforming the dissemination pattern of social video contents. We can utilize the social information propagation pattern to improve the efficiency of social video distribution. In this paper, motivated by the social community classification, we present a social video replication and user request dispatching mechanism in the cloud content delivery network architecture to reduce the system operational cost, while guaranteeing the averaged service latency. Specifically, we first present a community classification method that clusters social users with social relationships, close geolocations, and similar video watching interests into various communities. Then, we conduct a large-scale measurement on a real OSN system to study the diversities of social video propagation and the effectiveness of our communities on smoothing the diversity. Finally, we propose the community-based video replication and request dispatching strategy and formulate it as a constrained optimization problem. Based on a stochastic optimization framework, we derive an online solution and rigorously prove the optimality. We evaluate our algorithm on a real trace under realistic settings and demonstrate that our algorithm can reduce the monetary cost by 30% against traditional approaches with the same service latency.
Han Hu 0003, Yonggang Wen 0001, Tat-Seng Chua, Wenwu Zhu 0001, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.1
2015 Cost-efficient and QoS-aware content management in media cloud: Implementation and evaluation
abstract
Adaptive bitrate streaming has been proposed to encode video contents into multiple versions for device heterogeneity and changing network conditions. This solution, however, could consume enormous computing and storage resource. In fact, only a small fraction of videos are frequently requested. Thus, caching multiple versions for unpopular contents is not cost efficient. In this paper, we design a cost-efficient and QoS-aware content management system for video streaming. The system consists of a set of streaming servers and a computing cluster, where streaming servers can cache video contents or transcode them in real time, and the computing cluster can perform transcoding tasks on behalf of streaming servers. Based on this architecture, to provide cost-efficient and QoS-aware video service, first, we design a cost-efficient content cache management module to minimize the operational cost, by dynamically determining whether a segment should be cached or transcoded on fly according to their popularity. Second, to reduce transcoding latency, we design a QoS-aware transcoding task delegation module to determine whether a transcoding task in streaming server should be delegated to the computing cluster according to the streaming server's workload. We implement the system and evaluate the performance in a real environment. The results demonstrate that our method can greatly reduce the operational cost and guarantee the QoS in providing video services.
Guanyu Gao, Yonggang Wen 0001, Han Hu 0003
ICC4
2014 Community based effective social video contents placement in cloud centric CDN network
abstract
The increasing popularity of online social networks (OSNs) has been transforming the dissemination pattern of social video contents. Considering the unique features of social videos, e.g., huge volume, long-tailed, and short length, how to utilize the information propagation pattern to improve the efficiency of content distribution for social videos attracts more and more attention. In this paper, we first conduct a large scale measurement to explore the social video viewing behavior under the community classification. Based on the measurement, we investigate the community driven sharing video distribution problem under the cloud-centric content delivery network (CDN) architecture. In particular, we formulate it as a constrained optimization problem with the objective to minimize the operational cost. The constraint is the averaged transmission delay. Following that, we propose a dynamic algorithm to seek the optimal solution. Our trace-driven experiments further demonstrate our algorithm can make a better tradeoff between monetary cost and QoS, and outperforms the traditional method with less operational cost while satisfying the QoS requirement.
Han Hu 0003, Yonggang Wen 0001, Tat-Seng Chua, Zhi Wang 0001, Wenwu Zhu 0001, Di Wu 0001
ICME1
2014 Social TV analytics: a novel paradigm to transform TV watching experience
abstract
The blooming online social networks have revolutionized the way information is created, disseminated and consumed, positing significant challenges to the conventional information propagation carriers, especially for the television land-scape. In this paper, we design and develop a multi-screen cloud social TV integrated with social media via a second screen as a novel paradigm in response to this trend. Our system comprises three building blocks, including a cloud based social TV system, a social TV analytics system, and a multi-screen orchestration system. In particular, we leverage the cloud infrastructure to improve the system scalability, and design intelligent social media collection & analysis mechanisms to mine deeper social perception. Furthermore, we demonstrate two key features of our system based on a real user case.
Han Hu 0003, Yonggang Wen 0001, Chang Wen Chen, Tat-Seng Chua
MMSys1
2014 MUTAS: Multi-screen TV experience as a service through cloud centric media network
abstract
Recently, the TV landscape is rapidly shifting from the traditional “laid-back” experience to a “lean-forward” multiscreen experience. In this paper, we propose MUTAS (MUltiscreen TV experience As a Service), a novel cloud-based service delivery model, in response to this trend. The design objective is to facilitate the development process of new multi-screen features, and improve user experiences by offering an all-in-one solution. The enabling technology is to encapsulate basic functions into a unified cloud platform, and expose divergent multi-screen services through a cloud clone per user. Based on MUTAS, we will use one system to demonstrate four different multi-screen experiences (i.e., synchronized social TV watching, video teleportation, social networking integration, and advertising re-distribution). This demo provides a reference to build cloud-based frameworks for flexible, extensible, and scalable multi-screen TV experience.
Yichao Jin 0002, Han Hu 0003, Yonggang Wen 0001
SECON2
2014 Toward a biometric-aware cloud service engine for multi-screen video applications
abstract
The emergence of portable devices and online social networks (OSNs) has changed the traditional video consumption paradigm by simultaneously providing multi-screen video watching, social networking engagement, etc. One challenge is to design a unified solution to support ever-growing features while guarantee system performance. In this demo, we design and implement a multi-screen technology to provide multi-screen interactions over wide area network (WAN). Furthermore, we incorporate face-detection technology into our system to identify users' bio-features and employ a machine learning based traffic scheduling mechanism to improve the system performance.
Han Hu 0003, Yichao Jin 0002, Yonggang Wen 0001, Tat-Seng Chua, Xuelong Li 0001
SIGCOMM1
2014 Reducing Operational Costs in Cloud Social TV: An Opportunity for Cloud Cloning
abstract
The emergence of social TV has transformed TV experiences, providing a unified media experience across different devices. In response to this trend, we have implemented a multi-screen social TV system, offering video teleportation as an attractive feature. The enabling technology is instantiating a cloud clone to support all media outlets of each user. As the user shifts his attention from one device to the other, the cloud clone might migrate to a better location to reduce its operational cost. This paper investigates this cloud clone migration problem, aiming to minimize the monetary cost on operating video teleportation. Specifically, we formulate it into a Markov Decision Problem, to balance the trade-off between the migration cost and the content transmission cost. Under this framework, four algorithms are proposed to solve this optimization problem. We first characterize an upper and a lower bound for the optimal cost, by considering a random fixed placement and an offline algorithm. We then present a semi-online and a more practical Q-learning approach to make online decisions. Their performances are evaluated based on both simulated and real user traces. The results show that the Q-learning method achieves up to 25% cost compared to random fixed placement in typical scenarios. The savings are affected by the delivery path length, the migration size, and the user behavior pattern. Moreover, our investigations reveal the optimal cloud clone location is either at the nearest or the furthest node to the user along the content delivery path for a single user scenario.
Yichao Jin 0002, Yonggang Wen 0001, Han Hu 0003, Marie-José Montpetit
IEEE Trans. Multim.3
2013 Minimizing monetary cost via cloud clone migration in multi-screen cloud social TV system
abstract
The emergence of multi-screen cloud social TV has the potential to transform TV experience, providing a unified media experience across a diverse set of devices at an affordable cost. One key technology to support unified media experience across multiple screens is to instantiate a virtual machine (VM) as a cloud clone of the user, to manage all his/her media outlets (e.g., TV and smartphone), as implemented in our Cloud-Centric Media Network (CCMN). In this case, as the user shifts his attention from one device to another, the cloud clone can migrate to another location for better quality of experience. In this paper, we investigates the problem of cloud-clone migration for the multi-screen social TV application, minimizing its monetary cost. This problem can be cast into the Markov Decision Process (MDP) framework, to balance a trade-off between the migration cost and the transmission cost. Under this framework, we first derive an upper and lower bound for the optimal monetary cost, by considering a fixed placement policy and an offline policy. We then follow up with an online policy using a dynamic programming approach. Our numerical results indicate, up to 10% monetary cost can be saved, by optimally migrating the cloud clone. Moreover, the cost reduction depends on the length of content-delivery path, the data size associated with VM migration, and the user behavior pattern. These insights would offer operational guidelines to deliver cost effective multi-screen social TV services over CCMN, potentially easing its adoption.
Yichao Jin 0002, Yonggang Wen 0001, Han Hu 0003
GLOBECOM3
2013 Weighted non-linear criterion-based adaptive generalised eigendecomposition
abstract
Generalised eigendecomposition problem for a symmetric matrix pencil is reinterpreted as an unconstrained minimisation problem with a weighted non‐linear criterion. The analytical results show that the proposed criterion has a unique global minimum which corresponds to the principal generalised eigenvectors, thus guaranteeing the global convergence via iterative methods to search the minimum. A gradient‐based adaptive algorithm and a fixed point iteration‐based adaptive algorithm are derived for the generalised eigendecomposition, which both work in parallel and avoid the error propagation effect of sequential‐type algorithms. By applying the stochastic approximation theory, the global convergence of the proposed adaptive algorithm is proved. The performance of the proposed method is evaluated by simulations in terms of convergence rate, estimation accuracy as well as tracking capability.
Jian Yang 0014, Han Hu 0003, Hongsheng Xi
IET Signal Process.2
2012 Dynamic Cluster Reconfiguration for Energy Conservation in Computation Intensive Service
abstract
This paper considers the problem of dynamic cluster reconfiguration for computation intensive services. In order to provide a quality-of-service in terms of overload probability, we formulate the problem of energy consumption as a constrained optimization problem, i.e., minimizing the number of active servers to reduce the energy consumption while keeping the overload probability below a desired threshold. An overload probability estimation model is derived by applying large deviation principle, and an online measurement based algorithm is developed to decide the number of servers to power on/off, which makes decision based on current workload without any prior knowledge of the workload statistics. Moreover, the proposed dynamic cluster reconfiguration algorithm iteratively adjusts the number of the active servers, instead of directly determining the number of active servers that is hard to guarantee optimality for the nonstationary workloads. Since the distribution of the workloads among the servers has an impact on potential active servers to turn off, a server scheduling strategy is proposed to collaborate with the proposed decision algorithm to achieve better energy conservation. In order to provide an integrated solution, we present an event model-based implementation to demonstrate the practical application of the proposed approach. Finally, we evaluate the performance of the scheme by using real workloads. The experimental results show the adaptability of the proposed approach to the variations in the workload and robustness of quality-of-service for nonstationary workloads.
Jian Yang 0014, Han Hu 0003, Hongsheng Xi
IEEE Trans. Computers3
2011 Online Buffer Fullness Estimation Aided Adaptive Media Playout for Video Streaming
abstract
Adaptive media playout (AMP) control is proposed in order to compensate for the bit-rate fluctuation of networks, which may result in playout interruptions in video streaming application. Most AMP algorithms found in the literature trigger playout-rate adjustments based on the buffer fullness or its variation. However, the challenge of the threshold based methods is to select the appropriate threshold for triggering a playout-rate adjustment owing to the unknown fluctuation of the channel quality and the video bitrate. We conceive an adaptive media playout regime based on underflow probability estimation, which requires no significant statistical knowledge of the previous tele-traffic load. To achieve this, we present an underflow probability estimation model based on large deviation theory relying on the buffer fullness and on its variation. We will then directly use the underflow probability to trigger the actions of playout control, instead of using indirect methods based on a buffer fullness threshold or buffer fullness variation threshold. Experiments based on MPEG-4 Variable Bit-Rate encoded video and VBR channels associated with Adaptive Modulation and Coding are conducted in order to investigate the achievable performance of the proposed algorithm. Our simulation results demonstrate an improved performance in comparison to other recent AMP algorithms.
Jian Yang 0014, Han Hu 0003, Hongsheng Xi, Lajos Hanzo
IEEE Trans. Multim.2